Purdue University / Findings of EMNLP 2026

BIRDS

Characterizing and Understanding Biodiversity Impact of Large Language Model Serving

AI runs on infrastructure.
Its footprint reaches into ecosystems.

Tianyao Shi & Yi Ding

An ecological lens on computing

01 / Beyond carbon and water

What does an AI request
mean for the living world?

The infrastructure behind a response has a life beyond the screen. Electricity generation, chip manufacturing, transport, and disposal all contribute to environmental pressures.

Large language models (LLMs) are artificial intelligence (AI) systems that generate language. Serving means running a model to answer requests. BIRDS connects those activities to modeled ecosystem damage, then asks how much impact is associated with a completed request and the quality of its answer.

A completed request

A defined task and model, running on a specified hardware setup.

A physical footprint

Electricity use and the lifecycle of serving hardware.

An ecological impact

Climate, water, pollution, and other ecosystem pathways.

02 / Explore the evidence

Different choices. Different footprints.

Selected results from the final accepted paper. Biodiversity is the variety of living organisms and the ecosystems they form. Biodiversity impact (BI) is the modeled damage to ecosystems associated with activities such as electricity use and hardware production. Each comparison keeps the task and measurement assumptions in view.

BI per request
Modeled biodiversity impact of one completed request. Its unit, species·year, expresses potential species loss integrated over time. It is an assessment indicator, not a count of observed local extinctions.

Quality-Normalized Biodiversity Impact (QNBI)
BI per request divided by the model's task-specific quality score. Quality measures how well the answer performs on that task. Lower means less impact per unit of achieved quality.

FIGURE 4

Where does the impact come from?

Graphics processing units (GPUs) are the chips that run the models; H100 and A100 are GPU types. Operation means using electricity to serve requests, and it dominates total BI on average in these configurations. An assessment horizon is the period over which environmental effects are counted. The contributions of pathways such as climate change and water use change with that horizon.

Interpretation

Mean contribution shares across the evaluated serving configurations. This view summarizes Figure 4; its lifecycle panel uses means rather than reproducing the paper's boxplots. Lifecycle shares use the 100-year horizon. An assessment horizon is a modeling choice, not a forecast date.

FIGURES 5 & 6

The smallest model is not always the best tradeoff.

Lower impact per request does not guarantee lower quality-normalized impact. Compare models within one task to see how accounting for quality changes the picture.

A workload is a type of task, such as chat or summarization. QA means question answering; retrieval-augmented generation (RAG) supplies reference context for the answer.

In model names, B denotes billions of parameters, the learned values that determine model behavior. For example, 26B-A4B indicates total and active parameter counts. Dense models use all their parameters for each generated token (a piece of text); mixture-of-experts (MoE) models activate a subset of experts. Chat denotes conversation tuning; Instruct and it denote instruction tuning; Thinking denotes extended reasoning before answering. A logarithmic scale shows equal ratios as equal distances; a linear scale shows equal differences as equal distances.

QNBI is comparable within a workload and quality metric. Quality scores across different tasks are not interchangeable. Only models with available quality scores are included in this view.

FIGURE 8

More reasoning comes with a tradeoff.

MMLU-Pro is a multiple-choice knowledge and reasoning benchmark. Instruct variants are tuned to follow instructions; Thinking variants generate extended reasoning before answering. Thinking variants improve MMLU-Pro accuracy, but their additional serving impact can outweigh the quality gain. The appropriate choice depends on the task.

Interpretation

Selected Qwen3 configurations on MMLU-Pro. Quality is normalized benchmark accuracy. This view shows the BI, quality, and QNBI comparisons; response-length distributions remain in the paper.

FIGURE 10

Hardware efficiency changes the outcome.

GPU generations differ in the useful work they sustain. The selected configurations show how hardware choice changes the impact of serving the same model and workload. This comparison uses ShareGPT, a dataset of user-chatbot conversations.

Interpretation

ShareGPT is a dataset of user-chatbot conversations. This view uses the most energy-efficient evaluated configuration for each model and GPU. The paper also reports throughput (requests completed per unit time) and tensor parallelism (splitting a model's computations across GPUs).

Measurement scope and figure provenance

Results correspond to the EMNLP camera-ready paper (arXiv v3). Model comparisons primarily use H100 GPUs, with A100 results for some smaller models; model tooltips identify the GPU used. Operational modeling uses US grid conditions; reported impact includes hardware lifecycle contributions.

Energy accounting uses the additional GPU energy consumed while serving requests, above the energy used when idle. The modeling also accounts for datacenter overhead such as cooling. The figures show processed biodiversity-impact results. The paper provides the complete measurement and modeling methodology. Model names follow the appendix overview, including Chat, Instruct, Thinking, and it (instruction-tuned) variants. Chat variants are tuned for conversation. Qwen3 Instruct and Thinking variants use the 2507 release.

Species·year is a lifecycle assessment indicator of ecosystem damage integrated over time, not a literal count of locally observed extinctions. QNBI depends on the task-specific quality evaluation. These results do not cover every effect of data centers, model training, or agentic systems that perform multistep tasks and use external tools.

Read the complete methodology

03 / Follow the sources

The evidence behind the estimates.

The sources in Appendix B.3, organized by their role in the lifecycle model. Links lead to the original providers or publications; source datasets are not redistributed here.

04 / Build on this work

Research is a shared effort.

Tianyao Shi and Yi Ding
Purdue University

Research & media inquiries

Cite BIRDS

@inproceedings{shi2026birds,
  title = {{BIRDS}: Characterizing and Understanding
    Biodiversity Impact of Large Language Model Serving},
  author = {Shi, Tianyao and Ding, Yi},
  booktitle = {Findings of the Association for
    Computational Linguistics: EMNLP 2026},
  year = {2026},
  eprint = {2605.27480},
  archivePrefix = {arXiv}
}