What does an AI request mean for the living world?
The infrastructure behind a response has a life beyond the screen.
Electricity generation, chip manufacturing, transport, and
disposal all contribute to environmental pressures.
Large language models (LLMs) are artificial intelligence (AI)
systems that generate language. Serving means running a model to
answer requests. BIRDS connects those activities to modeled ecosystem damage, then
asks how much impact is associated with a completed request and
the quality of its answer.
A completed request
A defined task and model, running on a specified hardware setup.
A physical footprint
Electricity use and the lifecycle of serving hardware.
An ecological impact
Climate, water, pollution, and other ecosystem pathways.
02 / Explore the evidence
Different choices. Different footprints.
Selected results from the final accepted paper. Biodiversity is the
variety of living organisms and the ecosystems they form.
Biodiversity impact (BI) is the modeled damage to
ecosystems associated with activities such as electricity use and
hardware production. Each comparison keeps the task and measurement
assumptions in view.
BI per request Modeled biodiversity impact
of one completed request. Its unit, species·year,
expresses potential species loss integrated over time. It is an
assessment indicator, not a count of observed local extinctions.
Quality-Normalized Biodiversity Impact (QNBI)
BI per request divided by the model's task-specific quality score.
Quality measures how well the answer performs on that task.
Lower means less impact per unit of achieved quality.
FIGURE 4
Where does the impact come from?
Graphics processing units (GPUs) are the chips that run the
models; H100 and A100 are GPU types. Operation means using
electricity to serve requests, and it dominates total BI on
average in these configurations. An assessment horizon is the
period over which environmental effects are counted. The
contributions of pathways such as climate change and water use
change with that horizon.
Interpretation
Mean contribution shares across the evaluated serving
configurations. This view summarizes Figure 4; its lifecycle
panel uses means rather than reproducing the paper's boxplots.
Lifecycle shares use the 100-year horizon. An assessment horizon
is a modeling choice, not a forecast date.
FIGURES 5 & 6
The smallest model is not always the best tradeoff.
Lower impact per request does not guarantee lower
quality-normalized impact. Compare models within one task to see
how accounting for quality changes the picture.
A workload is a type of task, such as chat or summarization.
QA means question answering; retrieval-augmented generation
(RAG) supplies reference context for the answer.
In model names, B denotes billions of parameters, the learned
values that determine model behavior. For example, 26B-A4B
indicates total and active parameter counts. Dense models use
all their parameters for each generated token (a piece of text);
mixture-of-experts (MoE) models activate a subset of experts.
Chat denotes conversation tuning; Instruct and it denote
instruction tuning; Thinking denotes extended reasoning before
answering.
A logarithmic scale shows equal ratios as equal distances;
a linear scale shows equal differences as equal distances.
QNBI is comparable within a workload and quality metric. Quality
scores across different tasks are not interchangeable.
Only models with available quality scores are included in this view.
FIGURE 8
More reasoning comes with a tradeoff.
MMLU-Pro is a multiple-choice knowledge and reasoning benchmark.
Instruct variants are tuned to follow instructions; Thinking
variants generate extended reasoning before answering.
Thinking variants improve MMLU-Pro accuracy, but their
additional serving impact can outweigh the quality gain. The
appropriate choice depends on the task.
Interpretation
Selected Qwen3 configurations on MMLU-Pro. Quality is normalized
benchmark accuracy. This view shows the BI, quality, and QNBI
comparisons; response-length distributions remain in the paper.
FIGURE 10
Hardware efficiency changes the outcome.
GPU generations differ in the useful work they sustain. The
selected configurations show how hardware choice changes the
impact of serving the same model and workload.
This comparison uses ShareGPT, a dataset of user-chatbot conversations.
Interpretation
ShareGPT is a dataset of user-chatbot conversations. This view
uses the most energy-efficient evaluated configuration for each
model and GPU. The paper also reports throughput (requests
completed per unit time) and tensor parallelism (splitting a
model's computations across GPUs).
Measurement scope and figure provenance
Results correspond to the EMNLP camera-ready paper (arXiv v3).
Model comparisons primarily use H100 GPUs, with A100 results for
some smaller models; model tooltips identify the GPU used.
Operational modeling uses US grid conditions; reported impact
includes hardware lifecycle contributions.
Energy accounting uses the additional GPU energy consumed while
serving requests, above the energy used when idle. The modeling
also accounts for datacenter overhead such as cooling.
The figures show processed biodiversity-impact results. The paper
provides the complete measurement and modeling methodology.
Model names follow the appendix overview, including Chat,
Instruct, Thinking, and it (instruction-tuned) variants.
Chat variants are tuned for conversation. Qwen3 Instruct and Thinking
variants use the 2507 release.
Species·year is a lifecycle assessment indicator of ecosystem
damage integrated over time, not a literal count of locally
observed extinctions. QNBI depends on the task-specific quality
evaluation. These results do not cover every effect of data
centers, model training, or agentic systems that perform multistep
tasks and use external tools.
The sources in Appendix B.3, organized by their role in the lifecycle
model. Links lead to the original providers or publications; source
datasets are not redistributed here.
@inproceedings{shi2026birds,
title = {{BIRDS}: Characterizing and Understanding
Biodiversity Impact of Large Language Model Serving},
author = {Shi, Tianyao and Ding, Yi},
booktitle = {Findings of the Association for
Computational Linguistics: EMNLP 2026},
year = {2026},
eprint = {2605.27480},
archivePrefix = {arXiv}
}