Non-LLM
Exact Match
Single-turn
The Exact Match metric measures whether your LLM application's actual_output matches the expected_output exactly.
Required Arguments
To use the ExactMatchMetric, you'll have to provide the following arguments when creating an LLMTestCase:
input-
actual_output -
expected_output
Read the How Is It Calculated section below to learn how test case parameters are used for metric calculation.
Usage
from deepeval.metrics import ExactMatchMetric
from deepeval.test_case import LLMTestCase
from deepeval import evaluate
metric = ExactMatchMetric(
threshold=1.0,
verbose_mode=True,
)
test_case = LLMTestCase(
input="Translate 'Hello, how are you?' in french",
actual_output="Bonjour, comment รงa va ?",
expected_output="Bonjour, comment allez-vous ?"
)
# To run metric as a standalone
# metric.measure(test_case)
# print(metric.score, metric.reason)
evaluate(test_cases=[test_case], metrics=[metric])There are THREE optional parameters when creating an ExactMatchMetric:
- [Optional]
threshold: a number representing the minimum passing threshold. Can also be set toNoneto run the metric in score-only mode. Defaulted to1.0. - [Optional]
verbose_mode: a boolean which when set toTrue, prints the intermediate steps used to calculate said metric to the console, as outlined in the How Is It Calculated section. Defaulted toFalse. - [Optional]
flaky: a boolean which when set toTrue, marks the metric as flaky. Defaulted toFalse.
As a Standalone
You can also run the ExactMatchMetric on a single test case as a standalone, one-off execution.
...
metric.measure(test_case)
print(metric.score, metric.reason)How Is It Calculated?
The ExactMatchMetric score is calculated according to the following equation:
The ExactMatchMetric performs a strict equality check to determine if the actual_output matches the expected_output.
FAQs
Does the Exact Match metric call an LLM or cost money?
No. The
ExactMatchMetric is a plain string equality check between actual_output and expected_output โ no model, no API key, zero token cost, fully deterministic.Is the comparison case-sensitive and whitespace-sensitive?
Yes. Any difference โ casing, whitespace, punctuation, accents โ scores
0. For example, "Bonjour, comment รงa va ?" won't match "Bonjour, comment allez-vous ?".When should I use Exact Match instead of an LLM-judge metric?
When there's exactly one acceptable answer โ classification labels, enum values, canned responses. For open-ended outputs with many valid phrasings, use a semantic metric like Answer Relevancy.
Why does my output fail even though it looks correct?
It requires a character-for-character match. Set
verbose_mode=True to print the compared strings and spot invisible differences like trailing newlines, smart quotes, or leading whitespace.