Turn Faithfulness
The turn faithfulness metric is a conversational metric that determines whether your LLM chatbot generates factually accurate responses grounded in the retrieval context throughout a conversation.
Required Arguments
To use the TurnFaithfulnessMetric, you'll have to provide the following arguments when creating a ConversationalTestCase:
turns
You must provide the role, content, and retrieval_context for evaluation to happen. Read the How Is It Calculated section below to learn more.
Usage
First, set the eval mode:
deepeval set-eval-mode llm # LLM-as-a-judge (default)
deepeval set-eval-mode hybrid # LLM extracts, Jev decides
deepeval set-eval-mode system_one # Jev-as-a-judge, no LLMThe TurnFaithfulnessMetric() can be used for end-to-end multi-turn evaluation:
from deepeval.test_case import Turn, ConversationalTestCase
from deepeval.metrics import TurnFaithfulnessMetric
from deepeval import evaluate
convo_test_case = ConversationalTestCase(
turns=[
Turn(role="user", content="...", retrieval_context=["..."]),
Turn(role="assistant", content="...", retrieval_context=["..."])
]
)
metric = TurnFaithfulnessMetric(threshold=0.5)
# To run metric as a standalone
# metric.measure(convo_test_case)
# print(metric.score, metric.reason)
evaluate(test_cases=[convo_test_case], metrics=[metric])There are TWELVE optional parameters when creating a TurnFaithfulnessMetric:
- [Optional]
threshold: a number representing the minimum passing threshold. Can also be set toNoneto run the metric in score-only mode. Defaulted to0.5. - [Optional]
model: a string specifying which of OpenAI's GPT models to use, OR any custom LLM model of typeDeepEvalBaseLLM. Defaulted togpt-5.4. - [Optional]
include_reason: a boolean which when set toTrue, will include a reason for its evaluation score. Defaulted toTrue. - [Optional]
strict_mode: a boolean which when set toTrue, enforces a binary metric score: 1 for perfection, 0 otherwise. It also overrides the current threshold and sets it to 1. Defaulted toFalse. -
[Optional]
async_mode: a boolean which when set toTrue, enables concurrent execution within themeasure()method. Defaulted toTrue. - [Optional]
verbose_mode: a boolean which when set toTrue, prints the intermediate steps used to calculate said metric to the console, as outlined in the How Is It Calculated section. Defaulted toFalse. - [Optional]
truths_extraction_limit: an optional integer to limit the number of truths extracted from retrieval context per document. Defaulted toNone. - [Optional]
penalize_ambiguous_claims: a boolean which when set toTrue, penalizes claims that cannot be verified as true or false. Defaulted toFalse. - [Optional]
window_size: an integer which defines the size of the sliding window of turns used during evaluation. Defaulted to10. - [Optional]
flaky: a boolean which when set toTrue, marks the metric as flaky. Defaulted toFalse. - [Optional]
system_one_model: the Jev model to use, as a string or aDeepEvalBaseSystemOneModel. Only used underhybridorsystem_oneeval_mode. Defaulted tojev-latest. - [Optional]
eval_mode:llm,hybridorsystem_one, choosing whether an LLM, Jev, or both judge. Defaulted to the configured eval mode (llmunless set).
As a standalone
You can also run the TurnFaithfulnessMetric on a single test case as a standalone, one-off execution.
...
metric.measure(convo_test_case)
print(metric.score, metric.reason)How Is It Calculated?
You can change how the TurnFaithfulnessMetric is calculated by setting the eval mode.
LLM-as-a-judge
The TurnFaithfulnessMetric score is calculated according to the following equation:
The TurnFaithfulnessMetric first constructs a sliding windows of turns. For each window, it:
- Extracts truths from the retrieval context provided in the turns
- Generates claims from the assistant's responses in the interaction
- Evaluates verdicts by checking if each claim contradicts the truths
- Calculates the interaction score as the ratio of faithful claims to total claims
The final score is the average of all interaction faithfulness scores across the conversation.
Hybrid
Under the hybrid eval mode, step 3 is answered by Jev, a System One model, instead: one yes/no question per claim, with P(yes) > 0.65 counted as truthful, < 0.35 as contradictory, and anything in between as borderline. The extraction, the equation and the reason are unchanged. If a Jev call fails, the LLM makes that decision instead.
Jev-as-a-judge
Under the system_one eval mode, Jev judges the whole metric in one request. It is sent the whole conversation, each turn carrying its role, content and retrieval_context, and asked three questions:
| Question | Type | Weight |
|---|---|---|
Every factual claim made in an assistant turn in turns is supported by the retrieval_context of that turn or of an earlier assistant turn. | Noul | 2 |
No assistant turn in turns contradicts the retrieval_context available to it. | Noul | 1 |
Across turns, how much of what the assistant states is grounded in the retrieval_context available to it? (Fabricated → Fully grounded) | Score | 1 |
Each answer becomes a value in and the score is their weighted mean. No LLM is called: the reason lists each answer with its probability, and metric.confidence reports how decisive Jev was.
FAQs
When should I use TurnFaithfulnessMetric instead of the single-turn FaithfulnessMetric?
retrieval_context. The single-turn FaithfulnessMetric grades one output against one context and can't pinpoint which turn hallucinated.Does every turn need its own retrieval context?
retrieval_context to turns generated from retrieval. The metric extracts truths from that context and verifies the assistant's claims against them — turns with no grounding context are where faithfulness problems surface.My answers are factually correct but faithfulness is low — how do I debug it?
retrieval_context, not claims that merely happen to be true — correct-but-unretrieved facts count as unfaithful. Inspect per-claim verdicts with verbose_mode, and try penalize_ambiguous_claims for vague claims.How do window size and truths extraction limit affect long conversations?
window_size (default 10) groups turns per evaluation; truths_extraction_limit caps truths pulled from each document. Tune them together when long conversations make scoring slow or noisy.