🔥 DeepEval 4.0 just got released. Read the announcement.
Multi-Turn

Topic Adherence

LLM-as-a-judge
Multi-turn
Referenceless
Agent
Multimodal

The Topic Adherence metric is a multi-turn agentic metric that evaluates whether your agent has answered questions only if they adhere to relevant topics. It is a self-explaining eval, which means it outputs a reason for its metric score.

Required Arguments

To use the TopicAdherenceMetric, you'll have to provide the following arguments when creating a ConversationalTestCase:

  • turns

You can learn more about how it is calculated here.

Usage

The TopicAdherenceMetric() can be used for end-to-end multi-turn evaluations of agents.

from deepeval.test_case import Turn, ConversationalTestCase, ToolCall
from deepeval.metrics import TopicAdherenceMetric
from deepeval import evaluate

convo_test_case = ConversationalTestCase(
    turns=[
        Turn(role="...", content="..."), 
        Turn(role="...", content="...", tools_called=[...])
    ],
)
metric = TopicAdherenceMetric(threshold=0.5)

# To run metric as a standalone
# metric.measure(convo_test_case)
# print(metric.score, metric.reason)

evaluate(test_cases=[convo_test_case], metrics=[metric])

There is ONE mandatory and SEVEN optional parameters when creating a TopicAdherenceMetric:

  • relevant_topics: a list of strings that define what topics your LLM agent can answer. Any answers that don't adhere to this topic will penalise the score this metric.
  • [Optional] threshold: a number representing the minimum passing threshold. Can also be set to None to run the metric in score-only mode. Defaulted to 0.5.
  • [Optional] model: a string specifying which of OpenAI's GPT models to use, OR any custom LLM model of type DeepEvalBaseLLM. Defaulted to gpt-5.4.
  • [Optional] include_reason: a boolean which when set to True, will include a reason for its evaluation score. Defaulted to True.
  • [Optional] strict_mode: a boolean which when set to True, enforces a binary metric score: 1 for perfection, 0 otherwise. It also overrides the current threshold and sets it to 1. Defaulted to False.
  • [Optional] async_mode: a boolean which when set to True, enables concurrent execution within the measure() method. Defaulted to True.

  • [Optional] verbose_mode: a boolean which when set to True, prints the intermediate steps used to calculate said metric to the console, as outlined in the How Is It Calculated section. Defaulted to False.
  • [Optional] flaky: a boolean which when set to True, marks the metric as flaky. Defaulted to False.

As a standalone

You can also run the TopicAdherenceMetric on a single test case as a standalone, one-off execution.

...

metric.measure(convo_test_case)
print(metric.score, metric.reason)

How Is It Calculated

The TopicAdherenceMetric score is calculated through the following process:

  • Find question-answer pairs from the entire conversation, where question is taken from user and answered by the LLM agent.
  • Find the truth table values for all the question-answer pairs.
    • True Positives: Question is relevant and the response correctly answers it.
    • True Negatives: Question is NOT relevant, and the assistant correctly refused to answer.
    • False Positives: Question is NOT relevant, but the assistant still gave an answer.
    • False Negatives: Question is relevant, but the assistant refused or gave an irrelevant response.

Now, the metric uses the following formula to find the final score:

Topic Adherence Score=Number of True Positives and True NegativesTotal Number of QA Pairs\text{Topic Adherence Score} = \frac{\text{Number of True Positives and True Negatives}}{\text{Total Number of QA Pairs}}

The TopicAdherenceMetric converts turns into individual unit interactions and iterates over each interaction to find the question-answer pairs separately, which are also evaluated individually for more accurate results.

FAQs

How do I stop my support bot from drifting off-topic?
Define allowed topics in relevant_topics and run TopicAdherenceMetric. Any turn answering a question outside those topics is penalized, pinpointing where the bot wandered off-script.
How are off-topic turns actually scored?
Each user question is paired with the agent's answer: answering an off-topic question is a false positive (bad), refusing one is a true negative (good). The score is the share of true positives and true negatives across all QA pairs.
Will my agent be penalized for refusing irrelevant questions?
No — refusing a question outside relevant_topics is a true negative that helps the score. You're only penalized for answering off-topic questions or refusing on-topic ones.
What should I put in relevant topics?
A list of strings describing the subject areas your agent covers (for example, "billing" or "shipping"). Keep them specific but not so narrow that legitimate questions get flagged.

On this page