Skip to main content

Evals for Fin [beta]

Test Fin's behavior before it reaches your customers

Written by Beth-Ann Sher

Evals is currently in closed beta. If you're interested in early access, please fill out this beta request form. To use Evals via Operator, you also need access to Operator.


What is Fin Evals?

Use Fin Evals to build realistic test conversations, run them against your Fin configuration, and get automatic pass/fail results — before any change reaches a real customer. This article explains how to create Evals and Simulations, run them manually or via Operator, read results with Scorecards, and build regression suites that catch problems over time.

Instead of guessing how Fin will handle a tricky refund request, an angry customer, or a change you just made to a Procedure, you can build realistic test conversations, run them against Fin, and get an automatic pass/fail result. Run the same set of tests again any time you make a change, and you'll know in minutes whether Fin still behaves the way you expect.

Fin Evals is built around two simple concepts:

  • Evals are a container (a themed group of tests). Think "Refund requests," "Escalation scenarios," or "Tone and friendliness."

  • Simulations are the individual tests inside an Eval. Each Simulation is a realistic, multi-turn conversation between a simulated customer and Fin, along with the criteria you want Fin's behavior judged against.

When you run an Eval, every Simulation inside it runs automatically, and you get a pass/fail result for each one, scored by a combination of deterministic checks and an AI judge, based on the criteria you set.

Overview of the Fin Evals interface showing a list of Evals with pass/fail results for each Simulation


Why, when, and how you should use Evals

Why use Evals?

Fin's behavior isn't fixed, it depends on your content, your Procedures, your guidance, and the data it has access to at the time it responds. Every time you make a change to any of these, there's a chance it shifts how Fin behaves somewhere else, in ways that are easy to miss until a real customer hits the issue.

Evals give you a repeatable way to check Fin's behavior across realistic, multi-turn conversations, and catch problems before they reach a live conversation.

When should you use Evals?

Reach for an Eval whenever you want confidence in how Fin behaves:

  • Before shipping a change: to content, guidance, Procedures, or anything else in Fin's configuration.

  • After a fix: to confirm you actually solved the problem you set out to solve.

  • On an ongoing basis: re-run existing Evals periodically as a regression suite, to catch unexpected drift before customers do.

  • Before going live with Fin: to check your configuration is working in the way you expect before your customers interact with Fin for the first time.

How to organize your Evals

Group Simulations into an Eval by whatever theme makes sense for your business. Common starting points beta customers have used:

  • Escalation scenarios: situations where Fin should (or shouldn't) hand off to a human.

  • Topic scenarios: a specific topic like refund requests, tested across a range of variations.

  • Behavioral scenarios: checking that Fin stays on-tone and on-brand across different situations, regardless of topic.


Creating and analyzing Evals with Operator

If you have access to Operator, it can create your Evals and Simulations for you, and analyze the results once they've run.

How to create an Eval with Operator

Go to Operator and ask it to create an Eval for your use case.

Operator chat interface showing a prompt asking Operator to create an Eval for refund request scenarios

You can guide it on what the Simulations, Simulation checks, and handoff checks should be, or leave it up to Operator to infer them. Once it's created your Eval and Simulations, they'll appear for you to review and approve.

Operator displaying a proposed Eval with Simulations and handoff checks ready to review and approve

How to run an Eval with Operator

When you’re ready, ask Operator to run the Eval for you.

Operator running an Eval, showing each Simulation being processed in real time with a progress indicator

It will show you the results as each Simulation is run.

Operator results table showing Simulations with Passed, Failed, and Error statuses and the checks behind each result

Once every Simulation has run, Operator lays out the results and explains them for you:

  • A results table — each Simulation with its overall result (Passed, Failed, or Error) and the individual checks behind it, so you can see why it landed that way. In the example above, "Bulk workspace export" failed because the handoff check failed (Fin escalated when it shouldn't have), even though the Fin reply check passed.

  • A summary — the headline counts (3 passed, 1 failed, 2 errored) followed by a plain-language breakdown of what went wrong and why. Here, Operator traces the failed export back to a guidance rule that routes to a human whenever the word "failure" appears.

Operator doesn't just report the results — it proposes the next actions and offers to take them for you.

You don't need Operator to run an Eval. Head to Fin AI Agent > Test > Evals and run it yourself.

Fin Evals tab showing an Eval with Simulations ready to run, with a Run button in the top right


Step by step: creating and running an Eval

1. Create your Eval

Go to Fin AI Agent > Test > Evals and click New eval. Give it a clear, descriptive name (up to 250 characters). You and your colleagues will be re-running this Eval later, so it should be obvious what it covers at a glance (e.g. "Refund requests - happy path and edge cases").

You can also optionally attach a Scorecard at this stage to measure the quality of Fin's responses against criteria you define, on top of pass/fail testing. See Scoring quality with Scorecards.

New eval creation screen with a name field, optional Scorecard selector, and a description field

2. Add Simulations to your Eval

Inside your Eval, select New Simulation. For each Simulation, you'll fill in:

If you don't add a title, the Simulation is automatically named from the first customer message in the conversation.

New Simulation editor showing fields for the simulated customer message, follow-up instructions, and context

  • What the simulated customer says: the conversation the simulated customer has with Fin. Add follow up instructions if you want the conversation to be multi-turn, so there is enough context to generate the customer side of the conversation.

    Simulation editor with the customer message field filled in with a refund request scenario and follow-up instructions

  • Any context Fin should have: for example, customer attributes or the state of a data connector, so the test reflects a realistic real-world scenario.

    Simulation editor showing the context field where customer attributes or data connector state can be added

  • How Fin's behavior should be evaluated:

    • Simulation checks — what Fin should say or do to pass.

      Add one or more of:

      • Fin reply — Fin's answer meets a condition you describe.

      • Procedure triggered — Fin started a specific Procedure.

      • Procedure switched — Fin moved from one Procedure to another mid-conversation.

      • Data connector — Fin called a specific data connector.

      Simulation editor showing Simulation checks (Fin reply, Procedure triggered) and Handoff check options
    • Handoff checks — what should be true about handoff by the end of the conversation. Choose one:

      • No handoff — Fin resolved it without handing off to a team or workflow.

      • Handed off to team or teammate — Fin couldn't answer, or your handoff guidance applied.

      • Handed off to a workflow — Fin passed the conversation to a workflow. (Note: the conversation isn't simulated after the handoff.)

Note:

  • Each Eval currently supports up to 50 Simulations. You won't be able to add more until you delete existing ones.

  • You can also create a Simulation from a real inbox conversation — see the section creating a Simulation from a real conversation below.

  • Coming soon:

    • Bulk import from a CSV, so you don't have to create every Simulation manually, one at a time.

3. Run the Eval

Once your Simulations are in place, run the Eval. Fin works through every Simulation in the group, and you'll get back:

  • A pass/fail result for each Simulation, checked against the Simulation checks and Handoff checks you defined.

  • The full conversation transcript for every Simulation run.

  • The event log showing FIn’s thinking and the tools and information it used at every turn.

  • The Handoff check result for each one: whether Fin handled the query or handed off to your team or a workflow. Alongside this, your Simulation checks confirm the specifics you set, such as whether a Procedure was triggered or a data connector was called.

If a single conversation needs Fin to work across multiple Procedures, Evals handle that too. You'll be able to see Fin switch between Procedures as the simulated conversation unfolds, just as it would for a real customer.

If a Simulation shows Error instead of Passed or Failed, the conversation couldn't be completed — check the Simulation's setup (customer message, context, or checks) for misconfiguration. If it Failed, open the transcript and event log to see where Fin's behavior diverged from the checks you set.

4. Re-run on demand

After making any change to Fin's content, Procedures, or guidance, re-run your existing Evals to check for regressions. Because every Simulation is stored and reusable, this takes seconds rather than requiring you to recreate your test scenarios from scratch, making it easy to build genuine regression suites over time.


Creating a Simulation from a real conversation

Some of the most valuable tests come from conversations that have already happened. When you spot a real conversation where Fin's answer wasn't good enough, you can turn that exact moment into a replay — and keep re-running it as you improve Fin, until the answer is right.

A replay captures a snapshot of the conversation up to the Fin answer you selected, and freezes the earlier turns in place. Fin is then re-run against that frozen starting point so you can see how it answers now. Change Fin's configuration, run it again, and compare — because everything before that answer stays fixed, you're testing the one response you care about, not a moving target.

How to create a replay

  1. In the inbox, open the conversation and find the Fin answer you want to test.

  2. Open the overflow menu (...) on that Fin answer and select Add Fin answer to Evaluation. (This action only appears on Fin answers that can be replayed.)

    Fin inbox conversation with the overflow menu open on a Fin answer, showing the Add Fin answer to Evaluation option

  3. Choose an existing Eval to add it to, or create a new one by giving it a name.

    Modal showing a list of existing Evals to add the replay to, with an option to create a new Eval by name
  4. The replay is saved as a Simulation inside that Eval, titled with the customer's opening message. Run it right away, or run it later as part of the whole Eval. To run it later, go to Fin AI Agent > Test > Evals, open the Eval, and click Run.

    Eval detail page showing the replay saved as a Simulation titled with the customer's opening message

How to use a replay to fix and re-test Fin's answers

Once you've saved a replay Simulation from a real conversation, use this workflow to diagnose and fix the issue:

  • Run it to see how Fin answers the frozen conversation with your current configuration.

  • Make a change — update content, a Procedure, or guidance — then re-run to see whether the answer improved.

  • Add Simulation checks, Handoff checks, or a Scorecard to define what 'fixed' looks like, so you get a clear pass/fail rather than a judgment call.

  • Keep the Simulation once it passes. It now doubles as a regression test: re-run it after future changes to make sure the fix holds.

This is the fastest way to turn a real-world miss into a permanent test — instead of writing a scenario from scratch, you start from something that genuinely happened.

Note: A replay re-runs the single Fin answer you selected, with the earlier turns of the conversation frozen in place. If you want to simulate what happens next in the conversation, add 'follow-up instructions' in the Simulation editor — these tell the system what the customer would say next so Fin has turns to respond to.


Scoring quality with Scorecards

Handoff checks and Simulation checks tell you whether Fin did the right thing. A Scorecard tells you how well it did it.

By default, every Simulation is scored on its Handoff checks (did Fin answer, or hand off to your team or a workflow) and any Simulation checks you set (for example, a specific Procedure triggered or a data connector called). That's a pass/fail read on Fin's behavior.

A Scorecard adds a qualitative scoring layer on top of your Eval’s pass/fail checks — measuring dimensions like tone, brand safety, or efficiency that don’t map cleanly to a binary outcome. Scorecards are shared with Monitors, so you’re testing against the same quality rubric before a change goes live that your monitors enforce after it does.

What’s in a Scorecard

A Scorecard is made up of one or more criteria — the dimensions you care about (e.g. "Efficiency", "Clarification", "Escalation ease"). Each criterion has:

  • A name: a short label that appears in your results.

  • A description: what you're evaluating and how it should be judged. This is the instruction the AI judge follows — use one of the pre-existing options, or be specific if creating your own.

  • Rating options: the possible scores (at least two), each with a name (e.g. "Good", "Okay", "Poor") and a numerical value (e.g. 100%, 50%, 0%).

In an Eval, criteria are scored automatically by an AI judge, alongside the rest of the Simulation.

How to configure Scorecard scoring

When attaching a Scorecard to an Eval, you can configure how each criterion contributes to the overall score:

  • Weighting: give each criterion a weight to reflect how much it matters. Weights are proportional, so a criterion with weight 2 counts twice as much as one with weight 1.

  • Critical criteria: mark a criterion as Critical when it's non-negotiable (e.g. for compliance or safety). A failing rating on a Critical criterion fails the whole qualitative review, whatever the other scores are.

  • Pass threshold: set the minimum overall score a Simulation needs to pass on quality.

How to read Scorecard results

When an Eval runs, each Simulation shows its Handoff check and Simulation check results as before, plus its Scorecard score and the rating and reason the judge gave for each criterion. Because the judge shows its reasoning, you can see why a response scored the way it did, and turn that straight into a fix.

Simulation result showing Scorecard scores per criterion with the AI judge's rating and reasoning for each

Tip: Vague criteria produce vague scores. Write each criterion the way you'd brief a new reviewer — spell out what 'good' looks like and what should tank the score. See how to write effective Monitor & Scorecard criteria.


How Evals complement Procedure Simulations

Procedure Simulations aren't going anywhere, and they're still the right tool for a specific job: testing one Procedure in isolation. They require you to define success and outcome criteria for every test, and they're ideal for validating that an individual Procedure behaves correctly on its own.

Evals pick up where Procedure Simulations leave off. Where a Procedure Simulation is scoped to a single Procedure, an Eval Simulation can exercise Fin's entire configuration end to end: moving across multiple Procedures, content, guidance, and data connectors within a single realistic conversation, the same way a real customer conversation would.

Use them together:

  • Use Procedure Simulations to build and validate an individual Procedure while you're working on it.

  • Use Evals to validate the full customer journey once that Procedure is live, including how it interacts with everything else in Fin's configuration.


Using Evals instead of Batch Test

Today, many teams use Batch Test to validate Fin which involves asking a batch of single-turn, informational questions and manually rating each response as Good, Acceptable, or Poor.

Fin Evals is designed to do everything Batch Test does, and considerably more:

The table below compares Batch Test and Fin Evals across four dimensions: conversation type, scoring, regression testing, and grouping.

Batch Test

Fin Evals

Conversation type

Single-turn, informational questions only

Full multi-turn conversations

Scoring

Manual rating (Good/Acceptable/Poor)

Automatic — deterministic checks plus AI-as-judge against criteria you set

Regression testing

Re-run manually

Re-run on demand as a regression suite

Grouping

Grouped, up to 50 questions

Grouped into Evals of up to 50 Simulations, organized however suits you

If you're currently using Batch Test, we'd encourage you to start building out equivalent test scenarios as Evals during the beta. Evals gives you a faster, more automated, and more realistic way to get the same confidence and then some.


Simulation usage limits

There is a limit to the number of Simulations you can run within Evals each month. This limit is applied at the workspace level and resets on the first day of every calendar month.

Each workspace receives a monthly allowance of simulation runs. The allowance is based on your workspace’s conversation-volume segment, with larger customers receiving higher allowances.

Simulation allowance is based on your workspace’s conversation volume in Intercom.

  • We assign your workspace to a segment using the number of conversations in the last calendar month.

  • Your segment is re-evaluated monthly and your allowance will reflect your most recent month’s conversation volume.

  • If your conversation volume increases or decreases, your allowance may change in the next monthly cycle.

The table below shows the monthly Simulation run allowance by workspace conversation-volume segment.

Conversation Volume Segment

Simulation Limit per month

Under 1K

250

1K–15K

1,000

15K–100K

1,750

100K–1M

5,000

1M+

12,500

Monitoring your usage

To help you manage your testing, you'll see visual indicators within the Evaluations tab:

Usage warning

When your workspace reaches 80% of its monthly limit, a yellow warning banner will appear. It displays your current usage (e.g., "850/1000") and reminds you when the limit will reset.

Evals tab showing a yellow warning banner at 850/1000 Simulations used with the monthly reset date displayed

Limit reached

Once you hit 100% of your monthly limit, a red error message will appear. You will be unable to run further Simulations within Evals until the start of the next month.

Evals tab showing a red error banner indicating the monthly Simulation limit has been reached and no more runs are possible


Key things to know about the beta

Fin Evals is in Closed Beta, and we're actively building it out. Here's what to keep in mind right now:

  • Organizing beyond an Eval itself isn't available yet. There's currently no folder structure for grouping Evals by team or ownership.

  • In-product annotation is coming. You will soon be able to add simple notes to a Simulation, but there's no dedicated way yet to record whether a reviewer agrees or disagrees with a result.

  • There aren't pre-built templates yet. Common test categories like prompt injection or generic edge-case questions don't have a ready-made starting point — you'll be building these from scratch for now.

  • Workflows are out of scope for now. Evals currently cover Fin's answer configuration — content, Procedures, guidance, audiences, and data connectors — but not Workflows.

  • Up to 50 Simulations per Eval, and up to 50 rows per CSV import.


Understanding pass/fail: the default "no handoff" behavior

If you add a Simulation with no explicit Handoff check and no Simulation checks, Evals defaults to evaluating it against a "No handoff" criterion — meaning the Simulation passes only if Fin resolved the conversation without handing off to a team or workflow.

This default isn't shown in the Simulation setup UI, which means Simulations can fail in unexpected ways if Fin escalated and you didn't intend to test for that. If a Simulation fails and you can't see why, check whether a Handoff check was set — if none was, the default no-handoff rule is what failed it.

Tip: Always add an explicit Handoff check to each Simulation so the pass/fail result reflects your intent, not the default.

Have feedback on Fin Evals? We'd love to hear it so please reach out to your account manager.


💡Tip

Need more help? Get support from our Community Forum
Find answers and get help from Intercom Support and Community Experts


Did this answer your question?