Today we’re announcing Evals and Releases. Paired with Monitors, they give you an evaluation system for Fin, so you can test changes before they go live, roll them out with control, evaluate every live conversation, and have confidence in the experience Fin delivers.
Evals tests Fin’s behavior at scale using simulated scenarios built from your own conversations. You group these simulated conversations into an Eval around a theme – such as refund requests, escalation rules, or Fin’s tone of voice – and each one is scored automatically against criteria you set. Releases gives your team a dedicated space to build changes away from the live version of Fin, test them on real conversations, and roll them out safely.
Once live, Monitors keeps the process continuous: assessing the quality of every conversation, and catching issues you can work through in your next Eval and Release.
We call this “Eval-driven delivery.” It’s the discipline our AI research team uses to build and tune AI products, and it’s now in the hands of the teams running support with Fin.

Here’s what’s new:
- Evals tests Fin end to end. Create an Eval for a topic, add Simulations that mirror real customer conversations, and each one is scored pass/fail against your success and qualitative criteria. Re-run the Eval after any change to Fin to catch unexpected regressions in performance.
- Releases gives you a safe version of Fin to work in. Bundle changes to content, Procedures, or Guidance into one Release, collaborate on it with your team, and test performance with Evals before going live. Then publish to everyone, ramp traffic gradually, or run an A/B test against Fin’s current configuration.
- Evals, Releases, and Monitors work as one system. Monitors checks every live conversation against your standards and flags the ones that fall short. Turn any flagged conversation into your next set of improvements and run the system again.
Why we built this
As AI Agents take on more volume and complex work, ensuring they meet your standards every time is a tough task.
Fin, like any AI Agent, is probabilistic. It reasons through each conversation as it happens, so it can answer the same question in different ways – and your customers will ask the same question in countless ways too. That creates thousands of possible scenarios that no team can validate by hand.
And success isn’t just about whether the final answer was right. It depends on whether Fin followed your Procedures, used the right content and data sources, handed off at the right moment, and sounded like your brand while doing it.
Teams also continuously make changes to Fin’s content, updating its context when they launch a product, adjust a policy, rewrite a help article, or add a Procedure. That can introduce drift, where changing one thing can impact something else without you noticing. At the scale some of our customers run Fin, a 1% regression could affect thousands of conversations a day.
Current oversight and testing solutions weren’t built for this scale and complexity. Meaning, it can be really hard to always know how Fin will handle a conversation.
We think every support team should have the ability to rigorously test changes and monitor performance without needing a developer or ML specialist.
Evals: Know how Fin will behave before you go live
An Eval is a named group of Simulations – multi-turn test conversations – that validates Fin’s behavior on one theme, topic, or issue. You might build one for your most common queries, one for how Fin handles refund requests inside and outside your policy, and another for tone of voice across a range of situations.

Build out Simulations from real conversations
Simulations are made up of three actors:
- The simulated customer – their opening message, and the context they reveal as the conversation goes on, like account details, order information, or a change in mood.
- What Fin has access to – including attributes and data connectors.
- The criteria the LLM judge scores against – whether Fin replied and what it needed to say, whether the right Procedure triggered, or whether a data connector was called.
You can create Simulations manually, upload them from existing conversations, or generate them based on real inbox conversations.

Score every Simulation the way your best reviewers would
Running an Eval runs every Simulation inside it and returns a pass or fail for each one. Scoring combines deterministic checks with an LLM judge, and if any single criterion fails, the Simulation fails.
Every run shows the full conversation transcript, the event log, and the outcome – answered by Fin, handed off to your team, or handed off to a workflow – so you can see exactly how Fin got there. Where a scenario needs more than one Procedure, the Simulation shows Fin switching between them mid-conversation.


Re-run an Eval after any change to catch regressions
Once an Eval exists, it becomes a way to test for regressions on an ongoing basis. Re-run it before or after committing a change to Fin so you can ensure a fix to one thing doesn’t break another.
“We group our Evals around different topics, so every time we make a change, we rerun the whole set and see immediately whether anything regressed.”
— Hila Horenshtein, CX AI Operations Team Lead, AutoDS
Releases: Control how a change reaches customers
Releases gives you a dedicated space to plan, collaborate, and test changes to Fin so nothing reaches customers until you decide it’s ready.
Work on changes without touching live Fin
In a Release, you can edit content, add or update Procedures, and change or delete Guidance, all bundled together. This simplifies big changes like new product launches, or optimizations to a specific workflow – like moving Fin from explaining your refund policy to processing refunds with a Procedure.
Your teammates can work in the same Release – reviewing edits, adding their own, or testing what’s there. Every change is listed in one place, so you can see exactly what’s in the Release and who made each edit.
You can also run an Eval against any changes within the Release before you publish it. Build the change, run your Eval, see what fails, make adjustments, and run it again – without touching live Fin.

A/B test a change before it goes wide
When a Release is ready you can publish it to everyone or run an A/B test against Fin’s current configuration. Experiment reporting uses metrics you care about, like resolution rate, escalation rate and CSAT, so you can be confident in every new release.

Roll back instantly
If something looks wrong partway through a rollout, pause it. Fix the specific issue, re-run the Eval to confirm the fix worked, then pick the rollout back up. If you need to undo a Release entirely, roll it back in one step.

How Evals, Releases, and Monitors work together
Each of these is useful alone. Together, they strengthen the Fin Flywheel, so only the best version of Fin reaches your customers.
You create a change in a Release and validate it with Evals before going live. You roll it out as an experiment, then publish to everyone once the results hold up. From there, a Monitor keeps watch – checking live conversations against your quality standards so the change keeps performing after it’s fully live. When a Monitor flags a conversation that fell short, you turn it into a Simulation, add it to the relevant Eval, and use it as a benchmark to test a new Release and drive better performance. That’s how you can have confidence in the experience Fin delivers at scale.

Operator runs Eval-driven delivery for you
You don’t have to work through Evals and Releases step by step yourself. Operator, our Agent for customer operations, can use all of these capabilities on your behalf.
For example, you can ask Operator to:
- Update your refund policy to say that customer refunds can be applied as a credit on their account.
- Then, add a test for this to the ‘refund policy’ Eval.
- And finally, put this change into a new Release.
And it will handle all three. It drafts the content change, builds the Eval, and bundles the work into a Release. Every step comes back as a proposal, so you can review it and give final approval.

Maintaining customer experience as your business evolves
You can build a very good AI Agent on day one. Maintaining this performance while your products, policies, and content keep changing is a different problem.
Evals and Releases give your team the same technical rigor an engineering team would build for itself – without needing an engineering team to run it. You set the standards, build the tests, and control the rollout. When integrated with Monitors, you get the oversight you need to maintain your standards as you grow.
Learn more here.