Video: How to evaluate voice agents: execution, outcomes, and experience | Duration: 2724s | Summary: How to evaluate voice agents: execution, outcomes, and experience | Chapters: Introducing LangSmith (1.68s), Voice Agent Challenges (95.215s), Voice Eval Challenges (283.575s), Three Evaluation Dimensions (766.885s), LangSmith Demo (1301.96s), Evaluation Output Types (2397.645s), Voice Tone Evaluation (2501.465s), Golden Dataset Structure (2600.935s), Closing Remarks (2701.055s)
Transcript for "How to evaluate voice agents: execution, outcomes, and experience": Agents are a new kind of software. Endless inputs and non deterministic outputs. Build them the old way, and they break. Enter Lingsmith, the agent engineering It's model agnostic, cloud agnostic, and framework agnostic. Organization is shipping the best agents iterate with a system, and we call that system the agent development life cycle. Build, test, deploy, and monitor. You build with our open source frameworks. Deep agents, langchain, or langgraph. LangSmith Fleet enables anyone to build agents without code. Test agents using evals and experiments. Deploy changes in one click with LangSmith Deployment. Monitor every interaction from a single dashboard. Governance is built into every stage of the agent development cycle with Lanesmith LLM Gateway. Lanesmith Engine helps you improve agents autonomously and move through the development lifecycle more quickly. Blanksmith Engine is your agent for agent engineering. Welcome, everyone. Today, we're going to talk about evaluating voice agents. Voice agents are tricky. Evaluating them is important. And so we're going to dive into the details of how to do that effectively in order to build robust, reliable production agents. Let's get it done. So before we get into it, let's give a little bit of context on Linkedin. Linkedin is the platform for the agent development life cycle, and we sit at the core of how a lot of teams are building, deploying, and managing agents in production. Lynx Smith is our platform for agent engineering, and it has over 7,000 active customers. We also have a number of open source frameworks, which are incredibly widespread with monthly downloads in the hundreds of millions. We work with a lot of incredible companies from high growth startups like Klarna and Rippling to large enterprises like NVIDIA and ServiceNow, and we are used by half of the Fortune 10 cutting across many regulated industries from HR to legal to financial services. I'm your speaker today. My name is Caroline De Vittorio. I'm a software engineer at Langchain, and I lead our observability and evals work for voice agents at Langchain. So the number of voice agents being built and deployed is growing incredibly fast, especially when it comes to building customer experience or support agents. There's a number of use cases for voice agents, and there's a number of places where they are a good fit. But customer support support is one that fits really well for voice agents because voice agents are able to resolve customer issues quickly, which is good for a business. They're able to work around the clock. And ideally, they're able to resolve issues without escalating to a human, which can kinda help cut call center costs. But in order for them to be effective at this, the customer experience needs to be very good. You know, we've all interacted with a voice agent that had issues, and we all know how frustrating that is. And when this happens, it's incredibly counterproductive because, you know, users request to speak to a human, which defeats the purpose of having a voice agent in the first place, or they simply drop off the call, which is just bad for your brand or missed opportunity or revenue. Now we all dream of building great voice agents, but actually getting to a good voice agent experience is really hard. I talk to a number of people on a daily basis who are building voice agents, And one thing that comes up a lot is, like, being able to master all of these different pieces. You know, having a natural sounding voice, having the low latency, having the high accuracy, you know, putting all of these guardrails in place is something that's incredibly challenging. And both from, like, an engineering or traditional engineering standpoint, but also from an applied AI or an AI engineering standpoint. And so we at LandChain believe that, you know, making these improvements, getting to reliable good production agents starts with understanding what failures are occurring in your agents and iterating from there, which is basically saying, like, we need good evals. And so let's talk about evals and voice agents. So before we get into what I think we should do for building good evals, let's first talk about why evaluating voice agents is so hard to do. And there are three reasons that voice agent evals are incredibly hard. The first is that the thing that you're evaluating, the source of truth for your interaction between your voice agent and your user is an audio conversation. It's audio. And audio has so much more signal than text. When you're evaluating a voice agent, you're evaluating the conversation itself. So did the agent do the right thing? Did it call the right tools? Did any failures occur? But on top of that, you're also evaluating a number of failures and a number of engineering problems that only occur because of the audio. Did the agent talk over the human? Was there a lot of dead air or awkward pauses? Was the tone of the conversation appropriate both on the agent side, but also, you know, was the user frustrated? You know, did the agent pronounce all the words correctly? And the reverse of that, did the agent hear all of the user's words correctly? Those are issues that just don't occur in a text agent. And there are also things that just don't show up in a transcript, and so you have to evaluate this audio, and that's a tricky problem. The second thing is that a voice agent has a lot more turns than a text agent. So, you know, in in a voice agent, you're gonna have a lot of back and forth between the agent and the user. The agent asks a question, the user answers, the agent asks another question, the user answers again, somebody interrupts, somebody coughs, the user the agent pauses, the agent continues, etcetera. And so you can't really look at things on individual turn by turn basis. You need to look at the conversation holistically and ask at the end of the call, like, how was the voice agent? What did it do? Was this whole conversation from beginning to end an appropriate conversation? And doing this is a lot more difficult because, you know, there there there are a lot of things that are occurring. There are a lot of paths that the agents can take. Some of them can be good. Some of them can be bad. And so understanding and evaluating that nuance is tricky. And lastly, and this is also true for text agents, there isn't a single metric that you can measure that'll tell you whether or not a conversation was successful. There's a number of different things that you need to look at, and all of them in order, together will give you a picture of your voice agent, your voice agent's interaction with the customer, and whether or not it can be considered successful. So some example of this is, you know, the agent needs to follow all of its instructions. It should resolve a user's issues. It should be a good experience. It should be natural. It should sound good. It should avoid awkward pauses. It should call the right tools. It shouldn't throw any errors, and it should help move business metrics. And defining, you know, what success looks like looks like is going to depend on your use case, but there's there's a number of things that you're gonna need to measure in order to determine whether this this voice agent was successful. So the first thing I touched upon on on why agent, evals are hard was that the thing you're measuring is audio. And I wanna double down on that a little bit because we you you can't use the transcript as a proxy for the conversation. You really do have to to evaluate that audio. And there's a number of reasons why, that's the case. The first is that if you just evaluate the transcript, you're not gonna catch transcription errors. So, you know, the agent may have responded correctly given what it thought the user said, given what the transcript told the agent that the that the user had said. But if that transcript doesn't accurately represent what the user actually said, then, you know, your user's gonna get frustrated, and you might think that you had a good conversation when actually your conversation, you know, wasn't successful from the perspective of your user. The second is that the transcript doesn't capture, you know, tone and sentiment. These are huge pieces of a conversation. You know, sarcastic customer, a frustrated customer, all of these will play a huge part in, the user's experience of the agent, and they won't come across on the transcript. They'll only come across in the audio, and so we want to evaluate that. And then lastly, audio failures. Like, the failures related to the audio part of the conversation won't be captured either. So the agent speaking over the user, interruptions, awkward pauses, mispronunciations, audio generation issues, etcetera. So none of these failures are going to be visible in the transcript, and so but yet they play a big part in the user's experience. And so it's really important that we actually evaluate the audio directly. Okay. So what we're proposing today is a bit of structure and a sort of systematic way of thinking about how do you define your voice evals. And we propose thinking about this across three dimensions, essentially three types of evaluators that you should create in order to successfully evaluate and measure the success of your voice agent. The first is execution. And execution asks, you know, did the agent follow its instructions? Did it call the right tools? Did it use those tools correctly? Did it maintain a coherent conversation flow? And did it follow all of its instructions accurately? The second is outcomes. And outcomes asks, you know, did the caller like, did the conversation resolve what the caller asked for? So this can be, you know, if the caller is, calling about a particular issue, like, did that was that issue resolved at the end of the call? This can also be looking towards, like, the business KPIs. Is this voice agent actually moving the needle for the business in the right direction? If, you know, if if every conversation ends with, a a ticket that a human has to go take a look at, that might not actually be a a particularly great voice agent because you're not actually reducing kind of any of the work that people need to do. And the last piece is experience. And experience asks whether the conversation itself felt good. So that's the latency, the pacing, the interruptions, the audio, all of the things we've talked about. And these dimensions are orthogonal. So a call can perform well in one dimension and very poorly in another. So we wanna think about these dimensions independently, and we wanna measure, you know, holistically over all three of these dimensions in order to get a good sense for the accuracy and the success of our voice agent. So, for example, you know, an agent could follow all of its instructions, but maybe it doesn't move the needle on any of the outcomes because the instructions, you know, or the flow of the conversation needs updating. Or, you know, the voice agent actually does achieve all of its outcomes, and it does follow all of its instructions, but maybe the the latency is terrible and all of your users are really frustrated, and and really not happy with you. So in order to evaluate these three dimensions, you need tracing. And a trace is basically a full record of what happened under the hood of your voice agent. It's not a transcript because a transcript just tells you, you know, what the user said, what the agent said, and then maybe, you know, some of the tool calls. But the idea of a trace is that it gives you, you know, all of the, all of the, you know, tool behavior, all the model behavior, any reasoning, any timing. So, you know, how long did each of the pieces of your pipelines take? They help you differentiate between different failure modes. You know, was this a transcription problem? Was this a reasoning problem? Was this a a tool error? And then it helps you understand, you know, everything that happened under the hood so that when you see something happen in a conversation, you can kinda go back and say, okay. I can pinpoint where where this came from. So the point here is, like, in order to do our, you know, three eval dimensions, we also need to have good tracing because without tracing, you have nothing to evaluate. Or maybe beyond having nothing to evaluate, you have no way of understanding why that evaluation failed. Alright. So let's dive into the three dimensions specifically and get a little bit of detail on what we can do for each of them. So the first was execution. And here we were asking, did the agent do the right things? So, basically, we wanna ask, you know, did the agent follow its prompt throughout the call rather than simply, you know, the final response sounding reasonable? If we just ask, you know, an LLM as a judge, hey. You know, does does this conversation appear good? It may answer yes, but that's not really indicative of, like, did it actually do what it was supposed to do? And so we wanna check, you know, whether it followed all of its instructions. This can be things like, you know, did it stay within its scope? So Did it not promise things it wasn't allowed to promise? Did it complete all of the required steps, including things like identity verification, disclosures? And then did it follow all of those requirements at the right point in the conversation? You also wanna make sure that the, you know, agent followed its sort of style instructions and conversation flow instructions, so things like, you know, speaking to the user in a particular way, avoiding very long responses, etcetera. The second way you can evaluate execution or whether the agent followed its instructions is on tool use. So you wanna make sure that the agent, you know, used the correct tool, that that tool was called at the correct time, that important details like dates, names, IDs were all transcribed or passed to the agent or to the tool call correctly. And then you also wanna look at failure handling. You know, if a tool call fails and ideally a tool call doesn't fail, it's probably something we want to evaluate as well. Did the agent recover correctly from that tool call error? So did it not, you know hopefully, it didn't invent a response. Hopefully, it didn't promise things it can't promise. And then hopefully, it, you know, followed its instructions accurately in terms of how to handle that tool call error. And then the last part is, you know, the conversation flow. The agent should retain context across turns, so avoid asking the same questions many times over. It should base all of its answers on the context that it has provided to it. So, ideally, it doesn't hallucinate a number of other, conversation flow, evals there too. Okay. So the way you would evaluate execution in Lang Smith or or otherwise is with evaluators, code evaluators and LLM judges in particular. So the code evaluators are valuable where there are deterministic assertions that you can test for. So, for example, you know, tool a should be called before tool b. Tools shouldn't have errors. You can check for the appearance of disclosures in transcripts, something that, you know, you know is is expected to be there. Or, for example, you can check for heuristics, like the number of times a particular tool was called. If you're seeing, you know, this tool a is being called five times, you might start to to look into that or flag that as as something to look into. And LLM judges are important where you need, interpretation over natural language. So, if you have, you know, an agent and you're asking it to follow a series of steps, you can ask an LLM judge to check that all of the steps were completed in order. You can also ask you to check for hallucinations. You know, was this answer that the agent provided backed up by the context, that it had provided to it at the moment that it gave the response? Or, you know, did it say something that it wasn't supposed to say? Okay. The second dimension is outcomes. You know, did the caller get what they needed? What was the resolution at the end of the call? And so here, the question shifts from, you know, did the agent follow its instructions correctly to what did the call accomplish? Maybe even evaluating, you know, are these instructions the right ones? I like to think of it in two categories. So the first is, task resolution. Did the caller's stated reason for calling the agent actually get resolved over the course of the interaction? And so here, we care about the agent's capabilities and whether they are meeting the needs of your users. So if people are, you know, interacting with your agent to do thing a, but your agent can't actually do thing a, you would want to know about this. Maybe this is a gap in your agent's abilities. Maybe this is a gap in some other part of your business. That's something that you would want to to react on. And maybe if your agent can do thing a, but, you know, your customers are dropping off and they're not finishing the flow to actually complete thing a, that's also something worth investigating. And then the second part to it is looking at moving business metrics, moving business KPIs. And here the KPIs are going to depend a lot on what the purpose and the intent of your agent is and where it sits in in your business. But the idea here is that, you know, if you have an appointment booking agent, you wanna track, you know, how much revenue is that agent booking. How does that compare to maybe what you had before? Did it get better? Did it get worse? Why? So, you know, if you have a call center agent, for example, you could track, you know, how many minutes of call center time is it saving or, you know, how many tasks was the percent of tasks that it's completing without any form of human intervention. And these all matter a ton because they determine you know, at the end of the day, they determine, like, whether your agent is working for you or maybe if it's working against you. Great. Okay. So the last dimension that we wanna talk about is and that we wanna evaluate is experience. Did the conversation feel like a good conversation? And this is maybe the one that I think comes to mind, the first when we talk about voice agents, which is looking at, you know, the voice the actual voice agent interaction. Was that part successful? And so the first thing that comes to mind here is always latency. You know, if if the call is really slow, if it gets very laggy, callers will notice that. And it will end up with awkward pauses. It might end up with interruptions. It makes for very unsuccessful conversation. The second part you wanna look in, look at is interruptions and turn taking and just barging in general. So, you know, do the agent and the caller talk over each other? If so, for how long? How many times over the course of the conversation? How long does it take for the agent to stop talking when a user interrupts it, and then how does it recover from, you know, interjections, umms, ahs, coughs, background noises, etcetera. And then the last part we care about is the audio generation, the agent's speech. So we care about pronunciation, so particularly names, numbers, addresses, as well as, you know, issues over clarity. So, you know, did the caller have to repeat themselves? I did they have to give their phone number four different times before they got to the right, information? Did the agent say something that the caller couldn't understand and they had to to clarify themselves? And these signals all help identify sort of moments of friction. And oftentimes, these pieces of friction are particularly important because they block maybe other parts of, our voice agent. You know, maybe the agent can't complete its task because we never get past to collecting the phone number. So this part is also really key to unlocking a good voice agent. Okay. So concretely, when you're managing a a voice agent, evals aren't something that you just do once before you launch, but rather they're something that you iterate on consistently. So what this looks like in practice is that, you know, you start with tracing every call. You capture every single interaction. You capture the transcript. You capture the tool calls. You tap capture the timing information, and then you capture, you know, the final outcome. And then next, you're gonna run you're gonna define evaluators, and then you're gonna run those evaluators on those traces. So you're gonna define code evaluators. You're gonna define LLM judges for different pieces of the behavior, and then you're gonna measure a number of voice specific metrics. And then based on the errors that you see based on the conversations that get flagged, you're going to improve things. You're gonna improve your prompts. You're gonna improve your tools. You're gonna improve, you know, maybe the routing. You're gonna improve the voice settings. You're gonna try and reduce the latency, etcetera. And you're going to continue to measure, the performance of your agent against your evaluator based on the traces that you continue to collect as you improve your agent. And so the key takeaway here is that it's it's an iterative process. You know, you're not gonna define all of your evaluators at once and then just sort of see things happen. You're gonna continue to notice issues, define evaluators for those issues in order to track the changes, and you're gonna continue to measure your evaluator's performance over time. And now you might be thinking, okay. That's great, Caroline. And so far, this is all theory. It's all slides. How do I actually put this into practice? And so in order to make this more tangible, I wanna now walk through how you do all of these different things in LangSmith. Okay. So the first step is tracing. This is what a trace looks like in LangSmith. On the right, you have a full transcript of the conversation, which captures both what the AI said as well as what the user said. On the left, you have the audio recording. It highlights when the AI was speaking, when the user was speaking, and it allows you to play back any part of the conversation. Below the audio recording is the full trace of the conversation with details about all of the events that were happening under the hood while the agent and the user were speaking. This allows you to inspect any of these events further. So for example, this end of user detection has valuable metadata, and you can look at the details of LLM requests as well as tool calls among many other things. Okay. So now that we've traced our agents, the next step is to evaluate them. Again, here, I wanna highlight that the evaluations are going to depend a lot on what you're building and what you need to check for. So the the framework that I presented is not about giving you all the answers of, like, here are the 15 evaluators that you should create, but rather about how to think about, you know, how to which evaluators you wanna create in the context of the types of agents that you're building and what your goals what you're trying to achieve with that voice agent. And so the next step is creating pre actually creating those evaluators, and you can do that in LYNX with our code and LM as a judge evaluators. Now that we've talked through the kinds of evals that matter for voice agents, let me show you how to actually build them in LangSmith. Evaluators in LangSmith run on your traces and attach a score. We have two kinds of evaluators, code evaluators and LLM as a judge evaluators. With code evaluators, you write a function that runs against your traces. So for example here, we could check that a particular disclaimer appears in the transcript, or we could check that a particular tool was called. With LLM as a judge evaluators, you write a prompt, and it runs against your traces in the same way. And you'll wanna use these when scoring requires actually reading, understanding, and interpreting the conversation. You don't have to start from scratch because we ship a number of template evaluators, including ones that are built specifically for voice and that evaluate the audio of the conversation directly. Okay. So now that we have our evaluators, the last piece is to continue to improve this voice agent and then specifically be able to measure the improvements that we're making, be able to track progress over time. And that's something that we would do with monitoring in LangSmith. As you make changes to your agents, you're going to wanna keep track of how your agents are improving, and that's where Lanesmith's monitoring comes in. We ship a number of prebuilt dashboards, but you can also create your own custom dashboards with any of the graphs that you care about. On my dashboard here, I'm able to keep track of trends, including conversation volume, latency, as well as feedback scores from the evaluators that I created. So for example here, I have an evaluator tracking whether an appointment was booked in a call as well as an evaluator tracking whether an appointment was requested by the customer in a call, and this allows me to keep track of how my agent is doing at booking appointments for my customers. Okay. And so this sort of wraps up, how we think about evaluating voice agents and how we recommend that you think about evaluating voice agents. It's an iterative process. It requires putting a lot of different pieces together and figuring it out as you go. But I think if you think of things through the the the three dimensions that I shared, if you think of things through execution, through outcomes, and then through experience, you'll be able to get a holistic picture of how your agent is interacting with customers, whether that is going successfully, and then more specifically, about which areas of your voice agent need improvement that you can then go and dive into in order to, make the experience better. That's it for me. We're gonna open it up to questions now. Okay. So the question here on the screen is, in a VAPI bring your own SAP integration, we experience a cross tenant routing incident. Inbound calls to our number were answered by assistance belonging to another organization, while no call ID logs, webhook events, or CRM events appeared in our tenant. Support letter confirmed that the number was associated with another customer and removed the association, but we still lack a root cause analysis and evidence regarding possible data exposure. I sorry. I lost I can't see the end of this. Okay. From an evaluation perspective, how would you design preproduction and continuous tests for tenant isolation when the failure occurs outside the affected tenant's own observability boundary? This question is gonna be specific to, VAPI itself, and it sounds like this was a bug in their system, so I'm not really sure, that there's really anything that, I'm not really sure that there is anything that we can do here that would really help mitigate this issue. This sounds like an an issue with Fabi, specifically. Should we use satisfaction? If it was a great experience and the problem was solved, or long and hard and the problem was also solved, like quality, right conversation path. So satisfaction is also an important metric, and it's one that we track that that a lot of, like, call centers will track for, human agents as well. I think satisfaction in my experience tends to be lower with voice agents because a lot of people just don't wanna talk to voice agents in general. And so, it can be a little bit biased, you know. If you have an issue with the business or if you have you know, the customer was building inadvertently, they're gonna be very frustrated customer, and that's going to be reflected in the satisfaction part of things. That doesn't necessarily mean that the voice agent did the wrong thing or that it was a poor interaction. It can be biased by other issues as well. So I think satisfaction is an is an interesting score, and it's an important one to measure as well. But it's not one that is going to reflect the quality of your voice agent, like the engineering work and its ability to resolve customer issues. It's going to, be biased with certain other parts of your business as well. Okay. Is there safety concern with voice agents given that a lot of places authentication is on voice agents? Yes. There's always there's always a safety concern, and this is going to be something that, you need to think through very thoroughly when you're building the design of your voice agent to make sure that, only authorized or authenticated users are able to access, you know, their own resources. That depends a lot on your use case. I realize it's a bit of a non answer, but that depends a lot on your use case. It depends a lot on, like, how you're able to verify, the agent. There's a number of ways you can mitigate this issue. For example, you can validate you can validate with information on their part by not exposing that information with the voice agent. So if you give the voice agent the answer, they might be able to tell the agent the answer. But if you force the the user to give you the answer and then the only the voice agent can evaluate whether that is a correct answer with a tool call, that's one way to work around this. Another way to work around this as well is to, you know, allow for, some to to to to reverse potential, actions if something goes wrong and to build that into your system so that, you know, if a customer complains or if it's clear that upon, you know, further review something was wrong, you can actually roll back, the customer action. How the voice evaluators actually manipulate the voice records for scoring? Does it send the record to a multimodal LLM, or is it using other kinds of diagnosis tools? Yes. So if this is asking about, like, how does a an LLM as a judge evaluate the voice, if I'm understanding that correctly, then the answer is that you can send the voice, the audio recording to a multimodal LLM, and it will be able to provide, an evaluation of the, audio recording. So that's a great place to ask about things like tone, you know, frustrated users, any issues with, the audio quality, pronunciation, and another another a number of other, sort of what I would call it vibe based metrics, like, how did it feel scoring metrics. You can ask an LMS as a judge evaluator, to do that for you. As with all LMS as a judge evaluators, it's important to keep them very specific. You don't wanna ask broad questions like, you know, was this conversation good? Did the user seem happy? But rather ask very pointed questions where you would expect that, you know, several LLM as a judge or several even humans evaluating the conversation would be able to give the same answer every time. So a number of ways of doing this are asking, you know, more pointed questions. Did the user express frustration? Providing examples of what frustration might look like, as well as providing a really nice rubric for how to score the conversation so that two LLMs as a judge can see, you know, a one or five or three or two as being the same thing every time rather than having to guess at what the scale means. How can we measure the end to end latency between a voice agent generating audio and that audio being delivered to the caller over a telephony service such as Twilio. So Twilio will provide a recording of the conversation, which is going to be the best approximation you're gonna have for exactly what was said over the course of the conversation. And so the best way of measuring the end to end latency between the voice agent and, oh, you're asking about the voice agent generating the audio and the audio being delivered. That is a great question. Twilio might provide you with logs that would be able to, help diagnose, what that latency is. But, otherwise, a good proxy, as I was sort of alluding to earlier is that audio recording where you'll be able to see the latency between what the user said and then what the, agent said. And that should give you a pretty good proxy not for, the latency between the generation and the playing, but it'll give you a pretty good sense for how laggy or how slow or how awkward the conversation was. Yeah. Can you share an example eval to help improve interruptions and turn taking? Yes. So I think a great proxy for, interruptions is looking at the number of interruptions and then looking at how long it took the agent to stop speaking after an interruption occurred. And this is a great metric because it's usually a sign of how healthy other parts of your pipeline are as well. So if your voice agent has really high like, we talk about latency, but the question that we really need to answer is, well, what is too much latency? What becomes too slow? And you often know that something is too much latency when you start seeing those interruptions go, the the the interruption count go high because it's usually that usually occurs when the latency is so long that the user goes, hey. You know, are you still there? Starts speaking right as the as the agent was about to take its turn. It's also a good proxy for, whether your agent is saying the right thing. You know, people tend to interrupt their agent when they're not interested in the response or the response is too lengthy. And so that interruption count is a like, counting the number of times there was an interruption is is a really nice metric to have on hand. For turn taking, the things that you wanna look at are going to be, you know, how was how did the agent resume its turn after an interruption? Did it repeat a bunch of things that it already say, that it already said, or did it pick up naturally? Was it able to pick up on whether an interruption was actually an interruption versus, you know, background noise, a cough? And so you can do this based on, you know, after the fact of the conversation by looking at was, you know, if if the agent said, uh-huh. Sorry. If the user said, or mhmm, or whatever, did the agent, you know, pick up where it had left off and continue in a seamless way? And that's something that you would ask like an LLM as a judge. Do thinking models matter in voice agents? So this is always a question of a trade off between, you know, if you put too much thinking, your conversation is gonna get very slow. And if you don't put enough thinking, your agent might start doing things that aren't what you actually want it to do or or that reliability of your agent is going to decrease. And this is where evals are critically important because you can test them. And you can see, you know, how is your agent performing, when you put a lot of thinking on, how is your agent performing when you lower the thinking, and then what is the right middle in terms of, like, how you're seeing it perform, on your evals. Can we track in LangSmith also user disconnections during the call? Example, for an AI web interview. This is gonna depend a little bit on how you've built your agent. So if your disconnection is, I mean so your disconnection is going to end the presumably end the call, and so you will be able to see, that the call has ended, and you'll be able to evaluate on, why or, you know, when the call at what point in the in the call did the call end and then try to guess at why did the call end. So one thing that you can do here if you're seeing a bunch of disconnects is is use the context of, like, what was happening until that point in order to figure out try to guess at why the user disconnected. So, for example, if, you know, the, user had to repeat their phone number six times and then they were like, ah, you know, never mind, and then hung up, that would be a pretty good indication that the inaccuracy of the agent at capturing the phone number was the reason that the call didn't complete. You know, for a web interview, you know, maybe the agent is is asking the same question three times, and then the user just gives up. Maybe the agent, is asking questions that the user doesn't really care to answer, and that's what they gave up. You can kinda guess at that based on on the conversation. And so, yes, you can track in La Lanesmith user disconnections, and you can also try and understand, where they're coming from and what you can do about them. What to do after running the evaluations? Is there a tangible way to improve the agent prompt, a knowledge base, or tool definitions and test them before publishing the changes? Yes. There are. You can. So there's, so in terms of, improving the agent prompt, there there's a number of things that you can do. We have a product called Engine that sort of tries to identify failures and suggest edits, but a lot of it is going to be looking at the result of the evaluation, looking at the trace that backs up the the failed evaluation, and then, making a change that looks to address that particular issue. And oftentimes, those will be very, you know, self explanatory. If the agent called the wrong tools, you can go look at the prompt and say, well, why did it call the wrong tools? Maybe there's there's an instruction there that's, counter counterintuitive or not clear, or, maybe there's some, issues with the instructions. In terms of, like, testing changes before publishing changes, you can run simulations on your agent. So try to have another voice agent simulate a user and call go through the call flow and then evaluate that conversation output, or you can do this manually too by yourself calling the agent and and, observing the changes as well as running evaluations on that conversation that you, as a developer, simulated with your agent in order to, understand, whether or not it has moved the needle on on your evaluation. Can users jailbreak voice agents? This question is I mean, the answer is always the answer is always yes. I think there I would be a lie to say no, but that's true of any agent. And so just like any any agent, you want to guard against this as much as possible. At the end of the day, a voice agent is just an agent under the hood with with some voice. So, all of the different, you know, security concerns that occur for text agents also occur from voice agents, and so this is also something that you wanna test for and that you want to try and mitigate as much as possible. How should we think about the output of an eval? Example, using binary pass fail versus an integer, number of interruptions. Which types of eval outputs can language support? Are there any best practices on which types of outputs to use? Yes. So you wanna use, you wanna use binary pass fail where there's very clearly a pass fail. So if you're looking for something that, you know, you know it must be there or you wanna check whether it is there, then that's naturally a, you know, a binary, a pass fail metric. So here you would want to look at, you know, was the disclosure in the transcript? Yes? No? Did the user, swear at any point in the conversation? Yes, no. Did the user make threats against the agent at any point in the conversation? Yes, no. And then when it comes to numbers, you would wanna use an integer whenever this can be either a scale and you're trying to measure, you know, an average or an improvement or whether this is something that you're actually trying to count. So number of interruptions as you state. But, you know, if you wanna look at, if you want, you know, if you want to look at, different pieces of, you know, user sentiment, you might use, like, a zero one two scale. So, you know, zero being the user's frustrated, one being, you know, is neutral, two being, you know, the user is, does not, you know, seems, happy about the conversation. And the outputs just depend on what it is that you're measuring and, like, what the natural scale of those things is going to be. A lot of evals in general tend to be more on the binary side of things, and then, but some evals, especially ones that use, scales are, on a scale. Zero zero to three, zero to five, depends on depends on what you're measuring. How do we do voice, so how do we do evals for tone of the customer's voice? For instance, handling angry customers versus a sarcastic one. Yeah. Great question. So this is where, the multimodal, LLM evals are going to be particularly handy. You can run the audio through a multimodal LLM, and that will be able to infer a lot of these questions for you and do so, reasonably accurately. The, sorry. The main way you're gonna want to do this is to, you know, ask maybe provide you to provide a list of different things that you wanna check for. You know, was the was the customer sarcastic at any point in the conversation? Was the customer angry at any point in the cost in the conversation? Or you can try to do this on a scale. You know, what is the user sentiment about the voice agent at any point? And then provide examples for what the rubric of that scale would look like. So, you know, negative two might be the, agent seems deeply unhappy with the conversation and has, you know, swore swore the agent. Number, you know, one above that might be, you know, the agent expressed some frustration. Zero is neutral, and then and then, you get the picture. So it's going to be this one's one where for me, it's going to be about a scale and then providing, different examples of, like, where different pieces on the scale fall and, being able to track that inferred sentiment that average inferred sentiment over time. What does a golden dataset look like for agent evals? This is a great question. I don't so I think a golden dataset actually looks like a a scenario and an expected behavior or outcome. It's going to be really hard to have, you know, a a golden, voice conversation and be able to say, like, this is the perfect conversation given an input and an output because of the things we've talked about in terms of, you know, there being a lot of turns and, there being, you know, a lot of potential ways for the agent to get to the right answer and, all the different things that can happen over the course of a voice conversation. And so, a golden dataset is really, like, when the user calls about this and behaves in this particular way with a pretty in-depth description of of how they are behaving, the expected outcome from the voice agent is this. And that expected outcome will usually have two pieces to it. The first is the agent should be robust to, the different interruptions with the different issues that occurred over the course of the conversation. So, you know, it should be able to handle failures. It should be able to handle interruptions. But then the second part is, you know, if the cop if the customer calls about this particular issue, I want to you know, at the end of the conversation, I want to have resolved this particular outcome. And that outcome can is gonna depend on on your instructions and what matters. But, that that instructions can depend on on, yeah, the the convert the your instructions, but you should be able to have a a a goal that you expect the conversation to get to, and your agent should achieve that particular goal. I'm being told that that was our last question and that we are at time. But, we will continue to go through the questions that you left, and, we'll be in touch. Thank you so much for joining everyone, and, I hope this was useful. And, would love to would love to hear your thoughts and feedback.