Video: Towards Automating Eval & Environment Engineering | Duration: 2788s | Summary: Towards Automating Eval & Environment Engineering | Chapters: Welcome and Introductions (0s), Panel Introductions (64.24s), Automating Post Training (227.275s), Eval Structure (368.045s), Robust Evaluation Design (534.405s), Building Eval Environments (783.31s), Evaluation Calibration (1090.525s), Agent-Led Questioning (1271.225s), Continual Learning Loop (1449.62s), Multi-Agent Systems (1629.795s), Pipeline Specifications (1892.28s), Evaluation Specifications (2050.325s), Evaluating Hallucinations (2243.47s), Prompt Variation Strategy (2407.805s), Sensitive Data Evals (2510.095s), Q&A and Closing (2560.5s)
Transcript for "Towards Automating Eval & Environment Engineering": Really excited for this, for for this webinar. We were a little bit late because we were chatting backstage about all the cool things that we wanna talk about. So we've got a really fun, agenda coming up. Few, like, minor, housekeeping notes. We will be taking a bunch of questions at the end. So there's a q and a section in particular, so it's not the chat. There is a q and a section. Please both add questions there as we're going throughout, but then also upvote the ones that you wanna hear answered. So I'll be going through them in order of ones that have kind of, like, the most upvotes. And and so plea please do that. We will be spending ten to fifteen minutes, if not longer, at the end on that. And then other minor, thing, this is being recorded, and we'll post afterwards as well. With that, I'd like to welcome, everyone, to to this panel for joining me. Maybe we can start by just going around and doing quick intros. I can start my name is Harrison, cofounder, CEO of Linkchain. Do background in ML, but do a lot of, product related work these days. Will, do you maybe wanna do an intro next? Sure. Hey. I'm Will. I lead applied research at Prime and Elect and do. a lot of work around post training, RL, evals, environments, etcetera, and, very excited to to have all those things set. Arin, do you wanna go next? Yep. Hi, everyone. I'm Aaron. I'm an applied researcher at Base ten, on the post training team. Yeah. I help our customers post train, and spend my day thinking about evals and harnesses. Super excited for this talk. Cool. I can go next. Hey, everyone. I'm Vivek. I lead a bunch of our applied research on our labs team here at LinkedIn. And similar to all the other guys here thinking a bunch about environments, post training, and then turning all of our really good data into into environments so we can make our agents better. Cool. So in terms of format, what I figured I'd do is I'd start with maybe a a a minute or two on kind of just, how we think about, or or the overall process for kind of, like, building agents. And then everyone here has thought a lot about how to automate more parts of that process, in particular, the parts around evals. And so I figured we'd go around. Everyone can share kind of, like, what they've been working on, and we can dive into that. And then we'll and then we'll end with q and a from from folks as well. So sharing my screen, briefly. Here we go. So so so this is roughly how we think about what we call kind of, like, the agent development life cycle in in general. And we think there's a bunch of different stages. First, you build a prototype of your agent, then you kind of test it and and start to iterate on it and make sure it's at the quality that you want it to be, then you deploy it. Then and then once it's deployed, you monitor it in production. And then once it's being monitored, you wanna you wanna take all those signals that you're getting and and and continue to iterate on the agent and pass it back into build. And this is a process, that that, a, really centers around kind of, like, traces as a center of of of of record for everything that kind of, like, happens. But it's also a process you wanna do as early and as quickly as possible. And so the best teams we see, they they launch, they launch early and they iterate quickly. And so, that's what everyone on this panel is working on in some form or another is basically thinking about how to speed up this iteration cycle as much as possible. A lot of that iteration cycle comes down to kind of like evals and testing things. So knowing where things fail, making changes to kind of improve there, and then and then making sure to evaluate it in in in the right manner. So so that's the general kind of, like, background about why I'm interested in it. The the iteration cycle is, you know, it's hard. It's important. How can we help people do it faster? And so, Aron, I wanna maybe hand it off first. How how do you think about this cycle, and and how do you think we can make this faster, easier, better? Sure. So just for context, you know, my day to day job is really this loop end to end. Like, customers will come to base ten with production traces or a harness as well. And, yeah, it's my job to turn those into evals and then into training data and then those into models that they can own and and serve. You know, there's some limiting factor here in terms of human capacity. So what I've been working on recently is building a a post training copilot that can automate parts of this workflow. So, yeah, I have strong opinions about what parts should be sort of automated and what parts should probably stay human. I think one way to answer this question is to look at where, customers really, like, struggle with this pipeline. For example, co designing evals is really what underpins the success of this entire loop. You can't do amazing post training or amazing model selection or even understand, like, whether a harness change is improving the end result of a model's rollout or not. So I can sort of get more into that. But in terms of just, like, data ingestion from traces from, say, Langchains, Lang Smith, down into building up a big suite of sort of principles and data mining from these traces really gives, like, a full list of, what can eventually go into a Rubrik based eval, which should probably be co designed with the human in an iteration loop between, actually, during rollouts from frontier models to smaller, less frontier models. And then once you get a bunch of different, like, post training science around, like, separation of different models by the judge, and fully reflecting the preferences of the human, then you can go and build up and curate training data for SFT or harness for r l RL. And I think that the right level of extraction, that should still have humans in the loop is really this co designing of the evals and then just being, like, really, really exact on your harness matching production. So, Aaron, maybe maybe, like, so eval's super important. You can you can, measure harness changes, do model selection, post train on them. What what do these evals actually look like in practice? Like, how many do you need? Are you testing, like, a whole agent end to end? Are they at specific parts? Like, what what what like, I I I think we can talk a bunch about how to generate these evals more effectively, but, like, what is the end result that we're even trying to can you maybe paint a picture for people here about what the best eval sets look like? Yeah. Exactly. Like, it's it's very rare for an LLM to be able to after data mine just one shot amazing evals. Typically, how I think of good and evals is to have, you know, somewhere between, like, three to five axes that sort of, at least orthogonal to each other so that you test non overlapping parts of performance of good and bad. So, you know, if you're having, say, a summarization task, like, maybe one axis would be completeness, right, that you actually extract all of the relevant information from some, like, messy, nonstructured data into something that's more structured. Right? Maybe something else would be, like, that you have the right formatting, that the LLM that's generating this output actually, like, adheres to the instructions of the task. And the key point I'm trying to make here is that these two axes are actually orthogonal. So once you have, And for every for for every, like, task in your dataset, do you test on all of these axes, or do some some data points you test on some, others you test on others? That's right. So each axis should be, generally, like, either iterated on sequentially. I think most customers have some idea of their preference of what's most important to get right. So what's least important to get right, sometimes we'll work on down that list and iterate on the the evals until there's generally, like, a bunch of subchecks, which are either, like, pass or fail. Right? Something that's binary is very easy for an LLM as a judge to do. So for each axis, typically, what can work is having, say, 10 or more subchecks, which are somewhat verifiable by a judge. And if different frontier judge models disagree, then that may mean that either your your judge prompt may be a little bit ambiguous or the system prompt for the task itself, it has issues with it. A a key part of this is probably, like, the the verifiability, that that you mentioned. Will, I know PrimeIntellect has a library literally called verifiers. How do you think about kind of, like, verifiers for these tasks? What's good practices? What do you see working? Yeah. I mean, it's it's definitely evolved a lot as a library since, the the old days before people were doing crazy agent environments. And so it's maybe a bit of a legacy term in terms of the the library origin, but it's it is broadly about thinking about what does good look like and what is a signal that is kind of robustly able to capture what we mean. And I think, like, it kind of is both exactly what you want for RL, but it's also, like, the robust version of evals where you kind of, I think there's a lot of things you can do as evals that are, like, correlated with correct, that are maybe useful offline, that are, like, deterministic text for keywords. But these aren't necessarily, like, robust evals where, like, if you wanted to kind of find backdoors in them or hack them, you could. And so I think that this is the way I think about it is, like, one is it's gotta be a signal that captures what you care about, but it's also gotta be one that is, like, robust to, like, adversarial, friction against the the edges of it. So, like, really good l m judges combined with rule based systems, I think, are often a really powerful approach, especially things that allow you to kind of incorporate trade offs between different, metrics. So you might care about efficiency. You might care about, kind of style. And so these are not necessarily zero one correctness things, but they're things that you want to be able to influence the the resulting behavior with. And I think another piece of it in terms of, like, capturing signal is also thinking about, like, what should be deterministic, what should be inferred with a model or simulated. And I think this is true both at the verification level as well as, like, the environment itself. There's this, a sudden a less, like, well known sudden essay called, like, the big world hypothesis or something like this, talking about kind of how and it it comes up again and again in other pieces of writing, but the world itself is often too complex to literally model with a a model. And so in some cases, you want the real world, where you're, like, you're using real code execution. In some cases, you want to have a user simulated by an LM because you can't have literally the user sitting in the eval. And so I think this is a design trade off choice that comes up again and again, especially when you're dealing with, like, more complex sedentary environments, things like computer use or, like, interacting with, like, large datasets or, like, complex applications, is deciding what can you get away with simulating and still have a, like, robust eval that is, like, full fidelity, versus what needs to be, captured by the human, in terms needs to be, like, designed in a way that is, like, more clever, but also clever in a way that is aware of the world beyond the eval. And you can do some of this from traces. I think there's a lot of, like, trace mining you can do to kind of get these good. And there's also, like, choices about, like, should I use a real tool here or a synthetic tool with a fake database behind it is a choice that, like, depends on many things that are beyond the scope of, like, just the eval loop itself. Can you maybe give an example of some of the places where you've chosen to either use a real tool or build a synthetic one just to make it a little bit more concrete? Yeah. So I think one thing we've seen is that, like, browser use is something that we initially were like, oh, we'll just go you browse the web. Turns out browse the web is one very nasty. There's all sorts of bot blockers. It doesn't scale very well with most tooling, and, also, it's hard to program in verifiability. And so something we've, like, done quite a lot of as, like, a standard practice whenever we're working on product related to computer use is, like, leaning into, like, just build a simulator, and getting good at, like, simulating very complex applications, from being able to, like, have data that represents what you see in the applications. And so that's one where I think, like, there you don't necessarily have to simulate the whole thing, but you wanna simulate the part that matters for the task. Yep. And so I think figuring out, like, how much of it you need and, like, when do you shop? Like, you can always make a more high fidelity simulator by, like, approaching rebuilding the real thing. But there's, like, this exploration exploitation trade off of the environment simulation itself that we see coming up quite a lot. Interesting. And and for folks, like, how how hard is it to build these, like, simulators or build these environments? Like, is is that real like, what what and and part of where I'm going with this is, like, what is the blocker for more evals? Like, when we talk about automating this loop, like, how much of that is automating the the creation of that that kind of, like, simulator or environment versus, like, you know, figuring out what good inputs for the tasks are or things like that? Yeah. I would say the cases where it's simplest are ones that are, like, either read only actions or where it's, where the thing itself is relatively cheap. So, like, coding is one where, like, creating coding evals, like, you have to design a good task. And I think for coding, designing a good task that you can then verify is the hard part. But the environment itself is not necessarily hard because it's just kind of where you're coding already. And I'd say search is, like, kind of one of the easiest in that. If you have a good, like, stable search endpoint that's gonna be used by your application anyways, that you can kind of, like, version and and stuff, you can then, like there's a lot of, like, kind of low hanging fruit tricks for generating tasks from, like, grounding in real world queries as well as, synthetic sampling of documents where then you can use the document to kind of back out queries that are maybe useful, and then do some kind of validation with passive one versus passive k testing. And so these are the sort of things that are very useful in kind of constructing tasks in settings where, like, the environment problem is not super hard. I'd say read, write tool use and user simulators are some of the first that people interact where there actually is kind of this question of, like, do you want to, like, you probably need to start simulating something, because you don't really wanna be spinning up, like, a real dev account for every single rollout. You want something that's gonna be cheaper to interact with and more reliable, to do rollbacks, etcetera. And so that the as well as, like, when the task itself is, like, kind of hard to get a ground truth answer for and deals with lots of messy role role systems, it's useful to kind of have the programmability of a simulator so that you can ensure that the task itself has some state that can be checked against. And. so it's I think that is very much like advanced version where I wouldn't necessarily recommend people start there, but it's useful to know that, like, that's kind of where we've seen a lot of this stuff go. It's like, you have to choose between, do you do it in, like, you do it in prod? Not not prod prod, but do you do it in the real thing, or. do you simulate? And both of them have their payoffs. Do you have a flowchart or anything that people can, like, follow to figure out when to decide? That would be very We should make one, but it also changes every day because, like, the frontier of, like, what we see is I think there's been lots of things where. we initially would have said, no. This doesn't make sense. This is not practical. But then it we actually realized it is practical. What changes to make it practical? Are, like, coding agents getting better so it's easier to to generate things? Is that it, or are there other things? We we do spend a lot of time working on, like, four methods for this. Like, it's a big focus of, like, I'd say a large fraction of the applied leadership team here is thinking about, this problem very systematically, and we approach it from many different angles. And sometimes someone will be like, hey. I got a thing. I found a new way to do this. And it's a mix of maybe we found a really useful tool on the Internet that we actually find, like, some linters, like, actually really good for code, style checking. And we don't need to judge for everything because we can get, pretty good rules through, like, existing utilities, or it's something that allows us to massively reduce the cost of bringing up a simulator for a certain type of application, which then makes this practical. And it but it also interplays with, like, your task design. Like, I think agents in production ideally have a very clear start and stop. I think when people are iterating on what their agent should even be, like, it's kind of a chicken and egg problem a little bit where when I'm trying to think what I'm building, like, I'm discovering use cases in coding agents oftentimes. And so a lot of these things are now emerging out of coding agents from people realizing a thing works in their coding agent. And they're like, oh, I my coding agent was able to do this. I wonder if I can productionize this. And then the coding agent scaffold ends up being a good starting point, but often you need to, like, then think about your skills and your other tool to tools you need and what sorts of system access. And, like, do you need, like, live applications? Like, do what sorts of documents does it need to be able to to interact with? Do you have enough of these documents, or do you have to synthesize them? How do you make sure documents are synthesized well? Like, there is a very long tail. And so I'd say, like, currently, like, large scale document synthesis is something that, like, I think is at this boundary where, like, it's it is very hard to do well. I think there's evidence it can be done well, but I I wouldn't say it's, like, robustly solved, versus, like, read write tool use for, like, local state based applications that feel, like, I don't know, like, linear. I think something that's, like, a mini database that doesn't need to be large and can just be, like, state for a user. I'd say this is, like, pretty robustly solved, as a way to do some layers. And then there's a structure in between. Vivek, anything to add in from what you've seen around what good evals, verifiers, or environments end up end up looking like? Yeah. I think, like, Harrison and I chat about this every day. So, like, that's sort of helpful. I think one one sort of, like, interesting thing that both, like, Will and Arin said is kind of this is fuzzy, but figuring out what's good. And I think, like, we we obviously get a bunch of traces. There's, like, one thing that we do a bunch, to figure out. Like, this eval, like, whatever it ends up becoming when we hill climb it, like, the same behavior that the eval measures is gonna get, like, imbued into our agent. So it's like one thing is essentially figuring out what our users actually asking, like, our real agent in prod. Right? So there's, like, a big task of okay. We get, like, millions of traces. Like, essentially, clustering a bunch of, hey. These are, like, five topics that users are sort of, like, asking about because there's, like, infinite emails that we can make. We need to, like, scope down. Like, Will was talking about, you can, like, simulate the entire world, but that's not, like, feasible or maybe even, like, useful. So kind of, like, one thing that we try to do a lot is look at a bunch of traces, figure out, like, what are maybe two or three axes that we wanna get better on. So, like, though those might be, like, multi hop reasoning. Or the like, for example, if you have, like, a bunch of tables in, like, Salesforce. Right? And it's like, oh, we see, like, as soon as the agent has, like, reason over, like, more than two tables, right, it's like it gets, like, super dumb. Right? So it's, like, kind of getting, like, super concrete by looking at things that actually happen in traces and, like, sort of just turning those into, like, little mini tasks that we can then, like, make a little bit more more difficult. I think, like, that's that's really helpful. Another, like, interesting I'd love to get other people's thoughts on this. Like, one is essentially calibration of what makes a good task. Because I think there's, like, models come out every few months, and it's not super useful if, like, all of your tasks just, like, pass. Like, that that's if you're not useful at all. So, like, one thing that we do is, like, we sort of, like, calibrate over different model strengths. So, like, Luna compared to SOL, like, one interesting thing is, like, is there a gap in between their pass rates? Like, that's immediately interesting. If there's not, then that might sort of tell someone that, like, I can maybe use LUNA for my task, and that that's great for customers. Right? If, like, you can do your task like that. The other is, like, if there's something that Sol is doing that, like, we don't, like, know a priority, can we, like, distill those same behaviors and get to Luna via, like, prompting or something like that? So I think, like, calibrating task is super important and, like, doing a different model's trends. And then maybe I'll just, like, leave one more thing. We can, like, Harrison, if you have thoughts, like, Will or Aaron, is, like, one that I'm super interested in the UX as well of making these. I think we all talked about, like, where humans need to be in the loop. Like, we use coding agents a bunch to help us make evals, like, understand how to make evals. And I think if anyone has thoughts on, like, what should the UX be? Like, how do I best put myself in the loop to review the work for generating evals? So, yeah, a a bunch of stuff there, maybe, like, open questions, like, Harrison will or, like, Aron have thoughts on that. I mean, I think our incident had some hot takes on on and strong. opinions on where humans should be in the loop. So I'd love I'd love to hear that. Yeah. Yeah. I I think right now, a lot of people, when they want to go to a coding agent and ask it, can you help me build an eval? They don't exactly know what to prompt the agent with. And I think that maybe one hot take is and I know that Viv had may have the same opinion, is that you can flip this the other way around, is that you can make the process of co designing with a coding agent be such that the agent asks you questions. Right? Like, you have the domain knowledge on this particular slice of task. And, if given that it's your, like, your business. Right? And, so it's completely fine, I think, for a coding agent to, like, mine a bunch of traces. Maybe do some sort of analysis like, like, Haiku, which is a dumb model, is making these kinds of errors compared to a smart model, and there's, like, this large separation, does that particular kind of gap matter? Or any sort of question around just really ensuring that this coding agent has the right knowledge of the tasks specifics, like, ensuring that it it essentially, like, what I'm building here is a way for which, model knows via, like, my own skills to get it to ask you the right sort of questions that are actually going to matter downstream for building evals and post training. So in terms of, do, you how do you get it to ask the right types of questions? Is that a skill that you give it, or how do you think about that? Yeah. Essentially, like, the in my opinion, the right way of getting models to do this correctly is just, like, imbuing my own knowledge into skills, which become, like, in context for the model. So, for example, something that I think if you just go and give it a bunch of traces and just say, hey. Like, can you build me an eval for this, which is, like, super general. My feeling is that a lot of coding agents right now, like, won't do that first step of proper proper exploratory data analysis. Like, they won't go and spend upfront compute on really understanding different clusterings, really understanding, like, the like, is there any errors in, like, your traces, such as just, like, wrong formats, tools that don't exist? Like, what is the full exhaustive list of tools available for this task? Like, what is the right distribution? How was these traces sampled from, like, production distributions? All of these pieces of information is generally, like, not something that users will typically go and ask models to do and, something that models are just, like, going to do on their own. So that's one example of something that I see day to day that I will, like, put into this skill and say, like, yeah. You need to, like, go do this thing first so that a nontechnical user who doesn't exactly know how to prompt can go and say, hey. How can I just, like, build an eval? And then the model will actually go and do this upfront work. Sure. Will, maybe switching to some of the stuff that you're working on, what are what are some of the things in the area of kind of, like, automating eval or this eval loop that you're most excited about that you or Prime Mint Elect are doing? Yeah. I mean, I think, so a lot of people, like, talk about manual learning, as, like, this new hot buzzword, but I think, like, it's starting to become clearer, like, what practical versions of it really are that you can do and deploy live. And so I think part of this is, like, constructing evals as a substrate for continual learning, where if this can become dynamic and you can find regimes where the the tasks are like their own evals and the human is in the loop at the right degree, but not necessarily, like, driving, but they're kind of responding. Like, essentially, you want the human to be a, like, meta reward model, like a meta judge, where the human is kind of steering, proposals of what good looks like. And they're saying, hey. This is actually not a good rubric. That's overfitting to this edge case or, like, this thing failed. That was actually the user being silly, and the user has asked to have been kind of wrong here, and so we shouldn't kind of bake that into the reward. And if you can do this from a trace pipeline, then you kind of have this flywheel where you can go from the stream of traces. You have either a a static, like, environment artifact that you can plug into, or you have pieces where the the the nonfixed components can be readily generated. And you also can then, have the the verification, the grading be the sort of thing that is grounded in kind of higher level specs. Like, I think the constitutional AI principles are still very relevant here where you can have these these meta rules. And, like, I think this is a pretty useful interface for people to think about when it's like I'm like, the same way we use skills, the same way we, like, give, like, specs to a coding agent. Here, what we're building is not a code base. Like, we're building an agent, and we're building the agent. And you can do you can do a lot of this in context. You can I'd say if it does right now, it requires a lot of scaling of self verification, and, like, it's not the most token efficient today from what from what I've seen. And, like, then the other angle is, like, once you know that you're gonna be doing something a lot, you can amortize this out by doing post training where, you can start, like, having the stream of new failure modes you see from a model that keep getting discovered over time. Because I think this is the thing that people see with with any new model that's deployed is, like, once you make a model better, it is something that you can deploy in the world, and then people will use it in new ways because the things that didn't work now work. And now there's a new thing that doesn't work. And so this is kind of what you want to to get, to work with manual learning. I think as a first approximation is like, it's just it's like faster model releases that packs the issues from the previous one. And so this is the sort of thing we're, like, doing, turning new traces into RL tasks for you to identify the player modes, then allows you to just kind of turn the crank. So so maybe one question there, Will. Like, conceptually, it makes sense. You see these tasks come in. You get a good sense of how it's being used in production. Totally agree that you're not gonna know everything that goes on. How do you end up knowing what, like, the right things actually to generate, end up being? Like, do you, like yeah. Is it is it some hard coded rubric that just always applies and you just notice that it, like, fails that rubric and so you're like, hey. We've already we've already got the verifier. We don't need to create a new verifier for each instance. Or is this, like does this have to be human generated? Can you look at, like, the back end like, how do you come up with the verifiers for these new data points you're adding? I mean, you ideally want it to be as scalable as possible, but, also, you need to trust it. So I think, like, there is a with with agents in general, there's always this, like, dance you're doing of, like, how much can you trust them, and how much could they prove to you that what they're doing is reasonable. And so you do kinda need to start small and get a sense of, like, okay. I know it can just do this thing. And then you're, like, you let it go, and you get a little more confident in it. At some point, you'll be like you'll get a little overconfident. You'll realize, hey. Wait. This whole thing isn't working. What's going on? You'll look under the hood, and you'll realize two steps back, it made some wrong turn that is now throwing everything off. And so you kind of want to kind of pace your, trust in the agent, with, like, what you see work in practice. And I think it's the sort of thing where you discover a lot of structure as you try to do this that is like like, it's not just an like, for the most complex stuff, it doesn't work for it to be just an agent. It's gotta be a structured pipeline. The same way that, like, back in the day, people were, like, doing pipelines for everything even, like, at the local search level. They'd, like, have each search step would be like a pipeline like a piece of code. I think we actually are re rediscovering this now at a much bigger scale because we are pushing at the limits of what the agent's timetable can do kind of in one go. And, for example, goal, I'd say, is really powerful, but it's also the sort of thing that can get steered off track really quickly. And you end up with something that's really, really, like, overbuilt that's, like, not quite the direction you're pointing at. And so I think to correct against this, you want other systems where you have more agents. And so I think multi agent systems are where we see this going. And you can kind of just throw more apes to the problem by having them check each other's work. And this allows you to kind of pull in multiple directions at once. And so as you overshoot in one direction, you can pull back in another. And some of these are, like, construction pipelines. Some of these are, like, more amorphous, but I think, eventually, you hope the models get good enough that they can just do this. You can train them for this as well. And so that's something where we've been exploring is, like, how do we train our own models to to, be able to do this work for us? And but it is still, like, a very nascent, space in terms of, like, the the full complexity. Vivek, I'll maybe go to you for the last question before we go to q and a from from from the audience. But what are some of the things that you're thinking about or working on, in the in this space? Yeah. I think, like, one super late to what Will said. I think, like, one thing that we actually all do, on this call is, like, we sort of make some sort of, like, spec, essentially, for an agent to go and execute as a way of generating tasks. And, like, specs are basically you, like, you go back and forth to the human and make I think it's kinda like a research question. Like, what should actually live in that spec, and then, like, how should humans give feedback on it? I think one other, like, super related thing is how do you self verify that a spec that you created actually produces, like, a useful task? And I think, like, one thing that we're working on a bunch is, once you have a spec, like, human sort of, like, roughly agree that that's good, that's still not, like, an environment. So there is this sort of, like, spec to task, like, creation pipeline. And, like, a lot of that is maybe just, like, guidance on, like, hey. If this is, like, a tabular dataset, like, use SQLite. This is how you, like, generate good looking data for it. And I think another part is, like, run an agent, like, on that actual task, and, like, have another agent. This is, like, sort of multi agent systems or, like, have another agent sort of, like, judge how well that agent did in that task, and, like, that tells you, is there something wrong in your environment? Is there, like, something wrong in your verifier? Like, does a human need to give feedback in there? So it's, like, there's a lot you kind of discover by just running the agent, like, in your proposed task and then, like, having another agent look if the trajectory is, like, correct. So I think, like, self verification of, like, that agent system of creating tasks, like, that's probably a super exciting, like, future direction. Like, it will help us, like, scale task creation. Maybe, okay. So maybe two questions. One, you mentioned this idea of a spec. Will and, Aron, is this something that you guys also think of or have in these pipelines? And if so, like, what is the spec actually end up looking like? Alright. If you wanna I can Yeah. I'll I'll just I'll just say quickly. Yeah. So in in what I'm building, the pipeline generally goes as goes through a bunch of different phases. You know, generally, the first one will be some kind of, like, data ingestion and then building up this spec. Like I've already said before, this is multi turn ensuring that the LLM has, like, a full picture of the scope of the task, where the boundaries are, etcetera. And then that will essentially go once decided, will go to state, in in the actual harness, not just, like, living in the context history of of of the chat. And, if you ever, like, close your pipeline and reopen it again, it exists on state. And then that really, like, forms the basis for how you co design, like, a harness for this task. Like, what elements of of the harness are you like, what tools are there, like, what skills and, like, state of the world, and and really, like, stuff sets the backbone for the rest of the engagement. But, yeah, typically done right at the beginning. And just very practically so is this just like a markdown file with all that information in it in natural language? Yeah. Sometimes. Yeah. To for for at least what I'm building, this is just done in. a in a in a JSONL, that has to fit some sort of, like, spec contract rather than just, like, open form in an MD. What's the what's the is there a standard schema for that JSONL? Yeah. There there is. So, yeah, this would be things like what is, sort of, like, natural language description of the task. So, like, let's say we wanna build up, like, a judge classification problem. Like, you would sort of, like, build up, okay, let this is kinda, like, what the system problem should look like. This is maybe what any, like, skills and databases it has access to, and and those tool descriptions and so forth. So it's really just like a full contract or, like, spectrum of of the actual task itself, rather than just, like, in a in a specific, like, MG with headers. But in reality, these are these are really much like the same thing. Just a different way of storing the data. Will, anything to add on spec shape? Yeah. I mean, I think something we see a lot is we we spend a lot of time, like, working with the broader, like, Evals community and ecosystem to, like, one, make sure that we can support the kind of popular popular Evals out of the box. But this is also a really useful source of ideas for specs because a lot of things people are doing, like, do kind of look like an eval that's out there in the wild or they look like some paper that was written where someone figured out a thing. Like, SuiteGrep, for example, is the sort of thing. There's, like, a lot of people who have, like, done SuiteGrep like things or, like, how to bench like things, or, like, any any sort of, like, search benchmark you want. There's, or, terminal bench or, all the three bench labels. Like, there are these shapes of eval problems that you do see come up time and time and again. And so I think that's kind of, like, one of the easier cases is when, like, the real world problem actually does map to a thing that has a pretty well understood spec. But then, also, there's cases where this isn't true. And then you have to think a little out. And I think especially in cases that are, like, less verifiable or that are, like, balancing multiple concerns, you kind of need to it's it ends up being more interactive because you can write down a list of rules of, like, the prompt the answer shouldn't be too long or it, should, like, not use too many tools. But, really, these are, like, continuous objectives. There's, like, a trade off between, efficiency and quality, and you want kind of to be able to sweep this Pareto curve. And so this is the sort of thing where you generally wanna, like, look at examples. And so I think approaching it from the few shot example age example like, I think few shot examples have also become, useful again in the sense of finding, like like even if you don't know what the the spec is, like, come up with one really good task. Come up with one perfect task that you think is for sure in the right email set and a grader that you really, really like, then do another one and start pointing at the axes of variation. And so I think once you start getting, like, canonical examples and axes of variation nailed down, then it's easier for the models to kind of move in the space and, like, create new things because they can see the kinds of moves that are allowed. Yeah. I think, like, maybe, like, one one interesting thing related to that is, like, when when, like, Will is talking about, like, these, like, specs, I think one really interesting way to, like, get specs is just, like, indexing a bunch of existing, like, really good benchmarks of, like, this is what a good environment looks like for, like, this type of knowledge work task. Like, this is, like, what Notion sort of looks like, what a Salesforce database sort of looks like. And then, like, you don't have to, like, shoot yourself in the foot on, like, trying to teach the agents, like, reinvent this thing, which, like, already exists somewhere else. And instead, you can, like, focus a lot of the reasoning effort on make this, like, particular flavor of the task good, but, like, copy a bunch of the stuff that lives in, like, automation bench or, like, terminal bench two for these, like, coding style tasks. I think, like, there's a lot of learning that can be done from, like, existing benchmarks that, like, get the get the proper shape. And, And, like, we have a skill to do that. And, like, we literally in the skill say, like, if something sort of fits this, here's a good example of what that look like, and then try to make, like, one task that sort of looks like that. Maybe jump into some of the questions from the audience and reminder this is in q and a, and people should upvote the ones they want because we'll be going in order of most uploaded. How do you go about evaluating hallucinations given that the evaluator needs to know the context and the knowledge base of the agent to determine what answer is right and wrong? Great. Yeah. It's a good question. I I'll talk for some of the cases that I've I've run-in to. So I've done some post training work for tasks like self summarization, tasks like voice agents. In the case of voice agents, it's interesting. Like, if a model at some point in, like, an agentic trace goes and reads some data from a database and then writes an answer back to, the chatbot with information that's, like, just not in its context window, I think that, like, post hoc, LLM as a judge is, like, very good at doing at, actually, like, spotting out whether this information is in the actual history or not. Hopefully, like, in a lot of cases, these can be, like, verified just with, actual code. But, yeah, like, my my basic opinion here is that models that are good general reasoners can very much easily look at, like, traces and then see what is actually in, a hallucination or not. So you can absolutely have, like, an LLM as a judge hallucination eval, and I typically have that for most of my post training engagements. Actually, maybe question on that, Arin. Is that is that an alum as a judge, or is that, like, an agent as a judge? Because I can imagine for some of these big trajectories where it's interacting with a ton of contacts, like, how do you like, can you even represent that in a way where it can go in the context window, and do you end up with more of this, like, agent as a judge type thing? Yeah. So, I mean, if the trace is short enough that it can fit within the context window of the judge, then, you can just dump it, and get a judge to give its answer. If it isn't, then you'll have to have some system of either, like, splicing the traces into, into different, parts and then have a window judge go look at each of these and kinda, like, concatenate, etcetera. But, yeah, for the most part, if if your trace can fit within the contact like, a 1,000,000 context window, which I think a lot of the time it can, then, having, like, a either a a panel of judges and seeing whether they agree or disagree is, like, a a good way to check that, the outcome of all of these judges are are right or not to to establish that gold standard, like, ground truth on hallucination. When you run evals, how do you, do you run them on one prompt or multiple versions of a prompt? How do you approach adjusting a prompt for a failed eval to not mess up other evals which already passed? Yeah. That's a good question, and I think it's a thing that you definitely need to have variation for. And I think one of the easiest ways to do this is if you have production inputs. Like, the best prompts are gonna be the real ones. And so if you can kind of mirror the situation in which a kind of initial prompt occur or if you, like, can maybe you want to you need to kind of fuzz it to kind of remove FPI or something like this. If you want it to, like, be something that is shared internally, this very much depends on the situation you're in, of course. But I'd say there are cases where, like, kind of dumb variation is fine. I'd say for RL particularly, it is easier to overfit to a prompt format. I think if you're doing prompt optimization or harness optimization, you can overfit maybe, but I think it's the overfitting is more possible when the LOM has decided to write all the prompts in a very kind of stilted formulaic pattern that a human would never write. And maybe that's actually totally fine for a sub agent. Like, I think, like, if you have subroutines in a larger system where the prompts are always gonna come in a very strict format, then it's like you kind of want to, like, just work with the format. You don't necessarily need to poke at the boundaries of it. And so I think it does depend on, like, how much native variation do you expect in production. How would you approach improving agents via evals if your agents work on sensitive data so customer slash user traces can't be used for improving the system? Yeah. I mean, talk to the end user. Like, I there's so much value in, like, eval like, traces are, like, one source of signal, but a lot of times, what you need is just an understanding of the problem shape. And you wanna be able to like, if you know, let's say, the tool definitions and you have kind of illustrative fake examples from the person you're talking to of the kinds of prompts, the kinds of outputs they expect, and you can iterate with them on kind of synthetic versions that converge on something where they're like, yes. This looks like the real data even though I can't show you the real data. Like, that, I think, has been very useful for us in these cases. There's two questions that are kind of related, so I'm gonna combine them. One one they're both around basically involving nonengineers or domain experts in eval creation. So, yeah, can nonengineers define what correct means for an agent? What tooling has to exist to make that work? How much does this deep vertical domain knowledge actually matter? Like, does really understanding one vertical's nuances make the environment better? Any thoughts on this general, like, yeah, domain expert involvement? It's crucial. Like, I can also let someone else jump in too. But just to be quick, I'd say, yes. Absolutely. Very useful. And I think the way we interact with agents in general and, like, chat bots is the sort of feedback mode you want. And so I think you wanna domain expert to be, like, able to use TrashDBT or cloud as, like, a. precondition. Like, I think that's the mode that is gonna unlock a lot of the utility for them to be able to, like, yell at the model and be like, hey. This thing is weird. Like, like, redlining the answer, basically. And I think a lot of people are have the feeling of, like, AI is good at everything except what I do. Like, we find in our personal workflows, it's not actually that useful. But then we see some release from some other thing, and we're like, oh, I can't believe it did that. And it's like, the models are, like, great at a lot of stuff, but, also, like, there's a lot of things they just don't know. And if you feel like you have to correct the models a lot, that's because you're an expert at something, and the models don't know what you know. And so your input as a driver of the model is useful as a user. And that the same thing goes for emails and experts. Yeah. I I think there's a kinda, like, a mom test around this, which is, like, like, your mom shouldn't really know about, like, what the verifier is, but, like, I think we're super bullish on, like, interviewing the user in some way. So, like, if your mom is able to sort of, like, say, like, this is the job that I do and, like, this is what good looks like, I think the world has not cracked that, like, UX yet of, like, getting, like, domain expert knowledge, like, into a spec. But, like, we have people who are, like, nontechnical, and, like, they give feedback on, like, hey. Like, this is actually my job. And I think, like, right now, the technical and, like, nontechnical pairing is, like, super powerful because, like like, I we know what verifiers are, but, like, not everyone needs to know. So I think, yeah, that that might end up being in the short term, like, really powerful, like, pairing these two these two, like, roles together. Yep. Hello? Alright. Any final words there on just the the right UX to to to build these evals if if if you're nontechnical domain expert? Yeah. I mean, I think for Evals, even if you're nontechnical, having strong opinions on what you think an end user would consider good or bad is really, like, the crux of what matters. So, you know, like, a product manager who's not an engineer but has strong opinions on what's going to matter for the end user is really, like, a a fantastic baseline to start working with any, LLM on building and co designing Evals. In terms of, like, the UX, yeah, I agree with Vivek that typically, like, it would be fantastic to have some way where an LLM, like, knows really pointed good questions to ask you with some sort of, like, dreaming of the future that this answer to this question is actually going to matter for training. I don't think we're, like, quite there yet, but, like, excited to be working on that space. I think we're all excited to be working on that space. And so thank you guys for joining this webinar. Thank you everyone for tuning in. Hopefully, it was fun, useful, interesting. We'll put up a recording. But, thank you. Thank you, everyone. I guess including Will's cat, who I guess showed up in the background a little bit and apparently is very cute, for joining this webinar. See you on Twitter or in the future.