Video: Building Governed Agents: A Framework for Cost, Control, and Compliance | Duration: 3148s | Summary: Building Governed Agents: A Framework for Cost, Control, and Compliance | Chapters: Introducing LangSmith (10.96s), Agent Governance Introduction (106.065s), Governance Framework (318.95s), LLM Call Governance (492.345s), Tool and Agent Governance (699.23s), Enforcing Governance (841.86s), Model Management Systems (1112.815s), Routing and Limits (1351.545s), Governance and Cost Control (1591.15s), Data Compliance & Guardrails (1815.69s), LMM Gateway Architecture (2122.71s), Agent Security Q&A (2406.37s), Latency Overhead Analysis (2534.305s), Governance Middleware (2621.775s), Agent Identity & Governance (2749.88s), Gateway Policies (2820.145s), Limit Enforcement Details (2930.585s), Closing Remarks (3110.99s)
Transcript for "Building Governed Agents: A Framework for Cost, Control, and Compliance": Agents are a new kind of software. Endless inputs and non deterministic outputs. Build them the old way, and they break. Enter Lingsmith, the agent engineering It's model agnostic, cloud agnostic, and framework agnostic. Organization is shipping the best agents iterate with a system, and we call that system the agent development life cycle. Build, test, deploy, and monitor. You build with our open source frameworks. Deep agents, langchain, or langgraph. LangSmith Fleet enables anyone to build agents without code. Test agents using evals and experiments. Deploy changes in one click with LangSmith Deployment. Monitor every interaction from a single dashboard. Governance is built into every stage of the agent development cycle with Lanesmith LLM Gateway. Lanesmith Engine helps you improve agents autonomously and move through the development lifecycle more quickly. Blanksmith Engine is your agent for agent engineering. Alright. Hi, everyone. So excited to have you all here. We're excited to talk about governance and building governed agents using Linkchain and Linksmith, and the overall frameworks that you should consider as you think about cost, control, and compliance. I'm the product manager for our governance and l m gateway offering. I wanna give a little bit of context for what we do here, who we are. So before we get into it, Langchain sits at the center of how a lot of teams are building agents. LangSmith specifically is our platform for agent engineering. So we have over 7,000 active customers. Our open source frameworks are widespread with monthly downloads in the hundreds of millions. We work with also incredible companies as you can see all of the logos here. So anything from high growth startups like Klarna and Rippling to large enterprises, like NVIDIA and ServiceNow. We're used by half of the Fortune 10 companies cutting across many regulated industries from, HR to legal to financial services. So you can definitely trust us with your agent engineering. So today, we're gonna talk about why agent governance matters at all. Why are we here? Why are we thinking about this? Why are we, Linkchain, working on this? So we know that production agents introduce a different risk profile than traditional LLM apps. They have greater autonomy, meaning they have a greater need for visibility and control. Governance should enable teams to move faster, more safely. And it also with with production agents in the mix, there's three main pressure costs that come up. There's cost, there's reliability, and then there's compliance. So central policy enforcement becomes increasingly critical in this loop. We know that there is unpredictable spend that comes into the mix. So, you know, that agents can loop, they can retry, they can consume large contexts unexpectedly. So that makes cost harder to forecast. We also know that production agents need reliability. You need to make sure that when you send them out to the real world that they do things that you want them to do. So they need fallbacks, rate limits, and clear failure behavior that you can trace down and track to make sure that if something comes up, you can easily find the find what caused it and then fix it. And then finally, organizations need consistent documented policy enforcement permissions. You need compliance in the loop to make sure that you're meeting all of your expected requirements wherever you are. So to sum it up, the key takeaway here is that these pressures create a need for centralized runtime controls rather than application by application policies, especially as you scale. And Langchain builds particularly towards that to make sure that at every step of the way, from the beginning of your project as a solo developer, all the way to a fortune 10 company. We're by your side helping you on all three all three all three pillars. So governance as a whole, what are the starting points to think about? We think about it in five different areas. So on the first, there's governance, there's the govern action, there's ownership, identity, risk tiers, policies to think about. Next, you're gonna want to decide, you know, you're gonna select models, you're gonna figure out how to escalate, you're gonna figure out when to fail over. So you're gonna make some decisions around how these models are going to act and how you're going to interact with them. Next, you need to think about protect, protecting and enforcing controls at those interaction boundaries. After you do that, you're gonna observe. You're gonna measure behavior and outcomes. You're going to understand, understand what happened. And then finally assure. So especially for those who are more regulated industries, you need to make sure that there's evidence to manage changes over time. So even if you're a small engineer, you wanna have a trace of what happened. If you're a large enterprise and you get audits or you have compliance requirements, you wanna be able to assure those people who care that your agents are doing what they're expected to be doing. We also think of this as, like, part of the life cycle and your agent development life cycle. Right? You build something, you test it, you deploy it, you monitor, and governance falls within that entire cycle from the first step all the way to the end and as you loop back around to build some more. It touches on a lot of areas right across this life cycle. You're gonna have things like cost control as you're gonna have to figure out tool access. You're gonna think of audit trails. Human in the loop also needs governance and permissions, discoverability, and who can use which tools, which I which agents, which patterns. And then you wanna share a contact skills across an organization, across your team in a way that is safe, audible, auditable, and permissible. So we wanna make sure is that the key point here that falls in is that governance should really be part of how agents are built and how they're operated. It's really important that this falls into that main infrastructure. Problems come in when it's just bolted on because you miss that entire knowledge of step one all the way to step four. And in that, you might miss really critical pathways. So you might miss really critical pathways that your engineers need to fix and deploy later. In our case, we really think of it holistically and that the governance factors into every single aspect of how we build tools. So there's many areas of governance. Right? I mean, all the way from how your company signs on to how you measure permissions. All of those are really important areas of just using links as a whole. But then once you build those agents, you need to make sure that those interaction points with the outside world are measured and governed appropriately. So while your ecosystem might be reliable and you know who's who and who's what with a LangSmith, for instance, or Langchain. As you start interacting with the outside world with an LLM, with an outside tool, with an MCP server, and as agents go talk to other agents, you need to make sure that you're authorizing and you're certain of how those calls are being made, especially when you don't have moment by moment a human in the loop. Right? As those become more automated with AI, that's where really the pressure points start to come in. And as you start working at scale, that's where you start to see those moments where things could break or things could be unauthorized. And so the way that we think about it is across some various factors, and we either have robust offerings here or building towards them. So the primary one is l m calls. This is the bread and butter of how AI works. It's calling out to the provider and saying, you know, making that LLM call to get, get some some AI back. Right? When you're doing so, there's many things at risk. A way that we summarize it, and this is obviously not comprehensive, but to give examples and be as comprehensive as possible, we think about it on the size scale of cost. So, obviously, LLMs can become more expensive. As context grows, those calls can become more expensive. As models become more robust, they also the call gets more and more expensive as we've seen across some of the frontier providers. Model availability becomes also a crux. Your agent might become unavailable to your, to your customers, to your users, to your internal staff because the provider went down, the model went down, credentials weren't verified. There's many reasons, and they can have real production consequences. You think about private data. Right? And you wanna make sure that the LLM, the external model, gets only the data that it should. And especially if you're working in a regulated industry or with sensitive data, you need to really be thinking about what data is going to these external providers. Particularly as your teams start to use a broader breadth of LMS and providers using open weight, not open weight, that data can really propagate across many different ecosystems. And so you wanna be sure that only the things that you want to be going out that are unidentifiable are going out into the world. Typically, if you're trying to govern these, you're gonna think of things like spend limits, redaction, routing. You're gonna think about fallbacks and rate limits. You're gonna wanna make sure that your model your models are available when you need them and that you're staying within budget and you're being you're thinking about token optimization and economics as as those costs grow and as they become a more critical pathway in in decisions that you're making and access for your customers. Next, on the scale, to a smaller extent, but just as important, tool calls also bring on risk and require governance. Right? So there's unintended actions and production systems, an unintended call or unauthorized call to a tool. This person should not access to this tool or, you know, this type of prompt should only call this kind of tool. Those unintended actions can also bring in risks and are areas that you might wanna think about governing. So typical typically, what you might think about here is permissioning, having an audit trail to make sure that you can see what tools were available, who had permission to use them and access them, and then once they were accessed, how are they accessed. Another area is MCP calls. So similarly to LLM calls, this is once again data leaving your infrastructure boundary. And you need to also be sure that the data that is leaving is both authorized, who has access to which MCPs and how are they being used, what MCPs are even available. Those are all critical areas of governance. And you're going to think about access control and logging once again. And then finally, agents are starting to call more agents, not just calling a provider directly, but your agent might call another agent or a sub agent. And you need to figure out what is the identity allowed? What is identity of this agent? What does this agent allowed to do? Which other agents is it allowed to call? Do they take on the identities of the existing agents that you that exist, the existing person who's making a call, or do they have other permissions altogether? And this is where you might start seeing compounding errors because this really can scale up and snowball. So, you know, those errors grow and snowball. You might get unauthorized access across those agent chains. And especially as they start talking to each other, you can imagine that chain continuing for many agents before, before somebody catches it if you don't have the right governance layers in place. So here you might think about tracing, policy enforcement, and thinking what agent identity of how how that plays out for your individual agents and what permissions they take on. Now enforcing governance. Right? So those are we've talked about, like, the areas where you might wanna think about this. We've talked about, the types where this falls into the agent development life cycle. But now you kinda wanna actually enforce it and what are those the areas and how to think about it. So three main considerations are around cost, integration, and adoption. So to start, there's areas around cost and risk. Right? So well, building a simple proxy is really easy. Maintaining controls and integrations and guardrails is harder. So you might wanna think about, building versus maintaining. Right? So there's ways to build in governance across the areas that I talked about previously. And many companies build just their own proxies, build their own gateways, for instance, to manage those interactions, or build them even into their agents themselves. And so, you know, a basic forwarder is simple. Basic permissions are simple. When your company is three people, it's really easy to make sure that you're doing the right thing. But as you scale, that's where the real effort can come in and making sure that across your your company, across your system, you've really thought about across the board who is allowed to do what, and that becomes just much harder once you reach the 100 to 200 person threshold. And you have many different styles, you have many different models at play, you have many different agents at play, many different development styles. So thinking about that as as you kind of consider how to apply this governance. Then with guardrails themselves, not everyone cares about them. You know, it's really the more regulated industries typically or or customers who are in, in more regulated parts of the world. They they might require guardrails. So thinking about, for instance, secret detection and redaction, PII detection and redaction, other areas where you might wanna make sure that your agents are behaving appropriately, various guardrails. Those require precise optimization to eliminate false positives on those critical data. So you'd have to think about how to tune, how to make that available, the models themselves, even though there are many models out there. It takes work to actually make sure that they're doing things correctly and that you're you're able to adjust them to your needs. And finally, there's just the operational risk in setting this kind of governance up. Right? If you're running a custom control plane, this well, it might be easy to build it at the beginning. It introduces long term costs, operational overhead. You have to have a team that's always monitoring it because this sits in your run time. You know, you you need to be always on top of it if you're managing this kind of centralized tool. And we've seen this ourselves. Internally, we built a gateway that we've been using ourselves, and there's a lot of work that goes into just making sure that even works for our own staff. Next, you're gonna think about integration and integration in the broader agent stack. Right? So you're gonna wanna have full context on what calls happen, what agents said. Traces are a great way of managing that and noticing and debugging. You're able to set alerts when thing when certain things happen, when thresholds are passed, and able to refer afterwards and see an aggregate potentially through evaluations, how your agents behave and if there are certain patterns that are inappropriate. You also might wanna match this into monitoring and make sure that you pull this into insights. So evals are one way, general pattern noticing, using tools within Lingsmith for instance, like engine, if if you're doing so, gives you the ability to notice those patterns as effectively as possible. And then finally, there's instant debugging. So if there's an issue, if you suddenly have a runaway loop, you're gonna wanna be able to quickly and easily discover it. Notice those policy violations and inspect what happened. And especially if you're using centralized tool and centralized governance, you're gonna wanna make sure that it's not just the agents themselves that have that sort of monitoring. But those centralized tools should emit their own their own logs or traces to ensure that you can see whether the tool itself is also malfunctioning. And then finally, there's adoption. If you're just one person or a small team of three, then it's relatively easy to make sure you have access to the models that you need. And typically, you're gonna have a smaller scale of what those models are. But as you scale, you're gonna really want to make sure that your system is configured for all of the types of models and all the configurations that might be possible. Right? So you might wanna have a seamless endpoint swap. Right? So beginning adoption by updating the provider endpoint point directly to administrative tools. Right? Make sure that it points to a gateway to other tools that centralize that. And you're gonna wanna make sure that it's as easy as possible for engineers to apply it to agents that they're building. The models themselves, you wanna make sure that securely you're managing their credentials, that your systems are fully accustomed and able to take in a variety of models. What we've seen across our customers and our use cases is that it's not just the frontier models that people are using. Increasingly, there's desire for open weight models or swapping or even swapping based on certain signals. Right? So this model for this or cheap model for this type of work, and you wanna make sure that's as easy as possible. So saving the credentials in some centralized way within a centralized platform really helps with that. And then finally, there's cohesive governance. Right? So you wanna make sure that that system where all this is playing out is trustworthy. You wanna make sure that you're not duplicating governance, that your roles and responsibilities that you set up, you don't have to duplicate in multiple systems. Ideally, you set this up once and then you have all the tools that you need in one place to apply them to all the different settings and situations where they might apply. So critical points there, in terms of how to think about basically, how to apply this kind of governance in your ecosystem, how to enforce it, and some considerations for, you know, how that might apply to your system or to your structure and ecosystem based on the size you are, the scale you are, and how much effort it might take to apply that governance. Next is just reliability. Right? You wanna make sure that a gateway or a central management structure is reliable under your production loads. And this obviously matters with scale and especially with many agents running. You wanna make sure that everything that you're running through, any sort of governance platform can take the load. So there's resilience, and this comes through, like, many factors. Right? So there's resilience, there's routing limits, there's cost and access and making sure that everything's accurate and how you're measuring it. So on the resilient side, you don't wanna have a single point of failure. Because it sits in the critical path, you need to make sure that there's redundancy, that there's failovers, and there's mechanics to ensure that your production agents will never go down, that your internal agents keep on working so your employees can work, that your engineers don't get stopped with unnecessary provider outages or cost controls that are blocking them from using it. So this can be done through things like timeouts, through balancing loads, you know, being explicit about which types of measurements should fail open or fail closed. Adding those types of determinations at the front helps you both be predictable in what works when with your teams that they understand what mechanics are at play. But then also make sure that you're ready for all scenarios regardless of what will happen. And then finally, it's just main maintaining effective available centralized governance and making sure that it's always up and running, and you've talked through that as well. So with routing and limits, you know, I what I mentioned in terms of resilience applies here too. So you have automatic failovers. If a model provider is out, if they're timed out, if they've reached a budget, if they've been rate limited. You need to make sure there's one or two models that are ready and tuned to the use case and ready to go as a backup in those scenarios, even sometimes the same model through a different provider. Right? So often what we see is customers who have Frontier models that they're using for, like, a production agent, but they might have something through, like, Bedrock for instance that, that acts as a backup. Exact same model, tune in the same way, but just coming from a different source just in case something goes out. You also wanna set rate limits. So a lot of a lot of customers, we see just set them within the agents themselves. But if you have many going out of time or you have entire teams that are working together, you wanna make sure that you you enforce those partly, obviously, to catch runaway loops or agents that are misbehaving or sudden spikes of of traffic. But you actually also wanna make sure that you don't hit the rate limits of your providers. Right? You wanna make sure that those continue to be accessible because there's many providers out there that when you hit their limit rate limit, you're out for some amount of time, And that could be a real problem when you're trying to provide a product to your customers. And then finally, there's cost and access. Right? So there's you wanna make sure that what you're if you're especially for managing budgets and you care about budget management through something that's through centralized governments, you wanna make sure that how you're measuring that is accurate. And it comes a little bit harder than one might expect. It's not just the main token cost, but there's factors such as caching costs, of various types. There are different types of interactions, compression, that come into play that change how much that call might actually cost. And if you're looking at just pure token count for a call, there you're often likely to miss it. And, you know, in building our own tools and our own experience with using a gateway internally, and trying to govern and set budgets internally, we found that, we have to put significant amount of work to make sure that, at all at all phases, in all scenarios, our costs were accurate, and that we are taking into consideration all particular aspects of how a model call is made. And then keeping that updated. So as model providers make changes, release new models with new capabilities, with new, saving techniques, that those are applied within our central within our central management system. And we're quite proud of how far we've come in terms of ensuring accurate model counting and configurations within our own ecosystem. And then finally, you wanna make sure that there's really wide model access. We've heard from a lot of folks that provider lock in is a big fear, and Langchain really stands on that ground that having access to a wide range of models is critical and very important. We internally have access built in access to support frontier models, open weight models across the board through providers such as Fireworks and Base ten, all the frontier models that you can think of and a variety of hosting model hosting providers as well to make sure that you can easily switch models as needed. You can route flexibly across all the models that you would need and for different uses. Right? So choosing a cheap model for a particular task should be just as easy as using, you know, a cloud code, for your main your main work. So then finally, governing on top of that. Right? So you've allowed yourself the flexibility to set up this ecosystem and make sure that you allow your engineers as flexible an environment as possible. But there's many ways that you wanna make sure this is in check based on what your internal policies and thought processes are. You have budgets and investor dollars that you need to make sure are used efficiently. And as we've said, as models get more expensive, as as, and as context grow larger and prompts grow larger, you need to make sure that you're ensuring availability and that you're staying within budget and that token economics are taken into consideration. So multi level limits help you find help you ensure that at all levels from the individuals, the full organization, you've caught any sort of runaway loops. You've made sure to layer them based on daily, weekly, or monthly limits for instance. So you might say that during a single day, you you spend $5 or $10, but then also over the month, you wanna make sure you don't go over some amount. And you wanna make sure that these are, like, early indicators. Right? So it's really easy to see when you've hit limits because that might be really useful productive information for agents that are misbehaving, for agents that are working unexpectedly. And you have a bunch of factors as you build agents to do so. But you want those early indicators, and these and policy violations can can be one of them, including if there's a learning associated with it. You also wanna think about, you know, which model is used for which purposes. You might use a cheap model for a really cheap retrieval task or summary task, or you might use a more expensive model for truly the the high level thinking tasks that require that kind of model. Increasingly, what we're hearing from folks is that it's not really just a one model, one size fits all, but that there's different models for different purposes. And you wanna be able to switch easily between them and then eventually run evaluations and see how they're performing differently, and have that all be within one ecosystem because you wanna make sure that your quality isn't suffering while you're optimizing for cost. So, you know, areas where we see this is like matching on types of workload, on task delegation, and really reserving the most expensive safe frontier models and the most highly capable ones to those tasks that really require it and have complex reasoning needs, but potentially going to, you know, a cheaper open weight model for some of the easier easier tasks that don't require the highest capabilities, or don't need to be forward facing or front facing and can be something running behind the scenes. And then finally, you wanna think about context, and being thoughtful about caching and token pruning. So you wanna prune excess data to minimize cost, to minimize latency, to minimize exposure. So thinking about are there ways in a central way to notice those types of patterns and reduce their existence in what you're calling to an LM to save on cost and save on save on how long it takes to make the call. Evals and tracing are really great tools in this regard as well. Running online or offline evals for instance, to understand, you know, where your money is going, running insights on how to establish safe thresholds for that context reduction, for latency, maintenance, and minimizing exposure to threats. As we think about this in central governance, we also think about how sensitive data at runtime plays into things. Right? So, there's a few regulation, and this is just a short list, but these are the main ones that customers have come to us with. And this is really just to provide an example of areas that we ourselves consider, and we think about how we have how this works within, not just for instance, a central gateway, but also within governance as a whole. Right? We wanna make sure that we're compliant with all the newest regulations that our customers are gonna be facing. So some examples I'll walk through just to give you a taste are the CCPA, the California Consumer Privacy Act. This gives consumers the right to know what personal data is collected. And it it's it runs within the state of California. There's GDPR that's been pretty famous for some time and a lot of people have worked towards meeting its its expectations. It applies to citizens of the European Union residents. And it governs basically the collection, storage, and processing of personal data. A new one that a lot of our customers are working with us to make sure that they're meeting the demands of our is the EUAI Act. Once again, it applies to the European Union. And it classifies AI systems by risk level, imposes obligations, and it really is the next level of thinking about how to minimize the risks of AI systems, not just the data and how it's stored. And then finally, for our more regulated colleagues, especially in the health care industries, HIPAA is the Health Insurance Portability and Accountability Act. Within The United States, it governs protected health information. So make sure there's safeguards for how people access and store it, and how it's transmitted. We think of this in, you know, as new policies and as new regulations come out, we realize that they have different data handling requirements. And often we think about where sensitive data can go, who can access it, and what controls apply before it leaves the system. And so within Langchain, this applies both in central governance layers, but also across the full LangSmith platform. So we have many tools in place to make sure that roles are defined in ways to limit access that you think about are able to limit PII and tracing, for instance, in particular ways, that there's, audit logs to see how things and policies are changed so that any sort of suspicious changes are tracked in case there's enforcement requirements. And really the thought thinking is we're thinking about this in terms of risk based prevention versus trying to build out every possible control for every workload. We know that controls introduce trade offs around latency, cost, and complexity, and we wanna make sure to apply the right level of protection based on the agent data involved. From this, customers have talked to us about guardrails, and we've built some basic ones and are broadening our our our our offering here to make sure that PII and secrets are protected. And thinking about the broader range of what kind of guardrails should be available in a centralized way, especially for our customers in protected industries. So there's many different ways to apply guardrails and to think about them. So if you're building this yourself or using existing models, this is where this might apply. So think about it in structured PII detection. So it's the stuff like regex and pattern based matching, stuff like Social Security numbers, phone numbers, or, like, very formatted identifiers that have predictable patterns. There's also unstructured PII detection. And for this, you might use something that's more LLM based to be able to identify things within context. So stuff like named, entity recognition for, like, names, locations, affiliations. There's not really a fixed pattern or not, like, something that you could name very easily. An LLM is gonna be much better at defining and noticing those patterns in a way that just basic code might not. And then finally, secrets detection. So, this really even applies to any engineers who are working and making sure that you're not sending an LLM, your API keys, your tokens, your credentials, those can have real security consequences if if API keys are exposed and can run up a big dollar bill, if if not if left unattended, by, you know, by third parties who don't have good intentions. And so yeah. I mean, long story short, I mean, we'd see these as reducing risk. They're not an entire governance system, but they help you make sure that you can give confidence to your customers, that their data is gonna be protected when that's necessary, and that your engineers can work freely. And that, you can pass on, you know, high impact actions to human approval as needed, and figure out sort of where LLMs fall into that pattern, where tools fall into that pattern, and what you wanna be sharing externally outside of your safe ecosystem. So internally, we have built an LMM gateway, and it does a lot of the things that we've talked about. It isn't the only solution, but it is one that is built into Lingsmith. So if you're one of our customers, this is something that is available to to you today in public beta. And this is an example of how a gateway might work. And there's, you know, we've talked to customers, some of whom built it themselves, some of whom have been interested in ours. And there's some basic functions that exist here. And then obviously, we're expanding to bigger and better things, to make sure that we cover all of the needs of our broader range of customers as you mentioned at the beginning. So, the agent makes a request, and that request is intercepted by the gateway before it actually, which was the provider, whoever they may be. And there's various policies that run against that call to make sure that it is authorized to go to the provider. And if so, that the right data and the right structure of data is going to that provider. So there's elements such as fallback routing with retry policies, circuit breakers. There's spend limits. So making sure that across your organization, workspace, API keys, and users, or even custom headers if you're maybe charging customers, that those spend limits are applied and that your budgets are maintained. So whether it's your budget for individual customers and how how much calls they're about allowed to make or your own internal developers who are running coding agents or your production agent or internal agent and making sure that teams are not overspending their allotted abilities. With rate limits, you wanna make sure that, you know, a retry loop is caught, that sudden spikes in usage are caught, and that you might wanna fall back if to a different model if, for instance, one of your LLMs is overloaded with traffic. This might this might mean that you, this might also allow you to prevent your LLM from making that decision for you and blocking you for some amount of time, only extending the amount of time that, for instance, the customer doesn't have access to your product. And then finally, data reduction. We have it for our enterprise customers, and you can redact PII and secrets and make sure that the data that's reaching the provider is only the stuff that you actually wanna be sharing and doesn't have any negative consequences for your work. Within this, this the these actions create a trace, and these can be used in in tools such as engine. So running an agent over your agents to find insights you didn't even know were there to improve them, create PRs, and improve how your agents are running and how your gateway is even running. There is insights. So a lower scale here to find certain patterns and identify patterns. So one use case that one of our engineers use is he asked the insights tool to tell him why his all cost was so high, what he was doing, and where were there ways that he could optimize how he was using LLMs and then one. And then finally evaluations give you a really broad breadth of abilities to detect outliers, to measure performance, to measure quality of your agents. And this is an area that we're very interested in because, there's a lot of there's a lot of availability here to really have a gateway be a central place to improve your agents that doesn't have to require every single agent being developed in a particular way. And then finally, for those who care, who wanna see aggregate spending, aggregate behavior, we have dashboards within this gateway. So you can see in granular ways reports on who's spending what and how much. And then just to bring it all together, you know, the gateway is a runtime control plane. That centralized model access, spend controls, routing, data protection, and failure behavior. It separates governance logic from individual applications. And more broadly, strategically, you get model provider optionality in the ecosystem so you're not locked in to any particular providers and their budget their budget limit tools or, or their capabilities or their cost structures. Good governance should make switching models and providers safer, not harder, and should be built in to where you already are running and tracing your agents to make sure that you have the highest quality available to and the least headache available, and the entire process so you can focus on your customers, and not all the mechanics behind the scenes. Alright. I think we have some time for questions, and I've been seeing a lot of things coming in through. So I'm wondering. Alright. Should I be an I think I should be answering the one that's on the screen right now. Right, Angeline? Yes. Okay. Great. Thank you. Okay. So and the question is, if an agent thinks a wrong decision or gets compromised, how can we make sure that this mistake doesn't spread to the tools, MCP servers, or other agents it's interacting with? So at the end of the day, this depends on the agents and depends on you, the engineer. So there's there's many different ways. There's it depends on what the signals might be in terms of that decision or compromise. So you can run evaluations, both within a gateway, but frankly also even outside of it to notice patterns and how your agents are running to prevent them ahead of to prevent them in the future. You could also add governance of tool access. So for instance, adding a human in the loop element before certain calls are made. You might also think about, guardrails. So you might want to avoid sharing PII or secrets so that even if it gets compromised, you're all that's releasing is information that is safe to release outside of your ecosystem. But I think evaluations and having tracing and evaluations are your first stop shop to notice what actually happened and to be able to catch in the future and then eventually to add runtime runtime monitoring for those types of behaviors, stuff like, you know, jail breaks or, or hallucinations or that sort of thing. So that in real time, they can catch those those behaviors. I'm ready for the next one. Alright. The question is, what is the measured latency overhead introduced by Lang chain compared to direct API API calls to each l m provider? And how does that overhead vary by model, workflow complexity, and scale? So I think the answer is in the question, which is that it depends. So we work hard. So within a I don't know if the question is just Langchain as a whole. I mean, I think really Langchain is a broad set of tools and ecosystems, and so it's hard to answer one particular way. The gateway we keep, it also depends on how it's set up and all the different policies you're running through. If you're running, for instance, data protection models against your calls, those we're gonna require are gonna add more latency even like one or two or it could add many seconds. And we have tools internally to make decisions ahead of time that say, after this amount of time, maybe say five seconds, either fail open or fail closed. Right? So for the if you're using a data protection, for instance. And then for other tools, otherwise, our policy, you know, running through actual policy, like a stun policy, things that are pretty, low resource, those are pretty fast. And so our numbers are, you know, industry standard in terms of what you'd expect. Alright. Wait for the next one. What is the best governance middleware at runtime? Do you follow a framework, policy as code, harness code? Well, I think it depends what you're trying to do, which of the governance policies you're trying to apply. So we actually have a robust set. If you're building especially with lang chain tools, we have a robust set of middleware, for instance, even to do stuff like fallbacks. So we have, you can find in our docs, you can find, middleware to set a fallback, to set a fallback within your agent. So you wouldn't even need to get if you're like a one off engineer or you're on a small team, very honestly, we would probably recommend just using our middleware because it's much much much easier to measure and understand, and you can make changes as you need. It's really at scale of or if you have many agents running at a time that you might wanna start using a gateway. But I I recommend I recommend looking at our docs to seeing all of our availability. And I think it just at the end of the day, it depends on what exactly you're trying to apply. What would be your suggested governance MVP? Oh, thank you, Alfredo. My suggested governance MVP. I'm not sure I fully understand the question. I think as far as, your stack, maybe that's the question. I do I do think that, the you know, frankly, if you're using if you're building with Langchain or LangSmith, the OLM gateway is right there. It's really easy to set up. It just requires a few changes in your in your code, and you automatically get policies and governance infrastructure built in. So if you're starting from scratch, that's a great place to start. And we're trying to make it as easy as possible to apply all of these governance policies that I talked about throughout the presentation as easily as possible in the click of a button as much as possible, within the within your agents and ecosystem. Alright. The Langchain blog highlights the need for an LLM gateway to enforce centralized policies on agent to agent interactions. How can you effectively enforce these centralized gateway policies and emergency centralized swarm architecture? At work, we are facing these issues. So I think this probably requires more one on one, so feel free to send me an email, Claire. We can think about how we can monitor and govern your your agent to agent interactions. But starting to think about just identity. So we're working on such act aspects and thinking about how to apply agent identity within our interactions. And we have some tools within our within our existing ecosystem to think about agent identity and how how it's applied. So I would start with looking through our docs to see what's available. And then if you have any more questions, just send me an email. We can think through, you know, how this applies to your exact use case. Interesting. Okay. When is Langdon going to be self healing from these problems? So at the end of the day, governance is the point of governance is to make sure that agents are not just working in agent world. You wanna make sure that there's a human decision of what these agents are allowed to do. And so engine is a great place to kind of automate that process of understanding what should be happening and when. But I would say that you actually probably want a human setting some of these policies. You don't want them to be completely determined by agentic systems. You wanna make sure that the agents are responding to human signals, of what you want them to be doing. Alright. Next question is, if we have a lane graph with a predefined set of models, how will LMG Gateway apply its fallback policies on top of it? So, basically, the the agent will make an LMM call, and it'll you'll you'll set the the URL that it's supposed to call. Instead of calling the provider directly, you're gonna put a gateway URL, And then it'll run the LMM call through the gateway, and then the gateway will apply that fallback policy. It'll basically monitor. It'll say, before I make this call, is this provider actually available? Is there something else happening? Does the spend cap was the spend cap already reached? If that's the case, then it'll either block the call to that provider and send it to someone somewhere else that you've defined. I don't know that I can read this whole question. Oh, there it is. Okay. So my first your first point notes that an unmonitored loop can consume thousands of dollars in a single session essentially within minutes. The second proposes daily, weekly, and monthly caps. These timelines do not align a daily limit, does not prevent a loop from exhausting the entire month's budget in ten minutes. What is the actual granularity of limit enforcement? And, crucially, what is the late latency of usage tracking? If your counters are distributed across gateway nodes with eventual consistency, what risk of overage exists during a simultaneous activity spike before the cap takes effect? Finally, when a limit is reached, what happens to running agents? Are they abruptly cut off mid transaction, or do they shut down gracefully? Great. This is a very well thought out question. So the the point here is that you have I think also frankly, in this case, a rate limit for the minute by minute big spike in traffic is probably your best bet to stop that first big runaway loop. And then within if you can set dollar limits on minutes, on days, on weeks, on months. The point here with the monthly, for instance, is that, the the the spend limit should interact with one another. Right? So, the you would likely set a daily limit that would not be just as high as your monthly limit. Right? You would make sure that it's much lower and so that you're actually pacing yourself against your monthly limit. In terms of enforcement, the latency of usage tracking is, I honestly, it depends, of how you have it set up, what policies you're applying. But if you're just doing, like, very basic, like, spend limits, it's relatively low, and kind of within industry standards. And then what risk of overage existing or simultaneous activities before the count your counters. So you would apply you could apply the counters not just across your whole org, but you can apply them on API keys. So e what we have, for instance, a default, so saying each user is only allowed to spend x, y, and z per hour, per month, per per day. Similarly, each API key also has that default set. So it applies per API key. It's not just per for the entire org, which I think is what maybe your question implies. And then when a limit is reached, what happens to running agents? So right now they are shut down. You know, they are stopped. Like, that l m the l m call itself is blocked before it's made. But in these cases, a good idea is to set a fallback, for instance. So if that spend limit is reached, you fall you fall back to another model to make sure you're not losing activity. Alright. I think that was maybe our last question, and it was great chatting with you guys. Thank you for the really well thought out questions. Feel free to find me on LinkedIn. I think my email might have been shared somewhere. But also find me on LinkedIn, send me a message. Feel free to add me. Really happy to talk about this kind of stuff. And we learn so much from our customers and our listeners who are sharing questions and asking all of these really thoughtful things. They help us understand where we need to be focusing our attention. So look forward to hearing from all of you.