Talk
Deterministic code meets probabilistic agents: Testing and evaluating WebMCP Tools
WebMCP is a proposed web standard that aims to expose structured tools to AI agents on currently existing websites. WebMCP tools are pieces of code - Javascript code and HTML form based. We should test them the way we would test any JS/HTML. Then we need to understand what can go wrong when probabilistic agents start interacting with our tools. What are the touchpoints with agents powered by LLMs? What are the failure modes? Evaluations are here to help :)
About Kasper Kulikowski
Kasper Kulikowski is a Senior Developer Relations Engineer on Google’s Chrome AI DevRel team, which enables web developers to build the next generation of experiences with AI. He has a passion for performance, future of the web and working with coding agents. Before joining Google Kasper spent ~10 years in Dynatrace building Log Monitoring solutions. Outside of work, he enjoys powerlifting.