KeppyLab
Applied AI consulting
Verifiable agents for companies that aren't AI companies.
A production agent on your actual workflow in 2–4 weeks, with an eval harness included, so you know whether to trust it.
How The Pilot Works
One workflow, scoped
We pick the single workflow with the best effort-to-value ratio, then quote scope and timeline in writing. No hourly meter.
An agent on your existing tools
Weekly demos on your actual data, wired into the systems you already run; not a slide deck, and not mock data.
An eval harness by default
An automated test suite that scores the agent against your real cases, so you know its accuracy before you rely on it.
An honest recommendation
A plain-English score report and a go/no-go call, including "don't automate this" when that's the truthful answer.
See the full engagement shapes
From The Lab
KeppyLab builds small, sharp AI systems: developer tools, research infrastructure, agent evaluation workflows, and experiments that make software easier to operate through language.
gonogo - The eval harness from our implementations, open sourced. Scores an agent on your real cases and returns a deployment decision, including "not enough evidence yet."
cobol-reporter - RAG and report generation for understanding COBOL systems.
Disease Lab - Knowledge graph AI for rare disease literature and discovery workflows.
WorldEnder.ai - RAG-powered text adventures with coherent long-horizon world state.
describe - An MCP capability manager: discover servers, write client config, read the capability map back as resources and prompts.
Stay In The Loop
@yok0zuna | GitHub | LinkedIn | Hugging Face