AI / Developer Tool 2026
LLM Reliability Lab
A local-first lab for running models and prompts against the same datasets, and evaluating which configuration actually did better.
- Year
- 2026
- Status
- In development
- Platform
- Web · API · Local models
- Scope
- Web application, API, Evaluation engine
- 01DatasetItems imported and validated
- 02ExperimentModel, prompt template, parameters
- 03RunBounded concurrency, streamed progress
- 04EvaluationPluggable evaluators
- 05ComparisonMetrics and regression detection
- 06RAG and tracingPlanned
Overview
LLM Reliability Lab is an engineering tool for developers building on language models. It records datasets, experiments and runs, then scores them — so a comparison between two models or two prompts is reproducible rather than a feeling.
Challenge
Comparing LLM configurations is usually ad hoc: a notebook here, a spreadsheet there, and no shared record of what was tried or why one version beat another.
Direction
Treat it as a developer tool, not a chatbot. Application code depends on a provider interface rather than on any one model API. Experiments are fixed configurations, runs are recorded per item, and evaluation is pluggable and kept separate from generation.
What was built
- A provider-agnostic LLM interface, implemented for local models through Ollama
- Model discovery with distinct loading, empty and unreachable states
- Playground with streamed output, or structured JSON validated against a schema
- Datasets with JSON and JSONL import and per-line validation errors
- Experiments run across a dataset with bounded concurrency, live progress and cancellation
- Side-by-side run comparison with a word-level diff
- Evaluators: exact match, contains, local-embedding similarity and LLM-as-judge
- Aggregate metrics and baseline-versus-candidate regression detection
Technology
- Next.js
- TypeScript
- Tailwind CSS v4
- FastAPI
- Pydantic v2
- SQLAlchemy (async)
- PostgreSQL
- Ollama
Outcome
Four of ten planned phases are complete: foundation, execution, experiments and evaluation. A model or prompt change can now be run across a dataset and checked for regressions against a baseline.
Technical notes
Local by default
Models run through a local Ollama instance and semantic similarity uses a local embedding model, so evaluation needs no paid API.
Not yet built
Retrieval (RAG), tracing and model routing are on the roadmap. The RAG package exists as a placeholder and is not presented as a feature.