Early-stage, running a pilot

Know whether your LLM tools are good enough to ship, and to stay shipped.

Minos is an evaluation and monitoring platform for LLM-based tools deployed inside enterprises, from chatbots to code review bots. Plug in your tool with minimal configuration. Minos builds the evaluation pipeline, acts as a deployment gate in your CI/CD, and keeps watching quality in production with human feedback in the loop.

Pre-deploymentevaluate, then gate
ReleaseCI/CD checkpoint
Productioncontinuous monitoring
Humansreview what is uncertain

How it works

From first connection to continuous quality, in four steps.

One trace store spans pre-deployment and post-deployment, so what you test before a release and what you see after it are directly comparable.

  1. Connect

    Point Minos at your LLM-powered tool with minimal configuration. No rewrite of your application and no labeled dataset required to begin.

  2. Induce

    Minos mines your existing tickets, logs and documentation to build an evaluation suite tailored to your tool and your domain.

  3. Gate

    The evaluation pipeline runs in CI/CD as a deployment gate. A change that degrades quality does not reach your users.

  4. Monitor

    In production, Minos keeps scoring real traffic and sends uncertain cases to people, so quality is tracked and reviewed over time.

Our approach

Three methods that make automated evaluation trustworthy.

LLM-as-a-judge is only useful if you can trust the judge. Each component below targets a specific reason it usually cannot be trusted.

Component 01

Cold-start pipeline induction

When no labeled data exists, we mine the artifacts a company already has, such as tickets, logs and documentation, to build an evaluation suite from scratch.

Component 02

Active learning-based judge calibration

We align an LLM-as-a-judge with domain experts using only a few labels, choosing the most informative cases for experts to review.

Component 03

Judge reliability layer

Every verdict carries a confidence estimate. Confident verdicts flow through automatically, and uncertain cases are routed to human reviewers.

Current pilot

Evaluating a customer-facing chatbot for a major airline.

We are running our first pilot with a major airline, evaluating its customer-facing chatbot against a curated golden dataset. Results are reviewed in weekly cycles with the people who own the tool, which keeps the judge calibrated and the findings grounded in how the business actually defines a good answer.

  • ToolCustomer-facing chatbot
  • CustomerA major airline
  • ReferenceCurated golden dataset
  • CadenceWeekly review cycles
  • StatusIn progress

Team and contact

Built by researchers who study LLM evaluation.

Minos comes from a team of computer engineering researchers at Bilkent University, working on LLM evaluation, code review and software analytics. We are based in Ankara, Turkey.

If you run an LLM-powered tool inside your company and want to know how it really performs, we would like to hear from you.

Get in touch

veli.karakaya@ug.bilkent.edu.tr Tell us what tool you are running and how you evaluate it today.