surety
Open source Apache-2.0 Python library

Automate only the decisions you can prove.

surety finds the confidence threshold where your Jev or Laya decisions are wrong at most 2% of the time (you pick the number), with a statistical guarantee. Everything below it goes to a human. It watches for drift and keeps the audit trail.

Open-source core (Apache-2.0) · Works with Jev, Laya and any /v1/systemone server · One email at launch, no spam.

Decisions by confidenceautoreviewwrong
Decisions plotted by model confidence36 decisions fall below the certified threshold t* and go to human review. 79 decisions above it are automated; 1 of those 79 is wrong, within a 2% error target.t* · certified0.5confidence →1.0REVIEW → HUMANAUTOMATE
automated 79 · wrong 1 (1.3%)target α = 2%

The problem

Your model says 0.93. Can you act on it?

A confidence score is a starting point, not a decision policy.

probabilities ≠ guarantees

Scores aren't promises

Decision models give you probabilities, not guarantees. Even the vendors tell you to validate thresholds on your own labelled data.

jev-latest → ?

Thresholds don't transfer

A threshold tuned for one model doesn't transfer to another, or to the next version behind jev-latest.

logs · oversight

Someone will ask for proof

Automated decisions increasingly need logs and human oversight.

How it works

Four steps from labels to safe automation

Per question and per segment, with α and δ chosen by you.

  1. 01

    Label

    A few hundred real requests, labelled by your team.

  2. 02

    Certify

    The loosest threshold that keeps automated errors ≤ α with probability ≥ 1−δ, per question and per segment.

  3. 03

    Gate

    Confident decisions run automatically and the rest go to review. If the question or model changes, everything goes to review until you re-certify.

  4. 04

    Monitor

    Random audits feed a drift alarm, and automation pauses itself if accuracy slips. Every decision lands in a tamper-evident log.

See it run

Same labels, two models, different thresholds

Real output from the repo's jev_vs_laya example, run offline on synthetic data.

examples/jev_vs_layareal output · synthetic data
$ python examples/jev_vs_laya/run.py --fake
400 synthetic tickets, alpha=0.1, delta=0.05 (split across all certificates)

question                         Jev (simulated)                    Laya (simulated)
------------------------------------------------------------------------------------
department              t=0.575  93.2% automated             not certifiable (n=400)
urgency                 t=0.5    96.8% automated            t=0.5    91.0% automated
is_refund               t=0.7    97.2% automated             not certifiable (n=400)

Thresholds differ per model: never copy one model's threshold to another.

Where a model can't meet the target on your labels, nothing is automated for that question.

Hosted version · coming

The library, plus the workflow around it

  • Review queue
  • Reviewer labels feed automatic re-certification
  • Drift dashboard and alerts
  • Compliance evidence export

10

design-partner spots

We'll certify one of your decision flows with you, free.

Apply for a spot

FAQ

Questions

Is it open source?

Yes. The core library is Apache-2.0 on GitHub.

Which models does it work with?

Jev (hosted), Laya (self-hosted), and any server that speaks POST /v1/systemone.

What does the guarantee mean?

With probability ≥ 1−δ over your labelled sample, the error rate among automated decisions is ≤ α, assuming future traffic looks like that sample. That assumption is why the drift monitor exists.

Do you see my data?

No. The library runs on your infrastructure, and its log stores hashes of inputs, not the inputs.

Who's behind it?

An independent project, not affiliated with TypeSafe AI or Convai Innovations.

Automate only the decisions you can prove.

Get one email when surety launches. Design-partner spots are limited.

Join the waitlist