Skip to content

, Open source

Artefacts you can run, instead of claims you have to trust.

We publish the tooling we grade ourselves with, because an artefact somebody can run against their own data settles a question that a case study only asserts.

01Release plan

What is coming, and in what order.

These are commitments with dates attached internally. Status here mirrors reality, an entry moves to a public repository only when there is something worth cloning.

rudvanth-eval

Prototype

The evaluation harness we use on every engagement. Declarative suite definitions, pluggable graders, CI-friendly output, and regression tracking across model and prompt changes.

Licence · Apache-2.0

grounding-bench

Research

A benchmark for citation faithfulness: does a generated statement actually follow from the passage it cites? Includes an annotation protocol so others can extend it to their own domain.

Licence · CC BY-SA 4.0

refusal-suite

Research

Test sets for calibrated refusal, measuring whether a system stays quiet when the corpus does not contain the answer, which is the behaviour enterprise buyers care most about and public benchmarks measure least.

Licence · Apache-2.0

indic-domain-corpora

Research

Openly licensed domain evaluation sets for Indian-language and code-switched enterprise text, starting with healthcare and financial services terminology.

Licence · CC BY 4.0

02Policy

How we approach publishing.

01

Publish the measurement, not just the claim

Anyone can say their retrieval is accurate. A harness someone can run against their own corpus is a different kind of statement, because it can be checked. That is the only kind of claim worth making about engineering rigour.

02

Permissive licences by default

Code under Apache-2.0, data and benchmarks under Creative Commons. We are not trying to build a moat out of a licence, the moat is knowing how to use the thing.

03

Maintained or archived, never abandoned

A repository with no commits for a year and eleven open issues is worse for our credibility than no repository. Anything we stop maintaining gets archived with a note saying so.

04

No client data, ever

Nothing derived from a client corpus is published in any form, including as synthetic data generated from it. Benchmarks are built from public or purpose-created material.

Contributing

Issues and pull requests are welcome on every public repository. The contributions we value most are new hard cases for the evaluation suites, a document, a query and an expected behaviour that current systems get wrong.

We also review security reports through a coordinated disclosure process.

Research enquiries

research@rudvanth.com

Security reports

security@rudvanth.com