Data for the SI era

Human reasoning,
in layers.

Ethora Labs produces expert-verified social science data for training and evaluating AI. Every claim traced to its source. Every disagreement kept with all its sides. Every judgment with the expert's reasoning.

Eight layers per record

    08

    Expert rationale

    Illustrative record · Herodotus, Histories 7.60

    An abridged, illustrative record. Hover or select a layer to see what it holds.

    01Thesis

    Technique compounds.
    Society does not.

    Illustration: technical capability versus social understanding over time A conceptual drawing, not measured data. A technical capability curve rises steeply and accelerates; a social understanding line rises slowly. The gap between them widens over time. Time Capability Technical capability models train their successors Social understanding follows, built by people
    Illustration of the thesis, not measured data.

    Models now write code, prove theorems and generate the data that trains their successors. Technical capability feeds on itself.

    Social and moral understanding does not work that way. In ethics, politics, religion and history there is rarely a single answer a machine can check. When models learn these questions from other models' output, diversity thins out and reasoning gets shallower.

    Human thought follows technical change; it cannot generate its own ground truth. So the social and moral training of advanced models will keep depending on data made by people: sourced, argued, contested and checked by specialists. That is the data we build.

    Technique

    • Outputs can be verified: code runs, proofs check.
    • Models produce training data for the next generation.
    • Progress compounds recursively.

    Society

    • Questions have several defensible answers.
    • Synthetic data narrows perspectives instead of adding them.
    • Understanding is earned from sources, arguments and experts.

    02What we deliver

    Data that teaches models to reason about people.

    Every product rests on one schema, so a history dataset and an ethics benchmark share the same provenance, quality metrics and delivery formats.

    • 01

      Gold datasets

      Expert-verified, multi-layer datasets licensed for training and fine-tuning. Non-exclusive, versioned and refreshed over time.

      Training · Fine-tuning

    • 02

      Evaluation benchmarks

      Tests of pluralistic, reasoned social judgment that show where a model is one-sided, shallow or simply wrong.

      Evaluation · Red-teaming

    • 03

      Expert rationale programs

      Custom collection of step-by-step reasoning from vetted specialists, in the disciplines and questions your model needs.

      Custom programs

    • 04

      Grounded knowledge API In development

      Sourced claims and perspectives your models can query, cite and stay current with.

      Retrieval · Citation

    03Method

    Gold is a measurement, not an adjective.

    Every release ships with the numbers behind it. This is how a record moves from source to delivery, and every change is logged: who made it, when, with which model and which guideline version.

    1. 1

      Versioned guidelines

      A written annotation guide per discipline. Every record points to the version it was made under.

    2. 2

      Model drafts, expert decides

      Models propose structure and candidates. Specialists accept, correct or reject each one.

    3. 3

      Double-blind annotation

      Gold items are annotated twice, independently, by experts who cannot see each other's work.

    4. 4

      Senior adjudication

      A senior scholar resolves errors. Real scholarly disagreement is kept and labelled, not erased.

    5. 5

      Measured agreement

      Inter-annotator agreement and error rates are reported per release. Test sets stay separate from training data.

    6. 6

      Datasheet

      Every dataset ships with its sources, rights, methods and known limits documented.

    04Disciplines

    One method, across the human sciences.

    We start where argument, interpretation and evidence are densest, then extend the same schema discipline by discipline. Only the domain ontology changes.

    Starting with

    • History
    • Philosophy and logic
    • Ethics

    Next

    • Sociology and social problems
    • Theology and religious studies
    • Literature

    Planned

    • Political theory
    • Legal theory
    • Anthropology

    With a deliberate focus on traditions today's models know poorly.

    • Islamic thought
    • Ottoman and Middle Eastern history
    • Eastern Mediterranean literatures

    These are often thin in training data and reach models mostly through translation. They are close to where our team works.

    Building social or moral reasoning into your models?

    Tell us where your models fall short. We will show you what a layered dataset looks like for that problem.

    The expert network is for historians, philosophers, sociologists, theologians and literary scholars who want their reasoning to shape how models think.

    hello@ethoralabs.com