Skip to content

Aurora HardCase

Platform direction

Turn failures into training data.

The most valuable training scenarios are the ones a model currently gets wrong. HardCase closes the loop between evaluation and generation — automatically expanding failures into targeted synthetic datasets. This is failure-driven synthetic data generation.

The loop

  1. 01

    Model

    A perception model runs against an evaluation suite.

  2. 02

    Evaluation

    Scenarios are scored; weak spots surface as measurable failures.

  3. 03

    Failure discovered

    A failure is a coordinate in scenario space — a condition the model gets wrong.

  4. 04

    Generate variants

    Aurora expands that condition into thousands of controlled variants around it.

  5. 05

    Dataset

    The variants become a targeted, labeled synthetic dataset.

  6. 06

    Retrain

    The model retrains on exactly the distribution it was failing on.

Worked example

From one weak frame to ten thousand.

A detector is unsure about a partially occluded pedestrian. Instead of hand-collecting similar footage, HardCase treats the failure as a seed and generates a neighborhood of related scenarios — varying occlusion, pose, lighting, and distance.

The numbers below are illustrative, shown to demonstrate the workflow.

Evaluation frame — illustrative
detection / pedestrianoccluded

Model confidence

42%

Variants generated

10k

Generating targeted variants…

Aurora HardCase

Close the loop from failure to fix.

If you already run structured evaluation and want failures to feed the next dataset automatically, we would like to compare notes.