AB.A model is useful only when its limits are measurable.All projects

Data Science

FaultForge

A model-evaluation workbench that measures how classification pipelines behave under missingness, noise, distribution shift, and other real-world stresses.

Robustness TestingBootstrap EvidenceCalibrationModel Reliability
View source repository↗
FaultForge interface

Why it matters

Replaces a single accuracy score with reproducible evidence showing where a model remains reliable and where it does not.

My role

Data scientist, model-evaluation designer, and solution developer

  1. 01Defined reproducible stress plans, metric guards, deployment ranges, adequacy rules, and confirmed breaking-point logic.
  2. 02Implemented repeated-seed evaluation, paired bootstrap evidence, calibration, prediction margins, flip diagnostics, PSI screening, and comparable version ranking.
  3. 03Designed evidence outputs that distinguish pass, performance failure, execution failure, flakiness, and inconclusive results instead of forcing false certainty.
01

Situation

Strong held-out accuracy does not reveal when a classifier stops working in production. Teams need evidence of where a model remains reliable, not just a single number.

02

Task

Build a model-evaluation workbench that deliberately stresses a classifier with missingness, numerical noise, distribution shift, categorical drift, unseen categories, and prevalence change, then turns the results into reproducible, honest pass or fail evidence instead of false certainty.

03

Action

  • Every severity level runs across multiple deterministic seeds, reporting mean, standard deviation, minimum, and maximum. Conservative guards compare mean minus standard deviation against user-defined requirements.
  • A failure only becomes a confirmed breaking point once all later severities also fail. Isolated drops are marked flaky, and the threshold is reported as an interval rather than a false-precise number.
  • The baseline covers accuracy, balanced accuracy, recall, precision, F1, ROC-AUC, PR-AUC, and Brier score, backed by paired class-stratified bootstrap evidence so unsupported drops are never promoted as confirmed failures.
  • Calibration, probability margins, per-row prediction flips, and PSI feature-drift rankings diagnose how degradation actually appears, while immutable model and dataset fingerprints keep every decision traceable to its exact inputs.
04

Result

Deployment ranges and adequacy rules (fewer than 30 minority-class examples is inconclusive, fewer than 10 is rejected) keep promotion decisions honest. A single accuracy score is replaced with reproducible evidence of exactly where a model remains reliable and where it does not.

Technical implementation

How the solution was built.

Each layer connects an implementation choice to the decision or workflow it supports.

06 layers
LayerImplementationOperational purpose
Experiment definitionVersioned model artifacts, dataset fingerprints, target semantics, stress plans and metric guardsMake every evaluation reproducible and tied to the exact model and data used
Stress generationMissingness, numerical noise, distribution shift, categorical drift, unseen categories and prevalence changeMeasure behaviour under realistic data-quality and population changes
Repeated evaluationDeterministic multi-seed trials with mean, standard deviation, minimum and maximum summariesSeparate stable degradation from one-off variation
Statistical evidencePaired class-stratified bootstrap intervals and conservative guard valuesAvoid declaring failure when the observed difference is not adequately supported
DiagnosticsCalibration, confusion matrices, probability margins, prediction flips, PSI and compound-stress ablationsExplain how and where model behaviour changes
Decision outputPass, failure, flaky, inconclusive and execution-failure states with JSON and HTML evidenceTranslate evaluation results into reviewable promotion decisions

Product walkthrough

Screens connected to decisions.

08 screens
FaultForge: robustness confusion calibration
01
robustness confusion calibration

Combines confusion-matrix, calibration, and robustness curves into a single evaluation view.

FaultForge: stress severity results
02
stress severity results

Shows how a metric degrades across increasing stress severity levels, with confirmed breaking points.

FaultForge: model registry
03
model registry

Organizes versioned model assets and keeps lineage and reliability evidence attached to each model.

FaultForge: model comparison
04
model comparison

Compares reference and challenger versions using evidence relevant to promotion decisions.

FaultForge: new experiment
05
new experiment

Turns a robustness evaluation into a guided, reproducible configuration workflow.

FaultForge: model detail
06
model detail

Documents a single model version's lineage, configuration, and evaluation history.

FaultForge: stress matrix
07
stress matrix

Maps pass, fail, and flaky outcomes across every stress type and severity level tested.

FaultForge: seaborn line
08
seaborn line

Plots metric trends across severity levels to visualize where performance actually breaks down.

Continue exploring02 / 11
Next project · Financial AnalyticsCasablanca RiskMeasure exposure. Understand uncertainty. Decide with context.↗
← Previous · Recruitment ATS ProView project archive
Project archiveGitHub repository ↗Discuss this work ↗