Comparable evidence

Same label.
Different evidence.

HumanoidUptime only compares results when metric definitions, tasks and observation conditions align. In the current dataset, no cross-model pair clears that gate.

Comparison gate

Five conditions.
All must pass.

A matching headline number is insufficient. Every comparison needs a shared denominator and preserved operating context.

01Same metric definition
02Comparable task
03Known observation window
04Known fleet and attempts
05Comparable autonomy mode
0

cross-model pairs currently qualify for a performance ranking

Blocked comparisons

What looks comparable—but isn't.

Same model, different deploymentNot comparable

Digit × GXO

100,000+ totes

Manufacturer-reported cumulative output

Digit × Schaeffler

>1 year public span

Calendar evidence span; runtime undisclosed

Why blockedA cumulative task count and a calendar deployment span have different units, tasks and denominators—even for the same robot model.

Cumulative outputNot comparable

Figure 02 × BMW

90,000+ sheet-metal parts

Ten-month automotive production period

Digit × GXO

100,000+ totes

Warehouse tote-transfer workflow

Why blockedDifferent tasks, output units and missing attempt or assisted-cycle denominators.

RuntimeNot comparable

Figure 02 × BMW

1,250+ cumulative hours

Production record; schedule and downtime missing

Unitree G1 EDU-4

1 h 49 min per charge

Single-unit controlled laboratory scenario

Why blockedCumulative operating time and per-charge endurance answer different questions.

Energy continuityNot comparable

Walker S2

Battery swap within 3 minutes

Manufacturer product claim

Unitree G1 EDU-4

239 W mean power

Independent test in a defined scenario

Why blockedA swap-duration claim cannot be compared with measured power draw or endurance.

The useful result

No false leaderboard.

As comparable measurements arrive, this page will show them without changing the admission rules.

Read the methodology