Experience-guided robot manipulation

TraceFlow

Guiding Frozen Flow-Matching Robot Policies
with Success and Failure Traces

Reuse what worked. Steer away from what did not.
A bounded action guidance field, with no deployment-time policy updates.

1 The University of Hong Kong2 Southern University of Science and Technology

* Corresponding author

Explore the paper, code, models, and real-world dataset.

01 / On the real robot

Same policy. Different execution.

Three tasks. Two behaviors.
Selected rollout comparisons at 4× speed.

Compare each task vertically: baseline above, TraceFlow below.

T1 · Precise pickup

T2 · Ordered packing

T3 · Drawer selection

π₀.₅ Baseline4× speed
T1 · Precise pickup: Pickup stalls
Pickup stalls
π₀.₅ Baseline4× speed
T2 · Ordered packing: Wrong sequence
Wrong sequence
π₀.₅ Baseline4× speed
T3 · Drawer selection: Wrong drawer
Wrong drawer
TraceFlow4× speed
T1 · Precise pickup: Successful pickup
Successful pickup
TraceFlow4× speed
T2 · Ordered packing: Correct sequence
Correct sequence
TraceFlow4× speed
T3 · Drawer selection: Correct drawer
Correct drawer

All clips play at 4× speed; loops are not time-synchronized. The TraceFlow drawer clip uses only source 00:00–01:17. On smaller screens, swipe horizontally to compare all three columns.

Selected qualitative demonstrations, not a success-rate estimate. Aggregate trial results are reported below.

The idea

Experience changes.
Weights stay fixed.

A frozen robot policy does not automatically benefit from its previous successes and failures. TraceFlow gives that experience a route back into action generation.

A trace is a time-ordered state–action record with one terminal success/failure label. The TraceBank starts from task-training traces and can later admit the robot’s own completed rollouts. At each policy query, the system retrieves comparable moments and uses their action windows to guide the frozen flow-matching action expert.

Frozen policyBinary outcome labelsNo learned critic

02 / How it works

From past outcomes to the next action.

A retrieval path alongside the policy,
not a replacement for it.

TraceFlow pipeline: frozen VLM states query a TraceBank, progress-aligned candidates form an action guidance field, a frozen flow-matching expert generates actions, and completed rollouts are stacked for later rounds.
The guidance field lives in action-chunk space. Green indicates attraction toward successful action regions; red indicates repulsion from failed ones. Guidance is bounded and applied only during early integration.
01

Store experience

Keep state anchors, action records, episode metadata, and a terminal outcome bit.

02

Retrieve & align

Match the current state and task progress to relevant success and failure action windows.

03

Guide the action

Build local action densities and bound their signed correction to the expert’s velocity.

04

Stack new traces

Align completed deployment traces and admit them for subsequent rounds, without policy updates.

03 / Evidence

Better ordering. More completed tasks.

Measured on the real robot.
Not inferred from the example videos.

Ordered fruit packing · T2

One additional TraceBank stack.

On the same frozen base policy, ordered task success rises from 21/50 to 39/50 with TraceFlow, then to 47/50 after one stacking round.

Baseline
42%
TraceFlow
78%
+ TraceBank-Stack
94%
Real-world task performance
TaskMethodOrdered success ↑Wrong sequence ↓Subtask success ↑
T1 · Tape + hammer2 subtasksBaseline8/500/5034/100
TraceFlow16/500/5052/100
T2 · Ordered fruits3 subtasksBaseline21/5020/50123/150
TraceFlow39/502/50129/150
TraceFlow + TraceBank-Stack47/500/50145/150
T3 · Tape + cable6 subtasksBaseline0/108/108/60
TraceFlow1/100/1031/60

Ordered success requires all stages in sequence. Subtask success counts completed subtasks irrespective of order. T3 has only ten trials per setting; its result is descriptive, not a general reliability claim. The videos above are not labeled as the stacked setting.

Where does it help?

Simulation evaluates LIBERO, LIBERO-Plus (Long), and RoboMemArena. Gains are strongest in the studied Sequence and Transferring settings; Counting and Occlusion remain unresolved. Stacking is not monotonically beneficial, and bank outcome counts do not establish a universal positive/negative retrieval-budget rule.

Simulation results · Paper Table I
EvaluationBackboneBaseTraceFlowΔ (percentage points)
LIBERO four-suiteπ₀.₅96.8598.30+1.45
LIBERO-Plus (Long)π₀.₅79.8381.10+1.27
Arena Sequence + Transferringπ₀.₅56.75 / 62.9257.75 / 63.42+1.00 / +0.50
Arena SequencePrediMem78.92 / 83.0991.50 / 95.58+12.58 / +12.49
Arena TransferringPrediMem54.41 / 66.3462.00 / 69.08+7.59 / +2.74
Arena CountingPrediMem26.61 / 56.1225.49 / 53.50−1.12 / −2.62
Arena OcclusionPrediMem17.11 / 42.3415.69 / 40.60−1.42 / −1.74
Arena Full26PrediMem34.92 / 56.0134.99 / 55.04+0.07 / −0.97

LIBERO reports SR (%); Arena reports TSR / CSR (%). Bold marks the better value. Each row compares the same checkpoint and evaluator. These are selected settings, not one shared configuration: LIBERO averages suite-wise selections (98.00% with one shared configuration); Sequence uses Upper–Direct with K₊ = 50; Transferring uses Joint round 2 with K₊/K₋ = 8/8. Full26 is a separate suite-conditioned aggregate, not the average of the selected rows. LIBERO-Plus (Long): 2,519 paired trials, p = 0.0733.

Experience accumulation

What happens when the TraceBank grows?

Only the admission rule changes: Success-only adds new successes, Failure-only adds new failures, and Joint adds both. Previously stored traces remain retrievable; policy weights stay frozen. R denotes the collection round.

LIBERO-10 ten-round stacking curves for Success-only, Failure-only and Joint admission

LIBERO-10 · 500 rollouts per round

Success-only reaches 96.8% SR at R7, up from 94.8% at R1. Failure-only and Joint peak at R9 (95.6% and 96.2%), then finish at 95.2% and 94.4%. More experience does not guarantee a higher success rate.

Arena Transferring ten-round stacking curves and the shared 54 percent collector baseline

Arena Transferring · 200 rollouts per round

Joint peaks early at R2 with 62.0% TSR, compared with the 54.0% shared collector (dashed line). Success-only peaks at R7 (59.0%); Failure-only at R1 (58.0%). Combining outcomes gives the highest observed peak, not a universal stopping rule.

Observed single-lineage curves from the paper, not uncertainty estimates or predictions of an optimal future round. The two benchmarks have different collection budgets and should not be compared as equal-cost updates.

04 / Takeaway

Conclusion

TraceFlow improves long-horizon task success and action ordering by reusing successful and failed traces to guide a frozen policy without weight updates, but stacking gains are finite and task-dependent, with no improvement on Counting or Occlusion tasks.