OpsRoute v0.1.0
Published experimentOperational routing under strict contracts.
Two policy families test whether a successor can choose the correct decision, tool, arguments, approval state, policy code, and reason code—without repairing invalid output.
Refund policy routing
Duplicate payments, approval thresholds, fraud review, authorization, evidence completeness, settlement state, and refund windows.
Subscription cancellation and retention
Cancellation confirmation, locked contracts, balances, pauses, retention eligibility, requester authorization, and explicit intent.
Six systems · two surfaces
Capability and resilience are not the same result.
Confirmatory evidence contains 64 clean records. Adversarial evidence contains 32 untouched stress cases.
Separate frozen surfaces. No blended score.
Qwen base
source_base_supporting
Qwen adapted
source_adapted_full
OLMo untouched
target_untouched
OLMo full retrain
target_full_retrain
OLMo limited retrain
target_limited_retrain_10pct
OLMo anchored transfer
target_hybrid_anchored_distillation_10
| System | Semantic exactness | Strict validity |
|---|---|---|
| Qwen base | 0.0% | 51.6% |
| Qwen adapted | 54.7% | 87.5% |
| OLMo untouched | 0.0% | 0.0% |
| OLMo full retrain | 79.7% | 100.0% |
| OLMo limited retrain | 65.6% | 90.6% |
| OLMo anchored transfer | 85.9% | 100.0% |
Migration recommendations
The best successor depends on the constraint.
Anchored transfer leads the clean confirmatory surface. Full retraining leads adversarial semantic performance.
Minimum Direct Labels
OLMo anchored transfer
Eligible target systems: 3. Qwen remains a reference.
Maximum Confirmed Capability
OLMo anchored transfer
Eligible target systems: 3. Qwen remains a reference.
Maximum Adversarial Resilience
OLMo full retrain
Eligible target systems: 3. Qwen remains a reference.
Minimum Complexity
OLMo full retrain
Eligible target systems: 3. Qwen remains a reference.
No Source Teacher
OLMo full retrain
Eligible target systems: 2. Qwen remains a reference.
Original Labels Unavailable
No viable trained migration
Pure synthetic transfer never produced a balanced trainable target, so no completed zero-direct-label migration exists in this published case.
Eligible target systems: 0. Qwen remains a reference.