Validated analyst memo
Authoritative structured outputValidated migration recommendation.
A constraint-aware product recommendation generated from the frozen evidence graph. The authoritative GPT-5.6 Sol memo remains unchanged beneath this executive layer.
Recommendation
OLMo anchored transfer
- Applies when
- Prioritize maximum confirmed capability among safety-eligible target migration methods.
- Primary reason
- Canonical lexicographic selection. Confirmatory semantic exact is 85.9%, the highest among target candidates, and confirmatory strict validity is 100.0% with zero unauthorized actions and zero approval bypasses.
- Critical tradeoff
- OLMo full retrain performed better under adversarial pressure (68.8% versus 62.5% semantic exactness).
- Requirements
- Trained source teacher; 10 direct original anchors; 214 teacher outputs; target retraining; TEACHER HYBRID LORA pipeline.
Model
GPT-5.6 Sol
Mode
Structured output
Attempt policy
One repair
Validation
Offline passed
Executive readout
The highest confirmatory semantic-exact value among target migration candidates was 85.9%.
The highest adversarial semantic-exact value among target migration candidates was 68.8%.
The canonical migration analysis recommends different trained migrations for maximum confirmed capability and maximum adversarial resilience; no viable trained migration is recorded when original labels are unavailable.
Transfer assessment
Confirmatory semantic exact was 85.9%, 79.7%, 65.6%, and 0.0% across the four target migration candidates, in descending order.
Two trained target candidates tied for the highest confirmatory strict-validity value at 100.0%.
The 24-direct-label trained candidate exceeded the fully adapted source reference on confirmatory semantic exact and strict validity and on adversarial semantic exact and strict validity.
The untouched target recorded 0.0% for semantic exact and strict validity on both confirmatory and adversarial surfaces.
On the adversarial surface, two trained target candidates tied at 93.8% strict validity and 71.9% argument F1; their semantic-exact values were 68.8% and 62.5%.
Adversarial weaknesses
Each trained target candidate recorded one unauthorized action and one approval bypass on 32 adversarial examples, despite recording zero false actions.
The untouched target recorded two unauthorized actions and two false actions on 30 adversarial examples, while its approval-bypass count was zero.
Each trained target candidate recorded lower counts than the fully adapted source reference for unauthorized actions, approval bypasses, and false actions on the adversarial surface.
The representative-case set includes cross-system disagreement, parser/schema failure, prompt-injection resilience, refund-family contrast, subscription-family contrast, and hybrid-versus-direct-training contrast cases.
Constraint-aware decisions
Migration recommendations
OLMo anchored transfer
Canonical lexicographic selection among trained candidates. Method accounting records 10 directly used original labels, compared with 24 and 224 for the other eligible candidates; it also records 224 upstream original labels.
OLMo anchored transfer
Canonical lexicographic selection. Confirmatory semantic exact is 85.9%, the highest among target candidates, and confirmatory strict validity is 100.0% with zero unauthorized actions and zero approval bypasses.
OLMo full retrain
Canonical lexicographic selection. Adversarial semantic exact is 68.8%, the highest among target candidates; adversarial strict validity and argument F1 are tied for the highest observed values at 93.8% and 71.9%.
OLMo full retrain
Canonical lexicographic selection. The selected workflow is classified as DIRECT_TARGET_LORA; the hybrid alternative is classified as TEACHER_HYBRID_LORA.
OLMo full retrain
Canonical selection among the two direct target LoRA candidates. The selected candidate records 79.7% confirmatory semantic exact and 68.8% adversarial semantic exact, compared with 65.6% and 40.6% for the other eligible candidate.
No viable trained migration
The canonical profile records no eligible trained system. All three trained migration candidates use original labels either directly or upstream.
Tradeoffs
Direct-original-label accounting across the trained target candidates is 10, 24, and 224 labels. The smallest direct-label count is paired with upstream use of 224 original labels.
Hybrid accounting records 10 direct original labels, 224 upstream original labels, 10 selected anchor labels, and 214 synthetic labels within 224 unique target-training examples. It also records 379768 source-teacher training tokens and 323601 teacher-generation processed tokens.
The hybrid lineage additionally records 224 original labeled records used to design the distribution, 768 previously generated synthetic candidates, and an accepted synthetic pool of 719 examples.
Recorded target-training accounting is 272568 processed tokens and 504.76380866600084 seconds, 272643 tokens and 617.3140285830013 seconds, and 272634 tokens and 635.4282235830033 seconds across the three trained target candidates.
The hybrid workflow is classified as TEACHER_HYBRID_LORA, while both direct retraining workflows are classified as DIRECT_TARGET_LORA. Its lineage separately records 437.86 seconds of source-teacher training and 1122.69 seconds of teacher generation.
Scientific limitations
- Canonical restriction: one seed supports reproducibility, not statistical significance. Source: evidence pack phase4-evidence-4587ec90e2e79eb4.
- Canonical restriction: findings apply only to the pinned Qwen-to-OLMo pair and OpsRoute v0.1.0. Source: evidence pack phase4-evidence-4587ec90e2e79eb4.
- Canonical restriction: the adversarial split was evaluated exactly once per system without tuning. Source: evidence pack phase4-evidence-4587ec90e2e79eb4.
- Source systems are scientific references rather than target migration recommendations; all six frozen profile records recommend target candidates or no viable trained migration. Evidence IDs: migration_profile:minimum_direct_labels, migration_profile:maximum_confirmed_capability, migration_profile:maximum_adversarial_resilience, migration_profile:minimum_complexity, migration_profile:no_source_teacher, migration_profile:original_labels_unavailable.
- No weighted composite score is used; recommendations are frozen profile-specific selections. Evidence IDs: migration_profile:minimum_direct_labels, migration_profile:maximum_confirmed_capability, migration_profile:maximum_adversarial_resilience, migration_profile:minimum_complexity, migration_profile:no_source_teacher, migration_profile:original_labels_unavailable.
Evidence-backed next steps
- 1Repeat the frozen evaluation across additional seeds before making statistical-significance claims.
- 2Evaluate additional source-target model pairs and task versions to assess scope beyond the pinned pair and OpsRoute v0.1.0.
- 3Preserve one-shot, no-tuning handling for any new adversarial split and pre-register comparison rules.
- 4Track direct original labels, upstream original labels, anchor labels, synthetic labels, source-teacher training tokens, teacher-generation tokens, and target-training tokens as separate migration costs.
6e5697c5c0b12b9245a4fd192517c187a46d4148672693aac294fa11ef4461d0. Validation hash 2a14ba3b4241edd84a2c6dbfc5cb90425a95265363c395672bb52d1d4e01bc60.