Same model, different target
Action 0
Action 0
Your Frozen World Model Can Plan Better Than You Think
1Tsinghua University 2Manifold AI
†Equal contribution *Corresponding authors
Action 0
Action 0
Latent world models plan toward goal images with a frozen pretrained predictor, without task rewards or extra trained heads. However, their planners struggle with long-range goals, and prior work addresses this by training extra components such as value functions or subgoal models. We show that the planning target itself can cause this failure: even with exact dynamics and globally optimal short-horizon search, scoring predictions by their distance to the final goal rejects the first steps of a route that initially moves away from the goal. Building on this insight, we propose Anchored Planning (AP), a training-free method that reuses the world model's own offline trajectories. AP retrieves a segment that leads from the current observation toward the goal and aims the frozen planner at an observation shortly after the segment's start. Across four diverse tasks, AP substantially improves frozen LeWM planners for both action synthesis and action ranking, and it outperforms both additional final-goal search and the LeWM planner on long-range goals.
Paired rollouts share their start and goal. Every video shows its recorded outcome and action count.
Both planners use the same frozen model and five-action CEM search from standard starts. Only the scoring target changes.
Action 0
Action 0
Action 0
Action 0
Action 0
Action 0
Action 0
Action 0
Action 0
Action 0
Action 0
Action 0
Action 0
Action 0
Action 0
Action 0
Both planners rank the same eight retrieved action blocks from the same perturbed start. Final-goal ranking aims at the goal image. AP-rank aims at an observed intermediate target.
Action 0
Action 0
Action 0
Action 0
Action 0
Action 0
Action 0
Action 0
Action 0
Action 0
Action 0
Action 0
Action 0
Action 0
Action 0
Action 0
We show the first two queries per method and task where AP succeeds, the final-goal baseline fails, and AP takes more than ten actions. AP-CEM uses standard starts. AP-rank uses the first assigned perturbation. These selected examples illustrate behavior. The full evaluation below reports success rates over all queries.
A short search can reject a useful detour because it ends farther from the goal. This can happen even with an exact world model and globally optimal search. An observed target on a recorded route gives those first steps somewhere useful to aim.
The planner selects u₁.
Its predicted endpoint is closest to the final goal, so it receives the lowest cost.
Illustrative geometry. The dashed route comes from recorded experience, and the three candidate predictions stay fixed.
Anchored Planning retrieves a segment whose start resembles the current image and whose distant endpoint resembles the goal. The observation five steps after the retrieved start becomes the target.
A recorded observation gives the planner a concrete place to aim. Retrieval looks farther along the trajectory than the model needs to predict.
AP-CEM searches for a new action block. AP-rank chooses among recorded blocks. Both score predictions from the current state, and neither trains an additional model.
Changing only the target raises mean success on Cube, PushT, Reacher, and TwoRoom. The comparisons below use standard starts.
Final-goal target
Observed target
Final-goal target
Observed target
Mean success over four tasks. Each comparison keeps the frozen model and action rule fixed.
Success (%) over 128 queries per task. Perturbed results average the two assigned perturbations for each query.
As the goal moves from 25 to 100 actions away, the average gain from observed targets grows from 33.4 to 47.9 percentage points.
On every task, two AP-CEM iterations outperform thirty iterations of final-goal CEM. The improvement comes from changing the target, with fewer search iterations.
The experiments separate target choice, search effort, and prediction. The following comparisons examine what the target changes during control and which parts of memory matter.
We measure detours by an increase above a previous minimum in normalized physical error. Stalling means little task-relevant motion near the end of a failed rollout. On PushT, all 89 AP-CEM successes take a detour, while 100 of 125 final-goal CEM failures stall.
The learned target predicts the recorded successor more accurately on every task. Observed targets nevertheless achieve higher CEM success on PushT and Reacher. On Cube and TwoRoom, the learned target performs better.
| Task | Successor error ↓ | CEM success (%) ↑ | ||
|---|---|---|---|---|
| Learned | Observed | Learned | Observed | |
Mean squared distance to the recorded successor on held-out memory transitions, paired with control success at standard starts.
The target can lie one, three, five, or ten actions after the retrieved start, while prediction and execution still cover five actions. AP-CEM reaches its highest mean at a three-action target. AP-rank reaches its highest mean at the default five-action target.
| Offset | AP-CEM | AP-rank |
|---|
| Memory | AP-CEM | AP-rank |
|---|
As execution advances, AP shortens the span used to find a matching route. Holding it fixed at the original span reduces AP-rank success from 83.8% to 9.6% at standard starts.
Anchoring uses the observed endpoint as the target. Displacement transport adds the recorded movement to the current state instead. The two can differ when the current state is offset from the recorded source.
Under the same CEM search, anchoring improves mean success from 53.7% to 59.8% at standard starts and from 46.0% to 55.1% at perturbed starts. Reacher is the exception: displacement transport performs better there.
| Start | Transport | Anchored |
|---|
These experiments use frozen LeWM models on four simulated visual-control tasks. Learned intermediate targets also improve control, and LeWM leads at some shorter Reacher offsets. Testing other world-model families would show how far the gains extend.
Paper, implementation, and per-episode results are available below.
@misc{liu2026aimshort,
title = {Aim Short to Reach Far: Your Frozen World Model Can Plan Better Than You Think},
author = {Xvyuan Liu and Jianjie Fang and Wei Wu and Chen Gao and Yong Li},
year = {2026},
eprint = {2609.30036},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2609.30036}
}