Aim Short to Reach Far

Your Frozen World Model Can Plan Better Than You Think

Xvyuan Liu1,2,†Jianjie Fang1,†Wei Wu2Chen Gao1,*Yong Li1,*

1Tsinghua University 2Manifold AI
†Equal contribution *Corresponding authors

Same model, different target

Final-goal CEM

Action 0

AP-CEM

Action 0

Action 0

Abstract

Latent world models plan toward goal images with a frozen pretrained predictor, without task rewards or extra trained heads. However, their planners struggle with long-range goals, and prior work addresses this by training extra components such as value functions or subgoal models. We show that the planning target itself can cause this failure: even with exact dynamics and globally optimal short-horizon search, scoring predictions by their distance to the final goal rejects the first steps of a route that initially moves away from the goal. Building on this insight, we propose Anchored Planning (AP), a training-free method that reuses the world model's own offline trajectories. AP retrieves a segment that leads from the current observation toward the goal and aims the frozen planner at an observation shortly after the segment's start. Across four diverse tasks, AP substantially improves frozen LeWM planners for both action synthesis and action ranking, and it outperforms both additional final-goal search and the LeWM planner on long-range goals.

Four tasks. The same target comparison.

Paired rollouts share their start and goal. Every video shows its recorded outcome and action count.

Find actions with AP-CEM

Both planners use the same frozen model and five-action CEM search from standard starts. Only the scoring target changes.

PushT

Final-goal CEM

Action 0

AP-CEM

Action 0

Action 0

Cube

Final-goal CEM

Action 0

AP-CEM

Action 0

Action 0

Reacher

Final-goal CEM

Action 0

AP-CEM

Action 0

Action 0

TwoRoom

Final-goal CEM

Action 0

AP-CEM

Action 0

Action 0

More examples

PushT

Final-goal CEM

Action 0

AP-CEM

Action 0

Action 0

Cube

Final-goal CEM

Action 0

AP-CEM

Action 0

Action 0

Reacher

Final-goal CEM

Action 0

AP-CEM

Action 0

Action 0

TwoRoom

Final-goal CEM

Action 0

AP-CEM

Action 0

Action 0

Choose recorded actions with AP-rank

Both planners rank the same eight retrieved action blocks from the same perturbed start. Final-goal ranking aims at the goal image. AP-rank aims at an observed intermediate target.

PushT

Final-goal rank

Action 0

AP-rank

Action 0

Action 0

Cube

Final-goal rank

Action 0

AP-rank

Action 0

Action 0

Reacher

Final-goal rank

Action 0

AP-rank

Action 0

Action 0

TwoRoom

Final-goal rank

Action 0

AP-rank

Action 0

Action 0

More examples

PushT

Final-goal rank

Action 0

AP-rank

Action 0

Action 0

Cube

Final-goal rank

Action 0

AP-rank

Action 0

Action 0

Reacher

Final-goal rank

Action 0

AP-rank

Action 0

Action 0

TwoRoom

Final-goal rank

Action 0

AP-rank

Action 0

Action 0

We show the first two queries per method and task where AP succeeds, the final-goal baseline fails, and AP takes more than ten actions. AP-CEM uses standard starts. AP-rank uses the first assigned perturbation. These selected examples illustrate behavior. The full evaluation below reports success rates over all queries.

A good first step can lead away from the goal.

A short search can reject a useful detour because it ends farther from the goal. This can happen even with an exact world model and globally optimal search. An observed target on a recorded route gives those first steps somewhere useful to aim.

Switch the target. Keep the predictions.

The selected action changes when the target changesThree action endpoints stay fixed. The final goal favors action one. An observed target on a recorded route favors action two. Current state u₁ u₂ u₃ Goal Observed target

The planner selects u₁.

Its predicted endpoint is closest to the final goal, so it receives the lowest cost.

Illustrative geometry. The dashed route comes from recorded experience, and the three candidate predictions stay fixed.

Use a route that was actually taken.

Anchored Planning retrieves a segment whose start resembles the current image and whose distant endpoint resembles the goal. The observation five steps after the retrieved start becomes the target.

Match the current and goal images to the start and distant endpoint of a recorded segment. Use its five-step successor as the target for CEM or recorded-action ranking.

Memory supplies the target.

A recorded observation gives the planner a concrete place to aim. Retrieval looks farther along the trajectory than the model needs to predict.

The model evaluates the actions.

AP-CEM searches for a new action block. AP-rank chooses among recorded blocks. Both score predictions from the current state, and neither trains an additional model.

What changes across the full evaluation?

Changing only the target raises mean success on Cube, PushT, Reacher, and TwoRoom. The comparisons below use standard starts.

Action synthesis with CEM

9.2%

Final-goal target

59.8%

Observed target

Ranking recorded actions

67.0%

Final-goal target

83.8%

Observed target

Mean success over four tasks. Each comparison keeps the frozen model and action rule fixed.

Every task, both starting conditions.

Success (%) over 128 queries per task. Perturbed results average the two assigned perturbations for each query.

Success rates by task and planning rule

The gain grows with goal distance.

As the goal moves from 25 to 100 actions away, the average gain from observed targets grows from 33.4 to 47.9 percentage points.

Mean success at goal offsets of 25, 50 and 100 actions. The observed-target advantage grows from 33.4 to 47.9 percentage points.

More search does not close the gap.

On every task, two AP-CEM iterations outperform thirty iterations of final-goal CEM. The improvement comes from changing the target, with fewer search iterations.

Observed-target CEM with two iterations already beats final-goal CEM with thirty. At thirty iterations, their mean success rates are 59.8 and 9.2 percent.

What makes an observed target useful?

The experiments separate target choice, search effort, and prediction. The following comparisons examine what the target changes during control and which parts of memory matter.

Successful routes can begin with a detour.

We measure detours by an increase above a previous minimum in normalized physical error. Stalling means little task-relevant motion near the end of a failed rollout. On PushT, all 89 AP-CEM successes take a detour, while 100 of 125 final-goal CEM failures stall.

PushT error traces, stalling among failures, and detours among successes. Reacher has zero stalling under both plotted methods.
Behavior at standard starts. Reacher’s zero stalling rate means its failed rollouts do not become stationary under the study’s definition.

Lower target-prediction error does not always mean higher control success.

The learned target predicts the recorded successor more accurately on every task. Observed targets nevertheless achieve higher CEM success on PushT and Reacher. On Cube and TwoRoom, the learned target performs better.

Successor error and CEM success at standard starts
TaskSuccessor error ↓CEM success (%) ↑
LearnedObservedLearnedObserved

Mean squared distance to the recorded successor on held-out memory transitions, paired with control success at standard starts.

Target placement and retrieval timing both matter.

The target can lie one, three, five, or ten actions after the retrieved start, while prediction and execution still cover five actions. AP-CEM reaches its highest mean at a three-action target. AP-rank reaches its highest mean at the default five-action target.

Mean success over four tasks

How far ahead is the target?

OffsetAP-CEMAP-rank

How much memory is used?

MemoryAP-CEMAP-rank

Keep the retrieval span moving.

As execution advances, AP shortens the span used to find a matching route. Holding it fixed at the original span reduces AP-rank success from 83.8% to 9.6% at standard starts.

Aim at the recorded endpoint.

Anchoring uses the observed endpoint as the target. Displacement transport adds the recorded movement to the current state instead. The two can differ when the current state is offset from the recorded source.

Under the same CEM search, anchoring improves mean success from 53.7% to 59.8% at standard starts and from 46.0% to 55.1% at perturbed starts. Reacher is the exception: displacement transport performs better there.

Mean CEM success (%)
StartTransportAnchored

What the evidence covers

These experiments use frozen LeWM models on four simulated visual-control tasks. Learned intermediate targets also improve control, and LeWM leads at some shorter Reacher offsets. Testing other world-model families would show how far the gains extend.

Citation

Paper, implementation, and per-episode results are available below.

@misc{liu2026aimshort,
  title  = {Aim Short to Reach Far: Your Frozen World Model Can Plan Better Than You Think},
  author = {Xvyuan Liu and Jianjie Fang and Wei Wu and Chen Gao and Yong Li},
  year   = {2026},
  eprint = {2609.30036},
  archivePrefix = {arXiv},
  primaryClass = {cs.LG},
  url    = {https://arxiv.org/abs/2609.30036}
}