WM-Craftnet: World Synesthesia Model for Robust and Generalizable Dexterous In-Hand Manipulation

CoRL 2026

We will release the codes soon!
supplementary_ablation

World Synesthesia Model for Visuotactile Dexterity

An action-conditioned recurrent state that turns depth, touch, proprioception, and control history into a reusable prior for perturbation-robust, object-ID-free in-hand manipulation.

TL;DR

WM-Craftnet learns a Dreamer-style World Synesthesia Model from proprioception, wrist depth, tactile signals, actions, and rewards, then exposes its deterministic recurrent feature to a PPO actor. Rather than optimizing through imagined rollouts, the world model acts as a deployable visuotactile state estimator, helping one policy infer object geometry, contact evolution, drift, and slip under pose shifts, disturbances, unseen objects, different rotation axes, and tool-use-style translation.

WM-Craftnet

Pipeline

Overview of WM-Craftnet pipeline
Figure: Overview of WM-Craftnet. A Dreamer-style RSSM reconstructs depth and low-dimensional proprio-tactile state, predicts reward during world-model training, and provides its deterministic recurrent feature to the manipulation policy.

Ablation Evidence: Why WSM Helps

These compact ablations isolate the mechanism behind WM-Craftnet: the performance gain comes from clean geometric supervision coupled with action-conditioned recurrent inference, not from simply adding raw depth, tactile inputs, or a per-frame learned feature. More detailed rollout and recovery videos are presented in the real-world, multi-axis, and robustness sections below.

Best z-axis WSM 753.3 Return

Full WSM with prop+tac+clean-depth heads improves over raw-sensor baselines and scratch training.

Key controls +45.3 Return

Clean-depth WSM beats noisy-depth supervision, showing that denoising is a task-relevant target.

Real robot 16.18 rad vs. 2.76 rad

On duck rotation, WM-Craftnet outperforms IHR even when IHR receives WSM-denoised depth.

Latent structure Object-aware ht

t-SNE shows recurrent states organize interaction regimes without object identifiers.

Sensor-head ablation: clean depth and tactile prediction make the WSM state strongest
Method Return ↑ EpLen ↑ RotR ↑ Fall ↓
Best raw-sensor baseline 386.9 362.5 1.018 0.047
WM-Craftnet from scratch 414.3 264.2 0.742 0.002
WSM pretrained, prop-only 684.5 424.0 1.191 0.009
WSM pretrained, prop+tac 698.9 435.4 1.192 0.004
WSM pretrained, prop+tac+depth 753.3 435.2 1.293 0.005

Conclusion. The full predictive state is not just a larger observation vector: pretraining and clean-depth supervision give the actor a reusable, denoised interaction state for rotation.

Training reward curves and real robot z-axis evaluation
Training curves from the main paper: WM-Craftnet stays above the baseline policies, with the WSM ablation curve isolating the representation gain.
Latent-state diagnostic: WSM organizes object-dependent interaction regimes
t-SNE visualization of WSM recurrent states
t-SNE of WSM recurrent states ht collected from z-axis rollouts over nine objects. Clusters and overlaps reflect object geometry and shared contact affordances, even though the deployable actor does not receive object IDs.

Conclusion. The recurrent WSM state is not a generic visual embedding; it separates and blends interaction histories according to geometry/contact regimes that matter for control.

Controlled WSM variants: clean-depth supervision and recurrence are both needed
WSM Variant EpLen ↑ Return ↑ RotR ↑ OffAxis ↓ AngVar ↓
Noisy-depth supervision 431.1 708.0 1.265 1.299 1.456
No previous h recurrent input 434.3 705.4 1.282 1.125 1.185
Encoder-decoder z policy 417.4 667.6 1.254 1.323 1.343
Full WSM 435.2 753.3 1.293 1.225 1.324
Supplementary WSM ablation training curves
The supplementary control curves show the full WSM learning faster and reaching higher final reward than each isolated variant.

Conclusion. A per-frame latent feature is insufficient; the best control performance comes from combining clean geometric targets with action-conditioned recurrent memory.

Real-world clear-depth control: denoising helps, but does not fully explain WM-Craftnet
Method Duck RR/SR Cross RR/SR Corner RR/SR Unseen RR/SR
In-Hand Rotation 1.83 / 5/10 0.45 / 1/10 0.36 / 1/10 0.39 / 0/10
IHR with WSM-denoised depth 2.76 / 8/10 1.21 / 8/10 2.43 / 10/10 0.80 / 0/10
WM-Craftnet 16.18 / 10/10 8.01 / 10/10 8.48 / 10/10 4.32 / 8/10

Conclusion. Cleaner depth substantially improves IHR, but WM-Craftnet remains far ahead because its WSM state also carries temporal contact and motion information needed for closed-loop adjustment.

Reusable prior: the WSM state transfers to 49 new objects
49-object downstream rotation with a WSM pretrained on the nine-object z-axis setting.

Conclusion. The pretrained WSM serves as reusable physical structure: after 3000 epochs, the downstream 49-object policy reaches 9.37 ± 0.13 rad per episode versus 3.28 rad without the prior, while the fall rate drops from 6% to 0.3%. The video shows the resulting broad object coverage.

WSM Latent Depth Reconstruction

To visualize what the World Synesthesia Model learns, we reconstruct wrist-camera depth from the WSM latent state during duck z-axis rotation rollouts. Although the model receives a noisy depth input, the reconstructed predicted depth suppresses sensor noise and preserves hand-object geometry in both real and simulated rollouts, supporting the role of WSM as a deployable visuotactile state estimator.

Real Hardware

Real Noisy Wrist Depth

Real Predicted Depth from WSM Latent

On real hardware, the wrist depth stream is noisy and incomplete. The WSM latent reconstructs a cleaner depth representation that preserves the hand-object geometry.

Simulation

Third-Person RGB: duck z-axis rotation

Ground-Truth Wrist Depth

Noisy Depth Input to WSM

Predicted Depth from WSM Latent

The WSM latent reconstructs a clean wrist-depth stream from noisy observations. Compared with the noisy input, the predicted depth preserves the duck and hand-object workspace while removing substantial sensor noise, showing that the recurrent latent captures task-relevant geometry for downstream control.

One Policy for Multi-Object Rotation

WM-Craftnet evaluates whether one closed-loop policy can rotate objects with different size, shape, mass, curvature, and contact geometry without object identifiers or an object-specific finger gait.

Sim-to-Real Transfer

On the real Sharpa platform, WM-Craftnet is deployed with wrist depth sensing and tactile feedback under real sensing noise, latency, and unmodeled contacts. The real-hardware rollouts below show that one policy can adapt across objects in a continuous run, recover from out-of-distribution hand-object states, and sustain z-axis rotation on objects with different shapes and contact geometry.

Qualitative Generalizability and Robustness Demos

Continuous multi-object real-world rollout. The same policy rotates several objects in one uninterrupted run: 0-24 s duck, 24-35 s cylinder, 35-50 s cross block, 50-73 s steamed bun, and 73-94 s an unseen double-notched block. In the unseen-object segment, the policy initially struggles, adapts online through closed-loop sensing, and gradually reaches stable rotation.

OOD-to-duck recovery. An unseen object first drives the policy into an out-of-distribution interaction state. After switching to the duck, WM-Craftnet recovers a controllable grasp and resumes successful z-axis rotation on the physical hand.

Real-world perturbation recovery. 3-5 s and 22-25 s: a human perturbs the corner block's position; the policy adjusts the object back into a controllable state and continues smooth rotation.

Challenging-start recovery. At 0 s, the switch starts in a side-fall state; the policy recovers from the challenging initial pose and brings the object back into a controllable rotation mode.

Bucket palm-region recovery. During bucket rotation, at 6 s the object pops into the palm region; WM-Craftnet pushes it back toward the comfortable workspace and continues stable rotation.

Cylinder unexpected-contact recovery. Around 6 s and 15 s, unexpected contacts push the cylinder toward the palm; the policy adjusts it back into the comfortable region and keeps the rotation going.

Real-World Z-Axis Rotation across Objects

Duck. WM-Craftnet keeps stable finger contacts while rotating a curved, asymmetric object.

Corner block. WM-Craftnet maintains stable z-axis rotation for over one minute without side-fall.

Cross block. Closed-loop visuotactile feedback supports rotation with protruding arms and changing contact patches.

Cylinder. The hand produces smooth z-axis spin on a rounded object with limited tactile landmarks.

Piggy bank. WM-Craftnet preserves controllable rotation on an anisotropic everyday object with direction-dependent contact geometry.

Strawberry. The policy maintains rotation on an anisotropic tapered object whose irregular profile creates shifting contact patches.

Fire hydrant. WM-Craftnet handles protruding side structures and changing contact geometry while sustaining z-axis rotation.

Blue block. The policy keeps the block upright and achieves stable continuous rotation for more than six full turns.

Bulb. Despite an unstable center of mass, closed-loop sensing lets the policy adjust contacts and recover balanced rotation.

Baseline Comparison in Real World

Open-loop Replay

open-loop

The replayed gait has no feedback or adjustment ability; even with the block near the palm center, it cannot recover a controllable rotation mode.

Touch Dexterity

tactile-only feedback

The tactile policy can produce rotation, but without explicit geometry the object becomes unstable and frequently tips into side-fall states.

In-Hand Rotation

tactile+noisy real depth

Real sensor noise quickly pushes the depth observation out of the simulated training distribution, causing an OOD interaction state and loss of stable rotation.

IHR + WSM Depth

tactile+denoised depth

WSM-reconstructed depth gives a cleaner geometry estimate and slightly improves behavior, but depth denoising alone cannot sustain stable rotation.

Ours

WM-Craftnet

With WSM recurrent task context, the policy keeps the corner block upright and maintains stable continuous z-axis rotation for over one minute.

Tests in Simulation

In simulation, initial object positions are uniformly perturbed within an 18 cm2 feasible reset region. The main z-axis benchmark rotates nine everyday objects: corner block, duck, apple, piggy bank, strawberry, cross block, coca can, bulb, and double-notched cylinder.

Corner block

Duck

Apple

Piggy Bank

Strawberry

Cross block

Coca Can

Bulb

Double-notched Cylinder

Gravity-Invariant Rotation

For gravity-invariant rotation, we adapt WM-Craftnet with a closer hand-object grasp inspired by the HORA implementation, together with modified initial conditions and reward design. This compact grasp keeps the object more securely inside the hand, so the policy can maintain contact and stable rotation when gravity acts under different palm orientations.

Compact-grasp rotation. A single policy rotates a duck, a strawberry, and a cylinder using the closer hand posture.

Successful sim-to-real transfer. The object remains grasped and keeps rotating as palm orientation changes under gravity.

Rotation around Diverse Axis

Beyond the z-axis benchmark, WM-Craftnet is trained and evaluated for axis-specific rotation around x and y. Both settings use the same 18 cm2 randomized initial-position protocol to test whether the representation remains effective under different commanded rotation goals.

These axes are not just alternative target directions; they create harder contact-rich motion modes. The z-axis setting is comparatively stable because gravity helps keep the object seated in the palm while the fingers accumulate spin around the vertical direction. In contrast, y-axis rotation often involves elongated or tool-like objects rotating around a shorter object direction, not along the long body axis. This puts mass and contact points farther from the commanded axis, requiring control of larger moment arms, asymmetric contact patches, and sideward torques. The x-axis setting is even more challenging because rotation relies on rolling and fingertip-contact reallocation, while gravity can drive the object out of the stable palm-supported region.

Sim-to-Real Transfer

We first validate that the learned y-axis controller transfers from simulation to the real hand on elongated and edge-rich objects. These real-world rollouts highlight whether WM-Craftnet can preserve closed-loop target-axis rotation while coping with sensing noise, contact asymmetry, and tool-like geometries outside the simulator.

Toothpaste. The policy keeps an elongated object controllable while rotating around its shorter axis under asymmetric leverage.

Screwdriver. Closed-loop visuotactile feedback stabilizes a tool-like object and maintains y-axis rotation despite the handle-shaft geometry.

Corner block. WM-Craftnet sustains real-world y-axis rotation despite sharp edges, intermittent contacts, and changing sideward torques.

Tests in Simulation

In simulation, we next evaluate both alternative rotation goals under the same randomized initialization protocol. We first present the broader y-axis benchmark and then the more contact-sensitive x-axis benchmark.

Y-Axis Rotation

The y-axis object set contains toy trashcan, corner block, bulb, corner cylinder, toothpaste, coca can, screwdriver, flashlight, and bottle. This setting emphasizes whether the policy can maintain sustained target-axis rotation while rotating elongated bodies about a shorter axis and regulating the resulting moment arms and sideward contact torques.

Toy Trashcan

Corner block

Bulb

Corner Cylinder

Toothpaste

Coca Can

Screwdriver

Flashlight

Bottle

X-Axis Rotation

The x-axis rollouts below show four representative objects: block, corner block, stepped block, and lego brick. Here the target axis changes the required contact sequence, gravity interaction, and feasible finger-coordination pattern, making stable rotation harder than simply replaying a z-axis finger gait.

block

Corner block

Stepped block

Lego Brick

Application: Tool Use

Finally, we examine whether WM-Craftnet can support tool-use-style manipulation beyond benchmark object rotation. Using a screwdriver as an example, the policy must reason about the tool's elongated geometry and contact state while preserving a stable in-hand grasp. This setting includes both goal-conditioned translation, where the screwdriver is adjusted to a target pose, and fast axial rotation, where the hand spins the tool while maintaining control.

Goal-conditioned translation. The hand adjusts the screwdriver toward a target position while keeping stable contact along the handle.

Fast screwdriver rotation. The same tool-use setup demonstrates rapid in-hand rotation while preventing the tool from slipping out of the grasp.

Failure Case Analysis and Limitations

The central failure case for dexterous in-hand rotation is an overfit, nearly open-loop finger gait. Policies trained on a narrow object set or fixed initial pose can rotate familiar objects, but small pose offsets, force disturbances, or object drift may push the interaction outside the trained distribution and lead to stuck states or drops.

WM-Craftnet addresses this by using a learned recurrent state to improve task-relevant inference over object pose, geometry, contact, and slip. However, the current quantitative evaluation is still focused on short-horizon in-hand rotation. Translation, tool-use-like behavior, grasp changes, and broader real-world sensing variation are evaluated qualitatively or left as future work.

Representative Failure Cases

The examples below separate Baseline failures from Ours failures on both real hardware and simulation. Baseline policies can quickly leave the observation distribution under noisy depth and fail to produce useful rotation actions. WM-Craftnet reduces these issues, but can still fail under especially challenging initial poses or long-horizon contact drift.

Real Hardware

Baseline

Block, real z-axis

Noisy real depth rapidly drives the baseline out of distribution, preventing it from producing a coherent rotation strategy.

Ours

Corner block, challenging start

WM-Craftnet attempts to adjust from a difficult initial pose, but cannot recover a stable controllable configuration.

Ours

Duck, long-horizon drift

After completing 5 successful spins, the duck tips toward one side and the policy loses the stable contact mode needed to continue at 29s.

Simulation

Baseline

Bulb, y-axis

The bulb starts on the palm; during adjustment, the policy pushes it out of the hand and drops it.

Baseline

Bulb, z-axis

The object rotates, but the hand uses an unnatural posture and the bulb drifts far from the intended rotation axis.

Baseline

Duck, z-axis

The baseline can induce rotation, but with an awkward finger gait and substantial off-axis object displacement.

Baseline

Notched block, x-axis

During x-axis rotation of the notched block, contact becomes unstable and the object drops midway through the rollout.

Baseline

block, x-axis

The baseline struggles to keep a stable axis-aligned grasp, causing the block to drift and lose controllable rotation.

Baseline

Flashlight, y-axis

The flashlight enters a stuck contact state, preventing the policy from continuing useful in-hand rotation.

Baseline

Toothpaste, y-axis

While rotating the toothpaste-like object, the policy fails to explore a useful contact mode and cannot complete rotation.

Ours

Bulb, z-axis

WM-Craftnet can still fail on difficult bulb manipulation when contact recovery is insufficient and the object falls.

BibTeX

@inproceedings{yin2026wmcraftnet,
  author    = {Jie Yin and Zeyuan Zhao and Xiaojing Tan and Yang Liu and Chiyu Wang and Xinyang Gu},
  title     = {WM-Craftnet: World Synesthesia Model for Robust and Generalizable Dexterous In-Hand Manipulation},
  booktitle = {Conference on Robot Learning (CoRL)},
  year      = {2026},
}