SkillWeave: Weaving Heterogeneous Demonstrations into Long-Horizon Manipulation Skills

Anonymous
Under review

Overview video placeholder

static/videos/teaser.mp4

SkillWeave uses teleoperation for long-horizon reaching and transport, kinesthetic teaching for precise contact-rich skills, and steers the handoff between the resulting policies.

Abstract

Dexterous manipulation requires both large-scale task progression and precise contact-rich interaction, making it challenging to collect demonstrations that effectively support both regimes. We present SkillWeave, a heterogeneous demonstration framework for long-horizon dexterous manipulation that combines teleoperation for coarse reaching and transport with kinesthetic teaching for precise, contact-rich skills.

To address the visual mismatch introduced by the demonstrator's presence during kinesthetic data collection, we propose an object-mask-conditioned diffusion policy that uses offline object segmentation for training supervision and a lightweight learned mask predictor at deployment, avoiding online segmentation and image inpainting. To mitigate distribution shift between independently trained sub-task policies, we introduce successor-aware terminal steering, which selects among actions sampled from the predecessor policy to guide the system toward states supported by the successor's demonstrated initial-state distribution.

Across three real-world long-horizon tasks, SkillWeave achieves 27% average end-to-end success. Mask-conditioned kinesthetic policies improve dexterous sub-task success to an average of 65%, while successor-aware handoffs achieve an average composition efficiency of 87%. These results show that matching demonstration modality to interaction regime, explicitly addressing kinesthetic visual mismatch, and steering policy handoffs toward successor-supported states substantially improves long-horizon dexterous manipulation.

Overview

Teleoperated reaching frames (blue), a policy-composition panel steering predecessor reachable states into successor start states, and kinesthetic teaching frames (red).

Teleoperation is used for long-horizon coverage and free-space reaching (blue), kinesthetic teaching for precise contact-rich interaction (red). Policy composition with inference-time steering combines policies trained on individual sub-tasks into long-horizon task success.

Contributions

  • Multi-modal demonstration framework. Teleoperation for long-horizon coverage and free-space reaching, kinesthetic teaching for precise contact-rich interaction. Collecting each sub-task separately also makes data cheaper to gather: a failed attempt costs one sub-task rather than the whole episode.
  • Successor-aware terminal steering. Guides predecessor policies toward the successor's start-state distribution, reducing distribution shift at policy boundaries without additional transition demonstrations or a separately trained transition policy.
  • Object-mask-conditioned diffusion policy. SAM 3 provides offline training masks while a learned predictor supplies masks at deployment, improving robustness across backgrounds and scenes with no online segmentation or inpainting in the control loop.

Method

Mask-conditioned kinesthetic policies

Each sub-task policy is a diffusion policy conditioned on an external RGB view, a wrist RGB view and proprioception (end-effector pose, arm joint angles, hand joint angles). Both views are encoded by a ResNet-18; the visual features are concatenated with proprioception to condition a 1-D conditional U-Net that denoises an action chunk.

Kinesthetic demonstrations put the demonstrator's hand and arm in frame during collection, but they are absent during autonomous execution. Rather than inpainting the demonstrator away, we keep the original RGB observations and condition visual feature aggregation on a mask of the manipulated object: features pooled inside the mask form a query, and cross-attention pools the spatial feature map against it. The attention distribution is additionally supervised by the normalized object mask. During training, masks come from SAM 3 offline; at deployment, a lightweight per-view U-Net predicts them directly from RGB, so neither SAM 3 nor inpainting runs online.

Mask-conditioned diffusion policy architecture.

Mask-conditioned diffusion policy. Object masks guide cross-attention pooling over the ResNet-18 feature maps; pooled features plus proprioception condition a 1-D U-Net that predicts action chunks.

Video placeholder

static/videos/mask-conditioned-rollout.mp4

Autonomous rollout with no demonstrator in frame.

Successor-aware terminal steering

Independently trained policies do not automatically compose: small prediction errors accumulate, so the predecessor's terminal state drifts away from the states the successor was trained to start from. We leave both policies unchanged and instead modify how the predecessor finishes.

For each successor policy, the first states of its training demonstrations form an empirical start-state set (wrist pose, arm joints, hand joints, each normalized). A state's compatibility is its mean distance to its K nearest neighbours in that set. Once the sub-task enters its terminal phase, the predecessor samples L candidate action chunks per replanning step, predicts each endpoint, and executes the candidate with the smallest distance, repeating until the score drops below a threshold or a maximum steering horizon is reached. Because every candidate is sampled from the predecessor itself, the system stays within behaviour the predecessor supports while moving toward states familiar to the successor.

Predecessor policy, successor-aware terminal steering (sample, score, select), and handoff inside the successor start-state set.

Policy composition and handoff. Candidate action chunks sampled from the predecessor are scored by proximity to the successor's demonstrated start states; the best candidate is executed, closed-loop, until handoff.

Video placeholder

static/videos/steering-without.mp4

Without steering: the predecessor ends far from the successor's start states.

Video placeholder

static/videos/steering-with.mp4

With steering: the handoff lands inside the successor's start-state region.

Terminal hand pose without steering, an example successor start state, and the terminal hand pose with steering.

Without terminal steering (left) the predecessor ends at finger joint states far from the successor's start states (middle); with terminal steering (right) it ends much closer.

Demonstration Collection

Demonstrations are collected on a 7-DoF xArm7 with a 16-DoF LEAP Hand, observed by a third-person RGB camera and a wrist-mounted RGB camera facing the palm. For each sub-task we collect 100 demonstrations with the modality that fits it. During teleoperation the operator drives the arm with a 6-DoF SpaceMouse while finger motion is captured by a camera, tracked with MediaPipe and retargeted to the robot hand by joint-angle mapping. During kinesthetic teaching the arm runs in drag/teach mode while the operator positions the fingers by hand under a compliant controller that yields to manual adjustment and holds the resulting pose. Synchronized RGB observations and proprioception (6D end-effector pose, 7D arm joints, 16D hand joints) are recorded at 30 Hz, with object positions randomized across episodes.

Video placeholder

static/videos/data-teleoperation.mp4

Teleoperation: SpaceMouse + MediaPipe hand retargeting.

Video placeholder

static/videos/data-kinesthetic.mp4

Kinesthetic teaching: drag/teach arm, compliant finger posing.

Example RGB observations: teleoperation, KineDex inpainting, kinesthetic teaching, and the wrist view during kinesthetic teaching.

Example RGB observations. Raw teleoperation and kinesthetic observations are fed directly to the policies; KineDex inpaints the image to match inference-time observations, but often blurs the robot hand and hallucinates content.

Long-Horizon Tasks

Each task is decomposed into sub-tasks, each demonstrated with the modality suited to its interaction regime: Teleoperation for arm-level reaching and placement, Kinesthetic for finger-level, contact-rich skills.

Task 1: Cube Reorientation

Video placeholder

static/videos/task1-cube-full.mp4

Full rollout End-to-end SkillWeave execution.

Video placeholder

static/videos/task1-sub1-reach-grasp-reorient.mp4

Teleoperation Reach, grasp and reorient cube. Success: cube securely grasped without dropping.

Video placeholder

static/videos/task1-sub2-in-hand-rotation.mp4

Kinesthetic Rotate cube in-hand. Success: cube rotated by at least 90°.

Task 2: Nut Removal and Storage

Video placeholder

static/videos/task2-nut-full.mp4

Full rollout End-to-end SkillWeave execution.

Video placeholder

static/videos/task2-sub1-reach-nut.mp4

Teleoperation Reach for nut. Success: hand positioned over the nut.

Video placeholder

static/videos/task2-sub2-unscrew-nut.mp4

Kinesthetic Unscrew nut. Success: nut fully detached from screw.

Video placeholder

static/videos/task2-sub3-place-nut.mp4

Teleoperation Place nut in container. Success: nut released inside the container.

Task 3: Kettle Preparation

Video placeholder

static/videos/task3-kettle-full.mp4

Full rollout End-to-end SkillWeave execution.

Video placeholder

static/videos/task3-sub1-reach-grasp-position.mp4

Teleoperation Reach, grasp and position kettle. Success: kettle grasped by the handle and positioned under the faucet.

Video placeholder

static/videos/task3-sub2-open-lid.mp4

Kinesthetic Open lid using release button. Success: button pressed and lid opened.

Video placeholder

static/videos/task3-sub3-place-kettle.mp4

Teleoperation Place kettle down. Success: kettle released and stable on the table.

Frame strips for the three long-horizon tasks; blue borders mark teleoperation-trained policies, red borders mark mask-conditioned kinesthetic policies.

Long-horizon rollouts. Top to bottom: Cube Reorientation, Nut Removal and Storage, Kettle Preparation. Blue frames are policies trained with teleoperation; red frames are mask-conditioned kinesthetic policies.

Results

Every policy is evaluated over 20 rollouts with randomized object configurations, autonomously and with no demonstrator in frame.

Baseline comparison

Video placeholder

static/videos/baseline-teleop-only.mp4

Teleoperation only — fails at the contact-rich stage.

Video placeholder

static/videos/baseline-kinesthetic-only.mp4

Kinesthetic only — visual shift breaks large-workspace motion.

Video placeholder

static/videos/baseline-kinedex.mp4

KineDex — inpainting artifacts near contact regions.

Video placeholder

static/videos/baseline-skillweave.mp4

SkillWeave — modality matched per sub-task, handoffs steered.

Sub-task success

Sub-task Teleop. Kin. KineDex Mask-Kin. (ours)
Reach, grasp & reorient70%0%0%—
In-hand rotation10%30%0%70%
Reach for nut85%0%40%—
Unscrew nut5%0%60%75%
Place nut in container30%0%5%—
Reach, grasp & position kettle80%0%0%—
Open lid with button0%0%0%50%
Place kettle down55%0%0%—

Individual sub-task policies across demonstration and visual-processing strategies; the rules separate the three tasks (top to bottom: Cube Reorientation, Nut Removal & Storage, Kettle Preparation). Mask-conditioned policies were trained only for the dexterous sub-tasks.

End-to-end long-horizon success

Task Teleop. only Kin. only KineDex SkillWeave w/o steering SkillWeave
Cube Reorientation0%0%0%10%45%
Nut Removal & Storage0%0%0%0%15%
Kettle Preparation0%0%0%5%20%
Average0%0%0%5%27%

Composition efficiency

Task Teleop. S / S* / C KineDex S / S* / C w/o steering S / S* / C SkillWeave S / S* / C
Cube Reorientation0 / 7 / 0%0 / 0 / –10 / 49 / 20%45 / 49 / 92%
Nut Removal & Storage0 / 1 / 0%0 / 1 / 0%0 / 19 / 0%15 / 19 / 79%
Kettle Preparation0 / 0 / –0 / 0 / –5 / 22 / 23%20 / 22 / 91%
Average0 / 3 / 0%0 / 1 / 0%5 / 30 / 14%27 / 30 / 87%

S is end-to-end success, S* = ∏j S(j) is the success expected from isolated sub-task performance, and C = S / S* is composition efficiency. Values near 1 mean policy handoffs add little degradation beyond the sub-tasks' own failure rates. Kinesthetic-only results are omitted because S* = 0 for all tasks.

  • Teleoperation averages 64% on arm-level reaching, reorientation and placement sub-tasks but only 5% on the dexterous ones, where retargeting and indirect contact feedback break down.
  • On dexterous sub-tasks, raw kinesthetic observations reach 10% and KineDex's segment-and-inpaint reaches 20%, while mask conditioning reaches 65%: explicit object-centric conditioning beats reconstructing occluded pixels.
  • The same sub-task policies composed naively average 5% end-to-end; with successor-aware terminal steering they average 27%. Strong isolated sub-tasks do not imply reliable long-horizon execution.