DreamLake

Subtask Annotation

A worked example: the Macrodata Labs subtask-annotation pipeline — which turns long robot and egocentric manipulation videos into timestamped subtask annotations — traced into a PipelineGraph. It's a real-world instance of the AI auto-label → review → sink loop the Pipeline Graph page describes, and a good stress test for the layout: two VLM stages, a fan-out of contact sheets, a review gate, and a rework loop.

The whole thing is two VLM passes over timestamped contact sheets (frame grids with the time drawn on each tile) using Gemini 3.5 Flash: Part I segments each video into subtask spans; Part II labels every span. Headline numbers from the post: 0.306 segmentation F1 and 61.0% labeling accuracy at $2.64 per hour of video on batch pricing — roughly 19× cheaper than human annotation.

videos
framestimestamps
sheets
instructions
spans
frames
windows
instructions
labels
rows
rows
rows
rows
load_videos
source · 0→1
videos_to_frames
transform · 1→1
frames_to_sheets
transform · 2→1
sheets_to_segments
model · 2→1
frames_to_windows
transform · 2→1
windows_to_labels
model · 2→1
score_segments
review · 1→1
save_dataset
sink · 1→0
rework
sink · 1→0
This graph is a hand-authored trace

The reference pipeline ships in Macrodata's Refiner framework. The JSON here is authored to match what the dl_trace tracer would emit for the equivalent @ls.udf stages — same node/edge shape as the other demos — so the structure is faithful even though the tracer wasn't run against the original source.

How to read it

Left to right, the DAG is the pipeline's dataflow. The two model nodes (sheets_to_segments, windows_to_labels) are the Gemini calls; score_segments is the review gate (its output is a mask, drawn as dashed edges); the two sinks split accepted annotations from spans sent back for rework.

  • load_videos (source) — a batch of robot / ego clips, each with its high-level task instruction.
  • videos_to_frames — sample every video at 2 fps (a frame every 0.5 s), keeping per-frame timestamps.
  • frames_to_sheets — tile 20 frames into a 4×5 contact sheet at 224 px, with the timestamp drawn onto each tile.
  • sheets_to_segments (Part I — Gemini) — read the sheets and emit subtask spans (start, end).
  • frames_to_windows — for each span, build a previous / current / next window of contact sheets, 5 frames each.
  • windows_to_labels (Part II — Gemini) — label each span from its window.
  • score_segments (review) — WGO-Bench gate on span F1 + label accuracy.
  • save_dataset / rework (sinks) — accepted spans ship; short, low-recall ones requeue.

Click any stage below to read the exact @ls.udf it traced from — the graph and the source rail share one selection.

videos
framestimestamps
sheets
instructions
spans
frames
windows
instructions
labels
rows
rows
rows
rows
load_videos
source · 0→1
videos_to_frames
transform · 1→1
frames_to_sheets
transform · 2→1
sheets_to_segments
model · 2→1
frames_to_windows
transform · 2→1
windows_to_labels
model · 2→1
score_segments
review · 1→1
save_dataset
sink · 1→0
rework
sink · 1→0
subtask_annotation
9idle0running0waiting0ok0error0stale
9 nodes · 14 edges
"""Annotate robot-video subtasks — segment, then label.
Two VLM stages over timestamped contact sheets (Gemini 3.5 Flash):
Part I segments each video into subtask spans from full-video contact
sheets; Part II labels every span from a prev/current/next window of
per-segment sheets (seeded relabeling). A WGO-Bench review gates what
ships; short, low-recall spans go back for rework. EVERY function the
pipeline calls is a `@ls.udf` — the source and the two sinks included.
"""
import dreamlake as dl
import lakeshore as ls
from dreamlake import batch, requeue, to_dataset
from lakeshore.types import Mask, String, Tensor, Tuple
@ls.udf(kind="source")
def load_videos() -> Tuple["videos", "instructions"]:
"""Pull the batch of robot / ego videos with their high-level task instruction."""
...
@ls.udf
def videos_to_frames(videos: Tensor["N"]) -> Tuple["frames", "timestamps"]:
"""Sample each video at 2 fps (every 0.5 s); keep the per-frame timestamps."""
...
@ls.udf
def frames_to_sheets(frames, timestamps) -> Tuple["sheets"]:
"""Tile frames into timestamped contact sheets — 20 frames, 4x5, 224 px, times drawn on each frame."""
...
@ls.udf(kind="model")
def sheets_to_segments(sheets, instructions: String["N"]) -> Tuple["spans"]:
"""Gemini 3.5 Flash reads the sheets and emits subtask spans (start, end). GEPA-tuned prompt."""
...
@ls.udf
def frames_to_windows(frames, spans) -> Tuple["windows"]:
"""Per span, build a prev / current / next window of contact sheets — 5 frames each, uniform."""
...
@ls.udf(kind="model")
def windows_to_labels(windows, instructions: String["N"]) -> Tuple["label", "confidence"]:
"""Gemini labels each span from its window (seeded relabeling off the segmentation prior)."""
...
@ls.udf(kind="review")
def score_segments(labels) -> Mask["N"]:
"""WGO-Bench review: F1 on the spans + label accuracy; pass the ones that clear threshold."""
...
@ls.udf(kind="sink")
def save_dataset(rows):
"""Write the accepted subtask annotations to the dataset (wraps dreamlake.to_dataset)."""
to_dataset(rows)
@ls.udf(kind="sink")
def rework(rows):
"""Send short (<2 s) / low-recall spans back for rework (wraps dreamlake.requeue)."""
requeue(rows)
@dl.pipeline
def subtask_annotation():
src = load_videos()
for items in batch(src, n=16): # batch = chunked/streamed run
frames = videos_to_frames(items.videos) # Part I —
sheets = frames_to_sheets(frames.frames, frames.timestamps) # segmentation
spans = sheets_to_segments(sheets.sheets, items.instructions)
windows = frames_to_windows(frames.frames, spans.spans) # Part II —
labels = windows_to_labels(windows.windows, items.instructions) # labeling
ok = score_segments(labels) # WGO-Bench review
save_dataset(labels[ok]) # sink: accepted spans
rework(labels[~ok]) # short / low-recall → rework

Part I — segmentation

The load-bearing design choice is the visual input: a timestamped contact sheet beats every alternative the post tried. Drawing the timestamp on the frame — rather than mapping frames to times in text — is worth ~6 F1 points, and per-frame individual images do worse still.

Visual inputSegmentation F1
Timestamped contact sheet0.263
Text-only frame→time mapping0.201
Per-frame individual images0.193
Fixed-length baseline0.070

(224 px tiles, 20 frames per sheet in a 4×5 grid, one frame every 0.5 s. Higher tile resolutions did not help.)

Model choice mattered as much as input format: Gemini 3.5 Flash led the board, beating the best non-Gemini model (GPT-5.5) by 24.5%. A round of GEPA prompt search (completed_events_duration_prior_v1) then lifted the best model from 0.290 → 0.306 F1 by teaching it to cut only on completed manipulation events — object grasped, placed, released, a state transition — and not to segment approach, grasp adjustment, small repositioning, or retreat unless the world state changes.

Part II — labeling

Labeling runs on the segmentation output as a prior ("seeded relabeling"): the model is handed the span's first-pass label and the surrounding context, and mostly verifies or minimally corrects it. Neighboring-segment context is what moves the needle.

Context given to the labelerAccuracy
Target segment only (5 frames)56.1%
Previous / current / next segment61.0%
Prev / current / next + episode overview60.8%
Seeded relabeling on predicted spans78.1%

Where it breaks

The pipeline is bottlenecked by short subtasks: segments under 2 seconds are recalled only 7.4% of the time, and recall stays under 50% for everything shorter than 10 seconds. Labeling errors are overwhelmingly grounding errors — wrong target or direction (42.5%) and right-verb-wrong-object (25.0%) dominate. End-to-end (segment and label correct) semantic F1 lands at 0.168.

Performance is also strongly viewpoint-dependent: robot head-camera video is far easier than egocentric human video.

WGO-Bench sourceViewpointEpisodesSegmentsSeg. F1
Galaxearobot head camera251230.589
DROIDexternal robot camera501500.292
HomERegocentric (human)254700.227

100 episodes · 743 segments · 62 unique task instructions. Dataset: macrodata/WGO-Bench.


Source: Annotating Robot Video Subtasks, Macrodata Labs · See also: Pipeline Graph (the component) · Anatomy (field-to-pixel map) · Pipeline Graph JSON (the data model).