← 返回 traj 列表

11_insert_pencil_case · insert the pen into the pencil case

skill = insert · target = pencil case · demo episode_00290240 frame 11385 · GT mask 7612px(GT mask 来源: 开合/按钮=sim link, place-in=整物体)

输入

全图 720×720。绿=GT mask,橙框=crop 范围(GT bbox 外扩 10%)
crop(喂给 SAM3/DINO/网格点的视野)
GT mask(评估基准 & arm3 的交集 mask)

Arm1 · UAD — 原句→UAD 热图

UAD 热图 @ 全图(UAD 输入=全图+原句,无 GPT)
crop 视野

Arm2 · OpenSeg 原句 — 原句→OpenSeg (raw, 不取交集)

OpenSeg(原句) raw @ 全图
raw @ crop(对比行里显示的就是这个 raw 版)
参考: ×GT mask 后

Arm3 · GPT短语+OpenSeg — GPT part 短语→OpenSeg→×GT mask

Step 1 · GPT 出 part 短语(输入=全图+crop)

GPT A-openseg — 完整 prompt
[system] You are helping build part-level annotations for robot manipulation data. Answer in JSON only.

[user] The robot's current subgoal is: "insert the pen into the pencil case". The target object is "pencil case".
Image 1 is the full scene; image 2 is a close-up of the target object.
Name the subpart of the target object that should currently be visually focused for this subgoal.
Rules for the phrase (it will be fed to an open-vocabulary segmentation model that expects category-style queries):
- 1-4 words, nouns and adjectives ONLY, compound-noun style: "trash can opening", "door handle", "toggle button", "jar lid".
- No articles, no verbs, no clauses, no "that/which", no gerunds.
Output JSON (reason first): {"reason": "<one short sentence>", "part_phrase": "..."}
GPT: pencil case opening — The pen should be inserted through the pencil case opening.

Step 2 · OpenSeg(短语) → Step 3 · ×GT mask

OpenSeg(GPT短语) raw @ crop
×GT mask 后(= 对比行显示版)
masked @ 全图

Arm4 · mask+DINO+GPT — mask 内 DINO 聚类→GPT 选簇

Step 1 · GT mask 内 DINO 特征 KMeans-6 聚类

聚类伪彩图(只画 mask 内,编号喂给 GPT)

Step 2 · GPT 选簇(输入=全图+crop+聚类图)

GPT B-cluster — 完整 prompt
[system] You are helping build part-level annotations for robot manipulation data. Answer in JSON only.

[user] The robot's current subgoal is: "insert the pen into the pencil case". The target object is "pencil case".
Image 1 is the full scene, image 2 is a close-up of the target object, image 3 shows the same close-up divided into numbered colored regions (clusters).
Which cluster is the subpart that should currently be the visual focus for this subgoal?
Rules:
- Pick 1-2 cluster numbers. Prefer the single best cluster; add a second only if the focused subpart clearly spans two.
- If no cluster matches, return an empty list.
Output JSON (reason first): {"reason": "<one short sentence>", "clusters": [..]}
GPT 选簇: [2] — The case opening or zipper area is represented by cluster 2.

Step 3 · 选中簇 → mask

arm4 结果 @ crop

Arm5 · SAM3 原句 — 原句直接喂 SAM3 (crop)

SAM3(crop, 原句) 结果(无 GPT)

Arm6 · GPT短语+SAM3 — GPT part 短语→SAM3 (crop)

Step 1 · GPT 出 SAM3 风格短语(输入=全图+crop)

GPT A-sam3 — 完整 prompt
[system] You are helping build part-level annotations for robot manipulation data. Answer in JSON only.

[user] The robot's current subgoal is: "insert the pen into the pencil case". The target object is "pencil case".
Image 1 is the full scene; image 2 is a close-up of the target object.
Give a short plain noun phrase naming the subpart of this object that should currently be visually focused for this subgoal, in the style of segmentation prompts like "door handle", "bottle cap", "lid of the jar".
Rules: 1-5 words; plain noun phrase only ("X" / "adjective X" / "X of the Y"); no "that/which" clauses, no gerunds, no past participles.
Output JSON (reason first): {"reason": "<one short sentence>", "part_phrase": "..."}
GPT: pencil case opening — The pen should enter the pencil case opening.

Step 2 · SAM3(crop, 短语)

arm6 结果 @ crop

Arm7 · GPT点+SAM1 — GPT 网格点→SAM1 point prompt

Step 1 · GPT 在坐标网格上给点(输入=全图+网格 crop)

GPT C-points — 完整 prompt
[system] You are helping build part-level annotations for robot manipulation data. Answer in JSON only.

[user] The robot's current subgoal is: "insert the pen into the pencil case". The target object is "pencil case".
Image 1 is the full scene. Image 2 is a close-up of the target object with a coordinate grid: x runs 0 (left) to 10 (right), y runs 0 (top) to 10 (bottom). Grid lines are only visual aids - coordinates are continuous, decimals encouraged (e.g. x=3.7).
Give 1-3 points that lie ON the subpart of this object that should currently be the visual focus for this subgoal.
Rules:
- Points must be inside the subpart, not on its boundary or on other parts.
- First point = the most central/confident location.
Output JSON (reason first): {"reason": "<one short sentence>", "points": [[x, y], ...]}
GPT 点: [[5.5, 2.7], [6.8, 2.6]](红=第一/最自信点) — The pencil case opening is the dark interior slit where the pen should be inserted.
网格图+GPT 点位

Step 2 · SAM1 point-prompt,多点取并集

arm7 结果 @ crop
← 上一例 10_insert_pencil_case↑ 返回 traj 列表下一例 12_insert_pencil_case →