Automatic masks → Layer extraction
This experiment processed 48 Vietnamese marketing posters. Grounding DINO and SAM 2.1 proposed object masks; Hi-SAM proposed text strokes and line groups. Each assigned candidate was then sent to Qwen-Image-Layered-Control-V2.
| Candidate type | Attempts | Nearly blank outputs |
|---|---|---|
| Object | 457 | 115 |
| Text line | 694 | 298 |
Another 33 unassigned stroke groups were retained for inspection. “Nearly blank” means no alpha value exceeds 32/255. Nonempty outputs can also be incorrect.
Objects can be selected separately
In VI_033, four overlapping cars were extracted individually. Hidden parts remain missing, and small lettering or other details may change.
Text generation remains unreliable
The mask cutouts retain the original poster pixels. Qwen can still add a dark fringe, return a nearly blank layer, or produce a colored block instead of the lettering.
These are selected diagnostic examples, not a measured success rate. Masks are unverified predictions and may miss letters or diacritics. Binary mask cutouts still need edge refinement. The outputs do not guarantee a reconstructable layer stack or a repaired background.
Run settings: 30 steps, seed 777, area resolution 640, three H100 GPUs. All 1,151 requested output files passed artifact-integrity checks.