🤖 DM0.5 — Vision-Language-Action robot control
Dexmal/DM05 is an open-world Vision-Language-Action (VLA) foundation model (Gemma3-4B backbone + 680M action expert). Given three robot camera views and a natural-language instruction, it predicts a chunk of 50 future bimanual actions (6 joint targets + 1 gripper command per arm) via flow-matching.
Provide the three camera images and a task instruction, then inspect the predicted joint / gripper trajectories. Built with OpenDM. Robot type: DOS W1 (14-dim state).
Examples