Fine-Tuning OpenVLA-OFT for Household Manipulation Task
We fine-tuned OpenVLA-OFT on 100 expert demonstrations collected in an Isaac Sim kitchen environment, enabling a Franka robot to pick up a cup, transport it to the sink target region, and release it. The dataset contains 34,237 timesteps, and every demonstration was independently verified through action replay in simulation. The policy takes two RGB observations—from a front-side camera and a wrist-mounted camera—together with an 8-dimensional proprioceptive state and a language instruction. It predicts a 7-dimensional action consisting of end-effector translation deltas, rotation deltas, and a gripper command. In preliminary closed-loop evaluations, the fine-tuned policy performed well and successfully executed the complete grasping, transporting, and releasing sequence.
- Zhiyuan Gao
Email: gao@uni-bremen.de
Profile: Zhiyuan Gao
- Mohammad Khoshnazar
Email: khoshnazar@uni-bremen.de
Profile: Mohammad Khoshnazar
- Yanxiang Zhan
Email: yanxiang@uni-bremen.de
Profile: Yanxiang Zhan

