Prebot
Learning-based robotic manipulation for café order preparation.
Overview
Prebot is being developed as a graduation project through Gachon University’s P-CareerCatchⅡ program. It aims to automate the preparation of cups, lids, straws, and napkins by combining order information, visual observations, and learning-based robotic manipulation.
The project uses an OpenMANIPULATOR-X arm and explores imitation learning, reinforcement learning, and vision-language-action models for pick-and-place tasks.
Initial validation is planned in a miniature café workspace, focusing on task success and robustness to changes in item placement.
Problem
Preparing supplies for each order is a repetitive part of café operations that competes with drink preparation and customer service during busy periods. Automating this work requires a robot to handle different order requirements and changes in the workspace.
Fixed-position routines are difficult to reuse when items move. The project therefore explores policies that use current visual observations and task instructions to perform pick-and-place operations across varying item arrangements.
My role
- My work focuses on the robot manipulation and VLA components, connecting visual observations and task instructions to executable robot behavior.
- The manipulation work covers the OMX control environment, demonstration-based imitation learning, and subsequent reinforcement learning for policy improvement.
- The VLA work explores SmolVLA and diffusion-based action generation, with evaluation planned across changes in object positions and workspace layouts.
Approach
- The planned approach begins with establishing the robot control environment and collecting demonstrations for basic pick-and-place tasks.
- Imitation learning will provide an initial manipulation policy, followed by reinforcement learning to explore improvements in task success and adaptation.
- In parallel, the VLA component will connect camera observations and task instructions to robot actions using SmolVLA, while exploring diffusion-based action generation.
- Evaluation will compare performance across item arrangements in simulation and on the physical robot.
System architecture
- Proposed flow: order instruction → LLM task interpretation → SmolVLA manipulation inference → OMX pick-and-place control.
- Order interface: Receives order requirements and displays task progress and results.
- Visual input: A fixed camera observes the items and their arrangement on the workspace.
- Robot learning: Manipulation policies and VLA models determine actions for selecting and placing the required items.
- Execution: The OMX arm performs pick-and-place tasks through the robot control and motion-planning stack.
- Deployment and experiment management: Supporting components handle model optimization, deployment, experiment tracking, and task logs.
Technologies
Current work
- Current work focuses on establishing the OMX manipulation environment and developing baseline learning policies.
- The robot and VLA components are being developed together, with imitation learning as the starting point and reinforcement learning and diffusion-based action generation as subsequent directions.