MaGe, the robotic arm that combines vision and language
A robotic system capable of performing manipulation operations based on instructions expressed in natural language: this is the goal of MaGe – Make Grasping Easy, a research project conducted by the Technologies of Vision (TeV) Unit at Fondazione Bruno Kessler.
The project focuses on the interaction between three-dimensional machine vision, precision robotics and large language models, with the aim of improving the autonomy and flexibility of grasping systems in unstructured environments.
Starting from a textual description and a visual representation of the scene, the system is able to identify relevant objects, estimate their spatial position and plan the movements necessary to interact with them safely. The robotic arm, guided by this data, can thus perform manipulation actions based on commands expressed in natural language—both written and spoken—autonomously determining the most suitable grasping strategy.
“Our goal is to study how language can become an effective means of controlling physical systems, even in complex contexts,” explained Fabio Poiesi with TeV. “The integration between semantic understanding and spatial perception represents an open challenge in robotics.”