The invention discloses a human-computer
interaction method and
system based on a vision-language-
action model, and belongs to the field of human-computer interaction. According to the method, an anchoring ring strategy is adopted to collect human teaching data to finely adjust the VLA model, and then the trained model is applied to an actual human-computer interaction scene. In the
data acquisition stage, a first operator guides a master
robot to execute task actions, and a slave
robot synchronously moves and interacts with a second operator, and returns to a predefined initial position after each interaction; a teaching sample is formed by recording a
robot state, an environment image and an instruction text, and a high-
quality data set is generated through data enhancement. In the application stage, the real-time robot state, the environment image and the instruction text serve as input, an action instruction is generated through the VLA model, and the robot is driven to complete a cooperation task. According to the method, the data utilization efficiency and the model generalization ability are remarkably improved, the difference between
simulation and the real environment is effectively overcome, and efficient, safe and natural man-
machine cooperation is achieved.