The invention discloses a vision-language-action
robot control method and
system based on task stage
semantic enhancement, and the method comprises the steps: firstly constructing a task stage semantic
system, and disassembling a task into stages with definite
semantics; constructing a task stage observation module based on a pre-trained vision-
language model, deducing a task state and
abnormality in real time, and generating stage semantic features; modulation parameters are generated through a lightweight network, and
dynamic modulation is carried out on the vision-language features to strengthen the vision feedback weight; and finally, fusing the modulated characteristics and the body state characteristics, and outputting a control instruction through an action generation network to form closed-
loop control. According to the method, the multi-
modal information weight can be dynamically balanced, task abnormity can be identified in real time, error correction is triggered, the false completion phenomenon is remarkably reduced, the task execution reliability, robustness and generalization ability of the
robot in a complex dynamic environment are improved, and the method is suitable for various
robot systems needing high-credibility operation.