The invention discloses a visual language action
large model design method for enhancing
spatial perception ability, which comprises the following steps of: acquiring a plurality of frames of task scene images, inputting the images into a two-dimensional image
encoder and a visual geometry guidance
Transformer VGGT
encoder, extracting image features and performing
time sequence spatial embedding, encoding a
natural language instruction into language embedding representation, and extracting a visual language action
large model. Image features and language embedding are fused through a cross attention mechanism, a pre-training
large model is input to generate cross-
modal representation, the fusion features are input into an action expert module to output an
action control sequence in combination with
robot body state information, and a
robot is driven to execute an operation task. Therefore, damage of additional information to an original pre-training model is avoided, and compared with an original combined structure of a pre-training visual
language model and an action expert, utilization of multi-view picture information is enhanced, so that higher understanding ability on space depth is achieved in the task execution process, the task success rate is increased, and the task execution efficiency is improved. And a more efficient and more robust
robot sensing and decision-making integrated
system is realized.