The application provides a multimodal embodied intelligent
robot control method and device, wherein the method comprises: acquiring a current
modal instruction, the current
modal instruction comprising at least one of a
video instruction, a text instruction, a picture instruction and an audio instruction; based on a pre-trained multimodal
large model, acquiring a collapsed representation according to the current
modal instruction; based on a pre-trained
robot imitation learning network, predicting an output
target action according to the collapsed representation and a current environment observation; controlling the multimodal embodied intelligent
robot to operate according to the
target action; wherein the robot
imitation learning network is obtained by training and optimizing a training sample set composed of single
modal data samples, corresponding environment
observation data and real robot actions. The method acquires a collapsed representation by using a pre-trained multimodal
large model, and predicts a
target action by using the collapsed representation, which avoids the need for multimodal artificial
labeled data when training the robot
imitation learning network, and can achieve multimodal reasoning effect only by training single
modal data, thereby reducing data collection and labeling costs, and improving the reasoning ability of the model.