Multimodal embodied intelligent robot control method and device
By combining a pre-trained multimodal large model with a robot imitation learning network, the collapsed representation is obtained, which solves the problem of multimodal embodied intelligent robots relying on massive training samples, and achieves efficient multimodal reasoning and reduced costs.
Patent Information
- Application Number
- CN202411363124.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2044-09-27
AI Technical Summary
Existing multimodal embodied intelligent robot reasoning relies on massive, paired multimodal data sample sets, resulting in high data collection and annotation costs, and traditional methods are unable to fully capture the interdependencies and semantic integration between different modalities.
By obtaining the collapsed representation through pre-trained multimodal large model and using a robot imitation learning network trained on single-modal data, the need for manually labeled multimodal data is avoided, and multimodal reasoning can be achieved using only single-modal data.
It reduces the cost of data collection and labeling while improving the model's reasoning ability, enabling efficient control of multimodal embodied intelligent robots.
Smart Images

Figure CN119141538B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of embodied intelligent robots, and in particular to a multi-modal embodied intelligent robot control method and device. BACKGROUND
[0002] The existing multi-modal embodied intelligent robot reasoning relies on a large number of paired one-to-one multi-modal data sample sets for training, and collecting, labeling and storing these data samples requires a large amount of cost, which is a major obstacle for the application of the robot field.
[0003] Traditional multi-modal learning methods usually process information of different modalities in an independent segmented manner, which is difficult to fully capture the mutual dependence relationship and semantic integration between different modalities. Therefore, it is still a challenge to train a multi-modal embodied robot model that can effectively reason in a complex real-world scenario.
[0004] Therefore, it is necessary to solve the problem that the existing multi-modal embodied intelligent robot reasoning relies on a large number of training samples and requires high cost. SUMMARY
[0005] The present application provides a multi-modal embodied intelligent robot control method and device to overcome the defects of the existing multi-modal embodied intelligent robot reasoning relying on a large number of training samples and requiring high cost, reduce the cost of data collection and labeling, and improve the reasoning ability of the model.
[0006] In one aspect, the present application provides a multi-modal embodied intelligent robot control method, comprising: obtaining a current modal instruction, the current modal instruction comprising at least one of a video instruction, a text instruction, a picture instruction and an audio instruction; based on a pre-trained multi-modal large model, obtaining a collapsed representation according to the current modal instruction; based on a pre-trained robot imitation learning network, predicting an output target action according to the collapsed representation and a current environment observation; controlling the multi-modal embodied intelligent robot to operate according to the target action; wherein the robot imitation learning network is obtained by training and optimizing a training sample set composed of a single modal data sample, its corresponding environment observation data and a real robot action.
[0007] Further, based on the pre-trained multi-modal large model, the collapsed representation is obtained according to the current modal instruction, comprising: inputting the current modal instruction into the pre-trained multi-modal large model to obtain a connected representation; determining a target modal mean corresponding to the current modal instruction; and subtracting the connected representation from the target modal mean to obtain the collapsed representation.
[0008] Further, the training optimization robot imitation learning network specifically comprises: constructing a training sample set, the training sample set comprising single-modal data samples and corresponding environment observation data and robot real actions; based on the single-modal data samples in the training sample set, obtaining training input representations of the robot imitation learning network; taking the training input representations and the environment observation data as model inputs, taking predicted actions as model outputs, and taking differences between the predicted actions and the robot real actions as training losses, and iteratively optimizing the robot imitation learning network.
[0009] Further, based on the single-modal data samples in the training sample set, the training input representations of the robot imitation learning network are obtained, comprising: inputting the single-modal data samples into a pre-trained multi-modal large model to obtain connected representations; subtracting the connected representations from the corresponding modal mean values of the single-modal data samples to obtain collapsed representations; normalizing the collapsed representations and adding a set noise to obtain damaged representations, the damaged representations being the training input representations.
[0010] Further, the set noise comprises cosine noise and / or Gaussian noise.
[0011] Further, the corresponding modal mean values of the single-modal data samples are obtained by the following steps: obtaining a preset number of single-modal data samples; extracting sample features of all single-modal data samples; summing the sample features according to feature dimensions and dividing by the preset number to obtain the corresponding modal mean values of the single-modal data samples.
[0012] In a second aspect, the present application further provides a multi-modal embodied intelligent robot control device, comprising: a current modal instruction obtaining module for obtaining a current modal instruction, the current modal instruction comprising at least one of a video instruction, a text instruction, a picture instruction and an audio instruction; a collapsed representation obtaining module for obtaining a collapsed representation based on a pre-trained multi-modal large model according to the current modal instruction; a target action prediction module for predicting and outputting a target action based on a pre-trained robot imitation learning network and the collapsed representation and a current environment observation; and a target action operation module for controlling a multi-modal embodied intelligent robot to operate according to the target action.
[0013] In a third aspect, the present application further provides a multi-modal embodied intelligent robot, which executes the multi-modal embodied intelligent robot control method according to any one of the above aspects, or comprises the multi-modal embodied intelligent robot control device according to the above aspect.
[0014] In a fourth aspect, the present application also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the multi-modal embodied intelligent robot control method according to any one of the above.
[0015] In a fifth aspect, the present application also provides a non-transitory computer-readable storage medium, having a computer program stored thereon, wherein the computer program is executable on a processor to implement the multi-modal embodied intelligent robot control method according to any one of the above.
[0016] The multi-modal embodied intelligent robot control method provided by the present application comprises the following steps: obtaining a current modal instruction, wherein the current modal instruction comprises at least one of a video instruction, a text instruction, a picture instruction, and an audio instruction; obtaining a collapsed representation according to the current modal instruction; predicting an output target action according to the collapsed representation and a current environment observation based on a pre-trained robot imitation learning network; and controlling a multi-modal embodied intelligent robot to operate according to the target action. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0018] Figure 1 is a flowchart of the multi-modal embodied intelligent robot control method provided by the embodiments of the present application.
[0019] Figure 2 is a training and optimization diagram of the multi-modal large model provided by the embodiments of the present application.
[0020] Figure 3 is a training and optimization diagram of the robot imitation learning network provided by the embodiments of the present application.
[0021] Figure 4 is a whole training and testing diagram of the robot imitation learning network provided by the embodiments of the present application.
[0022] Figure 5It is the whole flow schematic diagram of the multi-modal embodied intelligent robot control method provided by the embodiment of the application.
[0023] Figure 6 It is the structure schematic diagram of the multi-modal embodied intelligent robot control device provided by the embodiment of the application.
[0024] Figure 7 It is the entity structure schematic diagram of the electronic device provided by the embodiment of the application. DETAILED DESCRIPTION
[0025] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described below in connection with the drawings in the present application. Obviously, the described embodiments are some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0026] It should be noted that the robot has the ability to understand human multi-modal instructions (voice instructions, text instructions, picture instructions and video instructions). At present, to train a multi-modal embodied intelligent robot, a large amount of paired one-to-one multi-modal data set is needed, but such data set needs manual annotation and high collection cost.
[0027] In view of this, the present application provides a multi-modal embodied intelligent robot control method, which avoids the need for multi-modal manual annotation data when training a multi-modal embodied intelligent robot, and can achieve multi-modal reasoning effect through single modal training.
[0028] Figure 1 The flow schematic diagram of the multi-modal embodied intelligent robot control method provided by the embodiment of the application is shown. As shown in Figure 1 The method comprises steps S110-S140, which will be described in detail below.
[0029] S110, acquiring a current modal instruction, the current modal instruction comprising at least one of a video instruction, a text instruction, a picture instruction and an audio instruction.
[0030] It is easy to understand that, to make the robot execute an operation, the user needs to input a current modal instruction first. The current modal instruction can be any one of a video instruction, a text instruction, a picture instruction and an audio instruction, or any combination of the video instruction, the text instruction, the picture instruction and the audio instruction. It can be adjusted according to actual conditions.
[0031] The video instruction includes video data, the text instruction includes text data, the picture instruction includes image data, and the audio instruction includes audio data.
[0032] For example, in a specific embodiment, the current modal instruction is a video instruction, the video instruction includes video data, and the robot is prompted to operate according to the video data.
[0033] In another specific embodiment, the current modal instruction includes a video instruction and a text instruction, and the video data included in the video instruction corresponds in content to the text data included in the text instruction.
[0034] On the basis of obtaining the current modal instruction, further, step S120 is executed.
[0035] S120, obtaining a collapsed representation according to the current modal instruction.
[0036] Specifically, first, a pre-trained multi-modal large model can be used to extract useful features from the current modal instruction. Since the multi-modal large model is trained and optimized using all open-source robot operation data sets of different modalities, the multi-modal large model can obtain a connected representation as output when the current modal instruction is input into the multi-modal large model. The connected representation not only contains information of the current modality, but also contains a close relationship between the information of the current modality and the information of other modalities corresponding to the current modal instruction.
[0037] The current modality refers to the modality corresponding to the current modal instruction, such as video, text, image, and audio.
[0038] Then, the connected representation is subtracted by the modality mean corresponding to the current modality, and a collapsed representation is obtained. The modality mean can be obtained by collecting a large number of data samples of the current modality and calculating the average value according to the feature dimension.
[0039] Subtracting the modality mean corresponding to the current modality from the connected representation is also called centering or de-meaning. The centered feature representation (i.e., the collapsed representation) is more stable and is not affected by the overall data shift. At the same time, the risk of overfitting caused by data shift can be reduced, and the generalization ability of the model can be improved. Of course, the centered feature representation is also easier to compare and fuse across modalities.
[0040] That is, the collapsed representation in the present embodiment is a centered feature representation, which is convenient for subsequent machine learning model processing.
[0041] After obtaining the collapsed representation, further, step S130 is executed.
[0042] S130, predicting an output target action based on the pre-trained robot imitation learning network and the collapsed representation and the current environment observation; wherein the robot imitation learning network is obtained by training and optimizing a training sample set composed of single modal data samples, corresponding environment observation data and real robot actions.
[0043] The robot imitation learning network is a deep learning model that has been trained on a large number of example data, which can learn from examples how to perform a specific task.
[0044] The current environment observation is the state information of the environment in which the multi-modal embodied intelligent robot is currently located, including sensor data and other environmental parameters. Among them, the sensor data such as image data obtained by the camera, laser radar data obtained by the laser radar, various force data read by the force sensor, and other environmental parameters such as the position, speed and acceleration of the multi-modal embodied intelligent robot.
[0045] After obtaining the collapsed representation and the current environment observation, the collapsed representation and the current environment observation are input into the pre-trained robot imitation learning network, and the target action that the multi-modal embodied robot should take next can be predicted.
[0046] Among them, the target action can be a continuous action, such as joint angle change, speed adjustment, etc., or a discrete action, such as moving to a certain position, grabbing an object, etc., which is not limited here.
[0047] After determining the target action of the multi-modal embodied intelligent robot, further, step S140 is performed.
[0048] S140, controlling the multi-modal embodied intelligent robot to operate according to the target action.
[0049] It is easy to understand that the target action is converted into specific control instructions that can be directly executed by the multi-modal embodied intelligent robot, and then the execution mechanism of the multi-modal embodied intelligent robot controls the multi-modal embodied intelligent robot to operate according to the converted control instructions.
[0050] It should be noted that the multi-modal embodied intelligent robot in the embodiment includes but is not limited to household robots, companion robots, factory robots, catering robots and cleaning robots.
[0051] In the embodiment, by acquiring the current modal instruction, the current modal instruction at least includes one of a video instruction, a text instruction, a picture instruction and an audio instruction, and according to the current modal instruction, a collapsed representation is acquired, and then based on a pre-trained robot imitation learning network, a target action is predicted according to the collapsed representation and a current environment observation, so that the multi-modal embodied intelligent robot operates according to the target action. The method acquires the collapsed representation by pre-training a multi-modal large model, and predicts the target action by using the collapsed representation. This process avoids the need for multi-modal artificial annotation data when training the robot imitation learning network, and can achieve multi-modal reasoning effect by training only single modal data, which not only reduces the data collection and annotation cost, but also improves the reasoning ability of the model.
[0052] On the basis of the above embodiment, further, the process of acquiring the collapsed representation will be described in detail.
[0053] Based on the pre-trained multi-modal large model, the collapsed representation is acquired according to the current modal instruction, including: inputting the current modal instruction into the pre-trained multi-modal large model to obtain a connected representation; determining a target modal mean corresponding to the current modal instruction; and performing difference processing on the connected representation and the target modal mean to obtain the collapsed representation.
[0054] It can be understood that after acquiring the current modal instruction, the current modal instruction is input into the pre-trained multi-modal large model, and a connected representation can be obtained through feature extraction.
[0055] The pre-trained multi-modal large model is a deep learning model capable of processing multiple types of data (such as text, image, audio, video, etc.), which is pre-trained on a large-scale dataset to learn cross-modal data representation and association. The pre-trained multi-modal large model performs well in various tasks, such as image-text matching, video understanding, cross-modal retrieval, etc.
[0056] Figure 2 A training and optimization schematic diagram of the multi-modal large model provided by the embodiment of the application is shown. As Figure 2 shown, various different modal robot operation datasets (including but not limited to text data, image data, video data and audio data) are used to train and optimize the multi-modal large model, so that the pre-trained multi-modal large model can understand multiple types of input at the same time and learn complex interaction relationships between different modal data, thereby improving the model understanding and prediction ability.
[0057] The connected representation not only contains the information of the current modal, but also contains the close relationship between the information of the current modal and the information of other modal corresponding to the current modal instruction.
[0058] Meanwhile, a target modality mean corresponding to the current modality instruction is obtained. Specifically, different modality instructions have corresponding modality means (for example, image mean, text mean, video mean, and audio mean), and these modality means are pre-calculated and stored. In the case of obtaining the current modality instruction, the current modality corresponding to the current modality instruction can be determined, so that the modality mean corresponding to the current modality can be directly obtained, that is, the modality mean corresponding to the current modality instruction, denoted as the target modality mean.
[0059] Among them, the modality mean corresponding to different modality instructions can be obtained by averaging a large number of single modality data samples in the feature dimension.
[0060] Then, the target modality mean corresponding to the current modality is subtracted from the connected representation, that is, the collapsed representation is obtained.
[0061] Subtracting the target modality mean corresponding to the current modality from the connected representation can also be called centering or de-meaning. The centered feature representation (that is, the collapsed representation) is more stable and is not affected by the overall data offset. At the same time, it can also reduce the risk of overfitting caused by data offset and improve the generalization ability of the model. Of course, the centered feature representation is also easier to compare and fuse across modalities.
[0062] After obtaining the collapsed representation, the collapsed representation and the current environment observation are input into the pre-trained robot imitation learning network to obtain the predicted target action.
[0063] Further, the multi-modal embodied intelligent robot is controlled to operate according to the target action to achieve the effect of controlling the robot operation.
[0064] In this embodiment, by inputting the current modality instruction into the pre-trained multi-modal large model, the connected representation is obtained, and the target modality mean corresponding to the current modality instruction is determined. Then, the connected representation and the target modality mean are processed by difference to obtain the collapsed representation. Then, based on the pre-trained robot imitation learning network, the target action is predicted and output according to the collapsed representation and the current environment observation. Therefore, the multi-modal embodied intelligent robot is controlled to operate according to the target action. This method obtains the collapsed representation by using the pre-trained multi-modal large model, and predicts the target action by using the collapsed representation. This process avoids the need for multi-modal artificial annotation data when training the robot imitation learning network. The model can achieve multi-modal reasoning effect only by training single modality data, which not only reduces the data collection and annotation cost, but also improves the reasoning ability of the model.
[0065] On the basis of the above embodiment, further, the training and optimization process of the robot imitation learning network will be described in detail.
[0066] The training optimization robot imitation learning network specifically comprises: constructing a training sample set, the training sample set comprising single-modal data samples and corresponding environment observation data and robot real actions; based on the single-modal data samples in the training sample set, obtaining training input representations of the robot imitation learning network; taking the training input representations and the environment observation data as model inputs, taking predicted actions as model outputs, and taking differences between the predicted actions and the robot real actions as training losses to iteratively optimize the robot imitation learning network.
[0067] First, a training sample set is constructed. In the embodiment, when training the robot imitation learning network, no multi-modal artificial annotation data is needed, and only single-modal data samples are collected.
[0068] The single-modal data samples and corresponding environment observation data and robot real actions included in the training sample set can be derived from actual robot operations, or from simulation environments or artificial demonstrations, which are not specifically limited here.
[0069] In a specific embodiment, the training sample set is collected from multiple robot platforms, such as the RoboNet project jointly researched by Carnegie Mellon University, the University of Pennsylvania and Stanford University, which provides a video frame dataset from different robot platforms for learning a vision-based robot operation model.
[0070] After the training sample set is constructed, based on the single-modal data samples in the training sample set, the training input representations of the robot imitation learning network are obtained, including: inputting the single-modal data samples into a pre-trained multi-modal large model to obtain connected representations; subtracting the modal mean corresponding to the single-modal data samples from the connected representations to obtain collapsed representations; normalizing the collapsed representations and adding a set noise to obtain destroyed representations, which are the training input representations.
[0071] The modal mean corresponding to the single-modal data samples is obtained, specifically, first a preset number of single-modal data samples are obtained, then sample features of all single-modal data samples are extracted, then all sample features are summed according to feature dimensions and divided by the preset number, and the modal mean corresponding to the single-modal data samples is obtained. The preset number can be set according to actual needs, which is not specifically limited here.
[0072] The set noise can be cosine noise or Gaussian noise, which is not specifically limited here.
[0073] In a specific embodiment, a process of adding cosine noise to the collapsed representations is given.
[0074] Suppose the original vector (i.e., the collapsed representation) is with size [10, 1024]. First, the collapsed representation is normalized to obtain a normalized representation, which can be seen from the following formula (1).
[0075] (1).
[0076] In formula (1), represents the normalized result of the th vector, i.e., the normalized representation.
[0077] Then, a random orthogonal vector is generated for the normalized representation. First, a random vector is generated, then the projection of the random vector on the normalized representation is subtracted , and then the result is normalized, which can be seen from the following formula (2).
[0078] (2).
[0079] Next, a target cosine similarity is given to generate a target vector . The broken representation is a linear combination of the normalized representation and the random orthogonal vector, which can be seen from the following formula (3).
[0080] (3).
[0081] The final target vector has a target cosine similarity with the original vector .
[0082] The target vector in the above formula (3) is the broken representation, which is a random number generated after the cosine noise size is determined. For example, if the cosine noise is set to 0.7, a floating-point number between 0.7-1.0 will be randomly sampled as the value of the target cosine similarity each time the cosine noise is added.
[0083] In another specific embodiment, a process of adding Gaussian noise to the collapsed representation is given.
[0084] First, the original vector (i.e., the collapsed representation) is normalized to obtain a unit vector , which can be seen from the following formula (4).
[0085] (4).
[0086] Then, Gaussian noise is generated according to the standard deviation of each dimension. Assuming represents the standard deviation of each dimension, with a size of
[1024] , the generated Gaussian noise can be represented by the following formula (5).
[0087] (5).
[0088] Next, the generated Gaussian noise is added to the normalized unit vector , which can be seen in the following formula (6).
[0089] (6).
[0090] Finally, the vector after adding noise is normalized to obtain the final vector , which can be seen in the following formula (7).
[0091] (7).
[0092] The vector in formula (7) is the destroyed representation.
[0093] After obtaining the destroyed representation (training input representation), the training input representation and its corresponding environmental observation data are used as the input of the robot imitation learning network, the predicted action is output, and the difference between the predicted action and the real action of the robot is used as the training loss. The robot imitation learning network is iteratively optimized, so that the robot imitation learning network trained to convergence is obtained.
[0094] In addition, Figure 3 a training optimization diagram of the robot imitation learning network provided by the embodiment of the application is shown.
[0095] As Figure 3 shown, first, a single modal data sample (including but not limited to video data, text data, image data, audio data) is input into a multi-modal large model (pre-trained multi-modal large model) to obtain an output connected representation. Then, the modal mean of the corresponding modal of the single modal data sample is subtracted from the connected representation to obtain a collapsed representation. Then, the collapsed representation is normalized and noise (cosine noise in Figure 3 , which can also be other noise) is added to obtain a destroyed representation.
[0096] Finally, the destroyed representation and the environmental observation data are used as the input of the robot imitation learning network for imitation learning. After multiple iterations of optimization, a trained robot imitation learning network can be obtained.
[0097] Figure 4A schematic diagram of the overall training and testing of the robot imitation learning network provided by the embodiment of the present application is shown.
[0098] As shown in Figure 4 When training the optimized robot imitation learning network, any single modal data sample can be used as a training sample without the need for a large number of multi-modal data sample annotations. When testing and applying the robot imitation learning network, any single modal data can be used as model input for target action prediction. As can be seen, the robot imitation learning network provided by the embodiment can achieve multi-modal data reasoning effect based on single modal data samples.
[0099] It should be noted here that using any single modal data sample as a training sample can avoid the need for multi-modal artificial annotation data when training the robot imitation learning network, greatly reducing data collection and annotation costs.
[0100] It should also be noted that when testing and applying, a plurality of corresponding modal data combinations can also be used as input to the robot imitation learning network, which is not limited here.
[0101] In the present embodiment, a training sample set is constructed, and based on the single modal data samples in the training sample set, a training input representation of the robot imitation learning network is obtained, and then the training input representation and the environment observation data are used as model input, the predicted action is used as model output, the difference between the predicted action and the real action of the robot is used as training loss, the robot imitation learning network is iteratively optimized, and then based on the pre-trained robot imitation learning network, the target action is predicted and output according to the collapsed representation and the current environment observation, so as to control the multi-modal embodied intelligent robot to operate according to the target action. The method obtains the collapsed representation by pre-training the multi-modal large model, and predicts the target action using the collapsed representation. This process avoids the need for multi-modal artificial annotation data when training the robot imitation learning network, and can achieve multi-modal reasoning effect based on only single modal data, which not only reduces data collection and annotation costs, but also improves the reasoning ability of the model.
[0102] In addition, in some embodiments, Figure 5 A schematic diagram of the overall flow of the multi-modal embodied intelligent robot control method provided by the embodiment of the present application is shown.
[0103] As shown in Figure 5 First, the current modal instruction is obtained. The current instruction can be any one or a combination of multiple items of video instructions, text instructions, picture instructions, and audio instructions.
[0104] Then, the obtained current modal instruction is input into the pre-trained multi-modal large model to obtain the connected representation.
[0105] Subsequently, the connection after the feature is subtracted from the target modal average corresponding to the current modal instruction to obtain the collapsed feature.
[0106] Finally, the collapsed feature is input into the robot imitation learning network together with the current environment observation to obtain the output target action, and then the multi-modal embodied intelligent robot is controlled to operate according to the target action, thereby realizing precise control of the multi-modal embodied intelligent robot.
[0107] Corresponding to the multi-modal embodied intelligent robot control method described in the above embodiments, the present application also provides a multi-modal embodied intelligent robot control device. Specifically, Figure 6 The structure of the multi-modal embodied intelligent robot control device provided by the embodiment of the present application is shown.
[0108] As Figure 6 shown, the device comprises: a current modal instruction acquisition module 610 for acquiring a current modal instruction, the current modal instruction comprising at least one of a video instruction, a text instruction, a picture instruction and an audio instruction; a collapsed feature acquisition module 620 for acquiring a collapsed feature based on a pre-trained multi-modal large model according to the current modal instruction; a target action prediction module 630 for predicting an output target action based on a pre-trained robot imitation learning network according to the collapsed feature and a current environment observation; and a target action operation module 640 for controlling a multi-modal embodied intelligent robot to operate according to the target action; wherein the robot imitation learning network is obtained by training and optimizing a training sample set composed of a single modal data sample, its corresponding environment observation data and a real robot action.
[0109] In this embodiment, the current modal instruction acquisition module 610 acquires the current modal instruction, the current modal instruction comprising at least one of a video instruction, a text instruction, a picture instruction and an audio instruction, the collapsed feature acquisition module 620 acquires the collapsed feature based on a pre-trained multi-modal large model according to the current modal instruction, and then the target action prediction module 630 predicts an output target action based on a pre-trained robot imitation learning network according to the collapsed feature and a current environment observation, so that the target action operation module 640 controls the multi-modal embodied intelligent robot to operate according to the target action. The device acquires the collapsed feature through the pre-trained multi-modal large model, and predicts the target action using the collapsed feature. This process avoids the need for multi-modal artificial annotation data when training the robot imitation learning network, and can achieve multi-modal reasoning effect only by training single modal data, thereby reducing the data collection and annotation cost, and improving the reasoning ability of the model.
[0110] It should be noted that the multi-modal embodied intelligent robot control device provided by the embodiments of the present application can be correspondingly referred to the multi-modal embodied intelligent robot control method described in the above embodiments, and will not be described here.
[0111] In addition, the present application also provides a multi-modal embodied intelligent robot, which includes the multi-modal embodied intelligent robot control device described in the above embodiments, and executes the multi-modal embodied intelligent robot control method described in the above embodiments during operation. The robot obtains the collapsed representation by pre-training the multi-modal large model, and predicts the target action by using the collapsed representation. This process avoids the need for multi-modal artificial annotation data when training the robot imitation learning network, and can achieve multi-modal reasoning effect by training only according to single modal data, which not only reduces the data collection and annotation cost, but also improves the reasoning ability of the model.
[0112] Figure 7 An example of an entity structure diagram of an electronic device is shown in Figure 7 As shown, the electronic device can include a processor 710, a communications interface 720, a memory 730, and a communications bus 740, wherein the processor 710, the communications interface 720, and the memory 730 can communicate with each other through the communications bus 740. The processor 710 can invoke the logic instructions in the memory 730 to execute the multi-modal embodied intelligent robot control method, which includes: obtaining a current modal instruction, the current modal instruction at least including one of a video instruction, a text instruction, a picture instruction, and an audio instruction; obtaining a collapsed representation according to the current modal instruction based on a pre-trained multi-modal large model; predicting an output target action according to the collapsed representation and a current environment observation based on a pre-trained robot imitation learning network; controlling a multi-modal embodied intelligent robot to operate according to the target action; wherein the robot imitation learning network is obtained by training and optimizing according to a training sample set composed of single modal data samples, corresponding environment observation data, and real robot actions.
[0113] In addition, the logic instructions in the memory 730 described above can be implemented in the form of software functional units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0114] In another aspect, the present application also provides a non-transitory computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the multi-modal embodied intelligent robot control method provided by the above method, the method comprising: obtaining a current modal instruction, the current modal instruction comprising at least one of a video instruction, a text instruction, a picture instruction, and an audio instruction; based on a pre-trained multi-modal large model, obtaining a collapsed representation according to the current modal instruction; based on a pre-trained robot imitation learning network, predicting an output target action according to the collapsed representation and a current environment observation; controlling a multi-modal embodied intelligent robot to operate according to the target action; wherein the robot imitation learning network is obtained by training and optimizing according to a training sample set composed of a single modal data sample, corresponding environment observation data, and a real robot action.
[0115] The device embodiments described above are only schematic, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment. Those skilled in the art can understand and implement without creative labor.
[0116] Those skilled in the art can clearly understand the technical solutions of the various embodiments from the above description of the embodiments, and the various embodiments can be implemented by means of software with the necessary general hardware platforms, and of course, can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that makes a contribution, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0117] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features therein; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A control method for a multimodal embodied intelligent robot, characterized in that, include: Obtain the current modal instruction, which includes at least one of video instructions, text instructions, image instructions, and audio instructions; Based on the pre-trained multimodal large model, the collapsed representation is obtained according to the current modal instruction; Based on a pre-trained robot imitation learning network, the target action is predicted and output according to the collapsed representation and current environmental observations. Control the multimodal embodied intelligent robot to operate according to the target action; The robot imitation learning network is obtained by training and optimizing a training sample set consisting of single-modal data samples, their corresponding environmental observation data, and the robot's actual actions. Training and optimizing the robot's imitation learning network specifically includes: Construct a training sample set, which includes single-modal data samples and their corresponding environmental observation data and robot real actions; Based on the single-modal data samples in the training sample set, the training input representation of the robot imitation learning network is obtained; The robot imitation learning network is iteratively optimized by using training input representations and environmental observation data as model inputs, predicting actions as model outputs, and the difference between the predicted actions and the robot's actual actions as training loss.
2. The multimodal embodied intelligent robot control method according to claim 1, characterized in that, Based on a pre-trained multimodal large model, and according to the current modal instruction, the collapsed representation is obtained, including: The current modal command is input into a pre-trained multimodal large model to obtain the connected representation; Determine the target modal mean corresponding to the current modal command; The collapsed representation is obtained by subtracting the mean of the target mode from the post-connection representation.
3. The multimodal embodied intelligent robot control method according to claim 1, characterized in that, Based on single-modal data samples in the training sample set, the training input representation of the robot imitation learning network is obtained, including: Single-modal data samples are input into a pre-trained multimodal large model to obtain the concatenated representation; Subtracting the modal mean corresponding to the single modal data sample from the concatenated representation yields the collapsed representation. The collapsed representation is normalized and a set noise is added to obtain the destroyed representation, which is the training input representation.
4. The multimodal embodied intelligent robot control method according to claim 3, characterized in that, The set noise includes cosine noise and / or Gaussian noise.
5. The multimodal embodied intelligent robot control method according to claim 3, characterized in that, The modal mean corresponding to the single modal data sample is obtained through the following steps: Obtain a preset number of single-modality data samples; Extract sample features from all single-modality data samples; The modal mean of the single modal data sample is obtained by summing all the features according to the feature dimension and dividing by a preset number.
6. A control device for a multimodal embodied intelligent robot, characterized in that, include: The current modal instruction acquisition module is used to acquire the current modal instruction, which includes at least one of video instructions, text instructions, image instructions, and audio instructions; The post-collapse representation acquisition module is used to acquire the post-collapse representation based on the pre-trained multimodal large model and according to the current modality instruction. The target action prediction module is used to predict the output target action based on the pre-trained robot imitation learning network, according to the collapsed representation and the current environmental observation. The target action operation module is used to control the multimodal embodied intelligent robot to perform operations according to the target action; Training and optimizing the robot's imitation learning network specifically includes: Construct a training sample set, which includes single-modal data samples and their corresponding environmental observation data and robot real actions; Based on the single-modal data samples in the training sample set, the training input representation of the robot imitation learning network is obtained; The robot imitation learning network is iteratively optimized by using training input representations and environmental observation data as model inputs, predicting actions as model outputs, and the difference between the predicted actions and the robot's actual actions as training loss.
7. A multimodal embodied intelligent robot, characterized in that, The method of controlling a multimodal embodied intelligent robot as described in any one of claims 1-5 is executed during runtime, or it includes the multimodal embodied intelligent robot control device as described in claim 6.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the multimodal embodied intelligent robot control method as described in any one of claims 1 to 5.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the multimodal embodied intelligent robot control method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Intelligent robot with garbage classification function
CN117140470A