A training method for generating a multimodal large model of the maneuverable area of ​​an embodied robot.

By pre-training and fine-tuning the multimodal large model and integrating various data and expert model instructions, the robot's ability to accurately extract operable areas has been improved. This solves the generalization and accuracy problems of the multimodal large model in the robot, and achieves higher operational accuracy and environmental perception capabilities.

CN120181127BActive Publication Date: 2026-05-26BEIJING ACAD OF ARTIFICIAL INTELLLIGENCE
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING ACAD OF ARTIFICIAL INTELLLIGENCE
Filing Date
2025-03-21
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing multimodal large models have low generalization and accuracy in embodied robots, leading to operational errors in variable environments and preventing their widespread application.

Method used

By acquiring multimodal data from the embodied robot and instruction data from the expert model, the multimodal large model is pre-trained and fine-tuned. Visual data, language instruction data, and robot body and sensor data are integrated, and the expert model is used to improve the model's ability to accurately extract operable areas.

Benefits of technology

It has improved the accuracy and precision of embodied robots in changing environments, enhanced their perception of complex environments, reduced operational errors, and laid the foundation for the widespread application of embodied robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181127B_ABST
    Figure CN120181127B_ABST
Patent Text Reader

Abstract

This invention discloses a training method for a multimodal large model used to generate the operable areas of an embodied robot, relating to the field of artificial intelligence technology. The method includes: pre-training the multimodal large model using the multimodal data to obtain a pre-trained multimodal large model; and fine-tuning the pre-trained multimodal large model using instruction data from an expert model to obtain a trained multimodal large model. This allows the method to output the operable key object parts of the embodied robot and the semantic relationships between them, and / or, by calling an expert model, output visualized operable key points and their positional coordinates. This improves the embodied robot's operational capability and flexibility in complex environments; enhances its ability to process multimodal information and improves the accuracy of environmental perception; and increases the accuracy of object segmentation and localization, reducing operational errors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method for training a multimodal large model for generating operable areas of an embodied robot. Background Technology

[0002] With the rapid development of artificial intelligence technology, the demand for embodied robots in various application scenarios is increasing. Embodied robots not only need to possess autonomous navigation and environmental perception capabilities, but also need to understand and generate their operable areas in order to perform tasks efficiently in complex environments. However, traditional environmental modeling methods often rely on single-modal data, such as two-dimensional images or LiDAR data, which limits the robot's perception capabilities and operational flexibility.

[0003] Currently, with the successful theoretical research and practical application of the large model scale law, researchers both domestically and internationally have begun to introduce multimodal large models into the field of embodied robotics. Through learning from massive amounts of multimodal data, multimodal large models can capture more complex and richer environmental features, thus demonstrating a stronger ability to understand and generate operable regions. However, existing multimodal large models still suffer from low generalization and low accuracy, often leading to operational errors in embodied robots in changing environments, hindering their widespread application. Summary of the Invention

[0004] In order to solve the problems existing in the prior art, the present invention provides the following technical solution.

[0005] This invention provides a method for training a multimodal large model to generate the operable area of ​​an embodied robot, comprising:

[0006] Acquire multimodal data of the embodied robot and instruction data for calling expert models;

[0007] The multimodal large model is pre-trained using the aforementioned multimodal data to obtain a pre-trained multimodal large model;

[0008] The pre-trained multimodal large model is fine-tuned using the instruction data of the expert model to obtain a trained multimodal large model. The trained multimodal large model can then use the input multimodal data of the embodied robot to output the semantic relationships between the operable key object parts of the embodied robot and the operable key object parts, or output the semantic relationships between the operable key object parts of the embodied robot and the operable key object parts, and output visualized operable key points and the position coordinates of the operable key points by calling the expert model.

[0009] The multimodal data includes visual data, language command data, and robot body and sensor data.

[0010] The expert model is used to support the accurate extraction of operable regions by the multimodal large model.

[0011] Preferably, the multimodal large model training method for generating the operable area of ​​the embodied robot further includes the step of aligning the multimodal data of the embodied robot.

[0012] Preferably, obtaining the instruction data for calling the expert model includes: generating instruction data for calling the expert model using the generative large model GPT-4.

[0013] Preferably, the expert model includes an image segmentation model and an image localization model.

[0014] Preferably, the multimodal large model is an LLaVA multimodal large model.

[0015] Preferably, the visual data includes images, videos, and point clouds.

[0016] A second aspect of the present invention provides a method for generating the operable area of ​​a embodied robot based on a multimodal large model, comprising:

[0017] The multimodal data of the embodied robot is input into the trained multimodal large model, and the semantic relationship between the operable key object parts and the operable key object parts is output, or the semantic relationship between the operable key object parts and the operable key object parts is output, and the visualized operable key points and the position coordinates of the operable key points are output by calling the expert model.

[0018] The trained multimodal large model is pre-trained using the method described in the first aspect.

[0019] A third aspect of the present invention provides a multimodal large model training apparatus for generating operable areas of an embodied robot, comprising:

[0020] The training data acquisition module is used to acquire multimodal data of the embodied robot and instruction data for calling the expert model;

[0021] The pre-training module is used to pre-train the multimodal large model using the multimodal data to obtain the pre-trained multimodal large model.

[0022] The fine-tuning module is used to fine-tune the pre-trained multimodal large model using the instruction data of the expert model to obtain a trained multimodal large model; so that the trained multimodal large model can use the input multimodal data of the embodied robot to output the semantic relationship between the operable key object parts of the embodied robot and the operable key object parts, or output the semantic relationship between the operable key object parts of the embodied robot and the operable key object parts, and output the visualized operable key points and the position coordinates of the operable key points by calling the expert model;

[0023] The multimodal data includes visual data, language command data, and robot body and sensor data.

[0024] The expert model is used to support the accurate extraction of operable regions by the multimodal large model.

[0025] A fourth aspect of the present invention provides a memory characterized in that it stores a plurality of instructions for implementing the multimodal large model training method for generating operable areas of a hymened robot as described in the first aspect.

[0026] The fifth aspect of the present invention provides an electronic device including a processor and a memory connected to the processor, the memory storing a plurality of instructions which can be loaded and executed by the processor to enable the processor to perform a multimodal large model training method for generating operable areas of an embodied robot as described in the first aspect.

[0027] The beneficial effects of this invention are as follows: The multimodal large-scale model training method for generating the operable area of ​​an embodied robot provided by this invention comprehensively improves the embodied robot's ability to operate and perceive complex environments by training the large model using multimodal data; by fine-tuning the pre-trained multimodal large-scale model using instruction data from expert models, the accuracy of object segmentation and localization of the embodied robot is improved, and operational errors caused by insufficient model generalization are reduced. Therefore, adopting the technical solution provided by this invention can improve the accuracy and precision of embodied robots operating in changing environments, laying the foundation for their widespread application. Attached Figure Description

[0028] Figure 1 This is a flowchart illustrating the multimodal large model training method for generating the operable area of ​​an embodied robot as described in this invention.

[0029] Figure 2 This is a schematic diagram of the structural framework for generating a multimodal large model of the operable area of ​​a embodied robot, as described in this invention.

[0030] Figure 3This is a functional structural diagram of the multimodal large model training device for generating operable areas of a embodied robot, as described in this invention. Detailed Implementation

[0031] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0032] The method provided by this invention can be implemented in a terminal environment that may include one or more of the following components: a processor, a memory, and a display screen. The memory stores at least one instruction, which is loaded and executed by the processor to implement the method described in the following embodiments.

[0033] A processor may include one or more processing cores. The processor connects various parts of the terminal using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets or instruction sets stored in memory, and by calling data stored in memory.

[0034] Memory can include random access memory (RAM) or read-only memory (ROM). Memory can be used to store instructions, programs, code, code sets, or instructions.

[0035] The display screen is used to show the user interface of each application.

[0036] In addition, those skilled in the art will understand that the above-described structure of the terminal does not constitute a limitation on the terminal. The terminal may include more or fewer components, or combine certain components, or have different component arrangements. For example, the terminal may also include radio frequency circuits, input units, sensors, audio circuits, power supplies, and other components, which will not be described in detail here.

[0037] Example 1

[0038] like Figure 1 As shown, this embodiment of the invention provides a multimodal large model training method for generating the operable area of ​​an embodied robot, including:

[0039] S101, acquire multimodal data of the embodied robot and instruction data for calling the expert model;

[0040] S102, The multimodal large model is pre-trained using the multimodal data to obtain a pre-trained multimodal large model;

[0041] S103, the pre-trained multimodal large model is fine-tuned using the instruction data of the expert model to obtain a trained multimodal large model. The trained multimodal large model can then use the input multimodal data of the embodied robot to output the semantic relationship between the operable key object parts of the embodied robot and the operable key object parts, or output the semantic relationship between the operable key object parts of the embodied robot and the operable key object parts, and output the visualized operable key points and the position coordinates of the operable key points by calling the expert model.

[0042] The multimodal data includes visual data, language command data, and robot body and sensor data.

[0043] The expert model is used to support the accurate extraction of operable regions by the multimodal large model.

[0044] The core concepts of the solution provided by this invention include the following aspects:

[0045] Fusion of multimodal data from embodied robots and pre-training of multimodal large models: To address the problem of insufficient utilization of multimodal information in existing technologies, this invention integrates embodied robot data information from multiple modalities, such as images, videos, point clouds, and language, and uses the advanced multimodal large model framework LLaVA for data modeling and feature extraction to comprehensively improve the multimodal large model's ability to operate robots and perceive the environment.

[0046] Construction of Expert Model Instruction Data and Fine-tuning of Multimodal Large Model Instructions: While multimodal large models excel at understanding visual input and human instructions, they are not adept at precise localization and segmentation of manipulable regions. Therefore, this invention fine-tunes the multimodal large model by constructing instructions to "call the expert model," enabling it to "call the expert model when deemed necessary." This leverages the strengths of expert models to compensate for their weaknesses. With this capability, when faced with the task of precise localization and segmentation of manipulable regions, the multimodal large model can understand visual input and human instructions, obtain a "language description" of the manipulable key points, and then directly call the expert model to generate the position coordinates and visualization of the manipulable key points based on the image input and the "language description." Specifically, to address the performance disadvantages of existing technologies in object segmentation and localization, this invention constructs instruction data to call the expert model and further fine-tunes the pre-trained multimodal large model to improve its ability to call the expert model. This fully utilizes the advantages of expert models in object segmentation and localization, avoiding the accuracy degradation caused by directly using the multimodal large model for object segmentation and localization.

[0047] Robot Propriometry and Scene Understanding: To address the problem of low generalization caused by the lack of proprioception in existing technologies, this invention inputs information such as robot model and robotic arm type into a multimodal large model via text to ensure the generation of a realistic and reliable operable area for the current robot.

[0048] Through these technological innovations, the solution provided by this invention will effectively enhance the operational capabilities of embodied robots in diverse environments, solve the problems of low generalization and low precision faced by existing technologies, and provide a solid technical foundation for the widespread application of embodied robots.

[0049] In this embodiment of the invention, the structural framework for generating a multimodal large model of the operable area of ​​the embodied robot can be as follows: Figure 2 As shown. It mainly consists of the following modules:

[0050] (1) Multimodal data input layer. This layer is responsible for aligning input data from different modalities. For example, for the same input data sample, visual input (images, point clouds, video data), language input (language commands to the embodied robot), and robot input (including robot body parameters and sensor parameters) need to correspond one-to-one.

[0051] (2) Multimodal data mapping layer. It is responsible for mapping data of different modalities to high-dimensional space through multilayer perceptron, representing them as a set of vectors, and meeting the dimensionality requirements that multimodal large models can handle.

[0052] (3) Multimodal Large Model Modeling Layer. This layer is responsible for uniformly modeling the vectors obtained from the multimodal data mapping layer, enabling the large model to understand the input data information. The modeling process includes the multimodal large model modeling process and the expert model invocation modeling process.

[0053] (4) Operational Area Output Layer. This layer is responsible for outputting the modeled data. The specific output includes visualization of operational key points (output by the expert model), location coordinates of operational key points (output by the expert model), and semantic relationships between operational key object parts and operational key object parts (output by the multimodal large model).

[0054] exist Figure 2 In this system, given visual input, human command input (language input), and robot input, the multimodal large model can understand the commands and output "language descriptions" and semantic relationships of the robot's operable key points through visual and robot inputs. Simultaneously, the multimodal large model can call upon expert models when appropriate to generate the position coordinates and visualization of the operable key points.

[0055] exist Figure 2In the example, the multimodal large model understands that water from the teapot needs to be poured into the cup, so it can provide a linguistic description of the grasping order and the pouring action. This includes the "linguistic description" and semantic relationships of the key operable points, such as the "teapot handle," "teapot spout," and "cup rim," as well as their relationships, such as "the teapot spout should be aligned with the cup rim." Then, the expert model uses the linguistic description and visual input to accurately output the position coordinates and visualization of the key operable points.

[0056] In this embodiment of the invention, the training for generating a multimodal large model of the operable area of ​​the embodied robot can be performed using the following steps:

[0057] The first step is data acquisition and preprocessing. This involves integrating multimodal data such as images, videos, point clouds, and speech data to meet the input format requirements of a large multimodal model. Simultaneously, the acquired multimodal data is cleaned to ensure data quality.

[0058] The second step is the pre-training of the multimodal large model. Based on the data collected and preprocessed in the first step, the multimodal large model is pre-trained on a large scale using the LLaVA framework.

[0059] The third step is to construct expert model instruction data. The generative large model GPT-4 is used to generate call data for expert models (such as the DINOv2 model and the SAM model), forming a standardized call instruction fine-tuning dataset.

[0060] The fourth step is fine-tuning the multimodal large model instructions. Based on the expert model invocation instructions generated in the third step, the dataset is fine-tuned, and training continues using the multimodal large model framework LlaVA to enhance the multimodal large model's ability to actively invoke expert models when facing object segmentation and localization tasks.

[0061] In this embodiment of the invention, when the trained multimodal large model receives multimodal information such as language commands and visual information, it outputs the corresponding result depending on whether an expert model is invoked. If the expert model is not invoked, only the answer based on the multimodal large model is output, including the operable key object parts and the semantic relationships between them. If the expert model is invoked, the answer based on the multimodal large model is output, and the expert model is invoked. The expert model performs image segmentation and localization based on the answer from the multimodal large model, utilizing the operable key object parts and the semantic relationships between them, and outputs the operable key points and their position coordinates. That is, if the expert model is invoked, the operable key object parts of the embodied robot and the semantic relationships between them are output, and the visualized operable key points and their position coordinates are output through the invocation of the expert model. Therefore, in this invention, there is no need to train an expert model. Instead, the expert model is used as a tool. After the trained multimodal large model calls the expert model, the expert model can directly use the operable key object parts and the semantic relationships between the operable key object parts output by the multimodal large model to perform image segmentation and localization, and output the operable key points and the position coordinates of the operable key points.

[0062] In this embodiment of the invention, expert models are used to support the accurate extraction of operable regions by a multimodal large model. They typically have the function of localization and segmentation based on images and verbal descriptions. For example, the DINOv2 model has the function of feature extraction and localization based on images of embodied robots; the SAM model has the function of segmentation based on verbal descriptions and localized coordinates. Different expert models are called to output different content, achieving different functions. In this invention, to solve the problems of image segmentation and localization, the expert models used include image segmentation models and image localization models. For example, the DINOv2 model is used to output the features of the image in the embodied robot's visual information and aggregate them to form key points and their position coordinates; the SAM model is used to further segment the object parts where the key points are located in the image based on the key points and their position coordinates. That is, the DINOv2 model is used for localization, and the SAM model is used for segmentation.

[0063] Example 2

[0064] This invention provides a method for generating the operable area of ​​a embodied robot based on a multimodal large model, comprising:

[0065] The multimodal data of the embodied robot is input into the trained multimodal large model, and the semantic relationship between the operable key object parts and the operable key object parts is output, or the semantic relationship between the operable key object parts and the operable key object parts is output, and the visualized operable key points and the position coordinates of the operable key points are output by calling the expert model.

[0066] The trained multimodal large model is pre-trained using the method described in Example 1.

[0067] Example 3

[0068] like Figure 3 As shown, the present invention also includes a functional module architecture that is completely consistent with the aforementioned method flow. That is, the embodiments of the present invention also provide a multimodal large model training device for generating the operable area of ​​the embodied robot, including:

[0069] The training data acquisition module 201 is used to acquire multimodal data of the embodied robot and instruction data for calling the expert model;

[0070] The pre-training module 202 is used to pre-train the multimodal large model using the multimodal data to obtain the pre-trained multimodal large model.

[0071] The fine-tuning module 203 is used to fine-tune the pre-trained multimodal large model using the instruction data of the expert model to obtain a trained multimodal large model; so that the trained multimodal large model can use the input multimodal data of the embodied robot to output the semantic relationship between the operable key object parts of the embodied robot and the operable key object parts, or output the semantic relationship between the operable key object parts of the embodied robot and the operable key object parts, and output the visualized operable key points and the position coordinates of the operable key points by calling the expert model;

[0072] The multimodal data includes visual data, language command data, and robot body and sensor data.

[0073] The expert model is used to support the accurate extraction of operable regions by the multimodal large model.

[0074] Furthermore, the device also includes a data processing module for aligning the multimodal data of the embodied robot.

[0075] Furthermore, in the training data acquisition module, acquiring the instruction data for calling the expert model includes: generating instruction data for calling the expert model using the generative large model GPT-4.

[0076] Furthermore, the expert model includes an image segmentation model and an image localization model. The multimodal large model adopts the LLaVA multimodal large model. The visual data includes images, videos, and point clouds.

[0077] This device can be implemented using the multimodal large model training method for generating the operable area of ​​the embodied robot provided in Embodiment 1 above. The specific implementation method can be found in the description in Embodiment 1, and will not be repeated here.

[0078] The present invention also provides a memory that stores multiple instructions for implementing the multimodal large model training method for generating the operable area of ​​a hymenoid robot as described in Embodiment 1.

[0079] The present invention also provides an electronic device, including a processor and a memory connected to the processor, the memory storing a plurality of instructions which can be loaded and executed by the processor to enable the processor to perform a multimodal large model training method for generating operable areas of a body robot as described in Embodiment 1.

[0080] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its spirit and scope. Thus, if these modifications and modifications of the invention fall within the scope of the claims and their equivalents, the invention is also intended to include these modifications and modifications.

Claims

1. A method for training a multimodal large model for generating the operable area of ​​an embodied robot, characterized in that, include: Acquire multimodal data of the embodied robot and instruction data for calling expert models; The multimodal large model is pre-trained using the aforementioned multimodal data to obtain a pre-trained multimodal large model; The pre-trained multimodal large model is fine-tuned using the instruction data of the expert model to obtain a trained multimodal large model. The trained multimodal large model can then use the input multimodal data of the embodied robot to output the semantic relationships between the operable key object parts of the embodied robot and the operable key object parts, or output the semantic relationships between the operable key object parts of the embodied robot and the operable key object parts, and output visualized operable key points and the position coordinates of the operable key points by calling the expert model. The multimodal data includes visual data, language command data, and robot body and sensor data. The expert model is used to support the accurate extraction of operable regions by the multimodal large model; The expert model has the function of localization and segmentation based on image and language descriptions; the expert model includes an image segmentation model and an image localization model. The multimodal large model adopts the LLaVA multimodal large model; When the trained multimodal large model receives multimodal data, it outputs the corresponding results depending on whether the expert model is invoked. If the expert model is not invoked, it only outputs the answer based on the multimodal large model, including the operable key object parts and the semantic relationships between them. If the expert model is invoked, it outputs the answer based on the multimodal large model and invokes the expert model. The expert model uses the operable key object parts and the semantic relationships between them in the answer to perform image segmentation and localization, and outputs the operable key points and their position coordinates.

2. The multimodal large model training method for generating the operable area of ​​a unibody robot as described in claim 1, characterized in that, It also includes the step of aligning the multimodal data of the embodied robot.

3. The multimodal large model training method for generating the operable area of ​​a unibody robot as described in claim 1, characterized in that, Obtaining the instruction data for calling the expert model includes: generating instruction data for calling the expert model using the generative large model GPT-4.

4. The multimodal large model training method for generating the operable area of ​​a unibody robot as described in claim 1, characterized in that, The visual data includes images, videos, and point clouds.

5. A method for generating the operable area of ​​a embodied robot based on a multimodal large model, characterized in that, include: The multimodal data of the embodied robot is input into the trained multimodal large model, and the semantic relationship between the operable key object parts and the operable key object parts is output, or the semantic relationship between the operable key object parts and the operable key object parts is output, and the visualized operable key points and the position coordinates of the operable key points are output by calling the expert model. The trained multimodal large model is pre-trained using the method described in any one of claims 1-4.

6. A multimodal large-scale model training device for generating operable areas of a embodied robot, characterized in that, include: The training data acquisition module is used to acquire multimodal data of the embodied robot and instruction data for calling the expert model; The pre-training module is used to pre-train the multimodal large model using the multimodal data to obtain the pre-trained multimodal large model. The fine-tuning module is used to fine-tune the pre-trained multimodal large model using the instruction data of the expert model to obtain the trained multimodal large model. So that the trained multimodal large model can use the input multimodal data of the embodied robot to output the semantic relationship between the operable key object parts of the embodied robot and the operable key object parts, or output the semantic relationship between the operable key object parts of the embodied robot and the operable key object parts, and output the visualized operable key points and the position coordinates of the operable key points by calling the expert model. The multimodal data includes visual data, language command data, and robot body and sensor data. The expert model is used to support the accurate extraction of operable regions by the multimodal large model; The expert model has the function of localization and segmentation based on image and language descriptions; the expert model includes an image segmentation model and an image localization model. The multimodal large model adopts the LLaVA multimodal large model; When the trained multimodal large model receives multimodal data, it outputs the corresponding results depending on whether the expert model is invoked. If the expert model is not invoked, it only outputs the answer based on the multimodal large model, including the operable key object parts and the semantic relationships between them. If the expert model is invoked, it outputs the answer based on the multimodal large model and invokes the expert model. The expert model uses the operable key object parts and the semantic relationships between them in the answer to perform image segmentation and localization, and outputs the operable key points and their position coordinates.

7. A memory, characterized in that, The system stores multiple instructions for implementing the multimodal large model training method for generating the operable area of ​​a embodied robot as described in any one of claims 1-4.

8. An electronic device, characterized in that, The system includes a processor and a memory connected to the processor, the memory storing multiple instructions that can be loaded and executed by the processor to enable the processor to perform the multimodal large model training method for generating operable areas of a body robot as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Multi-modal large model robot control method based on meta-learning fine tuning

    CN119610132A