Body-equipped robot operation track generation method and system based on multi-modal large model

Through the multimodal large model training of vision, language and robot data, the embodied robot operation trajectory is generated, which solves the problems of low generalization and low precision in the existing technology, and achieves more efficient and adaptable robot operation.

CN120023826AActive Publication Date: 2025-05-23BEIJING MALIMAO TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510397773.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-05-23
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The prior art has low generalization and low precision in operation in embodied robot operations, making it difficult to adapt to changing environments and task requirements.

Method used

By obtaining visual input data, language input data and robot input data, preprocessing and mapping, and training with multimodal large models, robot operation trajectory is generated to achieve simultaneous generation of text planning and operation trajectory.

Benefits of technology

It improves the accuracy and efficiency of robot operation, enhances environmental understanding and adaptability, reduces programming and control difficulties, and improves human-computer interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120023826A_ABST
    Figure CN120023826A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a system for generating an operation track of a robot with a body based on a multi-modal large model, and belongs to the technical field of artificial intelligence. The method comprises the steps that vision, language and robot input data are obtained, preprocessed and then mapped to a high-dimensional space, and vision, language and robot input vectors are obtained; respectively inputting the input vectors into an original multi-modal large model and an expert model for training to obtain a preliminarily trained multi-modal large model and a trained expert model; generating calling data for the trained expert model by using the generative large model, forming an instruction fine tuning data set, inputting the instruction fine tuning data set into the preliminarily trained multi-modal large model for training, and obtaining a finally trained multi-modal large model; and generating a robot operation track based on robot operation data and the finally trained multi-modal large model. The problems that an existing robot is low in operation generalization, low in precision and difficult to adapt to variable environments and task requirements are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and system for generating an embodied robot operation trajectory based on a multimodal large model. Background Art

[0002] In recent years, multimodal large models have been widely used in the field of embodied robots, especially in robot task planning and execution. However, there are two main strategies in the existing technology: one is to directly output text planning (Text Planning), and the other is to directly output actions (Action). Although the method of directly outputting text planning can generate high-level task planning with the help of the powerful reasoning ability of large models, it often lacks applicability in specific real-world scenarios. For example, when the instruction requires "tidying up the desktop", the text planning may generate a high-level task planning similar to "organizing the items". However, the robot cannot determine the specific operation steps through this general instruction, such as how to classify items, what order to execute, or what tools to use. This makes the robot lack targeted operation guidance when performing tasks, and it is difficult to complete the details of the task. On the other hand, although the method of directly outputting actions can generate specific operation instructions, it largely loses the reasoning ability of the large model, resulting in the lack of generalization of the generated operations and difficulty in adapting to changing environments and task requirements.

[0003] Therefore, in order to address the shortcomings of these two extreme strategies, there is an urgent need for a method that can find a balance between the reasoning ability of large models and the execution of specific operations, so as to retain the generalization ability of large models in environmental understanding and task planning while providing sufficiently detailed guidance for downstream models to support robots in performing fine-grained operations. Summary of the invention

[0004] In order to solve the problems existing in the prior art, the present invention provides the following technical solutions.

[0005] A first aspect of the present invention provides a method for generating an embodied robot operation trajectory based on a multimodal large model, comprising:

[0006] Obtaining visual input data, language input data and robot input data and preprocessing them, aligning the preprocessed visual input data, language input data and robot input data and mapping them to high-dimensional space respectively, to obtain a visual input vector, a language input vector and a robot input vector;

[0007] Inputting the visual input vector, the language input vector and the robot input vector into the original multimodal large model and the original expert model for training respectively, to obtain a preliminarily trained multimodal large model and a trained expert model;

[0008] Use the generative big model to generate call data for the trained expert model and form a standardized instruction fine-tuning dataset:

[0009] Inputting the instruction fine-tuning data set into the preliminarily trained multimodal large model for training to obtain a final trained multimodal large model;

[0010] Generate the robot operation trajectory based on the robot operation data and the final trained multimodal large model.

[0011] Preferably, preprocessing the acquired visual input data, language input data and robot input data includes: cleaning the acquired visual input data, language input data and robot input data.

[0012] Preferably, aligning the preprocessed visual input data, language input data and robot input data and mapping them to high-dimensional space respectively includes: aligning the preprocessed visual input data, language input data and robot input data and mapping them to high-dimensional space through a multi-layer perceptron respectively to obtain a visual input vector, a language input vector and a robot input vector.

[0013] Preferably, the visual input data includes picture data, point cloud data and video data; the language input data includes language instruction data for the robot; and the robot input data includes robot body parameters and sensor parameters.

[0014] Preferably, the original multimodal large model is an LLaVA model.

[0015] Preferably, generating the robot operation trajectory includes generating a robot visualization trajectory, robot operation trajectory coordinates and robot operation instruction planning.

[0016] Preferably, the robot visualization trajectory and the robot operation trajectory coordinates are generated and output by a trained expert model, and the robot operation instruction planning is generated and output by a finally trained multimodal large model.

[0017] A second aspect of the present invention provides an embodied robot operation trajectory generation system based on a multimodal large model, comprising:

[0018] A multimodal data acquisition and mapping module, used to acquire and preprocess visual input data, language input data and robot input data, align the preprocessed visual input data, language input data and robot input data and map them to a high-dimensional space to obtain a visual input vector, a language input vector and a robot input vector;

[0019] The first model training module is used to input the visual input vector, the language input vector and the robot input vector into the original multimodal large model and the original expert model for training, so as to obtain a preliminarily trained multimodal large model and a trained expert model;

[0020] The expert model call data generation module is used to generate call data for the trained expert model using the generative large model and form a standardized instruction fine-tuning data set;

[0021] The second model training module is used to input the instruction fine-tuning data set into the preliminarily trained multimodal large model for training to obtain a final trained multimodal large model;

[0022] The operation trajectory generation module is used to generate the robot operation trajectory based on the robot operation data and the final trained multimodal large model.

[0023] A third aspect of the present invention provides a memory storing a plurality of instructions, wherein the instructions are used to implement the method for generating an embodied robot operation trajectory based on a multimodal large model as described in the first aspect.

[0024] The fourth aspect of the present invention provides an electronic device, comprising a processor and a memory connected to the processor, the memory storing a plurality of instructions, the instructions being loaded and executed by the processor so that the processor can execute the method for generating operation trajectories of an embodied robot based on a multimodal large model as described in the first aspect.

[0025] The beneficial effects of the present invention are as follows: the method and system for generating operation trajectories of embodied robots based on a multimodal large model provided by the present invention, by using multimodal data such as vision, language, and robots to train the multimodal large model, enables the trained multimodal large model to simultaneously perform text planning and operation trajectory generation for robot operations, effectively solving the problems of low generalization, low precision, and difficulty in adapting to changing environments and task requirements of embodied robot operations in the prior art, and provides solid technical support for the operation of embodied robots in changing environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a flow chart of the method for generating operation trajectory of an embodied robot based on a multi-modal large model according to the present invention;

[0027] Figure 2 This is a functional structure diagram of the embodied robot operation trajectory generation system based on the multimodal large model described in the present invention. DETAILED DESCRIPTION

[0028] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0029] The method provided by the present invention can be implemented in the following terminal environment, and the terminal may include one or more of the following components: a processor, a memory, and a display screen. The memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the method described in the following embodiment.

[0030] The processor may include one or more processing cores. The processor uses various interfaces and lines to connect various parts of the entire terminal, and executes various functions of the terminal and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory, and calling data stored in the memory.

[0031] The memory may include random access memory (RAM) or read-only memory (ROM). The memory may be used to store instructions, programs, codes, code sets or instructions.

[0032] The display screen is used to display the user interface of each application.

[0033] In addition, those skilled in the art can understand that the structure of the above terminal does not constitute a limitation on the terminal, and the terminal may include more or fewer components, or combine certain components, or arrange the components differently. For example, the terminal also includes components such as a radio frequency circuit, an input unit, a sensor, an audio circuit, and a power supply, which will not be described in detail here.

[0034] Embodiment 1

[0035] like Figure 1 As shown, an embodiment of the present invention provides a method for generating an embodied robot operation trajectory based on a multimodal large model, comprising:

[0036] S101, obtaining visual input data, language input data and robot input data and preprocessing them, aligning the preprocessed visual input data, language input data and robot input data and mapping them to a high-dimensional space respectively, to obtain a visual input vector, a language input vector and a robot input vector;

[0037] S102, inputting the visual input vector, the language input vector and the robot input vector into the original multimodal large model and the original expert model for training, respectively, to obtain a preliminarily trained multimodal large model and a trained expert model;

[0038] S103, using the generative large model to generate call data for the trained expert model, and forming a standardized instruction fine-tuning data set;

[0039] S104, inputting the instruction fine-tuning data set into the preliminarily trained multimodal large model for training, to obtain a final trained multimodal large model;

[0040] S105, generating a robot operation trajectory based on the robot operation data and the finally trained multimodal large model.

[0041] Specifically, in step S101, the visual input data includes image data, point cloud data and video data; the language input data includes language instruction data for the robot, such as the language instruction issued to the robot "Please pour the water in the teapot into the cup, be careful not to knock over the surrounding things"; the robot input data includes robot body parameters (robot model, robot arm type, etc.), sensor parameters, etc. Preprocessing the acquired visual input data, language input data and robot input data includes: cleaning the acquired visual input data, language input data and robot input data, so as to ensure the quality of the acquired multimodal data. After aligning the pre-processed visual input data, language input data and robot input data, they are mapped to high-dimensional space respectively to obtain visual input vectors, language input vectors and robot input vectors, including: after aligning the pre-processed visual input data, language input data and robot input data, they are mapped to high-dimensional space respectively through a multi-layer perceptron to obtain visual input vectors, language input vectors and robot input vectors that meet the processing dimension requirements of the multimodal large model. It should be noted that, for example, the visual input data obtained by the visual encoder is mapped through a visual mapping layer to obtain a visual input vector, the language input data is mapped through a text mapping layer to obtain a language input vector, and the robot input data is mapped through an ontology mapping layer to obtain a robot input vector. In this step, by fusing the multimodal data such as visual input data, language input data and robot input data, the reasoning ability, scene perception ability and generalization ability of the multimodal large model in the field of robots can be comprehensively improved.

[0042] In step S102, the original multimodal large model can select the LLaVA model, which is a multimodal large model designed to integrate and understand multiple data types including text, pictures, and videos; the original expert model can select the DINOv2 model, SAM model, etc. It should be noted that expert model is one of the terms in the field of artificial intelligence. In this application, the expert model generates operation trajectories such as robot visualization trajectory and robot operation trajectory coordinates based on the text planning (i.e., robot operation instruction planning) output by the multimodal large model, thereby achieving the goal of simultaneously outputting text planning and operation trajectory.

[0043] In step S103, the generative big model can use GPT-4. The robot input vector is input into the original multimodal big model for training, so that the model can generate a real and credible operable area based on the current robot body, ensuring that the model has a more accurate understanding of its own capabilities and limitations, thereby improving the executability and accuracy of the operation.

[0044] In step S104, the instruction fine-tuning data set obtained in step S103 is input into the preliminarily trained multimodal large model to continue training, mainly for enhancing the ability of the finally trained multimodal large model to actively call the trained expert model when facing the object operation trajectory generation task.

[0045] The robot operation data in step S105 is consistent with the visual input data, language input data and robot input data format of step S101, and can be used as test data to test the finally trained multimodal large model; in addition, the robot operation data can also be used as an evaluation indicator during testing to judge the generation effect of the operation trajectory of the finally trained multimodal large model; it should be noted that through detector tracking and environmental depth information estimation, a very accurate 3D operation trajectory can be obtained, and the robot operation data is generated based on the 3D operation trajectory as the true value in the model training process; generating the robot operation trajectory includes generating a robot visualization trajectory, a robot operation trajectory coordinate and a robot operation instruction plan, the generated robot visualization trajectory and the robot operation trajectory coordinates are generated and output by the trained expert model, and the robot operation instruction plan is generated and output by the finally trained multimodal large model.

[0046] This application trains a large multimodal model and generates text planning of the operation steps of the embodied robot and the operation trajectory corresponding to the text planning. It can use both the high-level general guidance of the text planning and the low-level fine-grained guidance of the operation trajectory. Through the above method, this application can achieve the following effects:

[0047] (1) Improving operational accuracy: Using multimodal data to enhance the understanding of scenes and target objects can generate more accurate operational trajectories and improve the success rate of robots in delicate tasks (such as assembly and grasping).

[0048] (2) Improve operational efficiency: By integrating multimodal data (such as vision, language, and robotics) and the powerful learning capabilities of large models, the robot's operation trajectory can be generated more efficiently, reducing the time for trajectory planning and improving the robot's response speed in actual tasks;

[0049] (3) Enhanced intelligent decision-making capabilities: Based on multimodal data, task scenarios can be analyzed from multiple perspectives, including vision, language, and robotics, enabling robots to have a higher level of environmental understanding, thereby improving their intelligent decision-making capabilities in complex environments;

[0050] (4) Strong adaptability: Through large model learning and fusion of multimodal data, it can better cope with diverse tasks and environmental changes, and has strong adaptability, thereby improving the task execution ability of embodied robots. It has broad application prospects in multiple industries, such as intelligent manufacturing, home services, and medical assistance.

[0051] (5) It can reduce the difficulty of programming and control: The multimodal large model reduces the dependence on programming and control by traditional robot engineers through automatic learning and generation of operation trajectories, allowing non-professionals to perform simple robot control through natural language instructions;

[0052] (6) Improving the human-computer interaction experience: By combining multimodal inputs such as language and vision, robots can understand human intentions more naturally, thereby achieving a smoother interactive experience, which enables robots to better collaborate with humans;

[0053] (7) Task automation and generalization capabilities: Multimodal large models have powerful task transfer and generalization capabilities. They can apply generated trajectories to new tasks by learning from existing data, thereby reducing the need for retraining and debugging.

[0054] Embodiment 2

[0055] like Figure 2 As shown, another aspect of the present invention also includes a functional module architecture that is completely consistent with the aforementioned method flow, that is, an embodiment of the present invention also provides an embodied robot operation trajectory generation system based on a multimodal large model, including:

[0056] The multimodal data acquisition and mapping module 201 is used to acquire and preprocess visual input data, language input data and robot input data, align the preprocessed visual input data, language input data and robot input data, and map them to a high-dimensional space to obtain a visual input vector, a language input vector and a robot input vector;

[0057] The first model training module 202 is used to input the visual input vector, the language input vector and the robot input vector into the original multimodal large model and the original expert model for training, so as to obtain a preliminarily trained multimodal large model and a trained expert model;

[0058] The expert model call data generation module 203 is used to generate call data for the trained expert model using the generative large model and form a standardized instruction fine-tuning data set;

[0059] The second model training module 204 is used to input the instruction fine-tuning data set into the preliminarily trained multimodal large model for training to obtain a final trained multimodal large model;

[0060] The operation trajectory generation module 205 is used to generate the robot operation trajectory based on the robot operation data and the finally trained multi-modal large model.

[0061] The system can be implemented through the method for generating operation trajectories of an embodied robot based on a multimodal large model provided in the first embodiment. The specific implementation method can be found in the description of the first embodiment and will not be repeated here.

[0062] The present invention also provides a memory storing a plurality of instructions, wherein the instructions are used to implement the method for generating an embodied robot operation trajectory based on a multimodal large model as described in Example 1.

[0063] The present invention also provides an electronic device, comprising a processor and a memory connected to the processor, wherein the memory stores a plurality of instructions, and the instructions can be loaded and executed by the processor so that the processor can execute the method for generating an embodied robot operation trajectory based on a multimodal large model as described in Example 1.

[0064] Although preferred embodiments of the present invention have been described, additional changes and modifications may be made to these embodiments by those skilled in the art once the basic inventive concepts are known. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention. Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A method for generating operation trajectories of an embodied robot based on a multimodal large model, characterized in that: include: Obtaining visual input data, language input data and robot input data and preprocessing them, aligning the preprocessed visual input data, language input data and robot input data and mapping them to high-dimensional space respectively, to obtain a visual input vector, a language input vector and a robot input vector; Inputting the visual input vector, the language input vector and the robot input vector into the original multimodal large model and the original expert model for training respectively, to obtain a preliminarily trained multimodal large model and a trained expert model; Use the generative big model to generate call data for the trained expert model and form a standardized instruction fine-tuning dataset; Inputting the instruction fine-tuning data set into the preliminarily trained multimodal large model for training to obtain a final trained multimodal large model; Generate the robot operation trajectory based on the robot operation data and the final trained multimodal large model.

2. The method for generating operation trajectories of an embodied robot based on a multimodal large model as claimed in claim 1, characterized in that: Preprocessing the acquired visual input data, language input data and robot input data includes: cleaning the acquired visual input data, language input data and robot input data.

3. The method for generating operation trajectories of an embodied robot based on a multimodal large model as claimed in claim 1, characterized in that: The step of aligning the preprocessed visual input data, language input data and robot input data and mapping them to high-dimensional space includes: aligning the preprocessed visual input data, language input data and robot input data and mapping them to high-dimensional space through a multi-layer perceptron.

4. The method for generating operation trajectories of an embodied robot based on a multimodal large model as claimed in claim 1, characterized in that: The visual input data includes picture data, point cloud data and video data; the language input data includes language instruction data for the robot; and the robot input data includes robot body parameters and sensor parameters.

5. The method for generating operation trajectories of an embodied robot based on a multimodal large model as claimed in claim 1, characterized in that: The original multimodal large model is the LLaVA model.

6. The method for generating operation trajectories of an embodied robot based on a multimodal large model as claimed in claim 1, characterized in that: The generating of the robot operation trajectory includes generating the robot visualization trajectory, the robot operation trajectory coordinates and the robot operation instruction planning.

7. The method for generating operation trajectories of an embodied robot based on a multimodal large model as claimed in claim 6, characterized in that: The robot visualization trajectory and robot operation trajectory coordinates are generated and output by the trained expert model, and the robot operation instruction planning is generated and output by the finally trained multi-modal large model.

8. A system for generating operation trajectories of an embodied robot based on a multimodal large model, characterized in that: include: A multimodal data acquisition and mapping module, used to acquire and preprocess visual input data, language input data and robot input data, align the preprocessed visual input data, language input data and robot input data and map them to a high-dimensional space to obtain a visual input vector, a language input vector and a robot input vector; The first model training module is used to input the visual input vector, the language input vector and the robot input vector into the original multimodal large model and the original expert model for training, so as to obtain a preliminarily trained multimodal large model and a trained expert model; The expert model call data generation module is used to generate call data for the trained expert model using the generative large model and form a standardized instruction fine-tuning data set; The second model training module is used to input the instruction fine-tuning data set into the preliminarily trained multimodal large model for training to obtain a final trained multimodal large model; The operation trajectory generation module is used to generate the robot operation trajectory based on the robot operation data and the final trained multimodal large model.

9. A memory, characterized in that: A plurality of instructions are stored, and the instructions are used to implement the method for generating operation trajectories of an embodied robot based on a multimodal large model as described in any one of claims 1-7.

10. An electronic device, characterized in that: It includes a processor and a memory connected to the processor, the memory stores a plurality of instructions, and the instructions can be loaded and executed by the processor, so that the processor can execute the method for generating operation trajectories of an embodied robot based on a multimodal large model as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Apparatus and method for robotic display choreography

    CA2684192A1

  • Method for kinematic calibration of six-degree-of-freedom robot based on monocular vision

    CN107175660A

  • Robot dynamic capture method and system

    CN107992881A

  • Physical scanning modeling method and application thereof

    CN108748137A

  • Robot action generation method and system combining universal and special models

    CN119550345A