Intelligent method of multi-mode large model with body
Through multimodal large model and SLAM algorithm, combined with large language model and robotic arm control, the adaptability problem of embodied intelligent systems in fuzzy instructions and unseen environments is solved, and efficient task execution in complex environments is achieved.
Patent Information
- Application Number
- CN202510269660.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
Existing embodied intelligent systems are difficult to cope with fuzzy instructions or unseen environments, and show poor adaptability and flexibility.
The embodied intelligent method of multimodal large model is adopted to analyze user instructions through the large language model to generate navigation and crawling tasks; three-dimensional map scene information is constructed based on point cloud data and camera data, and real-time positioning and path planning is used using SLAM algorithm; for the crawling task, the optimal crawling position and path are determined, and executed through the robotic arm control module.
Enables robots to understand fuzzy language instructions, adapt and react quickly, and perform multiple tasks in complex and dynamic environments, improving system adaptability and task completion rate.
Smart Images

Figure CN120197726A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to an embodied intelligence method for multi-modal large models. Background Art
[0002] In recent years, multi-modal large language models and robotics technology have developed rapidly, and embodied intelligence has become a new focus of global technology and industrial competition. Different from traditional "brain-type" intelligence, embodied intelligence emphasizes that robots obtain information, make decisions, and adapt to changes by interacting with the environment.
[0003] Traditional embodied intelligence systems rely on preset rules and instructions. Their working processes generally include: perceiving and understanding the environment through sensor data (such as visual, tactile, or sound signals); performing path planning or action design based on preset rules or task objectives; and completing specific operations through actuators. However, this method relies on prior knowledge of tasks and environments and is difficult to cope with environments with uncertainty, complexity, and dynamic changes, thus showing poor adaptability and flexibility. For example, traditional embodied intelligence systems will appear inadequate when dealing with unseen environments or ambiguous instructions.
[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide an embodied intelligence method for multi-modal large models in view of the above-mentioned defects of the existing technology, aiming to solve the problem that existing embodied intelligence systems are difficult to cope with ambiguous instructions or unseen environments.
[0006] The technical solution adopted by the present invention to solve the problem is as follows:
[0007] In a first aspect, an embodiment of the present invention provides an embodied intelligence method for multi-modal large models, and the method includes:
[0008] Obtain a user instruction, and generate a number of subtasks to be executed according to the user instruction and prompt data through a large language model; each of the subtasks includes a navigation task and / or a grasping task;
[0009] Obtain point cloud data and camera data of the current scene, and construct the three-dimensional map scene information and obtain real-time positioning according to the point cloud data and the camera data;
[0010] For the navigation task, through a preset navigation algorithm, generate a target path according to the navigation task, the three-dimensional map scene information, and the real-time positioning, and control the machine motion module based on the target path;
[0011] For the grasping task, based on a preset robotic arm control algorithm, determine the optimal grasping pose and grasping path for each target object to be grasped according to the grasping task and the three-dimensional map scene information, and control the robotic grasping module based on the optimal grasping pose and the grasping path.
[0012] In one implementation, use a large language model to generate several subtasks to be executed according to the user instruction and the prompt data, including:
[0013] Use a large language model to parse the semantic information corresponding to the user instruction according to the user instruction and the prompt data;
[0014] Determine the target task according to the semantic information, and generate several subtasks to be executed according to the three-dimensional map scene information and the target task.
[0015] In one implementation, determine the target task according to the semantic information, and generate several subtasks to be executed according to the three-dimensional map scene information and the target task, including:
[0016] Determine the target task according to the semantic information, and obtain matching memory data from the vector memory database according to the three-dimensional map scene information and the target task, where the memory data is used to reflect historical three-dimensional map scene information and / or historical task data;
[0017] Generate several subtasks to be executed according to the target task, the three-dimensional map scene information, and the memory data.
[0018] In one implementation, construct the three-dimensional map scene information and obtain real-time positioning according to the point cloud data and the camera data, including:
[0019] Through a SLAM algorithm based on 3D Gaussian sputtering, perform three-dimensional mapping and positioning according to the point cloud data and the camera data to obtain the three-dimensional map scene information and the real-time positioning of the current scene.
[0020] In one implementation, generate a target path according to the navigation task, the three-dimensional map scene information, and the real-time positioning through a preset navigation algorithm, including:
[0021] Determine the target position according to the navigation task;
[0022] Plan a target path from the real-time positioning to the target position through a preset navigation algorithm and the three-dimensional map scene information.
[0023] In one implementation, determine the optimal grasping pose and grasping path for each target object to be grasped according to the grasping task and the three-dimensional map scene information, including:
[0024] Determine a number of target objects to be grasped according to the grasping task;
[0025] For each of the target objects, generate a number of potential grasping points for the target object according to the three-dimensional map scene information, and calculate the success rate corresponding to each of the potential grasping points;
[0026] Determine the optimal grasping pose and grasping path of the target object according to the success rate of each of the potential grasping points.
[0027] In one implementation, the method further includes:
[0028] For each of the target objects, analyze the attribute information of the object, and adjust the force control strategy of the robotic arm according to the attribute information of the object.
[0029] In a second aspect, an embodiment of the present invention further provides an embodied intelligent system for a multimodal large model, the system includes:
[0030] A large language model, which generates a number of subtasks to be executed according to the user instruction and the prompt data through the large language model; each of the subtasks includes a navigation task and / or a grasping task;
[0031] A positioning and mapping module, configured to obtain point cloud data and camera data of the current scene, and construct the three-dimensional map scene information and obtain real-time positioning according to the point cloud data and the camera data;
[0032] A navigation decision module, configured to, for the navigation task, generate a target path through a preset navigation algorithm according to the navigation task, the three-dimensional map scene information, and the real-time positioning;
[0033] A grasping decision module, configured to, for the grasping task, determine the optimal grasping pose and grasping path of each target object to be grasped according to the grasping task and the three-dimensional map scene information through a preset robotic arm control algorithm;
[0034] A machine motion module, configured to execute the target path;
[0035] A machine grasping module, configured to execute the optimal grasping pose and the grasping path.
[0036] In a third aspect, an embodiment of the present invention further provides a terminal, the terminal includes a memory and more than one processor; the memory stores more than one program; the program includes instructions for executing the embodied intelligent method of the multimodal large model as described in any one of the above; the processor is configured to execute the program.
[0037] Fourthly, an embodiment of the present invention also provides a computer-readable storage medium, on which multiple instructions are stored. The instructions are suitable for being loaded and executed by a processor to implement the steps of the embodied intelligence method of the multi-modal large model as described in any one of the above.
[0038] Advantages of the present invention: By embedding a large language model in the embodiments of the present invention, which has particular advantages in dealing with language uncertainty and environmental complexity, the robot is enabled to understand ambiguous language instructions and helps the system quickly adapt and respond, so that it can execute various tasks in complex and dynamic environments. For example, when the robot receives an ambiguous verbal instruction, it can perform semantic parsing through the large model and use tools such as a robotic arm to complete specific operations, such as picking up an object to open a door or other similar tasks. Description of the Drawings
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0040] Figure 1 It is a schematic flowchart of the embodied intelligence method of the multi-modal large model provided by the embodiment of the present invention.
[0041] Figure 2 It is a schematic diagram of a quadruped robot product provided by the embodiment of the present invention.
[0042] Figure 3 It is a schematic diagram of a robotic arm product provided by the embodiment of the present invention.
[0043] Figure 4 It is a schematic diagram of the overall framework provided by the embodiment of the present invention.
[0044] Figure 5 It is a schematic diagram of the quadruped robot control design provided by the embodiment of the present invention.
[0045] Figure 6 It is a schematic diagram of the modules of the embodied intelligence system of the multi-modal large model provided by the embodiment of the present invention.
[0046] Figure 7 It is a schematic block diagram of the terminal provided by the embodiment of the present invention. Detailed Embodiments
[0047] The present invention discloses an embodied intelligence method for multimodal large models. To make the objectives, technical solutions, and effects of the present invention clearer and more explicit, the following further elaborates on the present invention with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0048] Those skilled in the art of this technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0049] Those skilled in the art of this technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention pertains. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as herein.
[0050] In view of the above defects of the prior art, the present invention provides an embodied intelligence method for a multimodal large model. The method obtains a user instruction, and uses a large language model to generate a number of subtasks to be executed according to the user instruction and prompt data; each of the subtasks includes a navigation task and / or a grasping task; obtains point cloud data and camera data of the current scene, constructs the three-dimensional map scene information according to the point cloud data and the camera data, and obtains real-time positioning; for the navigation task, through a preset navigation algorithm, generates a target path according to the navigation task, the three-dimensional map scene information, and the real-time positioning, and controls the machine motion module based on the target path; for the grasping task, through a preset robotic arm control algorithm, determines the optimal grasping pose and grasping path of each target object to be grasped according to the grasping task and the three-dimensional map scene information, and controls the machine grasping module based on the optimal grasping pose and the grasping path. By embedding a large language model, which has particular advantages in dealing with language uncertainty and environmental complexity, the present invention enables the robot to have the ability to understand ambiguous language instructions, and helps the system to quickly adapt and respond, so that it can perform multiple tasks in complex and dynamic environments. For example, when the robot receives an ambiguous verbal instruction, it can perform semantic parsing through the large model and use tools such as a robotic arm to complete specific operations, such as picking up an object to open a door or other similar tasks.
[0051] As Figure 1 shown, the method includes:
[0052] Step S100, obtain a user instruction, and use a large language model to generate a number of subtasks to be executed according to the user instruction and prompt data; each of the subtasks includes a navigation task and / or a grasping task.
[0053] Specifically, the user instruction can be a task description in natural language form, which reflects the ultimate goal or desired operation that the user wants to achieve. The large language model will analyze according to the received user instruction and prompt data (Prompt), and decompose the task assigned by the user instruction into a number of subtasks to be executed. The prompt data can be a carefully designed text prompt to guide the large language model to output the required content. The types of subtasks include, but are not limited to, navigation tasks and grasping tasks. If the user instruction involves moving between different locations or locating the position of an object, the decomposed subtasks may include a navigation task. If the user instruction has a requirement to operate a specific object, the decomposed subtasks may include a grasping task. The relevant system of this embodiment can be constructed by an agent. By introducing an advanced large language model, the understanding depth and parsing accuracy of the agent for ambiguous and complex natural language instructions can be enhanced, enabling it to extract the core intention from a higher-level semantics, efficiently convert the language instruction into an executable operation, and at the same time perform task decomposition.
[0054] In one implementation, the large language model generates a number of subtasks to be executed based on the user instruction and the prompt data, including:
[0055] The large language model parses the semantic information corresponding to the user instruction according to the user instruction and the prompt data;
[0056] Determine the target task according to the semantic information, and generate a number of subtasks to be executed according to the three-dimensional map scene information and the target task.
[0057] The large language model (LLM) in this embodiment is responsible for performing the core task understanding and distribution functions. Its main function is to receive natural language instructions from humans and parse the semantic content of complex instructions through carefully designed prompt data. This parsing ability can not only process highly abstract semantics, but also identify the ambiguity in the instructions and interpret it through the context. After parsing, the large language model decomposes the task into multiple executable subtasks and distributes them to the subsequent intermediate layer modules for execution. In order to adapt to complex and changing scenarios, the large language model also has the ability of real-time feedback, dynamically adjusting according to the task execution status, so as to ensure that the system can flexibly handle the needs of uncertainty and multi-task interaction.
[0058] For example, the large language model performs lexical analysis on the user instruction through the prompt data, decomposes the sentence corresponding to the user instruction into several words, and marks the part of speech of each word (noun, verb, adjective, etc.). Identify the relationship between words and the sentence structure according to the part of speech of each word, such as the subject-predicate-object structure, etc. Determine several semantic roles according to the identified word relationship and sentence structure, such as the subject performing the action, the object of the action, time, place, etc. Determine the semantic information of the user instruction according to each semantic role, and determine an overall target task according to the semantic information, covering the core requirements of the user instruction. In order to better complete the target task, decompose the target task according to the logical order and execution steps of the task, and combine the current three-dimensional map scene information to obtain several more specific and operable subtasks, so as to gradually parse and refine the complex user instruction and more effectively complete the complete target task.
[0059] In one implementation, determine the target task according to the semantic information, and generate a number of subtasks to be executed according to the three-dimensional map scene information and the target task, including:
[0060] Determine the target task according to the semantic information, and obtain matching memory data from the vector memory database according to the three-dimensional map scene information and the target task, where the memory data is used to reflect historical three-dimensional map scene information and / or historical task data;
[0061] A plurality of subtasks to be executed are generated according to the target task, the three-dimensional map scene information and the memory data.
[0062] In order to demonstrate higher intelligence and efficiency when performing long-term tasks, this embodiment uses a vector memory database to build a long-term semantic memory capability to store data related to long-term tasks. The vector memory database stores data in the form of vectors, which can better capture the semantic relationship between data. This embodiment defines the data stored in the vector memory database as memory data, which is mainly used to reflect past map scene data and / or past task data, where the map scene data can be generated based on a visual model. The large language model can quickly retrieve past tasks and observations through the vector memory database, thereby giving the system / agent stronger semantic reasoning and long-term decision-making capabilities, enabling it to make effective associations based on past experience and current task requirements, and improve its ability to handle complex, multi-step tasks. Suitable for scenarios that require long-term deployment, such as disaster relief or industrial inspections.
[0063] In actual application scenarios, when the system receives a new user instruction, it determines the target task based on the semantic information corresponding to the new user instruction, and initiates a query request to the vector memory database through the current 3D map scene information and the target task. The vector memory database uses the current 3D map scene information and the target task as query conditions to retrieve and return the matching memory data. The system uses the returned memory data as a reference, combined with the current 3D map scene information and the target task, to gradually decompose the target task into multiple subtasks to be executed.
[0064] For example, the current three-dimensional map scene information is Factory A, and the target task is "to find specific parts in the warehouse area of Factory A and move them to the assembly area". The system may obtain the historical map scene data of Factory A and the historical task data of parts handling in the warehouse area of Factory A from the vector memory database. The historical task data contains relevant information about various tasks that have been performed in the past, such as task objectives, execution steps, problems encountered and solutions. These data can help the system understand the handling methods and experience of similar tasks. For example, there have been many parts handling tasks in Warehouse A before, and the historical task data can record the best path for each handling and the obstacles that may be encountered. When a new handling task arrives, the system can refer to these historical task data to optimize the decomposition process of the current target task and the construction process of each subtask.
[0065] In one implementation, a plurality of subtasks to be executed are generated according to the target task, the three-dimensional map scene information and the memory data, including:
[0066] Determine an initial spatial partitioning method and an initial operation partitioning method according to the memory data;
[0067] Adjust the initial spatial partitioning method according to the three-dimensional map scene information to obtain a target spatial partitioning method;
[0068] Adjust the initial operation partitioning method according to the target task to obtain a target operation partitioning method;
[0069] Perform spatial decomposition on the three-dimensional map scene information according to the target spatial partitioning method;
[0070] Perform a first decomposition on the target task according to the spatial decomposition result to obtain a number of initial subtasks;
[0071] Perform a second decomposition on each of the initial subtasks according to the target operation partitioning method to obtain a number of subtasks.
[0072] Specifically, the initial space partitioning method refers to the space partitioning method for similar three-dimensional map scenarios obtained based on memory data before task decomposition, considering factors such as the functions of different regions and the distribution of objects. The initial operation partitioning method refers to the arrangement of operation steps for similar tasks obtained based on memory data, reflecting the overall process and main operation links of the tasks. By referring to the memory data in this embodiment, the subsequent task decomposition efficiency can be greatly improved. The three-dimensional map scenario information provides specific information about the current environment, including the actual object positions, obstacle distributions, and spatial dimensions, etc., which may be different from the situations in the memory data. Therefore, it is necessary to make corresponding adjustments to the initial space partitioning method based on the three-dimensional map scenario information. The adjustment process includes: comparing the three-dimensional map scenario information with the initial space partitioning method, determining the abnormal partitioning positions therein in combination with the preset space partitioning constraint conditions, and correcting each abnormal partitioning position to obtain the target space partitioning method that conforms to the current environment. The target task has specific requirements and conditions, which may be different from the historical tasks in the memory data. Therefore, it is necessary to optimize and adjust the initial operation partitioning method according to the specific content of the target task to ensure that the task can be successfully completed. The adjustment process includes: analyzing the completion objectives and constraint conditions of the target task, and modifying and supplementing the initial operation partitioning method according to the completion objectives and constraint conditions to obtain the target operation partitioning method suitable for the current task. Then, the three-dimensional map scenario information is further refined into smaller spatial units according to the target space partitioning method. The target task is then decomposed into several initial subtasks related to different spatial units. In other words, each subtask has its specific spatial range. The target operation partitioning method provides more specific operation guidance for each initial subtask. Therefore, according to the target operation partitioning method, each initial subtask can be further refined and decomposed into smaller and more easily executable subtasks to ensure that each operation link can be accurately executed.
[0073] Furthermore, the method further includes:
[0074] Calculating the similarity degree between each pair of the subtasks, and merging the subtasks with a similarity degree higher than a preset threshold.
[0075] Specifically, in this embodiment, similar or duplicate subtasks will also be merged to reduce redundant operations. And the execution priorities of each subtask can also be determined according to the task dependency relationships and resource constraints. In addition, the execution results of the subtasks (such as success, failure, time consumption, etc.) can be fed back to the vector memory database, so as to optimize the memory data according to the feedback results and improve the accuracy of future task segmentation.
[0076] Step S200: Obtain the point cloud data and camera data of the current scene, and construct the three-dimensional map scene information and obtain real-time positioning based on the point cloud data and the camera data.
[0077] Specifically, the point cloud data can be obtained by a lidar, which can provide rich environmental geometric information to help the system accurately perceive the three-dimensional structure of the scene, identify the shape, size, and positional relationship of objects. The camera data can be obtained by shooting with a camera module. The camera data contains the visual information of the scene, such as color, texture, and the outline of objects, which can provide rich visual details and help identify the category of objects. The point cloud data and the camera data have different characteristics and advantages. To describe the scene more comprehensively and accurately, in this embodiment, the two types of data will be combined to jointly construct the three-dimensional map scene information of the current environment. For example, through the feature matching method, corresponding feature points can be found in the point cloud data and the camera data, and then the information of these feature points can be associated and integrated to obtain more accurate scene information for constructing more accurate three-dimensional map scene information. Real-time positioning refers to determining the current position and / or pose (direction) of the device in the three-dimensional map scene.
[0078] In one implementation manner, constructing the three-dimensional map scene information and obtaining real-time positioning based on the point cloud data and the camera data includes:
[0079] Perform three-dimensional mapping and positioning according to the point cloud data and the camera data through the SLAM algorithm based on 3D Gaussian sputtering to obtain the three-dimensional map scene information and the real-time positioning of the current scene.
[0080] Specifically, the system of this embodiment integrates SLAM (Simultaneous Localization and Mapping) technology for real-time positioning and path planning. The SLAM algorithm combines multi-sensor data of a monocular camera and a lidar, improving the navigation efficiency of the robot in a dynamic environment. In addition, the SLAM algorithm also combines the 3D Gaussian sputtering algorithm. By using the 3D Gaussian sputtering algorithm for three-dimensional modeling of the monocular camera, it can provide efficient and fine small environment perception capabilities. The 3D Gaussian sputtering algorithm can real-time position complex scenes and generate high-fidelity three-dimensional map scene information to provide accurate environmental information for the robot.
[0081] For example, this embodiment can support efficient scene perception and real-time positioning. By integrating innovative scene perception technologies, such as high-precision three-dimensional scene real-time positioning algorithms, the system can quickly understand the spatial layout of the environment, providing a reliable perception basis for the planning and execution of complex tasks, especially showing higher adaptability in dynamic and unknown environments. Specifically, the SLAM algorithm is used to model the large environment of the monocular camera and lidar, providing key data for indoor navigation. The 3D Gaussian sputtering algorithm is used to perform three-dimensional modeling of the monocular camera to provide efficient and fine small environment perception capabilities. 3D Gaussian sputtering can real-time locate complex scenes, thereby efficiently generating high-fidelity three-dimensional map scene information and providing accurate environmental information for the robot.
[0082] Step S300: For the navigation task, through a preset navigation algorithm, generate a target path based on the navigation task, the three-dimensional map scene information, and the real-time positioning, and control the machine motion module based on the target path.
[0083] In this embodiment, smooth movement on the scene is achieved through the machine motion module. The machine motion module can be a bipedal robot or a quadruped robot dog (as Figure 2 shown). This embodiment does not limit the shape / structure / number of legs of the machine motion module. The navigation task refers to the movement task that the robot needs to complete from a starting position to a target position, which may include some specific constraints and requirements, such as avoiding obstacles, following specific path rules, arriving within a specified time, etc. The preset navigation algorithm can be an optimized navigation algorithm, which can be obtained by changing the Dijkstra algorithm to the A* algorithm, and can effectively improve the navigation efficiency of the robot in a dynamic environment. In an actual application scenario, the navigation task, the current three-dimensional map scene information, and the real-time positioning are input into the navigation algorithm, and the navigation algorithm comprehensively plans a target path from the current position to the target position. Finally, the target path is sent to the machine motion module for execution.
[0084] In one implementation, generating a target path based on the navigation task, the three-dimensional map scene information, and the real-time positioning through a preset navigation algorithm includes:
[0085] Determine the target position according to the navigation task;
[0086] Plan a target path from the real-time positioning to the target position through the preset navigation algorithm and the three-dimensional map scene information.
[0087] Specifically, the navigation task provides a goal orientation for the navigation algorithm, indicating the target location that the user expects to reach; the three-dimensional map scene information provides detailed information about the environment, such as the space size and obstacle distribution; the real-time positioning provides the current location of the device. Based on these input information, the navigation algorithm comprehensively considers factors such as distance, obstacles, and traffic rules, and plans a target path from the current location to the target location. This embodiment combines the three-dimensional map scene information, real-time positioning, and navigation algorithm, enhancing the system / agent's path planning and motion control capabilities in complex environments, improving the robot's precise navigation ability in complex terrains, providing solid technical support for flexible environment adaptation and exploration, enabling the robot to move precisely whether in narrow passages or multi-obstacle environments.
[0088] For example, as Figure 5 shown, the user inputs the instruction "Where is the nearest door", and the system obtains video source data (i.e., camera data) through the camera. And obtains point cloud data through the lidar. Then, based on the point cloud data, the aforementioned SLAM algorithm is used to establish and save the map (i.e., three-dimensional map scene data). The target location is determined by combining the user input instruction, video source data, and map through the ReMEmbR project (construction and reasoning of long-range spatio-temporal memory for robot navigation). And the target location is input into the go2-ros2-sdk toolkit (supporting the development of Unitree go2 system based on ROS2 (Robot Operating System2)), and the navigation and motion control of the quadruped robot are realized through the navigation framework Nav2 in ROS2 (Nav2 is Navigation2, supporting various path planning algorithms, map generation, obstacle avoidance, and real-time execution of navigation tasks).
[0089] Step S400: For the grasping task, through a preset robotic arm control algorithm, determine the optimal grasping pose and grasping path of each target object to be grasped according to the grasping task and the three-dimensional map scene information, and control the robotic grasping module based on the optimal grasping pose and the grasping path.
[0090] Specifically, the grasping task refers to the operation target that the robotic grasping module needs to complete, that is, grasping a specific object. The robotic grasping module can be a robotic arm (such as Figure 3As shown, the shape / structure of the robotic grasping module is not restricted in this embodiment. In this embodiment, the object to be grasped is defined as the target object. The grasping task may include some specific requirements, such as the error range of the placement position of the object after being grasped, the force requirement for grasping, the time limit for grasping, and so on. The robotic arm control algorithm is responsible for calculating the optimal grasping pose of each target object according to the given grasping task and the three-dimensional map scene information. The optimal grasping pose of each target object is sent to the robotic grasping module, and the robotic grasping module can then know the actions it should take.
[0091] In an actual application scenario, the robotic arm control algorithm can adopt a multi-modal robotic arm control model or an integrated intelligent grasping optimization model to optimize the target positioning and pose planning of the robotic arm. Through the high-precision analysis of the three-dimensional scene, this model generates the optimal grasping pose for each object, and then the vision-language model determines the specific object. Finally, it generates the optimal grasping path for the robotic arm and ensures the accuracy and stability of the actions. The system in this embodiment empowers the robotic arm through an advanced instruction set, enabling the dexterous robotic arm to exhibit excellent performance in grasping and operation tasks and efficiently complete various tasks, such as object handling, switch operation, etc.
[0092] In one implementation, determining the optimal grasping pose and grasping path of each target object to be grasped according to the grasping task and the three-dimensional map scene information includes:
[0093] Determining several target objects to be grasped according to the grasping task;
[0094] For each of the target objects, generating several potential grasping points for the target object according to the three-dimensional map scene information, and calculating the success rate corresponding to each of the potential grasping points;
[0095] Determining the optimal grasping pose and grasping path of the target object according to the success rate of each of the potential grasping points.
[0096] Specifically, taking one target object as an example, through the robotic arm control algorithm analyzing the grasping task and the current three-dimensional map scene information, multiple potential grasping points of the target object can be sampled, and their success rates can be evaluated, so as to select the optimal grasping action, that is, obtain the optimal grasping pose of the target object. The grasping path is the motion trajectory planning of the robotic arm moving from the current position to a certain potential grasping point. Each potential grasping point corresponds to a grasping path, and multiple potential grasping points can obtain multiple candidate grasping paths, and finally execute the grasping path corresponding to the optimal grasping pose. This strategy enhances the flexibility and adaptability of the robotic arm in a dynamic scene, improves the accuracy and stability of the robotic arm's grasping and operation of objects, enables it to flexibly handle various object types and task requirements in a high-density and complex layout environment, and significantly improves the task completion rate.
[0097] In one implementation, the method further includes:
[0098] For each of the target objects, analyze the attribute information of the object, and adjust the force control strategy of the robotic arm according to the attribute information of the object.
[0099] Specifically, by acquiring the image information of each target object, the object name corresponding to the target object can be identified and the texture information of the target object can be extracted. According to the object name and texture information, the attribute information of the target object can be analyzed. For example, the attribute information includes that the surface of the target object is soft or hard. Finally, the force control strategy of the robotic arm is accurately adjusted through the attribute information, so as to further optimize the grasping action of the robotic arm and enable it to adapt to various target shapes and materials. For example, flexible objects have certain requirements for the grasping force of the robotic arm, and the force control strategy of the robotic arm needs to be adjusted according to the grasping force requirements of the flexible objects; rigid objects have certain requirements for the grasping accuracy and stability of the robotic arm, and the force control strategy of the robotic arm needs to be adjusted according to the grasping accuracy and stability requirements of the rigid objects.
[0100] In one implementation, this embodiment can also strengthen the multi-task processing and dynamic allocation capabilities: through dynamic resource management and task scheduling optimization, improve the flexibility and efficiency of the system in multi-task scenarios, enable it to quickly adjust the strategy according to the task priority and environmental changes, and ensure the coordination and stability during multi-task parallelism.
[0101] Furthermore, the currently used SLAM technology can be replaced with a solution compatible with multi-modal sensors to further improve the environmental perception ability. Secondly, the existing 3D Gaussian sputtering technology can be optimized and improved to enhance the accuracy of the system. In addition, the architecture and parameters of the large language model can be adjusted according to different task requirements. Also, more precise sensors can be integrated on the dexterous robotic arm to improve the accuracy of micro-operation tasks. Additionally, the multi-modal perception ability of the system can be enhanced by adding audio and tactile signal processing modules. Or, the combination of the vector memory database system and the image generation model can be explored for virtual scene prediction in complex environments. Moreover, more low-power designs can be added at the hardware layer to optimize the energy efficiency of the overall system. Further, a machine learning model can be introduced for dynamically adjusting the Prompt generation strategy. In addition, a force control feedback function can be integrated on the robotic arm to enhance its performance in high-precision tasks. Finally, the application of this system to new fields such as medical robots and educational robots can be explored.
[0102] The advantages of the present invention are:
[0103] 1. By designing a framework that integrates large language models, 3D Gaussian sputtering algorithms, robot bases, and dexterous robotic arms, the performance and efficiency of embodied intelligent systems in natural language understanding and task execution have been significantly improved. This framework parses natural language instructions through large language models, constructs semantic memories to support long-term task reasoning, and combines robot and robotic arm hardware to achieve highly intelligent human-computer interaction.
[0104] 2. Through the monocular three-dimensional scene reconstruction technology based on 3D Gaussian sputtering, the Prompt design driven by large models and its decision feedback loop, the collaborative task execution architecture between robotic arms and mobile platforms, and the multi-modal semantic memory system for complex tasks, a multi-modal collaborative framework that combines large models with SLAM and multi-modal robotic arm control can be realized, and long-term task reasoning and decision-making mechanisms based on semantic memory, multi-modal navigation and path planning mechanisms based on scene changes can be developed to improve the grasping efficiency of robotic arms in complex dynamic environments and realize the functions of task dynamic allocation and real-time optimization.
[0105] Based on the above embodiments, the present invention also provides an embodied intelligent system for multi-modal large models, as Figure 6 shown, the system includes:
[0106] A large language model 01, which generates a number of subtasks to be executed according to the user instructions and prompt data through the large language model; each of the subtasks includes a navigation task and / or a grasping task;
[0107] A positioning and mapping module 02, which is used to obtain the point cloud data and camera data of the current scene, and construct the three-dimensional map scene information and obtain real-time positioning according to the point cloud data and the camera data;
[0108] A navigation decision module 03, which is used for the navigation task, and generates a target path according to the navigation task, the three-dimensional map scene information, and the real-time positioning through a preset navigation algorithm;
[0109] A grasping decision module 04, which is used for the grasping task, and determines the best grasping pose and grasping path of each target object to be grasped according to the grasping task and the three-dimensional map scene information through a preset robotic arm control algorithm;
[0110] A machine motion module 05, which is used to execute the target path;
[0111] A machine grasping module 06, which is used to execute the best grasping pose and the grasping path.
[0112] Specifically, as Figure 4As shown, in this embodiment, the top-level module of the system is composed of a large language model (LLM) and is responsible for executing the core task understanding and distribution functions. Among them, the ReMEmbR project is used for the construction and reasoning of long-term visual-spatial memory for robot navigation; the non-navigation task large model segmentation module refers to tasks that do not involve navigation, which are directly handed over to the following execution / control module after being segmented by the top-level large model. The main function of the top-level module is to receive natural language instructions from humans. The core is to enable the robot to behave more naturally and intelligently in human-computer interaction. 3D modeling, robot movement, and robotic arm grasping planning are executed by the middle layer. The middle layer completes 3D scene modeling and robotic arm grasping planning tasks through integrated small models. Among them, MonoGS refers to a simultaneous localization and mapping (SLAM) technology based on a Gaussian distribution model; the world model generates 3D mapping of 3DGS by referring to the work of MonoGS and reconstructs a high-fidelity 3D scene in real time; GraspSplats refers to a technology for robotic arm grasping tasks. The middle layer mainly solves two core problems: one is to control the behavioral memory and movement of the robot; the other is to optimize the grasping action planning of the robotic arm through technologies such as 3D Gaussian sputtering and multi-modal robotic arm control. The modules in the middle layer cooperate closely with the large model in the top layer, complete task calls through the API interface, and provide clear goals and path planning for the actual execution of the underlying hardware. The bottom layer module is mainly responsible for hardware interfaces and task execution. The architecture of the bottom layer module includes the hardware interfaces of the quadruped robot / robot and the dexterous robotic arm, which are responsible for the execution of specific tasks. This layer directly interacts with physical devices and realizes the motion control of the robot and the grasping operation of the robotic arm by calling the underlying drivers. This layer is the direct interaction interface between the entire system and the real environment, and its performance directly determines the final task completion quality of the system.
[0113] Based on the above embodiment, the present invention also provides a terminal, and its principle block diagram can be as Figure 7 shown. The terminal includes a processor, a memory, a network interface, and a display screen connected through a system bus. Among them, the processor of the terminal is used to provide computing and control capabilities. The memory of the terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the terminal is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes the embodied intelligence method of the multi-modal large model. The display screen of the terminal can be a liquid crystal display screen or an electronic ink display screen.
[0114] Those skilled in the art can understand, Figure 7The principle block diagram shown only shows the block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the terminal to which the solution of the present invention is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0115] In one implementation, more than one program is stored in the memory of the terminal, and is configured to be executed by more than one processor. The more than one program includes instructions for performing the embodied intelligence method of the multimodal large model.
[0116] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0117] In summary, the present invention discloses an embodied intelligence method for multi-modal large models. The method obtains a user instruction, and a large language model generates a number of subtasks to be executed according to the user instruction and prompt data; each of the subtasks includes a navigation task and / or a grasping task; obtains point cloud data and camera data of the current scene, constructs the three-dimensional map scene information according to the point cloud data and the camera data, and obtains real-time positioning; for the navigation task, through a preset navigation algorithm, generates a target path according to the navigation task, the three-dimensional map scene information, and the real-time positioning, and controls the machine motion module based on the target path; for the grasping task, through a preset robotic arm control algorithm, determines the optimal grasping pose and grasping path of each target object to be grasped according to the grasping task and the three-dimensional map scene information, and controls the machine grasping module based on the optimal grasping pose and the grasping path. By embedding a large language model, which has particular advantages in dealing with language uncertainty and environmental complexity, the present invention enables the robot to have the ability to understand ambiguous language instructions and helps the system quickly adapt and respond, so that it can execute multiple tasks in complex and dynamic environments. For example, when receiving an ambiguous verbal instruction, the robot can perform semantic parsing through the large model and use tools such as a robotic arm to complete specific operations, such as picking up an object to open a door or other similar tasks.
[0118] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. A multi-modal large model embodied intelligent method, characterized in that: The method comprises: Obtaining user instructions, and generating a plurality of subtasks to be executed according to the user instructions and prompt data through a large language model; each of the subtasks includes a navigation task and / or a crawling task; Acquire point cloud data and camera data of the current scene, construct the three-dimensional map scene information and obtain real-time positioning according to the point cloud data and the camera data; For the navigation task, a target path is generated according to the navigation task, the three-dimensional map scene information and the real-time positioning through a preset navigation algorithm, and a machine motion module is controlled based on the target path; For the grasping task, through the preset robot arm control algorithm, the optimal grasping posture and grasping path of each target object to be grasped are determined according to the grasping task and the three-dimensional map scene information, and the machine grasping module is controlled based on the optimal grasping posture and the grasping path.
2. The embodied intelligent method of a multimodal large model according to claim 1, characterized in that: Generate several subtasks to be executed according to the user instructions and prompt data through the large language model, including: Parsing semantic information corresponding to the user instruction according to the user instruction and prompt data through a large language model; A target task is determined according to the semantic information, and a plurality of subtasks to be executed are generated according to the three-dimensional map scene information and the target task.
3. The embodied intelligent method of a multimodal large model according to claim 2, characterized in that: Determine a target task according to the semantic information, and generate a plurality of subtasks to be executed according to the three-dimensional map scene information and the target task, including: Determine the target task according to the semantic information, and obtain matching memory data from a vector memory database according to the three-dimensional map scene information and the target task, wherein the memory data is used to reflect historical three-dimensional map scene information and / or historical task data; A plurality of subtasks to be executed are generated according to the target task, the three-dimensional map scene information and the memory data.
4. The embodied intelligent method of a multimodal large model according to claim 1, characterized in that: Constructing the three-dimensional map scene information and obtaining real-time positioning according to the point cloud data and the camera data includes: Through the SLAM algorithm based on 3D Gaussian sputtering, three-dimensional mapping and positioning are performed according to the point cloud data and the camera data to obtain the three-dimensional map scene information and the real-time positioning of the current scene.
5. The embodied intelligent method of a multimodal large model according to claim 1, characterized in that: Generate a target path according to the navigation task, the three-dimensional map scene information and the real-time positioning through a preset navigation algorithm, including: determining a target location according to the navigation task; A target path from the real-time positioning to the target location is planned using a preset navigation algorithm and the three-dimensional map scene information.
6. The embodied intelligent method of a multimodal large model according to claim 1, characterized in that: Determining the optimal grasping posture and grasping path for each target object to be grasped according to the grasping task and the three-dimensional map scene information includes: Determining a number of target objects to be grasped according to the grasping task; For each of the target objects, a number of potential grasping points of the target object are generated according to the three-dimensional map scene information, and the success rate corresponding to each of the potential grasping points is calculated; According to the success rate of each potential grasping point, the optimal grasping posture and grasping path of the target object are determined.
7. The embodied intelligent method of a multimodal large model according to claim 1, characterized in that: The method further comprises: For each of the target objects, the attribute information of the object is analyzed, and the force control strategy of the robot arm is adjusted according to the attribute information of the object.
8. A multi-modal large model embodied intelligent system, characterized in that: The system comprises: A large language model, generating a plurality of subtasks to be executed according to the user instructions and prompt data through the large language model; each of the subtasks includes a navigation task and / or a crawling task; A positioning and mapping module, used to obtain point cloud data and camera data of the current scene, construct the three-dimensional map scene information according to the point cloud data and the camera data, and obtain real-time positioning; A navigation decision module, for generating a target path for the navigation task, according to the navigation task, the three-dimensional map scene information and the real-time positioning through a preset navigation algorithm; A grasping decision module, for determining the optimal grasping posture and grasping path for each target object to be grasped according to the grasping task and the three-dimensional map scene information through a preset manipulator control algorithm for the grasping task; A machine motion module, used for executing the target path; A machine grasping module is used to execute the optimal grasping posture and the grasping path.
9. A terminal, characterized in that: The terminal includes a memory and one or more processors; the memory stores one or more programs; the program contains instructions for executing the embodied intelligent method of the multimodal large model as described in any one of claims 1-7; and the processor is used to execute the program.
10. A computer-readable storage medium having a plurality of instructions stored thereon, characterized in that: The instructions are suitable for being loaded and executed by a processor to implement the steps of the embodied intelligent method of a multimodal large model as described in any one of claims 1-7.
Citation Information
Patent Citations
Mobile robot positioning method and system based on three-dimensional point cloud and vision fusion
CN111429574A
Intelligent vehicle based on panoramic vision and laser radar fusion SLAM system
CN114092551A
Grabbing force self-adaptive adjusting manipulator grabbing method and system based on large language model
CN117773914A
Robot intelligent control method and system, electronic equipment and medium
CN118123799A
Concrete compressive strength test method based on full-process scheduling system
CN118347849A
Cited By
Autonomous planning method based on large language model
CN121179443A
Multi-object scene robot reasoning method based on 3DGS modeling and diffusion repairing
CN121212195A
Autonomous navigation method, system and equipment based on multi-modal large model and medium
CN121230741A
Inspection method based on Nano multi-agent framework
CN122454652A