Multi-task operation planning method and system for autonomous robot driven by body cognition large model

The method integrates visual and language modalities with a diffusion strategy to enhance robot operation planning, addressing limitations in traditional methods by enabling precise and adaptive multi-task execution in dynamic environments.

CN120307299AActive Publication Date: 2025-07-15TONGJI UNIV

Patent Information

Application Number
CN202510687989.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-07-15
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

Traditional robot operation planning methods have problems such as task planning limitations, insufficient modal information fusion, lack of accurate real-time action trajectory prediction and poor adaptability to environmental changes in complex environments, especially inadequate performance in multi-task planning.

Method used

The embodied cognitive big model-driven method is adopted to deeply integrate vision and language modality, combine diffusion strategies to generate accurate action trajectories, and update task execution strategies in real time to adapt to environmental changes.

Benefits of technology

It significantly improves the robot's multi-task autonomous planning and execution capabilities in dynamic and complex environments, has better generalization capabilities and real-time adaptability, and can be applied in fields such as automated production, intelligent assembly and service robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120307299A_ABST
    Figure CN120307299A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-task operation planning method and system for an autonomous robot driven by a body cognition large model. The method comprises the steps that S1, coding is carried out based on an RGB image and a depth image which are obtained in real time, and body visual representation is obtained; s2, acquiring a natural language instruction, performing cross-modal fusion on the natural language instruction and the body visual representation to obtain a fusion feature, and performing multi-task decomposition on the basis of the fusion feature; s3, on the basis of a multi-task decomposition scheme, a diffusion strategy is utilized to generate a continuous action track of a robot end effector; and S4, acquiring a second RGB image and a second depth map after the robot executes according to the continuous action track, and taking the second RGB image and the second depth map as closed-loop feedback signals. The system is used for realizing the method. Compared with the prior art, the multi-task autonomous planning and accurate execution capability of the robot in a dynamic complex environment is remarkably improved by deeply fusing vision and language modalities based on the biont cognition large model and predicting an accurate action trajectory in combination with the diffusion strategy action decision module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous robot operation, and in particular, to a method and system for autonomous robot multi-task operation planning driven by an embodied cognition large model. Background Art

[0002] With the rapid development of technologies such as artificial intelligence, big data, and deep learning, robot technology has gradually shifted from traditional pre-programmed operations to intelligent and autonomous operation modes. Especially in complex and dynamic real-world environments, the ability of robots to autonomously execute multiple tasks has become a research hotspot. However, the current traditional robot operation planning methods have the following technical bottlenecks and deficiencies: 1) Task planning limitations: Traditional robot systems often rely on manual definitions or simple rules for task decomposition and planning, lacking flexibility and generalization, and it is difficult to achieve efficient and reliable planning when facing complex environments and multi-task collaborative operations; 2) Insufficient fusion of modal information: Most existing methods usually only rely on single-modal data (such as vision or force sense), and the understanding of environmental information is not comprehensive and in-depth enough, affecting the stability and accuracy of robot task execution; 3) Lack of accurate real-time action trajectory prediction, limiting the success rate of robot operations; 4) Lack of real-time adaptability to environmental changes: In dynamic environments, traditional robot systems are difficult to perceive environmental changes in real time and dynamically adjust task execution strategies accordingly, resulting in poor environmental adaptability and robustness. Although the rise of embodied cognition large models provides potential solutions to the above problems, such as the Chinese patent application 《CN118036750A》, which provides an embodied intelligent task planning method through a multi-modal large model and a behavior tree structure, combining environmental feedback and multi-modal information, although it solves the problem of insufficient planning reliability in the prior art, it still has the following disadvantages: a) It is more suitable for single-task planning and has limitations for multi-task planning; b) The actions rely on sampling of probabilities for generation, resulting in inaccurate prediction results; c) Only a one-time environmental assessment is carried out in the planning stage and it is unable to respond to dynamic changes during the execution process in real time.

[0003] Therefore, it is a technical problem to be solved to provide a robot multi-task planning method based on an embodied cognition large model. Summary of the Invention

[0004] The objective of the present invention is to overcome the deficiencies of the above-mentioned existing technologies and provide a method and system for autonomous robot multi-task operation planning driven by an embodied cognitive large model, which deeply integrates visual and language modalities, and combines a diffusion strategy action decision module to predict accurate action trajectories, effectively solving the problems of insufficient accuracy, poor generalization, and low real-time performance in traditional robot multi-task planning, significantly improving the multi-task autonomous planning and precise execution capabilities of robots in dynamic complex environments, having better generalization ability and real-time adaptability, and being able to be widely applied to multiple fields such as automated production, intelligent assembly, and service robots.

[0005] The objective of the present invention can be achieved through the following technical solutions:

[0006] According to the first aspect of the present invention, there is provided a method for autonomous robot multi-task operation planning driven by an embodied cognitive large model, including:

[0007] S1. Real-time obtain the first RGB image and the first depth map when the robot executes multi-task operations, and encode them based on the RGB map and the depth map to obtain an embodied visual representation;

[0008] S2. Obtain the natural language instructions of the multi-task operations, perform cross-modal fusion on the natural language instructions and the embodied visual representation to obtain a fusion feature, and decompose the multi-task based on the fusion feature to obtain a multi-task decomposition scheme; the multi-task decomposition scheme includes the execution order and priority of single tasks;

[0009] S3. Generate a continuous action trajectory of the robot end effector based on the multi-task decomposition scheme by using a diffusion strategy;

[0010] S4. Obtain the second RGB image and the second depth map after the robot executes according to the continuous action trajectory, compare the first RGB image and the first depth map with the second RGB image and the second depth map respectively. If they are the same, jump to S5; otherwise, generate a prompt message based on the differences and return to S2 to update the multi-task decomposition scheme;

[0011] S5. Determine whether the continuous motion trajectory is successfully executed. If it is successful, determine whether the multi-task is completed. If not, return to S1; if it is not executed successfully, generate a feedback packet and return to S2 to guide the update of the multi-task decomposition strategy and S4 as an attachment condition.

[0012] As a preferred technical solution, the method for obtaining the embodied visual representation includes:

[0013] Preprocess the first RGB image and the first depth map;

[0014] The pre-processed first RGB image and the first depth map are encoded using a ViT encoder to obtain an embodied visual representation; the visual representation includes semantic features and spatial geometric features.

[0015] As a preferred technical solution, the cross-modal fusion method is: encoding the natural language instruction to obtain text features, and using a cross-modal feature attention mechanism to perform feature fusion on the text features and the embodied visual representation.

[0016] As a preferred technical solution, the method for obtaining the text features includes:

[0017] Splitting the natural language instruction into multiple tokens to obtain a discrete index sequence;

[0018] Mapping the discrete index sequence to a low-dimensional dense vector space to form an embedding matrix;

[0019] Performing sine-cosine position encoding on the discrete index sequence, and adding the result of the sine-cosine position encoding to the embedding matrix to obtain the text features.

[0020] As a preferred technical solution, the method for feature fusion includes:

[0021] Mapping the embodied visual representation and the text features to a unified dimension;

[0022] Calculating the first attention of the text features to the embodied visual representation after dimension unification, and its expression is: Calculating the second attention of the embodied visual representation to the text features, and its expression is: Where, and respectively represent the normalized attention weights of the text features to the embodied visual representation and the normalized attention weights of the embodied visual representation to the text features, V j represents the j-th embodied visual representation, T j represents the j-th text feature; N v represents the number of embodied visual representations; N T represents the number of text features;

[0023] Adding the first attention and the second attention to the text features and the embodied visual representation respectively to obtain intermediate features;

[0024] Performing residual processing and normalization on the intermediate features to obtain fusion features.

[0025] As a preferred technical solution, the method for calculating the priority includes:

[0026] Insert a learnable task-level CLS summary token into the fusion feature to form a complete input sequence;

[0027] Process the input sequence with a high-level Transformer to obtain a CLS vector, and calculate the task priority based on the CLS vector. The expression is:

[0028] P = Softmax(h CLS W s )

[0029] where P represents the priority; h CLS represents the CLS vector; W s represents the task classification head for multi-task decomposition.

[0030] As a preferred technical solution, the method for generating the continuous action trajectory includes:

[0031] Obtain the continuous action space of the robot end effector based on the multi-task decomposition scheme, and discretize the continuous action space into multiple action units to construct a discrete action space;

[0032] Sample a continuous initial trajectory from a Gaussian distribution, and then map each element of the trajectory according to the discrete interval to convert it into an initial discrete sequence;

[0033] Use the initial discrete sequence as the initial value of the diffusion model, and perform iterative denoising using the denoising network of the diffusion model based on the embodied visual representation. Adjust the action unit at each iteration to obtain an optimized continuous action trajectory.

[0034] According to the second aspect of the present invention, there is provided an embodied cognitive large model-driven autonomous robot multi-task operation planning system for implementing the above method.

[0035] According to the third aspect of the present invention, there is provided an embodied cognitive large model-driven autonomous robot multi-task operation planning device, including a memory and a processor. A computer program is stored on the memory, and the processor implements the above method when executing the program.

[0036] According to the fourth aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, and the program implements the above method when executed by a processor.

[0037] Compared with the prior art, the present invention performs cross-modal fusion of image features based on real-time acquired images and text instructions of robot tasks, decomposes multiple tasks based on the fused features, explores the execution order and its priority of each subtask, and realizes obtaining the optimal execution scheme of multiple tasks in one inference process. Considering that the continuous action trajectory obtained by inference may have large uncertainties and errors, the present invention also introduces a diffusion strategy to perform iterative convergence on the conditional probability flow, aiming to significantly reduce the impact on trajectory generation in the morning and realize online correction of abnormal trajectories. In addition, the present invention also collects the environmental information after the robot executes the predicted continuous action trajectory, and uses this environmental information as a feedback signal for closed-loop feedback, significantly improving the adaptability and robustness of the robot in the face of environmental changes, and providing an efficient, accurate and robust solution for autonomous robot operation in complex scenarios. Brief Description of the Drawings

[0038] Figure 1 It is a flowchart of the method of the present invention. Detailed Embodiments

[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0040] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the ordinary meanings understood by those with ordinary skills in the technical field to which this application belongs. The words such as "a", "an", "one kind", "the" and the like involved in this application do not indicate a quantity limitation and may represent a singular or plural number. The terms "including", "comprising", "having" and any variations thereof involved in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or units, but may further include unlisted steps or units, or may further include other steps or units inherent to these processes, methods, products or devices. The similar words such as "connected", "linked", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "plurality" involved in this application means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the front and rear associated objects. The terms "first", "second", "third" and the like involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.

[0041] The present invention provides a robot multi-task operation planning method that uses an embodied cognitive large model to drive a vision Transformer and a GPT-4V model for deep fusion of vision and language modalities, and combines a diffusion strategy action decision module to predict accurate action trajectories. It effectively solves the problems of insufficient accuracy, poor generalization, and low real-time performance in traditional robot multi-task planning, significantly improves the multi-task autonomous planning and accurate execution capabilities of robots in dynamic and complex environments, has better generalization ability and real-time adaptability, and can be widely applied to multiple fields such as automated production, intelligent assembly, and service robots.

[0042] The detailed method process is as Figure 1 shown and includes:

[0043] S1. Real-time obtain the first RGB image and the first depth map when the robot performs multi-task operations, and encode them based on the RGB image and the depth map to obtain an embodied visual representation.

[0044] Real-time obtain the RGB image and depth information in the environment through the RGB-D visual sensor carried on the robot to form an environmental perception input. Specifically, use an RGB-D camera such as an Intel RealSense or an Azure Kinect installed on the robot to capture images of the environment in real time and obtain the corresponding depth map data to provide rich environmental information.

[0045] S11. Preprocess the first RGB image and the first depth map.

[0046] S111. Crop the first RGB image and the first depth map to the same size for subsequent fusion.

[0047] S112. Normalize the first RGB image and the first depth map after S111 to eliminate the difference in pixel value distributions.

[0048] S113. Denoise the first RGB image and the first depth map after normalization to reduce sensor noise.

[0049] S12. Encode the preprocessed first RGB image and the first depth map using the ViT encoder to obtain an embodied visual representation; the visual representation includes semantic features and spatial geometric features.

[0050] The preprocessed first RGB image and the first depth map are jointly fed into the Vision Transformer (ViT) encoder to obtain an embodied visual representation that is more friendly to the robot, and a high-dimensional feature vector is output as the embodied visual representation.

[0051] Step S1 enables the robot to accurately identify obstacles and target objects during environment recognition. For example, when dealing with an object recognition task on a table, the model can accurately identify and locate multiple objects such as cups and books, and extract their spatial positions and semantic features.

[0052] S2. Obtain natural language instructions for multi-task operations, perform cross-modal fusion on the natural language instructions and the embodied visual representation to obtain a fused feature, and decompose the multi-task based on the fused feature to obtain a multi-task decomposition plan; the multi-task decomposition plan includes the execution order and priority of individual tasks.

[0053] S21. Obtain text features:

[0054] S211. Split the natural language instructions into multiple tokens to obtain a discrete index sequence where represents the discrete index of the N t th token.

[0055] S212. Map the discrete index sequence to a position dense vector space to form an embedding matrix.

[0056] S213. Perform sine-cosine position encoding on the discrete index sequence, and add the result of the sine-cosine position encoding to the embedding matrix to obtain text features. The expression is: T = W + P, where T represents text features, W represents the embedding matrix, and P represents sine-previous position encoding.

[0057] S22. Cross-modal fusion:

[0058] Encode the natural language instruction to obtain text features, and use the cross-modal feature attention mechanism to fuse the text features with the embodied visual representation.

[0059] Specifically, it includes:

[0060] S221. Map the embodied visual representation and the text features to the same dimension, so that they are in the same semantic space for subsequent interaction.

[0061] S222. Calculate the first attention of the text features to the embodied visual representation after dimension unification. The expression is:

[0062]

[0063] Calculate the second attention of the embodied visual representation to the text features. The expression is:

[0064]

[0065] Among them, and respectively represent the attention weights after normalization of the text features to the embodied visual representation and the attention weights after normalization of the embodied visual representation to the text features. V j represents the j-th embodied visual representation, and T j represents the j-th text feature; N v represents the number of embodied visual representations; N T represents the number of text features.

[0066] Through the above operations, each text token weights all the embodied visual representations, and finds the image region most relevant to its own semantics, such as the specific pixel block on the desktop corresponding to the cup.

[0067] S223. Add the first attention and the second attention to the text features and the embodied visual representation respectively to obtain intermediate features, so that the embodied visual representation pays reverse attention to the text tokens, injects spatial-geometric information into the text sequence, and enables the text to understand the spatial relationship in the text instruction.

[0068] S224. After performing residual processing and normalization on the intermediate features, fused features are obtained. The two-way features after interaction are stably fused through residual connection and layer normalization to suppress noise and stabilize gradients. Finally, fused features are obtained, where each element carries both language and visual contexts.

[0069] S3. Based on the multi-task decomposition scheme, a diffusion strategy is used to generate the continuous action trajectory of the robot end effector.

[0070] S31. Insert a learnable task-level CLS summary token into the fused features to form a complete input sequence.

[0071] S32. Process the input sequence using a high-level Transformer to obtain the CLS vector. Calculate the task priority based on the CLS vector, and its expression is:

[0072] P = Softmax(h CLS W s )

[0073] where h CLS represents the CLS vector; W s represents the task classification head for multi-task decomposition; P represents the priority, and the larger the value, the higher the priority. The final multi-task execution sequence can be obtained by sorting P in descending order. For example: pick up the cup, move to the right side of the table, and put down the cup to grab the book.

[0074] S4. Obtain the second RGB image and the second depth map after the robot executes according to the continuous action trajectory. Compare the first RGB image and the first depth map with the second RGB image and the second depth map respectively. If they are the same, jump to S5; otherwise, generate prompt information based on the differences and return to S2 to update the multi-task decomposition scheme.

[0075] The trajectory consists of countless continuous control quantities, such as position, attitude, joint angle, speed, etc., and cannot be enumerated by rules. Moreover, if directly predicting the continuous trajectory based on the multi-task decomposition scheme obtained in S2, there may be large uncertainties and errors. Therefore, the present invention introduces a diffusion strategy to optimize the continuous trajectory prediction process.

[0076] S41. Based on the multi-task decomposition scheme, obtain the continuous action space of the robot end effector, and discretize the continuous action space into multiple action units to construct a discrete action space.

[0077] S42. Sample a continuous initial trajectory from a Gaussian distribution, and then map each element of the trajectory to the discrete interval to convert it into an initial discrete sequence.

[0078] S43. Use the initial discrete sequence as the initial value of the diffusion model, and perform iterative denoising using the denoising network of the diffusion model based on the embodied visual representation. Adjust the action unit at each iteration to obtain an optimized continuous action trajectory.

[0079] The diffusion strategy regards the entire trajectory as a random variable, and infers the curve that best matches the current state, environment, and task through conditional probability. After the processing in step S4, the specific position, optimal angle, and optimal movement path for the robot to grasp the cup are obtained.

[0080] During the process of the robot executing the task, the embodied cognitive large model analyzes the changes in the environment in real time and dynamically updates the task execution strategy. Specifically, it includes real-time detection of changes in objects or states in the environment, such as the movement of items on the table or the sudden appearance of obstacles. It analyzes and identifies the environmental change prompt information through visual feature analysis, dynamically adjusts the execution order and priority of subtasks according to the prompt information, as well as subsequent action planning, and guides the model to recalculate the priorities and parameters of relevant subtasks (if it is a blocking-level event, insert a "clear obstacle" or "reposition" task and lower the priority of the original task; if it is an optimization-level event, only update the target pose while retaining the order). The updated plan immediately relies on the diffusion strategy to truncate the executed trajectory, uses the current end state as the starting point, and uses the new target constraint g′ as the condition to hot-start denoising for 10 - 20 steps to obtain the remaining trajectory, and seamlessly splices it to the historical trajectory after a quick collision check, achieving millisecond-level online incremental replanning. By continuously cycling through "environmental change detection → prompt generation → context refresh → local re-reasoning → hot-start replanning", the robot can adaptively adjust the subtask order and action trajectory without pausing in the evolving real scenario, ensuring the coherence, safety, and robustness of multi-task execution.

[0081] S5. Determine whether the continuous motion trajectory is successfully executed. If successful, determine whether the multi-task is completed. If not, return to S1; if the execution is not successful, generate a feedback packet and return to S2 to guide the update of the multi-task decomposition strategy and S4 as an additional condition.

[0082] Specifically, immediately after the completion of the actions of each subtask in multitasking, the visual feedback closed-loop is started. The robot captures the current RGB-D image frame_t and stitches it together with the end-effector pose to form the observation data. Subsequently, the object detection and pose estimation network regresses the three-dimensional position and pose of the target object from frame_t, denoted as detected_pose_t. The detection result is differentiated from the expected pose in the multitasking decomposition strategy to calculate the error vector. If the magnitude of the error vector is less than the threshold, the status flag "success" is generated, and the action sequence continues to advance along the current trajectory; if it exceeds the threshold, the flag "fail" is marked, and a feedback packet is constructed. The feedback packet is first sent to the diffusion strategy module in step S4, merged with the original target constraint as an additional condition, and the unfinished trajectory is refreshed during the next denoising warm start to output the corrected remaining trajectory. At the same time, the feedback packet is sent back to S2 to guide the update of the multitasking decomposition strategy for the dynamic task rearrangement logic to determine whether to insert subtasks such as "repositioning" or "obstacle clearing". The entire link sequentially experiences "observation acquisition → object detection → error evaluation → feedback generation → diffusion replanning → dynamic rearrangement", and the end-to-end delay is controlled within approximately 20 milliseconds, which can continuously correct the placement accuracy and improve the robustness and accuracy of the overall motion planning while the robot is continuously moving.

[0083] The present invention also provides an embodied cognition large model-driven autonomous robot multitasking operation planning system for implementing the above method.

[0084] Moreover, the embodied cognition large model-driven autonomous robot multitasking operation planning device provided by the present invention includes a central processing unit (CPU), which can execute various appropriate actions and processes according to the computer program instructions stored in the read-only memory (ROM) or the computer program instructions loaded from the storage unit into the random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other through a bus. The input / output (I / O) interface is also connected to the bus.

[0085] Multiple components in the device are connected to the I / O interface, including: an input unit, such as a keyboard, a mouse, etc.; an output unit, such as various types of displays, speakers, etc.; a storage unit, such as a disk, an optical disc, etc.; and a communication unit, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0086] The processing unit executes the various methods and processes described above, such as methods S1 to S5. For example, in some embodiments, methods S1 to S5 may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of methods S1 to S5 described above may be executed. Alternatively, in other embodiments, the CPU may be configured to execute methods S1 to S5 by any other suitable means (such as, by means of firmware).

[0087] The functions described above herein may be performed at least in part by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0088] The program code for implementing the methods of the present invention may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on the remote machine or server.

[0089] In the context of the present invention, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0090] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. An autonomous robot multi-task operation planning method driven by an embodied cognitive large model, characterized in that, Including: S1. Obtain the first RGB image and the first depth map in real time when the robot performs multi-task operations, and encode them based on the RGB image and the depth map to obtain an embodied visual representation; S2. Obtain the natural language instructions of the multi-task operations, perform cross-modal fusion on the natural language instructions and the embodied visual representation to obtain a fusion feature, and decompose the multi-task based on the fusion feature to obtain a multi-task decomposition scheme; the multi-task decomposition scheme includes the execution order and priority of a single task; S3. Generate a continuous action trajectory of the robot end effector based on the multi-task decomposition scheme by using a diffusion strategy; S4. Obtain the second RGB image and the second depth map after the robot executes according to the continuous action trajectory, and compare the first RGB image and the first depth map with the second RGB image and the second depth map respectively. If they are the same, jump to S5; Otherwise, generate a prompt message based on the differences and return to S2 to update the multi-task decomposition scheme; S5. Determine whether the continuous motion trajectory is successfully executed. If it is successful, determine whether the multi-task is completed. If not, return to S1; if it is not executed successfully, generate a feedback packet and return to S2 to guide the update of the multi-task decomposition strategy and S4 as an attachment condition.

2. The method for autonomous robot multi-task operation planning driven by an embodied cognition large model according to claim 1, wherein The method for obtaining the embodied visual representation includes: Preprocess the first RGB image and the first depth map; Encode the preprocessed first RGB image and first depth map using a ViT encoder to obtain an embodied visual representation; the visual representation includes semantic features and spatial geometric features.

3. A method for autonomous robot multi-task operation planning driven by an embodied cognition large model according to claim 1, characterized in that, The method for cross-modal fusion is: Encode the natural language instructions to obtain text features, and use a cross-modal feature attention mechanism to fuse the text features with the embodied visual representation.

4. A method for autonomous robot multi-task operation planning driven by an embodied cognitive large model according to claim 3, characterized in that The method for obtaining the text features includes: Split the natural language instructions into multiple tokens to obtain a discrete index sequence; Map the discrete index sequence to a low-dimensional dense vector space to form an embedding matrix; Perform sine-cosine position encoding on the discrete index sequence, and add the result of the sine-cosine position encoding to the embedding matrix to obtain the text features.

5. A method for autonomous robot multi-task operation planning driven by an embodied cognitive large model according to claim 3, characterized in that, The method for feature fusion includes: Map the embodied visual representation and the text features to a unified dimension; Calculate the first attention of the text features to the embodied visual representation after the calculation dimension is unified. Its expression is: Calculate the second attention of the embodied visual representation to the text features. Its expression is: Among them, and respectively represent the attention weight after normalizing the text features to the embodied visual representation and the attention weight after normalizing the embodied visual representation to the text features. V j represents the j-th embodied visual representation, and T j represents the j-th text feature; N v represents the number of embodied visual representations; N T represents the number of text features; Add the first attention and the second attention to the text features and the embodied visual representation respectively to obtain intermediate features; Perform residual processing and normalization on the intermediate features to obtain a fusion feature.

6. The method for autonomous robot multi-task operation planning driven by an embodied cognitive large model according to claim 1, wherein The method for calculating the priority includes: Insert a learnable task-level CLS summary token into the fusion feature to form a complete input sequence; Process the input sequence using a high-level Transformer to obtain a CLS vector, and calculate the task priority based on the CLS vector. Its expression is: P = Softmax(h CLS W s ) where P represents the priority; h CLS represents the CLS vector; W s represents the task classification head for multi-task decomposition.

7. A method for autonomous robot multi-task operation planning driven by an embodied cognitive large model according to claim 1, characterized in that, The method for generating the continuous action trajectory includes: Obtain the continuous action space of the robot end effector based on the multi-task decomposition scheme described above, and discretize the continuous action space into multiple action units to construct a discrete action space; Sample a continuous initial trajectory from a Gaussian distribution, and then map each element of the trajectory at discrete intervals to convert it into an initial discrete sequence; Use the initial discrete sequence as the initial value of the diffusion model, and perform iterative denoising using the denoising network of the diffusion model based on the embodied visual representation. Adjust the action units at each iteration to obtain an optimized continuous action trajectory.

8. An autonomous robot multi-task operation planning system driven by an embodied cognitive large model, characterized in that, The system described above is used to implement the method described in any one of claims 1 to 7.

9. An apparatus for autonomous robot multi-task operation planning driven by an embodied cognitive large model, comprising a memory and a processor, wherein a computer program is stored on the memory, and is characterized in that, When the processor executes the program, it implements the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • End-to-end learning method, system and equipment based on multi-modal large model

    CN118211643A

  • Online reinforcement learning data increasing and expanding method based on diffusion model

    CN119476372A

  • Natural language control method for humanoid robot

    CN119610090A

  • Action planning for robot control

    US20250162150A1

  • Thermodynamic artificial intelligence for generative diffusion models and bayesian deep learning

    WO2024118915A1

Cited By

  • Reconfigurable robot autonomous splicing method and system based on visual language large model

    CN120563624A

  • Method for predicting directed availability of object and action of robot and related equipment

    CN121132651A

  • Action planning method and device based on visual language action model and storage medium

    CN122416055A