Multi-modal data processing method supporting multi-task switching and related equipment

By generating fusion features and multi-task conditional features, combining action expert model and denoising flow matching technology, the complexity problem during robot task switching is solved, the action accuracy and robustness are improved, and task switching is adapted to complex scenarios.

CN120245085AActive Publication Date: 2025-07-04BEIJING HUMANOID ROBOTICS INNOVATION CENTER CO LTD

Patent Information

Application Number
CN202510711329.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-04
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

The existing robot multi-task switching methods fail to effectively handle the complexity of task switching in real scenarios, resulting in poor operational accuracy, especially when task switching, the execution status and position changes of the previous and subsequent tasks are not considered.

Method used

By obtaining the robot's current environment image, historical task instructions, current task instructions and robot status, fusion features and multi-task condition features are generated, and action expert models are used to make action decisions to generate optimal action sequences, combining multimodal large model and denoising flow matching technology to process action noise.

Benefits of technology

It improves the accuracy of the robot's movement during task switching, can handle complex task switching scenarios, and enhances the robustness and task coherence of the robot in arbitrary interruption and irregular states.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120245085A_ABST
    Figure CN120245085A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal data processing method supporting multi-task switching and related equipment, the method is applied to the technical field of robots, and the method comprises the following steps: acquiring an environment image, a historical task instruction, a current task instruction, a robot state and robot action noise of an environment where a robot is located at present; generating fusion features and multi-task condition features by using the environment image, the historical task instruction and the current task instruction; robot state features are extracted from the robot states; and the robot action noise, the fusion features, the multi-task condition features and the robot state features are input into a preset action expert model for action decision making, so that an optimal action sequence of the robot is obtained. According to the scheme, the optimal action sequence of the robot is obtained according to the historical task instruction, the current task instruction and the environment image, the problem of complexity during task switching of the robot is effectively solved, and the accuracy of robot actions output during task switching is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robots, and in particular, to a multi-modal data processing method and related devices that support multi-task switching. Background Art

[0002] In recent years, end-to-end imitation learning has developed rapidly. With the robustness of large models, methods such as DP (Diffusion Policy) and OpenVLA have been proposed, and most of these existing methods support multi-task training.

[0003] However, in the real scenario of robot applications, task switching is random and at any time. However, in the training of existing methods such as DP (Diffusion Policy) and OpenVLA, each task data obtained is from start to end and is an independent and complete operation. When the robot executes tasks, these existing methods rely on the successful execution of the previous task and fully return to the initial position when collecting data for the next task, without considering the complexity of task switching, which results in poor accuracy of the robot actions output when task switching occurs. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a multi-modal data processing method and related devices that support multi-task switching to solve the complexity problem when the robot switches tasks.

[0005] To achieve the above object, embodiments of the present invention provide the following technical solutions:

[0006] A first aspect of an embodiment of the present invention discloses a multi-modal data processing method that supports multi-task switching, and the method includes:

[0007] Obtain the environmental image, historical task instruction, current task instruction, robot state, and robot action noise of the environment where the robot is currently located;

[0008] Generate a fusion feature and a multi-task conditional feature by using the environmental image, the historical task instruction, and the current task instruction; wherein, the fusion feature is generated based on the following features: visual features extracted from the environmental image and the current environmental state features, and task instruction features extracted from the historical task instruction and the current task instruction; the multi-task conditional feature is generated based on the current environmental state features and the task instruction features;

[0009] Extract robot state features from the robot state;

[0010] Input the robot motion noise, the fusion features, the multi-task conditional features, and the robot state features into a preset motion expert model for motion decision-making to obtain the optimal motion sequence of the robot.

[0011] Preferably, using the environmental image, the historical task instruction, and the current task instruction to generate fusion features and multi-task conditional features, including:

[0012] Use an image encoding model to extract visual features from the environmental image;

[0013] Input the environmental image and a preset prompt word into a vision-language large model to obtain the current environmental state;

[0014] Input the historical task instruction and the current task instruction into a large language model to obtain corresponding task instruction features;

[0015] Input the current environmental state into the large language model to obtain corresponding current environmental state features;

[0016] Use a specified encoder to fuse the visual features, the task instruction features, and the current environmental state features to obtain fusion features;

[0017] Input the task instruction features and the current environmental state features into a multi-task conditional fusion model for information integration to obtain multi-task conditional features.

[0018] Preferably, extracting robot state features from the robot state, including:

[0019] Use a robot state encoding model to extract robot state features from the robot state.

[0020] Preferably, after obtaining the multi-task conditional features, it further includes:

[0021] Distill the multi-task conditional features to obtain the compressed multi-task conditional features.

[0022] Preferably, the process of obtaining the environmental image of the environment where the robot is currently located includes:

[0023] Call the first camera and the second camera to obtain the environmental image of the environment where the robot is currently located. The first camera has a first perspective and is mounted on the robot, and the second camera has a global perspective and is set in the environment where the robot is located.

[0024] Preferably, the process of obtaining the robot state includes:

[0025] Call the sensors of the robot body to obtain the robot state, where the robot state at least includes: the joint angles of the robot and the end effector state.

[0026] Preferably, after obtaining the optimal action sequence of the robot, it further includes:

[0027] Perform denoising processing on the optimal action sequence.

[0028] A second aspect of the embodiments of the present invention discloses a multi-modal data processing system supporting multi-task switching, and the system includes:

[0029] An acquisition unit, configured to acquire an environmental image of the current environment where the robot is located, a historical task instruction, a current task instruction, a robot state, and robot action noise;

[0030] A generation unit, configured to generate a fusion feature and a multi-task conditional feature by using the environmental image, the historical task instruction, and the current task instruction; wherein, the fusion feature is generated based on the following features: visual features and current environmental state features extracted from the environmental image, and task instruction features extracted from the historical task instruction and the current task instruction; the multi-task conditional feature is generated based on the current environmental state feature and the task instruction feature;

[0031] An extraction unit, configured to extract robot state features from the robot state;

[0032] A decision-making unit, configured to input the robot action noise, the fusion feature, the multi-task conditional feature, and the robot state features into a preset action expert model for action decision-making to obtain the optimal action sequence of the robot.

[0033] A third aspect of the embodiments of the present invention discloses an electronic device, including: a processor and a memory, and the processor and the memory are connected through a communication bus; wherein, the processor is configured to call and execute a program stored in the memory; the memory is configured to store a program, and the program is used to implement the multi-modal data processing method supporting multi-task switching disclosed in the first aspect of the embodiments of the present invention.

[0034] A fourth aspect of the embodiments of the present invention discloses a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it implements the multi-modal data processing method supporting multi-task switching disclosed in the first aspect of the embodiments of the present invention.

[0035] Based on the multi-modal data processing method and related device for supporting multi-task switching provided in the embodiments of the present invention, the optimal action sequence of the robot is obtained according to the historical task instruction, the current task instruction and the environmental image, effectively solving the complexity problem when the robot performs task switching and improving the accuracy of the robot actions output during task switching. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0037] Figure 1 It is a flowchart of a multi-modal data processing method for supporting multi-task switching provided in the embodiments of the present invention;

[0038] Figure 2 It is a flowchart of generating fused features and multi-task conditional features provided in the embodiments of the present invention;

[0039] Figure 3 It is a general process example diagram provided in the embodiments of the present invention;

[0040] Figure 4 It is a detailed example diagram of each part provided in the embodiments of the present invention;

[0041] Figure 5 It is an example diagram of a multi-task conditional fusion model provided in the embodiments of the present invention;

[0042] Figure 6 It is a structural block diagram of a multi-modal data processing system for supporting multi-task switching provided in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0044] In this application, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0045] In recent years, end-to-end imitation learning has developed rapidly. Thanks to the robustness of large models, methods such as DP (DiffusionPolicy), RDT, OpenVLA, and Pi0 have been proposed. Most of these existing methods support multi-task training, that is, each task has a corresponding language instruction during training, and the task is switched by adjusting the corresponding language instruction when the robot executes the task.

[0046] However, when the above-mentioned existing methods are training, each task data obtained is from start to end and operates independently and completely. Therefore, when the robot executes, these existing methods rely on the successful execution of the previous task and fully return to the initial position of the next task when collecting data, without considering the complexity of task switching.

[0047] In the real scenario of robot applications, task switching is random and arbitrary at any time. There may be situations where "the user switches to the next task during the execution of the previous task", or there may be situations where "the robot switches the current task when it is not at the position for collecting data of the current task".

[0048] In addition, the actions that should be taken between the state of these tasks during operation and the start of the execution of the next task also need to be concerned. For example, when performing the task of "picking up an apple", it can be divided into the following three states: 1) The robot has not yet touched the apple, 2) The robot picks up the apple, 3) The robot has finished picking up the apple. When switching to the task of "opening the drawer" (i.e., switching to the next task) during the execution of these three states, the corresponding operations should be different, which is a problem not considered by existing methods.

[0049] Therefore, how to process the previous and next tasks and the execution state of the previous task when the user switches tasks, model these combination situations in a certain form, and enable the multi-modal embodied operation model to understand and give the correct action output is an issue that cannot be ignored in the development of robot technology.

[0050] In view of the above problems, this solution proposes a multi-modal data processing method and related devices that support multi-task switching. The optimal action sequence of the robot is determined based on historical task instructions, current task instructions, and environmental images, effectively solving the complexity problem during robot task switching and improving the accuracy of the robot actions output during task switching.

[0051] It should be noted that the multi-modal data processing method and related devices proposed in this solution can be implemented in a multi-modal embodied intelligent operation framework that supports multi-task switching. The following will detail this solution through various embodiments.

[0052] To better understand the content of each embodiment of this solution, a brief explanation of the general concept of this solution is given here first:

[0053] This solution uses the encoder (i.e., the encoder specified later) in a pre-trained multi-modal large model (such as any model like a vision-language large model) as the backbone network, so as to have the ability to fuse multiple modalities, including visual features, task instruction features, and current environmental state features.

[0054] Among them, the multi-view images (hand-eye camera images + third-view camera images) are extracted by the image encoder to obtain corresponding visual features. At the same time, a vision-language large model (VLM) is used to parse the multi-view images and return the current environmental state. For example, it returns current environmental states such as "whether holding an apple", "approaching the apple", "already completed placing the apple", etc. The current environmental state will be input into the large language model (LLM) to extract the current environmental state features, and the current environmental state features are input into the pre-trained "specified encoder".

[0055] The previous task instruction (historical task instruction in the following text) and the current task instruction respectively pass through the large language model (LLM) to extract task instruction features and input them into the "specified encoder". This "specified encoder" extracts corresponding fusion features from all inputs.

[0056] The multi-task conditional fusion model further fuses the previous task instruction features (historical task instruction features in the following text), current task instruction features, and current environmental state features extracted by the vision-language large model (VLM) and the large language model (LLM), and distills the obtained multi-task conditional features to obtain compressed multi-task conditional features.

[0057] The robot state (such as the angle values of each joint and the state of the gripper) is extracted by the robot state encoding model to obtain robot state features.

[0058] The motion expert model serves as the decoder of the vision-language large model (VLM). Through the cross-attention mechanism, it performs state prediction and future motion decision-making for the robot by combining the fused features extracted by the "specified encoder", the multi-task conditional features extracted by the multi-task conditional fusion model, the robot state features extracted by the robot state encoding model, and the robot motion noise.

[0059] Among them, the robot motion noise is Gaussian noise. After passing through the motion expert model and the motion decoder to decide the future motion of the robot, this training method is denoising flow-matching. Ultimately, the robot can make corresponding and effective motion decisions according to the complex conditions of the previous task instruction, the current task instruction, and the current environmental image.

[0060] See Figure 1 , which shows a flowchart of a multi-modal data processing method for supporting multi-task switching provided by an embodiment of the present invention. The multi-modal data processing method includes the following steps:

[0061] Step S101: Obtain the environmental image of the current environment where the robot is located, the historical task instruction, the current task instruction, the robot state, and the robot motion noise.

[0062] In the process of specifically implementing step S101, obtain the environmental image of the current environment where the robot is located, obtain the robot state and the robot motion noise, and receive the task instruction input, which includes the historical task instruction (the previous task instruction) and the current task instruction.

[0063] It should be noted that the received task instruction input (historical task instruction and current task instruction) can provide the robot with multi-level composite condition execution goals; the robot motion noise can be Gaussian noise.

[0064] In some specific embodiments, the specific implementation method for obtaining the environmental image of the current environment where the robot is located is: call the first camera and the second camera to obtain the environmental image of the current environment where the robot is located. The first camera has a first perspective and is mounted on the robot, and the second camera has a global perspective and is set in the environment where the robot is located.

[0065] That is, the first camera is mounted on the robot, and the second camera is in the environment where the robot is located.

[0066] It can be understood that the robot obtains the environmental image (or environmental information) through multi-source sensors. The environmental image includes the hand-eye camera image (first perspective) and the third perspective camera image (global perspective or third perspective), thus comprehensively covering the working scene.

[0067] Among them, the multi-source sensors at least include a first camera and a second camera. The hand-eye camera image is captured by the first camera with a first viewing angle, and the third-view camera image is captured by the second camera with a global viewing angle. The first camera and the second camera can use high-resolution depth cameras.

[0068] For example: The first camera and the second camera use a specified model of depth camera, and the resolution of this depth camera is "1280*720@30fps".

[0069] In some other specific embodiments, the specific implementation method for obtaining the robot state is: calling the robot body sensor to obtain the robot state, and the robot state at least includes: the joint angles of the robot and the end effector state.

[0070] That is to say, the real-time robot state is collected through the robot body sensor.

[0071] Step S102: Generate a fusion feature and a multi-task conditional feature by using the environmental image, the historical task instruction, and the current task instruction.

[0072] In the process of specifically implementing step S102, an image encoding model, a vision-language large model, a large language model, and a multi-task conditional fusion model are used to process the environmental image, the historical task instruction, and the current task instruction, so as to generate a fusion feature and a multi-task conditional feature.

[0073] It should be noted that the fusion feature is generated based on the following features: the visual feature and the current environmental state feature extracted from the environmental image, and the task instruction feature extracted from the historical task instruction and the current task instruction; the multi-task conditional feature is generated based on the current environmental state feature and the task instruction feature.

[0074] Step S103: Extract the robot state feature from the robot state.

[0075] In the process of specifically implementing step S103, a robot state encoding model is used to extract the robot state feature from the robot state.

[0076] For example: The robot state encoding model can be a multi-layer perceptron (MLPs). The robot state (i.e., the original sensor data) is encoded through this multi-layer perceptron to convert the robot state into a high-dimensional feature representation, so as to extract the robot state feature.

[0077] Step S104: Input the robot action noise, the fusion feature, the multi-task conditional feature, and the robot state feature into a preset action expert model for action decision-making to obtain the optimal action sequence of the robot.

[0078] In the process of specifically implementing step S104, the robot action noise, fused features, multi-task conditional features, and robot state features are input into a preset action expert model for action decision-making to obtain the optimal action sequence of the robot (the action decision result, that is, the future actions of the robot).

[0079] Specifically, the robot action noise, fused features, multi-task conditional features, and robot state features are input into the action expert model for action decision-making to obtain the output features output by the action expert model; the action decoder (robot action decoder) is used to decode the output features, thereby obtaining the optimal action sequence of the robot.

[0080] In some specific embodiments, after obtaining the optimal action sequence of the robot, denoising processing is performed on the optimal action sequence, thereby gradually eliminating decision-making noise and improving the smoothness and accuracy of the action trajectory.

[0081] It should be noted that the action expert model adopts a decoder-only transformer architecture (for example, with a parameter quantity of 500M), and performs weighted processing on various types of information through the cross-attention mechanism, and finally outputs the action sequence (that is, the optimal action sequence) for the next 30 steps (for example only, can be adjusted according to actual needs) and the immediate state evaluation.

[0082] Furthermore, it should be noted that when performing denoising processing on the obtained optimal action sequence, Diffusion Flow-Matching (DFM) can be used for denoising processing, and the decision-making noise is gradually eliminated through multiple (for example, 10 times) iterative optimization to improve the smoothness and accuracy of the action trajectory.

[0083] The core goal of DFM is to gradually transform the noise distribution (such as random Gaussian noise) into the target distribution (such as the optimal action sequence of the robot) by learning a probability path. Among them, Flow-Matching directly learns the conversion path from noise to target data by defining a vector field, which is more efficient than traditional diffusion models.

[0084] In the embodiments of the present invention, the optimal action sequence of the robot is obtained based on historical task instructions, current task instructions, and environmental images, effectively solving the complexity problem when the robot switches tasks and improving the accuracy of the robot actions output during task switching.

[0085] Regarding the above embodiments of the present invention Figure 1 For the fused features and multi-task conditional features involved in step S102, refer to Figure 2, which shows the steps of generating fusion features and multi-task conditional features provided by the embodiments of the present invention, including the following steps:

[0086] Step S201: Extract visual features from the environmental image using an image encoding model.

[0087] In the process of specifically implementing step S201, the environmental image (multi-view image, that is, hand-eye camera image + third-view camera image) is input into the image encoding model, and the corresponding visual features (also called image features) are extracted through this image encoding model.

[0088] It should be noted that the image encoding model can be an image encoder based on transformer (for example, the number of parameters can reach 300M). This image encoding model is pre-trained on a large amount of image data and can effectively capture the key visual features in the scene.

[0089] Step S202: Input the environmental image and a preset prompt into the vision-language large model to obtain the current environmental state.

[0090] In the process of specifically implementing step S202, the environmental image and the prompt (used to obtain the current environmental state) are input into the vision-language large model (VLM) for processing, so as to extract the semantic information of the current environmental state and form a multi-modal understanding of the environment.

[0091] It should be noted that this vision-language large model is pre-trained with large-scale multi-modal (text + image). This vision-language large model is based on the transformer architecture and understands the relevance between the visual scene and the text prompt through the self-attention mechanism, and outputs a concise description of the environmental state within 20 words (the number of words is only an example) (that is, the current environmental state).

[0092] Step S203: Input the historical task instruction and the current task instruction into the large language model to obtain the corresponding task instruction features.

[0093] In the process of specifically implementing step S203, the historical task instruction and the current task instruction are input into the large language model, and the corresponding task instruction features are extracted through this large language model (LLM). Among them, the task instruction features include: historical task instruction features corresponding to the historical task instruction (providing historical task context), and current task instruction features corresponding to the current task instruction (clarifying the immediate task goal).

[0094] It should be noted that the parameter scale of this large language model can be 7B parameter scale (for example only). This large language model is responsible for parsing the natural language instructions input by the user (historical task instructions and current task instructions), and converting the vague user instructions into a structured task representation through semantic understanding ability, providing a clear action goal for subsequent decision-making.

[0095] Step S204: Input the current environmental state into the large language model to obtain the corresponding current environmental state features.

[0096] In the process of specifically implementing step S204, input the current environmental state into the large language model to extract the corresponding current environmental state features.

[0097] Step S205: Use the specified encoder to fuse visual features, task instruction features, and current environmental state features to obtain fused features.

[0098] In the process of specifically implementing step S205, use the specified encoder to fuse visual features, task instruction features, and current environmental state features to obtain fused features.

[0099] It should be noted that the specified encoder can be a transformer encoder; the specified encoder deeply fuses visual features, task instruction features, and current environmental state features to generate fused features containing multi-modal information.

[0100] Step S206: Input the task instruction features and current environmental state features into the multi-task conditional fusion model for information integration to obtain multi-task conditional features.

[0101] In the process of specifically implementing step S206, input the task instruction features and current environmental state features into the multi-task conditional fusion model, and extract multi-task conditional features through this multi-task conditional fusion model.

[0102] It should be noted that the multi-task conditional fusion model adopts an architecture that combines the Transformer cross-attention mechanism and a multi-layer perceptron (MLP) to achieve intelligent integration of robot task-related information.

[0103] In some specific embodiments, after extracting the multi-task conditional features, distill the multi-task conditional features to obtain compressed multi-task conditional features; in this case, the "multi-task conditional features" input into the action expert model are the "compressed multi-task conditional features".

[0104] It should be noted that the specified encoder (transformer encoder) is used for multimodal fusion (i.e., the fusion of language and vision), and the multi-task conditional fusion model is used to fuse and compress text information (historical task instructions, current task instructions, current environmental status). There are three reasons for such settings:

[0105] First: The fused features are doped with visual features, which may affect the multi-task conditional features;

[0106] Second: The fused features are passed into the action expert model, and cross-attention mechanism calculations are performed with multi-task conditional features, robot state features, and robot action noise. It is not a relationship of parallel input;

[0107] Third: The action expert model is a Diffusion model, which has a large demand for computing resources (it is best to provide a small number of tokens). The multi-task conditional fusion model is equivalent to compressing information. Directly inputting the multi-task conditional features into the action expert model can reduce the amount of calculation.

[0108] The above is the related description of how to generate fused features and multi-task conditional features.

[0109] It should be noted that in the above embodiments of the present invention Figure 2 the execution order of each step is only for illustrative purposes. In actual execution, one execution order is: sequentially execute step S201-step S206; another execution order is: parallelly execute step S201, step S202, and step S203, then execute step S204, and then parallelly execute step S205 and step S206; another execution order is: parallelly execute step S201, step S202, and step S203, then execute step S204, and then sequentially execute step S205 and step S206. Here, no Figure 2 specific limitation is imposed on the execution order of each step in it, and it can be adjusted according to the actual situation.

[0110] To better understand the content of the above various embodiments, the following combines Figures 3 to 5 to give an illustrative example from three aspects: overall process, details of each part, and multi-task conditional fusion model.

[0111] I. Illustrative example for the overall process:

[0112] For example Figure 3As can be seen from the overall process example diagram shown, the robot obtains environmental images through sensors. The environmental images include hand-eye camera images (first perspective) and third-perspective camera images (global perspective or third perspective). The sensors at least include a first camera and a second camera, and the first camera and the second camera can use depth cameras with a resolution of "1280*720@30fps".

[0113] Receive task instruction input, which includes historical task instructions (previous task instructions) and current task instructions.

[0114] The Visual Language Model (VLM) jointly encodes multi-perspective images (hand-eye camera images + third-perspective camera images) and prompt words (used to obtain the current environmental state), extracts semantic information of the current environmental state, and forms a multi-modal understanding of the environment.

[0115] The obtained multi-modal data (such as visual features, task instruction features, and current environmental state features) are uniformly processed by a feature fusion Transformer encoder (designated encoder). The feature fusion Transformer encoder uses a self-attention mechanism to dynamically fuse visual features, temporal context (historical task instruction features + current task instruction features), and current environmental state features of the current task executed by the robot, thereby extracting fusion features. The fusion features are input into an action expert model, and combined with the robot state features to decide the future actions of the robot.

[0116] The above process effectively solves the problem of difficult compound condition training during multiple task switches, and enhances the adaptability to complex scene human-computer interaction. Based on the various features input into the action expert model, the action expert model generates the optimal action sequence (future actions of the robot) of the robot through an imitation learning strategy. The robot performs specific operations (such as grasping, moving, etc.) according to the instruction requirements, thereby forming a "perception - decision - execution" closed loop. In addition, it also supports real-time feedback adjustment to ensure the accuracy and robustness of task execution.

[0117] The above is the specific content regarding "I. Example Explanation of the Overall Process".

[0118] II. Example Explanation of Each Part's Details:

[0119] Such as Figure 4 As can be seen from the example diagram of each part's details shown, a multi-perspective visual acquisition device is configured in the perception layer to obtain environmental images (that is, hand-eye camera images + third-perspective camera images, which is also equivalent to visual data) from different angles. These obtained environmental images and the pre-designed prompt words are jointly input into a Visual Language Model (VLM) that has undergone large-scale multi-modal pre-training.

[0120] The visual language model (VLM) is based on the transformer architecture. It understands the correlation between visual scenes and text prompts through the self-attention mechanism, and outputs a concise description of the environment state within 20 words (i.e. the current environment state).

[0121] At the same time, a special transformer-based image encoder (i.e., image coding model) is used to perform deep feature extraction on the original visual data (i.e., environmental images) to obtain corresponding visual features. The image encoder is pre-trained based on a large amount of image data and can effectively capture the key visual features in the scene.

[0122] At the instruction understanding layer, a large language model (LLM) with 7B parameters is integrated. This large language model is responsible for parsing the natural language instructions input by the user (historical task instructions and current task instructions), and converts vague user instructions into structured task representations through semantic understanding capabilities, providing clear action goals for subsequent decisions.

[0123] The robot state (joint angle, end effector state) is obtained through the robot body sensor, and encoded through a multi-layer fully connected network (MLPs, i.e., the robot state encoding model), converting the original sensor data into a high-dimensional feature representation, thereby extracting the robot state features. Subsequently, these features are processed by the state decoder, which consists of multi-level MLPs and GeLU activation functions, and can restore the encoded features to a readable state description.

[0124] At the decision-making and planning layer, this scheme adopts a multi-task conditional fusion mechanism to obtain multi-task conditional features; through the transformer encoder, the visual features, task instruction features and current environment state features are deeply fused to generate fused features containing multimodal information (i.e., task conditional representation).

[0125] The fusion features, multi-task condition features, robot state features and robot action noise are input into the action expert model. The action expert model adopts a decoder-only transformer architecture (with 500M parameters) and performs weighted processing on various types of information through a cross-attention mechanism. Finally, it outputs the action sequence of the next 30 steps (i.e., the optimal action sequence) and the immediate state evaluation. The action decision results are denoised based on DFM, and the decision noise is gradually eliminated through 10 iterations of optimization to improve the smoothness and accuracy of the action trajectory.

[0126] Specifically, the robot motion noise, fused features, multi-task conditional features, and robot state features are input into the motion expert model for motion decision-making to obtain the output features output by the motion expert model; the motion decoder (robot motion decoder) is used to decode the output features, thereby obtaining the optimal motion sequence of the robot.

[0127] For the prediction of the current environmental state, during the training phase, the accurate state annotations provided by the vision-language large model (VLM) are used to ensure the basic performance of each module.

[0128] In the inference phase, a flexible operation path selection is provided: when dealing with a completely new environment, the complete state evaluation process of the language large model (VLM) is enabled; when pursuing real-time performance and the environment is known, the state prediction function of the motion expert model can be directly called to significantly improve the response speed.

[0129] The above is the specific content regarding "II. Illustrative Examples for Each Particular Detail".

[0130] III. Illustrative Examples for the Multi-Task Conditional Fusion Model:

[0131] As Figure 5 As can be seen from the example diagram of the multi-task conditional fusion model shown, the multi-task conditional fusion model adopts an architecture that combines the Transformer cross-attention mechanism and the multi-layer perceptron (MLP) to achieve the intelligent integration of robot task-related information.

[0132] The multi-task conditional fusion model receives three key inputs: the previous task instruction features (i.e., historical task instruction features, providing historical task context) encoded by a large language model (LLM) with a parameter scale of 7B, the current task instruction features (clarifying the immediate task goal), and the current environmental state features.

[0133] Using the current environmental state features as the query vector, and the current task instruction and the previous task instruction as the key-value pair, the multi-task conditional fusion model calculates the correlation weights between features using the cross-attention mechanism to achieve the adaptive focusing of task-related information. This design can automatically ignore irrelevant information and focus on "the environmental state and historical instructions closely related to the current operation".

[0134] The high-dimensional features after attention fusion are then compressed and non-linearly enhanced through a multi-layer MLP network containing the GeLU activation function, and finally, unified multi-task conditional features are output.

[0135] The multi-task conditional feature not only preserves the semantic integrity of the original information, but also improves the operation efficiency of the subsequent action expert model through dimensionality reduction. As the core information hub, the multi-task conditional features generated by the multi-task conditional fusion model will be input into the action expert model together with the fusion features, the robot state features, and the robot action noise, ensuring that the finally generated optimal action sequence simultaneously satisfies the alignment of the task context instructions and the robot environmental state, and providing key support for the intelligent decision-making of the robot in complex scenarios.

[0136] The above is the specific content of "III. Example Illustration of the Multi-Task Conditional Fusion Model".

[0137] Generally speaking, this solution aims to further improve the robot's understanding and execution capabilities of complex tasks through deep language instructions and current task status information, enabling the robot to handle complex situations encountered during task switching.

[0138] Furthermore, this solution effectively solves the complexity problem during robot task switching by integrating the encoding capabilities, state parsing, and task instruction fusion of multi-modal large models, and can dynamically generate adapted actions based on the previous task instruction, the current task instruction, and the current environmental state (such as the execution stage, environmental feedback). Specifically, through the fusion of multi-perspective image parsing and language instruction feature extraction, combined with task conditional distillation and state encoding, it can accurately identify task stage differences such as "not in contact / picked up / placed", and use the denoising flow matching technology to decide the optimal action that conforms to the context, thus overcoming the limitations of traditional methods that rely on independent task execution and fixed initial positions, and significantly improving the robustness and task coherence of the robot under arbitrary interruptions and non-regular state switches.

[0139] Corresponding to the multi-modal data processing method for supporting multi-task switching provided in the above embodiments of the present invention, see Figure 6 , the embodiments of the present invention also provide a structural block diagram of a multi-modal data processing system for supporting multi-task switching. The multi-modal data processing system includes: an acquisition unit 100, a generation unit 200, an extraction unit 300, and a decision unit 400;

[0140] The acquisition unit 100 is used to acquire the environmental image of the current environment where the robot is located, the historical task instruction, the current task instruction, the robot state, and the robot action noise.

[0141] In a specific implementation, the process of the acquisition unit 100 acquiring the environmental image of the current environment where the robot is located includes: calling the first camera and the second camera to acquire the environmental image of the current environment where the robot is located. The first camera has a first perspective and is mounted on the robot, and the second camera has a global perspective and is set in the environment where the robot is located.

[0142] In another specific implementation, the process of the acquisition unit 100 for acquiring the robot state includes: calling the sensors of the robot body to obtain the robot state, where the robot state at least includes: the joint angles of the robot and the state of the end effector.

[0143] The generation unit 200 is configured to generate a fusion feature and a multi-task conditional feature by using the environmental image, the historical task instruction, and the current task instruction.

[0144] Among them, the fusion feature is generated based on the following features: the visual feature extracted from the environmental image and the current environmental state feature, and the task instruction feature extracted from the historical task instruction and the current task instruction; the multi-task conditional feature is generated based on the current environmental state feature and the task instruction feature.

[0145] The extraction unit 300 is configured to extract the robot state feature from the robot state.

[0146] In a specific implementation, the extraction unit 300 is specifically configured to: use a robot state encoding model to extract the robot state feature from the robot state.

[0147] The decision-making unit 400 is configured to input the robot action noise, the fusion feature, the multi-task conditional feature, and the robot state feature into a preset action expert model for action decision-making to obtain the optimal action sequence of the robot.

[0148] Preferably, the decision-making unit 400 is further configured to: perform denoising processing on the optimal action sequence.

[0149] In the embodiment of the present invention, obtaining the optimal action sequence of the robot based on the historical task instruction, the current task instruction, and the environmental image effectively solves the complexity problem when the robot switches tasks, and improves the accuracy of the robot action output during task switching.

[0150] Preferably, in combination with Figure 6 As shown in the content, the generation unit 200 includes: an extraction subunit, a first input subunit, a second input subunit, a third input subunit, a fusion subunit, and an integration subunit; the execution principles of each subunit are as follows:

[0151] The extraction subunit is configured to extract the visual feature from the environmental image by using an image encoding model.

[0152] The first input subunit is configured to input the environmental image and a preset prompt word into a vision-language large model to obtain the current environmental state.

[0153] The second input subunit is configured to input the historical task instruction and the current task instruction into a large language model to obtain the corresponding task instruction feature.

[0154] A third input subunit, configured to input the current environmental state into a large language model to obtain corresponding current environmental state features.

[0155] A fusion subunit, configured to use a specified encoder to fuse visual features, task instruction features, and current environmental state features to obtain fused features.

[0156] An integration subunit, configured to input task instruction features and current environmental state features into a multi-task conditional fusion model for information integration to obtain multi-task conditional features.

[0157] Preferably, the integration subunit is further configured to: distill the multi-task conditional features to obtain compressed multi-task conditional features.

[0158] Preferably, an embodiment of the present invention further provides an electronic device, including: a processor and a memory, the processor and the memory are connected through a communication bus; wherein, the processor is configured to call and execute a program stored in the memory; the memory is configured to store a program, and the program is used to implement the multi-modal data processing method supporting multi-task switching provided in the above method embodiment.

[0159] Preferably, an embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it implements the multi-modal data processing method supporting multi-task switching provided in the above method embodiment.

[0160] In summary, an embodiment of the present invention provides a multi-modal data processing method and related devices supporting multi-task switching, which obtain the optimal action sequence of the robot based on historical task instructions, current task instructions, and environmental images, effectively solve the complexity problem when the robot performs task switching, and improve the accuracy of the robot actions output during task switching.

[0161] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for a system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The systems and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.

[0162] Those skilled in the art may further realize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.

[0163] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multimodal data processing method supporting multitask switching, characterized in that, The method includes: Obtaining an environmental image of the current environment where the robot is located, historical task instructions, current task instructions, robot state, and robot action noise; Generating a fusion feature and multi-task conditional features by using the environmental image, the historical task instructions, and the current task instructions; wherein, the fusion feature is generated based on the following features: visual features and current environmental state features extracted from the environmental image, and task instruction features extracted from the historical task instructions and the current task instructions; the multi-task conditional features are generated based on the current environmental state features and the task instruction features; Extracting robot state features from the robot state; Inputting the robot action noise, the fusion feature, the multi-task conditional features, and the robot state features into a preset action expert model for action decision-making to obtain the optimal action sequence of the robot.

2. The method according to claim 1, characterized in that, Generating a fusion feature and multi-task conditional features by using the environmental image, the historical task instructions, and the current task instructions, including: Extracting visual features from the environmental image by using an image encoding model; Inputting the environmental image and a preset prompt into a vision-language large model to obtain the current environmental state; Inputting the historical task instructions and the current task instructions into a large language model to obtain corresponding task instruction features; Inputting the current environmental state into the large language model to obtain corresponding current environmental state features; Using a specified encoder to fuse the visual features, the task instruction features, and the current environmental state features to obtain a fusion feature; Inputting the task instruction features and the current environmental state features into a multi-task conditional fusion model for information integration to obtain multi-task conditional features.

3. The method according to claim 1, characterized in that, Extracting robot state features from the robot state, including: Extracting robot state features from the robot state by using a robot state encoding model.

4. The method according to claim 2, characterized in that After obtaining the multi-task conditional features, it further includes: Distilling the multi-task conditional features to obtain the compressed multi-task conditional features.

5. The method according to claim 1, wherein The process of obtaining the environmental image of the current environment where the robot is located includes: Invoking a first camera and a second camera to obtain the environmental image of the current environment where the robot is located, the first camera has a first view angle and is mounted on the robot, and the second camera has a global view angle and is set in the environment where the robot is located.

6. The method according to claim 1, wherein The process of obtaining the robot state includes: Invoking robot body sensors to obtain the robot state, the robot state at least includes: joint angles of the robot, end effector state.

7. According to the method described in any one of claims 1-6, characterized in that, After obtaining the optimal action sequence of the robot, it further includes: Performing denoising processing on the optimal action sequence.

8. A multi-modal data processing system supporting multi-task switching, characterized in that, The system includes: An acquisition unit for obtaining an environmental image of the current environment where the robot is located, historical task instructions, current task instructions, robot state, and robot action noise; A generation unit, configured to generate a fusion feature and a multi-task conditional feature by using the environmental image, the historical task instruction, and the current task instruction; wherein, the fusion feature is generated based on the following features: a visual feature and a current environmental state feature extracted from the environmental image, and a task instruction feature extracted from the historical task instruction and the current task instruction; the multi-task conditional feature is generated based on the current environmental state feature and the task instruction feature; An extraction unit, configured to extract a robot state feature from the robot state; A decision-making unit, configured to input the robot action noise, the fusion feature, the multi-task conditional feature, and the robot state feature into a preset action expert model for action decision-making, so as to obtain an optimal action sequence of the robot.

9. An electronic device, characterized in that, Comprising: A processor and a memory, the processor and the memory are connected through a communication bus; wherein, the processor is configured to call and execute a program stored in the memory; The memory is configured to store a program, and the program is used to implement the multi-modal data processing method supporting multi-task switching as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the multi-modal data processing method supporting multi-task switching as described in any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Method and device for inspecting the condition of vacuum suction units of gripping device

    CN108621190A

  • Controller task switching method, device and equipment and readable storage medium

    CN110597601A

  • Microgrid cluster system task switching smooth transition control method and system

    CN118074100A

  • Terminal equipment and retrieval method

    CN118779402A

  • Motion mode switching method and device, service robot and medium

    CN119115978A

Cited By

  • Robot control method and device and storage medium

    CN120921403A

  • Modal fusion-based end-to-end robot arm claw control method

    CN121061869A

  • Robot control method and system based on visual language action intelligent agent

    CN122210661A