A service robot long-sequence complex task planning method and device based on multi-modal thought chain reasoning

CN122807853APending Publication Date: 2026-09-25TONGJI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610758042.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-29
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

1.长序列信息处理效率低:基于Transformer的架构在处理长时间视频流时计算复杂度呈二次方增长,难以实时融合长时序观测数据;

Benefits of technology

1. 本发明采用Mamba架构替代传统Transformer,将长序列建模的计算复杂度从二次方降至线性,在有限算力下支持更长时间窗口的视觉信息融合,提升对动态环境变化的感知效率,具有高效的长序列感知能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122807853A_ABST
    Figure CN122807853A_ABST
Patent Text Reader

Abstract

The application relates to a service robot long-sequence complex task planning method and equipment based on a multi-modal thought chain reasoning, which comprises the following steps: extracting an image embedding sequence and a text embedding sequence; mapping each sequence to a unified input space by using a projection layer to generate a multi-modal unified input; performing structured reasoning based on a multi-modal mixed expert model to generate multiple action branches and perform comprehensive scoring, and selecting the highest score as the current optimal planning; converting the current optimal planning into a robot action trajectory based on a diffusion model, and comparing the visual feedback with a preset intermediate state in the thought chain in real time during the execution process; if a deviation is detected, performing logical backtracking and dynamic re-planning; repeating the above steps until the long-sequence task is completed. Compared with the prior art, the application improves the semantic understanding depth, reasoning efficiency and execution robustness of the robot in a complex dynamic environment, and is suitable for autonomous decision-making and task execution of a household service robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot task planning, and in particular to a method and device for planning long-sequence complex tasks for service robots based on multimodal thought chain reasoning. Background Technology

[0002] Currently, service robots primarily rely on visual perception and simple instruction parsing to perform basic tasks such as object grasping and movement. Traditional task planning methods are mostly based on symbolic programming or finite state machines, requiring predefined strict rules and environmental templates. While these methods can provide deterministic execution flows, they struggle to adapt to unpredictable factors in unstructured environments such as homes, including randomly placed objects and dynamic obstacles, thus limiting task success rates. On the other hand, while end-to-end large-scale model-based planning methods can improve the understanding of natural language instructions, the lack of effective physical constraints and long-range logical reasoning mechanisms can easily lead to illusory outputs that do not conform to physical laws, making it difficult to guarantee reliable closed-loop execution of long, complex tasks.

[0003] In recent years, multimodal large-scale models have integrated visual, linguistic, and robot control information, enabling more comprehensive environmental perception and operational reasoning. However, existing technologies still have the following problems: 1. Low efficiency in processing long sequence information: The computational complexity of the Transformer-based architecture increases quadratically when processing long video streams, making it difficult to fuse long time-series observation data in real time; 2. Lack of interpretability and task specialization in reasoning process: General large models are often "large but not precise" in robot operation tasks. They lack focused reasoning ability for tasks such as navigation and grasping, and the reasoning process is not transparent, making it difficult to trace and correct errors. 3. Poor physical consistency of planning results: Directly deriving action sequences from instructions can easily overlook robot kinematic constraints, leading to infeasible planning or collisions during execution.

[0004] Chinese patent application CN121625159A discloses a robot control method and device, a robot, and a storage medium. The method involves using a target visual language model to perform semantic analysis on the current long-sequence task and current multi-view images to obtain multimodal semantic information; using the target visual language model and multimodal semantic information to break down the current long-sequence task into sub-steps to obtain at least one predicted sub-step and step description information for the predicted sub-step; using a target motion planning model to plan the motion trajectory of each predicted sub-step to obtain motion trajectory planning information; and controlling the target robot based on the motion trajectory planning information. This invention improves the stability, accuracy, and coherence of robot motion execution, but it still relies on a traditional Transformer-like structure, resulting in high computational overhead and high inference latency. It cannot support infinitely long temporal environmental perception and is an open-loop task planning method without visual feedback verification during execution, making it unable to detect sudden environmental changes or motion execution deviations.

[0005] Existing methods mostly employ open-loop programming, generating all steps at once without fully considering the dynamic changes in the environment during execution. For long-sequence complex tasks, how to achieve dynamic environment perception and planning correction through multimodal fusion and thought chain reasoning to ensure the robustness and adaptability of the system remains a key research challenge. Therefore, there is an urgent need for a service robot long-sequence complex task planning method based on multimodal thought chain reasoning, which can achieve more intelligent and reliable execution of long-sequence home service tasks. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of the existing technology and provide a method and device for planning long-sequence complex tasks for service robots based on multimodal thinking chain reasoning.

[0007] The objective of this invention can be achieved through the following technical solutions: A method for planning long-sequence complex tasks for service robots based on multimodal thought chain reasoning, the method comprising: Step S1: Obtain visual images of the unstructured home environment and the user's natural language commands through the visual acquisition device and human-computer interaction interface, and extract image embedding sequences and text embedding sequences based on the visual encoder and language encoder respectively. Step S2: Using a preset projection layer, the image embedding sequence and the text embedding sequence are mapped to a unified input space to obtain visual shared representations and text shared representations and splice them together to generate a multimodal unified input. Step S3: Based on the pre-built multimodal hybrid expert model and the multimodal unified input, structured reasoning is performed through the thinking chain prompting mechanism to obtain long sequence tasks and their sub-goals, generate multiple action branches, and comprehensively score the logical rationality, consistency of multimodal observations and estimation success rate of each action branch, and select the branch with the highest score as the current optimal plan. Step S4: Based on the diffusion model, the current optimal plan is transformed into a continuous and smooth robot motion trajectory, which is sent to the robot's underlying controller for execution. During the execution process, the visual feedback and the preset intermediate states in the thought chain are compared in real time. If an execution deviation or environmental change is detected, logical backtracking and dynamic replanning are performed. Step S5: Repeat steps S1-S4 until all long sequence tasks are completed; wherein, the visual image is updated in real time after each action is performed, and the reasoning context of the thought chain prompting mechanism is updated according to the new visual image.

[0008] Furthermore, both the visual encoder and the language encoder adopt the Mamba architecture, which includes a hardware-aware selective scanning algorithm. This algorithm uses discretization parameters to map a continuous-time system to a discrete-time system, and processes long video stream inputs with constant memory usage during inference.

[0009] Furthermore, the projection layer includes a perceptual resampler structure, and the specific process of step S2 includes: By using a pre-defined set of learnable latent query vectors to perform cross-attention calculation with the image embedding sequence, variable-length visual features are compressed into fixed-length visual representations, filtering out visual redundancy information in the environment. The visual representation is dynamically filtered and transformed using a preset adaptive gated linear unit to generate a visual shared representation that is compatible with the input space of a multimodal hybrid expert model. The text embedding sequence is processed by linear projection and layer normalization to generate a text shared representation; The visual shared representation and the text shared representation are concatenated to obtain a multimodal unified input.

[0010] Furthermore, the multimodal hybrid expert model includes a visual understanding expert network, a logical reasoning expert network, a planning and generation expert network, and a gating network. The gating network adopts a Top-K routing strategy, activating the K expert networks with the highest weights for each layer's input token to participate in the calculation of expert weights.

[0011] Furthermore, the structured reasoning process using the thought chain prompting mechanism in step S3 includes: Based on the aforementioned visual understanding expert network, a natural language description of the current environment is generated according to the multimodal unified input; the natural language description of the current environment includes scene objects, spatial relationships, and task-related initial states; Based on the aforementioned logical reasoning expert network, long sequence tasks and their sub-goals are parsed according to the natural language description of the current environment, the dependencies between sub-tasks are identified, physical constraints are evaluated, and intermediate causal reasoning is output. Based on the planning and generation expert, multiple action branches are generated in parallel using a tree-like search strategy based on intermediate causal reasoning. Based on the gating network, the logical rationality of each action branch, its consistency with multimodal observations, and the estimation success rate are comprehensively scored, and the branch with the highest score is selected as the current optimal plan.

[0012] Furthermore, in step S3, when the multimodal hybrid expert model performs structured reasoning, it also uses a pre-built neural symbol memory bank to retrieve the most relevant historical cases in the neural symbol memory bank as contextual hints through maximum inner product search, thereby guiding the expert network to generate a plan that conforms to the constraints of the current scenario. The neural symbolic memory bank is a vector database that stores the thought chain trajectories of successful historical interactions.

[0013] Furthermore, the diffusion model is a conditional denoising diffusion probability model, which uses the current optimal plan and the current visual image obtained by structured reasoning as conditions to generate a multimodal motion distribution in the robot joint space or Cartesian space by repeatedly denoising from Gaussian noise.

[0014] Furthermore, in step S4, the alignment score between the semantic features of visual observation and the semantic features of the preset intermediate state of the thought chain is calculated in real time. The alignment score is compared with a preset score threshold. If the alignment score is lower than the preset score threshold, it is determined to be an execution deviation or environmental change. The deviation signal is fed back to the multimodal hybrid expert model, and a tree search is used to make local corrections or global replanning based on the original optimal action sequence to generate a new optimal action sequence.

[0015] Furthermore, in the process of repeating steps S1-S4 in step S5, adaptive feedback update based on semantic difference detection is also included. The adaptive feedback update process includes: By comparing the semantic features of real-time visual observation with the semantic features of the thought chain prediction state, the semantic offset is calculated. When the semantic offset exceeds the threshold, the parameters of the gating network in the multimodal hybrid expert model are adjusted using the dominance function in reinforcement learning to enhance the logical robustness of the model when handling similar long sequence tasks.

[0016] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the service robot long-sequence complex task planning method based on multimodal thought chain reasoning as described above.

[0017] Compared with the prior art, the beneficial effects of the present invention include: 1. This invention uses the Mamba architecture to replace the traditional Transformer, reducing the computational complexity of long sequence modeling from quadratic to linear. It supports visual information fusion for longer time windows under limited computing power, improves the perception efficiency of dynamic environmental changes, and has a highly efficient long sequence perception capability.

[0018] 2. This invention integrates expert networks for visual understanding, logical reasoning, and planning generation through a multimodal hybrid expert model, and uses a gating mechanism to activate task-oriented experts, thereby improving the accuracy and interpretability of reasoning.

[0019] 3. This invention combines tree-based search and value assessment to ensure that the output action sequence conforms to logical constraints and scenario feasibility. It also introduces historical successful experiences through a neural symbolic memory bank to enhance the reusability and generalization ability of the planning, thus achieving physically consistent planning generation.

[0020] 4. This invention establishes a closed-loop verification mechanism based on visual feedback, which detects planning execution deviations in real time and triggers replanning, realizing a dynamic adjustment cycle of perception-planning-execution-verification, improving the robustness of the system in unstructured environments, and achieving closed-loop adaptive planning adjustment.

[0021] 5. This invention, through a retrieval-enhanced generation mechanism, can quickly retrieve the optimal planning trajectory of similar historical tasks, reduce the overhead of repetitive reasoning, improve the planning efficiency and learning ability of long-sequence complex tasks, and accelerate experience-guided planning. Attached Figure Description

[0022] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0024] Example 1 This embodiment discloses a method for planning long-sequence complex tasks for service robots based on multimodal thought chain reasoning. The method is as follows: Figure 1 As shown, steps S1-S5 are included, and each step is described in detail below, including: Step S1: Obtain visual images of the unstructured home environment and the user's natural language commands through a visual acquisition device and a human-computer interaction interface, and extract image embedding sequences and text embedding sequences based on the visual encoder and the language encoder, respectively.

[0025] Specifically, environmental video streams are acquired through visual acquisition devices. ,in This refers to the timing length. The system receives user natural language commands via the voice interaction module. .

[0026] Both the visual encoder and the language encoder adopt a Mamba-based architecture, mapping visual input and text input to image embedding sequences, respectively. With text embedding sequence To extract long-term temporal environmental features and task semantic features, the visual encoder employs a two-dimensional flattening strategy to process video frame sequences into long-sequence tokens.

[0027] The Mamba architecture includes a hardware-aware selective scan algorithm that uses discretization parameters to map continuous-time systems to discrete-time systems, processing long video stream inputs with constant memory usage during inference.

[0028] The visual feature encoding process can be formally represented as: in, This indicates that video frames are mapped to the initially embedded linear projection layer. The parameter is A selective state-space model encoder is used to capture the spatiotemporal dynamic features of visual sequences.

[0029] The language feature encoding process can be formally represented as: in, This indicates that text instructions are being segmented and embedded initialized. The parameter is A selective state-space model encoder is used to extract semantic features and long-distance dependencies from instructions.

[0030] Step S2: Using a preset projection layer, the image embedding sequence and the text embedding sequence are mapped to a unified input space to obtain visual shared representations and text shared representations and splice them together to generate a multimodal unified input.

[0031] Specifically, a projection layer is used to embed the visual sequence. With text embedding sequence Mapping to a unified semantic space of large language models .

[0032] This projection layer combines a perceptual resampler with an adaptive gated linear unit (GLU) to filter and compress long sequences of visual features, generating fixed-length visual representations. and text representation Perform semantic alignment.

[0033] The specific process of step S2 includes: By using a pre-defined set of learnable latent query vectors and image embedding sequences to perform cross-attention calculation, variable-length visual features are compressed into fixed-length visual representations, filtering out redundant visual information in the environment. By using a preset adaptive gated linear unit to dynamically filter and transform visual representations, a shared visual representation compatible with the input space of a multimodal hybrid expert model is generated. Text embedding sequences are processed by linear projection and layer normalization to generate shared text representations; By concatenating visual shared representations and text shared representations, a unified multimodal input can be obtained.

[0034] Specifically, the perceptual resampler uses a set of learnable query vectors. Cross-attention compression of visual embeddings: in, To ensure a fixed number of output tokens, For feature dimension, These are the compressed visual features.

[0035] Adaptive gated linear units dynamically filter compressed features to achieve cross-modal alignment: in, For learnable projection matrices, It is the Sigmoid activation function. Representation layer normalization, This represents element-wise multiplication. The final multimodal unified input is... .

[0036] Step S3: Based on the pre-built multimodal mixture of experts (MoE) model and multimodal unified input, structured reasoning is performed through the thought chain prompting mechanism to obtain long sequence tasks and their sub-goals, generate multiple action branches, and comprehensively score the logical rationality, consistency of multimodal observations and estimation success rate of each action branch, and select the branch with the highest score as the current optimal plan.

[0037] Specifically, driven by the multimodal thinking chain prompting mechanism, the multimodal hybrid expert model executes an explicit, structured observation-reasoning-planning multimodal reasoning loop to generate hierarchical task planning.

[0038] The multimodal hybrid expert model includes a visual understanding expert network, a logical reasoning expert network, a planning and generation expert network, and a gating network. The gating network adopts a Top-K routing strategy, which activates the K expert networks with the highest weights for each layer of input tokens to participate in the calculation of expert weights.

[0039] Step S3, which involves structured reasoning through the thought chain prompting mechanism, includes: Based on a visual understanding expert network, a natural language description of the current environment is generated from a unified multimodal input. The natural language description of the current environment includes scene objects, spatial relationships, and task-related initial states. Based on a logical reasoning expert network, it parses long sequence tasks and their sub-goals, identifies sub-task dependencies, evaluates physical constraints, and outputs intermediate causal reasoning based on the natural language description of the current environment. Based on planning and generation experts, multiple action branches are generated in parallel using a tree-like search strategy based on intermediate causal reasoning. The gated network comprehensively scores the logical rationality of each action branch, its consistency with multimodal observations, and the estimation success rate, and selects the branch with the highest score as the current optimal plan.

[0040] In step S3, when the multimodal hybrid expert model performs structured reasoning, it also uses a pre-built neural symbol memory bank to retrieve the most relevant historical cases in the neural symbol memory bank as contextual cues through maximum inner product search, guiding the expert network to generate a plan that conforms to the constraints of the current scenario.

[0041] The neural symbolic memory bank is a vector database that stores the thought chain trajectories of successful historical interactions.

[0042] Specifically, the model is first based on visual features. With semantic features Generate a natural language description of the current environment. This process is led by a visual understanding expert network, whose output integrates objects in the scene, spatial relationships, and initial states related to the instructions. With instructions and observation description Given the conditions, the logic reasoning expert network is activated to perform causal and constraint reasoning. It resolves the task objective, identifies potential subtask dependencies, evaluates physical constraints, and outputs intermediate reasoning steps. : Planning Generates Expert Network Integration and This generates the final action plan. This stage employs a tree-based search strategy: first, multiple candidate atomic action sequence branches are generated. Each branch This represents a complete action plan. The value network within the model. In other words, a gated network evaluates each branch, and the scoring function comprehensively considers logical rationality, consistency with multimodal observations, and the success rate of estimation. Finally, the branch with the highest score is selected as the optimal program output for the current step: This thought process It is fully recorded for subsequent verification and memorization.

[0043] Step S4: Based on the diffusion model, the current optimal plan is transformed into a continuous and smooth robot motion trajectory, which is sent to the robot's underlying controller for execution. During the execution process, the visual feedback and the preset intermediate states in the thought chain are compared in real time. If an execution deviation or environmental change is detected, logical backtracking and dynamic replanning are performed.

[0044] The diffusion model is a conditional denoising diffusion probability model. It uses the current optimal plan and the current visual image obtained by structured reasoning as conditions, and generates a multimodal motion distribution in the robot joint space or Cartesian space by repeatedly denoising from Gaussian noise.

[0045] In step S4, the alignment score between the semantic features of visual observation and the semantic features of the preset intermediate state of the thought chain is calculated in real time. The alignment score is compared with the preset score threshold. If the alignment score is lower than the preset score threshold, it is determined to be an execution deviation or environmental change. The deviation signal is fed back to the multimodal hybrid expert model. Tree search is used to make local corrections or global replanning based on the original optimal action sequence to generate a new optimal action sequence.

[0046] Specifically, the atomic action sequence instructions generated by S3 This is transformed into specific actions that the robot can perform, a process achieved through a conditional denoising diffusion probability model.

[0047] The model obtains multimodal observations from previous steps. With multimodal thinking chain Based on these conditions, two main categories of control commands are generated: one is discrete symbolic actions (such as grabbing or opening), and the other is continuous, smooth motion trajectory parameters in joint space or task space. Specifically, for the commands... The diffusion model generates the corresponding action parameter set. This includes the target pose, force control parameters, and motion velocity. The inverse denoising process can be formally represented as: Finally, physically consistent and executable action instructions are sent to the underlying controller.

[0048] After the action is executed, the system verifies the execution effect through real-time visual feedback, and the visual encoder captures the new environmental state. And extract features The state verification module calculates the expected state semantic description. semantic features of actual observation Alignment score between .like ( If the threshold is set to a preset value, it is considered an execution deviation, triggering a replanning mechanism: the system feeds back the deviation signal to the MoE model, and the gating network reroutes accordingly, activates different expert combinations, and uses tree search to make local corrections or global replanning based on the original plan, generating a new action sequence to deal with emergencies.

[0049] Step S5: Repeat steps S1-S4 until all long sequence tasks are completed; wherein, the visual image is updated in real time after each action is performed, and the reasoning context of the thought chain prompting mechanism is updated according to the new visual image.

[0050] In step S5, during the repetition of steps S1-S4, an adaptive feedback update based on semantic difference detection is also performed. The adaptive feedback update process includes: By comparing the semantic features of real-time visual observation with the semantic features of the thought chain prediction state, the semantic offset is calculated. When the semantic offset exceeds the threshold, the parameters of the gating network in the multimodal hybrid expert model are adjusted using the advantage function in reinforcement learning to enhance the logical robustness of the model when handling similar long sequence tasks.

[0051] Each action execution and environmental state update constitutes a complete perception-planning-execution-verification loop. Specifically, after each action execution, the environmental state is updated in real time, and the system uses this updated multimodal observation as the input context for the next round of inference. Simultaneously, the memory stores the successful thought chain trajectory and corresponding environmental state features as a success case. In subsequent task planning, when the MoE model performs inference, it can quickly retrieve the most similar historical success case from the memory through maximum inner product search, injecting it as a contextual cue to guide the expert network to generate plans more efficiently and reliably. Furthermore, the system performs adaptive optimization based on long-term execution feedback: by continuously comparing the semantic differences between the predicted and actual states, the average semantic offset is calculated. When this offset consistently exceeds a threshold in a specific task category, the parameters of the MoE gating network are fine-tuned, dynamically adjusting expert selection preferences and enhancing the model's cumulative robustness and generalization ability in handling complex long-sequence tasks.

[0052] It should be noted that the selective state-space architecture employed by the visual and language encoder in step S1 above is crucial, as it utilizes a hardware-aware sequence modeling method. This method transforms a continuous system into a recursive computational form through a discretization process, enabling the model to achieve linear computational efficiency when processing long sequences of data such as video streams and long text. This addresses the real-time performance degradation problem caused by the dramatic increase in computational load when traditional attention mechanisms handle long-duration environmental information and complex user commands, providing an efficient perceptual foundation for subsequent steps.

[0053] In step S2 above, the adaptive gating mechanism and the perceptual resampling structure employed by the projection layer work together. The perceptual resampler filters and condenses the lengthy visual feature sequence using a set of learnable vectors, extracting a fixed number of key information markers. The adaptive gating unit dynamically adjusts the flow of visual information into the language model space, achieving deep fusion and alignment of cross-modal features, effectively filtering out a large number of irrelevant details in the environmental visual information, and improving the representation quality of task-related features.

[0054] In step S3 above, the collaborative design of the multimodal hybrid expert model and the tree-structured thought chain search is the core of achieving complex task decomposition and deep planning. The hybrid expert model intelligently calls dedicated modules such as visual understanding, logical calculation, and action planning for different input content through a gating network. The tree-structured search simulates a multi-branch reasoning process during the planning phase, and uses an internal value evaluation network to score the feasibility of each possible action path and select the best one, thereby ensuring that each generated atomic instruction is logically rigorous and conforms to the actual scenario.

[0055] In step S4 above, the conditional diffusion model transforms symbolic atomic instructions into specific action parameters that the robot can directly execute. This model not only generates smoothly varying motion trajectories within joints or task space but also determines the type, force, and target parameters of the action. Combined with a vision-based closed-loop verification mechanism, the system can determine in real time whether the state after action execution meets expectations. When deviations occur, it quickly triggers local adjustments in the planning module, forming an adaptive loop of "planning-execution-verification-correction" to cope with uncertainties in dynamic environments.

[0056] In step S5 above, the entire system completes long-sequence tasks by iteratively executing the above process. The key lies in the continuous updating of the state and the introduction of a neural symbolic memory. After each loop, new environmental observations are incorporated into the context, serving as the starting point for the next round of reasoning. The memory continuously accumulates historically successful task planning cases. When encountering similar scenarios, the system can quickly retrieve and call upon past experiences, providing a reference for the hybrid expert model, thereby significantly reducing repetitive reasoning and improving overall planning efficiency and task success rate.

[0057] Example 2 Based on Embodiment 1, this embodiment provides an electronic device, including: one or more processors and a memory, wherein the memory stores one or more programs, the one or more programs including instructions for executing the service robot long sequence complex task planning method based on multimodal thought chain reasoning as described above.

[0058] At the hardware level, the electronic device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to implement the aforementioned service robot long-sequence complex task planning method based on multimodal thought chain reasoning. Of course, in addition to software implementation, this invention does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution subject of the following processing flow is not limited to individual logic units, but can also be hardware or logic devices.

[0059] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0060] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0061] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for planning long-sequence complex tasks for service robots based on multimodal thought chain reasoning, characterized in that, The method includes: Step S1: Obtain visual images of the unstructured home environment and the user's natural language commands through a visual acquisition device and a human-computer interaction interface, and extract image embedding sequences and text embedding sequences based on the visual encoder and the language encoder, respectively. Step S2: Using a preset projection layer, the image embedding sequence and the text embedding sequence are mapped to a unified input space to obtain visual shared representations and text shared representations and splice them together to generate a multimodal unified input. Step S3: Based on the pre-built multimodal hybrid expert model and the multimodal unified input, structured reasoning is performed through the thinking chain prompting mechanism to obtain long sequence tasks and their sub-goals, generate multiple action branches, and comprehensively score the logical rationality, consistency of multimodal observations and estimation success rate of each action branch, and select the branch with the highest score as the current optimal plan. Step S4: Based on the diffusion model, the current optimal plan is transformed into a continuous and smooth robot motion trajectory, which is sent to the robot's underlying controller for execution. During the execution process, the visual feedback and the preset intermediate states in the thought chain are compared in real time. If an execution deviation or environmental change is detected, logical backtracking and dynamic replanning are performed. Step S5: Repeat steps S1-S4 until all long sequence tasks are completed; wherein, the visual image is updated in real time after each action is performed, and the reasoning context of the thought chain prompting mechanism is updated according to the new visual image.

2. The method for planning long-sequence complex tasks for service robots based on multimodal thought chain reasoning according to claim 1, characterized in that, Both the visual encoder and the language encoder adopt the Mamba architecture, which includes a hardware-aware selective scanning algorithm. The selective scanning algorithm uses discretization parameters to map the continuous-time system to a discrete-time system and processes long video stream inputs with constant memory usage during inference.

3. The method for planning long-sequence complex tasks for service robots based on multimodal thought chain reasoning according to claim 1, characterized in that, The projection layer includes a perceptual resampler structure, and the specific process of step S2 includes: By using a pre-defined set of learnable latent query vectors to perform cross-attention calculation with the image embedding sequence, variable-length visual features are compressed into fixed-length visual representations, filtering out visual redundancy information in the environment. The visual representation is dynamically filtered and transformed using a preset adaptive gated linear unit to generate a visual shared representation that is compatible with the input space of a multimodal hybrid expert model. The text embedding sequence is processed by linear projection and layer normalization to generate a text shared representation; The visual shared representation and the text shared representation are concatenated to obtain a multimodal unified input.

4. The method for planning long-sequence complex tasks for service robots based on multimodal thought chain reasoning according to claim 1, characterized in that, The multimodal hybrid expert model includes a visual understanding expert network, a logical reasoning expert network, a planning and generation expert network, and a gating network. The gating network adopts a Top-K routing strategy, which activates the K expert networks with the highest weights for each layer of input tokens to participate in the calculation of expert weights.

5. The method for planning long-sequence complex tasks for service robots based on multimodal thought chain reasoning according to claim 4, characterized in that, The structured reasoning process using the thought chain prompting mechanism in step S3 includes: Based on the aforementioned visual understanding expert network, a natural language description of the current environment is generated according to the multimodal unified input; the natural language description of the current environment includes scene objects, spatial relationships, and task-related initial states; Based on the aforementioned logical reasoning expert network, long sequence tasks and their sub-goals are parsed according to the natural language description of the current environment, the dependencies between sub-tasks are identified, physical constraints are evaluated, and intermediate causal reasoning is output. Based on the planning and generation expert, multiple action branches are generated in parallel using a tree-like search strategy based on intermediate causal reasoning. Based on the gating network, the logical rationality of each action branch, its consistency with multimodal observations, and the estimation success rate are comprehensively scored, and the branch with the highest score is selected as the current optimal plan.

6. The method for planning long-sequence complex tasks for service robots based on multimodal thought chain reasoning according to claim 1, characterized in that, In step S3, when the multimodal hybrid expert model performs structured reasoning, it also uses a pre-built neural symbol memory bank to retrieve the most relevant historical cases in the neural symbol memory bank as contextual hints through maximum inner product search, guiding the expert network to generate a plan that conforms to the constraints of the current scenario. The neural symbolic memory bank is a vector database that stores the thought chain trajectories of successful historical interactions.

7. The method for planning long-sequence complex tasks for service robots based on multimodal thought chain reasoning according to claim 1, characterized in that, The diffusion model is a conditional denoising diffusion probability model. It uses the current optimal plan and the current visual image obtained by structured reasoning as conditions to generate a multimodal motion distribution in the robot joint space or Cartesian space by repeatedly denoising from Gaussian noise.

8. The method for planning long-sequence complex tasks for service robots based on multimodal thought chain reasoning according to claim 1, characterized in that, In step S4, the alignment score between the semantic features of visual observation and the semantic features of the preset intermediate state of the thought chain is calculated in real time. The alignment score is compared with the preset score threshold. If the alignment score is lower than the preset score threshold, it is determined to be an execution deviation or environmental change. The deviation signal is fed back to the multimodal hybrid expert model. Tree search is used to make local corrections or global replanning based on the original optimal action sequence to generate a new optimal action sequence.

9. The method for planning long-sequence complex tasks for service robots based on multimodal thought chain reasoning according to claim 1, characterized in that, In the process of repeating steps S1-S4 in step S5, adaptive feedback update based on semantic difference detection is also included. The adaptive feedback update process includes: By comparing the semantic features of real-time visual observation with the semantic features of the thought chain prediction state, the semantic offset is calculated. When the semantic offset exceeds the threshold, the parameters of the gating network in the multimodal hybrid expert model are adjusted using the dominance function in reinforcement learning to enhance the logical robustness of the model when handling similar long sequence tasks.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the service robot long sequence complex task planning method based on multimodal thought chain reasoning as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Robot control method and device, robot and storage medium

    CN121625159A