A somatic intelligent task planning basic model for humanoid robots
Patent Information
- Application Number
- CN202610732455.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-26
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]1、语义理解肤浅:无法解析SOP中隐含的时序、因果与空间逻辑关系
[0023]1、提升任务成功率与鲁棒性:通过端到端的深度理解、精确关联、可行性验证和自适应执行,显著降低任务失败率,尤其在复杂、动态环境中。
Smart Images

Figure CN122829798A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of humanoid robots and artificial intelligence, and in particular to a basic model for embodied intelligent task planning for humanoid robots. Background Technology
[0002] In the fields of intelligent manufacturing, service robots, and special operations, humanoid robots need to perform complex tasks described by Standard Operating Procedures (SOPs) written in natural language. However, existing technologies struggle to achieve end-to-end automatic conversion from unstructured SOP text to sequences of robot-executable actions. The core bottleneck lies in the insufficient capabilities of existing systems in several areas, including deep semantic understanding, multimodal environmental perception and physical modeling, precise association between symbols and entities, adaptive execution of skills, and verification of the physical feasibility of planning schemes. This results in low success rates, robustness, and efficiency for robots when facing real, dynamic, and uncertain environments.
[0003] The existing technology has the following problems and drawbacks:
[0004] 1. Superficial semantic understanding: Unable to interpret the temporal, causal, and spatial logical relationships implied in the SOP.
[0005] 2. The environmental model is static and lacks physical properties: the perceived results are disconnected from the physical simulation, and the consequences of interaction cannot be predicted.
[0006] 3. Poor robustness of cross-modal association: It is prone to incorrect matching in complex scenarios.
[0007] 4. Rigidity of skill execution: Predefined skills cannot cope with changes in environmental disturbances and task constraints.
[0008] 5. Lack of physical verification for planning: The planned scheme may fail in actual implementation due to reasons such as collisions or inaccessibility.
[0009] 6. Disconnect between code generation and planning process: The generated code has poor readability, is difficult to debug, and cannot guarantee its alignment with the planning intent. Figure 1 To. Summary of the Invention
[0010] The purpose of this invention is to address the shortcomings of existing technologies and provide a basic model for embodied intelligent task planning for humanoid robots. This model can deeply understand SOP semantics, accurately perceive and physically model the environment, intelligently associate symbols and entities, adaptively adjust skill parameters, perform feasibility verification and replanning, and ultimately generate reliable and traceable platform-specific code, thereby achieving high success rate and highly adaptive autonomous execution of complex tasks for humanoid robots.
[0011] To achieve the above objectives, the present invention adopts the following technical solution:
[0012] A basic model for embodied intelligent task planning for humanoid robots includes, in sequence, a SOP deep structured semantic parsing module, a multimodal environment perception and modeling module, a cross-modal semantic-entity association module, a parameterized atomic skill library module, a hierarchical task planning module, a feasibility verification and replanning module, a complex geometric constraint solving module, a spatiotemporal trajectory optimization module, and an intelligent code generation and verification module. The SOP deep structured semantic parsing module transforms unstructured text SOPs into directed semantic graphs rich in semantic logic that can be directly computed by computers. The multimodal environment perception and modeling module constructs a dynamic digital environment model that can be used for physical simulation and collision detection. The cross-modal semantic-entity association module interprets the semantic symbols in the SOPs. The system includes: precise matching with entities in the physical environment; a parameterized atomic skill library module for defining intelligent skills that can adapt to environmental changes online and meet complex constraints; a hierarchical task planning module for generating high-level task sequences and sub-task decompositions with self-reflection and correction capabilities; a feasibility verification and replanning module for evaluating the feasibility of planning schemes in a physical simulation environment and triggering replanning when infeasibility occurs; a complex geometric constraint solving module for converting complex geometric constraints described in natural language into mathematical optimization problems and solving them; a spatiotemporal trajectory optimization module for generating dynamically feasible robot motion trajectories that adapt to dynamic environments; and an intelligent code generation and verification module for generating platform-specific code rich in semantic information and performing simulation verification.
[0013] Preferably, the SOP deep structured semantic parsing module includes: a predicate-argument structure extraction and relational reasoning unit, which uses a large language model fine-tuned with domain process knowledge to identify basic arguments such as core actions, tools, and objects, and infers the implicit temporal, causal, and spatial relationships between arguments; a process parameter and constraint extraction unit, which adopts a hybrid extraction strategy, using predefined rules and NER models to quickly extract explicit numerical parameters, and calling a large language model to perform semantic disambiguation and formal transformation for fuzzy natural language constraints; and a directed semantic graph construction and embedding unit, which constructs a directed semantic graph from the extracted entities and relations, and uses a semantic graph neural network to embed nodes and edges to generate a vectorized semantic knowledge base that can be used for subsequent logical reasoning.
[0014] Preferably, the multimodal environment perception and modeling module includes: an instance-level scene deconstruction and state classification unit, which uses an instance segmentation network to identify objects and simultaneously determines the dynamic state of the objects through classification branches; a six-degree-of-freedom pose estimation unit, which integrates multi-view geometry and neural radiation field technology to solve for the precise six-degree-of-freedom pose of the objects; and a dynamic scene semantic map construction unit, which imports the identified objects and their poses and attributes into a lightweight physics engine to construct a scene semantic map with physical simulation capabilities.
[0015] Preferably, the cross-modal semantic-entity association module includes: a multi-feature similarity calculation unit, which constructs a visual language model-assisted similarity evaluation framework and uses the common sense reasoning ability of the visual language model to obtain a comprehensive matching confidence score; an association matrix construction unit, which introduces the association history context when constructing the association matrix and integrates the order constraints in the SOP into the association process; and an optimal allocation strategy learning unit, which models association matching as a sequential decision problem and uses reinforcement learning to learn the optimal allocation strategy.
[0016] Preferably, the parameterized atomic skill library module includes: a skill-model mapping and parameter interface design unit, which designs a learnable parameter interface for each skill; a dynamic motion primitive enhancement unit, which introduces a simplified robot dynamics model as a constraint on the basis of the standard DMP and uses reinforcement learning to learn the residual adjustment amount; and a skill parameter instantiation unit, which introduces an online imitation learning mechanism to adjust the learnable parameters of the skill through behavior cloning or inverse reinforcement learning.
[0017] Preferably, the hierarchical task planning module includes: a high-level task sequence generation unit, in which the large language model actively queries an external process knowledge base using RAG technology when generating high-level task sequences; a subtask decomposition unit, in which the large language model generates explicit thought chains when decomposing subtasks, and performs self-criticism and correction on each step of the thought chain; and an atomic skill invocation and interface verification unit, which, by drawing on the idea of combinatorial reinforcement learning, verifies the interface compatibility between subtasks.
[0018] Preferably, the feasibility verification and replanning module includes: a feasibility assessment unit, which places the sub-task and its corresponding initial skill parameters in a scene semantic map with a physics engine for forward simulation; a multi-dimensional feasibility judgment unit, which introduces a semi-online reinforcement learning judgment mechanism to make a decision by comprehensively considering simulation results, visual language model confidence, and long-term value assessment; and a semantically guided replanning unit, which initiates a targeted exploration strategy when the sub-task is determined to be infeasible and feeds back successful exploration experiences to the large language model.
[0019] Preferably, the complex geometric constraint solving module includes: a constraint understanding and optimization objective formalization unit, in which the large language model transforms the natural language constraints in the SOP into a computable optimization objective description; an optimization algorithm guidance unit, in which the large language model provides a high-quality initial population or search direction for the optimization algorithm based on its understanding of the problem; and an optimization result verification unit, in which the large language model performs semantic interpretation and feasibility posterior verification on the optimal operation sequence obtained from the optimization solution.
[0020] Preferably, the spatiotemporal trajectory optimization module includes: an online motion trajectory generation unit that uses a model predictive control framework to resolve the optimal control problem in the finite time domain in each control cycle; a trajectory optimization strategy learning unit that uses deep reinforcement learning to train a neural network and outputs the optimal potential field parameters or direct trajectory adjustment; and a spatiotemporal joint optimization unit that performs spatiotemporal joint optimization considering the robot's differential dynamics constraints.
[0021] Preferably, the intelligent code generation and verification module includes: an intermediate representation generation unit, which generates an intermediate representation containing action instructions, parameters, and semantic information from the original SOP and planning process; an API mapping and code generation unit, which uses the code understanding and generation capabilities of a large language model to automatically generate or verify the corresponding API call code fragments based on the intermediate representation; and a code-level behavior verification unit, which runs the generated executable code in a simulation environment and verifies the results by comparing the simulation execution results with the expected task objectives.
[0022] Compared with the prior art, the beneficial effects of this invention are as follows:
[0023] 1. Improve task success rate and robustness: Through end-to-end deep understanding, precise correlation, feasibility verification and adaptive execution, significantly reduce task failure rate, especially in complex and dynamic environments.
[0024] 2. Enhance system versatility and automation: Get rid of dependence on fixed templates and a large number of predefined rules, be able to handle unseen SOP descriptions and scenario changes, and achieve a higher degree of autonomy.
[0025] 3. Improved Development and Debugging Efficiency: The intelligent code generation and simulation verification closed loop significantly reduces the time and cost of manual programming and on-machine debugging. Semantic intermediate representations and annotations facilitate problem tracing.
[0026] 4. Optimize task execution performance: Through spatiotemporal joint optimization of trajectory and adaptive skills, the robot's movement becomes more efficient, energy-saving, and stable.
[0027] 5. Reduce reliance on domain experts: The common sense and knowledge retrieval capabilities of large language models and visual language models partially replace the role of domain experts in rule making and knowledge injection.
[0028] 6. Possesses continuous learning and evolution capabilities: Through online imitation learning, reinforcement learning strategies, and experience feedback during replanning, the system can learn from the execution results and continuously improve performance.
[0029] In summary, the model in this invention can deeply understand SOP semantics, accurately perceive and physically model the environment, intelligently associate symbols and entities, adaptively adjust skill parameters, perform feasibility verification and replanning, and ultimately generate reliable and traceable platform-specific code, thereby achieving high success rate and highly adaptive autonomous execution of complex tasks for humanoid robots. Attached Figure Description
[0030] Figure 1 This is a logical framework diagram of the basic model for embodied intelligent task planning for humanoid robots proposed in this invention. Detailed Implementation
[0031] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0032] Reference Figure 1 A basic model for embodied intelligent task planning for humanoid robots includes, in sequence, a SOP deep structured semantic parsing module, a multimodal environment perception and modeling module, a cross-modal semantic-entity association module, a parameterized atomic skill library module, a hierarchical task planning module, a feasibility verification and replanning module, a complex geometric constraint solving module, a spatiotemporal trajectory optimization module, and an intelligent code generation and verification module.
[0033] Example 1: Deep Structured Semantic Parsing of SOP
[0034] The SOP deep structured semantic parsing module includes a predicate-argument structure extraction and relational reasoning unit, a process parameter and constraint extraction unit, and a directed semantic graph construction and embedding unit.
[0035] First, the predicate-argument structure extraction and relational reasoning unit uses a large language model (such as GPT-4, Llama3, etc.) fine-tuned with domain-specific technical knowledge as the core engine for semantic role labeling. For input unstructured SOP text, such as "Use a torque wrench to tighten four M8 bolts in a cross-shaped sequence with a torque of 15 N·m, and then check the seal," the large language model not only identifies the core actions "tighten" and "check," the tool "torque wrench," and the objects "four M8 bolts" and "seal," but also infers the implicit temporal relationships ("tighten" before "check"), spatial relationships ("cross-shaped sequence"), and causal relationships ("tighten" is a prerequisite for "check the seal") between arguments, and represents these relationships as directed edges.
[0036] Next, the process parameter and constraint extraction unit adopts a hybrid extraction strategy. For the explicit numerical parameter "15N·m", it is quickly extracted using predefined rules and the NER model; for the fuzzy natural language constraint "cross-shaped sequence", a large language model is called to perform semantic disambiguation and formal transformation, outputting it as a machine-readable logical expression of {sequence pattern: cross-shaped, geometric constraint: vertical angle priority}.
[0037] Finally, the extracted entities and relations are constructed into a directed semantic graph. Then, semantic graph neural networks are used to learn the embeddings of nodes and edges. Here, graph convolutional neural networks (GCNs) are used to achieve feature aggregation. The core formula is:
[0038]
[0039] in Let v be the feature vector of the l-th node. For the normalized adjacency matrix, , For convolution parameters, The ReLU activation function is used. This embedding not only encodes textual information but also captures the topological relationships in the graph structure, generating a vectorized semantic knowledge base that can be used for subsequent logical reasoning.
[0040] Example 2: Multimodal Environment Perception and Modeling
[0041] The multimodal environment perception and modeling module includes an instance-level scene deconstruction and state classification unit, a six-degree-of-freedom pose estimation unit, and a dynamic scene semantic map construction unit.
[0042] First, the instance-level scene deconstruction and state classification unit uses instance segmentation networks (such as Mask R-CNN, YOLOv8, etc.) to identify objects in the scene, and simultaneously determines the dynamic state of the objects through a classification branch. For example, for a bolt, the classification branch will output a state label of "tightened" or "loose". This perception of object state provides crucial environmental context for task planning.
[0043] Next, the six-DOF pose estimation unit integrates multi-view geometry and neural radiation field techniques. For simple objects, traditional multi-view geometry methods are used to solve the pose; for complex or heavily occluded objects, multi-view visual fusion techniques are employed. The pose transformation of the object from the camera coordinate system C to the world coordinate system W follows the formula:
[0044]
[0045] in It is a 3×3 rotation matrix. It is a 3×1 translation vector. , These are the coordinates of the object in the two coordinate systems, respectively.
[0046] Furthermore, neural radiation field (NeRF) technology can be introduced, through... (x is a spatial point, d is the direction of the light ray) Reconstruct a detailed three-dimensional geometric model of the object, thereby solving for a more accurate six-degree-of-freedom pose, providing a foundation for precise operation.
[0047] Finally, the dynamic scene semantic map construction unit imports the identified objects, their poses, and attributes into a lightweight physics engine (such as a simplified version of NVIDIA PhysX or Bullet) to construct a scene semantic map with physical simulation capabilities. The dynamic state updates of objects in the map follow Newton's second law:
[0048]
[0049] in Let i be the mass of object i. For acceleration, For collision force, This map provides a framework for robot interaction. It not only describes "where" objects are, but also predicts the physical reactions (such as collisions and pushes) that the robot will exhibit when interacting with them, thus providing an environment for subsequent simulation training.
[0050] Example 3: Cross-modal semantic-entity association
[0051] The cross-modal semantic-entity association module includes a multi-feature similarity calculation unit, an association matrix construction unit, and an optimal allocation strategy learning unit.
[0052] First, a multi-feature similarity calculation unit constructs a similarity evaluation framework assisted by a visual language model (such as CLIP, BLIP-2, etc.). When traditional features (text, attributes) cannot distinguish between objects with similar appearances, the scene image and candidate objects are sent to the visual language model, which poses questions such as "Which object in the image is most likely the 'M8 bolt used to fix the cover plate' described in the SOP?" The common sense reasoning ability of the visual language model is used to obtain a comprehensive matching confidence score, which serves as an important feature for similarity calculation.
[0053] Next, in constructing the association matrix When M is the number of SOP symbols and N is the number of physical entities, an associated historical context is introduced, and the matrix elements are defined as follows:
[0054]
[0055] in For feature similarity, Scoring is based on historical context. To balance the weights. For example, when the "first bolt" in the SOP is successfully associated, the system records this information and, when associating the "second bolt", prioritizes other bolt instances that conform to the "cross" order in space and have not been associated before, thus incorporating the order constraints in the SOP into the association process.
[0056] Finally, the optimal allocation policy learning unit 33 models association matching as a sequence decision problem and uses reinforcement learning to learn the optimal allocation policy. The agent's policy network outputs action probabilities:
[0057]
[0058] in For state, The action of "symbol m associating with entity n".
[0059] The reward is set based on the final success rate of the entire task. (T is the number of task steps,) (This is an indicator function). Through RL training, the model can learn to make globally optimal association decisions in complex scenarios, rather than just locally optimal ones.
[0060] Example 4: Definition of Parameterized Atomic Skills
[0061] The parameterized atomic skill library module includes a skill-model mapping and parameter interface design unit, a dynamic motion primitive enhancement unit, and a skill parameter instantiation unit.
[0062] First, the skill-model mapping and parameter interface design unit defines various skill models (such as DMP, LQR, etc.) in the skill library and designs a set of learnable parameter interfaces for each skill. For example, for the MoveTo skill, in addition to the target pose, the gain parameters K and D of its DMP can also be used as interfaces, allowing subsequent optimization modules to adjust them to change the agility or compliance of the movement.
[0063] Next, the dynamic motion primitive enhancement unit introduces a simplified robot dynamics model as a constraint based on the standard DMP. The core dynamic equation of the DMP is:
[0064]
[0065] Where y represents the trajectory position, K and D are the stiffness / damping coefficients, and g is the target position. (Residual acceleration learned by RL). Forward dynamics calculations ensure that the trajectory generated by DMP is dynamically feasible. Furthermore, using the MoRe-ERL (Motion Residual Reinforcement Learning) approach, reinforcement learning is used to learn residual adjustment amounts to fine-tune the initial DMP trajectory, compensating for inaccuracies in the dynamics model or responding to unknown environmental disturbances.
[0066] Finally, the skill parameter instantiation unit incorporates an online imitation learning mechanism. The system demonstrates the optimal execution of the skill in simulation, generating expert teaching data. Then, through behavior cloning or inverse reinforcement learning, the difference between the model and the expert's actions is minimized, with the loss function being:
[0067]
[0068] in For skill parameters, This is done by adapting the robot's actions to those of an expert. The learnable parameters of the skill are then adjusted to make the robot's execution style and effects closely resemble expert demonstrations, thus enabling online optimization of the skill.
[0069] Example 5: Hierarchical Task Planning
[0070] The hierarchical task planning module includes a high-level task sequence generation unit, a subtask decomposition unit, and an atomic skill invocation and interface verification unit.
[0071] First, the large language model in the high-level task sequence generation unit can actively query an external process knowledge base (such as through RAG technology) when generating high-level task sequences. For example, when parsing the "welding" task, the large language model can automatically retrieve key parameters such as "preheating temperature" and "shielding gas flow rate" from the welding process manual and inject them as constraints into the task description to ensure the rigor of the planning.
[0072] Next, the large language model in the subtask decomposition unit generates an explicit chain of thought when decomposing subtasks. The system requires the large language model to self-criticize each step in the chain of thought, for example: "Must step 'place the sealing ring' follow 'clean the sealing groove'?" When a logical contradiction is found, a self-correction mechanism is triggered, and the decomposition is repeated to improve the logical rationality of the plan.
[0073] Finally, the atomic skill invocation and interface verification unit draws on the idea of combinatorial reinforcement learning, treating each atomic skill as a sub-policy. When invoking a skill sequence, it not only performs mapping but also verifies the interface compatibility between subtasks (e.g., whether the output state of the previous skill satisfies the input prerequisites of the next skill). A high-level coordinator ensures that the overall combined policy meets global task specifications.
[0074] Example 6: Feasibility Verification and Replanning
[0075] The feasibility verification and replanning module includes a feasibility assessment unit, a multi-dimensional feasibility judgment unit, and a semantically guided replanning unit.
[0076] First, the feasibility assessment unit places the sub-tasks and their corresponding initial skill parameters into the scene semantic map with a physics engine constructed in step two for forward simulation. Through simulation, it intuitively observes whether collisions occur, whether the target is reachable, and whether force control requirements are met during the execution process, thereby obtaining a more accurate and physically realistic feasibility assessment than the confidence level of a visual language model.
[0077] Next, the multi-dimensional feasibility decision unit introduces a semi-online reinforcement learning decision mechanism. The system not only judges whether a single subtask is feasible, but also evaluates the contribution of that subtask to completing the entire long-term task through a value function V(s). The feasibility decision integrates simulation results, visual language model confidence, and long-term value assessment to form a multi-dimensional decision.
[0078] Finally, when a subtask is deemed infeasible, the semantically guided replanning unit initiates a directional exploration strategy, such as slightly adjusting the robot's approach angle or the tool's pose in the simulation, to attempt to find a new feasible path. These successful exploration experiences are recorded and fed back to the large language model, serving as a crucial basis for its semantic replanning and enabling a leap from "learning through thinking" to "learning through action."
[0079] Example 7: Solving Complex Geometric Constraints
[0080] The complex geometric constraint solving module includes a constraint understanding and optimization objective formalization unit, an optimization algorithm guidance unit, and an optimization result verification unit.
[0081] First, the large language model in the constraint understanding and optimization objective formalization unit directly participates in the constraint formalization process. The system inputs the natural language constraint "cross-shaped sequence tightening" from the SOP and real-world environmental information (such as bolt point clouds) into the large language model, requiring the model to output a computable optimization objective description. The large language model interprets "cross-shaped sequence" as "minimizing the dot product of consecutive operation points on the spatial vector," and the corresponding optimization objective function is:
[0082]
[0083] in This is the bolt operation sequence. Let L_k be the spatial vector of the L_k-th bolt. This transforms a fuzzy technological requirement into a clear mathematical objective.
[0084] Next, before launching the optimization algorithm (such as a genetic algorithm), the large language model in the optimization algorithm guidance unit provides a high-quality initial population or search direction for the algorithm based on its understanding of the problem. For example, the large language model might generate an initial sequence based on the distribution of bolts using the following formula:
[0085]
[0086] in This is the minimum bolt spacing threshold. This greatly accelerates the convergence process of the optimization algorithm.
[0087] Finally, optimize the solution to obtain the optimal operation sequence. The results will be fed back to the LLM for semantic interpretation and feasibility post-hoc verification. The LLM needs to determine whether this mathematically optimal solution is reasonable in terms of process logic. For example, "Although the sequence [1,3,2,4] satisfies the minimum vector dot product, it may lead to uneven stress on the cover plate. Fine-tuning is recommended." This will form a closed loop of "LLM defining the problem, algorithm solving, and LLM verification".
[0088] Example 8: Spatiotemporal Trajectory Optimization
[0089] The spatiotemporal trajectory optimization module includes an online motion trajectory generation unit, a trajectory optimization strategy learning unit, and a spatiotemporal joint optimization unit.
[0090] First, the online trajectory generation unit employs a model predictive control (MPC) framework for online trajectory generation. In each control cycle, based on the current state and the dynamic environment (such as moving obstacles), the optimal control problem in the finite-time domain is resolved, with the optimization objective being:
[0091]
[0092] Where H represents the prediction time domain and u represents the control variable. Using Q and R as the reference trajectory, and Q and R as penalty matrices, the motion trajectory for the next time period is generated, enabling dynamic obstacle avoidance and real-time tracking.
[0093] Next, the trajectory optimization strategy learning unit uses parameters from traditional optimization methods such as artificial potential fields (e.g., repulsive field strength, gravitational field gain) as learnable strategies. A neural network is trained using deep reinforcement learning in a large number of random environments. The network's input is the current environment state, and its output is the optimal potential field parameters or a direct trajectory adjustment. This learned "optimization strategy" is more adaptable than the potential field method with fixed parameters.
[0094] Finally, in the time optimization, the spatiotemporal joint optimization unit no longer only optimizes the total time T, but performs spatiotemporal joint optimization, and the optimization problem satisfies the robot's differential dynamics constraints:
[0095]
[0096] in Let M(q) be the joint torque, and M(q) be the inertia matrix. , (Time / energy weights). The optimization objective considers multiple indicators such as shortest time, minimum energy consumption, and end effector stability to ensure that the generated trajectory is not only time-optimal but also kinetically feasible.
[0097] Example 9: Intelligent Code Generation and Verification
[0098] The intelligent code generation and verification module includes an intermediate representation generation unit, an API mapping and code generation unit, and a code-level behavior verification unit.
[0099] First, the intermediate representation generated by the intermediate representation generation unit not only includes action instructions and parameters, but also embeds semantic information from the original SOP and planning process as annotations. The priority of semantic associations is calculated through attention weights:
[0100]
[0101] in For action commands, This is a fragment of the original SOP. (For similarity functions). For example, commenting the source of the Screw command's parameters as "torque value 15 N·m from SOP 3" provides rich context for subsequent debugging, tracing, and code verification.
[0102] Next, the API mapping and code generation unit leverages the code understanding and generation capabilities of the large language model to assist in completing the API mapping. Using the official API documentation of the target robot platform as a knowledge base, the large language model automatically generates or verifies the corresponding API call code snippets based on the intermediate representation. The large language model can identify and avoid some common programming pitfalls, such as unit conversions and coordinate system transformations.
[0103] Finally, before being deployed to the physical robot, the code-level behavior verification unit runs the generated executable code in a simulation environment. By comparing the simulation results with the expected task objectives, code-level behavior verification is performed. If the simulation results are unsatisfactory, the system records the deviation and traces back to the code or planning stage for correction, forming a reliable closed loop of "generation-verification-correction".
[0104] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A basic model for embodied intelligent task planning for humanoid robots, characterized in that: It includes the following modules connected in sequence: SOP deep structured semantic parsing module, multimodal environment perception and modeling module, cross-modal semantic-entity association module, parameterized atomic skill library module, hierarchical task planning module, feasibility verification and replanning module, complex geometric constraint solving module, spatiotemporal trajectory optimization module, and intelligent code generation and verification module; The SOP deep structured semantic parsing module is used to transform unstructured text SOPs into directed semantic graphs rich in semantic logic that can be directly computed by computers. The multimodal environment perception and modeling module is used to construct dynamic digital environment models that can be used for physical simulation and collision detection. The cross-modal semantic-entity association module is used to accurately match semantic symbols in the SOP with entities in the physical environment; The parameterized atomic skill library module is used to define intelligent skills that can adapt to environmental changes online and meet complex constraints. The hierarchical task planning module is used to generate high-level task sequences and sub-task decompositions with self-reflection and correction capabilities. The feasibility verification and replanning module is used to evaluate the feasibility of the planning scheme in a physical simulation environment and trigger replanning when it is not feasible. The complex geometric constraint solving module is used to transform complex geometric constraints described in natural language into mathematical optimization problems and solve them. The spatiotemporal trajectory optimization module is used to generate a dynamically feasible robot motion trajectory that adapts to the dynamic environment. The intelligent code generation and verification module is used to generate platform-specific code rich in semantic information and perform simulation verification.
2. The basic model for embodied intelligent task planning for humanoid robots according to claim 1, characterized in that, The SOP deep structured semantic parsing module includes: The predicate-argument structure extraction and relational reasoning unit uses a large language model finely tuned with domain process knowledge to identify basic arguments such as core actions, tools, and objects, and infers the implicit temporal, causal, and spatial relationships between arguments; The process parameter and constraint extraction unit adopts a hybrid extraction strategy. For explicit numerical parameters, it uses predefined rules and NER models for rapid extraction. For fuzzy natural language constraints, it calls a large language model for semantic disambiguation and formal transformation. The directed semantic graph construction and embedding unit constructs a directed semantic graph from the extracted entities and relations, and uses a semantic graph neural network to learn the embedding of nodes and edges, generating a vectorized semantic knowledge base that can be used for subsequent logical reasoning.
3. The basic model for embodied intelligent task planning for humanoid robots according to claim 1, characterized in that, The multimodal environment perception and modeling module includes: The instance-level scene deconstruction and state classification unit uses an instance segmentation network to identify objects and simultaneously determines the dynamic state of the objects through classification branches. A six-DOF pose estimation unit, integrating multi-view geometry and neural radiation field technology, solves the precise six-DOF pose of an object; The dynamic scene semantic map construction unit imports the identified objects, their poses, and attributes into a lightweight physics engine to build a scene semantic map with physical simulation capabilities.
4. The basic model for embodied intelligent task planning for humanoid robots according to claim 1, characterized in that, The cross-modal semantic-entity association module includes: A multi-feature similarity calculation unit is used to construct a similarity evaluation framework assisted by a visual language model, and to obtain a comprehensive matching confidence score by utilizing the common sense reasoning ability of the visual language model. The association matrix construction unit introduces the association history context when constructing the association matrix, and integrates the order constraints in the SOP into the association process; The optimal allocation strategy learning unit models association matching as a sequence decision problem and uses reinforcement learning to learn the optimal allocation strategy.
5. The basic model for embodied intelligent task planning for humanoid robots according to claim 1, characterized in that, The parameterized atomic skill library module includes: The skill-model mapping and parameter interface design unit designs a learnable parameter interface for each skill; The dynamic motion primitive enhancement unit introduces a simplified robot dynamics model as a constraint on the basis of the standard DMP, and uses reinforcement learning to learn the residual adjustment amount; The skill parameter instantiation unit introduces an online imitation learning mechanism, which adjusts the learnable parameters of the skill through behavior cloning or inverse reinforcement learning.
6. The basic model for embodied intelligent task planning for humanoid robots according to claim 1, characterized in that, The hierarchical task planning module includes: The high-level task sequence generation unit, when generating high-level task sequences, actively queries the external process knowledge base through RAG technology. Subtask decomposition unit: When decomposing subtasks, the large language model generates explicit thought chains and performs self-criticism and correction on each step of the thought chain. The Atomic Skill Invocation and Interface Verification Unit, drawing on the idea of combinatorial reinforcement learning, verifies the interface compatibility between subtasks.
7. The basic model for embodied intelligent task planning for humanoid robots according to claim 1, characterized in that, The feasibility verification and replanning module includes: The feasibility assessment unit places the sub-tasks and their corresponding initial skill parameters in a scene semantic map with a physics engine for forward simulation; A multi-dimensional feasibility decision unit is introduced, which adopts a semi-online reinforcement learning decision mechanism to make decisions based on simulation results, visual language model confidence, and long-term value assessment. The semantically guided replanning unit initiates a targeted exploration strategy when a subtask is deemed infeasible, and feeds back successful exploration experiences to the large language model.
8. The basic model for embodied intelligent task planning for humanoid robots according to claim 1, characterized in that, The complex geometric constraint solving module includes: The constraint understanding and optimization objective formalization unit, the large language model, transforms the natural language constraints in the SOP into a computable description of the optimization objective; The optimization algorithm guidance unit, based on the understanding of the problem, provides a high-quality initial population or search direction for the optimization algorithm; The optimization result verification unit performs semantic interpretation and feasibility posterior analysis on the optimal operation sequence obtained from the optimization solution using the large language model.
9. The basic model for embodied intelligent task planning for humanoid robots according to claim 1, characterized in that, The spatiotemporal trajectory optimization module includes: The online trajectory generation unit adopts a model predictive control framework to resolve the optimal control problem in the finite time domain in each control cycle; The trajectory optimization strategy learning unit uses deep reinforcement learning to train a neural network and outputs the optimal potential field parameters or the direct trajectory adjustment amount. The spatiotemporal joint optimization unit performs spatiotemporal joint optimization considering the constraints of robot differential dynamics.
10. The basic model for embodied intelligent task planning for humanoid robots according to claim 1, characterized in that, The intelligent code generation and verification module includes: The intermediate representation generation unit generates an intermediate representation that includes action instructions, parameters, and semantic information from the original SOP and planning process; The API mapping and code generation unit utilizes the code understanding and generation capabilities of a large language model to automatically generate or verify corresponding API call code snippets based on intermediate representations. The code-level behavior verification unit runs the generated executable code in a simulation environment and compares the simulation execution results with the expected task objectives to verify them.