Instruction understanding and task execution method and device, equipment and medium

By generating structured task descriptions through collaborative processing of voice and text commands, and dynamically adjusting them in conjunction with real-time environment models and sensor information, this technology solves the problems of inaccuracy and inflexibility in task execution of intelligent agents in complex environments in existing technologies, and achieves efficient and stable task execution and model optimization.

CN121009987APending Publication Date: 2025-11-25PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 9 Cited by

Patent Information

Application Number
CN202511112513.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing technologies have limitations in multimodal instruction parsing, dynamic environment modeling, task strategy generation and execution adjustment, and are difficult to adapt to the high flexibility and high accuracy requirements of intelligent agents in complex business environments. In particular, they are prone to comprehension and execution errors when faced with semantic ambiguity, complex spaces and dynamic changes.

Method used

By acquiring voice and text commands and performing collaborative processing to generate structured task descriptions, and combining them with real-time environment models to generate task execution strategies, the system adjusts actions based on real-time sensor information during execution, while simultaneously collecting task execution data and user feedback to update the model.

Benefits of technology

It achieves deep fusion of multimodal information, improves the ability to perceive dynamic environments and the accuracy and flexibility of task execution, and enhances the intelligent agent's understanding, autonomous decision-making and adaptability in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009987A_ABST
    Figure CN121009987A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to service scenes of pension service, financial science and technology, medical health and the like, and discloses an instruction understanding and task execution method, device, equipment and medium, and the method comprises the steps: receiving a voice instruction and a text instruction, and carrying out the cooperative processing through an instruction understanding model, and generating a structured task description; collecting environment data to construct a real-time environment model; generating a task execution strategy by utilizing a task execution model based on the structured task description and the real-time environment model; controlling the intelligent agent to execute the task according to the task execution strategy, and dynamically adjusting the action in combination with real-time sensor information; task execution data and user feedback information are collected, and the instruction understanding model and the task execution model are updated. According to the method, multi-modal information is fused through structural description, an execution strategy is generated in combination with real-time environment perception, actions are dynamically adjusted, model self-optimization is further achieved through execution data and feedback, and the understanding, decision-making and adaptive capacity of an intelligent agent is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and storage medium for instruction understanding and task execution. Background Technology

[0002] In the technologies related to intelligent agent task understanding and autonomous execution, existing solutions have many limitations in multimodal instruction parsing, dynamic environment modeling, task strategy generation and execution adjustment, and model iterative updates, making it difficult to meet the high flexibility and high accuracy task execution requirements of intelligent agents in complex business environments.

[0003] In the fintech sector, intelligent assistants or service robots are widely used to assist with tasks such as customer service, document delivery, and office guidance. However, existing systems generally rely on a single text input method, making it difficult to accurately understand customer intent, especially when dealing with semantic ambiguity or vague expressions in user language. This is particularly true for instructions involving synonyms, word order changes, or ambiguous semantics, which easily lead to task parsing errors. Furthermore, in financial office settings, the complex and frequently changing spatial layout, such as office reorganization or the appearance of temporary obstacles, can interfere with the robot's task path. Existing robots lack sufficient accuracy in environmental modeling and the ability to model complex spatial structures and dynamic changes, resulting in unstable execution paths and high task failure rates.

[0004] In the healthcare sector, intelligent robots or nursing assistive devices are commonly used for tasks such as directional rounds, material delivery, and patient interaction. When elderly people, patients, or medical staff give instructions, they may exhibit unclear expression, uneven speech rate, or accents, affecting comprehension. Current technologies mostly employ shallow fusion methods, failing to fully utilize the complementary information from speech and text, leading to semantic comprehension biases. Simultaneously, the medical environment is complex and constantly changing, with frequent occurrences such as patient movement, bed adjustments, and equipment obstruction. Existing robots lack a complete perception-mapping-execution closed-loop mechanism, making it difficult to adjust actions promptly to adapt to real-time environmental changes. Furthermore, when robots encounter path deviations or grasping anomalies during task execution, existing systems often fail to make effective adjustments based on real-time feedback and lack the ability to continuously optimize models using execution data and user feedback. This results in robots remaining in a suboptimal state for extended periods, lacking adaptive learning capabilities.

[0005] In the field of elderly care services, health and wellness robots, as assistive care tools, are gradually taking on tasks such as retrieving and placing items, reminding patients of medications, and providing companionship and interaction. However, in practical applications, the elderly may use colloquial, emotional, or ambiguous expressions for voice or text interaction. Existing systems cannot accurately interpret the deeper meanings of these non-standardized expressions, leading to frequent errors in instruction comprehension or execution. Simultaneously, the living environments of the elderly often present complex factors such as irregular furniture placement, cramped spaces, and insufficient lighting. Current robots' limitations in environmental perception and 3D mapping capabilities prevent them from accurately modeling, thus affecting the stability of path planning and operational execution. Furthermore, existing robots lack dynamic adjustment mechanisms during task execution, failing to make fine-grained corrections to motion parameters based on real-time sensor information. This easily leads to errors in grasping and movement, reducing service quality. More importantly, traditional systems generally neglect the value of user satisfaction feedback in model optimization, failing to achieve continuous adaptation and self-evolution based on the individual needs of the elderly. Summary of the Invention

[0006] The main objective of this invention is to provide a method, apparatus, device, and storage medium for instruction understanding and task execution, aiming to solve the technical problems in the prior art where the collaborative processing of voice and text instructions is shallow, lacks a structured fusion mechanism, and lacks the ability to dynamically adjust actions and continuously optimize models based on real-time sensor feedback during task execution.

[0007] To achieve the above objectives, the present invention provides a method for instruction understanding and task execution, comprising:

[0008] Acquire voice commands and text commands, and use a command understanding model to collaboratively process the voice commands and text commands to generate a structured task description;

[0009] Collect environmental data and construct a real-time environmental model based on the environmental data;

[0010] Based on the structured task description and the real-time environment model, a task execution strategy is generated through the task execution model;

[0011] The agent is controlled to perform tasks according to the task execution strategy, and the agent's actions are adjusted according to real-time sensor information during task execution.

[0012] Collect task execution data and user feedback information, and update the instruction understanding model and task execution model based on the task execution data and user feedback information.

[0013] Furthermore, to achieve the above objectives, the present invention provides an instruction understanding and task execution apparatus, comprising:

[0014] The instruction understanding module is used to acquire voice instructions and text instructions, and to perform collaborative processing of the voice instructions and text instructions through the instruction understanding model to generate a structured task description;

[0015] The environmental modeling module is used to collect environmental data and construct a real-time environmental model based on the environmental data.

[0016] The task strategy generation module is used to generate a task execution strategy based on the structured task description and the real-time environment model through the task execution model.

[0017] The task execution control module is used to control the intelligent agent to execute tasks according to the task execution strategy, and to adjust the actions of the intelligent agent according to real-time sensor information during the task execution process;

[0018] The model adaptive update module is used to collect task execution data and user feedback information, and update the instruction understanding model and task execution model based on the task execution data and user feedback information.

[0019] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and an instruction understanding and task execution program stored in the memory and executable on the processor, wherein when the instruction understanding and task execution program is executed by the processor, it implements the steps of the instruction understanding and task execution method as described above.

[0020] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing an instruction understanding and task execution program, wherein the instruction understanding and task execution program, when executed by a processor, implements the steps of the instruction understanding and task execution method as described above.

[0021] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as elderly care services, fintech, and healthcare. It discloses a method, apparatus, device, and medium for instruction understanding and task execution, including: receiving voice and text instructions; generating a structured task description through collaborative processing using an instruction understanding model; collecting environmental data to construct a real-time environment model; generating a task execution strategy based on the structured task description and the real-time environment model using a task execution model; controlling the agent to execute tasks according to the task execution strategy and dynamically adjusting actions based on real-time sensor information; collecting task execution data and user feedback information to update the instruction understanding model and the task execution model. This invention achieves deep fusion of multimodal information by constructing a structured task description, enhances the perception of dynamic environments through a real-time environment model, improves the accuracy and flexibility of task execution through task execution strategy-driven and dynamic action adjustment, and further optimizes the model by combining task execution data and user feedback, thereby enhancing the agent's understanding, autonomous decision-making, and adaptability in complex scenarios. Attached Figure Description

[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:

[0023] Figure 1 This is a schematic diagram of an application environment for the instruction understanding and task execution method in one embodiment of the present invention;

[0024] Figure 2 This is a flowchart illustrating an embodiment of the instruction understanding and task execution method of the present invention;

[0025] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the instruction understanding and task execution device of the present invention;

[0026] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0027] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0028] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0029] The instruction understanding and task execution method provided in this embodiment of the invention can be applied to, for example... Figure 1In this application environment, the user terminal communicates with the server via a network. The server can receive voice and text commands from the user terminal, and use a command understanding model to collaboratively process and generate a structured task description; collect environmental data to build a real-time environment model; generate a task execution strategy based on the structured task description and the real-time environment model using a task execution model; control the agent to execute tasks according to the task execution strategy, and dynamically adjust actions based on real-time sensor information; collect task execution data and user feedback information to update the command understanding model and the task execution model. This invention achieves deep fusion of multimodal information by constructing a structured task description, enhances the perception of dynamic environments through a real-time environment model, improves the accuracy and flexibility of task execution through task execution strategy-driven and dynamic action adjustment, and further optimizes the model by combining task execution data and user feedback, thereby enhancing the agent's understanding, autonomous decision-making, and adaptability in complex scenarios. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.

[0030] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the instruction understanding and task execution method provided by the present invention. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0031] like Figure 2 As shown, the instruction understanding and task execution method proposed in this invention includes the following steps:

[0032] S10, acquire voice commands and text commands, and use the command understanding model to collaboratively process the voice commands and text commands to generate a structured task description;

[0033] In this embodiment, the process of acquiring voice and text commands involves receiving and formatting natural language content input by the user via both voice and text. During voice command acquisition, external language signals are typically acquired using a microphone array or a single microphone. The acquired audio signals undergo noise reduction filtering to remove background noise and environmental interference signals, and then are converted from audio to text using a speech recognition model. This model can be built based on a Long Short-Term Memory network, an end-to-end Transformer structure, or a connection-time classification model, outputting a stable voice-text sequence. Text commands, on the other hand, receive character information input by the user through a graphical user interface, touchscreen, or remote input device, and undergo language structure preprocessing. Processing methods include word segmentation, part-of-speech tagging, and lexical disambiguation, with the goal of generating a standardized text sequence with consistent structure and clear semantics.

[0034] The role of the instruction understanding model is to achieve semantic complementarity and collaborative processing between speech and text instructions. Its implementation mechanism includes two key parts: multimodal encoding of the input information and semantic aggregation of the two types of encoded information. Speech and text sequences are encoded using phonemes and embedded semantically to generate acoustic and semantic vectors. Text sequences, on the other hand, are processed using context-based language models (such as BERT) to extract context-related semantic representations. Both types of vectors undergo alignment mapping and unified embedding space transformation before being input to the fusion module. The fusion module can employ gated attention mechanisms, self-attention networks, cross-modal contrastive learning, and other structures to achieve vector-level weight enhancement and information integration, thereby generating intermediate feature representations that express the task intent.

[0035] The generation process of structured task descriptions is based on fused semantic representation vectors, combined with predefined task templates or contextual reasoning mechanisms. Task templates include elements such as target action type, object category, spatial location information, and expected results. Each element must obtain a corresponding matching position and strongly related content in the fused vector. To achieve the mapping from vectors to structured descriptions, a multi-label classifier or sequence generation module can be constructed. Structured data representations can be generated through template filling, slot matching, or graph structure construction. This representation not only includes basic information about actions and objects but can also cover information dimensions such as condition triggers, constraints, and execution priorities, forming a machine-parseable task expression structure that serves as input conditions for subsequent task execution models.

[0036] For the voice input section, a Transformer-based speech recognition model can be used, coupled with a multi-microphone array acquisition device with direction awareness to handle far-field voice input. To address accent interference, a speaker-adaptive training mechanism can be introduced, allowing for minor fine-tuning of the recognition model using a small number of user voice samples to improve the accuracy of speech-to-text output. For the text input section, a natural language processing engine with entity recognition capabilities can be used to perform structured analysis of user commands, extracting key fields such as task objectives, objects, and modifiers.

[0037] The fusion processing module can employ a dual-stream Transformer architecture, encoding the speech-text sequence and the standard text sequence separately. A gating mechanism is then used to achieve weighted fusion of multimodal features. This gating mechanism controls the attention intensity of different modalities by introducing contextual memory states, thus forming more task-oriented fusion features. The generation of structured task descriptions can be achieved by constructing a task parser based on a graph attention network, mapping the fusion features to nodes and edges in the task graph, representing the relationships between task objectives, actions, and constraints. The task graph can be further encoded into structured JSON or a tree-like description for use by subsequent modules.

[0038] The system can also use multi-round dialogue guidance to supplement task details. When there is missing information in the structured description, the system can guide the user to complete the conditions through automatically generated questions. For example, in the medical and health scenario, when the system identifies "helping patients take medication" but lacks information on the type of medication, it can guide the user to provide the specific medication. In the financial scenario, when the system identifies "generating customer reports" but lacks customer identification information, it can guide the user to confirm further.

[0039] Example: In the healthcare business, an elderly person can say "I feel a little dizzy, please get me my medicine" via voice, and at the same time click the "Medication Reminder" function in the application interface. The system will extract the physical state and intention mentioned in the voice, and integrate it with the text operation to generate a structured description "Task Type: Medication Retrieval; Target: Emergency Medication for Fainting; Execution Location: Medicine Cabinet Area", providing input for subsequent path planning and action generation.

[0040] In the fintech business, users can say "Please check my customers' credit data from last week" via voice, while simultaneously entering the customer ID into the system. The system combines the task objective "query credit data" from the voice with the customer identifier from the text command to construct a structured description: "Operation Task: Data Query; Query Scope: Last Week; Customer ID: 123456; Data Type: Credit Transaction." This allows the system to invoke a backend model to retrieve and visualize the data. This approach effectively integrates voice and text information, even if there are discrepancies in their descriptions, for task execution.

[0041] In the field of elderly care services, elderly users might express their needs via voice, such as "Help me watch TV programs," while simultaneously selecting the "entertainment" category through a remote control. The system recognizes the user's ambiguous request, "watch TV programs," and combines this with the textual context of "entertainment content." After semantic fusion, a structured task description is generated, including "Task type: information browsing; Content type: TV program recommendation; Display method: list; Priority: normal." Supported by this structured description, the system can accurately determine the task execution logic, automatically retrieve available channels for the current time slot, and display them in an interface suitable for elderly users' reading habits. This avoids misunderstandings caused by semantic ambiguity and improves the responsiveness and user-friendliness of smart devices. By integrating voice and text, the system effectively alleviates the problems of unclear expression or operational difficulties for the elderly, enhancing the service adaptability and humanization of elderly care equipment.

[0042] This embodiment unifies the encoding of speech and text commands and integrates semantic features to generate structured task descriptions. This allows information from different input formats to effectively complement each other, improving the completeness of task generation while ensuring accuracy in understanding, and avoiding execution deviations caused by misunderstandings in traditional speech recognition. The generated structured task descriptions have a standardized format and semantic hierarchy, providing a clear input basis for subsequent task reasoning and action planning, thereby enhancing the automation and robustness of the entire task chain.

[0043] S20, Collect environmental data and construct a real-time environmental model based on the environmental data;

[0044] In this embodiment, the process of acquiring environmental data involves the collaborative work of multiple sensors. Environmental data refers to a set of information reflecting the characteristics of the surrounding physical space, the state of the spatial structure, the existence of objects, and their spatial poses. Its sources may include, but are not limited to, vision-based, distance-based, and spatial structure-based devices. Vision-based sensors, such as RGB cameras, can provide color information, edge contours, and texture variations to capture the appearance of objects, lighting conditions, and spatial contours in a scene. Depth cameras can acquire distance information between each pixel and the imaging device, forming a depth image that reflects the distribution of objects in three-dimensional space. LiDAR provides high-precision point cloud data, constructing a dense three-dimensional environmental representation through high-speed scanning and returning distance measurements. Ultrasonic sensors provide directional distance information for nearby obstacles, suitable for supplementing near-field detection blind spots.

[0045] After being collected uniformly, this environmental data will be processed by a multimodal data fusion module. During the fusion process, edge detection and texture analysis are first performed on the color images to extract the main structural contours. Common methods include using the Sobel operator to extract edge features and the gray-level co-occurrence matrix to extract texture distribution. For depth images, coordinate transformation methods (such as point cloud reconstruction based on camera intrinsics) are used to map two-dimensional pixels to three-dimensional spatial coordinates, thereby restoring the spatial position and distance relationships of objects. Point cloud data can then be divided into independent object regions in the scene using clustering methods (such as Euclidean distance clustering or the DBSCAN algorithm) to identify interactive targets or obstacle structures. Obstacle distance information can be weighted during the fusion process and superimposed on low-density regions in the point cloud for judgment, thereby improving near-field accuracy.

[0046] Multi-source fusion data is used to form a complete scene description through temporal alignment and spatial registration, upon which a 3D spatial map is constructed. The 3D spatial map can be dynamically maintained using voxel meshes, octrees, or sparse graph representations to support efficient updates to spatial structures and region indexing. The construction of the real-time environment model includes not only the maintenance of spatial geometric information but also the overlay of dynamically changing states related to the time axis, such as recording dynamic object trajectories and identifying static structural changes, achieving spatiotemporal joint modeling of the environmental state. Ultimately, the real-time environment model will serve as the foundational information source for subsequent task execution strategy generation and action planning.

[0047] In some scenarios, RGB-D cameras can be used to simultaneously acquire color and depth images. A unified driver program can then synchronize and register the two data streams, reducing the impact of parallax errors. Alternatively, LiDAR can be combined with a depth camera. Calibration can be used to obtain their relative position and orientation, and an extrinsic parameter matrix can be used to project the LiDAR point cloud for fusion with the image domain. Furthermore, multiple sets of low-cost ultrasonic sensors can be configured to cover blind spots, making it particularly suitable for close-range operations or enclosed spaces.

[0048] In terms of data processing strategies, a frame fusion algorithm based on a sliding time window can be selected, combined with Kalman filtering or Bayesian fusion strategies to perform temporal modeling of multi-source data, thereby improving the model's adaptability to dynamically changing scenarios. For the update mechanism of the 3D map, the local update frequency can be set according to the degree of environmental change. For example, when a high-density object change area is detected, the accuracy of that area can be improved and the map refreshed first, while a low-frequency refresh can be maintained in static areas to save computing resources.

[0049] Example: In a healthcare setting, rehabilitation assistive robots need to identify the bed layout, obstacle locations, and patient movement in real time when delivering medication. By acquiring depth images to identify the bed contours, dividing the ward layout into point cloud areas, and using ultrasound to sense the patient's approach, the robot adjusts its path to ensure a safe, efficient, and dynamically adaptable medication delivery route.

[0050] In financial business scenarios, when intelligent guidance equipment assists customers in navigating to counters within branches, it can identify personnel density and queue structure through visual images and sense whether the channels are unobstructed through laser point clouds, thus constructing a dynamic environmental model of the current business hall. This allows the equipment to reasonably guide customers to avoid congested areas during peak hours, improving service efficiency and customer experience.

[0051] In the field of elderly care services, home assistant robots use visual cameras to identify the outlines of furniture and passageways in the home environment, use point cloud technology to locate the positions of major objects such as sofas and coffee tables, and use obstacle distance data to determine in real time whether the passage is blocked by clutter, thereby ensuring safety and accuracy when delivering items to the elderly or accompanying them on walks, and enhancing the reliable execution capability of elderly care service equipment in dynamic home scenarios.

[0052] This embodiment fuses data from multiple types of sensors to establish a dynamically updated 3D spatial map. Combined with environmental perception and structural reconstruction technologies, it constructs a real-time environmental model reflecting the true spatial state. The system can accurately perceive the physical environment in which the agent is currently located. The model's spatiotemporal update capability enables the agent to adapt to complex and changing environmental conditions, providing accurate and high-dimensional scene references for subsequent task planning and execution. The fusion of multi-source sensor data not only improves the model's spatial resolution of the environment but also enhances its robustness to dynamic changes and occlusion interference, effectively solving the problems of incomplete environmental perception and lagging model updates in traditional systems.

[0053] S30, Based on the structured task description and the real-time environment model, generate a task execution strategy through the task execution model;

[0054] In this embodiment, the structured task description is a set of semantic expressions generated by the preceding processing steps, containing target actions, operation objects, target attributes, and temporal constraints. After unified encoding, it can be used as the input representation of task instructions. The structured task description typically uses a nested structure to represent the logical relationship between multiple subtasks, including action keywords (such as "grab", "deliver", "move"), category labels of target objects (such as "water cup", "medicine box", "receipt"), and attribute labels related to operation conditions (such as "near the desktop", "located at specified coordinates", "complete within 30 seconds", etc.).

[0055] A real-time environment model is a scene representation structure that expresses the current physical space state, encompassing static spatial structures (such as walls, furniture, and obstacles) and dynamic element states (such as moving targets, temporarily obstructed areas, and current path reachability). In practical applications, this model is typically represented in the form of a graph structure or mesh voxels, and includes semantic annotations (such as "sofa," "elevator entrance," and "waiting area") and topological relationships (such as "accessible paths" and "action collision risk zones").

[0056] The task execution model, acting as an information fusion and decision engine, receives two types of input: a structured task description and a real-time environment model. It then generates a task execution strategy through a joint modeling approach. The strategy generation process begins with instruction-environment matching, where the target location in the structured description is matched and calibrated with the object location information in the environment model to resolve the alignment between the target instruction and the real target in the current environment. Subsequently, the model derives the task execution chain based on the operational objective and the environmental state, constructing an action sequence graph. Each node in this graph corresponds to a physical action unit (e.g., "approaching the target," "grabbing action," "obstacle avoidance"), and edges represent the dependency order and conditional triggering relationships between actions.

[0057] Execution models are typically implemented based on a hybrid decision architecture, combining symbolic logic reasoning with machine learning prediction. For example, a behavior tree structure is used to decompose and manage tasks, while a reinforcement learning policy network is used to select paths and dynamically adjust low-level action nodes. In some implementations, attention mechanisms can be introduced to prioritize multi-objective, multi-constraint tasks, thereby achieving optimal path generation and task flow planning for complex tasks.

[0058] The generated task execution strategy is ultimately output in the form of a structured action graph, sequence control code, or real-time instruction stream, and can be interfaced with the execution control module for real-time task scheduling. The strategy includes operation action type, target pose information, time window, path instructions, and action conditions, and has a clear execution sequence, target constraints, and dynamic adjustability.

[0059] In some implementations, structured task descriptions can generate task graph structures through natural language processing modules, embedding these graphs using graph neural networks, and engaging in multimodal interactive reasoning with entity nodes in the environment graph. Alternatively, tasks can be represented as solvable logical task structures using action planning languages ​​(such as PDDL), and then graph search algorithms such as A* or Dijkstra's algorithm can be used to search for execution paths that meet the conditions on the environment model.

[0060] The task execution model can employ a modular neural network architecture, where high-level modules are responsible for parsing the logical relationship between the task graph and the environment graph, while low-level modules generate motion control parameters, such as grasping posture estimation and path navigation parameter optimization. Alternatively, a policy model based on imitation learning can be used, learning the execution order and adjustment logic of complex tasks from human demonstration trajectories, and combining this with an online evaluation module to provide real-time scoring and adjustment of the generated strategy.

[0061] For scenarios involving conflicting or real-time changes in task priorities, a task replanning module can be integrated. When the target state changes or execution is interrupted, the current task graph structure is automatically updated and the strategy path is regenerated to ensure the continuity and robustness of the task execution process.

[0062] Example Description: In the healthcare field, after receiving the instruction "hand the water cup to the patient," the ward assistive robot parses the three action objectives: "grab the water cup – navigate to the patient's bedside – release the item." The relationship between the action nodes and the objectives is represented by a structured task description. By combining the environmental model to analyze the current orientation of the patient's bed and the position of the water cup, an execution path is generated that avoids the passageway of medical staff and bedside equipment. A posture adjustment strategy for grasping and placing is also formed to ensure the successful completion of the task.

[0063] In the financial sector, branch service robots analyze the task of "guiding customers to the VIP window" and convert it into a structured description that includes sub-tasks such as identifying the customer's current location, planning the route, and indicating the direction. Combining the current traffic status, queuing status, and channel congestion data in the branch, dynamic guidance paths and instruction action sequences are generated to achieve intelligent guidance services and avoid high-density areas.

[0064] In elderly care service scenarios, home assistant robots generate sequential actions such as "approaching the elderly – assisting movement – ​​locating the medicine – guiding the elderly to retrieve the medicine" based on the instruction "help the elderly to the kitchen to get medicine". They also combine home environment models to analyze furniture distribution, lighting conditions and passage width, and generate auxiliary paths and stable support actions through adaptable strategies to improve the safety and autonomy of the elderly in their daily lives.

[0065] This embodiment inputs a structured task description and a real-time environment model into the task execution model, and constructs a task execution strategy with logical timing, target alignment, and environmental adaptability. The system achieves a closed-loop inference link from semantic understanding to physical action generation. This strategy generation method considers both high-level task semantic understanding and low-level physical control elements, enabling the agent to maintain stable and efficient task execution capabilities even in complex environments facing dynamic changes and uncertainties. Especially in scenarios with multiple objectives and constraints, it can perform task decomposition, priority ranking, and path planning, effectively improving task success rate and execution efficiency.

[0066] S40, control the intelligent agent to execute the task according to the task execution strategy, and adjust the actions of the intelligent agent according to real-time sensor information during the task execution process;

[0067] In this embodiment, the task execution strategy is typically represented as a structured sequence of action instructions, including task action type (such as grab, move, avoid), target parameters (such as pose, grab point, release position), execution order, expected state constraints, and time control elements. This strategy structure can describe the temporal progression logic and control requirements of the task in space, forming the input basis for generating subsequent task execution actions.

[0068] Controlling an intelligent agent to execute a task refers to mapping the abstract action goals expressed in the task execution strategy into executable low-level control commands, driving the agent's execution units to complete the actual physical actions. The execution process involves the coordinated operation of multiple subsystems, including the path navigation module, the robotic arm control module, the grasping module, and the execution status monitoring module. The high-level actions in the execution strategy need to be refined into continuous pose sequences by the motion planning module, and further converted into motor control signals by the controller to achieve specific behavioral outputs.

[0069] During execution, real-time sensor information, including visual data, depth information, inertial measurement unit (IMU) signals, force feedback data, and tactile sensor data, reflects the current environmental state and action execution feedback of the agent. After collecting this information, the action adjustment mechanism compares the current sensor state with the expected action state, calculates the error, and dynamically fine-tunes the control parameters during execution.

[0070] For example, when performing a grasping task, if the force sensor detects an abnormal contact force on the target (too large or too small), the grasping module will adjust the closing angle or the applied force of the gripper in real time to prevent the target from slipping or being damaged. During path navigation, if the vision system detects a dynamic obstacle intruding into the original path, the path module will recalculate the obstacle avoidance path based on the latest environmental model and interrupt the original control flow to switch to a new motion trajectory.

[0071] Real-time action adjustment is typically achieved based on feedback control and adaptive modeling mechanisms, including but not limited to proportional-integral-derivative (PID) controllers, model predictive control (MPC), reinforcement learning policy updates, and impedance control. Through these methods, the agent can not only accurately reproduce the predetermined behavior in the task strategy but also possess the ability to autonomously adjust to cope with environmental uncertainties and sudden state changes, ensuring the integrity and safety of the task.

[0072] In some implementations, the task execution strategy is parsed as a finite state machine structure, with each state corresponding to a control module's call action. Once a sensor confirms the success of a state's execution, the system automatically jumps to the next action state. Alternatively, a behavior tree model can be used to drive task execution, with a state monitor attached to each leaf node to detect whether an action was successful or needs to be retried, and to call the action adjustment module to correct parameters in the failure branch.

[0073] To achieve high-frequency control and low-latency response, some systems introduce edge computing modules to offload computational tasks such as image recognition, force estimation, and path determination to local processing units, shortening the time cycle from sensor sensing to control command issuance. Prediction modules can also be deployed to model the sensor state change trends over a future period and pre-generate multiple alternative trajectory options, improving the foresight of decision-making.

[0074] In terms of dynamic environment adaptation, the sensor perception and motion adjustment mechanism can be configured as a sliding time window mechanism to perform average filtering on continuously sampled data, extract stable state indicators, and avoid motion jitter caused by short-term interference. In the implementation of multimodal data fusion, Bayesian filtering or Kalman filtering methods can be used to integrate state information from multiple sensors to improve the stability and robustness of motion adjustment.

[0075] Example: In a healthcare scenario, a rehabilitation assistive robot performs the task of "delivering a water cup to a patient." The control system breaks down the path planning action into multiple point-to-point movement and grasping / releasing operations. During execution, visual sensors identify the patient's position and the distribution of obstacles within the action area in real time. If the system detects that the patient's hand movement exceeds the predetermined receiving range, it can adjust the ending posture of the delivery action based on sensor feedback to avoid collisions or misplacement.

[0076] In financial settings, branch service robots perform the task of "collecting printed vouchers and delivering them to the counter." When the control system grasps the voucher in the printing area, it uses a camera to locate the paper's position and a force sensor to monitor the contact state of the gripper. If the printed voucher is not fully delivered or becomes damp and adheres, causing the gripper to fail to make initial contact, the robot can determine that the grasp was unsuccessful based on the contact force feedback. It then initiates a fine-tuning path strategy, adjusting its posture to attempt a second grasp without compromising the voucher's integrity.

[0077] In elderly care service scenarios, home assistive robots help seniors walk according to task strategies, monitoring their posture stability and gait rhythm through lower limb pressure sensors and gyroscopes. If the robot detects that the senior is leaning to the side or pausing during the process, it can immediately adjust its assistive posture and traction force, slowing down the speed and providing support by replanning the gait guidance path, thus achieving more humane and safer assistive actions.

[0078] This embodiment introduces a dynamic action adjustment mechanism based on real-time sensor feedback, enabling the agent to no longer rely on preset static paths and parameters during task execution. Instead, it can fine-tune its actions and switch strategies according to environmental changes and its own state, significantly enhancing the system's adaptability to complex environments. This mechanism ensures high stability and fault tolerance in task execution, effectively reducing the risk of task failure due to unforeseen circumstances (such as obstacle intrusion, target deviation, or action error), and improving overall execution efficiency and the agent's reliability.

[0079] S50: Collect task execution data and user feedback information, and update the instruction understanding model and task execution model based on the task execution data and user feedback information.

[0080] In this embodiment, collecting task execution data refers to recording the agent's behavior and execution status throughout the entire process of performing actions. This type of data includes, but is not limited to, path execution time, operation trajectory deviation, grasping or operation success rate, force control response value, obstacle avoidance trigger count, and resource consumption statistics. This data is collected in real time by monitoring units embedded in the execution module (such as encoders, sensors, and feedback control loops), and is used to form a multi-dimensional task execution log aligned with timestamps for subsequent modeling and analysis.

[0081] User feedback typically originates from human-computer interaction channels, such as satisfaction assessments following voice commands, touchscreen clicks, app ratings, emotion recognition outputs, or physiological state assessments. Feedback reflects the user's subjective evaluation of task completion quality, and its collection process must maintain temporal relevance to the execution data to ensure a causal link between state, behavior, and evaluation during model training.

[0082] Updating the instruction understanding model refers to optimizing the semantic parsing parameters, alignment mechanism weights, modality fusion structure, or task intent mapping strategy within the understanding model based on task execution data and user feedback. This update is typically achieved through supervised learning, combining the error between the sampled structured instructions and the final task execution result to backpropagate and update the model's internal parameters, thereby improving the accuracy of extracting task semantics from speech or text input.

[0083] Updating the task execution model primarily involves retraining the path planning strategy, action generation strategy, and control parameter decoding network. Especially after accumulating data from multiple rounds of task execution, strategy replay, success sample augmentation, or failure case comparison learning can be performed to make the model's action strategies more aligned with the dynamic changes in the target environment and user behavioral preferences. Furthermore, based on the annotation information of abnormal execution behaviors, distorted judgments in the model's state-action mapping can be corrected, improving robustness.

[0084] The update process can be implemented using an online learning mechanism to quickly fine-tune the results after each task execution, or using batch training to periodically retrain and replace the old model after collecting a certain amount of data. Furthermore, threshold indicators can be set before model updates, such as triggering the update process when user satisfaction or task completion rate falls below a set standard.

[0085] In the actual implementation, task execution data acquisition can be automatically completed by the embedded data recording module, and uniformly encoded using a standard data structure, including operation action categories, control parameter settings, target status records, actual sensor feedback, and execution time statistics. The system can analyze the current strategy execution deviation in real time through offline log analysis or the embedded evaluation module to determine whether model optimization is needed.

[0086] User feedback can be collected using a multimodal interface fusion approach. For example, the satisfaction score can be determined by combining the emotion recognition scores of voice evaluation words such as "too slow" and "unsatisfactory" with facial motion units extracted from visual expression recognition. In some implementations, to improve feedback accuracy, users can be guided to perform a rating operation after completing the task, and a summary information can be generated by combining the semantic restatement of the agent's behavior for user confirmation, thereby improving feedback quality.

[0087] The instruction understanding model can be updated by fine-tuning the weight parameters of the pre-trained model. Fine-tuning can employ a differential parameter update strategy, performing local optimization only on key parameters to improve efficiency. In certain implementations, to prevent model degradation, an elastic weight fixation mechanism can be used to retain the original model's basic generalization ability.

[0088] The task execution model can be updated based on a behavior cloning mechanism, by collecting historical successful task samples to train a new policy network. Furthermore, in multi-task systems, a sub-task priority ranking model can be constructed to adjust the policy invocation order and action target weight allocation, automatically replanning the execution path when task conflicts occur or user goals change.

[0089] Example: In a healthcare scenario, a companion robot performs the task of "reminding patients to take their medication." It collects task execution data such as arrival time at the bedside, whether the voice broadcast is interrupted, and patient response delay. At the same time, it combines the patient's feedback via button to see if they have taken their medication and the level of satisfaction expressed in the voice to update the robot's semantic recognition module and reminder process model, making subsequent broadcasts more concise and timely, and reducing the cognitive burden on the patient.

[0090] In financial scenarios, when a branch service robot handles the task of "printing and submitting transaction vouchers," it records information such as the voucher printing waiting time and whether the user's receiving action was successful. It also combines this information with whether the user clicked the "service evaluation" button or expressed satisfaction on the interface to adjust the judgment logic of its instruction recognition model for business process keywords, thereby enhancing the accuracy of subsequent processing of similar business.

[0091] In elderly care service scenarios, when assistive robots perform tasks such as "guiding the elderly to exercise," they collect execution data, such as deviations in walking paths, frequency of abnormal posture recognition, and deviations in exercise time. Simultaneously, they collect feedback from the elderly through voice or information on fluctuations in vital signs. Based on this data, the system updates its path stability control logic and voice guidance rhythm adjustment strategy, enabling the robot to better adapt to the slower pace and greater emotional fluctuations of the elderly, thereby improving the care experience and safety assurance capabilities.

[0092] This embodiment establishes a correlation mechanism between task execution data and user feedback information, and uses this correlation feedback to continuously optimize the model's understanding and execution capabilities, enabling the system to continuously learn, autonomously adjust, and dynamically adapt. Compared to traditional static control logic, it can respond to user preferences and adapt to environmental changes more quickly, improve the matching degree of execution strategies and the accuracy of semantic understanding, and enhance task completion efficiency and user satisfaction. Especially in scenarios with high user behavior complexity or changing needs, this mechanism significantly enhances the system's scalability and human-machine collaborative intelligence level.

[0093] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as elderly care services, fintech, and healthcare. It discloses a method, apparatus, device, and medium for instruction understanding and task execution, including: receiving voice and text instructions; generating a structured task description through collaborative processing using an instruction understanding model; collecting environmental data to construct a real-time environment model; generating a task execution strategy based on the structured task description and the real-time environment model using a task execution model; controlling an intelligent agent to execute tasks according to the task execution strategy and dynamically adjusting actions based on real-time sensor information; collecting task execution data and user feedback information to update the instruction understanding model and the task execution model. This invention achieves deep fusion of multimodal information by constructing a structured task description, enhances the perception of dynamic environments through a real-time environment model, improves the accuracy and flexibility of task execution through task execution strategy-driven and dynamic action adjustment, and further optimizes the model by combining task execution data and user feedback, thereby enhancing the intelligent agent's understanding, autonomous decision-making, and adaptability in complex scenarios.

[0094] In one embodiment, step S10 above includes:

[0095] S101 performs noise reduction processing on voice commands to generate noise-reduced voice signals;

[0096] S102, convert the noise-reduced speech signal into a speech-text sequence;

[0097] S103, performs word segmentation and part-of-speech tagging on the text instructions to generate a standardized text sequence;

[0098] S104, extract the acoustic and semantic features of the speech text sequence to generate a speech command feature vector;

[0099] S105, Extract the semantic features of the standardized text sequence to generate a text instruction feature vector;

[0100] S106, the voice command feature vector and the text command feature vector are fused to generate an initial fused feature vector;

[0101] S107, perform gated weighted enhancement on the initial fused feature vector to generate an enhanced fused feature vector;

[0102] S108, Align the speech text sequence with the standardized text sequence based on semantic similarity processing to generate alignment instruction data;

[0103] S109, the enhanced fusion feature vector and alignment instruction data are fused to generate a structured task description.

[0104] In this embodiment, noise reduction of voice commands refers to using signal enhancement methods to increase the proportion of effective components, targeting non-semantic elements such as background noise, echo interference, or multi-source aliasing in the voice signal. Applicable noise reduction methods include, but are not limited to, spectral subtraction, Wiener filtering, adaptive noise cancellation, and neural network speech enhancement. During implementation, the time-frequency energy distribution in the voice waveform and the distinguishing characteristics of the dominant frequency of human voice need to be considered. Time-domain or frequency-domain filter structures should be established to remove non-human voice segments or non-linguistic feature peak regions. To maintain semantic integrity, a perceptual loss function constraint is introduced during the spectral reconstruction stage to ensure consistency in semantic restoration of the reconstructed speech.

[0105] The denoised speech signal is converted into a speech-to-text sequence using an automatic speech recognition module. This module maps acoustic features to text representations and includes a joint processing chain of an acoustic model, a language model, and a pronunciation dictionary. The acoustic model typically employs a sequence modeling structure supported by Connectionist Temporal Classification (CTC) or attention mechanisms to achieve a many-to-one mapping between input speech frames and output characters. The language model provides contextual constraints at the output end to improve semantic coherence. The text output needs to be adapted to the user dictionary to ensure accurate recognition of domain keywords such as device names, location terms, and operation commands.

[0106] Word segmentation and part-of-speech tagging of text instructions involves dividing a continuous sequence of characters into semantically bounded word units and assigning a grammatical category label to each unit. This process is based on dictionary-driven methods, statistical learning, or neural models, with commonly used tools including Conditional Random Fields (CRFs) and BiLSTM+CRF sequence labeling networks. In the part-of-speech tagging stage, contextual information is used to identify the role of words in the semantic structure, providing interpretable structured information for subsequent semantic parsing.

[0107] Extracting acoustic and semantic features from speech-text sequences involves tracing back the original audio signal from the text obtained through speech recognition and performing multi-scale encoding on it. Acoustic features include Mel-frequency cepstral coefficients (MFCC), log-Mel, and pitch contour, reflecting the manner of speech production and rhythmic intensity. Semantic features are modeled from the text content using encoders such as Transformers to generate context-sensitive instruction embeddings. The combination of these two features forms a speech instruction feature vector capable of recognizing the subjective intent of the speech, possessing a dual representation capability of voiceprint emotion and language structure.

[0108] Extracting semantic features from standardized text sequences refers to vectorizing the segmented and part-of-speech-tagged text sequences using language modeling networks to form a dense set of vectors representing the meaning of each word in the context. A common approach is to use deep language models based on self-attention mechanisms, such as BERT and RoBERTa, to extract multi-layered semantic relationships, preserving features such as command intent, entity invocation, and parameter pointing. The generated text instruction feature vectors possess semantic comparability and the ability to deconstruct operational intent within the embedding space.

[0109] The feature vectors of voice commands and text commands are fused using a cross-modal fusion mechanism for vector-level synthesis. Fusion strategies may include concatenation, weighted averaging, attention fusion, or graph neural network projection fusion. In implementation, a shared representation space is designed or a matching loss function is introduced to spatially approximate the two modal representations, thereby obtaining an initial fused feature vector. The fusion process must maintain a common focus on the semantic core while preserving the supplementary information carried by the modal differences.

[0110] Gated weighted enhancement of the initial fused feature vector refers to introducing a gating mechanism (such as a sigmoid gating unit or a GRU gating unit) onto the fused multimodal features, adjusting the weights of the information channels based on context or input signal quality. This mechanism resolves intermodal conflicts or redundancy by dynamically controlling the degree of preservation of different modal features. The gating weights can be automatically learned from training data, making the enhanced fused feature vector more stable and providing a more accurate semantic representation of instructions.

[0111] Aligning speech-text sequences with standardized text sequences based on semantic similarity involves using semantic matching algorithms to identify contextual consistency and semantically equivalent regions between the two text sequences, thereby generating alignment instruction data. This can be achieved using methods such as cosine similarity, sentence vector matching, and contrastive learning matching networks to ensure a semantic mapping between the text generated from speech recognition and the instruction information in the standardized structure. In cases of slips of the tongue, synonyms, or word order changes, this alignment mechanism can help correct semantic shifts and improve the stability of downstream task structure parsing.

[0112] The process of fusing enhanced feature vectors and aligned instruction data to generate a structured task description is a joint modeling of modal fusion representation and semantic alignment results. A structured task description is a set of well-defined task semantic units, including fields such as action category, target object, parameter configuration, and execution order. This structure is generated through decoding by a task encoder and stored using a graph structure, key-value pairs, or semantic frame structure for subsequent use by the execution module. Residual connections and contextual attention mechanisms can be introduced during the fusion process to preserve the original semantics and improve the integrity of the structured representation.

[0113] This embodiment integrates and processes both voice and text commands simultaneously, leveraging the advantages of both modalities in terms of expression, information density, and semantic redundancy to significantly improve the accuracy of command recognition and the precision of task intent parsing. Multi-stage feature extraction and fusion operations introduce a learnable structural control mechanism during semantic modeling, effectively enhancing tolerance to non-standard inputs, complex language structures, and ambiguous expressions. Simultaneously, a semantic alignment mechanism compensates for speech-to-text errors, preventing task misunderstandings due to recognition mistakes. Finally, the generation of a structured task description enables subsequent system execution based on a unified task semantic format. Overall, this achieves high fault tolerance, high fidelity, and high-structure conversion of complex natural language inputs, providing a perceptual foundation with scalability and generalization capabilities for the semantic understanding module in multimodal human-computer interaction.

[0114] In one embodiment, step S108 includes:

[0115] S1081, using words in the speech-text sequence as nodes and subject-verb-object relationships generated by semantic role annotations in the speech-text sequence as edges, generate a speech-semantic graph;

[0116] S1082, using the words in the standardized text sequence as nodes and the subject-verb-object relationship generated by the semantic role annotation in the standardized text sequence as edges, a text semantic graph is generated;

[0117] S1083, Analyze the cross-modal association between the speech semantic graph and the text semantic graph using a graph convolutional network to generate a semantic association weight matrix;

[0118] S1084, Filter the semantic association weight matrix based on the semantic similarity threshold to generate valid semantic association data;

[0119] S1085, Based on the effective semantic association data, the semantic information of the speech text sequence and the standardized text sequence is fused to generate alignment instruction data.

[0120] In this embodiment, constructing a speech semantic graph using words in the speech text sequence as nodes refers to extracting all valid words from the speech text generated by speech recognition and mapping them to nodes in the graph structure. Words include semantic load items such as entity nouns, action verbs, and instruction adverbs. During node construction, elements with low semantic contribution, such as stop words and punctuation, need to be removed, while retaining keywords that appear frequently in the semantic expression of instructions or are closely related to the task intent. The speech semantic graph not only records the word content in the speech but also reconstructs and explicitly expresses semantic information in the graph space by introducing hierarchical logic of language structure.

[0121] Generating subject-verb-object relationships using semantic role labeling (SRL) in speech-text sequences refers to parsing the syntactic and semantic relationships between words in speech-text, particularly the mapping structure between the agent (subject), the action itself (verb), and the target (object), based on SRL methods. SRL models often employ Transformer-based syntactic enhancement models, identifying the predicate argument structure of each verb headword and labeling its corresponding semantic participants as roles such as ARG0 (Agent) and ARG1 (Theme). Finally, by connecting subject-verb-object triples to generate edge structures, a speech-semantic graph with semantic coherence and task orientation is constructed.

[0122] Using words in a standardized text sequence as nodes and subject-verb-object relationships generated from semantic role annotations in the standardized text sequence as edges, a text semantic graph is generated, which is completely consistent with the construction logic of a speech semantic graph. The difference lies in that standardized text sequences typically come from rule bases, template structures, or structured input instructions, resulting in more standardized language expression, clear structural boundaries, and less ambiguity. During graph construction, the same SRL processing logic is applied to the words contained in the standardized text, using the semantically clear subject-verb-object structure as the graph's topological edges to construct a graph structure capable of expressing task standards. The text semantic graph serves as an important reference benchmark for subsequent semantic alignment, localization, and semantic error correction.

[0123] Analyzing cross-modal associations between speech and text semantic graphs using graph convolutional networks (GCNs) involves establishing a multi-level nested propagation mechanism between two graph structures using the graph convolutional network (GCN) structure within a graph neural network (GNN) to calculate semantic similarity and structural matching between nodes. In implementation, each graph node is initialized with its word embedding vector, and edge structures guide the information propagation path. GCN weights and aggregates the feature vectors of adjacent nodes at each layer and projects them into a common embedding space, bringing the structures of speech text and standardized text closer together in the deeper space. Through multiple rounds of convolutional operations, a semantic association weight matrix containing cross-graph node correspondences is generated. This matrix reflects the semantic consistency, pragmatic relevance, and structural nesting degree between different nodes.

[0124] Filtering the semantic association weight matrix based on a semantic similarity threshold involves applying a threshold to the weight matrix output by the graph convolutional network, retaining only node pairs with weights greater than a set threshold as valid semantic correspondences. This threshold can be set empirically, learned through adaptive training, or a dynamic discrimination strategy. This process aims to eliminate weak association pairs caused by factors such as speech recognition errors, word order changes, and linguistic redundancy, ensuring that the associated data used in subsequent fusion stages possesses stable semantic consistency and sufficient pragmatic interpretability. The filtered, valid semantic association data can be considered a set of linguistic units capable of mapping operational intentions, providing high-quality input support for instruction semantic alignment.

[0125] Generating aligned instruction data by fusing semantic information from effective semantic association data to speech-text sequences and standardized text sequences is a process of unifying the semantic content of two input sequences after semantic mapping and structural normalization. Fusion methods can be based on graph alignment transformation, coreference resolution, or semantic frame merging, aiming to preserve, to the maximum extent possible, the true intent extension information contained in the speech input, such as additional instruction parameters, tone cues, and instruction hierarchy, while maintaining the stability of the standard instruction structure. The fusion process requires information integration and annotation based on semantic context, structural priority, and the interpretability of the execution path, ultimately forming aligned instruction data with unified semantic interpretation, structured data representation, and task adaptability, providing a semantic foundation for the generation of structured task descriptions.

[0126] This embodiment constructs a dual-graph structure of speech text and standardized text, and introduces a graph convolutional network to jointly model the semantic space. This enables deep semantic relationship mining and accurate alignment between instructions from different sources, effectively improving the fault tolerance capability in cases of word order differences, ambiguous expressions, or pronunciation abnormalities. The combination of cross-modal graph structure processing and semantic role labeling technology enables the instruction understanding system to not only recognize language content but also model the internal structural logic of language. In multimodal input collaborative understanding tasks, it provides a more clearly structured and semantically stable expression. The final generated aligned instruction data has a unified structure, accurate semantics, and executable parsing, providing high-precision and robust semantic support for subsequent structured instruction generation and task execution path planning, and realizing the effective standardized expression of complex natural language instructions in multi-source input scenarios.

[0127] In one embodiment, step S20 above includes:

[0128] S201 acquires color and depth images of the environment through a visual sensor;

[0129] S202 uses lidar to scan the environmental space and generate environmental point cloud data;

[0130] S203 detects obstacles using ultrasonic sensors and generates obstacle distance information;

[0131] S204, Extract the edge and texture features of the color image to generate environmental structure features;

[0132] S205, determine the three-dimensional coordinate information of the depth image and generate spatial depth features;

[0133] S206, perform clustering and segmentation processing on the environmental point cloud data to generate object segmentation data;

[0134] S207, integrate the environmental structural features, spatial depth features, object segmentation data and obstacle distance information to generate multi-source fusion data;

[0135] S208, Construct a dynamically updated three-dimensional spatial map based on the multi-source fusion data;

[0136] S209, Associate the three-dimensional spatial map with real-time sensor data to generate a real-time environment model.

[0137] In this embodiment, acquiring color and depth images of the environment via a visual sensor refers to continuously capturing images of the surrounding environment using an RGB camera and an RGB-D camera mounted on the intelligent agent. The color images provide color information in a two-dimensional plane, reflecting the shape, boundaries, and color composition of objects in the environment, while the depth images record the spatial distance between each pixel and the sensor, forming a preliminary bridge between the two-dimensional image and the three-dimensional structure. The visual sensor possesses high-frequency sampling capabilities and low power consumption, enabling it to continuously provide environmental perception data with stable information density under varying lighting and occlusion conditions.

[0138] Generating environmental point cloud data by scanning the environment with lidar involves using a laser ranging system to emit laser beams in all directions at high speed, receiving the reflected light, and calculating the reflection time to obtain a set of three-dimensional coordinates of a target point in space. Point cloud data is a collection of numerous dense or sparse spatial coordinate points, accurately reflecting structural information such as object surfaces, environmental boundaries, open areas, and narrow passages. LiDAR possesses millimeter-level accuracy and environmental robustness, making it one of the core data sources for spatial positioning, map building, and navigation.

[0139] Obstacle detection using ultrasonic sensors, generating obstacle distance information, involves emitting high-frequency sound waves and measuring the time difference of the echoes to determine the presence and distance of obstacles ahead. Ultrasonic sensors are insensitive to material properties and can effectively detect soft obstacles, transparent objects, and complex edges in low-visibility environments. Their data structure is typically a distance vector centered on the sensor, exhibiting high directionality and real-time performance, making them suitable for spatial fusion with image and point cloud data.

[0140] Extracting edge and texture features from color images to generate environmental structure features involves using image processing algorithms such as Canny edge detection, Sobel operators, or Gabor filters to identify object contours, boundary transitions, and repetitive texture regions from color images. These features characterize the physical structures of the environment, such as wall boundaries, door frame outlines, and furniture surfaces, forming the basis for constructing environmental topology and recognizing static targets. Environmental structure features can be further encoded into feature points using Scale Invariant Feature Transform (SIFT) or Speed-Up Robust Feature Transform (SURF) to improve the stability of image matching across time series.

[0141] Determining the 3D coordinates of a depth image and generating spatial depth features involves mapping the depth value of each pixel in the depth image back to a 3D coordinate system and calculating the true spatial location of each pixel through back projection. This process is typically based on the inverse solution of an intrinsic parameter matrix and a distortion model, and can be represented as a depth point cloud or a spatial heatmap. Spatial depth features can be used to identify structural types such as vertical planes, horizontal planes, and inclined planes, aiding in subsequent 3D modeling and target localization.

[0142] Clustering and segmenting environmental point cloud data to generate object segmentation data involves using point cloud clustering algorithms such as DBSCAN, Euclidean Cluster Extraction, or Region Growing to group closely spaced points in the point cloud data into the same cluster and extract point cloud features within the cluster, such as geometric shape, boundary lines, and size, thereby identifying independent objects or structural units in the environment. This segmentation result allows the environment to be resolved into multiple identifiable functional regions or interactive objects, serving as a necessary preprocessing step for behavioral path planning and task target localization.

[0143] Integrating environmental structural features, spatial depth features, object segmentation data, and obstacle distance information to generate multi-source fused data involves mapping these four types of information to a unified three-dimensional spatial reference coordinate system. Spatiotemporal errors are corrected using data registration and alignment methods (such as ICP, NDT, or Kalman filtering), and semantic-level information fusion processing is then performed. The fusion result is a high-dimensional representation vector set or raster map containing spatial geometry, visual semantics, interactive capability boundaries, and dynamic change features, possessing multimodal consistency and task adaptability.

[0144] Building dynamically updated 3D spatial maps based on multi-source fusion data involves projecting the fusion results onto a 3D spatial raster or point cloud structure, constructing the structure using OctoMap, VoxelMap, or TSDF methods, and then updating the layer content in real time using a time-series dynamic update strategy. Dynamic maps possess the capabilities of incremental updates, local repair, and real-time reconstruction. They can adapt to non-static changes such as object movement, occlusion changes, and fluctuations in ambient lighting, providing a stable environmental model to support task path planning and behavior adjustment.

[0145] Connecting 3D spatial maps with real-time sensor data to generate real-time environment models involves bidirectionally comparing and jointly updating the structural states in the map with real-time sensor feedback data. This allows the environment model to not only record previously known structures but also dynamically reflect the current environmental state. For example, if ultrasound detects a temporarily appearing obstacle, a temporary occupying voxel can be generated on the map; if vision identifies a new target object, semantic labels can be added in real time. Real-time environment models possess continuous adaptability, perceptual consistency, and predictive capabilities, providing a stable, interactive, and reasoning-enabled operational space foundation for intelligent agents.

[0146] This embodiment utilizes multi-sensor collaborative acquisition of environmental data across different modalities and constructs a 3D environmental model with unified semantic representation and spatial coordinate structure. This significantly enhances the cognitive representation and planning capabilities in dynamic, complex, and multi-task environments. Visual sensors provide rich structural and texture information, LiDAR and depth images offer high-precision spatial modeling capabilities, and ultrasound enhances robustness against low-visibility obstacles. This information is spatiotemporally registered and semantically fused, and a dynamic map mechanism maintains the timeliness and accuracy of the environmental representation, enabling the agent to respond instantly to environmental changes and adjust its path during task execution. The resulting real-time environmental model not only supports accurate modeling of static structures but also timely adaptation to temporary changes, achieving highly robust environmental representation capabilities for semantic task understanding. This effectively overcomes the information incompleteness and mapping delays caused by the reliance on a single modality in existing methods.

[0147] In one embodiment, step S30 above includes:

[0148] S301, parse the structured task description to obtain the task target location and operation object information;

[0149] S302, Determine the current position of the agent based on the real-time environment model;

[0150] S303, Based on the current location and the target location, plan a mobile navigation path;

[0151] S304, Identify the spatial pose information of the manipulated object in the real-time environment model;

[0152] S305, Generate a robotic arm joint motion sequence based on the spatial pose information;

[0153] S306, Detect potential obstacles in the real-time environment model and generate unexpected response strategies;

[0154] S307, integrate the mobile navigation path, robotic arm joint motion sequence and unexpected response strategy to generate a task execution strategy.

[0155] In this embodiment, parsing the structured task description to obtain the task target location and operation object information refers to using the semantic parsing module in the task execution model to accurately identify and decode each field in the structured representation, transforming the multi-dimensional vector containing information such as target area identifier, object category, behavior action, and temporal relationship into operation entities in physical space. The task target location indicates the three-dimensional coordinate position of the target area or object, which may be expressed in the form of relative coordinates, absolute position, or environmental reference points. The operation object information includes the object's category (such as switch, medicine box, door handle, etc.), material properties, and interactive methods, which are used for action planning and interaction mode generation in subsequent decision-making. This parsing process needs to be combined with semantic graph structure or task template for alignment matching, supporting the generation of differentiated behavioral targets for the same semantic mapping in different scenarios.

[0156] Determining the agent's current position based on a real-time environment model refers to estimating the agent's spatial position and orientation angle in real time using a localization module within a coordinate system jointly calibrated by a 3D dynamic map and sensors. This localization includes not only traditional x, y, and z coordinates but also attitude information (such as roll, pitch, and yaw), typically achieved through collaborative techniques such as SLAM (Simultaneous Localization and Mapping), IMU (Inertial Measurement Unit), and visual odometry. This position result serves as the starting point for subsequent path planning and must possess centimeter-level accuracy and millisecond-level timeliness to support stable path updates under high-frequency control commands.

[0157] Planning a mobile navigation path based on the current location and the target location involves using the current pose and the target pose as the start and end points, respectively, and then employing a path search algorithm based on a real-time environment map (such as A*, Dijkstra's algorithm, RRT, or Hybrid A*) to generate a set of feasible paths, taking into account obstacles, dynamic objects, path drivability, and local optima. Paths are typically represented as a sequence of trajectory points or a motion vector field, possessing attributes such as speed limits, directionality, and path smoothness. During execution, the path must also have a dynamically adjustable structure to cope with real-time events such as sudden obstacles and navigation deviations.

[0158] Identifying the spatial pose information of the manipulated object in a real-time environment model refers to analyzing the specific position (x, y, z) and orientation (roll, pitch, yaw) of the manipulated object in space based on point cloud, image, and semantic label information in the environment model, using object detection models, pose estimation algorithms, and 3D reconstruction techniques, to form a target reference frame that can be used by the robotic arm actuator. The pose estimation process typically includes four steps: object recognition, boundary fitting, surface normal extraction, and coordinate system transformation, ensuring that the accuracy of the captured object in space meets the operational tolerances of the end effector.

[0159] Generating a robotic arm joint motion sequence based on spatial pose information refers to using the target pose of the manipulated object as the end effector target, and calling an inverse kinematics solver (such as the Jacobian matrix method, numerical optimization method, or neural network solver) to calculate the sequence of joint angles required for the robotic arm to transition from its current position to the target operating pose. This sequence must satisfy constraints such as kinematic reachability, obstacle avoidance, time constraints, and stability, and requires trajectory optimization and force control adjustment in a highly redundant multi-degree-of-freedom system, ultimately forming a command flow that the control and drive system can directly execute.

[0160] Detecting potential obstacles in a real-time environment model and generating contingency response strategies involves identifying temporary obstacles (such as moving people, fallen objects, pets, etc.), potential cross-path risks, and unmapped areas outside the visible range based on fused map data and sensor streams. This detection is achieved by combining visual semantic segmentation, ultrasonic anomaly detection, and laser reflection holes. Contingency response strategies include modules for path deviation detours, task interruption and recovery decisions, dynamic priority switching, and replanning triggers, with strategy distribution accomplished through an event-driven state machine or a reinforcement learning policy execution engine.

[0161] Integrating mobile navigation paths, robotic arm joint motion sequences, and contingency response strategies to generate task execution strategies involves constructing a complete execution chain within the task execution model. This involves combining the path trajectory output by the navigation module, the sequence of actions output by the robotic arm module, and the multi-branch response strategies generated by the risk identification module in a time-synchronized and spatially coordinated manner, forming a strategy tree structure that includes phased goals, triggering conditions, and feedback constraints. The strategy can be dynamically parsed into action units, enabling real-time branch jumps, weight adjustments, and subtask switching based on the agent's state and environmental feedback, ensuring adaptability and robustness during task execution.

[0162] This embodiment employs a joint modeling mechanism of structured task semantics and 3D environmental state, and introduces a multi-path fusion strategy of spatial perception, action generation, and risk response into the task execution model. This enables the agent to autonomously complete task path planning, pose calculation, and action control in a dynamic environment with a clear goal-driven approach. Spatial understanding of target locations and interactive objects not only improves the accuracy of instruction execution but also provides a foundation for behavioral continuity and feedback control. Furthermore, by constructing a complete strategy path and embedding real-time response capabilities to unexpected events, the system's robustness and practicality in scenarios such as elderly care services, medical assistance, and unmanned delivery are significantly enhanced, solving the problems of unclear task structure, fragmented path planning, and weak dynamic obstacle avoidance capabilities in traditional voice task systems. The resulting integrated task execution strategy provides multimodal agents with efficient, controllable, and adjustable autonomous execution capabilities.

[0163] In one embodiment, step S40 above includes:

[0164] S401, Generate drive motor control instructions based on the mobile navigation path in the task execution strategy;

[0165] S402, based on the drive motor control command, control the drive motor to perform motion, and monitor the actual position data in real time;

[0166] S403, determine the deviation between the actual location data and the expected location data of the mobile navigation path, and generate location deviation information;

[0167] S404, Adjust the output torque parameters of the drive motor according to the position deviation information;

[0168] S405, when the intelligent agent reaches the task target position in the mobile navigation path, it generates joint control commands according to the robotic arm joint motion sequence in the task execution strategy.

[0169] S406, Based on the joint control commands, control the robotic arm to perform operations and collect actual force data in real time;

[0170] S407, determine the deviation between the actual force data and the preset force threshold, and generate force deviation information;

[0171] S408, Adjust the joint control command according to the force deviation information;

[0172] S409, if the robotic arm grasping operation fails, update the robotic arm joint motion sequence in the task execution strategy based on the real-time environment model.

[0173] In this embodiment, generating drive motor control commands based on the mobile navigation path in the task execution strategy refers to converting the navigation trajectory information output by the path planning module into execution commands for driving the underlying hardware. The navigation path is typically represented by a series of discrete spatial coordinate points or velocity vectors, each point associated with parameters such as timestamp, velocity constraints, acceleration limits, and path curvature. To enable the drive motor to execute these trajectories, the navigation path needs to be sampled and interpolated, smoothed, and a trajectory tracking controller (such as PurePursuit, MPC, or a PID deformation structure) needs to be invoked to generate speed and steering control commands. The motor control commands are transmitted to the controller in PWM duty cycle, current / voltage target values, or CAN bus message format, ultimately driving the wheeled chassis or tracked system to produce the desired motion.

[0174] Controlling the drive motor to execute motion based on drive motor control commands and monitoring actual position data in real time refers to simultaneously collecting current position information as feedback input for closed-loop control while controlling the motor's movement. Position data can be collected through sensors such as encoders, inertial measurement units (IMUs), visual odometry, GPS, or UWB, fusing multi-source information to form a high-precision pose estimation result. This monitoring process requires a high sampling rate and low latency to ensure the real-time performance and accuracy of subsequent deviation assessments.

[0175] Determining the deviation between the actual location data and the expected location data of the mobile navigation path, and generating position deviation information, involves comparing the real-time acquired position coordinates with the corresponding trajectory points at the current moment using differential methods to calculate information such as offset distance, angle error, and velocity error. This deviation information includes not only scalar differences in the position dimension but also multi-dimensional state deviations such as orientation difference and path curvature alignment error. Error estimation can be stabilized using sliding window filtering or Kalman filtering to generate deviation feedback quantities suitable for dynamic compensation.

[0176] Adjusting the output torque parameters of the drive motor based on position deviation information refers to inputting deviation information into the controller based on the feedback loop of the control system, and dynamically adjusting the motor control parameters through a proportional-integral-derivative (PID) regulator, adaptive control algorithm, or fuzzy logic control system. These parameters mainly include motor torque, current limit, voltage output, or encoder pulse frequency, thereby correcting path deviation phenomena generated during navigation and improving trajectory tracking accuracy and motion stability.

[0177] When the agent reaches the target position in the mobile navigation path, it generates joint control commands based on the robotic arm joint motion sequence in the task execution strategy. This means that once the pose error meets the navigation termination condition, the system switches to operation mode, extracts the robotic arm motion trajectory from the strategy (including angle changes, velocity curves, acceleration limits, etc. for each joint), and calls the joint-level control module to generate the corresponding control commands. These control commands are typically joint angular velocity target values ​​or end effector position targets, and need to be converted into low-level drive commands through inverse kinematics calculations or control mapping to adapt to serial or parallel mechanical structures with different degrees of freedom.

[0178] Controlling a robotic arm to perform operations based on joint control commands and collecting real-time force data refers to executing a target action through a joint drive controller, while simultaneously collecting data such as contact force and shear torque during the operation via end-effector force sensors, joint torque sensors, or a six-axis force sensing module. This data is used to determine whether the grip is stable, whether the applied force is appropriate, and whether there is abnormal friction, slippage, or overload.

[0179] Determining the deviation between actual force data and a preset force threshold, and generating force deviation information, involves comparing the real-time detected end-effector force with a predefined target force, calculating the differences in forward pressure, lateral friction, and torsional force, and outputting a multi-dimensional force deviation matrix through an error function. This matrix reflects the degree of deviation between the current operational action and the expected strategy, and is particularly suitable for safety adjustments in tasks such as compliant grasping and precise switching.

[0180] Adjusting joint control commands based on force deviation information refers to real-time fine-tuning based on the original motion trajectory. This includes reducing end-effector velocity, limiting joint torque, extending motion duration, or increasing the contact transition phase. This ensures that the operation achieves the mission objective without causing damage or loss of control due to excessive force or slippage. This adjustment mechanism is implemented in conjunction with adaptive impedance control, force feedback control, or composite force-position control strategies.

[0181] If the robotic arm's grasping operation fails, updating the joint motion sequence in the task execution strategy based on the real-time environment model means that after detecting the grasping failure, the re-sensing module is invoked to perform local environmental mapping and pose re-estimation of the target area, identify the new position and posture of the object being manipulated, and replan the execution path and grasping method. This process may introduce strategies such as different end-effector contact points, gripper deformation compensation, and approach path adjustment, ensuring the feasibility of re-executing the grasping operation by updating the joint sequence. The entire operation has a closed-loop structure of error detection, strategy correction, and target re-identification, improving the system's stability and fault tolerance in continuous operation in complex environments.

[0182] This embodiment organically couples mobile navigation, operation control, and multi-dimensional sensor feedback, and constructs a dynamic closed-loop relationship between motor control commands, pose feedback, joint motion control, and force adjustment, enabling the agent to autonomously execute from global path planning to local motion adjustment. During task execution, it not only maintains precise alignment between the navigation trajectory and the target action, but also performs fine-grained real-time corrections when deviations occur, improving path tracking accuracy and the safety of interactive actions. Furthermore, in the event of abnormal situations such as grasping failure, the execution trajectory can be reconstructed through a dynamic update mechanism of the environment model, providing self-recovery capability in the event of task failure.

[0183] In one embodiment, step S50 above includes:

[0184] S501 records the path execution time, grasping force data, and operation result status during the task execution process;

[0185] S502, associate the path execution time, grasping force data and operation result status to generate training sample data;

[0186] S503, use the training sample data to optimize the instructions to understand the model's weight parameters;

[0187] S504, based on the difference between the operation result status and the real-time environment model, update the three-dimensional map data;

[0188] S505 collects user satisfaction scores for task execution results;

[0189] S506, Adjust the semantic similarity threshold in the instruction understanding model according to the satisfaction score information;

[0190] S507, Analyze the path planning efficiency based on the path execution time and update the path cost function parameters;

[0191] S508, update the instruction understanding model and task execution model with the optimized weight parameters, updated 3D map data, adjusted semantic similarity threshold in the instruction understanding model, and updated path cost function parameters.

[0192] In this embodiment, recording the path execution time, grasping force data, and operation result status during task execution refers to collecting data on key process variables during each actual task execution to construct an execution sample set for model optimization. Path execution time refers to the cumulative time from the initiation of the execution strategy to the achievement of the task objective, which needs to be collected through a high-precision clock and timestamp recording system and associated with the start and end markers of the trajectory tracking controller. Grasping force data includes contact force, shear force, action time, and stability indicators during the robotic arm's operation, derived from feedback from a six-dimensional force sensor or joint actuator. The operation result status is a structured judgment result indicating whether the task was successful, interrupted, or timed out, which needs to be analyzed and output by the task monitoring module or end effector. The above data is labeled and stored according to task identifiers, constituting a multimodal synchronous execution trajectory record.

[0193] Associating path execution time, grasping intensity data, and operation result status to generate training sample data refers to constructing a standardized supervised learning training set based on time alignment, task labeling, and result attribution analysis. Sample data is typically represented as a vector structure, containing the mapping relationship between instruction encoding, operation actions, execution performance, and result feedback. Preprocessing operations such as outlier removal, time normalization, numerical standardization, and missing data completion are required during sample construction to ensure the stability of the subsequent optimization process.

[0194] Optimizing the weight parameters of a model using training sample data involves fine-tuning the trainable parameters in the word embedding layer, semantic attention mechanism, and similarity matching structure of the model, while preserving the original semantic encoding capabilities, through backpropagation algorithms or other gradient descent-like optimization strategies. This makes the model more closely resemble the semantics and operational intentions expressed by the user. This process can be performed using mini-batch iteration, mixed-precision training, and regularization strategies to avoid overfitting while maintaining training efficiency.

[0195] Updating 3D map data based on the discrepancy between the operational result state and the real-time environment model refers to the local reconstruction of the 3D environment map constructed before task execution, in cases of task failure, target positioning deviation, or inaccurate grasping posture. Discrepancy analysis requires extracting the offset from the actual perceived result of the target object and the expected pose information in the execution strategy, and then combining pose tracking algorithms and point cloud transformation algorithms (such as ICP or NDT) to complete map adjustments. The update process maintains map structure consistency using incremental or sliding window mechanisms without affecting the overall map coherence.

[0196] Collecting user satisfaction scores for task execution results refers to obtaining subjective evaluation data of the execution results through user feedback collection interfaces (such as HMI interfaces, voice evaluation systems, or multi-turn question-and-answer modules) and mapping it into structured scoring information. The scoring data can be continuous variables (e.g., 0-100 points) or discrete labels (e.g., "Satisfied," "Neutral," "Failed"), and needs to be normalized to adapt to the input requirements of subsequent model parameter adjustment modules.

[0197] Adjusting the semantic similarity threshold in the instruction comprehension model based on satisfaction score information refers to dynamically updating the semantic matching threshold in the similarity judgment module by associating user feedback with instruction comprehension results. This improves the system's adaptability to ambiguous semantic instructions or personalized expressions. For example, a dynamic threshold adjustment function can be introduced into the similarity matching module, combining historical satisfaction trends for trend convergence or divergence, automatically adjusting the judgment boundary, and optimizing semantic alignment accuracy.

[0198] Analyzing path planning efficiency based on execution time and updating path cost function parameters refers to evaluating the actual execution performance of the path planning algorithm output using time performance as an indicator, and fine-tuning the weight parameters of different factors in the cost function (such as path length, curvature, turning angle change, and drivability score). This process, through constructing a regression model or reinforcement learning feedback path between the path cost function and execution time, enables the path planning module to have dynamic parameter adaptation capabilities, thereby improving task completion efficiency.

[0199] Updating the instruction understanding model and task execution model with optimized weight parameters, updated 3D map data, adjusted semantic similarity thresholds, and updated path cost function parameters means synchronously applying various optimization results to all sub-modules of the system. This ensures that the model behavior reflects the optimization results in the next task, forming a complete model update loop. This process requires a model parameter management mechanism and a version synchronization mechanism to avoid inconsistencies in model states or loss of training states during task switching, and to ensure that the inference speed and system stability are maintained after the model update.

[0200] Example Description: In a scenario where an elderly person lives alone at home, the user issues a service command to a senior care assistant robot via voice and mobile phone input: "Pass me the medicine on the coffee table." After receiving the command, the system initiates a command parsing process, filtering the voice signal for noise and transcribing it into a speech-to-text sequence, while simultaneously recognizing the text command input from the mobile phone. During preprocessing, the text content undergoes word segmentation and part-of-speech tagging, enabling the robot to clearly identify "medicine" as a noun entity and "pass" as an action verb. Subsequently, the system extracts semantic vector representations from both input sources, generating voice command feature vectors and text command feature vectors respectively, and then fuses them to generate enhanced fusion features.

[0201] Because elderly users' speech may have unclear phrasing or ambiguous meanings, the system further utilizes a semantic similarity mechanism to semantically align the speech and text expressions. In this process, the system constructs semantic graphs for both the speech-text sequence and the standardized text sequence. Nodes in the graphs represent words, and edges represent semantic roles such as subject, verb, and object. Semantic matching relationships are then mined using a graph convolutional neural network. After filtering invalid semantic links based on a preset similarity threshold, the system generates alignment instruction data and fuses it with enhanced fusion features to ultimately obtain a structured task description, clearly identifying the task objective as "the medicine on the coffee table" and the action objective as "hand it to me."

[0202] After understanding the user's instructions, the elderly care robot immediately activates its perception module to collect data from the home environment. Its onboard visual sensors acquire color and depth images, LiDAR scans the spatial point cloud, and ultrasonic sensors detect potential obstacles, such as table corners, chairs, or walls. The system extracts texture edges from the environmental images, 3D coordinate information from the depth map, and object outlines and segmented regions from the point cloud data, combining this with ultrasonic distance information to create multi-source data. Based on this fused data, the system constructs a dynamic 3D spatial map and continuously updates it to ensure the robot's accurate recognition of furniture, passageways, and obstacles.

[0203] Subsequently, the system analyzes the target location and object information in the structured task description, locating the medicine and the elderly person's position on a 3D map. The path planning module generates feasible movement paths based on the map and the agent's current position, avoiding obstacles and narrow passages, and combines this with the object recognition module to determine the medicine's pose. The robotic arm planning module generates a grasping trajectory and end-effector parameters based on the identified pose, while also identifying potential influencing factors such as cats and slippers and developing corresponding strategies, such as temporary detours or task pauses, comprehensively constructing the final task execution strategy.

[0204] During the execution phase, the agent moves to the vicinity of the coffee table according to the navigation path, while simultaneously monitoring positional deviations and fine-tuning motor outputs. Upon reaching the designated position, the robotic arm unfolds its movements according to a predetermined grasping strategy, adjusting the gripping force based on force feedback to ensure that the medicine is neither crushed nor dropped. When the actual force data deviates significantly from the preset grasping threshold, the system automatically adjusts the joint control commands or replans the robotic arm trajectory to perform grasping compensation operations. In the event of a grasping failure, the system remodels the grasping posture based on the updated environment model and retrys.

[0205] After the entire task is completed, the system records the path time, operation results, and changes in intensity during the execution process, and collects user satisfaction evaluations, such as asking "How do you feel about the grasping speed and accuracy?" via voice Q&A. This data, after being structured, is used to optimize the model. Path execution time is used to evaluate navigation efficiency and update the path cost function; grasping intensity data and success rate are used together to improve the stability of the execution model; user satisfaction is used to calibrate the semantic similarity threshold in the semantic understanding model, enabling the system to better adapt to personalized language styles.

[0206] Meanwhile, if discrepancies arise between the environmental map and the actual environment during the task (such as a shift in the location of medicines), the system will dynamically correct the map based on the operation results and sensor feedback, enabling the next task to adapt more quickly to new real-world scenarios. All model updates are ultimately written synchronously into the model management module, ensuring that instruction recognition and action planning can continuously evolve and become more adaptable for the next task.

[0207] Through the above operation process, the elderly care assistant robot has achieved a closed-loop processing capability from understanding user instructions, environmental perception, path and action planning, task execution to feedback learning, effectively improving the service response quality, environmental adaptability and user interaction intelligence level in elderly care scenarios.

[0208] In a rehabilitation hospital, a doctor commands a smart service robot in a ward via voice and mobile terminal input: "Bring the blood pressure monitor from the bedside table." The robot receives the voice signal and simultaneously parses the synonymous text command entered by the doctor in the system interface, then begins to execute the task. The system first performs noise reduction processing on the raw voice data, filtering out background noise such as fan sounds and patient voices to generate a clean voice signal. This voice signal is then converted into a speech-to-text sequence by an automatic speech recognition module. The system then performs word segmentation and part-of-speech analysis on the doctor's input text command in parallel, constructing a standardized text sequence and identifying "blood pressure monitor" as the target object and "bring" as the action verb.

[0209] Subsequently, the system extracts semantic information from the speech-text sequence and the standardized text sequence respectively, constructs speech feature vectors and text feature vectors, and generates fused features through a fusion mechanism. Since medical instructions often contain proper nouns or technical terms (such as "cuff-type blood pressure monitor" or "ambulatory blood pressure monitor"), the system further mines potential correspondences between texts using semantic graphs. By comparing subject-verb-object role edges between the two semantic graphs through a graph convolutional network, the system generates a semantic association weight matrix and filters out high-confidence semantic alignment relationships using a semantic similarity threshold. Finally, the aligned instruction data is synthesized and fused into the fused features to generate a structured task description.

[0210] The structured task description clearly defines the task objective as "the blood pressure monitor on the bedside table," located on the right front side of the hospital bed, and the target object is a medical device. The service robot then initiates its environmental modeling module. Its vision system acquires color and depth images of the ward interior and generates a spatial point cloud model including the bed, cabinet, and chair using LiDAR. Ultrasonic sensors detect potential obstacles in the space, such as bedside handrails and patient walking aids. Through edge and texture extraction from the image data, 3D reconstruction of the depth map, and point cloud clustering, the system identifies the location and shape of various objects in the ward, integrates them into a unified multi-source dataset, and generates a 3D spatial map.

[0211] This 3D map is continuously updated and bound to sensor data to determine the current position of the agent and the orientation of the target object. The system obtains the target pose of the blood pressure monitor and specific grasping requirements through a structured task description, and generates the optimal navigation path based on the agent's current position information in the map. Before reaching the bedside table, the robot detects that there are patient's daily necessities piled up on one side of the table. To prevent collisions, the system dynamically generates an obstacle avoidance trajectory based on the point cloud, while simultaneously calculating the grasping posture for the robotic arm, determining the action sequence of each joint, and generating an execution strategy.

[0212] The control module translates the navigation path into low-latency control commands for the drive motors based on the task execution strategy, propelling the robot's motion unit forward slowly. During movement, the robot acquires its own position data in real time and calculates the positional deviation from the planned path, dynamically adjusting the motor output power to ensure precise control in complex spatial environments. When the robotic arm performs a grasping operation in the target area, force sensors detect in real time whether the force applied to the blood pressure monitor exceeds the maximum load-bearing threshold of the medical device. If excessive force or positional deviation occurs, the system immediately corrects the joint control parameters to ensure safe and reliable grasping.

[0213] After grasping the instrument, the service robot delivers it to the doctor's designated location. Throughout the process, key parameters of task execution, such as navigation time, grasping success rate, number of robotic arm movements, and force control feedback values, are fully recorded. Simultaneously, the system accesses the patient care evaluation interface to collect feedback from doctors on the execution effectiveness, such as scores for operation speed, path rationality, and grasping accuracy.

[0214] This execution data, along with user feedback, will be used synchronously for model updates. Path execution time is used to evaluate path efficiency, thereby adjusting the cost function in the path planning module to make future planning more inclined towards efficient paths; grasping strength and failure rates are used to feed back into the action model part of the training task execution model; and doctors' satisfaction evaluations are fed back into the instruction understanding model, updating semantic thresholds to better adapt to the language style of medical professional terminology. Furthermore, if discrepancies are found between the map and the real environment during a task, such as a shifted bed or cabinet, the system will update the environmental map based on operation failure points and sensor feedback, improving the environmental accuracy for the next task.

[0215] Inside a smart bank branch, a customer approaches an intelligent reception robot and says, "Please print my account statement from last month," while simultaneously clicking the "Print Statement" shortcut on their mobile banking app, creating a multimodal concurrent command input. The system collects voice data via a microphone array and receives structured text commands from the mobile device. To eliminate background noise interference, such as conversations from other customers and air conditioning operation, the system performs spectral analysis and filtering on the raw voice data to generate a noise-reduced voice signal, which is then converted into a speech-to-text sequence, "Please print my account statement from last month," using the speech recognition module.

[0216] In parallel, the mobile command undergoes keyword extraction, word segmentation, part-of-speech tagging, and semantic standardization to generate a standardized text sequence "Print July 2025 account statement". The system extracts named entities (such as month and business type) from both the spoken text and the standardized text, generates semantic vectors, and fuses them into an initial feature vector. To eliminate ambiguity caused by semantic inconsistencies, such as "last month" and "July 2025", the system constructs two semantic graphs, building node-edge structures based on subject-verb-object relationships. It then uses a graph convolutional network to infer cross-modal semantic similarity between these graphs, extracting the optimal matching pair to achieve semantic-level alignment and generate aligned fused task command data.

[0217] The system generates a structured task description based on the alignment results: the business objective is "account transaction record printing," the object is the current customer's linked account, the target time period is "July 2025," and the service location is the "printing area." Subsequently, the environmental perception system begins collecting environmental data within the bank branch. High-definition images and depth information are acquired through visual sensors deployed on the ceiling and intelligent agents to identify the layout of physical facilities such as the front desk, self-service terminals, and waiting seating areas. Spatial paths are detected using LiDAR point clouds, and temporary obstacles, such as customers lingering or children running, are sensed through ultrasonic sensors.

[0218] The system integrates image edge features, depth coordinates, point cloud structure, and obstacle distance information to update a 3D spatial map in real time and binds it to the current sensor status, forming a dynamic real-time environment model. Based on the printing area target in the task description, the system plans the shortest obstacle avoidance path according to the environment model and, combined with the location of the printing terminal and the customer queuing status, generates a movement path and operating parameters for the service robot.

[0219] When the service robot begins to move, its drive control system generates chassis motor drive parameters based on path planning instructions and collects its own position coordinates in space in real time. If the robot detects that its current position deviates from the path centerline or exhibits veer-off behavior, the system immediately adjusts the motor torque output to ensure that it accurately reaches the target area even in congested environments. Upon reaching the printing area, the system invokes the robotic arm module according to the "print operation" in the structured task, pressing the printer button, placing paper, or removing the item to be printed. Throughout the entire operation, force control sensors monitor the gripping force in real time to prevent damage to the printing equipment or paper jams.

[0220] If the grasping fails, such as when the printed paper gets stuck in the slot, the system dynamically updates the robotic arm's pose and generates corrective control parameters based on visual and force control anomaly data, then re-executes the grasping action to ensure continuous and stable operation. After the operation is completed, the service robot delivers the printed invoice to the customer and enters the learning and optimization phase based on the path time, force control deviation, and customer feedback data during the task execution.

[0221] The system records path execution time, operation success rate, and execution duration, and combines these with the current 3D map and instruction understanding data to form training samples. If the printing area device position is offset or the desktop is congested during the task, the system automatically updates the spatial map structure, making the next task planning more accurate. Simultaneously, customers can rate the service through an evaluation interface, assessing aspects such as operation speed, path rationality, and output clarity. The system uses this rating to adjust the semantic matching threshold in the instruction understanding model, making subsequent models more suitable for the customer's expression style in financial business.

[0222] Furthermore, the system analyzes the differences between task path execution time and historical records, dynamically adjusting path cost function parameters. For example, it prioritizes routes through less crowded areas or avoids high-frequency interaction zones, optimizing subsequent navigation strategies. All optimized parameters are uniformly updated in the instruction understanding and task execution model, enabling the service robot to self-evolve and continuously adapt.

[0223] This embodiment collects multi-source feedback information during task execution and establishes a data-driven execution behavior learning mechanism based on operation results and user evaluations, achieving the co-evolution capability of the language instruction understanding model and the task execution model. Path execution time and operation performance are mapped as optimization factors to update model weights, path cost functions, and semantic matching thresholds, thereby improving understanding accuracy and execution efficiency. Driven by feedback, it can not only dynamically reconstruct 3D environment maps and identify environmental changes, but also respond to user preferences to achieve personalized optimization of semantic understanding and action execution strategies. This structure supports long-term continuous online model updates and has the ability to gradually adapt to new data, new expressions, and new goals in complex task scenarios, significantly improving the system's task completion quality and user satisfaction in real elderly care, medical, or service environments.

[0224] In one embodiment, an instruction understanding and task execution apparatus is provided, which corresponds one-to-one with the instruction understanding and task execution methods described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the instruction understanding and task execution device of the present invention. The modules include an instruction understanding module 10, an environment modeling module 20, a task strategy generation module 30, a task execution control module 40, and a model adaptive update module 50. Detailed descriptions of each functional module are as follows:

[0225] The instruction understanding module 10 is used to acquire voice instructions and text instructions, and to perform collaborative processing on the voice instructions and text instructions through the instruction understanding model to generate a structured task description;

[0226] The environment modeling module 20 is used to collect environmental data and construct a real-time environment model based on the environmental data.

[0227] The task strategy generation module 30 is used to generate a task execution strategy based on the structured task description and the real-time environment model through the task execution model.

[0228] The task execution control module 40 is used to control the intelligent agent to execute tasks according to the task execution strategy, and to adjust the actions of the intelligent agent according to real-time sensor information during the task execution process;

[0229] The model adaptive update module 50 is used to collect task execution data and user feedback information, and update the instruction understanding model and task execution model according to the task execution data and user feedback information.

[0230] In one embodiment, the instruction understanding module 10 is specifically used for:

[0231] The voice commands are processed to reduce noise, generating a noise-reduced voice signal;

[0232] The noise-reduced speech signal is converted into a speech-text sequence;

[0233] The text instructions are segmented and part-of-speech tagged to generate a standardized text sequence;

[0234] The acoustic and semantic features of the speech text sequence are extracted to generate a speech command feature vector;

[0235] Extract the semantic features of the standardized text sequence to generate a text instruction feature vector;

[0236] The voice command feature vector and the text command feature vector are fused to generate an initial fused feature vector;

[0237] The initial fused feature vector is subjected to gated weighting enhancement to generate an enhanced fused feature vector;

[0238] Based on semantic similarity processing, the speech text sequence is aligned with the standardized text sequence to generate alignment instruction data;

[0239] The enhanced fusion feature vector and alignment instruction data are fused to generate a structured task description.

[0240] In one embodiment, the instruction understanding module 10 is specifically used for:

[0241] A speech semantic graph is generated using words in the speech text sequence as nodes and subject-verb-object relationships generated by semantic role annotations in the speech text sequence as edges.

[0242] A text semantic graph is generated by using words in the standardized text sequence as nodes and subject-verb-object relationships generated by semantic role annotations in the standardized text sequence as edges.

[0243] The cross-modal associations of the speech semantic graph and the text semantic graph are analyzed using graph convolutional networks to generate a semantic association weight matrix.

[0244] The semantic association weight matrix is ​​filtered based on a semantic similarity threshold to generate valid semantic association data;

[0245] Based on the effective semantic association data, the semantic information of the speech text sequence and the standardized text sequence is fused to generate alignment instruction data.

[0246] In one embodiment, the environment modeling module 20 is specifically used for:

[0247] Acquire color and depth images of the environment using a visual sensor;

[0248] Environmental point cloud data is generated by scanning the environmental space with lidar;

[0249] Obstacles are detected using ultrasonic sensors, and obstacle distance information is generated.

[0250] Extract edge and texture features from the color image to generate environmental structure features;

[0251] Determine the three-dimensional coordinate information of the depth image to generate spatial depth features;

[0252] The environmental point cloud data is clustered and segmented to generate object segmentation data;

[0253] By fusing the environmental structural features, spatial depth features, object segmentation data, and obstacle distance information, multi-source fusion data is generated.

[0254] A dynamically updated 3D spatial map is constructed based on the multi-source fusion data;

[0255] By linking the 3D spatial map with real-time sensor data, a real-time environment model is generated.

[0256] In one embodiment, the task strategy generation module 30 is specifically used for:

[0257] Parse the structured task description to obtain the task target location and operation object information;

[0258] The current position of the agent is determined based on the real-time environment model;

[0259] Based on the current location and the target location, plan a mobile navigation path;

[0260] Identify the spatial pose information of the manipulated object in the real-time environment model;

[0261] Generate a sequence of joint motions for the robotic arm based on the spatial pose information;

[0262] Detect potential obstacles in the real-time environment model and generate contingency response strategies;

[0263] The mobile navigation path, robotic arm joint motion sequence, and unexpected response strategy are integrated to generate a task execution strategy.

[0264] In one embodiment, the task execution control module 40 is specifically used for:

[0265] Generate drive motor control commands based on the mobile navigation path in the task execution strategy;

[0266] The drive motor is controlled to perform motion based on the drive motor control command, and the actual position data is monitored in real time;

[0267] Determine the deviation between the actual location data and the expected location data of the mobile navigation path, and generate location deviation information;

[0268] Adjust the output torque parameters of the drive motor based on the position deviation information;

[0269] When the agent reaches the task target location in the mobile navigation path, it generates joint control commands based on the robotic arm joint motion sequence in the task execution strategy.

[0270] The robotic arm is controlled to perform operations based on the joint control commands, and actual force data is collected in real time.

[0271] Determine the deviation between the actual force data and the preset force threshold, and generate force deviation information;

[0272] The joint control command is adjusted based on the force deviation information;

[0273] If the robotic arm grasping operation fails, the robotic arm joint motion sequence in the task execution strategy is updated based on the real-time environment model.

[0274] In one embodiment, the model adaptive update module 50 is specifically used for:

[0275] Record the path execution time, grasping force data, and operation result status during the task execution process;

[0276] The path execution time, grasping force data, and operation result status are correlated to generate training sample data;

[0277] The training sample data is used to optimize instructions to understand the model's weight parameters;

[0278] Update the 3D map data based on the difference between the operation result status and the real-time environment model;

[0279] Collect user satisfaction scores for task execution results;

[0280] The semantic similarity threshold in the instruction understanding model is adjusted based on the satisfaction score information;

[0281] Analyze the path planning efficiency based on the path execution time, and update the path cost function parameters accordingly;

[0282] The optimized weight parameters, updated 3D map data, adjusted semantic similarity threshold in the instruction understanding model, and updated path cost function parameters are then used to update the instruction understanding model and task execution model.

[0283] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements a server-side function or step of an instruction understanding and task execution method.

[0284] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements a user-side function or step of an instruction understanding and task execution method.

[0285] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0286] Acquire voice commands and text commands, and use a command understanding model to collaboratively process the voice commands and text commands to generate a structured task description;

[0287] Collect environmental data and construct a real-time environmental model based on the environmental data;

[0288] Based on the structured task description and the real-time environment model, a task execution strategy is generated through the task execution model;

[0289] The agent is controlled to perform tasks according to the task execution strategy, and the agent's actions are adjusted according to real-time sensor information during task execution.

[0290] Collect task execution data and user feedback information, and update the instruction understanding model and task execution model based on the task execution data and user feedback information.

[0291] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0292] Acquire voice commands and text commands, and use a command understanding model to collaboratively process the voice commands and text commands to generate a structured task description;

[0293] Collect environmental data and construct a real-time environmental model based on the environmental data;

[0294] Based on the structured task description and the real-time environment model, a task execution strategy is generated through the task execution model;

[0295] The agent is controlled to perform tasks according to the task execution strategy, and the agent's actions are adjusted according to real-time sensor information during task execution.

[0296] Collect task execution data and user feedback information, and update the instruction understanding model and task execution model based on the task execution data and user feedback information.

[0297] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0298] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0299] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0300] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for instruction understanding and task execution, characterized in that, Includes the following steps: Acquire voice commands and text commands, and use a command understanding model to collaboratively process the voice commands and text commands to generate a structured task description; Collect environmental data and construct a real-time environmental model based on the environmental data; Based on the structured task description and the real-time environment model, a task execution strategy is generated through the task execution model; The agent is controlled to perform tasks according to the task execution strategy, and the agent's actions are adjusted according to real-time sensor information during task execution. Collect task execution data and user feedback information, and update the instruction understanding model and task execution model based on the task execution data and user feedback information.

2. The instruction understanding and task execution method as described in claim 1, characterized in that, Acquire voice and text commands, and use a command understanding model to collaboratively process the voice and text commands to generate a structured task description, including: The voice commands are processed to reduce noise, generating a noise-reduced voice signal; The noise-reduced speech signal is converted into a speech-text sequence; The text instructions are segmented and part-of-speech tagged to generate a standardized text sequence; The acoustic and semantic features of the speech text sequence are extracted to generate a speech command feature vector; Extract the semantic features of the standardized text sequence to generate a text instruction feature vector; The voice command feature vector and the text command feature vector are fused to generate an initial fused feature vector; The initial fused feature vector is subjected to gated weighting enhancement to generate an enhanced fused feature vector; Based on semantic similarity processing, the speech text sequence is aligned with the standardized text sequence to generate alignment instruction data; The enhanced fusion feature vector and alignment instruction data are fused to generate a structured task description.

3. The instruction understanding and task execution method as described in claim 2, characterized in that, Aligning the speech-text sequence with the standardized text sequence based on semantic similarity processing, and generating alignment instruction data, including: A speech semantic graph is generated using words in the speech text sequence as nodes and subject-verb-object relationships generated by semantic role annotations in the speech text sequence as edges. A text semantic graph is generated by using words in the standardized text sequence as nodes and subject-verb-object relationships generated by semantic role annotations in the standardized text sequence as edges. The cross-modal associations of the speech semantic graph and the text semantic graph are analyzed using graph convolutional networks to generate a semantic association weight matrix. The semantic association weight matrix is ​​filtered based on a semantic similarity threshold to generate valid semantic association data; Based on the effective semantic association data, the semantic information of the speech text sequence and the standardized text sequence is fused to generate alignment instruction data.

4. The instruction understanding and task execution method as described in claim 1, characterized in that, Collect environmental data and construct a real-time environmental model based on the environmental data, including: Acquire color and depth images of the environment using a visual sensor; Environmental point cloud data is generated by scanning the environmental space with lidar; Obstacles are detected using ultrasonic sensors, and obstacle distance information is generated. Extract edge and texture features from the color image to generate environmental structure features; Determine the three-dimensional coordinate information of the depth image to generate spatial depth features; The environmental point cloud data is clustered and segmented to generate object segmentation data; By fusing the environmental structural features, spatial depth features, object segmentation data, and obstacle distance information, multi-source fusion data is generated. A dynamically updated three-dimensional spatial map is constructed based on the multi-source fusion data; By linking the 3D spatial map with real-time sensor data, a real-time environment model is generated.

5. The instruction understanding and task execution method as described in claim 1, characterized in that, Based on the structured task description and the real-time environment model, a task execution strategy is generated through the task execution model, including: Parse the structured task description to obtain the task target location and operation object information; The current position of the agent is determined based on the real-time environment model; Based on the current location and the target location, plan a mobile navigation path; Identify the spatial pose information of the manipulated object in the real-time environment model; Generate a sequence of joint motions for the robotic arm based on the spatial pose information; Detect potential obstacles in the real-time environment model and generate contingency response strategies; The mobile navigation path, robotic arm joint motion sequence, and unexpected response strategy are integrated to generate a task execution strategy.

6. The instruction understanding and task execution method as described in claim 1, characterized in that, The agent is controlled to perform tasks according to the task execution strategy, and the agent's actions are adjusted based on real-time sensor information during task execution, including: Generate drive motor control commands based on the mobile navigation path in the task execution strategy; The drive motor is controlled to perform motion based on the drive motor control command, and the actual position data is monitored in real time; Determine the deviation between the actual location data and the expected location data of the mobile navigation path, and generate location deviation information; Adjust the output torque parameters of the drive motor based on the position deviation information; When the agent reaches the task target location in the mobile navigation path, it generates joint control commands based on the robotic arm joint motion sequence in the task execution strategy. The robotic arm is controlled to perform operations based on the joint control commands, and actual force data is collected in real time. Determine the deviation between the actual force data and the preset force threshold, and generate force deviation information; The joint control command is adjusted based on the force deviation information; If the robotic arm grasping operation fails, the robotic arm joint motion sequence in the task execution strategy is updated based on the real-time environment model.

7. The instruction understanding and task execution method as described in claim 1, characterized in that, Collecting task execution data and user feedback information, and updating the instruction understanding model and task execution model based on the task execution data and user feedback information, including: Record the path execution time, grasping force data, and operation result status during the task execution process; The path execution time, grasping force data, and operation result status are correlated to generate training sample data; The training sample data is used to optimize instructions to understand the model's weight parameters; Update the 3D map data based on the difference between the operation result status and the real-time environment model; Collect user satisfaction scores for task execution results; The semantic similarity threshold in the instruction understanding model is adjusted based on the satisfaction score information; Analyze the path planning efficiency based on the path execution time, and update the path cost function parameters accordingly; The optimized weight parameters, updated 3D map data, adjusted semantic similarity threshold in the instruction understanding model, and updated path cost function parameters are then used to update the instruction understanding model and task execution model.

8. An instruction understanding and task execution device, characterized in that, The instruction understanding and task execution device includes: The instruction understanding module is used to acquire voice instructions and text instructions, and to perform collaborative processing of the voice instructions and text instructions through the instruction understanding model to generate a structured task description; The environmental modeling module is used to collect environmental data and construct a real-time environmental model based on the environmental data. The task strategy generation module is used to generate a task execution strategy based on the structured task description and the real-time environment model through the task execution model. The task execution control module is used to control the intelligent agent to execute tasks according to the task execution strategy, and to adjust the actions of the intelligent agent according to real-time sensor information during the task execution process; The model adaptive update module is used to collect task execution data and user feedback information, and update the instruction understanding model and task execution model based on the task execution data and user feedback information.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and an instruction understanding and task execution program stored in the memory and executable on the processor, wherein the instruction understanding and task execution program, when executed by the processor, implements the steps of the instruction understanding and task execution method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores an instruction understanding and task execution program, which, when executed by a processor, implements the steps of the instruction understanding and task execution method as described in any one of claims 1-7.

Citation Information

Cited By

  • Method, device and equipment for controlling mechanical arm and storage medium

    CN121315988A

  • Task processing method and system for robot with body based on context self-adaption

    CN121650020A

  • Context-adaptive embodied robot task processing method and system

    CN121650020B

  • Cluster fixed-wing unmanned aerial vehicle operation method, system, equipment and medium

    CN121936609A

  • Community service robot intelligent decision-making system based on large model and behavior tree

    CN122021835A