Context-adaptive embodied robot task processing method and system

By classifying user commands and generating intent-guided suggestions, embodied robots can accurately understand and execute ambiguous commands, overcoming the obstacle of embodied robots in understanding ambiguous commands and improving the reliability and efficiency of task processing.

CN121650020BActive Publication Date: 2026-05-19WOCAO TECH (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WOCAO TECH (SHENZHEN) CO LTD
Filing Date
2026-02-06
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Embossed robots cannot effectively understand and execute ambiguous instructions, resulting in low task processing efficiency. Furthermore, existing technologies lack the ability to understand open contexts, making it difficult to guarantee the reliability and safety of task execution.

Method used

By receiving user commands, classifying command types, generating intent guidance suggestions, reconstructing the context based on command context information, performing semantic analysis, and generating target action sequences to execute tasks.

Benefits of technology

It enhances the embodied robot's ability to understand ambiguous instructions, ensuring an accurate correlation between task intent and user needs, avoiding task errors, and improving task processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121650020B_ABST
    Figure CN121650020B_ABST
Patent Text Reader

Abstract

The application relates to a context-adaptive embodied robot task processing method and system. The method comprises the following steps: receiving a user instruction, classifying the user instruction to obtain an instruction type; in the case that the instruction type is an ambiguous instruction type, outputting an intention guidance suggestion for the user instruction; the intention guidance suggestion is generated based on instruction context information corresponding to the user instruction, and the intention guidance suggestion is used for guiding the user to clarify the task intention; based on the intention guidance suggestion, the instruction context information is reconstructed to obtain reconstructed context information; receiving a feedback instruction of the user for the intention guidance suggestion, performing semantic analysis on the feedback instruction based on the reconstructed context information to obtain a feedback task intention; in the case that the feedback task intention indicates an execution target task, generating a target action sequence based on the target task, and controlling the embodied robot to execute the target action sequence. The method can improve the task processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of embodied robot technology, and in particular to a context-adaptive embodied robot task processing method and system. Background Technology

[0002] With the development of artificial intelligence and robotics, embodied robots are gradually integrating into daily life. For example, in daily life, embodied robots can perform corresponding tasks according to user instructions.

[0003] In traditional technologies, embodied robots can only understand and execute user instructions of a specific type, such as "pick up an apple" or "go to the kitchen," which are preset or have fixed syntax. They lack the ability to understand user instructions of a vague type (i.e., non-preset or non-fixed syntax), which can easily lead to errors in task execution or failure to perform tasks, resulting in low task processing efficiency. Summary of the Invention

[0004] Therefore, it is necessary to provide a context-adaptive embodied robot task processing method, embodied robot system, computer device, computer-readable storage medium, and computer program product that can improve task processing efficiency in response to the above-mentioned technical problems.

[0005] On the one hand, this application provides a context-adaptive embodied robot task processing method, including: receiving user instructions and classifying the user instructions to obtain instruction types; when the instruction type is an ambiguous instruction type, outputting intent guidance suggestions for the user instructions; the intent guidance suggestions are generated based on the instruction context information corresponding to the user instructions, and are used to guide the user to clarify the task intent; based on the intent guidance suggestions, reconstructing the instruction context information to obtain reconstructed context information; receiving user feedback instructions for the intent guidance suggestions, performing semantic analysis on the feedback instructions based on the reconstructed context information to obtain the feedback task intent; when the feedback task intent indicates the execution of a target task, generating a target action sequence based on the target task, and controlling the embodied robot to execute the target action sequence.

[0006] On the other hand, this application also provides an embodied robot system, which is used to: receive user instructions and classify the user instructions to obtain instruction types; when the instruction type is an ambiguous instruction type, output intention guidance suggestions for the user instructions; the intention guidance suggestions are generated based on the instruction context information corresponding to the user instructions, and are used to guide the user to clarify the task intention; based on the intention guidance suggestions, reconstruct the instruction context information to obtain reconstructed context information; receive user feedback instructions for the intention guidance suggestions, and perform semantic analysis on the feedback instructions based on the reconstructed context information to obtain the feedback task intention; when the feedback task intention indicates the execution of a target task, generate a target action sequence based on the target task, and control the embodied robot to execute the target action sequence.

[0007] On the other hand, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described context-adaptive embodied robot task processing method.

[0008] On the other hand, this application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps in the above-described context-adaptive embodied robot task processing method.

[0009] On the other hand, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps in the above-described context-adaptive embodied robot task processing method.

[0010] The aforementioned context-adaptive embodied robot task processing method, embodied robot system, computer device, computer-readable storage medium, and computer program product improve the reliability of intent guidance suggestions because these suggestions are generated based on the instruction context information corresponding to the user's command. Intent guidance suggestions guide the user to clarify the task intent, transforming vague user needs into clear and executable target tasks, thus clarifying the target task the user wants to perform (i.e., the correct one). Based on the intent guidance suggestions, the instruction context information is reconstructed to obtain reconstructed context information. Semantic analysis of the feedback command is then performed based on this reconstructed context information to obtain the feedback task intent. This ensures an accurate association between the feedback task intent and the user's needs, avoiding interference from historical context information. Only when the user's feedback task intent instructs the execution of the target task will the corresponding target action sequence be executed. This ensures the executed task is correct (i.e., what the user wants to perform), avoiding task execution errors and preventing situations where user commands based on vague command types cannot execute tasks, thereby improving task processing efficiency. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is an application environment diagram of a context-adaptive embodied robot task processing method in one embodiment;

[0013] Figure 2 This is a flowchart illustrating a context-adaptive embodied robot task processing method in one embodiment;

[0014] Figure 3 This is a flowchart illustrating a context-adaptive embodied robot task processing method in another embodiment;

[0015] Figure 4 This is an overall flowchart of a context-adaptive embodied robot task processing method in one embodiment;

[0016] Figure 5 This is an internal structural diagram of a computer device in one embodiment;

[0017] Figure 6 This is a diagram of the internal structure of a computer device in another embodiment. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0019] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0020] The context-adaptive embodied robot task processing method provided in this application can be applied to, for example... Figure 1In the application environment shown, the application scenario includes an embodied robot 102 and a target environment, with the embodied robot 102 residing in the target environment. The embodied robot 102 possesses sensing, movement, and interaction capabilities, and can interact with its environment in real time. It can capture information about its surrounding environment through sensory organs such as cameras, LiDAR, or tactile LiDAR installed on the embodied robot 102. The embodied robot 102 can be a home service robot, including but not limited to: cleaning robots, companion robots, and humanoid robots. Cleaning robots can be, but are not limited to: sweeping robots, mopping robots, and combined sweeping and mopping robots. The target environment can be any environment where the embodied robot 102 can be applied, including but not limited to: home indoor environments, industrial manufacturing environments, medical environments, warehousing and logistics environments, agricultural production environments, public service environments, educational and research environments, or cultural and entertainment performance environments.

[0021] In related technologies, human-computer task interaction schemes can be mainly divided into two categories: (1) Interaction schemes based on fixed instruction sets: Embodied robots can only understand and execute a limited set of preset or grammatically fixed instructions (such as "pick up an apple" or "go to the kitchen"). This type of scheme usually lacks the ability to understand user instructions and contexts with ambiguous instruction types, such as being unable to handle user instructions such as "I am thirsty". (2) Interaction schemes based on end-to-end models: Using large-scale language models or visual language models, user instructions are directly received and responses or preliminary action plans are generated. Although this type of scheme improves the naturalness of dialogue, its decision-making process is like a "black box", lacking controllable and structured task decomposition logic. Especially when it is necessary to combine real-time environmental perception to execute multi-step physical tasks, problems such as logical confusion, redundant steps, or non-compliance with physical world constraints often occur, making it difficult to guarantee the reliability and safety of task execution.

[0022] This application aims to address the deficiencies in the aforementioned related technologies, and the technical problems it addresses include: how to enable embodied robots to accurately understand vague or specific user instructions, and to dynamically orchestrate user instructions into executable action sequences that conform to the real-time physical environment through a controllable, reliable, and efficient interactive process. Specifically, it aims to overcome the embodied robot's obstacle to understanding open contexts, achieve context-aware intelligent dialogue management, and generate task plans that adapt to environmental states.

[0023] Based on this, in an exemplary embodiment, such as Figure 2 As shown, a context-adaptive embodied robot task processing method is provided, executed by embodied robot 102 (hereinafter referred to as the robot), including the following steps:

[0024] Step 202: Receive user instructions and classify the user instructions to obtain the instruction type.

[0025] User commands are instructions issued by the user. User commands can be in voice or text form. For example, user commands can be voice-triggered commands such as "Bring me an apple," "I'm thirsty," "Find me something good to eat," and "What's the weather like today?" User commands can also be commands triggered by the user through the app controlling the robot, or button commands triggered by buttons on the robot.

[0026] Understandably, user instructions can refer to commands from a user to instruct a robot to perform a task. In one approach, when the robot can only understand a limited set of pre-defined or grammatically fixed instructions, after receiving a user instruction, the robot needs to analyze it. If the user instruction is pre-defined or grammatically fixed, it is a descriptive instruction type for the robot. If the user instruction is not pre-defined or grammatically ambiguous, it is a vague instruction type for the robot. In other words, when the user instruction is incomprehensible to the robot, its instruction type is a model instruction type; when the user instruction is comprehensible to the robot, its instruction type is a descriptive instruction type.

[0027] Of course, when the user command is not preset or has an inconsistent syntax, or when there are no preset or fixed syntax commands, the robot can parse the user command. If the user command contains a specific object to be manipulated and / or an action to be performed on that object, then the command type can be determined to be a definite command type. If the user command lacks a specific object to be manipulated and / or an action to be performed on that object, then the command type can be determined to be a fuzzy command type. For example, if the user command is "Help me find something delicious," although it indicates searching for food, the user command does not have a specific object to be manipulated (i.e., it does not indicate searching for bread or apples). In this case, the command type is a fuzzy command type.

[0028] When a user instruction is a vague instruction type, its task intent may be ambiguous; in this case, the user instruction is a vague task instruction. In one scenario, the user instruction expresses a need or state but lacks a specific object or action; in this case, the task intent is vague. For example, the user instruction "I'm thirsty" expresses the state of "thirsty," but lacks a specific or clear object (e.g., does it want water or a beverage) and action (e.g., should it be given to the user or placed somewhere). In another scenario, the user instruction directly requests to find a certain type of item (i.e., without specifying a particular item); in this case, the task intent is vague. For example, the user instruction "Find me something delicious" requests to find food but does not specify what kind of food, therefore the instruction type is vague. User instructions that directly request to find a certain type of item can be considered active search instructions; that is, active search instructions are also a vague instruction type.

[0029] When a user instruction is an explicit instruction type, its task intent is clear. For example, if a user instruction contains a specific or concrete object and action, then its task intent is clear. For instance, a user instruction like "Bring me an apple" clearly indicates the instruction to find an apple and bring it to the user; "apple" is the operable object, and "bring" is the action. In this case, the user instruction's intent is clear (i.e., "bring the apple to the user"). Of course, an explicit instruction type can also be an instruction where the user asks the robot a question (i.e., a consultation instruction). In this case, the task intent of the consultation instruction is clear (i.e., answering the user's question). Furthermore, an explicit instruction type can be a casual conversation instruction where the user chats with the robot. For example, a user instruction like "Who are you?" has a clear task intent (i.e., chatting with the user). It's understandable that an explicit task intent refers to an intent that the robot can understand and execute.

[0030] For example, the robot can classify user instructions to determine their instruction type. Specifically, the robot can utilize natural language processing technology to extract key information from the user instructions. The instruction type is then determined based on this key information. Key information may include action verbs (such as "take," "put," or "find"), object entities (such as "apple" or "cup"), location information (such as "on the table"), and / or state description information (such as "hungry" or "messy").

[0031] In an exemplary embodiment, classifying user instructions to obtain the specific content of the instruction type may include: extracting key information from the user instructions; if the key information contains demand reflection information but lacks an operation object or an execution action for the operation object, then the instruction type of the user instruction is determined to be an ambiguous instruction type, and the demand reflection information is information that presents the demand; if the key information contains an operation object and an execution action for the operation object, then the instruction type of the user instruction is determined to be an explicit instruction type.

[0032] For example, when the key information representation task intent is to answer a user's inquiry or engage in casual conversation, the instruction type of the user instruction is determined to be an explicit instruction type, in which case the user instruction is a social dialogue instruction.

[0033] In this embodiment, by clarifying the instruction type through key information, the "initial screening" of user intent can be effectively completed, laying a clear and reliable interactive foundation for task orchestration.

[0034] Step 204: If the instruction type is an ambiguous instruction type, output intent guidance suggestions for the user instruction.

[0035] The intent guidance suggestion is generated based on the instruction context information corresponding to the user instruction, and is used to guide the user to clarify the task intent. The instruction context information includes one or more of the user instruction, the context preceding the user instruction, and the context following the user instruction. For example, if the context preceding and following the user instruction are empty, only the user instruction may be included.

[0036] The instruction context information of a user instruction may include one or more of the following: the user instruction itself, the object preceding the user instruction, and the response content to the user instruction. Specifically, the instruction context information may include the user instruction (e.g., "I'm hungry"); and / or, the instruction context information may include the user instruction and the dialogue preceding the user instruction, such as a historical dialogue: "User asks: What food do you have at home? Robot answers: There are apples, ...", user instruction: "I'm hungry"; and / or, the instruction context information may include the user instruction and the response content to the user instruction; and / or, the instruction context information may include the user instruction, the dialogue preceding the user instruction, and the response content to the user instruction.

[0037] Based on perception results, robots can generate intent-guided suggestions. In one scenario, these suggestions can be item-oriented, for example, recommending specific food items available in the environment when the potential task intent is "eating." Item-oriented suggestions can be in the form of option-based questions, providing clear options for the user to choose from, reducing decision-making difficulty. For example, an item-oriented suggestion could be, "I see apples and milk in the fridge, and bread on the table. Which would you like?" In another scenario, these suggestions can be task-oriented, for example, proposing specific cleaning or tidying solutions when the potential task intent is "cleaning." Task-oriented suggestions can be in the form of affirmative guidance, using affirmative language to guide the user to confirm the suggestion. For example, a task-oriented suggestion could be, "There's some clutter on the floor that needs cleaning. Can I help you clean it up?" Intent-guided suggestions can also employ emotional expression to maintain a warm or caring tone, enhancing the user experience. Robots can output intent-guided suggestions to users via a display screen or voice.

[0038] For example, the robot can output intent guidance suggestions in a target structured format, such as JSON format. The intent guidance suggestion can be: {"response": "I see apples and milk in the refrigerator, and bread on the table. Which one do you want?"}, and can provide feedback on the content of the intent guidance suggestion to the user through a display screen or voice.

[0039] For example, when the user command is of a fuzzy command type, the robot can generate a potential task intent based on the command context information, and then generate intent guidance suggestions based on the potential task intent and real-time environmental perception data. Since the user command is of a fuzzy command type, the task intent may be vague, i.e., without a clear task intent. Therefore, the robot can perform intent analysis on the user command to obtain the user's potential task intent. The potential task intent can be understood as the type of the user's potential needs, such as, but not limited to, needs like "eating," "cleaning," or "organizing." It is understood that the potential task intent also includes potential objects and potential actions. For example, when the potential task intent is of the "eating" type, the potential object could be "food" (such as apples, bread, etc.), and the potential action could be "give it to the user." The robot can then utilize its environmental perception skills to perform object recognition based on the potential objects in the potential task intent and real-time environmental perception data to obtain the perception results. The perception results include object information of potential operational objects identified from real-time environmental perception data, identifying the identified potential operational objects as associated operational objects, and identifying the object information of the potential operational objects as associated object information of the associated operational objects. For example, when the potential operational object can be "food," the robot can perform object recognition on the real-time environmental perception data. When it identifies food such as apples, bananas, and bread from the real-time environmental perception data, it will identify the identified apples, bananas, bread, etc., as associated operational objects.

[0040] For example, intent guidance suggestions can be generated based on instruction context information and real-time environmental perception data of the robot, or based on instruction context information, real-time environmental perception data of the robot, and user instructions. Real-time environmental perception data includes at least one of global environmental perception data or local environmental perception data. Global environmental perception data can be obtained through a vision module in the robot, which may include environmental perception skills. Global environmental perception data can be image data acquired from the environment by invoking environmental perception skills. Local environmental perception data are images captured in real-time using camera components. It can be understood that global environmental perception data can refer to three-dimensional image data of the target environment in which the robot is located, acquired through environmental perception skills, while local environmental perception data is two-dimensional image data acquired using camera components.

[0041] Step 206: Based on the intent guidance suggestion, the instruction context information is reconstructed to obtain the reconstructed context information.

[0042] For example, the robot can perform a clearing operation on historical context information, which is the instruction context information about the user command before generating intent guidance suggestions. Reconstructed context information is then generated based on the intent guidance suggestions and feedback instructions. The historical context information may include the text corresponding to the user command and / or the user's response to the user command. The reconstructed context information includes the intent guidance suggestions and feedback instructions.

[0043] For example, after generating intent guidance suggestions, after outputting intent guidance suggestions to the user, or after receiving feedback instructions, a clearing operation is performed on historical context information. Reconstructed context information is generated based on the intent guidance suggestions and feedback instructions. Based on the reconstructed context information, the feedback instructions are parsed to obtain the feedback task intent corresponding to the feedback instructions.

[0044] Step 208: Receive feedback instructions from the user regarding the intent guidance suggestion, and perform semantic analysis on the feedback instructions based on the reconstructed context information to obtain the feedback task intent.

[0045] Feedback instructions are the user's responses to the intention-guided suggestions. The task intent corresponding to the feedback instruction can be generated by the robot through semantic analysis (such as deep semantic analysis) of the feedback instruction.

[0046] The feedback task intent can be any of the following: "Confirm Acceptance," "Select Preference," "Reject / Cancel," and "New Request." "Confirm Acceptance" means the user explicitly agrees to the intention guidance suggestion. For example, the intention guidance suggestion could be "Do you need me to get you an apple?", guiding the user to define the task intent as "get an apple for the user." When the user's feedback task intent is "Confirm Acceptance," such as "Okay" or "I want an apple," the user explicitly agrees to the intention guidance suggestion. "Select Preference" means the user makes a specific choice from multiple options provided by the intention guidance suggestion. For example, the intention guidance suggestion is "I see apples and milk in the refrigerator, and bread on the table. Which one would you like?", and the feedback instruction is "I want an apple." "Reject / Cancel" means the user rejects the intention guidance suggestion or cancels the task, for example, the feedback instruction is "No, thank you" or "Cancel." "New Request" means the user makes a completely new request that is different from the intention guidance suggestion. For example, the intention guidance suggestion is "Do you need me to get you an apple?", but the feedback instruction is "No, I want water." In this way, by providing intention-guided suggestions, we can improve our understanding of user instructions with ambiguous command types, and transform vague user needs (such as "I am hungry") into specific and executable operation instructions (such as "I want an apple").

[0047] Step 210: When the feedback task intent indicates that the target task should be executed, a target action sequence is generated based on the target task, and the embodied robot is controlled to execute the target action sequence.

[0048] The target action sequence includes one or more action elements corresponding to each action; "multiple" means at least two. The action elements corresponding to each action in the target action sequence are arranged according to the execution order of the actions. An action element may contain an action ID, action description information, and action implementation instructions. The action description information describes the action. The action implementation instructions are used to implement the action; for example, they can be skill function call instructions, i.e., basic skill units, which may include, but are not limited to, `robot.search` (search for skills) and `robot.pick` (pick up skills). The action IDs of the action elements can be sequentially incremented. The action description information can be information conforming to natural language rules. The target task includes a specific operation object and a task that performs actions on that operation object.

[0049] The target action sequence can be in a structured format, such as a JSON array, where each element of the JSON array is an action element. For example, the target action sequence could be:

[0050] {"stage1": "Find dirty clothes", "function1": "robot.search('dirty clothes')"},

[0051] {"stage2": "Navigate to Dirty Clothes", "function2": "robot.goto('Dirty Clothes')"},

[0052] {"stage3": "Pick up dirty clothes", "function3": "robot.pick('dirty clothes')"},

[0053] {"stage4": "Find the washing machine", "function4": "robot.search('washing machine')"},

[0054] {"stage5": "Navigate to the washing machine", "function5": "robot.goto('washing machine')"},

[0055] {"stage6": "Open the washing machine", "function6": "robot.open('washing machine')"},

[0056] {"stage7": "Put dirty clothes into the washing machine", "function7": "robot.place('washing machine')"},

[0057] {"stage8": "Close the washing machine", "function8": "robot.close('washing machine')"},

[0058] {"stage9": "Start the washing machine", "function9": "robot.activate('washing machine')"}

[0059] Each {} and its contents constitute an action element, meaning each line in the target action sequence is an action element, for a total of 9 action sequences. `stagei` is the action ID, where 1 ≤ i ≤ 9, indicating that the action IDs are incrementing. The action ID and action description form key-value pairs, for example, the key-value pair "stage1": "Find dirty clothes", where "Find dirty clothes" is the action description. Action elements can also contain instruction IDs. The action implementation instruction can carry parameters. The instruction ID and the action implementation instruction form key-value pairs, for example, the key-value pair "function1": "robot.search('dirty clothes')", where "function1" is the instruction ID, and "robot.search('dirty clothes')" is the action implementation instruction carrying the parameter "dirty clothes".

[0060] For example, when the feedback instruction corresponds to a feedback task intent indicating the execution of a target task, the robot can generate an initial action sequence based on the target task. This initial action sequence contains action elements corresponding to actions typically required to achieve the target task. For instance, when the target task is "washing clothes," the initial action sequence can include a sequence of actions typically required for "washing clothes," such as: "stage1": "find dirty clothes"; "stage2": "navigate to dirty clothes"; "stage3": "pick up dirty clothes"; "stage4": "find the washing machine"; "stage5": "navigate to the washing machine"; "stage6": "turn on the washing machine"; "stage7": "put the dirty clothes into the washing machine"; "stage8": "turn off the washing machine"; "stage9": "start the washing machine." Further, the robot can use the initial action sequence as the target action sequence to achieve the target task. Alternatively, the robot can remove action elements corresponding to omissionable actions from the initial action sequence, and use the resulting initial action sequence as the target action sequence. Omitted actions can be, but are not limited to, actions that do not need to be performed. For example, if the washing machine door is open, the action "stage 6": "open the washing machine" can be omitted, meaning there's no need to open the washing machine door. This avoids logical inconsistencies, redundant steps, or violations of physical constraints when executing the target task, ensuring the reliability and security of task execution.

[0061] In the aforementioned context-adaptive embodied robot task processing method, the reliability of intent guidance suggestions is improved because they are generated based on the instruction context information corresponding to the user's command. Intent guidance suggestions guide the user to clarify the task intent, transforming vague user needs into clear and executable target tasks, thus clarifying the target task the user wants to perform (i.e., the correct one). Based on the intent guidance suggestions, the instruction context information is reconstructed to obtain reconstructed context information. Semantic analysis is then performed on the feedback command based on this reconstructed context information to obtain the feedback task intent. This ensures an accurate association between the feedback task intent and the user's needs, avoiding interference from historical context information. The target action sequence is only executed when the user's feedback task intent indicates the execution of the target task. This ensures the executed task is correct (i.e., what the user wants to perform), avoiding task execution errors and preventing situations where user commands based on vague command types cannot execute tasks, thereby improving task processing efficiency.

[0062] In an exemplary embodiment, the command context information is reconstructed based on the intent guidance suggestion to obtain the specific content of the reconstructed context information. This may include: clearing all content in the command context information to obtain the cleared command context information; adding the intent guidance suggestion to the cleared command context information to obtain the reconstructed context information; and the intent guidance suggestion serving as the dialogue starting point in the reconstructed context information.

[0063] For example, reconstructed context information can be generated based on intent guidance suggestions and feedback instructions. For instance, intent guidance suggestions and feedback instructions can be added to the cleared instruction context information to obtain reconstructed context information. For example, the step of clearing all content in the instruction context information can be performed after generating intent guidance suggestions, after outputting intent guidance suggestions to the user, or after receiving feedback instructions. In this embodiment, since the task intent of the user instruction is ambiguous, clearing the instruction context information before the intent guidance suggestion and reconstructing the context information can reduce the interference of the instruction context information before the intent guidance suggestion on intent parsing and improve the accuracy of the feedback task intent. At the same time, it can also improve the robot's ability to perceive and utilize context, enabling intelligent maintenance and utilization of historical context information (such as instruction context information) in multi-turn dialogues with the user, thereby improving the accuracy of understanding user instructions or feedback instructions.

[0064] In an exemplary embodiment, when the instruction type is an fuzzy instruction type, the specific content of the intention guidance suggestion output for the user instruction may include: when the instruction type is an fuzzy instruction type, performing a context saving operation for the user instruction to obtain the instruction context information of the user instruction; parsing the instruction context information to obtain the instruction parsing result of the user instruction; if the instruction parsing result includes state description information, determining the potential task intention corresponding to the user instruction based on the state description information, the potential task intention being used to change the state presented by the state description information; generating an intention guidance suggestion for the user instruction based on the potential task intention, and outputting the intention guidance suggestion.

[0065] Among these, the state description information reflects the state of the object described in the user's instruction. This can be information reflecting the user's current state or the current state of things in the environment. Things in the environment can be, but are not limited to, other people, animals, or inanimate objects. For example, the state description information could be "I'm hungry" or "The room is messy." The potential task intent can be understood as the type of potential need, such as, but not limited to, needs like "eating," "cleaning," or "tidying up."

[0066] For example, a robot can determine the intent to change the state presented by the state description information based on the state description information, thereby obtaining the potential task intent. For instance, if the state description information is "I am hungry", the potential task intent is "eating"; if the state description information is "the room is messy", the potential task intent is "cleaning" or "tidying up".

[0067] In an exemplary embodiment, the specific content of generating intent guidance suggestions for user instructions based on potential task intents may include: performing object recognition on real-time environmental perception data based on potential task intents to obtain perception results, the perception results including associated object information of associated operation objects identified from the real-time environmental perception data; the associated operation objects are used to realize the potential task intents; generating intent guidance suggestions based on the perception results; the intent guidance suggestions include associated object information, and the intent guidance suggestions are used to guide the user to clarify the task intent based on the associated object information.

[0068] The real-time environmental perception data includes at least one of global environmental perception data or local environmental perception data. The number of associated operation objects can be one or more. Associated object information refers to the information of the associated operation objects, which may include object identification information (e.g., object name, object ID, etc.) and associated object characteristics. These characteristics include one or more of the following: object location information, object status information, and object acquisition difficulty information. Object status information reflects the current state of the associated operation object. The object status information may differ at different times. For example, when the associated operation object is a washing machine, at the first moment, the washing machine's object status information is "door open," and at the second moment, the washing machine's object status information is "door closed." Different types of associated operation objects may have different object status information. For example, object status information may be the temperature and / or freshness of the associated operation object, the opening / closing status of the associated operation object, or the placement status of the associated operation object. Object acquisition difficulty information reflects the acquisition difficulty of the associated operation object. When the acquisition difficulty of the associated operation object is greater than a difficulty threshold, it can be determined that the associated operation object is not accessible; when the acquisition difficulty of the associated operation object is less than or equal to the difficulty threshold, it can be determined that the associated operation object is accessible.

[0069] For example, if the type of potential task intention is "eating", the perceived result could be: finding "apples" and "milk" in the refrigerator, and finding "bread" and "biscuits" on the table; as another example, if the type of potential task intention is "cleaning", the perceived result could be: finding "clutter" and "dust" on the ground.

[0070] For example, a robot can determine the potential type of a potential operational object based on its potential task intent. The potential type is the type to which the associated operational object to be identified belongs. For instance, if the potential task intent of a potential operational object is "eating," then the potential type is "food." The robot can identify associated operational objects with the potential type "food" from real-time environmental perception data and obtain the perception result.

[0071] For example, when there are multiple associated operation objects, the robot can generate intent guidance suggestions based on the associated object information corresponding to each of the multiple associated operation objects, and the intent guidance suggestions are used to guide the user to select the required associated operation object from the multiple associated operation objects to clarify the task intent.

[0072] In an exemplary embodiment, the associated object information includes associated object features and object identification information. When there are multiple associated operation objects, the recommendation priority of each associated operation object is determined based on its features and the user's interest features. Based on these recommendation priorities, M associated operation objects (where M is a positive integer) are selected for recommendation to the user. Intent guidance suggestions are generated based on the identification information and recommendation priorities of the M associated operation objects. The order of the identification information in the intent guidance suggestions matches the recommendation priorities of the M associated operation objects. Thus, when there are multiple associated operation objects, the associated operation objects are selected for recommendation according to their recommendation priorities. The order of the object identification information in the intent guidance suggestions matches the recommendation priorities of the associated operation objects. Since the recommendation priorities are determined based on object status information and user interest features, objects with better status and greater user interest are prioritized for recommendation. This improves recommendation accuracy and increases the likelihood that the user will clearly understand the task intent based on the intent guidance suggestions, thereby improving task execution efficiency and success rate.

[0073] User interest features are characteristics that reflect user interests, specifically the degree of interest in related operation objects. User interest features can be determined based on user-inputted interest descriptions, which may refer to the user's hobbies and interests. Alternatively, they can be determined based on the user's historical interaction behavior, such as historical intent-guided suggestions delivered to the user over a historical time period and / or the user's responses to those suggestions. Specifically, based on the user's responses to historical intent-guided suggestions, the operation objects selected by the user and those not selected are identified from the historical suggestions. The more times a user selects a particular operation object, the higher their level of interest in that object; conversely, the fewer times a user selects an operation object, the lower their level of interest. Currently, user interest features can also be determined jointly based on user-inputted interest descriptions and the user's historical interaction behavior.

[0074] The associated object features include one or more of the following: object location information, object status information, and / or object acquisition difficulty information. The object status information of the associated operation object can include a status representation value, which reflects the quality of the associated operation object's status. For example, a larger status representation value indicates a better status, and a smaller value indicates a worse status. The robot can determine the recommendation priority of multiple associated operation objects based on their respective associated object features and the user's interest features. Recommendation priority can be positively correlated with the user's level of interest; the higher the user's interest in a particular associated operation object, the higher its recommendation priority. Alternatively, recommendation priority can be negatively correlated with the acquisition difficulty information of the associated operation object; the higher the acquisition difficulty, the lower the recommendation priority. Recommendation priority can also be positively correlated with the status representation value of the associated operation object; a larger status representation value indicates a better status, and thus a higher recommendation priority. Here, M is a positive integer, such as 1, 2, 3, ...

[0075] The higher the recommendation priority of the associated operation object, the higher its object identification information appears in the intent guidance suggestion. For example, if the potential task intent is "food," and the associated operation objects are "apple," "milk," and "bread," and if the recommendation priority of "apple" is higher than that of "milk," and the recommendation priority of "milk" is higher than that of "bread," then the intent guidance suggestion would be: "I see apples and milk in the refrigerator, and bread on the table. Which one would you like?"

[0076] For example, when the number n of multiple associated operation objects is greater than N, N associated operation objects are selected from these multiple associated operation objects in descending order of recommendation priority as associated operation objects to be recommended to the user. In this case, M=N. The value of N can be set according to actual needs, and N≥2.

[0077] For example, when the number of associated operation objects n is equal to or less than N, the n associated operation objects are used as associated operation objects to recommend to the user, in which case M=n.

[0078] In this embodiment, state description information is determined based on the instruction context information of the user's instruction, potential task intent is determined based on the state description information, and then intent guidance suggestions are generated based on the user's potential task intent, which can improve the accuracy of intent guidance suggestions.

[0079] In an exemplary embodiment, the specific content after parsing the instruction context information to obtain the instruction parsing result of the user instruction may include: generating search response content when the instruction parsing result indicates that an object of the target type is to be searched; the search response content is used to indicate the processing flow of entering the object of the target type; invoking environmental perception skills to obtain global environmental perception data of the current environment of the embodied robot; searching out object information of objects of the target type from the global environmental perception data as associated object information of the associated operation object; and generating intent guidance suggestions for the user instruction based on the associated object information of the associated operation object.

[0080] The phrase "user instruction specifies the type of object to search" indicates that the user instruction is an active search instruction. Environmental perception skills are only invoked when the user instruction is an active search instruction, allowing the robot to acquire global environmental perception data of its current environment. This means that environmental perception skills are only invoked when necessary to obtain surrounding images, thus conserving bandwidth resources.

[0081] Among them, associated object information refers to the object information of the associated operation object, which may include the associated object characteristics and object identification information (such as name, ID). The associated object characteristics may include one or more of the following: object location information, object status information, and acquisition difficulty information.

[0082] For example, when the user command is a fuzzy command type and is an active search command, the user command is parsed to obtain the target type of the search object indicated by the user command. For example, if the user command is "Find me something delicious," then the target type of the search object is "food." With the target type determined, object information of objects with the target type is searched from the global environment awareness data and used as the associated object information for the associated operation object.

[0083] For example, the robot can generate skill invocation instructions, which include a skill identifier and input parameters. For instance, the robot can generate an environment invocation instruction to invoke an environmental awareness skill. This instruction includes the skill identifier and output parameters of the environmental awareness skill, with input parameters including the target type of the search object parsed from the user instruction. The robot can execute the environment invocation instruction to obtain global environmental awareness data. Of course, skill invocation instructions can also be `robot.goto` (i.e., navigation skill), `robot.pick` (i.e., pickup skill), etc. The robot uses a standard instruction format to generate all skill invocation instructions and strictly outputs skill invocation instructions with a standard instruction format, ensuring seamless integration with subsequent execution or system components and improving applicability.

[0084] For example, the search response can be a direct and definitive reply, indicating that the task has been accepted. The robot can provide the search response to the user via a display screen or voice. The robot can generate the search response and skill invocation instructions using a structured format, which can be, but is not limited to, JSON. For example, the search response and skill invocation instructions can be generated in the following ways:

[0085] {"response": "Of course, let me help you find some good food."}

[0086] {"function": "robot.search_VQA(1, 'food')"}.

[0087] In this context, `{"response": "Of course, let me help you find some good food"}` represents the search response as "Of course, let me help you find some good food". `{"function": "robot.search_VQA(1, 'food')"}` represents the skill invocation information. "robot.search_VQA" is the environmental awareness skill, "food" is the search target, and `robot.search_VQA(1, 'food')` is the environmental invocation command.

[0088] For example, when the user command is a vague command type and not an actively searched command, the robot can determine the potential task intent corresponding to the user command. It can also identify the associated operation objects used to achieve the potential task intent from real-time environmental perception data, obtain the associated object information of the associated operation objects, and generate intent guidance suggestions based on the associated object information. In addition, the robot can also output helpful responses to the user command. These responses can be guiding or caring, proactively offering assistance to clarify the user's needs. For example, a helpful response could be: {"response": "Are you thirsty? Do you need me to find you something to drink?"}

[0089] In this embodiment, when the user instructs to search for an object of the target type, the global environmental perception data of the current environment of the embodied robot is obtained through environmental perception skills. The associated object information is obtained from the global environmental perception data, and the intention guidance suggestion is generated based on the associated object information, which can improve the accuracy of the intention guidance suggestion.

[0090] In an exemplary embodiment, after obtaining the feedback task intent by performing semantic analysis on the feedback instruction based on the reconstructed context information, the method further includes: if the feedback task intent indicates rejection of the intent guidance suggestion or cancellation of the potential task intent corresponding to the user instruction, generating a task cancellation instruction and cancellation response content with a standard instruction format; based on the task cancellation instruction, clearing the reconstructed context information and ending the user instruction processing flow; the cancellation response content is used to remind users that the user instruction processing flow has ended; if the feedback task intent indicates that the potential task intent corresponding to the user instruction has changed, generating a task change instruction and change response content with a standard instruction format; based on the task change instruction, clearing the reconstructed context information and entering the feedback instruction processing flow; the change response content is used to remind users to enter the feedback instruction processing flow.

[0091] The phrase "after obtaining the feedback task intent by performing semantic analysis on the feedback instruction based on the reconstructed context information" means any time after the timestamp of obtaining the feedback task intent. Changing the response content serves as a notification that the processing flow for the user instruction has ended and the processing flow for the feedback instruction has begun.

[0092] The cancellation message is used to indicate that the processing of a user's instruction has ended. It can be an understanding response to show respect for the user's decision. The task cancellation instruction is the command to cancel the task. For example, the following content can be generated:

[0093] {"response": "Okay, call me again if you need anything"},

[0094] {"uncommand": "User cancels task"}.

[0095] In this context, "Okay, call me again if needed" is the message to cancel the reply, and "User cancels task" is the command to cancel the task.

[0096] For example, the robot can end the processing flow of the user's instruction based on the task cancellation instruction, and provide the cancellation response to the user through a display screen or voice, and enter the waiting instruction stage.

[0097] For example, when the feedback task intent indicates that the potential task intent has changed, such as when the feedback task intent is "a new requirement has been proposed", a change response and task change instruction can be generated. For example, the following content can be generated:

[0098] {"response": "I want some water. Let me see where there is water."}

[0099] {"uncommand": "User submits new request: Drink water"}.

[0100] Among them, "I want to drink water, let me see where there is water" is the modified response content. "User puts forward a new request: drink water" is the task modification instruction.

[0101] In this embodiment, by canceling or changing the response content, users can clearly understand the current stage of the robot, thereby improving their understanding of the task execution process.

[0102] In an exemplary embodiment, after controlling the embodied robot to execute the target action sequence, the method further includes any one of the following: if the execution result corresponding to the target action sequence is execution failure or the system returns an error code, then store the current instruction context information; if the execution result is execution success, then clear the current instruction context information; if a task termination instruction is received from the user during the execution of the target action sequence, then stop executing the target action sequence and clear the current instruction context information.

[0103] The current instruction context information is obtained by updating the context based on the reconstructed context information. For example, at least one of the following can be added to the reconstructed context information: feedback instructions, instructions generated by the robot after generating the reconstructed context information, or content fed back to the user, and execution records of the target action sequence, in order to obtain the current instruction context information through context update.

[0104] In this embodiment, dynamic adjustment of the current instruction context information is realized, achieving intelligent and adaptive context management.

[0105] In an exemplary embodiment, when the feedback task intent indicates that a target task should be executed, the specific content of generating a target action sequence based on the target task may include: generating a task execution instruction with a standard instruction format based on the target task; performing semantic parsing on the task execution instruction to obtain a parsing result, the parsing result including object identification information corresponding to multiple task operation objects associated with the target task, and object association relationships between the multiple task operation objects; generating an initial action sequence corresponding to the target task based on the object identification information and object association relationships corresponding to the multiple task operation objects, the initial action sequence including action elements corresponding to the actions required to achieve the target task; and removing action elements corresponding to actions that can be omitted from the initial action sequence to obtain the target action sequence.

[0106] The task operation object is used to achieve the target task. Object identification information includes the object's ID, name, and other information. The initial action sequence is an executable action sequence, which can be a sequence of basic skill units dynamically constructed by calling a predefined task strategy library.

[0107] Object relationships include, but are not limited to, spatial relationships, physical and mechanical relationships, or functional and usage relationships. Spatial relationships include, for example, "a cup next to a plate," "a knife on a cutting board," "a kettle in the center of an induction cooker," or "a refrigerator in a corner of the kitchen." Physical and mechanical relationships include, for example, "a plate supporting an apple," "a hook hanging a spatula," "a kettle lid attached to a kettle," or "a drawer embedded in a cabinet." Functional and usage relationships include, for example, "a knife used to cut vegetables on a cutting board," "a kettle used to pour water into a cup," or "a dishcloth used to clean a countertop."

[0108] For example, action elements corresponding to omissionable actions are removed from the initial action sequence to obtain an adjusted action sequence. The adjusted action sequence is then subjected to a rationality check, and the rationally validated adjusted action sequence is determined as the target action sequence. Specifically, based on real-time environmental perception data, object feature information of the task operation object associated with the target task is determined. Based on the object feature information of the task operation object, action elements corresponding to omissionable actions are removed from the initial action sequence to obtain the adjusted action sequence.

[0109] The object feature information may include at least one of state information and position information, and the object feature information is extracted from real-time environmental perception data. The state information may include at least one of existence state, open / closed state, or spatial pose. For container-type task operation objects, such as boxes or refrigerators, the spatial pose in the state information may include open / closed state (e.g., open or closed state), and the object feature information may also include container type (e.g., open or closed).

[0110] Reasonableness testing can include one or more of the following: goal consistency verification, repeatability verification, logical verification, and executability verification. Goal consistency verification verifies whether the task objective of the adjusted action sequence is consistent with the task objective of the target task; repeatability verification verifies whether there are duplicate action elements in the adjusted action sequence, which can be action elements with consistent objectives; executability verification verifies the executability of the action elements in the adjusted action sequence; logical verification verifies whether the execution logic of the adjusted action sequence is reasonable. If the adjusted action sequence satisfies the following conditions: the task objective of the adjusted action sequence is consistent with the task objective of the target task, there are no duplicate action elements in the adjusted action sequence, and the execution logic of the adjusted action sequence is reasonable, then the adjusted action sequence is determined to be reasonable.

[0111] Reasonableness checks are used to ensure the logical integrity of the target action sequence and determine whether the target action sequence conforms to the physical operation logic of "search-pick-place". For task operation objects that need to be closed or stopped (such as containers or devices), the corresponding action sequence for the closing action can be added to the adjustment action sequence. The adjustment action sequence after adding the action elements corresponding to the closing action is used as the target action sequence.

[0112] For example, there may be no task operation object associated with the target task in the environment. Therefore, the robot needs to determine whether there is a task operation object in the environment based on real-time environmental perception data. If there is, the robot adjusts the initial action sequence to obtain the adjusted action sequence based on the object feature information of the task operation object.

[0113] For example, based on the location information of the task operation object and the location information of the embodied robot, range indication information is determined to reflect whether the task operation object is within the operating range of the embodied robot. If the range indication information indicates that the task operation object is within the operating range, action elements corresponding to search and navigation actions targeting the task operation object are removed from the initial action sequence to obtain a revised initial action sequence. From the revised initial action sequence, action elements corresponding to actions that contradict the state information of the task operation object are removed to obtain an adjusted action sequence. In this way, by removing action elements corresponding to search and navigation actions targeting the task operation object, and removing action elements corresponding to actions that contradict the state information of the task operation object, unnecessary action elements can be eliminated, thereby reducing redundancy in the target action sequence and improving task execution efficiency.

[0114] For example, if the object to be operated on is already within the robot's operating range, the action elements corresponding to its "search" and "navigation" actions are omitted.

[0115] For example, the specific content of removing action elements corresponding to actions that contradict the state information of the task operation object may include: when the task operation objects associated with the target task are a first object and a second object, the target task instructs the second object to process the first object, and the second object is a container, the opening / closing state of the second object is determined based on its state information. If the opening / closing state is open, the action elements corresponding to the opening action for the second object are removed. For example, if the container is already in an open state, the "open" step is omitted. For example, if the first object is "clothes" and the second object is "washing machine," and the opening / closing state of the "washing machine" is open, then the action elements corresponding to the opening action for the "washing machine" can be removed.

[0116] In this embodiment, removing action elements corresponding to omissionable actions from the initial action sequence can eliminate redundant actions, prevent the robot from performing unnecessary actions, and determine the reasonable adjustment action sequence as the target action sequence, thus ensuring the reasonableness of the target action sequence.

[0117] In an exemplary embodiment, removing action elements corresponding to omissionable actions from the initial action sequence to obtain the specific content of the target action sequence may include: calling a camera component to obtain local environmental perception data of the embodied robot; if global environmental perception data of the embodied robot exists, and the timestamp of the global environmental perception data is within a preset time period before the current timestamp, then based on the local environmental perception data and the global environmental perception data, determining the action elements corresponding to omissionable actions in the initial action sequence, removing the action elements corresponding to omissionable actions from the initial action sequence, and obtaining the specific content of the target action sequence.

[0118] Local environmental perception data can include images captured by camera components (e.g., images taken by a camera). Global environmental perception data is obtained by invoking environmental perception skills to capture surrounding environmental data.

[0119] For example, the "real-time environmental perception data" in the above-mentioned "determining the object feature information of the task operation object associated with the target task based on real-time environmental perception data" can include at least one of local environmental perception data and global environmental perception data. The duration of the preset time period can be set according to actual needs, such as 1 hour or 3 hours. The timestamp of the global environmental perception data represents the time when the global environmental perception data was acquired. The current timestamp represents the current time. For example, if global environmental perception data exists within 3 hours before the current time, then the action elements corresponding to the omissionable actions in the initial action sequence are determined by combining the local environmental perception data and the global environmental perception data.

[0120] For example, if there is no global environment perception data for the embodied robot, or if the timestamp of the global environment perception data is outside a preset time period before the current timestamp, then the action elements corresponding to the omissionable actions in the initial action sequence are determined based on the local environment perception data.

[0121] Of course, the robot can also determine the required level of environmental perception skills based on the task characteristics of the target task. If the required level is greater than or equal to the preset required level, the robot invokes environmental perception skills to obtain global environmental perception data and local environmental perception data through its camera components. The global and local environmental perception data are then used to determine the robot's real-time environmental perception data. Based on this real-time environmental perception data, the robot determines the action elements corresponding to omissionable actions in the initial action sequence. The task characteristics include one or more of the following: task importance, task execution difficulty, task urgency, the location of the target task's operational object, and dependencies on other objects (such as dependent equipment, data, personnel, etc.). For example, a higher task importance, greater task execution difficulty, greater task urgency, and / or the object being outside the robot's field of vision results in a higher required level. The preset required level can be specifically set based on the specific needs.

[0122] During the execution of a target action sequence, the robot can also remove omissions from the action sequence based on real-time environmental perception data. This real-time environmental perception data can include global environmental perception data and / or local environmental perception data. The global environmental perception data can be based on global environmental perception data acquired over a preset historical time period, or it can be global environmental perception data acquired in real-time by invoking environmental perception skills. This allows the target task to be transformed into an executable and reasonable sequence of actions, while also ensuring the rationality of the sequence during execution, thereby improving task execution efficiency and success rate.

[0123] For example, the initial task sequence or target action sequence can be generated using a task generation model, which can be an artificial intelligence-based model, including but not limited to a neural network model. Task generation prompts can be generated and input into the task generation model to obtain the initial task sequence or target action sequence. The task generation prompts may include, for example, the following:

[0124] "You are a wheeled, dual-armed humanoid robot responsible for providing services in a home environment. Based on the user's task and the current image information, break down the task into specific sub-tasks and output them in JSON format. Important requirements: - The user explicitly uses a single skill; only the single skill needs to be output in JSON format, with the highest priority. - Strictly adhere to the JSON array format; do not add JSON tags or other formats. - Item naming accuracy: Strictly use the specific item names mentioned in the user's instructions; do not use vague terms. - Combine with image information: Intelligently adjust task steps based on the status of items in the current image. - Only use the provided skills; using non-existent skills is strictly prohibited. - IDs must be sequentially incremented: 1, 2, 3, 4... Available skill functions: 1. robot.search(id, object): Search skill - Searches for the approximate location of an item; 2. robot.goto(id, object): Navigation skill - Moves to the location of an item; 3. robot.pick(id, object): Pick-up skill - Visually detects and precisely picks up an item; 4. robot.place(id, object): Placement skill - Visually detects and precisely places an item in the target location; 5. robot.open(id, object) 6. robot.close(id, object): Activate skill - Open boxes, lids, doors, etc.; 7. robot.activate(id, object): Activate skill - Activate home appliances; 8. robot.deactivate(id, object): Deactivate skill - Deactivate home appliances; 9. robot.clean(id): Clean skill - Clean surfaces with a cloth.

[0125] In this embodiment, by combining local environment perception data and global environment perception data, action elements corresponding to actions that can be omitted can be selected more accurately.

[0126] In an exemplary embodiment, after classifying user instructions to obtain instruction types, the method further includes: if the instruction type is a specific instruction type and the user instruction indicates casual conversation, controlling the embodied robot to output casual conversation response content in response to the user instruction; initiating a context saving operation for the user instruction, generating a casual conversation context based on the user instruction and the casual conversation response content, and storing the casual conversation context.

[0127] The chat context is used for the robot to engage in casual conversation with the user. Chat responses can be warm and / or playful, fostering a friendly human-robot relationship. After outputting the chat response, the robot waits for the next valid instruction. The robot can generate chat responses using a structured format, such as: {"response": "Hello! I am your home service robot. How can I help you?"}. The robot can provide feedback to the user via a display screen or voice.

[0128] In this embodiment, a chat context is generated based on user commands and chat replies, and the chat context is stored to facilitate smooth chatting with the user based on the chat context.

[0129] In an exemplary embodiment, the method further includes: controlling the embodied robot to execute a second action sequence when the user instruction is of an explicit instruction type and the user instruction indicates the execution of a second task; the second action sequence is generated based on the second task. The second task is a task that includes a specific operation object and an action to be performed on that operation object; the second task may be the same as the target task or may be different from the target task.

[0130] For example, when the user instruction is of an explicit instruction type and instructs the execution of a second task, the robot can generate a structured and unambiguous task execution instruction in a standard instruction format. This task execution instruction is used to instruct the execution of the second task. In response to this task execution instruction, the robot generates a second sequence of actions to implement the second task.

[0131] For example, when the user instruction is an explicit instruction type and instructs the execution of a second task, the robot can also generate an explicit response corresponding to the user instruction and provide this response to the user via a display screen or voice. This explicit response can be a concise and friendly confirmation to enhance the naturalness of the interaction.

[0132] For example, the robot can generate explicit response content and task execution instructions corresponding to the second task using a structured format. Taking the user instruction "Bring me an apple" and the structured format as JSON as an example, the robot can output the following:

[0133] {"response": "Okay, I'll get you an apple right away"},

[0134] {"command": "Give the apple to the user"}.

[0135] In this context, {"response": "Okay, I'll get you the apple right away"} represents a clear response that says "Okay, I'll get you the apple right away", and {"command": "Give the apple to the user"} represents the task execution instruction for the second task that is "Give the apple to the user".

[0136] In this embodiment, when the instruction type is explicit and the user instruction indicates the execution of the second task, the robot is controlled to execute the second action sequence, thereby directly executing the task and ensuring the efficiency of task execution.

[0137] In one exemplary embodiment, such as Figure 3 The diagram illustrates a flow chart of a context-adaptive embodied robot task processing method, including:

[0138] Step 302: Receive user instructions and classify the user instructions to obtain the instruction type.

[0139] Step 304: Determine whether the instruction type is an explicit instruction type, a vague instruction type, or a dialog instruction type.

[0140] Step 306: When the instruction type is an explicit instruction type and the user instruction indicates casual conversation, control the embodied robot to output casual conversation reply content in response to the user instruction, initiate the context saving operation for the user instruction, generate a casual conversation context based on the user instruction and the casual conversation reply content, and store the casual conversation context.

[0141] Step 308: If the user instruction is of a specified type and the user instruction indicates that the second task should be performed, a second action sequence is generated based on the second task, and the embodied robot is controlled to perform the second action sequence.

[0142] Step 310: When the instruction type is fuzzy, perform a context saving operation for the user instruction to obtain the instruction context information. Parse the instruction context information to obtain the instruction parsing result. If the instruction parsing result includes state description information, determine the potential task intent corresponding to the user instruction based on the state description information. Based on the potential task intent, perform object recognition on the real-time environmental perception data to obtain the associated object information of the associated operation object. Based on the associated object information of the associated operation object, generate an intent guidance suggestion for the user instruction, output the intent guidance suggestion, clear all content in the instruction context information to obtain the cleared instruction context information, and add the intent guidance suggestion to the cleared instruction context information to obtain the reconstructed context information.

[0143] Real-time environmental perception data includes at least one of local environmental perception data and global environmental perception data.

[0144] Step 312: Receive feedback instructions from the user regarding the intent guidance suggestion, and perform semantic analysis on the feedback instructions based on the reconstructed context information to obtain the feedback task intent.

[0145] Step 314: When the feedback task intent indicates that the target task should be executed, a target action sequence is generated based on the target task, and the embodied robot is controlled to execute the target action sequence.

[0146] Step 316: In the case of feedback task intent indication rejection intent guidance suggestion or cancellation of potential task intent corresponding to user instruction, generate task cancellation instruction with standard instruction format and cancellation response content, and based on task cancellation instruction, clear and rebuild context information and end user instruction processing flow.

[0147] Step 318: If the potential task intent corresponding to the user instruction has changed, the task change instruction and change response content with standard instruction format are generated. Based on the task change instruction, the context information is cleared and rebuilt, and the feedback instruction processing flow is entered.

[0148] The specific content of steps 302-318 can be referred to the specific content of steps 202-210 above, and will not be repeated here in the embodiments of this application.

[0149] In one exemplary embodiment, such as Figure 4 As shown, a flowchart of a context-adaptive embodied robot task processing method is presented, applied to an interactive task orchestration system in a robot. It comprises four stages: Stage 1 (user instruction analysis), Stage 2 (requirement clarification and suggestion generation), Stage 3 (task instruction generation), and Stage 4 (action sequence generation). Stage 1, as the main entry point of the interactive task orchestration system, is responsible for real-time parsing and classification of user instructions. Based on the parsing and analysis results, it initiates corresponding downstream processing flows to achieve instruction classification and process routing. Its core is a semantic and context-based multiplexer. Stages 2 and 3 constitute a "clarification loop" for handling ambiguous requirements. Stage 4, as the final output, transforms explicit instructions into executable actions and executes those actions. Specifically:

[0150] Phase 1 (User Command Analysis): If the user command is unrelated to the task, simply reply to the user, engage in casual conversation, and activate the context dialogue management mechanism to save the topic content and the complete context of the historical dialogue, then loop back to Phase 1; if the user command is a specific command type (i.e., a concrete or explicit task command), reply to the user and generate a task execution command with a standard command format, then directly proceed to Phase 4; if the user command is a vague command type (i.e., a vague task command), reply to the user, invoke environmental awareness skills, and activate the context dialogue management mechanism to save the topic content, then proceed to Phase 2.

[0151] Phase 2 (Needs Clarification and Suggestion Generation): Receives the dialogue context from Phase 1 and invokes environmental awareness skills; based on the dialogue context from Phase 1, processes real-time environmental awareness data to obtain perception results; after obtaining the perception results, combines the dialogue history to generate intent-guided suggestions; after outputting intent-guided suggestions but before waiting for user response, a context reconstruction operation can be performed: clear the previous dialogue history, save the generated intent-guided suggestions as the new dialogue starting point, thereby building a focused and concise context for Phase 3 (i.e., reconstructing the context), ensuring accurate association between perceived information and user needs, and avoiding interference from historical dialogue.

[0152] Phase 3 (Task Instruction Generation): Based on the simplified context built in Phase 2, the user's response to the intent guidance suggestion (i.e., feedback instruction) is parsed; if the user refuses, cancels, or makes a new request, the user is responded to and a task cancellation instruction uncommand is generated. This task cancellation instruction is used to trigger the clearing of the current context and return to Phase 1; if the user agrees to the intent guidance suggestion, the user is responded to and a task execution instruction command with a standard instruction format is generated, and Phase 4 is entered.

[0153] Phase 4 (Action Sequence Generation): Based on the task execution instructions with standard instruction formats and combined with real-time environmental awareness data, the action sequence is arranged and instructions are generated; after the task is successfully executed, all context related to the user instructions is cleared and the system returns to standby state.

[0154] The environmental perception skill is invoked only in Phase 1 and during active search to acquire surrounding environmental data (i.e., global environmental perception data). At other times, image data is acquired via a camera to obtain local environmental perception data. If global environmental perception data exists within a preset time period prior to the current time (e.g., the previous 3 hours), the robot's current real-time environmental perception data includes both local and global environmental perception data. The environmental perception skill is only invoked in Phase 1 and during active search requests; in other phases, the results of the environmental perception skill invocation are directly used.

[0155] Upon entering Phase 4, the system directly accesses camera-captured image data (i.e., local environmental perception data, not global environmental perception data obtained through pre-defined environmental perception skills). For the first three hours, global environmental perception data is available. This data is then combined with the local and global environmental perception data to generate the action sequence. Thus, environmental perception skills are only invoked when necessary to acquire surrounding images, aiding in decision-making. This ensures that the system is used only when needed, conserving bandwidth resources.

[0156] For example, throughout the task orchestration process, a time-window-based long-dialogue management mechanism is employed to maintain the state and track the history of multi-turn dialogues. The specific implementation is as follows:

[0157] 1. The context-based long dialogue management mechanism runs through the four core stages of task processing, with each stage employing a differentiated context processing strategy:

[0158] Phase 1: Employing a conditional context injection mechanism: When in pure chat mode, a complete context containing historical dialogues is constructed using the get_context_prompt() method; when in task execution mode, only the instruction context of the current user command is used for analysis to avoid interference from historical dialogues; the context usage strategy is dynamically switched using the in_pure_chat status flag.

[0159] Phase 2: Implementing a context reconstruction mechanism: After invoking the environment-aware skill, clear the original dialogue history using the clear_context() method; reconstruct a simplified context (i.e., rebuild the context), retaining only the perception results and subsequent user commands; ensure that the perception results are accurately associated with the user's needs, and avoid interference from historical dialogues.

[0160] Phase 3: Maintaining Contextual Integrity: Use get_history_for_vlm() to obtain the complete dialogue history; maintain the continuity between the guiding dialogue and user responses to ensure the accuracy of intent understanding; dynamically generate task execution instructions or task cancellation instructions with standard instruction formats based on user selection.

[0161] Phase 4: Execute context cleanup operations: After the task is successfully executed, clear_context() completely clears all dialogue history related to the current user command; resets the system state to prepare to receive new independent task commands; and handles the state transition after the task is completed uniformly through _handle_task_success.

[0162] 2. The following specific methods are used to achieve intelligent maintenance of context: (1) Structured storage of dialogue records: Each dialogue is recorded using a triplet structure, which includes role identifier, content entity and timestamp information; role identifier can include user 1 identifier, robot identifier, user 2 identifier, etc. (2) Dynamic pruning of capacity control: Automated maintenance of historical records is achieved to ensure that the storage capacity of the repository used to store context information is within the preset range, and context information in the repository can be removed or deleted periodically or quantitatively; (3) Intelligent selection of context use: The most recent relevant dialogue record is dynamically selected according to the configuration parameters (in_pure_chat status flag) to balance context depth and processing efficiency.

[0163] 3. State-aware context strategy switching, dynamic adjustment of context strategy through state awareness: (1) Pure chat mode (such as casual chat or asking questions): maintain the continuity of the complete dialogue history (i.e. historical context) and support multi-round natural language interaction; activate the context preservation function through set_pure_chat_mode(True); (2) Task execution mode: clear the context at key task nodes to avoid historical interference and focus on the accurate execution of the current task instructions; ensure the independence of task execution through clear_context().

[0164] 4. Context recovery mechanism under abnormal conditions, with complete exception handling capabilities: (1) Task failure recovery: After the TaskFailedException is caught, the context is maintained or cleared, and the dialogue is restarted according to the specific scenario. (2) System error handling: In general abnormal situations, the system recoverability is ensured by clear_context(). (3) User interruption response: Supports the recognition of natural language exit commands and properly handles context cleanup.

[0165] The context has consistency guarantee, specifically: (1) Requirement traceability: ensure that the generated task execution instructions and the original fuzzy requirements of the user instructions are logically consistent; (2) Item name standardization: uniformly use the accurate item names identified in the environment perception; (3) Instruction format standardization: strictly follow the standard output instructions of command / uncommand.

[0166] The context-based long dialogue management mechanism established in this application can intelligently maintain the coherence of dialogue in complex home service scenarios, while ensuring the cleanliness of the context during task execution, effectively improving the naturalness of human-computer interaction and the reliability of task execution. This mechanism achieves intelligent and adaptive context management through multiple technical means such as state awareness, dynamic adjustment, and anomaly recovery.

[0167] The multimodal perception-based embodied robot task processing method provided in this application uses a contextual dialogue management mechanism as a supporting module, which runs through all stages and is responsible for maintaining and switching dialogue states. Each stage communicates through standardized instructions (such as uncommand and command) and data formats (JSON) to jointly achieve intelligent conversion from user natural language input to ordered robot action output.

[0168] This application implements a multi-stage routing and context-adaptive interactive task orchestration architecture, specifically including: (1) a multi-path distribution mechanism based on intent classification: through accurate semantic decoding and intent recognition, user instructions are dynamically routed to different processing pipelines such as chat, direct execution, and environment-aware guidance, achieving efficient task diversion from the source. (2) a "perception-guidance-confirmation" fuzzy requirement clarification loop: creatively combining environment awareness (such as visual question answering) with dialogue guidance, constructing a closed-loop sub-process consisting of "stage 2" and "stage 3", actively guiding users to concretize fuzzy requirements. (3) a state-aware dynamic context management mechanism: a context-long dialogue management method that runs throughout the process is designed, which can intelligently determine the saving, reconstruction, and clearing strategies of the context according to the dialogue mode (chat / task) and task stage, maintaining both continuity and avoiding information interference. (4) Environmental state adaptive action sequence planning: In the final action arrangement, real-time environmental perception data is deeply integrated, and the preset task strategy is optimized based on dynamic information such as item location and container status. Redundant steps are intelligently omitted to generate the most efficient executable sequence. (5) System integration standardization: Communication between each stage is carried out through standardized and structured data interfaces (such as JSON format instructions), forming a loosely coupled and highly cohesive modular architecture. This not only enhances the maintainability and scalability of the system, but also greatly improves the reliability and security of the entire task execution process due to its clear responsibility boundaries and controllable process.

[0169] This application, through a multi-stage interactive task orchestration process of "analysis-guidance-confirmation," can gradually transform vague or abstract user needs into precise and executable task instructions, improving the robot's understanding accuracy in open environments. The contextual long-dialogue management mechanism, through state awareness and dynamic strategy switching, maintains the coherence and naturalness of dialogue during multiple rounds of casual conversation while clearing irrelevant history at key task execution nodes, avoiding information interference, thereby optimizing the interaction path, reducing unnecessary dialogue rounds, and making the interaction process more efficient. In the action orchestration stage, state-adaptive task sequence planning based on real-time environmental awareness can intelligently omit redundant steps (such as searching for nearby objects or opening already opened containers), generating better execution paths, improving task execution efficiency, and reducing mechanical wear and time costs. Data transmission between stages uses a strictly standardized JSON format, ensuring decoupling and efficient collaboration between system modules, enhancing system maintainability and scalability, and achieving standardized system integration.

[0170] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0171] In one exemplary embodiment, an embodied robot system is provided, which is used to: receive user instructions and classify the user instructions to obtain instruction types; when the instruction type is an ambiguous instruction type, output intent guidance suggestions for the user instructions; reconstruct the instruction context information based on the intent guidance suggestions to obtain reconstructed context information; receive user feedback instructions for the intent guidance suggestions and perform semantic analysis on the feedback instructions based on the reconstructed context information to obtain feedback task intent; when the feedback task intent indicates the execution of a target task, generate a target action sequence based on the target task and control the embodied robot to execute the target action sequence.

[0172] In one exemplary embodiment, a computer device is provided, the internal structure of which can be as shown in the figure. Figure 5As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores at least some of the data involved in the context-adaptive embodied robot task processing method. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a context-adaptive embodied robot task processing method.

[0173] In one exemplary embodiment, a computer device is provided, the internal structure of which can be as shown in the figure. Figure 6 As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When executed by the processor, the computer program implements a context-adaptive embodied robot task processing method. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0174] Those skilled in the art will understand that Figure 5 and Figure 6The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0175] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0176] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0177] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0178] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0179] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0180] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0181] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A context-adaptive embodied robot task processing method, characterized in that, The method includes: Receive user instructions and classify the user instructions to obtain the instruction type; When the instruction type is an ambiguous instruction type, the instruction context information of the user instruction is parsed. If the parsing result includes state description information, the potential task intent is determined based on the state description information. Based on the potential task intent, an intent guidance suggestion for the user instruction is output. The intent guidance suggestion is used to guide the user to clarify the task intent, and the potential task intent is used to change the state presented by the state description information. Receive user feedback instructions for the intent guidance suggestion, clear the instruction context information, generate reconstructed context information based on the intent guidance suggestion and the feedback instruction, perform semantic analysis on the feedback instruction based on the reconstructed context information to obtain the feedback task intent, and the intent guidance suggestion is the dialogue starting point in the reconstructed context information; When the feedback task intent indicates that a target task should be executed, a target action sequence is generated based on the target task, and the embodied robot is controlled to execute the target action sequence. If the execution result corresponding to the target action sequence is execution failure or the system returns an error code, the current instruction context information is stored. The current instruction context information is obtained by updating the reconstructed context information.

2. The method according to claim 1, characterized in that, The step of outputting intent guidance suggestions for the user instruction based on the potential task intent includes: Acquire real-time environmental perception data of the embodied robot; Based on the potential task intent and the real-time environmental awareness data, intent guidance suggestions are output for the user command.

3. The method according to claim 1, characterized in that, In the case where the instruction type is an ambiguous instruction type, the instruction context information of the user instruction is parsed. If the parsing result includes state description information, the potential task intent is determined based on the state description information, including: If the instruction type is an ambiguous instruction type, perform a context saving operation for the user instruction to obtain the instruction context information of the user instruction; The instruction context information is parsed to obtain the instruction parsing result of the user instruction. If the instruction parsing result includes state description information, the potential task intent corresponding to the user instruction is determined based on the state description information.

4. The method according to claim 3, characterized in that, After parsing the instruction context information to obtain the instruction parsing result of the user instruction, the method further includes: If the instruction parsing result indicates that an object of the target type is to be searched, a search response is generated; the search response is used to indicate the process of entering the search for the object of the target type. The environmental perception skill is invoked to obtain global environmental perception data of the current environment of the embodied robot. From the global environment awareness data, object information of objects with the target type is searched out and used as associated object information of the associated operation object. Based on the associated object information of the associated operation object, intent guidance suggestions for the user command are generated.

5. The method according to any one of claims 1-4, characterized in that, After performing semantic analysis on the feedback instruction based on the reconstructed context information to obtain the feedback task intent, the method further includes: If the feedback task intent indicates rejection of the intent guidance suggestion or cancellation of the potential task intent corresponding to the user instruction, a task cancellation instruction with a standard instruction format and a cancellation response content are generated. Based on the task cancellation instruction, the reconstruction context information is cleared and the processing flow of the user instruction is terminated. The cancellation response content is used to remind that the processing flow of the user instruction has been terminated. When the feedback task intent indicates that the potential task intent corresponding to the user instruction has changed, a task change instruction and change response content with the standard instruction format are generated. Based on the task change instruction, the reconstruction context information is cleared and the processing flow of the feedback instruction is entered. The change response content is used to remind users to enter the processing flow of the feedback instruction.

6. The method according to any one of claims 1-4, characterized in that, After controlling the android to execute the target action sequence, the method further includes any one of the following: If the execution result corresponding to the target action sequence is execution failure or the system returns an error code, then the current instruction context information is stored; If the execution result is successful, then clear the current instruction context information; If a task termination command is received from the user during the execution of the target action sequence, the execution of the target action sequence will be stopped and the current command context information will be cleared. The current instruction context information is obtained by updating the context based on the reconstructed context information.

7. The method according to any one of claims 1-4, characterized in that, When the feedback task intent indicates that a target task should be performed, generating a target action sequence based on the target task includes: When the feedback task intent indicates that a target task should be executed, a task execution instruction with a standard instruction format is generated based on the target task. The task execution instructions are semantically parsed to obtain the parsing results. The parsing results include object identification information corresponding to multiple task operation objects associated with the target task, as well as the object association relationships between the multiple task operation objects. Based on the object identification information corresponding to the multiple task operation objects and the object association relationship, an initial action sequence corresponding to the target task is generated. The initial action sequence contains action elements corresponding to the actions required to implement the target task. The target action sequence is obtained by removing action elements corresponding to omissionable actions from the initial action sequence.

8. The method according to claim 7, characterized in that, The step of removing action elements corresponding to omissionable actions from the initial action sequence to obtain the target action sequence includes: The camera component is invoked to obtain local environmental perception data of the embodied robot's current location; If global environment perception data of the embodied robot exists, and the timestamp of the acquisition of the global environment perception data is within a preset time period before the current timestamp, then the action elements corresponding to the omissionable actions in the initial action sequence are determined based on the local environment perception data and the global environment perception data. The target action sequence is obtained by removing the action elements corresponding to the omissionable actions from the initial action sequence.

9. The method according to any one of claims 1-4, characterized in that, After classifying the user instructions to obtain the instruction type, the method further includes: When the instruction type is an explicit instruction type, and the user instruction indicates casual conversation, the avatar robot is controlled to output casual conversation response content in response to the user instruction; Initiate a context saving operation for the user command, generate a chat context based on the user command and the chat reply content, and store the chat context; the chat context is used for the avatar robot to chat with the user.

10. The method according to any one of claims 1-4, characterized in that, The process of classifying the user instructions to obtain instruction types includes: Extract key information from the user instructions; If the key information includes demand reflection information, but lacks an operation object or an execution action for the operation object, then the instruction type of the user instruction is determined to be a vague instruction type, and the demand reflection information is information that presents the demand. If the key information includes the operation object and the action to be performed on the operation object, then the instruction type of the user instruction is determined to be an explicit instruction type.

11. A hymenoidae robot system, characterized in that, The embodied robotic system is used for: Receive user instructions and classify the user instructions to obtain the instruction type; When the instruction type is an ambiguous instruction type, the instruction context information of the user instruction is parsed. If the parsing result includes state description information, the potential task intent is determined based on the state description information, and the intent guidance suggestion for the user instruction is output based on the potential task intent. The intent guidance suggestion is used to guide the user to clarify the task intent, and the potential task intent is used to change the state presented by the state description information; Receive user feedback instructions for the intent guidance suggestion, clear the instruction context information, generate reconstructed context information based on the intent guidance suggestion and the feedback instruction, perform semantic analysis on the feedback instruction based on the reconstructed context information to obtain the feedback task intent, and the intent guidance suggestion is the dialogue starting point in the reconstructed context information; When the feedback task intent indicates that a target task should be executed, a target action sequence is generated based on the target task, and the embodied robot is controlled to execute the target action sequence. If the execution result corresponding to the target action sequence is execution failure or the system returns an error code, the current instruction context information is stored. The current instruction context information is obtained by updating the reconstructed context information.