Task execution method of intelligent agent with body
Through predictive action data analysis based on user voice and historical data, dynamically determine the target task in combination with current environmental data, and real-time detection during execution, the problem of inaccurate target tasks in the existing technology is solved, and the accuracy and robustness of the execution of embodied agent tasks is improved.
Patent Information
- Application Number
- CN202510119277.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-01-24
AI Technical Summary
In the prior art, the target tasks determined by the robot task execution model are not accurate enough, resulting in low robustness in the execution of embodied agent tasks.
By determining predicted action data based on the user's voice text data and/or historical execution data, determining multiple tasks to be selected based on the predicted action data, determining the target tasks in the to be selected based on the current environment data collected by the embodied agent, and real-time detection during the task execution process, improving the accuracy of task execution.
Dynamically determine the target task from multiple to-select tasks and detect it in real time during the task execution, which significantly improves the accuracy and robustness of the execution of embodied agent tasks.
Smart Images

Figure CN120134299A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of data processing, and in particular, to a method for task execution of an embodied intelligent agent. Background Art
[0002] With the development of artificial intelligence technology, embodied intelligent agents have also been continuously developed in terms of task execution.
[0003] Currently, the target task of an embodied intelligent agent can be determined through a robot task execution model, and the embodied intelligent agent can be controlled to execute the target task. Among them, the robot task execution model can be trained through training data sets in multiple scenarios, and the target task can be a task with long time series and high complexity.
[0004] In the above task execution process, the target task determined by the robot task execution model may not be accurate enough. Summary of the Invention
[0005] The embodiments of the present application provide a method for task execution of an embodied intelligent agent, which is used to solve the defect that the target task determined by the robot task execution model in the prior art is not accurate enough.
[0006] In a first aspect, the present application provides a method for task execution of an embodied intelligent agent, and the method includes:
[0007] Determine predicted action data according to the voice text data and / or historical execution data of the user, where the voice text data is obtained by analyzing the voice data collected by the embodied intelligent agent from the user, and the historical execution data is the data generated when executing the previous task;
[0008] Determine multiple candidate tasks according to the predicted action data;
[0009] Determine a target task from the multiple candidate tasks according to the current environmental data collected by the embodied intelligent agent;
[0010] Execute the target task and detect the execution process of the target task to obtain a task detection result.
[0011] In a possible implementation manner, executing the target task and detecting the execution process of the target task to obtain a task detection result includes:
[0012] Start to execute the target task;
[0013] Execute at least one detection operation until the target task is completed, or when the number of times of executing the target task is greater than or equal to a preset threshold, obtain a task detection result, and the task detection result includes successful task execution and failed task execution;
[0014] The detection operation includes: determining a current detection result according to the task execution information of the target task in the current detection cycle; if the current detection result is an abnormal detection, restarting the execution of the target task; if the current detection result is a normal detection, determining whether the target task is completed, and if not, performing the detection operation in the next detection cycle.
[0015] In a possible implementation manner, determining whether the target task is completed includes:
[0016] determining whether there is task execution information of at least one historical detection cycle, where the historical detection cycle is a detection cycle after starting to execute the target task and before the current detection cycle;
[0017] if so, determining multiple first execution images according to the task execution information of at least one historical detection cycle and the task execution information of the current detection cycle; performing multiple Fourier transform processes on the multiple first execution images to obtain a first frequency domain signal;
[0018] if not, determining multiple second execution images according to the task execution information of the current detection cycle; performing multiple Fourier transform processes on the multiple second execution images to obtain a second frequency domain signal;
[0019] inputting the first frequency domain signal or the second frequency domain signal into a first preset model to obtain an output result, where the output result is used to indicate that the target task is completed or indicate that the target task is not completed.
[0020] In a possible implementation manner, the method further includes:
[0021] acquiring multiple images and the text corresponding to each image;
[0022] performing multiple Fourier transform processes on each image to obtain a frequency domain signal corresponding to each image;
[0023] generating multiple multimodal training data according to the frequency domain signal corresponding to each image and the text corresponding to each image, where the multimodal training data includes a frequency domain signal and text;
[0024] acquiring a multimodal preset model, where the multimodal preset model includes a visual encoder, a language model, and a projector;
[0025] performing an adjustment operation on the multimodal preset model to obtain a multimodal intermediate model, where the adjustment operation includes at least one of the following: freezing the visual encoder, performing a freezing process on the language model, replacing the activation function in the projector, and adding a normalization layer operation;
[0026] Train the multimodal intermediate model based on the multiple multimodal training data to obtain the first preset model.
[0027] In one possible implementation, determining a target task from the multiple candidate tasks according to the current environmental data collected by the embodied agent includes:
[0028] Input the current environmental data into the first preset model to obtain current scenario data; input the current scenario data and the multiple candidate tasks into the second preset model to obtain the target task; or,
[0029] Input the current environmental data and the multiple candidate tasks into the first preset model to obtain the target task.
[0030] In one possible implementation, determining multiple candidate tasks according to the predicted action data includes:
[0031] Perform vector embedding processing on the predicted action data to obtain a first vector;
[0032] Determine the similarity between the first vector and each vector in the preset database, where the preset database includes multiple vectors and the task corresponding to each vector;
[0033] Sort the multiple vectors in the preset database in descending order of similarity to obtain a vector sequence;
[0034] Determine the first n vectors in the vector sequence as multiple candidate vectors, where n is a preset quantity;
[0035] Determine the multiple candidate tasks corresponding to the multiple candidate vectors respectively.
[0036] In one possible implementation, the determining a target task from the multiple candidate tasks according to the current environmental data collected by the embodied agent includes:
[0037] Identify the user's hand movement according to the current environmental data collected by the embodied agent;
[0038] When the user's hand movement meets the first preset condition, obtain the current time;
[0039] When the current time meets the preset time period, obtain the note information associated with the task execution object in the target task;
[0040] Determine the target task according to the note information.
[0041] In a possible implementation manner, determining a target task according to the matter information includes:
[0042] Identifying the user's action intention according to the user's hand actions;
[0043] Obtaining user preference information associated with the action intention;
[0044] Determining a target task according to the matter information and the user preference information.
[0045] In a possible implementation manner, determining a target task from the multiple candidate tasks according to the current environmental data collected by the embodied agent includes:
[0046] When the current time meets a preset time period, identifying the user's action intention according to the user's hand actions;
[0047] Obtaining the correct number of items corresponding to the action intention;
[0048] Obtaining the number of items triggered by the user's hand actions;
[0049] When the correct number of items does not match the triggered number of items, the target task further includes controlling the embodied agent to send a voice prompt message to the user, and the voice prompt message includes the correct number of items.
[0050] In a possible implementation manner, after executing the target task and detecting the execution process of the target task to obtain a task detection result, the method further includes:
[0051] If the task detection result is that the task execution fails, generating a task exception indication according to the target task, and the task exception indication is used to indicate not to execute the target task within a preset duration;
[0052] Updating the historical execution data according to the task exception indication and at least one task execution information corresponding to the target task.
[0053] In a second aspect, the present application provides a task execution device for an embodied agent, and the device includes:
[0054] A prediction module, configured to determine predicted action data according to the user's voice text data and / or historical execution data, where the voice text data is obtained by analyzing the voice data collected by the embodied agent from the user, and the historical execution data is data generated when executing the previous task;
[0055] A first determination module, configured to determine multiple candidate tasks according to the predicted action data;
[0056] A second determination module, configured to determine a target task from the multiple candidate tasks according to the current environmental data collected by the embodied intelligent agent;
[0057] An execution module, configured to execute the target task and detect the execution process of the target task to obtain a task detection result.
[0058] In a possible implementation manner, the execution module is specifically configured to:
[0059] Start to execute the target task;
[0060] Execute at least one detection operation until the target task is completed, or when the number of times of executing the target task is greater than or equal to a preset threshold, obtain a task detection result, where the task detection result includes successful task execution and failed task execution;
[0061] The detection operation includes: determining a current detection result according to the task execution information of the target task in the current detection period; if the current detection result is an abnormal detection, restart the execution of the target task; if the current detection result is a normal detection, determine whether the target task is completed, and if not, execute the detection operation in the next detection period.
[0062] In a possible implementation manner, the execution module is specifically configured to:
[0063] Determine whether there is task execution information of at least one historical detection period, where the historical detection period is a detection period after starting to execute the target task and before the current detection period;
[0064] If so, determine multiple first execution images according to the task execution information of at least one historical detection period and the task execution information of the current detection period; perform multiple Fourier transform processes on the multiple first execution images to obtain a first frequency domain signal;
[0065] If not, determine multiple second execution images according to the task execution information of the current detection period; perform multiple Fourier transform processes on the multiple second execution images to obtain a second frequency domain signal;
[0066] Input the first frequency domain signal or the second frequency domain signal into a first preset model to obtain an output result, where the output result is used to indicate that the target task is completed or indicate that the target task is not completed.
[0067] In a possible implementation manner, the device further includes a training module, and the training module is used for:
[0068] Obtain multiple images and the text corresponding to each image;
[0069] Perform multiple Fourier transform processes on each image to obtain the frequency-domain signals corresponding to each image;
[0070] Generate multiple multimodal training data according to the frequency-domain signals corresponding to each image and the text corresponding to each image, where the multimodal training data includes frequency-domain signals and text;
[0071] Obtain a multimodal preset model, where the multimodal preset model includes a visual encoder, a language model, and a projector;
[0072] Perform an adjustment operation on the multimodal preset model to obtain a multimodal intermediate model, and the adjustment operation includes at least one of the following: freezing the visual encoder, performing a freezing process on the language model, replacing the activation function in the projector, and adding a normalization layer operation;
[0073] Train the multimodal intermediate model according to the multiple multimodal training data to obtain the first preset model.
[0074] In a possible implementation manner, the second determination module is specifically configured to:
[0075] Input the current environmental data into the first preset model to obtain the current scene data; input the current scene data and the multiple candidate tasks into the second preset model to obtain the target task; or,
[0076] Input the current environmental data and the multiple candidate tasks into the first preset model to obtain the target task.
[0077] In a possible implementation manner, the first determination module is specifically configured to:
[0078] Perform vector embedding processing on the predicted action data to obtain a first vector;
[0079] Determine the similarity between the first vector and each vector in the preset database, where the preset database includes multiple vectors and the task corresponding to each vector;
[0080] Sort the multiple vectors in the preset database in descending order of similarity to obtain a vector sequence;
[0081] Determine the first n vectors in the vector sequence as multiple candidate vectors, where n is a preset quantity;
[0082] Determine the multiple candidate tasks corresponding to the multiple candidate vectors respectively.
[0083] In a possible implementation manner, the second determination module is specifically configured to, according to the current environment data collected by the embodied intelligent agent:
[0084] Identify the user's hand movement according to the current environment data collected by the embodied intelligent agent;
[0085] Obtain the current time when the user's hand movement meets the first preset condition;
[0086] Obtain the matter information associated with the task execution object in the target task when the current time meets the preset time period;
[0087] Determine the target task according to the matter information.
[0088] In a possible implementation manner, the second determination module is specifically configured to:
[0089] Identify the user's action intention according to the user's hand movement;
[0090] Obtain the user preference information associated with the action intention;
[0091] Determine the target task according to the matter information and the user preference information.
[0092] In a possible implementation manner, the second determination module is specifically configured to:
[0093] Identify the user's action intention according to the user's hand movement when the current time meets the preset time period;
[0094] Obtain the correct number of items corresponding to the action intention;
[0095] Obtain the number of items triggered by the user's hand movement;
[0096] When the correct number of items does not match the triggered number of items, the target task further includes controlling the embodied intelligent agent to send a voice prompt message to the user, and the voice prompt message includes the correct number of items.
[0097] In a possible implementation manner, the device further includes an update module, and the update module is configured to:
[0098] If the task detection result is that the task execution fails, generate a task exception indication according to the target task, and the task exception indication is used to indicate not to execute the target task within a preset duration;
[0099] Update the historical execution data according to the task exception indication and at least one task execution information corresponding to the target task.
[0100] In a third aspect, the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0101] The memory stores computer-executable instructions;
[0102] The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of the first aspect.
[0103] In a fourth aspect, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the method according to any one of the first aspect.
[0104] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a computer, it implements the method according to any one of the first aspect.
[0105] A task execution method for an embodied agent provided by an embodiment of the present application determines predicted action data according to the user's voice text data and / or historical execution data, determines multiple candidate tasks according to the predicted action data, determines a target task from the multiple candidate tasks according to the current environment data collected by the embodied agent, executes the target task, and detects the execution process of the target task to obtain a task detection result. In this way, the target task can be dynamically determined from multiple candidate tasks according to the user's voice text data and / or historical execution data and the current environment data, and real-time detection is performed during the task execution process, improving the accuracy of the task execution of the embodied agent. BRIEF DESCRIPTION OF THE DRAWINGS
[0106] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings according to these drawings without creative efforts.
[0107] Figure 1 It is a schematic diagram of an application scenario provided by an embodiment of the present application;
[0108] Figure 2 It is a schematic flowchart of a task execution method for an embodied agent provided by an embodiment of the present application;
[0109] Figure 3A schematic flowchart of another method for a task execution of an embodied agent provided by an embodiment of the present application;
[0110] Figure 4 A schematic flowchart of yet another method for a task execution of an embodied agent provided by an embodiment of the present application;
[0111] Figure 5 A schematic structural diagram of a task execution device of an embodied agent provided by an embodiment of the present application;
[0112] Figure 6 A schematic structural diagram of an electronic device provided by an embodiment of the present application.
[0113] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and more detailed descriptions will be provided hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed Embodiments
[0114] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0115] It should be noted that although the terms "first", "second", etc. are used in the embodiments of the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. Optionally, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information.
[0116] It should be understood that the terms "comprising" and "including" indicate the presence of the previously mentioned features, steps, operations, but do not exclude the presence, appearance, or addition of one or at least one other feature, step, operation. The term "and / or" used in the present application may be interpreted as inclusive, or mean any one or any combination. Optionally, "A and / or B" means "any one of the following: A; B; A and B". Additionally, the character " / " in this text generally indicates an "or" relationship between the associated objects before and after.
[0117] With the development of artificial intelligence technology, embodied agents have also been continuously evolving in task execution.
[0118] In related technologies, a robotic task execution model can be used to determine the target task of an embodied agent and control the embodied agent to execute the target task. Among them, the robotic task execution model can be trained using training data sets in multiple scenarios, and the target task can be a long-term and high-complexity task.
[0119] During the above task execution process, there may be situations where the target task determined by the robotic task execution model is not accurate enough, or abnormalities occur during the process of the embodied agent executing the target task and cannot be corrected in a timely manner, resulting in low robustness of the task execution of the embodied agent.
[0120] To solve the above technical problems, an embodiment of the present application provides a method for task execution of an embodied agent. By determining predicted action data based on the user's voice text data and / or historical execution data, determining multiple candidate tasks based on the predicted action data, determining a target task from the multiple candidate tasks according to the current environmental data collected by the embodied agent, executing the target task, and detecting the execution process of the target task to obtain a task detection result. In this way, the target task can be dynamically determined from multiple candidate tasks based on the user's voice text data and / or historical execution data, as well as the current environmental data, and real-time detection can be performed during the task execution process to improve the accuracy of the task execution of the embodied agent.
[0121] Next, in combination with Figure 1 , an example of the structure of the embodied agent will be described.
[0122] Figure 1 FIG. Figure 1 , Figure 1 shows a schematic structural diagram of an embodied agent provided by an embodiment of the present application. Please refer to
[0123] The embodied agent 100 may include a perception module 101, a cognition module 102, a decision module 103, an execution module 104, and a communication module 105.
[0124] The perception module 101 may include a visual perception sub-module, an auditory perception sub-module, a tactile perception sub-module, and so on.
[0125] The visual perception sub-module may include a camera, which is similar to a human eye and can be used to capture image information of the surrounding environment. The embodied agent can analyze and process these images to identify features such as the shape, color, and position of objects, thereby perceiving the surrounding environment.
[0126] Tactile sensors can be distributed on the body surface or actuators of an embodied agent to sense information such as pressure, magnitude, and direction of force when contacting an object. This is very important for tasks such as achieving fine operations and avoiding collisions.
[0127] The cognitive module 102 can include a storage sub-module, an algorithm sub-module, and so on.
[0128] The storage sub-module can be used to store information acquired by the embodied agent during the learning and task execution processes, including environmental information, task objectives, historical decisions, etc. This information can provide a reference basis for the agent's decision-making and behavior.
[0129] The algorithm sub-module can enable the embodied agent to learn the laws of the environment, task patterns, and reasoning abilities from a large amount of data. The reasoning ability can be used to enable the embodied agent to perform logical thinking and decision-making based on existing knowledge and current perception information.
[0130] The decision-making module 103 can include a policy sub-module, a task planning sub-module, and so on.
[0131] The decision-making sub-module can generate specific action strategies based on the information provided by the perception module and the analysis results of the cognitive module. The decision-making sub-module can be implemented by a neural network, and through learning and training on a large amount of data, it can make optimal decisions according to different situations.
[0132] The task planning sub-module can decompose complex tasks into a series of atomic tasks and determine the execution order and time arrangement of the atomic tasks.
[0133] The execution module 104 can include a motion component and a control sub-module.
[0134] The motion component can be used to implement the physical movement of the embodied agent. For example, the embodied agent can move by rotating the wheels, and can walk by relying on the joint movements of the legs, etc.
[0135] The control sub-module can be used to convert the action instructions generated by the decision-making module into specific control signals for the motion mechanism, enabling the embodied agent to execute tasks according to the predetermined strategy and plan.
[0136] Next, the technical solutions shown in this application will be described in detail through specific embodiments. It should be noted that the following several embodiments can exist independently or be combined with each other. For the same or similar content, it will not be repeated in different embodiments.
[0137] Figure 2The flowchart shows a method for a task execution of an embodied agent provided by an embodiment of this application. The execution subject of the embodiment of this application can be an embodied agent or a decision-making module of the embodied agent. Refer to Figure 2 , the method includes:
[0138] S201. Determine predicted action data according to the user's voice text data and / or historical execution data.
[0139] The voice text data is obtained by analyzing the voice data collected by the embodied agent from the user, and the historical execution data is the data generated by executing the previous task.
[0140] Optionally, the user's voice text data and / or historical execution data can be input into a second preset model to determine the predicted action data; among them, the user voice text data is obtained by performing text recognition processing on the user's voice data.
[0141] Optionally, if there is the user's voice text data, the user's voice text data can be input into the second preset model to determine the first predicted action data. If there is historical execution data, the historical execution data can be input into the second prediction model to determine the second predicted action data. The predicted action data is determined according to the first predicted action data and / or the second predicted action data.
[0142] Among them, the second preset model can be a pre-set model or a model optimized by multiple task description information in the skill library, which is not limited here.
[0143] In this way, the optimized second preset model can better fit the known skills during task prediction. In this way, the second preset model not only improves the prediction ability for specific tasks, but also enhances its coordination with the existing skill library, thus ensuring the efficiency and accuracy of task execution.
[0144] It should be noted that the predicted action data can be determined according to any feasible implementation method, and the embodiment of this application does not limit this.
[0145] S202. Determine multiple candidate tasks according to the predicted action data.
[0146] The candidate tasks can be determined from a preset database or determined according to a third preset model, which is not limited here.
[0147] The preset database includes multiple vectors and the tasks corresponding to each vector.
[0148] Multiple candidate tasks with higher similarity can be determined according to the similarity with the predicted action data.
[0149] Optionally, the predicted action data can be subjected to vector embedding processing to obtain a first vector, and multiple candidate tasks are determined according to the similarity between the first vector and multiple vectors in the preset database.
[0150] Optionally, the predicted action data can be input into a third preset model to determine multiple candidate tasks.
[0151] It should be noted that multiple candidate tasks can be determined according to any feasible implementation manner, and the embodiments of the present application do not limit this.
[0152] S203. According to the current environment data collected by the embodied agent, a target task is determined from multiple candidate tasks.
[0153] The current environment data can include at least one current environment image and a current environment video, which is not limited here.
[0154] Optionally, description information corresponding to multiple candidate tasks can be determined, and according to the current environment data collected by the embodied agent, target description information is determined from the multiple pieces of description information, and the candidate task corresponding to the target description information is determined as the target task.
[0155] Optionally, the description information corresponding to multiple candidate tasks and the current environment data can be input into a first preset model to determine the target task.
[0156] Optionally, the current environment data is input into a first preset model to obtain current scene data; the current scene data and multiple candidate tasks are input into a second preset model to obtain the target task; or, the current environment data and multiple candidate tasks are input into a first preset model to obtain the target task.
[0157] It should be noted that the target task can be determined according to any feasible implementation manner, and the embodiments of the present application do not limit this.
[0158] S204. Execute the target task and detect the execution process of the target task to obtain a task detection result.
[0159] The task detection result can include successful task execution and failed task execution.
[0160] Among them, the failed task execution can be determined when the number of times of executing the target task is greater than or equal to a preset threshold.
[0161] Optionally, the target task can be executed, and the execution process of the target task is detected through the first preset model every detection period to obtain a task detection result.
[0162] Optionally, the target task can be executed, task execution information can be obtained in real time, and the task execution information can be detected in real time to obtain a task detection result.
[0163] It should be noted that the task detection result can be determined according to any feasible implementation manner, and the embodiments of the present application do not limit this.
[0164] For example, in a life scenario, assume the user issues a request: "My Coke has spilled. Please help me clean it up." After the user's voice is recognized, voice text data is obtained. According to the voice text data, the predicted action data is determined to be "I should first throw the Coke can into the trash can." Vector embedding processing is performed on the task description of the predicted action data to obtain a first vector, and the five candidate tasks most similar to the first vector are determined through cosine similarity calculation.
[0165] The five candidate tasks include information such as task name, unique identifier, task description, and their similarity.
[0166] One of the candidate tasks can be represented in the following way: [{"action": "put_coke_in_trash", "uuid": "xxxxxxxxxx", "description": "Throw the Coke into the trash can", "similarity": 0.92}
[0167] In addition, the current environmental data collected by the embodied agent is obtained. For example, the description information of the current environmental data is "I see a white table, and there is a bottle of spilled Coke on the table." According to the current environmental data, the previously determined five candidate tasks, and the voice text data, they are input into the language model. The language model selects the target task that best meets the user's needs according to the current environmental data and executes it. During the execution process, the generated execution data can be verified by the first preset model to determine whether the task is completed.
[0168] Optionally, after executing the target task and detecting the execution process of the target task to obtain a task detection result, the method further includes:
[0169] If the task detection result is that the task execution fails, a task exception indication is generated according to the target task. The task exception indication is used to indicate that the target task is not executed within a preset duration; according to the task exception indication and at least one task execution information corresponding to the target task, the historical execution data is updated;
[0170] If the task detection result is that the task execution is successful, the historical execution data is updated according to at least one task execution information corresponding to the target task.
[0171] For example, assume that the target task is Task 1. If the task detection result indicates that the task execution fails, a task exception indication is generated. The task exception indication is used to indicate that Task 1 should not be executed within a preset duration. Update this task exception indication to the historical execution data. During the determination process of the next target task, based on the task exception indication, it is possible to avoid predicting that the action data contains the task description information corresponding to Task 1.
[0172] The task execution method of the embodied agent provided in this embodiment determines the predicted action data based on the user's speech text data and / or historical execution data, determines multiple candidate tasks based on the predicted action data, determines the target task from the multiple candidate tasks according to the current environmental data collected by the embodied agent, executes the target task, and detects the execution process of the target task to obtain the task detection result. In this way, based on the user's speech text data and / or historical execution data, as well as the current environmental data, the target task can be dynamically determined from multiple candidate tasks, and real-time detection can be performed during the task execution process, improving the accuracy of the task execution of the embodied agent.
[0173] Next, in conjunction with Figure 3 , the process of executing the target task and detecting the execution process of the target task to obtain the task detection result (S204) will be explained.
[0174] Figure 3 It is a schematic flowchart of another task execution method of the embodied agent provided in the embodiment of the present application. Based on the above embodiment, refer to Figure 3 for a detailed description of this method. The method includes:
[0175] S301. Start executing the target task.
[0176] A task instruction for the target task can be sent to the execution module to start executing the target task.
[0177] S302. Determine whether the number of times of executing the target task is greater than or equal to a preset threshold.
[0178] If so, execute S307;
[0179] If not, execute S303.
[0180] It is possible to determine the number of times corresponding to the sent task instruction. If this number is greater than or equal to the preset threshold, it is determined that the task detection result is that the task execution fails; otherwise, execute S303.
[0181] This can avoid the abnormal repetition of executing this task, improving the flexibility of the task execution of the embodied agent.
[0182] S303. Determine the current detection result according to the task execution information of the target task in the current detection cycle.
[0183] The task execution information can be used to record the information recorded by the embodied agent when executing the target task.
[0184] The task execution information can include environmental perception data, time data of task execution, the state of the agent itself, etc., which are not limited here.
[0185] The environmental perception data can include multiple environmental videos or multiple environmental images.
[0186] For example, the environmental image can be the image information of surrounding objects collected by the camera when the embodied agent moves indoors.
[0187] Optionally, the target task and the task execution information can be input into the first preset model to obtain the detection result.
[0188] Optionally, multi-modal feature extraction can be performed on the task execution information to obtain the visual semantic features of the task execution information and the body pose features of the embodied agent; the visual semantic features and the body pose features are fused to obtain the fusion features; according to the fusion features, task recognition of the task execution information and behavior intention recognition of the embodied agent are performed to obtain the current task information of the embodied agent and the current behavior intention information of the embodied agent; according to the target task, the task information and the behavior intention information are verified for correctness to obtain the detection result.
[0189] Among them, the task execution information can be input into a pre-trained multi-modal large model, and in the multi-modal large model, multi-modal feature extraction is performed on the task execution information to obtain the visual semantic features of the task execution information and the body pose features of the embodied agent.
[0190] In this way, through the pre-trained multi-modal large model, the accuracy of multi-modal feature extraction is improved.
[0191] The visual semantic features and the body pose features are input into a deep learning model using the multi-head attention mechanism, and in the deep learning model, the visual semantic features and the body pose features are fused to obtain the fusion features.
[0192] The fusion features can reflect the visual semantics of the task video and can also reflect the body pose of the embodied agent in the task video.
[0193] Based on the fusion features for feature recognition, task recognition of the task execution information and behavior intention recognition of the embodied agent can be realized, and the current task information of the embodied agent and the current behavior intention information of the embodied agent are obtained.
[0194] If the task information of the target task is consistent with the current task information of the embodied intelligent agent and the intention information of the target task is consistent with the behavior intention information, it is determined that the detection result is normal; otherwise, it is determined that the detection result is abnormal.
[0195] It should be noted that the current detection result can be determined according to any feasible implementation manner, and the embodiments of the present application do not limit this.
[0196] S304. Determine whether the current detection result is normal.
[0197] If so, execute S305;
[0198] If not, execute S301.
[0199] S305. Determine whether the target task is completed.
[0200] If so, execute S306;
[0201] If not, execute S303.
[0202] It is possible to determine whether the target task is completed according to the task execution information.
[0203] Optionally, determine whether there is task execution information for at least one historical detection period; if so, determine multiple first execution images according to the task execution information for at least one historical detection period and the task execution information for the current detection period; perform multiple Fourier transform processes on the multiple first execution images to obtain a first frequency domain signal; if not, determine multiple second execution images according to the task execution information for the current detection period; perform multiple Fourier transform processes on the multiple second execution images to obtain a second frequency domain signal; input the first frequency domain signal or the second frequency domain signal into a first preset model to obtain an output result, and the output result is used to indicate that the target task is completed or indicate that the target task is not completed.
[0204] Among them, the historical detection period is the detection period after starting to execute the target task and before the current detection period.
[0205] The following method can be used to determine multiple first execution images according to the task execution information for at least one historical detection period and the task execution information for the current detection period: Image extraction can be performed on the task execution information for at least one historical detection period and the task execution information for the current detection period to obtain multiple first execution images.
[0206] The following method can be used to determine multiple second execution images according to the task execution information for the current detection period: Image extraction can be performed on the task execution information for the current detection period to obtain multiple first execution images.
[0207] The multi-Fourier transform processing may include the following steps:
[0208] Perform the first Fourier transform processing on the execution image to obtain the first frequency domain feature;
[0209] Perform the second Fourier transform processing on the read first frequency domain feature to obtain the second frequency domain feature;
[0210] Extract the real part feature from the second frequency domain feature to obtain the frequency domain signal.
[0211] For example, the Fourier transform processing can be implemented by the following formula:
[0212]
[0213] M and N are respectively the two-dimensional dimensions corresponding to the execution image.
[0214] x m,n is the (m, n) -th element in the two-dimensional dimension corresponding to the execution image.
[0215] X k,l is the (k, l) -th feature of the frequency domain feature;
[0216] i is the imaginary unit.
[0217] The real part feature extraction may be to extract the real part of the frequency domain feature.
[0218] In this way, through the multi-Fourier transform processing, the time domain feature in the execution image can be converted to the frequency domain, and the periodic feature of the image can be effectively retained, so as to more accurately judge whether the target task has been completed.
[0219] Optionally, a trained first preset model can be obtained in the following way: Obtain multiple images and the text corresponding to each image; Perform multi-Fourier transform processing on each image to obtain the frequency domain signal corresponding to each image; Generate multiple multi-modal training data according to the frequency domain signal corresponding to each image and the text corresponding to each image, where the multi-modal training data includes the frequency domain signal and the text; Obtain a multi-modal preset model, and the preset model includes a visual encoder, a language model, and a projector; Perform adjustment operations on the multi-modal preset model to obtain a multi-modal intermediate model, and the adjustment operations include at least one of the following: Freezing the visual encoder, freezing the language model, replacing the activation function in the projector, adding a normalization layer; Train the multi-modal intermediate model according to multiple multi-modal training data to obtain the first preset model.
[0220] Among them, the replacement operation of the activation function in the projector can be implemented in the following way: Determine the first activation function as the activation function in the projector.
[0221] For example, the first activation function can be represented by the following formula:
[0222]
[0223] Among them, Swish(x) can represent the first activation function, and x can represent the input value.
[0224] The operation of adding a normalization layer can be implemented in the following way:
[0225] In this way, a relatively small positive value can be output through the first activation function. This processing method can maintain a certain information transmission in the negative value region, thereby helping to better retain the details of the frequency domain signal after Fourier transform, avoid the attenuation of the frequency domain signal, and further better retain the periodic details of the image and avoid information loss.
[0226] Adding a normalization layer can perform the normalization operation after the data passes through the activation function.
[0227] Adding a normalization layer can standardize each feature, making the distribution of features more consistent, thereby accelerating the gradient propagation during training and reducing the training instability problem caused by uneven data distribution. Such an operation can not only help the network maintain efficient learning ability on the frequency domain signal, but also improve the convergence speed and stability of the model.
[0228] S306. Determine that the task detection result is that the task is completed.
[0229] Optionally, when determining that the task detection result is that the task is completed, record the completion status of the target task and at least one task execution information to obtain historical execution data, and refer to the historical execution data to select and execute the next task.
[0230] Among them, at least one task execution information can be embedded to obtain multiple information such as the embedded description content, task name, serial number, etc., and store the multiple information in a preset format to form a data set for subsequent content recall.
[0231] S307. Determine that the task detection result is that the task is not completed.
[0232] For the implementation content of each step in the embodiments of the present application, reference can be made to the description of the corresponding steps or operations in the above method embodiments, and repeated content will not be elaborated.
[0233] A task execution method for an embodied agent provided in this embodiment includes starting to execute a target task; performing at least one detection operation until the target task is completed, or when the number of times of executing the target task is greater than or equal to a preset threshold, obtaining a task detection result; the detection operation includes: determining a current detection result according to the task execution information of the target task in the current detection cycle; if the current detection result is an abnormal detection, restart executing the target task; if the current detection result is a normal detection, determine whether the target task is completed, and if not, perform the detection operation in the next detection cycle. In this way, detection can be performed in real time during the task execution process, improving the accuracy of the task execution of the embodied agent.
[0234] Next, in combination with Figure 4 , the process of determining multiple candidate tasks according to the predicted action data (S202) will be explained.
[0235] Figure 4 It is a schematic flowchart of another task execution method for an embodied agent provided in an embodiment of the present application. Based on the above embodiment, refer to Figure 4 , this method includes:
[0236] S401. Perform vector embedding processing on the predicted action data to obtain a first vector.
[0237] The first vector can be used to represent low-dimensional words.
[0238] The predicted action data can be input into a vector embedding model to obtain a first vector.
[0239] It should be noted that the vector embedding model is not limited.
[0240] S402. Determine the similarity between the first vector and each vector in the preset database.
[0241] The preset database includes multiple vectors and the tasks corresponding to each vector.
[0242] Obtain the preset database and calculate the similarity between the first vector and each vector in the preset database.
[0243] Optionally, the similarity can be calculated in the following manner:
[0244]
[0245] Among them, Q is the first vector;
[0246] D i is the vector corresponding to the i-th task in the preset database;
[0247] Cosine Similarity(Q,Di ) is the similarity between the first vector and the vector corresponding to the i-th task in the preset database;
[0248] Q·D i is the dot product of the two vectors;
[0249] ‖Q‖ and ‖D i ‖ are the Euclidean norms of vectors Q and D respectively i .
[0250] S403. Sort the multiple vectors in the preset database in descending order of similarity to obtain a vector sequence.
[0251] The multiple vectors in the preset database can be sorted in descending order of similarity in sequence to determine a vector sequence including the first n vectors.
[0252] S404. Determine the first n vectors in the vector sequence as multiple candidate vectors.
[0253] n is a preset quantity.
[0254] For example, n can be 5 or 3, which is not limited herein.
[0255] S405. Determine multiple candidate tasks corresponding to the multiple candidate vectors respectively.
[0256] For the implementation content of each step in the embodiments of the present application, reference can be made to the description of the corresponding steps or operations in the above method embodiments, and repeated content will not be elaborated.
[0257] The task execution method of the embodied intelligent agent provided in this embodiment obtains the first vector by performing vector embedding processing on the predicted action data; determines the similarity between the first vector and each vector in the preset database, where the preset database includes multiple vectors and the tasks corresponding to each vector; sorts the multiple vectors in the preset database in descending order of similarity to obtain a vector sequence; determines the first n vectors in the vector sequence as multiple candidate vectors, where n is a preset quantity; determines multiple candidate tasks corresponding to the multiple candidate vectors respectively. Among them, the candidate vectors and the candidate tasks are in one-to-one correspondence. In this way, multiple candidate tasks can be determined according to the similarity between the first vector corresponding to the predicted action data and each vector in the preset database, improving the accuracy of the task execution of the embodied intelligent agent.
[0258] In a possible implementation manner, based on any of the above embodiments, the above step S203 includes:
[0259] S2031. Identify the user's hand actions according to the current environmental data collected by the above-mentioned embodied intelligent agent.
[0260] S2032. When the above-mentioned user's hand movement meets the first preset condition, obtain the current time.
[0261] S2033. When the above-mentioned current time meets the preset time period, obtain the precautions information associated with the task execution object in the target task.
[0262] S2034. Determine the target task according to the above-mentioned precautions information.
[0263] Specifically, the above-mentioned current environmental data includes the image acquisition data of the environment. According to this image acquisition data, the user's hand movement can be recognized. The first preset condition can be, for example, that the hand has just picked up the medicine. That is, when the user has just picked up the medicine, generally speaking, the user is about to take the medicine, so drinking water is needed. In this case, obtain the current time. The above-mentioned preset time period corresponds to the user's hand movement. For example, when the hand movement is a medicine-picking movement, the preset time period can be the user's medicine-taking time period within the last 3 days. If the current time is within this preset time period, then combined with the aforementioned user's hand movement, that is, based on two factors of the current time and the user's hand movement, the user's action intention is comprehensively determined, which makes the accuracy of task determination higher. In this example, it is determined that the user's action intention is to take medicine, so the next target task is to get drinking water.
[0264] The above-mentioned precautions information can be, for example, the water temperature for taking the medicine determined according to the corresponding medicine instruction manual or the doctor's written instructions on the electronic medical record. And determine the final water-taking task according to this water temperature.
[0265] When the water temperature for taking the medicine is greater than the first preset threshold, then control the embodied intelligent agent to heat the water using the heating device, and hand it to the user after heating is completed. The first preset threshold can be 20 degrees Celsius or room temperature.
[0266] This embodiment can accurately predict the user's next need without the user issuing a voice command, giving the user a better experience, especially suitable for users who cannot speak.
[0267] In some alternative embodiments, the above-mentioned step S2034 includes:
[0268] According to the above-mentioned user's hand movement, recognize the user's action intention.
[0269] Obtain the user preference information associated with the above-mentioned action intention.
[0270] Determine the target task according to the above-mentioned precautions information and the above-mentioned user preference information.
[0271] Exemplarily, for instance, if the user's habit (i.e., user preference information) is to rinse their mouth with water after taking medicine, then the embodied agent will divide the fetched water into portions, with one portion for taking the medicine and the other for rinsing the mouth after the user takes the medicine. Then the determined target task will accordingly be, for example, to fetch water first and then divide the water.
[0272] In some alternative embodiments, based on the above embodiments, step S203 includes:
[0273] When the above current time meets the preset time period, identify the user's action intention according to the above user hand movement.
[0274] Obtain the correct quantity of items corresponding to the above action intention.
[0275] Obtain the quantity of items triggered by the above user hand movement.
[0276] When the above correct quantity of items does not match the triggered quantity of items, the above target task further includes controlling the embodied agent to send a voice prompt message to the above user, and the voice prompt message includes the above correct quantity of items.
[0277] Specifically, when the embodied agent detects that the quantity of medicine in the user's hand is incorrect based on the acquired image data (which can be judged according to the medicine instruction manual or the electronic medical record), it sends a voice prompt message to the user, and the voice prompt message includes the correct quantity of medicine, so as to prompt the user to take the medicine according to the correct quantity of medicine, thus improving the user experience.
[0278] Figure 5 It is a schematic structural diagram of a task execution device of an embodied agent provided by an embodiment of the present application. Please refer to Figure 5 , the task execution device 500 of the embodied agent includes a prediction module 501, a first determination module 502, a second determination module 503, and an execution module 504, wherein,
[0279] The prediction module 501 is configured to determine prediction action data according to the user's voice text data and / or historical execution data, where the voice text data is obtained by analyzing the voice data collected by the embodied agent from the user, and the historical execution data is the data generated by executing the previous task;
[0280] The first determination module 502 is configured to determine multiple candidate tasks according to the prediction action data;
[0281] The second determination module 503 is configured to determine a target task from the multiple candidate tasks according to the current environment data collected by the embodied agent;
[0282] An execution module 504 is configured to execute the target task and detect the execution process of the target task to obtain a task detection result.
[0283] In a possible implementation manner, the execution module 504 is specifically configured to:
[0284] Start to execute the target task;
[0285] Execute at least one detection operation until the target task is completed, or when the number of times of executing the target task is greater than or equal to a preset threshold, obtain a task detection result, where the task detection result includes successful task execution and failed task execution;
[0286] The detection operation includes: determining a current detection result according to the task execution information of the target task in the current detection period; if the current detection result is an abnormal detection, restart the execution of the target task; if the current detection result is a normal detection, determine whether the target task is completed, and if not, execute the detection operation in the next detection period.
[0287] In a possible implementation manner, the execution module 504 is specifically configured to:
[0288] Determine whether there is task execution information of at least one historical detection period, where the historical detection period is a detection period after starting to execute the target task and before the current detection period;
[0289] If so, determine multiple first execution images according to the task execution information of at least one historical detection period and the task execution information of the current detection period; perform multiple Fourier transform processes on the multiple first execution images to obtain a first frequency domain signal;
[0290] If not, determine multiple second execution images according to the task execution information of the current detection period; perform multiple Fourier transform processes on the multiple second execution images to obtain a second frequency domain signal;
[0291] Input the first frequency domain signal or the second frequency domain signal into a first preset model to obtain an output result, where the output result is used to indicate that the target task is completed or indicate that the target task is not completed.
[0292] In a possible implementation manner, the device further includes a training module 505, and the training module 505 is configured to:
[0293] Obtain multiple images and the text corresponding to each image;
[0294] Perform multiple Fourier transform processes on each image to obtain a frequency domain signal corresponding to each image;
[0295] Generate a plurality of multimodal training data according to the frequency-domain signals corresponding to the respective images and the texts corresponding to the respective images, where the multimodal training data includes frequency-domain signals and texts;
[0296] Obtain a multimodal preset model, where the multimodal preset model includes a visual encoder, a language model, and a projector;
[0297] Perform an adjustment operation on the multimodal preset model to obtain a multimodal intermediate model, where the adjustment operation includes at least one of the following: performing a freezing operation on the visual encoder, performing a freezing process operation on the language model, performing a replacement operation on the activation function in the projector, and adding a normalization layer operation;
[0298] Train the multimodal intermediate model according to the plurality of multimodal training data to obtain the first preset model.
[0299] In a possible implementation manner, the second determination module 503 is specifically configured to:
[0300] Input the current environmental data into the first preset model to obtain current scene data; input the current scene data and the multiple candidate tasks into the second preset model to obtain the target task; or,
[0301] Input the current environmental data and the multiple candidate tasks into the first preset model to obtain the target task.
[0302] In a possible implementation manner, the first determination module 502 is specifically configured to:
[0303] Perform vector embedding processing on the predicted action data to obtain a first vector;
[0304] Determine the similarity between the first vector and each vector in a preset database, where the preset database includes multiple vectors and the task corresponding to each vector;
[0305] Sort the multiple vectors in the preset database in descending order of similarity to obtain a vector sequence;
[0306] Determine the first n vectors in the vector sequence as multiple candidate vectors, where n is a preset quantity;
[0307] Determine the multiple candidate tasks corresponding to the multiple candidate vectors respectively.
[0308] In a possible implementation manner, according to the current environmental data collected by the embodied intelligent agent, the second determination module 503 is specifically configured to:
[0309] Identify the user's hand movements based on the current environmental data collected by the embodied agent;
[0310] When the user's hand movements meet the first preset condition, obtain the current time;
[0311] When the current time meets the preset time period, obtain the attention information associated with the task execution object in the target task;
[0312] Determine the target task according to the attention information;
[0313] In a possible implementation manner, the second determination module 503 is specifically configured to:
[0314] Identify the user's action intention according to the user's hand movements;
[0315] Obtain the user preference information associated with the action intention;
[0316] Determine the target task according to the attention information and the user preference information;
[0317] In a possible implementation manner, the second determination module 503 is specifically configured to:
[0318] When the current time meets the preset time period, identify the user's action intention according to the user's hand movements;
[0319] Obtain the correct number of items corresponding to the action intention;
[0320] Obtain the number of items triggered by the user's hand movements;
[0321] When the correct number of items does not match the number of triggered items, the target task further includes controlling the embodied agent to send a voice prompt message to the user, and the voice prompt message includes the correct number of items;
[0322] In a possible implementation manner, the device further includes an update module 506, and the update module 506 is used for:
[0323] If the task detection result is that the task execution fails, generate a task exception indication according to the target task, and the task exception indication is used to indicate that the target task is not executed within a preset duration;
[0324] Update the historical execution data according to the task exception indication and at least one task execution information corresponding to the target task;
[0325] Figure 6The structural schematic diagram of the electronic device provided by an embodiment of the present application. Exemplarily, the electronic device may be provided as a server or a computer. Refer to Figure 6 , the electronic device 600 includes a processing component 601, which further includes one or more processors, and memory resources represented by a memory 602 for storing instructions executable by the processing component 601, such as application programs. The application programs stored in the memory 602 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 601 is configured to execute instructions to perform any of the above method embodiments.
[0326] The electronic device 600 may further include a power supply component 603 configured to perform power management of the electronic device 600, a wired or wireless network interface 604 configured to connect the electronic device 600 to a network, and an input / output (I / O) interface 605. The electronic device 600 may operate based on an operating system stored in the memory 602, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSD TM or the like.
[0327] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When the processor executes the computer-executable instructions, the solution of the task execution method of the embodied intelligent agent as described above is implemented.
[0328] The present application also provides a computer program product, including a computer program, which implements the solution of the task execution method of the embodied intelligent agent as described above when executed by the processor.
[0329] The above computer-readable storage medium may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disk. The readable storage medium may be any available medium accessible by a general-purpose or special-purpose computer.
[0330] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in the task execution device of the embodied agent.
[0331] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0332] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A task execution method of an embodied intelligent agent, characterized in that: The method comprises: Determine the predicted action data according to the user's voice text data and / or historical execution data, wherein the voice text data is obtained by analyzing the voice data of the user collected by the embodied intelligent agent, and the historical execution data is data generated by executing the last task; Determining a plurality of tasks to be selected according to the predicted action data; Determining a target task from the plurality of candidate tasks according to current environment data collected by the embodied intelligent agent; The target task is executed, and the execution process of the target task is detected to obtain a task detection result.
2. The method according to claim 1, characterized in that Executing the target task and detecting the execution process of the target task to obtain a task detection result includes: Start executing the target task; Perform at least one detection operation until the target task is completed, or the number of times the target task is executed is greater than or equal to a preset threshold, and obtain a task detection result, wherein the task detection result includes a successful task execution and a failed task execution; The detection operation includes: determining the current detection result based on the task execution information of the target task in the current detection cycle; if the current detection result is a detection abnormality, restarting the execution of the target task; if the current detection result is a detection normal, judging whether the target task is completed, if not, executing the detection operation in the next detection cycle.
3. The method according to claim 2, characterized in that Determining whether the target task is completed includes: Determine whether there is task execution information of at least one historical detection cycle, where the historical detection cycle is a detection cycle after the target task is started and before the current detection cycle; If yes, determine a plurality of first execution images according to the task execution information of at least one historical detection cycle and the task execution information of the current detection cycle; perform a plurality of Fourier transform processes on the plurality of first execution images to obtain a first frequency domain signal; If not, determining a plurality of second execution images according to the task execution information of the current detection cycle; performing a plurality of Fourier transform processes on the plurality of second execution images to obtain a second frequency domain signal; The first frequency domain signal or the second frequency domain signal is input into a first preset model to obtain an output result, where the output result is used to indicate that the target task is completed or that the target task is not completed.
4. The method according to claim 3, characterized in that The method further comprises: Get multiple images and text corresponding to each image; Performing multiple Fourier transform processes on each image to obtain the frequency domain signal corresponding to each image; Generate a plurality of multimodal training data according to the frequency domain signals corresponding to each image and the text corresponding to each image, wherein the multimodal training data includes the frequency domain signals and the text; Acquire a multimodal preset model, wherein the multimodal preset model includes a visual encoder, a language model, and a projector; Performing an adjustment operation on the multimodal preset model to obtain a multimodal intermediate model, wherein the adjustment operation includes at least one of the following: performing a freezing operation on the visual encoder, performing a freezing processing operation on the language model, performing a replacement operation on the activation function in the projector, and adding a normalization layer operation; The multimodal intermediate model is trained according to the multiple multimodal training data to obtain the first preset model.
5. The method according to any one of claims 1 to 4, characterized in that: Determining a target task from the plurality of candidate tasks according to the current environment data collected by the embodied intelligent agent includes: Input the current environment data into the first preset model to obtain the current scene data; input the current scene data and the multiple tasks to be selected into the second preset model to obtain the target task; or, The current environment data and the plurality of tasks to be selected are input into the first preset model to obtain the target task.
6. The method according to any one of claims 1 to 4, characterized in that: According to the predicted action data, a plurality of tasks to be selected are determined, including: Performing vector embedding processing on the predicted action data to obtain a first vector; Determine a similarity between the first vector and each vector in a preset database, wherein the preset database includes a plurality of vectors and a task corresponding to each vector; Sorting the multiple vectors in the preset database in descending order of similarity to obtain a vector sequence; Determine the first n vectors in the vector sequence as a plurality of candidate vectors, where n is a preset number; Determine a plurality of tasks to be selected corresponding to the plurality of vectors to be selected respectively.
7. The method according to any one of claims 1 to 4, characterized in that: The step of determining a target task from the plurality of tasks to be selected based on the current environment data collected by the embodied intelligent agent comprises: recognizing a user's hand movements based on current environment data collected by the embodied intelligent agent; When the user's hand movement satisfies a first preset condition, obtaining the current time; When the current time meets the preset time period, obtaining the precaution information associated with the task execution object in the target task; The target task is determined according to the precaution information.
8. The method according to claim 7, characterized in that The determining of the target task according to the precaution information includes: According to the hand movement of the user, identifying the movement intention of the user; Acquire user preference information associated with the action intention; A target task is determined according to the precaution information and the user preference information.
9. The method according to claim 7, characterized in that: The step of determining a target task from the plurality of tasks to be selected based on the current environment data collected by the embodied intelligent agent comprises: When the current time meets the preset time period, identifying the user's action intention according to the user's hand movement; Obtaining the correct number of items corresponding to the action intention; Obtaining the number of items triggered by the user's hand motion; In the case where the correct number of items does not match the triggered number of items, the target task also includes controlling the embodied intelligent body to issue a voice prompt message to the user, wherein the voice prompt message includes the correct number of items.
10. The method according to any one of claims 1 to 4, characterized in that: After executing the target task and detecting the execution process of the target task to obtain the task detection result, the method further includes: If the task detection result is that the task execution fails, a task abnormality indication is generated according to the target task, and the task abnormality indication is used to indicate that the target task will not be executed within a preset time period; The historical execution data is updated according to the task abnormality indication and at least one task execution information corresponding to the target task.
Citation Information
Patent Citations
Robot control method and device, storage medium and electronic equipment
CN117444956A
Intelligent reception robot device with body and working method thereof
CN117484518A
Large model body perception and decision-making integrated method, system and device
CN118332298A
Control method of intelligent robot with body
CN119200585A
Information interaction method of intelligent robot with body
CN119260754A