Intelligent data processing method and equipment for body

By extracting and fusion multimodal feature of the task reference video of the embodied intelligent body, identifying the task and behavioral intentions, and determining effective video clips, the problems of high labor costs and low efficiency in embodied intelligent data processing are solved, and automated processing and efficient and accurate data processing are achieved.

CN120147918AActive Publication Date: 2025-06-13人形机器人(上海)有限公司

Patent Information

Application Number
CN202510119577.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-13
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

There are problems of high labor costs and low efficiency in embodied intelligent data processing, especially in the process of task data collection and processing, which requires a lot of manual intervention.

Method used

By obtaining the task reference video of the embodied agent, multimodal feature extraction is performed, visual semantic features and body posture features are integrated, task recognition and behavioral intention recognition are performed, and effective video clips are determined for operation verification and reference during the execution of tasks by the embodied agent.

Benefits of technology

It improves the accuracy of task recognition and behavioral intention recognition, realizes the automated processing of embodied intelligent data, reduces labor costs, and improves data processing efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147918A_ABST
    Figure CN120147918A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent data processing method and device, and relates to the intelligent field, and the intelligent data processing method comprises the steps: obtaining a task reference video of an intelligent agent; performing multi-modal feature extraction on the task reference video to obtain visual semantic features of the task reference video and body posture features of the task execution object; fusing the visual semantic features and the body posture features to obtain fused features; performing task identification on the task reference video and performing behavior intention identification on the task execution object according to the fusion feature to obtain an identification result; according to the recognition result, effective video clips of the task reference video are determined, and the effective video clips are used for verification and / or reference of operation in the task execution process of the agent with the body. According to the method, automatic processing of the intelligent data is realized through a multi-mode technology, the data processing efficiency and accuracy are improved, and the labor cost is saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of embodied intelligence technology, and in particular, to an embodied intelligence data processing method and device. Background Art

[0002] With the rapid development of industrial automation and artificial intelligence technologies, human-machine collaboration through embodied intelligent agents has become one of the important research directions in intelligent manufacturing.

[0003] In related technologies, the task data collection of embodied intelligent agents requires a large amount of manual intervention: one person is responsible for the teaching operation of the embodied intelligent agent, and another person is responsible for the collection and processing of data.

[0004] The above method increases the labor cost and has a low data processing efficiency. Summary of the Invention

[0005] This application provides an embodied intelligence data processing method and device to solve the problems of high labor cost and low efficiency in embodied intelligence data processing.

[0006] In a first aspect, this application provides an embodied intelligence data processing method, including: obtaining a task reference video of an embodied intelligent agent, where the task reference video includes the behavior information of a task execution object and the state information of task-related objects; performing multi-modal feature extraction on the task reference video to obtain the visual semantic feature and the body pose feature of the task reference video; fusing the visual semantic feature and the body pose feature to obtain a first fusion feature; performing task recognition on the task reference video and behavior intention recognition on the task execution object according to the first fusion feature to obtain a recognition result; determining an effective video segment of the task reference video according to the recognition result, and the effective video segment is used for the verification and / or reference of the operations during the task execution process of the embodied intelligent agent.

[0007] In some embodiments, fusing the visual semantic feature and the body pose feature to obtain a first fusion feature includes: inputting the visual semantic feature and the body pose feature into a deep learning model using a multi-head attention mechanism, and in the deep learning model, fusing the visual semantic feature and the body pose feature to obtain a first fusion feature.

[0008] In some embodiments, the recognition result includes the task category of the task reference video and the behavioral intention categories of the task execution object at multiple time points. According to the first fusion feature, task recognition is performed on the task reference video and behavioral intention recognition is performed on the task execution object to obtain the recognition result, including: in a deep learning model using a multi-head attention mechanism, according to the first fusion feature, task recognition is performed on the task reference video and behavioral intention recognition is performed on the task execution object to obtain the output data of the deep learning model, where the output data includes the confidence of the candidate task category and the confidence of the candidate behavioral intention categories corresponding to multiple time points respectively; according to the confidence of the candidate task category, determine the task category of the task reference video; according to the confidence of the candidate behavioral intention categories corresponding to multiple time points respectively, determine the behavioral intention categories of the task execution object at multiple time points.

[0009] In some embodiments, the recognition result includes the task category of the task reference video and the behavioral intention categories of the task execution object at multiple time points. The valid video segment includes the video segment from the start to the end of the task. According to the recognition result, determining the valid video segment of the task reference video includes: according to the task category, obtaining the task actions corresponding to multiple task stages, where the multiple task stages include the task start stage, the task execution stage, and the task end stage; matching the task actions corresponding to multiple task stages with the behavioral intention categories of the task execution object at multiple time points to obtain the time information corresponding to multiple task stages respectively; according to the time information corresponding to multiple task stages respectively and the time information corresponding to the video frames in the task reference video, determining the valid video segment in the task reference video; after determining the valid video segment in the task reference video according to the time information corresponding to multiple task stages respectively and the time information corresponding to the video frames in the task reference video, it further includes: marking multiple task stages for the valid video segment; and / or, marking the task category of the valid video segment; and / or, saving the task actions corresponding to multiple task stages in the valid video segment.

[0010] In some embodiments, after obtaining the task reference video corresponding to the embodied agent, it further includes: in the task reference video, extracting the body key points of the task execution object to obtain the pose data of the task execution object, where the pose data includes the sequence of body key points of the task execution object sorted by time; aligning the pose data, the agent operation data, and the valid video segment in time to obtain the offline data, which is used for the verification and / or reference of the operations during the task execution of the embodied agent, and the agent operation data is the sensing data, motion trajectory data, and / or control instruction data of the robot participating in the task collaboration in the task reference video.

[0011] In some embodiments, the verification and / or reference process of the operations during the execution of tasks by an embodied agent includes: obtaining the current task video of the embodied agent; performing multi-modal feature extraction on the task video to obtain the second visual semantic feature of the task video and the second body pose feature of the embodied agent; fusing the second visual semantic feature and the second body pose feature to obtain a second fused feature; based on the second fused feature, performing task recognition on the task video and behavior intention recognition on the embodied agent to obtain the current task information of the embodied agent and the current behavior intention information of the embodied agent; based on the valid video segment, the current task information of the embodied agent, and the current behavior intention information of the embodied agent, verifying the correctness of the current task actions of the embodied agent, predicting the movement path of the embodied agent, and / or predicting the next task actions of the embodied agent.

[0012] In some embodiments, the current task information of the embodied agent includes the current task category and / or task phase executed by the embodied agent, the current behavior intention information of the embodied agent includes the current behavior intention category of the embodied agent, and the valid video segment is marked with the task category and / or task phase; based on the valid video segment, the current task information of the embodied agent, and the current behavior intention information of the embodied agent, verifying the correctness of the current task actions of the embodied agent includes at least one of the following: matching the task category marked in the valid video segment with the current task category executed by the embodied agent to determine whether the current task executed by the embodied agent is correct; according to the current time and the time information corresponding to the task phase marked in the valid video segment, matching the task phase marked in the valid video segment with the current task phase executed by the embodied agent to determine whether the current task phase executed by the embodied agent is correct; according to the current time and the time information corresponding to the task phase marked in the valid video segment, determining the task phase corresponding to the current time, and matching the task actions of the task phase corresponding to the current time with the current behavior intention category of the embodied agent to determine whether the current behavior intention category of the embodied agent is correct.

[0013] In some embodiments, the current task information of the embodied agent includes the task category and / or task phase that the embodied agent is currently executing, and the current behavioral intention information of the embodied agent includes the category of the current behavioral intention of the embodied agent. The valid video segment is marked with the task category and / or task phase. Based on the valid video segment, the current task information of the embodied agent, and the current behavioral intention information of the embodied agent, predicting the movement path of the embodied agent and / or predicting the next task action of the embodied agent includes: matching the task category that the embodied agent is currently executing with the task category marked in the valid video segment, matching the category of the current behavioral intention of the embodied agent with the task actions of the task phase marked in the valid video segment, and determining the task phase that the embodied agent is currently executing. Based on the task phase that the embodied agent is currently executing and the task actions of the task phase marked in the valid video segment, predicting the movement path of the embodied agent and / or predicting the next task action of the embodied agent.

[0014] In some embodiments, the method further includes:

[0015] Verifying the correctness of the task execution of the embodied agent according to the valid video segment to obtain a verification result; the verification result is correct or incorrect;

[0016] In the case where the verification result of the embodied agent is incorrect, identifying the user's hand actions in the video according to the valid video segment;

[0017] In the case where the user's hand actions meet the first preset condition, obtaining the corresponding time of the video when the user's hand actions occur;

[0018] In the case where the corresponding time of the video meets the preset time period, obtaining the information on precautions associated with the task execution object in the task;

[0019] Determining and executing a new target task for the embodied agent according to the information on precautions.

[0020] In some embodiments, the determining and executing a new target task for the embodied agent according to the information on precautions includes:

[0021] Identifying the action intention of the user according to the user's hand actions;

[0022] Obtaining the user preference information associated with the action intention;

[0023] Determining a new target task for the embodied agent according to the information on precautions and the user preference information.

[0024] In a second aspect, the present application provides an embodied intelligence data processing device, including: a video data acquisition unit configured to acquire a task reference video of an embodied intelligent agent, where the task reference video includes behavior information of a task execution object and status information of task-related objects; a multi-modal feature extraction unit configured to perform multi-modal feature extraction on the task reference video to obtain visual semantic features and body pose features of the task reference video; a multi-modal feature fusion unit configured to fuse the visual semantic features and the body pose features to obtain a first fused feature; a task behavior recognition unit configured to perform task recognition on the task reference video and behavior intention recognition on the task execution object according to the first fused feature to obtain a recognition result; and a valid data determination unit configured to determine a valid video segment of the task reference video according to the recognition result, where the valid video segment is used for verification and / or reference of operations during the task execution process of the embodied intelligent agent.

[0025] In a third aspect, the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0026] The memory stores computer-executable instructions;

[0027] The processor executes the computer-executable instructions stored in the memory to implement the embodied intelligence data processing method as described in the first aspect of the present application.

[0028] In a fourth aspect, the present application provides a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the embodied intelligence data processing method as described in the first aspect of the present application.

[0029] In a fifth aspect, the present application provides a computer program product including a computer program, which, when executed by a processor, implements the embodied intelligence data processing method as described in the first aspect of the present application.

[0030] The embodied intelligence data processing method and device provided by the present application perform multi-modal feature extraction on the task reference video of the embodied intelligent agent to obtain video semantic features of the task reference video and body pose features of the task execution object. Based on the fused features of the video semantic features and the body pose features, task recognition is performed on the task reference video and behavior intention recognition is performed on the task execution object, and the accuracy of task recognition and behavior intention recognition is improved by using multi-modal technology; based on the results of task recognition and behavior intention recognition, a valid video segment of the task reference video is determined, realizing the automated processing of video data related to embodied intelligence and improving the data processing efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0032] Figure 1 It is an example diagram of the application scenario applicable to the embodiments of the present application;

[0033] Figure 2 It is a flowchart showing the process of the embodied intelligence data processing method provided by the embodiments of the present application Figure 1 ;

[0034] Figure 3 It is a flowchart showing the process of the embodied intelligence data processing method provided by the embodiments of the present application Figure 2 ;

[0035] Figure 4 It is a flowchart showing the verification and / or reference process of the operations during the task execution of the embodied intelligent agent in the embodied intelligence data processing method provided by the embodiments of the present application;

[0036] Figure 5 It is a schematic structural diagram of the embodied intelligence data processing device provided by the embodiments of the present application;

[0037] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0038] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.

[0039] An embodied intelligent agent is an intelligent system. During the task process, the embodied intelligent agent imitates the human perception ability, decision-making ability, and behavior ability by perceiving the environment, understanding the environment, and interacting with the environment. In the embodiments of the present application, the embodied intelligent agent can be one or more of intelligent agents such as humanoid embodied intelligent agents, animal-shaped embodied intelligent agents, multi-legged embodied intelligent agents, wheeled or tracked embodied intelligent agents, flying embodied intelligent agents, modular embodied intelligent agents, etc. In terms of existence form, the embodied intelligent agent can be a virtual object, a digital entity, or a machine entity (such as a robot).

[0040] During the process of collecting teaching data for embodied agents, usually one person performs the teaching operation, and another person is responsible for collecting and processing the data, resulting in low efficiency and high labor costs.

[0041] To solve the above problems, the embodiments of this application provide an embodied intelligence data processing method and device. In the embodied intelligence data processing method, for the task reference video of the embodied agent, multi-modal technology is used for processing to obtain the multi-modal features of the task reference video, and the multi-modal features are used for task recognition and behavior intention recognition, improving the accuracy of task recognition and behavior intention recognition; based on the results of task recognition and behavior intention recognition, the effective video segments of the task reference video are determined. Thus, without manual participation, the collection of effective video segments in the task reference video is automatically completed, improving the accuracy and collection efficiency of the effective video segments and reducing the labor cost.

[0042] Among them, the technical means adopted by the device, equipment and storage medium and the technical effects produced can refer to the technical means adopted by the above solution and the technical effects produced, which will not be elaborated here.

[0043] Figure 1 This is an example diagram of the application scenario applicable to the embodiments of this application. As Figure 1 shown, the application scenario includes a data processing device 101, a data verification device 102, and an embodied agent 103.

[0044] In the application scenario, the data processing device 101 can process the task reference video of the embodied agent 103 through multi-modal technology to obtain the effective video segments of the task reference video; the data verification device 102 can verify the operations during the task execution of the embodied agent 103 or provide a reference for the operations during the task execution of the embodied agent 103 based on the effective video segments and the real-time task video of the embodied agent.

[0045] Among them, the data processing device 101 and the data verification device 102 can be electronic devices, and the electronic devices include servers or terminals. The terminal can be a personal digital assistant (PDA) device, a handheld device with wireless communication function (such as a smart phone, a tablet computer), a computing device (such as a personal computer (PC)), a wearable device (such as a smart watch, a smart bracelet), and a smart home device (such as a smart speaker, a smart display device), etc. The server can be an independent server or a server cluster, and can be a local server or a cloud server. Figure 1 Taking the data processing device 101 and the data verification device 102 as servers as an example.

[0046] It should be noted thatFigure 1 This is only a schematic diagram of an application scenario provided by the embodiments of the present application. The embodiments of the present application do not Figure 1 limit the devices included therein, nor do they Figure 1 limit the positional relationship between the devices therein.

[0047] Next, the embodied intelligent data processing method will be introduced through specific embodiments.

[0048] Figure 2 This is a flowchart of the embodied intelligent data processing method provided by the embodiments of the present application. Figure 1 As Figure 2 shown, the method of the embodiments of the present application includes:

[0049] S201. Obtain the task reference video of the embodied intelligent agent. The task reference video contains the behavior information of the task execution object and the state information of the task-related objects.

[0050] Among them, the task reference video, such as the task teaching video, can be obtained by shooting the process of a person and / or an embodied intelligent agent correctly executing a task (such as a grasping task, a handling task).

[0051] Among them, in the task reference video, the task execution object can be a person and / or an embodied intelligent agent. For example, the task in the task reference video is a human-robot collaboration task, and the task execution objects include a person and an embodied intelligent agent. The task-related objects refer to the objects that need to be processed during the task, such as the objects that need to be moved, the objects that need to be packaged, etc.

[0052] In this embodiment, the task reference video of the embodied intelligent agent can be obtained from the database, or the task reference video of the embodied intelligent agent input by the user or other devices can be received, or the task reference video collected by the sensing device (such as a camera) can be directly obtained.

[0053] S202. Perform multi-modal feature extraction on the task reference video to obtain the visual semantic features of the task reference video and the body pose features of the task execution object.

[0054] Among them, the visual semantic features refer to the semantic features conveyed by the visual picture of the task reference video. In the task reference video, the visual semantic features are related to the task execution object and the task-related objects, and can reflect the characteristics of the task execution object, the characteristics of the task-related objects, and the relationship between the task execution object and the task-related objects. The body pose features of the task execution object can reflect the actions of the task execution object. When the task execution object is a person, the body pose features are the human body pose features.

[0055] In this embodiment, multi-modal feature extraction can be performed on the video frames in the task reference video (which can be each video frame or the video frames that have been screened, for example, the video frames whose picture content contains the task execution object and / or task-related objects after screening), to obtain the visual semantic features of the reference task video and the body posture features of the task execution object.

[0056] Optionally, before performing multi-modal feature extraction on the task reference video, preprocess the task reference video to improve the accuracy of multi-modal feature extraction.

[0057] Among them, preprocessing the task reference video may include: denoising the task reference video and / or enhancing the images of the video frames in the task reference video.

[0058] Among them, enhancing the images of the video frames may include: adjusting at least one of the color, brightness, contrast or saturation of the video frames. For example, the brightness and contrast of the video frames in the task reference video can be randomly adjusted to simulate different lighting conditions.

[0059] S203, fuse the visual semantic features and the body posture features to obtain a first fused feature.

[0060] In this embodiment, the visual semantic features and the body posture features can be fused through a feature fusion technique to obtain a first fused feature.

[0061] In a possible implementation manner, weighted summation is performed on the visual semantic features and the body posture features according to the weight coefficients to obtain a first fused feature. In this way, the fusion effect of the visual semantic features and the body posture features can be improved by adjusting the weight coefficients.

[0062] It should be noted that in addition to the weighted method, other fusion methods can also be used, and other formulas can be designed to perform operations on the visual semantic features and the body posture features to obtain a first fused feature.

[0063] S204, based on the first fused feature, perform task recognition on the task reference video and behavior intention recognition on the task execution object to obtain recognition results.

[0064] In this embodiment, since the first fusion feature is obtained by fusing multi-modal features, it can reflect both the visual semantics of the task reference video and the body posture of the task execution object in the task reference video. Feature recognition based on the first fusion feature can achieve task recognition of the task reference video and recognition of the behavior intention of the task execution object, and obtain a recognition result. The recognition result may include a task recognition result and a behavior intention recognition result. The task recognition result indicates the task information corresponding to the task reference video, and the behavior intention recognition result indicates the behavior intention corresponding to the task execution object.

[0065] S205. According to the recognition result, determine the effective video segment of the task reference video, where the effective video segment is used for verification and / or reference of the operations during the task execution of the embodied agent.

[0066] Among them, the effective video segment is a video segment related to task execution.

[0067] Among them, during the task execution of the embodied agent, the correctness of the task execution of the embodied agent can be verified based on the effective video segment, and / or the effective video segment can provide guidance for the embodied agent to execute the task, that is, provide reference.

[0068] In this embodiment, the task reference video may include video segments unrelated to task execution, such as a video segment recording the preparation work before task execution, or a video segment recording the task-unrelated operations performed by the task execution object during the task execution, or a video segment recording the operations related to the task execution object after the task ends. In order to accurately collect video data related to task execution, according to the recognition result, determine the effective video segment of the task reference video. In the process of determining the effective video segment of the task reference video according to the recognition result, the recognition result reflects the task information of the task reference video and the behavior intention of the task execution object. The task stage of the task reference video can be analyzed according to the recognition result to determine the video segment located in the task stage of the task reference video, that is, determine the effective video segment of the task reference video.

[0069] In the embodiment of the present application, for the task reference video, by extracting multi-modal features and performing task recognition and behavior intention recognition based on the multi-modal features, the accuracy of task recognition and behavior intention recognition is improved; based on the task recognition result and the behavior intention recognition result, the effective video segment of the task reference video is determined, and the accuracy of obtaining the effective video segment from the task reference video is improved. Thus, without manual participation, automatic processing of the task reference video is realized, and the data processing efficiency and accuracy are improved.

[0070] Figure 3 Flowchart schematic of the embodied intelligent data processing method provided by the embodiment of the present application Figure 2 .

[0071] As shown Figure 3 in the figure, the method of the embodiment of the present application may include:

[0072] S301, Obtain a task reference video of the embodied intelligent agent, where the task reference video contains the behavior information of the task execution object and the state information of the task-related objects.

[0073] S302, Perform multi-modal feature extraction on the task reference video to obtain the visual semantic features of the task reference video and the body pose features of the task execution object.

[0074] Among them, the implementation principles and technical effects of S301 to S302 can refer to the foregoing embodiments and will not be elaborated herein.

[0075] In a possible implementation, input the task reference video into a pre-trained multi-modal large model. In the multi-modal large model, perform multi-modal feature extraction on the task reference video to obtain the visual semantic features of the task reference video and the body pose features of the task execution object. Thus, the accuracy of multi-modal feature extraction is improved through the pre-trained multi-modal large model.

[0076] In a possible implementation, the task reference video can be input into a pre-trained convolutional neural network (CNN). In the CNN, perform feature extraction on the video frames of the task reference video to obtain the visual semantic features of the task reference video. Among them, the pre-training objective of the CNN is to improve the accuracy of the CNN in extracting visual semantic features.

[0077] Optionally, the pre-trained CNN adopts at least one of the following: Residual Network (ResNet), Visual Geometry Group network, or Inception network. These networks can be directly used for image feature extraction.

[0078] In yet another possible implementation, divide the video frames of the task reference video into multiple image patches, input the multiple image patches into a Vision Transformer (ViT) network, and in the ViT network, extract the visual semantic features by learning the relationships between the multiple image patches. Thus, the global visual semantic features are extracted using the ViT network, improving the accuracy of visual semantic feature extraction.

[0079] S303, Input the visual semantic features and the body pose features into a deep learning model using a multi-head attention mechanism. In the deep learning model, fuse the visual semantic features and the body pose features to obtain a first fusion feature.

[0080] In this embodiment, the visual semantic features and the body pose features are input into a deep learning model adopting a multi-head attention mechanism. In the deep learning model, through the multi-head attention mechanism, the features of different modalities can be adaptively weighted, that is, the visual semantic features and the body pose features are adaptively weighted, the mutual relationship between the visual semantic features and the body pose features is mined, the complex spatio-temporal dependence between the visual semantic features and the body pose features is captured, and more accurate fused features are obtained.

[0081] Optionally, the deep learning model adopting the multi-head attention mechanism is a Transformer architecture. Among them, the Transformer architecture is a neural network model based on the self-attention mechanism, which includes a self-attention mechanism and a multi-head attention mechanism. The processing effect of multi-modal features can be effectively improved through the self-attention mechanism and the multi-head attention mechanism in the Transformer architecture.

[0082] S304. In the deep learning model, according to the first fused feature, task recognition is performed on the task reference video and behavior intention recognition is performed on the task execution object to obtain the output data of the deep learning model. Among them, the output data of the deep learning model includes the confidence of the candidate task category and the confidence of the candidate behavior intention categories corresponding to multiple time points respectively.

[0083] Among them, the multiple time points can be multiple timestamps on the time axis of the task reference video.

[0084] Among them, there can be multiple candidate task categories. In the output data of the deep learning model, the confidence corresponding to each of the multiple candidate task categories can be included. Among the multiple time points, each time point can correspond to multiple candidate behavior intention categories, and the multiple candidate behavior intention categories respectively correspond to their own confidence.

[0085] It can be understood that the output data of the deep learning model includes a task recognition result and behavior intention recognition results corresponding to multiple time points respectively. In the task recognition result, the confidence corresponding to each of the multiple candidate task categories is included. In the behavior intention recognition result corresponding to the time point, the confidence corresponding to each of the multiple candidate behavior intention categories at this time point is included.

[0086] As an example, the output data of the deep learning model includes: a single task category recognition result, and the task category result is, for example, {"assembly task": 0.9, "handling task": 0.1}, where the confidence of the "assembly task" is the highest, indicating that the task reference video as a whole belongs to the "assembly task"; multiple behavior intention category recognition results, and one behavior intention category recognition result is, for example, {"grasping": 0.8, "handling": 0.1, "assembly": 0.1}, where the confidence of "grasping" is the highest, indicating that the behavior intention of the task execution object at the time point corresponding to this behavior intention category recognition result is "grasping".

[0087] In this embodiment, in the deep learning model, according to the first fusion feature, the task category of the task reference video is recognized to obtain a task category recognition result, and the task category recognition result includes the confidences corresponding to multiple candidate task categories respectively. The behavior intention category of the task execution object is recognized to obtain behavior intention recognition results corresponding to multiple time points respectively, and the behavior intention recognition result corresponding to the time point includes the confidences corresponding to multiple candidate behavior intention categories at this time point. The deep learning model outputs the task category recognition result and the behavior intention recognition result.

[0088] Among them, the deep learning model is obtained after being trained.

[0089] Optionally, during the training process of the deep learning model, behavior intention category labels and task category labels are used as training labels to perform multi-task learning on the deep learning model, so as to improve the recognition accuracy of the deep learning model for behavior intentions and task categories.

[0090] S305. Determine the task category of the task reference video according to the confidence of the candidate task category.

[0091] In this embodiment, in the task category recognition result, according to the confidence corresponding to the candidate task category, select the candidate task category with the highest confidence or a confidence greater than the first threshold, and determine this candidate task category as the task category of the task reference video.

[0092] S306. Determine the behavior intention categories of the task execution object at multiple time points according to the confidences of the candidate behavior intention categories corresponding to multiple time points respectively.

[0093] In this embodiment, for each of the multiple time points, among the candidate behavior intention categories corresponding to the time point, select the candidate behavior intention category with the highest confidence or a confidence greater than the second threshold, and determine this candidate behavior intention category as the behavior intention category of the task execution object at this time point.

[0094] S307. Determine the valid video segments of the task reference video according to the recognition result. The valid video segments are used for the verification and / or reference of the operations during the task execution of the embodied agent.

[0095] In this embodiment, analyze the task phases of the task reference video according to the task category of the task reference video and the behavior intention categories of the task execution object at multiple time points, and determine the video segments in the task phases of the task reference video, that is, determine the valid video segments of the task reference video.

[0096] In a possible implementation, as Figure 3 shown, S307 includes:

[0097] S3071. According to the task category of the task reference video, obtain the task actions corresponding to multiple task phases respectively. The multiple task phases include the task start phase, the task execution phase, and the task end phase.

[0098] In this implementation, different task categories may correspond to different task actions in the same task phase. The task actions of the task category in multiple task phases can be found in the task information library according to the task category of the task reference video, including the task actions of the task category in the task start phase, the task actions of the task category in the task execution phase, and the task actions of the task category in the task end phase.

[0099] S3072. Match the task actions corresponding to multiple task phases respectively with the behavior intention categories of the task execution object at multiple time points to obtain the time information corresponding to multiple task phases respectively.

[0100] In this implementation, match the task actions corresponding to multiple task phases respectively with the behavior intention categories of the task execution object at multiple time points to determine the task phases corresponding to the behavior intention categories of the task execution object at multiple time points; according to the task phases corresponding to the behavior intention categories of the task execution object at multiple time points, obtain the time information corresponding to multiple task phases respectively. Among them, the time information corresponding to the task phase includes the start time of the task phase and / or the end time of the task phase.

[0101] As an example, if the behavior intention category of the task execution object at time point A matches the task action of the task start phase, the time information of the task start phase includes time point A. In this way, the multiple time points included in the time information of the task start phase can be obtained, and the start time and / or end time of the task start phase can be determined according to the time order of the multiple time points.

[0102] S3073. Determine valid video segments in the task reference video according to the time information corresponding to multiple task stages and the time information corresponding to video frames in the task reference video.

[0103] In this implementation, the time information corresponding to multiple task stages and the time information corresponding to video frames in the task reference video can be matched to obtain the task stages corresponding to the video frames in the task reference video, that is, obtain the video frame ranges corresponding to multiple task stages respectively. According to the video frame ranges corresponding to multiple task stages respectively, valid video segments in the task reference video are obtained.

[0104] In this implementation, by analyzing the task stages corresponding to the task category of the task reference video and the behavior intention categories of the task execution object at multiple time points, the task actions in the task stages are analyzed, and then the valid video segments of the task reference video are determined, effectively improving the accuracy of determining the valid video segments of the task reference video.

[0105] In the embodiments of the present application, multi-modal feature extraction is performed on the task reference video of the embodied intelligent agent. The extracted multi-modal features are feature fused by a deep learning model adopting a multi-head attention mechanism, and task recognition and behavior intention recognition are performed based on the fused features, improving the accuracy of task recognition and behavior intention recognition; the valid video segments of the task reference video are determined based on the recognition results, improving the accuracy of the valid video segments. Therefore, while improving the efficiency of embodied intelligent data processing, the accuracy of embodied intelligent data processing is improved, providing more accurate reference data for the embodied intelligent agent to execute tasks, and improving the success rate and accuracy of the embodied intelligent agent to execute tasks.

[0106] In some embodiments, the embodied intelligent data processing method further includes: marking the valid video segments in the task reference video. Thus, the valid video segments in the task reference video are distinguished by marking. The marking of the valid video segments in the task reference video can be achieved by adding marks at the start time and end time of the valid video segments.

[0107] In some embodiments, the embodied intelligent data processing method further includes: marking multiple task stages for the valid video segments, and / or marking the task category of the valid video segments, and / or saving the task actions corresponding to multiple task stages respectively in the valid video segments.

[0108] In this embodiment, according to the foregoing embodiments, the time information corresponding to multiple task stages of the task reference video can be determined. The multiple task stages can be marked in the valid video segments according to the time information corresponding to multiple task stages and the time information of the valid video segments.

[0109] As an example, in the task reference video, the operator reaches out to grab an object. Through the deep learning model, the behavior intention category of the operator is identified as "grabbing". "Grabbing" is the task action in the task start stage and can mark the task start. In the task reference video, the operator moves the object to the target position. Through the deep learning model, the behavior intention category of the operator is identified as "carrying". "Carrying" is the task action in the task progress stage and can mark the task in progress. In the task reference video, the operator accurately places the object. Through the deep learning model, the behavior intention category of the operator is identified as "placing". "Placing" is the task action in the task end stage and can mark the task completion.

[0110] In this embodiment, according to the foregoing embodiment, the task category of the task reference video can be determined. The task category of the valid video segment is the same as that of the task reference video. The task category of the valid video segment can be marked by adding a task label to the valid video segment.

[0111] In this embodiment, according to the foregoing embodiment, the behavior intention category of the task execution object at multiple time points and the time information corresponding to multiple task stages of the task reference video can be determined. By performing time matching on the behavior intention category of the task execution object at multiple time points and the time information corresponding to multiple task stages of the task reference video, the behavior intention category corresponding to each of the multiple task stages, that is, the task action corresponding to each of the multiple task stages, can be determined, and the task actions corresponding to each of the multiple task stages are saved.

[0112] Among them, the multiple task stages marked for the valid video segment, the task category marked for the valid video segment, and / or the task actions corresponding to each of the multiple task stages can be used for the verification and / or reference of the operations during the task execution of the embodied intelligent agent, improving the success rate and accuracy of the task execution of the embodied intelligent agent.

[0113] In some embodiments, the embodied intelligent data processing method further includes: in the task reference video, extracting the body key points of the task execution object to obtain the pose data of the task execution object. The pose data includes the sequence of body key points of the task execution object sorted by time; aligning the pose data, the intelligent agent operation data, and the valid video segment in time to obtain the offline data. The offline data is used for the verification and / or reference of the operations during the task execution of the embodied intelligent agent. Thus, outside the valid video segment, the pose data of the task execution object and the intelligent agent operation data are also provided for the verification and / or reference of the operations during the task execution of the embodied intelligent agent. Combining these three types of data, the task execution process of the embodied intelligent agent can be analyzed more precisely in terms of behavior and the task progress can be tracked, and the success rate and accuracy of the task execution of the embodied intelligent agent can be improved.

[0114] Among them, the agent operation data includes the sensing data, motion trajectory data, and / or control instructions of the robots participating in task collaboration in the task reference video. The sensing data is recorded by the sensing devices on the robots participating in task collaboration in the task reference video. For example, when the robot is performing a task, it records its own motion state and sensor data (such as position, speed, joint angle, force sensor readings, etc.) in real time, and the motion data and sensor data are automatically saved by the robot control system.

[0115] Among them, the body key point extraction can be performed on the video frames of the task reference video through a pose estimation algorithm to obtain the body key point data of the task execution object at multiple time points, and the pose information of the task execution object can be obtained from the body key point data of the task execution object at multiple time points. When the task execution object is a relevant person, the human body key point extraction is performed on the video frames of the task reference video through a human body pose estimation algorithm to obtain the human body key point data at multiple time points, and these human body key point data are saved as a time series. The saved human key point data can be synchronized with the task reference video in time.

[0116] Optionally, combined with the depth information corresponding to the video frames in the task reference video, the body key point extraction is performed on the video frames of the task reference video to improve the accuracy of body key point extraction. Among them, the depth information corresponding to the video frames can be collected by a depth camera or an RGB (Red, Green, Blue)-D (Depth) sensor (this sensor combines RGB color information and depth information).

[0117] Optionally, the offline data also includes: multiple task stages corresponding to the valid video segments, the task categories of the valid video segments, and the task actions respectively corresponding to the multiple task stages in the valid video segments. These information can be referred to the description of the foregoing embodiments and will not be elaborated herein. Thus, the richness of the offline data is improved to perform more accurate behavior analysis and task progress tracking on the task execution process of the embodied agent, so as to improve the success rate and accuracy of the embodied agent in performing tasks.

[0118] In traditional automation systems, the execution of tasks depends on preset rules and programming instructions. For complex manufacturing scenarios, it is difficult for preset rules to cope with unstructured or frequently changing environments, resulting in the robot's difficulty in understanding or adapting to the dynamic changes in tasks when collaborating with humans to perform tasks, resulting in low efficiency and accuracy in the task execution process.

[0119] Compared with relying on preset rules and programming instructions to execute tasks, relying on single-modal data to identify and predict task execution behaviors can improve the flexibility of the robot in performing tasks, but this method cannot capture spatio-temporal information and has insufficient recognition and prediction accuracy.

[0120] To solve the above problems, the offline data (valid video segments, task categories of valid video segments, task actions corresponding to multiple task stages in valid video segments, pose data, and / or agent operation data) generated by the foregoing embodiments can be used to verify and / or predict the operations of the embodied agent during the task execution, so that the embodied agent can cope with complex and dynamically changing environments, and improve the efficiency, success rate, and accuracy of the embodied agent in task execution.

[0121] Figure 4 The figure is a flowchart of the verification and / or reference process of the operations of the embodied agent during the task execution in the embodied intelligent data processing method provided by the embodiments of the present application. As Figure 4 shown, taking the offline data including valid video segments as an example, the verification and / or reference process of the operations of the embodied agent during the task execution may include:

[0122] S401, obtain the current task video of the embodied agent.

[0123] In this embodiment, the current task video of the embodied agent can be obtained by acquiring the real-time acquisition data of the sensing device (including the imaging device).

[0124] S402, perform multi-modal feature extraction on the task video to obtain the second visual semantic feature of the task video and the second body pose feature of the embodied agent.

[0125] In a possible implementation manner, S402 includes: inputting the task video into a pre-trained multi-modal large model, and in the multi-modal large model, perform multi-modal feature extraction on the task video to obtain the visual semantic feature of the task video and the body pose feature of the embodied agent. Thus, the accuracy of multi-modal feature extraction is improved through the pre-trained multi-modal large model.

[0126] In a possible implementation manner, input the task video into a pre-trained CNN, and in the CNN, perform feature extraction on the video frames of the task video to obtain the visual semantic feature of the task video. Among them, the pre-training objective of the CNN is to improve the accuracy of the CNN in visual semantic feature extraction.

[0127] In yet another possible implementation manner, divide the video frames of the task video into multiple image patches, input the multiple image patches into the ViT network, and in the ViT network, extract the visual semantic feature by learning the relationship between the multiple image patches. Thus, the global visual semantic feature is extracted by using the ViT network, and the accuracy of visual semantic feature extraction is improved.

[0128] Among them, the implementation principles and technical effects of S402 and the above implementation manners can refer to the description of multi-modal feature extraction of the task reference video in the foregoing embodiments, and will not be elaborated here.

[0129] S403. Fuse the second visual semantic feature of the task video and the second body pose feature of the embodied agent to obtain a second fusion feature.

[0130] In a possible implementation, S403 includes: inputting the second visual semantic feature of the task video and the second body pose feature of the embodied agent into a deep learning model using a multi-head attention mechanism, and in the deep learning model, fusing the second visual semantic feature and the second body pose feature to obtain a second fusion feature.

[0131] Among them, the implementation principle and technical effect of S403 and the above implementation can refer to the description of multi-modal feature extraction of the task reference video in the foregoing embodiments, and will not be elaborated here.

[0132] S404. According to the second fusion feature, perform task recognition on the task video and behavior intention recognition on the embodied agent to obtain the current task information of the embodied agent and the current behavior intention information of the embodied agent.

[0133] In this embodiment, since the second fusion feature is obtained by multi-modal feature fusion, it can not only reflect the visual semantics of the task video but also reflect the body pose of the embodied agent in the task video. Based on the second fusion feature for feature recognition, task recognition of the task video and behavior intention recognition of the embodied agent can be realized, and the current task information and behavior intention information of the embodied agent can be obtained.

[0134] In a possible implementation, in the deep learning model, according to the second fusion feature, perform task recognition on the task video and behavior intention recognition on the embodied agent to obtain the output data of the deep learning model. Among them, the output data of the deep learning model includes the confidence of the candidate task category and the confidence of the candidate behavior intention category of the embodied agent at the current time; according to the confidence of the candidate task category, determine the task category of the task video; according to the confidence of the candidate behavior intention category of the embodied agent at the current time, determine the behavior intention category of the embodied agent at the current time.

[0135] In this embodiment, according to the confidence corresponding to the candidate task category, select the candidate task category with the highest confidence or the confidence greater than the first threshold, and determine this candidate task category as the task category of the task video; among the candidate behavior intention categories corresponding to the current time point, select the candidate behavior intention category with the highest confidence or the confidence greater than the second threshold, and determine this candidate behavior intention category as the behavior intention category of the embodied agent at this time point. Thus, the recognition accuracy of the task category and the behavior intention category is improved by using multi-modal features and the deep learning model.

[0136] S405. Verify the correctness of the current task action of the embodied agent, predict the movement path of the embodied agent, and / or predict the next task action of the embodied agent based on the valid video segments of the task reference video, the current task information of the embodied agent, and the current behavioral intention information of the embodied agent.

[0137] In this embodiment, the valid video segments reflect the task execution process in video form. The correctness of the current task action of the embodied agent can be verified, the movement path of the embodied agent can be predicted, and / or the next task action of the embodied agent can be predicted by comparing or matching the valid video segments of the task reference video with the current task information of the embodied agent and the current behavioral intention information of the embodied agent.

[0138] In a possible implementation, the current task information of the embodied agent includes the current task category and / or task phase being executed by the embodied agent, the current behavioral intention information of the embodied agent includes the current behavioral intention category of the embodied agent, and the valid video segments are marked with the task category and / or task phase. Based on this, verifying the correctness of the current task action of the embodied agent according to the valid video segments, the current task information of the embodied agent, and the current behavioral intention information of the embodied agent includes at least one of the following implementation methods:

[0139] (1) Match the task category marked in the valid video segments with the current task category being executed by the embodied agent to determine whether the current task being executed by the embodied agent is correct.

[0140] In this implementation method, if the task category marked in the valid video segments is consistent with the current task category being executed by the embodied agent, it is determined that the current task being executed by the embodied agent is correct; otherwise, it is determined that the current task being executed by the embodied agent is incorrect, thus realizing the verification of the current task being executed by the embodied agent.

[0141] (2) Match the task phase marked in the valid video segments with the current task phase being executed by the embodied agent according to the current time and the time information corresponding to the task phase marked in the valid video segments to determine whether the current task phase being executed by the embodied agent is correct.

[0142] In this implementation method, according to the current time and the time information corresponding to the task phase marked in the valid video segments, the task phase corresponding to the current time is determined among the task phases marked in the valid video segments. If the task phase corresponding to the current time is consistent with the current task phase being executed by the embodied agent, it is determined that the current task phase being executed by the embodied agent is correct; otherwise, it is determined that the current task phase being executed by the embodied agent is incorrect, thus realizing the verification of the current task phase being executed by the embodied agent.

[0143] (3) Determine the task phase corresponding to the current time based on the current time and the time information corresponding to the marked task phases in the valid video segment. Match the task actions of the task phase corresponding to the current time with the current behavior intention category of the embodied agent to determine whether the current behavior intention category of the embodied agent is correct.

[0144] In this implementation, if the task actions of the task phase corresponding to the current time are consistent with the current behavior intention category of the embodied agent, it is determined that the current behavior intention category of the embodied agent is correct; otherwise, it is determined that the current behavior intention category of the embodied agent is incorrect, thus realizing the verification of the current behavior intention category of the embodied agent.

[0145] Optionally, in the above implementation, if the task currently executed by the embodied agent is incorrect, the task phase currently executed by the embodied agent is incorrect, or the current behavior intention category of the embodied agent is incorrect, an exception reminder message is output or an exception is marked (for example, marking the behavior intention category of the embodied agent as abnormal) to remind relevant personnel to check the situation of the embodied agent in a timely manner.

[0146] For example, according to the valid video segment, it is determined that the embodied agent should perform a "grasp" action at the current time. If it is detected in real time that the behavior of the embodied agent is "adjustment", an exception alarm will be triggered.

[0147] Another example, according to the valid video segment, it is determined that the current time should be in the task completion phase. If it is detected in real time that the embodied agent is still adjusting the object, the behavior of adjusting the object is marked as an abnormal behavior.

[0148] In a possible implementation, the current task information of the embodied agent includes the task category and / or task phase currently executed by the embodied agent, the current behavior intention information of the embodied agent includes the current behavior intention category of the embodied agent, and the valid video segment is marked with the task category and / or task phase; based on the valid video segment, the current task information of the embodied agent, and the current behavior intention information of the embodied agent, predict the movement path of the embodied agent and / or predict the next task action of the embodied agent, including: matching the task category currently executed by the embodied agent with the task category marked in the valid video segment, matching the current behavior intention category of the embodied agent with the task actions of the task phase marked in the valid video segment, and determining the task phase currently executed by the embodied agent; predicting the movement path of the embodied agent and / or predicting the next task action of the embodied agent based on the task phase currently executed by the embodied agent and the task actions of the task phase marked in the valid video segment. Thus, the guidance of the movement path and / or the next task action of the embodied agent is realized, and the accuracy of the movement path and / or the next task action of the embodied agent is improved.

[0149] In this implementation manner, before predicting the movement path, the map information of the working area of the embodied agent can be established based on the environmental information collected by sensors (such as lidar, camera devices). The position information of task-related objects, the position information of the embodied agent, and the address information of obstacles are marked in the map information. Based on the map information, a path planning algorithm and a deep learning algorithm are used for path planning to obtain the optimal path corresponding to the embodied agent. Then, according to the above process, the optimal path is adjusted, that is: the task category currently executed by the embodied agent is matched with the task category marked in the valid video segment, and the current behavior intention category of the embodied agent is matched with the task actions in the task stage marked in the valid video segment to determine the task stage currently executed by the embodied agent. According to the task stage currently executed by the embodied agent and the task actions in the task stage marked in the valid video segment, the next N movement positions of the embodied agent are obtained, where N is greater than or equal to 1. According to the next N movement positions of the embodied agent, the optimal path of the embodied agent is adjusted. Thus, the movement path of the embodied agent is optimized, and the efficiency and accuracy of the embodied agent in executing tasks are improved.

[0150] In this implementation manner, the task category currently executed by the embodied agent is matched with the task category marked in the valid video segment, and the current behavior intention category of the embodied agent is matched with the task actions in the task stage marked in the valid video segment to determine the task stage currently executed by the embodied agent. According to the task stage currently executed by the embodied agent, the operation instructions corresponding to this task stage are determined based on the task database. For example, the "grasping stage" corresponds to operation instructions such as "adjust the position of the end effector of the robot" and "open the gripper"; the operation instructions corresponding to this task stage are generated to control the embodied agent to execute the next task action according to this operation instruction. Thus, the accuracy of the embodied agent in executing task actions is effectively improved.

[0151] In a possible implementation manner, based on any of the above embodiments, the method further includes the steps of:

[0152] S206, verifying the correctness of the task execution of the embodied agent according to the above valid video segment to obtain a verification result. The above verification result is correct or incorrect.

[0153] S207, when the verification result of the embodied agent is incorrect, identifying the hand movement of the user in the video according to the valid video segment.

[0154] S208, when the hand movement of the user meets the first preset condition, obtaining the corresponding time of the video when the hand movement of the user occurs.

[0155] S209. When the corresponding time of the video meets the preset time period, obtain the attention information associated with the task execution object in the task.

[0156] S210. Determine and execute a new target task for the embodied intelligent agent according to the attention information.

[0157] Specifically, the user can be a teaching person. For example, in the effective video segment, the user issues an instruction: fetch water. In the effective video segment, the teaching action is originally to fetch drinking water. The embodied intelligent agent may observe that the user is in the bathroom and then fetches tap water instead of drinking water. Then this action execution is incorrect. In this case, identify the user's hand movement. The first preset condition can be, for example, that the hand has just taken medicine. That is, when the user has just taken medicine, generally speaking, the user is about to take the medicine, so drinking water is needed. In this case, obtain the corresponding time of the video when the user's hand movement, such as the medicine-taking movement, occurs. This time is the natural time (that is, timed in 24 hours a day), not the video playback duration. The above task execution object can be, for example, medicine. The above preset time period corresponds to the user's hand movement. For example, when the hand movement is a medicine-taking movement, the preset time period can be the user's medicine-taking time period within the last 3 days. If the corresponding time of the video is within this preset time period, then combined with the aforementioned user's hand movement, that is, based on two factors of time and the user's hand movement, comprehensively determine the user's action intention, which makes the accuracy of task determination higher. In this example, it is determined that the user's action intention is to take medicine, so the next target task is to fetch drinking water.

[0158] The above attention information can be, for example, the water temperature for taking medicine determined according to the corresponding medicine instruction manual or the doctor's written instructions on the electronic medical record. And determine the final water fetching task according to this water temperature.

[0159] In some alternative embodiments, the above step S210 includes:

[0160] Identify the user's action intention according to the user's hand movement.

[0161] Obtain the user preference information associated with the action intention.

[0162] Determine a new target task for the embodied intelligent agent according to the attention information and the user preference information.

[0163] Exemplarily, for example, if the user's habit (that is, the user preference information) is to rinse the mouth with water after taking the medicine, then the embodied intelligent agent will divide the fetched water into portions, one portion for taking the medicine and the other portion for rinsing the mouth after the user takes the medicine. Then the determined new target task will be, for example, to fetch water first and then divide the water.

[0164] The following is an embodiment of the apparatus of the present application, which can be used to execute the method embodiment of the present application. For details not disclosed in the apparatus embodiment of the present application, please refer to the method embodiment of the present application.

[0165] Figure 5 It is a schematic structural diagram of an embodied intelligent data processing apparatus provided by an embodiment of the present application. As Figure 5 shown, the embodied intelligent data processing apparatus 500 of the embodiment of the present application includes: a video data acquisition unit 501, a multimodal feature extraction unit 502, a multimodal feature fusion unit 503, a task behavior recognition unit 504, and a valid data determination unit 505. Among them:

[0166] The video data acquisition unit 501 is used to acquire a task reference video of the embodied intelligent agent, and the task reference video includes the behavior information of the task execution object and the state information of the task-related object; the multimodal feature extraction unit 502 is used to perform multimodal feature extraction on the task reference video to obtain the visual semantic feature and the body posture feature of the task reference video; the multimodal feature fusion unit 503 is used to fuse the visual semantic feature and the body posture feature to obtain a first fusion feature; the task behavior recognition unit 504 is used to perform task recognition on the task reference video and behavior intention recognition on the task execution object according to the first fusion feature to obtain a recognition result; the valid data determination unit 505 is used to determine a valid video segment of the task reference video according to the recognition result, and the valid video segment is used for verification and / or reference of the operations during the task execution process of the embodied intelligent agent.

[0167] In some embodiments, the multimodal feature fusion unit 503 is specifically used to: input the visual semantic feature and the body posture feature into a deep learning model using a multi-head attention mechanism, and in the deep learning model, fuse the visual semantic feature and the body posture feature to obtain a first fusion feature.

[0168] In some embodiments, the recognition result includes the task category of the task reference video and the behavior intention categories of the task execution object at multiple time points. The task behavior recognition unit 504 is specifically used to: in the deep learning model, perform task recognition on the task reference video and behavior intention recognition on the task execution object according to the first fusion feature to obtain the output data of the deep learning model, where the output data includes the confidence of the candidate task category and the confidence of the candidate behavior intention categories corresponding to multiple time points respectively; determine the task category of the task reference video according to the confidence of the candidate task category; determine the behavior intention categories of the task execution object at multiple time points according to the confidence of the candidate behavior intention categories corresponding to multiple time points respectively.

[0169] In some embodiments, the recognition result includes the task category of the task reference video and the behavior intention categories of the task execution object at multiple time points. The valid video segment includes the video segment from the start to the end of the task. The valid data determination unit 505 is specifically configured to: according to the task category, obtain the task actions corresponding to multiple task stages respectively. The multiple task stages include the task start stage, the task execution stage, and the task end stage; match the task actions corresponding to the multiple task stages respectively with the behavior intention categories of the task execution object at multiple time points to obtain the time information corresponding to the multiple task stages respectively; determine the valid video segment in the task reference video according to the time information corresponding to the multiple task stages respectively and the time information corresponding to the video frames in the task reference video; according to the time information corresponding to the multiple task stages respectively and the time information corresponding to the video frames in the task reference video, the embodied intelligent data processing device further includes: a task stage marking unit (not shown in the figure), configured to mark the multiple task stages of the valid video segment; and / or, a task category marking unit (not shown in the figure), configured to mark the task category of the valid video segment; and / or, a behavior intention storage unit (not shown in the figure), configured to store the task actions corresponding to the multiple task stages in the valid video segment.

[0170] In some embodiments, the embodied intelligent data processing device further includes: a key point extraction unit (not shown in the figure), configured to extract the body key points of the task execution object in the task reference video to obtain the pose data of the task execution object. The pose data includes the sequence of body key points of the task execution object sorted by time; a time alignment unit (not shown in the figure), configured to perform time alignment on the pose data, the agent operation data, and the valid video segment to obtain the offline data. The offline data is used for the verification and / or reference of the operations during the task execution of the embodied intelligent agent. The agent operation data is the sensing data, the motion trajectory data, and / or the control instruction data of the robot participating in the task cooperation in the task reference video.

[0171] In some embodiments, the process of verifying and / or referencing the operations during the task execution of the embodied intelligent agent includes: obtaining the current task video of the embodied intelligent agent; performing multi-modal feature extraction on the task video to obtain the visual semantic feature and the body pose feature of the task video; fusing the visual semantic feature and the body pose feature to obtain the second fusion feature; according to the second fusion feature, performing task recognition on the task video and behavior intention recognition on the embodied intelligent agent to obtain the current task information of the embodied intelligent agent and the current behavior intention information of the embodied intelligent agent; according to the valid video segment, the current task information of the embodied intelligent agent, and the current behavior intention information of the embodied intelligent agent, verifying the correctness of the current task actions of the embodied intelligent agent, predicting the movement path of the embodied intelligent agent, and / or predicting the next task actions of the embodied intelligent agent.

[0172] In some embodiments, the current task information of the embodied agent includes the task category and / or task stage currently executed by the embodied agent, the current behavior intention information of the embodied agent includes the current behavior intention category of the embodied agent, and the valid video segment is marked with the task category and / or task stage; according to the valid video segment, the current task information of the embodied agent, and the current behavior intention information of the embodied agent, verify the correctness of the current task action of the embodied agent, including at least one of the following: match the task category marked in the valid video segment with the task category currently executed by the embodied agent to determine whether the task currently executed by the embodied agent is correct; according to the current time and the time information corresponding to the task stage marked in the valid video segment, match the task stage marked in the valid video segment with the task stage currently executed by the embodied agent to determine whether the task stage currently executed by the embodied agent is correct; according to the current time and the time information corresponding to the task stage marked in the valid video segment, determine the task stage corresponding to the current time, and match the task action of the task stage corresponding to the current time with the current behavior intention category of the embodied agent to determine whether the current behavior intention category of the embodied agent is correct.

[0173] In some embodiments, the current task information of the embodied agent includes the task category and / or task stage currently executed by the embodied agent, the current behavior intention information of the embodied agent includes the current behavior intention category of the embodied agent, and the valid video segment is marked with the task category and / or task stage; according to the valid video segment, the current task information of the embodied agent, and the current behavior intention information of the embodied agent, predict the movement path of the embodied agent and / or predict the next task action of the embodied agent, including: match the task category currently executed by the embodied agent with the task category marked in the valid video segment, match the current behavior intention category of the embodied agent with the task action of the task stage marked in the valid video segment, and determine the task stage currently executed by the embodied agent; according to the task stage currently executed by the embodied agent and the task action of the task stage marked in the valid video segment, predict the movement path of the embodied agent and / or predict the next task action of the embodied agent.

[0174] The device of this embodiment can be used to execute the technical solutions of any of the above - shown method embodiments, and its implementation principle and technical effects are similar, so they will not be elaborated here.

[0175] Figure 6 It is a schematic structural diagram of an electronic device provided in an embodiment of the present application. Exemplarily, the electronic device can be provided as a server or a computer. Refer to Figure 6, the electronic device 600 includes a processing component 601, which further includes one or more processors, and memory resources represented by a memory 602 for storing instructions executable by the processing component 601, such as application programs. The application programs stored in the memory 602 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 601 is configured to execute instructions to perform any of the above method embodiments.

[0176] The electronic device 600 may further include a power supply component 603 configured to perform power management of the electronic device 600, a wired or wireless network interface 604 configured to connect the electronic device 600 to a network, and an input / output (I / O) interface 605. The electronic device 600 may operate based on an operating system stored in the memory 602, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSD TM or the like.

[0177] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the solution of the above-mentioned embodied intelligent data processing method.

[0178] This application also provides a computer program product including a computer program, which, when executed by a processor, implements the solution of the above-mentioned embodied intelligent data processing method.

[0179] The above-mentioned computer-readable storage medium may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disk. The readable storage medium may be any available medium accessible by a general-purpose or special-purpose computer.

[0180] An exemplary readable storage medium is coupled to the processor so that the processor can read information from and write information to the readable storage medium. Of course, the readable storage medium may also be a component of the processor. The processor and the readable storage medium may be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium may also exist as discrete components in the embodied intelligent data processing device.

[0181] Those of ordinary skill in the art will understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps included in the above method embodiments; and the aforementioned storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0182] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. An embodied intelligent data processing method, characterized in that: include: Obtaining a task reference video of the embodied intelligent agent, wherein the task reference video contains behavior information of a task execution object and state information of task-related objects; Performing multimodal feature extraction on the task reference video to obtain visual semantic features of the task reference video and body posture features of the task execution object; fusing the visual semantic feature and the body posture feature to obtain a first fused feature; According to the first fusion feature, performing task recognition on the task reference video and performing behavior intention recognition on the task execution object to obtain a recognition result; Based on the recognition result, a valid video segment of the task reference video is determined, and the valid video segment is used for verification and / or reference of operations performed by the embodied intelligent agent during the execution of the task.

2. The embodied intelligence data processing method according to claim 1, characterized in that: The recognition result includes the task category of the task reference video and the behavior intention category of the task execution object at multiple time points. The task recognition is performed on the task reference video and the behavior intention is performed on the task execution object according to the first fusion feature to obtain the recognition result, including: In a deep learning model using a multi-head attention mechanism, according to the first fusion feature, task identification is performed on the task reference video and behavior intention identification is performed on the task execution object to obtain output data of the deep learning model, wherein the output data includes confidences of candidate task categories and confidences of candidate behavior intention categories corresponding to the multiple time points respectively; Determining the task category of the task reference video according to the confidence of the candidate task category; The behavior intention categories of the task execution object at the multiple time points are determined according to the confidences of the candidate behavior intention categories respectively corresponding to the multiple time points.

3. The embodied intelligence data processing method according to claim 1, characterized in that: The recognition result includes the task category of the task reference video and the behavioral intention category of the task execution object at multiple time points, the valid video segment includes the video segment from the beginning to the end of the task, and determining the valid video segment of the task reference video according to the recognition result includes: According to the task category, task actions corresponding to a plurality of task stages are obtained, wherein the plurality of task stages include a task start stage, a task progress stage, and a task end stage; Matching the task actions corresponding to the multiple task stages respectively with the behavioral intention categories of the task execution object at multiple time points to obtain the time information corresponding to the multiple task stages respectively; Determining the valid video segment in the task reference video according to the time information respectively corresponding to the multiple task stages and the time information corresponding to the video frames in the task reference video; Wherein, after determining the valid video segment in the task reference video according to the time information respectively corresponding to the multiple task stages and the time information corresponding to the video frames in the task reference video, the method further includes: Marking the multiple task stages on the valid video clips; and / or, marking the task category of the valid video clip; And / or, save the task actions corresponding to multiple task stages in the effective video clip.

4. The embodied intelligent data processing method according to any one of claims 1 to 3, characterized in that: After obtaining the task reference video corresponding to the embodied intelligent agent, the method further includes: In the task reference video, body key points of the task execution object are extracted to obtain posture data of the task execution object, wherein the posture data includes a body key point sequence of the task execution object sorted in time; The posture data, agent operation data and the valid video clip are time-aligned to obtain offline data, and the offline data is used to verify and / or refer to the operations of the embodied agent during the execution of the task. The agent operation data includes sensor data, motion trajectory data and / or control instruction data of the robot participating in the task collaboration in the task reference video.

5. The embodied intelligent data processing method according to any one of claims 1 to 3, characterized in that: The verification and / or reference process of the operation of the embodied intelligent agent during the execution of the task includes: Obtaining a current task video of the embodied intelligent agent; Performing multimodal feature extraction on the task video to obtain a second visual semantic feature of the task video and a second body posture feature of the embodied intelligent agent; fusing the second visual semantic feature and the second body posture feature to obtain a second fused feature; According to the second fusion feature, performing task recognition on the task video and behavioral intention recognition on the embodied intelligent agent, to obtain current task information of the embodied intelligent agent and current behavioral intention information of the embodied intelligent agent; Based on the valid video clip, the current task information of the embodied intelligent agent and the current behavioral intention information of the embodied intelligent agent, the correctness of the current task action of the embodied intelligent agent is verified, the moving path of the embodied intelligent agent is predicted and / or the next task action of the embodied intelligent agent is predicted.

6. The embodied intelligence data processing method according to claim 5, characterized in that: The current task information of the embodied intelligent agent includes the task category and / or task stage currently performed by the embodied intelligent agent, the current behavior intention information of the embodied intelligent agent includes the current behavior intention category of the embodied intelligent agent, and the valid video clip is marked with the task category and / or task stage; Verifying the correctness of the current task action of the embodied intelligent agent according to the valid video clip, the current task information of the embodied intelligent agent, and the current behavior intention information of the embodied intelligent agent includes at least one of the following: Matching the task category marked by the valid video clip with the task category currently performed by the embodied intelligent agent to determine whether the task currently performed by the embodied intelligent agent is correct; According to the time information corresponding to the current time and the task stage marked in the valid video segment, the task stage marked in the valid video segment is matched with the task stage currently executed by the embodied intelligent agent to determine whether the task stage currently executed by the embodied intelligent agent is correct; According to the current time and the time information corresponding to the task stage marked in the valid video clip, the task stage corresponding to the current time is determined, and the task action of the task stage corresponding to the current time is matched with the current behavior intention category of the embodied intelligent agent to determine whether the current behavior intention category of the embodied intelligent agent is correct.

7. The embodied intelligence data processing method according to claim 5, characterized in that: The current task information of the embodied intelligent agent includes the task category and / or task stage currently performed by the embodied intelligent agent, the current behavior intention information of the embodied intelligent agent includes the current behavior intention category of the embodied intelligent agent, and the valid video clip is marked with the task category and / or task stage; According to the valid video clip, the current task information of the embodied intelligent agent and the current behavior intention information of the embodied intelligent agent, the moving path of the embodied intelligent agent and / or the next task action of the embodied intelligent agent are predicted, including: Matching the task category currently performed by the embodied intelligent agent with the task category marked by the valid video clip, matching the current behavior intention category of the embodied intelligent agent with the task action of the task stage marked by the valid video clip, and determining the task stage currently performed by the embodied intelligent agent; According to the task stage currently being performed by the embodied intelligent agent and the task action of the task stage marked in the effective video clip, the movement path of the embodied intelligent agent is predicted and / or the next task action of the embodied intelligent agent is predicted.

8. The embodied intelligence data processing method according to claim 1, characterized in that: The method further comprises: Verifying the correctness of the task execution of the embodied intelligent agent according to the valid video clip to obtain a verification result; the verification result is correct or wrong; In a case where the verification result of the embodied intelligent agent is wrong, identifying a user's hand motion in the video according to the valid video clip; When the user's hand movement satisfies a first preset condition, obtaining a video corresponding time when the user's hand movement occurs; When the time corresponding to the video meets the preset time period, obtaining the precaution information associated with the task execution object in the task; Based on the precaution information, a new target task of the embodied intelligent body is determined and executed.

9. The embodied intelligence data processing method according to claim 8, characterized in that: Determining and executing a new target task of the embodied intelligent agent according to the precaution information includes: According to the hand movement of the user, identifying the movement intention of the user; Acquire user preference information associated with the action intention; Determine a new target task for the embodied intelligent agent based on the precaution information and the user preference information.

10. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the embodied intelligent data processing method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Man-machine cooperation method and system based on multi-modal behavior online prediction

    CN113524175A

  • Motion capture method and device, electronic equipment, and storage medium

    CN114078279A

  • Video clip positioning method and system, control device and readable storage medium

    CN114896451A

  • Crown block scheduling method, device and equipment and storage medium

    CN117764293A

  • Brain-computer fusion enhanced visual navigation method, electronic equipment and storage medium

    CN118840515A

Cited By

  • Personal intelligent multi-source data quality evaluation and verification method, device, medium and product

    CN121188440A

  • Body intelligence multi-source data quality evaluation and verification method, device, medium and product

    CN121188440B

  • Data construction method, equipment, medium and product for intelligent training

    CN122391793A

  • Data construction method, device, medium and product for embodied intelligence training

    CN122391793B

  • Task instruction hierarchical reconstruction-based video generation model fine tuning method

    CN122496691A