Body-aware data processing method and apparatus
By extracting and fusing multimodal features from the task reference videos of embodied intelligent agents, and using a deep learning model with a multi-head attention mechanism for task recognition and behavioral intent recognition, the effective video segments are automatically determined, solving the problems of high labor costs and low efficiency in embodied intelligent data processing, and achieving efficient and accurate data processing.
Patent Information
- Application Number
- CN202510119577.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-01-24
AI Technical Summary
Embodied intelligence data processing suffers from high labor costs and low efficiency, with existing technologies requiring extensive human intervention for data collection and processing.
Multimodal feature extraction and fusion techniques are employed to process the task reference video of the embodied intelligent agent. A deep learning model with multi-head attention mechanism is used for task recognition and behavioral intent recognition to automatically determine the effective video segments of the task reference video.
It improves the accuracy of task recognition and behavioral intent recognition, enables effective video clip acquisition without human intervention, reduces labor costs, and improves data processing efficiency and accuracy.
Smart Images

Figure CN120147918B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of embodied intelligence technology, and in particular to an embodied intelligence data processing method and device. Background Technology
[0002] With the rapid development of industrial automation and artificial intelligence technologies, human-machine collaboration through embodied intelligent agents has become one of the important research directions in intelligent manufacturing.
[0003] In related technologies, the collection of task data for embodied intelligent agents requires a large amount of human intervention: one person is responsible for teaching the embodied intelligent agent, while another person is responsible for data collection and processing.
[0004] The above methods increase labor costs and have low data processing efficiency. Summary of the Invention
[0005] This application provides a method and apparatus for embodied intelligence data processing to solve the problems of high labor costs and low efficiency in embodied intelligence data processing.
[0006] In a first aspect, this application provides an embodied intelligence data processing method, comprising: acquiring a task reference video of an embodied intelligent agent, the task reference video containing behavioral information of the task execution object and state information of task-related objects; performing multimodal feature extraction on the task reference video to obtain visual semantic features and body posture features of the task execution object; fusing the visual semantic features and body posture features to obtain a first fused feature; performing task recognition on the task reference video and behavioral intent recognition on the task execution object based on the first fused feature to obtain a recognition result; and determining valid video segments of the task reference video based on the recognition result, the valid video segments being used for verification and / or reference of operations performed by the embodied intelligent agent during task execution.
[0007] In some embodiments, fusing visual semantic features and body posture features to obtain a first fused feature includes: inputting visual semantic features and body posture features into a deep learning model employing a multi-head attention mechanism, and fusing the visual semantic features and body posture features in the deep learning model to obtain the first fused feature.
[0008] In some embodiments, the identification result includes the task category of the task reference video and the behavioral intent category of the task execution object at multiple time points. Based on a first fusion feature, task identification is performed on the task reference video and behavioral intent identification is performed on the task execution object to obtain the identification result. This includes: in a deep learning model employing a multi-head attention mechanism, task identification is performed on the task reference video and behavioral intent identification is performed on the task execution object based on the first fusion feature to obtain output data of the deep learning model. The output data includes the confidence scores of candidate task categories and the confidence scores of candidate behavioral intent categories corresponding to multiple time points. Based on the confidence scores of candidate task categories, the task category of the task reference video is determined. Based on the confidence scores of candidate behavioral intent categories corresponding to multiple time points, the behavioral intent category of the task execution object at multiple time points is determined.
[0009] In some embodiments, the identification result includes the task category of the task reference video and the behavioral intent category of the task execution object at multiple time points. Valid video segments include video segments from the start to the end of the task. Determining valid video segments of the task reference video based on the identification result includes: obtaining task actions corresponding to multiple task stages according to the task category, where the multiple task stages include a task start stage, a task progress stage, and a task end stage; matching the task actions corresponding to the multiple task stages with the behavioral intent categories of the task execution object at multiple time points to obtain time information corresponding to the multiple task stages; determining valid video segments in the task reference video based on the time information corresponding to the multiple task stages and the time information corresponding to video frames in the task reference video; after determining valid video segments in the task reference video based on the time information corresponding to the multiple task stages and the time information corresponding to video frames in the task reference video, the method further includes: marking the valid video segments with multiple task stages; and / or marking the task category of the valid video segments; and / or saving the task actions corresponding to the multiple task stages in the valid video segments.
[0010] In some embodiments, after obtaining the task reference video corresponding to the embodied agent, the method further includes: extracting the body key points of the task execution object in the task reference video to obtain the posture data of the task execution object, the posture data including the sequence of body key points of the task execution object sorted by time; aligning the posture data, agent operation data and effective video segments in time to obtain offline data, the offline data being used for verification and / or reference of the operations of the embodied agent in the process of performing the task, the agent operation data including the sensor data, motion trajectory data and / or control command data of the robot participating in the task collaboration in the task reference video.
[0011] In some embodiments, the verification and / or reference process for the actions performed by the embodied agent during task execution includes: acquiring the current task video of the embodied agent; extracting multimodal features from the task video to obtain a second visual semantic feature of the task video and a second body posture feature of the embodied agent; fusing the second visual semantic feature and the second body posture feature to obtain a second fused feature; performing task recognition on the task video and behavioral intent recognition on the embodied agent based on the second fused feature to obtain the current task information and current behavioral intent information of the embodied agent; and verifying the correctness of the current task action of the embodied agent, predicting the movement path of the embodied agent, and / or predicting the next task action of the embodied agent based on the valid video segment, the current task information of the embodied agent, and the current behavioral intent information of the embodied agent.
[0012] In some embodiments, the current task information of the embodied agent includes the task category and / or task stage currently being performed by the embodied agent, and the current behavioral intent information of the embodied agent includes the current behavioral intent category of the embodied agent. A valid video segment is marked with a task category and / or task stage. The correctness of the current task action of the embodied agent is verified based on the valid video segment, the current task information, and the current behavioral intent information, including at least one of the following: matching the task category marked in the valid video segment with the task category currently being performed by the embodied agent to determine whether the task currently being performed by the embodied agent is correct; matching the task stage marked in the valid video segment with the task stage currently being performed by the embodied agent according to the current time and the time information corresponding to the task stage marked in the valid video segment to determine whether the task stage currently being performed by the embodied agent is correct; determining the task stage corresponding to the current time according to the current time and the time information corresponding to the task stage marked in the valid video segment, and matching the task action of the task stage corresponding to the current time with the current behavioral intent category of the embodied agent to determine whether the current behavioral intent category of the embodied agent is correct.
[0013] In some embodiments, the current task information of the embodied agent includes the task category and / or task stage currently being performed by the embodied agent, and the current behavioral intention information of the embodied agent includes the current behavioral intention category of the embodied agent. The effective video segment is marked with the task category and / or task stage. Based on the effective video segment, the current task information of the embodied agent, and the current behavioral intention information of the embodied agent, the movement path of the embodied agent and / or the next task action of the embodied agent are predicted, including: matching the task category currently being performed by the embodied agent with the task category marked in the effective video segment; matching the current behavioral intention category of the embodied agent with the task action of the task stage marked in the effective video segment to determine the current task stage being performed by the embodied agent; and predicting the movement path of the embodied agent and / or the next task action of the embodied agent based on the current task stage being performed by the embodied agent and the task action of the task stage marked in the effective video segment.
[0014] In some embodiments, the method further includes:
[0015] Based on the valid video clips, the correctness of the embodied intelligent agent's task execution is verified, and a verification result is obtained; the verification result is either correct or incorrect.
[0016] If the verification result of the embodied intelligent agent is incorrect, the user's hand movements in the video are identified based on the valid video segments;
[0017] If the user's hand gesture meets the first preset condition, obtain the video time corresponding to the occurrence of the user's hand gesture.
[0018] If the video time corresponds to a preset time period, obtain the precaution information associated with the task execution object in the task;
[0019] Based on the aforementioned precautions, the embodied intelligent agent determines and executes a new target task.
[0020] In some embodiments, determining and executing a new target task for the embodied agent based on the precaution information includes:
[0021] Based on the user's hand movements, identify the user's intention.
[0022] Obtain user preference information associated with the stated action intent;
[0023] Based on the aforementioned precautions and user preference information, a new target task is determined for the embodied intelligent agent.
[0024] Secondly, this application provides an embodied intelligence data processing device, comprising: a video data acquisition unit for acquiring a task reference video of an embodied intelligent agent, the task reference video containing behavioral information of the task execution object and state information of task-related objects; a multimodal feature extraction unit for extracting multimodal features from the task reference video to obtain visual semantic features of the task reference video and body posture features of the task execution object; a multimodal feature fusion unit for fusing the visual semantic features and body posture features to obtain a first fused feature; a task behavior recognition unit for performing task recognition on the task reference video and behavioral intent recognition on the task execution object based on the first fused feature to obtain a recognition result; and a valid data determination unit for determining valid video segments of the task reference video based on the recognition result, the valid video segments being used for verification and / or reference of operations performed by the embodied intelligent agent during task execution.
[0025] Thirdly, this application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0026] The memory stores instructions that the computer executes;
[0027] The processor executes computer execution instructions stored in memory to implement the embodied intelligence data processing method as described in the first aspect of this application.
[0028] Fourthly, this application provides a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the embodied intelligent data processing method as described in the first aspect of this application.
[0029] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the embodied intelligent data processing method as described in the first aspect of this application.
[0030] The embodied intelligence data processing method and apparatus provided in this application perform multimodal feature extraction on the task reference video of the embodied intelligent agent to obtain the video semantic features of the task reference video and the body posture features of the task execution object. Based on the fusion features of the video semantic features and body posture features, task recognition is performed on the task reference video and behavioral intention recognition is performed on the task execution object. Multimodal technology is used to improve the accuracy of task recognition and behavioral intention recognition. Based on the results of task recognition and behavioral intention recognition, the effective video segments of the task reference video are determined, realizing the automated processing of video data related to embodied intelligence and improving data processing efficiency and accuracy. Attached Figure Description
[0031] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0032] Figure 1 This diagram illustrates an application scenario applicable to the embodiments of this application.
[0033] Figure 2 A flowchart illustrating the embodied intelligence data processing method provided in the embodiments of this application. Figure 1 ;
[0034] Figure 3 A flowchart illustrating the embodied intelligence data processing method provided in the embodiments of this application. Figure 2 ;
[0035] Figure 4 A flowchart illustrating the verification and / or reference process of the operations performed by the embodied intelligent agent during the task execution process in the embodied intelligent data processing method provided in this application embodiment;
[0036] Figure 5 A schematic diagram of the structure of the embodied intelligent data processing device provided in the embodiments of this application;
[0037] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0038] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0039] An embodied intelligent agent is an intelligent system that, during a task, mimics human perception, decision-making, and behavioral abilities by perceiving, understanding, and interacting with the environment. In the embodiments of this application, the embodied intelligent agent can be one or more of the following: humanoid embodied intelligent agent, animal-shaped embodied intelligent agent, multi-legged embodied intelligent agent, wheeled or tracked embodied intelligent agent, flying embodied intelligent agent, modular embodied intelligent agent, etc. In terms of its form, the embodied intelligent agent can be a virtual object, a digital entity, or a machine entity (such as a robot).
[0040] In the process of teaching data collection for embodied intelligent agents, one person usually performs the teaching operation while another person is responsible for data collection and processing, which is inefficient and has high labor costs.
[0041] To address the aforementioned issues, embodiments of this application provide an embodied intelligence data processing method and apparatus. In the embodied intelligence data processing method, multimodal technology is employed to process the task reference video of the embodied intelligent agent, obtaining multimodal features of the task reference video. These multimodal features are then used for task recognition and behavioral intent recognition, improving the accuracy of task recognition and behavioral intent recognition. Based on the results of task recognition and behavioral intent recognition, valid video segments of the task reference video are determined. Thus, without human intervention, the acquisition of valid video segments from the task reference video is completed automatically, improving the accuracy and efficiency of valid video segment acquisition and reducing labor costs.
[0042] The technical means and effects of the devices, equipment and storage media can be referred to the technical means and effects of the above-mentioned schemes, and will not be repeated here.
[0043] Figure 1 The diagram illustrates an application scenario applicable to the embodiments of this application. For example... Figure 1 As shown, the application scenario includes a data processing device 101, a data verification device 102, and an embodied intelligent agent 103.
[0044] In the application scenario, the data processing device 101 can process the task reference video of the embodied intelligent agent 103 through multimodal technology to obtain valid video segments of the task reference video; the data verification device 102 can verify the operation of the embodied intelligent agent 103 in the process of performing the task or provide a reference for the operation of the embodied intelligent agent 103 in the process of performing the task based on the valid video segments and the real-time task video of the embodied intelligent agent.
[0045] The data processing device 101 and the data verification device 102 can be electronic devices, including servers or terminals. Terminals can be personal digital assistant (PDA) devices, handheld devices with wireless communication capabilities (such as smartphones and tablets), computing devices (such as personal computers (PCs)), wearable devices (such as smartwatches and smart bracelets), and smart home devices (such as smart speakers and smart display devices). Servers can be independent servers or server clusters, and can be local servers or cloud servers. Figure 1 Taking a server as an example, the data processing device 101 and the data verification device 102 are used.
[0046] It should be noted that, Figure 1 This is merely a schematic diagram illustrating one application scenario provided by an embodiment of this application. This embodiment of the application does not necessarily represent an application scenario. Figure 1 The included equipment is not limited, nor is it restricted. Figure 1 The positional relationships between the devices are defined.
[0047] Next, we will introduce the embodied intelligence data processing method through specific embodiments.
[0048] Figure 2 A flowchart illustrating the embodied intelligence data processing method provided in the embodiments of this application. Figure 1 .like Figure 2 As shown, the method in this application embodiment includes:
[0049] S201, Obtain the task reference video of the embodied intelligent agent. The task reference video contains behavioral information of the task execution object and state information of the task-related objects.
[0050] Among them, task reference videos, such as task tutorial videos, can be obtained by filming the process of personnel and / or embodied intelligent agents correctly performing tasks (such as grasping tasks or carrying tasks).
[0051] In the task reference video, the task execution objects can be humans and / or embodied intelligent agents. For example, the task in the task reference video is a human-machine collaborative task, and the task execution objects include humans and embodied intelligent agents. Task-related objects refer to objects that need to be handled during the task, such as objects that need to be moved or packaged.
[0052] In this embodiment, the task reference video of the embodied intelligent agent can be obtained from the database, or the task reference video of the embodied intelligent agent input by the user or other devices can be received, or the task reference video collected by the sensing device (such as a camera) can be directly obtained.
[0053] S202, perform multimodal feature extraction on the task reference video to obtain the visual semantic features of the task reference video and the body posture features of the task execution object.
[0054] Among them, visual semantic features refer to the semantic features conveyed by the visual images in the task reference video. In the task reference video, visual semantic features are related to the task execution object and task-related objects, reflecting the characteristics of the task execution object, the characteristics of task-related objects, and the relationship between the task execution object and task-related objects. The body posture features of the task execution object can reflect the actions of the task execution object. When the task execution object is a person, the body posture features are the human posture features.
[0055] In this embodiment, multimodal feature extraction can be performed on video frames in the task reference video (which can be each video frame or selected video frames, such as video frames whose screen content contains the task execution object and / or task-related objects after selection) to obtain the visual semantic features of the reference task video and the body posture features of the task execution object.
[0056] Optionally, the task reference video can be preprocessed before multimodal feature extraction to improve the accuracy of multimodal feature extraction.
[0057] Preprocessing the task reference video may include: denoising the task reference video and / or enhancing the video frames in the task reference video.
[0058] Image enhancement of video frames may include adjusting at least one of the following: color, brightness, contrast, or saturation. For example, the brightness and contrast of video frames in a task reference video may be randomly adjusted to simulate different lighting conditions.
[0059] S203, the visual semantic features and body posture features are fused to obtain the first fused feature.
[0060] In this embodiment, visual semantic features and body posture features can be fused using feature fusion technology to obtain the first fused feature.
[0061] In one possible implementation, the visual semantic features and body posture features are weighted and summed according to weight coefficients to obtain the first fused feature. In this approach, the fusion effect of visual semantic features and body posture features can be improved by adjusting the weight coefficients.
[0062] It should be noted that, in addition to the weighted method, other methods can be used for fusion, and other formulas can be designed to calculate the visual semantic features and body posture features to obtain the first fused feature.
[0063] S204. Based on the first fusion feature, perform task recognition on the task reference video and behavioral intent recognition on the task execution object to obtain the recognition result.
[0064] In this embodiment, since the first fusion feature is obtained by fusing multimodal features, it can reflect both the visual semantics of the task reference video and the body posture of the task execution object in the task reference video. Based on the first fusion feature, feature recognition can be performed to realize task recognition of the task reference video and behavioral intention recognition of the task execution object, and obtain recognition results. The recognition results can include task recognition results and behavioral intention recognition results. The task recognition results indicate the task information corresponding to the task reference video, and the behavioral intention recognition results indicate the behavioral intention corresponding to the task execution object.
[0065] S205, Based on the recognition results, determine the valid video segments of the task reference video. The valid video segments are used for verification and / or reference of the operations performed by the embodied intelligent agent during the task execution process.
[0066] Among them, valid video clips are those that are relevant to the execution of the task.
[0067] In particular, during the execution of tasks by the embodied intelligent agent, the correctness of the task execution can be verified based on valid video clips, and / or, guidance can be provided to the embodied intelligent agent to execute tasks based on valid video clips, that is, reference can be provided.
[0068] In this embodiment, the task reference video may include video segments unrelated to task execution. For example, it could be a video segment recording preparation work before task execution, a video segment recording tasks unrelated to the task execution object during task execution, or a video segment recording tasks related to the task execution object after the task is completed. To accurately collect video data related to task execution, valid video segments of the task reference video are determined based on the recognition results. In determining valid video segments of the task reference video based on the recognition results, the recognition results reflect the task information and behavioral intent of the task execution object. The task stages of the task reference video can be analyzed based on the recognition results to determine the video segments located within the task stages, i.e., the valid video segments of the task reference video.
[0069] In this embodiment, for the task reference video, multimodal features are extracted, and task recognition and behavioral intent recognition are performed based on these features, thereby improving the accuracy of task recognition and behavioral intent recognition. Based on the task recognition results and behavioral intent recognition results, valid video segments are determined from the task reference video, improving the accuracy of obtaining valid video segments from the task reference video. Thus, automatic processing of the task reference video is achieved without manual intervention, improving data processing efficiency and accuracy.
[0070] Figure 3 A flowchart illustrating the embodied intelligence data processing method provided in the embodiments of this application. Figure 2 .
[0071] like Figure 3 As shown, the method in this application embodiment may include:
[0072] S301, Obtain the task reference video of the embodied intelligent agent. The task reference video contains behavioral information of the task execution object and state information of the task-related objects.
[0073] S302, perform multimodal feature extraction on the task reference video to obtain the visual semantic features of the task reference video and the body posture features of the task execution object.
[0074] The implementation principles and technical effects of S301 to S302 can be referred to in the aforementioned embodiments, and will not be repeated here.
[0075] In one possible implementation, a task reference video is input into a pre-trained multimodal large-scale model. Within this model, multimodal feature extraction is performed on the task reference video to obtain its visual semantic features and the body pose features of the task-performing object. Thus, the accuracy of multimodal feature extraction is improved through the pre-trained multimodal large-scale model.
[0076] In one possible implementation, the task reference video can be input into a pre-trained convolutional neural network (CNN). The CNN extracts features from the video frames of the task reference video to obtain its visual semantic features. The goal of pre-training the CNN is to improve the accuracy of its visual semantic feature extraction.
[0077] Optionally, the pre-trained CNN may employ at least one of the following: a residual network (ResNet), a visual geometry group network, or an Inception network. These networks can be directly used for image feature extraction.
[0078] In another possible implementation, the video frames of the task reference video are divided into multiple image blocks. These image blocks are then input into a vision transformer (ViT) network. The ViT network learns the relationships between these image blocks to extract visual semantic features. Thus, the ViT network is used to extract global visual semantic features, improving the accuracy of visual semantic feature extraction.
[0079] S303 inputs visual semantic features and body posture features into a deep learning model using a multi-head attention mechanism. In the deep learning model, the visual semantic features and body posture features are fused to obtain the first fused feature.
[0080] In this embodiment, visual semantic features and body posture features are input into a deep learning model employing a multi-head attention mechanism. In the deep learning model, the multi-head attention mechanism can adaptively weight features of different modalities, that is, adaptively weight visual semantic features and body posture features, explore the relationship between visual semantic features and body posture features, capture the complex spatiotemporal dependence between visual semantic features and body posture features, and obtain more accurate fused features.
[0081] Optionally, the deep learning model employing a multi-head attention mechanism is the Transformer architecture. The Transformer architecture is a neural network model based on a self-attention mechanism, incorporating both self-attention and multi-head attention mechanisms. These mechanisms effectively improve the processing performance for multimodal features.
[0082] S304, In the deep learning model, based on the first fusion feature, task recognition is performed on the task reference video and behavioral intent recognition is performed on the task execution object to obtain the output data of the deep learning model. The output data of the deep learning model includes the confidence of the candidate task category and the confidence of the candidate behavioral intent category corresponding to multiple time points.
[0083] Among them, multiple time points can be multiple timestamps on the timeline of the task reference video.
[0084] There can be multiple candidate task categories, and the output data of the deep learning model can include the confidence scores corresponding to each of these candidate task categories. At multiple time points, each time point can correspond to multiple candidate behavioral intent categories, and each candidate behavioral intent category corresponds to its own confidence score.
[0085] Understandably, the output data of a deep learning model store includes a task recognition result and behavioral intent recognition results corresponding to multiple time points. The task recognition result includes the confidence scores corresponding to multiple candidate task categories, and the behavioral intent recognition results corresponding to a time point include the confidence scores corresponding to multiple candidate behavioral intent categories at that time point.
[0086] As an example, the output data of the deep learning model includes: a single task category identification result, such as {"assembly task": 0.9, "transportation task": 0.1}, where "assembly task" has the highest confidence, indicating that the overall reference video of the task belongs to the "assembly task" category; and multiple behavioral intent category identification results, such as {"grabbing": 0.8, "transportation": 0.1, "assembly": 0.1}, where "grabbing" has the highest confidence, indicating that the behavioral intent of the task execution object at the time point corresponding to this behavioral intent category identification result is "grabbing".
[0087] In this embodiment, the deep learning model identifies the task category of the task reference video based on the first fusion feature, obtaining a task category identification result. This result includes the confidence scores corresponding to multiple candidate task categories. The model also identifies the behavioral intent category of the task execution object, obtaining behavioral intent identification results for multiple time points. Each time point's behavioral intent identification result includes the confidence scores corresponding to multiple candidate behavioral intent categories at that time point. The deep learning model outputs both the task category identification result and the behavioral intent identification result.
[0088] Among them, deep learning models are obtained after training.
[0089] Optionally, during the training of the deep learning model, behavioral intent category labels and task category labels can be used as training labels to perform multi-task learning on the deep learning model, so as to improve the accuracy of the deep learning model in recognizing behavioral intent and task category.
[0090] S305, determine the task category of the task reference video based on the confidence level of the candidate task category.
[0091] In this embodiment, based on the confidence level of the candidate task category in the task category identification results, the candidate task category with the highest confidence level or a confidence level greater than the first threshold is selected, and the candidate task category is determined as the task category of the task reference video.
[0092] S306, Based on the confidence level of the candidate behavioral intent categories corresponding to multiple time points, determine the behavioral intent category of the task execution object at multiple time points.
[0093] In this embodiment, for each of the multiple time points, among the candidate behavioral intent categories corresponding to the time point, the candidate behavioral intent category with the highest confidence or with a confidence greater than the second threshold is selected, and the candidate behavioral intent category is determined as the behavioral intent category of the task execution object at that time point.
[0094] S307, Based on the recognition results, determine the valid video segments of the task reference video. The valid video segments are used for verification and / or reference of the operations performed by the embodied intelligent agent during the task execution process.
[0095] In this embodiment, the task stages of the task reference video are analyzed based on the task category and the behavioral intent category of the task execution object at multiple time points to determine the video segments in the task reference video that are located in the task stages, that is, to determine the valid video segments of the task reference video.
[0096] In one possible implementation, such as Figure 3 As shown, S307 includes:
[0097] S3071, based on the task category of the task reference video, obtain the task actions corresponding to multiple task stages, including the task start stage, the task progress stage, and the task end stage.
[0098] In this implementation, different task categories may correspond to different task actions in the same task stage. Based on the task category of the task reference video, the task actions of that task category in multiple task stages can be found in the task information database, including the task actions of that task category at the beginning of the task, the task actions of that task category during the task, and the task actions of that task category at the end of the task.
[0099] S3072, Match the task actions corresponding to multiple task stages with the behavioral intent categories of the task execution object at multiple time points to obtain the time information corresponding to each of the multiple task stages.
[0100] In this implementation, the task actions corresponding to multiple task stages are matched with the behavioral intent categories of the task execution object at multiple time points to determine the task stages corresponding to the behavioral intent categories of the task execution object at multiple time points. Based on the task stages corresponding to the behavioral intent categories of the task execution object at multiple time points, the time information corresponding to each of the multiple task stages is obtained. The time information corresponding to each task stage includes the start time and / or end time of the task stage.
[0101] As an example, if the behavioral intent category of the task execution object at time point A is successfully matched with the task action at the beginning of the task, then the time information of the beginning of the task includes time point A. In this way, multiple time points can be obtained, and the start time and / or end time of the beginning of the task can be determined according to the time order of these multiple time points.
[0102] S3073, Based on the time information corresponding to each of the multiple task stages and the time information corresponding to the video frames in the task reference video, determine the valid video segments in the task reference video.
[0103] In this implementation, the time information corresponding to multiple task stages can be matched with the time information corresponding to video frames in the task reference video to obtain the task stages corresponding to the video frames in the task reference video, that is, to obtain the range of video frames corresponding to multiple task stages. Based on the range of video frames corresponding to multiple task stages, the effective video segments in the task reference video can be obtained.
[0104] In this implementation, by analyzing the task category of the task reference video and the behavioral intent category of the task execution object at multiple time points, the task stage corresponding to the task category of the task reference video and the task actions in the task stage are analyzed, thereby determining the valid video segments of the task reference video, which effectively improves the accuracy of determining the valid video segments of the task reference video.
[0105] In this embodiment, multimodal features are extracted from the task reference video of the embodied agent. A deep learning model with a multi-head attention mechanism is used to fuse the extracted multimodal features, and task recognition and behavioral intent recognition are performed based on the fused features, thus improving the accuracy of task recognition and behavioral intent recognition. Based on the recognition results, valid video segments of the task reference video are determined, improving the accuracy of valid video segments. Therefore, while improving the efficiency of embodied intelligence data processing, the accuracy of embodied intelligence data processing is also improved, providing more accurate reference data for the embodied agent to perform tasks, and increasing the success rate and accuracy of task execution.
[0106] In some embodiments, the embodied intelligence data processing method further includes: marking valid video segments in the task reference video. This marks the valid video segments in the task reference video. Marking of valid video segments in the task reference video can be achieved by adding markers at the start and end times of the valid video segments.
[0107] In some embodiments, the embodied intelligence data processing method further includes: marking multiple task stages of a valid video segment, and / or marking the task category of the valid video segment, and / or saving the task actions corresponding to the multiple task stages in the valid video segment.
[0108] In this embodiment, the time information corresponding to multiple task stages of the task reference video can be determined according to the foregoing embodiments. Based on the time information corresponding to multiple task stages and the time information of the effective video segments, multiple task stages can be marked in the effective video segments.
[0109] As an example, in the task reference video, when a worker reaches out to grab an object, the deep learning model identifies the worker's behavioral intention as "grabbing," which is the task action at the beginning of the task and can be marked as task initiation. In the task reference video, when a worker moves an object to a target location, the deep learning model identifies the worker's behavioral intention as "carrying," which is the task action during the task execution phase and can be marked as task in progress. In the task reference video, when a worker accurately places an object, the deep learning model identifies the worker's behavioral intention as "placing," which is the task action at the end of the task and can be marked as task completion.
[0110] In this embodiment, the task category of the task reference video can be determined according to the foregoing embodiments. The task category of the effective video segment is the same as the task category of the task reference video. The task category of the effective video segment can be marked by adding task tags to the effective video segment.
[0111] In this embodiment, based on the foregoing embodiments, the behavioral intent categories of the task execution object at multiple time points and the time information corresponding to multiple task stages of the task reference video can be determined. By matching the behavioral intent categories of the task execution object at multiple time points with the time information corresponding to multiple task stages of the task reference video, the behavioral intent categories corresponding to multiple task stages can be determined, that is, the task actions corresponding to multiple task stages, and the task actions corresponding to multiple task stages can be saved.
[0112] Among them, the multiple task stages marked for valid video segments, the task categories marked for valid video segments, and / or the task actions corresponding to the multiple task stages can be used for verification and / or reference of the operations performed by the embodied intelligent agent during task execution, thereby improving the success rate and accuracy of the embodied intelligent agent in performing tasks.
[0113] In some embodiments, the embodied intelligence data processing method further includes: extracting body key points of the task execution object from the task reference video to obtain the posture data of the task execution object, the posture data including a sequence of body key points of the task execution object ordered by time; and aligning the posture data, agent operation data, and effective video segments in time to obtain offline data. The offline data is used for verification and / or reference of the operations performed by the embodied intelligence agent during task execution. Therefore, in addition to effective video segments, posture data of the task execution object and agent operation data are provided for verification and / or reference of the operations performed by the embodied intelligence agent during task execution. Combining these three types of data allows for more accurate behavioral analysis and task progress tracking of the embodied intelligence agent's task execution process, improving the success rate and accuracy of task execution by the embodied intelligence agent.
[0114] The intelligent agent's operational data includes sensor data, motion trajectory data, and / or control commands from the robots participating in the task collaboration within the task reference video. The sensor data is recorded by sensing devices on the robots participating in the task collaboration within the task reference video. For example, when performing a task, the robot records its own motion state and sensor data (such as position, velocity, joint angles, force sensor readings, etc.) in real time, and this motion data and sensor data are automatically saved by the robot's control system.
[0115] Specifically, a pose estimation algorithm can be used to extract body key points from video frames of the task reference video, obtaining body key point data of the task execution object at multiple time points. From this data, the pose information of the task execution object can be derived. When the task execution object is a relevant person, a human pose estimation algorithm can be used to extract human skeletal key points from video frames of the task reference video, obtaining human skeletal key point data at multiple time points. This skeletal key point data is then saved as a time series, allowing for temporal synchronization with the task reference video.
[0116] Optionally, body keypoints can be extracted from the video frames of the task reference video by combining the depth information corresponding to the video frames in the task reference video, thereby improving the accuracy of body keypoint extraction. The depth information corresponding to the video frames can be obtained by a depth camera or an RGB (red, green, blue)-D (depth) sensor (which combines RGB color information and depth information).
[0117] Optionally, the offline data also includes: multiple task stages corresponding to the valid video segments, the task categories of the valid video segments, and the task actions corresponding to each of the multiple task stages in the valid video segments. This information can be referred to the description in the foregoing embodiments and will not be repeated here. This improves the richness of the offline data, enabling more accurate behavioral analysis and task progress tracking of the embodied agent's task execution process, thereby improving the success rate and accuracy of the embodied agent's task execution.
[0118] In traditional automation systems, task execution relies on preset rules and programming instructions. For complex manufacturing scenarios, preset rules are difficult to cope with unstructured or frequently changing environments, making it difficult for robots to understand or adapt to dynamic changes in tasks when collaborating with humans, resulting in low efficiency and accuracy in task execution.
[0119] Compared to relying on preset rules and programming instructions to perform tasks, relying on single-modal data to identify and predict task execution behavior can improve the flexibility of robot task execution. However, this method cannot capture spatiotemporal information, resulting in insufficient identification and prediction accuracy.
[0120] To address the aforementioned issues, offline data generated in the aforementioned embodiments (valid video segments, task categories of valid video segments, task actions, posture data, and / or agent operation data corresponding to multiple task stages within the valid video segments) can be used to verify and / or predict the operations of the embodied agent during task execution. This enables the embodied agent to cope with complex and dynamically changing environments, improving the efficiency, success rate, and accuracy of task execution.
[0121] Figure 4 This is a flowchart illustrating the verification and / or reference process of the operations performed by the embodied intelligent agent during task execution in the embodied intelligent data processing method provided in this application embodiment. For example... Figure 4 As shown, taking offline data including valid video clips as an example, the verification and / or reference process of operations performed by the embodied intelligent agent during task execution may include:
[0122] S401, Obtain the current task video of the embodied intelligent agent.
[0123] In this embodiment, the current task video of the embodied intelligent agent can be obtained by acquiring real-time data collected by sensing devices (including camera devices).
[0124] S402, perform multimodal feature extraction on the task video to obtain the second visual semantic features and the second body posture features of the embodied agent in the task video.
[0125] In one possible implementation, S402 includes: inputting the task video into a pre-trained multimodal large model, and extracting multimodal features from the task video within the multimodal large model to obtain the visual semantic features and the body posture features of the embodied agent in the task video. Thus, the accuracy of multimodal feature extraction is improved through the pre-trained multimodal large model.
[0126] In one possible implementation, the task video is input into a pre-trained CNN, where the CNN extracts features from the video frames to obtain the visual semantic features of the task video. The goal of pre-training the CNN is to improve the accuracy of the CNN in extracting visual semantic features.
[0127] In another possible implementation, the video frames of the task video are divided into multiple image blocks, which are then input into a ViT network. The ViT network learns the relationships between these image blocks to extract visual semantic features. This allows for the extraction of global visual semantic features, improving the accuracy of visual semantic feature extraction.
[0128] The implementation principles and technical effects of S402 and the above implementation method can be referred to the description of multimodal feature extraction of the task reference video in the foregoing embodiments, and will not be repeated here.
[0129] S403, the second visual semantic features of the task video and the second body posture features of the embodied agent are fused to obtain the second fused feature.
[0130] In one possible implementation, S403 includes: inputting the second visual semantic features of the task video and the second body posture features of the embodied agent into a deep learning model employing a multi-head attention mechanism; and fusing the second visual semantic features and the second body posture features in the deep learning model to obtain a second fused feature.
[0131] The implementation principles and technical effects of S403 and the above implementation method can be referred to the description of multimodal feature extraction of the task reference video in the foregoing embodiments, and will not be repeated here.
[0132] S404, based on the second fusion feature, perform task recognition on the task video and behavioral intent recognition on the embodied intelligent agent to obtain the current task information and current behavioral intent information of the embodied intelligent agent.
[0133] In this embodiment, since the second fusion feature is obtained by fusing multimodal features, it can reflect both the visual semantics of the task video and the body posture of the embodied agent in the task video. Based on the second fusion feature, feature recognition can be performed to realize task recognition of the task video and behavioral intention recognition of the embodied agent, thereby obtaining the current task information and behavioral intention information of the embodied agent.
[0134] In one possible implementation, within the deep learning model, task identification is performed on the task video and behavioral intent identification is performed on the embodied agent based on the second fusion feature, resulting in the output data of the deep learning model. The output data of the deep learning model includes the confidence scores of candidate task categories and the confidence scores of candidate behavioral intent categories of the embodied agent at the current time. Based on the confidence scores of candidate task categories, the task category of the task video is determined; based on the confidence scores of candidate behavioral intent categories of the embodied agent at the current time, the behavioral intent category of the embodied agent at the current time is determined.
[0135] In this embodiment, based on the confidence level corresponding to the candidate task category, the candidate task category with the highest confidence level or a confidence level greater than a first threshold is selected, and this candidate task category is determined as the task category of the task video. Among the candidate behavioral intent categories corresponding to the current time point, the candidate behavioral intent category with the highest confidence level or a confidence level greater than a second threshold is selected, and this candidate behavioral intent category is determined as the behavioral intent category of the embodied agent at that time point. Thus, by utilizing multimodal features and deep learning models, the accuracy of task category and behavioral intent category recognition is improved.
[0136] S405, based on the valid video segments of the task reference video, the current task information of the embodied agent, and the current behavioral intention information of the embodied agent, verify the correctness of the current task action of the embodied agent, predict the movement path of the embodied agent, and / or predict the next task action of the embodied agent.
[0137] In this embodiment, the effective video segments reflect the task execution process through video. The correctness of the current task action of the embodied intelligent agent can be verified, the movement path of the embodied intelligent agent can be predicted, and / or the next task action of the embodied intelligent agent can be predicted by comparing or matching the effective video segments of the task reference video with the current task information and the current behavioral intention information of the embodied intelligent agent.
[0138] In one possible implementation, the embodied agent's current task information includes the task category and / or task stage currently being performed by the embodied agent, and the embodied agent's current behavioral intent information includes the category of the embodied agent's current behavioral intent. Valid video clips are marked with task category and / or task stage. Based on this, the correctness of the embodied agent's current task action is verified according to the valid video clips, the embodied agent's current task information, and the embodied agent's current behavioral intent information, including at least one of the following implementation methods:
[0139] (1) Match the task category marked by the valid video segment with the task category currently being executed by the embodied agent to determine whether the task currently being executed by the embodied agent is correct.
[0140] In this implementation, if the task category marked on the valid video segment is consistent with the task category currently being executed by the embodied agent, then the task currently being executed by the embodied agent is determined to be correct; otherwise, the task currently being executed by the embodied agent is determined to be incorrect, thus realizing the verification of the task currently being executed by the embodied agent.
[0141] (2) Based on the current time and the time information corresponding to the task stage marked in the valid video segment, match the task stage marked in the valid video segment with the task stage currently being executed by the embodied agent to determine whether the task stage currently being executed by the embodied agent is correct.
[0142] In this implementation, based on the current time and the time information corresponding to the task stage marked in the valid video segment, the task stage corresponding to the current time is determined from the task stages marked in the valid video segment. If the task stage corresponding to the current time is consistent with the task stage currently being executed by the embodied intelligent agent, then the task stage currently being executed by the embodied intelligent agent is determined to be correct; otherwise, the task stage currently being executed by the embodied intelligent agent is determined to be incorrect, thus realizing the verification of the task stage currently being executed by the embodied intelligent agent.
[0143] (3) Based on the time information corresponding to the task stage marked in the effective video segment, determine the task stage corresponding to the current time, match the task action of the task stage corresponding to the current time with the current behavioral intention category of the embodied agent, and determine whether the current behavioral intention category of the embodied agent is correct.
[0144] In this implementation, if the task action of the task stage corresponding to the current time is consistent with the current behavioral intention category of the embodied agent, then the current behavioral intention category of the embodied agent is determined to be correct; otherwise, the current behavioral intention category of the embodied agent is determined to be incorrect, thus realizing the verification of the current behavioral intention category of the embodied agent.
[0145] Optionally, in the above implementation, if the task being performed by the embodied agent is incorrect, the stage of the task being performed by the embodied agent is incorrect, or the category of the embodied agent's current behavioral intent is incorrect, then an error reminder message is output or an error is marked (for example, the category of the embodied agent's behavioral intent is marked as abnormal) to remind relevant personnel to check the status of the embodied agent in a timely manner.
[0146] For example, based on valid video clips, if it is determined that the embodied agent should perform the "grab" action at the current time, and the behavior of the embodied agent is detected in real time as "adjustment", an abnormal alarm will be triggered.
[0147] For example, based on valid video clips, if it is determined that the current time should be in the task completion stage, and if it is detected in real time that the embodied intelligent agent is still adjusting the object, then the behavior of adjusting the object is marked as abnormal behavior.
[0148] In one possible implementation, the embodied agent's current task information includes the task category and / or task stage currently being performed, and the embodied agent's current behavioral intent information includes the category of the current behavioral intent. Valid video clips are marked with task categories and / or task stages. Based on the valid video clips, the embodied agent's current task information, and the embodied agent's current behavioral intent information, the embodied agent's movement path and / or its next task action is predicted. This includes: matching the embodied agent's current task category with the task categories marked in the valid video clips; matching the embodied agent's current behavioral intent category with the task actions of the task stages marked in the valid video clips to determine the current task stage; and predicting the embodied agent's movement path and / or its next task action based on the current task stage and the task actions of the task stages marked in the valid video clips. This provides guidance for the embodied agent's movement path and / or next task action, improving the accuracy of the embodied agent's movement path and / or next task action.
[0149] In this implementation, before predicting the movement path, a map of the embodied agent's working area is created based on environmental information collected by sensors (such as LiDAR and cameras). This map marks the locations of task-related objects, the embodied agent's location, and obstacles. Based on this map, path planning and deep learning algorithms are used to plan the optimal path for the embodied agent. Then, this optimal path is adjusted according to the above process: the task category currently being executed by the embodied agent is matched with the task category marked in the effective video clip; the current behavioral intent category of the embodied agent is matched with the task actions of the task stage marked in the effective video clip to determine the current task stage. Based on the current task stage and the task actions of the task stage marked in the effective video clip, the next N-step movement position of the embodied agent is obtained, where N is greater than or equal to 1. The optimal path of the embodied agent is then adjusted based on these next N-step movement positions. This optimizes the embodied agent's movement path and improves the efficiency and accuracy of its task execution.
[0150] In this implementation, the task category currently being executed by the embodied agent is matched with the task category marked on the valid video clip, and the current behavioral intent category of the embodied agent is matched with the task actions of the task stage marked on the valid video clip to determine the current task stage being executed by the embodied agent. Based on the current task stage being executed by the embodied agent, the corresponding operation instructions for that task stage are determined from the task database, such as "adjust the position of the robot's end effector" and "open the gripper" for the "grasping stage". The corresponding operation instructions for that task stage are generated to control the embodied agent to execute the next task action according to the operation instructions. Thus, the accuracy of the embodied agent in executing task actions is effectively improved.
[0151] In one possible implementation, based on any of the above embodiments, the step further includes:
[0152] S206, Based on the aforementioned valid video clips, verify the correctness of the embodied intelligent agent's task execution and obtain the verification result. The verification result is either correct or incorrect.
[0153] S207, if the verification result of the embodied agent is incorrect, identify the user's hand movements in the video based on valid video segments.
[0154] S208, if the user's hand movement meets the first preset condition, obtain the video time corresponding to the occurrence of the user's hand movement.
[0155] S209, if the video time corresponds to a preset time period, obtain the precaution information associated with the task execution object in the task.
[0156] S210, based on the precautions information, determine the new target task of the embodied intelligent agent and execute it.
[0157] Specifically, the user can be the instructor. For example, in a valid video clip, the user issues the instruction: "Fetch water." The instruction in the valid video clip is to fetch drinking water, but the embodied agent might observe that the user is in a bathroom and fetches tap water instead of drinking water, making the action incorrect. In this case, the user's hand gesture is identified. The first preset condition could be that the user has just taken medication, meaning they are about to take it and therefore need drinking water. In this case, the corresponding video time when the user's hand gesture, such as taking medication, occurs is obtained. This time is natural time (i.e., 24 hours in a day), not the video playback duration. The task execution object could be, for example, the medication. The preset time period corresponds to the user's hand gesture. For example, if the hand gesture is taking medication, the preset time period could be the user's medication taking period within the last 3 days. If the corresponding video time falls within this preset time period, then combining the aforementioned user hand gesture—that is, determining the user's action intent based on both time and user hand gesture—makes the task determination more accurate. In this example, the user's intention is determined to take medication, so the next objective is to obtain drinking water.
[0158] The aforementioned precautions may include determining the appropriate water temperature for taking the medication based on the corresponding instructions for use or the doctor's written instructions in the electronic medical record. The final water collection task should then be determined based on this water temperature.
[0159] In some optional embodiments, step S210 above includes:
[0160] Based on the user's hand gestures, identify the user's intention.
[0161] Obtain user preference information associated with action intentions.
[0162] Based on the information regarding precautions and user preferences, determine the new target tasks for the embodied intelligent agent.
[0163] For example, if a user's habit (i.e., user preference information) is to rinse their mouth with water after taking medication, then the embodied intelligent agent would collect water and divide it into portions, one portion for taking the medication and the other portion for rinsing the mouth after taking the medication. The corresponding new objective task would then be, for example, to collect water and then divide it into portions.
[0164] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.
[0165] Figure 5 This is a schematic diagram of the structure of the embodied intelligent data processing device provided in the embodiments of this application, such as... Figure 5 As shown, the embodied intelligent data processing device 500 of this application embodiment includes: a video data acquisition unit 501, a multimodal feature extraction unit 502, a multimodal feature fusion unit 503, a task behavior recognition unit 504, and an effective data determination unit 505. Wherein:
[0166] The video data acquisition unit 501 is used to acquire a task reference video of the embodied intelligent agent. The task reference video contains behavioral information of the task execution object and state information of task-related objects. The multimodal feature extraction unit 502 is used to extract multimodal features from the task reference video to obtain visual semantic features and body posture features of the task execution object. The multimodal feature fusion unit 503 is used to fuse the visual semantic features and body posture features to obtain a first fused feature. The task behavior recognition unit 504 is used to perform task recognition on the task reference video and behavioral intent recognition on the task execution object based on the first fused feature to obtain a recognition result. The valid data determination unit 505 is used to determine the valid video segments of the task reference video based on the recognition result. The valid video segments are used for verification and / or reference of the operations performed by the embodied intelligent agent during task execution.
[0167] In some embodiments, the multimodal feature fusion unit 503 is specifically used to: input visual semantic features and body posture features into a deep learning model employing a multi-head attention mechanism, and fuse the visual semantic features and body posture features in the deep learning model to obtain a first fused feature.
[0168] In some embodiments, the identification results include the task category of the task reference video and the behavioral intent category of the task execution object at multiple time points. The task behavior identification unit 504 is specifically used to: in a deep learning model, perform task identification on the task reference video and behavioral intent identification on the task execution object based on the first fusion feature to obtain the output data of the deep learning model, wherein the output data includes the confidence scores of candidate task categories and the confidence scores of candidate behavioral intent categories corresponding to multiple time points respectively; determine the task category of the task reference video based on the confidence scores of candidate task categories; and determine the behavioral intent category of the task execution object at multiple time points based on the confidence scores of candidate behavioral intent categories corresponding to multiple time points respectively.
[0169] In some embodiments, the identification result includes the task category of the task reference video and the behavioral intent category of the task execution object at multiple time points. Valid video segments include video segments from the start to the end of the task. The valid data determination unit 505 is specifically used for: obtaining task actions corresponding to multiple task stages according to the task category, the multiple task stages including the task start stage, task progress stage, and task end stage; matching the task actions corresponding to the multiple task stages with the behavioral intent categories of the task execution object at multiple time points to obtain time information corresponding to the multiple task stages; determining valid video segments in the task reference video based on the time information corresponding to the multiple task stages and the time information corresponding to video frames in the task reference video; and further, the embodied intelligent data processing device includes: a task stage marking unit (not shown in the figure) for marking multiple task stages on the valid video segments; and / or a task category marking unit (not shown in the figure) for marking the task category of the valid video segments; and / or a behavioral intent storage unit (not shown in the figure) for storing the task actions corresponding to the multiple task stages in the valid video segments.
[0170] In some embodiments, the embodied intelligent data processing device further includes: a key point extraction unit (not shown in the figure), used to extract the body key points of the task execution object in the task reference video to obtain the posture data of the task execution object, the posture data including the sequence of body key points of the task execution object sorted by time; and a time alignment unit (not shown in the figure), used to time-align the posture data, agent operation data and effective video segments to obtain offline data, the offline data being used for verification and / or reference of the operations of the embodied intelligent agent in the process of performing the task, the agent operation data including the sensor data, motion trajectory data and / or control command data of the robot participating in the task collaboration in the task reference video.
[0171] In some embodiments, the verification and / or reference process for the actions performed by the embodied agent during task execution includes: acquiring the current task video of the embodied agent; extracting multimodal features from the task video to obtain visual semantic features of the task video and body posture features of the embodied agent; fusing the visual semantic features and body posture features to obtain a second fused feature; performing task recognition on the task video and behavioral intent recognition on the embodied agent based on the second fused feature to obtain current task information and current behavioral intent information of the embodied agent; and verifying the correctness of the current task action of the embodied agent, predicting the movement path of the embodied agent, and / or predicting the next task action of the embodied agent based on valid video segments, the current task information of the embodied agent, and the current behavioral intent information of the embodied agent.
[0172] In some embodiments, the current task information of the embodied agent includes the task category and / or task stage currently being performed by the embodied agent, and the current behavioral intent information of the embodied agent includes the current behavioral intent category of the embodied agent. A valid video segment is marked with a task category and / or task stage. The correctness of the current task action of the embodied agent is verified based on the valid video segment, the current task information, and the current behavioral intent information, including at least one of the following: matching the task category marked in the valid video segment with the task category currently being performed by the embodied agent to determine whether the task currently being performed by the embodied agent is correct; matching the task stage marked in the valid video segment with the task stage currently being performed by the embodied agent according to the current time and the time information corresponding to the task stage marked in the valid video segment to determine whether the task stage currently being performed by the embodied agent is correct; determining the task stage corresponding to the current time according to the current time and the time information corresponding to the task stage marked in the valid video segment, and matching the task action of the task stage corresponding to the current time with the current behavioral intent category of the embodied agent to determine whether the current behavioral intent category of the embodied agent is correct.
[0173] In some embodiments, the current task information of the embodied agent includes the task category and / or task stage currently being performed by the embodied agent, and the current behavioral intention information of the embodied agent includes the current behavioral intention category of the embodied agent. The effective video segment is marked with the task category and / or task stage. Based on the effective video segment, the current task information of the embodied agent, and the current behavioral intention information of the embodied agent, the movement path of the embodied agent and / or the next task action of the embodied agent are predicted, including: matching the task category currently being performed by the embodied agent with the task category marked in the effective video segment; matching the current behavioral intention category of the embodied agent with the task action of the task stage marked in the effective video segment to determine the current task stage being performed by the embodied agent; and predicting the movement path of the embodied agent and / or the next task action of the embodied agent based on the current task stage being performed by the embodied agent and the task action of the task stage marked in the effective video segment.
[0174] The apparatus of this embodiment can be used to execute the technical solutions of any of the method embodiments shown above. Its implementation principle and technical effect are similar, and will not be repeated here.
[0175] Figure 6 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of this application. Exemplarily, the electronic device may be provided as a server or a computer. (Refer to...) Figure 6The electronic device 600 includes a processing component 601, which further includes one or more processors, and memory resources represented by memory 602 for storing instructions, such as application programs, that can be executed by the processing component 601. The application programs stored in memory 602 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 601 is configured to execute instructions to perform any of the method embodiments described above.
[0176] Electronic device 600 may also include a power supply component 603 configured to perform power management of electronic device 600, a wired or wireless network interface 604 configured to connect electronic device 600 to a network, and an input / output (I / O) interface 605. Electronic device 600 may operate on an operating system stored in memory 602, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or similar.
[0177] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described embodied intelligent data processing method.
[0178] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described embodied intelligent data processing method.
[0179] The aforementioned computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0180] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an application-specific integrated circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in an embodied intelligent data processing device.
[0181] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0182] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A method for embodied intelligence data processing, characterized in that, include: Obtain a task reference video of the embodied intelligent agent, wherein the task reference video contains behavioral information of the task execution object and state information of task-related objects; Multimodal feature extraction is performed on the task reference video to obtain the visual semantic features of the task reference video and the body posture features of the task execution object; The visual semantic features and the body posture features are fused to obtain a first fused feature; Based on the first fusion feature, task recognition is performed on the task reference video and behavioral intent recognition is performed on the task execution object to obtain the recognition result; Based on the recognition results, valid video segments of the task reference video are determined, and the valid video segments are used for verification and / or reference of the operations performed by the embodied intelligent agent during the task execution process; The method further includes: Based on the valid video clips, the correctness of the embodied intelligent agent's task execution is verified, and a verification result is obtained; the verification result is either correct or incorrect. If the verification result of the embodied intelligent agent is incorrect, the user's hand movements in the video are identified based on the valid video segments; If the user's hand gesture meets the first preset condition, obtain the video time corresponding to the occurrence of the user's hand gesture. If the video time corresponds to a preset time period, obtain the precaution information associated with the task execution object in the task; Based on the aforementioned precautions, the embodied intelligent agent determines and executes a new target task.
2. The embodied intelligence data processing method according to claim 1, characterized in that, The recognition result includes the task category of the task reference video and the behavioral intent category of the task execution object at multiple time points. The step of performing task recognition on the task reference video and behavioral intent recognition on the task execution object based on the first fusion feature to obtain the recognition result includes: In a deep learning model employing a multi-head attention mechanism, based on the first fusion feature, task identification is performed on the task reference video and behavioral intent identification is performed on the task execution object to obtain the output data of the deep learning model. The output data includes the confidence scores of candidate task categories and the confidence scores of candidate behavioral intent categories corresponding to the multiple time points. Based on the confidence level of the candidate task categories, determine the task category of the task reference video; Based on the confidence level of the candidate behavioral intent categories corresponding to the multiple time points, the behavioral intent categories of the task execution object at the multiple time points are determined.
3. The embodied intelligence data processing method according to claim 1, characterized in that, The identification result includes the task category of the task reference video and the behavioral intent category of the task execution object at multiple time points. The valid video segments include video segments from the start to the end of the task. Determining the valid video segments of the task reference video based on the identification result includes: Based on the task category, obtain the task actions corresponding to multiple task stages, including the task start stage, the task progress stage, and the task end stage. The task actions corresponding to the multiple task stages are matched with the behavioral intention categories of the task execution object at multiple time points to obtain the time information corresponding to the multiple task stages. Based on the time information corresponding to the multiple task stages and the time information corresponding to the video frames in the task reference video, the effective video segments are determined in the task reference video. The step of determining the valid video segment in the task reference video based on the time information corresponding to the multiple task stages and the time information corresponding to the video frames in the task reference video further includes: The valid video segments are marked according to the multiple task stages; And / or, the task category of the valid video segment is marked; And / or, save the task actions corresponding to multiple task stages in the effective video segment.
4. The embodied intelligence data processing method according to any one of claims 1 to 3, characterized in that, After obtaining the task reference video corresponding to the embodied intelligent agent, the process also includes: In the task reference video, the body key points of the task execution object are extracted to obtain the posture data of the task execution object. The posture data includes the sequence of body key points of the task execution object sorted by time. The posture data, agent operation data, and effective video segments are time-aligned to obtain offline data. The offline data is used for verification and / or reference of the operations performed by the embodied agent during task execution. The agent operation data includes sensor data, motion trajectory data, and / or control command data of the robot participating in task collaboration in the task reference video.
5. The embodied intelligence data processing method according to any one of claims 1 to 3, characterized in that, The verification and / or reference process for the operations performed by the embodied intelligent agent during task execution includes: Obtain the current task video of the embodied intelligent agent; Multimodal feature extraction is performed on the task video to obtain the second visual semantic features of the task video and the second body posture features of the embodied agent; The second visual semantic feature and the second body pose feature are fused to obtain the second fused feature; Based on the second fusion feature, task recognition is performed on the task video and behavioral intent recognition is performed on the embodied intelligent agent to obtain the current task information and the current behavioral intent information of the embodied intelligent agent; Based on the valid video clips, the current task information of the embodied intelligent agent, and the current behavioral intention information of the embodied intelligent agent, the correctness of the current task action of the embodied intelligent agent is verified, the movement path of the embodied intelligent agent is predicted, and / or the next task action of the embodied intelligent agent is predicted.
6. The embodied intelligence data processing method according to claim 5, characterized in that, The current task information of the embodied intelligent agent includes the task category and / or task stage currently being performed by the embodied intelligent agent; the current behavioral intention information of the embodied intelligent agent includes the current behavioral intention category of the embodied intelligent agent; and the effective video segment is marked with the task category and / or task stage. Based on the valid video clip, the current task information of the embodied intelligent agent, and the current behavioral intention information of the embodied intelligent agent, the correctness of the current task action of the embodied intelligent agent is verified, including at least one of the following: The task category tagged in the valid video segment is matched with the task category currently being executed by the embodied intelligent agent to determine whether the task currently being executed by the embodied intelligent agent is correct; Based on the current time and the time information corresponding to the task stage marked in the effective video segment, the task stage marked in the effective video segment is matched with the task stage currently being executed by the embodied intelligent agent to determine whether the task stage currently being executed by the embodied intelligent agent is correct. Based on the current time and the time information corresponding to the task stage marked in the effective video segment, determine the task stage corresponding to the current time, match the task action of the task stage corresponding to the current time with the current behavioral intention category of the embodied intelligent agent, and determine whether the current behavioral intention category of the embodied intelligent agent is correct.
7. The embodied intelligence data processing method according to claim 5, characterized in that, The current task information of the embodied intelligent agent includes the task category and / or task stage currently being performed by the embodied intelligent agent; the current behavioral intention information of the embodied intelligent agent includes the current behavioral intention category of the embodied intelligent agent; and the effective video segment is marked with the task category and / or task stage. Based on the valid video clips, the current task information of the embodied agent, and the current behavioral intent information of the embodied agent, the movement path of the embodied agent is predicted and / or the next task action of the embodied agent is predicted, including: The task category currently being executed by the embodied intelligent agent is matched with the task category marked by the effective video segment, and the current behavioral intention category of the embodied intelligent agent is matched with the task action of the task stage marked by the effective video segment to determine the task stage currently being executed by the embodied intelligent agent. Based on the task stage currently being executed by the embodied agent and the task actions of the marked task stages in the effective video segment, the movement path of the embodied agent is predicted and / or the next task action of the embodied agent is predicted.
8. The embodied intelligence data processing method according to claim 1, characterized in that, The step of determining and executing a new target task for the embodied intelligent agent based on the aforementioned precautions information includes: Based on the user's hand movements, identify the user's intention. Obtain user preference information associated with the stated action intent; Based on the aforementioned precautions and user preference information, a new target task is determined for the embodied intelligent agent.
9. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the embodied intelligent data processing method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Man-machine cooperation method and system based on multi-modal behavior online prediction
CN113524175A