Intention understanding method based on multimodal large model in human-computer collaborative environment

By using multimodal large models in a human-computer collaborative environment, combining visual and auditory information to perceive the environment and task execution process in real time, the problem that agents in the prior art are difficult to adapt to in complex dynamic scenarios, and more efficient and accurate task execution and autonomy are achieved.

CN119785276BActive Publication Date: 2025-05-09ZHONGKE NANJING SOFTWARE TECH RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510289601.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-05-09
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

In human-computer collaborative tasks in complex dynamic scenarios, existing large language models are difficult to learn and evolve behavioral strategies from incomplete information game scenarios, and cannot effectively adapt in a multi-task environment.

Method used

Using the intention understanding method based on the multimodal big model, a multimodal big model integrating behavior recognition, action prediction and environmental understanding is built by combining visual and auditory information to perceive the environment and task execution process in real time, so as to realize the agent's autonomous inference of the next task and respond.

Benefits of technology

It improves the task recognition accuracy and adaptability of the agent in complex environments, realizes the agent's autonomous execution of actions and optimizes task execution efficiency, and enhances the robot's autonomy and system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785276B_ABST
    Figure CN119785276B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and to a method for understanding intentions in a human-machine collaborative environment based on a multimodal large model. It includes the following specific steps: video analysis and task planning; preprocessing the video using key frame extraction and image segmentation methods; constructing a directed acyclic graph of tasks, memorizing feasible task paths; real-time intention judgment; processing multimodal data, splicing data interception pictures of different modes together in a fixed manner; using a task directed acyclic graph to screen the subtasks that the large model needs to face when making a judgment, and combing some more likely subtask sequences for the large model; mechanical arm command generation and feedback; issuing corresponding instructions according to the task directed acyclic graph, executing corresponding steps, and generating feedback data. The present invention successfully realizes accurate recognition and task inference of human behavior in a complex environment by combining multimodal information such as vision and hearing, and perceiving the environment and task execution process in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method for understanding intentions in a human-computer collaborative environment based on a multimodal large model. Background Art

[0002] Intent understanding is a critical and challenging task in the field of artificial intelligence. Its goal is to enable machines to accurately infer user intent in order to provide solutions for complex tasks. Intent understanding aims to infer the user's true intent or goal from input natural language or multimodal data. It is widely used in many fields such as dialogue systems, recommendation systems, and game AI. In this task, the model needs to deal with the ambiguity and contextual dependence of language, especially when faced with complex tasks or dynamic scenarios, the difficulty of intent understanding increases significantly.

[0003] Designing an intelligent agent with a strong ability to understand human intentions to handle human-machine collaboration problems in complex dynamic scenarios has always been a vision of the academic community. This requires the agent to have learning and generalization capabilities in a variety of tasks. The emergence of large language models reveals this vision, especially that they can be quickly generalized in a range of tasks with just a few demonstrations. Thanks to this, many systems based on large models show significantly enhanced performance, such as question answering, code generation, and real-world applications.

[0004] At present, the methods to achieve intent understanding mainly include intent understanding based on deep learning and intent understanding based on large models.

[0005] Recurrent neural networks and long short-term memory networks significantly improve the ability to capture language context by efficiently processing time series data. This method is particularly suitable for tasks that are sensitive to semantic order. The Transformer architecture introduces a self-attention mechanism, which greatly improves the efficiency of modeling long sequence dependencies and becomes an important cornerstone of intent understanding technology.

[0006] Pre-training-fine-tuning paradigm: Through unsupervised pre-training on massive data, the model learns a wide range of language representations; then fine-tunes with a small amount of task-related data to achieve adaptation to specific tasks. Cue word optimization: Through careful design and improvement of cue words, large language models can perform well on multiple tasks without fine-tuning, providing a lightweight solution for task generalization.

[0007] The intent understanding technology of large language models in the prior art has the following disadvantages:

[0008] (1) Most agents are designed for specific tasks through complex prompts, including detailed task descriptions and behavioral specifications. However, many incomplete information games, such as poker, require painstaking efforts to design strategic behaviors because the opponent’s hand and the opponent’s game strategy cannot be directly obtained.

[0009] (2) In new scenarios, most agents do not consider interactions with incomplete information game scenarios, and more importantly, cannot learn from past experience and evolve their behavioral strategies during the interaction process. In contrast, humans usually learn and adjust their behavior through interaction.

[0010] Based on this, a method for intention understanding in a human-computer collaborative environment based on a multimodal large model is proposed. Summary of the invention

[0011] The purpose of the present invention is to address the problems existing in the background technology and propose a method for understanding intentions in a human-computer collaborative environment based on a multimodal large model. The present invention combines multimodal information such as vision and hearing to perceive the environment and task execution process in real time, and designs an intelligent task planning method based on a large model, which successfully realizes the accurate recognition and task inference of human behaviors in complex environments. By constructing a multimodal large model that integrates behavior recognition, action prediction, and environmental understanding, the ability of existing intelligent agents to effectively adapt in dynamic environments is improved, and the intelligent agent can autonomously infer the next task and respond.

[0012] The first aspect of the present invention provides a method for understanding intention in a human-computer collaborative environment based on a multimodal large model, comprising the following specific steps:

[0013] S1, video analysis and task planning;

[0014] S11, preprocessing the video using key frame extraction and image segmentation methods;

[0015] S12. Build a directed acyclic graph of tasks and memorize feasible task paths;

[0016] S2, real-time intention judgment;

[0017] S21, processing the multimodal data, and splicing the data capture images of different modalities together in a fixed manner;

[0018] S22. Use the task directed acyclic graph to screen the subtasks that the large model needs to face when making a judgment, and sort out some more likely subtask sequences for the large model;

[0019] S3, generation and feedback of robot arm motion instructions;

[0020] S31. Issue corresponding instructions and execute corresponding steps according to the task directed acyclic graph to generate feedback data.

[0021] The extraction of key frames in step S11 includes the following steps:

[0022] The target action in the video is detected, the moving speed of the target is calculated, and the moving speed curve of the target is drawn; frames are extracted at the trough and peak sections of the speed curve.

[0023] The image segmentation in step S11 includes the following steps:

[0024] First, the Grounding DINO method is used to realize the transformation from text description to object recognition. The position of the object in the picture is identified according to the specific text description, and the position coordinates of the object are returned. Then, the SAM2 model is used to segment the image according to the position coordinates of the object.

[0025] In step S12, the nodes in the task directed acyclic graph are used to represent the numerous actions in the task, and the arrows are used to represent the temporal logical sequence between the actions.

[0026] Constructing a task directed acyclic graph includes the following specific steps:

[0027] By inputting the guidance video into the multimodal agent, the task nodes and the time sequence of the nodes about the task are constructed according to the video information;

[0028] Find the dependency set of all task nodes, and the prerequisites required by the task no matter which path is taken;

[0029] Find the dependent set of all task nodes, and the task nodes that will be executed no matter which path is taken;

[0030] Remove irrelevant nodes, that is, nodes that are in the dependency set of the same task node and in its dependent set;

[0031] Mark potential dependencies and potential dependencies, that is, the preconditions and postconditions that exist in a certain path;

[0032] Mark mutually exclusive nodes;

[0033] All dependent, potentially dependent and mutually exclusive relationships are listed as negation terms in the logical expression, and dependencies are listed as positive terms.

[0034] The steps to merge task paths into a directed acyclic graph are as follows:

[0035] definition Different mission paths , the path includes Different task nodes ;

[0036] According to a task path , get the node The corresponding dependency set and the dependent set , as shown below:

[0037]

[0038] in, Represents a node N Split task paths for the center R ;

[0039] For a node , find the intersection of the dependency sets of all task paths The intersection of all dependent sets , get the task dependencies based on all task paths, as shown below:

[0040]

[0041] Since there may be multiple dependency paths between two nodes, it is necessary to clear redundant dependencies to obtain the nodes. The minimum set of dependencies and the minimum dependent set , as shown below:

[0042]

[0043] in, , What is contained in is the edge of the task directed acyclic graph, and a task directed acyclic graph based on all current task paths is obtained.

[0044] In step S21, image data is generated every 5 seconds; a picture is captured every 1 second;

[0045] For vertical screen videos, use time-lapse sequence to stitch pictures from left to right;

[0046] For horizontal screen videos, first splice the first 4 seconds of video into a field pattern clockwise from the upper left, and then splice the last second to the right side of the field pattern.

[0047] Step S22 includes the following steps:

[0048] (1) Mark the task nodes with in-degree 0 in the task directed acyclic graph as executable nodes;

[0049] (2) When an executable node is executed, remove the node from the task directed acyclic graph, and then reselect a node with an in-degree of 0 from the new task directed acyclic graph;

[0050] (3) When all nodes have been executed, the task is completed; otherwise, return to step (1).

[0051] A second aspect of the present invention provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned method for understanding intentions based on a multimodal large model in a human-computer collaborative environment.

[0052] A third aspect of the present invention provides an electronic device, comprising:

[0053] one or more processors;

[0054] A storage device for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the above-mentioned intention understanding method based on the multimodal large model in a human-computer collaborative environment.

[0055] Compared with the prior art, the present invention has the following beneficial technical effects:

[0056] 1. The present invention combines visual, auditory and other multimodal information, uses deep learning models to analyze character behavior and infer task progress in real time, thereby accurately monitoring and predicting task status in complex environments. It is beneficial to improve task recognition accuracy: it can accurately identify the current task status based on the details of the environment and character behavior, avoiding the recognition errors of traditional methods. Enhance task adaptability: it can dynamically adjust the inference results, adapt to complex and changing environments, and provide more intelligent task execution support.

[0057] 2. The present invention designs a multimodal large model architecture, which, combined with data analysis, can infer the next step of the character's behavior and guide the intelligent agent to complete the corresponding task. Improve robot autonomy: The robot can autonomously perform actions according to the task flow inferred by the model, reducing human intervention. Optimize task execution efficiency: By adjusting the task plan in real time, ensure that the robot arm is more efficient and accurate in performing tasks.

[0058] 3. The present invention uses a directed acyclic graph to represent the logical relationship between tasks, making the task planning process clearer and the task dependencies effectively managed; task management is clearer: the directed acyclic graph structure can clearly describe the order and dependencies between tasks, facilitating task planning and scheduling. Reduce errors and conflicts: by clarifying the task relationship, conflicts and errors in task execution are avoided, improving system stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 It is a key frame extraction effect and hand speed curve diagram in an embodiment of the present invention;

[0060] Figure 2 A schematic diagram of a simple directed non-changing graph in an embodiment of the present invention;

[0061] Figure 3 is a logical expression that satisfies the conditions in the embodiment of the present invention;

[0062] Figure 4 A schematic diagram of a directed acyclic graph of a multimodal agent based on video learning to construct a task in an embodiment of the present invention;

[0063] Figure 5 Schematic diagram of vertical screen video splicing in an embodiment of the present invention;

[0064] Figure 6 Schematic diagram of horizontal screen video splicing in an embodiment of the present invention. DETAILED DESCRIPTION Example

[0065] The present invention proposes a method for understanding intentions in a human-machine collaborative environment based on a multimodal large model, which includes the following specific steps:

[0066] S1, video analysis and task planning;

[0067] S11, preprocessing the video using key frame extraction and image segmentation methods;

[0068] The extraction of key frames in step S11 includes the following steps:

[0069] The target action in the video is detected, the moving speed of the target is calculated, and the moving speed curve of the target is drawn; frames are extracted at the trough and peak sections of the speed curve.

[0070] The image segmentation in step S11 includes the following steps:

[0071] First, the Grounding DINO method is used to realize the transformation from text description to object recognition. The position of the object in the picture is identified according to the specific text description, and the position coordinates of the object are returned. Then, the SAM2 model is used to segment the image according to the position coordinates of the object.

[0072] In this embodiment, Figure 1For example, since the current multimodal large model reads videos by extracting frames from the video, and according to current relevant research, the longer the video read by the multimodal large model, the more frames it has, which may enhance the information density, but also bring redundancy and noise. Therefore, in order to reduce the redundancy and noise of information, it is necessary to extract key frames from the video, and the effect is shown in Figure 1. At the same time, in order to facilitate the multimodal large model to understand the changes in the action sequence, it is also necessary to supplement frames near the key frames. In order to accurately extract the key actions of the characters in the video, the hands of the characters in the video will be used for target detection, and the hand movement speed will be calculated and the hand movement speed curve will be drawn. It can be found that when the hand movement speed is at the trough of the curve, it is easy to show the interactive behavior of the characters, and when it is at the peak of the curve, it is easy to show the movement trend of the characters. Based on this discovery, it is only necessary to extract frames around the trough and peak.

[0073] Image segmentation: In order to improve the understanding ability and task effect of large multimodal models when they need to accurately analyze objects, actions or backgrounds in complex scenes, image segmentation preprocessing of videos is indispensable. Image segmentation can filter out irrelevant information in video frames, retain task-related areas, and significantly reduce data redundancy: regional focusing and feature compression of video information can be achieved; at the same time, image segmentation can provide more refined visual semantic information for multimodal models, such as category information for objects in video frames and global context of the background environment in the frame provided by the model. In order to obtain an excellent segmentation result, the Grounded-SAM2 method is used. This method first uses the Grounding DINO method to realize the transformation from text description to object recognition, identifies the position of objects in the picture according to the specific text description, and returns the position coordinates of the object; then uses the SAM2 model to segment the image according to the position coordinates of the object.

[0074] S12. Build a directed acyclic graph of tasks and memorize feasible task paths;

[0075] In step S12, the nodes in the task directed acyclic graph are used to represent the numerous actions in the task, and the arrows are used to represent the temporal logical order between the actions. In order to enhance the multimodal intelligent agent's understanding of human intentions in actual work, a method is designed in this embodiment to help the model identify and understand tasks by constructing a directed acyclic graph that describes the relationship between subtasks. The present invention memorizes feasible task paths through a task directed acyclic graph. Compared to directly using the entire task process to judge intent, while providing reasonable subtask options, a large number of redundant and inappropriate options for the current scenario are removed, thereby improving the judgment ability and efficiency of the multimodal intelligent agent.

[0076] Compared with the prior art, the present invention innovatively introduces a multi-task path intersection analysis mechanism when constructing a task directed acyclic graph. The multi-task path intersection analysis mechanism is a method of accurately describing the logical dependencies between tasks by analyzing the dependencies and dependencies of nodes in different task paths and finding their intersections. The key is to use logical expressions to describe the relationship between subtasks. The use of this method can overcome the problem that a single directed acyclic graph cannot be used to describe the relationship between subtasks, such as mutual exclusion. For example, in Figure 2, if a single directed acyclic graph is used to describe it, it will be found that nodes 2 and 3 are both necessary prerequisites for node 4; but using a directed acyclic graph with a logical expression can more flexibly describe the dependency relationship between tasks, for example, "the execution of node 4 requires the completion of either node 2 or node 3", and this relationship can be described by the logical expression in Figure 3. This method not only retains the intuitiveness of the directed acyclic graph in describing the task flow, but also enhances the ability to express complex task dependencies through logical expressions.

[0077] Based on this concept, this embodiment designs a learning method based on video input, which aims to enable the multimodal agent to understand the basic composition of the task. The process of the method is shown in Figure 2. In the figure, G represents a graph describing the task relationship summarized by the multimodal agent, and video represents a video used to guide the multimodal agent. In order to describe the directed acyclic graph of tasks constructed by the multimodal agent, the present invention defines a special node for describing subtasks - the task node, which contains the job description, tools used, processed items information of the characters in the video, and information of the previous and subsequent nodes.

[0078] The present invention allows a multimodal agent to construct a basic understanding of the task based on the video information by inputting a guidance video demonstrated by a human being. In the content of the video, the multimodal agent does not need to directly participate in the task, but only needs to watch humans work as an observer, and then summarize the task nodes and the time sequence of the nodes based on the human work of this video. Since the multimodal agent will derive different descriptions of the same task in multiple summaries, which will lead to repeated task nodes, it is necessary to identify the repeated task nodes in the results obtained by the large model and eliminate redundant nodes. At the same time, in order to enable the multimodal agent to learn more sufficient sub-task relationships, it is necessary to perform simple manual screening on the guidance video so that the task steps it contains are as different as possible. After obtaining the task steps, the relationship between the tasks can be judged according to the order in which the task steps appear.

[0079] Constructing a task directed acyclic graph includes the following specific steps:

[0080] (a) By inputting the guidance video into the multimodal agent, the task nodes and the temporal sequence of nodes about the task are constructed according to the video information;

[0081] (b) Find the dependency set of all task nodes, no matter which path is followed, the prerequisites required by the task. For example: Path 1: A→B→C; Path 2: A→D→C. Then the dependency set of task C is {A}.

[0082] (c) Find the dependent set of all task nodes, which are the task nodes that will be executed no matter which path is taken. For example: Path 1: A→B→D→C; Path 2: A→D→C. Then the dependent set of task A is {C, D}.

[0083] (d) Remove irrelevant nodes, that is, nodes in the dependency set of the same task node and in its dependent set. For example: Path 1: A→B→C; Path 2: A→C→B. Then the dependency set of task B is {A, C}, and the dependent set is {C}, so B and C are irrelevant nodes.

[0084] (e) Mark potential dependencies and potential dependencies, that is, the preconditions and postconditions that exist in a certain path. For example: Path 1: A→B→C→D; Path 2: A→C. Then the potential dependency set of task C is {B}, and the potential dependency set is {D}.

[0085] (f) Mark mutually exclusive nodes. For example: Path 1: A→B→D; Path 2: A→C→D. Then B and C maintain a mutually exclusive relationship.

[0086] (g) All dependencies, potential dependencies and mutually exclusive relationships are listed as non-terms in the logical expression, and dependencies are listed as positive terms. For example: for path 1: A→B→C→D→E; path 2: A→D→C→F, its construction method is shown in Table 1 and Table 2.

[0087] Table 1 Node relationship table

[0088]

[0089] Table 2 Node relationship construction results table

[0090]

[0091] As shown in Figure 4, the steps of constructing a directed acyclic graph of tasks based on video learning by multimodal agents are as follows:

[0092] definition Different mission paths , the path includes Different task nodes ;

[0093] According to a task path , get the node The corresponding dependency set and the dependent set , as shown below:

[0094]

[0095] in, Represents a node N Split task paths for the center R ;

[0096] For a node , find the intersection of the dependency sets of all task paths The intersection of all dependent sets , get the task dependencies based on all task paths, as shown below:

[0097]

[0098] Since there may be multiple dependency paths between two nodes, it is necessary to clear redundant dependencies to obtain the nodes. The minimum set of dependencies and the minimum dependent set , as shown below:

[0099]

[0100] in, , What is contained in is the edge of the task directed acyclic graph, and a task directed acyclic graph based on all current task paths is obtained.

[0101] S2, real-time intention judgment;

[0102] S21, processing the multimodal data, and splicing the data capture images of different modalities together in a fixed manner;

[0103] S22. Use the task directed acyclic graph to screen the subtasks that the large model needs to face when making a judgment, and sort out some more likely subtask sequences for the large model;

[0104] In current multimodal data processing, most multimodal agents mainly support images and text as input. A few models, such as ChatGPT 4v and Gemini 1.5 Pro, can also process video stream input. However, considering the possible delay of multimodal agents when generating answers, in order to ensure the real-time nature of the experiment and the smoothness of content interaction, this embodiment chooses a method to generate image data every 5 seconds. The specific video processing method is as follows:

[0105] (1) For a single image as the data source: capture an image every 5 seconds.

[0106] (2) For data sources consisting of multiple images: capture an image every 1 second, and then stitch the images together in a fixed manner.

[0107] Considering that most videos are currently divided into horizontal and vertical videos, in order to ensure that the multimodal intelligent agent can read the pictures smoothly, the present invention adopts different splicing methods. For vertical videos, the present invention uses a delayed time sequence to splice pictures from left to right; for horizontal videos, the present invention first splices the first 4 seconds of video from the upper left clockwise into a field picture, and then splices the last second to the right of the field picture. The splicing format is as follows: Figure 3 , Figure 4 shown.

[0108] This processing method not only increases the frequency of image data generation, but also enhances the visualization effect of video content in the time dimension, which helps to better understand the performance of multimodal data in the experiment. At the same time, in order to facilitate the large model to understand the meaning of the spliced ​​pictures, the present invention designs the following prompt words to help the multimodal agent understand the meaning of the spliced ​​pictures:

[0109] # Vertical video:

[0110] The following images are stitched together from left to right at one-second intervals. This is the n-th set of images I am sending.

[0111] # Horizontal video:

[0112] The following set of images is cropped every second and stitchedtogether in the order of top left, top center, bottom left, bottom center, and bottom right. This is the n-th set of images I am sending.

[0113] In the intent recognition task, the multimodal agent needs to select the most reasonable task from a large number of subtasks based on the image. However, since the information in the image is discrete and fragmented, it may lead to a one-to-many relationship where one image corresponds to multiple subtasks, which is inconsistent with the one-to-one relationship where one image corresponds to one subtask in the present invention. Therefore, the present invention introduces a task directed acyclic graph to screen the subtasks that the large model needs to face when making a judgment at one time, and sort out some more likely subtask sequences for the large model.

[0114] At the same time, the present invention restricts and regulates the output of the multimodal agent so that it will ideally output a JSON-formatted answer as shown below:

[0115] {

[0116] "Name":<name of the task> ,

[0117] "items":<tools and items needed> ,

[0118] "description":<steps required for the task> ,

[0119] “Response”:<response spoken by the LLM> ,

[0120] "Step": {

[0121] “CurrentStep”:<the step that the student is currentlyworking on> ,

[0122] “NextStep”:<the next steps that the student needs to take>

[0123] }

[0124] }

[0125] In order to apply the directed acyclic graph to the task of intent recognition, the present invention stipulates the following operation steps:

[0126] (1) First, mark the task nodes with in-degree 0 in the task directed acyclic graph as executable nodes. Because the execution of these nodes does not depend on other nodes, when the large model performs the intent recognition task, it should make a choice from these nodes.

[0127] (2) When an executable node is executed, remove the node from the task directed acyclic graph, and then reselect a node with an in-degree of 0 from the new task directed acyclic graph;

[0128] (3) If all nodes have been executed, the task is completed; otherwise, return to step (1).

[0129] By using a directed acyclic graph containing logical expressions as input information, the multimodal large model can be assisted in making judgments about intent understanding, which improves the inference accuracy by about 15% compared to directly making judgments based on action information. In addition, combining spliced ​​image input containing temporal relationships can further improve the inference accuracy, with an overall increase of about 20%.

[0130] S3, generation and feedback of robot arm motion instructions;

[0131] S31. Issue corresponding instructions and execute corresponding steps according to the task directed acyclic graph to generate feedback data.

[0132] Since the relationship between collaborative actions and human intentions is not always a simple one-to-one mapping, the present invention achieves a more flexible and accurate alignment of intention understanding and human-machine collaborative actions by constructing a many-to-many mapping relationship between human intentions and robot arm actions. For example, when making a sandwich, the robot arm actions corresponding to the intention of "prepare ingredients" are "take bread", "take lettuce", etc.

[0133] However, due to the complexity of the environment and the uncertainty of the task, the actions of the robot arm may not always fully meet human expectations. In this case, if the execution results of the robot arm fail to meet human needs or goals, the system will start the reflection mechanism. By recombining the previous inputs (including directed acyclic graphs, spliced ​​image data, action recognition and intention judgment results, and human feedback), the information will be passed to the multimodal large model for reflection and re-evaluation. The multimodal large model needs to re-analyze and evaluate multiple factors: first, re-examine the accuracy of action recognition to ensure that the recognized actions are consistent with expectations; second, evaluate the rationality of intention inference to ensure that the inferred task goals meet human actual needs. Based on these evaluation results, the system will further improve the logical relationship between the nodes in the directed acyclic graph and optimize the many-to-many mapping relationship between human intentions and robot arm actions.

[0134] The present invention combines multimodal information such as vision and hearing to perceive the environment and task execution process in real time, designs an intelligent task planning method based on a large model, and successfully realizes the accurate recognition of human behavior and task inference in complex environments.

[0135] By building a large multimodal model that integrates behavior recognition, action prediction, and environmental understanding, the ability of existing intelligent agents to effectively adapt in dynamic environments has been improved, enabling the intelligent agents to autonomously infer the next task and respond.

[0136] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited thereto, and various changes can be made within the knowledge scope of technicians in the relevant technical field without departing from the purpose of the present invention.

Claims

1. A method for understanding intentions in a human-machine collaborative environment based on a multimodal large model, characterized in that: The specific steps include: S1, video analysis and task planning; S11, preprocessing the video using key frame extraction and image segmentation methods; S12. Build a directed acyclic graph of the task to memorize feasible task paths; the nodes in the directed acyclic graph of the task are used to represent the numerous actions in the task, and the arrows are used to represent the temporal logical order between the actions; Constructing a task directed acyclic graph includes the following specific steps: By inputting the guidance video into the multimodal agent, the task nodes and the time sequence of the nodes about the task are constructed according to the video information; Find the dependency set of all task nodes, and the prerequisites required by the task no matter which path is taken; Find the dependent set of all task nodes, and the task nodes that will be executed no matter which path is taken; Remove irrelevant nodes, that is, nodes that are in the dependency set of the same task node and in its dependent set; Mark potential dependencies and potential dependencies, that is, the preconditions and postconditions that exist in a certain path; Mark mutually exclusive nodes; All dependent, potential dependent and mutually exclusive relationships are listed as negative terms in the logical expression, and dependencies are listed as positive terms; S2, real-time intention judgment; S21, processing the multimodal data, and splicing the data capture images of different modalities together in a fixed manner; S22. Use the task directed acyclic graph to screen the subtasks that the large model needs to face when making a judgment, and sort out some more likely subtask sequences for the large model; Step S22 includes the following steps:

1. Mark the task nodes with in-degree 0 in the task directed acyclic graph as executable nodes; 2. When an executable node is executed, remove the node from the task directed acyclic graph, and then reselect a node with an in-degree of 0 from the new task directed acyclic graph; 3. Once all nodes have been executed, the task is completed; Otherwise, return to step 1; S3, generation and feedback of robot arm motion instructions; S31. Issue corresponding instructions and execute corresponding steps according to the task directed acyclic graph to generate feedback data.

2. The method for understanding intention based on a multimodal large model in a human-computer collaborative environment according to claim 1, characterized in that: The extraction of key frames in step S11 includes the following steps: The target action in the video is detected, the moving speed of the target is calculated, and the moving speed curve of the target is drawn; frames are extracted at the trough and peak sections of the speed curve.

3. The method for understanding intention based on a multimodal large model in a human-computer collaborative environment according to claim 1, characterized in that: The image segmentation in step S11 includes the following steps: First, the Grounding DINO method is used to realize the transformation from text description to object recognition. The position of the object in the picture is identified according to the specific text description, and the position coordinates of the object are returned. Then, the SAM2 model is used to segment the image according to the position coordinates of the object.

4. The method for understanding intention based on a multimodal large model in a human-computer collaborative environment according to claim 1, characterized in that: The steps to merge task paths into a directed acyclic graph are as follows: definition Different mission paths , the path includes Different task nodes ; According to a task path , get the node The corresponding dependency set and the dependent set , as shown below: in, Represents a node N Split task paths for the center R ; For a node , find the intersection of the dependency sets of all task paths The intersection of all dependent sets , get the task dependencies based on all task paths, as shown below: Since there may be multiple dependency paths between two nodes, it is necessary to clear redundant dependencies to obtain the nodes. The minimum set of dependencies and the minimum dependent set , as shown below: in, , What is contained in is the edge of the task directed acyclic graph, and a task directed acyclic graph based on all current task paths is obtained.

5. The method for understanding intention based on a multimodal large model in a human-computer collaborative environment according to claim 1, characterized in that: In step S21, image data is generated every 5 seconds; a picture is captured every 1 second; For vertical screen videos, use time-lapse sequence to stitch pictures from left to right; For horizontal screen videos, first splice the first 4 seconds of video into a field pattern clockwise from the upper left, and then splice the last second to the right side of the field pattern.

6. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for understanding intention based on a multimodal large model in a human-computer collaborative environment as described in any one of claims 1 to 5 is implemented.

7. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the method for understanding intentions based on a multimodal large model in a human-computer collaborative environment as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Directed acyclic graph based framework for training models

    CN111860753A

  • Voice Interaction Method and Apparatus, Terminal, and Storage Medium

    US20210183386A1