Operator intention recognition method based on vision-language large model

By combining a large vision-language model with multimodal interaction and multi-source perception data fusion, the problem of inaccurate positioning of the assembly area of ​​interest in human-machine collaboration was solved, enabling accurate recognition of operator intentions and efficient collaboration, thus improving the efficiency of human-machine collaboration.

CN120850197APending Publication Date: 2025-10-28DONGHUA UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510852677.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

In human-machine collaboration, inaccurate positioning of the assembly area of ​​interest makes it difficult to identify the operator's intentions. Existing large-scale language models cannot accurately analyze the operator's intentions. In particular, during visual analysis, the two operational behaviors show strong similarity and the objects in the workspace are small in size, which makes recognition difficult.

Method used

Design a semantic model for assembly tasks. By combining a large vision-language model with multimodal interaction, collect task instruction information, utilize object availability detection and real-time sensor fusion, and combine a multimodal human behavior recognition network. Through a gating mechanism, fuse multi-source perception data, analyze operator behavior and environmental changes, generate collaborative strategy suggestions, and adjust and optimize strategies in real time.

Benefits of technology

It improves the accuracy of operator intent recognition and the efficiency of human-machine collaboration, solves the problem of inaccurate positioning of the assembly area of ​​interest, and realizes efficient and stable complex human-machine collaboration tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120850197A_ABST
    Figure CN120850197A_ABST
Patent Text Reader

Abstract

The invention provides an operator intention recognition method based on a vision-language large model, and the method comprises the steps: constructing a standardized cue word template according to a target assembly object, operation demands and steps, environment information, cooperation requirements, natural expression habits and part features; task link reasoning is carried out in combination with task instruction information collected by an operator to generate task execution steps and assembly area visual information, dynamic cooperation data is generated based on the object availability detection technology in combination with real-time sensor fusion and environment perception, and then part control information and robot cooperation instructions are generated to serve as dynamic control information; by combining operator skeleton information and appearance texture features and fusing multi-source sensing data through a gating mechanism, feature weighted fusion is realized, operator intentions are obtained, internal correlation between human behaviors and assembly tasks is enhanced, irrelevant information interference is inhibited, accurate operator intention recognition is realized, man-machine cooperation efficiency is improved, and the method is suitable for popularization and application. The method has important theoretical significance and practical value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of large-scale language model analysis, human behavior recognition, and human-computer collaboration, specifically to a method for operator intent recognition based on a large vision-language model. Background Technology

[0002] The rapid development of the manufacturing industry has positioned human-machine collaboration as the dominant paradigm, combining the flexibility of human cognition with the precision of robots to improve efficiency.

[0003] In customized production, robots must actively perceive operators and components, typically using visual perception systems for scene understanding. However, visual perception systems often fail to accurately analyze the context of similar tasks or focus on the most relevant information, especially when two operational behaviors exhibit strong similarities during visual analysis. Furthermore, the small size of objects in the workspace makes operator intent recognition challenging.

[0004] Currently, the powerful analysis and reasoning capabilities of large language models have provided new insights into operator intent recognition in human-computer collaboration. Research generally emphasizes the potential of large language models to enable task-centered and efficient collaboration through intuitive language commands. However, the uncertainty of operator intent still requires dynamic guidance to meet the requirements of personalized product assembly, which is why pre-trained models are still unable to proactively adapt to specific tasks. Summary of the Invention

[0005] The technical problem that the present invention aims to solve is the difficulty in recognizing the operator's intentions due to inaccurate positioning of the assembly area of ​​interest.

[0006] The present invention provides a method for operator intent recognition based on a large vision-language model, comprising the following steps:

[0007] Based on the target assembly object, operational requirements and steps, environmental information, collaboration requirements, natural expression habits and part characteristics, design assembly task prompts, construct assembly task semantic model, and design standardized prompt templates to guide the visual perception system to focus on key components and work areas related to the operator's current task.

[0008] Based on standardized prompt word templates, task instruction information is collected through language or multimodal interaction with operators. This information is then input into a large language model and combined with an assembly process knowledge base for task link reasoning. The core semantic information in the operator's language, gestures, and actions is analyzed, and contextual information in the task instruction information is extracted. In conjunction with assembly process requirements, the task steps, component status, and operating environment during the assembly process are dynamically analyzed to generate task execution steps and visual information of the assembly area.

[0009] Based on the task execution steps and visual information of the assembly area, and using object availability detection technology combined with real-time sensor fusion and environmental perception, the system extracts the operational semantics of operator behavior, assembly task status, environmental changes, and data on parts and the operating environment in real time. It identifies whether the potential operational functions of the operator and robot are suitable for executing the current task, generates dynamic collaborative data from the data suitable for executing the current task, and then generates part control information and robot collaborative instructions in the assembly process. This clarifies the operator's and robot's manipulation intentions and collaborative needs as dynamic manipulation information.

[0010] Dynamic manipulation information is input into a multimodal human behavior recognition network. Combined with operator skeletal information and appearance texture features, the network uses a gating mechanism to fuse multi-source perception data to achieve feature weighted fusion. This allows for a comprehensive analysis of operator actions and behavior patterns, inference of the operator's specific action intentions and needs, collaboration requests and environmental changes in a particular task, and ultimately, the operator's intentions.

[0011] Based on the operator's intentions, the task object, operation requirements, and collaboration methods are extracted, and the assembly step sequence and collaboration strategy suggestions are output. The strategy is then adjusted and optimized in real time by combining dynamic environmental information and operator feedback.

[0012] Preferably, the standardized prompt template includes the operation objective, specific steps, tool requirements, and collaboration method.

[0013] Preferably, the key components include a control component, a movable hand, and an important assembly structure.

[0014] Preferably, the task instruction information includes operation steps, part information, and collaboration requirements.

[0015] Preferably, the dynamic collaborative data includes object characteristics, target assembly actions, and assembly sequence.

[0016] Preferably, the object characteristics include shape, position, and state.

[0017] Preferably, the combination of operator skeletal information and appearance texture features includes: extracting motion structure features through skeletal key points and capturing posture details and contact conditions using surface texture information.

[0018] Preferably, the assembly step sequence and collaboration strategy recommendations include a coverage of the part assembly sequence, operation action instructions, and robot collaboration path.

[0019] Preferably, the gating mechanism uses a gating matrix generated by 1×1 convolution, batch normalization, and the Sigmoid activation function.

[0020] Preferably, the optimization objective of the gating mechanism is:

[0021] P = -logw s

[0022] Among them, w s The weights assigned to the gating matrix represent the skeleton information, and P represents the priority of the gating mechanism in learning skeleton features.

[0023] The present invention proposes an operator intent recognition method based on a large vision-language model. By integrating the semantic analysis and reasoning capabilities of a large language model, object availability detection, and multimodal human behavior recognition technology, it improves the accuracy of operator intent recognition and the efficiency of system collaboration in human-machine collaboration. It solves the problems of inaccurate positioning of the assembly area of ​​interest and difficulty in operator intent recognition in human-machine collaboration scenarios. Attached Figure Description

[0024] Figure 1 A flowchart illustrating an operator intent recognition method based on a large vision-language model, provided for an embodiment of the present invention;

[0025] Figure 2 This invention provides a method for designing prompt words for assembly tasks.

[0026] Figure 3 The task analysis and reasoning process based on a large language model provided in this embodiment of the invention;

[0027] Figure 4 This invention provides a method for recognizing operator intent based on dynamic collaborative information. Detailed Implementation

[0028] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading the teachings of this invention, those skilled in the art can make various alterations or modifications to the invention, and these equivalent forms also fall within the scope defined by the appended claims.

[0029] like Figure 1 As shown, this embodiment of the invention provides a method for operator intent recognition based on a large vision-language model, characterized by comprising:

[0030] Step 1: As Figure 2As shown, by combining the target assembly object, operation requirements and steps, environmental information, collaboration requirements, natural expression habits and part characteristics, assembly task prompt words are designed, and an assembly task semantic model is constructed. Based on the task semantic model, a standardized prompt word template containing operation objectives, specific steps, tool requirements and collaboration forms is designed to guide operators to express themselves in a standardized manner, reduce the risk of language ambiguity and information omission, guide the visual perception system to focus on key components and work areas related to the operator's current task, and clarify the basic components and expression structure of the task.

[0031] The assembly task prompts, by combining natural language interaction with task semantic information, highlight key components (such as manipulators, active hands, and important assembly structures), guide the visual perception system to focus on the region of interest, and use the work area to represent the main work scope of human-machine collaboration. This reduces interference from irrelevant information in complex environments, clarifies the main work scope of human-machine collaboration, ensures that the prompts are concise, clear, and operationally relevant, and improves the semantic understanding effect of the visual perception system and the efficiency of human-machine collaboration.

[0032] The assembly task prompts are used to guide operators to clearly describe the target component, assembly steps, and / or collaboration requirements. For example, prompts may include instructions such as "insert the screw into the specified hole" or "select a wrench to tighten." In case of abnormal situations or during task confirmation, prompts can also guide operators to provide feedback, such as "Adjust robot position?" or "Continue to the next assembly step?". The prompt design combines natural language interaction, visual perception, and task semantic information. For example, when the camera detects a missing part, the system automatically issues a prompt, "Part missing detected, do you need to search?", ensuring clear task information expression, accurate intent recognition, and improved collaboration efficiency.

[0033] Step 2: Task semantic parsing and assembly process analysis. Based on standardized prompt word templates, the system interacts with the operator via language or multimodal methods to collect complete task instruction information, including operation steps, part information, and collaboration requirements, ensuring the integrity and reliability of the foundational data for subsequent intelligent analysis.

[0034] like Figure 3As shown, the collected task instruction information is input into a large-scale language model (LLM). Combined with an assembly process knowledge base, task chain reasoning is performed based on natural language understanding and reasoning. The core semantic information in the operator's language, gestures, and actions is analyzed, and contextual information such as assembly operations, target components, and collaboration requirements is extracted. Furthermore, combined with assembly process requirements, the task steps, component states, and operating environment during the assembly process are dynamically analyzed to generate task execution steps and visual information of the assembly area, assisting the system in understanding the operator's intentions and subsequent actions. This task semantic parsing is based on multimodal input, integrating language instructions, visual information, and environmental perception data to deeply mine the intention information behind the operator's expressions, improving the accuracy of task semantic mapping.

[0035] Step 3: Based on the task execution steps and visual information of the assembly area, and using object availability detection technology combined with real-time sensor fusion and environmental perception, the system extracts operational semantics of operator behavior, assembly task status, environmental changes, and component and operating environment data in real time. This generates dynamic collaborative data containing information such as component characteristics, status, target assembly actions, and assembly sequence. This data then generates component manipulation information and robot collaboration instructions during the assembly process, clarifying the operator's and robot's manipulation intentions and collaboration needs. As dynamic manipulation information, this information, through multimodal perception and analysis, supports real-time task allocation, action adjustment, and collaboration optimization, thereby achieving efficient and precise assembly.

[0036] For example, during the assembly of an aircraft wing, if the system detects that the operator is preparing to fix a screw, it can generate information such as the target position of the screw, the assembly sequence, the type of action and the tool requirements, and issue robot collaboration instructions, such as handing over tools or assisting in positioning, thereby improving the efficiency of human-machine collaboration.

[0037] The object availability detection technology senses components, tools, and operating space in the environment, analyzes object characteristics (such as shape, position, and state), identifies their potential operational functions for operators and robots, and determines whether they are suitable for performing the current task. This includes detecting the graspability of components, the usage status of tools, the accessibility of the workspace, and whether assembly actions meet the execution conditions. It supports real-time task allocation and strategy adjustment, and availability detection is a crucial foundation for optimizing collaborative task allocation and dynamic adjustment. Taking a screw as an example, the system detects its size, position, and state to confirm whether it meets assembly requirements; simultaneously, it determines whether the tool (such as an electric screwdriver) matches the task to ensure smooth operation.

[0038] Step 4: As Figure 4As shown, dynamic manipulation information is input into a multimodal human behavior recognition network. Combining operator skeletal information and appearance texture features, action structure features are extracted through skeletal key points. At the same time, surface texture information is used to capture posture details and contact situations. Through a gating mechanism, multi-source perception data is fused to achieve feature weighted fusion, comprehensively analyze operator action and behavior patterns, accurately infer the operator's specific action intentions and needs, collaboration requests and environmental changes in a specific task, improve the recognition accuracy of operator posture, action and interaction details, and obtain the operator's intention.

[0039] This method extracts the task object, operational requirements, collaboration methods, and potential intentions based on operator intent. It outputs a reasonable assembly step sequence and collaborative strategy suggestions, covering part assembly order, operational instructions, and robot collaboration paths. By combining dynamic environmental information and operator feedback, the strategy is adjusted and optimized in real time to ensure real-time understanding of operator intent, enhancing the inference effect of operator intent and making the collaboration process efficient, stable, and safe. This method integrates multimodal data, enabling more accurate identification of operator intent, thereby improving collaboration efficiency and accuracy.

[0040] For example, in the assembly of electronic devices, the system combines the skeletal motion path with the grasping state to identify when the operator is about to tighten a screw and adjust the robot strategy accordingly.

[0041] The gating mechanism employs a gating matrix generated by 1×1 convolution, batch normalization, and a sigmoid activation function. It dynamically allocates weights between skeleton and appearance information, emphasizing skeletal features to generate fused features. This ensures accurate intent recognition and improves the efficiency of multimodal information integration and recognition robustness. The optimization objective of the gating mechanism is:

[0042] P = -logw s

[0043] Where w s The gating matrix represents the weights assigned to the skeleton information, where P indicates the priority of the gating mechanism in learning skeleton features. The gating module only allows appearance features as input to the gating matrix when human actions are blurred based on given skeleton features.

[0044] The operator's intent refers to the goal or action plan that the operator expects to achieve in a specific assembly task, such as picking up a component, performing assembly steps, or adjusting the robot's task. By sensing the operator's behavior and changes in the environment, and combining this with task semantic analysis, the operator's intent can be accurately predicted, thereby achieving efficient collaboration between the robot and the operator.

[0045] In summary, to address the problem of inaccurate positioning of the assembly area of ​​interest, which leads to difficulty in recognizing operator intentions, this invention proposes an operator intention recognition method based on a vision-language large model. Assembly prompts are designed as input to the LLM (Learning-Language Model), the assembly task is analyzed, and positioning and guidance information for the assembly area is generated. Visual information of the assembly area is input into a dynamic collaboration information extraction model, and then the generated dynamic collaboration information and human behavior are jointly input into the intention recognition model. This solves the uncertainty of operator behavior, achieves operator intention recognition, and improves human-machine collaboration efficiency, possessing significant theoretical and practical value.

[0046] This invention optimizes the human-computer interaction information collection process through structured prompt word design, laying the foundation for subsequent task analysis and intent recognition. It improves the accuracy and efficiency of the system's understanding of assembly tasks. Based on prompt word guidance, it leverages a large language model to enhance the system's semantic understanding, logical reasoning, and strategy generation capabilities for assembly tasks. The dynamic collaborative information support mechanism and multimodal human behavior recognition technology achieve intelligent support for complex human-computer collaborative tasks, offering the following beneficial effects:

[0047] 1. To address the problem of traditional visual perception's difficulty in accurately locating assembly areas of interest, the design of task prompts and language-based large-scale model reasoning significantly improves the accuracy of operator intent recognition;

[0048] 2. To address the issues of randomness and susceptibility to environmental interference in operator behavior, a dynamic collaborative information support mechanism is proposed to analyze task context and operational requirements in real time, thereby mitigating the impact of behavioral uncertainty.

[0049] 3. By integrating multimodal information and gating mechanisms, the system can comprehensively capture details of operator behavior, improve its ability to accurately identify operator intentions, and help ensure the efficient and stable operation of human-machine collaboration.

[0050] The discussion of any of the above embodiments is merely for illustrative purposes and explanation, and is not intended to limit the invention. Those skilled in the art should understand that the technical features in the above embodiments of the invention can be combined, modified, equivalently substituted, or improved without departing from the spirit and principles of the invention, and should all be covered within the scope of protection of the claims of the invention.

Claims

1. A method for operator intent recognition based on a large vision-language model, characterized in that, Includes the following steps: Based on the target assembly object, operational requirements and steps, environmental information, collaboration requirements, natural expression habits and part characteristics, design assembly task prompts, construct assembly task semantic model, and design standardized prompt templates to guide the visual perception system to focus on key components and work areas related to the operator's current task. Based on standardized prompt word templates, task instruction information is collected through language or multimodal interaction with operators. This information is then input into a large language model and combined with an assembly process knowledge base for task link reasoning. The core semantic information in the operator's language, gestures, and actions is analyzed, and contextual information in the task instruction information is extracted. In conjunction with assembly process requirements, the task steps, component status, and operating environment during the assembly process are dynamically analyzed to generate task execution steps and visual information of the assembly area. Based on the task execution steps and visual information of the assembly area, and using object availability detection technology combined with real-time sensor fusion and environmental perception, the system extracts the operational semantics of operator behavior, assembly task status, environmental changes, and data on parts and the operating environment in real time. It identifies whether the potential operational functions of the operator and robot are suitable for executing the current task, generates dynamic collaborative data from the data suitable for executing the current task, and then generates part control information and robot collaborative instructions in the assembly process. This clarifies the operator's and robot's manipulation intentions and collaborative needs as dynamic manipulation information. Dynamic manipulation information is input into a multimodal human behavior recognition network. Combined with operator skeletal information and appearance texture features, the network uses a gating mechanism to fuse multi-source perception data to achieve feature weighted fusion. This allows for a comprehensive analysis of operator actions and behavior patterns, inference of the operator's specific action intentions and needs, collaboration requests and environmental changes in a particular task, and ultimately, the operator's intentions. Based on the operator's intentions, the task object, operation requirements, and collaboration methods are extracted, and the assembly step sequence and collaboration strategy suggestions are output. The strategy is then adjusted and optimized in real time by combining dynamic environmental information and operator feedback.

2. The operator intent recognition method based on a large vision-language model as described in claim 1, characterized in that, The standardized prompt template includes the operation objective, specific steps, tool requirements, and collaboration methods.

3. The operator intent recognition method based on a large vision-language model as described in claim 1, characterized in that, The key components include the control components, the movable hand, and important assembly structures.

4. The operator intent recognition method based on a large vision-language model as described in claim 1, characterized in that, The task instruction information includes operation steps, part information, and collaboration requirements.

5. The operator intent recognition method based on a large vision-language model as described in claim 1, characterized in that, The dynamic collaborative data includes object characteristics, target assembly actions, and assembly sequence.

6. The operator intent recognition method based on a large vision-language model as described in claim 1, characterized in that, The object's characteristics include shape, position, and state.

7. The operator intent recognition method based on a large vision-language model as described in claim 1, characterized in that, The combination of operator skeletal information and appearance texture features includes: extracting motion structure features through skeletal key points and capturing posture details and contact conditions using surface texture information.

8. The operator intent recognition method based on a large vision-language model as described in claim 1, characterized in that, The assembly step sequence and collaboration strategy recommendations include part assembly sequence, operation instructions and robot collaboration paths.

9. The operator intent recognition method based on a large vision-language model as described in claim 1, characterized in that, The gating mechanism uses a gating matrix generated by 1×1 convolution, batch normalization, and the Sigmoid activation function.

10. The operator intent recognition method based on a large vision-language model as described in claim 9, characterized in that, The optimization objective of the gating mechanism is: P=-logw s Among them, w s The weights assigned to the gating matrix represent the skeleton information, and P represents the priority of the gating mechanism in learning skeleton features.