Cooperative interaction method and device based on augmented reality and related equipment

By identifying critical task nodes in the AR environment and obtaining multimodal information, it transforms them into visual information display, and solves the problem of incomplete intention expression in AR collaboration and improves the efficiency of collaborative tasks.

CN120469575AActive Publication Date: 2025-08-12SUZHOU I MUSEUM DIGITAL TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510564356.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-12
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

During the AR collaboration process, the user's operation intention cannot be fully expressed, and the lack of a structured information transmission mechanism leads to inefficient collaboration tasks.

Method used

By identifying key task nodes of collaborative tasks in an augmented reality environment, obtaining user's operational behavior and voice information, using multimodal information to identify operational intentions, and converting them into visual information to display in other users' AR environments.

Benefits of technology

It realizes the accurate communication of collaboration intentions at key nodes of collaboration tasks, and improves the efficiency of collaboration tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120469575A_ABST
    Figure CN120469575A_ABST
Patent Text Reader

Abstract

The invention provides a cooperative interaction method and device based on augmented reality and related equipment, and relates to the technical field of augmented reality. According to the technical scheme provided by the invention, the cooperative task is acquired in the augmented reality environment, the key task node in the cooperative task is identified, when the user arrives at the key task node, the operation behavior information and the voice information of the user are acquired in time, and the operation intention information of the user is acquired through multi-modal information identification; meanwhile, operation intention information is converted into visual information and displayed in augmented reality environments of other users, accurate collection, understanding and visual transmission of operation intentions are achieved at key task nodes, and finally accurate transmission of cooperation intentions at key nodes of cooperation tasks is achieved; therefore, the completion efficiency of the cooperation task can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of augmented reality technology, and in particular to a collaborative interaction method, apparatus and related equipment based on augmented reality. Background Art

[0002] Augmented reality (AR) technology overlays virtual information onto the real world, providing users with real-time information enhancement and interactive experiences. With the development of AR technology, multi-user AR collaboration applications are becoming increasingly widespread, such as remote maintenance guidance, collaborative design, and telemedicine. During AR collaboration, multiple users share the same augmented reality environment through AR devices, enabling real-time information exchange and task collaboration.

[0003] Currently, during AR collaboration, users primarily communicate their operational intentions through voice calls, gesture annotations, and other methods. However, this simple interaction method cannot fully convey the operator's specific intentions and lacks a structured information transmission mechanism. Other users need to repeatedly confirm the operator's intentions to accurately understand them, resulting in inefficient information transmission during the collaboration process and low efficiency in completing collaborative tasks. Summary of the Invention

[0004] The present application provides a collaborative interaction method, apparatus, and related equipment based on augmented reality, which can accurately convey collaborative intentions at key nodes of collaborative tasks, thereby improving the efficiency of completing collaborative tasks.

[0005] In a first aspect, the present application provides a collaborative interaction method based on augmented reality, the method comprising: Acquire a collaborative task in an augmented reality environment, and identify key task nodes in the collaborative task; When the user's task progress reaches the key task node, the user's operation behavior information and voice information are obtained; Identify the operation behavior information and the voice information to obtain the user's operation intention information at the key task node; Visualization information is generated according to the operation intention information, and the visualization information is displayed in an augmented reality environment of other users.

[0006] By adopting the above technical solution, collaborative tasks are acquired in an augmented reality environment and key task nodes therein are identified. When a user reaches a key task node, the user's operation behavior information and voice information are acquired in a timely manner, and the user's operation intention information is obtained through multimodal information recognition. At the same time, the operation intention information is converted into intuitive visual information and displayed in the augmented reality environment of other users. By accurately collecting, understanding and visually conveying the operation intention at the key task nodes, the collaborative intention is finally accurately conveyed at the key nodes of the collaborative task, thereby improving the efficiency of completing the collaborative task.

[0007] Optionally, acquiring a collaborative task in an augmented reality environment and identifying key task nodes in the collaborative task include: Acquire collaborative tasks in augmented reality environments; Decomposing the collaborative task to obtain multiple task nodes; A dependency graph between the task nodes is constructed, and key task nodes in the collaborative task are determined according to the topological structure of the dependency graph.

[0008] By employing this technical solution, collaborative tasks are decomposed into multiple specific task nodes, and a dependency graph is constructed between these task nodes. Key task nodes are then identified based on the topological structure of the dependency graph. This method of task decomposition and dependency analysis makes the collaborative task structure clearer, and topological analysis of the dependency graph accurately identifies the key nodes crucial to task completion.

[0009] Optionally, constructing a dependency graph between the task nodes and determining key task nodes in the collaborative task according to the topological structure of the dependency graph includes: Obtaining node attributes of the task node; Constructing a dependency graph between the task nodes according to the task flow of the collaborative task and the node attributes of the task nodes; Traversing the dependency graph to determine a key task path in the dependency graph; The number of predecessor nodes of the task node on the critical task path is recorded, and the task node whose number of predecessor nodes is greater than a preset threshold is determined as the critical task node in the collaborative task.

[0010] By employing this technical solution, a dependency graph between task nodes is constructed based on the task flow of collaborative tasks. This dependency graph is then traversed to determine the critical task path. Finally, critical task nodes are identified based on a comparison of the number of predecessor nodes with a preset threshold. This analysis method, based on node attributes and dependency relationships, makes the identification of critical task nodes more objective and quantitative. Statistical analysis of the number of predecessor nodes can effectively identify nodes with high levels of dependency within the task flow.

[0011] Optionally, the identifying the operation behavior information and the voice information to obtain the operation intention information of the user at the key task node includes: Formatting the operation behavior information to obtain a plurality of first operation information in the form of event streams; Matching the first operation information with the key task node, and filtering out second operation information related to the key task node; extracting a first operation type and a first operation object according to voice keywords in the voice information; extracting a second operation type and a second operation object from the second operation information; Identify the first operation type, the first operation object, the second operation type, and the second operation object, and obtain the user's operation intention information at the key task node, where the operation intention information is structured data.

[0012] By adopting the above technical solution, since the operation behavior information is converted into a standard event stream format, the system can effectively filter out the operation information related to the current key task node, avoiding the interference caused by irrelevant operations; at the same time, by simultaneously analyzing the two different dimensions of input, voice information and operation behavior information, they can verify and complement each other, thereby improving the accuracy and completeness of operation intention recognition.

[0013] Optionally, the identifying the first operation type, the first operation object, the second operation type, and the second operation object to obtain the user's operation intention information at the key task node includes: Determining whether the first operation object and the second operation object are the same operation object; If the first operation object and the second operation object are not the same operation object, determining the first operation object or the second operation object as the target operation object according to the standard operation object corresponding to the key task node; Determining, according to the operation type corresponding to the target operation object, a main operation type and an auxiliary operation type in the first operation type and the second operation type; The operation intention information of the user at the key task node is determined according to the main operation type, the auxiliary operation type and the target operation object.

[0014] By adopting the above technical solution, it is determined whether the first operation object and the second operation object extracted from the voice information and the operation behavior information are consistent. When inconsistency is found, the judgment is made based on the standard operation object corresponding to the key task node, so as to accurately determine the target operation object, avoiding the intention understanding error caused by the inconsistency of multimodal information; then, according to the determined target operation object and its corresponding operation type, the first operation type and the second operation type are divided into the main operation type and the auxiliary operation type. This hierarchical operation type division enables the system to highlight the key operations while retaining the auxiliary information; finally, by combining the main operation type, the auxiliary operation type and the target operation object, a complete operation intention information is formed, which realizes the accurate understanding and structured expression of the user's intention, improves the interaction accuracy and efficiency in the AR collaboration process, and reduces the collaboration errors caused by the deviation of intention understanding.

[0015] Optionally, determining the main operation type and the auxiliary operation type in the first operation type and the second operation type according to the operation type corresponding to the target operation object includes: If the target operation object is the first operation object, determining the first operation type as a primary operation type and determining the second operation type as an auxiliary operation type; If the target operation object is the second operation object, the second operation type is determined as the main operation type, and the first operation type is determined as the auxiliary operation type.

[0016] By adopting the above technical solution, the operation type associated with the target operation object is determined as the primary operation type, and the other operation type is determined as the auxiliary operation type, thereby establishing a mechanism for determining the primary and secondary relationships of operation types. When the target operation object is the first operation object, the corresponding first operation type is determined as the primary operation type, and the second operation type is determined as the auxiliary operation type, and vice versa. This method of dividing the primary and secondary relationships of operation types based on the relevance of the target operation object ensures logical consistency in the operation intention recognition process, while retaining the association characteristics of different modal information, thereby being able to more accurately restore the user's true operation intention and improve the accuracy of intention understanding in the AR collaboration process.

[0017] Optionally, the visual information includes an operation intention animation, an operation intention subtitle, and an operation intention mark. Converting the operation intention information into visual information and displaying the visual information in an augmented reality environment of other users includes: generating the operation intention animation, the operation intention subtitles, and the operation intention mark according to the operation intention information; Determining target users who participate in the collaborative task and are associated with the key task nodes; The operation intention animation, the operation intention subtitles and the operation intention mark are sent to the augmented reality device of the target user, so that the augmented reality device of the target user displays the operation intention animation, the operation intention subtitles and the operation intention mark in an augmented reality environment.

[0018] By adopting the above technical solution, the operation intention information is converted into multi-dimensional visual information including operation intention animation, operation intention subtitles and operation intention marks, thereby realizing the three-dimensional expression of the operation intention; at the same time, by identifying the target users associated with key task nodes, it is ensured that the visual information can be accurately pushed to the relevant collaborators, avoiding the excessive transmission of information; when these visual information are sent to the target user's augmented reality device and displayed in its AR environment, the target user can intuitively understand the operation action through the animation, obtain detailed operation instructions through the subtitles, and quickly locate the focus through the marks. This multi-dimensional visual presentation method makes the transmission of operation intention clearer and more intuitive, significantly improves the information transmission efficiency and understanding accuracy in the AR collaboration process, and reduces the communication cost and understanding deviation in the collaboration process.

[0019] In a second aspect, the present application provides a collaborative interaction device based on augmented reality, the device comprising: A node identification module is used to obtain collaborative tasks in an augmented reality environment and identify key task nodes in the collaborative tasks; The acquisition module is used to obtain the user's operation behavior information and voice information when the user's task progress reaches the key task node; An intention recognition module is used to recognize the operation behavior information and the voice information to obtain the user's operation intention information at the key task node; The display module is configured to generate visual information according to the operation intention information and display the visual information in an augmented reality environment of other users.

[0020] In a third aspect, the present application provides a computer storage medium, which stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing any one of the above methods.

[0021] In a fourth aspect, the present application provides an electronic device comprising a processor, a memory and a transceiver, wherein the memory is used to store instructions, the transceiver is used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device performs any one of the above methods.

[0022] In summary, the beneficial effects brought about by the technical solution of this application include: By adopting the above technical solution, collaborative tasks are acquired in an augmented reality environment and key task nodes therein are identified. When a user reaches a key task node, the user's operation behavior information and voice information are acquired in a timely manner, and the user's operation intention information is obtained through multimodal information recognition. At the same time, the operation intention information is converted into intuitive visual information and displayed in the augmented reality environment of other users. By accurately collecting, understanding and visually conveying the operation intention at the key task nodes, the collaborative intention is finally accurately conveyed at the key nodes of the collaborative task, thereby improving the efficiency of completing the collaborative task. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 This is a flowchart of a collaborative interaction method based on augmented reality according to an embodiment of the present application; Figure 2 This is a schematic structural diagram of a collaborative interaction device based on augmented reality according to an embodiment of the present application; Figure 3 This is a structural diagram of an electronic device provided in an embodiment of the present application.

[0024] Description of reference numerals: 300, electronic device; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. DETAILED DESCRIPTION

[0025] In order to enable people skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.

[0026] In the description of the embodiments of this application, words such as "exemplary," "for example," or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary," "for example," or "for example" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary," "for example," or "for example" is intended to present the relevant concepts in a concrete manner.

[0027] In the description of the embodiments of the present application, the term "multiple" means two or more. For example, multiple devices refer to two or more devices, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.

[0028] See Figure 1 , which is a flow chart of a collaborative interaction method based on augmented reality provided in an embodiment of the present application. This method can be implemented using a computer program, a single-chip microcomputer, or run on an augmented reality-based collaborative interaction device based on a von Neumann architecture. The computer program can be integrated into an application or run as a standalone tool application. The specific steps of the collaborative interaction method based on augmented reality are described in detail below.

[0029] S101: Acquire collaborative tasks in an augmented reality environment and identify key task nodes in the collaborative tasks; An augmented reality environment refers to a hybrid environment created by overlaying virtual information onto the real world through computer technology. This environment includes a fusion display of the real physical environment and computer-generated digital information (such as 3D models, text, images, audio, etc.). In the embodiments of this application, it can be understood as a hybrid visual environment perceived by users through augmented reality devices (such as AR glasses, head-mounted displays, etc.). This environment includes the real physical scene, the virtual interactive interface, digital information related to the collaborative task, and visualization information of the operational intentions of other collaborating users.

[0030] A collaborative task refers to a task process that requires the participation and completion of multiple users. During this process, each user completes their respective task responsibilities based on a common task goal through information sharing and interactive cooperation. In the embodiments of this application, it can be understood as a task with a clear workflow performed by multiple users in an augmented reality environment. The task can be broken down into multiple interdependent task nodes, each of which may involve the operation behavior and voice interaction of one or more users.

[0031] Among them, key task nodes refer to important task execution points in the collaborative task execution process that play a key role in task completion and require the cooperation of multiple users. In the embodiments of this application, it can be understood as: execution nodes that require special attention in the collaborative task process and have a significant impact on the progress of subsequent tasks.

[0032] In this embodiment, since collaborative tasks typically involve multiple execution steps, each with varying degrees of impact on task completion, it is necessary to identify the task nodes that are critical to task completion. Specifically, the collaborative tasks in the augmented reality environment are first acquired. The acquired collaborative tasks are then analyzed and decomposed, breaking the overall task into multiple specific execution steps. By analyzing the importance and impact of these execution steps within the task, the key task nodes are identified.

[0033] Based on the above embodiment, as an optional implementation, the identification of key mission nodes may be implemented using steps S201 to S203.

[0034] S201: Acquire collaborative tasks in an augmented reality environment; A collaborative task is one that requires multiple users to complete together through augmented reality devices. In practice, the augmented reality device receives task information input by the user, or retrieves a preset collaborative task from a task management system. The retrieved collaborative task information includes basic information such as the task type, task description, and participating users.

[0035] S202: Decompose the collaborative task to obtain multiple task nodes; In this embodiment, since collaborative tasks are often complex and directly executing the entire task is difficult, it is necessary to decompose the task into smaller, more easily executed task nodes. In specific implementations, based on the task's execution flow and logical relationships, the overall task is split into multiple relatively independent execution units, each of which constitutes a task node. The task node contains basic attribute information about the execution unit, such as execution order and completion conditions.

[0036] S203: Construct a dependency graph between task nodes, and determine key task nodes in the collaborative task based on the topological structure of the dependency graph.

[0037] Among them, the dependency graph between task nodes refers to a data structure that describes the execution order and mutual relationship between each task node in a collaborative task. In the embodiment of the present application, it can be understood as: a graph structure composed of nodes and directed edges, wherein the nodes represent task execution units, and the directed edges represent the execution order between task nodes. Each node contains node attribute information, which is used to characterize the characteristics of the node in the entire task process. The dependency graph is used to analyze the node correlation in the task process, determine the critical path of task execution, and thus find the task nodes that play a key role in task completion.

[0038] After obtaining the task nodes, it is necessary to construct a dependency graph between the task nodes and, based on its topological structure, identify the key task nodes in the collaborative task. Because task nodes have an execution order and interdependencies, a dependency graph is needed to express this correlation and identify the nodes that have a significant impact on task completion. In specific implementation, a dependency graph between task nodes is constructed based on the task flow of the collaborative task. This graph structure represents task execution units using nodes and the execution order between nodes using directed edges. Based on the constructed dependency graph, the key task nodes in the collaborative task can be identified by analyzing its topological structure.

[0039] Specifically, step S203 also includes S301-S304.

[0040] S301: Obtain node attributes of the task node; Node attributes are a set of data describing the characteristics of a task node. In practice, these data represent basic characteristic parameters for each task node, such as its execution order, completion conditions, and time limits. By acquiring these data, we can accurately characterize the characteristics of each task node within the overall task flow.

[0041] S302: Constructing a dependency graph between task nodes based on the task flow of the collaborative task and the node attributes of the task nodes; During implementation, since collaborative tasks have specific execution processes and each task node has unique attribute characteristics, this information needs to be comprehensively considered to construct an accurate dependency graph. During implementation, the overall execution process of the collaborative task is first analyzed to clarify the execution order requirements for each link, and then the task nodes are arranged as nodes in the graph. The execution order between nodes is then determined based on the task flow. Nodes with a sequential execution relationship are connected via directed edges, and the specific association method between nodes is determined based on the node attributes. In the resulting dependency graph, each node represents a task execution unit, and the directed edges between nodes represent the execution order and dependency relationships.

[0042] S303: traverse the dependency graph to determine the key task path in the dependency graph; Because the dependency graph contains multiple possible paths from the starting node to the ending node, some of which have a decisive impact on the completion of the entire collaborative task, it is necessary to identify these critical task paths. In its implementation, we first begin at the starting node of the dependency graph and use a depth-first search to traverse the entire graph structure, recording each complete path from the starting node to the ending node. During this traversal, we comprehensively consider the strength of the dependencies between the nodes on the path and the importance of the nodes in the overall task process to evaluate the impact of each path on task completion. By comparing the impact of different paths, we determine the paths that play a key role in task completion, namely the critical task paths.

[0043] S304: Record the number of predecessor nodes of the task node on the critical task path, and determine the task node with the number of predecessor nodes greater than a preset threshold as the critical task node in the collaborative task.

[0044] Wherein, the number of preceding nodes refers to the number of other nodes pointing to a certain task node in the dependency graph. In the embodiment of the present application, it can be understood as the total number of all task nodes that need to be completed before any task node starts executing.

[0045] Since the number of predecessor nodes reflects the dependency complexity of the task node, the larger the number, the more constraints the node is subject to, and the greater the impact on the smooth completion of the entire collaborative task. In the specific implementation, first traverse each task node on the critical task path and count the number of other nodes pointing to the node, that is, calculate the number of predecessor nodes of each node. Then compare the number of predecessor nodes of each node with a pre-set threshold. When the number of predecessor nodes of a node exceeds the preset threshold, it indicates that the node has a high dependency complexity and requires special attention and management. Therefore, it is identified as a critical task node in the collaborative task. The critical task nodes identified in this way are not only on the critical path of task execution, but also have a high dependency complexity. The execution of these nodes will directly affect the progress of the entire collaborative task.

[0046] S102: When the user's task progress reaches a critical task node, obtain the user's operation behavior information and voice information; Operational behavior information refers to the specific behavioral data recorded by task participants during the execution of task nodes. In the present application, this information can be understood as key behavioral characteristics such as the operation type, operation time, and operation results during the execution of key task nodes. Operational behavior information is used to record and reflect the actual situation of task execution.

[0047] Among them, voice information refers to the verbal communication content generated by the task participants during the task execution process. In the embodiment of the present application, it can be understood as: voice interaction data such as voice dialogue, instruction transmission, and exchange of opinions between participants during the task execution process.

[0048] Since key task nodes have a significant impact on the completion of the entire collaborative task, it is necessary to comprehensively record the user's execution status at these nodes. In specific implementation, the system monitors the user's task progress in real time. When it detects that the user has reached a key task node, it begins to collect the user's operational behavior information, including key behavioral characteristics such as operation type, operation time, and operation results. At the same time, the voice collection function is activated to record the user's voice commands in the AR environment, voice conversations with other participants, and other voice interaction content. By simultaneously collecting operational behavior information and voice information, it is possible to comprehensively record the user's execution process at key task nodes, including both specific operational actions and voice interactions during the decision-making process.

[0049] S103: Identify the operation behavior information and voice information to obtain the user's operation intention information at the key task node; Among them, operation intention information refers to the inference of the user's desired goal or the next action to be performed by analyzing the user's performed operation behavior and voice information. This intention information reflects the purpose of the user's behavior, rather than the behavior process itself. In the embodiment of the present application, operation intention information can be understood as identifying the real goal that the user wants to achieve by analyzing the combination of the user's actual operation behavior and voice information at key task nodes.

[0050] In the embodiment of the present application, the purpose of identifying the operational behavior information and voice information is to improve the accuracy of understanding the user's behavioral intentions by combining and analyzing these two types of information. In actual applications, when the user's operational behavior information and voice information are obtained at the same time, these two types of information contain the user's current task intentions that they want to perform. By identifying and analyzing these two types of information, the operation types and operation objects contained in the operational behavior information and voice information are extracted, and then the extracted information is associated with the key task nodes for analysis, thereby obtaining the user's operational intention information at the key task nodes. This identification method can not only accurately understand the user's current operational intentions, but also effectively avoid the identification bias that may be caused by a single information source.

[0051] Based on the above embodiment, as an optional implementation, step S103 further includes steps S401-S405.

[0052] S401: Formatting the operation behavior information to obtain a plurality of first operation information in the form of event streams; The first operation information refers to a formatted data structure of operation behavior, which contains information about independent operation units performed by the user at a specific time point. In the embodiments of the present application, it can be understood as event stream data obtained by decomposing the user's continuous operation behavior in chronological order. Each event stream contains attribute information such as the timestamp of the operation, the operation type, and the operation object.

[0053] In an embodiment of the present application, the operation behavior information is formatted to obtain the first operation information in the form of several event streams, mainly to convert the user's continuous operation behavior into a standardized data format to facilitate subsequent processing and analysis. In actual applications, the user's operation behavior information is often a continuous and complex combination of actions. By formatting these operation behavior information, it can be split into multiple independent event streams. The formatting process includes timing analysis and action decomposition of the operation behavior information, and dividing the continuous operation behavior into independent operation units in chronological order. Each operation unit contains attribute information such as operation time, operation type and operation object. These independent operation units constitute the first operation information in the form of event streams.

[0054] S402: Match the first operation information with the key task node, and filter out the second operation information related to the key task node; The second operation information refers to a set of operation information that has been matched and filtered by the key task node, and it retains all operation behavior data related to the current task node. In the embodiment of the present application, it can be understood as the operation information related to the current key task node that is retained after the operation type and operation object in the first operation information are matched with the standard operation type and standard operation object preset by the key task node.

[0055] In an embodiment of the present application, the purpose of matching the first operation information with the key task node is to filter out the operation information related to the current task and avoid interference from irrelevant operation information. In actual applications, the user may perform some operation behaviors that are not related to the current task, and these operation behaviors will also be converted into the first operation information. By matching the operation type and operation object in the first operation information with the standard operation type and standard operation object preset in the key task node, the successfully matched first operation information can be filtered out as the second operation information. This matching and screening method can effectively filter out operation information that is not related to the current task, and the obtained second operation information more accurately reflects the user's relevant operation behavior at the key task node. S403: Extracting a first operation type and a first operation object based on voice keywords in the voice information; In an embodiment of the present application, the first operation type and the first operation object are extracted based on the voice keywords in the voice information, with the purpose of understanding the type of operation the user wants to perform and the target object of the operation by analyzing the user's voice instructions. In actual applications, the user's voice information usually contains a verb phrase representing the operation behavior and a noun phrase representing the operation target. By performing semantic analysis on the voice information, the verb phrase is identified as the first operation type, and the noun phrase is identified as the first operation object. This extraction method can obtain the key elements of the operation intention from the user's voice description, so that the voice information can be converted into structured data that can be used for intent recognition.

[0056] S404: Extracting the second operation type and the second operation object from the second operation information; In an embodiment of the present application, the second operation type and the second operation object are extracted from the second operation information in order to obtain the key elements of the operation intention from the user's actual operation behavior. In actual applications, the second operation information has been screened by key task nodes and contains operation behavior data related to the current task. By analyzing the structural characteristics of these operation behavior data, the action features therein can be identified as the second operation type, and the target of the action can be identified as the second operation object. This extraction method can obtain the key information of the operation intention from the user's actual operation behavior, so that the operation behavior data can be converted into structured information that can be used for intent identification.

[0057] S405: Identify the first operation type, the first operation object, the second operation type, and the second operation object, and obtain the user's operation intention information at the key task node, where the operation intention information is structured data.

[0058] In an embodiment of the present application, the user's operation intention information at the key task node is obtained by identifying the first operation type, the first operation object, the second operation type, and the second operation object. The purpose is to integrate and analyze the operation information at the voice level and the behavior level, so as to more accurately understand the user's true intention. In actual applications, the first operation type and the first operation object reflect the user's intention expressed through voice, and the second operation type and the second operation object reflect the user's intention expressed through actual behavior. Comprehensively identifying this information can construct a structured data containing the operation type and the operation object. This form of structured data can not only fully describe the user's operation intention, but also facilitate data transmission and processing in a collaborative environment.

[0059] Based on the above embodiment, as an optional implementation, step S405 further includes steps S501-505.

[0060] S501: Determine whether the first operation object and the second operation object are the same operation object.

[0061] In an embodiment of the present application, the purpose of determining whether the first operation object and the second operation object are the same operation object is to verify whether the operation object expressed by the user through voice is consistent with the operation object in the actual operation behavior, thereby ensuring the accuracy of the recognized operation intention. In actual applications, by comparing the feature information of the first operation object and the second operation object, such as object type, object attributes, etc., it is determined whether the two operation objects point to the same entity. When the two operation objects are consistent, it indicates that the user's voice expression and actual operation behavior are pointing to the same target. This consistency can enhance the credibility of the operation intention recognition; when the two operation objects are inconsistent, it indicates that there is a difference between the user's voice description and the actual operation, and further analysis and processing of this inconsistency is required to ensure that the final recognized operation intention can accurately reflect the user's true intention.

[0062] S502: If the first operation object and the second operation object are not the same operation object, the first operation object or the second operation object is determined as the target operation object according to the standard operation object corresponding to the key task node.

[0063] In an embodiment of the present application, when the first operation object is inconsistent with the second operation object, it is necessary to determine the target operation object through the standard operation object corresponding to the key task node. This is because in actual application scenarios, the user's voice expression and actual operation behavior may be inconsistent. By comparing the matching degree of the first operation object and the second operation object with the standard operation object preset by the key task node, the operation object with higher matching degree is selected as the target operation object. This method can provide a reliable basis for judgment when the operation objects are inconsistent. In specific implementation, the feature similarity between the first operation object and the second operation object and the standard operation object can be calculated, such as object type similarity, attribute feature similarity, etc., and the operation object with higher similarity can be determined as the target operation object.

[0064] S503: Determine a main operation type and an auxiliary operation type in the first operation type and the second operation type according to the operation type corresponding to the target operation object.

[0065] In the embodiment of the present application, the main operation type and the auxiliary operation type are determined according to the operation type corresponding to the target operation object, with the purpose of performing a hierarchical analysis on the user's operation behavior and identifying the core operation and supporting operations. In actual applications, the target operation object usually has a standard operation type set associated with it. By matching and analyzing the first operation type and the second operation type with this standard operation type set, it can be determined which operation types are the main operations for the target operation object and which are auxiliary operations. In specific implementation, the standard operation type set corresponding to the target operation object is first obtained, and then the first operation type and the second operation type are calculated for similarity with the operation types in the set. The operation type with higher similarity and directly acting on the target operation object is determined as the main operation type, while the operation type with lower similarity or playing an auxiliary role is determined as the auxiliary operation type.

[0066] Based on the above embodiment, as an optional implementation manner, if the target operation object is the first operation object, the first operation type is determined as the main operation type, and the second operation type is determined as the auxiliary operation type; If the target operation object is the second operation object, the second operation type is determined as the main operation type, and the first operation type is determined as the auxiliary operation type.

[0067] In an embodiment of the present application, the main operation type and the auxiliary operation type are determined according to whether the target operation object comes from voice expression or actual operation behavior. This determination method grades the operation type based on the source of the operation object. When the target operation object is the first operation object, it means that the user's voice expression more accurately points to the task goal, so the first operation type is determined as the main operation type, and the second operation type is determined as the auxiliary operation type; when the target operation object is the second operation object, it means that the user's actual operation behavior is more in line with the task requirements, so the second operation type is determined as the main operation type, and the first operation type is determined as the auxiliary operation type. In actual applications, this grading method can automatically adjust the primary and secondary relationships of the operation types according to the source of the operation object, ensuring that the final recognized operation intention is more in line with the user's true intention.

[0068] S504: Determine the user's operation intention information at the key task node according to the main operation type, the auxiliary operation type, and the target operation object.

[0069] In the embodiment of the present application, the user's operation intention information at the key task node is determined based on the main operation type, auxiliary operation type and target operation object, with the aim of constructing a complete and structured description of the operation intention. In actual applications, by taking the main operation type as the core action, the auxiliary operation type as the coordinated action, and the target operation object as the operation target, these elements are combined according to the preset intention template to form a structured data describing the user's complete operation intention. In specific implementation, a template structure of "main operation type + auxiliary operation type + target operation object" can be adopted, in which the main operation type determines the main purpose of the operation, the auxiliary operation type provides auxiliary information of the operation, and the target operation object clarifies the action target of the operation.

[0070] S505: If the first operation object and the second operation object are the same operation object, determine the user's operation intention information at the key task node according to the first operation type, the second operation type and the same operation object.

[0071] In an embodiment of the present application, when the first operation object and the second operation object are the same operation object, the user's operation intention information is determined directly based on the first operation type, the second operation type and the same operation object. This is because the consistency of the operation objects indicates that the user's voice expression and actual operation behavior have a good correspondence. In actual applications, combining the first operation type and the second operation type can more comprehensively understand the user's operation behavior, wherein the first operation type reflects the user's operation intention expressed through voice, and the second operation type reflects the user's operation intention expressed through actual behavior. The two act together on the same operation object to form a complete description of the operation intention. In specific implementation, the first operation type and the second operation type can be fused and analyzed, and combined with the feature information of the same operation object to construct a structured data containing an operation type combination and an operation object. This structured data can accurately describe the user's operation intention at the current task node.

[0072] S104: Generate visualization information according to the operation intention information, and display the visualization information in the augmented reality environment of other users.

[0073] Visual information refers to a collection of information that intuitively displays the operation intent through visual elements and graphical identifiers. It converts abstract operation intent into a visible graphical representation. In the embodiments of this application, it can be understood as a visual effect combination consisting of a primary operation type identifier, an auxiliary operation type identifier, and a target operation object identifier. These identifiers may include specific visual elements such as an operation direction arrow, a dotted line along the operation path, and a highlighted box for the target object.

[0074] In an embodiment of the present application, visualization information is generated based on the operation intention information and displayed in the augmented reality environment of other users. The purpose is to convert the user's operation intention into an intuitive and visible visual effect, so that other users in the collaborative environment can understand and grasp the current user's operation behavior. In actual applications, the system first parses the structured operation intention information into visualization elements, including converting the main operation type into the main visual identifier, converting the auxiliary operation type into an auxiliary visual identifier, and converting the position and characteristics of the target operation object into spatial positioning information, and then superimposes and displays these visualization effects in the augmented reality devices of other users. In specific implementation, different visual styles can be used to distinguish between main operations and auxiliary operations, such as using eye-catching arrows to indicate the main operation direction, using dotted lines or translucent effects to indicate the auxiliary operation path, and adding highlight prompts or interactive marks around the target operation object. This visualization method can help other users quickly understand the current user's operation intention.

[0075] Based on the above embodiment, as an optional implementation, an operation intention animation, an operation intention subtitle, and an operation intention mark are generated according to the operation intention information; Identify target users who will participate in collaborative tasks and be associated with key task nodes; The operation intention animation, operation intention subtitles and operation intention marks are sent to the target user's augmented reality device, so that the target user's augmented reality device displays the operation intention animation, operation intention subtitles and operation intention marks in the augmented reality environment.

[0076] In the embodiment of the present application, the operation intention information is converted into multiple visual forms such as operation intention animation, operation intention subtitles and operation intention marks, with the purpose of enhancing the effect of collaborative interaction through a multi-dimensional information display method. In actual application, first, a dynamic operation intention animation is generated based on the operation intention information to show the process of the operation behavior, and at the same time, a text-based operation intention subtitle is generated to provide a semantic description, and a spatially positioned operation intention mark is generated to indicate the specific operation location. Then, the target users related to the current task node are identified. These users are collaborators who participate in the collaborative task and are associated with the current key task node. These visual information is then sent to the target user's augmented reality device so that it can see the complete expression of the operation intention in their respective augmented reality environments. It is worth noting that when different users arrive at the same key task node one after another, the operation intention information of other users will be automatically identified and displayed in their respective augmented reality environments, realizing information sharing and interactive synchronization at the task node level. In specific implementation, the operation intention animation can show the dynamic process of the operation, the operation intention subtitle can provide a text description, and the operation intention mark can indicate the spatial position. The three visual forms complement each other and jointly build a comprehensive operation intention display system.

[0077] This multi-dimensional visual display method not only helps target users understand the operational intentions of other users from different perspectives, but also ensures the accuracy and completeness of information transmission during the collaboration process, significantly improving the interaction effect and collaboration efficiency in the augmented reality collaboration environment.

[0078] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0079] See Figure 2 , which shows a schematic diagram of the structure of an augmented reality-based collaborative interaction device provided by an exemplary embodiment of the present application. The device can be implemented as all or part of the device through software, hardware, or a combination of both. The augmented reality-based collaborative interaction device includes: Node identification module, used to obtain collaborative tasks in the augmented reality environment and identify key task nodes in the collaborative tasks; The acquisition module is used to obtain the user's operation behavior information and voice information when the user's task progress reaches the key task node; The intention recognition module is used to identify operation behavior information and voice information to obtain the user's operation intention information at key task nodes; The display module is used to generate visual information based on the operation intention information and display the visual information in the augmented reality environment of other users.

[0080] Based on the above embodiment, as an optional embodiment, the node identification module is also used to obtain collaborative tasks in an augmented reality environment; decompose the collaborative tasks to obtain multiple task nodes; construct a dependency graph between task nodes, and determine the key task nodes in the collaborative tasks based on the topological structure of the dependency graph.

[0081] Based on the above embodiments, as an optional embodiment, the node identification module is also used to obtain the node attributes of the task nodes; construct a dependency graph between the task nodes based on the task flow of the collaborative task and the node attributes of the task nodes; traverse the dependency graph to determine the critical task path in the dependency graph; record the number of predecessor nodes of the task node on the critical task path, and determine the task node with a number of predecessor nodes greater than a preset threshold as the critical task node in the collaborative task.

[0082] Based on the above embodiments, as an optional embodiment, the intention recognition module is also used to format the operation behavior information to obtain several first operation information in the form of event streams; match the first operation information with the key task node to filter out the second operation information related to the key task node; extract the first operation type and the first operation object based on the voice keywords in the voice information; extract the second operation type and the second operation object in the second operation information; identify the first operation type, the first operation object, the second operation type and the second operation object to obtain the user's operation intention information at the key task node, and the operation intention information is structured data.

[0083] Based on the above embodiments, as an optional embodiment, the intention recognition module is also used to determine whether the first operation object and the second operation object are the same operation object; if the first operation object and the second operation object are not the same operation object, the first operation object or the second operation object is determined as the target operation object according to the standard operation object corresponding to the key task node; according to the operation type corresponding to the target operation object, the main operation type and the auxiliary operation type in the first operation type and the second operation type are determined; according to the main operation type, the auxiliary operation type and the target operation object, the user's operation intention information at the key task node is determined.

[0084] Based on the above embodiments, as an optional embodiment, the intention recognition module is also used to determine the first operation type as the main operation type and the second operation type as the auxiliary operation type if the target operation object is the first operation object; if the target operation object is the second operation object, the second operation type is determined as the main operation type and the first operation type is determined as the auxiliary operation type.

[0085] Based on the above embodiments, as an optional embodiment, the display module is also used to generate operation intention animation, operation intention subtitles and operation intention marks based on operation intention information; determine the target users who participate in the collaborative task and are associated with the key task nodes; and send the operation intention animation, operation intention subtitles and operation intention marks to the target user's augmented reality device, so that the target user's augmented reality device displays the operation intention animation, operation intention subtitles and operation intention marks in the augmented reality environment.

[0086] An embodiment of the present application also provides a computer storage medium, which can store multiple instructions. The instructions are suitable for being loaded by a processor and executed by a collaborative interaction method based on augmented reality as described in the above embodiment. The specific execution process can be found in the specific description of the embodiment and will not be repeated here.

[0087] See Figure 3, is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 3 As shown, the electronic device 300 may include: at least one processor 301 , at least one network interface 304 , a user interface 303 , a memory 305 , and at least one communication bus 302 .

[0088] The communication bus 302 is used to implement the connection and communication between these components.

[0089] The user interface 303 may include a standard wired interface or a wireless interface.

[0090] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).

[0091] The processor 301 may include one or more processing cores. Using various interfaces and circuits, the processor 301 connects to various components within the server. It executes instructions, programs, code sets, or instruction sets stored in the memory 305, as well as accesses data stored in the memory 305, to perform various server functions and process data. Optionally, the processor 301 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 301 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing content displayed on the display screen; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 301 but implemented as a separate chip.

[0092] Among them, the memory 305 may include a random access memory (RAM) or a read-only memory (Read-Only Memory). Optionally, the memory 305 includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 305 may also be optionally at least one storage device located away from the aforementioned processor 301. As Figure 3 As shown, the memory 305 as a computer storage medium may include an operating system, a network communication module, a user interface module, and an application program for a collaborative interaction method based on augmented reality.

[0093] exist Figure 3 In the electronic device 300 shown, the user interface 303 is mainly used to provide an input interface for the user and obtain data input by the user; and the processor 301 can be used to call an application stored in the memory 305 that stores an augmented reality-based collaborative interaction method. When executed by one or more processors, the electronic device executes one or more methods in the above-mentioned embodiments.

[0094] An electronic device readable storage medium stores instructions, which, when executed by one or more processors, enable the electronic device to execute one or more methods in the above embodiments.

[0095] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for this application.

[0096] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0097] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely schematic, such as the division of units, which is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interface, and the indirect coupling or communication connection of devices or units can be electrical or other forms.

[0098] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0099] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0100] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of this application. The aforementioned memory includes various media that can store program code, such as USB flash drives, mobile hard drives, magnetic disks, or optical disks.

[0101] The above are merely exemplary embodiments of the present disclosure and are not intended to limit the scope of the present disclosure. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure are still within the scope of the present disclosure. After considering the disclosure of the specification and the truth of practice, those skilled in the art will easily think of other embodiments of the present disclosure. This application is intended to cover any variations, uses or adaptive changes of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary technical means in the field of the present disclosure that are not recorded in the present disclosure.

Claims

1. A collaborative interaction method based on augmented reality, characterized in that: The method comprises: Acquire a collaborative task in an augmented reality environment, and identify key task nodes in the collaborative task; When the user's task progress reaches the key task node, the user's operation behavior information and voice information are obtained; Identify the operation behavior information and the voice information to obtain the user's operation intention information at the key task node; Visualization information is generated according to the operation intention information, and the visualization information is displayed in an augmented reality environment of other users.

2. The method according to claim 1, characterized in that The acquiring of the collaborative task in the augmented reality environment and identifying key task nodes in the collaborative task include: Acquire collaborative tasks in augmented reality environments; Decomposing the collaborative task to obtain multiple task nodes; A dependency graph between the task nodes is constructed, and key task nodes in the collaborative task are determined according to the topological structure of the dependency graph.

3. The method according to claim 2, characterized in that The step of constructing a dependency graph between the task nodes and determining key task nodes in the collaborative task according to the topological structure of the dependency graph includes: Obtaining node attributes of the task node; Constructing a dependency graph between the task nodes according to the task flow of the collaborative task and the node attributes of the task nodes; Traversing the dependency graph to determine a key task path in the dependency graph; The number of predecessor nodes of the task node on the critical task path is recorded, and the task node whose number of predecessor nodes is greater than a preset threshold is determined as the critical task node in the collaborative task.

4. The method according to claim 1, wherein The identifying of the operation behavior information and the voice information to obtain the operation intention information of the user at the key task node includes: Formatting the operation behavior information to obtain a plurality of first operation information in the form of event streams; Matching the first operation information with the key task node, and filtering out second operation information related to the key task node; extracting a first operation type and a first operation object according to voice keywords in the voice information; extracting a second operation type and a second operation object from the second operation information; Identify the first operation type, the first operation object, the second operation type, and the second operation object, and obtain the user's operation intention information at the key task node, where the operation intention information is structured data.

5. The method according to claim 1, wherein The identifying the first operation type, the first operation object, the second operation type, and the second operation object, and obtaining the operation intention information of the user at the key task node includes: Determining whether the first operation object and the second operation object are the same operation object; If the first operation object and the second operation object are not the same operation object, determining the first operation object or the second operation object as the target operation object according to the standard operation object corresponding to the key task node; Determining, according to the operation type corresponding to the target operation object, a main operation type and an auxiliary operation type in the first operation type and the second operation type; The operation intention information of the user at the key task node is determined according to the main operation type, the auxiliary operation type and the target operation object.

6. The method according to claim 5, characterized in that The determining, according to the operation type corresponding to the target operation object, a main operation type and an auxiliary operation type in the first operation type and the second operation type, includes: If the target operation object is the first operation object, determining the first operation type as a primary operation type and determining the second operation type as an auxiliary operation type; If the target operation object is the second operation object, the second operation type is determined as the main operation type, and the first operation type is determined as the auxiliary operation type.

7. The method according to claim 1, characterized in that The visualization information includes an operation intention animation, an operation intention subtitle, and an operation intention mark. The converting the operation intention information into visualization information and displaying the visualization information in an augmented reality environment of other users includes: generating the operation intention animation, the operation intention subtitles, and the operation intention mark according to the operation intention information; Determining target users who participate in the collaborative task and are associated with the key task nodes; The operation intention animation, the operation intention subtitles and the operation intention mark are sent to the augmented reality device of the target user, so that the augmented reality device of the target user displays the operation intention animation, the operation intention subtitles and the operation intention mark in an augmented reality environment.

8. A collaborative interactive device based on augmented reality, characterized in that: The device comprises: A node identification module is used to obtain collaborative tasks in an augmented reality environment and identify key task nodes in the collaborative tasks; The acquisition module is used to obtain the user's operation behavior information and voice information when the user's task progress reaches the key task node; An intention recognition module is used to recognize the operation behavior information and the voice information to obtain the user's operation intention information at the key task node; The display module is configured to generate visual information according to the operation intention information and display the visual information in an augmented reality environment of other users.

9. A computer storage medium, characterized in that The computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the method according to any one of claims 1 to 7.

10. An electronic device, characterized in that: The electronic device comprises a processor, a memory and a transceiver, wherein the memory is used to store instructions, the transceiver is used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Interaction method and system based on network dynamic Gantt chart

    CN106302788A

  • Task processing method and device, electronic equipment and storage medium

    CN113761127A

  • Multi-agent cooperative task execution method, device and equipment and storage medium

    CN119904069A

  • Simulating communication expressions using virtual objects

    US10467792B1

  • Augmented Reality Coordination Of Human-Robot Interaction

    US20210094180A1