An augmented reality-based collaborative interaction method, apparatus and related device

By identifying key task nodes and acquiring multimodal information in the AR environment, and transforming it into visual information display, the problem of incomplete expression of intent in AR collaboration is solved, and the efficiency of collaborative tasks is improved.

CN120469575BActive Publication Date: 2026-01-06SUZHOU I MUSEUM DIGITAL TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510564356.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2026-01-06
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

In AR collaboration, existing interaction methods cannot fully express the operator's specific intentions, resulting in low information transmission efficiency and low efficiency in completing collaborative tasks.

Method used

By identifying key task nodes of collaborative tasks in augmented reality environments, user operation behavior information and voice information are obtained. Multimodal information is used to identify user operation intentions and transform them into visual information to be displayed in the augmented reality environments of other users.

Benefits of technology

It enables accurate communication of collaborative intentions at key nodes in collaborative tasks, thereby improving the efficiency of task completion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120469575B_ABST
    Figure CN120469575B_ABST
Patent Text Reader

Abstract

The application provides an augmented reality-based cooperative interaction method and device and related equipment, and relates to the technical field of augmented reality. The technical scheme provided by the application obtains a cooperative task in an augmented reality environment and identifies key task nodes therein. When a user reaches a key task node, the operation behavior information and voice information of the user are obtained in a timely manner, the operation intention information of the user is obtained through multi-modal information recognition, the operation intention information is converted into intuitive visual information, and the visual information is displayed in the augmented reality environment of other users. Through the accurate collection, understanding and visual communication of the operation intention at the key task node, the cooperative intention is accurately communicated at the key node of the cooperative task, so that the completion efficiency of the cooperative task can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of augmented reality technology, specifically to an augmented reality-based collaborative interaction method, apparatus, and related equipment. Background Technology

[0002] Augmented Reality (AR) technology is a technique that overlays virtual information onto the real world, providing users with real-time information enhancement and interactive experiences. With the development of AR technology, multi-user AR collaborative applications are becoming increasingly widespread, such as remote repair guidance, collaborative design, and telemedicine. In AR collaboration, multiple users share the same augmented reality environment through AR devices, enabling real-time information exchange and task collaboration.

[0003] Currently, in AR collaboration, users primarily convey their intentions through voice calls and gesture annotations. However, this simple interaction method cannot fully express the operator's specific intentions and lacks a structured information transmission mechanism. Other users need to repeatedly confirm to accurately understand the operator's intentions, resulting in low information transmission efficiency during the collaboration process and consequently, low completion efficiency of collaborative tasks. Summary of the Invention

[0004] This application provides an augmented reality-based collaborative interaction method, apparatus, and related equipment that can accurately convey collaborative intentions at key nodes of a collaborative task, thereby improving the efficiency of completing the collaborative task.

[0005] In a first aspect, this application provides a collaborative interaction method based on augmented reality, the method comprising:

[0006] Acquire collaborative tasks in an augmented reality environment and identify key task nodes within those tasks;

[0007] When the user's task progress reaches the key task node, the user's operation behavior information and voice information are obtained;

[0008] The user's operational intent information at the key task node is obtained by recognizing the operational behavior information and the voice information.

[0009] Visual information is generated based on the operation intent information and displayed in the augmented reality environment of other users.

[0010] By adopting the above technical solution, collaborative tasks are acquired and key task nodes are identified in the augmented reality environment. When a user reaches a key task node, the user's operation behavior information and voice information are acquired in a timely manner, and the user's operation intention information is obtained through multimodal information recognition. At the same time, the operation intention information is transformed into intuitive visual information and displayed in the augmented reality environment of other users. By accurately collecting, understanding and visually communicating operation intentions at key task nodes, the collaborative intentions are accurately communicated at key nodes of collaborative tasks, thereby improving the efficiency of completing collaborative tasks.

[0011] Optionally, acquiring the collaborative task in the augmented reality environment and identifying key task nodes in the collaborative task includes:

[0012] Acquire collaborative tasks in augmented reality environments;

[0013] The collaborative task is decomposed into multiple task nodes;

[0014] Construct a dependency graph between the task nodes, and determine the key task nodes in the collaborative task based on the topology of the dependency graph.

[0015] By employing the above technical solution, the collaborative task is decomposed into multiple specific task nodes, and a dependency graph between these nodes is constructed. Based on the topological structure of this dependency graph, key task nodes are identified. This method of task decomposition and dependency analysis makes the structure of the collaborative task clearer, and the topological analysis of the dependency graph accurately identifies the key nodes crucial to task completion.

[0016] Optionally, constructing the dependency graph between the task nodes and determining the key task nodes in the collaborative task based on the topology of the dependency graph includes:

[0017] Obtain the node attributes of the task node;

[0018] Based on the task flow of the collaborative task and the node attributes of the task nodes, construct a dependency graph between the task nodes;

[0019] Traverse the dependency graph to determine the critical task paths in the dependency graph;

[0020] Record the number of preceding nodes of the task nodes on the critical task path, and determine the task nodes whose number of preceding nodes is greater than a preset threshold as critical task nodes in the collaborative task.

[0021] By employing the above technical solution, a dependency graph between task nodes is constructed based on the task flow of collaborative tasks. This graph is then traversed to determine critical task paths. Finally, critical task nodes are identified by comparing the number of preceding nodes with a preset threshold. This analysis method, based on node attributes and dependencies, makes the identification of critical task nodes more objective and quantifiable. Statistical analysis of the number of preceding nodes effectively identifies nodes with high dependencies in the task flow.

[0022] Optionally, the step of recognizing the operation behavior information and the voice information to obtain the user's operation intent information at the key task node includes:

[0023] The operation behavior information is formatted to obtain several first operation information in the form of event streams;

[0024] The first operation information is matched with the key task node to filter out the second operation information related to the key task node;

[0025] Based on the voice keywords in the voice information, the first operation type and the first operation object are extracted;

[0026] Extract the second operation type and the second operation object from the second operation information;

[0027] Identify the first operation type, the first operation object, the second operation type, and the second operation object to obtain the user's operation intent information at the key task node, wherein the operation intent information is structured data.

[0028] By adopting the above technical solution, the operation behavior information is converted into a standard event stream format, which enables the system to effectively filter out operation information related to the current key task node and avoid interference from irrelevant operations. At the same time, by analyzing the inputs of voice information and operation behavior information from two different dimensions, they can corroborate and complement each other, improving the accuracy and completeness of operation intent recognition.

[0029] Optionally, the step of identifying the first operation type, the first operation object, the second operation type, and the second operation object to obtain the user's operation intent information at the critical task node includes:

[0030] Determine whether the first operation object and the second operation object are the same operation object;

[0031] If the first operation object and the second operation object are not the same operation object, then the first operation object or the second operation object is determined as the target operation object according to the standard operation object corresponding to the key task node.

[0032] Based on the operation type corresponding to the target operation object, determine the first operation type and the main operation type and auxiliary operation type in the second operation type;

[0033] Based on the primary operation type, the auxiliary operation type, and the target operation object, the user's operation intent information at the critical task node is determined.

[0034] By adopting the above technical solution, it is determined whether the first operation object extracted from voice information and operation behavior information are consistent with the second operation object. When an inconsistency is found, the system makes a judgment based on the standard operation object corresponding to the key task node, thereby accurately identifying the target operation object and avoiding intention misunderstanding caused by inconsistency in multimodal information. Then, based on the identified target operation object and its corresponding operation type, the first operation type and the second operation type are distinguished into primary operation type and auxiliary operation type. This hierarchical division of operation types enables the system to highlight key operations while retaining auxiliary information. Finally, by combining the primary operation type, auxiliary operation type and target operation object, complete operation intention information is formed, realizing accurate understanding and structured expression of user intention, improving the interaction accuracy and efficiency in AR collaboration process, and reducing collaboration errors caused by intention misunderstanding deviations.

[0035] Optionally, determining the primary operation type and auxiliary operation type in the first operation type and the second operation type according to the operation type corresponding to the target operation object includes:

[0036] If the target operation object is the first operation object, then the first operation type is determined as the primary operation type, and the second operation type is determined as the auxiliary operation type;

[0037] If the target operation object is the second operation object, then the second operation type is determined as the primary operation type, and the first operation type is determined as the auxiliary operation type.

[0038] By adopting the above technical solution, the operation type associated with the target operation object is identified as the primary operation type, while the other operation type is identified as the auxiliary operation type, thus establishing a primary and secondary relationship determination mechanism for operation types. When the target operation object is the first operation object, the corresponding first operation type is identified as the primary operation type, and the second operation type is identified as the auxiliary operation type, and vice versa. This method of dividing the primary and secondary relationship of operation types based on the relevance of the target operation object ensures the logical consistency in the operation intent recognition process, while retaining the correlation features of different modal information. This enables a more accurate reconstruction of the user's true operation intent and improves the accuracy of intent understanding in the AR collaboration process.

[0039] Optionally, the visualization information includes operation intent animations, operation intent captions, and operation intent markers. Converting the operation intent information into visualization information and displaying the visualization information in the augmented reality environment of other users includes:

[0040] Generate the operation intent animation, the operation intent subtitle, and the operation intent marker based on the operation intent information;

[0041] Identify the target users who will participate in the collaborative task and are associated with the key task node;

[0042] The operation intention animation, the operation intention subtitle, and the operation intention marker are sent to the target user's augmented reality device so that the target user's augmented reality device displays the operation intention animation, the operation intention subtitle, and the operation intention marker in an augmented reality environment.

[0043] By adopting the above technical solution, the operational intent information is transformed into multi-dimensional visual information including operational intent animation, operational intent subtitles, and operational intent markers, thus achieving a three-dimensional expression of the operational intent. Simultaneously, by identifying target users associated with key task nodes, it ensures that the visual information can be accurately pushed to relevant collaborators, avoiding information overload. When this visual information is sent to the target user's augmented reality device and displayed in their AR environment, the target user can intuitively understand the operational actions through animation, obtain detailed operational instructions through subtitles, and quickly locate points of interest through markers. This multi-dimensional visualization presentation makes the transmission of operational intent clearer and more intuitive, significantly improving the efficiency and accuracy of information transmission in the AR collaboration process, and reducing communication costs and misunderstandings during collaboration.

[0044] Secondly, this application provides a collaborative interaction device based on augmented reality, the device comprising:

[0045] The node identification module is used to acquire collaborative tasks in an augmented reality environment and identify key task nodes in the collaborative tasks.

[0046] The acquisition module is used to acquire the user's operation behavior information and voice information when the user's task progress reaches the key task node;

[0047] An intent recognition module is used to recognize the operation behavior information and the voice information to obtain the user's operation intent information at the key task node;

[0048] The display module is used to generate visual information based on the operation intention information and display the visual information in the augmented reality environment of other users.

[0049] Thirdly, this application provides a computer storage medium storing a plurality of instructions adapted for loading by a processor and executing any of the methods described above.

[0050] Fourthly, this application provides an electronic device including a processor, a memory, and a transceiver, wherein the memory is used to store instructions, the transceiver is used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform any of the methods described above.

[0051] In summary, the beneficial effects of the technical solution of this application include:

[0052] By adopting the above technical solution, collaborative tasks are acquired and key task nodes are identified in the augmented reality environment. When a user reaches a key task node, the user's operation behavior information and voice information are acquired in a timely manner, and the user's operation intention information is obtained through multimodal information recognition. At the same time, the operation intention information is transformed into intuitive visual information and displayed in the augmented reality environment of other users. By accurately collecting, understanding and visually communicating operation intentions at key task nodes, the collaborative intentions are accurately communicated at key nodes of collaborative tasks, thereby improving the efficiency of completing collaborative tasks. Attached Figure Description

[0053] Figure 1 This is a flowchart illustrating a collaborative interaction method based on augmented reality according to an embodiment of this application.

[0054] Figure 2 This is a schematic diagram of the structure of an augmented reality-based collaborative interaction device according to an embodiment of this application;

[0055] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0056] Explanation of reference numerals in the attached drawings: 300, electronic device; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. Detailed Implementation

[0057] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0058] In the description of the embodiments of this application, words such as "illustrative," "for example," or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "illustrative," "for example," or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Rather, the use of words such as "illustrative," "for example," or "for example" is intended to present the relevant concepts in a specific manner.

[0059] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple devices refer to two or more devices, and multiple screen terminals refer to two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0060] Please see Figure 1 This document presents a flowchart illustrating an augmented reality-based collaborative interaction method, which can be implemented using a computer program, a microcontroller, or run on an augmented reality-based collaborative interaction device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application. The specific steps of the augmented reality-based collaborative interaction method are described in detail below.

[0061] S101: Acquire collaborative tasks in an augmented reality environment and identify key task nodes in the collaborative tasks;

[0062] Augmented reality (AR) environments refer to hybrid environments created by overlaying virtual information onto the real world using computer technology. These environments include a fusion of the real physical environment and computer-generated digital information (such as 3D models, text, images, and audio). In this embodiment, it can be understood as a hybrid visual environment perceived by a user through augmented reality devices (such as AR glasses and head-mounted displays). This environment includes real physical scenes, virtual interactive interfaces, digital information related to collaborative tasks, and visualizations of other collaborative users' operational intentions.

[0063] In this context, a collaborative task refers to a task process that requires the joint participation and completion of multiple users. During this process, each user, based on a shared task objective, collaborates through information sharing and interaction to fulfill their respective task responsibilities. In this embodiment, it can be understood as a task with a clearly defined workflow executed jointly by multiple users in an augmented reality environment. This task can be decomposed into multiple interdependent task nodes, each of which may involve the operational actions and voice interactions of one or more users.

[0064] In this context, a critical task node refers to a crucial execution point in the collaborative task process that plays a key role in task completion and requires the cooperation of multiple users. In this embodiment, it can be understood as an execution node that requires close attention in the collaborative task flow and has a significant impact on the progress of subsequent tasks.

[0065] In this embodiment, since collaborative tasks typically involve multiple execution stages, and different stages have varying degrees of impact on task completion, it is necessary to identify the key task nodes that play a crucial role in task completion. Specifically, the collaborative task in the augmented reality environment is first acquired. Then, the acquired collaborative task is analyzed and decomposed, breaking down the overall task into multiple specific execution stages. By analyzing the importance and scope of influence of these execution stages within the task, the key task nodes are identified.

[0066] Based on the above embodiments, as an optional implementation method, the identification of key task nodes can be achieved using steps S201-S203.

[0067] S201: Obtain collaborative tasks in an augmented reality environment;

[0068] Collaborative tasks refer to tasks that require multiple users to complete together using augmented reality devices. In practice, the augmented reality device receives task information input from users or retrieves pre-set collaborative tasks from a task management system. The retrieved collaborative task information includes basic details such as task type, task description, and participating users.

[0069] S202: Decompose the collaborative task to obtain multiple task nodes;

[0070] In this embodiment, since collaborative tasks are typically complex, directly executing the entire task is difficult. Therefore, it is necessary to decompose the task into smaller, more easily executed task nodes. In specific implementation, based on the task's execution flow and logical relationships, the overall task is divided into multiple relatively independent execution units, each constituting a task node. Each task node contains basic attribute information of that execution unit, such as execution order and completion conditions.

[0071] S203: Construct a dependency graph between task nodes and determine the key task nodes in the collaborative task based on the topology of the dependency graph.

[0072] In this context, the dependency graph between task nodes refers to a data structure that describes the execution order and interrelationships among task nodes in a collaborative task. In this embodiment, it can be understood as a graph structure composed of nodes and directed edges, where nodes represent task execution units, directed edges represent the execution order between task nodes, and each node contains node attribute information to characterize its features within the entire task flow. The dependency graph is used to analyze the node relationships in the task flow, determine the critical path of task execution, and thus identify task nodes that play a crucial role in task completion.

[0073] After obtaining the task nodes, a dependency graph needs to be constructed to identify the key task nodes in the collaborative task based on its topology. Since there are execution orders and interdependencies among the task nodes, a dependency graph is needed to express these relationships, thereby identifying nodes that significantly impact task completion. In practice, a dependency graph is constructed based on the task flow of the collaborative task. This graph structure uses nodes to represent task execution units and directed edges to represent the execution order between nodes. Based on the constructed dependency graph, the key task nodes in the collaborative task can be determined by analyzing its topology.

[0074] Specifically, step S203 also includes S301-S304.

[0075] S301: Get the node attributes of the task node;

[0076] Node attributes refer to a set of data information describing the characteristics of a task node. In specific implementations, basic characteristic parameters such as the execution order, completion conditions, and time limits of each task node are obtained; this data information constitutes the node attributes. By obtaining node attributes, the characteristics of each task node in the entire task flow can be accurately characterized.

[0077] S302: Based on the task flow of the collaborative task and the node attributes of the task nodes, construct a dependency graph between task nodes;

[0078] In practical implementation, because collaborative tasks have specific execution flows and each task node has unique attributes, it is necessary to comprehensively consider this information to construct an accurate dependency graph. In the actual implementation, the overall execution flow of the collaborative task is first analyzed to clarify the execution order requirements of each stage. Then, the task nodes are arranged as nodes in the graph. Next, the execution order between nodes is determined according to the task flow. For nodes with sequential execution relationships, directed edges are used to connect them, and the specific association methods between nodes are determined based on node attributes. In the final dependency graph, each node represents a task execution unit, and the directed edges between nodes represent the execution order and dependency relationships.

[0079] S303: Traverse the dependency graph and determine the critical task paths in the dependency graph;

[0080] Since the dependency graph contains multiple possible paths from the start node to the end node, some of which have a decisive impact on the completion of the entire collaborative task, it is necessary to identify these critical task paths. In the implementation, starting from the start node of the dependency graph, a depth-first search is used to traverse the entire graph structure, recording each complete path from the start node to the end node. During the traversal, the strength of the dependencies between nodes on the path and the importance of each node in the overall task flow are considered to evaluate the impact of each path on task completion. By comparing the impact of different paths, the paths that play a crucial role in task completion, i.e., the critical task paths, are determined.

[0081] S304: Record the number of preceding nodes of the task nodes on the critical task path, and identify the task nodes with a number of preceding nodes greater than a preset threshold as critical task nodes in the collaborative task.

[0082] Here, the number of preceding nodes refers to the number of other nodes in the dependency graph that point to a certain task node. In this embodiment, it can be understood as the total number of all task nodes that need to be completed before any given task node begins execution.

[0083] Since the number of preceding nodes reflects the dependency complexity of a task node, a larger number indicates that the node is subject to more constraints, and thus has a greater impact on the successful completion of the entire collaborative task. In practice, we first traverse each task node on the critical task path, counting the number of other nodes pointing to that node, i.e., calculating the number of preceding nodes for each node. Then, we compare the number of preceding nodes for each node with a pre-set threshold. When the number of preceding nodes for a node exceeds the threshold, it indicates that the node has high dependency complexity and requires close attention and management; therefore, it is identified as a critical task node in the collaborative task. Critical task nodes identified in this way not only occupy the critical path of task execution but also have high dependency complexity; the execution status of these nodes directly affects the progress of the entire collaborative task.

[0084] S102: When the user's task progress reaches a critical task node, obtain the user's operation behavior information and voice information;

[0085] Operational behavior information refers to the specific behavioral data records generated by task participants during the execution of task nodes. In this embodiment, it can be understood as key behavioral characteristic information such as operation type, operation time, and operation result during the execution of key task nodes. Operational behavior information is used to record and reflect the actual situation of task execution.

[0086] Here, voice information refers to the oral communication content generated by task participants during task execution. In the embodiments of this application, it can be understood as: voice interaction data such as voice dialogues, instruction transmissions, and opinion exchanges between participants during task execution.

[0087] Because key task nodes have a significant impact on the completion of the entire collaborative task, it is necessary to comprehensively record the user's performance at these nodes. In practice, the system monitors the user's task progress in real time. When a user reaches a key task node, it begins collecting the user's operational behavior information, including key behavioral characteristics such as operation type, operation time, and operation result. Simultaneously, it activates the voice acquisition function to record the user's voice commands and voice interactions with other participants in the AR environment. By simultaneously collecting operational behavior information and voice information, the system can comprehensively record the user's execution process at key task nodes, including both specific operational actions and voice interactions during the decision-making process.

[0088] S103: Recognize the operation behavior information and voice information to obtain the user's operation intention information at key task nodes;

[0089] Operational intent information refers to inferring the user's desired goal or next action by analyzing the user's executed actions and voice information. This intent information reflects the user's behavioral purpose, rather than the behavioral process itself. In this embodiment, operational intent information can be understood as identifying the user's true goal by analyzing the combination of actual operational actions and voice information at key task nodes.

[0090] In this embodiment, the purpose of recognizing operation behavior information and voice information is to improve the accuracy of understanding user intent by combining and analyzing these two types of information. In practical applications, when user operation behavior information and voice information are acquired simultaneously, both types of information contain the user's current task intent. By recognizing and analyzing these two types of information, the operation type and operation object contained in the operation behavior information and voice information are extracted. Then, the extracted information is correlated with key task nodes to obtain the user's operation intent information at that key task node. This recognition method can not only accurately understand the user's current operation intent, but also effectively avoid recognition bias that may be caused by a single information source.

[0091] Based on the above embodiments, as an optional implementation method, step S103 further includes steps S401-S405.

[0092] S401: Format the operation behavior information to obtain several first operation information in the form of an event stream;

[0093] The first operation information refers to the formatted operation behavior data structure, which contains information about the independent operation units performed by the user at a specific point in time. In this embodiment, it can be understood as event stream data obtained by decomposing the user's continuous operation behavior in chronological order. Each event stream contains attribute information such as the timestamp of the operation, the operation type, and the operation object.

[0094] In this embodiment, the operation behavior information is formatted to obtain several event streams of first operation information. This is mainly to transform the user's continuous operation behavior into a standardized data format, facilitating subsequent processing and analysis. In practical applications, user operation behavior information is often a continuous and complex combination of actions. By formatting this operation behavior information, it can be broken down into multiple independent event streams. The formatting process includes performing temporal analysis and action decomposition on the operation behavior information, dividing the continuous operation behavior into independent operation units according to time sequence. Each operation unit contains attribute information such as operation time, operation type, and operation object. These independent operation units constitute the first operation information in the form of event streams.

[0095] S402: Match the first operation information with the key task node and filter out the second operation information related to the key task node;

[0096] The second operation information refers to the set of operation information after matching and filtering by the key task node. It retains all operation behavior data related to the current task node. In this embodiment, it can be understood as the operation information related to the current key task node that is retained after matching the operation type and operation object in the first operation information with the preset standard operation type and standard operation object of the key task node.

[0097] In this embodiment, matching the first operation information with the key task node aims to filter out operation information relevant to the current task and avoid interference from irrelevant operation information. In practical applications, users may perform some operations unrelated to the current task, and these operations will also be converted into the first operation information. By matching the operation type and operation object in the first operation information with the preset standard operation type and standard operation object of the key task node, the successfully matched first operation information can be filtered out as the second operation information. This matching and filtering method can effectively filter out operation information unrelated to the current task, and the obtained second operation information more accurately reflects the user's relevant operation behavior at the key task node.

[0098] S403: Extract the first operation type and the first operation object based on the voice keywords in the voice information;

[0099] In this embodiment, a first operation type and a first operation object are extracted based on voice keywords in the voice information. The purpose is to understand the type of operation the user wants to perform and the target object of the operation by analyzing the user's voice commands. In practical applications, the user's voice information typically contains verb phrases representing the operation behavior and noun phrases representing the operation target. By performing semantic analysis on the voice information, the verb phrases are identified as the first operation type, and the noun phrases are identified as the first operation object. This extraction method can obtain the key elements of the operation intention from the user's voice description, enabling the voice information to be transformed into structured data that can be used for intention recognition.

[0100] S404: Extract the second operation type and the second operation object from the second operation information;

[0101] In this embodiment, the extraction of the second operation type and the second operation object from the second operation information aims to obtain key elements of the user's operational intent from their actual operational behavior. In practical applications, the second operation information has already been filtered by key task nodes and contains operational behavior data related to the current task. By analyzing the structural characteristics of this operational behavior data, the action features can be identified as the second operation type, and the target of the action can be identified as the second operation object. This extraction method can obtain key information about the user's operational intent from their actual operational behavior, enabling the operational behavior data to be transformed into structured information that can be used for intent recognition.

[0102] S405: Identify the first operation type, the first operation object, the second operation type, and the second operation object to obtain the user's operation intent information at the key task node. The operation intent information is structured data.

[0103] In this embodiment, user intent information at key task nodes is obtained by identifying a first operation type, a first operation object, a second operation type, and a second operation object. The aim is to integrate and analyze operation information at both the voice and behavioral levels to more accurately understand the user's true intent. In practical applications, the first operation type and the first operation object reflect the user's intent expressed through voice, while the second operation type and the second operation object reflect the user's intent expressed through actual behavior. Comprehensive identification of this information can construct structured data containing operation types and operation objects. This structured data format not only fully describes the user's operation intent but also facilitates data transmission and processing in collaborative environments.

[0104] Based on the above embodiments, as an optional implementation method, step S405 further includes steps S501-505.

[0105] S501: Determine whether the first operand and the second operand are the same operand.

[0106] In this embodiment, determining whether the first and second operation objects are the same object aims to verify whether the operation object expressed by the user through voice matches the operation object in the actual operation, thereby ensuring the accuracy of the identified operation intent. In practical applications, by comparing the feature information of the first and second operation objects, such as object type and object attributes, it is determined whether the two operation objects point to the same entity. When the two operation objects are consistent, it indicates that the user's voice expression and actual operation both point to the same target, and this consistency enhances the credibility of the operation intent recognition. When the two operation objects are inconsistent, it indicates that there is a difference between the user's voice description and the actual operation, and this inconsistency needs to be further analyzed and processed to ensure that the finally identified operation intent can accurately reflect the user's true intent.

[0107] S502: If the first operation object and the second operation object are not the same operation object, then the first operation object or the second operation object shall be determined as the target operation object according to the standard operation object corresponding to the key task node.

[0108] In this embodiment, when the first operation object and the second operation object are inconsistent, the target operation object needs to be determined through the standard operation object corresponding to the key task node. This is because in real-world application scenarios, the user's voice expression and actual operation behavior may be inconsistent. By comparing the matching degree of the first and second operation objects with the preset standard operation objects of the key task node, and selecting the operation object with the higher matching degree as the target operation object, this method can provide a reliable judgment basis when the operation objects are inconsistent. Specifically, the feature similarity between the first and second operation objects and the standard operation object can be calculated, such as object type similarity, attribute feature similarity, etc., and the operation object with higher similarity can be determined as the target operation object.

[0109] S503: Based on the operation type corresponding to the target operation object, determine the primary operation type and auxiliary operation type in the first operation type and the second operation type.

[0110] In this embodiment, primary and secondary operation types are determined based on the operation type corresponding to the target operation object. The aim is to perform hierarchical analysis of user behavior and identify core and supporting operations. In practical applications, the target operation object typically has a set of associated standard operation types. By matching and analyzing the first and second operation types with this set of standard operation types, it can be determined which operation types are primary operations for the target operation object and which are secondary operations. Specifically, the set of standard operation types corresponding to the target operation object is first obtained. Then, the first and second operation types are compared with the operation types in this set to calculate their similarity. Operation types with high similarity that directly affect the target operation object are identified as primary operation types, while operation types with low similarity or those that play a supporting role are identified as secondary operation types.

[0111] Based on the above embodiments, as an optional implementation, if the target operation object is a first operation object, then the first operation type is determined as the main operation type, and the second operation type is determined as the auxiliary operation type.

[0112] If the target operation object is the second operation object, then the second operation type is determined as the primary operation type, and the first operation type is determined as the auxiliary operation type.

[0113] In this embodiment, the primary and secondary operation types are determined based on whether the target operation object originates from voice expression or actual operation behavior. This determination method categorizes operation types based on the source of the operation object. When the target operation object is the first operation object, it indicates that the user's voice expression more accurately points to the task target; therefore, the first operation type is determined as the primary operation type, and the second operation type is determined as the secondary operation type. When the target operation object is the second operation object, it indicates that the user's actual operation behavior better meets the task requirements; therefore, the second operation type is determined as the primary operation type, and the first operation type is determined as the secondary operation type. In practical applications, this categorization method can automatically adjust the primary and secondary relationships of operation types according to the source of the operation object, ensuring that the ultimately recognized operation intent more closely matches the user's true intent.

[0114] S504: Determine the user's operational intent information at key task nodes based on the primary operation type, auxiliary operation type, and target operation object.

[0115] In this embodiment, the user's operational intent information at key task nodes is determined based on the primary operation type, auxiliary operation type, and target operation object, with the aim of constructing a complete and structured description of the operational intent. In practical applications, the primary operation type is used as the core action, the auxiliary operation type as the cooperating action, and the target operation object as the operation target. These elements are combined according to a preset intent template to form structured data describing the user's complete operational intent. Specifically, a template structure of "primary operation type + auxiliary operation type + target operation object" can be adopted, where the primary operation type determines the main purpose of the operation, the auxiliary operation type provides auxiliary information for the operation, and the target operation object clarifies the objective of the operation.

[0116] S505: If the first operation object and the second operation object are the same operation object, then determine the user's operation intention information at the critical task node based on the first operation type, the second operation type, and the same operation object.

[0117] In this embodiment, when the first operation object and the second operation object are the same operation object, the user's operation intent information is directly determined based on the first operation type, the second operation type, and this same operation object. This is because the consistency of the operation object indicates a good correspondence between the user's voice expression and actual operation behavior. In practical applications, combining the first operation type and the second operation type can provide a more comprehensive understanding of the user's operation behavior. The first operation type reflects the user's operation intent expressed through voice, while the second operation type reflects the user's operation intent expressed through actual behavior. Both work together on the same operation object to form a complete description of the operation intent. In specific implementation, the first operation type and the second operation type can be fused and analyzed, combined with the feature information of the same operation object, to construct structured data containing operation type combinations and operation objects. This structured data can accurately describe the user's operation intent at the current task node.

[0118] S104: Generate visual information based on the operation intention information and display the visual information in the augmented reality environment of other users.

[0119] Visualized information refers to a set of information that intuitively displays operational intentions through visual elements and graphic symbols, transforming abstract operational intentions into visible graphical representations. In this embodiment, it can be understood as a combination of visual effects consisting of primary operation type symbols, auxiliary operation type symbols, and target operation object symbols. These symbols may include specific visual elements such as operation direction arrows, dashed operation path lines, and highlighted boxes for target objects.

[0120] In this embodiment, visualization information is generated based on the user's operational intent and displayed in the augmented reality environment of other users. The aim is to transform the user's operational intent into an intuitive and visible visual effect, facilitating understanding and comprehension of the current user's actions by other users in the collaborative environment. In practical applications, the system first parses the structured operational intent information into visual elements, including converting primary operation types into primary visual identifiers, secondary operation types into secondary visual identifiers, and the location and features of the target operation object into spatial positioning information. These visualizations are then overlaid and displayed on the augmented reality devices of other users. Specifically, different visual styles can be used to distinguish between primary and secondary operations. For example, prominent arrows can be used to indicate the primary operation direction, while dashed lines or semi-transparent effects can be used to represent secondary operation paths. Highlights or interactive markers can also be added around the target operation object. This visualization method helps other users quickly understand the current user's operational intent.

[0121] Based on the above embodiments, as an optional implementation, operation intention animation, operation intention subtitles, and operation intention markers are generated according to operation intention information;

[0122] Identify the target users who will participate in collaborative tasks and are associated with key task nodes;

[0123] The operation intent animation, operation intent caption, and operation intent marker are sent to the target user's augmented reality device so that the target user's augmented reality device can display the operation intent animation, operation intent caption, and operation intent marker in the augmented reality environment.

[0124] In this embodiment, operation intent information is transformed into various visualization forms such as operation intent animation, operation intent subtitles, and operation intent markers. The aim is to enhance the effect of collaborative interaction through multi-dimensional information display. In practical applications, firstly, dynamic operation intent animations are generated based on the operation intent information to demonstrate the operation process. Simultaneously, text-based operation intent subtitles are generated to provide semantic descriptions, and spatially positioned operation intent markers are generated to indicate the specific operation location. Next, target users related to the current task node are identified. These users are collaborators participating in the collaborative task and associated with the current key task node. Subsequently, this visualization information is sent to the target users' augmented reality devices, enabling them to see the complete expression of operation intent in their respective augmented reality environments. Notably, when different users arrive at the same key task node successively, the operation intent information of other users is automatically identified and displayed in their respective augmented reality environments, achieving information sharing and interactive synchronization at the task node level. Specifically, operation intent animations can demonstrate the dynamic process of the operation, operation intent subtitles can provide textual explanations, and operation intent markers can indicate spatial locations. These three visualization forms complement each other, jointly constructing a comprehensive operation intent display system.

[0125] This multi-dimensional visualization method not only helps target users understand the operational intentions of other users from different perspectives, but also ensures the accuracy and completeness of information transmission during collaboration, significantly improving the interactive effects and collaboration efficiency in augmented reality collaborative environments.

[0126] The following are embodiments of the apparatus of this application, which can be used to execute the embodiments of the method of this application. For details not disclosed in the embodiments of the apparatus of this application, please refer to the embodiments of the method of this application.

[0127] Please see Figure 2 This illustration shows a schematic diagram of an augmented reality-based collaborative interaction device provided in an exemplary embodiment of this application. The device can be implemented entirely or partially through software, hardware, or a combination of both. The augmented reality-based collaborative interaction device includes:

[0128] The node recognition module is used to acquire collaborative tasks in an augmented reality environment and identify key task nodes in the collaborative tasks.

[0129] The data acquisition module is used to acquire user operation behavior information and voice information when the user's task progress reaches a critical task node;

[0130] The intent recognition module is used to recognize operation behavior information and voice information to obtain the user's operation intent information at key task nodes;

[0131] The display module is used to generate visual information based on the user's intention and display the visual information in the augmented reality environment of other users.

[0132] Based on the above embodiments, as an optional embodiment, the node identification module is also used to obtain collaborative tasks in an augmented reality environment; decompose the collaborative tasks to obtain multiple task nodes; construct a dependency graph between task nodes, and determine the key task nodes in the collaborative tasks according to the topology of the dependency graph.

[0133] Based on the above embodiments, as an optional embodiment, the node identification module is further used to obtain the node attributes of the task nodes; construct a dependency graph between task nodes according to the task flow of the collaborative task and the node attributes of the task nodes; traverse the dependency graph to determine the critical task path in the dependency graph; record the number of predecessor nodes of the task nodes on the critical task path, and determine the task nodes with a number of predecessor nodes greater than a preset threshold as critical task nodes in the collaborative task.

[0134] Based on the above embodiments, as an optional embodiment, the intent recognition module is further configured to format the operation behavior information to obtain a number of first operation information in the form of event streams; match the first operation information with key task nodes to filter out second operation information related to the key task nodes; extract the first operation type and the first operation object based on the voice keywords in the voice information; extract the second operation type and the second operation object from the second operation information; identify the first operation type, the first operation object, the second operation type, and the second operation object to obtain the user's operation intent information at the key task node, wherein the operation intent information is structured data.

[0135] Based on the above embodiments, as an optional embodiment, the intent recognition module is further used to determine whether the first operation object and the second operation object are the same operation object; if the first operation object and the second operation object are not the same operation object, then the first operation object or the second operation object is determined as the target operation object according to the standard operation object corresponding to the key task node; according to the operation type corresponding to the target operation object, the primary operation type and the auxiliary operation type in the first operation type and the second operation type are determined; according to the primary operation type, the auxiliary operation type and the target operation object, the user's operation intent information at the key task node is determined.

[0136] Based on the above embodiments, as an optional embodiment, the intent recognition module is further configured to determine the first operation type as the primary operation type and the second operation type as the auxiliary operation type if the target operation object is the first operation object; and to determine the second operation type as the primary operation type and the first operation type as the auxiliary operation type if the target operation object is the second operation object.

[0137] Based on the above embodiments, as an optional embodiment, the display module is further configured to generate an operation intention animation, an operation intention subtitle, and an operation intention marker according to the operation intention information; determine the target user who participates in the collaborative task and is associated with the key task node; and send the operation intention animation, the operation intention subtitle, and the operation intention marker to the target user's augmented reality device, so that the target user's augmented reality device displays the operation intention animation, the operation intention subtitle, and the operation intention marker in the augmented reality environment.

[0138] This application also provides a computer storage medium that can store multiple instructions. The instructions are adapted to be loaded and executed by a processor using the augmented reality-based collaborative interaction method as described above. For details of the execution process, please refer to the specific description of the embodiments, which will not be repeated here.

[0139] Please see Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 3 As shown, the electronic device 300 may include: at least one processor 301, at least one network interface 304, user interface 303, memory 305, and at least one communication bus 302.

[0140] The communication bus 302 is used to enable communication between these components.

[0141] The user interface 303 may include a standard wired interface and a wireless interface.

[0142] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0143] The processor 301 may include one or more processing cores. The processor 301 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 305, and by calling data stored in memory 305. Optionally, the processor 301 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 301 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 301 and may be implemented as a separate chip.

[0144] The memory 305 may include random access memory (RAM) or read-only memory. Optionally, the memory 305 may include a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 305 may also be at least one storage device located remotely from the aforementioned processor 301. Figure 3 As shown, the memory 305, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application based on an augmented reality-based collaborative interaction method.

[0145] exist Figure 3In the electronic device 300 shown, the user interface 303 is mainly used to provide an input interface for the user and to obtain the user input data; while the processor 301 can be used to call an application program stored in the memory 305 that is an augmented reality-based collaborative interaction method. When executed by one or more processors, the electronic device executes one or more methods as described in the above embodiments.

[0146] An electronic device readable storage medium stores instructions that, when executed by one or more processors, cause the electronic device to perform one or more methods as described in the above embodiments.

[0147] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0148] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0149] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the shown or discussed mutual couplings or direct couplings or communication connections may be through some service interfaces; indirect couplings or communication connections between apparatuses or units may be electrical or other forms.

[0150] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0151] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0152] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, portable hard drives, magnetic disks, or optical disks.

[0153] The above are merely exemplary embodiments of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of other embodiments of this disclosure upon considering the specification and the disclosure of practical truths. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure.

Claims

1. A method of augmented reality based collaborative interaction, characterized in that, The method comprises: obtaining a collaborative task in an augmented reality environment, identifying a key task node in the collaborative task; when the task progress of a user reaches the key task node, obtaining operation behavior information and voice information of the user; identifying the operation behavior information and the voice information to obtain operation intention information of the user at the key task node; wherein the operation behavior information and the voice information are identified to obtain the operation intention information of the user at the key task node, comprising: formatting the operation behavior information to obtain first operation information in the form of event stream; matching the first operation information with the key task node to filter out second operation information related to the key task node; extracting a first operation type and a first operation object according to a voice keyword in the voice information; extracting a second operation type and a second operation object in the second operation information; identifying the first operation type, the first operation object, the second operation type and the second operation object to obtain the operation intention information of the user at the key task node, the operation intention information being structured data; wherein the identification of the first operation type, the first operation object, the second operation type and the second operation object to obtain the operation intention information of the user at the key task node comprises: determining whether the first operation object and the second operation object are the same operation object; if the first operation object and the second operation object are not the same operation object, determining the first operation object or the second operation object as a target operation object according to a standard operation object corresponding to the key task node; determining a main operation type and an auxiliary operation type in the first operation type and the second operation type according to an operation type corresponding to the target operation object; determining the operation intention information of the user at the key task node according to the main operation type, the auxiliary operation type and the target operation object; generating visual information according to the operation intention information, and displaying the visual information in the augmented reality environment of other users.

2. The method of claim 1, wherein, The obtaining of the collaborative task in the augmented reality environment and the identification of the key task node in the collaborative task comprise: obtaining a collaborative task in an augmented reality environment; decomposing the collaborative task to obtain a plurality of task nodes; constructing a dependency graph between the task nodes, and determining a key task node in the collaborative task according to a topological structure of the dependency graph.

3. The method of claim 2, wherein, The construction of the dependency graph between the task nodes and the determination of the key task node in the collaborative task according to the topological structure of the dependency graph comprise: obtaining node attributes of the task nodes; constructing a dependency graph between the task nodes according to a task flow of the collaborative task and the node attributes of the task nodes; traversing the dependency graph to determine a key task path in the dependency graph; Record the number of preceding nodes of the task nodes on the critical task path, and determine the task nodes with the number of preceding nodes greater than a preset threshold as critical task nodes in the collaborative task.

4. The method of claim 1, wherein, The determining the main operation type and the auxiliary operation type from the first operation type and the second operation type according to the operation type corresponding to the target operation object comprises: If the target operation object is the first operation object, determining the first operation type as the main operation type and the second operation type as the auxiliary operation type; If the target operation object is the second operation object, determining the second operation type as the main operation type and the first operation type as the auxiliary operation type.

5. The method of claim 1, wherein, The visual information comprises operation intention animation, operation intention subtitle and operation intention mark, and the generating visual information according to the operation intention information and displaying the visual information in the augmented reality environment of other users comprises: Generating the operation intention animation, the operation intention subtitle and the operation intention mark according to the operation intention information; Determining a target user participating in the collaborative task and associated with the critical task node; Sending the operation intention animation, the operation intention subtitle and the operation intention mark to the augmented reality device of the target user, so that the augmented reality device of the target user displays the operation intention animation, the operation intention subtitle and the operation intention mark in the augmented reality environment.

6. An augmented reality based collaborative interaction device, characterized by, The device comprises: A node identification module configured to acquire a collaborative task in an augmented reality environment and identify a critical task node in the collaborative task; A collection module configured to acquire operation behavior information and voice information of a user when a task progress of the user reaches the critical task node. An intention recognition module is configured to recognize the operation behavior information and the voice information to obtain operation intention information of the user at the key task node. The operation behavior information is formatted to obtain first operation information in the form of event flow. The first operation information is matched with the key task node to filter out second operation information related to the key task node. A first operation type and a first operation object are extracted from the voice keywords in the voice information. A second operation type and a second operation object are extracted from the second operation information. The first operation type, the first operation object, the second operation type, and the second operation object are recognized to obtain the operation intention information of the user at the key task node. The operation intention information is structured data. The first operation object and the second operation object are determined to be the same operation object. If the first operation object and the second operation object are not the same operation object, the first operation object or the second operation object is determined to be a target operation object according to a standard operation object corresponding to the key task node. A main operation type and an auxiliary operation type are determined from the first operation type and the second operation type according to the target operation object. The operation intention information of the user at the key task node is determined according to the main operation type, the auxiliary operation type, and the target operation object. A display module is configured to generate visual information according to the operation intention information and display the visual information in an augmented reality environment of other users.

7. A computer storage medium, characterized in that The computer storage medium stores a plurality of instructions adapted to be loaded and executed by a processor to perform the method of any one of claims 1-5.

8. An electronic device, comprising: An electronic device includes a processor, a memory, and a transceiver. The memory is configured to store instructions. The transceiver is configured to communicate with other devices. The processor is configured to execute the instructions stored in the memory to cause the electronic device to perform the method of any one of claims 1-5.

Citation Information

Patent Citations

  • Interaction method and system based on network dynamic Gantt chart

    CN106302788A

  • Task processing method and device, electronic equipment and storage medium

    CN113761127A

  • Multi-agent cooperative task execution method, device and equipment and storage medium

    CN119904069A