Expert-based guidance through virtual avatars in augmented reality and virtual reality environments
By capturing the motion data of individual experts and utilizing the semantic action database and work context knowledge graph, virtual avatars are generated to provide real-time and effective expert guidance to AR and VR users, solving the problem of lack of expert guidance in existing technologies and improving the accuracy and efficiency of user task execution.
Patent Information
- Application Number
- CN202380094816.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-23
- Publication Date
- 2025-10-03
AI Technical Summary
Existing AR and VR technologies lack real-time, effective expert guidance when providing user task execution assistance. In particular, AI-based virtual assistants are unable to provide practical guidance from domain experts, and conventional training models have limited effectiveness in different environments.
The learning engine captures the motion data of individual experts, utilizes the semantic action database and work context knowledge graph, and generates virtual avatars to provide users with precise expert-based guidance, including verbal guidance and virtual demonstrations, adapting to various environments and task complexities.
It provides real-time and effective expert guidance in AR and VR environments, improves the accuracy and efficiency of user task execution, and adapts to the needs of different environments and task complexities.
Smart Images

Figure CN120752601A_ABST
Abstract
Description
Background Art
[0001] Computer systems can be used to create, use, and manage data for nearly any type of process or purpose. Virtual reality (VR) and augmented reality (AR) technologies allow users to access and use data in increasingly sophisticated ways within increasingly digital environments. AR and VR users can benefit from the enhanced capabilities and resources within AR and VR environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0002] Certain examples are described in the following detailed description and with reference to the accompanying drawings.
[0003] Figure 1 An example of a system that supports expert-based guidance through virtual avatars in AR and VR environments is shown.
[0004] Figure 2 An example of expert knowledge capture supporting expert-based guidance according to the present disclosure is shown.
[0005] Figure 3 Another example of expert knowledge capture supporting expert-based guidance according to the present disclosure is shown.
[0006] Figure 4 An example of providing expert-based guidance to individuals performing tasks in an environment according to the present disclosure is shown.
[0007] Figure 5 Examples of logic that the system may implement to support expert-based guidance in AR and VR environments are shown.
[0008] Figure 6 An example of a computing system that supports expert-based guidance in AR and VR environments is shown. DETAILED DESCRIPTION
[0009] With the advancement of modern technology, the feasibility and adoption of AR and VR technologies continue to increase. By overlaying digital data on a physical environment (e.g., through AR devices), AR technology provides users with enhanced availability of data collection, analysis, and display capabilities superimposed on the real world (physical context). VR technology can support virtual gatherings in a common virtual site to work together, allowing training, problem solving, and deeper collaboration between users dispersed in vastly different geographic locations, time zones, and physical contexts. Virtual worlds are being created and populated, allowing users to virtually gather in almost any type of context to train, learn, collaborate, and perform complex tasks in virtual gatherings.
[0010] As the capabilities of AR and VR technologies increase, it becomes increasingly feasible to provide assistance to users in performing tasks. Such guidance may be particularly relevant to assisting users in performing complex tasks in different environments. Virtual environments may be particularly suitable for performing complex industrial tasks, for example allowing users to first virtually train to operate industrial machinery or perform complex tasks in a virtual background before attempting to perform such tasks in a physical environment. Conventional forms of user assistance for performing complex tasks may be in the form of training videos (e.g., recorded demonstrations of performing the task or through instructional videos and training slides). However, this mode of training provides little feedback or real-time guidance to the individual performing the task, and is often performed in a different background or under different environmental conditions than the recorded video.
[0011] Digital assistants provide another form of assistance to users in performing tasks. Some forms of digital assistants may incorporate artificial intelligence (AI) learning techniques to provide feedback to users based on predictions of user interactions. Continued research into AI-based chatbots, virtual assistants, and AI avatars may result in improved user interactions with AI-trained virtual beings in virtual contexts. However, AI-based training may require large amounts of training data to operate effectively and, at best, may only provide learned predictions for user assistance rather than actual guidance (e.g., demonstrations) from experts in a given field or trained to perform specific tasks.
[0012] The disclosure herein can provide systems, methods, devices, and logic for expert-based guidance in AR and VR environments. At a high level, the expert-based guidance technology disclosed herein can provide the ability to capture the knowledge and actions of an expert and transfer them to another individual to perform a specific task. As used herein, an expert can refer to any individual with a threshold level of experience, knowledge, or expertise to perform a task. Therefore, capturing the real-world experience of an expert and transferring it to a less experienced user can provide directly relevant guidance to the individual performing the task, whether in an AR or VR context. As described herein, expert-based guidance can be provided via an avatar, which can refer to any digital or virtual representation of a person, entity, logic, agent, or organism. An avatar can be controlled, rendered, and driven by the expert-based guidance technology disclosed herein and can therefore represent the expert-based guidance technology disclosed herein (as opposed to an avatar representing a human expert). In other words, the avatars described herein can represent digital assistant agents generated and controlled by the expert-based guidance technology disclosed herein. The avatars disclosed herein (including their underlying expert-based guidance technology) can be easily replicated and readily available in all types of contexts and environments, providing support to users. Thus, the replicable avatars of the present disclosure can provide expert support without spatial or temporal limitations that would constrain the availability of human experts located in fixed geographic locations with limited temporal availability.
[0013] Compared to AI-based virtual assistant technologies that attempt to guess user interactions and predict relevant feedback, the expert-based guidance technology of the present disclosure can semantically classify user motions, actions, environmental conditions, and any other factors related to task execution in order to accurately interpret user actions and generate guidance accordingly. Along a similar direction, the present disclosure considers capturing and classifying the precise motions and actions of experts in performing tasks, thereby allowing direct comparison between target actions (e.g., actions captured for experts) and actual actions performed by users in AR or VR environments. In addition, the actions performed by experts and users can be enhanced with the working context of the user and expert actions, allowing for more comprehensive comparisons to provide users with more relevant and effective expert-based guidance. The working context can be captured through a knowledge graph, which can support the dissemination of relevant guidance even if there are deviations in the working context and environmental conditions in the user environment.
[0014] The expert-based guidance technology disclosed herein can support virtual 3D avatars that can provide relevant expert-based guidance to any individual performing any task of any type or complexity. The expert-based guidance provided by the present disclosure can take a variety of forms, from verbal guidance (e.g., via a natural language interface) to demonstrations by virtual avatars of steps in performing complex tasks, and so on. These and other expert-based guidance features and technical advantages will be described in more detail herein.
[0015] Figure 1 An example of a computing system 100 that supports expert-based guidance in AR and VR environments is shown. The computing system 100 can take the form of a single or multiple computing devices, such as an application server, a computing node, a desktop or laptop computer, a smartphone or other mobile device, a tablet device, an embedded controller, and any related or applicable technology devices. In some implementations, the computing system 100 hosts, supports, executes, or implements a digital assistant system that can implement any of the various features described herein, including building and using a 3D digital assistant as an avatar in VR and AR environments that can provide expert-based guidance according to the present disclosure.
[0016] As an example implementation supporting any combination of the expert-based guidance features described herein, Figure 1 The illustrated computing system 100 includes a learning engine 110 and an expert avatar engine 112. The computing system 100 can implement the engines 110 and 112 (including their components) in various ways, such as hardware and programming. The programming of the engines 110 and 112 can take the form of processor-executable instructions stored on a non-transitory machine-readable storage medium, and the hardware of the engines 110 and 112 can include a processor for executing these instructions. The processor can take the form of a single processor or a multi-processor system, and in some examples, the computing system 100 implements multiple engines using the same computing system features or hardware components (e.g., a general-purpose processor or a general-purpose storage medium).
[0017] In operation, the learning engine 110 can capture the expertise of an expert individual who performs a given task. The learning engine 110 can do this in any of the various ways described herein, such as by determining an action set for the expert individual to perform the task, storing the action set as a target action for the task in a semantic action database, and inserting the actions in the action set, the environmental conditions for the action set, or a combination of the two as entries into a work context knowledge graph. As described herein, the semantic action database can be configured to reference a work context knowledge graph to specify target actions based on the task and the environmental conditions of the environment in which the individual (e.g., an expert or an AR user or a VR user) performs the task.
[0018] In operation, the expert avatar engine 112 can access a gesture set from a digital data stream of a target individual performing a task in an environment, wherein gestures of the gesture set are represented by joint positions (e.g., body joints) of the target individual, classify the gestures of the gesture set into discrete actions, retrieve target actions for performing the task in the environment from a semantic action database, and generate guidance for the target individual based on a comparison between the discrete actions classified for the target individual and the target actions retrieved from the semantic action database. The expert avatar engine 112 can also provide guidance to the target individual to assist the target individual in performing the task, such as in the form of a virtual 3D avatar in an AR or VR environment, in any of the manners described herein.
[0019] These and other expert-based guidance features and technical advantages are described in more detail below. Many of the examples and descriptions provided herein are explained with respect to tasks that are specific to individual performance. Thus, the expert-based guidance technology of the present disclosure can be implemented to support and assist in the performance of a single task, and a task can refer to any job to be performed. In an industrial context, the complexity of tasks can vary to almost any degree, from simple tasks such as inserting a screw into a threaded opening on a metal frame to complex tasks such as assembling a vehicle engine, etc. The expert-based guidance technology described herein is flexible in that it can be adapted and applied to tasks of any complexity and difficulty, thereby having broad applicability and having expert avatars available for any type of need, task, or project.
[0020] Figure 2 An example of capturing expert knowledge to support expert-based guidance according to the present disclosure is shown. Specifically, Figure 2 An illustrative example is provided by which the learning engine 110 can capture expert knowledge for performing a task by observing and analyzing the movements of an expert individual in a physical (e.g., non-virtual) environment. In general, the learning engine 110 can process the movement data of the expert individual while performing the task to accurately classify and categorize the expert individual's actions while performing the task. The learning engine 110 can then semantically classify the expert individual's movements to determine a target set of actions taken by the expert to perform the task.
[0021] In order to pass Figure 2 To illustrate, an environment 200 is shown in which an expert individual 202 performs a task. Figure 2In the specific example of , the tasks performed by expert individual 202 include operating manufacturing equipment in a production line at a factory, although any suitable tasks are contemplated herein. As noted herein, an expert (e.g., expert individual 202) can refer to any person who includes or possesses a threshold amount of knowledge, experience, or ability to perform a given task. Such thresholds can be configured or measured in any relevant or meaningful manner. An expert can be a person with a certain amount of work experience in a particular field, a person with specific educational requirements, a person with a particular familiarity with a given procedure or workflow, or other person designated by any suitable entity (e.g., a company, a certification board, an industry expert group, etc.). Accordingly, an expert can possess knowledge and "real world experience" in performing a particular task that can be captured and form the basis upon which expert-based guidance can be provided to other individuals performing a particular task.
[0022] The environment 200 can be a physical environment, such as a non-virtual background, such as an actual workshop or field service location where an expert individual 202 operates a machine to perform a task. To support expert knowledge capture, the environment 200 can include any number of sensors to capture motion data of the expert individual 202 as the expert individual 202 performs the task. The sensor can take the form of any device capable of capturing data about the actions or movements of the individual expert 202. As an example, the environment 200 shown in Figure 200 includes a camera, such as camera 204, to capture motion data of the individual expert 202. The camera 204 can be an RGBD camera that can track additional depth information of the expert individual 202. The video stream of the expert individual 202 can capture the motion data of the expert individual 202, and the learning engine 110 can access the video stream captured by the camera 204 or other data streams captured by the sensors of the environment 200.
[0023] The captured sensor data of the expert individual 202 may be processed using gesture recognition technology. A gesture recognizer may be implemented as a software component that can calculate the body posture of a human body based on a kinematic human body model (e.g., based on the joints and limb connections of the human body model). Figure 2 In the example of FIG, a gesture recognizer can generate a gesture set for an expert individual 202 from video data (e.g., video frames) captured by camera 204. The gesture set can specify a sequence of gestures of expert individual 202, such as by positions of joints of expert individual 202 in consecutively sampled video frames from camera 204 or other time-sequential sensor data.
[0024] For joint recognition and pose calculation, the pose recognizer can utilize any number of software libraries or AI technologies, such as deep learning neural networks such as HRNet, MediaPipe, OpenPose, PoseNet, etc. In some implementations, the learning engine 110 can cascade or otherwise combine joint recognition technology with finger tracking technology, as doing so can provide a broader or more complete view of the expert's movements when performing a task. Finger tracking technology can also allow expert movements to be presented to AR and VR users with greater efficiency based on expert guidance (e.g., provided by a virtual 3D avatar). Thus, the learning engine 110 can support the generation or access of a calculated pose set with finger joint positions.
[0025] In any manner described herein, the learning engine 110 may access the gesture sets of individual experts performing the task. The learning engine 110 itself may implement any suitable gesture recognition technology to determine the gesture set or otherwise receive a gesture set computed by a gesture recognizer external to (e.g., remote or logically separate from) the learning engine 110. Figure 2 In the example of , the learning engine 110 accesses a set of gestures 210 computed from sensor data captured from an expert individual 202 performing a task in an environment 200 .
[0026] To further support knowledge capture by the expert individuals 202, the learning engine 110 can classify the gesture set 210 into discrete actions. Discrete actions can refer to any form of grouping of a set of human gestures into a finite or semantically atomic classification, referred to herein as actions. Examples of actions can include semantic terms such as "stand," "bend," "reach," "walk," "sit," "lift," "push," "pull," etc. In an industrial context for performing a specific task, the learning engine 110 can limit the classification to a limited number of actions, as many industrial tasks only require a limited set (e.g., a few dozen) of actions to achieve satisfactory performance.
[0027] The action classifier technique can be implemented as a software component that receives a stream of body gestures (e.g., gesture set 210) and classifies the body gestures into discrete actions. An example of such a component is shown as Figure 2 The action classifier 220 in the learning engine 110 can be implemented by a neural network (NN) architecture, such as a long short-term memory (LSTM) network, a TransformerNN, a deep NN, a small sample learning, etc. The learning engine 110 itself can implement the action classifier 220 (or any suitable action classifier technology) to classify the gesture set. In other implementations, the learning engine 110 can classify the gesture set by receiving the classified actions from an action classifier component external to the learning engine 110 (e.g., remote or logically separated).
[0028] In some implementations, the learning engine 110 (e.g., via the action classifier 220) can further classify actions into combinations of actions in a gesture set. Such a combined action can be specified as a combination of other actions, such as a "stand_arms extended_overhead" action, which can be a combination of "stand" and "arms extended" actions. The actions classified by the learning engine 110 can be discrete, in that gestures (e.g., a subset of gestures in the gesture set 210) can be classified as separate and distinct actions. The sequence of actions classified by the learning engine 110 can form a set of target actions that the expert individual 202 takes to perform a task. The target actions caused by the expert individual 202 can semantically precisely define the set (and sequence) of actions taken to perform a task. The actions of such expert individuals 202 can be referred to as "target" actions because they represent exemplary or model action sequences that an expert takes in order to perform a given task.
[0029] The learning engine 110 can use the semantic action database to store the captured expert knowledge for performing tasks. Figure 2 In the example shown, the learning engine 110 can implement or otherwise access a semantic action database 230. The semantic action database 230 can store target actions for performing specific tasks, and such target actions can be derived from actual performance of specific tasks by expert individuals. Thus, the learning engine 110 can store target actions classified from a gesture set 210 of an expert individual 202 performing a task in an environment 200 in the semantic action database 230. In some implementations, the learning engine 110 can also store the gesture set 210 in the semantic action database 230, and can further link a particular gesture subset to a given action to which the gesture subset is classified.
[0030] Note that semantic action database 230 does not need to store video data of expert individuals 202 performing a task. Instead, entries in semantic action database 230 capture or semantically characterize the movements of expert individuals 202 through classified actions (and, in some implementations, corresponding gesture sets) without the need for video data. Consequently, the amount of data required to characterize the movements of an expert performing a task can be relatively compact (and significantly smaller in the absence of video data), while still maintaining sufficient semantic clarity to support the generation and provision of guidance to other non-expert individuals performing the task.
[0031] As yet another example feature, the learning engine 110 can store the working context of the environment 200 in which the expert individual 202 performs a task, as well as the classified actions used to perform the task. The working context of task performance can refer to any quantifiable aspect of the environment in which the individual performs the task, the task itself, or the individual performing the task. Therefore, the working context of task performance can be measured and specified in almost unlimited ways. By considering the working context, the learning engine 110 can learn, track, and process various factors that may affect task performance, which can allow relevant guidance to be generated when other (non-expert) individuals other than the expert perform the task in different environments. Various examples of working contexts are presented herein.
[0032] The work context for a given task may include component data for any components involved in the task. Learning engine 110 can capture physical component dimensional values, structural features, batch numbers, component tolerances, and any other component data values as work context for performing the task. Similarly, the work context for a given task may include tool data for any tools used to perform the task, such as tool parameters, maintenance schedules, machine type, and any other quantifiable tool values.
[0033] As another example, the learning engine 110 can also quantify environmental conditions as a work context for performing a given task. Environmental conditions can include any features in the environment in which the task is performed, and therefore can include component data and tool data. Other environmental conditions can include ambient temperature, weather characteristics (e.g., outdoor environment), pressure level, humidity, resource consumption level (e.g., power consumption, network bandwidth, memory storage level, processor utilization, etc.), etc. Such environmental conditions can be captured by sensor data in the environment (e.g., the environment 200 in which the expert individual 202 performs the task). For a virtual environment, environmental conditions can be tracked, extracted or otherwise obtained by software (e.g., by using specific parameters, features and settings of the virtual environment in which the task is performed in VR). As another example, any quantifiable aspect of the individual performing the task can be tracked as a work context for performing the task. These aspects include the height or age of the individual, whether the individual is right-handed or left-handed, or any other aspect of the individual.
[0034] While some non-exhaustive examples of work context are presented herein, the work context of a given task may include any aspect related to the task, and the learning engine 110 may track the work context accordingly. The learning engine 110 may track the work context of a task via a knowledge graph. A knowledge graph may refer to a data model for integrating a graph structure of data. Thus, a knowledge graph may specify a collection of interconnected descriptions of entities, objects, relationships, events, abstract concepts, and the like. A knowledge graph may specify the context in which data objects exist by specifying the semantics of node links or semantic metadata. Accordingly, a knowledge graph may be a particularly suitable data structure by which the learning engine 110 may track the work context of task execution.
[0035] The learning engine 110 may build or otherwise maintain a work context knowledge graph to track the work context of a task. Figure 2 In the example of , the learning engine 110 maintains a work context knowledge graph 240 and inserts entries (e.g., tuples) into the work context knowledge graph 240 to store context data. The nodes and edges of the work context knowledge graph 240 can be constructed by tuple insertion, where the edges specify the semantic relationships between objects. Through the work context knowledge graph 240, the learning engine 110 can achieve a general semantic description and understanding of any aspect of the task and the environment in which the task is performed. In this regard, the work context knowledge graph 240 can summarize expert knowledge and work context data to communicate to others. In addition, the learning engine 110 can make full use of the reasoning power of the knowledge graph to learn new relationships in the work context.
[0036] In some implementations, the learning engine 110 may link the work context knowledge graph 240 to the semantic action database 230. By doing so, the semantic action database 230 may store or otherwise reference the work context conditions, values, and any relevant aspects of performing an action for a given task. The learning engine 110 may implement the link from the semantic action database 230 to the work context knowledge graph 240 as a reference from a specific target action in the semantic action database 230 to a specific node or edge in the work context knowledge graph 240. Such links may provide insight and semantic understanding into the environmental conditions, tools, components, and other relevant contextual information for performing the specific steps, actions, and motions of a task, which may allow for more detailed and relevant guidance to be provided to other individuals performing the task.
[0037] As described herein, the learning engine 110 may maintain a work context knowledge graph to track any relevant aspects of the work context in which a task is performed. To maintain the work context knowledge graph, the learning engine 110 may populate or otherwise insert entries into the work context knowledge graph in various ways. For example, for expert knowledge captured through video recordings of tasks performed by individual experts in an entity context (e.g., Figure 2 ), the learning engine 110 can extract any relevant work context data from the video stream and insert it into the work context knowledge graph as corresponding nodes and edges. For example, depth information between a user and various parts or tools may be contained in a video stream (e.g., captured by an RGBD camera). The learning engine 110 can process the video stream data to determine the corresponding depth values and insert such work context data into the work context knowledge graph.
[0038] As another example, the learning engine 110 may explicitly insert tuples or relations, e.g., via input from the expert individuals 202 themselves through an I / O interface to the learning engine 110. As another example, the learning engine 110 supports extracting engineering data from an engineering tool (e.g., a computer-aided design (CAD) system, a computer-aided engineering (CAD) tool, a computer-aided manufacturing (CAM) application, a product lifecycle management (PLM) system, or any other engineering system or tool). Figure 3 Describe in more detail example features of expert knowledge capture and work context tracking through engineering tools.
[0039] Figure 3 Another example of capturing expert knowledge to support expert-based guidance according to the present disclosure is shown. Figure 3 In the example of , the learning engine 110 can capture expert knowledge for storage in the semantic action database 230 and the work context knowledge graph 240. Specifically, the learning engine 110 can achieve this by extracting expert knowledge and context data from engineering tools. Engineering tools (which can include CAD, CAM, CAE, and systems as non-exhaustive examples) can specify various characteristics of parts, products, tools, manufacturing processes, and other related data in a digital format. While each respective engineering tool can implement and store data according to a specific (sometimes proprietary) data format, the learning engine 110 can support the extraction of engineering data from engineering tools into a common semantic and ontological understanding, i.e., the work context knowledge graph 240.
[0040] exist Figure 3In the example shown, the learning engine 110 extracts expert knowledge from a CAD application 300. The CAD application 300 is shown merely as an example of an engineering tool from which the learning engine 110 can extract data for storage in the semantic action database 230 or the work context knowledge graph 240. For example, the learning engine 110 can extract the engineering designs (e.g., CAD models) of any relevant parts or tools of the environment in which the process is performed. The CAD engineering data may include part dimensions, tolerances, material properties, etc. The extracted engineering designs (and the underlying engineering data) can then be converted into tuples supported by the knowledge graph and thus inserted into the work context knowledge graph 240.
[0041] Many modern engineering tools support extracting engineering data into semantic formats supported by knowledge graphs, and the learning engine 110 can leverage any engineering tool's supported or pre-existing data export tools. Additionally or alternatively, the learning engine 110 can apply any data extraction, information processing, and cross-domain link discovery techniques to process data from the CAD application 300 and insert it into the work context knowledge graph 240.
[0042] The learning engine 110 may also support extracting expert knowledge from engineering tools for storage in the semantic action database 230. In some examples, the CAD application 300 or other engineering tool may store or specify an instruction set for performing a task. The instruction set may include any textual or visual instructions for the engineering tool, such as an instruction manual for using a particular machine or industrial tool. The learning engine 110 may extract the instruction set from the engineering tool and convert the instruction set into a semantic format suitable for the semantic action database 230. In this regard, the learning engine 110 may classify the derived instruction set into discrete actions that fit into the semantic framework of the target action stored in the semantic action database 240. The method used by the learning engine 110 to do this may vary based on how the engineering tool stores or provides the instruction set.
[0043] For text-based instruction sets, the learning engine 110 can parse the text of the instruction set and extract the relevant actions used to execute the instructions. In a sense, the learning engine 110 can translate or convert the text of the instruction set (e.g., manual) of the engineering tool into atomic actions of a semantic framework, and the semantic action database 240 stores the actions for it. Typically, in an industrial context, the overall steps to perform a task are limited, so the instruction manual can be translated or converted into cost-disclosed semantic actions with improved efficiency and speed. The learning engine 110 can implement any suitable technology to support this conversion.
[0044] As another example, an engineering tool may provide a virtual instruction video, for example, where a virtual character performs a task as part of the instruction video. Such instruction videos or virtual instructions may include a set of gestures and categorized actions of an expert performing the task. In this case, the learning engine 110 may extract the gesture set, action sequence, or a combination of the two from the engineering tool itself.
[0045] In other implementations, the learning engine 110 can extract expert knowledge from these engineering tools in a manner consistent with video data of individual experts performing tasks in a physical environment. Instead of sensor data in the form of a video stream, the learning engine 110 can provide virtual learning videos as input to the gesture recognizer to access the gesture set of a virtual avatar performing a task in the virtual environment. The virtual video can be processed in a manner consistent with processing a video stream of a physical environment, with gesture recognition performed on a virtual 3D avatar in the video stream rather than a human. The learning engine 110 can then classify the gesture set of the virtual 3D avatar in the learning video and store the classified actions as target actions in the semantic action database 240. In this case, the "expert" from whom the learning engine 110 captures expert knowledge can be the virtual avatar virtually performing the task in the instruction video. The work context of the virtual instruction video can also be derived from the engineering tool and stored as a data entry in the work context knowledge graph 240.
[0046] In any of the manners described herein, the learning engine 110 can capture the knowledge of the expert performing the task and store the captured knowledge in a common semantic format. Through knowledge graph technology, the learning engine 110 can track the work context in which the expert performs the task and provide a more comprehensive understanding of the various environmental conditions and individual factors that contribute to successful task execution. Extracting instruction sets and work context from engineering tools can provide an additional or alternative mechanism by which the learning engine 110 can populate the work context knowledge graph 240 and the semantic action database 230.
[0047] Expert knowledge captured in a semantic action database (e.g., in the form of target action sequences for performing a task), together with the work context in which the target action sequences are performed, can provide an accurate and flexible definition of successful task performance, against which the action sequences of other individuals can be compared. Through such comparisons, expert-based guidance can be provided to other individuals attempting to perform a given task, for example through avatars that can interact with these other individuals to provide verbal guidance or visual demonstrations. Figure 4 Example features for generating and providing expert-based guidance using the semantic action database 230 and the work context knowledge graph 240 are described.
[0048] Figure 4An example of providing expert-based guidance to individuals performing tasks in an environment according to the present disclosure is shown. Figure 4 The example features of the embodiment are described using the expert avatar engine 112 as an example, but any implementation consistent with the present disclosure is contemplated herein. The expert avatar engine 112 can leverage the expert knowledge captured in the semantic action database 230 and the work context knowledge graph 240, such as through a virtual 3D avatar, to provide guidance to an individual in performing a given task in a given environment.
[0049] To illustrate, Figure 4 The environment 400 in which the target individual performs the task is included. Please note that the environment 400 in which the target individual performs the task does not need to be the same as Figure 2 The environment 200 in which the target individual 202 performs the task is exactly the same as the environment 200 in which the target individual 202 performs the task. For example, environment 400 can be a virtual environment of an industrial virtual reality background, and the target individual 402 can perform the task virtually in the virtual reality background. There may be any number of changes in environmental conditions between environment 200 and environment 400, but the expert avatar engine 112 can still provide relevant expert-based guidance. The expert avatar engine 112 can provide guidance for performing the task through a virtual 3D avatar rendered in the virtual reality background. As another example, example 400 can be a physical environment in which the target individual 402 physically performs the task, and therein, the expert avatar engine 112 can provide guidance through AR technology, for example, through a virtual 3D avatar superimposed in the field of view of the target individual 402 via an AR device.
[0050] To provide expert-based guidance, the avatar engine 112 can identify and track the movement of a target individual 402 performing a task in the environment 400. To this end, the environment 400 can include any number of sensors to capture movement data of the target individual 402. The sensors can include Figure 2 Any sensor described herein, such as a camera or other sensor. The expert avatar engine 112 can access the set of gestures of the target individual 402 performing the task from the motion data of the target individual 402. In this regard, the expert avatar engine 112 can implement or otherwise access gesture recognizer technology in any manner consistent with the descriptions herein. Figure 4 In the example of , the avatar engine 112 accesses a gesture set 410 of a target individual 402 that performs a task in the environment 400 , and the gesture set 410 can be represented by joint positions (e.g., including finger joint positions) of the target individual 402 .
[0051] The avatar engine 112 may also access environmental conditions 412 for the target individual 402 performing the task in the environment 400. The environmental conditions 412 may specify any quantifiable aspect of the environment in which the target individual 402 performs the task, and thus may include component dimensions, tool parameters, and any other aspect of task performance as described herein. The avatar engine 112 may access the environmental conditions 412 in a variety of ways. The environment 400 may include any suitable sensors through which the expert avatar engine 112 may access relevant environmental conditions, such as temperature, pressure, humidity, resource availability, and the like. As an additional or alternative example, the expert avatar engine 112 may support direct input of the environmental conditions 412 by the target individual 402, for example, by engaging in a natural language conversation with a virtual 3D avatar generated by the expert avatar engine 112 for the environment 400.
[0052] Expert avatar engine 112 can itself derive any number of environmental conditions for target individual 402 and environment 400, such as by processing gesture set 410 to determine whether target individual 402 is performing a task with a particular dominant hand, or the target individual's height or relative position with respect to other objects in environment 400. In any manner described herein, expert avatar engine 112 can access gesture set 410 and environmental conditions 412 for target individual 402 performing a task in environment 400.
[0053] In a manner consistent with the description herein, the expert avatar engine 112 can classify the gestures of the gesture set 412 into discrete actions, and do so via the action classifier technology described herein. The expert avatar engine 112 can then retrieve the target action from the semantic action database 420 to perform the task. By comparing the action sequence classified for the target individual 402 with the target action captured for the expert individual 202 performing the task, the expert avatar engine 112 can determine the deviation between the target individual 402's performance of the task and the expert's performance of the task through the action comparison.
[0054] The expert avatar engine 112 can compare the action sequence of the target individual 402 with the target action sequence of the retrieved expert in various ways. In some implementations, the expert avatar engine 112 can synchronize the two action sequences based on the initial action sequence detected for the target individual 402, the target action of the expert individual retrieved from the semantic action database 230, or a combination of the two. For example, the target action of the expert performing the task may start with a specific action sequence (e.g., action 1-action 2-action 3). The expert avatar engine 112 can synchronize the action sequence classified for the target individual 402 when the action 1-action 2-action 3 sequence of the target individual 402 is detected. Any threshold for matching actions or action subsequences can be used to synchronize the two action streams for comparison. As another example, the expert avatar engine 112 can synchronize the action sequence of the target individual 402 and the target action of the retrieved expert based on timestamps or by any suitable time-based synchronization.
[0055] When comparing the action sequence of the target individual 402 and the target action sequence of the expert, the expert avatar engine 112 can determine any deviation between the two action sequences as a difference between the target individual 402 performing the task and the expert's task performance. Deviation can refer to any difference between the classified action sequence of the target individual 402 and the target action sequence performed by the expert. The expert avatar engine 112 can take action (e.g., generate guidance) based on the degree of deviation between the two action sequences. For deviations that are determined to be minor deviations that will not affect the target individual 402's performance of the task, the expert avatar engine 112 may not take any action. For major deviations that are different between the action sequences, the expert avatar engine 112 can intervene by providing guidance (including occasionally requesting the target individual 402 to stop the action).
[0056] In some implementations, the expert avatar engine 112 can consider the work context for performing the task to determine the deviation between the target individual's 402 action sequence and the expert's target action sequence (and the extent of such deviation). To this end, the expert avatar engine 112 can query the work context knowledge graph 240 using the specific action performed by the target individual 402 and the work conditions 412 for the specific action. The work context knowledge graph 420 can specify certain constraints, restrictions, or allowed deviations for the target individual to perform the specific action. Through these constraints, restrictions, or allowed deviations, the expert avatar engine 112 can characterize the extent to which any determined deviation between the action sequence and / or the work context affects the performance of the task.
[0057] The expert avatar engine 112 can classify the deviations between the action sequence and the work context into major deviations and minor deviations based on any number of deviation criteria. In some cases, the deviation criteria can specify that certain actions in the target action sequence are key actions, and when the action sequence of the target individual 402 deviates from the key instructions in the expert's target action sequence, it is determined to be a major deviation. Minor deviations can be characterized as minor differences in the target individual's posture or environmental conditions that do not affect the actual performance of the task. For example, the target individual 402 performs the task with its left hand, while the target action performs the task with its right hand, which may be characterized as a minor deviation in the action sequence. In some cases, the work context data of the work context knowledge graph 420 can specify a criticality measure of the context data, so that a query to the work context knowledge graph 240 can indicate whether the difference in specific work context data or corresponding actions is classified as a major deviation or a minor deviation.
[0058] The expert avatar engine 112 may generate guidance for the target individual 402 based on a comparison between the discrete actions classified for the target individual 402 and the target actions retrieved from the semantic action database 240. The comparison by the expert avatar engine 112 may indicate a deviation classification, which may indicate the degree of the deviation and the impact on performing the task, e.g., whether it is major or minor on a criticality scale or according to any suitable and configurable classification scheme.
[0059] In some implementations, the expert avatar engine 112 may implement a guidance generator that drives the feedback and guidance that the virtual 3D avatar can provide to the target individual 402 performing the task. An example of a virtual 3D avatar that the expert avatar engine 112 can render is shown in FIG. Figure 4 4. Target individual 402 is shown as an expert avatar 430, which can be any virtual avatar generated and controlled by the expert avatar engine 112 to provide the expert-based coaching functionality of the present disclosure. For minor deviations in the action sequence or work context (or no deviation at all), the expert avatar engine 112 does not need to utilize the guidance generator and determines not to provide guidance to the target individual 402. For major deviations, the expert avatar engine 112 can generate guidance to assist the target individual 402 in performing the task. In some implementations, the guidance generator can provide verbal feedback, such as in the form of natural language that the expert avatar engine 112 can provide to the target individual 402.
[0060] As another form of guidance, the guidance generator can generate guidance in the form of a demonstration. For example, the expert avatar engine 112 can drive the expert avatar 430 to virtually perform a deviation action for the target individual 402, either in a virtual environment where the target individual 402 performs the task or as a virtual overlay in a physical environment. By doing so, the expert avatar engine 112 can utilize the joint positions of a subset of postures for the deviation action and drive the expert avatar 430 based on the posture subset to virtually demonstrate the deviation action to the target individual 402. This form of guidance can be combined with a natural language dialogue, which can provide a conversation and collaborative experience for the target individual 402. In some implementations, the expert avatar engine 112 can provide such a dialogue to convey any relevant or additional information to the target individual 402, and can use voice to implement this functionality through a text-to-speech (TTS) component, thereby providing a natural interactive environment.
[0061] In a consistent manner, the expert avatar 430 provided by the expert avatar engine 112 can answer questions of the target individual 402, which can include querying the work context knowledge graph 240 to provide answers to any questions that the target individual may ask. When providing guidance, the expert avatar engine 112 can animate or otherwise render the expert avatar 430 in the field of view of the target individual 402, for example, through an AR or VR device (e.g., a head-mounted device). This rendering of a virtual 3D avatar does not require any artificial intelligence to implement, which can reduce the complexity and computational requirements of the expert-based guidance technology of the present disclosure compared to AI-driven virtual assistants. In addition, the expert avatar engine 112 can position the rendered expert avatar 430 in close proximity to the target individual 402 for a more effective knowledge transfer experience with the target individual 402.
[0062] By any of the methods described herein, the expert avatar engine 112 can provide guidance to the target individual 402 to assist the target individual in performing the task. Examples of such guidance are provided as guidance 420. Figure 4 As shown, it can take the form of textual guidance provided by the voice and TTS functions of the virtual 3D avatar, animation and demonstration of any action or sub-step of performing the task, or any other form of animated guidance provided by the virtual avatar to assist the target individual 402. Based on the deviation of the action sequence, the expert avatar engine 112 can identify the missing part that the target individual forgot, and the generated guidance can include the virtual 3D avatar's identification of the missing part (e.g., pointing to it). Other forms of animated guidance can guide the target individual in a human-like manner, imitate the (one or more) actions required to perform the specific task step where the deviation occurred, or any other form of appropriate assistance to provide the target individual 402 with the task.
[0063] In any of the various ways described herein, the expert avatar engine 112 can generate guidance 420 and provide guidance 420 to a target individual 402 performing a task in the environment 400. As described herein, the guidance can be generated based on a direct comparison between the target action sequences of the expert performing the task. Through the classified action sequences, the expert avatar engine 112 can have a consistent semantic understanding of the actions performed by the target individual 402 and the target action sequences performed by the expert performing the task. This direct comparison along a consistent semantic framework can achieve an efficient and accurate comparison, allowing the expert avatar engine 112 to generate guidance based on actual expert actions (rather than predictions like AI-based virtual assistants). In addition, the working context knowledge graph applied by the expert avatar engine 112 can allow the expert avatar engine 112 to determine whether the deviation in the action is minor or major, and customize the generated guidance accordingly.
[0064] In some implementations, the expert avatar engine 112 or the learning engine 110 can update the work context knowledge graph 240. Since the expert avatar engine 112 provides guidance for multiple different individuals who perform tasks in different environments with different work contexts, the expert avatar engine 112 can track the various executed action sequences of the individuals. Each action and its corresponding work context can be inserted into the work context knowledge graph 240 as an entry. The expert avatar engine 112 or the learning engine 110 can analyze the work context knowledge graph 240 and / or the action sequence through various analysis techniques to evaluate the effectiveness of the executed action sequence. In some cases, the learning engine 110 can, for example, determine that a different action sequence may be optimal compared to the target action captured for the expert. In this case, the learning engine 110 can update the semantic action database 230 with the updated target action sequence (e.g., learned through the analysis process and optimization analysis). Any suitable form of feedback loops, knowledge collection, analytical processing, optimization techniques, knowledge graph reasoning techniques, etc. is contemplated herein to continuously update (e.g., improve or optimize) the semantic action database 230, the work context knowledge graph 240, or the avatar itself.
[0065] In some implementations, the work context knowledge graph 240 can capture any knowledge related to the task, the individuals performing the task, and the various environments in which the task is performed, and the learning engine 110 can continuously update the work context knowledge graph 240. Real-time context and execution data from individuals performing the task can be captured, analyzed, evaluated, and / or stored in the work context knowledge graph 240. The analysis can include any type of metrics or estimates of executed process steps, effectiveness, efficiency, KPIs, or any other form of measurement to assess the performance of the task, which the learning engine 110 can capture in the work context knowledge graph 240. Thus, the work context knowledge graph 240 can support the various expert-based coaching techniques presented herein.
[0066] Figure 5 An example of logic 500 that the system may implement to support expert-based guidance in AR and VR environments is shown. For example, the computing system 100 may implement the logic 500 as hardware, executable instructions stored on a machine-readable medium, or a combination of both. The computing system 100 may implement the logic 500 via the learning engine 110, the expert avatar engine 112, or a combination of both, through which the computing system 100 may perform or execute the logic 500 as a method to support providing expert-based guidance in accordance with the present disclosure. A description of the logic 500 is provided below using the expert avatar engine 112 as an example. However, the computing system may also implement various other options.
[0067] When implementing logic 500, the expert avatar engine 112 can access a gesture set (502) from a digital data stream of a target individual performing a task in an environment. As noted herein, the gestures of the gesture set can be represented by the joint positions of the target individual. The expert avatar engine 112 can further classify the gestures of the gesture set into discrete actions (504) and retrieve target actions for performing the task in the environment from a semantic action database. The expert avatar engine 112 can then generate guidance (506) for the target individual based on a comparison between the discrete actions classified for the target individual and the target actions retrieved from the semantic action database, and provide guidance to the target individual to assist the target individual in performing the task (508).
[0068] Figure 5 The logic 500 shown in provides an illustrative example by which the computing system 100 can support expert-based guidance in AR and VR environments. Additional or alternative steps in the logic 500 are contemplated herein, including any of the various features described herein for the learning engine 110, the expert avatar engine 112, or any combination thereof. For example, the method 500 can additionally or alternatively include any of the expert knowledge capture features described herein for the learning engine 110.
[0069] Figure 6 An example of a computing system 600 that supports expert-based guidance in AR and VR environments is shown. The computing system 600 may include a processor 610, which may take the form of a single processor or multiple processors. The processor(s) 610 may include a central processing unit (CPU), a microprocessor, or any hardware device suitable for executing instructions stored on a machine-readable medium. The computing system 600 may include a machine-readable medium 620. The machine-readable medium 620 may take the form of any non-transitory electronic, magnetic, optical, or other physical storage device that stores executable instructions, such as Figure 6 Shown are learning instructions 622 and expert avatar instructions 624. Thus, the machine-readable medium 620 can be, for example, a random access memory (RAM), such as dynamic RAM (DRAM), flash memory, spin transfer torque memory, electrically erasable programmable read-only memory (EEPROM), a storage drive, an optical disk, or the like.
[0070] The computing system 600 may execute instructions stored on the machine-readable medium 620 via the processor 610. Executing the instructions (e.g., the learning instructions 622 and / or the expert avatar instructions 624) may cause the computing system 600 to perform any of the expert-based guidance features described herein, including any features according to the learning engine 110, the expert avatar engine 112, or a combination of both.
[0071] For example, the processor 610 executing the learning instructions 622 can cause the computing system 600 to capture the expert knowledge of an expert individual performing a given task, for example, by determining an action set for the expert individual to perform the task, storing the action set as a target action for the task in a semantic action database, and inserting the actions in the action set, the environmental conditions for the action set, or a combination of the two as entries into a work context knowledge graph. As described herein, the semantic action database can be configured to reference the work context knowledge graph to specify the target action based on the task and the environmental conditions of the environment in which the individual (e.g., an expert or an AR or VR user) performs the task.
[0072] Processor 610 executing expert avatar instructions 624 may cause computing system 600 to access a gesture set from a digital data stream of a target individual performing a task in an environment, classify gestures of the gesture set into discrete actions, retrieve target actions for performing the task in the environment from a semantic action database, and generate guidance for the target individual based on a comparison between the discrete actions classified for the target individual and the target actions retrieved from the semantic action database. Processor 610 executing expert avatar instructions 624 may also cause computing system 600 to provide guidance to the target individual to assist the target individual in performing the task, for example, in the form of a virtual 3D avatar rendered in an AR or VR environment, in any manner described herein.
[0073] Any additional or alternative expert-based guidance features described herein may be implemented via learning instructions 622 , expert avatar instructions 624 , or a combination of both.
[0074] The systems, methods, devices, and logic described above, including the learning engine 110 and the expert avatar engine 112, may be implemented in many different ways in many different combinations of hardware, logic, circuitry, and executable instructions stored on a machine-readable medium. For example, the learning engine 110, the expert avatar engine 112, or a combination thereof may include circuitry in a controller, microprocessor, or application-specific integrated circuit (ASIC), or may be implemented with a combination of discrete logic or components or other types of analog or digital circuitry, combined on a single integrated circuit or distributed across multiple integrated circuits. A product (e.g., a computer program product) may include a storage medium and machine-readable instructions stored on the medium that, when executed in an endpoint, computer system, or other device, cause the device to perform operations according to any of the descriptions above, including according to any of the features of the learning engine 110, the expert avatar engine 112, or a combination thereof.
[0075] The processing power of the systems, devices, and engines described herein (including the learning engine 110 and the expert avatar engine 112) can be distributed across multiple system components, such as across multiple processors and memories, optionally including multiple distributed processing systems or cloud / network elements. Parameters, databases, and other data structures can be stored and managed separately, can be combined into a single memory or database, can be logically and physically organized in a variety of different ways, and can be implemented in a variety of ways, including data structures such as linked lists, hash tables, or implicit storage mechanisms. Programs can be parts of a single program (e.g., subroutines), can be separate programs, distributed across multiple memories and processors, or implemented in a variety of different ways, such as in libraries (e.g., shared libraries).
[0076] While various examples are described above, many more implementations are possible.
Claims
1. A method comprising: By computing the system: accessing a pose set from a digital data stream of a target individual performing a task in an environment, wherein poses of the pose set are represented by joint positions of the target individual; classifying the gestures of the gesture set into discrete actions; Retrieving a target action for performing the task in the environment from a semantic action database, wherein the semantic action database is configured to reference a work context knowledge graph to specify the target action based on the task and environmental conditions of the environment in which the target individual performs the task; generating guidance for the target individual based on a comparison between the discrete actions classified for the target individual and the target actions retrieved from the semantic action database; The guidance is provided to the target individual to assist the target individual in performing the task.
2. The method according to claim 1, wherein The environment includes a physical environment, and wherein the posture set is determined from a video stream of the target individual performing the task in the physical environment; and The guidance is provided via an augmented reality (AR) device used by the target individual or another individual in the physical environment.
3. The method according to claim 1, wherein The environment comprises a virtual reality environment, and wherein the target individual comprises an avatar of a user in the virtual reality environment, and wherein the set of gestures is determined from the user avatar performing the task in the virtual environment; and This includes providing the guidance through a virtual avatar in the virtual reality environment.
4. The method according to any one of claims 1 to 3, further comprising: Capturing expert knowledge to store in the semantic action database, the work context knowledge graph, or a combination thereof, includes: Determine the set of actions for the expert individual to perform the task; storing the action set as the target action of the task in the semantic action database; and The actions in the action set, the environmental conditions for the action set, or a combination of the two are inserted as entries into the work context knowledge graph.
5. The method according to claim 4, wherein Determining the set of actions for the individual expert to perform the task includes deriving a set of instructions from an engineering tool.
6. The method according to claim 4, wherein: Determining the action set for the expert to perform the task includes: accessing an expert pose set from a digital data stream of an expert individual performing the task, wherein poses of the expert pose set are represented by joint positions of the expert; and The gestures of the expert gesture set are classified into discrete actions to form the action set of the individual expert.
7. The method according to any one of claims 1 to 6, further comprising: The work context knowledge graph or the semantic action database is updated based on an analysis process performed by analyzing the work context data stored in the work context knowledge graph.
8. A system comprising: a semantic action database configured to reference a work context knowledge graph to specify target actions for performing a task and environmental conditions of an environment in which an individual performs the task; as well as An expert avatar engine, which is configured to: accessing a pose set from a digital data stream of a target individual performing a task in an environment, wherein poses of the pose set are represented by joint positions of the target individual; classifying the gestures of the gesture set into discrete actions; retrieving a target action for performing the task in the environment from a semantic action database; generating guidance for the target individual based on a comparison between the discrete actions classified for the target individual and the target actions retrieved from the semantic action database; The guidance is provided to the target individual to assist the target individual in performing the task.
9. The system according to claim 8, wherein: The environment includes a physical environment, and wherein the posture set is determined from a video stream of the target individual performing the task in the physical environment; and The instructions cause the computing system to provide the guidance via an augmented reality (AR) device used by the target individual or another individual in the physical environment.
10. The system according to claim 8, wherein: The environment comprises a virtual reality environment, and wherein the target individual comprises an avatar of a user in the virtual reality environment, and wherein the set of gestures is determined from a user avatar performing the task in the virtual environment; and The instructions cause the computing system to provide the guidance through a virtual avatar in the virtual reality environment.
11. The system according to any one of claims 8 to 10, further comprising a learning engine configured to capture expert knowledge to be stored in the semantic action database, the work context knowledge graph, or a combination thereof, comprising: Determine the set of actions for the expert individual to perform the task; storing the action set as the target action of the task in the semantic action database; and The actions in the action set, the environmental conditions for the action set, or a combination of the two are inserted as entries into the work context knowledge graph.
12. The system according to claim 11, wherein The learning engine is configured to determine the set of actions for the expert individual to perform the task by deriving an instruction set from an engineering tool.
13. The system according to claim 11, wherein: The learning engine is configured to determine the set of actions for the expert to perform the task by: accessing an expert pose set from a digital data stream of an expert individual performing the task, wherein poses of the expert pose set are represented by joint positions of the expert; and The gestures of the expert gesture set are classified into discrete actions to form the action set of the individual expert.
14. The system according to any one of claims 8 to 13, wherein: The expert avatar engine is further configured to update the work context knowledge graph or the semantic action database based on an analysis process performed on the work context data stored in the work context knowledge graph.
15. A non-transitory machine-readable medium comprising instructions that, when executed by a processor, cause a computing system to perform the method according to any one of claims 1 to 7.