Systems and methods for ai-driven training and assessment in extended reality
An automated system using LLMs and digital twins addresses the inefficiencies in immersive training content creation by generating scenarios from video input and adapting to trainee performance, improving learning outcomes.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- THE ARIZONA BOARD OF REGENTS ON BEHALF OF THE UNIV OF ARIZONA
- Filing Date
- 2026-01-20
- Publication Date
- 2026-07-23
AI Technical Summary
The process of creating immersive training content in extended reality environments is manual and time-consuming, requiring specialized skills and is not easily scalable or adaptable to different scenarios, despite the benefits of immersive training.
An automated workflow using large language models (LLMs) and digital twins to generate immersive content from video input, with integrated assessment and grading features that adapt to trainee performance and incorporate expert feedback.
Facilitates efficient, adaptable, and cost-effective creation of immersive training scenarios across diverse fields, enhancing learning outcomes through automated content generation and performance evaluation.
Smart Images

Figure US2026011814_23072026_PF_FP_ABST
Abstract
Description
Atty. Docket: UA25-088PCTSYSTEMS AND METHODS FOR AI-DRIVEN TRAINING AND ASSESSMENT IN EXTENDED REALITYCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] The present application claims the benefit of the filing date of U.S. Provisional Application Serial No. 63 / 746,400, filed January 17, 2025, the entire teachings of which application is hereby incorporated herein by reference.TECHNICAL FIELD
[0002] The present disclosure relates generally to training and, more particularly, to a system and method for artificial intelligence (Al)-driven training and assessment in extended reality.BACKGROUND
[0003] A digital twin is a virtual representation, or model, of a physical object or system that uses real-time data to accurately reflect its real-world counterpart’s behavior, performance, and conditions. The virtual model allows for continuous monitoring, simulation, analysis, and prediction of outcomes, enabling better decision-making, optimization, and proactive maintenance for the physical asset without direct interaction.
[0004] Virtual reality (VR) is increasingly utilized in training environments across industries such as manufacturing, assembly lines, and laboratories. However, adapting VR applications often requires a manual, time-consuming process that must be repeated for each new scenario. There exists a need for automating adapting of VR applications to training environments.Atty. Docket: UA25-088PCTBRIEF DESCRIPTION OF THE DRAWINGS
[0005] Reference should be made to the following detailed description which should be read in conjunction with the following figures, wherein like numerals represent like parts.
[0006] FIG. 1 is a functional block diagram illustrating a system for Al-driven training and assessment in extended reality consistent with the present disclosure.
[0007] FIG. 2 is a system overview illustrating key modules for the system of FIG. 1, consistent with the present disclosure.
[0008] FIG. 3 illustrates the supported actions for an example application in a nanotechnology lab for the system of FIG. 1, consistent with the present disclosure.
[0009] FIG. 4 is an example platform pipeline for the system of FIG. 1 , consistent with the present disclosure.
[0010] FIG. 5 depicts Table 1, a summary of various actions and their attributes for the system of FIG. 1, consistent with the present disclosure.
[0011] FIG. 6 is a flowchart diagram depicting operations for one illustrative example embodiment of the algorithm for smart instruction and assessment rule generation on the system of FIG. 1, consistent with the present disclosure.
[0012] FIG. 7 is an example smart workflow for the system of FIG. 1, consistent with the present disclosure.
[0013] FIG. 8 is an example of an assessment performance and grading panel for the system of FIG. 1, consistent with the present disclosure.
[0014] FIG. 9 is an example of an assessment performance and grading panel for sub-steps for the system of FIG. 1, consistent with the present disclosure.
[0015] FIG. 10 is an example of an expert performance and grading panel for the system of FIG. 1, consistent with the present disclosure.Atty. Docket: UA25-088PCT
[0016] FIG. 11 is an example of an expert performance and grading panel for sub-steps for the system of FIG. 1, consistent with the present disclosure.
[0017] FIG. 12 is an example of an expert overview of trainee performance data panel for the system of FIG. 1, consistent with the present disclosure.
[0018] FIG. 13 depicts a block diagram of components of the computing device executing the process within the system of FIG. 1, consistent with the present disclosure.DETAILED DESCRIPTION
[0019] Immersive training is widely used and particularly valuable in fields where practical experience is essential, such as training in industrial and laboratory environments. It allows trainees to simulate real-world scenarios in a controlled environment, enabling them to practice critical tasks without the risks associated with real-life operations. In a factory, for example, workers can train in VR or Augmented Reality (AR) simulations to practice assembling complex products, learn to operate machinery, troubleshoot issues, and understand workflow processes without halting production, risking injury, or risking damage to equipment. AR refers to the realtime integration of digital information onto the real world in real-time, blending digital content with the physical environment. In laboratories, scientists and technicians can simulate experiments, chemical processes, or equipment handling in a virtual environment. This helps them master technical skills, explore different outcomes of procedures, and minimize the chances of accidents or costly mistakes in a real lab setting.
[0020] Furthermore, with an immersive training environment, trainees can access training anytime and anywhere, providing several key benefits. It allows for on-demand, self-paced learning and 24 / 7 availability, ensuring flexibility for trainees with varying schedules. Immersive training systems may significantly enhance knowledge retention, leading to fewer errors and improved performance compared to conventional methods.
[0021] Conventionally, instructional content for industrial machinery is created using written manuals, training videos, and well-defined Computer-Aided Design (CAD) models, which can beAtty. Docket: UA25-088PCTtime-consuming and lack interactivity. Immersive training offers numerous benefits, but the process of immersive authoring — creating the training content — remains largely manual and timeconsuming. Developing high-quality, interactive experiences often requires specialized skills in three-dimensional (3D) design, programming, and content creation. This can slow down the process, especially when customizing complex scenarios for specific industries. As a result, despite the advantages of immersive training, the creation of training modules can be a bottleneck, making it more challenging to scale or rapidly update training content.
[0022] In recent years, various immersive authoring methods have been proposed to simplify the creation of immersive instructional content. Some of these methods provide user-friendly interfaces that allow experts to create training environments without any programming knowledge. Others enable experts to perform demonstrations, from which tutorials are automatically generated. However, these approaches often still require extensive manual input even though a wealth of instructional content is already available in text or video formats.
[0023] Disclosed herein is an automated workflow that leverages large language models (LLMs) and digital twins to facilitate immersive content creation that can be generalized across multiple scenarios. The disclosed system requires only a video input to autonomously generate digital twins of training procedures, leveraging LLMs to extract instructional content and action-related requirements. Additionally, the system incorporates an assessment feature that evaluates and grades trainee performance using text analysis and external inputs, refining grading criteria through expert feedback. This adaptive training system adjusts its difficulty based on trainee performance, thereby enhancing learning outcomes.
[0024] With the advancement of LLMs. understanding and automating actions described in instructional content has significantly improved. The disclosed system simplifies content authoring by abstracting training behaviors into a set of predefined actions, such as moving, clicking, and checking. LLMs interpret instructional content, map required actions to system-supported behaviors and implement them in a manner that is agnostic to the task. This approach enables efficient, adaptable immersive authoring across diverse fields.
[0025] Beyond training, the disclosed system integrates a digital twin framework for comprehensive assessment. The system monitors trainee behavior, evaluates errors, and automatesAtty. Docket: UA25-088PCTgrading based on rules extracted from expert demonstration videos. LLMs enhance this process by filling in missing details and generating grading rubrics. Using both common and domain- specific knowledge embedded in LLMs. along with external sources integrated through in-context learning, the system ensures logical supplementation of the missing information.
[0026] To further improve reliability, a human-in-the-loop interface is incorporated, allowing experts to refine assessments through direct feedback. This integration with the digital twin model continuously enhances the system’s accuracy and effectiveness over time.
[0027] The disclosed system introduces a novel method for immersive authoring through descriptions, enabling the disclosed system to automatically construct training and assessment scenarios from readily accessible instructional content in video or text formats.
[0028] The disclosed system supports a range of extendable general actions commonly required in training scenarios and the action implementation is designed in a task-agnostic way, making it capable of being transferred and applied to a different application scenario seamlessly.
[0029] The disclosed system can fully automatically evaluate and grade trainee performance using its embedded common or domain-specific knowledge, along with external sources though in-context learning. The system may also increase its reliability over time by incorporating a human-in-the-loop mode that integrates expert feedback to refine the assessment criteria.
[0030] The system presents an automated workflow leveraging LLMs to streamline immersive authoring, demonstrating the capability of the LLMs to understand training tasks and fulfill various roles within the system.
[0031] Creating immersive training environments is typically a costly and time-consuming process that requires the expertise of skilled programmers. However, domain experts, who possess the necessary knowledge, often lack the programming skills needed to build these environments themselves. To overcome this challenge, several approaches have been developed to simplify the authoring process through graphical user interfaces (GUIs). One example proposed method enables experts to create training scenarios without the need for coding by developing a tool which facilitates the creation of new training scenarios for different assemblies.Atty. Docket: UA25-088PCT
[0032] In addition to the immersive authoring by doing, recent methods allow experts to create training steps by directly interacting with the environment. One example introduced an immersive authoring method that relies on content creation through demonstration. Experts perform tasks, such as assembly procedures, within an immersive interface based on their experience, and the system uses process mining algorithms to derive assembly models from event logs, capturing expert knowledge through natural interaction.
[0033] Unlike conventional methods that require experts to perform tasks in a VR environment, the disclosed system enables immersive authoring through descriptive input. The disclosed system bridges this gap by transforming text and video-based instructional materials into immersive training experiences.
[0034] Emerging immersive technologies such as VR and AR have transformed traditional training paradigms, necessitating sophisticated performance evaluation methods. The disclosed system focuses on the evaluations with rule-based or deep learning-based methods during the simulation as the learner solves tasks or missions.
[0035] Rule-based evaluation systems in VR and AR training hinge on predefined criteria and logical constructs. Traditional metrics such as completion times, error rates, and adherence to procedures are often used in rule-based evaluations. These metrics are straightforward and provide quantifiable data that reflects the trainee’s performance in achieving specific objectives outlined in the training program. While rule-based systems ensure consistency, they may not fully capture the nuances of human performance in high-fidelity simulations.
[0036] Conversely, deep learning methodologies excel at processing complex datasets derived from trainee interactions, particularly in fields such as medical education, where evaluation processes are inherently complex. Studies have highlighted the potential of these methodologies to provide nuanced insights into behavioral patterns and predictive outcomes. This approach not only enhances the adaptability of training systems through real-time feedback mechanisms but also facilitates personalized training adjustments based on dynamic performance analysis.
[0037] In the disclosed system, LLMs are used to determine scoring metrics based on video analysis and the domain knowledge of the LLM. This approach includes a human-in-the-loopAtty. Docket: UA25-088PCTsystem, allowing the LLM to iteratively learn and refine scoring metrics through in-context learning, thereby enhancing the reliability and effectiveness of the training assessments.
[0038] Generative artificial intelligence (Al) models, particularly LLMs, are transforming the landscape of VR training systems by automating the creation of immersive instructional content. This automation extends the versatility of VR training systems across various domains without the need for extensive reprogramming. Studies, such as the VRCopilot system, explore the integration of pretrained generative Al models into immersive authoring workflows for 3D layout design in VR. VRCopilot is a mixed-initiative system that supports rapid prototyping and iterative design modifications through multimodal interactions, including voice commands and gestural inputs. The system facilitates human- Al co-creation by allowing users to directly manipulate or modify Al-generated layout suggestions, enhancing user agency and creative exploration in design tasks.
[0039] Generative models enrich the VR environments with dynamic, context-aware elements. They can adapt training scenarios in real time, responding to the trainee’s interactions by adjusting the difficulty level or providing customized feedback, thereby enhancing the learning experience.
[0040] In the disclosed system, LLMs are leveraged to generate step-by-step training guides from expert video feeds, interpret instructional content, and autonomously create corresponding VR training actions and grading rubrics. This demonstrates the potential of generative Al to streamline complex training setups in extended Reality (XR).
[0041] FIG. 1 is a functional block diagram illustrating a system, generally designated 100, suitable for Al-driven training and assessment in extended reality consistent with the present disclosure. FIG. 1 provides only an illustration of one implementation and does not imply any limitations with regard to the environments in which different embodiments may be implemented. Many modifications to the depicted environment may be made by those skilled in the art without departing from the scope of the disclosure as recited by the claims.
[0042] System 100 includes computing device 110 optionally connected to network 120. Network 120 can be, for example, a telecommunications network, a personal area network (PAN), such as Bluetooth, a local area network (LAN), a wide area network (WAN), such as the Internet, or a combination of the three, and can include, but is not limited to, wired, wireless, or fiber opticAtty. Docket: UA25-088PCTconnections. In general, network 120 can be any combination of connections and protocols that will support communications between computing device 110 and other computing devices (not shown) within system 100.
[0043] In an embodiment, computing device 110 can be a standalone computing device, a management server, a web server, a mobile computing device, or any other electronic device or computing system capable of receiving, sending, and processing data. In another embodiment, computing device 110 can represent a server computing system utilizing multiple computers as a server system, such as in a cloud computing environment. In yet another embodiment, computing device 110 represents a computing system utilizing clustered computers and components (e.g., database server computers, application server computers) that act as a single pool of seamless resources when accessed within system 100.
[0044] System 100 optionally includes Al circuitry 112 communicatively coupled with the computing device 110. In some embodiments, the Al circuitry 112 may be incorporated into the computing device 110. In some other embodiments, the Al circuitry 112 may be external to the computing device 110. In those embodiments where the Al circuitry 112 is external to the computing device 110, the Al circuitry 112 may be directly coupled to the computing device 110 or may be coupled to the computing device 110 remotely, for example, over the network 120.
[0045] Al can be defined as the theory and development of computer systems able to perform tasks that normally require human intelligence, such as speech recognition, visual perception, decision-making, and translation between languages. The term Al is often used to describe systems that mimic cognitive functions of the human mind, such as learning and problem solving.
[0046] Machine learning (ML) is an application of Al that creates systems that have the ability to automatically learn and improve from experience. ML involves the development of computer programs that can access data and learn based on that data. ML algorithms typically build mathematical models based on sample, or training, data in order to make predictions or decisions without being explicitly programmed to do so.
[0047] Deep learning is a sub-field of ML that automates much of the feature extraction, eliminating some of the manual human intervention required and enabling the use of larger dataAtty. Docket: UA25-088PCTsets. Deep learning typically uses neural networks, which are highly interconnected entities, called nodes. Each node, or artificial neuron, connects to another and has an associated weight and threshold. A node multiplies the input data with the weight, which either amplifies or dampens that input, thereby assigning significance to inputs with regard to the task the algorithm is trying to learn. If the output of any individual node is above the specified threshold value, that node is activated, sending data to the next layer of the network. Otherwise, no data is passed along to the next layer of the network. A neural network that consists of more than three layers can be considered a deep learning algorithm or a deep neural network.
[0048] LLMs are a category of deep learning models trained on immense amounts of data, making them capable of understanding and generating natural language and other types of content to perform a wide range of tasks. LLMs are built on a type of neural network architecture called a transformer which excels at handling sequences of words and capturing patterns in text. In some embodiments, the Al circuitry 112 consists of one or more LLMs.
[0049] System 100 optionally includes a display 114. The display 114 provides a mechanism to display data to a user and may be, for example, a computer monitor. Display 114 can also function as a touchscreen, such as a display of a tablet computer. In some embodiments, the display 114 may be a VR headset.
[0050] FIG. 2 is an example platform pipeline diagram 200 for the system of FIG. 1, consistent with the present disclosure. The diagram 200 illustrates the comprehensive workflow of the disclosed ecosystem. A snapshot of the real-world video 202 of the workflow is presented on the left, capturing the performance and the explanation of the expert. Panel A 204 showcases the initial stage where the system processes a video of the expert performing a task. This component prepares the digital twin and 3D CAD models for building scenes and intelligently formulates a training and assessment plan by analyzing the video content with an LLM. Panel B 206 details the Adaptive Training System, which customizes the instruction plan based on the trainee’s experience level, dynamically and selectively adjusting assistance based on the trainees’ historical performance errors. Panel C 208 presents the Assessment System, powered by the Assessment Plan, which enables automatic evaluation of the trainee’s performance. This system incorporates a ‘human-in-the-loop’ mode, allowing expert feedback to be captured and utilized as suggestions for the smartAtty. Docket: UA25-088PCTagent, thereby informing future planning. The rightmost section of the image depicts a trainee in the virtual environment 210, undergoing training that mirrors the expert’s task shown in the real world on the real-world video 202.
[0051] FIG. 3 is a system overview 300 illustrating the smart instruction and assessment ecosystem 302 for the system of FIG. 1. The smart instruction and assessment ecosystem 302 may consist of four key modules, each interconnected through the sharing of content of serialized objects. Initially, the smart instruction and assessment module 304 generates an instruction plan 306 and an assessment plan 310. The instruction plan 306 is utilized by the adaptive training module 308 to construct the training environment, while the assessment plan 310 informs the assessment module 312, which establishes the rules for evaluating trainee performance. Following participation in assessments, trainee performance data is collected and stored within a trainee profile 314 for the trainee. This data is subsequently accessed by the trainee profile 314 and performance module 316 for performance review and analysis. Additionally, the adaptive training module 308 utilizes this performance data to adjust the training level dynamically, ensuring that the training intensity and complexity are tailored to the individual trainee’s learning progress
[0052] The first module, the smart instruction and assessment module 304, takes a video as input and generates two serialized objects as instruction plan 306 and an assessment plan 310. These objects are then used to build the training workflow and define the grading rubrics for the assessment process.
[0053] The second module, the adaptive training module 308, constructs the training scene based on the generated serialized object, i.e., the instruction plan 306, by enabling different actions for different objects according to the role of the object in the instruction plan. A hints system adds hints and restrictions to each step, ensuring the steps are conducted in correct order and a correct way. If a mistake is made, the system provides appropriate warnings to guide the trainee.
[0054] In the third module, the assessment module 312, trainees are given full freedom to perform tasks without any hints or warnings. The system monitors and detects their behavior and the sequence of actions they take. At the end of the session, the assessment module analyzes the trainee’s performance data and creates a performance report. Next, the assessment module generates grading based on the trainee’s performance using the generated assessment plan 310.Atty. Docket: UA25-088PCTAfter each assessment, the performance data is stored in the corresponding trainee profile 314 in the system. The performance data is used to automatically adjust future training sessions. The performance data helps determine the level of assistance provided and allows the system to offer options for trainees to skip steps they have already mastered and focus on those that need improvement. The data are also used for later review by trainees themselves or experts to adjust assessment rules.
[0055] The forth module, the user profiles and performance management module 316, allows experts to review the performance history of all trainees by reading data from the trainee profiles 314. It also provides a user-friendly interface where experts can adjust assessment rules easily. Feedback from experts is then fed back into the smart instruction and assessment module 304, enabling incremental improvement of the system based on past experiences.
[0056] Terms that are used to illustrate the various system components are defined below. An illustrative example use case study in a nanotechnology lab is used to illustrate the terminology in an intuitive manner.
[0057] In the disclosed system, users are divided into two categories: trainees and experts. The role of each user is determined during registration. A trainee profile, e.g., trainee profile 314 from FIG. 3, is created for every individual, which records the performance history and grading for the trainee across all tests. Experts have additional permissions, allowing them to review the performance of all trainees. This helps the experts adjust assessment plan based on overall trainee performance.
[0058] As used herein, a training task consists of multiple steps, which can typically be further broken down into sub-steps that can be replayed and validated in the disclosed system. Each step represents a smaller, focused task. For example, in the case study of a nanotechnology lab, producing a wafer is a training task. This task is completed by going through steps such as a gowning room procedure, operating the spin coater, performing photolithography, and using a decomposer. Each step can be divided into smaller actions; for instance, in the spin coater step, tasks may include placing the wafer into the spin coater and applying photoresist onto the wafer, etc.Atty. Docket: UA25-088PCT
[0059] As used herein, objects are provided by the system and may include prepared CAD models and / or digital twins created through scanning. CAD models generally support more interactive operations compared to scanned models due to their inherent nature of being clearly defined and meticulously built. To ensure these models function correctly within the disclosed system, each part of the machine must be clearly identified and labeled. For the scanned model, in some embodiments human input with labeling may be used. In some other embodiments, this setup may be enhanced with advanced scene understanding capabilities.
[0060] As used herein, actions refer to the set of operations that can be detected and monitored by the disclosed system. These actions are abstractions of real-world operations performed during training. In the illustrative example of the nanotechnology lab discussed herein, the disclosed system supports four common actions in training tasks. The system is designed to be extendable, and in other embodiments may allow for the addition of new actions, such as rotation, without impacting the existing functionality. FIG. 4 illustrates the supported actions for the example application in the nanotechnology lab for the system of FIG. 1.
[0061] In the “Move” column 402, there are two actions: “Move Wafer” and “Apply Photoresist.” The “Check” column 404 depicts a trainee checking himself in a mirror. The “Value Input” column 406 lists three distinct actions: inputting speed values into the spin coater, setting timer values, and entering values on a photolithography monitor using a virtual keyboard. The “Click” column 408 highlights the most common action — clicking a button. Lastly, the “Wear” action 410 demonstrates gloves being placed on the avatar following the action’s completion.
[0062] In an embodiment, the system may include, but is not limited to, the following four actions, the move action, the check action, the value input action, and the touch action.
[0063] As used herein, the move action involves moving an object from one place to another. The “place” can be a specific part of an object and located through the part name. For instance, the trainee might be asked to place a wafer on the holder of a spin coater. Then a search for “holder” inside the spin coater would locate the position. An indicator line may also be used to remind the trainee of the difference from the target position.Atty. Docket: UA25-088PCT
[0064] As used herein, the check is triggered when the trainee needs to gaze at or check something. For example, check the value on the monitor, or look at the mirror to make sure all the gears are worn correctly. The check action is detected as a success if the trainee focuses on the same object continuously for more than, for example, three seconds.
[0065] As used herein, the value input action is required when the trainee needs to input a value as part of the training task. The value can be in numeric, text format. In the example embodiment of the nanotechnology lab, the system supports at least two different styles of input, a panel with buttons, or a virtual keyboard, depending on the needs. In other embodiments, the system may support any other appropriate style of input as would be known to one skilled in the art.
[0066] As used herein, the touch action is the most common action, typically involving the trainee clicking a button or touching an object. Depending on the outcome of the touch action, this can be used to execute the wear action, for example, putting gloves on the avatar after the trainee touches them.
[0067] As used herein, attribution refers to the specific aspects used to measure an action. Different actions are monitored and graded based on various criteria, as shown in Table 1 in FIG.5. Table 1 categorizes various actions by their focus on specific attributes such as Tool, Value Correctness, Accuracy, and Time, illustrating how each action type focuses on different aspects of performance.
[0068] As used herein, performance refers to the recorded performance data in the trainee assessment. The performance of the trainee in all assessments is stored in the trainee profile of the trainee and may be used for further analysis.
[0069] As used herein, the instruction plan and the assessment plan are generated by the smart instruction and assessment module, which encompass the processes through which trainees are guided through tasks and how their performance is evaluated. The assessment plan includes assessment rubrics for each step, which outline the penalties for errors made by trainees and detected by the system.
[0070] The smart instruction and assessment rules generation module 304 follows a structured workflow. This module takes video or text as an input and outputs two serialized files, theAtty. Docket: UA25-088PCTinstruction plan, and the assessment plan. FIG. 6 is a flowchart diagram depicting operations for one illustrative example embodiment of the algorithm for smart instruction and assessment rule generation on the system of FIG. 1. consistent with the present disclosure. It should be appreciated that embodiments of the present disclosure provide at least for smart instruction and assessment rule generation. However, FIG. 6 provides only an illustration of one implementation and does not imply any limitations with regard to the environments in which different embodiments may be implemented. Many modifications to the depicted environment may be made by those skilled in the art without departing from the scope of the disclosure as recited by the claims.
[0071] In the illustrated example embodiment, process 600 includes receiving data in operation 602. The data received may include videos or text data. In operation 604, process 600 extracts the video frames and audio channel from the video. In operation 606, process 600 generates a transcript from the audio channel. In operation 608, process 600 sends the transcript to an LLM. In operation 610, process 600 divides the training task into a plurality of steps.
[0072] In operation 612, for each step defined in operation 610 process 600 identifies sub-steps for the step. In operation 614, for each sub-step defined in operation 612 process 600 identifies actions of the sub-step. In operation 616, process 600 specifies attributions of the action. In operation 618, process 600 adds the attributions to the corresponding part of the instruction.
[0073] In decision block 620, process 600 determines if the step is necessary. If process 600 determines that the step is necessary (“yes” branch, decision block 620), then process 600 proceeds to operation 622. If process 600 determines that the step is not necessary (“no” branch, decision block 620), then process 600 proceeds to decision block 626.
[0074] In operation 622, process 600 specifies a grading rubric for the step. In operation 624, process 600 adds the attributions to the corresponding part of the assessment.
[0075] In decision block 626, process 600 determines if there are additional steps. If process 600 determines that there are additional steps (“yes” branch, decision block 626), then process 600 returns to operation 612 to identify sub-steps for the next step. If process 600 determines that there are no additional steps (“no” branch, decision block 626), then process 600 proceeds to operation 628.Atty. Docket: UA25-088PCT
[0076] In operation 628, process 600 generates an instruction plan. In operation 630, process 600 generates an assessment plan. In operation 632, process 600 returns the instruction plan and the assessment plan.
[0077] FIG. 7 is an example smart instruction and assessment workflow 700 for the process 600 of FIG. 6. As is shown in FIG. 7, the general workflow 700 begins by taking a video as input 702, extracting the video frames and audio channel, and generating a transcript 704 from the audio (using, for example, Whisper, an OpenAI model, to perform speech to text). The transcript 704 is then fed into an LLM. Utilizing the information extraction 706 capabilities of LLMs, the system divides the training task into several steps 708, with each step corresponding to a smaller, distinct task. For each of these tasks, the specific sub-steps 710-1 through 710-n are outlined. For later analysis, the LLM is also tasked to attach the original text for each of the corresponding steps. The corresponding prompt, for example, Prompt for Information Extraction, is used for this step to generate the following parameters in the instruction plan 716.
[0078] The Step parameter is filled with the task name from a prompt, for example, Prompt for Information Extra ction(lE). A step is typically composed of several sub-steps, each containing different attributes. For any attribute that is not applicable to the current action, “N / A” may be used. The sub-step here corresponds to the step in a prompt, for example, Prompt for Information Extraction ( E). According to the definition in the prompt, each sub-step should represent an atomic action.
[0079] The instruction video is generated automatically. The start and end points of the video are determined by the original text for each step, which is marked by the LLM. The portion of the video between the time stamp when the first sentence is spoken and the time stamp when the last sentence is spoken is extracted as the instruction video for the step.
[0080] In the disclosed system, the step-level instruction video is generated automatically from the expert demonstration video by first transcribing it to obtain a time-aligned transcript, then using the LLM pipeline to clean the transcript and segment it into discrete procedural steps (and substeps) that correspond to observable actions. For each step, the system aligns the step’s text span (e.g., the first and last sentence associated with that step) back to the transcript and uses the corresponding timestamps to extract the matching video segment as the “instruction video” forAtty. Docket: UA25-088PCTthat step; these per-step clips are stored and retrieved by the training runtime as step-specific video guidance aligned with the automatically generated procedure. The digital twin (implemented as scanned / CAD-based 3D environments with semantic object labels) grounds the plan and the video guidance: object references produced by the LLM are matched to labeled scene entities, and the training system tracks the trainee’s current step by executing the instruction plan in the twin, enabling it to retrieve and present the corresponding step video clip in context.
[0081] For each sub-step, the system checks if the object involved in the action is available in the system’s provided models. If so, the system will then identify which actions from the system can be used to replay this sub-step or return “N / A” if there is no provided action that can be mapped to it. To facilitate this, the system prompts the Al model to determine the appropriate action from the supported action set, which can be framed as a relation extraction task in Natural Language Processing (NLP). In this context, once the target object of an action is identified, the action itself can be treated as a relationship between the trainee (performing the action) and the target object (being operated on).
[0082] However, relation extraction tasks remains a challenge for generative Al models. To address this, a method has been proposed that reformulates the prompt into a multiple-choice question format, which has been shown to improve Al performance. The disclosed system adopts this method. The options for the questions are generated by the system simply using the format <Researcher> (performing the action) + <action>+ <target object >(being operated on).
[0083] For the Instruction Plan 716, once the action is identified, different actions will focus on different attributes. To get this information, a one-shot example prompt is provided for each action type. For example, for the action move Prompt for Instruction Details Supplement(action: move), an example is provided that contains all the required information for the attribution and shows an expected format it should return. This approach makes sure the model can focus on corresponding attributions for different actions.
[0084] The parameters that can be extracted for each sub- step in the Instruction Plan may include, but are not limited to, action, determined by relation extraction 712; speech instruction, a brief description to guide the trainee on what to do; target object, the object that is operated on; target position, applicable only to the “move” action, indicating to where the object should beAtty. Docket: UA25-088PCTmoved; tool, the actions move and click may require the use of a specific tool, e.g. use a tweezer to move the cube; value format, specifies the format required for value input, e.g., timer, number, and text, the system provides the appropriate input component based on the required format; and parameters, if the action is “move,” this field specifies the parameters for the movement, and if the action is “value input,” this field defines the value that should be input.
[0085] With the above prompts, a serialized file may be generated, which can be interpreted by the disclosed system. The format of the serialized file is defined in the following format: Instruction Plan, e.g., instruction plan 716.
[0086] For the Assessment Plan 720, for each action’s attributes, if explicit instructions are provided, the system will also search for corresponding grading rubrics 718 mentioned by the expert in the text. If such grading information is found, it is extracted and included as part of the serialized object Assessment Plan 720.
[0087] Similar to the action selection process in the Instruction Plan 716, for assessment grading rubrics, the system provides predefined grading functions. In an embodiment, three different functions may be provided to cover common grading scenarios. The first function may be Grading function [1], Proportional deduction. For example, if an expert specifies: “The expected time is 5 seconds, and the acceptable time is 10 seconds,” the longer the time taken, the greater the penalty in points. The second function may be Grading function [2], Fixed point deduction within a specific error range. For instance, an expert might state: “The expected time is 5 seconds.” Points are deducted according to defined intervals, such as from 5-10 seconds, 10-15 seconds, and more than 15 seconds. The third function may be Grading function [3], Full point deduction. For example, in a surgical procedure, an expert might specify: “You will lose all points if accuracy is not met.”
[0088] Below are the definitions for the parameters in the Assessment Plan 720 and how the above grading functions are applied. Action is determined by relation extraction from the instructions. Tool, if a tool is explicitly required in the Instruction Plan 716, the attribute status is marked as Yes / No (binary value). For this attribute, only grading function [3] is applicable. Value Correctness is treated as a binary attribute (Yes / No). Therefore, grading function [3] is applied here as well. Time Duration, if the time duration is specified in the Instruction Plan 716, thisAtty. Docket: UA25-088PCTattribute applies to all actions. As time is a continuous value, grading functions [1], [2], or [3] can be used depending on the type of grading information provided through a prompt, for example, Prompt for Assessment Details .S' / / / ? / cmc / 7 / (Attribution:Time duration)(ADS:attri). Accuracy applies to movement actions. Similar to time duration, it is a continuous numeric value, and grading functions [1], [2], or [3] can be used, depending on the grading criteria provided. Necessity, this attribute indicates whether this step is optional or not to this training task. For this attribute, the generative Al model can be used through a prompt, for example, Prompt for Assessment Details supplement(Attribution:Necessity)(ADS:necessity) to fill in the parameter as Yes / No.
[0089] At times, grading information may be absent from the instruction video. In such cases, as required by a prompt, for example, Prompt for Assessment Details Supplement^ AttributiomTime duration)(ADS:attri), the LLM provides an estimated value based on its accumulated knowledge and marks that value as (speculate number).
[0090] For example, for a sentence, “please place the cube on the top of the table,” with a context of “Please place the cube on the top of the table. This action is expected to be finished within 30 seconds. Points deduction begins after 30 seconds. The more time you spend, the more points you will lose correspondingly.” The LLM will reply with Grading Function 1: 30 seconds, 60 seconds (speculate).
[0091] To enhance the reliability of the LLM’s estimations, the system incorporates additional information or necessary domain knowledge into the context of the prompt if such external sources are available. For instance, the initial example might be augmented to include: “Please place the cube on the top of the table. This action is expected to be finished within 30 seconds. Points deduction begins after 30 seconds. The more time you spend, the more points you will lose correspondingly. + <extemal knowledge To optimize the LLM’s performance and also avoid exceeding prompt length limitations, the disclosed system employs a strategy to extract only the most relevant information from the external source. For a short content, a simple prompt is used to identify pertinent details. For a larger documents, a pre-trained sentence transformer model, such as all-MiniLM-L6-v2, may be utilized to filter and select the most relevant sentences to be used as the external source in the context.Atty. Docket: UA25-088PCT
[0092] The nanotechnology lab case study of FIG. 3 is used here to illustrate the workflow 700. The case study will focus on the introduction to the spin coater part of the expert video. The text extracted from the video is. “Expert: We put a wafer into this machine we called spin coater. So, you put your wafer right on there. And you put your photoresist on top. And what this does is spinning this wafer around like, say its 2000 RPM and 45 seconds. And it will coat the top of the wafer evenly with photoresist. And then after you do that... ”
[0093] According to the workflow 700, the next step is information extraction by adding the experts words to a prompt of, for example, Information Extractions text part. The result is a serialized object Toy Example: step and sub-steps, which includes 4 steps in this case, place the wafer on the spin coater, apply photoresist to the wafer, set the spin speed to 2,000 RPM, and set the spin time to 45 seconds.
[0094] Then focusing on sub-step 2, apply the photoresist to wafer, a prompt is generated. Toy Example: Relation Extraction to decide that the action which should be used for this step is “Move”. Then the Prompt for Instruction Details Supplement(action move) is used to generate the instruction plan Toy Example: Instruction Plan. Then since there is no instruction specified in the instruction, the workflow 700 only needs to fill in the “necessity” using a prompt, for example, Prompt for Assessment Details SM / ?pZem (Attribution:necessity) to get the assessment plan Toy Example: Assessment Plan.
[0095] The adaptive training, e.g., Adaptive Training System from panel B 206 in FIG. 2, initiates by preparing the virtual training environment, which involves loading provided models into the scene and enabling objects corresponding function according to their role indicated in the instruction plan.
[0096] The system develops the training scene relying on either scanned models or CAD models. Therefore, for each identified sub-step, the system needs to check the availability of the target object and the required tools. If either is missing, the step is currently unsupported, and the system will note the missing objects and skip this step. Otherwise, the system assigns appropriate functionalities based on their designated roles. The appropriate functionalities based on their designated roles may include, but are not limited to, the following. First, for objects identified as tools, scripts are added to enable them to function as tools within the system. For objects associatedAtty. Docket: UA25-088PCTwith move actions, scripts are added to make them movable and their locations detectable along with some assist functions e.g. indicator line. For objects that are the targets of click actions, scripts are implemented to make them clickable and to trigger appropriate reactions after they are clicked. For objects associated with value input, the system adds a value input object as a child to the original object. For the example of the nanotechnology lab. two types of value input interfaces are provided: a virtual keyboard and button controls (see FIG. 4). It should be noted, however, that other embodiments of the system may have other numbers of value input interfaces. The choice of interface depends on the type of machine to which they are attached. For instance, smart machines equipped with monitors or touch screens are better suited for virtual keyboards, whereas older machines with limited button interfaces may benefit more from button-style inputs. Check action is controlled by a gazing system, where a reticle will appear in front of the trainee when performing an action “check” as is shown in FIG. 4. The reticle continues to be filled as long as the trainee maintains a gaze on the same object. After approximately three seconds, the system will verify whether the object the trainee is looking at is the intended target.
[0097] The positioning of objects within the scene is adjustable, allowing experts to place objects as needed. The layout is recorded and then used for later training.
[0098] After building the basic training scene, the adaptive training system includes several key components that make the training process intuitive and adaptive to the needs of trainees. The system adjusts the training environment dynamically to suit the trainee’s abilities and learning pace that was learned from trainee performance history. This personalization helps in maximizing effectiveness by addressing individual needs, which can vary widely among trainees. This is performed by the hints system, which provides various types of hints, including text and voice reminders, a timer displaying the expected time for an action, and an animated “ghost hand” to demonstrate the correct hand posture for specific tasks. Additionally, the system may include an animated arrow to guide trainees to the next task location.
[0099] To prevent trainees from skipping steps or making mistakes, the system incorporates a warning system. This system alerts trainees if they deviate from the correct sequence and reminds them to return to the current step, showing which tool should be used. The warnings are deliveredAtty. Docket: UA25-088PCTin both text and audio formats, and are head anchored, ensuring they remain in the trainee’s field of view and cannot be ignored.
[0100] To further enhance the trainee’s understanding of the overall procedure, a performance panel is provided. This panel displays the trainee’s current step and progress, allowing the trainee to see how many steps remain. The performance panel also allows the trainee to play instructional videos to help them understand the current steps better.
[0101] The hints system is adaptive and adjusts based on the trainee’s performance history. It operates at two levels, step-level and sub-step level. If a trainee has completed a step perfectly without any mistakes in a previous assessment session, the system will offer an option to skip that step in next training session. This allows trainees to bypass steps they have already mastered, making training more efficient.
[0102] The system includes partial hints for sub-steps. Each action in the system is evaluated based on different attributes, which correspond to the aspects checked during the assessment phase. In an embodiment, there may be three types of hints for each action in the hint system including text and voice instructions to provide information about the correctness of the attribute values and tool usage, a timer indicating the expected time for completing the action, and a “ghost hand” demonstrating the correct hand posture for accuracy, to help the trainee to perform an action more accurately. In other embodiments, there may be any number of types of hints for each action in the hint system.
[0103] If a step was not completed perfectly and a mistake was made in one specific aspect, the system will provide corresponding hints for that sub-step in the training. For example, if an action is completed correctly in all aspects except time, the timer will appear, but other hints are suppressed. If only the accuracy is incorrect, the ghost hand is shown, while other hints are removed. Similarly, if the issue lies with tool usage or value correctness, only the text and voice commands are displayed.
[0104] In the assessment module, e.g., the assessment system of panel C 208 in FIG. 2, the system evaluates the trainee’s performance by collecting and analyzing performance data. TheAtty. Docket: UA25-088PCTsystem detects the condition of the objects in the scene to determine which action the trainee is performing.
[0105] Once the conducted actions (sub-steps) are identified, the system then compares these steps with the instructional plan, which contains the steps in the correct order. To validate whether the steps were performed in the correct sequence, the system uses the Longest Common Subsequence (LCS) algorithm. This allows the system to identify steps that were conducted but in the wrong order. For example, if the correct order is 1, 2, 3, 4, 5, and the trainee performs them in the order 2, 3, 4, 5, 1, the validated steps would be 2, 3, 4, and 5, with step 1 marked as done but in the wrong order.
[0106] By applying the LCS algorithm to the collected data, a performance report as shown in Fig. 8 is generated. In FIG. 8, each row represents a distinct step, and each column represents a sub-step. For each step, information is divided into three components: the action and target object’s name, a screenshot of the target object, and the step’s score. The report contains grades for each step and classifies these steps into three categories, 1) Done (shown by color green), 2) Not done (shown by color red), and 3) Done but in the wrong order (shown by color red-yellow). For the steps marked as “done,” the system also records its attribution value, where accuracy is determined by the closest distance to the target object during movement actions. The time duration for each step is calculated by measuring the time difference between the start and completion of consecutive steps. For example, if sub-step 1 of step 2 is selected, the system records which tool was used for the action and the time the trainee spent on that action (see FIG. 9). Only steps marked as “done” are graded. The system uses the grading function in the assessment plan to calculate the score based on the step’s attribution value.
[0107] The User Profiles and Performance Management module, e.g., module 316 from FIG.3, serves as the starting point for the entire system. It allows users to register, select their role (trainee or expert), and choose between training mode and assessment mode.
[0108] Additionally, the module enables trainees to review their performance history, while experts are granted permission to access all trainees’ performance data, allowing experts to easily adjust the assessment plan that was previously generated by the generative Al model, i.e., the LLM.Atty. Docket: UA25-088PCT
[0109] To assist experts to perform their work, an interface is designed (FIG. 10) which shows all steps and their sub-steps as well as their corresponding points. In FIG. 10, “10 / 10” represents the grading for the current selected trainee and average score of all trainees. With this information, experts can select sub-steps they wish to modify. For example, if they select sub-step 2 of step 2, experts can see all the attributes relevant to sub-step 2 of step 2 (see FIG. 11). Upon selecting an attribute, the corresponding grading function is displayed and can be adjusted.
[0110] As described for the Smart Instruction and Assessment Rules Generation module, each grading function includes several adjustable parameters. For example, with a linear grading function, the expert can modify the maximum score deduction and the acceptable error range. To assist in this process, experts can also view individual trainee data in a statistic panel (see FIG. 12) by selecting a specific trainee through a drop down menu “Data Selection” provided in the grading function interface.
[0111] FIG. 12 shows a panel that consists of two rows. The first row 1202 displays three charts depicting the selected trainee’s attribute scores, sub-step scores, and overall assessment scores across multiple attempts (x-axis) with corresponding points earned (y-axis). The second row 1204 presents the average scores for all trainees, displaying the distribution of points within intervals (x-axis) and the corresponding points earned (y-axis).
[0112] When an expert adjusts the grading parameters, the modified grading function is immediately updated in the assessment plan, and the data visualization changes in real time. This allows the expert to instantly observe how their adjustments impact overall performance.
[0113] The system allows for a Human In the Loop. The adjustment made by experts are fed back to the Smart Instruction and Assessment Generation module as an external source in a text format. The text can be easily generated by combining parameter information. For example, if an expert adjusts an attribution time duration that uses grading function [1], the generated information may be, “In the step <step name > of training task <Training task name >. The sub-step is <sub-step>. The grading function <number > is used, where the expected time is <expected time > and largest acceptable error range is <maximum acceptable error rangeAtty. Docket: UA25-088PCT
[0114] A pilot study was conducted as a preliminary step in the development of the system and proof of the concept. The experiments were designed to assess the effectiveness and intuitiveness of the disclosed system. The pilot involved four users who completed training and assessment tasks across three trials. This allowed for necessary adjustments to parameters, such as interaction of different action types, attribution of each action, hint type and interaction way and better prepare for the main study. The study was designed to answer the following research questions, (1) how does the application address current challenges in authoring effective manuals in nanotechnology lab; (2) how natural and intuitive are trainee interactions within the VR training system; and (3) how effective is the system’s mechanism in real-time skill adaptation and trainee guidance.
[0115] The experiment involved four participants, one of whom was an expert in nanotechnology labs, while the other three were novices (2 males and 1 female, age: 20-30) in the nanotechnology sector. These participants were chosen to provide a balanced perspective on the usability and educational impact of the system. Equipment used in the tests included a commercial VR headset, running the VR training application developed in the Unity3D and MRTK packages on a laptop computer. Model GPT-4 from OpenAI was used for the module Smart Instruction and Assessment Generation. Participants performed a series of tasks that were outlined in the instruction plan. Interviews were conducted to allow for subjective measures of the user experience. The subjective measure of the user experience included, (1) Interaction of Different Action Types, whether interactions with the system felt natural and intuitive was assessed; (2) Measurement on Attribution of Each Action, whether it made sense to grade actions based on these attributions was evaluated (Expert only); (3) Hints Function, the relevance and clarity of the hints provided was assessed; and (4) Other Suggestions, participants were invited to offer any additional feedback or suggestions.
[0116] The experimental procedure is as follows. First, orientation, where participants were first acquainted with the VR equipment and the experiment’s objectives. Next, Task Execution, where participants carried out wafer production tasks within the VR environment across training and assessment three times, and Feedback Collection where, after completing the tasks, participants were interviewed for feedback on their experience.Atty. Docket: UA25-088PCT
[0117] Participants provided qualitative feedback through semi-structured interviews that focused on their subjective experiences with the VR system. This feedback was instrumental in identifying features of the VR training that either facilitated or hindered their learning. Below are the interview responses from the four participants.
[0118] Expert: The expert did not participate in the training exercises but reviewed the steps provided in the training scene. He confirmed that the steps and the attributions assigned to each action for measurement were logical and appropriate.
[0119] Novice A: Novice A was generally satisfied with the system interaction. However, he noted that the reticle for the check action was too small to be easily noticed. While the hints for each step were overall helpful, he reported that sometimes the placement of text was obstructed by machinery, making it difficult to follow.
[0120] Novice B: Novice B experienced difficulty with value input due to the small size of the buttons. He suggested to make the button larger. Additionally, he suggested the training mode should provide an overview of all tasks at the beginning. Despite this, he expressed satisfaction with the overall hints system.
[0121] Novice C: Novice C suggested that interactions with the model, particularly during clicking actions, should provide more feedback to enhance intuitiveness, e.g. the sounds of click. Without such feedback, the lack of sensory cues diminishes the immersive quality of the experience.
[0122] The research questions section of the interviews presents initial results relevant to the three research questions. The expert confirmed that the division of steps within the training scenarios is both appropriate and logical. This endorsement suggests that the disclosed system’s framework aligns well with industry standards and expert expectations, providing a solid foundation for immersive training.
[0123] Feedback indicates that the system effectively supports task completion, especially for novices. The interactive, immersive features were well-received, though some users noted a need for better hand-tracking and interaction feedback.Atty. Docket: UA25-088PCT
[0124] According to the observation on trainee performance and their feedback, trainees are able to leam tasks quickly using the disclosed system. This indicates that the system not only supports effective learning but does so in an efficient manner, potentially reducing the time needed to train novices in complex nanotechnology tasks.
[0125] In an embodiment, the user interactions are refined with action values and checks to enhance the overall user experience, including providing suitable corresponding sound effects when an action is performed to make interactions more intuitive and easier to follow. Additionally, the placement of hints has been adjusted to ensure they are accurately positioned and do not obstruct user interactions, thus improving the clarity and effectiveness of guidance provided during training. In an embodiment, an overview is introduced at the start of each training session to familiarize users with the system’s features and functionalities. In an embodiment, a pretraining scene is included that allows users to become acquainted with all actions and attributes mentioned in the training, ensuring they are comfortable and well-prepared before starting the actual training tasks.
[0126] During development, the GPT-4 model, a highly capable large language model (LLM) available at the time of the development, was used. GPT-4 has been trained on a broad range of tasks and can adapt to domain- specific requirements through in-context learning (i.e.. by following instructions and examples provided in the prompt). To assess the model’s output predictability, token-level perplexity was computed from the model’s log-probabilities over its generated responses. Perplexity is a standard NLP metric that quantifies how well a probabilistic language model predicts a sequence of tokens; lower perplexity indicates that the model assigns higher probability to the observed text. Across the prompt set, the estimated perplexity values were close to 1, suggesting that — under this measurement — the model assigned high likelihood to its own generated outputs. For instance, the Information Extraction (IE) prompt yielded a perplexity of approximately 1.04, and the Relation Extraction (RE) prompt yielded approximately 1.001. These values are reported as indicators of output predictability (and they should not be interpreted as a direct measure of task correctness).
[0127] FIG. 13 is a block diagram depicting components of one example of the computing device 110 suitable for Al-driven training and assessment in extended reality, within the system ofAtty. Docket: UA25-088PCTFIG. 1, consistent with the present disclosure. FIG. 13 displays the computing device or computer 1300, one or more processor(s) 1304 (including one or more controllers or computer processors), a communications fabric 1302, a memory 1306 including, a random-access memory (RAM) 1316 and a cache 1318, a persistent storage 1308, a communications unit 1312, I / O interfaces 1314, a display 1322, and external devices 1320. It should be appreciated that FIG. 13 provides only an illustration of one embodiment and does not imply any limitations with regard to the environments in which different embodiments may be implemented. Many modifications to the depicted environment may be made.
[0128] As depicted, the computer 1300 operates over the communications fabric 1302, which provides communications between the computer processor(s) 1304, memory 1306, persistent storage 1308, communications unit 1312, and input / output (VO) interface(s) 1314. The communications fabric 1302 may be implemented with an architecture suitable for passing data or control information between the processors 1304 (e.g.. microprocessors, communications processors, and network processors), the memory 1306, the external devices 1320, and any other hardware components within a system. For example, the communications fabric 1302 may be implemented with one or more buses.
[0129] The memory 1306 and persistent storage 1308 are computer readable storage media. In the depicted embodiment, the memory 1306 comprises a RAM 1316 and a cache 1318. In general, the memory 1306 can include any suitable volatile or non-volatile computer readable storage media. Cache 1318 is a fast memory that enhances the performance of processor(s) 1304 by holding recently accessed data, and near recently accessed data, from RAM 1316.
[0130] Program instructions for Al-driven training and assessment in extended reality may be stored in the persistent storage 1308, or more generally, any non-transitory computer readable storage media, for execution by one or more of the respective computer processors 1304 via one or more memories of the memory 1306. The persistent storage 1308 may be a magnetic hard disk drive, a solid-state disk drive, a semiconductor storage device, flash memory, read only memory (ROM), electronically erasable programmable read-only memory (EEPROM), or any other computer readable storage media that is capable of storing program instruction or digital information.Atty. Docket: UA25-088PCT
[0131] The media used by persistent storage 1308 may also be removable. For example, a removable hard drive may be used for persistent storage 1308. Other examples include optical and magnetic disks, thumb drives, and smart cards that are inserted into a drive for transfer onto another computer readable storage medium that is also part of persistent storage 1308.
[0132] The communications unit 1312, in these examples, provides for communications with other data processing systems or devices. In these examples, the communications unit 1312 includes one or more network interface cards. The communications unit 1312 may provide communications through the use of either or both physical and wireless communications links. In the context of some embodiments of the present disclosure, the source of the various input data may be physically remote to the computer 1300 such that the input data may be received, and the output similarly transmitted via the communications unit 1312.
[0133] The I / O interface(s) 1314 allows for input and output of data with other devices that may be connected to computer 1300. For example, the I / O interface(s) 1314 may provide a connection to external device(s) 1320 such as a keyboard, a keypad, a touch screen, a microphone, a digital camera, and / or some other suitable input device. External device(s) 1320 can also include portable computer readable storage media such as, for example, thumb drives, portable optical or magnetic disks, and memory cards. Software and data used to practice embodiments of the present disclosure can be stored on such portable computer readable storage media and can be loaded onto persistent storage 1308 via the I / O interface(s) 1314.
[0134] I / O interface(s) 1314 may also connect to a display 1322. Display 1322 provides a mechanism to display data to a user and may be, for example, a computer monitor. Display 1322 can also function as a touchscreen, such as a display of a tablet computer.
[0135] The disclosed system introduces a novel approach to immersive training by leveraging the capabilities of LLMs for automatic content generation within adaptive training systems. The system takes advantage of the readily accessible input from videos and CAD models (or scanned models), powered by LLM-driven action recognition and assessment criteria generation. This integration significantly streamlines the authoring process. Through a detailed case study in a nanotechnology lab, the disclosed system’s ability to enhance learning outcomes by providing personalized training solutions that dynamically adapt to the trainee’s performance wasAtty. Docket: UA25-088PCTdemonstrated. The integration of expert feedback and a human-in-the-loop system further ensures the accuracy and relevance of the training content and assessments.
[0136] According to one aspect of the disclosure there is thus provided a system for Al-driven training and assessment in extended reality. The system includes a computing device; and Al circuitry, where the Al circuitry includes a Large Language Model (LLM). The computing device is configured to: receive video input data; prepare the digital twin and one or more three-dimensional (3D) models from the video input data; generate an instruction plan from the video input data; generate an assessment plan from the video input data; extract an audio channel and one or more video frames from the video input data; generate a transcript from the audio channel; extract a plurality of steps from the transcript using the LLM; extract one or more sub-steps from each of the plurality of steps; generate an instruction video for each step and each sub-step using the digital twin and the one or more 3D models; perform adaptive training of a trainee based on the instruction plan; and assess a performance of the trainee based on the assessment plan.
[0137] According to another aspect of the disclosure, there is thus provided a method of AI-driven training and assessment in extended reality. The method includes receiving video input data; preparing the digital twin and one or more three-dimensional (3D) models from the video input data; generating an instruction plan from the video input data; generating an assessment plan from the video input data; extracting an audio channel and one or more video frames from the video input data; generating a transcript from the audio channel; extracting a plurality of steps from the transcript using a Large Language Model (LLM); extracting one or more sub-steps from each of the plurality of steps; generating an instruction video for each step and each sub-step using the digital twin and the one or more 3D models; performing adaptive training of a trainee based on the instruction plan; and assessing a performance of the trainee based on the assessment plan.
[0138] According to yet another aspect of the disclosure, there is thus provided a non-transitory storage device that includes machine-readable instructions that, when executed by one or more processors, cause the one or more processors to perform operations, including receive video input data; prepare a digital twin and one or more three-dimensional (3D) models from the video input data; generate an instruction plan from the video input data; generate an assessment plan from the video input data; extract an audio channel and one or more video frames from the video input data;Atty. Docket: UA25-088PCTgenerate a transcript from the audio channel; extract a plurality of steps from the transcript using a Large Language Model (LLM); extract one or more sub-steps from each of the plurality of steps; generate an instruction video for each step and each sub-step using the digital twin and the one or more 3D models; perform adaptive training of a trainee based on the instruction plan; and assess a performance of the trainee based on the assessment plan.
[0139] Although the methods and systems have been described relative to a specific embodiment thereof, they are not so limited. Obviously, many modifications and variations may become apparent in light of the above teachings. Many additional changes in the details, materials, and arrangement of parts, herein described and illustrated, may be made by those skilled in the art. Also, it may be appreciated that the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting as such may be understood by one of skill in the art. Throughout the present disclosure, like reference characters may indicate like structure throughout the several views, and such structure need not be separately discussed. Furthermore, any particular feature(s) of a particular exemplary embodiment may be equally applied to any other exemplary embodiment(s) of this disclosure as suitable. In other words, features between the various exemplary embodiments described herein are interchangeable, and not exclusive.
[0140] As used in this application and in the claims, a list of items joined by the term “and / or” can mean any combination of the listed items. For example, the phrase “A, B and / or C” can mean A; B; C; A and B; A and C; B and C; or A, B and C. As used in this application and in the claims, a list of items joined by the term “at least one of’ can mean any combination of the listed terms. For example, the phrases “at least one of A, B or C” can mean A; B; C; A and B; A and C; B and C; or A, B and C.
[0141] “Circuitry,” as used in any embodiment herein, may comprise, for example, singly or in any combination, hardwired circuitry, programmable circuitry such as processors comprising one or more individual instruction processing cores, state machine circuitry, and / or firmware that stores instructions executed by programmable circuitry and / or future computing circuitry including, for example, massive parallelism, analog or quantum computing, hardware embodiments of accelerators such as neural net processors and non- silicon implementations of the above. The circuitry may, collectively or individually, be embodied as circuitry that forms part ofAtty. Docket: UA25-088PCTa larger system, for example, an integrated circuit (IC), system on-chip (SoC), application-specific integrated circuit (ASIC), programmable logic devices (PLD), digital signal processors (DSP), field programmable gate array (FPGA), logic gates, registers, semiconductor device, chips, microchips, chip sets, etc.
[0142] The term “coupled" as used herein refers to any connection, coupling, link, or the like by which signals carried by one system element are imparted to the "coupled" element. Such “coupled" devices, or signals and devices, are not necessarily directly connected to one another and may be separated by intermediate components or devices that may manipulate or modify such signals.
[0143] The programs described herein are identified based upon the application for which they are implemented in a specific embodiment of the disclosure. However, it should be appreciated that any particular program nomenclature herein is used merely for convenience, and thus the disclosure should not be limited to use solely in any specific application identified and / or implied by such nomenclature.
[0144] The present disclosure may be a system and / or a method. The system may include one or more non-transitory computer readable storage media having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.
[0145] The one or more non-transitory computer readable storage media can be any tangible device that can retain and store instructions for use by an instruction execution device. The one or more non-transitory computer readable storage media may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-transitory computer readable storage media, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.Atty. Docket: UA25-088PCT
[0146] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from one or more non-transitory computer readable storage media or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in one or more non-transitory computer readable storage media within the respective computing / processing device.
[0147] The computer readable program instructions may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a LAN or a WAN, or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, Field-Programmable Gate Arrays (FPGA), or other Programmable Logic Devices (PLD) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0148] It will be appreciated by those skilled in the art that any block diagrams herein represent conceptual views of illustrative circuitry embodying the principles of the disclosure. Similarly, it will be appreciated that any block diagrams, flow charts, flow diagrams, state transition diagrams, pseudocode, and the like represent various processes which may be substantially represented in computer readable medium and so executed by a computer or processor, whether or not such computer or processor is explicitly shown. Software modules, or simply modules which are implied to be software, may be represented herein as any combination of flowchart elements or other elements indicating performance of process steps and / or textual description. Such modules may be executed by hardware that is expressly or implicitly shown.Atty. Docket: UA25-088PCT
[0149] The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems and methods according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, a segment, or a portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the Figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions.
Claims
Atty. Docket: UA25-088PCTCLAIMSWhat is claimed is:
1. A system for artificial intelligence (Al)-driven training and assessment in extended reality, the system comprising:a computing device; andartificial intelligence (Al) circuitry, wherein the Al circuitry includes a Large Language Model (LLM);wherein the computing device configured to:receive video input data;prepare the digital twin and one or more three-dimensional (3D) models from the video input data;generate an instruction plan from the video input data;generate an assessment plan from the video input data;extract an audio channel and one or more video frames from the video input data; generate a transcript from the audio channel;extract a plurality of steps from the transcript using the LLM;extract one or more sub-steps from each of the plurality of steps;generate an instruction video for each step and each sub-step using the digital twin and the one or more 3D models;perform adaptive training of a trainee based on the instruction plan; and assess a performance of the trainee based on the assessment plan.
2. The system of claim 1, wherein generate the instruction plan from the video input data further comprises:for each sub-step:identify one or more actions of the sub-step;specify one or more attributions for each of the actions; andadd the one or more attributions of each of the actions to a corresponding part of the instruction plan.Atty. Docket: UA25-088PCT3. The system of claim 1, wherein generate the assessment plan from the video input data further comprises:for each sub-step:identify one or more actions of the sub-step;determine whether the sub-step is necessary;responsive to determining that the sub-step is necessary:specify a grading rubric for the sub- step;specify one or more attributions for each of the actions; andadd the one or more attributions for each of the actions to a corresponding part of the assessment.
4. The system of claim 1, wherein generate the transcript from the audio channel further comprises:generate the transcript using the LLM in the Al circuitry.
5. The system of claim 1, wherein perform the adaptive training of the trainee based on the instruction plan further comprises:construct a training scene from the instruction plan by enabling one or more different actions for each different object in the instruction plan according to a role of the different object in the instruction plan;add one or more hints and one or more restrictions to each of the plurality of steps in the instruction plan; andresponsive to detecting that the trainee made a mistake, provide one or more appropriate warnings to guide the trainee.
6. The system of claim 1, wherein assess the performance of the trainee based on the assessment plan further comprises:detect a behavior and a sequence of one or more actions taken by the trainee; responsive to the trainee completing a session:analyze the performance of the trainee based on the assessment plan to create a performance report for the trainee;Atty. Docket: UA25-088PCTgenerate a grading based on the performance of the trainee;store performance data in a trainee profile corresponding to the trainee; and adjust future training sessions based on the performance data.
7. The system of claim 6, wherein a level of assistance provided by the system is determined from the performance data.
8. The system of claim 6, wherein the trainee is offered options to skip any steps that the trainee has already mastered and focus on the steps that need improvement based on the performance data.
9. A method for artificial intelligence (Al)-driven training and assessment in extended reality, the method comprising:receiving video input data;preparing the digital twin and one or more three-dimensional (3D) models from the video input data;generating an instruction plan from the video input data;generating an assessment plan from the video input data;extracting an audio channel and one or more video frames from the video input data; generating a transcript from the audio channel;extracting a plurality of steps from the transcript using a Large Language Model (LLM); extracting one or more sub-steps from each of the plurality of steps;generating an instruction video for each step and each sub-step using the digital twin and the one or more 3D models;performing adaptive training of a trainee based on the instruction plan; andassessing a performance of the trainee based on the assessment plan.
10. The method of claim 9, wherein generating the instruction plan from the video input data further comprises:for each sub-step:identifying one or more actions of the sub-step;Atty. Docket: UA25-088PCTspecifying one or more attributions for each of the actions; andadding the one or more attributions of each of the actions to a corresponding part of the instruction plan.
11. The method of claim 9, wherein generating the assessment plan from the video input data further comprises:for each sub-step:identifying one or more actions of the sub-step;determining whether the sub-step is necessary;responsive to determining that the sub-step is necessary:specifying a grading rubric for the sub-step;specifying one or more attributions for each of the actions; and adding the one or more attributions for each of the actions to a corresponding part of the assessment plan.
12. The method of claim 9, wherein generating the transcript from the audio channel further comprises:generating the transcript using the LLM in an artificial intelligence (Al) circuitry.
13. The method of claim 9, wherein performing the adaptive training of the trainee based on the instruction plan further comprises:constructing a training scene from the instruction plan by enabling one or more different actions for each different object in the instruction plan according to a role of the different object in the instruction plan;adding one or more hints and one or more restrictions to each of the plurality of steps in the instruction plan; andresponsive to detecting that the trainee made a mistake, providing one or more appropriate warnings to guide the trainee.
14. The method of claim 9, wherein assessing the performance of the trainee based on the assessment plan further comprises:Atty. Docket: UA25-088PCTdetecting a behavior and a sequence of one or more actions taken by the trainee; responsive to the trainee completing a session:analyzing the performance of the trainee based on the assessment plan to create a performance report for the trainee;generating a grading based on the performance of the trainee;storing performance data in a trainee profile corresponding to the trainee; and adjusting future training sessions based on the performance data.
15. A non-transitory storage device that includes machine-readable instructions that, when executed by one or more processors, cause the one or more processors to perform operations, comprising:receive video input data;prepare a digital twin and one or more three-dimensional (3D) models from the video input data;generate an instruction plan from the video input data;generate an assessment plan from the video input data;extract an audio channel and one or more video frames from the video input data; generate a transcript from the audio channel;extract a plurality of steps from the transcript using a Large Language Model (LLM); extract one or more sub-steps from each of the plurality of steps;generate an instruction video for each step and each sub-step using the digital twin and the one or more 3D models;perform adaptive training of a trainee based on the instruction plan; andassess a performance of the trainee based on the assessment plan.
16. The non-transitory storage device of claim 15, wherein the machine-readable instructions that, when executed by one or more processors, cause the one or more processors to perform the operations to generate the instruction plan from the video input data further comprises:for each sub-step:identify one or more actions of the sub-step;specify one or more attributions for each of the actions; andAtty. Docket: UA25-088PCTadd the one or more attributions of each of the actions to a corresponding part of the instruction plan.
17. The non-transitory storage device of claim 15, wherein the machine-readable instructions that, when executed by one or more processors, cause the one or more processors to perform the operations to generate the assessment plan from the video input data further comprises:for each sub-step:identify one or more actions of the sub-step;determine whether the sub-step is necessary;responsive to determining that the sub-step is necessary:specify a grading rubric for the sub- step;specify one or more attributions for each of the actions; andadd the one or more attributions for each of the actions to a corresponding part of the assessment plan.
18. The non-transitory storage device of claim 15, wherein the machine-readable instructions that, when executed by one or more processors, cause the one or more processors to perform the operations to generate the transcript from the audio channel further comprises:generate the transcript using the LLM in an artificial intelligence (Al) circuitry.
19. The non-transitory storage device of claim 15, wherein the machine-readable instructions that, when executed by one or more processors, cause the one or more processors to perform the operations to perform the adaptive training of the trainee based on the instruction plan further comprises:construct a training scene from the instruction plan by enabling one or more different actions for each different object in the instruction plan according to a role of the different object in the instruction plan;add one or more hints and one or more restrictions to each of the plurality of steps in the instruction plan; andresponsive to detecting that the trainee made a mistake, provide one or more appropriate warnings to guide the trainee.Atty. Docket: UA25-088PCT20. The non-transitory storage device of claim 15, wherein the machine-readable instructions that, when executed by one or more processors, cause the one or more processors to perform the operations to assess the performance of the trainee based on the assessment plan further comprises:detect a behavior and a sequence of one or more actions taken by the trainee; responsive to the trainee completing a session:analyze the performance of the trainee based on the assessment plan to create a performance report for the trainee;generate a grading based on the performance of the trainee;store performance data in a trainee profile corresponding to the trainee; and adjust future training sessions based on the performance data.