Automated documentation generation from multimodal recordings of processes and operations

An AI-assisted system using multimodal data streams and modular AI roles automatically generates high-quality documentation, addressing expertise gaps and resource challenges in task documentation, enhancing learning transfer and evaluation.

WO2025217574A1PCT designated stage Publication Date: 2025-10-16NAKAMIR INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/024357
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-12
Filing Date
2025-04-11
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

Existing methods for documenting tasks and processes face challenges such as lower-quality documentation due to expertise gaps, translation errors, and resource-intensive face-to-face instruction, especially in augmented reality and virtual reality environments, where manual authoring of steps and content creation are prevalent, and existing AI-based solutions struggle with capturing nuances and integrating diverse data streams.

Method used

An AI-assisted system that automatically generates documentation from multimodal recordings, using multiple heterogeneous data streams including audio, video, hand tracking, eye tracking, and IoT sensors, with modular AI roles to learn and document processes, reducing the need for manual assistance and enhancing context-awareness.

Benefits of technology

The system produces comprehensive documentation in various formats, including static and interactive content, while reducing manpower requirements and improving learning transfer by automatically evaluating student performance, thus overcoming limitations of existing AI-based procedures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025024357_16102025_PF_FP_ABST
    Figure US2025024357_16102025_PF_FP_ABST
Patent Text Reader

Abstract

A method for analysis and summarization of task includes recording two distinct data streams with time stamps representing the physical actions of a user and an environment while the user performs the task; processing the first and second data streams with first and second AI models [128, 130] to produce first and second higher-level representations [132, 134] of the data streams, where the processing of the second data stream with the second AI model uses the first higher-level representation of the first data stream as input; processing [136] the first higher-level representation of the first data stream and the second higher-level representation of the second data stream to produce a summarization [110] of the task; and generating an output comprising the summarization of the task.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] AUTOMATED DOCUMENTATION GENERATION FROM MULTIMODAL

[0002] RECORDINGS OF PROCESSES AND OPERATIONS

[0003] FIELD OF THE INVENTION

[0004] The present invention relates generally to methods for automatically documenting tasks performed by a user in an environment.

[0005] BACKGROUND OF THE INVENTION

[0006] Existing methods for documenting tasks and processes vary significantly depending on the field and intended audience, ranging from simple instruction manuals to complex technical documentation. Creating comprehensive and effective documentation, however, presents numerous challenges. Often, the individuals with the expertise to perform a process may not possess the necessary skills in writing, figure creation, image editing, and education required to produce high-quality documentation. This can lead to lower-quality documentation that may be difficult for novices to understand due to jargon or missing fundamental knowledge.

[0007] When documentation is delegated to individuals other than the process experts, a disconnect in knowledge can introduce "translation errors" and necessitate time-consuming iterations to ensure accuracy. Furthermore, even with perfect documentation, the challenge of "learning transfer" remains a significant concern across educational fields. Understanding a process at a high level from documentation does not guarantee the development of low- level practical skills, often requiring resource-intensive face-to-face instruction and practice. The creation of more complex learning materials like online courses and simulations further compounds these logistical issues.

[0008] In the realm of augmented reality (AR) and virtual reality (VR), while offering more immersive learning experiences, many current solutions still rely heavily on manual authoring of steps and content. Even with tools designed to aid in AR tutorial creation, significant manual effort is often required to define process steps, annotate 3D models, and link digital content to the real world. The competitive landscape in AR and Al for business applications reveals that while various solutions exist for enhancing operational efficiency and providing AR visualization, a common limitation is the dependence on manual processes for content creation, documentation, and updates. The general idea of automatically generating documentation has been explored, often leveraging video and basic Al techniques, but these methods still require significant manual assistance and face challenges in capturing the nuances of complex tasks and may not effectively integrate diverse data streams to provide comprehensive and context-aware documentation. Quality assurance of automatically generated documentation also remains a significant hurdle.

[0009] SUMMARY OF THE INVENTION

[0010] Herein is disclosed an Al-assisted automated system for generating documentation from multimodal recordings of processes and operations, offering significant advantages over existing techniques that require manual assistance in documentation creation. A key innovation is its ability to automatically learn and document a process by observing a professional's actions within an environment using multiple heterogeneous data streams. These streams can include audio, video (potentially with depth information), hand tracking, eye tracking, head tracking, camera tracking, virtual annotations, and even loT sensor data.

[0011] A further key feature is a technique for Al-processing heterogeneous data streams to produce higher-level representations, where the higher-level representation of one data stream is used in generating the higher-level representation for another data stream, thereby dramatically reducing computational complexity. For instance, hand tracking or gaze location data can focus the Al's analysis on specific objects or areas of interest within the video stream, significantly limiting the solution space and simplifying subsequent Al analysis. This multimodal fusion allows the system to pinpoint the times, locations, and objects a user interacts with, a capability that addresses the difficulty prior Al-based procedures face when confronted with unknown situations or multiple similar-looking objects. Unlike existing AR solutions that require manual authoring of steps and linking of digital content, this invention automatically converts recordings into comprehensive documentation. This documentation can take various forms, including static documents (like PDFs) and interactive content (like online courses), and can include features such as video playback, 3D model viewing, and interactive Al and Augmented Reality assistants. Another feature is the use of multiple interacting Al processing modules (e.g., recording assistant, documentation Al, localization Al, editing assistant, learning / evaluation assistant) that collaborate by learning and sharing context to perform different tasks in the documentation pipeline. This compartmentalized approach enhances efficiency and addresses the limitations of general-purpose Al models when handling complex and diverse tasks. Moreover, the system facilitates automated evaluation of a student's performance by identifying the goals and observable criteria for success within each step, enabling the generation of tools like checklists and the training of an Al learning assistant to provide feedback and assessment. This focus on automatic generation from diverse data streams and intelligent Al collaboration represents a key advancement over the state of the art in process documentation and training.

[0012] In one implementation, the invention provides a method for automatically documenting or summarizing a task and making documentation such as instruction manuals, simulations, online courses. Al is used to learn the process itself and generate documentation of the task, thus reducing the manpower required for the documentation to be sufficient and complete.

[0013] A system is provided that uses Al and multimodal summarization techniques to create both interactive (e.g., online course-like) and static (e.g., PDF) documentation of some process that is recorded with multiple heterogeneous data streams (i.e., some combination of audio, video, depth, hand-tracking, camera tracking).

[0014] The multiple data streams are processed by sharing context and semantic understanding of the process between streams. The pipeline is modular, thus it can work if the process is recorded with a smartphone or a Microsoft HoloLens, by finding as much semantic information and context as is feasible given whatever data streams are available and whatever information the Al is capable of discovering or extrapolating.

[0015] As a result, instead of requiring a team to create sufficient documentation for a process, the professional simply records the process as they perform it, using some device (e.g., a phone, AR headset such as Microsoft HoloLens), and the recording is automatically converted to documentation with our pipeline, with options for both static display (e.g., a PDF, Word, PowerPoint similar to traditional methods) and interactive display (e.g., as a 3D digital twin in Augmented Reality, an editing interface, playing videos, automated evaluation of a student).

[0016] The main use cases that are seen in the state-of-the-art focus on content creation. However, our pipeline avoids any need for content creation to be done manually at all, while still providing optional editing features for optional human intervention. Thus, a key advantage is the automatic nature of the pipeline that receives some number of data streams and outputs complete documentation automatically, which is supported by a set of novel Al interaction methods and natural language processing, sensor fusion, and computer vision.

[0017] Another key feature is that the data streams need not be synchronous. For example, in addition to a recording of the professional completing the process, we may also pass to the Al some seed context, i.e., the instruction manual for specific parts of the machine. Although a completely asynchronous use case degrades the results into what is practically the same as GPT / NLP-based Al’s summarization features, our pipeline has the benefit of more context-awareness and semantic understanding of these data streams. For example, current large language models (LLMs) and multi-modal (MMLLMs) would not know what to do with depth information, but because our pipeline includes features for reconstruction and object segmentation, we can still infer a substantial amount of context from that data. Another example would be prescans, e.g., 3D CAD models of the machinery — these could theoretically be provided to GPT somehow to aid its recognition of steps in the video recording, but as of yet, MMLLMs methods cannot understand them. While a user can feed a large language model multiple sets of instruction manuals to further complete its understanding thus allowing it to provide more context-relevant answers, the state-of-the-art is still rather limited to text, and to some extent, images. Thus, another key advantage of the present invention is the wide range of types of data streams that our pipeline can make sense of and blend together into a context. In the case of state-of-the-art Al chatbots and large language models, context tends to be a set of variables that includes primitives such as strings, numbers, and Boolean variables, and when an answer is requested, it searches this context as necessary. Our context augments this with a 3D scene, process (e.g., state machine or pipeline), per-step evaluation criteria, and multiple other components discussed in detail later. From this context, we gain more semantic understanding, i.e., our Al can infer "this step operates on these parts," "the business end of this tool is on that side, and the tool is held from this side," which is used for an Al in the role of a learning assistant that the human student may ask questions. The technical implementation of this context as it is provided to the Al involves injecting context indicators into the tasks we provide to various roles of Al such that the textual prompt itself, as well as other inputs such as images, can gain a complete understanding of the process, operation area, etc. without needing to create entirely different Al models for specific types of data streams.

[0018] A significant upgrade over the state-of-the-art for 3D documentation that we provide is not only the use of Al, but the use of multiple Al in different roles, operating at different points of the teaching / documenting / learning pipeline. They share context-awareness, semantic understanding, etc., and essentially operate as a faculty or Al support team, better compartmentalizing the tasks required at different stages. This is not only sensible at a high- level, but in terms of the known technical limitations of Al such as ChatGPT, which function more consistently when their instructions are better divided and better defined, this is more efficient and effective. For example, there are distinct Al for:

[0019] • Recording assistant: helps the person recording know what to record, when to expand on the documentation process itself, when what they are recording will not be sufficient for a subsequent viewer to understand • Documentation Al: after the data streams are processed into higher-level representations (i.e., depth->segmented 3D models, audio->speech transcription), the documentation Al creates the document itself, creating an understanding of the process, finding keyframes, organizing the text, and finding any extra relevant information given any seed context (i.e., instruction manuals for machine parts)

[0020] • Localization Al: translates and localizes the document such that it is appropriate to the field (e.g., medical process vs CNC machine vernacular) and in a target language

[0021] • Editing Assistant: if the user wants to edit the resulting document manually, the editing assistant provides suggestions, answers questions about the process (i.e., in a typical documentation process, the editor is not the professional who knows the process), etc.

[0022] • Learning / Evaluation Assistant: similar to the Editing Assistant, this is an Al that answers questions that a student may have about the process, i.e., the student may ask "where does this go", "why is he doing this." In a situation where the student is replicating the process using an AR headset (i.e., NARA), they may also ask this assistant questions like "what do I do next", "was that accurate enough." In this way, the Assistant can also evaluate the user automatically, i.e., grade them, check items on a checklist, etc.

[0023] The automatic evaluation of a student's performance of a physical task is another significant improvement over the state-of-the-art. With many documented processes, the student’s only way of knowing if they are improving or doing something correctly is by comparing their results to the teacher themselves, which is biased and inaccurate. Additionally, they may not even know the criteria for correctness. In finding the context of the process, we include information about what is the actual goal of a step, what does the operator do to accomplish it, and what information from the recording or seed context actually tells a viewer that the goal was reached (i.e., what are the observable criteria for success). This allows us to automatically generate completion tools like checklists, as well as train the Learning / Evaluation Assistant to evaluate a student observed replicating the process. Finally, the output of the automated documentation process (the Documentation Al) surpasses many other methods of encapsulating such a document. It is structured somewhat similarly to a zipped website directory, while maintaining portability and easy editing tools. It has the basic functionality of simple document-editing platforms, but additionally provides widgets for video and audio playback, 3D viewing, and tools to interact with the various Al assistants who can help a human user make further edits. The original data streams are all kept intact as well, so more advanced users can visualize them in our interface or use the data for their own purposes. Thus, our output is not difficult to maintain or edit, and is also more dynamic and interactive than something like a PDF, striking a balance that allows multiple types of human users to work on the final document they would like to be produced. Also, edits made to the output and subsequent sessions can feed back into the pipeline to update the information available to the various Al roles.

[0024] In one aspect, the invention provides a method for analysis and summarization of physical actions of a user and an environment while the user performs a task, the method comprising: recording a first data stream representing the physical actions of a user and an environment while the user performs the task; wherein the first data stream is selected from the group consisting of a position data stream of position and / or orientation of the user (including the user’s eyes, head, and hands) in the environment, and a video data stream of images of bodily actions performed by the user; recording a second data stream representing the physical actions of a user and an environment while the user performs the task; wherein the first data stream and second data stream are mutually distinct data streams, wherein the second data stream is selected from the group consisting of an audio data stream of sounds produced by the user and / or by objects in the environment, a video data stream of images of the object in the environment, a depth video data stream of depth images that contain the depth information to objects in the environment, an loT sensor data stream of changes of state of the object in the environment; processing the first data stream with a first Al model to produce a first higher-level representation of the first data stream; processing the second data stream with a second Al model to produce a second higher-level representation of the second data stream; wherein the processing of the second data stream with the second Al model uses the first higher-level representation of the first data stream as input; wherein the first Al model is distinct from the second Al model; wherein the first higher-level representation of the first data stream contains timestamps; wherein the second higher- level representation of the second data stream contains timestamps; processing the first higher-level representation of the first data stream and the second higher-level representation of the second data stream to produce a summarization of the task; generating an output comprising the summarization of the task.

[0025] The method may further include recording a third data stream representing the physical actions of a user and an environment while the user performs the task; processing the third data stream with a third Al model to produce a third higher-level representation of the third data stream; wherein the processing of the third data stream with the second Al model uses the second higher-level representation of the second data stream as input; wherein processing the first higher-level representation of the first data stream and the second higher-level representation of the second data stream to produce the summarization of the task further comprises processing the third data stream to produce the summarization.

[0026] The method may further include recording four or more data streams representing physical actions of a user and an environment while the user performs the task; wherein the summarization of the task is produced by processing the four or more data streams.

[0027] The first higher-level representation may be selected from the group consisting of user hand recognition data produced from the video data stream and / or position data stream; user attention data from eye tracking data streams.

[0028] The second higher-level representation may be selected from the group consisting of speech recognition text produced from the audio data stream; object recognition data produced from the video data stream; camera motion data produced from a position data stream; shape information from a depth data stream (e.g., LIDAR); temporal / sequential data that informs on relation between objects based on data series ofvideo / hand / eye. The first Al model and second Al model may be Al models chosen from the group consisting of LLM, MMLLM, image segmentation network. The first Al model and second Al model may be Al models designed to perform object detection, object tracking, or machine learning performed on data. The first Al model and second Al model may be trained to identify objects using predetermined tagged data. The first Al model and second Al model may be trained in real time using data from the first data stream and / or second data stream. The first Al model and second Al model may be trained using supervised or unsupervised learning, or by fine-tuning an existing neural network.

[0029] The method may further include comparing the generated summarization of the task with a prior summarization of the task to produce an assessment of the task performed by the user. The summarization may be a checklist of steps performed.

[0030] BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Fig. 1A is a flow diagram illustrating a processing pipeline of a method of recording, generating documentation, and displaying documentation of a procedure performed by a person, according to an embodiment of the invention.

[0032] Fig. IB is a flow diagram illustrating details of a step processing pipeline of a method of generating documentation of a procedure performed by a person, according to an embodiment of the invention.

[0033] Fig. 2 is a schematic diagram illustrating how higher level representations of audio and hand tracking data may be used in more efficiently processing a video data stream, according to an embodiment of the invention.

[0034] Fig. 3 is a schematic diagram illustrating a user performing a task while wearing an AR headset that includes audio, video, depth, and inertial measurement unit (IMU) sensors, according to an embodiment of the invention.

[0035] Fig. 4 is a schematic diagram illustrating a user performing a task with an AR headset that includes audio, video, depth, and inertial measurement unit (IMU) sensors, according to an embodiment of the invention. Fig. 5 is a schematic diagram illustrating a processing pipeline performed by components of a device (e.g., smart phone) used to record audio and video while a user performs a task, according to an embodiment of the invention.

[0036] Fig. 6 is a schematic diagram illustrating a processing pipeline performed by components of a device (e.g., AR headset) used to record audio, video, and depth while a user performs a task, according to an embodiment of the invention.

[0037] Fig. 7 is a schematic diagram illustrating a processing pipeline performed by components of a device (e.g., AR headset and loT device) used to record audio, video, depth, and object properties while a user performs a task, according to an embodiment of the invention.

[0038] DETAILED DESCRIPTION

[0039] In this disclosure, we present an efficient and hands-free way to automatically document work processes and hands-free create documentation from recordings of user interactions with equipment.

[0040] In order to train novices how to complete some process involving performing physical actions on objects in an environment (e.g., repairing machinery, doing a medical operation like CPR, inspecting sensors on a vehicle or other machine, assembling something), good documentation (i.e., instruction manuals, technical documentation, video recordings, online courses) is crucial to an effective learning process. Additionally, we may want to document a performed procedure for record-keeping purposes, e.g., to log maintenance procedures and the results of completing them each time or to document a quality control process and the results of a particular instance of QC evaluations.

[0041] While a professional may know the actual process, they may or may not be the person who typically creates the documentation. Documentation creation is often a different, not- necessarily-overlapping skillset than that needed to complete the process, which introduces challenges in the process of transcribing everything necessary for a novice to learn from.

[0042] The skills needed to create a high-quality document include writing, figure creation, image editing, quality assurance, education, etc., meaning that an entire team may even be necessary for this process. This is of course a major logistical issue in itself. When the professional is the documenter, saving time and resources, we often see lower-quality documentation which may be indecipherable due to jargon or terminology not known to anyone else. When the documentation is left to others, there is a disconnect between the professional knowledge which introduces translation error in the process if not handled carefully and meticulously, thus adding more time from iterations.

[0043] These concerns are compounded when considering more complex but more helpful learning materials such as online courses and simulations. Additionally, even if the documentation is perfect, there are further concerns on the student’s end, such as learning transfer — even if the student understands the general process at a high-level from the documentation, there is often no guarantee that the high-level knowledge translates to low-level skill (i.e., the actual physical operation of a machine). This is often covered by face-to-face instruction and practice, but this of course introduces limitations related to manpower resources.

[0044] Additionally, many more complex processes, such as military aircraft operation and medical skill practice, are typically taught through simulations as well as documentations, in some cases, through tools like augmented reality (AR) and virtual reality (VR), which introduce more data streams to the process, such as depth sensing, hand tracking, spatial tracking of the head-mounted display (HMD) as well as the equipment, and much more, depending on the application.

[0045] Recent advances in these areas also include complex evaluation methods to determine if a student is learning and improving, which further increases the required resource overhead just for this documentation process.

[0046] At some point, it becomes infeasible for any team of limited resources to document such complicated training procedures, and even now, there are few methods of automatically documenting most processes. With recent advances in artificial intelligence (Al), we have more options for automating the process of documenting almost anything. We have Al such as ChatGPT which can summarize long passages with linguistic styles closer to real human vernacular than ever before, and it can also read sensor data from images, infer context, determine objects of interest, and much more. We can even define the Al’s roles, personalities, and many other parameters practically guaranteeing a certain type and quality of output that has never been seen before in computer-aided technology.

[0047] We augment the current process of making documentation like instruction manuals, simulations, online courses, etc. by using Al to learn the process itself and document it for the professional instead, thus reducing the manpower required for the documentation to be sufficient and complete.

[0048] To meet this end, we have developed a system that that uses Al and multimodal summarization techniques to create both interactive (e.g., online course-like, 3D AR guidance) and static (e.g., PDF, PowerPoint) documentation of some process that is recorded with some number of data streams (e.g., some combination of audio, video, depth, handtracking, camera tracking).

[0049] The system uses different Al roles that learn and share context and semantic understanding of the process such that the Al itself essentially becomes the professional, e.g., some digital representation like an avatar of the person who recorded it, who additionally has the documentation skillset that the professional is often lacking or might not have the time to do, like writing, assisting the student, etc.

[0050] Our pipeline is modular, thus it can work if the process is recorded with a phone or something as advanced as a Microsoft HoloLens, by finding as much semantic information and context as is feasible given whatever data streams are available and whatever information the Al is capable of discovering or extrapolating. As a result, instead of requiring a team to create sufficient documentation for a process, the professional simply records the process themself with some device (i.e., a phone, AR headset such as Microsoft HoloLens), and the recording is simply converted to documentation with our pipeline, with options for both static display (e.g., a PDF similar to traditional methods) and interactive display (e.g., an editing interface, playing videos, automated evaluation of a student).

[0051] Embodiments of the invention provide methods for an Al-assisted automatic documentation creation. This method automatically generates comprehensive documentation of a task or process performed by a user by leveraging artificial intelligence to process multiple heterogeneous data streams captured during the performance.

[0052] Following is an overview of some key feature of a method for automatically generating documentation:

[0053] 1. DATA ACQUISITION VIA MULTIMODAL RECORDING

[0054] The process begins with the recording of a user performing a task or operation, acquiring data of multiple heterogeneous types. The system is designed to be compatible with a variety of sensor types, capturing multiple synchronized or asynchronous data streams. These data streams may include, but are not limited to:

[0055] • Audio: Captured via a microphone, recording the user's speech, verbalizations, and ambient sounds related to the process, such as equipment noises. This audio can be processed for speech-to-text transcription, speaker annotations, and the identification of sound events indicative of process steps or evaluation criteria.

[0056] • Video: Standard two-dimensional video recordings of the user and the operation area, capturing visual aspects of the task being performed. This stream can be analyzed for object detection, object tracking, and activity recognition to understand the actions and interactions within the scene. There may be multiple video streams, e.g., Augmented Reality headsets may include a color camera stream, and the raw video streams of one or more wide-field of view cameras for room tracking, hand tracking, and eye tracking. • Depth: Video data augmented with depth information, providing the distance of objects from the recording device for each pixel. Depth sensors may include, for example, infrared and time-of-flight (ToF) techniques. This data enables 3D reconstruction of the operation area and the objects within it, facilitating object segmentation and the understanding of spatial relationships.

[0057] • Hand Tracking: Data capturing the position, orientation, and articulation of the user's hands. This can be achieved through dedicated hand-tracking sensors or inferred from video and depth data using computer vision techniques. Hand tracking allows the system to determine which objects the user is interacting with and the gestures they are performing.

[0058] • Eye Tracking: Data indicating the direction of the user's gaze, typically captured using small cameras embedded in AR headsets. This information reveals where the user is focusing their attention during the process, providing valuable context for understanding their actions and intentions.

[0059] • Head Tracking / Camera Tracking: Data on the position and orientation of the recording device (e.g., AR headset or camera) in the environment. This is often achieved using inertial measurement units (IMUs) and simultaneous localization and mapping (SLAM) techniques that combine data from multiple sensors. This allows for stabilization and alignment of different data streams and the creation of a consistent spatial frame of reference.

[0060] • Virtual Annotations: User-generated virtual labels, highlights, or other graphical elements overlaid on the real-world scene during recording (e.g., via an AR interface) . These annotations provide direct indications of points of interest or specific components relevant to the process, significantly focusing subsequent Al analysis.

[0061] • loT Sensor Data: Readings from sensors integrated with the equipment or environment involved in the process. This data can provide information on the state of the equipment, environmental conditions, or other relevant parameters that change during the task. 2. AI-POWERED MULTIMODAL DATA PROCESSING:

[0062] The captured multimodal / heterogeneous data streams are then processed by different Al models tailored to each data type. This processing aims to extract higher-level representations of the information contained in each stream. Crucially, the representation of one data stream can be used as input in the processing of another, enabling a deeper understanding of the relationships and context within the recorded process.

[0063] Examples of this cross-modal processing include:

[0064] • Using hand tracking and gaze data to focus object detection in the video stream on the specific areas or objects the user is interacting with.

[0065] • Correlating speech transcripts with visual actions and object interactions to understand the steps being performed and the objects being manipulated.

[0066] • Leveraging depth information to segment objects in the video and understand the 3D structure of the operation area.

[0067] • Synchronizing events detected in different data streams based on timestamps to establish temporal relationships and cause-and-effect.

[0068] • Using virtual annotations to identify key features or components that the Al should focus on in other data streams.

[0069] This processing is facilitated by a "meta-prompt", a dynamic data structure that accumulates metadata and intermediate processing results from the different data streams. This metaprompt acts as a central repository of contextual information that is shared and utilized by the various Al roles within the system. It can be structured in a format like JSON, allowing for iterative expansion and refinement of the system's understanding of the process. For example, instead of finalizing a step like "[Step 1] Hit the button to turn the light on", the meta-prompt provides the context needed for the Al to further understand the key characteristics of the step, i.e., a JSON-like structure that can be expanded over multiple data processing iterations and that lists an image of the button, and the location of button and user’s hand. 3. COLLABORATIVE Al ROLES ("Al FACULTY"):

[0070] The method employs a modular pipeline using multiple Al agents, each with a specific role and focused on a particular aspect of the documentation process. These Al roles collaborate and share information (primarily through the meta-prompt) to generate the final documentation. The key Al roles include:

[0071] • Recording Assistant Al: Provides real-time feedback to the user during the recording process, guiding them on what to record, when to provide additional explanation, and ensuring the captured data will be sufficient for generating comprehensive documentation. This can involve suggesting different angles, prompting for verbal explanations of actions, or ensuring all relevant steps are captured. This role might leverage prompt engineering techniques to guide the user effectively. An example of a recording assistant Al can provide real-time feedback about camera movement, warning the user about shaky video quality during too much and rapid camera movement and if certain objects mentioned by the user are not detected in the camera data.

[0072] • Documentation Al: This core Al agent is responsible for creating the documentation itself after the initial data processing. It analyzes the higher-level representations of the multimodal data to understand the process flow, identify individual steps, extract keyframes and relevant segments from video and audio, organize the content logically, and generate textual descriptions and summaries. It can also integrate any provided "seed context," such as existing instruction manuals, to enhance the documentation. Techniques like natural language processing (NLP), including summarization and relationship extraction, are employed.

[0073] • Localization Al: Adapts the generated documentation to specific fields (e.g., using appropriate terminology for medical procedures versus industrial machinery) and translates the content into target languages. This ensures the documentation is accessible and relevant to the intended audience.

[0074] • Editing Assistant Al: Assists human users in manually editing the generated documentation. It can provide suggestions for rewording or restructuring content, answer questions about the documented process (leveraging the shared context), and ensure consistency. This role understands the underlying process and can guide edits based on that understanding.

[0075] • Learning / Evaluation Assistant Al: Designed to aid students learning from the documentation. It can answer student questions about the process ("where does this part go?", "why is this step necessary?"), provide guidance during a student's attempt to replicate the process (potentially in an AR environment), and automatically evaluate the student's performance. This evaluation is based on the identified goals, observable criteria for success, and cause-and-effect relationships learned from the professional's recording. This Al can generate checklists and compare a student's actions and the resulting state of the operation area to the original recording.

[0076] 4. DOCUMENTATION OUTPUT AND FORMATS:

[0077] The Documentation Al compiles the processed information into a structured representation of the process. This representation serves as the basis for generating the final documentation in various formats:

[0078] • Static Documentation: Traditional formats such as PDF documents, similar to instruction manuals, containing text, images, and potentially 3D model snapshots.

[0079] • Interactive Documentation: More dynamic formats such as online courses or interactive tutorials. These can include embedded video and audio playback, interactive 3D viewers, and access to the Al assistants for questions and guidance.

[0080] • AR / VR Integration: The system can generate content suitable for display in augmented or virtual reality environments, potentially creating a "digital twin" of the instructor performing the process or providing step-by-step guidance overlaid on real equipment.

[0081] • Editable Interface: The output can be structured to facilitate easy editing and customization by human users through a dedicated interface. The original data streams are preserved, allowing advanced users to access and utilize them for further analysis or customization. 5. ITERATIVE LEARNING AND IMPROVEMENT:

[0082] The system is designed to learn and improve over time. Edits made to the documentation by human users and data from subsequent sessions (e.g., student attempts and evaluations) can be fed back into the pipeline to update the information available to the various Al roles and refine the system's understanding of the process. This iterative feedback loop enhances the accuracy and effectiveness of future documentation and evaluations.

[0083] Fig. 1A is a Processing Flow Diagram illustrating the steps performed in one embodiment of the invention.

[0084] Step 102: Multimodal Recording.

[0085] A user performs a task / process while multiple heterogeneous data streams record the user and environment. Optionally, a recording assistant Al 100 helps the person recording know what to record, when to expand on the documentation process itself, when what they are recording will not be sufficient for a subsequent viewer to understand.

[0086] Each recorded data stream is pre-processed to generate a low-level representation of the data stream using an Al model designed specifically to process its specific type of data. For example, an audio data stream is pre-processed to produce a text transcription, a video data stream is pre-processed to produce boundaries of unidentified objects. The output of the recording process 102 may include these low-level data streams representing audio, video, depth, hand tracking, eye tracking, head tracking, virtual annotations, loT sensor data, and metadata, preferably with timestamps to facilitate synchronization.

[0087] Step 108: Automatic documentation pipeline.

[0088] The recorded multimodal data streams are processed by an automatic documentation pipeline to produce a summary of the recorded task in the form of static or interactive documentation 110. Details of step 108 are shown in Fig. IB, which illustrates for the sake of simplicity the processing of two example data streams. Low-level representations of first and second data streams are input into first and second Al models 128 and 130, respectively. These models generate first and second higher-level representations of the data streams 132 and 134. Significantly, this involves cross-modal fusion, where the higher- level representation 132 of the first data stream is used as input to the second Al model that generates the higher-level representation 134 of the second data stream. For example, a higher-level representation of an eye tracking data stream (the direction the user is gazing) is used to generate a higher-level representation of a video data stream (e.g., an object detected in the camera image in the direction the user is looking). Other higher-level, time- stamped representations 132 may include, for example, identified user actions, recognized objects and their states, 3D models, and speech segments associated with specific actions.

[0089] A key feature of this technique is that higher-level representations of data streams representing the position or orientation of the user's body are used in generating higher- level representations of other data streams. This can dramatically reduce processing load and accuracy of generating the higher-level representations by strongly limiting the possible solution space, making further data and Al analysis much easier. For example, the location the user is touching / looking at / virtually labeling (which is derived from hand tracking, head and eye tracking data stream and annotations) can be used to narrow down the search for an object when generating higher-level representations from a video stream. In this example, the hand tracking result, gaze location, or virtual label is a higher level representation. The hand tracking is based on the video or depth camera stream, gaze is based on a combination of eye tracking and head tracking using simultaneous localization and mapping (SLAM), which is a combination of multiple video streams and an inertial measurement unit. The virtual labels are a result of hand tracking results during a certain time frame in a frame of reference that is setup by the SLAM.

[0090] These higher-level data streams are input to a processing module 136 that includes a documentation Al and localization Al 106 which generate the summary 110 in the form of documentation. First, the documentation Al creates the raw documentation, which involves developing an understanding of the process, finding keyframes, organizing the text, and finding any extra relevant information given any seed context such as instruction manuals or specifications of machine parts and operation 104. Next, the output from the Documentation Al is used to create the final documentation in various formats (static PDF, interactive online course, AR / VR content). This involves structuring the content, embedding relevant media (e.g., images of the object the user was looking at in the example above, video, audio, 3D models), and integrating interactive elements. Finally, the localization Al translates and localizes the document such that it is appropriate to the field of use. This may involve translating into a target language and including appropriate technical terminology.

[0091] Step 114: Interactive display and editing interface.

[0092] After the documentation 110 is generated, it is input to an interactive display and editing interface where the user can review, interact with, and edit the generated documentation. An optional editing assistant Al 112 can assist the user with manual editing and can answer queries about the process.

[0093] Step 120: Documentation viewer.

[0094] The edited documentation 116 from step 114 can be viewed by an appropriate user interface 120. For example, a static PDF document or dynamic video could be viewed on a display screen.

[0095] Step 126: Learning and evaluation.

[0096] If a student attempts to replicate the process while being recorded, their performance can be evaluated by a basic evaluation tool 124, or automatically evaluated by an automated evaluation tool 118. Optionally, a learning / evaluation Assistant Al 122 can compare their actions and the resulting state to an original recording and identified evaluation criteria. The Learning / Evaluation Assistant Al identifies evaluation criteria and develops tools for assessing student performance. Student performance data can be fed back into the system to refine the meta-prompt and improve the Al models for future documentation and evaluation tasks.

[0097] In an illustrative example, a user explains how to connect a pump with a manifold.

[0098] 1.) The expert user looks at the pump and manifold while verbally explaining that this is a pump and a manifold. User looks at pump while explaining the pump and points with the hand to the manifold while explaining the manifold. The user places a virtual label at one of the connectors on the manifold and another virtual label at the pump connector while explaining that these are the connectors where both parts will be connected later.

[0099] 2.) Data streams: a. Video that shows the pump and manifold. b. Audio that includes the words pump and manifold at certain timestamps. c. Gaze tracking during the time when the user explains the manifold limits the possible number of frames and the number of pixels in these frames that contain the manifold. Object detection and segmentation can be performed on small number of frames and pixels (regions of interest) to identify the manifold. d. Hand tracking during the time when the user explains the pump limits the possible number of frames and the number of pixels in these frames that contain the pump. Object detection and segmentation can be performed on small number of frames and pixels (regions of interest) to identify the pump. e. Location of the virtual labels very accurately pinpoints the location of both the connector on the manifold and the connector on the pump. Timestamps of virtual label placement and audio and spatial location of virtual labels and manifold and pump help to identify, which connector belongs to the manifold and which connector belongs to the pump.

[0100] 3.) User connects a yellow tube to the connector on the manifold and to the pump while explaining "This yellow tube is connected from manifold connector to pump.".

[0101] Data streams: a. Video shows the user’s hand holding the tube. b. Hand and eye tracking identify when the hands are close to the manifold and manifold connector, helping to select the frames(timestamps) and pixels (location) where the user is connecting the tube. Similar thing for the connection to the pump. c. Camera and audio register that the tube connected to the respective connectors needs to be yellow. 4.) Automated documentation creates instructions containing the video, summarization of the audio, a parts list, and relevant areas in the video or extracted frames that highlight different parts described during the process. The instructions can be shown on a video, printed out via step-by-step instructions (e.g., pdf, word) or played in 3D on an AR headset as a virtual digital twin of the instructor performing the same procedures as the real instructor.

[0102] 5.) A student connects a red tube to a different connector on the manifold and the same connector on the pump.

[0103] Data streams: a. Video shows the user’s hands holding a tube. b. Hand tracking limits the frames and pixels on the video to identify the tube. Red tube is detected. c. Hand tracking and gaze tracking identifies that the tube is connected to a wrong connector on the manifold. d. Hand tracking and gaze tracking identifies that the tube is connected to the pump.

[0104] 6.) Same system can now be compared against existing instructions to identify incorrect execution of procedure.

[0105] A significant advantage over existing Al based procedures is that the higher-level representation of multiple data streams including gaze, hand tracking and virtual labels allows the system to pinpoint times, locations and objects the user is interacting with, which may be a difficult task for Al if encountered with an unknown setup that contains multiple unknown objects of the same type (e.g., multiple connectors on a manifold, multiple similar looking buttons, levers).

[0106] A higher-level representation of a data stream is defined as the result received from processing that reduces the available data in the data stream to a lower number of relevant datapoints that can be in a different representation. A higher-level representation can be used as input into another processing method. A higher-level representation can be an abstraction or filtering of a lower-level representation. It represents the same type of data in a more compressed or abstract form. Illustrative examples are as follows:

[0107] As shown in Fig. 2, a video data stream 200 is recorded along with an audio data stream 202. The audio stream 202 is processed with speech-to-text to produce higher-level representation of time-stamped text data, "turn on this button" synchronized with a times ti to ts. This higher-level representation of the audio is used in the processing of the video data stream 200 to limit the processing to the specific frame 204 when the words are spoken. Limiting to video frames when words are spoken dramatically reduces the number of video frames that need to be processed for object recognition. Moreover, at time ts when video frame 204 is captured, a user is pointing with their hand 206 at an object 208 in a region of interest (ROI) 210 within the video frame 204. The region 210 is specified by a higher-level representation of hand tracking data that results from processing the hand tracking data to identify the region 210. This higher-level representation of the hand tracking data can then be used in the processing of the video frame 204 to limit the data in the frame to the region of interest 210, dramatically reducing the number of pixels that need to be analyzed by a computer vision or Al algorithm for detection of the object 208 to produce a higher level representation of the video. This allows efficient coordination of the button object 208 derived from the video stream and the words "turn on this button" derived from the audio stream. If, at a later time t6 the text "then this" is extracted from the audio, and another region of interest is extracted from the hand tracking data, these can be used to reduce the computational burden of object recognition in video processing. If processing of the region of interest identifies multiple objects, then the word "button" extracted from the audio processing can help select the correct object in the region. The higher-level representation of the video processing can be a label, time stamp, and pixel locations. This higher-level representation can then be plugged into e.g. a large language model together with another object at another time for processing of the relative order of interaction with different objects. The documentation created automatically for this example might include a still frame or a video clip from times ti to ts together with written instructions "Press this button." An example assessment created by system could include asking the user to identify and press the button, and testing via hand tracking and object detection whether the button was pressed. 3D instructions created by the system could include a 3D reconstruction of the button (e.g. using NERF, Gaussian Splatting, 2D to 3D model reconstruction), animated 3D user interaction showing the user recording the procedure as a digital avatar guiding through the process, and a 3D animation of a virtual hand pressing the button. If more than one button is present, the examples above could include similar instructions and evaluations, taking into account the proper sequence of pushing the different buttons.

[0108] Representations of Data

[0109] We can process audio data to gain higher-level representation such as individual speakers, ambient noise, speech-to-text, etc. Simple operations like removing ambient noises is often accomplished with noise gates, which essentially delete sound frequencies in a certain range. To determine whether or not some low-pitched sound is actually noise rather than a novel non-ambient sound source, we use more complex noise gates which only turn on after the same sound is heard over some time interval, after which we assume it is simply some continuous ambient noise.

[0110] Segmenting speech is a more difficult challenge involving multiple steps. We begin with denoising operations like removing ambient noise so that human voices, which tend to be in the mid-range of audio frequencies, are as clear as possible. We then use some program like a trained neural net to recognize specific parts spectrograms as syllables, possibly normalizing the spectrogram first so that it is voice-agnostic. These details and mappings are often encoded as Mel-frequency cepstral coefficients (MFCCs), which help us look up what a particular syllable should look like in a spectrogram. This general process is called spectral analysis.

[0111] In some cases, like with homophones (words that sound the same but are spelled differently), we can use the context of the sentence with NLP to determine the correct option. To recognize different speakers, a process called diarization, the patterns of speech, pitch of voices, etc. are used to identify speakers, which is much easier if the speakers are alternating. To recognize accents, vernacular, and other more individualistic speech qualities, speech-to-text networks are simply trained with more data on these characteristics, perhaps storing MFCCs particular to the localization.

[0112] Images are recorded by lenses which contain optics more or less representing the human eye, with lenses, irises, and various components representing cones and rods. There is often a process to undistort the image, because the more non-spherical that a lens is, the more incorrect the shapes are of the objects in the video. For example, some types of lenses with an extremely high field-of-view (FOV) like fisheye lenses introduce extreme bends of what should be straight lines, especially near the edges of the lens. This is caused partially by the structure of the lens vs. the sensor as well as physical effects such as the Fresnel effect which distorts light rays near the edges of translucent surfaces.

[0113] The general processing method for images has been discussed elsewhere in this document, so instead, we focus on higher-level data inferences such as reconstruction of 3D surfaces. This is typically achieved one of two ways: directly (using a device called a depth sensor, which emits infrared light into the scene and senses the reflected light to determine how far the reflection happened, thus how far the surface was), and indirectly (typically using photogrammetry, which finds patterns between multiple images in the scene to localize the different viewpoints and determine the structure of the scene and its main features by triangulating multiple camera viewpoints).

[0114] With either of these methods, we receive a point cloud, which is the set of features used to determine the structure of the seen surfaces, triangulated such that they are represented now as points in 3D space. From here, we perform multiple operations to interpolate the structure of the 3D scene by connecting the points (triangulation) and texturing them (determining the colors of points or surfaces). The specific details of this 3D reconstruction process are found in various 3D graphics and computer vision material.

[0115] The result of the 3D reconstruction is a textured 3D mesh / model, which we can further process to do complex tasks like segment / split the mesh into individually recognized objects so that they can be made individually interactive, as a full non-segmented mesh from a 3D reconstruction is typically not useful for much more than collision detection in interaction scenarios like video games and simulations.

[0116] From this process, we also gain information about the camera trajectory, albeit not very accurately in most cases, especially if it needs to be done in real time (i.e., when the recording is happening). To gain information about camera position in real time, the standard method is for the recording device to include an inertial measurement unit (IMU), which provides 6 degree-of-freedom (DoF) acceleration values, which can be added to estimate velocity values, which can be added to estimate the position of the device relative to its initial position when the recording started. The 6D0F transform includes 3D position (X, Y, Z) and 3D rotation (pitch, yaw, roll). This general process is called simultaneous localization and mapping (SLAM), with the localization referring to finding the camera trajectory and the mapping referring to finding the 3D structure of the environment.

[0117] A common issue with IMUs is that they rely on a process called dead reckoning, where you add acceleration or velocity values to try to estimate a position. However, no sensor is free of error, and with each acceleration measurement having some error, these will compound the overall error unless mitigated somehow (called cumulative error). A related concept in calculus would be using the Euler method to estimate an integral based on samplings of the derivative. However, since IMUs are so fast compared to other methods like triangulation using photogrammetry, IMU usage is helpful in most practical circumstances.

[0118] We mitigate this by finding global indicators of position, albeit at lower sampling rates than IMU sensors, and mitigate the local (IMU) and global estimates with a process called sensor fusion. The global indicators of position may be trajectory estimated from photogrammetry, as discussed above, or even global positioning system (GPS) data, depending on the accuracy. As a real life parallel, it is almost impossible for a human being to walk in a straight line with their eyes closed (equivalent to IMU only), but very easy if their eyes are open because their estimates of trajectory are grounded in the real world. A common technique to accomplish this is called Kalman filtering, which takes multiple estimates of the same data point, albeit at different frequencies and accuracies, and combines them by making corrections to the more inaccurate estimates periodically using the more accurate estimates. For example, we can first take an image from the camera at t=0, then we can take IMU data over 10 timestamps (which will result in relatively high error estimates of position with dead reckoning), then take another image at t=10, perform triangulation, and observe how much worse the IMU estimate is from the triangulation estimate, then modify the IMU’s understanding of its own accuracy with that discrepancy. A similar technique with slight modifications is called the particle filter, which works better for non-Gaussian estimates of inaccuracy. For example, most sensors like IMUs might say they are +- X m / sA2 inaccurate, which is Gaussian as this is like drawing a circle around the data point, so Kalman filtering would work well. For something like a magnet whose accuracy is dependent on the nearby magnetic field, or electricity, a particle filter may be more appropriate.

[0119] From images and videos especially, we can also determine gestures. This is typically done by first recognizing a hand or other limb, then fitting bones to the observed fingers, then assuming details about the 3D structure of the hand using inferences about human hand size (many of these inferences can be skipped or trivialized when using a depth sensor), then observing the hand over time.

[0120] Many gestures need to be seen to be recognized. For example, if a user’s index finger is already pointing at something before we see their hand closed, this is often not recognized as a gesture. Instead, for better accuracy, it is often better to observe a state machine, i.e., observe the closed hand, then the bent finger, then the finger pointing straight, because there are many reasons that a hand may be in a particular configuration and we want more granular control over when the gesture is actually recognized, because gestures are often meant to represent discrete concepts or commands rather than continuous (i.e., pointing at something is one of relatively few continuous motions to recognize). Artificial Intelligence

[0121] Artificial Intelligence (Al) is a general set of techniques for getting a computer to act smart, which, in technical terms, typically means programs that can automatically create higher- level inferences on some set of data. The term Al means different things to different fields and people, but in computer science, it is most often used to refer to any program which uses machine learning (ML) techniques, as discussed later.

[0122] For example, a non-AI program that processes text might do basic operations like word count, letter count, frequency of each letter, redundancy checking (frequency of words).

[0123] A program that appears closer to Al as in it seems very advanced to humans, but is still not typically Al (as in, there is not much smart about how it is done at a technical level), would be grammar checking. Since human language always has grammatical structures and rules defined by the language itself (as in, the language could not be formalized without a set of rules for how words must interact, parts of speech, etc.), a computer program must simply have these rules hard-coded into itself to know whether grammar is correct, it does not actually need to understand the content.

[0124] We transcend this barrier into more Al-like programs when we require an understanding of the content itself, i.e., when determining if the language is appropriate (e.g., formal / informal, inclusive) or when determining if there is a better way for the text to be written.

[0125] In order to accomplish this, we need examples, i.e., we must train the program by having it evaluate multiple human-doctored examples of what it looks like to correct text well. After enough of these examples, the program learns the patterns to correct, and with enough examples, it may also gain information about the relationship between the content itself and the corrections.

[0126] To relate with examples from a different field, we can apply the same idea to image processing, i.e., detecting if a dog is in an image. We train the program to recognize dogs by simply providing it enough pictures of dogs and not-dogs, and then it can determine if some image that it has never seen before contains a dog by comparing the new image’s content to the images it was trained on.

[0127] However, there is a real risk of overfitting, i.e., in the above example, if the program is not built carefully, it will simply determine that a new image is not a dog because it does not exactly match the images of dogs it was trained on, pixel for pixel. Thus, in addition to providing more examples, we also need to make the program more complex by allowing it to recognize features, as in interesting content, that tell us whether or not something is a dog rather than by directly comparing images.

[0128] We accomplish this by distorting the original data and putting the results through a set of nodes or neurons which perform different functions until it can find specific features.

[0129] In the case of processing language, we may want operators which remove prepositions, provide synonyms, etc.

[0130] In the case of processing speech from an audio clip (mp3, wav, etc.), we may want to cull sound frequencies which are not typically found in human speech, or the specific syllable being recognized (e.g., to recognize a "k" sound, we may want to observe only higher frequencies of human speech). However, due to the variety of human voice pitches and other features, we may first want to normalize the audio so that our program does not overfit on a specific voice.

[0131] In the case of processing images, we may want to use kernels to perform convolutions which result in modified images, i.e., an image with clearer edges and less blur, an image describing when the color gradients are more extreme, etc. After a few convolutions, we may be able to recognize contours, shapes, and other higher-level features like ears and mouths that better define what a dog actually looks like rather than a group of pixels. These higher-level features may also be thought of as descriptions of the data.

[0132] A pipeline of these neurons / nodes creates a neural network. These are typically structured so that you provide a set of data to it (X), it performs a set of convolutions or other operations to the data through nodes in hidden layers (any layer that is not input or output, but rather some function that modifies the data, like f(X) ), and then returns a result (Y). This result may be a binary label (i.e., dog or not dog, encoded as 1 or 0), category (e.g., part of speech: adjective, noun, verb), value (e.g., how inclusive is the speech in this sentence on a scale of 0.0 to 1.0), or many others depending on the application.

[0133] Additionally, there are some types of networks which can complete data or even generate new data that looks similar to the original data.

[0134] An example of the former would be an autoencoder, which is a type of network used to perform operations like denoising (i.e., remove empty pixels from a rendered image, upscale to a higher resolution by filling in the missing pixels). These often function as a form of interpolator (i.e., using observable patterns to complete the data, often by fitting a function to it).

[0135] An example of the latter would be a generative adversarial network (GAN), which can generate new images like human faces which do not exist, by learning patterns in human faces that it is trained on. Various hyperparameters (human-editable parameters) can also be modified to change specific details about the output (e.g., wrinkles, skin color, sex, hair color and detail). GANs often involve some form of autoencoder which fills in details about the image to be generated.

[0136] Another type of network which completes similar tasks is the transformer architecture, which is similar to a GAN, but rather than operating on the entire data (e.g., an entire image like a face), it seeks to understand and complete a sequence, such as a sentence, often by storing a set of rules like grammatical structures that help it understand the context and purpose of specific elements of a sequence like a sentence, video, etc.

[0137] A related concept in image processing is the recurrent neural network (RNN) and long shortterm memory network (LSTM), which attempt to find patterns over time, such as detecting an object as it moves over multiple image frames, i.e., a video. However, in a network like an RNN, the temporal variable is consistent, i.e., timestamps or frames which proceed at a particular rate in a linear fashion. With transformer architectures and language processing in general, the progression of the text is determined by the rules, the content itself (e.g., keeping track of the story, the characters / subjects, etc.), and context that can be determined over variable amounts of space between processed components. Transformers accomplish this by performing various convolutions, embeddings (i.e., some way to encode words in a normalized way), and determining attention or context (i.e., what is the relationship between different words or concepts in the passage).

[0138] Combining concepts from GANs, RNNs, CNNs, and transformers, we see more recent developments in this space in the form of chatbots such as GPT (Generative pre-trained transformer), which process text (e.g., human conversation input, books, webpages), and then generate responses appropriate to the conversation. What is determined to be appropriate is defined by the chatbot’s role (i.e., informational Al, a friend, fictional character), the content that has been provided to it already, and whatever data it was pretrained on (e.g., Github repos, web scraping results — as in almost anything on the internet, Reddit forum posts). With all of this information, the chatbot can now conversate like a human and can be asked to do specific tasks.

[0139] Unless specified otherwise, the response by GPT-like Al is at the mercy of the Al, but if one desires a very specific output (such as website code), more or less detail, specific language, etc., then they will need to change their prompt to the Al, requesting that it respond in a specific way or process the human’s messages in a particular order or with specific implementation details. This process is called prompt engineering and is becoming increasingly crucial for people, especially programmers, who wish to work side-by-side with an Al who can help them in specific ways.

[0140] Implementation details

[0141] We begin with a recording of the process to be documented. We assume that there is at least one sensor that records a user while they are performing their procedure. This sensor can be at least one of the following: • A microphone to record the user’s speech while they are verbalizing their task. The audio could also be used by the Al to learn about what the process sounds like, i.e., audio cues that indicate a button was pressed or lever was pulled

[0142] • A color camera to film the user’s actions, e.g., a head-mounted camera that records the first-person view of the user, or a phone camera

[0143] • A depth camera to film the user’s actions, e.g., a head-mounted infrared / time-of- flight (ToF) camera that records the first-person view of the user, or a phone Lidar camera such as those on newer Apple and Samsung phones

[0144] • loT sensor data of the equipment that is recorded while the user is interacting with the equipment, i.e., internal components that broadcast when a button is pressed, sensor readings shared over Bluetooth or COM ports

[0145] • An eye tracking camera that tracks where the user is looking, i.e., those including on the HoloLens or MagicLeap

[0146] • Etc.

[0147] Each sensor data package is labeled with a timestamp when that sensor data was produced. This allows us to synchronize multiple sensors temporally to extract relevant information. Because different sensors have different frequencies of data observations, this timestamp is with respect to real dates and times that can be synchronized e.g. through a common NTP (network-time protocol) server.

[0148] While the user is interacting with the equipment to e.g., perform a maintenance procedure, the sensor or sensors are recording the user interaction. Once the user is done with their interaction, the sensor data is uploaded to a server via a network connection. On the server, the data is processed using a neural network to extract relevant information and derive higher-level representations of the data, i.e., audio to text transcriptions. Depending on the sensor data, a different neural network is used.

[0149] While it is possible for the pipeline to function with no actual recording of a process, i.e., an instruction manual is fed to the pipeline directly, this limits the utility of the pipeline in general by reducing it to a summarization task that is possible for many other Al-based systems as well, albeit we also support streams other Al summarization systems do not at the moment, such as 3D models.

[0150] The processing pipelines for each data stream are described below.

[0151] Audio

[0152] An audio recording will typically include user speech (which may include multiple users that can be separated), as well as ambient noise normal for the operation area and the noises generated throughout the process (audio events). In a machinery maintenance task, these might be audio feedback from pressing buttons. For a medical task, this may be cues from the patient (i.e., coughing).

[0153] MAKING SENSE OF RAW AUDIO DATA

[0154] The audio is processed through spectral analysis, which helps separate and categorize sources. For example, human speech tends to be low to lower midrange frequency sounds (85 to 155 Hz for adult males, 165 to 255 Hz for adult females), with individual letters and syllables having specific frequency ranges and phase lengths that can be learned by neural networks. Typically, a Mel Frequency Cepstral Coefficients (MFCC) — a type of audio feature extractor — is used for this process, especially in pipelines typical of Natural Language Processing (NLP). Additionally, if multiple human voices are present in the recording, modern NLP techniques, such as those found in Adobe Premiere’s transcription service and Google’s MixIT Al, can isolate the vocals in a process called speaker diarization, which sometimes involves a set of audio unmixing methods if multiple people are speaking simultaneously. This is possible because an individual will have a pattern of speech recognizable in the data — not only will their overall voice be a different frequency, but if the data is observed long enough, a neural net will recognize specific frequency ranges in the letters and syllables spoken by a particular speaker.

[0155] Ambient audio tends to be lower range frequency (1 Hz to 100 kHz, which is often easily segmented through denoising methods such as noise gates or other isolation techniques (e.g., FFT). After isolating vocals and ambient audio, the remainder of the audio stream will likely be random noises (e.g., keyboard taps are a common audio source isolated in online video platforms like Zoom), but in a process, these are more likely to be some feature of the process, such as equipment responses. If the operation area is the same between the professional recording the initial training video and the student, then these noises may be a strong indicator of successfully completing a step, thus it is important for the Al to make sense of these sound cues as well.

[0156] With this first step of converting the raw audio data to higher-level representations — sound source separation, transcriptions, and sound events, we move on to a higher-level stage of the pipeline that extracts features more indicative of the actual context of the process.

[0157] EXTRACTING STEPS FROM AUDIO

[0158] The primary source of identifying steps from the audio stream is from the transcription. Since each sentence and word is also timestamped (a handy feature of many speech-to-text libraries), when we plot out the speech on a timeline, there are typically noticeable gaps in the speech itself that indicate when discrete steps begin and end. With this information, we can split the speech into paragraphs as an initial indicator of these separations.

[0159] Of course, there are many different reasons why a gap may exist, for example, the speaker may simply be considering what to say next, so Al is helpful for making more sense of the content itself.

[0160] We label each individual sentence and provide delimiters between paragraphs in order to give an Al like GPT an initial understanding of our proposals of where separations are. This can be done in the form of stage direction-like organization, e.g., "[1] This is sentence one of paragraph one. [2] This is sentence two of paragraph one. [new line, \n, or ] [3] This is sentence one of paragraph two....," which is consistent with recommended methods for successful prompt engineering. With this, the Al proposes a separation of steps, and we also request it to reword / summarize each step, then translate. The translation should occur at the end because the Al will likely lose some information in translation, so it is better that the intermediate steps are in the same language as the source speech. This translation error is common in humans translating languages, but with Al, the issue is compounded because of various levels of machine translation (i.e., the way that information is encoded to a computer must be translated to other forms of machine language and to human language).

[0161] EXTRACTING EVALUATION CRITERIA FROM AUDIO

[0162] With Al like GPT, we can trivially request it to provide evaluation criteria, because often, this is a dual of the summarized points. However, the sound events discussed earlier may also be helpful for this purpose by adding them to the stage directions, i.e., "[1] Now we’re going to turn on this machine by hitting this red button, [sound of button being pressed] [2] You can see that the screen turned on when I did that." An Al is capable of automatically inferring a cause and effect relationship between the sound of the button and the screen turning on, which is aided by the use of the square bracket "[" delimiter that GPT-like Al are programmed to recognize and methods such as relational extraction which the Al uses to infer relationships between concepts in a sentence. Thus, when the student is completing the process, if they are recording themselves with their phone, an AR headset, e.g., the evaluator Al can use the context of the process shared between the Al to determine that if it hears the same button noise in the recording, it is likely a criteria for completing a step was met. Paired with the video recording, depth information, etc. as discussed later, this is quite robust.

[0163] Images

[0164] MAKING SENSE OF RAW IMAGE DATA

[0165] Images recorded by cameras have traditionally been standard 2D videos (RGB), but many devices also contain sensors that can determine depth and store it as a 4thchannel (RGB+D), which can be thought of as: for each pixel in the RGB video, it is D meters away from the camera. There are many types of depth sensors, but the most common methods use infrared sensors and emitters and time-of-flight" (ToF) techniques, which essentially emit a grid of infrared lights into the environment and measure the time it takes the light to travel to the points in the environment and back to the camera sensor.

[0166] Both the standard RGB video and the depth channels can be used for object detection in different ways.

[0167] From the RGB video, we can use object detection algorithms like YOLO and convolutional neural nets (CNNs) to detect objects in an image, and with recurrent neural nets (RNNs) and long short-term memory architectures (LSTMs), we can track an object throughout the video. Typically, with RGB video, we can do little more than create a 2D bounding box around a detected object, because due to camera distortions and various camera parameters, it may not be feasible to accurately estimate the depth of the object. With structure-from-motion (SfM) techniques, edge / plane detection, and other computer vision (CV) techniques, we may be able to infer depth from the motion of the camera over multiple frames. Modern Al models can estimate the depth of each RGB pixel from a large amount of prior training data.

[0168] With a depth stream, we know how far a particular pixel is from the camera lens, which allows us to project that bounding box into 3D and gather the relevant 3D points that should belong to that object. This may still not be accurate, as a bounding box is a poor estimate of the shape of most objects, so some common techniques used to pin the object to a 3D world include: use the middle point of the bounding box only, cluster the points such that we try to find a centroid representing the object, use computer vision techniques such as interpolation, radial basis functions (RBFs), or signed distance functions, etc. More recently, the most accurate approach is automatic segmentation of objects in an image using Al networks such as SAM (Segment Anything), which allows to easily segment the 3D points belonging to a chair from a depth stream.

[0169] EXTRACTING STEPS & EVALUATION CRITERIA FROM IMAGES

[0170] Once objects are detected, we can classify them and use Al to figure out how they fit into the overall process, especially if it is obvious that the professional is currently working on the object. Using methods similar to before / after comparisons, we can also determine if the professional’s actions actually modified an object in the operation area, which we can also use as evaluation criteria. For example, if the operator hits a button (i.e., the camera image first shows the machine with a button on it, then we recognize that a hand has entered the field of view and touched the button, then the hand leaves, thus indicating an event), we can then assume something should have changed and seek the result of hitting that button, e.g., if hitting the button turned a screen on, and the camera image indicates a screen has turned on immediately after the button press, an Al can establish a cause / effect relationship that is then used as both context and criteria. Thus, we build a knowledge graph or other form of relational data structure which Al can use to understand the process. In combination with additional sensors such as EMG wrist-bands it is possible to accurately identify whether the hand performed an action during a certain time-frame.

[0171] If a student is replicating a process while recording themselves with a phone camera, AR headset, etc., then we can seek similar changes in the operation area to evaluate them.

[0172] Additionally, if the operation area is always the same between multiple sessions pertaining to the process, each session can share a world map of the operation area, which is a standard technique in extended reality (XR) for determining that a space is the same between multiple sessions, and also identifies the key features that identify the space. Many of these systems, such as the Microsoft HoloLens’ tracking system, automatically segment the static surfaces of the environment, which can also be used to determine when parts of the workspace are dynamic and should be tracked.

[0173] Location

[0174] A basic function of most augmented reality (AR) libraries is the ability to track the camera device in 6 degrees of freedom (6D0F), which include location (X, Y, Z) and rotation (pitch, yaw, roll). Location data is usually recorded with respect to the beginning of the tracking session (where the camera is located in the real world when the AR session starts), and rotation data is usually recorded with respect to the direction of gravity and where the camera is pointing when the AR session starts. With this information, we know the trajectory of the user as they record the process to be documented.

[0175] An AR-capable device is not strictly necessary. With a regular camera and only RGB images, the trajectory can be inferred using computer vision techniques, in particular, photogrammetry, which uses changes in the camera images over time to estimate how much the camera must have moved between image frames (e.g., if in frame 1, we see the front of a table, and in frame 2, we see more of the left side of the table, we know the camera must have moved left). This is paired with structure from motion (SfM) processes, which build a 3D understanding of the environment while also estimating the camera position, returning both a 3D representation of the environment and the camera trajectory simultaneously. If the camera device contains any sensors such as accelerometers, compasses, or inertial measurement units (IMUs), then this simplifies the process as these sensors provide some initial data about the motion of the recording device. This process is called SLAM (simultaneous localization and mapping).

[0176] Given the trajectory of the camera and 3D structural data for the environment, we can determine where the user was located at different steps of the process, which parts of the environment they were focusing on, and behavioral characteristics (e.g., the velocity of the camera may indicate a pause in the process or a narration). Since this data is also timestamped, for any given point of interest indicated by other streams, such as video and audio, we can look up their user’s 3D location and essentially pin every stream to the relevant 3D positions of interest, building a 3D context that can be used for powerful inference about the process.

[0177] Additionally, as mentioned above, the world map that is also compiled by many AR libraries is helpful for aligning multiple sessions of the same process, even at different physical locations. For example, if the process involves assembling an engine, the recording by the professional can take place in one location and the student’s replication and evaluation can occur at another because the operation area can be recognized as the same. For some use cases, geographical positioning system (GPS) data may also be helpful, for example, in processes involving inspections of particular properties or many other types of location-based contract work.

[0178] User interaction

[0179] Many modern AR libraries also provide additional user interaction data, i.e., hand and eye tracking, through extra sensors such as EMG sensors and / or processes running during the recording process. Hand-tracking is inferred by fitting a skeleton to some structure that a computer vision technique determines is a hand, which can be done in 2D or 3D. Eyetracking is inferred by embedding small eye-facing camera within the front visor of an AR headset, and then using computer vision to identify and segment specific parts of the eye, such as the iris, resulting in a 2DoF rotation value (yaw and pitch) that can be used in unison with other sensors, such as environment-facing cameras, to determine the exact 3D point or direction the user is looking at.

[0180] From hand-tracking data, we can recognize specific gestures (i.e., pointing, grabbing something) that can be used to infer attention, intent, and actions / events that can help identify discrete steps. It can also be used to aid object detection and segmentation, as the hand-tracking will clearly identify the 3D object of interest and make any segmentation of that object simpler and more accurate.

[0181] From eye-tracking data, we can similarly determine attention and intent. For example, if we detect the user hit a button, and then we see in the eye-tracking data that they were first looking at the button, and then a screen with some information on it, we infer a relationship between the button and sensor readings (i.e., the button should have changed a value on that sensor, turned a screen on, etc.). Due to saccades and other factors related to attention inferred from eye data, smoothing operations, behavioral analysis, and Al-assisted inferences may be necessary to improve the quality of these results.

[0182] From EMG sensors we can understand whether the user performed a certain action with their hands, e.g. grabbed an object, interacted with equipment or performed a hand gesture. In an operation with more loT features, e.g., sensor readings can be received on the recording device somehow, the hand and eye-tracking data can be used to easily tell which sensors to read from, further improving the quality of the documentation.

[0183] Synchronization and key data identification

[0184] The time scales of the different sensor streams and the data required for procedure documentation are usually very different for different sensors, requiring some operations to introduce synchronization and / or time invariance depending on the context.

[0185] For example, if it takes the professional 1 second to complete a step, we should not assume that it takes 1 second to complete that step, but rather identify the key characteristics necessary to determine the start / end of the event and criteria to determine the cause / effect relationships between user interaction and observable events. Otherwise, we risk overfitting, over-describing, or otherwise making the detection of a step too strict such that it will not be consistently detected in future sessions. For example, if the professional took 1 second to complete a step, and everything else indicates that an observed student completed the same step but it took 2 seconds, and we use the time as too strict of an evaluation criteria, then the student’s replication of the step will be miscategorized.

[0186] As such, time is used in different ways throughout the pipeline. For example, if we want to determine a good keyframe in the video stream for a step that was identified in the audio stream, they must be synchronized (i.e., the transcription of a step from timestamps 00:00:01 to 00:00:05 is "Now, you flip this switch labelled 'POWER' to the ON position, and you’ll see a light right next to it", so the relevant video keyframe should also come from that timestamp range). However, in the final documentation, there is typically no need to include timestamps.

[0187] Also, if an Al summarizes a step, the original timestamps are no longer helpful in some cases. Expanding on the previous example, the Al may rephrase the transcription to "Flip the POWER switch to the ON position and observe the light adjacent to the switch." This rephrased transcription still refers to the same timestamps (00:00:01 to 00:00:05), but the timestamps of individual words are no longer accurate, so if we were using the word "light" to find a keyframe that shows the light that the narrator refers to, we must use the original transcription’s per-word timestamps, not the summarized words.

[0188] For evaluation criteria, we generally abandon timestamps entirely in favor of searching for the events that indicate the steps, allowing for time-invariant evaluation that is more flexible and applies to future sessions. Part of this process is removing irrelevant information and ensuring that only key features of a step are used to describe the events. Additionally, this helps ensure that the content in the final documentation is more relevant, accurate, and professional.

[0189] For the audio data, the full description of what the user was doing, or an important user observation should generally be included in the documentation. This can vary from a single word (e.g., "ok") to whole paragraphs describing the procedure or observation in detail. However, irrelevant words and phrases, e.g., "Huh?" should be removed and not considered for event identification.

[0190] For image data, a single image or a few key images often contain all the necessary information required for documentation, e.g., for a maintenance procedure, the documentation may include an image of the equipment before maintenance and after maintenance, or, with more granularity, before and after each individual event.

[0191] For video data, the video should only include the relevant part showing a step of the procedure, not parts showing different steps before or after the step. This may also include the removal of image frames that are too blurry to discern any useful content, and frames that do not actually display the operation area.

[0192] For loT sensor data, relevant data included in the process documentation can range from only a single value to a complete time series over the measurement period. Since many loT sensors continuously stream data, we must be careful to not always assume it is relevant, using information from other streams to determine when to observe the sensors. Making sense of multiple data streams and finalizing the context & process

[0193] After all of the individual streams are appropriately processed into higher-level representations which can serve as components in the final document, the next task is to finalize our understanding of the context, that is:

[0194] • What IS the process? What are its goals?

[0195] • What are the individual steps of the process?

[0196] • For each step: o What is the goal (as in, what should change to indicate the step was completed successfully)? o Which parts of the operation area does it involve? o What is the action that the user does to complete the step? o What changes before and after the user did the step? o How do we know when the step begins and ends? o Why did they do the step with respect to the overall process? o Where should the operator be located?

[0197] • What is the operation area? o Description (i.e., a workbench, a machine, a medical operation room) o What are its individual parts / tools to be operated on?

[0198] ■ Description (i.e., "a table", "a wrench", "a button")

[0199] ■ Purpose (i.e., "a table to set the tools up on")

[0200] ■ Segmented 3D model (to be used to recognize it and evaluate changes due to actions)

[0201] ■ Are they static or dynamic?

[0202] • If they are dynamic, what are its different stages?

[0203] Notice that these are all questions that should be answered by the professional, and should be answerable by an Al representing the professional as well. GPT-like Al should be able to answer questions like these given primarily textual and image information, thus, for data streams that are not so simple (such as video, depth, loT sensors, hand-tracking, etc.), much of this process is converting all data streams possible into text or image inputs to an Al.

[0204] The general process for this might look like: All data is cleaned up, i.e., audio data denoised, blurry video frames deleted, tracking errors removed. A process runs on the video and / or depth data to reconstruct the entire scene in 3D, segment objects as much as possible (e.g., chairs, tables), and label as much of these objects as possible. This creates an initial 3D context that does not yet have a concept of the relationship between any of these objects or the operation area — that information builds up over subsequent steps. Audio is used to get timestamped transcriptions + events (extraneous sound cues), which are compiled into an initial data structure (called a meta-prompt) with stage directions or context indicators that include sounds indicating certain events, multiple speakers, gaps in speech, separators between steps, etc. When possible, sensor readings are pinned to their 3D counterparts. If the sensor data comes from loT technology, that is more reliable and preferred. Otherwise, infer the data from video / images when possible. Video is used to find objects of interest in the step, which is connected to their 3D counterparts to build relationships between the information and also added to the meta-prompt. These will also be labelled in the keyframes. Connect audio transcription to the observed video, i.e., if the audio indicates a button, the Al searches the video segment for that step for the relevant button (proposing multiple options if necessary). Add this information + metadata such as confidence to the meta-prompt. Hand-tracking and gesture detection is used to create events in the meta-prompt, and when it makes sense (e.g., user pointing at something), increase a value in the meta-prompt indicating the relevance of an object of interest to the step, and provide textual descriptions of the event (e.g., the operator is pointing to the red button in the 2ndrow of the console). A similar process is used for eye-tracking. Compare the gestures from hand or eye-tracking with the audio transcription and video object detections to select keyframes. E.g., when the audio describes a known part of the operation area, like a button, we search for the closest gesture that relates to that object (e.g., when the user pointed at a button) and / or when a button was found in the video frame. When the stars align best, so to speak, use this moment in time as a keyframe representing the step.

[0205] 9. Location information used to add context to the meta-prompt, e.g., where the user should be standing. Convert this information to a textual description (i.e., the operator should be standing in front of the machine where the big screen is)

[0206] 10. Make a pass generating structures such as relationship graphs and pipelines typical of context-finding Al.

[0207] 11. Convert the meta-prompt into one complete prompt for the Documentation Al, requesting it to return a specific structure of document. This would not be the complete document itself (i.e., an HTML document), but the information that will be in the final document.

[0208] 12. Request the Al to build evaluation tools such as checklists, which are stored for use later. a. (at this point, the Al’s understanding from the automatic documentation process is complete, and will only be updated if a human intervenes in a particular steps, such as editing the output, from which the Al studies and learns for subsequent uses)

[0209] 13. Convert the Documentation Al’s summary of the process into a document that can be edited, e.g., a skeleton of an interactive and editable interface that we build separately.

[0210] With each step, the shared context and meta-prompt shared by all of the Al becomes increasingly complex, and only for the purposes of final documentation and display is it simplified. When a human user intervenes, e.g., edits the data or documentation, the context is not erased, but updated, as is the meta-prompt. Thus, future documentation iterations become more useful and powerful.

[0211] Creating the document

[0212] The Al provides the information to be displayed in the form of the meta-prompt — from there, the actual document structure would be application- and client-specific — they would need to provide the specific HTML structure, desired widgets and data representation, etc., which technically boils down to representing specific fields of the meta-prompt and documentation Al results in a particular way.

[0213] Examples

[0214] Fig. 3 is a schematic diagram illustrating a user performing a task while wearing an AR headset 300 that includes audio, video, depth, and inertial measurement unit (IMU) sensors. The raw sensor streams recorded by the headset include an inertial data stream, an audio data stream, and a video data stream including color and depth information. The video information preferably includes wide angle views of the environment, eye tracking of the user, and a forward view of the user's hands and nearby objects. During the task, the user's finger 302 points toward a sequence of locations 304. The audio data stream is processed to produce time stamped descriptions of speech. The camera sensor data streams are processed to produce time stamped descriptions of hand, head, and eye tracking. The hand projections are processed to define a region of interest (ROI) that can be used to extract or highlight relevant objects on a camera image. These time stamped descriptions are time synchronized and then used to generate documentation.

[0215] Fig. 4 is a schematic diagram illustrating a user performing a task with an AR headset 400 that includes audio, video, depth, and inertial measurement unit (IMU) sensors. The raw sensor streams recorded by the headset include an inertial data stream, an audio data stream, and a video data stream including color and depth information. The video information preferably includes wide angle views of the environment, eye tracking of the user, and a forward view of the user's hands and nearby objects. During the task, the user's finger 402 points toward an object 404. The audio data stream is processed to produce time stamped descriptions of speech. The camera sensor data streams are processed to produce time stamped descriptions of hand, head, eye tracking, as well as object detection to identify object 404. These time stamped descriptions are time synchronized and then used to generate documentation.

[0216] Fig. 5 is a schematic diagram illustrating a processing pipeline performed by components of a device (e.g., smart phone) used to record audio and video while a user performs a task. A microphone and audio recorder 510 produces an audio data stream 512 that is processed by an LLM audio transcription module 514 to produce a time stamped text transcription 516 of the recorded speech. The transcript might include, for example, text such as "00:03:23 to 00:07:22, Ok, see here, now this lever opens the door. Yes, take care." Meanwhile, a camera and video recorder 500 produces a video data stream 502 that is processed by an Al video processing module 504 that performs object segmentation and detection to produce a time stamped text description 506 of objects appearing in the video. The video description might include, for example, text such as "Between 00:03:23 to 00:07:22, Images showing lever, hand close to one lever and door." The time stamped audio description 516 is advantageously time synchronized with the video data and used to improve efficiency of the video processing 504. The audio and video descriptions 506 and 516 are then processed by an LLM module 518 to produce a text summary 520 that documents the task performed by the user. For example, the summary 520 may contain text such as "This lever opens the door." The summary 520 may also contain a video sequence and screenshots of the lever and door that opens. The summary of instructions may also contain a bill of materials that includes all relevant objects, in this case the lever and the door.

[0217] Fig. 6 is a schematic diagram illustrating a processing pipeline performed by components of a device (e.g., AR headset) used to record audio, video, and depth while a user performs a task. A microphone and audio recorder 610 produces an audio data stream 612 that is processed by an LLM audio transcription module 614 to produce a time stamped text transcription 616 of the recorded speech. The transcript might include, for example, text such as "00:03:23 to 00:07:22, This lever opens the door." Meanwhile, a camera video depth recorder 620 produces a hand tracking data stream 622 that is processed by a hand tracking module 624 that performs hand tracking detection to produce a time stamped location of the hand and direction the hand is pointing 626. Meanwhile, a camera and video recorder 600 produces a video data stream 602 that is processed by an Al video processing module 604 that performs object segmentation and detection to produce a time stamped text description 606 of objects identified in the video and their locations. The video description might include, for example, text such as "Between 00:03:23 to 00:07:22, Object detection close to the hand tracking results show a lever and that the lever moves from A to B." The time stamped audio description 616 is advantageously time synchronized with the video depth data and used to improve efficiency of the hand tracking processing 624. The time stamped hand-tracking description 626 is advantageously time synchronized with the video data and used to improve efficiency of the object detection processing 604.

[0218] The descriptions 606, 616 and 626 are then processed by an LLM module 630 to produce a text summary 632 that documents the task performed by the user. Examples of text that may be included in the documentation description include "This lever opens the door" and video sequence and screenshot of respective lever being pulled by the hand; or a 3D Augmented Reality work instruction where a virtual avatar explains "This lever opens the door", a virtual hand moves towards the lever, and a virtual label highlights the lever to be pulled; or an Al assistant that replies to the user prompt "Which lever opens the door?" on a display either an image or video of the respective lever or displays a virtual label highlighting the lever in an AR headset. Examples other possible output workflows include: 2D instructions containing step description "This lever opens the door" and video sequence and screenshot of respective lever being pulled by the hand; LLM to translate the output into other language; Multi-modal LLM that translates the label on the video sequence and screenshot into other language; and 2D instructions containing step description "This lever opens the door" and video sequence and screenshot of respective lever being pulled by the hand in another language.

[0219] Fig. 7 is a schematic diagram illustrating a processing pipeline performed by components of a device (e.g., AR headset and loT device) used to record audio, video, depth, and object properties while a user performs a task. A microphone and audio recorder 710 produces an audio data stream 712 that is processed by an LLM audio transcription module 714 to produce a time stamped text transcription 716 of the recorded speech. The transcript might include, for example, text such as "00:03:23 to 00:07:22, This lever releases the gas." Meanwhile, an loT sensor and recorder 730 produces an loT sensor data stream 732 that is processed by an Al analysis module 734 to produce a time stamped text description 736 of loT sensor data. The description might include, for example, text such as "Pressure falls after 00:07:22." Meanwhile, a camera video depth recorder 720 produces a hand tracking data stream 722 that is processed by an Al hand tracking module 724 that performs hand tracking detection to produce a time stamped location and pointing direction 726 of hands appearing in the video. Meanwhile, a camera and video recorder 700 produces a video data stream 702 that is processed by an Al video processing module 704 that performs object segmentation and detection to produce a time stamped text description 706 of objects appearing in the video close to the user’s hand. The video description might include, for example, text such as "Between 00:03:23 to 00:07:22, Object detection close to the hand tracking results show a lever and that the lever moves from A to B." The time stamped audio description 716 is advantageously time synchronized with the video depth data and used to improve efficiency of the hand tracking processing 724. The time stamped hand-tracking description 726 is advantageously time synchronized with the video data and used to improve efficiency of the object detection processing 704.

[0220] The descriptions 706, 716, 726 and 726 are then processed by an LLM module 740 to produce a text summary 742 that documents the task performed by the user. Examples of text that may be included in the documentation is a description of the steps the user completed during the task. The output of the LLM process 740 may also include more sophisticated video tutorials and interactive training. For example, an Al agent can detect a user pointing to a lever and asking "What does this lever do?" The Al agent may respond, for example, by performing hand tracking to understand where the user is pointing, object detection and object segmentation to detect objects where the user is pointing, LLM to understand that user is asking for a lever, filtering the object detection results around the pointing location for a lever, identify lever and understand that this lever releases the gas and leads to drop in gas pressure. The Al can then tell the user "That lever opens the gas valve and results in a drop in gas pressure."

[0221] Applications

[0222] In this section, we describe specific examples of how our pipeline may be used in different scenarios.

[0223] The pipeline would generally be tailored to a specific client in practice rather than utilizing a general-purpose software, but due to the multimodal and modular nature of the pipeline and its inputs, in many cases, simply adding or removing data streams is often sufficient to customize results for a specific use case.

[0224] PROCESS REPLICATION:

[0225] These are essentially processes that are followed exactly, step-by-step, with practically no deviation or situation-specific learning necessary. These usually involve machinery that always operates exactly the same way regardless of who is using it or when, such as CNC machine, laser printers, 3D printers, and various manufacturing machines you might find in factories. As an analogy, consider how LEGO block sets are built: the exact instructions should always be followed for the particular set and there’s no real reason to deviate.

[0226] As a technical note, there is always some deviation involved, because different iterations and different users result in different sensor data, thus slightly different results are possible when processing it through the AL For example, while at a high-level, it seems that there is only one way to press a button, consider that different people will hold the button down for longer, be closer to it, etc. However, since the Al is determining key step features, this should not be an issue as far as documentation and evaluation is concerned. In general, the more loT involved, the better (this seems to also be consistent among prior work), because then we do not need to guess the moment that buttons are pressed, for example but can rely on loT data that results from the button press.

[0227] When no button press loT is provided, how does the Al mitigate this variation between users? Essentially, by relying on descriptions of what it sees (i.e., person pushing button with finger) and hears (person just described that they will press a button, and we also recognize a button sound cue), and establishing cause / effect relationships between sensor streams. Initializing the Al with an instruction manual of the machine may certainly help here.

[0228] Generally speaking, for this scenario, the Al is less likely to need to add anything as the process is exact apart from details about the equipment, i.e., labelling it with names or describing a step in more detail or providing tips. The narration and instruction manual may help it better extrapolate any helpful information about elements like sensor readings. As a more advanced example, consider the disassembly and reassembly of a vehicle, which is another step-by-step process which is followed exactly. The main difference between this situation and the above examples is that car reassembly relies less on sensor readings and more on establishing cause / effect relationships between the user’s actions and movements, and the state of the operation area (vehicle). Since car parts are rigid and easily identifiable, this is not particularly difficult for the Al to figure out.

[0229] When determining how to evaluate the user in these situations, it is often some matter of detecting if a student replicated the instructor physically, and observing the same cause / effect relationships we saw in the original recording. The result of the operation area should be the same as in the original recording.

[0230] MAINTENANCE:

[0231] A maintenance task generally consists of checking the state of a machine, i.e., its physical state or sensor readings, and filling in a checklist. A secondary task often involves recording broken or damaged parts.

[0232] The recording process can depend more on the state and content of the checklist as well as following the reasoning of the narrator.

[0233] In the case of evaluating a student, we mostly observe if they are checking for the same factor that the instructor did, but do note that in this case, a subsequent inspection may actually result in a different checklist, unlike the replication tasks, which always have the same result. Thus, a student is not necessarily rewarded for having the same or different checklist, but for diligent observations that match the behavior of the instructor.

[0234] INSPECTIONS:

[0235] Inspection scenarios are similar to checklists, except instead of simply completing a checklist, the user is often recording evidence like pictures and videos which need to be examined for correctness (i.e., is the evidence sufficient to prove the state of a particular item). Some specific examples of inspections might be: an apartment tenant recording the state of the rental, a factory worker inspecting the state of the operation area, a food inspector evaluating a kitchen, or a car renter evaluating a rental for damage.

[0236] In these cases, the evidence is often received similar to a button press — as an event which provides a sensor reading (in these cases, usually a picture) which the Al evaluates. Thus, the fundamentals have not changed significantly, and with the addition of the narration of the recorder telling the Al whether or not the evidence is good, we can still document these situations quite well. The recorder may vocally approve the evidence, or the acceptance may be more implicit, i.e., the user moves on to the next step, in which case it is obvious that they’ve accepted the evidence.

[0237] To evaluate a student, we would pay more attention to the evidence they provide — the sensor readings. The Al is not looking so much for a cause / effect relationship in the way that it does with buttons, but a simpler connection along the lines of a classifier: image->good / bad.

[0238] MEDICAL:

[0239] Medical tasks are more difficult because of both variability and more challenging observations involving soft bodies (i.e., deformable objects such as human bodies, blankets, and tools). Some simple examples (as in, more predictable results) might be biopsies or checkups, while more complex examples include surgeries an laparoscopic operations.

[0240] In these cases, we may rely more heavily on observations of a 3D scan than direct image observations. A collection of 3D scans of the same object over time is often called a geometry cache, which is essentially the 3D model equivalent of a video as it stores 1 copy of the model for each frame just as a video stores 1 image per frame to simulate motion. Additionally, there is more focus on the human factor, i.e., teamwork, positioning and roles of various members of the medical staff, handing off tools, and so on.

[0241] Additionally, medical situations may have more awkward angles for the recording equipment due to physical interference with the machinery or operation area. Thankfully, with our multimodal pipeline, the Al should be able to adjust automatically by simply lacking or having weaker observations / events in the sensor streams which are less reliable due to positioning.

[0242] A challenge that may be mitigated with loT is the state of the patient — for some observations, the doctor and nurses may rely on haptics (e.g., feeling for certain symptoms), or observations in smaller areas like the throat which must be accomplishes with special small-FOV cameras like otoscopes. In these cases, an integrated loT network is the most effective way for our pipeline to help observe the patient and provide information about what a new student should be looking for, especially since medical observations are often not observable to the naked eye.

[0243] Evaluation in these cases can be quite difficult as the patient (who essentially IS part of the operation area) will likely be completely different. However, evaluation is still possible by essentially mixing the checklist and replication scenarios — maintaining a checklist of steps the student should be doing to inspect the state of the patient and observing what evidence is sufficient to prove that the student inspected the patient properly (regardless of the status of the patient), and keeping track of the overall process of the observation.

Claims

CLAIMS1. A method for analysis and summarization of physical actions of a user and an environment while the user performs a task, the method comprising: recording a first data stream representing the physical actions of a user and an environment while the user performs the task; wherein the first data stream is selected from the group consisting of a position data stream of position and / or orientation of the user (including the user’s eyes, head, and hands) in the environment, and a video data stream of images of bodily actions performed by the user; recording a second data stream representing the physical actions of a user and an environment while the user performs the task; wherein the first data stream and second data stream are mutually distinct data streams, wherein the second data stream is selected from the group consisting of an audio data stream of sounds produced by the user and / or by objects in the environment, a video data stream of images of the object in the environment, a depth video data stream of depth images that contain the depth information to objects in the environment, an loT sensor data stream of changes of state of the object in the environment; processing the first data stream with a first Al model to produce a first higher-level representation of the first data stream; processing the second data stream with a second Al model to produce a second higher-level representation of the second data stream; wherein the processing of the second data stream with the second Al model uses the first higher-level representation of the first data stream as input; wherein the first Al model is distinct from the second Al model; wherein the first higher-level representation of the first data stream contains timestamps; wherein the second higher-level representation of the second data stream contains timestamps;processing the first higher-level representation of the first data stream and the second higher-level representation of the second data stream to produce a summarization of the task; generating an output comprising the summarization of the task.

2. The method of claim 1 further comprising: recording a third data stream representing the physical actions of a user and an environment while the user performs the task; processing the third data stream with a third Al model to produce a third higher-level representation of the third data stream; wherein the processing of the third data stream with the second Al model uses the second higher-level representation of the second data stream as input; wherein processing the first higher-level representation of the first data stream and the second higher-level representation of the second data stream to produce the summarization of the task further comprises processing the third data stream to produce the summarization.

3. The method of claim 1 further comprising: recording four or more data streams representing physical actions of a user and an environment while the user performs the task; wherein the summarization of the task is produced by processing the four or more data streams.

4. The method of claim 1 wherein the first higher-level representation is selected from the group consisting of user hand recognition data produced from the video data stream and / or position data stream; user attention data from eye tracking data streams.

5. The method of claim 1 wherein the second higher-level representation is selected from the group consisting of speech recognition text produced from the audio data stream; object recognition data produced from the video data stream; camera motion data produced from a position data stream; shape information from a depth data stream (e.g., LIDAR); temporal / sequential data that informs on relation between objects based on data series of video / hand / eye.

6. The method of claim 1 wherein the first Al model and second Al model are Al models chosen from the group consisting of LLM, MMLLM, image segmentation network.

7. The method of claim 1 wherein the first Al model and second Al model are Al models designed to perform object detection, object tracking, or machine learning performed on data.

8. The method of claim 1 wherein the first Al model and second Al model are trained to identify objects using predetermined tagged data.

9. The method of claim 1 wherein the first Al model and second Al model are trained in real time using data from the first data stream and / or second data stream.

10. The method of claim 1 wherein the first Al model and second Al model are trained using supervised or unsupervised learning, or by fine-tuning an existing neural network.

11. The method of claim 1 wherein the summarization is a checklist of steps performed.

12. The method of claim 1 further comprising: comparing the generated summarization of the task with a prior summarization of the task to produce an assessment of the task performed by the user.

Citation Information

Patent Citations

  • Machine learning image processing

    US20180012110A1

  • Systems and methods for generating a content item based on a status of a node profile determined using electronic activities

    US20230325381A1

  • Systems, methods, and apparatus for enhanced cameras

    US20240046642A1