Systems and methods for using a machine learning model and head mounted device for medical procedures
The use of a head-mounted device with machine learning models to capture and analyze gaze and audio data during medical procedures addresses inefficiencies and inaccuracies in conventional methods, providing enhanced context and accuracy in medical procedure documentation.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SENTIAR INC
- Filing Date
- 2025-10-31
- Publication Date
- 2026-05-07
AI Technical Summary
Conventional medical procedures, such as electrophysiology mapping and ablation, are time-consuming and may miss relevant details, leading to inefficiencies and increased risks due to missed communication events, compromised sterile fields, and delays, while existing AI systems are prone to hallucinations and inaccuracies in transcription and reporting.
A method utilizing a head-mounted device (HMD) with sensors to capture gaze direction, audio, and environmental data, combined with machine learning models, to generate accurate and comprehensive medical procedure annotations and reports, incorporating manual feedback for improved accuracy and efficiency.
Enhances the capture of medical procedure context and operator intent, reducing human error, improving workflow efficiency, and ensuring accurate reporting and billing by automating annotation and transcription processes.
Smart Images

Figure US2025053628_07052026_PF_FP_ABST
Abstract
Description
Atty Docket No.: 34237-64696 / WOSYSTEMS AND METHODS FOR USING A MACHINE LEARNING MODEL AND HEAD MOUNTED DEVICE FOR MEDICAL PROCEDURESCROSS REFENCE TO RELATED APPLICATION
[0001] This application claims the benefit of priority to U.S. Provisional Application No.63 / 714,697, filed on October 31, 2024, which is incorporated herein by reference in its entirety for all purposes.TECHNICAL FIELD
[0002] This disclosure generally relates to training machine learning models within a medical computing environment.BACKGROUND
[0003] Conventional methods of summarizing a medical procedure, such as an Electrophysiology mapping and ablation procedure, are time consuming and may not capture enough relevant detail for medical or billing purposes.[0004J There may also be circumstances (which are currently unknown or routinely missed) during the medical procedure that are related to opportunities to improve patient outcomes (for example, missed communication events, staff unknowingly compromising the sterile field leading to an increased risk of infection, distracting unrelated conversion causing staff to miss clinically important events) or increase efficiency (for example, an missed request to record an event, or the position of a certain piece of large equipment causing staff to have to frequently wait for each other to walk around it, causing delays in the procedure). The results of such medical-procedure analysis may also be useful for training staff and physicians.[0005J Medical procedures are performed with the aid of imaging and other systems that present data on 2D display screens (e.g., ultrasound, fluoroscopy, endoscopy, navigation, electroanatomic mapping, electrogram recording systems, electronic medical record systems, etc.). Thus, physicians and clinical staff are often looking at a variety of 2D displays and real -world objects (e.g., the patient and procedure site, the physicians’ hands, medical instruments, other people in the procedure room, etc.) while performing the medical procedure.
[0006] During many medical procedures, the physician is aided by staff members (assistants). Some of these assistants operate computer-based systems (e.g., electronic medical record systems, ultrasound scanners, electrogram signal recording systems, surgicalAtty Docket No.: 34237-64696 / WO navigation systems, medical-image viewers, cardiac stimulators, etc.) on behalf of the physician, whose hands are sterile and / or occupied. By partially or fully automating the staff’s computer-based tasks, the clinic may reduce staff workload, thereby reducing costs, time, and error rates.
[0007] Conventional artificial-intelligence / machine-learning in clinical use for tasks such as transcription and report generation are susceptible to “hallucinate” and generate transcription text or report text that does not faithfully represent things that were actually said or done or happened during the medical procedure. A human reviewing the transcript or report may miss these hallucination artifacts and thus not remove or correct them in the submitted medical report.SUMMARY
[0008] In an embodiment, a method comprises receiving sensor data from a headmounted device (HMD) during a medical procedure. The method further comprises determining a gaze direction of a wearer of the HMD using the sensor data. The method further comprises determining that the gaze direction is directed to a content of a plurality of content displayed for the medical procedure. The method further comprises determining context for the medical procedure based on the content. The method further comprises generating, using a trained machine learning model (e.g., having the structure described with respect to FIG. 7) taking the context as input, an annotation for a record of the medical procedure.
[0009] In an embodiment, the method further comprises receiving an input from the wearer of the HMD using additional sensor data from the HMD; generating a manual annotation using the input; and further training the trained machine learning model using the manual annotation.
[0010] In an embodiment, the method further comprises determining one or more words spoken by the wearer of the HMD by processing the input; and responsive to determining that the one or more words are associated with the medical procedure, generating the manual annotation to include the one or more words.[0001.1] In an embodiment, the method further comprises determining a ranking of a plurality of annotations including at least the manual annotation and the annotation generated using the trained machine learning model, wherein the record of the medical procedure reflects the ranking. In an embodiment, the record of the medical procedure is a timeline including graphical representations of the manual annotation and the annotation generatedAtty Docket No.: 34237-64696 / WO using the trained machine learning model, and wherein the graphical representations are displayed with emphasis to reflect the ranking. In an embodiment, the plurality of annotations further includes an additional annotation generated using input from another user different than the wearer of the HMD.
[0012] In an embodiment, the method further comprises determining one or more words spoken by the wearer of the HMD by processing the input; and responsive to determining that the one or more words are associated with the medical procedure, generating the manual annotation to include the one or more words.
[0013] In an embodiment, the method further comprises providing a virtual graphic for the medical procedure for display by the HMD; identifying 2D content for the medical procedure displayed by a device separate from the HMD; and wherein determining that the gaze direction is directed to the content of the plurality of content displayed for the medical procedure comprises determining that the gaze direction is directed to the virtual graphic or the 2D content.
[0014] In an embodiment, the method further comprises requesting, from the wearer of the HMD, confirmation of the annotation generated using the trained machine learning model; and responsive to receiving the confirmation from the wearer of the HMD, including the annotation in the record of the medical procedure.
[0015] In an embodiment, the method further comprises receiving, from the wearer of the HMD, a modification of the annotation generated using the trained machine learning model; and further training the trained machine learning model using the modification.
[0016] In an embodiment, the method further comprises providing electrogram data for the medical procedure for display by the HMD; determining a measurement of the electrogram data; and wherein the trained machine learning model generates the annotation by taking the measurement as input.
[0017] In an embodiment, the method further comprises determining that the gaze direction is directed to a virtual content displayed by the HMD; responsive to determining that the virtual content is at least partially obstructed from a point of view of the wearer of the HMD, modifying a position of the virtual content; and further training the trained machine learning model using the modified position of the virtual content. In an embodiment, the trained machine learning model is trained using previous inputs from the wearer of the HMD and from other medical procedures.Atty Docket No.: 34237-64696 / WO
[0018] In various embodiments, a non-transitory computer-readable storage medium storing instructions that when executed by one or more processors cause the one or more processors to perform steps of any of the methods described herein.
[0019] In various embodiments, a system comprises a head-mounted device (HMD); and a non-transitory computer-readable storage medium storing instructions that when executed by one or more processors cause the one or more processors to perform steps of any of the methods described herein.
[0020] Embodiments of the present invention provide a method for processing system inputs and sensor data to extract medical procedure context, and operator intent using machine learning models. Using one or more machine learning models, medical procedure context and operator intent is manually captured or automatically processed intra- procedurally to generate automated procedure annotations and post-procedurally to generate annotation suggestions, procedure report suggestions, and a medical report summarizing the procedure (e.g., for review by the physician) with appropriate context for each recorded event and corresponding billing code. In some embodiments, the machine learning model is further trained using one or more of feedback from manual annotation events, physician report review, internal billing review of procedure report, and reimbursement review for procedure, among other available training data.
[0021] Embodiments of the present invention process any combination of imaging, positioning, aural, or ocular sensor data from any number of integrated medical information systems and a head-mounted device (HMD) during a medical procedure to generate a continually updated model of medical procedure context and operator intent.
[0022] Embodiments of the present invention provide a mechanism for hands-free interaction with a processing system in a sterile environment. In some embodiments, the HMD sensors provide position and orientation data to provide a gaze-based user interface. In some embodiments, the HMD sensors provide eye-tracking to provide eye-tracking to improve accuracy and user interface interaction intent. In some embodiments, the HMD microphones provide voice control of system functions.
[0023] Embodiments of the present invention use an HMD to provide information not available to a human transcriptionist or other machine learning models (similar to how some digital pen-based handwriting recognition use pen velocity, pressure and tilt-angle, at points along the pen strokes, and how these data channels are not available to a human looking at handwritten text on paper). A processing system uses one or more of camera video from the HMD user’s point-of-view, head position, head velocity, spatial aspects of audio, eye gazeAtty Docket No.: 34237-64696AVO direction, tracking, pupil dilation, audio from the point-of-view of the HMD user, and corresponding video of any combination of 2D, 3D, or 4D data or displays that the HMD user is looking at. For example, 2D data corresponds to conventional 2D displays such as a computer monitor, laptop screen, or tablet screen. For example, 3D data corresponds to a hologram, virtual reality graphic, augmented reality, or mixed reality graphic displayed by an HMD. For example, 4D data corresponds to 2D or 3D data and their trajectory over time as the fourth dimension. This allows the processing system to better discern the operator’s intent than other clinical procedure machine-learning systems.[00024| Embodiments of the present invention use audio monitoring, transcription, speaker identification, and speech or language processing to extract verbal features or other medical context, which can be used as training data for a machine learning model. Language processing further refines speech content to extract operator intent or medical procedure context data responsive to whether communication can be classified as information seeking or information producing. In some embodiments, HMD microphone spatial differentiation provides additional feature information for improved accuracy in speaker identification (e.g., HMD Wearer 1) and correlation of speaker and listener to medical information sources (e.g., EAMS Operator 1). In some embodiments, spatial differentiation by wearer and environment microphones provides discretion between the HMD wearer speaker and other speakers in the room. In some embodiments, the array of external microphones uses phase information to determine relative positions between multiple speakers with respect to the HMD wearer. In various embodiments, one or more HMDs communicate known relationships between speakers to identify individual speakers. In some embodiments, HMD wearer’s unique identity as an individual (e.g., HMD Wearer 1 is Doctor Smith) is assigned and communicated between HMDs according to user selected at HMD login or automatically determined responsive to biometric login (e.g., retina recognition, face recognition, etc.).
[0025] Embodiments of the present invention use temporal analysis of external medical information sources to extract medical procedure context and operator intent for non-HMD wearers participating in the medical procedure. For example, mouse cursor extraction and monitoring, keyboard cursor extraction and text change detection processing, touchscreen inputs, or voice inputs provide a time correlated record of medical procedure context and operator intent activity to provide as inputs to the processing system.[00026J In various embodiments, the processing system improves its performance using one or more of feedback from clinical users, continued feedback from manual operator annotations, operator procedure report review, billing review, and reimbursement outcomeAtty Docket No.: 34237-64696 / WO evidence in an online or offline process. For example, the processing system, when annotations are manually added during report generation, uses temporally relevant operator intent and procedure context to further train machine learning models of the processing system for providing automated annotation intra-procedurally or during report generation. In some embodiments, the processing system may solicit specific training feedback or specific confirmation if an annotation should be used to further improve processing system performance. Responsive to receiving the feedback or confirmation, the processing system uses the annotation as training data for a machine learning model.[00027| In various embodiments, the processing system uses any combination of application programming interface (API) connections, audio monitoring, HMD sensor data, eye tracking sensor data, HMD video and the corresponding full-resolution 2D video streams that were displayed on 2D monitors to process medical procedure context and operator intent of HMD and non-HMD wearers. The processing system can use HMD-captured video and receive input from an operator consciously pointing using a visible head-pose gaze cursor, eye tracking, physical gestures, or verbal description to highlight specific medical procedure context information for emphasis.BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure (FIG.) 1 is an illustration of a processing system according to various embodiments.
[0029] FIG. 2 is a flowchart of a processing system sequence according to various embodiments.[00030| FIG. 3 illustrates sources of automated medical procedure context and operator intent using an HMD according to various embodiments.1 0031] FIG. 4 illustrates sources of automated medical procedure context and operator intent using external video capture according to various embodiments.
[0032] FIG. 5 is a diagram of a hands-free user interface of manual annotation according to various embodiments.
[0033] FIG. 6 is a diagram of a procedure report navigation and review interface according to various embodiments.
[0034] FIG. 7 illustrates a structure of an example neural network for a machine learning model of a processing system according to various embodiments.Atty Docket No.: 34237-64696 / WODETAILED DESCRIPTION
[0035] FIG.l is an illustration of a processing system 100 according to various embodiments. A data processing computer 110 communicates with one or more HMDs 120 using a wireless connection 130 (e.g., 802.11, 5G, etc.). Medical information sources connect to the processing system 100 through the data processing computer 110. Medical information sources include any number of video sources 140 providing medical images or conventional 2D monitor display outputs (e.g., VGA, DVI, HDMI, etc.). Other medical information sources, such as an Electronic Health Record (EHR) 150 or Electro- Anatomic Mapping System (EAMS) 160 may connect to the processing system 100 through an Application Programming Interface (API) over a computer connection (e.g., 802.3, RS232, USB, etc.). The processing system 100 additionally connects with an Intercom 170 through an audio interface (e.g., RCA, XLR, etc.) or digital API (e.g., USB, PCM, etc.). Any number of additional users 180 may optionally interact with a user interface, e.g., an HMD or intercom system interface, and operate external medical information source devices. One or more of the medical information sources may not be otherwise connected to the processing system 100 but may still be used to generate operator intent and medical procedure context through HMD image sensors, a microphone or phased array of microphones, or other means. The HMD wearer interacts with the medical information sources and the processing system 100 using during the medical procedure using a gaze-based interface 190. Medical information source data collected are stored and analyzed on the data processing computer 110. In some embodiments, the connection and processing may be performed on HMD 120 directly. In some embodiments, the processing may be performed on a network connected resource (e.g., one or more servers) in a cloud computing environment such as using a remote data center.
[0036] In some embodiments, an HMD 120 may not include a display (e.g., a head mounted earpiece). In some embodiments, the HMD 120 includes a monoscopic or stereoscopic display. In some embodiments, the HMD 120 includes one or more cameras with visible, infrared, or ultraviolet sensing. The HMD 120 can capture video of real-world objects in a procedure room including one or more of a patient, a physician’s hands, a medical instrument, other people, and a 2D display screen. In addition, the HMD 120 can capture video of virtual objects displayed by the HMD 120 a wearer of the HMD 120 including one or more of a 3D model, a virtual display screen, a textual annotation, an arrow annotation, a medical image, and a remote collaborator.Atty Docket No.: 34237-64696 / WO
[0037] In some embodiments, the HMD 120 includes one or more depth sensors including time-of-flight or structure-light projection varieties. In some embodiments, the HMD 120 includes one or more electroencephalogram (EEG) or electromyography (EMG) sensors used to sense a mental state of a user wearing the HMD 120. In some embodiments, the HMD 120 streams camera or audio data in real-time to another device and captures the camera or audio data locally. In some embodiments, the HMD 120 includes one or more light sources to increase quality of video captured by the HMD 120. In some embodiments, the HMD 120 includes one or more speakers or haptic vibration components to convey information to a user wearing the HMD 120.
[0038] FIG. 2 is a flowchart of a processing system 100 sequence according to various embodiments. FIG. 2 illustrates how Intraprocedural Collection 200 of Manual one or more Annotations 210 and one or more Automated Annotations 220 are processed within the processing system 100. Automated Annotations 220 reduce the likelihood of missed communication events that may not be captured manually due to human error. Any combination of annotations generated from manual input through the hands-free interface and annotations automatically generated by processing operator intent and medical procedure context are combined with the relevant medical procedure context to generate a candidate procedure report entry. This intraprocedural collection is repeated any number of times throughout the procedure to generate Medical Procedure Context 230. In some embodiments, the processing system 100 suggests automated annotations, e.g., based on Medical Procedure Context 230, to the operator and requests operator confirmation before recording the suggested annotation as a procedure report entry. In some embodiments, the candidate procedure report entries are presented post-procedure; the processing system 100 makes this presentation responsive to processing the entire procedure during Procedure Report Guidance 240 for the operator to review and provides feedback, e.g., on the relevance of automated annotation, report entries, and captured medical procedure context. After operator review, Automated Report Generation 250 generates a complete report of the entire procedure or partial report of a portion of the procedure. In various embodiments, the processing system 100 presents the report on a 2D desktop display, mobile tablet, mobile phone, large-format TV or boom display, or a HMD 120. Physician Review 260 and authorization for entry into the electronic health record and provides feedback and training on the medical accuracy of the generated report. In some instances, the physician removes entire annotations based on relevance to treatment. In other examples, the physician may add additional written detail or context to the generated report for a specific annotation. In another example, the physicianAtty Docket No.: 34237-64696 / WO may associate different contexts to a specific event as more appropriate evidence. Each of these examples of corrections are used by the processing system 100 as feedback to improve future associations between context in the report for the specific procedure and procedure events performed. Trained machine learning models (e.g., having the structure described with respect to FIG. 7) can take the context as input to generate outputs such as automated annotations for records of a medical procedure or relevant predictions. Further Billing Review 270 is optionally performed by medical provider coding specialists to ensure that adequate evidence for reported procedure events is included in the procedure report. In one example, a billing specialist may request the report include an additional image from an imaging system to provide adequate evidence for reimbursement. The processing system 100 uses this feedback to suggest this evidence to be automatically included for future reports at the institution, for the specific and similar procedures being performed. Reimbursement Feedback 280 provides feedback and training on the billing performance of the generated report responsive to billing metrics provided when reimbursement is adjudicated. This feedback provides an external, numeric scoring metric for performance of the system with respect to binary success or fail for each billing code, as well as overall success ratios for procedures. In some embodiments, feedback provided at each stage is provided as training data to a machine learning model online. In some embodiments, feedback is recorded and stored for future offline training of the machine learning model.
[0039] FIG. 3 illustrates sources of automated medical procedure context and operator intent using an HMD according to various embodiments. The processing system 100 uses the sources for automated annotation for a procedure report entry. An HMD wearer 300 communicating with other personnel 310 in the medical procedure can communicate a voice request 320 to perform an action on a medical device or system. The voice request 320 is detected by the HMD microphone and is processed using any combination of speech detection (e.g., automatic speech recognition of words spoken), speaker identification, and language processing to detect one or more characteristics of communication. In some embodiments, the processing system 100 uses the type and frequency of the communication as input features to determine if an event in progress should be recorded as an automated annotation. In some communications, the HMD wearer 300 requests specific procedure actions trained to be recognized as recordable events, (e.g., ablation start, stop, etc.), which may be used to record an annotation when other sources of this information are not available (e.g., EAMS not API enabled). The processing system 100 can also use cursor location and eye-gaze direction 330 detected by the HMD to determine the relevance of medicalAtty Docket No.: 34237-64696 / WO information source data for emphasis or scoring into the automated annotation. In some embodiments, audio data from other personnel 310 wearing HMDs is used to improve accuracy of speaker identification or speech processing. In some embodiments, the processing system 100 uses any combination of camera, gaze, and eye sensor data of other personnel 310 wearing HMDs to improve medical information source emphasis when receiving voice requests 320.
[0040] FIG. 4 illustrates sources of automated medical procedure context and operator intent using external video capture according to various embodiments. FIG. 4 illustrates an example of methods for automated medical procedure context an operator intent from a connected medical information source. On the captured display 400 (e.g., on a computer monitor or tablet display separate from an HMD) of the connected system, the processing system 100 identifies and monitors activity of a mouse cursor 420 and keyboard entry cursor 450 on the display 400. Frequency and content of activity from the mouse and keyboard is recorded as inputs for determining operator intent and procedure context. Additional procedure and connected display-specific training models are used to extract additional procedure context. In an embodiment, the processing system 100 identifies electrogram data 410 in the display 400, and monitors cursor activity controlling vertical cursor elements 430 and 440 to record electrogram interval measurement activity. In this embodiment, position of the mouse cursor 420 also provides intent information on the signals of interest for performing the interval measurement. The processing system 100 may be further configured to determine that electrogram measurement activity before ablation is more likely to be relevant evidence for performance of an electrophysiology study. In this embodiment, if the corresponding interval is associated with text entry, the contents of the text entry and the electrogram cursor screenshot with appropriate mouse cursor positions. The processing system 100 monitors text input entry (e.g., via a keyboard, touchscreen, or transcribed voice input) to record frequency and emphasis of specific recorded electrogram intervals. In other embodiments, the system is configured to identify reported ablation parameters (e.g., radiofrequency power, duration, impedance, pulse frequency, force, etc.) within a screen or display. These ablation specific parameters and events are extracted, counted, processed, and recorded to emphasize ablation events in the procedure. In this embodiment, in the processing system 100 creates report timeline annotations to note the number and characteristics of ablation events performed during the treatment period of the procedure.
[0041] FIG. 5 is a diagram of a hands-free user interface of manual annotation according to various embodiments. For example, a gaze-based interface enables hands-free userAtty Docket No.: 34237-64696 / WO interaction. During medical procedures, a physician may explicitly want to record the context of the procedure within the virtual environment 500 for a specific point in the procedure. The processing system 100 determines to activate a user interface element 530 to bookmark an annotation responsive to tracking a cursor controlled by the physician’s gaze. This triggers a time windowed annotation of captured display 510 (e.g., including what is currently shown on one or more displays A, B, C, D, and E) and integrated API data 520 (e.g., a 3D graphic) for review and export into a procedure report.
[0042] FIG. 6 is a diagram of a procedure report navigation and review interface according to various embodiments. FIG. 6 diagrams an example of a user interface for reviewing the activity or events of a medical procedure for identifying, modifying, or removing annotations from a procedure report. A timeline 600 of the procedure is displayed on the x-axis representing time, with frequency of automatically detected context activity displayed on the positive y-axis. Periods of high context activity on a specific information source display a corresponding graphic (e.g., a thumbnail or other visual indicator) to be included in the report, such as ultrasound 610 during needle access, electro-anatomic map 620 during periods of ablation, and electrograms 630 after a waiting period. On the negative y-axis, automatic and manual intent activity is displayed, identifying periods of high verbal intent activity 640 as well as specific manual annotations 650. The processing system 100 can train a machine learning model using user feedback on the procedure report. For example, the user indicates via selection of a graphic (610, 620, 630, 640, or 650) that a different source should be emphasized.
[0043] In various embodiments, the processing system 100 combines collected procedure data with text and image data from the patient and billing records (records of previous procedures, metrics of care-quality and outcomes, billing records and billing outcomes, previous electrogram recordings, previous ultrasound, x-ray and MRI images, etc.), and uses this data to both train and query machine learning models, such as large language models and small language models (LLM and SML), large video models and small video models (LVM and SVM). In some embodiments, the results of training and queries improve accuracy of system-specific processing such as improved electrogram processing for a specific system or physician workflow. In some embodiments, training improves accuracy of automatically generated reports and included automatic annotations, tailored to the specific workflows and billing practices of a specific institution or physician. In some embodiments, training feedback from multiple physicians within an institution or multiple institutions across a provider system improves consistency in procedure workflow and billing practices across aAtty Docket No.: 34237-64696 / WO provider system or network by consistently suggesting specific workflows and report generation.
[0044] In some embodiments, the processing system 100 may apply machine learning models trained from the medical or billing record to detect and rank automated annotations for the purpose of reducing the physician’s time to generate the report and / or the overall completeness or compliance of the report. The report reflects the ranking, e.g., by providing emphasis on higher-ranked annotations (automatically generated or manual) or other information included in the report. In various embodiments, the processing system 100 determines which particular 2D display (among multiple 2D displays) a physician is looking at, a position of the physician’s gaze direction on the 2D display, and content that was being displayed on the 2D display at that moment when the physician’s gaze was directed to it. Based on this information, the processing system 100 is better able to identify areas of interest or relevance than conventional systems (which examine content from only the HMD, only room-fixed cameras, or only the 2D video signals). The physician often makes verbal requests to assistants to enter data or manipulate interfaces on computer-based systems (e.g., annotating an electrogram signal, start or stopping the collection of EAMS mapping data, changing the depth, gain or frequency of an ultrasound scan, dictating a note in the patient record, changing a view parameter on a pre-op CT scan, etc.). In some embodiments, the keyboard and mouse inputs of those computer-based systems that are operated by the physician’s assistants are also recorded and time-synchronized. The processing system 100 correlates the pattern of mouse and keyboard inputs with the verbal audio requests (from the physician) that immediately precede (e.g., within a predetermined threshold amount of time in milliseconds) mouse and keyboard events. The processing system 100 can predict what mouse and keyboard events (or other types of user input events) follow verbal requests by the physician, and either suggest them to the human assistants or emit those mouse and keyboard signals directly. Thus, the processing system 100 enables the physician to more quickly train operators on specific workflows and operations, improve skill transferability, and reduce training burden or manual overhead. This in turn enables the integration of multiple clinical computer-based systems into a single system that requires fewer people to operate and less space within the procedure room.
[0045] In various embodiments, the processing system 100 automatically detects and flags automated annotations within the procedure of particular interest. Examples include: (1) when the physician spent significant time looking at a particular screen such as the ultrasound scans; (2) when the physician’s eye pupil dilation rapidly increasedAtty Docket No.: 34237-64696 / WO(physiologically indicating particular interest, attention, and mental effort); (3) when videos or images looked different from what is typically observed in procedures of a certain type; and (4) specific evidence to support appropriate billing. The processing system 100 also includes a real-time manual-annotation feature. The HMD wearer triggers (via a voice command, foot pedal or other interface) the processing system 100 to mark a specific timepoint in the collected audio and video data (e.g., “System: note the distance from this polyp here and this lesion here”). The processing system 100 includes a visible head-pose gaze-cursor so that the HMD user can consciously point at features and locations (on the 2D displays, such as a peak on a signal graph, or particular location in an X-ray image; or on other objects in the room, such as a particular instrument on the sterile tray, or a feature on the patient’s skin) and trains a machine learning model to account for the user’s one or more inputs. In some embodiments, the processing system 100 uses both eye-gaze tracking and a visible head-pose gaze-cursor: the processing system 100 uses eye-gaze tracking to confirm that the HMD wearer is actively looking at the visible head-pose gaze-cursor (i.e., they are looking where they are pointing) to reduce unintentionally directing the machine learning model’s attention to errant inputs.
[0046] The processing system 100 improves upon conventional methods of developing machine learning models by leveraging speech processing, intent detection, and activity detection to bootstrap generalized training with robust manual review and correction tools. Through use, training of a machine learning model to any combination of procedure-, physician-, or institution-specific workflows and practices is performed through physician and billing feedback.
[0047] In various embodiments, the processing system 100 includes a user-interface for reviewing and editing a medical report automatically generated by the processing system 100. This report-user-interface is linked to the video-review user-interface. The user (e.g., operator or other personnel) can highlight any item in the medical report and the processing system 100 highlights the corresponding evidence in the video-review user-interface. Likewise, the user can highlight annotations or timepoints in the video-review user-interface, and the processing system 100 in response to receiving this user input highlights those portions of the automatically generated medical report that correspond to the annotation or timepoint. The processing system 100 tracks the history of which report items and video annotations have been highlighted by the user, to help the user confirm that the medical report is thorough and accurate, thus mitigating the risk of machine-learning hallucinations being included in the submitted medical report. In some embodiments, the user wears theAtty Docket No.: 34237-64696 / WOHMD while reviewing the automatically generated transcription and medical report. The processing system 100 is thereby able to monitor which portions of the transcription and medical report the user has read, reviewed, and edited.
[0048] In some embodiments, the HMD includes a display that is visible by the HMD- wearer. The processing system 100 may display virtual objects (such as virtual 2D displays, 3D models of anatomy, text, arrows, or other graphics annotating real-world or virtual objects, etc.) that appear to the HMD-wearer to be fixed relative to their head (i.e., they move with the wearer’s head) or fixed relative to the procedure room (i.e., a virtual object that is position over a doorway appears to remain over the doorway no matter how the HMD-wearer moves within the room). In some embodiments, the processing system 100 receives input that a virtual object obstructs the view of the HMD-wearer of a target real -world object (e.g., the procedure-site on the patient) or another virtual object (e.g., one virtual 2D screen obscuring the view of another virtual 2D screen). When this happens, the processing system 100 can process a command from the HMD-wearer to move one or more virtual objects or otherwise modify the HMD display to reduce or eliminate the obstruction. The processing system 100 captures the history of these manual-move commands, along with the history of what real -world objects the HMD-wearer spends time looking at (measured using eye-gaze tracking) to predict which positionings of virtual object would cause obstructions. The processing system 100 trains a machine learning model to avoid those positionings in the future to reduce the likelihood of unwanted obstructions in a HMD display.
[0049] In various embodiments, the history of what real -world and virtual objects the HMD-wearer spends significant time looking at (e.g., exceeding a threshold amount of time over a certain time duration), timing patterns of gaze, and the distances between those real- world objects are analyzed by the processing system 100 in order to suggest, to clinical staff, rearrangements of the physical equipment and virtual objects within a procedure room to reduce the eye motion of the physician during the medical procedure. This can help reduce physician fatigue and human error.
[0050] In some embodiments, the processing system 100 captures and analyzes the history of how the HMD-wearer interacts with the displayed virtual objects (e.g., menus, buttons, modes, etc.). The processing system 100 uses this data to train a machine learning model to predict future user interactions with virtual objects.[00051 j In some embodiments, the processing system 100 compresses HMD-captured video streams to reduce transmission bandwidth and storage space. High-efficiency encoding techniques for compression use image-entropy (such as AVI, H.265, AVC, etc.).Atty Docket No.: 34237-64696 / WOIn various embodiments, the HMD includes sensors to determine a video camera’s translational and rotational velocity. These signals are used by the compression algorithm to predict how the image changes between subsequent frames, thereby reducing the computational resources and time needed to perform compression. The HMD may also use sensor data to determine where the user is looking and / or pointing, and the processing system 100 may selectively encode the video such that encoding-decoding degradation is reduced in those regions where the user is looking and / or pointing. In other words, a technical advantage of the processing system 100 is the ability to prioritize the highest quality (or higher quality) video signal for where the user is looking and / or pointing, and permit more degradation in other regions. This reduces the bandwidth to transfer video and storage space based on priority of different regions of the video.
[0052] FIG. 7 illustrates a structure of an example neural network for a machine learning model of a processing system 100 according to various embodiments. The neural network 700 may receive inputs 710 and generate an output 720. While inputs 710 is graphically illustrated as having two dimensions in FIG. 7, the inputs 710 may be in any dimension. For example, the neural network 700 may be a one-dimensional convolutional network.
[0053] The neural network 700 may include different kinds of layers, such as convolutional layers 730, pooling layers 740, recurrent layers 750, fully connected layers 760, and custom layers 770. A convolutional layer 730 convolves the input of the layer (e.g., a matrix of any dimension ) with one or more weight kernels to generate different types of sequences that are filtered by the kernels to generate feature spaces. Each convolution result may be associated with an activation function. A convolutional layer 730 may be followed by a pooling layer 740 that selects the maximum value (max pooling) or average value (average pooling) from the portion of the input covered by the kernel size. The pooling layer 740 reduces the spatial size of the extracted features. In some embodiments, a pair of convolutional layer 730 and pooling layer 740 may be followed by a recurrent layer 750 that includes one or more feedback loops 755. The feedback 755 may be used to account for spatial relationships of the features in an image or temporal relationships in sequences. The layers 730, 740, and 750 may be followed in multiple fully connected layers 760 that have nodes (represented by squares in FIG. 7) connected to each other. The fully connected layers 760 may be used for classification and object detection. In one embodiment, one or more custom layers 770 may also be presented for the generation of a specific format of output 720. For example, a custom layer may be used for image segmentation for labeling pixels of an image input with different segment labels.Atty Docket No.: 34237-64696 / WO
[0054] The order of layers and the number of layers of the neural network 700 in FIG. 7 is for example only. In various embodiments, a neural network 700 includes one or more convolutional layer 730 but may or may not include any pooling layer 740 or recurrent layer 750. If a pooling layer 740 is present, not all convolutional layers 730 are always followed by a pooling layer 740. A recurrent layer may also be positioned differently at other locations of the neural network. For each convolutional layer 730, the sizes of kernels (e.g., 1x1, 1x2, 3x3, 5x5, 7x7, NxM, where N or M = 1,2,3, . . . , etc.) and the numbers of kernels allowed to be learned may be different from other convolutional layers 730.
[0055] A machine learning model may include certain layers, nodes, kernels and / or coefficients. Training of the neural network 700 may include forward propagation and backpropagation. Each layer in a neural network may include one or more nodes, which may be fully or partially connected to other nodes in adjacent layers. In forward propagation, the neural network performs the computation in the forward direction based on outputs of a preceding layer. The operation of a node may be defined by one or more functions. The functions that define the operation of a node may include various computation operations such as convolution of data with one or more kernels, pooling, recurrent loop in RNN, various gates in LSTM, etc. The functions may also include an activation function that adjusts the weight of the output of the node. Nodes in different layers may be associated with different functions.
[0056] Each of the functions in the neural network may be associated with different coefficients (e.g., weights and kernel coefficients) that are adjustable during training. In addition, some of the nodes in a neural network may also be associated with an activation function that decides the weight of the output of the node in forward propagation. Common activation functions may include step functions, linear functions, sigmoid functions, hyperbolic tangent functions (tanh), and rectified linear unit functions (ReLU). After input is provided into the neural network and passes through a neural network in the forward direction, the results may be compared to the training labels or other values in the training set to determine the neural network’s performance. The process of prediction may be repeated for other inputs in the training sets to compute the value of the objective function in a particular training round. In turn, the neural network performs backpropagation by using gradient descent such as stochastic gradient descent (SGD) or other optimization techniques to adjust the coefficients in various functions to improve the value of the objective function.
[0057] Multiple rounds of forward propagation and backpropagation may be performed. Training may be completed when the objective function has become sufficiently stable (e.g.,Atty Docket No.: 34237-64696 / WO the machine learning model has converged) or after a predetermined number of rounds for a particular set of training samples. The trained machine learning model can be used for performing various machine learning tasks as discussed in this disclosure.
[0058] The foregoing description of the embodiments of the invention has been presented for the purpose of illustration; it is not intended to be exhaustive or to limit the invention to the precise forms disclosed. Persons skilled in the relevant art can appreciate that many modifications and variations are possible in light of the above disclosure.
[0059] Some portions of this description describe the embodiments of the invention in terms of algorithms and symbolic representations of operations on information. These algorithmic descriptions and representations are commonly used by those skilled in the data processing arts to convey the substance of their work effectively to others skilled in the art. These operations, while described functionally, computationally, or logically, are understood to be implemented by computer programs or equivalent electrical circuits, microcode, or the like. Furthermore, it has also proven convenient at times, to refer to these arrangements of operations as modules, without loss of generality. The described operations and their associated modules may be embodied in software, firmware, hardware, or any combinations thereof.
[0060] Any of the steps, operations, or processes described herein may be performed or implemented with one or more hardware or software modules, alone or in combination with other devices. In one embodiment, a software module is implemented with a computer program product including a non-transitory computer-readable storage medium storing instructions (e.g., computer program code), which can be executed by one or more computer processors for performing any or all of the steps, operations, or processes described.
[0061] Embodiments of the invention may also relate to a product that is produced by a computing process described herein. Such a product may include information resulting from a computing process, where the information is stored on a non-transitory, tangible computer readable storage medium and may include any embodiment of a computer program product or other data combination described herein.
[0062] Finally, the language used in the specification has been principally selected for readability and instructional purposes, and it may not have been selected to delineate or circumscribe the inventive subject matter. It is therefore intended that the scope of the invention be limited not by this detailed description, but rather by any claims that issue on an application based hereon. Accordingly, the disclosure of the embodiments of the invention isAtty Docket No.: 34237-64696 / WO intended to be illustrative, but not limiting, of the scope of the invention, which is set forth in the following claims.
Claims
Atty Docket No.: 34237-64696 / WOWhat is claimed is:
1. A method comprising: receiving sensor data from a head-mounted device (HMD) during a medical procedure; determining a gaze direction of a wearer of the HMD using the sensor data; determining that the gaze direction is directed to a content of a plurality of content displayed for the medical procedure; determining context for the medical procedure based on the content; and generating, using a trained machine learning model taking the context as input, an annotation for a record of the medical procedure.
2. The method of claim 1, further comprising: receiving an input from the wearer of the HMD using additional sensor data from the HMD; generating a manual annotation using the input; and further training the trained machine learning model using the manual annotation.
3. The method of claim 2, further comprising: determining one or more words spoken by the wearer of the HMD by processing the input; and responsive to determining that the one or more words are associated with the medical procedure, generating the manual annotation to include the one or more words.
4. The method of any of claims 2-3, further comprising: determining a ranking of a plurality of annotations including at least the manual annotation and the annotation generated using the trained machine learning model, wherein the record of the medical procedure reflects the ranking.
5. The method of claim 4, wherein the record of the medical procedure is a timeline including graphical representations of the manual annotation and the annotation generated using the trained machine learning model, and wherein the graphical representations are displayed with emphasis to reflect the ranking.
6. The method of any of claims 4 or 5, wherein the plurality of annotations further includes an additional annotation generated using input from another user different than the wearer of the HMD.
7. The method of any of claims 1-6, further comprising: determining one or more words spoken by the wearer of the HMD by processing the input; andAtty Docket No.: 34237-64696 / WO responsive to determining that the one or more words are associated with the medical procedure, generating a manual annotation to include the one or more words.
8. The method of any of claims 1-7, further comprising: providing a virtual graphic for the medical procedure for display by the HMD; identifying 2D content for the medical procedure displayed by a device separate from the HMD; and wherein determining that the gaze direction is directed to the content of the plurality of content displayed for the medical procedure comprises determining that the gaze direction is directed to the virtual graphic or the 2D content.
9. The method of any of claims 1-8, further comprising: requesting, from the wearer of the HMD, confirmation of the annotation generated using the trained machine learning model; and responsive to receiving the confirmation from the wearer of the HMD, including the annotation in the record of the medical procedure.
10. The method of any of claims 1-9, further comprising: receiving, from the wearer of the HMD, a modification of the annotation generated using the trained machine learning model; and further training the trained machine learning model using the modification.
11. The method of any of claims 1-10, further comprising: providing electrogram data for the medical procedure for display by the HMD; determining a measurement of the electrogram data; and wherein the trained machine learning model generates the annotation by taking the measurement as input.
12. The method of any of claims 1-11, further comprising: determining that the gaze direction is directed to a virtual content displayed by the HMD; responsive to determining that the virtual content is at least partially obstructed from a point of view of the wearer of the HMD, modifying a position of the virtual content; and further training the trained machine learning model using the modified position of the virtual content.
13. The method of any of claims 1-10, wherein the trained machine learning model is trained using previous inputs from the wearer of the HMD and from other medical procedures.Atty Docket No.: 34237-64696 / WO14. A non-transitory computer-readable storage medium storing instructions that when executed by one or more processors cause the one or more processors to perform steps of any of the methods of claims 1-13.
15. A system comprising: a head-mounted device (HMD); and a non-transitory computer-readable storage medium storing instructions that when executed by one or more processors cause the one or more processors to perform steps of any of the methods of claims 1-13.
Citation Information
Patent Citations
System and method for interactive event timeline
US20190164633A1
Domain-specific human-model collaborative annotation tool
US20220222952A1
Electrogram Annotation System
US20230000419A1
Annotation data collection using gaze-based tracking
US20230266819A1
Two-way communication between head-mounted display and electroanatomic system
US20230341932A1