Methods and systems for generating data from a surgical procedure using generative artificial intelligence

The system employs a generative large language model to automatically generate detailed surgical notes from image data, addressing the issue of incomplete documentation and enhancing reimbursement by providing accurate and comprehensive procedural records.

WO2025122673A1PCT designated stage expired Publication Date: 2025-06-12ACTIV SURGICAL INC +1

Patent Information

Application Number
PCT/US2024/058546
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-05
Filing Date
2024-12-04
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

The increasing burden of documentation in surgical procedures leads to incomplete and inaccurate recording, potentially resulting in medical errors and decreased reimbursement for hospitals and physicians.

Method used

A system and method using a multimodal, generative large language model connected to a visual encoder to automatically generate written descriptions of anatomy, pathology, and physiology observed during surgical procedures, based on image data and multi-turn question and answer dialogue.

Benefits of technology

Improves documentation accuracy and completeness, potentially reducing medical errors and increasing reimbursement by providing detailed, real-time descriptions of surgical procedures, including reasoning for specific steps and potential complications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024058546_12062025_PF_FP_ABST
    Figure US2024058546_12062025_PF_FP_ABST
Patent Text Reader

Abstract

A method for annotating data generated from a surgical procedure may include: providing image data of a surgical procedure; implementing a large language and vision assistant, wherein the large language and vision assistant comprises a large language model and an image processing model; and using the large language and vision assistant to generate a written description of anatomy, pathology, and / or physiology observed during the surgical procedure.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND SYSTEMS FOR GENERATING DATA FROM A SURGICAL PROCEDURE USING GENERATIVE ARTIFICIAL INTELLIGENCECROSS-REFERENCE[1] This application claims the benefit of U.S. Provisional Application No. 63 / 606,533, filed December 5, 2023, which application is incorporated herein by reference in its entirety for all purposes.BACKGROUND[2] Paperwork may comprise routine work involving written documents such as forms, records, or letters. While it is clear that appropriate documentation (written and / or electronic) is critical to patient safety, continuity of care, and both individual and system quality improvement, it seems that the working lives of surgeons are increasingly filled with documentation tasks that move physicians farther away from high-quality bedside patient care. The increasing burden of documentation incentivizes incomplete and inaccurate recording preceding medical events. This can result in medical errors and denials of insurance coverage damaging patient health and financial wellbeing.SUMMARY[3] Systems and methods of the present disclosure help to ameliorate at least some of the above identified drawbacks by populating image snapshots and a corresponding written description of complex and abnormal anatomy, pathology and physiology observed during a surgical case. This is particularly useful when unforeseen intraoperative complications or delays occur, and the reimbursement opportunity substantially increases (2-3X) for hospitals and physicians. In an example, laparoscopic cholecystectomy reimbursements may be improved by an average of $8,000 per procedure and may avoid bile duct injuries and visceral and vascular injuries and may increase reimbursements consistent with the level of complication when they do. In another example, laparoscopic colorectal surgery reimbursements may be improved by an average of $21,000 per procedure and may avoid bowl injury, obstruction, perforation, ischemia, fistulae, and hemorrhage events and may increase reimbursements consistent with the level of complication when they do.[4] In an aspect, the present disclosure provides a method for annotating data generated from a surgical procedure. The method may comprise providing image data of a surgical procedure, wherein the image data is based at least in part on one or more frames of video data of the surgical procedure; implementing a multimodal, generative large language model, wherein the multimodal, generative large language model is connected to a visual encoder; and generating awriten description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedure based at least in part on the multimodal, generative large language model, wherein the written description is based at least on multi-turn question and answer dialogue with the multimodal, generative large language model.[5] In some embodiments, the written description comprises reasoning for particular surgical steps or deduction of next steps. In some embodiments, the multimodal, generative large language model is trained based at least in part on a library of annotations of surgical procedures. In some embodiments, the library of annotations of surgical procedures comprises action triplets. In some embodiments, the action triplets comprise an indication of an instrument, a verb, and a target in a surgical scene. In some embodiments, the action triplets are further associated with a surgical phase. In some embodiments, the multimodal, generative large language model is finetuned based on a ground truth dataset comprising ground truth image data and the action triplets.[6] In some embodiments, the image data comprises one or more members selected from the group consisting of white light imaging, fluorescence imaging, and perfusion imaging. In some embodiments, the multimodal, generative large language model is a LLAVA, LLAVA-Med, or LLAVA Surg. In some embodiments, the visual encoder is configured to (i) receive the one or more frames of video data obtained using one or more image and / or video acquisition modules and transform the one or more frames of video data into an input into the multimodal, generative large language model. In some embodiments, the multimodal, generative large language model is configured (ii) generate the written description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedure based on (A) the one or more frames of video data and (B) one or more training data sets comprising surgical or medical data associated with a patient or a surgical procedure. In some embodiments, the one or more training data sets comprise anatomical and physiological data obtained using a plurality of imaging modalities.[7] In some embodiments, the written description is generated in substantially real time with the surgical procedure. In some embodiments, the annotations comprise billing coding of surgical procedures. In some embodiments, the multimodal, generative large language model is a multimodal video to text model. In some embodiments, In some embodiments, the method further comprises annotating one or more frames of the image data with action triplets based at least in part on an output of the multimodal, generative large language model.[8] In another aspect, the present disclosure provides a method for annotating data generated from a surgical procedure. The method may comprise providing image data of a surgical procedure, wherein the image data is based at least in part on one or more frames of video data of the surgical procedure; implementing a multimodal, generative large language model, whereinthe multimodal, generative large language model is connected to a visual encoder, wherein the large language model is trained based at least in part on a library of annotations of surgical procedures; and generating a written description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedure based at least in part on the multimodal, generative large language model, and wherein the written description comprises reasoning for particular surgical steps or deduction of next steps.[9] In some embodiments, the written description is based at least on multi-turn question and answer dialogue with the large language and vision assistant. In some embodiments, the written description comprises reasoning for particular surgical steps or deduction of next steps. In some embodiments, the library of annotations of surgical procedures comprises action triplets. In some embodiments, the action triplets comprise an indication of an instrument, a verb, and a target in a surgical scene. In some embodiments, the action triplets are further associated with a surgical phase. In some embodiments, the multimodal, generative large language model is fine-tuned based on a ground truth dataset comprising ground truth image data and the action triplets.

[0010] In some embodiments, the method further comprises using the multimodal, generative large language model to annotate frames of the image data with action triplets. In some embodiments, the image data comprises one or more members selected from the group consisting of white light imaging, fluorescence imaging, and perfusion imaging. In some embodiments, the multimodal, generative large language model is LLAVA, LLAVA-Med, or LLAVA Surg. In some embodiments, the visual encoder is configured to (i) receive the one or more frames of video data obtained using one or more image and / or video acquisition modules and transform the one or more frames of video data into an input into the multimodal, generative large language model. In some embodiments, the multimodal, generative large language model is configured (ii) generate the written description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedure based on (A) the one or more frames of video data and (B) one or more training data sets comprising surgical or medical data associated with a patient or a surgical procedure. In some embodiments, the one or more training data sets comprise anatomical and physiological data obtained using a plurality of imaging modalities.

[0011] In some embodiments, the written description is generated in substantially real time with the surgical procedure. In some embodiments, the annotations comprise billing coding of surgical procedures. In some embodiments, the multimodal, generative large language model is a multimodal video to text model. In some embodiments, the annotations comprise billing coding of surgical procedures.

[0012] In another aspect, the present disclosure provides a system for annotating data generated from a surgical procedure. The system may comprise an imaging module configured to collectimage data of the surgical procedure, wherein the image data is based at least in part on one or more frames of video data of the surgical procedure; and a processor configured to: implement a multimodal, generative large language model, wherein the multimodal, generative large language model is connected to a visual encoder; and generate a written description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedure based at least in part on the multimodal, generative large language model, wherein the written description is based at least on multi-turn question and answer dialogue with the multimodal, generative large language model.

[0013] In some embodiments, the processor is further configured to annotate one or more frames of the image data with action triplets based at least in part on an output of the multimodal, generative large language model. In some embodiments, the image data comprises one or more members selected from the group consisting of white light imaging, fluorescence imaging, and perfusion imaging. In some embodiments, the multimodal, generative large language model is a LLAVA, LLAVA-Med, or LLAVA Surg. In some embodiments, the visual encoder is configured to (i) receive the one or more frames of video data obtained using one or more image and / or video acquisition modules and transform the one or more frames of video data into an input into the multimodal, generative large language model. In some embodiments, the multimodal, generative large language model is configured (ii) generate the written description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedure based on (A) the one or more frames of video data and (B) one or more training data sets comprising surgical or medical data associated with a patient or a surgical procedure. In some embodiments, the one or more training data sets comprise anatomical and physiological data obtained using a plurality of imaging modalities. In some embodiments, the one or more image and / or video acquisition modules comprise (i) a first image and / or video acquisition module configured to capture images and / or videos using a first imaging modality and (ii) a second image and / or video acquisition module configured to capture images and / or videos using a second imaging modality.

[0014] In some embodiments, the one or more medical predictions or assessments comprise an identification or a classification of one or more critical structures. In some embodiments, the written description comprises a location of one or more critical structures. In some embodiments, the written description is configured provide the surgical operator with predictive clinical decision support before and during a surgical procedure.

[0015] In some embodiments, the multimodal, generative large language model is configured to update or refine the written description in real time based on additional image data obtained during a surgical procedure. In some embodiments, the one or more training data sets comprisemedical data associated with one or more reference surgical procedures. In some embodiments, the one or more training data sets comprise medical data associated with (i) one or more critical phases or scenes of a surgical procedure or (ii) one or more views of a critical structure that is visible or detectable during the surgical procedure. In some embodiments, the one or more training data sets comprise medical data obtained using a laparoscope, a robot assisted imaging unit, or an imaging sensor configured to generate anatomical images and / or videos in a red- green-blue visual spectrum. In some embodiments, the one or more training data sets comprise medical data obtained using an imaging sensor configured to generate physiologic or functional imaging data based on chemical signal enhancers. In some embodiments, the chemical signal enhancers comprise ICG, fluorescent, or radiolabeled dyes. In some embodiments, the one or more training data sets comprise medical data obtained using an imaging sensor configured to generate physiologic or functional imaging data based on laser speckle patterns. In some embodiments, the one or more training data sets comprise medical data obtained using an imaging sensor configured to generate physiologic or functional imaging data, wherein said physiologic or functional imaging data comprises two or more co-registered images or videos. In some embodiments, the one or more training data sets comprise medical data obtained using an imaging sensor configured to generate physiologic or functional imaging data, wherein said physiologic or functional imaging data comprises segmented anatomic positions of critical structures of surgical interest. In some embodiments, the one or more training data sets comprise medical data obtained using a plurality of imaging modalities that enable direct objective correlation between physiologic images or videos and corresponding RGB images or videos, wherein the plurality of imaging modalities comprises depth imaging, distance mapping, intraoperative cholangiograms, angiograms, ductograms, ureterograms, or lymphangiograms. In some embodiments, the training data set comprises perfusion data. In some embodiments, the image data obtained using one or more image and / or video acquisition modules comprises perfusion data.

[0016] In another aspect, the present disclosure provides a system for annotating data generated from a surgical procedure. The system may comprise an imaging module configured to collect image data of the surgical procedure, wherein the image data is based at least in part on one or more frames of video data of the surgical procedure; and a processor configured to: implement a multimodal, generative large language model, wherein the multimodal, generative large language model is connected to a visual encoder, wherein the large language model is trained based at least in part on a library of annotations of surgical procedures; and generate a written description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedurebased at least in part on the multimodal, generative large language model, and wherein the written description comprises reasoning for particular surgical steps or deduction of next steps.

[0017] In some embodiments, the processor is further configured to annotate one or more frames of the image data with action triplets based at least in part on an output of the multimodal, generative large language model. In some embodiments, the image data comprises one or more members selected from the group consisting of white light imaging, fluorescence imaging, and perfusion imaging. In some embodiments, the multimodal, generative large language model is a LLAVA, LLAVA-Med, or LLAVA Surg. In some embodiments, the visual encoder is configured to (i) receive the one or more frames of video data obtained using one or more image and / or video acquisition modules and transform the one or more frames of video data into an input into the multimodal, generative large language model. In some embodiments, the multimodal, generative large language model is configured (ii) generate the written description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedure based on (A) the one or more frames of video data and (B) one or more training data sets comprising surgical or medical data associated with a patient or a surgical procedure. In some embodiments, the one or more training data sets comprise anatomical and physiological data obtained using a plurality of imaging modalities. In some embodiments, the one or more image and / or video acquisition modules comprise (i) a first image and / or video acquisition module configured to capture images and / or videos using a first imaging modality and (ii) a second image and / or video acquisition module configured to capture images and / or videos using a second imaging modality.

[0018] In some embodiments, the one or more medical predictions or assessments comprise an identification or a classification of one or more critical structures. In some embodiments, the written description comprises a location of one or more critical structures. In some embodiments, the written description is configured provide the surgical operator with predictive clinical decision support before and during a surgical procedure.

[0019] In some embodiments, the multimodal, generative large language model is configured to update or refine the written description in real time based on additional image data obtained during a surgical procedure. In some embodiments, the one or more training data sets comprise medical data associated with one or more reference surgical procedures. In some embodiments, the one or more training data sets comprise medical data associated with (i) one or more critical phases or scenes of a surgical procedure or (ii) one or more views of a critical structure that is visible or detectable during the surgical procedure. In some embodiments, the one or more training data sets comprise medical data obtained using a laparoscope, a robot assisted imaging unit, or an imaging sensor configured to generate anatomical images and / or videos in a red-green-blue visual spectrum. In some embodiments, the one or more training data sets comprise medical data obtained using an imaging sensor configured to generate physiologic or functional imaging data based on chemical signal enhancers. In some embodiments, the chemical signal enhancers comprise ICG, fluorescent, or radiolabeled dyes. In some embodiments, the one or more training data sets comprise medical data obtained using an imaging sensor configured to generate physiologic or functional imaging data based on laser speckle patterns. In some embodiments, the one or more training data sets comprise medical data obtained using an imaging sensor configured to generate physiologic or functional imaging data, wherein said physiologic or functional imaging data comprises two or more co-registered images or videos. In some embodiments, the one or more training data sets comprise medical data obtained using an imaging sensor configured to generate physiologic or functional imaging data, wherein said physiologic or functional imaging data comprises segmented anatomic positions of critical structures of surgical interest. In some embodiments, the one or more training data sets comprise medical data obtained using a plurality of imaging modalities that enable direct objective correlation between physiologic images or videos and corresponding RGB images or videos, wherein the plurality of imaging modalities comprises depth imaging, distance mapping, intraoperative cholangiograms, angiograms, ductograms, ureterograms, or lymphangiograms. In some embodiments, the training data set comprises perfusion data. In some embodiments, the image data obtained using one or more image and / or video acquisition modules comprises perfusion data.

[0020] In another aspect, the present disclosure provides a method of image-to-text generation. The method may comprise providing one or more images of a surgical procedure to a generative artificial intelligence (Al) model, wherein the one or more images are based at least in part on a video of the surgical procedure; and generating text associated with the one or more images using the generative Al model, wherein the text comprises a description of one or more intraoperative scenes in the one or more images, wherein the written description comprises reasoning for particular surgical steps or deduction of next steps, and wherein the written description is based at least on multi -turn question and answer dialogue with the generative Al model.

[0021] In some embodiments, the description comprises a descriptor of type(s) of surgical tools used, and a manner in which the surgical tools are being used during the surgical procedure. In some embodiments, the description comprises a descriptor of one or more anatomical features within the one or more intraoperative scenes, and changes to the one or more anatomical features during the surgical procedure. In some embodiments, the description comprises a descriptor of one or more intraoperative complications that occurred during the surgical procedure.

[0022] In some embodiments, wherein the method further comprises identifying one or more billing codes from a plurality of billing codes for the surgical procedure, based at least in part on the text comprising the description of the one or more intraoperative scenes. In some embodiments, the one or more billing codes are identified at an accuracy of greater than about 80%. In some embodiments, the generative Al model comprises a large language model (LLM). In some embodiments, the generative Al model is used for processing visual data and / or text from the one or more images.

[0023] In some embodiments, the generative Al model automatically generates the text in response to one or more questions provided via a graphical interface. In some embodiments, the one or more questions are provided by one or more users via the graphical interface. In some embodiments, the one or more questions are automatically generated based at least in part on a stored library of questions associated with the surgical procedure. In some embodiments, the one or more questions are selected from a plurality of questions presented to one or more users via the graphical interface.

[0024] In some embodiments, the one or more images are annotated. In some embodiments, the one or more images are non-annotated. In some embodiments, the one or images are extracted from a video of the surgical procedure. In some embodiments, the one or images are taken at different points in time during the surgical procedure. In some embodiments, the one or images are taken at different locations and / or perspectives during the surgical procedure. In some embodiments, the one or images are taken using two or more different imaging modalities. In some embodiments, the two or more different imaging modalities comprise at least one of white light imaging or fluorescence imaging.

[0025] In some embodiments, the one or images comprises a plurality of images captured at different wavelengths. In some embodiments, the one or more images comprises perfusion data. In some embodiments, the surgical procedure comprises a laparoscopic procedure. In some embodiments, the generative Al model is trained based at least in part on a library of annotations of surgical procedures.

[0026] In another aspect, the present disclosure provides a system for video-to-text generation. The system may comprise a processor configured to: receive one or more images of a surgical procedure and direct the one or more images to a generative artificial intelligence (Al) model, wherein the one or more images are based at least in part on a video of the surgical procedure; and implement the generative Al model to automatically generate text associated with the one or more images, wherein the text comprises a description of one or more intraoperative scenes in the one or more images, wherein the written description comprises reasoning for particular surgicalsteps or deduction of next steps, and wherein the written description is based at least on multiturn question and answer dialogue with the large language and vision assistant.

[0027] In some embodiments, the generative Al model is trained based at least in part on a library of annotations of surgical procedures. In some embodiments, the description comprises a descriptor of type(s) of surgical tools used, and a manner in which the surgical tools are being used during the surgical procedure. In some embodiments, the description comprises a descriptor of one or more anatomical features within the one or more intraoperative scenes, and changes to the one or more anatomical features during the surgical procedure. In some embodiments, the description comprises a descriptor of one or more intraoperative complications that occurred during the surgical procedure. In some embodiments, the method further comprises: identifying one or more billing codes from a plurality of billing codes for the surgical procedure, based at least in part on the text comprising the description of the one or more intraoperative scenes. In some embodiments, the one or more billing codes are identified at an accuracy of greater than about 80%.

[0028] In some embodiments, the generative Al model comprises a large language model (LLM). In some embodiments, the generative Al model is used for processing visual data and / or text from the one or more images. In some embodiments, the generative Al model automatically generates the text in response to one or more questions provided via a graphical interface. In some embodiments, the one or more questions are provided by one or more users via the graphical interface. In some embodiments, the one or more questions are automatically generated based at least in part on a stored library of questions associated with the surgical procedure. In some embodiments, the one or more questions are selected from a plurality of questions presented to one or more users via the graphical interface.

[0029] In some embodiments, the one or more images are annotated. In some embodiments, the one or more images are non-annotated. In some embodiments, the one or images are extracted from a video of the surgical procedure. In some embodiments, the one or images are taken at different points in time during the surgical procedure. In some embodiments, the one or images are taken at different locations and / or perspectives during the surgical procedure. In some embodiments, the one or images are taken using two or more different imaging modalities. In some embodiments, the two or more different imaging modalities comprise at least one of white light imaging or fluorescence imaging.

[0030] In some embodiments, the one or images comprises a plurality of images captured at different wavelengths. In some embodiments, the one or more images comprises perfusion data. In some embodiments, the surgical procedure comprises a laparoscopic procedure.

[0031] In another aspect, the present disclosure provides, a method for annotating data generated from a surgical procedure. The method may comprise: providing image data of a surgical procedure; implementing a large language and vision assistant, wherein the large language and vision assistant comprises a large language model and an image processing model; and using the large language and vision assistant to generate a written description of anatomy, pathology, and / or physiology observed during the surgical procedure.

[0032] In some embodiments, the method further comprises using the image processing module to annotate frames of the image data with action triplets. In some embodiments, the image data comprises one or more members selected from the group consisting of white light imaging, fluorescence imaging, and perfusion imaging. In some embodiments, the large language model is a generative Al. In some embodiments, the image processing model comprises a surgical guidance system comprising an image processing module configured to (i) receive image or video data obtained using one or more image and / or video acquisition modules and (ii) generate one or more medical predictions or assessments based on (a) the image or video data or physiological data associated with the image or video data and (b) one or more training data sets comprising surgical or medical data associated with a patient or a surgical procedure, wherein the one or more training data sets comprise anatomical and physiological data obtained using a plurality of imaging modalities; and a visualization module configured to provide a surgical operator with an enhanced view of a surgical scene, based on the one or more augmented data sets.

[0033] In another aspect, the present disclosure provides a system for annotating data generated from a surgical procedure. The system may comprise: an imaging module configured to collect image data of the surgical procedure; and a processor configured to: implement a large language and vision assistant, wherein the large language and vision assistant comprises a large language model and an image processing model, and use the large language and vision assistant to generate a written description of anatomy, pathology, and / or physiology observed during the surgical procedure.

[0034] In some embodiments, the processor further configured to use the image processing module to annotate frames of the image data with action triplets. In some embodiments, the image data comprises one or more members selected from the group consisting of white light imaging, fluorescence imaging, and perfusion imaging. In some embodiments, the large language model is a generative Al. In some embodiments, the image processing model comprises a surgical guidance system comprising an image processing module configured to (i) receive image or video data obtained using one or more image and / or video acquisition modules and (ii) generate one or more medical predictions or assessments based on (a) the image or video data orphysiological data associated with the image or video data and (b) one or more training data sets comprising surgical or medical data associated with a patient or a surgical procedure, wherein the one or more training data sets comprise anatomical and physiological data obtained using a plurality of imaging modalities; and visualization module configured to provide a surgical operator with an enhanced view of a surgical scene, based on the one or more augmented data sets.

[0035] In some embodiments, the one or more image and / or video acquisition modules comprise (i) a first image and / or video acquisition module configured to capture images and / or videos using a first imaging modality and (ii) a second image and / or video acquisition module configured to capture images and / or videos using a second imaging modality.

[0036] In some embodiments, the one or more medical predictions or assessments comprise an identification or a classification of one or more critical structures. In some embodiments, the one or more medical predictions or assessments comprise a predicted location of one or more critical structures. In some embodiments, the one or more medical predictions or assessments provide the surgical operator with predictive clinical decision support before and during a surgical procedure. In some embodiments, the one or more medical predictions or assessments are generated or updated in part based on an output of one or more computer vision algorithms. In some embodiments, the one or more medical predictions or assessments comprise medical information or live guidance that is provided to a user or an operator through one or more notifications or message banners. In some embodiments, the one or more medical predictions or assessments comprise tissue viability for one or more tissue regions in the surgical scene.

[0037] In some embodiments, the image processing module is configured to update or refine the one or more medical predictions or assessments in real time based on additional image data obtained during a surgical procedure. In some embodiments, the image processing module is configured to update or refine a predicted location of one or more critical structures in real time based on additional image data obtained during a surgical procedure.

[0038] In some embodiments, the one or more training data sets comprise medical data associated with one or more reference surgical procedures. In some embodiments, the one or more training data sets comprise medical data associated with (i) one or more critical phases or scenes of a surgical procedure or (ii) one or more views of a critical structure that is visible or detectable during the surgical procedure. In some embodiments, the one or more training data sets comprise medical data obtained using a laparoscope, a robot assisted imaging unit, or an imaging sensor configured to generate anatomical images and / or videos in a red-green-blue visual spectrum. In some embodiments, the one or more training data sets comprise medical data obtained using an imaging sensor configured to generate physiologic or functional imaging databased on chemical signal enhancers. In some embodiments, the chemical signal enhancers comprise ICG, fluorescent, or radiolabeled dyes. In some embodiments, the one or more training data sets comprise medical data obtained using an imaging sensor configured to generate physiologic or functional imaging data based on laser speckle patterns. In some embodiments, the one or more training data sets comprise medical data obtained using an imaging sensor configured to generate physiologic or functional imaging data. The physiologic or functional imaging data may comprise two or more co-registered images or videos. In some embodiments, the one or more training data sets comprise medical data obtained using an imaging sensor configured to generate physiologic or functional imaging data. The physiologic or functional imaging data may comprise segmented anatomic positions of critical structures of surgical interest. In some embodiments, the one or more training data sets comprise medical data obtained using a plurality of imaging modalities that enable direct objective correlation between physiologic images or videos and corresponding RGB images or videos. The plurality of imaging modalities may comprise depth imaging, distance mapping, intraoperative cholangiograms, angiograms, ductograms, ureterograms, or lymphangiograms. In some embodiments, the training data set comprises perfusion data.

[0039] In some embodiments, the image processing module is configured to update or refine the one or more medical predictions or assessments in real time based on additional physiologic data obtained during a surgical procedure. In some embodiments, the visualization module is configured to display and track a position and an orientation of one or more medical tools or critical features in real time. In some embodiments, the visualization module is configured to display a visual outline of the one or more critical features corresponding to predicted contours or boundaries of the critical features. In some embodiments, the image data obtained using one or more image and / or video acquisition modules comprises perfusion data. In some embodiments, the visualization module is configured to display a visual outline indicating a predicted location of one or more critical features before, during, or after the surgical operator performs one or more steps of a surgical procedure. In some embodiments, the visualization module is configured to autonomously or semi-autonomously track mobile and deformable tissue targets as tissue maneuvers are performed during a surgical procedure.

[0040] In some embodiments, the image processing module comprises a data multiplexer configured to combine anatomical and / or physiological data obtained using the one or more image and / or video acquisition modules. In some embodiments, the image processing module further comprises one or more feature extractors on spatial or temporal domains. The one or more feature extractors may be trained using the one or more training data sets and configured to extract one or more features from the combined anatomical and / or physiological data. In someembodiments, the one or more feature extractors comprise a spatial feature extractor configured to detect spatial features and generate a feature set for every image frame. In some embodiments, the spatial features comprise textures, colors, or edges of one or more critical structures. In some embodiments, the one or more feature extractors comprise a temporal feature extractor configured to detect a plurality of temporal features or feature sets over a plurality of image frames. In some embodiments, the temporal features correspond to changes in contrast, perfusion, or perspective.

[0041] In some embodiments, the image processing module further comprises a view classifier configured to use the one or more extracted features to determine a current surgical view relative to the surgical scene. In some embodiments, the image processing module further comprises a tissue classifier configured to use the one or more extracted features to identify, detect, or classify one or more tissues. In some embodiments, the image processing module further comprises a phase detector configured to use the one or more extracted features to determine a surgical phase and generate guidance based on the surgical phase. In some embodiments, the image processing module further comprises a critical structure detector configured to locate and identify one or more critical features based on (i) the one or more tissues identified, detected, or classified using the tissue classifier, (ii) a current surgical view determined using the view classifier, and (iii) the surgical phase determined using the phase detector. In some embodiments, the image processing module further comprises an augmented view generator configured to display guidance and metrics associated with the surgical procedure, based on the one or more critical features located and identified using the critical structure detector.

[0042] In another aspect, a method of image-to-text generation is provided. The method may comprise: providing one or more images of a surgical procedure to a generative artificial intelligence (Al) model; and using the generative Al model to automatically generate text associated with the one or more images, wherein the text comprises a description of one or more intraoperative scenes in the one or more images.

[0043] In some embodiments, the description comprises a descriptor of type(s) of surgical tools used, and a manner in which the surgical tools are being used during the surgical procedure. In some embodiments, the description comprises a descriptor of one or more anatomical features within the one or more intraoperative scenes, and changes to the one or more anatomical features during the surgical procedure. In some embodiments, the description comprises a descriptor of one or more intraoperative complications that occurred during the surgical procedure. In some embodiments, the method further comprises identifying one or more billing codes from a plurality of billing codes for the surgical procedure, based at least in part on the text comprising the description of the one or more intraoperative scenes. In some embodiments, the one or morebilling codes are identified at an accuracy of greater than about 80%. In some embodiments, the generative Al model comprises a large language model (LLM). In some embodiments, the generative Al model is used for processing visual data and / or text from the one or more images.

[0044] In some embodiments, the generative Al model automatically generates the text in response to one or more questions provided via a graphical interface. In some embodiments, the one or more questions are provided by one or more users via the graphical interface. In some embodiments, the one or more questions are automatically generated based at least in part on a stored library of questions associated with the surgical procedure. In some embodiments, the one or more questions are selected from a plurality of questions presented to one or more users via the graphical interface.

[0045] In some embodiments, the one or more images are annotated. In some embodiments, the one or more images are non-annotated. In some embodiments, the one or images are extracted from a video of the surgical procedure. In some embodiments, the one or images are taken at different points in time during the surgical procedure. In some embodiments, the one or images are taken at different locations and / or perspectives during the surgical procedure. In some embodiments, the one or images are taken using two or more different imaging modalities. In some embodiments, the two or more different imaging modalities comprise at least one of white light imaging or fluorescence imaging. In some embodiments, the one or images comprises a plurality of images captured at different wavelengths. In some embodiments, the one or more images comprises perfusion data. In some embodiments, the surgical procedure comprises a laparoscopic procedure.

[0046] Another aspect of the present disclosure provides a non-transitory computer readable medium comprising machine executable code that, upon execution by one or more computer processors, graphical processing units and / or digital signal processors, implements any of the methods above or elsewhere herein.

[0047] Another aspect of the present disclosure provides a system comprising one or more computer processors and computer memory coupled thereto. The computer memory comprises machine executable code that, upon execution by the one or more computer processors, implements any of the methods above or elsewhere herein.

[0048] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure.Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.INCORPORATION BY REFERENCE

[0049] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also “Figure” and “FIG.” herein), of which:

[0051] FIG. l is a flowchart of an example of a method for annotating data generated from a surgical procedure, in accordance with some embodiments.

[0052] FIG. 2 is a flowchart of another example of a method for annotating data generated from a surgical procedure, in accordance with some embodiments.

[0053] FIG. 3 is a flowchart of an example of a method for method of image-to-text generation, in accordance with some embodiments.

[0054] FIG. 4 is a box diagram of an example of a system for annotating data generated from a surgical procedure of the present disclosure, in accordance with some embodiments.

[0055] FIG. 5 illustrates an overlay of three types of imaging data: white light, at the back, fluorescence in the middle, and perfusion imaging at the front, in accordance with some embodiments.

[0056] FIG. 6 is a box diagram of a structure of multi-modal generative model, in accordance with some embodiments.

[0057] FIG. 7 is a flowchart of a method of training a multi-modal generative model of the present disclosure, in accordance with some embodiments.

[0058] FIG. 8A illustrates an image of a surgical video; a caption for the image; and an expert generated multi-tum question and answer for the training dataset.

[0059] FIG. 8B illustrates an image of a surgical video; a caption for the image; and a GPT-4 generated multi-tum question and answer for the training dataset.

[0060] FIG. 8C illustrates another image of a surgical video; a caption for the image; and a GPT-4 generated multi-turn question and answer for the training dataset.

[0061] FIG. 8D illustrates another image of a surgical video; a caption for the image; and a GPT-4 generated multi-turn question and answer for the training dataset.

[0062] FIG. 9A illustrates six examples of generated action triplets for individual frames of a laparoscopic cholecystectomy.

[0063] FIG. 9B illustrates an example of how the prompt to the Al may be modified to call a particular action triplet in the data.

[0064] FIG. 10A illustrates an output from a multi-modal generative Al of the present disclosure.

[0065] FIG. 10B illustrates an example multi turn question and answer exchange with a model as disclosed herein.

[0066] FIG. 10C is a flow chart of a method of the present disclosure, in accordance with some embodiments.

[0067] FIG. 10D is a schematic of a system of the present disclosure, in accordance with some embodiments.

[0068] FIG. 11 illustrates three Al models which may be used in concert with systems and methods of the present disclosure.

[0069] FIG. 12 illustrates a computer system that is programmed or otherwise configured to implement systems and methods of the present disclosure.

[0070] FIG. 13 illustrates a non-limiting example of testing GPT models with a triplet dataset.

[0071] FIG. 14 is a comparison of VILA 1.5 with LLAVA for triplet evaluation and phase evaluation on a representative dataset of surgical images.DETAILED DESCRIPTION

[0072] While various embodiments of the invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.

[0073] Whenever the term “at least,” “greater than,” or “greater than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “at least,” “greater than” or “greater than or equal to” applies to each of the numerical values in that series of numerical values. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.

[0074] Whenever the term “no more than,” “less than,” or “less than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “no more than,” “less than,” or “less than or equal to” applies to each of the numerical values in that series of numerical values. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.

[0075] Certain inventive embodiments herein contemplate numerical ranges. When ranges are present, the ranges include the range endpoints. Additionally, every sub range and value within the range is present as if explicitly written out.

[0076] The term “about” or “approximately” may mean within an acceptable error range for the particular value, which will depend in part on how the value is measured or determined, e g., the limitations of the measurement system. For example, “about” may mean within 1 or more than 1 standard deviation, per the practice in the art. Alternatively, “about” may mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Where particular values are described in the application and claims, unless otherwise stated the term “about” meaning within an acceptable error range for the particular value may be assumed.

[0077] The term “real time” or “real-time,” as used interchangeably herein, generally refers to an event (e.g., an operation, a process, a method, a technique, a computation, a calculation, an analysis, a visualization, an optimization, etc.) that is performed using recently collected or received data. In some cases, a real time event may be performed almost immediately or within a short time span, such as within at least 0.0001 millisecond (ms), 0.0005 ms, 0.001 ms, 0.005 ms, 0.01 ms, 0.05 ms, 0.1 ms, 0.5 ms, 1 ms, 5 ms, 0.01 seconds, 0.05 seconds, 0.1 seconds, 0.5 seconds, 1 second, or more. In some cases, a real time event may be performed almost immediately or within a short enough time span, such as within at most 1 second, 0.5 seconds, 0.1 seconds, 0.05 seconds, 0.01 seconds, 5 ms, 1 ms, 0.5 ms, 0.1 ms, 0.05 ms, 0.01 ms, 0.005 ms, 0.001 ms, 0.0005 ms, 0.0001 ms, or less. Real-time may also be real time and may not indicate an instantaneous processing.Annotating Surgical Data

[0078] Disclosed herein are systems and methods for annotating surgical data. The present disclosure generally relates to systems and methods for providing written descriptions of surgical scenes using generative artificial intelligence (Al). More specifically, the present disclosure relates to systems and methods for generating surgical notes relating to images extracted from surgical video. The surgical notes may be useful for reducing time spent by clinician generating reports of a particular procedure. The notes may be then better correlated to classifications of a difficulty of particular procedure.

[0079] FIG. l is a flowchart of an example of a method 100 for annotating data generated from a surgical procedure. At an operation 110, method 100 may comprise providing image data of a surgical procedure. The image data may be based at least in part on one or more frames of video data of the surgical procedure. At an operation 120, method 100 may comprise implementing a multimodal, generative large language model. The multimodal, generative large language model may be connected to a visual encoder. At an operation 130, method 100 may comprise generating a written description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedure based at least in part on the multimodal, generative large language model. The written description may be based at least on multi -turn question and answer dialogue with the multimodal, generative large language model.

[0080] FIG. 2 is a flowchart of an example of a method 200 for annotating data generated from a surgical procedure. At an operation 210, method 200 may comprise providing image data of a surgical procedure. The image data may be based at least in part on one or more frames of video data of the surgical procedure. At an operation 220, method 200 may comprise implementing a multimodal, generative large language model. The multimodal, generative large language model may be connected to a visual encoder. The large language model may be trained based at least in part on a library of annotations of surgical procedures. At an operation 230, method 200 may comprise generating a written description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedure based at least in part on the multimodal, generative large language model. The written description may comprise reasoning for particular surgical steps or deduction of next steps.

[0081] FIG. 3 is a flowchart of an example of a method 300 for method of image-to-text generation. At an operation 310, method 300 may comprise providing one or more images of a surgical procedure to a generative artificial intelligence (Al) model. The one or more images may be based at least in part on a video of the surgical procedure. At an operation 320, method 300 may comprise generating text associated with the one or more images using the generative Al model. In some cases, the text comprises a description of one or more intraoperative scenes in the one or more images. In some cases, the written description comprises reasoning for particular surgical steps or deduction of next steps. In some cases, the written description is based at least on multi -turn question and answer dialogue with the generative Al model.

[0082] FIG. 4 is a box diagram of an example 400 of a system for annotating data generated from a surgical procedure of the present disclosure. System 400 may comprise one or more imaging modules 410. The one or more imaging modules may be configured to collect image data of the surgical procedure. The image data may be based at least in part on one or more frames of video data of the surgical procedure. System 400 may comprise one or more processors420. A processor 420 may comprise any embodiment, variation, or example described herein with respect to the section “Computer Systems” herein.

[0083] For example, a processor may be configured to: implement a multimodal, generative large language model, such as any multi-modal, generative large language model herein. The multimodal, generative large language model may be connected to a visual encoder. For example, a processor may be configured to generate a written description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedure based at least in part on the multimodal, generative large language model. The written description may be based at least on multi-tum question and answer dialogue with the multimodal, generative large language model.

[0084] For example, a processor may be configured to: implement a multimodal, generative large language model. The multimodal, generative large language model may be connected to a visual encoder. The large language model may be trained based at least in part on a library of annotations of surgical procedures. For example, a processor may be configured to generate a written description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedure based at least in part on the multimodal, generative large language model. The written description comprises reasoning for particular surgical steps or deduction of next steps.

[0085] In some cases, systems for annotating data of the present disclosure may be implemented without an imaging module. For example, the processor may instead be configured to receive one or more images of a surgical procedure and direct the one or more images to a generative artificial intelligence (Al) model.

[0086] The present disclosure provides systems for video-to-text generation. A system may comprise a processor configured to receive one or more images of a surgical procedure and direct the one or more images to a generative artificial intelligence (Al) model. The one or more images may be based at least in part on a video of the surgical procedure, system may comprise a processor configured to implement the generative Al model to automatically generate text associated with the one or more images. The text may comprise a description of one or more intraoperative scenes in the one or more images. In some cases, the written description comprises reasoning for particular surgical steps or deduction of next steps. In some cases, the written description is based at least on multi-tum question and answer dialogue with the large language and vision assistant. In some cases, the generative Al model comprises a large language model (LLM). In some cases, the generative Al model is used for processing visual data and / or text from the one or more images.Data

[0087] Systems and method of the present disclosure may leverage machine learning models which integrate data of more than one modality. A machine learning model that leverages data from more than one modality may be a multi-modal model. For example, data modalities may comprise images, videos, text, audio, speech, numerical data, etc. It may be important to develop datasets for model training and model fine tuning which correspond to the modalities used when the model is used in an inference mode. For example, method and systems of the present disclosure may take in surgical videos and output a text based on the surgical videos. Thus, it may be important to develop a library of surgical videos and corresponding annotations which are similar to the annotations that one wishes for the model to generate.

[0088] Each of these various types of data may be data herein. Data may contain information useful for carrying out a process (such as training data) or resulting from a process (such as information collected from a sensor, a camera, an output of a machine learning model, statistics). Data used for training a machine learning model may comprise an input. Data used for training a machine learning model may comprise an expected output. Data may be an image, a video, text or any other format or modality suitable for a machine learning model. Data may be multiple modalities (such as image and text). Data for machine learning may comprise multiple fields. Data for machine learning may be tabular.

[0089] Preprocessing - Normalization - Data may be normalized. Data normalization may comprise scaling. Scaling may comprise linear scaling. Scaling may comprise converting floating-point values to a standard range (such as 0-1, or -1 to +1). Scaling may comprise z-score scaling. Z-score scaling may use a probability distribution. Z-score scaling may transform data values to their standard deviations (z-score). Data may be log scaled. Log scaling may be used to scale data that conforms to the power law. Values of data may be counted. The counts of data may have a distribution where the smaller values of the data a much greater count than larger values of the data. Log scaling may transform the data values to a log value. Data may be scaled using clipping. Clipping may cap the data to be within a given value range. Values greater the given range may be assigned the greatest possible value within that range. Values smaller than the given range may be assigned the smallest possible value within that range.

[0090] Preprocessing - Whitening - The data may be whitened. Whitening may be a linear transformation. Whitening may comprise using a covariance matrix. Whitening may decorrelate data. Whitening may zero the center of data. Whitening may apply an eigenvalue. Whitening may apply an eigenvector. Whitening may comprise principal component analysis. Whitening may comprise zero component analysis.

[0091] Preprocessing - Binning / Bucketing - Data may be binned. Binning may also be called bucketing. Binning may subdivide the data into ranges. Ranges may be of a fixed size. Ranges may be variable size. Ranges may be predetermined. Ranges may be determined based upon values of the data. Ranges may be based upon quantiles. Data within a range of a bin may be grouped into that bin. Binned data may be further processed. Data may be discretized. Data may be smoothed.

[0092] Data splitting - Data may be split into different sets. One such set may be a training set that is used during the training phase as input to the model. Optionally, other sets may be made such as a validation set which may be used during training time to provide an indication of the model’s performance on previously untrained data, and / or a test set which may be used after training is complete to test the trained model at inference time.

[0093] Data may be from one or more sources (such as proteomic, genomic, epigenomic, metabolomic, clinical, image, text, and / or coordinates). Data may be from one or more sources. Data may be from multiple sources (or modalities). Data may be from a single source (or modality). Data may be raw. Data may be processed. Data may be preprocessed. Data may be filtered.Image Data

[0094] At an operation 110, method 100 may comprise providing image data of a surgical procedure. At an operation 210, method 200 may comprise providing image data of a surgical procedure. The image data may be based at least in part on one or more frames of video data of the surgical procedure. At an operation 310, method 300 may comprise providing one or more images of a surgical procedure to a generative artificial intelligence (Al) model. The one or more images may be based at least in part on a video of the surgical procedure. Further, system 400 may comprise a processor 420 configured to receive data.

[0095] In some cases, the one or images are extracted from a video of the surgical procedure. For example a video of a surgical procedure may comprise a number of frames of image data. In some cases, individual frames or collections of multiple frames may be used as input data.

[0096] The data may relate to images of a surgical procedure. For example, the images may comprise videos of appendectomy, mastectomy, hysterectomy, colectomy, coronary artery bypass, open heart surgery, cataract surgery, cesarean section, gallbladder removal, hip replacement, knee replacement, angioplasty, gastric bypass, kidney transplant, liver transplant, lung transplant, heart transplant, spinal fusion, lumpectomy, prostatectomy, vasectomy, tubal ligation, thyroidectomy, rhinoplasty, tonsillectomy, breast augmentation, liposuction, face lift, bariatric surgery, hernia repair, skin grafting, hair transplant, Lasik eye surgery, dental implant surgery, hand surgery, appendectomy, cholecystectomy, craniotomy, adenoidectomy, heart valverepair, hemorrhoidectomy, brain tumor surgery, plastic and reconstructive surgery, neck dissection, ovarian cyst removal, sinus surgery, tympanoplasty, carpal tunnel release surgery, laparoscopic surgery, transurethral resection of the prostate, etc.

[0097] In some cases, the surgical procedure may comprise a laparoscopic procedure. For example, the data may relate to a video of laparoscopic cholecystectomy. For example, the images may comprise videos of laparoscopic: appendectomy, cholecystectomy, hernia repair, hysterectomy, colectomy, prostatectomy, gastric bypass, bariatric surgery, nephrectomy, liver biopsy, spleen removal (splenectomy), adrenalectomy, anti-reflux surgery (fundoplication), cystectomy, pyloroplasty, pancreatectomy, oophorectomy, myomectomy, tubal ligation, endometrial ablation, hysterectomy, salpingectomy, ovarian drilling, lymph node biopsy, lobectomy, etc.

[0098] In some cases, the one or images are taken at different points in time during the surgical procedure. For example, a surgical procedure may be broken up into surgical phases. Each surgical phase may be mapped to a particular subset of images. For example, a surgical phase may be determined by when a particular tool is in use, when a particular action is being taken with a particular tool, a particular response of the tissue, etc. A surgical phase may comprise an opening phase, a cutting phase, a cleaning phase, a closing phase, etc.

[0099] In some cases, the one or images are taken at different locations and / or perspectives during the surgical procedure. For example, a surgery may be performed with one or more imaging devices, thereby providing a plurality of different views from each of several imaging devices. For example, a single imaging device may be moved thereby providing a plurality of different views from a single imaging device.

[0100] FIG. 5 shows an overlay of three types of imaging data: white light, at the back, fluorescence in the middle, and perfusion imaging at the front. In some cases, the image data comprises one or more members selected from the group consisting of white light imaging, fluorescence imaging, and perfusion imaging. In some cases, the one or images are taken using two or more different imaging modalities. In some cases, the two or more different imaging modalities comprise at least one of white light imaging or fluorescence imaging.

[0101] In some cases, the one or more images comprises white light data. White light imaging may be a name for imaging with a visible color spectrum. A white light image may comprise an image taken with a continuous or semicontinuous spectrum within a visible range. For example, white light may be between 400 and 700 nm. A white light source may comprise a tungsten bulb. A white light source may comprise a multicolor LED with a diffuser to create a continuous spectrum. White light imaging may be a relatively standard imaging modality which may bridge the gap between perfusion imaging (e g., ActivSight) and more generally applicable applications.

[0102] In some cases, the one or more images comprises fluorescence data. Florescence imaging may comprise a type of imaging based on the ability of a fluorophore in the image to fluoresce in response to an excitation. Fluorescence imaging may comprise an excitation beam. An excitation beam may be absorbed by a fluorophore in the tissue. The fluorophore may then emit light in response. The emitted light fluorescence may be at a lower energy / longer wavelength than the excitation beam. Fluorescence imaging may help show critical structures (e.g., blood vessels, oxygenated tissues, etc.) outside of the capability of human sight and may be specifically tuned to anatomy. For example, a fluorophore may be injected into a patient blood stream in order to allow for imaging of the fluorescing tissue. Fluorophore may comprise indocyanine green, other cyanine derivatives, riboflavin, fluorescein, etc.

[0103] In some cases, the one or more images comprises perfusion data. Perfusion imaging may be used to show real-time blood flow and tissue perfusion. In some cases, perfusion imaging may comprise laser speckle contrast imaging. In a laser speckle contrast image, speckles are produced when coherent light scattered back from biological tissue. Mobile scatterers cause the speckle pattern to blur. In biological tissue, red blood cells may be useful scatters. Thus, where blood is flowing, a laser speckle contrast signal may be produced. Laser speckle imaging may be non- invasive. Imaging that shows where blood is flowing may help clearly indicate the presence of a blood vessel, which may otherwise be hard to see, and it may help allow a practitioner to see when a blood vessel has been adequately clamped before performing a procedure. This data may be useful for a broad range of surgical procedures. In some cases, a coherent light source is provided to generate the laser speckle data. The coherent light source may be an 800 nm light source. The coherent light sources may be a longer than 700 nm light source.

[0104] In some cases, the one or images comprises a plurality of images captured at different wavelengths. For example, an image may be captured within a white light region (for white light data) and within an emission region for fluorescence data. For example, an image may be captured within a white light region (for white light data) and at the wavelength of the scattered light for perfusion data.

[0105] The image data based at least in part on one or more frames of video data of the surgical procedure may be supplemented by other forms of data. For example, a particular frame data may be coded with a time stamp, an indication of the type of procedure, an indication of the patient or particular instance of the procedure, etc.Machine Learning

[0106] Machine learning may be used to implement the systems and methods of the present disclosure. For example, at an operation 120, method 100 may comprise implementing a multimodal, generative large language model. The multimodal, generative large language modelmay be connected to a visual encoder. For example, at an operation 220, method 200 may comprise implementing a multimodal, generative large language model. The multimodal, generative large language model may be connected to a visual encoder. The large language model may be trained based at least in part on a library of annotations of surgical procedures. For example, at an operation 320, method 300 may comprise generating text associated with the one or more images using the generative Al model. In some cases, the text comprises a description of one or more intraoperative scenes in the one or more images. In some cases, the written description comprises reasoning for particular surgical steps or deduction of next steps. In some cases, the written description is based at least on multi-turn question and answer dialogue with the generative Al model.

[0107] As used herein, machine learning may refer to various functions, logic, and / or algorithms to teach a model to adapt, modify, and refine its decision-making capabilities as the model is exposed to new data. In some cases, a model or rule set may be built and used to predict a result based on values of one or more features. One or more computing devices may be used to implement machine learning techniques and methods to build and / or train one or more models, functions or algorithms from an example training data set of input observations in order to make data-driven predictions or decisions expressed as outputs based on known properties learned from a training data set (rather than strictly following static programming instructions).

[0108] Machine learning may comprise a learning phase (i.e., a training phase) and an inference phase. During a training phase, one or more training data sets may be presented to the computing device for classification. In some cases, the one or more training data sets may comprise medical images and / or videos taken using a plurality of different imaging modalities. The actual result produced by the computing device may be compared against a known correct result, in some cases with reference to a function. A known correct result may be, e.g., a result that is predetermined to be correct by an expert in the field or based on evidence or collective agreement or concordance between anatomic and physiologic data.

[0109] One objective of the training phase is to minimize discrepancies between known correct results and outputs by the computing device based on the one or more training data sets. Results from an output of the computing device may then be used to adjust certain parameters of the function and the computing device, in such a way that if another training data set were presented to the computing device another time, the computing device would theoretically produce a different output consistent with known correct results. Training of a computing device using machine learning methods may be complete when subsequent test data is presented to the computing device, the computing device generates an output on that test data, and a comparisonbetween the output and known correct results yields a difference or value that is within a predetermined acceptable margin.

[0110] During the inference phase, one or more models trained during the training phase may be used to classify previously unclassified data. Such classifications may be performed automatically based on data input during a supervised portion of the machine learning process (i.e., the training phase). In the inference phase, medical data (e.g., live surgical anatomic and physiologic images and / or videos) may be parsed in accordance with a model developed during the training phase. Once the medical data has been parsed, the model may be used to classify one or more features contained or represented within the medical data. In some embodiments, this classification may be verified using supervised classification techniques (i.e., an editor may validate the classification). This classification may also be used as additional training data that is input into the training phase and used to train a new model and / or refine a previous model.

[0111] During the inference phase, medical data such as medical images and / or videos may be provided to the computing device. The computing device may be programmed to identify possible regions of interest, having been trained with a plurality of different training data sets comprising various medical images and / or videos. In some cases, the computing device may scan the medical images and / or videos, retrieve features from the medical images and / or videos, extract values from the medical images and / or videos, and match the values to predetermined values that the computing device has been programmed to recognize as being associated with various critical anatomic and physiologic features.

[0112] Types of training - A machine learning model such as those disclosed here may comprise hyperparameters (such as layer size, number of layers, choice of optimizer, learning rate, etc.), parameters (such as weights, biases, or coefficients), one or more processing steps (such as layers), and may produce one or more outputs and have one or more inputs. Hyperparameters may be optimized, in a process called hyperparameter optimization but may be set during training may not change. Parameters may be changed during training. During training the machine learning model may calculate a loss useful for calculating the error between the real output of the model and the expected output of the model (for example, labels). A loss may measure a portion of the model, such as information in the model and / or the learned distribution of samples. Some set of the model parameters may be updated based at least in part on the loss calculation. The model may perform multiple rounds, or epochs, of training wherein an input or set of inputs is given and processed by the model which may produce an output or set of outputs which may then the basis for updating the weights. The updated weights may be used in the next epoch. Some training may comprise more steps. Training may occur in different environments such as supervised, unsupervised, semi-supervised, self-supervised or some combination thereof.

[0113] In a supervised environment the expected output may be provided for each input during training. The training data may have labels associated with each sample of the training data. The labels are an indication of the desired output of the model when the corresponding input is given.

[0114] In an unsupervised method the training set does not have corresponding labels. In some cases the input is the desired output of the model and may be used in place of a label. In other cases, the desired output is communicated through a score which may be related to some other output indication.

[0115] The model may be trained using self-supervised learning (SSL). SSL may use no labels. SSL may use some labels. Self-supervised methods may generate implicit labels from the unstructured data. In SSL, tasks may fall into two categories: pretext tasks and downstream tasks. In a pretext task, SSL may be used to train an Al system to leam meaningful representations of unstructured data. Those learned representations can be subsequently used as input to a downstream task, like a supervised learning task or reinforcement learning task. The reuse of a pre-trained model on a new task is referred to as “transfer learning.”

[0116] SSL may be used in the training of a diverse array of sophisticated deep learning architectures for a variety of tasks, from transformer-based large language models (LLMs) like BERT and GPT to image synthesis models like variational autoencoders (VAEs) and generative adversarial networks (GANs) to computer vision models like SimCLR and Momentum Contrast (MoCo). These methods may use other types of learning such as semi-supervised learning, supervised learning, and / or unsupervised learning.

[0117] Semi supervised may combine unsupervised and supervised tasks by using labeled and unlabeled data. In some cases, there may be datasets where some samples are labeled and others are not. In these cases, it may be desirable to have a fully labeled dataset but producing labels for large datasets is time consuming and expensive. Semi supervised learning first trains on the labeled data of the set of data and may then be used to produce pseudo-labels, or labels that are not validated.

[0118] Labels / Ground Truth - Labels may be in various forms. Labels may be in a continuous range, for example 0 to 1. A label may use a confidence threshold. A confidence value may be associated with a label. A confidence value above a confidence threshold may be used along with the labeled data to retrain the model to improve the overall performance of the model. A label may be binary. Labels may be ordinal. Labels may be cardinal. Labels may be discrete. Labels may be vectors. Labels may be scalars. Labels may be incomplete (e.g., not all labels are present).

[0119] Classification - A machine learning model may be trained as a classifier. A classifier may perform multiclass classification where more than one class is indicated. A classifier may be amulticlass multilabel, where more than one class may be output as present at one time. This may be useful in settings where classes may co-exist in the input. For example, an image segmentation model or object detection model may indicate the presence of multiple objects in an image and output an indication in its output for each of the detected objects. This may also be useful when the model is used to detect either multiple classes in the input and / or where some other label is desired such as a contextual output.

[0120] Regression - A machine learning model may be trained as a regression model. A regression model may be used in a predictive fashion, whereas a classifier is used to place input or portions of input into classes that are predefined. Regression models may take an input and output a continuous value as a prediction or forecast score. As an example, a regression model may take an image and predict a desired set of values describing a shape of a new object to be placed in the image. In this example the output, or a portion of the output, of a regression may be used as an input to another model.

[0121] Once training is completed, a model may be used to infer on a set of inputs. The model output may be the desired output for the use of the model or there may be some portion of the model that is used for a desired output different than the output that was used during training time. At inference time the model’s weights may be static.Machine Learning Algorithms

[0122] One or more machine learning algorithms may be used to implement the systems and methods of the present disclosures. The machine learning algorithm may be, for example, an unsupervised learning algorithm, supervised learning algorithm, or a combination thereof. The unsupervised learning algorithm may be, for example, clustering, hierarchical clustering, k- means, mixture models, DBSCAN, OPTICS algorithm, anomaly detection, local outlier factor, neural networks, autoencoders, deep belief nets, Hebbian learning, generative adversarial networks, self-organizing map, expectation-maximization algorithm (EM), method of moments, blind signal separation techniques, principal component analysis, independent component analysis, non-negative matrix factorization, singular value decomposition, or a combination thereof. In some embodiments, the supervised learning algorithm may be, for example, support vector machines, linear regression, logistic regression, linear discriminant analysis, decision trees, k-nearest neighbor algorithm, neural networks, similarity learning, or a combination thereof.

[0123] In some embodiments, the machine learning algorithm may comprise a deep neural network (DNN). In other embodiments, the deep neural network may comprise a convolutional neural network (CNN). The CNN may be, for example, U-Net, ImageNet, LeNet-5, AlexNet, ZFNet, GoogleNet, VGGNet, ResNetl8, or ResNet, etc. In some cases, the neural network maybe, for example, a deep feed forward neural network, a recurrent neural network (RNN), LSTM (Long Short Term Memory), GRU (Gated Recurrent Unit), Auto Encoder, variational autoencoder, adversarial autoencoder, denoising auto encoder, sparse auto encoder, Boltzmann machine, RBM (Restricted BM), deep belief network, generative adversarial network (GAN), deep residual network, capsule network, or attention / transformer networks, etc. In some embodiments, the neural network may comprise neural network layers. The neural network may have at least about 2 to 1000 or more neural network layers. In some cases, the machine learning algorithm may be, for example, a random forest, a boosted decision tree, a classification tree, a regression tree, a bagging tree, a neural network, or a rotation forest.

[0124] The machine learning algorithms may be used to generate one or more medical models for predictive visualizations and / or predictive support during a surgical procedure. In some cases, the one or more medical models may be trained using neural networks or convolutional neural networks. In some cases, the one or more medical models may be trained using deep learning. In some cases, the deep learning may be supervised, unsupervised, and / or semi-supervised. In some cases, the one or more medical models may be trained using reinforcement learning and / or transfer learning. In some cases, the one or more medical models may be trained using image thresholding and / or color-based image segmentation. In some cases, the one or more medical models may be trained using clustering. In some cases, the one or more medical models may be trained using regression analysis. In some cases, the one or more medical models may be trained using support vector machines. In some cases, the one or more medical models may be trained using one or more decision trees or random forests associated with the one or more decision trees. In some cases, the one or more medical models may be trained using dimensionality reduction. In some cases, the one or more medical models may be trained using one or more recurrent neural networks. In some cases, the one or more recurrent neural networks may comprise a long short-term memory neural network. In some cases, the one or more medical models may be trained using one or more classical algorithms. The one or more classical algorithms may be configured to implement exponential smoothing, single exponential smoothing, double exponential smoothing, triple exponential smoothing, Holt-Winters exponential smoothing, autoregressions, moving averages, autoregressive moving averages, autoregressive integrated moving averages, seasonal autoregressive integrated moving averages, vector autoregressions, or vector autoregression moving averages.Large Language Model

[0125] In some cases, a generative Al model comprises a large language model (LLM). Any of the methods or systems disclosed herein may include, incorporate, or utilize a large language model (LLM). For example, at an operation 120, method 100 may comprise implementing amultimodal, generative large language model. The multimodal, generative large language model may be connected to a visual encoder. For example, at an operation 220, method 200 may comprise implementing a multimodal, generative large language model. The multimodal, generative large language model may be connected to a visual encoder. In some cases, the multi modal model may also comprise three components the base LLM, the vision encoder (such as CLIP or BLIP herein) and a linear or multilayer perceptron layer(s).

[0126] For example, a generative artificial intelligence (Al) model may include an LLM. An LLM, such as LLM 610 may be a deep learning algorithm that can recognize, summarize, translate, predict and / or generate text and / or other content based on knowledge gained from massive datasets. A LLM may be trained on large sets of data. For example, training sets for training the LLM may include greater than 500,000 images and / or videos containing intraoperative scenes of a variety of different surgical procedures. The training sets may be drawn from diverse sets of data such as medical images (captured using white light imaging, fluorescence imaging, laser speckle contrast imaging, ultrasound images, X-rays, CT-scans, etc.), electronic medical records (EMRs), electronic health records (EHRs), and the like. In some embodiments, the training sets may include other types of data such as medical test data and the like. In some embodiments, the training sets of LLM may be obtained from an expert database. For example, the training sets may include information relating to one or more disease states / conditions, which may be stored in the expert database. In some embodiments, training sets may include surgical practices correlated to alleviation of disease states.

[0127] In some embodiments, a LLM may be generally trained. For example, generally trained may mean that the LLM is trained on a general training set comprising a variety of subject matters, data sets, and fields. In some embodiments, LLM may be initially generally trained. In some embodiments, the LLM may be specifically trained. For example, specifically trained may mean that LLM is trained on a specific training set, wherein the specific training set includes data including specific correlations for the LLM to learn.

[0128] An LLM in some embodiments may include Generative Pretrained Transformer (GPT), GPT-2, GPT-3, GPT-4, and the like. The LLM may include a text prediction-based algorithm configured to receive one or more images and / or videos and process the one or more images and / or videos to generate a textual description of the one or more images and / or videos. In some cases, the LLM may apply a probability distribution to the words already typed in a sentence (e.g., a question, query, or input topic sentence) to determine the most likely word to be next when generating the textual description of the of the one or more images and / or videos. The LLM may include an encoder component and a decoder component.

[0129] A LLM may include a transformer architecture. In some embodiments, encoder component of the LLM may include transformer architecture. A transformer architecture may be a neural network architecture that uses self-attention and positional encoding. The transformer architecture may be designed to process sequential input data, such as natural language, with applications towards tasks such as translation and text summarization.

[0130] A LLM and / or transformer architecture may include an attention mechanism. An attention mechanism as used herein may be a part of a neural architecture that enables a system to dynamically extract one or more relevant features of the input data (e.g., the input data may include one or more images and / or videos comprising intraoperative scenes of a surgical procedure). Natural language processing and the LLM can be applied to the extracted relevant features to generate a sequence of textual elements descriptive of those features.

[0131] Multimodal Models - FIG. 6 is a box diagram of a structure of multi-modal generative model 600. Model 600 comprises a large language model 610, a vision / visual encoder 620, and optionally multilayer perceptron layer 630. In some cases, the LLM 610 may be part of a multimodal large language model. Multimodal learning may be a type of deep learning that integrates and processes multiple types of data, referred to as modalities, such as text, audio, images, or video. For example, a multi modal model herein may comprise video to text generation or image to text generation. In some cases, the generative Al model may be used for processing visual data and / or text from the one or more images. In some cases, the multimodal, generative large language model is a multi-modal video to text model.

[0132] Various generative models may be implemented as multi-modal models. For example, GPT-4 may be able to implement both text and images as input. Gemini and Mistral may also be multimodal. However, video to text multimodal models may be rarer. In some cases, the multimodal model is LLAVA (large language and vision assistant). LLAVA may comprise a vision encoder and an LLM for general purposes visual and language understanding. The LLM may comprise any LLM herein. In some cases, the LLM is a text only version of GPT-4.

[0133] In some cases, the multimodal model is NVIDIA’s Vision Language (VILA) model. VILA may comprise a text encoder for text data and a vision encoder for image and / or video data. Some models may be employ a separate module for each of understanding and generating visual content, which can lead to misalignment and increased complexity. In contrast, VILA may employ a single autoregressive next-token prediction framework for both tasks, eliminating the need for additional components like diffusion models. This approach may simplifies the model and may achieve near state-of-the-art performance in visual language understanding and generation. The success of VILA may be attributed to one or more of: the unified vision tower that aligns discrete visual tokens with textual inputs during pretraining, which enhances visualperception, and autoregressive image generation can achieve similar quality as diffusion models with high-quality dataset. VILA may perform comparably to more complex models using a fully token-based autoregressive framework. FIG. 14 is a comparison of VILA 1.5 with LLAVA for triplet evaluation and phase evaluation on a representative dataset of surgical images. As shown VILA may perform better than LLAVA at most phases and most triplets.

[0134] As described herein, it may be helpful for a model to be specifically trained. For example, a specifically trained model may be domain specific. Systems and the methods of the present disclosure may employ a domain specific model for medical applications, such as LLAVA-Med (a domain specific version of LLAVA trained on 15M caption image / text pairs on PubMed). Systems and the methods of the present disclosure may employ a domain specific model for surgical applications, such as LLAVA-surg.

[0135] Visual Encoder: In some cases, the multi-modal generative model comprises a visual encoder 620. One method to create a multimodal model is to tokenize the output of trained encoder. In some cases, the visual encoder is configured to receive the one or more frames of video data obtained using one or more image and / or video acquisition modules and transform the one or more frames of video data into an input into the multimodal, generative large language model. A vision encoder may be a version of (contrastive language-image pre-training) CLIP or Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation (BLIP). BLIP may generate a set of candidate captions for a given image. CLIP may ranks the generated captions based on accuracy. CLIP may train a pair of neural networks, one for image understanding and one for text understanding using a contrasting objective. A text encoding model may be a transformer.

[0136] The image encoding models used may be vision transformers (ViT). For example, a vision transformer (ViT) is a transformer-like model that handles vision processing tasks. While CNNs use convolution, a local operation bounded to a small neighborhood of an image, ViTs use self-attention, a global operation, since the ViT draws information from the whole image. This allows the ViT to capture distant semantic relevances in an image effectively. Advantageously, ViTs may be well-suited catching long-term dependencies. In some cases, ViTs may be a competitive alternative to convolutional neural networks as ViTs may outperform the current state-of-the-art CNNs by almost four times in terms of computational efficiency and accuracy. ViTs may be well-suited to object detection, image segmentation, image classification, and action recognition. Moreover, ViTs may be applied in generative modeling and multi-model tasks, including visual grounding, visual-question answering, and visual reasoning. In some cases, ViTs may represent images as sequences, and class labels for the image are predicted, which allows models to learn image structure independently. Input images may be treated as asequence of patches where every patch is flattened into a single vector by concatenating the channels of all pixels in a patch and then linearly projecting it to the desired input dimension. For example, a ViT architecture may include the following operations: (A) split an image into patches; (B) flatten the patches; (C) generate lower-dimensional linear embeddings from the flattened patches; (D) add positional embeddings; (E) provide the sequence as an input to a standard transformer encoder; (F) pretrain a model with image labels (e.g., fully supervised on a huge dataset); and (G) finetune on the downstream dataset for image classification.

[0137] In some cases, there may be multiple blocks in a ViT encoder, with each block comprising three major processing elements: (1) Layer Norm; (2) Multi-head Attention Network; and (3) Multi-Layer Perceptrons. The Layer Norm may keep the training process on track and enable the model to adapt to the variations among the training images. The Multi-head Attention Network may be a network responsible for generating attention maps from the given embedded visual tokens. These attention maps may help the network focus on the most critical regions in the image, such as object(s). The Multi-Layer Perceptrons may be a two-layer classification network with a Gaussian Error Linear Unit at the end. The final Multi-Layer Perceptrons block may be used as an output of the transformer.

[0138] Computer vision - In some non-limiting embodiments, one or more computer vision algorithms may be used to implement the systems and methods of the present disclosure. The one or more computer vision algorithms may be used in combination with machine learning to enhance the quality of annotations generated.

[0139] In some cases, the one or more computer vision algorithms may comprise, for example, an object recognition or object classification algorithm, an object detection algorithm, a shape recognition algorithm and / or an object tracking algorithm. In some cases, the computer vision algorithm may comprise computer implemented automatic image recognition with pixel and pattern analysis. In other cases, the computer vision algorithm may comprise machine vision processes that involve different kinds of computer processing such as object and character recognition and / or text and visual content analysis. In some embodiments, the computer vision algorithm may utilize machine learning or may be trained using machine learning.The systems, the methods, the computer-readable media, and the techniques disclosed herein may implement one or more computer vision techniques. Computer vision is a field of artificial intelligence that uses computers to interpret and understand the visual world at least in part by processing one or more digital images from cameras and videos. In some instances, computer vision may use deep learning models (e g., convolutional neural networks). Bounding boxes may be used in object detection techniques within computer vision. Bounding boxes may be annotation markers drawn around objects in an image. Bounding boxes, are often, although notalways, may be rectangularly shaped. Bounding boxes may be applied by humans to training data sets. However, bounding boxes may also be applied to images by a trained machine learning that is trained to detect one or more different objects (e g., humans, hands, faces, cars, etc.). In addition or in alternative to bounding boxes detection and tracking techniques may use any object detection annotation techniques, such as semantic segmentation, instance segmentation, polygon annotation, non-polygon annotation, landmarking, 3D cuboids, etc.Training

[0140] A model may be trained. Model training may involve an optimization step wherein model parameters (such as weights or biases) may be altered based on the optimizer. Model training may involve a loss function which calculates a score based on the output of the model and the expected output of the model (such as a ground truth or labels). Model training may involve a dataset. The dataset may be split into one or more subsets. The subsets may be of different sizes. The subsets may be used for training, validation, testing or any combination thereof.

[0141] A model may be trained more than once. A model may be trained one a different dataset than was used in a previous training round (e.g., transfer learning, fine-tuning of the model, integration of the model into a larger model, continuous training, or some combination thereof). During training, whether in the first round or subsequent rounds, a subset of the model parameters may be untrainable during fine-tuning.

[0142] During fine-tuning a trained model may be trained on a different set of data, a subset of the original data or some combination thereof. Fine-tuning may cause the model to improve its performance on a given task or subtask. During transfer learning a model may be trained to improve performance on a task similar to the task the model was previously trained on. For example, a model may be trained to answer questions based on images or video and later finetuned to improve on a specific task such as answering questions about medical video and images, such as illustrated in FIG. 10A. During transfer learning a model may be trained to improve performance on a task not similar to the task the model was previously trained on. During transfer learning a model may be trained to learn a different task it was not previously trained for, for example a model may be trained to answer questions about a video in the first round of training, in a subsequent round of training the model may then be trained to answer questions where the question or answer contains (or is expected -based on the labels - contain in the case of the answer). In such an example the model may be pretrained to understand the relationship between a video or image and a question answer pair, in the subsequent rounds of training the model is then able to use the information learned to carry out a more difficult task of identifying a bounding box and processing what is in the segment of the image identified by the bounding box to give an answer.

[0143] FIG. 7 shows a method of training a multi-modal generative model of the present disclosure. Systems and method of the present disclosure may comprise three training stages: pre-training (feature alignment) 710, visual instruction tuning 720, and purpose-specific instruction fine tuning 730.

[0144] Pretraining step 710 may comprise use of a generative model with some level of pretraining. For example, LLAVA may be pre-trained from to process textual instructions and the vision encoder from pre-trained CLIP, a ViT model, to process image information. In this way, the training of a generally trained model may be leveraged. While pretraining or concept alignment may be helpful, it may not be necessary.

[0145] FIG. 12 shows an image from the CholecT50 dataset with the phase, bounding boxes and triplets superimposed on the image. A pretrained GPT (GPT-4v) was evaluated using 100 images from the CholecT50 dataset. The dataset supports surgical action triplet recognition; surgical action triplet detection / localization; surgical tool presence detection; surgical tool detection / localization; surgical action / verb recognition; surgical target recognition; surgical phase recognition or any combination thereof. The data in FIG. 12 is shown prior to model fine tuning and with an existing database of white light images for evaluation purposes. The second triplet indicates there is a bipolar (the instrument), coagulating (the verb), and liver (the target) for the bounding box. The pretrained GPT model was evaluated for performance across multiple portions of the dataset. It predicted phase with an accuracy of 24.0%, triplet accuracy at 9.4%, target accuracy 21.37%, verb accuracy 22.22%, instrument accuracy 48.72%, instrument count accuracy 49.0%.

[0146] Visual instruction tuning step 720 may comprise a process of training multimodal Al models to understand and respond to text-based instructions paired with visual inputs, such as images or videos. This technique aligns visual understanding with natural language processing capabilities, enabling the model to perform tasks like image captioning, visual question answering, object recognition, and information extraction.

[0147] A first layer of visual instruction training may be more tailored to task at hand. Examples of datasets of visual instruction tuning may comprise common objects in context (COCO) a database of objects and captions; the GQA dataset (a Stanford dataset comprising 22 million questions about various day-to-day images); the OCR-VQA database (a dataset of a total of 207,572 images along with their associated question-answer pairs); the visualgenome dataset (U Washington dataset with 100k images and 1.7 million visual question and answers; and LLAVA instruct (GPT4 generated instruction set from COCO).

[0148] Purpose specific instruction fine tuning 730 may comprise further fine tuning the model to the surgical domain. Purpose specific fine tuning may comprise one or more of: expertgenerated tuning and GPT-4 generated multi-turn Q and A on a database of surgical procedures. The GPT-4 generated responses may be generated from an initial round of expert training or not. The GPT-4 generated responses may be rated based on accuracy or filtered for accuracy by an expert trainer to improve quality.

[0149] In some cases, the large language model is trained based at least in part on a library of annotations of surgical procedures. In some cases, the multimodal, generative large language model is configured generate the written description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedure based on the one or more frames of video data and one or more training data sets comprising surgical or medical data associated with a patient or a surgical procedure. In some cases, the one or more training data sets comprise anatomical and physiological data obtained using a plurality of imaging modalities.

[0150] The following shows example correspondence between a number a surgeries and a corresponding number of annotated frames using in fine tuning systems and methods of the present disclosure.

[0151] Action triplets - In some cases, the fine tuning comprises action triplets. For example, the library of annotations may be formulated as action triplets. In some cases, the action triplets comprise an indication of an instrument, a verb, and a target in a surgical scene. Action triplet may be advantageous in that they may provide structure to the output of the surgical summaries. For example, a particular step of a surgical phase may comprise a description of an instrument, a verb, and a target in surgical scene. The action triplets may reduce hallucinations and improve the quality of the annotations. In some cases, the action triplets are further associated with a surgical phase.

[0152] In some cases the large language and vision assistant is fine-tuned based on a ground truth dataset comprising ground truth image data and the action triplets. In some cases, the ground truth dataset comprises expert annotations. In some cases, the ground truth dataset comprises expert rated Q and A pairs.

[0153] In some cases, a method described herein such as method 100, 200, or 300 may further comprise annotating one or more frames of the image data with action triplets based at least in part on an output of the multimodal, generative large language model. In some cases, a methoddescribed herein such as method 100, 200, or 300 may further comprise using the multimodal, generative large language model to annotate frames of the image data with action triplets.

[0154] FIG. 8A illustrates an image of a surgical video; a caption for the image; and an expert generated multi-tum question and answer for the training dataset. FIG. 8B illustrates an image of a surgical video; a caption for the image; and a GPT-4 generated multi-tum question and answer for the training dataset. FIG. 8C illustrates another image of a surgical video; a caption for the image; and a GPT-4 generated multi-turn question and answer for the training dataset. FIG. 8D illustrates another image of a surgical video; a caption for the image; and a GPT-4 generated multi-turn question and answer for the training dataset.

[0155] As shown in the illustrated examples, bounding boxes may be used for tool localization or localization of features of interested. Bounding boxes may be generated by computer vision techniques as disclosed herein. As shown in the illustrated examples, the captions for the images and the question and answer may use the action triplets to format responses and captions. FIG. 8A-FIG. 8D each show white light image data; however, other imaging modalities as disclosed herein may be used. As shown during the fine-tuning process, the written description comprises reasoning for particular surgical steps or deduction of next steps.

[0156] In some cases, a model may be trained on a large corpus of images and text, using image as input, and outputting a string of text. During training the model may learn to identify objects in various regions of images, such as within a bounding box. The model may learn to describe a process within a bounding box, point out where in the image an object is by outputting, as a part of a greater output, a description of a bounding box containing the object being described, etc.

[0157] FIG. 9A shows six examples of generated action triplets for individual frames of a laparoscopic cholecystectomy. As shown, the instruments in the procedure may comprise a hook, grasper, clipper, bipolar, irrigator, scissors, etc. As shown, the various images may be coded to a surgical phase. The target may comprise the item the tool is performing an action on. The action may be the thing that the tool is doing. An action may be dissecting, coagulating, retracting, grasping, aspirating, cutting, etc. A target may be an organ, a tissue, a fluid, a specimen bag, a cystic duct, etc. FIG. 9B illustrates an example of how the prompt to the Al may be modified to call a particular action triplet in the data.

[0158] A model may learn emergent properties. Emergent properties may be behaviors or capabilities that are not explicitly trained for. Emergent properties may be tasks (such as reasoning in LLMs, development of strategies in game-playing, obstacle avoidance in robot swarms, or edge detection in convolutional neural networks). Emergent properties may go beyond simple interactions in data and may involve interactions between components of the model. Emergent properties may arise from a self-organization wherein complex informationstructures and patterns are formed. Emergent properties may be used to improve a task (see FIG 9), such as a question answering task being applied to images or video where an emergent property, such as the ability to identify objects within a user defined bounding box and / or identifying a bounding box for a given object that has been described in an input text, the model may then be trained to improve the question answer task by using the emergent property . A model may be trained (either in a first round of training or subsequent rounds of training) to improve performance of an emergent property.

[0159] In some cases, the written description in the training data comprises reasoning for particular surgical steps or deduction of next steps. A written description may comprise generating an annotation for a particular frame. A written description may comprise generating a multi-turn dialog, such as a question-and-answer dialog. In some cases, the annotations in the training data comprise billing coding of surgical procedures. In some cases, the written description in the training data comprises a descriptor of type(s) of surgical tools used, and a manner in which the surgical tools are being used during the surgical procedure. In some cases, the written description in the training data comprises a descriptor of one or more anatomical features within the one or more intraoperative scenes, and changes to the one or more anatomical features during the surgical procedure. In some cases, the written description in the training data comprises a descriptor of one or more intraoperative complications that occurred during the surgical procedure.Surgical Annotations

[0160] Methods and systems disclosed herein may provide a written description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedure. For example, at an operation 130, method 100 may comprise generating a written description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedure based at least in part on the multimodal, generative large language model. For example, at an operation 230, method 200 may comprise generating a written description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedure based at least in part on the multimodal, generative large language model. For example, at an operation 320, method 300 may comprise generating text associated with the one or more images using the generative Al model. In some cases, the text comprises a description of one or more intraoperative scenes in the one or more images. In some cases, the written description comprises reasoning for particular surgical steps or deduction of next steps. In some cases, the written description is based at least on multi-turn question and answer dialogue with the generative Al model.

[0161] The writen description may be generated at least in part by a machine learning algorithm herein. The writen description may comprise reasoning for particular surgical steps or deduction of next steps. Methods and systems disclosed herein may automatically generate text associated with the one or more images. The text may comprise a description of one or more intraoperative scenes in the one or more images. In some cases, the written description comprises reasoning for particular surgical steps or deduction of next steps. In some cases, the written description is based at least on multi-turn question and answer dialogue with the large language and vision assistant.

[0162] Thus, it may be useful to supplement existing large language models with a library of annotated surgical data which may be used to help tune the model to improve the generation surgical notes while mitigating hallucinations.

[0163] FIG. 10A, FIG. IOC, and FIG. 10D are a schematic illustrations of the systems and methods of the present disclosure (e.g., Activ Co-pilot). In some cases, systems and methods disclosed herein may be useful for more accurate billing. FIG. 10A shows an output from a multi-modal generative Al of the present disclosure. Systems and methods herein may populate image snapshots and a corresponding writen description of complex and abnormal anatomy, pathology and physiology observed during a surgical case. This may be particularly useful when unforeseen intraoperative complications or delays occur, and the reimbursement opportunity substantially increases (2-3X) for hospitals and physicians.

[0164] FIG. 10A shows an example prompt sent to the model with specific descriptions of the requested output and a call for the action triplet. The call for the action triplet may be similar to that described above with respect to FIG. 9B. As shown the output of the model shows a high degree of surgical reasoning and shows intuition for the next step of the procedure and the relationship to the surgical procedure in progress. FIG. 10B shows an example multi turn question and answer exchange with a model as disclosed herein.

[0165] FIG. IOC shows a flow chart of a method 1000 of the present disclosure. The method may comprise providing a video library of a surgical procedure 1010, annotating frames with action triplets 1020, implementing a large language and vision assistant comprising a large language model and an image processing model 1030, and using the large language and vision assistant to generate a written description of anatomy, pathology, and / or physiology observed during the surgical procedure 1040. Method 1000 may comprise any embodiment, variation, or example of any of method 100, 200, or 300 herein.

[0166] FIG. 10D shows a schematic of a system 1050 of the present disclosure. System 1050 may comprise any embodiment, variation, or example of a system 400 herein. System 400 may comprise one or more imaging modules 410. The one or more imaging modules may be configured to collect image data of the surgical procedure. The image data may be based at leastin part on one or more frames of video data of the surgical procedure. System 400 may comprise one or more processors 420. A processor 420 may comprise any embodiment, variation, or example described herein with respect to the section “Computer Systems” herein.

[0167] The system 1050 may comprise an imaging module 1051 configured to collect a video library of a surgical procedure and a processor 1052 configured to implement a large language and vision assistant of the present disclosure. Processor 1052 may be configured to use the large language and vision assistant to generate a written description of anatomy, pathology, and / or physiology observed during the surgical procedure. Imaging module 1051 may comprise any embodiment, variation, or example of one or more imaging modules 410. Processor 1052 may comprise any embodiment, variation, or example of one or more processors 420. Processor 1052 may comprise any embodiment, variation, or example described herein with respect to the section “Computer Systems” herein.

[0168] Methods and systems for image-to-text generation are provided. A method may comprise providing one or more images of a surgical procedure to a generative artificial intelligence (Al) model; and using the generative Al model to automatically generate text associated with the one or more images, wherein the text comprises a description of one or more intraoperative scenes in the one or more images.

[0169] In some cases, the description comprises a descriptor of type(s) of surgical tools used, and a manner in which the surgical tools are being used during the surgical procedure. In some cases, the description comprises a descriptor of one or more anatomical features within the one or more intraoperative scenes, and changes to the one or more anatomical features during the surgical procedure. In some cases, the description comprises a descriptor of one or more intraoperative complications that occurred during the surgical procedure. In some cases, the method further comprises identifying one or more billing codes from a plurality of billing codes for the surgical procedure, based at least in part on the text comprising the description of the one or more intraoperative scenes. In some cases, the one or more billing codes are identified at an accuracy of greater than about 80%. In some cases, the generative Al model comprises a large language model (LLM). In some cases, the generative Al model is used for processing visual data and / or text from the one or more images.

[0170] In some cases, the generative Al model automatically generates the text in response to one or more questions provided via a graphical interface. In some cases, the one or more questions are provided by one or more users via the graphical interface. In some cases, the one or more questions are automatically generated based at least in part on a stored library of questions associated with the surgical procedure. In some cases, the one or more questions are selected from a plurality of questions presented to one or more users via the graphical interface.

[0171] In some cases, the one or more images are annotated. In some cases, the one or more images are non-annotated. In some cases, the one or images are extracted from a video of the surgical procedure. In some cases, the one or images are taken at different points in time during the surgical procedure. In some cases, the one or images are taken at different locations and / or perspectives during the surgical procedure. In some cases, the one or images are taken using two or more different imaging modalities. In some cases, the two or more different imaging modalities comprise at least one of white light imaging or fluorescence imaging. In some cases, the one or images comprises a plurality of images captured at different wavelengths. In some cases, the one or more images comprises perfusion data. In some cases, the surgical procedure comprises a laparoscopic procedure.

[0172] In some cases, data for machine learning may comprise descriptions of an image and / or video. For example, any of the image data disclosed herein above may be annotated. In some cases, the one or more images are annotated. In some cases, the one or more images are nonannotated. Data for machine learning may comprise descriptions of surgical phase. In some cases, the written description is generated in substantially real time with the surgical procedure.

[0173] Data may be structured as an input and expected output as a pair (such as an image as input and a set of action triplet, surgical phase (phase), and bounding box coordinate as expected output for that input). For example, data for machine learning may comprise bounding box coordinates. Data for machine learning may comprise action triplets. Ground truth data may comprise surgeon annotated frames from prior surgeries. Data for machine learning may comprise question-answer pairs.

[0174] In some cases, the written description comprises reasoning for particular surgical steps or deduction of next steps. In some cases, the annotations comprise billing coding of surgical procedures. In some cases, the description comprises a descriptor of type(s) of surgical tools used, and a manner in which the surgical tools are being used during the surgical procedure. In some cases, the description comprises a descriptor of one or more anatomical features within the one or more intraoperative scenes, and changes to the one or more anatomical features during the surgical procedure. In some cases, the description comprises a descriptor of one or more intraoperative complications that occurred during the surgical procedure.

[0175] In some cases, the written description in the model output comprises reasoning for particular surgical steps or deduction of next steps. A written description may comprise generating an annotation for a particular frame. A written description may comprise generating a multi-turn dialog, such as a question-and-answer dialog. In some cases, the annotations in the model output comprise billing coding of surgical procedures. In some cases, the written description model output comprises a descriptor of type(s) of surgical tools used, and a manner inwhich the surgical tools are being used during the surgical procedure. In some cases, the written description in the model output comprises a descriptor of one or more anatomical features within the one or more intraoperative scenes, and changes to the one or more anatomical features during the surgical procedure. In some cases, the written description in the model output comprises a descriptor of one or more intraoperative complications that occurred during the surgical procedure.

[0176] In some cases, a method of the present disclosure such as method 100, 200, 300, 1000, etc. further comprises: identifying one or more billing codes from a plurality of billing codes for the surgical procedure, based at least in part on the text comprising the description of the one or more intraoperative scenes. In some cases, the one or more billing codes are identified at an accuracy of greater than about 80%. In some cases, the generative Al model is trained based at least in part on a library of annotations of surgical procedures. In some cases, written description comprises a descriptor of type(s) of surgical tools used, and a manner in which the surgical tools are being used during the surgical procedure. In some cases, the written description comprises a descriptor of one or more anatomical features within the one or more intraoperative scenes, and changes to the one or more anatomical features during the surgical procedure. In some cases, the written description comprises a descriptor of one or more intraoperative complications that occurred during the surgical procedure.

[0177] Question and answer - In some cases, a multimodal generative model of the present disclosure may be configured to output response based on a multi-turn dialog, such as a question- and-answer dialog. For example, the written description may be based at least on multi-turn question and answer dialogue with the large language and vision assistant. For example, the written description may be based at least on multi-turn question and answer dialogue with any example of a multi-modal generative model herein.

[0178] In some cases, the generative Al model automatically may generate text in response to one or more questions provided via a graphical interface, such as a graphical user interface described herein with respect to the section “Computer Systems.” In some cases, in the one or more questions are provided by one or more users via the graphical interface. In some cases, the one or more questions are automatically generated based at least in part on a stored library of questions associated with the surgical procedure. In some cases, the one or more questions are selected from a plurality of questions presented to one or more users via the graphical interface. For example, it may be helpful to populate example prompts for particular use cases, such as generating a surgical summary, generating bill codes, generating a description of a particular tool, calling one or more action triplets for any of the preceding, etc. In some cases, an output from a generative Al may be improved by prompt engineering. A preformed list of questions orsuggestion to improve questions may thus improve the prompt sent to the multi-modal generative Al.

[0179] Video Acquisition - As shown in FIG. 10D, the system 1050 may comprise an imaging module 1051 configured to collect a video library of a surgical procedure. Imaging module 1051 may comprise any embodiment, variation, or example of one or more imaging modules 410. In some cases, the imaging module is coupled to a processor comprising the generative Al with a wired connection. In some cases, the imaging module is coupled to a processor comprising the generative Al with a wireless connection. The imaging module may be communicatively coupled to the processor comprising the generative Al model.

[0180] The imaging model may be configured to collect any form of imaging / image data herein. In some cases, the imaging module comprises one or more members selected from the group consisting of a white light imaging module, fluorescence imaging module, and perfusion imaging module. In some cases, the one or images are taken using two or more different imaging modalities. In some cases, the two or more different imaging modalities comprise at least one of white light imaging or fluorescence imaging.

[0181] In some cases, the one or more images from the imaging module comprises white light data. White light imaging may be a name for imaging with a visible color spectrum. A white light image may comprise an image taken with a continuous or semicontinuous spectrum within a visible range. For example, white light may be between 400 and 700 nm. A white light source may comprise a tungsten bulb. A white light source may comprise a multicolor LED with a diffuser to create a continuous spectrum. White light imaging may be a relatively standard imaging modality which may bridge the gap between perfusion imaging (e.g., ActivSight) and more generally applicable applications.

[0182] In some cases, the one or more images from the imaging module comprises fluorescence data. Florescence imaging may comprise a type of imaging based on the ability of a fluorophore in the image to fluoresce in response to an excitation. Fluorescence imaging may comprise an excitation beam. An excitation beam may be absorbed by a fluorophore in the tissue. The fluorophore may then emit light in response. The emitted light fluorescence may be at a lower energy / longer wavelength than the excitation beam. Fluorescence imaging may help show critical structures (e.g., blood vessels, oxygenated tissues, etc.) outside of the capability of human sight and may be specifically tuned to anatomy. For example, a fluorophore may be injected into a patient blood stream in order to allow for imaging of the fluorescing tissue. Fluorophore may comprise indocyanine green, other cyanine derivatives, riboflavin, fluorescein, etc.

[0183] In some cases, the one or more images from the imaging module comprises perfusion data. Perfusion imaging may be used to show real-time blood flow and tissue perfusion. In somecases, perfusion imaging may comprise laser speckle contrast imaging. In a laser speckle contrast image, speckles are produced when coherent light scattered back from biological tissue. Mobile scatterers cause the speckle pattern to blur. In biological tissue, red blood cells may be useful scatters. Thus, where blood is flowing, a laser speckle contrast signal may be produced. Laser speckle imaging may be non-invasive. Imaging that shows where blood is flowing may help clearly indicate the presence of a blood vessel, which may otherwise be hard to see, and it may help allow a practitioner to see when a blood vessel has been adequately clamped before performing a procedure. This data may be useful for a broad range of surgical procedures. In some cases, a coherent light source is provided to generate the laser speckle data. The coherent light source may be an 800 nm light source. The coherent light sources may be a longer than 700 nm light source.

[0184] In some cases, the one or images from the imaging module comprises a plurality of images captured at different wavelengths. For example, an image may be captured within a white light region (for white light data) and within an emission region for fluorescence data. For example, an image may be captured within a white light region (for white light data) and at the wavelength of the scattered light for perfusion data.

[0185] The image data from the imaging module based at least in part on one or more frames of video data of the surgical procedure may be supplemented by other forms of data. For example, a particular frame data may be coded with a time stamp, an indication of the type of procedure, an indication of the patient or particular instance of the procedure, etc.

[0186] The imaging module may interface with an example of a generative Al herein. For example, a visual encoder may be configured to (i) receive the one or more frames of video data obtained using one or more image and / or video acquisition modules and transform the one or more frames of video data into an input into the multimodal, generative large language model. In some cases, the one or more training data sets comprise anatomical and physiological data obtained using a plurality of imaging modalities, such as from an imaging module capable of multiple image types.

[0187] An imaging device may comprise (i) a scope, such as a laparoscope (as shown in FIG. 10D), (ii) a robot assisted imaging device, or (iii) any digital anatomic visual sensors in the red- green-blue visual, near-infrared, and / or short-wave infrared spectrums, and may include, in some non-limiting examples, physiologic or functional imaging data. In some cases, the imaging data may be based on chemical signal enhancers like indocyanine green (ICG), fluorescent or radiolabeled dyes, and / or optical signal enhancers based on coherent light transmissions, such as laser speckle imaging. In some cases, the imaging data may comprise two or more medical images and / or videos that are co-registered.Integration with other Models

[0188] In some cases, a model herein can then generate additional Al models to solve specific tasks intraoperatively. For example, a generative multi-modal Al model herein may be integrated with another Al in order to improve that Al. FIG. 11 shows three Al models which may be used in concert with systems and methods of the present disclosure. They comprise Colorectal Al which limits perfusion information to display on target anatomy, Lap Chole Al which segments anatomy during a Lap Chole to prevent common bile duct injuries, Out of Body Detector which removes frames with PHI, and DLMed annotation tools which train Clara DLMed. FIG. 11 illustrates before and after photos showing improvement of the image processing model.

[0189] In an example, a multi-modal generative Al may be integrated with a surgical guidance system. Surgical guidance systems may provide predictions of anatomic locations and contours of tissue structures based on standard intraoperative laparoscopic and robot assisted endoscopic datasets in the visual spectrum. Given the variations in patient anatomy, pathology, surgical techniques, and surgeon style, competence and proficiency, the precision and accuracy of such systems may be insufficient for clinical applications, and the feasibility of such systems may be impractical. A surgical guidance system may provide artificial intelligence (Al)-based systems that can be utilized to improve a priori or anticipatory surgical outcomes.

[0190] A surgical guidance system may provide augmented visuals, real-time guidance, and decision support to surgeons and other operating assistants using artificial intelligence (Al). More specifically, a surgical guidance system may provide intraoperative real-time prediction of the location of, and identification of critical structures and / or tissue viability before, during, and / or after surgical procedures using machine learning algorithms trained on any combination of white light-based images and / or videos, chemical signal enhancers such as fluorescent near-infrared (NIR) images and / or videos or radio-labelled dye imaging, indigenous chromophores, autofluorescence, and optically enhanced signals such as laser speckle-based images and / or videos, for the purpose of executing, analyzing, identifying, anticipating, and planning intended steps and avoiding unintended steps. A surgical guidance system may provide real-time intraoperative machine learning-based guidance systems that may provide real time visual support before and / or during surgical procedures (e.g., dissection) on or near critical structures, so that surgical operators can identify relevant critical structures, assess tissue viability, and / or make surgical decisions during the operation. In some embodiments, such real-time machine learning-based guidance systems may also be implemented as a tool for ex post facto medical or surgical training. A surgical guidance system may provide the capacity to autonomously or semi- autonomously track and update mobile and deformable tissue targets as tissue maneuvers are performed during a surgical operation or intervention.

[0191] Tissue viability may relate to the ability of a tissue (or the cells of the tissue) to receive blood, nutrients, or other diffusible materials or substances which allow the tissue to stay alive and function in a normal manner. The machine learning algorithms may be trained on any combination of white light-based images and / or videos, chemical signal enhancers such as fluorescent near-infrared (NIR) images and / or videos or radio-labelled dye imaging, indigenous chromophores, autofluorescence, and optically enhanced signal such as laser speckle-based images and / or videos. The systems and methods of the present disclosure may be implemented for the purpose of identifying, planning, analyzing, and executing intended steps of a surgical procedure while avoiding unintended or unnecessary steps. The systems and methods of the present disclosure may be implemented to facilitate autonomous and / or semi-autonomous detection and tracking of mobile and deformable tissue targets over time as tissue maneuvers are performed during a surgical procedure.

[0192] The machine learning algorithms may be used to implement a real-time intraoperative Al guidance system that provides surgeons and medical personnel with visual support before, during, and / or after a surgical procedure. In one example, the real-time intraoperative guidance system may be used during a dissection or resection of tissues to aid a surgeon with anticipating and identifying relevant critical structures, evaluating, or visualizing tissue viability, and / or providing options of surgical decisions during the surgical procedure. In another example, the real-time Al-based guidance system may provide anticipated identification of critical anatomy as an intraoperative decision support to the surgeon and / or the operating assistant(s) before and / or during a Calot triangle dissection during the removal of a medical subject’s gall bladder(cholecystectomy). In another example, the real-time Al-based guidance system may be configured and / or used to provide information about tissue viability during surgical procedures.

[0193] The Al guidance system may be configured to obtain and evaluate new information about a patient’s anatomy and / or physiology in real-time as dissection or resection progresses. The Al guidance system may be configured to make intraoperative predictions or assessments of (i) the location of relevant critical structures and / or (ii) tissue viability during one or more surgical procedures. The predictions of critical structure locations and / or tissue viability may begin prior to a dissection or resection of critical areas and may be refined based on new information acquired by the system as further dissection or resection occurs.

[0194] In some cases, the Al guidance system may use perfusion data obtained using laser speckle contrast imaging to provide a surgeon with medical decision support during an operation. The system may also be used to predict the locations of critical structures and / or determine tissue viability for a given operative procedure. For example, in laparoscopic cholecystectomy, the system may be used to predict the location of one or more of the extrahepatic biliary structures,such as the cystic duct, cystic artery, common bile duct, common hepatic duct, and / or gallbladder. In other surgical procedures (e.g., surgical procedures on an organ, colorectal procedures, bariatric procedures, vascular procedures, endocrine procedures, otolaryngological procedures, urological procedures, gynecological procedures, neurosurgical procedures, orthopedics procedures, lymph node dissection procedures, procedures to identify a subject’s sentinel lymph nodes, etc.), the system may be configured to provide surgical guidance and medical data analytics support by predicting and outlining the locations of critical structures and / or organs before, during, and / or after the surgical procedures, as well as tissue viability.

[0195] In some cases, a real-time guidance and data analytics support system may be based on Al models that have been trained using one or more medical datasets corresponding to entire surgical procedures (or portions thereof), such as those described herein for surgical annotations. In some cases, the one or more medical data sets may comprise medical data pertaining to one or more critical phases or scenes of a surgical procedure, such as pre-dissection and post-dissection views of critical or vital structures.

[0196] Such medical data may be obtained using an imaging device herein. An imaging device may comprise (i) a scope, such as a laparoscope, (ii) a robot assisted imaging device, or (iii) any digital anatomic visual sensors in the red-green-blue visual, near-infrared, and / or short-wave infrared spectrums, and may include, in some non-limiting examples, physiologic or functional imaging data. In some cases, the imaging data may be based on chemical signal enhancers like indocyanine green (ICG), fluorescent or radiolabeled dyes, and / or optical signal enhancers based on coherent light transmissions, such as laser speckle imaging. In some cases, the imaging data may comprise two or more medical images and / or videos that are co-registered.

[0197] In some cases, the imaging data may comprise any type of imaging data disclosed herein. For example, imaging data may comprise one or more segmented medical images and / or videos indicating (1) anatomic positions of critical structures of surgical interest and / or (2) tissue viability. In any of the examples described herein, the medical data and / or the imaging data used to train the Al or machine learning models may be obtained using one or more imaging modalities, such as, for example, depth imaging or distance mapping. In some cases, information derived from depth imaging or distance mapping may be used to train the Al or machine learning models. The information derived from depth imaging may include, for example, distance maps that may be (1) calculated from disparity maps (e.g., maps showing the apparent pixel difference or motion between a pair of images or videos), (2) estimated using Al based monocular algorithms, or (3) measured using technologies such as time of flight imaging. In some cases, the one or more imaging modalities may alternatively or further include, for example, fluorescence imaging, speckle imaging, infrared imaging, UV imaging, X-ray imaging, intraoperativecholangiograms, angiograms, ductograms, ureterograms, and / or lymphangiograms. The use of such imaging modalities may allow for direct objective correlations between physiologic features and corresponding RGB images and / or videos.

[0198] The medical models described herein may be trained using physiologic data (e.g., perfusion data). The medical models may be further configured to process physiologic data and to provide real time updates. The addition of physiologic data in both training datasets and real time updates may improve an accuracy of medical data analysis, predictions for critical structure locations, evaluations of tissue viability, planning for surgical scene and content recognition, and / or execution of steps and procedural planning.

[0199] In some cases, the system may be configured to identify critical structures and visually outline the predicted contours and / or boundaries of the critical structures with a statistical confidence. In some cases, the system may be further configured to display, track, and update critical structure locations in real time a priori, at the onset of and / or during any surgical procedure based on a machine learning model-based recognition of surgical phase and content (e.g., tissue position, tissue type, or a selection, location, and / or movement of medical tools or instruments). The display and tracking of critical structure locations may be updated in real time based on medical data obtained during a surgical procedure, based on cloud connectivity and edge computing. As described elsewhere herein, the system may also be used to evaluate, assess, visualize, quantify, or predict tissue viability. In some cases, the system may be further configured to display notifications and / or informational banners providing medical inferences, information, and / or guidance to a surgeon, doctor, or any other individual participating in or supporting the surgical procedure (either locally or remotely). The notifications or message banners may provide information relating to a current procedure or step of the procedure to a surgeon or doctor as the surgeon or doctor is performing the step or procedure. The notifications or message banners may provide real time guidance or alerts to further inform or guide the surgeon or the doctor. The information provided through the notifications or message banners may be derived from the machine learning models described herein.

[0200] In some cases, the system may be configured to track a movement or a deformation of mobile and deformable tissue targets as tissue maneuvers are being performed during a surgical procedure. Such tracking may occur in an autonomous or semi-autonomous manner.

[0201] The system may be configured to use one or more procedure-specific machine learning models to provide surgical guidance. Each model may use a multi-stage approach, starting with a data multiplexer to combine anatomical and / or physiological data obtained from one or more medical imaging systems. The data multiplexer may be configured to combine all available data streams (e.g., data streams comprising medical information derived from medical imagery), witheach stream allocated to its own channel. The system may further comprise one or more machine-learning based feature extractors. The one or more machine-learning based feature extractors may comprise a spatial feature extractor and / or a temporal feature extractor. The one or more machine-learning based feature extractors may be in a spatial and / or temporal domain and may be trained on pre-existing surgical data. The ML based spatial feature extractor may be used on each individual channel to find relevant features that may be combined to generate a feature set for every frame (which may correspond to one or more distinct time points). The ML based temporal feature extractor may be configured to take multiple such feature sets and extract the relevant features for the whole feature set. The ML based spatial feature extractor may be used to identify textures, colors, edges etc., while the ML based temporal feature extractor can indicate frame to frame changes, such as changes in contrast, perfusion, perspective, etc.

[0202] The extracted features described above may be used by ML classifiers trained on preexisting data to (i) determine a surgical phase of an ongoing surgical procedure, (ii) predict or infer when a surgeon needs guidance, (iii) infer tissue types as a surgeon is viewing and / or operating on a target tissue or other nearby regions, and / or (iv) register the current surgical view to a known set of coordinates or reference frame. The classifications provided by the ML classifiers may be fed into a machine-learning based critical structure detector or tissue viability evaluator that can provide an augmented view generator with the required information to display guidance and metrics to a surgeon and / or one or more operating assistant(s) in real time during critical phases of a surgical procedure.

[0203] A surgical guidance system comprising an image processing module configured to (i) receive image or video data obtained using one or more image and / or video acquisition modules and (ii) generate one or more medical predictions or assessments based on (a) the image or video data or physiological data associated with the image or video data and (b) one or more training data sets comprising surgical or medical data associated with a patient or a surgical procedure. The one or more training data sets may comprise anatomical and physiological data obtained using a plurality of imaging modalities. The system may further comprise a visualization module configured to provide a surgical operator with an enhanced view of a surgical scene, based on the one or more augmented data sets. In some embodiments, the enhanced view provided by the visualization module comprises information or data corresponding to the tissue viability for one or more tissue regions. In some embodiments, the information or data comprises quantitative measurements of tissue viability. In some embodiments, the visualization module is integrated with or operatively coupled to the image processing module.

[0204] In some embodiments, the one or more image and / or video acquisition modules comprise (i) a first image and / or video acquisition module configured to capture images and / or videosusing a first imaging modality and (ii) a second image and / or video acquisition module configured to capture images and / or videos using a second imaging modality.

[0205] In some embodiments, the one or more medical predictions or assessments comprise an identification or a classification of one or more critical structures. In some embodiments, the one or more medical predictions or assessments comprise a predicted location of one or more critical structures. In some embodiments, the one or more medical predictions or assessments provide the surgical operator with predictive clinical decision support before and during a surgical procedure. In some embodiments, the one or more medical predictions or assessments are generated or updated in part based on an output of one or more computer vision algorithms. In some embodiments, the one or more medical predictions or assessments comprise medical information or live guidance that is provided to a user or an operator through one or more notifications or message banners. In some embodiments, the one or more medical predictions or assessments comprise tissue viability for one or more tissue regions in the surgical scene.

[0206] In some embodiments, the image processing module is configured to update or refine the one or more medical predictions or assessments in real time based on additional image data obtained during a surgical procedure. In some embodiments, the image processing module is configured to update or refine a predicted location of one or more critical structures in real time based on additional image data obtained during a surgical procedure.

[0207] In some embodiments, the one or more training data sets comprise medical data associated with one or more reference surgical procedures. In some embodiments, the one or more training data sets comprise medical data associated with (i) one or more critical phases or scenes of a surgical procedure or (ii) one or more views of a critical structure that is visible or detectable during the surgical procedure. In some embodiments, the one or more training data sets comprise medical data obtained using a laparoscope, a robot assisted imaging unit, or an imaging sensor configured to generate anatomical images and / or videos in a red-green-blue visual spectrum. In some embodiments, the one or more training data sets comprise medical data obtained using an imaging sensor configured to generate physiologic or functional imaging data based on chemical signal enhancers. In some embodiments, the chemical signal enhancers comprise ICG, fluorescent, or radiolabeled dyes. In some embodiments, the one or more training data sets comprise medical data obtained using an imaging sensor configured to generate physiologic or functional imaging data based on laser speckle patterns. In some embodiments, the one or more training data sets comprise medical data obtained using an imaging sensor configured to generate physiologic or functional imaging data. The physiologic or functional imaging data may comprise two or more co-registered images or videos. In some embodiments, the one or more training data sets comprise medical data obtained using an imaging sensorconfigured to generate physiologic or functional imaging data. The physiologic or functional imaging data may comprise segmented anatomic positions of critical structures of surgical interest. In some embodiments, the one or more training data sets comprise medical data obtained using a plurality of imaging modalities that enable direct objective correlation between physiologic images or videos and corresponding RGB images or videos. The plurality of imaging modalities may comprise depth imaging, distance mapping, intraoperative cholangiograms, angiograms, ductograms, ureterograms, or lymphangiograms. In some embodiments, the training data set comprises perfusion data.

[0208] In some embodiments, the image processing module is configured to update or refine the one or more medical predictions or assessments in real time based on additional physiologic data obtained during a surgical procedure. In some embodiments, the visualization module is configured to display and track a position and an orientation of one or more medical tools or critical features in real time. In some embodiments, the visualization module is configured to display a visual outline of the one or more critical features corresponding to predicted contours or boundaries of the critical features. In some embodiments, the image data obtained using one or more image and / or video acquisition modules comprises perfusion data. In some embodiments, the visualization module is configured to display a visual outline indicating a predicted location of one or more critical features before, during, or after the surgical operator performs one or more steps of a surgical procedure. In some embodiments, the visualization module is configured to autonomously or semi-autonomously track mobile and deformable tissue targets as tissue maneuvers are performed during a surgical procedure.

[0209] In some cases, the image processing module comprises a data multiplexer configured to combine anatomical and / or physiological data obtained using the one or more image and / or video acquisition modules. In some embodiments, the image processing module further comprises one or more feature extractors on spatial or temporal domains. The one or more feature extractors may be trained using the one or more training data sets and configured to extract one or more features from the combined anatomical and / or physiological data. In some embodiments, the one or more feature extractors comprise a spatial feature extractor configured to detect spatial features and generate a feature set for every image frame. In some embodiments, the spatial features comprise textures, colors, or edges of one or more critical structures. In some embodiments, the one or more feature extractors comprise a temporal feature extractor configured to detect a plurality of temporal features or feature sets over a plurality of image frames. In some embodiments, the temporal features correspond to changes in contrast, perfusion, or perspective.

[0210] In some cases, the image processing module further comprises a view classifier configured to use the one or more extracted features to determine a current surgical view relative to the surgical scene. In some embodiments, the image processing module further comprises a tissue classifier configured to use the one or more extracted features to identify, detect, or classify one or more tissues. In some embodiments, the image processing module further comprises a phase detector configured to use the one or more extracted features to determine a surgical phase and generate guidance based on the surgical phase. In some embodiments, the image processing module further comprises a critical structure detector configured to locate and identify one or more critical features based on (i) the one or more tissues identified, detected, or classified using the tissue classifier, (ii) a current surgical view determined using the view classifier, and (iii) the surgical phase determined using the phase detector. In some embodiments, the image processing module further comprises an augmented view generator configured to display guidance and metrics associated with the surgical procedure, based on the one or more critical features located and identified using the critical structure detector.[2H] A surgical guidance system may provide real-time augmented surgical guidance to surgeons or medical personnel. A surgical guidance system may be implemented to provide predictive visualizations of critical structure locations and / or tissue viability based on machine learning (ML) recognition of surgical phase and content. In some cases, such predictive visualizations may be updated and / or refined based on newly acquired information (e.g., visual, anatomic, and / or physiological data). A surgical guidance system may also be implemented to provide real-time predictive decision support for analyzing, anticipating, planning, and executing steps in a surgical procedure while avoiding unintended, unrecognized, or unnecessary steps. In some cases, such predictive decision support may be based on ML recognition of surgical phase and content, and such ML recognition may be based on training using (i) data sets obtained using a plurality of imaging modalities and (ii) physiological data of a patient. A surgical guidance system may be used for autonomous or semi-autonomous tracking of mobile and deformable tissue targets as tissue maneuvers are performed during a surgical procedure. In some cases, A surgical guidance system may be used to generate one or more procedure-specific ML-based models that are configured to identify or anticipate critical features, surgical phases, and / or procedure-specific guidance based on (i) imaging data obtained using a plurality of different imaging modalities and (ii) anatomical and / or physiological data of a patient. In some cases, such procedure-specific ML models may be further configured to generate an augmented view of a surgical scene that displays guidance / metrics to a surgeon or operating assistant.Computer Systems

[0212] The present disclosure provides computer systems that are programmed or otherwise configured to implement methods of the disclosure, e.g., method 100, method 200, method 300, or method 400 herein. A computer system herein may be an example of a processor disclosed herein with respect to system 1050 or system 400 herein.

[0213] FIG. 12 shows a computer system 1201 that is programmed or otherwise configured to implement systems and methods of the present disclosure. For example, computer system 1201 may be an example of one or more processors 420. For example, computer system 1201 may be an example of one or more processors 1052.

[0214] The computer system 1201 may be configured to, for example, annotate one or more frames of the image data with action triplets based at least in part on an output of the multimodal, generative large language model. The computer system may be operable to receive image data herein. For example, the image data may comprise one or more members selected from the group consisting of white light imaging, fluorescence imaging, and perfusion imaging. In some cases, the one or more image and / or video acquisition modules comprise (i) a first image and / or video acquisition module configured to capture images and / or videos using a first imaging modality and (ii) a second image and / or video acquisition module configured to capture images and / or videos using a second imaging modality. In some cases, the image data obtained using one or more image and / or video acquisition modules comprises perfusion data.

[0215] The computer system 1201 may comprise or be configured to implement a multimodal, generative Al herein, such as a generative large language model. In some cases, the multimodal, generative large language model is a LLAVA, LLAVA-Med, or LLAVA Surg.

[0216] The computer system 1201 may comprise or be configured to implement a visual encoder as described herein. In some cases, the visual encoder is configured to (i) receive the one or more frames of video data obtained using one or more image and / or video acquisition modules and transform the one or more frames of video data into an input into the multimodal, generative large language model. In some cases, a multimodal, generative large language model is configured (ii) generate the written description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedure based on (A) the one or more frames of video data and (B) one or more training data sets comprising surgical or medical data associated with a patient or a surgical procedure.

[0217] The computer system 1201 may comprise or be configured to operate on one or more training datasets herein. For example, the one or more training data sets comprise anatomical and physiological data obtained using a plurality of imaging modalities. In some cases, the one or more training data sets comprise medical data associated with one or more reference surgicalprocedures. In some cases, the one or more training data sets comprise medical data associated with (i) one or more critical phases or scenes of a surgical procedure or (ii) one or more views of a critical structure that is visible or detectable during the surgical procedure. In some cases, the one or more training data sets comprise medical data obtained using a laparoscope, a robot assisted imaging unit, or an imaging sensor configured to generate anatomical images and / or videos in a red-green-blue visual spectrum. In some cases, the one or more training data sets comprise medical data obtained using an imaging sensor configured to generate physiologic or functional imaging data based on chemical signal enhancers. In some cases, the chemical signal enhancers comprise ICG, fluorescent, or radiolabeled dyes. In some cases, the one or more training data sets comprise medical data obtained using an imaging sensor configured to generate physiologic or functional imaging data based on laser speckle patterns. In some cases, the one or more training data sets comprise medical data obtained using an imaging sensor configured to generate physiologic or functional imaging data, wherein said physiologic or functional imaging data comprises two or more co-registered images or videos. In some cases, the one or more training data sets comprise medical data obtained using an imaging sensor configured to generate physiologic or functional imaging data, wherein said physiologic or functional imaging data comprises segmented anatomic positions of critical structures of surgical interest. In some cases, the one or more training data sets comprise medical data obtained using a plurality of imaging modalities that enable direct objective correlation between physiologic images or videos and corresponding RGB images or videos, wherein the plurality of imaging modalities comprise depth imaging, distance mapping, intraoperative chol angiograms, angiograms, ductograms, ureterograms, or lymphangiograms. In some cases, the training data set comprises perfusion data.

[0218] In some cases, the computer system 1201 may configured to co-implement or train together or generate other Al models. For example, a surgical annotation system may be configured to fee a surgical guidance system. A surgical guidance system may form one or more medical predictions or assessments. In some cases, the one or more medical predictions or assessments comprise an identification or a classification of one or more critical structures.

[0219] A multi-modal generative Al maybe configured to generate a written description of one or more frames of video data. In some cases, the written description comprises a location of one or more critical structures. In some cases, the written description is configured provide the surgical operator with predictive clinical decision support before and during a surgical procedure. In some cases, the multimodal, generative large language model is configured to update or refine the written description in real time based on additional image data obtained during a surgical procedure.

[0220] The computer system 1201 can be an electronic device of a user or a computer system that is remotely located with respect to the electronic device. The electronic device can be a mobile electronic device. In some instances, a single computer system or computing device may be used for delivering information to a surgeon, performing image and / or video acquisition, training or running one or more machine learning models, and / or generating medical inferences based on the one or more machine learning models. In other instances, a plurality of computer systems or computing devices may be used to perform different functions (including, for example, delivering information to a surgeon, performing image and / or video acquisition, training or running one or more machine learning models, and / or generating medical inferences based on the one or more machine learning models). Utilizing two or more computer systems or computing devices may enable parallel processing of tasks or information to provide real time surgical inferences and / or machine learning based surgical guidance.

[0221] The computer system 1201 may include a central processing unit (CPU, also "processor" and "computer processor" herein) 1205, which can be a single core or multi core processor, or a plurality of processors for parallel processing. The computer system 1201 may also include one or more graphical processing units (GPUs) 406 and / or digital signal processors (DSPs) 407, memory or memory location 1210 (e.g., random-access memory, read-only memory, flash memory), electronic storage unit 1215 (e.g., hard disk), communication interface 1220 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 1225, such as cache, other memory, data storage and / or electronic display adapters. The memory 1210, storage unit 1215, interface 1220 and peripheral devices 1225 are in communication with the CPU 1205 through a communication bus (solid lines), such as a motherboard. The storage unit 1215 can be a data storage unit (or data repository) for storing data. The computer system 1201 can be operatively coupled to a computer network ("network") 1230 with the aid of the communication interface 1220. The network 1230 can be the Internet, an internet and / or extranet, or an intranet and / or extranet that is in communication with the Internet. The network 1230 in some cases is a telecommunication and / or data network. The network 1230 can include one or more computer servers, which can enable distributed computing, such as cloud computing. The network 1230, in some cases with the aid of the computer system 1201, can implement a peer-to- peer network, which may enable devices coupled to the computer system 1201 to behave as a client or a server.

[0222] The CPU 1205 can execute a sequence of machine-readable instructions, which can be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 1210. The instructions can be directed to the CPU 1205, which can subsequently program or otherwise configure the CPU 1205 to implement methods of the present disclosure.Examples of operations performed by the CPU 1205 can include fetch, decode, execute, and writeback.

[0223] The CPU 1205 can be part of a circuit, such as an integrated circuit. One or more other components of the system 1201 can be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).

[0224] The storage unit 1215 can store files, such as drivers, libraries, and saved programs. The storage unit 1215 can store user data, e.g., user preferences and user programs. The computer system 1201 in some cases can include one or more additional data storage units that are located external to the computer system 1201 (e.g., on a remote server that is in communication with the computer system 1201 through an intranet or the Internet).

[0225] The computer system 1201 can communicate with one or more remote computer systems through the network 1230. For instance, the computer system 1201 can communicate with a remote computer system of a user (e.g., a doctor, a surgeon, a medical worker assisting or performing a surgical procedure, etc.). Examples of remote computer systems include personal computers (e.g., portable PC), slate or tablet PC's (e.g., Apple® iPad, Samsung® Gala4 Tab), telephones, Smart phones (e.g., Apple® iPhone, Android-enabled device, Blackberry®), or personal digital assistants. The user can access the computer system 1201 via the network 1230.

[0226] Methods as described herein can be implemented by way of machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system 1201, such as, for example, on the memory 1210 or electronic storage unit 1215. The machine executable or machine-readable code can be provided in the form of software. During use, the code can be executed by the processor 1205. In some cases, the code can be retrieved from the storage unit 1215 and stored on the memory 1210 for ready access by the processor 1205. In some situations, the electronic storage unit 1215 can be precluded, and machine-executable instructions are stored on memory 1210.

[0227] The code can be pre-compiled and configured for use with a machine having a processor adapted to execute the code or can be compiled during runtime. The code can be supplied in a programming language that can be selected to enable the code to execute in a pre-compiled or as- compiled fashion.

[0228] Aspects of the systems and methods provided herein, such as the computer system 1201, can be embodied in programming. Various aspects of the technology may be thought of as "products" or "articles of manufacture" typically in the form of machine (or processor) executable code and / or associated data that is carried on or embodied in a type of machine readable medium. Machine-executable code can be stored on an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk."Storage" type media can include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, disk drives and the like, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer into the computer platform of an application server. Thus, another type of media that may bear the software elements includes optical, electrical, and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links or the like, also may be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.

[0229] Hence, a machine readable medium, such as computer-executable code, may take many forms, including but not limited to, a tangible storage medium, a carrier wave medium or physical transmission medium. Non-volatile storage media including, for example, optical or magnetic disks, or any storage devices in any computer(s) or the like, may be used to implement the databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include for example: a floppy disk, a flexible disk, hard disk, magnetic tape, any other magnetic medium, a CD-ROM, DVD or DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium with patterns of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, cables or links transporting such a carrier wave, or any other medium from which a computer may read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0230] The computer system 1201 can include or be in communication with an electronic display 1235 that comprises a user interface (UI) 1240 for providing, for example, an augmented surgical view portal for a doctor or a surgeon to view surgical guidance or metrics in real time during asurgical procedure. The portal may be provided through an application programming interface (API). A user or entity can also interact with various elements in the portal via the UI. Examples of UI's include, without limitation, a graphical user interface (GUI) and web-based user interface.

[0231] Methods and systems of the present disclosure can be implemented by way of one or more algorithms. An algorithm can be implemented by way of software upon execution by the central processing unit 1205. For example, the algorithm may be configured to combine data streams obtained from one or more medical imaging units, extract one or more spatial or temporal features from the combined data streams, classify the features, and generate an augmented surgical view using the classifications to provide surgeons and medical personnel with real time guidance and metrics during critical phases of a surgical procedure.

[0232] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to the aforementioned specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it shall be understood that all aspects of the invention are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is therefore contemplated that the invention shall also cover any such alternatives, modifications, variations, or equivalents. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.

Claims

CLAIMSWHAT IS CLAIMED IS:

1. A method for annotating data generated from a surgical procedure, the method comprising: providing image data of a surgical procedure, wherein the image data is based at least in part on one or more frames of video data of the surgical procedure; implementing a multimodal, generative large language model, wherein the multimodal, generative large language model is connected to a visual encoder; and generating a written description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedure based at least in part on the multimodal, generative large language model, wherein the written description is based at least on multi -turn question and answer dialogue with the multimodal, generative large language model.

2. The method of claim 1, wherein the written description comprises reasoning for particular surgical steps or deduction of next steps.

3. The method of claim 1, wherein the multimodal, generative large language model is trained based at least in part on a library of annotations of surgical procedures.

4. The method of claim 3, wherein the library of annotations of surgical procedures comprises action triplets.

5. The method of claim 4, wherein the action triplets comprise an indication of an instrument, a verb, and a target in a surgical scene.

6. The method of claim 5, wherein the action triplets are further associated with a surgical phase.

7. The method of claim 4, wherein the multimodal, generative large language model is fine-tuned based on a ground truth dataset comprising ground truth image data and the action triplets.

8. The method of claim 1, wherein the image data comprises one or more members selected from the group consisting of white light imaging, fluorescence imaging, and perfusion imaging.

9. The method of claim 1, wherein the multimodal, generative large language model is a LLAVA, LLAVA-Med, or LLAVA Surg.

10. The method of claim 1, wherein the visual encoder is configured to (i) receive the one or more frames of video data obtained using one or more image and / or video acquisition modules and transform the one or more frames of video data into an input into the multimodal, generative large language model.

11. The method of claim 10, wherein the multimodal, generative large language model is configured (ii) generate the written description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedure based on (A) the one or more frames of video data and (B) one or more training data sets comprising surgical or medical data associated with a patient or a surgical procedure.

12. The method of claim 11, wherein the one or more training data sets comprise anatomical and physiological data obtained using a plurality of imaging modalities.

13. The method of claim 1, wherein the written description is generated in substantially real time with the surgical procedure.

14. The method of claim 1, wherein the annotations comprise billing coding of surgical procedures.

15. The method of claim 1, wherein the multimodal, generative large language model is a multi-modal video to text model.

16. The method of claim 1, further comprising annotating one or more frames of the image data with action triplets based at least in part on an output of the multimodal, generative large language model.

17. A method for annotating data generated from a surgical procedure, the method comprising: providing image data of a surgical procedure, wherein the image data is based at least in part on one or more frames of video data of the surgical procedure; implementing a multimodal, generative large language model, wherein the multimodal, generative large language model is connected to a visual encoder, wherein the large language model is trained based at least in part on a library of annotations of surgical procedures; and generating a written description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedure based at least in part on the multimodal, generative large language model, and wherein the written description comprises reasoning for particular surgical steps or deduction of next steps.

18. The method of claim 17, wherein the written description is based at least on multiturn question and answer dialogue with the large language and vision assistant.

19. The method of claim 17, wherein the written description comprises reasoning for particular surgical steps or deduction of next steps.

20. The method of claim 17, wherein the library of annotations of surgical procedures comprises action triplets.

21. The method of claim 20, wherein the action triplets comprise an indication of an instrument, a verb, and a target in a surgical scene.

22. The method of claim 21, wherein the action triplets are further associated with a surgical phase.

23. The method of claim 22, wherein the multimodal, generative large language model is fine tuned based on a ground truth dataset comprising ground truth image data and the action triplets.

24. The method of claim 17, further comprising using the multimodal, generative large language model to annotate frames of the image data with action triplets.

25. The method of claim 17, wherein the image data comprises one or more members selected from the group consisting of white light imaging, fluorescence imaging, and perfusion imaging.

26. The method of claim 17, wherein the multimodal, generative large language model is LLAVA, LLAVA-Med, or LLAVA Surg.

27. The method of claim 17, wherein the visual encoder is configured to (i) receive the one or more frames of video data obtained using one or more image and / or video acquisition modules and transform the one or more frames of video data into an input into the multimodal, generative large language model.

28. The method of claim 27, wherein the multimodal, generative large language model is configured (ii) generate the written description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedure based on (A) the one or more frames of video data and (B) one or more training data sets comprising surgical or medical data associated with a patient or a surgical procedure.

29. The method of claim 28, wherein the one or more training data sets comprise anatomical and physiological data obtained using a plurality of imaging modalities.

30. The method of claim 17, wherein the written description is generated in substantially real time with the surgical procedure.

31. The method of claim 17, wherein the annotations comprise billing coding of surgical procedures.

32. The method of claim 17, wherein the multimodal, generative large language model is a multi-modal video to text model.

33. The method of claim 17, wherein the annotations comprise billing coding of surgical procedures.

34. A system for annotating data generated from a surgical procedure, the system comprising:an imaging module configured to collect image data of the surgical procedure, wherein the image data is based at least in part on one or more frames of video data of the surgical procedure; and a processor configured to: implement a multimodal, generative large language model, wherein the multimodal, generative large language model is connected to a visual encoder; and generate a written description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedure based at least in part on the multimodal, generative large language model, wherein the written description is based at least on multi-turn question and answer dialogue with the multimodal, generative large language model.

35. A system for annotating data generated from a surgical procedure, the system comprising: an imaging module configured to collect image data of the surgical procedure, wherein the image data is based at least in part on one or more frames of video data of the surgical procedure; and a processor configured to: implement a multimodal, generative large language model, wherein the multimodal, generative large language model is connected to a visual encoder, wherein the large language model is trained based at least in part on a library of annotations of surgical procedures; and generate a written description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedure based at least in part on the multimodal, generative large language model, and wherein the written description comprises reasoning for particular surgical steps or deduction of next steps.

36. The system of claim 34 or 35, wherein the processor is further configured to annotate one or more frames of the image data with action triplets based at least in part on an output of the multimodal, generative large language model.

37. The system of claim 34 or 35, wherein the image data comprises one or more members selected from the group consisting of white light imaging, fluorescence imaging, and perfusion imaging.

38. The system of claim 34 or 35, wherein the multimodal, generative large language model is a LLAVA, LLAVA-Med, or LLAVA Surg.

39. The system of claim 34 or 35, wherein the visual encoder is configured to (i) receive the one or more frames of video data obtained using one or more image and / or videoacquisition modules and transform the one or more frames of video data into an input into the multimodal, generative large language model.

40. The system of claim 39, wherein the multimodal, generative large language model is configured (ii) generate the written description of anatomy, pathology, physiology, or a combination thereof observed during the surgical procedure based on (A) the one or more frames of video data and (B) one or more training data sets comprising surgical or medical data associated with a patient or a surgical procedure.

41. The system of claim 40, wherein the one or more training data sets comprise anatomical and physiological data obtained using a plurality of imaging modalities.

42. The system of claim 39, wherein the one or more image and / or video acquisition modules comprise (i) a first image and / or video acquisition module configured to capture images and / or videos using a first imaging modality and (ii) a second image and / or video acquisition module configured to capture images and / or videos using a second imaging modality.

43. The system of claim 34 or 35, wherein the one or more medical predictions or assessments comprise an identification or a classification of one or more critical structures.

44. The system of claim 34 or 35, wherein the written description comprises a location of one or more critical structures.

45. The system of claim 34 or 35, wherein the written description is configured provide the surgical operator with predictive clinical decision support before and during a surgical procedure.

46. The system of claim 34 or 35, wherein the multimodal, generative large language model is configured to update or refine the written description in real time based on additional image data obtained during a surgical procedure.

47. The system of claim 40, wherein the one or more training data sets comprise medical data associated with one or more reference surgical procedures.

48. The system of claim 47, wherein the one or more training data sets comprise medical data associated with (i) one or more critical phases or scenes of a surgical procedure or (ii) one or more views of a critical structure that is visible or detectable during the surgical procedure.

49. The system of claim 47, wherein the one or more training data sets comprise medical data obtained using a laparoscope, a robot assisted imaging unit, or an imaging sensor configured to generate anatomical images and / or videos in a red-green-blue visual spectrum.

50. The system of claim 47, wherein the one or more training data sets comprise medical data obtained using an imaging sensor configured to generate physiologic or functional imaging data based on chemical signal enhancers.

51. The system of claim 50, wherein the chemical signal enhancers comprise ICG, fluorescent, or radiolabeled dyes.

52. The system of claim 47, wherein the one or more training data sets comprise medical data obtained using an imaging sensor configured to generate physiologic or functional imaging data based on laser speckle patterns.

53. The system of claim 47, wherein the one or more training data sets comprise medical data obtained using an imaging sensor configured to generate physiologic or functional imaging data, wherein said physiologic or functional imaging data comprises two or more coregistered images or videos.

54. The system of claim 47, wherein the one or more training data sets comprise medical data obtained using an imaging sensor configured to generate physiologic or functional imaging data, wherein said physiologic or functional imaging data comprises segmented anatomic positions of critical structures of surgical interest.

55. The system of claim 47, wherein the one or more training data sets comprise medical data obtained using a plurality of imaging modalities that enable direct objective correlation between physiologic images or videos and corresponding RGB images or videos, wherein the plurality of imaging modalities comprises depth imaging, distance mapping, intraoperative cholangiograms, angiograms, ductograms, ureterograms, or lymphangiograms.

56. The system of claim 40, wherein the training data set comprises perfusion data.

57. The system of claim 34 or 35, wherein the image data obtained using one or more image and / or video acquisition modules comprises perfusion data.

58. A method of image-to-text generation, the method comprising: providing one or more images of a surgical procedure to a generative artificial intelligence (Al) model, wherein the one or more images are based at least in part on a video of the surgical procedure; and generating text associated with the one or more images using the generative Al model, wherein the text comprises a description of one or more intraoperative scenes in the one or more images, wherein the written description comprises reasoning for particular surgical steps or deduction of next steps, and wherein the written description is based at least on multi-turn question and answer dialogue with the generative Al model.

59. The method of claim 58, wherein the description comprises a descriptor of type(s) of surgical tools used, and a manner in which the surgical tools are being used during the surgical procedure.

60. The method of claim 58, wherein the description comprises a descriptor of one or more anatomical features within the one or more intraoperative scenes, and changes to the one or more anatomical features during the surgical procedure.

61. The method of claim 58, wherein the description comprises a descriptor of one or more intraoperative complications that occurred during the surgical procedure.

62. The method of claim 58, further comprising: identifying one or more billing codes from a plurality of billing codes for the surgical procedure, based at least in part on the text comprising the description of the one or more intraoperative scenes.

63. The method of claim 62, wherein the one or more billing codes are identified at an accuracy of greater than about 80%.

64. The method of claim 58, wherein the generative Al model comprises a large language model (LLM).

65. The method of claim 58, wherein the generative Al model is used for processing visual data and / or text from the one or more images.

66. The method of claim 58, wherein the generative Al model automatically generates the text in response to one or more questions provided via a graphical interface.

67. The method of claim 66, wherein the one or more questions are provided by one or more users via the graphical interface.

68. The method of claim 67, wherein the one or more questions are automatically generated based at least in part on a stored library of questions associated with the surgical procedure.

69. The method of claim 67, wherein the one or more questions are selected from a plurality of questions presented to one or more users via the graphical interface.

70. The method of claim 58, wherein the one or more images are annotated.

71. The method of claim 58, wherein the one or more images are non-annotated.

72. The method of claim 58, wherein the one or images are extracted from a video of the surgical procedure.

73. The method of claim 58, wherein the one or images are taken at different points in time during the surgical procedure.

74. The method of claim 58, wherein the one or images are taken at different locations and / or perspectives during the surgical procedure.

75. The method of claim 58, wherein the one or images are taken using two or more different imaging modalities.

76. The method of claim 75, wherein the two or more different imaging modalities comprise at least one of white light imaging or fluorescence imaging.

77. The method of claim 58, wherein the one or images comprises a plurality of images captured at different wavelengths.

78. The method of claim 58, wherein the one or more images comprises perfusion data.

79. The method of claim 58, wherein the surgical procedure comprises a laparoscopic procedure.

80. The method of claim 58, wherein the generative Al model is trained based at least in part on a library of annotations of surgical procedures.

81. A system for video-to-text generation, the system comprising: a processor configured to: receive one or more images of a surgical procedure and direct the one or more images to a generative artificial intelligence (Al) model, wherein the one or more images are based at least in part on a video of the surgical procedure; and implement the generative Al model to automatically generate text associated with the one or more images, wherein the text comprises a description of one or more intraoperative scenes in the one or more images, wherein the written description comprises reasoning for particular surgical steps or deduction of next steps, and wherein the written description is based at least on multi-turn question and answer dialogue with the large language and vision assistant.

82. The system of claim 81, wherein the generative Al model is trained based at least in part on a library of annotations of surgical procedures.

83. The system of claim 81, wherein the description comprises a descriptor of type(s) of surgical tools used, and a manner in which the surgical tools are being used during the surgical procedure.

84. The system of claim 81, wherein the description comprises a descriptor of one or more anatomical features within the one or more intraoperative scenes, and changes to the one or more anatomical features during the surgical procedure.

85. The system of claim 81, wherein the description comprises a descriptor of one or more intraoperative complications that occurred during the surgical procedure.

86. The system of claim 81, further comprising: identifying one or more billing codes from a plurality of billing codes for the surgical procedure, based at least in part on the text comprising the description of the one or more intraoperative scenes.

87. The system of claim 86, wherein the one or more billing codes are identified at an accuracy of greater than about 80%.

88. The system of claim 81, wherein the generative Al model comprises a large language model (LLM).

89. The system of claim 81, wherein the generative Al model is used for processing visual data and / or text from the one or more images.

90. The system of claim 81, wherein the generative Al model automatically generates the text in response to one or more questions provided via a graphical interface.

91. The system of claim 90, wherein the one or more questions are provided by one or more users via the graphical interface.

92. The system of claim 91, wherein the one or more questions are automatically generated based at least in part on a stored library of questions associated with the surgical procedure.

93. The system of claim 92, wherein the one or more questions are selected from a plurality of questions presented to one or more users via the graphical interface.

94. The system of claim 81, wherein the one or more images are annotated.

95. The system of claim 81, wherein the one or more images are non-annotated.

96. The system of claim 81, wherein the one or images are extracted from a video of the surgical procedure.

97. The system of claim 81, wherein the one or images are taken at different points in time during the surgical procedure.

98. The system of claim 81, wherein the one or images are taken at different locations and / or perspectives during the surgical procedure.

99. The system of claim 81, wherein the one or images are taken using two or more different imaging modalities.

100. The system of claim 99, wherein the two or more different imaging modalities comprise at least one of white light imaging or fluorescence imaging.

101. The system of claim 81, wherein the one or images comprises a plurality of images captured at different wavelengths.

102. The system of claim 81, wherein the one or more images comprises perfusion data.

103. The system of claim 81, wherein the surgical procedure comprises a laparoscopic procedure.

Citation Information

Patent Citations

  • Language model-based image report generation method and system

    CN116884559A

  • Sequencing medical codes methods and apparatus

    US20180081859A1

  • Surgery planning

    US20180360543A1

  • System and method for augmented reality guidance for use of equipment systems

    US20210295048A1

  • Intraoperative clinical decision support system

    US20220157423A1

Cited By

  • Operation video question and answer task-oriented training data set construction and intelligent reasoning method

    CN121616911A