Deep learning systems and methods for integrating, displaying, and generating inferences from surgical information from disparate medical devices

EP4750413A1Pending Publication Date: 2026-06-03DIGITAL SURGERY SYSTEMS INC

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
DIGITAL SURGERY SYSTEMS INC
Filing Date
2024-07-26
Publication Date
2026-06-03

AI Technical Summary

Technical Problem

Deploying deep learning models in a surgical context is challenging due to the vast amounts of data from disparate medical devices that need to be captured, filtered, processed, and integrated in real-time or near real-time, while also being burdensome for surgeons and medical staff to track the overall status of the patient and procedure progress.

Method used

The system integrates data from disparate medical devices using sensors and data gathering modules to process and transmit image, video, and audio data to a computing device, where deep learning models are applied to generate inferences and display data in a surgical cockpit environment.

Benefits of technology

This approach effectively transforms and integrates data from disparate medical devices, enabling the training and application of deep learning models in real-time surgical environments, thereby improving patient outcomes and reducing the burden on surgical staff.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024039786_30012025_PF_FP_ABST
    Figure US2024039786_30012025_PF_FP_ABST
Patent Text Reader

Abstract

Deep learning based systems and methods are disclosed for integrating, displaying, and drawing inferences from surgical information from disparate medical devices. An example system comprises: a plurality of medical devices, each comprising an interface for outputting image, video, and / or audio data; a sensor associated with each medical device configured to capture raw image, video, and / or audio data from the respective interface; a plurality of data gathering modules configured to transmit processed image, video, and / or audio data from the plurality of medical devices to a computing device; the computing device comprising: a processor; and memory storing computer-executable instructions that, when executed by the processor, causes the system to: receive, in real-time, the processed image, video, and / or audio data from each of the plurality of medical devices; and apply a deep learning model to the processed image, video, and / or audio data to generate an output outcome.
Need to check novelty before this filing date? Find Prior Art

Description

DEEP LEARNING SYSTEMS AND METHODS FOR INTEGRATING, DISPLAYING, AND GENERATING INFERENCES FROM SURGICAL INFORMATION FROM DISPARATE MEDICAL DEVICESCROSS REFERENCE TO RELATED APPLICATION

[0001] The application claims priority to U.S. Provisional Patent Application No. 63 / 529,097, filed July 26, 2023, entitled “Deep Learning Systems And Methods For Integrating, Displaying, And Generating Inferences From Surgical Information from Disparate Medical Devices,” the contents of which are incorporated by reference herein in their entirety.BACKGROUND

[0002] Effectively deploying deep learning models in a surgical context is particularly challenging. For surgical procedures, the digitization of the operating room and support structures and processes has proceeded rapidly, resulting in an explosion of data and the opportunity to capture even more data. Capturing, filtering, and processing these vast amounts of data, and executing on such processed data is a challenge even for the best of computer systems, especially while doing so in real time or near real time while a surgical procedure is underway. Moreover, medical devices are also typically disparate in their physical locations relative to the surgeon at any given time. This arrangement is burdensome for surgeons and medical staff to keep track of the overall status of the patient and the progress of the procedure as collectively represented by all such devices.

[0003] Accordingly, there is a desire and need to more effectively transform and integrate various data from disparate medical devices and locations to more positively and effectively contribute to the training and application (e.g., inferencing) of deep learning models, especially in a surgical context.SUMMARY

[0004] Among other features, the present disclosure provides new and innovative deep learning based systems and methods for integrating, displaying, and drawing inferences from surgical information from disparate medical devices. Various embodiments of the present disclosure further describe systems and methods for receiving large data sets from medical devices, trainingand applying (via e.g., inferencing) deep learning models, and displaying such data (e g., in a surgical cockpit), for use in surgical environments.

[0005] In one embodiment, an example system may comprise a plurality of medical devices configured to measure patient-specific data from a patient, each medical device comprising an interface for outputting image, video, and / or audio data; a plurality of sensors comprising at least one sensor for each of the plurality of medical devices, the at least one sensor configured to capture raw image, video, and / or audio data from the respective interface of the respective medical device; a plurality of data gathering modules corresponding to the plurality of medical devices, each data gathering module configured to: generate, from the raw image, video, and / or audio data from the respective medical device, processed image, video, and / or audio data from the respective medical device; and transmit the processed image, video, and / or audio data from the respective medical device to a computing device; a computing device comprising: a processor; and memory storing computer-executable instructions that, when executed by the processor, causes the system to: receive, in real-time, processed image, video, and / or audio data from each of the plurality of medical devices; and apply a deep learning model to the processed image, video, and / or audio data to generate an output outcome.

[0006] In an embodiment, the plurality of medical devices may include a digital surgical microscope system. The digital surgical microscope system comprises at least one camera configured to capture raw image, video, and / or audio data from a surgical site of the patient. The data gathering module of the digital surgical microscope system may be further configured to: generate, from the raw image, video, and / or audio data from the digital surgical microscope, processed image, video, and / or audio data from the digital surgical microscope; and transmit the processed image, video, and / or audio data from the respective medical device to a computing device.

[0007] In an embodiment, the system may further comprise: one or more monitors accessible to a surgeon. The one or more monitors may be configured to display the processed image, video, and / or audio data from one or more of the plurality of medical devices.

[0008] In an embodiment, the computer-executable instructions, when executed by the processor, further causes the system to, prior to applying the deep learning model: train the deep learning model using a training dataset comprising a plurality of reference input image, video, and / or audio data and reference output outcomes.

[0009] In an embodiment, the at least one sensor comprises one or more of a camera or a microphone.

[0010] In an embodiment, the at least one of the plurality of medical devices may comprise a QR code located within a field of view of the respective sensor of the medical device, wherein the respective sensor comprises a camera.

[0011] In an example, an apparatus for integrating, displaying, and generating inferences from surgical information from a digital surgical microscope (DSM) system is disclosed. The apparatus comprises the DSM system, which comprises at least one camera and a data gathering module. The at least one camera is configured to capture raw image, video, and / or audio data from a surgical site of the patient. The data gathering module of the DSM system is configured to: generate, from the raw image, video, and / or audio data from the digital surgical microscope, processed (e.g., unwarped, digitized, quantified, etc.) image, video, and / or audio data from the digital surgical microscope; and transmit the processed image, video, and / or audio data from the respective medical device to a computing device. The apparatus further comprises a computing device comprising: a processor; and memory storing computer-executable instructions that. When executed by the processor, the computer-executable instructions cause the computing device to: receive, in real-time, processed image, video, and / or audio data from each of the plurality of medical devices; and apply, a deep learning model to the processed image, video, and / or audio data to generate an output outcome.

[0012] In an embodiment, the computer-executable instructions, when executed by the processor, further causes the computing device to, prior to applying the deep learning model: train the deep learning model using a training dataset comprising a plurality of reference input image, video, and / or audio data and reference output outcomes.

[0013] In an embodiment, the apparatus further comprises: one or more additional cameras associated with the DSM system. The one or more additional cameras are configured to capture additional raw image, video, and / or audio data from a field of view that is greater than the surgical site of the patient. The processed image, video, and / or audio data from the digital surgical microscope is further generated from the additional image, video, and / or audio data captured by the one or more additional cameras.

[0014] In an example, computer-executable methods are disclosed to describing one or more steps, methods, or processes described herein.

[0015] In an example, a non-transitory computer-readable medium for use on a computer system is disclosed. The non-transitory computer-readable medium may contain computer-executable programming instructions may cause processors to perform one or more steps or methods described herein.BRIEF DESCRIPTION OF THE FIGURES

[0016] FIG. 1 illustrates a diagram of a system for integrating, displaying, and generating inferences from surgical information from disparate medical devices using advanced deep learning techniques, according to an example embodiment of the present disclosure.

[0017] FIG. 2 shows a diagram of a surgical environment including a digital surgical microscope camera and deep learning based surgical system, according to an example embodiment of the present disclosure.

[0018] FIG. 3 is an example of a digital surgical microscope system configured to include data collection sensors, in accordance with a non-limiting embodiment of the present disclosure.

[0019] FIG. 4A shows a side view of an augmented medical device used for the deep learning based surgical system, in accordance with a non-limiting embodiment of the present disclosure.

[0020] FIG. 4B shows an isometric view of an augmented medical device used for the deep learning based surgical system, in accordance with a non-limiting embodiment of the present disclosure.

[0021] FIG. 5 shows a pictorial representation of an example method of processing data (e.g., image, video, and / or audio data) captured from human-intended interfaces on medical devices, in accordance with a non-limiting embodiment of the present disclosure.

[0022] FIG. 6 shows an example labeling process of a human intended interface, according to an example embodiment of the present disclosure.

[0023] FIG. 7 shows digitized medical data superimposed on a monoscopical or stereoscopical view of a surgical site, according to an example embodiment of the present disclosure.DETAILED DESCRIPTION

[0024] Various aspects of the present disclosure will be described with reference to the drawings, wherein like reference numerals are used to refer to like elements throughout. In the following description, for purposes of explanation, numerous specific details are set forth in order to promote a thorough understanding of one or more aspects of the present disclosure. It may be evident insome or all instances, however, that any aspects described below can be practiced without adopting the specific design details described below.

[0025] Referring to FIG. 1, in accordance with aspects, a system 10 deployed within a serverbased computing environment and communication network 20 may be configured to obtain, transform, integrate, and display various data from disparate medical devices and locations to train and apply deep learning models, especially in a surgical context. The system 10 may be configured to ingest any existing centralized data e.g. from a “tower” of devices from the same manufacturer that already shares digitized data. For example, scope towers may be used in various medical procedures, such as endoscopy, laparoscopy, and arthroscopy. These towers provide medical professionals with a comprehensive solution that enables them to achieve visualization during surgeries. Medical scope towers may be equipped with high-resolution displays, light sources, and cameras that provide high-quality images to doctors and surgeons. These towers may also have the capability to record and store images and videos, which can be used for diagnostic and educational purposes.

[0026] In one embodiment, one or more individuals 102 (e.g., medical professionals, medical researchers, health providers, caregivers, trained staff (e.g., nurses, support staff, monitor technicians), and other end-users) may use or access at least one computing system 14 configured to exchange information and data with a data acquisition system 16 and a computing server 22 based on advanced artificial intelligence and machine learning / deep learning techniques and models.

[0027] Artificial intelligence, such as deep learning, may involve training a mathematical model to calculate a set of output data from a set of input data. The output data may represent a desired processing of the input data, such as detection and classification in a manner more rapid and accurate than prior implementations.

[0028] This set of calculations may be referred to as inferencing. Such inferencing offers a means to process a large set of often real-time, high bandwidth data to extract useful information that might otherwise be missed or may be too costly to duplicate in one or more ways. The information thus extracted can help to improve patient outcomes to avoid a degradation in care.

[0029] The task of developing a deep learning system may comprise one or more of the following steps: determining input data to be processed; accessing or making the input data available as needed (e.g., transforming, vectorizing, and / or digitizing the input data to render it as “compute-friendly” (i.e., computer-friendly and deep learning-friendly) as appropriate); specifying a desired output data (e.g., transforming, vectorizing, and / or digitizing the output data to be computefriendly); designing the form of a deep learning model (e.g., by examining the respective forms of the input data and output data and designing a network of mathematical operations with a plurality of coefficients (i.e., weights) that operate on the inputs to calculate the outputs); gathering representative sets of input data; processing and / or associating the input data sets with output data sets (“labeling”) (e.g., such that the results may be considered as ground truth data sets); dividing the input data sets into training and test portions (training dataset and testing dataset, respectively); training the deep learning model using the training data sets (e.g., resulting in a set of coefficients or weights for the deep learning model); testing the deep learning model using the testing data sets; iterating, as necessary and until sufficiently acceptable results are achieved in the testing data set; and deploying the deep learning model using the trained deep learning model, coefficients, and weights thus found (e g., by feeding input data similar to that which was used to train the deep learning model).

[0030] Effectively deploying deep learning models in a surgical context has been proven to be particularly challenging. For example, deep learning systems often receive, access, and / or make available vast amounts of appropriate data (e.g., transforming, vectorizing, and / or digitizing the input data to render it as “compute-friendly” (i.e., computer-friendly and deep learning-friendly) as appropriate)). The deep learning systems are then often required to use such data to train a deep learning model, and then receive same or similar data and supply such data to the deep learning model during runtime (i.e., during the inferencing phase). Furthermore, during a typical surgical procedure, there are usually many disparate existing medical devices in an operating room that may each generate large datasets (e.g., by monitoring patient biometric statistics). As will be described fully below, the systems and methods of the present disclosure may be configured to more effectively transform and integrate such data from disparate medical devices to more positively and effectively contribute to the training and application (e.g., inferencing) of selected deep learning models.

[0031] Furthermore, medical devices are typically disparate in their physical locations relative to the surgeon at any given time. This arrangement is burdensome for surgeons and medical staff to keep track of the overall status of the patient and the progress of the procedure as collectively represented by all such devices. The present disclosure effectively addresses particular challengesof developing and deploying deep learning models in a surgical context. The deep learning based system of the present disclosure may be configured to receive, access and / or make available vast amounts of appropriate data (e.g., transforming, vectoring, and / or digitizing the input data to render it as “compute-friendly” (i.e., computer-friendly and deep learning-friendly) as appropriate)), use the data to train a deep learning model, and receive same or similar data and supplying them to the deep learning model (i.e., during the inferencing phase), especially in real time or near real-time during a surgical procedure. As a result, the present disclosure overcomes the burden to surgeons and medical staff of attending to, making sense of, and tracking the data provided by such disparate medical devices situated in disparate physical locations relative to the surgeon.

[0032] In some aspects, the system 10 of the present disclosure may be configured to obtain large data sets from surgical and medical devices, train and apply (via e.g., inferencing) selected deep learning models, and display such data, for use in surgical environments. As shown in FIG. 1, a data acquisition system 16 of the system 10 may be configured to implement one or more sensors (e.g., image sensors, microphones, temperature sensors, humidity sensors, light sensors) 18 a2, 18 b2, ... 18 n2 to capture at least raw image, video, and / or audio data from a corresponding interface of each medical device. In one embodiment, the system 10 of the present disclosure may augment a digital surgical microscope (DSM) to serve as a comprehensive surgical data gathering and inferencing platform (e.g., part of the data acquisition system 16 or the computing system 14). The data obtained by the data acquisition system 16 may be used to train and develop one or more deep learning models hosted or incorporated by the computing server 22. In certain implementations, the deep learning models may be deployed on the inferencing portion of the system 10 and are supplied with data obtained via the data acquisition system 16 to generate outcomes via inferencing. Furthermore, the system 10 may enable an individual 102 to manage, select, and display any desired output data via e.g., a user interface associated with the computing system 14. The systems and methods of the present disclosure may enable the capture, management, and display of data from the many various medical devices in use during the surgical procedure.

[0033] In one aspect, the data acquisition system 16 may be configured to transmit data to a selected computing device or system (e.g., the computing system 14) and / or the computing server 22 via the communication network 20 and communication protocols 20a, 20b, 20c. According to one embodiment, an application, which may be a mobile or web-based application (e.g., nativeiOS or Android Apps), may be downloaded and installed on the selected computing device or system 14 for instantiating various modules for processing and analyzing various data received from the data acquisition system 16, and interacting with the user 12 of the application, among other features. For example, such an application may be used by surgeons and medical staff or team in real time or near real-time during a surgical procedure. Such a user-facing application of the system 10 may include a plurality of modules executed and controlled by the processor of the computing system 14 for processing various signals received from the data acquisition system 16 using various algorithms, as will be described fully below. The computing system 14 hosting the mobile or web-based application may be configured to connect, using a suitable communication protocols 20a, 20b, 20c and network 20, to the computing server 22. Here, communication network 20 may generally include a geographically distributed collection of computing devices or data points interconnected by communication links and segments for transporting signals and data therebetween. Communication protocol(s) 20a, 20b, 20c may generally include a set of rules defining how computing devices and networks may interact with each other, such as frame relay, Internet protocol (IP), transmission control protocol (TCP), user datagram protocol (UDP), hypertext transfer protocol (HTTP). It should be appreciated that the system 100 of the present disclosure may use any suitable communication network, ranging from local area networks (LANs), wide area networks (WANs), cellular networks, to overlay networks and software-defined networks (SDNs), a packet data network (e.g., the Internet), mobile telephone networks (e.g., cellular networks, such as 4G or 5G), Plain Old Telephone (POTS) networks, and wireless data networks (e.g., Institute of Electrical and Electronics Engineers (IEEE) 802.11 family of standards known as Wi-Fi®, WiGig®, IEEE 802.16 family of standards known as WiMax®), IEEE 802.15.4 family of standards, a Long Term Evolution (LTE) family of standards, a Universal Mobile Telecommunications System (UMTS) family of standards, peer-to-peer (P2P) networks, virtual private networks (VPN), Bluetooth, Near Field Communication (NFC), or any other suitable network.

[0034] In some embodiments, the computing server 22 may be Cloud-based or an on-site server. The term “server” generally refers to a computing device or system, including processing hardware and process space(s), an associated storage medium such as a memory device or database, and, in some instances, at least one database application as is well known in the art. The computing server 22 may provide functionalities for any connected devices such as sharing data or provisioningresources among multiple client devices or performing computations for each connected client device. According to one embodiment, within a Cloud-based computing architecture, the computing server 22 may provide various Cloud computing services using shared resources. Cloud computing may generally include Internet-based computing in which computing resources are dynamically provisioned and allocated to each connected computing device or other devices on-demand, from a collection of resources available via the network or the Cloud. Cloud computing resources may include any type of resource, such as computing, storage, and networking. For instance, resources may include service devices (firewalls, deep packet inspectors, traffic monitors, load balancers, etc.), computing / processing devices (servers, central processing units (CPUs), graphics processing units (GPUs), random access memory, caches, etc.), and storage devices (e.g., network attached storages, storage area network devices, hard disk drives, solid-state devices, etc.). In addition, such resources may be used to support virtual networks, virtual machines, databases, applications, etc. The term “database,” as used herein, may refer to a database (e.g., relational database management system (RDBMS) or structured query language (SQL) database), or may refer to any other data structure, such as, for example a comma separated values (CSV), tab-separated values (TSV), JavaScript Object Notation (JSON), eXtendible markup language (XML), TeXT (TXT) file, flat file, spreadsheet file, and / or any other widely used or proprietary format. In some embodiments, one or more of the databases or data sources may be implemented using one of relational databases, flat file databases, entity-relationship databases, object-oriented databases, hierarchical databases, network databases, NoSQL databases, and / or record-based databases.

[0035] Cloud computing resources accessible using any suitable communication network (e.g., Internet) may include a private Cloud, a public Cloud, and / or a hybrid Cloud. Here, a private Cloud may be a Cloud infrastructure operated by an enterprise for use by the enterprise, while a public Cloud may refer to a Cloud infrastructure that provides services and resources over a network for public use. In a hybrid Cloud computing environment which uses a mix of on-premises, private Cloud and third-party, public Cloud services with orchestration between the two platforms, data and applications may move between private and public Clouds for greater flexibility and more deployment options. Some example public Cloud service providers may include Amazon (e.g., Amazon Web Services® (AWS)), IBM (e.g., IBM Cloud), Google (e.g., Google Cloud Platform), and Microsoft (e.g., Microsoft Azure®). These providers provide Cloud services using computingand storage infrastructures at their respective data centers and access thereto is generally available via the Internet. Some Cloud service providers e.g., Amazon AWS Direct Connect, Microsoft Azure ExpressRoute) may offer direct connect services and such connections typically require users to purchase or lease a private connection to a peering point offered by these Cloud providers.

[0036] In certain implementations, the computing server 22 (e.g., Cloud-based or an on-site server) of the present disclosure may be configured to connect with various data sources or services 24a, 24b, 24c, . . . 24n. For example, in addition to obtaining data from the computing system 14 and the data acquisition system 16, the computing server 22 may be configured to collect data from one or more of 24a, 24b, 24c, . .. 24n to form a dataset for training the machine learning model of the present disclosure. Exemplary data source may include existing company data banks, company systems in the field, and 3rdpart data repositories. All data used for the machine learning development may be subject to the following considerations for quality and efficacy: anonymization, validity, completeness, permission, and anti-bias.

[0037] In one aspect, the computing server 22 may be configured to host, train, operate, and / or incorporate any suitable type of deep learning models (e.g., at least one of 24a, 24b, 24c, . .. 24n) to process image, video, and / or audio data received from each of the plurality of medical devices 18 al, 18 b 1, ... 18 nl and corresponding sensors 18 a2, 18 b2, . . . 18n2, and generate an output outcome. In one embodiment, the computing server 22 may include an application programming interface (API) interface configured to make one or more API calls therethrough. On the other hand, the computing server 22 may include an API gateway device (not shown) configured to receive and process API calls from various connected computing devices deployed within the system 10 (e.g., an operating system, a library, a device driver, an API, an application program, software or other module). Such an API gateway device may specify one or more functions, methods, classes, objects, protocols, data structures, formats and / or other features of the computing server 20 that may be used by the mobile or web-based application of the computing system 14. For example, the API interface or gateway device may define at least one calling convention that specifies how a function associated with the computing server 22 receives data and parameters from a requesting device / system and how the function returns a result to the requesting device / system. It should be appreciated that the computing server 22 may include additional functions, methods, classes, data structures, and / or other features that are not specified through the API interface or gateway device and are not available to a requesting computing device.

[0038] In accordance with certain aspects, the computing server 22 may be configured to apply a deep learning model (e.g., one of 24a, 24b, 24c, . . . 24n) for each medical device 18 al, 18 b 1 , ... 18 nl of the data acquisition system 16. That is, each individual medical device from which the system 10 gathers data may have a deep learning model apply thereto in order to convert data obtained from the associated human intended interface to computer readable (and AI / DL-ready) digitized data. For example, convolutional neural networks (CNNs) may be used to process image data including object recognition, graphical display conversion, text recognition. For text recognition, machine learning optical character recognition techniques may be used. Other example text recognition models may include MMOCR, PaddleOCR and CRNN.

[0039] In accordance with additional aspects, either or both the computing system 14 and the data acquisition system 16 may include a patient monitor or user interface to provide e.g., electronic monitoring and display of a patient’s physiological conditions during a surgical process, electronic management of all related medical devices 18 al, 18 bl, . . . 18 nl, and display an output outcome generated by the deep learning model based at least upon the processing results of the image, video, and / or audio data from each of the plurality of medical devices 18 al, 18 bl, . .. 18 nl of the data acquisition system 16. For example, electronic monitoring of basic patient vital signs may include blood pressure, blood oxygen saturation (oximetry), carbon dioxide levels in the patient’s inhaled and exhaled gases (capnometry), etc. According to some embodiments, a DSM, utilized as a comprehensive surgical data gathering and inferencing platform, may be configured to guide a surgeon, via the patient monitor or user interface, to navigate to a specific surgical site of a patient. The generated output outcome may be a combination of the digitized medical data from the sensors 18 a2, 18 b2, ... 18 n2 overlaid on video recorded by the stereoscopic camera of the DSM. The computing server 22 or the surgical system 120 of FIG. 2 may be configured to add the digitized medical data to the video before it is displayed. As a result, the surgeon may rely upon digital information obtained before surgery such as MRI or CAT scans and / or real-time image, video, and / or audio data (e.g., fluoroscopic X-rays) to achieve more accurate surgical procedures with minimal incisions and fewer surgical complications. In addition to providing a live stereoscopic view, relevant and useful patient information may be provided on the same display. As such, a surgeon does not have to look around the operating room for certain medial data. Since the images are presented monoscopically or stereoscopically, in one embodiment, the computing server 22 or the surgical system 120 may be configured to render the medical data so it is also viewedmonoscopically or stereoscopically in correspondence with the image. In yet another embodiment, the output outcome may include surgical device placement recommendations or surgical procedure recommendations. For example, the deep learning model of the present disclosure may be configured to analyze received surgical procedure data, patient data, targeted postoperative outcomes and other relevant information to identify correlations between surgical device placement and specific surgical procedures and adverse events, or device placement or procedures and positive post-operative outcomes. Such recommendations may be presented to the surgeon via a selected output format provided by the patient monitor or user interface. Example output formats may include visual, audio and / or other sensorial or perceptive (e.g., tactile, gustatory, haptic, pressure-sensing-based or electromagnetic (e.g., neurostimulation) communications (e.g., via a computer, a smartphone, or a tablet). According to one embodiment, natural sounding audio signals may be generated to represent the recommendations. For example, the computing server 22 or the computing system 120 may be configured to apply the deep learning model(s) to analyze received data based on a known surgical procedure to provide recommendations on the display screen overlaid on the stereoscopic video. In some embodiments, alarm signals may be only enabled and displayed after detecting from the video that a patient’s internal tissue has been exposed, indicating the start of the critical phase of the surgery.

[0040] According to further embodiments, a record may be created of the video data in conjunction with the medical data to enable e.g., playback as well as comprehensive recording of all data of interest during the procedure time-synchronized to the surgical video and including said video. Such a record may be stored in the system 10, shared with other computing devices deployed within the system 10, and / or used to form a dataset for training and testing a macro deep learning model of the present disclosure. An example dataset may include intraoperative surgical site video recordings of various procedures undergone by various patients. In an embodiment, the patient monitor or user interface may be configured to allow a surgeon to highlight and annotate a video recording in real time as it is being recorded. In addition, the computing server 22 or the computing system 120 may apply the deep learning model to search through individual patient or surgical records, identify relevant records, and provide a surgeon with this relevant information for a particular surgical procedure.

[0041] FIG. 2 shows a diagram of a surgical environment 100 including a digital surgical microscope camera 106 and deep learning based surgical system 120 (e.g., computing system 14or a combination of computing system 14 and computing server 22 in FIG. 1), according to an example embodiment of the present disclosure. The surgical environment 100 may include a patient 102 that is subjected to or intended to be subjected to a surgical procedure, surgical tools 104A and 104B, a surgical camera 106) (also referred to as “digital surgical microscope (DSM)” “surgical camera,” or “camera”), a computing system 120 configured for performing deep learning based methods described herein (also referred to herein as “deep learning system” or “deep learning based surgical system”) robotic controls (e.g., a foot pedals 108), one or more robotic arms 114 (including monitor boom mast 116 and / or any boom arm 118), surgical cart 124, a surgeon 112, one or more surgeon assistants (e.g., nurses) 122, surgical tools (e.g., 104A, 104B, etc.), data collection sensors (e.g., additional cameras 126A, microphones 126B, etc.), and one or more displays 132A-132B. The data collection sensors of the surgical environment may correspond to the sensors 18 a2, 18 b2, ... 18n2 of the data acquisition system 16 in FIG. 1.

[0042] In some embodiments, the digital surgical microscope system that includes the DSM 106 can be augmented for data capture for the deep learning based system 120. For surgical procedures in which the surgical camera 106 is used, the video of the surgical site 128 during the procedure is a particularly relevant input data for the deep learning based system 120. The video may show the surgical tools 104A and 104B used and the manner in which they are deployed during the surgical procedure. The DSM 106 may be an important tool for data capture for the deep learning based system 120.

[0043] The magnification range of DSM 106 may limit the field of view of the DSM 106. Further valuable input data for the deep learning based system 120 may be obtained by capturing image data from the region around the surgical site 128. This image and / or video capture may involve a wider field of view. The wider field of view may include but is not limited to the surgeon’s hands 130, additional surgical tools and controls, for example, the surgical assistant’s tools 104B and hands, and etc.

[0044] Expanding the field of view to enable surgical cameras to receive additional valuable input data may be available by “zooming out” again to generate the wider field of view to capture views, for example, of the scrub nurse table of tools, various operating room actors. Additionally, or alternatively, surgical cameras (the DSM 106 or other surgical cameras (e.g., camera 126A) as will be described herein) may be enabled to capture views from various poses, such as to form a 360 degree view of an area of interest.

[0045] The DSM head 106, one or more robot arms 114, monitor boom mast 1 16, boom arm 118, and cart base 124, and their respective poses in the operating room before, during and after surgery may be important locale for the placement of data collection sensors (e.g., cameras 106 and 126A and microphone(s) 126B). It should be appreciated that such sensors may include but are not limited to cameras and microphones.

[0046] In some embodiments, the DSM camera 106 may be the primary, predominant, or only data collection sensor. The DSM camera 106 may thus be used to provide image data (e.g., past image data) that can be relied on for the training and / or testing of one or more deep learning models. For example, prior to model training, the anonymized, validated, and labelled data may be split into separate repositories for model training / testing, and evaluation. Data splits may be made in a way that no unit of data exists in multiple splits. In one embodiment, training / test data may be the portion of data used during model training and testing. This data set may be split into training and testing datasets during the training and tuning process. The training data may include the portion of data that the model will ingest during model training. The test data may be used after each training session to evaluate model hyperparameters and perform tuning. Moreover, the model evaluation data may be used to independently test the tuned model.

[0047] The trained deep learning models may then be applied, in real-time, to incoming image data from the DSM camera 106 to provide recommendations to the surgeon. In embodiments, where the DSM camera 106 is a primary or predominant data collection sensor, additional sensors (e.g., cameras and / or microphones) may be added to areas around the DSM head 106 and the cart base 124, as will be described in relation to FIG. 3.

[0048] FIG. 3 is an example of a digital surgical microscope system configured to include data collection sensors (e.g., the data acquisition system 16 of FIG. 1), in accordance with a nonlimiting embodiment of the present disclosure. As used herein, a digital surgical microscope system may include but is not limited to the DSM camera head itself 106, and one or more of the robot arms 114, monitor boom mast 116, boom arm 118, cart base 124, the computing system 120 (e.g., deep learning based surgical system) having a processor 120A and memory 120B, surgical devices around the surgical site (e.g., surgical device 230), and their respective poses in the operating room before, during, and after surgery. As shown in FIG. 3, data collection sensors placed or embedded within various components of the digital surgical microscope system may include but are not limited to microphones and additional cameras.

[0049] In at least one embodiment, cameras placed on or embedded within the digital surgical camera system may include but are not limited to: the stereoscopic DSM camera 106 viewing the surgical site 128, one or more wider field of view cameras 202 on the underside of the DSM head 106 facing downward, one or more wider field of view cameras 204 on handles 206 of the DSM camera 106 facing downward, “surround” cameras 208 mounted around the DSM head 106 facing sideways (e.g., providing a 360 degree coverage through purposed drape window(s)) one or more camera(s) 210 on a robotic arm, one or more camera(s) 212 on a boom arm (e.g., a “room view” camera 212 at a top of a mast and providing a 360 degree coverage of the surroundings). Also or alternatively, a “room view” camera 220 providing a 360 degree coverage of the surroundings may be located on top of the monitor 132A at the end of the boom arm. In some embodiments, one or more of the above described cameras may be supported with lighting (e.g., near infrared lighting or visible lighting). For example, to minimize distraction and noise for image data capture, near infrared lighting may be used, resulting in deep learning model that may be trained using input data captured under infrared imaging. In some aspects, one or more of these cameras may be supported by a height extender. Furthermore, as will be described in relation to FIGS. 4A and 4B, additional cameras may also or alternatively be placed or affixed to other medical devices in the operating room, for example, to capture measurements recorded by the medical device.

[0050] In at least one embodiment, microphones placed on or embedded within the digital surgical camera system may include one or more microphones 126B, as shown in FIG. 2, dedicated to capturing surgeon voice. For example, such one or microphones may comprise a wearable, wireless microphone that communicates with the deep learning based system 120 (e.g., via a base station communication module connected to a main data gathering module). Additionally or alternatively, such one or more microphones may comprise a shotgun microphone or similar directional microphone aimed at a region where the surgeon is located for the majority of the procedure. Also or alternatively, such one or more microphones may comprise a microphone mounted on the DSM camera head 106. In some aspects, the one or more microphones may be configured to detect and / or capture voice commands intended to control voice-activated modules in the DSM 106.

[0051] The microphones placed on or embedded within the digital surgical camera system may further include but are not limited to: one or more microphones 214, as shown in FIG. 3, mounted on the underside of the DSM head 106 to capture sounds occurring in an around the surgical site;one or more microphones mounted at each camera location, to capture audio from the area(s) being captured on video; or dedicated “field” microphones 216 and 218 capturing warning and other indicator sounds from surgical devices in the operating room. In some aspects, one or more of the “field” microphones may comprise a shotgun microphone or a similar directional microphone mounted on the cart (e.g., “field” microphone 218) and pointed at a said surgical device, or may comprise a remote microphone mounted closer to the surgical device (e.g., as in “field” microphone 216) and connected to the DSM cart data gathering module, shown in FIGS. 3A and 3B.

[0052] In some embodiments, a set of data collection sensors (e.g., “field” cameras, “field” microphones, etc.) may be placed on, affixed to, or embedded within medical devices or instruments used during a surgical procedure and / or located in the operating room. The combined system may be referred to herein as an augmented medical device.

[0053] FIGS. 4A and 4B respectively show a side view and an isometric view of an augmented medical device used for the deep learning based surgical system, in accordance with a non-limiting embodiment of the present disclosure. As previously discussed, further valuable data may be gathered from existing medical devices or instruments involved directly or indirectly in the surgical procedure. The augmentation of existing medical device or instruments enables the ability to capture relevant data (e.g., to form deep learning training datasets) from a disparate array of existing involved devices without relying on any common or predetermined communication protocol or connection thereto.

[0054] In some embodiments, such predetermined communication protocols and connections may exist and may be used. However, in the commercial world of many product manufacturers and devices, there may be a lack of comprehensive standards to connect the wide range of devices present in a given operating room for a given procedure.

[0055] It is contemplated that various medical devices and instruments are typically designed with a “human-intended" interface, such as by having warning lights, one or more digital displays, an analog gauge or meter, and audio output. This characteristic of such medical devices and instruments may be relied on for developing augmented medical devices as described herein and as shown in FIGS. 4A and 4B. Furthermore, as will described herein, the augmentation of medical devices and instruments may enable the integration and / or unification all or a vast amount ofinformation collectively available from the existing medical devices and instruments into a single system (e.g., the computing system 120) for the operating room team.

[0056] An example augmented medical device, as shown in FIGS. 4A and 4B may include but is not limited to: a camera 308, which may comprise an overhanging lens on an adjustable but fixable arm 310, and an illuminator, WiFi (or other suitable networking functionality), a status LED light 312, a battery pack or other power management device 314, one or more or a swarm of data gathering modules 302, designed to capture data (e.g., from a human intended display 304); the medical device 306 or instrument being augmented; and a QR-code 316 or other image recognition indicia added to the medical device 306 being monitored to determine automatically identify or track feeds. In some embodiments, the medical device 306, the augmented medical device, or a component thereof may be easily retrievable for recharge. In some aspects, a recharge station may be built into the surgical cart 124. The augmented medical device may be sterilizable and wipeable as per medical sterilization requirements (e.g., in both a sterile field and not in a sterile field).

[0057] As shown in FIGS. 4A and 4B, a data gathering field module 302 may be deployed to capture output from “human-intended" interfaces of medical devices and instruments (e.g., humanintended display 304 of existing surgical procedure related device 306). The data gathering field module 302 may be a separate component affixed to the medical device 306. Also or alternatively, an augmented medical device may be designed to have the data gathering field module 302 embedded and / or integrated with the medical device 306. The data gathering field module 302 may comprise data gathering sensors (e g., a camera and / or a microphone).

[0058] In some aspects, the data gathering field module 302 may include but is not limited a power management component, an optional battery and / or a backup power component, and an optional data storage capability. In some aspects, the data gathering field module 302 may further include functionalities for communicating data (e.g., a network interface, a wired or wireless communication port, etc.) to the computing system 120. The computing system 120 (deep learning based system) may integrate the surgical data captured from the various disparate augmented medical devices 306, train and test deep learning models based on the surgical data, and apply and generate inferences using the deep learning models. The data gathering field module 302 may include various mechanisms for mounting and maintaining pose relative to its target device.

[0059] In some embodiments, the data gathering field module 302 may include a deep learning inferencing capability, for example, to offload computational processes and resources from themain deep learning based system in the computing system 120. In such embodiments, the data gathering field module 302 may be equipped with one or more processors and memory that are configured to process training datasets, and apply inferences from trained deep learning models.

[0060] In at least one embodiment, the data collection sensor associated with the data gathering module 302 may comprise a camera 308, for example, as shown in FIGS. 4A and 4B. The camera 308 may be aimed at the human-intended interface 304 for each device of interest (e.g., medical device 306), causing the camera to capture or record image and / or video output from the humanintended interface 304. For example, the camera 308 may record at 30 frames per second at a resolution of 1920x 1080 pixels in color. In some aspects, a microphone may be placed nearby to record audio output from the medical device 306 and / or its human-intended interface 304. Furthermore, the camera 308 may be positioned to identify the medical device 306 by capturing a QR code 316 placed on the medical device 306.

[0061] In some embodiments, at least one data gathering field module may be dedicated or assigned to each medical device of interest and mounted in a manner that does not interfere with the normal operation of the medical device (e.g., as shown in the way data gathering module 302 is mounted on medical device 306). The mounting may be temporary. Alternatively, the mounting may be semi-permanent or permanent. To ease the addition of the presence of one or more data gathering modules 302 to the medical device 302, the data collection sensors (e.g., camera 308, microphone, etc.) may be designed or configured to be unobtrusive so as not to interfere with the medical device or instrument’s original use and design, or the user’s normal behavior or interaction with or around the medical device or instrument 306.

[0062] In one embodiment, a nonobtrusive design or configuration for cameras may be achieved by designing the camera 308 as a special optical system mounted on a short arm 310 protruding slightly outward from an edge (e.g., the top) of an area of interest (e g., human-intended display 304) of the medical device 306. Despite the protrusion of the short arm 310, the camera 308 may be situated such that it can view the area of interest (e.g., the human-intended display 304) in its entirety including physical peaks and valleys. Associated image processing may be applied to the individual frames of the video stream captured or recorded by the camera 308 to correct any distortions introduced by the optical system and its positioning. Also or alternatively, the camera 308 may be fitted with long-range (i.e., zoom or high magnification) optics and placed remotely to the medical device or instrument that the camera is intended to monitor and / or capture datafrom. The camera’s view of the medical device or instrument can be maintained, for example, by permanently or semi-permanently fixing the camera and the medical device or instrument each in place at least for the course of the surgical procedure.

[0063] Various parameters controlling the creation of images of the video stream of each camera, as well as various parameters controlling the creation of a sound recording via the microphone may be set up manually, adjusted manually during the procedure, and / or set up automatically and adjusted automatically during the procedure. In some aspects, the various parameters controlling the creation of images of the video stream of each camera and / or the various parameters controlling the creation of a sound recording may adapt (e.g., automatically) to any changes to enable the respective sensor to capture the data of interest effectively.

[0064] The output data from the augmented medical device may be communicated to the surgical data integration and inferencing module of computing system 120 in one of various ways, depending on the setup and needs of the individual installation. In some embodiments the data is communication via wireless communications such as Wi-Fi or Bluetooth. In some embodiments the data is communicated via light-based communications such as Li-Fi. In some embodiments, particularly dedicated “ground-up” installations, the communications link is wired.

[0065] FIG. 5 shows a pictorial representation of an example method of processing data (e.g., image, video, and / or audio data) captured from human-intended interfaces on medical devices (e.g., human-intended interface 304 of medical device 306 of FIG. 4B), in accordance with a nonlimiting embodiment of the present disclosure. The representative image, video and / or audio data captured from the human-intended interfaces by data collection sensors (e.g., one or more cameras and microphones dedicated to data capture for an augmented medical device) may then be used to train one or more deep learning models.

[0066] Although the method shown in FIG. 5 pertains to image and video data captured by a camera, it is contemplated that other forms of data (e.g., audio, temperature, humidity, light, etc.) output by medical devices may be similarly captured by other data collection sensors (e.g., microphones, temperature sensors, humidity sensors, light sensors, etc.) and processed in a similar manner as exemplified by method 400. As shown in FIG. 5, an example method 400 may begin with the camera 308 associated with the data gathering field module 302 capturing raw image and / or video data (“raw view”) of the human intended interface 304 of the medical device 306 (block 402). As shown in the raw view 403, the image of the human-intended interface 304 maybe reversed and may include the QR code 316 that helps to identify, and / or track data output by, the medical device 306. At block 404, the data gathering field module 302 and / or the computing system 120 may unwarp the raw image and / or video data. For example, unwarped image and / or video data may show the image or video in the correct orientation, dimension, and / or size. Also or alternatively, the unwarped image and / or video data enable the capture of relevant quantifiable or vectorizable information (e.g., measurements, values, etc.). For example, each frame of the training set of the unwarped image and / or video data may be processed by a labeling function or module associated with the computing system 120. In some embodiments, the labeling module may be configured to operate on an inferencing model to make suggestions to a user, such that the user can affirm the suggestion or make corrections to it. Referring now to FIG. 6, although labeling may differ for each target device, there may be a series of common tasks including partitioning the target device human intended interface into regions such as Region 602, Region 604, Region 606, Region 608, Region 610, and border 612; and attaching labels to each region such as Region 602 corresponding to a graph of blood pressure, and Region 604 correspond to Pulse, and so on. The labeling module may be further configured to indicate whether a region may be used for inferencing or for other subsequent processing or should be used as is. When a region is marked for inferencing, the labeling module may indicate what type of existing and / or custom inferencing models are to be used on that region’s information to extract inferred information. For example, the labeling module may determine that a text recognition model is to be applied to Region 604 to convert the image to a numeric value, in this case “89.”

[0067] Furthermore, the raw image or video data may be unwarped by or unwarped from an image processing module that may be associated with or a part of the data gathering field module 302 and / or the computing system 120. For example, one or more camera parameter adjustment and image processing techniques may be used to ensure an inferencing model -friendly image including but not limited to auto-focus, auto-exposure, contrast and brightness adjustment, automatic gain adjustment, histogram stretching, histogram equalization, and high dynamic range imaging. According to one implementation, inferencing models may be configured to identify differences in image quality especially during the training phase where such parameter variation may be simulated or achieved in actual training data. In yet another embodiment, a portion of the target devices may be configured to use light-emitting displays, such that no light source may be used by the corresponding sensors (e g., observing cameras). In cases where the human-intended interfaceis not self-lighted, a light source may be used to illuminate the display of the device such that the captured data contains sufficient information for the inferencing model to detect and analyze.

[0068] In one aspect, during the training phase of the deep learning models of the present disclosure (e.g., at least one of 24a, 24b, 24c, ... 24n of FIG. 1), data may be stored in various persistent memory devices and locations of the system 10 of FIG. 1. Example options may include but are not limited to: on an embedded processing unit (EPU) mass storage device; on a removable hard drive; on a networked storage device deployed within the network (e.g., communication network 20 of FIG. 1) or on an external network (variously known as “the Cloud”). Such training data is then typically used to train the model on offline computers separate from the runtime EPU.

[0069] During the inference phase, the captured data may be stored and moved temporarily in various random access memory (RAM) present in the EPU such as main computer memory and GPU (general processing unit also known as graphics processing unit) memory for a selected period of time and may be accessed by all inferencing algorithm steps to allow inferencing to proceed on said data.

[0070] This temporary data storage may be implemented in a ring buffer of finite size some multiple (e.g., N = 10) of each individual data unit size which adds incoming samples to the next available slot and when filled with arriving data returns to the beginning address of the buffer and repeats the cycle, overwriting the oldest unit or data, then next-oldest and so on.

[0071] The size of the buffer may be determined to allow the processing of incoming data and providing sufficient time for processing steps to make local copies. If local copies are not made, then the buffer is made big enough to allow enough processing time for all inferencing steps to finish. After all inferencing steps are completed, local copies, if used, may be destroyed. The data stored in the ring buffer may be eventually destroyed by overwriting.

[0072] Additionally, snapshots of said data may be stored optionally during the inferencing phase in a separate persistent memory location of the type listed above, for the purpose of later retrieval for use as further training data in updating the inferencing model. Such data may be retrieved by other computing devices deployed within the system 10 of FIG. 1.

[0073] In some embodiments, at block 402, the data gathering field module 302 may also capture (e.g., via camera 308) settings adjustments made by operators of the medical device 306, for example, desired units used for measurements, desired thresholds pertaining to functions performed by the medical device 306 that the medical device 306 may abide by, and / or statuses(e g., active, idle, etc.) and configurations for the medical device 306 set by the operator. Moreover, data from even the oldest instrument associated with the procedure either in the operating room or outside of it can optionally be captured, interpreted, and used in the training, testing, application, and inferences created by deep learning models using computing system 120. Further, raw data captured by data collection sensors (e.g., camera 308) can be converted into a form readily usable by the computing system 120 (e.g., at block 404).

[0074] In some embodiments, the digital surgical microscope system may display the unwarped image or video data (e.g., via one or both of monitors 132A or 132B). The display of received data from the disparate medical devices into a presentable form for the surgeon (e.g., on monitors 132A-132B) can thus allow the surgeon to seamlessly view information from the various medical devices in one general area (a “surgical cockpit”) (e.g., monitors 132A-132B) (block 406). The surgical cockpit is thus a substantial auxiliary benefit provided to surgeons, based on the above described systems and methods for receiving, processing, and displaying, in one general area that is easily accessible to the surgeon, the data output from numerous and disparate operating room medical devices in real time during the surgery.

[0075] In some embodiments, the augmented medical device may be configured to apply deep learning models to the processed image and / or video data (e.g., via a deep learning module in the augmented medical device) (block 408). The application of the deep learning model by the augmented medical device may generate an output indicating what the data pertaining to the augmented medical device represents or recommends. This output, which pertains to the locally processed image and / or video data captured by the medical device, may be referred to as a micro output to distinguish from a macro output generated by the computing system 120. In one aspect, a micro output may convert data captured from a target medical device into a digitized form that is human readable as well as machine readable, but may not connect information among devices other than displaying and storing it possibly adjacent to each other. Macro outputs may be configured to connect information from multiple devices such that for example, the cause or solution of an emergent situation in a surgery can be connected to a change in some biological parameter of the patient such as heart rate, or to some action of someone in the surgical scene. As shown in FIG. 5, an example micro output 409 may comprise a digitized presentation of surgical information represented by the human intended interface 403 of the augmented medical device. According to some embodiments, referring to FIG. 7, digitized medical data 702 (e.g., microoutput 409) may be superimposed on a monoscopical or stereoscopical real time view or recording 704 of a surgical site captured by e.g., the stereoscopic camera of the DSM. For example, the computing system 120 may include superimposing circuitry (not shown) configured to determine a location for overlying the digitized medical data 702 onto the monoscopical or stereoscopical view without interfering with the surgeon’s view of the surgical site. Further, the superimposing circuitry may be configured to enable and disable the overlay function in accordance with a user’s command.

[0076] Further, as shown in block 410, a macro output may be generated by the computing device 120 by applying a deep learning model to image, video, and / or audio data received from a plurality of disparate medical devices. By applying deep learning models locally at the augmented medical device, the augmented medical device may be configured to offload computing resources from the main computing device 120. The augmented medical device may then transmit the processed image and / or video data or an outcome of the applied deep learning model to the computing system 120. At block 410, the computing system 120 may receive the unwarped image and / or video data from various augmented medical devices (if not already received) and may apply one or more deep learning models based on the image and / or video data for a desired macro output. The various augmented medical devices may include the augmented medical device on which processes described in blocks 402 through 408 are performed. Other augmented medical devices that transmit the processed data to computing system 120 may perform similar processes as described in blocks 402 through 408.

[0077] In some aspects, prior to applying the one or more deep learning models, the one or more deep learning models may be trained and tested using reference training data. For example, the training data may comprise a plurality of reference input image, video, and / or audio data associated with (e.g., labeled with) a plurality of reference outcome outputs (e.g., an indication of what the composite input data represents or a surgical decision recommended from the composite input data). The training may comprise quantifying, vectorizing and / or digitizing the received data (e.g., to form feature vectors for the training of a neural network). Furthermore, the training may involve receiving output data and / or labeling input data with markers or metadata to indicate an output for supervised learning.

[0078] At block 410, the training may thus comprise performing one or more iterations (e.g., feedforward and backpropagation) to cause the deep learning-based system to determine or “learn”the parameters, weights, and / or biases associating input data to the output data (desired macro output). In another embodiment, an extended specialty-focused large language model (LLM) in medicine may be used for generating macro outputs. Such an LLM may be based on reliable peer- reviewed surgery specific data sources. Further, retrieval-augmented generation (RAG) techniques may be used by the model to mitigate risks and ensure the highest standards of output. In one implementation, the system 10 of FIG. 1 may be configured to combine the generative strength of LLMs with an external data retrieval mechanism. For example, upon receiving any prompt, the computing server 22 may be configured to initiate a process to source and retrieve pertinent data and sources (e.g., at least one of 24a, 24b, 24c, ... 24n), predetermined during the training phase of model development. After retrieving relevant articles, passages, or partitions from a large, predefined language corpus, the computing server 22 may integrate the information with the initial prompt to provide a sufficient context for delivering a truthful and accurate answer. In addition, the complexity and tone of the response may be adjusted according to the level of medical knowledge of the user or target audience 12. The implementation architecture used by the system 10 may be scaled and integrated easily via REST API. Automated ingestion pipelines may also in place to ensure that the system 10 remains up to date and relevant over time.

[0079] It is contemplated that in various embodiments, patient consent is obtained prior to using patient-sensitive data in the training of the deep learning models. Furthermore, the handling of patient-identifiable data may be left to the policies of the institution employing the use of the surgical data gathering, data display and inferencing system described herein.

[0080] In various embodiments, the captured image, video, and / or audio data (e.g., by the “field” cameras 308 of the augmented medical devices) may not necessarily be re-used except, for example, in the case of a crisis where injury or death occurs to a patient, or as input data for continued deep learning model training. Such decisions concerning the re-use of captured image, video, and / or audio data may be left to the policies of the institution employing the use of the surgical data gathering, data display and inferencing system described herein.

[0081] Furthermore, it is contemplated that any deep learning model output generated by the surgical data gathering, data display and inferencing system may be suggestive only, allowing the surgeon to be a final arbiter for surgical decisions.

[0082] Unless specifically stated otherwise as apparent from the foregoing disclosure, it is appreciated that, throughout the present disclosure, discussions using terms such as “processing,”“computing,” “calculating,” “determining,” “displaying,” or the like, refer to the action and processes of a computer system, or similar electronic computing device, that manipulates and transforms data represented as physical (electronic) quantities within the computer system's registers and memories into other data similarly represented as physical quantities within the computer system memories or registers or other such information storage, transmission or display devices.

[0083] One or more components may be referred to herein as "configured to," "configurable to," "operable / operative to," "adapted / adaptable," "able to," "conformable / conformed to," etc. Those skilled in the art will recognize that "configured to" can generally encompass active-state components and / or inactive-state components and / or standby-state components, unless context requires otherwise.

[0084] Those skilled in the art will recognize that, in general, terms used herein, and especially in the appended claims (e.g., bodies of the appended claims) are generally intended as "open" terms (e.g., the term "including" should be interpreted as "including but not limited to," the term "having" should be interpreted as "having at least," the term "includes" should be interpreted as "includes but is not limited to," etc.). It will be further understood by those within the art that if a specific number of an introduced claim recitation is intended, such an intent will be explicitly recited in the claim, and in the absence of such recitation no such intent is present. For example, as an aid to understanding, the following appended claims may contain usage of the introductory phrases "at least one" and "one or more" to introduce claim recitations. However, the use of such phrases should not be construed to imply that the introduction of a claim recitation by the indefinite articles "a" or "an" limits any particular claim containing such introduced claim recitation to claims containing only one such recitation, even when the same claim includes the introductory phrases "one or more" or "at least one" and indefinite articles such as "a" or "an" (e.g., "a" and / or "an" should typically be interpreted to mean "at least one" or "one or more"); the same holds true for the use of definite articles used to introduce claim recitations.

[0085] In addition, even if a specific number of an introduced claim recitation is explicitly recited, those skilled in the art will recognize that such recitation should typically be interpreted to mean at least the recited number (e.g., the bare recitation of "two recitations," without other modifiers, typically means at least two recitations, or two or more recitations). Furthermore, in thoseinstances where a convention analogous to "at least one of A, B, and C, etc." is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (c.g., "a system having at least one of A, B, and C" would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). In those instances where a convention analogous to "at least one of A, B, or C, etc. " is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g, "a system having at least one of A, B, or C" would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, c / c. ). It will be further understood by those within the art that typically a disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms unless context dictates otherwise. For example, the phrase "A or B" will be typically understood to include the possibilities of "A" or "B" or "A and B."

[0086] With respect to the appended claims, those skilled in the art will appreciate that recited operations therein may generally be performed in any order. Also, although various operational flow diagrams are presented in a sequence(s), it should be understood that the various operations may be performed in other orders than those which are illustrated, or may be performed concurrently. Examples of such alternate orderings may include overlapping, interleaved, interrupted, reordered, incremental, preparatory, supplemental, simultaneous, reverse, or other variant orderings, unless context dictates otherwise. Furthermore, terms like "responsive to," "related to," or other past-tense adjectives are generally not intended to exclude such variants, unless context dictates otherwise.

[0087] It is worthy to note that any reference to "one aspect," "an aspect," "an exemplification," "one exemplification," and the like means that a particular feature, structure, or characteristic described in connection with the aspect is included in at least one aspect. Thus, appearances of the phrases "in one aspect," "in an aspect," "in an exemplification," and "in one exemplification" in various places throughout the specification are not necessarily all referring to the same aspect. Furthermore, the particular features, structures or characteristics may be combined in any suitable manner in one or more aspects.

[0088] As used herein, the singular form of "a", "an", and "the" include the plural references unless the context clearly dictates otherwise.

[0089] As used herein, the term "comprising" is not intended to be limiting, but may be a transitional term synonymous with "including," "containing," or "characterized by." The term "comprising" may thereby be inclusive or open-ended and does not exclude additional, un-recited elements or method steps when used in a claim. For instance, in describing a method, "comprising" indicates that the claim is open-ended and allows for additional steps. In describing a device, "comprising" may mean that a named element(s) may be essential for an embodiment or aspect, but other elements may be added and still form a construct within the scope of a claim. In contrast, the transitional phrase "consisting of' excludes any element, step, or ingredient not specified in a claim. This is consistent with the use of the term throughout the specification.

[0090] Any patent application, patent, non-patent publication, or other disclosure material referred to in this specification and / or listed in any Application Data Sheet is incorporated by reference herein, to the extent that the incorporated materials is not inconsistent herewith. As such, and to the extent necessary, the disclosure as explicitly set forth herein supersedes any conflicting material incorporated herein by reference. Any material, or portion thereof, that is said to be incorporated by reference herein, but which conflicts with existing definitions, statements, or other disclosure material set forth herein will only be incorporated to the extent that no conflict arises between that incorporated material and the existing disclosure material. None is admitted to be prior art.

[0091] In summary, numerous benefits have been described which result from employing the concepts described herein. The foregoing description of the one or more forms has been presented for purposes of illustration and description. It is not intended to be exhaustive or limiting to the precise form disclosed. Modifications or variations are possible in light of the above teachings. The one or more forms were chosen and described in order to illustrate principles and practical application to thereby enable one of ordinary skill in the art to utilize the various forms and with various modifications as are suited to the particular use contemplated. It is intended that the claims submitted herewith define the overall scope.- 21 -

Claims

What Is Claimed Is:

1. A system for integrating, displaying, and generating inferences from surgical information from disparate medical devices, the system comprising: a plurality of medical devices configured to measure patient-specific data from a patient, each medical device comprising an interface for outputting image, video, and / or audio data; a plurality of sensors comprising at least one sensor for each of the plurality of medical devices, the at least one sensor configured to capture raw image, video, and / or audio data from the interface of each medical device; a plurality of data gathering modules corresponding to the plurality of medical devices, each data gathering module configured to: generate, from the raw image, video, and / or audio data from each medical device, processed image, video, and / or audio data from each medical device; and transmit the processed image, video, and / or audio data from each medical device to a computing device; wherein the computing device comprises: a processor; and memory storing computer-executable instructions that, when executed by the processor, causes the system to: receive, in real-time, the processed image, video, and / or audio data from each of the plurality of medical devices; and apply a plurality of deep learning models to the processed image, video, and / or audio data to generate an output outcome.

2. The system of claim 1, wherein the plurality of medical devices includes a digital surgical microscope system, wherein the digital surgical microscope system comprises at least one camera configured to capture raw image, video, and / or audio data from a surgical site of the patient, wherein a data gathering module of the digital surgical microscope system is configured to: generate, from the raw image, video, and / or audio data, processed image, video, and / or audio data; andtransmit the processed image, video, and / or audio data from the respective medical device to the computing device.

3. The system of claim 1, further comprising: one or more monitors accessible to a surgeon; wherein the one or more monitors are configured to display the processed image, video, and / or audio data from one or more of the plurality of medical devices.

4. The system of claim 1, wherein the computer-executable instructions that, when executed by the processor, further causes the system to, prior to applying the plurality of deep learning models: train the plurality of deep learning models using a training dataset comprising a plurality of reference input image, video, and / or audio data and reference output outcomes.

5. The system of claim 1, wherein the at least one sensor comprises one or more of a camera or a microphone.

6. The system of claim 1, wherein at least one of the plurality of medical devices further comprises a QR code located within a field of view of a respective sensor of each medical device, wherein the respective sensor comprises a camera.

7. The system of claim 1, wherein the computer-executable instructions that, when executed by the processor, further causes the system to apply each of the plurality of deep learning models to process data received from each of the plurality of data gathering modules corresponding to each of the plurality of medical devices.

8. The system of claim 2, wherein the data gathering module of the digital surgical microscope system is configured to convert the raw image, video, and / or audio data into digitized medical data that is at least machine readable, wherein the output outcome comprises the digitized medical data overlaid on at least the raw video data captured by the at least one camera, wherein the output outcome is presented monoscopically or stereoscopically.

9. The system of claim 8, wherein the computer-executable instructions that, when executed by the processor, further causes the system to render the digitized medical data of the output outcome monoscopically or stereoscopically in correspondence with the raw image, video, and / or audio data.

10. The system of claim 2, wherein the computer-executable instructions that, when executed by the processor, further causes the system to apply the plurality of deep learning models to provide recommendations on a display screen of the digital surgical microscope system based on a known surgical procedure, wherein the recommendations are overlaid on the raw image, video, and / or audio data captured in real time during a surgical procedure.

11. The system of claim 10, wherein the recommendations include an alarm signal enabled and displayed on the display screen of the digital surgical microscope system in response to detecting an exposure of a patient’s tissue based on the raw image, video, and / or audio data.

12. The system of claim 10, wherein the computer-executable instructions that, when executed by the processor, further causes the system to generate natural sounding audio signals to represent the recommendations.

13. An apparatus for integrating, displaying, and generating inferences from surgical information from a digital surgical microscope (DSM) system, the apparatus comprising: the DSM system comprising at least one camera and a data gathering module, wherein the at least one camera is configured to capture raw image, video, and / or audio data from a surgical site of the patient, wherein the data gathering module of the DSM system is configured to: generate, from the raw image, video, and / or audio data, processed image, video, and / or audio data; and transmit the processed image, video, and / or audio data to a computing device; wherein the computing device comprises: a processor; andmemory storing computer-executable instructions that, when executed by the processor, causes the computing device to: receive, in real-time, processed image, video, and / or audio data from each of a plurality of medical devices; and apply a deep learning model to the processed image, video, and / or audio data to generate an output outcome.

14. The apparatus of claim 13, wherein the computer-executable instructions that, when executed by the processor, further causes the computing device to, prior to applying the deep learning model: train the deep learning model using a training dataset comprising a plurality of reference input image, video, and / or audio data and reference output outcomes.

15. The apparatus of claim 13, further comprising: one or more additional cameras associated with the DSM system; wherein the one or more additional cameras are configured to capture additional raw image, video, and / or audio data from a field of view that is greater than the surgical site of the patient; wherein the processed image, video, and / or audio data from the DSM system is further generated from additional image, video, and / or audio data captured by the one or more additional cameras.

16. The apparatus of claim 13, wherein the data gathering module is further configured to covert the raw image, video, and / or audio data into digitized medical data that is at least machine readable, wherein the output outcome comprises the digitized medical data overlaid on at least the raw video data captured by the at least one camera, wherein the output outcome is presented monoscopically or stereoscopically.

17. The apparatus of claim 14, wherein the computer-executable instructions that, when executed by the processor, further causes the computing device to render the digitized medicaldata of the output outcome monoscopically or stereoscopically in correspondence with the raw image, video, and / or audio data.

18. The apparatus of claim 14, wherein the computer-executable instructions that, when executed by the processor, further causes the computing device to create and store a record including at least the digitized medical data and the raw video data captured by the at least one camera.

19. The apparatus of claim 13, wherein the computer-executable instructions that, when executed by the processor, further causes the computing device to apply the deep learning model to provide recommendations on a display screen of the DSM system based on a known surgical procedure, wherein the recommendations are overlaid on the raw image, video, and / or audio data captured in real time during a surgical procedure.

20. The apparatus of claim 19, wherein the recommendations include an alarm signal enabled and displayed on the display screen of the digital surgical microscope system in response to detecting an exposure of a patient’s tissue based on the raw image, video, and / or audio data, wherein the computer-executable instructions that, when executed by the processor, further causes the computing device to generate natural sounding audio signals to represent the recommendations.