Potential spatial training of operating room data

By training a machine learning system to identify semantic relationships in operating room data and employing latent space representation and masking strategies, the problem of complex labeling of operating room data was solved, enabling effective training and downstream applications under low-labeling conditions and improving the analytical capabilities of operating room data.

CN121986383APending Publication Date: 2026-05-05INTUITIVE SURGICAL OPERATIONS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480064734.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-08-28
Filing Date
2024-08-26
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

The multimodal characteristics and unique conditions of operating room data make manual labeling and annotation impractical. Machine learning systems are complex to train, lack resources for additional training, and rely on a large number of human annotators, making it difficult to achieve effective downstream applications.

Method used

By training a machine learning system to identify semantic relationships within operating room data, and employing latent spatial representation and masking strategies, combined with loss determination methods and multimodal data grouping, operating room data is compressed and reconstructed, reducing reliance on labeled data.

Benefits of technology

Under low-label data conditions, efficient training and downstream applications of machine learning systems were achieved, improving the availability and analytical capabilities of operating room data and reducing reliance on human annotators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121986383A_ABST
    Figure CN121986383A_ABST
Patent Text Reader

Abstract

Various disclosed embodiments provide systems and methods for training a machine learning system to anticipate semantic features of operating room data even in the absence of fully labeled data. Specifically, embodiments may compress and then reconstruct available operating room data, even if the data is not tagged, and, in doing so, instruct a coding and decoding machine learning system to identify semantic features that are significant for downstream training and applications. Further, many disclosed embodiments accommodate a wide variety of forms of operating room data, such as pairs of visual intensity video operating room data and depth frame video operating room data. Masking, loss, and other public features may further improve the ability of the system to infer semantic relevance. Using various disclosed embodiments, machine learning applications can be made possible, which would otherwise be infeasible in low data operating room data conditions.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications This application claims the benefit and priority of U.S. Provisional Application No. 63 / 535,060, filed August 28, 2023, entitled “LATENTSPACE TRAINING FOR SURGICAL THEATER DATA,” which is incorporated herein by reference in its entirety for all purposes. Technical Field

[0002] Various disclosed embodiments relate to systems and methods for improving the training of machine learning systems based on operating room data. Background Technology

[0003] Recent advances in machine learning, particularly deep learning, have brought immense promise for a wide range of operating room applications. Combined with data acquisition from sensors within the operating room, such applications can identify or predict adverse configurations, identify or predict adverse patient conditions, provide guidance for improving team workflows, identify operating room conditions in preferred taxonomy of state representations, provide comparative analysis with other hospitals and surgical teams, guide teams transitioning from non-robotic to robotic operating rooms, and vice versa. These applications not only improve patient outcomes but also enhance surgical team efficiency, help reduce costs, and make healthcare more predictable, consistent, and cost-effective.

[0004] Unfortunately, the unique conditions of the operating room and the abundance of data collected often make manual labeling and annotation impractical. This lack of labeled data, in turn, complicates the training of machine learning systems. For example, if a machine learning system is expected to identify one of multiple surgical tasks from operating room data (such as surgical videos), then trained expert annotators familiar with the details of the surgery and tasks are obligated to manually examine the operating room data and manually annotate each time segment to correspond to one of the multiple tasks. Naturally, such human-in-the-loop annotation risks subjective labeling variations, is limited by annotator fatigue, and is constrained by the number of professionally trained annotators available to review and label such data.

[0005] In reality, the situation becomes even more complex due to the multimodal nature of much operating room data. It's difficult enough to ask professional human annotators to simply segment surgical video data into discrete taxonomic states, because videos, unlike still images, possess dense time-varying information. Furthermore, expecting such professional annotators to annotate the diverse auditory, kinematic, and depth data acquired concurrently with video annotation in the operating room, while simultaneously identifying the unique features of those different data modalities, is likely simply impossible. This is particularly unfortunate because such information-dense data types can be especially useful for machine learning.

[0006] Even when annotators manage to complete the arduous task of annotating a sufficient amount of operating room data for training machine learning systems, many of the aforementioned applications require or greatly benefit from the ability to perform subsequent online training of the machine learning system as additional operating room data becomes available. Such new data can be specific to the particular conditions of the healthcare environment in which the machine learning system is now deployed (e.g., a specific hospital or operating room where the system has been deployed). Therefore, training on this new data can greatly facilitate the localization of the deployed machine learning system to specific characteristics of its local environment. Unfortunately, few hospitals have the resources to re-annotate this newly acquired data for additional training of the machine learning system.

[0007] Therefore, systems and methods are needed to overcome challenges and difficulties such as those mentioned above. For example, systems are needed that facilitate a variety of downstream operating room analytical applications, without always requiring the heavy involvement of a large number of human annotators. Attached Figure Description

[0008] The various embodiments described herein can be better understood by referring to the following detailed description and the accompanying drawings, in which similar reference numerals indicate the same or similarly functional elements: Figure 1A This is a schematic diagram of various elements that may occur in the operating room during a surgical procedure, as can happen in some embodiments; Figure 1B This is a schematic diagram of various elements that may appear in the operating room during a surgical procedure using a robotic surgical system, which may occur in some embodiments; Figure 2A A schematic depth map rendered from the perspective of an example operating room global sensor, which may be used in some embodiments; Figure 2B For use in some embodiments, with corresponding sensor locations Figure 2A A schematic top view of objects in the operating room; Figure 2CTo depict a pair of images capturing a grid pattern of orthogonal rows and columns at a viewing angle from an operating room global vision image sensor with a straight-line view and an operating room global vision image sensor with a fisheye view, each image may be used in conjunction with some embodiments; Figure 3 This is a schematic depiction of data acquired during the surgical procedure, which may occur in conjunction with some embodiments; Figure 4 This is a schematic depiction of data acquired during the non-operational period of surgery, which may occur in conjunction with some embodiments; Figure 5A This is a schematic representation of multiple operating room status identifications that can be determined in some embodiments; Figure 5B This is a schematic representation of the segmentation results of multiple objects that may be determined in some embodiments; Figure 6A This is a schematic representation of an example set of time intervals that can be used in some embodiments to evaluate activities in the operating room; Figure 6B A schematic block diagram illustrating example activity analysis category groupings that can be used in conjunction with some embodiments; Figure 7A A schematic block diagram illustrating components in an example training or inference data item that may be used in some embodiments; Figure 7B A schematic block diagram illustrating components in labeled, partially labeled, and unlabeled datasets that can be used in various embodiments; Figure 8A A flowchart illustrating various operations in an example process for performing supervised training of a machine learning system, which may be performed in some embodiments; Figure 8B A flowchart illustrating various operations in an example process for performing unsupervised training and then task-specific supervised training of a machine learning system, which may be performed in some embodiments; Figure 8C This is a schematic block diagram illustrating the components in an example deployed classifier application machine learning system, supervised training process, unsupervised-to-supervised architecture adaptation, and unsupervised training system that can be implemented in combination with various embodiments; Figure 9 A schematic block diagram illustrating example loss determination iterations during unsupervised training of the reconstruction system, which may be implemented in some embodiments; Figure 10A To illustrate a schematic tensor decomposition of component masking values ​​according to a first masking strategy that may be implemented in some embodiments; Figure 10B To illustrate a schematic tensor decomposition of component masking values ​​according to a second masking strategy that may be implemented in some embodiments; Figure 10C To illustrate a schematic tensor decomposition of component masking values ​​according to a third masking strategy that may be implemented in some embodiments; Figure 10D To illustrate a schematic tensor decomposition of component masking values ​​according to a fourth masking strategy that may be implemented in some embodiments; Figure 10E This is a pair of schematic intensity and depth tensors and their corresponding masking representations that can be implemented in some embodiments; Figure 11 A schematic block diagram illustrating the data flow during loss calculation in multimodal deep and intensive training of a machine learning system, which may be implemented in some embodiments; Figure 12 This is a schematic block diagram illustrating the data flow during multimodal training rounds, which may be implemented in some embodiments; Figure 13 A flowchart illustrating various operations in an example process for performing a round of loss determination during training, which may be implemented in some embodiments; Figure 14 A flowchart illustrating various operations performed in some embodiments for preparing an example process for applying a specific machine learning system; Figure 15 A table depicting comparative results generated from example prototype implementations in conjunction with the embodiments; and Figure 16 This is a block diagram of an example computer system that can be used in conjunction with some embodiments.

[0009] Specific examples depicted in the accompanying drawings have been selected for ease of understanding. Therefore, the disclosed embodiments should not be limited to the specific details in the drawings or the corresponding disclosure. For example, the drawings may not be drawn to scale, the dimensions of some elements may have been adjusted for ease of understanding, and the operations of the embodiments associated with the flowcharts may include additional, alternative, or fewer operations compared to those depicted herein. Thus, some components and / or operations may be divided into different blocks or combined into a single block in a manner different from that depicted. The embodiments are intended to cover all modifications, equivalents, and substitutions falling within the scope of the disclosed examples, and not to limit the embodiments to the specific examples described or depicted. Detailed Implementation

[0010] Implementation Examples Overview Various disclosed embodiments seek to train machine learning systems to identify semantic relationships within operating room data (e.g., relationships within a single data modality or across two or more modalities). Machine learning systems trained to identify these semantic relationships may be more readily trained for specific applications with relatively little or no labeled data. More specifically, embodiments may compress and then reconstruct available operating room data to learn a latent spatial representation of the operating room data that captures significant semantic relationships. By further employing certain masking strategies, comparative loss determination methods, and multimodal operating room data groupings disclosed herein, embodiments can significantly improve downstream training for a variety of applications. Thus, by using various disclosed embodiments, it is possible to create operating room machine learning applications that were previously infeasible in the domain of low-labeled operating room data.

[0011] Example Operating Room Overview Figure 1A This is a schematic diagram of various elements present in operating room 100a during surgical procedures, which may occur in some embodiments. In particular, Figure 1A A non-robotic operating room 100a is depicted, in which a patient-side surgeon 105a operates on a patient 120 with the assistance of one or more assistants 105b, who may themselves be a surgeon, physician assistant, nurse, technician, etc. The surgeon 105a may use a variety of tools for operation, including visualization tools 110b (such as laparoscopic ultrasound, visual image acquisition endoscopes, etc.) and mechanical instruments 110a (such as scissors, retractors, dissecting instruments, etc.).

[0012] The visualization tool 110b provides the surgeon 105a with an internal view of the patient 120, for example, by displaying visualization output from an imaging device mechanically and electrically connected to the visualization tool 110b. The surgeon can view the visualization output, for example, through an eyepiece connected to the visualization tool 110b or on a display 125 configured to receive the visualization output. For example, in the case where the visualization tool 110b acquires an endoscope for visual images, the visualization output may be a color or grayscale image. The display 125 can allow the assistant 105b to monitor the progress of the surgeon 105a during the procedure. The visualization output from the visualization tool 110b can be recorded and stored for future review, for example, by capturing the visualization output in parallel with the visualization output provided to the display 125 using hardware or software on the visualization tool 110b itself, or by capturing the output of the display 125 as soon as the output appears on the screen, etc. While this document can broadly discuss two-dimensional video capture using visualization tool 110b, such as when visualization tool 110b is a visual image endoscope, it will be understood that in some embodiments, visualization tool 110b can capture depth data instead of two-dimensional image data, or capture depth data in addition to two-dimensional image data (e.g., using a laser rangefinder, stereoscope, etc.).

[0013] A single surgical procedure may include performing multiple sets of actions, each set forming a discrete unit, referred to herein as a task. For example, locating a tumor may constitute a first task, resecting the tumor a second task, and closing the surgical site a third task. Each task may include multiple actions; for example, a tumor resection task may require several cutting actions and several cauterizing actions. While some surgeries require tasks to be performed in a specific order (e.g., resection before closure), in some surgeries, the order and presence of some tasks may be allowed to change (e.g., canceling preventative tasks or reordering resection tasks if the order is invalid). Transitions between tasks may require surgeon 105a to remove a tool from the patient, replace it with a different tool, or introduce a new tool. Some tasks may require visualization tools 110b to be removed and repositioned relative to their position in a previous task. While some assistants 105b may assist with surgery-related tasks, such as administering anesthesia 115 to patient 120, assistants 105b may also assist with these task transitions, such as predicting the need for new tools 110c.

[0014] Technological advancements have made it possible to use robotic systems to perform tasks such as Figure 1A The described program can execute programs that cannot be executed in a non-robotic operating room 100a. Specifically, Figure 1B In some embodiments, this can occur when using a robotic surgical system (such as da Vinci). TMThis is a schematic diagram of various elements present in operating room 100b during surgical procedures using a surgical system. Here, a patient-side trolley 130, with tools 140a, 140b, 140c, and 140d respectively attached to each of multiple arms 135a, 135b, 135c, and 135d, may occupy the position of the patient-side surgeon 105a. As previously mentioned, one or more of tools 140a, 140b, 140c, and 140d may include visualization tools (here, visualization tool 140d), such as a visual imaging endoscope, laparoscopic ultrasound, etc. The operator 105c, who may be the surgeon, can view the output of the visualization tool 140d via a display 160a on a surgeon's console 155. By manipulating the handheld input mechanism 160b and the pedal 160c, the operator 105c can remotely communicate with the tools 140a-140d on the patient-side trolley 130 to perform surgical procedures on the patient 120. In practice, because communication between the surgeon's console 155 and the patient-side trolley 130 can be via a communication network in some embodiments, the operator 105c may or may not be in the same physical location as the patient-side trolley 130 and the patient 120. The electronics / console 145 may also include a display 150 that depicts the patient's vital signs and / or the output of the visualization tool 140d.

[0015] Similar to the task transitions in non-robotic operating room 100a, surgical procedures in operating room 100b may require the removal or replacement of tools 140a-140d for various tasks, including visualization tools 140d, and the introduction of new tools, such as new tool 165. As previously mentioned, one or more assistants 105d can now anticipate such changes and work with operator 105c to make any necessary adjustments as the surgery progresses.

[0016] Similar to the non-robotic operating room 100a, the output from the visualization tool 140d can be recorded at locations such as the patient-side trolley 130, surgeon's console 155, and monitor 150. In the non-robotic operating room 100a, some tools 110a, 110b, and 110c can record additional data, such as temperature, motion, conductivity, and energy levels. The presence of the surgeon's console 155 and patient-side trolley 130 in operating room 100b facilitates the recording of far more data than just the output from the visualization tool 140d. For example, operator manipulation of the handheld input mechanism 160b, activation of the pedal 160c, and eye movement relative to the monitor 160a can all be recorded. Similarly, the patient-side trolley 130 can record tool activation (e.g., application of radiant energy, closure of scissors), instrument movement, etc., throughout the surgical procedure. In some embodiments, an operating room recording device may be used to record data, which may capture and store sensor data (e.g., software, firmware, or hardware configured to record surgeon kinematics data, console kinematics data, instrument kinematics data, system event data, patient status data, etc. during surgery) locally or at a network location.

[0017] Computer systems 190a and 190b can communicate with the operating rooms via a network, either within each of the operating rooms 100a and 100b respectively, or from an external location (in some embodiments, computer system 190b can be integrated with a robotic surgical system rather than used as a standalone workstation). As will be discussed in more detail herein, computer systems 190a and 190b can facilitate, for example, data collection and data processing.

[0018] Similarly, many operating rooms 100a, 100b may include sensors placed around the operating room, such as sensors 170a and 170c, which are configured to record activities within the operating room from their respective fields of view 170b and 170d. Sensors 170a and 170c may be, for example, visual image sensors (e.g., color or grayscale image sensors), depth acquisition sensors (e.g., visual image pairs acquired via a real stereoscope, via time-of-flight with a laser rangefinder, structured light, etc.), or a combination of visual image and depth acquisition sensors (e.g., RGB-D sensors for red, green, and blue depth). In some embodiments, sensors 170a and 170c may also include audio acquisition sensors, or sensors specifically designed for audio acquisition may be placed around the operating room. Multiple such sensors may be placed within operating rooms 100a, 100b, potentially with overlapping fields of view and sensing ranges, to enable a more comprehensive assessment of the operating room. For example, depth acquisition sensors may be strategically placed around the operating room such that the depth frames they generate at each moment can be incorporated into a single three-dimensional virtual element model depicting objects within the operating room. Similarly, sensors can be strategically placed in the operating room to focus on areas of interest. For example, sensors can be attached to monitors 125, 150, or a patient-side trolley 130, with their field of view focused on the surgical site of the patient 120, which may be attached to a wall or ceiling. Likewise, sensors can be placed on a control console 155 to monitor the operator 105c. Sensors can also be placed on movable platforms specifically designed to facilitate sensor orientation in various postures within the operating room.

[0019] For clarity, as used herein, "pose" refers to the translational position and rotational orientation of the body. For example, in three-dimensional space, a pose can be represented by a total of six degrees of freedom. It will be readily apparent that various data structures can be used to represent poses, such as matrices, quaternions, vectors, combinations thereof, etc. Therefore, in some cases, a pose may consist only of translational components when there is no rotation. Conversely, a pose may consist only of rotational components when there is no translation.

[0020] Similarly, for clarity, “theater-wide” sensor data as used herein refers to data acquired from one or more sensors configured to monitor a specific area of ​​the operating room (including all or part of the operating room) outside the operating room, such that the sensors can sense the presence or passage of the patient, personnel, equipment, or other object within at least a portion of that area throughout the procedure. Sensors thus configured to collect such “theater-wide” data are referred to herein as “theater-wide sensors.” For clarity, it will be understood that the specific area does not need to be strictly fixed throughout the procedure, as some sensors may, for example, cyclically pan their fields of view to increase the size of the specific area, even if this results in temporal gaps in the sensor data (gaps can be remedied by cooperative panning or field of view from other nearby sensors). Similarly, in some cases, a human or robotic system may be able to reposition the thermometer-wide sensor throughout the procedure, changing the specific area, for example, to better capture different tasks. Therefore, sensors 170a and 170c are thermometer-wide sensors configured to generate thermometer-wide data. In this paper, "visualization data" refers to visual intensity or depth image data captured from sensors. Therefore, visualization data may or may not be operating room-wide data. For example, visualization data captured at sensors 170a and 170c is operating room-wide data, while visualization data captured via visualization tool 140d will not be operating room-wide data (at least because the data is not outside the patient's field of vision).

[0021] Example Operating Room Global Sensor Topology To further clarify the deployment of sensors throughout the operating room Figure 2A This is a schematic depth map rendered from an example operating room global sensor viewpoint 205, which may be used in some embodiments. Specifically, this example depicts depth values ​​corresponding to electronics / consoles 205a (e.g., electronics / consoles 145) and nearby trays 205b and cabinets 205c. Also within the field of view are depth values ​​associated with a first technician 205d, who is currently adjusting a robotic arm on a robotic surgical system (associated with depth value 205e) (associated with depth value 205f). Team members with corresponding depth values ​​205g, 205h, and 205i are also present in the field of view, as is a portion of the operating table 205j. Depth values ​​205l corresponding to the movable platform cart and boom, and depth value 205k corresponding to the lighting system, are also present in the field of view.

[0022] The operating room-wide sensor capturing a 205-degree field of view may be just one of several sensors placed throughout the operating room. For example, Figure 2BThis is a schematic top view of objects in the operating room at a given moment during a surgical procedure. Specifically, viewpoint 205 may have been captured by an operating room global sensor 220a with a corresponding field of view 225a. Thus, for clarity, cabinet depth value 205c may correspond to cabinet 210c, electronics / console depth value 205a may correspond to electronics / console 210a, and tray depth value 205b may correspond to tray 210b. Robotic system 210e may correspond to depth value 205e, and each individual team member 210d, 210g, 210h, and 210i may correspond to depth values ​​205d, 205g, 205h, and 205i, respectively. Similarly, platform cart 210l may correspond to depth value 205l. Depth value 205j may correspond to operating table 210j (for clarity, the outline of the patient is shown here, although the patient has not yet been placed on the operating table corresponding to depth value 205j in example viewpoint 205). For clarity, a top view of the boom corresponding to a depth of 205k is not shown, but it will be understood that the boom can also be considered in various embodiments.

[0023] As noted, each of sensors 220a, 220b, and 220c is associated with a different field of view 225a, 225b, and 225c, respectively. Fields of view 225a-c can sometimes have complementary characteristics, providing different perspectives of the same object, offering a view of the object from another perspective when the object is occluded outside or within one perspective. The complementarity between perspectives can be dynamic both spatially and temporally. Such dynamic characteristics may be caused by the movement of the tracked object, but may also be caused by the movement of an intermediate occluded object (and in some cases, by the movement of the sensor itself). For example, in Figure 2A and Figure 2BAt the moment depicted, field of view 225a has only a limited view of the operating table 210j because the electronics / console 210a essentially obscures that portion of field of view 225a. Therefore, at the moment depicted, field of view 225b is better positioned to view the operating table 210j. However, neither field of view 225b nor 225a provides sufficient view of the operator 210n in console 210k. Field of view 225c may be more suitable for observing the operator 210n (e.g., when they move their head according to a “head protrusion” event). However, these complementary relationships may change during data acquisition. For example, before the procedure begins, the electronics / console 210a may be removed, and the robotic system 210e moves to position 210m. In this configuration, field of view 225a may be better suited to view the patient table 210j than field of view 225b. As another example, the posture of console 210k moving to the electronics / console 210a described in this disclosure may make field of view 225a better suited to view the operator 210n than field of view 225c. Therefore, the suitability of the field of view may depend on the amount and duration of occlusion, the quality of the field of view (e.g., how close the object of interest is to the sensor), and the movement of the object of interest within the operating room. Such changes may be temporary and short-lived, such as when a team member moving within the operating room briefly obstructs the sensor, or they may be long-term or continuous, such as when the device is moved to a fixed position throughout the duration of the procedure.

[0024] As mentioned, operating room omnidirectional sensors can take many forms and, for example, can be configured to acquire visual image data, depth data, and both visual and depth data. It will be understood that visual and depth image capture can also take various forms, for example, to provide increased visibility into different parts of the operating room. For example, Figure 2CImages 250b and 255b depict a grid pattern of orthogonal rows and columns captured from an operating room omnidirectional sensor with a straight-line view and an operating room omnidirectional sensor with a fisheye view, respectively. More specifically, some operating room omnidirectional sensors can capture straight-line visual images or straight-line depth frames, for example, via appropriate lenses, post-processing, or a combination of lenses and post-processing, while others can instead capture fisheye or distorted visual images or straight-line depth frames, for example, via appropriate lenses, post-processing, or a combination of lenses and post-processing. For clarity, image 250b depicts a checkerboard pattern from the viewpoint of the straight-line operating room omnidirectional sensor. Thus, the orthogonal rows and columns 250a shown here maintain a linear relationship with their vanishing points. In contrast, image 255b depicts the same checkerboard pattern from the same viewpoint, but from the viewpoint of the fisheye operating room omnidirectional sensor. Therefore, the orthogonal rows and columns 255a, although actually linearly related to their vanishing points (as they appear in image 250b), appear here to have a curvilinear relationship with their vanishing points according to the sensor data. Thus, by combining various embodiments, each type of sensor and other types of sensors can be used individually or in some cases in combination.

[0025] Similarly, it will be known that not all sensors can acquire perfect straight lines, fisheye, or other desired mappings. Therefore, a grid pattern or other calibration reference (such as the known shape of a depth system) can be helpful in determining the inherent parameters of a given operating room global sensor. For example, the focal point of a fisheye lens and other details of the operating room global sensor (base point, distortion coefficients, etc.) may vary from device to device and even change over time on the same device. Therefore, it may be necessary to recalibrate various processing methods for the specific device in question, anticipating device variations when training and configuring a system for machine learning tasks. Furthermore, it will be known that once the inherent parameters of the camera are known, a straight-line view can be achieved by making the fisheye view undistorted (this can be used, for example, to normalize different sensor systems to a similar form recognized by the machine learning architecture). Thus, while a fisheye view can allow the system and user to perceive a wider field of view more easily than in the case of a straight-line view, different perspectives can be normalized to a common perspective form (e.g., mapping all straight-line data to a fisheye representation and vice versa) when the processing system is considering data from some sensors acquiring undistorted perspectives and others acquiring distorted perspectives.

[0026] Overview of operative and non-operative operating room data Figure 3This is a schematic depiction of data acquired during the surgical procedure period, which may occur in conjunction with some embodiments. Specifically, over time 305, for example throughout the day, the operating room, whether robotic operating room 100a or non-robotic operating room 100b, may alternate between an operational period with surgical activity and a non-operational period without surgery. In this example, during the preoperative preliminary preparation period 310a at the beginning of the day, team members may prepare the operating room for the day's surgery. Then, during the interval 315a, the first surgery may be performed by the same or different team members. This may be followed by an interoperative period 310b, during which the team (which may or may not include the same composition as the team members in the previous period) prepares the operating room for the next surgery, such as changing equipment configuration, moving the patient in and out of the operating room, etc. The alternating operational and non-operational periods may then continue throughout the day, for example, as reflected in surgical period 315b, the nth surgery 315c, and the intermediate ellipsis 310e (which may indicate the presence of additional surgical procedure intervals). At the end of the day, during the 3-10 day post-operative period, team members can, for example, clean the operating room, set up equipment, and ensure that necessary items are available and used in an orderly manner for the next day's surgeries.

[0027] During each surgical phase 315a-c, operating room operations typically generate data divided into two parallel sets: data acquired by the surgical instruments 375a and operating room-wide data 375b. Each of these data sets can correspond to the execution of a specific task. For example, data acquired by the surgical instruments 375a, as mentioned, does not necessarily originate from the robotic system and can be acquired in conjunction with surgical tasks 370a, 370b, 370c, and 370d (the ellipsis 370e indicates the possibility of additional intermediate tasks). As shown, each of tasks 370a-d can be associated with corresponding surgical data, such as visual image video data 320a-d (in some cases, depth video data may also be available or alternatively). Similarly, kinematic data 325a-d can be acquired from the robotic platform, surgeon's console, instruments themselves, etc. This type of data can indicate the movement and posture of various instruments, end effectors, tools, etc., throughout each corresponding task. Similarly, system event data 335a-d can be acquired in conjunction with each task, indicating when instruments are activated (e.g., endoscope position lock, cauterization tool, surgical scissors activation, etc.). For clarity, the ellipses 320e, 325e, and 335e also indicate the possibility of additional intermediate tasks and corresponding data.

[0028] Therefore, surgical data 375a typically corresponds to data generated by the actions of one or more surgeons. In contrast, operating room global data 75b can be acquired from one or more operating room global sensors placed around the operating room and can typically depict the actions of team members within the operating room, particularly during the execution of operating room tasks 330a-d (ellipsis 330e indicates the possibility of additional tasks). Since the tasks are distinct, although sometimes related, tasks 330a-d and 370a-d. In this example, there are three data streams 340a-d, 350a-d, and 360a-d (ellipsis 380 indicates the possibility of additional data streams from other sensors in some embodiments, although it will be understood that fewer than three streams may also occur in some embodiments), corresponding to three different operating room global sensor systems with different corresponding postures within the operating room (ellipsis 340e, 350e, and 360e similarly indicate the possibility of additional datasets for the stream data of each corresponding row). Although streams 340a-d, 350a-d, and 360a-d are shown herein as visual image data, as discussed earlier, it will be understood that the stream may additionally or alternatively include, for example, depth data.

[0029] While many tasks 330a-d are performed in conjunction with or before tasks 370a-d (but of course, not all tasks are), tasks 330a-d performed in the operating room may or may not correspond to surgical tasks 370a-d in time, their number may differ, and their start and stop times may differ. That is, although each of the surgical data 375a and the operating room-wide data 375b can be acquired relative to a common timing device, the start and end times of the various tasks 330a-d and 370a-d may not correspond. For example, here, surgical task 370c may require an imaging system. Therefore, as indicated by video data image 340b of task 330b, team members have begun moving the imaging system to the location of task 370c. Therefore, regardless of their relationship, the start and end times of task 330b may be earlier than those of task 370c. Thus, the start and end times, durations, and number of tasks 330a-d and 370a-d, while sometimes related, may often differ. Finally, for clarity, note that although shown here as linearly sequential in time, one will understand that some tasks can be performed in parallel, such as multiple team members simultaneously performing their own role-specific tasks in the operating room.

[0030] Therefore, if the goal is to identify tasks 330a-d or specific actions from the operating room global sensor data stream 375b, and surgical data 375a is available, then a reviewer or machine learning system might be able to make partial inferences and correspondingly identify and label at least some sets of the operating room global data 375b, associating them with one of the tasks 330a-d (or other corresponding operating room states, patient states, etc. of interest). Unfortunately, the certainty in the annotations may require human review, and such manual checks may, for example, limit the size of the annotated dataset, slow down such annotation, and introduce the risk of human error.

[0031] During the interoperative period 310a-d, such annotation may be more difficult because surgical data 375a may no longer be available. For example, Figure 4 This is a schematic depiction of data that can be acquired during the non-surgical operation period 310c. Here, three data streams remain active, acquiring the corresponding datasets 425a-e, 430a-e, and 435a-e during the execution of various in-operation tasks 420a-e (again, each of the ellipses 420f, 425f, 430f, and 435f indicates the possibility of additional intermediate tasks and datasets). Because the corresponding surgical data 375a may not be available during these non-operational periods, the annotations may rely entirely on human or machine learning examination of the operating room-wide sensor datasets 425a-f, 430a-f, and 435a-f.

[0032] Example downstream applications People will become aware of many downstream machine learning applications that can be performed on operating room data, especially in low-data situations, as enabled by various publicly available embodiments. For example, Figure 5A This is a schematic representation of multiple operating room state recognitions that can be performed in the embodiments. A downstream machine learning system can receive operating room global data and is required to identify the operating room state from one or more frames. Here, for example, six operating room global visual intensity video frames 505a-f have been presented to the system, and the system is required to identify the corresponding operating room state from a taxonomy (e.g., “initial” setup, “aseptic preparation”, “patient entry”, etc.). Despite the low data conditions, performing such taxonomy recognition can facilitate various additional downstream applications and analyses (e.g., once the video data has been segmented into its corresponding taxonomy categories, evaluation of the surgical team's performance on various tasks relative to other hospitals, teams, configurations, etc., may be readily feasible).

[0033] As another example of a downstream application Figure 5BThis is a schematic representation of multiple object segmentation results that can be performed in some embodiments. Specifically, it is similar to the "You Only See Once" (YOLO) object detection system (in fact, the neural network can be modified to resemble the YOLO architecture). Figure 5B Various operating room-wide video depth frames are depicted, in which team members are identified, and their corresponding depth values ​​are highlighted in frames 550a-c. In depth frame 550d, the system identifies a specific piece of equipment, in this case, the patient bed. As in the example of 5A, such a system can facilitate a variety of downstream analyses. For example, knowing when team members arrive, when equipment is used, the relative orientation of team members and equipment, and the patient's location and orientation, can facilitate a more refined assessment of the surgical team's performance.

[0034] although Figure 5A and Figure 5B The example refers to downstream classification applications, but the reader will realize that applications other than classification can also be enabled through the disclosed embodiments. For example, once the neural network has been pre-trained to understand general semantic patterns, predictive applications seeking to predict surgical procedures, and generative applications that may seek to generate synthetic operating room data, may become more feasible.

[0035] Example downstream application of operating room identification and classification method Therefore, downstream operating room machine learning applications include, for example: segmenting images or videos based on objects depicted in them; recognizing the roles of team members; identifying the actions of team members; determining the patient's condition; and detecting adverse events in the operating room. Many of these downstream operating room applications may seek to classify the data according to taxonomy. For example, surgical activity recognition may attempt to classify the data into one or more of a plurality of actions depicting events occurring in the operating room.

[0036] As an example taxonomy (primarily focusing on non-operational activities). Figure 6A This is a schematic representation of a set of time intervals that can be used in some embodiments to evaluate activities in the operating room. Figure 6A The grouping depicted in Figure B is particularly useful for processing multimodal operating room data. Specifically, Figure 6A The grouping provides example "syntax" describing operating room workflows. In this example, activities typically do not overlap, although substitution syntax can allow for such time overlap.

[0037] People will know, according to the above Figure 3 and Figure 4The alternation of operational and non-operational periods in the operating room, as described, is applied cyclically through intervals. For example, initially, surgical operation 315b could correspond to interval 650e. After operation 315b is completed, the actions and corresponding data in the operating room can be assigned to consecutive intervals 650a-d during the subsequent non-operational period 310c. Then, the data and actions in the next operation (e.g., operation 315c if there is no intermediate period in ellipsis 310e) can similarly be assigned to a second instance of interval 650e, and so on (therefore, data from each of the non-operational periods 310b, 310b can be assigned to instances of intervals 650a-d). Intervals can also be grouped into larger intervals, such as the "pull-to-advance" interval 650f here, which groups intervals 650b and 650c, sharing the start time of interval 650b and the end time of interval 650c. The ability to automatically incorporate operating room data into this example taxonomy can benefit a variety of downstream applications. Supervised learning methods can be sufficient to enable such applications when such labeled data is readily available, but as will be discussed, such labeled data is often unavailable.

[0038] To provide further background, Figure 6B Depicting based on Figure 6A Another, more refined action classification of the example taxonomy. Specifically, the duration of each interval 650a-e can be determined based on the corresponding start and end times of various tasks or actions within the operating room. Naturally, when intervals 650a-e are used consecutively, the end time of the previous interval (e.g., the end of interval 650c) can be the start time of the subsequent interval (e.g., the start of interval 650d). When combined with a task-action grouping ontology, the operating room-wide data can be easily grouped into meaningful partitions for downstream analysis. This can be beneficial, for example, in verifying the consistency of team members' adherence to proposed feedback, and in conducting computer-based validation across different operating rooms, team configurations, etc. As will be explained, some task actions can occur over a period of time (e.g., cleaning), while other task actions can occur at a specific moment (e.g., the entry of a team member).

[0039] Specifically, Figure 6BFour high-level task action classes are described: Postoperative 620, Turnover 625, Preoperative 610, and Surgery 615. Surgery 615 may include tasks or actions 615a-i. Specifically, the task “First Cut” 615a indicates the time when the first incision occurs on the patient. The task “Port Placement” 615b indicates the time when the first port is placed into the patient. The task “Rollup” 615c is the duration from when the team members begin moving the robotic system to when the robotic system adopts a posture it will use during at least the initial part of the surgical procedure. The task “Room Preparation” 615d begins with the first surgical preparation action for the ongoing surgery and ends with the last preparation action for the ongoing surgery. The task “Docking” 615e begins when the team members begin docking the robotic system and ends when the robotic system docks. The task “Surgery” 615f begins with the first incision and ends with the final closure of the patient. Naturally, in many taxonomics, block 615f can be further broken down into more actions and tasks (this can be facilitated, for example, by the availability of surgical data 375a). Task “Disengagement” 615g begins when team members begin disengaging the robotic system and ends when the robotic system is disengaged. Task “Rollback” 615h begins when team members begin removing the robotic system from the patient and ends when the robotic system assumes a posture it will maintain until the turnaround begins. Task “Patient Closure” 615a begins and ends with the final suturing of the patient.

[0040] In postoperative category group 620, task "Robotic Removal of Covers" 620a begins when a team member first begins removing the robotic system's covers and ends when the robotic system removes the covers. Task "Patient Discharge" 620b begins and ends when the patient leaves the operating room. Task "Patient Removal of Covers" 620c begins when a team member first begins removing the patient's covers and ends when the patient's covers are removed.

[0041] In the turnaround category group 625, task "Cleaning" 625a begins when the first team member begins cleaning the equipment in the operating room and ends when the last team member (who may be the same team member) completes the final cleaning of any equipment. Task "Idle" 625b begins when team members are not performing any other tasks and ends when they begin performing another task. Task "Turnaround" 605a begins when the first team member resets the operating room from the last procedure and ends when the last team member (who may be the same team member) completes the reset. Task "Setup" 605b begins when the first team member begins changing the position of the equipment to be used in the surgery and ends when the last team member (who may be the same team member) completes the final equipment position adjustment. Task "Sterile Preparation" 605c begins when the first team member begins cleaning the surgical area and ends when the last team member (who may be the same team member) completes cleaning the surgical area. Furthermore, although shown here in a linear order, it will be understood that the tasks within a category may be performed in a different order than shown, or in some cases, may refer to potentially overlapping time periods.

[0042] In the preoperative category grouping 610, task "Patient Entry" 610a begins and ends when the patient first enters the operating room. Task "Robotic Coverage" 610b begins when team members begin covering the robotic system and ends when coverage is complete. Task "Intubation" 610c begins when patient intubation begins and ends when intubation is complete. Task "Patient Preparation" 610d begins when team members begin preparing the patient for surgery and ends when preparation is complete. Task "Patient Coverage" 610e begins when team members begin covering the patient and ends when the patient is covered.

[0043] Therefore, as Figure 6B As indicated by the corresponding arrow in the image. Figure 6A The intervals can be allocated as follows. "Skin closure to patient departure" 650a can begin at the last closure operation 615i of the previous surgical interval and end with the patient leaving the operating room (e.g., from the last suture at block 615i until the patient leaves at block 620b). Similarly, the interval "patient departure to case-open" 650b can begin at block 620b when the patient leaves the operating room and end at block 605c when aseptic preparation for the next case begins.

[0044] Interval 650c, "Patient preparation to patient entry," can begin at block 605c and end at block 610a when a new patient enters the operating room. Interval 650d, "Patient entry to skin incision," can begin at block 610a when a new patient enters the operating room and end at block 615a when the first incision begins. As shown, the surgery itself can occur during interval 650e. As previously discussed, the "pull-out to advance" interval 650f begins at the start of "Patient departure to patient preparation" 650b and ends at the end of "Patient preparation to patient entry" 650c.

[0045] and Figure 6A Like a more inclusive interval, it has identifiable Figure 6B More fine-grained activity-based machine learning classifiers are often desired, for example, from operating room-wide sensor data. In fact, machine learning systems capable of identifying personnel, equipment, adverse events, operating room-wide video data, operating room configurations, etc., in the operating room during operative periods, non-operative periods, and both operative and non-operative periods would benefit many downstream applications. However, preparing such machine learning systems solely through supervised methods may be impractical without adequately labeled data. Similarly, while the above and other examples in this paper may focus on operating room-wide data for ease of understanding, various embodiments can be applied to non-operating room-wide data (e.g., in-patient videos from visualization tools, kinematic data, etc.), or combinations of operating room-wide and non-operating room-wide operating room data.

[0046] Example dataset properties As mentioned earlier, many machine learning applications using operating room data remain infeasible because they rely on large amounts of labeled operating room data. For clarity, Figure 7A and Figure 7B Various types of labeled dataset instances are illustrated.

[0047] Specifically, Figure 7A An example of entering operating room data item 730 is illustrated in the context of an exemplary classification system. Here, data item 730 includes operating room data 730a to be classified (here, indicated by a question mark before it reaches a classifier with unknown categorization), which, as mentioned, can be a single visual image or depth frame, a video of each, etc. The classifier machine learning system can then, for example, seek to determine which of several categories data 730a should be classified into (e.g., ...). Figure 5A , Figures 6A-6B It is one of the operating room status categories, or it may be Figure 5BThe probability 730b is associated with one of the object categories. Although only three categories A, B, and C are shown here for the reader's understanding, the reader will realize that the probability can be distributed across only two categories, or across more categories. Similarly, although this article will periodically mention classification applications as recurring contextual examples to aid the reader's understanding, as mentioned, the reader will realize that downstream machine learning systems can perform other applications besides classification, such as synthetic data generation, data extrapolation, and future prediction.

[0048] In order to train a classifier in a supervised manner, to assign probabilities to, for example Figures 5A-5B and Figures 6A-6B The categories shown, ideally, can yield results such as Figure 7B The “fully labeled dataset” 735 shown is used to perform supervised training. In such datasets, each training instance 735a-c presents operating room data with a known classification (ellipsis 735d indicates the possibility of additional intermediate terms). Thus, data instance 735a includes operating room data known to be in class C, and therefore, the known probability of this instance is 1.0 for class C, and 0.0 for the other classes. Similarly, instance 735b includes operating room data known to be in class B, so the known probability of this instance is 1.0 for class B, and so on.

[0049] Unfortunately, as mentioned above regarding Figure 3 and Figure 4 The “fully labeled datasets” 735 of the operating room data 375a-b are generally difficult to obtain. Operating room datasets, even those that only include surgical data across the entire operating room, typically contain millions of images from various surgical settings, procedures, ORs, etc., making manual annotation by experts difficult, if not infeasible.

[0050] As mentioned, with the aid of surgical data 375a, it is possible to infer the labeling of some portions of data 375a-b simply by examining them (e.g., some tools will only be activated during surgery or certain surgical tasks). Such inference could allow for some automated classification, providing some labeled data instances, which could then be supplemented by expensive and time-consuming expert manual annotation. However, as... Figure 7B As shown, this may only result in a “partially labeled dataset” 740, where only a few instances have the correct labels. Here, for example, it is known that instance 740a describes category C and should receive the corresponding labels for training. However, it is unknown which category and therefore what labels are suitable for instances 740b and 740c (the ellipsis 740d indicates the possibility of additional intermediate instances).

[0051] In practice, given the difficulty of labeling operating room data, the following situation may sometimes occur: no available data is labeled, resulting in an unlabeled dataset 745, in which the nature of the global operating room data for all instances 745a-d is unknown, and therefore it is unclear what classification they should accept. Indeed, as mentioned above, expert annotation can be cumbersome and expensive, and so much of the acquired operating room data will be in the form of an unlabeled dataset 745.

[0052] The disclosed embodiments reduce the amount of annotated data required to achieve the desired level of machine learning system performance in downstream applications. As will be discussed in the following sections, the various embodiments employ coupled encoder-decoder or autoencoder methods that leverage a large amount of available unlabeled data to infer semantic content suitable for downstream training. Such downstream training can then utilize any available labeled data, even if the amount is small. In this way, downstream applications can achieve performance better than that achievable in such low-labeled data conditions.

[0053] Example of unsupervised pre-trained machine learning methods Since most operating room data 375a-b is in the form of unlabeled datasets 745, various publicly disclosed embodiments envision a variety of unsupervised learning operations to adapt machine learning systems to the required syntax of a variety of downstream applications. Intuitively, one would expect that various machine learning applications in the operating room would rely on similar "semantic knowledge" unique to the operating room environment (e.g., "what the robotic surgical system looks like," "where the surgeon is when surgery begins," "how people move around in the operating room," "how patients appear in the OR and what orientation they typically take," etc.).

[0054] More specifically, various publicly available embodiments can train machine learning systems to reduce operating room data to a concise, “compressed” form and then re-drive the original surgical data from that concise form. Training in this way forces the system to identify the most salient semantic features of the operating room data, and such identification can then be leveraged in downstream training to perform more specific application tasks against a much smaller labeled dataset that might otherwise be needed. In a sense, the initial unsupervised training “prepares” the machine learning system for downstream tasks, making it require less labeled data to achieve the desired performance during training. Various clustering methods for unlabeled semantic content recognition, such as Swap Assignment (SwAV) across multiple views of the same image, are illustrated in Caron, Mathilde, et al.’s article, “Unsupervised Learning of Visual Features by Contrasting Cluster Assignments,” arXiv.TM preprint arXiv TM As discussed in 2006.09882 (2020), designers are required to predict the number of prototype clusters in advance. Various publicly available reconstruction methods do not impose this constraint on designers, but rather better allow the system to organically adapt itself to the true semantic content behind the data. Therefore, it has been found that the various publicly available embodiments are easier to generalize than certain corresponding SwAV implementations.

[0055] Similarly, although specific categories are often presented here for ease of understanding, readers will realize that the disclosed methods can be used in a wide range of downstream applications, such as 2D visual image or depth data segmentation, 3D visual image / video or depth video data segmentation, surgical activity recognition, semantic segmentation, and so on. Furthermore, readers will learn that the various disclosed methods also address privacy concerns, a crucial issue in healthcare settings, as the data and methods used maintain the anonymity of individuals present in the operating room.

[0056] Therefore, with the labeled dataset 735 available, the classifier can be trained in a supervised manner, such as... Figure 8A The process 805 is shown. In such a process 805, supervised training can be simply performed at block 805a using labeled data, and then the classifier is deployed for classification at block 805b. While this is appropriate when a large amount of labeled data is available, as mentioned, it is generally not feasible to bring an initially untrained classifier to a training state with only a small amount of labeled data using only the process 805.

[0057] However, when partially labeled data (740) is available, a combination of unsupervised and supervised training can sometimes be applied to produce a machine learning classifier with sufficient performance. Specifically, such as... Figure 8B As shown in process 810, instead of directly performing the supervised classification of process 805, preliminary context training 810a can be performed first, which can be unsupervised (or, in some cases, semi-supervised if some labeled data is available).

[0058] During the initial context training 710a, at block 715a, the system can perform unsupervised training of the classifier using unlabeled data. Specifically, it is expected that the data will fall into one of the groups of downstream tasks (e.g., Figures 6A-6BEven if the explicit identifier of such semantic patterns is unknown, unsupervised training can seek to first train a machine learning system to recognize general operating room semantic content. For example, patient movement as a semantic pattern in the data may only occur during the actions of "patient enters" 610a and "patient leaves" 620b. Therefore, a machine learning system trained to coarsely distinguish between data where the patient is in motion and data where the patient is not in motion will have been better prepared in advance to distinguish the actions "patient enters" 610a and "patient leaves" 620b from other actions. Thus, after pre-training at block 815a, at block 815b, the classifier can be adapted to a form suitable for supervised training (e.g., as described in more detail below), and this modified architecture is used for task-specific training 810b, which, like process 805, performs supervised training at block 815c using available known labeled data, and then deploys the classifier at block 815d for its desired downstream purpose.

[0059] In practice, using some of the embodiments described herein, unsupervised methods can be very successful for certain applications where the system can be deployed without performing task-specific training 810b. This may occur, for example, when the predictions of unsupervised discrimination sufficiently correspond to the desired downstream behavior, allowing human reviewers to simply annotate the downstream behavior by checking it with their appropriate labels. In these cases, some minimal architectural tweaks or post-processing can be used to refine the unsupervised groupings to the desired form. However, for many applications, it is necessary to perform additional training according to task-specific training 810b.

[0060] Before describing in detail the various encoder / decoder embodiments used for initial unsupervised context training... Figure 8C As may be employed in some embodiments, the reader is first provided with directional background by describing the general adaptation process. Specifically, Figure 8C The text describes a machine learning system during deployment 820 (e.g., the classifier following block 815d), supervised training 825 (e.g., the state of the machine learning system at block 815c), adaptive classification 830 (e.g., the state of the machine learning system at block 815b), and supervised training state 870 (e.g., the state of the machine learning system at block 815a). The reader will appreciate that in some implementations and embodiments, training 870 may include some labeled data and is therefore “semi-supervised.” For example, instead of sequentially, as will be discussed in this example, training may alternate between unsupervised and supervised training, resulting in a “semi-supervised” training approach (e.g., continuously applying and removing adaptations 830 during training).

[0061] In this example, training processes 870 and 825 will produce a machine learning system, such as a classifier 820b, which, when presented with operating room data 820a (kinematic data, system event data, operating room global visual intensity and depth data, combinations of operating room global, kinematic, and system event data, etc.) of unknown category (here, "C" only indicates correct classification), is able to perform its desired task, in this case, meaningful classification prediction. For example, classifier 820b successfully identifies category C as the most probable category in its probability output 820c with a probability of 0.7. Such a classifier can be created using supervised training 825, as discussed in block 815c of processes 805 and 810. Specifically, when known instances of classified data are available, such as instances 825a associated with category C, then during training, the system can compare the classifier's output 825c with known labels 825d of the labeled data and update the weights of classifier 825b using the differences 825e, for example, using known backpropagation techniques. The completion of training 825 can produce a trained and deployed classifier 820b, 840c.

[0062] As mentioned, however, when the amount of available labeled data is insufficient, or when one expects to maximize the use of the labeled data available during supervised training 825 (initially or during subsequent online training), then initial unsupervised training 870 can be used to put the classifier into a state where it can be trained in a supervised manner with a limited number of labels. For example, from Figure 6A In this context, one can recognize that the data will come from one of five categories 650a-e, from... Figure 6B One of the finite activity tasks in the dataset, and so on, can each be associated with latent semantic features. Some of these latent semantic features may be unique to a single classification (e.g., patient absence during certain non-operational periods), while others may be common to all classifications but with some variation (e.g., team member movement). Therefore, during unsupervised training, a machine learning system can be trained to discern latent encoded representations that can be compressed relative to the input tensor of the unlabeled data (e.g., having fewer dimensions) to fit the reconstruction of the original unlabeled data, indirectly capturing various of these latent semantic features in the encoded representation (e.g., patient motion, equipment motion, surgeon inactivity, instrument size and rigidity, robot system posture patterns, etc.).

[0063] In this example, the unsupervised initial training 870 includes an encoder-decoder machine learning system 850, such as an autoencoder or different encoders and decoders. Here, unlabeled data is shown as multiple tensors, represented by tensor 855a, tensor 855c, and an ellipsis 855b indicating the possibility of additional tensors. For clarity, "tensor" here simply refers to a data input with one or more dimensions and should not be confused with, for example, a "metric tensor" in differential geometry (therefore, no corresponding metric or manifold is implied for interpreting tensors). Thus, a "matrix" is simply a two-dimensional tensor. For example, tensor 855c may include consecutive video image frames, thus constituting a three-dimensional tensor (each frame has two dimensions, and the third temporal dimension corresponds to the relative acquisition time of each frame).

[0064] Encoder 860a can be configured to receive tensor 855c (e.g., in the case where the input tensor has three dimensions, receiving a tensor in a Conv3d layer, reformatting the tensor to receive it in a linear layer, where the input nodes correspond to the product of all tensor dimensions, etc.). Encoder 860a can then produce a compressed coded representation 860c. For example, as in an autoencoder, the number of nodes in the layers between encoder 860a and decoder 860b can simply be less than the number at the input of encoder 860a, thus forcing the input data to be represented in a “reduced latent space” form. Presented in this simplified form, decoder 860b can then seek to recover the data of the original tensor 855c, shown here as the reconstructed tensor 855d.

[0065] By determining the difference 875a between the original tensor 855c and its recovered counterpart 855d, the quality of the current operation of the encoder 860a and decoder 860b (and implying the quality of the underlying representation 860c) can be evaluated. Initially during training, one would not expect the reconstructed counterpart tensor 855d to be very similar to the original tensor 855c, so the difference 875a and the associated loss 875b may be large, which are then used to update the encoder 860a and decoder 860b, as indicated by arrow 875c (e.g., via backpropagation). Naturally, although only a reconstruction of a single instance is shown here, the reader will understand that, as will be described in more detail, the loss 875b can incorporate the differences 875a between multiple tensors within a training batch.

[0066] Such iterative reconstruction, loss determination, and updates of encoder 860a and decoder 860b (again, which can be a single autoencoder neural network system, rather than two distinct systems) can ultimately result in encoder / decoder system 850 being able to reconstruct the input tensor relatively consistently using its corresponding compressed latent space representation. Once in this satisfactory state, as indicated by arrow 840a, system 850 can then be adapted 830 for downstream applications, such as as part of supervised training 825 for a classifier. Such adaptation can be achieved, for example, via structural modifications at block 830b (e.g., attaching one or more layers to decoder 860b, consistent with the number of classes used for classification or the type of desired output, removing decoder 860b and attaching a new neural network to encode the interpretation of representation 860c, etc.). Such adaptation can then yield machine learning system 825b suitable for supervised training on available labeled data.

[0067] Example of unsupervised training based on reconstruction Although some embodiments may employ, as Figure 8C The encoding and decoding operations described herein do not involve masking; however, many embodiments may employ masking, or otherwise force reconstruction from less than the entire input data, in order to better encourage the system to infer semantic relationships between the data. Specifically, Figure 9 A schematic block diagram illustrating example loss determination iterations during unsupervised training of the reconstruction system (again, for ease of understanding, for a single instance), as may be implemented in some embodiments, is provided. Specifically, in this example, it is represented as... V Furthermore, tensor 905a, corresponding to input tensor 855c, includes, for example, visual intensity video image frames of the operating room global data captured by one of sensors 220a-c. Similarly, although visual intensity video is used in this example, it should be understood that depth or other data, including non-operating room global surgical data, may be used alternatively. Here, consecutive frames of video (possibly downsampled by decimation or upsampled by interpolation) appear consecutively in the temporal dimension of the tensor, providing their respective two-dimensional intensity values ​​at each time instance (in some embodiments, each image may be associated with a color pixel, which itself comprises three dimensions, thus the number of tensor dimensions can be increased accordingly).

[0068] Typically, the latent space feature optimization training system 905 corresponding to the encoder-decoder machine learning system 850 can tokenize the tensor 905a 910a to produce n tokens 950a, represented in the figure as follows (although it goes without saying, for clarity, like other boundary variables used in this article for the reader's understanding, the n tokens mentioned here do not need to be compared with...). Figure 3 (The number of surgeries N cited in the text is the same). When partitioning along the spatial dimension of the frame, lexical unit 950a may include, for example, discrete portions of tensor 905a, resulting in columns, for example, having length along the temporal dimension, but with a height and width less than the full height and width of the frame. These column segments may or may not extend to the full temporal duration of the tensor. For example, if the spatial dimensions are equal and divided into four equal regions, and each column segment is extended only half the total temporal length of tensor 905a, then eight segments will be produced (N=8).

[0069] These spatiotemporal samples of 910b can then lead to selection 950b, representing less than all of these 950a segments. For example, spatiotemporal sampling 910b can randomly select word segments, select word segments based on a pre-selected pattern or logic, select word segments based on pattern rotation, and so on. Such sampling can be part of a masking operation, for example, as will be referenced in this paper. Figures 10A-10E More detailed description. The system can then present word selections 950b to the encoder system 910c (e.g., to a “cube embedding layer” or Conv3D layer, etc.), in this example, using a visual transducer (ViT) model pre-trained on, for example, the Kinetics-400 dataset or the ImageNet-21k dataset, as described by Dosovitskiy, Alexey et al. in “An image is worth 16x16 words: Transformers for image recognition at scale” arXiv TM As described in the preprint arXiv:2010.11929 (2020). The encoder can then produce the latent spatial embedding 950c of the sampled word segmentation 950c, denoted here as... .

[0070] The system then combines the latent space embeddings 950c of the sampled word segments with the representations of the word segments not selected by spatiotemporal sampling 910b to produce a "complete" latent representation 950d for decoding by the decoder 910d. The representations of the word segments not selected by spatiotemporal sampling (here denoted as...) ) can simply be a space word (e.g., a mask corresponding to input 950a, and can be "zeroed out" accordingly). Therefore, This indicates that the patch will be unmasked (patch) Embeddings of ) are appended to the "masking tokens", for example, initialized with zero, although they may be learnable parameters ( ).

[0071] Therefore, for clarity, if the image has 100 blocks and the system masks 80%, then... This corresponds to the embedding of 20 blocks, and uses zeros. The initial masked word segmentation corresponds to 80 masked blocks. Therefore, the append operation can produce 100 blocks corresponding to the dimensions of the original input, and thus, in this example, is suitable for the decoder to create the reconstructed input.

[0072] Like the encoder 910c, the decoder 910d can also be a transducer-based architecture. For example, the decoder 910d could be a vision transducer-based model with decoder depths of four and six attention heads. (See reference...) Figure 11 As discussed, when the system considers multiple modalities, the decoder 910d can also have two fully connected layers for each modality of the reconstructed output (e.g., projecting the latent representation into a two-dimensional pixel space). In some embodiments, the decoder 910d can have an embedding dimension of 384.

[0073] Decoder 910d can be a series of neural network layers that can then generate a reconstructed tensor word segmentation 950e from input tensor 905a, represented as... V Rec If the output token 950e is not yet in a form corresponding to the input tensor 905a, the system can reorganize the output token 950e 920 into an appropriate form 905b for comparison with the input 905a. As discussed, the difference 925b between the original input tensor 905a and the reconstructed output tensor 905b can then be used to determine the loss 925a, which is used to update the encoder 910c and decoder 910d, for example, corresponding to the difference 875a and the loss 875b. Iterative determination of the loss for various tensor inputs in this way will ultimately produce a system capable of reconstructing the input tensor with the desired fidelity, and, by implication, prepare a compressed representation 950c capable of capturing the significant semantic content of the data.

[0074] Example masking operation Various embodiments may employ various strategies to perform masking, such as in spatiotemporal sampling 910b. In some embodiments, when processing operating room global data, the segmentation or “tube” of unmasked data in the tensor may include 16 frames after downsampling to every four frames, as this has been found to be effective for the system in inferring desired semantic relationships within the data in some cases. For example, in some embodiments, the raw frame rate of operating room video (interior camera video, operating room global visual intensity video, operating room global depth frame video, etc.) is five frames per second. The reader will appreciate that the depth frame and visual intensity capture rates may be the same or different (where they are different, downsampling or interpolation may be applied). However, while masking can be specifically chosen to construct local tubes of such characteristics, various embodiments may employ other masking strategies to achieve various objectives. In fact, in some embodiments, different masking strategies may be employed throughout training to fine-tune the semantic awareness of the system. For example, the expectation of identifying more spatiotemporally local or more spatiotemporally global semantic relationships via masking may be communicated by anticipated downstream applications. Ablation studies can be conducted on various masking options to determine which masking strategy or combination of strategies yields optimal results for the intended application and low-data conditions. For example, depending on the nature of the collected data, a masking method used to detect the semantic properties of motion in a robotic surgical system may differ from a masking method focused on the motion of team members.

[0075] As an example of a masking strategy, Figure 10A A schematic tensor decomposition illustrating component masking values ​​is depicted. In this schematic 3D tensor 1005, which could be, for example, the global depth of the operating room or a visual intensity video (the reader will know that for other surgical data modalities, there may be other suitable tensors with different dimensions), there are only four consecutive frames 1005a-d. In this example, different masks are applied to each consecutive frame. Such fine-grained strategies may not be suitable when attempting to have the system infer semantic temporal relationships, but may be useful when attempting to identify semantic spatial relationships (e.g., if the downstream application will focus on the segmentation of individual video frames). Therefore, depending on the nature of the downstream application, emphasizing spatial relationships over temporal relationships may be appropriate. The expected movement speed, the size of objects suitable for the downstream application, temporal relevance, the expected number of people in the operating room, etc., can therefore serve as priors informing the masking selection. Figure 10A Masking strategies may be suitable for image-based data (e.g., when the input is a single 2D visual or depth frame image).

[0076] As another example, in Figure 10BIn this context, the same masking has been applied to each frame 1010a-d of tensor 1010. This continuous application of the same masking over all or part of the temporal duration of tensor 1010 can precipitate "tubes" of the aforementioned properties. Such temporal tubes can be used to encourage the system to recognize temporal relationships between spatially local elements of frames. For example, one would expect a person pushing a cart to always maintain a certain posture, and the cart to always exhibit a corresponding relative linear motion. (via...) Figure 10B Masking strategies, which encourage the system to recognize relationships between descriptions of pushing team members and descriptions of pushing carts over time without complete access to one or the other, can teach the system to understand associations between two spatial and temporal semantic patterns (although this is an example of discussing relationships within a modality, one will understand that masking can be similarly tuned to identify relationships between modalities, e.g., when depth and visual image data share relationships with kinematic data). Therefore, Figure 10A Masking strategies may be suitable for video-based data (e.g., receiving 3D vision or depth frame video as input).

[0077] Despite targeting Figure 10B Example strategies have been discussed, but for clarity, Figure 10C An example is provided that explicitly indicates how to apply the same mask to a consecutive set of frames within tensor 1015, and how to apply a second mask to a second set of frames within tensor 1015 (here, the first mask is used for frames 1015a and 1015b, and then the second mask is used for frames 1015c and 1015d). This naturally leads to... Figure 10B The strategy involves shorter-duration tubes.

[0078] In some cases, encouraging the system to specifically perform time interpolation can facilitate the system's representation of long-running relations. For example, in Figure 10D In this strategy, the entire frames 1020b and 1020c are masked by tensor 1020, forcing the system to infer intermediate values ​​from frames 1020a and 1020d. This strategy may be suitable for inferring certain types of motion. The reader will be aware of the variations, as only the last few frames or the first few frames are not masked, and the system must predict the preceding or subsequent frames separately. Therefore, temporally concentrated masking can encourage the system to identify temporal relationships between clustered phenomena (within a modality or across modalities), for example, to characterize the motion and momentum of objects commonly found in an operating room. In some embodiments, when... Figures 10A-10COne of the methods in the text is applied after the previous training batch. This temporal masking can be applied to the training batch, for example, to encourage the system to recognize global temporal relationships between previously identified local or spatial semantic patterns (naturally, this policy change may or may not be applied to peer modalities during the same training period, depending on the nature of the relationship being identified).

[0079] For clarity, the reader will understand that the semantic knowledge obtained through the disclosed tensor masking can readily facilitate downstream applications and architectures (according to Adaptation 830) that can receive and produce inputs and outputs with dimensions very different from the tensors applied during pre-training in that latent space. For example, a two-dimensional visual intensity image or depth image segmentation application can receive and output a two-dimensional tensor. Since consecutive temporal frames may not be provided to the system in this example, but only to individual images, the system can utilize less of the semantic knowledge acquired during pre-training. However, due to various masking methods, such as... Figures 10A-10D While masking may still emphasize the identification of local spatial relationships, latent spatial pre-training can still greatly improve subsequent supervised training of segmentation architectures.

[0080] This serves as a specific example of masking for operating room data across the entire operating room. Figure 10E A pair of intensity tensors 1025a and depth tensors 1030a, and their corresponding masking representations 1025c and 1030c, which can be implemented in some embodiments, are illustrated. In some embodiments, masking operations 1025b and 1030b can apply the same mask to each of the respective tensors, or scale the mask according to different magnitudes of the respective dimensions of the tensors. As will be discussed in the next section, various embodiments have been found to achieve the desired results by employing an architecture that adapts to operating room data considering different modalities, such as operating room-wide data of intensity 1025a and depth 1030a tensors. For clarity, it can be understood that a masking strategy applied when only one modality is considered can also be applied to that modality when it is considered in series with other modalities (instead, for example, complementary masking operations between modalities as discussed elsewhere herein).

[0081] Although, for ease of understanding, we have so far primarily discussed three-dimensional tensors that depict the intensity of visual images in video frames, various embodiments can also employ masking strategies for other types of data and tensors other than three dimensions, which will be discussed in more detail here. In practice, different details of different modalities can be masked either individually and unrelatedly (so that semantic relevance can be inferred randomly) or by designing corresponding masks (so that specific semantic relevance can be emphasized).

[0082] Therefore, masking strategies for visual image videos or depth frame videos can be selected based on temporal redundancy and temporal correlation typically found in video modalities. As mentioned, Figure 10B The "pipe" concealment, or more precisely, Figure 10C The "pipe" masking may therefore be applicable to these modalities. In some embodiments, in these and other modalities, the entropy in the data can also be used to inform the masking strategy selection (e.g., as in...). Figure 12 (As discussed in the text). While in some embodiments the selected masking strategy may not change during training, in other embodiments "smart" masking may change with the training data (e.g., when motion in a given batch of visual or depth video, such as that evaluated via optical flow or temporal variations, is notified). Figure 10C When considering the length of the tube in the video). For non-video modalities, random masking (such as in...). Figure 10A In this context, although, as the reader will understand, the masking size (adjusted according to the expected modality) may initially be appropriate, it can then be intelligently changed if the result is not as desired. In a further embodiment, where, for example, available depth data is scarcer than visual intensity data and visual image videos corresponding to missing depth data are available, the sparse regions of depth data can be masked, while the corresponding regions of the visual intensity modality can be deliberately unmasked (or vice versa). In these complementary cases, the reconstruction loss of the sparse modality can be adjusted in the same way (e.g., compared to the interpolation of the missing regions, scaled down, etc.). Therefore, if the two modalities correspond in time and space, a masking strategy for a modality with a set of dimensions can be used directly or complementaryly with data or data availability for different modalities (e.g.,...). Figure 10E (The case of visual images and depth frame video modalities).

[0083] Example of strength and depth multimodal unsupervised training Many publicly disclosed embodiments consider various combinations of multimodal operating room data to capture more robust semantic representations, which can further advance downstream application-specific training. Certain cross-modal loss determinations may also be considered to better ensure the recognition of these relationships. For example, Figure 11 This is a schematic block diagram illustrating the data flow during loss computation during training in the multimodal depth and intensity unsupervised (or semi-supervised) latent space of a machine learning system 1170, which may be implemented in some embodiments. Similarly, although loss computation will be discussed here considering individual instances of the input data for the reader's understanding, the reader will realize that in practice, these instances can be merged into blocks and considered during training epochs. The system can then be updated using one or more batches of merged loss, rather than at the instance level described herein.

[0084] Here, during the latent space training (e.g., pre-training) of system 1170, instances of operating room global data tensors, including a pair of time-dependent visual intensity 1105a and depth 1110a, will be considered. The corresponding masking logics 1105c and 1110c (e.g., as referenced herein) Figures 10A-10E The described approach can be applied to each instance tensor to infer the masked counterparts 1105d and 1110d, respectively. The masked counterparts 1105d and 1110d can then be formatted, for example, by word segmentation (as previously described, this can be part of the masking operation), to form inputs 1110e and 1110f, respectively, for reception by encoder 1115, which in this example is a transducer encoder, such that they can be jointly encoded. Although in some embodiments, encoder 1115 and decoders 1145a, 1145b can be associated with them in... Figure 9 The corresponding architecture is the same or similar, but in some embodiments, the multimodal topology may use, for example, a visual transformer (ViT) infrastructure for the encoder and decoder; a fully convolutional neural network for each encoder and decoder, such as a ConvNext layer instead of a transformer; and so on. In some implementations, the visual transformer infrastructure may have 12 blocks (depth), and each block may have 12 multi-head attention layers. The decoders (one or more) may have fewer layers and are therefore “shallower” than the encoder. For example, each decoder may have four blocks, and each block may have six multi-head attention layers with 384 dimensions. The last layer of each decoder may be a single fully connected layer called a prediction head. For example, the output dimension of the prediction head of an intensity decoder may be 1536, while the output dimension of a depth decoder may be 512.

[0085] Since both intensity and depth frames can be in two-dimensional form (e.g., each intensity value is represented by a single grayscale or synthetic color, and each depth frame value is associated with a depth determination), some embodiments can provide raw values ​​to the encoder, while others can scale these values ​​to a common range (e.g., mapping each grayscale and depth value to a range between 0 and 1, where 0 is black and 1 is white, 0 is the closest recognizable depth value, and 1 is the most probable recognizable depth value, etc.). Therefore, some embodiments first use a cube embedding layer (e.g., a Conv3D layer) to convert the intensity and depth video into tokens. In some such implementations, each cube is 2 × 16 × 16 in size and corresponds to a token embedding. Masking can then be applied to these tokens, as discussed. As mentioned, the reader will be familiar with tensor transformations, where, for example, color pixel values ​​are considered. In the case of video data, each modality can have its own chunk embedding layer (e.g., a Conv3D layer). The system can then, for example, concatenate the chunk embeddings and pass them to the encoder.

[0086] A portion of the output 1120a of encoder 1115 can be designated as representing a compressed latent spatial visual intensity representation. For example, a portion of the encoder output can be designated for one modality (e.g., visual intensity video), while other portions are used for other modalities (e.g., the remainder of depth frame video), and passed to the corresponding head layer for reconstruction. Thus, in some embodiments, the latent space size can be the same for each modality (e.g., in some implementations, the latent space is 768 dimensions for both visual images and depth frame modalities). Another portion of output 1120b is here designated as representing a compressed latent spatial depth modal representation. Similar to the previously described embodiments, each of the latent spatial representations 1120a and 1120b can be passed to their respective decoders 1145a and 1145b to produce 1145c and 1145d reconstructed visual intensity 1105b and reconstructed depth 1110b representations, respectively. In particular, the latent representation from the encoder is here passed to a modality-specific decoder to reconstruct missing chunks for both intensity and depth. For clarity, as discussed elsewhere herein, not every embodiment has a separately specified decoder architecture, as in this example, but rather, for example, a single decoder with a prediction head for each modality, and the various modal latent representations are concatenated into a single tensor and passed to the single decoder architecture. As mentioned above, the prediction heads can be fully concatenated. In some embodiments, the single decoder architecture can be coupled with... Figure 11The examples use the same decoder as those discussed in this paper (e.g., one of the converter architectures) to preserve adjustments to the prediction head. These embodiments that combine decoders into a single structure can reduce the overall model size and also reduce training time.

[0087] In the depicted embodiments, the total loss 1140e used to update encoder 1115 and decoders 1145a, 1145b may include various components. Specifically, as previously described, the difference 1130b between the original intensity tensor input 1105a and the reconstructed intensity tensor output 1105b can be used to determine the visual intensity modality reconstruction loss 1140c. Similarly, the difference 1130c between the original depth tensor input 1110a and the reconstructed depth tensor output 1110b can be used to determine the depth modality reconstruction loss 1140d. However, in addition to these reconstruction losses 1140c, 1140d, the total loss 1140e may also include one or both of the matching loss 1140b and the contrast loss 1140a, or alternatively may include one or both of the matching loss 1140b and the contrast loss 1140a. Likewise, while specific data examples are discussed herein for the reader's understanding, the reader will understand that these operations (e.g., loss determination) can be performed at the batch level. Contrast loss, reconstruction loss, and matching loss can be particularly useful in enhancing cross-modal information (e.g., in...). Figure 11 It appeared in, but Figure 12 (Even further in multimodal contexts).

[0088] More specifically, regarding the visual intensity reconstruction loss 1140c, given a video clip of intensity modality V of size T×3×H×W, where T is the number of frames in the clip, H and W are the height and width of the frame, respectively, and 3 corresponds to the number of channels in the frame (e.g., where the frame depicts red, green, and blue pixel values), the loss 1140c can be determined, for example, according to Equation 1. .

[0089] (1) in p It is a word segmentation index, and It is a set of masking word segments. The reconstruction of the model corresponds to the prediction.

[0090] Similarly, regarding the depth reconstruction loss 1140d, given a video clip of depth modality D of size T×1×H×W, where T is the number of frames in the clip, H and W are the height and width of the frame, and 1 corresponds to the number of channels in the frame (here, a single distance-depth value in each entry), the loss 1140d can be determined, for example, according to Equation 2. .

[0091] (2) Where m is the word segmentation index. It is a masked word segmentation set, and The reconstruction of the model corresponds to the prediction.

[0092] If only these reconstruction losses are applied, the system may only discern the semantics of each modality individually, rather than learning any cross-modal relationships. Instead, when data from modalities correspond semantically (e.g., if they belong to the same video clip), for example, if they are considered a pair in the embedding space (e.g., their data values ​​correspond), the contrastive and matching losses described here can be used, instead of reconstruction losses or in addition to reconstruction losses, to "pull" modalities closer together. For example, intensity-depth contrastive loss learning (or between additional or alternative modalities as discussed herein) can align visual scenes and corresponding depths by "pulling" temporally paired intensity-depth data closer while "repelling" temporally unpaired intensity and depth data.

[0093] Regarding the contrastive loss 1140a, each compressed latent space representation 1120a and 1120b can be passed through a corresponding average pooling network 1125a or 1125b, respectively, to produce corresponding cumulative outputs 1125c and 1125d (e.g., corresponding averages). That is, each average pooling layer can receive encoder embeddings and output the average of the embeddings for each modality (instead of neural network layers, the reader will know that, for example, logic can be used alternatively to infer the average of the features). Since the average pooling layer is used to determine the average of the features, which is then used to infer the contrastive loss, max pooling layers or other neural networks (e.g., consisting of two or three fully connected layers mapping to a projection layer) can be used alternatively to map the latent space features to different dimensions to achieve the same or similar functional results.

[0094] Therefore, given a set of masking intensity-depth training data, various embodiments use an encoder to compute the corresponding features, followed by a global average pooling layer, and then a contrastive loss 1140a can be applied to these features, for example, as determined according to Equation 3. .

[0095] (3) in N This represents the number of instances in the batch (here, the visual intensity and depth tensor pairs). (for example, schematically represented by difference 1130a), and It's temperature. and For training instances iFor example, these correspond to encoder features for intensity and depth, respectively.

[0096] Regarding the matching loss 1140b, a fully connected multilayer perceptron (MLP) 1135a can receive each latent space representation 1120a and 1120b, respectively, to produce the matching loss 1140b. The MLP 1135a can be a fully connected single layer with an output dimension of two, followed by SoftMax as a classifier for the pair representations to predict the probabilities of the two classes. That is, the matching loss can provide a binary classification loss, indicating "positive" and "negative" modality pairs. Like an encoder and decoder, the MLP itself can be updated by the total loss 1140b (e.g., as indicated by arrow 1195). In some embodiments, the latent space representations for each modality (e.g., latent space representations 1120a and 1120b) can be concatenated before being passed to the MLP layer 1135a. As mentioned, in some implementations, the latent space can be 768 dimensions, and the MLP 1135a can correspondingly receive a single fully connected layer with 768 dimensions and output, for example, two dimensions.

[0097] As described here, using a matching loss can better facilitate cross-modal training because the matching loss can predict whether a pair of modalities is "matched" (reflected by a positive value in the example of Equation 4) or "not matched" (reflected by a negative value in the example of Equation 4). Specifically, these example embodiments reuse features from the encoder, pass them to a linear layer, and then a SoftMax classifier to solve a 2-class classification problem, thereby finding the matching loss, for example, as shown in Equation 4. .

[0098] (4) in This is a sign function; if t is 1, then output 1; otherwise, output 0. The probability score for t represents the value from SoftMax. and Let be the latent visual intensity and depth representations of the i-th pair, respectively, and M be the total number of intensity-depth instance pairs in this batch. Matching loss can be particularly useful when considering two or more data modalities, as will be referenced in this paper. Figure 12 Let's discuss this in more detail.

[0099] Therefore, in these examples, the final total loss 1140a can be represented as shown in Equation 5.

[0100] (5) in and These are hyperparameters that can be tuned during pre-training. In some implementations, the system can select from the set {0.1, 0.2, 0.3, 0.4, 0.5}. and (For example, the highest performing result is then selected at the end of training). Similarly, in some implementations, the system can select a masking ratio from the set {0.75, 0.85, 0.90, 0.95} (the numerator is the size of the masked portion of the tensor, and the denominator is the entire size of the tensor). The optimal choice of hyperparameters can be found by evaluating the pre-trained model under, for example, a 5% labeled data setting. For further clarity, in some implementations, the base learning rate is selected from {1.5e-4, 1.5e-5}.

[0101] The disclosed latent space training strategies are highly data-efficient, not only for visual and deep modalities, but also for online learning settings, such as as part of a progressive, iterative online approach, adapting deep learning systems as more data becomes available. For example, the applied neural network can be modified back into an encoder-decoder form (the same or different from the previously used one), the disclosed unsupervised training performed with newly available unlabeled data, and the application-specific supervised training performed again with any new labeled data available. Such online training (e.g., in a hospital setting) can be useful because more cases are performed, and therefore more data becomes available, facilitating continuous improvement in the understanding of the "semantic context" local to that unique deployment, while supervised training continuously improves the "application-specific foreground" adaptation, which may also be specific to that deployment environment.

[0102] Another variation of the above embodiments, in some embodiments, multimodal (e.g.) Figure 11The depth and intensity embodiments, as well as those involving additional or alternative data modalities (as discussed herein), can also be pre-trained in two or more stages (e.g., unsupervised 870 training). For example, in the first stage, one can pre-train the encoder (e.g., encoder 1115, such as the visual converter base) using the contrastive and matching losses discussed herein, but without applying any decoder or any masking (i.e., the total loss will be derived solely from the latent space representation). Following this first stage, in the second stage, the encoder can be initialized with the weights determined from the first stage (if not already present), a modality-specific decoder can be introduced (or, as discussed herein, a single merged decoder with different prediction heads for the modality), the MLP and average pooling layers can be removed, and then additional unsupervised training can be performed, but this time only using the reconstruction loss of the total loss. Masking can be employed in this second stage. In this way, reconstruction and cross-modal learning can be performed sequentially rather than simultaneously, providing specific learning pressure on their semantic content. While some embodiments may apply the first and second phases only once, the reader will appreciate that in some embodiments, the first and second phases may be applied iteratively. For example, in cases where the trainer wants to monitor the training progress of the semantic content of a particular dataset, the training of multiple copies of the training system may be performed in parallel, with the number and duration of the first and second phases varying (as the reader will appreciate, they can thus serve as training hyperparameters themselves).

[0103] Examples of different multimodal unsupervised latent space training Various embodiments can be further extended Figure 11 In some embodiments, one or more of reconstruction, matching, and contrastive losses are used when training the system, but with alternative or additional, and potentially entirely different, operating room data modalities. Figure 11 The loss, algorithm, and training strategy are not limited to Figure 11 Visual intensity and depth modalities. Modal versatility can be useful in various surgical settings, such as when considering robotic and non-robotic operating rooms (e.g., when training continuously with data from different types of operating rooms, data availability may vary).

[0104] For example, Figure 12 This is a schematic block diagram illustrating the data flow during multimodal training epochs, which may be implemented in some embodiments. That is, Figure 12 The embodiment further configures the encoder 1245 and prepares a corresponding decoder for receiving additional types of peer-to-peer operating room data modalities. Although for economic reasons, it is not included... Figure 12As shown, but as the reader will understand, each of the reconstruction loss, matching loss, and contrastive loss discussed earlier, along with various possible architectures, can be reused for three or more data modalities after necessary modifications. Therefore, training can identify cross-modal correspondences by aligning modalities by bringing semantically related modality pairs closer together within the latent feature space. For example, other modalities can be aligned with the visual intensity modality, for instance, by creating pairs such as intensity + depth, intensity + text, intensity + audio, etc., using the same contrastive loss method discussed herein. While many embodiments may employ a single encoder architecture within their topology, as discussed herein, modality-specific or modality-agnostic encoders may also be used in some embodiments. However, it has been found that, during testing, it is sufficient for various applications to use only a single shared encoder across multiple modalities.

[0105] Here, as Figure 11 In this training instance, each of the visual image intensity tensor 1205a and the temporally corresponding depth frame tensor 1230a can be considered. Corresponding masking logics 1205b and 1230b are used for each generated mask's corresponding entity 1205c and 1230c. Latent space representations 1205d and 1230d are determined via an encoder 1245 (here, a transducer encoder), and then corresponding decoders 1205e and 1230e are applied to generate reconstructed tensors 1205f and 1230f, whose corresponding losses can be used to inform the total loss, as shown in... Figure 11 The situation in the embodiments. For example... Figure 11 As described in the paper, the contrast and matching losses can also be determined between the intensity mode and each of the other modes.

[0106] Similarly, the visual intensity video tensor 1210a (or, with necessary modifications, depth values ​​acquired within the patient's body) can represent video data captured from within the patient's body (e.g., as seen wholly or partially on displays 125, 150 or within the surgeon's console 155). Appropriate corresponding masking logic 1210b can generate a masked representation 1210c, which can also be received by encoder 1245 and then reproduced as a reconstructed visual image 1210f by applying a decoder 1210e dedicated to this purpose to the latent representation 1210d. Thus, encoder 1245 can learn to generate a joint representation derived from each of the multiple data modalities.

[0107] As another example, kinematic data that also temporally corresponds to other data modalities (e.g., captured in parallel with operating room global data 1205a, 1230a, internal data 1210a, etc.) can be presented in an appropriate tensor 1215a. For example, where the kinematic data is captured as multiple waveforms over time (e.g., sampled poses of the end effector, rotational values ​​of degrees of freedom in the manipulator, etc.), tensor 1215a can be, for example, two-dimensional, with a first dimension associated with the number of waveforms and a second dimension associated with the value (or multiple values, with additional suitable dimensions) of each corresponding waveform value at each point in time. The reader will understand that interpolation, smoothing, downsampling, etc., can be applied to the kinematic data to correspond substantially concurrently with the values ​​of tensors of other equivalent data modalities. As with other modalities, tensor 1215a can also be masked via logic 1215b to form a masked representation 1215c, and a reconstructed counterpart 1215f can be generated by applying a kinematic-specific decoder 1215e to the latent space representation 1215d.

[0108] Similar to kinematic data, auditory data, such as auditory waveforms, can also be represented in tensor 1220a, similarly masked via logic 1220b to form a masked representation 1220c, and a reconstructed counterpart 1220f produced by applying an auditory-specific decoder 1220e to the latent spatial representation 1220d. While contemporaneous audio in waveform form can be represented in tensor 1220a and similarly downsampled, smoothed, interpolated, etc., to match other data modalities, in some embodiments, audio and kinematic data can be represented in alternative formats, although mapped to time points corresponding to other multimodal peer data tensors. For example, waveforms can be transformed into their frequency counterparts, and a sliding window of the dominant frequency is instead represented within the waveform. Similarly, discrete counterparts presented over time can also be determined from the waveform and interpolated into the tensor, as when using natural language processing tools to infer spoken words, spoken phrases, non-linguistic codebook values ​​(such as code-excited linear prediction (CELP) codes, code division multiple access (CDMA) values, etc.). It can also include text data that may be part of the system data. For example, if the operating room configuration, surgical procedure type, and team composition are known from the healthcare staff scheduling roster, a dictionary of these values ​​can also be included as part of a tensor (e.g., an additional tensor for operating room metadata). Similarly, medical codes, such as the International Classification of Diseases, Tenth Revision (ICD 10) code, can appear in such metadata tensors or in “text” tensors, and can also be temporarily masked.

[0109] By maintaining such discrete values ​​over time intervals of other multimodal data equivalent tensors, this type of data can be made suitable for parallel consideration of the encoder system. In fact, system event data (e.g., tool activation, operator head removal from the surgical console, the state of the robotic surgical system, the state of the surgeon's console, etc.), although often presented with discrete values ​​over time, can be organized into tensor 1225a by extending their values ​​over time if not presented in their original form. Similar to waveforms, tensor 1225a can be, for example, two-dimensional, with one dimension associated with each type of system event data and another dimension associated with the value of the event data. The corresponding masking logic 1225b can prepare a masked counterpart 1225c, from which the encoder 1245 can infer a latent space representation 1225d suitable for the decoder 1225e to form the reconstructed tensor 1225f.

[0110] Readers will see that even though latent space training can utilize different types of data itself during operation, such as Figure 3 Such pre-training can still be applied to non-surgical applications, such as... Figure 4 And vice versa, provided that the relevant semantic associations with the application under consideration have been roughly captured. For example, kinematic data could initially help an aid system identify semantic associations between the movement of a cart and movement in operating room-wide surgical data during pre-training. Once this is recognized, application-specific training might then strengthen these associations, but now only operating room-wide surgical data is considered specifically. Conversely, a system capable of identifying such associations will be better positioned to anticipate the associations expected from operating room-wide data when trained solely from operating room-wide data. Figure 3 Further correlation of kinematic data obtained from additional data modes. The reader will see that fewer than... Figure 12 All modalities described herein, or additional or alternative modalities (such as ellipses 1235a-f indicating; for example, text, as discussed, such as the phrase "operating room staff are preparing or cleaning the room," medical codes and descriptions, identified colloquial terms, etc.). Similarly, though... Figure 12 The reconstruction of only one instance is shown, but these instances can be considered in batches to infer the combined loss value.

[0111] like Figure 11 Similar to Figure 12Implementations can determine the loss based solely on the difference between the original and reconstructed tensors, solely on contrastive loss (e.g., summing between each modality pair), solely on matching loss (e.g., summing between each modality pair), or a combination of these types of losses. Indeed, in some cases, and for certain downstream applications, the presence of multiple modalities may render any single type of component loss sufficient for training. However, combining loss types, as in Equation 5 (adjusted according to additional modalities), can facilitate more nuanced semantic determination and better facilitate training for a wider range of downstream applications.

[0112] The disclosed different encoder and decoder architectures or merged autoencoder embodiments can therefore leverage multiple modalities (visual intensity, depth values, text, kinematic data, etc.) to extract discriminative representations for semantic understanding. The resulting latent space representation can narrow the gap with subsequent supervised training, thereby reducing the need for acquiring large amounts of labeled data. Unlike methods limited to a single modality, the various disclosed multimodal approaches are better suited to a holistic consideration of the operating room state in order to better understand semantic relevance that arises within and between modalities. Thus, by employing multiple data modalities, some embodiments are used to pre-train models for many different downstream tasks, such as scene understanding, audio classification, dense depth estimation, visual question answering, cross-modal (e.g., video-text, audio-text) retrieval, and even in some zero-shot settings.

[0113] Example of unsupervised training loss determination Figure 13 The flowchart illustrates various operations in an example process 1300, which may be implemented in some embodiments, for performing a round of loss determination on one or more data modalities during a round of training. Specifically, at block 1305a, the system may receive or select data specified for this round of latent space training. As previously mentioned, this data may include one or more modalities from operating room data that does not have operating room global data, operating room global data only, or operating room data that includes operating room global data (naturally, it will be understood that the disclosed methods can be applied to robotic or non-robotic operating rooms, and in fact, latent space training from data from one type of operating room can be used for another type of application in some cases).

[0114] At blocks 1305b and 1305c, the system can iterate over various envisioned modalities, determine the corresponding masking strategy at block 1305d, and generate the masked representation of the instance at block 1305e (again, as discussed, this can be performed during word segmentation). For example, as regarding... Figure 12The different masking logics discussed (e.g., logics 1205b, 1210b, 1215b, 1220b, 1225b, and 1230b) can be applied to each of the corresponding modalities of the data received at block 1305a, and the masking strategy applied to the modality can be changed during training (e.g., in a way that complements the changes in masking of other modalities).

[0115] At block 1310a, the system can construct an input tensor (e.g., a corresponding word segmentation representation) to receive the masked representation generated at block 1305e at the encoder (e.g., encoder 1115, encoder 1245, etc.), and obtain one or more reduced representations at block 1310b by applying the tensor to the encoder. Each of encoders 1115 and 1245 can receive the input tensor, for example, via a single convolutional layer.

[0116] As mentioned, while some embodiments may determine only some component losses, here the system can determine each of the contrast and match losses as well as the reconstruction loss. According to Equations 3 and 4, the corresponding component losses can be determined at the batch level using the latent space results of specific instances of these modalities. Therefore, at block 1310c, the system can record the latent space derivation results of the considered multimodal instances to be used for the contrast loss, and at block 1310d, the system can record the latent space derivation results of the considered multimodal instances to be used for the match loss (again, for clarity, the reader will understand that the operations do not necessarily need to be performed in the depicted chronological order, nor need to be divided into the depicted tissue blocks).

[0117] At block 1310e, the system can generate a reconstructed representation for each modality, for example, such as Figure 11 and Figure 12 As described in [the document]. At blocks 1315a and 1315b, the system can iterate again on the envisioned modes, and at block 1315c determine the corresponding reconstruction loss. In example procedure 1300, these results can be similarly recorded to determine the total loss at the batch level.

[0118] When all instances in the current batch (e.g., a set of temporally corresponding data for different data modalities) have been considered at block 1320 (otherwise, the process can proceed to the next instance), the total loss for the batch can be determined and applied to various neural networks (e.g., as discussed above). Specifically, when all batch instances have been considered at block 1320, the system can determine the total loss, for example, according to Equation 5. The system can, for example, sum the reconstruction losses determined for each modality and each instance at block 1315c to determine the ensemble reconstruction loss for the batch at block 1325a. Similarly, at block 1325b, the system can use the composition results determined for each instance at block 1310c to determine the contrastive loss for the batch. At block 1325c, the system can use the composition results determined for each instance at block 1310d to determine the matching loss for the batch. Finally, for example, according to Equation 5, the system can determine the total loss for the batch at block 1325d, and then use the loss at block 1325e to update various machine learning systems as described herein (e.g., encoder 1115; encoder 1245; each of decoders 1205e, 1210e, 1215e, 1220e, 1225e and 1230e; decoders 1145a and 1145b; MLP 1135a; etc.).

[0119] Example of downstream application training process Figure 14 This is a flowchart illustrating various operations in an example process 1400 for preparing an application of a specific machine learning system, which may be performed in some embodiments. At block 1405a, the training computer system may receive partially labeled operating room data. At block 1405b, the system may determine which data to use for supervised or unsupervised training (if not specified by a human designer), such as distinguishing between labeled and unlabeled data. For example, the system may identify data manually labeled by human annotators, along with data 375a about the surgical state obtained from surgical equipment and other contextual data or metadata, to infer labels. Depending on the heterogeneous nature of the data, in some embodiments, labels may be removed from the labeled data (or ignored) to provide more unsupervised training data, for example, creating a more balanced supervised and unsupervised training dataset.

[0120] Then, at block 1405c, the data for unsupervised (or, with necessary modifications, semi-supervised) training can be divided into training blocks. At block 1405d, the system can perform image processing on these blocks or batches. For example, as described herein, RGB and grayscale visual image pixel values ​​can be mapped to a common range with depth distance values, corrupted data can be removed (and the supervised and unsupervised training datasets adjusted accordingly), downsampling or interpolation upsampling can be performed, etc. At block 1405e, the system or human designer can select appropriate hyperparameters (e.g., weights of the total loss, masking strategy, training epochs, backpropagation settings, etc.) for unsupervised training (e.g., experience-derived lookup tables, interpolation from historical selections based on the properties of the provided operating room dataset, the goals or requirements of downstream applications, parallel operation of different training systems, etc.).

[0121] The system can then perform multiple rounds of initial latent space training at block 1410a to generate a loss score, which is used to update one or more encoders, one or more decoders, and any other relevant machine learning systems. However, if all rounds of training are completed, or if the performance of one or more encoders / decoders is satisfactory at block 1410b, the system can perform an adaptation process at block 1420a (e.g., according to arrow 840a).

[0122] However, before the initial latent space training is complete, during each training round, the system can iterate over data blocks (also known as batches) of unsupervised data, for example, at blocks 1410c and 1410d, and then iterate over each instance of the data block at blocks 1410e and 1415a. At block 1415b, an appropriate masking representation for the data modality can be determined based on the currently applicable masking strategy for the corresponding modality. Using the masking representation, at block 1415c, the system can determine various components of the loss derived from the currently considered instance (e.g., the latent space average pooling value for each modality, the MLP result for each modality pair in the instance, the reconstruction loss for each modality, etc.).

[0123] Once all instances of the block have been considered, the total loss for the block (e.g., total loss 1140e) (e.g., from cumulative reconstruction losses 1140c, 1140d, matching loss 1140b, and contrastive loss 1140a) can then be determined at block 1415d and used to update various machine learning systems (e.g., one or more encoders, one or more decoders, MLPs, etc.) at block 1415e.

[0124] Once the initial latent space training is complete, portions of one or both of the encoders and decoders can be used at block 1420a to prepare an architecture suitable for the application, thus corresponding, for example, to adaptation 830. For example, a new output layer can be attached to the combined encoder / decoder system (e.g., as part of modification 830b), attached only to the encoder, and so on. This can produce a structure suitable for training downstream applications (e.g., supervised training) (e.g., the architecture of classifier 825b) (e.g., as shown by arrow 840b).

[0125] Therefore, at block 1420b, the system can divide the application data (e.g., labeled data) into batches for application-specific training, and at block 1420c, select appropriate hyperparameters for the dedicated training (again, the reader will understand that application training may not be performed exactly as described here for ease of understanding). Then, application-specific training can be performed (e.g., identifying from labeled operating room global training data). Figure 6A or Figure 6B The system continues in its current state until all training epochs are completed at block 1425a, or until the performance of the machine learning system is acceptable at block 1425b. After this, the application-specific machine learning system can be released for use at block 1435a. For each epoch of application-specific training, the system can consider the application data blocks at blocks 1425c and 1425d, and for each block, consider the corresponding data instances at blocks 1425e and 1430a to determine the current loss of the machine learning system. At block 1430b, the application-specific machine learning system is adjusted based on those results, for example, as discussed in 825 regarding supervised training (naturally, one will know that different application-specific training methods may not be performed exactly as shown here, but may have different loss determination methods, some updated by instance, some by batch, some only updated throughout the epoch, etc.). Therefore, it will be understood that the resulting machine learning system does not need to output the probability of discrete categories, but can be any suitable downstream application, such as generative applications or object recognition and visual image segmentation applications, whose output may not be a single prediction set, but rather, for example, probability tensors, synthetic data values, projected data values, etc.

[0126] Example prototype of the embodiment Implementation The result Figure 15 A table illustrating the comparative results generated by the example prototype implementations of the embodiments. Specifically, it describes the example multimodal data implementations of the disclosed embodiments (using visual and depth data, such as...). Figure 11This is compared with state-of-the-art techniques. Here, "MICCAI 2022-SwAV (Intensity + Depth)" refers to the method proposed by Jamal, Muhammad Abdullah, and Omid Mohareri in the arXiv preprint "Multi-Modal Unsupervised Pre-Training for Surgical Operating Room Workflow Analysis" (arXiv:2207.07894 (2022) using visual intensity images and depth frame modalities). Here, "SurgMAE" refers to the method proposed by Jamal, Muhammad Abdullah, and Omid Mohareri in the arXiv preprint "SurgMAE: Masked Autoencoders for Long Surgical Video Analysis" (arXiv:2305.11451 (2023)). Results are presented in... Figure 15 The table illustrates the superior performance of the example implementations, as demonstrated by the average accuracy values ​​in each table cell of the rightmost three columns. Specifically, as shown in the figure, the corresponding systems were evaluated based on low data condition settings (e.g., 5%, 10%, and 20% of labeled data accessible to downstream tasks). For all experiments, the pre-trained model was fine-tuned for 75 epochs using a base learning rate of 6e-4. A cosine learning scheduler with a termination learning rate of 1e-5 was used, and the fine-tuning was warmed up for 5 epochs with a learning rate of 1e-8. For models configured to use time data, a learning rate of 1e-3 was used for a total of 15 epochs.

[0127] Computer System Figure 16 This is a block diagram of an example computer system that can be used in conjunction with some embodiments. The computing system 1600 may include an interconnect 1605 that connects several components, such as, for example, one or more processors 1610, one or more memory components 1615, one or more input / output systems 1620, one or more storage systems 1625, one or more network adapters 1630, etc. The interconnect 1605 may be, for example, one or more bridges, traces, buses (e.g., ISA, SCSI, PCI, I2C, FireWire bus, etc.), wires, adapters, or controllers.

[0128] One or more processors 1610 may be included, such as Intel. TMProcessor chips, math coprocessors, graphics processors, etc. One or more memory components 1615 may include, for example, volatile memory (RAM, SRAM, DRAM, etc.), non-volatile memory (EPROM, ROM, flash memory, etc.), or similar devices. One or more input / output devices 1620 may include, for example, display devices, keyboards, pointing devices, touch screen devices, etc. One or more storage devices 1625 may include, for example, cloud-based storage, removable universal serial bus (USB) storage, disk drives, etc. In some systems, memory component 1615 and storage device 1625 may be the same component. Network adapter 1630 may include, for example, a wired network interface, a wireless interface, Bluetooth, etc. TM Adapters, line-of-sight interfaces, etc.

[0129] People will recognize that, in some embodiments, only Figure 16 Some of the components, alternative components, or additional components described herein. Similarly, in some systems, these components may be combined or used for dual purposes. These components may be implemented using dedicated hardwired circuit systems, such as one or more ASICs, PLDs, FPGAs, etc. Therefore, some embodiments may be implemented in programmable circuit systems (e.g., one or more microprocessors) programmed, for example, with software and / or firmware, or in systems entirely in dedicated hardwired (non-programmable) circuitry, or in combinations thereof.

[0130] In some embodiments, data structures and message structures may be stored or transmitted via network adapter 1630, via a data transmission medium, such as signals on a communication link. Transmission can be performed via various media, such as the Internet, local area network, wide area network, or point-to-point dial-up connection. Therefore, "computer-readable medium" can include computer-readable storage media (e.g., "non-transitory" computer-readable media) and computer-readable transmission media.

[0131] One or more memory components 1615 and one or more storage devices 1625 may be computer-readable storage media. In some embodiments, one or more memory components 1615 or one or more storage devices 1625 may store instructions that can perform or cause the performance of various operations discussed herein. In some embodiments, the instructions stored in memory 1615 may be implemented as software and / or firmware. These instructions can be used to perform operations on one or more processors 1610 to perform the processes described herein. In some embodiments, such instructions may be provided to one or more processors 1610 by, for example, downloading instructions from another system via network adapter 1630.

[0132] For clarity, it will be understood that although a computer system can be a single machine residing in a single location, it has... Figure 16 One or more components, but not necessarily. For example, a distributed networked computer system may include multiple individual processing workstations, each with... Figure 16 Some or all of the components described herein. Therefore, the processes and various operations described herein can be distributed across one or more workstations of such a computer system. For example, it will be appreciated that a process suitable for running in a single thread on a single workstation can instead be divided into any number of sub-threads on one or more workstations, which then run serially or in parallel to achieve the same or substantially similar results as a process running in a single thread. Similarly, it will be appreciated that while non-transitory computer-readable media can be standalone (e.g., in a single USB storage device) or reside in a single workstation (e.g., in the workstation's random access memory or disk storage), such media does not need to reside in a single geographic location, but can include, for example, multiple memory storage units residing on geographically separated workstations of a computer system that communicates with each other via a network, or residing on geographically separated storage devices.

[0133] Notes The accompanying drawings and descriptions are illustrative. Therefore, neither the specification nor the drawings should be construed as limiting this disclosure. For example, headings or subheadings are merely for convenience and to aid understanding. Therefore, headings or subheadings should not be construed as limiting the scope of this disclosure, such as by grouping features presented in a particular order, or simply by presenting them together for the purpose of aiding understanding. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. In case of conflict, this document shall prevail, including any definitions provided herein. The use of one or more synonyms herein does not preclude the use of other synonyms. The use of examples anywhere in this specification, including examples of any terms discussed herein, is illustrative only and is not intended to further limit the scope and meaning of this disclosure or any exemplary terms.

[0134] Similarly, although specific representations are made in the figures herein, those skilled in the art will recognize that the actual data structures used to store information can differ from those shown. For example, data structures may be organized differently, may contain more or less information than shown, may be compressed and / or encrypted, and so on. To avoid confusion, common or well-known details may be omitted from the figures and disclosures. Similarly, the figures may depict specific sets of operations to aid understanding, which are merely examples of a broader category of such sets of operations. Thus, it will be readily apparent that additional, alternative, or fewer operations can often be used to achieve the same purpose or effect described in some flowcharts. For example, data may be encrypted, although not so shown in the figures, but items may be considered different loop patterns (“For” loops, “while” loops, etc.) or stored differently to achieve the same or similar effects, etc.

[0135] The terms "an embodiment" or "one embodiment" as used herein refer to at least one embodiment of this disclosure that includes a particular feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrase "in one embodiment" in various places herein does not necessarily refer to the same embodiment in each of those places. Separate or alternative embodiments may not be mutually exclusive with other embodiments. It will be appreciated that various modifications can be made without departing from the scope of the embodiments.

Claims

1. A method for preparing a machine learning system configured to receive operating room data, the method comprising: A first training session is performed on the machine learning system using a first set of operating room data, the machine learning system comprising: The first part is configured to create a latent spatial representation from operating room data input; and The second part is configured to create a reconstructed representation of the operating room data from the potential spatial representation.

2. The method according to claim 1, wherein the operating room data input includes: Data input for the first operating room in the first mode; as well as Second modality second operating room data input.

3. The method according to claim 2, wherein the method further comprises: Mask at least a portion of the data input from the first operating room; as well as At least a portion of the data input to the second operating room is masked.

4. The method according to claim 3, wherein, The first modality is the full-area visual intensity image and video data of the operating room, and wherein, The second modality is the full-area depth frame video data of the operating room.

5. The method according to claim 3, wherein, Each of the first mode and the second mode is different from one of the following modes: Comprehensive depth data of the entire operating room; Operating room full-area visual intensity data; Operating room kinematic data: Operating room auditory data; Operating room system event data; Operating room text data; Operating room visual intensity data, which depicts the interior of the patient; and Operating room depth data, which depicts the interior of the patient.

6. The method according to claim 3, wherein, Masking at least the portion of the first operating room data input generates a first masking representation, wherein, The first masking representation includes a first tensor, wherein at least one dimension of the first tensor corresponds to time, wherein... Masking at least the portion of the second operating room data input generates a second masking representation, and wherein, The second masking representation includes a second tensor, at least one dimension of which corresponds to time.

7. The method of claim 6, wherein performing the first training session comprises: The first masking representation is provided as input to the machine learning system; as well as The second masking representation is provided as input to the machine learning system.

8. The method according to claim 7, wherein, The masking of at least the portion of the first operating room data input includes masking at least one complete temporal data frame of the first tensor, and wherein, The masking of the second operating room data input includes at least one complete time data frame of the masking of the second tensor.

9. The method according to claim 7, wherein, The masking of at least the portion of the first operating room data input includes masking the same first spatial portion of the first tensor across consecutive temporal data frames of the first tensor, and wherein, The masking of at least the portion of the second operating room data input includes masking the same second spatial portion of the second tensor across consecutive temporal data frames of the second tensor.

10. The method according to claim 9, wherein, The first spatial portion and the second spatial portion are the same spatial portion.

11. The method according to claim 3, wherein the method comprises: Modify the machine learning system using one or more neural network layers; as well as Perform a second training session on the modified machine learning system.

12. The method of claim 11, wherein the second training session is for an application involving the identification of situational states in an operating room.

13. The method of claim 11, wherein the second training session is for an application involving the detection of objects appearing in the operating room.

14. The method according to any one of claims 2 to 13, wherein, The first part of the machine learning system includes an encoder neural network, and wherein, The second part of the machine learning system includes a decoder neural network.

15. The method of claim 14, wherein the first part of the machine learning system and the second part of the machine learning system are portions of the same autoencoder neural network.

16. The method of claim 14, wherein the method further comprises: Determine the total loss used to update the machine learning system, wherein determining the total loss includes determining: The first difference between the first operating room input and the first reconstructed operating room data instance; and The second difference between the second operating room input and the second reconstructed operating room data instance.

17. The method of claim 16, wherein determining the total loss further comprises: The contrast loss is determined at least in part based on the differences identified from the following: The first potential spatial representation of the data input from the first operating room; as well as The second potential space representation of the second operating room data input.

18. The method of claim 17, wherein determining the contrast loss comprises: Determine a first average value of the first potential space representation; Determine the second average value of the second latent space representation; as well as The contrast loss is determined based on the first average value and the second average value.

19. The method of claim 18, wherein determining the total loss further comprises: Determine the matching loss, wherein determining the matching loss includes: Determine whether the first latent space representation and the second latent space representation correspond to a binary indicator.

20. The method of claim 19, wherein determining the binary indication comprises: The SoftMax classifier is applied, and the SoftMax classifier itself is updated at least in part based on the total loss.

21. A non-transitory computer-readable medium including instructions configured to cause one or more computer systems to perform a method for preparing a machine learning system configured to receive operating room data, the method comprising: A first training session is performed on the machine learning system using a first set of operating room data, the machine learning system comprising: The first part is configured to create a latent spatial representation from operating room data input; and The second part is configured to create a reconstructed representation of the operating room data from the potential spatial representation.

22. The non-transitory computer-readable medium of claim 21, wherein the operating room data input comprises: Data input for the first operating room in the first mode; as well as Second modality second operating room data input.

23. The non-transitory computer-readable medium of claim 22, wherein the method further comprises: Mask at least a portion of the data input from the first operating room; as well as At least a portion of the data input to the second operating room is masked.

24. The non-transitory computer-readable medium according to claim 23, wherein, The first modality is the full-area visual intensity image and video data of the operating room, and wherein, The second modality is the full-area depth frame video data of the operating room.

25. The non-transitory computer-readable medium according to claim 23, wherein, Each of the first mode and the second mode is different from one of the following modes: Comprehensive depth data of the entire operating room; Operating room full-area visual intensity data; Operating room kinematic data: Operating room auditory data; Operating room system event data; Operating room text data; Operating room visual intensity data, which depicts the interior of the patient; and Operating room depth data, which depicts the interior of the patient.

26. The non-transitory computer-readable medium according to claim 23, wherein, Masking at least the portion of the first operating room data input generates a first masking representation, wherein, The first masking representation includes a first tensor, wherein at least one dimension of the first tensor corresponds to time, wherein... Masking at least the portion of the second operating room data input generates a second masking representation, and wherein, The second masking representation includes a second tensor, at least one dimension of which corresponds to time.

27. The non-transitory computer-readable medium of claim 26, wherein performing the first training session comprises: The first masking representation is provided as input to the machine learning system; as well as The second masking representation is provided as input to the machine learning system.

28. The non-transitory computer-readable medium according to claim 27, wherein, The masking of at least the portion of the first operating room data input includes masking at least one complete temporal data frame of the first tensor, and wherein, The masking of the second operating room data input includes at least one complete time data frame of the masking of the second tensor.

29. The non-transitory computer-readable medium according to claim 27, wherein, The masking of at least the portion of the first operating room data input includes masking the same first spatial portion of the first tensor across consecutive temporal data frames of the first tensor, and wherein, The masking of at least the portion of the second operating room data input includes masking the same second spatial portion of the second tensor across consecutive temporal data frames of the second tensor.

30. The non-transitory computer-readable medium according to claim 29, wherein, The first spatial portion and the second spatial portion are the same spatial portion.

31. The non-transitory computer-readable medium of claim 23, wherein the method comprises: Modify the machine learning system using one or more neural network layers; as well as Perform a second training session on the modified machine learning system.

32. The non-transitory computer-readable medium of claim 31, wherein the second training session is for an application involving the identification of situational states in an operating room.

33. The non-transitory computer-readable medium of claim 31, wherein the second training session is for an application involving the detection of objects appearing in an operating room.

34. The non-transitory computer-readable medium according to any one of claims 22 to 33, wherein, The first part of the machine learning system includes an encoder neural network, and wherein, The second part of the machine learning system includes a decoder neural network.

35. The non-transitory computer-readable medium of claim 34, wherein the first portion of the machine learning system and the second portion of the machine learning system are portions of the same autoencoder neural network.

36. The non-transitory computer-readable medium of claim 34, wherein the method further comprises: Determine the total loss used to update the machine learning system, wherein determining the total loss includes determining: The first difference between the first operating room input and the first reconstructed operating room data instance; and The second difference between the second operating room input and the second reconstructed operating room data instance.

37. The non-transitory computer-readable medium of claim 36, wherein determining the total loss further comprises: The contrast loss is determined at least in part based on the differences identified from the following: The first potential spatial representation of the data input from the first operating room; as well as The second potential space representation of the second operating room data input.

38. The non-transitory computer-readable medium of claim 37, wherein determining the contrast loss comprises: Determine a first average value of the first potential space representation; Determine the second average value of the second latent space representation; as well as The contrast loss is determined based on the first average value and the second average value.

39. The non-transitory computer-readable medium of claim 38, wherein determining the total loss further comprises: Determine the matching loss, wherein determining the matching loss includes: Determine whether the first latent space representation and the second latent space representation correspond to a binary indicator.

40. The non-transitory computer-readable medium of claim 39, wherein determining the binary indication comprises: The SoftMax classifier is applied, and the SoftMax classifier itself is updated at least in part based on the total loss.

41. A computer system, the computer system comprising: At least one processor; as well as At least one memory, the at least one memory including instructions configured to cause the computer system to perform a method for preparing a machine learning system configured to receive operating room data, the method comprising: A first training session is performed on the machine learning system using a first set of operating room data, the machine learning system comprising: The first part is configured to create a latent spatial representation from operating room data input; and The second part is configured to create a reconstructed representation of the operating room data from the potential spatial representation.

42. The computer system of claim 41, wherein the operating room data input includes: Data input for the first operating room in the first mode; as well as Second modality second operating room data input.

43. The computer system of claim 42, wherein the method further comprises: Mask at least a portion of the data input from the first operating room; as well as At least a portion of the data input to the second operating room is masked.

44. The computer system according to claim 43, wherein, The first modality is the full-area visual intensity image and video data of the operating room, and wherein, The second modality is the full-area depth frame video data of the operating room.

45. The computer system according to claim 43, wherein, Each of the first mode and the second mode is different from one of the following modes: Comprehensive depth data of the entire operating room; Operating room full-area visual intensity data; Operating room kinematic data: Operating room auditory data; Operating room system event data; Operating room text data; Operating room visual intensity data, which depicts the interior of the patient; and Operating room depth data, which depicts the interior of the patient.

46. ​​The computer system according to claim 43, wherein, Masking at least the portion of the first operating room data input generates a first masking representation, wherein, The first masking representation includes a first tensor, wherein at least one dimension of the first tensor corresponds to time, wherein... Masking at least the portion of the second operating room data input generates a second masking representation, and wherein, The second masking representation includes a second tensor, at least one dimension of which corresponds to time.

47. The computer system of claim 46, wherein performing the first training session comprises: The first masking representation is provided as input to the machine learning system; as well as The second masking representation is provided as input to the machine learning system.

48. The computer system according to claim 47, wherein, The masking of at least the portion of the first operating room data input includes masking at least one complete temporal data frame of the first tensor, and wherein, The masking of the second operating room data input includes at least one complete time data frame of the masking of the second tensor.

49. The computer system according to claim 47, wherein, The masking of at least the portion of the first operating room data input includes masking the same first spatial portion of the first tensor across consecutive temporal data frames of the first tensor, and wherein, The masking of at least the portion of the second operating room data input includes masking the same second spatial portion of the second tensor across consecutive temporal data frames of the second tensor.

50. The computer system according to claim 49, wherein, The first spatial portion and the second spatial portion are the same spatial portion.

51. The computer system according to claim 43, wherein the method comprises: Modify the machine learning system using one or more neural network layers; as well as Perform a second training session on the modified machine learning system.

52. The computer system of claim 51, wherein the second training session is for an application involving the identification of situational states in an operating room.

53. The computer system of claim 51, wherein the second training session is for an application involving the detection of objects appearing in the operating room.

54. The computer system according to any one of claims 42 to 53, wherein, The first part of the machine learning system includes an encoder neural network, and wherein, The second part of the machine learning system includes a decoder neural network.

55. The computer system of claim 54, wherein the first part of the machine learning system and the second part of the machine learning system are portions of the same autoencoder neural network.

56. The computer system of claim 54, wherein the method further comprises: Determine the total loss used to update the machine learning system, wherein determining the total loss includes determining: The first difference between the first operating room input and the first reconstructed operating room data instance; and The second difference between the second operating room input and the second reconstructed operating room data instance.

57. The computer system of claim 56, wherein determining the total loss further comprises: The contrast loss is determined at least in part based on the differences identified from the following: The first potential spatial representation of the data input from the first operating room; as well as The second potential space representation of the second operating room data input.

58. The computer system of claim 57, wherein determining the contrast loss comprises: Determine a first average value of the first potential space representation; Determine the second average value of the second latent space representation; as well as The contrast loss is determined based on the first average and the second average.

59. The computer system of claim 58, wherein determining the total loss further comprises: Determine the matching loss, wherein determining the matching loss includes: Determine whether the first latent space representation and the second latent space representation correspond to a binary indicator.

60. The computer system of claim 59, wherein determining the binary indication comprises: The SoftMax classifier is applied, and the SoftMax classifier itself is updated at least in part based on the total loss.