Computer vision architecture to identify and monitor sterility states in medical procedures

The computer vision architecture generates a 3D representation of medical environments to monitor sterility by classifying and tracking objects, addressing the challenge of maintaining sterile fields and reducing contamination risks.

WO2025184384A1PCT designated stage Publication Date: 2025-09-04INTUITIVE SURGICAL OPERATIONS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/017661
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-01
Filing Date
2025-02-27
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

In medical environments, objects such as people, furniture, and medical instruments can obscure or be obscured, making it difficult to effectively monitor and maintain sterility during procedures, which can lead to potential breaches and complications.

Method used

A computer vision architecture that uses a cascading rendering architecture to generate a 3D representation of the medical environment, classifies objects as sterile or non-sterile, and monitors their movement to detect potential sterility breaches, providing real-time alerts and visual indications.

Benefits of technology

Enhances sterility monitoring by minimizing contamination risks through real-time detection and alerting, reducing the risk of surgical site infections by maintaining sterile fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025017661_04092025_PF_FP_ABST
    Figure US2025017661_04092025_PF_FP_ABST
Patent Text Reader

Abstract

Aspects of this technical solution can include identifying, using a first model configured to detect image features, an object within an image of a medical environment, classifying, using a second model, the object to associate a sterility status with the object, the sterility status indicative of a first state of the object as sterile or a second state of the object as non-sterile, monitoring, using the first model, movement of the object within the medical environment, determining, based on the movement of the object within the medical environment, a metric indicative of a risk of breach of sterility of the object, and generating, in response to the metric meeting a threshold indicative of the second state of the object as non-sterile, an output in real¬ time indicative of the metric meeting the threshold.
Need to check novelty before this filing date? Find Prior Art

Description

COMPUTER VISION ARCHITECTURE TO IDENTIFY AND MONITOR STERILITY STATES IN MEDICAL PROCEDURESTECHNICAL FIELD

[0001] The present implementations relate generally to medical devices, including but not limited to a computer vision architecture to identify and monitor sterility states in medical environments and / or during medical procedures.INTRODUCTION

[0002] Awareness of objects, people, and movements associated with a medical procedure is crucial to effective and efficient completion of the medical procedure. A medical environment can include a medical environment with a large number of people, pieces of furniture, medical instruments, and medical devices. Each of these people, pieces of furniture, medical instruments, and medical devices can potentially obscure or be obscured by others of the people, pieces of furniture, medical instruments, and medical devices in the medical environment. Effective support of the medical procedure during the medical procedure and evaluation of the medical procedure subsequent to the medical procedure can be negatively impacted or otherwise impossible to perform where people, pieces of furniture, medical instruments, or medical devices are obscured or not visible.SUMMARY

[0003] Systems, methods, apparatuses, and non-transitory computer-readable media are provided for predicting and indicating sterility states of one or more objects in a sterile field. For example, a sterile field is a physical volume in a medical environment arranged or prepared for a given medical procedure. For example, a system can execute a model that receives input from various sensors located in a medical environment, and can identify and track one or more objects and / or spaces in the medical environment. Objects, as discussed herein, can include, but are not limited to, people, furniture, medical tools, trays, or surgical robots. Sensors can include, but are not limited to, cameras configured to capture images or video of the medical environment, or any portion thereof. With this input, the model can identify one or more objects as being sterile or non-sterile, and can determine an actual breach of sterility or a risk of breach of sterility for particular objects in the medical environment. Thus, a technical solution for a computer vision architecture to identify and monitor sterility states in medical procedures is provided.

[0004] At least one aspect is directed to a method. The method can include identifying, using a first model configured to detect image features, an object within an image of a medical environment. The method can include classifying, using a second model, the object to associate a sterility status with the object, the sterility status indicative of a first state of the object as sterile or a second state of the object as non-sterile. The method can include monitoring, using the first model, movement of the object within the medical environment. The method can include determining, based on the movement of the object within the medical environment, a metric indicative of a risk of breach of sterility of the object. The method can include generating, in response to the metric meeting a threshold indicative of the second state of the object as non-sterile, an output in real-time indicative of the metric meeting the threshold.

[0005] At least one aspect is directed to a system. The system can include a memory and one or more processors. The system can identify, using a first model configured to detect image features, an object within an image of a medical environment. The system can classify, using a second model, the object to associate a sterility status with the object, the sterility status indicative of a first state of the object as sterile or a second state of the object as non-sterile. The system can monitor, using the first model, movement of the object within the medical environment. The system can determine, based on the movement of the object within the medical environment, a metric indicative of a risk of breach of sterility of the object. The system can generate, in response to the metric meeting a threshold indicative of the second state of the object as non-sterile, an output in real-time indicative of the metric meeting the threshold.

[0006] At least one aspect is directed to a non-transitory computer readable medium can include one or more instructions stored thereon and executable by a processor. The processor can identify, using a first model configured to detect image features, an object within an image of a medical environment. The processor can classify, using a second model, the object to associate a sterility status with the object, the sterility status indicative of a first state of the object as sterile or a second state of the object as non-sterile. The processor can monitor, using the first model, movement of the object within the medical environment. The processor can determine, based on the movement of the object within the medical environment, a metric indicative of a risk of breach of sterility of the object. The processor can generate, in response to the metric meeting a threshold indicative of the second state of the object as non-sterile, an output in real-time indicative of the metric meeting the threshold.BRIEF DESCRIPTION OF THE FIGURES

[0007] These and other aspects and features of the present implementations are depicted by way of example in the figures discussed herein. Present implementations can be directed to, but are not limited to, examples depicted in the figures discussed herein. Thus, this disclosure is not limited to any figure or portion thereof depicted or referenced herein, or any aspect described herein with respect to any figures depicted or referenced herein.

[0008] FIG. 1 A depicts an example architecture of a system according to this disclosure.

[0009] FIG. IB depicts an example environment of a system according to this disclosure.

[0010] FIG. 2A depicts an example model in first rendering iteration according to this disclosure.

[0011] FIG. 2B depicts an example model in second rendering iteration according to this disclosure.

[0012] FIG. 2C depicts an example model in final rendering iteration according to this disclosure.

[0013] FIG. 3 depicts an example motion estimation architecture according to this disclosure.

[0014] FIG. 4 depicts an example first sterility presentation according to this disclosure.

[0015] FIG. 5 depicts an example second sterility presentation according to this disclosure.

[0016] FIG. 6 depicts an example layer model architecture according to this disclosure.

[0017] FIG. 7 depicts an example user interface with sterile person presentation according to this disclosure.

[0018] FIG. 8 depicts an example user interface with sterile object presentation according to this disclosure.

[0019] FIG. 9 depicts an example user interface for sterility training according to this disclosure.

[0020] FIG. 10 depicts an example method of a computer vision architecture to identify and monitor sterility states in medical procedures, according to this disclosure.

[0021] FIG. 11 depicts an example method of a computer vision architecture to identify and monitor sterility states in medical procedures, according to this disclosure.DETAILED DESCRIPTION

[0022] Aspects of this technical solution are described herein with reference to the figures, which are illustrative examples of this technical solution. The figures and examples below are not meant to limit the scope of this technical solution to the present implementations or to a single implementation, and other implementations in accordance with present implementations are possible, for example, by way of interchange of some or all of the described or illustrated elements. Where certain elements of the present implementations can be partially or fully implemented using known components, only those portions of such known components that are necessary for an understanding of the present implementations are described, and detailed descriptions of other portions of such known components are omitted to not obscure the present implementations. Terms in the specification and claims are to be ascribed no uncommon or special meaning unless explicitly set forth herein. Further, this technical solution and the present implementations encompass present and future known equivalents to the known components referred to herein by way of description, illustration, or example.

[0023] A system can generate a 3D representation of the medical environment based on input data including a plurality of 2D images of the medical environment that each capture the medical environment from differing viewpoints. For example, the system can include a cascading rendering architecture that iteratively generates the 3D representation of the medical environment. For example, at a first stage of an object segmentation engine, the system can receive a first image of the medical environment from a first viewpoint as a first prompt to a 2D-to-3D transformer (e.g., a NeRF engine). The 2D-to-3D transformer can then generate a 3D representation of an image, and a mask engine can generate a mask of the 3D representation that corresponds to an object in the medical environment. The system can provide the mask as input to a second stage of the object segmentation engine. The second stage can obtain the mask as a second prompt to the 2D-to-3D transformer. The second stage can obtain a second image of the medical environment from a second viewpoint as part of the second prompt to the 2D-to-3D transformer. The system can generate or update the mask based on the second prompt to expand on the 3D representation corresponding to the object, or to refine the mask based on the second viewpoint. The object segmentation engine is not limited to two stages as discussedherein, and can include any number of stages to generate a 3D mask of a given object in the medical environment.

[0024] In one aspect, the classification model can be trained to classify identified objects as sterile or non-sterile (e.g., a sterility status). The sterility status of an object can be determined by the model based on one or more types of input data from one or more data sources. These data sources can include, but are not limited to, for example, visual indicia, distances between various objects (e.g., determined using the 3D segmentation model), relative positions of objects within the medical environment, system data of one or more components of robotic devices associated with the medical environment, motion data from medical environment sensors, procedure type data, or 3D room layout data. System data can include event data or kinematics data from a medical platform or medical device (e.g., computer-assisted medical system, robotic surgical system, etc.). The system as discussed herein can thus determine sterility status based at least on any of the above-noted features, and by tracking of these objects over time in the medical environment. For example, the object’s visual indicia, position within the medical environment at a given time or over time through timestamps in video input data, position relative to other identified objects in the medical environment, etc. For example, the system can detect, by the classification model, that a person is wearing scrubs of a given color, and can determine, by a machine learning model configured to determine a sterility state for one or more objects, that the person is sterile by identifying that the given color indicates a sterile person. The machine learning model can also identify one or more objects as non-sterile, based on particular visual indicia or lack of any indicia of sterility. For example, the machine learning model can track movement of a plurality of objects over time, and can change state of an object in the medical environment from sterile to non-sterile, due to contact with non-sterile objects, proximity with non-sterile objects, or movement that exceeds a threshold for an economy of motion indicative of a sterile object.

[0025] In certain implementations, the machine learning model can be used to provide realtime or near real-time indications or warnings relating to potential sterility breaches in the medical environment. In addition, or as an alternative, the model can be used to generate objective performance indicators (OPIs) relating to performance of medical staff in the medical environment, operations of a robotic medical device (e.g., surgical robot), and effectiveness of layout of the medical environment in maintaining sterile fields within the medical environment.

[0026] For instance, the model can monitor the identified objects as they move or are moved within the medical environment to determine either of the actual breach or risk of breach on a per-object basis and provide indications of an actual breach or risk of breach concurrently or in real-time with the occurrence of the physical states that are linked with an actual breach or risk of breach. The technical solution can implement and monitor a sterile field according to rigorous guidelines to provide information that medical environment supervisors, risk management, and surgical team members can use in the development and implementation of policies and procedures for establishing the sterile field in the medical environment. The model can incorporate temporal information collected from input data (e.g., case video) to determine sterility status of various objects on an object-by-object basis. For example, an instrument might be sterile before the case begins and non-sterile after being used, even though an appearance of the instrument remains materially the same throughout the case. Thus, this technical solution can provide a technical improvement to increase sterility of a medical environment by providing real-time monitoring and alerts for sterility for each sterile object associated with a given sterile field.

[0027] In some embodiments, the model can identify whether an actual breach of sterility or a risk of breach of sterility occurs for a given object, based on one or more of a position of the object, a motion of the object, or an aggregate of motion of a plurality of objects. For example, the model can determine a position of a given person with respect to the sterile field, including whether the person has left the sterile field or is at a position at a perimeter of the sterile field that risks contamination. The system can determine position of a person within the medical environment, and can also determine a position of one or more limbs of the person, with respect to one or more sterile objects or non-sterile objects. For example, the model can determine a movement of a given person with respect to the sterile field, including whether the person has left the sterile field or is at a position at a perimeter of the sterile field that risks contamination. The model can determine one or more of a motion of an object or a portion of the object, and displacement of the object or the portion of the object. For example, the model can determine that a person is moving in the medical environment at a speed exceeding an economy of motion threshold that maintains sterility of the sterile field or the person in the sterile field. For example, the model can determine that a person is moving through a medical environment by an amount exceeding an economy of motion threshold that maintains sterility of the sterile field or the person in the sterile field. Here, a person could be walking within the sterile field a distance that exceeds an economy of motion parameter, or could be moving one or more limbsat a speed that exceeds the economy of motion parameter. The economy of motion parameter can include one or more sub-parameters to detect one or more of the above-noted positions or movements, and is not limited thereto.

[0028] Thus, the technical solution discussed herein can provide technical improvements including monitoring and alerting to minimize traffic and movement in a sterile field, to keep air movement to a minimum, thus reducing the microorganisms and other contaminants, such as dust and debris, from entering the atmosphere. This can reduce the patient’s risk for acquiring a surgical site infection (SSI). The system can provide various outputs that indicate whether an actual sterility breach has occurred or whether a risk of a sterility breach is increased, with respect to a given object. For example, the system can provide a visual indication on a video stream, to annotate a particular person who is sterile. For example, an annotation can include a color overlay over a portion of a video corresponding to a sterile person. The annotation can include one or more of text and color indicative of a breach, a type of breach, or a risk of breach.

[0029] FIG. 1A depicts an example architecture of a system according to this disclosure. As illustrated by way of example in FIG. 1 A, an architecture of a system 100 A can include at least a data processing system 110, communication bus 120, and a robotic system 130.

[0030] In some embodiments, the system 100 A can configure multiple sensors in the medical environment based on detection of s state or scene corresponding to the medical environment as a whole. The system can detect, for example, a robot docking scene as discussed above, and can configure multiple sensors in the medical environment according to the field of view of the sensor or the location of the sensor in the medical environment. For example, the system can enter a training mode in which a model is trained with machine learning from input including video from a plurality of camera sensors distributed within the medical environment. The machine learning model can optimize the allocation of resources to each of the cameras during corresponding modes, using a loss function that is based on at least one of the video data, parameters assigned to each sensor, positions assigned to each sensor, compute resources allocated to each sensor, or energy resources allocated to each sensor. This way, the machine learning model can be updated (e.g., trained) to provide a technical improvement to optimize sensor placement and resource consumption on an individualized basis for each sensor, in view of a state of the medical environment as a whole. The system can combine input from a plurality of sensors to provide “fused” sensor input to the machine learning model. For example, thesystem can designate two sensors placed near each other and having a field of view oriented toward an operating table as a stereoscopic fused sensor input. For example, the system can designate two sensors placed near each other and having a field of view oriented toward a medical environment as a panoramic fused sensor input. The machine learning model can treat video input from each of these fused sensors as a combined input for determining optimized allocation or a loss. This way, the machine learning model can be updated (e.g., trained) to provide a technical improvement to optimize sensor placement and resource consumption on a combined basis for sensors having coordinated roles, to increase optimization and efficiency of resource allocation beyond that provided by individualized loss determinations for all sensors.

[0031] The data processing system 110 can include a physical computer system operatively coupled or that can be coupled with one or more components of the system 100 A, either directly or directly through an intermediate computing device or system. The data processing system 110 can include a virtual computing system, an operating system, and a communication bus to effect communication and processing. The data processing system 110 can include a system processor 112 and a system memory 114.

[0032] The system processor 112 can execute one or more instructions associated with the system 100A. The system processor 112 can include an electronic processor, an integrated circuit, or the like including one or more of digital logic, analog logic, digital sensors, analog sensors, communication buses, volatile memory, nonvolatile memory, and the like. The system processor 112 can include, but is not limited to, at least one microcontroller unit (MCU), microprocessor unit (MPU), central processing unit (CPU), graphics processing unit (GPU), physics processing unit (PPU), embedded controller (EC), or the like. The system processor 112 can include a memory operable to store or storing one or more instructions for operating components of the system processor 112 and operating components operably coupled to the system processor 112. The one or more instructions can include at least one of firmware, software, hardware, operating systems, embedded operating systems, and the like. The system processor 112 can include at least one communication bus controller to effect communication between the system processor 112 and the other elements of the system 100.

[0033] The system memory 114 can store data associated with the data processing system 110. The system memory 114 can include one or more hardware memory devices to store binary data, digital data, or the like. The system memory 114 can include one or more electricalcomponents, electronic components, programmable electronic components, reprogrammable electronic components, integrated circuits, semiconductor devices, flip flops, arithmetic units, or the like. The system memory 114 can include at least one of a non-volatile memory device, a solid-state memory device, a flash memory device, or a NAND memory device. The system memory 114 can include one or more addressable memory regions disposed on one or more physical memory arrays. A physical memory array can include a NAND gate array disposed on, for example, at least one of a particular semiconductor device, integrated circuit device, and printed circuit board device. The system memory 114 can correspond to a non-transitory computer readable medium. For example, the non-transitory computer readable medium can include one or more instructions executable by a processor. The processor can generate, based on a second image corresponding to a second two-dimensional view of the object, the three- dimensional object model corresponding to the object from the view and the second view. The processor can cause a user interface to present the output can include a region of the image having a visual property based on the metric, the region of the image defined by the object model, and the visual property corresponding to a color based on the metric.

[0034] The communication bus 120 can communicatively couple the data processing system 110 with the robotic system 130. The communication bus 120 can communicate one or more instructions, signals, conditions, states, or the like between one or more of the data processing system 110 and components, devices, blocks operatively coupled or couplable therewith. The communication bus 120 can include one or more digital, analog, or like communication channels, lines, traces, or the like. As one example, the communication bus 120 can include at least one serial or parallel communication line among multiple communication lines of a communication interface. The communication bus 120 can include one or more wireless communication devices, systems, protocols, interfaces, or the like. The communication bus 120 can include one or more logical or electronic devices including but not limited to integrated circuits, logic gates, flip flops, gate arrays, programmable gate arrays, and the like. The communication bus 120 can include one or more telecommunication devices including but not limited to antennas, transceivers, packetizers, and wired interface ports.

[0035] The robotic system 130 can include one or more robotic devices configured to perform one or more actions of a medical procedure (e.g., a surgical procedure). For example, a robotic device can include, but is not limited to a surgical device that can be manipulated by robotic device. For example, a surgical device can include, but is not limited to, a scalpel or acauterizing tool. The robotic system 130 can include various motors, actuators, or electronic devices whose position or configuration can be modified according to input at one or more robotic interfaces. For example, a robotic interface can include a manipulator with one or more levers, buttons, or grasping controls that can be manipulated by pressure or gestures from one or more hands, arms, fingers, or feet. The robotic system 130 can include a surgeon console in which the surgeon can be positioned (e.g., standing or seated) to operate the robotic system 130. However, the robotic system 130 is not limited to a surgeon console co-located or on-site with the robotic system 130.

[0036] In some examples, the robotic system 130 can provide its robotic system data, including kinematics data and system events data, to the data processing system 110. The robotic system data can be used in combination with other data described herein by the data processing system 110 to determine one or more metrics (e.g., aggregate or individual metrics) described herein. The kinematics data describe kinematics of robotic manipulator(s) of the robotic system 130). For example, the kinematics data can indicate configuration(s) of one or more manipulators or manipulator assemblies of the robotic system 130 as well as the instruments attached or operatively coupled to the manipulators or manipulator assemblies over time throughout the medical procedure.

[0037] The system events data can be generated by the robotic system 130 and can indicate system events of the robotic system 130. Examples of system events can include, for example, a docking event (e.g., in which manipulator arms are docked to cannulas inserted into a patient anatomy), operator (e.g., surgeon) head-in or head-out event (e.g., indicating a surgeon’s head being present or absent at a viewer on a input or control console of the robotic system), an instrument attachment or removal event (e.g., indicating attachment or removal of an instrument, such as a medical instrument or an imaging instrument, on a manipulator of the robotic system, a tool exchange event), an instrument change event (e.g., indicating performance of an exchange of one instrument for another instrument for attachment on a manipulator on the robotic system), a draping-start event or a sterile adapter attachment event (e.g., which may indicate beginning of a sterile draping process), and the like. The robotic system 130 can include one or more sensors (e.g., camera, infrared sensor, ultrasonic sensors, etc.), actuators, interfaces, consoles, that can output information used to detect such a system event. Although system events such as instrument attachment or removal event and instrument change event can explicitly inform the installation or removal of an instrument, which may besubjected to SP, other system events can also provide implicit indication of the installation, use, removal of an instrument. Such robotic system data can be in natural language (e.g., a string of characters).

[0038] In other words, the robotic system data can be indicative of one or more states of one or more components of the robotic system 130. Components of the robotic system 130 can include, but are not limited to, actuators on the robotic system and instruments coupled or attached to the actuators of the robotic system. For example, the robotic system data can include one or more data points indicative of one or more of an activation state (e.g., activated or deactivated), a position, or orientation of a component of the robotic system 130. For example, the robotic system data can be linked with or correlated with one or more medical procedures, one or more phases of a given medical procedure, one or more tasks of a given phase of a given medical procedure, and steps of sterile processing. For example, a type of robotic system data can be indicative of a position of one or more actuators of a given arm (or an instrument attached thereto) of the robotic system 130 at a given time or over a given time interval. Such time interval can be associated with a given part, step, or sub-step of a sterile processing procedure. In some examples, the robotic system data can include a timeline based on timestamps of system events that are time-aligned.

[0039] FIG. IB depicts an example environment of a system according to this disclosure. As illustrated by way of example in FIG. IB, an environment 100B of a system 100A can include at least the robotic system 130 having a field of view 132, a first sensor system 140, a second sensor system 150, persons 160, and objects 170. For example, the environment 100B is illustrated by way of example as a plan view of a medical environment having the robotic system 130, the first sensor system 140, the second sensor system 150, the persons 160, and the objects 170 disposed therein or thereabout. The presence, placement, orientation, and configuration, for example, of one or more of the robotic system 130, the first sensor system 140, the second sensor system 150, the persons 160, and the objects 170 can correspond to a given medical procedure or given type of medical procedure that is being performed, is to be performed, or can be performed in the medical environment corresponding to the environment 100B. This disclosure is not limited to the presence, placement, orientation, or configuration of the robotic system 130, the first sensor system 140, the second sensor system 150, the persons 160, or the objects 170, or any other element, illustrated herein by way of example. The field of view 132 of the robotic system 130 can correspond to a physical volume withinthe environment 100B that is within range of detection of one or more sensors of the robotic system 130. For example, the field of view 132 is positioned above a surgical site of a patient. For example, the field of view 132 is oriented toward a surgical site of a patient.

[0040] The first sensor system 140 can include one or more sensors oriented to a first portion of the environment 100B. For example, the first sensor system 140 can include one or more cameras configured to capture images or video in visual or near-visual spectra and / or one or more depth-acquiring sensors for capturing depth data (e.g., three-dimensional point cloud data). For example, the first sensor system 140 can include a plurality of cameras configures to collectively capture images or video in a stereoscopic view. For example, the first sensor system 140 can include a plurality of cameras configures to collectively capture images or video in a panoramic view. The first sensor system 140 can include a field of view 142. The field of view 142 can correspond to a physical volume within the environment 100B that is within range of detection of one or more sensors of the first sensor system 140. For example, the field of view 142 is oriented toward a surgical site of a patient. For example, the field of view 152 is located behind a surgeon at surgical site of a patient.

[0041] The second sensor system 150 can include one or more sensors oriented to a second portion of the environment 100B. For example, the second sensor system 150 can include one or more cameras configured to capture images or video in visual or near-visual spectra and / or one or more depth-acquiring sensors for capturing depth data (e.g., three-dimensional point cloud data). For example, the second sensor system 150 can include a plurality of cameras configures to collectively capture images or video in a stereoscopic view. For example, the second sensor system 150 can include a plurality of cameras configures to collectively capture images or video in a panoramic view. The second sensor system 150 can include a field of view 152. The field of view 152 can correspond to a physical volume within the environment 100B that is within range of detection of one or more sensors of the second sensor system 150. For example, the field of view 152 is oriented toward the robotic system 130. For example, the field of view 152 is located adjacent to the robotic system 130.

[0042] The persons 160 can include one or more individuals present in the environment 100B. For example, the persons can include, but are not limited to, assisting surgeons, supervising surgeons, specialists, nurses, or any combination thereof. The objects 170 can include, but are not limited to, one or more pieces of furniture, instruments, or any combination thereof. For example, the objects 170 can includes tables and surgical instruments.

[0043] In some implementations, the environment 100B includes one or more audio sensors (e.g., microphones), e.g., on the ceiling, on the robotic system 130, on the system 140, on the system 150, and so on to record real-time audio data in the environment 100B. The audio data can be used in combination with other data described herein by the data processing system 110 to determine one or more metrics (e.g., aggregate or individual metrics) described herein.

[0044] FIG. 2A depicts an example model in first rendering iteration according to this disclosure. As illustrated by way of example in FIG. 2A, a model in first rendering iteration 200 A can include at least a physical environment 210, a first mask processor 220, and a first 3D object model 230. For example, the system can generate, based on an image corresponding to a two-dimensional view of the object, a three-dimensional object model corresponding to the object from the view.

[0045] The physical environment 210 can correspond at least partially in one or more of structure and operation to a depiction of the environment 100B, or a portion thereof, but is not limited thereto. The physical environment 210 can include a first viewpoint 212. The first viewpoint 212 can correspond to a field of view within the environment 100B, but is not limited thereto. For example, the first viewpoint 212 can correspond to the field of view 142, and can at least partially depict the one or more persons 160 or objects 170 present at the central area of the environment 100B from the first viewpoint 212.

[0046] The first mask processor 220 can execute a model configured to transform at least a portion of a 2D image into a 3D object having a structure corresponding to an object in the 2D image. For example, the first mask processor 220 can identify a portion of an object in the physical environment image 210 that is depicted (e.g., visible and not occluded) in the physical environment image 210 from the first viewpoint 212. The first viewpoint 212 can isolate the object that is depicted (or a plurality of objects). The first mask processor 220 can receive as input a first 2D object image 222, and provide as output a first 2D object mask 224. The first 2D object image 222 can depict at least a portion of the physical environment image 210 from the first viewpoint 212. For example, the first mask processor 220 can depict a portion of the environment 100B from the first sensor system 140. For example, the portion of the environment 100B can correspond to a central area of the environment 100B facing a patient site. The first mask processor 220 can depict one or more persons 160 or objects 170 present at the central area of the environment 100B. The first 2D object mask 224 can correspond to a mapping of one or more points or portions of a surface onto the portion of the first 2D objectimage 222 corresponding to the object depicted in the first 2D object image 222. For example, the first mask processor 220 can generate a mask to identify one or more positions of a surface of the object in the first 2D object image 222. The first mask processor 220 can correlate the positions in a 2D coordinate space for the 2D object image 222 to a 3D coordinate space.

[0047] The first 3D object model 230 can correspond to a 3D object having at least one of a shape, volume, position, or orientation in the coordinate space that corresponds the shape, volume, position, or orientation of the depiction of the object from the first viewpoint 212. The first mask processor 220 can generate the 3D object with respect to a coordinate system corresponding to at least one of the environment 100B or the first viewpoint 212. The first partial 3D object 232 can correspond to a 3D structure for the portion of the object in the physical environment 210 that is visible and not occluded in the first 2D object image 222. For example, the first partial 3D object 232 is a 3D object that is a partial reconstruction of the object depicted in the first 2D object image 222 from only the first viewpoint 212. The first 3D object model 230 can include a first partial 3D object 232, and a first 3D view direction 234. The first 3D view direction 234 can correspond to a direction of the first viewpoint 212. For example, the first 3D view direction 234 is a first vector having an origin at the first viewpoint 212 (e.g., view of first sensor system 140) and extending from a central axis or direction of gaze from the first viewpoint 212.

[0048] FIG. 2B depicts an example model in second rendering iteration according to this disclosure. As illustrated by way of example in FIG. 2B, a model in second rendering iteration 200B can include at least a second viewpoint 214, a second mask processor 240, and a second 3D object model 250. The second viewpoint 214 can correspond to a field of view within the environment 100B that is distinct from the first viewpoint 212, but is not limited thereto. For example, the second viewpoint 214 can correspond to the field of view 152, and can at least partially depict the one or more persons 160 or objects 170 present at the central area of the environment 100B from the second viewpoint 214. For example, the system can generate, based on a second image corresponding to a second two-dimensional view of the object, the three-dimensional object model corresponding to the object from the view and the second view.

[0049] The second mask processor 240 can correspond at least partially in one or more of structure and operation to the first mask processor 220. For example, the second mask processor 240 can be integrated with the first mask processor 220, and is not limited to a distinct structure as discussed herein by way of example. The second mask processor 240 can receiveas input a second 2D object image 242, and provide as output a second 2D object mask 244. The second 2D object image 242 can depict at least a portion of the physical environment 210 from the second viewpoint 214. For example, the second mask processor 240 can depict the portion of the environment 100B from the second sensor 150. The first mask processor 220 can depict the one or more persons 160 or objects 170 present at the central area of the environment 100B from a distinct viewpoint to capture features of various objects that are occluded or not visible in the depiction of the first 2D object image 222 from the first viewpoint 212. The second 2D object image 242 can include a second view portion 246. The second view portion 246 can correspond to a portion of the object in the environment 100B that is depicted in the second 2D object image 242 but not visible or occluded in the first object image 222. Thus, the second 2D object image 242 can incorporate 3D model or mask information from at least one of the first 2D object mask 224 or the first 3D object model 230 to augment a depiction of the object. The augmented depiction can render visible one or more portions of the object that are not visible in or are occluded in the first 2D object image 222. The second 2D object mask 244 can correspond to a mapping of one or more points or portions of a surface onto the portion of the second 2D object image 242 corresponding to the object depicted in the first 2D object image 222 and the second 2D object image 242. For example, the second mask processor 240 can generate a mask to identify one or more positions of a surface of the object in the second 2D object image 242. The second mask processor 240 can import the positions from the first 2D object mask 224 via the first 3D object model 230, and correlate the positions in a 2D coordinate space for the second 2D object image 242 to a 3D coordinate space.

[0050] The second 3D object model 250 can correspond to a 3D object having at least one of a shape, volume, position, or orientation in the coordinate space that corresponds the shape, volume, position, or orientation of the depiction of the object from the first viewpoint 212 and from the second viewpoint 214. The second mask processor 240 can generate the 3D object with respect to the coordinate system corresponding to at least one of the environment 100B or the second viewpoint 214. The second 3D object model 250 can include a second partial 3D object 252, and a second 3D view direction 254. The second partial 3D object 252 can correspond to a 3D structure for the portion of the object in the physical environment 210 that is visible and not occluded in at least one of the first 2D object image 222 or the second 2D object image 242. For example, the second partial 3D object 252 is a 3D object that is a partial reconstruction of the object depicted in the first 2D object image 222 from the first viewpoint 212, and the second 2D object image 242 from the second viewpoint 214. The second 3D viewdirection 254 can correspond to a direction of the second viewpoint 214. For example, the second 3D view direction 254 is a second vector having an origin at the second viewpoint 214 (e.g., view of second sensor system 150) and extending from a central axis or direction of gaze from the second viewpoint 214.

[0051] FIG. 2C depicts an example model in final rendering iteration according to this disclosure. As illustrated by way of example in FIG. 2C, a model in final rendering iteration 200C can include at least a final viewpoint 216, a final mask processor 260, and a final 3D object model 270. For example, the final rendering iteration 200C can correspond to an iteration after an arbitrary number of preceding iterations subsequent to the second rendering iteration 200B. For example, the final rendering iteration 200C can correspond to an nth iteration that results in a 3D object model that is complete. For example, a complete object model has no portions that are occluded, or has incorporated viewpoints from all sensors that depict a given object in a given environment. The final viewpoint 216 can correspond to a field of view within the environment 100B that is distinct from the first viewpoint 212 and the second viewpoint 214, but is not limited thereto. For example, the final viewpoint 216 can correspond to the field of view 132 of a camera of the robotic system 130, and can at least partially depict the one or more persons 160 or objects 170 present at the central area of the environment 100B from the second viewpoint 214.

[0052] The final mask processor 260 can correspond at least partially in one or more of structure and operation to the first mask processor 220. For example, the final mask processor 260 can be integrated with the first mask processor 220, and is not limited to a distinct structure as discussed herein by way of example. The final mask processor 260 can receive as input a final 2D object image 262, and provide as output a final 2D object mask 264. The final 2D object image 262 can depict at least a portion of the physical environment 210 from the final viewpoint 216. For example, the final mask processor 260 can depict the portion of the environment 100B from the sensor of the robotic system 130. The final mask processor 260 can depict the one or more persons 160 or objects 170 present at the central area of the environment 100B from a distinct viewpoint to capture features of various objects that are occluded or not visible in the depiction of the first 2D object image 222 from the first viewpoint 212, or the second 2D object image 242 from the second viewpoint 214. The final 2D object mask 264 can correspond to a mapping of one or more points or portions of a surface onto the portion of the second 2D object image 242 corresponding to the object depicted in the first 2Dobject image 222, the second 2D object image 242, and the final 2D object image 262. For example, the final mask processor 260 can generate a mask to identify one or more positions of a surface of the object in the final 2D object image 262. The final mask processor 260 can import the positions from the first 2D object mask 224 and the positions from the second 2D object mask 244 via the second 3D object model 250, and correlate the positions in a 2D coordinate space for the second 2D object image 242 to a 3D coordinate space.

[0053] The final 3D object model 270 can correspond to a 3D object having at least one of a shape, volume, position, or orientation in the coordinate space that corresponds the shape, volume, position, or orientation of the depiction of the object from the first viewpoint 212, the second viewpoint 214, and the final viewpoint 216. The final mask processor 260 can generate the 3D object with respect to the coordinate system corresponding to at least one of the environment 100B or the final viewpoint 216. The final 3D object model 270 can include a final 3D object 272. The final 3D object 272 can correspond to a 3D structure for the portion of the object in the physical environment 210 that is visible and not occluded in at least one of the first 2D object image 222, the second 2D object image 242, or the final 2D object image 262. For example, the final 3D object 272 is a 3D object that is a complete reconstruction of the object depicted in the first 2D object image 222 from the first viewpoint 212, the second 2D object image 242 from the second viewpoint 214, and the final 2D object image 262 from the final viewpoint 216. Thus, the final 3D object model 270 can correspond to a specific instance or type of person 160 or object 170 across one or more field of view or viewpoints as discussed herein. For example, the object corresponding to the final 3D object model 270 can include people, furniture, and robots. The system can apply a second machine learning model to each of the objects identified in the environment 100B via the iterative process of 200 A-C, to associate sterility metrics with one or more of the persons 160 or the objects 170.

[0054] FIG. 3 depicts an example motion estimation architecture according to this disclosure. As illustrated by way of example in FIG. 3, a motion estimation architecture 300 can include at least an environment sensor system 310, a pose estimation system 320, and a motion estimation system 330.

[0055] The environment sensor system 310 can provide instructions to one or more sensor systems according to or in response to one or more metrics corresponding to the robotic system 130, the environment 100B, or a medical procedure of the environment 100B, or any combination thereof. For example, the environment sensor system 310 can include one or morelogical or electronic devices including but not limited to integrated circuits, logic gates, flip flops, gate arrays, programmable gate arrays, and the like. One or more electrical, electronic, or like devices, or components associated with the environment sensor system 310 can also be associated with, integrated with, integrable with, replaced by, supplemented by, complemented by, or the like, the data processing system 110 or any component thereof. The environment sensor system 310 can include the first model configured to detect image features. For example, the first model can execute the architecture according to the iterations of 200A-C to generate a 3D object model.

[0056] The environment sensor system 310 can resample the environment 100B at a given frequency (e.g., 24 Hz to 60Hz) to monitor one or more persons 160 and objects 170 in the environment 100B, and to detect any changes in the environment 100B (including the changes of the positions and orientations of the persons 160 and the objects 170 in the environment 100B over time). For example, the system 310 can monitor, using the first model, a position and an orientation of the object within the medical environment (e.g., the environment 100B). The system 310 can determine, based on the position and orientation of the object within the medical environment, the metric corresponding to the object.

[0057] The pose estimation system 320 can execute a model to determine properties of one or more persons 160 in the environment 100B. For example, the pose estimation system 320 can receive one or more 3D objects from the environment sensor system 310. For example, the pose estimation system 320 can receive the final 3D object model 270 from the environment sensor system 310. The pose estimation system 320 can execute the model to determine one or more metrics indicative of orientation of one or more limbs of one or more persons 160 at one or more times. For example, the pose estimation system 320 can generate a wireframe for a torso and one or more limbs for a person 160 in the environment 100B, and can determine positions of each limb at one or more times. Thus, the pose estimation system 320 can track limb position and movement of one or more persons 160. Collectively a pose can correspond to an arrangement of one or more limbs and a torso of a person 160.

[0058] The motion estimation system 330 can execute a model to determine one or more metrics indicative of economy of motion of one or more of the persons 160 or the objects 170 in the environment 100B. For example, an economy of motion can correspond to an aggregate displacement of one or more limbs of one or more people within the environment 100B. For example, an economy of motion can be used to determine the risk of sterility breach, due tomotions that can displace volumes of air and transmit contaminants suspending in air onto surfaces of various persons 160 or objects 170 in the environment 100B, or any surface in the environment 100B. This air flow (e.g., doors open / closed, fast movement in sterile field) can increase risk of break of sterility of one or more surfaces in the environment 100B. Thus, the more aggregate displacement of the persons 160 or objects 170 in the environment 100B, the greater the risk of sterility breach. For example, the second machine learning model can apply different thresholds of movement that exceed economy of motion for people and for objects (e.g., medical tool, instrument, furniture). For example, the second machine learning model can apply a first threshold is applied to something in the video or image identified by the first machine learning model as a person. For example, the second machine learning model can apply a second threshold, indicating an amount of motion lower than an amount of motion indicated by the first threshold, to something in the video or image identified by the first machine learning model as furniture. For example, the system can determine, using the first model, the movement of the object within the medical environment. The system can determine, using the second model, the metric. Here, the metric can correspond to a metric indicative of economy of motion.

[0059] For example, a threshold corresponding to an economy of motion metric can be indicative of a specific clinical constraint (e.g., maximum aggregate motion of a person or persons to maintain sterility for a given medical procedure), which may differ for different types of medical procedures. For example, a threshold may be associated with a lower value (e.g., corresponding to a lower maximum aggregate motion) for an open surgery, and may be associated with a higher value (e.g., corresponding to a higher maximum aggregate motion) for a laparoscopic surgery. An economy of motion metric crossing over a threshold can indicate sterilization breach or a greater risk of sterilization breach. For example, economy of motion metric for medical staff member being greater than a threshold indicating the maximum allowable aggregate motion for a person can indicate sterilization breach or a greater risk of sterilization breach, whereas economy of motion metric for medical staff member being less than a threshold indicating the maximum allowable aggregate motion for a person can indicate no sterilization breach or a lesser risk of sterilization breach.

[0060] The motion estimation system 330 can determine a displacement estimation metric indicative of distances between people and specific sterile / non-sterile objects, and can also be used determine a risk of breach of sterility. In some examples, the displacement estimationmetric represents the proximity between a person and an object, between two objects, or between two people, where the closer proximity indicates a higher risk ot breach of sterility. The displacement estimation metric can be or can be determined based on an average displacement in term of pixels (in the image or video) between a person and an object, between two objects, or between two people. Thus, the motion estimation system 330 can determine, in addition to whether motion by the medical staff member is sufficiently great, whether the medical staff member is sufficiently near to a sterile object or sterile object or site to cause breach of sterility, at a level of accuracy beyond the capability of manual processes (largely dependent on human observation) to achieve.

[0061] For example, a threshold corresponding to a displacement estimation metric can be indicative of a specific clinical constraint (e.g., minimum distance between persons / objects to maintain sterility for a given medical procedure), which may differ for different object types. For example, a threshold may be associated with a lower value (e.g., corresponding to a lower minimum distance) for a distance between a sterile and a non-sterile object or between a person and a sterile object, and may be associated with a higher value (e.g., corresponding to a higher minimum distance) for a distance between two sterile objects, between two non-sterile objects, or between two people. A displacement estimation metric crossing over a threshold can indicate sterilization breach or a greater risk of sterilization breach. For example, a displacement estimation metric for a distance between the medical staff member and a sterile object being less than a threshold indicating the minimum allowable distance can indicate sterilization breach or a greater risk of sterilization breach, whereas a displacement estimation metric for a distance between the medical staff member and the sterile object being greater than the threshold indicating the minimum allowable distance can indicate no sterilization breach or a less risk of sterilization breach.

[0062] In some implementation, the risk of breach of sterility can be evaluated based on a time of exposure. The time dimension can be obtained by analyzing the video (e.g., visual videos, depth videos, and so on) collected using the sensor systems described herein, including the sensor systems 140, 150, a camera of the robotic system 130, and so on, based on frame-by- frame analysis. For example, a long exposure of a sterile object to a non-sterile object or a person may significantly increase the risk of breach of sterility, and on the other hand, the impact on sterility can be low due to a short exposure of sterile object to a non-sterile object or a person. Specifically, a greater aggregate motion of a person over a long period of time maysignificantly increase the risk of breach of sterility, and on the other hand, the impact on sterility can be low due to the same aggregate motion of a person over a very short period of time. Similarly, a displacement estimation metric for a distance between the medical staff member and a sterile object being less than a threshold indicating the minimum allowable distance over a longer period of time can indicate sterilization breach or a greater risk of sterilization breach, whereas a displacement estimation metric for the same distance between the medical staff member and the sterile object over a very short period of time can indicate no sterilization breach or a less risk of sterilization breach. In some examples, a metric, e.g., the economy of motion metric and the displacement estimation metric, can be a time-based metric determined using the metric as described herein and a time factor. Any suitable function can be used to determine the time-based metric, e.g., the time-based metric equals to the metric multiplied by a time factor, where the time factor is or is determined based at least in part on the time period by which the metric is valid.

[0063] In some examples, the threshold (e.g., an economy of motion metric threshold or a displacement estimation metric threshold) can be evaluated over time. For example, an economy of motion metric crossing over a metric threshold for a period of time over a maximum time threshold can indicate sterilization breach or a greater risk of sterilization breach. For example, economy of motion metric for medical staff member being greater than a threshold indicating the maximum allowable aggregate motion for a person, for a period of time over a time threshold can indicate sterilization breach or a greater risk of sterilization breach, whereas the economy of motion metric for medical staff member being greater than the threshold, for a period of time less a minimum time threshold can indicate no sterilization breach or a lesser risk of sterilization breach. The maximum time threshold and the minimum time threshold can be the same or different thresholds.

[0064] For example, a displacement estimation metric crossing over a metric threshold for a period of time over a maximum time threshold can indicate sterilization breach or a greater risk of sterilization breach. For example, a displacement estimation metric for a distance between the medical staff member and a sterile object being less than a threshold indicating the minimum allowable distance for a period of time over a maximum time threshold can indicate sterilization breach or a greater risk of sterilization breach, whereas the displacement estimation metric for the same distance between the medical staff member and a sterile object being less than the threshold for a period of time less than a minimum time threshold can indicate nosterilization breach or a less risk of sterilization breach. The maximum time threshold and the minimum time threshold can be the same or different thresholds.

[0065] In some implementations, whether a sterilization breach has occurred can be outputted (e.g., at 1120), in the form of texts, graphics, colors, visual overlay, visual alarms, audio alarms, tactile alarms, etc. indicating or explaining the sterilization breach. In some examples multiple levels (e.g., three or more levels) of sterilization breach can be defined by the data processing system 110 and outputted (e.g., visual, audio, tactile alarms), indicating different levels or risks of sterilization breach. In some examples, different thresholds can be used to define different ranges, each range corresponding to a level of sterilization breach (or a risk thereof), and the different levels define a highest risk or seriousness of sterilization breach, a lowest risk or seriousness of sterilization breach, and at least one level therebetween. A metric disclosed herein that is within a range is assigned a corresponding level of sterilization breach. In some examples, the different thresholds can be defined for an aggregate metric, which can be the aggregate weighted sum or another combination of different metrics (time-based and otherwise) disclosed herein, such as the economy of motion metric, the displacement estimation metric, metric indicative of risk of breach, metric based on movement of objects, and so on.

[0066] In some examples, different thresholds can be defined for each of one or more different metrics disclosed herein, such as the economy of motion metric, the displacement estimation metric, metric indicative of risk of breach, metric based on movement of objects, and so on, to define ranges for multiple levels (e.g., three or more levels) of sterilization breach. For example, multiple economy of motion metric thresholds can be defined, including the maximum allowable aggregate motion threshold and two or more additional aggregate motion thresholds indicating successively lower aggregate motion of a person or object. An economy of motion disclosed herein that is within a range is assigned a corresponding level of sterilization breach. For example, multiple displacement estimation metric thresholds can be defined, including the minimum allowable distance and two or more additional distance thresholds indicating successively greater distances between persons / objects. A displacement estimation metric disclosed herein that is within a range is assigned a corresponding level of sterilization breach.

[0067] In some examples, different thresholds can be defined for the aggregate metric and each of one or more different metrics disclosed herein (such as the economy of motion metric, thedisplacement estimation metric, metric indicative of risk of breach, metric based on movement of objects, and so on) to define ranges for multiple levels (e.g., three or more levels) of sterilization breach based on time. For example, multiple ranges of time (e.g., less than 0.1 second, 0.1-0.5 second, 0.5 - 1 second, 1-2 seconds, etc.) can be defined for a metric (e.g., aggregate, economy of motion metric, the displacement estimation metric, metric indicative of risk of breach, metric based on movement of objects, and so on). A metric calculated for a time period within a range is assigned a corresponding level of time-based sterilization breach.

[0068] In some examples, different classifications of sterilization breach can be defined the data processing system 110 and outputted (e.g., visual, audio, tactile alarms). A class of breach can be defined based on the types of people / objects in the environment 100B for which the metrics are determined. In some implementations, for a same metric determined for different types of people / objects, different classes of breach can be identified, leading to different levels of serious of the breach or the risk thereof. For example, a same economy of motion metric determined for a surgeon, a nurse, robotic system, or a patient may correspond to different classes of breach (e.g., respective one of surgeon breach, nurse breach, robot breach, and patient breach), where surgeon breach, robot breach, and patient breach may correspond to a more serious classification of breach than nurse breach. For example, a displacement estimation metric determined for a sterile object and a non-sterile object or determined for a person and a sterile object, crossing over a threshold as described, can indicate a serious breach. For example, a displacement estimation metric determined for a sterile object and a robotic arm, crossing over the same threshold as described, can indicate a serious breach. For example, a displacement estimation metric determined for a sterile object and medical staff member, crossing over the same threshold, can indicate a slight breach. For example, an economy of motion metric determined for a medical staff member, crossing over a threshold, can indicate a medium breach. For example, displacement estimation metrics determined for a same sterile object with respect to each of multiple medical staff members or objects, crossing over the same threshold, can indicate a serious breach.

[0069] FIG. 4 depicts an example first sterility presentation according to this disclosure. As illustrated by way of example in FIG. 4, a first sterility presentation 400 can include at least sterile object overlays 410, and non-sterile object overlays 420. For example, the system can cause a user interface to present the output. The output can include a region of the image having a visual property based on the metric, the region of the image defined by the object model, andthe visual property corresponding to a color based on the metric. For example, the system can classify, using the second model receiving the object model as input, the object to associate the sterility status with the object.

[0070] The sterile object overlays 410 can include one or more annotations to one or more images or videos of a medical environment during a medical procedure. The sterile object overlays 410 can indicate a status of an object identified according to the iterations of 200 A- C, and can identify a sterility state corresponding to a sterile object or person, according to the motion estimation architecture 300. For example, the sterile object overlays 410 can include a text object indicating that an object is “Sterile” and a color object providing a color overlay corresponding to boundaries of the object. For example, the color object can have a color indicating a risk of breach of sterility. For example, a green color of a given transparency or tone can indicate a magnitude of a lower risk of breach, and an orange color of a given transparency or tone can indicate a magnitude of a higher risk of breach. For example, the color object can have a color indicating a boundary between an object and another object. For example, each object in the first sterility presentation 400 can have a distinct color.

[0071] The non-sterile object overlays 420 can include one or more annotations to one or more images or videos of a medical environment during a medical procedure. The non-sterile object overlays 420 can indicate a status of an object identified according to the iterations of 200 A- C, and can identify a sterility state corresponding to a non-sterile object or person, according to the motion estimation architecture 300. For example, the non-sterile object overlays 420 can include a text object indicating that an object is “Non-Sterile” and a color object providing a color overlay corresponding to boundaries of the object. For example, the color object can have a color indicating a risk of breach of sterility. For example, a red color of a given transparency or tone can indicate a current breach. For example, the color object can have a color indicating a boundary between an object and another object. For example, each object in the first sterility presentation 400 can have a distinct color.

[0072] Thus, the data processing system 110 can identify sterility states based on the type of the object, or characteristics specific to the type of the object (e.g., a system identifies a person in scrubs in the video, and identifies that the person is sterile based on the color of the scrubs). For example, the data processing system 110 can identify specific sub-objects for specific objects (e.g., arms and legs of people, tabletop surfaces of furniture, bases and arms of robots) and apply sterility metrics to the sub-objects (e.g., different thresholds of movement ordisplacement for limbs of a person vs torso or body of a person as a whole). For example, the data processing system 110 can generate annotations with specific colors that indicate an object is sterile or non-sterile. For example, the data processing system 110 can generate annotations with specific colors that varying levels of risk that a sterile object is a varying risk levels of becoming non-sterile through contamination.

[0073] FIG. 5 depicts an example second sterility presentation according to this disclosure. As illustrated by way of example in FIG. 5, a second sterility presentation 500 can include at least sterile object overlays 510, and non-sterile object overlays 520. The sterile object overlays 510 can correspond at least partially in one or more of structure and operation to the sterile object overlays 410. The non-sterile object overlays 520 can correspond at least partially in one or more of structure and operation to the non-sterile object overlays 420. For example, the first sterility presentation 400 and the second sterility presentation 500 can each include the overlays 410, 420, 510 and 520 according to a plurality of 3D object models generated as discussed herein, where each of the 3D object models corresponds to a distinct person or object tin each of the presentations 400 and 500. Thus, the system can provide a technical improvement o provide overlays with sterility states in concurrently and in real-time for image or video data for a number of viewpoints or fields of view, beyond the capability of manual processes to achieve.

[0074] FIG. 6 depicts an example layer model architecture according to this disclosure. As illustrated by way of example in FIG. 6, a layer model architecture 600 can include at least a first layer 610, a second layer 612, a third layer 614, a fourth layer 616, a first neural network 630, a second neural network 632, a third neural network 634, a fourth neural network 636, and a mixer 650. For example, the mask processors 220, 240 and 260 can each be trained according to the layer model architecture 600. For example, the environment sensor system 310, the pose estimation system 320, and the motion estimation system 330 can each generate metrics according to the architecture of the layer model architecture 600, while being trained on distinct training data.

[0075] The first layer 610 can correspond to an instance of a vision architecture as discussed herein. The first layer 610 can include a first clip model 620, a first layer processor 630, and a first feature processor 640, and can provide output to a layer output 654. The first clip model 620 can include one or more instructions to divide a video into one or more frames, and to select one or more frames corresponding to one or more timestamps or times of captureassociated with those one or more frames. For example, the first layer processor 630 can correspond to a first viewpoint or view of view of the environment 100B. For example, the first layer processor 630 can include a recursive neural network (RNN). The first feature processor 640 can be configured to receive one or more features from the first layer processor 630. The second layer 612 can correspond to an instance of a vision architecture as discussed herein. The second layer 612 can include a second clip model 622, a second layer processor 632, and a second feature processor 642. The second layer processor 632 can correspond to a second viewpoint or view of view of the environment 100B. The second feature processor 642 can correspond at least partially in one or more of structure and operation to the first feature processor 640. The third layer 614 can correspond to an instance of a vision architecture as discussed herein. The third layer 614 can include a third clip model 624, a third layer processor 634, and a third feature processor 644. The third layer processor 634 can correspond to a third viewpoint or view of view of the environment 100B. The third feature processor 644 can correspond at least partially in one or more of structure and operation to the first feature processor 640. The fourth layer 616 can correspond to an instance of a vision architecture as discussed herein. The fourth layer 616 can include a fourth clip model 626, a fourth layer processor 636, and a fourth feature processor 646. The fourth layer processor 636 can correspond to a fourth viewpoint or view of view of the environment 100B. The fourth feature processor 646 can correspond at least partially in one or more of structure and operation to the first feature processor 640.

[0076] The mixer 650 can aggregate output from each of the first, second, third, and fourth layers 610, 612, 614 and 616. Thus, the mixer 650 can provide a fused output 652 based on predictions output by each of the first, second, third, and fourth layers 610, 612, 614 and 616. The layer output 654 can correspond to an output of the first layer 610. For example, the layer output 654 can correspond to a prediction output by the first layer 630. The layer output 654 is not limited to the example illustrated herein. For example, one or more of the second, third and fourth layers 612, 614 and 616 can provide layer outputs that correspond at least partially in one or more of structure and operation to the layer output 654. Thus, the layer model architecture 600 can identify features from one or more individual viewpoints or sets of viewpoints. This disclosure is not limited to the number of viewpoints illustrated in any given example.

[0077] In some implementations, one or more of the mask processors 220, 240 and 260, the environment sensor system 310, the pose estimation system 320, the motion estimation system 330, and the layer model architecture 600 can receive, in addition to other inputs described herein, one or more of the audio data or the robotic system data as additional inputs, to generate the outputs as described herein. For example, the environment sensor system 310, the pose estimation system 320, and the motion estimation system 330 can each generate metrics described herein using, in addition to other inputs described herein, one or more of the audio data or the robotic system data as additional inputs. The robotic system data can provide realtime information regarding the operations of the robotic systems 130, which can be indicative of the activities of the robotic systems 130. For instance, robotic system data that indicates frequent and numerous tasks being performed using the robotic system 130 over a short period of time may indicate a greater aggregate motion of the robotic system 130 and a lesser aggregate motion of the medical staff members, which can be used in conjunction with the sensor data to determine the economy of motion metric for the robotic system 130 and the persons in the environment 100B. On the other hand, robotic system data that indicates no or few tasks being performed over a period of time using the robotic system 130 may indicate a lesser aggregate motion of the robotic system 130 and a greater aggregate motion of medical staff members, which can be used in conjunction with the sensor data to determine the economy of motion metric for the robotic system 130 and the persons in the environment 100B. The captured audio data can capture in real time conversations that evidence or related to activities being performed within the environment 100B. In some examples, the robotic system data and the audio data can be provided as further inputs to the first model and the second model.

[0078] FIG. 7 depicts an example user interface with sterile person presentation according to this disclosure. As illustrated by way of example in FIG. 7, a user interface 700 with sterile person presentation can include at least a case presentation region 702, and a case navigation pane 704. The case presentation region 702 can correspond to a first portion of the user interface 700. The case presentation region 702 can include filter control affordances 706, a case data presentation 710, a case metrics presentation 720, and a case video content presentation 730. The filter control affordances 706 can receive input to modify case data available at or presented at the case presentation region 702. For example, the filter control affordances 706 can present one or more cases have metrics corresponding to one or more filters selected at the filter control affordances 706. For example, filters can indicate proceduretype, patient metrics, sterility states of one or more objects, or any combination thereof, but are not limited thereto.

[0079] The case data presentation 710 can present indications of one or more case data that satisfy one or more of the filters. For example, the case data presentation 710 can include one or more thumbnail images or a selectable gallery of instances of cases each corresponding to various medical procedure, or phases or tasks within various medical procedures. The case metrics presentation 720 can present indications of one or more metrics associated with given case data for a given medical procedure. For example, the case metrics presentation 720 can include a stacked bar chart indicative of a sterility state or risk of breach of sterility for one or more persons or objects identified in the case data. For example, the case metrics presentation 720 can include a stacked bar chart indicative of a displacement or economy of motion for one or more persons or objects identified in the case data.

[0080] The case video content presentation 730 can correspond to a portion of the case data of data presentation 710. For example, the case video content presentation 730 depicts the environment 100B from a first viewpoint at a given time. The case video content presentation 730 can include a sterile person video annotation 732. The sterile person video annotation 732 can correspond at least partially in one or more of structure and operation to an instance of the sterile object overlays 410, and can indicate sterility of a person 160 according to a 3D object model corresponding to the person 160 and a sterility state threshold corresponding to the person 160. The case navigation pane 704 can include various control affordances to select various data associated with one or more medical procedures or medical environments, or any set thereof (e.g., all medical environments of a hospital).

[0081] FIG. 8 depicts an example user interface with sterile object presentation according to this disclosure. As illustrated by way of example in FIG. 8, a user interface 800 with sterile object presentation can include at least a case data presentation 810, and a case video content presentation 820. The case data presentation 810 can correspond at least partially in one or more of structure and operation to the case data presentation 710. The case video content presentation 820 can correspond to a portion of the case data of data presentation 810. For example, the case video content presentation 820 depicts the environment 100B from a second viewpoint at a given time, during the same medical procedure as depicted in the case video content presentation 730. The case video content presentation 820 can include a sterile object video annotation 822. The sterile object video annotation 822 can correspond at least partiallyin one or more of structure and operation to an instance of the sterile object overlays 410, and can indicate sterility of a person 160, an object 170 according to a 3D object model corresponding to the person 160 and a sterility state threshold corresponding to the object 170.

[0082] FIG. 9 depicts an example user interface for sterility training according to this disclosure. As illustrated by way of example in FIG. 9, a user interface 900 for sterility training can include at least a sterility configuration presentation 902, an image feature presentation 910, a feature identification presentation 912, and a feature classification presentation 920. The sterility configuration presentation 902 can include a first portion of the user interface 900 presenting one or more options for configuration of control affordances of the user interface 800 to receive input parameters for sterility training. For example, the sterility configuration presentation 902 can receive input to set a default sterility state for an object in a video (e.g., one or more video frames) or image of training data. In some examples, the training data can include a video or an image (e.g., visual or near-visual spectra and / or one or more depthacquiring video or image) that can be provided to the data processing system 110 to determine a metric for one or more persons or objects in the manner described. The data processing system 110 can determine an output including whether a person or an object is sterile or non- sterile, determine a level of sterilization breach, or determine a classification of breach, in the manner described. The output can be used to compare with the output indicated by a user of the user interface 900 for training purposes.

[0083] The image feature presentation 910 can correspond to a person or object in the video frame or image of the training data. The feature identification presentation 912 can include a control affordance to identify the image feature presentation 910 as a person or object (or a type of person or object) in the video frame or image of the training data. For example, the feature identification presentation 912 can include an editable selection box control affordance that can be drawn around a person or object depicted in a video frame or an image. The feature classification presentation 920 can receive input to identify a sterility state of the identified person or object. For example, the feature classification presentation 920 can be displayed a pop-up window in some examples, and in other examples, the feature classification presentation 920 can be displayed in another screen or overlay. The feature classification presentation 920 can include a feature classification affordance 922. The feature classification affordance 922 can receive input from a user to select a sterility state (e.g., sterile, non-sterile, a level of sterilization breach, or a classification of breach). For example, the featureclassification affordance 922 can present options of “sterile” or “non-sterile” classifications, and can receive selection input of either the “sterile” classification or “non-sterile” classification. For example, the feature classification affordance 922 can present different levels of sterilization breach or classifications of breach, and can receive selection input of one of the levels or one of the classifications.

[0084] In some examples, in response to the user selection, the user interface 900 can present the output determined by the data processing system 110 including whether a person or an object is sterile or non-sterile, determine a level of sterilization breach, or determine a classification of breach. The output of the data processing system 110 can be displayed as a pop-up window, an overlay window, or in another screen in the user interface 900. In some examples, the output of the data processing system 110 and the user input can be displayed side-by-side, in a same screen of the user interface 900, to allow the user to view the user’s selection and the information determined by the data processing system 110 simultaneously. In some examples, the user interface 900 can display the aggregate metric, the economy of motion metric, the displacement estimation metric, metric indicative of risk of breach, metric based on movement of objects, and so on to the user for further information. In some examples, the user interface 900 can display any threshold disclosed herein to the user for further information.

[0085] FIG. 10 depicts an example method of a computer vision architecture to identify and monitor sterility states in medical procedures, according to this disclosure. At least one of the data processing system 110, the mask processor 220, the mask processor 240, or the mask processor 260 can perform method 1000.

[0086] At 1010, the method 1000 can identify an object within an image of a medical environment. At 1012, the method 1000 can identify the object using a first model configured to detect image features. For example, the method can include generating, based on an image corresponding to a two-dimensional view of the object, a three-dimensional object model corresponding to the object from the view. For example, the method can include generating, based on a second image corresponding to a second two-dimensional view of the object, the three-dimensional object model corresponding to the object from the view and the second view.

[0087] At 1020, the method 1000 can classify the object to associate a sterility status with the object. At 1022, the method 1000 can classify the object using a second model. For example,the method can include classifying, using the second model receiving the object model as input, the object to associate the sterility status with the object. At 1024, the method 1000 can classify the sterility status indicative of a first state of the object as sterile. At 1026, the method 1000 can classify the sterility status indicative of a second state of the object as non-sterile.

[0088] At 1030, the method 1000 can monitor movement of the object within the medical environment. At 1032, the method 1000 can monitor the movement using the first model. For example, the method can include monitoring, using the first model, a position of the object within the medical environment. The method can include determining, based on the position of the object within the medical environment, the metric corresponding to the object. For example, the method can include monitoring, using the first model, an orientation of the object within the medical environment. The method can include determining, based on the orientation of the object within the medical environment, the metric corresponding to the object.

[0089] FIG. 11 depicts an example method of a computer vision architecture to identify and monitor sterility states in medical procedures, according to this disclosure. At least one of the data processing system 110, the mask processor 220, the mask processor 240, or the mask processor 260 can perform method 1100.

[0090] At 1110, the method 1100 can determine a metric for the object. For example, the metric can be indicative of at least one of an economy of motion of at least one person in the medical environment, or a pose of at least one person in the medical environment. At 1112, the method 1100 can determine the metric indicative of a risk of breach of sterility of the object. At 1114, the method 1100 can determine the metric based on the movement of the object within the medical environment. For example, the method can include determining, using the first model, the movement of the object within the medical environment. The method can include determining, using the second model, the metric.

[0091] At 1120, the method 1100 can generate an output indicative of the metric meeting the threshold. For example, the method can include causing a user interface to present the output. The output can include a region of the image having a visual property based on the metric, the region of the image defined by the object model, and the visual property corresponding to a color based on the metric. At 1122, the method 1100 can generate the output according to the metric meeting a threshold indicative of the second state of the object as non-sterile. At 1124, the method 1100 can generate the output in real-time.

[0092] Having now described some illustrative implementations, the foregoing is illustrative and not limiting, having been presented by way of example. In particular, although many of the examples presented herein involve specific combinations of method acts or system elements, those acts and those elements may be combined in other ways to accomplish the same objectives. Acts, elements and features discussed in connection with one implementation are not intended to be excluded from a similar role in other implementations.

[0093] The phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. The use of "including," "comprising," "having," "containing," "involving," "characterized by," "characterized in that," and variations thereof herein, is meant to encompass the items listed thereafter, equivalents thereof, and additional items, as well as alternate implementations consisting of the items listed thereafter exclusively. In one implementation, the systems and methods described herein consist of one, each combination of more than one, or all of the described elements, acts, or components.

[0094] References to "or" may be construed as inclusive so that any terms described using "or" may indicate any of a single, more than one, and all of the described terms. References to at least one of a conjunctive list of terms may be construed as an inclusive medical environment to indicate any of a single, more than one, and all of the described terms. For example, a reference to "at least one of 'A' and 'B'" can include only 'A', only 'B', as well as both 'A' and 'B'. Such references used in conjunction with "comprising" or other open terminology can include additional items. References to "is" or "are" may be construed as nonlimiting to the implementation or action referenced in connection with that term. The terms "is" or "are" or any tense or derivative thereof, are interchangeable and synonymous with "can be" as used herein, unless stated otherwise herein.

[0095] Directional indicators depicted herein are example directions to facilitate understanding of the examples discussed herein, and are not limited to the directional indicators depicted herein. Any directional indicator depicted herein can be modified to the reverse direction, or can be modified to include both the depicted direction and a direction reverse to the depicted direction, unless stated otherwise herein. While operations are depicted in the drawings in a particular order, such operations are not required to be performed in the particular order shown or in sequential order, and all illustrated operations are not required to be performed. Actions described herein can be performed in a different order. Where technical features in the drawings, detailed description or any claim are followed by reference signs, the reference signshave been included to increase the intelligibility of the drawings, detailed description, and claims. Accordingly, neither the reference signs nor their absence have any limiting effect on the scope of any clam elements.

[0096] Scope of the systems and methods described herein is thus indicated by the appended claims, rather than the foregoing description. The scope of the claims includes equivalents to the meaning and scope of the appended claims.

Claims

WHAT IS CLAIMED IS:

1. A method, comprising: identifying, using a first model configured to detect image features, an object within an image of a medical environment; classifying, using a second model, the object to associate a sterility status with the object, the sterility status indicative of a first state of the object as sterile or a second state of the object as non-sterile; monitoring, using the first model, movement of the object within the medical environment; determining, based on the movement of the object within the medical environment, a metric indicative of a risk of breach of sterility of the object; and generating, in response to the metric meeting a threshold indicative of the second state of the object as non-sterile, an output in real-time indicative of the metric meeting the threshold.

2. The method of claim 1, further comprising: monitoring, using the first model, a position of the object within the medical environment; and determining, based on the position of the object within the medical environment, the metric corresponding to the object.

3. The method of claim 1, further comprising: monitoring, using the first model, an orientation of the object within the medical environment; and determining, based on the orientation of the object within the medical environment, the metric corresponding to the object.

4. The method of claim 1, further comprising: determining, using the first model, the movement of the object within the medical environment; and determining, using the second model, the metric.

5. The method of claim 1, further comprising: generating, based on an image corresponding to a two-dimensional view of the object, a three-dimensional object model corresponding to the object from the view.

6. The method of claim 5, further comprising: generating, based on a second image corresponding to a second two-dimensional view of the object, the three-dimensional object model corresponding to the object from the view and the second view.

7. The method of claim 5, further comprising: classifying, using the second model receiving the object model as input, the object to associate the sterility status with the object.

8. The method of claim 5, further comprising: causing a user interface to present the output including a region of the image having a visual property based on the metric, the region of the image defined by the object model, and the visual property corresponding to a color based on the metric; or causing the user interface to present the output.

9. The method of claim 1, the metric indicative of at least one of: an economy of motion of at least one person in the medical environment; a pose of at least one person in the medical environment; a displacement estimation metric related to a distance between a person and an object, between two objects, or between two people in the medical environment.

10. The method of claim 1, wherein the metric is determined based at least in part on a period of time in which the metric is determined.

11. The method of claim 1, wherein the output comprises one of a plurality of levels of sterility breach or one of a plurality of classifications of sterilization breach.

12. The method of claim 1, wherein at least one audio data of the medical environment or robotic system data of at least one robotic system in the medical environment is provided as input to the first model or the second model.

13. A system, comprising: a memory and one or more processors to: identify, using a first model configured to detect image features, an object within an image of a medical environment; classify, using a second model, the object to associate a sterility status with the object, the sterility status indicative of a first state of the object as sterile or a second state of the object as non-sterile; monitor, using the first model, movement of the object within the medical environment; determine, based on the movement of the object within the medical environment, a metric indicative of a risk of breach of sterility of the object; and generate, in response to the metric meeting a threshold indicative of the second state of the object as non-sterile, an output in real-time indicative of the metric meeting the threshold.

14. The system of claim 13, the processors to: monitor, using the first model, a position of the object within the medical environment; anddetermine, based on the position of the object within the medical environment, the metric corresponding to the object.

15. The system of claim 13, the processors to: monitor, using the first model, an orientation of the object within the medical environment; and determine, based on the orientation of the object within the medical environment, the metric corresponding to the object.

16. The system of claim 13, the processors to: determine, using the first model, the movement of the object within the medical environment; and determine, using the second model, the metric.

17. The system of claim 13, the processors to: generate, based on an image corresponding to a two-dimensional view of the object, a three-dimensional object model corresponding to the object from the view.

18. The system of claim 17, the processors to: generate, based on a second image corresponding to a second two-dimensional view of the object, the three-dimensional object model corresponding to the object from the view and the second view.

19. The system of claim 17, the processors to: classify, using the second model receiving the object model as input, the object to associate the sterility status with the object.

20. A non-transitory computer readable medium including one or more instructions stored thereon and executable by a processor to: identify, by the processor using a first model configured to detect image features, an object within an image of a medical environment;classify, by the processor using a second model, the object to associate a sterility status with the object, the sterility status indicative of a first state of the object as sterile or a second state of the object as non-sterile; monitor, by the processor using the first model, movement of the object within the medical environment; determine, by the processor and based on the movement of the object within the medical environment, a metric indicative of a risk of breach of sterility of the object; and generate, by the processor and in response to the metric meeting a threshold indicative of the second state of the object as non-sterile, an output in real-time indicative of the metric meeting the threshold.

Citation Information

Patent Citations

  • Sterility in an operation room

    EP3975201A1

  • Tracking performance of medical procedures

    WO2024015620A1