Sensor optimization architecture for medical procedures
Patent Information
- Application Number
- PCT/US2024/062138
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-12-27
- Publication Date
- 2025-08-07
AI Technical Summary
In medical environments like operating rooms, objects, people, and movements can obscure or be obscured, leading to inefficient and ineffective support and evaluation of medical procedures due to visibility issues.
A sensor optimization architecture that uses machine learning models to identify states in a medical environment, adjusting sensor configurations such as pose, modality, and parameters to optimize data capture, including multi-modal data acquisition and fusion, to enhance visibility and resource efficiency.
Improves detection accuracy and resource utilization by dynamically adjusting sensor configurations based on environmental states, reducing occlusions and enhancing data quality for medical procedures.
Smart Images

Figure US2024062138_07082025_PF_FP_ABST
Abstract
Description
SENSOR OPTIMIZATION ARCHITECTURE FOR MEDICAL PROCEDURESCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of, and priority to, U.S. Provisional Patent Application No. 63 / 615,722, filed December 28, 2023, the full disclosure of which is incorporated herein in its entirety.TECHNICAL FIELD
[0002] The present implementations relate generally to medical devices, including but not limited to a sensor optimization architecture for medical procedures.INTRODUCTION
[0003] Awareness of objects, people, and movements associated with a medical procedure is crucial to effective and efficient completion of the medical procedure. A medical environment can include an operating room (OR) with a large number of people, pieces of furniture, medical instruments, and medical devices. Each of these people, pieces of furniture, medical instruments, and medical devices can potentially obscure or be obscured by others of the people, pieces of furniture, medical instruments, and medical devices in the medical environment. Effective support of the medical procedure during the medical procedure and evaluation of the medical procedure subsequent to the medical procedure can be negatively impacted or otherwise impossible to perform where people, pieces of furniture, medical instruments, or medical devices are obscured or not visible.SUMMARY
[0004] Systems, methods, apparatuses, and non-transitory computer-readable media are provided for updating one or more sensors in a medical environment (e.g., an OR) based on, for example, a state of the medical environment. For example, a model trained with machine learning can receive video from one or more sensors, and can identify a state of the medical environment based on one or more features extracted from the video by the model. For example, a model can correspond to a machine learning model including a neural network model that can analyze input provided as video of a given procedure performed at a given OR. The user could select or reference the video from a database of recordings or menu of available recordings, for example. For example, the neural network can generate features describingobjects or motion in or across video frames. A system can determine, based on a given state of the OR, a configuration for the sensor and can update the sensor to operate in accordance with the determined configuration. The configuration for the sensor may include one or more sensing parameters, a pose of the sensor within the medical environment, and / or a sensing modality. Sensing parameters of a sensor can affect the manner in which a sensor generates sensor output data in response to detecting input from an environment. For example, sensing parameters of a sensor can include video resolution, video framerate, ISO (sensitivity), dynamic range, gain, video contrast, subsampling (binning) ratio, laser projection framerate, laser projection power, or any combination thereof, but are not limited thereto. For example, pose of a sensor can include a position in the medical environment, an orientation of the sensor, or any combination thereof, but are not limited thereto. The system can modify the sensor by modifying one or more of the parameters or the positions of the sensors. For example, the system can modify the sensing parameters of a first sensor viewing an OR room to a lower resolution or power, and can modify the parameters of a second sensor viewing a robot of the OR room to a higher resolution or power, in response to detecting a robot docking state of the OR. The system can thus provide at least a technical improvement to increase efficiency in both computational resources and energy resources, by allocation of resources to sensors best positioned for view or detect activity corresponding to given states of the OR. Thus, a technical solution for a multi-sensor optimization architecture for medical procedures is provided.
[0005] In certain embodiments, the sensor(s) for capturing data of the medical environment is a multi-modal sensor configured to capture data in multiple modalities. For example, a sensor in the medical environment may be configured to capture video data and / or depth data (e.g., three-dimensional point cloud data). The depth data may be captured using time-of-flight, structured light, stereoscopic sensing, or any other appropriate technique for acquiring depth and three-dimensional data of the medical environment. Furthermore, the system may be configured to adjust a sensing modality of the sensor based on the detected state of the OR. The system may be further configured to modify the sensing modality of the sensor(s) monitoring the medical environment based on, for example, the determined state of the medical environment. For example, in response to determining that a state of the medical environment that corresponds to a patient being wheeled into the medical environment, the system can modify the sensing modality of the sensor(s) in the medical environment such that only depth data is generated by the sensor(s) (e.g., turning off generation of video data by the sensor(s).In this manner, protected health information (PHI) may be protected while data of the medical environment is still captured while the patient is being wheeled into the medical environment.
[0006] At least one aspect is directed to a system. The system can include one or more processors, coupled with memory. The system can receive, via a sensor located about a medical environment, video of a medical procedure, where the video can include at least one of medical staff, a patient, a robotic system, instrument, or the medical environment. The system can identify, using a model configured to detect one or more features in the video, a state indicative of at least one of the robotic system, the instrument, or the medical environment. The system can determine, based at least on the state, a configuration for the sensor. The system can update the sensor to operate in accordance with the determined configuration for the sensor.
[0007] At least one aspect is directed to a system. The system can include one or more processors, coupled with memory. The system can receive, via a sensor located about a medical environment, video of a medical procedure, where the video can include at least one of medical staff, a patient, a robotic system, instrument, or the medical environment. The system can identify, using a model configured to detect one or more features in the video, a state indicative of at least one of the robotic system, the instrument, or the medical environment. The system can determine, based at least on the state, a parameter of the sensor, where the parameter corresponds to a configuration of the sensor. The system can update the model based at least on a loss determined with respect to the parameter.
[0008] At least one aspect is directed to a non-transitory computer readable medium can include one or more instructions stored thereon and executable by a processor. The processor can receive, via a sensor located about a medical environment, video of a medical procedure, where the video can include at least one of medical staff, a patient, a robotic system, instrument, or the medical environment. The processor can identify, using a model configured to detect one or more features in the video, a state indicative of at least one of the robotic system, the instrument, or the medical environment. The processor can determine, based at least on the state, a parameter of the sensor, where the parameter corresponds to a configuration of the sensor. The processor can update, based at least on the parameter, the sensor to correspond to the configuration.
[0009] At least one aspect is directed to a non-transitory computer readable medium can include one or more instructions stored thereon and executable by a processor. The processorcan receive, via a sensor located about a medical environment, video of a medical procedure, where the video can include at least one of medical staff, a patient, a robotic system, instrument, or the medical environment. The processor can identify, using a model configured to detect one or more features in the video, a state indicative of at least one of the robotic system, the instrument, or the medical environment. The processor can determine, based at least on the state, a parameter of the sensor, where the parameter corresponds to a configuration of the sensor. The processor can update the model based at least on a loss determined with respect to the parameter.
[0010] At least one aspect is directed to a method. The method can include receiving, via a sensor located about a medical environment, video of a medical procedure, where the video can include at least one of medical staff, a patient, a robotic system, instrument, or the medical environment. The method can include identifying, using a model configured to detect one or more features in the video, a state indicative of at least one of the robotic system, the instrument, or the medical environment. The method can include determining, based at least on the state, a parameter of the sensor, where the parameter corresponds to a configuration of the sensor. The method can include updating, based at least on the parameter, the sensor to correspond to the configuration.
[0011] At least one aspect is directed to a method. The method can include receiving, via a sensor located about a medical environment, video of a medical procedure, where the video can include at least one of medical staff, a patient, a robotic system, instrument, or the medical environment. The method can include identifying, using a model configured to detect one or more features in the video, a state indicative of at least one of the robotic system, the instrument, or the medical environment. The method can include determining, based at least on the state, a parameter of the sensor, where the parameter corresponds to a configuration of the sensor. The method can include updating the model based at least on a loss determined with respect to the parameter.BRIEF DESCRIPTION OF THE FIGURES
[0012] These and other aspects and features of the present implementations are depicted by way of example in the figures discussed herein. Present implementations can be directed to, but are not limited to, examples depicted in the figures discussed herein. Thus, this disclosure is not limited to any figure or portion thereof depicted or referenced herein, or any aspect described herein with respect to any figures depicted or referenced herein.
[0013] FIG. 1 A depicts an example architecture of a system according to this disclosure.
[0014] FIG. IB depicts an example environment of a system according to this disclosure.
[0015] FIG. 2 depicts an example sensor control system according to this disclosure.
[0016] FIG. 3 depicts an example single-view vision architecture according to this disclosure.
[0017] FIG. 4A depicts an example early fusion vision architecture according to this disclosure.
[0018] FIG. 4B depicts an example mid-fusion vision architecture according to this disclosure.
[0019] FIG. 4C depicts an example late-fusion vision architecture according to this disclosure.
[0020] FIG. 5 depicts an example second-stage fusion vision architecture according to this disclosure.
[0021] FIG. 6 depicts an example layer model architecture according to this disclosure.
[0022] FIG. 7A depicts an example operating environment before optimization according to this disclosure.
[0023] FIG. 7B depicts an example operating environment after parameter optimization according to this disclosure.
[0024] FIG. 7C depicts an example operating environment after position optimization according to this disclosure.
[0025] FIG. 8 depicts an example method of a sensor optimization architecture for medical procedures according to this disclosure.
[0026] FIG. 9 depicts an example method of a sensor optimization architecture for medical procedures according to this disclosure.
[0027] FIG. 10 depicts an example method of a sensor optimization architecture for medical procedures according to this disclosure.DETAILED DESCRIPTION
[0028] Aspects of this technical solution are described herein with reference to the figures, which are illustrative examples of this technical solution. The figures and examples below are not meant to limit the scope of this technical solution to the present implementations or to a single implementation, and other implementations in accordance with present implementations are possible, for example, by way of interchange of some or all of the described or illustrated elements. Where certain elements of the present implementations can be partially or fully implemented using known components, only those portions of such known components that are necessary for an understanding of the present implementations are described, and detailed descriptions of other portions of such known components are omitted to not obscure the present implementations. Terms in the specification and claims are to be ascribed no uncommon or special meaning unless explicitly set forth herein. Further, this technical solution and the present implementations encompass present and future known equivalents to the known components referred to herein by way of description, illustration, or example.
[0029] Aspects of this disclosure are directed to systems and methods to automatically propose number and location / pose of multiple sensors in an OR, to maximize detection of people, objects, and motions relevant to a given medical procedure. For example, a system according to this disclosure can provide applications such as understanding aspects of a workflow (time, motion, etc.) for a medical procedure, detection of or tracking of patients, clinicians, or objects, 3D reconstruction of the medical environment at one or more points in time during the medical procedure, and inventory tracking. To create a holistic understanding of activities, clinicians and equipment, an OR can include multiple sensors at various positions and configured to capture aspects of the OR as various resolutions. The arrangement and configuration of the sensors can collectively reduce or eliminate occlusion, and can address clutter in the OR, sensor resolution or field-of-view (FOV) limitations, and variety in downstream tasks / objectives to be activated at different times. Thus, a multi-sensor technical solution as discussed herein can provide at least a technical improvement to create more robustness in detection of objects in the OR by providing redundancy of sensor coverage within the OR environment. As discussed herein, a pose can correspond to a location of an object in a 3D coordinate space, an orientation of an object in the 3D coordinate space, or a combination or aggregation thereof.
[0030] Determining the number of sensors needed and the location and orientation of this sensors can be based on a type of downstream task that informs the activity performed by thesensor, performance requirements for the task (e.g., resolution, frame rate), limitations on material cost, limitations of available computing resources (e.g., available local memory, available network bandwidth, available CPU / GPU processing cores), freedom of movement of mechanical fixtures for locating or orienting sensors in the room, and the type / modality of surgical procedure being done in the room (e.g., robotic, open, lap, prostatectomy, gynecological). Operating room size and layout can also affect sensor system configuration. For example, OR imaging sensors can be wall mounted, ceiling mounted, attached to ceiling light equipment, or placed on mobile carts. Scene detection can be based on using a fixed type and location of sensors being used throughout the medical procedure. This limits the capability to enable different applications (e.g., lack of high-quality data) or creates redundant data, through coverage of a given location by multiple sensors. A pan, tilt, and zoom (PTZ) camera is a robotic video camera that can be controlled by a remote operator. A PTZ camera can pan horizontally, tilt vertically and zoom in on a subject to enhance image quality without digital pixelation. However, there are conventional systems cannot automatically change orientation of one or more sensors (e.g., PTZ cameras) to enable a specific Al perception task.
[0031] Systems and methods according to this disclosure can include a sensor optimization system for generating multi-modal data of a medical environment (e.g., an OR) that can automatically adjust configuration (e.g., sensing parameters, pose, and / or sensing modality) of one or more sensors to maximize data quality for specific applications and save cost and computing resources associated with power, data processing, transfer, storage. For example, sensors can be wall-mounted or ceiling-mounted on one or more rails that can enable sensor movement along and / or rotation about one or more degrees of freedom. For example, sensors can be provided on a motorized cart or piece of furniture, to enable sensor movement along one or more degrees of freedom.
[0032] Systems and methods according to this disclosure can process data from a single sensor for applications like activity / phase detection, object / human detection, and 3D reconstruction. Systems and methods according to this disclosure can process data and fuse information from multiple sensors. Fusing sensor input for on clinical data (different ORs, procedures, workflows) can provide increased resolution and accuracy compared to sensor systems without fused input. Further, the pose (e.g., position and / or orientation) of sensors plays an important role in the performance of multi-view fusion Al models according to this disclosure. For example, multi-view sensor fusion for specific procedure phases, procedure types or ORlayouts, or different sensor locations result in better performance in accuracy, resource consumption, or both. For example, a multi-view fusion Al system is deployed in the OR, consuming data from all sensors, performing activity or object recognition in real-time. The multi-view fusion Al system can automatically understand input of each sensor, assign a data quality score to each, and automatically adjust pose(s) (e.g., re-position and / or re-orient) one or more sensors to improve data quality and overall performance. For example, through analysis of large clinical and in-house data that contain various discrete sensor poses, system and methods according to this disclosure can re-position sensors to one or more ideal sensor locations corresponding to various procedures, activity phases or OR layouts. A system can automatically adjust sensor locations for each procedure type, phase, or OR layout, based on the data indicating the discrete sensor locations. For example, the system can automatically disable one or more given sensors (or adjust a sensing modality) according to a determination that input from the one or more given sensors is not useful. Thus, deactivation of the one or more given sensors according to the determination can result in a reduction in consumption of resources including power, computation resources, and data transfer costs.
[0033] Systems and methods according to this disclosure can communicate data from a surgical robotic device with one or more sensors of the OR. For example, communication of event information from a robotic surgical device in the OR can cause sensor movement (e.g., location adjustment) to enable a sensor to capture an unobscured view of specific actions like intraoperative disruptions driven by the robotic surgical device. For example, a system as discussed herein can perform multiple tasks at the same time. For example, during the medical procedure, a subset of sensors are allocated for detecting overall activity of the OR, while others are located and oriented to detect collisions between robot or patient, understand bedside activities or look for other intraoperative disruptions. Thus, for example, a multi-view fusion Al system can identify and allocate a first subset of sensors for a first task having a first priority, can identify and allocate a second subset of sensors for a second task having a second priority. For example, a system according to this disclosure can complete 3D reconstruction of the OR. For example, the system can move sensors while saving their location data, to position sensors to detect or capture input from portions of the OR that would otherwise result in missing data. Sensor data and location can be used in a 3D reconstruction / stitching algorithm to partially or fully reconstruct the OR at one or more times.
[0034] FIG. 1A depicts an example architecture of a system according to this disclosure. As illustrated by way of example in FIG. 1 A, an architecture of a system 100 A can include at least a data processing system 110, communication bus 120, and a robotic system 130.
[0035] In some embodiments, the system 100 A can configure multiple sensors in the OR based on detection of s state or scene corresponding to the OR as a whole. The system can detect, for example, a robot docking scene as discussed above, and can configure multiple sensors in the OR according to the field of view of the sensor or the location of the sensor in the OR. For example, the system can enter a training mode in which a model is trained with machine learning from input including video from a plurality of camera sensors distributed within the OR. The machine learning model can optimize the allocation of resources to each of the cameras during corresponding modes, using a loss function that is based on at least one of the video data, parameters assigned to each sensor, positions assigned to each sensor, compute resources allocated to each sensor, or energy resources allocated to each sensor. This way, the machine learning model can be updated (e.g., trained) to provide a technical improvement to optimize sensor placement and resource consumption on an individualized basis for each sensor, in view of a state of the OR as a whole. The system can combine input from a plurality of sensors to provide “fused” sensor input to the machine learning model. For example, the system can designate two sensors placed near each other and having a field of view oriented toward an operating table as a stereoscopic fused sensor input. For example, the system can designate two sensors placed near each other and having a field of view oriented toward an operating room as a panoramic fused sensor input. The machine learning model can treat video input from each of these fused sensors as a combined input for determining optimized allocation or a loss. This way, the machine learning model can be updated (e.g., trained) to provide a technical improvement to optimize sensor placement and resource consumption on a combined basis for sensors having coordinated roles, to increase optimization and efficiency of resource allocation beyond that provided by individualized loss determinations for all sensors.
[0036] The data processing system 110 can include a physical computer system operatively coupled or that can be coupled with one or more components of the system 100 A, either directly or directly through an intermediate computing device or system. The data processing system 110 can include a virtual computing system, an operating system, and a communication bus toeffect communication and processing. The data processing system 110 can include a system processor 112 and a system memory 114.
[0037] The system processor 112 can execute one or more instructions associated with the system 112. The system processor 112 can include an electronic processor, an integrated circuit, or the like including one or more of digital logic, analog logic, digital sensors, analog sensors, communication buses, volatile memory, nonvolatile memory, and the like. The system processor NNN can include, but is not limited to, at least one microcontroller unit (MCU), microprocessor unit (MPU), central processing unit (CPU), graphics processing unit (GPU), physics processing unit (PPU), embedded controller (EC), or the like. The system processor 112 can include a memory operable to store or storing one or more instructions for operating components of the system processor 112 and operating components operably coupled to the system processor 112. The one or more instructions can include at least one of firmware, software, hardware, operating systems, embedded operating systems, and the like. The system processor 112 can include at least one communication bus controller to effect communication between the system processor 112 and the other elements of the system 100.
[0038] The system memory 114 can store data associated with the data processing system 110. The system memory 114 can include one or more hardware memory devices to store binary data, digital data, or the like. The system memory 114 can include one or more electrical components, electronic components, programmable electronic components, reprogrammable electronic components, integrated circuits, semiconductor devices, flip flops, arithmetic units, or the like. The system memory 114 can include at least one of a non-volatile memory device, a solid-state memory device, a flash memory device, or a NAND memory device. The system memory 114 can include one or more addressable memory regions disposed on one or more physical memory arrays. A physical memory array can include a NAND gate array disposed on, for example, at least one of a particular semiconductor device, integrated circuit device, and printed circuit board device.
[0039] The communication bus 120 can communicatively couple the data processing system 110 with the robotic system 130. The communication bus 120 can communicate one or more instructions, signals, conditions, states, or the like between one or more of the data processing system 110 and components, devices, blocks operatively coupled or couplable therewith. The communication bus 120 can include one or more digital, analog, or like communication channels, lines, traces, or the like. As one example, the communication bus 120 can include atleast one serial or parallel communication line among multiple communication lines of a communication interface. The communication bus 120 can include one or more wireless communication devices, systems, protocols, interfaces, or the like. The communication bus 120 can include one or more logical or electronic devices including but not limited to integrated circuits, logic gates, flip flops, gate arrays, programmable gate arrays, and the like. The communication bus 120 can include one or more telecommunication devices including but not limited to antennas, transceivers, packetizers, and wired interface ports.
[0040] The robotic system 130 can include one or more robotic devices configured to perform one or more actions of a medical procedure (e.g., a surgical procedure). For example, a robotic device can include, but is not limited to a surgical device that can be manipulated by robotic device. For example, a surgical device can include, but is not limited to, a scalpel or a cauterizing tool. The robotic system 130 can include various motors, actuators, or electronic devices whose position or configuration can be modified according to input at one or more robotic interfaces. For example, a robotic interface can include a manipulator with one or more levers, buttons, or grasping controls that can be manipulated by pressure or gestures from one or more hands, arms, fingers, or feet. The robotic system 130 can include a surgeon console in which the surgeon can be positioned (e.g., standing or seated) to operate the robotic system 130. However, the robotic system 130 is not limited to a surgeon console co-located or on-site with the robotic system 130.
[0041] FIG. IB depicts an example environment of a system according to this disclosure. As illustrated by way of example in FIG. IB, an environment 100B of a system 100A can include at least the robotic system 130 having a field of view 132, a first sensor system 140, a second sensor system 150, persons 160, and objects 170. For example, the environment 100B is illustrated by way of example as a plan view of an OR having the robotic system 130, the first sensor system 140, the second sensor system 150, the persons 160, and the objects 170 disposed therein or thereabout. The presence, placement, orientation, and configuration, for example, of one or more of the robotic system 130, the first sensor system 140, the second sensor system 150, the persons 160, and the objects 170 can correspond to a given medical procedure or given type of medical procedure that is being performed, is to be performed, or can be performed in the OR corresponding to the environment 100B. This disclosure is not limited to the presence, placement, orientation, or configuration of the robotic system 130, the first sensor system 140, the second sensor system 150, the persons 160, or the objects 170, or any other element,illustrated herein by way of example. The field of view 132 of the robotic system 130 can correspond to a physical volume within the environment 100B that is within range of detection of one or more sensors of the robotic system 130. For example, the field of view 132 is positioned above a surgical site of a patient. For example, the field of view 132 is oriented toward a surgical site of a patient.
[0042] The first sensor system 140 can include one or more sensors oriented to a first portion of the environment 100B. For example, the first sensor system 140 can include one or more cameras configured to capture images or video in visual or near-visual spectra and / or one or more depth-acquiring sensors for capturing depth data (e.g., three-dimensional point cloud data). For example, the first sensor system 140 can include a plurality of cameras configures to collectively capture images or video in a stereoscopic view. For example, the first sensor system 140 can include a plurality of cameras configures to collectively capture images or video in a panoramic view. The first sensor system 140 can include a field of view 142. The field of view 142 can correspond to a physical volume within the environment 100B that is within range of detection of one or more sensors of the first sensor system 140. For example, the field of view 142 is oriented toward a surgical site of a patient. For example, the field of view 152 is located behind a surgeon at surgical site of a patient.
[0043] The second sensor system 150 can include one or more sensors oriented to a second portion of the environment 100B. For example, the second sensor system 150 can include one or more cameras configured to capture images or video in visual or near-visual spectra and / or one or more depth-acquiring sensors for capturing depth data (e.g., three-dimensional point cloud data). For example, the second sensor system 150 can include a plurality of cameras configures to collectively capture images or video in a stereoscopic view. For example, the second sensor system 150 can include a plurality of cameras configures to collectively capture images or video in a panoramic view. The second sensor system 150 can include a field of view 152. The field of view 152 can correspond to a physical volume within the environment 100B that is within range of detection of one or more sensors of the second sensor system 150. For example, the field of view 152 is oriented toward the robotic system 130. For example, the field of view 152 is located adjacent to the robotic system 130.
[0044] The persons 160 can include one or more individuals present in the environment 100B. For example, the persons can include, but are not limited to, assisting surgeons, supervising surgeons, specialists, nurses, or any combination thereof. The objects 170 can include, but arenot limited to, one or more pieces of furniture, instruments, or any combination thereof. For example, the objects 170 can includes tables and surgical instruments.
[0045] FIG. 2 depicts an example sensor control system according to this disclosure. As illustrated by way of example in FIG. 2, a sensor control system 200 can include at least a sensor mode scheduling system 210, and an environment processing system 220. For example, the sensor control system 200 can be at least partially housed in data processing system 110, but is not limited thereto. The sensor control system 200 can communicate with the first sensor system 140 and the second sensor system 150 by a wired communication interface or wireless communication interface. For example, the communication interface can correspond to or be a component of the communication bus 120.
[0046] The sensor mode scheduling system 210 can provide instructions to one or more sensor systems according to or in response to one or more metrics corresponding to the robotic system 130, the environment 100B, or a medical procedure of the environment 100B, or any combination thereof. For example, the sensor mode scheduling system 210 can include one or more logical or electronic devices including but not limited to integrated circuits, logic gates, flip flops, gate arrays, programmable gate arrays, and the like. One or more electrical, electronic, or like devices, or components associated with the sensor mode scheduling system 210 can also be associated with, integrated with, integrable with, replaced by, supplemented by, complemented by, or the like, the data processing system 110 or any component thereof.
[0047] The sensor mode scheduling system 210 can provide instructions to one or more sensors or sensor system as discussed herein, to change location, orientation or configuration of a given sensor or sensor system as discussed herein, according to one or more input metrics 212. The input metrics 212 can be indicative of a state of the environment 100B, the environment 100B, or the robotic system 130, or any component, person, or object thereof, any combination thereof. For example, the input metrics 212 can include a workflow phase metric that indicates a current phase of a medical procedure. For example, the input metrics 212 can include a robot data metric that indicates telemetry of one or more components of the robotic device 130. For example, the input metrics 212 can include a room motion metric that indicates aggregate motion of one or more of the persons 160 or objects 170 in the environment 100B. For example, the input metrics 212 can include one or more distance metrics that each indicates distance traveled by of one or more of the persons 160 or objects 170 in the environment 100B during a given phase. For example, the input metrics 212 can include a task metric that indicates acurrent task of a medical procedure. For example, the input metrics 212 can include a manual input metric that indicates an instruction for changing a given location, orientation or configuration of a given sensor or sensor system as discussed herein. For example, the sensor mode scheduling system 210 can provide instructions to one or more sensors of the first sensory system 140 or the second sensor system 150.
[0048] The environment processing system 220 can identify one or more characteristics of the environment 100B. For example, the environment processing system 220 can include a vision architecture as discussed herein. The environment processing system 220 can generate one or more output metrics 222. For example, the environment processing system 220 can include one or more logical or electronic devices including but not limited to integrated circuits, logic gates, flip flops, gate arrays, programmable gate arrays, and the like. One or more electrical, electronic, or like devices, or components associated with the environment processing system 220 can also be associated with, integrated with, integrable with, replaced by, supplemented by, complemented by, or the like, the data processing system 110 or any component thereof.
[0049] The output metrics 222 can be indicative of a state of a medical procedure of the environment 100B, or any component, person, or object thereof, or any combination thereof. For example, the output metrics 222 can include an activity detection metric that indicates an action being performed by one or more persons in the environment 100B. For example, the activity detection metric can indicate that a person 160, corresponding to a surgeon, is seated at the robotic device 130 and is performing a surgical task. For example, the output metrics 222 can include a reconstruction output that indicates a structure of at least a portion of the environment 100B. For example, the reconstruction output can include a three-dimensional model of at least a portion of the environment 100B during the medical procedure. For example, the output metrics 222 can include an object detection metric that indicates a state of one or more objects in the environment 100B. For example, the object detection metric can indicate that a first object 170, corresponding to a medical instrument (e.g., forceps), are located on a second object 170 corresponding to a table. For example, the environment processing system 220 can first identify one or more objects, and can subsequently identify corresponding states for one or more of the identified objects, via one or more of the object detection metrics. For example, the output metrics 222 can include a gesture detection metric that indicates a state of one or more body parts of one or more persons 160 of the environment 100B. For example, thegesture metric can indicate that a person 160, corresponding to a surgeon, is holding one or more manipulators of the robotic device 130 by one or more fingers or hands.
[0050] FIGs. 3-5 are directed to example architectures of machine learning models that can be integrated as or into the environment processing system 220. The architectures of FIGs. 3-5 can be integrated as or at least partially into the data processing system 110. For example, the elements of FIGs. 3-5 can be embodied as machine-readable instructions of a non-transitory memory that can be executed by a processor. Both the memory and processor can be components of the data processing system 110. As described herein, the field of view 142 from the first sensor system 140 can include video generated by one or more image sensors or video sensors of the first sensor system 140. The field of view 142 can further include depth data (e.g., three-dimensional point cloud data) captured by one or more depth-acquiring sensors of the first sensor system. In some examples, the depth data may be converted to a two- dimensional representation (e.g., in which depth or distance is represented by color or hue). Thus, any of the architectures of FIGs. 3-5 may be configured to process a video stream (e.g., captured by an image or video sensor of the first sensor system 140) of the medical environment, a depth video stream (e.g., depth data captured over time converted into a two- dimensional representation in the format of a video in which depth or distance is represented by color or hue), or both.
[0051] FIG. 3 depicts an example single-view vision architecture according to this disclosure. As illustrated by way of example in FIG. 3, a single-view vision architecture 300 can include at least a cascade input 302, a cascade output 304, a tokenizer 310, vision layer models 320, a temporal model 340, and a prediction 350.
[0052] The cascade input 302 can correspond to a feedback input to the temporal model 340. For example, the cascade input 302 can be coupled with a second cascade output of a second instance of the single-view vision architecture 300. For example, the second cascade output can correspond at least partially in one or more of structure and operation to the cascade output 304. For example, the second instance of the single-view vision architecture 300 can correspond to a second view of the environment 100B distinct from the view 142 from the first sensor system 140. The cascade output 304 can correspond to a feedforward output from the temporal model 340. For example, the cascade output 304 can be coupled with a second cascade input of a third instance of the single-view vision architecture 300. For example, the second cascade input can correspond at least partially in one or more of structure and operationto the cascade input 302. For example, the third instance of the single-view vision architecture 300 can correspond to a third view of the environment 100B distinct from the view 142 from the first sensor system 140.
[0053] The tokenizer 310 can divide input into a plurality of components, where each of the components can be provided as input to a machine learning model. For example, the tokenizer 310 can divide an image into a plurality of portions each corresponding to a subset of the image. For example, the tokenizer 310 can divide a video into a plurality of images, and divide one or more of the plurality of images into corresponding pluralities of portions, where of the portions each correspond to a subset of an image of the video. Thus, the of portions each corresponding to a subset of the image can segment video data. The tokenizer 310 can include segments 312. The segments 312 can include a plurality of portions each corresponding to a subset of the image, as discussed herein.
[0054] The vision layer models 320 can include a machine learning model to identify characteristics of an image. For example, the first layer of the vision layer models 320 can receive as input one or more of the segments 312, and can generate output corresponding to identification of objects in the image at a first level of granularity. The first layer of the vision layer models 320 can provide the output to a second layer of the vision layer models 320, and the second layer can generate output corresponding to identification of objects in the image at a second level of granularity. For example, a first layer LI can have a first layer of granularity indicative of a feature within a first portion of the image than a second layer L2 that has a second layer of granularity within a second portion of the image larger than the first portion of the image. The vision layer models 320 can include a plurality of layers, and are not limited to the layers LI, L2, L3 and L4 illustrated herein by way of example. The vision layer models 320 can include features 322. The features 322 can include data indicative of the characteristics of the image at a highest level of granularity of the image. For example, the features 322 can be data structures including vectors output by the fourth layer L4. For example, the features can be indicative of one or more shapes, objects, types of shapes, or other data indicative of the environment 100B or any persons 160 or objects 170 therein.
[0055] The temporal model 340 can receive the cascade input 302 and generate the cascade output 304 and the prediction 350. For example, the temporal model 340 can receive the cascade input corresponding to an output of a second temporal model 340 of the second instance of the single-view vision architecture 300. For example, the second instance of thesingle-view vision architecture 300 receives as input an image corresponding a second frame of video from the same view of the same sensor system or sensor device as the single-view vision architecture 300. The second frame can correspond to a frame of the video that was captured at a time earlier than a time of capture of an image corresponding to the image provided as input to the single-view vision architecture 300. Thus, the temporal model 340 can modify one or more of the features 322 to provide a technical improvement to increase accuracy of the features 322 with respect to the characteristics of the image provided as input to the single-view vision architecture 300.
[0056] The prediction 350 can correspond to a prediction of a change in the environment 100B or in one or more of the sensors of the environment 100B. For example, the prediction 350 can indicate a pose of one or more of the sensors of the environment 100B. The prediction 350 can include a corresponding prediction to modify one or more of the orientation, location, or configuration of one or more sensor systems as discussed herein to maintain a predetermined level of visibility of a portion of the environment 100B associated with the corresponding sensor systems. For example, the prediction 250 can include be data structures including vectors that describe predicted change in position, location, brightness of any person 160 or object 170 in the field of view of a given sensor system. For example, brightness can correspond to the luminosity of a portion of the image corresponding to a person 160 or object 170. Thus, the prediction 350 can provide or be transformed into one or more instructions to modify the
[0057] FIG. 4A depicts an example early fusion vision architecture according to this disclosure. As illustrated by way of example in FIG. 4A, an early fusion vision architecture 400 A can include at least a cascade input 402, a cascade output 404, a second tokenizer 410, first segments 412A, second segments 414A, a segment aggregator circuit 416, vision layer models 420A, features 430A, and a prediction 440A.
[0058] The cascade input 402 can correspond to a feedback input to the temporal model 340. For example, the cascade input 402 can be coupled with a second cascade output of a second instance of the early fusion vision architecture 400A. For example, the second cascade output can correspond at least partially in one or more of structure and operation to the cascade output 402. For example, the second instance of the early fusion vision architecture 400A can correspond to a second view of the environment 100B distinct from the view 142 from the first sensor system 140. The cascade output 404 can correspond to a feedforward output from the temporal model 340. For example, the cascade output 404 can be coupled with a secondcascade input of a third instance of the early fusion vision architecture 400A. For example, the second cascade input can correspond at least partially in one or more of structure and operation to the cascade input 402. For example, the third instance of the early fusion vision architecture 400A can correspond to a third view of the environment 100B distinct from the view 142 from the first sensor system 140.
[0059] The second tokenizer 410 can correspond at least partially in one or more of structure and operation to the tokenizer 410. For example, the second tokenizer 410 can divide an image into a plurality of portions each corresponding to a subset of the image. For example, in the early fusion vision architecture 400A the tokenizer 310 can receive a first image corresponding to a frame captured by a first sensor of a sensor system having a panoramic or stereoscopic capture configuration, and the second tokenizer 410 can receive a second image corresponding to a frame captured by a second sensor of the sensor system having the panoramic or stereoscopic capture configuration. The early fusion vision architecture 400A is not limited to the number of tokenizers 310 and 410 illustrated herein by way of example. The first segments 412A can correspond at least partially in one or more of structure and operation to the segments 312, and can be output by the tokenizer 310 based on input corresponding to or including the first image. The second segments 414A can correspond at least partially in one or more of structure and operation to the segments 312, and can be output by the second tokenizer 410 based on input corresponding to or including the second image. The segment aggregator circuit 416 can combine the first segments 412A and the second segments 414A into an aggregate input that can be provided to the vision layer models 420A. For example, the segment aggregator circuit 416 can concatenate the first segments 412A and the second segments 414A. The aggregate input is not limited to the number of segments illustrated herein by way of example.
[0060] The vision layer models 420A can correspond at least partially in one or more of structure and operation to the vision layer models 320. For example, the vision layer models 420A can be configured to receive as input the aggregate input corresponding to the first segments 412A and the second segments 414. The features 430A can correspond at least partially in one or more of structure and operation to the features 322, and can be based on the aggregate input to the vision layer models 420A. The prediction 440A can correspond at least partially in one or more of structure and operation to prediction 350, and can be based on the aggregate input to the vision layer models 420A.
[0061] FIG. 4B depicts an example mid-fusion vision architecture according to this disclosure. As illustrated by way of example in FIG. 4B, a mid-fusion vision architecture 400B can include at least first segments 412B, second segments 414B, first vision layer models 420B, second vision layer models 422B, an inter-layer models 424, features 430B, and a prediction 440B.
[0062] The first segments 412B can correspond at least partially in one or more of structure and operation to the first segments 412A. The second segments 414B can correspond at least partially in one or more of structure and operation to the second segments 414A. However, the first segments 412B and the second segments 414B can remain distinct from each other and be provided as separate input respectively to the first vision layer models 420B and the second vision layer models 422B.
[0063] The first vision layer models 420B can correspond at least partially in one or more of structure and operation to the vision layer models 320, and can receive the first segments 412B as input. The second vision layer models 422B can correspond at least partially in one or more of structure and operation to the vision layer models 320, and can receive the second segments 414B as input. The first vision layer models 420B and the second vision layer models 422B can output the features 430B via a fusion of machine learning processing at higher layers. For example, fusion between the first vision layer models 420B and the second vision layer models 422B can occur at layers L3 and L4 via the inter-layer models 424. Fusion between the first vision layer models 420B and the second vision layer models 422B is not limited to the layers illustrated herein by way of example, and can include fusion between lower layers of the models 420B and 422B.
[0064] The inter-layer models 424 can communicate layer features or data indicative of output of the first vision layer models 420B and the second vision layer models 422B between the first vision layer models 420B and the second vision layer models 422B. For example, the inter-layer models 424 can distribute or swap subsets of input to the L3 layers of the first vision layer models 420B and the second vision layer models 422B, and can distribute or swap subsets of input to the L4 layers of the first vision layer models 420B and the second vision layer models 422B. The features 430B can correspond at least partially in one or more of structure and operation to the features 322, and can be based on the inter-layer models 424. Thus, the first vision layer models 420B and the second vision layer models 422B can generate fused features using the inter-layer models 424. The prediction 440B can correspond at least partiallyin one or more of structure and operation to the prediction 350, and can be based on the features 430B.
[0065] FIG. 4C depicts an example late-fusion vision architecture according to this disclosure. As illustrated by way of example in FIG. 4C, a late-fusion vision architecture 400C can include at least features 430C, and a prediction 440C. The features 430C can correspond at least partially in one or more of structure and operation to the features 322, and can be based on an aggregation of the features output by the first vision layer models 420B and the second vision layer models 422B as discussed herein. For example, the features 430C can correspond to a concatenation of first features output by the first vision layer models 420B and second features output by the second vision layer models 422B Thus, the first vision layer models 420B and the second vision layer models 422B can generate fused features in the absence of the interlayer models 424. The prediction 440C can correspond at least partially in one or more of structure and operation to the prediction 350, and can be based on the features 430C.
[0066] FIG. 5 depicts an example second-stage fusion vision architecture according to this disclosure. As illustrated by way of example in FIG. 5, a second-stage fusion vision architecture 500 can include at least first vision layer models 510, second vision layer models 512, features 520, features 522, and a prediction 530. The first vision layer models 510 can correspond at least partially in one or more of structure and operation to the first vision layer models 420C, and can output the features 520 to the temporal model 340. The second vision layer models 512 can correspond at least partially in one or more of structure and operation to the second vision layer models 422C, and can output the features 522 to the temporal model 340. The features 520 can correspond at least partially in one or more of structure and operation to the features 322, and can be based exclusively on the output of the first vision layer models 510. The features 522 can correspond at least partially in one or more of structure and operation to the features 322, and can be based exclusively on the output of the second vision layer models 512.
[0067] The prediction 530 can correspond at least partially in one or more of structure and operation to the prediction 350, and can be based on input received at the temporal model 340 including the features 510 and the features 520 as distinct components. For example, the temporal model 340 can receive the features 510 and the features 520, aggregate the features 510 and the features 520, and generate the prediction based on the aggregation of the features 510 and the features 520. For example, the temporal model 340 can receive the features 510and the features 520, and generate the prediction based on the features 510 and the features 520 independently.
[0068] FIG. 6 depicts an example layer model architecture according to this disclosure. As illustrated by way of example in FIG. 6, a layer model architecture 600 can include at least a first layer 610, a second layer 612, a third layer 614, a fourth layer 616, a first neural network 630, a second neural network 632, a third neural network 634, a fourth neural network 636, and a mixer 650.
[0069] The first layer 610 can correspond to an instance of a vision architecture as discussed herein. For example, the first layer 610 can correspond to a first instance of the single-view vision architecture 300, but is not limited thereto. The first layer 610 can include a first clip model 620, a first layer processor 630, and a first feature processor 640, and can provide output to a layer output 654. The first clip model 620 can correspond at least partially in one or more of structure and operation to a first instance of the tokenizer 310 or 410. The first clip model 620 can include one or more instructions to divide a video into one or more frames, and to select one or more frames corresponding to one or more timestamps or times of capture associated with those one or more frames. The first layer processor 630 can correspond at least partially in one or more of structure and operation to a first instance of the vision layer models 320, 420A-C, 422B-C, 510 or 512. For example, the first layer processor 630 can be a first instance of the vision layer models 320. For example, the first layer processor 630 can include a recursive neural network (RNN). The first feature processor 640 can correspond at least partially in one or more of structure and operation to a first instance of the temporal model 340 configured to receive one or more features from the first layer processor 630.
[0070] The second layer 612 can correspond to an instance of a vision architecture as discussed herein. For example, the second layer 612 can correspond to a second instance of the singleview vision architecture 300, but is not limited thereto. The second layer 612 can include a second clip model 622, a second layer processor 632, and a second feature processor 642. The second clip model 622 can correspond at least partially in one or more of structure and operation to the tokenizer 310 or 410. The second layer processor 632 can correspond at least partially in one or more of structure and operation to a first instance of the vision layer models 320, 420A-C, 422B-C, 510 or 512. For example, the second layer processor 632 can be a second instance of the vision layer models 320. The second feature processor 642 cancorrespond at least partially in one or more of structure and operation to the first feature processor 640.
[0071] The third layer 614 can correspond to an instance of a vision architecture as discussed herein. For example, the third layer 614 can correspond to a third instance of the single-view vision architecture 300, but is not limited thereto. The third layer 614 can include a third clip model 624, a third layer processor 634, and a third feature processor 644. The third clip model 624 can correspond at least partially in one or more of structure and operation to the tokenizer 310 or 410. The third layer processor 634 can correspond at least partially in one or more of structure and operation to a first instance of the vision layer models 320, 420A-C, 422B-C, 510 or 512. For example, the third layer processor 634 can be a third instance of the vision layer models 320. The third feature processor 644 can correspond at least partially in one or more of structure and operation to the first feature processor 640.
[0072] The fourth layer 616 can correspond to an instance of a vision architecture as discussed herein. For example, the fourth layer 616 can correspond to a fourth instance of the single-view vision architecture 300, but is not limited thereto. The fourth layer 616 can include a fourth clip model 626, a fourth layer processor 636, and a fourth feature processor 646. The fourth clip model 626 can correspond at least partially in one or more of structure and operation to the tokenizer 310 or 410. The fourth layer processor 636 can correspond at least partially in one or more of structure and operation to a first instance of the vision layer models 320, 420A-C, 422B-C, 510 or 512. For example, the fourth layer processor 636 can be a fourth instance of the vision layer models 320. The fourth feature processor 646 can correspond at least partially in one or more of structure and operation to the first feature processor 640.
[0073] The mixer 650 can aggregate output from each of the first, second, third, and fourth layers 610, 612, 614 and 616. Thus, the mixer 650 can provide a fused output 652 based on predictions output by each of the first, second, third, and fourth layers 610, 612, 614 and 616. The layer output 654 can correspond to an output of the first layer 610. For example, the layer output 654 can correspond to a prediction output by the first layer 630. The layer output 654 is not limited to the example illustrated herein. For example, one or more of the second, third and fourth layers 612, 614 and 616 can provide layer outputs that correspond at least partially in one or more of structure and operation to the layer output 654.
[0074] FIG. 7A depicts an example operating environment before optimization according to this disclosure. As illustrated by way of example in FIG. 7A, an operating environment before optimization 700A can include at least a first sensor system 710A, and a second sensor system 740A.
[0075] The first sensor system 710A can correspond to a sensor system having a first location. For example, the first location is behind a surgeon at the robotic system 130. The first sensor system 710A can include a first sensor 720 A, and a second sensor 730 A. For example, the first sensor 720 A and the second sensor 730 A are arranged in a stereoscopic arrangement adjacent to each other in the first sensor system 710A. The first sensor 720 A can correspond to a first camera as discussed herein. The first sensor 720A can have a first configuration. For example, the first configuration is a first bitrate or first framerate of images captured by the first sensor 720A in a first video stream. The first sensor 720A can include a field of view 722A. The field of view 722A can correspond to a first orientation. For example, the first orientation is toward the surgeon at the robotic system 130. The second sensor 730A can correspond to a second camera as discussed herein. The first sensor 720A can have the first configuration. The second sensor 730A can include a field of view 732A. The field of view 732A can correspond to the second orientation and be offset from the field of view 722A according to the stereoscopic arrangement.
[0076] The second sensor system 740A can correspond to a sensor system having a second location. For example, the second location is behind the robotic system 130. The second sensor system 740A can include a first sensor 750A, and a second sensor 760A. For example, the first sensor 750A and the second sensor 760A are arranged in a stereoscopic configuration adjacent to each other in the first sensor system 710A. The first sensor 750A can correspond to a third camera as discussed herein. The first sensor 750A can have a second configuration. For example, the second configuration is a first bitrate or first framerate of images captured by the first sensor 750A in a first video stream from a focal distance greater than that of the first sensor 720A. The first sensor 750A can include a field of view 752A. The field of view 752A can correspond to a second orientation. For example, the second orientation is toward the robotic system 130. The second sensor 760A can correspond to a fourth camera as discussed herein. The second sensor 760A can have the second configuration. The second sensor 730A can include a field of view 762A. The field of view 762A can correspond to the second orientation and be offset from the field of view 752 A according to the stereoscopic arrangement.
[0077] FIG. 7B depicts an example operating environment after parameter optimization according to this disclosure. As illustrated by way of example in FIG. 7B, an operating environment after parameter optimization 700B can include at least a first reconfigured sensor system 71 OB and a second reconfigured sensor system 740B.
[0078] The first reconfigured sensor system 71 OB can correspond at least partially in one or more of structure and operation to the first sensor system 710A, and can have a second configuration of one or more components at least partially distinct from the first configuration of one or more components of the first sensor system 710A. For example, the first reconfigured sensor system 71 OB can have one or more sensors having one or more configurations distinct from corresponding configuration of sensors of the first sensor system 710A. The first reconfigured sensor system 71 OB can include a first reconfigured sensor 720B and a second reconfigured sensor 73 OB.
[0079] The first reconfigured sensor 720B can have the second configuration and include the reconfigured field of view 722B. For example, the second configuration is a second bitrate or second framerate of images captured by the first sensor 720A in the first video stream that is lower than the first bitrate or the first framerate. For example, the first reconfigured sensor 720B can have the second configuration with the second bitrate or second framerate to accommodate a phase of surgery or pause of surgery in which surgeon activity is lower or when the surgeon is inactive. The reconfigured field of view 722B can correspond at least partially in one or more of structure and operation to the field of view 722A, and can have a depth or viewing angle distinct from the field of view 722A according to the second configuration.
[0080] The second reconfigured sensor 730B can The second reconfigured sensor 730B can have the second configuration and include the reconfigured field of view 732B. The reconfigured field of view 732B can correspond at least partially in one or more of structure and operation to the field of view 732A, and can have a depth or viewing angle distinct from the field of view 732A according to the second configuration. The second configuration is not limited to include all of the bitrate, framerate, and field of view modifications discussed herein, and is not limited to the modification discussed herein by way of example.
[0081] The second reconfigured sensor system 740B can correspond at least partially in one or more of structure and operation to the second sensor system 740A, and can have a third configuration of one or more components at least partially distinct from the first configurationof one or more components of the first sensor system 710A. For example, the second reconfigured sensor system 740B can have one or more sensors having one or more configurations distinct from corresponding configuration of sensors of the second sensor system 740A. The second reconfigured sensor system 740B can include a first reconfigured sensor 75 OB and a second reconfigured sensor 760B.
[0082] The first reconfigured sensor 750B can have the third configuration and include the reconfigured field of view 722B. For example, the third configuration is a third bitrate or third framerate of images captured by the first sensor 720A in the first video stream that is lower than the second bitrate or the second framerate. For example, the second reconfigured sensor 750B can have the third configuration with the third bitrate or third framerate to accommodate a phase of surgery or pause of surgery in which surgeon activity is lower or when the surgeon is inactive, and activity is expected to be minimal behind the robotic device 130. The reconfigured field of view 752B can correspond at least partially in one or more of structure and operation to the field of view 752A, and can have a depth or viewing angle distinct from the field of view 752 A according to the third configuration.
[0083] The second reconfigured sensor 760B can The second reconfigured sensor 760B can have the third configuration and include the reconfigured field of view 762B. The reconfigured field of view 762B can correspond at least partially in one or more of structure and operation to the field of view 762A, and can have a depth or viewing angle distinct from the field of view 762A according to the third configuration. The third configuration is not limited to include all of the bitrate, framerate, and field of view modifications discussed herein, and is not limited to the modification discussed herein by way of example. Thus, this technical solution can provide a technical improvement to reconfigure one or more sensors and sensor system during a medical procedure, including in real-time, according to a task or phase in the medical procedure as detected by a machine vision architecture.
[0084] FIG. 7C depicts an example operating environment after position optimization according to this disclosure. As illustrated by way of example in FIG. 7C, an operating environment after position optimization 700C can include at least a first reconfigured sensor system 710C and a second reconfigured sensor system 740C. The first relocated sensor system 710C can correspond at least partially in one or more of structure and operation to the first sensor system 710A, and can have a second location distinct from the first location of the first sensor system 710A. For example, the first relocated sensor system 710C can be placed at alocation at a corner of the OR to reduce obstruction by the surgeon or other persons 160 or objects 170. The first relocated sensor system 710C can include a first relocated sensor 720C and a second relocated sensor 730C. The first relocated sensor 720C and the second relocated sensor 730C can be located to the second location. The first relocated sensor 720C can have a field of view 730C that correspond at least partially in one or more of structure and operation to the field of view 730 A and can be located at the second location. The second relocated sensor 730C can have a field of view 732C that correspond at least partially in one or more of structure and operation to the field of view 732A and can be located at the second location.
[0085] FIG. 8 depicts an example method for a sensor optimization architecture for medical procedures according to this disclosure. At least the data processing system 110 or the environment processing system 220, or the layer model architecture 600, or any component thereof, can perform method 800.
[0086] At 810, the method 800 can receive video of a medical procedure. At 812, the method 800 can receive video via a sensor located about a medical environment. At 814, the method 800 can receive video that includes at least one of medical staff, a patient, a robotic system, an instrument, or the medical environment. At 820, the method 800 can identify a state. For example, a state can correspond to a medical procedure, a phase of a medical procedure, or a task being performed during a medical procedure. For example, a phase can describe or correspond to a set of tasks or describe a portion of a medical procedure. At 822, the method 800 can identify a state using a model configured to detect one or more features in the video. At 824, the method 800 can identify a state indicative of at least one of the medical staff, the patient, the robotic system, the instrument, or the medical environment. For example, a layer model architecture as discussed herein can identify pose, position, or movement of one or more of the persons 160, the objects 170, and the robotic system 130 in the environment 100B. At 826, the method 800 can identify a state based on multi-modal data include at least one oof depth data or a 2D representation of depth data.
[0087] At 830, the method 800 can determine a configuration for the sensor. At 832, the method 800 can determine the configuration according to a parameter based at least on the state. For example, a parameter can correspond to a configuration or a portion of a configuration as discussed herein. For example, a parameter can be a bit rate, frame rate, focal length, or brightness sensitivity. For example, the method 800 can determine a parameter that corresponds to a configuration of the sensor. For example, a parameter for lower bit rate cancorrespond to a configuration for lower bit rate data. At 840, the method 800 can update the sensor to operate according to the determined configuration. At 842, the method 800 can update the sensor based at least on the parameter.
[0088] FIG. 9 depicts an example method for a sensor optimization architecture for medical procedures according to this disclosure. At least the data processing system 110 or the environment processing system 220, or the layer model architecture 600, or any component thereof, can perform method 900. At 910, the method 900 can receive video of a medical procedure. At 912, the method 900 can receive via a sensor located about a medical environment. At 914, the method 900 can receive video that includes at least one of medical staff, a patient, a robotic system, an instrument, or the medical environment. At 920, the method 900 can identify a state. At 922, the method 900 can identify a state using a model configured to detect one or more features in the video. At 924, the method 900 can identify a state indicative of at least one of the medical staff, the patient, the robotic system, the instrument, or the medical environment.
[0089] FIG. 10 depicts an example method for a sensor optimization architecture for medical procedures according to this disclosure. At least the data processing system 110 or the environment processing system 220, or the layer model architecture 600, or any component thereof, can perform method 1000. At 1010, the method 1000 can determine a parameter of the sensor. At 1012, the method 1000 can determine based at least on the state. At 1014, the method 1000 can determine a parameter that corresponds to a configuration of the sensor. At 1020, the method 1000 can update the model. At 1022, the method 1000 can update the model based at least on a loss. At 1024, the method 1000 can update the model based on loss determined with respect to the parameter. For example, the system can receive a second video via a second sensor located about the medical environment, where the second video can include at least one of the medical staff, the patient, the robotic system, the instrument, or the medical environment.
[0090] For example, the system can identify the state using the model, the model configured to detect one or more features in the second video of the medical procedure. For example, the system can update, based at least on a second parameter of the second sensor, the second sensor to correspond to a second configuration of the second sensor. For example, the system can determine, based at least on the state, the second parameter, where the second parameter corresponds to the second configuration. For example, the system can fuse the video with the second video into a fused video.
[0091] For example, the system can provide the fused video as input to a vision layer model of the model. The system can detect, using the vision layer model, one or more features in the fused video. The system can identify the state according to the one or more features in the fused video. For example, the system can provide the video as input to a first vision layer model of the model. The system can provide the second video as input to a second vision layer model of the model. For example, the system can detect, using the first vision layer model and the second vision layer model, one or more features in at least one of the video or the second video. For example, the system can identify the state according to the one or more features in at least one of the video or the second video.
[0092] For example, the system can detect, using an inter-layer model that couples the first vision layer model and the second vision layer model, the one or more features in at least one of the video or the second video. For example, the system can detect, using the first vision layer model, one or more first features in the video. The system can detect, using the second vision layer model, one or more second features the second video. For example, the system can identify the state according to at least one of the one or more first features or the second features.
[0093] For example, the system can receive a second video via a second sensor located about the medical environment, where the second video can include at least one of the medical staff, the patient, the robotic system, the instrument, or the medical environment. For example, the system can update, based at least on a second parameter of the second sensor, the model based at least on the loss determined with respect to the second parameter. For example, the system can determine, based at least on the state, the second parameter, where the second parameter corresponds to a second configuration of the second sensor.
[0094] For example, the non-transitory computer readable medium further can include one or more instructions executable by the processor. The processor can receive a second video via a second sensor located about the medical environment, where the second video can include at least one of the medical staff, the patient, the robotic system, the instrument, or the medical environment. For example, the non-transitory computer readable medium further can include one or more instructions executable by the processor. The processor can identify the state using the model, the model configured to detect one or more features in the second video of the medical procedure.
[0095] For example, the non-transitory computer readable medium further can include one or more instructions executable by the processor. The processor can update, based at least on a second parameter of the second sensor, the second sensor to correspond to a second configuration of the second sensor. For example, the non-transitory computer readable medium further can include one or more instructions executable by the processor. The processor can determine, based at least on the state, the second parameter, where the second parameter corresponds to the second configuration.
[0096] For example, the non-transitory computer readable medium further can include one or more instructions executable by the processor. The processor can fuse the video with the second video into a fused video. For example, the non-transitory computer readable medium further can include one or more instructions executable by the processor. The processor can provide the fused video as input to a vision layer model of the model. The processor can detect using the vision layer model, one or more features in the fused video. The processor can identify the state according to the one or more features in the fused video.
[0097] For example, the non-transitory computer readable medium further can include one or more instructions executable by the processor. The processor can provide the video as input to a first vision layer model of the model. The processor can provide the second video as input to a second vision layer model of the model. For example, the non-transitory computer readable medium further can include one or more instructions executable by the processor. The processor can detect using the first vision layer model and the second vision layer model, one or more features in at least one of the video or the second video. For example, the non-transitory computer readable medium further can include one or more instructions executable by the processor. The processor can identify the state according to the one or more features in at least one of the video or the second video.
[0098] For example, the non-transitory computer readable medium further can include one or more instructions executable by the processor. The processor can detect, using an inter-layer model that couples the first vision layer model and the second vision layer model, the one or more features in at least one of the video or the second video. For example, the non-transitory computer readable medium further can include one or more instructions executable by the processor. The processor can detect, using the first vision layer model, one or more first features in the video. The processor can detect, using the second vision layer model, one or more second features of the second video.
[0099] For example, the non-transitory computer readable medium further can include one or more instructions executable by the processor. The processor can identify the state according to at least one of the one or more first features or the second features. For example, the non- transitory computer readable medium further can include one or more instructions executable by the processor. The processor can receive a second video via a second sensor located about the medical environment, where the second video can include at least one of the medical staff, the patient, the robotic system, the instrument, or the medical environment. For example, the non-transitory computer readable medium further can include one or more instructions executable by the processor. The processor can update, based at least on a second parameter of the second sensor, the model based at least on the loss determined with respect to the second parameter.
[0100] For example, the non-transitory computer readable medium further can include one or more instructions executable by the processor. The processor can determine, based at least on the state, the second parameter, where the second parameter corresponds to a second configuration of the second sensor. For example, the method can include receive a second video via a second sensor located about the medical environment, where the second video can include at least one of the medical staff, the patient, the robotic system, the instrument, or the medical environment. For example, the method can include identifying the state using the model, the model configured to detect one or more features in the second video of the medical procedure. For example, the method can include updating, based at least on a second parameter of the second sensor, the second sensor to correspond to a second configuration of the second sensor. For example, the method can include determining, based at least on the state, the second parameter, where the second parameter corresponds to the second configuration.
[0101] For example, the method can include fusing the video with the second video into a fused video. For example, the method can include providing the fused video as input to a vision layer model of the model. The method can include detecting, using the vision layer model, one or more features in the fused video. The method can include identifying the state according to the one or more features in the fused video. For example, the method can include providing the video as input to a first vision layer model of the model. The method can include providing the second video as input to a second vision layer model of the model. For example, the method can include detecting, using the first vision layer model and the second vision layer model, one or more features in at least one of the video or the second video. For example, the method caninclude identifying the state according to the one or more features in at least one of the video or the second video.
[0102] For example, the method can include detecting, using an inter-layer model that couples the first vision layer model and the second vision layer model, the one or more features in at least one of the video or the second video. For example, the method can include detecting, using the first vision layer model, one or more first features in the video. The method can include detecting, using the second vision layer model, one or more second features the second video. For example, the method can include identifying the state according to at least one of the one or more first features or the second features.
[0103] For example, the method can include receiving a second video via a second sensor located about the medical environment, where the second video can include at least one of the medical staff, the patient, the robotic system, the instrument, or the medical environment. For example, the method can include updating the model, based at least on a second parameter of the second sensor, and based at least on the loss determined with respect to the second parameter. For example, the method can include determining, based at least on the state, the second parameter, where the second parameter corresponds to a second configuration of the second sensor.
[0104] Having now described some illustrative implementations, the foregoing is illustrative and not limiting, having been presented by way of example. In particular, although many of the examples presented herein involve specific combinations of method acts or system elements, those acts and those elements may be combined in other was to accomplish the same objectives. Acts, elements and features discussed in connection with one implementation are not intended to be excluded from a similar role in other implementations.
[0105] The phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. The use of "including," "comprising," "having," "containing," "involving," "characterized by," "characterized in that," and variations thereof herein, is meant to encompass the items listed thereafter, equivalents thereof, and additional items, as well as alternate implementations consisting of the items listed thereafter exclusively. In one implementation, the systems and methods described herein consist of one, each combination of more than one, or all of the described elements, acts, or components.
[0106] References to "or" may be construed as inclusive so that any terms described using "or" may indicate any of a single, more than one, and all of the described terms. References to at least one of a conjunctive list of terms may be construed as an inclusive OR to indicate any of a single, more than one, and all of the described terms. For example, a reference to "at least one of 'A' and 'B'" can include only 'A', only 'B', as well as both "A1and 'B'. Such references used in conjunction with "comprising" or other open terminology can include additional items. References to "is" or "are" may be construed as nonlimiting to the implementation or action referenced in connection with that term. The terms "is" or "are" or any tense or derivative thereof, are interchangeable and synonymous with "can be" as used herein, unless stated otherwise herein.
[0107] Directional indicators depicted herein are example directions to facilitate understanding of the examples discussed herein, and are not limited to the directional indicators depicted herein. Any directional indicator depicted herein can be modified to the reverse direction, or can be modified to include both the depicted direction and a direction reverse to the depicted direction, unless stated otherwise herein. While operations are depicted in the drawings in a particular order, such operations are not required to be performed in the particular order shown or in sequential order, and all illustrated operations are not required to be performed. Actions described herein can be performed in a different order. Where technical features in the drawings, detailed description or any claim are followed by reference signs, the reference signs have been included to increase the intelligibility of the drawings, detailed description, and claims. Accordingly, neither the reference signs nor their absence have any limiting effect on the scope of any clam elements.
[0108] Scope of the systems and methods described herein is thus indicated by the appended claims, rather than the foregoing description. The scope of the claims includes equivalents to the meaning and scope of the appended claims.
Claims
WHAT IS CLAIMED IS:
1. A system, comprising: one or more processors, coupled with memory, to: receive, via a sensor located about a medical environment, video of a medical procedure, wherein the video includes at least one of medical staff, a patient, a robotic system, instrument, or the medical environment; identify, using a model configured to detect one or more features in the video, a state indicative of at least one of the robotic system, the instrument, or the medical environment; determine, based at least on the state, a configuration for the sensor; and update the sensor to operate in accordance with the determined configuration for the sensor.
2. The system of claim 1, further comprising: causing the sensor to re-position or to re-orient in accordance with the new pose, wherein the determined configuration for the sensor comprises a new pose of the sensor within the medical environment.
3. The system of claim 1, wherein the determined configuration for the sensor comprises a sensing parameter of the sensor.
4. The system of claim 1, wherein the determined configuration for the sensor comprises a sensing modality of the sensor.
5. The system of claim 1, wherein the sensor is configured to generate three-dimensional point cloud data and wherein the video is a depth video generated using the three-dimensional point cloud data.
6. The system of claim 1, wherein the sensor is configured to generate multi-modal data including visual data and three-dimensional point cloud data, and wherein the video includes a first video generated using the visual data and a second depth video generated using the three-dimensional point cloud data.
7. The system of claim 1, the processors to: receive a second video via a second sensor located about the medical environment, wherein the second video includes at least one of the medical staff, the patient, the robotic system, the instrument, or the medical environment.
8. The system of claim 7, the processors to: identify the state using the model, the model configured to detect one or more features in the second video of the medical procedure.
9. The system of claim 7, the processors to: update, based at least on a second parameter of the second sensor, the second sensor to correspond to a second configuration of the second sensor.
10. The system of claim 9, the processors to: determine, based at least on the state, the second parameter, wherein the second parameter corresponds to the second configuration.
11. The system of claim 7, the processors to: fuse the video with the second video into a fused video.
12. The system of claim 11, the processors to: provide the fused video as input to a vision layer model of the model; detect, using the vision layer model, one or more features in the fused video; and identify the state according to the one or more features in the fused video.
13. The system of claim 7, the processors to: provide the video as input to a first vision layer model of the model; and provide the second video as input to a second vision layer model of the model.
14. The system of claim 13, the processors to: detect, using the first vision layer model and the second vision layer model, one or more features in at least one of the video or the second video.
15. The system of claim 14, the processors to: identify the state according to the one or more features in at least one of the video or the second video.
16. The system of claim 14, the processors to: detect, using an inter-layer model that couples the first vision layer model and the second vision layer model, the one or more features in at least one of the video or the second video.
17. The system of claim 13, the processors to: detect, using the first vision layer model, one or more first features in the video; and detect, using the second vision layer model, one or more second features the second video.
18. The system of claim 17, the processors to: identify the state according to at least one of the one or more first features or the second features.
19. A system, comprising: one or more processors, coupled with memory, to: receive, via a sensor located about a medical environment, video of a medical procedure, wherein the video includes at least one of medical staff, a patient, a robotic system, instrument, or the medical environment; identify, using a model configured to detect one or more features in the video, a state indicative of at least one of the robotic system, the instrument, or the medical environment; determine, based at least on the state, a parameter of the sensor, wherein the parameter corresponds to a configuration of the sensor; and update the model based at least on a loss determined with respect to the parameter.
20. The system of claim 19, the processors to: receive a second video via a second sensor located about the medical environment, wherein the second video includes at least one of the medical staff, the patient, the robotic system, the instrument, or the medical environment.
21. The system of claim 20, the processors to: update, based at least on a second parameter of the second sensor, the model based at least on the loss determined with respect to the second parameter.
22. The system of claim 21, the processors to: determine, based at least on the state, the second parameter, wherein the second parameter corresponds to a second configuration of the second sensor.
23. A non-transitory computer readable medium including one or more instructions stored thereon and executable by a processor to: receive, by the processor via a sensor located about a medical environment, video of a medical procedure, wherein the video includes at least one of medical staff, a patient, a robotic system, instrument, or the medical environment; identify, by the processor using a model configured to detect one or more features in the video, a state indicative of at least one of the robotic system, the instrument, or the medical environment; determine, by the processor and based at least on the state, a parameter of the sensor, wherein the parameter corresponds to a configuration of the sensor; and update, by the processor and based at least on the parameter, the sensor to correspond to the configuration.
24. The non-transitory computer readable medium of claim 23, the non-transitory computer readable medium further including one or more instructions executable by the processor to: receive, by the processor, a second video via a second sensor located about the medical environment, wherein the second video includes at least one of the medical staff, the patient, the robotic system, the instrument, or the medical environment.
25. The non-transitory computer readable medium of claim 24, the non-transitory computer readable medium further including one or more instructions executable by the processor to: identify, by the processor, the state using the model, the model configured to detect one or more features in the second video of the medical procedure.
26. The non-transitory computer readable medium of claim 24, the non-transitory computer readable medium further including one or more instructions executable by the processor to: update, by the processor and based at least on a second parameter of the second sensor, the second sensor to correspond to a second configuration of the second sensor.
27. The non-transitory computer readable medium of claim 26, the non-transitory computer readable medium further including one or more instructions executable by the processor to: determine, by the processor and based at least on the state, the second parameter, wherein the second parameter corresponds to the second configuration.
28. The non-transitory computer readable medium of claim 24, the non-transitory computer readable medium further including one or more instructions executable by the processor to: fuse, by the processor, the video with the second video into a fused video.
29. The non-transitory computer readable medium of claim 23, the non-transitory computer readable medium further including one or more instructions executable by the processor to: provide, by the processor, the fused video as input to a vision layer model of the model; detect, by the processor, using the vision layer model, one or more features in the fused video; and identify, by the processor, the state according to the one or more features in the fused video.
30. The non-transitory computer readable medium of claim 24, the non-transitory computer readable medium further including one or more instructions executable by the processor to: provide, by the processor, the video as input to a first vision layer model of the model; and provide, by the processor, the second video as input to a second vision layer model of the model.
31. The non-transitory computer readable medium of claim 30, the non-transitory computer readable medium further including one or more instructions executable by the processor to: detect, by the processor, using the first vision layer model and the second vision layer model, one or more features in at least one of the video or the second video.
32. The non-transitory computer readable medium of claim 31, the non-transitory computer readable medium further including one or more instructions executable by the processor to: identify, by the processor, the state according to the one or more features in at least one of the video or the second video.
33. The non-transitory computer readable medium of claim 31, the non-transitory computer readable medium further including one or more instructions executable by the processor to: detect, by the processor using an inter-layer model that couples the first vision layer model and the second vision layer model, the one or more features in at least one of the video or the second video.
34. The non-transitory computer readable medium of claim 30, the non-transitory computer readable medium further including one or more instructions executable by the processor to: detect, by the processor using the first vision layer model, one or more first features in the video; and detect, by the processor using the second vision layer model, one or more second features of the second video.
35. The non-transitory computer readable medium of claim 34, the non-transitory computer readable medium further including one or more instructions executable by the processor to: identify, by the processor, the state according to at least one of the one or more first features or the second features.
36. A non-transitory computer readable medium including one or more instructions stored thereon and executable by a processor to: receive, by the processor via a sensor located about a medical environment, video of a medical procedure, wherein the video includes at least one of medical staff, a patient, a robotic system, instrument, or the medical environment; identify, by the processor using a model configured to detect one or more features in the video, a state indicative of at least one of the robotic system, the instrument, or the medical environment; determine, by the processor and based at least on the state, a parameter of the sensor, wherein the parameter corresponds to a configuration of the sensor; and update, by the processor, the model based at least on a loss determined with respect to the parameter.
37. The non-transitory computer readable medium of claim 36, the non-transitory computer readable medium further including one or more instructions executable by the processor to: receive, by the processor, a second video via a second sensor located about the medical environment, wherein the second video includes at least one of the medical staff, the patient, the robotic system, the instrument, or the medical environment.
38. The non-transitory computer readable medium of claim 37, the non-transitory computer readable medium further including one or more instructions executable by the processor to: update, by the processor and based at least on a second parameter of the second sensor, the model based at least on the loss determined with respect to the second parameter.
39. The non-transitory computer readable medium of claim 38, the non-transitory computer readable medium further including one or more instructions executable by the processor to: determine, by the processor and based at least on the state, the second parameter, wherein the second parameter corresponds to a second configuration of the second sensor.
40. A method, comprising: receiving, via a sensor located about a medical environment, video of a medical procedure, wherein the video includes at least one of medical staff, a patient, a robotic system, instrument, or the medical environment; identifying, using a model configured to detect one or more features in the video, a state indicative of at least one of the robotic system, the instrument, or the medical environment; determining, based at least on the state, a parameter of the sensor, wherein the parameter corresponds to a configuration of the sensor; and updating, based at least on the parameter, the sensor to correspond to the configuration.
41. The method of claim 40, further comprising: receive a second video via a second sensor located about the medical environment, wherein the second video includes at least one of the medical staff, the patient, the robotic system, the instrument, or the medical environment.
42. The method of claim 41, further comprising: identifying the state using the model, the model configured to detect one or more features in the second video of the medical procedure.
43. The method of claim 41, further comprising: updating, based at least on a second parameter of the second sensor, the second sensor to correspond to a second configuration of the second sensor.
44. The method of claim 43, further comprising: determining, based at least on the state, the second parameter, wherein the second parameter corresponds to the second configuration.
45. The method of claim 41, further comprising: fusing the video with the second video into a fused video.
46. The method of claim 45, further comprising: providing the fused video as input to a vision layer model of the model;detecting, using the vision layer model, one or more features in the fused video; and identifying the state according to the one or more features in the fused video.
47. The method of claim 41, further comprising: providing the video as input to a first vision layer model of the model; and providing the second video as input to a second vision layer model of the model.
48. The method of claim 47, further comprising: detecting, using the first vision layer model and the second vision layer model, one or more features in at least one of the video or the second video.
49. The method of claim 48, further comprising: identifying the state according to the one or more features in at least one of the video or the second video.
50. The method of claim 48, further comprising: detecting, using an inter-layer model that couples the first vision layer model and the second vision layer model, the one or more features in at least one of the video or the second video.
51. The method of claim 47, further comprising: detecting, using the first vision layer model, one or more first features in the video; and detecting, using the second vision layer model, one or more second features the second video.
52. The method of claim 48, further comprising: identifying the state according to at least one of the one or more first features or the second features.
53. A method, comprising: receiving, via a sensor located about a medical environment, video of a medical procedure, wherein the video includes at least one of medical staff, a patient, a robotic system, instrument, or the medical environment;identifying, using a model configured to detect one or more features in the video, a state indicative of at least one of the robotic system, the instrument, or the medical environment; determining, based at least on the state, a parameter of the sensor, wherein the parameter corresponds to a configuration of the sensor; and updating the model based at least on a loss determined with respect to the parameter.
54. The method of claim 53, further comprising: receiving a second video via a second sensor located about the medical environment, wherein the second video includes at least one of the medical staff, the patient, the robotic system, the instrument, or the medical environment.
55. The method of claim 54, further comprising: updating the model, based at least on a second parameter of the second sensor, and based at least on the loss determined with respect to the second parameter.
56. The method of claim 55, further comprising: determining, based at least on the state, the second parameter, wherein the second parameter corresponds to a second configuration of the second sensor.
Citation Information
Patent Citations
Robotic surgical systems with multi-modality imaging for performing surgical steps
US11672614B1
Method of using imaging devices in surgery
US20210307868A1
Surgical visualization and monitoring
US20220323066A1
Systems and methods for tracking objects crossing body wall for operations associated with a computer-assisted system
WO2022147074A1
AU2021210962A1