Depth-Based OR Video De-Identification Under PPE and Poor Lighting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing OR video management systems face challenges in efficiently tracking and de-identifying OR personnel due to the unreliability of RGB images under PPE and poor lighting conditions, leading to increased costs and complexity in maintaining privacy and workflow efficiency.
Innovation Solution
Utilizing depth images from depth cameras to generate 3D point clouds, apply machine-learning techniques for human-body detection, and project 3D body contours and joints onto RGB images for de-identification, while also tracking target objects like patient beds and surgical tables.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If wireless electronic tags or trackers are attached to patients, then patient tracking accuracy is improved, but system complexity and cost increase
Solution Approach 1:
The patent extracts the tracking function from physical tags attached to patients and relocates it to the video analysis system. Depth cameras capture geometric information, and machine learning algorithms automatically detect and track personnel without any physical attachments, thereby eliminating the complexity of wireless tag systems while maintaining tracking accuracy.
Solution Approach 2:
The patent replaces the mechanical/electronic tag-based tracking system with an optical-based depth imaging system combined with machine learning. Instead of using wireless electronic tags that require power and communication hardware, the system uses depth cameras to capture 3D geometric data and processes it through algorithms, substituting a simpler optical-mechanical system for the complex electronic system.
2Loss of information
If RGB cameras are used to capture OR videos, then visual feedback is obtained, but privacy protection becomes difficult due to PII in images
Solution Approach 1:
The patent applies different quality characteristics to different parts of the image processing pipeline. Depth images are processed with high precision for tracking purposes, while RGB images are used only for visual feedback with selective blurring applied only to regions containing identifiable features. This local differentiation allows full utilization of visual information where needed while protecting privacy where required.
Solution Approach 2:
The patent introduces depth image processing and machine learning-based person detection as an intermediary between raw RGB video capture and final video output. The system first detects persons in depth images, then uses this information to selectively blur regions in the RGB video, creating an intermediate processed version that maintains both visual utility and privacy protection.
3Ease of operation
If standard RGB video images are used for personnel detection, then de-identification can be performed, but detection reliability deteriorates under PPE and poor lighting
Solution Approach 1:
The patent transitions from 2D RGB image analysis to 3D depth image analysis for personnel detection. By capturing and processing depth information that provides 3D geometric data about personnel positions and shapes, the system achieves reliable detection even when facial features are obscured by PPE or lighting conditions are poor, while still maintaining the ability to perform de-identification on the original RGB images.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enhances OR efficiency by providing reliable personnel and object tracking with high privacy protection, reducing costs by avoiding the need for expensive wireless tags and complex RGB-based solutions.
Implementation Method 1
Depth sensors or depth cameras are imaging devices that produce two-dimensional (2D) images by casting lights (typically in infrared wavelengths) and measuring distances of points in a scene based on the travel time or intensity of the reflected light.
Data Source
AI summary
Embodiments described herein provide systems and techniques for tracking and de-identifying person in a captured operating room (OR) video. In one aspect, a process for de-identifying OR personnel in an OR video begins by simultaneously receiving a color image captured by an RGB camera and a depth image captured by a depth camera installed in the vicinity of the RGB camera. The process then generates a 3D point cloud based on the depth image. Next, the process applies a human-body detector to the 3D point cloud to detect a set of 3D bodies in the 3D point cloud, wherein each detected 3D body corresponds to a detected person in the OR. The process next projects each detected 3D body into a 2D body outline in the color image to represent the same detected person in the color image. The process subsequently de-identifies the detected people in the color image based on the projected 2D body outlines.


