Depth-Based OR Video De-Identification Under PPE and Poor Lighting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing OR video management systems face challenges in efficiently tracking and de-identifying OR personnel due to the unreliability of RGB images under PPE and poor lighting conditions, leading to increased costs and complexity in maintaining privacy and workflow efficiency.

Innovation Solution

Utilizing depth images from depth cameras to generate 3D point clouds, apply machine-learning techniques for human-body detection, and project 3D body contours and joints onto RGB images for de-identification, while also tracking target objects like patient beds and surgical tables.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If wireless electronic tags or trackers are attached to patients, then patient tracking accuracy is improved, but system complexity and cost increase

Engineering Contradiction:
Improvetracking accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts the tracking function from physical tags attached to patients and relocates it to the video analysis system. Depth cameras capture geometric information, and machine learning algorithms automatically detect and track personnel without any physical attachments, thereby eliminating the complexity of wireless tag systems while maintaining tracking accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces the mechanical/electronic tag-based tracking system with an optical-based depth imaging system combined with machine learning. Instead of using wireless electronic tags that require power and communication hardware, the system uses depth cameras to capture 3D geometric data and processes it through algorithms, substituting a simpler optical-mechanical system for the complex electronic system.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of information

If RGB cameras are used to capture OR videos, then visual feedback is obtained, but privacy protection becomes difficult due to PII in images

Engineering Contradiction:
Improvevisual informationVSAvoidprivacy protection
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The patent applies different quality characteristics to different parts of the image processing pipeline. Depth images are processed with high precision for tracking purposes, while RGB images are used only for visual feedback with selective blurring applied only to regions containing identifiable features. This local differentiation allows full utilization of visual information where needed while protecting privacy where required.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent introduces depth image processing and machine learning-based person detection as an intermediary between raw RGB video capture and final video output. The system first detects persons in depth images, then uses this information to selectively blur regions in the RGB video, creating an intermediate processed version that maintains both visual utility and privacy protection.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If standard RGB video images are used for personnel detection, then de-identification can be performed, but detection reliability deteriorates under PPE and poor lighting

Engineering Contradiction:
Improvede-identification capabilityVSAvoidpersonnel detection reliability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent transitions from 2D RGB image analysis to 3D depth image analysis for personnel detection. By capturing and processing depth information that provides 3D geometric data about personnel positions and shapes, the system achieves reliable detection even when facial features are obscured by PPE or lighting conditions are poor, while still maintaining the ability to perform de-identification on the original RGB images.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enhances OR efficiency by providing reliable personnel and object tracking with high privacy protection, reducing costs by avoiding the need for expensive wireless tags and complex RGB-based solutions.

Implementation Method 1

Depth sensors or depth cameras are imaging devices that produce two-dimensional (2D) images by casting lights (typically in infrared wavelengths) and measuring distances of points in a scene based on the travel time or intensity of the reflected light.

Methodology Applied
Scientific EffectTime of Flight: Time of Flight

Data Source

PatentUS12354329B2Automatic de-identification of operating room (OR) videos based on depth images
Publication Date: 2025.07.08 AURIS HEALTH INC
  • US12354329B2 patent drawing
  • US12354329B2 patent drawing
  • US12354329B2 patent drawing

AI summary

Embodiments described herein provide systems and techniques for tracking and de-identifying person in a captured operating room (OR) video. In one aspect, a process for de-identifying OR personnel in an OR video begins by simultaneously receiving a color image captured by an RGB camera and a depth image captured by a depth camera installed in the vicinity of the RGB camera. The process then generates a 3D point cloud based on the depth image. Next, the process applies a human-body detector to the 3D point cloud to detect a set of 3D bodies in the 3D point cloud, wherein each detected 3D body corresponds to a detected person in the OR. The process next projects each detected 3D body into a 2D body outline in the color image to represent the same detected person in the color image. The process subsequently de-identifies the detected people in the color image based on the projected 2D body outlines.