Anonymizing faces in surgical environments

EP4720979A1Pending Publication Date: 2026-04-08VERB SURGICAL INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-24
Publication Date
2026-04-08

AI Technical Summary

Technical Problem

Conventional face detection techniques in surgical environments are ineffective due to occlusions from Personal Protective Equipment (PPE) and equipment, as well as off-angle poses, low resolutions, and varying illuminations, leading to cumbersome manual intervention and inefficiencies in de-identifying faces in OR videos.

Method used

A 3D model is generated from multiple video streams to identify and anonymize faces by projecting 3D face mesh models into raw video streams, utilizing RGB-D data and depth maps to determine face visibility and perform realistic face replacement, even in occluded or partially visible conditions.

Benefits of technology

This method provides a robust and automated solution for de-identifying faces in surgical videos, reducing manual intervention and ensuring effective anonymization across various challenging conditions, enabling efficient processing of large video datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024055085_28112024_PF_FP_ABST
    Figure IB2024055085_28112024_PF_FP_ABST
Patent Text Reader

Abstract

A system can anonymize faces in a surgical environment. The system may generate a three-dimensional (3D) model of persons in a surgical environment based on a plurality of video streams received from a plurality of cameras. Each camera may provide a different view of the surgical environment. The 3D model may indicate positions of the persons in a global coordinate space. The system may determine first locations comprising faces of persons in the 3D model based on the positions of the persons in the global coordinate space. The system may match the first locations to second locations in each video stream. The system may then anonymize faces in each video stream based on visibility of the faces at the second locations. Other aspects are also described and claimed.
Need to check novelty before this filing date? Find Prior Art

Description

ANONYMIZING FACES IN SURGICAL ENVIRONMENTSCROSS-REFERENCE AND RELATED APPLICATIONS

[0001] This patent application claims the benefit of priority of U.S. Provisional Application No. 63 / 468,546, filed May 24, 2023, which is incorporated herein by reference in its entirety.FIELD

[0002] This invention relates generally to surgical environments, and more specifically to systems and methods for anonymizing faces in surgical environments. Other aspects are also described.BACKGROUND

[0003] Many operating rooms (ORs) have cameras installed for monitoring OR workflows. OR videos captured by the cameras can provide visual feedback from the events taking place during a surgery, and hence analyzing and mining recorded OR videos can lead to improved OR efficiency which can subsequently reduce the costs for both patients and hospitals. However, OR videos need to be de-identified first by removing Personally Identifiable Information (PII), so that the de-identified OR videos can be stored and passed to postprocessing services without exposing PII of the patients and OR personnel.

[0004] The primary sources of PII in OR videos are the faces of patients and OR staff. To de-identify captured faces in the videos, the faces must first be detected. Conventional face detection techniques are generally constructed to utilize a bottom-up approach that relies on the detection of facial features, such as the nose and eyes, to build up and infer the face locations in the videos.SUMMARY

[0005] Implementations of this disclosure include utilizing a three-dimensional (3D) model, generated from video streams corresponding to different views of a surgical environment, to identify locations of faces of persons in the video streams and anonymize them when visible. In some implementations, a system may generate a 3D model of persons in a surgical environment based on a plurality of video streams received from a plurality of cameras arranged in the surgical environment. Each camera may provide a different view of the surgical environment. The 3D model may indicate positions of the persons in a global coordinate space. The system may determine first locations including faces of persons in the3D model based on the positions of the persons in the global coordinate space. The system may then match the first locations to second locations in each video stream. The system may then anonymize faces in video streams based on identifying the second locations. Other aspects are also described and claimed.

[0006] The above summary does not include an exhaustive list of all aspects of the present disclosure. It is contemplated that the disclosure includes all systems and methods that can be practiced from all suitable combinations of the various aspects summarized above, as well as those disclosed in the Detailed Description below and particularly pointed out in the Claims section. Such combinations may have particular advantages not specifically recited in the above summary.BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Several aspects of the disclosure herein are illustrated by way of example and not by way of limitation in the figures of the accompanying drawings in which like references indicate similar elements. It should be noted that references to “an” or “one” aspect in this disclosure are not necessarily to the same aspect, and they mean at least one. Also, in the interest of conciseness and reducing the total number of figures, a given figure may be used to illustrate the features of more than one aspect of the disclosure, and not all elements in the figure may be required for a given aspect.

[0008] FIG. 1 is an example of a system for anonymizing faces in surgical environments.

[0009] FIG. 2 is an example of a 3D model generated from video streams from different cameras arranged in a surgical environment.

[0010] FIG. 3 is an example of matched locations of faces in raw video streams.

[0011] FIG. 4 is an example of anonymized faces in anonymized video streams.

[0012] FIG. 5 is a block diagram of an example internal configuration of a computer system for anonymizing faces in surgical environments.

[0013] FIG. 6 is a flowchart of an example of a process for anonymizing faces in surgical environments.DETAILED DESCRIPTION

[0014] Several aspects of the disclosure with reference to the appended drawings are now explained. Whenever the shapes, relative positions and other aspects of the parts described are not explicitly defined, the scope of the invention is not limited only to the parts shown, which are meant merely for the purpose of illustration. Also, while numerous details are set forth, it is understood that some aspects of the disclosure may be practiced without these details. In other instances, well-known circuits, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description.

[0015] It is often desirable to perform de -identification of OR videos from a surgical environment to obscure or remove the faces of patients and OR staff from the videos. However, features of people’s faces in an OR are often heavily covered by Personal Protective Equipment (PPE) such as face masks, face shields, goggles, and glasses, and can also be occluded by other OR staff and OR equipment, which can make existing face detection techniques less effective. These challenges may be exacerbated by off-angle poses of faces, backward-facing faces, small faces with low resolutions, low illuminations, and in some cases, illuminations that are too strong. Conventional automated two-dimensional (2D) anonymization systems may have difficulty when such occlusions / obstructions are present, sometimes requiring manual intervention. Further complicating this is the large quantity of video data with images to be anonymized, making manual anonymization cumbersome. What is needed is a robust and effective OR video de-identification technique without the drawbacks of existing techniques.

[0016] Implementations of this disclosure address problems such as these by utilizing a 3D model, generated from video streams corresponding to different views of a surgical environment, to identify visible locations of faces of persons in the video streams and anonymize them when visible. In some implementations, a system may generate a 3D model of persons in a surgical environment based on a plurality of video streams (raw video streams) received from a plurality of cameras arranged in the surgical environment. Each camera may provide a different view of the surgical environment. The 3D model may indicate positions of the persons in a global coordinate space, such as by utilizing a Cartesian or spherical coordinate system. The system may then determine first locations including faces of persons in the 3D model based on the positions of the persons in the global coordinate space. The system may then match the first locations to second locations in each video stream, such as by projecting 3D face mesh models to the second locations in the raw videostreams. The system may then anonymize faces in each video stream, based on based on visibility of faces at the second locations, to produce anonymized video streams in place of the raw video streams.

[0017] In some implementations, the system can determine, from multiple raw video streams corresponding to different views of cameras, a global 3D face mesh model that overlaps with each person in a scene. The system can perform a 2D anonymization of faces in the raw video streams by projecting the meshes from the 3D model back into the raw video streams from the cameras. In some cases, a face might be occluded in a particular view, for example, due to an OR light, and therefore not visible. The system can determine whether a 3D face is visible in a 2D image by searching for differences between a depth map produced by the camera and the 3D face mesh that is generated. The system can then perform anonymization of faces that are visible, including 2D rendering adjustments to obtain a more realistic face replacement. For example, the system can merge the 3D face mesh model with a background image. In some cases, the system can utilize a face replacement template which can be individualized and changed for each person, including to influence factors such as age, gender, and / or ethnicity.

[0018] In some implementations, generating a 3D model of a person may include combining 2D poses of the person from the plurality of video streams into a 3D pose. In some implementations, determining positions of persons in the global coordinate space may include regressing a human model onto detected human key points in the 3D model. In some implementations, the 3D model may include a 3D point cloud, and each video stream may include a partial point cloud registered in the global coordinate space of the 3D point cloud. In some implementations, the 3D model may be generated based on a combination of red, green, blue (RGB) images and depth maps (collectively, red, green, blue, depth (RGB-D) data) from the plurality of video streams. In some implementations, matching the first locations to the second locations may include projecting 3D face models into video streams to determine location of the faces. In some implementations, anonymizing faces at the second locations may include determining faces in the 3D model to be visible in the video streams. In some implementations, anonymizing the faces may include replacing the faces with artificially generated faces. In some implementations, anonymizing faces may include utilizing face replacement templates individualized for the persons.

[0019] Referring now to the drawings, FIG. 1 is an example of a system 100 for anonymizing faces in surgical environments. For example, it may be desirable to perform de-identification of OR videos from a surgical environment to obscure or remove faces of patients and OR staff from the videos. The system 100 may receive a plurality of video streams (e.g., raw video streams) from a plurality of cameras arranged in the surgical environment. There may be n number of cameras present in the surgical environment, providing n number of video streams, respectively, where n is an integer greater than one. For example, the system 100 may receive video stream 1 (VS1) from camera 1, video stream 2 (VS2) from camera 2, and so forth.

[0020] Each camera may provide a video stream including multi-view RGB-D images with color and depth. Further, each video stream may correspond to a different view of the surgical environment, such as different angle or perspective of the scene. For example, a first set of one or more cameras could be surgical cameras acquiring the scene from a narrower vantage points above a surgical procedure, and a second set of one or more cameras could be workflow cameras acquiring the scene from broader vantage points in the OR. In some cases, the system 100 may receive the video streams in real time from the cameras, and in other cases, the system 100 may receive the video streams from a storage device holding recordings from the cameras.

[0021] The system 100 may utilize a 3D model generator 102 to generate a 3D model of persons in the surgical environment. For example, the 3D model generator 102 may generate a 3D mesh model of persons in the environment. The 3D model generator 102 can generate the 3D model based on the video streams (e.g., VS1, VS2) from the cameras. The 3D model generator 102 can detect 2D human key points in each video stream and regress a human model in a global coordinate space. The 3D model may indicate positions of the persons in the global coordinate space, such as a Cartesian coordinate space with XYZ axes, a spherical coordinate space, or other. For example, with additional reference to FIG. 2, the 3D model generator 102 can generate a 3D model 120 of persons in a global coordinate space, such as persons 122 and 124 corresponding to OR staff.

[0022] In some implementations, the 3D model 120 may represent a point cloud with certain data points in the point cloud corresponding to the persons in the environment (e.g., persons 122 and 124). Each video stream may include a partial point cloud registered in the global coordinate space of the 3D point cloud. For example, the two or more partial point clouds from VS1, VS2 could be registered in the one global coordinate space of the 3D model 120 by minimizing a photometric reprojection error over key points on a large visual marker.

[0023] In some implementations, determining the positions of persons in the 3D model 120 may include regressing, by the generator 102, a human model (e.g., a statistical parametric human mesh model) onto detected human key points. The generator 102 may then align a face of the human model onto the key points in the 3D point cloud of the 3D model 120. Thus, the 3D model 120 may be generated based on utilizing raw video streams from cameras having different angles, perspectives, and / or views of the scene, such as VS 1 and VS2. The 3D model 120 may be generated based on a combination of RGB-D images from each video streams. For example, RGB images and depth maps / images from the cameras, or multi-view RGB-D data, may be fused into a 3D point cloud of the 3D model 120 to represent the scene in the surgical environment.

[0024] In some implementations, generating a 3D model of a person for the 3D model 120 may include combining 2D poses of a persons from the video streams into a 3D pose of the person in the 3D model 120. For example, with additional reference to FIG. 3, generating the 3D model 120 of person 122 may include combining 2D poses of person 122 from video streams VS1 and VS2 (shown in a focused area of VS1 and VS2 in FIG. 3, highlighting person 122 in different 2D poses) into a single 3D pose of person 122 (shown in the 3D model 120 of FIG. 2). In some implementations, 3D poses from the 3D model 120 can be projected back into each 2D image of the video streams as input to guide a 3D model regression (reprojection). Thus, the system 100 can combine 2D human poses from multiple views into one joint 3D human pose for each person in the scene. The system 100 can perform temporal smoothing on each 3D human-pose sequence to interpolate missing poses and / or to reduce noise.

[0025] Referring again to FIG. 1, the system 100 may then utilize a face locator 104 to determine first locations that include faces of persons in the 3D model 120 based on the positions of the persons in the global coordinate space. The face locator 104 may determine the first locations based on the determined positions of persons as constructed in the global coordinate space. For example, referring again to FIG. 2, the face locator 104 can determine first location 132, registered in the global coordinate space of the 3D model 120, as corresponding to a face of person 122. The face locator 104 can also determine another first location 134, registered in the global coordinate space of the 3D model 120, as corresponding to a face of person 124, and so forth.

[0026] The system 100 may then utilize a matching system 106 to match the first locations to second locations in each video stream (e.g., the raw video streams). The matching system106 can project 3D face models from the first locations into video streams to determine locations of the faces at the second locations. For example, with additional reference to FIG. 3, the matching system 106 can match first location 132, corresponding to a face of person 122, to second locations 142 and 144 in each video stream. To perform the matching, the matching system 106 can project a 3D face model of person 122 from first location 132 into video streams VS1 and VS2 to determine locations of a faces of person 122 at second locationsl42 and 144, respectively. The matching system 106 can similarly match first location 134, corresponding to a face of person 124 in the 3D model 120, to other second locations in the video streams VS1 and VS1. In some implementations, matching the first locations to the second locations may include comparing the face in the 3D model 120 with a face in a depth map or color image from a video stream.

[0027] The system 100 may then utilize an anonymizer 108 to anonymize faces in each video stream, based on based on visibility of faces at the second locations, to produce anonymized video streams. For example, the system 100 may produce anonymized video stream 1 (AS1, an anonymized version of VS1 corresponding to camera 1), anonymized video stream 2 (AS2, an anonymized version of VS2 corresponding to camera 2), and so forth. To perform the anonymization, the anonymizer 108 may determine a face in the 3D model 120 is visible at a second location in a video stream (e.g., not occluded by other OR staff or equipment in a view). When a face is visible, the anonymizer 108 can render a face of a human model (e.g., a statistical parametric human mesh model) in the video stream to replace the face with an artificially generated face (e.g., changing facial features, such as adding / removing glasses, changing eyes, ears, nose, hair, etc.).

[0028] For example, referring again to FIG. 3, the anonymizer 108 can anonymize the face of person 122 based on visibility of the face in video streams VS1 and VS2, at the second locations 142 and 144. With additional reference to FIG. 4, the anonymizer 108 can anonymize the face of person 122 by replacing the face with artificially generated faces 152 and 154 (e.g., adding glasses) at the second locations in anonymized video streams AS1 and AS2, respectively.

[0029] In some implementations, the system 100 can determine from the video streams (e.g., VS1 and VS2) a global 3D face mesh model that overlaps with each person in a scene. Then, the system 100 (e.g., the anonymizer 108) can perform a 2D anonymization of faces by projecting the meshes back into the video streams to produce the anonymized video streams AS1 and AS2 (reprojection). In some cases, a 3D face might be occluded in a particular view,for example, due to an OR light, and therefore not visible. The system 100 can determine whether a 3D face is visible in 2D image by searching for differences between a camera’s produced depth map and the 3D face mesh. The system 100 can also perform 2D rendering adjustments to obtain a more realistic face replacement, such as by merging the 3D face mesh model with a background image. In some implementations, the anonymizer 108 can utilize a face replacement template which can be individualized and changed for each person (e.g., to generate the artificially generated faces 152 and 154). This may enable adjusting factors such as age, gender, and / or ethnicity. In some implementations, anonymizing the faces by the anonymizer 108 may include pixelating, blurring, and / or covering faces that are visible at the second locations.

[0030] In some implementations, the system 100 may initially perform multi -view 2D face localization by utilizing 3D information. This may enable consistent anonymization across multiple camera views. In some cases, the system 100 may utilize a mesh-based anonymization to control 3D face replacements. The images anonymized by the system 100 can then enable utilization by other systems further downstream.

[0031] FIG. 5 is a block diagram of an example internal configuration of a computer system 500 for anonymizing faces in surgical environments. For example, the computer system 500 could be a client, server, computer, smartphone, PDA, laptop, or tablet computer with one or more processors embedded therein or coupled thereto, or another computing device. The computer system 500 may include different types of computer-readable media and interfaces for various other types of computer-readable media. The computer system 500 includes a bus 502, one or more processors 504, system memory 506, read-only memory (ROM) 508, a storage device 510, input devices 512, output devices 514, and a network interface 516. In some embodiments, the computer system 500 may be part of a robotic surgical system.

[0032] The bus 502 collectively represents all system, peripheral, and chipset buses that communicatively connect the numerous internal devices of the computer system 500. For example, the bus 502 can communicatively connect the one or more processors 504 with the ROM 508, system memory 506, and the storage device 510.

[0033] From the various storage / memory units, the one or more processors 504 retrieves instructions to execute and data to process in order to execute various processes described in herein, including the above-described processes of anonymizing faces in surgical environments. The one or more processors 504 can include any type of processor, including,but not limited to, a microprocessor, a graphics processing unit (GPU), a tensor processing unit (TPU), an intelligent processor unit (IPU), a digital signal processor (DSP), a field- programmable gate array (FPGA), and an application-specific integrated circuit (ASIC). The one or more processors 504 can be a single processor or a multi -core processor in different implementations, including distributed computing.

[0034] The ROM 508 stores static data and instructions that may be utilized by the one or more processors 504 and other modules of the computer system 500. The storage device 510, on the other hand, is a read-and-write memory device. The storage device 510 may be a permanent, non-volatile memory unit that stores instructions and data even when the computer system 500 is turned off. Some implementations of the subject disclosure use a mass-storage device (such as a magnetic or optical disk and its corresponding disk drive) as the storage device 510. Other implementations use a removable storage device (such as a floppy disk, flash drive, and its corresponding disk drive) as the storage device 510.

[0035] Like the storage device 510, the system memory 506 is a read-and-write memory device. However, unlike the storage device 510, the system memory 506 is a volatile read- and-write memory, such as a random access memory (RAM). The system memory 506 can store some of the instructions and data that the one or more processors 504 utilize at runtime. In some implementations, various processes described herein, including the above-described processes of anonymizing faces in surgical environments, may be stored in the system memory 506, the ROM 508, and / or the storage device 510. From these the various storage / memory units, the one or more processors 504 can retrieve instructions to execute and data to process in order to execute the processes of some implementations.

[0036] The bus 502 may also connect to input devices 512 via input device interfaces and may connect to output devices 514 via output device interfaces. The input devices 512 may enable a user to communicate information to and select commands for the computer system 500. The input devices 512 can include, for example, alphanumeric keyboards and pointing devices (also called “cursor control devices”). The output devices 514 may enable, for example, the display of images, generated by the computer system 500, such as to the user. The output devices 514 can include, for example, printers and display devices, such as cathode ray tubes (CRT) or liquid crystal displays (LCD). Some implementations include devices such as a touchscreen that functions as both input and output devices.

[0037] The bus 502 may also couple the computer system 500 to a network through a network interface 516. In this manner, the computer system 500 can be a part of a network ofcomputers (such as a local area network (“LAN”), wide area network (“WAN”), or intranet) or a network of networks (such as the Internet). Any or all components of the computer system 500 can be used in conjunction with the disclosure herein.

[0038] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software may depend upon the application and / or design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each application. Such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.

[0039] The hardware used to implement the various illustrative logic, logical blocks, modules, and / or circuits described in connection with the aspects disclosed herein may be implemented or performed with a general-purpose processor, a DSP, ASIC, FPGA, or other programmable-logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general- purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of receiver devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Alternatively, some steps or methods may be performed by circuitry that is specific to a given function.

[0040] In one or more exemplary aspects, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as one or more instructions or code on a non-transitory computer- readable storage medium. The steps of a method or algorithm disclosed herein may be embodied in processor-executable instructions that may reside on a non-transitory computer- readable. Non-transitory computer-readable or processor-readable storage media may be any storage media that may be accessed by a computer or processor. By way of example but not limitation, such non-transitory computer-readable storage media may include RAM, ROM, EEPROM, flash memory, CD-ROM or other optical disk storage, magnetic disk storage orother magnetic storage devices, or any other medium that may be used to store desired program code in the form of instructions or data structures and that may be accessed by a computer. Disk and disc, as used herein, include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks usually reproduce data magnetically, while discs may reproduce data optically with lasers. Combinations of the above are also included within the scope of non-transitory computer-readable and processor- readable media. Additionally, the operations of a method or algorithm may reside as one or any combination or set of codes and / or instructions on a non-transitory computer-readable storage medium, which may be incorporated into a computer-program product.

[0041] FIG. 6 is a flowchart of an example of a process 600 for anonymizing faces in surgical environments (e.g., OR video de -identification). The process 600 can be executed using computing devices, such as the systems, hardware, and software described with respect to FIGS. 1-5. The process 600 can be performed, for example, by executing a machine- readable program or other computer-executable instructions, such as routines, instructions, programs, or other code. The operations of the process 600 or another technique, method, process, or algorithm described in connection with the implementations disclosed herein can be implemented directly in hardware, firmware, software executed by hardware, circuitry, or a combination thereof.

[0042] For simplicity of explanation, the process 600 is depicted and described herein as a series of operations. However, the operations in accordance with this disclosure can occur in various orders and / or concurrently. Additionally, other operations not presented and described herein may be used. Furthermore, not all illustrated operations may be required to implement a technique in accordance with the disclosed subject matter.

[0043] At operation 602, a system may generate a 3D model of persons in a surgical environment based on video streams received from cameras in the surgical environment. For example, the system 100 can generate the 3D model 120 of persons in an OR based on video streams VS1 and VS2 received from cameras 1 and 2, respectively. Each camera may provide a different view of the surgical environment, such as different angles, perspectives, and / or views of the scene, based on their different positions and / or orientations in the environment. The 3D model may indicate positions of the persons in a global coordinate space, e.g., utilizing a Cartesian or spherical coordinate system. In some cases, the 3D model may include a 3D point cloud, and each video stream may include a partial point cloud registered in the global coordinate space of the 3D point cloud.

[0044] At operation 604, the system may determine first locations including faces of persons in the global coordinate space based on the positions. For example, the system 100 may determine first locations 132 and 134 which include faces of persons 122 and 124 in the 3D model 120, respectively. The system 100 may determine the first locations based on regressing a human model onto detected human key points in the 3D model and aligning faces of the human model onto detected human key points of faces in the 3D model.

[0045] At operation 606, the system may match the first locations to second locations in each video stream of the plurality of video stream. For example, the system 100 may match first locations 132 and 134 in the 3D model 120 to second locations 142 and 144 in video streams VS1 and VS2. The system 100 may perform the matching based on projecting 3D face models into video streams to determine locations of the faces, and / or by comparing a face in the 3D model with a face in a depth map / image from a video stream. In some cases, the system may utilize bounding boxes to perform the matching.

[0046] At operation 608, the system may anonymize faces in each video stream based on visibility of the faces at the second locations (e.g., at a second location, where a face is not occluded by other OR staff or equipment in a view). For example, the system 100 may anonymize faces of persons 122 and 124 in video streams VS1 and VS2 based on visibility of the faces at second locations 142 and 144. The system 100 may then produce anonymized video streams AS 1 and AS2 based on performing anonymizations of the faces that are visible. In various implementations, anonymizing the faces may include replacing the faces with artificially generated faces (e.g., changing facial features, such as adding / removing glasses, changing eyes, ears, nose, hair, etc.), pixelating, blurring, and / or covering faces, or a combinations thereof. In some implementations, anonymizing the faces may include utilizing a face replacement template that may be individualized for each person.

[0047] While certain aspects have been described and shown in the accompanying drawings, it is to be understood that such are merely illustrative of and not restrictive on the broad invention, and that the invention is not limited to the specific constructions and arrangements shown and described, since various other modifications may occur to those of ordinary skill in the art. The description is thus to be regarded as illustrative instead of limiting.

Claims

CLAIMSWhat is claimed is:

1. A method for anonymizing faces in a surgical environment, comprising: generating a three-dimensional (3D) model of persons in a surgical environment based on a plurality of video streams received from a plurality of cameras, wherein each camera provides a different view of the surgical environment, and wherein the 3D model indicates positions of the persons in a global coordinate space; determining first locations comprising faces of persons in the 3D model based on the positions of the persons global coordinate space; matching the first locations to second locations in the plurality of video streams; and anonymizing faces in the plurality of video streams based on visibility of the faces at the second locations.

2. The method of claim 1, wherein generating the 3D model of a person includes combining two-dimensional (2D) poses of the person from the plurality of video streams into a 3D pose.

3. The method of claim 1, wherein determining the positions includes regressing a human model onto detected human key points in the 3D model.

4. The method of claim 1, wherein the 3D model comprises a 3D point cloud, and wherein each video stream comprises a partial point cloud registered in the global coordinate space of the 3D point cloud.

5. The method of claim 1, wherein the 3D model is generated based on a combination of red, green, blue (RGB) images and depth maps from the plurality of video streams.

6. The method of claim 1, wherein matching the first locations to the second locations includes projecting 3D face models into video streams to determine locations of the faces.

7. The method of claim 1, wherein anonymizing the faces includes replacing the faces with artificially generated faces.

8. The method of claim 1, wherein anonymizing faces includes utilize face replacement templates individualized for the persons.

9. A system for anonymizing faces in a surgical environment, comprising: a memory; and a processor configured to execute instructions stored in the memory to: generate a three-dimensional (3D) model of persons in a surgical environment based on a plurality of video streams received from a plurality of cameras, wherein each camera provides a different view of the surgical environment, and wherein the 3D model indicates positions of the persons in a global coordinate space; determine first locations comprising faces of persons in the 3D model based on the positions of the persons global coordinate space; match the first locations to second locations in the plurality of video streams; and anonymize faces in the plurality of video streams based on visibility of the faces at the second locations.

10. The system of claim 9, wherein generating the 3D model includes generating a 3D mesh model for each person in the surgical environment.

11. The system of claim 9, wherein determining the positions includes aligning a face of a human model onto detected human key points of faces in the 3D model.

12. The system of claim 9, wherein partial point clouds from each video stream are registered in the global coordinate space.

13. The system of claim 9, wherein the 3D model is generated based on a red, green, blue, depth (RGB-D) data from images of the plurality of video streams.

14. The system of claim 9, wherein matching the first locations to the second locations includes comparing a face in the 3D model with a face in a depth map from a video stream.

15. A non-transitory computer readable medium storing instructions operable to cause one or more processors to perform operations comprising: generating a three-dimensional (3D) model of persons in a surgical environment based on a plurality of video streams received from a plurality of cameras, wherein each camera provides a different view of the surgical environment, and wherein the 3D model indicates positions of the persons in a global coordinate space; determining first locations comprising faces of persons in the 3D model based on the positions of the persons global coordinate space; matching the first locations to second locations in the plurality of video streams; and anonymizing faces in the plurality of video streams based on visibility of the faces at the second locations.

16. The non-transitory computer readable medium of claim 15, wherein generating the 3D model includes utilizing red, green, blue, depth (RGB-D) data from each camera.

17. The non-transitory computer readable medium of claim 15, wherein determining the positions includes regressing a statistical parametric human mesh model onto detected human key points in the 3D model.

18. The non-transitory computer readable medium of claim 15, wherein matching the first locations to the second locations includes projecting 3D face mesh models into video streams to determine locations of the faces.

19. The non-transitory computer readable medium of claim 15, wherein anonymizing the faces includes merging a 3D face mesh model with a background image.

20. The non-transitory computer readable medium of claim 15, wherein partial point clouds of video streams are registered into the global coordinate space by minimizing a photometric reprojection error.