Methods and systems for anonymizing original surgical procedure videos

By using machine learning-based methods to automatically detect and de-identify text and image information in surgical videos, this technology solves the problem of time-consuming and labor-intensive manual anonymization in existing technologies, and achieves efficient automatic anonymization and privacy protection for surgical videos.

CN122091114APending Publication Date: 2026-05-26VERB SURGICAL INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
VERB SURGICAL INC
Filing Date
2019-05-24
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In existing technologies, the anonymization process of surgical videos mainly relies on manual operation, which is time-consuming and labor-intensive, making it difficult to meet the large-scale video processing needs of machine learning tools, and lacking automated tools for de-identifying sensitive information.

Method used

Using a machine learning-based approach, text and image information in surgical videos, including filenames, text in video frames, out-of-bounds events, and facial images, are detected and de-identified. The machine learning model automatically identifies and blurs or edits out personally identifiable information.

Benefits of technology

It achieves efficient and automatic anonymization of surgical videos, ensuring that no personally identifiable information is present in the videos, complying with HIPAA regulations, and is suitable for various research purposes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122091114A_ABST
    Figure CN122091114A_ABST
Patent Text Reader

Abstract

This patent disclosure provides various embodiments for anonymizing raw surgical procedure videos recorded by recording devices, such as endoscopic cameras, during surgical procedures performed on a patient in an operating room (OR). In one aspect, a method for anonymizing raw surgical procedure videos recorded by recording devices within an OR is disclosed. This method can begin by receiving a set of raw surgical videos corresponding to a surgical procedure performed within the OR. The method then merges the set of raw surgical videos to generate a surgical procedure video corresponding to the surgical procedure. Next, the method detects image-based personally identifiable information embedded in a set of raw video images of the surgical procedure video. When image-based personally identifiable information is detected, the method automatically de-identifies the detected image-based personally identifiable information in the surgical procedure video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates generally to constructing surgical procedure video analysis tools, and more specifically, to systems, apparatuses, and techniques for anonymizing raw surgical procedure videos to de-identify personally identifiable information and providing anonymized surgical procedure videos for various research purposes. Background Technology

[0002] Recorded videos of medical procedures, such as surgical procedures, contain invaluable and rich information for medical education and training, for evaluating and analyzing the quality of surgical procedures and the skills of surgeons, and for improving surgical outcomes and surgeon skills. Many surgical procedures involve displaying and capturing video images of the surgical process. For example, almost all minimally invasive procedures (MIS), such as endoscopy, laparoscopy, and arthroscopy, involve the use of cameras and video images to assist surgeons. Furthermore, existing robot-assisted surgeries require capturing intraoperative video images and displaying them to the surgeon on a monitor. Therefore, for many of the aforementioned surgical procedures, such as gastric sleeve or cholecystectomy, a large number of surgical videos already exist and continue to be created due to the large number of surgical cases performed by many different surgeons from different hospitals.

[0003] The sheer volume (and ever-increasing volume) of surgical videos for specific surgical procedures makes processing and analyzing these videos a potential machine learning problem. However, raw surgical videos recorded in the operating room (OR) can contain all sorts of patient information in the form of text-based identifiers, including patient name, medical record number, age, sex, demographics, date and time of surgery, and so on. Furthermore, some surgical procedure videos may contain sensitive and private information captured within the OR, such as information written on whiteboards within the OR and the faces of surgical staff. Therefore, raw surgical procedure videos need to be anonymized to remove personally identifiable information and comply with HIPAA regulations and procedures before they can be used for various research purposes, such as building machine learning tools.

[0004] Several automated anonymization tools exist for removing text identifiers from files and for detecting and removing sensitive information from medical image files such as patient CT scans and X-rays. However, existing techniques for anonymizing sensitive information embedded in raw surgical procedure videos are typically manual, requiring a human operator to review individual videos to identify sensitive information in video frames and then manually anonymize (e.g., by removal or deletion). This manual video anonymization process is both laborious and time-consuming. In particular, building machine learning tools requires first anonymizing a large volume of raw surgical procedure videos, making manual video anonymization impractical for machine learning purposes. Unfortunately, no existing automated anonymization tools exist for anonymizing sensitive information embedded in raw surgical procedure videos. Summary of the Invention

[0005] This patent disclosure provides various embodiments for anonymizing raw surgical procedure videos recorded by recording devices, such as endoscopic cameras, during surgical procedures performed on a patient in an operating room (OR). In one aspect, a method for anonymizing raw surgical procedure videos recorded by recording devices within an OR is disclosed. This method can begin by receiving a set of raw surgical videos corresponding to a surgical procedure performed within the OR. The method then merges the set of raw surgical videos to generate a surgical procedure video corresponding to the surgical procedure. Next, the method detects image-based personally identifiable information embedded in a set of raw video images of the surgical procedure video. When image-based personally identifiable information is detected, the method automatically de-identifies the detected image-based personally identifiable information in the surgical procedure video.

[0006] In some implementations, the method merges the set of original surgical videos to generate the surgical procedure video by analyzing a set of filenames associated with the set of original surgical videos to determine the correct order relative to the surgical procedure; and then splices the set of original surgical videos together based on the determined order.

[0007] In some implementations, the method detects image-based personally identifiable information embedded in one or more video images of the surgical procedure video by detecting one or more forms of personally identifiable text that are part of the set of original video images. If a form of personally identifiable text is detected in one or more original video images within the set of original video images, the method then deidentifies the detected image-based personally identifiable information by blurring or otherwise making the detected text in the one or more original video images difficult to recognize.

[0008] In some implementations, the one or more forms of personally identifiable text also include a form of recorded text captured by the recording device used to record the set of original surgical videos.

[0009] In some embodiments, the recorded text may include text printed on one or more surgical instruments used in the surgical procedure and recorded by the recording device positioned inside the patient. The recording may also include text displayed inside the OR performing the surgical procedure, wherein the text is accidentally recorded by the recording device during an external (OOB) event when the recording device is removed from the patient.

[0010] In some implementations, the method detects recorded text printed on the one or more surgical instruments by: using a machine learning-based tool detection and recognition model to detect the surgical instruments within the one or more raw video images; and processing a portion of the one or more raw video images containing the detected surgical instruments to detect any personally identifiable text within that portion of the one or more raw video images.

[0011] In some implementations, the one or more forms of personally identifiable text also include text boxes inserted into the surgical procedure video to display information related to the surgical procedure.

[0012] In some implementations, the method detects the text box in the surgical procedure by: using a machine learning-based text box detection model to detect the text box at or near a predetermined location within the one or more original video images; and processing a portion of the one or more original video images containing the detected text box to detect any personally identifiable text within that portion of the one or more original video images.

[0013] In some implementations, the method detects image-based personally identifiable information embedded in the set of original video images of the surgical procedure video by scanning the set of original video images to detect video segments corresponding to an out-of-body (OOB) event when the recording device is removed from the patient. It should be noted that personally identifiable information may be inadvertently captured by the recording device during the OOB event. If an OOB event is detected within the set of original video images, the method deidentifies the detected image-based personally identifiable information by automatically blurring or otherwise editing each video image in the detected video segment corresponding to the detected OOB event, making it impossible to identify any personally identifiable information embedded within the detected video segment.

[0014] In some implementations, the method detects video segments corresponding to OOB events in the set of raw video images by: using a machine learning-based OOB event detection model to detect the beginning phase of the OOB event when the endoscope is being removed from the patient; using the same machine learning-based OOB event detection model to detect the end phase of the OOB event when the endoscope is being reinserted into the patient; and labeling a set of video images between the detected beginning and end phases in the set of raw video images as the video segment corresponding to the detected OOB event.

[0015] In some implementations, the method further includes training the OOB event detection model based on a set of labeled video segments of a set of OOB events extracted from actual surgical procedure videos, before using the machine learning-based OOB event detection model to detect OOB events.

[0016] In some implementations, before merging the set of original surgical videos, the method further includes processing the received set of original surgical videos to detect text-based personally identifiable information embedded in file structure data associated with the set of original surgical videos. If text-based personally identifiable information is detected, the method may further include removing or otherwise de-identifying the detected text-based personally identifiable information from the file structure data.

[0017] In some implementations, this file structure data includes: a file identifier associated with the set of original surgical videos; a folder identifier for the folder containing the set of original surgical videos; file attributes associated with the set of original surgical videos; and other metadata associated with the set of original surgical videos.

[0018] In some implementations, after deidentifying the detected image-based personally identifiable information in the surgical procedure video, the method further includes the steps of: performing random sampling within the deidentified surgical procedure video to randomly select multiple video segments within the deidentified surgical procedure video; and verifying that the randomly selected video segments do not contain any personally identifiable information.

[0019] In some implementations, after deidentifying the detected image-based personally identifiable information in the surgical procedure video, the method further includes the steps of: permanently removing the set of original surgical videos from a surgical video repository; and replacing the removed original surgical videos with the deidentified surgical procedure videos.

[0020] In some implementations, the personally identifiable information includes both patient identifiable information associated with the patient receiving the surgical procedure and surgical staff identifiable information associated with the surgical staff performing the surgical procedure. Attached Figure Description

[0021] The structure and operation of this disclosure will be understood by reviewing the following detailed description and accompanying drawings, in which similar reference numerals refer to similar components, and wherein: Figure 1 A block diagram of an exemplary original surgical video anonymization system according to some embodiments described herein is shown.

[0022] Figure 2 A flowchart illustrating an exemplary process for anonymizing raw surgical videos to de-identify personally identifiable information embedded in video images, according to some embodiments described herein, is presented.

[0023] Figure 3 A flowchart illustrating an exemplary process for detecting and removing in vitro (OOB) video segments from an original surgical video, according to some embodiments described herein, to deidentify personally identifiable information embedded in the associated OOB video image.

[0024] Figure 4 A computer system is conceptually illustrated that can be used to implement some of the embodiments of the techniques in this subject matter. Detailed Implementation

[0025] The specific embodiments listed below are intended to describe various configurations of the subject matter and are not intended to represent the only configuration in which the subject matter can be practiced. The accompanying drawings are incorporated herein and form part of the specific embodiments. The specific embodiments include particular details intended to provide a thorough understanding of the subject matter. However, the subject matter is not limited to the particular details listed herein and can be practiced without these particular details. In some cases, structures and components are shown in block diagrams to avoid obscuring the concept of the subject matter.

[0026] Throughout this specification, the terms "anonymization" and "de-identification" are used interchangeably to refer to the de-identification of personally identifiable information. Furthermore, the terms "anonymize" and "de-identify" are used interchangeably to refer to the act of de-identifying personally identifiable information. Additionally, the terms "anonymized" and "de-identified" are used interchangeably to refer to the result of de-identifying personally identifiable information.

[0027] Raw surgical videos typically include all types of personally identifiable information, including both patient-identifiable information and surgical staff-identifiable information. Patient-identifiable information (or "patient data" below) is any information that can be used to identify a patient, and may include, but is not limited to, the patient's name, date of birth (DOB), Social Security number (SSN), age, sex, address, Medical Record Number (MRN), and the time of the surgery. Surgical staff-identifiable information (or "staff data" below) is any information that can be used to identify a given surgical staff member, such as the name of the surgeon performing the procedure. The aforementioned personally identifiable information may be in text format. For example, after a surgical video is recorded, some personally identifiable information may be embedded in the metadata associated with the video file and the folder containing the surgical video file. It should be noted that any standard text data analysis techniques can be used to anonymize or de-identify the text-based patient-identifiable information associated with the recorded raw surgical video.

[0028] In some implementations, the aforementioned personally identifiable information may be in image format and embedded in several video frames within a given surgical procedure video. Image-based personally identifiable information may include text recorded in various ways during the surgical procedure. For example, text printed on surgical instruments used inside the patient may be recorded during an endoscopic procedure. Such text may identify the surgeon's name and the type of instrument being used. For example, video images may capture text such as "Dr. Hogan's scissors" or "Dr. Hogan's sutures" on the corresponding surgical instrument. Image-based personally identifiable information may also include text boxes inserted into the recorded procedure video that display identifiable information such as the surgeon's name and the hospital's name. Furthermore, image-based personally identifiable information may also include patient and / or staff data written on a whiteboard or displayed on monitors in the OR room. It should be noted that such information is often accidentally recorded during extracorporeal events of the endoscopic procedure (described in more detail below). It should be noted that personally identifiable information may also include non-textual information. In particular, non-textual personally identifiable information may include facial images of the patient and / or surgical staff. Similarly, such facial images may be inadvertently recorded during extracorporeal events in endoscopic procedures. Non-textual, personally identifiable information may also include recorded audio tracks embedded in the original surgical video.

[0029] It should be noted that each raw endoscopic video may include multiple out-of-body (OOB) events. An OOB event is generally defined as the period during which the endoscope is removed from the patient for one of a variety of reasons during a surgical procedure while the endoscopic camera continues recording, or the period before and / or immediately after a surgical procedure when the endoscope is removed from the patient while the endoscopic camera is recording. During a surgical procedure, OOB events can occur for a variety of reasons. For example, an OOB event may occur when the endoscopic lens must be cleaned. It should be noted that multiple surgical events can cause partial or complete obstruction of endoscopic vision, hindering the surgeon's observation of anatomical structures. These surgical events may include, but are not limited to: (a) the endoscopic lens is covered with blood (e.g., due to bleeding complications); (b) the endoscopic lens fogs due to condensation; and (c) the endoscopic lens is covered with cauterized tissue particles that adhere to the lens and ultimately obstruct endoscopic vision. In each of the above scenarios, the endoscopic camera needs to be removed from the body to allow for cleaning of the endoscopic lens to restore visibility or heating for condensation removal. Following cleaning and / or other necessary procedures, the endoscope camera typically requires recalibration, including performing white balance before it can be returned to the patient. It should be noted that this type of lens cleaning OOB event can take several minutes to complete. Furthermore, the initial OOB time / event can occur at the start of the surgical procedure, when the endoscope camera is turned on before insertion into the patient; and the final OOB time / event can occur at the end of the surgical procedure, when the endoscope camera has been removed from the patient after the surgical procedure has been completed and kept on for a certain period of time.

[0030] However, whenever the endoscopic camera is removed from the patient, the surgeon may inadvertently point the camera at someone in the OR (Occupational Orbit), such as the patient or surgical staff, including the surgeon themselves, resulting in facial images of one or more people in the OR being captured in the original surgical video. Furthermore, during OOB (Out-of-Body) events, the surgeon may inadvertently point the camera at the OR whiteboard displaying personally identifiable information such as the patient's name and DOB, the names of surgical staff, protocols, and the hospital name. Both the textual personally identifiable information and the facial images in the video images captured during these OOB events must be anonymized / de-identified.

[0031] The disclosed original surgical video anonymization technique can be used to detect and anonymize / de-identify each type of the aforementioned personally identifiable information, either in the form of text information embedded in video frames or in the form of facial images within video frames. For example, text information embedded in video frames may include text / dialogue panels / boxes inserted into the video frame, text printed on surgical instruments, and text information on whiteboards or monitors inside the OR accidentally captured during an OOB event; while facial images in video frames may include the faces of patients and surgical staff inside the OR accidentally captured during an OOB event. After processing a given surgical procedure video using the disclosed video anonymization technique, the given surgical video becomes completely anonymized, making it impossible to identify the patient or surgical staff from the anonymized video images.

[0032] Figure 1 A block diagram of an exemplary raw surgical video anonymization system 100 according to some embodiments described herein is shown. (As in...) Figure 1 As can be seen, the original surgical video anonymization system 100 (or "video anonymization system 100" below) includes a text data de-identification module 102, a file merging module 112, an image data de-identification module 104, a verification module 106, and an original data cleaning module 108, which are coupled to each other in the order shown.

[0033] Generally, the file merging module 112 is configured to stitch together a set of video clips / files to form a complete procedural video; the text data de-identification module 102 is configured to process the original surgical video to detect and de-identify text-based personally identifiable information embedded in file and folder identifiers; the image data de-identification module 104 is configured to process the original surgical video to detect and de-identify various types of image-based personally identifiable information embedded in the original video images; the verification module 106 is configured to ensure that the anonymized surgical video output by the image data de-identification module 104 does not contain any personally identifiable information; and the original data erasure module 108 is configured to permanently remove the original surgical video and associated file identifiers from the original surgical video repository and replace the removed original video with the de-identified video. Each component of the video anonymization system 100 will now be described in more detail.

[0034] like Figure 1As shown, the text data de-identification module 102 of the video anonymization system 100 is coupled to a surgical video repository 130, which is typically not part of the video anonymization system 100. In some embodiments, the surgical video repository 130 is a HIPAA-compliant video repository. In some embodiments, the surgical video repository 130 may temporarily store raw surgical procedure videos recorded for surgical procedures performed in an OR, wherein the surgical procedures may include open surgical procedures, endoscopic surgical procedures, or robotic surgical procedures. Therefore, the raw surgical procedure videos may include various types of raw surgical videos, including but not limited to raw open surgical videos, raw endoscopic surgical videos, and raw robotic surgical videos.

[0035] It should be noted that if a complete surgical procedure were recorded into a single video file, the video length could be several hours (e.g., 2-2.5 hours) and the file size could be multiple gigabytes (e.g., 4-5 GB). However, hospital IT departments typically impose some limitations on the actual file size because these recorded raw video files must be transferred to different storage devices. Therefore, long surgical procedures are often broken down and recorded as a set of shorter video segments, such as a set of 500-megabyte (MB) video files each. For example, a complete surgical procedure corresponding to a 4-gigabyte (GB) procedure video would be recorded as eight 500MB video files instead of a single 4-GB video. However, the original order or some time-series information of this set of recorded segments needs to be known so that the complete surgical procedure can be reconstructed later, for example, through the file merging module 112.

[0036] In the illustrated embodiment, the text data de-identification module 102 receives a set of raw surgical video files 120 from a surgical video storage library 130, corresponding to a set of recorded video clips / edits of a complete surgical procedure. The text data de-identification module 102 is configured to process the received set of raw video files to detect text-based personally identifiable information embedded in file identifiers (e.g., filenames), folder identifiers (e.g., folder names), file attributes, and other metadata associated with the set of raw surgical video files 120. The text data de-identification module 102 then removes or otherwise de-identifies the detected text-based personally identifiable information from the corresponding file identifiers, folder identifiers, and other metadata associated with the set of raw video files 120. The text data de-identification module 102 then outputs a set of partially processed raw video files 122. It should be noted that the text data de-identification module 102 can be implemented using open-source text detection and removal tools or based on conventional text detection techniques.

[0037] In some implementations, the file merging module 112 receives from the text data de-identification module 102 the set of partially processed original video files 122 corresponding to the set of recorded video clips / edits of the complete surgical procedure, and then splices the set of partially processed original video files together to recreate the partially processed complete procedure video 124 (or "complete procedure video 124"). It should be noted that in order to merge the set of original video files 122 into a single video file, the set of original video files need to have the same format. It should also be noted that for a surgical procedure comprising a set of original video clips, the set of video clips will have different filenames and are typically stored in a single folder with a folder name. Unfortunately, different recording devices often have different naming conventions: for example, some may use identifiers such as "A, B, C, D, E" to name video clips, while others may use identifiers such as "1A, 1B, 1C, 1D" to name video clips. Therefore, the file merging module 112 should be configured to analyze different naming conventions to determine the correct order of the received set of video segments associated with a particular recording device, so as to merge these segments back into the appropriate full-length procedure video. It should be noted that if a given surgical procedure consists of a single video segment, file merging will not actually occur.

[0038] exist Figure 1 In an alternative embodiment of the illustrated implementation, instead of receiving the set of partially processed original video files 122 from the text data de-identification module 102, the file merging module 112 can independently receive the set of original video files 120 from the surgical video storage 130 and subsequently merge the set of original video files to recreate the merged (i.e., complete procedure) original surgical video.

[0039] Next, the image data de-identification module 104 receives the partially processed complete procedure video 124. In some embodiments, the image data de-identification module 104 is configured to detect various types of image-based personally identifiable information embedded in the original video images / frames of the complete procedure video 124, and subsequently de-identify the detected personally identifiable information in the corresponding video images / frames. For example, in Figure 1As can be seen, the image data de-identification module 104 may include a set of data anonymization submodules, such as the image text de-identification submodule 104-1 and the OOB event removal submodule 104-2. More specifically, the image text de-identification submodule 104-1 is configured to detect various personally identifiable texts that are recorded or automatically inserted into the original video images to make them part of the original video images. For example, the recorded text may include text printed on surgical instruments used during laparoscopic or endoscopic surgical procedures. The recorded text may also include various text-based information displayed inside the OR but accidentally recorded by a laparoscopic or endoscopic camera during an OOB event, such as text written on a whiteboard inside the OR, text printed on the surgical gowns or uniforms of surgical staff, text displayed on monitors inside the OR, or other text-bearing objects inside the OR. Personally identifiable text may also include inserted text automatically inserted into the recorded procedural video, such as standard user interface (UI) text boxes / panels that display identifiable information such as the surgeon's name and the hospital's name.

[0040] After detecting such personally identifiable text in one or more video frames of the partially processed complete procedure video 124, the image text de-identification submodule 104-1 can also be configured to automatically blur or otherwise render the detected text unreadable using other special effects or techniques in the corresponding video frame. In some embodiments, for detected text boxes within video frames not considered part of the surgical video, the entire text box can be blurred or otherwise edited away. In some other embodiments, the text within the detected text boxes can be identified first, and then the identified text can be divided into sensitive text and informational text. Next, only the identified sensitive text, such as the surgeon's name, will be blurred or otherwise edited away, while the identified informational text, such as the type / name of surgical instruments, can remain unchanged in the video image.

[0041] In various implementations, the image text de-identification submodule 104-1 may include one or more machine learning models trained to detect and identify various recorded texts and text boxes within a given video image, such as text printed on surgical instruments or displayed within text boxes. Therefore, the image text de-identification submodule 104-1 may use one or more machine learning models to automatically identify different types of personally identifiable text embedded in video images of the partially processed video 124, and automatically blur / make unrecognizable or otherwise de-identify the detected text data in the corresponding video images.

[0042] For example, the image text de-identification submodule 104-1 may include a surgical tool text detection model configured to first detect surgical tools within a video image using a machine learning-based tool detection and recognition model. Next, the surgical tool text detection model further processes the detected tool images to detect any personally identifiable text within the boundaries of the detected tool images. If such text is detected, the image text de-identification submodule 104-1 is configured to blur or otherwise render the detected text difficult to identify.

[0043] For example, the image text de-identification submodule 104-1 may include a text box detection model configured to detect standard UI panels or standard dialog boxes within video images. In some embodiments, this text box detection model may be trained based on a set of video images containing inserted text panels. To prepare training data, appropriately sized boxes may be drawn around the text panels within the training images to indicate that the content inside the text boxes is not part of the surgical video. The text box detection model may then be trained based on these generated text boxes to teach the model to look for such standard UI panels in other video frames that may contain personally identifiable text that needs to be edited out. Note that a single video frame may contain more than one such standard UI panel / text box that should be detected by the model. In some embodiments, the text box detection model may also learn the potential locations of potential UI panels / text boxes within a given video frame from the training data to facilitate the detection of these standard UI panels / text boxes with higher accuracy and faster speed.

[0044] In some implementations, multiple machine learning-based image text detection models may be configured such that each detection model is used to detect a specific type of image text embedded in a video image. For example, in addition to the text box detection models described above, there may be a surgical tool text detection model configured to detect text printed on surgical tools, a whiteboard text detection model configured to detect text captured from an OR whiteboard, and a monitor text detection model configured to detect text captured from an OR monitor. In various implementations, the aforementioned image text detection models used to detect text embedded in video images may include regression models, deep neural network-based models, support vector machines, decision trees, Naive Bayes classifiers, Bayesian networks, or K-Nearest Neighbors (KNN) models. In some implementations, each of these machine learning models is built based on a convolutional neural network (CNN) architecture, a recurrent neural network (RNN) architecture, or another form of deep neural network (DNN) architecture.

[0045] Re-reference Figure 1The OOB event removal submodule 104-2 of the data de-identification module 104 is configured to detect each OOB event within the original surgical video and subsequently remove the identified OOB segments from the original surgical video. In some embodiments, the OOB event removal submodule 104-2 may include an OOB event detector configured to scan the procedure video for OOB events. For example, such an OOB event detector may include an image processing unit configured to detect when the endoscope is removed from the patient, i.e., the start of an OOB event. The image processing unit is also configured to detect when the endoscope is reinserted into the patient, i.e., the end of an OOB event. Therefore, the sequence of video frames between the detected start and end of an OOB event corresponds to the detected OOB segments within the complete procedure video. In some implementations, for each detected OOB segment, each video frame within that segment may be completely blurred or otherwise edited away (e.g., replaced with a black screen) to make it impossible to identify any personally identifiable information within the edited video frames, such as all recorded text and inserted text boxes, as well as non-textual personally identifiable information such as faces. It should be noted that typically only the video frames associated with each detected OOB segment are blurred or otherwise edited away, but the actual frames are not removed from the video, thus preserving the original timing information of the detected OOB event.

[0046] In various implementations, the OOB event removal submodule 104-2 may include a machine learning-based OOB event detection model trained to detect OOB segments within the original surgical procedure video. This OOB event detection model may be trained based on a set of labeled video segments of a set of actual OOB events extracted from the actual surgical procedure video. Alternatively, the OOB event detection model may be trained based on a set of video segments of a set of simulated OOB events extracted from the actual procedure video or a training video. For example, if the training OOB video segments correspond to a 15-second segment of the procedure video, all video frames from the 15-second segment can be labeled with the "OOB" identifier. The OOB event detection model may then be trained based on these labeled video frames to teach the model to detect and identify similar events in the original surgical procedure video.

[0047] As mentioned above, different OOB events may be caused by different reasons; for example, one reason might be a switch from a robotic procedure to a laparoscopic procedure, while another might be lens cleaning. However, there is a strong similarity between different OOB events because they typically include an initial phase of removing the endoscope camera from the patient and an ending phase of placing the endoscope camera back into the patient. Therefore, a single trained OOB event detection model can be used by the OOB event removal submodule 104-2 to detect various OOB events caused by different reasons within the complete procedure video, and then blur, mask, or otherwise de-identify each detected OOB video segment. However, in some implementations, multiple machine learning-based OOB event detection models can be configured such that each OOB event detection model is used to detect one type of OOB event caused by a specific triggering cause / event.

[0048] In various implementations, the aforementioned single or multiple OOB event detection models for detecting OOB segments in a complete surgical procedure video may include regression models, deep neural network-based models, support vector machines, decision trees, Naive Bayes classifiers, Bayesian networks, or K-nearest neighbor (KNN) models. In some implementations, each of these machine learning models is built upon a convolutional neural network (CNN) architecture, a recurrent neural network (RNN) architecture, or another form of deep neural network (DNN) architecture.

[0049] It should be noted that personally identifiable information embedded in the original video images may also include non-textual personally identifiable information. Specifically, non-textual personally identifiable information may include facial images of patients and / or surgical staff. For example, such facial images may be accidentally recorded during an OOB event. However, any facial images captured during an OOB event can be effectively removed using the OOB event removal submodule 104-2 described above. However, in some embodiments, the image data de-identification module 104 may also include a facial image removal submodule (…). Figure 1 (Not shown in the image), the face image removal submodule is configured to detect faces embedded in the original video image (e.g., using conventional face detection techniques) and subsequently blur or otherwise de-identify each detected face from the corresponding video image. In some embodiments, to de-identify the original video image, the image data de-identification module 104 may first apply the OOB event removal submodule 104-2 to the original surgical video to remove OOB events from the original surgical video. Next, the image data de-identification module 104 applies the face image removal submodule described herein to the processed video image to search for any faces in the video frames outside of the detected OOB segments, and subsequently blur any detected faces.

[0050] Although not explicitly shown, the video anonymization system 100 may include additional modules for detecting and de-identifying other types of personally identifiable information within the original surgical video that are not described by the combined text data de-identification module 102 and image data de-identification module 104. For example, the video anonymization system 100 may also include an audio data de-identification module configured to detect and de-identify audio data containing personally identifiable information, such as surgical staff communications recorded during surgical procedures. In some embodiments, the audio data de-identification module is configured to completely remove all audio tracks embedded in a given original surgical video.

[0051] As in Figure 1 As can be seen, the image data de-identification module 104 outputs the fully processed complete procedure video 126 received by the verification module 106. In some embodiments, the verification module 106 is configured to ensure that the given fully processed surgical video does indeed contain no personally identifiable information. In some embodiments, the verification module 106 is configured to perform random sampling on the fully processed surgical video 126, that is, to randomly select multiple video segments within the fully processed surgical video and verify that the randomly selected video segments do not contain any personally identifiable information.

[0052] In some implementations, instead of selecting video segments for verification using complete randomness, verification module 106 may perform strategic random sampling to select a set of video segments from the processed surgical video that have a high probability of containing personally identifiable information for verification. For example, if it is known from statistical data that the surgeon typically or almost always removes the camera for cleaning or other reasons during one or more specific steps / phases of a given surgical procedure, it becomes more predictable approximately when one or more OOB events may have occurred during the recorded procedure. Therefore, instead of randomly sampling the entire video to verify the anonymization results of image data de-identification module 104, verification module 106 may first determine one or more time periods associated with one or more specific procedural steps / phases that have a high probability of containing OOB events. Verification module 106 then selects a set of video segments within or around the determined high-probability time periods to verify that the selected video segments do not contain any personally identifiable information.

[0053] In some implementations, the verification module 106 may collaborate with a surgical stage segmentation engine configured to segment a surgical procedure video into a set of predefined stages, where each stage represents a specific stage of the associated surgical procedure for a unique and distinguishable purpose throughout the surgical procedure. Further details of the surgical stage segmentation technique based on surgical video analysis are described in the related patent application Serial No. 15 / 987,782, filed May 23, 2018, the contents of which are incorporated herein by reference.

[0054] More specifically, the fully processed procedure video 126, or even the partially processed procedure video 124, can first be sent to a stage segmentation engine configured to identify different stages of the surgical procedure video. The output from the stage segmentation engine includes a set of predefined stages and associated timing information relative to the complete procedure video. Knowing which predefined stages (or those stages) have a high probability of including an OOB event, the verification module 106 can then "amplify" each "high OOB probability" stage of the fully processed procedure video 126 and strategically select a set of video segments from these high-probability stages of the fully processed video for verification that the selected video segments do not contain any personally identifiable information.

[0055] In some implementations, the verification module 106 may also reuse the OOB event detection model described above to identify the exact segment of the fully processed procedure video 126 corresponding to the detected OOB event. Each identified OOB video segment is then directly verified (e.g., by a human operator) to determine whether the OOB video segment does not contain any personally identifiable information.

[0056] In some implementations, if the verification module 106 determines that a given sampled video segment is not entirely free of personally identifiable information, the video anonymization system 100 may be configured to reapply the image data de-identification module 104 to a portion of the fully processed procedural video 126 containing the problematic video segment in an attempt to de-identify any remaining personally identifiable information in that portion of the video. Figure 1 (Indicated by the arrow returning from verification module 106 to module 104). Alternatively, a manual de-identification process can be used to de-identify any remaining personally identifiable information within the sampled video segment that is determined to contain personally identifiable information. In some embodiments, after reapplying image data de-identification module 104 to the problematic video segment or performing manual de-identification on the problematic video segment, verification module 106 can be reapplied to the surgical video for further processing to perform another pass of the above verification operation.

[0057] Re-reference Figure 1 After verification module 106 has verified the anonymization result of the fully processed complete procedure video 126, verification module 106 outputs a de-identified surgical video 128 containing no personally identifiable information; that is, any patient or surgical staff member in the video is completely unidentifiable. Upon receiving the de-identified surgical video 128, raw data erasure module 108 can be configured to permanently remove any raw and partially processed video files, as well as any detected text-based personal identifiers. For example, raw data erasure module 108 can permanently remove the raw surgical video file 120, the partially processed raw video file 122, the partially processed complete procedure video file 124, and the fully processed complete procedure video file 126.

[0058] In some embodiments, after erasing the original and partially processed surgical videos, the original data erasure module 108 is further configured to store the de-identified surgical video 128 back into the surgical video repository 130. If the surgical video repository 130 also stores the original surgical video file 120 of the de-identified surgical video 128, the original data erasure module 108 may be configured to permanently remove any copies of the original surgical video file 120 from the surgical video repository 130. In some embodiments, the original data erasure module 108 is also configured to create a database of the de-identified surgical videos. In some embodiments, the database entries in the surgical video repository 130 associated with the de-identified surgical video 128 may store filenames, as well as other statistics and attributes of the de-identified surgical video 128 extracted during the aforementioned video anonymization process. For example, these statistics and attributes may include, but are not limited to, the type of surgical procedure, the identifier of a complete video / video clip, the identifier of a laparoscopic procedure / robotic procedure, the number of surgeons who contributed to a given procedure, the number of OOB events during the procedure, and so on.

[0059] Next, the de-identified surgical videos stored in the surgical video repository 130 can be published to or otherwise made available to clinical experts for various research purposes. For example, the de-identified surgical videos can be published to surgical video data experts who can use the de-identified surgical videos to build various machine learning-based analytics tools. More specifically, de-identified surgical videos can be used to: establish machine learning objectives to prepare surgical data for mining from surgical videos of a given surgical procedure, as described in the relevant patent application with serial number 15 / 987,782 and filed on May 23, 2018; construct machine learning-based surgical tool inventory and tool usage tracking tools, as described in the relevant patent application with serial number 16 / 129,607 and filed on September 12, 2018; or construct surgical stage segmentation tools, also described in the relevant patent application with serial number 15 / 987,782 and filed on May 23, 2018, the contents of which are incorporated herein by reference.

[0060] Figure 2 A flowchart illustrating an exemplary process 200 for anonymizing raw surgical videos to de-identify personally identifiable information embedded in the video images, according to some embodiments described herein, is presented. In one or more embodiments, the process may be omitted, repeated, and / or performed in a different order. Figure 2 One or more steps in the process. Therefore, Figure 2 The specific arrangement of the steps shown should not be construed as limiting the scope of this technology.

[0061] Process 200 may begin by receiving a set of original surgical videos corresponding to a set of recorded video clips / clips of a complete surgical procedure performed in the OR (step 202). As described above, this set of original surgical videos corresponds to a set of shorter video clips / clips of a complete surgical procedure, wherein each original surgical video in the set may be generated as a result of a maximum file size constraint (e.g., 500 MB / video). In some embodiments, process 200 may receive the set of original surgical videos from a HIPAA-compliant video repository. In these embodiments, the set of recorded original surgical videos is first transferred from the OR to a HIPAA-compliant repository for temporary storage. In other embodiments, process 200 may receive the set of original surgical videos directly from the recording device within the OR via a secure network connection, without having to retrieve stored videos from a HIPAA-compliant video repository. Note that when receiving the set of original surgical video files, process 200 may receive the set of original video files along with a folder.

[0062] Next, process 200 performs text-based de-identification on the received set of original surgical videos and folders to detect and de-identify text-based personally identifiable information associated with the original surgical videos and associated folders (step 204). As described above, the text-based de-identification operation analyzes those text data associated with the original surgical videos and folders (containing the original surgical videos) that are not embedded in the video images. More specifically, the text-based de-identification operation detects text-based personally identifiable information embedded in file identifiers (e.g., filenames), folder identifiers (e.g., folder names), file metadata such as file attributes, and other metadata associated with the set of original video files and folders. As described above, the text-based personally identifiable information may include both text-based patient data and text-based staff data, and the text-based patient data may include, but is not limited to, the patient's name, DOB, age, gender, SSN, address, and MRN embedded in the text-based data of the video files and folders.

[0063] Process 200 then removes or otherwise de-identifies the detected text-based personally identifiable information from the file identifiers, folder identifiers, and other metadata associated with the set of original video files and folders. As mentioned above, process 200 may use open-source text detection and removal tools or conventional text detection techniques in step 204. It should be noted that at the end of step 204, the set of original video files is considered to be partially processed: that is, although the text-based personally identifiable information has been de-identified, the potential personally identifiable information embedded in the video images has not yet been de-identified.

[0064] Next, process 200 stitches the partially processed video files together to recreate a complete procedure video corresponding to the full surgical procedure (step 206). As mentioned above, different recording devices often have different naming conventions. In some embodiments, merging the partially processed video files involves first analyzing the specific naming convention used by the group of video files to determine the correct order of the video segments relative to the complete surgical procedure, and then placing the partially processed video files back into the complete procedure video.

[0065] Next, process 200 performs image-based de-identification on the partially processed complete procedure video to detect and de-identify image-based personally identifiable information embedded in the corresponding video images (step 208). As described above, image-based personally identifiable information may include various texts recorded or automatically inserted into video frames so that they can be displayed together with the video images. Therefore, process 200 can use the aforementioned image text de-identification submodule 204-1 to automatically identify different types of personally identifiable text within the original video images of the complete procedure video, and automatically blur / make the detected text data embedded in the corresponding video images difficult to recognize or otherwise de-identify it.

[0066] Furthermore, for personally identifiable information captured during an OOB event, process 200 can use the aforementioned OOB event removal submodule 104-2 to automatically detect various OOB segments within the full procedure video, and subsequently blur, mask, or otherwise de-identify the identified OOB segments from the full procedure video. Additionally, process 200 can use a dedicated face image removal submodule to search for any faces not removed by the OOB event removal submodule 104-2 in video frames outside of the detected OOB segments, and blur each detected face. It should be noted that at the end of step 208, the full procedure video is considered fully processed.

[0067] Next, process 200 performs a verification operation on the fully processed procedure video to verify that the fully processed procedure video does not contain any personally identifiable information (step 210). As described above, process 200 can use the verification module 106 described above to automatically perform random sampling or strategic sampling of the fully processed procedure video.

[0068] Next, process 200 determines whether the verification operation was successful (step 212). If one or more sampled video segments are found to be incompletely devoid of personally identifiable information, process 200 may perform another automatic or manual de-identification operation for each failed video segment (step 214). After each problematic video segment has been properly processed, the verification operation may be repeated (i.e., process 200 returns to step 210). If the verification operation is determined to be successful at step 212, the video de-identification process is complete. Then, process 200 permanently removes the original surgical video and related file data from the video repository and stores the de-identified procedural video in place of the removed original surgical video (step 216).

[0069] It should be noted that while the disclosed anonymization techniques have been described within the scope of anonymizing original surgical procedure videos, these techniques can also be used to anonymize still images captured within the surgical procedure (OR) to de-identify personally identifiable information embedded in the still images. Furthermore, although some of the disclosed anonymization techniques have been described within the scope of anonymizing a set of original surgical procedure videos corresponding to long surgical procedures, these techniques can also be applied to a single surgical procedure video to de-identify personally identifiable information embedded in that single surgical procedure video.

[0070] Figure 3 A flowchart illustrating an exemplary process 300 for detecting and removing OOB video segments from an original surgical video to de-identify personally identifiable information embedded in the associated OOB video image, according to some embodiments described herein. In one or more embodiments, the process may be omitted, repeated, and / or performed in a different order. Figure 3 One or more steps in the process. Therefore, Figure 3 The specific arrangement of the steps shown should not be construed as limiting the scope of this technology.

[0071] Process 300 can begin by training an OOB event detection model (step 302). In some implementations, this OOB event detection model can be trained based on a set of labeled video clips of a set of actual OOB events extracted from a video of an actual surgical procedure. For example, if the training OOB video clips correspond to a 15-second segment of the procedure video, then all video frames from the 15-second segment can be labeled with the “OOB” identifier. Next, the OOB event detection model can be trained based on these labeled video frames to teach the model to detect and identify similar events in the original surgical video. In particular, the OOB event detection model can be trained to detect the start phase of an OOB event, i.e., a sequence of video images corresponding to an event in which an endoscope is being removed from the patient; and the end phase of an OOB event, i.e., a sequence of video images corresponding to an event in which an endoscope is being reinserted into the patient. The OOB event detection model is also trained to associate the detected start phase and the detected end phase with the same OOB event. For example, if the detected end phase of an OOB event immediately follows the detected start phase of an OOB event, and the time interval between the two detections is less than a predetermined threshold, then the two detected phases can be considered to belong to the same OOB event.

[0072] Next, process 300 applies the trained OOB event detection model to the original procedure video to detect the start phase of the OOB event (step 304). If the start phase of the OOB event is detected, process 300 then applies the trained OOB event detection model to the original procedure video to detect the end phase of the OOB event (step 306). Next, process 300 determines whether the detected start and end phases belong to the same OOB event (step 308), and if so, process 300 blurs, masks, or otherwise de-identifies all video frames between the detected start and end phases of the OOB event (step 310). However, if process 300 determines that the detected start and end phases do not belong to the same OOB event, process 300 may generate an alarm indicating that an anomaly has been detected (step 312).

[0073] Figure 4 A computer system conceptually illustrated is provided for some embodiments that can be used to implement the techniques of this subject matter. The computer system 400 may be a client, server, computer, smartphone, PDA, laptop, or tablet computer with one or more processors embedded therein or coupled thereto, or any other type of computing device. Such computer systems include various types of computer-readable media and interfaces for various other types of computer-readable media. The computer system 400 includes a bus 402, a processing unit 412, system memory 404, read-only memory (ROM) 410, persistent storage device 408, input device interface 414, output device interface 406, and network interface 416. In some embodiments, the computer system 400 is part of a robotic surgical system.

[0074] Bus 402 collectively represents all system buses, peripheral buses, and chipset buses that communicatively connect multiple internal devices of computer system 400. For example, bus 402 communicatively connects processing unit 412 to ROM 410, system memory 404, and permanent storage device 408.

[0075] Processing unit 412 retrieves instructions to be executed and data to be processed from these various memory units in order to perform the various processes described in this patent disclosure, including combining Figures 1 to 3The process described above involves anonymizing the original surgical video to de-identify personally identifiable information embedded in the corresponding file and folder identifiers and in the corresponding video images. The processing unit 412 may include any type of processor, including but not limited to microprocessors, graphics processing units (GPUs), tensor processing units (TPUs), intelligent processor units (IPUs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), and application-specific integrated circuits (ASICs). In different specific implementations, the processing unit 412 may be a single processor or a multi-core processor.

[0076] ROM 410 stores static data and instructions required by processing unit 412 and other modules of the computer system. On the other hand, persistent storage device 408 is a read-write memory device. This device is a non-volatile memory unit that stores instructions and data even when the computer system 400 is turned off. Some specific embodiments of this subject matter disclosure use mass storage devices (such as disks or optical discs and their corresponding disk drives) as persistent storage device 408.

[0077] Other embodiments use removable storage devices (such as floppy disks, flash drives, and their corresponding disk drives) as permanent storage device 408. Similar to permanent storage device 408, system memory 404 is a read-write memory device. However, unlike storage device 408, system memory 404 is volatile read-write memory, such as random access memory. System memory 404 stores some of the instructions and data required by the processor during operation. In some embodiments, various processes described in this patent disclosure (including combinations) Figures 1 to 3 The process of anonymizing the original surgical video to de-identify personally identifiable information embedded in the corresponding file and folder identifiers and in the corresponding video images is stored in system memory 404, permanent storage device 408, and / or ROM 410. Processing unit 412 retrieves instructions to be executed and data to be processed from these various memory units in order to execute some specific implementation method.

[0078] Bus 402 is also connected to input device 414 and output device 406. Input device 414 enables a user to transmit information to the computer system and select commands for the computer system. Input device 414 may include, for example, an alphanumeric keypad and a pointing device (also known as a "cursor control device"). Output device 406 enables, for example, the display of images generated by computer system 400. Output device 406 may include, for example, a printer and a display device such as a cathode ray tube (CRT) or liquid crystal display (LCD). Some embodiments include devices that function as both input and output devices, such as a touch screen.

[0079] Finally, as Figure 4 As shown, bus 402 also couples computer system 400 to a network (not shown) via network interface 416. Thus, the computer can be part of a network of computers (such as a local area network (“LAN”), wide area network (“WAN”), intranet) or a network of network groups (such as the Internet). Any or all components of computer system 400 can be used in conjunction with the disclosure of this subject matter.

[0080] The various exemplary logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed in this patent disclosure can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this disclosure.

[0081] Hardware for implementing the various exemplary logics, logic blocks, modules, and circuits described in conjunction with the aspects disclosed herein may be implemented or performed using a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic components, discrete hardware components, or any combination thereof designed to perform the functions described herein. The general-purpose processor may be a microprocessor, but in alternative embodiments, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of receiver devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Alternatively, some steps or methods may be performed by circuitry specific to a given function.

[0082] In one or more exemplary aspects, the functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored as one or more instructions or code on a non-transitory computer-readable storage medium or a non-transitory processor-readable storage medium. The steps of the methods or algorithms disclosed herein may be embodied in processor-executable instructions that may reside on a non-transitory computer-readable or processor-readable storage medium. A non-transitory computer-readable or processor-readable storage medium may be any storage medium accessible by a computer or processor. By way of example, but not limitation, such non-transitory computer-readable or processor-readable storage media may include RAM, ROM, EEPROM, flash memory, CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired program code in the form of instructions or data structures and is accessible by a computer. As used herein, magnetic disks and optical disks include compact discs (CDs), laser discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, wherein magnetic disks typically reproduce data magnetically, while optical discs utilize lasers to reproduce data optically. The combinations described above are also included within the scope of non-transitory computer-readable and processor-readable media. Additionally, the operation of a method or algorithm may reside as one or any combination or set of code and / or instructions on a non-transitory processor-readable storage medium and / or computer-readable storage medium, thereby being incorporated into a computer program product.

[0083] While this patent document contains numerous details, these details should not be construed as limiting the scope of any disclosed technology or content protected by the claims, but rather as descriptions of features that may be specific to particular embodiments of a particular technology. Certain features described in this patent document in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any sub-combination in multiple embodiments. Furthermore, while features may be described above as functioning in certain combinations and even initially so protected by the claims, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may involve sub-combinations or variations thereof.

[0084] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order shown or sequentially, or requiring the performance of all illustrated operations to achieve the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.

[0085] Only a few specific implementations and examples are described, but other specific implementations, enhancements and variations can be derived based on the content described and illustrated in this patent document.

Claims

1. A computer-implemented method for verifying that anonymized surgical procedure videos do not contain personally identifiable information, the method comprising: Receive surgical videos corresponding to surgical procedures; The personally identifiable information (PII) is removed from the surgical video to generate anonymized surgical videos; Select a set of verification video clips from the anonymized surgical videos, including: Identify one or more time periods in the anonymized surgical video that are associated with a higher probability of containing a PII; and The set of verification video segments is randomly selected from one or more time periods determined from the anonymized surgical videos; Determine whether each segment in the set of verification video clips does not contain a PII; and If so, the anonymized surgical video will be used to replace the surgical video for storage. If not, an additional PII removal step is performed on the anonymized surgical procedure video to generate an updated anonymized surgical procedure video.

2. The computer-implemented method of claim 1, wherein determining the one or more time periods includes: The anonymized surgical video is segmented into a set of video segments containing time information based on a set of predefined surgical stages. as well as Identify a subset of the predefined surgical stages that is statistically known to contain a PII, wherein The set of verification video clips is randomly selected from a subset of the set of video segments corresponding to the identified subset of the set of predefined surgical stages.

3. The computer-implemented method according to claim 2, wherein, Performing additional PII removal steps on the anonymized surgical video includes: Each video segment in the subset of video segments identified as containing randomly selected verification video segments is again determined to contain a PII; and For each identified video portion that is determined to contain a PII, a PII removal step is performed on the identified video portion to deidentify the remaining PII within the identified video portion.

4. The computer-implemented method of claim 3, wherein identifying the subset of the set of predefined surgical stages that are statistically known to contain a PII includes identifying one or more surgical stages in the set of predefined surgical stages that have a high probability of containing one or more extracorporeal events, wherein the extracorporeal events correspond to the time period during which the endoscope is outside the patient in the surgical procedure.

5. The computer-implemented method according to claim 4, wherein, The time period during which the endoscope is outside the patient's body may include: The period during which the endoscope is temporarily removed from the patient during the surgical procedure; The period of time prior to the insertion of the endoscope into the patient at the beginning of the surgical procedure; and The period of time after the endoscope is removed from the patient at the end of the surgical procedure.

6. The method according to claim 1, wherein, Selecting the set of verification video segments from the anonymized surgical videos includes performing a completely random sampling by randomly selecting the set of verification video segments from the entire anonymized surgical videos.

7. The method of claim 1, wherein performing additional PII removal steps on the anonymized surgical video comprises performing a manual PII removal procedure on each of the set of verified video segments that are determined to still contain PII.

8. The method of claim 1, wherein after performing additional PII removal steps on the anonymized surgical video, the method further comprises randomly sampling another set of verification video segments from the updated anonymized surgical video for additional verification.

9. The computer-implemented method according to claim 1, wherein, Removing the PII from the surgical video includes: Detect one or more forms of personally identifiable text within each video image of the surgical video; and When personally identifiable text is detected in a video image, it is de-identified by blurring it to make it unreadable in the video image.

10. The computer-implemented method according to claim 9, wherein, The one or more forms of personally identifiable text include: Text printed on surgical instruments captured in one or more video images in the surgical video; and Text displayed in the operating room that was accidentally captured during the surgical procedure; and Insert a text box into the surgical video to display information related to the surgical procedure.

11. An apparatus for verifying that anonymized surgical procedure videos do not contain personally identifiable information, the apparatus comprising: One or more processors; as well as A memory coupled to the one or more processors, the memory storing instructions that, when executed by the one or more processors, cause the device to perform the following operations: Receive surgical videos corresponding to surgical procedures; The personally identifiable information (PII) is removed from the surgical video to generate anonymized surgical videos; Select a set of verification video clips from the anonymized surgical videos, including: Identify one or more time periods within the anonymized surgical video that are associated with a higher probability of containing the PII; and The set of verification video segments is randomly selected from one or more time periods determined within the anonymized surgical video; Determine whether each segment in the set of verification video clips is free of PII; and If so, the anonymized surgical video will be used to replace the surgical video for storage. If not, an additional PII removal step is performed on the anonymized surgical video to generate an updated anonymized surgical video.

12. The apparatus according to claim 11, wherein, The memory also stores instructions that, when executed by the one or more processors, cause the device to determine the one or more time periods in the following manner: Based on a set of predefined surgical stages, the anonymized surgical video is segmented into a set of video segments containing time information; as well as The identifier is a subset of the predefined surgical stages that is statistically known to contain PII, wherein The set of verification video clips is randomly selected from a subset of video segments corresponding to the identified set of predefined surgical stages.

13. The apparatus according to claim 12, wherein, The memory also stores instructions that, when executed by the one or more processors, cause the device to perform additional PII removal operations on the anonymized surgical video in the following manner: The subset identifying the set of predefined video segments contains each video segment that is determined to contain a randomly selected verification video segment containing a PII; as well as For each identified video portion that is determined to contain a PII, a PII removal step is performed on the identified video portion to deidentify the remaining PII within the identified video portion.

14. The apparatus according to claim 13, wherein, The memory also stores instructions that, when executed by the one or more processors, cause the device to identify a subset of the set of predefined surgical stages that are statistically known to contain a PII by identifying one or more surgical stages within the set of predefined surgical stages that have a higher probability of containing one or more extracorporeal events, wherein the extracorporeal event corresponds to the time period during which the endoscope is outside the patient in the surgical procedure.

15. The apparatus according to claim 11, wherein, The memory also stores instructions that, when executed by the one or more processors, cause the device to select the set of verification video segments from the anonymized surgical video by randomly selecting the set of verification video segments throughout the anonymized surgical video.

16. The apparatus according to claim 11, wherein, The memory also stores instructions that, when executed by the one or more processors, cause the device to remove the PII from the surgical video in the following manner: Detect one or more forms of personally identifiable text within each video image of the surgical video; as well as When personally identifiable text is detected in a video image, it is de-identified by blurring it to make it unreadable in the video image.

17. A system for verifying that anonymized surgical procedure videos do not contain personally identifiable information, the system comprising: One or more processors; as well as A memory coupled to the one or more processors, the memory storing instructions that, when executed by the one or more processors, cause the system to: Receive anonymized surgical videos, wherein the anonymized surgical videos are obtained by removing personally identifiable information (PII) from the original surgical videos of the surgical procedure; Select a set of verification video clips from the anonymized surgical videos, including: Identify one or more time periods in the anonymized surgical video that are associated with a higher probability of containing a PII; and The set of verification video segments is randomly selected from one or more time periods determined from the anonymized surgical videos; Determine whether each segment in the set of verification video segments does not contain a PII; and If so, the anonymized surgical video will replace the original surgical video for storage. If not, an additional PII removal step is performed on the anonymized surgical video to generate an updated anonymized surgical video.

18. The system according to claim 17, wherein, The system is configured to determine the one or more time periods in the following manner: The anonymized surgical video is segmented into a set of video segments containing time information based on a predefined set of surgical stages; and Identify a subset of the predefined set of surgical stages, which are statistically known to contain PII, wherein The set of verification video segments is randomly selected from the subset of the set of video segments corresponding to the identified subset of the set of predefined surgical stages.