Method and system for anonymizing raw surgical procedure videos

By using machine learning-based methods to detect and de-identify text and image information in surgical videos, this approach solves the problem of time-consuming and labor-intensive manual anonymization in existing technologies, achieving efficient anonymization of surgical videos and preparation of machine learning data.

CN113853658BActive Publication Date: 2026-03-27AURIS HEALTH INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-05-24
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies struggle to automatically anonymize sensitive information in surgical videos, making manual anonymization time-consuming and labor-intensive, which fails to meet the needs of machine learning tools.

Method used

A machine learning-based approach is used to detect and de-identify text and image information in surgical videos, including text on surgical tools, text during out-of-bounds (OOB) events, and facial images. The machine learning model automatically identifies and blurs or edits out sensitive information.

Benefits of technology

It achieves efficient anonymization of surgical videos, ensuring that no personally identifiable information is contained in the videos, complies with HIPAA regulations, and is suitable for data preparation for machine learning tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113853658B_ABST
    Figure CN113853658B_ABST
Patent Text Reader

Abstract

The present patent disclosure provides various embodiments for anonymizing raw surgical procedure videos recorded by recording devices, such as endoscope cameras, during surgical procedures performed on patients within an operating room (OR). In one aspect, a method for anonymizing raw surgical procedure videos recorded by recording devices within an OR is disclosed. The method can begin by receiving a set of raw surgical videos corresponding to a surgical procedure performed within the OR. The method next merges the set of raw surgical videos to generate a surgical procedure video corresponding to the surgical procedure. Next, the method detects image-based personally identifiable information embedded in a set of raw video images of the surgical procedure video. Upon detecting image-based personally identifiable information, the method automatically de-identifies the detected image-based personally identifiable information in the surgical procedure video.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates generally to building surgical procedure video analysis tools, and more specifically, to systems, devices, and techniques for anonymizing raw surgical procedure videos to de-identify personal identifiable information and provide anonymized surgical procedure videos for various research purposes. BACKGROUND

[0002] Recorded videos of medical procedures, such as surgeries, contain very valuable and rich information for medical education and training, evaluating and analyzing the quality of surgeries and skills of surgeons, and for improving outcomes of surgeries and skills of surgeons. There are many surgical procedures that involve displaying and capturing video images of surgical procedures. For example, almost all minimally invasive procedures (MIS) such as endoscopy, laparoscopy, and arthroscopy involve the use of cameras and video images to assist surgeons. In addition, existing robotic-assisted surgeries require intraoperative video images to be captured and displayed on a monitor to the surgeon. Thus, for many of the above surgical procedures, e.g., gastric sleeve or cholecystectomy, a large amount of surgical videos already exist and continue to be created due to a large number of surgical cases performed by many different surgeons from different hospitals.

[0003] The simple fact of the existence of a large (and growing) amount of surgical videos of a particular surgical procedure makes it a potential machine learning problem to process and analyze surgical videos of a given procedure. However, raw surgical videos from recordings in the operating room (OR) can contain all kinds of patient information in the form of text-based identifiers, including the patient’s name, medical record number, age, gender, demographics, date and time of surgery, and so on. In addition, some surgical procedure videos can also contain sensitive and private information captured inside the OR, such as information written on whiteboards in the OR and faces of surgical staff. Therefore, before raw surgical procedure videos can be used for various research purposes such as for building machine learning tools, raw surgical procedure videos need to be anonymized so as to be free of personal identifiable information and to comply with HIPAA regulations and processes.

[0004] There are several automated anonymization tools available for removing text identifiers from files and for detecting and removing sensitive information from medical image files such as CT scans, X-rays, etc. of patients. However, existing techniques for anonymizing sensitive information buried in raw procedural videos are typically manual-based, which requires a human operator to review individual videos to identify sensitive information in video frames and then manually anonymize (e.g., by removing or redacting) the sensitive information. Manual-based video anonymization processes are both laborious and time-consuming. In particular, building a machine learning tool requires first anonymizing a large number of raw surgical procedural videos, which makes manual-based video anonymization impractical for machine learning purposes. Unfortunately, there are no existing automated anonymization tools for anonymizing sensitive information buried in raw surgical procedural videos. SUMMARY

[0005] The present patent disclosure provides various embodiments for anonymizing raw surgical procedural videos recorded by recording devices such as endoscope cameras during surgical procedures performed on patients within an operating room (OR). In one aspect, a method for anonymizing raw surgical procedural videos recorded by recording devices within an OR is disclosed. The method can begin by receiving a set of raw surgical videos corresponding to a surgical procedure performed within the OR. The method next merges the set of raw surgical videos to generate a surgical procedural video corresponding to the surgical procedure. Next, the method detects image-based personally identifiable information embedded in a set of raw video images of the surgical procedural video. When image-based personally identifiable information is detected, the method automatically de-identifies the detected image-based personally identifiable information in the surgical procedural video.

[0006] In some embodiments, the method merges the set of raw surgical videos to generate the surgical procedural video by analyzing a set of file names associated with the set of raw surgical videos to determine a correct order with respect to the surgical procedure; and subsequently stitches the set of raw surgical videos together based on the determined order.

[0007] In some embodiments, the method detects image-based personally identifiable information embedded in one or more video images of the surgical procedural video by detecting one or more forms of personally identifiable text as part of the set of raw video images. If one form of personally identifiable text is detected in one or more raw video images within the set of raw video images, the method then de-identifies the detected image-based personally identifiable information by blurring out or otherwise rendering unreadable the detected text in the one or more raw video images.

[0008] In some embodiments, the one or more forms of personally identifiable text further includes a form of recorded text captured by the recording device used to record the set of raw surgical procedure videos.

[0009] In some embodiments, the recorded text can include text printed on one or more surgical tools used in the surgical procedure and recorded by the recording device positioned inside the patient. The recorded can also include text displayed inside the OR where the surgical procedure is performed, where the text is inadvertently recorded by the recording device during an out-of-body (OOB) event when the recording device is removed from the patient.

[0010] In some embodiments, the method detects the recorded text printed on the one or more surgical tools by: using a machine learning-based tool detection and recognition model to detect a surgical tool within the one or more raw video images; and processing a portion of the one or more raw video images containing the detected surgical tool to detect any personally identifiable text within the portion of the one or more raw video images.

[0011] In some embodiments, the one or more forms of personally identifiable text further includes a text box inserted into the surgical procedure video to display surgical procedure related information.

[0012] In some embodiments, the method detects the text box in the surgical procedure by: using a machine learning-based text box detection model to detect a text box at or near a predetermined location within the one or more raw video images; and processing a portion of the one or more raw video images containing the detected text box to detect any personally identifiable text within the portion of the one or more raw video images.

[0013] In some embodiments, the method detects image-based personally identifiable information embedded in the set of raw video images of the surgical procedure video by scanning the set of raw video images to detect a video segment corresponding to an OOB event when the recording device is removed from the patient. Note that the personally identifiable information can be inadvertently captured by the recording device during the OOB event. If an OOB event is detected within the set of raw video images, the method de-identifies the detected image-based personally identifiable information by automatically blurring out or otherwise editing out each video image in the detected video segment corresponding to the detected OOB event so that any personally identifiable information embedded within the detected video segment cannot be identified.

[0014] In some embodiments, the method detects a video segment in the set of raw video images corresponding to an OOB event by: detecting, using the machine learning-based OOB event detection model, a beginning phase of an OOB event when an endoscope camera is being removed from a patient; detecting, using the machine learning-based OOB event detection model, an ending phase of the OOB event when the endoscope camera is being inserted back into the patient; and labeling a set of video images in the set of raw video images between the detected beginning phase and the detected ending phase as the video segment corresponding to the detected OOB event.

[0015] In some embodiments, prior to detecting an OOB event using the machine learning-based OOB event detection model, the method further comprises training the OOB event detection model based on a set of labeled video segments of a set of OOB events extracted from actual surgical procedure videos.

[0016] In some embodiments, prior to merging the set of raw surgical procedure videos, the method further comprises processing the received set of raw surgical procedure videos to detect text-based personally identifiable information embedded in file structure data associated with the set of raw surgical procedure videos. If text-based personally identifiable information is detected, the method can further comprise removing or otherwise de-identifying the detected text-based personally identifiable information from the file structure data.

[0017] In some embodiments, the file structure data comprises: file identifiers associated with the set of raw surgical procedure videos; folder identifiers of folders containing the set of raw surgical procedure videos; file attributes associated with the set of raw surgical procedure videos; and other metadata associated with the set of raw surgical procedure videos.

[0018] In some embodiments, after de-identifying the detected image-based personally identifiable information in the surgical procedure video, the method further comprises the steps of: performing random sampling within the de-identified surgical procedure video to randomly select a plurality of video segments within the de-identified surgical procedure video; and verifying that the randomly selected video segments are free of any personally identifiable information.

[0019] In some embodiments, after de-identifying the detected image-based personally identifiable information in the surgical procedure video, the method further comprises the steps of: permanently removing the set of raw surgical procedure videos from a surgical video repository; and replacing the removed raw surgical procedure videos with the de-identified surgical procedure video.

[0020] In some embodiments, the personally identifiable information includes both patient personally identifiable information associated with a patient receiving the surgical procedure and surgical staff personally identifiable information associated with surgical staff performing the surgical procedure. BRIEF DESCRIPTION OF DRAWINGS

[0021] The structure and operation of the present disclosure will be understood by reviewing the following detailed description in conjunction with the drawings, in which like reference characters refer to like parts, and in which:

[0022] Figure 1 A block diagram illustrating an exemplary raw surgical video anonymization system is shown, in accordance with some embodiments described herein.

[0023] Figure 2 A flowchart illustrating an exemplary process for anonymizing raw surgical video to de-identify personally identifiable information embedded in video images is presented, in accordance with some embodiments described herein.

[0024] Figure 3 A flowchart illustrating an exemplary process for detecting out-of-body (OOB) video segments from raw surgical video and removing them from the raw surgical video to de-identify personally identifiable information embedded in associated OOB video images is presented, in accordance with some embodiments described herein.

[0025] Figure 4 A computer system that can be used to implement some embodiments of the subject technology is conceptually illustrated. DETAILED DESCRIPTION

[0026] The detailed description set forth below is intended as a description of various configurations of the subject technology and is not intended to represent the only configurations in which the subject technology can be practiced. The appended drawings are incorporated herein and constitute a part of the detailed description. The detailed description includes specific details for the purpose of providing a thorough understanding of the subject technology. However, the subject technology is not limited to the specific details set forth herein and can be practiced with variations without these specific details. In some instances, structures and components are shown in block diagram form in order to avoid obscuring the concepts of the subject technology.

[0027] Throughout this specification, the terms "anonymization" and "de-identification" are used interchangeably to mean de-identification of personally identifiable information. Further, the terms "anonymize" and "de-identify" are used interchangeably to mean the act of de-identification of personally identifiable information. Further, the terms "anonymized" and "de-identified" are used interchangeably to mean the result of de-identification of personally identifiable information.

[0028] Original surgical procedure videos often include all sorts of personally identifiable information, including both patient identifiable information and surgical staff identifiable information. Patient identifiable information (or "patient data" hereinafter) is any information that can be used to identify a patient, which can include, but is not limited to, the patient's name, date of birth (DOB), social security number (SSN), age, gender, address, medical record number (MRN), and time of surgery. Surgical staff identifiable information (or "staff data" hereinafter) is any information that can be used to identify a given surgical staff member, such as the name of the surgeon performing a procedure. The above-mentioned personally identifiable information can be in textual format. For example, after a surgical procedure video is recorded, some of the personally identifiable information can be embedded in the metadata associated with the video file and the folder containing the surgical procedure video file. Note that any conventional textual data analysis techniques can be used to anonymize or de-identify the text-based patient identifiable information associated with a recorded original surgical procedure video.

[0029] In some embodiments, the above-mentioned personally identifiable information can be in image format and embedded in some of the video frames within a given surgical procedure video. Image-based personally identifiable information can include text recorded in various ways during a surgical procedure. For example, text printed on surgical tools used in a patient's body during an endoscopic procedure can be recorded. Such text can identify the name of the surgeon and the type of tool being engaged. For example, a video image can capture text such as "Dr. Hogan's scissors" or "Dr. Hogan's stapler" on the corresponding surgical tool. Image-based personally identifiable information can also include text boxes inserted into a recorded procedure video that show identifiable information such as the name of the surgeon and the name of the hospital. In addition, image-based personally identifiable information can also include patient and / or staff data written on whiteboards or displayed on monitors within the OR room. Note that such information is often inadvertently recorded during the extracorporeal events of an endoscopic procedure (described in more detail below). Note that personally identifiable information can also include non-textual information. In particular, non-textual personally identifiable information can include facial images of patients and / or surgical staff. Again, such facial images can be inadvertently recorded during the extracorporeal events of an endoscopic procedure. Non-textual personally identifiable information can also include recorded audio tracks embedded in the original surgical procedure video.

[0030] It is noted that each raw endoscopic video can include multiple out-of-body (OOB) events. OOB events are generally defined as periods of time during a surgical procedure when the endoscope is removed from the patient for one of a variety of reasons while the endoscope camera continues to record, or periods of time immediately before and / or after a surgical procedure when the endoscope is removed from the patient while the endoscope camera is recording. During a surgical procedure, OOB events can occur for a variety of reasons. For example, an OOB event would occur in the case where the endoscope lens must be cleaned. It is noted that a number of surgical events can cause partial or complete obstruction of the endoscope view, preventing the surgeon from viewing the anatomy. These surgical events can include, but are not limited to: (a) the endoscope lens is covered with blood (e.g., due to a bleeding complication); (b) the endoscope lens is fogged due to condensation; and (c) the endoscope lens is covered with tissue particles generated by cauterization, which adhere to the lens and eventually obstruct the endoscope view. In each of the above scenarios, the endoscope camera needs to be removed from the body so that the endoscope lens can be cleaned to restore visibility or warmed up for condensation removal. After cleaning and / or other necessary processing, the endoscope camera typically needs to be recalibrated, including performing a white balance before it can be put back into the patient. It is noted that this type of OOB event for lens cleaning can take several minutes to complete. In addition, an initial OOB time / event can exist at the beginning of a surgical procedure, in the case where the endoscope camera is turned on before being inserted into the patient; and a final OOB time / event can exist at the end of a surgical procedure, in the case where the endoscope camera remains on for a period of time after the surgical procedure is completed while the endoscope camera is removed from the patient.

[0031] However, each time the endoscope camera is removed from the patient, the surgeon can inadvertently point the camera at someone in the OR, such as the patient or surgical staff including the surgeon himself, so that the facial images of one or more people in the OR can be captured in the raw surgical video. In addition, during OOB events, the surgeon can accidentally point the camera at the OR whiteboard showing personally identifiable information such as the patient’s name and DOB, the names of surgical staff, the names of the procedure and hospital, and so on. Both the textual personally identifiable information and the facial images in the video images captured during these OOB events must be anonymized / de-identified.

[0032] The disclosed raw surgical procedure video anonymization techniques can be used to detect each type of the above personal identifiable information in the form of text information embedded in the video frames or in the form of face images in the video frames and anonymize / de-identify them. For example, the text information embedded in the video frames can include text / conversation panels / boxes inserted in the video frames, text printed on surgical tools, and text information on whiteboards or monitors inside the OR accidentally captured during OOB events; while the face images in the video frames can include faces of patients and surgical staff inside the OR accidentally captured during OOB events. After a given surgical procedure video is processed using the disclosed video anonymization techniques, the given surgical video becomes fully anonymized such that the identities of the patients or surgical staff cannot be identified from the anonymized video images.

[0033] Figure 1 A block diagram of an exemplary raw surgical procedure video anonymization system 100 according to some embodiments described herein is shown. As can be seen in Figure 1 The raw surgical procedure video anonymization system 100 (or “video anonymization system 100” hereinafter) includes a text data de-identification module 102, a file merging module 112, an image data de-identification module 104, a verification module 106, and a raw data purge module 108, which are coupled to each other in the order shown.

[0034] In general, the file merging module 112 is configured to stitch together a set of video clips / files into a complete procedure video; the text data de-identification module 102 is configured to process the raw surgical procedure video to detect and de-identify text-based personal identifiable information embedded in file and folder identifiers; the image data de-identification module 104 is configured to process the raw surgical procedure video to detect and de-identify various types of image-based personal identifiable information embedded in the raw video images; the verification module 106 is configured to ensure that the anonymized surgical procedure video output by the image data de-identification module 104 is indeed free of any personal identifiable information; and the raw data purge module 108 is configured to permanently remove the raw surgical procedure video and associated file identifiers from the raw surgical procedure video repository and replace the removed raw video with the de-identified video. Each component of the video anonymization system 100 is now described in more detail.

[0035] As Figure 1As shown, the textual data de-identification module 102 of the video anonymization system 100 is coupled to a surgical video repository 130, which is typically not part of the video anonymization system 100. In some embodiments, the surgical video repository 130 is a HIPAA-compliant video repository. In some embodiments, the surgical video repository 130 can temporarily store raw surgical procedure videos recorded for a surgical procedure performed in an OR, where the surgical procedure can include an open surgical procedure, an endoscopic surgical procedure, or a robotic surgical procedure. Thus, the raw surgical procedure videos can include various types of raw surgical videos, including but not limited to raw open surgical videos, raw endoscopic surgical videos, and raw robotic surgical videos.

[0036] Note that if the entire surgical procedure is recorded into a single video file, the length of the video can be several hours (e.g., 2-2.5 hours) and the file size can be multiple gigabytes (e.g., 4-5 GB). However, hospital IT departments often impose some limit on how large the actual file size can be, as these recorded raw video files must be transferred to different storage devices. Thus, long surgical procedures are often broken up and recorded as a set of shorter video segments, such as a set of 500-megabyte (MB) video files per file. For example, a complete surgical procedure corresponding to a 4-gigabyte (GB) procedure video would be recorded as eight 500-MB video files, rather than a single 4-GB video. However, the original order or some temporal sequence information of the set of recorded segments needs to be known in order to later reconstruct the complete surgical procedure, e.g., by the file merging module 112.

[0037] In the illustrated embodiment, the textual data de-identification module 102 receives from the surgical video repository 130 a set of raw surgical video files 120 corresponding to a set of recorded video segments / clips of a complete surgical procedure. The textual data de-identification module 102 is configured to process the received set of raw video files to detect text-based personally identifiable information embedded in file identifiers (e.g., file names), folder identifiers (e.g., folder names), file attributes, and other metadata associated with the set of raw surgical video files 120. The textual data de-identification module 102 then removes or otherwise de-identifies the detected text-based personally identifiable information from the corresponding file identifiers, folder identifiers, and other metadata associated with the set of raw video files 120. The textual data de-identification module 102 then outputs a set of partially processed raw video files 122. Note that the textual data de-identification module 102 can be implemented with open-source text detection and removal tools or based on conventional text detection techniques.

[0038] In some embodiments, the file merging module 112 receives the set of partially processed raw video files 122 corresponding to the set of recorded video segments / clips of the complete surgical procedure from the textual data de-identification module 102, and then stitches the set of partially processed raw video files together to recreate a partially processed complete procedure video 124 (or "complete procedure video 124"). Note that in order to be able to merge the set of raw video files 122 into a single video file, the set of raw video files need to have the same format. Also note that for a surgical procedure that includes a set of raw video segments, the set of video segments will have different file names, and are typically stored within a single folder having a folder name. Unfortunately, different recording devices often have different naming conventions: for example, some can name video segments with identifiers "A, B, C, D, E," etc., and others can name video segments with identifiers "1A, 1B, 1C, 1D," etc. Thus, the file merging module 112 should be configured to analyze the different naming conventions to determine the correct order of the received set of video segments associated with a particular recording device in order to merge the segments back into the proper full-length procedure video. Note that if a given surgical procedure consists of a single video segment, no file merging actually occurs.

[0039] In Figure 1 In an alternative embodiment to the illustrated embodiment, instead of receiving the set of partially processed raw video files 122 from the textual data de-identification module 102, the file merging module 112 can independently receive the set of raw video files 120 from the surgical video repository 130, and then merge the set of raw video files to recreate a merged (i.e., complete procedure) raw surgical video.

[0040] Next, the image data de-identification module 104 receives the partially processed complete procedure video 124. In some embodiments, the image data de-identification module 104 is configured to detect various types of image-based personally identifiable information embedded in the raw video images / frames of the complete procedure video 124, and then de-identify the detected personally identifiable information in the corresponding video images / frames. As in the textual data de-identification module 102, the image data de-identification module 104 can employ various image processing techniques to detect and de-identify personally identifiable information in the video images / frames of the complete procedure video 124. For example, the image data de-identification module 104 can employ facial recognition techniques to detect and de-identify faces in the video images / frames of the complete procedure video 124. In some embodiments, the image data de-identification module 104 can employ other image processing techniques to detect and de-identify other types of personally identifiable information in the video images / frames of the complete procedure video 124. Figure 1As can be seen, the image data de-identification module 104 can include a set of data anonymization sub-modules, such as an image text de-identification sub-module 104-1 and an OOB event removal sub-module 104-2. More specifically, the image text de-identification sub-module 104-1 is configured to detect various personally identifiable text that is recorded or automatically inserted into the original video images so that they are part of the original video images. For example, the recorded text can include text printed on surgical tools used during a laparoscopic or endoscopic surgical procedure. The recorded text can also include various text-based information displayed inside the OR but accidentally recorded by the laparoscope or endoscope camera during OOB events, e.g., text written on a whiteboard inside the OR, text printed on surgical staff’s scrubs or uniforms, text displayed on a monitor inside the OR, or other text-bearing objects inside the OR. The personally identifiable text can also include inserted text automatically inserted into the recorded procedure video such as standard user interface (UI) text boxes / panels that show identifiable information such as the name of the surgeon and the name of the hospital.

[0041] After detecting such personally identifiable text in one or more video frames of the partially processed complete procedure video 124, the image text de-identification sub-module 104-1 can also be configured to automatically blur out or otherwise render indiscernible the detected text using other special effects or techniques in the corresponding video frames. In some embodiments, for detected text boxes within video frames that are not considered part of the surgical video, the entire text box can be blurred out or otherwise edited out. In some other embodiments, the text within the detected text box can first be identified, and then the identified text can be separated into sensitive text and informational text. Next, only the determined sensitive text such as the name of the surgeon will be blurred out or otherwise edited out, while the determined informational text such as the surgical tool type / name can be left untouched in the video images.

[0042] In various embodiments, the image text de-identification sub-module 104-1 can include one or more machine learning models trained to detect and identify various recorded text and text boxes within a given video image, such as text printed on surgical tools or displayed within text boxes. Thus, the image text de-identification sub-module 104-1 can use one or more machine learning models to automatically identify different types of personally identifiable text embedded in the video images of the partially processed video 124, and automatically blur out / render indiscernible or otherwise de-identify the detected text data in the corresponding video images.

[0043] For example, the image text de-identification sub-module 104-1 can include a surgical tool text detection model configured to first detect surgical tools within a video image using a machine learning-based tool detection and recognition model. Next, the surgical tool text detection model further processes the detected tool images to detect any personally identifiable text within the boundaries of the detected tool images. If such text is detected, the image text de-identification sub-module 104-1 is configured to blur out or otherwise render the detected text unrecognizable.

[0044] As another example, the image text de-identification sub-module 104-1 can include a text box detection model configured to detect standard UI panels or standard dialog boxes within a video image. In some embodiments, this text box detection model can be trained based on a set of video images containing inserted text panels. To prepare the training data, appropriately sized boxes can be drawn around the text panels within the training images, indicating that the contents inside the text boxes are not part of the surgical video. Next, the text box detection model can be trained based on these generated text boxes to teach the model to look for such standard UI panels within other video frames that can contain personally identifiable text that needs to be redacted. Note that a single video frame can contain more than one such standard UI panel / text box that should be detected by the model. In some embodiments, the text box detection model can also learn from the training data the potential locations of potential UI panels / text boxes within a given video frame to facilitate detecting these standard UI panels / text boxes with higher accuracy and faster speed.

[0045] In some embodiments, multiple machine learning-based image text detection models can be structured such that each detection model is used to detect a particular type of image text embedded in a video image. For example, in addition to the text box detection model described above, there can be a surgical tool text detection model structured to detect text printed on surgical tools, a whiteboard text detection model structured to detect text captured from an OR whiteboard, and a monitor text detection model structured to detect text captured from an OR monitor. In various embodiments, the one or more image text detection models described above for detecting text embedded in a video image can include a regression model, a deep neural network-based model, a support vector machine, a decision tree, a Naive Bayes classifier, a Bayesian network, or a K-Nearest Neighbor (KNN) model. In some embodiments, each of these machine learning models is built based on a convolutional neural network (CNN) architecture, a recurrent neural network (RNN) architecture, or another form of deep neural network (DNN) architecture.

[0046] Referring back to Figure 1The OOB event removal submodule 104-2 of the data de-identification module 104 is configured to detect each OOB event within the original surgical procedure video and subsequently remove the identified OOB segment from the original surgical procedure video. In some embodiments, the OOB event removal submodule 104-2 can include an OOB event detector configured to scan the procedure video for OOB events. For example, such an OOB event detector can include an image processing unit configured to detect when the endoscope is removed from the patient, i.e., the start of an OOB event. The image processing unit is also configured to detect when the endoscope is inserted back into the patient, i.e., the end of an OOB event. Thus, the sequence of video frames between the detected start and end of an OOB event corresponds to a detected OOB segment within the full procedure video. In some embodiments, for each detected OOB segment, each video frame in the detected OOB segment can be completely blurred out or otherwise edited out (e.g., replaced with a black frame) such that any personally identifiable information within the edited video frames cannot be identified, such as all recorded text and inserted text boxes, as well as non-text personally identifiable information such as human faces. Note that typically only the video frames associated with each detected OOB segment are blurred out or otherwise edited out, but the actual frames are not deleted from the video, thereby maintaining the original timing information of the detected OOB event.

[0047] In various embodiments, the OOB event removal submodule 104-2 can include a machine learning-based OOB event detection model trained to detect OOB segments within an original surgical procedure video. Such an OOB event detection model can be trained based on a set of labeled video segments of a set of actual OOB events extracted from actual surgical procedure videos. The OOB event detection model can also be trained based on a set of video segments of a set of simulated OOB events extracted from actual procedure videos or training videos. For example, if a training OOB video segment corresponds to a 15-second segment of a procedure video, then all video frames from the 15-second segment can be labeled with an “OOB” identifier. Next, the OOB event detection model can be trained based on these labeled video frames to teach the model to detect and identify similar events in an original surgical procedure video.

[0048] As noted above, different OOB events can be caused by different reasons, e.g., one reason can be due to a switch from a robotic procedure to a laparoscopic procedure, while another reason can be due to lens cleaning. However, there is a strong similarity between different OOB events in that they typically all include a beginning phase when the endoscope camera is taken out of the patient and an ending phase when the endoscope camera is put back into the patient. Thus, a single trained OOB event detection model can be used by the OOB event removal submodule 104-2 to detect various OOB events caused by different reasons within a complete procedure video, and then to blur out, mask out, or otherwise de-identify each detected OOB video segment. However, in some embodiments, multiple machine learning based OOB event detection models can be constructed such that each OOB event detection model is used to detect one type of OOB event caused by a particular triggering reason / event.

[0049] In various embodiments, the above described single or multiple OOB event detection models used to detect OOB segments in a complete surgical procedure video can include a regression model, a deep neural network based model, a support vector machine, a decision tree, a Naive Bayes classifier, a Bayesian network, or a K-Nearest Neighbor (KNN) model. In some embodiments, each of these machine learning models is built based on a convolutional neural network (CNN) architecture, a recurrent neural network (RNN) architecture, or another form of deep neural network (DNN) architecture.

[0050] It is noted that personally identifiable information embedded in the original video images can also include non-textual personally identifiable information. In particular, non-textual personally identifiable information can include facial images of patients and / or surgical staff. For example, such facial images can be inadvertently recorded during an OOB event. However, any facial images captured during an OOB event can be effectively removed using the above described OOB event removal submodule 104-2. However, in some embodiments, the image data de-identification module 104 can also include a facial image removal submodule (not shown) configured to detect faces embedded in the original video images (e.g., by using conventional facial detection techniques), and then to blur out or otherwise de-identify each detected face from the corresponding video images. In some embodiments, to de-identify the original video images, the image data de-identification module 104 can first apply the OOB event removal submodule 104-2 to the original surgical procedure video to remove OOB events from the original surgical procedure video. Next, the image data de-identification module 104 applies the facial image removal submodule described herein to the processed video images to search for any faces in video frames outside of the detected OOB segments, and then to blur out any detected faces. Figure 1 In various embodiments, the above described single or multiple OOB event detection models used to detect OOB segments in a complete surgical procedure video can include a regression model, a deep neural network based model, a support vector machine, a decision tree, a Naive Bayes classifier, a Bayesian network, or a K-Nearest Neighbor (KNN) model. In some embodiments, each of these machine learning models is built based on a convolutional neural network (CNN) architecture, a recurrent neural network (RNN) architecture, or another form of deep neural network (DNN) architecture.

[0050] It is noted that personally identifiable information embedded in the original video images can also include non-textual personally identifiable information. In particular, non-textual personally identifiable information can include facial images of patients and / or surgical staff. For example, such facial images can be inadvertently recorded during an OOB event. However, any facial images captured during an OOB event can be effectively removed using the above described OOB event removal submodule 104-2. However, in some embodiments, the image data de-identification module 104 can also include a facial image removal submodule (not shown) configured to detect faces embedded in the original video images (e.g., by using conventional facial detection techniques), and then to blur out or otherwise de-identify each detected face from the corresponding video images. In some embodiments, to de-identify the original video images, the image data de-identification module 104 can first apply the OOB event removal submodule 104-2 to the original surgical procedure video to remove OOB events from the original surgical procedure video. Next, the image data de-identification module 104 applies the facial image removal submodule described herein to the processed video images to search for any faces in video frames outside of the detected OOB segments, and then to blur out any detected faces.

[0051] Although not explicitly shown, the video anonymization system 100 can include additional modules for detecting and de-identifying other types of personally identifiable information within the original surgical video that are not described by the unbound text data de-identification module 102 and the image data de-identification module 104. For example, the video anonymization system 100 can also include an audio data de-identification module configured to detect and de-identify audio data containing personally identifiable information, such as surgical staff communications recorded during a surgical procedure. In some embodiments, the audio data de-identification module is configured to completely remove all voice tracks embedded in a given original surgical video.

[0052] As can be seen in Figure 1 the image data de-identification module 104 outputs a fully processed complete procedure video 126 that is received by a verification module 106. In some embodiments, the verification module 106 is configured to ensure that a given fully processed surgical video does not contain any personally identifiable information. In some embodiments, the verification module 106 is configured to perform random sampling on the fully processed surgical video 126, i.e., by randomly selecting a plurality of video segments within the fully processed surgical video and verifying that the randomly selected video segments do not contain any personally identifiable information.

[0053] In some embodiments, instead of selecting video segments for verification with complete randomness, the verification module 106 can perform strategic random sampling to select a set of video segments for verification from portions of the processed surgical video that have a higher probability of containing personally identifiable information. For example, if it is known from statistical data that, during one or more particular steps / phases of a given surgical procedure, the surgeon typically or almost always removes the camera for cleaning or other reasons, it becomes more predictable approximately when during the recorded procedure one or more OOB events can have occurred. Thus, instead of randomly sampling the complete video to verify the anonymization results of the image data de-identification module 104, the verification module 106 can first determine one or more time periods associated with one or more particular procedure steps / phases that have a high probability of containing OOB events. The verification module 106 then selects a set of video segments within or around the determined high probability time periods to verify that the selected video segments do not contain any personally identifiable information.

[0054] In some embodiments, the verification module 106 can cooperate with a surgical stage segmentation engine configured to segment a surgical procedure video into a set of predefined stages, where each stage represents a particular stage of the associated surgical procedure that is used for a unique and distinguishable purpose throughout the surgical procedure. Further details of surgical stage segmentation techniques based on surgical video analysis have been described in related patent application Serial No. 15 / 987,782 and filing date of May 23, 2018, the contents of which are incorporated by reference herein.

[0055] More specifically, the fully processed procedure video 126 or even the partially processed procedure video 124 can first be sent to a stage segmentation engine configured to identify different stages of the surgical procedure video. The output from the stage segmentation engine includes a set of predefined stages and associated timing information relative to the complete procedure video. In conjunction with knowing which predefined stage(s) have a high probability of including OOB events, the verification module 106 can then “zoom in” on each “high OOB probability” stage of the fully processed procedure video 126 and strategically select a set of video segments from these high probability stages of the fully processed video for verification that the selected video segments do not contain any personally identifiable information.

[0056] In some embodiments, the verification module 106 can also reuse the above-described OOB event detection model to identify exact segments of the fully processed procedure video 126 that correspond to detected OOB events. Each identified OOB video segment is then directly verified (e.g., by a human operator) to determine whether the OOB video segment does not contain any personally identifiable information.

[0057] In some embodiments, if the verification module 106 determines that a given sampled video segment does not completely lack personally identifiable information, the video anonymization system 100 can be configured to reapply the image data de-identification module 104 on the portion of the fully processed procedure video 126 that contains the problematic video segment in an attempt to de-identify any remaining personally identifiable information in that portion of the video (illustrated by the arrow from the verification module 106 back to the module 104 in Figure 1 Alternatively, a manual de-identification process can be used to de-identify any remaining personally identifiable information within the sampled video segment that was determined to contain personally identifiable information. In some embodiments, after reapplying the image data de-identification module 104 to the problematic video segment or performing manual de-identification on the problematic video segment, the verification module 106 can be reapplied to the further processed surgical video to perform another pass of the verification operations described above.

[0058] Referring backFigure 1 After the verification module 106 has verified the anonymization results of the fully processed complete procedure video 126, the verification module 106 outputs the de-identified surgical video 128 that does not contain any personally identifiable information, i.e., any patient or surgical staff in the video is completely unidentifiable. Upon receiving the de-identified surgical video 128, the raw data purge module 108 can be configured to permanently remove any raw and partially processed video files and detected text-based personal identifiers. For example, the raw data purge module 108 can permanently remove the raw surgical video file 120, the partially processed raw video file 122, the partially processed complete procedure video file 124, and the fully processed complete procedure video file 126.

[0059] In some embodiments, after purging the raw and partially processed surgical videos, the raw data purge module 108 is further configured to store the de-identified surgical video 128 back into the surgical video repository 130. If the surgical video repository 130 also stores the raw surgical video file 120 of the de-identified surgical video 128, the raw data purge module 108 can be configured to permanently remove any copies of the raw surgical video file 120 from the surgical video repository 130. In some embodiments, the raw data purge module 108 is further configured to create a database of the de-identified surgical videos. In some embodiments, the database entry associated with the de-identified surgical video 128 in the surgical video repository 130 can store the file name, as well as other statistical data and attributes of the de-identified surgical video 128 extracted during the video anonymization process described above. For example, these statistical data and attributes can include, but are not limited to, the type of surgical procedure, the complete video / video segment identifier, the laparoscopic procedure / robotic procedure identifier, the number of surgeons who contributed to the given procedure, the number of OOB events during the procedure, and the like.

[0060] Next, the de-identified surgical procedure videos stored in the surgical procedure video repository 130 can be published to or otherwise provided to clinical experts for various research purposes. For example, the de-identified surgical procedure videos can be published to a surgical procedure video data specialist who can use the de-identified surgical procedure videos to build various machine learning based analysis tools. More specifically, the de-identified surgical procedure videos can be used to: build machine learning targets to prepare for mining surgical data from surgical procedure videos of a given surgical procedure, as described in related patent application having serial number 15 / 987,782 and filing date of May 23, 2018; build machine learning based surgical procedure tool inventory and tool usage tracking tools, as described in related patent application having serial number 16 / 129,607 and filing date of September 12, 2018; or build a surgical procedure stage segmentation tool, which is also described in related patent application having serial number 15 / 987,782 and filing date of May 23, 2018, the contents of which are incorporated by reference herein.

[0061] Figure 2 A flowchart showing an exemplary process 200 for anonymizing original surgical procedure videos to de-identify personally identifiable information embedded in the video images, in accordance with some embodiments described herein, is presented. In one or more embodiments, one or more of the steps in Figure 2 may be omitted, repeated, and / or performed in a different order. As such, the specific arrangement of steps shown should not be construed as limiting the scope of the present technology. Figure 2 The specific arrangement of steps shown should not be construed as limiting the scope of the present technology.

[0062] The process 200 can begin by receiving a set of original surgical procedure videos corresponding to a set of recorded video segments / clips of a complete surgical procedure performed in an OR (step 202). As described above, the set of original surgical procedure videos correspond to a set of shorter video segments / clips of a complete surgical procedure, where each of the original surgical procedure videos can be generated as a result of a maximum file size constraint (e.g., 500 MB / video). In some embodiments, the process 200 can receive the set of original surgical procedure videos from a HIPAA compliant video repository. In these embodiments, the set of recorded original surgical procedure videos are first transmitted from the OR to the HIPAA compliant repository for temporary storage. In other embodiments, the process 200 can receive the set of original surgical procedure videos directly from a recording device within the OR over a secure network connection without having to retrieve stored videos from a HIPAA compliant video repository. Note that when receiving the set of original surgical procedure video files, the process 200 can receive the set of original video files as well as the folder.

[0063] Next, the process 200 performs text-based de-identification on the set of received raw surgical videos and folders to detect and de-identify text-based personally identifiable information associated with the raw surgical videos and associated folders (step 204). As described above, the text-based de-identification operation analyzes those text data associated with the raw surgical videos and folders containing the raw surgical videos that are not embedded in the video images. More specifically, the text-based de-identification operation detects text-based personally identifiable information embedded in file identifiers (e.g., file names), folder identifiers (e.g., folder names), file metadata such as file attributes, and other metadata associated with the set of raw video files and folders. As described above, the text-based personally identifiable information can include both text-based patient data and text-based staff data, and the text-based patient data can include, but is not limited to, the patient's name, DOB, age, gender, SSN, address, MRN embedded in the text-based data of the video files and folders.

[0064] The process 200 then removes or otherwise de-identifies the detected text-based personally identifiable information from the file identifiers, folder identifiers, and other metadata associated with the set of raw video files and folders. As described above, the process 200 can use open source text detection and removal tools or conventional text detection techniques in step 204. Note that at the end of step 204, the set of raw video files is considered partially processed: that is, while the text-based personally identifiable information has been de-identified, the potential personally identifiable information embedded in the video images has not yet been de-identified.

[0065] Next, the process 200 stitches together the set of partially processed videos to recreate a complete procedure video corresponding to the complete surgical procedure (step 206). As described above, different recording devices often have different naming conventions. In some embodiments, merging the set of partially processed video files involves first analyzing the particular naming convention used by the set of video files to determine the correct order of the set of video segments relative to the complete surgical procedure, and then placing the set of partially processed videos back into the complete procedure video.

[0066] Next, process 200 performs image-based de-identification on the partially processed complete procedure video to detect image-based personally identifiable information de-identification embedded in the corresponding video images (step 208). As described above, image-based personally identifiable information can include various text recorded or automatically inserted into the video frames such that they can be displayed along with the video images. Accordingly, process 200 can use image text de-identification submodule 204-1 described above to automatically identify different types of personally identifiable text within the original video images of the complete procedure video and automatically blur out / obscure or otherwise de-identify the detected text data embedded in the corresponding video images.

[0067] Further, for personally identifiable information captured during OOB events, process 200 can use OOB event removal submodule 104-2 described above to automatically detect various OOB segments within the complete procedure video and subsequently blur out, obscure, or otherwise de-identify the identified OOB segments from the complete procedure video. Further, process 200 can use a dedicated facial image removal submodule to search video frames outside of the detected OOB segments for any faces that were not removed by OOB event removal submodule 104-2 and blur out each detected face. Note that at the end of step 208, the complete procedure video is considered to be fully processed.

[0068] Next, process 200 performs a verification operation on the fully processed procedure video to verify that the fully processed procedure video does not contain any personally identifiable information (step 210). As described above, process 200 can use verification module 106 described above to automatically perform random sampling of the fully processed procedure video or strategic sampling of the fully processed procedure video.

[0069] Next, process 200 determines whether the verification operation was successful (step 212). If one or more of the sampled video segments is found to not be completely free of personally identifiable information, process 200 can perform another automatic de-identification operation or a manual de-identification operation on each failed video segment (step 214). After each problematic video segment has been properly processed, the verification operation can be repeated (i.e., process 200 returns to step 210). If the verification operation is determined to be successful at step 212, the video de-identification process is complete. Process 200 then permanently removes the original surgical procedure video and related file data from the video repository and stores the de-identified procedure video in place of the removed original surgical procedure video (step 216).

[0070] It is noted that while the disclosed anonymization techniques have been described in the context of anonymizing original surgical procedure videos, the disclosed techniques can also be used to anonymize still images captured within an OR to de-identify personal identifiable information embedded in the still images. Further, while some of the disclosed anonymization techniques have been described in the context of anonymizing a set of original surgical procedure videos corresponding to a long surgical procedure, the disclosed anonymization techniques can also be applied to a single surgical procedure video to de-identify personal identifiable information embedded in the single surgical procedure video.

[0071] Figure 3 A flowchart showing an exemplary process 300 for detecting OOB video segments from original surgical procedure videos and removing them from the original surgical procedure videos to de-identify personal identifiable information embedded in associated OOB video images in accordance with some embodiments described herein is presented. In one or more embodiments, one or more of the steps in Figure 3 may be omitted, repeated, and / or performed in a different order. As such, Figure 3 The specific arrangement of steps shown should not be construed as limiting the scope of the present technology.

[0072] Process 300 can begin by training an OOB event detection model (step 302). In some embodiments, such an OOB event detection model can be trained based on a set of labeled video segments of a set of actual OOB events extracted from actual surgical procedure videos. For example, if the training OOB video segments correspond to a 15-second segment of a procedure video, all video frames from the 15-second segment can be labeled with an “OOB” identifier. Next, the OOB event detection model can be trained based on these labeled video frames to teach the model to detect and identify similar events in original surgical procedure videos. In particular, the OOB event detection model can be trained to detect a beginning phase of an OOB event, i.e., a sequence of video images corresponding to the event of a scope being removed from a patient’s body, and an ending phase of an OOB event, i.e., a sequence of video images corresponding to the event of a scope being inserted back into a patient’s body. The OOB event detection model is also trained to associate a detected beginning phase with a detected ending phase to the same OOB event. For example, if a detected ending phase of an OOB event immediately follows a detected beginning phase of an OOB event, and the time interval between the two detections is below a predetermined threshold, the two detected phases can be considered to belong to the same OOB event.

[0073] The process 300 next applies the trained OOB event detection model to the original procedure video to detect a start phase of an OOB event (step 304). If a start phase of an OOB event is detected, the process 300 then applies the trained OOB event detection model to the original procedure video to detect an end phase of the OOB event (step 306). Next, the process 300 determines whether the detected start phase and end phase belong to the same OOB event (step 308), and if so, the process 300 obscures, masks, or otherwise de-identifies all video frames between the detected start phase and end phase of the OOB event (step 310). However, if the process 300 determines that the detected start phase and end phase do not belong to the same OOB event, the process 300 can generate an alert indicating that an anomaly was detected (step 312).

[0074] Figure 4 A computer system that can be used to implement some embodiments of the subject technology is conceptually illustrated. Computer system 400 can be a client, a server, a computer, a smart phone, a PDA, a laptop, or a tablet computer with one or more processors embedded therein or coupled thereto, or any other type of computing device. Such a computer system includes various types of computer readable media and interfaces for various other types of computer readable media. Computer system 400 includes a bus 402, a processing unit 412, a system memory 404, a read-only memory (ROM) 410, a permanent storage device 408, an input device interface 414, an output device interface 406, and a network interface 416. In some embodiments, computer system 400 is part of a robotic surgical system.

[0075] Bus 402 collectively represents all system and peripheral buses that communicatively connect the various internals of the computer system 400. For instance, bus 402 communicatively connects the processing unit 412 with the ROM 410, the system memory 404, and the permanent storage device 408.

[0076] Processing unit 412 retrieves from these various memory units instructions to be executed and data to be processed, in order to execute the various processes described in this patent disclosure, including in connection with Figures 1 to 3The above process anonymizes the original surgical procedure videos to de-identify personal identifiable information embedded in corresponding file and folder identifiers and in corresponding video images. Processing unit 412 can include any type of processor, including but not limited to microprocessors, graphics processing units (GPUs), tensor processing units (TPUs), intelligent processor units (IPUs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), and application-specific integrated circuits (ASICs). In different implementations, processing unit 412 can be a single processor or a multi-core processor.

[0077] ROM 410 stores static data and instructions that are needed by processing unit 412 and other modules of the computer system. In another aspect, permanent storage 408 is a read-and-write memory device. This device is a non-volatile memory unit that stores instructions and data even when computer system 400 is off. Some implementations of the subject disclosure use a mass-storage device (such as a magnetic or optical disk and its corresponding disk drive) as permanent storage 408.

[0078] Other implementations use a removable storage device (such as a floppy disk, flash drive, and its corresponding disk drive) as permanent storage 408. Like permanent storage 408, system memory 404 is a read-and-write memory device. However, unlike storage device 408, which is a non-volatile memory, system memory 404 is a volatile read-and-write memory, such as a random access memory. System memory 404 stores some of the instructions and data that the processor needs at runtime. In some implementations, the various processes described in this patent disclosure are implemented as program modules that are stored in system memory 404 and that are executed by the processor. Bus 402 provides a communication path for the processor 412, the system memory 404, the read-only memory 410, the permanent storage 408, the input devices 414, and the output devices 406. Examples of bus 402 include, but are not limited to, a system bus, a video bus, a midi bus, and an audio bus. Figures 1 to 3 The above process anonymizes the original surgical procedure videos to de-identify personal identifiable information embedded in corresponding file and folder identifiers and in corresponding video images. Processing unit 412 can include any type of processor, including but not limited to microprocessors, graphics processing units (GPUs), tensor processing units (TPUs), intelligent processor units (IPUs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), and application-specific integrated circuits (ASICs). In different implementations, processing unit 412 can be a single processor or a multi-core processor.

[0079] Bus 402 also connects to the input and output devices 414 and 406. Input device 414 enables the user to communicate information and select commands to the computer system. Input devices 414 can include, for example, alphanumeric keyboards and pointing devices (also referred to as “cursor control devices”). Output device 406 enables, for example, the display of images generated by the computer system 400. Output devices 406 can include, for example, printers and display devices, such as cathode ray tubes (CRT) or liquid crystal displays (LCD). Some implementations include devices that function as both input and output devices, such as a touch screen.

[0080] Finally, as shown in FIG. 4, bus 402 also couples computer system 400 to a network through a network adapter 416. Network adapter 416 enables computer system 400 to exchange data with the network, which can be an intranet, Local Area Network (LAN), or a Wide Area Network (WAN). Examples of networks include Ethernet, Token Ring, and the Internet. Wireless networks, such as IEEE 802.11, 802.16, 802.20, and other forms of mobile or wireless data networks, are also included.Figure 4 As shown, bus 402 also couples computer system 400 to a network (not shown) through network interface 416. In this manner, the computer can be a part of a network of computers such as a local area network ("LAN"), a wide area network ("WAN"), an inter network (e.g., the Internet), or an Intranet, or a network of networks, such as the Internet. Any or all components of computer system 400 can be used in conjunction with the subject disclosure.

[0081] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.

[0082] The hardware used to implement the various illustrative logics, logical blocks, modules, and circuits described in connection with the aspects disclosed herein can be implemented or performed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field

[0083] In one or more exemplary aspects, the functions described can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored as one or more instructions or code on a non-transitory computer-readable storage medium or a non-transitory processor-readable storage medium. The steps of a method or algorithm described herein can be embodied in a processor-executable instruction that can reside on a non-transitory computer- or processor-readable storage medium. Non-transitory computer- or processor-readable storage media can be any storage media that can be accessed by a computer or processor. By way of example but not limitation, such non-transitory computer- or processor-readable storage media can include RAM, ROM, EEPROM, FLASH memory, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Disk and disc, as used herein, includes compact discs (CD), laser discs, optical discs, digital versatile discs (DVD), floppy disks and blu-ray discs where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above are also included within the scope of non-transitory computer- and processor-readable media. Additionally, the operations of a method or algorithm can reside in one or any combination of the above memory hardware, as one or any combination of the code and / or instructions that can reside on a non-transitory processor-readable storage medium and / or computer-readable storage medium, thereby making a computer program product.

[0084] Although the patent document contains many details, these should not be understood as limiting the scope of any disclosed technologies or claimable content to the specific embodiments described, but rather as a description of features that can be particular to the specific implementation of the particular technology. Certain features described in the context of separate embodiments in this patent document can also be implemented in a combined form in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately or in any subcombination in multiple embodiments. Also, while features can be described above as acting in certain combinations and even initially so claimed by the claims, one or more features from a claimed combination can in some cases be excised from the combination and the claimed combination can then be directed to a subcombination or variation of a subcombination.

[0085] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring such order, nor that all illustrated operations be performed, to achieve desirable results. Moreover, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.

[0086] Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in this patent document.

Claims

1. A computer-implemented method for anonymizing raw surgical procedure video recorded by a recording device within an operating room (OR) of a surgical procedure performed within the OR, the method comprising: receiving a set of raw surgical videos generated by the recording device and corresponding to a surgical procedure performed within the OR; merging the set of raw surgical videos to generate a surgical procedure video corresponding to the surgical procedure; and detecting image-based personally identifiable information embedded in a set of raw video images of the surgical procedure video; wherein detecting the image-based personally identifiable information embedded in the set of raw video images of the surgical procedure video further comprises: scanning the set of raw video images to detect video segments corresponding to an extracorporeal event at a time when the recording device is removed from a patient's body, wherein personally identifiable information can be inadvertently captured by the recording device during the extracorporeal event; and wherein if an extracorporeal event is detected within the set of raw video images, de-identifying the detected image-based personally identifiable information by automatically blurring out each video image in the detected video segments corresponding to the detected extracorporeal event such that any personally identifiable information embedded within the detected video segments cannot be identified.

2. The computer-implemented method of claim 1, wherein merging the set of raw surgical videos to generate the surgical procedure video comprises: analyzing a set of file names associated with the set of raw surgical videos to determine a correct order with respect to the surgical procedure; and splicing the set of raw surgical videos together based on the determined order.

3. The computer-implemented method of claim 1, wherein detecting the image-based personally identifiable information embedded in the set of raw video images of the surgical procedure video comprises: detecting one or more forms of personally identifiable text as part of the set of raw video images; wherein if one or more forms of personally identifiable text are detected in one or more raw video images within the set of raw video images, de-identifying the detected image-based personally identifiable information by blurring out the detected text in the one or more raw video images.

4. The computer-implemented method of claim 3, wherein the one or more forms of personally identifiable text further comprise a form of recording text captured by the recording device used to record the set of raw surgical videos.

5. The computer-implemented method of claim 4, wherein the recording text can comprise one or more of: text printed on one or more surgical tools used in the surgical procedure and recorded by the recording device within a patient's body in which the surgical procedure is performed; and ​ ​ text displayed inside the surgical room where the surgical procedure is performed, wherein the text is accidentally recorded by the recording device during an extracorporeal OOB event when the recording device is being removed from the patient’s body.

6. The computer-implemented method of claim 5, wherein detecting the recorded text printed on the one or more surgical procedure tools comprises: detecting surgical procedure tools within the one or more raw video images using a machine learning based tool detection and recognition model; and processing a portion of the one or more raw video images containing the detected surgical procedure tools to detect any personally identifiable text within the portion of the one or more raw video images.

7. The computer-implemented method of claim 3, wherein the one or more forms of personally identifiable text further includes text boxes inserted into the surgical procedure video to display surgical procedure related information.

8. The computer-implemented method of claim 7, wherein detecting the text boxes in the surgical procedure video comprises: detecting text boxes at or near predetermined locations within the one or more raw video images using a machine learning based text box detection model; and processing a portion of the one or more raw video images containing the detected text boxes to detect any personally identifiable text within the portion of the one or more raw video images.

9. The computer-implemented method of claim 1, wherein detecting a video segment in the set of raw video images corresponding to an extracorporeal event comprises: detecting a beginning phase of an extracorporeal event where an endoscope camera is being removed from a patient’s body using a machine learning based extracorporeal event detection model; detecting an ending phase of the extracorporeal event where the endoscope camera is being inserted back into the patient’s body using the machine learning based extracorporeal event detection model; and labeling a set of video images in the set of raw video images located between the detected beginning phase and the detected ending phase as the video segment corresponding to the detected extracorporeal event.

10. The computer-implemented method of claim 9, wherein prior to using the machine learning based extracorporeal event detection model to detect an extracorporeal event, the method further comprises: training the extracorporeal event detection model based on a set of labeled video segments of a set of extracorporeal events extracted from actual surgical procedure videos.

11. The computer-implemented method of claim 1, wherein prior to merging the set of raw surgical procedure videos, the method further comprises: processing the received set of raw surgical procedure videos to detect text-based personally identifiable information embedded in file structure data associated with the set of raw surgical procedure videos; and if text-based personally identifiable information is detected, removing the detected text-based personally identifiable information from the file structure data.

12. The computer-implemented method of claim 11, wherein the file structure data comprises: file identifiers associated with the set of raw surgical procedure videos; folder identifiers of folders containing the set of raw surgical procedure videos; file attributes associated with the set of original surgical procedure videos; and other metadata associated with the set of original surgical procedure videos.

13. The computer-implemented method of claim 1, wherein after de-identifying the detected image-based personally identifiable information in the surgical procedure video, the method further comprises: performing random sampling within the de-identified surgical procedure video to randomly select a plurality of video segments within the de-identified surgical procedure video; and verifying that the randomly selected video segments are free of any personally identifiable information.

14. The computer-implemented method of claim 1, wherein after de-identifying the detected image-based personally identifiable information in the surgical procedure video, the method further comprises: permanently removing the set of original surgical procedure videos from a surgical procedure video repository; and replacing the removed original surgical procedure videos with the de-identified surgical procedure video.

15. The computer-implemented method of claim 1, wherein the personally identifiable information includes both patient identifiable information associated with a patient receiving the surgical procedure and surgical staff identifiable information associated with surgical staff performing the surgical procedure.

16. A system for anonymizing an original surgical procedure video recorded by a recording device within an operating room (OR), the system comprising: one or more processors; a memory coupled to the one or more processors; and wherein the memory stores a set of instructions which, when executed by the one or more processors, cause the system to: receive a set of original surgical procedure videos produced by the recording device and corresponding to a surgical procedure performed within the OR; merge the set of original surgical procedure videos to generate a surgical procedure video corresponding to the surgical procedure; and detect personally identifiable information embedded in a set of original video images of the surgical procedure video; wherein detecting the personally identifiable information embedded in the set of original video images of the surgical procedure video further comprises: scanning the set of original video images to detect a video segment corresponding to an extracorporeal event at a time when the recording device is removed from a patient’s body, wherein personally identifiable information can be inadvertently captured by the recording device during the extracorporeal event; and wherein if an extracorporeal event is detected within the set of original video images, de-identifying the detected image-based personally identifiable information by automatically blurring out each video image in the detected video segment corresponding to the detected extracorporeal event such that any personally identifiable information embedded within the detected video segment is rendered unidentifiable.

17. The system of claim 16, wherein detecting the personally identifiable information embedded in the one or more video images of the surgical procedure video comprises: detecting one or more forms of personally identifiable text as part of the set of original video images; And wherein if a form of personally identifiable text is detected in one or more raw video images within the set of raw video images, de-identifying the detected personally identifiable information includes blurring the detected text in the one or more raw video images.

18. The system of claim 16, wherein after de-identifying the detected personally identifiable information in the surgical procedure video, the system is configured to: permanently remove the set of raw surgical procedure videos from a surgical video repository; and replace the removed raw surgical procedure videos with the de-identified surgical procedure video.

19. The system of claim 16, wherein the system is configured to de-identify the detected personally identifiable information in the surgical procedure video by: blurring the detected personally identifiable information in the surgical procedure video.

20. The system of claim 16, wherein the system is configured to de-identify the detected personally identifiable information in the surgical procedure video by: replacing the detected personally identifiable information in the surgical procedure video with a generic placeholder.

21. The system of claim 16, wherein the system is configured to de-identify the detected personally identifiable information in the surgical procedure video by: replacing the detected personally identifiable information in the surgical procedure video with a generic placeholder.

22. The system of claim 16, wherein the system is configured to de-identify the detected personally identifiable information in the surgical procedure video by: replacing the detected personally identifiable information in the surgical procedure video with a generic placeholder.

23. The system of claim 16, wherein the system is configured to de-identify the detected personally identifiable information in the surgical procedure video by: replacing the detected personally identifiable information in the surgical procedure video with

Citation Information

Patent Citations

  • Method and system for automatically tracking and managing inventory of surgical tools in operating rooms

    US10679743B2

  • Machine-learning-oriented surgical video analysis system

    US11205508B2

  • Operating room black-box device, system, method and computer readable medium

    CN107615395A

  • De-identification in visual media data

    US20130182006A1

  • Automated system for medical video recording and storage

    US20180110398A1