Generating motor enhanced and de-identified video data from a patient video and generating a training data set of de-identified patient video data
The method addresses precision and privacy challenges in motor symptom assessment by generating de-identified video data with face swapping and blending, enhancing motor analysis and reducing bias in AI training data.
Patent Information
- Application Number
- PCT/EP2025/067109
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-25
- Filing Date
- 2025-06-18
- Publication Date
- 2026-01-02
AI Technical Summary
Existing methods for assessing motor symptoms in patients with conditions like Parkinson's Disease and Huntington's Disease face precision issues due to psychometric limitations and variability in human ratings, while maintaining patient privacy is crucial for compliance with regulations.
A computer-implemented method using face swapping and blending algorithms to generate motor enhanced and de-identified video data, ensuring patient anonymity while preserving relevant medical information for motor assessment.
The method effectively de-identifies patient faces, reduces human rater bias, and maintains essential facial features for accurate motor analysis, enabling the creation of unbiased training data sets for AI/ML algorithms.
Smart Images

Figure EP2025067109_02012026_PF_FP_ABST
Abstract
Description
[0001] Description
[0002] Generating motor enhanced and de-identified video data from a patient video and generating a training data set of de-identified patient video data
[0003] Technical Field
[0004] [1] The present disclosure relates generally to a computer-implemented method for generating a motor enhanced and de-identified video data from a patient video, a computer-implemented method for generating a training video data set for training an, e. g., Al incorporating, algorithm for automized video analysis of a patient’s motor capabilities, and a system for generating a motor enhanced and de-identified video data for training an Al incorporating algorithm for automized video analysis of a patient’s motor capabilities data. More generally, the present disclosure relates to providing de-identified training video data of patients for training a software-based movement analysis within a patient monitoring project, thereby simplifying the monitoring of motion affected patients.
[0005] Background
[0006] [2] The monitoring of patients participating in disease studies and their reaction to treatment is a fundamental aspect when developing new drugs for, e. g., neurodegenerative diseases such as Parkinson’s Disease (PD) and Huntington’s Disease (HD) that result in neurological motor disorders (motor symptoms). Specifically, continuous monitoring of the development of the severeness of the disease during the study requires repeatedly assessing any changes within the patient’s health and motor condition.
[0007] [3] For example, for the evaluation of PD, a Unified Parkinson’s Disease Rating Scale (MDS- UPDRS) was developed for assessing non-motor and motor experiences of daily living and motor complications. MDS-UPDRS includes an evaluation of motor capabilities and is used in a clinical setting as well as in research as MDS-UPDRS is acknowledged as a “gold standard measure” in Parkinson’s disease studies. Despite being a highly valid and reliable instrument, Parts II and III of the MDS-UPDRS testing include psychometric limitations. Psychometric limitations reduce, for example, the precision of measurement of motor symptoms and impact that are present in early PD. In addition, change scores of MDS-UPDRS Part I - III may contain a substantial amount of error variance because a rater’s background and experience may affect at least some of the MDS-UPDRS ratings. Therefore, it will be acknowledged that any improvement of disease ratings may affect and improve the outcome of disease studies.
[0008] [4] For background on the above considerations regarding MDS-UPDRS ratings, see, for example, Brooks C., et al.; “Quantification of discrete behavioral components of the MDS- UPDRS”; Journal of Clinical Neuroscience, 2019, 61 : 174-179, https: / / doi.Org / 10.1016 / j.jocn.2018.10.043 and Post B, et al.; “Unified Parkinson ’s Disease Rating Scale Motor Examination: Are Ratings of Nurses, Residents in Neurology, and Movement Disorders Specialists Interchangeable?”; Movement Disorders, 2005, 20(12): 1577-1584, DOI: 10.1002 / mds.20640 and Holden SK, et al. “Progression of MDS-UPDRS Scores Over Five Years in De Novo Parkinson Disease from the Parkinson ’s Progression Markers Initiative Cohort”; Movement Disorders Clinical Practice, 2018, 5(1): 47-53. doi: 10.1002 / mdc3.12553 and Horvath K, et al.; “Minimal clinically important difference on the Motor Examination part of MDS-UPDRS” Parkinsonism and Related Disorders, 2015, 21 : 1421-1426, http: / / dx.doi.Org / 10.1016 / j.parkreldis.2015.10.006; Evers L. J. W ., et al.; “Measuring Parkinson ’s Disease Over Time: The Real-World Within-Subject Reliability of the MDS-UPDRS”; Movement Disorders, 2019, 34 (10), DOI: 10.1002 / mds.27790.
[0009] [5] Similar ratings of motor symptoms are used for HD studies and are summarized, for example, in Schobel et al.; “Motor, cognitive, and functional declines contribute to a single progressive factor in early HD”; Neurology 2017, 89:2495-2502.
[0010] [6] In the medical field such as for drug development and patient monitoring, the protection of patients' privacy and ensuring compliance with local regulations (including, e. g., respective company policies) require defined procedures for the anonymization of a patient’s sensitive personal data. “Personal data” refers herein to any information relating to an identified or identifiable natural person ('data subject'). Generally, an identifiable person is a person that can be identified, directly or indirectly, in particular by reference to an identification number or by reference to one or more factors specific to the person’s identity such as, for example, physical, physiological, mental, economic, cultural or social identity. These factors include also biometric data, such as a patient’s face. Accordingly, the protection of a patient’s privacy in a video - e. g., an anonymization of patent data in a recorded video - includes the anonymization of a patient’s face in addition to any meta data associated with the recorded video. The anonymization of biometric data such as a patient’s face is generally also referred to as “de-identification”.
[0011] [7] “A digital mask to safeguard patient privacy” by Yahan Yang et al., Nature Medicine, [ONLINE], vol. 28, no. 9, pages!883-1892, Sep. 15, 2022 discloses a digital mask technology that is based on three-dimensional reconstruction and deep-learning algorithms to irreversibly erase identifiable features, while retaining disease-relevant features needed for diagnosis. Specifically, a digital mask is based on a face model that represents the overall geometry of the face. Deep learning achieves feature extraction from different facial parts of the patient’s face, and 3D reconstruction automatically digitalizes the shapes and motions of 3D faces, eyelids and eyeballs based on the extracted facial features. The final result of the derived digital mask is put onto the patient’s face, to completely cover the same.
[0012] [8] “Deepfakes for Medical Video De-Identification: Privacy Protection and Diagnostic Information Preservation" by Bingquan Zhu et al., AIES '20: Proceedings of the AAAI / ACM Conference on Al, Ethics, and Society, pages 414-420, Feb. 07, 2020 discloses a solution for de-identification using a face swapping technique to protect privacy in medical video. Specifically, patients' faces are swapped to a proper target face. The patient’s face becomes unrecognizable. The proposed swapping de-identification method showed that face-swapping as a de-identification approach is reliable, and keeps keypoints almost invariant, significantly better than traditional methods.
[0013] [9] “Artificial Intelligence-Based Face Transformation in Patient Seizure Videos for Privacy Protection" by Jen-Cheng Hou, Mayo Clinic Proceedings, Digital Health, Vol. 1, Issue 4, pages 619-628, Dec. 2023 discloses a face-swapping tool called MobileF aceSwap to substitute the faces of patients with desired faces. Furthermore, the model VToonify was used to transform portrait videos into cartoon-like styles. Facial deidentification and preservation of clinically relevant facial detail were calculated based on: (1) scoring by 5 independent expert clinicians and (2) objective computation. According to the clinician scoring of 26 facial frames in 16 patients, the best compromise between deidentification and preservation of facial semiology was the cartoonization model.
[0014]
[0010] With respect to maintaining facial attribute information, it is referred to “SF-GAN: Face DeIdentification Method Without Losing Facial Attribute Information" by Yongxiang Li et al., IEEE Signal Processing Letters, Vol. 28, pages 1345-1349, March 19, 2021.
[0015]
[0011] In the context of the present disclosure, “video frame de-identification” - also context depending “de-identification” in short - refers to anonymization of video frame data of a patient with respect to the patient’s face. Similarly, “video de-identification” refers to anonymization of video data, i. e., a plurality / sequence of video frames.
[0016]
[0012] A first generic object of the invention disclosed herein is improving drug development and specifically simplify disease studies and their assessment. In a more specifical object, the invention aims to support the development of an, e. g., Al (artificial intelligence) / ML (machine learning) incorporating algorithm for automized analysis of videos of patients, such as PD or HD patients. In an even more specific object, the invention aims to provide anonymization of video data (video streams) and accordingly video frames within the medical field. In other words, there is a need of a de-identification method of video data that can be applied to, for example, a video stream generated in a medical context, the medical context making the video data subject to protection of a patient’s privacy by de-identification and additionally requiring the de-identification to maintain the medical information of interest.
[0017]
[0013] Thus, the present disclosure is directed, at least in part, to improving or overcoming one or more aspects of prior systems, and, in particular, to providing a solution of one or more of the above summarized objectives.
[0018]
[0014] Some of the above objects may be achieved by a method as recited in claim 1 or in claim 14, a computer program as recited in claim 17, a system as recited in claim 18, and a use as recited in claim 19. Further aspects and developments are given in the dependent claims.
[0019]
[0015] In a first aspect, there is disclosed a computer-implemented method for generating a motor enhanced and de-identified video data from a patient video data. The method comprises:
[0020] - receiving, by a computer device, the patient video data. The patient video data comprises a plurality of frames, wherein the frames of the plurality of frames include respective image data of a face of a patient; the face of the patient has, for example, facial features for a medical motor assessment;
[0021] - receiving, by the computer device, image data of a first computer-generated head image (e. g., from a library of computer-generated head images) and identifying the image data of a face of the first computer-generated head image as image data of a source face;
[0022] - in each frame of the plurality of frames: identifying, by the computer device, the image data of the face of the patient as image data of a target face; based on the image data of the source face, performing, by the computer device, a de-identification process of the frame that includes an anonymization process, a blend process, and a re-identification validation process. The anonymization process includes a face swapping process configured to replace the image data of the target face with image data of a modified source face, wherein the image data of the modified source face is derived from the image data of the source face and attribute feature data of the target face, thereby generating a de-identified frame including image data of the modified source face. The blend process generates a motor enhanced frame by blending the image data of the target face onto the de-identified frame, thereby forming image data of a motor enhanced face that includes facial features of the target face for enhancing the medical motor assessment, and the re-identification validation process is configured to calculate a re- identification score for the motor enhanced frame that indicates to what extent the motor enhanced face resembles the target face of the patient; and
[0023] - based on the re-identification scores, generating, by the computer device, the motor enhanced and de-identified video data from the motor enhanced frames.
[0024]
[0016] In some embodiments of the method, the blend process, in particular in combination with the face swapping process, may be configured to introduce a dissimilarity between the image data of at least one motor enhanced face and the image data of the target face to be a above a preset threshold value. The preset threshold value may in particular correspond to a Cosine Distance of 0.4. The Cosine Distance may be based on the cosine of the angle between two non-zero vectors in a multi-dimensional space, in particular an RGB color image data space, for the feature vector representing the image data of the motor enhanced face and the image data of the target face.
[0025]
[0017] In some embodiments, the method may further comprise setting the blend process such that, for at least one of the plurality of the motor enhanced frames, the calculated re-identification score passes a preset threshold value. Furthermore, an adjustable parameter of the blend process, such as a weight parameter and / or an erosion parameter of the target face when overlaid on the modified source face, may in particular be set to adapt the contribution of the image data of the target face to the image data of the motor enhanced face.
[0026]
[0018] In some embodiments of the method, the re-identification score may be determined by calculating a dissimilarity between the image data of the motor enhanced face and the image data of the target face. The dissimilarity can in particular be determined based on a Cosine Distance that is based on the cosine of the angle between two non-zero vectors in a multidimensional space, in particular an RGB color image data space, for the feature vector representing the image data of the motor enhanced face and the image data of the target face. In particular the preset threshold value can be set to correspond to a Cosine Distance in a range from 0 to 2, in particular from 0.4 to 2, such as >0.4.
[0027]
[0019] In some embodiments, when it is determined that the calculated re-identification score of one of the motor enhanced frames does not pass the preset threshold value, the method may further comprise replacing the frame, which does not pass the preset threshold value, with a motor enhanced neighboring frame, which does pass the preset threshold value and is next to or close to in a sequence of motor enhanced frames with respect to the motor enhanced frame, which does not pass the preset threshold.
[0028]
[0020] In some embodiments, when it is determined that that the calculated re-identification score for none or a preset maximum number of frames of the motor enhanced frames does not pass the preset threshold value, the method may further comprise setting the blend process, in particular an adjustable parameter, such as a weight parameter and / or an erosion parameter of the target face when overlaid on the modified source face, to increase a Cosine Distance between the image data of the motor enhanced faces and the image data of the target face, respectively.
[0029]
[0021] In some embodiments, when it is determined that the calculated re-identification score for none or a preset maximum number of frames of the motor enhanced frames does not pass the preset threshold value, the method may further comprise receiving / selecting, e. g., from the library of computer-generated head images, image data of a second computer-generated head image and identifying the image data of a face of the second computer-generated head image as image data of the source face, and performing the de-identification process of the target face based on that source face.
[0030]
[0022] In some embodiments, the face swapping process may include a face swapping algorithm that is selected from the group of deep fake face swapping algorithms comprising the GHOST algorithm, the DeepFaceLab algorithm, the SimSwap algorithm, the MobileF aceSwap algorithm, the FaceShifter algorithm, and the HifiFace algorithm. In addition or alternatively, the face swapping process, in particular a face swapping algorithm, may be configured to keep a dissimilarity between the image data of the modified source face and the image data of the target face (i. e., before the blend process) below a Cosine Distance of 0.4. The Cosine Distance may be based on the cosine of the angle between two non-zero vectors in a multidimensional space, in particular an RGB color image data space, for the feature vector representing the image data of the motor enhanced face and the image data of the target face.
[0031]
[0023] In some embodiments of the method, the plurality of computer-generated head images may in particular be generated by generative artificial intelligence and may be based on Caucasian and non-Caucasian face images as well female and male face images in a plurality of age groups. In addition or alternatively, computer-generated head images of the plurality of computer-generated head images may differ in positions of face landmarks such as eye shape and nose size, relative positions between the face landmarks, and / or sizes of the face landmarks. In addition or alternatively, computer-generated head images of the plurality of computer-generated head images may differ in appearance features such as skin texture and color.
[0032]
[0024] In some embodiments of the method, a gender of the patient and a gender associated with the source face may coincide.
[0025] In some embodiments of the method, the re-identification validation process may include, for calculating the re-identification score, a face detection step, optionally an alignment step, and a feature vector representation step. The feature vector representation step may be based on a feature vector defined in the RGB color image data space for the respective face image.
[0033]
[0026] In some embodiments, when it is determined that the calculated re-identification score for each frame passes a preset threshold, the method may further comprise the step of outputting the motor enhanced and de-identified video data as de-identified patient video data for anonymous motor assessment of a disease stage by a human rater or an Al incorporating algorithm for automized video analysis.
[0034]
[0027] In some embodiments, when it is determined that the calculated re-identification score for each frame passes a preset threshold, the method may further comprise the step of outputting the motor enhanced and de-identified video data as de-identified training video data for training an Al incorporating algorithm for automized video analysis of a patient’s motor capabilities within the motor assessment.
[0035]
[0028] In another aspect, a computer-implemented method for generating a training video data set for training an Al incorporating algorithm for automized video analysis of a patient’s motor capabilities comprises, for de-identifying each of the plurality of patient video data, performing the method as described above, thereby generating a plurality of motor enhanced video data as the training video data set.
[0036]
[0029] In some embodiments of the computer-implemented method for generating a training video data set, at least one patient video data may be de-identified at least twice based on respective different computer-generated head images such that the training video data set includes at least two motor enhanced and de-identified video data that are based on the same patient video data.
[0037]
[0030] In some embodiments of the computer-implemented method for generating a training video data set, a machine learning model or an Al algorithm for automized video analysis of a patient’s motor capabilities may be trained using the de-identified training video data set.
[0038]
[0031] In another aspect, a computer program comprises instructions which, when the program is executed by a computer device, cause the computer device to carry out the method as described above.
[0039]
[0032] In another aspect, a system for generating a motor enhanced and de-identified video data for training an Al incorporating algorithm for automized video analysis of a patient’s motor capabilities is disclosed. The system comprises a storage medium storing an image data library comprising a plurality of computer-generated head images, wherein the computer- generated head images include respectively image data of a face for being used as image data of a source face in an anonymization process, and a data processing device, such as a computer device, including a non-transitory computer readable medium storing instructions that are executable by the data processing device and that, upon such execution, cause the data processing device to perform the method as described herein.
[0040]
[0033] In another aspect, a use of motor enhanced and de-identified video data output by the method as described above is disclosed for training an Al incorporating algorithm for automized video analysis of a patient’s motor capabilities or for anonymously assessing a disease stage of a patient.
[0041]
[0034] Specifically, with respect to generating de-identified video data for training an Al incorporating algorithm for automized video analysis of a patient’s motor capabilities, it is noted that the de-identification process of the frame may include an anonymization process and a re-identification validation process only, as long as the de-identification is sufficiently ensured. This may already improve over the prior art as will be acknowledged by the skilled person. Thus, in particular in the context of generating de-identified training video data, the blend process can be considered optional as the concepts and methods, which do not use any blending, may already have the herein discussed advantages over the prior art. For example, in a further aspect, a computer-implemented method for generating a training video data set for training an Al incorporating algorithm for automized video analysis of a patient’s motor capabilities is disclosed, wherein the training video data set comprises a plurality of de- identified video data generated from a plurality of patient video data. That method comprises, for de-identifying each of the plurality of patient video data:
[0042] - receiving a patient video data that comprise a plurality of frames, wherein the frames of the plurality of frames include respective image data of a face of a patient; the face of the patient has, for example, facial features for a medical motor assessment;
[0043] - receiving, e. g., from a library of computer-generated head images, image data of a (first) computer-generated head image and identifying the image data of a face of the (first) computer-generated head image as image data of a source face;
[0044] - in each frame of the plurality of frames: identifying the image data of the face of the patient as image data of a target face; based on the source face, performing a de-identification process of the target face that includes an anonymization process and a re-identification validation process, wherein the anonymization process includes a face swapping process configured to replace the image data of the target face with image data of a modified source face, wherein attribute feature data are extracted from the image data of the target face and the image data of the modified source face is derived from the image data of the source face and the attribute feature data; for example, the modified source face shows facial features of the face of the patient for the medical motor assessment, thereby generating a de-identified frame including image data of the modified source face (for a de-identified patient), and the re-identification validation process is configured to calculate a re-identification score for the de-identified frame that indicates to what extent the modified source face resembles the target face of the patient;
[0045] - based on the re-identification scores, generating, by the computer device, the de-identified video data based on the de-identified frames. For example, in some embodiments, the de- identified frames may be output as de-identified video data of a de-identified patient, if the calculated re-identification score for each de-identified frame passes a preset threshold, (thereby ensuring that the de-identified patient in each de-identified frame does not resemble the patient). Other approaches, in case not all de-identified frames can be used due to not passing the threshold, are disclosed herein as well.
[0046]
[0035] It is noted further with respect to the generation of de-identified video data for training purposes, the generation of a plurality of “motor enhanced” video data (in this case enhanced without the blend process, thus, primarily by the face swapping process) may use the image data of the modified source face for the training video data set. Generally, the herein discussed aspects that do not explicitly relate to the blending, similarly apply to the anonymization process not including the blending process. For example, when the blending process can be considered optionally in the respective context, aspects or steps disclosed for the image data of the motor enhanced face may apply similarly to the image data of the modified source face (e. g., within the re-identification validation process), as will be acknowledged by the skilled person. In some embodiments, for example, the training video data set includes preferably at least two de-identified video data that are based on the same patient video data, however, respectively processed with the de-identification process based on different computergenerated head images; respective advantages resulting therefrom are disclosed herein.
[0047]
[0036] The herein disclosed inventive concepts and their various aspects can provide for specific aspects, such as an irreversibility of the de-identification, a training data specificgeneralization when generating training data for medical applications, an effectiveness of the de-identification, and data utility preservation of the training data in medical applications:
[0048]
[0037] “Irreversibility” The face swapping process and the blend process (as proposed herein) can generate and blend a modified source face with the original patient’s face and, thereby, blend existing features of the patient’s face onto the modified source face to form the motor enhanced face. The face swapping and blending is a highly nonlinear image processing function that can be machine learned partly through a training process involving many iterations and adjustments to model parameters. This makes it impossible or extremely difficult to reverse the de-identification process and reconstruct the patient’s face, once the face has been anonymized. Additional steps can be taken, such as adding a random noise to the resulting image or a region of the image throughout the deepfake process, which makes the reversibility of the results even harder.
[0049]
[0038] “Effectiveness”'. In the re-identification verification step, each frame can be verified (in line with the inventive concepts after anonymization and (optionally) blending, if applied) to make sure that patient-specific facial features are not recognizable in the de-identified frame. Thus, the de-identification fulfils the criteria - that is required for a suitable anonymization method
[0050] - to reliably obscure or alter facial features, thereby preventing the identification of a patient. This effectiveness of the de-identification applies to automated facial recognition systems as well as to a re-identification by humans who may know the patient, for example.
[0051]
[0039] “Training data specific-generalization in medical application”'. A face swapping process (as applied in the invention) can be based on recognized facial landmarks. Thus, the face swapping process performs the anonymization independently of any racial or ethnicity features and fulfils the criteria - that is required for a suitable anonymization method - to operate well across different demographics, including various ages, genders, and ethnicities as well as in different lighting and environmental conditions.
[0052]
[0040] “Training data utility preservation in medical application”'. The face swapping and (optionally) blending, if applied, (as disclosed herein) preserves the utility (usability) of the training data for its intended purpose, such as extracting facial features important for PD or HD evaluation such as evaluating PD hypomimia. Specific features, which are preferably preserved, relate to the preservation / transfer of dynamic motoric features onto the training video data such as:
[0053] - a blink rate: a low blink rate can contribute to the appearance of a "staring" expression (for example, PD patients may develop a reduced blink rate);
[0054] - a smile: a decreased formation of a smile can be observed by the movement of the mussels around the mouth for example, PD patients may develop a decreased ability to smile spontaneously or voluntarily);
[0055] - a general expressiveness: neurological motor disorders such as a reduced activity of the face mussels can be observed in a video stream (for example, PD and HD affect the overall expressiveness of a patient’s face, including the ability to raise eyebrows, wrinkle the forehead, and other spontaneous mimics / face expressions).
[0056]
[0041] With respect to the training data utility preservation, it is noted that the inventors realized that these dynamic motoric features may be assessed anonymously by a human rater or an Al incorporating algorithm for automized video analysis even if individual frames may not be characterized by an optimal face swapping result. In other words, the blend process may decrease quality of an individual frame when look at it singularly; however, anonymous motor assessment of a disease stage may be enhanced due to the added facial features of the target face overlaid onto the face swapping result (modified source face).
[0057]
[0042] The herein disclosed procedure allows to achieve inter alia the creation of a more balanced training data sets of de-identified patient video data in terms of age, gender, ethnicity, thereby limiting a potential bias of the AI / ML incorporating algorithm. For example, this may address technically a bias that otherwise would occur. For example, if there are only two female participants in the training data set, both of which have weak symptoms, the trained algorithm may then derive that women have weak symptoms in general, thus consistently underestimating their disease severity.
[0058]
[0043] Furthermore, the herein disclosed procedure may address technically the creation of a large training data set (i. e., data augmentation) by turning the video of one person into multiple people with different faces. A training data set technically enlarged in this manner may help an Al algorithm to learn to focus on the movement, rather than on the characteristics of a person performing the movement.
[0059]
[0044] Moreover, the herein disclosed procedure may address technically a rater’s bias. If such a bias is identified, it may be corrected for by presenting to the rater the same patient video multiple times, however each time process with a different source face.
[0060]
[0045] Even furthermore, as outlined above, the original patient videos may need to be stored and processed with the highest confidentiality protection, which makes any analysis cumbersome. As the de-identification is a one-time activity, any de-identified training video data that has passed the Re-ID evaluation process can be moved to a storage and processing platform that is installed for de-identified patient data. Thereby, the de-identified training video data can be made accessible, for example, to a project team for generating an Al incorporating algorithm for automated video analysis of patients such as PD and HD patients using the de-identified training video data to train the Al algorithm. It will be understood that this is much more practical, as the bulk of human-derived pharmaceutical research data is de-identified, and thus more powerful computing platforms are available at this lower level of data confidentiality.
[0046] Besides the above indicated use of the herein disclosed concepts in the field of development of AI / ML incorporating algorithms for automized / automated video analysis of patients, the inventive concepts may further be applied, for example, in de-centralized cellphone-based video monitoring of patients. For example, if a patient performs a remote test including, e. g., tipping on cellphone the execution of that test can at the same time be further observed by an online video acquisition. A physician, generally a human or non-human observer, can study the acquired video. However, to fulfill privacy requirements, a de-identified video stream may be generated, showing the overall scenery of the remote test, however, the remote test being executed by an anonymized patient.
[0061]
[0047] Other features and aspects of this disclosure will be apparent from the following description and the accompanying drawings.
[0062] Brief Description of the Drawings
[0063]
[0048] The accompanying drawings, which are incorporated herein and constitute a part of the specification, illustrate exemplary embodiments of the disclosure and, together with the description, serve to explain the principles of the disclosure. In the drawings:
[0064] Fig. 1 is an overview illustrating generally the de-identification of a frame of a patient video in line with the inventive concepts;
[0065] Fig. 2 is an exemplary illustration of training an Al incorporating algorithm for automized video analysis of patients based on a training video data set of de-identified videos of, e. g., PD or HD patients;
[0066] Fig. 3 is an exemplary illustration of the effect of a blend process as used in line with the inventive concepts;
[0067] Fig. 4 is an exemplary illustration of de-identification of one patient video frame based on two different artificial head images for generating a more un-biased training video data set;
[0068] Fig. 5 is a flowchart of an exemplary detailed de-identification process including a face swapping process, a blend process, and a re-identification validation process; and
[0069] Fig. 6 is a generic flowchart of using de-identification of patient video data for assessing an, e g., PD or HD disease stage.
[0070] Detailed Description
[0071]
[0049] The following is a detailed description of exemplary embodiments of the present disclosure. The exemplary embodiments described therein and illustrated in the drawings are intended to teach the principles of the present disclosure, enabling those of ordinary skill in the art to implement and use the present disclosure in many different environments and for many different applications. While, for example, method steps are shown in the drawings in a specific order, this should not be understood as requiring that such operations to be performed in that order as the skilled person may consider different orders suitable as well. Therefore, the exemplary embodiments are not intended to be, and should not be considered as, a limiting description of the scope of patent protection. Rather, the scope of patent protection shall be defined by the appended claims.
[0072]
[0050] The disclosure is based in part on the realization by the inventors that an automatic, AI / ML- based system can help addressing the before mentioned issues pertaining to, for example, bias and variability in human raters’ performance.
[0073]
[0051] Moreover, the inventors realized that for motor assessment and ensuring patient privacy, standard procedures of anonymization or de-identification in computer vision such as pixelization and blurring techniques (where faces in the frames are replaced with a coarsely pixelated or a smoothly blurred version of the original data) as well as redaction techniques (where faces are blanked in the video frames, for example, with black geometrical shapes) are either destroying any information needed for motor assessment or lack sufficient patient privacy.
[0074]
[0052] In contrast, synthesis techniques in line with the inventive concepts (where a target face of a patient can be replaced in a sequence of frames with a different (motor enhanced source) face having the “essentially” same mimic of the patient) may maintain the information required for motor assessment, while ensuring patient privacy.
[0075]
[0053] In this context, the inventors considered anonymization methods that rely on generative adversarial networks (GANs) as preferred implementations for the face swapping algorithm. For example, the inventors propose to design and modify GANs specifically to protect an individual’s privacy while also preserving and preferably enhancing important facial features. In particular, the facial features were identified as crucial for assessing motor capabilities, such as observing facial expression and / or the analysis of eye blinking frequency and lip movement. The inventors realized that these types of attribute features - while being important for analyzing facial expressions - are not inherently associated with a person's identity. Therefore, by retaining these features, GAN-based anonymization techniques were identified as a well-suited start point for allowing motor analysis while still maintaining the privacy of patients.
[0076]
[0054] As will be apparent from the following description, the de-identification method proposed herein for generating the de-identified training video data in line with the inventive concepts can be configured to ensure that (1) the video data is modified in such a way that a patient (the original patient) is no longer identifiable to a human observer,
[0077] (2) the motor disease symptoms are preserved, including, for example, any (reduced) facial expressions (mimic), and
[0078] (3) artifacts are not introduced that may impact the subsequent automated video analysis.
[0079]
[0055] As proposed herein, swapping a face of a patient with a computer-generated face in combination with the blending of the original patient face will allow preserving facial features in an interpretable way within de-identified training video data, while making the patient, on which the de-identified training video data is based, unrecognizable.
[0080]
[0056] The overview shown in Fig. 1 illustrates generally the de-identification of a frame 1 of a patient video that results in a de-identified frame, herein referred to as a motor enhanced frame 3. The de-identification is performed specifically to de-identify the image data 1A of a face of a patient (the target face) such that the motor enhanced frame 3 includes image data
[0081] 3 A of a motor enhanced face. The de-identification is specifically performed in a manner that facial features of the target face are transferred into the motor enhanced frame 3. The motor enhanced frame 3 - as part of a de-identified video data - can then be provided within a training video data set for training a respective algorithm or performing a medical motor assessment. Fig. 2 illustrates exemplarily the training of an Al incorporating algorithm 5 with a training video data set 7 created by respectively de-identified video data 9, each comprising motor enhanced frames 3 with image data 3 A of motor enhanced faces.
[0082]
[0057] The de-identification uses a face swapping process 11, a blend process 13, and a subsequent frame specific re-identification (re-ID) evaluation 15 with the aim to generate the motor enhanced face as “unidentifiable” by machine or human with respect to the respective target face. For that purpose, human differences can be introduced during the face swapping process 11 by variation of face landmarks such as eye shape and nose size variation as well as artificial differences such as change in skin texture and color. Additionally, the target face can be “mixed” into the motor enhanced face as explained below within the blend process 13.
[0083]
[0058] As illustrated in Fig. 1, the face swapping process 11 can be based, for example, on a known face swapping algorithm (such as the above-mentioned GHOST algorithm) and a computergenerated (artificially generated) head image. Specifically, image data of a face of the head image is used as image data 17A of a source face in the face swapping process 11. The face swapping process 11 uses further the image data 1 A of the patient’s face to derive facial features and transfer these facial features from the target face to the source face as illustrated by arrow 19. The face swapping process 11 generates image data 21 A of a modified source face within a frame 21 that otherwise includes the image content of frame 1.
[0084]
[0059] By performing the blend process 13 after the face swapping process 11, the generation of the image data 3A of the motor enhanced face can be based on the image data 21 A of the modified source face (output from the face swapping process 11) and the - original - image data 1A of the patient’s face as illustrated by arrow 23.
[0085]
[0060] The frame specific re-ID evaluation 15 compares the image data 1A of the target face with the image data 3 A of the motor enhanced face to ensure sufficient de-identification as illustrated by arrow 25. For example, there can be calculated a distance between identity feature vectors of the frame 1 (here limited to the image data 1 A of the face region within the frame 1) and the motor enhanced frame 3 (similarly limited to the image data 3 A of the face region within the “artificial” replacement frame).
[0086]
[0061] The face swapping process 11, the blend process 13, and the re-ID evaluation procedure 15 are exemplary described below in more detail in connection with the remaining drawings. As will be understood by the skilled person, the face swapping process 11, the blend process 13, as well as the re-ID evaluation 15 can be specifically adapted to respective diseases to be monitored, such as PD or HD. As it is known, observable changes of the motor characteristics (in short motor changes) of PD or HD patients may relate to the eye regions of a patient’s face (such as for assessing any changes in eye blinking frequency) and the mouth region of a patient’s face (such as for assessing any changes in lip movement). Accordingly, for those respective regions - being essential for PD or HD evaluation, facial attributes such as eye blinking and lip movement can be preserved (transferred to the motor enhanced (modified source) image) when de-identifying a video data (video stream) with the face swapping and blend processes. The blend process 13 may emphasize the presence of those facial attributes for the data analysis, although the overall appearance of the anonymized person may become visually less appealing when looked at a single (static) frame. However, in a sequence of frames that is to be assessed, the access to dynamic aspects can be emphasized by the blend process 13 such that an automated as well as a human-based identification of the dynamic aspects are improved.
[0087]
[0062] To improve the information (facial attributes) of those essential facial regions, an increased similarity between the original mimic of the patient and a mimic of a de-identified face can, for example, be reached in the de-identified face by adjusting blend parameters of weight and erosion of the modified source face and the patient’s face. Blend parameter(s) may adjust the overlay and reduction of intensity opacity of the modified source face generated by the face swapping process 11. These blend parameters can be set to provide a dissimilarity between the generated face and the original face, for example, above or equal a Cosine Distance of, e. g., 0.4. The blend parameters generally allow increasing or decreasing the Cosine Distance, whereby the Cosine Distance corresponds to “1 -Cosine Similarity” and the Cosine Similarity is calculated using a mathematical formula that involves the dot product of the two respective identity feature vectors and their magnitudes. Generally, the Cosine Distance is based on the cosine of the angle between two non-zero vectors in a multi-dimensional space (here the RGB color image data space) given for the identity feature vectors representing the motor enhanced face image and target face image (specifically, representing the image data of the motor enhanced face and the image data of the target face). For the concept of Cosine Distance, see also https: / / en.wikipedia.org / wiki / Cosine_similarity and https: / / www.itl.nist.gov / div898 / software / dataplot / refman2 / auxillar / cosdist.htm.
[0088]
[0063] Fig. 3 illustrates how the blending of a target face 27 (specifically image data 1 A) onto the modified source face (image data 21 A) can be varied, thereby affecting the Cosine Distance to 0.42 for an image 27A and to 0.83 for an image 27B.
[0089]
[0064] For the face swapping process 11, for example, a library of computer-generated (such as Al generated) source heads / faces can be provided. Generally, computer-generated source heads / faces can be used to generate de-identified training video data that are in line with an intended / desired characteristic of the (de-identified) training video data set. The characteristics of the training video data set may include, e. g., diversity requirements with respect to gender, age, or ethnicity. For example, the image data 17A of the source faces may be generated using generative artificial intelligence, based on Caucasian and non-Caucasian, female and male face images. Respective libraries of Al generated source faces are available, for example, from public data banks or may be generated specifically for a project.
[0090]
[0065] Referring to the support of the development of an algorithm for automized analysis of videos of patients (see also Fig. 2), the inventive concepts can contribute by providing a training video data set that is based on “real” patient videos; however, the videos are anonymized with respect to any patient data. Specifically, training video data in line with the inventive concepts can include anonymized (de-identified) video data that are generated from patient videos. The training video data set can then be used, for example, for developing and testing the AI- incorporating algorithm for automized video analysis of motor capabilities of PD or HD patients.
[0091]
[0066] It is noted that the herein disclosed concepts can be specifically suitable for providing a sufficiently large de-identified training video data set under circumstances where the sample size of available patient videos is limited. As training data in the medical field often originate only from a few patients, the herein disclosed concepts may specifically be of relevance for the development of a bias-free / bias-reduced AI / ML incorporating algorithm for automized video analysis of patients. Specifically, using the herein disclosed concepts, an AI / ML incorporating algorithm can be developed that has a reduced bias regarding gender, age, or ethnicity etc. despite of having available only a limited number of sample “real” patient video data. For this, the herein disclosed concepts allow technically generating a set of sample training data in a sufficient amount and in un-biased (reduced bias) manner to ensure proper training of the Al incorporating algorithm.
[0092]
[0067] To illustrate the concept of designing desired training data for generating a more un-biased training video data set, Fig. 4 illustrates how one patient video frame 29 can be de-identified based on two different artificial head images 31 A, 3 IB at first into frames 33 A, 33B with image data of modified source faces. Using the respective head images 31 A, 3 IB in the blend process, the image data of modified source faces can optionally be further processed to result in differing motor enhanced frames 35 A, 35B that can be associated with faces of the same or different age groups, the same or different general appearances etc. Each of the motor enhanced frames 35 A, 35B may be characterized by a Cosine Distance larger than 0.4 with respect to the patient video frame 29 (as indicated by arrow 37).
[0093]
[0068] Moreover, the automized use of de-identified training video data generated in line with the inventive concepts can allow error detection and respective adaptation during the development of an AI / ML incorporating algorithm. For example, an AI / ML incorporating algorithm may be configured to identify motor characteristics of a patient to monitor and assess the patient’s temporal changes of the motor characteristics when, for example, comparing videos of a tapping movement repeatedly taken at different stages during a drug development study.
[0094]
[0069] Fig. 5 shows a schematic flowchart of an exemplary embodiment of a computer-implemented method for generating a motor enhanced and de-identified video data 9 from a patient video data 43. For this purpose, the method includes a de-identification process performed, e. g., in a computer device 45 having, for example, a non-transitory computer readable medium, a processor, a storage medium etc. and optionally having access to an external storage medium. The de-identification process includes, for facial anonymization, a frame specific anonymization process that is based on the face swapping process 11 (for replacing a respective target face with an artificial face in a frame 1) and the blend process 13 (for optionally enhancing facial motor features in the artificial face, resulting in the motor enhanced frame 3). The de-identification process includes further, for approving the achieved anonymization, the re-identification process 15 (for measuring a (dis-)similarity). The (dis-)similarity can be validated to ensure a sufficient difference between the target face and the motor enhanced faces in the de-identified frames of the motor enhanced and de-identified video data.
[0095]
[0070] Fig. 5 shows several patient video data 43 that are received, e. g., by the computer 45 device (and stored optionally on an internal storage medium) and each comprises a plurality of frames 1. The video data 43 recorded, for example, a MDS-UPDRS test of the patient and the frames 1 include respective image data 1 A of a face of the patient. The face of the patient shows facial features for a medical motor assessment in line with the MDS-UPDRS testing of inter alia static and dynamic facial features.
[0096]
[0071] Fig. 5 shows further a library 47 of computer-generated head images 47A. The library 47 may be stored on a storage medium 46 and comprises, preferably, a plurality of head images 47A with a large variety in appearance of respective “source faces” in age (e. g., “young” and “old” face images), gender (e. g., female and male face images), ethnicity (e. g., Caucasian and nonCaucasian face images), etc. The computer-generated head images 47A may differ in positions of face landmarks such as eye shape and nose size, relative positions between the face landmarks, and / or sizes of the face landmarks. In addition, the computer-generated head images may differ in appearance features such as skin texture and color. For the deidentification process, a first computer-generated head image 17 is received, e. g., by the computer device 45. In the head image 17, the computer device 45 can identify the image data of a face that is then used as an image data 17A of a source face for the face swapping process 11.
[0097]
[0072] Generally, the de-identification is performed for each frame. The de-identification, instructions executed by the computer device 45, operates primarily on the image data 17A of the face of the “artificial” head (illustrated as a line through square in Fig. 5) and uses the target face, specifically, the image data 1A of the patient’s original face (illustrated as a dashed line square in Fig. 5) for enhancing motor features. The patient’s original face may be identified by, e.g., the computer device 45. The de-identification process outputs a frame 3 showing a person with an anonymized face, specifically, image data 3 A of a motor enhanced face (illustrated as a dash-dotted line square in Fig. 5).
[0098]
[0073] The anonymization process comprises, for deriving the image data 1 A of the target face, the steps of face detection 11 A and face segmentation 1 IB that are performed on the original frames 1. The steps 11A, 1 IB identify one target / patient face (or more “target faces”) within the video data 43, specifically frames 1. The target face is to be processed for anonymization of each frame 1. For the steps 11A, 1 IB, well established mature solutions exist.
[0099]
[0074] In the face swapping process 11, the image data 1A of the target face are replaced with image data 21A of a modified source face. The image data 21A of the modified source face is derived from the image data 17A of the source face under consideration of attribute feature data derived from the target face 1 A. The face swapping process 11 generates a de-identified frame 21 including image data 21 A of the modified source face. In other words, a source face 17A is received, such as selected from the library 47 of computer-generated source heads or otherwise derived, and is processed with, e. g., a face swapping algorithm 11C to generate the respective modified source face. For example, depending on the gender of a patient, either a male or female artificial head may be received, such as selected from the library 47, and then may be processed to replace the target face. In some implementations, the face swapping process 11 can be based on a Generative High-Fidelity One Shot Transfer (GHOST) algorithm as described in “GHOST — A New Face Swap Approach for Image and Video Domains” by A. Groshev, et al., IEEE Access, vol. 10, pp. 83452-83462, 2022, doi:
[0100] 10.1109 / ACCESS.2022.3196668. The GHOST algorithm is a DeepFake algorithm with high accuracy in preserving eye gaze and facial expressions of the original subject / target face. Specifically, the GHOST algorithm can maintain facial features in a video stream such as eye blinks, gaze, and mouth movements, being important for identifying symptoms of motion affecting diseases such as PD and HD. Besides the above-mentioned GHOST algorithm, other face swapping algorithms such as “deepfake algorithms” can be implemented in the anonymization process. For example, the following “deepfake algorithm” may preserve the facial attributes (using attribute feature data derived from the target face) while replacing the style of the face:
[0101] - DeepFaceLab (presenting an integrated, flexible and extensible face-swapping framework);
[0102] - SimSwap (presenting an Efficient Framework For High Fidelity Face Swapping);
[0103] - MobileF aceSwap (presenting a Lightweight Framework for Video Face Swapping);
[0104] - FaceShifter (presenting a High Fidelity And Occlusion Aware Face Swapping).
[0105]
[0075] The above algorithms provide control on the source face that is projected on the target face in an original image, while maintaining the faci al features / texture. Thus, independent of the type of patient be it old, young, male, or female, the lighting conditions, and the face view (front facing image or profile image or occluded image), a textured face can be output in a realistic video stream providing the underlying motoric features for video analysis.
[0076] The anonymization process can further use landmark tracking 1 ID that accounts for discrepancies in facial shape between the source face and the target face. Optionally, the anonymization process can additionally apply super-resolution post processing 1 IE for each frame to remove artifacts and, thereby, increase video quality. The output of the anonymization process is the de-identified frame 21 that has replaced therein the target face with a modified source face.
[0106]
[0077] Generally, the face swapping process 11, in particular the face swapping algorithm 11C, can be configured to keep a desired minimum dissimilarity between the image data 21 A of the modified source face and the image data 1 A of the target face (i. e., even before the blend process). A dissimilarity before the blend process may, for example, be below / above a Cosine Distance of 0.4 calculated for respective image vectors, see also below for further distance metrics.
[0107]
[0078] Generally, the blend process 13 may further increase the Cosine Distance. In the blend process 13, the motor enhanced frame 3 is generated. In the motor enhanced frame 3, the image data
[0108] 1 A of the target face are overlaid onto the de-identified frame 21, thereby forming the image data 3 A of the motor enhanced face. The motor enhanced face includes facial features of the target face for enhancing the medical motor assessment.
[0109]
[0079] Specifically, and in particular in combination with the face swapping process 11, the blend process 13 is configured to introduce a dissimilarity between the image data 3 A of at least one motor enhanced face and the image data 1 A of the target face to be a above a preset threshold value T as determined in the re-ID validation process 15. For example, an adjustable parameter of the blend process 13, such as a weight parameter and / or an erosion parameter of the target face when overlaid on the modified source face, is set to adapt the contribution of image data 1 A of the target face to the image data 3 A of the motor enhanced face
[0110]
[0080] In the de-identification process, the motor enhanced frame is then provided to the re-ID validation process 15 for further evaluation of the degree of anonymization. In an exemplary embodiment, the re-ID validation process 15 is configured to ensure, for each output de- identified video data 9 that the face swapping process 11 has succeeded in a sufficient privacy. For example, the re-identification evaluation process 15 can be configured to avoid accidentally replacing the patient’s face with a face that looks very similar.
[0111]
[0081] For example, a re-identification score - calculated by, e. g., the computer device 45 when performing the re-ID validation process 15 for the motor enhanced frame 3 - may indicate to what extent the motor enhanced face resembles the target face of the patient. For example, depending on the calculated re-ID values, the de-identification can be considered complete, and the computer device 45 may generate the motor enhanced and de-identified video data 9 from the plurality of the motor enhanced frames 3. Alternatively, further steps (indicated by arrow 49) may need to be taken, if one or more motor enhanced frames have an insufficient re-ID value.
[0112]
[0082] Generally, the re-ID validation process 15 includes, for each frame, a comparison of the generated artificial faces with the original faces. The comparison may be based on a calculation of, for example, a Re-ID score. The Re-ID score is selected to enable determining whether the artificial person in the de-identified video data resembles the patient to some degree, i. e., the image data would be perceived as the patient. The Re-ID score can be based on the concept of visual similarity; for example, a high Re-ID score (e. g., a Cosine Distance) corresponds to successfully having de-identified the face and, thus, the video data.
[0113]
[0083] Specifically, an exemplary face re-identification (verification) is a process that allows confirming whether two face images (given in two frames to be compared) represent the same individual. It is noted that similar technologies are known for applications, such as within security systems, authentication systems, and identity verification systems. A re-identification process can involve a sequence of steps of detection, alignment, representation, and verification as illustrated below in more detail.
[0114]
[0084] In the first step 15A (detection), a face is detected within each frame. Face detection algorithms -also used for the face swapping process 5 described above - are well known and can accurately detect a bounding box around a face such that any surrounding noise can be removed from the frame. An exemplary face detection algorithm is the “mediapipe” (https: / / arxiv.org / pdf / 1906.08172) that outputs a bounding box containing a detected (cropped) face within a noise removed frame, as well as six landmark information, such as the position in the frame of the left eye, the right eye, the nose tip, the mouth, the left eye tragion, and the right eye tragion.
[0115]
[0085] In the next step 15B (alignment), for each frame, the detected face can be aligned to ensure that the landmarks (generally facial features) are given in a standardized orientation within a noise removed frame. For example, the image of the detected face in the noise removed frame can be rotated and scaled based on the positions of the landmarks extracted in the detection step to result in an aligned face within the noise removed frame. It is noted that the alignment step can significantly improve the accuracy of the subsequent steps of the re-identification process.
[0116]
[0086] In the next step 15C (representation), the noise removed frame with the aligned face image is processed to create a numerical representation of the face in the form of a feature vector. Face representations as feature vectors are well known. For example, a convolutional neural network (CNN) can be trained to extract features from a face image that are useful for distinguishing between different faces. The CNN may represent the frame with the face image as a feature vector in a high-dimensional space, where each dimension represents a particular aspect of facial features. An exemplary representation as a feature vector can be generated with state of the art algorithms such as the algorithm “FaceNet” (https: / / arxiv.org / pdf / 1503.03832) that outputs, for example, a feature vector of the size 512.
[0117]
[0087] In the next step 15D (verification), the feature vectors of the two faces are compared. Specifically, the feature vector determined for the motor enhanced source face generated by the anonymization process for the de-identified patient is compared with the feature vector determined for the target face of the patient (the patient’s face before anonymization). As an exemplary re-identification score, a comparison of feature vectors for the image data 3 A of the motor enhanced face and the image data 1 A of the target face can use a distance metrics between those feature vectors, such as the well-known Cosine (Dis-)Similarity / Cosine Distance, the Euclidean distance, or the Euclidean L2 distance; see, for example, https: / / www.itl.nist.gov / div898 / software / dataplot / refman2 / auxillar / cosdist.htm and “Choice of distance metrics for RGB color image analysis” by Amadou T. SANDA MAHAMA et al., IS&T International Symposium on Electronic Imaging 2016 Color Imaging XXI: Displaying, Processing, Hardcopy, and Applications, pages COLOR-349.1-4.
[0118]
[0088] In the final step 15E (decision), it is then decided based on the verification, whether the frames 3 can be set together to form the de-identified video data 9 or whether further revision / improvement steps are needed. For example, a threshold value T for the reidentification score can be set for the decision such that if a re-identification score calculated for two frames passes the preset threshold value T, it is ensured that the “artificial” patient with the anonymized face does not resemble the patient with the face before anonymization. In other words, if a distance between the two feature vectors is below a certain threshold distance, the faces are considered to be similar, meaning that the anonymization was not sufficient, while a distance above the threshold distance provides for sufficient anonymization.
[0119]
[0089] Alternative steps indicated by arrow 49 may include, for example, that the de-identification may be performed for a second source face, or generally a different source face may be applied. For example, when it is determined in the re-ID evaluation that the calculated reidentification score for none or a preset maximum number of motor enhanced frames generated from a patient video data 43 does not fulfill the distance requirement, the computer device may receive, in particular select from the library 47, image data of a second computer- generated head image 47A. It may then identify the image data of a face of the second computer-generated head image as image data 17 A, and perform the de-identification process of the target face based on that source face.
[0120]
[0090] Alternatively, when the re-ID verification determines that a re-identification score calculated for one of the motor enhanced frames 3 (within a sub-sequence of frames generated from a patient video data 43) does not pass the preset threshold value T, that frame, which does not pass the preset threshold value T, can be replaced with a motor enhanced neighboring frame, which passes the preset threshold value T and is next to or close to in a sequence of motor enhanced frames 3 with respect to the motor enhanced frame, which does not pass the preset threshold.
[0121]
[0091] Alternatively, the calculated re-identification score may fulfill the requirement for none or a preset maximum number of frames of motor enhanced frames 3, the blend process 13 can be adjusted, e. g., by the computer device 45 or a user, in its parameter. Exemplary parameters include a weight parameter and / or an erosion parameter of the target face when overlaid on the modified source face. By the adjustment, the parameter may specifically be set to generally increase a Cosine Distance between the image data 3 A of the motor enhanced faces and the image data 1 A of the target face.
[0122]
[0092] If the Re-ID value is good for all frames or a respective replacement was performed, a de- identified video-data 9 including the motor enhanced source face within the sequence of frames 3 is output, whereby the re-ID values are above the predetermined threshold value T for all the de-identified motor enhanced video frames 3.
[0123]
[0093] Fig. 6 summarizes in a generic flowchart the medical use of the herein identified deidentification of patient video data for assessing, e. g., PD or HD disease stages.
[0124]
[0094] Patient video data 43 can be de-identified using the computer device 45. Specifically, the computer device 45 includes a non-transitory computer readable medium that stores instructions that are executable by the computer device 45 and that, upon such execution, cause the data processing device to perform the method described herein (including the face swapping process 11, the blend process 13, and the re-identification evaluation 15) for generating the motor enhanced and de-identified video data 9 from the patient video data 43 using, for example, a library 47 of computer-generated head images.
[0125]
[0095] Moreover, the de-identified video data 9 or the patient video data 43 directly can be evaluated by the Al incorporating algorithm 5 for automized video analysis of a patient’s motor capabilities within the motor assessment. The algorithm 5 may, for example, output a result 51 of an MDS-UPDRS test of the patient. Alternatively or in addition, a human rater 53 my evaluate the motor enhanced and de-identified video data 9 to provide a respective result 51 of an MDS-UPDRS test for the patient.
[0126]
[0096] Thus, regarding the improvement of drug development, the inventive concepts can enable the use of (de-identified) videos of, e. g. PD, patients under less restrictive privacy requirements when developing an AI / ML incorporating algorithm for automized analysis of video recordings of patients as it can be used, for example, when monitoring drug efficiency during a drug study. Moreover, the inventive concepts can contribute, by enabling the development of such an algorithm, to reducing / avoiding any human uncertainty that unavoidably is introduced by motor assessment of patients by human raters such as physicians or researchers.
[0127]
[0097] In summary, the motor enhanced and de-identified video data 9 may be generated for uses such using de-identified training video data for training an Al incorporating algorithm for automized video analysis of a patient’s motor capabilities within the motor assessment or performing anonymous motor assessment of a disease stage by a human rater or an Al incorporating algorithm for automized video analysis.
[0128]
[0098] As will be understood by the skilled person, the original patient videos are to be stored and processed with the highest confidentiality protection, which makes any analysis cumbersome. As the de-identification is a one-time activity, any de-identified training video data that has passed the Re-ID evaluation process can be moved to a storage and processing platform for de-identified patient data. Thereby, the de-identified training video data can be made accessible to a project team for generating an Al incorporating algorithm for automated video analysis of patients such as PD and HD patients using the de-identified training video data to train the Al algorithm. It will be understood that this is much more practical, as the bulk of human-derived pharmaceutical research data is de-identified, and thus more powerful computing platforms are available at this lower level of data.
[0129]
[0099] Embodiments of the subject matter and the operations described herein can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described herein can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions, encoded on computer storage medium for execution by, or to control the operation of, data processing apparatus. A computer storage medium can be, or be included in, a computer-readable storage device, a random or serial access memory array or device, or a combination of one or more of them. Moreover, a computer storage medium can be a source or destination of computer program instructions encoded in an artificially-generated propagated signal. The method aspects described herein can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0130]
[0100] The term "data processing device" encompasses all kinds of devices, apparatus, and machines for processing data, including by way of example a programmable processor, a computer (device), a system on a chip, or multiple ones, or combinations, of the foregoing. The apparatus can include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[0131]
[0101] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program can be stored in a portion of a file that holds other programs or data (e. g., one or more scripts stored in a markup language document), in a single file dedicated to the computer program in question, or in multiple coordinated files.
[0132]
[0102] The method described herein can be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input data and generating output data. Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a non-transitory computer readable medium. The primary elements of a computer device are a processor for performing actions in accordance with instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, storage devices. Moreover, the computer device can be embedded in another device, e.g., a mobile telephone / device used for interacting with a patient. To provide for interaction with a user, patient, programmer etc., the methods described herein can be implemented on a computer device having a display device for displaying information to the user and generally an input device such as a keyboard by which the user, patient, programmer etc. can provide input to the computer.
[0133]
[0103] It is explicitly stated that all features disclosed in the description and / or the claims are intended to be disclosed separately and independently from each other for the purpose of original disclosure as well as for the purpose of restricting the claimed invention independent of the composition of the features in the embodiments and / or the claims. It is explicitly stated that all value ranges or indications of groups of entities disclose every possible intermediate value or intermediate entity for the purpose of original disclosure as well as for the purpose of restricting the claimed invention, in particular as limits of value ranges.
[0134]
[0104] Although the preferred embodiments of this invention have been described herein, improvements and modifications may be incorporated without departing from the scope of the following claims.
Claims
Claims1. A computer-implemented method for generating a motor enhanced and de-identified video data (9) from a patient video data (43), the method comprising:- receiving, by a computer device (45), the patient video data (43) that comprises a plurality of frames (1), wherein the frames of the plurality of frames (1) include respective image data of a face of a patient, the face of the patient having facial features for a medical motor assessment;- receiving, by the computer device (45), in particular from a library (47) of computer-generated head images (47 A), image data of a first computer-generated head image and identifying the image data of a face of the first computer-generated head image as image data (17A) of a source face;- in each frame of the plurality of frames (1):- identifying, by the computer device (5), the image data of the face of the patient as image data (1A) of a target face;- based on the image data (17A) of the source face, performing, by the computer device, a de-identification process of the frame that includes an anonymization process, a blend process (13), and a re-identification validation process (15), wherein- the anonymization process includes a face swapping process (11) configured to replace the image data (1A) of the target face with image data (21 A) of a modified source face, wherein the image data (21 A) of the modified source face is derived from the image data (17A) of the source face and attribute feature data of the target face, thereby generating a de-identified frame (21) including image data (21 A) of the modified source face,- the blend process (13) generates a motor enhanced frame (3) by blending the image data (1A) of the target face onto the de-identified frame (21), thereby forming image data (3 A) of a motor enhanced face that includes facial features of the target face for enhancing the medical motor assessment, and- the re-identification validation process (15) is configured to calculate a reidentification score for the motor enhanced frame (3) that indicates to what extent the motor enhanced face resembles the target face of the patient; and- based on the re-identification scores, generating, by the computer device (45), the motor enhanced and de-identified video data (9) from the motor enhanced frames (3).
2. The method of claim 1, wherein the blend process (13), in particular in combination with the face swapping process (11), is configured to introduce a dissimilarity between the image data (3 A) of at least one motor enhanced face and the image data (1A) of the target face to be a above a preset threshold value (T), the preset threshold value (T) in particular corresponding to a Cosine Distance of 0.4, wherein the Cosine Distance is based on the cosine of the angle between two non-zero vectors in a multi-dimensional space, in particular an RGB color image data space, for the feature vector representing the image data (3 A) of the motor enhanced face and the image data (1 A) of the target face.
3. The method of claim 1 or claim 2, further comprising:- setting the blend process (13) such that, for at least one of the plurality of the motor enhanced frames (3), the calculated re-identification score passes a preset threshold value (T), and wherein in particular an adjustable parameter of the blend process (13), such as a weight parameter and / or an erosion parameter of the target face when overlaid on the modified source face, is set to adapt the contribution of the image data (1 A) of the target face to the image data (3 A) of the motor enhanced face.
4. The method of any one of the preceding claims, wherein the re-identification score is determined by calculating a dissimilarity between the image data (3 A) of the motor enhanced face and the image data (1A) of the target face, wherein the dissimilarity is in particular determined based on a Cosine Distance that is based on the cosine of the angle between two non-zero vectors in a multi-dimensional space, in particular an RGB color image data space, for the feature vector representing the image data (3 A) of the motor enhanced face and the image data (1 A) of the target face, and wherein in particular the preset threshold value (T) is set to correspond to a Cosine Distance in a range from 0 to 2, in particular from 0.4 to 2.
5. The method of any one of the preceding claims, wherein, when determining that the calculated re-identification score of one of the motor enhanced frames (3) does not pass the preset threshold value (T), further comprising:- replacing the frame (3), which does not pass the preset threshold value (T), with a motor enhanced neighboring frame (3), which does pass the preset threshold value (T) and is next to or close to in a sequence of motor enhanced frames (3) with respect to the motor enhanced frame, which does not pass the preset threshold.
6. The method of any one of the preceding claims, wherein, when determining that the calculated re-identification score for none or a preset maximum number of frames of the motor enhanced frames (3) does not pass the preset threshold value (T), setting the blend process, in particular an adjustable parameter, such as a weight parameter and / or an erosion parameter of the target face when overlaid on the modified source face, to increase a Cosine Distance between the image data (3A) of the motor enhanced faces and the image data (1A) of the target face, respectively.
7. The method of any one of the preceding claims, wherein, when determining that the calculated re-identification score for none or a preset maximum number of frames of the motor enhanced frames (3) does not pass the preset threshold value (T), receiving, in particular selecting from the library (47) of computer-generated head images (47 A), image data of a second computer-generated head image (47 A) and identifying the image data of a face of the second computer-generated head image as image data (17A) of the source face, and performing the de-identification process of the target face based on that source face.
8. The method of any one of the preceding claims, wherein the face swapping process (11) includes a face swapping algorithm (11 C) that is selected from the group of deep fake face swapping algorithms comprising the GHOST algorithm, the DeepFaceLab algorithm, the SimSwap algorithm, the MobileFaceSwap algorithm, the FaceShifter algorithm, and the HifiFace algorithm, and / or wherein the face swapping process (11), in particular a face swapping algorithm (11 A), is configured to keep a dissimilarity between the image data (21 A) of the modified source face and the image data (1 A) of the target face below a Cosine Distance of 0.4, wherein the Cosine Distance is based on the cosine of the angle between two non-zero vectors in a multidimensional space, in particular an RGB color image data space, for the feature vector representing the image data (3 A) of the motor enhanced face and the image data (1 A) of the target face.
9. The method of any one of the preceding claims, wherein the computer-generated head image (47 A) is in particular generated by generative artificial intelligence and is based on Caucasian and non-Caucasian face images as well female and male face images in a plurality of age groups, and / orwherein computer-generated head images, in particular of the plurality of computer-generated head images (47 A), differ in positions of face landmarks such as eye shape and nose size, relative positions between the face landmarks, and / or sizes of the face landmarks and / or wherein computer-generated head images, in particular of the plurality of computer-generated head images (47 A), differ in appearance features such as skin texture and color.
10. The method of any one of the preceding claims, wherein a gender of the patient and a gender associated with the source face coincide.
11. The method of any one of the preceding claims, wherein the re-identification validation process (15) includes, for calculating the re-identification score, a face detection step (15 A), optionally an alignment step (15B), and a feature vector representation step (15C), and wherein the feature vector representation step (15C) is based on a feature vector defined in the RGB color image data space for the respective face image.
12. The method of any one of the preceding claims, wherein, when determined that the calculated re-identification score for each frame (3) passes a preset threshold, the method further comprises:- outputting the motor enhanced and de-identified video data as de-identified patient video data (9) for anonymous motor assessment of a disease stage by a human rater or an Al incorporating algorithm (5) for automized video analysis.
13. The method of any one of claims 1 to 11, wherein, when determined that the calculated re-identification score for each frame (3) passes a preset threshold (T), further comprising:- outputting the motor enhanced and de-identified video data as de-identified training video data for training an Al incorporating algorithm (5) for automized video analysis of a patient’s motor capabilities within the motor assessment.
14. A computer-implemented method for generating a training video data set (7) for training an Al incorporating algorithm (5) for automized video analysis of a patient’s motor capabilities, the method comprising:for de-identifying each of the plurality of patient video data (43), performing the method of any one of claims 1 to 11, thereby generating a plurality of motor enhanced video data (9) as the training video data set (7).
15. The method of claim 14, wherein at least one patient video data (43) is de-identified at least twice based on respective different computer-generated head images (47 A) such that the training video data set (7) includes at least two motor enhanced and de-identified video data (9) that are based on the same patient video data (43).
16. The method of claims 14 or 15, further comprising: training a machine learning model or an Al algorithm (5) for automized video analysis of a patient’s motor capabilities using the de-identified training video data set (7).
17. A computer program comprising instructions which, when the program is executed by a computer device (5), cause the computer device (5) to carry out the method of any one of the preceding claims.
18. A system for generating a motor enhanced and de-identified video data (9) for training an Al incorporating algorithm (5) for automized video analysis of a patient’s motor capabilities, the system comprising:- a storage medium (46) storing computer-generated head images (47 A), in particular an image data library (47) comprising a plurality of computer-generated head images (47 A), wherein the computer-generated head images (47 A) include respectively image data of a face for being used as image data (17A) of a source face in an anonymization process, and- a data processing device, such as a computer device (45), including a non-transitory computer readable medium storing instructions that are executable by the data processing device and that, upon such execution, cause the data processing device to perform the method of any one of claims 1 to 16.
19. Use of motor enhanced and de-identified video data (9) output by the method of any one of claims 1 to 11 for training an Al incorporating algorithm (5) for automized video analysis of a patient’s motor capabilities or for anonymously assessing a disease stage of a patient.