Facial Image De-identification for AI Training Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Access to Personal Health Information (PHI) for training artificial intelligence algorithms in healthcare is hindered by the need for patient consent, which is time-consuming, expensive, and complex to obtain, limiting the creation of large databases necessary for training.
Innovation Solution
The development of a system, AnonVisioGuard (AVG), that de-identifies facial images, allowing for the use of de-identified PHI data without patient consent, enabling unsupervised machine learning and machine vision training.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If patient consent is obtained for accessing PHI to train AI algorithms, then data access is authorized and training can proceed, but the process becomes time-consuming, expensive, and complex
Solution Approach 1:
The patent segments PHI into identified and de-identified portions. Facial images are extracted and processed separately from other patient data, with de-identification techniques applied specifically to facial features while preserving other necessary clinical data for training purposes.
Solution Approach 2:
De-identification techniques serve as an intermediary process between raw PHI and training data. By introducing this intermediate processing step, the system enables data usage without requiring direct patient consent, as the de-identified data no longer contains personally identifiable information.
2Object-affected harmful factors
If patient consent is obtained for accessing PHI, then data privacy is protected, but the time and cost to obtain consent increases
Solution Approach 1:
De-identification is performed preliminarily on all facial images before they are used for training purposes. This advance processing ensures that privacy protection is built into the data preparation stage, eliminating the need for ongoing consent management and reducing time delays.
Solution Approach 2:
Facial images are extracted and de-identified separately from the main PHI dataset. This extraction and separate processing allows privacy protection to be implemented independently, reducing the administrative burden and time required for consent acquisition.
3Productivity
If de-identified facial images are used for training, then patient consent is not required and datasets can be expanded, but data privacy protection mechanisms must be implemented
Solution Approach 1:
De-identified facial images serve as copies that preserve the visual information needed for training while removing personally identifiable features. These copied and modified images can be used freely for training purposes without requiring patient consent, thereby expanding dataset size and speed.
Data Source
AI summary
A method or system uses de-identified images collected from patients in association with de-identified data collected as part of medical care to train machine learning, machine vision, deep learning, or other algorithms that associate an outcome variable with an input image or video image(s).


