Secure Reference Dataset Portal for AI Model Validation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge in the medical field is that there are no common, ground-truth reference data to test and compare new or updated AI algorithms for oncological imaging, making it difficult to determine the generalizability and performance of AI products, especially when faced with varying patient demographics, imaging techniques, and noise processes, which are time-consuming and costly to source.
Innovation Solution
A system providing secure, limited access to curated reference datasets with confirmed truth, allowing users to evaluate AI models independently trained on different data, using a portal system and evaluation algorithms to assess performance, bias, fairness, and response to data variation, while ensuring the datasets are not used for training AI algorithms directly.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If representative data sets are made publicly available for AI training, then data accessibility and training efficiency are improved, but data security and prevention of data leakage are worsened
Solution Approach 1:
The patent introduces a secure evaluation environment as an intermediary between the reference data set and the AI model evaluation process. This environment allows models to be evaluated against the reference data without direct access to the data itself, preventing data leakage while maintaining evaluation integrity. The intermediary layer enables controlled access where only evaluation results, not raw data, are exposed to external systems.
Solution Approach 2:
The patent creates a synthetic reference data set that replicates the statistical properties and characteristics of real patient data without containing actual patient information. This copy allows extensive sharing and access for training and evaluation purposes while eliminating the security risks associated with real data. The synthetic data maintains the necessary statistical properties for valid AI model evaluation without exposing sensitive information.
2Reliability
If comprehensive validation data sets are collected to represent all patient variations, then evaluation thoroughness and generalizability assessment are improved, but time and economic resources required are worsened
Solution Approach 1:
The patent performs preliminary action by pre-generating and pre-curating the synthetic reference data set with representative variations of patient demographics, disease characteristics, and imaging parameters before actual model evaluation begins. This advance preparation eliminates the need for time-consuming data collection during the evaluation process, as the comprehensive reference data is already ready for immediate use in assessing model generalizability across different patient populations.
Solution Approach 2:
The patent uses synthetic data that copies the statistical properties and variations of real patient populations without requiring actual patient data collection. This approach maintains evaluation thoroughness by including diverse patient variations while dramatically reducing the time and resources needed for data collection, as the synthetic data can be generated efficiently without clinical infrastructure requirements.
3Adaptability or versatility
If reference data sets are made accessible for model evaluation, then performance comparison capability is improved, but risk of data misuse for training models is worsened
Solution Approach 1:
The patent implements a secure evaluation environment as an intermediary that enables performance comparison capability while preventing data misuse. This environment provides controlled access where reference data can be used for model evaluation and comparison against predicate devices, but the data remains isolated and cannot be downloaded or used for training external models. The intermediary layer maintains the necessary adaptability for evaluation while blocking harmful data transfer.
Solution Approach 2:
The patent creates a synthetic reference data set that copies only the necessary statistical properties and image characteristics required for evaluation, excluding any information that could be misused for training. This selective copying approach enables performance comparison capability while minimizing the risk of data misuse, as the synthetic data lacks the specific identifiers and detailed information needed for effective model training.
4Adaptability or versatility
If AI models are trained on diverse real-world data, then model generalizability is improved, but difficulty in controlling data quality and representativeness is worsened
Solution Approach 1:
The patent uses synthetic data that copies representative variations of patient demographics, disease phenotypes, and imaging characteristics in a controlled manner. This approach achieves model generalizability by including diverse patient representations while eliminating the data curation complexity associated with real-world data, as the synthetic data can be generated with precise control over statistical properties and representativeness without requiring complex human review and validation processes.
Data Source
AI summary
A method includes providing a computer system which includes a processor system and a memory system in operative connection with the processor system. The memory system has one or more reference datasets which have confirmed truth stored thereon. Access to data of the one or more reference datasets is secured such that, at least one of: (i) the data of the one or more reference datasets cannot be accessed for use in training artificial intelligence algorithms or (ii) the data of the one or more reference datasets is altered to diminish the viability thereof in training artificial intelligence algorithms. The method further includes providing a portal system to access the computer system over a network via a computing device.


