Anonymized Data Combination for Disease Risk Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing disease development risk prediction systems fail to predict disease risk using existing information without special testing and do not protect personal information when combining data from different sources.
Innovation Solution
A disease development risk prediction system that generates combination data by combining medical, dispensing, and nursing care insurance data using a combination key that includes anonymized insured person numbers, birth dates, and gender, and uses this data to create a prediction model that predicts disease risk while protecting personal information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If data from multiple sources is combined to improve prediction accuracy, then prediction precision is improved, but personal information protection deteriorates
Solution Approach 1:
The patent extracts and removes personally identifiable information from the data before combining multiple data sources. By taking out the harmful element (personal information) while retaining the useful data for prediction, the system achieves both high prediction accuracy and personal information protection simultaneously
Solution Approach 2:
The patent introduces an intermediary processing step that anonymizes data before combination. This intermediary layer allows multiple data sources to be combined for improved prediction while preventing direct exposure of personal information, effectively mediating between the conflicting requirements
2Loss of energy
If existing information is used without special testing to reduce costs, then loss of energy is reduced, but measurement precision deteriorates
Solution Approach 1:
The patent enables the system to use existing information that is already collected for other purposes (self-service data) without requiring additional special testing. This approach reduces the loss of energy (testing costs) while the sophisticated combination and analysis methods ensure adequate prediction accuracy is maintained
3Productivity
If multiple types of receipt data are combined to improve prediction capability, then productivity is improved, but device complexity increases
Solution Approach 1:
The patent segments the complex task of combining multiple data types into distinct, manageable steps: data extraction, anonymization, combination, and analysis. This segmentation improves prediction capability by systematically processing each data type while reducing the perceived complexity through structured organization of the combination process
Data Source
AI summary
A disease development risk prediction system 10 includes: a data generation means 11 which generates combination data by combining at least two different types of receipt data using a combination key, wherein the receipt data includes an insured person number for an insured person which was converted using a predetermined method, a birth date or birth year and month which are both age-identifiable items, and gender, and the combination key combines the converted insured person number, age-identifiable items, and gender; and a model generation means 12 which uses the generated combination data to generate a prediction model predicting a risk of the insured person of developing a predetermined disease.


