Common Key Generation for Cross-Company Data Sharing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data sharing technologies face challenges in obtaining statistical information from datasets of different companies, requiring pre-prepared metadata and raising concerns about personal information protection, especially when companies use different data management systems.
Innovation Solution
An information processing system that generates a 'common key' by training a model on feature information from multiple companies' databases, allowing for the classification and calculation of statistical information without exposing identifiable personal data, enabling seamless data sharing while maintaining privacy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If statistical information is obtained from datasets of different companies, then data sharing capability is improved, but complexity of data integration increases due to different platforms and metadata requirements
Solution Approach 1:
The patent introduces an intermediary system that automatically generates and manages metadata schemas, acting as a mediator between different company data platforms. This intermediary component translates and harmonizes diverse data formats without requiring direct complex integration between companies, thereby improving data sharing capability while managing integration complexity through a dedicated mediation layer.
Solution Approach 2:
The patent applies preliminary action by automatically generating metadata schemas and data structure definitions before actual data sharing occurs. This pre-processing step establishes standardized frameworks in advance, eliminating the need for complex manual metadata preparation and integration planning when companies share their datasets, thus resolving the complexity issue while enabling versatile data sharing.
2Adaptability or versatility
If metadata is prepared and shared between companies, then data compatibility is improved, but time and resources required for metadata preparation increase
Solution Approach 1:
The patent implements self-service by enabling the system to automatically generate, validate, and maintain metadata schemas without requiring manual intervention from company representatives. The system autonomously performs metadata preparation tasks, including schema generation from sample data, compatibility validation, and automatic updates, thereby achieving data compatibility while eliminating time-consuming manual metadata preparation.
Solution Approach 2:
The patent replaces the mechanical manual process of metadata preparation with an automated computational system. Instead of humans manually creating and managing metadata, the system uses algorithms and machine learning models to automatically generate metadata schemas, validate data compatibility, and maintain metadata repositories, thus achieving compatibility while dramatically reducing preparation time and resource requirements.
3Measurement precision
If identifiable user data is shared between companies, then statistical accuracy is improved, but personal information protection is compromised
Solution Approach 1:
The patent applies the extraction principle by systematically removing identifiable personal information from datasets before sharing. The system extracts and eliminates direct identifiers (names, IDs) and indirect identifiers (quasi-identifiers that could lead to re-identification) while retaining the statistical and analytical value of the remaining data. This extraction process enables statistical accuracy to be maintained through preserved data patterns while preventing personal information exposure by removing identification capabilities.
Solution Approach 2:
The patent implements parameter changes by transforming data parameters from identifiable to non-identifiable states. This includes techniques such as data aggregation (grouping individual records into statistical categories), generalization (replacing specific values with broader categories), and perturbation (adding controlled noise to prevent exact matching). These parameter transformations maintain the statistical integrity and accuracy of the data while fundamentally changing its identification parameters to protect personal information.
Data Source
AI summary
An information processing device according to the present application includes a generation unit and a providing unit. The generation unit uses a model that is trained to learn a relationship between a criterion for classifying users of a first company and a criterion for classifying users of a second company to generate a criterion (common key) for classifying the users of the second company into a first category, from the criterion for classifying the users of the first company into the first category. The providing unit provides a criterion generated by the generation unit.


