Incremental Metadata Sharing for Privacy-Preserving Kinship Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for collaborative data analysis, particularly in genomic studies, often overlook the privacy preservation of sample identification and relatedness determination, leading to potential breaches of sensitive information.
Innovation Solution
A method and system that iteratively selects and synchronizes a subset of data units, generates metadata, and calculates relatedness measures across datasets while obfuscating identities, using an incremental update kinship identification framework to enhance privacy and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If all data units are shared to determine relatedness, then measurement precision is improved, but loss of information increases due to privacy breaches
Solution Approach 1:
The patent divides the complete set of data units into multiple subsets and processes them in sequential iterations. Each iteration handles a portion of the data, allowing relatedness determination to converge gradually without requiring immediate access to all data. This segmentation maintains measurement precision through cumulative results while minimizing privacy breach risk by limiting exposure of individual data units at any given time.
Solution Approach 2:
The patent performs preliminary actions by pre-processing and preparing data units before the main relatedness determination process. Metadata is generated and prepared in advance, and the iterative framework is established beforehand to systematically process data portions. This preliminary preparation enables efficient processing while maintaining privacy, as the framework is designed to handle data in controlled portions rather than requiring full data exposure.
2Measurement precision
If data is shared to determine relatedness, then measurement precision is improved, but loss of time increases due to computational complexity
Solution Approach 1:
The patent segments the computational task into multiple iterations, each processing a subset of data units. This divides the large computational problem into smaller, manageable portions that can be processed in parallel or sequentially. By breaking down the computation rather than processing all data at once, the patent reduces time loss while maintaining measurement precision through cumulative results from each iteration.
Solution Approach 2:
The patent performs preliminary computations to prepare metadata and establish the iterative framework before the main processing begins. This preliminary action includes pre-processing data units, defining the iteration structure, and preparing computational resources. By preparing everything in advance, the patent reduces the time required during actual relatedness determination, as the framework is already in place to efficiently process data portions.
3Measurement precision
If more data units are processed, then measurement precision is improved, but device complexity increases
Solution Approach 1:
The patent segments the data processing system into modular components that handle specific portions of data independently. Each iteration operates on a subset of data units using the same standardized process, creating a modular architecture. This segmentation reduces device complexity by allowing the system to handle large datasets through repeated application of simple, well-defined operations rather than requiring a single complex system to process all data at once.
Solution Approach 2:
The patent performs preliminary actions to establish a standardized framework and metadata structure before processing begins. This preliminary setup includes defining the iteration protocol, preparing data formats, and establishing processing parameters. By preparing the framework in advance, the patent simplifies the actual processing steps, reducing device complexity during execution as the system follows a pre-established pattern rather than requiring complex adaptive logic.
Data Source
AI summary
An example method can include generating first metadata representative of a first proper subset of data units in each of a plurality of samples stored in a first dataset, in which each sample of the plurality of samples stored in the first dataset includes a plurality of data units. The method can also include generating second metadata representative of a second proper subset of data units in each of a plurality of samples stored in a second dataset, in which each sample of the plurality of samples stored in the second dataset includes a plurality of data units. The method can also include sending the first and second metadata to a computer. The method can also include calculating a measure of relatedness between the samples stored in the first dataset and the samples stored in the second dataset based on the first metadata and the second metadata.


