Genotyping Data Merging with Conflict Resolution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current genotyping data processing methods face challenges in efficiently merging and analyzing data from multiple platforms, leading to inconsistencies and inefficiencies in genotype calling, particularly due to duplicate or conflicting data, which affects the accuracy and speed of DNA analysis for individuals.
Innovation Solution
A system and process for processing genotyping data that merges data from multiple platforms into a unified data structure, resolving duplicates and inconsistencies by using techniques such as consensus calls, majority voting, and incorporating contextual information like family inheritance, population-specific data, and linkage disequilibrium to improve genotype calling accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data from multiple genotyping platforms is merged, then the quantity and comprehensiveness of genetic data increases, but data inconsistencies and duplicates increase leading to reduced accuracy
Solution Approach 1:
The patent segments the data merging process into distinct stages: receiving data from multiple platforms, storing it in structured data sets, merging the data sets while resolving conflicts, and producing unified genotype calls. This segmentation allows systematic handling of inconsistencies through defined conflict resolution rules at each stage.
Solution Approach 2:
The patent implements feedback mechanisms where genotype calls are validated against known genetic principles and family inheritance patterns. When conflicts are detected during merging, the system uses feedback from familial relationships and population data to resolve inconsistencies and improve accuracy.
2Measurement precision
If comprehensive conflict resolution methods are applied, then genotype calling accuracy improves, but processing time and computational complexity increase
Solution Approach 1:
The patent performs preliminary actions by pre-processing data from each platform individually before merging, validating data quality and structure in advance. It also pre-establishes conflict resolution rules and hierarchical priorities for resolving conflicts, which speeds up the actual merging process.
Solution Approach 2:
The patent applies different conflict resolution strategies to different types of data conflicts based on their local characteristics. Simple conflicts are resolved quickly using predefined rules, while complex conflicts requiring familial analysis or population data are handled with more intensive computation only when necessary.
3Adaptability or versatility
If multiple data platforms are integrated, then the comprehensiveness of genetic analysis improves, but data inconsistency and errors increase
Solution Approach 1:
The patent creates a universal data structure and processing framework that can handle multiple genotyping platforms simultaneously. The unified data structure accommodates different platform formats while maintaining consistency, and the multi-functional conflict resolution system handles various types of conflicts using the same coherent approach.
Solution Approach 2:
The patent introduces an intermediary processing layer that translates and harmonizes data from different platforms before integration. This intermediary layer acts as a buffer that resolves format differences and potential inconsistencies, ensuring reliable data consistency across all platforms.
Data Source
AI summary
Processing genetic data includes receiving two or more genetic data sets for an individual from one or more genetic data sources, wherein the genetic data sets comprises data pertaining to the individual's deoxyribonucleic acid (DNA); merging the genetic data sets from the one or more genetic data sources to obtain a set of merged genetic data for the individual, including: identifying data in the genetic data sets that is conflicting, the identified data corresponding to a genetic marker associated with a variation that occurs at a region in the individual's genome; analyzing the identified data to resolve a discrepancy attributed to the identified conflicting data and automatically determine an appropriate value that corresponds to the genetic marker, the analysis and the determination being based at least in part on contextual information; and storing the appropriate value in the set of merged genetic data.


