SNP Data Block Segmentation for Genomic Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional genomic analysis methods are computationally intensive due to the need for serial comparison of millions of SNPs between DNA kits and database genomes, making efficient genealogical mapping and relatedness analysis challenging.
Innovation Solution
The method involves converting DNA chip data into a standardized data file format using a comprehensive SNPs template, allowing for efficient comparison by packaging data into blocks that match the processor's word length, and performing bitwise operations for half-matching analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional genomic analysis methods are used to compare millions of SNPs between DNA kits and database genomes, then comprehensive genealogical analysis can be performed, but the computational load and time required become excessively high
Solution Approach 1:
The patent segments the genome into smaller units called bins, where each bin contains a specific number of SNPs. Instead of comparing all millions of SNPs sequentially across the entire genome, the system processes these segmented bins independently and in parallel. This segmentation allows the computational task to be divided into manageable chunks that can be processed more efficiently, reducing the overall computation time while maintaining the ability to perform comprehensive genealogical analysis across the entire genome.
2Measurement precision
If conventional genomic analysis methods are used to compare millions of SNPs between DNA kits and database genomes, then comprehensive genealogical analysis can be performed, but the computational complexity increases significantly
Solution Approach 1:
The patent segments the genome into smaller units called bins, where each bin contains a specific number of SNPs. Instead of comparing all millions of SNPs sequentially across the entire genome, the system processes these segmented bins independently and in parallel. This segmentation allows the computational task to be divided into manageable chunks that can be processed more efficiently, reducing the overall computation time while maintaining the ability to perform comprehensive genealogical analysis across the entire genome.
Solution Approach 2:
The patent creates a compressed binary representation (copy) of the SNP data within each bin, where matching SNPs are represented by binary values. This copied binary format simplifies the comparison operation, allowing for efficient bitwise operations instead of complex sequential comparisons. The binary copy maintains the essential information needed for genealogical analysis while dramatically reducing the computational complexity of the matching process.
Data Source
AI summary
The present invention relates to an improved system and method for analyze data from submitted DNA kit and/or genome data form database records so as to compare sequences for determining the level of SNP homology between the two tested sequences. The DNA calling data is compared in a stepwise block by block manner, where the blocks for the compared sequences have data blocks in bit-word lengths of the processor performing the sequence comparison analysis. The blocks are compared in block sets having a minimum cM length, where the block comparisons are initiated at the last block of the minimum length block sets and proceed in a retrograde manner.


