Genomic Identity Scores via Local Differential String Sets
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for determining genetic distances between biological samples are inefficient due to large data files and inability to accurately reflect meaningful heterogeneity, often underestimating or overestimating genetic differences, and struggling with differential weighing of sequence variations.
Innovation Solution
The method involves incremental synchronous sequence alignments using a sequence analysis engine to generate local differential string sets, which are used to calculate difference scores and determine genetic distances, allowing for the certification of homogeneity within a group of biological samples by synchronizing genetic sequence strings and identifying outlier samples.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional pairwise genetic distance methods are used to compare DNA sequences, then comprehensive genetic comparison can be performed, but the analysis becomes slow and difficult to manage due to hundreds of gigabytes of sequence information data
Solution Approach 1:
The patent segments the genome into multiple windows or regions, allowing parallel processing of different genomic segments. Instead of loading entire genomes into memory for comparison, the system divides the genome into manageable chunks that can be processed independently and concurrently, significantly reducing memory requirements and enabling faster processing while maintaining comprehensive genetic comparison capabilities
Solution Approach 2:
The patent introduces a new computational dimension by using coordinate-based indexing and spatial data structures to organize genomic data. This allows the system to access and compare specific genomic regions efficiently without loading entire genomes, transforming the problem from a memory-intensive operation to a computation-intensive operation that can be parallelized across multiple processors
2Device complexity
If all differences in sequence alignments are weighed equally to determine genetic distances, then the calculation process is simple, but the method underestimates or overestimates genetic differences among samples
Solution Approach 1:
The patent applies local quality weighting by assigning different importance weights to different genomic regions based on their biological significance. Functional regions such as coding sequences, regulatory elements, and conserved regions receive higher weights, while non-functional regions receive lower weights. This allows the system to accurately reflect the biological reality that not all genetic differences are equally important, thereby improving genetic distance measurement accuracy
Solution Approach 2:
The patent implements differential weighting schemes that dynamically adjust the importance of different sequence variations based on multiple factors including functional annotation, conservation status, and variant frequency. By changing the weighting parameters according to these factors, the system transforms the simple equal-weighting approach into a sophisticated multi-factor weighting system that accurately captures meaningful genetic heterogeneity
3Loss of information
If whole genome sequencing is performed to obtain comprehensive genetic data, then complete genetic information is available, but the very large data files are difficult to transport and process
Solution Approach 1:
The patent extracts only the essential genetic information needed for comparison by identifying and isolating variant positions and their coordinates. Instead of storing and transporting complete genome sequences, the system extracts minimal sufficient data including variant positions, reference alleles, alternative alleles, and quality scores. This extraction approach maintains genetic information completeness for analysis while reducing data size by orders of magnitude, making data transport and processing tractable
Data Source
AI summary
Methods for analyzing omics data and using the omics data to determine genetic distances and/or difference scores among a plurality of biological samples to so further determine the homogeneity of a group having a plurality of biological samples and/or exclude an individual biological sample from a group of biological samples as an outlier are presented. In preferred methods, a plurality of local differential string sets among the plurality of sequence strings is generated using a plurality of local alignments. The local different string is an indicative of genetic difference between one sequence string and one of the rests of the sequence strings among the plurality of sequence strings. From the plurality of local differential string sets, a plurality of difference scores among the plurality of sequence strings can be determined.


