Allelotyping Method for MPS Data Compression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing use of massively parallel sequencing (MPS) in forensic DNA analysis generates large, unwieldy nucleotide sequence data files that are difficult to transmit and store, and require human-readable formats for effective exploitation and preservation.
Innovation Solution
An allelotyping method that selects and compares text strings representing nucleotide sequences from MPS instruments, determines unique alleles by abundance count, and generates a smaller, human-readable file format, such as the Sequence Evidence Format (SEF), which includes forensic metadata and quality scores, to compress data while retaining essential information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If massively parallel sequencing is used to generate nucleotide sequence data, then the information content and analytical capability are improved, but the file size becomes excessively large making transmission and storage difficult
Solution Approach 1:
The patent extracts only the essential information from the raw nucleotide sequence data by identifying and retaining unique alleles at each locus. Instead of preserving all sequence reads, the method extracts the distinct allele sequences and their abundance counts, eliminating redundant data while maintaining the forensic analytical value needed for human identification and genetic mapping.
Solution Approach 2:
The patent segments the large nucleotide sequence dataset into discrete loci and identifies unique alleles at each locus. By dividing the data into manageable units (unique alleles rather than individual sequence reads), the method creates a compressed representation that is much smaller in size but retains all necessary genetic information for forensic analysis.
2Reliability
If all nucleotide sequence data is preserved for forensic analysis, then the completeness of forensic information is improved, but the complexity of data management and human readability deteriorates
Solution Approach 1:
The patent extracts the essential forensic information by identifying unique alleles and their abundance counts at each locus. This extraction process removes the overwhelming volume of raw sequence data while preserving the critical genetic markers needed for human identification, making the data manageable and interpretable without sacrificing forensic reliability.
Solution Approach 2:
The patent changes the data representation parameters by converting raw nucleotide sequences into a standardized format that includes locus identifiers, unique allele sequences, and abundance counts. This parameter transformation maintains the scientific integrity of the forensic data while making it far more manageable and potentially more readable for analysis and reporting.
3Loss of information
If raw nucleotide sequence data is transmitted and stored, then the data completeness for future analysis is improved, but the transmission time and storage requirements increase significantly
Solution Approach 1:
The patent extracts only the unique allele information and abundance counts from the complete raw sequencing data. This extracted subset contains all the essential genetic variation information needed for forensic analysis, allowing for rapid transmission and storage while maintaining data completeness for future analytical purposes.
Solution Approach 2:
The patent segments the comprehensive nucleotide sequence data into discrete unique alleles at each locus, creating a compressed data structure that preserves all necessary genetic information in a much smaller format, thereby reducing transmission time and storage requirements without sacrificing analytical completeness.
Data Source
AI summary
In one illustrative embodiment, an allelotyping method may include selecting a plurality of text strings that each represent a nucleotide sequence that was read by a massively parallel sequencing (MPS) instrument, where the nucleotide sequences represented by the selected plurality of text strings each correspond to a particular locus, comparing the selected plurality of text strings to one another to determine an abundance count for each unique text string included in the selected plurality of text strings, and determining one or more alleles for the particular locus by comparing the abundance count for each unique text string included in the selected plurality of text strings to an abundance threshold.


