Allelotyping Method for MPS Forensic Data Compression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing use of massively parallel sequencing (MPS) in forensic DNA analysis generates large, text-based nucleotide sequence data files that are difficult to transmit and store, and require human-readable formats for effective exploitation and preservation in forensic applications.
Innovation Solution
An allelotyping method that selects and compares text strings representing nucleotide sequences to determine abundance counts, identifies unique alleles by comparing these counts to an abundance threshold, and generates a smaller human-readable file containing only the relevant allele information, reducing file size by at least ten-thousand times.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If massively parallel sequencing is used to generate nucleotide sequence data, then forensic DNA analysis capability is improved, but file size increases making transmission and storage difficult
Solution Approach 1:
The patent extracts only the essential allele information from the complete nucleotide sequence data. By identifying and retaining only the unique text strings representing alleles and their abundance counts, the method removes redundant sequencing data while preserving the forensic analysis capability, reducing file size by at least ten-thousand times.
2Loss of information
If complete nucleotide sequence data is preserved, then data completeness is improved, but human readability deteriorates
Solution Approach 1:
The patent creates a simplified copy of the nucleotide sequence data that contains only the essential allele information in a human-readable format. This copy includes unique text strings representing alleles and their abundance counts, making the data accessible to human analysts while maintaining the complete forensic information needed for analysis.
3Loss of information
If all nucleotide sequences are retained, then sequence information completeness is improved, but data processing complexity increases
Solution Approach 1:
The patent extracts only the essential allele information from the complete nucleotide sequence data. By identifying and retaining only the unique text strings representing alleles and their abundance counts, the method removes redundant sequencing data while preserving the forensic analysis capability, reducing file size by at least ten-thousand times.
Solution Approach 2:
The patent segments the nucleotide sequence data by locus, organizing the information into discrete allelic units. Each unique text string represents a distinct allele at a specific locus, allowing for systematic processing and analysis of individual alleles rather than handling the entire sequence dataset as a single complex unit.
Data Source
AI summary
In one illustrative embodiment, an allelotyping method may include selecting a plurality of text strings that each represent a nucleotide sequence that was read by a massively parallel sequencing (MPS) instrument, where the nucleotide sequences represented by the selected plurality of text strings each correspond to a particular locus, comparing the selected plurality of text strings to one another to determine an abundance count for each unique text string included in the selected plurality of text strings, and determining one or more alleles for the particular locus by comparing the abundance count for each unique text string included in the selected plurality of text strings to an abundance threshold.


