Allelotyping Method for MPS Forensic Data Compression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing use of massively parallel sequencing (MPS) in forensic DNA analysis generates large, text-based nucleotide sequence data files that are difficult to transmit and store, and require human-readable formats for effective exploitation and preservation in forensic applications.

Innovation Solution

An allelotyping method that selects and compares text strings representing nucleotide sequences to determine abundance counts, identifies unique alleles by comparing these counts to an abundance threshold, and generates a smaller human-readable file containing only the relevant allele information, reducing file size by at least ten-thousand times.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If massively parallel sequencing is used to generate nucleotide sequence data, then forensic DNA analysis capability is improved, but file size increases making transmission and storage difficult

Engineering Contradiction:
Improveforensic DNA analysis capabilityVSAvoidfile size
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential allele information from the complete nucleotide sequence data. By identifying and retaining only the unique text strings representing alleles and their abundance counts, the method removes redundant sequencing data while preserving the forensic analysis capability, reducing file size by at least ten-thousand times.

Inventive Principle:
Principle #2Taking out (Extraction)

2Loss of information

If complete nucleotide sequence data is preserved, then data completeness is improved, but human readability deteriorates

Engineering Contradiction:
Improvedata completenessVSAvoidhuman readability
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent creates a simplified copy of the nucleotide sequence data that contains only the essential allele information in a human-readable format. This copy includes unique text strings representing alleles and their abundance counts, making the data accessible to human analysts while maintaining the complete forensic information needed for analysis.

Inventive Principle:
Principle #26Copying

3Loss of information

If all nucleotide sequences are retained, then sequence information completeness is improved, but data processing complexity increases

Engineering Contradiction:
Improvesequence information completenessVSAvoiddata processing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent extracts only the essential allele information from the complete nucleotide sequence data. By identifying and retaining only the unique text strings representing alleles and their abundance counts, the method removes redundant sequencing data while preserving the forensic analysis capability, reducing file size by at least ten-thousand times.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the nucleotide sequence data by locus, organizing the information into discrete allelic units. Each unique text string represents a distinct allele at a specific locus, allowing for systematic processing and analysis of individual alleles rather than handling the entire sequence dataset as a single complex unit.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20230197196A1Allelotyping Methods for Massively Parallel Sequencing
Publication Date: 2023.06.22 BATTELLE MEMORIAL INST
  • US20230197196A1 patent drawing
  • US20230197196A1 patent drawing
  • US20230197196A1 patent drawing

AI summary

In one illustrative embodiment, an allelotyping method may include selecting a plurality of text strings that each represent a nucleotide sequence that was read by a massively parallel sequencing (MPS) instrument, where the nucleotide sequences represented by the selected plurality of text strings each correspond to a particular locus, comparing the selected plurality of text strings to one another to determine an abundance count for each unique text string included in the selected plurality of text strings, and determining one or more alleles for the particular locus by comparing the abundance count for each unique text string included in the selected plurality of text strings to an abundance threshold.