Nucleic Acid Sequencing Data Analysis for Stutter Product Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current genetic analysis methods, particularly those using capillary electrophoresis, fail to accurately determine the sequence of nucleotides in short tandem repeats (STRs), leading to potential misidentification of alleles and quality control issues such as primer dimer formation and chimeric amplicons, which can result in unreliable genetic profiles.

Innovation Solution

A method involving the analysis of sequencing data to assign sample reads to genetic loci based on nucleotide sequences, identify regions-of-interest (ROIs) with repeat motifs, and differentiate potential alleles by their sequences, with additional steps to distinguish stutter products and noise, using a programmed computer to sort and count reads and determine genotypes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If capillary electrophoresis systems are used to analyze STR alleles, then the analysis process is simplified and faster, but the sequence information of alleles is lost leading to potential misidentification

Engineering Contradiction:
Improveanalysis speedVSAvoidallele sequence identification
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the analysis process into two distinct parts: (1) initial rapid screening using capillary electrophoresis to separate alleles by length, and (2) subsequent targeted sequencing of specific alleles to obtain precise sequence information. This segmentation allows the system to benefit from both the speed of CE and the accuracy of sequencing without requiring full sequencing of all samples.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Instead of sequencing all alleles completely (excessive action), the patent applies partial sequencing only to alleles that require further verification or show ambiguous results in the initial CE analysis. This partial action approach maintains high productivity while improving measurement precision only where necessary.

Inventive Principle:
Principle #16Partial or excessive action

2Loss of information

If all sequencing data is retained for analysis, then comprehensive genetic information is obtained, but noise and chimeric data reduce the reliability of results

Engineering Contradiction:
Improvegenetic information retentionVSAvoidgenetic profile accuracy
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The patent performs preliminary quality control assessments on sequencing data before full analysis, identifying and flagging potential noise, primer dimers, and chimeric sequences. By detecting these issues in advance, the system can either correct them or exclude them from final analysis, thereby maintaining comprehensive genetic information while ensuring high reliability of results.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback mechanisms where initial analysis results inform subsequent filtering decisions. Quality metrics from preliminary analysis feed back into the data processing pipeline, dynamically adjusting which data points are retained or discarded. This feedback loop ensures that comprehensive information is preserved while systematically removing unreliable data.

Inventive Principle:
Principle #23Feedback

3Quantity of substance

If stutter products are not distinguished from true alleles, then all detected sequences are retained, but false allele identification occurs reducing analysis accuracy

Engineering Contradiction:
Improvedetected allelesVSAvoidallele identification accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent applies different quality assessment criteria to different regions and characteristics of the sequencing data. Specifically, it analyzes local sequence patterns around repeat regions to identify stutter artifacts, while applying different evaluation standards to unique flanking regions. This localized quality assessment allows the system to distinguish true alleles from stutter products based on their distinct local characteristics.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system changes multiple parameters simultaneously to differentiate stutter products from true alleles, including but not limited to: sequence composition ratios, repeat unit uniformity, flanking region similarity, and read depth distributions. By monitoring multiple parameters rather than a single metric, the patent achieves high precision in allele identification while retaining comprehensive detection capability.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP3194627B1Methods and systems for analyzing nucleic acid sequencing data
Publication Date: 2023.08.16 ILLUMINA INC
  • EP3194627B1 patent drawingFigure 1~2
  • EP3194627B1 patent drawingFigure 3~5
  • EP3194627B1 patent drawingFigure 4

AI summary

Method includes receiving sequencing data including a plurality of sample reads that have corresponding sequences of nucleotides and assigning the sample reads to designated loci. The method also includes analyzing the assigned reads for each designated locus to identify corresponding regions-of-interest (ROIs) within the assigned reads. Each of the ROIs has one or more series of repeat motifs. The method also includes sorting the assigned reads based on the sequences of the ROIs such that the ROIs with different sequences are assigned as different potential alleles. The method also includes analyzing, for designated loci having multiple potential alleles, the sequences of the potential alleles to determine whether a first allele of the potential alleles is suspected stutter product of a second allele of the potential alleles.