DNA Sequence Annotation Confidence via Centroid Distance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current DNA sequence databases lack reliability due to sequencing errors and outdated annotations, making it difficult for lab technicians to accurately identify microorganisms, as incorrect annotations can lead to erroneous identifications.

Innovation Solution

A computer-implemented method assesses classification annotations by grouping DNA sequences based on established schemes, calculating a measure of distance between sequences, determining a centroid sequence, and assigning a confidence level to each sequence, which helps identify outliers and provide reliable annotations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If scientists add entries to public repositories with fair quality in terms of sequence content and annotation, then the database grows and covers more life-forms, but sequencing errors and incorrect annotations accumulate, reducing reliability

Engineering Contradiction:
Improvenumber of DNA sequencesVSAvoidannotation accuracy
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The system performs preliminary assessment of classification annotations by calculating confidence levels before sequences are used for identification. By pre-evaluating annotation quality using centroid-based confidence level calculation, the system prevents unreliable annotations from compromising identification accuracy, thus resolving the contradiction between database growth and reliability maintenance

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback mechanisms where confidence levels are calculated and displayed to users, enabling them to assess annotation reliability. This feedback loop allows the system to adapt to annotation quality variations and maintain reliable identification results even as the database expands with contributions from multiple sources

Inventive Principle:
Principle #23Feedback

2Measurement precision

If a reference database contains correct sequences with accurate annotations, then identification accuracy improves, but verifying and maintaining annotation accuracy requires extensive expert knowledge that routine lab technicians lack

Engineering Contradiction:
Improveidentification accuracyVSAvoiduser expertise requirement
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system introduces confidence levels as an intermediary metric that bridges the gap between complex annotation verification and user-friendly identification. Instead of requiring users to directly assess annotation accuracy using expert knowledge, the system calculates and displays confidence levels that automatically indicate reliability, making accurate identification accessible to routine lab technicians without specialized training

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs self-assessment of annotation quality by automatically calculating confidence levels based on sequence analysis and classification schemes. This self-service mechanism eliminates the need for manual expert verification of each annotation, maintaining high identification accuracy while reducing the expertise burden on users

Inventive Principle:
Principle #25Self-service

3Loss of information

If sequence database searches display all matching references, then comprehensive results are provided, but correct and incorrect matches become indistinguishable, requiring user expertise to interpret

Engineering Contradiction:
Improvecompleteness of resultsVSAvoidannotation reliability assessment
Core Design Contradiction:
Loss of informationVSDifficulty of detecting and measuring

Solution Approach 1:

The system applies local quality differentiation by associating distinct confidence levels with individual sequence annotations within the results list. Instead of treating all matches uniformly, the system locally tags each annotation with its assessed reliability, enabling users to quickly distinguish trustworthy from questionable matches without reducing result completeness or requiring expert interpretation skills

Inventive Principle:
Principle #3Local quality

Data Source

PatentEP2215578B1Method and computer system for assessing classification annotations assigned to DNA sequences
Publication Date: 2014.03.26 SMARTGENE
  • EP2215578B1 patent drawingFigure 1
  • EP2215578B1 patent drawingFigure 2
  • EP2215578B1 patent drawingFigure 3

AI summary

For assessing classification annotations assigned to DNA sequences stored in a reference database, the DNA sequences are grouped by species (S1) using established classification schemes. Subsequently, a measure of distance between pairs of DNA sequences is determined (S41) by aligning (S31) the respective sequences and determining the measure of distance (S41 ) based on a score of similarity between the aligned DNA sequences. Determined are one or more centroid sequences (S42) which have the shortest aggregate measure of distance to the other DNA sequences in the respective group (species). Assigned to the DNA sequences (S5) as a quantitative confidence level for their classification annotations is in each case the measure of distance between the respective DNA sequence and the centroid sequence. The assessment and rating of the classification annotations with these confidence levels make it possible to provide to a user a quantitative indication of the degree of representativeness of a DNA sequence for a particular species.