Automated Pathogen Identification via Sequence Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for pathogen detection are limited by the need for foreknowledge of potential pathogens, require significant manpower and technical resources, and are impractical for resource-poor locations or field use, especially due to the manual and computationally intensive nature of data analysis in Next Generation Sequencing (NGS) systems.

Innovation Solution

A computerized system with an electronic filtering subsystem and an electronic mapping subsystem that automatically compares genetic sequence reads to known sequences, calculates hit and distance scores, and clusters data to identify pathogens without prior knowledge of the infecting organism, enabling efficient and accurate pathogen identification in resource-constrained settings.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual data analysis is used in NGS systems, then pathogen identification accuracy is improved, but computational load and time requirements increase significantly

Engineering Contradiction:
Improvepathogen identification accuracyVSAvoiddata analysis time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs self-service through automated algorithms that independently complete data analysis without human intervention. The computational system automatically processes sequence reads, compares them against pathogen databases, calculates hit scores, and generates identifications, eliminating the need for manual analysis while maintaining accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical data analysis with automated electronic computing systems. The process substitutes human technicians with computer algorithms that perform sequence comparison, scoring, and identification tasks, thereby reducing time requirements while maintaining precision through standardized computational methods.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If highly skilled technicians perform data analysis, then measurement precision is improved, but device complexity and operational difficulty increase

Engineering Contradiction:
Improvepathogen identification accuracyVSAvoidoperational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system eliminates dependency on highly skilled technicians by implementing self-service automation. The computational algorithms automatically handle complex data processing tasks, and the interface allows operators with basic health skills to simply load samples and receive results, removing the complexity barrier while maintaining precision.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces an intermediary automated software layer between the sequencing instrument and the operator. This intermediary system handles the complex computational tasks, translating raw sequence data into pathogen identifications without requiring operators to understand or manually process the complexity, thereby simplifying operation while maintaining accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If comprehensive pathogen screening is performed, then adaptability is improved, but device complexity and resource requirements increase

Engineering Contradiction:
Improvepathogen detection coverageVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system achieves universality by using a single automated platform that can detect multiple pathogen types through standardized sequence comparison algorithms. The same hardware and software infrastructure handles diverse pathogen screening tasks, eliminating the need for separate specialized systems for each pathogen type and thereby reducing overall complexity while maintaining broad adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The automated self-service system handles comprehensive pathogen screening independently without requiring complex manual coordination. The algorithms automatically compare sequences against comprehensive databases, filter results, and generate identifications, managing the complexity of broad screening internally while presenting a simple interface to users.

Inventive Principle:
Principle #25Self-service

4Productivity

If automated data analysis is implemented, then productivity is improved, but measurement precision may be compromised

Engineering Contradiction:
Improveanalysis speedVSAvoidpathogen identification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system incorporates feedback mechanisms where computational algorithms continuously evaluate hit scores, compare results against threshold criteria, and adjust identifications based on predefined accuracy parameters. This feedback loop ensures that automated processing maintains precision by validating results against established scientific criteria while achieving high productivity through automation.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20220392576A1Method for detection and identification of known and emergent pathogens
Publication Date: 2022.12.08 THE GOVERNMENT OF THE UNITED STATES AS REPRESENTED BY THE SECRETARY OF THE AIR FORCE
  • US20220392576A1 patent drawing
  • US20220392576A1 patent drawing
  • US20220392576A1 patent drawing

AI summary

A method of detecting and identifying pathogens in a sample comprising a plurality of genetic sequences. A plurality of electronic sequence reads corresponding to the plurality of genetic sequences is received and sampled to form a sample set. The sample set is iteratively and electronically compared to a plurality of pathogen sequences to create a detection group, which populates a putative genome data structure. A distance score is measured between each electronic sequence read of the sampled set to each pathogen sequence of the putative genome data structure. A hit score is calculated by comparing the distance score to a threshold value. A plurality of clusters of the electronic sequence reads of the sample set is formed to maximize the cluster hit score and to minimize a difference in distance scores of the cluster. A respective taxonomic group assigned to electronic reads of the sample set after clustering is displayed.