Computational Phylogenetic Analysis Using Entropy Distribution Functions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for analyzing viruses and tracking gene sequence changes are labor-intensive and computationally demanding, leading to delays in identifying viruses, which can have significant implications for public health.

Innovation Solution

A computational phylogenetic analysis system that automates the process of identifying mutations by generating binary sequences from genetic sequences, partitioning them into binary strings, calculating entropy values, and using entropy distribution functions to compare and categorize genomes, leveraging machine learning concepts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual research and string comparison algorithms are used to analyze genetic sequences, then analysis can be performed with simple tools, but the process is labor-intensive and computationally demanding, leading to delays in identifying viruses

Engineering Contradiction:
Improvespeed of virus identificationVSAvoidtime required for manual analysis
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical analysis methods with an automated computational system that uses entropy distribution functions and machine learning algorithms to analyze genetic sequences, eliminating the need for labor-intensive manual research and string comparison while significantly improving analysis speed

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-service by automatically generating binary sequences from genetic data, calculating entropy values, generating entropy distribution functions, and identifying mutations without requiring manual intervention, thereby reducing both labor intensity and analysis time

Inventive Principle:
Principle #25Self-service

2Measurement precision

If comprehensive genetic sequence analysis is performed to accurately identify mutations, then mutation detection precision is improved, but computational resources and complexity increase

Engineering Contradiction:
Improveaccuracy of mutation identificationVSAvoidcomplexity of analysis system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transforms genetic sequence data into binary sequences and calculates entropy distribution functions as intermediate parameters, creating a simplified representation that maintains mutation detection accuracy while reducing the complexity of direct genetic sequence comparison

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The entropy distribution function serves as an intermediary between the raw genetic sequence and the mutation identification process, enabling accurate mutation detection through a computationally manageable intermediate representation rather than direct complex sequence analysis

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10808280B2Computational phylogenetic analysis
Publication Date: 2020.10.20 COLOSSIO INC
  • US10808280B2 patent drawing
  • US10808280B2 patent drawing
  • US10808280B2 patent drawing

AI summary

Systems and a method for computationally analyzing genetic base pairs are provided. In one or more aspects, a system includes a memory and a processor coupled to the memory. The processor is configured to receive a number of genetic sequences from a genetic sequencer device. The processor can generate, for each genetic sequence, a binary sequence. Each binary sequence is partitioned into a set of binary strings. Each binary string includes multiple binary base pairs. A set of entropy values are determined, each entropy value is associated with a binary string, and an entropy distribution function (EDF) is generated based on the set of entropy values.