Neural Network Polynucleotide Property Evaluation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for understanding the functions of DNA and RNA sequences, particularly non-coding regions, are time-consuming and labor-intensive, and existing computational models struggle with accurately predicting properties like promoter identification, viral sequence classification, and RNA stability.

Innovation Solution

A computer-implemented method using a neural network model, specifically a convolutional neural network with a self-attention mechanism, to determine properties of polynucleotide sequences, capable of capturing local and global dependencies, and trained in supervised, unsupervised, or semi-supervised environments to identify promoters, viral genes, and assess stability under environmental parameters.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If experimental approaches are used to understand DNA and RNA sequence functions through genetic mutation, then accurate functional information can be obtained, but the process becomes time-consuming and labor-intensive

Engineering Contradiction:
Improvefunctional information accuracyVSAvoidevaluation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent creates computational copies of experimental evaluation processes through machine learning models. Instead of physically performing genetic mutations and experiments on DNA/RNA sequences, the system uses trained models that replicate the functional evaluation capability, providing accurate predictions without the time and labor costs of actual experiments

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical experimental process (wet lab procedures, genetic manipulation, physical measurement) with a computational system. Machine learning models process sequence data algorithmically to predict functional properties, substituting physical experimentation with computational analysis while maintaining prediction accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of time

If computational models are used to predict polynucleotide properties, then evaluation time is reduced, but prediction accuracy and reliability deteriorate

Engineering Contradiction:
Improveevaluation timeVSAvoidprediction accuracy
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The patent transforms the input parameters from raw sequence data alone to include multiple processed features such as k-mer frequencies, positional information, and structural predictions. This parameter transformation enables the model to capture complex sequence properties while maintaining computational efficiency and improving prediction accuracy

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent combines multiple computational approaches and data sources into a composite prediction system. By integrating different feature representations and model architectures (sequence-based features, structure-based features, and their interactions), the system achieves higher prediction accuracy than any single method alone while maintaining fast evaluation

Inventive Principle:
Principle #40Composite materials

3Ease of operation

If simple computational models are used for polynucleotide analysis, then ease of operation is improved, but the ability to capture complex local and global dependencies is lost

Engineering Contradiction:
Improvemodel simplicityVSAvoiddependency capture capability
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent segments the polynucleotide sequence into multiple k-mer units (overlapping subsequences of fixed length). This segmentation allows the model to process complex sequence information in manageable chunks, capturing local dependencies within each k-mer while the aggregation of multiple k-mer representations captures global sequence patterns

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the one-dimensional sequence data into multi-dimensional feature space by extracting k-mer frequencies, positional encodings, and structural features. This dimensional transformation enables the model to capture complex dependencies that cannot be represented in the original sequence space, while the underlying architecture remains computationally efficient

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20240087685A1Systems and methods for evaluation of structure and property of polynucleotides
Publication Date: 2024.03.14 TEXAS A&M UNIVERSITY
  • US20240087685A1 patent drawing
  • US20240087685A1 patent drawing
  • US20240087685A1 patent drawing

AI summary

Provided here are neural network-based methods and systems that process nucleic acid sequences to identify, characterize, and interpret specific properties, like characterization of promoters of gene sequences, identification of viral genomes, and stability of the nucleic acids. The neural network models can be trained in supervised, unsupervised, and semi-supervised environments.