Computational Model for Polyadenylation Site Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods lack an effective way to determine the impact of genetic variations on polyadenylation site selection, which is crucial for understanding diseases and developing targeted therapies.

Innovation Solution

A computational model that uses machine learning to analyze genomic sequences and identify the effects of antisense oligonucleotides on polyadenylation site selection by processing genomic sequences, extracting feature vectors, and applying trained algorithms to predict preferences for candidate sites.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If machine learning models are trained to predict polyadenylation site selection, then prediction accuracy is improved, but computational complexity and data requirements increase

Engineering Contradiction:
Improveprediction accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The model processes genomic sequences by dividing them into manageable feature vectors representing specific polyadenylation site characteristics. The sequence is segmented into regions around candidate polyA sites, with features extracted from specific distance ranges (e.g., 50bp upstream, 100bp downstream), allowing the complex prediction problem to be broken into tractable computational units.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The model performs preliminary feature extraction and encoding of genomic sequences before the main prediction step. Candidate polyA sites are pre-identified and characterized with feature vectors containing sequence motifs, conservation scores, and structural features, preparing the data in advance for efficient model processing and reducing computational burden during prediction.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If comprehensive genomic sequences are analyzed to identify polyadenylation sites, then measurement precision is improved, but processing time increases

Engineering Contradiction:
Improvepolyadenylation site identification accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The model focuses computational resources on local regions around candidate polyadenylation sites rather than analyzing entire genomes uniformly. Feature extraction is performed within specific distance windows from each candidate site (e.g., 50bp upstream, 100bp downstream), concentrating analysis where it is most needed while reducing unnecessary processing of distant regions.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

Candidate polyadenylation sites are pre-identified using sequence motifs and conservation patterns before detailed feature extraction and prediction. This preliminary filtering step reduces the number of sites requiring comprehensive analysis, significantly reducing processing time while maintaining identification accuracy.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If multiple candidate polyadenylation sites are evaluated, then prediction completeness is improved, but computational resources required increase

Engineering Contradiction:
Improveprediction completenessVSAvoidcomputational resources
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The model applies different feature extraction and analysis strategies to different candidate sites based on their local characteristics. High-confidence candidates identified by strong sequence motifs receive more comprehensive analysis, while lower-confidence sites receive streamlined processing, optimizing resource allocation across multiple candidates.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The model performs full feature extraction and prediction for a subset of high-priority candidate sites while using simplified criteria for lower-priority sites. This partial action approach ensures complete evaluation of the most important candidates while reducing computational resources spent on less likely candidates, maintaining prediction completeness for critical sites.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20240412818A1Systems and methods for determining effects of therapies and genetic variation on polyadenylation site selection
Publication Date: 2024.12.12 DEEP GENOMICS INC
  • US20240412818A1 patent drawing
  • US20240412818A1 patent drawing
  • US20240412818A1 patent drawing

AI summary

The present disclosure provides systems and methods for determining effects of genetic variants on selection of polyadenylation sites (PAS) during polyadenylation processes. In an aspect, the present disclosure provides a polyadenylation code, a computational model that can predict alternative polyadenylation patterns from transcript sequences. A score can be calculated that describes or corresponds to the strength of a PAS, or the efficiency in which it is recognized by the 3′-end processing machinery. The polyadenylation model may be used, for example, to assess the effects of anti-sense oligonucleotides to alter transcript abundance. As another example, the polyadenylation model may be used to scan the 3′-UTR of a human genome to find potential PAS.