Splice Site Prediction via Computational Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for predicting pre-mRNA splicing and alternative splicing are limited in accuracy and fail to explain the diversity, specificity, and fidelity of splice site choice, particularly in identifying introns and exon boundaries, which is crucial for understanding genetic diseases and protein regulation.

Innovation Solution

A computational method that identifies introns and exons by detecting characteristic splicing junctions in genomic DNA or pre-mRNA sequences, using databases and bioinformatics tools to predict novel splice sites and isoforms associated with diseases, and constructing splicing code tables to analyze and verify these sites.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If current computational methods are used to predict pre-mRNA splicing, then the process is simple and fast, but the accuracy and reliability of identifying introns and exon boundaries is insufficient

Engineering Contradiction:
Improveaccuracy of identifying introns and exon boundariesVSAvoidcomplexity of computational method
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the pre-mRNA sequence into distinct regions (exons, introns, splice sites) and applies specific computational analyses to each segment. The method divides the prediction task into multiple stages: initial sequence analysis, splice site identification, intron/exon boundary determination, and validation against known databases. This segmentation allows each component to be optimized independently, improving overall accuracy without requiring a single overly complex algorithm.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent employs preliminary action by first gathering and preprocessing reference data from known splice sites and intron/exon boundaries before performing the actual prediction. The method pre-establishes criteria for splice site recognition, pre-identifies conserved sequences, and pre-validates prediction algorithms against training datasets. This preliminary preparation creates a robust framework that enhances prediction accuracy while keeping the actual prediction process relatively simple.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If comprehensive database searches are performed to verify splice sites, then the reliability of predictions improves, but the time and computational resources required increase

Engineering Contradiction:
Improvereliability of splice site predictionVSAvoidtime for verification process
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies local quality by focusing verification efforts on specific critical regions rather than uniformly analyzing the entire sequence. The method concentrates computational resources on splice site regions, exon-intron boundaries, and conserved sequences where accuracy is most critical. Less critical regions receive minimal or no verification, optimizing the balance between reliability and time consumption by applying quality control where it matters most.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements partial action by performing verification only on predictions that meet certain confidence thresholds or exhibit specific characteristics. High-confidence predictions from initial analysis may skip extensive verification, while low-confidence or ambiguous cases receive full verification scrutiny. This selective verification approach maintains reliability for critical predictions while reducing overall time investment.

Inventive Principle:
Principle #16Partial or excessive action

3Manufacturing precision

If multiple verification methods are used to characterize introns and exons, then the manufacturing precision of gene annotation improves, but the ease of operation decreases

Engineering Contradiction:
Improveprecision of gene annotationVSAvoidease of using the method
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The patent merges multiple verification methods into a unified computational pipeline that automatically performs sequence analysis, splice site prediction, intron/exon identification, and database verification in a single integrated process. The method combines homology searching, conserved sequence analysis, statistical prediction, and experimental validation criteria into one cohesive workflow, maintaining high annotation precision while simplifying user operation through automation.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements self-service by designing the system to automatically perform all verification steps without requiring manual intervention at each stage. The computational method self-validates predictions by comparing against multiple databases, self-corrects errors through iterative refinement, and self-documented results with confidence scores. This automation maintains high precision while greatly improving ease of operation, as users simply need to input the sequence and receive comprehensive results.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10446261B1System and method for analyzing splicing codes of spliceosomal introns
Publication Date: 2019.10.15 BIOTAILOR
  • US10446261B1 patent drawing
  • US10446261B1 patent drawing
  • US10446261B1 patent drawing

AI summary

A system and method for analyzing splicing codes of spliceosomal introns is disclosed. One embodiment comprises methods of identifying introns and exons in genomic DNA or pre-mRNA sequences by locating characteristic markers in splicing junctions by computation and/or manually. Exon sequences predicted by computation can be verified and characterized by employing standard amplification methods, such as comparative genomic, RNA-seq, next-generation sequencing, RT-PCR. DNA/RNA/oligo, electrophoretic or protein chip technologies. If a given sample is verified, its polypeptide can be translated based on genetic codons. Its functions can be deduced based on its characteristics, computation predictions and related knowledge databases. These data can be used to compare databases which correlate the characterized intron or exon or gene to characterized diseases or genetic mutations. Isoforms can be detected and analyzed at mRNA and protein levels alone and with other isoforms predicted by computation, characterized by experiments and stored in existing databases.