Splice Site Prediction via Computational Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for predicting pre-mRNA splicing and alternative splicing are limited in accuracy and fail to explain the diversity, specificity, and fidelity of splice site choice, particularly in identifying introns and exon boundaries, which is crucial for understanding genetic diseases and protein regulation.
Innovation Solution
A computational method that identifies introns and exons by detecting characteristic splicing junctions in genomic DNA or pre-mRNA sequences, using databases and bioinformatics tools to predict novel splice sites and isoforms associated with diseases, and constructing splicing code tables to analyze and verify these sites.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If current computational methods are used to predict pre-mRNA splicing, then the process is simple and fast, but the accuracy and reliability of identifying introns and exon boundaries is insufficient
Solution Approach 1:
The patent segments the pre-mRNA sequence into distinct regions (exons, introns, splice sites) and applies specific computational analyses to each segment. The method divides the prediction task into multiple stages: initial sequence analysis, splice site identification, intron/exon boundary determination, and validation against known databases. This segmentation allows each component to be optimized independently, improving overall accuracy without requiring a single overly complex algorithm.
Solution Approach 2:
The patent employs preliminary action by first gathering and preprocessing reference data from known splice sites and intron/exon boundaries before performing the actual prediction. The method pre-establishes criteria for splice site recognition, pre-identifies conserved sequences, and pre-validates prediction algorithms against training datasets. This preliminary preparation creates a robust framework that enhances prediction accuracy while keeping the actual prediction process relatively simple.
2Reliability
If comprehensive database searches are performed to verify splice sites, then the reliability of predictions improves, but the time and computational resources required increase
Solution Approach 1:
The patent applies local quality by focusing verification efforts on specific critical regions rather than uniformly analyzing the entire sequence. The method concentrates computational resources on splice site regions, exon-intron boundaries, and conserved sequences where accuracy is most critical. Less critical regions receive minimal or no verification, optimizing the balance between reliability and time consumption by applying quality control where it matters most.
Solution Approach 2:
The patent implements partial action by performing verification only on predictions that meet certain confidence thresholds or exhibit specific characteristics. High-confidence predictions from initial analysis may skip extensive verification, while low-confidence or ambiguous cases receive full verification scrutiny. This selective verification approach maintains reliability for critical predictions while reducing overall time investment.
3Manufacturing precision
If multiple verification methods are used to characterize introns and exons, then the manufacturing precision of gene annotation improves, but the ease of operation decreases
Solution Approach 1:
The patent merges multiple verification methods into a unified computational pipeline that automatically performs sequence analysis, splice site prediction, intron/exon identification, and database verification in a single integrated process. The method combines homology searching, conserved sequence analysis, statistical prediction, and experimental validation criteria into one cohesive workflow, maintaining high annotation precision while simplifying user operation through automation.
Solution Approach 2:
The patent implements self-service by designing the system to automatically perform all verification steps without requiring manual intervention at each stage. The computational method self-validates predictions by comparing against multiple databases, self-corrects errors through iterative refinement, and self-documented results with confidence scores. This automation maintains high precision while greatly improving ease of operation, as users simply need to input the sequence and receive comprehensive results.
Data Source
AI summary
A system and method for analyzing splicing codes of spliceosomal introns is disclosed. One embodiment comprises methods of identifying introns and exons in genomic DNA or pre-mRNA sequences by locating characteristic markers in splicing junctions by computation and/or manually. Exon sequences predicted by computation can be verified and characterized by employing standard amplification methods, such as comparative genomic, RNA-seq, next-generation sequencing, RT-PCR. DNA/RNA/oligo, electrophoretic or protein chip technologies. If a given sample is verified, its polypeptide can be translated based on genetic codons. Its functions can be deduced based on its characteristics, computation predictions and related knowledge databases. These data can be used to compare databases which correlate the characterized intron or exon or gene to characterized diseases or genetic mutations. Isoforms can be detected and analyzed at mRNA and protein levels alone and with other isoforms predicted by computation, characterized by experiments and stored in existing databases.


