Methods, systems, and apparatuses for calculating SNP-induced mRNA allosteric scores
By calculating the allosteric probability of splicing-related elements in precursor mRNA sequence pairs, the influence of SNPs on RNA splicing was quantified, solving the problem of insufficient SNP influence analysis in existing technologies. This enabled the precise localization and quantification of RNA secondary structures and provided a pathway for the discovery of disease-related SNPs.
Patent Information
- Application Number
- CN202211392287.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-08
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2042-11-08
AI Technical Summary
Existing technologies are insufficient to effectively quantify the impact of single nucleotide polymorphisms (SNPs) on the RNA splicing process, especially regarding their influence on disease at the secondary structure level.
By extracting the RNA secondary structure (RSS) from the splicing-related elements of precursor mRNA sequence pairs, the allosteric probability of the splicing-related elements is calculated, the risk of splicing allosteric changes caused by SNPs is quantified, and the ratio of allosteric nucleotides is calculated by using RNA structure prediction algorithms and genome annotation information, combined with the location and type of splicing-related elements, to output the SNP-induced mRNA allosteric risk score.
This study achieved precise localization and mismatch-free matching of SNPs across the entire genome, quantified the impact of SNPs on RNA splicing, provided a new approach for the discovery of disease-related SNPs, and improved the ability to analyze RNA secondary structures.
Smart Images

Figure CN117174168B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of bioinformatics, and more particularly, to a method, system, diagnostic device and computer readable storage medium for calculating SNP-induced mRNA allosteric score. BACKGROUND
[0002] RNA splicing is an essential biological process in post-transcriptional regulation and gene expression of eukaryotes, which matures the precursor mRNA (pre-mRNA) and diversifies the proteome of multicellular organisms. Meanwhile, a large number of transcriptional variations in human cells are derived from genetic interference in RNA splicing. Studies have shown that some SNPs may cause splicing variations, and abnormal splicing significantly affects the occurrence and development of diseases by regulating expression. On the other hand, more and more evidence shows that RNA secondary structure (RSS) is widely involved in a large number of biological functional processes, including the occurrence and development of diseases. The present application analyzes the influence of SNPs on the splicing process from the level of RNA secondary structure, and proposes a method for calculating SNP-induced mRNA allosteric score. SUMMARY
[0003] The present application aims to provide a method for calculating SNP-induced mRNA allosteric score, comprising:
[0004] Step one: obtaining a SNP and its corresponding pair of pre-mRNA sequences, the pair of pre-mRNA sequences being a pair of pre-mRNA sequences with and without the SNP;
[0005] Step two: extracting M splicing-related elements from the pair of pre-mRNA sequences, and extracting RSS in the M splicing-related elements, to obtain RSS of the splicing-related elements, wherein M is a natural number of 2-10;
[0006] Step three: calculating the allosteric probability of the RSS of the splicing-related elements, the allosteric probability being the ratio of the total number of allosteric nucleotides in the RSS of the splicing-related elements to the total number of nucleotides in the RSS of the splicing-related elements;
[0007] Step four: outputting SNP-induced mRNA allosteric risk score, the allosteric risk score being the allosteric probability of the RSS of the splicing-related elements.
[0008] Optionally, the pair of pre-mRNA sequences is one or more pairs of pre-mRNA sequences with and without the SNP; RNA splicing can produce one or several alternative splicings, when there is only one alternative splicing, the pair of pre-mRNA sequences is one pair of pre-mRNA sequences with and without the SNP, and when there are multiple alternative splicings, the pair of pre-mRNA sequences is multiple pairs of pre-mRNA sequences with and without the SNP;
[0009] Further, the method for obtaining SNP and its corresponding precursor mRNA sequence pair is:
[0010] Obtaining SNP flanking sequence, and cutting N bases upstream and downstream of the SNP as seed sequence, wherein N is a natural number of 15-50;
[0011] Downloading transcript information as reference sequence in database;
[0012] Obtaining reference sequence position corresponding to SNP seed sequence based on short sequence alignment tool;
[0013] According to SNP matched positive and negative strand and starting base position information corresponding to upstream and downstream fragments, screening SNP located fragments in reference sequence, obtaining SNP and its corresponding precursor mRNA sequence pair;
[0014] Preferably, screening SNP located fragments PM in reference sequence according to the following principles:
[0015]
[0016] Ini up ,Ini down respectively represent base starting position of SNP upstream and downstream sequence matched to reference sequence, Seed up ,Seed down respectively correspond to SNP upstream and downstream flanking sequence, Ref forward ,Ref reverse respectively indicate seed sequence located in positive and negative strand of reference sequence.
[0017] Further, the second step can be replaced by the following steps: extracting RSS of the precursor mRNA sequence and its position in the precursor mRNA sequence, extracting M kinds of splicing related elements of the precursor mRNA sequence, based on the position information of RSS, retaining RSS located in the splicing related element, obtaining RSS of the splicing related element, wherein M is a natural number of 2-10.
[0018] Further, the M kinds of splicing related elements include one or more of the following splicing related elements: 5' splice site, 3' splice site, branch point, polypyrimidine tract mRNA splicing region, splicing regulatory element; preferably, the splicing regulatory element includes exon splicing enhancer (ESE), intron splicing enhancer (ISE), exon splice silencer (ESS), intron splice silencer;
[0019] Optionally, the following methods are used to extract M splicing-related elements of the precursor mRNA sequence: extracting the 5' splicing site and 3' splicing site of the precursor mRNA sequence based on genome annotation information; preferably, one or more of the following methods are used to extract the 5' splicing site and 3' splicing site of the precursor mRNA sequence: Deep Splicer, Splice Finder, Splice2Deep, SpliceRover, DeepSS, SpliceAI; one or more of the following methods are used to extract the branch points and polypyrimidine tract mRNA splicing regions of the precursor mRNA sequence: SVM-Bpfinder, BPP, Branchpointer, LaBranchoR, RNABPS; one or more of the following methods are used to extract the splicing regulatory elements of the precursor mRNA sequence: HSF, Protein-Specific Prediction of RNA-Binding Sites Based on InformationEntropy, RBPMmap, GraphProt, RNA-binding protein targets, iONMF, iDeep, CircRNA-RBPWeb Server.
[0020] Further, the RSS comprises a stem (S), a hairpin loop (H), an inner loop (I), an outer loop (E), a raised loop (B), and a multi-branched loop (M);
[0021] Optionally, the extracting of RSS from the M splicing-related elements is to predict the secondary structure of the splicing-related element RNA using an RNA structure prediction algorithm, and then extracting the RSS from the secondary structure using an RNA motif prediction algorithm;
[0022] Preferably, the RNA structure prediction algorithm includes one or more of the following algorithms: RNA structure, RNAfold, Mfold, Sfold, MaxExpect; preferably, the RNA motif prediction algorithm includes one or more of the following algorithms: bpRNA, DotAligner, Cmfinder, RNAz, QRNA.
[0023] The RNA structure, RNAfold, Mfold, Sfold, and MaxExpect include the above-mentioned known algorithm software and upgraded versions thereof or other commercial software that solves the same problem (predicting the secondary structure of RNA); the bpRNA, DotAligner, Cmfinder, RNAz, and QRNA include the above-mentioned known algorithm software and upgraded versions thereof or other commercial software that solves the same problem (extracting RSS from the secondary structure).
[0024] Optionally, the step three is replaced by the following steps: calculating the allosteric probability of RSS of each nucleic acid site in each splicing-related element, the allosteric probability being the ratio of the total number of allosteric nucleotides of RSS of each nucleic acid site in each splicing-related element to the total number of RSS nucleotides of each nucleic acid site in each splicing-related element;
[0025] Step four: outputting the SNP-induced mRNA allosteric risk score, the allosteric risk score being the set of allosteric probabilities of RSS of each splicing-related element.
[0026] Further, the method for calculating the SNP-induced mRNA allosteric score further comprises calculating the distance between the SNP and the base where allosterism occurs, and when the distance is lower than a threshold, adding a distance risk prompt item to the SNP-induced mRNA allosteric risk score;
[0027] Optionally, the distance between the SNP and the base where allosterism occurs is the distance between the absolute position of the SNP on the genome and the absolute position of the base where allosterism occurs on the genome; and optionally, the threshold is 350 bp.
[0028] The purpose of the present application is to provide a device for calculating a SNP-induced mRNA allosteric score, the device comprising: a memory and a processor;
[0029] The memory is used to store program instructions;
[0030] The processor is used to invoke the program instructions, and when the program instructions are executed, the method for calculating the SNP-induced mRNA allosteric score described above is executed.
[0031] The purpose of the present application is to provide a system for calculating a SNP-induced mRNA allosteric score, comprising:
[0032] An acquisition unit is configured to acquire a SNP and a corresponding precursor mRNA sequence pair;
[0033] An extraction unit is configured to extract M splicing-related elements of the precursor mRNA sequence pair and extract RSS in the M splicing-related elements, to obtain RSS of the splicing-related elements, wherein M is a natural number of 2-10;
[0034] A calculation unit is configured to calculate an allosteric probability of RSS of the splicing-related elements, the allosteric probability being the ratio of the total number of allosteric nucleotides in RSS of the splicing-related elements to the total number of nucleotides in RSS of the splicing-related elements;
[0035] An output unit is configured to output a SNP-induced mRNA allosteric risk score, the allosteric risk score being the allosteric probability of RSS of the splicing-related elements.
[0036] A computer readable storage medium, having stored thereon a computer program, the computer program being executed by a processor to implement the method for calculating SNP-induced mRNA allosteric score.
[0037] Advantages of the present application:
[0038] 1. A method for calculating SNP-induced mRNA allosteric score is provided, the RSS of splicing-related elements is extracted from the pair of pre-mRNA sequences containing or not containing SNP, and the allosteric probability of the RSS of splicing-related elements caused by SNP is calculated, which quantifies the influence of SNP on RNA splicing, and provides a new way for discovering disease-related SNPs;
[0039] 2. The present application proposes an accurate positioning algorithm of SNP in the coding, non-coding genes, and functional elements such as exons and introns in the whole genome, so as to realize the accurate and error-free alignment of SNP on the reference sequence.
[0040] 3. In the research process, it is proved that the importance of genomic position, and when the distance between SNP and the base that occurs allosteric is less than the threshold value, it has a higher risk of inducing mRNA allosteric. BRIEF DESCRIPTION OF DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0042] Figure 1 is a schematic flow chart of the method for calculating SNP-induced mRNA allosteric score provided by the embodiments of the present application;
[0043] Figure 2 is a schematic diagram of the system for calculating SNP-induced mRNA allosteric score provided by the embodiments of the present application;
[0044] Figure 3 is a schematic diagram of the device for calculating SNP-induced mRNA allosteric score provided by the embodiments of the present application;
[0045] Figure 4Figure is a myopia-related SNP data collection diagram provided by specific embodiments of the present application; (A) Venn diagram of data sources; (B) SNP gene location; (C) SNP location of gene regulatory elements; (D) gene locus distribution in eye tissue; (E) gene locus distribution in cornea; (F) gene locus distribution in iris; (G) gene locus distribution in retina; (H) gene locus distribution in sclera and choroid.
[0046] Figure 5 Figure is a myopia-related SNP causing a large amount of structural heterogeneity in the splicing-related elements of pre-mRNA provided by specific embodiments of the present application; (A) An example of SNP-induced pre-mRNA RNA secondary structure variation. The red box indicates the position where local structural heterogeneity exists. The blue box indicates the position where global structural heterogeneity exists. The black triangle represents the positions of SNP, start site and end site, respectively; (B) Distribution of P values in RNAsnp. RNAsnp website suggests p value > 0.2 as less obvious structural change; (C) RNAsmc score distribution (score range 0-10). RNAsmc score reflects the overall structural heterogeneity. The lower the score, the higher the global structural heterogeneity. Scores less than 9 are considered significant structural heterogeneity; (D, E) Tables and Venn diagrams show the number of RSSs with structural changes and pre-mRNA pairs that occur in specific regions;
[0047] Figure 6 Figure is a structural context feature of splicing-related elements involved in explaining their structural change probability provided by specific embodiments of the present application; A, B: RNA structural atlas of splicing-related elements. Blue boxes represent exons. Light blue boxes represent introns. Dotted lines represent 5'SS, BP site and 3'SS, respectively. Statistical test of Mann-Whitney U non-parametric test; C RNA structural analysis of splicing-related elements. Statistical test of Mann-Whitney U non-parametric test; D Scatter plot describing the linear relationship between structural change probability and structural motif. r: Pearson correlation coefficient. E Correlation matrix of structural change probability and structural motif using Pearson correlation test. *: p value ≤ 0.05, **: p value ≤ 0.01, ***: p value ≤ 0.001, ****: p value ≤ 0.0001.
[0048] Figure 7 Figure is a structural atlas of splicing-related elements provided by specific embodiments of the present application; A-E Overall structural change feature atlas of RSSs; The vertical axis in the figure represents the structural change probability (%), and the upper limit is 2%, and the horizontal axis represents each type of structural change; F-H Ranking score of each type of structural change in the three structural change modes, and the statistical test is performed by Wilcoxon rank sum test. *: p value < 0.05, **: p value < 0.01, ***: p value < 0.001 and ****: p value < 0.0001.
[0049] Figure 8 Figure 13 is a schematic diagram of factors that can be involved in regulating SNP allosteric effect by distance and genomic position; A a schematic diagram of factors that can be involved in regulating allosteric probability (AP); B a histogram showing the allosteric probability under different distance intervals. The horizontal axis represents the absolute distance (nt) between nucleotides and SNPs. The vertical axis represents the percentage of allosteric probability. The horizontal dotted line represents 5% allosteric probability. The vertical pink dotted line divides the distance effect into distal and proximal dominance; C shows the odds ratio of proximal and distal allosteric features. Black boxes represent no significant odds ratio. Blue boxes represent odds ratio (P < 0.05) less than 1. Yellow boxes represent odds ratio (P < 0.01) greater than 1; D Violin plot showing the allosteric probability under different genomic fragments (exons vs introns). Paired T test was performed to analyze the difference of AP in each allosteric motif. **: p value < 0.01, ***: p value < 0.001, ****: p value < 0.0001
[0050] Figure 9 Figure 14 is a molecular interaction map showing the SNPs with high structural heterogeneity risk widely affect the molecular interaction between RNAs and splicing-related proteins; A a Venn diagram showing the intersection of RNAsnp, RNAsmc and 5 splicing-related elements screening; B top 10 high-risk SNPs; C shows the docking information scatter plot of RBP-RNA molecular docking state change, green points represent HDOCK score shift, purple and pink points represent the mean and median shift of RNA binding residues, respectively;
[0051] Figure 10 Figure 15 is a schematic diagram of RNA structure features at LIM2 first intron 5' s that can regulate splicing; A, B experimental design schematic diagram to verify the role of secondary structure of 5' SS in splicing process, in order to form a firm hairpin structure with the entire 5' SS, a short sequence (blue box and circle) is immediately inserted upstream of the original 5' SS sequence (pink box and circle); C, D determine the splicing events by minigene splicing analysis in HEK293T, use the same pair of primers to identify unspliced and spliced products by RT-PCR, and show different lengths of bands, measure the average gray value of uncut, spliced products and background by image J / FIJI to calculate splicing efficiency, E shows the radar chart of the docking state change amplitude between LIM2 RNA and U1 snRNP;
[0052] Figure 11 Figure 16 is a schematic diagram of allosteric probability of RSS of each nucleic acid site of specific splicing-related elements of myopia A. DETAILED DESCRIPTION
[0053] In order to better understand the technical scheme of the present application, the technical scheme in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present application.
[0054] In some of the processes described in this specification, including the description of the drawings, the operations in the processes are not necessarily performed in the order indicated. The sequences of operations are illustrative only. Depending on the implementation, certain operations can be performed in a different order, or omitted, or in parallel. Also, the illustrations in the drawings are for simplicity and clarity and are not necessarily drawn to scale. The description and drawings are to be regarded as illustrative in nature and its contents not as restrictive.
[0055] The technical scheme in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.
[0056] Figure 1 is a schematic flow chart of a method for calculating SNP-induced mRNA allosteric score provided by an embodiment of the present application. Specifically, the method comprises the following steps:
[0057] S101: obtaining a SNP and a corresponding pair of precursor mRNA sequences;
[0058] In an embodiment, the pair or pairs of precursor mRNA sequences with and without SNP are one pair or multiple pairs of precursor mRNA sequences with and without SNP. RNA splicing produces one or several alternative splicing. When there is only one alternative splicing, the pair of precursor mRNA sequences with and without SNP is one pair of precursor mRNA sequences with and without SNP. When there are multiple alternative splicing, the pair of precursor mRNA sequences with and without SNP is multiple pairs of precursor mRNA sequences with and without SNP.
[0059] In an embodiment, the method for obtaining a SNP and a corresponding pair of precursor mRNA sequences is:
[0060] Obtaining a SNP flanking sequence, and taking each N bases upstream and downstream of the SNP as a seed sequence, wherein N is a natural number of 15-50;
[0061] Downloading transcript information in a database as a reference sequence;
[0062] obtaining reference sequence positions corresponding to the SNP seed sequences respectively based on a short sequence alignment tool;
[0063] According to the start base position information of the positive and negative strands matched by the SNP and the upstream and downstream fragments, the fragment in which the SNP is located on the reference sequence is screened, and a SNP and a corresponding precursor mRNA sequence pair are obtained.
[0064] In one specific embodiment, the method for obtaining a SNP and a corresponding precursor mRNA sequence pair is as follows: first, we obtain the SNP flanking sequences in the dbSNP database, and to ensure the accuracy of the alignment results, 30 bases upstream and downstream of the SNP are respectively taken as seed sequences; second, the transcript information corresponding to the gene or functional element is downloaded from databases such as ENSMBL and GENCODE as reference sequences. Finally, based on the short sequence alignment tool-Bowtie 2, strict alignment parameters such as “--n-ceil C,3--np0--end-to-end-a--score-min C,0” are set, and reference sequence positions corresponding to the seeds upstream and downstream of the SNP are obtained respectively. According to the start base position information of the positive and negative strands matched by the SNP and the upstream and downstream fragments, the fragment in which the SNP is accurately located on the reference sequence is screened according to the following principles:
[0065]
[0066] Ini up ,Ini down represent the base start positions of the upstream and downstream sequences of the SNP matched to the reference sequence, Seed up ,Seed down correspond to the upstream and downstream flanking sequences of the SNP, Ref forward ,Ref reverse indicate the positive and negative strand cases of the seed sequence located on the reference sequence. PM is the SNP accurately aligned to the gene or other reference sequence, thereby realizing accurate and error-free alignment of the SNP on the reference sequence.
[0067] S102: Extracting M splice-related elements of the precursor mRNA sequence pair and extracting RSS in the M splice-related elements to obtain RSS of the splice-related elements, wherein M is a natural number from 2 to 10;
[0068] In one embodiment, another method of S102 is as follows: extracting RSS of the precursor mRNA sequence and its position in the precursor mRNA sequence, extracting M splice-related elements of the precursor mRNA sequence, and based on the position information of the RSS, retaining the RSS located in the splice-related element to obtain RSS of the splice-related element, wherein M is a natural number from 2 to 10.
[0069] In one embodiment, the M splicing-related elements include one or more of the following splicing-related elements: 5' splice site, 3' splice site, branch point, polypyrimidine tract mRNA splice region, splicing regulatory element; preferably, the splicing regulatory element includes exon splice enhancer (ESE), intron splice enhancer (ISE), exon splice silencer (ESS), intron splice silencer;
[0070] In one specific embodiment, the M splicing-related elements of the pre-mRNA sequence are extracted by the following methods: 5' splice site and 3' splice site of the pre-mRNA sequence are extracted based on genomic annotation information; preferably, one or more of the following methods are used to extract the 5' splice site and 3' splice site of the pre-mRNA sequence: Deep Splicer, SpliceFinder, Splice2Deep, SpliceRover, DeepSS, SpliceAI; one or more of the following methods are used to extract the branch point and polypyrimidine tract mRNA splice region of the pre-mRNA sequence: SVM-Bpfinder, BPP, Branchpointer, LaBranchoR, RNABPS; one or more of the following methods are used to extract the splicing regulatory element of the pre-mRNA sequence: HSF, Protein-Specific Prediction of RNA-Binding Sites Based on Information Entropy, RBPMmap, GraphProt, RNA-binding protein targets, iONMF, iDeep, CircRNA-RBP Web Server.
[0071] In one embodiment, the RSS includes stem (S), hairpin loop (H), internal loop (I), external loop (E), bulge loop (B), multi-branch loop (M); optionally, the RSS in the M splicing-related elements is predicted by the RNAfold algorithm to predict the secondary structure of the splicing-related element RNA, and then the bpRNA is used to extract the RSS in the secondary structure.
[0072] S103: Calculate the allosteric probability of the RSS of the splicing-related element, which is the ratio of the total number of allosteric nucleotides in the RSS of the splicing-related element to the total number of nucleotides in the RSS of the splicing-related element;
[0073] In one specific embodiment, the allosteric probability AP of the RSS of the splicing-related element is calculated, and for the pre-mRNA sequence region, the calculation formula of the overall allosteric probability AP is:
[0074]
[0075] wherein, AP R represents the probability of RSS change in the sequence region, n1 represents the total number of RSS change nucleotides in the sequence region, and n2 represents the total number of nucleotides in the sequence region.
[0076] In one embodiment, the other method of S103 is to calculate the probability of RSS change of each nucleic acid site in each splicing-related element, which is the ratio of the total number of RSS change nucleotides of each nucleic acid site in each splicing-related element to the total number of RSS nucleotides of each nucleic acid site in each splicing-related element.
[0077] In one specific embodiment, the probability of RSS change of each nucleic acid site in each splicing-related element is calculated, and the calculation formula of the probability of RSS change of each nucleic acid site in the precursor mRNA sequence region is:
[0078]
[0079] wherein, AP N represents the probability of RSS change of each nucleic acid site in the sequence region, n1 represents the total number of RSS change nucleotides in the sequence region, and n2 represents the total number of nucleotides in the sequence region.
[0080] In one specific embodiment, the probability of RSS change of each nucleic acid site in each splicing-related element is calculated, and the results are shown in Figure 11 .
[0081] The RSS change nucleotide refers to a nucleotide that changes the RSS of the precursor mRNA sequence with SNP relative to the precursor mRNA sequence without SNP.
[0082] S104: Output SNP-induced mRNA change risk score, which is the probability of RSS change of the splicing-related element.
[0083] In one embodiment, the SNP-induced mRNA change risk score is output, which is the set of the probability of RSS change of each splicing-related element; and the specific method is as follows:
[0084] SNP risk ∈{AP 5′ss ,AP 3′ss ,AP BP ,AP PPT ,AP SRE}
[0085] wherein, SNP riskAP is the SNP-induced mRNA misfolding risk score 5′ss AP is the misfolding probability of 5' splice site 3′ss AP is the misfolding probability of 3' splice site BP AP is the misfolding probability of branch point PPT AP is the misfolding probability of polypyrimidine tract mRNA splicing region SRE AP is the misfolding probability of splicing regulatory element.
[0086] In one embodiment, the method of calculating SNP-induced mRNA misfolding score further comprises calculating the distance between the SNP and the misfolded base, when the distance is below a threshold, adding a distance risk prompt item in the SNP-induced mRNA misfolding risk score; optionally, the distance between the SNP and the misfolded base is the distance between the absolute position of the SNP on the genome and the absolute position of the misfolded base on the genome; optionally, the threshold is 350bp.
[0087] In one embodiment, the method further comprises: evaluating the structural heterogeneity of the pair of pre-mRNA sequences with and without the SNP, obtaining a structural heterogeneity score; outputting the SNP-induced mRNA misfolding risk score, which is the collection of the misfolding probability of the RSS of the splicing-related element and the structural heterogeneity score; optionally, the structural heterogeneity score is obtained by evaluating the degree of local structural heterogeneity impact of the SNP on the pre-mRNA sequence, preferably by RNAsnp; optionally, the structural heterogeneity score is obtained by evaluating the degree of global structural heterogeneity impact of the SNP on the pre-mRNA sequence, preferably by RNAsmc.
[0088] In one embodiment, the inventors investigated SNP-induced mRNA misfolding risk scores using myopia as an example. Human myopia-associated SNPs were collected from the genotype and phenotype database (dbGaP), NHGRI-EBI published genome-wide association studies (GWAS) and manual literature mining. The final set contained a total of 1145 SNPs. Ensembl variant effect predictor (VEP) was used to annotate the SNPs in the genome. Allelic information was retrieved from the single nucleotide polymorphism database (dbSNP), and for quality control, SNPs were strictly filtered by the following steps: removing SNPs without reference, removing SNPs with multiple loci variations, removing SNPs located in intergenic regions. A total of 1541 pairs of wild-type (WT) and mutant (MT) pre-mRNAs (pre-mRNA sequence pairs) involving 806 myopia-associated SNPs were used as FASTA sequences from Ensemble (GRCh38) (see Figure 4 ). Splice-related elements consist of 5' splice site (5'SS), 3' splice site (3'SS), branch point (BP), polypyrimidine tract (Py-tract), and splice regulatory element (SRE) ( Figure 5 ). Among them, the genomic information of 5'SS and 3'SS comes from the genome annotation of Ensemble (GRCh38). BP and Py-tract are identified by Branchpointer, SRE is determined by functional screening from UniProt, and is detected by RBPmap. We mapped the nucleotide sites of RNA misfolding to each splice-related element and calculated the involved pre-mRNA sequence pairs in the misfolding region. Among the 1541 pairs of pre-mRNAs, 121 involved 5'SS misfolding, 102 involved 3'SS misfolding, 78 involved BP misfolding, 118 involved Py-tract misfolding, and 979 involved SRE misfolding ( Figure 5 E). Our research results show that the RNA structure of splice-related elements is widely disturbed. RNA secondary structure is predicted by RNAfold in the viennaRNA package (version 2.4.18). RNA secondary structure motifs are mined by bpRNA. These RNA substructures can be divided into two categories: (1) paired state (Pair): stem (S); (2) unpaired state (Unpair): hairpin loop (H), internal loop (I), external loop (E), bulge loop (B), and multi-branch loop (M). We calculated the misfolding probability (AP) of each splice-related element, including 5'ss, 3'ss, branch point site, py-tract, and 59 SRE binding regions ( Figure 6). For the splice sites, 5'SS-1 (5'ss region upstream-1 position nucleotide), 5'SS+1 and 3'SS-1 have better anti-allosteric properties than the surrounding nucleotides. BP is the most sensitive to RNA allosteric effect in the 5 regions, with AP up to 4.2%. For Py-tract, AP fluctuates around (3%~3.5%), without obvious peaks and valleys. At the same time, we also plotted the RNA structure profile of these splicing-related motifs on the WT type transcript precursor Figure 6 ). It is not difficult to find that there is a certain relationship between allosteric probability and RSS Figure 6 ). Pearson correlation test Figure 6 ) shows that MotifS is highly negatively correlated with allosteric effect (r = -0.812; r = Pearson correlation coefficient), MotifH (r = 0.790) and M (r = 0.753) are highly positively correlated, MotifI (r = 0.669) and E (r = 0.447) are moderately positively correlated, and MotifB is lowly positively correlated. The results of correlation analysis show that motif S is conducive to the stability of RNA structure and helps the splicing-related region resist SNP-induced allosteric. The significant correlation between AP and structure feature profile shows that nucleotides in different RSS states may have different allosteric probabilities, so we further study the allosteric feature profile of each splicing-related element. In order to observe the allosteric characteristics, we further divide the 20 specific allosteric types into 3 allosteric modes. Pair>Unpair mode includes S>B, S>H, S>I, S>M; Unpair>Pair mode includes B>S, H>S, I>S, M>S; Unpair>Unpair mode contains all the remaining motif allosteric types. Regardless of the region allosteric, the AP of Unpair>Pair is the largest, and the AP of Pair>Unpair is the lowest, which means that the nucleotides in Unpair state are more susceptible to the impact of secondary structure caused by SNPs and transitions. To the paired state. In addition, the maximum AP in 3'SS, BP and Py-tract is I>S except 5SS. Specifically, the first three APs of 5'SS are B>S (4.6%), M>S (3.7%) and I>S (3.5%) Figure 7 ). It is worth noting that the +1 site downstream of 5'SS (5'SS+1) is particularly special in 5'SS sequence, and the AP of B>S is more than 2 times higher than that of other 5'SS sites. The first three APs of 3'SS are I>S (4.6%), B>S (3.5%) and M>S (3%) Figure 7 ). We studied the relationship between AP and the distance between allosteric nucleotide and SNP. The statistical results show that the RNA allosteric effect gradually weakens with the increase of the distance between nucleotide and SNP. Take 350 nt as the boundary Figure 8B), we divided the SNP allostery into proximal allostery and distal allostery. Through risk assessment, we found that the two kinds of allostery have different regulatory effects on RNA structure. Compared with distal allostery, proximal allostery has stronger allostery risk for allostery derived from S (S > H except) and M > S. However, the distal allostery has higher allostery risk than proximal allostery for allostery derived from H (H > B except), B (B > S except) and S > H. Second, given that SRE binding regions distributed on introns or exons exhibit large differences in AP, we performed paired t-tests to determine whether the differences were statistically significant. The results showed that there were significant differences in the allostery probability of certain motif-derived allostery between introns and exons, including “S > B”, “S > H”, “S > I”, “S > M”, “B > I”, “B > S”, “H > I”, “H > S”, “I > B”, “I > H”, “I > S”, “M > H”, “M > S”. We combined the two structural heterogeneity scores (RNAsnp, RNAsmc) and five splicing-related elements to obtain SNPs with high structural heterogeneity risk of splicing-related motifs. To further explore the influence of SNPs on the docking ability of splicing-related RBPs and pre-mRNA, we simulated the docking process of pre-mRNA and top10 risk score SNPs (see Figure 9 ) and splicing-related proteins by HDOCK Server. The HDOCK evaluation results found that the top10 risk score SNPs widely interfered with the docking score and docking site between pre-mRNA and splicing-related proteins. To visually present the binary interaction between structural interference in splicing-related elements and splicing efficiency, we used RNA structure interference in the 5'SS of the first exon-intron-exon region of the myopia-related gene LIM2 for functional verification. Take ENST00000596399.2) as an example. The LIM2 gene has been reported to be associated with axial elongation and cataract. For this purpose, 1 control group and 3 experimental groups were designed. To avoid the disorder of base pairing between U1 snRNA and 5'ss in the splicing process, the 5'ss sequence content was preserved in each experimental group. To form a complete base-paired firm hairpin structure in the entire 5's, a short sequence was inserted immediately upstream of the 5's sequence. The -1 and -2 nucleotides upstream of the 5'ss are considered sufficient to regulate splicing in Arabidopsis thaliana, and it has been assessed to be in motif B (unpaired state) in the native sequence of LIM2. Then, we created two mutations that base-paired with the 5'ss sequence in the inserted sequence to disrupt the base-pairing state of the -1 and -2 nucleotides upstream of the 5'ss( Figure 10 ). We evaluated the splicing events on these design constructs by small gene splicing assays in HEK293T cells. Figure 10First, we verified that the native sequence construct was fully spliced in HEK293T cells ( Figure 10 , lane ENST00000596399.2: (1) When the entire sequence of 5 is perfectly base-paired with the upstream inserted sequence, splicing is significantly inhibited ( Figure 10 , lane 2 of ENST00000596399.2), by introducing the mutations “AA” or “GC” to force the -1 and -2 positions upstream of the 5'ss to be unpaired, we found that the splicing event was partially rescued ( Figure 10 , Lane 3 and Lane 4 of ENST00000596399; (2) Compared with the stem design group, the splicing efficiency increased from 8% to 30.83% or 50.08%, respectively ( Figure 10 ).
[0089] Figure 2 A system for calculating SNP-induced mRNA allosteric scores provided in an embodiment of the present invention includes:
[0090] An acquisition unit, used to acquire a pair of SNPs and their corresponding precursor mRNA sequences;
[0091] an extraction unit, configured to extract M splicing-related elements of the pre-mRNA sequence pair, and extract RSSs from the M splicing-related elements to obtain the RSSs of the splicing-related elements, wherein M is a natural number ranging from 2 to 10;
[0092] a calculation unit, configured to calculate an allosteric probability of the RSS of the splicing-related element, wherein the allosteric probability is a ratio of the total number of allosteric nucleotides in the RSS of the splicing-related element to the total number of nucleotides in the RSS of the splicing-related element;
[0093] The output unit is used to output the SNP-induced mRNA allosteric risk score, where the allosteric risk score is the allosteric probability of the RSS of the splicing-related element.
[0094] Figure 3 Schematic diagram of a device for calculating SNP-induced mRNA allosteric score provided by an embodiment of the present invention, wherein the device includes: a memory and a processor;
[0095] The memory is used to store program instructions;
[0096] The processor is used to call program instructions, and when the program instructions are executed, it is used to execute the above-mentioned method for calculating the SNP-induced mRNA allosteric score
[0097] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements the above-mentioned method for calculating the SNP-induced mRNA allosteric score.
[0098] The verification result of the verification embodiment shows that assigning inherent weights to indications can moderately improve the performance of the method compared with the default setting.
[0099] Those skilled in the art can clearly understand the specific working processes of the system, device and unit described above for the convenience and brevity of description, which can refer to the corresponding processes in the foregoing method embodiments, and will not be described here.
[0100] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the described device embodiments are merely schematic, and the division of the units is merely a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or other forms.
[0101] The units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, that is, can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0102] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically, or two or more units can be integrated into one unit. The integrated unit can be realized in the form of hardware, or in the form of a software functional unit.
[0103] Those skilled in the art can understand that all or part of the steps in the above-mentioned embodiments of various methods can be completed by a program instructing relevant hardware, and the program can be stored in a computer readable storage medium, which can include read only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.
[0104] Those skilled in the art can understand that all or part of the steps in the above-mentioned embodiment methods can be completed by a program instructing relevant hardware, and the program can be stored in a computer readable storage medium, which can include read only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.
[0105] The computer device provided by the present application is described in detail above, and for the general technical personnel in the field, the specific implementation manners and application ranges will be changed according to the idea of the embodiment of the present application, and the above description should not be understood as the limitation of the present application.
Claims
1. A method for calculating a SNP-induced mRNA allosteric score, comprising: Step 1: Obtain SNP and its corresponding pre-mRNA sequence pair; Step 2: extracting M splicing-related elements of the precursor mRNA sequence pair, and extracting RSSs from the M splicing-related elements to obtain the RSSs of the splicing-related elements, wherein M is a natural number of 2-10; the RSS is an RNA secondary structure; the RSS includes a stem, a hairpin loop, an inner loop, an outer loop, a protruding loop, and a multi-branched loop; the M splicing-related elements include one or more of the following splicing-related elements: a 5' splice site, a 3' splice site, a branch point, a polypyrimidine tract mRNA splicing region, and a splicing regulatory element; Step 3: Calculate the allosteric probability of the RSS of the splicing-related element, where the allosteric probability is the ratio of the total number of allosteric nucleotides in the RSS of the splicing-related element to the total number of nucleotides in the RSS of the splicing-related element; the allosteric nucleotide refers to a nucleotide whose RSS of the precursor mRNA sequence with the SNP is changed relative to the precursor mRNA sequence without the SNP; Step 4: Output the SNP-induced mRNA allosteric risk score, which is the allosteric probability of the RSS of the splicing-related element.
2. The method for calculating the SNP-induced mRNA allosteric score according to claim 1, wherein The method for obtaining the SNP and its corresponding precursor mRNA sequence pair is: Obtain the SNP flanking sequence and extract N bases upstream and downstream of the SNP as the seed sequence, where N is a natural number between 15 and 50; Download transcript information from the database as a reference sequence; Based on the short sequence alignment tool, the reference sequence positions corresponding to the SNP seed sequences were obtained respectively; According to the starting base position information corresponding to the positive and negative strands and upstream and downstream fragments of the SNP match, the fragments where the SNP is located in the reference sequence are screened to obtain the SNP and its corresponding precursor mRNA sequence pair.
3. The method for calculating the SNP-induced mRNA allosteric score according to claim 2, wherein: Screen the PM fragments where SNPs are located in the reference sequence according to the following principles: , Respectively represent the base starting position of the upstream and downstream sequences of the SNP matched to the reference sequence, Corresponding to the upstream and downstream flanking sequences of the SNP, , Respectively indicate whether the seed sequence is located on the positive and negative strands of the reference sequence.
4. The method for calculating the SNP-induced mRNA allosteric score according to claim 1, wherein Another method for step 2 is: extracting the RSS of the precursor mRNA sequence and its position in the precursor mRNA sequence, extracting M splicing-related elements of the precursor mRNA sequence, retaining the RSS located in the splicing-related elements based on the position information of the RSS, and obtaining the RSS of the splicing-related elements, wherein M is a natural number of 2-10, and the RSS is an RNA secondary structure; the RSS includes a stem, a hairpin loop, an inner loop, an outer loop, a protruding loop, and a multi-branched loop; the M splicing-related elements include one or more of the following splicing-related elements: a 5' splice site, a 3' splice site, a branch point, a polypyrimidine tract mRNA splicing region, and a splicing regulatory element.
5. The method for calculating the SNP-induced mRNA allosteric score according to claim 4, wherein The following method is used to extract M splicing-related elements of the pre-mRNA sequence: based on genome annotation information, the 5' splicing site and the 3' splicing site of the pre-mRNA sequence are extracted.
6. The method for calculating the SNP-induced mRNA allosteric score according to claim 5, wherein: Use one or more of the following methods to extract the 5' splice site and 3' splice site of the pre-mRNA sequence: Deep Splicer, Splice Finder, Splice2Deep, SpliceRover, DeepSS, SpliceAI; use one or more of the following methods to extract the branch point and polypyrimidine tract mRNA splicing region of the pre-mRNA sequence: SVM-Bpfinder, BPP, Branchpointer, LaBranchoR, RNABPS; use one or more of the following methods to extract the splicing regulatory elements of the pre-mRNA sequence: HSF, Protein-Specific Prediction of RNA-Binding Sites Based on Information Entropy, RBPMmap, GraphProt, RNA-binding protein targets, iONMF, iDeep, CircRNA-RBP Web Server.
7. The method for calculating the SNP-induced mRNA allosteric score according to claim 1, wherein: The extracting of RSS from the M splicing-related elements is to use an RNA structure prediction algorithm to predict the secondary structure of the splicing-related element RNA, and then use an RNA motif prediction algorithm to extract the RSS from the secondary structure.
8. The method for calculating the SNP-induced mRNA allosteric score according to claim 7, wherein: The RNA structure prediction algorithm includes one or more of the following algorithms: RNA structure, RNAfold, Mfold, Sfold, and MaxExpect.
9. The method for calculating the SNP-induced mRNA allosteric score according to claim 7, wherein: The RNA motif prediction algorithm includes one or more of the following algorithms: bpRNA, DotAligner, Cmfinder, RNAz, and QRNA.
10. The method for calculating the SNP-induced mRNA allosteric score according to claim 1, wherein: The third step is: calculating the allosteric probability of the RSS of each nucleic acid site in each splicing-related element, wherein the allosteric probability is the ratio of the total number of RSS allosteric nucleotides at each nucleic acid site in each splicing-related element to the total number of RSS nucleotides at each nucleic acid site in each splicing-related element, wherein the allosteric nucleotide refers to a nucleotide whose RSS of the precursor mRNA sequence with the SNP is changed relative to the precursor mRNA sequence without the SNP; Step 4: Output the SNP-induced mRNA allosteric risk score, which is the set of allosteric probabilities of the RSS of each splicing-related element.
11. The method for calculating the SNP-induced mRNA allosteric score according to claim 1, 4 or 10, wherein: The method for calculating the SNP-induced mRNA allosteric score further includes calculating the distance between the SNP and the base where the allosteric change occurs. When the distance is lower than a threshold, a distance risk prompt item is added to the output SNP-induced mRNA allosteric risk score.
12. The method for calculating the SNP-induced mRNA allosteric score according to claim 11, wherein: The distance between the SNP and the base undergoing conformational change is the distance between the absolute position of the SNP on the genome and the absolute position of the base undergoing conformational change on the genome.
13. The method for calculating the SNP-induced mRNA allosteric score according to claim 11, wherein: The threshold is 350 bp.
14. A device for calculating a SNP-induced mRNA allosteric score, the device comprising: memory and processor; The memory is used to store program instructions; The processor is used to call program instructions, and when the program instructions are executed, is used to execute the method for calculating the SNP-induced mRNA allosteric score according to any one of claims 1 to 13.
15. A system for calculating a SNP-induced mRNA allosteric score, comprising: An acquisition unit, used to acquire a pair of SNPs and their corresponding precursor mRNA sequences; an extraction unit, configured to extract M splicing-related elements from the precursor mRNA sequence pair, and extract RSSs from the M splicing-related elements to obtain the RSSs of the splicing-related elements, wherein M is a natural number ranging from 2 to 10, and the RSSs are RNA secondary structures; the RSSs include stems, hairpin loops, inner loops, outer loops, raised loops, and multi-branched loops; and the M splicing-related elements include one or more of the following splicing-related elements: a 5' splice site, a 3' splice site, a branch point, a polypyrimidine tract mRNA splicing region, and a splicing regulatory element; a calculation unit, configured to calculate an allosteric probability of the RSS of the splicing-related element, wherein the allosteric probability is a ratio of the total number of allosteric nucleotides in the RSS of the splicing-related element to the total number of nucleotides in the RSS of the splicing-related element, wherein the allosteric nucleotide refers to a nucleotide whose RSS of the precursor mRNA sequence having the SNP is changed relative to the precursor mRNA sequence not having the SNP; The output unit is used to output the SNP-induced mRNA allosteric risk score, where the allosteric risk score is the allosteric probability of the RSS of the splicing-related element.
16. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for calculating the SNP-induced mRNA allosteric score according to any one of claims 1 to 13 is implemented.
Citation Information
Patent Citations
Method for identifying altered vitamin D metabolism
CN101448955A
Splice modulating oligonucleotides and methods of use thereof
WO2016027168A2