Prediction method for RNA variable region structure, device, medium and program product
By predicting RNA variable region structure using cluster analysis and the minimum free energy principle, and combining this with Euclidean distance to quantify the influence of SNPs, the problems of RNA secondary structure prediction and SNP change identification were solved, enabling accurate assessment of RNA structural changes and support for therapeutic methods.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- THE EYE HOSPITAL OF WENZHOU MEDICAL UNIVERSITY
- Filing Date
- 2025-04-30
- Publication Date
- 2026-07-30
AI Technical Summary
Existing technologies struggle to accurately predict RNA secondary structure and recognize the impact of single nucleotide polymorphisms (SNPs) on RNA structural changes, hindering our understanding of RNA function and biological significance, as well as the development of RNA-based therapeutics.
RNA variable region structure is predicted by cluster analysis and the minimum free energy principle. The influence of SNPs on RNA structure is quantified by Euclidean distance. The minimum free energy structure is adopted as the optimal structure, and RNA structural changes are identified based on SNPs.
It enables accurate prediction of RNA variable region structure and quantification of SNP impact, providing a comprehensive assessment of RNA structural changes and supporting RNA function research and the development of RNA-based therapeutics.
Smart Images

Figure CN2025092626_30072026_PF_FP_ABST
Abstract
Description
A method, apparatus, medium, and program product for predicting the structure of RNA variable regions. Technical Field
[0001] This invention relates to the field of bioanalysis, and more specifically, to a method, apparatus, medium, and program product for predicting the structure of variable regions of RNA. Background Technology
[0002] RNA structure plays a crucial role in many biological processes, including gene regulation and signal transduction. Therefore, determining the structure-function relationship of RNA is both necessary and a significant challenge for better understanding the mechanisms of biological processes. Nucleotides in an RNA molecule are arranged in different sequences to form the RNA sequence, which is the primary structure of RNA. RNA molecules contain many planar structures composed of complementary base pairs, such as single-stranded regions, stem-loop structures, and double-stranded structures. These structures undergo self-folding motion, forming the secondary structure (RSS) of RNA. The tertiary structure of RNA is a higher-level three-dimensional structure built upon the RNA secondary structure. Besides the base pairing interactions, there are also interactions within the RNA molecule, such as interactions between the main strands, between the main strand and bases, and between isolated hydrogen bonds. These interactions cause the planar RNA secondary structure to fold into a compact spatial structure. RNA secondary structure motifs are fundamental building blocks for studying structural biology mechanisms.
[0003] The structure of RNA is crucial to its function and biological significance; the structural diversity of RNA determines the variety of biological functions it can perform in cells. Research on RNA structure not only helps us understand its biological functions but also contributes to the development of RNA-based therapeutics, such as RNA interference (RNAi) and RNA-based vaccines. Therefore, predicting the actual structure of RNA has vital research and application value in fields such as molecular biology and disease research. Summary of the Invention
[0004] This invention aims to at least solve one of the technical problems existing in the prior art. To this end, this invention provides a method, device, medium, and program product for predicting the structure of RNA variable regions, as well as a method, device, medium, and program product for identifying RNA structural changes based on SNPs. The method of this invention achieves optimization through the step of selecting the optimal structure. Based on the hypothesis that the structure with the lowest free energy is the most stable, it selects the structure with the lowest free energy from among many suboptimal structures as the optimal structure, thus obtaining an RNA structure that is closest to the true structure.
[0005] The first aspect of this application discloses a method for predicting the structure of RNA variable regions, the method comprising:
[0006] 101. Obtain the target RNA to be predicted;
[0007] 102. Based on the target RNA, a set of RNA secondary structures is obtained;
[0008] 103. Perform cluster analysis on the set of RNA secondary structures to obtain at least two clusters, and select the centroid structure and its energy value for each cluster; the number of centroid structures is consistent with the number of clusters;
[0009] 104. Compare the energy values of the centroid structures and select the centroid structure with the lowest energy value as the predicted RNA variable region structure.
[0010] In some embodiments, the method for selecting the centroid structure with the lowest energy value includes: performing pairwise cyclic comparisons or traversal comparisons on the centroid structures in each cluster to determine the centroid structure with the lowest minimum free energy as the optimal structure.
[0011] In some embodiments, the energy value is a comprehensive assessment of the free energy of the overall structure based on the thermodynamic stability principles recognized in structural prediction; the thermodynamic stability principles include any one or more of the following structural combinations: base pair energy, unpaired regions, and ring structures.
[0012] In some embodiments, the method for selecting the centroid structure of each cluster includes: obtaining the structure with the highest frequency of occurrence in each cluster based on the statistical Boltzmann sampling method, as the representative structure of each cluster, i.e., the centroid structure.
[0013] The second aspect of this application discloses a method for identifying RNA structural changes based on SNPs, the method comprising:
[0014] 201. Obtain the SNPs to be tested;
[0015] 202, based on the SNPs to be tested, retrieve the Ref and Alt sequence pairs, the sequences corresponding to the target RNA;
[0016] 203. Based on the RNA variable region structure prediction method described in the first aspect of this application, RNA secondary structure features of the target RNA structure are extracted; the secondary structure features include B, E, H, I, M, and S subunits, as well as the number, length, and position of each subunit;
[0017] 204. The structural differences between the Ref and Alt sequences are quantified by calculating the B, E, H, I, M, and S subunits, as well as the number, length, and position of each subunit, using Euclidean distance to obtain the difference values.
[0018] 205, based on the difference values, output the results of the influence of SNPs on RNA structure changes.
[0019] In some embodiments, between 203 and 204, the method further includes standardizing the number, length, and position of each subunit, wherein the standardization method is as follows:
[0020] Where, N i,j Represents any 6×2 vector, where Nori∈(Num,len,Loc), j∈(B,E,H,I,M,S) are standardized values ranging from 0 to 1; B, E, H, I, M, and S represent convex loop, outer loop, hairpin loop, inner loop, multi-branched loop, and stem, respectively; R i∈(Num,Len,Loc) and A i∈(Num,Len,Loc) These represent the secondary structures of ref and alt RNA induced by SNPs, respectively.
[0021] In some embodiments, the quantization method in 204 includes:
[0022] M represents the characteristic of each RNA subunit, N refers to the B, E, H, I, M, and S subunits, and Nor(R) and Nor(A) represent the normalized values of the characteristics of each RNA subunit; Euc R,A The values represent the differences obtained after quantization; B, E, H, I, M, and S represent the convex ring, outer ring, hairpin ring, inner ring, multi-branched ring, and stem, respectively.
[0023] A third aspect of this application discloses a computer device, the device comprising: a memory and a processor; the memory being used to store a computer program; and the processor executing the computer program to implement the steps of the above-described method.
[0024] The fourth aspect of this application discloses a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.
[0025] The fifth aspect of this application discloses a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.
[0026] This application has the following beneficial effects:
[0027] 1. This application innovatively discloses a method for predicting the structure of RNA variable regions, assessing changes in RNA subunits in regions of significant global RNA structural variation caused by SNPs. First, RNA secondary structures are predicted to be as consistent as possible with the actual folding state. Cluster analysis is performed on 1000 possible structures, resulting in multiple clusters. By comparing the minimum free energy (MFE) of RNA structures in each cluster, the structure with the lowest MFE is selected as the most likely structure. Optimization is achieved in the optimal structure selection step. Based on the hypothesis that the minimum free energy structure is the most stable, the minimum free energy structure among numerous suboptimal structures is selected as the optimal structure; that is, the centroid structures are compared pairwise in a cyclical manner, ultimately determining the minimum free energy structure as the optimal structure. Experimentally validated stRNA structures are obtained from the RNASTRAND public database. The correlation between the selected minimum free energy structure and the experimentally validated structure is compared, and a structure comparison algorithm confirms that the two structures are completely identical.
[0028] 2. This application primarily addresses the problem of obtaining the optimal structure. The software (Sfold) performs statistical sampling on 1000 structures, providing the centroid structure and its energy value. Our approach involves extracting the energy value of the centroid structure, comparing the free energies of the structures, and selecting the structure with the lowest free energy as the optimal structure. Here, the lowest free energy is based on the generally accepted thermodynamic stability principle in structure prediction. It considers the free energy of various structural combinations, including base pair energies, unpaired regions, and ring structures, comprehensively evaluating the free energy of the overall structure. Based on the hypothesis that a thermodynamically stable state typically has the lowest free energy, the structure with the lowest free energy is obtained.
[0029] 3. When selecting the optimal structure, i.e. the final RNA structure, the Sfold algorithm does not explicitly indicate which type of structure set represents the optimal structure. In this case, the statistical Boltzmann sampling method is mainly used to obtain the cluster centroid structure with the highest frequency in each cluster as the representative structure set of each cluster, and further screening is used to obtain the optimal RNA structure.
[0030] 4. This application innovatively discloses a method for identifying RNA structural changes at the post-transcriptional level of the genome. This method retrieves Ref and Alt sequence pairs based on SNPs, extracts secondary structure features of target RNA based on sequence pairs, quantifies the secondary structures in the Ref and Alt sequences using Euclidean distance to obtain difference values, identifies the influence of SNPs on RNA structural changes based on the difference values, and explores interference with post-transcriptional regulation.
[0031] 5. This application innovatively develops a comprehensive process for evaluating the global and local effects of SNPs on variable regions of RNA structure, providing an effective method for assessing RNA accessibility and stability. Attached Figure Description
[0032] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0033] Figure 1 is a schematic flowchart of a method provided in the first aspect of an embodiment of the present invention;
[0034] Figure 2 is a schematic flowchart of a method provided in the second aspect of an embodiment of the present invention;
[0035] Figure 3 is a schematic diagram of the RNA variable region structure prediction system provided in an embodiment of the present invention;
[0036] Figure 4 is a schematic diagram of the recognition system for identifying RNA structural changes based on SNPs provided in an embodiment of the present invention;
[0037] Figure 5 is a schematic diagram of a computer device provided in an embodiment of the present invention;
[0038] Figure 6 is a schematic diagram of the architecture of an exemplary computing device provided in an embodiment of the present invention;
[0039] Figure 7 is a schematic diagram of the storage medium provided in an embodiment of the present invention;
[0040] Figure 8 is a fine localization map of myopia-related SNPs on genome-wide regulatory elements provided in the embodiments of the present invention;
[0041] Figure 8A shows the distribution of each genomic regulatory element during transcription and post-transcriptional processes; Figure 8B shows the percentage of relation pairs during transcription and post-transcriptional processes; Figure 8C shows the length of genomic regulatory elements covering myopia-related SNPs, with the black-marked horizontal lines in each violin representing the median length of each element; Figure 8D shows the SNP density of each regulatory element at the transcriptional and post-transcriptional levels. The dotted lines indicate the average SNP density across the entire human genome.
[0042] Figure 9 is a schematic diagram illustrating the heterogeneity of SNP-mediated TF binding among the four regulatory elements during transcription provided in this embodiment of the invention; wherein, Figure 9A shows the heterogeneity of TF binding to enhancers, Figure 9B shows the heterogeneity of TF binding to OCRs, Figure 9C shows the heterogeneity of TF binding to promoters, and Figure 9D shows the heterogeneity of TF binding to PFRs; P T The value was calculated using Fisher's exact test, and FC was calculated using multiple changes.
[0043] Figure 10 illustrates the SNP-mediated global RNA structural variable regions on genomic regulatory elements during post-transcriptional processes, as provided in this embodiment of the invention. Figure 10A shows the length of the RNA structural variable regions within the genomic regulatory elements; Figure 10B shows the percentage of relation pairs illustrating the proximal and distal effects mediated by SNPs on genomic regulatory elements; Figure 10C shows the percentage of relation pairs with significant RNA structural variable regions on genomic regulatory elements; Figure 10D shows the percentage of relation pairs in the CM covering significant RNA structural variable regions; and Figure 10E shows the percentage of relation pairs in the HM covering significant RNA structural variable regions.
[0044] Figure 11 illustrates the analysis of important RNA structural variable regions based on RNA subunits provided in this embodiment of the invention. Figure 11A shows the RNA secondary structure of selenocysteine transfer RNA (stRNA) obtained by Sfold, with the stRNA structural conformation predicted by Sfold (left) and downloaded via RNASTRAND (right). Figure 11B visualizes the six subunits of the RNA secondary structure. Figure 11C shows the changes in the single-stranded or double-stranded state of RNA in important RNA structural variable regions at SNP sites on RNA subunits. Figure 11D shows the detection of cfSNPs that can affect RNA subunits based on Euclidean distance. The threshold (1.65) is represented by a horizontal line. Detailed Implementation
[0045] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0046] In some of the processes described in the specification, claims, and accompanying drawings of this invention, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0047] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0048] Figure 1 is a schematic flowchart of a method for predicting the structure of RNA variable regions according to an embodiment of the present invention. Specifically, the method includes the following steps:
[0049] 101: Obtain the target RNA to be predicted;
[0050] RNA is an abbreviation for ribonucleic acid, a biological macromolecule closely related to DNA. Both are composed of nucleotides, but differ in structure and function. RNA plays multiple roles in cells, including the transmission of genetic information, protein synthesis, and the regulation of gene expression.
[0051] The main types of RNA include: Messenger RNA (mRNA): It carries genetic information from DNA and guides protein synthesis. mRNA is synthesized from a DNA template through transcription, then transported to the cytoplasm and translated into protein on ribosomes. Ribosomal RNA (rRNA): It is an important component of ribosomes, the site of protein synthesis. rRNA, along with mRNA and transfer RNA (tRNA), participates in protein synthesis. Transfer RNA (tRNA): It is responsible for transporting amino acids to ribosomes and linking them together according to the codon sequence on mRNA to form a protein chain. MicroRNA (miRNA): It is a class of small RNA molecules that mainly regulate gene expression post-transcriptionally. By binding to mRNA, it affects its stability and translation efficiency, thereby regulating protein synthesis. Long non-coding RNA (lncRNA): These are longer RNA molecules that do not encode proteins but play an important role in gene expression regulation, including processes such as chromatin remodeling, transcription, and splicing.
[0052] 102: Obtain a set of RNA secondary structures based on the target RNA;
[0053] In some embodiments, the software for predicting the set of RNA secondary structures includes any one or more of the following: RNAfold, Sfold, CONTRAfold, with Sfold being the preferred software.
[0054] 103: Cluster analysis was performed on the RNA secondary structure set to obtain at least two clusters. The centroid structure and its energy value of each cluster were selected. The number of centroid structures was consistent with the number of clusters. The centroid structures were Cluster 1 centroid, Cluster 2 centroid and Ensemble centroid in Figure 11A. Centroid can be translated as centroid. The centroid structure of each cluster was marked with red dots in the figure.
[0055] In some embodiments, the method for selecting the centroid structure of each cluster includes: obtaining the structure with the highest frequency of occurrence in each cluster based on the statistical Boltzmann sampling method, as the representative structure of each cluster, i.e., the centroid structure.
[0056] 104: Compare the energy values of the centroid structures and select the centroid structure with the lowest energy value as the predicted RNA variable region structure.
[0057] In some embodiments, the method for selecting the centroid structure with the lowest energy value includes: performing pairwise cyclic comparisons or traversal comparisons on the centroid structures in each cluster to determine the centroid structure with the lowest minimum free energy as the optimal structure.
[0058] In some embodiments, the energy value is a comprehensive assessment of the free energy of the overall structure based on recognized thermodynamic stability principles in structural prediction; these thermodynamic stability principles include any one or more of the following structural combinations: base pair energy, unpaired regions, and ring structures. The focus is on defining the optimal structure based on the cluster centroid structure set provided by Sfold, using the minimum free energy principle.
[0059] As shown in Figure 2, the second aspect of this application discloses a method for identifying RNA structural changes based on SNPs, the method comprising:
[0060] 201. Obtain the SNPs to be tested;
[0061] In some embodiments, SNP stands for Single Nucleotide Polymorphism, which refers to a variation in a single nucleotide (base pair) in a genomic DNA sequence. This variation can be a substitution (e.g., A to G), an insertion, or a deletion. SNPs are one of the most common forms of genetic variation in the genome, and their distribution in a population is polymorphic, meaning that different individuals may have different nucleotides at the same location. In some embodiments, the SNP to be tested is derived from a subject. The terms "subject," "test subject," or "sample to be tested" as used herein refer to any animal (e.g., a mammal), including but not limited to humans, non-human primates, rodents, etc., who will become the recipient of a specific treatment. Generally, the terms "subject" and "patient" are used interchangeably herein when referring to human subjects. Preferably, the subject is a human.
[0062] 202, based on the SNPs to be tested, retrieve the Ref and Alt sequence pairs, the sequences corresponding to the target RNA;
[0063] In some embodiments, the Ref and Alt sequence pairs corresponding to each SNP are retrieved. Specifically, several sequence pairs are obtained by retrieving Mbp upstream and downstream of the Ref and Alt sequences, where M is a natural number greater than 1; the range of M is 10-35, preferably 20. Specifically, Ref represents the wild type, and Alt represents the mutant type. The "several" refers to a natural number greater than 1.
[0064] 203. Based on the RNA variable region structure prediction method described in the first aspect of this application, RNA secondary structure features of the target RNA structure are extracted; the secondary structure features include B, E, H, I, M, and S subunits, as well as the number, length, and position of each subunit;
[0065] 204. The structural differences between the Ref and Alt sequences are quantified by calculating the B, E, H, I, M, and S subunits, as well as the number, length, and position of each subunit, using Euclidean distance to obtain the difference values.
[0066] 205, based on the difference values, output the results of the influence of SNPs on RNA structure changes.
[0067] In some embodiments, between 203 and 204, the method further includes standardizing the number, length, and position of each subunit, wherein the standardization method is as follows:
[0068] Where, N i,j Represents any 6×2 vector, where Nori∈(Num,len,Loc), j∈(B,E,H,I,M,S) are standardized values ranging from 0 to 1; B, E, H, I, M, and S represent convex loop, outer loop, hairpin loop, inner loop, multi-branched loop, and stem, respectively; R i∈(Num,Len,Loc) and A i∈(Num,Len,Loc) These represent the secondary structures of ref and alt RNA induced by SNPs, respectively.
[0069] In some embodiments, the quantization method in 204 includes:
[0070] M represents the characteristic of each RNA subunit, N refers to the B, E, H, I, M, and S subunits, and Nor(R) and Nor(A) represent the normalized values of the characteristics of each RNA subunit; Euc R,A The values represent the differences obtained after quantization; B, E, H, I, M, and S represent the convex ring, outer ring, hairpin ring, inner ring, multi-branched ring, and stem, respectively.
[0071] In some embodiments, if the difference value exceeds a threshold, the result indicating that the SNP has an impact on RNA structural changes is output. The threshold is determined by assessing the lower quartile of RNA subunits where significant changes occur.
[0072] In some embodiments, the posttranscriptional genomic regulatory elements include any one or more of the following: 3'UTR, 5'UTR, exons, introns, and long non-coding RNAs.
[0073] In some embodiments, the method further includes: identifying RNA structurally variable regions of posttranscribed gene regulatory elements in the Ref and Alt sequence pairs; extracting RNA secondary structure features from the RNA structurally variable regions; wherein the RNA structurally variable regions include global RNA structurally variable regions and / or local RNA structurally variable regions.
[0074] In some embodiments, SNPs in the globally variable RNA structural region are classified into proximal allosteric effects and distal allosteric effects. If the SNP is located in the globally variable RNA structural region and affects the folding of that region, the output is classified as a proximal allosteric effect; if the SNP is not located in the globally variable RNA structural region but affects the folding of that region, the output is classified as a distal allosteric effect. Specifically, a proximal allosteric effect refers to affecting a small surrounding region, and in molecular experiments, only the gene currently affected by the SNP or other genes are considered; a distal allosteric effect refers to affecting genes other than the current gene, possibly other genes, and in molecular experiments, other genes are considered first.
[0075] In some embodiments, the evaluation results or prediction results may be in the form of paper or electronic reports, but are not limited to those obtained by intelligent machines based on the relevant data of the test subjects. These results are for reference only and are not considered as final diagnostic results.
[0076] In this embodiment, the threshold is obtained through training on training set samples. It can be a specific threshold or an interval range. The specific form is not specifically limited in this embodiment.
[0077] Figure 5 is a schematic diagram of a computer device provided in an embodiment of the present invention. As shown in Figure 5, the device 2000 may include: one or more processors 2010 and one or more memories 2020; wherein, the memory stores computer-readable code, which, when run by the one or more processors, can execute the method described above.
[0078] The processor in this embodiment can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, operations, and logic block diagrams disclosed in this embodiment. The general-purpose processor can be a microprocessor or any conventional processor, and can be based on an x86 or ARM architecture.
[0079] In general, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of embodiments of this disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0080] For example, the methods or apparatus according to embodiments of this disclosure can also be implemented using the architecture of the computing device 3000 shown in FIG. 6. As shown in FIG. 6, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. Storage devices in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for processing and / or communication of the methods provided in this disclosure, as well as program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, the architecture shown in FIG. 6 is merely exemplary, and one or more components in the computing device shown in FIG. 6 may be omitted as needed when implementing different devices.
[0081] This invention also provides a computer-readable storage medium, as shown in FIG7, which is a schematic diagram of a storage medium 4000 provided in an embodiment of this invention. The computer storage medium 4020 stores computer-readable instructions 4010. When the computer-readable instructions 4010 are executed by a processor, the method according to the embodiments of this disclosure described with reference to the above figures can be performed. The computer-readable storage medium in the embodiments of this disclosure can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), Synchronous Link Dynamic Random Access Memory (SLDRAM), and Direct Memory Bus Random Access Memory (DR RAM). It should be noted that the memory used in the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0082] This disclosure also provides a computer program product or system, including a computer program that, when executed by a processor, implements the steps of the above-described method.
[0083] In some embodiments, this embodiment also discloses a prediction system for RNA variable region structure, as shown in Figure 3, the system comprising:
[0084] RNA acquisition module 301, used or configured to acquire target RNA to be predicted;
[0085] RNA secondary structure extraction module 302 is used or configured to obtain a set of RNA secondary structures based on the target RNA;
[0086] The centroid structure extraction module 303 is used or configured to perform cluster analysis on the RNA secondary structure set to obtain at least two clusters, and select the centroid structure and its energy value for each cluster; the number of centroid structures is consistent with the number of clusters;
[0087] RNA structure prediction module 304 is used or configured to compare the energy values of the centroid structures and select the centroid structure with the lowest energy value as the predicted RNA structure.
[0088] In some embodiments, this embodiment also discloses a recognition system for identifying RNA structural changes based on SNPs, as shown in Figure 4. The system includes:
[0089] SNP acquisition module 401 is used or configured to acquire SNPs to be tested;
[0090] Sequence pair extraction module 402 is used or configured to retrieve Ref and Alt sequence pairs based on the SNPs to be tested, wherein the sequences correspond to the target RNA;
[0091] RNA secondary structure feature extraction module 403 is used or configured to extract RNA secondary structure features of a target RNA structure based on the RNA variable region structure prediction method described in the first aspect of this application; the secondary structure features include B, E, H, I, M, and S subunits, as well as the number, length, and position of each subunit;
[0092] The difference value calculation module 404 is used or configured to calculate the structural differences between the Ref sequence and the Alt sequence by using Euclidean distance to calculate the B, E, H, I, M, and S subunits, as well as the number, length, and position of each subunit;
[0093] RNA structure change influence module 405 is used or configured to output the influence results of SNPs on RNA structure changes based on difference values.
[0094] Specific implementation examples:
[0095] 1. Materials and Methods:
[0096] 1.1 Data Collection:
[0097] Myopia-related SNPs were collected from dbGap, the GWAS catalog, and published literature. Based on the population distribution of myopia-related SNPs in the 1000 Genomes Project (GRCh38), we performed quality control (QC) on the samples and genotypes. SNPs were filtered according to the following selection criteria: East Asian, two alleles in NCBI, minor allele frequency (MAF) > 5%, Hardy-Weinberg equilibrium P < 0.01, recall > 75%, and genotyping rate > 75%. 343 SNPs were obtained for the following analysis. The start and termination sites of eight genomic regulatory elements, including open chromatin regions (OCRs), CTCF binding sites (CTCFBSs), enhancers, promoters, promoter flanking regions (PFRs), exons, introns, and non-coding RNA (ncRNA) transcripts, were derived from ENSEMBL (v102), and their sequences were extracted using NCBI's Refseq. The 3' UTR and 5' UTR sequences were obtained using the BioMart tool in ENSEMBL.
[0098] 1.2 Fine localization of myopia-related SNPs on 10 genomic regulatory elements:
[0099] Fine mapping was performed using the classic read alignment tool Bowtie2. Primary sequences of genomic regulatory elements were considered long reference reads. The 30bp flanking regions upstream and downstream of these SNPs were considered short alignment sequences. Strict parameters were then set using "--n-ceil C,3--np 0--end-to-end-a--score-minC,0" to avoid mismatches for each seed. After fine mapping, paired reference (Ref) and alternative (Alt) sequences were constructed based on the alleles of the SNPs. The alignment software included Bowtie1, Bowtie2, and BLAST, with Bowtie2 being preferred.
[0100] 1.3 Assessing changes in molecular binding at the transcriptional level:
[0101] We extracted 20 bp upstream and downstream of the SNP-induced paired Ref and Alt sequences and obtained the promoter, open chromatin region, promoter, and promoter flanking region binding motifs of these sequences from the HumanTFDB database, while CTCF was obtained from the CTCFBSDB database. For protein binding, we obtained the number of regulatory proteins (TF and CTCF), protein symbols, and the start and end positions of the binding motifs. First, we used fold change (FC) to obtain the changes in the number of protein binding sites mediated by myopia-related SNPs:
[0102] Num Ref and Num AltFC represents the number of proteins bound to the Ref and Alt sequences, respectively. We define FC > 1 as a protein binding loss, and conversely, FC < 1 as a protein binding gain. We then apply Fisher's exact test to determine changes in the protein binding motif, as shown below: P T =(a / b) / (c / d)
[0103] a represents the number of binding motifs that overlap only with the Ref sequence, b represents the number of binding motifs that do not overlap with the Ref sequence, c represents the number of binding motifs that overlap only with the Alt sequence, and d represents the number of binding motifs that do not overlap with the Alt sequence. Here, when FC is not equal to 1 and P T A score <0.05 indicates that myopia-related SNPs significantly affect protein binding, and these SNPs are identified as cfSNPs. Furthermore, we define the protein with the highest score as the "leader" protein.
[0104] 1.4 Identifying global RNA structural variable regions after transcription:
[0105] At the post-transcriptional level, we used RNAsnp to identify the structural variable regions of five regulatory elements, including the 3'UTR, 5'UTR, exons, introns, and lncRNA transcripts. Based on thresholds set by RNAsnp, P... PT A value <0.2 indicates that SNPs may have a significant impact on local RNA structure. Here, we define these regions as "global RNA structurally variable regions". If an SNP is located in a RNA structurally variable region and affects the folding of that region, these SNPs are defined as "proximal allosteric effects". Otherwise, if an SNP is not located in a RNA structurally variable region but affects the folding of surrounding areas, these SNPs are defined as "distal allosteric effects".
[0106] 1.5 Quantifying changes in the variable region of RNA structure using RNA subunits:
[0107] We evaluated the changes in RNA subunits within key global RNA structural variable regions induced by SNPs. First, we predicted RNA secondary structures that were as consistent as possible with the actual folding states. Cluster analysis was performed based on 1000 possible structures to obtain representative structures from multiple clusters. By comparing the minimum free energy (MFE) of RNA structures within each cluster, we selected the structure with the lowest MFE as the most probable structure. Then, we used RNAsmc to extract six RNA subunits for each RNA: the convex loop (B), outer loop (E), hairpin loop (H), inner loop (I), multi-branched loop (M), and stem (S). Next, we developed a computational pipeline and quantified the structural heterogeneity (Euclidean distance) of the global RNA structural variable regions using Euclidean distance (based on RNA subunits). R,AWe obtained the features of RNA subunits in the global RNA structural variable region, namely the number, length, and position of each subunit in the paired Ref and Alt structures. For these three dimensions, we constructed two corresponding 6×1 vectors: R i∈(Num,Len,Loc) =(R iB ,R iE ,R iH ,R iI ,R iM ,R is A i∈(Num,Len,Loc) =(A iB A iE A iH A iI A iM A is )
[0108] Here, 'i' represents the characteristic of the RNA subunit. B, E, H, I, M, and S refer to the six RNA subunits: convex loop, outer loop, hairpin loop, inner loop, multi-branched loop, and stem, respectively. R and A represent the secondary structures of Ref and AltRNA mediated by SNPs, respectively. Then, to normalize the size of the three characteristics of the RNA subunits, each type of subunit is normalized as follows:
[0109] N i,j Let represent any 6×2 vector, where Nori ∈ (Num, len, Loc), j ∈ (B, E, H, I, M, S) ranging from 0 to 1. Finally, we use Euclidean distance (Euclidean distance). R,A To quantify the differences between Ref and Alt structures:
[0110] M represents the feature of each sub-unit. N refers to the six sub-units. Nor(R) and Nor(A) represent the standardized values of each feature, respectively. In this study, Euc R,A The lower quartile was set as the significance threshold. If Euc R,A Above the lower quartile, we consider RNA subunits to exhibit significant changes. VARNAs are used to visualize RNA secondary structure.
[0111] 1.6 Same annotations for evaluating SNPs based on HM queues and computational pipelines:
[0112] A total of 10,348 highly myopic participants (worst eye SE < -6.00D) in the CAMS study were sequenced using the Twist HumanCoreExomeKit on a Berry Genomics Illumina NovaSeq 6000 sequencer. We obtained the genotypes and phenotypes of myopia-related SNPs from the CAMS study. Next, we used the Combined Annotation-Dependent Depletion score (CADD) to score the harmfulness of SNPs based on genotype and phenotype in the highly myopic cohort. Here, an SNP with a CADD score ≥ 10 was defined as a harmful SNP, potentially leading to loss of gene function. The Combined Annotation-Dependent Depletion score (CADD score) is used to assess and quantify single nucleotide variants (SNVs); a higher score indicates more harmful variants, i.e., a higher probability of pathogenicity.
[0113] 2. Results:
[0114] 2.1 Myopia-related SNPs are widely distributed at the transcriptional and post-transcriptional levels:
[0115] We obtained myopia-related SNPs from public resources, and after quality control, 343 SNPs were used for subsequent analysis (see Methods). To reveal a complete map of SNPs at the transcriptional and post-transcriptional levels, fine mapping was used to identify 636 relationship pairs formed by 263 SNPs, 10 genomic regulatory elements, and myopia (Figure 8A). Furthermore, these SNPs were associated with five phenotypes: refractive error (RE), common myopia (CM), high myopia (HM), pathological myopia (PM), and visual impairment (VD). During transcription, 84 SNPs were located in enhancers, open chromatin regions (OCRs), CTCF binding sites (CTCFBSs), promoters, and promoter flanking regions (PFRs), forming 90 relationship pairs. A total of 244 SNPs were located on five genomic regulatory elements: the 5'UTR, exons, introns, the 3'UTR, and lncRNAs, constituting 546 relationship pairs at the post-transcriptional level. Next, we found that 14.15% (90 / 636) and 85.85% (546 / 636) of the pairs were enriched during transcription and post-transcriptional processes, respectively (Fig. 8B).
[0116] To further assess the average distribution of SNPs across all genomic regulatory elements, we analyzed the transcriptional length of these elements and the density of SNPs within each element. The results showed that the median length of the longest transcript, LncRNA, was 95.53 times that of the shortest transcript, CTCFBS (Fig. 8C). We then further evaluated the enrichment of SNPs within each 1000 bp of each regulatory element. Compared to the average of one SNP per 1000 bp in the human genome, we observed highly distributed myopia-associated SNPs in the OCR, CTCFBS, 5'UTR, and exons, respectively (Fig. 8D). Analysis of myopia revealed that 81.76% (520 / 636) of the relationships were associated with myopia (CM). Furthermore, little was known about HM, PM, RE, and VD, with approximately 12.11% (77 / 636), 5.19% (33 / 636), 0.63% (4 / 636), and 0.31% (2 / 636), respectively. These data indicate that the distribution of myopia-related SNPs varies at the transcriptional and post-transcriptional levels as well as among genomic regulatory elements. These differences may be related to the severity of myopia and underlying molecular regulation.
[0117] 2.2 Scoring of SNP-induced molecular binding heterogeneity during transcription:
[0118] To investigate SNP-mediated molecular binding heterogeneity at the transcriptional level, we developed a computational pipeline that uses the fold-over test (FC) to assess changes in the number of binding proteins and applies Fisher's exact test to assess changes in the number of binding protein molecules. Here, a threshold P is used. T <0.05, FC not equal to 1, and found that 38.46% (5 / 13), 80% (4 / 5), 57.89% (11 / 19), and 44.73% (17 / 38) of SNPs could disrupt transcription factor binding to enhancers, OCR, promoters, and PFRs, respectively (Figure 9). For the CTCF protein, we found that approximately 13.33% (2 / 15) of SNPs could disrupt the CTCFBS interaction. In total, 43.33% (39 / 90) of one-to-one pairings could disrupt binding affinity, and 46.43% (39 / 84) of SNPs were recognized as “cfSNPs” during transcription. In conclusion, the OCR was significantly enriched with myopia-associated SNPs and cfSNPs, exhibiting high-density distribution and disrupted molecular binding.
[0119] To determine the potential influence of cfSNP-related regulatory proteins associated with or potentially associated with myopia, we explored the molecular functions of the “leader” proteins in each relationship pair in published studies. Interestingly, most “leader” proteins were able to influence ocular tissues or structures. For example, approximately 7.69% (3 / 39) of cfSNPs induced changes in the REST binding motif, thereby affecting the fate of retinal ganglion cells (RGCs) in the developing retina. Approximately 7.69% (3 / 39) of cfSNPs induced changes in IRF1 binding, IRF1 being known to be expressed in retinal microglia and play a key role in microglia activation and retinal inflammation. Furthermore, 15.28% (6 / 39) of cfSNPs altered the binding motif site of SPI1, which has been reported to regulate microglia in the retina. In summary, we found abundant leader proteins in retinal inflammation that have been confirmed to be involved in the occurrence and development of myopia. Moreover, this can help us reveal potential protein regulators and understand how SNPs participate in the molecular regulation of myopia. Finally, we assessed the distribution of significantly altered relationships between myopia types at the transcriptional level. Over 75% (3 / 4) were associated with PM, while approximately 44.44% (4 / 9) and 41.56% (32 / 77) were associated with HM and CM (Figures 9A-D). The results also indicated that the impact of SNPs on myopia increases with the severity of myopia.
[0120] 2.3 Identification of SNP-mediated global RNA structural variable regions during post-transcriptional processes:
[0121] To identify SNP-induced RNA secondary structure heterogeneity during post-transcriptional processing, we initially used RNA snp to detect potential global RNA structural variable regions. PT <0.2. Here, since there is only one relational pair in the 5'UTR, this exon was chosen as a reference to compare the significance of differences in the average length of RNA structural variable regions. As shown in Figure 10A, the RNA structural variable regions of the 5'UTR, 3'UTR, and exon are relatively longer than those of introns and lncRNAs. Notably, myopia-related SNPs are not only highly enriched in the 5'UTR and exons, but also cause significant structural damage to these regulatory regions. Although myopia-related SNPs are not significantly enriched in the 3'UTR region, they still cause significant structural effects. This may be because the 3'UTR region exhibits a highly structured feature, and once disrupted, it has a significant impact on RNA secondary structure.
[0122] Furthermore, we assessed whether SNPs were located in regions of RNA structural variability. Definitions are shown in the methods. Statistically, approximately 97.54% (238 / 244) of the SNPs exhibited proximal allosteric effects, mapping to 91.96% (502 / 546) of relation pairs and all five regulatory elements, while only 12.70% (31 / 244) of the SNPs showed proximal allosteric effects. SNPs exhibited distal allosteric effects, mapping to 8.06% (44 / 546) of pairwise relationships and three regulatory elements: exons, introns, and lncRNAs (Figure 10B). This finding is consistent with a previous study that SNPs primarily affect local regions rather than having a global impact.
[0123] To further explore which genomic regulatory elements were significantly disrupted by SNPs, we analyzed the proportion of important RNA structurally variable regions within these elements. As shown in Figure 10C, in 9 relation pairs in the 3'UTR, 33.33% (3 / 9) exhibited significant RNA structurally variable regions induced by 37.50% (3 / 8) of the SNPs, while 1 relation pair in the 5'UTR did not show significant RNA structurally variable regions. In 35 relation pairs in exons, 22.86% (8 / 35) showed significant changes, mediated by 20.69% (6 / 29) of the SNPs. Furthermore, in 308 relation pairs in introns, 14.61% (45 / 308) showed significant RNA structurally variable regions, which were affected by 16.52% (38 / 230) of the SNPs. Of the 193 relationship pairs of lncRNAs, 12.95% (25 / 193) showed significant RNA structural variable regions, induced by 14.56% (23 / 158) of the SNPs. Overall, we found that 14.84% (81 / 546) of the relationship pairs contained significant RNA structural variable regions induced by 20.08% (49 / 244) of the SNPs, which were identified as cfSNPs.
[0124] To characterize the role of RNA secondary structure in myopia, we analyzed the distribution of important variable regions of RNA structure. The three myopia types—RE, CM, and HM—exhibited significant SNP-mediated changes. Among these phenotypes, 1.23% (1 / 81) of the one-to-one correlations were associated with RE, 16.05% (13 / 81) with HM, and 82.72% (67 / 81) with CM. As shown in Figures 10D and E, compared to CM, HM exhibited a higher proportion of important regions, such as the 3'UTR, exons, and introns, that were more susceptible to the influence of cfSNPs in HM.
[0125] 2.4 Quantification of the local stability of important RNA structural variable regions based on RNA subunits:
[0126] To further evaluate the changes in RNA stability and accessibility caused by cfSNPs in important RNA structural variable regions, we revealed changes based on the single-stranded or double-stranded state of RNA subunits. First, we designed a computational pipeline using Sfold to identify the most probable structures of 81 important RNA structural variable regions. Here, we obtained experimentally determined RNA secondary structures from RNASTRAND and used these structures as references to evaluate the accuracy of the predicted structures. For example, Sfold and RNASTRAND showed high consistency in predicting the RNA secondary structure of stRNA, achieving an RNAsmc score of 10 (Fig. 11A). Then, we obtained the basic RNA subunits of the RNA structural variable regions: stem (S), inner loop (I), convex loop (B), multi-branched loop (M), hairpin loop (H), and outer loop (E) (Fig. 11B). The single-stranded and double-stranded states of RNA folds are closely related to RNA stability or RNA binding accessibility. By analyzing 81 paired Ref and Alt RNA subunits induced by 35 cfSNPs in significantly variable RNA structural regions, we found that 53.09% (43 / 81) of these pairs showed changed pairing states. Of these, 24.69% (20 / 81) changed from single-stranded to double-stranded, 28.40% (23 / 81) changed from double-stranded to single-stranded, and 46.91% (38 / 81) remained unchanged (Figure 11C). In the cases of these changed subunits, the double-stranded state (S represents the double-stranded state) and other single-stranded subunits (B, E, H, I, and M are single-stranded states) underwent significant changes.
[0127] Finally, to identify localized changes in significant RNA structural variable regions induced by cfSNPs at the posttranscriptional level, we performed a comprehensive comparison based on RNA subunits, including the number, length, and base composition of each subunit between Ref and Alt structures. Structural differences in computational analysis were quantified using Euclidean distance. The threshold for assessing significant changes in RNA subunits was defined as the lower quartile (Euclidean distance) of the observed structural changes. R,A =1.65). Here, 8.79% (48 / 546) of important RNA structural variable regions induced by 14.34% (35 / 244) of SNPs may undergo changes in RNA subunits, thereby affecting RNA secondary structure. These results suggest that myopia-related SNPs play an important role in the development of myopia-related diseases by affecting RNA stability.
[0128] This study establishes for the first time a comprehensive map of molecular dysregulation of genomic regulatory elements mediated by myopia-related SNPs. 636 one-to-one pairings were performed on 263 SNPs, 10 genomic regulatory elements, and 5 genomic regulatory elements. A total of 82 myopia-related cfSNPs were identified across the entire genome, including 39 cfSNPs with 39 pairs detected during transcription and 49 cfSNPs with 81 pairs detected during transcription. Transcription was performed using FC and P... T We quantified the gain or loss of SNP-mediated binding proteins. At the post-transcriptional level, we further investigated important variable regions of RNA structure from a global perspective using RNAsnp. Furthermore, we devised a novel method to quantify changes in RNA subunits from a local perspective, reflecting RNA accessibility and stability. Additionally, we obtained genotypic and phenotypic information from a previously established high myopia cohort and assessed the harmfulness of SNPs based on CADD. In summary, this study reveals potential molecular regulation, enhances the interpretability of SNPs, and provides new insights into the genetic mechanisms of myopia.
[0129] During transcription, results showed that, in computational analysis, myopia-associated SNPs may gain or lose T binding sites. In OCR, rs8110889 mapped to ENSR00001023878 weakened the interaction between FOXA1 and ENSR00001023878 (PT = 6.28e-14, FC = 1.38, Fig. 9B). Previous studies have shown that FoxA1 is closely associated with signaling pathways in the vertebrate retina, suggesting that rs8110889 may alter molecular binding, disrupt retinal-related signaling pathways, and lead to myopia susceptibility. In the promoter region, the alt allele (C) of rs7550232 in ENSR00000020131 increased the number of binding motifs on FLI1 (PT = 1.86e-26, FC = 0.75, Fig. 9C). Another previous study reported that fli1 can drive vascular endothelial gene expression and control eye development in zebrafish. These results suggest that the gain and loss of SNP-induced regulatory proteins in the transcriptome may affect gene function and contribute to the pathogenesis of myopia. Recent studies have also reported similar relationships. For example, the T allele of rs17079281 in the DCBLD1 promoter can create a YY1 binding site to suppress DCBLD1 gene expression, reducing the risk of lung cancer in the Chinese population. Furthermore, rs3101339 may disrupt the binding between the TF and NEGR1 genes, leading to gene dysregulation and contributing to major depressive disorder. The discovery of SNP-induced loss or gain of protein motifs can reveal molecular interactions and provide potential intervention targets.
[0130] During post-transcriptional processes, we discovered that myopia-associated SNPs may also disrupt the secondary structure of genomic regulatory elements and lead to molecular dysregulation. Previous studies reported a classic myopia-associated risk molecule, FGF10, whose secondary structure was disrupted by rs339501 (P...). PT =0.11, Euc R,A =1.99). We then obtained the most probable structures of local regions, with the C allele located in the inner loop and the U allele in the stem. The results indicate that rs339501 can alter the single- or double-stranded nature of FGF10 and affect RNA stability. Similarly, rs905224 located on the 3'UTR may disrupt the structural stability of ZNF891. We then downloaded two binding proteins of ZNF891, GAPDH and PSME3, from the IntAct database. By examining the molecular interactions induced by rs905224, we found that changes in the RNA secondary structure of ZNF891 may affect the binding of these two proteins in the 3D spatial conformation and result in different docking fractions. Likewise, numerous previous studies have shown that SNP-mediated changes in RNA secondary structure may contribute to disease development. For example, the two SNPs, SNP22G and A56U, in the 5'UTR of the FTL gene have been shown to alter the overall mRNA structure and are associated with ferritinemia-cataract syndrome. Another study showed that the rs27770 allele in the 3'UTR exhibits a distinct minimum free energy (MFE) structure, which can significantly affect mRNA stability and increase cancer risk. These observations support the role of RNA secondary structure of genomic regulatory elements as a key factor in myopia. We developed a computational pipeline to identify cfSNPs during transcription and post-transcriptional processes. This is significant for understanding the potential regulatory mechanisms of myopia development.
[0131] In summary, we have achieved precise localization for the first time and comprehensively established the relationship between SNPs, genomic regulatory elements, and myopia. Based on our self-designed molecular binding and classic structural heterogeneity assessment algorithms, we identified a series of cfSNPs that can disrupt the regulatory function of genomic regulatory elements at both the transcriptional and post-transcriptional levels. Furthermore, we developed an RNA subunit heterogeneity assessment algorithm to further evaluate the impact of cfSNPs on RNA structural accessibility and stability. These results provide a broad perspective for future research on molecular regulation.
[0132] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0133] In general, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of embodiments of this disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0134] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0135] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0136] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0137] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0138] The exemplary embodiments of this disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art will understand that various modifications and combinations can be made to these embodiments or their features without departing from the principles and spirit of this disclosure, and such modifications should fall within the scope of this disclosure.
Claims
1. A method of predicting an RNA variable region structure, characterized by, The method comprises: 101, obtaining a target RNA to be predicted; 102, obtaining a set of RNA secondary structures based on the target RNA; 103, performing cluster analysis on the set of RNA secondary structures to obtain at least two clusters, and selecting a centroid structure and an energy value of each cluster; the number of centroid structures is consistent with the number of clusters; 104, comparing the energy values of the centroid structures, and selecting a centroid structure with the lowest energy value as the predicted RNA variable region structure.
2. The method of predicting RNA variable region structure according to claim 1, wherein, The method of selecting a centroid structure with the lowest energy value comprises: performing pairwise comparison or traversal comparison on the centroid structures in each cluster to determine the centroid structure with the lowest free energy as the optimal structure.
3. The method of predicting RNA variable region structure according to claim 1, wherein, The energy value is based on the recognized thermodynamic stability principle in structure prediction to evaluate the free energy of the overall structure; the thermodynamic stability principle includes any one or several structure combinations: base pair energy, unpaired region and loop structure.
4. The method of predicting RNA variable region structure according to claim 1, wherein, The method of selecting a centroid structure of each cluster comprises: based on the statistical Boltzmann sampling method, obtaining the structure with the highest frequency in each cluster as the representative structure of each cluster, i.e. the centroid structure.
5. A method of identifying RNA structural changes based on SNPs, characterized in that, The method comprises: 201, obtaining SNPs to be tested; 202, based on the SNPs to be tested, calling Ref and Alt sequence pairs, the sequence corresponding to the target RNA; 203, extracting the RNA secondary structure features of the target RNA structure based on the RNA variable region structure prediction method of any one of claims 1-4; the secondary structure features include B, E, H, I, M, S subunits, and the number, length and position of each subunit; 204, using Euclidean distance to calculate the structure difference of B, E, H, I, M, S subunits, and the number, length, and position of each subunit to obtain the difference value of Ref sequence and Alt sequence; 205, based on the difference value, output the influence of SNPs on RNA structure change.
6. The method of claim 5, wherein the SNP-based identification of RNA structural changes is performed by the method of any one of claims 1-4. Between 203 and 204, the method further comprises normalizing the number, length and position of each subunit, the method for normalizing being: where N i,j represents any numerical 6x2 vector, Nori∈(Num, len, Loc), j∈(B, E, H, I, M, S) is the normalized numerical value, ranging from 0 to 1; B, E, H, I, M, S represent the bulge loop, the external loop, the hairpin loop, the internal loop, the multi-branch loop, and the stem, respectively; R i∈(Num,Len,Loc) and A i∈(Num,Len,Loc) represent the Ref and Alt RNA secondary structures induced by the SNP, respectively.
7. The method for identifying RNA structural changes based on SNPs according to claim 5, characterized in that, The method of quantizing in 204 includes: M represents each RNA subunit feature, N refers to B, E, H, I, M, S subunit, Nor(R) and Nor(A) represent the normalized value of each RNA subunit feature; Euc R,A The difference value obtained after quantization is represented; B, E, H, I, M, S represent the convex ring, the outer ring, the hairpin ring, the inner ring, the multi-branch ring, and the stem, respectively.
8. A computer device, comprising: The device comprises: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to realize the steps of the method of any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and the computer program is executed by a processor to realize the steps of the method of any one of claims 1-7.
10. A computer program product comprising a computer program, characterized in that, The computer program is executed by a processor to realize the steps of the method of any one of claims1-7.