A method, device, medium and program product for predicting RNA variable region structure
Through clustering analysis and optimization of the principle of minimum free energy, the center of mass structure with the lowest energy value of RNA is predicted, which solves the problem of inaccurate prediction of RNA structure in the prior art, and achieves high-accurate prediction of RNA structure, which supports molecular biology research and the development of RNA therapy.
Patent Information
- Application Number
- CN202510096489.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The prior art is difficult to effectively predict the true structure of RNA, which has affected molecular biology research and the development of RNA-based therapeutic methods.
Through cluster analysis of the RNA secondary structure set, the center of mass structure with the lowest energy value is selected as the predicted RNA variable region structure, and optimized in combination with the principle of minimum free energy to obtain the RNA structure closest to the real structure.
Accurate prediction of RNA structure is achieved, the secondary structure prediction capability of RNA consistent with the true folding state is improved, and it supports molecular biology research and the development of RNA therapy methods.
Smart Images

Figure CN119541629B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biological analysis, and more specifically, to a method, device, medium and program product for predicting RNA variable region structure. Background Art
[0002] RNA structure plays a vital role in many biological processes such as gene regulation and signal transduction. Therefore, determining the structure and function relationship of RNA is necessary and a major challenge to better understand the mechanism of biological processes. The nucleotides in RNA molecules are arranged in different orders to form RNA sequences, which is the primary structure of RNA; RNA molecules have many single-stranded region structures, stem-loop structures, double-stranded structures and other planar structures formed by various components composed of complementary base pairs, and these structures are used to perform self-folding movements, and the structure formed is the RNA secondary structure (RSS); the tertiary structure of RNA molecules is a high-level construction in the form of three-dimensional space. This three-dimensional structure is based on the RNA secondary structure. In addition to the interaction force generated by base pairing, there are also main chain-to-main chain interaction forces, main chain-to-base interaction forces, and isolated hydrogen bond interaction forces within the RNA molecule. These interaction forces cause the planar RNA secondary structure to fold into a compact spatial structure. RNA secondary structure motifs are basic components for studying the mechanism of structural biology.
[0003] The structure of RNA is crucial to its function and biological significance. The structural diversity of RNA determines that it can perform a variety of biological functions in cells. The structural study of RNA not only helps us understand its biological function, but also helps to develop RNA-based treatments, such as RNA interference (RNAi) and RNA-based vaccines. Therefore, predicting the true RNA structure has vital research and application value in molecular biology, disease research and other fields. Summary of the invention
[0004] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention provides a method, device, medium and program product for predicting RNA variable region structure and a method, device, medium and program product for identifying RNA structural changes based on SNPs; the method of the present invention achieves optimization in the step of selecting the optimal structure, and based on the hypothesis that the minimum free energy structure is the most stable, selects the minimum free energy structure among many suboptimal structures as the optimal structure, thereby obtaining the RNA structure closest to the real structure.
[0005] The first aspect of the present application discloses a method for predicting RNA variable region structure, the method comprising:
[0006] 101, obtaining a target RNA to be predicted;
[0007] 102, obtaining an RNA secondary structure set based on the target RNA;
[0008] 103. Perform cluster analysis on the RNA secondary structure set to obtain at least 2 clusters, and select the centroid structure and energy value of each cluster; the number of centroid structures is consistent with the number of clusters;
[0009] 104 , comparing the energy values of the centroid structures, and selecting the centroid structure with the lowest energy value as the predicted RNA variable region structure.
[0010] In some embodiments, the method of selecting the center of mass structure with the lowest energy value includes: performing a pairwise cyclic comparison or ergodic comparison on the center of mass structures in each cluster, and determining the center of mass structure with the lowest minimum free energy as the optimal structure.
[0011] In some embodiments, the energy value is a comprehensive assessment of the free energy of the overall structure based on the recognized thermodynamic stability principles in structure prediction; the thermodynamic stability principles include any one or a combination of the following structures: base pair energy, unpaired region, and loop structure.
[0012] In some embodiments, the method of selecting the centroid structure of each cluster includes: based on the statistical Boltzmann sampling method, obtaining the structure with the highest occurrence frequency in each cluster as the representative structure of each cluster, that is, the centroid structure.
[0013] The second aspect of the present application discloses a method for identifying RNA structural changes based on SNPs, the method comprising:
[0014] 201, obtaining SNPs to be tested;
[0015] 202, based on the SNPs to be detected, retrieve Ref and Alt sequence pairs, wherein the sequences correspond to target RNA;
[0016] 203. Extract RNA secondary structure features of the target RNA structure based on the RNA variable region structure prediction method described in the first aspect of the present application; the secondary structure features include B, E, H, I, M, S subunits, and the number, length and position of each subunit;
[0017] 204, the Euclidean distance was used to calculate the B, E, H, I, M, and S subunits, and the number, length, and position of each subunit were used to quantify the structural differences between the Ref sequence and the Alt sequence to obtain the difference value;
[0018] 205, output the effect of SNPs on RNA structure changes based on the difference value.
[0019] In some embodiments, between 203 and 204, the method further comprises standardizing the number, length and position of each subunit, wherein the standardization method is:
[0020]
[0021] in, represents any numerical 6 × 2 vector, is a normalized value ranging from 0 to 1; B, E, H, I, M, and S represent convex loop, outer loop, hairpin loop, inner loop, multi-branched loop, and stem, respectively; and Represent the ref and alt RNA secondary structures induced by SNP, respectively.
[0022] In some embodiments, the quantification method in 204 includes:
[0023]
[0024] M represents the characteristics of each RNA subunit, N refers to B, E, H, I, M, S subunits, and Respectively represent the normalized values of each RNA subunit feature; represents the difference value obtained after quantification; B, E, H, I, M, and S represent convex loop, outer loop, hairpin loop, inner loop, multi-branched loop, and stem, respectively.
[0025] The third aspect of the present application discloses a computer device, which includes: a memory and a processor; the memory is used to store a computer program; and the processor executes the computer program to implement the steps of the above method.
[0026] A fourth aspect of the present application discloses a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of the above method when executed by a processor.
[0027] A fifth aspect of the present application discloses a computer program product, including a computer program, which implements the steps of the above method when executed by a processor.
[0028] This application has the following beneficial effects:
[0029] 1. The present application innovatively discloses a method for predicting RNA variable region structure, which evaluates the changes of RNA subunits in the region of significant global RNA structural variation caused by SNP. First, the RNA secondary structure that is as consistent as possible with the true folding state is predicted. Cluster analysis is performed based on 1000 possible structures to obtain multiple clusters. By comparing the minimum free energy (MFE) of the RNA structure in each cluster, the structure with the lowest free energy is selected as the most likely structure. In the step of selecting the optimal structure, optimization is achieved. Based on the hypothesis that the minimum free energy structure is the most stable, the minimum free energy structure among many suboptimal structures is selected as the optimal structure, that is, the centroid structures are compared cyclically in pairs, and the minimum free energy structure is finally determined to be the optimal structure. Based on the RNASTRAND public database to obtain the experimentally confirmed stRNA structure, we compared the correlation between the selected minimum free energy structure and the experimentally confirmed structure, and confirmed that the two structures are completely consistent through the structural comparison algorithm.
[0030] 2. This application mainly solves the problem of how to obtain the optimal structure. The software (Sfold) will perform statistical sampling for 1,000 structures, give the center of mass structure, and the energy value of the center of mass structure. What we do is to extract the energy value of the center of mass structure, compare the free energy of the structure, and select the minimum free energy as the optimal structure. The minimum free energy here is based on the recognized principle of thermodynamic stability in structure prediction, taking into account the free energy of various structural combinations such as base pair energy, unpaired regions and ring structures, comprehensively evaluating the free energy of the overall structure, and obtaining the minimum free energy structure based on the hypothesis that the thermodynamically stable state usually has the lowest free energy.
[0031] 3. When screening the optimal structure, i.e. the final RNA structure, the Sfold algorithm does not clearly indicate which type of structure set’s representative structure is the optimal structure. This case is mainly based on the statistical Boltzmann sampling method to obtain the cluster centroid structure with the highest frequency in each cluster as the representative structure set of each cluster, and further screen to obtain the optimal RNA structure.
[0032] 4. This application innovatively discloses a method for identifying RNA structural changes at the post-transcriptional level of the genome. The method retrieves Ref and Alt sequence pairs based on SNPs, extracts target RNA secondary structural features based on the sequence pairs, quantifies the secondary structures in the Ref sequence and the Alt sequence using Euclidean distance to obtain difference values, identifies the effects of SNPs on RNA structural changes based on the difference values, and explores the interference of post-transcriptional regulation;
[0033] 5. This application innovatively develops a process to comprehensively evaluate the global and local effects of SNPs on variable regions of RNA structure, providing an effective method for evaluating the accessibility and stability of RNA. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0035] Figure 1 is a schematic diagram of a method flow chart provided by the first aspect of an embodiment of the present invention;
[0036] Figure 2 is a schematic diagram of a method flow chart provided by the second aspect of an embodiment of the present invention; Figure 3 Schematic diagram of a prediction system for RNA variable region structure provided by an embodiment of the present invention;
[0037] Figure 4 Schematic diagram of a recognition system for recognizing RNA structural changes based on SNPs provided by an embodiment of the present invention; Figure 5 is a schematic diagram of a computer device provided by an embodiment of the present invention; Figure 6 is a schematic diagram of the architecture of an exemplary computing device provided by an embodiment of the present invention; Figure 7 is a schematic diagram of a storage medium provided by an embodiment of the present invention; Figure 8 is a fine-scale mapping map of myopia-related SNPs on genome-wide regulatory elements provided by an embodiment of the present invention; wherein, Figure 8 A is the distribution of each genomic regulatory element in transcription and post-transcriptional processes; Figure 8 B is the percentage of relationship pairs in transcription and posttranscriptional processes; Figure 8 C is the length of the genomic regulatory elements covering the myopia-associated SNPs, and the black-marked horizontal line in each violin represents the median length of each element; Figure 8 D is the SNP density of each regulatory element at the transcriptional and post-transcriptional levels. The dotted line shows the average SNP density of the entire human genome; Fig. 9 is a schematic diagram of the heterogeneity of TF binding mediated by SNPs between four regulatory elements in the transcription process provided by an embodiment of the present invention; wherein, Fig. 9 A is the heterogeneity of binding between TF and enhancer, Fig. 9 B is the heterogeneity of the binding between TF and OCR, Fig. 9 C is the heterogeneity of binding between TF and promoter, Fig. 9 D is the heterogeneity of the binding between TF and PFR; P T Values were calculated using Fisher’s exact test, and FC was calculated by fold change;
[0038] Fig.10It is a global RNA structure variable region mediated by SNP on genomic regulatory elements in the post-transcriptional process provided by an embodiment of the present invention; wherein, Fig.10 A is the length of the variable region of RNA structure in the genomic regulatory element; Fig.10 B is the percentage of relationship pairs showing proximal and distal effects mediated by SNPs on genomic regulatory elements; Fig.10 C is the percentage of relationship pairs with important RNA structural variable regions on genomic regulatory elements; Fig.10 D is the percentage of significant RNA structure variable region relationship pairs covered in CM; Fig.10 E is the percentage of relationship pairs covering significant RNA structural variable regions in HM;
[0039] Fig.11 The present invention provides an important RNA structure variable region analysis based on RNA subunits; wherein, Fig.11 A is the RNA secondary structure of selenocysteine transfer RNA (stRNA) obtained by Sfold, and the structural conformation of stRNA predicted by Sfold (left) and downloaded by RNASTRAND (right). Fig.11 B is the visualization of the six subunits of RNA secondary structure. Fig.11 C is the change in the single-stranded or double-stranded state of the RNA in the important RNA structural variable region at the SNP site based on the RNA subunit. Fig.11 D is the detection of cfSNPs that can affect RNA subunits based on Euclidean distance. The threshold (1.65) is represented by the horizontal line. DETAILED DESCRIPTION
[0040] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0041] In some of the processes described in the specification and claims of the present invention and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent the order of precedence, and do not limit the "first" and "second" to be different types.
[0042] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0043] Figure 1 1 is a schematic flow chart of a method for predicting RNA variable region structure provided by an embodiment of the present invention. Specifically, the method comprises the following steps: 101: obtaining a target RNA to be predicted;
[0044] RNA is the abbreviation of ribonucleic acid, a biological macromolecule closely related to DNA. They are both composed of nucleotides, but they are different in structure and function. RNA plays multiple roles in cells, including the transmission of genetic information, protein synthesis and the regulation of gene expression.
[0045] The main types of RNA include: Messenger RNA (mRNA): It carries the genetic information on DNA and guides the synthesis of proteins. mRNA is synthesized from a DNA template through the process of transcription, then transported to the cytoplasm and translated into proteins on the ribosomes. Ribosomal RNA (rRNA): It is an important component of the ribosome, which is the site of protein synthesis. rRNA, together with mRNA and transfer RNA (tRNA), participates in the protein synthesis process. Transfer RNA (tRNA): It is responsible for transporting amino acids to the ribosome and connecting amino acids according to the codon sequence on the mRNA to form a protein chain. MicroRNA (miRNA): It is a class of small molecule RNA that mainly regulates gene expression after transcription. It regulates protein synthesis by binding to mRNA, affecting its stability and translation efficiency. Long non-coding RNA (lncRNA): They are a class of longer RNA molecules that do not encode proteins, but play an important role in the regulation of gene expression, including processes such as chromatin remodeling, transcription, and splicing.
[0046] 102: obtaining an RNA secondary structure set based on the target RNA;
[0047] In some embodiments, the software for predicting RNA secondary structure sets includes any one or more of the following: RNAfold, Sfold, CONTRAfold, and the software is preferably Sfold.
[0048] 103: Perform cluster analysis on the RNA secondary structure set to obtain at least 2 clusters, select the centroid structure and its energy value of each cluster; the number of centroid structures is consistent with the number of clusters; the centroid structure is Fig.11 Cluster 1 centroid, Cluster 2 centroid and Ensemble centroid in A, where centroid can be translated as center of mass, and the center of mass structure of each cluster is marked with a red dot in the figure; in some embodiments, the method for selecting the center of mass structure of each cluster includes: based on the statistical Boltzmann sampling method, obtaining the structure with the highest occurrence frequency in each cluster as the representative structure of each cluster, that is, the center of mass structure.
[0049] 104: Compare the energy values of the centroid structures, and select the centroid structure with the lowest energy value as the predicted RNA variable region structure.
[0050] In some embodiments, the method of selecting the center of mass structure with the lowest energy value includes: performing a pairwise cyclic comparison or ergodic comparison on the center of mass structures in each cluster, and determining the center of mass structure with the lowest minimum free energy as the optimal structure.
[0051] In some embodiments, the energy value is a comprehensive evaluation of the free energy of the overall structure based on the recognized thermodynamic stability principle in structure prediction; the thermodynamic stability principle includes any one or a combination of the following structures: base pair energy, unpaired region and loop structure. The key point is to define the optimal structure based on the cluster centroid structure set given by Sfold using the minimum free energy principle.
[0052] like Figure 2 As shown, the second aspect of the present application discloses a method for identifying RNA structural changes based on SNPs, the method comprising:
[0053] 201, obtaining SNPs to be tested;
[0054] In some embodiments, the full name of SNP is Single Nucleotide Polymorphism, which refers to the variation of a single nucleotide (base pair) in a genomic DNA sequence. This variation may be a substitution (e.g., A becomes G), an insertion or a deletion. SNP is the most common form of genetic variation in the genome, and their distribution in the population is polymorphic, that is, different individuals may have different nucleotides at the same position. In some embodiments, the SNP to be tested comes from a subject, and the terms "subject" or "person to be tested" or "sample to be tested" used herein refer to any animal (e.g., mammal), including but not limited to humans, non-human primates, rodents, etc., which will become recipients of specific treatments. Typically, the terms "subject" and "patient" are used interchangeably herein when referring to human subjects. Preferably, the subject is a human.
[0055] 202, based on the SNPs to be detected, retrieve Ref and Alt sequence pairs, wherein the sequences correspond to target RNA;
[0056] In some embodiments, based on each SNP, the Ref and Alt sequence pairs corresponding to the SNP are retrieved. Specifically, the Ref and Alt sequences are retrieved each Mbp upstream and downstream to obtain several sequence pairs, where M is a natural number greater than 1; the range of M is 10-35, preferably 20. Specifically, Ref is a wild type, and Alt is a mutant. The number refers to a natural number greater than 1.
[0057] 203. Extract RNA secondary structure features of the target RNA structure based on the RNA variable region structure prediction method described in the first aspect of the present application; the secondary structure features include B, E, H, I, M, S subunits, and the number, length and position of each subunit;
[0058] 204, the Euclidean distance was used to calculate the B, E, H, I, M, and S subunits, and the number, length, and position of each subunit were used to quantify the structural differences between the Ref sequence and the Alt sequence to obtain the difference value;
[0059] 205, output the effect of SNPs on RNA structure changes based on the difference value.
[0060] In some embodiments, between 203 and 204, the method further comprises standardizing the number, length and position of each subunit, wherein the standardization method is:
[0061]
[0062] in, represents any numerical 6 × 2 vector, is a normalized value ranging from 0 to 1; B, E, H, I, M, and S represent convex loop, outer loop, hairpin loop, inner loop, multi-branched loop, and stem, respectively; and Represent the ref and alt RNA secondary structures induced by SNP, respectively.
[0063] In some embodiments, the quantification method in 204 includes:
[0064]
[0065] M represents the characteristics of each RNA subunit, N refers to B, E, H, I, M, S subunits, and Respectively represent the normalized values of each RNA subunit feature; represents the difference value obtained after quantification; B, E, H, I, M, and S represent convex loop, outer loop, hairpin loop, inner loop, multi-branched loop, and stem, respectively.
[0066] In some embodiments, if the difference value exceeds a threshold, a result indicating that the SNP has an effect on RNA structural change is outputted. The threshold is determined by evaluating the lower quartile of significant changes in RNA subunits.
[0067] In some embodiments, the post-transcriptional genomic regulatory element comprises any one or more of the following: 3'UTR, 5'UTR, exon, intron, long non-coding RNA.
[0068] In some embodiments, the method further includes: identifying the RNA structure variable region of the post-transcriptional gene regulatory element in the Ref and Alt sequence pairs; extracting RNA secondary structure features from the RNA structure variable region; the RNA structure variable region includes a global RNA structure variable region and / or a local RNA structure variable region.
[0069] In some embodiments, the SNPs in the global RNA structure variable region are classified into proximal allosteric effects and distal allosteric effects; if the SNPs are located in the global RNA structure variable region and affect the folding of the region, the output belongs to the proximal allosteric effect; if the SNPs are not located in the global RNA structure variable region and affect the folding of the region, the output belongs to the distal allosteric effect. Specifically, the proximal allosteric effect refers to the effect on the surrounding small segment area, and only the current gene or other of the SNP is considered when doing molecular experiments; the distal allosteric effect refers to the effect not on the current gene, but may be on other genes, and other genes are considered first when doing molecular experiments.
[0070] In some embodiments, the evaluation results or prediction results include but are not limited to paper or electronic report forms. The results are only obtained by the intelligent machine based on the analysis of relevant data of the subjects, and are only used as a reference, not as the final diagnosis result.
[0071] In this embodiment, the threshold is obtained through training of training set samples, and it can be a specific threshold or an interval range. The specific form is not specifically limited in this embodiment.
[0072] Figure 5 is a schematic diagram of a computer device provided by an embodiment of the present invention, such as Figure 5 As shown, the device 2000 may include: one or more processors 2010, and one or more memories 2020; wherein the memories store computer-readable codes, and when the computer-readable codes are run by the one or more processors, the method described above may be executed.
[0073] The processor in this embodiment can be an integrated circuit chip with signal processing capabilities. The above processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The disclosed methods, operations and logic block diagrams in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc., which can be an X86 architecture or an ARM architecture.
[0074] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the disclosed embodiments are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general purpose hardware or controllers or other computing devices, or some combination thereof as non-limiting examples.
[0075] For example, the method or device according to the embodiment of the present disclosure may also be implemented by Figure 6 The architecture of the computing device 3000 shown in FIG. Figure 6 As shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as ROM 3030 or hard disk 3070, may store various data or files used for processing and / or communication of the method provided by the present disclosure and program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 6 The architecture shown is only exemplary and can be omitted according to actual needs when implementing different devices. Figure 6 One or more components of a computing device are shown.
[0076] The embodiment of the present invention also provides a computer-readable storage medium, such as Figure 7As shown, it is a schematic diagram of a storage medium 4000 provided in an embodiment of the present invention, and a computer readable instruction 4010 is stored on the computer storage medium 4020. When the computer readable instruction 4010 is executed by a processor, the method according to the embodiment of the present disclosure described with reference to the above figures can be executed. The computer readable storage medium in the embodiment of the present disclosure can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory. It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0077] The embodiments of the present disclosure further provide a computer program product or system, including a computer program, which implements the steps of the above method when executed by a processor.
[0078] In some embodiments, this embodiment also discloses a prediction system for RNA variable region structure, such as Figure 3 As shown, the system comprises:
[0079] RNA acquisition module 301, used for or configured to acquire target RNA to be predicted;
[0080] RNA secondary structure extraction module 302, used for or configured to obtain an RNA secondary structure set based on the target RNA;
[0081] The centroid structure extraction module 303 is used or configured to perform cluster analysis on the RNA secondary structure set to obtain at least 2 clusters, and select the centroid structure and energy value of each cluster; the number of centroid structures is consistent with the number of clusters;
[0082] The RNA structure prediction module 304 is used for or configured to compare the energy values of the centroid structures and select the centroid structure with the lowest energy value as the predicted RNA structure.
[0083] In some embodiments, this embodiment also discloses a recognition system based on SNPs to identify RNA structural changes, such as Figure 4 As shown, the system comprises:
[0084] SNP acquisition module 401, used for or configured to acquire SNPs to be tested;
[0085] A sequence pair extraction module 402 is used or configured to retrieve Ref and Alt sequence pairs based on the SNPs to be tested, wherein the sequences correspond to the target RNA;
[0086] RNA secondary structure feature extraction module 403, used or configured to extract RNA secondary structure features of the target RNA structure based on the RNA variable region structure prediction method described in the first aspect of the present application; the secondary structure features include B, E, H, I, M, S subunits, and the number, length and position of each subunit;
[0087] A difference value calculation module 404 is used or configured to use Euclidean distance to calculate the number, length and position of B, E, H, I, M, S subunits, and each subunit to quantify the structural difference between the Ref sequence and the Alt sequence to obtain a difference value;
[0088] The RNA structure change impact module 405 is used for or configured to output the impact results of SNPs on RNA structure changes based on the difference values. Specific embodiment:
[0090] 1 Materials and Methods
[0091] 1.1 Collecting Data:
[0092] Myopia-associated SNPs were collected from dbGap, GWAS catalogs, and published literature. Based on the population distribution of myopia-associated SNPs in the 1000 Genomes Project (GRCh38), we performed quality control (QC) on samples and genotypes. SNPs were filtered according to the following selection criteria: East Asian, two alleles in NCBI, minor allele frequency (MAF)>5%, Hardy-Weinberg equilibrium P<0.01, recall rate>75%, and genotyping rate>75%. 343 SNPs were obtained for the following analysis. The start and end positions of eight genomic regulatory elements, including open chromatin regions (OCRs), CTCF binding sites (CTCFBS), enhancers, promoters, promoter flanking regions (PFRs), exons, introns, and noncoding RNA (ncRNA) transcripts, were derived from ENSEMBL (v102), and their sequences were extracted according to Refseq from NCBI. The sequences of 3'UTR and 5'UTR were obtained using the BioMart tool from ENSEMBL.
[0093] 1.2 Fine positioning of myopia-related SNPs on 10 genomic regulatory elements:
[0094] Fine mapping was performed using the classic read alignment tool Bowtie2. The primary sequences of genomic regulatory elements were considered as long reference reads. The 30bp flanking regions upstream and downstream of these SNPs were considered as short alignment sequences. Then, strict parameters were set using "--n-ceil C,3--np 0--end-to-end -a --score-minC,0" to avoid mismatches for each seed. After fine positioning, we constructed paired reference (Ref) and alternative (Alt) sequences based on the alleles of the SNP. The alignment software includes Bowtie1, Bowtie2, BLAST, preferably Bowtie2.
[0095] 1.3 Assessing changes in molecular binding at the transcriptional level:
[0096] We extracted the 20 bp upstream and downstream of the paired Ref and Alt sequences induced by the SNP, and obtained the binding motifs of the promoters, open chromatin regions, promoters, and promoter flanking regions of these sequences from the HumanTFDB database, while CTCF was from the CTCFBSDB database. For protein binding, we obtained the number of regulatory proteins (TF and CTCF), protein symbols, and the start and end positions of the binding motifs, respectively. First, we used fold change (FC) to obtain the change in the number of protein binding sites mediated by myopia-associated SNPs:
[0097]
[0098] and Represents the number of proteins bound by the Ref and Alt sequences, respectively. We defined FC>1 as a loss of protein binding, and conversely, FC<1 as a gain of protein binding. We then applied Fisher's exact test to determine changes in protein binding motifs as follows:
[0099]
[0100] a is the number of binding motifs that overlap only with the Ref sequence, b is the number of binding motifs that do not overlap with the Ref sequence, c is the number of binding motifs that overlap only with the Alt sequence, and d is the number of binding motifs that do not overlap with the Alt sequence. T <0.05, indicating that the myopia-associated SNPs significantly affected protein binding, and these SNPs were identified as cfSNPs. In addition, we defined the proteins with the highest scores as “leader” proteins.
[0101] 1.4 Identification of variable regions of global RNA structure after transcription:
[0102] At the post-transcriptional level, we used RNAsnp to identify structurally variable regions of five regulatory elements, including 3'UTR, 5'UTR, exons, introns, and lncRNA transcripts. PT <0.2 indicates that the SNP may have a significant effect on the local RNA structure. Here, we define these regions as "global RNA structure variable regions". If the SNP is located in the RNA structure variable region and affects the folding of this region, these SNPs are defined as "proximal allosteric effects". Otherwise, if the SNP is not located in the RNA structure variable region and affects the surrounding folding, these SNPs are defined as "distal allosteric effects".
[0103] 1.5 Using RNA subunits to quantify changes in RNA variable regions:
[0104] We evaluated SNP-induced changes in RNA subunits within important global RNA structural variable regions. First, we predicted RNA secondary structures that were as consistent as possible with the true folding state. Cluster analysis was performed based on 1,000 possible structures to obtain representative structures of multiple clusters. By comparing the minimum free energy (MFE) of the RNA structures in each cluster, we selected the structure with the lowest free energy as the most likely structure. Then, RNAsmc was used to extract the six RNA subunits of each RNA, namely the bulge loop (B), outer loop (E), hairpin loop (H), inner loop (I), multi-branched loop (M) and stem (S) of each RNA. Next, we developed a computational pipeline and quantified the structural heterogeneity of the variable regions of the global RNA structure using Euclidean distance (based on RNA subunits) ( ). We obtained the characteristics of RNA subunits in the variable region of the global RNA structure, namely the number, length and position of each subunit in the paired Ref and Alt structures. For these three dimensions, we constructed two corresponding 6×1 vectors:
[0105]
[0106]
[0107] Among them, i represents the characteristics of RNA subunits. B, E, H, I, M and S refer to the six RNA subunits, namely the convex loop, outer loop, hairpin loop, inner loop, multi-branched loop and stem. R and A represent the Ref and AltRNA secondary structures mediated by SNP, respectively. Then, in order to standardize the size of the three characteristics of RNA subunits, each type of subunit is standardized as follows:
[0108]
[0109] represents any numerical 6×2 vector, ranges from 0 to 1. Finally, we use the Euclidean distance ( ) to quantify the difference between the Ref and Alt structures:
[0110]
[0111] M represents the characteristics of each subunit. N refers to six subunits. and denotes the standardized value of each feature. The lower quartile of was set as the significance threshold. Above the lower quartile, we considered RNA subunits to show significant changes. VARNA was used to visualize RNA secondary structure.
[0112] 1.6 Evaluate the same annotation of SNPs based on the HM cohort and computational pipeline:
[0113] A total of 10,348 high myopic participants (worst eye SE <-6.00D) in the CAMS study were sequenced on the Illumina NovaSeq6000 sequencer from Berry Genomics using the Twist Human Core Exome Kit. We obtained the genotypes and phenotypes of myopia-associated SNPs from the CAMS study. Next, we used combined annotation-dependent deletion (CADD) to score the harmfulness of SNPs based on the genotypes and phenotypes in the high myopia cohort. Here, SNPs with a CADD score ≥ 10 are defined as harmful SNPs that may lead to loss of gene function. Combined Annotation Dependent Depletion score (CADD score): used to evaluate and quantify single nucleotide variations (SNVs). The higher the score, the more harmful variants there are, that is, the higher the possibility of pathogenicity.
[0114] 2 Results:
[0115] 2.1 Myopia-related SNPs are widely distributed at the transcriptional and post-transcriptional levels:
[0116] We obtained myopia-associated SNPs from public resources and, after quality control, used 343 SNPs for subsequent analysis (see Methods). To reveal the complete map of SNPs at the transcriptional and post-transcriptional levels, we used fine mapping and identified 636 relationship pairs formed by 263 SNPs, 10 genomic regulatory elements, and myopia ( Figure 8A). In addition, these SNPs were associated with five phenotypes, namely refractive error (RE), common myopia (CM), high myopia (HM), pathological myopia (PM) and visual impairment (VD). In the transcriptional process, 84 SNPs were located in enhancers, open chromatin regions (OCRs), CTCF binding sites (CTCFBS), promoters and promoter flanking regions (PFRs), forming 90 relationship pairs. A total of 244 SNPs were located on five genomic regulatory elements, namely 5'UTR, exons, introns, 3'UTR and LncRNA, constituting 546 relationship pairs at the post-transcriptional level. Next, we found that 14.15% (90 / 636) and 85.85% (546 / 636) of the pairs were enriched in the transcriptional and post-transcriptional processes, respectively ( Figure 8 B).
[0117] To further evaluate the average distribution of SNPs in all genomic regulatory elements, we analyzed the transcript lengths of these elements and the density of SNPs in each genomic regulatory element. The results showed that the median length of the longest transcript LncRNA was 95.53 times the median length of the shortest transcript CTCFBS ( Figure 8 C). We then further evaluated the enrichment of SNPs in each regulatory element within every 1000 bp. Compared with the average of 1 SNP per 1000 bp in the human genome, we observed that myopia-associated SNPs were highly distributed in OCR, CTCFBS, 5'UTR, and exons ( Figure 8 D). By analyzing myopia, 81.76% (520 / 636) of the relationship pairs were found to be associated with CM. In addition, little is known about HM, PM, RE, and VD, which are approximately 12.11% (77 / 636), 5.19% (33 / 636), 0.63% (4 / 636), and 0.31% (2 / 636), respectively. These data suggest that the distribution of myopia-associated SNPs differs at the transcriptional and post-transcriptional levels and between genomic regulatory elements. These differences may be related to the severity of myopia and potential molecular regulation.
[0118] 2.2 Scoring SNP-induced molecular binding heterogeneity during transcription:
[0119] To investigate SNP-mediated molecular binding heterogeneity at the transcriptional level, we developed a computational pipeline that used the fold test (FC) to assess changes in the number of bound proteins and applied Fisher’s exact test to assess changes in the number of bound proteins. T<0.05, FC was not equal to 1, and 38.46% (5 / 13), 80% (4 / 5), 57.89% (11 / 19), and 44.73% (17 / 38) of the SNPs were found to be likely to disrupt transcription factor binding at enhancers, OCRs, promoters, and PFRs, respectively ( Fig. 9 ). For CTCF protein, we found that approximately 13.33% (2 / 15) of SNPs could disrupt CTCFBS interactions. In total, 43.33% (39 / 90) of one-to-one pairings were able to disrupt binding affinity, and 46.43% (39 / 84) of SNPs were identified as “cfSNPs” during transcription. In summary, OCR was significantly enriched in myopia-associated SNPs and cfSNPs, showing high-density distribution and molecular binding disruption.
[0120] To determine the possible effects of cfSNP-associated regulatory proteins associated with or potentially associated with myopia, we explored the molecular functions of the “leader” proteins of each relationship pair in published studies. Interestingly, most of the “leader” proteins were able to affect ocular tissues or structures. For example, about 7.69% (3 / 39) of cfSNPs caused changes in the REST binding motif, thereby affecting the fate of RGC retinal ganglion cells (RGCs) in the developing retina. About 7.69% (3 / 39) of cfSNPs caused changes in IRF1 binding, which is known to be expressed in retinal microglia and plays a key role in microglial activation and retinal inflammation. In addition, 15.28% (6 / 39) of cfSNPs changed the binding motif site of SPI1, which has been reported to regulate microglia in the retina. In summary, we found that there are abundant leader proteins in retinal inflammation, which have been shown to be associated with the occurrence and development of myopia. In addition, it can help us reveal potential protein regulators and understand how SNPs are involved in the molecular regulation of myopia. Finally, we evaluated the distribution of myopia types among the significantly changed relationship pairs at the transcriptional level. More than 75% (3 / 4) of them were related to PM, while approximately 44.44% (4 / 9) and 41.56% (32 / 77) were related to HM and CM ( Fig. 9 AD). The results also showed that the effect of SNP on myopia increased with the severity of myopia.
[0121] 2.3 Identification of SNP-mediated global RNA structural variable regions during post-transcriptional processes:
[0122] To identify SNP-induced RNA secondary structural heterogeneity during post-transcriptional processes, we initially employed RNA SNPs to detect potential global RNA structural variable regions, P PT<0.2. Here, since there is only one relative pair in the 5'UTR, this exon was selected as a reference to compare the significance of the average length differences of the variable regions of RNA structures. Fig.10 As shown in A, the RNA structural variable regions of 5'UTR, 3'UTR and Exon are relatively longer than introns and LncRNA. It is worth noting that the SNPs associated with myopia are not only highly enriched in 5'UTR and exons, but also cause huge structural damage to these regulatory regions. Although myopia-related SNPs are not significantly enriched in the 3'UTR region, they still cause significant structural effects. This may be due to the fact that the 3'UTR region presents highly structured characteristics, and once it is damaged, it has a great impact on the RNA secondary structure.
[0123] In addition, we evaluated whether the SNPs were located in the RNA structural variable regions. The definitions are shown in the Methods. Statistically, approximately 97.54% (238 / 244) of the SNPs exhibited proximal allosteric effects, mapping to 91.96% (502 / 546) of the relationship pairs and all five regulatory elements, while only 12.70% (31 / 244) of the SNPs exhibited proximal allosteric effects. The SNPs showed distal allosteric effects, mapping to 8.06% (44 / 546) of the relationship pairs and three regulatory elements, namely, exons, introns, and lncRNAs ( Fig.10 B). This finding is consistent with a previous study showing that SNPs affect mainly local regions rather than having a global effect.
[0124] To further explore which genomic regulatory elements are significantly disrupted by SNPs, we analyzed the proportion of important RNA structural variable regions within these elements. Fig.10As shown in Figure C, among the 9 relationship pairs in 3'UTR, 33.33% (3 / 9) were induced by 37.50% (3 / 8) of SNPs and showed significant RNA structure variable regions, while 1 relationship pair in 5'UTR did not show significant RNA structure variable regions. Among the 35 relationship pairs in exons, 22.86% (8 / 35) showed significant changes, which were mediated by 20.69% (6 / 29) of SNPs. In addition, among the 308 relationship pairs in introns, 14.61% (45 / 308) showed significant RNA structure variable regions, which were affected by 16.52% (38 / 230) SNPs. Among the 193 relationship pairs in LncRNA, 12.95% (25 / 193) showed significant RNA structure variable regions, which were induced by 14.56% (23 / 158) of SNPs. Overall, we found that 14.84% (81 / 546) of relationship pairs contained significant RNA structural variable regions induced by 20.08% (49 / 244) of the SNPs, which were identified as cfSNPs.
[0125] To characterize the role of RNA secondary structure in myopia, we analyzed the distribution of important RNA structural variable regions. Three myopia phenotypes, RE, CM, and HM, showed significant SNP-mediated changes. Among these phenotypes, 1.23% (1 / 81) were one-to-one associated with RE, 16.05% (13 / 81) were associated with HM, and 82.72% (67 / 81) were associated with CM. Fig.10 As shown in D and E, HM exhibited a higher proportion of important regions compared with CM, such as 3'UTR, exons, and introns that were more susceptible to cfSNPs in HM.
[0126] 2.4 Quantifying the local stability of important RNA structural variable regions based on RNA subunits:
[0127] To further evaluate changes in RNA stability and accessibility caused by cfSNPs in important RNA structural variable regions, we revealed changes based on the single-stranded or double-stranded state of the RNA subunit. First, we designed a computational pipeline to identify the most likely structures of 81 important RNA structural variable regions using Sfold. Here, we obtained experimentally determined RNA secondary structures from RNASTRAND and used these structures as references to evaluate the accuracy of the predicted structures. For example, Sfold and RNASTRAND showed high consistency in predicting RNA secondary structures of stRNAs, achieving an RNAsmc score of 10 ( Fig.11A). Then, we get the basic RNA subunits of the variable region of RNA structure, which are stem (S), inner loop (I), bulge loop (B), multi-branched loop (M), hairpin loop (H) and outer loop (E) ( Fig.11 B). The single-stranded and double-stranded states of RNA folding are closely related to RNA stability or accessibility of RNA binding. By analyzing 81 pairs of paired Ref and Alt structures of RNA subunits induced by 35 cfSNPs in the significant RNA structural variable regions, we found that 53.09% (43 / 81) of these pairs showed changes in pairing state. Among them, 24.69% (20 / 81) changed from single-stranded to double-stranded, 28.40% (23 / 81) changed from double-stranded to single-stranded, and 46.91% (38 / 81) remained unchanged ( Fig.11 C). In the case of these altered subunits, the double-stranded state (S represents the double-stranded state) and other single-stranded subunits (B, E, H, I and M are single-stranded states) are significantly changed.
[0128] Finally, to identify local changes in RNA structural variable regions that are significant at the posttranscriptional level due to cfSNPs, we performed a comprehensive comparison based on RNA subunits, including the number, length, and base composition of each subunit between the Ref and Alt structures. The structural differences in the in silico analysis were quantified using Euclidean distance. The threshold for evaluating significant changes in RNA subunits was defined as the lower quartile of the observed structural changes ( =1.65). Here, 8.79% (48 / 546) of the important RNA structural variable regions induced by 14.34% (35 / 244) of the SNPs may undergo changes in RNA subunits, thereby affecting the secondary structure of RNA. These results suggest that myopia-associated SNPs play an important role in the occurrence of myopia-related diseases by affecting RNA stability.
[0129] This study established for the first time a comprehensive map of molecular dysregulation of genomic regulatory elements mediated by myopia-associated SNPs. A total of 636 one-to-one pairings were performed on 263 SNPs, 10 genomic regulatory elements, and 5 genomic regulatory elements. We identified a total of 82 myopia cfSNPs on genome-wide regulatory elements, of which 39 cfSNPs in 39 pairs were detected in the transcription process and 49 cfSNPs in 81 pairs were detected in the transcription process. TSNP-mediated gain or loss of binding proteins was quantified. At the post-transcriptional level, we used RNAsNPs to further investigate important RNA structural variable regions from a global perspective. In addition, we designed a new method to quantify changes in RNA subunits from a local perspective, which can reflect the accessibility and stability of RNA. In addition, we obtained genotype and phenotypic information from a previously established high myopia cohort and assessed the harmfulness of SNPs based on CADD. In summary, this study revealed potential molecular regulation, enhanced the interpretability of SNPs, and provided new insights into the genetic mechanism of myopia.
[0130] In the transcriptional process, the results showed that the SNPs associated with myopia may gain or lose binding sites for T in the computational analysis. In OCR, rs8110889 mapped to ENSR00001023878 could weaken the interaction between FOXA1 and ENSR00001023878 (PT = 6.28e-14, FC = 1.38, Fig. 9 B). Previous studies have shown that FoxA1 is closely related to signaling pathways in the vertebrate retina, suggesting that rs8110889 may alter molecular binding, disrupt retinal-related signaling pathways, and lead to myopia susceptibility. In the promoter region, the alt allele of rs7550232 in ENSR00000020131 (C) can increase the number of binding motifs on FLI1 (PT=1.86e-26, FC=0.75, Fig. 9 C). Another previous study reported that fli1 can drive endothelial gene expression and control eye development in zebrafish. These results suggest that the gain and loss of regulatory proteins induced by SNPs in the transcriptome may affect gene function and contribute to the pathogenic factors of myopia. Recent studies have also reported similar relationship pairs. For example, the allele T of rs17079281 in the DCBLD1 promoter can create a YY1 binding site to inhibit the expression level of the DCBLD1 gene and reduce the risk of lung cancer in the Chinese population. In addition, rs3101339 may disrupt the binding between TF and NEGR1 genes, leading to dysregulated gene expression and causing major depression. The discovery of SNP-induced protein motif loss or gain can reveal molecular interactions and provide potential intervention targets.
[0131] In the post-transcriptional process, we found that myopia-associated SNPs may also disrupt the secondary structure of genomic regulatory elements and lead to molecular dysregulation. A previous study reported that a classic myopia-associated risk molecule FGF10 had its secondary structure disrupted by rs339501 (P PT =0.11, =1.99). Then, we also obtained the most likely structure of the local region, with the C allele located in the inner loop and the U allele located on the stem. The results indicate that rs339501 can change the single or double strands of FGF10 and have an impact on RNA stability. Similarly, rs905224 located on the 3'UTR may disrupt the structural stability of ZNF891. Then, two binding proteins of ZNF891, GAPDH and PSME3, were downloaded from the IntAct database. By detecting the molecular interactions induced by rs905224, we found that changes in the RNA secondary structure of ZNF891 may affect the binding of these two proteins in the 3D spatial conformation and have different docking scores. Similarly, a large number of previous studies have shown that SNP-mediated changes in RNA secondary structure may contribute to the development of the disease. For example, two SNPU22G and A56U in the 5'UTR of the FTL gene were confirmed to change the overall mRNA structure and were associated with hyperferritinemia cataract syndrome. Another study showed that alleles of rs27770 in the 3'UTR exhibited different minimum free energy (MFE) structures that could significantly affect mRNA stability and increase cancer risk. These observations support that RNA secondary structures of genomic regulatory elements are key factors in myopia. We developed a computational pipeline to identify cfSNPs in transcriptional and post-transcriptional processes. This is of great significance for understanding the potential regulatory mechanisms underlying the development of myopia.
[0132] In conclusion, we have precisely mapped for the first time and comprehensively established the relationship between SNPs, genomic regulatory elements, and myopia. Based on self-designed molecular binding and classical structural heterogeneity assessment algorithms, we identified a series of cfSNPs that can disrupt the regulatory functions of genomic regulatory elements at the transcriptional and post-transcriptional levels. In addition, we developed an RNA subunit heterogeneity assessment algorithm to further evaluate the effects of cfSNPs on RNA structural accessibility and stability. These results provide a broad perspective on molecular regulation in future studies.
[0133] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment, or a part of a code, and the module, program segment, or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0134] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the disclosed embodiments are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general purpose hardware or controllers or other computing devices, or some combination thereof as non-limiting examples.
[0135] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0136] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0137] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0138] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0139] The exemplary embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. It should be understood by those skilled in the art that various modifications and combinations may be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.
Claims
1. A method for identifying RNA structural changes based on SNPs, characterized in that: The method comprises: 201, obtaining SNPs to be tested; 202, based on the SNPs to be detected, retrieve Ref and Alt sequence pairs, wherein the sequences correspond to target RNA; 203, extracting RNA secondary structure features of the target RNA structure; the secondary structure features include B, E, H, I, M, S subunits, and the number, length and position of each subunit; standardizing the number, length and position of each subunit, and the standardization method is: in, represents any numerical 6 × 2 vector, is a normalized value ranging from 0 to 1; B, E, H, I, M, and S represent convex loop, outer loop, hairpin loop, inner loop, multi-branched loop, and stem, respectively; and They represent the secondary structures of Ref and Alt RNA induced by SNP, respectively; 204, the Euclidean distance was used to calculate the B, E, H, I, M, and S subunits, and the number, length, and position of each subunit were used to quantify the structural differences between the Ref sequence and the Alt sequence to obtain the difference value; 205, output the effect of SNPs on RNA structure changes based on the difference value; The method for extracting RNA secondary structure features of a target RNA structure comprises: 101, obtaining a target RNA to be predicted; 102, obtaining an RNA secondary structure set based on the target RNA; 103, performing cluster analysis on the RNA secondary structure set to obtain at least 2 clusters, and selecting a centroid structure and its energy value of each cluster; the number of centroid structures is consistent with the number of clusters; 104, comparing the energy values of the centroid structures, and selecting the centroid structure with the lowest energy value as the predicted RNA variable region structure.
2. The method for identifying RNA structural changes based on SNPs according to claim 1, characterized in that: The quantification method in 204 includes: M represents the characteristics of each RNA subunit, N refers to B, E, H, I, M, S subunits, and Respectively represent the normalized values of each RNA subunit feature; represents the difference value obtained after quantification; B, E, H, I, M, and S represent convex loop, outer loop, hairpin loop, inner loop, multi-branched loop, and stem, respectively.
3. The method for identifying RNA structural changes based on SNPs according to claim 1, characterized in that: The method for selecting the center of mass structure with the lowest energy value includes: performing a pairwise cyclic comparison or traversal comparison on the center of mass structures in each cluster, and determining the center of mass structure with the lowest minimum free energy as the optimal structure.
4. The method for identifying RNA structural changes based on SNPs according to claim 3, characterized in that: The energy value is a comprehensive evaluation of the free energy of the overall structure based on the recognized thermodynamic stability principle in structure prediction; the thermodynamic stability principle includes any one or a combination of the following structures: base pair energy, unpaired region and loop structure.
5. The method for identifying RNA structural changes based on SNPs according to claim 3, characterized in that: The method for selecting the centroid structure of each cluster includes: based on the statistical Boltzmann sampling method, obtaining the structure with the highest occurrence frequency in each cluster as the representative structure of each cluster, that is, the centroid structure.
6. A computer device, characterized in that: The device comprises: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the steps of the method according to any one of claims 1 to 5.
7. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Method, system and equipment for comparing RNA structures based on RNA motif vectors
CN113936737A
Method of identifying functional section of RNA from base sequence
JP2003242153A