A method, device, medium and program product for identifying RNA structural changes at the post-transcriptional level of the genome
By identifying and analyzing the RNA structural changes after genome transcription, the problem of difficulty in systematically regulating the impact of SNP on RNA structure and function in the prior art is solved, and efficient identification and analysis of RNA structural changes is achieved, and the interpretability of disease-related SNPs is enhanced.
Patent Information
- Application Number
- CN202510096381.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-22
AI Technical Summary
It is difficult for the prior art to effectively identify and analyze RNA structural changes after genome transcription, especially how to systematically regulate the impact of SNP on RNA structure and function.
By obtaining the SNPs to be tested, ref and Alt sequence pairs are retrieved, the secondary structural characteristics of RNA are extracted, and the structural differences are calculated using Euclidean distance to identify the impact of SNP on RNA structural changes.
Efficient identification and analysis of RNA structural changes at the post-genomic transcription level was achieved, and processes were developed to evaluate the global and local effects of SNPs on variable regions of RNA structural were enhanced, enhancing the interpretability of disease-related SNPs.
Smart Images

Figure CN119541628B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biological analysis, and more particularly, to a method, device, medium and program product for identifying RNA structural changes at the post-transcription level of the genome. Background Art
[0002] Post-transcriptional processes are key steps in regulating gene expression and are essential for the normal function and health of organisms. Post-transcriptional processes include modifications and regulation of mRNA, such as splicing, capping and tailing, and editing, which ensure the stability and correctness of mRNA and affect its translation efficiency. For example, the splicing process can remove introns (non-coding regions) in mRNA and connect exons (coding regions) to form mature mRNA, which is essential for the production of functional proteins. Post-transcriptional regulation also allows cells to respond quickly to environmental changes and change protein synthesis by adjusting the expression level of specific mRNA. In the post-transcriptional process, 3'UTR, 5'UTR, exons, introns and multiple RNAs can form complex spatial conformations, acting as "scaffolds" to regulate RNA stability, expression, translation and decay. Abnormalities in the post-transcriptional process may lead to mRNA instability or incorrect splicing, which in turn affects the structure and function of RNA.
[0003] Genomic regulatory elements are key regulators in the processes of gene transcription, expression, and translation. Molecular dysregulation of genomic regulatory elements may lead to human diseases, however, the functions of these elements remain poorly understood. Previous studies have shown that genomic regulatory elements cover the vast majority of genetic variations, especially single nucleotide polymorphisms (SNPs), which are prone to abnormal regulation and are likely to modulate disease susceptibility. The above-mentioned functional SNPs are also defined as "candidate functional SNPs (cfSNPs)". On the one hand, cfSNPs can significantly disrupt the molecular binding between cis-acting elements and trans-acting factors, which can affect transcriptional activity and may lead to complex diseases. On the other hand, cfSNPs can change RNA secondary structure and affect molecular function, which is further involved in a variety of human diseases. Therefore, it is crucial to dissect the systemic regulation between SNPs, target gene regulatory elements, and the original phenotype at the post-transcriptional level of genes. Summary of the invention
[0004] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention provides a method, device, medium and program product for identifying RNA structural changes at the post-transcriptional level of the genome; the method of the present invention identifies RNA structural changes through SNP at the post-transcriptional level of the gene, and develops a process to study the system regulation between SNP and post-transcriptional genome regulatory elements.
[0005] The first aspect of the present application discloses a method for identifying RNA structural changes at the post-transcription level of the genome, the method comprising: 101, obtaining SNPs to be tested; 102, retrieving Ref and Alt sequence pairs based on the SNPs to be tested; 103, extracting RNA secondary structural features based on the Ref and Alt sequence pairs; the secondary structural features include B, E, H, I, M, S subunits, and the number, length and position of each subunit; 104, using Euclidean distance to calculate B, E, H, I, M, S subunits, and the number, length and position of each subunit to quantify the structural differences between the Ref sequence and the Alt sequence to obtain difference values; based on the difference values, identifying the impact of SNPs on RNA structural changes.
[0006] In some embodiments, between 103 and 104, the method further comprises standardizing the number, length and position of each subunit, wherein the standardization method is:
[0007]
[0008] in, represents any numerical 6 × 2 vector, is a normalized value ranging from 0 to 1; B, E, H, I, M, and S represent convex loop, outer loop, hairpin loop, inner loop, multi-branched loop, and stem, respectively; and They represent the ref and alt RNA secondary structures induced by SNPs, respectively. RNA secondary structure specifically refers to the number, length, and position of each type of subunit in RNA secondary structure.
[0009] In some embodiments, the quantification method in 104 includes:
[0010]
[0011] M represents the characteristics of each RNA subunit, N refers to B, E, H, I, M, S subunits, and Respectively represent the normalized values of each RNA subunit feature; represents the difference value obtained after quantification; B, E, H, I, M, and S represent convex loop, outer loop, hairpin loop, inner loop, multi-branched loop, and stem, respectively.
[0012] In some embodiments, the method further includes 105, identifying the effect of SNPs on RNA structural changes based on the difference values, and then sorting the SNPs.
[0013] In some embodiments, the method further includes: identifying the RNA structure variable region of the post-transcriptional gene regulatory element in the Ref and Alt sequence pairs; extracting RNA secondary structure features from the RNA structure variable region; the RNA structure variable region includes a global RNA structure variable region and / or a local RNA structure variable region.
[0014] In some embodiments, the SNPs in the global RNA structure variable region are classified into proximal allosteric effects and distal allosteric effects; if the SNPs are located in the global RNA structure variable region and affect the folding of the region, the output belongs to the proximal allosteric effect; if the SNPs are not located in the global RNA structure variable region and affect the folding of the region, the output belongs to the distal allosteric effect.
[0015] In some embodiments, the post-transcriptional gene regulatory elements include: 3'UTR, 5'UTR, exons, introns, and long non-coding RNA (LncRNA).
[0016] The second aspect of the present application discloses a computer device, which includes: a memory and a processor; the memory is used to store a computer program; and the processor executes the computer program to implement the steps of the above method.
[0017] A third aspect of the present application discloses a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of the above method when executed by a processor.
[0018] A fourth aspect of the present application discloses a computer program product, including a computer program, which implements the steps of the above method when executed by a processor.
[0019] The present application has the following beneficial effects: 1. The present application innovatively discloses a method for identifying RNA structural changes at the post-transcriptional level of the genome. The method retrieves Ref and Alt sequence pairs based on SNPs, extracts target RNA secondary structural features based on the sequence pairs, quantifies the secondary structures in the Ref sequence and the Alt sequence using Euclidean distance to obtain difference values, identifies the effects of SNPs on RNA structural changes based on the difference values, and explores the interference of post-transcriptional regulation;
[0020] 2. This application innovatively develops a process to comprehensively evaluate the global and local effects of SNPs on variable regions of RNA structure, providing an effective method for evaluating the accessibility and stability of RNA.
[0021] 3. The method proposed in this application can be used to study the possible genetic mechanisms of SNPs with undefined functions for other human diseases, thereby enhancing the interpretability of disease-related SNPs. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0023] Figure 1 is a schematic diagram of a method flow chart provided by the first aspect of an embodiment of the present invention;
[0024] Figure 2 Schematic diagram of a system for identifying RNA structural changes at the post-transcription level of a genome provided by an embodiment of the present invention;
[0025] Figure 3 is a schematic diagram of a computer device provided by an embodiment of the present invention;
[0026] Figure 4 is a schematic diagram of the architecture of an exemplary computing device provided by an embodiment of the present invention;
[0027] Figure 5 is a schematic diagram of a storage medium provided by an embodiment of the present invention;
[0028] Figure 6 is a fine-scale mapping map of myopia-related SNPs on genome-wide regulatory elements provided by an embodiment of the present invention; wherein, Figure 6 A is the distribution of each genomic regulatory element in transcription and post-transcriptional processes; Figure 6 B is the percentage of relationship pairs in transcription and posttranscriptional processes; Figure 6 C is the length of the genomic regulatory elements covering the myopia-associated SNPs, and the black-marked horizontal line in each violin represents the median length of each element; Figure 6 D is the SNP density of each regulatory element at the transcriptional and post-transcriptional levels. The dotted line shows the average SNP density of the entire human genome;
[0029] Figure 7 It is a global RNA structure variable region mediated by SNP on genomic regulatory elements in the post-transcriptional process provided by an embodiment of the present invention; wherein, Figure 7 A is the length of the variable region of RNA structure in the genomic regulatory element; Figure 7 B is the percentage of relationship pairs showing proximal and distal effects mediated by SNPs on genomic regulatory elements; Figure 7 C is the percentage of relationship pairs with important RNA structural variable regions on genomic regulatory elements; Figure 7 D is the percentage of relationship pairs covering significant RNA structural variable regions in CM; Figure 7E is the percentage of relationship pairs in HM that cover significant RNA structural variable regions. DETAILED DESCRIPTION
[0030] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0031] In some of the processes described in the specification and claims of the present invention and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent the order of precedence, and do not limit the "first" and "second" to be different types.
[0032] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0033] Figure 1 1 is a flow chart of a method for identifying RNA structural changes at the post-transcription level of a genome provided by an embodiment of the present invention. Specifically, the method comprises the following steps: 101: obtaining SNPs to be tested;
[0034] The full name of SNP is Single Nucleotide Polymorphism, which refers to the variation of a single nucleotide (base pair) in a genomic DNA sequence. This variation may be a substitution (e.g., A becomes G), an insertion or a deletion. SNP is the most common form of genetic variation in the genome, and their distribution in the population is polymorphic, that is, different individuals may have different nucleotides at the same position. In some embodiments, the SNP to be tested comes from a subject, and the terms "subject" or "person to be tested" or "sample to be tested" used herein refer to any animal (e.g., mammal), including but not limited to humans, non-human primates, rodents, etc., which will become the recipient of a specific treatment. Generally, the terms "subject" and "patient" are used interchangeably herein when referring to human subjects. Preferably, the subject is a human.
[0035] 102: Retrieving Ref and Alt sequence pairs based on the SNPs to be tested;
[0036] In some embodiments, based on each SNP, the Ref and Alt sequence pairs corresponding to the SNP are retrieved. Specifically, the Ref and Alt sequences are retrieved each Mbp upstream and downstream to obtain several sequence pairs, where M is a natural number greater than 1; the range of M is 10-35, preferably 20. Specifically, Ref is a wild type, and Alt is a mutant. The number refers to a natural number greater than 1.
[0037] 103: extracting RNA secondary structure features based on the Ref and Alt sequence pairs; the secondary structure features include B, E, H, I, M, S subunits, and the number, length and position of each subunit; in some embodiments, between 103 and 104, the method further includes standardizing the number, length and position of each subunit, and the standardization method is:
[0038]
[0039] in, represents any numerical 6 × 2 vector, is a normalized value ranging from 0 to 1; B, E, H, I, M, and S represent convex loop, outer loop, hairpin loop, inner loop, multi-branched loop, and stem, respectively; and They represent the ref and alt RNA secondary structures induced by SNPs, respectively. RNA secondary structure specifically refers to the number, length, and position of each type of subunit in RNA secondary structure.
[0040] 104: Use Euclidean distance to calculate the B, E, H, I, M, and S subunits, as well as the number, length, and position of each subunit to quantify the structural differences between the Ref sequence and the Alt sequence to obtain the difference value; identify the impact of SNP on RNA structural changes based on the difference value.
[0041] In some embodiments, the quantification method in 104 includes:
[0042]
[0043] M represents the characteristics of each RNA subunit, N refers to B, E, H, I, M, S subunits, and Respectively represent the normalized values of each RNA subunit feature; represents the difference value obtained after quantification; B, E, H, I, M, and S represent convex loop, outer loop, hairpin loop, inner loop, multi-branched loop, and stem, respectively.
[0044] In some embodiments, if the difference value exceeds a third threshold, the result that the SNP has an effect on the RNA structure change is output. The third threshold is obtained by evaluating the lower quartile of the RNA subunits that undergo significant changes. In some embodiments, the third threshold is obtained by training with training set samples, which can be a specific threshold or an interval range, and the specific form is not specifically limited in this embodiment.
[0045] In some embodiments, the method further includes 105, identifying the effect of SNPs on RNA structural changes based on the difference values, and then sorting the SNPs.
[0046] In some embodiments, the method further includes: identifying the RNA structure variable region of the post-transcriptional gene regulatory element in the Ref and Alt sequence pairs; extracting RNA secondary structure features from the RNA structure variable region; the RNA structure variable region includes a global RNA structure variable region and / or a local RNA structure variable region.
[0047] In some embodiments, the SNPs in the global RNA structure variable region are classified into proximal allosteric effects and distal allosteric effects; if the SNPs are located in the global RNA structure variable region and affect the folding of the region, the output belongs to the proximal allosteric effect; if the SNPs are not located in the global RNA structure variable region and affect the folding of the region, the output belongs to the distal allosteric effect. Specifically, the proximal allosteric effect refers to the effect on the surrounding small segment area, and only the current gene or other of the SNP is considered when doing molecular experiments; the distal allosteric effect refers to the effect not on the current gene, but may be on other genes, and other genes are considered first when doing molecular experiments.
[0048] In some embodiments, the post-transcriptional gene regulatory elements include: 3'UTR (3' untranslated region), 5'UTR (5' untranslated region), exons, introns, long non-coding RNA (LncRNA). Studying genomic regulatory elements helps to gain a deeper understanding of how genes are regulated and how these regulations affect the physiological and pathological processes of organisms.
[0049] In some embodiments, the evaluation results or prediction results include but are not limited to paper or electronic report forms. The results are only obtained by the intelligent machine based on the analysis of relevant data of the subjects, and are only used as a reference, not as the final diagnosis result.
[0050] Figure 3 is a schematic diagram of a computer device provided by an embodiment of the present invention, such as Figure 3As shown, the device 2000 may include: one or more processors 2010, and one or more memories 2020; wherein the memories store computer-readable codes, and when the computer-readable codes are run by the one or more processors, the method described above may be executed.
[0051] The processor in this embodiment can be an integrated circuit chip with signal processing capabilities. The above processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The disclosed methods, operations and logic block diagrams in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc., which can be an X86 architecture or an ARM architecture.
[0052] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the disclosed embodiments are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general purpose hardware or controllers or other computing devices, or some combination thereof as non-limiting examples.
[0053] For example, the method or device according to the embodiment of the present disclosure may also be implemented by Figure 4 The architecture of the computing device 3000 shown in FIG. Figure 4 As shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as ROM 3030 or hard disk 3070, may store various data or files used for processing and / or communication of the method provided by the present disclosure and program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 4 The architecture shown is only exemplary and can be omitted according to actual needs when implementing different devices. Figure 4 One or more components of a computing device are shown.
[0054] The embodiment of the present invention also provides a computer-readable storage medium, such as Figure 5As shown, it is a schematic diagram of a storage medium 4000 provided in an embodiment of the present invention, and a computer readable instruction 4010 is stored on the computer storage medium 4020. When the computer readable instruction 4010 is executed by a processor, the method according to the embodiment of the present disclosure described with reference to the above figures can be executed. The computer readable storage medium in the embodiment of the present disclosure can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory. It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0055] The embodiments of the present disclosure also provide a computer program product or system, including a computer program, which implements the steps of the above method when executed by a processor.
[0056] In some embodiments, this embodiment also discloses a system for identifying RNA structural changes at the post-transcription level of the genome, such as Figure 2 As shown, the system comprises:
[0057] A first acquisition module 201 is used or configured to acquire SNPs to be tested;
[0058] A sequence pair retrieval module 202 is used or configured to retrieve Ref and Alt sequence pairs based on the SNPs to be tested;
[0059] The RNA secondary structure feature extraction module 203 is used or configured to extract RNA secondary structure features based on the Ref and Alt sequence pairs; the secondary structure features include B, E, H, I, M, S subunits, and the number, length and position of each subunit;
[0060] The RNA structure change identification module 204 is used or configured to use Euclidean distance to calculate the B, E, H, I, M, S subunits, as well as the number, length and position of each subunit to quantify the structural differences between the Ref sequence and the Alt sequence to obtain difference values; based on the difference values, the impact of SNP on RNA structure changes is identified.
[0061] In some embodiments, the system further comprises a standardization module located between the RNA secondary structure feature extraction module and the RNA structure change identification module, which is used or configured to standardize the number, length and position of each subunit, and the standardization method is:
[0062]
[0063] in, represents any numerical 6 × 2 vector, is a normalized value ranging from 0 to 1; B, E, H, I, M, and S represent convex loop, outer loop, hairpin loop, inner loop, multi-branched loop, and stem, respectively; and Represent the Ref and Alt RNA secondary structures induced by SNP, respectively. RNA secondary structure specifically refers to the number, length and position of each type of subunit in RNA secondary structure. Specific embodiment:
[0065] 1 Materials and Methods
[0066] 1.1 Collecting Data:
[0067] Myopia-associated SNPs were collected from dbGap, GWAS catalogs, and published literature. Based on the population distribution of myopia-associated SNPs in the 1000 Genomes Project (GRCh38), we performed quality control (QC) on samples and genotypes. SNPs were filtered according to the following selection criteria: East Asian, two alleles in NCBI, minor allele frequency (MAF)>5%, Hardy-Weinberg equilibrium P<0.01, recall rate>75%, and genotyping rate>75%. 343 SNPs were obtained for the following analysis. The start and end positions of eight genomic regulatory elements, including open chromatin regions (OCRs), CTCF binding sites (CTCFBS), enhancers, promoters, promoter flanking regions (PFRs), exons, introns, and noncoding RNA (ncRNA) transcripts, were derived from ENSEMBL (v102), and their sequences were extracted according to Refseq from NCBI. The sequences of 3'UTR and 5'UTR were obtained using the BioMart tool from ENSEMBL.
[0068] 1.2 Fine positioning of myopia-related SNPs on 10 genomic regulatory elements:
[0069] Bowtie2, a classic read alignment tool, was used to perform precise positioning. The primary sequence of the genomic regulatory element was considered as a long reference. The 30 bp flanking regions upstream and downstream of these SNPs were considered as short alignment sequences. Then, strict parameters were set using "--n-ceil C,3 --np 0 --end-to-end -a --score-min C,0" to avoid seed sequence mismatches. After fine positioning, we constructed paired reference (Ref) and alternative (Alt) sequences based on the alleles of the SNP. The alignment software includes Bowtie1, Bowtie2, BLAST, preferably Bowtie2.
[0070] 1.3 Assessing changes in molecular binding at the transcriptional level:
[0071] We extracted the 20bp upstream and downstream of the paired Ref and Alt sequences induced by the SNP, and obtained the binding motifs of Enhancer, OpenchromatinRegion (OCR), Promoter, and Promoterflankingzone (PFR) of these sequences from the HumanTFDB database, while CTCF was from the CTCFBSDB database. For protein binding, we obtained the number of regulatory proteins (TF and CTCF), protein symbols, and the start and end positions of the binding motifs, respectively. First, we used fold change (FC) to obtain the change in the number of protein binding sites mediated by myopia-associated SNPs:
[0072]
[0073] and Represents the number of proteins bound by the Ref and Alt sequences, respectively. We defined FC>1 as a loss of protein binding, and conversely, FC<1 as a gain of protein binding. We then applied Fisher's exact test to determine changes in protein binding motifs as follows:
[0074]
[0075] a is the number of binding motifs that overlap only with the Ref sequence, b is the number of binding motifs that do not overlap with the Ref sequence, c is the number of binding motifs that overlap only with the Alt sequence, and d is the number of binding motifs that do not overlap with the Alt sequence. Here, when FC is not equal to 1 and P T <0.05, indicating that the myopia-associated SNPs significantly affected protein binding, and these SNPs were identified as cfSNPs. In addition, we defined the proteins with the highest scores as “leader” proteins.
[0076] 1.4 Identification of variable regions of global RNA structure after transcription:
[0077] At the post-transcriptional level, we used RNAsnp to identify structurally variable regions of five regulatory elements, including 3'UTR, 5'UTR, exons, introns, and lncRNA transcripts. PT <0.2 indicates that the SNP may have a significant effect on the local RNA structure. Here, we define these regions as "global RNA structure variable regions". If the SNP is located in the RNA structure variable region and affects the folding of this region, these SNPs are defined as "proximal allosteric effects". Otherwise, if the SNP is not located in the RNA structure variable region and affects the surrounding folding, these SNPs are defined as "distal allosteric effects".
[0078] 1.5 Using RNA subunits to quantify changes in RNA variable regions:
[0079] We evaluated SNP-induced changes in RNA subunits within important global RNA structural variable regions. First, we predicted RNA secondary structures that were as consistent as possible with the true folding state. Cluster analysis was performed based on 1,000 possible structures to obtain representative structures of multiple clusters. By comparing the minimum free energy (MFE) of the representative RNA structures in each cluster, we selected the structure with the lowest free energy as the most likely structure. Then, RNAsmc was used to extract the six RNA subunits of each RNA, namely the bulge loop (B), outer loop (E), hairpin loop (H), inner loop (I), multi-branched loop (M) and stem (S) of each RNA. Next, we developed a computational pipeline and quantified the structural heterogeneity of the variable regions of the global RNA structure using Euclidean distance (based on RNA subunits) ( ). We obtained the characteristics of RNA subunits in the variable region of the global RNA structure, namely the number, length and position of each subunit in the paired Ref and Alt structures. For these three dimensions, we constructed two corresponding 6×1 vectors:
[0080]
[0081]
[0082] Among them, i represents the characteristics of RNA subunits. B, E, H, I, M and S refer to the six RNA subunits, namely the convex loop, outer loop, hairpin loop, inner loop, multi-branched loop and stem. R and A represent the Ref and Alt RNA secondary structures mediated by SNPs, respectively. Then, in order to standardize the size of the three characteristics of RNA subunits, each type of subunit is standardized as follows:
[0083]
[0084] represents any numerical 6×2 vector, ranges from 0 to 1. Finally, we use the Euclidean distance ( ) to quantify the difference between the ref and alt structures:
[0085]
[0086] M represents the characteristics of each subunit. N refers to six subunits. and denotes the standardized value of each feature. The lower quartile of was set as the significance threshold. Above the lower quartile, we considered the RNA subunits to exhibit significant changes. VARNA was used to visualize RNA secondary structure.
[0087] 1.6 Evaluate the same annotation of SNPs based on the HM cohort and computational pipeline:
[0088] A total of 10,348 high myopic participants (worst eye SE < -6.00D) in the CAMS study were sequenced on the Illumina NovaSeq6000 sequencer from Berry Genomics using the Twist Human Core Exome Kit. We obtained the genotypes and phenotypes of myopia-associated SNPs from the CAMS study. Next, we used combined annotation-dependent deletion (CADD) to score the harmfulness of SNPs based on the genotypes and phenotypes in the high myopia cohort. Here, SNPs with a CADD score ≥ 10 are defined as harmful SNPs that may lead to loss of gene function. Combined Annotation Dependent Depletion score (CADD score): used to evaluate and quantify single nucleotide variations (SNVs). Higher scores indicate more harmful variants, that is, higher pathogenicity potential.
[0089] 2 Results:
[0090] 2.1 Myopia-related SNPs are widely distributed at the transcriptional and post-transcriptional levels:
[0091] We obtained myopia-associated SNPs from public resources and, after quality control, used 343 SNPs for subsequent analysis (see Methods). To reveal the complete map of SNPs at the transcriptional and post-transcriptional levels, we used precise mapping and identified 636 relationship pairs formed by 263 SNPs, 10 genomic regulatory elements, and myopia ( Figure 6A). In addition, these SNPs were associated with five phenotypes, namely refractive error (RE), common myopia (CM), high myopia (HM), pathological myopia (PM) and visual impairment (VD). In the transcriptional process, 84 SNPs were located in enhancers, open chromatin regions (OCRs), CTCF binding sites (CTCFBS), promoters and promoter flanking regions (PFRs), forming 90 relationship pairs. A total of 244 SNPs were located on five genomic regulatory elements, namely 5'UTR, exons, introns, 3'UTR and LncRNA, constituting 546 relationship pairs at the post-transcriptional level. Next, we found that 14.15% (90 / 636) and 85.85% (546 / 636) of the pairs were enriched in the transcriptional and post-transcriptional processes, respectively ( Figure 6 B).
[0092] To further evaluate the average distribution of SNPs in all genomic regulatory elements, we analyzed the transcript lengths of these elements and the density of SNPs in each genomic regulatory element. The results showed that the median length of the longest transcript LncRNA was 95.53 times the median length of the shortest transcript CTCFBS ( Figure 6 C). We then further evaluated the enrichment of SNPs in each regulatory element within every 1000 bp. Compared with the average of 1 SNP per 1000 bp in the human genome, we observed that myopia-associated SNPs were highly distributed in OCR, CTCFBS, 5'UTR, and exons ( Figure 6 D). By analyzing myopia, 81.76% (520 / 636) of the relationship pairs were found to be associated with CM. In addition, little is known about HM, PM, RE, and VD, which are approximately 12.11% (77 / 636), 5.19% (33 / 636), 0.63% (4 / 636), and 0.31% (2 / 636), respectively. These data suggest that the distribution of myopia-associated SNPs differs at the transcriptional and posttranscriptional levels and between genomic regulatory elements. These differences may be related to the severity of myopia and potential molecular regulation.
[0093] 2.2 Scoring SNP-induced molecular binding heterogeneity during transcription:
[0094] To investigate SNP-mediated molecular binding heterogeneity at the transcriptional level, we developed a computational pipeline that used the fold test (FC) to assess changes in the number of bound proteins and applied Fisher’s exact test to assess changes in the number of bound proteins. T<0.05, FC was not equal to 1, and 38.46% (5 / 13), 80% (4 / 5), 57.89% (11 / 19), and 44.73% (17 / 38) of the SNPs were found to be likely to disrupt transcription factor binding at enhancers, open chromatin regions, promoters, and PFRs, respectively ( Figure 7 ). For CTCF protein, we found that about 13.33% (2 / 15) of SNPs could disrupt CTCFBS interactions. In total, 43.33% (39 / 90) of the relationship pairs were able to disrupt binding affinity, and 46.43% (39 / 84) SNPs were identified as “cfSNPs” during transcription. In summary, open chromatin regions were significantly enriched with myopia-associated SNPs and cfSNPs, showing high-density distribution and disruption of molecular binding.
[0095] To determine the possible effects of regulatory proteins associated with cfSNPs that are associated or potentially associated with myopia, we explored the molecular functions of the “leader” proteins of each relationship pair in published studies. Interestingly, most of the “leader” proteins are able to affect ocular tissues or structures. For example, about 7.69% (3 / 39) of cfSNPs cause changes in the REST binding motif, thereby affecting the fate of RGC retinal ganglion cells (RGCs) in the developing retina. About 7.69% (3 / 39) of cfSNPs cause changes in IRF1 binding, which is known to be expressed in retinal microglia and plays a key role in microglial activation and retinal inflammation. In addition, 15.28% (6 / 39) of cfSNPs change the binding motif site of SPI1, which has been reported to regulate microglia in the retina. In summary, we found that there are abundant leader proteins in retinal inflammation, which have been shown to be associated with the occurrence and development of myopia. In addition, it can help us reveal potential protein regulators and understand how SNPs are involved in the molecular regulation of myopia. Finally, we evaluated the distribution of myopia types in pairs with significant changes at the transcriptional level. More than 75% (3 / 4) of them were related to PM, while about 44.44% (4 / 9) and 41.56% (32 / 77) were related to HM and CM ( Figure 7 AD). The results also showed that the effect of SNP on myopia increased with the severity of myopia.
[0096] 2.3 Identification of SNP-mediated global RNA structural variable regions during post-transcriptional processes:
[0097] To identify SNP-induced RNA secondary structural heterogeneity during post-transcriptional processes, we initially employed RNA SNPs to detect potential global RNA structural variable regions, P PT<0.2. Here, since there is only one-to-one pair in the 5'UTR, this exon was chosen as a reference to compare the significance of the average length difference of the variable region of RNA structure. Figure 7 As shown in A, the RNA structural variable regions of 5'UTR, 3'UTR and Exon are relatively longer than introns and LncRNA. It is worth noting that the SNPs associated with myopia are not only highly enriched in 5'UTR and exons, but also cause huge structural damage to these regulatory regions. Although myopia-related SNPs are not significantly enriched in the 3'UTR region, they still cause significant structural effects. This may be due to the fact that the 3'UTR region presents highly structured characteristics, and once it is damaged, it has a great impact on the RNA secondary structure.
[0098] In addition, we evaluated whether the SNPs were located in the variable regions of RNA structure. The definitions are shown in the Methods. Statistically, approximately 97.54% (238 / 244) of the SNPs exhibited proximal allosteric effects, mapping to 91.96% (502 / 546) of the relationship pairs and all five regulatory elements, while only 12.70% (31 / 244) of the SNPs exhibited proximal allosteric effects. The SNPs showed distal allosteric effects, mapping to 8.06% (44 / 546) of the relationship pairs and three regulatory elements, namely, exons, introns, and lncRNAs ( Figure 7 B). This finding is consistent with a previous study showing that SNPs affect mainly local regions rather than having a global effect.
[0099] To further explore which genomic regulatory elements are significantly disrupted by SNPs, we analyzed the proportion of important RNA structural variable regions within these elements. Figure 7As shown in Figure C, among the 9 pairs in 3'UTR, 33.33% (3 / 9) were induced by 37.50% (3 / 8) of the SNPs and showed significant RNA structural variable regions, while 1 pair in 5'UTR did not show significant RNA structural variable regions. Among the 35 pairs in exons, 22.86% (8 / 35) of the pairs showed significant changes mediated by 20.69% (6 / 29) of the SNPs. In addition, among the 308 pairs in introns, 14.61% (45 / 308) contained significant RNA structural variable regions, which were affected by 16.52% (38 / 230) of the SNPs. Among the 193 pairs in LncRNA, 12.95% (25 / 193) showed significant RNA structural variable regions, which were induced by 14.56% (23 / 158) of the SNPs. Overall, we found that 14.84% (81 / 546) of relationship pairs contained significant RNA structural variable regions induced by 20.08% (49 / 244) of the SNPs, which were identified as cfSNPs.
[0100] To characterize the role of RNA secondary structure in myopia, we analyzed the distribution of important RNA structural variable regions. Three myopia phenotypes, RE, CM, and HM, showed significant SNP-mediated changes. Among these phenotypes, 1.23% (1 / 81) of the relationship pairs were associated with RE, 16.05% (13 / 81) with HM, and 82.72% (67 / 81) with CM. Figure 7 As shown in D and E, HM exhibited a higher proportion of important regions compared with CM, such as 3'UTR, exons, and introns that were more susceptible to cfSNPs in HM.
[0101] 2.4 Quantifying the local stability of important RNA structural variable regions based on RNA subunits:
[0102] To further evaluate the changes in RNA stability and accessibility caused by cfSNPs in important RNA structural variable regions, we revealed changes based on the single-stranded or double-stranded state of RNA subunits. First, we designed a computational pipeline to identify the most likely structures of 81 important RNA structural variable regions using Sfold. Here, we obtained experimentally determined RNA secondary structures from RNASTRAND and used these structures as references to evaluate the accuracy of the predicted structures. For example, Sfold and RNASTRAND showed high consistency in predicting RNA secondary structures of stRNAs, achieving an RNAsmc score of 10. Then, we obtained the basic RNA subunits of RNA structural variable regions, which are stem (S), internal loop (I), bulge loop (B), multi-branched loop (M), hairpin loop (H), and external loop (E). The single-stranded and double-stranded states of RNA folding are closely related to RNA stability or accessibility of RNA binding. By analyzing the RNA subunits of the paired Ref and Alt structures of 81 relationship pairs induced by 35 cfSNPs in significant RNA structural variable regions, we found that 53.09% (43 / 81) of these relationship pairs showed changes in pairing state. Among them, 24.69% (20 / 81) changed from single-stranded to double-stranded, 28.40% (23 / 81) changed from double-stranded to single-stranded, and 46.91% (38 / 81) remained unchanged. In the case of subunit changes, the double-stranded state (S represents the double-stranded state) and other single-stranded subunits (B, E, H, I and M are single-stranded states) changed significantly.
[0103] Finally, to identify local changes in RNA structural variable regions that are significant due to cfSNPs at the post-transcriptional level, we performed a comprehensive comparison based on RNA subunits, including the number, length, and base composition of each subunit between the Ref and Alt structures. The structural differences were quantified using Euclidean distance. The threshold for evaluating significant changes in RNA subunits was defined as the lower quartile of the observed structural change values ( =1.65). Here, 8.79% (48 / 546) of the important RNA structural variable regions induced by 14.34% (35 / 244) of the SNPs may undergo changes in RNA subunits, thereby affecting the secondary structure of RNA. These results suggest that myopia-associated SNPs play an important role in the occurrence of myopia-related diseases by affecting RNA stability.
[0104] This study established for the first time a comprehensive map of molecular dysregulation of genomic regulatory elements mediated by myopia-associated SNPs. A total of 636 relationship pairs were identified for 263 SNPs, 10 genomic regulatory elements, and 5 genomic regulatory elements. We identified a total of 82 myopia cfSNPs on genome-wide regulatory elements, of which 39 cfSNPs with 39 relationship pairs were detected at the transcriptional level and 49 cfSNPs with 81 relationship pairs were detected at the transcriptional level. Using FC and P T SNP-mediated gain or loss of binding proteins was quantified. At the post-transcriptional level, we used RNAsNPs to further investigate important RNA structural variable regions from a global perspective. In addition, we designed a new method to quantify changes in RNA subunits from a local perspective, which can reflect the accessibility and stability of RNA. In addition, we obtained genotype and phenotypic information from a previously established high myopia cohort and assessed the harmfulness of SNPs based on CADD. In summary, this study revealed potential molecular regulation, enhanced the interpretability of SNPs, and provided new insights into the genetic mechanism of myopia.
[0105] In the post-transcriptional process, we found that myopia-associated SNPs may also disrupt the secondary structure of genomic regulatory elements and lead to molecular dysregulation. A previous study reported that a classic myopia-associated risk molecule FGF10 had its secondary structure disrupted by rs339501 (P PT =0.11, =1.99). Then, we also obtained the most likely structure of the local region, with the C allele located in the inner loop and the U allele located on the stem. The results indicate that rs339501 can change the single or double strands of FGF10 and have an impact on RNA stability. Similarly, rs905224 located on the 3'UTR may disrupt the structural stability of ZNF891. Then, two binding proteins of ZNF891, GAPDH and PSME3, were downloaded from the IntAct database. By detecting the molecular interactions induced by rs905224, we found that changes in the RNA secondary structure of ZNF891 may affect the binding of these two proteins in the 3D spatial conformation and have different docking scores. Similarly, a large number of previous studies have shown that SNP-mediated changes in RNA secondary structure may contribute to the development of the disease. For example, two SNPU22G and A56U in the 5'UTR of the FTL gene were confirmed to change the overall mRNA structure and were associated with hyperferritinemia cataract syndrome. Another study showed that alleles of rs27770 in the 3'UTR exhibited different minimum free energy (MFE) structures that could significantly affect mRNA stability and increase cancer risk. These observations support that RNA secondary structures of genomic regulatory elements are key factors in myopia. We developed a computational pipeline to identify cfSNPs in transcriptional and post-transcriptional processes. This is of great significance for understanding the potential regulatory mechanisms underlying the development of myopia.
[0106] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment, or a part of a code, and the module, program segment, or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0107] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the disclosed embodiments are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general purpose hardware or controllers or other computing devices, or some combination thereof as non-limiting examples.
[0108] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0109] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0110] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0111] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0112] The exemplary embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. It should be understood by those skilled in the art that various modifications and combinations may be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.
Claims
1. A method for identifying RNA structural changes at the post-transcriptional level of the genome, characterized in that: The method comprises: 101, obtaining SNPs to be tested; 102, retrieve Ref and Alt sequence pairs based on the SNPs to be tested; 103. Extracting RNA secondary structure features based on the Ref and Alt sequence pairs; the secondary structure features include B, E, H, I, M, S subunits, and the number, length and position of each subunit; standardizing the number, length and position of each subunit, and the standardization method is: in, represents any numerical 6 × 2 vector, is a normalized value ranging from 0 to 1; B, E, H, I, M, and S represent convex loop, outer loop, hairpin loop, inner loop, multi-branched loop, and stem, respectively; and They represent the RNA secondary structures corresponding to Ref and Alt induced by SNP, respectively; 104. The Euclidean distance was used to calculate the B, E, H, I, M, and S subunits, and the number, length, and position of each subunit were used to quantify the structural differences between the Ref sequence and the Alt sequence to obtain the difference value; based on the difference value, the effect of SNP on RNA structural changes was identified.
2. The method for identifying RNA structural changes at the post-transcriptional level of a genome according to claim 1, characterized in that: The quantification method in 104 includes: M represents the characteristics of each RNA subunit, N refers to B, E, H, I, M, S subunits, and Respectively represent the normalized values of each RNA subunit feature; represents the difference value obtained after quantification; B, E, H, I, M, and S represent convex loop, outer loop, hairpin loop, inner loop, multi-branched loop, and stem, respectively.
3. The method for identifying RNA structural changes at the post-transcription level of a genome according to claim 2, characterized in that: The method further includes 105, identifying the influence of SNP on RNA structure change based on the difference value, and then sorting the SNPs.
4. The method for identifying RNA structural changes at the post-transcription level of a genome according to claim 1, characterized in that: The method also includes: identifying the RNA structure variable region of the post-transcriptional gene regulatory element in the Ref and Alt sequence pair; extracting RNA secondary structure features from the RNA structure variable region; the RNA structure variable region includes a global RNA structure variable region and / or a local RNA structure variable region.
5. The method for identifying RNA structural changes at the post-transcriptional level of a genome according to claim 4, characterized in that: The SNPs in the global RNA structure variable region are classified into proximal allosteric effects and distal allosteric effects; if the SNPs are located in the global RNA structure variable region and affect the folding of the region, the output belongs to the proximal allosteric effect; if the SNPs are not located in the global RNA structure variable region and affect the folding of the region, the output belongs to the distal allosteric effect.
6. The method for identifying RNA structural changes at the post-transcriptional level of a genome according to claim 4, characterized in that: The post-transcriptional gene regulatory elements include: 3'UTR, 5'UTR, exons, introns, and long non-coding RNA.
7. A computer device, characterized in that: The device comprises: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the steps of the method according to any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Method of identifying functional section of RNA from base sequence
JP2003242153A