Method for identifying candidate functional snps, device, medium, and program product
By screening candidate functional SNPs at the transcriptional and post-transcriptional levels and using computer analysis to assess the heterogeneity of molecular binding and RNA secondary structure, this method solves the problem of inaccurate localization of genomic regulatory elements and SNPs in existing technologies, enabling a deeper understanding of the pathogenesis of myopia and more accurate SNP identification.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- THE EYE HOSPITAL OF WENZHOU MEDICAL UNIVERSITY
- Filing Date
- 2025-04-30
- Publication Date
- 2026-07-30
AI Technical Summary
Current technologies lack precise localization and functional identification of genomic regulatory elements and single nucleotide polymorphisms (SNPs), especially during transcription and post-transcriptional processes, leading to an incomplete understanding of the pathogenesis of diseases such as myopia.
By screening candidate functional SNPs at the transcriptional and post-transcriptional levels, using computer analysis to assess the heterogeneity of molecular binding and RNA secondary structure, and combining Euclidean distance and Fisher's exact test, we can precisely locate SNPs on transcriptional and post-transcriptional regulatory elements and identify candidate functional SNPs.
It provides insights into the potential pathogenesis of myopia, enhances our understanding of molecular regulation in other diseases, enables more precise SNP identification and microscopic localization of genomic regulatory elements, and improves the accuracy of analytical results.
Smart Images

Figure CN2025092627_30072026_PF_FP_ABST
Abstract
Description
A method, device, medium, and program product for identifying candidate functional points (SNPs). Technical Field
[0001] This invention relates to the field of bioanalysis, and more specifically, to a method, apparatus, medium, and procedure for identifying candidate functional points (SNPs). Background Technology
[0002] Abnormalities in transcription and post-transcriptional biological processes can lead to serious health problems, including genetic diseases and cancer. For example, transcriptional errors can result in incorrect protein synthesis, while post-transcriptional abnormalities can lead to mRNA instability or incorrect splicing, thereby affecting RNA structure and function. Therefore, these processes are a key area of focus in biological research and medical applications.
[0003] Genomic regulatory elements are key regulators in gene transcription, expression, and translation. Dysfunction of these elements can lead to human diseases; however, their functions remain poorly understood. Previous research has shown that genomic regulatory elements overlap with the vast majority of genetic variations, particularly single nucleotide polymorphisms (SNPs), which are prone to aberrant regulation and likely modulate disease susceptibility. These functional SNPs are defined as "candidate functional SNPs (cfSNPs)." On one hand, cfSNPs can significantly disrupt the molecular binding between cis-acting elements and trans-acting factors, affecting transcriptional activity and potentially leading to the development of complex diseases. On the other hand, cfSNPs can alter RNA secondary structure, affecting molecular function and thus further participating in various human diseases. Therefore, analyzing the systemic regulation between SNPs, target gene regulatory elements, and the original phenotype at the post-transcriptional level is crucial.
[0004] Myopia has increased rapidly over the past few decades and is projected to continue to increase in the coming decades. Notably, approximately 80% to 90% of young people in East Asia have myopia. Individuals with common myopia (CM) may develop high myopia (HM), and a significant proportion of these individuals with high myopia will develop pathological myopia (PM) or even eye complications in the coming decades. Recent studies have shown that myopic individuals have a 100-fold increased risk of developing myopic macular degeneration (MMD), a three-fold increased risk of retinal detachment (RD), a three-fold increased risk of developing posterior subcapsular cataract (PSC), and almost double the risk of open-angle glaucoma (OAG). Research indicates that the etiology of myopia involves both environmental and genetic factors, with genetic factors influencing the pathophysiology of myopia at different levels. Recent studies have identified a large number of SNPs and genes associated with myopia. Despite these advances, a comprehensive understanding of the molecular regulation of myopia at the genetic level remains lacking. Summary of the Invention
[0005] This invention aims to address at least one of the technical problems existing in the prior art. To this end, this invention provides a method, apparatus, medium, and program product for identifying candidate functional SNPs; the method of this invention, by identifying cfSNPs during transcription and post-transcriptional processes, provides valuable insights into the potential pathogenesis of myopia and also offers opportunities to understand the molecular regulation of other diseases.
[0006] The first aspect of this application discloses a method for identifying candidate functional points (SNPs), the method comprising:
[0007] 101. Obtain the SNPs to be tested;
[0008] 102. The SNPs to be tested are located in the Ref and Alt sequence pairs for transcriptional and / or post-transcriptional genomic regulatory elements, resulting in several sequence pairs;
[0009] 103. When the SNPs are located at the transcriptional genomic regulatory elements, SNPs at the transcriptional level are extracted; when the SNPs are located at the post-transcriptional gene regulatory elements, SNPs at the post-transcriptional level are extracted.
[0010] 104, output the SNPs at the transcriptional level and / or the SNPs at the post-transcriptional level as the candidate functional SNPs.
[0011] In some embodiments, the method for obtaining the transcriptional SNPs includes: extracting the number of regulatory proteins that bind to the sequence pairs and Ref and Alt sequences; and calculating the number of proteins that bind to the Ref sequence, denoted as Num. Ref The number of proteins bound to the Alt sequence is denoted as Num. Alt According to the Num Ref and Num Alt The SNPs at the transcriptional level were obtained by screening the ratio of the two values and were denoted as the first cfSNPs.
[0012] In some embodiments, the method for obtaining SNPs at the transcriptional level further includes: screening sequence pairs located at binding motifs and their position information from the plurality of sequence pairs; calculating a first ratio of the number of binding motifs in the Ref sequence that overlap only with the binding motif and the number of binding motifs that do not overlap with the binding motif, based on the position information; calculating a second ratio of the number of binding motifs in the Alt sequence that overlap only with the binding motif and the number of binding motifs that do not overlap with the binding motif; calculating a change value based on the first ratio and the second ratio; and screening SNPs with a change value less than a second threshold as SNPs at the transcriptional level, denoted as the second cfSNPs. The plurality of sequences refers to natural numbers greater than 1.
[0013] In some embodiments, the method for obtaining the transcriptional SNPs further includes: taking the intersection of the first cfSNPs and the second cfSNPs to obtain the transcriptional SNPs.
[0014] In some embodiments, the method for obtaining the SNPs at the post-transcriptional level includes: 201, extracting RNA secondary structure features based on Ref and Alt sequence pairs; the secondary structure features include B, E, H, I, M, and S subunits, as well as the number, length, and position of each subunit; 202, calculating the structural differences between the Ref and Alt sequences using Euclidean distance to obtain difference values for the B, E, H, I, M, and S subunits, as well as the number, length, and position of each subunit; 203, selecting the SNPs at the post-transcriptional level based on the difference values.
[0015] In some embodiments, the transcriptional genome regulatory elements include any one or more of the following: promoters, enhancers, promoter flanking regions, open chromatin regions, and CTCF binding sites; the posttranscriptional genome regulatory elements include any one or more of the following: 3'UTR, 5'UTR, exons, introns, and long non-coding RNAs.
[0016] A second aspect of this application discloses a method for identifying candidate functional mutation sites related to myopia, the method comprising:
[0017] 301. Obtain myopia-related mutation sites or SNPs;
[0018] 302, The mutation site or SNP is located in the Ref and Alt sequence pairs of transcriptional and / or posttranscriptional genomic regulatory elements to obtain several sequence pairs;
[0019] 303. When located at transcriptional genomic regulatory elements, extract mutation sites or SNPs at the transcriptional level; when located at post-transcriptional genomic regulatory elements, extract mutation sites or SNPs at the post-transcriptional level.
[0020] 304, output the mutation sites or SNPs at the transcriptional and / or post-transcriptional levels as the candidate functional mutation sites or SNPs.
[0021] In some embodiments, the type of myopia includes any one or more of the following: ordinary myopia, high myopia, pathological myopia, refractive error, and visual impairment; the transcriptional genome regulatory element includes any one or more of the following: promoter, enhancer, promoter flanking region, open chromatin region, and CTCF binding site; the posttranscriptional genome regulatory element includes any one or more of the following: exon, intron, and long non-coding RNA.
[0022] A third aspect of this application discloses a computer device, the device comprising: a memory and a processor; the memory being used to store a computer program; and the processor executing the computer program to implement the steps of the above-described method.
[0023] The fourth aspect of this application discloses a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.
[0024] The fifth aspect of this application discloses a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.
[0025] This application has the following beneficial effects:
[0026] 1. This application innovatively discloses a method for identifying candidate functional SNPs at the transcriptional and / or post-transcriptional levels, comprehensively integrating transcriptional and post-transcriptional regulation, and providing for the first time a macroscopic perspective on potential genetic molecular disorders of myopia. At the transcriptional level, FC value and P... T cfSNPs were obtained by screening based on the values. To address the heterogeneity of molecular binding during transcription, an index (Fold Chang, Fisher's exact test) was designed to quantify protein gain or loss, thereby identifying candidate functional SNPs (cfSNPs). At the post-transcriptional level, Euclidean distance was used to quantify the RNA secondary structure in the Ref and Alt sequences to obtain difference values, and cfSNPs were obtained based on these difference values.
[0027] 2. This application innovatively develops a comprehensive process for assessing the global and local effects of SNPs on variable regions of RNA structure, providing an effective method for evaluating RNA accessibility and stability.
[0028] 3. The method proposed in this application can be used to study the possible genetic mechanisms of SNPs with undefined functions in other human diseases, thereby enhancing the interpretability of disease-related SNPs.
[0029] 4. This application innovatively develops a new genome mapping method, which screens sequence pairs that bind to motifs located on transcriptional gene regulatory elements within sequence pairs. This method abandons the macroscopic mapping method in existing databases that only locates the genome. Compared with the mapping method in existing databases that only locates the genome without locating genomic regulatory elements, the mapping method in this application is more microscopic, precise, and accurate, and the results of subsequent analysis and evaluation are also more accurate.
[0030] In summary, this application provides a detailed analysis of the one-to-one relationships between SNPs, genomic regulatory elements, and myopia based on precise alignments. SNP-mediated molecular dysregulation was further investigated by assessing structural heterogeneity in molecular binding and RNA secondary structure through computational analysis. cfSNPs were also identified during transcription and post-transcriptional processes. This study provides valuable insights into the potential pathogenesis of myopia and offers opportunities to understand the molecular regulation of other diseases. Attached Figure Description
[0031] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0032] Figure 1 is a schematic flowchart of a method provided in the first aspect of an embodiment of the present invention;
[0033] Figure 2 is a schematic flowchart of a method provided in the second aspect of an embodiment of the present invention;
[0034] Figure 3 is a schematic diagram of the candidate function (SNP) identification system provided in an embodiment of the present invention;
[0035] Figure 4 is a schematic diagram of the identification system for myopia-related candidate functional mutation sites provided in an embodiment of the present invention;
[0036] Figure 5 is a schematic diagram of a computer device provided in an embodiment of the present invention;
[0037] Figure 6 is a schematic diagram of the architecture of an exemplary computing device provided in an embodiment of the present invention;
[0038] Figure 7 is a schematic diagram of the storage medium provided in an embodiment of the present invention;
[0039] Figure 8 is a fine-map of myopia-related SNPs on whole-genome regulatory elements provided in this embodiment of the invention; wherein, Figure 8A shows the distribution of each genomic regulatory element during transcription and post-transcriptional processes; Figure 8B shows the percentage of one-to-one pairings during transcription and post-transcriptional processes; Figure 8C shows the length of the genomic regulatory elements covering myopia-related SNPs, with the black horizontal lines in each violin representing the median length of each element; Figure 8D shows the SNP density of each regulatory element at the transcriptional and post-transcriptional levels. The dotted lines show the average SNP density of the entire human genome;
[0040] Figure 9 is a schematic diagram illustrating the heterogeneity of SNP-mediated transcription factor binding among the four regulatory elements during transcription provided in this embodiment of the invention; wherein, Figure 9A shows the heterogeneity of binding between TF and enhancer, Figure 9B shows the heterogeneity of binding between TF and open chromatin, Figure 9C shows the heterogeneity of binding between TF and promoter, and Figure 9D shows the heterogeneity of binding between TF and promoter flanking regions; P T The value was calculated using Fisher's exact test, and FC was calculated using multiple changes.
[0041] Figure 10 illustrates the SNP-mediated global RNA structural variable regions on genomic regulatory elements during post-transcriptional processes, as provided in this embodiment of the invention. Figure 10A shows the length of the RNA structural variable regions within the genomic regulatory elements; Figure 10B shows the percentage of relation pairs illustrating the proximal and distal effects mediated by SNPs on genomic regulatory elements; Figure 10C shows the percentage of relation pairs with significant RNA structural variable regions on genomic regulatory elements; Figure 10D shows the percentage of relation pairs in the CM covering significant RNA structural variable regions; and Figure 10E shows the percentage of relation pairs in the HM covering significant RNA structural variable regions.
[0042] Figure 11 illustrates the analysis of important RNA structural variable regions based on RNA subunits provided in this embodiment of the invention. Figure 11A shows the RNA secondary structure of selenocysteine transfer RNA (stRNA) obtained by Sfold, with visualizations of the stRNA structural conformations predicted by Sfold (left) and downloaded from RNASTRAND (right). Figure 11B visualizes the six subunits of the RNA secondary structure. Figure 11C shows the changes in the single-stranded or double-stranded state of RNA in important RNA structural variable regions at SNP sites on RNA subunits. Figure 11D shows the detection of cfSNPs that can affect RNA subunits based on Euclidean distance. The threshold (1.65) is represented by a horizontal line.
[0043] Figure 12 is a schematic diagram of the structural heterogeneity of FGF10 mediated by rs339501 provided in the embodiment of the present invention, wherein Figure 12A shows the base pair probabilities of the Ref and Alt structures; Figure 12B shows the Ref and Alt structures of FGF10;
[0044] Figure 13 is a schematic diagram of the changes in the spatial conformation and molecular binding of ZNF891 mediated by rs905224 according to the embodiments of the present invention. Figure 13A shows the secondary structures of ZNF891's Ref and Alt RNAs, with the Ref and Alt alleles marked in blue and red, respectively. Figure 13B shows the molecular binding of ZNF891 with GAPDH and PSME3, and Figure 13C shows the molecular binding of ZNF891 with GAPDH and PSME3. The black dashed box represents the actual structure of ZNF891, and the molecules marked in yellow represent regulatory proteins. Detailed Implementation
[0045] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0046] In some of the processes described in the specification, claims, and accompanying drawings of this invention, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0047] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0048] Figure 1 is a schematic flowchart of a method for identifying candidate functional points (SNPs) according to an embodiment of the present invention. Specifically, the method includes the following steps:
[0049] 101: Obtain the SNPs to be tested;
[0050] In some embodiments, SNP stands for Single Nucleotide Polymorphism, which refers to a variation in a single nucleotide (base pair) in a genomic DNA sequence. This variation can be a substitution (e.g., A to G), an insertion, or a deletion. SNPs are one of the most common forms of genetic variation in the genome, and their distribution in a population is polymorphic, meaning that different individuals may have different nucleotides at the same location. In some embodiments, the SNP to be tested is derived from a subject. The terms "subject," "test subject," or "sample to be tested" as used herein refer to any animal (e.g., a mammal), including but not limited to humans, non-human primates, rodents, etc., who will become the recipient of a specific treatment. Generally, the terms "subject" and "patient" are used interchangeably herein when referring to human subjects. Preferably, the subject is a human.
[0051] 102: Locate the SNPs to be tested to transcriptional and / or post-transcriptional genomic regulatory elements in the Ref and Alt sequence pairs to obtain several sequence pairs;
[0052] In some embodiments, the Ref and Alt sequence pairs corresponding to each SNP are retrieved. Specifically, several sequence pairs are obtained by retrieving Mbp upstream and downstream of the Ref and Alt sequences, where M is a natural number greater than 1; the range of M is 10-35, preferably 20. Specifically, Ref represents the wild type, and Alt represents the mutant type.
[0053] 103: When located at transcriptional genomic regulatory elements, extract SNPs at the transcriptional level; when located at post-transcriptional gene regulatory elements, extract SNPs at the post-transcriptional level.
[0054] In some embodiments, the method for obtaining the transcriptional SNPs includes: extracting the number of regulatory proteins that bind to the sequence pairs and Ref and Alt sequences; and calculating the number of proteins that bind to the Ref sequence, denoted as Num. Ref The number of proteins bound to the Alt sequence is denoted as Num. Alt According to the Num Ref and Num Alt The SNPs at the transcriptional level are obtained by screening based on the ratio, and are denoted as the first cfSNPs. If the ratio is not equal to a first threshold, the SNP to be tested is output as the first cfSNP. In a more specific embodiment, the first threshold ranges from 0.8 to 1.5; preferably 1.
[0055] In some embodiments, the method for obtaining SNPs at the transcriptional level further includes: screening sequence pairs located at binding motifs and their position information from the plurality of sequence pairs; calculating a first ratio of the number of binding motifs in the Ref sequence that overlap only with the binding motif and the number of binding motifs that do not overlap with the binding motif, based on the position information; calculating a second ratio of the number of binding motifs in the Alt sequence that overlap only with the binding motif and the number of binding motifs that do not overlap with the binding motif; calculating a change value based on the first ratio and the second ratio; and screening SNPs with a change value less than a second threshold as SNPs at the transcriptional level, denoted as second cfSNPs.
[0056] In some embodiments, the method for obtaining the transcriptional SNPs further includes: taking the intersection of the first cfSNPs and the second cfSNPs to obtain the transcriptional SNPs.
[0057] In some embodiments, the transcriptional genome regulatory element includes any one or more of the following: promoter, enhancer, promoter flanking region, open chromatin region, CTCF binding site;
[0058] In some embodiments, the method for obtaining the SNPs at the post-transcriptional level includes: 201, extracting RNA secondary structure features based on Ref and Alt sequence pairs; the secondary structure features include B, E, H, I, M, and S subunits, as well as the number, length, and position of each subunit; 202, calculating the structural differences between the Ref and Alt sequences using Euclidean distance to obtain difference values for the B, E, H, I, M, and S subunits, as well as the number, length, and position of each subunit; 203, selecting the SNPs at the post-transcriptional level based on the difference values.
[0059] In some embodiments, between 201 and 202, the method further includes standardizing the number, length, and position of each subunit, wherein the standardization method is as follows:
[0060] Where, N i,j Represents any 6×2 vector, where Nori∈(Num,len,Loc), j∈(B,E,H,I,M,S) are standardized values ranging from 0 to 1; B, E, H, I, M, and S represent convex loop, outer loop, hairpin loop, inner loop, multi-branched loop, and stem, respectively; R i∈(Num,Len,Loc) and A i∈(Num,Len,Loc) These represent the ref and alt RNA secondary structures induced by SNPs, respectively. RNA secondary structure specifically refers to the number, length, and position of each type of subunit within the RNA secondary structure.
[0061] In some embodiments, the quantization method in 202 includes:
[0062] M represents the feature of each RNA subunit, N refers to subunits B, E, H, I, M, and S, Nor(R) and Nor(A) represent the standardized value of each RNA subunit feature, respectively; and represent the difference value obtained after quantification. B, E, H, I, M, and S represent convex loop, outer loop, hairpin loop, inner loop, multi-branched loop, and stem, respectively.
[0063] In some embodiments, if the difference value exceeds a third threshold, the result indicating that the SNP has an impact on RNA structural changes is output. The third threshold is determined by assessing the lower quartile of RNA subunits where significant changes occur.
[0064] In some embodiments, the posttranscriptional genomic regulatory elements include any one or more of the following: 3'UTR, 5'UTR, exons, introns, and long non-coding RNAs.
[0065] 104: Output the SNPs at the transcriptional level and / or the SNPs at the post-transcriptional level as the candidate functional SNPs.
[0066] In some embodiments, the method further includes: identifying RNA structurally variable regions of posttranscribed gene regulatory elements in the Ref and Alt sequence pairs; extracting RNA secondary structure features from the RNA structurally variable regions; wherein the RNA structurally variable regions include global RNA structurally variable regions and / or local RNA structurally variable regions.
[0067] In some embodiments, SNPs in the globally variable RNA structural region are classified into proximal allosteric effects and distal allosteric effects. If the SNP is located in the globally variable RNA structural region and affects the folding of that region, the output is classified as a proximal allosteric effect; if the SNP is not located in the globally variable RNA structural region but affects the folding of that region, the output is classified as a distal allosteric effect. Specifically, a proximal allosteric effect refers to affecting a small surrounding region, and in molecular experiments, only the gene or other genes currently affected by the SNP are considered; a distal allosteric effect refers to affecting genes other than the current gene, and in molecular experiments, other genes are considered first.
[0068] In some embodiments, the evaluation results or prediction results may be in the form of paper or electronic reports, but are not limited to those obtained by intelligent machines based on the relevant data of the test subjects. These results are for reference only and are not considered as final diagnostic results.
[0069] In this embodiment, the threshold is obtained through training on training set samples. It can be a specific threshold or an interval range. The specific form is not specifically limited in this embodiment.
[0070] Figure 2 shows a flowchart of a method for identifying myopia-related candidate functional mutation sites disclosed in the second aspect of this application. The method includes:
[0071] 301. Obtain myopia-related mutation sites or SNPs;
[0072] 302, The mutation site or SNP is located in the Ref and Alt sequence pairs of transcriptional and / or posttranscriptional genomic regulatory elements to obtain several sequence pairs;
[0073] 303. When located at transcriptional genomic regulatory elements, extract mutation sites or SNPs at the transcriptional level; when located at post-transcriptional genomic regulatory elements, extract mutation sites or SNPs at the post-transcriptional level.
[0074] 304, output the mutation sites or SNPs at the transcriptional and / or post-transcriptional levels as the candidate functional mutation sites or SNPs.
[0075] In some embodiments, the type of myopia includes any one or more of the following: common myopia (CM), high myopia (HM), pathological myopia (PM), refractive error (RE), and visual impairment (VD); the transcriptional genomic regulatory element includes any one or more of the following: promoter, enhancer, promoter flanking region, open chromatin region, and CTCF binding site; the posttranscriptional genomic regulatory element includes any one or more of the following: exon, intron, and long non-coding RNA.
[0076] Figure 5 is a schematic diagram of a computer device provided in an embodiment of the present invention. As shown in Figure 5, the device 2000 may include: one or more processors 2010 and one or more memories 2020; wherein, the memory stores computer-readable code, which, when run by the one or more processors, can execute the method described above.
[0077] The processor in this embodiment can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, operations, and logic block diagrams disclosed in this embodiment. The general-purpose processor can be a microprocessor or any conventional processor, and can be based on an x86 or ARM architecture.
[0078] In general, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of embodiments of this disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0079] For example, the methods or apparatus according to embodiments of this disclosure can also be implemented using the architecture of the computing device 3000 shown in FIG. 6. As shown in FIG. 6, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. Storage devices in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for processing and / or communication of the methods provided in this disclosure, as well as program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, the architecture shown in FIG. 6 is merely exemplary, and one or more components in the computing device shown in FIG. 6 may be omitted as needed when implementing different devices.
[0080] This invention also provides a computer-readable storage medium, as shown in FIG7, which is a schematic diagram of a storage medium 4000 provided in an embodiment of this invention. The computer storage medium 4020 stores computer-readable instructions 4010. When the computer-readable instructions 4010 are executed by a processor, the method according to the embodiments of this disclosure described with reference to the above figures can be performed. The computer-readable storage medium in the embodiments of this disclosure can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), Synchronous Link Dynamic Random Access Memory (SLDRAM), and Direct Memory Bus Random Access Memory (DR RAM). It should be noted that the memory used in the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0081] This disclosure also provides a computer program product or system, including a computer program that, when executed by a processor, implements the steps of the above-described method.
[0082] In some embodiments, this embodiment also discloses a candidate function (SNP) identification system, as shown in FIG3, the system comprising:
[0083] The first acquisition module 401 is used or configured to acquire SNPs to be tested;
[0084] The first sequence pair extraction module 402 is used or configured to locate the SNPs to be tested in the Ref and Alt sequence pairs of transcriptional and / or post-transcriptional genomic regulatory elements to obtain several sequence pairs;
[0085] SNP extraction module 403 is used or configured to extract SNPs at the transcriptional level when located at transcriptional genomic regulatory elements; and to extract SNPs at the post-transcriptional level when located at post-transcriptional gene regulatory elements.
[0086] The candidate functional SNP output module 404 is used or configured to output the SNPs at the transcriptional level and / or the SNPs at the post-transcriptional level as the candidate functional SNPs.
[0087] In some embodiments, this embodiment also discloses a system for identifying myopia-related candidate functional mutation sites, as shown in Figure 4. The system includes:
[0088] The second acquisition module 501 is used or configured to acquire myopia-related mutation sites or SNPs.
[0089] The second sequence pair extraction module 502 is used or configured to locate the mutation site or SNP in the Ref and Alt sequence pairs for transcriptional and / or post-transcriptional genomic regulatory elements, thereby obtaining several sequence pairs;
[0090] The mutation site or SNP extraction module 503 is used or configured to extract mutation sites or SNPs at the transcriptional level when located at transcriptional genomic regulatory elements; and to extract mutation sites or SNPs at the post-transcriptional level when located at post-transcriptional genomic regulatory elements.
[0091] The candidate functional mutation site or SNP extraction module 504 is used or configured to output the mutation sites or SNPs at the transcriptional level and / or post-transcriptional level as the candidate functional mutation sites or SNPs.
[0092] Specific implementation examples:
[0093] 1. Materials and Methods:
[0094] 1.1 Data Collection:
[0095] Myopia-related SNPs were collected from dbGap, the GWAS catalog, and published literature. Based on the population distribution of myopia-related SNPs in the 1000 Genomes Project (GRCh38), we performed quality control (QC) on the samples and genotypes. SNPs were filtered according to the following selection criteria: East Asian, two alleles in NCBI, minor allele frequency (MAF) > 5%, Hardy-Weinberg equilibrium P < 0.01, recall > 75%, and genotyping rate > 75%. 343 SNPs were obtained for the following analysis. The start and termination sites of eight genomic regulatory elements, including open chromatin regions (OCRs), CTCF binding sites (CTCFBSs), enhancers, promoters, promoter flanking regions (PFRs), exons, introns, and non-coding RNA (ncRNA) transcripts, were derived from ENSEMBL (v102), and their sequences were extracted using NCBI's Refseq. The 3' UTR and 5' UTR sequences were obtained using the BioMart tool in ENSEMBL.
[0096] 1.2 Fine localization of myopia-related SNPs on 10 genomic regulatory elements:
[0097] Fine mapping was performed using the classic read alignment tool Bowtie2. Primary sequences of genomic regulatory elements were considered long reference reads. The 30bp flanking regions upstream and downstream of these SNPs were considered short alignment sequences. Strict parameters were then set using "--n-ceil C,3--np 0--end-to-end-a--score-min C,0" to avoid mismatches for each seed. After fine mapping, paired reference (Ref) and alternative (Alt) sequences were constructed based on the alleles of the SNPs. The alignment software included Bowtie1, Bowtie2, and BLAST, with Bowtie2 being preferred.
[0098] 1.3 Assessing changes in molecular binding at the transcriptional level:
[0099] We extracted 20 bp upstream and downstream of the SNP-induced paired Ref and Alt sequences and obtained the enhancers, open chromatin, promoters, and promoter flanking regions of these sequences from the HumanTFDB database, while CTCF was obtained from the CTCFBSDB database. For protein binding, we obtained the number of regulatory proteins (TF and CTCF), protein symbols, and the start and end positions of the binding motifs. First, we used fold change (FC) to obtain the changes in the number of protein binding sites mediated by myopia-related SNPs:
[0100] Num Ref and Num AltFC represents the number of proteins bound to the Ref and Alt sequences, respectively. We define FC > 1 as a protein binding loss, and conversely, FC < 1 as a protein binding gain. We then apply Fisher's exact test to determine changes in the protein binding motif, as shown below: P T =(a / b) / (c / d)
[0101] a represents the number of binding motifs that overlap only with the Ref sequence, b represents the number of binding motifs that do not overlap with the Ref sequence, c represents the number of binding motifs that overlap only with the Alt sequence, and d represents the number of binding motifs that do not overlap with the Alt sequence. Here, when FC is not equal to 1 and P T A score <0.05 indicates that myopia-related SNPs significantly affect protein binding, and these SNPs are identified as cfSNPs. Furthermore, we define the protein with the highest score as the "leader" protein.
[0102] 1.4 Identifying global RNA structural variable regions after transcription:
[0103] At the post-transcriptional level, we used RNAsnp to identify the structural variable regions of five regulatory elements, including the 3'UTR, 5'UTR, exons, introns, and lncRNA transcripts. Based on thresholds set by RNAsnp, P... PT A value <0.2 indicates that SNPs may have a significant impact on local RNA structure. Here, we define these regions as "global RNA structurally variable regions". If an SNP is located in a RNA structurally variable region and affects the folding of that region, these SNPs are defined as "proximal allosteric effects". Otherwise, if an SNP is not located in a RNA structurally variable region but affects the folding of surrounding areas, these SNPs are defined as "distal allosteric effects".
[0104] 1.5 Quantifying changes in the variable region of RNA structure using RNA subunits:
[0105] We evaluated the changes in RNA subunits within important global RNA structural variable regions induced by SNPs. First, we predicted RNA secondary structures that were as consistent as possible with the actual folding state. Cluster analysis was performed based on 1000 possible structures to obtain representative structures for multiple clusters. By comparing the minimum free energy (MFE) of representative RNA structures in each cluster, we selected the structure with the lowest MFE as the most probable structure. Then, the six RNA subunits for each RNA were extracted using the RNAsmc method: the convex loop (B), outer loop (E), hairpin loop (H), inner loop (I), multi-branched loop (M), and stem (S). Next, we developed a computational pipeline and quantified the structural heterogeneity (Euclidean distance) of the global RNA structural variable regions using Euclidean distance (based on RNA subunits). R,AWe obtained the features of RNA subunits in the global RNA structural variable region, namely the number, length, and position of each subunit in the paired Ref and Alt structures. For these three dimensions, we constructed two corresponding 6×1 vectors: R i∈(Num,Len,Loc) =(R iB ,R iE ,R iH ,R iI ,R iM ,R is A i∈(Num,Len,Loc) =(A iB A iE A iH A iI A iM A is )
[0106] Here, 'i' represents the characteristic of the RNA subunit. B, E, H, I, M, and S refer to the six RNA subunits: convex loop, outer loop, hairpin loop, inner loop, multi-branched loop, and stem, respectively. R and A represent the secondary structures of Ref and AltRNA mediated by SNPs, respectively. Then, to normalize the size of the three characteristics of the RNA subunits, each type of subunit is normalized as follows:
[0107] N i,j Let represent any 6×2 vector, where Nori ∈ (Num, len, Loc), j ∈ (B, E, H, I, M, S) ranging from 0 to 1. Finally, we use Euclidean distance (Euclidean distance). R,A To quantify the difference between ref and alt structures:
[0108] M represents the feature of each sub-unit. N refers to the six sub-units. Nor(R) and Nor(A) represent the standardized values of each feature, respectively. In this study, Euc R,A The lower quartile was set as the significance threshold. If Euc R,A Above the lower quartile, we consider RNA subunits to exhibit significant changes. VARNAs are used to visualize RNA secondary structure.
[0109] 1.6 Same annotations for evaluating SNPs based on HM queues and computational pipelines:
[0110] A total of 10,348 highly myopic participants (worst eye SE < -6.00D) in the CAMS study were sequenced using the Twist HumanCoreExomeKit on a Berry Genomics Illumina NovaSeq 6000 sequencer. We obtained the genotypes and phenotypes of myopia-related SNPs from the CAMS study. Next, we used the Combined Annotation-Dependent Depletion score (CADD) to score the harmfulness of SNPs based on genotype and phenotype in the highly myopic cohort. Here, an SNP with a CADD score ≥ 10 was defined as a harmful SNP, potentially leading to loss of gene function. The Combined Annotation-Dependent Depletion score (CADD score) is used to assess and quantify single nucleotide variants (SNVs); a higher score indicates more harmful variants, i.e., a higher probability of pathogenicity.
[0111] 2. Results:
[0112] 2.1 Myopia-related SNPs are widely distributed at the transcriptional and post-transcriptional levels:
[0113] We obtained myopia-related SNPs from public resources, and after quality control, 343 SNPs were used for subsequent analysis (see Methods). To reveal the complete map of SNPs at the transcriptional and post-transcriptional levels, precise mapping was used to identify 636 relationship pairs formed by 263 SNPs, 10 genomic regulatory elements, and myopia (Figure 8A). Furthermore, these SNPs were associated with five phenotypes: refractive error (RE), common myopia (CM), high myopia (HM), pathological myopia (PM), and visual impairment (VD). During transcription, 84 SNPs were located in enhancers, open chromatin regions (OCRs), CTCF binding sites (CTCFBSs), promoters, and promoter flanking regions (PFRs), forming 90 relationship pairs. A total of 244 SNPs were mapped to five genomic regulatory elements: the 5'UTR, exons, introns, the 3'UTR, and lncRNAs, constituting 546 relationship pairs at the post-transcriptional level. Next, we found that 14.15% (90 / 636) and 85.85% (546 / 636) of the pairs were enriched during transcription and post-transcriptional processes, respectively (Fig. 8B).
[0114] To further assess the average distribution of SNPs across all genomic regulatory elements, we analyzed the transcriptional length of these elements and the density of SNPs within each element. The results showed that the median length of the longest transcript, LncRNA, was 95.53 times that of the shortest transcript, CTCFBS (Fig. 8C). We then further evaluated the enrichment of SNPs within each 1000 bp of each regulatory element. Compared to the average of one SNP per 1000 bp in the human genome, we observed highly distributed myopia-associated SNPs in the OCR, CTCFBS, 5'UTR, and exons, respectively (Fig. 8D). Analysis of myopia revealed that 81.76% (520 / 636) of the relationships were associated with myopia (CM). Furthermore, little was known about HM, PM, RE, and VD, with approximately 12.11% (77 / 636), 5.19% (33 / 636), 0.63% (4 / 636), and 0.31% (2 / 636), respectively. These data indicate that the distribution of myopia-related SNPs varies at the transcriptional and post-transcriptional levels as well as among genomic regulatory elements. These differences may be related to the severity of myopia and underlying molecular regulation.
[0115] 2.2 Scoring of SNP-induced molecular binding heterogeneity during transcription:
[0116] To investigate SNP-mediated molecular binding heterogeneity at the transcriptional level, we developed a computational pipeline that uses the fold-over test (FC) to assess changes in the number of binding proteins and applies Fisher's exact test to assess changes in the number of binding protein molecules. Here, a threshold P is used. T <0.05, FC not equal to 1, and found that 38.46% (5 / 13), 80% (4 / 5), 57.89% (11 / 19), and 44.73% (17 / 38) of SNPs could disrupt transcription factor binding to enhancers, open chromatin regions, promoters, and PFRs, respectively (Figure 9). For the CTCF protein, we found that approximately 13.33% (2 / 15) of SNPs could disrupt the CTCFBS interaction. In total, 43.33% (39 / 90) of the relationship pairs were able to disrupt binding affinity and recognize 46.43% (39 / 84) of SNPs as “cfSNPs” during transcription. In conclusion, open chromatin regions were significantly enriched with myopia-associated SNPs and cfSNPs, exhibiting high-density distribution and disrupted molecular binding.
[0117] To determine the potential influence of regulatory proteins associated with or potentially associated with myopia-related cfSNPs, we explored the molecular functions of the “leader” proteins in each relationship pair in published studies. Interestingly, most “leaders” were able to influence regulators of ocular tissues and structures. For example, approximately 7.69% (3 / 39) of cfSNPs induced changes in the REST binding motif, thereby affecting the fate of retinal ganglion cells (RGCs) in the developing retina. Approximately 7.69% (3 / 39) of cfSNPs induced changes in IRF1 binding, IRF1 being known to be expressed in retinal microglia and play a key role in microglia activation and retinal inflammation. Furthermore, 15.28% (6 / 39) of cfSNPs altered the binding motif site of SPI1, which has been reported to regulate microglia in the retina. In summary, we found abundant leader proteins in retinal inflammation that have been confirmed to be involved in the occurrence and development of myopia. Moreover, this can help us reveal potential protein regulators and understand how SNPs participate in the molecular regulation of myopia. Finally, we assessed the distribution of significantly altered relationships between myopia types at the transcriptional level. Over 75% (3 / 4) were associated with PM, while approximately 44.44% (4 / 9) and 41.56% (32 / 77) were associated with HM and CM (Figures 9A-D). The results also indicated that the impact of SNPs on myopia increases with the severity of myopia.
[0118] 2.3 Identification of SNP-mediated global RNA structural variable regions during post-transcriptional processes:
[0119] To identify SNP-induced RNA secondary structure heterogeneity during post-transcriptional processing, we initially used RNA snp to detect potential global RNA structural variable regions. PT <0.2. Here, since there is only one pair of relationships in the 5'UTR, this exon was chosen as a reference to compare the significance of differences in the average length of RNA structural variable regions. As shown in Figure 10A, the RNA structural variable regions of the 5'UTR, 3'UTR, and exon are relatively longer than those of introns and lncRNAs. Notably, myopia-related SNPs are not only highly enriched in the 5'UTR and exons, but also cause significant structural damage to these regulatory regions. Although myopia-related SNPs are not significantly enriched in the 3'UTR region, they still cause significant structural effects. This may be because the 3'UTR region exhibits a highly structured feature, and once disrupted, it has a significant impact on RNA secondary structure.
[0120] Furthermore, we assessed whether SNPs were located in regions of RNA structural variability. Definitions are shown in the methods. Statistically, approximately 97.54% (238 / 244) of the SNPs exhibited proximal allosteric effects, mapping to 91.96% (502 / 546) of relation pairs and all five regulatory elements, while only 12.70% (31 / 244) of the SNPs showed proximal allosteric effects. SNPs exhibited distal allosteric effects, mapping to 8.06% (44 / 546) of relation pairs and three regulatory elements: exons, introns, and lncRNAs (Figure 10B). This finding is consistent with a previous study that SNPs primarily affect local regions rather than having a global impact.
[0121] To further explore which genomic regulatory elements were significantly disrupted by SNPs, we analyzed the proportion of important RNA structurally variable regions within these elements. As shown in Figure 10C, among the nine pairs in the 3'UTR, 33.33% (3 / 9) exhibited significant RNA structurally variable regions induced by 37.50% (3 / 8) of the SNPs, while one pair in the 5'UTR did not show significant RNA structurally variable regions. Among the 35 pairs in exons, 22.86% (8 / 35) showed significant changes, mediated by 20.69% (6 / 29) of the SNPs. Furthermore, among the 308 pairs in introns, 14.61% (45 / 308) showed significant RNA structurally variable regions, which were affected by 16.52% (38 / 230) of the SNPs. Of the 193 relationship pairs of lncRNAs, 12.95% (25 / 193) showed significant RNA structural variable regions, induced by 14.56% (23 / 158) of the SNPs. Overall, we found that 14.84% (81 / 546) of the relationship pairs contained significant RNA structural variable regions induced by 20.08% (49 / 244) of the SNPs, which were identified as cfSNPs.
[0122] To characterize the role of RNA secondary structure in myopia, we analyzed the distribution of important variable regions of RNA structure. The three myopia types—RE, CM, and HM—exhibited significant SNP-mediated changes. Among these phenotypes, 1.23% (1 / 81) of the relationships were associated with RE, 16.05% (13 / 81) with HM, and 82.72% (67 / 81) with CM. As shown in Figures 10D and E, compared to CM, HM exhibited a higher proportion of important regions, such as the 3'UTR, exons, and introns, that were more susceptible to the influence of cfSNPs in HM.
[0123] 2.4 Quantification of the local stability of important RNA structural variable regions based on RNA subunits:
[0124] To further evaluate the changes in RNA stability and accessibility induced by cfSNPs in important RNA structural variable regions, we revealed variations based on the single-stranded or double-stranded state of RNA subunits. First, we designed a computational pipeline using Sfold to identify the most probable structures of 81 important RNA structural variable regions. Here, we obtained experimentally determined RNA secondary structures from RNASTRAND and used these structures as references to evaluate the accuracy of the predicted structures. For example, Sfold and RNASTRAND showed high consistency in predicting the RNA secondary structure of stRNA, achieving an RNAsmc score of 10 (Fig. 11A). Then, we derived the basic RNA subunits of the RNA structural variable regions: stem (S), inner loop (I), convex loop (B), multi-branched loop (M), hairpin loop (H), and outer loop (E) (Fig. 11B). The single-stranded and double-stranded states of RNA folds are closely related to RNA stability or RNA binding accessibility. By analyzing 81 RNA subunits with paired Ref and Alt structures induced by 35 cfSNPs in significantly variable RNA structural regions, we found that 53.09% (43 / 81) of these pairs showed changed pairing states. Specifically, 24.69% (20 / 81) changed from single-stranded to double-stranded, 28.40% (23 / 81) changed from double-stranded to single-stranded, and 46.91% (38 / 81) remained unchanged (Figure 11C). In the cases of these changed subunits, the double-stranded state (S represents the double-stranded state) and other single-stranded subunits (B, E, H, I, and M are single-stranded states) underwent significant changes.
[0125] Finally, to identify localized changes in significant RNA structural variable regions induced by cfSNPs at the posttranscriptional level, we performed a comprehensive comparison based on RNA subunits, including the number, length, and base composition of each subunit between Ref and Alt structures. Structural differences in computational analysis were quantified using Euclidean distance. The threshold for assessing significant changes in RNA subunits was defined as the lower quartile (Euclidean distance) of the observed structural changes. R,A =1.65). Here, 8.79% (48 / 546) of important RNA structural variable regions induced by 14.34% (35 / 244) of SNPs may undergo changes in RNA subunits, thereby affecting RNA secondary structure. These results suggest that myopia-related SNPs play an important role in the development of myopia-related diseases by affecting RNA stability.
[0126] This study establishes for the first time a comprehensive map of molecular dysregulation of genomic regulatory elements mediated by myopia-related SNPs. 636 relationship pairs were established for 263 SNPs, 10 genomic regulatory elements, and 5 genomic regulatory elements. A total of 82 myopia-related cfSNPs were identified across the entire genome, with 39 cfSNPs in 39 relationship pairs detected during transcription and 49 cfSNPs in 81 relationship pairs detected during transcription. Transcription was performed using FC and P... T We quantified the gain or loss of SNP-mediated binding proteins. At the post-transcriptional level, we further investigated important variable regions of RNA structure from a global perspective using RNAsnp. Furthermore, we devised a novel method to quantify changes in RNA subunits from a local perspective, reflecting RNA accessibility and stability. Additionally, we obtained genotypic and phenotypic information from a previously established high myopia cohort and assessed the harmfulness of SNPs based on CADD. In summary, this study reveals potential molecular regulation, enhances the interpretability of SNPs, and provides new insights into the genetic mechanisms of myopia.
[0127] During transcription, the results showed that, in computational analysis, myopia-associated SNPs may gain or lose T binding sites. In OCR, rs8110889 mapped to ENSR00001023878 weakened the interaction between FOXA1 and ENSR00001023878 (P < 0.05). T =6.28e-14, FC=1.38 (Figure 9B). Previous studies have shown that FoxA1 is closely related to signaling pathways in the vertebrate retina, suggesting that rs8110889 may alter molecular binding, disrupt retinal-related signaling pathways, and lead to myopia susceptibility. In the promoter region, the Alt allele (C) of rs7550232 in ENSR00000020131 can increase the number of binding motifs on FLI1 (P). T =1.86e-26, FC=0.75 (Figure 9C). Another previous study reported that fli1 can drive vascular endothelial gene expression and control eye development in zebrafish. These results suggest that the acquisition and loss of SNP-induced regulatory proteins in the transcriptome may affect gene function and contribute to the pathogenesis of myopia. Recent studies have also reported similar relationships. For example, the allele T of rs17079281 in the DCBLD1 promoter can create a YY1 binding site to suppress DCBLD1 gene expression levels, reducing the risk of lung cancer in the Chinese population. Furthermore, rs3101339 may disrupt the binding between TF and NEGR1 genes, leading to gene dysregulation and contributing to major depressive disorder. The discovery of SNP-induced protein motif loss or gain can reveal molecular interactions and provide potential intervention targets.
[0128] At the post-transcriptional level, we found that myopia-associated SNPs may also disrupt the secondary structure of genomic regulatory elements and lead to molecular dysregulation. Previous studies reported that FGF10, a well-known molecule associated with myopia risk factors, had its secondary structure disrupted by rs339501 (P...). PT =0.11, Euc R,A =1.99). As shown in Figure 12A, it is clear that U can reduce spatial accessibility, conformation, and promote stable structures. We then obtained the most probable structures in local regions. As shown in Figure 12B, the C allele is located in the inner loop, while the U allele is located on the stem. The results indicate that rs339501 can alter the single- or double-stranded nature of FGF10 and affect RNA stability. Similarly, rs905224 located on the 3'UTR may disrupt the structural stability of ZNF891 (Figure 13A). We then downloaded two binding proteins of ZNF891, GAPDH and PSME3, from the IntAct database. By examining the molecular interactions induced by rs905224, we found that changes in the RNA secondary structure of ZNF891 may affect the binding of these two proteins in 3D spatial conformation and result in different docking fractions (Figures 13B and 13C). Similarly, numerous previous studies have shown that SNP-mediated changes in RNA secondary structure may contribute to disease development. For example, the two SNPU22G and A56U cells in the 5'UTR of the FTL gene have been shown to alter the overall mRNA structure and are associated with ferritinemia-associated cataract syndrome. Another study showed that the alleles of rs27770 in the 3'UTR exhibit distinct minimum free energy (MFE) structures, significantly affecting mRNA stability and increasing cancer risk. These observations support the role of RNA secondary structure in genomic regulatory elements as a key factor in myopia. We developed a computational pipeline to identify cfSNPs during transcription and post-transcriptional processes. This is significant for understanding the potential regulatory mechanisms of myopia development.
[0129] In summary, we have achieved, for the first time, fine-mapped SNPs and established a comprehensive relationship between SNPs, genomic regulatory elements, and myopia. Based on our self-designed molecular binding and classic structural heterogeneity assessment algorithms, we identified a series of cfSNPs that can disrupt the regulatory function of genomic regulatory elements at both the transcriptional and post-transcriptional levels. Furthermore, we developed an RNA subunit heterogeneity assessment algorithm to further evaluate the impact of cfSNPs on RNA structural accessibility and stability. These results provide a broad perspective for future research on molecular regulation.
[0130] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0131] In general, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of embodiments of this disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0132] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0133] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0134] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0135] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0136] The exemplary embodiments of this disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art will understand that various modifications and combinations can be made to these embodiments or their features without departing from the principles and spirit of this disclosure, and such modifications should fall within the scope of this disclosure.
Claims
1. A method for identifying candidate functional points (SNPs), characterized in that, The method includes:
101. Obtain the SNPs to be tested; 102. The SNPs to be tested are located in the Ref and Alt sequence pairs for transcriptional and / or post-transcriptional genomic regulatory elements, resulting in several sequence pairs; 103. When the SNPs are located at the transcriptional genomic regulatory elements, SNPs at the transcriptional level are extracted; when the SNPs are located at the post-transcriptional gene regulatory elements, SNPs at the post-transcriptional level are extracted. 104, output the SNPs at the transcriptional level and / or the SNPs at the post-transcriptional level as the candidate functional SNPs.
2. The method for identifying candidate functional points (SNPs) according to claim 1, characterized in that, The method for obtaining the SNPs at the transcriptional level includes: extracting the number of regulatory proteins that bind to the Ref and Alt sequences of the sequence pairs; calculating the number of proteins that bind to the Ref sequence and denoting it as Num. Ref The number of proteins bound to the Alt sequence is denoted as Num. Alt According to the Num Ref and Num Alt The SNPs at the transcriptional level were obtained by screening the ratio of the values, and were denoted as the first cfSNPs.
3. The method for identifying candidate functional points (SNPs) according to claim 2, characterized in that, The method for obtaining SNPs at the transcriptional level further includes: screening sequence pairs located at binding motifs and their position information from the plurality of sequence pairs; calculating a first ratio of the number of binding motifs in the Ref sequence that overlap only with the binding motif and the number of binding motifs that do not overlap with the binding motif, based on the position information; calculating a second ratio of the number of binding motifs in the Alt sequence that overlap only with the binding motif and the number of binding motifs that do not overlap with the binding motif; calculating a change value based on the first ratio and the second ratio; and screening SNPs with a change value less than a second threshold as SNPs at the transcriptional level, denoted as the second cfSNPs.
4. The method for identifying candidate functional points (SNPs) according to claim 3, characterized in that, The method for obtaining SNPs at the transcriptional level further includes: taking the intersection of the first cfSNPs and the second cfSNPs to obtain the SNPs at the transcriptional level.
5. The method for identifying candidate functional points (SNPs) according to claim 1, characterized in that, The method for obtaining SNPs at the post-transcriptional level includes: 201, extracting RNA secondary structure features based on Ref and Alt sequence pairs; the secondary structure features include B, E, H, I, M, and S subunits, as well as the number, length, and position of each subunit; 202, using Euclidean distance to calculate the structural differences between the Ref and Alt sequences for the B, E, H, I, M, and S subunits, as well as the number, length, and position of each subunit, to obtain difference values; 203, selecting SNPs at the post-transcriptional level based on the difference values. The transcriptional genomic regulatory elements include any one or more of the following: promoters, enhancers, promoter flanking regions, open chromatin regions, and CTCF binding sites; optionally, the posttranscriptional genomic regulatory elements include any one or more of the following: 3'UTR, 5'UTR, exons, introns, and long non-coding RNAs.
6. A method for identifying candidate functional mutation sites related to myopia, characterized in that, The method includes:
301. Obtain myopia-related mutation sites or SNPs; 302, The mutation site or SNP is located in the Ref and Alt sequence pairs of transcriptional and / or posttranscriptional genomic regulatory elements to obtain several sequence pairs; 303. When located at transcriptional genomic regulatory elements, extract mutation sites or SNPs at the transcriptional level; when located at post-transcriptional genomic regulatory elements, extract mutation sites or SNPs at the post-transcriptional level. 304, output the mutation sites or SNPs at the transcriptional and / or post-transcriptional levels as the candidate functional mutation sites or SNPs.
7. The method for identifying myopia-related candidate functional mutation sites according to claim 6, characterized in that, The type of myopia includes any one or more of the following: ordinary myopia, high myopia, pathological myopia, refractive error, and visual impairment; optionally, the transcriptional genomic regulatory element includes any one or more of the following: promoter, enhancer, promoter flanking region, open chromatin region, and CTCF binding site; optionally, the posttranscriptional genomic regulatory element includes any one or more of the following: exon, intron, and long non-coding RNA.
8. A computer device, characterized in that, The device includes: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the steps of the method according to any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1-7.
10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method described in any one of claims 1-7.