A method, device, medium and program product for identifying candidate functional SNPs
By identifying candidate functional SNPs in the transcription and post-transcriptional process, and evaluating the heterogeneity of molecular binding and RNA secondary structure using computer analysis, the problem of insufficient understanding of myopia molecular regulation in the prior art is solved, and an in-depth understanding of the potential pathogenesis of myopia and the evaluation of the impact on variable regions of RNA structure is achieved.
Patent Information
- Application Number
- CN202510096430.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The prior art lacks a comprehensive understanding of the regulation of myopia molecular, especially at the genetic level, and it is difficult to identify and understand the potential effects of candidate functional SNPs (cfSNPs) on myopia.
By identifying cfSNPs in transcription and post-transcriptional processes, using computer analysis to evaluate structural heterogeneity of molecular binding and RNA secondary structures, identify candidate functional SNPs, and provide insights into the potential pathogenesis of myopia.
A macro perspective on the potential genetic molecular dysregulation of myopia is achieved, providing a process for evaluating the global and local effects of variable regions of RNA structure, enhancing the interpretability of disease-related SNPs.
Smart Images

Figure CN119560011B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biological analysis, and more specifically, to a method, device, medium and program product for identifying candidate functional SNPs. Background Art
[0002] Abnormalities in transcriptional and post-transcriptional biological processes may lead to serious health problems, including genetic diseases and cancer. For example, transcriptional errors may lead to incorrect protein synthesis, while post-transcriptional abnormalities may lead to mRNA instability or incorrect splicing, which in turn affects the structure and function of RNA. Therefore, these processes are key areas of research in biology and medical applications.
[0003] Genomic regulatory elements are key regulators in the processes of gene transcription, expression, and translation. Molecular dysfunction of genomic regulatory elements may lead to human diseases, however, the functions of these elements remain poorly understood. Previous studies have shown that genomic regulatory elements overlap with the vast majority of genetic variations, especially single nucleotide polymorphisms (SNPs), which are prone to abnormal regulation and are likely to modulate disease susceptibility. The above-mentioned functional SNPs are also defined as "candidate functional SNPs (cfSNPs)". On the one hand, cfSNPs can significantly disrupt the molecular binding between cis-acting elements and trans-acting factors, which affects transcriptional activity and may lead to the development of complex diseases. On the other hand, cfSNPs can change RNA secondary structure and affect molecular function, which is further involved in a variety of human diseases. Therefore, it is crucial to dissect the systemic regulation between SNPs, target gene regulatory elements, and the original phenotype at the post-transcriptional level of genes.
[0004] Myopia has increased rapidly over the past few decades and is expected to continue to increase in the coming decades. Notably, approximately 80% to 90% of young people in East Asia suffer from myopia. Individuals with normal myopia (CM) may develop high myopia (HM), and a large proportion of these high myopic individuals will develop pathological myopia (PM) or even ocular complications in the coming decades. Recent studies have shown that myopic patients have a 100-fold higher risk of myopic macular degeneration (MMD), a three-fold higher risk of retinal detachment (RD), a three-fold higher risk of posterior subcapsular cataract (PSC), and an almost doubled risk of open-angle glaucoma (OAG). Studies have shown that the etiology of myopia involves both environmental and genetic factors, and genetic factors can affect the pathophysiology of myopia at different levels. Recent studies have discovered a large number of SNPs and genes associated with myopia. Despite these advances, a comprehensive understanding of the molecular regulation of myopia at the genetic level is still lacking. Summary of the invention
[0005] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention provides a method, device, medium and program product for identifying candidate functional SNPs; the method of the present invention provides valuable insights into the potential pathogenesis of myopia by identifying cfSNPs in transcriptional and post-transcriptional processes, and also provides opportunities for understanding the molecular regulation of other diseases.
[0006] In a first aspect, the present application discloses a method for identifying candidate functional SNPs, the method comprising: 101, obtaining SNPs to be tested; 102, locating the SNPs to be tested at transcriptional and / or post-transcriptional genomic regulatory elements in Ref and Alt sequence pairs to obtain several sequence pairs; 103, when located at transcriptional genomic regulatory elements, extracting SNPs at the transcriptional level; when located at post-transcriptional gene regulatory elements, extracting SNPs at the post-transcriptional level; 104, outputting the SNPs at the transcriptional level and / or the SNPs at the post-transcriptional level as the candidate functional SNPs.
[0007] In some embodiments, the method for obtaining SNPs at the transcriptional level includes: extracting the number of regulatory proteins bound to the sequence pair and the Ref and Alt sequences; calculating the number of proteins bound to the Ref sequence and recording it as Num Ref Calculate the number of proteins bound to the Alt sequence and record it as Num Alt ; According to the Num Ref and Num Alt The SNPs at the transcriptional level are obtained by ratio screening and recorded as the first cfSNPs.
[0008] In some embodiments, the method for obtaining SNPs at the transcriptional level further includes: screening sequence pairs located at the binding motif and their position information from the several sequence pairs; calculating a first ratio of the number of binding motifs that only overlap with the binding motif and those that do not overlap with the binding motif in the Ref sequence according to the position information, and calculating a second ratio of the number of binding motifs that only overlap with the binding motif and those that do not overlap with the binding motif in the Alt sequence; calculating a change value according to the first ratio and the second ratio; screening SNPs whose change values are less than the second threshold as the SNPs at the transcriptional level, recorded as the second cfSNPs. The number refers to a natural number greater than 1.
[0009] In some embodiments, the method for obtaining the SNPs at the transcriptional level further includes: taking the intersection of the first cfSNPs and the second cfSNPs to obtain the SNPs at the transcriptional level.
[0010] In some embodiments, the method for obtaining SNPs at the post-transcriptional level includes: 201, extracting RNA secondary structure features based on Ref and Alt sequence pairs; the secondary structure features include B, E, H, I, M, S subunits, and the number, length and position of each subunit; 202, using Euclidean distance to calculate B, E, H, I, M, S subunits, and the number, length and position of each subunit to quantify the structural differences between the Ref sequence and the Alt sequence to obtain difference values; 203, based on the difference values, screening to obtain the SNPs at the post-transcriptional level.
[0011] In some embodiments, the transcriptional genomic regulatory elements include any one or more of the following: promoters, enhancers, promoter flanking regions, open chromatin regions, CTCF binding sites; the post-transcriptional genomic regulatory elements include any one or more of the following: 3'UTR, 5'UTR, exons, introns, long non-coding RNA.
[0012] The second aspect of the present application discloses a method for identifying candidate functional mutation sites associated with myopia, the method comprising:
[0013] 301, obtaining myopia-related mutation sites or SNPs;
[0014] 302, locating the mutation site or SNP to the transcriptional and / or post-transcriptional genomic regulatory elements in the Ref and Alt sequence pairs to obtain a plurality of sequence pairs;
[0015] 303, when located in the transcriptional genomic regulatory element, extract the mutation site or SNP at the transcriptional level; when located in the post-transcriptional genomic regulatory element, extract the mutation site or SNP at the post-transcriptional level;
[0016] 304, output the mutation site or SNP at the transcriptional level and / or post-transcriptional level as the candidate functional mutation site or SNP.
[0017] In some embodiments, the type of myopia includes any one or more of the following: common myopia, high myopia, pathological myopia, refractive error, visual impairment; the transcriptional genomic regulatory elements include any one or more of the following: promoters, enhancers, promoter flanking regions, open chromatin regions, CTCF binding sites; the post-transcriptional genomic regulatory elements include any one or more of the following: exons, introns, long non-coding RNA.
[0018] The third aspect of the present application discloses a computer device, which includes: a memory and a processor; the memory is used to store a computer program; and the processor executes the computer program to implement the steps of the above method.
[0019] A fourth aspect of the present application discloses a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of the above method when executed by a processor.
[0020] A fifth aspect of the present application discloses a computer program product, including a computer program, which implements the steps of the above method when executed by a processor.
[0021] The present application has the following beneficial effects: 1. The present application innovatively discloses a method for identifying candidate functional SNPs at the transcriptional and / or post-transcriptional level, which comprehensively integrates transcriptional and post-transcriptional regulation and provides for the first time a macroscopic perspective of the potential genetic molecular disorders of myopia. T The cfSNPs were obtained by screening based on the difference values. Indicators (Fold Chang, Fisher's exact test) were designed to quantify the gain or loss of proteins in view of the heterogeneity of molecular binding during transcription, thereby identifying candidate functional SNPs (cfSNPs). At the post-transcriptional level, the Euclidean distance was used to quantify the RNA secondary structure in the Ref sequence and the Alt sequence to obtain the difference value, and cfSNPs were obtained based on the difference value.
[0022] 2. This application innovatively develops a process to comprehensively evaluate the global and local effects of SNPs on variable regions of RNA structure, providing an effective method for evaluating the accessibility and stability of RNA.
[0023] 3. The method proposed in this application can be used to study the possible genetic mechanisms of SNPs with undefined functions for other human diseases, thereby enhancing the interpretability of disease-related SNPs.
[0024] 4. The present application innovatively develops a new genome positioning method, which screens sequence pairs located on binding motifs on transcriptional gene regulatory elements within sequence pairs, abandoning the macro positioning method in existing databases that only locates the genome. Compared with the positioning method in existing databases that only locates the genome but not the genome regulatory elements, the positioning method in the present application is more microscopic, precise and accurate, and the results of subsequent analysis and evaluation are also more accurate.
[0025] In summary, this application provides a detailed analysis of the one-to-one relationship between SNPs, genomic regulatory elements, and myopia based on precise alignment. SNP-mediated molecular dysregulation was further investigated by evaluating structural heterogeneity in molecular binding and RNA secondary structure in silico analysis. cfSNPs were also identified in transcriptional and post-transcriptional processes. This study provides valuable insights into the potential pathogenesis of myopia and also provides opportunities to understand the molecular regulation of other diseases. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0027] Figure 1 is a schematic diagram of a method flow chart provided by the first aspect of an embodiment of the present invention;
[0028] Figure 2 is a schematic diagram of a method flow chart provided by the second aspect of an embodiment of the present invention;
[0029] Figure 3 Schematic diagram of a system for identifying candidate functional SNPs provided by an embodiment of the present invention;
[0030] Figure 4 Schematic diagram of a system for identifying candidate functional mutation sites associated with myopia provided by an embodiment of the present invention;
[0031] Figure 5 is a schematic diagram of a computer device provided by an embodiment of the present invention;
[0032] Figure 6 is a schematic diagram of the architecture of an exemplary computing device provided by an embodiment of the present invention;
[0033] Figure 7 is a schematic diagram of a storage medium provided by an embodiment of the present invention;
[0034] Figure 8 is a fine-scale mapping map of myopia-related SNPs on genome-wide regulatory elements provided by an embodiment of the present invention; wherein, Figure 8 A is the distribution of each genomic regulatory element in transcription and post-transcriptional processes; Figure 8 B is the percentage of one-to-one pairing during transcription and post-transcription; Figure 8 C is the length of the genomic regulatory elements covering the myopia-associated SNPs, and the black-marked horizontal line in each violin represents the median length of each element; Figure 8 D is the SNP density of each regulatory element at the transcriptional and post-transcriptional levels. The dotted line shows the average SNP density of the entire human genome;
[0035] Fig. 9 is a schematic diagram of the heterogeneity of SNP-mediated transcription factor binding between four regulatory elements in the transcription process provided by an embodiment of the present invention; wherein, Fig. 9 A is the heterogeneity of binding between TF and enhancer, Fig. 9 B is the heterogeneity of TF binding to open chromatin, Fig. 9 C is the heterogeneity of binding between TF and promoter, Fig. 9 D is the heterogeneity of binding between TF and promoter flanking regions; P T Values were calculated using Fisher’s exact test, and FC was calculated by fold change;
[0036] Fig.10 It is a global RNA structure variable region mediated by SNP on genomic regulatory elements in the post-transcriptional process provided by an embodiment of the present invention; wherein, Fig.10 A is the length of the variable region of RNA structure in the genomic regulatory element; Fig.10 B is the percentage of relationship pairs showing proximal and distal effects mediated by SNPs on genomic regulatory elements; Fig.10 C is the percentage of relationship pairs with important RNA structural variable regions on genomic regulatory elements; Fig.10 D is the percentage of relationship pairs covering significant RNA structural variable regions in CM; Fig.10 E is the percentage of relationship pairs covering significant RNA structural variable regions in HM;
[0037] Fig.11 The present invention provides an important RNA structure variable region analysis based on RNA subunits; wherein, Fig.11 A is the RNA secondary structure of selenocysteine transfer RNA (stRNA) obtained by Sfold, and the structural conformation of stRNA predicted by Sfold (left) and downloaded by RNASTRAND (right) is visualized. Fig.11 B is the visualization of the six subunits of RNA secondary structure. Fig.11 C is the change in the single-stranded or double-stranded state of the RNA in the important RNA structural variable region at the SNP site based on the RNA subunit. Fig.11 D is the detection of cfSNPs that can affect RNA subunits based on Euclidean distance. The threshold (1.65) is represented by the horizontal line;
[0038] Fig.12 Schematic diagram of the structural heterogeneity of FGF10 mediated by rs339501 provided in the embodiment of the present invention, wherein Fig.12 A is the base pair probability of Ref and Alt structures; Fig.12 B is the Ref and Alt structures of FGF10;
[0039] Fig.13 Schematic diagram of the change of ZNF891 spatial conformation and molecular binding mediated by rs905224 provided in the embodiment of the present invention, wherein Fig.13 A shows the Ref and Alt RNA secondary structures of ZNF891, with the Ref and Alt alleles marked in blue and red, respectively. Fig.13 B is the molecular binding of ZNF891 with GAPDH and PSME3. Fig.13 C shows the molecular binding of ZNF891 to GAPDH and PSME3. The black dotted box shows the actual structure of ZNF891, and the yellow-labeled molecules represent regulatory proteins. DETAILED DESCRIPTION
[0040] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0041] In some of the processes described in the specification and claims of the present invention and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent the order of precedence, and do not limit the "first" and "second" to be different types.
[0042] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0043] Figure 1 1 is a flow chart of a method for identifying a candidate functional SNP provided by an embodiment of the present invention. Specifically, the method comprises the following steps: 101: obtaining SNPs to be tested;
[0044] In some embodiments, the full name of SNP is Single Nucleotide Polymorphism, which refers to the variation of a single nucleotide (base pair) in a genomic DNA sequence. This variation may be a substitution (e.g., A becomes G), an insertion or a deletion. SNP is the most common form of genetic variation in the genome, and their distribution in the population is polymorphic, that is, different individuals may have different nucleotides at the same position. In some embodiments, the SNP to be tested comes from a subject, and the terms "subject" or "person to be tested" or "sample to be tested" used herein refer to any animal (e.g., mammal), including but not limited to humans, non-human primates, rodents, etc., which will become recipients of specific treatments. Typically, the terms "subject" and "patient" are used interchangeably herein when referring to human subjects. Preferably, the subject is a human.
[0045] 102: locating the SNPs to be tested to transcriptional and / or post-transcriptional genomic regulatory elements in the Ref and Alt sequence pairs to obtain a plurality of sequence pairs;
[0046] In some embodiments, based on each SNP, the Ref and Alt sequence pairs corresponding to the SNP are retrieved. Specifically, the Ref and Alt sequences are retrieved Mbp upstream and downstream to obtain several sequence pairs, where M is a natural number greater than 1; the range of M is 10-35, preferably 20. Specifically, Ref is a wild type, and Alt is a mutant.
[0047] 103: When located in transcriptional genomic regulatory elements, extract SNPs at the transcriptional level; when located in post-transcriptional gene regulatory elements, extract SNPs at the post-transcriptional level;
[0048] In some embodiments, the method for obtaining SNPs at the transcriptional level includes: extracting the number of regulatory proteins bound to the sequence pair and the Ref and Alt sequences; calculating the number of proteins bound to the Ref sequence and recording it as Num Ref Calculate the number of proteins bound to the Alt sequence and record it as Num Alt ; According to the Num Ref and Num Alt The ratio of is screened to obtain the SNPs at the transcriptional level, which are recorded as the first cfSNPs. If the ratio is not equal to the first threshold, the SNP to be tested is output as the first cfSNPs. In a more specific embodiment, the range of the first threshold includes 0.8-1.5; preferably 1.
[0049] In some embodiments, the method for obtaining SNPs at the transcriptional level also includes: screening sequence pairs located at the binding motif and their position information from the several sequence pairs; calculating a first ratio of the number of binding motifs that only overlap with the binding motif and that do not overlap with the binding motif in the Ref sequence based on the position information, and calculating a second ratio of the number of binding motifs that only overlap with the binding motif and that do not overlap with the binding motif in the Alt sequence; calculating a change value based on the first ratio and the second ratio; screening SNPs whose change values are less than a second threshold as SNPs at the transcriptional level, recorded as second cfSNPs.
[0050] In some embodiments, the method for obtaining the SNPs at the transcriptional level further includes: taking the intersection of the first cfSNPs and the second cfSNPs to obtain the SNPs at the transcriptional level.
[0051] In some embodiments, the transcriptional genomic regulatory element includes any one or more of the following: a promoter, an enhancer, a promoter flanking region, an open chromatin region, a CTCF binding site;
[0052] In some embodiments, the method for obtaining SNPs at the post-transcriptional level includes: 201, extracting RNA secondary structure features based on Ref and Alt sequence pairs; the secondary structure features include B, E, H, I, M, S subunits, and the number, length and position of each subunit; 202, using Euclidean distance to calculate B, E, H, I, M, S subunits, and the number, length and position of each subunit to quantify the structural differences between the Ref sequence and the Alt sequence to obtain difference values; 203, based on the difference values, screening to obtain the SNPs at the post-transcriptional level.
[0053] In some embodiments, between 201 and 202, the method further comprises standardizing the number, length and position of each subunit, and the standardization method is:
[0054]
[0055] in, represents any numerical 6 × 2 vector, is a normalized value ranging from 0 to 1; B, E, H, I, M, and S represent convex loop, outer loop, hairpin loop, inner loop, multi-branched loop, and stem, respectively; and They represent the ref and alt RNA secondary structures induced by SNPs, respectively. RNA secondary structure specifically refers to the number, length, and position of each type of subunit in RNA secondary structure.
[0056] In some embodiments, the quantification method in 202 includes:
[0057]
[0058] M represents the feature of each RNA subunit, N refers to B, E, H, I, M, and S subunits, Nor(R) and Nor(A) represent the standardized values of each RNA subunit feature, respectively; represents the difference value obtained after quantification; B, E, H, I, M, and S represent bulge loop, outer loop, hairpin loop, inner loop, multi-branch loop, and stem, respectively.
[0059] In some embodiments, if the difference value exceeds a third threshold, a result indicating that the SNP has an effect on RNA structural change is output. The third threshold is determined by evaluating the lower quartile of significant changes in RNA subunits.
[0060] In some embodiments, the post-transcriptional genomic regulatory element comprises any one or more of the following: 3'UTR, 5'UTR, exon, intron, long non-coding RNA.
[0061] 104: Output the SNPs at the transcriptional level and / or the SNPs at the post-transcriptional level as the candidate functional SNPs.
[0062] In some embodiments, the method further includes: identifying the RNA structure variable region of the post-transcriptional gene regulatory element in the Ref and Alt sequence pairs; extracting RNA secondary structure features from the RNA structure variable region; the RNA structure variable region includes a global RNA structure variable region and / or a local RNA structure variable region.
[0063] In some embodiments, the SNPs in the global RNA structure variable region are classified into proximal allosteric effects and distal allosteric effects; if the SNPs are located in the global RNA structure variable region and affect the folding of the region, the output belongs to the proximal allosteric effect; if the SNPs are not located in the global RNA structure variable region and affect the folding of the region, the output belongs to the distal allosteric effect. Specifically, the proximal allosteric effect refers to the effect on the surrounding small segment area, and only the current gene or other of the SNP is considered when doing molecular experiments; the distal allosteric effect refers to the effect on other genes other than the current gene, and other genes are considered first when doing molecular experiments.
[0064] In some embodiments, the evaluation results or prediction results include but are not limited to paper or electronic report forms. The results are only obtained by the intelligent machine based on the analysis of relevant data of the subjects, and are only used as a reference, not as the final diagnosis result.
[0065] In this embodiment, the threshold is obtained through training of training set samples, and it can be a specific threshold or an interval range. The specific form is not specifically limited in this embodiment.
[0066] like Figure 2 As shown, a flow chart of a method for identifying candidate functional mutation sites related to myopia is disclosed in the second aspect of the present application, and the method comprises:
[0067] 301, obtaining myopia-related mutation sites or SNPs;
[0068] 302, locating the mutation site or SNP to the transcriptional and / or post-transcriptional genomic regulatory elements in the Ref and Alt sequence pairs to obtain a plurality of sequence pairs;
[0069] 303, when located in the transcriptional genomic regulatory element, extract the mutation site or SNP at the transcriptional level; when located in the post-transcriptional genomic regulatory element, extract the mutation site or SNP at the post-transcriptional level;
[0070] 304, output the mutation site or SNP at the transcriptional level and / or post-transcriptional level as the candidate functional mutation site or SNP.
[0071] In some embodiments, the type of myopia includes any one or more of the following: common myopia (CM), high myopia (HM), pathological myopia (PM), refractive error (RE), visual impairment (VD); the transcriptional genomic regulatory elements include any one or more of the following: promoters, enhancers, promoter flanking regions, open chromatin regions, CTCF binding sites; the post-transcriptional genomic regulatory elements include any one or more of the following: exons, introns, long non-coding RNA.
[0072] Figure 5 is a schematic diagram of a computer device provided by an embodiment of the present invention, such as Figure 5 As shown, the device 2000 may include: one or more processors 2010, and one or more memories 2020; wherein the memories store computer-readable codes, and when the computer-readable codes are run by the one or more processors, the method described above may be executed.
[0073] The processor in this embodiment can be an integrated circuit chip with signal processing capabilities. The above processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The disclosed methods, operations and logic block diagrams in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc., which can be an X86 architecture or an ARM architecture.
[0074] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the disclosed embodiments are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general purpose hardware or controllers or other computing devices, or some combination thereof as non-limiting examples.
[0075] For example, the method or device according to the embodiment of the present disclosure may also be implemented by Figure 6 The architecture of the computing device 3000 shown in FIG. Figure 6 As shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as ROM 3030 or hard disk 3070, may store various data or files used for processing and / or communication of the method provided by the present disclosure and program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 6 The architecture shown is only exemplary and can be omitted according to actual needs when implementing different devices. Figure 6 One or more components of a computing device are shown.
[0076] The embodiment of the present invention also provides a computer-readable storage medium, such as Figure 7As shown, it is a schematic diagram of a storage medium 4000 provided in an embodiment of the present invention, and a computer readable instruction 4010 is stored on the computer storage medium 4020. When the computer readable instruction 4010 is executed by a processor, the method according to the embodiment of the present disclosure described with reference to the above figures can be executed. The computer readable storage medium in the embodiment of the present disclosure can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory. It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0077] The embodiments of the present disclosure also provide a computer program product or system, including a computer program, which implements the steps of the above method when executed by a processor.
[0078] In some embodiments, this embodiment also discloses a system for identifying candidate functional SNPs, such as Figure 3 As shown, the system comprises:
[0079] A first acquisition module 401 is used or configured to acquire SNPs to be tested;
[0080] A first sequence pair extraction module 402 is used or configured to locate the SNPs to be tested at transcriptional and / or post-transcriptional genomic regulatory elements in the Ref and Alt sequence pairs to obtain a plurality of sequence pairs;
[0081] SNP extraction module 403, used or configured to extract SNPs at the transcriptional level when located at the transcriptional genomic regulatory element; extract SNPs at the post-transcriptional level when located at the post-transcriptional gene regulatory element;
[0082] The candidate functional SNP output module 404 is used for or configured to output the SNPs at the transcriptional level and / or the SNPs at the post-transcriptional level as the candidate functional SNPs.
[0083] In some embodiments, this embodiment also discloses a system for identifying candidate functional mutation sites associated with myopia, such as Figure 4 As shown, the system comprises:
[0084] A second acquisition module 501 is used or configured to acquire myopia-related mutation sites or SNPs;
[0085] A second sequence pair extraction module 502 is used or configured to locate the mutation site or SNP to the transcriptional and / or post-transcriptional genomic regulatory elements in the Ref and Alt sequence pairs to obtain a plurality of sequence pairs;
[0086] A mutation site or SNP extraction module 503 is used or configured to extract mutation sites or SNPs at the transcriptional level when located at a transcriptional genomic regulatory element; and to extract mutation sites or SNPs at the post-transcriptional level when located at a post-transcriptional genomic regulatory element;
[0087] The candidate functional mutation site or SNP extraction module 504 is used for or configured to output the mutation site or SNP at the transcriptional level and / or post-transcriptional level as the candidate functional mutation site or SNP. Specific embodiment:
[0089] 1 Materials and Methods
[0090] 1.1 Collecting Data:
[0091] Myopia-associated SNPs were collected from dbGap, GWAS catalogs, and published literature. Based on the population distribution of myopia-associated SNPs in the 1000 Genomes Project (GRCh38), we performed quality control (QC) on samples and genotypes. SNPs were filtered according to the following selection criteria: East Asian, two alleles in NCBI, minor allele frequency (MAF)>5%, Hardy-Weinberg equilibrium P<0.01, recall rate>75%, and genotyping rate>75%. 343 SNPs were obtained for the following analysis. The start and end positions of eight genomic regulatory elements, including open chromatin regions (OCRs), CTCF binding sites (CTCFBS), enhancers, promoters, promoter flanking regions (PFRs), exons, introns, and noncoding RNA (ncRNA) transcripts, were derived from ENSEMBL (v102), and their sequences were extracted according to Refseq from NCBI. The sequences of 3'UTR and 5'UTR were obtained using the BioMart tool from ENSEMBL.
[0092] 1.2 Fine positioning of myopia-related SNPs on 10 genomic regulatory elements:
[0093] Fine mapping was performed using the classic read alignment tool Bowtie2. The primary sequences of genomic regulatory elements were considered as long reference reads. The 30bp flanking regions upstream and downstream of these SNPs were considered as short alignment sequences. Then, strict parameters were set using "--n-ceil C,3--np 0--end-to-end -a --score-min C,0" to avoid mismatches for each seed. After fine positioning, we constructed paired reference (Ref) and alternative (Alt) sequences based on the alleles of the SNP. The alignment software includes Bowtie1, Bowtie2, BLAST, preferably Bowtie2.
[0094] 1.3 Assessing changes in molecular binding at the transcriptional level:
[0095] We extracted 20bp upstream and downstream of the paired Ref and Alt sequences induced by the SNP, and obtained the binding motifs of enhancers, open chromatin, promoters, and promoter flanking regions of these sequences from the HumanTFDB database, while CTCF was from the CTCFBSDB database. For protein binding, we obtained the number of regulatory proteins (TF and CTCF), protein symbols, and the start and end positions of the binding motifs, respectively. First, we used fold change (FC) to obtain the change in the number of protein binding sites mediated by myopia-associated SNPs:
[0096]
[0097] and Represents the number of proteins bound by the Ref and Alt sequences, respectively. We defined FC>1 as a loss of protein binding, and conversely, FC<1 as a gain of protein binding. We then applied Fisher’s exact test to determine changes in protein binding motifs as follows:
[0098]
[0099] a is the number of binding motifs that overlap only with the Ref sequence, b is the number of binding motifs that do not overlap with the Ref sequence, c is the number of binding motifs that overlap only with the Alt sequence, and d is the number of binding motifs that do not overlap with the Alt sequence. T When the score was < 0.05, it indicated that the myopia-associated SNPs significantly affected protein binding, and these SNPs were identified as cfSNPs. In addition, we defined the proteins with the highest scores as “leader” proteins.
[0100] 1.4 Identification of variable regions of global RNA structure after transcription:
[0101] At the post-transcriptional level, we used RNAsnp to identify structurally variable regions of five regulatory elements, including 3'UTR, 5'UTR, exons, introns, and lncRNA transcripts. PT <0.2 indicates that the SNP may have a significant effect on the local RNA structure. Here, we define these regions as "global RNA structure variable regions". If the SNP is located in the RNA structure variable region and affects the folding of this region, these SNPs are defined as "proximal allosteric effects". Otherwise, if the SNP is not located in the RNA structure variable region and affects the surrounding folding, these SNPs are defined as "distal allosteric effects".
[0102] 1.5 Using RNA subunits to quantify changes in RNA variable regions:
[0103] We evaluated SNP-induced changes in RNA subunits within important global RNA structural variable regions. First, we predicted RNA secondary structures that were as consistent as possible with the true folding state. Cluster analysis was performed based on 1,000 possible structures to obtain representative structures of multiple clusters. By comparing the minimum free energy (MFE) of the representative RNA structures in each cluster, we selected the structure with the lowest free energy as the most likely structure. Then, the RNAsmc method was used to extract six RNA subunits of each RNA, namely, the bulge loop (B), outer loop (E), hairpin loop (H), inner loop (I), multi-branched loop (M) and stem (S) of each RNA. Next, we developed a computational pipeline and quantified the structural heterogeneity of the variable regions of the global RNA structure using Euclidean distance (based on RNA subunits) ( ). We obtained the characteristics of RNA subunits in the variable region of the global RNA structure, namely the number, length and position of each subunit in the paired Ref and Alt structures. For these three dimensions, we constructed two corresponding 6×1 vectors:
[0104]
[0105]
[0106] Among them, i represents the characteristics of RNA subunits. B, E, H, I, M and S refer to the six RNA subunits, namely the convex loop, outer loop, hairpin loop, inner loop, multi-branched loop and stem. R and A represent the Ref and AltRNA secondary structures mediated by SNP, respectively. Then, in order to standardize the size of the three characteristics of RNA subunits, each type of subunit is standardized as follows:
[0107]
[0108] represents any numerical 6×2 vector, ranges from 0 to 1. Finally, we use the Euclidean distance ( ) to quantify the difference between the ref and alt structures:
[0109]
[0110] M represents the characteristics of each subunit. N refers to six subunits. and denotes the standardized value of each feature. The lower quartile of was set as the significance threshold. Above the lower quartile, we considered the RNA subunits to exhibit significant changes. VARNA was used to visualize RNA secondary structure.
[0111] 1.6 Evaluate the same annotation of SNPs based on the HM cohort and computational pipeline:
[0112] A total of 10,348 high myopic participants (worst eye SE <-6.00D) in the CAMS study were sequenced on the Illumina NovaSeq6000 sequencer from Berry Genomics using the Twist Human Core Exome Kit. We obtained the genotypes and phenotypes of myopia-associated SNPs from the CAMS study. Next, we used combined annotation-dependent deletion (CADD) to score the harmfulness of SNPs based on the genotypes and phenotypes in the high myopia cohort. Here, SNPs with a CADD score ≥ 10 are defined as harmful SNPs that may lead to loss of gene function. Combined Annotation Dependent Depletion score (CADD score): used to evaluate and quantify single nucleotide variations (SNVs). The higher the score, the more harmful variants there are, that is, the higher the possibility of pathogenicity.
[0113] 2 Results:
[0114] 2.1 Myopia-related SNPs are widely distributed at the transcriptional and post-transcriptional levels:
[0115] We obtained myopia-associated SNPs from public resources and, after quality control, used 343 SNPs for subsequent analysis (see Methods). To reveal the complete map of SNPs at the transcriptional and post-transcriptional levels, we used precise mapping and identified 636 relationship pairs formed by 263 SNPs, 10 genomic regulatory elements, and myopia ( Figure 8A). In addition, these SNPs were associated with five phenotypes, namely refractive error (RE), common myopia (CM), high myopia (HM), pathological myopia (PM) and visual impairment (VD). In the transcriptional process, 84 SNPs were located in enhancers, open chromatin regions (OCRs), CTCF binding sites (CTCFBS), promoters and promoter flanking regions (PFRs), forming 90 relationship pairs. A total of 244 SNPs were located on five genomic regulatory elements, namely 5'UTR, exons, introns, 3'UTR and LncRNA, constituting 546 relationship pairs at the post-transcriptional level. Next, we found that 14.15% (90 / 636) and 85.85% (546 / 636) of the pairs were enriched in the transcriptional and post-transcriptional processes, respectively ( Figure 8 B).
[0116] To further evaluate the average distribution of SNPs in all genomic regulatory elements, we analyzed the transcript lengths of these elements and the density of SNPs in each genomic regulatory element. The results showed that the median length of the longest transcript LncRNA was 95.53 times the median length of the shortest transcript CTCFBS ( Figure 8 C). We then further evaluated the enrichment of SNPs in each regulatory element within every 1000 bp. Compared with the average of 1 SNP per 1000 bp in the human genome, we observed that myopia-associated SNPs were highly distributed in OCR, CTCFBS, 5'UTR, and exons ( Figure 8 D). By analyzing myopia, 81.76% (520 / 636) of the relationship pairs were found to be associated with CM. In addition, little is known about HM, PM, RE, and VD, which are approximately 12.11% (77 / 636), 5.19% (33 / 636), 0.63% (4 / 636), and 0.31% (2 / 636), respectively. These data suggest that the distribution of myopia-associated SNPs differs at the transcriptional and posttranscriptional levels and between genomic regulatory elements. These differences may be related to the severity of myopia and potential molecular regulation.
[0117] 2.2 Scoring SNP-induced molecular binding heterogeneity during transcription:
[0118] To investigate SNP-mediated molecular binding heterogeneity at the transcriptional level, we developed a computational pipeline that used the fold test (FC) to assess changes in the number of bound proteins and applied Fisher’s exact test to assess changes in the number of bound proteins. T<0.05, FC was not equal to 1, and 38.46% (5 / 13), 80% (4 / 5), 57.89% (11 / 19), and 44.73% (17 / 38) of the SNPs were found to be likely to disrupt transcription factor binding at enhancers, open chromatin regions, promoters, and PFRs, respectively ( Fig. 9 ). For CTCF protein, we found that about 13.33% (2 / 15) of SNPs could disrupt CTCFBS interactions. In total, 43.33% (39 / 90) of the relationship pairs were able to disrupt binding affinity, and 46.43% (39 / 84) SNPs were identified as “cfSNPs” during transcription. In summary, open chromatin regions were significantly enriched with myopia-associated SNPs and cfSNPs, showing high-density distribution and disruption of molecular binding.
[0119] To determine the possible effects of regulatory proteins associated with cfSNPs that are associated or potentially associated with myopia, we explored the molecular functions of the “leader” proteins of each relationship pair in published studies. Interestingly, most of the “leaders” were able to affect regulators of ocular tissues and structures. For example, about 7.69% (3 / 39) of cfSNPs caused changes in the REST binding motif, thereby affecting the fate of RGC retinal ganglion cells (RGCs) in the developing retina. About 7.69% (3 / 39) of cfSNPs caused changes in IRF1 binding, which is known to be expressed in retinal microglia and plays a key role in microglial activation and retinal inflammation. In addition, 15.28% (6 / 39) of cfSNPs changed the binding motif site of SPI1, which has been reported to regulate microglia in the retina. In summary, we found that there are abundant leader proteins in retinal inflammation, which have been confirmed to be associated with the occurrence and development of myopia. In addition, it can help us reveal potential protein regulators and understand how SNPs are involved in the molecular regulation of myopia. Finally, we evaluated the distribution of myopia types among the significantly changed relationship pairs at the transcriptional level. More than 75% (3 / 4) of them were related to PM, while approximately 44.44% (4 / 9) and 41.56% (32 / 77) were related to HM and CM ( Fig. 9 AD). The results also showed that the effect of SNP on myopia increased with the severity of myopia.
[0120] 2.3 Identification of SNP-mediated global RNA structural variable regions during post-transcriptional processes:
[0121] To identify SNP-induced RNA secondary structural heterogeneity during post-transcriptional processes, we initially employed RNA SNPs to detect potential global RNA structural variable regions, P PT<0.2. Here, since there is only one set of related pairs in the 5'UTR, this exon was selected as a reference to compare the significance of the average length differences of the variable regions of RNA structures. Fig.10 As shown in A, the RNA structural variable regions of 5'UTR, 3'UTR and Exon are relatively longer than introns and LncRNA. It is worth noting that the SNPs associated with myopia are not only highly enriched in 5'UTR and exons, but also cause huge structural damage to these regulatory regions. Although myopia-related SNPs are not significantly enriched in the 3'UTR region, they still cause significant structural effects. This may be due to the fact that the 3'UTR region presents highly structured characteristics, and once it is damaged, it has a great impact on the RNA secondary structure.
[0122] In addition, we evaluated whether the SNPs were located in the variable regions of RNA structure. The definitions are shown in the Methods. Statistically, approximately 97.54% (238 / 244) of the SNPs exhibited proximal allosteric effects, mapping to 91.96% (502 / 546) of the relationship pairs and all five regulatory elements, while only 12.70% (31 / 244) of the SNPs exhibited proximal allosteric effects. The SNPs showed distal allosteric effects, mapping to 8.06% (44 / 546) of the relationship pairs and three regulatory elements, namely, exons, introns, and lncRNAs ( Fig.10 B). This finding is consistent with a previous study showing that SNPs affect mainly local regions rather than having a global effect.
[0123] To further explore which genomic regulatory elements are significantly disrupted by SNPs, we analyzed the proportion of important RNA structural variable regions within these elements. Fig.10As shown in C, among the 9 pairs in 3'UTR, 33.33% (3 / 9) were induced by 37.50% (3 / 8) of SNPs and showed significant RNA structure variable regions, while 1 pair in 5'UTR did not show significant RNA structure variable regions. Among the 35 pairs in exons, 22.86% (8 / 35) showed significant changes, which were mediated by 20.69% (6 / 29) of SNPs. In addition, among the 308 pairs in introns, 14.61% (45 / 308) showed significant RNA structure variable regions, which were affected by 16.52% (38 / 230) SNPs. Among the 193 pairs in LncRNA, 12.95% (25 / 193) showed significant RNA structure variable regions, which were induced by 14.56% (23 / 158) of SNPs. Overall, we found that 14.84% (81 / 546) of relationship pairs contained significant RNA structural variable regions induced by 20.08% (49 / 244) of the SNPs, which were identified as cfSNPs.
[0124] To characterize the role of RNA secondary structure in myopia, we analyzed the distribution of important RNA structural variable regions. Three myopia phenotypes, RE, CM, and HM, showed significant SNP-mediated changes. Among these phenotypes, 1.23% (1 / 81) of the relationship pairs were associated with RE, 16.05% (13 / 81) with HM, and 82.72% (67 / 81) with CM. Fig.10 As shown in D and E, HM exhibited a higher proportion of important regions compared with CM, such as 3'UTR, exons, and introns that were more susceptible to cfSNPs in HM.
[0125] 2.4 Quantifying the local stability of important RNA structural variable regions based on RNA subunits:
[0126] To further evaluate changes in RNA stability and accessibility caused by cfSNPs in important RNA structural variable regions, we revealed changes based on the single-stranded or double-stranded state of the RNA subunit. First, we designed a computational pipeline to identify the most likely structures of 81 important RNA structural variable regions using Sfold. Here, we obtained experimentally determined RNA secondary structures from RNASTRAND and used these structures as references to evaluate the accuracy of the predicted structures. For example, Sfold and RNASTRAND showed high consistency in predicting RNA secondary structures of stRNAs, achieving an RNAsmc score of 10 ( Fig.11A). Then, we get the basic RNA subunits of the variable region of RNA structure, which are stem (S), inner loop (I), bulge loop (B), multi-branched loop (M), hairpin loop (H) and outer loop (E) ( Fig.11 B). The single-stranded and double-stranded states of RNA folding are closely related to RNA stability or accessibility of RNA binding. By analyzing the paired Ref and Alt structured RNA subunits of 81 related pairs induced by 35 cfSNPs in the significant RNA structure variable regions, we found that 53.09% (43 / 81) of these pairs showed a change in pairing state. Among them, 24.69% (20 / 81) changed from single-stranded to double-stranded, 28.40% (23 / 81) changed from double-stranded to single-stranded, and 46.91% (38 / 81) remained unchanged ( Fig.11 C). In the case of these altered subunits, the double-stranded state (S represents the double-stranded state) and other single-stranded subunits (B, E, H, I and M are single-stranded states) are significantly changed.
[0127] Finally, to identify local changes in RNA structural variable regions that are significant at the posttranscriptional level due to cfSNPs, we performed a comprehensive comparison based on RNA subunits, including the number, length, and base composition of each subunit between the Ref and Alt structures. The structural differences in the in silico analysis were quantified using Euclidean distance. The threshold for evaluating significant changes in RNA subunits was defined as the lower quartile of the observed structural changes ( =1.65). Here, 8.79% (48 / 546) of the important RNA structural variable regions induced by 14.34% (35 / 244) of the SNPs may undergo changes in RNA subunits, thereby affecting the secondary structure of RNA. These results suggest that myopia-associated SNPs play an important role in the occurrence of myopia-related diseases by affecting RNA stability.
[0128] This study established for the first time a comprehensive map of molecular dysregulation of genomic regulatory elements mediated by myopia-associated SNPs. A total of 636 relationship pairs were identified for 263 SNPs, 10 genomic regulatory elements, and 5 genomic regulatory elements. We identified a total of 82 myopia cfSNPs on regulatory elements across the genome, of which 39 cfSNPs with 39 relationship pairs were detected in the transcription process and 49 cfSNPs with 81 relationship pairs were detected in the transcription process. TSNP-mediated gain or loss of binding proteins was quantified. At the post-transcriptional level, we used RNAsNPs to further investigate important RNA structural variable regions from a global perspective. In addition, we designed a new method to quantify changes in RNA subunits from a local perspective, which can reflect the accessibility and stability of RNA. In addition, we obtained genotype and phenotypic information from a previously established high myopia cohort and assessed the harmfulness of SNPs based on CADD. In summary, this study revealed potential molecular regulation, enhanced the interpretability of SNPs, and provided new insights into the genetic mechanism of myopia.
[0129] In transcription, the results showed that the myopia-associated SNPs may gain or lose binding sites for T in computational analysis. In OCR, rs8110889 mapped to ENSR00001023878 could weaken the interaction between FOXA1 and ENSR00001023878 (P T =6.28e-14, FC=1.38, Fig. 9 B). Previous studies have shown that FoxA1 is closely associated with signaling pathways in the vertebrate retina, suggesting that rs8110889 may alter molecular binding, disrupt retina-related signaling pathways, and lead to myopia susceptibility. In the promoter region, the Alt allele of rs7550232 in ENSR00000020131 (C) can increase the number of binding motifs on FLI1 (P T =1.86e-26, FC=0.75, Fig. 9 C). Another previous study reported that fli1 can drive endothelial gene expression and control eye development in zebrafish. These results suggest that the gain and loss of regulatory proteins induced by SNPs in the transcriptome may affect gene function and contribute to the pathogenic factors of myopia. Recent studies have also reported similar relationship pairs. For example, the allele T of rs17079281 in the DCBLD1 promoter can create a YY1 binding site to inhibit the expression level of the DCBLD1 gene and reduce the risk of lung cancer in the Chinese population. In addition, rs3101339 may disrupt the binding between TF and NEGR1 genes, leading to dysregulated gene expression and causing major depression. The discovery of SNP-induced protein motif loss or gain can reveal molecular interactions and provide potential intervention targets.
[0130] At the post-transcriptional level, we found that myopia-associated SNPs may also disrupt the secondary structure of genomic regulatory elements and lead to molecular dysregulation. A well-known molecule, FGF10, which is a risk factor associated with myopia, was previously reported to have its secondary structure disrupted by rs339501 (P PT =0.11, =1.99). Fig.12 As shown in A, it is obvious that U can reduce spatial accessibility, conformation and favor stable structures. Then, we also obtained the most likely structure of the local area. Fig.12 As shown in B, the C allele is located in the inner loop, while the U allele is located on the stem. The results show that rs339501 can change the single-stranded or double-stranded structure of FGF10 and affect RNA stability. Similarly, rs905224 located in the 3'UTR may disrupt the structural stability of ZNF891 ( Fig.13 A). Then, two binding proteins of ZNF891, GAPDH and PSME3, were downloaded from the IntAct database. By detecting the molecular interactions induced by rs905224, we found that changes in the RNA secondary structure of ZNF891 may affect the binding of these two proteins in 3D spatial conformation and have different docking scores ( Fig.13 B and Fig.13 C). Similarly, a large number of previous studies have shown that SNP-mediated changes in RNA secondary structure may contribute to the development of the disease. For example, two SNPs U22G and A56U in the 5'UTR of the FTL gene were shown to change the overall mRNA structure and were associated with hyperferritinemia cataract syndrome. Another study showed that alleles of rs27770 in the 3'UTR exhibited different minimum free energy (MFE) structures, which could significantly affect mRNA stability and increase cancer risk. These observations support that the RNA secondary structure of genomic regulatory elements is a key factor in myopia. We developed a computational pipeline to identify cfSNPs in transcriptional and post-transcriptional processes. This is of great significance for understanding the potential regulatory mechanisms of myopia development.
[0131] In conclusion, we have performed fine mapping for the first time and established a comprehensive relationship between SNPs, genomic regulatory elements, and myopia. Based on self-designed molecular binding and classical structural heterogeneity assessment algorithms, we identified a series of cfSNPs that can disrupt the regulatory functions of genomic regulatory elements at the transcriptional and post-transcriptional levels. In addition, we developed an RNA subunit heterogeneity assessment algorithm to further evaluate the effects of cfSNPs on RNA structural accessibility and stability. These results provide a broad vision for molecular regulation in future studies.
[0132] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment, or a part of a code, and the module, program segment, or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0133] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the disclosed embodiments are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general purpose hardware or controllers or other computing devices, or some combination thereof as non-limiting examples.
[0134] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0135] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0136] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0137] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0138] The exemplary embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. It should be understood by those skilled in the art that various modifications and combinations may be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.
Claims
1. A method for identifying candidate functional SNPs, characterized in that: The method comprises: 101, obtaining SNPs to be tested; 102, locating the SNPs to be tested to transcriptional and / or post-transcriptional genomic regulatory elements in the Ref and Alt sequence pairs to obtain a plurality of sequence pairs; 103. When the transcriptional genome regulatory element is located, extracting SNPs at the transcriptional level; the method for obtaining SNPs at the transcriptional level includes: extracting the number of regulatory proteins bound to the sequence pair and the Ref and Alt sequences; calculating the number of proteins bound to the Ref sequence and recording it as Num Ref Calculate the number of proteins bound to the Alt sequence and record it as Num Alt ; According to the Num Ref and Num Alt The first cfSNPs are obtained by screening the ratio of the first cfSNPs; the sequence pairs located in the binding motif and their position information are screened from the several sequence pairs; the first ratio of the number of binding motifs that only overlap with the binding motif and the number of binding motifs that do not overlap with the binding motif in the Ref sequence is calculated according to the position information, and the second ratio of the number of binding motifs that only overlap with the binding motif and the number of binding motifs that do not overlap with the binding motif in the Alt sequence is calculated; the change value is calculated according to the first ratio and the second ratio; the SNPs whose change value is less than the second threshold are screened as the second cfSNPs; the intersection of the first cfSNPs and the second cfSNPs is taken to obtain the SNPs at the transcriptional level; when located at the post-transcriptional gene regulatory element, the SNPs at the post-transcriptional level are extracted; 104, output the SNPs at the transcriptional level and / or the SNPs at the post-transcriptional level as the candidate functional SNPs.
2. The method for identifying candidate functional SNPs according to claim 1, characterized in that: The method for obtaining SNPs at the post-transcriptional level includes: 201, extracting RNA secondary structure characteristics based on Ref and Alt sequence pairs; the secondary structure characteristics include B, E, H, I, M, S subunits, and the number, length and position of each subunit; 202, using Euclidean distance to calculate B, E, H, I, M, S subunits, and the number, length and position of each subunit to quantify the structural differences between the Ref sequence and the Alt sequence to obtain difference values; 203, screening to obtain the SNPs at the post-transcriptional level based on the difference values.
3. The method for identifying candidate functional SNPs according to claim 1, characterized in that: The transcriptional genomic regulatory elements include any one or more of the following: promoter, enhancer, promoter flanking region, open chromatin region, CTCF binding site; the post-transcriptional genomic regulatory elements include any one or more of the following: 3'UTR, 5'UTR, exon, intron, long non-coding RNA.
4. A method for identifying candidate functional mutation sites associated with myopia, characterized in that: The method comprises: 301, obtaining myopia-related mutation sites or SNPs; 302, locating the mutation site or SNP to the transcriptional and / or post-transcriptional genomic regulatory elements in the Ref and Alt sequence pairs to obtain a plurality of sequence pairs; 303. When located at a transcriptional genomic regulatory element, extracting a mutation site or SNP at the transcriptional level; when located at a post-transcriptional genomic regulatory element, extracting a mutation site or SNP at the post-transcriptional level; the method for obtaining SNPs at the transcriptional level includes: extracting the number of regulatory proteins bound by the sequence pair to the Ref and Alt sequences; calculating the number of proteins bound to the Ref sequence and recording it as Num Ref Calculate the number of proteins bound to the Alt sequence and record it as Num Alt ; According to the Num Ref and Num Alt The first cfSNPs are obtained by screening the ratio of the ratio; the sequence pairs located in the binding motif and their position information are screened from the several sequence pairs; according to the position information, a first ratio of the number of binding motifs that only overlap with the binding motif and the number of binding motifs that do not overlap with the binding motif in the Ref sequence is calculated, and a second ratio of the number of binding motifs that only overlap with the binding motif and the number of binding motifs that do not overlap with the binding motif in the Alt sequence is calculated; the change value is calculated according to the first ratio and the second ratio; SNPs with a change value less than a second threshold are screened as the second cfSNPs; the intersection of the first cfSNPs and the second cfSNPs is taken to obtain the SNPs at the transcriptional level; 304, output the mutation site or SNP at the transcriptional level and / or post-transcriptional level as the candidate functional mutation site or SNP.
5. The method for identifying myopia-related candidate functional mutation sites according to claim 4, characterized in that: The types of myopia include any one or more of the following: common myopia, high myopia, pathological myopia, refractive error, and visual impairment; the transcriptional genomic regulatory elements include any one or more of the following: promoters, enhancers, promoter flanking regions, open chromatin regions, and CTCF binding sites; the post-transcriptional genomic regulatory elements include any one or more of the following: exons, introns, and long non-coding RNA.
6. A computer device, characterized in that: The device comprises: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the steps of the method according to any one of claims 1 to 5.
7. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Method for improving GWAS cause mutation positioning efficiency
CN111986731A
Method, system and equipment for comparing RNA structures based on RNA motif vectors
CN113936737A