Method for evaluating protein binding efficiency at gene transcription level, device, medium, and program product

By obtaining the Ref and Alt sequence pairs of SNPs, calculating the protein binding ratio and motif location information, the problem of accurately evaluating protein binding efficiency at the gene transcription level is solved, improving the accuracy of microscopic localization of genomic regulatory elements and the understanding of molecular regulation.

WO2026157070A1PCT designated stage Publication Date: 2026-07-30THE EYE HOSPITAL OF WENZHOU MEDICAL UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
THE EYE HOSPITAL OF WENZHOU MEDICAL UNIVERSITY
Filing Date
2025-04-30
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately evaluate protein binding efficiency at the gene transcription level, leading to molecular dysregulation of genomic regulatory elements that may cause human diseases. Furthermore, existing database localization methods lack microscopic precision.

Method used

By acquiring the SNP to be tested, retrieving the Ref and Alt sequence pairs, calculating the ratio of the number of regulatory proteins in the sequence pairs, and combining the motif position information, the protein binding efficiency change values ​​are screened and calculated to provide an accurate evaluation of protein binding efficiency.

Benefits of technology

It enables precise evaluation of protein binding efficiency at the gene transcription level, enhances the understanding of molecular regulation at the gene level, and improves the microscopic precision and accuracy of the localization method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025092628_30072026_PF_FP_ABST
    Figure CN2025092628_30072026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of biological analysis, and provides a method for evaluating protein binding efficiency at a gene transcription level, a device, a medium, and a program product. The method comprises: acquiring at least one SNP to be evaluated; retrieving a Ref and Alt sequence pair based on the SNP; extracting a count of regulatory proteins binding to Ref and Alt sequences of the sequence pair; calculating a count of proteins binding to the Ref sequence and denoting same as NumRef, and calculating a count of proteins binding to the Alt sequence and denoting same as NumAlt; and evaluating an impact of the SNP on protein binding efficiency according to a ratio of the NumRef to the NumAlt. A development process of the present invention studies systematic regulation between SNPs and transcriptional genomic regulatory elements, exploring in-depth how SNPs can be used to evaluate protein binding efficiency in a gene transcription context, identifying biological principles underlying biological data, and resolving related life science problems.
Need to check novelty before this filing date? Find Prior Art

Description

A method, apparatus, medium, and procedure for evaluating protein binding efficiency at the gene transcription level. Technical Field

[0001] This invention relates to the field of bioanalysis, and more specifically, to a method, apparatus, medium, and procedure for evaluating protein binding efficiency at the gene transcription level. Background Technology

[0002] In biology, transcription is a crucial step in regulating gene expression and is essential for controlling cell development and fate. Abnormalities in transcription can lead to serious health problems, including genetic diseases and cancer. For example, errors in transcription can result in incorrect protein synthesis, affecting protein structure and function. Transcription is the process of converting DNA into RNA, during which DNA segments are replicated into messenger RNA (mRNA). Transcription determines the genes expressed in a cell, thus determining the cell type and function. For instance, in muscle cells, genes related to muscle contraction are transcribed, while in nerve cells, genes related to transmitting nerve signals may be expressed.

[0003] Genomic regulatory elements are key regulators in gene transcription, expression, and translation. At the transcriptional level, enhancers, promoters, promoter flanking regions (PFRs), open chromatin regions (OCRs), and CTCFBS (CCCTC-binding factors) can bind to specific proteins, such as transcription factors (TFs) and CTCFs, thereby influencing transcriptional activity and regulating gene function. Molecular dysregulation of genomic regulatory elements can lead to human diseases; however, the functions of these elements remain poorly understood. Previous research has shown that genomic regulatory elements overlap with the vast majority of genetic variations, particularly single nucleotide polymorphisms (SNPs), which are prone to aberrant regulation and likely modulate disease susceptibility. Therefore, dissecting the systemic regulation between SNPs, target gene regulatory elements, and the original phenotype at the transcriptional level is crucial. Summary of the Invention

[0004] This invention aims to address at least one of the technical problems existing in the prior art. To this end, this invention provides a method, apparatus, medium, and procedure for evaluating protein binding efficiency at the gene transcription level. The method development process of this invention studies the systematic regulation between SNPs and transcriptional genomic regulatory elements, thereby deeply exploring how SNPs evaluate protein binding efficiency at the gene transcription level, uncovering the life laws hidden behind biological data, and solving related life science problems.

[0005] The first aspect of this application discloses a method for evaluating protein binding efficiency at the gene transcription level, the method comprising:

[0006] 101, Obtain at least one SNP to be tested;

[0007] 102, Retrieving Ref and Alt sequence pairs based on SNP;

[0008] 103, extract the number of regulatory proteins that bind to the Ref and Alt sequences of the sequence pairs;

[0009] 104. The number of proteins bound to the Ref sequence is denoted as Num. Ref The number of proteins bound to the Alt sequence is denoted as Num. Alt ;

[0010] 105, according to the stated Num Ref and Num Alt The ratio of SNPs to protein binding efficiency is used to evaluate the effect of SNPs on protein binding efficiency.

[0011] In some embodiments, 105 includes: if the ratio is not equal to a first threshold, outputting an evaluation result of the impact of the SNP to be tested on protein binding efficiency; the SNP to be tested is denoted as the first SNP.

[0012] In some embodiments, 105 includes: if the ratio is greater than a first threshold, outputting an evaluation result indicating that the SNP to be tested has a low impact on protein binding efficiency; the SNP to be tested is denoted as the second SNP; if the ratio is less than the first threshold, outputting an evaluation result indicating that the SNP to be tested has a high impact on protein binding efficiency; the SNP to be tested is denoted as the third SNP.

[0013] In some embodiments, the method further includes: 106, screening sequence pairs from the plurality of sequence pairs that bind motifs located on transcriptional gene regulatory elements; 107, calculating a first ratio of the number a of binding motifs that overlap only with the Ref sequence and the number b of binding motifs that do not overlap with the Ref sequence based on the position information of the binding motifs, and calculating a second ratio of the number c of binding motifs that overlap only with the Alt sequence and the number d of binding motifs that do not overlap with the Alt sequence; calculating the change value of the protein binding motif based on the first ratio and the second ratio; screening SNPs whose change value is less than a second threshold as fourth SNPs, and outputting the evaluation result of the fourth SNP's impact on protein binding efficiency.

[0014] In some embodiments, the method further includes: taking the intersection between the first SNP and the fourth SNP, and outputting the evaluation result of the SNP in the intersection affecting the protein binding efficiency.

[0015] In some embodiments, the method further includes: taking the intersection between the second SNP and the fourth SNP, and outputting a prediction result that the SNP in the intersection has low protein binding efficiency; taking the intersection between the third SNP and the fourth SNP, and outputting a prediction result that the SNP in the intersection has high protein binding efficiency.

[0016] In some embodiments, the gene regulatory element includes: an enhancer, an open chromatin region, a promoter, and a promoter flanking region; the regulatory protein includes transcription factors and CTCFs.

[0017] A second aspect of this application discloses a computer device, the device comprising: a memory and a processor; the memory being used to store a computer program; and the processor executing the computer program to implement the steps of the above-described method.

[0018] A third aspect of this application discloses a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.

[0019] The fourth aspect of this application discloses a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.

[0020] This application has the following beneficial effects:

[0021] 1. This application innovatively discloses a method for evaluating protein binding efficiency at the gene transcription level. This method retrieves Ref and Alt sequence pairs based on SNPs and calculates the FC value based on the number of regulatory proteins binding to the Ref and Alt sequence pairs. The evaluation result of the SNP on protein binding efficiency is given based on the FC value. Furthermore, the quantitative result of whether the SNP has a high or low impact on protein binding efficiency is given based on the magnitude of the FC value. Even further, this application also calculates P based on the number of binding motifs that overlap only with or do not overlap with the Ref and Alt sequences. T Value, based on FC value and P T The relationships between values ​​are further used to screen for factors affecting protein binding efficiency. This approach allows for a better understanding of the relationship between SNPs and protein binding at the transcriptional level, enhancing our understanding of molecular regulation at the gene level.

[0022] 2. This application innovatively develops a new genome mapping method to screen sequence pairs that bind to motifs located on transcriptional gene regulatory elements. It abandons the macroscopic mapping method in existing databases that only locates the genome. Compared with the mapping method in existing databases that only locates the genome without locating genomic regulatory elements, the mapping method in this application is more microscopic, precise and accurate, and the results of subsequent analysis and evaluation are also more accurate. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 is a schematic flowchart of a method provided in the first aspect of an embodiment of the present invention;

[0025] Figure 2 is a schematic diagram of a system for evaluating protein binding efficiency at the gene transcription level provided in an embodiment of the present invention.

[0026] Figure 3 is a schematic diagram of a computer device provided in an embodiment of the present invention;

[0027] Figure 4 is a schematic diagram of the architecture of an exemplary computing device provided in an embodiment of the present invention;

[0028] Figure 5 is a schematic diagram of the storage medium provided in an embodiment of the present invention;

[0029] Figure 6 is a fine localization map of myopia-related SNPs on genome-wide regulatory elements provided in the embodiments of the present invention;

[0030] Figure 6A shows the distribution of each genomic regulatory element during transcription and post-transcriptional processes; Figure 6B shows the percentage of one-to-one pairings during transcription and post-transcriptional processes; Figure 6C shows the length of genomic regulatory elements covering myopia-related SNPs, with the black-marked horizontal lines in each violin representing the median length of each element; Figure 6D shows the SNP density of each regulatory element at the transcriptional and post-transcriptional levels. The dotted lines indicate the average SNP density across the entire human genome.

[0031] Figure 7 is a schematic diagram illustrating the heterogeneity of SNP-mediated TF binding among the four regulatory elements during transcription provided in this embodiment of the invention; wherein, Figure 7A shows the heterogeneity of TF binding to enhancers, Figure 7B shows the heterogeneity of TF binding to OCRs, Figure 7C shows the heterogeneity of TF binding to promoters, and Figure 7D shows the heterogeneity of TF binding to PFRs; P T The values ​​are calculated using Fisher's exact test, and FC is calculated by multiple changes. Detailed Implementation

[0032] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0033] In some of the processes described in the specification, claims, and accompanying drawings of this invention, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0035] Figure 1 is a schematic flowchart of a method for evaluating protein binding efficiency at the gene transcription level according to an embodiment of the present invention. Specifically, the method includes the following steps:

[0036] 101: Obtain at least one SNP to be tested;

[0037] SNP stands for Single Nucleotide Polymorphism, which refers to a variation in a single nucleotide (base pair) in the sequence of genomic DNA. This variation can be a substitution (e.g., A to G), an insertion, or a deletion. SNPs are one of the most common forms of genetic variation in the genome, and their distribution in the population is polymorphic, meaning that different individuals may have different nucleotides at the same location. In some embodiments, the SNP to be tested is derived from a subject. The terms "subject," "test subject," or "sample to be tested" as used herein refer to any animal (e.g., a mammal), including but not limited to humans, non-human primates, rodents, etc., who will become a recipient of a specific treatment. Generally, the terms "subject" and "patient" are used interchangeably herein when referring to a human subject. Preferably, the subject is a human.

[0038] 102: Retrieving Ref and Alt sequence pairs based on SNP;

[0039] In some embodiments, the Ref and Alt sequence pairs corresponding to each SNP are retrieved. Specifically, several sequence pairs are obtained by retrieving Mbp upstream and downstream of the Ref and Alt sequences, where M is a natural number greater than 1; the range of M is 10-35, preferably 20. Specifically, Ref represents the wild type, and Alt represents the mutant type.

[0040] 103: Extract the number of regulatory proteins that bind to the Ref and Alt sequences of the sequence pairs;

[0041] 104: The number of proteins bound to the Ref sequence is denoted as Num. Ref The number of proteins bound to the Alt sequence is denoted as Num. Alt ;

[0042] 105: According to the stated Num Ref and Num Alt The ratio of SNPs to protein binding efficiency is used to evaluate the effect of SNPs on protein binding efficiency.

[0043] In some embodiments, step 105 includes: if the ratio is not equal to a first threshold, outputting an evaluation result of the impact of the SNP to be tested on protein binding efficiency; the SNP to be tested is denoted as the first SNP. In a more specific embodiment, the first threshold ranges from 0.8 to 1.5; preferably 1.

[0044] In some embodiments, 105 includes: if the ratio is greater than a first threshold, outputting an evaluation result indicating that the SNP to be tested has a low impact on protein binding efficiency; the SNP to be tested is denoted as the second SNP; if the ratio is less than the first threshold, outputting an evaluation result indicating that the SNP to be tested has a high impact on protein binding efficiency; the SNP to be tested is denoted as the third SNP.

[0045] In this embodiment, the first threshold is obtained through training on training set samples. It can be a specific threshold or an interval range. The specific form is not specifically limited in this embodiment.

[0046] In some embodiments, the method further includes: 106, screening sequence pairs from the plurality of sequence pairs that bind motifs located on transcriptional gene regulatory elements; 107, calculating a first ratio of the number a of binding motifs that overlap only with the Ref sequence and the number b of binding motifs that do not overlap with the Ref sequence based on the position information of the binding motifs, and calculating a second ratio of the number c of binding motifs that overlap only with the Alt sequence and the number d of binding motifs that do not overlap with the Alt sequence; calculating the change value of the protein binding motif based on the first ratio and the second ratio; screening SNPs whose change value is less than a second threshold as fourth SNPs, and outputting the evaluation result of the fourth SNP's impact on protein binding efficiency. The plurality of sequences refers to natural numbers greater than 1.

[0047] In some embodiments, the method further includes: taking the intersection between the first SNP and the fourth SNP, and outputting the evaluation result of the SNP in the intersection affecting the protein binding efficiency.

[0048] In some embodiments, the method further includes: taking the intersection between the second SNP and the fourth SNP, and outputting a prediction result that the SNP in the intersection has low protein binding efficiency; taking the intersection between the third SNP and the fourth SNP, and outputting a prediction result that the SNP in the intersection has high protein binding efficiency.

[0049] In some embodiments, the gene regulatory element includes: an enhancer, an open chromatin region, a promoter, and a promoter flanking region; the regulatory protein includes transcription factors and CTCFs.

[0050] In some embodiments, the evaluation results or prediction results may be in the form of paper or electronic reports, but are not limited to those obtained by intelligent machines based on the relevant data of the test subjects. These results are for reference only and are not considered as final diagnostic results.

[0051] Figure 3 is a schematic diagram of a computer device provided in an embodiment of the present invention. As shown in Figure 3, the device 2000 may include: one or more processors 2010 and one or more memories 2020; wherein, the memory stores computer-readable code, which, when run by the one or more processors, can execute the method described above.

[0052] The processor in this embodiment can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, operations, and logic block diagrams disclosed in this embodiment. The general-purpose processor can be a microprocessor or any conventional processor, and can be based on an x86 or ARM architecture.

[0053] In general, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of embodiments of this disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.

[0054] For example, the methods or apparatus according to embodiments of this disclosure can also be implemented using the architecture of the computing device 3000 shown in FIG. 4. As shown in FIG. 4, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage devices in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for processing and / or communication of the methods provided in this disclosure, as well as program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, the architecture shown in FIG. 4 is only exemplary, and one or more components in the computing device shown in FIG. 4 may be omitted as needed when implementing different devices.

[0055] This invention also provides a computer-readable storage medium, as shown in FIG5, which is a schematic diagram of a storage medium 4000 provided in an embodiment of this invention. The computer storage medium 4020 stores computer-readable instructions 4010. When the computer-readable instructions 4010 are executed by a processor, the method according to the embodiments of this disclosure described with reference to the above figures can be performed. The computer-readable storage medium in the embodiments of this disclosure can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), Synchronous Link Dynamic Random Access Memory (SLDRAM), and Direct Memory Bus Random Access Memory (DR RAM). It should be noted that the memory used in the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0056] This disclosure also provides a computer program product or system, including a computer program that, when executed by a processor, implements the steps of the above-described method.

[0057] In some embodiments, this embodiment also discloses a system for evaluating protein binding efficiency at the gene transcription level, as shown in Figure 2, the system comprising:

[0058] Acquisition module 201 is used or configured to acquire at least one SNP to be tested;

[0059] Sequence pair retrieval module 202 is used or configured to retrieve Ref and Alt sequence pairs based on SNP;

[0060] Protein quantity extraction module 203 is used or configured to extract the quantity of regulatory proteins that bind to the sequence pairs with the Ref and Alt sequences;

[0061] Calculation module 204 is used or configured to calculate the number of proteins bound to the Ref sequence, denoted as Num. Ref The number of proteins bound to the Alt sequence is denoted as Num. Alt ;

[0062] Evaluation result output module 205, used or configured to output results based on the Num Ref and Num Alt The ratio of SNPs to protein binding efficiency is used to evaluate the effect of SNPs on protein binding efficiency.

[0063] In some embodiments, the evaluation result output module includes: if the ratio is not equal to a first threshold, outputting an evaluation result of the impact of the SNP to be tested on protein binding efficiency; the SNP to be tested is denoted as the first SNP.

[0064] In some embodiments, the evaluation result output module includes: if the ratio is greater than a first threshold, outputting an evaluation result indicating that the SNP to be tested has a low impact on protein binding efficiency; the SNP to be tested is denoted as the second SNP; if the ratio is less than the first threshold, outputting an evaluation result indicating that the SNP to be tested has a high impact on protein binding efficiency; the SNP to be tested is denoted as the third SNP.

[0065] In some embodiments, the system further includes: a sequence pair screening module, configured to screen sequence pairs from the plurality of sequence pairs that are binding motifs located on transcriptional gene regulatory elements; an evaluation result calculation module, configured to calculate, based on the position information of the binding motifs, a first ratio of the number of binding motifs a that overlap only with the Ref sequence and the number of binding motifs b that do not overlap with the Ref sequence, and a second ratio of the number of binding motifs c that overlap only with the Alt sequence and the number of binding motifs d that do not overlap with the Alt sequence; calculate the change value of the protein binding motif based on the first ratio and the second ratio; screen SNPs whose change value is less than a second threshold as fourth SNPs, and output the evaluation result of the fourth SNP's impact on protein binding efficiency.

[0066] In some embodiments, the system further includes: a first result prediction module, configured to take the intersection between the first SNP and the fourth SNP, and output an evaluation result of the SNPs in the intersection affecting protein binding efficiency; a second result prediction module, configured to take the intersection between the second SNP and the fourth SNP, and output a prediction result that the SNPs in the intersection have low protein binding efficiency; and take the intersection between the third SNP and the fourth SNP, and output a prediction result that the SNPs in the intersection have high protein binding efficiency.

[0067] Specific implementation examples:

[0068] 1. Materials and Methods:

[0069] 1.1 Data Collection: Myopia-related SNPs were collected from dbGap, the GWAS catalog, and published literature. Based on the population distribution of myopia-related SNPs in the 1000 Genomes Project (GRCh38), we performed quality control (QC) on the samples and genotypes. SNPs were filtered according to the following selection criteria: East Asian, two alleles in NCBI, minor allele frequency (MAF) > 5%, Hardy-Weinberg equilibrium P < 0.01, recall > 75%, and genotyping rate > 75%. 343 SNPs were obtained for the following analysis. The start and termination sites of eight genomic regulatory elements, including open chromatin regions (OCRs), CTCF binding sites (CTCFBSs), enhancers, promoters, promoter flanking regions (PFRs), exons, introns, and non-coding RNA (ncRNA) transcripts, were derived from ENSEMBL (v102), and their sequences were extracted using NCBI Refseq. Sequences of the 3'UTR and 5'UTR were obtained using the BioMart tool of ENSEMBL.

[0070] 1.2 Fine localization of myopia-related SNPs on 10 genomic regulatory elements:

[0071] Precise localization was performed using the classic sequence alignment tool Bowtie 2. Primary sequences of genomic regulatory elements were considered long reference reads. The 30bp flanking regions upstream and downstream of these SNPs were considered short alignment sequences. Strict parameters were then set using "--n-ceil C,3--np 0--end-to-end-a--score-min C,0" to avoid seed sequence mismatches. After precise localization, paired reference (Ref) and alternative (Alt) sequences were constructed based on the alleles of the SNPs. Alignment software included Bowtie1, Bowtie2, and BLAST, with Bowtie2 being preferred.

[0072] 1.3 Assessing changes in molecular binding at the transcriptional level:

[0073] We extracted 20 bp upstream and downstream of the SNP-induced paired Ref and Alt sequences and obtained the enhancers, open chromatin regions, promoters, and promoter flanking regions of these sequences from the HumanTFDB database, while CTCF was obtained from the CTCFBSDB database. For protein binding, we obtained the number of regulatory proteins (TF and CTCF), protein identifiers, and the start and end positions of the binding motifs. First, we used fold change (FC) to obtain the changes in the number of protein binding sites mediated by myopia-related SNPs:

[0074] Num Ref and NumAlt FC represents the number of proteins bound to the Ref and Alt sequences, respectively. We define FC > 1 as a protein binding loss, and conversely, FC < 1 as a protein binding gain. We then apply Fisher's exact test to determine changes in the protein binding motif, as shown below: P T =(a / b) / (c / d)

[0075] a represents the number of binding motifs that overlap only with the Ref sequence, b represents the number of binding motifs that do not overlap with the Ref sequence, c represents the number of binding motifs that overlap only with the Alt sequence, and d represents the sequence of binding motifs that do not overlap with Alt. Here, when FC is not equal to 1 and P T A score <0.05 indicates that SNPs associated with myopia significantly affect protein binding, and these SNPs are identified as cfSNPs. Furthermore, we define the protein with the highest score as the "leader" protein.

[0076] 2. Results:

[0077] 2.1 Myopia-related SNPs are widely distributed at the transcriptional and post-transcriptional levels:

[0078] We obtained myopia-related SNPs from public resources, and after quality control, 343 SNPs were used for subsequent analysis (see Methods). To reveal the complete map of SNPs at the transcriptional and post-transcriptional levels, 636 pairs were identified using precise alignment of 263 SNPs, 10 genomic regulatory elements, and myopia (Figure 6A). Furthermore, these SNPs were associated with five phenotypes: refractive error (RE), common myopia (CM), high myopia (HM), pathological myopia (PM), and visual impairment (VD). During transcription, 84 SNPs were located in enhancers, open chromatin regions (OCRs), CTCF binding sites (CTCFBSs), promoters, and promoter flanking regions (PFRs), forming 90 pairs. A total of 244 SNPs were located on five genomic regulatory elements: the 5'UTR, exons, introns, the 3'UTR, and lncRNAs, forming 546 pairs at the post-transcriptional level. Next, we found that 14.15% (90 / 636) and 85.85% (546 / 636) of the pairs were enriched during transcription and post-transcriptional processes, respectively (Figure 6B).

[0079] To further assess the average distribution of SNPs across all genomic regulatory elements, we analyzed the transcriptional length of these elements and the density of SNPs in each regulatory element. The results showed that the median length of the longest transcript, LncRNA, was 95.53 times that of the shortest transcript, CTCFBS (Figure 6C). We then further evaluated the enrichment of SNPs within each regulatory element per 1000 bp. Compared to an average of one SNP per 1000 bp in the human genome, we observed that myopia-associated SNPs were highly distributed in open chromatin, CTCF binding sites, the 5' untranslated region, and exons (Figure 6D). Analysis of myopia revealed that 81.76% (520 / 636) of the one-to-one associations were with myopia (CM). Furthermore, little is known about HM, PM, RE, and VD, which account for approximately 12.11% (77 / 636), 5.19% (33 / 636), 0.63% (4 / 636), and 0.31% (2 / 636), respectively. These data suggest that the distribution of myopia-related SNPs varies at the transcriptional and post-transcriptional levels as well as among genomic regulatory elements. These differences may be related to myopia severity and underlying molecular regulation.

[0080] 2.2 Scoring of SNP-induced molecular binding heterogeneity during transcription:

[0081] To investigate SNP-mediated molecular binding heterogeneity at the transcriptional level, we developed a computational pipeline that uses fold change (FC) to assess changes in the number of binding proteins and applies Fisher's exact test to assess changes in the number of binding protein molecules. Here, a threshold P is used. T <0.05, FC not equal to 1, and found that 38.46% (5 / 13), 80% (4 / 5), 57.89% (11 / 19), and 44.73% (17 / 38) of SNPs could disrupt transcription factor binding in enhancers, open chromatin regions, promoters, and promoter flanking regions, respectively (Figure 7). For the CTCF protein, we found that approximately 13.33% (2 / 15) of SNPs could disrupt the interaction at the CTCF binding site. In total, 43.33% (39 / 90) of the relationship pairs were able to disrupt binding affinity, and 46.43% (39 / 84) of the SNPs were recognized as “cfSNPs” during transcription. In conclusion, open chromatin was significantly enriched with myopia-associated SNPs and cfSNPs, exhibiting high density distribution and disrupted molecular binding.

[0082] To determine the potential influence of regulatory proteins associated with or potentially associated with myopia-related cfSNPs, we explored the molecular functions of the “leader” proteins in each relationship pair in published studies. Interestingly, most “leader” proteins were able to modulate ocular structures. For example, approximately 7.69% (3 / 39) of cfSNPs induced changes in the REST binding motif, thereby affecting the fate of retinal ganglion cells (RGCs) in the developing retina. Approximately 7.69% (3 / 39) of cfSNPs induced changes in IRF1 binding, which is known to be expressed in retinal microglia and play a key role in microglia activation and retinal inflammation. Furthermore, 15.28% (6 / 39) of cfSNPs altered the binding motif site of SPI1, which has been reported to modulate microglia in the retina. In summary, we found an abundance of “leader” proteins in retinal inflammation, which have been confirmed to be involved in the occurrence and development of myopia. Moreover, this can help us reveal potential protein regulators and understand how SNPs participate in the molecular regulation of myopia. Finally, we assessed the distribution of significantly altered relationships between myopia types at the transcriptional level. Over 75% (3 / 4) were associated with PM, while approximately 44.44% (4 / 9) and 41.56% (32 / 77) were associated with HM and CM (Figures 7A-D). The results also indicated that the impact of SNPs on myopia increases with the severity of myopia.

[0083] At the transcriptional level, computational analysis revealed that myopia-associated SNPs may gain or lose T binding sites. In open chromatin, rs8110889 located on ENSR00001023878 can weaken the interaction between FOXA1 and ENSR00001023878 (P < 0.05). T =6.28e-14, FC=1.38 (Figure 7B). Previous studies have shown that FoxA1 is closely related to signaling pathways in the vertebrate retina, suggesting that rs8110889 may disrupt retinal-related signaling pathways and lead to myopia susceptibility by altering molecular binding. In the promoter region, the Alt allele (C) of rs7550232 in ENSR00000020131 can increase the number of binding motifs on FLI1 (P). T=1.86e-26, FC value = 0.75 (Figure 7C). Another previous study reported that fli1 can drive vascular endothelial gene expression and control eye development in zebrafish. These results suggest that the acquisition and loss of SNP-induced regulatory proteins in the transcriptome may affect gene function and contribute to the pathogenesis of myopia. Recent studies have also reported similar relationships. For example, the T allele of rs17079281 in the DCBLD1 promoter can create a YY1 binding site to suppress DCBLD1 gene expression levels, reducing the risk of lung cancer in the Chinese population. Furthermore, rs3101339 may disrupt the binding between transcription factors and the NEGR1 gene, causing gene dysregulation and leading to major depressive disorder. The discovery of SNP-induced protein motif loss or gain can reveal molecular interactions and provide potential intervention targets.

[0084] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0085] In general, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of embodiments of this disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.

[0086] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0087] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0088] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0089] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0090] The exemplary embodiments of this disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art will understand that various modifications and combinations can be made to these embodiments or their features without departing from the principles and spirit of this disclosure, and such modifications should fall within the scope of this disclosure.

Claims

1. A method for evaluating protein binding efficiency at the gene transcription level, characterized in that, The method includes: 101, Obtain at least one SNP to be tested; 102, Retrieving Ref and Alt sequence pairs based on SNP; 103, extract the number of regulatory proteins that bind to the Ref and Alt sequences of the sequence pairs; 104. The number of proteins bound to the Ref sequence is denoted as Num. Ref The number of proteins bound to the Alt sequence is denoted as Num. Alt ; 105, according to the stated Num Ref and Num Alt The ratio of SNPs to protein binding efficiency is used to evaluate the effect of SNPs on protein binding efficiency.

2. The method for evaluating protein binding efficiency at the gene transcription level according to claim 1, characterized in that, The 105 includes: if the ratio is not equal to the first threshold, outputting the evaluation result of the SNP to be tested affecting the protein binding efficiency; the SNP to be tested is denoted as the first SNP.

3. The method for evaluating protein binding efficiency at the gene transcription level according to claim 2, characterized in that, The 105 includes: if the ratio is greater than a first threshold, outputting an evaluation result that the SNP to be tested has a low impact on protein binding efficiency; the SNP to be tested is denoted as the second SNP; optionally, if the ratio is less than the first threshold, outputting an evaluation result that the SNP to be tested has a high impact on protein binding efficiency; the SNP to be tested is denoted as the third SNP.

4. The method for evaluating protein binding efficiency at the gene transcription level according to claim 3, characterized in that, The method further includes: 106, screening sequence pairs located on transcriptional gene regulatory elements from the plurality of sequence pairs; 107, calculating a first ratio of the number a of binding motifs that overlap only with the Ref sequence and the number b of binding motifs that do not overlap with the Ref sequence based on the position information of the binding motifs, and calculating a second ratio of the number c of binding motifs that overlap only with the Alt sequence and the number d of binding motifs that do not overlap with the Alt sequence; calculating the change value of the protein binding motif based on the first ratio and the second ratio; screening SNPs whose change value is less than a second threshold as fourth SNPs, and outputting the evaluation result of the fourth SNP's impact on protein binding efficiency.

5. The method for evaluating protein binding efficiency at the gene transcription level according to claim 4, characterized in that, The method further includes: taking the intersection between the first SNP and the fourth SNP, and outputting the evaluation results of the SNPs in the intersection affecting protein binding efficiency.

6. The method for evaluating protein binding efficiency at the gene transcription level according to claim 4, characterized in that, The method further includes: taking the intersection between the second SNP and the fourth SNP, and outputting the prediction result that the SNP in the intersection has low protein binding efficiency; optionally, taking the intersection between the third SNP and the fourth SNP, and outputting the prediction result that the SNP in the intersection has high protein binding efficiency.

7. The method for evaluating protein binding efficiency at the gene transcription level according to claim 1, characterized in that, The gene regulatory elements include: enhancers, open chromatin regions, promoters, and promoter flanking regions; optionally, the regulatory proteins include transcription factors and CTCFs.

8. A computer device, characterized in that, The device includes: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the steps of the method according to any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1-7.

10. A computer program product comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method described in any one of claims 1-7.