Screening method for target genes of susceptible SNP loci for AML diagnosis and its application

By systematically integrating multidimensional omics data from AML-related cell lines, preprocessing three-dimensional genomic data to identify DNA interaction segments, predicting and screening target genes of AML susceptible SNP sites, the shortcomings in target gene identification in the prior art are solved, and high-reliability target gene prediction and AML diagnosis support are achieved.

CN119580833BActive Publication Date: 2025-05-27HAIHE LAB OF CELL ECOSYSTEM +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510116879.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify target genes at AML-susceptible SNP sites, especially sites in non-coding regions, resulting in a limited breadth of target gene identification and accompanied by a higher risk of false positives.

Method used

By obtaining AML susceptible SNP sites, three-dimensional genomic data and AML patient cohort data, pre-processing is performed to obtain DNA interaction segments, predict potential target genes, and obtain a highly reliable target gene set through highly reliable transcriptional regulatory target gene screening methods.

Benefits of technology

A comprehensive prediction of the transcriptional regulatory target genes of AML-susceptible SNP sites has been achieved, and a high-reliability transcriptional regulatory target genes have been verified, and new biomarkers and drug targets have been provided, which makes up for the shortcomings of the existing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119580833B_ABST
    Figure CN119580833B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of biotechnology, and provides a method for screening target genes of susceptible SNP sites for AML diagnosis and its application. The method includes: obtaining AML susceptible SNP sites, three-dimensional genomic data, and AML patient cohort data; preprocessing the three-dimensional genomic data to obtain DNA interaction segments; predicting potential target genes regulated by AML susceptible SNP sites based on the DNA interaction segments; screening highly credible transcriptional regulatory target genes according to the potential target genes regulated by AML susceptible SNP sites to obtain a highly credible target gene set; and verifying the transcriptional regulatory activity of SNPs in the highly credible target gene set by in vitro cell experiments. The present invention expands the understanding of the regulatory network between AML susceptible SNPs and target genes, and provides new biomarkers and drug targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biotechnology, and particularly to a method for screening target genes of susceptible SNP loci for AML diagnosis and its application. Background Art

[0002] Acute myeloid leukemia (AML) is a hematological malignancy originating from bone marrow hematopoietic stem cells, characterized by excessive proliferation and differentiation arrest of myeloblasts, resulting in impaired normal hematopoiesis. Although certain progress has been made in the treatment of AML in recent years, with approximately 75% of newly diagnosed AML patients achieving complete remission after 1 - 2 courses of induction chemotherapy, the problems of recurrence and drug resistance remain prominent, leading to poor long-term prognosis for AML patients, with a 5-year overall survival rate of less than 50%. The high genomic heterogeneity of AML patients is an important factor affecting recurrence and drug resistance. This heterogeneity is not only reflected in the diversity of somatic gene mutations and gene fusions but also involves a large number of genetic variations, especially single nucleotide polymorphisms (SNPs). With the widespread application of high-throughput sequencing technology, more and more studies have revealed the important role of genetic variations in the occurrence and development of AML. It is estimated that approximately 5% - 10% of AML cases are related to susceptibility caused by genetic variations.

[0003] In recent years, researchers have identified more than 94 SNP loci related to AML disease risk, treatment response, or prognosis based on Genome-Wide Association Studies (GWAS). These AML-susceptible SNP loci can regulate gene expression through interaction with target genes, thereby affecting the occurrence, diagnosis, and prognosis of the disease. However, only the molecular mechanisms of a very small number of AML-susceptible SNP loci have been elucidated, and the specific functions of a large number of susceptible loci in AML etiology remain to be explored. Accurately identifying the target genes of these SNP loci, especially those located in non-coding regions, is a major challenge currently faced. Existing methods for predicting target genes of SNP loci mainly rely on linear distance, predicting the gene closest to the SNP locus on the linear genome as its target gene. This method is limited to neighboring genes and fails to consider the possibility that SNP loci regulate distal genes through long-range chromatin interactions, resulting in a significant limitation in the breadth of target gene identification and a relatively high false positive risk.

[0004] In recent years, the rapid development of three-dimensional genomics technology has enabled researchers to predict the target genes of disease-susceptible SNP sites by analyzing the spatial contact relationships between disease-susceptible SNP sites and genes. However, existing methods can only confirm the physical contact between SNP sites and potential target genes in the spatial structure, and it is still impossible to clarify whether this physical contact can effectively change the transcriptional levels of target genes in the AML patient cohort. This limitation leads to great uncertainty in using three-dimensional genomics methods to analyze the specific regulatory mechanisms of SNP sites on their potential target genes, hindering our in-depth understanding of the functions of disease-susceptible SNP sites. In addition, there is currently no research on the precise and systematic identification and analysis of the transcriptional regulatory targets and their functions of AML-susceptible SNP sites. Therefore, existing methods have problems such as high false positives in predicting target genes of SNP sites, inability to identify distal target genes, and inability to clarify the transcriptional regulatory relationships between SNP sites and target genes. Summary of the Invention

[0005] The present invention aims to at least solve one of the technical problems existing in the related art. For this purpose, the present invention provides a method for screening target genes of susceptible SNP sites for AML diagnosis and its application, which realizes people's understanding of the regulatory network between AML-susceptible SNPs and target genes, and provides new biomarkers and drug targets.

[0006] The present invention provides a method for screening target genes of susceptible SNP sites for AML diagnosis, including:

[0007] S1: Obtain AML-susceptible SNP sites, three-dimensional genomic data, and AML patient cohort data;

[0008] S2: Preprocess the three-dimensional genomic data to obtain DNA interaction segments;

[0009] S3: Predict potential target genes regulated by AML-susceptible SNP sites according to the DNA interaction segments;

[0010] S4: Screen highly credible transcriptional regulatory target genes according to the potential target genes and AML patient cohort data to obtain a set of highly credible target genes.

[0011] Further, step S1 includes:

[0012] S11: Collect currently published AML-susceptible SNP sites;

[0013] S12: Collect three-dimensional genomic data from multiple myeloid cell lines, including Hi-C, ChIA-PET, BL-Hi-C;

[0014] S13: Collect data of the AML patient cohort, including genotype data and RNA-seq data of the TCGA-LAML, TARGET-AML, and BeatAML cohorts.

[0015] Further, step S2 includes:

[0016] For ChIA-PET and BL-Hi-C data, adopt the CHIAPET2 process. After linker sequence cleavage, quality control, data alignment, and binding peak finding, identify the interaction information between protein binding sites, and obtain DNA interaction segments based on the interaction information between protein binding sites;

[0017] For Hi-C data, use the HICCUPS software to identify DNA interaction segments.

[0018] Further, in step S3,

[0019] Compare the positions of AML-susceptible SNP sites with the DNA interaction segments identified in the three-dimensional genomics data to determine the DNA segments containing the susceptible SNP sites;

[0020] Based on the interaction pairing relationship between DNA segments provided in the three-dimensional genomic data, locate the DNA segments that interact with the DNA segments of the susceptible SNP sites, and obtain the potential target genes regulated by the AML-susceptible SNP sites.

[0021] Further, classify the target genes according to the interaction mode between the susceptible SNP sites and the target genes,

[0022] The susceptible SNP site is located at one end of a pair of interacting DNA segments, and the other end overlaps with the promoter region of the gene. Genes in this positional relationship with the susceptible SNP site are defined as type 1 target genes;

[0023] The susceptible SNP site is located at one end of a pair of interacting DNA segments, and the other end overlaps with the gene body region of the gene. Genes in this positional relationship with the susceptible SNP site are defined as type 2 target genes.

[0024] Further, compare the bases at the SNP site with the alleles in the hg38 reference genome. When both chromosomes of an AML patient carry the reference allele at this SNP site, the AML patient has the AA genotype; when the two chromosomes of an AML patient carry the reference allele and the mutant allele at this SNP site respectively, the AML patient has the AB genotype; when both chromosomes of an AML patient carry the mutant allele at this SNP site, the AML patient has the BB genotype.

[0025] Further, step S4 includes:

[0026] S41: Construct a linear regression model for the target gene based on the genotype, broad factors that have a wide impact on mRNA expression, and the mRNA expression level of the target gene;

[0027] S42: Use the least squares method to calculate the regression coefficient between the genotype of the target gene linear regression model and the mRNA expression level of the target gene, and obtain the regression coefficient between each genotype and the mRNA expression level of the target gene;

[0028] S43: According to the regression coefficient between the genotype and the mRNA expression level of the target gene, use the t-test to calculate the statistic and obtain the P-value for each genotype;

[0029] S44: Use Fisher's method to calculate the combined P-value for each genotype;

[0030] S45: Perform Benjamini-Hochberg correction on the combined P-value to obtain the false discovery rate FDR;

[0031] S46: Set three thresholds based on the P-value, the regression coefficient between the genotype and the mRNA expression level of the target gene, and the false discovery rate FDR to filter the target gene set, and screen the target genes that simultaneously meet the three thresholds to obtain a target gene set with high confidence.

[0032] Further, step S45 includes:

[0033] S451: Sort all the obtained combined P-values in ascending order and record the original position of each combined P-value;

[0034] S452: Calculate the adjusted value of each combined P-value based on the sorted combined P-values and their rankings, and the adjusted value of the combined P-value is the FDR value;

[0035] S453: Adjust the order of the FDR values so that the adjusted FDR values satisfy the principle of monotonic decrease;

[0036] S454: Reverse map the adjusted FDR values to the original positions of each combined P-value so that each combined P-value matches the corresponding adjusted FDR value to obtain the false discovery rate FDR.

[0037] Further, in S46,

[0038] Threshold 1: In at least one patient cohort, P-value < 0.05;

[0039] Threshold 2: The directions of multiple β 1 values are the same, all being positively correlated or negatively correlated;

[0040] Threshold 3: False Discovery Rate (FDR) < 0.05.

[0041] Application of a screening method for target genes of susceptibility SNP loci for AML diagnosis, wherein the screening method for target genes of susceptibility SNP loci for AML diagnosis is the above-mentioned screening method for target genes of susceptibility SNP loci for AML diagnosis.

[0042] One or more of the above technical solutions in the embodiments of the present invention have at least one of the following technical effects:

[0043] The present invention provides a screening method for target genes of susceptibility SNP loci for AML diagnosis and its application, systematically integrating multi-dimensional omics data of AML-related cell lines, and achieving a comprehensive prediction of transcriptional regulatory target genes of AML susceptibility SNP loci. Based on the AML patient cohort, the present invention verifies highly credible transcriptional regulatory target genes of AML susceptibility SNP loci, providing potential biomarkers and drug targets for the precise diagnosis and treatment of AML.

[0044] The present invention reveals various forms of regulatory target genes of AML susceptibility SNP loci, effectively making up for the deficiency of predicting SNP locus target genes only based on linear distance, clarifying that these susceptibility loci can regulate the expression of distal targets through chromatin interactions, and providing a theoretical basis for understanding the complex regulatory network between AML susceptibility SNPs and target genes.

[0045] Additional aspects and advantages of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0047] Figure 1 is a schematic flowchart of a screening method for target genes of susceptibility SNP loci for AML diagnosis provided by the present invention.

[0048] Figure 2 is an integrated view of AML susceptibility SNP rs187084 and its adjacent region observed in HL-60 cells by IGV in a screening method for target genes of susceptibility SNP loci for AML diagnosis provided by the present invention.

[0049] Figure 3 It is a result graph of eQTL analysis between the mRNA expression levels of rs187084 and ITIH4 genes based on three AML patient cohorts, namely TCGA-LAML, TARGET-AML, and BeatAML, for the screening method of susceptible SNP locus target genes for AML diagnosis provided by the present invention.

[0050] Figure 4 It is a statistical graph of the in vitro cell dual-luciferase reporter experiment to verify the regulatory effect of SNP rs187084 on gene expression for the screening method of susceptible SNP locus target genes for AML diagnosis provided by the present invention. Detailed implementation manners

[0051] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope protected by the present invention. The following embodiments are used to illustrate the present invention but cannot be used to limit the scope of the present invention.

[0052] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present invention. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without conflict, those skilled in the art can combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.

[0053] The following Figures 1 to 4 describes a screening method for susceptible SNP locus target genes for AML diagnosis and its application of the present invention.

[0054] As Figure 1 shown, a screening method for susceptible SNP locus target genes for AML diagnosis includes:

[0055] S1: Obtain AML susceptible SNP loci, three-dimensional genomic data, and AML patient cohort data;

[0056] S11: Collect 94 currently published AML-susceptible SNP loci, all of which are from literature reports and identified by genome-wide association studies.

[0057] S12: Collect three-dimensional genomic data from multiple myeloid cell lines, from multiple technical types, including Hi-C, ChIA-PET, and BL-Hi-C. All data are from the public database GEO, the DNA three-dimensional genomics sequencing set released in the third phase of the ENCODE project, or published literature.

[0058] S13: Collect AML patient cohort data, including genotype data and RNA-seq data of acute myeloid leukemia (TCGA-LAML), targeted acute myeloid leukemia (TARGET-AML), and the BeatAML cohort.

[0059] S2: Preprocess the three-dimensional genomic data to obtain DNA interaction segments.

[0060] For ChIA-PET and BL-Hi-C data, use the CHIAPET2 pipeline. After steps such as linker sequence cleavage, quality control, data alignment, and binding peak finding, identify the interaction information between specific protein-binding sites for subsequent analysis. The specific parameters of the CHIAPET2 pipeline are -m 1 -e 3 -k 2 -t 6 -l 12 -C 1. For Hi-C data, use the HICCUPS software to identify DNA interactions in different segments. The specific parameters of the HICCUPS software are -k KR--cpu -m 1024 --threads 8 -r 5000,10000 --ignore_sparsity.

[0061] S3: Predict the potential target genes regulated by AML-susceptible SNP loci based on the DNA interaction segments.

[0062] Predict the potential transcriptional regulatory target genes of AML-susceptible SNP loci based on the interaction information between different DNA segments provided by genomics data. Specifically, first compare the positions of AML-susceptible SNP loci with the DNA interaction segments identified in the three-dimensional genomic data to determine the DNA segments containing the susceptible SNP loci. Then, based on the interaction pairing relationship between DNA segments provided in the three-dimensional genomic data, locate the DNA segments that interact with the above-mentioned susceptible SNP loci, thereby inferring the potential target genes regulated by the susceptible SNP loci. According to the different interaction modes between the susceptible SNP loci and the target genes, the target genes are divided into the following two categories:

[0063] The susceptible SNP loci interact with the promoter region of genes. Specifically, the SNP is located at one end of a pair of interacting DNA segments, while the other end overlaps with the promoter region of the gene. Genes in this positional relationship with the susceptible SNP loci are defined as type 1 target genes;

[0064] The susceptible SNP loci interact with the gene body region of genes. Specifically, the SNP is located at one end of a pair of interacting DNA segments, while the other end overlaps with the gene body region of the gene. Genes in this positional relationship with the susceptible SNP loci are defined as type 2 target genes;

[0065] In addition, if a target gene is simultaneously identified as a type 1 target gene and a type 2 target gene of the AML susceptible SNP loci, it is preferentially classified as a type 1 target gene. The present invention predicted 192 potential target genes for rs187084, including 93 type 1 target genes and 99 type 2 target genes. The prediction results of potential transcriptional regulatory target genes of the AML susceptible SNP locus rs187084 based on three-dimensional genomic data are shown in Table 1.

[0066] Table 1 Prediction results of potential transcriptional regulatory target genes of the AML susceptible SNP locus rs187084 based on three-dimensional genomic data

[0067]

[0068]

[0069]

[0070] S4: Screen high-confidence transcriptional regulatory target genes based on the potential target genes regulated by the AML susceptible SNP loci and the AML patient cohort data to obtain a high-confidence target gene set;

[0071] Conduct Expression Quantitative Trait Loci (eQTL) analysis to compare the differences in the transcriptional levels of target genes in patients with different genotypes, and further clarify the effect of the AML susceptible SNP loci on target gene transcription. Considering that the effect of SNPs on target genes is small and the uneven sample size may cause bias in statistical significance, on the basis of conventional eQTL analysis, different statistical test methods are combined to optimize the effectiveness of eQTL verification;

[0072] First, according to the differences in the SNP loci of the patient gene sequences, the patients are divided into three genotypes: AA, AB, and BB.

[0073] By comparing the bases at the SNP locus with the alleles in the hg38 reference genome, when both chromosomes of an AML patient carry the reference allele at this SNP locus, it is classified as the AA genotype; when the two chromosomes of an AML patient carry the reference allele and the mutant allele respectively at this SNP locus, it is classified as the AB genotype; when both chromosomes of an AML patient carry the mutant allele at this SNP locus, it is classified as the BB genotype;

[0074] S41: Construct a linear regression model for the target gene based on the genotype, the broad factors that have a wide impact on mRNA expression, and the mRNA expression level of the target gene;

[0075] Using the genotype at the SNP locus as the independent variable and the mRNA expression level of the target gene as the dependent variable, construct a linear regression model for the target gene. The calculation expression of the linear regression model for the target gene is:

[0076]

[0077] Where, is the mRNA expression value of the target gene, are different genotypes at the SNP locus, AA = 0, AB = 1, BB = 2, is the regression coefficient between the genotype and the mRNA expression level of the target gene, is the broad factor, is the random error, is the mRNA expression value of the target gene when both the SNP genotype and the broad factor are 0, is the regression coefficient between the broad factor and the mRNA expression level of the target gene;

[0078] To avoid the influence of different sample sources (such as the gender, age, ancestry, etc. of the samples) on the identification of SNP genotypes and the calculation of mRNA expression levels, the broad factors that have a wide impact on mRNA expression are included as covariates in the regression model to remove their interference with the results;

[0079] The broad factors include the gender, age, ethnic background of the samples, and hidden factors; hidden factors refer to factors related to gene expression but cannot be directly measured, such as batch effects, environmental factors, and technical noise, etc. Calculate using the R package PEER, and according to the sample size of the three AML patient cohorts (150 ≤ N < 250), calculate 30 hidden factors for each AML patient cohort and add them as covariates to ensure the accuracy of the analysis;

[0080] The ethnic background refers to using principal component analysis (PCA) to reduce the dimensionality of the SNP genotype data, extract the main variation components between the samples, and select the first 5 principal components to add as covariates;

[0081] Evaluate the variation law of the mRNA expression level of the target gene in patients with different genotypes according to the linear regression model of the target gene;

[0082] S42: Calculate the regression coefficient between the genotype of the linear regression model of the target gene and the mRNA expression level of the target gene using the least squares method, and obtain the regression coefficient between each genotype and the mRNA expression level of the target gene;

[0083] Calculate the regression coefficients between each genotype AA, AB, and BB and the mRNA expression level of the target gene according to the linear regression model of the target gene; through the regression coefficients, judge the influence of different genotypes (AA, AB, BB) on the expression level of the target gene; a positive regression coefficient indicates that the genotype is associated with a higher mRNA expression level of the target gene, and a negative regression coefficient indicates that the genotype is associated with a lower mRNA expression level of the target gene;

[0084] S43: Calculate the statistic using the t-test according to the regression coefficient between the genotype and the mRNA expression level of the target gene, and obtain the P-value for each genotype;

[0085] Through three AML patient cohorts, obtain three according to the linear regression model of the target gene β 1 , according to β 1 Calculate the two-tailed significance level P,

[0086] According to distribution (degrees of freedom is the sample size minus the number of model parameters), calculate value, The calculation expression of the value is:

[0087]

[0088] where, is β 1 The standard error of, which measures the uncertainty of the estimate.

[0089] Calculate The two-tailed significance level corresponding to the value, and obtain the P-value.

[0090] If the P-value is less than the set significance threshold, reject the null hypothesis and consider the association between the SNP and the mRNA expression level of the target gene to be significant.

[0091] The P-value is used to evaluate the statistical significance of the relationship. Based on the P-value and the FDR value after BH correction, it is evaluated whether the association between the SNP and gene expression is significant. If the P-value is less than the preset significance level (e.g., 0.05), it indicates that there is a significant association between the SNP genotype and the expression of the target gene.

[0092] Three β 1 The values represent the regression coefficients of the SNP and the mRNA expression level of the target gene under different patient cohorts. The three P-values represent the statistical significance under different patient cohorts.

[0093] S44: Calculate the combined P-value by applying Fisher's method to the P-values of each genotype;

[0094] To obtain a reliable SNP-target gene set and eliminate the influence of uneven sample sizes on the statistical test significance, the combined P-value is calculated using Fisher's method for the above three P-values:

[0095]

[0096] Among them, is the combined P-value, is the P-value obtained from the th individual test, and

[0097] is the number of P-values.

[0098] To control the false positive rate in multiple hypothesis testing, the combined P-values of type 1 target genes and type 2 target genes are subjected to Benjamini-Hochberg (BH for short) correction.

[0099] S451: Sort all the obtained combined P-values in ascending order and record the original positions of each combined P-value.

[0100] S452: Calculate the adjusted value of each combined P-value based on the sorted combined P-values and their rankings. The adjusted value of the combined P-value is the FDR value.

[0101]

[0102] Among them, is the adjusted FDR value, is the th combined P-value after sorting, m is the total number of hypothesis tests, is the sorting position of the combined P-value,

[0103] S453: Adjust the order of FDR values so that the adjusted FDR values satisfy the principle of monotonic decrease;

[0104] If an adjusted FDR value is greater than the next adjusted FDR value, adjust it to the same value;

[0105] S454: Reverse map the adjusted FDR values to the original positions of each combined P-value so that each combined P-value matches the corresponding adjusted FDR value, obtaining the false discovery rate FDR.

[0106] S46: Set three thresholds according to the P-value, the regression coefficient between the genotype and the mRNA expression level of the target gene, and the false discovery rate FDR to filter the target gene set, screen the target genes that simultaneously meet the three thresholds, and obtain a highly credible target gene set;

[0107] Set three thresholds to filter the target gene set according to the performance of the target genes of the susceptibility SNP loci in different patient cohorts,

[0108] Threshold 1: In at least one patient cohort, P-value < 0.05;

[0109] Threshold 2: The directions of multiple β 1 values are the same, all being positively correlated or negatively correlated;

[0110] Threshold 3: False discovery rate FDR < 0.05.

[0111] If a target gene simultaneously meets the three criteria of Threshold 1, Threshold 2, and Threshold 3, it is a highly credible target gene.

[0112] Verify the transcriptional regulatory activity of the SNPs in the highly credible target gene set using in vitro cell experiments;

[0113] Detect the regulatory function of the SNP locus on the gene using a dual-luciferase reporter assay. In this experiment, the PGL3-REPORT vector is used, which contains a firefly luciferase reporter gene controlled by an SV40 promoter. At the multiple cloning site upstream of the promoter, nucleotide sequences of 150 bp in size upstream and downstream of the selected wild-type and its mutant SNP loci are inserted respectively to recombine the target plasmid. At the same time, the Renilla luciferase plasmid pRL-CMV is used as an internal reference plasmid. The target plasmid and the internal reference plasmid are co-transfected into 293T cells for 24 h, and the luciferase activity is measured using a dual-luciferase reporter detection kit. The regulatory effect of the SNP on the gene is judged by the level of luciferase activity, as Figure 4 shown.

[0114] The present invention successfully identified the highly credible transcriptional regulatory target gene ITIH4 for the AML-susceptible SNP locus rs187084. As shown in Table 2, the cross-validation results of rs187084-ITIH4 in three independent AML patient cohorts.

[0115] Table 2. Cross-validation results of rs187084-ITIH4 in three independent AML patient cohorts

[0116]

[0117] As Figure 2 shown, by observing through the Integrative Genomics Viewer (IGV for short), it was found that in HL-60 cells, rs187084 can physically contact the gene ITIH4 through long-range chromatin interactions.

[0118] Figure 2 shows multi-omics data of chromosome 3 (Chr3) of the HL-60 cell line in the region from 52,180 kb to 52,850 kb. The horizontal axis represents the genomic coordinates, with a coverage range of 670 kb. The blue annotations represent RefSeq gene annotations, indicating the positions of the key genes TLR9 and ITIH4 in this region; the red short lines represent the position of the SNP locus rs187084; the purple arcs represent the chromatin loops formed by all the anchor peaks where rs187084 is located and their interacting anchor peaks in the region from 52,180 kb to 52,850 kb, and the dark purple arc highlights the chromatin loop between the anchor peak where rs187084 is located and the regulatory region of the ITIH4 gene; Figure 2 shows the signal distribution and intensity of RNASeq (transcriptome sequencing), H3K27ac (enhancer mark), H3K4me3 (promoter mark), CTCF (chromatin structure protein), and RNA polymerase II (RNA Pol II phosphoS5). [0-3676] shows the numerical range of the protein peak intensity of RNASeq, [0-300] shows the numerical range of the protein peak intensity of H3K27ac, [0-515] shows the numerical range of the protein peak intensity of H3K4me3, [0-15] shows the numerical range of the protein peak intensity of CTCF, [0-102] shows the numerical range of the protein peak intensity of RNA polymerase II. The larger the value, the stronger the signal.

[0119] As Figure 3 shown, based on the three AML patient cohorts of TCGA-LAML, TARGET-AML, and BeatAML, rs187084 can upregulate the mRNA expression of the ITIH4 gene.

[0120] As Figure 4 shown, the results showed that the luciferase activity of the mutant rs187084 increased, indicating that the mutant rs187084 in AML patients could regulate gene promoter expression, had enhancer activity, and affected the up-regulation of oncogene expression.

[0121] A screening method for target genes of susceptible SNP loci for AML diagnosis was used to screen target genes of susceptible SNP loci for AML, providing new biomarkers and drug targets for the accurate diagnosis and treatment of AML.

[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for screening a susceptible SNP site target gene for AML diagnosis, characterized in that: include: S1: Obtain AML susceptibility SNP loci, three-dimensional genome data, and AML patient cohort data; S2: Preprocess the 3D genome data to obtain DNA interaction segments; S3: Prediction of potential target genes regulated by AML susceptibility SNP loci based on DNA interaction segments; S4: Screen high-confidence transcriptional regulatory target genes based on potential target genes and AML patient cohort data to obtain a high-confidence target gene set; S41: construct a linear regression model of the target gene based on the genotype, broad factors that have a broad impact on mRNA expression, and the mRNA expression level of the target gene; The genotype at the SNP site was used as the independent variable and the mRNA expression level of the target gene was used as the dependent variable to construct a linear regression model of the target gene. The calculation expression of the linear regression model of the target gene is: in, is the mRNA expression value of the target gene, For different genotypes of SNP sites, is the regression coefficient between genotype and target gene mRNA expression, is a broad factor, is a random error, is the target gene mRNA expression value when both the SNP genotype and the broad factor are 0, is the regression coefficient between the broad factor and the target gene mRNA expression; Broad factors include the sample's gender, age, ethnic background, and latent factors; latent factors are factors that are related to gene expression but cannot be directly measured. Ethnic background refers to the dimensionality reduction of SNP genotype data using principal component analysis to extract the main variation components between samples; S42: using the least squares method to calculate the regression coefficient between the genotype of the target gene linear regression model and the target gene mRNA expression level, and obtaining the regression coefficient between each genotype and the target gene mRNA expression level; S43: Based on the regression coefficient between the genotype and the target gene mRNA expression level, the t-test was used to calculate the statistic and obtain the P value of each genotype; S44: Fisher's method was used to calculate the P value of each genotype to obtain the joint P value; S45: Benjamini-Hochberg correction was performed on the joint P value to obtain the false discovery rate FDR; S46: According to the P value, the regression coefficient between the genotype and the target gene mRNA expression, and the false discovery rate FDR, three thresholds are set to filter the target gene set, and the target genes that meet the three thresholds are selected to obtain a high-confidence target gene set; Threshold 1: P value < 0.05 in at least one patient cohort; Threshold 2: The direction of the regression coefficients between multiple genotypes and target gene mRNA expression is the same, all positively or negatively correlated; Threshold 3: False Discovery Rate FDR < 0.

05.

2. The method for screening a susceptible SNP site target gene for AML diagnosis according to claim 1, characterized in that: The S1 step includes: S11: Collect the currently published AML susceptibility SNP sites; S12: Collect 3D genomic data from various myeloid cell lines, including Hi-C, ChIA-PET, and BL-Hi-C; S13: Collect AML patient cohort data, including genotype data and RNA-seq data of TCGA-LAML, TARGET-AML, and BeatAML cohorts.

3. The method for screening a susceptible SNP site target gene for AML diagnosis according to claim 2, characterized in that: The S2 step includes: For ChIA-PET and BL-Hi-C data, the CHIAPET2 process was used to identify the information of the interaction between protein binding sites after linker sequence cutting, quality control, data alignment, and binding peak search. Based on the information of the interaction between protein binding sites, the DNA interaction segments were obtained. For Hi-C data, HICCUPS software was used to identify DNA interacting segments.

4. The method for screening a susceptible SNP site target gene for AML diagnosis according to claim 1, characterized in that: In step S3, Compare the positions of AML susceptibility SNP sites with DNA interaction segments identified in three-dimensional genomics data to determine the DNA segments containing the susceptibility SNP sites; Based on the interaction pairing relationship between DNA segments provided in the three-dimensional genome data, the DNA segments that interact with the DNA segments of the susceptible SNP sites are located, and the potential target genes regulated by the AML susceptible SNP sites are obtained.

5. The method for screening a susceptible SNP site target gene for AML diagnosis according to claim 1, characterized in that: The target genes are classified according to the interaction between the susceptible SNP sites and the target genes. The susceptible SNP site is located at one end of a pair of interacting DNA segments, while the other end overlaps with the promoter region of the gene. The gene in this positional relationship with the susceptible SNP site is defined as a type 1 target gene; The susceptible SNP site is located at one end of a pair of interacting DNA segments, while the other end overlaps with the gene body region of the gene. The gene in this positional relationship with the susceptible SNP site is defined as a type 2 target gene.

6. The method for screening a susceptible SNP site target gene for AML diagnosis according to claim 1, characterized in that: Compare the bases at the SNP site with the alleles in the hg38 reference genome. When both chromosomes of an AML patient carry the reference allele at the SNP site, the AML patient has an AA genotype; when both chromosomes of an AML patient carry the reference allele and the mutant allele at the SNP site, the AML patient has an AB genotype; when both chromosomes of an AML patient carry the mutant allele at the SNP site, the AML patient has a BB genotype.

7. The method for screening a susceptible SNP site target gene for AML diagnosis according to claim 1, characterized in that: Step S45 includes: S451: sort all obtained joint P values ​​in ascending order and record the original position of each joint P value, S452: Calculate an adjusted value of each joint P value based on the sorted joint P values ​​and rankings, where the adjusted value of the joint P value is an FDR value; S453: Adjust the order of FDR values ​​so that the adjusted FDR values ​​meet the principle of monotonically decreasing; S454: Reversely map the adjusted FDR value to the original position of each joint P value, so that each joint P value matches the corresponding adjusted FDR value to obtain the false discovery rate FDR.

8. An application of a method for screening a susceptible SNP site target gene for AML diagnosis, characterized in that: The method for screening a susceptible SNP site target gene for AML diagnosis is the method for screening a susceptible SNP site target gene for AML diagnosis described in any one of claims 1-7.