Method for screening disease-related gene targets based on whole genome correlation analysis

By introducing covariates and maximum likelihood estimation methods to optimize model parameters, the shortcomings of SNP screening methods in the prior art are solved, and more accurate identification of disease-related gene targets is achieved, especially in the genetic risk assessment of type 2 diabetes.

CN120340596APending Publication Date: 2025-07-18SHENZHEN GIANT CROCODILE BIOTECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510444748.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing SNP screening methods usually ignore the effects of covariates, interactions between SNPs, and challenges of high-dimensional data, resulting in false positive or false negative results, and the P-value screening is not accurate enough.

Method used

Logistic regression model is used to introduce covariates such as age, gender, and weekly movements, and high-dimensional data is processed through regularization methods, combined with maximum likelihood estimation to optimize model parameters, capture the interaction between SNPs and reduce the risk of overfitting.

Benefits of technology

It improves the accuracy of identification of disease-related SNPs, reduces the influence of confounding factors, and enhances the reliability and applicability of analysis results, especially in the genetic risk assessment of type 2 diabetes, which provides more accurate gene target recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340596A_ABST
    Figure CN120340596A_ABST
Patent Text Reader

Abstract

The invention provides a method for screening disease-related gene targets based on whole genome association analysis, and belongs to the technical field of disease genetic risk assessment, the method comprises the following steps: collecting and recording a sample information data set; performing whole genome sequencing on the samples to obtain SNP data, and combining the SNP data of all the samples to obtain a training data set; performing principal component analysis on the SNP matrix, and screening the top 10-30 SNPs as covariables of the model; and carrying out whole genome association analysis on the training data set and constructing a logistic regression model, and screening out the SNP related to the disease according to the P value. According to the method, confounding factors can be effectively controlled by adopting the logistic regression model and introducing covariables (such as age, gender and the number of weekly exercises), high-dimensional data are processed through a regularization method, and the risk of over-fitting is reduced. The method can also capture interaction between SNPs, and provides more accurate disease-related SNP recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of disease genetic risk assessment, and particularly relates to a method for screening disease-related gene targets based on genome-wide association analysis. Background Art

[0002] Type 2 diabetes is a common metabolic disease, mainly characterized by insulin resistance and pancreatic islet β-cell function decline. Different from type 1 diabetes, type 2 diabetes usually occurs in adults and is closely related to genetic and environmental factors. Lifestyle factors such as obesity, unhealthy diet, and lack of exercise are its main risk factors. In addition, genetic susceptibility also plays a key role in the occurrence of type 2 diabetes, and certain gene mutations can increase an individual's susceptibility to insulin resistance.

[0003] With the development of genomics technology, more and more studies have revealed gene mutations related to type 2 diabetes through genome-wide association studies (GWAS). These studies help to more deeply understand the pathogenesis of type 2 diabetes and provide new theoretical bases for individualized treatment and early screening. Although type 2 diabetes cannot be completely cured at present, the occurrence and progression of diabetes can be effectively controlled by early identification of risk genes and adjustment of lifestyle. Therefore, further exploration of the genetic background of diabetes, especially the identification of new candidate genes and targets, has important clinical significance.

[0004] Existing SNP screening methods usually ignore the influence of covariates, the interaction between SNPs, and the challenges of high-dimensional data, and often rely on P-values for screening, which easily leads to false positives or false negatives. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a method for screening disease-related gene targets based on genome-wide association analysis. By introducing covariates (such as age, gender, lifestyle, etc.) using a logistic regression model, the method can effectively control confounding factors, and process high-dimensional data through a regularization method to reduce the risk of overfitting. The method can also capture the interaction between SNPs and optimize model parameters using maximum likelihood estimation to provide more accurate identification of disease-related SNPs.

[0006] The present invention provides a method for screening disease-related gene targets based on genome-wide association analysis, including the following steps:

[0007] 1) Randomly select samples, obtain primary hepatocytes of the samples, and divide the samples into a diseased group and a healthy group according to the disease status;

[0008] 2) Collect the age, gender, and weekly exercise frequency of each sample, and record them to form a sample information data set;

[0009] 3) Perform whole-genome sequencing on the primary hepatocytes of the samples to obtain sequencing data; align the whole-genome sequencing data of each sample with the reference genome to obtain data containing all SNPs of each sample, and combine the SNP data of the samples to obtain a training dataset;

[0010] 4) Perform count encoding on the SNPs in the training dataset described in step 3). Homozygous reference genes are recorded as 0, heterozygotes are recorded as 1, and mutant homozygotes are recorded as 2 to obtain an SNP matrix with SNP loci as rows and sample names as columns. Perform principal component analysis on the SNP matrix and screen the top 10 - 30 SNPs as covariates of the model; preferably, screen the top 20 SNPs as covariates of the model;

[0011] 5) Conduct a genome-wide association analysis and construct a logistic regression model:

[0012]

[0013] where P represents the probability of the sample being diseased, β0 is the intercept term, indicating the baseline level of the target phenotype in the case where there is no SNP effect (SNP = 0) and no other covariate effects; β1, β pc , β age , β gender and β exercise are regression coefficients, which represent the influence of each independent variable on the target variable; Xsnp is the count encoding of a specific SNP locus, PCi are the top 20 components obtained from the principal component analysis in step 4), Age is the age of the sample, Gender is the gender of the sample: males are encoded as 1 and females are encoded as 2; Exercise is the number of weekly exercise of the sample; Equation I is the expression of the logistic function, representing the logarithm of the probability ratio of occurrence and non-occurrence, used to optimize the coefficients.

[0014] Convert it into probability form P as the probability of being diseased:

[0015]

[0016] Equation II is the inverse function of Equation 1, calculating the probability of the sample being diseased, used to explain Equation 1. The other covariates among them are βpc, βage, βgender, and βexercise mentioned above.

[0017] Train the sample information dataset obtained in step 2) and the training dataset obtained in step 3) based on the maximum likelihood function;

[0018]

[0019] Take the logarithm to obtain the log-likelihood function:

[0020] logL(β) = ∑(from i = 1 to N)[yilogPi + (1 - yi)log(1 - Pi)], Equation IV

[0021] After importing the training data set, the ascending gradient of each data is calculated using the gradient ascent method:

[0022]

[0023] Set the learning rate η = 0.01, and the new coefficient β1 for each iteration is:

[0024]

[0025] When the gradient converges or all the data has been iterated, the optimal coefficient β1 is obtained; by calculating the P - value of the coefficient, it is determined whether this SNP is related to the disease;

[0026] When P ≤ 0.05, then this SNP is significantly related to the disease; if β1 > 0, then the SNP is positively related to the disease, if β1 ≤ 0, then the SNP is negatively related to the disease;

[0027] Repeat the above step 5) of constructing the logistic regression model. Build a logistic regression model for each SNP, and screen out all SNPs related to the disease according to the P - value;

[0028] After obtaining all the SNP sites related to the disease, according to the positions of the SNP sites corresponding in the genome, the genes significantly related to the disease can be determined.

[0029] Preferably, the disease includes type 2 diabetes.

[0030] Preferably, the total number of the samples is 500 - 2000.

[0031] Preferably, after obtaining the sequencing data in step 3), it further includes a step of performing quality control on the sequencing data.

[0032] Preferably, the reference genome in step 3) is SRX28141778.

[0033] Compared with the prior art, the present invention has the following beneficial effects:

[0034] The present invention provides a method for screening disease - related gene targets based on genome - wide association study, which combines the logistic regression model and the maximum likelihood estimation (MLE) optimization technology, and can effectively identify SNPs related to diseases such as diabetes. Compared with the traditional GWAS method, by considering multiple covariates, the present invention reduces the confounding factors in the samples and improves the accuracy of the association between SNPs and diseases.

[0035] Comprehensive analysis considering multiple covariates: The present invention introduces multiple covariates (such as age, gender, weekly exercise frequency, the top 10-30 principal components, etc.) to correct the relationship between SNPs and diseases. It can significantly reduce the influence of confounding factors, making the analysis results more reliable. Compared with the traditional method without adjusting covariates, the present invention has obvious advantages in controlling confounding factors.

[0036] The present invention uses the maximum likelihood estimation (MLE) and gradient ascent optimization method to ensure that the parameters of the regression model converge to the optimal values, avoiding the common problem of unstable model convergence. The gradient ascent method can quickly and efficiently adjust the model parameters during the optimization process to achieve better prediction ability.

[0037] The present invention performs data dimensionality reduction through PCA, reducing the computational complexity in high-dimensional data and making the processing of large-scale SNP datasets more efficient.

[0038] The GWAS analysis method of the present invention can flexibly adapt to different types of samples (such as case-control design, population samples, etc.), and has strong applicability. Whether in the study of a single disease or in the analysis of complex diseases involving multiple factors, the present invention can provide effective support. Description of the Drawings

[0039] Figure 1 It is a flowchart of the method for screening disease-related gene targets based on genome-wide association analysis of the present invention. Detailed Description of the Invention

[0040] The following combines examples to detail the technical solutions provided by the present invention, but they cannot be understood as limiting the protection scope of the present invention.

[0041] Example 1

[0042] 1. Sample Screening

[0043] Randomly select 1000 unrelated samples from the hospital, including 500 in the healthy group and 500 type 2 diabetes patients, and obtain a small amount of primary hepatocytes by liver puncture.

[0044] 2. Data Collection

[0045] Perform whole-genome sequencing on the collected primary hepatocytes of the samples, with a paired-end sequencing read length of 150 bp, and the calculated sequencing depth is 23.9x.

[0046] Collect the age, gender, and weekly exercise frequency of each sample and record them to form a dataset.

[0047] 3. Quality Control

[0048] Using Trimmomatic to remove low-quality reads and adapter sequences from sequencing data

[0049] 4. Sequence alignment

[0050] The template gene for human whole genome sequencing (SRX28141778) was obtained from NCBI. The whole genome sequencing data of 1000 samples were compared with the queried template gene using Hisat2 to obtain 1000 vcf files containing all SNPs of different individuals. They were then merged into one file as a training data set, and the sites with low frequency and high deletion were deleted using bcftools.

[0051] 5. Data Dimensionality Reduction

[0052] In order to eliminate the influence of population background information on model training and reduce model errors to prevent overfitting, the top 20 SNPs obtained by principal component analysis and the age, gender and weekly exercise times of the collected samples were added to the model as covariates.

[0053] Specific operation: All SNPs in the training data set are counted: reference gene homozygotes such as (AA) are recorded as 0, heterozygotes such as (AG) are recorded as 1, and mutation homozygotes such as (GG) are recorded as 2. In this way, a behavioral SNP site is obtained, and a SNP matrix with sample names is listed. According to this matrix, principal component analysis is performed. The SNPs of the first 20 components of the screening are set as covariates of the model.

[0054] 6. Model building and training

[0055] In this embodiment, since it is necessary to determine whether the SNP is related to type 2 diabetes, a logistic regression model is used to evaluate the association between the SNP and the phenotype.

[0056] Therefore, the following logistic regression model is established:

[0057]

[0058] Where P represents the probability of the sample suffering from type 2 diabetes, β0 is the intercept term, which represents the baseline level of the target phenotype in the absence of SNP influence (SNP = 0) and other covariates. pc , β age , β gender and β exerciseThe βs are regression coefficients, which represent the impact of each independent variable (such as SNP or covariate) on the target variable. Xsnp is the genotype coding of a specific SNP locus (e.g., 0, 1, 2), PCi are the top 20 components obtained from the principal component analysis in the previous step, Age is the age of the sample, Gender is the gender of the sample, coded as 1 for male and 2 for female. Exercise is the number of weekly exercise of the sample. Thus, a logistic regression model for a specific Xsnp is established.

[0059] Finally, it is converted into the probability form P, which is the probability of having the disease:

[0060]

[0061] Input the previously collected sample information data set and the VCF file of SNPs into the model, and train the model based on the maximum likelihood function. It is hoped to find the optimal regression coefficient β1 to make the observed data most likely to occur, that is, to maximize the likelihood function.

[0062]

[0063] Taking the logarithm gives the log-likelihood function:

[0064] logL(β) = ∑Ni = 1[yilogPi+(1 - yi)log(1 - Pi)] Equation IV;

[0065] Therefore, after importing the training data, the gradient ascent method is used to calculate the ascending gradient of each data:

[0066]

[0067] Set the learning rate η = 0.01, and the new coefficient β1 for each iteration is:

[0068]

[0069] When the gradient converges or all the data have been iterated, the optimal coefficient β1 is obtained (Equations III - VI are the process of obtaining the optimal β1, and β1 is the β1 in Equation I). By calculating the P-value of the coefficient, it is determined whether this SNP is related to type 2 diabetes.

[0070] Since then, repeat the above steps of constructing the logistic regression model to establish a logistic regression model for each SNP, and screen out all SNPs related to the disease according to the P-value;

[0071] When P ≤ 0.05, then this SNP is significantly related to type 2 diabetes; and if β1 > 0, the SNP is positively related to type 2 diabetes, otherwise it is negatively related. Since then, repeat the above steps to establish a logistic regression model for each SNP, and all SNPs related to type 2 diabetes can be screened out.

[0072] 7. Locus determination

[0073] After obtaining all the SNP loci related to type 2 diabetes, according to the positions of their genomes stored in the previously saved VCF file, the gene with the greatest association with type 2 diabetes can be determined.

[0074] Table 1 shows partial screening results

[0075] SNPs Classification Description Gene β1 P value rs4841132 Type 2 diabetes Non-coding RNA Loc157273 0.87 1.2e-05 rs2241758 Organism system Insulin secretion ADCY3 0.51 0.014 rs138370845 Metabolism Lipid metabolism SMPD3 -0.42 0.062 rs769412 Human disease Cancer MDM2 0.13 0.027 rs2076234 Human disease Insulin resistance PLCB4 0.62 0.0038 rs312480 Organism system Insulin secretion CACNA1D 0.39 0.041 rs874565 Metabolism Glycerolipid metabolism LIPG -0.19 0.03 rs690 Metabolism Glycerolipid metabolism LIPC 0.72 7.5e-04 rs896375 Iron-mediated programmed cell death Apoptosis SLC39A14 -0.55 0.022 rs140054426 Metabolism Glycerolipid synthesis HAO2 0.28 0.014

[0076] The results screened by the method of the present invention are partially shown to be related to other diseases. This is because type 2 diabetes has many complications, which also provides a certain explanation for gene-regulated diseases.

[0077] Traditional methods directly calculate the correlation by statistically analyzing the GC values of SNPs in diabetic and normal samples. The screening process of the method of the present invention is more accurate. For example, in the category of metabolism, the conventional screening method can only screen one locus, rs138370845, but the method of the present invention can find 4 SNP variants directly related to metabolism.

[0078] As can be seen from the above embodiments, by combining the logistic regression model and the maximum likelihood estimation (MLE) optimization technique, SNPs related to diseases such as diabetes can be effectively identified. Compared with the traditional GWAS method, the present invention reduces confounding factors in the sample by considering multiple covariates and improves the accuracy of the association between SNPs and diseases.

[0079] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for screening disease-related gene targets based on genome-wide association study, characterized in that It includes the following steps: 1) Randomly select samples, obtain primary hepatocytes of the samples, and divide the samples into a diseased group and a healthy group according to the disease status. 2) Collect the age, gender, and weekly exercise frequency of each sample, and record them to form a sample information data set. 3) Perform whole-genome sequencing on the primary hepatocytes of the samples to obtain sequencing data. Align the whole-genome sequencing data of each sample with the reference genome to obtain data containing all SNPs of each sample, and combine the SNP data of the samples to obtain a training data set. 4) Perform count encoding on the SNPs in the training data set described in step 3). Homozygous reference genes are recorded as 0, heterozygotes are recorded as 1, and mutant homozygotes are recorded as 2 to obtain an SNP matrix with row SNP loci and column sample names. Perform principal component analysis on the SNP matrix and screen the top 10 - 30 SNPs as covariates of the model. 5) Perform a genome-wide association analysis and construct a logistic regression model: Equation I; where P represents the probability of the sample being diseased, β0 is the intercept term, indicating the baseline level of the target phenotype in the absence of SNP influence (SNP = 0) and other covariates; β1, β pc , β age , β gender and β exercise are regression coefficients, which represent the influence of each independent variable on the target variable; Xsnp is the count coding of a specific SNP locus, PCi are the first 20 components obtained from the principal component analysis in step 4), Age is the age of the sample, Gender is the gender of the sample: coded as 1 for male and 2 for female; exercise is the number of weekly exercise of the sample; Convert it into the probability form P as the probability of having the disease: Based on the maximum likelihood function, train the sample information data set obtained in step 2) and the training data set obtained in step 3). Take the logarithm to obtain the log-likelihood function: logL(β) = ∑i = 1N[yilogPi+(1 - yi)log(1 - Pi)] Equation IV. After importing the training data set, use the gradient ascent method to calculate the ascending gradient of each data: Set the learning rate η = 0.01, and the new coefficient β1 for each iteration is: When the gradient converges or all data have been iterated, obtain the optimal coefficient β1. Determine whether this SNP is related to the disease by calculating the P value of the coefficient. When P ≤ 0.05, then this SNP is significantly related to the disease. If β1 > 0, the SNP is positively related to the disease; if β1 ≤ 0, the SNP is negatively related to the disease. Repeat the steps of constructing the logistic regression model in step 5) above to establish a logistic regression model for each SNP, and screen out all SNPs related to the disease according to the P value. After obtaining all SNP loci related to the disease, according to the positions of the SNP loci corresponding to the genome, the genes significantly related to the disease can be determined.

2. The method according to claim 1, wherein The disease includes type 2 diabetes.

3. The method according to claim 1, wherein The total number of the samples is 500 - 2000.

4. The method according to claim 1, characterized in that, After obtaining the sequencing data in step 3), it also includes the step of performing quality control on the sequencing data.

5. The method according to claim 4, characterized in that, The reference genome described in step 3) is SRX28141778.

Citation Information

Cited By

  • High-incidence disease gene analysis and prediction method and system and storage medium

    CN121583325A

  • Method and system for gene expression quantitative trait locus analysis

    CN121687177A