Marker combination and system for predicting risk of hyperuricemia turning into gout, readable storage medium and application

By constructing a biomarker combination of SNP sites and environmental factors, and combining it with a logistic regression model, the invasiveness of early gout diagnosis was solved, achieving non-invasive and efficient gout risk prediction and improving the accuracy of early detection and intervention.

CN121065330APending Publication Date: 2025-12-05TAIZHOU ENZE MEDICAL CENT GROUP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511335851.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing methods for diagnosing gout mainly rely on joint fluid aspiration, which carries invasive risks and is difficult to widely apply in primary hospitals. Furthermore, existing assessment methods for predicting the onset of gout are limited and fail to effectively combine the combined effects of genetic and environmental factors, leading to difficulties in early diagnosis.

Method used

A biomarker combination including 14 SNP loci and 5 environmental factors was constructed. The risk of hyperuricemia turning into gout was predicted by a logistic regression model. The SNP loci were detected by NGS-multiplex PCR targeted capture technology and the model was optimized by machine learning algorithms to provide a non-invasive and efficient screening method.

Benefits of technology

It enables accurate prediction of the risk of hyperuricemia turning into gout, reduces the risk of invasive diagnosis, increases the possibility of early detection and early intervention, reduces diagnostic costs, and has high accuracy and wide applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121065330A_ABST
    Figure CN121065330A_ABST
Patent Text Reader

Abstract

The invention provides a marker combination and system for predicting the risk that hyperuricemia is converted into gout, a readable storage medium and application, and belongs to the technical field of early diagnosis of gout. The marker combination is used for detecting the marker combination in a to-be-detected sample to construct a corresponding diagnosis system, the system is high in accuracy, sensitivity, specificity and precision in diagnosing the risk that hyperuricemia is converted into gout, accurate pre-judgment of conversion of hyperuricemia into gout can be realized, so that a patient is intervened in time, and the risk of conversion of hyperuricemia into gout is reduced. The method can efficiently predict that an individual is a gout high-risk person or a gout low-risk person, and meanwhile, non-invasive screening of gout risk prediction is realized. The gout risk is conveniently and quickly diagnosed by combining the mode of detecting the SNP combined genotype in the sample to be detected with environmental factors, the detection result is highly consistent with the clinical gold standard detection result, early gout patients can be accurately identified and diagnosed, and early discovery and early treatment of gout are promoted.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of early diagnosis of gout, and particularly relates to a marker combination for predicting the risk of hyperuricemia turning into gout, a system, a readable storage medium and application. BACKGROUND

[0002] The prevalence of gout is increasing year by year with the improvement of living standards, and the age of onset tends to be younger. Gout refers to intermittent onset and severe painful arthritis caused by deposition of urate crystals in joints and non-joint structures, and belongs to the category of metabolic rheumatic diseases. Gout patients are often accompanied by hyperlipidemia, hypertension, diabetes, arteriosclerosis and coronary heart disease, and gout may be complicated with kidney disease, and severe cases can cause joint damage and kidney function damage. In summary, gout invades multiple joints and even the kidney of the whole body, eventually leading to severe joint deformity and pain or kidney function damage, so early detection, early diagnosis and early prevention of gout are the key to controlling and reversing the disease.

[0003] The current gold standard for gout diagnosis is to extract joint fluid by puncture, and to check the joint puncture fluid under a polarized light microscope. If there are three-prism crystal urate, it is diagnosed as gout. Joint fluid puncture is a kind of invasive detection method, which is a non-direct-viewing invasive operation, so the puncture needle may damage the synovial membrane, causing synovial bleeding, and when the bleeding is more, it can cause traumatic synovitis, leading to joint pain and swelling, and joint movement is limited. At the same time, this diagnostic method is restricted by factors such as equipment, medical personnel's professional quality, etc., so that the method for diagnosing early gout often cannot be widely used in primary community hospitals. Once gout appears typical clinical symptoms, the joint has been damaged, the optimal intervention opportunity is missed, so early and effective screening is of great significance for gout disease management.

[0004] With the development of gout susceptibility genes and molecular markers, early prediction of gout has gradually become possible. However, the existing evaluation methods for predicting the onset of gout are very limited, usually only referring to basic information such as age, gender, BMI, and laboratory test indicators such as blood uric acid, triglycerides, etc.; on the other hand, the first large sample gout genome-wide association study (GWAS) in Asia found 3 gout susceptibility genes BCAS3, RFX3 and KCNQ1, and elucidated the genetic basis of gout, while existing GWAS analysis has also found multiple single nucleotide polymorphism sites associated with gout in East Asians. However, although genetic factors are the main cause of gout, environmental factors such as lifestyle and dietary habits still play an important triggering role in the onset of gout, for example, high purine diet (such as red meat, seafood), excessive alcohol consumption, obesity and lack of exercise and other unhealthy lifestyles can significantly increase the risk of gout. Studies have found that people with high genetic susceptibility and unhealthy lifestyles have a higher risk of gout. A study analyzed the effects of genetic susceptibility and lifestyle on gout through a multifactorial joint model, and the results showed that people with high genetic susceptibility and unhealthy lifestyles had a significantly increased relative risk of gout compared to people with low genetic susceptibility and healthy lifestyles. However, the combined effects of genetic factors and environmental factors on gout risk are still unknown.

[0005] In summary, it is urgent to explore the relationship between single nucleotide polymorphism sites (SNPs), environmental factors and gout risk, and based on this relationship, to construct a corresponding model to predict the risk of gout, so that more people can be diagnosed with gout as early as possible before the onset of the disease, so as to control the occurrence and development of the disease as early as possible. SUMMARY

[0006] In view of this, the purpose of the present application is to provide a marker combination, system, readable storage medium and application for predicting the risk of hyperuricemia turning into gout.

[0007] In order to achieve the above-mentioned purpose of the application, the present application provides the following technical solutions:

[0008] The present application provides a marker combination for predicting the risk of hyperuricemia turning into gout, the marker combination being reagents for detecting SNP combination genotypes and environmental factors; the SNP combination being one or more of SNP1-SNP14, the SNP1-SNP14 corresponding to reference sequence numbers rs938557, rs545854, rs6947309, rs7688672, rs7903456, rs2242206, rs3825017, rs55975541, rs435309, rs4684846, rs6837293, rs3751143, rs2941484 and rs208294, respectively;

[0009] The SNP site information of the SNP combination is as follows:

[0010]

[0011]

[0012] The environmental factor is one or more of exercise frequency, BMI index, alcohol drinking frequency, smoking frequency and staying up late frequency.

[0013] Preferably, the marker combination is any one of marker combination 1, marker combination 2 and marker combination 3.

[0014] The marker combination 1 is reagent 1 for detecting SNP combination genotype and an environmental factor, wherein the SNP combination is SNP1-SNP13, and the environmental factor is exercise frequency, BMI index, alcohol drinking frequency, smoking frequency and staying up late frequency.

[0015] The marker combination 2 is reagent 2 for detecting SNP combination genotype and an environmental factor, wherein the SNP combination is SNP1-SNP14, and the environmental factor is exercise frequency, BMI index, alcohol drinking frequency, smoking frequency and staying up late frequency.

[0016] The marker combination 3 is reagent 3 for detecting SNP combination genotype and an environmental factor, wherein the SNP combination is SNP1-SNP12, and the environmental factor is exercise frequency, BMI index, alcohol drinking frequency, smoking frequency and staying up late frequency.

[0017] In the present application, the SNP combination can be inquired according to the reference sequence number corresponding to each SNP in the 1000 Genome Database, NCBI Database, ensembl Database or ucsc Database.

[0018] In the present application, most of the above-mentioned SNP sites are present on genes related to uric acid metabolism and transport, inflammation regulation, metabolic regulation, etc. The above-mentioned SNP combination is composed of 12, 13 or 14 SNP sites located on the hg19 / GRCh37 reference genome version hg19 / GRCh37, which can be inquired in the 1000 Genome Database, NCBI Database, ensembl Database or ucsc Database.

[0019] In the present application, the method for detecting single nucleotide polymorphism (SNP) genotype such as microarray chip, TaqMan probe method, restriction fragment length polymorphism (RFLP), allele-specific PCR (AS-PCR), Sanger sequencing, high-throughput sequencing (NGS), MALDI-TOF mass spectrometry (such as Sequenom MassARRAY) or single molecule sequencing, etc. can be selected by those skilled in the art according to the experimental requirements and budget costs to obtain the SNP information of the detection object. In some embodiments, high-throughput sequencing is used to detect SNP sites, more specifically, NGS-multiplex PCR targeted capture technology, such as the reagent for detecting the above SNP combination genotype includes a primer set.

[0020] In some specific embodiments, the primer set includes a primer set for the first round of amplification and a primer pair for the second round of amplification;

[0021] The primer set for the first round of amplification is one or more of the upstream and downstream primers for amplifying SNP1-SNP14, wherein the upstream and downstream primers for amplifying SNP1 are as shown in SEQ ID NO. 76 and SEQ ID NO. 171; the upstream and downstream primers for amplifying SNP2 are as shown in SEQ ID NO. 53 and SEQ ID NO. 148; the upstream and downstream primers for amplifying SNP3 are as shown in SEQ ID NO. 64 and SEQ ID NO. 159; the upstream and downstream primers for amplifying SNP4 are as shown in SEQ ID NO. 72 and SEQ ID NO. 167; the upstream and downstream primers for amplifying SNP5 are as shown in SEQ ID NO. 74 and SEQ ID NO. 169; the upstream and downstream primers for amplifying SNP6 are as shown in SEQ ID NO. 28 and SEQ ID NO. 123; the upstream and downstream primers for amplifying SNP7 are as shown in SEQ ID NO. 91 and SEQ ID NO. 186; the upstream and downstream primers for amplifying SNP8 are as shown in SEQ ID NO. 54 and SEQ ID NO. 149; the upstream and downstream primers for amplifying SNP9 are as shown in SEQ ID NO. 47 and SEQ ID NO. 142; the upstream and downstream primers for amplifying SNP10 are as shown in SEQ ID NO. 48 and SEQ ID NO. 143; the upstream and downstream primers for amplifying SNP11 are as shown in SEQ ID NO. 59 and SEQ ID NO. 154; the upstream and downstream primers for amplifying SNP12 are as shown in SEQ ID NO. 43 and SEQ ID NO. 138; the upstream and downstream primers for amplifying SNP13 are as shown in SEQ ID NO. 39 and SEQ ID NO. 134; and the upstream and downstream primers for amplifying SNP14 are as shown in SEQ ID NO. 86 and SEQ ID NO. 181;

[0022] The upstream primer of the primer pair of the second round of amplification is AATGATACGGCGACCACCGAGATCTACACnnnnnnnnTCTTTCCCTACACGACGCTCTTCCGATCT, and the downstream primer is CAAGCAGAAGACGGCATACGAGATnnnnnnnnGTGACTGGAGTTCCTTGGCACCCGAGAAT.

[0023] In the above primer set, nnnnnnnn is a barcode, nnnnnnnn is an 8-base barcode of a sample, and each n is one of A, T, C, and G.

[0024] In the present application, the primer set of the first round of amplification is primer set 1 of the first round of amplification, primer set 2 of the first round of amplification, and primer set 3 of the first round of amplification; the primer set 1 of the first round of amplification is the upstream and downstream primers for amplifying SNP1-SNP12; the primer set 2 of the first round of amplification is the upstream and downstream primers for amplifying SNP1-SNP13; and the primer set 3 of the first round of amplification is the upstream and downstream primers for amplifying SNP1-SNP14.

[0025] In the present application, 14 SNP sites and 5 environmental factors are respectively input into different machine learning models to optimize the type of the model, and finally it is found that the prediction accuracy of the logistic regression model is the highest, so the logistic regression model is preferred. The 14 SNPs correspond to reference sequence numbers rs938557, rs545854, rs6947309, rs7688672, rs7903456, rs2242206, rs3825017, rs55975541, rs435309, rs4684846, rs6837293, rs3751143, rs2941484, and rs208294. Then, based on the logistic regression model, different SNP combinations are input to further improve the prediction effect of the model. The results show that the prediction accuracy of the combination of 13 SNPs+5 environmental factors (marker combination 1), 12 SNPs+5 environmental factors (marker combination 3), and 14 SNPs+5 environmental factors (marker combination 2) is higher, and the highest is the combination of 13 SNP sites and 5 environmental factors (marker combination 2).

[0026] The present application provides an application of the above marker combination in the preparation of a product for predicting the risk of hyperuricemia turning into gout.

[0027] In the present application, the product is a reagent or a drug.

[0028] The application provides a system for predicting the risk of hyperuricemia turning into gout, comprising a data detection module, a data output module and a data analysis module, the data detection module is used for detecting the SNP combination genotype in the sample to be detected;

[0029] The data output module is used for outputting the genotype of each SNP site of the SNP combination in the sample to be detected;

[0030] The data analysis module is used for receiving the genotype of each SNP site of the SNP combination genotype obtained by the data output module, assigning values to the genotype of each SNP site and environmental factors respectively, and then substituting the assigned values into a calculation formula to calculate the risk score, and comparing the risk score with a threshold value to predict the high and low risk of hyperuricemia turning into gout.

[0031] In the application, the detection object genome can be obtained from a sample containing complete genome information such as blood, hair, etc. of the detection object. In some specific embodiments, the sample to be detected is peripheral blood of the detection object, and further is serum of the detection object. The genotype of each SNP site is assigned as follows: if the alleles of the SNP site genotype of the sample to be detected are all mutated, the value is 2; if one allele is mutated, the value is 1; and if no allele is mutated, the value is 0. For example, the genotype of the reference genome of rs1481012 is AA, if the genotype of the SNP site of a certain sample is GG, the alleles are all mutated, and n is 2; if the genotype of the SNP site of a certain sample is GA, one allele is mutated, and n is 1; and if the genotype of the site of a certain sample is AA, no allele is mutated, and n is 0. The environmental factors are assigned as follows: exercise frequency: never exercise is assigned as 0, exercise 1-2 days per month is assigned as 1, exercise 1-2 days per week is assigned as 2, exercise 3-4 days per week is assigned as 3, and exercise more than 5 days per week is assigned as 4; smoking frequency: non-smoking is assigned as 0, smoking frequency of less than 5 cigarettes per day is assigned as 1, and smoking frequency of more than 5 cigarettes per day is assigned as 2; sleep deprivation frequency: never sleep deprivation is assigned as 0, sleep deprivation 1-2 days per week is assigned as 1, sleep deprivation 3-4 days per week is assigned as 2, and sleep deprivation more than 5 days per week is assigned as 3; and drinking frequency: non-drinking is assigned as 0, drinking frequency of less than or equal to 250 mL per day is assigned as 1, and drinking frequency of more than 250 mL per day is assigned as 2.

[0032] In the application, the calculation formula is: risk score = 1 ÷ (1 + e -LogitP );

[0033] When the risk score is greater than or equal to the threshold value, the sample to be detected is determined as a high-risk gout patient; and when the risk score is less than the threshold value, the sample to be detected is determined as a low-risk gout patient.

[0034] In the present application, in some specific embodiments, the threshold value is 0.6, i.e. when the risk score ≥ 0.6, the sample to be tested is determined as a high risk of gout; when the risk score < 0.6, the sample to be tested is determined as a low risk of gout.

[0035] In the present application, when the marker combination is the above-mentioned marker combination 1, LogitP = -0.51-1.36xsport+0.05xBMI+0.21xalcohol+0.08xstay_up_late+0.91xsmoking+0.78xrs938557+0.2xrs545854-0.39xrs6947309-11.1xrs7688672+0.33xrs7903456+0.11xrs2242206-0.37xrs3825017-1.42xrs55975541-0.46xrs435309+0.29xrs4684846+10.67xrs6837293-0.68xrs3751143+0.42xrs2941484.

[0036] In the present application, the prediction of the risk of gout from hyperuricemia is the risk of whether hyperuricemia will develop into gout within 10 years. Specifically, the 10 years refers to the risk of whether a hyperuricemia patient will develop into gout within 1 year, 2 years, 3 years, 4 years, 5 years, 6 years, 7 years, 8 years, 9 years, etc. of course, it can also be less than 1 year, for example, within 11 months, within 10 months, within 9 months, within 8 months, within 7 months, within 6 months, within 5 months, within 4 months, within 3 months, within 2 months or within 1 month, etc.

[0037] The present application provides a readable storage medium, which stores computer instructions, and the instructions are executed by a processor to perform the above-mentioned system.

[0038] The application firstly preliminarily screens 95 single nucleotide polymorphism (SNP) sites related to gout in East Asians through literature retrieval; then, on the one hand, the serum of clinically defined gout patients and hyperuricemia patients is collected, and the target 95 single nucleotide polymorphism (SNP) sites are detected by using NGS-multiplex PCR targeted capture detection means, and the genotypes of the SNP sites are assigned (0, 1, 2); on the other hand, the basic information of the above-mentioned patients is collected, and is assigned according to the relevant grades, and further 14 SNP sites and 6 environmental factors are screened; then, based on the distribution of the 14 SNP sites and 5 environmental factors in the two groups of gout patients and hyperuricemia patients, machine learning such as random forest, KNN, SVM and logistic regression is carried out, and a corresponding model is established, and it is found that the prediction result of the logistic regression model is the best; in order to further improve the prediction effect of the model, the input SNP site + environmental factor combination is optimized, that is, the 14 SNP sites and 5 environmental factors are randomly arranged and combined, and input into the logistic regression model, and finally a logistic regression model for predicting the risk of hyperuricemia turning into gout is obtained, which is based on the calculation score (Score) and needs to input 13 SNP sites + 5 environmental factors, and the diagnostic efficiency of the final model is evaluated by ROC analysis (the specific process is shown in Figure 1 The model can be used to distinguish between people with hyperuricemia who are not prone to developing gout and people with hyperuricemia who are prone to turning into gout, or to screen out early gout patients from the population, or to predict whether an individual is an early gout patient or the likelihood of an individual developing early gout.

[0039] The application provides a marker combination for predicting the risk of hyperuricemia turning into gout, and based on the marker combination 1, the marker combination 2 and the marker combination 3, a corresponding diagnostic model is constructed, which has high accuracy, sensitivity, specificity and precision, and can realize accurate prediction of hyperuricemia turning into gout, so as to intervene in patients in time;

[0040] In the application, the above-mentioned model is analyzed by using logistic regression analysis, a calculation formula for the risk of hyperuricemia turning into gout is provided, and efficient prediction of individuals as high-risk gouters or low-risk gouters is realized;

[0041] The application provides a high-efficiency, rapid and low-cost gout risk assessment system, realizes non-invasive screening of gout risk by collecting peripheral blood for detection, the system is simple to operate, the result is simple and easy to understand, and has excellent prediction performance;

[0042] The present application constructs a marker combination 1 prediction model, which is convenient and fast, and the detection result is highly consistent with the clinical gold standard detection result, can accurately differential diagnose early gout patients, is beneficial to early discovery and early intervention, promotes early discovery and early treatment of gout, meets the urgent needs of the clinic, significantly reduces the cost of diagnosing early gout, and has good application prospect.

[0043] The present application focuses on integrating multi-omics data (such as genome, metabolome, epigenome, etc.), and combining advanced algorithms such as machine learning, to construct a dynamic risk prediction model with high prediction accuracy and wide applicability, providing strong support for the precise prevention and treatment of gout. The model not only helps to identify high-risk groups of gout early, but also provides a scientific basis for individualized intervention, thereby reducing the incidence of gout and improving patient prognosis.

[0044] Compared with the prior art, the present application has the following beneficial effects:

[0045] The present application provides a marker combination, system, readable storage medium and application for predicting the risk of hyperuricemia turning into gout, which constructs a corresponding diagnostic system by detecting the above-mentioned marker combination in the test sample. The system has high accuracy, sensitivity, specificity and precision in diagnosing the risk of hyperuricemia turning into gout, can realize accurate prediction of hyperuricemia turning into gout, thereby intervening in patients in time, realizing efficient prediction of individuals as high-risk or low-risk gouters, and realizing non-invasive screening of gout risk prediction. The present application uses the above-mentioned SNP combination genotype combined with environmental factors to diagnose the risk of gout, which is convenient and fast, and the detection result is highly consistent with the clinical gold standard detection result, can accurately differential diagnose early gout patients, is beneficial to early discovery and early intervention, promotes early discovery and early treatment of gout, reduces the incidence of gout, improves patient prognosis, meets the urgent needs of the clinic, significantly reduces the cost of diagnosing early gout, and has good application prospect. BRIEF DESCRIPTION OF DRAWINGS

[0046] Figure 1 The flowchart for constructing the model of the present application;

[0047] Figure 2 The ROC curve of the risk model based on 13 SNPs + 5 environmental factors provided by the present application;

[0048] Figure 3 The ROC curve of the risk model based on only 13 SNPs. DETAILED DESCRIPTION

[0049] The prediction of the present application refers to detecting or assaying the marker in the sample, or the content of the target marker, such as absolute content or relative content, and then indicating whether the individual providing the sample is likely to have or suffer from a certain disease, or the possibility of having a certain disease, by whether the target marker exists or how much the amount is; and in the present application, the marker refers to a specific SNP site, and the genotype of the SNP site is detected. The result of such prediction is not a direct result of suffering from a disease, but an intermediate result, and if a direct result is obtained, it needs to be confirmed by other auxiliary means such as pathology or dissection to confirm that a certain disease is suffered from. For example, the present application only provides a plurality of markers (SNPs) associated with the conversion of hyperuricemia to gout, and the genotype of the marker has a direct correlation with the conversion of hyperuricemia to gout.

[0050] The correlation of the marker of the present application with the conversion of hyperuricemia to gout refers to that the presence or content change of a certain marker (SNP in the present application) in the sample has a direct correlation with a specific disease or the progress of the disease, for example, the more the difference between the genotype of the SNP and the reference genome, the higher the possibility of suffering from the disease relative to healthy people, or the progress of the disease develops to be more serious or develops from a certain stage to another stage.

[0051] If a plurality of different markers appear or have a relative change in content in the sample, the possibility of suffering from the disease relative to healthy people is also higher. That is, among the types of markers, some markers have a strong correlation with the disease, some markers have a weak correlation with the disease, or some markers have no correlation with a certain specific disease. One or more of those markers with strong correlation can be used as a marker for diagnosing the disease, and those markers with weak correlation can be combined with strong markers to diagnose a certain disease, thereby increasing the accuracy of the detection result. Here, the disease can be the progress or development of the disease, such as developing from a better stage of a certain disease to a more malignant or serious stage, or even death.

[0052] The gold standard for gout diagnosis of the present application: the gold standard for gout diagnosis is to draw joint fluid by puncture, and to check the joint puncture fluid under a polarized light microscope, and to confirm gout if there are trihedral prism crystals of urate.

[0053] Single nucleotide polymorphism (SNP) of the present application: DNA sequence polymorphism caused by single nucleotide variation at the genome level.

[0054] The whole genome association analysis (GWAS) of the present application: an analysis method widely used in genetic research, which can evaluate the association between genetic variation and complex traits at the whole genome level. The core idea of GWAS is to use a large number of single nucleotide polymorphisms in the genome as molecular genetic markers, by large-scale genetic typing of the population, comparing the differences in each genetic variation and its frequency between the abnormal and control groups, and statistically analyzing the correlation between each variation and the target trait.

[0055] The second generation sequencing (NGS) of the present application is also called high-throughput sequencing, which refers to the conversion of DNA molecules into a large number of fragments by PCR amplification or library construction, and parallel sequencing on a high-throughput platform.

[0056] The multiplex PCR (multiplex PCR) of the present application, also known as multiplex primer PCR or composite PCR, is a PCR reaction that simultaneously amplifies multiple nucleic acid fragments by adding two or more primers in the same PCR reaction system. Its reaction principle, reaction reagents and operation process are the same as general PCR.

[0057] In the present application, all raw material components are commercially available products well known to those skilled in the art unless otherwise specified.

[0058] The technical solutions provided by the present application will be described in detail below in conjunction with the examples, but they should not be understood as limiting the scope of protection of the present application.

[0059] Example 1 Screening of SNP sites related to gout risk

[0060] This example provides the screening process of SNP sites related to gout risk, and the flowchart of this example is shown in Figure 1 , and the specific steps are as follows:

[0061] 1.1 Retrieval and preliminary screening of SNP sites

[0062] Search for SNP sites related to gout risk in East Asian population, and search for information of the above related SNP sites in the 1000 Genomes database, and the filtering criteria are as follows:

[0063] A. Remove duplicate sites reported in the literature;

[0064] B. Keep SNP sites that exist in the 1000 Genomes database;

[0065] C. Remove SNP sites with too high mutation rate;

[0066] The information of the finally screened 95 SNP sites is shown in Table 1:

[0067] Table 1 SNP sites after literature retrieval and screening

[0068]

[0069]

[0070]

[0071] Note: CHROM is the chromosome number; POS is the physical position of the SNP on the chromosome (based on the hg19 / GRCh37 reference genome); REF is the reference allele.

[0072] 1.2 Detection of SNP sites

[0073] 1.2.1 Study subjects

[0074] Based on the inclusion and exclusion criteria, 159 gout patients and 86 hyperuricemia patients treated in Taizhou Hospital of Zhejiang Province were recruited from August 2022 to May 2024. The inclusion criteria include: patients diagnosed with gout according to the 2015 ACR / EULAR gout classification criteria and the 2019 Chinese guidelines for diagnosis and treatment of hyperuricemia and gout, i.e. patients who developed gout from hyperuricemia; people with non-same-day fasting blood uric acid > 420 μmol / L (regardless of gender) for more than 2 times and without any gout symptoms at present or previously, and who did not develop gout after 3 years of follow-up were defined as hyperuricemia patients. The exclusion criteria include: gout exclusion criteria and hyperuricemia exclusion criteria. The gout exclusion criteria are as follows: history of malignant disease; refusal to sign the informed consent form; chronic renal failure (eGFR < 15 ml / min / 1.73m 2 ); The exclusion criteria for hyperuricemia are as follows: taking drugs that may cause high uric acid within two weeks before enrollment, such as thiazide diuretics, furosemide, pyrazinamide, aspirin, etc.; receiving uric acid-lowering drug treatment within two weeks before enrollment; history of malignant disease; refusal to sign the informed consent form; chronic renal failure (eGFR < 15 mL / min / 1.73m 2 ).

[0075] At the same time, the 159 gout patients and 86 hyperuricemia samples (total 245) collected were randomly divided into a training set (n = 170) and a test set (n = 75). The training set was used for model training, adjustment of hyperparameters and monitoring of model performance, and the test set was used for final evaluation of model performance (used in Example 3, section 3.3: Evaluation of the diagnostic performance of the prediction of hyperuricemia to gout risk model).

[0076] 1.2.2 Detection of SNP sites

[0077] NGS-multiplex PCR targeted capture technology is a targeted genome analysis method combining multiplex PCR amplification and high-throughput sequencing (NGS), mainly used for high-depth sequencing of specific genomic regions. In this embodiment, NGS-multiplex PCR targeted capture technology is used for SNP detection, and the specific experimental steps are as follows:

[0078] (1) Collect peripheral blood samples of the recruited subjects (n = 245), and extract genomic DNA using the blood extraction kit (QIAamp DNA Blood Mini Kit) of Qiagen Company.

[0079] (2) Using the genomic DNA in step (1) as a template, perform the first PCR amplification; the design principles of the primers used in the first PCR amplification are as follows: the amplification product is within 150-250 bp, to ensure that the sequencing mode of PE125 of the high-throughput sequencer can be sequenced on both ends; to ensure that the SNP site is located after 30 bp of the 3' end of the upstream primer and before 30 bp of the 5' end of the downstream primer; at the same time, avoid the presence of other SNP sites in the primer region; the specific primer is 20-25 bp, the GC content is 40-60%, the T m value is 60±10℃, and there are no four consecutive identical bases, to ensure that the amplification efficiency of all primers is close during multiplex amplification. A universal sequence is added to the 5' end of the upstream and downstream primers (upstream primer: 5'-3': CCTACACGACGCTCTTCCGATCT (SEQ ID NO. 1); downstream primer: 5'-3': CCTACACGACGCTCTTCCGATCT (SEQ ID NO. 2)), to facilitate the introduction of sequencing adapters P5 (AATGATACGGCGACCACCGAGATCTACACCAGTCTGCACACTCTTTCCCTACACGACGCTCTTCCGATC (SEQ ID NO. 3)) and P7 (CAAGCAGAAGACGGCATACGAGATGCACACGTGTGACTGGAGTTCAGACGTGTGCTCTTCCGATC (SEQ ID NO. 4)) to the amplification product during the second amplification, for subsequent sequencing; the specific amplification primers are shown in Table 2, and the synthesis of the primers is entrusted to Shanghai Heiyuan Biotechnology Co., Ltd. to mix the 95 pairs of synthesized primers into a 10 μM mixed solution in equal amounts; the amplification system and the amplification program are shown in Tables 3 and 4, respectively. Then, the RealCap multiplex PCR technology of Shanghai Heiyuan Biotechnology Co., Ltd. (https: / / www.homgen.com / capture / ) is used to monitor the multiplex PCR amplification process based on TaqMan fluorescent probe in real time.

[0080] Table 2 First round of specific amplification primers for 95 SNPs

[0081]

[0082]

[0083]

[0084]

[0085] Table 3 PCR amplification system

[0086] Composition Volume (μL) Template DNA 2 10x PCR buffer 2.5 dNTP mix (10 mM) 0.5 95 mixture of primers (10 μM) 2.5 Hot-start DNA polymerase 0.5 MgCl2(25 mM) 1.5 Deionized water Make up to 25

[0087] Table 4 PCR amplification procedure

[0088]

[0089] (3) Product purification

[0090] The PCR product was purified using purified magnetic beads (Agencourt AMPure XP, manufacturer: Beckman Coulter). First, the PCR product was mixed with magnetic beads and incubated to allow the DNA to be adsorbed onto the magnetic beads. The magnetic beads were separated using a magnetic stand, and the supernatant was discarded. The magnetic beads were washed several times with washing buffer to remove impurities. Finally, the DNA was eluted from the magnetic beads using elution buffer to obtain the purified DNA product. The final product was dissolved in 50 μL of nuclease-free water. One-third of the volume of the purified PCR product was used for subsequent experiments.

[0091] (4) The second round of PCR amplification was performed using the amplification primers in Table 5, and the amplification system and procedure were shown in Tables 6 and 7, respectively. The PCR product was purified using purified magnetic beads, and the experimental steps were the same as described above. The concentration of the purified product was detected using a micro UV spectrophotometer, and the concentration was 5 ng / μL to 10 ng / μL, which was more appropriate.

[0092] Table 5 Universal amplification primer sequence

[0093]

[0094]

[0095] Note: Uni-F represents the upstream universal primer of the second round of amplification, TCTTTCCCTACACGACGCTCTTCCGATCT (SEQ ID NO. 192) can be combined with the first round of specific amplification product, nnnnnnnn is a barcode for distinguishing different samples, one barcode for each sample; Uni-R represents the downstream second round of universal amplification primer, GTGACTGGAGTTCCTTGGCACCCGAGAAT (SEQ ID NO. 194) can be combined with the first round of specific amplification product, nnnnnnnn is a barcode for distinguishing different samples, one barcode for each sample; all primers are synthesized by Shanghai Heiyin Biotechnology Co., Ltd.; and diluted to 20 pM with nuclease-free water, -20℃ for standby.

[0096] wherein nnnnnnnn is an 8-base barcode of the sample. Each n is one of A, T, C, G. By changing the combination of the 8 bases, hundreds of different indexes can be generated. In the dual index design, one sample has two independent barcodes (i7 and i5), which greatly improves the number of sample mixing and the accuracy of identification.

[0097] Table 6 Second round PCR amplification system

[0098] Composition System (total system 30 μL) Uni-F (10 μM) 1 μL Uni-R (10 μM) 1 μL DNA elution product 18 μL 3x enzyme premix 10 μL

[0099] Table 7 Second round PCR amplification program

[0100]

[0101] (5) Sequencing

[0102] According to the operation instruction of Illumina sequencing platform, the PE (Pair-End) sequencing mode of double-end 125bp or double-end 150bp is adopted, and the average sequencing depth of each sample is set to 200x; after sequencing, the low-quality sequences with quality value lower than Q30 are removed for quality control; then, the filtered high-quality sequences are compared and analyzed with the HG19 reference genome, and finally the mutation of each SNP (single nucleotide polymorphism) site in each sample is determined.

[0103] (6) Collect environmental factor variables of the detection object by questionnaire investigation and other methods, the variables include: basic information such as name, gender, age, mobile phone number, education, residence, height, weight, occupation, family income, BMI index; disease conditions such as past medical history (diabetes history, hypertension history, coronary heart disease history, liver and kidney basic disease history), family history (whether suffering from hyperuricemia or gout within three generations) and whether regular medication; living habits such as whether smoking, drinking, staying up late, exercising, sitting for a long time and the frequency of the above behaviors; eating habits such as the frequency of eating seafood, poultry meat (animal offal, meat), bean products, poultry eggs / milk and dairy products, vegetables / fruit, drinking water / beverage / coffee / wine; examination results such as whether there are tophi, whether it involves joints, final diagnosis, etc.

[0104] (7) Collect and clean the data obtained from steps (5) and (6), obtain the data set including target variables (whether gout) and characteristic variables (each SNP mutation, environmental variables), process missing values and outliers; identify and fill in missing values, and remove outliers; divide the data set into a training set (n = 170) and a test set (n = 75). The training set is used for model training, adjusting hyperparameters and monitoring model performance, and the test set is used for final evaluation of model performance. Then the SNP sites are numerically valued (assigned) according to the classification characteristics: record the number of risk bases as n, if the result is with one risk base, then n is 1, if with two risk bases, then n is 2, and without risk genes, then n is 0. Specifically, if the genotype of rs1481012 reference allele is AA, and the genotype of a certain sample at this site is GG, and the allele is all mutated, then n is 2; if the genotype of a certain sample at this site is GA, and one allele is mutated, then n is 1; if the genotype of a certain sample at this site is AA, and no allele is mutated, then n is 0. Exercise frequency: 0 for never exercising, 1 for 1-2 days of exercise per month, 2 for 1-2 days of exercise per week, 3 for 3-4 days of exercise per week, and 4 for more than 5 days of exercise per week, wherein the exercise frequency is a non-integer number of days, and the value is rounded, such as 2.5 days of exercise per week by default as 3 days, and 4.5 days of exercise per week by default as 5 days; Smoking frequency: 0 for not smoking, 1 for low smoking frequency (5 or less per day), and 2 for high smoking frequency (more than 5 per day); Night work frequency: 0 for never working overtime, 1 for rarely working overtime (1-2 days per week), 2 for moderate frequency of working overtime (3-4 days per week), and 3 for frequently working overtime (more than 5 days per week); Drinking frequency: 0 for not drinking, 1 for low drinking frequency (≤250mL per day), and 2 for high drinking frequency (>250mL per day); Seafood frequency: 0 for not eating seafood, 1 for low seafood frequency (>0 and ≤150g per week), and 2 for high seafood frequency (>150g per week); BMI is a continuous variable, and the specific value is the assigned value.

[0105] 1.3 Further screening of SNP sites related to the risk of gout from hyperuricemia

[0106] The SNP sites and environmental factors of the gout risk-related genes obtained through preliminary screening are more, which increases the time required for detection and analysis. Moreover, the number of SNP sites / environmental factors is not the more the better, which can introduce more heterogeneity and pleiotropy, which means that some SNP / environmental factors can be related to multiple traits at the same time, thereby leading to the decline of the explanatory ability of the model. In addition, too many SNP sites / environmental factors will increase the complexity of the model, leading to the risk of overfitting. Therefore, the feature (including SNP sites and environmental factors) extraction in this embodiment is based on the feature selection method of average impurity reduction of random forest, and the specific steps are as follows: the number of trees in the random forest ntree=500, the number of variables randomly selected when each tree is split mtry=sqrt(p), 5 times of 10-fold cross-validation; by randomly shuffling the values of the variables, the degree of performance (such as OOB error) reduction of the model is observed, the variables with importance higher than the average value are reserved, 21 sites and 6 environmental factors are reserved; then the optimal number of variables is selected through recursive feature elimination (RFE), and finally 14 SNP sites and 6 environmental factors are screened and included in the subsequent analysis. The 14 SNP sites include: rs938557, rs545854, rs6947309, rs7688672, rs7903456, rs2242206, rs3825017, rs55975541, rs435309, rs4684846, rs6837293, rs3751143, rs2941484, and rs208294; the 6 environmental factors are: sport, BMI, smoking, alcohol, stay_up_late, and seafood, wherein the assignment standards of each environmental factor are referred to the contents in the step (7) in the part of 1.2.2 of this embodiment. The detailed information of the 14 SNP sites related to gout risk genes and 6 environmental factors is shown in Table 8.

[0107] Table 8 Detailed information of SNP sites related to gout risk genes and environmental factors included in the model

[0108]

[0109]

[0110] Note: Estimate refers to the parameter estimate or effect size of the model output, the higher the value, the more significant the positive or negative impact of the variable on the dependent variable; Std. Error refers to the standard error, the smaller the standard error, the smaller the difference between the sample statistic and the population parameter, the higher the reliability of the estimate; z value refers to the standard score; AUC (Area under the curve) refers to the area under the ROC curve, the closer the AUC is to 0.5, the lower the diagnostic value of the single SNP; the closer the AUC is to 1, the higher the diagnostic value of the SNP; similarly, the 95% confidence interval of the AUC value, the closer to 1, the higher the diagnostic value and reliability of the SNP; the Youden index is the sum of sensitivity and specificity minus 1, which represents the total ability of the screening method to find true patients and non-patients when the harm of false negatives (missed diagnosis rate) and false positives (misdiagnosis rate) is assumed to be equal. The larger the index, the better the effect of the screening experiment and the greater the authenticity, therefore, the SNP corresponding to the maximum Youden index is the best SNP, and the closer the corresponding ROC sensitivity and specificity to 100%, the higher the diagnostic performance of the SNP.

[0111] As can be seen from Table 8, the Estimate values are mainly concentrated in -0.9-0.9, the standard errors are concentrated in 0.2-0.4, and the absolute values of the standard scores are concentrated in 0.01-3.91; the AUC values of the single SNPs obtained by screening are mostly low (0.60±0.11), and the maximum Youden index of most factors is less than 0.2; for the same SNP site, the higher the sensitivity, the lower the specificity; vice versa; it is worth noting that all the parameters of the exercise frequency factor are significantly better than those of other factors, which indicates that exercise can help reduce the risk of high uric acid patients developing into gout. In addition, since the data results of the seafood consumption frequency factor are negatively correlated with the incidence of gout, the seafood consumption factor is excluded, that is, the factor is not considered in the subsequent model construction. In general, the performance of single SNPs in predicting the risk of gout is generally low, so further exploration is still needed to further improve the diagnostic ability and diagnostic performance of the model for early gout.

[0112] Example 2 Preliminary construction of a model for predicting the risk of high uric acid developing into gout

[0113] Fourteen SNP sites (rs938557, rs545854, rs6947309, rs7688672, rs7903456, rs2242206, rs3825017, rs55975541, rs435309, rs4684846, rs6837293, rs3751143, rs2941484 and rs208294, see Table 1 for details) and five environmental factors (Sport, BMI, smoking, alcohol and stay_up_late) were screened in Example 1, wherein the assignment criteria of each environmental factor can be found in the content of step (7) in Section 1.2.2 of Example 1. Based on the above-mentioned fourteen SNP sites and five environmental factors, different machine learning models (including random forest, SVM, KNN and Logistic regression) were constructed in this example, the parameters of each model can be found in Table 9, and the test set was predicted using the above-mentioned models, and the specific prediction results can be found in Table 9.

[0114] Table 9 Hyperparameters of machine learning models

[0115] Model type Hyperparameter AUC Sensitivity Specificity Random forest mtry: 12; min_n: 500 0.810 0.864 0.633 SVM degree: 3; C: 1.0; kernel: rbf 0.787 0.600 0.903 KNN Neighbors: 5; weight_func: uniform; dist_power: 2 0.758 0.750 0.682 Logistic regression Penalty: l2; C = 1.0; max_iter: 100 0.854 0.818 0.800

[0116] As can be seen from Table 9, the overall performance of the logistic regression model is the best, and the AUC value, sensitivity and specificity are all above 0.8; followed by the random forest model, therefore the logistic regression model is preferred, which may be because the logistic regression performs more stably in the small sample or low effect SNP scenario.

[0117] Example 3 Further development and verification of the risk model for predicting the conversion of hyperuricemia to gout

[0118] 3.1 Development of logistic regression model

[0119] To further improve the prediction accuracy of the model, the embodiment combines different SNP sites and environmental factors based on a logistic regression model to screen the best combination of SNP sites and environmental factors. It should be understood that a single SNP site / environmental factor with higher prediction accuracy of gout risk may not necessarily play a greater role in the combination after being combined with one or more other SNP sites and environmental factors, and the number of different SNP sites and environmental factors is not necessarily more, and the prediction accuracy (AUC value) of the combination is not necessarily higher. Therefore, the embodiment randomly includes different SNP sites (the number of sites is 1-14) and environmental factors (the number of environmental factors is 1-5) based on a logistic regression model and taking the training set as the experimental object to construct a corresponding model for predicting the risk of hyperuricemia turning into gout, and the performance of different models is compared to screen the best combination of SNP and environmental factors. Since 19 factors (including 14 SNP sites and 5 environmental factors) can form thousands of combinations, and the present invention is limited in length, only the best three combinations are shown, and the specific results are shown in Table 10 and Figure 2 .

[0120] Table 10 Parameters of the prediction model based on the combination of SNP sites and environmental factors related to the risk of gout

[0121]

[0122] From Figure 2 and Table 10, it can be seen that the overall prediction effect of the model constructed based on multiple SNP and environmental factor combinations is better than that of a single SNP or a single environmental factor, and therefore, the SNP+environmental factor combination is preferred to construct a model for predicting the risk of hyperuricemia turning into gout, wherein the AUC value of model A is 0.856, the confidence interval is 0.806-0.908, and the specificity is 0.850. These parameters are better than those of model B and model C, which may be because too many SNP sites and environmental factors may introduce more heterogeneity and pleiotropy, leading to a decrease in the prediction ability of the model, which indicates that the diagnostic performance of model A constructed by including 13 SNP sites and 5 environmental factors is the highest.

[0123] In summary, the 13 SNP sites and 5 environmental factors are preferred to construct the model, and the SNP sites are specifically: rs938557, rs545854, rs6947309, rs7688672, rs7903456, rs2242206, rs3825017, rs55975541, rs435309, rs4684846, rs6837293, rs3751143, and rs2941484; and the environmental factors are sport, BMI, alcohol, stay_up_late, and smoking.

[0124] In order to more intuitively predict the risk of hyperuricemia developing into gout, the present embodiment is based on the optimization results of the combination of model types and input SNP sites + environmental factors, that is, on the basis of the combination of 13 SNP sites + 5 environmental factors and the logistic regression model, a stepwise logistic regression is used to obtain a formula for calculating the risk of hyperuricemia developing into gout, and the SNP sites and environmental factors are substituted into the formula by assignment, the Score value is calculated, and finally the risk of the detection object (i.e. the hyperuricemia patient) developing into gout is predicted by comparing the Score value with the threshold value.

[0125] Specifically, first, all possible independent variables (predictive variables) are put into the candidate pool, and the dependent variable (target variable) remains unchanged; when initializing the model, a model containing only the intercept term is selected as the starting point, and the variable with the smallest p value is selected to join the current model. If the p value of all candidate variables is greater than the set significance level, stop introducing variables, when the process of introducing new variables and removing variables no longer changes the model, the stepwise regression algorithm terminates, and a calculation formula is obtained; the calculation formula is: risk score = 1 / (1+e -LogitP ), wherein Score = 1 / (1+e -LogitP), wherein LogitP = -0.51 - 1.36 x sport + 0.05 x BMI + 0.21 x alcohol + 0.08 x stay_up_late + 0.91 x smoking + 0.78 x rs938557 + 0.2 x rs545854 - 0.39 x rs6947309 - 11.1 x rs7688672 + 0.33 x rs7903456 + 0.11 x rs2242206 - 0.37 x rs3825017 - 1.42 x rs55975541 - 0.46 x rs435309 + 0.29 x rs4684846 + 10.67 x rs6837293 - 0.68 x rs3751143 + 0.42 x rs2941484; the different factors are assigned and substituted into the calculation formula, and the assignment rules are as follows: rs938557, rs545854, rs6947309, rs7688672, rs7903456, rs2242206, rs3825017, rs55975541, rs435309, rs4684846, rs6837293, rs3751143, rs2941484 refer to the detection results of 13 SNPs, wherein the homozygous wild type, i.e. no mutation, is assigned as 0, the heterozygous mutation is assigned as 1, and the homozygous mutation is assigned as 2 (as described in Example 1, if all alleles are mutated, the value is 2; if one allele is mutated, the value is 1; if no allele is mutated, the value is 0); sport, BMI, alcohol, stay_up_late, smoking, five environmental factor indexes, wherein sport: never exercise is assigned as 0, 1-2 days of exercise per month is assigned as 1, 1-2 days of exercise per week is assigned as 2, 3-4 days of exercise per week is assigned as 3, and 5 days or more per week is assigned as 4, wherein the exercise frequency is a non-integer number of days, and the value is rounded off, such as 2.5 days of exercise per week is defaulted as 3 days, and 4.5 days of exercise per week is defaulted as 5 days; smoking: non-smoking is 0, low smoking frequency (5 or less per day) is assigned as 1, and high smoking frequency (more than 5 per day) is assigned as 2; stay_up_late: never stay up late is assigned as 0, rarely (1-2 days / week) is assigned as 1, moderately (1-2 days / week) is assigned as 1, and often (5 days or more / week) is assigned as 2; alcohol: non-drinking is assigned as 0, low drinking frequency (≤250 mL / day) is assigned as 1, and high drinking frequency (>250 mL / day) is assigned as 2; BMI is a continuous variable, and the specific value is the assignment value; the result judgment standard is set as P (threshold) = 0.6; when the data (Score) obtained by the calculation formula is ≥P, the subject is determined to be a high risk of gout; when the data (Score) obtained by the calculation formula is <P, the subject is determined to be a low risk of gout.

[0126] 3.2 Detailed description of the risk model for predicting the outcome of hyperuricemia to gout

[0127] The foregoing describes the construction process of the gout risk prediction model, and this embodiment will take a patient with hyperuricemia turning into gout as an example to fully describe the detailed implementation of the model, and the specific steps are as follows:

[0128] (1) Extract the genomic DNA of the gout patient according to the steps described in Embodiment 1, and detect the SNP sites using the NGS-multiplex PCR targeted capture technology, and the SNP site assignment method is the same as described in Embodiments 1-2 (if the alleles of the genotype of a certain sample are all mutated, then assign 2; if one allele is mutated, then assign 1; if no allele is mutated, then assign 0); and consult the patient's BMI index, frequency of staying up late, exercise, frequency of drinking and smoking, and the assignment rules are the same as described above (exercise frequency: never exercise is assigned 0, exercise 1-2 days per month is assigned 1, exercise 1-2 days per week is assigned 2, exercise 3-4 days per week is assigned 3, and exercise 5 days or more per week is assigned 4, wherein the exercise frequency is the number of non-integral days, and the value is rounded off, such as 2.5 days of exercise per week is defaulted to 3 days, and 4.5 days of exercise per week is defaulted to 5 days; smoking frequency: non-smoking is 0, low smoking frequency (5 or less per day) is assigned 1, and high smoking frequency (more than 5 per day) is assigned 2; staying up late frequency: never staying up late is assigned 0, rarely staying up late (1-2 days per week) is assigned 1, staying up late frequency is moderate (1-2 days per week) is assigned 1, and frequently staying up late (5 days or more per week) is assigned 2; drinking frequency: non-drinking is assigned 0, low drinking frequency (≤250 mL per day) is assigned 1, and high drinking frequency (>250 mL per day) is assigned 2, see Table 11 for specific assignment results. It should be understood that as long as the 13 SNP sites are detected, it is not necessary to use the NGS-multiplex PCR targeted capture technology, and this is only a scheme for obtaining SNP site information.

[0129] Table 11 Detailed information of SNP sites of gout patients

[0130] Exercise frequency BMI Frequency of alcohol consumption Frequency of staying up late Frequency of smoking 0 25.95 1 3 1 rs938557 rs545854 rs6947309 rs7688672 rs7903456 2 1 1 1 1 rs2242206 rs3825017 rs55975541 rs435309 rs4684846 2 2 0 1 1 rs6837293 rs3751143 rs2941484 1 1 1

[0131] (2) Substitute the data in Table 11 into the calculation formula, and the calculation formula is as follows: Score = 1 / (1+e -LogitP), wherein LogitP = -0.51 - 1.36 x sport + 0.05 x BMI + 0.21 x alcohol + 0.08 x stay_up_late + 0.91 x smoking + 0.78 x rs938557 + 0.2 x rs545854 - 0.39 x rs6947309 - 11.1 x rs7688672 + 0.33 x rs7903456 + 0.11 x rs2242206 - 0.37 x rs3825017 - 1.42 x rs55975541 - 0.46 x rs435309 + 0.29 x rs4684846 + 10.67 x rs6837293 - 0.68 x rs3751143 + 0.42 x rs2941484. Specifically, Score = 1 / (1 + e -LogitP ), wherein LogitP = -0.51 - 1.36 x 0 + 0.05 x 25.95 + 0.21 x 1 + 0.08 x 3 + 0.91 x 1 + 0.78 x 2 + 0.2 x 1 - 0.39 x 1 - 11.1 x 1 + 0.33 x 1 + 0.11 x 2 - 0.37 x 2 - 1.42 x 0 - 0.46 x 1 + 0.29 x 1 + 10.67 x 1 - 0.68 x 1 + 0.42 x 1; according to the judgment criteria (as described in section 3.1 of Example 3), Score = 0.9218, greater than 0.6, and the sample is considered to be at high risk of gout. The prediction result is consistent with the actual situation of the test object.

[0132] 3.3 Evaluation of the diagnostic performance of the gout risk prediction model

[0133] To further evaluate the prediction performance of the optimized model, the 245 samples collected in Example 1 were randomly divided into validation sample set 1 (n = 170) and validation sample set 2 (n = 75). According to the implementation described in "3.2", the gout risk of validation sample set 1 and validation sample set 2 was predicted respectively. When the prediction result is consistent with the actual situation, it is considered to be accurate, and the MAE and accuracy of different data sets are calculated. The specific results are shown in Table 12.

[0134] Table 12 Evaluation of the performance of the gout prediction model

[0135] Index Validation sample set 1 Validation sample set 2 RMSE (root mean square error) 0.479 0.385 MAE (mean absolute error) 0.229 0.122 Accuracy 0.771 0.878 Pearson correlation test 0.476 0.729

[0136] From the results of Table 12, the model constructed based on 13 SNP sites and 5 environmental factors combined to predict the RMSE value of different data sets is 0.385-0.479, the MAE value is 0.122-0.229, the accuracy is 0.771-0.878, and the p value is less than 0.05. These data all show that when facing different types of samples, the prediction results of the model have high reliability, which can be used to effectively predict the risk of high uric acid outcome to gout in clinical practice, provide personalized treatment for patients, and improve treatment effect.

[0137] Comparative Example 1

[0138] At present, there is no similar method based on SNP sites and environmental factors and other variables combined with machine learning technology to construct a big data model to predict the risk of high uric acid outcome to gout. On the other hand, there are also some patents that use other clinical routine variables for risk prediction, for example, patent 202210117699.4 (application number) screens out key variables closely related to gout as feature variables through Lasso regression, constructs sample data, and uses the Naive Bayes algorithm to train the classifier to predict the probability of gout in the subjects, thereby achieving the evaluation of the future gout incidence risk of the target population. However, this patent 202210117699.4 mainly focuses on the protection of the gout prediction model system, device and storage medium, and the feature variables lack innovation, and its value in clinical application is relatively limited. Compared with this patent 202210117699.4, the present application has more significant advantages in clinical application.

[0139] Comparative Example 2

[0140] Based on this, the present comparative example removes the environmental and lifestyle-related factors mentioned in the present application and only leaves the detection results of the 13 SNP sites of the detection object, with the same assignment rules as described in Example 1 (if all alleles are mutated, assign a value of 2; if one allele is mutated, assign a value of 1; if no allele is mutated, assign a value of 0); and also constructs a model through the logistic regression method to obtain a new formula and threshold value to determine the standard for judging the risk of developing gout in the detection object. The formula is: Score = 1 ÷ (1 + e -LogitP), wherein LogitP = 0.84 + 0.55xrs938557 + 0.17xrs545854 - 0.61xrs6947309 - 12.69xrs7688672 + 0.28xrs7903456 + 0.05xrs2242206 - 0.2xrs3825017 - 0.69xrs55975541 - 0.52xrs435309 + 0.18xrs4684846 + 12.35xrs6837293 - 0.38xrs3751143 + 0.16xrs2941484. The specific prediction results are shown in Table 13. Figure 3 and Table 13.

[0141] The detection object of the present comparative example is the 245 samples collected in Example 1, which are randomly divided into a training set (n = 170), i.e. 170 samples.

[0142] Performance comparison of different risk models

[0143] Detection index Risk model based on SNP site only (comparative example 2) Risk model based on SNP site and environmental factors Specificity 79.1% 85.0% Sensitivity 61.7% 76.4% Accuracy 0.676 0.771 Kappa value 0.230 0.463 AUC value 0.717(0.655-0.813) 0.856(0.801-0.918) Optimal critical value 0.600 0.600 Youden index 0.408 0.614

[0144] Note: The risk model based on SNP sites and environmental factors is the calculation formula described in Section 3.2 of Example 3 and the calculation formula of LogitP, and the SNP sites are rs938557, rs545854, rs6947309, rs7688672, rs7903456, rs2242206, rs3825017, rs55975541, rs435309, rs4684846, rs6837293, rs3751143 and rs2941484; the environmental factors are sport, BMI, alcohol, stay_up_late and smoking, wherein the assignment standards of each environmental factor are described in the content of Step (7) in Section 1.2.2 of Example 1.

[0145] From Table 13 and Figure 3The results show that the AUC (95% CI) of the model for predicting the risk of gout from hyperuricemia based on 13 SNP sites is 0.717 (0.655-0.813), the optimal critical value is 0.6, the sensitivity is 61.7%, the specificity is 79.1%, and the Youden index is 0.408; while the AUC value of the risk model combining genetics and environmental lifestyle of the application is 0.856 (0.801-0.918), the optimal critical value is 0.6, the sensitivity is 76.4%, the specificity is 85.0%, the accuracy is 0.771, the Youden index is 0.614, and the kappa value is 0.463, in other words, the indicators of the risk model of the application are all better than those of the risk model based on only genetic SNP data, indicating that the risk assessment model of pure genetic SNP is not as good as the genetic + environmental and lifestyle assessment model provided by the application in terms of accuracy, specificity or sensitivity, and therefore the genetic + environmental and lifestyle combination is preferred to construct the gout risk assessment model.

[0146] All patents and publications mentioned in the specification are indicative of the levels of those skilled in the art to which the application pertains, and are incorporated by reference. All patents and publications referenced in this specification are indicative of the level of those skilled in the art to which the application pertains, and are hereby incorporated by reference, each in its entirety. The application described herein can be implemented in the absence of any element or elements, or in the presence of one or more of the limitations described herein, unless otherwise specifically stated. For example, the terms "comprising," "consisting essentially of and "consisting of" as used herein can be replaced with either of the remaining two terms. The term "a" as used herein means "one," unless otherwise specifically stated. The term "one" as used herein means "one or more," unless otherwise specifically stated. The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention that in the use of such terms and expressions of describing the subject application the patentee hereby disclaim any equivalents of the features referred to except as set forth in the claims. It is recognized that certain embodiments described herein can be used in a variety of combinations, and it is intended that the application extend to all such combinations. It is also recognized that the examples described herein are intended to be exemplary only and that variations and modifications can be made by those skilled in the art without departing from the spirit and scope of the application.

Claims

1. A combination of biomarkers for predicting the risk of hyperuricemia progressing to gout, characterized in that, The biomarker combination comprises reagents and environmental factors used to detect SNP combination genotypes; the SNP combination comprises one or more of SNP1 to SNP14, and the reference sequence numbers corresponding to SNP1 to SNP14 are rs938557, rs545854, rs6947309, rs7688672, rs7903456, rs2242206, rs3825017, rs55975541, rs435309, rs4684846, rs6837293, rs3751143, rs2941484, and rs208294, respectively. The SNP locus information for the SNP combinations is as follows: The environmental factors mentioned are one or more of the following: exercise frequency, BMI index, alcohol consumption frequency, smoking frequency, and frequency of staying up late.

2. The marker combination according to claim 1, characterized in that, The combination of markers can be any one of marker combination 1, marker combination 2, and marker combination 3; The biomarker combination 1 consists of reagent 1 and environmental factors for detecting SNP combination genotypes, wherein the SNP combination is SNP1 to SNP13, and the environmental factors are exercise frequency, BMI index, alcohol consumption frequency, smoking frequency, and staying up late frequency. The biomarker combination 2 consists of reagent 2 and environmental factors for detecting SNP combination genotypes, wherein the SNP combination is SNP1 to SNP14, and the environmental factors are exercise frequency, BMI index, alcohol consumption frequency, smoking frequency, and staying up late frequency. The biomarker combination 3 consists of reagent 3 and environmental factors used to detect SNP combination genotypes, wherein the SNP combination is SNP1 to SNP12, and the environmental factors are exercise frequency, BMI index, alcohol consumption frequency, smoking frequency, and staying up late frequency.

3. The marker combination according to claim 1, characterized in that, The reagents include a primer set.

4. The marker combination according to claim 3, characterized in that, The primer set includes the primer set for the first round of amplification and the primer pair for the second round of amplification; The primer set for the first round of amplification consists of one or more upstream and downstream primers for amplifying SNP1 to SNP14. Specifically, the upstream and downstream primers for amplifying SNP1 are shown in SEQ ID NO. 76 and SEQ ID NO. 171; the upstream and downstream primers for amplifying SNP2 are shown in SEQ ID NO. 53 and SEQ ID NO. 148; the upstream and downstream primers for amplifying SNP3 are shown in SEQ ID NO. 64 and SEQ ID NO. 159; the upstream and downstream primers for amplifying SNP4 are shown in SEQ ID NO. 72 and SEQ ID NO. 167; the upstream and downstream primers for amplifying SNP5 are shown in SEQ ID NO. 74 and SEQ ID NO. 169; the upstream and downstream primers for amplifying SNP6 are shown in SEQ ID NO. 28 and SEQ ID NO. 123; the upstream and downstream primers for amplifying SNP7 are shown in SEQ ID NO. 91 and SEQ ID NO. 186; the upstream and downstream primers for amplifying SNP8 are shown in SEQ ID NO. 54 and SEQ ID NO. 149; and the upstream and downstream primers for amplifying SNP9 are shown in SEQ ID NO.

149. The primers for amplifying SNP10 are shown in SEQ ID NO. 47 and SEQ ID NO. 142; the primers for amplifying SNP10 are shown in SEQ ID NO. 48 and SEQ ID NO. 143; the primers for amplifying SNP11 are shown in SEQ ID NO. 59 and SEQ ID NO. 154; the primers for amplifying SNP12 are shown in SEQ ID NO. 43 and SEQ ID NO. 138; the primers for amplifying SNP13 are shown in SEQ ID NO. 39 and SEQ ID NO. 134; and the primers for amplifying SNP14 are shown in SEQ ID NO. 86 and SEQ ID NO.

181. The upstream primer for the second round of amplification was AATGATACGGCGACCACCGAGATCTACACnnnnnnnnTCTTTCCCTACACGACGCTCTTCCGATCT; the downstream primer was CAAGCAGAAGACGGCATACGAGATnnnnnnnnGTGACTGGAGTTCCTTGGCACCCGAGAAT.

5. The use of the biomarker combination according to any one of claims 1 to 4 in the preparation of a product for predicting the risk of hyperuricemia progressing to gout.

6. A system for predicting the risk of hyperuricemia progressing to gout, characterized in that, It includes a data detection module, a data output module, and a data analysis module. The data detection module is used to detect the SNP combination genotypes in any one of claims 1 to 4 in the sample to be tested. The data output module is used to output the genotype of each SNP locus in any one of the SNP combinations described in claims 1 to 4 in the sample to be tested; The data analysis module receives the genotypes of each SNP locus from the SNP combination genotype obtained from the data output module. It assigns values ​​to the genotypes and environmental factors of each SNP locus, and then substitutes these values ​​into the calculation formula to calculate the risk score. Based on the comparison between the risk score and the threshold, it predicts the risk of hyperuricemia turning into gout.

7. The system according to claim 6, characterized in that, The genotype values ​​for each SNP locus are assigned as follows: if all alleles of the SNP locus genotype in the sample to be tested are mutated, the value is 2; if one allele is mutated, the value is 1; if there is no allele mutation, the value is 0. The method for assigning values ​​to environmental factors is as follows: Exercise frequency: Never exercise is assigned a value of 0, exercising 1-2 times per month is assigned a value of 1, exercising 1-2 times per week is assigned a value of 2, exercising 3-4 times per week is assigned a value of 3, and exercising more than 5 days per week is assigned a value of 4. Smoking frequency: 0 for non-smokers, 1 for smoking frequency less than 5 cigarettes, and 2 for smoking frequency greater than 5 cigarettes. Frequency of staying up late: Never staying up late is assigned a value of 0, staying up late for 1-2 days / week is assigned a value of 1, staying up late for 3-4 days / week is assigned a value of 2, staying up late for more than 5 days / week is assigned a value of 3; Drinking frequency: No drinking is assigned a value of 0; drinking frequency ≤ 250mL / talent value is 1, drinking frequency > 250mL / talent value is 2.

8. The system according to claim 6, characterized in that, The calculation formula is: Risk Score = 1 ÷ (1 + e) -LogitP ); When the risk score is greater than or equal to the threshold, the sample is determined to be at high risk of gout; when the risk score is less than the threshold, the sample is determined to be at low risk of gout.

9. The system according to claim 8, characterized in that, When the combination of markers is the combination of markers 1 as described in claim 2, LogitP = -0.51 - 1.36 × sport + 0.05 × BMI + 0.21 × alcohol + 0.08 × stay_up_late + 0.91 × smoking + 0.78 × rs938557 + 0.2 × rs545854 - 0.39 × rs6947309 - 11. 1×rs7688672+0.33×rs7903456+0.11×rs2242206-0.37×rs3825017-1.42×rs55975541-0.46×rs435309+0.29×rs4684846+10.67×rs6837293-0.68×rs3751143+0.42×rs2941484.

10. A readable storage medium having computer instructions stored thereon, characterized in that, When the instruction is executed by the processor, the system described in any one of claims 6 to 9 is executed.

Citation Information

Patent Citations

  • Gout prediction model system, equipment and storage medium

    CN114512240A