Urine protein marker for diagnosing helicobacter pylori and application, kit and risk assessment model thereof

Through the detection of PHB1, C20orf202, CP and P0DOX5 protein markers in urine, combined with proteomics and machine learning, the high cost and complex operational problems of Helicobacter pylori diagnosis are solved, early diagnosis and risk assessment are achieved, and the risk of disease is reduced.

CN120254268APending Publication Date: 2025-07-04CIMING HEALTH CHECKUP MANAGEMENT GRP WUHAN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510196246.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, the invasive and non-invasive detection of Helicobacter pylori is expensive and cumbersome to operate, and is not suitable for a nationwide census, and lacks low-priced and effective census methods.

Method used

Proteomics combined with machine learning were used to identify the expression levels of PHB1, C20orf202, CP and P0DOX5 protein markers in urine, and early diagnosis of Helicobacter pylori was established by establishing a risk assessment model, and these protein markers were detected using immunoblotting, enzyme-linked immunosorbent assays and other methods.

Benefits of technology

The early diagnosis of Helicobacter pylori has been achieved, the detection ability and efficiency have been improved, the risks of diseases such as chronic gastritis, peptic ulcers, MALT lymphoma and gastric cancer have been reduced, and a non-invasive and easy-to-promote detection method has been provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120254268A_ABST
    Figure CN120254268A_ABST
Patent Text Reader

Abstract

The invention discloses a urine protein marker for diagnosing helicobacter pylori as well as application, a kit and a risk assessment model of the urine protein marker, and belongs to the technical field of biomedicine. And the protein marker is a combination of PHB1, C20orf202, CP and P0DOX5. Compared with a healthy control, the expression level of PHB1 in a helicobacter pylori patient is lowered, and the expression levels of C20orf202, CP and P0DOX5 in the helicobacter pylori patient are all increased. By detecting expression levels of PHB1, C20orf202, CP and P0DOX5, diagnosis of helicobacter pylori patients can be realized, detection capability and efficiency can be improved, intervention measures can be taken actively, and risks of diseases such as chronic gastritis, peptic ulcer, MALT lymphoma and gastric cancer can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biomedical technology, and particularly relates to a urine protein biomarker for diagnosing Helicobacter pylori, its application, a kit, and a risk assessment model. Background Art

[0002] Helicobacter pylori ( Helicobacter pylori,H.pylori ) is a Gram-negative spiral bacterium. The "Primary Care Diagnosis and Treatment Guidelines for Helicobacter pylori Infection" in 2019 determined that any one of the following 3 items can be judged as H.pylori current infection: ① Any one of the following three items in gastric mucosa tissue: positive RUT, tissue section staining, or bacterial culture; ② Positive for 13C or 14C-UBT; ③ Positive for HpSA detection (monoclonal antibody method verified clinically).

[0003] The global natural population Helicobacter pylori infection rate has exceeded 50%, and the infection rate in China is approximately 40% - 90%. Helicobacter pylori infection can almost always cause active gastric mucosal inflammation, and on the basis of chronic inflammatory activity, some patients may also develop a series of diseases such as peptic ulcer and gastric cancer. The above invasive and non-invasive diagnostic methods are actually relatively expensive and overly cumbersome in operation, and are not convenient for national-wide general surveys. Therefore, there is an urgent need to find a low-cost method suitable for national-wide general surveys.

[0004] Proteomics, as an emerging diagnostic tool, provides new possibilities for the diagnosis of Helicobacter pylori. Proteomics refers to the study of all proteins expressed in the genome and their characteristics, mainly including protein structure, protein abundance, protein-protein interactions, etc. Machine learning can be applied in various proteomics studies such as processing proteomic mass spectrometry data, screening protein biomarkers, and predicting protein-protein interactions. Machine learning algorithms can accurately identify biomarkers related to specific diseases or their subtypes by deeply analyzing proteomic data, and help scientists determine which proteins or protein combinations are highly relevant in the diagnosis and prediction of specific diseases by establishing models.

[0005] Helicobacter pylori has become a serious public health problem in China, and national-wide general survey is an important measure to prevent the occurrence of chronic infectious diseases. The combination of proteomics and machine learning helps researchers find more accurate and efficient protein biomarkers, establish diagnostic models, and provide new ideas and methods for the accurate diagnosis of Helicobacter pylori. Summary of the Invention

[0006] Based on the role of proteins in Helicobacter pylori patients and healthy individuals, the present invention studies protein markers related to the diagnosis of Helicobacter pylori, thereby providing a more universal detection method for the diagnosis of Helicobacter pylori.

[0007] On the one hand, the present invention provides a urinary protein marker for diagnosing Helicobacter pylori, and the protein marker is a combination of PHB1, C20orf202, CP, and P0DOX5. Among them, compared with healthy controls, the expression level of PHB1 is down-regulated in Helicobacter pylori patients, and the expression levels of C20orf202, CP, and P0DOX5 are all up-regulated in Helicobacter pylori patients. Among them, the source of the protein is urine. Urine is a non-invasive diagnostic tool that can provide useful information and is easy to carry out in outpatient clinics and hospitals.

[0008] On the other hand, the present invention discloses the application of the aforementioned protein marker in an early diagnosis kit for Helicobacter pylori.

[0009] Among them, the early diagnosis kit for detecting Helicobacter pylori includes a detection reagent capable of detecting the expression levels of PHB1, C20orf202, CP, and P0DOX5. Among them, the detection reagent includes a reagent for detecting the protein level by methods such as immunoblotting, enzyme-linked immunosorbent assay, mass spectrometry, radioimmunoassay, radioimmunodiffusion, immunoelectrophoresis, tissue immunostaining, immunoprecipitation analysis, complement fixation analysis, fluorescence-activated cell sorting, quality analysis, or protein microarray.

[0010] Specifically, the detection reagent is a binder that specifically binds to the proteins encoded by PHB1, C20orf202, CP, and P0DOX5. More specifically, the binder for the protein includes peptides, peptidomimetics, aptamers, spiegelmers, darpins, ankyrin repeat proteins, Kunitz-type domains, antibodies, single-domain antibodies, and / or monovalent antibody fragments, etc.

[0011] On yet another hand, the present invention discloses a risk assessment model for Helicobacter pylori. The risk assessment model uses the expression levels of the aforementioned protein markers (a combination of PHB1, C20orf202, CP, and P0DOX5) as variables (corresponding to corresponding scores), and uses a nomogram to visualize the model results. The sum of the individual scores corresponding to different values of each protein marker is obtained to get the total score, and then it corresponds to the probability of Helicobacter pylori events.

[0012] Those skilled in the art can implement and realize the steps of associating the biomarker level with a certain probability or risk in different ways. Preferably, the measured concentrations of one or more protein biomarkers are combined mathematically, and the combined value is associated with the early diagnosis problem. The biomarker values can be measured and combined by any suitable existing mathematical methods.

[0013] The function for associating the biomarker combination with the disease preferably adopts algorithms developed and obtained by applying statistical methods. For example, suitable statistical methods are discriminant analysis (DA) (i.e., linear, quadratic, regular DA), Kernel methods (i.e., SVM), non-parametric methods (i.e., k-nearest neighbor classifier), PLS (partial least squares), tree-based methods (i.e., logistic regression, CART, random forest method, boosting / bagging method), generalized linear models (i.e., log regression), principal component-based methods (i.e., SIMCA), generalized additive models, fuzzy logic-based methods, neural network- and genetic algorithm-based methods, etc. Skilled technicians will have no problem in selecting a suitable statistical method to evaluate the biomarker combination of the present invention and thus obtaining a suitable mathematical algorithm.

[0014] Specifically, in the nomogram, the score range of P0DOX5 is 12.5 - 16.5, corresponding to 0 - 32.5 in the score scale; the score range of C20orf202 is 6 - 16, corresponding to 0 - 25 in the score scale; the score range of CP is 7.5 - 14.5, corresponding to 0 - 100 in the score scale; the score range of PHB1 is 15 - 6, corresponding to 0 - 80 in the score scale; the occurrence probabilities of Helicobacter pylori in the population take values of 0.01, 0.1, 0.3, 0.5, 0.7, and 0.9, corresponding to 108, 128, 140, 148, 156, and 165 in the total score scale respectively.

[0015] By detecting the expression levels of PHB1, C20orf202, CP, and P0DOX5, the present invention can achieve the diagnosis of Helicobacter pylori, improve the detection ability and efficiency, actively take intervention measures, and reduce the risks of diseases such as chronic gastritis, peptic ulcer, MALT lymphoma, and gastric cancer. Description of the Drawings

[0016] Figure 1 is the flow chart for model construction; Figure 2 is the quantity distribution diagram of quantitative proteins in the Helicobacter pylori group and the healthy control group; Figure 3 is the quantity distribution diagram of quantitative peptide segments in the Helicobacter pylori group and the healthy control group; Figure 4 are the differentially expressed proteins screened out in the Helicobacter pylori group and the healthy control group; Figure 5 is the principal component analysis diagram; Figure 6 is the quantitative heat map of differentially expressed proteins; Figure 7 is the relationship diagram between fitting error and log(λ); Figure 8 is the LASSO coefficient distribution diagram of differentially expressed proteins; Figure 9 is the box plot of the protein expression level of the protein marker PHB1; Figure 10 is the box plot of the protein expression level of the protein marker C20orf202; Figure 11 is the box plot of the protein expression level of the protein marker CP; Figure 12 is the box plot of the protein expression level of the protein marker P0DOX5; Figure 13 is the receiver operating characteristic curve (ROC) diagram of the early diagnosis model; Figure 14 is the confusion matrix diagram of the early diagnosis model; Figure 15 is the calibration curve diagram of the early diagnosis model; Figure 16 is the decision curve (DCA) diagram of the early diagnosis model; Figure 17 is the nomogram for predicting the risk of Helicobacter pylori. Detailed implementation manners

[0017] In the present invention, PHB1 (UniprotID: P35232), C20orf202 (UniprotID: A1L168), CP (UniprotID: P00450); P0DOX5 (UniprotID: P0DOX5) include the gene, its encoded protein, and its homologs, mutations, and isoforms. This term covers full-length, unprocessed protein markers, as well as any form of protein marker derived from processing in cells. This term covers naturally occurring variants of the protein marker (such as splice variants or allelic variants). UniprotID can be obtained at https: / / www.uniprot.org / .

[0018] The term "expression level" generally refers to the amount of polynucleotide or amino acid product or protein in a biological sample. "Expression" generally refers to the process by which the information encoded by a gene is converted into a structure that exists and functions in a cell. Thus, as used herein, "expression" refers to transcription into polynucleotide, translation into protein, or post-translational modification of the protein. Fragments of the transcribed polynucleotide, translated protein, or post-translationally modified protein are also considered products of expression, whether they are derived from transcripts generated by alternative splicing or degraded transcripts, or from post-translational modification of the protein.

[0019] The term "early diagnosis" refers to making a diagnosis when a disease is asymptomatic or in the very early stages of having symptoms. For example, methods for early diagnosis of Helicobacter pylori can include measuring the expression of certain proteins in a biological sample from an individual.

[0020] The detection of the protein expression level described herein can be carried out using assay methods known in the art, including but not limited to methods for detecting the amount of a polypeptide encoded by a protein marker.

[0021] In this article, the amount of polypeptide can be detected by, for example, proteomics or reagents. Preferably, the sequencing technology is proteomic sequencing, more preferably targeted proteomic sequencing technology. Any suitable protein quantification method can be used in the methods provided herein. Exemplary methods that can be used include but are not limited to immunoblotting (western blot), enzyme-linked immunosorbent assay (ELISA), and mass spectrometry, etc.

[0022] A "kit" is a manufactured article (e.g., a package or container) that contains a probe for specifically detecting the protein marker of the present invention. Such a kit can include carrier means that are compartmentalized to tightly and restrictively accommodate one or more container means (such as vials, tubes, etc.), each container means containing one of the separate components to be used in the method. For example, one of the container segments can contain a probe that can carry out a detectable label or can be made to carry out a detectable label.

[0023] The present invention will be further described in detail below in conjunction with the drawings and examples. The following examples are only used to illustrate the present invention and are not used to limit the scope of the present invention. For experimental methods not specified with specific conditions in the examples, they are generally carried out under conventional conditions or according to the conditions recommended by the manufacturer.

[0024] 1. Sample collection and storage Urine samples were collected from 15 patients with Helicobacter pylori (Hp group) and 15 healthy control subjects in the general population (Un-Hp group). When collecting the samples, the basic information of the patients and the healthy control group, as well as relevant laboratory test indicators, including age, gender, height, weight, BMI, triglycerides, high-density lipoprotein cholesterol, fasting blood glucose, and medication history, etc., were recorded in detail. Among them, there were no statistically significant differences in age, gender, and height between the disease group and the healthy control group.

[0025] Inclusion criteria for the disease group: 1. Helicobacter pylori-infected patients who underwent gastroscopy and gastric mucosa biopsy due to a positive urea breath test during a health check; 2. The patient himself or his legal representative must sign a complete informed consent form before participating in the study.

[0026] Exclusion criteria for the disease group: Patients who had a positive urea breath test during a health check but did not undergo gastroscopy or gastric mucosa biopsy.

[0027] Inclusion criteria for the healthy group: 1. Adult patients 18 years old and above; 2. Based on the physical examination results, relevant indicators were obtained, and the urea breath test during the health check was negative; 3. The patient himself or his legal representative must sign a complete informed consent form before participating in the study.

[0028] Exclusion criteria for the healthy group: 1. History of serious diseases, such as cancer, autoimmune diseases, or other serious diseases; 2. Renal function impairment, positive urine protein, or albumin / creatinine ratio < 30 mg / g; 3. Pregnant or lactating women.

[0029] Take 10 mL of midstream fasting urine, centrifuge at 2000 g for 10 min at 4 °C; aspirate 2 mL of the supernatant and aliquot it into pre-numbered 2 mL sterile centrifuge tubes, and quickly freeze the aliquoted urine samples in a -20 °C refrigerator for storage.

[0030] 2. Experimental methods This patent uses the next-generation label-free quantitative proteomics technology to complete the analysis. In the data independent acquisition (DIA) mode, the latest high-resolution mass spectrometry is used to simultaneously collect the peptide ion characteristics in terms of mass number and retention time. Compared with the traditional method of extracting a single ion for fragmentation analysis, in the DIA mode, the mass spectrometry is set to cyclically collect in a wide parent ion window and simultaneously fragment multiple peptide ions. It realizes the complete collection of all detectable protein spectral peak information in the sample, enabling the highly reproducible analysis of a large number of samples. The DIA process provides an ideal platform for qualitative analysis of differentially expressed proteomes or quantitative analysis of proteomes of a large number of samples.

[0031] The detection and analysis process steps are as follows: 1) Sample preparation Add lysis buffer to the sample to make its final concentration 1% SDC / 100 mM Tris-HCl (pH = 8.5). After thorough mixing, determine the protein concentration using the BCA method. Take equal amounts of protein and use 1% SDC / 100 mM Tris-HCl (pH = 8.5) solution to make up all samples to the same volume. After adding TCEP and CAA, incubate at 60 °C for 30 min to complete reduction and alkylation. Add an equal volume of ddH2O to dilute the SDC concentration to less than 0.5%. Add trypsin according to the mass ratio of enzyme to protein of 1:50, and incubate and shake overnight at 37 °C for digestion. The next day, add TFA to terminate the digestion reaction, centrifuge the sample at 12,000 g, take the supernatant and desalt it with a self-made SDB desalting column, vacuum dry it and store it at -20 °C.

[0032] Reagent description: Tris-HCl: Tris(hydroxymethyl)aminomethane hydrochloride; TCEP: Tris(2-carboxyethyl)phosphine hydrochloride; CAA: Calcium acetylacetonate; ddH2O: Double deionized water; SDC: Sodium deoxycholate; TFA: Trifluoroacetic acid; SDB: Sulfonated styrene-vinylbenzene copolymer.

[0033] 2) Mass spectrometry detection The mass spectrometry detection of the sample was performed using an UltiMate 3000 RSLCnano nanoliter liquid phase (Thermo) tandem timsTOFPro mass spectrometer (Bruker). The peptide sample was injected through an autosampler, bound to a C18 trapping column (75 µm * 2 cm, 3 µm particle size, 100 Å pore size, Thermo), and then separated in an analytical column (75 µm * 15 cm, 1.7 µm particle size, 100 Å pore size, IonOpticks). An analytical gradient mass spectrometry was established using mobile phase A (0.1% formic acid) and mobile phase B (0.1% formic acid in ACN), and data was collected in diaPASEF mode with a flow rate set at 300 nL / min. The capillary voltage was set at 1500 V. The scanning ranges of MS1 and MS2 spectra were set at 100 - 1700 m / z. The ion mobility range was set at 0.6 - 1.6 Vs / cm 2 . The Accumulation time and ramp time were set at 50 ms. According to the distribution law of mass-to-charge ratio - ion mobility, the diaPASEF acquisition window was set using timsControl software. The collision energy was set to linearly decrease from 59 eV at 1 / K0 = 1.6 Vs / cm 2 to 20 eV at 1 / K0 = 0.6 Vs / cm 2 .

[0034] 3) Data analysis ① Database search The DIA raw data was analyzed using the library-free mode of DIA-NN (1.8.1) software. The database used for the search was the HUMAN protein sequence database downloaded from Uniprot. The search parameters mainly adopted the default settings, and the key parameters are described as follows: The Precursor ion generation option was enabled for the prediction of the theoretical spectral library; Trypsin / P was used with a maximum of 1 missed cleavage site allowed; Carbamidomethyl (C) modification was set as a fixed modification; Oxidaton (M) and Acetylation (protein N-terminal) were set as variable modifications; The MS1 and MS2 mass tolerances were set at 15 ppm; MBR was enabled; Heuristic protein inference was enabled; The FDR was set at 1%. The protein quantification information was normalized using the MaxLFQ algorithm.

[0035] ② Data cleaning, filtering, transformation, and filling Remove contaminating proteins from the quantitative protein matrix, filter out proteins with a missing ratio greater than 80% in the sample or a missing ratio exceeding 50% in both groups, and perform logarithmic transformation. Then, fill in the missing values with values representing the normal distribution near the detection limit of the mass spectrometer. To this end, we determined the mean and standard deviation of the actual intensity distribution, and then created a new distribution that was shifted down by 1.8 standard deviations and had a width of 0.25 standard deviations. These values were used to estimate the missing values in the protein matrix, and the remaining proteins were used for subsequent analysis.

[0036] ③ Differential expression analysis Compare the protein quantification data of the Helicobacter pylori patient group and the healthy control group, and screen for differentially expressed proteins in the comparison between the two groups using the fold change (FC) and differential test (two-sided P < 0.05) as criteria. The screening criteria were |log2FC| ≥ 0.58 and p < 0.05. The results are as Figure 4 shown.

[0037] ④ Key protein screening Randomly divide the total sample into a training set and a test set at a ratio of 3:1. Use the Lasso regression model for five-fold cross-validation in the training set to screen out N differential protein features; then, perform Logistic regression analysis on these N proteins in a random combination manner to construct a diagnostic model, and screen for the optimal diagnostic model by calculating the discrimination index AUC and the calibration index Brier score through three-fold cross-validation to select the optimal combination of protein markers. The differentially expressed proteins in the optimal combination are the disease-related protein markers.

[0038] ⑤ Model construction and evaluation Apply Logistic regression and cross-validation methods in the training set to randomly select protein combinations from the protein markers to construct multiple prediction models, and select the best according to the area under the receiver operating characteristic curve (AUC) and the Brier score. Subsequently, evaluate the selected optimal prediction model. Evaluate the discrimination, calibration, and clinical utility performance of the model in the training set and the test set through the receiver operating characteristic (ROC) curve, calibration curve, and decision curve in the test set and the training set, respectively. Generally, it is considered that 0.90 ≤ AUC < 1.00 indicates excellent model discrimination ability; 0.75 ≤ AUC < 0.90 indicates good model discrimination ability; 0.60 ≤ AUC < 0.75 indicates certain model discrimination ability, but it is not recommended; AUC < 0.60 indicates poor model discrimination ability. All statistical analyses were completed using R (version 4.3.2) and Python (version 3.10.12).

[0039] 3. Results and analysis 1) From the physical examination population, 15 pairs of Helicobacter pylori patients and healthy control groups with matched information were selected for proteomic research, and an early diagnosis model was established. Figure 1 The flow chart for constructing the early diagnosis model of Helicobacter pylori was described.

[0040] 2) The demographic and clinical characteristics of the study population were summarized, and the results are shown in Table 1.

[0041] Table 1 Demographic and clinical characteristics of the study population

[0042] (1) t: t-test, Z: Mann-Whitney test; (2) SD: standard deviation; M: median; Q1: first quartile; Q3: third quartile.

[0043] 3) Mass spectrometry quality control information (1) Using data-independent acquisition (DIA) technology, urine proteomic profiles of 15 Helicobacter pylori patients and 15 healthy control groups were obtained. The results are as Figure 2 and Figure 3 shown.

[0044] (2) The expression of 100 proteins in Helicobacter pylori patients was significantly different from that in the healthy control group. Compared with the healthy control group, the expression of 26 proteins was up-regulated and 74 proteins were down-regulated in the Helicobacter pylori group. The results are as Figure 4 shown.

[0045] (3) The results of principal component analysis (PCA) showed the intensities of 100 differentially expressed proteins, showing a significant difference between Helicobacter pylori patients and healthy control groups. The results are as Figure 5 shown.

[0046] (4) A heat map was made using the expression levels of 100 differentially expressed proteins, showing a significant difference in the expression levels of differentially expressed proteins between Helicobacter pylori patients and healthy control groups. The results are as Figure 6 shown.

[0047] 4) Protein screening The total samples were randomly divided into a training set and a test set at a ratio of 3:1. In the training set, a Lasso regression model was used for five-fold cross-validation to screen out 14 important proteins from 100 differentially expressed proteins. Logistic regression analysis was performed on these 14 important proteins by random combination to construct a diagnostic model, and the optimal diagnostic model was screened by the discrimination index AUC and the calibration index Brier score calculated by three-fold cross-validation. The screening results are as Figure 7 and Figure 8As shown. Finally, PHB1, C20orf202, CP, and P0DOX5 were selected as the best protein biomarker combination, and a clinical risk model was established. The results of the model are shown in Table 2.

[0048] Table 2 Logistic regression model

[0049] "Estimate" is the regression coefficient, "Std.Error" is the standard error of the regression coefficient, "z value" is the test statistic for the Z-test of the regression coefficient, and "Pr(>|z)" is the P-value for the Z-test of the regression coefficient.

[0050] The expression levels of the four protein biomarkers PHB1, C20orf202, CP, and P0DOX5 obtained by screening were plotted as box plots in two groups of people to visually transform the expression differences of these four protein biomarkers in the urine of Helicobacter pylori patients and healthy controls. The box plots showed that compared with the healthy control group, the expression level of PHB1 was down-regulated in Helicobacter pylori patients, and the expression levels of C20orf202, CP, and P0DOX5 were all up-regulated in Helicobacter pylori. The results are as Figures 9 - 12 shown.

[0051] 4. Model evaluation When evaluating the early diagnosis model composed of the protein biomarkers PHB1, C20orf202, CP, and P0DOX5, 30 samples were divided into two sets according to the ratio. Among them, 75% of the data was used as the training set (n = 22), and 25% of the data was used as the test set (n = 8). The training set and the test set were used for model evaluation.

[0052] 1) Receiver operating characteristic curve (ROC) graph The results are as Figure 13 shown. The ROC curve is used to evaluate the discrimination of the early diagnosis model. Generally, the area under the ROC curve (AUC) value is used to evaluate the model prediction performance. Evaluation criteria: AUC ≥ 0.90, the model has excellent discrimination ability; 0.90 > AUC ≥ 0.75, the model has good discrimination ability; 0.75 > AUC ≥ 0.60, the model has a certain discrimination ability, but it is not recommended to use; AUC < 0.60, the model has poor discrimination ability. The ROC curve graph showed that the AUC (95%CI) of the training set was 0.942 (0.847 - 1), and the AUC (95%CI) of the test set was 0.938 (0.764 - 1). The ROC curve showed that the early diagnosis model of Helicobacter pylori had good discrimination ability.

[0053] 2) Confusion matrix The results are as Figure 14As shown, the confusion matrix diagram shows the classification of the early diagnosis model for the test set. The horizontal axis represents the actual category, and the vertical axis represents the predicted category. Among them, the people in the upper left quadrant and the lower right quadrant are classified into the correct categories. The results of the confusion matrix diagram show that the accuracy rate of the model in the training set is 86%, and the accuracy rate of the model in the test set is 88%, indicating that the prediction ability of this early diagnosis model is good.

[0054] 3) Calibration curve The results are as Figure 15 shown. The calibration curve is used to evaluate the calibration degree of the early diagnosis model, that is, the consistency between the probability of the outcome prediction and the actual observed frequency. The dotted line is the ideal curve, the red line is the actual curve of the model, and the blue line is the curve after resampling calibration. The closer the calibration curve is to the ideal curve, the better the model effect. The results of the calibration curve show that the consistency between the predicted probability and the observed probability of the early diagnosis model of Helicobacter pylori is good.

[0055] 4) Decision curve The results are as Figure 16 shown. The decision curve is used to evaluate the clinical utility of the model. The net return rate reflects the clinical value of the model. The greater the net return rate, the better the model effect. In the decision curve diagram, the black horizontal line represents that no one is treated, and the net benefit is 0; the gray diagonal line represents that everyone is treated, and the net benefit is an inverse diagonal line with a negative slope; the red curve is the net benefit curve of the current early diagnosis model. The higher the net return rate of the model, the greater its clinical utility. The decision curve diagram shows that within a certain range, the net return rate of the early diagnosis model is relatively high and higher than the two schemes of all treatments and no treatments at all, proving that the clinical utility of the early diagnosis model of Helicobacter pylori is relatively large.

[0056] 5. Model visualization The results are as Figure 17As shown, a nomogram is used to visually display the model results for facilitating the prediction of the disease risk. The individual scores (Points) corresponding to each protein marker at different values are summed to obtain the total score (Total Points), and then mapped to the Risk score, which is the probability of Helicobacter pylori infection. Specifically, in the nomogram, the score range of P0DOX5 is 12.5 - 16.5, corresponding to 0 - 32.5 on the score scale; the score range of C20orf202 is 6 - 16, corresponding to 0 - 25 on the score scale; the score range of CP is 7.5 - 14.5, corresponding to 0 - 100 on the score scale; the score range of PHB1 is 15 - 6, corresponding to 0 - 80 on the score scale; the occurrence probabilities of Helicobacter pylori population taking values 0.01, 0.1, 0.3, 0.5, 0.7, and 0.9 respectively correspond to 108, 128, 140, 148, 156, and 165 on the total score scale. It can be seen from the figure that when the three markers are combined for risk prediction, the upper limit of the predicted probability can reach 0.9.

[0057] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A urinary protein biomarker for diagnosing Helicobacter pylori, characterized in that, The protein markers are a combination of PHB1, C20orf202, CP, and P0DOX5, and the source of the proteins is urine.

2. The protein marker according to claim 1, wherein Compared with healthy controls, the expression level of PHB1 is down-regulated in Helicobacter pylori patients, while the expression levels of C20orf202, CP, and P0DOX5 are up-regulated in Helicobacter pylori patients.

3. Use of the protein markers according to claim 1 in a diagnostic kit for metabolizing Helicobacter pylori.

4. An early diagnosis kit for detecting Helicobacter pylori, characterized in that, It includes detection reagents capable of detecting the expression levels of PHB1, C20orf202, CP, and P0DOX5.

5. An evaluation model for Helicobacter pylori risk, characterized in that, The risk assessment model uses the expression levels of the protein markers according to claim 1 as variables, and uses a nomogram to visualize the model results. The sum of the individual scores corresponding to different values of each protein marker is obtained to get the total score, and then it corresponds to the probability of healthy people developing Helicobacter pylori events.

6. The evaluation model according to claim 5, wherein In the nomogram, the score range of P0DOX5 is 12.5 - 16.5, corresponding to 0 - 32.5 on the score scale; the score range of C20orf202 is 6 - 16, corresponding to 0 - 25 on the score scale; the score range of CP is 7.5 - 14.5, corresponding to 0 - 100 on the score scale; the score range of PHB1 is 15 - 6, corresponding to 0 - 80 on the score scale; the probability of occurrence of Helicobacter pylori in the population takes values of 0.01, 0.1, 0.3, 0.5, 0.7, and 0.9, corresponding to 108, 128, 140, 148, 156, and 165 on the total score scale respectively.