Biomarker for detecting pancreatic cancer and application thereof

By screening and combining new biomarkers, using LC-MS/MS technology to detect pancreatic cancer markers, and constructing a pancreatic cancer detection system, the problem of insufficient accuracy in early diagnosis of pancreatic cancer in existing technologies has been solved, and efficient early prediction and staging diagnosis of pancreatic cancer has been achieved.

CN120761644AActive Publication Date: 2025-10-10HANGZHOU GUANGKE ANDE BIOTECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510970188.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-06-03
Filing Date
2025-07-15
Publication Date
2025-10-10
Estimated Expiration
2045-07-15

AI Technical Summary

Technical Problem

Existing methods for early diagnosis of pancreatic cancer lack accuracy, and common protein markers such as CA19-9 have poor specificity, resulting in most patients being diagnosed in the late stage and losing the opportunity for surgical resection.

Method used

A new group of biomarkers was screened out, including SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, ORM1, etc. The expression levels of these markers in body fluid samples were detected by LC-MS/MS technology, a pancreatic cancer detection system was constructed, and predictions were made using the data analysis module.

Benefits of technology

It has achieved non-invasive, convenient and efficient early prediction of pancreatic cancer, improved the accuracy and sensitivity of diagnosis, can distinguish the benign and malignant states of tumors, and effectively distinguish the clinical stages of pancreatic cancer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120761644A_ABST
    Figure CN120761644A_ABST
Patent Text Reader

Abstract

The invention discloses application of a biomarker in preparation of a pancreatic cancer detection product, the biomarker is selected from one or more of SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3 and ORM1, and the invention discloses a system and a kit for detecting pancreatic cancer, and the system and the kit comprise the biomarker. The biomarkers capable of predicting the tumor occurrence risk in the early stage of pancreatic cancer are screened out, a pancreatic cancer detection system is constructed based on the markers, the benign and malignant states of pancreatic tumors of individuals can be noninvasively, conveniently and efficiently predicted, the accuracy, sensitivity and specificity are high, the I stage, II stage, III stage and IV stage of pancreatic cancer can be effectively distinguished, and the pancreatic cancer detection method has good application prospects. The requirements of clinical detection and monitoring of the progress condition of the ovarian cancer are met, and the application prospect is relatively great.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of medical detection, and particularly relates to screening biomarkers of pancreatic cancer by proteomics for detecting pancreatic cancer and distinguishing between benign and malignant pancreatic tumors. BACKGROUND

[0002] Proteomics is a science that studies the composition, localization, changes and interaction rules of proteins in cells, tissues or organisms, including the study of protein expression patterns and protein function patterns. With the development of mass spectrometry technology, liquid chromatography-mass spectrometry (LC-MS / MS) has become a core tool for proteomics research, providing important support for the discovery of disease diagnostic markers, the screening of drug targets and toxicology research.

[0003] Pancreatic cancer is a highly malignant tumor that is difficult to detect in the early stage and has a poor prognosis, with a morbidity rate close to mortality rate. Currently, clinical diagnosis mainly relies on physical methods such as ultrasonic examination, magnetic resonance imaging and endoscopic ultrasonography. Although these methods have a certain diagnostic efficiency, more than 80% of pancreatic cancer patients are in the advanced stage at the time of diagnosis, losing the opportunity for surgical resection. Therefore, improving the accuracy of early diagnosis of pancreatic cancer and reducing the difficulty of diagnosis is the key to improving the survival rate of patients.

[0004] Currently, the protein marker CA1909 is the most common and widely used tumor marker for the diagnosis and prognosis monitoring of pancreatic cancer in clinical practice. However, CA19-9 as a biomarker still has some limitations, such as poor specificity, low expression in Lewis-negative phenotype, and high false positive rate in patients with benign diseases such as pancreatitis, cirrhosis and acute cholangitis. Other common protein markers such as CEA, TP53 and other single protein markers also have certain deficiencies in sensitivity and specificity.

[0005] Based on the above background, it is necessary to find new biomarkers of pancreatic cancer and their combinations for detecting pancreatic cancer and distinguishing between benign and malignant tumors, and to construct a prediction model for benign and malignant pancreatic cancer, which has important clinical value. SUMMARY

[0006] In view of the problems in the prior art, the present application provides a biomarker for detecting pancreatic cancer, a detection reagent, a detection kit and a pancreatic cancer risk prediction system. The present application screens a series of completely new biomarkers that can early predict the risk of pancreatic cancer, and can distinguish between benign and malignant pancreatic tumors.

[0007] The technical scheme adopted by the present application is: Use of a biomarker in preparing a pancreatic cancer detection product, wherein the biomarker is selected from one or more of the following: SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, ORM1; Preferably, the biomarker is one of the following combinations: (A) Combination of SERPINB1, CD74, KRT19, S100B, ORM1, and DEFA3; (ii) combination of KRT19, PGM5, SELL, DEFA3, S100B, CD74, and SERPINB1; (iii) combination of CD74, DEFA3, RNASE1, ORM1, SPINK5, KRT19, SERPINB1, and PGM5; (IV) combination of SERPINB1, ORM1, RNASE1, DEFA3, PGM5, SPINK5, KRT19, S100B, and CD74; (V) combination of S100B, SPINK5, SERPINB1, KRT19, ORM1, CD74, PGM5, SELL, DEFA3, and RNASE1; More preferably, it is a combination of SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, and ORM1.

[0008] The pancreatic cancer detection product can be a pancreatic cancer detection reagent or a pancreatic cancer detection kit.

[0009] The present invention also provides a kit for detecting pancreatic cancer, the kit comprising a reagent for detecting a biomarker, wherein the biomarker is selected from one or more of the following: SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, ORM1; Preferably, the biomarker is one of the following combinations: (A) Combination of SERPINB1, CD74, KRT19, S100B, ORM1, and DEFA3; (ii) combination of KRT19, PGM5, SELL, DEFA3, S100B, CD74, and SERPINB1; (iii) combination of CD74, DEFA3, RNASE1, ORM1, SPINK5, KRT19, SERPINB1, and PGM5; (IV) combination of SERPINB1, ORM1, RNASE1, DEFA3, PGM5, SPINK5, KRT19, S100B, and CD74; (V) combination of S100B, SPINK5, SERPINB1, KRT19, ORM1, CD74, PGM5, SELL, DEFA3, and RNASE1; More preferably, it is a combination of SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, and ORM1.

[0010] Furthermore, the reagents for detecting biomarkers can detect the expression levels of biomarkers. The reagents for detecting biomarkers can be sample pretreatment reagents, antigens or antibodies, and other biological reagents and kits suitable for biomarker detection; they can also be developed into standardized reagents or kits suitable for LC-UV or LC-MS detection of biomarkers.

[0011] In some embodiments, the reagent for detecting a biomarker is an antibody against the biomarker as described above, and the antibody is a monoclonal antibody.

[0012] The SELL is L-selectin, which is the protein or amino acid sequence with the UniProt database number P14151; The CD74 is an HLA-II class histocompatibility antigen γ chain, and is a protein or amino acid sequence with a UniProt database number of P04233; The SERPINB1 is a leukocyte elastase inhibitor, which is a protein or amino acid sequence with the UniProt database number P30740; The RNASE1 is a pancreatic ribonuclease, which is a protein or amino acid sequence with the UniProt database number P07998; The KRT19 is a type I cytoskeletal keratin 19, which is a protein or amino acid sequence with the UniProt database number P08727; The S100B is protein S100-B, which is the protein or amino acid sequence with the UniProt database number P04271; The PGM5 is a phosphoglucomutase 5-like protein or amino acid sequence with the UniProt database number Q15124; The SPINK5 is a Kazal-type serine protease inhibitor 5, which is a protein or amino acid sequence numbered Q9NQ38 in the UniProt database; DEFA3 is neutrophil defensin 3, which is the protein or amino acid sequence with the UniProt database number P59666; The ORM1 is α-1-acid glycoprotein 1, which is a protein or amino acid sequence with a UniProt database number of P02763. Furthermore, the reagent is used to detect biomarkers in a body fluid sample, and the body fluid sample includes any one of blood, urine, saliva, and sweat.

[0013] In some preferred embodiments, the biomarkers of the present invention are obtained by screening blood samples, and are particularly suitable for development into blood detection reagents or kits for predicting pancreatic cancer.

[0014] Furthermore, detecting a biomarker refers to detecting the presence or relative abundance or concentration of a biomarker in a body fluid sample of an individual.

[0015] In some approaches, relative abundance is preferably expressed as the peak area of ​​the biomarker in the detection spectrum obtained by high-performance liquid chromatography-tandem mass spectrometry. For example, if the average peak area of ​​a biomarker measured in a control sample (an individual without pancreatic cancer) is 500 and the average peak area measured in a pancreatic cancer sample is 3000, then the abundance of the biomarker in the pancreatic cancer sample is considered to be 6 times that in the control sample.

[0016] The present invention also provides a system for detecting pancreatic cancer, the system comprising a data analysis module, the data analysis module being configured to analyze the detection values ​​of biomarkers in a sample of a patient to be tested, the biomarkers being selected from one or more of the following: SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, ORM1; Preferably, the biomarker is one of the following combinations: (A) Combination of SERPINB1, CD74, KRT19, S100B, ORM1, and DEFA3; (ii) combination of KRT19, PGM5, SELL, DEFA3, S100B, CD74, and SERPINB1; (iii) combination of CD74, DEFA3, RNASE1, ORM1, SPINK5, KRT19, SERPINB1, and PGM5; (IV) combination of SERPINB1, ORM1, RNASE1, DEFA3, PGM5, SPINK5, KRT19, S100B, and CD74; (V) combination of S100B, SPINK5, SERPINB1, KRT19, ORM1, CD74, PGM5, SELL, DEFA3, and RNASE1; More preferably, it is a combination of SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, and ORM1.

[0017] The data analysis module includes an analysis model equation, which is as follows:

[0018] logit(Y)=log(Y / 1-Y). The formula for calculating Y is as follows:

[0019] Where Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers (m = 10), Xi represents the detection value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker, and b is a constant of -5.78.

[0020] The coefficients of Ki are shown in Table 6 below: Table 6: Coefficients of the 10 biomarkers in the model

[0021] The data analysis module calculates a prediction value of whether the patient to be tested has pancreatic cancer using an analysis model equation according to the detection value of the biomarker, and determines whether the patient to be tested has pancreatic cancer based on the prediction value.

[0022] The determination conditions are: When the predicted value is less than or equal to a preset threshold, it is determined that the patient to be tested is not a pancreatic cancer patient; When the predicted value is greater than a preset threshold, the patient to be tested is determined to be a pancreatic cancer patient; The preset threshold is 0.473.

[0023] Furthermore, the system also includes a data detection module, a data input module, and a data output module; the data detection module is used to detect biomarkers in the sample and obtain detection values; the data input module is used to input the detection values ​​of the biomarkers, and after the data analysis module analyzes the detection values, the data output interface is used to output the analysis results of whether the patient to be tested has pancreatic cancer.

[0024] The detection value is generally obtained by performing an enzyme-linked immunosorbent assay (ELISA) on the sample to obtain the concentration of the biomarker in the sample as the detection value, and the unit is μg / mL.

[0025] The present invention also provides the use of biomarkers in preparing a pancreatic cancer clinical staging diagnostic product, wherein the biomarkers are a combination of SERPINB1, ORM1, RNASE1, DEFA3, PGM5, SPINK5, KRT19, S100B, and CD74, or a combination of SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, and ORM1.

[0026] The technical solution of the present invention has the following beneficial effects: (1) The present invention uses proteomics technology to systematically screen differential proteins in blood samples of pancreatic cancer patients and healthy controls, screen out biomarkers that can indicate the risk of tumor development in the early stages of pancreatic cancer, and construct a pancreatic cancer detection system based on these markers, thereby achieving non-invasive, convenient, and efficient prediction of the benign and malignant status of individual pancreatic tumors, meeting clinical detection needs, and having great application prospects.

[0027] (2) The pancreatic cancer detection model of the present invention has an AUC of 0.948508, a sensitivity of 0.922, and a specificity of 0.864 in the model group, and an AUC of 0.91854, an accuracy of 0.861, a sensitivity of 0.88, and a specificity of 0.842 in the test group. It has high accuracy and discrimination ability and can more efficiently predict whether an individual has pancreatic cancer.

[0028] (3) The biomarker combination of the present invention can be used to construct a pancreatic cancer clinical staging prediction model, which can effectively distinguish between pancreatic cancer stages I, II, III and IV, and has the potential to diagnose the clinical staging of pancreatic cancer. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 Volcano plot for differential analysis of benign and malignant pancreatic tumors.

[0030] Figure 2 The following are the ROC and OPLS-DA analysis results of benign and malignant pancreatic tumors.

[0031] Figure 3 AUC bar chart comparing the performance of models built for different marker combinations.

[0032] Figure 4 AUC results for models built for different hyperparameters.

[0033] Figure 5This is the ROC curve of the combined diagnosis model in the model group.

[0034] Figure 6 This is the ROC curve of the combined diagnosis model in the test group. DETAILED DESCRIPTION

[0035] The present invention will be described in further detail below in conjunction with the accompanying drawings and Examples. It should be noted that the following examples are intended to facilitate understanding of the present invention and do not serve to limit the present invention in any way. The reagents used in this example are all known products and were obtained by purchasing commercially available products.

[0036] It should be noted that: (1) Diagnosis or testing The diagnosis or detection here refers to the detection or analysis of biomarkers in a sample, or the content of a target biomarker, such as the absolute content or relative content, and then the presence or amount of the target marker is used to indicate whether the individual providing the sample may have or suffer from a certain disease, or the possibility of having a certain disease. The meanings of diagnosis and detection here are interchangeable. The result of such a test or diagnosis cannot be directly used as a direct result of the disease, but rather an intermediate result. If a direct result is obtained, other auxiliary means such as pathology or anatomy are required to confirm that the patient has a certain disease. For example, the present invention provides a variety of new biomarkers associated with pancreatic cancer, and changes in the content of these markers are directly correlated with whether the patient has pancreatic cancer.

[0037] (2) Association between markers or biomarkers and pancreatic cancer The terms "marker" and "biomarker" have the same meaning in this disclosure. The association here refers to a direct correlation between the presence or change in the level of a biomarker in a sample and a specific disease. For example, a relative increase or decrease in the level indicates a higher likelihood of the individual having the disease compared to healthy individuals.

[0038] The simultaneous presence of multiple markers in a sample, or the relative changes in their levels, indicate a higher likelihood of the individual having the disease compared to healthy individuals. This means that among marker types, some are strongly associated with disease, while others are weakly associated, or even unrelated to a specific disease. One or more markers with strong correlations can be used as diagnostic markers, while markers with weaker correlations can be combined with stronger markers to diagnose a disease, increasing the accuracy of test results.

[0039] Regarding the numerous biomarkers found in serum by the present invention, these markers can be used to distinguish pancreatic cancer from healthy people. The markers here can be used alone as single markers for direct detection or diagnosis. The selection of such a marker indicates that the relative change in the content of the marker has a strong correlation with pancreatic cancer. Of course, it is understandable that one or more markers with a strong correlation with pancreatic cancer can be selected for simultaneous detection. It is normal to understand that in some ways, selecting biomarkers with strong correlation for detection or diagnosis can achieve a certain standard of accuracy, such as 60%, 65%, 70%, 80%, 85%, 90% or 95% accuracy, which means that these markers can obtain intermediate values ​​for diagnosing a certain disease, but it does not mean that a certain disease can be directly confirmed.

[0040] Of course, differentially expressed proteins with larger ROC values ​​can also be selected as diagnostic markers. The so-called strength and weakness are generally determined through calculations using algorithms, such as the contribution ratio of the marker to pancreatic cancer or weight analysis. Such calculation methods can include significance analysis (p-value or FDR value) and fold change. Multivariate statistical analysis primarily includes principal component analysis (PCA), partial least squares discriminant analysis (PLS-DA), and orthogonal partial least squares discriminant analysis (OPLS-DA), as well as other methods such as ROC analysis. Of course, other model prediction methods are also possible. When selecting biomarkers, the differentially expressed proteins disclosed herein can be selected, or other existing, well-known marker combinations can be selected or combined for prediction using model methods.

[0041] Example 1 Pancreatic cancer biomarker screening 1. Experimental Design The experiment was designed to collect plasma samples from patients with pancreatic tumors for the first time, enrich low-abundance proteins based on the method of removing high-abundance proteins by immunoaffinity chromatography, detect the abundance of proteins in the samples by high-performance liquid chromatography-mass spectrometry tandem equipment, analyze the differences in their abundance between patients with benign pancreatic tumors and patients with malignant pancreatic tumors, and analyze its diagnostic performance.

[0042] 2. Sample Collection Blood samples from 100 patients with benign and malignant pancreatic tumors were collected. All pancreatic tumors were biopsied and pathologically confirmed. Before chemotherapy, radiotherapy, and surgery, approximately 2 ml of peripheral blood was collected from the subjects. The samples were mixed in a vacuum tube containing EDTA anticoagulant and centrifuged twice at 120 g for 10 minutes at room temperature. The supernatant was removed and centrifuged at 360 g for 20 minutes. Platelet samples were then collected in centrifuge tubes and stored at -80°C until further use.

[0043] 3. Protein Sample Processing and Enzymatic Hydrolysis First, plasma samples were centrifuged for 15 minutes (15,000 g). The supernatant was filtered and subjected to immunoaffinity chromatography to isolate 14 highly abundant proteins. Low-abundance proteins were then concentrated to 350 μL using a 3 kDa cutoff concentrator at 4000 g for 1 hour. The recovered concentrate was then subjected to buffer exchange (AEX-A) using a 7 kDa cutoff desalting column at 1000 g for 2 minutes. The replacement buffer was AEX-A (20 mM Tris, 4 M urea, 3% isopropanol, pH 8.0). Protein concentrations were determined using the BCA assay using AEX-A as a blank. According to sample grouping, 25 mL of TCEP was added to the samples and incubated at 37°C for 30 minutes for protein reduction. TMT labeling was then performed by adding the corresponding TMT 16-plex reagent and incubating at room temperature in the dark for 1 hour. The sample was then buffer exchanged using a Zeba column with AEX-A. The TMT 16-plex labeled samples were mixed, and 2 mL of AEX-A was added to the mixed sample, bringing the final volume to 5.5 mL. The sample was filtered through a 0.22 µm filter and separated using a 2D-HPLC system. The collected fractions were freeze-dried, and finally, the Trypsin-Lysin C enzyme cocktail was added. The sample was digested by incubation at 37°C for 5 hours, and the digestion reaction was terminated by the addition of 5 μL of 10% TFA. A total of 60 2D-HPLC fractions were used for nanoLC-MS / MS analysis.

[0044] 4. LC-MS / MS Data Acquisition Each sample was separated using an Easy nLC-1200 nanoflow liquid chromatography system and connected online to a Q Exactive HF-X high-resolution mass spectrometer. Mobile phase A consisted of 0.1% formic acid in water, and mobile phase B consisted of 0.1% formic acid in acetonitrile (80% acetonitrile, 20% water). The chromatographic columns consisted of an enrichment column and an analytical column equilibrated with 100% mobile phase A. Samples were loaded via an autosampler onto the enrichment column (100 μm inner diameter (ID), 4 cm length (L), C18 packing, 3 μm particles, 100 Å pore size) and separated on the analytical column (75 μm inner diameter, 25 cm length, C18 packing, 3 μm particles, 100 Å pore size) at a flow rate of 300 nL / min. Following chromatographic separation, samples were analyzed by mass spectrometry on a Q Exactive HF-X mass spectrometer. The detection mode was positive ion, the parent ion scanning range was 350-1800 m / z, the primary mass spectrometer resolution was 120,000 at 200 m / z, and the AGC (Automatic gain control) target was 3×10 6 The maximum injection time was 50 ms, and the dynamic exclusion time was 40 s. The mass-to-charge ratios of peptides and peptide fragments were acquired using data dependent acquisition (DDA): 20 MS / MS (MS2) spectra were acquired after each full scan (MS / MS). The MS2 activation type was high energy collision dissociation (HCD), the selection window was 0.7 m / z, the MS2 resolution was 30 000 at 200 m / z, and the automatic gain control target was set to 1 × 10 5 , the maximum injection time was 65 ms, the fixed first mass was 110.0 m / z, the normalized collision energy was 32 eV, and the minimum automatic gain control target was 2.00×10 4 , exclude 1-valent, 6-8-valent, and >8-valent ions, allow only single charge states, set peptide matching to priority, and turn on the isotope exclusion function.

[0045] 5. Data Preprocessing Secondary mass spectrometry data were retrieved using Maxquant (v1.6.15.0). The data type is DIA proteomics data based on secondary reporter ion quantification. The secondary spectrum used for quantification requires that the parent ion account for more than 75% in the primary spectrum. The database comes from the Homo_sapiens_9606_proteome_gene (release: 2021-10-14, sequence: 20,437) of the Uniprot database, and a common contamination library is added to the database. Contaminating proteins are deleted during data analysis; the enzyme digestion method is set to Trypsin / P; the number of missed cut sites is set to 2; the parent ion mass error tolerance of the First search and Main search is set to 20 ppm and 5 ppm, respectively, and the mass error tolerance of the secondary fragment ion is 20 ppm. The fixed modification is cysteine ​​alkylation, and the variable modification is methionine oxidation and protein N-terminal acetylation. The FDR for protein identification and PSM identification is set to 1%.

[0046] 6. Variance Analysis We screened for differentially expressed proteins and transcripts using a combination of univariate and multivariate statistical analyses. Univariate analysis primarily included significance analysis (p-value or FDR value) and fold change analysis of signature molecules across different groups. Multivariate statistical analysis primarily included receiver operating characteristic (ROC) curve analysis and Boruta signature screening based on the random forest algorithm. All statistical analyses were performed using R. Detailed R information is provided in Table 1.

[0047] Table 1: R used in the present invention and its related information

[0048] The variable importance for the projection (VIP) was calculated to measure the influence and explanatory power of each protein expression pattern on the classification and discrimination of each group of samples. The Wilcoxon rank sum test was further performed to obtain the corrected p value (FDR). According to the conditions of FDR < 0.01 and Fold change > 2, 67 down-regulated and 66 up-regulated proteins were screened (see Figure 1 ).

[0049] In order to evaluate the role of each marker in the diagnosis and prediction of pancreatic cancer, we used ROC and Boruta analysis methods to evaluate each marker. The results are shown in Figure Figure 2, the horizontal coordinate is AUC obtained by ROC analysis, the vertical coordinate is -log10(FDR) calculated by Wilcoxon test, and the size of the point represents the VIP value obtained by Boruta analysis. According to VIP>3 and AUC>0.6, a total of 10 more significant candidate markers were found, as shown in Table 2.

[0050] Table 2: Differential markers of benign and malignant pancreatic tumors

[0051] The smaller the FDR value and / or the larger the VIP value in Table 2, the more significant the difference between the two groups, and the higher the diagnostic value of the protein.

[0052] Example 2: Classification model of 10 differential proteins for identifying benign and malignant pancreatic tumors and establishment thereof Although a single biomarker can also distinguish between serum samples of benign and malignant pancreatic tumors or predict pancreatic cancer, generally, a combination of multiple biomarkers has higher accuracy in distinguishing or predicting. However, a single biomarker with higher accuracy in predicting pancreatic cancer may not necessarily play a greater role in the combination after being combined with other one or more biomarkers, and the number of biomarkers is not necessarily the more, the higher the prediction accuracy (AUC value) of the combination, so a large number of verification experiments are still needed.

[0053] In this embodiment, a model constructed from 10 protein markers SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, and ORM1 was studied.

[0054] 1. Data acquisition Study population: A total of 1000 samples were collected, including 500 blood samples from patients with benign pancreatic tumors and 500 blood samples from patients with malignant pancreatic tumors. All samples were derived from patients with confirmed results by pathology. The enrolled personnel were divided into a model group and a test group in a ratio of 8:2.

[0055] Inclusion criteria for pancreatic cancer patients: (a) no history of other malignant tumors, (b) surgical treatment within one month after blood collection, and confirmed as pancreatic cancer by postoperative pathology. After informed consent, all collected serum samples were stored in a serum bank at -80°C.

[0056] In this example, the collected serum samples were subjected to enzyme-linked immunosorbent assay (ELISA) to obtain the concentrations of SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, and ORM1 in the serum.

[0057] 2. Statistical Analysis of Experimental Data In the model group, a combined diagnostic model of multiple pancreatic cancer markers was constructed using a combination of multiple machine learning methods. The predicted probability values ​​were used to estimate the area under the receiver operator characteristic (ROC) curve (AUC) with a 95% confidence interval (CI) to evaluate the discriminatory ability of the multivariate diagnostic model. Using the test group, the Youden index (YI) was calculated to determine the cut-off value for distinguishing primary pancreatic cancer patients from metastatic predictive probability. In addition, ROCs for single markers and different subgroups were constructed and compared. Standard descriptive statistics such as frequency, mean, median, positive predictive value (PPV), negative predictive value (NPV) and standard deviation (SD) were calculated to describe the experimental results of the study population. Statistical analysis was performed using R3.6.1, and a p-value less than 0.05 was considered statistically significant.

[0058] 3. Steps for building a joint diagnosis model S101: Randomly select 3 to 10 marker concentration matrices from 10 protein markers, including SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, and ORM1, in the samples of the model group as the original training data set.

[0059] S102: Select the generalized linear model (glmnet) algorithm for constructing the prediction model, and the grid search range for optimizing the algorithm's hyperparameters. In this step, the grid search range for model hyperparameter optimization is set for each algorithm as shown in Table 3.

[0060] Table 3: Parameter grid search ranges for the glmnet algorithm

[0061] S103: According to the algorithm and hyperparameter setting range set in step S102, one of the hyperparameter combinations is selected as the parameters for constructing the prediction model.

[0062] S104: Split the original dataset into K subsets using a K-fold cross validation mechanism. To ensure that the ratio of majority class samples to minority class samples in each subset is the same as in the original dataset, a Stratified K-Folds cross validation mechanism is used for data segmentation.

[0063] S105 , according to the K training data subsets obtained by segmentation in step S104 , one of the subsets is selected as a validation set Ddev.

[0064] S106: Combine the training data subsets not selected in step S105 to form a training data pool Dtrain.

[0065] S107 , building a prediction model based on the selected supervised classification algorithm and hyperparameters according to the training data set Dtrain obtained in step S106 .

[0066] S108: According to the prediction model obtained in step S107, the verification set Ddev is evaluated to obtain the AUC value, and the current prognosis prediction model and the corresponding AUC value are stored in the prediction model pool Pool.

[0067] Step S108 evaluates the prediction model obtained in step S107 on the validation set determined in the current iteration. The model and evaluation results are stored in the prediction model pool for future selection of prediction models. The evaluation mentioned in this step can be the AUC value or other reasonable indicators for evaluating model performance.

[0068] S109: Determine whether all subsets have been used as validation sets. Step S109 determines whether all K subsets obtained in step S104 have been used as validation sets and trained on the model. If all subsets have been used as validation sets and training has been completed, proceed to step S110; if any subsets have not been used as validation sets, proceed to step S105. This step ensures that every sample in the original dataset has been used as a validation set, improving model stability and preventing overfitting of the model to a particular subset.

[0069] S110: The average AUC value of all models in the prediction model pool Pool is used as the final performance evaluation value of the combined model. The model parameters and the final performance evaluation AUC value are stored in the optimal model pool Poolbest.

[0070] S111: Determine whether all hyperparameter combinations have been used to construct prediction models. Step S111 determines whether prediction models have been constructed for all algorithms and corresponding hyperparameter combinations obtained in step S102. If all combinations have been used to construct models, step S112 is executed. If any combination has not been used to construct models, step S103 is executed.

[0071] S113 , selecting the model with the largest AUC value from the model set Poolbest obtained in step S112 as the final prediction model for pancreatic cancer diagnosis.

[0072] S114, repeat all the above steps until all combinations of markers are modeled.

[0073] 4. Determination of the optimal combination of markers By executing the above model building steps, we obtained the optimal model constructed by all combinations of markers. In order to compare the performance of the models under these different marker combinations, we used the ROC method to evaluate the AUC values ​​of these models in the test group. As shown in Table 4 and Figure 3 As shown: Table 4: Comparison of the area under the ROC curve of models constructed with different marker combinations

[0074] As shown in Table 4, the AUCs for the 6MP, 7MP, 8MP, 9MP, and 10MP combinations were all greater than 0.65, with the maximum AUCs all greater than 0.75, demonstrating excellent performance. The AUC for the model (10MP) consisting of S100B, SPINK5, SERPINB1, KRT19, ORM1, CD74, PGM5, SELL, DEFA3, and RNASE1 was higher than the mean values ​​for the models with other marker combinations.

[0075] 5. 10MP model parameter optimization results Through the above analysis, we found that the optimal combination of markers is S100B+SPINK5+SERPINB1+KRT19+ORM1+CD74+PGM5+SELL+DEFA3+RNASE1. Based on this marker combination, we analyzed the models constructed under 9 different combinations of glmnet algorithm hyperparameters ( Figure 4 ), and the model performance was evaluated by AUC value. As shown in Table 5 and Figure 4 As shown in the figure, when the glmnet algorithm hyperparameter combination is alpha = 0.1, lambda = 0.0055, the AUC reaches a maximum value of 0.908 (the AUC is calculated using the 10-fold cross-validation method during the modeling process).

[0076] Table 5: AUC of the model constructed under different hyperparameter combinations of the glmnet algorithm

[0077] The equation for building a model based on the optimal hyperparameter combination is:

[0078] logit(Y)=log(Y / 1-Y). The formula for calculating Y is as follows:

[0079] Where Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers (m = 10), Xi represents the detection value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker (Table 6), and b is a constant of -5.78.

[0080] Table 6: Coefficients of the 10 biomarkers in the model

[0081] 6. Determination of the diagnostic threshold of the pancreatic cancer combined diagnostic model (10MP) The ROC curve was drawn with the predicted value in the model group, and the optimal diagnostic cutoff value of 0.473 was set according to the Youden index value. That is, when the predicted value of the diagnostic model is ≤0.473, the patient is considered not to be a pancreatic cancer patient; when the predicted value of the model is >0.473, the patient is considered to be a pancreatic cancer patient. The results are as follows Figure 5 As shown: The model has an AUC of 0.948508, a sensitivity of 0.922, and a specificity of 0.864 in the model group.

[0082] 7. Validation of the Pancreatic Cancer Combined Diagnostic Model (10MP) Draw the ROC curve with the predicted value in the test group, such as Figure 6 As shown in the figure, the AUC is 0.91854. The optimal diagnostic cutoff value is set to 0.473 based on the Youden index value. That is, when the diagnostic model prediction value is ≤0.473, the patient is considered not to be a pancreatic cancer patient; when the model prediction value is >0.473, the patient is considered to be a pancreatic cancer patient. The results are shown in the figure. Figure 6 As shown: The model has an accuracy of 0.861, a sensitivity of 0.88, and a specificity of 0.842 in the test group.

[0083] Example 3: Comparison of diagnostic value of different pancreatic cancer diagnostic models Table 7: Comparison of the area under the ROC curve of different diagnostic models

[0084] As shown in Table 7, the AUCs of our model (10MP) were 0.319 and 0.223 higher than those of traditional single markers, respectively. DeLong's test, using the AUC significance test, showed that the diagnostic value of our model (10MP) was significantly (p < 0.05) higher than that of traditional markers or traditional marker combination models.

[0085] Example 4: Construction of a pancreatic cancer clinical staging detection model This example attempts to construct a pancreatic cancer clinical staging diagnostic model based on seven protein markers in serum: CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1, and FABP4, hoping to use these markers to detect and diagnose the clinical staging of pancreatic cancer.

[0086] A total of 100 patients with clinically confirmed pancreatic cancer at stages I, II, III, and IV were enrolled, including 64 patients at stage I, 38 patients at stage II, 44 patients at stage III, and 54 patients at stage IV. All samples were obtained from biopsied patients and pathologically confirmed. The participants were divided into a model group and a test group at a ratio of 8:2.

[0087] The data of the model group is used to construct a classification model through a machine learning algorithm, and the model satisfies:

[0088] i represents the i-th biomarker, m represents the number of biomarkers (m=10), Xi represents the detection value of the i-th biomarker (μg / mL), and b and Ki are parameters to be optimized.

[0089] P represents the probability of pancreatic cancer being in any clinical stage: I, II, III, or IV.

[0090] The model is trained using logistic regression or neural network through maximum likelihood estimation or gradient descent to obtain the classification model parameters b and Ki for each clinical stage, thereby obtaining a prediction model for each clinical stage.

[0091] A confusion matrix was constructed using the prediction model for each clinical stage, a classification report was generated, and a receiver operating characteristic (ROC) curve was plotted. The optimal diagnostic cutoff value was set based on the Youden index. Specifically, when the model-predicted value P ≤ the cutoff value, the patient was deemed to be outside the clinical stage; when the model-predicted value > the cutoff value, the patient was deemed to be within the clinical stage.

[0092] The test set data were input into the model of each clinical stage for validation evaluation and parameter optimization.

[0093] If a certain indicator (such as AUC, accuracy) does not meet the preset requirements, the model will be iterated.

[0094] According to the best combination screened out in Table 4, the maximum AUC area under the ROC curve of the prediction model for clinical staging constructed for different marker combinations is shown in Table 8.

[0095] Table 8: Comparison of the maximum AUC area under the ROC curve for the prediction model of clinical stage constructed by different marker combinations

[0096] The results in Table 8 show that the maximum AUCs for both 9MP and 10MP were greater than 0.8, demonstrating good performance. 9MP performed best in the model predicting clinical phase III, while the 10MP combination performed best in predicting clinical phases I, II, and IV. Therefore, the 9MP combination was selected to construct a clinical phase III classification model, while the 10MP combination was selected to construct a clinical phase I, II, and IV classification model.

[0097] The test set data was input into the above four clinical staging models. The AUC area under the ROC curve, accuracy, sensitivity, and specificity are shown in Table 9 below.

[0098] Table 9 Test set performance verification results

[0099] The present invention uses 9 markers to establish a clinical stage III classification model and 10 markers to construct a clinical stage I, II, and IV classification model. The AUC is greater than 0.8, and the accuracy, sensitivity, and specificity are all greater than 70%. This shows that the pancreatic cancer clinical staging model of the present invention can effectively distinguish between pancreatic cancer stages I, II, III, and IV, and can therefore be used to monitor the progression of pancreatic cancer.

Claims

1. The application of biomarkers in the preparation of pancreatic cancer detection products is characterized by The biomarker is selected from one or more of the following: SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, ORM1.

2. The use according to claim 1, characterized in that The biomarker is one of the following combinations: (A) Combination of SERPINB1, CD74, KRT19, S100B, ORM1, and DEFA3; (ii) combination of KRT19, PGM5, SELL, DEFA3, S100B, CD74, and SERPINB1; (iii) combination of CD74, DEFA3, RNASE1, ORM1, SPINK5, KRT19, SERPINB1, and PGM5; (IV) combination of SERPINB1, ORM1, RNASE1, DEFA3, PGM5, SPINK5, KRT19, S100B, and CD74; (V) Combination of S100B, SPINK5, SERPINB1, KRT19, ORM1, CD74, PGM5, SELL, DEFA3, and RNASE1.

3. A kit for detecting pancreatic cancer, characterized in that The kit includes reagents for detecting biomarkers, and the biomarkers are selected from one or more of the following: SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, and ORM1.

4. The kit for detecting pancreatic cancer according to claim 3, wherein The biomarker is one of the following combinations: (A) Combination of SERPINB1, CD74, KRT19, S100B, ORM1, and DEFA3; (ii) combination of KRT19, PGM5, SELL, DEFA3, S100B, CD74, and SERPINB1; (iii) combination of CD74, DEFA3, RNASE1, ORM1, SPINK5, KRT19, SERPINB1, and PGM5; (IV) combination of SERPINB1, ORM1, RNASE1, DEFA3, PGM5, SPINK5, KRT19, S100B, and CD74; (V) Combination of S100B, SPINK5, SERPINB1, KRT19, ORM1, CD74, PGM5, SELL, DEFA3, and RNASE1.

5. A system for detecting pancreatic cancer, characterized in that The system includes a data analysis module, which is used to analyze the detection values ​​of biomarkers in the sample of the patient to be tested, and the biomarkers are selected from one or more of the following: SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, and ORM1.

6. The system for detecting pancreatic cancer according to claim 5, characterized in that The biomarkers are a combination of SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3 and ORM1.

7. The system for detecting pancreatic cancer according to claim 6, wherein The data analysis module includes an analysis model equation, which is as follows: logit(Y)=log(Y / 1-Y). The formula for calculating Y is as follows: Where Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers, Xi represents the detection value of the i-th biomarker, m = 10, the unit is μg / mL, Ki represents the coefficient of the i-th biomarker, and b is a constant of -5.78; The coefficients for each biomarker are as follows: The coefficient of S100B is 9.349; the coefficient of SPINK5 is 6.006; the coefficient of SERPINB1 is 1.465; the coefficient of KRT19 is 3.508; the coefficient of ORM1 is 6.933; the coefficient of CD74 is 2.159; the coefficient of PGM5 is 4.411; the coefficient of SELL is 2.325; the coefficient of DEFA3 is 6.300; and the coefficient of RNASE1 is 2.

931.

8. The system for detecting pancreatic cancer according to claim 5 or 6, characterized in that The data analysis module calculates a prediction value of whether the patient to be tested has pancreatic cancer using an analysis model equation according to the detection value of the biomarker, and determines whether the patient to be tested has pancreatic cancer based on the prediction value; The judgment conditions are: When the predicted value is less than or equal to a preset threshold, it is determined that the patient to be tested is not a pancreatic cancer patient; When the predicted value is greater than a preset threshold, the patient to be tested is determined to be a pancreatic cancer patient; The preset threshold is 0.

473.

9. The system for detecting pancreatic cancer according to claim 5 or 6, characterized in that The system includes a data detection module, a data input module, and a data output module; the data detection module is used to detect biomarkers in a sample and obtain a detection value; the data input module is used to input the detection value of the biomarker, and after the data analysis module analyzes the detection value, the data output interface is used to output the analysis result of whether the patient to be tested has pancreatic cancer.

10. Application of biomarkers in the preparation of pancreatic cancer clinical staging diagnostic products, characterized in that The biomarkers are a combination of SERPINB1, ORM1, RNASE1, DEFA3, PGM5, SPINK5, KRT19, S100B, and CD74, or a combination of SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, and ORM1.

Citation Information

Patent Citations

  • ELISA test kit of human-derived soluble CD74 protein and detection method thereof

    CN103207277A

  • Kit for detecting pancreatic cancer cells in peripheral blood

    CN107843731A

  • System for pancreatic cancer detection and reagent or kit thereof

    CN116626297A

  • Methods of diagnosis and prognosis of pancreatic cancer

    US20060269921A1

  • Exosome based gene expression analysis for cancer management

    WO2019008414A1