Biomarkers for detecting pancreatic cancer and their applications
By screening and combining new biomarker combinations, and using LC-MS/MS technology to detect pancreatic cancer biomarkers, a pancreatic cancer detection system was constructed. This system addresses the insufficient accuracy of early pancreatic cancer diagnosis in existing technologies, achieving efficient early prediction and differentiation between benign and malignant pancreatic cancer, and possessing high diagnostic accuracy and staging capabilities.
Patent Information
- Application Number
- CN202510970188.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2025-06-03
- Filing Date
- 2025-07-15
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-07-15
AI Technical Summary
Existing technologies for the early diagnosis of pancreatic cancer lack accuracy. Common biomarkers such as CA19-9 have poor specificity and are difficult to effectively distinguish between benign and malignant tumors, often resulting in late-stage diagnosis and loss of the opportunity for surgical resection.
A new set of biomarkers was screened, including SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, and ORM1. The expression levels of these biomarkers in body fluids were detected by LC-MS/MS technology. Combined with the data analysis module, a pancreatic cancer detection system was constructed to achieve early prediction and differentiation between benign and malignant pancreatic cancer.
The model improved the accuracy and sensitivity of early diagnosis of pancreatic cancer. The AUC of the constructed model in the model group and the test group were 0.948508 and 0.91854, respectively, and the sensitivity and specificity were 0.922 and 0.864, and 0.88 and 0.842, respectively. It can effectively distinguish between benign and malignant pancreatic cancer and clinical stage.
Smart Images

Figure CN120761644B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical testing technology, specifically involving the use of proteomics to screen biomarkers for pancreatic cancer, for the detection of pancreatic cancer, and to differentiate between benign and malignant pancreatic tumors. Background Technology
[0002] Proteomics is the science that studies the composition, location, changes, and interactions of proteins in cells, tissues, or organisms, including the study of protein expression patterns and proteome functional patterns. With the development of mass spectrometry, liquid chromatography-mass spectrometry (LC-MS / MS) has become a core tool in proteomics research, providing crucial support for the discovery of disease diagnostic biomarkers, the screening of drug targets, and toxicological studies.
[0003] Pancreatic cancer is a highly malignant tumor, difficult to detect in its early stages, with a very poor prognosis and a high incidence and mortality rate. Currently, clinical diagnosis mainly relies on physical diagnostic methods such as ultrasound, magnetic resonance imaging (MRI), and endoscopic ultrasound. Although these methods have some diagnostic efficiency, over 80% of pancreatic cancer patients are diagnosed at an advanced stage, losing the opportunity for surgical resection. Therefore, improving the accuracy of early diagnosis and reducing the difficulty of diagnosis are crucial to improving patient survival rates.
[0004] Currently, the protein biomarker CA1909 is the most common and widely used tumor marker in clinical practice for the diagnosis and prognostic monitoring of pancreatic cancer. However, CA1909 as a biomarker still has some limitations, such as poor specificity, low expression levels in Lewis-negative phenotypes, and increased false-positive rates in patients with benign diseases such as pancreatitis, cirrhosis, and acute cholangitis. Other common protein biomarkers, such as CEA and TP53, also have certain deficiencies in sensitivity and specificity.
[0005] Against this backdrop, the search for new pancreatic cancer biomarkers and their combinations for detecting pancreatic cancer and differentiating between benign and malignant tumors, as well as the construction of predictive models for benign and malignant pancreatic cancer, has significant clinical value. Summary of the Invention
[0006] To address the problems existing in the prior art, this invention provides a biomarker, detection reagent, detection kit, and pancreatic cancer risk prediction system for pancreatic cancer detection. This invention screens out a series of novel biomarkers that can predict the risk of pancreatic cancer at an early stage and can distinguish between benign and malignant pancreatic tumors.
[0007] The technical solution adopted in this invention is:
[0008] Application of biomarkers in the preparation of pancreatic cancer detection products, wherein the biomarkers are selected from one or more of the following: SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, ORM1;
[0009] Preferably, the biomarker is one of the following combinations:
[0010] (a) Combination of SERPINB1, CD74, KRT19, S100B, ORM1, and DEFA3;
[0011] (ii) A combination of KRT19, PGM5, SELL, DEFA3, S100B, CD74, and SERPINB1;
[0012] (iii) Combinations of CD74, DEFA3, RNASE1, ORM1, SPINK5, KRT19, SERPINB1, and PGM5;
[0013] (iv) Combinations of SERPINB1, ORM1, RNASE1, DEFA3, PGM5, SPINK5, KRT19, S100B, and CD74;
[0014] (v) Combinations of S100B, SPINK5, SERPINB1, KRT19, ORM1, CD74, PGM5, SELL, DEFA3, and RNASE1;
[0015] More preferably, it is a combination of SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3 and ORM1.
[0016] The pancreatic cancer detection product can be a pancreatic cancer detection reagent or a pancreatic cancer detection kit.
[0017] The present invention also provides a kit for pancreatic cancer detection, the kit comprising reagents for detecting biomarkers selected from one or more of the following: SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, ORM1.
[0018] Preferably, the biomarker is one of the following combinations:
[0019] (a) Combination of SERPINB1, CD74, KRT19, S100B, ORM1, and DEFA3;
[0020] (ii) A combination of KRT19, PGM5, SELL, DEFA3, S100B, CD74, and SERPINB1;
[0021] (iii) Combinations of CD74, DEFA3, RNASE1, ORM1, SPINK5, KRT19, SERPINB1, and PGM5;
[0022] (iv) Combinations of SERPINB1, ORM1, RNASE1, DEFA3, PGM5, SPINK5, KRT19, S100B, and CD74;
[0023] (v) Combinations of S100B, SPINK5, SERPINB1, KRT19, ORM1, CD74, PGM5, SELL, DEFA3, and RNASE1;
[0024] More preferably, it is a combination of SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3 and ORM1.
[0025] Furthermore, the reagents for detecting biomarkers can detect the expression level of biomarkers. These reagents can be sample pretreatment reagents, antigens or antibodies, or other biological reagents and kits suitable for biomarker detection; they can also be developed into standardized reagents or kits suitable for LC-UV or LC-MS detection of biomarkers.
[0026] In some methods, the reagent for detecting the biomarker is an antibody that is a biomarker as described above, and the antibody is a monoclonal antibody.
[0027] The SELL is an L-selectin, which is a protein or amino acid sequence with the UniProt database number P14151;
[0028] The CD74 is an HLA-II class histocompatibility antigen γ chain, which is a protein or amino acid sequence with UniProt database number P04233;
[0029] SERPINB1 is a leukocyte elastase inhibitor and is a protein or amino acid sequence with UniProt database number P30740.
[0030] The RNASE1 is an islet ribonuclease, and is a protein or amino acid sequence with the UniProt database number P07998;
[0031] The KRT19 is type I cytoskeleton keratin 19, which is the protein or amino acid sequence with UniProt database number P08727;
[0032] The S100B is protein S100-B, which is a protein or amino acid sequence with the UniProt database number P04271;
[0033] The PGM5 is a phosphoglucoside mutase 5, which is a protein or amino acid sequence with the UniProt database number Q15124;
[0034] The SPINK5 is a Kazal-type serine protease inhibitor 5, and is a protein or amino acid sequence with the UniProt database number Q9NQ38;
[0035] DEFA3 is neutrophil defensin 3, which is the protein or amino acid sequence with the UniProt database number P59666;
[0036] The ORM1 is α-1-acidic glycoprotein 1, which is the protein or amino acid sequence with the UniProt database number P02763.
[0037] Furthermore, the reagent is used to detect biomarkers in body fluid samples, which include any one of blood, urine, saliva, and sweat.
[0038] In some preferred embodiments, the biomarkers of the present invention are obtained through screening blood samples, and are particularly suitable for development into blood test reagents or kits for pancreatic cancer prediction.
[0039] Furthermore, biomarker detection refers to detecting the presence, relative abundance, or concentration of biomarkers in an individual's bodily fluid samples.
[0040] In some methods, relative abundance is preferred, which is the peak area of the biomarker in the detection chromatogram obtained by high performance liquid chromatography-tandem mass spectrometry. For example, if the average peak area of a biomarker measured in a control sample (an individual without pancreatic cancer) is 500, and the average peak area measured in a pancreatic cancer sample is 3000, then the abundance of the biomarker in the pancreatic cancer sample is considered to be 6 times that in the control sample.
[0041] The present invention also provides a system for detecting pancreatic cancer, the system including a data analysis module, the data analysis module being used to analyze the detection values of biomarkers in a sample of a patient to be tested, the biomarkers being selected from one or more of the following: SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, ORM1;
[0042] Preferably, the biomarker is one of the following combinations:
[0043] (a) Combination of SERPINB1, CD74, KRT19, S100B, ORM1, and DEFA3;
[0044] (ii) A combination of KRT19, PGM5, SELL, DEFA3, S100B, CD74, and SERPINB1;
[0045] (iii) Combinations of CD74, DEFA3, RNASE1, ORM1, SPINK5, KRT19, SERPINB1, and PGM5;
[0046] (iv) Combinations of SERPINB1, ORM1, RNASE1, DEFA3, PGM5, SPINK5, KRT19, S100B, and CD74;
[0047] (v) Combinations of S100B, SPINK5, SERPINB1, KRT19, ORM1, CD74, PGM5, SELL, DEFA3, and RNASE1;
[0048] More preferably, it is a combination of SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3 and ORM1.
[0049] The data analysis module includes an analysis model equation, which is as follows:
[0050]
[0051] logit(Y) = log(Y / 1 - Y), the formula for calculating Y is as follows:
[0052]
[0053] Where Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers (m=10), Xi represents the detection value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker, and b is a constant -5.78.
[0054] The coefficients of Ki are shown in Table 6 below:
[0055] Table 6: Coefficients of 10 biomarkers in the model
[0056]
[0057] The data analysis module calculates a predictive value for whether the patient under test has pancreatic cancer based on the detection value of the biomarker using an analytical model equation, and determines whether the patient under test has pancreatic cancer based on the predicted value.
[0058] The determination criteria are as follows:
[0059] When the predicted value is less than or equal to a preset threshold, the patient to be tested is determined not to be a pancreatic cancer patient.
[0060] When the predicted value is greater than a preset threshold, the patient to be tested is determined to be a pancreatic cancer patient;
[0061] The preset threshold is 0.473.
[0062] Furthermore, the system also includes a data detection module, a data input module, and a data output module; the data detection module is used to detect biomarkers in the sample and obtain detection values; the data input module is used to input the detection values of the biomarkers, and after the data analysis module analyzes the detection values, the data output interface is used to output the analysis results of whether the patient under test has pancreatic cancer.
[0063] The detection value is generally obtained by performing enzyme-linked immunosorbent assay (ELISA) on the sample to obtain the concentration of biomarkers in the sample, which is used as the detection value in μg / mL.
[0064] The present invention also provides the application of biomarkers in the preparation of clinical staging diagnostic products for pancreatic cancer, wherein the biomarkers are a combination of SERPINB1, ORM1, RNASE1, DEFA3, PGM5, SPINK5, KRT19, S100B, and CD74, or a combination of SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, and ORM1.
[0065] The technical solution of the present invention has the following beneficial effects:
[0066] (1) This invention uses proteomics technology to systematically screen differential proteins in blood samples of pancreatic cancer patients and healthy controls, screen out biomarkers that can predict the risk of tumor development in the early stage of pancreatic cancer, and construct a pancreatic cancer detection system based on these biomarkers, thereby realizing non-invasive, convenient and efficient prediction of the benign and malignant status of individual pancreatic tumors, meeting the needs of clinical testing, and having great application prospects.
[0067] (2) The pancreatic cancer detection model of the present invention has an AUC of 0.948508, a sensitivity of 0.922, and a specificity of 0.864 in the model group and an AUC of 0.91854, an accuracy of 0.861, a sensitivity of 0.88, and a specificity of 0.842 in the test group. It has high accuracy and discrimination ability and can more efficiently predict whether an individual has pancreatic cancer.
[0068] (3) The combination of biomarkers of the present invention can construct a clinical staging prediction model for pancreatic cancer, which can effectively distinguish between stage I, II, III and IV pancreatic cancer and has the potential to diagnose the clinical staging of pancreatic cancer. Attached Figure Description
[0069] Figure 1 Volcano plot for differential analysis of benign and malignant pancreatic tumors.
[0070] Figure 2 The graph shows the ROC and OPLS-DA analysis results for benign and malignant pancreatic tumors.
[0071] Figure 3 A bar chart comparing the performance AUC of models built for different combinations of markers.
[0072] Figure 4 AUC results for models with different hyperparameters.
[0073] Figure 5 The ROC curves of the combined diagnostic model in the model group are shown.
[0074] Figure 6 The ROC curve of the joint diagnostic model in the test group. Detailed Implementation
[0075] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be noted that the embodiments described below are intended to facilitate understanding of the present invention and are not intended to limit it in any way. The reagents used in this embodiment are all known products and were obtained by purchasing commercially available products.
[0076] It should be noted that:
[0077] (1) Diagnosis or testing
[0078] Here, diagnosis or testing refers to the detection or analysis of biomarkers in a sample, or the determination of the content of a target biomarker, such as its absolute or relative content. The presence or quantity of the target biomarker then indicates whether the individual providing the sample may have or suffer from a certain disease, or the likelihood of having a certain disease. The meanings of diagnosis and testing are interchangeable here. The results of such testing or diagnosis cannot be directly considered as a direct result of disease; rather, they are intermediate results. To obtain a direct result, further auxiliary methods such as pathology or anatomy are needed to confirm the presence of a certain disease. For example, this invention provides several novel biomarkers associated with pancreatic cancer, and changes in the levels of these biomarkers are directly related to the presence or absence of pancreatic cancer.
[0079] (2) The link between biomarkers or markers and pancreatic cancer
[0080] In this invention, "marker" and "biomarker" have the same meaning. Here, "linkage" refers to a direct correlation between the presence or change in the level of a certain biomarker in a sample and a specific disease; for example, a relative increase or decrease in the level indicates a higher likelihood of having the disease compared to healthy individuals.
[0081] If multiple different biomarkers are present simultaneously in a sample, or if their relative levels change, it indicates a higher likelihood of having the disease compared to healthy individuals. In other words, among biomarkers, some are strongly associated with the disease, some are weakly associated, and some may not even be associated with a specific disease. One or more of the strongly associated biomarkers can be used as diagnostic markers, while weakly associated biomarkers can be combined with strong biomarkers to diagnose a disease, increasing the accuracy of test results.
[0082] The numerous serum biomarkers discovered in this invention can be used to differentiate pancreatic cancer from healthy individuals. These biomarkers can be used individually for direct detection or diagnosis, indicating a strong correlation between the relative changes in their levels and pancreatic cancer. Of course, it is understood that one or more biomarkers strongly associated with pancreatic cancer can be detected simultaneously. It is generally understood that in some methods, selecting highly correlated biomarkers for detection or diagnosis can achieve a certain level of accuracy, such as 60%, 65%, 70%, 80%, 85%, 90%, or 95%. This indicates that these biomarkers can provide intermediate values for diagnosing a certain disease, but does not necessarily mean a direct diagnosis of that disease.
[0083] Of course, differentially expressed proteins with higher ROC values can also be selected as diagnostic biomarkers. The terms "strong" and "weak" are generally determined using algorithms, such as the contribution rate of the biomarker to pancreatic cancer or weighted analysis. Such calculation methods can include significance analysis (p-value or FDR value) and fold change. Multivariate statistical analyses mainly include principal component analysis (PCA), partial least squares discriminant analysis (PLS-DA), and orthogonal partial least squares discriminant analysis (OPLS-DA), as well as other methods such as ROC analysis. Other model prediction methods are also possible. When specifically selecting biomarkers, the differentially expressed proteins disclosed in this invention can be chosen, or combinations of other existing known biomarkers can be selected and predicted using model methods.
[0084] Example 1: Screening of pancreatic cancer biomarkers
[0085] 1. Experimental Design
[0086] This experiment was designed to collect plasma samples from newly diagnosed pancreatic tumor patients, enrich low-abundance proteins using immunoaffinity chromatography to remove high-abundance proteins, and detect the protein abundance in the samples using a high-performance liquid chromatography-mass spectrometry tandem device. The differences in protein abundance between patients with benign and malignant pancreatic tumors were analyzed, and the diagnostic performance was also analyzed.
[0087] 2. Sample Collection
[0088] Blood samples from 100 patients with benign and malignant pancreatic tumors were included. All pancreatic tumor patients underwent biopsy and were pathologically confirmed. Approximately 2 ml of peripheral blood was collected from each subject before chemotherapy, radiotherapy, and surgery. The blood was mixed in a vacuum tube containing EDTA anticoagulant and centrifuged at 120 g at room temperature for 10 minutes. The supernatant was collected, and this process was repeated twice. Subsequently, the centrifuge was performed at 360 g for 20 minutes. Platelet samples were then collected in centrifuge tubes and stored at -80°C for later use.
[0089] 3. Protein sample processing and enzymatic digestion
[0090] First, plasma samples were centrifuged for 15 minutes (15,000g), and the supernatant was collected, filtered, and subjected to immunoaffinity chromatography to remove 14 high-abundance proteins. Then, low-abundance proteins were concentrated to 350 μL using a 3 kDa cutoff concentration tube at 4000 g for 1 hour. The concentrate was recovered, and buffer exchange was performed using a 7 kDa cutoff desalting column at 1000 g for 2 minutes. The replacement buffer was AEX-A (20 mM Tris, 4 M Urea, 3% isopropanol, pH 8.0). Using AEX-A as a blank, protein concentration in the samples was determined using the BCA method. According to sample grouping, 25 mL of TCEP was added to the samples, and incubation was performed at 37°C for 30 minutes for protein reduction. Then, the corresponding TMT 16-plex reagent was added, and incubation was performed at room temperature in the dark for 1 hour for TMT labeling. Subsequently, the sample was buffer-replaced using a Zeba column with AEX-A as the replacement buffer. After mixing the TMT 16-plex labeled sample, 2 mL of AEX-A was added to the mixture, resulting in a final volume of 5.5 mL. The sample was filtered using a 0.22 μm filter, and the TMT 16-plex labeled sample was separated using a 2D-HPLC system. The collected fractions were lyophilized, and finally, a Trypsin-LysinC mixed enzyme was added, and the sample was incubated at 37°C for 5 hours to digest the enzymes. 5 μL of 10% TFA was added to terminate the digestion reaction. A total of 60 digested 2D-HPLC fractions were used for nanoLC-MS / MS analysis.
[0091] 4. LC-MS / MS Data Acquisition
[0092] Each sample was separated using an Easy nLC-1200 high-resolution liquid chromatography system with a flow rate of nanoliters and coupled online to a Q Exactive HF-X mass spectrometer. Mobile phase A was a 0.1% formic acid aqueous solution, and mobile phase B was a 0.1% formic acid-acetonitrile aqueous solution (acetonitrile 80%, water 20%). The chromatographic column consisted of an enrichment column and an analytical column, equilibrated with 100% mobile phase A. Samples were loaded onto the enrichment column (specifications: 100 μm inner diameter (ID), 4 cm length (L), C18 packing material, 3 μm particle size, 100 Å pore size) via an autosampler, and then separated on the analytical column (75 μm inner diameter, 25 cm length, C18 packing material, 3 μm particle size, 100 Å pore size) at a flow rate of 300 nL / min. After chromatographic separation, the samples were analyzed by mass spectrometry using a Q Exactive HF-X mass spectrometer. The detection method is positive ion, with a precursor ion scan range of 350-1800 m / z. The primary mass spectrometry resolution is 120,000 at 200 m / z. The automatic gain control (AGC) target is 3 × 10⁻⁶. 6 The maximum injection time was 50 ms, and the dynamic exclusion time was 40 s. The mass-charge ratio of the peptide and peptide fragments was acquired using a data-dependent acquisition (DDA) method: 20 secondary spectra (MS / MS, MS2 scans) were acquired after each full scan (primary mass spectrometry). The MS2 activation type was high-energy collisional dissociation (HCD), the selection window was 0.7 m / z, the secondary mass spectrometry resolution was 30,000 at 200 m / z, and the automatic gain control target was set to 1 × 10⁻⁶. 5 The maximum injection time is 65 ms, the fixed first term mass is 110.0 m / z, the normalized collision energy is 32 eV, and the minimum automatic gain control target is 2.00 × 10⁻⁶. 4 It excludes ions with 1-valent, 6–8-valent, and >8-valent charges, allowing only single charge states, sets peptide matching as a priority, and enables the exclusion of isotopes.
[0093] 5. Data Preprocessing
[0094] Secondary mass spectrometry (PMS) data were retrieved using Maxquant (v1.6.15.0). The data type was DIA proteomics data based on secondary reporter ion quantification; the secondary spectra used for quantification required the precursor ion to account for more than 75% of the primary spectrum. The database source was the Uniprot database's Homo_sapiens_9606_proteome_gene (release: 2021-10-14, sequence: 20,437), and common contamination libraries were added, with contaminating proteins removed during data analysis. Enzyme digestion was set to Trypsin / P; the number of missed cleavage sites was set to 2; the precursor ion mass error tolerance for First Search and Main Search was set to 20 ppm and 5 ppm, respectively, and the secondary fragment ion mass error tolerance was 20 ppm. Fixed modification was cysteine alkylation, and variable modifications included methionine oxidation and N-terminal acetylation of proteins. The free-dip ion ratio (FDR) for protein identification and PSM identification was set to 1%.
[0095] 6. Difference Analysis
[0096] A combination of univariate and multivariate statistical analyses was used to screen for differentially expressed proteins and transcripts. Univariate analysis primarily included significance analysis (p-value or FDR value) and fold change of characteristic molecules in different groups. Multivariate statistical analyses mainly included receiver operating characteristic (ROC) curve analysis and Boruta feature selection based on the random forest algorithm. All statistical analyses were performed using R; specific R-related information is shown in Table 1.
[0097] Table 1: R used in this invention and related information
[0098]
[0099] Variable Importance for the Projection (VIP) was calculated to measure the influence and explanatory power of each protein's expression pattern on the classification of each group of samples. A Wilcoxon rank-sum test was then performed to obtain the corrected p-value (FDR). Based on the criteria of FDR < 0.01 and Fold change > 2, 67 downregulated and 66 upregulated proteins were obtained (see details). Figure 1 ).
[0100] To evaluate the role of each biomarker in the diagnosis and prediction of pancreatic cancer, we used ROC and Boruta analyses to evaluate each biomarker. The results are shown in the figure below. Figure 2The x-axis represents the AUC obtained from ROC analysis, and the y-axis represents -log10(FDR) calculated by the Wilcoxon test. The size of the point represents the VIP value obtained from Boruta analysis. Further screening based on VIP>3 and AUC>0.6 identified a total of 10 more significant candidate biomarkers, as detailed in Table 2.
[0101] Table 2: Differential markers of benign and malignant pancreatic tumors
[0102]
[0103] In Table 2, a smaller FDR value and / or a larger VIP value indicate, to some extent, a more significant difference in protein between the two groups, and also suggest that the protein may have higher diagnostic value.
[0104] Example 2: A classification model for identifying benign and malignant pancreatic tumors using 10 differentially expressed proteins and its establishment.
[0105] While a single biomarker can differentiate between benign and malignant pancreatic tumors in serum samples or predict pancreatic cancer, combining multiple biomarkers generally yields higher accuracy in differentiation or prediction. However, a single biomarker with higher accuracy in predicting pancreatic cancer does not necessarily have a greater effect in a combination with one or more other biomarkers. Furthermore, a higher number of biomarkers does not necessarily lead to higher predictive accuracy (AUC value) for the combined combination; therefore, extensive validation experiments are still needed.
[0106] This embodiment studies a model constructed from 10 protein markers in serum: SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, and ORM1.
[0107] 1. Obtaining data
[0108] Study population:
[0109] A total of 1000 samples were collected, including 500 blood samples from patients with benign pancreatic tumors and 500 blood samples from patients with malignant pancreatic tumors. All samples were obtained from biopsies confirmed by pathology. Participants were divided into a model group and a test group at a ratio of 8:2.
[0110] Inclusion criteria for pancreatic cancer patients: (a) no history of other malignant tumors, and (b) surgical treatment performed within one month of blood collection, with postoperative pathological confirmation of pancreatic cancer. All collected serum samples were stored in a serum bank at -80°C after obtaining informed consent.
[0111] In this embodiment, the collected serum samples were subjected to enzyme-linked immunosorbent assay (ELISA) to obtain the concentrations of SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, and ORM1 in the serum.
[0112] 2. Statistical Analysis of Experimental Data
[0113] In the model cohort, a combined diagnostic model for multiple pancreatic cancer biomarkers was constructed using a combination of machine learning methods. The area under the receiver operator characteristic (ROC) curve (AUC) was estimated using predicted probability values with a 95% confidence interval (CI) to evaluate the discriminative power of the multivariate diagnostic model. In the test cohort, the Youden index (YI) was calculated to determine the cut-off value for differentiating between primary and metastatic pancreatic cancer patients. Furthermore, ROCs for individual biomarkers and different subgroups were constructed and compared. Standard descriptive statistics, such as frequency, mean, median, positive predictive value (PPV), negative predictive value (NPV), and standard deviation (SD), were calculated to describe the experimental results in the study population. Statistical analysis was performed using R3.6.1, and a p-value less than 0.05 was considered statistically significant.
[0114] 3. Steps for constructing a joint diagnostic model
[0115] S101: From the 10 protein biomarkers (SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, ORM1) in the model group samples, the concentration matrix of 3 to 10 biomarkers is randomly selected as the original training dataset.
[0116] S102, Select the generalized linear model (glmnet) algorithm to construct the prediction model, and define the grid search range during the hyperparameter optimization process. In this step, the grid search range for hyperparameter optimization of each algorithm is set as shown in Table 3.
[0117] Table 3: Parameter grid search range of the glmnet algorithm
[0118]
[0119] S103. Based on the algorithm and hyperparameter setting range set in step S102, select one of the hyperparameter combination methods as the parameters for constructing the prediction model.
[0120] S104. The original dataset is split into K subsets using a K-fold cross-validation mechanism. To ensure that the proportion of majority and minority class samples in each subset is the same as in the original dataset, a stratified K-fold cross-validation mechanism is used for data splitting.
[0121] S105. Based on the K training data subsets obtained from step S104, select one subset as the validation set Ddev.
[0122] S106, combine the unselected subset of training data from step S105 to form the training data pool Dtrain.
[0123] S107. Based on the training dataset Dtrain obtained in step S106, construct a prediction model based on the selected supervised classification algorithm and hyperparameters.
[0124] S108. Based on the prediction model obtained in step S107, evaluate it on the validation set Ddev to obtain the AUC value, and store the current prognostic prediction model and the corresponding AUC value in the prediction model pool Pool.
[0125] Step S108 involves evaluating the prediction model obtained in step S107 on the validation set determined in the current iteration, and storing both the model and the evaluation results in the prediction model pool for future use in selecting the baseline prediction model. The evaluation mentioned in this step can be the AUC value or other reasonable metrics for evaluating model performance.
[0126] S109, determine if each subset has been used as a validation set. Step S109 checks if all K subsets obtained in step S104 have been used as validation sets for model training. If all subsets have been used as validation sets and training has been completed, proceed to step S110; otherwise, proceed to step S105. This step ensures that every sample in the original dataset has been used as a validation set, improving model stability and preventing the model from overfitting to a particular subset.
[0127] S110: The average AUC of all models in the prediction model pool (Pool) is used as the final performance evaluation value of the model for this combination method. The model parameters and the final performance evaluation AUC value are then stored in the optimal model pool (Poolbest).
[0128] S111: Determine whether all hyperparameter combinations have constructed a prediction model. Step S111 involves determining whether all algorithms and corresponding hyperparameter combinations obtained in step S102 have been used to construct prediction models. If all combinations have completed model construction, proceed to step S112; otherwise, if any combination has not completed model construction, proceed to step S103.
[0129] S113. From the model set Poolbest obtained in step S112, select the model with the largest AUC value as the final prediction model for pancreatic cancer diagnosis.
[0130] S114, Repeat all the above steps until all combinations of the markers have been modeled.
[0131] 4. Determining the optimal combination of markers
[0132] By executing the model building steps described above, we obtained the optimal model for all combinations of markers. To compare the performance of these models under different marker combinations, we used the ROC method to evaluate the AUC values of these models in the test group. See Table 4 below. Figure 3 As shown:
[0133] Table 4: Comparison of the area under the ROC curve for models constructed with different combinations of biomarkers
[0134]
[0135] As shown in Table 4, the AUCs of the 6MP, 7MP, 8MP, 9MP, and 10MP combinations were all greater than 0.65, with the maximum AUC being greater than 0.75, demonstrating good performance. Among them, the model composed of S100B+SPINK5+SERPINB1+KRT19+ORM1+CD74+PGM5+SELL+DEFA3+RNASE1 (10MP) had a higher AUC than the models with other biomarker combinations.
[0136] 5. Optimization results of 10MP model parameters
[0137] Based on the above analysis, we obtained the optimal combination of biomarkers as S100B+SPINK5+SERPINB1+KRT19+ORM1+CD74+PGM5+SELL+DEFA3+RNASE1. Based on this biomarker combination, we analyzed the models constructed under nine different combinations of glmnet algorithm hyperparameters. Figure 4 The model performance was evaluated using AUC values, as shown in Table 5. Figure 4As shown: When the hyperparameter combination of the glmnet algorithm is alpha=0.1 and lambda=0.0055, the AUC reaches the maximum value of 0.908 (the AUC is calculated using the 10x cross-validation method during the modeling process).
[0138] Table 5: AUC of the model built under different hyperparameter combinations of the glmnet algorithm
[0139]
[0140] The equations for the model constructed based on the optimal hyperparameter combination are:
[0141]
[0142] logit(Y) = log(Y / 1 - Y), the formula for calculating Y is as follows:
[0143]
[0144] Where Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers (m=10), Xi represents the detection value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker (Table 6), and b is a constant -5.78.
[0145] Table 6: Coefficients of 10 biomarkers in the model
[0146]
[0147] 6. Determination of diagnostic thresholds for the combined diagnostic model for pancreatic cancer (10MP)
[0148] ROC curves were plotted using the predicted values from the model group, and the optimal diagnostic cutoff value of 0.473 was set based on the Youden index. That is, when the model's predicted value was ≤0.473, the subject was considered not to have pancreatic cancer; when the model's predicted value was >0.473, the subject was considered to have pancreatic cancer. The results are as follows... Figure 5 As shown, the model's AUC was 0.948508, sensitivity was 0.922, and specificity was 0.864 in the model group.
[0149] 7. Validation of the combined diagnostic model for pancreatic cancer (10MP)
[0150] Plot the ROC curve using the predicted values from the test group, such as... Figure 6 As shown, the AUC was 0.91854. The optimal diagnostic cutoff value was set to 0.473 based on the Youden index. That is, when the diagnostic model's predicted value was ≤0.473, the subject was considered not to have pancreatic cancer; when the model's predicted value was >0.473, the subject was considered to have pancreatic cancer. The results are as follows... Figure 6 As shown, the model achieved an accuracy of 0.861, a sensitivity of 0.88, and a specificity of 0.842 in the test group.
[0151] Example 3: Comparison of the diagnostic value of different pancreatic cancer diagnostic models
[0152] Table 7: Comparison of Area Under the ROC Curve for Different Diagnostic Models
[0153]
[0154] As shown in Table 7, the AUC of our model (10MP) was 0.319 and 0.223 higher than that of traditional single biomarkers, respectively. The DeLong's test, a method for testing the significance of AUC differences, showed that the diagnostic value of our model (10MP) was significantly higher (p<0.05) than that of traditional biomarkers or combinations of traditional biomarkers.
[0155] Example 4: Construction of a Clinical Staging Model for Pancreatic Cancer
[0156] This embodiment attempts to construct a clinical staging diagnostic model for pancreatic cancer based on seven protein biomarkers in serum: CTSG, PGM5, ORM2, TFF1, SFRP1, TALDO1, and FABP4. The aim is to use these biomarkers to detect and diagnose the clinical stage of pancreatic cancer.
[0157] One hundred patients clinically diagnosed with stage I, II, III, and IV pancreatic cancer were collected, with 64, 38, 44, and 54 cases in stage I, II, III, and IV, respectively. All samples were obtained from pathologically confirmed biopsies. The participants were divided into a model group and a test group at a ratio of 8:2.
[0158] The data from the model group is used to construct a classification model using machine learning algorithms, and the model satisfies:
[0159]
[0160] i represents the i-th biomarker, m represents the number of biomarkers (m=10), Xi represents the detection value of the i-th biomarker (μg / mL), and b and Ki are parameters to be optimized.
[0161] P represents the probability of any clinical stage of pancreatic cancer, namely stage I, II, III, or IV.
[0162] The model is trained using maximum likelihood estimation or gradient descent, logistic regression or neural network, to obtain the classification model parameters b and Ki for each clinical stage, thereby obtaining the prediction model for each clinical stage.
[0163] A confusion matrix is constructed using the prediction model for each clinical stage, a classification report is generated, an ROC curve is plotted, and an optimal diagnostic cutoff value is set based on the Youden index. That is, when the diagnostic model's predicted value P ≤ the cutoff value, the patient is determined not to belong to that clinical stage; when the model's predicted value > the cutoff value, the patient is determined to belong to that clinical stage.
[0164] Test set data is input into the model for each clinical stage for validation, evaluation, and parameter optimization.
[0165] If a certain metric (such as AUC or accuracy) fails to meet the preset requirements, the model will proceed to iteration.
[0166] Based on the optimal combinations selected in Table 4, the maximum AUC area under the ROC curve of the predictive model for clinical staging constructed for different biomarker combinations is shown in Table 8.
[0167] Table 8: Comparison of the maximum AUC area under the ROC curve for predictive models of clinical staging constructed with different biomarker combinations
[0168]
[0169] Table 8 shows that both 9MP and 10MP have a maximum AUC greater than 0.8, indicating good performance. 9MP showed the best predictive performance in clinical stage III, while the 10MP combination performed best in predicting clinical stages I, II, and IV. Therefore, the 9MP biomarker combination was chosen to construct the clinical stage III classification model, and 10MP was chosen to construct the clinical stage I, II, and IV classification model.
[0170] The test set data was input into the models of the four clinical staging processes mentioned above. The AUC area, accuracy, sensitivity, and specificity under the ROC curve are shown in Table 9 below.
[0171] Table 9 Performance Verification Results of the Test Set
[0172]
[0173] This invention utilizes nine biomarkers to establish a clinical stage III classification model and ten biomarkers to construct clinical stage I, II, and IV classification models. The AUC is greater than 0.8, and the accuracy, sensitivity, and specificity are all greater than 70%, indicating that the clinical staging model for pancreatic cancer of this invention can effectively distinguish between stages I, II, III, and IV of pancreatic cancer, and therefore can be used to monitor the progression of pancreatic cancer.
Claims
1. The application of biomarkers in the preparation of pancreatic cancer detection products, characterized in that... The biomarkers are a combination of SERPINB1, ORM1, RNASE1, DEFA3, PGM5, SPINK5, KRT19, S100B, and CD74, or a combination of SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, and ORM1.
2. A kit for pancreatic cancer detection, characterized in that... The kit includes reagents for detecting biomarkers, which are combinations of SERPINB1, ORM1, RNASE1, DEFA3, PGM5, SPINK5, KRT19, S100B, and CD74, or combinations of SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, and ORM1.
3. A system for detecting pancreatic cancer, characterized in that... The system includes a data analysis module, which is used to analyze the detection values of biomarkers in the samples of the patients to be tested. The biomarkers are a combination of SERPINB1, ORM1, RNASE1, DEFA3, PGM5, SPINK5, KRT19, S100B, and CD74, or a combination of SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, and ORM1.
4. The system for detecting pancreatic cancer as described in claim 3, characterized in that... The biomarkers are a combination of SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, and ORM1.
5. The system for detecting pancreatic cancer as described in claim 4, characterized in that... The data analysis module includes an analysis model equation, which is as follows: logit(Y) = log(Y / 1 - Y), the formula for calculating Y is as follows: Where Y is the predicted value, i represents the i-th biomarker, m represents the number of biomarkers, Xi represents the detection value of the i-th biomarker, m=10, the unit is μg / mL, Ki represents the coefficient of the i-th biomarker, and b is a constant -5.78; The coefficients for each biomarker are as follows: The coefficients for S100B, SPINK5, SERPINB1, KRT19, ORM1, CD74, SGM5, DEFA3, and RNASE1 are 9.349, 6.006, 1.465, 6.933, 4.411, 2.325, 6.300, and 2.931, respectively.
6. The system for detecting pancreatic cancer as described in claim 5, characterized in that... The data analysis module calculates a predictive value for whether the patient under test has pancreatic cancer based on the detection value of the biomarker using the analysis model equation, and determines whether the patient under test has pancreatic cancer based on the predicted value. The judgment criteria are: When the predicted value is less than or equal to a preset threshold, the patient to be tested is determined not to be a pancreatic cancer patient. When the predicted value is greater than the preset threshold, the patient to be tested is determined to be a pancreatic cancer patient; The preset threshold is 0.
473.
7. The system for detecting pancreatic cancer as described in claim 5, characterized in that... The system includes a data detection module, a data input module, and a data output module. The data detection module is used to detect biomarkers in a sample and obtain detection values. The data input module is used to input the detection values of the biomarkers. After the data analysis module analyzes the detection values, the data output interface is used to output the analysis results of whether the patient under test has pancreatic cancer.
8. The application of biomarkers in the preparation of clinical staging diagnostic products for pancreatic cancer, characterized in that... The biomarkers are a combination of SERPINB1, ORM1, RNASE1, DEFA3, PGM5, SPINK5, KRT19, S100B, and CD74, or a combination of SELL, CD74, SERPINB1, RNASE1, KRT19, S100B, PGM5, SPINK5, DEFA3, and ORM1.
Citation Information
Patent Citations
Kit for detecting pancreatic cancer cells in peripheral blood
CN107843731A
System for pancreatic cancer detection and reagent or kit thereof
CN116626297A