Bile protein marker combination for predicting liver metastasis of pancreatic cancer and application thereof
Patent Information
- Application Number
- CN202610762814.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-29
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2046-05-29
AI Technical Summary
胰腺癌细胞及其释放的生物标志物更容易直接渗透或脱落进入胆汁,可以在影像学可见病灶出现之前,反映出肝脏微环境的改变,但手术前预测可切除胰腺癌患者的胆汁标志物仍然匮乏
[0018]1.本申请技术方案提供了新的用于术前预测胰腺癌肝转移的胆汁蛋白标志物组合,在可切除的胰腺导管腺癌患者群体中,该胆汁蛋白标志物组合的表达水平对于早期区分肝转移风险大和肝转移风险小的患者具有高灵敏度、高特异度特点,为术前对于胰腺癌患者精准进行风险分层,筛选出早期肝转移高风险患者提供的新的思路。
Smart Images

Figure CN122283150B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomedical technology, and in particular to a combination of bile protein biomarkers for predicting liver metastasis of pancreatic cancer and their applications. Background Technology
[0002] Pancreatic ductal adenocarcinoma (PDAC) is one of the most malignant gastrointestinal tumors, with its incidence and mortality rates continuing to rise in recent years. The disease has an insidious onset, with early symptoms lacking specificity. The vast majority of patients are diagnosed at a locally advanced or metastatic stage, with a median survival of only 4 months. Despite progressive advancements in treatment through systemic chemotherapy and precision medicine, the overall five-year survival rate remains stagnant at around 13%. Once metastasis occurs, the prognosis deteriorates dramatically, with a five-year survival rate of only 3% for stage IV patients. As the most common histological subtype of pancreatic cancer, pancreatic ductal adenocarcinoma exhibits highly aggressive tumor biology, with the liver being its most common site of distant metastasis.
[0003] For patients with localized lesions who meet the surgical indications, surgical resection remains the only potentially curative treatment. However, the high postoperative recurrence rate severely limits its clinical benefit. Specifically, 60%–70% of patients with initially resectable pancreatic ductal adenocarcinoma will develop distant metastases within two years post-surgery, with the liver being the most common target organ. This postoperative liver metastasis is a key bottleneck restricting patient survival and, in most cases, deprives patients of the opportunity for further surgery.
[0004] In existing technologies, the following methods are commonly used to predict the risk of liver metastasis after surgery in patients with resectable pancreatic cancer: CA19-9 (carbohydrate antigen 19-9) is currently the only widely used circulating biomarker for pancreatic cancer in clinical practice, with an expression rate of 70-80% in patients with pancreatic ductal adenocarcinoma. When CA19-9 > 240 U / mL, the risk of liver metastasis increases threefold. CEA (carcinoembryonic antigen) has lower sensitivity and specificity than CA19-9, and its value when used alone is limited; it is generally used as an auxiliary biomarker in conjunction with CA19-9 for prediction. Enhanced CT can observe lesion enhancement characteristics, vascular invasion, and micrometastases in the liver, and is used to assess liver metastasis. In addition, high-throughput extraction of quantitative features such as texture, shape, and wavelets from images, combined with artificial intelligence algorithms, can be used to construct predictive models to predict metastasis risk.
[0005] However, CEA and CA19-9 have low detection rates in the early stages of tumors, often only significantly increasing after vascular invasion or micrometastasis. Benign diseases such as biliary tract infection, cholestasis, hepatitis, or pancreatitis can also cause a significant increase in CA19-9, easily leading to misdiagnosis. Imaging data has limited ability to detect small, occult metastases early and struggles to capture information about the tumor microenvironment. Current focus is mainly on predicting the risk of liver metastasis in pancreatic cancer patients, and there are currently no biomarkers for predicting liver metastasis preoperatively in patients with resectable pancreatic cancer. The pancreas and biliary system are anatomically closely connected. Pancreatic cancer cells and their released biomarkers can more easily penetrate or detach into the bile, reflecting changes in the liver microenvironment before radiographically visible lesions appear. However, bile biomarkers for predicting resectable pancreatic cancer patients preoperatively remain scarce.
[0006] Therefore, accurate risk stratification before surgery to screen out high-risk patients with early liver metastases is a pressing clinical need that is expected to improve patient survival outcomes. Summary of the Invention
[0007] In view of the shortcomings of the prior art described above, the purpose of this invention is to provide a bile protein biomarker composition for predicting liver metastasis of pancreatic cancer and its application, in order to solve the problems in the prior art.
[0008] To achieve the above and other related objectives, the present invention is obtained through the following technical solution.
[0009] The present invention provides a combination of bile protein biomarkers for predicting liver metastasis of pancreatic cancer, wherein the combination of bile protein biomarkers includes at least MMP9, PRTN3 and S100A12.
[0010] The present invention also provides the application of the combination of bile protein markers or their detectors as described above in the preparation of products for predicting liver metastasis of pancreatic cancer.
[0011] The present invention also provides a product for predicting liver metastasis of pancreatic cancer, the product comprising the combination of bile protein markers or their detectors as described above.
[0012] This invention also provides a system for predicting liver metastasis of pancreatic cancer, comprising the following modules: a data acquisition module, used to acquire expression level data of a combination of target biomarkers in the sample to be tested, wherein the combination of target biomarkers is any of the bile protein biomarker combinations described above for predicting liver metastasis of pancreatic cancer; and a data analysis module, used to calculate the disease risk level data of the corresponding test subject based on the expression level data of the combination of target biomarkers in the sample to be tested using a machine learning algorithm model, and output the prediction result, thereby determining the risk of liver metastasis after pancreatic cancer surgery before surgery.
[0013] This invention also provides a method for predicting liver metastasis of pancreatic cancer, which involves obtaining expression level data of a combination of target biomarkers in a sample to be tested, wherein the combination of target biomarkers is a combination of bile protein biomarkers for predicting liver metastasis of pancreatic cancer as described above; based on the expression level data of the combination of target biomarkers in the sample to be tested, a machine learning algorithm model is used to calculate the disease risk level data of the corresponding test subject, and the prediction result is output to determine the risk of liver metastasis after pancreatic cancer surgery.
[0014] The present invention also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement a method for predicting liver metastasis of pancreatic cancer. The method includes: acquiring expression level data of a combination of target biomarkers in a test sample, wherein the combination of target biomarkers is a combination of bile protein biomarkers for predicting liver metastasis of pancreatic cancer as described above; based on the expression level data of the combination of target biomarkers in the test sample, using a machine learning algorithm model to calculate the disease risk level data of the corresponding test subject, and outputting a prediction result to determine the risk of liver metastasis after pancreatic cancer surgery.
[0015] The present invention also provides an electronic terminal, the electronic terminal including a processor, a memory, an input / output interface, a communication port, and a bus; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the terminal performs the method for predicting liver metastasis of pancreatic cancer as described above in a computer-readable storage medium.
[0016] The present invention also provides a computer program product, comprising a computer program that, when executed by a processor, implements the method for predicting pancreatic cancer liver metastasis as described above in a computer-readable storage medium.
[0017] Beneficial effects:
[0018] 1. The technical solution of this application provides a new combination of bile protein biomarkers for preoperative prediction of liver metastasis in pancreatic cancer. In the patient population with resectable pancreatic ductal adenocarcinoma, the expression level of this combination of bile protein biomarkers has high sensitivity and high specificity in distinguishing between patients with high and low risk of liver metastasis in the early stage. It provides a new approach for accurate risk stratification of pancreatic cancer patients before surgery and screening out patients with high risk of early liver metastasis.
[0019] 2. This application makes full use of the routine preoperative jaundice reduction procedure performed on pancreatic cancer patients, transforming bile, which was previously regarded as biological waste, into a high-value diagnostic sample. This eliminates the need for additional liver biopsy, does not increase the patient's trauma risk or economic burden, and has excellent clinical operability and patient compliance.
[0020] 3. The technical solution of this application breaks through the diagnostic blind spots of conventional imaging, realizes the early identification of occult metastasis, overcomes the dilution effect and non-specific interference of blood indicators; the standardized preprocessing process ensures the biological authenticity of proteomic data; assists in precise clinical stratification, optimizes individualized treatment pathways, and improves the overall survival prognosis of pancreatic cancer patients. Attached Figure Description
[0021] Figure 1 The graph shows the correlation between bile proteins and clinical indicators in patients with resectable pancreatic cancer.
[0022] Figure 2 The figure shows the results of a differential analysis of bile proteins between patients with resectable pancreatic cancer and patients with benign diseases.
[0023] Figure 3 The figure shows the results of functional enrichment analysis of differentially expressed bile proteins between patients with resectable pancreatic cancer and patients with benign diseases.
[0024] Figure 4 This figure shows the difference in bile proteins between resectable pancreatic cancer patients who developed liver metastases after surgery and those who did not.
[0025] Figure 5 The distribution of important proteins related to tumor development in patients with benign diseases, patients without liver metastases after surgery, and patients with liver metastases after surgery are shown, along with GO enrichment analysis.
[0026] Figure 6 The graph shows the results of progressively increasing protein expression in patients with benign disease, patients without liver metastases after surgery, and patients with liver metastases after surgery.
[0027] Figure 7 This shows the expression of core bile proteins associated with resectable pancreatic cancer liver metastases in independent samples.
[0028] Figure 8 The graph shows the ROC curves of LTF, MMP9, PRTN3, and S100A12 in predicting liver metastasis after pancreatic cancer surgery.
[0029] Figure 9 The diagram shown is a schematic block diagram of the electronic terminal of this invention. Detailed Implementation
[0030] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification.
[0031] Before further describing specific embodiments of the present invention, it should be understood that the scope of protection of the present invention is not limited to the specific embodiments described below; it should also be understood that the terminology used in the embodiments of the present invention is for describing specific embodiments and not for limiting the scope of protection of the present invention. Test methods in the following embodiments that do not specify specific conditions are generally performed under conventional conditions or as recommended by the respective manufacturers.
[0032] When numerical ranges are given in the embodiments, it should be understood that, unless otherwise stated in the present invention, both endpoints of each numerical range and any value between the two endpoints may be selected. Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art. In addition to the specific methods, apparatus, and materials used in the embodiments, based on the knowledge of the prior art possessed by one of ordinary skill in the art and the description of this invention, any prior art methods, apparatus, and materials similar to or equivalent to those described, apparatus, and materials in the embodiments of this invention may be used to implement the present invention.
[0033] This application provides a combination of bile protein biomarkers for predicting liver metastasis of pancreatic cancer, the combination of bile protein biomarkers including at least MMP9, PRTN3 and S100A12.
[0034] In some embodiments, the bile protein marker combination further includes LTF.
[0035] Lactoferrin (LTF) in this application is a multifunctional iron-binding glycoprotein. The LTF gene is located on human chromosome 3p21.31 and belongs to the transferrin family. The encoded protein product is present in the secondary granules of neutrophils and various exocrine fluids. It plays an important role in antibacterial and antiviral activity, immune regulation and anti-inflammation, regulation of iron homeostasis and antioxidation.
[0036] Matrix metalloproteinase 9 (MMP9), also known as gelatinase B, is a zinc-dependent endopeptidase. The MMP9 gene is located on human chromosome 20q13.12 and belongs to the matrix metalloproteinase family. In the human body, MMP-9 is mainly produced by inflammatory cells such as neutrophils and macrophages, as well as various tumor cells. MMP9 is a key molecule in tissue remodeling and disease progression, serving as a central hub connecting tissue microenvironment remodeling with inflammation and immune signaling pathways.
[0037] Protease 3 (PRTN3), also known as myeloblastin, is a core member of the neutrophil serine protease family. The PRTN3 gene is located on human chromosome 19p13.3, is approximately 7 kb in length, and is specifically highly expressed in neutrophils and monocytes. Dysregulation of PRTN3 expression and activity is a driving force in many diseases, particularly autoimmune diseases, tumors, and respiratory diseases.
[0038] S100 calcium-binding protein A12 (S100A12) is a pro-inflammatory cytokine in the S100 protein family, primarily expressed by neutrophils. The S100A12 gene is located in the S100 protein gene cluster on human chromosome 1q21.3. S100A12 is a dual-function protein: intracellularly it participates in calcium signaling transduction, while extracellularly it acts as a diamine agonist (DAMP) to initiate and amplify inflammatory responses. Overexpression of S100A12 is a significant driver of various inflammatory diseases, and its levels are closely correlated with disease activity, making it a highly promising diagnostic biomarker and therapeutic target.
[0039] The combination of bile protein biomarkers for predicting liver metastasis in pancreatic cancer described in this application can specifically predict the risk probability of postoperative liver metastasis before surgical resection of pancreatic ductal adenocarcinoma, thereby enabling precise risk stratification of patients before surgery and screening out high-risk patients with early liver metastasis.
[0040] In some embodiments, the pancreatic cancer is pancreatic ductal adenocarcinoma (PDAC).
[0041] This application also provides the use of the combination of bile protein markers or their detectors as described above in the preparation of products for predicting liver metastasis of pancreatic cancer.
[0042] This application also provides a product for predicting liver metastasis of pancreatic cancer, the product comprising the combination of bile protein markers or their detectors as described above.
[0043] It should be noted that the product provided by this invention is mainly aimed at patients diagnosed with surgically resectable pancreatic ductal adenocarcinoma, and is used to identify the risk of pancreatic cancer liver metastasis after surgery in this patient group at an early stage.
[0044] In some embodiments, the detectable includes substances for detecting the expression levels of MMP9, PRTN3, and S100A12.
[0045] In some embodiments, the detectable includes substances for detecting the expression levels of LTF, MMP9, PRTN3, and S100A12.
[0046] In some embodiments, the detectable substance includes one or more of antibodies, membrane strips, chips, probes, or primer pairs.
[0047] The term "antibody" refers to a protective protein produced by the body in response to stimulation by an antigen. In some embodiments, antibodies can be used as a detection reagent to detect gene expression.
[0048] The term "membrane strip" refers to a diagnostic tool that utilizes the principle of specific biomolecular recognition. It involves immobilizing biomolecules such as antigens or antibodies on a membrane, allowing them to specifically bind to the analyte in a sample, and then using visualization or other signal detection methods to qualitatively or quantitatively analyze the target substance in the sample.
[0049] The term "chip" typically refers to a miniature device that integrates biosensors and microfluidics technology. It can perform various operations such as sample preparation, reaction, and detection in biological, chemical, and medical analysis processes at the microscopic level to achieve rapid and accurate detection of disease-related biomarkers.
[0050] The term "probe" refers to any molecule capable of selectively binding to a target biomolecule (e.g., a nucleic acid sequence that hybridizes with the probe). In some embodiments, the probe may be labeled, for example, with a fluorescent group and a quencher group. In some embodiments, the probe may be a Taqman probe with a fluorescent reporter group added to the 5' end and a fluorescent quencher group added to the 3' end.
[0051] The term "primer" refers to a naturally occurring oligonucleotide (e.g., a restriction fragment) or a synthetically produced oligonucleotide that can be used as a starting point for the synthesis of primer extension products. Under appropriate conditions (e.g., buffer, salt, temperature, and pH) and in the presence of nucleotides and reagents for nucleic acid polymerization (e.g., DNA-dependent or RNA-dependent polymerases), the primer extension product is complementary to the nucleic acid strand (template or target sequence). Typically, a primer set will consist of at least two primers, an "upstream primer" and a "downstream primer," which together define the amplicon (the sequence to be amplified using the primers).
[0052] In some embodiments, the product is selected from at least one of reagents, kits, membrane strips, chips, or detection systems.
[0053] In some embodiments, the test sample for the product is bile. Specifically, the test sample is taken from the bile of the object to which the product is applied.
[0054] This application also provides a system for predicting liver metastasis of pancreatic cancer, comprising the following modules: a data acquisition module, used to acquire expression level data of a combination of target biomarkers in the sample to be tested, wherein the combination of target biomarkers is the bile protein biomarker combination for predicting liver metastasis of pancreatic cancer as described above; and a data analysis module, used to calculate the disease risk level data of the corresponding test subject based on the expression level data of the combination of target biomarkers in the sample to be tested using a machine learning algorithm model, and output the prediction result, thereby determining the risk of liver metastasis after pancreatic cancer surgery before surgery.
[0055] Specifically, the disease risk level data is a risk score value of a combination of bile protein biomarkers. This risk score value is calculated by substituting the expression level data of the bile protein biomarkers for each test sample into a machine learning algorithm model. The machine learning algorithm model is obtained by training known samples' bile protein biomarker detection data using any one of the following algorithms: logistic regression, minimum absolute contraction, and selection operator. In the data analysis module, when the risk score value of the bile protein biomarker combination of the test sample is greater than or equal to the cutoff value, the test sample is output as positive; when the risk score value is less than the cutoff value, the test sample is output as negative.
[0056] More specifically, the known samples include bile from patients with benign diseases and bile from patients with resectable pancreatic cancer; the benign diseases include one or more of patients with chronic pancreatitis and patients with pancreatic cystic lesions; the resectable pancreatic cancer patients include patients who developed liver metastases after surgery and patients who did not develop liver metastases after surgery.
[0057] In some embodiments, a positive result indicates that the individual corresponding to the test sample has a high risk of liver metastasis after pancreatic cancer resection surgery, while a negative result indicates that the individual corresponding to the test sample has a relatively low risk of liver metastasis after pancreatic cancer resection surgery.
[0058] In some embodiments, the sample to be tested is selected from the bile of the object being tested.
[0059] In some embodiments, when the protein biomarker combination is MMP9, PRTN3, and S100A12, the cutoff value can be 0.5314 and the risk score can be 0.18541 + 0.12489 × MMP9 + 0.28349 × PRTN3 + 0.37622 × S100A12 based on the minimum absolute shrinkage and selection operator machine learning algorithm.
[0060] In some embodiments, when the protein biomarker combination is LTF, MMP9, PRTN3, and S100A12, the cutoff value can be 0.5314 and the risk score can be 0.18541 + 0.12489 × MMP9 + 0.28349 × PRTN3 + 0.37622 × S100A12 based on the minimum absolute shrinkage and selection operator machine learning algorithm.
[0061] In some embodiments, when the protein biomarker combination is LTF, MMP9, PRTN3, and S100A12, the cutoff value based on the logistic regression machine learning algorithm can be 0.5557, and the risk score is 0.2473002 - 0.1850809 × LTF + 0.2735794 × MMP9 + 0.5193074 × PRTN3 + 0.5405806 × S100A12.
[0062] Wherein LTF refers to the expression level data of LTF in the sample, MMP9 refers to the expression level data of MMP9 in the sample, PRTN3 refers to the expression level data of MMP9 in the sample, and S100A12 refers to the expression level data of MMP9 in the sample.
[0063] Specifically, the logistic regression model, minimum absolute shrinkage, and selection operator model are all constructed and trained using expression data of target biomarker combinations from known samples. The target biomarker combinations are LTF, MMP9, PRTN3, and S100A12.
[0064] Furthermore, the data analysis module includes a computer host, a central processing unit, a network server, and a monitor.
[0065] Furthermore, the data acquisition module for LTF, MMP9, PRTN3, and S100A12 expression levels and the data analysis module are connected via wired or wireless means.
[0066] Furthermore, wireless connection methods can include Wi-Fi, Bluetooth, infrared, etc.; wired connection methods can include landline networks, etc. Using these connection methods greatly facilitates the user's use of the detection system, and at the same time, it leverages the increasingly developed information technology and the increasingly widespread network resources to provide subjects with accurate predictions of their postoperative pancreatic cancer liver metastasis risk levels.
[0067] Furthermore, the data analysis module also includes information provided by the expression level data module based on the bile protein biomarker combination described above, to calculate the risk level data of postoperative pancreatic cancer liver metastasis, and to determine the risk of postoperative pancreatic cancer liver metastasis based on the risk level data.
[0068] It should be noted that the division of the various modules in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. These modules can be implemented entirely in software through processing element calls; they can also be implemented entirely in hardware; or some modules can be implemented by processing element calls to software, while others are implemented in hardware. For example, the database retrieval module can be a separate processing element, or it can be integrated into a chip. Alternatively, it can be stored in memory as program code, and its functions can be called and executed by a processing element. The implementation of other modules is similar. Furthermore, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element mentioned here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed through the integrated logic circuits in the hardware of the processor element or through software instructions.
[0069] For example, these modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more digital signal processors (DSPs), one or more Field Programmable Gate Arrays (FPGAs) or Graphics Processing Units (GPUs). As another example, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together as a system-on-a-chip (SOC).
[0070] This application also provides a method for predicting liver metastasis of pancreatic cancer, the method comprising: acquiring expression level data of a combination of target biomarkers in a sample to be tested, wherein the combination of target biomarkers is a combination of bile protein biomarkers for predicting liver metastasis of pancreatic cancer as described above; based on the expression level data of the combination of target biomarkers in the sample to be tested, using a machine learning algorithm model to calculate the disease risk level data of the corresponding test subject, and outputting the prediction result to determine the risk of liver metastasis after pancreatic cancer surgery.
[0071] This application also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement a method for predicting liver metastasis of pancreatic cancer. The method includes: acquiring expression level data of a combination of target biomarkers in a test sample, wherein the combination of target biomarkers is a combination of bile protein biomarkers for predicting liver metastasis of pancreatic cancer as described above; calculating the disease risk level data of the corresponding test subject using a machine learning algorithm model based on the expression level data of the combination of target biomarkers in the test sample; and outputting a prediction result to determine the risk of liver metastasis after pancreatic cancer surgery.
[0072] This application also provides an electronic terminal, such as... Figure 9 As shown, the electronic terminal includes a processor 11, a memory 12, an input / output interface 13, a communication port 14, and a bus 15; the memory 12 is used to store a computer program, and the processor 11 is used to execute the computer program stored in the memory, so that the terminal performs the method for predicting pancreatic cancer liver metastasis as described above in a computer-readable storage medium.
[0073] Processor 11 can execute computational instructions (program code) and perform the functions of the detection system described in this application. Computational instructions may include programs, objects, components, data structures, procedures, modules, and functions (functions refer to the specific functions described in this application). For example, processor 11 can process instructions for detecting pancreatic cancer liver metastases in a detection system for predicting pancreatic cancer liver metastases. In some embodiments, processor 11 may include a microcontroller, microprocessor, reduced instruction set computer (RISC), application-specific integrated circuit (ASIC), application-specific instruction set processor (ASIP), central processing unit (CPU), graphics processing unit (GPU), physical processing unit (PPU), microcontroller unit, digital signal processor (DSP), field-programmable gate array (FPGA), advanced RISC machine (ARM), programmable logic device, and any circuit and processor capable of performing one or more functions, or any combination thereof. This is for illustrative purposes only. Figure 9 Only one processor is described in this application, but it should be noted that this application may include multiple processors.
[0074] The memory 12 can store data / information obtained from any component of the pancreatic cancer liver metastasis prediction detection system. In some embodiments, the memory 12 may include mass storage, removable storage, volatile read-write storage, and read-only storage (ROM), or any combination thereof. Exemplary mass storage may include disks, optical disks, and solid-state drives. Removable storage may include flash drives, floppy disks, optical disks, memory cards, USB flash drives, compact disks, and portable hard drives. Volatile read-write storage may include random access memory (RAM). RAM may include dynamic RAM (DRAM), double-rate synchronous dynamic RAM (DDRSDRAM), static RAM (SRAM), thyristor RAM (T-RAM), and zero-capacitance RAM (Z-RAM), etc. ROM may include mask ROM (MROM), programmable ROM (PROM), erasable programmable ROM (PEROM), electrically erasable programmable ROM (EEPROM), optical disc ROM (CD-ROM), and digital universal disk ROM, etc.
[0075] The input / output interface 13 can be used to input or output signals, data, or information. In some embodiments, the input / output interface 13 can be used to enable interaction between a user (e.g., the person corresponding to the test sample, the user of the preoperative pancreatic cancer liver metastasis detection system, etc.) and the processor. In some embodiments, the user can input characteristic information of the person corresponding to the test sample through the input / output interface 13. In some embodiments, the input / output interface 13 may include an input device and an output device. Exemplary input devices may include a keyboard, mouse, touch screen, and microphone, or any combination thereof. Exemplary output devices may include a display device, speaker, printer, projector, or any combination thereof. Exemplary display devices may include a liquid crystal display (LCD), a light-emitting diode (LED) based display, a flat panel display, a curved display, a television device, a cathode ray tube (CRT), or any combination thereof.
[0076] Communication port 14 can be connected to a network for data communication. The connection can be wired, wireless, or a combination of both. Wired connections can include cables, fiber optic cables, or telephone lines, or any combination thereof. Wireless connections can include Bluetooth, WiFi, WiMax, WLAN, ZigBee, mobile networks (e.g., 3G, 4G, or 5G), or any combination thereof. In some embodiments, communication port 14 can be a standardized port, such as RS232 or RS485. In some embodiments, communication port 14 can be a specially designed port.
[0077] Bus 15 can connect processor 11, memory 12, and other components. Memory 12 communicates with processor 11 via bus 15. Processor 11 reads data from memory 12, executes programs, and may write results back to memory 12 via bus 15. Input / output interface 13 exchanges data with processor 11 and memory 12 via bus 15. For example, when a user enters data on the keyboard, the keyboard sends the data to the computer's input / output interface 13, and then it is transmitted to processor 11 or memory 12 via data bus 15. Communication port 14 exchanges data with processor 11 and memory 12 via bus 15. For example, when the computer receives network data through the Ethernet port, this data is first transmitted to the network interface card (NIC), and then transmitted to processor 11 or memory 12 via the network bus.
[0078] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the method for predicting pancreatic cancer liver metastasis as described above in a computer-readable storage medium.
[0079] The present invention can be implemented as a computer program product including a computer-readable storage medium and computer program code encoded on the medium for performing the above-described techniques.
[0080] The invention can be implemented on any electronic device, such as a handheld computer, a personal digital assistant (PDA), a personal computer, a public information kiosk, a cellular phone, etc.
[0081] Certain aspects of this application include the process steps and instructions described herein in algorithmic form. It should be noted that the process steps and instructions of this application may be embodied in software, firmware, or hardware, and when embodied in software, they may be downloaded and reside on various platforms used by different operating systems, and may be operated from said platforms.
[0082] This application also relates to apparatus for performing the operations described herein. This apparatus may be specifically constructed for the desired purpose, or it may comprise a general-purpose computer selectively activated or reconfigured by a computer program stored in a computer. This computer program may be stored in a computer-readable storage medium, such as, but not limited to: any type of disk, including floppy disks, optical disks, CD-ROMs, magneto-optical disks; read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic cards or optical cards, application-specific integrated circuits (ASICs), or any type of media suitable for storing electronic instructions, and may be coupled to a computer system bus. Furthermore, the computer referred to herein may comprise a single processor, or may be an architecture employing a multiprocessor design to seek increased computing power.
[0083] Unless otherwise stated, the experimental methods, detection methods, and preparation methods disclosed in this invention all employ conventional techniques in molecular biology, biochemistry, chromatin structure and analysis, analytical chemistry, cell culture, recombinant DNA technology, proteomics analysis, and related fields.
[0084] Example 1: Acquisition of Experimental Samples
[0085] 1. Standardized Procedure for Bile Sample Collection and Preservation
[0086] Resectable pancreatic cancer is defined as a tumor that does not invade the celiac trunk, superior mesenteric artery, or common hepatic artery. The tumor does not invade the superior mesenteric vein or portal vein, or if it does, the invasion is less than 180 degrees and the vein outline is regular. Subjects are confirmed to meet the clinical diagnosis of resectable pancreatic cancer and have been included in the preoperative jaundice reduction protocol (preoperative biliary drainage).
[0087] Prepare sterile sampling equipment, including a 10-50 ml sterile syringe or drainage bag, 1.5-2.0 ml low-absorption screw-cap cryovials, EDTA-free protease inhibitors, a cryogenic dual-label system, and a biosafety compliant transport container. Collect the sample before injecting contrast agent or saline flushing solution; if contrast agent or flushing solution has been pre-injected according to clinical procedures, discard the initial 3-5 ml of the mixture and wait until the drainage fluid returns to a pure yellow bile before collecting. Use a sterile syringe connected to the side hole of the drainage tube to slowly aspirate the bile at a uniform rate, avoiding vigorous aspiration that could cause bile duct epithelial cell breakage or erythrocyte hemolysis. Perform macroscopic quality control simultaneously during collection: a qualified sample should be golden yellow or dark green, clear or slightly turbid, and odorless. Collect the qualified bile in the same sterile container and gently invert 3-5 times to homogenize.
[0088] Dispense the samples. Dispensing must be performed in a sterile environment or clean bench to strictly avoid exogenous contamination. Immediately after dispensing, immerse the cryovials in liquid nitrogen for 15 to 30 seconds, then quickly transfer them to an -80°C ultra-low temperature freezer for long-term storage. All samples must be used only once; each sample is allowed only one freeze-thaw cycle, and repeated freeze-thaw cycles are strictly prohibited. Establish a complete sample identification and information traceability system. Use a dual-label system resistant to -196°C and organic solvents. Labels are simultaneously affixed to both the cryovial body and cap. Sample information is entered into the electronic sample management system, with records including at least the collection time, dispensing time, freezing time, operator's identity, reviewer's signature, and sample characteristics remarks, ensuring full traceability.
[0089] 2. Basic information of the subject sample in this application
[0090] In compliance with ethical requirements, this application collected and preserved a total of 96 samples using the methods described above, including bile samples from 40 patients with benign diseases (chronic pancreatitis, pancreatic cystic lesions) and 56 bile samples from patients with resectable pancreatic cancer. Among the 56 bile samples from patients with resectable pancreatic cancer, 26 patients did not develop liver metastasis within one year postoperatively, while 30 patients developed liver metastasis within six months postoperatively. The basic information of each patient is shown in Table 1.
[0091] Table 1 Patient Clinical Information
[0092]
[0093] Example 2 Screening of biomarkers
[0094] 1. High-throughput detection of bile proteins – DIA mass spectrometry data acquisition and evaluation
[0095] Step 1: Proteins were extracted from the bile of the 96 samples using SDT lysis buffer (containing 4% sodium dodecyl sulfate, 100 mM tris(hydroxymethyl)aminomethane-hydrochloric acid, pH 7.6). Protein concentrations were determined using a quinoline carboxylic acid kit (BCA method). After separation by sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE), the proteins were stained with Coomassie Brilliant Blue R-250 for visualization analysis. Equal volumes of proteins from each sample were mixed to prepare a pooled quality control (QC) sample.
[0096] Step 2: Filter-assisted proteomics (FASP) was used to prepare all samples via trypsin digestion. The trypsin-digested peptides were desalted using a C18 solid-phase extraction column, freeze-dried, and then reconstituted in 0.1% formic acid solution. The absorbance at 280 nm (OD) was measured. 280 To determine the peptide concentration, iRT standard peptides were added to each sample before analysis.
[0097] Step 3: Data Independent Acquisition (DIA) was performed on a Thermo Fisher Scientific Orbitrap Astral mass spectrometer coupled with a Vanquish Neo nano-liquid chromatography system in positive ion mode. The full-scan mass spectrometry (MS) range was set to 380–980 m / z with a resolution of 240,000 (at 200 m / z). In DIA mode, 299 isolation windows were set, each 2 m / z wide, with a high-energy collisional dissociation (HCD) collision energy of 25 eV.
[0098] Secondary mass spectrometry (MS) 2The normalized automatic gain control (AGC) target value was set to 500%, and the maximum injection time was set to 3 ms. Raw DIA data was processed using a spectral library-based search engine. Trypsin was specified as the digestive protease, allowing a maximum of one missed cleavage site. The fixed modification was carbamoyl methylation of cysteine; variable modifications included methionine oxidation and N-terminal acetylation of the protein. Peptide and protein identification results were filtered for a false detection rate (FDR) <1%.
[0099] Step 4: Assess data quality. Monitor the coefficient of variation (CV) of protein quantification in the QC samples to assess data quality, and remove protein features with a CV > 30%. Then, normalize the protein quantification values using median normalization, and complete missing values using the K-nearest neighbor (KNN) algorithm. Perform a log2 transformation on the normalized protein abundance to make the data distribution approximate normal. Finally, establish a bile protein database for the aforementioned 96 patient samples.
[0100] 2. Correlation analysis between protein expression levels and PDAC clinical indicators in patients with resectable pancreatic cancer
[0101] The applicant analyzed the correlation between protein expression levels and PDAC clinical indicators in 56 patients with resectable pancreatic cancer from the aforementioned bile protein database. Specific results are shown below. Figure 1 .in Figure 1 The redder the circle, the stronger the positive correlation between the protein's expression level and the corresponding PDAC clinical indicator. Figure 1 The bluer the circle, the stronger the negative correlation between the protein's expression level and the corresponding PDAC clinical indicator. The specific experimental method involved using Spearman rank correlation analysis to perform nonparametric correlation tests on each indicator. The correlation coefficient was calculated using the variable's rank rather than its original value to analyze the direction and strength of the monotonic association between the two variables.
[0102] Depend on Figure 1 The expression levels of epidermal growth factor-like domain protein 1 (EFEMP1), dipeptidyl peptidase 7 (DDP7), and membrane metalloendopeptidase (MME) were positively correlated with the levels of gamma-glutamyl transferase (GGT) and alanine aminotransferase (ALT), but negatively correlated with the level of cholinesterase (ChE). This suggests that EFEMP1, DDP7, and MME may be related to blood test indicators in pancreatic cancer patients.
[0103] 3. Screening for differentially expressed proteins in the bile protein database.
[0104] The differentially expressed proteins in the bile protein database of the 96 patient samples obtained above were statistically analyzed using R statistical software (version 4.4.3).
[0105] In this study, a linear model was constructed using the limma software package to analyze the differences in protein expression levels. The Benjamini-Hochberg (BH) method was used to perform multiple hypothesis testing to correct the original p-values, controlling for the false discovery rate. The adjusted p-value (or q-value) was used as the primary criterion for determining the statistical significance of differentially expressed proteins. A threshold of <0.05 was used to screen for statistically significant differentially expressed proteins, and then volcano plots, heatmaps, and statistical graphs of the most significantly upregulated and downregulated proteins in different groups were generated. Detailed results are shown below. Figure 2 .in, Figure 2 In the figure, 'a' represents a volcano plot of differentially expressed proteins between the benign disease group and the resectable pancreatic cancer group. Figure 2 In the image, b represents a heatmap showing the differentially expressed proteins between the benign disease group and the resectable pancreatic cancer group. Figure 2 In the figure, c represents the statistical graph of the most significantly upregulated and downregulated proteins in different groups (benign disease group and resectable pancreatic cancer group) among the 96 samples.
[0106] Depend on Figure 2 It can be seen that the volcano map ( Figure 2 a) and heatmap ( Figure 2 (b) in the diagram visually presents the differences between the two groups (benign disease group and resectable pancreatic cancer group). Meanwhile, [the following is a continuation of the previous sentence]... Figure 2 As shown in c, compared with the benign disease group, the five proteins with the most significant upregulation in the resectable pancreatic cancer group were: lipid carrier protein 2 (LCN2), C-reactive protein (CRP), interstitial α-trypsin inhibitor heavy chain H3 (ITIH3), carcinoembryonic antigen-associated cell adhesion molecule 6 (CEACAM6), and stem cell antigen 1 (PROM1); the five proteins with the most significant downregulation were: cysteine protease inhibitor B (CSTB), sodium-potassium pump α1 subunit (ATP1A1), sodium-potassium pump β1 subunit (ATP1B1), macrophage migration inhibitory factor (MIF), and glutathione S-transferase P1 (GSTP1).
[0107] This indicates that the expression levels of LCN2, CRP, ITIH3, CEACAM6, PROM1, CSTB, ATP1A1, ATP1B1, MIF, and GSTP1 proteins differ in patients with resectable pancreatic cancer compared to patients with benign diseases.
[0108] 4. Functional enrichment analysis
[0109] The applicant performed functional enrichment analysis on differentially expressed proteins in the bile protein database. The specific steps were as follows: Gene Set Enrichment Analysis (GSEA) was used to rank all proteins according to their fold change in expression, assess the enrichment of a predefined functional gene set, and set a false discovery rate (FDR) of less than 0.05 as the significance criterion. Gene Ontology (GO) biological process terms and KEGG metabolic pathways significantly enriched in the bile protein database were identified. A corrected p-value of less than 0.05 was set as the significant enrichment criterion. Detailed results are shown below. Figure 3 ,in Figure 3 Figure 3a shows the KEGG enrichment analysis diagram, and Figure 3b shows the GO biological pathway results.
[0110] Depend on Figure 3 As shown in 'a', it reveals the density distribution characteristics of the fold change (log2) of protein expression in each core pathway among different groups. Among them, compared with the benign disease group, the proteins involved in complement and coagulation cascade reactions in the resectable pancreatic cancer group showed a significant upregulation trend, and their density peaks shifted to the right; conversely, the protein expression distribution of those involved in carbon metabolism, material metabolism pathways, and chemical carcinogenesis-reactive oxygen species pathway was relatively balanced, and showed a slight downregulation trend.
[0111] Depend on Figure 3 As shown in b, compared with the benign disease group, the downregulated proteins in the resectable pancreatic cancer group were significantly enriched in energy metabolism-related pathways, including aerobic respiration, cellular respiration, and purine ribonucleotide metabolism; while the upregulated proteins were closely related to immune defense responses, such as adaptive immune responses, humoral immune responses, and defense responses against pathogens such as bacteria.
[0112] 5. Analysis of bile protein characteristics specific to patients with early liver metastasis after pancreatic cancer surgery
[0113] The applicant further constructed a linear model using the limma software package to identify bile protein characteristics specific to patients with early liver metastases after pancreatic cancer surgery and plotted differential protein expression scatter plots. See details below. Figure 4 'a' in 'a'.
[0114] Depend on Figure 4 As shown in 'a', there was no significant difference in protein expression among different patients in the pancreatic ductal adenocarcinoma group. Compared with patients who did not develop liver metastasis in the early stage, only 18 significantly upregulated proteins and 21 significantly downregulated proteins were screened in patients who developed liver metastasis in the early stage. Figure 4 (The 'a' in the text is not fully identified).
[0115] Simultaneously, based on the differentially expressed proteins identified above, particularly those significantly upregulated in patients with early liver metastases, the applicant constructed a protein-protein interaction (PPI) network. The STRING database was used to obtain the protein-protein interaction relationships, and network visualization tools such as Cytoscape were used for network construction and analysis. The connectivity (degree) of each node in the network was calculated, and key hub proteins located at the network center were identified. Specific results are detailed below. Figure 4 b in the text.
[0116] Depend on Figure 4 As can be seen from b, actin 4 (ACTN4), lactoferrin (LTF), matrix metalloproteinase 9 (MMP9), and S100 calcium-binding protein A12 (S100A12) are core hub proteins that interact with multiple other proteins.
[0117] This suggests that early liver metastasis compared to patients without liver metastasis may be related to the expression of ACTN4, LTF, MMP9, and S100A12 proteins in the patient's body.
[0118] 6. Distribution and GO enrichment analysis of key tumor-related proteins in the sample.
[0119] The applicant also analyzed the distribution of important proteins related to tumor development in the benign disease group and the resectable pancreatic cancer group (including the resectable pancreatic cancer without liver metastasis group and the resectable pancreatic cancer with liver metastasis group), and performed GO enrichment analysis on them. Specifically, over-representation analysis (ORA) was performed using the clusterProfiler software package. See the detailed results below. Figure 5 . Figure 5 The value of 'a' in the figure represents the distribution of important tumor-related proteins in patients with benign diseases, patients without liver metastases after surgery, and patients with liver metastases after surgery. Figure 5 b in the figure represents the GO enrichment analysis results of the differentially expressed proteins mentioned above.
[0120] Depend on Figure 5 As shown in a, compared to the benign disease group, regardless of whether liver metastasis occurred, the key functional proteins involved in angiogenesis, epithelial-mesenchymal transition (EMT), tumorigenesis, immune regulation, and extracellular matrix (ECM) remodeling processes in bile samples from the resectable pancreatic cancer group showed a moderate upregulation trend. Among them, matrix metalloproteinase 9 (MMP9) and vimentin (VIM) showed a more significant upregulation trend in patients with liver metastasis.
[0121] As shown in Figure 5b, the overexpression analysis of the above-upregulated genes revealed that the significantly enriched functional items included collagen catabolism, extracellular domain proteolysis of membrane proteins, and leukocyte migration.
[0122] 7. Identification of the core bile protein composition based on expression trends
[0123] In light of the biological characteristics of pancreatic cancer liver metastases, the applicant focused on proteins exhibiting progressively changing expression across three groups: benign disease, resectable pancreatic cancer without liver metastases, and resectable pancreatic cancer with liver metastases. The applicant compared the resectable pancreatic cancer group (PDAC) with the benign disease group (Benign), and the resectable pancreatic cancer with liver metastases (Metastasis) with the resectable pancreatic cancer without liver metastases (Disease-free), identifying proteins showing consistent expression trends in both groups. The focus was on identifying protein biomarkers that showed progressively upregulation or downregulation from benign disease to pancreatic cancer and then to liver metastases. Specific results are detailed below. Figure 6 .
[0124] Depend on Figure 6 It was found that the expression levels of lactoferrin (LTF), matrix metalloproteinase 9 (MMP9), proteinase 3 (PRTN3), and S100 calcium-binding protein A12 (S100A12) showed a gradual increasing trend in the three groups (benign disease group, resectable pancreatic cancer without liver metastasis group, and resectable pancreatic cancer with liver metastasis group).
[0125] In summary, the applicant identified four proteins that are most strongly associated with the progression of PDAC: lactoferrin (LTF), matrix metalloproteinase 9 (MMP9), protease 3 (PRTN3), and S100 calcium-binding protein A12 (S100A12), and used them together as a core biomarker composition for constructing a liver metastasis prediction model.
[0126] Example 3 Validation of Biomarker Compositions
[0127] The applicant validated the expression abundance of selected lactoferrin (LTF), matrix metalloproteinase 9 (MMP9), proteinase 3 (PRTN3), and S100 calcium-binding protein A12 (S100A12) proteins in another independent internal dataset (18 samples in total, including 10 samples from patients without liver metastasis after pancreatic cancer surgery and 8 samples from patients with liver metastasis after pancreatic cancer surgery). Specific results can be found in [link to relevant documentation]. Figure 7 .
[0128] Depend on Figure 7It can be seen that the differential expression trend of the four proteins is consistent with the previous research results, that is, compared with the group without liver metastasis after pancreatic cancer surgery, the expression of LTF, MMP9, PRTN3 and S100A12 proteins in the liver metastasis group were all upregulated.
[0129] Example 4: Predictive Model Construction and Predictive Performance Testing
[0130] The applicant used multiple machine learning algorithms to train the model and validated the model performance on the dataset. The training set consisted of 26 patients who did not develop liver metastases after pancreatic cancer surgery and 30 patients who did develop liver metastases after pancreatic cancer surgery. The additional independent validation set consisted of 10 patients who did not develop liver metastases after pancreatic cancer surgery and 8 patients who did develop liver metastases after pancreatic cancer surgery (not included in the training set data or in the aforementioned 96 samples).
[0131] The specific methods are as follows: 1) Data preprocessing: First, the abundance of the protein biomarkers is standardized and processed using log2. 2) Model construction: Based on the characteristics of binary classification prediction tasks and the characteristics of small-sample high-dimensional data, three machine learning algorithms with complementary advantages are selected for comparison: logistic regression, support vector machine (SVM), and least absolute shrinkage and selection operator (LASSO). The parameters of the three machine learning algorithms are set as follows: logistic regression model, using glmnet R package, with family = "binomial"; LASSO, using glmnet R package, with regularization parameter λ optimized through cross-validation; SVM, using e1071 R package, with radial basis function as the kernel function. 3) Model training and tuning: Then, 10-fold stratified cross-validation is used to evaluate the model performance, ensuring that the proportion of liver metastasis samples in each fold is consistent with the overall proportion. Within each fold, inner-layer cross-validation is further used, with the area under the curve of the inner-layer validation set as the optimization objective, and the hyperparameter combination that maximizes AUC is selected. 4) Formula determination: Based on the training set data and the maximized hyperparameter combination selected by 10-fold hierarchical cross-validation, the final model is retrained using logistic regression (glmnet package, family="binomial"), LASSO (glmnet package, λ cross-validation) and radial basis kernel SVM (e1071 package) to obtain the risk score calculation formula corresponding to each algorithm.
[0132] Subsequently, the risk score for each sample under different machine learning algorithms was calculated using the above model. The subject operating curve was plotted using the pROC package in R language with sensitivity (true positive rate) as the ordinate and 1-specificity (false positive rate) as the abscissa. The optimal cutoff value was determined by the Youden index.
[0133] Specifically, the Youden index = sensitivity - (1 - specificity). A corresponding Youden index can be calculated for each point on the ROC curve. The maximum value of the Youden index is taken as the optimal cutoff value, thereby achieving a balance between sensitivity and specificity.
[0134] The positivity or positivity of clinical samples is determined by using cutoff values from different models and risk scores for each sample. A sample is considered positive (i.e., a high-risk group for liver metastasis after pancreatic cancer resection) when its risk score is greater than the cutoff value of the corresponding model, and negative (i.e., a low-risk group for liver metastasis after pancreatic cancer resection) when its risk score is less than the cutoff value of the corresponding model.
[0135] For details, please see [link / details]. Figure 8 and Table 2, in which Figure 8 In the figure, 'a' represents the ROC curve (Receiver Operating Characteristic Curve) of LTF, MMP9, PRTN3, and S100A12 in the training set for jointly predicting liver metastasis after pancreatic cancer surgery. Figure 8 Table 2 shows the ROC curves of LTF, MMP9, PRTN3, and S100A12 in the validation set for predicting postoperative liver metastasis in pancreatic cancer. Table 2 shows the area under the curve (AUC) of the LTF, MMP9, PRTN3, and S100A12 protein marker combination for predicting postoperative liver metastasis.
[0136] Table 2. Results of area under the curve (AUC) for predicting postoperative liver metastasis using protein biomarker combinations.
[0137]
[0138] Note: The cutoff values for sensitivity and specificity in this table are the maximum values of the Youden index.
[0139] Depend on Figure 8 As shown in Table 2, the Support Vector Machine (SVM) model performs best in the training set, with an area under the curve (AUC) of 0.83, while the logistic regression model and the Least Absolute Shrinkage and Selection Operator (LASSO) model are next (both with an AUC of 0.77). Figure 8 As shown in b, on the validation set, the LASSO model performs best, followed by the logistic regression model, and the SVM model performs worst. Furthermore, comprehensive analysis of both the training and validation sets reveals that the LASSO model has the best overall performance.
[0140] Therefore, the risk score established using the logistic regression model is: 0.2473002 - 0.1850809 × LTF + 0.2735794 × MMP9 + 0.5193074 × PRTN3 + 0.5405806 × S100A12. The receiver operating procedure (ROC) curve for the composition was obtained using the R language ROC software package. The maximum Youden index was calculated, and the corresponding cutoff value was obtained. The cutoff value for this composition is 0.5557. When the risk score of the composition in a clinical sample is greater than or equal to 0.556, it is considered positive; otherwise, it is considered negative.
[0141] In the LASSO model, the LTF coefficient is 0 during feature selection. The risk score established using the LASSO model is: 0.18541 + 0.12489 × MMP9 + 0.28349 × PRTN3 + 0.37622 × S100A12. The receiver operating procedure (ROC) curve for the composition was obtained using the R language ROC software package. The maximum Youden index was calculated, and the corresponding cutoff value was obtained. The cutoff value for this composition was 0.5314. A clinical sample with a composition risk score greater than or equal to 0.5314 was considered positive, and vice versa.
[0142] In summary, the detection of protein expression levels of the combination of LTF, MMP9, PRTN3, and S100A12 to predict postoperative liver metastasis has good sensitivity and specificity. This biomarker combination has good diagnostic efficacy for postoperative liver metastasis of pancreatic ductal adenocarcinoma, and the LASSO model has the best overall performance.
[0143] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.
Claims
1. The application of a combination of bile protein markers or their detection substances in the preparation of products for predicting liver metastasis of pancreatic cancer; The combination of bile protein markers is selected from any one of the following combinations: A combination of MMP9, PRTN3 and S100A12; A combination of MMP9, PRTN3, S100A12, and LTF; The pancreatic cancer mentioned is pancreatic ductal adenocarcinoma.
2. A system for predicting liver metastasis of pancreatic cancer, characterized in that, Includes the following modules: The data acquisition module is used to acquire the expression level data of the target biomarker combination of the sample to be tested. The target biomarker combination is selected from any one of the following combinations: the combination of MMP9, PRTN3 and S100A12; the combination of MMP9, PRTN3, S100A12 and LTF. The data analysis module is used to calculate the disease risk level data of the corresponding test subjects based on the expression level data of the target biomarker combination of the test sample and the machine learning algorithm model, and output the prediction results.
3. The system according to claim 2, characterized in that, The disease risk level data is a risk score value of a combination of bile protein biomarkers, which is calculated by substituting the expression level data of bile protein biomarkers of each test sample into a machine learning algorithm model; And / or, the machine learning algorithm model is obtained by training bile protein biomarker detection data of known samples via any one of the following algorithms: logistic regression model, minimum absolute shrinkage, and selection operator algorithm; And / or, in the data analysis module, when the risk score value of the combination of bile protein markers in the test sample is greater than or equal to the cutoff value, the result of the test sample being positive is output. When the risk score of the combination of bile protein markers in the test sample is less than the cutoff value, the result of the test sample being negative is output. And / or, the sample to be tested is selected from the bile of the object being tested; And / or, when the protein biomarker combination is MMP9, PRTN3 and S100A12, the cutoff value based on the minimum absolute shrinkage and selection operator machine learning algorithm is 0.5314, and the risk score is 0.18541+0.12489×MMP9+0.28349×PRTN3+0.37622×S100A12; And / or, when the protein biomarker combination is LTF, MMP9, PRTN3 and S100A12, the cutoff value based on the minimum absolute shrinkage and selection operator machine learning algorithm is 0.5314, and the risk score is 0.18541+0.12489×MMP9+0.28349×PRTN3+0.37622×S100A12; And / or, when the protein biomarker combination is LTF, MMP9, PRTN3 and S100A12, the cutoff value based on the logistic regression machine learning algorithm is 0.5557, and the risk score is 0.2473002 - 0.1850809×LTF + 0.2735794×MMP9 + 0.5193074×PRTN3 + 0.5405806×S100A12.
4. A computer-readable storage medium, characterized in that, The storage medium stores computer instructions, which, when executed by a processor, implement a method for predicting liver metastasis of pancreatic cancer. The method includes: acquiring expression level data of a combination of target biomarkers in a sample to be tested, wherein the combination of target biomarkers is selected from any one of the following combinations: a combination of MMP9, PRTN3, and S100A12; or a combination of MMP9, PRTN3, S100A12, and LTF. Based on the expression level data of target biomarkers in the test sample, the disease risk level data of the corresponding test subjects are calculated using a machine learning algorithm model, and the prediction results are output.
5. The computer-readable storage medium according to claim 4, characterized in that, The disease risk level data is a risk score value of a combination of bile protein biomarkers, which is calculated by substituting the expression level data of bile protein biomarkers of each test sample into a machine learning algorithm model; The machine learning algorithm model is obtained by training known samples of bile protein biomarker detection data through any one of the following algorithms: logistic regression model, minimum absolute shrinkage, and selection operator algorithm. When the risk score of the combination of bile protein markers in the test sample is greater than or equal to the cutoff value, the test sample is output as positive. When the risk score of the combination of bile protein markers in the test sample is less than the cutoff value, the result of the test sample being negative is output. And / or, the sample to be tested is selected from the bile of the object being tested; And / or, when the protein biomarker combination is MMP9, PRTN3 and S100A12, the cutoff value based on the minimum absolute shrinkage and selection operator machine learning algorithm is 0.5314, and the risk score is 0.18541+0.12489×MMP9+0.28349×PRTN3+0.37622×S100A12; And / or, when the protein biomarker combination is LTF, MMP9, PRTN3 and S100A12, the cutoff value based on the minimum absolute shrinkage and selection operator machine learning algorithm is 0.5314, and the risk score is 0.18541+0.12489×MMP9+0.28349×PRTN3+0.37622×S100A12; And / or, when the protein biomarker combination is LTF, MMP9, PRTN3 and S100A12, the cutoff value based on the logistic regression machine learning algorithm is 0.5557, and the risk score is 0.2473002 - 0.1850809×LTF + 0.2735794×MMP9 + 0.5193074×PRTN3 + 0.5405806×S100A12.
6. An electronic terminal, the electronic terminal comprising a processor, a memory, an input / output interface, a communication port, and a bus; the memory for storing a computer program, the processor for executing the computer program stored in the memory, such that the terminal performs the method for predicting pancreatic cancer liver metastasis in a computer-readable storage medium as described in claim 4.
Citation Information
Patent Citations
Molecular typing of diffuse gastric cancer and protein markers for typing and screening method and application thereof
CN108445097A
Diagnosis of endotype and / or severity of sepsis
CN117751195A