Application of exosome miRNA marker in pancreatic cancer diagnosis

By screening and combining exosomal miRNA biomarkers, and integrating statistical and machine learning methods, a high-performance pancreatic cancer diagnostic model was constructed. This model addresses the shortcomings of existing diagnostic methods in terms of sensitivity and specificity, enabling early and accurate diagnosis and improving patients' quality of life.

CN121780692APending Publication Date: 2026-04-03THE FIRST PEOPLES HOSPITAL OF CHANGZHOU +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-26
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing methods for diagnosing pancreatic cancer, such as CT scans, endoscopic ultrasound-guided fine-needle aspiration biopsy, and blood tumor markers, have limited sensitivity and specificity, making early and accurate diagnosis impossible. Furthermore, existing studies on exosomal miRNA combinations lack large-sample validation, hindering accurate assessment of their clinical diagnostic value.

Method used

By screening and combining multiple exosomal miRNA biomarkers, and combining statistical and machine learning methods, a diagnostic model was constructed for the differential diagnosis of pancreatic cancer. Next-generation sequencing technology was used to detect the expression level of exosomal miRNAs in plasma, and a high-performance diagnostic model was established.

Benefits of technology

It improves the sensitivity and specificity of pancreatic cancer diagnosis, reduces overtreatment, improves patients' quality of life, reflects the complexities of real-world clinical samples, and is superior to existing imaging and hematological tumor marker methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121780692A_ABST
    Figure CN121780692A_ABST
Patent Text Reader

Abstract

The invention discloses an application of an exosome miRNA (micro Ribonucleic Acid) marker in pancreatic cancer diagnosis. The exosome miRNA marker is a combination of miR-574 to 5p, miR-7704, miR-106b to 3p, miR-877 to 5p, let-7a to 3p, miR-320b, miR-142 to 3p, miR-378i, miR-452 to 5p, miR-140 to 3p, miR-95 to 3p, miR-219a to 1 to 3p, miR-204 to 5p, miR-4429, miR-1246, miR-148b to 3p and miR-150 to 5p, and is a combination of other groups of different exosome miRNA markers. The blood exosome miRNA is used as a biomarker, a high-efficiency diagnosis model is established through statistics and a machine learning method and is applied to differential diagnosis of pancreatic cancer, and a verification queue verifies that the diagnosis model has good prediction efficiency in pancreatic cancer diagnosis; the sensitivity and the specificity of the kit are superior to those of existing clinical imaging, hematologic tumor markers, pathological detection and other methods, the phenomenon of excessive medical treatment is reduced, and the life quality of pancreatic cancer patients is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biological detection technology, specifically relating to the application of exosomal miRNA markers in the diagnosis of pancreatic cancer. Background Technology

[0002] Pancreatic cancer is a malignant tumor of the digestive system with high incidence and mortality rates. According to global cancer statistics, it is estimated that more than 400,000 people are newly diagnosed with pancreatic cancer each year, and this number is increasing annually, making it one of the cancers that seriously threaten human health. Pancreatic cancer is a type of malignant tumor that mainly originates from the exocrine pancreatic ducts or acinar cells, with pancreatic ductal adenocarcinoma (PDAC) accounting for about 80%. Due to the pancreas's unique anatomical location, early onset is often insidious and asymptomatic. 80%-85% of patients are diagnosed at a locally advanced stage, and only 15%-20% are diagnosed in the early stages and can be treated surgically. Furthermore, due to the tumor's aggressiveness and resistance to chemotherapy and radiotherapy, the prognosis for pancreatic cancer is poor, with a five-year survival rate of about 10%. If detected early, the five-year survival rate for pancreatic cancer patients can be increased to 30%-40%. Therefore, early screening and early diagnosis of pancreatic cancer are crucial for improving patient prognosis.

[0003] CT scans and endoscopic ultrasound-guided fine-needle aspiration biopsy (EUS-FNA) are the main diagnostic methods for pancreatic cancer. However, CT has limited sensitivity in diagnosing early lesions, especially stage I tumors, while histopathological biopsy is an invasive procedure with risks of complications. Liquid biopsy, due to its minimally invasive sampling method and repeatability, is used for early screening, diagnosis, monitoring, and treatment of pancreatic cancer. Commonly used blood tumor markers such as CA19-9, carcinoembryonic antigen (CEA), CA125, and CA242 have limited sensitivity and specificity for pancreatic cancer diagnosis and are prone to false negatives and false positives. Other blood component-based pancreatic cancer diagnostic markers such as CTC, ctDNA, cfDNA, miRNA, and lncRNA lack sufficient sensitivity and specificity, and are not validated internally or externally, thus failing to serve as effective differential diagnostic markers.

[0004] Small extracellular vesicles (sEVs) are small, double-membrane vesicles autonomously secreted and released by living cells, also known as exosomes. sEVs are rich in content and stable, making them advantageous in liquid biopsies. Compared to ctDNA and CTCs, exosomes exhibit higher sensitivity and specificity in the diagnosis of pancreatic cancer. sEVs miRNAs, lncRNAs, and proteins have been reported for the differential diagnosis of pancreatic cancer. Among these, single RNAs are more valuable than proteins in pancreatic cancer diagnosis, while miRNAs are abundant, resistant to degradation, and involved in tumorigenesis and development, making them more suitable as biomarkers for pancreatic cancer research. Existing studies combining sEVs miRNAs with other biomarkers have shown good diagnostic performance, achieving a sensitivity and specificity of 94% and 97% respectively for differentiating PDAC (PMID: 35850192). However, the classification of the control group samples in this study was not published, making it impossible to accurately assess its clinical diagnostic value. Therefore, it is necessary to use complex real-world sample cohorts to find validated and reliable exosomal miRNA biomarkers and explore their value in the diagnosis and clinical application of PDAC. Summary of the Invention

[0005] The purpose of this invention is to provide the application of exosomal miRNA biomarkers in the diagnosis of pancreatic cancer. This invention uses statistical and machine learning methods to screen exosomal miRNA biomarkers and construct diagnostic models for the differential diagnosis of pancreatic cancer.

[0006] The technical solution adopted by the present invention to achieve the above objectives is as follows: This invention provides the application of exosomal miRNA markers in the preparation of diagnostic reagents or kits for pancreatic cancer, wherein the exosomal miRNA markers are combinations of miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b, miR-142-3p, miR-378i, miR-452-5p, miR-140-3p, miR-95-3p, miR-219a-1-3p, miR-204-5p, miR-4429, miR-1246, miR-148b-3p, and miR-150-5p.

[0007] As a preferred embodiment, the combination of exosomal miRNA markers is a combination of miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b, miR-142-3p, miR-378i, miR-452-5p, miR-140-3p, miR-95-3p, miR-219a-1-3p, miR-204-5p, miR-4429, miR-1246, miR-148b-3p, miR-150-5p, let-7b-5p, miR-598-3p, and miR-625-3p.

[0008] This invention provides the application of exosomal miRNA markers in the preparation of diagnostic reagents or kits for pancreatic cancer, wherein the exosomal miRNA markers are combinations of miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b, miR-142-3p, miR-379-5p, miR-378i, miR-452-5p, miR-140-3p, miR-95-3p, miR-219a-1-3p, miR-4429, miR-1246, miR-148b-3p, miR-150-5p, let-7b-5p, miR-598-3p, and miR-625-3p.

[0009] This invention provides the application of exosomal miRNA markers in the preparation of diagnostic reagents or kits for pancreatic cancer, wherein the exosomal miRNA markers are combinations of miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b, miR-92a-3p, miR-6087, miR-142-3p, miR-423-5p, and miR-379-5p.

[0010] This invention provides the application of exosomal miRNA markers in the preparation of diagnostic reagents or kits for pancreatic cancer, wherein the exosomal miRNA markers are combinations of miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b, miR-142-3p, miR-423-5p, miR-378i, miR-452-5p, miR-140-3p, miR-95-3p, miR-219a-1-3p, miR-204-5p, miR-4429, miR-1246, let-7b-5p, miR-625-3p, miR-4532, miR-378c, miR-199a-3p, and miR-148a-3p.

[0011] The present invention also provides a kit for diagnosing pancreatic cancer, the kit comprising reagents for detecting the expression levels of each marker in the combination of exosomal miRNA markers described in any of the preceding claims.

[0012] As a preferred embodiment, the method for detecting the expression level of exosomal miRNA markers is as follows: using next-generation sequencing technology to detect the expression level of exosomal miRNA markers in plasma.

[0013] In a preferred embodiment, the kit further includes a diagnostic model, the formula of which is as follows:

[0014] X i This indicates that the model calculates the numerical result based on the biomarker expression of sample i, P(X). i ) represents the probability value of sample i to be cancer predicted by the diagnostic model; samples with a probability value greater than the reference value are considered to be pancreatic cancer patients and are represented by 1; samples with a probability value less than the reference value may be non-cancer individuals and are represented by 0, which are the final prediction results.

[0015] As a preferred embodiment, the diagnostic model uses a model algorithm selected from random forest, extreme random tree, neural network, gradient boosting machine or linear model algorithm.

[0016] The present invention also provides a system for diagnosing pancreatic cancer, the system comprising a detection unit and a diagnostic unit. The detection unit is used to detect the expression level of the exosomal miRNA markers described by the subject. The diagnostic unit is used to diagnose the subject based on the expression level of the subject's miRNA marker obtained from the detection unit, using a diagnostic model.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention establishes a high-performance diagnostic model for the differential diagnosis of pancreatic cancer based on blood exosome miRNA as a biomarker using statistical and machine learning methods. The model has been validated through a validation cohort, demonstrating excellent predictive efficacy for pancreatic cancer diagnosis. Its sensitivity and specificity are superior to existing clinical methods such as imaging, blood tumor markers, and pathological examinations, thereby reducing overtreatment and improving the quality of life for pancreatic cancer patients.

[0018] 2. Existing pancreatic cancer diagnostic models either focus on the differentiation between healthy individuals or pancreatitis and pancreatic cancer, or have small sample sizes and lack large-sample validation. In this invention, the non-pancreatic cancer group includes healthy individuals, patients with chronic pancreatitis, and patients with benign nodules, which better reflects the complex sample situation in the real clinical world. Attached Figure Description

[0019] Figure 1 This is a transmission electron microscope image of the plasma exosome detection results of this invention.

[0020] Figure 2 This is a graph showing the plasma exosome particle size distribution results of this invention.

[0021] Figure 3 This is the ROC curve of the risk scoring model constructed from 17 miRNA combinations consisting of miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b, miR-142-3p, miR-378i, miR-452-5p, miR-140-3p, miR-95-3p, miR-219a-1-3p, miR-204-5p, miR-4429, miR-1246, miR-148b-3p, and miR-150-5p in the training cohort.

[0022] Figure 4 This is the ROC curve of the risk scoring model constructed from 17 miRNA combinations of miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b, miR-142-3p, miR-378i, miR-452-5p, miR-140-3p, miR-95-3p, miR-219a-1-3p, miR-204-5p, miR-4429, miR-1246, miR-148b-3p, and miR-150-5p in the validation cohort.

[0023] Figure 5 This is the ROC curve of the risk scoring model constructed from 20 miRNA combinations consisting of miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b, miR-142-3p, miR-378i, miR-452-5p, miR-140-3p, miR-95-3p, miR-219a-1-3p, miR-204-5p, miR-4429, miR-1246, miR-148b-3p, miR-150-5p, let-7b-5p, miR-598-3p, and miR-625-3p in the training cohort.

[0024] Figure 6 This is the ROC curve of the risk scoring model constructed from 20 miRNA combinations consisting of miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b, miR-142-3p, miR-378i, miR-452-5p, miR-140-3p, miR-95-3p, miR-219a-1-3p, miR-204-5p, miR-4429, miR-1246, miR-148b-3p, miR-150-5p, let-7b-5p, miR-598-3p, and miR-625-3p in the validation cohort.

[0025] Figure 7 This is the ROC curve of the risk scoring model constructed from 20 miRNA combinations consisting of miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b, miR-142-3p, miR-379-5p, miR-378i, miR-452-5p, miR-140-3p, miR-95-3p, miR-219a-1-3p, miR-4429, miR-1246, miR-148b-3p, miR-150-5p, let-7b-5p, miR-598-3p, and miR-625-3p in the training cohort.

[0026] Figure 8This is the ROC curve of the risk scoring model constructed by combining 20 miRNAs of the present invention, consisting of miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b, miR-142-3p, miR-379-5p, miR-378i, miR-452-5p, miR-140-3p, miR-95-3p, miR-219a-1-3p, miR-4429, miR-1246, miR-148b-3p, miR-150-5p, let-7b-5p, miR-598-3p, and miR-625-3p, in the validation cohort.

[0027] Figure 9 This is the ROC curve of the risk scoring model constructed by the present invention, which consists of 11 miRNA combinations, namely miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b, miR-92a-3p, miR-6087, miR-142-3p, miR-423-5p, and miR-379-5p, in the training cohort.

[0028] Figure 10 This is the ROC curve of the risk scoring model constructed by the present invention, which consists of 11 miRNA combinations, namely miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b, miR-92a-3p, miR-6087, miR-142-3p, miR-423-5p, and miR-379-5p, in the validation cohort.

[0029] Figure 11This is the ROC curve of the risk scoring model constructed from 22 miRNA combinations consisting of miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b, miR-142-3p, miR-423-5p, miR-378i, miR-452-5p, miR-140-3p, miR-95-3p, miR-219a-1-3p, miR-204-5p, miR-4429, miR-1246, let-7b-5p, miR-625-3p, miR-4532, miR-378c, miR-199a-3p, and miR-148a-3p in the training cohort.

[0030] Figure 12 This is the ROC curve of the risk scoring model constructed from 22 miRNA combinations consisting of miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b, miR-142-3p, miR-423-5p, miR-378i, miR-452-5p, miR-140-3p, miR-95-3p, miR-219a-1-3p, miR-204-5p, miR-4429, miR-1246, let-7b-5p, miR-625-3p, miR-4532, miR-378c, miR-199a-3p, and miR-148a-3p in the validation cohort. Detailed Implementation

[0031] The technical solution of the present invention will be described in detail below with reference to the embodiments. Unless otherwise specified, all reagents and biological materials used below are commercial products.

[0032] Example 1: Screening of miRNA markers (1) Study cohort and clinical information A total of 60 patients clinically diagnosed with pancreatic cancer, 75 healthy individuals, 32 patients with pancreatitis, and 41 patients with benign pancreatic lesions were included in the study cohort. Blood samples were collected from pancreatic cancer patients before surgery. After enrollment, 30% of the 208 samples were randomly selected as the validation cohort, and the remainder as the training cohort. The training cohort consisted of 140 patients: 50 healthy individuals, 22 patients with pancreatitis, 28 patients with benign pancreatitis, and 40 patients with pancreatic cancer. The validation cohort consisted of 68 patients: 25 healthy individuals, 10 patients with pancreatitis, 13 patients with benign pancreatic lesions, and 20 patients with pancreatic cancer. Benign lesions included intraductal papillary myxomas (IPMNs), serous cystadenomas, solid pseudopapillary tumors, pancreatic cysts, and pancreatic abscesses. Table 1 shows the clinical information of the included samples, indicating no significant difference in the proportion of samples from different categories.

[0033] Table 1

[0034] (2) Extraction and characterization of plasma exosomes 1) Blood collection and plasma separation Blood samples from pancreatic cancer patients, healthy individuals, patients with pancreatitis, and patients with benign pancreatitis were collected and stored in 10 ml vacuum blood collection tubes (REF367525, BD, USA). After gentle inversion and mixing, plasma separation was performed using a two-step centrifugation method: First, the blood collection tubes were centrifuged at 1600 g and 4°C for 10 min to determine and record the plasma hemolysis grade. Samples with a grade of 4 or lower were used for subsequent studies. Second, the supernatant was transferred to 1.5 mL EP tubes and centrifuged at 16000 g for 15 min at 4°C to remove residual cell debris. Finally, the supernatant was aliquoted into 1.5 mL EP tubes at 1 ml each and stored at -80°C for later use.

[0035] 2) Extraction of exosomes Exosomes were extracted from the plasma of pancreatic cancer patients, healthy individuals, and patients with pancreatitis and benign pancreatitis using exosome separation reagent (L3525, 3DMed, Shanghai, China). The exosome extraction process was as follows: After the frozen plasma samples were taken out, they were first placed in a 37°C water bath to thaw. After the samples were completely thawed, they were centrifuged at 12000 g for 10 min at 4°C. The supernatant was collected and filtered sequentially through a 0.45 µm filter column (CLS8163-100EA, Corning, USA) and a 0.22 µm filter column (CLS8161-100EA, Corning, USA). The centrifugation conditions were 4°C and 12000 g for 5 min. The volume of the filtrate was collected and measured. 0.25 times the volume of L3525 reagent was added, and the mixture was thoroughly mixed. The mixture was then incubated at 4°C for 30 min. After incubation, the mixture was centrifuged at 4°C and 4700 g for 30 min. Discard the supernatant and resuspend the exosomes in 200 µL of phosphate-buffered saline (PBS, pH 7.4).

[0036] 3) Characterization of plasma exosomes The morphology of exosomes resuspended in PBS was examined using transmission electron microscopy (TEM). First, plasma exosomes were fixed in 4% paraformaldehyde and then transferred to a carbon-coated copper mesh for an electron microscope. The copper mesh was washed twice with PBS, then once each with PBS containing glycine (50 mM) and 0.5% BSA. The mesh was then stained with 2% uranyl acetate. Finally, the morphology of the plasma exosomes was characterized using a transmission electron microscope (H-7650, Hitachi High-Technologies, Japan). TEM results showed that the extracted exosomes exhibited a typical "horseshoe" morphology (see...). Figure 1 ).

[0037] The particle size distribution of plasma exosomes was detected using a nanoparticle tracking analysis (NTA) system: Resuspended plasma exosomes were diluted with PBS (1 × 10⁷ - 1 × 10⁹ / ml) and mixed thoroughly by pipetting. The sample was added to the sample chamber of the NTA instrument (NanoSight NS300, Malvern, UK). A 488 nm excitation module was used, and the camera lens parameters were set (shutter speed 890, gain 146, detection threshold 7). At least 200 complete tracks were analyzed for each video. The nanoparticle tracking data of plasma exosomes were analyzed using NTA software (version 2.3). NTA results showed that the main particle size peak of plasma exosomes was 103 nm, consistent with the particle size distribution of exosomes (see...). Figure 2 ).

[0038] (3) Extraction and expression level detection of plasma exosome miRNA 1) Extraction of exosomal miRNA from plasma Exosomal miRNAs were extracted from the plasma of pancreatic cancer patients, healthy individuals, patients with pancreatitis, and patients with benign pancreatitis using the miRNeasy Serum / Plasma Kit (217184, QIAGEN, Shanghai, China). The distribution and total amount of exosomal miRNA fragments were detected using a 2100 analyzer with accompanying chips and reagents (5067-1548, Agilent, USA). Detailed experimental procedures can be found in the product manual.

[0039] 2) Detection of plasma exosome miRNA expression The expression levels of exosomal miRNAs in the plasma of pancreatic cancer patients, healthy individuals, and patients with pancreatitis and benign pancreatitis were detected using small RNA sequencing technology. Plasma exosomal miRNA libraries were prepared using the NEBNext, Multiplex Small RNA Library Prep Set for Illumina (E7300L, NEB, USA) kit, and then purified using the NucleoSpin Geland PCR Clean-up kit (740609.50, MN, Germany). Detailed steps for library construction and purification can be found in the product manual, which briefly describes the process as follows: 100 ng of miRNA with a total volume not exceeding 6 µl was sequentially subjected to 3' adapter ligation, reverse transcription primer ligation, 5' adapter ligation, reverse transcription, PCR amplification, and library purification. The library concentration was then measured using a Qubit 4.0 quantitative PCR instrument and the accompanying Qubit dsDNA HS Assay Kit (Q32854, Thermofisher, USA). Peak shape was detected using a 2100 analyzer and the accompanying High Sensitivity DNA Kit & Reagents (5067-4626, Agilent, USA). Finally, sequencing was performed using the Illumina NovaSeq sequencing platform (PE150 sequencing strategy).

[0040] (4) Sequencing data analysis workflow The expression levels of exosomal miRNAs in the plasma of pancreatic cancer patients, healthy individuals, patients with pancreatitis, and benign pancreatitis were obtained using small RNA next-generation sequencing technology. The sequencing data analysis workflow is as follows: 1) Sequencing data alignment. After removing the sequencing adapters from the Small RNA-seq data, the sequencing data was aligned to the human reference genome hg19 (genome download link: http: / / hgdownload.soe.ucsc.edu / goldenPath / hg19 / bigZips / ) using BWA software (version: 0.7.12-r1039), and the number of reads aligned to miRNAs was counted.

[0041] 2) miRNA annotation. miRNAs were annotated using the Gencode v25 and miRBase v21 databases, retaining those annotated as known mature miRNAs for subsequent analysis.

[0042] 3) miRNA filtering. For the training cohort, mature miRNAs with a length of 30 nt or less and covering more than 10 reads in at least g samples in the training cohort data are retained for subsequent analysis, where g is the number of samples in the smallest group after cohort grouping; for the validation cohort, miRNAs selected from the training cohort are retained for subsequent analysis.

[0043] 4) miRNA expression level normalization. The original miRNA expression levels of the training cohort samples were normalized using the M-value weighted truncated mean (TMM) method, and the same parameters were used to process the validation cohort samples.

[0044] (5) Discovery of biomarkers Samples were grouped according to pathological examination results. Based on the expression levels of miRNAs in the training cohort, statistical and machine learning methods were used to discover miRNAs that could be used as biomarkers to distinguish between non-cancer controls (healthy individuals, pancreatitis patients, and benign pancreatitis patients) and pancreatic cancer patients. The process is as follows: 1) Training cohort grouping. Based on the pathological test results of the samples, the patient samples in the training cohort were divided into a non-cancer control group (healthy people, pancreatitis and benign patients) and a pancreatic cancer patient group, for a total of 2 groups.

[0045] 2) Statistical methods were used to screen biomarkers. First, the Kruskal-Wallis H test and U test were used to analyze the expression differences of all miRNAs between the two groups. Thirty-eight miRNAs with statistical differences between groups (p-value <= 0.05) and no significant differences between healthy individuals, pancreatitis patients, and benign patients were selected as candidate biomarkers for subsequent analysis.

[0046] 3) Machine learning methods were used to screen biomarkers. The miRNAs screened in step 2) were used as initial biomarkers. Five machine learning algorithms were employed: random forest, extremely randomized trees, neural network, gradient boosting machine (GBM), and linear model. The feature importance of each biomarker was calculated in each algorithm, and biomarkers with importance > 0 were retained. This process was repeated until the feature importance of all miRNAs included in each algorithm was > 0. Finally, 32 biomarkers were selected for subsequent analysis.

[0047] (6) Construction of diagnostic models for non-cancer control groups and pancreatic cancer patients Using non-cancer control groups and pancreatic cancer patients as classification prediction targets, the biomarker combination discovered in (5) was used, and five machine learning algorithms were employed: random forest, extremely randomized trees, neural network, gradient boosting machine (GBM), and linear model. Different hyperparameters were preset for each algorithm. Using training cohort sample data and pathological test results as the ground truth, multiple diagnostic models with different machine learning algorithms were trained. The model formulas are as follows:

[0048] X i This indicates that the model calculates the numerical result based on the biomarker expression of sample i, P(X). i The value represents the probability of sample i being cancer predicted by the diagnostic model. Samples with a probability value greater than the reference value are considered to be pancreatic cancer patients and are represented by 1; samples with a probability value less than the reference value may be non-cancer individuals and are represented by 0. This is the final prediction result. The trained model is saved on the hard drive as a file. When calling the model, inputting the sample marker expression value will yield the model's prediction result.

[0049] The classification performance of the model was evaluated using a 5-fold cross-validation method. The evaluation metrics included the area under the receiver operating characteristic curve (AUC, ranging from 0 to 1), sensitivity (ranging from 0 to 1), and specificity (ranging from 0 to 1), with higher values ​​indicating better performance. The evaluation results of biomarkers in the training cohort are shown in Table 2.

[0050] Among the 32 miRNAs selected using a combination of statistical methods and various machine learning algorithms, different machine learning algorithms selected different biomarkers, as shown in Table 2, which illustrates the biomarker combinations corresponding to different model algorithms. These combinations were selected by different model algorithms during model training based on the calculated miRNA feature weights. Each time, only miRNAs with weights > 0 were retained, and the AUC was re-evaluated. This process was iterated until the weights of all retained miRNAs in each model were > 0. As miRNAs with weights <= 0 were continuously removed, the model's AUC continuously improved. Therefore, the AUC corresponding to each model + biomarker combination in Table 2 represents the optimal performance of that model; adding or removing biomarkers will decrease the AUC.

[0051] Table 2

[0052]

[0053]

[0054] Example 2: Validation of the predictive efficacy of the diagnostic model in pancreatic cancer patients and non-cancer control populations. To validate the predictive efficacy of the established diagnostic model for pancreatic cancer, a separate validation cohort consisting of pancreatic cancer patients and non-cancer controls was selected to assess the model's classification performance. Pathological examination results were used as the true value. Model evaluation metrics included the area under the receiver operating characteristic curve (AUC, ranging from 0 to 1), sensitivity (ranging from 0 to 1), and specificity (ranging from 0 to 1), with higher values ​​indicating better performance. The evaluation results of biomarkers in the validation cohort are shown in Table 3.

[0055] Table 3

[0056]

[0057]

[0058] Based on the data in Tables 2 and 3, the risk scoring model constructed from five biomarker combinations that demonstrated good sensitivity and specificity (>0.8) in both the training and validation cohorts was selected as the final diagnostic model. The evaluation results of the five biomarker combinations are as follows: The risk scoring model constructed from 17 miRNA combinations consisting of miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b, miR-142-3p, miR-378i, miR-452-5p, miR-140-3p, miR-95-3p, miR-219a-1-3p, miR-204-5p, miR-4429, miR-1246, miR-148b-3p, and miR-150-5p showed AUC, sensitivity, and specificity of 0.939, 0.93, and 0.90, respectively, in the training cohort. (See [link to relevant documentation]). Figure 3 The ROC curve of the risk scoring model constructed for this biomarker combination is shown in the training cohort. The AUC, sensitivity, and specificity of the risk scoring model constructed from this 17 miRNA combination in the validation cohort were 0.951, 0.95, and 0.83, respectively. (See [link to relevant documentation]). Figure 4 The ROC curve of the risk scoring model constructed for this biomarker combination is shown in the validation cohort. The results indicate that this risk prediction model exhibits high AUC, sensitivity, and specificity in both the training and validation cohorts, demonstrating superior predictive performance.

[0059] The risk scoring model constructed from 20 miRNA combinations consisting of miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b, miR-142-3p, miR-378i, miR-452-5p, miR-140-3p, miR-95-3p, miR-219a-1-3p, miR-204-5p, miR-4429, miR-1246, miR-148b-3p, miR-150-5p, let-7b-5p, miR-598-3p, and miR-625-3p showed AUC, sensitivity, and specificity of 0.936, 0.90, and 0.91, respectively, in the training cohort. (See [link to relevant documentation]). Figure 5 The ROC curve of the risk scoring model constructed for this biomarker combination is shown in the training cohort. The AUC, sensitivity, and specificity of this risk scoring model constructed from the 20 miRNAs in the validation cohort were 0.943, 0.90, and 0.85, respectively. (See [link to relevant documentation]). Figure 6 The ROC curve of the risk scoring model constructed for this biomarker combination is shown in the validation cohort. The results indicate that this risk prediction model exhibits high AUC, sensitivity, and specificity in both the training and validation cohorts, demonstrating superior predictive performance.

[0060] The risk scoring model constructed from 20 miRNA combinations consisting of miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b, miR-142-3p, miR-379-5p, miR-378i, miR-452-5p, miR-140-3p, miR-95-3p, miR-219a-1-3p, miR-4429, miR-1246, miR-148b-3p, miR-150-5p, let-7b-5p, miR-598-3p, and miR-625-3p showed AUC, sensitivity, and specificity of 0.924, 0.88, and 0.91, respectively, in the training cohort. (See [link to relevant documentation]). Figure 7 The ROC curve of the risk scoring model constructed for this biomarker combination is shown in the training cohort. The AUC, sensitivity, and specificity of this risk scoring model constructed from the 20 miRNA combination in the validation cohort were 0.911, 0.80, and 0.88, respectively. (See [link to relevant documentation]). Figure 8 The ROC curve of the risk scoring model constructed for this biomarker combination is shown in the validation cohort. The results indicate that this risk prediction model exhibits high AUC, sensitivity, and specificity in both the training and validation cohorts, demonstrating superior predictive performance.

[0061] The risk scoring model constructed from an 11 miRNA combination consisting of miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b, miR-92a-3p, miR-6087, miR-142-3p, miR-423-5p, and miR-379-5p showed AUC, sensitivity, and specificity of 0.903, 0.88, and 0.84, respectively, in the training cohort. (See [link to relevant documentation]). Figure 9 The ROC curve of the risk scoring model constructed for this biomarker combination is shown in the training cohort. The AUC, sensitivity, and specificity of the risk scoring model constructed from this 11 miRNA combination in the validation cohort were 0.924, 0.90, and 0.88, respectively. (See [link to relevant documentation]). Figure 10 The ROC curve of the risk scoring model constructed for this biomarker combination is shown in the validation cohort. The results indicate that this risk prediction model exhibits high AUC, sensitivity, and specificity in both the training and validation cohorts, demonstrating superior predictive performance.

[0062] Consisting of miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b,miR-142-3p, miR-423-5p, miR-378i, miR-452-5p, miR-140-3p, miR-95-3p, miR-219a-1-3p, miR-204-5p,miR-4429,miR-1246,let-7b-5p,miR-625-3p,miR-4532,miR-378c, The risk scoring model constructed from a combination of 22 miRNAs, including miR-199a-3p and miR-148a-3p, achieved AUC, sensitivity, and specificity of 0.896, 0.93, and 0.81, respectively, in the training cohort. (See [link to relevant documentation]). Figure 11 The ROC curve of the risk scoring model constructed for this biomarker combination is shown in the training cohort. The AUC, sensitivity, and specificity of this risk scoring model constructed from the 22 miRNA combinations in the validation cohort were 0.920, 0.90, and 0.81, respectively. (See [link to relevant documentation]). Figure 12 The ROC curve of the risk scoring model constructed for this biomarker combination is shown in the validation cohort. The results indicate that this risk prediction model exhibits high AUC, sensitivity, and specificity in both the training and validation cohorts, demonstrating superior predictive performance.

[0063] The above are merely some preferred embodiments of the present invention, and the present invention is not limited to the contents of these embodiments. For those skilled in the art, various changes and modifications can be made within the scope of the present invention's technical solutions, and any changes and modifications made are within the protection scope of the present invention.

Claims

1. The application of exosomal miRNA markers in the preparation of diagnostic reagents or kits for pancreatic cancer, characterized in that: The exosomal miRNA markers are combinations of miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b, miR-142-3p, miR-378i, miR-452-5p, miR-140-3p, miR-95-3p, miR-219a-1-3p, miR-204-5p, miR-4429, miR-1246, miR-148b-3p, and miR-150-5p.

2. The application of the exosomal miRNA marker according to claim 1 in the preparation of detection reagents or kits for pancreatic cancer diagnosis, characterized in that: The combination of exosomal miRNA markers also includes let-7b-5p, miR-598-3p, and miR-625-3p.

3. The application of exosomal miRNA markers in the preparation of diagnostic reagents or kits for pancreatic cancer, characterized in that: The exosomal miRNA markers are combinations of miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b, miR-142-3p, miR-379-5p, miR-378i, miR-452-5p, miR-140-3p, miR-95-3p, miR-219a-1-3p, miR-4429, miR-1246, miR-148b-3p, miR-150-5p, let-7b-5p, miR-598-3p, and miR-625-3p.

4. The application of exosomal miRNA markers in the preparation of diagnostic reagents or kits for pancreatic cancer, characterized in that: The exosomal miRNA markers are a combination of miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b, miR-92a-3p, miR-6087, miR-142-3p, miR-423-5p, and miR-379-5p.

5. The application of exosomal miRNA markers in the preparation of diagnostic reagents or kits for pancreatic cancer, characterized in that: The exosomal miRNA markers are combinations of miR-574-5p, miR-7704, miR-106b-3p, miR-877-5p, let-7a-3p, miR-320b, miR-142-3p, miR-423-5p, miR-378i, miR-452-5p, miR-140-3p, miR-95-3p, miR-219a-1-3p, miR-204-5p, miR-4429, miR-1246, let-7b-5p, miR-625-3p, miR-4532, miR-378c, miR-199a-3p, and miR-148a-3p.

6. A reagent kit for diagnosing pancreatic cancer, characterized in that: The kit includes reagents for detecting the expression level of exosomal miRNA markers as described in any one of claims 1-5.

7. The reagent kit according to claim 6, characterized in that: The method for detecting the expression level of exosomal miRNA markers is as follows: the expression level of exosomal miRNA markers in plasma is detected by next-generation sequencing technology.

8. The reagent kit according to claim 6, characterized in that: The kit also includes a diagnostic model, the formula of which is as follows: ; X i This indicates that the model calculates the numerical result based on the biomarker expression of sample i, P(X). i ) represents the probability value of sample i to be cancer predicted by the diagnostic model; samples with a probability value greater than the reference value are considered to be pancreatic cancer patients and are represented by 1; samples with a probability value less than the reference value may be non-cancer individuals and are represented by 0, which are the final prediction results.

9. The reagent kit according to claim 7, characterized in that: The diagnostic model uses an algorithm selected from random forest, extreme random tree, neural network, gradient boosting machine or linear model algorithm.

10. A system for diagnosing pancreatic cancer, characterized in that: The system includes a detection unit and a diagnostic unit. The detection unit is used to detect the expression level of the miRNA biomarker as described in any one of claims 1-5 in the subject; The diagnostic unit is used to diagnose the subject based on the expression level of the subject's miRNA marker obtained from the detection unit, using the diagnostic model described in claim 8.