Pancreatic cancer diagnostic markers based on exosomal miRNAs and their applications

By constructing a pancreatic cancer diagnostic model using specific miRNA pairs in serum or plasma extracellular vesicles, the problem of instability of existing markers in different data centers is solved, and a high sensitivity and specific diagnosis of pancreatic cancer is achieved, improving the accuracy and stability of early diagnosis.

CN119082307BActive Publication Date: 2025-07-18PEKING UNION MEDICAL COLLEGE HOSPITAL +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411480045.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-22
Publication Date
2025-07-18
Estimated Expiration
2044-10-22

AI Technical Summary

Technical Problem

The sensitivity and specificity of existing pancreatic cancer diagnostic markers are insufficient, and the results in different data centers are unstable, resulting in poor repetitive diagnostic efficacy and lack of effective early-stage pancreatic cancer diagnostic kits.

Method used

Using specific types of miRNA pairs in extracellular vesicles of serum or plasma, such as hsa-miR-98/hsa-miR-25, hsa-miR-378i/hsa-miR-150, etc., a pancreatic cancer diagnosis model is established by constructing miRNA expression ratios as input characteristics, avoiding reliance on routine internal reference gene correction, and improving the stability and accuracy of diagnosis.

Benefits of technology

A highly sensitive and specific pancreatic cancer diagnosis is achieved, which can maintain stable diagnostic performance in different data centers, and improves the screening accuracy and prognosis evaluation ability of early-stage pancreatic cancer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005097217140000071
    Figure BDA0005097217140000071
  • Figure BDA0005097217140000091
    Figure BDA0005097217140000091
  • Figure BDA0005097217140000102
    Figure BDA0005097217140000102
Patent Text Reader

Abstract

The present invention provides a pancreatic cancer diagnostic marker based on extracellular vesicle miRNA and its application. The pancreatic cancer diagnostic or prognostic marker based on extracellular vesicle miRNA of the present invention relates to a miRNA combination in serum or plasma extracellular vesicles with biological significance. Without using a conventional reference gene as a reference, the pancreatic cancer diagnostic model constructed with the expression ratio of paired miRNAs in the combination as the input features of the model has high diagnostic sensitivity and specificity for pancreatic cancer, is applicable to different data centers, has good diagnostic stability, and has great potential in the screening of early pancreatic cancer, and has important practical significance for the early diagnosis or prognostic prediction of pancreatic cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical detection, and particularly relates to a pancreatic cancer diagnostic marker based on exosomal miRNA and its application. Background Art

[0002] Pancreatic cancer (PC) is one of the most lethal malignant tumors globally, with its incidence and mortality rates being almost equal. Due to the rapid progression of the disease and the insidious early symptoms, only 15%-20% of patients are diagnosed at an early stage. Currently, in the United States, Europe, Japan, and China, the incidence and mortality rates of pancreatic cancer are increasing year by year. It is estimated that by 2050, the global incidence of pancreatic cancer will reach 18.6 per 100,000 people, with an average annual growth rate of 1.1%, which will impose a huge burden on the public health system. The high fatality rate of pancreatic cancer is mainly attributed to the special anatomical location of the pancreas and the non-specific symptoms, resulting in many cases being discovered and diagnosed at an advanced stage, when obvious clinical symptoms usually appear. Globally, the 5-year survival rate of pancreatic cancer is generally less than 10%, with little difference between high-income and low- and middle-income countries, and the survival rate is the lowest among all cancers. Despite the progress in diagnosis and treatment in recent years, the high mortality rate of pancreatic cancer is still closely related to late detection and the limitations of existing treatment methods. Therefore, the development of new screening methods, early diagnostic tools, and treatment strategies is crucial for improving the prognosis of pancreatic cancer patients.

[0003] The diagnosis of pancreatic cancer mainly relies on imaging and serum tumor markers. Imaging methods include ultrasound, computed tomography (CT), and magnetic resonance imaging (MRI), which are commonly used to detect obvious pancreatic masses but have limitations in early diagnosis, especially for early cases without obvious masses. In addition, the diagnostic results of imaging are easily affected by the operator's experience and equipment performance. The serum marker CA19-9 is the most commonly used indicator in the diagnosis and monitoring of pancreatic cancer, but CA19-9 may show false positives in cases of biliary tract infection, inflammation, or obstruction, and may show false negatives in individuals negative for the Lewis antigen. Therefore, the early diagnostic effect of CA19-9 as a single indicator is not ideal. In contrast, molecular diagnostic techniques provide more accurate means for auxiliary diagnosis by detecting pancreatic cancer-related gene, protein, and RNA markers. Combining imaging, biomarker detection, and molecular diagnostic techniques helps to improve the accuracy of early screening for pancreatic cancer and provides strong support for improving the survival rate of patients.

[0004] Extracellular vesicles are small vesicles secreted by cells and enter the body fluids or extracellular environment, playing an important role in cell-to-cell communication. Due to the protection of the lipid bilayer, extracellular vesicles can stably exist in various biological body fluids and cell culture media; this property makes the contents in extracellular vesicles potentially become a reliable tumor marker, and miRNA is one of the contents of extracellular vesicles. It is reported that the levels of extracellular vesicle miRNA expression are different in different diseases and physiological conditions, and it also plays an important regulatory role in the process of malignant tumors. Therefore, extracellular vesicle miRNA may become a new form of biomarker for early cancer diagnosis and prognosis tracking.

[0005] However, the biomarkers for early screening of pancreatic cancer have problems such as low sensitivity, low specificity, or poor repeatability of diagnostic efficacy. For example, in the existing technology, relatively stable internal reference genes are often selected as references to detect the relative expression levels of biomarkers. However, existing literature has shown that although the expression of internal reference genes is relatively stable, there are still some differences in the expression levels of internal reference genes under different physiological states, resulting in inaccurate relative expression levels of biomarkers. The biomarkers or biomarker combinations screened in this way will make the repeatability of diagnostic efficacy worse (that is, the results are unstable). When the sample size, sample distribution, etc. change, the accuracy of diagnostic efficacy will decrease, making the reliability of the biomarkers decrease. In addition, different batches or different data centers from multiple sources will also lead to unstable results. For example, the biomarkers or biomarker combinations screened in a data set center (results detected in the same center or institution) are prone to unstable results (poor repeatability) when applied to other hospitals or data centers, and even unstable results will occur in different batches. Currently, there is still a lack of effective early pancreatic cancer diagnostic kits on the market. Therefore, it is urgent to screen out pancreatic cancer diagnostic biomarkers or biomarker combinations with good repeatability, high reliability, and high diagnostic efficacy. Summary of the Invention

[0006] Object of the Invention

[0007] Aiming at the problems existing in the above-mentioned existing technical methods, the present invention aims to provide a diagnostic or prognostic biomarker for pancreatic cancer with high sensitivity and specificity, an analysis system, a kit, or its application based thereon, including: an analysis system, a kit for predicting pancreatic cancer or evaluating the prognosis of pancreatic cancer, or an application in the preparation of a kit or analysis system for pancreatic cancer diagnosis, efficacy or prognosis evaluation, or related drug evaluation.

[0008] Solution

[0009] To achieve the above object, the present invention provides the following technical solutions:

[0010] In a first aspect, the present invention provides a combination of diagnostic or prognostic markers for pancreatic cancer, and the diagnostic or prognostic markers for pancreatic cancer comprise a marker combination composed of at least two or at least four of the following miRNAs:

[0011] hsa-miR-98, hsa-miR-25, hsa-miR-378i, hsa-miR-150, hsa-miR-501, hsa-miR-155, hsa-miR-193a, hsa-miR-1180.

[0012] Furthermore, the diagnostic or prognostic markers for pancreatic cancer comprise one or more of the following miRNA combinations:

[0013] hsa-miR-98 and hsa-miR-25, optionally, hsa-miR-25 is used to calibrate the relative expression level of hsa-miR-98 (i.e., the expression ratio of hsa-miR-98 and hsa-miR-25);

[0014] hsa-miR-378i and hsa-miR-150, optionally, hsa-miR-150 is used to calibrate the relative expression level of hsa-miR-378i (i.e., the expression ratio of hsa-miR-378i and hsa-miR-150);

[0015] hsa-miR-501 and hsa-miR-150, optionally, hsa-miR-150 is used to calibrate the relative expression level of hsa-miR-501 (i.e., the expression ratio of hsa-miR-501 and hsa-miR-150);

[0016] hsa-miR-378i and hsa-miR-155, optionally, hsa-miR-155 is used to calibrate the relative expression level of hsa-miR-378i;

[0017] hsa-miR-193a and hsa-miR-155, optionally, hsa-miR-155 is used to calibrate the relative expression level of hsa-miR-193a (i.e., the expression ratio of hsa-miR-193a and hsa-miR-155);

[0018] hsa-miR-150 and hsa-miR-1180, optionally, hsa-miR-1180 is used to calibrate the relative expression level of hsa-miR-150 (i.e., the expression ratio of hsa-miR-150 and hsa-miR-1180).

[0019] Furthermore, the detection of the expression level of the biomarker does not rely on a conventional reference gene as a relative reference or calibration.

[0020] Furthermore, the diagnostic or prognostic biomarker for pancreatic cancer further comprises CA19-9.

[0021] As a preference, the diagnostic or prognostic biomarker for pancreatic cancer comprises one or two of the following miRNA combinations 1) - 2):

[0022] 1) Combination 1: hsa-miR-98, hsa-miR-25, hsa-miR-378i, hsa-miR-150;

[0023] 2) Combination 2: hsa-miR-501, hsa-miR-150, hsa-miR-378i, hsa-miR-155, hsa-miR-193a, hsa-miR-1180.

[0024] As a further preference, the miRNA in the miRNA pair is the miRNA in serum or plasma.

[0025] Furthermore, the miRNA in the miRNA pair is the miRNA in serum or plasma glycosylated extracellular vesicles.

[0026] In a second aspect, there is provided an analysis system for predicting pancreatic cancer or evaluating the prognosis of pancreatic cancer, comprising a data analysis module, which is used for analyzing the expression levels of the biomarker combination of the target object to be predicted and calculating the ratio of the expression levels of the miRNA pair or directly using the ratio of the expression levels of the miRNA pair as an input feature, and is further used for calculating the prediction value of whether the target object is pancreatic cancer according to the ratio of the expression levels of the miRNA pair; the input feature is selected from at least one of the ratios of the expression levels of the following miRNA pairs:

[0027] hsa-miR-98 / hsa-miR-25;

[0028] hsa-miR-378i / hsa-miR-150;

[0029] hsa-miR-501 / hsa-miR-150;

[0030] hsa-miR-378i / hsa-miR-155;

[0031] hsa-miR-193a / hsa-miR-155;

[0032] hsa-miR-150 / hsa-miR-1180.

[0033] Further, the input features are selected from one or more of the following combinations:

[0034] 1) Combination 1: The expression level ratio of hsa-miR-98 / hsa-miR-25, hsa-miR-378i / hsa-miR-150;

[0035] 2) Combination 2: The expression level ratio of hsa-miR-501 / hsa-miR-150; hsa-miR-378i / hsa-miR-155; hsa-miR-193a / hsa-miR-155; hsa-miR-150 / hsa-miR-1180.

[0036] Further, a prediction model for outputting the probability value of pancreatic cancer is stored in the data analysis module, wherein the construction method of the prediction model includes: using the expression level ratio of the miRNA pairs in pancreatic cancer, benign, and healthy samples as input variables, and whether it is pancreatic cancer as the dependent variable, and obtaining the prediction model through multiple logistic regression calculation and testing it with a validation set.

[0037] Further, when the detection method is RT-qPCR, the expression level ratio is the CT difference of the miRNA pair (the CT value is the conversion of the expression level, so the CT difference represents the expression level ratio).

[0038] Further, the input features further include the expression level of CA19-9 of the target object to be predicted.

[0039] In a third aspect, there is provided an application of the combination of diagnostic or prognostic markers for pancreatic cancer described in the first aspect in the preparation of a kit or analysis system for diagnosing pancreatic cancer, evaluating the efficacy or prognosis, or evaluating related drugs.

[0040] In a fourth aspect, there is provided an application of detecting the relative expression level of the combination of diagnostic or prognostic markers for pancreatic cancer described in the first aspect in serum or plasma glycosylated extracellular vesicles in the preparation of a kit or analysis system for diagnosing pancreatic cancer, evaluating the efficacy or prognosis, or evaluating related drugs. Optionally, the marker combination is selected from one of the following combinations:

[0041] 1) Combination 1: The expression level of hsa-miR-98 relative to hsa-miR-25, the expression level of hsa-miR-378i relative to hsa-miR-150;

[0042] 2) Combination two: The expression level of hsa-miR-501 relative to hsa-miR-150; the expression level of hsa-miR-378i relative to hsa-miR-155; the expression level of hsa-miR-193a relative to hsa-miR-155; the expression level of hsa-miR-150 relative to hsa-miR-1180.

[0043] In a fifth aspect, a kit for diagnosing pancreatic cancer, evaluating therapeutic effect or prognosis, or evaluating related drugs is provided. The kit includes the diagnostic or prognostic marker combination of pancreatic cancer described in the first aspect and / or its detection reagents.

[0044] Furthermore, in patients with pancreatic cancer, the expression quantity ratio of the miRNA pairs in their serum glycosylated extracellular vesicles has a significant difference compared with that of healthy humans.

[0045] Furthermore, the detection reagents include primers and probes for detecting corresponding miRNAs. Among them, the primer and probe sequences of each miRNA are as follows:

[0046] The nucleotide sequence of the upstream primer of hsa-miR-150 is as shown in SEQ ID NO:1;

[0047] The nucleotide sequence of the upstream primer of hsa-miR-25 is as shown in SEQ ID NO:2;

[0048] The nucleotide sequence of the upstream primer of hsa-miR-378i is as shown in SEQ ID NO:3;

[0049] The nucleotide sequence of the upstream primer of hsa-miR-98 is as shown in SEQ ID NO:4;

[0050] The nucleotide sequence of the upstream primer of hsa-miR-1180 is as shown in SEQ ID NO:6;

[0051] The nucleotide sequence of the upstream primer of hsa-miR-501 is as shown in SEQ ID NO:7;

[0052] The nucleotide sequence of the upstream primer of hsa-miR-193a is as shown in SEQ ID NO:8;

[0053] The nucleotide sequence of the upstream primer of hsa-miR-155 is as shown in SEQ ID NO:9;

[0054] The nucleotide sequence of the probe is as shown in SEQ ID NO:5;

[0055] The reverse primer is a common downstream primer, and its nucleotide sequence is as shown in SEQ ID NO:10.

[0056] Furthermore, primers and probes that do not contain conventional reference genes.

[0057] Beneficial effects

[0058] The diagnostic or prognostic markers for pancreatic cancer provided by the present invention involve miRNA pairs in serum or plasma extracellular vesicles that have biological significance, especially miRNA pairs in specific types of glycosylated extracellular vesicles. Using the expression ratio of this miRNA pair as the input feature of the pancreatic cancer diagnostic model avoids the problem of unstable diagnostic performance of existing diagnostic markers caused by differences such as unstable reference genes and experimental batches. Thus, more accurate, stable, and reproducible diagnosis or prognostic evaluation can be achieved; the pancreatic cancer diagnostic model constructed with the expression ratio of the miRNA pair as the model input feature has high diagnostic sensitivity and specificity, and also has great potential in the screening of early pancreatic cancer. Description of the drawings

[0059] One or more embodiments are illustrated by way of example in the pictures of the corresponding drawings. These exemplary illustrations do not constitute a limitation on the embodiments. Here, the special word "exemplary" means "serving as an example, embodiment, or illustrative". Here, any embodiment illustrated as "exemplary" does not have to be construed as superior to or better than other embodiments.

[0060] Figure 1 Heat map of the expression ratios of 39 paired miRNAs screened in Example 1 on NGS samples.

[0061] Figure 2 LDA graph of the expression ratios of 39 paired miRNAs detected by RT-qPCR in Example 2.

[0062] Figure 3 Box plot of the CT value differences of RT-qPCR of 2 paired miRNAs in Example 3, where advanced and intermediate stages represent advanced and intermediate stages of pancreatic cancer, early stage represents early stage of pancreatic cancer, benign represents benign diseases, and healthy represents healthy population.

[0063] Figure 4 ROC curve of pancreatic cancer in different datasets for the pancreatic cancer diagnostic model constructed based on 2 paired miRNAs of Preferred Combination 1 in Example 3.

[0064] Figure 5 Box plot of the model scores for the model constructed based on 2 paired miRNAs of Preferred Combination 1 in Example 3; where advanced and intermediate stages represent advanced and intermediate stages of pancreatic cancer, early stage represents early stage of pancreatic cancer, benign represents benign diseases, and healthy represents healthy population.

[0065] Figure 6 Box plot of the CT value differences of the 4 paired miRNAs in Example 4 for RT-qPCR, where the middle and late stage represents the middle and late stage of pancreatic cancer, the early stage represents the early stage of pancreatic cancer, the benign represents benign diseases, and the healthy represents the healthy population.

[0066] Figure 7 ROC curve of pancreatic cancer of the pancreatic cancer diagnosis model constructed based on the 4 paired miRNAs of the second preferred combination in Example 4 in different datasets.

[0067] Figure 8 Box plot of the model scores of the model constructed based on the 4 paired miRNAs of the second preferred combination in Example 4; where the middle and late stage represents the middle and late stage of pancreatic cancer, the early stage represents the early stage of pancreatic cancer, the benign represents benign diseases, and the healthy represents the healthy population.

[0068] Figure 9 ROC curve of pancreatic cancer of the pancreatic cancer diagnosis model constructed based on the combination of 2 paired miRNAs of the first preferred combination and the pancreatic cancer index CA19-9 in Example 5 in different datasets.

[0069] Figure 10 ROC curve of pancreatic cancer of the pancreatic cancer diagnosis model constructed based on the combination of 4 paired miRNAs of the second preferred combination and the pancreatic cancer index CA19-9 in Example 6 in different datasets.

[0070] Figure 11 Box plot of the CT value differences of the 3 paired miRNAs in Comparative Example 1 for RT-qPCR, where the middle and late stage represents the middle and late stage of pancreatic cancer, the early stage represents the early stage of pancreatic cancer, the benign represents benign diseases, and the healthy represents the healthy population.

[0071] Figure 12 ROC curve of pancreatic cancer of the pancreatic cancer diagnosis model constructed based on the 3 paired miRNAs of the third combination in Comparative Example 1 in different datasets.

[0072] Figure 13 Box plot of the expression levels of hsa-miR-98 and hsa-miR-25 corrected by the internal reference gene U6 in Comparative Example 2 (CT value differences of RT-qPCR), where the middle and late stage represents the middle and late stage of pancreatic cancer, the early stage represents the early stage of pancreatic cancer, the benign represents benign diseases, and the healthy represents the healthy population. Detailed implementation manners

[0073] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Unless otherwise clearly stated, in the entire specification and claims, the term "comprise" or its variations such as "comprises" or "including" etc. shall be understood to include the stated elements or components, without excluding other elements or other components.

[0074] In addition, for better illustration of the present invention, numerous specific details are given in the following specific embodiments. Those skilled in the art should understand that the present invention can also be implemented without some specific details. In some embodiments, the raw materials, components, methods, means, etc. well-known to those skilled in the art are not described in detail to highlight the gist of the present invention.

[0075] The reagents, reagent kits, raw materials, equipment, etc. used in the following embodiments can be obtained through commercial channels without special instructions. The experimental or detection methods involved in the present invention are conventional experimental or detection methods in the art without special instructions, or are carried out with reference to the corresponding reagent kits or product specifications.

[0076] Example 1: Screening of characteristic miRNA pairs for pancreatic cancer based on serum extracellular vesicle data

[0077] 1. Collection and preparation of clinical data and samples

[0078] Serum samples of pancreatic cancer patients, patients with benign diseases, and healthy people were collected respectively, and the detailed information is shown in Table 1.

[0079] All pancreatic cancer patients were pathologically diagnosed with pancreatic cancer. Other systemic tumors, tumors metastasized to the pancreas, and other underlying diseases were excluded.

[0080] Patients with benign diseases include patients with pancreatitis, pancreatic benign tumors, cholecystolithiasis, etc.

[0081] Table 1 Statistical table of selected cases:

[0082]

[0083] In Table 1, the data of the NGS cohort and the data of the RT-qPCR confirmation cohort were the data obtained by the applicant's self-detection of the serum collected from the hospital.

[0084] In Table 1, the data of RT-qPCR validation cohorts 1 and 2 (including the detection data of miRNAs and CA19-9) were from two different hospitals respectively, and were the data obtained by entrusting the hospitals to conduct tests with the detection reagents provided by Beijing Hotgen Biotech Co., Ltd.

[0085] That is, the data of the RT-qPCR confirmation cohort, RT-qPCR validation cohort 1, and RT-qPCR validation cohort 2 were from different centers.

[0086] In Table 1, RT-qPCR validation cohort 1&2 is the combined dataset of RT-qPCR validation cohorts 1 and 2.

[0087] 2. Isolation and purification of extracellular vesicles

[0088] Venous blood was collected using a vacuum blood collection tube (coagulant / separator gel), and serum was separated by centrifugation at 2000×g for 10 min at 10℃ - 30℃ within 6 hours after blood collection. The serum samples after the above treatment should be immediately used for detection, or stored at -70℃ or below for no more than 12 months, and repeated freezing and thawing are strictly prohibited. GlyExo-Capture extracellular vesicle extraction reagent was used to isolate extracellular vesicles, and the magnetic beads used were "a lectin-magnetic carrier conjugate for isolating glycosylated extracellular vesicles from clinical samples" (Application No.: 202010060055.7).

[0089] 3. Extraction of extracellular vesicle miRNAs and NGS sequencing

[0090] Use The Mini Kit kit was used to extract the total RNA of extracellular vesicles, and the high-sensitivity RNA kit of the Qsep100 fully automatic nucleic acid analysis system was used to evaluate it. Then, using the Illumina NEBNext smallRNA library preparation kit, the extracellular vesicle small RNAs were transcribed into cDNA libraries, and the E-Gel Power Snap electrophoresis system and E-Gel SizeSelect II gel were used to select and purify the libraries of the target size fragments. After checking the quality and concentration of the cDNA libraries, 75nt, single-end sequencing was performed on the Illumina NextSeq 550 sequencing system, and the sequencing data volume of a single library was greater than 10M reads.

[0091] 4. Construct a miRNA interaction network, analyze the NGS sequencing data, and obtain characteristic miRNA pairs of pancreatic cancer through single-factor screening and genetic algorithm screening

[0092] (1) Construction of the miRNA interaction network

[0093] 1) Retrieve the target sites of miRNAs from the miRTarBase database;

[0094] 2) Based on the target site information of miRNAs obtained in step 1), screen for transcription factors that can serve as target sites of miRNAs from the human transcription factor databases hTFtarget and AnimalTFDB;

[0095] 3) Based on the transcription factors screened in step 2), obtain the miRNAs they further regulate through bioinformatics methods or public databases, thereby constructing the miRNA-TF-miRNA interaction relationship and obtaining the miRNA interaction network.

[0096] (2) Analyze NGS sequencing data and prepare miRNA quantification data

[0097] By analyzing the above NGS sequencing data, obtain the miRNA quantification data of each sample in the comprehensive sample set containing disease samples and control samples:

[0098] First, raw sequencing files in fastq format are required for quality control; for the quality control results, use the cutadapt software to remove adapters and low-quality reads; for the data passing quality control, use the exceRpt small RNA analysis pipeline for annotation and quantification to obtain an expression matrix; according to the expression levels in the expression matrix, filter out miRNAs with too low counts (specific cases need to be analyzed specifically) to obtain the corrected miRNA quantification data.

[0099] (3) Construction of miRNA pair expression ratio features

[0100] Based on the miRNA interaction network constructed above and the prepared miRNA quantification data, calculate the expression ratio of miRNA pairs with interaction relationships in each sample.

[0101] To avoid the phenomenon of a denominator of 0, when calculating the expression ratio of miRNA pairs, uniformly add "1" to the denominator, and the calculation formula is as follows:

[0102] miRNA_a / miRNA_b = counts a / (count b + 1)

[0103] (4) Screening of characteristic miRNA pairs

[0104] Through univariate screening and genetic algorithm screening, obtain the characteristic miRNA pairs of pancreatic cancer, and its process is basically as follows:

[0105] i) Single-factor screening: Compare the expression quantity ratios of each miRNA pair in the disease biological sample group relative to the normal biological sample group. Calculate the p-value using scipy.stats.ttest_ind in Python, and correct the p-value using statsmodels.stats.multitest.fdrcorrection. Based on the corrected p-value p-adjusted, perform screening based on the threshold of 0.05;

[0106] ii) For the screened miRNA pairs, calculate the logarithm of the fold change of their expression quantity ratios in the disease biological sample group relative to the expression quantity ratios in the normal biological sample group, that is, log2FoldChange, and select an appropriate threshold for log2FoldChange according to the actual situation to further screen suitable targets;

[0107] iii) Use the genetic algorithm to perform further screening by looping 100 times.

[0108] Collect the serum samples of 78 pancreatic cancer patients and 38 patients with benign diseases as the NGS cohort (as shown in Table 1). Measure the miRNA expression levels in glycosylated extracellular vesicles using NGS and construct paired miRNAs. After the above screening procedure, 39 paired miRNA expression quantity ratios are screened. And a heatmap of the expression quantity ratios of these 39 paired miRNAs on the NGS cohort is drawn, as shown in Figure 1 .

[0109] Figure 1 It shows that there is an obvious clustering effect between cancer samples and benign samples on the heatmap, that is, the expression quantity ratios of 39 paired miRNAs perform well on the NGS cohort.

[0110] Example 2: Verification of 39 paired miRNAs in RT-qPCR data

[0111] Collect the serum samples of 103 pancreatic cancer patients, 66 patients with benign diseases, and 56 healthy people as the RT-qPCR confirmation cohort (as shown in Table 1). Measure the CT values of the 39 paired miRNAs obtained in Example 1 using RT-qPCR to confirm whether the difference still exists at the RT-qPCR level.

[0112] The method of RT-qPCR is as follows:

[0113] 1) Use Mini Kit kit to extract the total RNA of extracellular vesicles, and perform first-strand cDNA synthesis by reverse transcription. The reaction system and reaction conditions are as follows:

[0114] Reaction system and volume:

[0115] Reaction system Volume (μL) Reverse transcription primer (20 μM) 1 5× Reverse transcription buffer 2 Poly A polymerase (5 U / μL) 0.5 Reverse transcriptase (200 U / μL) 0.5 ATP (10 mM) 0.5 dNTP (10 mM) 0.5 RNA template 5 Total volume 10 。

[0116] Reaction conditions and time:

[0117]

[0118]

[0119] Then, the miRNA RT-qPCR reaction was carried out on the ABI 7500 real-time fluorescence quantitative PCR system, and the reaction system and reaction conditions are as follows:

[0120] Reaction system and volume:

[0121] Reaction system Volume (μL) 2× RT-qPCR reaction solution 12.5 Forward primer (10 μM) 2 Reverse primer (10 μM) 2 Probe (10 μM) 1 Rox 0.5 Nuclease-free water 2 cDNA 5 Total volume 25 。

[0122] Reaction conditions and time:

[0123]

[0124] Among them, the primers can be designed according to conventional methods to design the corresponding primers.

[0125] The LDA diagram was drawn through the CT value detection results of RT-qPCR. The results of the LDA diagram of 39 paired miRNAs in the PCR data are as Figure 2 , it can be seen that these 39 paired miRNAs can significantly distinguish pancreatic cancer, benign diseases, and healthy samples. Therefore, these.

[0126] Example 3: Further screening and verification of characteristic miRNA pairs for pancreatic cancer (preferred combination one)

[0127] Based on the pancreatic cancer, benign disease, and healthy sample of the RT-qPCR confirmation cohort, a recursive feature elimination was performed on 39 paired miRNAs using a logistic regression model (using the diagnostic efficacy as the screening index, and screening out the marker combination with the optimal diagnostic performance by removing one or more markers), and 2 paired miRNAs (preferred combination one) were obtained:

[0128] hsa-miR-98 / hsa-miR-25

[0129] hsa-miR-378i / hsa-miR-150

[0130] Using the RT-qPCR quantitative data (the RT-qPCR method refers to Example 2), a box plot display of the preferred combination one (2 paired miRNAs) was performed.

[0131] Among them, in the RT-qPCR program of the preferred combination 1, the primer sequences used are as follows:

[0132]

[0133] The box plot is as Figure 3 shown: Pancreatic cancer has significant differences compared with both benign diseases and healthy samples (p < 0.05), especially in the early stage of pancreatic cancer, which also has significant differences compared with benign diseases and healthy samples (p < 0.05).

[0134] Based on this preferred combination 1 (hsa-miR-98 / hsa-miR-25, hsa-miR-378i / hsa-miR-150), a diagnostic model for pancreatic cancer was constructed and its diagnostic performance was verified: The RT-qPCR confirmation cohort shown in Table 1 was divided into a training set (24 cases of advanced pancreatic cancer, 38 cases of early pancreatic cancer, 39 cases of benign diseases, 34 cases of health) and a test set (16 cases of advanced pancreatic cancer, 25 cases of early pancreatic cancer, 27 cases of benign diseases, 22 cases of health) according to 6:4. A model was established using logistic regression on the training set, verified on the test set, and also verified on other external validation sets (RT-qPCR validation cohorts 1, 2, 1&2).

[0135] A logistic regression model was established for the 2 paired miRNAs (combination 1) in the training set, and the cutoff was determined following the principle of maximizing the log-likelihood index of the training set. The overall performance of the model was evaluated by the AUC, sensitivity, and specificity of the ROC curve.

[0136] The results of the AUC, sensitivity, and specificity of the ROC curve are as Figure 4 shown.

[0137] Figure 4The results showed that for the optimal combination 1 (hsa-miR-98 / hsa-miR-25, hsa-miR-378i / hsa-miR-150) in differentiating pancreatic cancer from other samples (benign diseases, healthy), the AUC values in the training set, test set, validation set 1 (i.e., validation cohort 1), validation set 2 (i.e., validation cohort 2), and validation set 1&2 (i.e., validation cohort 1&2) could reach 0.867 (sensitivity 0.806, specificity 0.808), 0.963 (sensitivity 0.878, specificity 0.918), 0.941 (sensitivity 0.849, specificity 0.898), 0.927 (sensitivity 0.848, specificity 0.845), and 0.934 (sensitivity 0.848, specificity 0.873) respectively. This indicated that the optimal combination 1 (hsa-miR-98 / hsa-miR-25, hsa-miR-378i / hsa-miR-150) not only had excellent diagnostic efficacy but also had stable diagnostic efficacy in the results of different data centers, without the situation of unstable results.

[0138] To better evaluate the diagnostic efficacy of the model for early-stage pancreatic cancer, under the condition of the optimal cutoff, the diagnostic sensitivity and specificity data for advanced and early-stage pancreatic cancer in samples from different datasets were obtained respectively, and the results are shown in Table 2.

[0139] Table 2. Sensitivity and specificity of the optimal combination 1 in different datasets

[0140]

[0141] The results in Table 2 showed that in the evaluation model, the recognition sensitivities for advanced and early-stage pancreatic cancer in different datasets were relatively high. The recognition sensitivity for advanced-stage pancreatic cancer was above 0.750 (i.e., 75.0%), and the recognition sensitivity for early-stage pancreatic cancer was above 0.763 (i.e., 76.3%). It could specifically distinguish other samples (benign diseases, healthy), and the diagnostic accuracy of pancreatic cancer was relatively high.

[0142] The box plot of the sample score distribution in this embodiment is as Figure 5 shown, and the results showed that there were significant differences between advanced and early-stage pancreatic cancer and other samples (benign diseases, healthy), and the differentiation effect was good.

[0143] In summary, it was shown that the optimal combination 1 (hsa-miR-98 / hsa-miR-25, hsa-miR-378i / hsa-miR-150) of this application could effectively distinguish pancreatic cancer from other samples (benign diseases, healthy), and the recognition sensitivity for early-stage pancreatic cancer was also very high. It presented stable diagnostic results in datasets from different sources, could improve the accuracy of early pancreatic cancer, and was of great significance for the diagnosis, treatment efficacy, or prognosis evaluation of pancreatic cancer.

[0144] Example 4: Further screening and verification of characteristic miRNAs for pancreatic cancer (preferred combination two)

[0145] Based on pancreatic cancer, benign, and healthy samples in the PCR confirmation cohort, 39 paired miRNAs were subjected to recursive feature elimination using an SVM model (using diagnostic efficacy as the screening index, and by removing one or more markers, the marker combination with the optimal diagnostic performance was screened out), resulting in 4 paired miRNAs (preferred combination two):

[0146] hsa-miR-501 / hsa-miR-150

[0147] hsa-miR-378i / hsa-miR-155

[0148] hsa-miR-193a / hsa-miR-155

[0149] hsa-miR-150 / hsa-miR-1180

[0150] Using RT-qPCR quantitative data (the RT-qPCR method refers to Example 2), box plots of the preferred combination two (4 paired miRNAs) were presented.

[0151] Among them, in the RT-qPCR program of the preferred combination two, the primer sequences used are as follows:

[0152]

[0153] The box plot is as Figure 6 shown: Pancreatic cancer showed significant differences compared to both benign diseases and healthy samples (p < 0.05).

[0154] Based on this preferred combination two (hsa-miR-501 / hsa-miR-150, hsa-miR-378i / hsa-miR-155, hsa-miR-193a / hsa-miR-155, hsa-miR-150 / hsa-miR-1180), a diagnostic model for pancreatic cancer was constructed and its diagnostic performance was verified: The RT-qPCR confirmation cohort shown in Table 1 was divided into a training set (24 patients with advanced pancreatic cancer, 38 patients with early pancreatic cancer, 39 patients with benign diseases, 34 healthy individuals) and a test set (16 patients with advanced pancreatic cancer, 25 patients with early pancreatic cancer, 27 patients with benign diseases, 22 healthy individuals) at a ratio of 6:4. A model was established using logistic regression on the training set, verified on the test set, and also verified on other external validation sets (RT-qPCR validation cohorts 1, 2, 1&2).

[0155] A logistic regression model was established for the 4 paired miRNAs (Combination 2) in the training set, and the cutoff was determined following the principle of maximizing the log-likelihood of the training set. The overall performance of the model was evaluated by the AUC, sensitivity, and specificity of the ROC curve.

[0156] The AUC, sensitivity, and specificity results of the ROC curve are as Figure 7 shown.

[0157] Figure 7 The results showed that for the optimized combination 2 (hsa-miR-501 / hsa-miR-150, hsa-miR-378i / hsa-miR-155, hsa-miR-193a / hsa-miR-155, hsa-miR-150 / hsa-miR-1180) in differentiating pancreatic cancer from other samples (benign diseases, healthy samples), the AUCs in the training set, test set, validation set 1 (i.e., validation cohort 1), validation set 2 (i.e., validation cohort 2), and validation set 1&2 (i.e., validation cohort 1&2) could reach 0.888 (sensitivity 0.806, specificity 0.849), 0.879 (sensitivity 0.780, specificity 0.816), 0.937 (sensitivity 0.935, specificity 0.805), 0.944 (sensitivity 0.834, specificity 0.900), and 0.931 (sensitivity 0.873, specificity 0.851), respectively. This indicated that the optimized combination 2 (hsa-miR-501 / hsa-miR-150, hsa-miR-378i / hsa-miR-155, hsa-miR-193a / hsa-miR-155, hsa-miR-150 / hsa-miR-1180) not only had excellent diagnostic efficacy but also had stable diagnostic efficacy in the results of different data centers, without the situation of unstable results.

[0158] To better evaluate the diagnostic efficacy of the model for early-stage pancreatic cancer, under the condition of the optimal cutoff, the sensitivity and specificity data for advanced and early-stage pancreatic cancer in the samples of different datasets were obtained respectively, and the results are shown in Table 3.

[0159] Table 3. Sensitivity and specificity of the optimized combination 2 in different datasets

[0160]

[0161] The results in Table 3 showed that in the evaluation model, the recognition sensitivities for advanced and early-stage pancreatic cancer in different datasets were relatively high. The recognition sensitivity for advanced-stage pancreatic cancer was above 0.875 (i.e., 87.5%), and the recognition sensitivity for early-stage pancreatic cancer was as high as above 0.72 (i.e., 72%). It could specifically distinguish other samples (benign diseases, health), and the correct diagnosis rate for early-stage pancreatic cancer was relatively high.

[0162] The box plot of the sample score distribution in this embodiment is as follows Figure 8 shown. The results show that there are significant differences in the model scores among the middle and late stages, early stage of pancreatic cancer and other samples (benign diseases, healthy), and the discrimination effect is good.

[0163] In summary, it shows that the second preferred combination of the present application (hsa-miR-501 / hsa-miR-150, hsa-miR-378i / hsa-miR-155, hsa-miR-193a / hsa-miR-155, hsa-miR-150 / hsa-miR-1180)) has a very high recognition sensitivity for early pancreatic cancer, and presents stable diagnostic results in datasets from different sources, which can improve the accuracy of early pancreatic cancer and is of great significance for the diagnosis, curative effect or prognosis evaluation of pancreatic cancer.

[0164] Example 5: Construct a pancreatic cancer diagnosis model based on the first preferred combination of paired miRNAs and the pancreatic cancer index CA19-9, and verify its diagnostic performance

[0165] In this embodiment, the first preferred combination (hsa-miR-98 / hsa-miR-25, hsa-miR-378i / hsa-miR-150) and the pancreatic cancer index CA19-9 are used for joint modeling (wherein, the expression level of CA19-9 is detected by a carbohydrate antigen 19-9 assay kit (magnetic particle chemiluminescence immunoassay method), and this kit is from Beijing Hotgen Biotech Co., Ltd. and can be purchased externally). The sample information of the training set and the validation set used is the same as that in Example 3 (the detection of CA19-9 is also carried out by the applicant providing the kit and entrusting the corresponding hospital for detection). The specific modeling method is as follows:

[0166] The RT-qPCR confirmation cohort shown in Table 1 is divided into a training set (24 cases of middle and late stage pancreatic cancer, 38 cases of early stage pancreatic cancer, 40 cases of benign diseases, 34 cases of healthy) and a test set (16 cases of middle and late stage pancreatic cancer, 25 cases of early stage pancreatic cancer, 26 cases of benign diseases, 22 cases of healthy) according to a ratio of 6:4. A model is established using logistic regression on the training set, verified on the test set, and also verified on other external validation sets (RT-qPCR validation cohorts 1, 2, 1&2).

[0167] A logistic regression model is established for the 2 paired miRNAs (combination one) and the pancreatic cancer index CA19-9 in the training set, and the cutoff is determined following the principle of maximizing the Youden index of the training set. The overall performance of the model is evaluated by the AUC, sensitivity and specificity of the ROC curve.

[0168] The results of the AUC, sensitivity and specificity of the ROC curve are as follows Figure 9as shown

[0169] Figure 9 The results showed that when the optimal combination I (hsa-miR-98 / hsa-miR-25, hsa-miR-378i / hsa-miR-150) was combined with the pancreatic cancer index CA19-9 to distinguish pancreatic cancer from other samples (benign diseases, healthy), the AUCs in the training set, test set, validation set 1 (i.e., validation cohort 1), validation set 2 (i.e., validation cohort 2), and validation set 1&2 (i.e., validation cohort 1&2) could reach 0.942 (sensitivity 0.839, specificity 0.959), 0.975 (sensitivity 0.902, specificity 0.896), 0.958 (sensitivity 0.839, specificity 0.924), 0.965 (sensitivity 0.795, specificity 0.936), and 0.961 (sensitivity 0.811, specificity 0.930) respectively. This indicated that the optimal combination I (hsa-miR-98 / hsa-miR-25, hsa-miR-378i / hsa-miR-150) and the pancreatic cancer index CA19-9 not only further improved the diagnostic efficiency, but also had stable diagnostic efficiency in the results of different data centers, without the situation of unstable results.

[0170] To better evaluate the diagnostic efficiency of the model for early pancreatic cancer, under the condition of the optimal cutoff, the diagnostic sensitivity and specificity data of advanced and early pancreatic cancer in samples of different datasets were obtained respectively, and the results are shown in Table 4.

[0171] Table 4. Sensitivity and specificity of the combined diagnosis of the optimal combination I and the pancreatic cancer index CA19-9 in different datasets

[0172]

[0173]

[0174] The results in Table 4 showed that in the evaluation model of the combined diagnosis of the optimal combination I and the pancreatic cancer index CA19-9, the recognition sensitivities of advanced and early pancreatic cancer in different datasets were relatively high. The recognition sensitivity for advanced pancreatic cancer was above 0.811 (i.e., 81.1%), and the recognition sensitivity for early pancreatic cancer was above 0.75 (i.e., 75.0%), which could specifically distinguish other samples (benign diseases, healthy), and the diagnostic accuracy of pancreatic cancer was relatively high.

[0175] Example 6: Construct a pancreatic cancer diagnostic model based on the paired miRNA optimal combination II and the pancreatic cancer index CA19-9, and verify its diagnostic performance

[0176] In this embodiment, a combined model is established using the preferred combination two (hsa-miR-501 / hsa-miR-150, hsa-miR-378i / hsa-miR-155, hsa-miR-193a / hsa-miR-155, hsa-miR-150 / hsa-miR-1180) and the pancreatic cancer indicator CA19-9 (wherein the expression level of CA19-9 is detected using a carbohydrate antigen 19-9 assay kit (magnetic particle chemiluminescence immunoassay method), and this kit is from Beijing Hotgen Biotech Co., Ltd. and can be purchased externally). The specific modeling method is as follows:

[0177] The RT-qPCR confirmation cohort shown in Table 1 is divided into a training set (24 cases of advanced and middle-stage pancreatic cancer, 38 cases of early-stage pancreatic cancer, 40 cases of benign diseases, 34 cases of health) and a test set (16 cases of advanced and middle-stage pancreatic cancer, 25 cases of early-stage pancreatic cancer, 26 cases of benign diseases, 22 cases of health) according to a ratio of 6:4. A model is established using logistic regression on the training set, verified on the test set, and also verified on other external validation sets (RT-qPCR validation cohorts 1, 2, 1&2).

[0178] A logistic regression model is established for the two paired miRNAs (combination two) in the training set and the pancreatic cancer indicator CA19-9, and the cutoff is determined following the principle of maximizing the training set's log-likelihood index. The overall performance of the model is evaluated by the AUC, sensitivity, and specificity of the ROC curve.

[0179] The results of the AUC, sensitivity, and specificity of the ROC curve are as Figure 10 shown.

[0180] Figure 10The results showed that when the optimal combination two (hsa-miR-501 / hsa-miR-150, hsa-miR-378i / hsa-miR-155, hsa-miR-193a / hsa-miR-155, hsa-miR-150 / hsa-miR-1180) was combined with the pancreatic cancer index CA19-9 to distinguish pancreatic cancer from other samples (benign diseases, healthy), the AUCs in the training set, test set, validation set 1 (i.e., validation cohort 1), validation set 2 (i.e., validation cohort 2), and validation set 1&2 (i.e., validation cohort 1&2) could reach 0.963 (sensitivity 0.952, specificity 0.878), 0.970 (sensitivity 1.000, specificity 0.792), 0.955 (sensitivity 0.946, specificity 0.814), 0.978 (sensitivity 0.934, specificity 0.891), and 0.964 (sensitivity 0.939, specificity 0.851), respectively. This indicated that the optimal combination two (hsa-miR-501 / hsa-miR-150, hsa-miR-378i / hsa-miR-155, hsa-miR-193a / hsa-miR-155, hsa-miR-150 / hsa-miR-1180) and the pancreatic cancer index CA19-9 not only further improved the diagnostic efficacy, but also had stable diagnostic efficacy in the results of different data centers, without the situation of unstable results.

[0181] To better evaluate the diagnostic efficacy of the model for early-stage pancreatic cancer, under the condition of the optimal cutoff, the diagnostic sensitivity and specificity data of advanced and early-stage pancreatic cancer in samples from different datasets were obtained respectively, and the results are shown in Table 5.

[0182] Table 5. Sensitivity and specificity of the combined diagnosis of the optimal combination two and the pancreatic cancer index CA19-9 in different datasets

[0183]

[0184] The results in Table 5 showed that in the evaluation model of the combined diagnosis of the optimal combination two and the pancreatic cancer index CA19-9, the recognition sensitivities of advanced and early-stage pancreatic cancer in different datasets were relatively high. The recognition sensitivity for advanced-stage pancreatic cancer was above 0.937 (i.e., 93.7%), and the recognition sensitivity for early-stage pancreatic cancer was above 0.925 (i.e., 92.5%). It could specifically distinguish other samples (benign diseases, healthy), and the diagnostic accuracy of pancreatic cancer was relatively high.

[0185] Comparative example 1:

[0186] Although the advent of big data has improved the convenience of biomarker screening, there are still great challenges in screening biomarker combinations applicable to different data centers (different hospitals). Most of the time, biomarker combinations with good diagnostic performance obtained through screening are difficult to apply to other data centers (other hospitals or institutions). For example, the following combination three of paired miRNAs screened by the applicant:

[0187] hsa-miR-501 / hsa-miR-25

[0188] hsa-miR-20b / hsa-miR-150

[0189] hsa-miR-193a / hsa-miR-150

[0190] Using RT-qPCR quantitative data (the RT-qPCR method refers to Example 2), a box plot of combination three (3 paired miRNAs) was presented.

[0191] Among them, in the RT-qPCR program of the preferred combination three, the primer sequences used are as follows:

[0192]

[0193]

[0194] The box plot is as Figure 11 shown: Pancreatic cancer has significant differences compared to both benign diseases and healthy samples (p < 0.05).

[0195] Based on this combination three (hsa-miR-501 / hsa-miR-25, hsa-miR-20b / hsa-miR-150, hsa-miR-193a / hsa-miR-150), a pancreatic cancer diagnosis model was constructed and its diagnostic performance was verified: The RT-qPCR confirmation cohort shown in Table 1 was divided into a training set (24 patients with advanced pancreatic cancer, 38 patients with early pancreatic cancer, 39 patients with benign diseases, 34 healthy individuals) and a test set (16 patients with advanced pancreatic cancer, 25 patients with early pancreatic cancer, 27 patients with benign diseases, 22 healthy individuals) according to a ratio of 6:4. A model was established using logistic regression on the training set, verified on the test set, and also verified on other external validation sets (RT-qPCR validation cohorts 1, 2, 1&2).

[0196] A logistic regression model was established for the 3 paired miRNAs (combination three) in the training set, and the cutoff was determined following the principle of maximizing the log-likelihood index of the training set. The overall performance of the model was evaluated by the AUC, sensitivity, and specificity of the ROC curve.

[0197] The AUC, sensitivity, and specificity results of the ROC curve are as Figure 12 shown.

[0198] Figure 12 The results show that for Panel 3 (hsa-miR-501 / hsa-miR-25, hsa-miR-20b / hsa-miR-150, hsa-miR-193a / hsa-miR-150) in differentiating pancreatic cancer from other samples (benign diseases, healthy), the AUCs in the training set and test set can reach 0.886 (sensitivity 0.871, specificity 0.795) and 0.881 (sensitivity 0.829, specificity 0.776) respectively, that is, it shows good diagnostic performance in one data center (RT-qPCR confirmation cohort). However, in other data centers, namely in Validation Set 1 (i.e., Validation Cohort 1), Validation Set 2 (i.e., Validation Cohort 2), and Validation Set 1&2 (i.e., Validation Cohort 1&2), the AUCs are only 0.607 (sensitivity 0.591, specificity 0.669), 0.696 (sensitivity 0.589, specificity 0.800), and 0.717 (sensitivity 0.669, specificity 0.751) respectively. This indicates that although Panel 3 (hsa-miR-501 / hsa-miR-25, hsa-miR-20b / hsa-miR-150, hsa-miR-193a / hsa-miR-150) shows good diagnostic performance in both the randomly divided training set and test set in one data center (RT-qPCR confirmation cohort), its diagnostic performance significantly decreases in other data centers (Validation Sets 1 and 2, i.e., RT-qPCR Validation Cohorts 1 and 2 from other hospitals for third-party testing), that is, the diagnostic results are unstable and it is difficult to apply to clinical practice.

[0199] The diagnostic efficacy of the model for early-stage pancreatic cancer was also evaluated. Under the condition of the optimal cutoff, the sensitivity and specificity data for advanced and early-stage pancreatic cancer in samples from different datasets were obtained respectively, and the results are shown in Table 6.

[0200] Table 6. Sensitivity and specificity of Panel 3 in different datasets

[0201]

[0202] The results in Table 4 show that in the evaluation model, the sensitivity and specificity of advanced and early-stage pancreatic cancer in the randomly divided training set and test set in the same data center are relatively high. Especially, the recognition sensitivity of advanced-stage pancreatic cancer in the training set is as high as 0.958 (95.8%), but the sensitivity significantly decreases in other data centers (Validation Sets 1 and 2 from other hospitals for third-party testing) (the recognition sensitivity is as low as below 0.5), and the diagnostic accuracy also decreases significantly, that is, the diagnostic results are unstable and it is difficult to apply to clinical practice.

[0203] Comparative Example 2:

[0204] As genes with relatively constant expression in various tissues and cells, internal reference genes are generally used as references for detecting the expression levels of target genes and play a certain calibration role. Commonly used internal reference genes include 18S rRNA, 28S rRNA, U6, etc. Although their expression is relatively constant, their expression may be unstable under different conditions such as physiological or pathological changes, which may have a great impact on the analysis of the expression levels of target genes, resulting in great difficulties in the selection and application of diagnostic markers. For example, the miRNA pair of hsa-miR-98 / hsa-miR-25 in Example 3 can significantly distinguish pancreatic cancer from other samples (see Figure 3 the left box plot), but when the conventional internal reference gene U6 is used to correct hsa-miR-98 and hsa-miR-25 respectively, and hsa-miR-98 and hsa-miR-25 are used as markers respectively, it will be found that: hsa-miR-98 can distinguish pancreatic cancer from healthy samples (p<0.05), but cannot distinguish pancreatic cancer from benign pancreatic diseases; hsa-miR-25 can distinguish pancreatic cancer from benign pancreatic diseases (p<0.05), but cannot distinguish pancreatic cancer from healthy samples (see Figure 13 the box plot shown). Specifically:

[0205] Using RT-qPCR quantitative data (the RT-qPCR method refers to Example 2), box plots of hsa-miR-98, hsa-miR-25, and hsa-miR-98 / hsa-miR-25 were shown. Among them, the box plots of hsa-miR-98 and hsa-miR-25 representing the expression levels of hsa-miR-98 and hsa-miR-25 corrected by the internal reference gene U6 are as shown in Figure 13 and the box plot of hsa-miR-98 / hsa-miR-25 is as shown in Figure 3 the left figure.

[0206] Among them, the primer sequences of hsa-miR-98, hsa-miR-25, and U6 are as follows:

[0207]

[0208] Figure 13The results showed that when the expression levels of hsa-miR-98 and hsa-miR-25 were calibrated with U6 as the internal reference gene, the expression levels in pancreatic cancer samples were sometimes not significantly correlated with those in other samples (for example, hsa-miR-98 could not significantly distinguish pancreatic cancer patients from benign patients, and hsa-miR-25 could not distinguish pancreatic cancer patients from healthy individuals), indicating that the diagnostic performance was unstable when calibrating the expression levels of hsa-miR-98 and hsa-miR-25 with a common internal reference gene (such as U6).

[0209] In summary, the preferred combination I (hsa-miR-98 / hsa-miR-25, hsa-miR-378i / hsa-miR-150) and the preferred combination II (hsa-miR-501 / hsa-miR-150, hsa-miR-378i / hsa-miR-155, hsa-miR-193a / hsa-miR-155, hsa-miR-150 / hsa-miR-1180) screened in the present invention were obtained by continuously adjusting the screening method or screening conditions. They are biomarker combinations that are applicable to different data centers, have good diagnostic performance and diagnostic stability, and can further improve the diagnostic performance when combined with the pancreatic cancer index CA19-9, showing good clinical application prospects.

[0210] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A diagnostic marker combination for pancreatic cancer, characterized in that, The diagnostic marker for pancreatic cancer consists of a combination of the following miRNA pairs: hsa-miR-501 / hsa-miR-150; hsa-miR-378i / hsa-miR-155; hsa-miR-193a / hsa-miR-155; hsa-miR-150 / hsa-miR-1180.

2. A diagnostic marker combination for pancreatic cancer, characterized in that, The diagnostic marker for pancreatic cancer consists of CA19-9 protein and a combination of the following miRNA pairs: hsa-miR-501 / hsa-miR-150; hsa-miR-378i / hsa-miR-155; hsa-miR-193a / hsa-miR-155; hsa-miR-150 / hsa-miR-1180.

3. The diagnostic marker combination for pancreatic cancer according to claim 1 or 2, characterized in that, Among the miRNA pairs, hsa-miR-150 is used to calibrate the relative expression level of hsa-miR-501; hsa-miR-155 is used to calibrate the relative expression level of hsa-miR-378i; hsa-miR-155 is used to calibrate the relative expression level of hsa-miR-193a; hsa-miR-1180 is used to calibrate the relative expression level of hsa-miR-150.

4. The diagnostic marker combination for pancreatic cancer according to claim 1 or 2, characterized in that, The miRNA in the miRNA pair is the miRNA in serum or plasma.

5. The diagnostic marker combination for pancreatic cancer according to claim 1 or 2, characterized in that, The miRNA in the miRNA pair is the miRNA in serum or plasma glycosylated extracellular vesicles.

6. An analysis system for predicting pancreatic cancer, characterized in that, It includes a data analysis module, which is used to analyze the expression levels of the marker combinations of the target object to be predicted and calculate the ratio of the expression levels of the miRNA pairs or directly use the ratio of the expression levels of the miRNA pairs as input features, and is also used to calculate the prediction value of whether the target object is pancreatic cancer according to the ratio of the expression levels of the miRNA pairs; The input features consist of a combination of the ratios of the expression levels of the following miRNA pairs: hsa-miR-501 / hsa-miR-150; hsa-miR-378i / hsa-miR-155; hsa-miR-193a / hsa-miR-155; hsa-miR-150 / hsa-miR-1180.

7. The analysis system according to claim 6, characterized in that, The data analysis module stores a prediction model for outputting the probability value of pancreatic cancer. Among them, the construction method of the prediction model includes: using the ratios of the expression levels of the miRNA pairs in pancreatic cancer, benign, and healthy samples as input variables, and whether it is pancreatic cancer as the dependent variable, calculating through multiple logistic regression to obtain the prediction model, and testing it with a validation set.

8. An analysis system for predicting pancreatic cancer, characterized in that, It includes a data analysis module, which is used to analyze the expression levels of the marker combinations of the target object to be predicted and calculate the ratio of the expression levels of the miRNA pairs or directly use the ratio of the expression levels of the miRNA pairs as input features, and is also used to calculate the prediction value of whether the target object is pancreatic cancer according to the ratio of the expression levels of the miRNA pairs and the expression level of CA19-9 protein; The input features consist of the expression level of CA19-9 protein of the target object to be predicted and the ratio of the expression levels of miRNA pairs in the following combinations: hsa-miR-501 / hsa-miR-150; hsa-miR-378i / hsa-miR-155; hsa-miR-193a / hsa-miR-155; hsa-miR-150 / hsa-miR-1180.

9. The analysis system according to claim 8, characterized in that, A prediction model for outputting the probability value of pancreatic cancer is stored in the data analysis module. Among them, the construction method of the prediction model includes: using the expression levels of CA19-9 protein in pancreatic cancer, benign, and healthy samples and the ratio of the expression levels of the miRNA pairs as input variables, and whether it is pancreatic cancer as the dependent variable, calculating through multiple logistic regression to obtain the prediction model, and testing it with a validation set.

10. The analysis system according to any one of claims 6 to 9, characterized in that, When the detection method is RT-qPCR, the ratio of the expression levels of the miRNA pair is the CT difference of the miRNA pair.

11. Use of a reagent for detecting the expression level of the diagnostic marker combination for pancreatic cancer according to any one of claims 1-5 in the preparation of a kit or an analysis system for diagnosing pancreatic cancer.

12. Use of a reagent for detecting the expression level of the diagnostic marker combination for pancreatic cancer according to any one of claims 1-5 in serum or plasma glycosylated extracellular vesicles in the preparation of a kit or an analysis system for diagnosing pancreatic cancer.

13. A kit for the diagnosis of pancreatic cancer, characterized in that, The kit includes a reagent for detecting the expression level of the diagnostic marker combination for pancreatic cancer according to any one of claims 1-5.

14. The kit according to claim 13, wherein The detection reagent includes primers and probes for the corresponding miRNAs. Among them, the primer and probe sequences for each miRNA are as follows: The nucleotide sequence of the upstream primer of hsa-miR-150 is as shown in SEQ ID NO:1; The nucleotide sequence of the upstream primer of hsa-miR-378i is as shown in SEQ ID NO:3; The nucleotide sequence of the upstream primer of hsa-miR-1180 is as shown in SEQ ID NO:6; The nucleotide sequence of the upstream primer of hsa-miR-501 is as shown in SEQ ID NO:7; The nucleotide sequence of the upstream primer of hsa-miR-193a is as shown in SEQ ID NO:8; The nucleotide sequence of the upstream primer of hsa-miR-155 is as shown in SEQ ID NO:9; The nucleotide sequence of the probe is as shown in SEQ ID NO:5; The reverse primer is a common downstream primer, and its nucleotide sequence is as shown in SEQ ID NO:

10.

15. The kit according to claim 14, wherein It does not contain primers and probes for conventional internal reference genes.

Citation Information

Patent Citations

  • Lectin-magnetic carrier coupled complex used for separating glycosylated exosomes in clinical samples

    CN111253491A

  • MiRNA group for early diagnosis and / or prognosis monitor of pancreatic caner and application of miRNA

    CN111172285A

  • Detection reagent of biomarker for early diagnosis of pancreatic cancer

    CN111748629A