Novel biomarker for diagnosing pancreatic cancer

A diagnostic composition using specific polypeptides and genes addresses the challenge of early pancreatic cancer detection, enhancing diagnostic accuracy and survival rates through blood-based mass spectrometry.

WO2025244187A1PCT designated stage Publication Date: 2025-11-27BERTIS INC
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/010792
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-22
Filing Date
2024-07-25
Publication Date
2025-11-27

AI Technical Summary

Technical Problem

Pancreatic cancer is challenging to detect early due to subtle symptoms and the lack of specific and sensitive biomarkers, leading to low survival rates.

Method used

Development of a diagnostic composition comprising specific polypeptides and their encoding genes, such as ANPEP, APOA4, APOC3, C9, CRP, HGFAC, IGFBP2, ITIH3, LRG1, ORM1, PFN1, PIGR, PON3, SERPINA3, and VWF, for early detection of pancreatic cancer through mass spectrometry methods using blood samples.

Benefits of technology

The proposed biomarkers enable accurate early diagnosis of pancreatic cancer, potentially increasing survival rates by identifying pancreatic ductal adenocarcinoma with high specificity and sensitivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024010792_27112025_PF_FP_ABST
    Figure KR2024010792_27112025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method for predicting the onset of pancreatic cancer with high accuracy by measuring the expression levels of genes or proteins associated with the development of pancreatic cancer. The present invention provides a novel, multifaceted, and comprehensive treatment strategy for the early diagnosis of pancreatic cancer, which includes a method for predicting the onset of pancreatic cancer at an early stage with high reliability by identifying effective biomarkers for pancreatic cancer, in particular pancreatic ductal adenocarcinoma. Ultimately, the present invention can be effectively applied to improve the survival rate of patients with pancreatic cancer.
Need to check novelty before this filing date? Find Prior Art

Description

A novel biomarker for the diagnosis of pancreatic cancer

[0001] The present invention relates to a method for predicting the onset of pancreatic cancer at an early stage with high accuracy by measuring the expression level of a factor specifically regulated in pancreatic cancer.

[0002] Pancreatic cancer is the seventh leading cause of cancer-related death worldwide, and ranks even higher in developed countries, ranking between the second and fifth leading causes. Despite its significant role in cancer-related death, pancreatic cancer remains challenging to detect and diagnose early due to its generally subtle symptoms and limited research on specific tumor markers. Due to the lack of early diagnostic methods, only 5 to 22% of pancreatic cancer cases are diagnosed early enough to be surgically resected. With a five-year survival rate of only 12.6% (2018 National Cancer Registry Statistics), pancreatic cancer is one of the most lethal cancers.

[0003] Currently, pancreatic cancer is commonly diagnosed clinically using imaging techniques such as ultrasound, CT scans, MRI, angiography, endoscopic retrograde cholangiopancreatography, and endoscopic ultrasound; as well as tumor markers (biomarkers). Carbohydrate antigen 19-9 (CA19-9) is the only pancreatic cancer biomarker known to have a proven diagnostic effect. However, this biomarker lacks sufficient sensitivity and specificity to effectively diagnose pancreatic cancer early. While CA19-9 is a useful tumor marker for predicting prognosis and tracking treatment progress in pancreatic cancer, its usefulness as a screening test is generally considered low due to the low incidence of pancreatic cancer. Therefore, to increase the survival rate of pancreatic cancer patients, the discovery of novel pancreatic cancer-specific biomarkers that enable early diagnosis with higher accuracy is urgently needed.

[0004] Meanwhile, blood is a promising biological sample for cancer biomarker research due to the minimally invasive nature of liquid biopsy. Therefore, the inventors of this study conducted a comparative analysis of blood samples from pancreatic cancer patients and controls to identify pancreatic cancer-specific biomarkers effective in the diagnosis of pancreatic cancer. Through this analysis, we aimed to identify biomarkers capable of predicting or diagnosing pancreatic cancer at an early stage with high accuracy, and to confirm their efficacy.

[0005] Numerous papers and patents are referenced and cited throughout this specification. The disclosures of these cited papers and patents are incorporated herein by reference in their entirety to provide a clearer understanding of the state of the art and the scope of the invention.

[0006]

[0007] The present inventors have diligently researched and developed effective diagnostic biomarkers for pancreatic cancer, a major cause of cancer-related death worldwide, but with a very low patient survival rate due to the lack of highly accurate early diagnosis methods. This ultimately aims to develop a new diagnostic method that can significantly reduce the mortality rate of pancreatic cancer, particularly pancreatic ductal adenocarcinoma. As a result, we have discovered proteins whose expression is specifically regulated in pancreatic cancer patients or the genes encoding them, and by measuring the expression levels of these markers, we have discovered that they are reliable markers that can not only diagnose pancreatic cancer early with high accuracy, but can also be applied to mass spectrometry diagnostic methods using blood, thereby completing the present invention.

[0008] Therefore, the purpose of the present invention is to provide a diagnostic composition for pancreatic cancer or a diagnostic kit containing the same.

[0009] Another object of the present invention is to provide a method for providing information necessary for diagnosing pancreatic cancer.

[0010] Another object of the present invention is to provide a method for screening a composition for preventing or treating pancreatic cancer.

[0011] Another object of the present invention is to provide a diagnosis system for pancreatic cancer.

[0012] Other objects and advantages of the present invention will become more apparent from the detailed description, claims and drawings below.

[0013]

[0014] According to one aspect of the present invention, the present invention provides a composition comprising at least one polypeptide selected from the group consisting of ANPEP (Aminopeptidase N), APOA4 (Apolipoprotein A-IV), APOC3 (Apolipoprotein C-III), C9 (Complement component C9), CRP (C-reactive protein), HGFAC (Hepatocyte growth factor activator), IGFBP2 (Insulin-like growth factor-binding protein 2), ITIH3 (Inter-alpha-trypsin inhibitor heavy chain H3), LRG1 (Leucine-rich alpha-2-glycoprotein), ORM1 (Alpha-1-acid glycoprotein 1), PFN1 (Profilin-1), PIGR (Polymeric immunoglobulin receptor), PON3 (Serum paraoxonase / lactonase 3), SERPINA3 (Alpha-1-antichymotrypsin), and VWF (von Willebrand factor), or a partial fragment thereof; Or, a composition for diagnosing pancreatic cancer is provided, which comprises as an active ingredient a preparation for measuring the expression level of a gene encoding the same.

[0015] The present inventors have diligently researched and developed effective diagnostic biomarkers for pancreatic cancer, a major cause of cancer-related death worldwide, but with a very low patient survival rate due to the lack of highly accurate early diagnosis methods. This ultimately aims to develop a new diagnostic method that can significantly reduce the mortality rate of pancreatic cancer, particularly pancreatic ductal adenocarcinoma. As a result, we have discovered proteins whose expression is specifically regulated in pancreatic cancer patients or the genes encoding them, and by measuring the expression levels of these markers, we have discovered that they are reliable markers that can not only diagnose pancreatic cancer early with high accuracy, but can also be applied to mass spectrometry diagnostic methods using blood, thereby completing the present invention.

[0016] All genes used for the diagnosis or early diagnosis of pancreatic cancer in this specification may be used as biomarkers for the diagnosis of pancreatic cancer independently or in combinations of two or more, and when used in combination, the set of genes may be a “biomarker panel.”

[0017] The term “biomarker panel” as used herein may also be referred to as “biomarker detection panel” and means a set of two or more biomarkers that can be used for detection; diagnosis; prognosis; staging; or monitoring of a disease or condition. The biomarker components of such a biomarker set may be physically associated by being packaged together or reversibly or irreversibly bound to a solid support. For example, the biomarker panel of the present invention may be provided through a separate tube that is sold or shipped together as part of a kit; or may be provided bound to the surface or interior of a chip, membrane, strip, filter, bead, particle, filament, fiber, gel, or matrix, a well of a multi-well plate, or other support, but is not limited to the above examples and includes any type of biomarker panel that can be used for the diagnosis of a disease as a combination of biomarkers.

[0018] In the present invention, the “diagnosis” includes determining the susceptibility of a subject to a specific disease or condition, determining whether a subject currently has a specific disease or condition, determining the prognosis of a subject suffering from a specific disease or condition (e.g., identifying a pre-metastatic or metastatic cancer state, determining the stage of cancer, or determining the responsiveness of cancer to treatment), or therametrics (e.g., monitoring the condition of a subject to provide information on the efficacy of treatment). For the purposes of the present invention, the diagnosis is to confirm whether or not the subject has developed the cancer or the possibility (risk) of developing the cancer.

[0019] The term “diagnostic composition” as used herein means an integrated mixture or device including a means for measuring the expression level of one or more genes selected from the group consisting of ANPEP, APOA4, APOC3, C9, CRP, HGFAC, IGFBP2, ITIH3, LRG1, ORM1, PFN1, PIGR, PON3, SERPINA3 and VWF or proteins encoded by them for determining whether pancreatic cancer has occurred in a subject or predicting the possibility of occurrence, and may also be expressed as a “diagnostic kit.”

[0020]

[0021] According to a specific embodiment of the present invention,

[0022] A fragment of the above ANPEP polypeptide has the amino acid sequence of SEQ ID NO: 1 (ALEQALEK);

[0023] A fragment of the above APOA4 polypeptide has the amino acid sequence of SEQ ID NO: 2 (LTPYADEFK);

[0024] A fragment of the above APOC3 polypeptide has the amino acid sequence of SEQ ID NO: 3 (GWVTDGFSSLK);

[0025] A fragment of the above C9 polypeptide has the amino acid sequence of SEQ ID NO: 4 (ALPTTYEK);

[0026] A fragment of the above CRP polypeptide has an amino acid sequence of SEQ ID NO: 5 (ESDTSYVSLK);

[0027] A fragment of the above HGFAC polypeptide has the amino acid sequence of SEQ ID NO: 6 (EALVPLVADHK);

[0028] A fragment of the above IGFBP2 polypeptide has the amino acid sequence of SEQ ID NO: 7 (LIQGAPTIR);

[0029] A fragment of the above ITIH3 polypeptide has the amino acid sequence of SEQ ID NO: 8 (ALDLSLK);

[0030] A fragment of the above LRG1 polypeptide has the amino acid sequence of SEQ ID NO: 9 (LHLEGNK);

[0031] A fragment of the above ORM1 polypeptide has the amino acid sequence of SEQ ID NO: 10 (SDVVYTDWK);

[0032] A fragment of the above PFN1 polypeptide has the amino acid sequence of SEQ ID NO: 11 (DSPSVWAAVPGK);

[0033] A fragment of the above PIGR polypeptide has the amino acid sequence of SEQ ID NO: 12 (VYTVDLGR);

[0034] A fragment of the above PON3 polypeptide has the amino acid sequence of SEQ ID NO: 13 (YVYVADVAAK);

[0035] A fragment of the above SERPINA3 polypeptide has the amino acid sequence of SEQ ID NO: 14 (EIGELYLPK);

[0036] A fragment of the above VWF polypeptide has the amino acid sequence of SEQ ID NO: 15 (ILAGPAGDSNVVK).

[0037]

[0038] According to a specific embodiment of the present invention, in an individual having pancreatic cancer, the expression level of one or more genes selected from the group consisting of ANPEP, APOA4, APOC3, C9, CRP, HGFAC, IGFBP2, ITIH3, LRG1, ORM1, PFN1, PIGR, PON3, SERPINA3 and VWF or the proteins encoded by them is increased.

[0039] The term “increased expression level” used when referring to the “composition for diagnosing pancreatic cancer” in the composition of the present invention means a case where the expression level of the corresponding gene or the protein encoded by the corresponding gene is significantly higher than that of the control group or normal group, and specifically means a case where the expression level is increased by about 10% or more, about 20% or more, about 30% or more, about 40% or more, about 50% or more, or about 60% or more compared to the control group or normal group, but does not exclude a range beyond this.

[0040] In the present invention, the term “reduced expression level” means a case where the expression level of a corresponding gene or a protein encoded by the corresponding gene is significantly lower than that of a control group or a normal group, and specifically means a case where the expression level is reduced by about 10% or more, about 20% or more, about 30% or more, about 40% or more, about 50% or more, or about 60% or more compared to the control group or normal group, but does not exclude a range exceeding this.

[0041]

[0042] According to a specific embodiment of the present invention, the pancreatic cancer is pancreatic ductal adenocarcinoma (PDAC).

[0043]

[0044] According to a specific embodiment of the present invention,

[0045] The preparation for measuring the expression level of the polypeptide includes at least one selected from the group consisting of an antibody, an antigen-binding fragment, a ligand, a peptide nucleic acid (PNA), and an aptamer that specifically binds to the polypeptide or a fragment thereof.

[0046] As used herein, the term "antibody" refers to a substance that specifically binds to an antigen and causes an antigen-antibody reaction. Specifically, for the purposes of the present invention, the antibody may be an antibody that specifically binds to the polypeptides mentioned herein.

[0047] The antibodies of the present invention include polyclonal antibodies, monoclonal antibodies, and recombinant antibodies. The antibodies can be readily produced using techniques well known in the art. For example, polyclonal antibodies can be produced by methods well known in the art, including a process of injecting an antigen of the protein into an animal and collecting blood from the animal to obtain serum containing the antibody. Such polyclonal antibodies can be produced from any animal, such as goats, rabbits, sheep, monkeys, horses, pigs, cows, and dogs. In addition, monoclonal antibodies can be produced using the hybridoma method (see Kohler and Milstein (1976) European Journal of Immunology 6:511-519) or the phage antibody library technique (see Clackson et al, Nature, 352:624-628, 1991; Marks et al, J. Mol. Biol., 222:58, 1-597, 1991), which are well known in the art. The antibody produced by the above method can be separated and purified using methods such as gel electrophoresis, dialysis, salt precipitation, ion exchange chromatography, and affinity chromatography. In addition, the antibody of the present invention includes not only a complete form having two full-length light chains and two full-length heavy chains, but also a functional fragment of the antibody molecule. A functional fragment of an antibody molecule means a fragment that possesses at least an antigen-binding function, and includes Fab, F(ab'), F(ab')2, and Fv.

[0048] As used herein, the term “antigen binding fragment” means a portion of a polypeptide in the overall structure of an immunoglobulin capable of binding an antigen, including, but not limited to, F(ab')2, Fab', Fab, Fv, and scFv.

[0049] The “PNA (Peptide Nucleic Acid)” in the present invention refers to an artificially synthesized polymer similar to DNA or RNA, and was first introduced in 1991 by Professors Nielsen, Egholm, Berg, and Buchardt of the University of Copenhagen, Denmark. While DNA has a phosphate-ribose sugar backbone, PNA has a repeated N-(2-aminoethyl)-glycine backbone linked by peptide bonds, which greatly increases its binding affinity and stability to DNA or RNA, and is thus used in molecular biology, diagnostic analysis, and antisense therapy. PNA is described in detail in the literature [Nielsen PE, Egholm M, Berg RH, Buchardt O (December 1991). “Sequence-selective recognition of DNA by strand displacement with a thymine-substituted polyamide”. Science 254(5037): 1497-1500].

[0050] In the present invention, the “aptamer” is an oligonucleotide or peptide molecule, and the general contents of the aptamer are disclosed in detail in the literature [Bock LC et al., Nature 355(6360):5646(1992); Hoppe-Seyler F, Butz K “Peptide aptamers: powerful new tools for molecular medicine”. J Mol Med. 78(8):42630(2000); Cohen BA, Colas P, Brent R. “An artificial cell-cycle inhibitor isolated from a combinatorial library”. Proc Natl Acad Sci USA. 95(24): 142727(1998)].

[0051]

[0052] According to a specific embodiment of the present invention,

[0053] A preparation for measuring the expression level of a gene encoding the above polypeptide or a fragment thereof comprises at least one selected from the group consisting of a primer, a probe and an antisense nucleotide that specifically bind to the above gene.

[0054] In the present invention, the "primer" refers to a fragment that recognizes a target gene sequence, and comprises a pair of forward and reverse primers, and is specifically a primer pair that provides analysis results with specificity and sensitivity. When the nucleic acid sequence of the primer is a sequence that does not match the non-target sequence present in the sample, and thus the primer only amplifies the target gene sequence containing the complementary primer binding site and does not cause non-specific amplification, high specificity can be imparted.

[0055] In the present invention, the “probe” refers to a substance that can specifically bind to a target substance to be detected in a sample, and refers to a substance that can specifically confirm the presence of the target substance in the sample through the binding. The type of probe is not limited to a substance commonly used in the art, and specifically may be PNA (peptide nucleic acid), LNA (locked nucleic acid), peptide, polypeptide, protein, RNA, or DNA, and most specifically PNA. More specifically, the probe includes a biomaterial derived from or similar to a living organism or manufactured in vitro, and may be, for example, an enzyme, a protein, an antibody, a microorganism, an animal or plant cell and organ, a nerve cell, DNA, and RNA. DNA includes cDNA, genomic DNA, and oligonucleotides, RNA includes genomic RNA, mRNA, and oligonucleotides, and examples of proteins may include antibodies, antigens, enzymes, peptides, etc.

[0056] In the present invention, the “LNA (Locked nucleic acids)” refers to a nucleic acid analog containing a 2’-O, 4’-C methylene bridge [J Weiler, J Hunziker and J Hall Gene Therapy (2006) 13, 496.502]. LNA nucleosides contain common nucleic acid bases of DNA and RNA and can form base pairs according to the Watson-Crick base pairing rule. However, due to the ‘locking’ of the molecule caused by the methylene bridge, LNA cannot form an ideal shape in Watson-Crick binding. When LNA is included in a DNA or RNA oligonucleotide, LNA can pair with a complementary nucleotide chain more quickly, thereby increasing the stability of the double helix. In the present invention, the term "antisense" refers to an oligomer having a sequence of nucleotide bases and a backbone between subunits that allows the antisense oligomer to hybridize with a target sequence within RNA by Watson-Crick base pairing, typically allowing the formation of an mRNA and RNA:oligomer heteroduplex within the target sequence. The oligomer may have exact sequence complementarity or approximate sequence complementarity to the target sequence.

[0057] Since the amino acid sequence information of the polypeptide according to the present invention and the nucleic acid sequence encoding the same are disclosed through various public data, those skilled in the art will be able to easily design primers, probes, or antisense nucleotides that specifically bind to the gene encoding the polypeptide based on this.

[0058] In this specification, the term “nucleic acid” may also be referred to as “nucleic acid molecule” and has a meaning that comprehensively includes DNA (gDNA and cDNA) and RNA molecules, and nucleotides, which are the basic structural units in nucleic acid molecules, include not only natural nucleotides but also analogues in which the sugar or base moiety is modified (Scheit, Nucleotide Analogs, John Wiley, New York (1980); Uhlman and Peyman, Chemical Reviews, 90:543-584 (1990)).

[0059]

[0060] According to one aspect of the present invention, the present invention provides a diagnostic kit comprising the diagnostic composition of the present invention.

[0061] In the present invention, the presence or absence of pancreatic cancer, the possibility of occurrence, treatment responsiveness, prognosis, stage, possibility of recurrence, etc. can be diagnosed using the above diagnostic kit.

[0062] Since the pancreatic cancer, which is the subject of the diagnosis in the present invention, has already been described above, its detailed description is omitted below to avoid excessive duplication.

[0063]

[0064] According to a specific embodiment of the present invention, the kit is an RT-PCR kit, a DNA chip kit, an ELISA kit, a protein chip kit, a rapid kit, or an MRM (Multiple reaction monitoring) kit.

[0065] In the present invention, the kit may be, but is not limited to, an RT-PCR kit, a DNA chip kit, an ELISA kit, a protein chip kit, a rapid kit, or an MRM (Multiple reaction monitoring) kit.

[0066] The diagnostic kit for pancreatic cancer of the present invention may further include one or more other component compositions, solutions or devices suitable for the analysis method.

[0067] For example, the diagnostic kit for pancreatic cancer in the present invention may further include essential components necessary for performing a reverse transcription polymerase reaction. The reverse transcription polymerase reaction kit includes a pair of primers specific for a gene encoding a marker protein. The primers are nucleotides having a sequence specific to the nucleic acid sequence of the gene, and may have a length of about 7 to 50 bp, more specifically, about 10 to 30 bp. It may also include a primer specific to the nucleic acid sequence of a control gene. In addition, the reverse transcription polymerase reaction kit may include a test tube or other appropriate container, a reaction buffer (with various pH and magnesium concentrations), deoxynucleotides (dNTPs), an enzyme such as Taq polymerase and reverse transcriptase, DNase, RNase inhibitor DEPC-water, sterile water, etc.

[0068] Additionally, the diagnostic kit of the present invention may include essential elements necessary for performing a DNA chip. The DNA chip kit may include a substrate to which cDNA or oligonucleotides corresponding to a gene or fragment thereof are attached, and reagents, preparations, enzymes, etc. for producing a fluorescently labeled probe. The substrate may also include cDNA or oligonucleotides corresponding to a control gene or fragment thereof.

[0069] Additionally, the diagnostic kit of the present invention may include essential components necessary for performing an ELISA. The ELISA kit includes an antibody specific for the protein. The antibody is an antibody with high specificity and affinity for the marker protein and little cross-reactivity with other proteins, and may be a monoclonal antibody, polyclonal antibody, or recombinant antibody. The ELISA kit may also include an antibody specific for a control protein. In addition, the ELISA kit may include reagents capable of detecting bound antibodies, such as labeled secondary antibodies, chromophores, enzymes (e.g., conjugated to antibodies), and their substrates or other substances capable of binding to antibodies.

[0070] In the diagnostic kit of the present invention, a fixative for the antigen-antibody binding reaction may be a nitrocellulose membrane, a PVDF membrane, a well plate synthesized from polyvinyl resin or polystyrene resin, a glass slide glass, etc., but is not limited thereto.

[0071] In addition, in the diagnostic kit of the present invention, the label of the secondary antibody is preferably a conventional chromogen that undergoes a color development reaction, and labels such as fluorescein and dyes such as HRP (horseradish peroxidase), alkaline phosphatase, colloid gold, FITC (poly L-lysine-fluorescein isothiocyanate), and RITC (rhodamine-B-isothiocyanate) can be used, but are not limited thereto.

[0072] In addition, in the diagnostic kit of the present invention, it is preferable to use a chromogenic substrate for inducing color development according to a marker that undergoes a color development reaction, and TMB (3,3',5,5'-tetramethyl bezidine), ABTS [2,2'-azino-bis(3-ethylbenzothiazoline-6-sulfonic acid)], OPD (o-phenylenediamine), etc. can be used. At this time, it is more preferable that the chromogenic substrate is provided in a state dissolved in a buffer solution (0.1 M NaAc, pH 5.5). A chromogenic substrate such as TMB is decomposed by HRP used as a marker of a secondary antibody conjugate to generate a chromogenic precipitate, and the presence or absence of the marker proteins is detected by visually confirming the degree of deposition of this chromogenic precipitate.

[0073] In the diagnostic kit of the present invention, the washing solution preferably contains phosphate buffer, NaCl, and Tween 20, and a buffer solution (PBST) composed of 0.02 M phosphate buffer, 0.13 M NaCl, and 0.05% Tween 20 is more preferred. After the antigen-antibody binding reaction, the washing solution reacts the antigen-antibody complex with a secondary antibody, and then adds an appropriate amount to the fixative and washes 3 to 6 times. The reaction stopping solution can preferably be a sulfuric acid solution (H2SO4).

[0074]

[0075] According to one aspect of the present invention, the present invention provides a biological sample isolated from a target object,

[0076] One or more polypeptides selected from the group consisting of ANPEP (Aminopeptidase N), APOA4 (Apolipoprotein A-IV), APOC3 (Apolipoprotein C-III), C9 (Complement component C9), CRP (C-reactive protein), HGFAC (Hepatocyte growth factor activator), IGFBP2 (Insulin-like growth factor-binding protein 2), ITIH3 (Inter-alpha-trypsin inhibitor heavy chain H3), LRG1 (Leucine-rich alpha-2-glycoprotein), ORM1 (Alpha-1-acid glycoprotein 1), PFN1 (Profilin-1), PIGR (Polymeric immunoglobulin receptor), PON3 (Serum paraoxonase / lactonase 3), SERPINA3 (Alpha-1-antichymotrypsin), and VWF (von Willebrand factor), or fragments thereof; Or, a method for providing information necessary for diagnosing pancreatic cancer is provided, which comprises a step of measuring the expression level of a gene encoding the same.

[0077] Since the polypeptides to be measured in the present invention have already been described above, their detailed description is omitted below to avoid excessive duplication.

[0078] The present inventors have discovered for the first time that the expression levels of ANPEP, C9, CRP, IGFBP2, ITIH3, LRG1, ORM1, PIGR, SERPINA3, and VWF proteins or the genes encoding them are positively correlated with the likelihood of developing pancreatic cancer in patients with pancreatic cancer, particularly pancreatic ductal adenocarcinoma, and that the expression levels of APOA4, APOC3, HGFAC, PFN1, and PON3 proteins or the genes encoding them are negatively correlated with the likelihood of developing pancreatic cancer. Accordingly, when a protein having the aforementioned positive correlation or a gene encoding it is highly expressed in a biological sample isolated from the subject or a biological sample isolated from the subject, or a protein having the aforementioned negative correlation or a gene encoding it is low-expressed, the subject is determined to have developed pancreatic cancer or to have a high likelihood of developing pancreatic cancer in the future.

[0079] In this specification, the term “high expression” means a case where the expression level of a gene or a protein encoded by the gene is significantly higher than that of a control group or a normal group, and specifically means a case where the expression level increases by about 10% or more, about 20% or more, about 30% or more, about 40% or more, about 50% or more, or about 60% or more compared to the control group or normal group, but does not exclude a range beyond this.

[0080] In this specification, the term “low expression” means a case where the expression level of a gene or a protein encoded by the gene is significantly lower than that of a control group or a normal group, and specifically means a case where the expression level is reduced by about 10% or more, about 20% or more, about 30% or more, about 40% or more, about 50% or more, or about 60% or more compared to the control group or normal group, but does not exclude a range beyond this.

[0081] As used herein, the term “subject” refers to a subject that provides a sample for measuring the expression level of the genes of the present invention or the proteins they encode, and ultimately becomes the subject of analysis for the development of pancreatic cancer. The subject includes, without limitation, a human, a mouse, a rat, a guinea pig, a dog, a cat, a horse, a cow, a pig, a monkey, a chimpanzee, a baboon, or a rhesus macaque, and is specifically a human. Since the composition of the present invention provides information for predicting not only the current development of pancreatic cancer but also the genetic risk of developing pancreatic cancer in the future, the subject of the present invention may be a pancreatic cancer patient or a healthy subject that has not yet developed pancreatic cancer.

[0082] In the present invention, the “biological sample” means any material, biological fluid, tissue or cell obtained from or derived from an individual, for example, whole blood, leukocytes, peripheral blood mononuclear cells, buffy coat, plasma, serum, sputum, tears, mucus, nasal washes, nasal aspirate, breath, urine, semen, saliva, peritoneal washings, ascites, cystic fluid, meningeal fluid, amniotic fluid, glandular fluid, pancreatic fluid, lymph fluid, pleural fluid, nipple aspirate, bronchial It includes bronchial aspirate, synovial fluid, joint aspirate, organ secretions, cells, cell extracts or cerebrospinal fluid, and more specifically, it may be a liquid biopsy (e.g., patient's tissue, cells, blood, serum, plasma, saliva, sputum or ascites, etc.) collected for histopathological examination by inserting a hollow needle or the like into an organ in a living body without incising the skin of a patient with a high risk of developing the disease.

[0083] The present invention may include a step of measuring the expression level of a polypeptide represented by SEQ ID NO: 1 to 15 or a gene encoding the same in a biological sample separated as described above.

[0084] In the present invention, the step of measuring the expression level may be a step of measuring the expression level of one or more proteins (polypeptides) selected from the group consisting of ANPEP (Aminopeptidase N), APOA4 (Apolipoprotein A-IV), APOC3 (Apolipoprotein C-III), C9 (Complement component C9), CRP (C-reactive protein), HGFAC (Hepatocyte growth factor activator), IGFBP2 (Insulin-like growth factor-binding protein 2), ITIH3 (Inter-alpha-trypsin inhibitor heavy chain H3), LRG1 (Leucine-rich alpha-2-glycoprotein), ORM1 (Alpha-1-acid glycoprotein 1), PFN1 (Profilin-1), PIGR (Polymeric immunoglobulin receptor), PON3 (Serum paraoxonase / lactonase 3), SERPINA3 (Alpha-1-antichymotrypsin), and VWF (von Willebrand factor) or a gene encoding the same.

[0085]

[0086] According to a specific embodiment of the present invention,

[0087] A fragment of the above ANPEP polypeptide has the amino acid sequence of SEQ ID NO: 1 (ALEQALEK);

[0088] A fragment of the above APOA4 polypeptide has the amino acid sequence of SEQ ID NO: 2 (LTPYADEFK);

[0089] A fragment of the above APOC3 polypeptide has the amino acid sequence of SEQ ID NO: 3 (GWVTDGFSSLK);

[0090] A fragment of the above C9 polypeptide has the amino acid sequence of SEQ ID NO: 4 (ALPTTYEK);

[0091] A fragment of the above CRP polypeptide has an amino acid sequence of SEQ ID NO: 5 (ESDTSYVSLK);

[0092] A fragment of the above HGFAC polypeptide has the amino acid sequence of SEQ ID NO: 6 (EALVPLVADHK);

[0093] A fragment of the above IGFBP2 polypeptide has the amino acid sequence of SEQ ID NO: 7 (LIQGAPTIR);

[0094] A fragment of the above ITIH3 polypeptide has the amino acid sequence of SEQ ID NO: 8 (ALDLSLK);

[0095] A fragment of the above LRG1 polypeptide has the amino acid sequence of SEQ ID NO: 9 (LHLEGNK);

[0096] A fragment of the above ORM1 polypeptide has the amino acid sequence of SEQ ID NO: 10 (SDVVYTDWK);

[0097] A fragment of the above PFN1 polypeptide has the amino acid sequence of SEQ ID NO: 11 (DSPSVWAAVPGK);

[0098] A fragment of the above PIGR polypeptide has the amino acid sequence of SEQ ID NO: 12 (VYTVDLGR);

[0099] A fragment of the above PON3 polypeptide has the amino acid sequence of SEQ ID NO: 13 (YVYVADVAAK);

[0100] A fragment of the above SERPINA3 polypeptide has the amino acid sequence of SEQ ID NO: 14 (EIGELYLPK);

[0101] A fragment of the above VWF polypeptide has the amino acid sequence of SEQ ID NO: 15 (ILAGPAGDSNVVK).

[0102]

[0103] According to a specific embodiment of the present invention, the biological sample is whole blood, leukocytes, peripheral blood mononuclear cells, buffy coat, plasma, serum, sputum, tears, mucus, nasal washes, nasal aspirate, breath, urine, semen, saliva, peritoneal washings, ascites, cystic fluid, meningeal fluid, amniotic fluid, glandular fluid, pancreatic fluid, lymph fluid, pleural fluid, nipple aspirate, bronchial aspirate, synovial fluid, joint Aspirates include joint aspirate, organ secretions, cells, cell extracts, or cerebrospinal fluid.

[0104] The term “blood” in this specification means “whole blood,” which includes plasma and serum.

[0105] The term "whole blood" as used herein generally refers to blood composed of uncoagulated plasma and cellular components. Plasma comprises approximately 50 to 60% of the volume of whole blood, and cellular components (e.g., red blood cells, white blood cells, or platelets) may comprise approximately 40 to 50%.

[0106] The term “plasma” as used herein refers to the liquid component of blood, which functions as a transport medium in supplying nutrients to the cells and organs of the body.

[0107] The term “serum” used in this specification refers to a pale yellow liquid collected from blood, and more specifically, it refers to a pale yellow body fluid component that remains when a red clot is removed after blood is collected and left to stand, as the fluidity of the blood decreases.

[0108] According to the present invention, pancreatic cancer can be diagnosed using the biomarkers of the present invention, and the diagnosis can be made through a liquid biopsy using a patient-derived body fluid such as blood as a biological sample. This liquid biopsy method has the advantage of being non-invasive compared to conventional tissue biopsy procedures, thereby minimizing patient pain, and also has the advantage of obtaining information about cancer more quickly.

[0109]

[0110] According to a specific embodiment of the present invention, the agent for measuring the expression level of the polypeptide comprises at least one selected from the group consisting of an antibody, an oligopeptide, a ligand, a peptide nucleic acid (PNA), and an aptamer that specifically binds to the polypeptide.

[0111]

[0112] According to a specific embodiment of the present invention, the measurement of the expression level of the polypeptide is performed by protein chip analysis, immunoassay, ligand binding assay, MALDI-TOF (Matrix Assisted Laser Desorption / Ionization Time of Flight Mass Spectrometry) analysis, SELDI-TOF (Sulface Enhanced Laser Desorption / Ionization Time of Flight Mass Spectrometry) analysis, radioimmunoassay, radioimmunodiffusion, aukteroni immunodiffusion, rocket immunoelectrophoresis, tissue immunostaining, complement fixation assay, two-dimensional electrophoresis analysis, liquid chromatography-mass spectrometry (LC-MS), liquid chromatography-mass spectrometry / mass spectrometry (LC-MS), liquid chromatography-mass spectrometry / mass spectrometry (LC-MS), western blotting or enzyme linked immunosorbent assay (ELISA).

[0113]

[0114] According to a specific embodiment of the present invention, the measurement of the expression level of the polypeptide is performed by a multiple reaction monitoring (MRM) method.

[0115] In the present invention, the multiple reaction monitoring method can be performed using a mass spectrometer (mass-spectrometry), specifically, a triple quadrupole mass spectrometer (triple quadrupole mass-spectrometry).

[0116] In the present invention, the multiple reaction monitoring (MRM) method using mass spectrometry is an analytical technique that can selectively isolate, detect, and quantify a specific analyte and monitor its concentration changes. MRM is a method that can quantitatively and accurately measure multiple substances, such as trace biomarkers, present in biological samples. It selectively transmits mother ions among the ion fragments generated from the ion source to the collision tube using the first mass filter (Q1). The mother ions that reach the collision tube then collide with the internal collision gas, splitting them into daughter ions and sending them to the second mass filter (Q2), from which only characteristic ions are transmitted to the detection unit. This method is a highly selective and sensitive analytical method that can detect only the information of the target component. MRM is utilized for the quantitative analysis of small molecules and is used to diagnose specific genetic diseases. The MRM method facilitates the simultaneous measurement of multiple peptides and has the advantage of identifying relative concentration differences in protein diagnostic marker candidates between normal individuals and cancer patients without the need for antibodies. Furthermore, due to its excellent sensitivity and selectivity, MRM analysis is being introduced for the analysis of complex proteins and peptides in blood, particularly in proteomic analyses using mass spectrometry (Anderson L. et al., Mol CellProteomics, 5: 375-88, 2006; DeSouza, LV et al., Anal. Chem., 81: 3462-70, 2009).

[0117] The term “mass spectrometry” as used herein may also be referred to as “mass spectrometry” and refers to a method of analyzing unknown compounds by their masses. Mass spectrometry is performed by filtering, detecting, and measuring ions using their mass-to-charge ratio (m / z) values. Generally, mass spectrometry includes the steps of (1) ionizing and charging a compound and (2) measuring the molecular weight of the charged compound and calculating the m / z value. The calculated m / z value is used as a reference to identify and quantify a target compound in a complex mixture. The mass spectrometry of this specification can be performed through any type of mass spectrometry that utilizes the above principles.

[0118] In the present invention, the expression level of the polypeptide mentioned in the present invention can be measured by the above multiple reaction monitoring method.

[0119] In order to analyze the expression level of the polypeptide mentioned in the present invention by the multiple reaction monitoring method in the present invention, the mass / charge value (m / z value) of the target peptide may be used, but is not limited thereto.

[0120] In this specification, the term "multiple reaction monitoring" refers to an analytical technology that can selectively isolate, detect, and quantify a specific analyte and monitor its concentration changes. MRM is a method that can quantitatively and accurately measure multiple substances such as trace biomarkers present in biological samples. It uses a first mass filter (Q1) to selectively transmit mother ions among the ion fragments generated from the ion source to the collision tube. The mother ions that reach the collision tube collide with the internal collision gas, splitting them to produce daughter ions that are then transmitted to the second mass filter (Q2), from which only characteristic ions are transmitted to the detection unit. This method is a highly selective and sensitive analytical method that can detect only the information of the target component. MRM is utilized for the quantitative analysis of small molecules and is used to diagnose specific genetic diseases. The MRM method is easy to measure multiple peptides simultaneously and has the advantage of being able to identify the relative concentration differences of protein diagnostic marker candidates between normal individuals and cancer patients without antibodies. In addition, because of its excellent sensitivity and selectivity, MRM analysis method is being introduced for the analysis of complex proteins and peptides in blood, especially in proteome analysis using mass spectrometry (Anderson L. et al., Mol CellProteomics, 5: 375-88, 2006; DeSouza, LV et al., Anal. Chem., 81: 3462-70, 2009).

[0121] The term “parallel reaction monitoring” in this specification refers to a method of applying MRM in parallel, which simultaneously analyzes all daughter ions generated from a selected parent ion, unlike MRM which analyzes one pair of parent ions / daughter ions at a time.

[0122] Mass spectrometry analysis of proteins used as biomarkers for pancreatic cancer in the present invention can be performed specifically through data-independent acquisition (DIA) analysis. As used herein, the term "data-independent acquisition" refers to a method of analyzing all ions within a selected m / z range without selecting a specific parent ion.

[0123]

[0124] According to a specific embodiment of the present invention,

[0125] The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 1 is 901.506 for a light peptide and 909.52 for a heavy peptide when the z value is 1; or a value within ±1 range of each value;

[0126] The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 2 is 1083.536 for a light peptide and 1091.550 for a heavy peptide when the z value is 1; or a value within ±1 range of each value;

[0127] The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 3 is 1196.595 for a light peptide and 1204.609 for a heavy peptide when the z value is 1; or a value within a range of ±1 of each value;

[0128] The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 4 is 922.496 for a light peptide and 930.508 for a heavy peptide when the z value is 1; or a value within a range of ±1 of each value;

[0129] The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 5 is 1128.542 for a light peptide and 1136.556 for a heavy peptide when the z value is 1; or a value within ±1 range of each value;

[0130] The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 6 is 1191.673 for a light peptide and 1209.708 for a heavy peptide when the z value is 1; or a value within a range of ±1 of each value;

[0131] The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 7 is 968.596 for a light peptide and 978.604 for a heavy peptide when the z value is 1; or a value within ±1 range of each value;

[0132] The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 8 is 759.468 for a light peptide and 767.482 for a heavy peptide when the z value is 1; or a value within a range of ±1 of each value;

[0133] The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 9 is 810.454 for a light peptide and 818.468 for a heavy peptide when the z value is 1; or a value within ±1 range of each value;

[0134] The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 10 is 1112.526 for a light peptide and 1120.540 for a heavy peptide when the z value is 1; or a value within a range of ±1 of each value;

[0135] The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 11 is 1213.621 for a light peptide and 1221.635 for a heavy peptide when the z value is 1; or a value within ±1 range of each value;

[0136] The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 12 is 922.506 for a light peptide and 932.508 for a heavy peptide when the z value is 1; or a value within ±1 range of each value;

[0137] The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 13 is 1098.59 for a light peptide and 1110.61 for a heavy peptide when the z value is 1; or a value within a range of ±1 of each value;

[0138] The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 14 is 1061.588 for a light peptide and 1069.602 for a heavy peptide when the z value is 1; or a value within a range of ±1 of each value;

[0139] The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 15 is 1240.696 for a light peptide and 1248.71 for a heavy peptide when the z value is 1; or a value within ±1 range of each value.

[0140] In the present invention, the term 'heavy peptide' means a synthetic peptide labeled with isotopes, etc., and may also be referred to as a SIL (synthetic isotopically labeled) peptide. A heavy peptide is bound to a non-radioactive stable isotope, and the stable isotope is generally 13 C (carbon-13), 15 N(nitrogen-15) or 2H (deuterium) is used, but is not limited thereto. In general, heavy peptides have the same physiochemical characteristics and chemical reactivity as the 'light peptide' that is the target of analysis, but behave differently from the 'light peptide' due to the mass difference by isotope. Such heavy peptides are used to quantify the absolute amount of the 'light peptide' that is the target of analysis. The term 'light peptide' as used herein means a peptide that is not labeled with an isotope, as opposed to a 'heavy peptide', and generally means the target peptide that is the target of quantitative analysis.

[0141]

[0142] According to a specific embodiment of the present invention,

[0143] When performing the above multiple reaction monitoring, a synthetic peptide in which a specific element of a specific amino acid constituting each polypeptide is substituted with an isotope is used as an internal standard material; or E. coli beta-galactosidase is used.

[0144] In the present invention, any internal standard material generally used in the above multiple reaction monitoring analysis may be used as the internal standard material, but for example, E. coli beta galactosidase may be used.

[0145] In addition, in the present invention, when a specific peptide in which some amino acids of the target peptide are substituted with stable isotopes is synthesized as an internal standard substance to measure the absolute amount of the polypeptide in blood, the isotope-substituted amino acid may be lysine or arginine, but is not limited thereto. Here, the synthesized peptide may be an isolated peptide with a purity of 95% or higher.

[0146] In the present invention, the internal standard material is 2 H, 3 H, 11 C, 13 C, 14 C, 13N, 15 N, 15 O, 17 O and 18 It may include one or more radioactive isotopes selected from the group consisting of O, but is not limited to the aforementioned radioactive isotopes, and may include any type of isotope that can be used as a comparison group in measuring the absolute amount of the polypeptide.

[0147]

[0148] According to a specific embodiment of the present invention, the synthetic peptide has a sequence identical to the sequence represented by SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 and includes a stable isotope.

[0149]

[0150] According to a specific embodiment of the present invention, the stable isotope is a stable isotope of at least one element selected from the group consisting of carbon and nitrogen.

[0151]

[0152] According to a specific embodiment of the present invention,

[0153] The expression level of the gene encoding the above polypeptide is measured by reverse transcription polymerase chain reaction (RT-PCR), competitive reverse transcription polymerase reaction (Competitive RT-PCR), real-time reverse transcription polymerase reaction (Real-time RT-PCR), RNase protection assay (RPA), Northern blotting, or DNA chip.

[0154] In the present invention, the presence or absence of a gene encoding the polypeptide and the level of expression thereof are confirmed, and the analytical methods for measuring the amount of mRNA include, but are not limited to, reverse transcription polymerase chain reaction (RT-PCR), competitive reverse transcription polymerase reaction (Competitive RT-PCR), real-time reverse transcription polymerase reaction (Real-time RT-PCR), RNase protection assay (RPA), Northern blotting, DNA chips, etc.

[0155] In the present invention, the expression level of the polypeptide or the gene encoding it measured for a biological sample of the target object can be measured to predict the treatment responsiveness to pancreatic cancer.

[0156] In the present invention, the prognosis of a subject, preferably a prognosis after surgical operation, can be predicted by measuring the expression level of one or more of the polypeptides represented by SEQ ID NOS: 1 to 15 or the genes encoding them, measured for a biological sample of the subject of interest. Here, the subject of interest may be a subject that has developed pancreatic cancer and has undergone surgical resection.

[0157] In the present invention, the stage of pancreatic cancer in an individual can be predicted by measuring the expression level of the polypeptide or the gene encoding it in a biological sample of the individual.

[0158] In the present invention, the above “stage” refers to the extent to which cancer cells have spread and the stage of cancer progression, and the international classification according to the progression of cancer generally follows the TNM staging classification. Here, ‘T (Tumor Size)’ is a classification according to the size of the primary tumor, ‘N (Lymph Node)’ is a classification according to the degree of lymph node metastasis, and ‘M (Metastasis)’ corresponds to a classification according to whether or not there is metastasis to other organs.

[0159] In the present invention, the possibility of recurrence of pancreatic cancer can be predicted by measuring the expression level of the polypeptide or the gene encoding it in a biological sample of the subject of interest.

[0160]

[0161] According to a specific embodiment of the present invention, the expression level of the ANPEP (Aminopeptidase N), C9 (Complement component C9), CRP (C-reactive protein), IGFBP2 (Insulin-like growth factor-binding protein 2), ITIH3 (Inter-alpha-trypsin inhibitor heavy chain H3), LRG1 (Leucine-rich alpha-2-glycoprotein), ORM1 (Alpha-1-acid glycoprotein 1), PIGR (Polymeric immunoglobulin receptor), SERPINA3 (Alpha-1-antichymotrypsin) or VWF (von Willebrand factor) polypeptide or a gene encoding the same measured for a biological sample of the target individual is increased compared to a normal control group;

[0162] If the expression level of APOA4 (Apolipoprotein A-IV), APOC3 (Apolipoprotein C-III), HGFAC (Hepatocyte growth factor activator), PFN1 (Profilin-1), or PON3 (Serum paraoxonase / lactonase 3) polypeptides or the genes encoding them is decreased compared to the normal control group,

[0163] It is predicted that the risk of developing the above pancreatic cancer is high.

[0164]

[0165] According to a specific embodiment of the present invention, the pancreatic cancer is pancreatic ductal adenocarcinoma (PDAC).

[0166]

[0167] According to another aspect of the present invention, the present invention provides a method for screening a composition for preventing or treating pancreatic cancer, comprising the following steps:

[0168] (a) a step of contacting a test substance with a biological sample containing ANPEP, APOA4, APOC3, C9, CRP, HGFAC, IGFBP2, ITIH3, LRG1, ORM1, PFN1, PIGR, PON3, SERPINA3 and VWF genes or proteins encoded by them or cells expressing them;

[0169] (b) a step of measuring the expression level of the protein or the gene in the biological sample;

[0170] The activity or expression level of the ANPEP, C9, CRP, IGFBP2, ITIH3, LRG1, ORM1, PIGR, SERPINA3 and VWF genes or proteins in the biological sample is reduced;

[0171] If the activity or expression level of APOA4, APOC3, HGFAC, PFN1, or PON3 genes or proteins increases, the composition is determined to be for the prevention or treatment of pancreatic cancer.

[0172] Since the sample in the present invention has already been described above, its detailed description is omitted to avoid excessive duplication.

[0173] The term "biological sample" in the present invention refers to any sample obtained from a mammal, including a human, that contains cells expressing the aforementioned gene or proteins produced by the expression of the aforementioned gene, including, but not limited to, tissues, organs, cells, or cell cultures. More specifically, the biological sample may be cancer tissue, cancer cells, a culture thereof, or blood.

[0174] The term "test substance" used when referring to the screening method of the present invention refers to an unknown substance used in screening to examine whether the substance added to a sample containing cells expressing the genes of the present invention affects the activity or expression level of these genes. The test substance includes, but is not limited to, compounds, nucleotides, peptides, and natural extracts. The step of measuring the expression level or activity of the gene in a biological sample treated with the test substance can be performed using various expression level and activity measurement methods known in the art.

[0175]

[0176] According to a specific embodiment of the present invention, the pancreatic cancer is pancreatic ductal adenocarcinoma (PDAC).

[0177] Since the meaning of the term “pancreatic ductal adenocarcinoma (PDAC)” in this specification has already been described above, its description is omitted to avoid excessive duplication.

[0178]

[0179] According to one aspect of the present invention, the present invention provides a diagnosis system for pancreatic cancer, comprising: an input unit for receiving an input value; a reading unit including a machine learning model pre-trained to read whether pancreatic cancer has occurred; and an output unit for outputting whether pancreatic cancer has occurred. The system provides that the above input value is a measurement value of the expression level of one or more polypeptides selected from the group consisting of SEQ ID NO: 1 (ALEQALEK), SEQ ID NO: 2 (LTPYADEFK), SEQ ID NO: 3 (GWVTDGFSSLK), SEQ ID NO: 4 (ALPTTYEK), SEQ ID NO: 5 (ESDTSYVSLK), SEQ ID NO: 6 (EALVPLVADHK), SEQ ID NO: 7 (LIQGAPTIR), SEQ ID NO: 8 (ALDLSLK), SEQ ID NO: 9 (LHLEGNK), SEQ ID NO: 10 (SDVVYTDWK), SEQ ID NO: 11 (DSPSVWAAVPGK), SEQ ID NO: 12 (VYTVDLGR), SEQ ID NO: 13 (YVYVADVAAK), SEQ ID NO: 14 (EIGELYLPK) and SEQ ID NO: 15 (ILAGPAGDSNVVK) in a biological sample.

[0180] The term "machine learning" in the present invention refers to algorithms and statistical models that a computer system uses to perform tasks without explicit external instructions, relying on patterns and inference. Specifically, the machine learning model in the present invention may use deep learning, logistic regression, support vector machine (SVM), random forest, and gradient boosting algorithm (GBM). More specifically, deep learning may be used. Most specifically, a stacking ensemble model may be used by combining a total of seven machine learning algorithms to maximize optimal performance and model robustness. This is a binary classification model that classifies pancreatic cancer, and the ensemble composition algorithm includes tree-based models such as Extra-trees (ET), Light-GBM (LGBM), random-forest (RF), gradient-boosting (GB), XGBoost (XGB), and Ada-boost (Ada) algorithms, and a neural-network-based multi-layer perceptron (MLP). The hyperparameters of each algorithm use default values, and the MLP model can be composed of one hidden layer consisting of a total of 16 neurons. The final prediction result can be evaluated by synthesizing the prediction results of each model with a logistic regression algorithm. However, without being limited to the above, any type of machine learning model that can diagnose breast cancer using data measuring the expression level of the biomarker of the present invention can be used.

[0181]

[0182] According to a specific embodiment of the present invention, the biological sample is blood.

[0183] In the present invention, the input value of the machine learning model may be a measurement value of the expression level of a polypeptide measured in “blood”, which is a biological sample, but is not limited thereto, and means any material, biological fluid, tissue or cell obtained from or derived from an individual, for example, whole blood, leukocytes, peripheral blood mononuclear cells, buffy coat, plasma, serum, sputum, tears, mucus, nasal washes, nasal aspirate, breath, urine, semen, saliva, peritoneal washings, ascites, cystic fluid, meningeal fluid, amniotic fluid, glandular fluid, pancreatic fluid, lymph. It may include lymph fluid, pleural fluid, nipple aspirate, bronchial aspirate, synovial fluid, joint aspirate, organ secretions, cells, cell extracts or cerebrospinal fluid, but preferably, it may be a liquid biopsy (e.g., tissue, cells, blood, serum, plasma, saliva, sputum or ascites of a patient) collected for histopathological examination by inserting a hollow needle or the like into an organ in a living body without incising the skin of a patient with a high risk of developing the disease.

[0184]

[0185] According to a specific embodiment of the present invention, the measurement value of the expression level of the polypeptide is a quantitative value according to mass spectrometry.

[0186] According to a specific embodiment of the present invention, the mass spectrometry is liquid chromatography-tandem mass spectrometry (LC-MS / MS).

[0187] According to a specific embodiment of the present invention, the pancreatic cancer is pancreatic ductal adenocarcinoma (PDAC).

[0188] Since the mass spectrometry for measuring the expression level of the polypeptide in the present invention and pancreatic ductal adenocarcinoma have already been described above, detailed description thereof will be omitted below to avoid excessive duplication.

[0189]

[0190] The features and advantages of the present invention are summarized as follows:

[0191] (a) The present invention provides a method capable of predicting the occurrence of pancreatic cancer with high accuracy by measuring the expression level of a gene or protein involved in the occurrence of pancreatic cancer.

[0192] (b) The present invention provides a multifaceted, comprehensive, and novel treatment strategy for the early diagnosis of pancreatic cancer, including a method for predicting the onset of pancreatic cancer at an early stage with high reliability by discovering an effective biomarker for pancreatic cancer, particularly pancreatic ductal adenocarcinoma, thereby ultimately being useful in improving the survival rate from pancreatic cancer.

[0193]

[0194] Figure 1 is a drawing showing chromatograms for 84 peptides (5 uL of 10 pg / uL in 0.1 % formic acid, respectively).

[0195] Figure 2 is a graph showing the results of analyzing the highest sensitivity distribution and analytical stability derived in the process of optimizing the column oven temperature.

[0196] Figure 3 is a graph showing the results of analyzing the highest sensitivity distribution and analysis stability by flow rate conditions.

[0197] Figure 4 is a drawing showing chromatograms according to the existing concentration gradient (top) and the concentration gradient after change according to optimization (bottom).

[0198] Figure 5 is a graph showing the distribution of maximum sensitivity by GS1 condition.

[0199] Figure 6 is a graph showing the distribution of maximum sensitivity by GS2 condition.

[0200] Figure 7 is a graph showing the distribution of maximum sensitivity according to TEM (Ion Source Temperature) conditions.

[0201] Figure 8 is a graph showing the distribution of maximum sensitivity by IS condition.

[0202] Figure 9 is a graph showing the distribution of maximum sensitivity by CUR condition.

[0203] Figure 10 is a graph showing the distribution of maximum sensitivity by CUR condition.

[0204] Figure 11 is a diagram illustrating an example of MRM and MS2 analysis for a light peptide standard.

[0205] Figure 12 is a diagram illustrating an example of optimization of conditions for a target peptide (CXP).

[0206] Figure 13 is a diagram illustrating an example of a standard spiking test for confirming the detection of an endogenous peptide.

[0207] Figure 14 is a diagram showing whether each candidate marker is detected according to the protein concentration in blood.

[0208] Figure 15 is a diagram showing an example of the results of a heavy interference test for isotope determination.

[0209] Figure 16 is a diagram showing an example of the results of a heavy interference test for isotope determination.

[0210] Figure 17 is a diagram illustrating the structure of a prediction model (Stacking ensemble).

[0211] Figure 18 is a diagram showing the performance results (ROC curve) of the prediction model (stacking ensemble).

[0212]

[0213]

[0214]

[0215] Hereinafter, the present invention will be described in more detail through examples. These examples are intended solely to illustrate the present invention more specifically, and it will be apparent to those skilled in the art that the scope of the present invention is not limited by these examples, in accordance with the gist of the present invention.

[0216]

[0217] Example

[0218] Experimental method

[0219] 1. Marker excavation

[0220] Sample collection method

[0221] For the present invention, a total of 80 serum samples collected from normal individuals and pancreatic cancer patients stored in the Seoul National University Hospital Biobank from 2015 to 2021 were studied. The normal samples used in the present invention were 40 cases collected from individuals who were confirmed to not have pancreatic cancer through a baseline examination at the time of blood collection and had no history of diagnosis of other cancers, including pancreatic cancer, within the past 10 years from the time of blood collection. The pancreatic cancer samples were samples from 40 patients with pathologically confirmed pancreatic cancer (Pancreatic Ductal Adenocarcinoma type) collected before surgery. The age range of the collected samples was 35 to 81 years, with adults aged 20 years or older being the majority. The distribution of the pancreatic cancer samples by stage was confirmed to be AJCC stage 1 in 9 cases, stage 2 in 6 cases, stage 3 in 15 cases, and stage 4 in 10 cases. Specimens were primarily collected from the study “Establishment of a Bio Repository for Liver, Bile Tract, Pancreas, and Tumor Research (Seoul National University Hospital IRB No. 0901-010-267)” and other specimens were collected and provided by the Seoul National University Hospital Human Resources Bank.

[0222]

[0223] Profiling analysis

[0224] - Blood sample preprocessing

[0225] Serum samples without removing high-concentration proteins (albumin, transferrin, immunoglobulins, etc.) were used. 5 μl of serum was added to a urea (8 M final) and dithiothreitol (19 mM final) reagent buffer and incubated at 37°C for 1 hour and 30 minutes. After the reaction, the solution was cooled to room temperature and incubated with iodoacetamide (IAA, 26 mM final) in the dark at room temperature for 30 minutes. The solution was then diluted to a urea concentration of less than 1 M for trypsin digestion, but with ammonium bicarbonate (100 mM final) buffer. After adding 5 μg of trypsin (~1:50 w / w), the solution was incubated at 37°C for 16 hours. Trifluoroacetic acid (TFA, 1% final TFA) was then added to terminate the trypsin reaction. Desalting was then performed using a Solid Phase Extraction (SPE) HRP Plate. The desalting process began with column activation with 800 μl of a 0.1% trifluoroacetic acid (TFA) and 80% acetonitrile (ACN) solution, followed by an equilibration step with 800 μl of 0.1% trifluoroacetic acid (TFA). The fragmented peptides were then loaded onto the column. After washing the column three times with 800 μl of 0.1% trifluoroacetic acid (TFA) each, the column was eluted with 300 μl of a 0.1% trifluoroacetic acid (TFA) and 80% acetonitrile (ACN) solution. The completely dried sample was stored at -80°C until mass spectrometry analysis. In brief, the desalting process consisted of column activation, equilibration, loading of the digested peptide sample, washing, peptide elution, and solvent drying. The dried samples were then reconstituted in 0.1% trifluoroacetic acid (TFA), and their concentrations were measured at a wavelength of 280 nm using a Nanodrop device. The concentrations of all samples were adjusted to 0.2 μg / μl based on the measured concentration values.

[0226]

[0227] - Data-independent mass spectrometry

[0228] DIonex Ultimate 3000 RSLCnano / DO-RPLC reversed-phase liquid chromatography and Orbitrap Exploris TM Data-independent mass spectrometry (DIA) was performed using a 480 mass spectrometer system. In the LC system, the sample was loaded onto a trap column (0.1 mm x 20 mm, 5 μm) equipped with Solvent A (0.1% formic acid, 5% dimethyl sulfoxide in water) 100%. Then, the sample was separated on an analytical column (50 cm x 75 μm, 2 μm) at 50°C with a gradient of 15-40% Solvent B (0.1% formic acid, 5% DMSO in 80% acetonitrile) for 80 min, and the separation was performed over a total of 120 min (0.3 μl / min).

[0229] The MS system was operated in positive mode, and the temperature and voltage of the spray voltage were set to 275°C and 2.5 kV, respectively. The full MS resolution was 60,000, the AGC target value was 500%, and the IT was 20 ms, and the range of 350 - 1500 m / z was analyzed. In MS2, the AGC target value was 300% and the IT was 22 ms, and the MS2 spectrum was obtained by collecting ions at 9 Da each in the precursor mass range of 385 - 1015 m / z. In addition, the resolution was set to 15,000, the collision energy in the HCD cell was set to 28%, and the other parameters were set to the default values. Detailed conditions are shown in [Table 1].

[0230]

[0231] LC-MS / MS SystemThermo Scientific Ultimate 3000 / Thermo Scientific Orbitrap Exploris쪠 480Trap ColumnAcclaim쪠 PepMap쪠 100 C18 HPLC Columns (0.1 ㎜ x 20 ㎜, 5 ㎛)Analytical ColumnEasy-spray PepMap RSLC C18 (50 ㎝ x 75 ㎛, 2 ㎛)Column Temp.50 °CFlow rate0.3 ㎕ / minIon source typeESI positive-ion modePositive ion Spray voltage2,500 VNegative ion Spray voltage600 VIon Transfer tube Temp.275 °CInjection volume5 ㎕Injection amount1 ugMobile phaseA: 0.1 % formic acid, 5 % DMSO in water,B: 0.1 % formic acid, 5 % DMSO in 80 % acetonitrileGradientTimeSolvent A (%)Solvent B (%)095559515856540909510095105955120955MS1 parameter - Full scanOrbitrap Resolution60,000Scan Range350 - 1500 m / zAutomatic gain control (AGC) Target500% (5 x 106)Maximum Injection Time (IT)20 msMS2 parameter - DIAOrbitrap Resolution15,000Scan Range350 - 1500 m / zPrecursor Mass Range385 -1015 m / zAutomatic gain control (AGC) Target300 % (3 x 105)Maximum Injection Time (IT)22 msHCD Collision Energy (NCE)28%Isolation Window9 m / z

[0232] Marker candidate selection

[0233] - Data processing

[0234] LC-DIA-MS / MS data were processed using DIA-NN (version 1.8) with an in-house serum spectral library. Spectra were searched using default settings except that match-between-runs (MBR) was enabled. Identification results were filtered with a false discovery rate (FDR) of 1% at the precursor and protein levels. Peptides and proteins were quantified using the MaxLFQ algorithm, and the quantitative data were quantile normalized.

[0235]

[0236] - Differential expression analysis

[0237] Differentially expressed proteins (DEPs) were defined using an integrative statistical method. Briefly, for each protein, test statistics were calculated using the Students' t test, Wilcoxon-Ranksum test, and log2-median ratio for each comparison. All samples were then randomly permuted 1,000 times to estimate the empirical distributions of the test statistics and log2-median ratio for the null hypothesis. Using the estimated empirical distributions, adjusted p-values ​​for the observed test statistics and log2-median ratio for each protein were calculated. These p-values ​​were then combined using Stouffer's method to calculate an overall p-value. Finally, for each comparison, DEPs were identified as having an overall p-value <0.05 and an absolute log2-median-ratio greater than the 5th and 95th percentile means of the empirical distribution. For each comparison, only proteins expressed at ≥25% in each of the two test groups were used for hypothesis testing.

[0238] Target peptide selection

[0239] To select target peptides, an in silico peptide library was generated for each protein candidate using the UniProtKB / SwissProt human protein database. Target peptide candidates were limited to peptides that met the following criteria: they were unique in the human database, were fully tryptic, had no missed cleavages, were composed of 7–25 amino acids, did not contain cysteine ​​or methionine, and did not correspond to the N-terminal sequence of a protein. In addition, a checklist was created to evaluate various characteristics of each peptide. For each peptide, the number of ragged tryptic sites, the number of asparagine / glutamine / proline amino acids, the presence of an N-glycosylation motif, and the presence of an N-terminal glutamine were considered. From this peptide list, three target peptides were selected for each protein candidate. To this end, peptides included in public MRM databases such as CPTAC assay and SRMAtlas were prioritized. Furthermore, the Human Plasma Peptide Atlas database was utilized to check whether peptides had previously been identified in the human plasma proteome. If target peptides were insufficient, additional peptides were selected based on the presence of ragged tryptic sites, the presence of N-glycosylation motifs and N-terminal glutamine, and the number of asparagine, glutamine, and proline amino acids.

[0240]

[0241] 2. Development of analytical methods

[0242] Establishing analysis conditions

[0243] - Setting analysis conditions

[0244] We aimed to establish optimal conditions for the various biomarkers included in PANCCHECK, namely, the flow rate of LC (Liquid chromatography), column oven temperature, and GS1 (Nebulizer gas), GS2 (Heating gas), TEM (Ion source temperature), IS (IonSpray voltage), CUR (Curtain gas), and CAD (Collision gas) in MS parameters.

[0245]

[0246] - Peptide selection

[0247] To optimize the above conditions, we attempted to select peptides to be used in setting the analysis conditions among 84 existing peptide standards with various compositions by length, hydrophobicity, and charge (Table 2).

[0248]

[0249]

[0250]

[0251] Under the existing analysis conditions (Table 3), 50 pg of each standard was injected into LC-MS / MS (Liquid chromatography-tandem mass spectrometry) and analyzed. 55 peptides with intensities higher than 10,000 cps (Counts per second), which is the intensity level that can generally produce reproducible analysis results, were selected as analysis peptides for setting the analysis conditions (Table 4, Fig. 1).

[0252] Existing analysis conditionsLC & LC-MS / MS systemSCIEX M5 micro LC & SCIEX QTRAP 5500+ColumnAgilent ZORBAX 300SB-C18(0.5 voltage5500Ion source Temp.500 ℃Ion source gas 1(GS1)15 psiIon source gas 2(GS2)50 psiInjection volume5 uLMobile phaseA: 0.1% formic acid in water, B: 0.1% formic acid in acetonitrileGradientTime(min)Solvent A(%)Solvent B(%)0955195525554525.1010028010028.195530955

[0253]

[0254] No.GeneSequenceIntensity(cps)1ALDOAAAQEEYVK45,4362ANPEPAEFNITLIHPK35,1223ANPEPNATLVNEADK49,5434C5FQNSAILTIQPK13,8035C5VFQFLEK238,5956C7LTPLYELVK426,8757CPB1AEDTVTVENVLK26,1128CPB1ELASLHGTK128,3899DOCK10LTGLSEISQR86,23010EVC2TSEGFQAFSK60,46411EVC2TVEDAGQYLHQK156,93112FERMT3VVLAGGVAPALFR18,90613FGAGSESGIFTNTK28,01414FGAESSSHHPGIAEFPSR83,24815FGBAHYGGFTVQNEANK44,03416FGGYEASILTHDSSIR22,00817GAPDHVGVNGFGR24,61118GAPDHGALQNIIPASTGAAK78,20719GP5YLGVTLSPR156,12720HPVGYVSGWGR64,47321HPHYEGSTVPEK78,00722HSPA2DAGTITGLNVLR28,71423HSPA2EIAEAYLGGK89,24024ICAM1LLGIETPLPK94,75725ICAM1VTLNGVPAQPLGPR96,66426IGFBP2HGLYNLK11,50227IGFBP2LIQGAPTIR186,00528ITGA2BALSNVEGFER58,66029ITGA2BVAIVVGAPR77,20430ITIH3EVSFDVELPK11,90231KLKB1IAYGTQGSSGYSLR11,30232LDHBLIAPVAEEEATVPNNK55,25433LDHBDDEVAQLK81,01534MBL2EEAFLGITDEK30,11635MBL2WLTFSLGK77,80636MMP9AVIDDAFAR161,15537MRC1LITASGSYHK33,82038PPIAVSFELFADK21,30839PPIAFEDENFILK92,45040PRDX2LSEDYGVLK53,45041PRDX2ATAVVDGAFK96,86442PRDX6LPFPIIDDR81,11543RAB10FHTITTSYYR27,91444S100A8ALNSIIDVYHK12,70345S100A8GADVWFK136,12446SE RPINA5AVVEVDESGTR32,21847SERPINA5EDQYHYLLDR72,09148TIMP1GFQALGDAADIR38,62649TP M4ASDAEGDVAALNR31,11750VCLSTVEGIQASVK31,31751VCLIPTISTQLK68,78352VWFDGTVTTDWK2 0,60753VWFHIVTFDGQNFK58,36054YWHAHNSVVEASEAAYK46,43855YWHAZFLIPNASQAESK85,528,

[0255] - Column oven temperature optimization

[0256] Column oven temperature is a factor that comprehensively affects peak resolution, peak shape, column lifetime, and sensitivity. The analytical column used was Agilent ZORBAX 300SB-C18 (0.5 x 150 mm, 3.5 um), and the column oven temperature was reviewed up to 60 °C. In general, tailing of peptide peaks may occur at low temperatures, and there is a possibility of variability due to external temperature according to season, etc., so temperatures below 30 °C were not reviewed.

[0257]

[0258] - Flow rate optimization

[0259] The flow rate was examined for a total of four conditions, from 10uL / min to 25uL / min, in 5uL / min intervals, considering the column pressure and configurable conditions. For the M5 microLC, a minimum of 5uL / min can be set; however, this condition was excluded from the review due to high variability in mobile phase ratio and flow.

[0260] When comparing and analyzing a total of four conditions at 5uL / min intervals from 10uL / min to 25uL / min, the greatest number of peptides showed the highest sensitivity at 10uL / min, and the coefficient of variation (CV %) calculated through 5 repeated analyses also showed that the number of peptides exceeding 20% ​​was the lowest at 10uL / min, so the precision was also the best. Therefore, the flow rate for the M5 microLC in use was set to 10uL / min.

[0261]

[0262] - Change the concentration gradient

[0263] In order to increase the resolution of peaks according to multi-peptide analysis while being suitable for the changed flow rate, the concentration gradient was changed. By changing the condition of increasing the ratio of solvent B, which corresponds to the organic solvent in the mobile phase, from the existing condition of 45% to the condition of increasing it to 35%, the elution intensity of the peptides from the column was lowered, which induced an increase in the resolution of each peak. As a result, it was confirmed that the peptide, which was eluted last at a retention time (Rt.) of 14 minutes under the existing condition, was eluted delayed to 18.5 minutes after the change, confirming that the resolution of each peak was improved. Although there is room for further improvement in separation by applying conditions to increase the concentration of organic solvent B to a level lower than 35%, considering the elution of matrix other than the target peptide, the possibility of adding peptides with high hydrophobicity in the future, and the occurrence of peak broadening, a concentration gradient condition to increase the concentration to 35% was ultimately applied (Table 5).

[0264]

[0265] Change existing gradientgradientTime(min)Solvent A(%)Solvent B(%)Flow rate(uL / min)Time(min)Solvent A(%)Solvent B(%)Flow rate(uL / min)0955200955101955195525554525653525.1010025.1010028010028010028.195528.19553095531955

[0266] - GS1 (Nebulizer gas) optimization

[0267] Nebulizer gas (GS1) is a gas that helps form droplets when the target peptides separated through the column are sprayed through the nebulizer. Generally, a higher pressure is used as the flow rate increases. The Qtrap 5500+ equipment can be reviewed from the lowest setting condition of 0 psi to the highest setting condition of 90 psi. However, since high pressure GS1 causes a stopping phenomenon due to an increase in pressure inside the source, a maximum of 70 psi was reviewed.

[0268] As a result of comparing and analyzing a total of 15 conditions at 5 psi intervals from GS1 0 psi to 70 psi, the peptides showing maximum sensitivity at 35 psi were most widely distributed, so the nebulizer gas GS1 was set to 35 psi (Fig. 5).

[0269]

[0270] - GS2 (Heating gas) optimization

[0271] GS2, a heating gas that is heated and injected from the source, promotes the evaporation of droplets containing ions, ultimately condensing the charged analytes. The pressure range of GS2 can be reviewed from the lowest setting condition of 0 psi to the highest setting condition of 90 psi of the Qtrap 5500+ equipment. However, since high pressure GS2 causes a stall phenomenon due to an increase in pressure inside the source, a maximum of 70 psi was reviewed.

[0272] As a result of comparing and analyzing GS2 for a total of 15 conditions at 5 psi intervals from 0 psi to 70 psi, the peptides showing maximum sensitivity at 20 psi were most distributed, so GS2, the heating gas, was set to 20 psi (Fig. 6).

[0273]

[0274] - TEM (Ion source temperature) optimization

[0275] It plays a role in heating the GS2, which is a heating gas, and can be set up to 750°C, but since the electrodes adjacent to the source and which can be affected by the temperature are made of some plastic material and may melt, only 0°C to a maximum of 550°C was examined. As a result of examining a total of 12 conditions at 50°C intervals from 0°C to 550°C, the number of peptides showing maximum sensitivity increased up to 500°C and then decreased at 550°C. Therefore, 500°C, where the number of peptides showing maximum sensitivity was the largest, was set as the optimal condition for TEM (Fig. 7).

[0276]

[0277] - IS (IonSpray Voltage) optimization

[0278] IS is the voltage applied between the electrode and the orifice plate (Corn), and is involved in introducing the analyte ions into the MS. The Qtrap 5500+ equipment used by our company can examine up to 5500 V, but since most peptides are not detected or show very low sensitivity below 2000 V, a total of eight conditions were examined for IS from 2000 V to 5500 V at 500 V intervals. Under the examined conditions, the sensitivity showed a pattern of gradually increasing as the voltage increased, and 5500 V, where the peptides showing the maximum sensitivity were most distributed, was set as the optimal IS condition (Fig. 8).

[0279]

[0280] - CUR (Curtain gas) optimization

[0281] CUR refers to the nitrogen gas flowing between the orifice and the curtain plate, and it prevents contamination of the source and keeps Q0 clean by removing some remaining droplets and neutral ions. A total of 10 conditions were examined at 5 psi intervals, from the lowest setting condition of 10 psi to the highest setting condition of 55 psi for the Qtrap 5500+ equipment. Among the examined conditions, peptides showing maximum sensitivity were most distributed at 45 psi, so the optimal CUR condition was set at 45 psi (Fig. 9).

[0282]

[0283] - CAD (Collision gas) optimization

[0284] Precursor ions collide with CAD, fragmenting and forming product ions. CAD was examined from the lowest setting condition of 2 psi to the highest setting condition of 12 psi, and a total of six conditions were compared and analyzed at 2 psi intervals. Among the examined conditions, peptides showing maximum sensitivity were most abundant at 8 psi, so the optimal CAD condition was set at 8 psi (Fig. 10).

[0285]

[0286] Optimization of analysis conditions

[0287] - MS2 analysis and selection of top 3 ions

[0288] (1) Synthesis of light peptide standard

[0289] To optimize individual analysis conditions for the PANCCHECK candidate markers, consisting of 53 genes and 159 peptides selected through the aforementioned discovery process, light peptides identical to endogenous peptides were utilized as standards for analysis. The light peptide standards were received and used after the sequences were transmitted to the Peptide Synthesis Service Team within the Bio-Production Technology Division of Vertis, which operates a Good Manufacturing Practice (GMP) facility, and a synthesis request was made.

[0290]

[0291] (2) MS2 analysis and selection of top 3 ions

[0292] The MS spectrum analyzed in the EPI (Enhanced product ion scan) mode of the Qtrap 5500+ instrument was used to confirm the equivalence between the received light peptide standard and the requested sequence, and the MS spectrum and MRM chromatogram were checked to select the top 3 ions to be used for MRM analysis. The top 3 ions were selected by comprehensively reviewing the ions detected at high concentrations in the MS spectrum and the ions detected at high concentrations in the MRM chromatogram (Fig. 11).

[0293]

[0294] - Optimization of target peptide conditions (optimization of individual parameters)

[0295] The parameters applied individually to each PANCCHECK candidate marker—declustering potential (DP), entrance potential (EP), collision energy (CE), and collision cell exit potential (CXP)—were compared and analyzed for all top three ions under each condition. Among the reviewed conditions, the condition exhibiting the highest sensitivity was set as the optimal condition (Table 6).

[0296] DP is the voltage applied to the orifice to minimize solvent clusters that may remain after the analyte ions are vacuum-driven into the MS. The adjustable range is 0–300, but an unnecessarily high DP value can cause the target peptide to go undetected, while an excessively low DP value can lead to orifice contamination as well as reduced sensitivity. Therefore, DP optimization was conducted across 13 conditions, with a range of 20–260, at 20-step intervals.

[0297] EP is a parameter involved in the passage of ions through Q0, and a total of 7 conditions were examined, from the lowest setting condition of 2 to the highest setting condition of 15, with intervals of 2 each.

[0298] CE is a parameter that directly affects the fragmentation of precursor ions. Precursor ions receive energy, accelerate into a collision cell, and collide with collision gas (CAD) to form product ions. For the Qtrap5500+, CE can be set from a minimum of 5 to a maximum of 180. While CE optimization can be achieved by examining all CEs within the configurable range, this process is time-consuming, so a method of examining multiple points based on the values ​​calculated by the Skyline program can also be applied. In this experiment, a total of 10 conditions were examined, including 4 conditions below and 5 conditions above the CE values ​​calculated by Skyline at intervals of 2 each. For markers where the sensitivity continuously increased or decreased as the CE increased within the examined CE range and the inflection point could not be identified, the examination interval was increased and reanalyzed to establish the optimal conditions.

[0299] CXP is a parameter involved in the transfer to Q3 of product ions formed by fragmentation in the collision cell. The optimal conditions were set by examining a total of 11 conditions at intervals of 5 from the lowest setting condition of 0 to the highest setting condition of 55 (Fig. 12).

[0300] IDDPEPCECXPA1BG.LLELTGPK.+2y6.light100522.435A1BG.LLELTGPK.+2y6.heavy100522.435A1BG.SSTSPDR.+2y5.light60515.435A1BG.SSTSPDR.+2y5.heavy60515.435A1BG.ATWSGAVLAGR.+2y8.light100531.740A1BG.ATWSGAVLAGR.+2y8.heavy100531.740APOA4.LTPYADEFK.+2y7+2.light80521.630APOA4.LTPYADEFK.+2y7+2.heavy80521.630APOA4.ISASAEELR.+2y8.light100528.940APOA4.ISASAEELR.+2y8.heavy100528.940APOA4.LAPLAEDVR.+2y5.light100527.140APOA4.LAPLAEDVR.+2y5.heavy100527.140APOC3.GWVTDGFSSLK.+2y9.light100926.345APOC3.GWVTDGFSSLK.+2y9.heavy100926.345APOC3.DYWSTVK.+2y4.light6022540APOC3.DYWSTVK.+2y4.heavy6022540APOL1.NEADELR.+2y5.light80523.845APOL1.NEADELR.+2y5.heavy80523.845APOL1.ALADGVQK.+2y6.light80218.735APOL1.ALADGVQK.+2y6.heavy80218.735ALDOA.ELSDIAHR.+3y6+2.light60515.145ALDOA.ELSDIAHR.+3y6+2.heavy60515.145AMY2A.NWGEGWGFVPSDR.+2y4.light1401545.955AMY2A.NWGEGWGFVPSDR.+2y4.heavy1401545.955CD14.VDADADPR.+2y2.light40526.115CD14.VDADADPR.+2y2.heavy40526.115CD14.ELTLEDLK.+2y6+2.light60220.625CD14.ELTLEDLK.+2y6+2.heavy60220.625CD163.GADLSLR.+2y5.light80514.945CD163.GADLSLR.+2y5.heavy80514.945CRP.ESDTSYVSLK.+2y3.light120524.725CRP.ESDTSYVSLK.+2y3.heavy120524.725CRP.GYSIFSYATK.+2y6.light100924.945CRP.GYSIFSYATK.+2y6.heavy100924.945FCGR3A.AVVFLEPQWYR.+2y5.light1201539.550FCGR3A.AVVFLEPQWYR.+2y5.heavy1201539.550FCGR3A.YFHHNSDFYIPK.+4y2.light80515.630FCGR3A.YFHHNSDFYIPK.+4y2.heavy80515.630HGFAC.TTDVTQTFGIEK.+2y7.light140929.840HGFAC.TTDVTQTFGIEK.+2y7.heavy140929.840ITIH3.SLPEGVANGIEVYSTK.+2y14+2.light1801537.855ITIH3.SLPEGVANGIEVYSTK.+2y14+2.heavy1801537.855PFN1.DSPSVWAAVPGK.+2y10+2.light100224.840PFN1.DSPSVWAAVPGK.+2y10+2.heavy100224.840PFN1.STGGAPTFNVTVTK.+2y9.light140530.845PFN1.STGGAPTFNVTVTK.+2y9.heavy140530.845APOA1.DLATVYVDVLK.+2y6.light201533.335APOA1.DLATVYVDVLK.+2y6.heavy201533.335APOA1.AHVDALR.+2y4.light60524.225APOA1.AHVDALR.+2y4.heavy60524.225APOA1.ATEHLSTLSEK.+3y3.light80221.540APOA1.ATEHLSTLSEK.+3y3.heavy80221.540C9.VVEESELAR.+2y7.light80222.345C9.VVEESELAR.+2y7.heavy80222.345C9.GEIHLGR.+3y5+2.light8021240C9.GEIHLGR.+3y5+2.heavy8021240C9.ALPTTYEK.+2y6+2.light80219.625C9.ALPTTYEK.+2y6+2.heavy80219.625ALDOA.ALQASALK.+2y6.light60218.740ALDOA.ALQASALK.+2y6.heavy60218.740HP.DYAEVGR.+2y5.light60222.935HP.DYAEVGR.+2y5.heavy60222.935GP5.TLPAAAFR.+2y6+2.light80917.825GP5.TLPAAAFR.+2y6+2.heavy80917.825MCAM.VSPAAPER.+2y6+2.light60217.325MCAM.VSPAAPER.+2y6+2.heavy60217.325ICAM1.LLGIETPLPK.+2y8.light1001123.545ICAM1.LLGIETPLPK.+2y8.heavy1001123.545SERPINA3.ITLLSALVETR.+2y7.light1201536.855SERPINA3.ITLLSALVETR.+2y7.heavy1201536.855CNDP1.AIHLDLEEYR.+2y2.light1601335.935CNDP1.AIHLDLEEYR.+2y2.heavy1601335.935CNDP1.HLEDVFSK.+3b4.light80211.630CNDP1.HLEDVFSK.+3b4.heavy80211.630CNDP1.DGSTIPIAK.+2y4.light60925.125CNDP1.DGSTIPIAK.+2y4.heavy60925.125ITIH3.ALDLSLK.+2y5.light80919.630ITIH3.ALDLSLK.+2y5.heavy80919.630PEPD.AFTPFSGPK.+2y7.light80918.335PEPD.AFTPFSGPK.+2y7.heavy80918.335NOTCH2.ALGTLLHTNLR.+3y9+2.light80915.445NOTCH2.ALGTLLHTNLR.+3y9+2.heavy80915.445HGFAC.YEYLEGGDR.+2y7.light10022645HGFAC.YEYLEGGDR.+2y7.heavy10022645HGFAC.EALVPLVADHK.+3y7+2.light80717.135HGFAC.EALVPLVADHK.+3y7+2.heavy80717.135MASP1.IEPSQAK.+2y5.light60215.935MASP1.IEPSQAK.+2y5.heavy60215.935IGFBP3.FHPLHSK.+3y5+2.light60215.935IGFBP3.FHPLHSK.+3y5+2.heavy60215.935IGFBP3.FLNVLSPR.+2y4.light80226.235IGFBP3.FLNVLSPR.+2y4.heavy80226.235IGFBP3.YGQPLPGYTTK.+2y8.light12022750IGFBP3.YGQPLPGYTTK.+2y8.heavy12022750SERPINA11.ITPTITNFALR.+2y9+2.light60227.625SERPINA11.ITPTITNFALR.+2y9+2.heavy60227.625PIGR.VYTVDLGR.+2y6.light100225.640PIGR.VYTVDLGR.+2y6.heavy100225.640PIGR.VLDSGFR.+2y6.light60222.535PIGR.VLDSGFR.+2y6.heavy60222.535PON3.STVEIFK.+2y5.light60923.230PON3.STVEIFK.+2y5.heavy60923.230PON3.YVYVADVAAK.+2y8.light80921.940PON3.YVYVADVAAK.+2y8.heavy80921.940PON3.ILIGTVFHK.+3y7+2.light80914.530PON3.ILIGTVFHK.+3y7+2.heavy80914.530SELL.AEIEYLEK.+2y6.light801119.425SELL.AEIEYLEK.+2y6.heavy801119.425SELL.SYYWIGIR.+2y6.light100228.935SELL.SYYWIGIR.+2y6.heavy100228.935SERPINA5.DFTFDLYR.+2y6.light1001531.450SERPINA5.DFTFDLYR.+2y6.heavy1001531.450SERPINA5.EDQYHYLLDR.+3y3.light60923.725SERPINA5.EDQYHYLLDR.+3y3.heavy60923.725VWF.EYAPGETVK.+2y6+2.light100223.440VWF.EYAPGETVK.+2y6+2.heavy100223.440ORM1.SDVVYTDWK.+2y5.light60224.340ORM1.SDVVYTDWK.+2y5.heavy60224.340PRDX2.ATAVVDGAFK.+2y6.light6022130PRDX2.ATAVVDGAFK.+2y6.heavy6022130PRDX2.TDEGIAYR.+2y6.light100225.740PRDX2.TDEGIAYR.+2y6.heavy100225.740PRDX2.GLFIIDGK.+2y6.light100720.235PRDX2.GLFIIDGK.+2y6.heavy100720.235SERPINA3.EIGELYLPK.+2y7.light8022345SERPINA3.EIGELYLPK.+2y7.heavy8022345SERPINA3.ADLSGITGAR.+2y7.light80226.650SERPINA3.ADLSGITGAR.+2y7.heavy80226.650LRG1.LHLEGNK.+2y5.light100222.935LRG1.LHLEGNK.+2y5.heavy100222.935LRG1.DLLLPQPDLR.+2y6.light1001529.945LRG1.DLLLPQPDLR.+2y6.heavy1001529.945LRG1.GQTLLAVAK.+2y7.light60219.145LRG1.GQTLLAVAK.+2y7.heavy60219.145LBP.ITLPDFTGDLR.+2y8.light1201535.650LBP.ITLPDFTGDLR.+2y8.heavy1201535.650LBP.LAEGFPLPLLK.+2y8.light1001528.450LBP.LAEGFPLPLLK.+2y8.heavy1001528.450TNXB.YEVTVVSVR.+2y7.light80924.850TNXB.YEVTVVSVR.+2y7.heavy80924.850GSN.TASDFITK.+2y6.light80920.640GSN.TASDFITK.+2y6.heavy80920.640GSN.AVEVLPK.+2y5.light100217.530GSN.AVEVLPK.+2y5.heavy100217.530GSN.TGAQELLR.+2y6.light100226.840GSN.TGAQELLR.+2y6.heavy100226.840CLEC3B.LDTLAQEVALLK.+2y8.light1001535.245CLEC3B.LDTLAQEVALLK.+2y8.heavy1001535.245CLEC3B.NWETEITAQPDGGK.+2y5.light160950.925CLEC3B.NWETEITAQPDGGK.+2y5.heavy160950.925PI16.WDEELAAFAK.+2y5.light1201531.955PI16.WDEELAAFAK.+2y5.heavy1201531.955IGFBP2.HGLYNLK.+2y6.light100223.745IGFBP2.HGLYNLK.+2y6.heavy100223.745GP5.YLGVTLSPR.+2y7.light80225.750GP5.YLGVTLSPR.+2y7.heavy80225.750VWF.DGTVTTDWK.+2y5.light60222.135VWF.DGTVTTDWK.+2y5.heavy60222.135VWF.HIVTFDGQNFK.+3b4.light60216.930VWF.HIVTFDGQNFK.+3b4.heavy60216.930GAPDH.GALQNIIPASTGAAK.+2y9.light140729.650GAPDH.GALQNIIPASTGAAK.+2y9.heavy140729.650HP.VGYVSGWGR.+2y5.light6092725HP.VGYVSGWGR.+2y5.heavy6092725IGFBP2.LIQGAPTIR.+2y7.light60226.835IGFBP2.LIQGAPTIR.+2y7.heavy60226.835SERPINA5.AVVEVDESGTR.+2y8.light100233.540SERPINA5.AVVEVDESGTR.+2y8.heavy100233.540TIMP1.GFQALGDAADIR .+2y7.light100725.235TIMP1.GFQALGDAADIR.+2y7.heavy100725.235HP.HYEGSTVPEK.+3y3.light60518.44 0HP.HYEGSTVPEK.+3y3.heavy60518.440ITIH3.EVSFDVELPK.+2y8.light1201529.555ITIH3.EVSFDVELPK.+2 y8.heavy1201529.555PPIA.VSFELFADK.+2y7.light1001526.945PPIA.VSFELFADK.+2y7.heavy1001526.945.

[0301] - Check whether endogenous peptides are detected

[0302] The detection of endogenous peptides in serum was confirmed by applying the optimal LC-MRM MS conditions established through the above-mentioned 'analysis condition setting' and 'individual parameter optimization' processes. The detection of endogenous peptides was confirmed by spiking peptide standards into pretreated serum samples (Fig. 13).

[0303] As a result of checking whether the endogenous peptides were detected by applying the optimal conditions, it was confirmed that endogenous peptides were detected in 88 peptides for 43 genes out of 159 peptides in 53 genes in total, showing a detection rate of 81.1% for genes. In addition, among the top 3 detected ions, ions that showed an intensity at which the endogenous peptides were detected more than 10 times compared to the noise and had no interference adjacent to the peak of the target peptide were selected for quantitative ions by marker (Table 7). In addition, when the detection of the reviewed candidate markers was checked by listing them by blood protein concentration, most of the undetected list was found to have relatively low blood concentrations (Fig. 14).

[0304]

[0305]

[0306]

[0307] - Isotope and Transition Selection

[0308] For method validation and verification analysis of the list of 88 peptides from 43 genes in which endogenous peptides were detected, synthesis of isotope-labeled peptide (heavy) standards was required, and a heavy interference test was performed to determine the isotope without interference. The test was performed with a sample pretreated with serum, with the light transition of the selected quantitative ion and the heavy transition using the isotope (Fig. 15), and the isotope without interference of the heavy peak at the retention time for each target peptide was finally selected (Table 9).

[0309]

[0310]

[0311]

[0312] method validation

[0313] - Purpose of method verification

[0314] This study was conducted to comprehensively evaluate the analytical stability and reproducibility of biomarkers included in PANCCHECK, a pancreatic cancer blood test, to secure the reliability of the analysis method. PANCCHECK is a screening method that can help diagnose pancreatic cancer by injecting 25 biomarkers (A1BG, APOA1, 3 types of APOA4, APOC3, APOL1, 3 types of C9, CRP, PFN1, HGFAC, GP5, 3 types of GSN, LRG1, ORM1, PIGR, PON3, SELL, VWF, SERPINA3, and ITIH3) in human serum into an LC and analyzing them in MRM (Multiple Reaction Monitoring) mode, and applying the resulting quantitative values ​​to an algorithm to predict the presence or absence of pancreatic cancer. In this study, we evaluated the reliability of the LC-MRM MS quantitative analysis method for 25 biomarkers to be used in pancreatic cancer screening by evaluating calibration, selectivity, accuracy, precision, matrix effect, carryover, recovery, and stability.

[0315]

[0316] - Reagent information

[0317] The isotope-labeled peptide standards (heavy peptides) for each biomarker used in this study were synthesized and purchased from ANYGEN and BIOSTEM. The stock solutions of each heavy peptide were prepared at a concentration of 1 mg / mL using 5% ACN containing 0.1% formic acid (FA), stored in a -80°C deep freezer, and diluted with 0.1% FA whenever needed. Water, acetonitrile (ACN), formic acid (FA), and trifluoroacetic acid (TFA) used for LC-MS / MS analysis were all LC / MS grade from Thermo Fisher Scientific (USA). Among the reagents used for sample pretreatment, urea and ammonium bicarbonate (ABC) were purchased from Sigma-Aldrich (USA) with a purity of 99% or higher, and dithiothreitol (DTT) and iodoacetamide (IAA) were Bioultra grade products with a purity of 99% or higher, sold by Sigma-Aldrich (USA). Trypsin for protein fragmentation in the sample was purchased from Promega (USA) with a purity of 99% or higher (Table 9).

[0318]

[0319] 표준품 및 시약정보표준품순도(%)Lot No.제조원A1BG: LLE{L(13C6,15N)}TGP{K(13C6,15N2)}98.7K231304ANYGEN, KoreaAPOA1: ATEHLSTLSE{K(13C6,15N2)}97.5K231307APOA4: LTPYADEF{K(13C6,15N2)}98.2K231312APOA4: ISASAEEL{R(13C6,15N4)}98.0K231310APOA4: LAPLAEDV{R(13C6,15N4)}97.0K231311APOC3: GWVTDGFSSL{K(13C6,15N2)}95.3K231314APOL1:AL{A(13C3,15N)}DG{V(13C5,15N)}Q{K(13C6,15N2)}98.1K231315C9: ALPTTYE{K(13C6,15N2)}98.1K231317C9: VVEESELA{R(13C6,15N4)}98.5K231318C9: GEIHLG{R(13C6,15N4)}98.6K231319CRP: ESDTSYVSL{K(13C6,15N2)}98.5K231323PFN1: DSPSVWAAVPG{K(13C6,15N2)}98.9K231329HGFAC: ALVPL{V(13C5,15N)}{A(13C3,15N)}DH{K(13C6,15N2)}95.0K231358GP5: TLPAAA{F(13C9,15N)}{R(13C6,15N4)}97.6K231362ITIH3: ALDLSL{K(13C6,15N2)}97.5K231367GSN: AVEV{L(13C6,15N)}P{K(13C6,15N2)}96.2JT-173112BIOSTEM, KoreaGSN: TASDF{I(13C6,15N)}T{K(13C6,15N2)}97.2JT-173113GSN: TGAQELL{R(13C6,15N4)}97.3JT-173114LRG1: LHLEGN{K(13C6,15N2)}98.3JT-173119ORM1: SDVVYTDW{K(13C6,15N2)}97.5JT-173120PIGR: VYTVDLG{R(13C6,15N4)}97.6JT-173122PON3: YVYVADVA{A(13C3,15N)}{K(13C6,15N2)}97.9JT-173125SELL: AEIEYLE{K(13C6,15N2)}97.6JT-173129VWF: EY{A(13C3,15N)}PGET{V(13C5,15N)}{K(13C6,15N2)}99.1JT-173136SERPINA3: EIGELYLP{K(13C6,15N2)}97.8JT-173132Reagent and Equipment ManufacturerDithiothreitol (DTT)Sigma, USAIodoacetamide (IAA)Ammonium bicarbonate (ABC)UreaTrypsin(V5113)Promega, USATrifluoroacetic acid (TFA)Thermo Fisher, USAFormic acid(FA)Acetonitrile(ACN)WaterC18 column (0.5 x 150 mm, 3.5 um, 300 Å)Agilent, USA100 mg C18 96 well plateWaters, USA.

[0320]

[0321] - Method verification results

[0322] Among the 43 genes and 88 peptides for which method validation was performed, 25 peptides corresponding to 19 genes ultimately passed all method validation criteria (Table 10), and the results are shown in Tables 11 to 14.

[0323]

[0324] 최종 바이오마커 목록No.GeneProteinSequenceAccession No.1ANPEPAminopeptidase NALEQALEKP151442APOA4Apolipoprotein A-IVLTPYADEFKP067273APOC3Apolipoprotein C-IIIGWVTDGFSSLKP026564C9Complement component C9ALPTTYEKP027485CRPC-reactive proteinESDTSYVSLKP027416HGFACHepatocyte growth factor activatorEALVPLVADHKQ047567IGFBP2Insulin-like growth factor-binding protein 2LIQGAPTIRP180658ITIH3Inter-alpha-trypsin inhibitor heavy chain H3ALDLSLKQ060339LRG1Leucine-rich alpha-2-glycoproteinLHLEGNKP0275010ORM1Alpha-1-acid glycoprotein 1SDVVYTDWKP0276311PFN1Profilin-1DSPSVWAAVPGKP0773712PIGRPolymeric immunoglobulin receptorVYTVDLGRP0183313PON3Serum paraoxonase / lactonase 3YVYVADVAAKQ1516614SERPINA3Alpha-1-antichymotrypsinEIGELYLPKP0101115VWFvon Willebrand factorILAGPAGDSNVVKP04275

[0325]

[0326] NO.GeneSequenceLLOQ (ng / uL)CalibrationMatrix effect(CV %)R2정확성(%)LowHigh1A1BGLLELTGPK0.1560.9999100.02.83.52APOA1ATEHLSTLSEK5.3130.999783.36.44.53APOA4ISASAEELR0.3130.9994100.06.83.24APOA4LAPLAEDVR0.1170.9990100.03.02.15APOA4LTPYADEFK0.1560.9999100.03.62.56APOC3GWVTDGFSSLK1.0160.9988100.07.73.77APOL1ALADGVQK0.0590.9997100.05.45.88C9ALPTTYEK0.0390.9999100.05.63.49C9VVEESELAR0.0780.9996100.05.52.510C9GEIHLGR0.0390.9996100.06.65.411CRPESDTSYVSLK0.0590.999583.35.43.412PFN1DSPSVWAAVPGK0.0980.998883.33.74.313HGFACEALVPLVADHK0.0390.9996100.04.45.114GP5TLPAAAFR0.0050.9998100.07.23.315GSNAVEVLPK0.1170.9997100.02.43.216GSNTASDFITK0.1170.9999100.04.03.117GSNTGAQELLR0.2730.9998100.04.94.218LRG1LHLEGNK0.1460.9999100.05.36.019ORM1SDVVYTDWK0.4690.9999100.03.26.320PIGRVYTVDLGR0.0200.999983.34.13.721PON3YVYVADVAAK0.0240.999183.37.53.022SELLAEIEYLEK0.0240.9999100.07.64.423VWFEYAPGETVK0.0440.9998100.08.26.724ITIH3ALDLSLK0.0780.9997100.04.34.425SERPINA3EIGELYLPK0.6250.9999100.07.95.3

[0327] In the above table, the selectivity was ≤ LLOQ 20% for interference.

[0328] NO.GeneSequenceCarry-over(%)Recovery (%)LowMidHigh평균CV평균CV평균CV1A1BGLLELTGPK2.592.010.294.23.197.81.02APOA1ATEHLSTLSEK2.985.59.086.55.097.92.93APOA4ISASAEELR1.989.63.188.46.396.43.34APOA4LAPLAEDVR2.494.23.595.62.696.70.85APOA4LTPYADEFK1.992.66.996.01.996.33.76APOC3GWVTDGFSSLK1.795.63.593.83.297.06.07APOL1ALADGVQK2.681.811.493.73.0105.41.58C9ALPTTYEK2.094.28.591.22.197.01.69C9VVEESELAR3.090.15.693.11.799.12.310C9GEIHLGR4.583.911.093.65.897.61.311CRPESDTSYVSLK0.095.99.491.61.9107.20.512PFN1DSPSVWAAVPGK3.498.85.992.54.799.34.313HGFACEALVPLVADHK2.387.63.195.93.5100.13.314GP5TLPAAAFR18.189.710.096.73.394.51.415GSNAVEVLPK2.197.35.995.04.296.02.116GSNTASDFITK2.596.74.697.93.497.33.617GSNTGAQELLR2.996.13.299.21.796.35.018LRG1LHLEGNK5.2103.09.286.612.399.21.419ORM1SDVVYTDWK5.698.28.792.813.897.13.220PIGRVYTVDLGR0.0105.911.299.41.390.73.521PON3YVYVADVAAK8.998.41.898.22.791.35.622SELLAEIEYLEK19.591.65.798.43.392.45.423VWFEYAPGETVK8.797.81.5102.13.2100.37.924ITIH3ALDLSLK1.490.25.189.11.894.91.425SERPINA3EIGELYLPK2.1101.22.397.82.293.93.0.

[0329] NoGeneSequenceAccuracy (Within-run / Between-run, %)LLOQLowMidHigh1A1BGLLELTGPK-9.1 / -4.9-2.2 / -0.8-8.3 / -4.5-3.6 / -2.42APOA1ATEHLSTLSEK-17 / -8.24.2 / 9.24.7 / 4-4.8 / -2.23APOA4ISASAEELR-11.7 / -19.31.4 / -20.5 / -1.2-3 / -1.74APOA4LAPLAEDVR18.6 / 7.69.3 / 3.84.2 / -0.3-4.8 / -5.25APOA4LTPYADEFK14.4 / 9.11.7 / 0.2-2.8 / -2.9-3 / -2.36APOC3GWVTDGFSSLK7.7 / 6.64.5 / 0.4-0.3 / -3.37.4 / 6.47APOL1ALADGVQK5.1 / 16.212.4 / 12.83.2 / 3.5-3.5 / -2.68C9ALPTTYEK-17.8 / -13.34.9 / 2.17.3 / 2.74.4 / 2.49C9VVEESELAR13.1 / 10.76.9 / 4.25.3 / 1.41.6 / -1.310C9GEIHLGR14 / 14.72.8 / 3.13.4 / 2-2.3 / -1.911CRPESDTSYVSLK-13 / -5.310.2 / 4.50.1 / -0.6-1.4 / -2.812PFN1DSPSVWAAVPGK8.1 / 92.1 / 2.40.2 / -0.8-0.7 / -0.413HGFACEALVPLVADHK9.4 / -2.32.4 / -0.9-1.7 / -1.4-3 / -0.814GP5TLPAAAFR-16.7 / -11.71.6 / -0.8-2.2 / -0.8-4 / -3.115GSNAVEVLPK19.4 / 14.95.8 / 54.1 / 5.1-1 / 1.316GSNTASDFITK-4.2 / 1.20.4 / 2.90.1 / 2.5-0.5 / 3.217GSNTGAQELLR6.7 / 164.6 / 6.7-1.2 / 1.4-1 / -1.218LRG1LHLEGNK-1.8 / 7.74.3 / 2.80.3 / -1.4-1.3 / -5.419ORM1SDVVYTDWK19.1 / 10.75.9 / 3.62.2 / 2-0.6 / 3.320PIGRVYTVDLGR7 / 161.2 / 1.91.8 / 1.2-0.9 / -0.821PON3YVYVADVAAK12.3 / 19.96.4 / 4.91.4 / 0.11.7 / -0.222SELLAEIEYLEK-9.9 / -6.13.6 / 2.52.8 / 1.43.1 / 3.223VWFEYAPGETVK14.7 / 1.34 / 25.6 / 3.7-1.5 / 2.324ITIH3ALDLSLK13.3 / 5.93.6 / 2.46.5 / 5.46.4 / 7.525SERPINA3EIGELYLPK-11.7 / -7.8-9.7 / -11-14.9 / -16-0.6 / 2.2.

[0330] NoGeneSequencePrecision (Within-run / Between-run, %)LLOQLowMidHigh1A1BGLLELTGPK1.2 / 6.32.7 / 213.6 / 5.63.9 / 1.72APOA1ATEHLSTLSEK2.6 / 13.76.3 / 6.41.9 / 0.914.5 / 3.73APOA4ISASAEELR3.7 / 13.33.7 / 4.91.9 / 2.47.5 / 24APOA4LAPLAEDVR3.5 / 14.51.8 / 7.61.6 / 6.44.8 / 0.65APOA4LTPYADEFK1 / 6.81.3 / 2.14.6 / 0.23.7 / 16APOC3GWVTDGFSSLK1.7 / 1.65.6 / 5.75 / 4.45.8 / 1.47APOL1ALADGVQK7.5 / 13.57.6 / 0.41.8 / 0.412.8 / 1.38C9ALPTTYEK3.7 / 7.31.4 / 3.83.8 / 6.37.5 / 2.89C9VVEESELAR3.2 / 32.4 / 3.62.7 / 5.43.4 / 4.210C9GEIHLGR4.3 / 0.96 / 0.41.7 / 1.98.7 / 0.511CRPESDTSYVSLK4.9 / 11.56.2 / 7.72 / 17.3 / 2.112PFN1DSPSVWAAVPGK4 / 1.26 / 0.56.3 / 1.43.9 / 0.413HGFACEALVPLVADHK3.4 / 16.94.2 / 4.77.2 / 0.45 / 3.114GP5TLPAAAFR7.7 / 7.93.2 / 3.53.5 / 1.93.9 / 1.415GSNAVEVLPK2.1 / 5.62.6 / 12.8 / 1.43.2 / 3.216GSNTASDFITK2.8 / 7.61.7 / 3.52 / 3.34 / 5.117GSNTGAQELLR6.2 / 11.42.2 / 2.81.1 / 3.74.9 / 0.218LRG1LHLEGNK4.6 / 12.43.1 / 25.6 / 2.48.1 / 6.119ORM1SDVVYTDWK2 / 10.72.1 / 3.24.2 / 0.14.7 / 5.420PIGRVYTVDLGR15.3 / 118.1 / 14.2 / 0.83.3 / 0.121PON3YVYVADVAAK11.3 / 9.12.1 / 2.13.3 / 1.85.1 / 2.722SELLAEIEYLEK13.3 / 5.75.8 / 1.56.2 / 24.9 / 0.223VWFEYAPGETVK15.9 / 18.74 / 2.66.6 / 2.74.1 / 5.224ITIH3ALDLSLK3.9 / 9.93.3 / 1.72.8 / 1.45.1 / 1.425SERPINA3EIGELYLPK0.9 / 6.12.7 / 2.14.4 / 1.84.7 / 4.

[0331] In the above Tables 13 and 14, the results of stability were for freeze and thaw stability, long-term stability, short-term stability, and stability of treated samples, and the accuracy was ≤ 15%.

[0332]

[0333] Verification

[0334] - Sample analysis

[0335] 438 samples were analyzed using MRM mode of LC-MS / MS for 15 biomarkers selected through method validation testing. Samples included 152 normal individuals, 154 pancreatic cancer cases, and 50 pancreatic benign cases.

[0336]

[0337] - Sample information

[0338] The serum samples used in this study were randomly selected from the collected samples, and were divided into one sample used for verification of calibration, accuracy, precision, and recovery, and six samples used for verification of selectivity, matrix effect, and stability. The one-sample sample was pooled from 22 samples, including 9 males and 13 females with an average age of 57.4 years ranging from 34 to 80 years. The sample included 18 healthy individuals, 3 cases of ovarian cancer, and 1 case of stomach cancer. The six-sample sample was pooled from 36 randomly selected healthy individuals with 6 cases each. The sample included 35 males and 1 female with an average age of 54.1 years ranging from 32 to 80 years.

[0339]

[0340] - Sample collection method

[0341] For this clinical trial, a total of 724 serum samples collected from 2005 to 2022 among the remaining serum samples stored at the Seoul National University Hospital Biobank were registered. However, only 438 serum samples were ultimately confirmed to meet all of the inclusion / exclusion criteria for this clinical trial. Normal samples used in this clinical trial were collected from individuals who were confirmed to be free of pancreatic cancer through screening at the time of blood collection and had no history of diagnosis of pancreatic cancer or other cancers within the past 10 years. Pancreatic cancer samples were collected before surgery from patients with pathologically confirmed pancreatic cancer (Pancreatic Ductal Adenocarcinoma type). The collected samples were adults aged 20 years or older, and the distribution of pancreatic cancer samples by stage was confirmed to be AJCC Stage I (17.5%), Stage II (59.1%), Stage III (16.2%), and Stage IV (7.2%). A total of 50 sera were used for benign pancreatic disease specimens. Specimens collected for the "Establishment of a Bio Repository for Liver, Biliary Tract, Pancreas, and Tumor Research (Seoul National University Hospital IRB No. 0901-010-267)" and "Genetic Analysis of Colorectal Cancer (Seoul National University Hospital IRB No. 1103-125-357)" studies were given priority. Other specimens were provided by the Seoul National University Hospital Biobank.

[0342]

[0343] GroupAge / Other Cancer TypeNumber of SamplesAJCC Stage1234Normal20 - 3922----40 - 4945----50 - 5934----60 - 6930----70 +21----Total152Pancreatic Ductal Adenocarcinoma20 - 392-1-140 - 49734--50 - 59287154260 - 695963912270 +58113296Total154Benign Pancreatic Disease50Grand total356

[0344] 3. Marker Panel Development

[0345] MRM mass spectrometry

[0346] - Data exploration and analysis

[0347] The analyzed quantitative values ​​were used as input data. By analyzing the data type, nature, and quantity, we conducted basic exploratory analysis to select a more efficient deep learning / machine learning model. The PCA analysis results revealed a pattern in which the pancreatic cancer and healthy populations were divided based on PC1 (Figure 16).

[0348]

[0349] - Selection of 12 markers

[0350] To exclude markers with low significance in distinguishing pancreatic cancer from healthy individuals among the 15 markers, we added 5-fold cross-validation to the LGBM model and examined the results of each validation. Table 13 ranks the markers with significance in each fold. Three markers (APOC3, CRP, ORM1) that were ranked low in all five folds were excluded. The final 12 markers (Table 17) were used to develop a model and test its performance using a machine learning algorithm.

[0351] 5-fold cross-validation results No. GeneSequenceFold1Fold2Fold3Fold4Fold51ANPEPALEQALEK37113102APOA4LTPYADEFK221623APOC3GWVTDGFSSLK10571354C9ALPTTYEK710912135CRPESDTSYVSLK1513151096HGFACEALVPLVADHK126125107IGFBP2LIQGAPT IR84141418ITIH3ALDLSLK486789LRG1LHLEGNK14141381510ORM1SDVVYTDWK111210151411PFN1DSPSVWAAVPGK611 211512PIGRVYTVDLGR133881213PON3YVYVADVAAK1131314SERPINA3EIGELYLPK51542715VWFILAGPAGDSNVVK99534

[0352]

[0353] 12 types of markers used in algorithm development No.GeneProteinSequenceAccession No.1ANPEPAminopeptidase NALEQALEKP151442APOA4Apolipoprotein A-IVLTPYADEFKP067273C9Complement component C9ALPTTYEKP027484HGFAChepatocyte growth factor activatorEALVPLVADHKQ047565IGFBP2Insulin-like growth factor-binding protein 2LIQGAPTIRP180656ITIH3Inter-alpha-trypsin inhibitor heavy chain H3ALDLSLKQ060337LRG1Leucine-rich alpha-2-glycoproteinLHLEGNKP027508PFN1Profilin-1DSPSVWAAVPGKP077379PIGRPolymeric immunoglobulin receptorVYTVDLGRP0183310PON3Serum paraoxonase / lactonase 3YVYVADVAAKQ1516611SERPINA3Alpha-1-antichymotrypsinEIGELYLPKP0101112VWFvon Willebrand factorILAGPAGDSNVVKP04275

[0354]

[0355] Algorithm development

[0356] - Data used to develop the predictive model

[0357] A total of 415 samples were used for model training and validation, including healthy individuals, pancreatic cancer, benign pancreatic cancer, and other cancers (Table 18). Predictor variables included CA19-9 levels and concentrations of 12 protein markers.

[0358]

[0359] #samplesratio (%)PDACstage-1271546.5%37.1%stage-29121.9%stage-3256.0%stage-4112.7%Benign5012.0%Healthy15236.6%Stomach20594.8%14.2%Liver204.8%Colorectal194.6%Total415100.0%

[0360]

[0361] - Prediction model

[0362] The prediction model was developed as a stacking ensemble model by combining a total of seven machine learning algorithms to maximize model robustness while achieving optimal performance (Fig. 17). It is a binary classification model that classifies pancreatic cancer, and the ensemble algorithms include tree-based models such as Extra-trees (ET), Light-GBM (LGBM), random-forest (RF), gradient-boosting (GB), XGBoost (XGB), and Ada-boost (Ada) algorithms, and a neural-network-based multi-layer perceptron (MLP). The hyperparameters of each algorithm use default values, and the MLP model consists of one hidden layer with a total of 16 neurons. The final prediction result is evaluated by synthesizing the prediction results of each model using a logistic regression algorithm.

[0363]

[0364] - Model training and validation

[0365] The prediction model was trained and validated using the 5-fold cross-validation method, and its performance was compared with that of the individual models of the stacking model and the logistic regression model using only CA19-9 values.

[0366] In performance comparison tests, predictive models including 12 markers showed better performance (AUC=0.868~0.923) than CA19-9 (AUC=0.826), and the stacking ensemble model designed as a predictive model showed better performance than the individual algorithms (Fig. 18). The predictive model utilizing 12 markers (stacking ensemble) showed higher specificity than CA19-9 in the high sensitivity area (Table 19).

[0367] Using a threshold value based on a model sensitivity of 84, we evaluated model sensitivity by pancreatic cancer stage and its specificity for pancreatic benign and other cancer samples. Sensitivity tended to increase with increasing pancreatic cancer stage (Table 20), and specificity ranged from 60% to 90% for both other cancers and pancreatic benign samples (Table 21).

[0368] Specificity of each model for a given sensitivity ROC AUC12 Specific marker+ CA19-90.92CA19-90.83

[0369] Sensitivity of the prediction model according to pancreatic cancer stage (Pancreatic cancer stage) (Number of samples) (84.4%) (Sensitivity) (Predicted value) (Sensitivity) (Healthy people) (Pancreatic cancer) (1279180.6729110810.893252230.92411380.73)

[0370] Number of samples: 84.4%; Sensitivity: Predicted value: Specificity: Healthy individuals; Pancreatic cancer: 503,218; Pancreatic benign: 0.64; Healthy individuals: 152,129,230.85

[0371]

[0372] 4. Conclusion

[0373] In this study, we identified differentially expressed proteins between normal and patient groups through LC-DIA MS analysis of serum samples without removing high-concentration proteins, and selected 12 final markers (ANPEP, APOA4, C9, HGFAC, IGFBP2, ITIH3, LRG1, PFN1, PIGR, PON3, SERPINA3, VWF) that passed method validation and verification analysis through LC-MRM MS after performing analysis method optimization as a pancreatic cancer diagnostic marker panel. In addition, we developed and verified an algorithm formula, and the combination of CA19-9, which is currently used for pancreatic cancer diagnosis, and the 12 proteins discovered in this study showed a high accuracy of 0.92 based on the AUC, making it a suitable marker combination for direct clinical application.

[0374]

[0375] While specific aspects of the present invention have been described in detail above, it should be apparent to those skilled in the art that these specific descriptions are merely preferred embodiments and do not limit the scope of the present invention. Therefore, the substantial scope of the present invention is defined by the appended claims and their equivalents.

Claims

One or more polypeptides selected from the group consisting of ANPEP (Aminopeptidase N), APOA4 (Apolipoprotein A-IV), APOC3 (Apolipoprotein C-III), C9 (Complement component C9), CRP (C-reactive protein), HGFAC (Hepatocyte growth factor activator), IGFBP2 (Insulin-like growth factor-binding protein 2), ITIH3 (Inter-alpha-trypsin inhibitor heavy chain H3), LRG1 (Leucine-rich alpha-2-glycoprotein), ORM1 (Alpha-1-acid glycoprotein 1), PFN1 (Profilin-1), PIGR (Polymeric immunoglobulin receptor), PON3 (Serum paraoxonase / lactonase 3), SERPINA3 (Alpha-1-antichymotrypsin), and VWF (von Willebrand factor), or a partial fragment thereof; Or a composition for diagnosing pancreatic cancer, comprising as an active ingredient a preparation for measuring the expression level of a gene encoding the same. In the first paragraph, A fragment of the above ANPEP polypeptide has the amino acid sequence of SEQ ID NO: 1 (ALEQALEK); A fragment of the above APOA4 polypeptide has the amino acid sequence of SEQ ID NO: 2 (LTPYADEFK); A fragment of the above APOC3 polypeptide has the amino acid sequence of SEQ ID NO: 3 (GWVTDGFSSLK); A fragment of the above C9 polypeptide has the amino acid sequence of SEQ ID NO: 4 (ALPTTYEK); A fragment of the above CRP polypeptide has an amino acid sequence of SEQ ID NO: 5 (ESDTSYVSLK); A fragment of the above HGFAC polypeptide has the amino acid sequence of SEQ ID NO: 6 (EALVPLVADHK); A fragment of the above IGFBP2 polypeptide has the amino acid sequence of SEQ ID NO: 7 (LIQGAPTIR); A fragment of the above ITIH3 polypeptide has the amino acid sequence of SEQ ID NO: 8 (ALDLSLK); A fragment of the above LRG1 polypeptide has the amino acid sequence of SEQ ID NO: 9 (LHLEGNK); A fragment of the above ORM1 polypeptide has the amino acid sequence of SEQ ID NO: 10 (SDVVYTDWK); A fragment of the above PFN1 polypeptide has the amino acid sequence of SEQ ID NO: 11 (DSPSVWAAVPGK); A fragment of the above PIGR polypeptide has the amino acid sequence of SEQ ID NO: 12 (VYTVDLGR); A fragment of the above PON3 polypeptide has the amino acid sequence of SEQ ID NO: 13 (YVYVADVAAK); A fragment of the above SERPINA3 polypeptide has the amino acid sequence of SEQ ID NO: 14 (EIGELYLPK); A composition characterized in that a fragment of the above VWF polypeptide has an amino acid sequence of SEQ ID NO: 15 (ILAGPAGDSNVVK). In the first paragraph, in an individual having pancreatic cancer, the expression level of one or more genes selected from the group consisting of ANPEP, C9, CRP, IGFBP2, ITIH3, LRG1, ORM1, PIGR, SERPINA3 and VWF or the proteins encoded by them is increased; A composition characterized in that the expression level of one or more genes selected from the group consisting of APOA4, APOC3, HGFAC, PFN1 and PON3 or the proteins encoded by them is reduced. A composition according to claim 1, characterized in that the pancreatic cancer is pancreatic ductal adenocarcinoma (PDAC). In the first paragraph, A composition characterized in that the agent for measuring the expression level of the polypeptide comprises at least one selected from the group consisting of an antibody, an antigen-binding fragment, a ligand, a peptide nucleic acid (PNA), and an aptamer that specifically binds to the polypeptide or a fragment thereof. In the first paragraph, A composition for measuring the expression level of a gene encoding the polypeptide or a fragment thereof, characterized in that it comprises at least one selected from the group consisting of a primer, a probe and an antisense nucleotide that specifically bind to the gene. A diagnostic kit comprising a diagnostic composition according to any one of claims 1 to 6. In paragraph 7, A diagnostic kit characterized in that the above kit is an RT-PCR kit, a DNA chip kit, an ELISA kit, a protein chip kit, a rapid kit, or an MRM (Multiple reaction monitoring) kit. In a biological sample isolated from the target organism, One or more polypeptides selected from the group consisting of ANPEP (Aminopeptidase N), APOA4 (Apolipoprotein A-IV), APOC3 (Apolipoprotein C-III), C9 (Complement component C9), CRP (C-reactive protein), HGFAC (Hepatocyte growth factor activator), IGFBP2 (Insulin-like growth factor-binding protein 2), ITIH3 (Inter-alpha-trypsin inhibitor heavy chain H3), LRG1 (Leucine-rich alpha-2-glycoprotein), ORM1 (Alpha-1-acid glycoprotein 1), PFN1 (Profilin-1), PIGR (Polymeric immunoglobulin receptor), PON3 (Serum paraoxonase / lactonase 3), SERPINA3 (Alpha-1-antichymotrypsin), and VWF (von Willebrand factor), or fragments thereof; Or a method for providing information necessary for diagnosing pancreatic cancer, comprising the step of measuring the expression level of a gene encoding it. In paragraph 9, A fragment of the above ANPEP polypeptide has the amino acid sequence of SEQ ID NO: 1 (ALEQALEK); A fragment of the above APOA4 polypeptide has the amino acid sequence of SEQ ID NO: 2 (LTPYADEFK); A fragment of the above APOC3 polypeptide has the amino acid sequence of SEQ ID NO: 3 (GWVTDGFSSLK); A fragment of the above C9 polypeptide has the amino acid sequence of SEQ ID NO: 4 (ALPTTYEK); A fragment of the above CRP polypeptide has an amino acid sequence of SEQ ID NO: 5 (ESDTSYVSLK); A fragment of the above HGFAC polypeptide has the amino acid sequence of SEQ ID NO: 6 (EALVPLVADHK); A fragment of the above IGFBP2 polypeptide has the amino acid sequence of SEQ ID NO: 7 (LIQGAPTIR); A fragment of the above ITIH3 polypeptide has the amino acid sequence of SEQ ID NO: 8 (ALDLSLK); A fragment of the above LRG1 polypeptide has the amino acid sequence of SEQ ID NO: 9 (LHLEGNK); A fragment of the above ORM1 polypeptide has the amino acid sequence of SEQ ID NO: 10 (SDVVYTDWK); A fragment of the above PFN1 polypeptide has the amino acid sequence of SEQ ID NO: 11 (DSPSVWAAVPGK); A fragment of the above PIGR polypeptide has the amino acid sequence of SEQ ID NO: 12 (VYTVDLGR); A fragment of the above PON3 polypeptide has the amino acid sequence of SEQ ID NO: 13 (YVYVADVAAK); A fragment of the above SERPINA3 polypeptide has the amino acid sequence of SEQ ID NO: 14 (EIGELYLPK); A method characterized in that a fragment of the above VWF polypeptide has an amino acid sequence of SEQ ID NO: 15 (ILAGPAGDSNVVK). In paragraph 9, The biological samples include whole blood, leukocytes, peripheral blood mononuclear cells, buffy coat, plasma, serum, sputum, tears, mucus, nasal washes, nasal aspirate, breath, urine, semen, saliva, peritoneal washings, ascites, cystic fluid, meningeal fluid, amniotic fluid, glandular fluid, pancreatic fluid, lymph fluid, pleural fluid, nipple aspirate, bronchial aspirate, synovial fluid, joint aspirate, trachea. A method characterized in that the secretion is an organ secretion, a cell, a cell extract or a cerebrospinal fluid. In paragraph 9, A method characterized in that the preparation for measuring the expression level of the polypeptide comprises at least one selected from the group consisting of an antibody, an oligopeptide, a ligand, a peptide nucleic acid (PNA), and an aptamer that specifically binds to the polypeptide. In paragraph 9, A method characterized in that the measurement of the expression level of the polypeptide is performed by protein chip analysis, immunoassay, ligand binding assay, MALDI-TOF (Matrix Assisted Laser Desorption / Ionization Time of Flight Mass Spectrometry) analysis, SELDI-TOF (Sulface Enhanced Laser Desorption / Ionization Time of Flight Mass Spectrometry) analysis, radioimmunoassay, radioimmunodiffusion, aukteroni immunodiffusion, rocket immunoelectrophoresis, tissue immunostaining, complement fixation assay, two-dimensional electrophoresis analysis, liquid chromatography-mass spectrometry (LC-MS), liquid chromatography-mass spectrometry / mass spectrometry (LC-MS), liquid chromatography-mass spectrometry / mass spectrometry (LC-MS), western blotting, or enzyme linked immunosorbent assay (ELISA). In paragraph 9, A method characterized in that the measurement of the expression level of the above polypeptide is performed using a multiple reaction monitoring (MRM) method. In paragraph 10, The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 1 is 901.506 for a light peptide and 909.52 for a heavy peptide when the z value is 1; or a value within ±1 range of each value; The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 2 is 1083.536 for a light peptide and 1091.550 for a heavy peptide when the z value is 1; or a value within ±1 range of each value; The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 3 is 1196.595 for a light peptide and 1204.609 for a heavy peptide when the z value is 1; or a value within a range of ±1 of each value; The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 4 is 922.496 for a light peptide and 930.508 for a heavy peptide when the z value is 1; or a value within a range of ±1 of each value; The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 5 is a light peptide when the z value is 1. 1128.542, heavy peptide is 1136.556; or a value within ±1 range of each value; The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 6 is 1191.673 for a light peptide and 1209.708 for a heavy peptide when the z value is 1; or a value within a range of ±1 of each value; The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 7 is 968.596 for a light peptide and 978.604 for a heavy peptide when the z value is 1; or a value within ±1 range of each value; The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 8 is 759.468 for a light peptide and 767.482 for a heavy peptide when the z value is 1; or a value within a range of ±1 of each value; The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 9 is 810.454 for a light peptide and 818.468 for a heavy peptide when the z value is 1; or a value within ±1 range of each value; The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 10 is 1112.526 for a light peptide and 1120.540 for a heavy peptide when the z value is 1; or a value within a range of ±1 of each value; The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 11 is 1213.621 for a light peptide and 1221.635 for a heavy peptide when the z value is 1; or a value within ±1 range of each value; The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 12 is 922.506 for a light peptide and 932.508 for a heavy peptide when the z value is 1; or a value within ±1 range of each value; The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 13 is 1098.59 for a light peptide and 1110.61 for a heavy peptide when the z value is 1; or a value within a range of ±1 of each value; The mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 14 is 1061.588 for a light peptide and 1069.602 for a heavy peptide when the z value is 1; or a value within a range of ±1 of each value; A method characterized in that the mass-to-charge ratio (m / z) of the polypeptide represented by the above sequence number 15 is 1240.696 for a light peptide and 1248.71 for a heavy peptide when the z value is 1; or a value within a range of ±1 of each value. In paragraph 14, A method characterized in that a synthetic peptide in which a specific element of a specific amino acid constituting each polypeptide is substituted with an isotope as an internal standard material when performing the above multiple reaction monitoring; or E. coli beta-galactosidase is used. A method according to claim 16, wherein the synthetic peptide has the same sequence as the sequence represented by SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 and includes a stable isotope. A method according to claim 17, characterized in that the stable isotope is a stable isotope of at least one element selected from the group consisting of carbon and nitrogen. In paragraph 9, A method characterized in that the measurement of the expression level of the gene encoding the above polypeptide is performed by reverse transcription polymerase reaction (RT-PCR), competitive reverse transcription polymerase reaction (Competitive RT-PCR), real-time reverse transcription polymerase reaction (Real-time RT-PCR), RNase protection assay (RPA), Northern blotting, or DNA chip. In paragraph 9, The expression level of the ANPEP (Aminopeptidase N), C9 (Complement component C9), CRP (C-reactive protein), IGFBP2 (Insulin-like growth factor-binding protein 2), ITIH3 (Inter-alpha-trypsin inhibitor heavy chain H3), LRG1 (Leucine-rich alpha-2-glycoprotein), ORM1 (Alpha-1-acid glycoprotein 1), PIGR (Polymeric immunoglobulin receptor), SERPINA3 (Alpha-1-antichymotrypsin) or VWF (von Willebrand factor) polypeptide or a gene encoding the same measured for a biological sample of the target individual is increased compared to a normal control group; A method characterized in that it is predicted that the possibility of developing pancreatic cancer is high when the expression level of APOA4 (Apolipoprotein A-IV), APOC3 (Apolipoprotein C-III), HGFAC (Hepatocyte growth factor activator), PFN1 (Profilin-1) or PON3 (Serum paraoxonase / lactonase 3) polypeptide or a gene encoding the same is decreased compared to a normal control group. A method according to claim 9, characterized in that the pancreatic cancer is pancreatic ductal adenocarcinoma (PDAC). A method for screening a composition for preventing or treating pancreatic cancer comprising the following steps: (a) a step of contacting a test substance with a biological sample containing ANPEP, APOA4, APOC3, C9, CRP, HGFAC, IGFBP2, ITIH3, LRG1, ORM1, PFN1, PIGR, PON3, SERPINA3 and VWF genes or proteins encoded by them or cells expressing them; and (b) a step of measuring the expression level of the protein or the gene in the biological sample; The activity or expression level of the ANPEP, C9, CRP, IGFBP2, ITIH3, LRG1, ORM1, PIGR, SERPINA3 and VWF genes or proteins in the biological sample is reduced; If the activity or expression level of APOA4, APOC3, HGFAC, PFN1, and PON3 genes or proteins increases, the composition is determined to be for the prevention or treatment of pancreatic cancer. A method according to claim 22, characterized in that the pancreatic cancer is pancreatic ductal adenocarcinoma (PDAC). As a diagnostic system for pancreatic cancer, Input section that receives input values; A reading unit including a pre-trained machine learning model to read whether pancreatic cancer has occurred; and Includes an output section that outputs whether pancreatic cancer has occurred; The system wherein the above input value is a measurement value of the expression level of one or more polypeptides selected from the group consisting of SEQ ID NO: 1 (ALEQALEK), SEQ ID NO: 2 (LTPYADEFK), SEQ ID NO: 3 (GWVTDGFSSLK), SEQ ID NO: 4 (ALPTTYEK), SEQ ID NO: 5 (ESDTSYVSLK), SEQ ID NO: 6 (EALVPLVADHK), SEQ ID NO: 7 (LIQGAPTIR), SEQ ID NO: 8 (ALDLSLK), SEQ ID NO: 9 (LHLEGNK), SEQ ID NO: 10 (SDVVYTDWK), SEQ ID NO: 11 (DSPSVWAAVPGK), SEQ ID NO: 12 (VYTVDLGR), SEQ ID NO: 13 (YVYVADVAAK), SEQ ID NO: 14 (EIGELYLPK), and SEQ ID NO: 15 (ILAGPAGDSNVVK) in a biological sample. A system characterized in that in claim 24, the machine learning model is a deep learning model. A system according to claim 24, wherein the biological sample is blood. A system characterized in that, in claim 24, the measurement value of the expression level of the polypeptide is a quantitative value according to mass spectrometry. A system according to claim 27, wherein the mass spectrometry is liquid chromatography-tandem mass spectrometry (LC-MS / MS). A method according to claim 24, characterized in that the pancreatic cancer is pancreatic ductal adenocarcinoma (PDAC).

Citation Information

Patent Citations

  • Markers for diagnosing pancreatic cancer and its use

    KR101384211B1

  • Pancreatic cancer biomarkers and uses thereof

    KR1020130100096A

  • Composition for diagnosing pancreatic cancer and method for diagnosing pancreatic cancer using the same

    KR1020160057352A

  • Fluoride and calcium preperation for remineralizing teeth

    KR1020230147937A

  • Disease determination system using exhaled gas component analysis

    KR1020240151540A