Pancreatic cyst classification

A computer-implemented method using a protein panel distinguishes benign from malignant pancreatic cysts, addressing the low sensitivity of current diagnostics and reducing unnecessary surgeries by accurately classifying cysts, thus improving patient outcomes and reducing costs.

WO2025191030A1PCT designated stage Publication Date: 2025-09-18UNIVERSITY OF HULL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/056791
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-14
Filing Date
2025-03-12
Publication Date
2025-09-18

AI Technical Summary

Technical Problem

Current diagnostic methods for pancreatic cysts have low sensitivity in determining their malignant potential, leading to incorrect assignments and unnecessary surgeries, and there is a need for improved accuracy in classifying cysts as benign, pre-malignant, or malignant to reduce morbidity and mortality.

Method used

A computer-implemented method using a key panel of differentially expressed proteins identified through machine learning, classifying pancreatic cysts as pre-malignant/malignant or benign based on protein abundance data from a liquid sample without the need for resection or histological assessment.

Benefits of technology

Accurately distinguishes between benign and malignant cysts, reducing unnecessary surgeries and enabling earlier treatment for pre-malignant/malignant cysts, thereby improving patient outcomes and reducing medical costs and discomfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025056791_18092025_PF_FP_ABST
    Figure EP2025056791_18092025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a computer-implemented method for determining if a pancreatic cyst of a subject is pre-malignant / malignant or benign, the method comprising a) providing protein abundance data obtained from a liquid sample from the subject ("sample data"); b) providing the sample data as input to a machine learning model trained on a training data set comprising a protein abundance profile of at least three pancreatic carcinomas and at least three presumed benign pancreatic cysts, the protein abundance profile comprising the abundance level of at least 15 proteins selected from Table 1 ("training data"), wherein each pancreatic carcinoma is a distinct pancreatic carcinoma subtype and each presumed benign pancreatic cyst is a distinct presumed benign pancreatic cyst subtype; c) receiving an output from the machine learning model, the output indicative of a pre-malignant / malignant classification or a benign pancreatic cyst classification for the sample data; and d) determining if the pancreatic cyst of the subject is pre-malignant / malignant or benign based on the output received in step c). Related methods and systems for use in performing the methods are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] PANCREATIC CYST CLASSIFICATION

[0002] Field of the Invention

[0003] The present invention relates to computer-implemented methods for determining if a pancreatic cyst of a subject is pre-malignant / malignant or benign using protein abundance data obtained from a liquid sample from the subject. Methods for selecting a treatment for a subject having a pancreatic cyst, and treating a subject having a malignant pancreatic cyst are also provided.

[0004] Background

[0005] Pancreatic cancer is an aggressive disease, which is persistently increasing in incidence and mortality. Many patients present with the disease at an advanced stage, which can reduce the efficacy of therapeutic options. Only 5% of patients survive pancreatic cancer for 10 or more years.

[0006] Pancreatic cancer can arise through a highly heterogeneous group of pancreatic cystic neoplasms (PCNs), including mucinous cystic neoplasms (MCNs) and intraductal papillary mucinous neoplasms (IPMNs). While the prevalence of pancreatic cysts in the general population is high, with a current rate of incidentally detected pancreatic lesions of 8%, it can be difficult to determine the subtype classification of the cyst. The identification of cysts which are likely to become or are malignant is also challenging.

[0007] Current diagnosis of cysts uses the Fukuoka guidelines, which are predominantly based on imaging, such as radiological imaging. These guidelines can only attribute ‘worrisome features’ or ‘high-risk stigmata’ to PCNs (Tanaka et al., 2017) and cannot accurately assign malignant potential. Other diagnosis strategies include the analysis of pancreatic cyst fluid (PCyF). This analysis, which is aimed at enhancing the radiological features diagnostics, is focused on the carcinoembryonic antigen (CEA), amylase and / or mucin levels, along with cytological examination, yet the whole diagnostic ‘package’ has low sensitivity for assigning malignant potential (Singhi et al., 2019 and Soreide et al., 2023).

[0008] The low sensitivity of current diagnosis methods means that many pancreatic cysts can be incorrectly assigned as having malignant or non-malignant potential. Once a pancreatic cyst is assigned as having malignant potential, the cyst is typically removed by resection. For pancreatic cysts which have been incorrectly assigned as having malignant potential, this is unnecessary surgery.

[0009] Recent studies using PCyF have identified distinct mutational profiles of various types of pre-malignant lesions and those that progressed to invasive carcinoma using genomic alteration panels (Paniccia et al., 2023). Findings from previous proteomic studies of PCyF suggested that large biomarker panels are required to decipher malignancy in IPMN-associated carcinomas (Do et al., 2018). However, these studies were based on either low proteomic coverage, or used samples collected only upon resection, and not EUS-guided aspiration. Moreover, the diagnostic accuracy of the proposed biomarker panels in relation to the risk of unnecessary surgery, an unmet clinical need, has not been addressed.

[0010] There remains a need for diagnostic methods having improved accuracy, particularly in relation to the classification of pancreatic cysts as benign, pre-malignant or malignant. Such improved accuracy will reduce the need for unnecessary surgery in cysts identified as benign, while ensuring that treatment can begin at an earlier stage for those cysts identified as pre-malignant or malignant, thereby reducing morbidity and mortality.

[0011] The present invention has been devised in light of the above considerations.

[0012] Summary of the Disclosure

[0013] Broadly, the present inventors have developed an improved method for determining the malignancy or pre-malignancy of a pancreatic cyst. In particular, the present inventors have defined a key panel of proteins which they have identified as being differentially expressed in pre-malignant or malignant cysts compared to presumed benign cysts. The differential expression of the key panel of proteins acts as a precise “fingerprint” or “signature” for premalignant / malignant versus benign cysts. This key panel is utilised by the present method to classify cysts into premalignant / malignant or benign. Advantageously, the present method can identify the likely malignancy of the cyst without the need for resection or histological assessment of the cyst, which previously was considered the only method by which a cyst could be definitively classified. This removes the need for unnecessary surgery for subjects determined to have benign pancreatic cysts, and means that other treatment options can be pursued for subjects who are identified as having premalignant / malignant cysts, which may have increased efficacy, particularly if administered when the malignancy is at an early stage. This also saves discomfort for the subject and reduces time and expense for medical practitioners.

[0014] Accordingly, in a first aspect the present invention provides a computer-implemented method for determining if a pancreatic cyst of a subject is pre-malignant / malignant or benign, the method comprising: a) Providing protein abundance data obtained from a liquid sample from the subject (“sample data”); b) Providing the sample data as input to a machine learning model trained on a training data set comprising a protein abundance profile of at least three pancreatic carcinomas and at least three presumed benign pancreatic cysts, the protein abundance profile comprising the abundance level of at least 15 proteins selected from Table 1 (“training data”), wherein each pancreatic carcinoma is a distinct pancreatic carcinoma subtype and each presumed benign pancreatic cyst is a distinct presumed benign pancreatic cyst subtype; c) Receiving an output from the machine learning model, the output indicative of a premalignant / malignant classification or a benign pancreatic cyst classification for the sample data; and d) Determining if the pancreatic cyst of the subject is pre-malignant / malignant or benign based on the output received in step c). The present inventors have advantageously found that protein abundance data of proteins from Table 1 can be used to accurately distinguish between pre-malignant / malignant and benign pancreatic cysts.

[0015] In some embodiments, the pancreatic cyst of the subject is a non-resected pancreatic cyst. In other words, the pancreatic cyst of the subject has not been resected, optionally either at the time of analysis of the protein abundance data and / or at the time when the liquid sample was obtained from the subject.

[0016] In some embodiments, the at least 15 proteins comprise at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80 or at least 85 proteins selected from Table 1 . In some embodiments the protein abundance profile comprises the abundance level of all 89 proteins of Table 1.

[0017] In some embodiments, the protein abundance profile comprises the abundance level of at least a further 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150 or 250 proteins selected from Table 2. In some embodiments, the protein abundance profile further comprises the abundance level of the 288 proteins of Table 2.

[0018] In some embodiments, the protein abundance profile comprises the abundance level of the 377 proteins of Table 3.

[0019] In some embodiments, the at least 15 proteins comprise one or more of the following proteins of Table 1 : Costars family protein ABRACL, Adenosine kinase, 4-trimethylaminobutyraldehyde dehydrogenase, Arachidonate 5-lipoxygenase-activating protein, ADP-ribosylation factor-like protein 1 , Calcium-binding protein 39, Coatomer subunit gamma-1 , Deaminated glutathione amidase, Pterin-4-alpha-carbinolamine dehydratase, NADPH-cytochrome P450 reductase, Proteasome subunit beta type-3, Resistin, Protein XRP2, 60S ribosomal protein L23, Small nuclear ribonucleoprotein Sm D2, Small nuclear ribonucleoprotein Sm D3, Thiopurine S-methyltransferase and WW domain-binding protein 2.

[0020] In some embodiments, the at least 15 proteins comprise two or more of the following proteins of Table 1 : Costars family protein ABRACL, Adenosine kinase, 4-trimethylaminobutyraldehyde dehydrogenase, Arachidonate 5-lipoxygenase-activating protein, ADP-ribosylation factor-like protein 1 , Calcium-binding protein 39, Coatomer subunit gamma-1 , Deaminated glutathione amidase, Pterin-4-alpha-carbinolamine dehydratase, NADPH-cytochrome P450 reductase, Proteasome subunit beta type-3, Resistin, Protein XRP2, 60S ribosomal protein L23, Small nuclear ribonucleoprotein Sm D2, Small nuclear ribonucleoprotein Sm D3, Thiopurine S-methyltransferase and WW domain-binding protein 2.

[0021] In some embodiments, the at least 15 proteins comprise three or more of the following proteins of Table 1 : Costars family protein ABRACL, Adenosine kinase, 4-trimethylaminobutyraldehyde dehydrogenase, Arachidonate 5-lipoxygenase-activating protein, ADP-ribosylation factor-like protein 1 , Calcium-binding protein 39, Coatomer subunit gamma-1 , Deaminated glutathione amidase, Pterin-4-alpha-carbinolamine dehydratase, NADPH-cytochrome P450 reductase, Proteasome subunit beta type-3, Resistin, Protein XRP2, 60S ribosomal protein L23, Small nuclear ribonucleoprotein Sm D2, Small nuclear ribonucleoprotein Sm D3, Thiopurine S-methyltransferase and WW domain-binding protein 2.

[0022] In some embodiments, the at least 15 proteins comprise four or more of the following proteins of Table 1 : Costars family protein ABRACL, Adenosine kinase, 4-trimethylaminobutyraldehyde dehydrogenase, Arachidonate 5-lipoxygenase-activating protein, ADP-ribosylation factor-like protein 1 , Calcium-binding protein 39, Coatomer subunit gamma-1 , Deaminated glutathione amidase, Pterin-4-alpha-carbinolamine dehydratase, NADPH-cytochrome P450 reductase, Proteasome subunit beta type-3, Resistin, Protein XRP2, 60S ribosomal protein L23, Small nuclear ribonucleoprotein Sm D2, Small nuclear ribonucleoprotein Sm D3, Thiopurine S-methyltransferase and WW domain-binding protein 2.

[0023] In some embodiments, the at least 15 proteins comprise five or more of the following proteins of Table 1 : Costars family protein ABRACL, Adenosine kinase, 4-trimethylaminobutyraldehyde dehydrogenase, Arachidonate 5-lipoxygenase-activating protein, ADP-ribosylation factor-like protein 1 , Calcium-binding protein 39, Coatomer subunit gamma-1 , Deaminated glutathione amidase, Pterin-4-alpha-carbinolamine dehydratase, NADPH-cytochrome P450 reductase, Proteasome subunit beta type-3, Resistin, Protein XRP2, 60S ribosomal protein L23, Small nuclear ribonucleoprotein Sm D2, Small nuclear ribonucleoprotein Sm D3, Thiopurine S-methyltransferase and WW domain-binding protein 2.

[0024] In some embodiments, the at least 15 proteins comprise at least 18 proteins selected from Table 1 , wherein the at least 18 proteins comprise Costars family protein ABRACL, Adenosine kinase, 4- trimethylaminobutyraldehyde dehydrogenase, Arachidonate 5-lipoxygenase-activating protein, ADP- ribosylation factor-like protein 1 , Calcium-binding protein 39, Coatomer subunit gamma-1 , Deaminated glutathione amidase, Pterin-4-alpha-carbinolamine dehydratase, NADPH-cytochrome P450 reductase, Proteasome subunit beta type-3, Resistin, Protein XRP2, 60S ribosomal protein L23, Small nuclear ribonucleoprotein Sm D2, Small nuclear ribonucleoprotein Sm D3, Thiopurine S-methyltransferase and WW domain-binding protein 2.

[0025] In some embodiments, the at least 15 proteins comprise protein XRP2, Thiopurine S-methyltransferase, resistin and at least 12 further proteins selected from Table 1 .

[0026] In some embodiments, the at least 15 proteins comprise four or more of the following proteins of Table 5: Phospholipase A2, Pancreatic alpha-amylase, Alpha-amylase 2B, Protein XRP2, Receptor-type tyrosineprotein phosphatase C, Eosinophil cationic protein, Catechol O-methyltransferase, Protein S100-P, Cytidine deaminase, Glutaredoxin-1 , Myeloid cell nuclear differentiation antigen, Thiopurine S- methyltransferase, Protein S100-A12, Apoptosis regulator BAX, Cadherin-17, Deaminated glutathione amidase, Resistin, Omega-amidase NIT2 and Septin-9. Table 5 contains 19 proteins which are present in Table 1 , Table 3 and Table 6. In some embodiments, the at least 15 proteins comprise five or more of the following proteins of Table 5: Phospholipase A2, Pancreatic alpha-amylase, Alpha-amylase 2B, Protein XRP2, Receptor-type tyrosineprotein phosphatase C, Eosinophil cationic protein, Catechol O-methyltransferase, Protein S100-P, Cytidine deaminase, Glutaredoxin-1 , Myeloid cell nuclear differentiation antigen, Thiopurine S- methyltransferase, Protein S100-A12, Apoptosis regulator BAX, Cadherin-17, Deaminated glutathione amidase, Resistin, Omega-amidase NIT2 and Septin-9.

[0027] In some embodiments, the at least 15 proteins comprise ten or more of the following proteins of Table 5: Phospholipase A2, Pancreatic alpha-amylase, Alpha-amylase 2B, Protein XRP2, Receptor-type tyrosineprotein phosphatase C, Eosinophil cationic protein, Catechol O-methyltransferase, Protein S100-P, Cytidine deaminase, Glutaredoxin-1 , Myeloid cell nuclear differentiation antigen, Thiopurine S- methyltransferase, Protein S100-A12, Apoptosis regulator BAX, Cadherin-17, Deaminated glutathione amidase, Resistin, Omega-amidase NIT2 and Septin-9.

[0028] In some embodiments, the at least 15 proteins comprise 15 or more of the following proteins of Table 5: Phospholipase A2, Pancreatic alpha-amylase, Alpha-amylase 2B, Protein XRP2, Receptor-type tyrosineprotein phosphatase C, Eosinophil cationic protein, Catechol O-methyltransferase, Protein S100-P, Cytidine deaminase, Glutaredoxin-1 , Myeloid cell nuclear differentiation antigen, Thiopurine S- methyltransferase, Protein S100-A12, Apoptosis regulator BAX, Cadherin-17, Deaminated glutathione amidase, Resistin, Omega-amidase NIT2 and Septin-9.

[0029] In some embodiments, the at least 15 proteins comprise the following proteins of Table 5: Phospholipase A2, Pancreatic alpha-amylase, Alpha-amylase 2B, Protein XRP2, Receptor-type tyrosine-protein phosphatase C, Eosinophil cationic protein, Catechol O-methyltransferase, Protein S100-P, Cytidine deaminase, Glutaredoxin-1 , Myeloid cell nuclear differentiation antigen, Thiopurine S-methyltransferase, Protein S100-A12, Apoptosis regulator BAX, Cadherin-17, Deaminated glutathione amidase, Resistin, Omega-amidase NIT2 and Septin-9. Since Table 5 contains 19 proteins which are all found in Table 1 , Table 3 and Table 6, it will be appreciated that the at least 15 proteins may comprise the following proteins of Table 1 : Phospholipase A2, Pancreatic alpha-amylase, Alpha-amylase 2B, Protein XRP2, Receptor-type tyrosine-protein phosphatase C, Eosinophil cationic protein, Catechol O-methyltransferase, Protein S100-P, Cytidine deaminase, Glutaredoxin-1 , Myeloid cell nuclear differentiation antigen, Thiopurine S-methyltransferase, Protein S100-A12, Apoptosis regulator BAX, Cadherin-17, Deaminated glutathione amidase, Resistin, Omega-amidase NIT2 and Septin-9.

[0030] In some embodiments, the protein abundance profile comprises the protein abundance level of the 19 proteins of Table 5 and the abundance level of at least a further 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150 or 250 proteins selected from Table 2. In some embodiments, the protein abundance profile comprises the protein abundance level of the 19 proteins of Table 5 and the abundance level of at least a further 5, 10, 20, 30, 40, 50, 60 or 70 proteins selected from Table 2. In some embodiments, the protein abundance profile comprises the protein abundance level of the 19 proteins of Table 5 and the abundance level of from 5 to 75, optionally of from 5 to 50 proteins selected from Table 2.

[0031] In some embodiments, the protein abundance profile comprises the protein abundance level of the 96 proteins of Table 6. Table 6 comprises 96 proteins, 19 of which are found in Table 5 and the remaining 77 are found in Table 2.

[0032] The at least 15 proteins may comprise no more than 4000 proteins, no more than 3500 proteins, no more than 3000 proteins, no more than 2500 proteins, no more than 2000 proteins, no more than 1500 proteins, no more than 1000 proteins or no more than 500 proteins in total. In some embodiments, the at least 15 proteins comprises of from 15 proteins to 4000 proteins, optionally of from 15 proteins to 1000 proteins, wherein at least 15 proteins are selected from Table 1. In some embodiments, the at least 15 proteins comprise no more than the 377 proteins of Table 3, optionally no more than the 89 proteins of Table 1. In some embodiments, the at least 15 proteins comprise no more than 100 of the 377 proteins of Table 3.

[0033] Typically, the protein abundance data obtained from the liquid sample comprises the abundance level of the same proteins as the protein abundance profile of the training data. Therefore, the protein abundance data will comprise the abundance level of the same at least 15 proteins of Table 1. It will, however, be appreciated that the abundance level of each of these proteins may differ, such that some of the proteins may have a low or undetectable abundance level, or indeed an increased level. It is this difference in abundance level of the specified proteins which is utilised by the present model to classify the pancreatic cyst as premalignant / malignant or benign.

[0034] The protein abundance data may have been obtained by mass spectrometry, aptamer assay-based proteomics, protein extension assay-based proteomics or immuno-based assays (e.g. immunoblotting or ELISA). For example, the protein abundance data may comprise label-free quantitation (LFQ) intensity of the relevant proteins. The mass spectrometry may comprise liquid chromatography-mass spectrometry (LC-MS) and / or data-independent acquisition (DIA) mass spectrometry. In some embodiments the protein abundance data was obtained using Olink.

[0035] In some embodiments, the protein abundance data obtained from the liquid sample is data obtained by mass spectrometry, aptamer assay-based proteomics or protein extension assay-based proteomics which has been normalised to derive protein abundance data. In some embodiments, the protein abundance data is obtained using a proximity extension assay (PEA). Alternatively, the protein abundance data may be data obtained using mass spectrometry.

[0036] The protein abundance profile of each of the at least three pancreatic carcinoma subtypes and at least three presumed benign pancreatic cysts may have been obtained or derived from mass spectrometry, aptamer assay-based proteomics, protein extension assay-based proteomics or immuno-based assays (e.g. immunoblotting or ELISA). For example, the protein abundance profile of the at least three pancreatic carcinoma subtypes and at least three presumed benign pancreatic cyst subtypes may have been obtained by normalisation of data obtained by mass spectrometry, aptamer assay-based proteomics or protein extension assay-based proteomics.

[0037] The protein abundance profile of each of the at least three pancreatic carcinoma subtypes and each of the at least three presumed benign pancreatic cyst subtypes may have been obtained from a liquid sample, as defined herein.

[0038] Reference to “at least three pancreatic carcinomas” or “at least three pancreatic carcinoma subtypes” will be understood to mean pancreatic carcinomas from at least three different subtypes. In some embodiments, at least one of the at least three pancreatic carcinoma subtypes comprises a pancreatic ductal adenocarcinoma (PDAC). In some embodiments, at least one of the at least three pancreatic carcinoma subtypes comprises a colloid carcinoma (CC). In some embodiments, at least one of the at least three pancreatic carcinoma subtypes comprises a basaloid squamous cell carcinoma (BSCC). In some embodiments the at least three pancreatic carcinoma subtypes comprise a colloid carcinoma (CC), a pancreatic ductal adenocarcinoma (PDAC) and / or a basaloid squamous cell carcinoma (BSCC). In some embodiments the at least three pancreatic carcinoma subtypes comprise a colloid carcinoma (CC), a pancreatic ductal adenocarcinoma (PDAC) and a basaloid squamous cell carcinoma (BSCC). It will be appreciated that a colloid carcinoma (CC), a pancreatic ductal adenocarcinoma (PDAC) and a basaloid squamous cell carcinoma (BSCC) are each considered a distinct pancreatic carcinoma subtype.

[0039] In some embodiments, the at least three pancreatic carcinoma subtypes comprise at least four pancreatic carcinoma subtypes.

[0040] In some embodiments, the at least three presumed benign pancreatic cyst subtypes comprise at least four, at least five, at least six or at least seven presumed benign pancreatic cyst subtypes. In some embodiments, the at least three presumed benign pancreatic cyst subtypes comprise at least four presumed benign pancreatic cyst subtypes. In some embodiments, the at least three presumed benign pancreatic cyst subtypes comprise at least seven individual presumed benign pancreatic cyst subtypes. In some embodiments the at least three presumed benign pancreatic cyst subtypes comprises at least one presumed benign pancreatic lymphoepithelial cyst. In some embodiments the at least three presumed benign pancreatic cyst subtypes comprises at least one presumed benign pseudocyst. In some embodiments the at least three presumed benign pancreatic cyst subtypes comprises at least one presumed benign pancreatic serous cystic neoplasm (SCN). In some embodiments the at least three presumed benign pancreatic cyst subtypes comprises at least one presumed benign pancreatic intraductal papillary mucinous neoplasm (IPMN). In some embodiments the at least three presumed benign pancreatic cyst subtypes comprises at least one presumed benign mucinous cystic pancreatic neoplasm (MCN).

[0041] In some embodiments the at least three presumed benign pancreatic cyst subtypes comprise: a) a presumed benign pseudocyst; b) a presumed benign pancreatic serous cystic neoplasm (SCN); and c) a presumed benign pancreatic intraductal papillary mucinous neoplasm (IPMN) or a presumed benign mucinous cystic pancreatic neoplasm (MCN).

[0042] In some embodiments the at least three presumed benign pancreatic cyst subtypes comprise at least four presumed benign pancreatic cyst subtypes comprising a presumed benign pseudocyst, a presumed benign pancreatic serous cystic neoplasm (SCN), a presumed benign pancreatic intraductal papillary mucinous neoplasm (IPMN) and a presumed benign mucinous cystic pancreatic neoplasm (MCN). It will be appreciated that a presumed benign pseudocyst, a presumed benign SCN, a presumed benign IPMN and a presumed benign MCN are each considered a distinct presumed benign pancreatic cyst subtype.

[0043] The pancreatic cyst of the subject may previously have been diagnosed by radiology, cytology and / or biochemistry. For example, the subject may previously have been diagnosed as having altered amylase, CEA levels and / or present or absent mucin from pancreatic cyst fluid (PCyF) sample. For example, the subject may previously have been diagnosed as having a CEA level of over 192 mg / ml from a PCyF sample. In some embodiments, the subject may previously have been diagnosed as having an amylase level from a PCyF sample of less than 250 U / L.

[0044] The machine learning model may comprise a support vector machine (SVM), a linear regression model, a decision tree, a logistic regression model, an artificial neural network, naive bayes or k-nearest neighbour algorithm. The decision tree may comprise a gradient boosting algorithm or a random forest algorithm.

[0045] In some embodiments, the machine learning model comprises a linear regression model. In some embodiments, the machine learning model comprises a support vector machine (SVM) algorithm. In some embodiments, the output comprises a malignancy score of between 0 and 1 . The present inventors have found that when the output comprises a malignancy score, the higher the score, the greater the chance of the cyst being pre-malignant or malignant. Thus, in some embodiments, in d), the pancreatic cyst of the subject is determined as pre-malignant / malignant when the output comprises a malignancy score of > 0.6 and the pancreatic cyst of the subject is determined as benign when the output comprises a malignancy score of < 0.6. In some embodiments, in d), the pancreatic cyst of the subject is determined as pre-malignant / malignant when the output comprises a malignancy score of > 0.7 and the pancreatic cyst of the subject is determined as benign when the output comprises a malignancy score of < 0.7. In some embodiments, in d), the pancreatic cyst of the subject is determined as pre- malignant / malignant when the output comprises a malignancy score of > 0.8 and the pancreatic cyst of the subject is determined as benign when the output comprises a malignancy score of < 0.8. In some embodiments, in d), the pancreatic cyst of the subject is determined as pre-malignant / malignant when the output comprises a malignancy score of > 0.9 and the pancreatic cyst of the subject is determined as benign when the output comprises a malignancy score of < 0.9.

[0046] In some embodiments, the method is for early detection of pre-malignancy / malignancy of the pancreas and / or monitoring progression of pancreatic cyst(s) from benign to pre-malignant / malignant.

[0047] In some embodiments, when the pancreatic cyst of the subject is determined to be benign in step d), the computer-implemented method may be repeated using one or more further liquid sample(s) obtained from the subject at later time point(s) to determine if the pancreatic cyst has developed into a premalignant / malignant pancreatic cyst. The later time point(s) may be at least three months, at least six months, at least 12 months, at least two years, at least three years or at least five years after the first liquid sample was obtained from the subject.

[0048] In some embodiments, the training data set further comprises imaging data of the at least three pancreatic carcinoma subtypes and at least three presumed benign pancreatic cyst subtypes. In such embodiments, step a) may comprise providing protein abundance data obtained from a liquid sample from the subject and imaging data obtained from the subject. The imaging data may comprise radiological images.

[0049] In a second aspect the present invention provides a method of determining if a pancreatic cyst of a subject is premalignant / malignant or benign, the method comprising: a) Obtaining protein abundance data from a liquid sample from the subject; and b) Performing the method of the first aspect to determine if the pancreatic cyst is premalignant / malignant or benign. Step a) may comprise obtaining protein abundance data from a liquid sample from the subject using a protein abundance data acquisition means. The protein abundance data acquisition means may comprise a mass spectrometer. In some embodiments the protein abundance data is obtained by mass spectrometry, aptamer assay-based proteomics, protein extension assay-based proteomics or immunobased assays (e.g. immunoblotting or ELISA). In some embodiments the protein abundance data is obtained by liquid chromatography-mass spectrometry (LC-MS) and / or data-independent acquisition (DIA) or other mass spectrometry. In some embodiments, the protein abundance data is obtained using a proximity extension assay (PEA). In some embodiments the protein abundance data is obtained using Olink.

[0050] In a third aspect the present invention provides a method for selecting a treatment for a subject having a pancreatic cyst, the method comprising: a) Performing the method of the first or second aspect to determine if the pancreatic cyst of the subject is premalignant / malignant or benign; and b) Selecting an anti-cancer treatment for the subject if the pancreatic cyst is determined to be premalignant / malignant.

[0051] In a fourth aspect the present invention provides a method of treatment of a subject having a malignant pancreatic cyst, the method comprising: a) Performing the method of any of the above aspects, wherein the pancreatic cyst of the subject is determined to be premalignant / malignant; and b) Administering an anti-cancer treatment to the subject.

[0052] The anti-cancer treatment of the third or the fourth aspect may comprise surgery, chemotherapy and / or radiotherapy. Preferably, the anti-cancer treatment of the third or the fourth aspect comprises chemotherapy and / or radiotherapy. The anti-cancer treatment of the third or the fourth aspect may comprise surgery. Surgery will be understood to comprise surgical removal of at least the pancreatic cyst. In some embodiments, the chemotherapy comprises gemcitabine, capecitabine, fluorouracil, irinotecan, oxaliplatin, nab-paclitaxel and / or cisplatin.

[0053] According to a fifth aspect, the present invention provides a method of measuring protein abundance in a pancreatic cyst fluid (PCyF) sample, wherein the method comprises measuring the abundance of at least 15 proteins selected from Table 1 in the PCyF sample to obtain a measured abundance level. In some embodiments, the at least 15 proteins selected from Table 1 comprise at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80 or at least 85 proteins selected from Table 1 . In some embodiments, the at least 15 proteins selected from Table 1 comprise all 89 proteins of Table 1. In some embodiments, the method further comprises measuring the abundance level of at least a further 10, 20, 30, 40, 50, 100, 150 or 250 proteins selected from Table 2. In some embodiments, the method further comprises measuring the abundance level of the 288 proteins of Table 2. In some embodiments the method comprises measuring the abundance level of at least the 377 proteins of Table 3 in the PCyF sample to obtain a measured abundance level. In some embodiments the method further comprises comparing the measured abundance level to a stored abundance level. The stored abundance level may comprise a reference abundance level. The method may comprise measuring the abundance level of no more than 4000 proteins, no more than 3500 proteins, no more than 3000 proteins, no more than 2500 proteins, no more than 2000 proteins, no more than 1500 proteins, no more than 1000 proteins or no more than 500 proteins in total. In some embodiments, the method comprises measuring the abundance level of from 15 proteins to 4000 proteins, optionally of from 15 proteins to 1000 proteins, wherein at least 15 proteins are selected from Table 1 . In some embodiments, the method comprises measuring the abundance level of no more than the 377 proteins of Table 3, optionally no more than the 89 proteins of Table 1 .

[0054] In some embodiments, the at least 15 proteins comprise protein XRP2, Thiopurine S-methyltransferase, resistin and at least 12 further proteins selected from Table 1. In some embodiments, the at least 15 proteins comprise four or more of the following proteins of Table 5: Phospholipase A2, Pancreatic alphaamylase, Alpha-amylase 2B, Protein XRP2, Receptor-type tyrosine-protein phosphatase C, Eosinophil cationic protein, Catechol O-methyltransferase, Protein S100-P, Cytidine deaminase, Glutaredoxin-1 , Myeloid cell nuclear differentiation antigen, Thiopurine S-methyltransferase, Protein S100-A12, Apoptosis regulator BAX, Cadherin-17, Deaminated glutathione amidase, Resistin, Omega-amidase NIT2 and Septin-9.

[0055] In some embodiments, the at least 15 proteins comprise the following proteins of Table 5: Phospholipase A2, Pancreatic alpha-amylase, Alpha-amylase 2B, Protein XRP2, Receptor-type tyrosine-protein phosphatase C, Eosinophil cationic protein, Catechol O-methyltransferase, Protein S100-P, Cytidine deaminase, Glutaredoxin-1 , Myeloid cell nuclear differentiation antigen, Thiopurine S-methyltransferase, Protein S100-A12, Apoptosis regulator BAX, Cadherin-17, Deaminated glutathione amidase, Resistin, Omega-amidase NIT2 and Septin-9.

[0056] In some embodiments, the method comprises measuring the protein abundance level of the 96 proteins of Table 6.

[0057] In a sixth aspect, the present invention provides a computer-implemented method for determining the subtype of a pancreatic cyst of a subject, the method comprising: a) Providing protein abundance data obtained from a liquid sample from the subject (“sample data”); b) Providing the sample data as input to a machine learning model trained on a training data set comprising a protein abundance profile of at least two pancreatic carcinomas and at least two presumed benign pancreatic cysts, the protein abundance profile comprising the abundance level of at least 10 proteins selected from Table 1 (“training data”), wherein each pancreatic carcinoma is a distinct pancreatic carcinoma subtype and each presumed benign pancreatic cyst is a distinct presumed benign pancreatic cyst subtype; c) Receiving an output from the machine learning model, the output indicative of a subtype classification for the sample data; and d) Determining the subtype of the pancreatic cyst of the subject based on the output received in step c).

[0058] The protein abundance data, pancreatic carcinoma subtypes and presumed benign pancreatic cyst subtypes may be as defined above for the first aspect.

[0059] The protein abundance profile may comprise the abundance level of at least 20, at least 25, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80 or all proteins from Table 1 . In some embodiments, the protein abundance profile comprises the abundance level of at least 50 proteins selected from Table 1 .

[0060] In some embodiments, the method is for determining the subtype of a presumed benign pancreatic cyst of a subject and the output in c) is indicative of a benign subtype classification. In some embodiments, the output in c) is indicative of a pancreatic cancer subtype classification, optionally a PDAC pancreatic cancer subtype or at least one other pancreatic cancer subtype classification. In some embodiments, the output in c) is indicative of a pancreatic cancer subtype classification or a benign classification.

[0061] According to a seventh aspect, the present invention provides a computer-implemented method of stratifying a population of subjects assigned for pancreatic surgery, to identify subjects who do not require pancreatic surgery, wherein each subject has or is suspected of having a pancreatic cyst, the method comprising: a) Performing the method of the first aspect to determine if the pancreatic cyst is premalignant / malignant or benign; and b) Identifying subjects as not requiring pancreatic surgery if the pancreatic cyst is determined as benign. Advantageously, this method avoids the need for unnecessary surgery in instances where subjects do not necessarily require treatment or surgery.

[0062] In some embodiments, the subject is suspected of having a premalignant or malignant cyst.

[0063] In a further aspect the present invention provides a system comprising: a) a processor; and b) a computer readable medium comprising instructions that, when executed by the processor, cause the processor to perform the steps of the method of the first or the sixth aspect.

[0064] The system may comprise a computing device communicably connected, such as e.g. through a network, to protein abundance data acquisition means, such as a mass spectrometer, and / or to one or more databases storing protein abundance data. The one or more databases may further store one or more of: training data, learned parameters (e.g. malignancy score thresholds), clinical and / or sample related information, etc. The computing device may be a smartphone, tablet, personal computer or other computing device. The computing device is configured to implement a method of the first or the sixth aspect of the invention. In alternative embodiments, a computing device may be configured to communicate with a remote computing device, which is itself configured to implement the method of the first or the sixth aspect of the invention. In such cases, the remote computing device may also be configured to send the result of the method to the computing device. Communication between the computing device and the remote computing device may be through a wired or wireless connection, and may occur over a local or public network such as e.g. over the public internet. The protein abundance data acquisition means may be in wired connection with the computing device, or may be able to communicate through a wireless connection, such as e.g. through Wi-Fi and / or over the public internet. The connection between the computing device and the protein abundance data acquisition means may be direct or indirect (such as e.g. through a remote computer). The protein abundance data acquisition means is configured to acquire protein abundance data. Any sample preparation process that is suitable for use in the determination of a protein abundance profile may be used within the context of the present invention. The protein abundance data acquisition means may comprise a system for high-throughput, parallel sequencing.

[0065] In a further aspect the present invention provides one or more computer readable media comprising instructions that, when executed by one or more processors, cause the one or more processors to performs the steps of the method of the first or the sixth aspect. Throughout this specification, including the claims which follow, unless the context requires otherwise, the word “comprise” and “include”, and variations such as “comprises”, “comprising”, and “including” will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps.

[0066] Any section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0067] Embodiments of the present invention will now be described by way of example and not limitation with reference to the accompanying figures. However various further aspects and embodiments of the present invention will be apparent to those skilled in the art in view of the present disclosure.

[0068] The present invention includes the combination of the aspects and preferred features described except where such a combination is clearly impermissible or is stated to be expressly avoided. These and further aspects and embodiments of the invention are described in further detail below and with reference to the accompanying examples and figures.

[0069] Brief Description of the Figures

[0070] Figure 1. Summary of the analytical approach used to stratify PCyF samples according to malignant potential based on signatures of cyst fluid proteome. Clinical assessment of cyst types was conducted based on radiological, cytological and biochemical characteristics, according to Fukuoka criteria. Cysts with high-risk features were resected and subsequently underwent pathological assessment, whilst patients with low-risk lesions (with non-malignant characteristics) either underwent follow-up surveillance or were discharged. Unresectable cysts due to advanced disease progression (i.e. those whose type couldn’t be determined due to the presence of high malignant components) are also mapped. MS-based proteomic profiling of PCyF samples was performed in Data-independent Acquisition (DIA) mode. The MS-based proteomic dataset was divided into a training set (11 PCyF samples; 7 presumed benign and 4 carcinoma samples), and a test set (20 samples; 6 presumed benign, 9 dysplastic with respective pathology available, 2 non-operative and 3 carcinomas). A signature of differentially expressed proteins (DEPs) between groups of training set was derived upon statistical analysis and inclusion of those proteins reduced in carcinomas and those specifically detected in all carcinomas and not detected in any presumed benign samples (see Figure 3). A machine learning approach based on Support Vector Machine (SVM) was used to build a classifier for cross-validation of training and prediction of test cohort, respectively. This model classifier used for assigning malignancy score to the samples and calculating the receiver operating characteristic (ROC) curve of the classifier. Abbreviations: R: radiology; C: cytology, B: biochemistry; IPMN: intraductal papillary mucinous neoplasm; MCN: mucinous cystic neoplasms; SCN: serous cystic neoplasm; Pseudo: pseudocyst; LE: lymphoepithelial cyst; N / A: Not available.

[0071] Figure 2. Comprehensive mapping and malignancy risk stratification of various pancreatic cysts using cystic fluid proteomic signatures. (A) Principal component analysis of pancreatic cystic fluid proteomic data. Principal component analysis (PCA) is presented in a 2D graph of principal components - 1 and -2 (PC1 and PC2). PCA was performed using the protein abundance values, represented by normalised label-free quantitation (LFQ) intensity, and assigned the largest variance to the difference between 31 anonymised PCyF samples acquired from various cyst types (Component 1 , 36.3%). Each circle represents the mean protein abundance per sample based on two technical replicates. The colour of each dot designates the pathology associated with each sample. For the training set, four types of presumed benign (light grey ellipses) and three types of carcinoma (dark grey ellipses) samples were used. (B) Cluster analysis of differential protein expression identified in pancreatic cystic fluid samples of training set. Clustered heatmap of 89 differentially expressed proteins (DEPs) between non-malignant and carcinoma samples of training set. Heatmap visualising proteins upon hierarchical clustering based on Z-scores. Data is presented as Iog2 label-free LFQ intensity values, which represent protein abundance. Values for each protein (rows) and for each sample (columns) are colored based on the protein abundance, in which high (red) and low (blue) values are indicated in the color scale bar shown at the bottom of the figure. The dendrogram was built using Euclidean distance in Perseus software. Boxplots of risk-score distribution and corresponding cyst type, resection and pathology status in (C) training and (D) test cohorts. (C and D) Malignancy score for each sample was assigned by a Support Vector Machine (SVM) model classifier generated based on top 89 DEP signature derived from between non-malignant and carcinoma samples of training set. The area under the receiver operator characteristic (ROC) curve was used to measure performance of Fukuoka high-risk features and 89 DEP panel on discrimination of not malignant cysts from others (dysplastic and carcinomas) and assignment of malignant potential of these lesions in (E) training and (F) test sets. Diagnostic performance (or accuracy) was assigned in terms of specificity ((True Negatives) / (True Negatives + False Positives)) and sensitivity ((True Positives) / (True Positives + False Negatives)) and area under the receiver operating characteristic curve (AUROC) values. Abbreviations: R: radiology; C: cytology, B: biochemistry; IPMN: intraductal papillary mucinous neoplasm; MCN: mucinous cystic neoplasms; SCN: serous cystic neoplasm; Pseudo: pseudocyst; LE: lymphoepithelial cyst; PDAC: Pancreatic ductal adenocarcinoma; BSCC: Basaloid squamous cell carcinoma; CC: Colloid carcinoma. GP: Groove pancreatitis; LGD: Low- grade dysplasia; HGD: High-grade dysplasia. N / A: Not available.

[0072] Figure 3. Flowchart for the discovery of a predictive signature for pancreatic cancer based on cyst fluid proteome. Potential markers of pancreatic cancer identified in cyst fluid were discovered by a 7- step procedure. More specifically, of the 3,469 proteins identified, 3,461 were quantifiable in individual pancreatic cyst fluid (PCyF) samples and had valid label-free quantitation (LFQ) intensity values in at both two technical replicates in any biological replicate. Next, of the 3,339 proteins quantified in training set (4 carcinoma vs 7 presumed benign samples), 1 ,385 had at least 75% valid label-free quantitation (LFQ) values within at least one sample group. Upon removal of proteins associated with contamination of samples with blood, 1 ,205 proteins used for statistical analysis and 377 differentially expressed proteins (DEPs) were identified as statistically significant (student’s t-test; FDR-adjusted p-value <0.05; s0=2.00) between the presumed benign vs carcinoma groups. Of the 377 DEPs, 374 were increased in carcinomas and 3 were decreased. DEPs reduced in carcinomas and DEPs specifically detected in all carcinomas and not detected in any presumed benign samples (n=86) were included in the final protein signature (n=89).

[0073] Figure 4. Hierarchical clustering of training cohort samples based on differential protein abundance. Protein abundance of samples was assigned by label-free quantitation (LFQ)-based on mass spectrometry analysis of pancreatic cystic fluid (PCyF) samples. (A) Volcano plot of protein abundance (represented by Iog2 LFQ intensity difference) in two comparison groups of training cohort (7 presumed benign and 4 carcinomas) plotted against the significance (showed by -log-io false discovery rate (FDR)-adjusted p-value< 0.05 and s0=1) identifying 377 differentially expressed proteins (DEPs) in two comparison groups of training cohort (7 presumed benign and 4 carcinomas). The DEPs that were significantly expressed in each comparison group are indicated as colored dots (dark grey for upregulated DEPs and light grey for downregulated DEPs). (B) Cluster analysis of DEPs identified in PCyF samples of training set. Clustered heatmap of 377 DEPs between presumed benign and carcinoma samples upon hierarchical clustering based on LFQ Z-scores. Data is presented as Iog2 label-free LFQ intensity values, which represent protein abundance. Values for each protein (rows) and for each sample (columns) are colored based on the protein abundance, in which high (red) and low (blue) values are indicated in the color scale bar shown at the bottom of the figure. The dendrogram was built using Euclidean distance in Perseus software.

[0074] Figure 5. Validation of mass spectrometry data from cohort A (cohort of Figures 1-3) in pancreatic cyst fluid using proximity extension assay (antibody-based technique). (A-B) Subsets of differentially expressed proteins (DEPs) previously identified in pancreatic cyst fluid (PCyF) samples using mass spectrometry (MS) data from training set (from cohort A) were analysed by proximity extension assay (PEA), an antibody-based technique with multiplexing capacity. Plots shown correlation analysis of the Iog2 fold-change of (A) 96 (proteins from Olink panels, Table 6 matched to 377-DEP list,) and (B) 19 (right; proteins from Olink panels, Table 5 matched to 89-DEP panel ASSIGN1 , Table 5) DEPs measured in PCyF samples of the training set in cohort A by data independent acquisition (DI A) mass spectrometry (MS; x-axis) and PEA (y-axis). Protein expression / abundance was measured as label-free quantitation (LFQ) or normalized protein expression (NPX) data for DIA MS and PEA analysis, respectively. Pearson correlation coefficient (R) with associated p value is shown. Figure 6. Prediction of malignant potential of pancreatic cystic lesions in cohort B (2ndvalidation cohort) based on multi-protein panel ASSIGN1 in pancreatic cyst fluid and comparison to Fukuoka guidelines-based diagnosis. (A) Dot plot of malignancy classification score distribution for heterogeneous pancreatic cystic lesions (PCLs) from the test set with corresponding cyst type, resection and pathology status indicated (Key). The malignancy risk classification score (0: low; 1 : high) for each sample was assigned by a support vector machine classifier using the 89-DEP panel of Table 1. Cancerous lesions had higher malignancy risk classification scores compared to others (far right dotted box). (B-C) The area under the receiver operator characteristic (AUROC) curve, sensitivity and specificity were used to measure the diagnostic performance (accuracy) from predicting dysplasia (B) and (C) highgrade dysplasia or carcinoma of PCLs based on histopathology of resected cysts and compared between preoperative assessments according to (B-top and C-top) Fukuoka guidelines (Fukuoka high-risk; top graphs) and (B-bottom and C-bottom) ASSIGN1 -based malignancy risk classification score (89 DEP; bottom graphs). Diagnostic accuracy was assigned in terms of specificity ([True Negatives] / [True Negatives + False Positives]) and sensitivity ([True Positives] / [True Positives + False Negatives]) and AUROC values. (A-C) Abbreviations: R: radiology; C: cytopathology, B: biochemistry; IPMN: intraductal papillary mucinous neoplasm; MCN: mucinous cystic neoplasms; SCN: serous cystic neoplasm; LE: lymphoepithelial cyst; PDAC: Pancreatic ductal adenocarcinoma; BSCC: Basaloid squamous cell carcinoma; CC: Colloid carcinoma. GP: Groove pancreatitis; LGD: Low-grade dysplasia; IGD: Intermediate-grade dysplasia; HGD: High-grade dysplasia. N / A: Not available.

[0075] Figure 7. Investigation of the utility of pancreatic cyst fluid-based biomarkers in serum. (A-C) Unsupervised principal component analysis (PCA) plot of 19 differentially expressed proteins (DEPs) panel (Table 6) identified and validated in PCyF by (A) data independent acquisition (DI A) mass spectrometry (MS) and (B) proximity extension assay PEA (middle), together with (C) data from PEA analysis of matched serum samples (right). Each circle represents the (A) mean label-free quantitation (LFQ) or (B-C) normalized protein expression (NPX; middle and right) values per sample. The colour of each dot designates the pathology associated with each sample (see Key).

[0076] Detailed Description

[0077] In describing the present invention, the following terms will be employed, and are intended to be defined as indicated below.

[0078] “and / or” where used herein is to be taken as specific disclosure of each of the two specified features or components with or without the other. For example “A and / or B” is to be taken as specific disclosure of each of (i) A, (ii) B and (iii) A and B, just as if each is set out individually herein. As used herein “malignant” refers to the presence of abnormal cells, which may otherwise be referred to as cancerous cells in or adjacent to the cyst, wherein the malignant cells have invaded nearby tissues in or outside of the pancreas. In malignant conditions, the abnormal cells also have the potential to metastasize or have already metastasized.

[0079] In the context of the present invention, “abnormal cells” will be understood to refer to cells which contain one or more mutations leading to a distinct abnormal histological phenotype, as compared to non- cancerous normal cells. Typically, “abnormal cells” will also be understood to have an excessive level of proliferation, relative to non-cancerous normal cells.

[0080] “Non-cancerous normal cells” will be understood to refer to cells which possess a normal histological phenotype for the pancreas or pancreatic cyst, and which proliferate at a normal rate for the pancreas. Non-cancerous normal cells may also lack mutations.

[0081] The term “pre-malignant”, as used herein, refers to the presence of abnormal cells in or adjacent to the pancreatic cyst that is likely to progress to become malignant. For example, the term “pre-malignant” may refer to the presence of cancerous abnormal cells which have not yet invaded nearby tissue or metastasized.

[0082] As used herein, “benign” refers to the absence of cancerous abnormal cells in or adjacent to the cyst. In other words, the cyst does not contain cancerous abnormal cells. The cyst may present low levels (grade) of or no dysplasia. The term “benign” may be used interchangeably with the term “non-malignant”.

[0083] In the context of the present invention, the term “presumed benign” as used herein, will be understood to refer to a cyst which is classified as high risk based on the Fukuoka guidelines, but which has not been histologically assessed. It will therefore be appreciated that in such embodiments, the level of dysplasia has not been assessed. The Fukuoka guidelines are detailed in Tanaka et al., 2017, which is herein incorporated by reference in its entirety. Briefly, the Fukuoka guidelines provide a risk stratification based on "high-risk stigmata" or "worrisome features". The clinical characteristics for high-risk stigmata comprise obstructive jaundice in a patient / subject with cystic lesion of the head of the pancreas, mural nodule size > 5mm and / or main pancreatic duct diameter > 10mm. For high-risk stigmata, immediate surgery is recommended. The clinical characteristics for worrisome features comprise patients with pancreatitis, cyst size > 3mm, mural nodule size < 5mm, thickened cyst walls, main duct diameter of 5-9 mm, abrupt change in the main duct caliber with distal pancreatic atrophy, lymphadenopathy and / or rapid rate of cyst growth > 5 mm within 2 years in combination with elevated serum level of carbohydrate antigen (CA)19-9. Patients with worrisome features are recommended for endoscopic ultrasound. If specific worrisome features are observed during endoscopic ultrasound, including definite mural nodes > 5 mm, main duct involvement and / or cytology suspicious or positive for malignancy, immediate surgery is recommended.

[0084] All patients with cysts of <3 cm in size without “worrisome features” undergo surveillance according to size stratification.

[0085] “Distinct subtype”, as used herein, will be understood to mean that each subtype of benign cyst / pancreatic carcinoma has a distinct clinical presentation and morphology. A pancreatic cyst may be referred to as a pancreatic cystic neoplasm if abnormal changes have previously been confirmed cytologically or histologically. Subtypes of pancreatic cysts include, but are not necessarily limited to intraductal papillary mucinous neoplasm (IPMN), mucinous cystic neoplasms (MCN), serous cystadenoma neoplasm (SCN), pseudocyst and lymphoepithelial cyst.

[0086] As used herein, “protein abundance data” refers to data which is representative of protein levels or abundance. The protein abundance data may comprise direct protein expression levels, for example as measured by mass spectrometry, aptamer assay-based proteomics or protein extension assay-based proteomics. Alternatively, the protein abundance data may be inferred by gene expression levels, for example by determining mRNA levels. mRNA levels may be measured by RT-PCR or Northern Blots. The protein abundance data may alternatively be inferred by gene expression levels obtained from genome sequencing such as whole genome sequencing data. Gene expression levels may be measured by analysing mutations and other alterations of DNA.

[0087] In embodiments where the protein abundance data is inferred by gene expression levels, it will be appreciated that the expression level of a gene encoding the protein will be used to infer the abundance of this protein in the liquid sample.

[0088] In the context of the present invention, it will be appreciated that the training data comprises a protein abundance profile for each of the at least three pancreatic carcinoma subtypes and for each of the at least three presumed benign pancreatic cyst subtypes. Each protein abundance profile will typically provide a profile of the abundance level of the same proteins.

[0089] As used herein “differentially expressed protein” refers to a protein identified by the inventors as having a differential level of abundance / expression between benign or presumed benign pancreatic cysts and pancreatic carcinomas. Each differentially expressed protein has been identified by the inventors as having a decreased level of abundance in pancreatic carcinomas compared to benign or presumed benign pancreatic cysts, or as being expressed / detectable and / or increased in pancreatic carcinomas but having an undetectable abundance level in benign or presumed benign pancreatic cysts.

[0090] The differentially expressed proteins are listed in Tables 1 , 2, 3, 5 and 6. Pancreatic carcinoma may otherwise be referred to as pancreatic cancer.

[0091] As used herein “liquid sample” refers to biological fluid. For example, the liquid sample may comprise urine, blood, serum, or pancreatic cyst fluid (PCyF). The sample may be one which has been freshly obtained from a subject or may be one which has been processed and / or stored prior to making a determination (e.g. frozen, fixed or subjected to one or more purification, enrichment or extractions steps). The liquid sample may comprise a PCyF, blood or serum sample. In some embodiments, the liquid sample comprises a serum sample. In particular, the liquid sample may comprise a PCyF sample. The liquid sample may comprise a PCyF sample previously obtained by endoscopic ultrasound-guided fine needle aspiration (EUS-FNA) or any other methods of acquisition.

[0092] The sample is preferably from a mammalian (such as e.g. a mammalian liquid sample or a liquid sample from a mammalian subject, including in particular a model animal such as mouse, rat, etc.), preferably from a human (such as e.g. a human liquid sample or a sample from a human subject). Further, the sample may be transported and / or stored, and collection may take place at a location remote from the protein abundance data acquisition location, and / or the computer-implemented method steps may take place at a location remote from the sample collection location and / or remote from the protein abundance data acquisition location (e.g. the computer-implemented method steps may be performed by means of a networked computer, such as by means of a “cloud” provider).

[0093] As used herein “subject” refers to an animal or human. The subject may be mammalian (e.g. a human, a non-human primate, a cat, dog, horse, donkey, sheep, pig, goat, cow, mouse, rat, rabbit or guinea pig). The subject is preferably a human (such as e.g. a human identified as having a pancreatic cyst). The subject is preferably a human identified as having a pancreatic cyst, wherein the pancreatic cyst has not been resected.

[0094] As used herein "treatment" and “therapy” refer to reducing, alleviating or eliminating one or more symptoms of the disease which is being treated, relative to the symptoms prior to treatment.

[0095] In the context of the present invention, “assigned for pancreatic surgery” refers to a subject being identified by a clinician as likely to require surgery to remove a pancreatic cyst or part of it. The clinician may have identified the subject as being likely to require surgery using the Fukuoka or other relevant guidelines and / or using radiological images of the cyst or analysis of amylase, CEA and / or mucin levels from a blood, serum and / or pancreatic cyst fluid (PCyF) sample from the subject. The systems and methods described herein may be implemented in a computer system, in addition to the structural components and user interactions described. As used herein, the term “computer system” includes the hardware, software and data storage devices for embodying a system or carrying out a method according to the above-described embodiments. For example, a computer system may comprise a processing unit such as a central processing unit (CPU) and / or graphics processing unit (GPU), input means, output means and data storage, which may be embodied as one or more connected computing devices. Preferably the computer system has a display or comprises a computing device that has a display to provide a visual output display. The data storage may comprise RAM, disk drives or other computer readable media. The computer system may include a plurality of computing devices connected by a network and able to communicate with each other over that network. It is explicitly envisaged that computer system may consist of or comprise a cloud computer.

[0096] The methods described herein may be provided as computer programs or as computer program products or computer readable media carrying a computer program which is arranged, when run on a computer, to perform the method(s) described herein. As used herein, the term “computer readable media” includes, without limitation, any non-transitory medium or media which can be read and accessed directly by a computer or computer system. The media can include, but are not limited to, magnetic storage media such as floppy discs, hard disc storage media and magnetic tape; optical storage media such as optical discs or CD-ROMs; electrical storage media such as memory, including RAM, ROM and flash memory; and hybrids and combinations of the above such as magnetic / optical storage media.

[0097] Tables

[0098] Table 1 : 89 proteins differentially expressed between pre-malignant / malignant and presumed benign pancreatic cysts.

[0099] ble 2: Additional 288 proteins differentially expressed between pre-malignant / malignant and presumed benign pancreatic cysts (377 proteins of Table - 89 proteins of Table 1 ).

[0100]

[0101]

[0102]

[0103]

[0104]

[0105]

[0106]

[0107]

[0108]

[0109]

[0110]

[0111]

[0112]

[0113] ble 3: 377 proteins differentially expressed between pre-malignant / malignant and presumed benign pancreatic cysts (includes the 89 proteins of Table + 288 proteins of Table 2).

[0114]

[0115]

[0116]

[0117]

[0118]

[0119]

[0120]

[0121]

[0122]

[0123]

[0124]

[0125]

[0126]

[0127]

[0128]

[0129]

[0130]

[0131]

[0132]

[0133] Table 4: Clinical and pathological characteristics of patient cohort A.

[0134] Table 5: 19 proteins differentially expressed between pre-malignant / malignant and presumed benign pancreatic cysts. These 19 proteins are present in Table 1, Table 3 and Table 6, below.

[0135] ble 6: 96 proteins differentially expressed between pre-malignant / malignant and presumed benign pancreatic cysts (19 proteins of Table 5 + 77 oteins from Table 2).

[0136]

[0137]

[0138]

[0139] ble 7. Patient cohort B (2ndvalidation cohort) and cyst clinical information. Pancreatic cyst fluid (PCyF) specimens (n=15) obtained by endoscopic rasound fine-needle aspiration along with corresponding clinical, imaging, and diagnostic surgical pathology follow-up (Methods). Cyst types were diagnosed sed on radiological, cytological and biochemical characteristics / features according to the 2017 International Association of Pancreatology (IAP) Fukuoka idelines3. All cysts were diagnosed with either high-risk stigmata (HRS) or worrisome features (WF), according to Fukuoka criteria3. Pancreatic cystic lesions with S and / or WF with suspicious or positive for malignancy cytology and / or biochemistry were recommended for resection and subsequent histopathological sessment, whilst patients with cysts with WF and non-suspicious or negative for malignancy cytology and / or biochemistry (“presumed benign” and therefore resected) either underwent follow-up surveillance or were discharged. Histological evaluation for specific grade (low, intermediate or high) of dysplasia or type of alignancy determined by an expert in gastrointestinal pathology. Abbreviations: F: Female; M: Male; EUS: endoscopic ultrasound; FNA: fine-needle aspiration; MN: intraductal papillary mucinous neoplasm; MCN: mucinous cystic neoplasms; SCN: serous cystic neoplasm; N / A: Not available.

[0140]

[0141] The following is presented by way of example and is not to be construed as a limitation to the scope of the claims.

[0142] Examples

[0143] INTRODUCTION

[0144] \Ne profiled a highly heterogeneous cohort of 31 PCyF-sampled pancreatic cysts (Table 4) by data- independent acquisition (DIA) mass spectrometry (MS)-based proteomics combined with machine learning analysis, generating proteomic signatures and a comprehensive multi-protein panel exposing PCN heterogeneity and enabling the assignment of malignant potential (Figure 1).

[0145] MATERIALS AND METHODS

[0146] Clinical cohort - Study samples

[0147] Pancreatic cyst fluid (PCyF) specimens (n=31 ; n=15 female; n=16 male; average age at diagnosis=63 years ± 12.9; Table 4) collected for the ethically-approved Tumour Regulatory Molecules as Markers of Malignancy in Pancreatic Cystic Lesions (TEM-PAC, NCT03536793; REC Reference: 18 / LO / 0736; IRAS PROJECT ID: 236870) project in the Queens Centre for Oncology and Haematology located in the Castle Hill Hospital (Hull, UK). PCyF specimens obtained by endoscopic ultrasound (EUS)-fine-needle aspiration (FNA) along with corresponding clinical, imaging, and diagnostic surgical pathology follow-up. Surgical pathology diagnoses were based on the 2017 International Association of Pancreatology (IAP) Fukuoka guidelines (Tanaka et al., 2017). Assessment of cyst types was conducted based on radiological, cytological and biochemical characteristics / features. Cysts with high-risk features (according to Fukuoka criteria as defined in Tanaka et al., 2017) were recommended for resection and subsequent pathological assessment, whilst patients with low-risk lesions (benign pathology) either underwent followup surveillance or were discharged. Unresectable cysts due to disease progression along with lesions which type couldn’t be assessed due to the presence of high malignant components were also included in the study. Samples with a confirmed diagnosis of the specific grade (low-, intermediate- or high-grade) of dysplasia or type of malignancy determined by an expert in gastrointestinal pathology were included in the study. PCyF from all patients, including those underwent surgical resection, were aspirated and stored at -80°C.

[0148] Proteomics sample processing.

[0149] Samples were diluted to 5% SDS and normalised for protein amount using BCA quantification. Samples underwent an S-Trap™ micro spin column digestion workflow (Wojtkiewicz et al., 2021) of reduction in 20 mM DTT for 30 min at room temperature (RT), followed by alkylation with 40 mM of iodoacetamide for 30 min at RT and in the dark. For S-Trap™ binding, proteins were acidified to 1 .2% phosphoric acid, and diluted 7-fold with 90% methanol / 100 mM TEAB. Protein was loaded onto the S-Trap™ cartridge and digested with a 1 :25 ratio of trypsin:protein for 2 hours at 47°C. Peptides were eluted in 50% ACN: 0.1% TFA, dried down, and stored at -80°C until further use.

[0150] LC-MS / MS

[0151] Peptides were resuspended in 2% acetonitrile (ACN), 0.1% formic acid and separated using a Dionex Ultimate 3000 nano-ultra high pressure reversed-phase chromatography system synced to an Orbitrap Ascend Tribrid Mass Spectrometer (Thermo Scientific), as previously described (O’Brien et al., 2023). Briefly, peptides were separated on an EASY-Spray PepMap RSLC C18 column (500 mm x 75 pm, 2 pm particle size; Thermo Scientific) over a 60 min gradient of 2-35% ACN in 5% DMSO, 0.1% formic acid and at 250 nL / min. The column temperature was maintained at 50°C with the aid of a column oven. The mass spectrometer was operated in positive polarity mode with a capillary temperature of 275°C. Data- independent acquisition (DIA) mode was utilised for automated switching between MS and MS / MS acquisition (O’Brien et al., 2023).

[0152] The Olink Explore 3072 proximity extension assay (PEA) platform is based on high-multiplex matched pairs of antibodies labelled with unique DNA oligonucleotides bound to their respective proteins (Wik et al., 2021). When both antibodies of a pair bind the target protein simultaneously, their respective conjugated oligonucleotides are brought into proximity, facilitating hybridization. The oligonucleotide sequence is then extended by DNA polymerase, amplified, and measured by quantitative polymerase chain reaction - PCR (Lundberg et al., 2011). Unique sample indexes were added to allow pooling of the DNA amplicons. Each Olink library was purified using Agentcourt AMPure XP magnetic beads. Quality was assessed using a Tapestation. Olink Explore 3072 consists of 8 panels of 384 assays analysed by next-generation sequencing using Illumina’s NovaSeq 6000 platform. Each assay is based on a pair of polyclonal antibodies. Four Olink Explore 3072 panels panels were selected based on the highest number proteins of interest: Olink Oncology (368 biomarkers), Olink Neurology (367 biomarkers), Olink Cardiometabolic (369 biomarkers) and CardiometabolicJI (367 biomarkers). Raw analyte expression values after PCR underwent multiple rounds of transformation by Olink and were returned as normalized protein expression (NPX) values. NPX values are not absolute quantifications, but an indication of relative concentration of each analyte.

[0153] All MS raw files were deposited in PRIDE under the unique identifier PXD045289. Data

[0154] Data were analyzed using DIA-NN (version 1 .8) with all settings as default (Demichev et al., 2020). Specifically, we allowed for a maximum of one tryptic missed cleavage and fixed modifications of N-term M excision and carbamidomethylation of Cys residues. No variable modifications were selected. A Homo sapiens UniProt database (20,370 entries, retrieved on April 16, 2021) was used for the analysis. A default threshold of 1% false discovery rate (FDR) was used at the peptide and protein levels. We used the library-free mode of DIA-NN to generate precursor and corresponding fragment ions in silico from the UniProt database (Demichev et al., 2020). The software also generates a library of decoy precursors (negative controls). Retention time alignment was performed using endogenous peptides, and peak scores were calculated by comparison of peak properties between observed and reference spectra. The ‘Match Between Runs’ feature was enabled to match high-resolution MS1 features between runs.

[0155] Statistical and bioinformatics analysis

[0156] All statistical and bioinformatics analyses were done using Perseus (v 2.0.10.0) (Tyanova et al., 2018). Protein abundance was calculated based on the label-free quantitation (LFQ) of normalized spectrum intensity, was Iog2-transformed, and the dataset potential contaminants were removed (Frankenfield et al., 2022). The steps followed to derive the 89-DEP signature are described in Figure 3. Briefly, dataset was filtered based on valid LFQ values in both technical duplicates of at least one biological sample and median used for normalization (z-score) and imputation. Missing values were imputed based on a normal distribution using default (width, 0.15; downshift, 1.8) settings in Perseus. Then, classical (high abundant) plasma proteins (Anderson & Anderson 2002), erythrocyte contamination-associated proteins (Geyer et al., 2019) and immunoglobulins were excluded from dataset containing at least 75% valid LFQ values within at least one sample group (benign vs carcinoma) in the training cohort. For pairwise MS-based proteomic data comparisons, two-tailed paired Student’s t-tests with a permutation-based on FDR- adjusted p-value of 1% (<0.05; applying 1000 randomizations) and SO value of 1 were conducted to identify statistically differential expressed proteins (DEPs). Next, only DEPs that were detected in all carcinomas and not detected in any presumed benign samples or DEPs reduced in carcinomas were included in final protein signature. Principal component analysis, heat maps with z-score values of Iog2 LFQ intensities and volcano plots were also generated in Perseus (Tyanova et al., 2018).

[0157] To investigate the prediction power of combinations of DEPs to distinguish malignant potential in PCyF samples we used a built-in Support Vector Machine (SVM) classifier in Perseus (Tyanova et al., 2016). Identified DEPs were for cross-validation and prediction using the training and test cohorts, respectively. Classification parameter optimization, classification feature optimization, and classification (cross- validation and prediction) were operated in sequence to identify the minimum number of DEPs that classified the samples with the smallest error. Next, DEPs selected as the optimal features for cross- validation and prediction runs and SVM classifier run based on leave-one-out cross-validation. Receiver operating characteristic (ROC) curves and corresponding area under curve (AUC) values were generated based on the SVM classifier scores of the training set that calculated during cross-validation on the selected DEP signatures. Then, we validated the model using the independent test set. Groups of signatures with both the highest accuracy and the highest AUC were selected as the most relevant features. The optimal cutoff for calculating sensitivity and specificity was determined as a value corresponding to the maximum value of balanced accuracy, defined as the average of the sensitivity and specificity. AUROC plots and optimal cutoff calculations were performed usng the R package pROC (Robin et al., 2011).

[0158] EXAMPLE 1: PROFILING OF HIGHLY HETEROGENEOUS COHORT OF PCYF-SAMPLED PANCREATIC CYSTS

[0159] \Ne first examined whether a simple, unsupervised analysis would be sufficient to expose PCyF sample heterogeneity and cluster them accordingly. Principal component analysis (PCA) of normalized abundance levels of 3,461 proteins quantified in individual 31 samples demonstrated that proteome profiling of PCyF has the capacity to discriminate malignant lesions from others (Figure 2A) and map various cyst types according to their individual comprehensive proteomic signatures. In particular, we observed that PCA was able to discriminate benign subtypes from malignant subtypes. Thus, the PCA mapped a benign SCN subtype, a benign pseudocyst subtype, a benign MCN subtype, a benign IPMN subtype, malignant IPMN subtypes and other malignant cysts.

[0160] EXAMPLE 2: GENERATION OF A COMPUTER-IMPLEMENTED MODEL FOR DETERMINING IF A PANCREATIC CYST OF A SUBJECT IS PRE-MALIGNANT / MALIGNANT OR BENIGN

[0161] Next, we used a training set selected from 31 samples to decipher a generic multi-protein panel which is associated with pancreatic carcinoma (Figure 3). For the training set, we used only presumed benign (four types) and carcinoma (three types) samples due to the high complexity / heterogeneity of intermediate (dysplastic) stages of pancreatic cancer. Through statistical analysis (student’s t-test; FDR- adjusted p-value <0.05; s0=2.00), we derived a signature of 377 differentially expressed proteins (DEPs) between malignant and non-malignant groups (Figure 4A and Table 3). This distinct multiple-protein panel was sufficient to detect proteomic differences, which could undeniably discriminate the two groups of the training set (Figure 4B).

[0162] Next, DEPs which abundance is reduced in carcinomas (n=3) and DEPs that were specifically detected in all carcinomas of training cohort but not detected in any presumed benign samples (n=86) were included in the final protein signature (n=89; Figure 2B and Table 1). The 99-DEP signature then was used to build and validate a Support Vector Machine (SVM) classifier which clearly separated the presumed benign from carcinomas by assigning a “malignancy” score (0: low; 1 : high) for the selected training cohort groups (Figure 2C). Then, we evaluated the capacity of this multi-protein model to predict malignancy and stratify low and high-risk PCNs using the test set (“cohort A”, Figure 2D). This analysis assigned the malignant score in presumed benign (not resected; Figure 2D-left) and Fukuoka high-risk (resected or unresectable or non-operable; Figure 2D-right) PCyF samples. Our 89 DEP signature also allowed the discrimination of cases resected, due to Fukuoka high-risk features, into two distinctive cohorts (Figure 2D-right). This demonstrates the potential of our 89 DEP model to discriminate cysts characterised by low dysplasia according to malignancy score, which could be potentially used to inform decision on surgery or surveillance. Importantly, our data indicates that higher malignancy scores were assigned to intermediate- and high-grade dysplastic cysts and carcinomas compared to low grade ones using the 89-DEP signature (Figure 2D-right).

[0163] Next, the performance of 89 DEP signature in diagnosis / prediction of malignant lesions was compared to initial clinical assessment performed according to Fukuoka high risk features using the training (Figure 2E) and test (Figure 2F) sets. For the training set, the sensitivity, specificity and mean AUROC values of our model were ideal (AUROC=1 .00), similar to Fukuoka high-risk assessment (Figure 2E), depicting the vast proteomic differences between carcinomas and non-malignant samples. In the test set, the 89-DEP signature had lower power in discriminating low grade dysplastic PCNs from non-malignant ones (mean AUROC of 0.75; Figure 2F-left bottom) compared to Fukuoka high-risk assessment (mean AUROC of 1 .00; Figure 2F-left top). At the same time, this DEP panel demonstrated its superiority over Fukuoka guidelines-based diagnosis (mean AUROC of 0.71 ; Figure 2F-right top) in assigning high malignant potential to intermediate- and high-grade dysplastic cysts along with carcinomas, culminating in a mean AUROC of 1.00 (Figure 2F-right bottom).

[0164] EXAMPLE 3: VALIDATION OF METHODS FOR OBTAINING PROTEIN ABUNDANCE DATA

[0165] \Ne then examined if a proximity extension assay (PEA), an antibody-based technique with multiplexing capacity, could be used to accurately obtain protein abundance data. An Olink proximity extension assay was used.

[0166] Briefly, subsets of the differentially expressed proteins (DEPs) previously identified in the pancreatic cyst fluid samples (PCyF) of the training set of the prior examples (previous obtained using mass spectrometry (MS) data) were analysed by PEA. Figure 5A shows a correlation analysis plot of the Iog2 fold-change of 96 proteins from Olink panels (listed in Table 6) matched to the 377-DEP list of the prior examples, while Figure 5B shows a correlation analysis of the Iog2 fold-change of 19 proteins (listed in Table 5) from Olink panels matched to the 89-DEP panel ASSIGN1) DEPs of the prior examples. Pearson correlation coefficient (R) analysis for PEA and MS data of PCyF were strongly correlated for both (A) 96 (r = 0.56, P < 0.001) and (B) 19 (r = 0.79, P < 0.001 ; right) DEPs analysed. This validates that mass spectrometry can be used to obtain protein abundance data and also confirms that Olink or other proximity extension assays are also suitable to obtain protein abundance data. EXAMPLE 4: VALIDATION OF COMPUTER-IMPLEMENTED MODEL

[0167] Next, the computer-implemented model trained on the 89 DEP signature of Example 2 was validated in a second validation cohort of patients (“cohort B”). Patient details are shown in Table 7. Performance was compared to initial clinical assessment performed according to Fukuoka high risk features.

[0168] As shown in Figure 6A, the multi-protein model was able to predict malignancy and stratify low and high- risk PCNs. As we observed with cohort A, the 89 DEP signature allowed the stratification of neoplastic PCLs with the same post-operative diagnosis (low grade dysplasia) and clinical management (resection / surgery) into two distinct cohorts (see middle two dotted boxes). Presumed benign cases are shown in the far left dotted box, while those considered high risk in the far right dotted box. In this validation cohort, two of the presumed benign cysts had higher malignancy scores (above 0.7), however, we believe that this is because these particular patients have a history of chronic pancreatitis (with possible associated markers of high inflammation). Pancreatitis is also known to predispose patients to the development of pancreatic cancer (Raimondi et al., 2009). Therefore, these two higher malignancy scores may be due to inflammation-associated proteins in the 89 DEP signature.

[0169] The performance of the computer-implemented model on cohort B was then compared to initial clinical assessment performed according to Fukuoka high risk features. The area under the receiver operator characteristic (AUROC) curve, sensitivity and specificity were used to measure the diagnostic performance (accuracy) from predicting dysplasia (Figure 6B) and (Figure 6C) high-grade dysplasia or carcinoma of PCLs based on histopathology of resected cysts and compared between preoperative assessments according to (Figure 6B-top and C-top) Fukuoka guidelines (Fukuoka high-risk; top graphs) and (Figure 6B-bottom and C-bottom) ASSIGN1 -based malignancy risk classification score (89 DEP; bottom graphs).

[0170] The 89-DEP signature had lower power in discriminating low grade dysplastic PCNs from non-malignant ones (mean AUROC of 0.73; Figure 6B-botom) compared to Fukuoka high-risk assessment (mean AUROC of 1.00; Figure 6B- top). However, as shown for cohort A, this DEP panel was superior to Fukuoka guidelines-based diagnosis (mean AUROC of 0.81 ; Figure 6C-bottom) in assigning high malignant potential to intermediate- and high-grade dysplastic cysts along with carcinomas, with a mean AUROC of 0.93 (Figure 6C- bottom).

[0171] EXAMPLE 5: INVESTIGATION OF UTILISATION OF PROTEIN ABUNDANCE DATA FROM SERUM

[0172] \Ne then investigated if protein abundance data could be obtained from a serum sample, rather than a PCyF sample. We also assessed the utility of using data independent acquisition (DIA) mass spectrometry, versus proximity extension assay (PEA) to obtain the protein abundance data from serum or a PCyF sample. Unsupervised principal component analysis (PCA) of the 19 differentially expressed proteins (DEPs) panel identified in Example 3 and validated in PCyF are shown in Figure 7. Figure 7(A) shows data independent acquisition (DIA) mass spectrometry (MS) in the validated PCyF sample. Figure 7B shows proximity extension assay PEA in the validated PCyF sample, while Figure 7C shows data from PEA analysis of matched serum samples. The 19 DEPs showed potential to discriminate malignant lesions (clustered together away from others) regardless of the method (MS or PEA) or liquid biopsy (PCyF or blood / serum) used.

[0173] DISCUSSION

[0174] In summary, we performed an in-depth label-free quantitative proteomic data analysis of PCyF samples using machine learning methods to provide evidence that a multiple-protein signature can successfully stratify high-risk patients with PCNs according to their malignant potential. The predictive ability of this PCyF proteome signature is dependent on multiple proteins accounting for the heterogeneity of PCN subtypes in various pre-invasive stages (van Huijgevoort et al., 2019). Our study demonstrates the clinical utility of deep proteome profiling of PCyF for risk stratification of different types of pancreatic cysts which could contribute to early diagnosis and guide management of pancreatic cancer. This could provide a tool / test for early detection of pancreatic cancer and eventually improve survival, avoid unnecessary surgery or excessive surveillance with detrimental effects on patients and reduce healthcare resource consumption. We have also demonstrated that proteome profiling can be achieved using serum samples, providing a minimally invasive, quick and accurate tool / test for early detection of pancreatic cancer.

[0175] References

[0176] A number of publications are cited above in order to more fully describe and disclose the invention and the state of the art to which the invention pertains. Full citations for these references are provided below. The entirety of each of these references is incorporated herein.

[0177] Anderson, N.L. & Anderson, N.G. The human plasma proteome: history, character, and diagnostic prospects. Mol Cell Proteomics 1 , 845-867 (2002).

[0178] Demichev, V., Messner, C.B., Vernardis, S.I., Lilley, K.S. & Raiser, M. DIA-NN: neural networks and interference correction enable deep proteome coverage in high throughput. Nat Methods 17, 41-44 (2020).

[0179] Do, M., et al. Quantitative proteomic analysis of pancreatic cyst fluid proteins associated with malignancy in intraductal papillary mucinous neoplasms. Clin Proteomics 15, 17 (2018).

[0180] Frankenfield, A.M., Ni, J., Ahmed, M. & Hao, L. Protein Contaminants Matter: Building Universal Protein Contaminant Libraries for DDA and DIA Proteomics. J Proteome Res 21 , 2104-2113 (2022).

[0181] Geyer, P.E., et al. Plasma Proteome Profiling to detect and avoid sample-related biases in biomarker studies. EMBO Mol Med 11 , e10427 (2019).

[0182] Lundberg, M., Eriksson, A., Tran, B., Assarsson, E., & Fredriksson, S.. Homogeneous antibody-based proximity extension assays provide sensitive and specific detection of low-abundant proteins in human blood. Nucleic acids research, 39(15), e102 (2011). O'Brien, D.P., et al. Structural Premise of Selective Deubiquitinase USP30 Inhibition by Small-Molecule Benzosulfonamides. Mol Cell Proteomics 22, 100609 (2023).

[0183] Paniccia, A., etal. Prospective, Multi-Institutional, Real-Time Next-Generation Sequencing of Pancreatic Cyst Fluid Reveals Diverse Genomic Alterations That Improve the Clinical Management of Pancreatic Cysts. Gastroenterology 164, 117-133.e117 (2023).

[0184] Raimondi, S., Maisonneuve, P., & Lowenfels, A. B. Epidemiology of pancreatic cancer: an overview. Nature Reviews. Gastroenterology & hepatology, 6(12), 699-708 (2009).

[0185] Robin, X., et al. pROC: an open-source package for R and S+ to analyze and compare ROC curves. BMC Bioinformatics 12, 77 (2011).

[0186] Singhi, A.D., Koay, E.J., Chari, S.T. & Maitra, A. Early Detection of Pancreatic Cancer: Opportunities and Challenges. Gastroenterology 156, 2024-2040 (2019).

[0187] Soreide, K., Ismail, W., Roalso, M., Ghotbi, J. & Zaharia, C. Early Diagnosis of Pancreatic Cancer: Clinical Premonitions, Timely Precursor Detection and Increased Curative-Intent Surgery. Cancer Control 30, 10732748231154711 (2023).

[0188] Tanaka, M., et al. Revisions of international consensus Fukuoka guidelines for the management of IPMN of the pancreas. Pancreatology 17 , 738-753 (2017).

[0189] Tyanova, S., et al. Proteomic maps of breast cancer subtypes. Nat Commun 7, 10259 (2016).

[0190] Tyanova, S. & Cox, J. Perseus: A Bioinformatics Platform for Integrative Analysis of Proteomics Data in Cancer Research. Methods Mol Biol 1711 , 133-148 (2018). van Huijgevoort, N.C.M., Del Chiaro, M., Wolfgang, C.L., van Hooft, J.E. & Besselink, M.G. Diagnosis and management of pancreatic cystic neoplasms: current evidence and guidelines. Nat Rev Gastroenterol Hepatol 16, 676-689 (2019).

[0191] Wik, L., Nordberg, N., Broberg, J., Bjorkesten, J., Assarsson, E., Henriksson, S., Grundberg, I., Pettersson, E., Westerberg, C., Liljeroth, E., Falck, A., & Lundberg, M. Proximity Extension Assay in Combination with Next-Generation Sequencing for High-throughput Proteome-wide Analysis. Molecular & cellular proteomics : MCP, 20, 100168 (2021).

[0192] Wojtkiewicz, M., Berg Luecke, L., Kelly, M.l. & Gundry, R.L. Facile Preparation of Peptides for Mass Spectrometry Analysis in Bottom-Up Proteomics Workflows. Curr Protoc 1 , e85 (2021).

[0193] For standard molecular biology techniques, see Sambrook, J., Russel, D.W. Molecular Cloning, A

[0194] Laboratory Manual. 3 ed. 2001 , Cold Spring Harbor, New York: Cold Spring Harbor Laboratory Press

Claims

Claims:1 . A computer-implemented method for determining if a pancreatic cyst of a subject is pre- malignant / malignant or benign, the method comprising: a. Providing protein abundance data obtained from a liquid sample from the subject (“sample data”); b. Providing the sample data as input to a machine learning model trained on a training data set comprising a protein abundance profile of at least three pancreatic carcinomas and at least three presumed benign pancreatic cysts, the protein abundance profile comprising the abundance level of at least 15 proteins selected from Table 1 (“training data”), wherein each pancreatic carcinoma is a distinct pancreatic carcinoma subtype and each presumed benign pancreatic cyst is a distinct presumed benign pancreatic cyst subtype; c. Receiving an output from the machine learning model, the output indicative of a pre- malignant / malignant classification or a benign pancreatic cyst classification for the sample data; and d. Determining if the pancreatic cyst of the subject is pre-malignant / malignant or benign based on the output received in step c).

2. The computer-implemented method of claim 1 , wherein the pancreatic cyst of the subject is a nonresected pancreatic cyst.

3. The computer-implemented method of claim 1 or claim 2, wherein the at least 15 proteins comprise protein XRP2, Thiopurine S-methyltransferase, resistin and at least 12 further proteins selected from Table 1 .

4. The computer-implemented method of any one of claims 1 to 3, wherein the at least 15 proteins comprise the following 19 proteins of Table 1 : Phospholipase A2, Pancreatic alpha-amylase, Alphaamylase 2B, Protein XRP2, Receptor-type tyrosine-protein phosphatase C, Eosinophil cationic protein, Catechol O-methyltransferase, Protein S100-P, Cytidine deaminase, Glutaredoxin-1 , Myeloid cell nuclear differentiation antigen, Thiopurine S-methyltransferase, Protein S100-A12, Apoptosis regulator BAX, Cadherin-17, Deaminated glutathione amidase, Resistin, Omega-amidase NIT2 and Septin-9.

5. The computer-implemented method of any one of the preceding claims, wherein the at least 15 proteins comprise at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80 or at least 85 proteins selected from Table 1 .

6. The computer-implemented method of any one of the preceding claims, wherein the protein abundance profile comprises the abundance level of the 89 proteins of Table 1.

7. The computer-implemented method of any one of claims 1 to 6, wherein the protein abundance profile comprises the abundance level of at least a further 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150 or 250 proteins selected from Table 2.

8. The computer-implemented method of any one of claims 1 to 4, wherein the protein abundance profile comprises the abundance level of the 96 proteins of Table 6.

9. The computer-implemented method of any one of the preceding claims, wherein the protein abundance profile comprises the abundance level of the 377 proteins of Table 3.

10. The computer-implemented method of any one of the preceding claims, wherein the at least three pancreatic carcinoma subtypes comprise a colloid carcinoma (CC) subtype, a pancreatic ductal adenocarcinoma (PDAC) subtype and / or a basaloid squamous cell carcinoma (BSCC) subtype.

11. The computer-implemented method of any one of the preceding claims, wherein the liquid sample comprises a pancreatic cyst fluid (PCyF), blood or serum sample, optionally a pancreatic cyst fluid (PCyF) sample.

12. The computer-implemented method of any one of the preceding claims, wherein the machine learning model comprises a support vector machine (SVM), a linear regression model, a decision tree, a logistic regression model, an artificial neural network, naive bayes or k-nearest neighbour algorithm, optionally wherein the decision tree comprises a gradient boosting algorithm or a random forest algorithm.

13. The computer-implemented method of any one of the preceding claims, wherein the machine learning model comprises a linear regression model.

14. The computer-implemented method of any one of the preceding claims, wherein the machine learning model comprises a support vector machine (SVM) algorithm.

15. The computer-implemented method of any one of the preceding claims, wherein the output comprises a malignancy score of between 0 and 1 .

16. The computer-implemented method of claim 15, wherein in d), the pancreatic cyst of the subject is determined as pre-malignant / malignant when the output comprises a malignancy score of > 0.7 and the pancreatic cyst of the subject is determined as benign when the output comprises a malignancy score of < 0.7.

17. The computer-implemented method of any one of the preceding claims, wherein the method is for early detection of pre-malignancy / malignancy of the pancreas and / or monitoring progression of pancreatic cyst(s) from benign to pre-malignant / malignant.

18. The computer-implemented method of any one of the preceding claims, wherein when the pancreatic cyst of the subject is determined to be benign in step d), the computer-implemented method may be repeated using one or more further liquid sample(s) obtained from the subject at later time point(s) to determine if the pancreatic cyst has developed into a premalignant / malignant pancreatic cyst.

19. The computer-implemented method of claim 18, wherein the later time point(s) is at least three months, at least six months, at least 12 months, at least two years, at least three years or at least five years after the first liquid sample was obtained from the subject.

20. A method of determining if a pancreatic cyst of a subject is premalignant / malignant or benign, the method comprising: a. Obtaining protein abundance data from a liquid sample from the subject; and b. Performing the method of any one of claims 1 to 19 to determine if the pancreatic cyst is premalignant / malignant or benign.

21. The method of claim 20, wherein the protein abundance data is obtained by mass spectrometry, aptamer assay-based proteomics or protein extension assay-based proteomics.

22. A method for selecting a treatment for a subject having a pancreatic cyst, the method comprising: a. Performing the method of any one of claims 1 to 21 to determine if the pancreatic cyst of the subject is premalignant / malignant or benign; and b. Selecting an anti-cancer treatment for the subject if the pancreatic cyst is determined to be premalignant / malignant.

23. A method of treatment of a subject having a malignant pancreatic cyst, the method comprising: a. Performing the method of any one of claims 1 to 22, wherein the pancreatic cyst of the subject is determined to be premalignant / malignant; and b. Administering an anti-cancer treatment to the subject.

24. The method of claim 22 or claim 23, wherein the anti-cancer treatment comprises surgery, chemotherapy and / or radiotherapy, preferably chemotherapy and / or radiotherapy.

25. The method of claim 24, wherein the chemotherapy comprises gemcitabine, capecitabine, fluorouracil, irinotecan, oxaliplatin, nab-paclitaxel and / or cisplatin.

26. A system comprising: a processor; and a computer readable medium comprising instructions that, when executed by the processor, cause the processor to perform the steps of the method of any one of claims 1 to 19.

27. One or more computer readable media comprising instructions that, when executed by one or more processors, cause the one or more processors to performs the steps of the method of any one of claims 1 to 19.

28. A method of measuring protein abundance in a pancreatic cyst fluid (PCyF) sample, wherein the method comprises measuring the abundance level of at least 15 proteins selected from Table 1 in the PCyF sample to obtain a measured abundance level.

29. The method of claim 28, wherein the method further comprises comparing the measured abundance level to a stored abundance level.

30. The method of claim 28 or claim 29, wherein the at least 15 proteins comprise no more than 4000 proteins, no more than 3500 proteins, no more than 3000 proteins, no more than 2500 proteins, no more than 2000 proteins, no more than 1500 proteins, no more than 1000 proteins or no more than 500 proteins in total.