Pancreas cancer assessment methods and devices

WO2026162808A1PCT designated stage Publication Date: 2026-08-06LUDWIG-MAXIMILIANS-UNIVERSITÄT MÜNCHEN IN VERTRETUNG DES FREISTAATES BAYERN
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
LUDWIG-MAXIMILIANS-UNIVERSITÄT MÜNCHEN IN VERTRETUNG DES FREISTAATES BAYERN
Filing Date
2026-02-02
Publication Date
2026-08-06

Smart Images

  • Figure IMGF000065_0001
    Figure IMGF000065_0001
  • Figure IMGF000065_0002
    Figure IMGF000065_0002
  • Figure IMGF000005_0001_TABLE
    Figure IMGF000005_0001_TABLE
Patent Text Reader

Abstract

The present invention relates to a method for assessing pancreatic cancer in a subject, said method comprising (a) determining the biomarkers Ceramide (d18:1;C24:0), Lysophosphatidylethanolamine (C18:0), Phosphatidylethanolamine (C18:0, C22:6), Sphingomyelin (d17:1; C16:0), and CA19.9 in a sample from said subject; (b) comparing the biomarkers determined in step (a) to at least one corresponding reference; and (c) assessing pancreatic cancer in said subject based on said comparing in step (b). The present invention further relates to a computer-implemented training method of training at least one trainable model for assessing pancreatic cancer, to an automated machine learning model, and to a system related thereto.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Ludwig-Maximilians-Universitat Munchen,

[0002] in Vertretung des Freistaates Bayern 1 LMU17275PC

[0003] Pancreas cancer assessment methods and devices

[0004] The present invention relates to a method for assessing pancreatic cancer in a subject, said method comprising (a) determining the biomarkers Ceramide (dl8:l;C24:0), Lysophosphatidylethanolamine (Cl 8:0), Phosphatidylethanolamine (Cl 8:0, C22:6), Sphingomyelin (dl7: 1; C16:0), and CA19.9 in a sample from said subject; (b) comparing the biomarkers determined in step (a) to at least one corresponding reference; and (c) assessing pancreatic cancer in said subject based on said comparing in step (b). The present invention further relates to a computer-implemented training method of training at least one trainable model for assessing pancreatic cancer, to an automated machine learning model, and to a system related thereto.

[0005] Pancreatic ductal adenocarcinoma (PDAC) is currently the third - and projected to be the second - leading cause of cancer deaths. In the US, its incidence rises by 1.1% annually and it accounts for 18.7% of all new gastrointestinal cancer cases. Five-year survival of all stages has been reported to be 13% and mortality has marginal improvement in recent decades. Only 15% of PDAC cases are diagnosed as localized cancers, but these patients have a 5 -year survival rate of 44%, highlighting that early diagnosis is key to improved survival. Balancing prevalence and cumulative risk of PDAC is indicative that a cohort of patients with an incidence increased over the average population, e.g. between 0.5 and 1%, such as chronic pancreatitis (CP, PDAC incidence: 0.49-1.04% (Kim et al., Sci Rep 2023; 13(1): 106)), new onset diabetes (NOD, PDAC incidence: 0.5-1% (Singhi et al., Gastroenterology 2019;156(7):2024-40)) and familial pancreatic cancer (3 or more first-degree relatives first-degree relatives with PDAC: 16%-40%, 2 first-degree relatives with PDAC: up to 12%, 1 first-degree relative with PDAC: up to 6% cumulative risk (Llach et al., Cancer Manag Res 2020;12:743-58)) populations might profit from surveillance. These arguments call for the development of a non-invasive test to diagnose PDAC early in at-risk cohorts.

[0006] Plasma metabolic signatures predicting PDAC were proposed earlier (Mayerle et al., Gut 2018;67(l): 128-37; Pepe et al., J Natl Cancer Inst 2008; 100(20): 1432-8; Mahajan et al., Gastroenterology 2022; 163(5): 1407-22). In particular Mahajan et al. (2022) reported highLudwig-Maximilians-Universitat Miinchen,

[0007] in Vertretung des Freistaates Bayern 2 LMU17275PC AUC values for the proposed diagnostic method, however required a pre-allocation of patients according to CA19.9 expression. Further known methods for diagnosing pancreatic cancer are summarized in Table 1.

[0008] Table 1 : Test methods for diagnosing pancreatic cancer

[0009]

[0010] Nonetheless, there is still a need for improved biomarkers and algorithms to reliably diagnose pancreatic cancer to improve survival. This problem is solved by the means and methods of the present invention, with the features of the independent claims. Preferred embodiments, which might be realized in an isolated fashion or in any arbitrary combination are listed in the dependent claims.

[0011] In accordance, the present invention relates to method for assessing pancreatic cancer in a subject, said method comprising

[0012] (a) determining the biomarkers Ceramide (dl8:l;C24:0), Lysophosphatidylethanolamine (Cl 8:0), Phosphatidylethanolamine (Cl 8:0, C22:6), Sphingomyelin (dl7:l; Cl 6:0), and CA19.9 in a sample from said subject;Ludwig-Maximilians-Universitat Munchen,

[0013] in Vertretung des Freistaates Bayern 3 LMU17275PC (b) comparing the biomarkers determined in step (a) to at least one corresponding reference; and

[0014] (c) assessing pancreatic cancer in said subject based on said comparing in step (b).

[0015] In general, terms used herein are to be given their ordinary and customary meaning to a person of ordinary skill in the art and, unless indicated otherwise, are not to be limited to a special or customized meaning. As used in the following, the terms “have”, “comprise” or “include” or any arbitrary grammatical variations thereof are used in a non-exclusive way. Thus, these terms may both refer to a situation in which, besides the feature introduced by these terms, no further features are present in the entity described in this context and to a situation in which one or more further features are present. As an example, the expressions “A has B”, “A comprises B” and “A includes B” may both refer to a situation in which, besides B, no other element is present in A (i.e. a situation in which A solely and exclusively consists of B) and to a situation in which, besides B, one or more further elements are present in entity A, such as element C, elements C and D or even further elements. Also, as is understood by the skilled person, the expressions "comprising a" and "comprising an" preferably refer to "comprising one or more", i.e. are equivalent to "comprising at least one". In accordance, expressions relating to one item of a plurality, unless otherwise indicated, preferably relate to at least one such item, more preferably a plurality thereof; thus, e.g. identifying "a cell" relates to identifying at least one cell, preferably to identifying a multitude of cells.

[0016] Further, as used in the following, the terms "preferably", "more preferably", "most preferably", "particularly", "more particularly", "specifically", "more specifically" or similar terms are used in conjunction with optional features, without restricting further possibilities. Thus, features introduced by these terms are optional features and are not intended to restrict the scope of the claims in any way. The invention may, as the skilled person will recognize, be performed by using alternative features. Similarly, features introduced by "in an embodiment" or similar expressions are intended to be optional features, without any restriction regarding further embodiments of the invention, without any restrictions regarding the scope of the invention and without any restriction regarding the possibility of combining the features introduced in such way with other optional or non-optional features of the invention.

[0017] The methods specified herein below, preferably, are in vitro methods. The method steps may, in principle, be performed in any arbitrary sequence deemed suitable by the skilled person, butLudwig-Maximilians-Universitat Munchen,

[0018] in Vertretung des Freistaates Bayern 4 LMU17275PC preferably are performed in the indicated sequence; also, one or more, preferably all, of said steps may be assisted or performed by automated equipment. Moreover, the methods may comprise steps in addition to those explicitly mentioned above. Also, the methods may be comprised in other methods; e.g. the method for assessing pancreatic cancer in a subject may be comprised in a method of monitoring a subject at risk of suffering from pancreatic cancer. Also, the methods are preferably at least partially, more preferably fully, computer-implemented, preferably as described in more detail herein below.

[0019] As used herein, if not otherwise indicated, the term "about" relates to the indicated value with the commonly accepted technical precision in the relevant field, preferably relates to the indicated value ± 20%, more preferably ± 10%, most preferably ± 5%. Further, the term "essentially" indicates that deviations having influence on the indicated result or use are absent, i.e. potential deviations do not cause the indicated result to deviate by more than ± 20%, more preferably ± 10%, most preferably ± 5%. Thus, “consisting essentially of’ means including the components specified but excluding other components except for materials present as impurities, unavoidable materials present as a result of processes used to provide the components, and components added for a purpose other than achieving the technical effect of the invention. For example, a composition defined using the phrase “consisting essentially of’ encompasses any known acceptable additive, excipient, diluent, carrier, and the like. Preferably, a composition consisting essentially of a set of components will comprise less than 5% by weight, more preferably less than 3% by weight, even more preferably less than 1% by weight, most preferably less than 0.1% by weight of non-specified component(s). As referred to herein, parameters and coefficients are preferably used with a number of significant digits deemed appropriate by the skilled person, and not necessarily the number of significant digits provided herein. Thus, e.g. a coefficient of 1.968635 provided for CA19.9 herein below, may be used as a coefficient 2, as 2.0, as 1.97, 1.969, and the like, as deemed appropriate by the skilled person. E.g., in case a %probability of a diagnosis is calculated, it may be sufficient to use values of coefficients only to the 2ndsignificant digit, e.g. a value of 2.0 for CA19.9 in the aforesaid example.

[0020] The term "fragment" of a biological or chemical molecule, e.g. of a biomarker as specified herein, is used herein in a wide sense relating to any sub-part of the respective molecule comprising the indicated sequence, structure and / or function. Thus, the term includes sub-parts generated by actual fragmentation of a molecule, but also sub-parts derived from the respectiveLudwig-Maximilians-Universitat Munchen,

[0021] in Vertretung des Freistaates Bayern 5 LMU17275PC molecule in an abstract manner, e.g. in silico. Thus, as used herein, e.g. an ion generated from a biological or chemical molecule, e.g. before or during a mass spectrometry (MS) measurement, may be referred to as a fragment of the parental molecule. Unless specifically indicated otherwise herein, the compounds specified, e.g. may be covalently or non-covalently linked to further atoms, molecules, and / or molecule complexes.

[0022] The term “assessing”, as used herein, refers to establishing information about the status of the indicated disease or condition, in particular its severity, symptoms, localization, prognosis and / or other relevant information. Said assessing preferably is an aid in diagnosing the indicated disease; as the skilled person will understand, establishing a diagnosis may be based on the aforesaid assessment, however, preferably is based on the aforesaid assessment in combination with further diagnostic information, such as anamnesis data, general physical, mental examination findings, and / or additional metabolic data. Thus, assessing pancreatic cancer may relate to assessing whether a subject suffers from pancreatic cancer, is at risk of suffering from pancreatic cancer, exhibits a medical condition which deteriorates with respect to pancreatic cancer, to establishing the disease stage of pancreatic cancer of a subject, and / or establishing a prognosis with respect to pancreatic cancer. Accordingly, assessing as used herein includes diagnosing pancreatic cancer, predicting the risk for developing pancreatic cancer, and / or predicting any deterioration of the health condition of the subject, in particular, with respect to signs and symptoms accompanying pancreatic cancer. Assessment referred to herein may also be the assessment of a risk of developing pancreatic cancer. In a further embodiment, assessment may be the prediction of the risk that the subject’s (health) condition of the subject will deteriorate. Moreover, it will be understood that if the risk of developing pancreatic cancer or risk of the deterioration of the health condition is predicted, typically, the prediction is made within a predictive window. More typically, said predictive window is of from 1 day to 5 years, in a further embodiment of from one week to 2 years.

[0023] As referred to herein, assessing pancreatic cancer preferably may also refer to establishing information about differentiating pancreatic cancer from a non-pancreatic cancer disorder as specified herein below, i.e. said assessing preferably is an aid in differentially diagnosing the indicated disease; as the skilled person will understand, establishing a differential diagnosis may be based on the aforesaid assessment, however, preferably is based on the aforesaid assessment in combination with further diagnostic information, preferably as specified herein above. Thus, assessing pancreatic cancer may relate to assessing whether a subject suffers fromLudwig-Maximilians-Universitat Munchen,

[0024] in Vertretung des Freistaates Bayern 6 LMU17275PC pancreatic cancer or from a non-pancreatic cancer disorder. Accordingly, assessing as used herein includes classifying pancreatic cancer from benign pancreatic disease as specified herein below. Preferably, in particular in case of differentiating pancreatic cancer from non-pancreatic cancer disorders, the subject was diagnosed to show at least one symptom indicative of pancreatic cancer and / or of a non-pancreatic cancer disorder, preferably as specified herein below. Thus, the method of assessing pancreatic cancer may in particular be applied to a sample of a subject from a pre-selected subgroup of subjects requiring or profiting from differential diagnosis, preferably as specified herein below.

[0025] Assessing as referred to herein may relate to a rule-in assessment, i.e. to identifying a subject as belonging to a group of subjects sharing a common feature, e.g. suffering from pancreatic cancer. Assessing, however, may also relate to a rule-out assessment, i.e. to identifying a subject as not belonging to a group of subjects sharing a common feature, e.g. as not suffering from pancreatic cancer. Thus, the assessment may aid in establishing a diagnosis; the assessment may, however, also aid in excluding a diagnosis. Thus, the method as specified may in particular be comprised in a method of monitoring pancreatic cancer. The above applies mutatis mutandis to differential diagnosis, i.e. the differential diagnosis may be a rule-in assessment, i.e. to identifying a subject as suffering from pancreatic cancer or as suffering from a non-pancreatic cancer disorder; it may, however, may also relate to a rule-out assessment, i.e. to identifying a subject as not suffering from pancreatic cancer or as not suffering from a non-pancreatic cancer disorder. Thus, the assessment may aid in establishing a differential diagnosis. As the skilled person is aware of, e.g. excluding pancreatic cancer in a subject showing symptoms thereof may be of help to decide on the further course of treatment.

[0026] In view of the above, assessing pancreatic cancer in particular may include or be aiding in diagnosis of pancreatic cancer; in an embodiment to be applied in a specialist or tertiary care setting, in particular with access to a laboratory environment where automated assays can be run; and / or an aid in the diagnosis and assessment of the severity of pancreatic cancer in subjects with signs and symptoms of pancreatic cancer, in an embodiment in conjunction with other laboratory findings and clinical assessments. Preferably, assessing comprises excluding pancreatic cancer, preferably comprises identifying a subject not suffering from pancreatic cancer; comprises diagnosing pancreatic cancer, preferably comprises diagnosing PDAC; and / or comprises staging pancreatic cancer. More preferably, assessing comprises differentiating between pancreatic cancer and non-pancreatic cancer disorders as specifiedLudwig-Maximilians-Universitat Munchen,

[0027] in Vertretung des Freistaates Bayern 7 LMU17275PC elsewhere herein, preferably pancreatitis, more preferably chronic pancreatitis, Intraductal Papillary Mucinous Neoplasia (IPMNs) and pancreatic cystic lesions (CPL). Even more preferably assessing is classifying pancreas cancer. Nonetheless, in an embodiment, the non-pancreatic cancer disorders do not comprise chronic pancreatitis; thus, in an embodiment, assessing does not comprise differentiating between pancreatic cancer and pancreatitis, preferably chronic pancreatitis.

[0028] As will be understood by those skilled in the art, the assessment made in accordance with the present invention, although usually preferred to be, may not be correct for 100% of the investigated subjects. However, the term typically requires that a statistically significant portion of subjects can be correctly assessed. Whether a portion is statistically significant can be determined without further ado by the person skilled in the art using various well known statistic evaluation tools, e.g., determination of confidence intervals, p-value determination, Student's t-test, Mann- Whitney test, etc. Details may be found in Dowdy and Wearden, Statistics for Research, John Wiley & Sons, New York 1983. Typically envisaged confidence intervals are at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95%. The p-values are, typically, 0.1, 0.05, or 0.01.

[0029] The terms “pancreatic cancer” and “pancreas cancer”, as used herein, relate equally to malignant neoplasms which are derived from pancreas cells and, preferably, from pancreatic epithelial cells. Thus, preferably, pancreatic cancer as used herein is pancreatic adenocarcinoma, more preferably pancreatic ductal adenocarcinoma (PDAC). Preferably, the pancreatic cancer is a resectable pancreatic cancer, i.e., preferably, is a pancreatic cancer at a tumor stage permitting, preferably complete, resection of the tumor from the subject. More preferably, said pancreatic cancer is a pancreatic cancer of tumor stage IA-IIB. The symptoms accompanying pancreatic cancer are well known from standard text books of medicine and may include severe abdominal pain, lower back pain, weight loss, new onset diabetes, and jaundice.

[0030] The term "non-pancreatic cancer disorder" as used herein, relates to any non-malignant disease of the pancreas. Preferably, the non-pancreatic cancer disorder has at least one symptom which is also a symptom of pancreatic cancer, e.g. presence of at least one pancreatic lesion in an affected subjectjaundice, HbAlc > 6.5%, weight loss in last 6 months, back pain in last weeks, exocrine insufficiency, and / or severe abdominal pain. Preferably, the non-pancreatic cancer disorder has formation of at least one pancreatic cystic lesion as a symptom, wherein said cysticLudwig-Maximilians-Universitat Munchen,

[0031] in Vertretung des Freistaates Bayern 8 LMU17275PC lesion preferably is detectable by an imaging technique such as radiography, computer tomography, magnetic resonance tomography, sonography, or the like. Thus, preferably, a non-pancreatic cancer disorder is pancreatitis, in particular chronic pancreatitis, intraductal papillary mucinous neoplasm (IPMN), cystic pancreatic lesions (CPL), and pancreatic metastases of extrapancreatic origin.

[0032] The term “subject”, as used herein, relates to an animal, preferably to a mammal. More preferably, the subject is a primate and, most preferably, a human. Preferably, the subject is an apparently healthy subject. Preferably, the subject is a subject at risk of suffering from pancreatic cancer. Risk factors for developing pancreatic cancer are known in the art, e.g. from Brand et al., Gut. 2007;56:1460-9; or Del Chiaro et al., World J Gastroenterol 2014; 20:12118-12131 and include new-onset diabetes, genetic factors, chronic disease, and age; more preferably, the risk factor is new-onset diabetes. Thus, preferably, the subject is at risk of suffering from pancreatic cancer, i.e., preferably, the subject is a subject from a population with a prevalence of pancreatic cancer of at least 0.5%, preferably at least 1%, more preferably at least 5%, even more preferably at least 10%, most preferably at least 20%. Corresponding populations are known in the art and include in particular populations having a genetic predisposition, preferably familiar pancreatic cancer, including Peutz-Jeghers Syndrome, BRCA1 positivity, or a genetic predisposition for developing pancreatitis; thus, the subject may be from a familial PDAC kinship. Also preferably, the subject at risk of suffering from pancreatic cancer is a subject at least 40 years old, more preferably, at least 50 years old. More preferably, the subject at risk of suffering from pancreatic cancer is a subject suffering from chronic pancreatitis and / or from pancreatic lesions, wherein said pancreatic lesions preferably have been confirmed by an imaging methods such as CT or MRT; thus, the subject preferably has been identified, e.g. by computer tomography (CT), to have pancreatic lesions necessitating further diagnostic assessment. Most preferably, the subject at risk of suffering from pancreatic cancer is a subject with new-onset diabetes, preferably new-onset diabetes type II. Preferably, the subject is a subject in whom new-onset diabetes was diagnosed at most three years, more preferably at most two years, most preferably at most one year, before the method as specified herein is performed on a sample of said subject. Thus, the method preferably is a method for diagnosing pancreatic cancer in a subject suffering from new-onset diabetes.

[0033] Also preferably, the subject is a subject in need of differential diagnosis of pancreatic cancer. Thus, the subject preferably is a subject showing at least one symptom of pancreatic cancerLudwig-Maximilians-Universitat Munchen,

[0034] in Vertretung des Freistaates Bayern 9 LMU17275PC and / or of a non-pancreatic cancer disorder, preferably as specified herein above. Preferably, the subject is known or suspected to be suffering from pancreatic cancer and / or is known or suspected to suffer from a non-pancreatic cancer disorder. The skilled person selects the appropriate assessment for a particular subject dependent on the subject's symptoms and existing diagnoses, including tentative diagnoses.

[0035] The term “sample”, as used herein, refers to a biological sample from a body fluid, preferably blood, plasma, serum, saliva or urine, or a sample derived from cells, tissues or organs, in particular from the pancreas, e.g., by biopsy. Preferably, the sample is a blood, plasma or serum sample, more preferably a serum or plasma sample. Even more preferably, the sample is a blood or plasma sample or is a serum or plasma sample, most preferably, is a plasma sample. Preferably, the sample is a citrate plasma sample, a heparin plasma sample, or an EDTA plasma sample.

[0036] Biological samples can be derived from a subject by techniques known in the art. For example, blood samples may be obtained by blood taking, while tissue or organ samples are to be obtained, e.g., by biopsy. In an embodiment, the sample is known or suspected to comprise biomarkers referred to herein. The aforementioned samples may be pre-treated before they are used for according to the present invention. Said pre-treatment may include treatments required to release or separate the biomarker(s) and / or the analyte(s) or to remove excessive material or waste. Suitable techniques comprise centrifugation, extraction, fractioning, ultrafiltration, protein precipitation followed by filtration and purification and / or enrichment of compounds. Moreover, other pre-treatments may be carried out in order to provide the biomarker and / or analyte in a form or concentration suitable for the intended determination. Suitable and necessary pre-treatments depend on the means used for carrying out the method of the invention and are well known to the person skilled in the art. Pre-treated samples as described before are also comprised by the term “sample” as used in accordance with the present invention.

[0037] Preferably, the sample is a fasting sample, in particular a fasting blood, plasma or serum sample. Thus, preferably, the sample is obtained from a fasting subject. A fasting subject, in particular, is a subject who refrained from food and beverages, except for water, prior to obtaining the sample to be tested. Preferably, a fasting subject refrained from food and beverages, except for water, for at least eight hours prior to obtaining the sample to be tested. More preferably, the sample has been obtained from the subject after an overnight fast. Preferably said fastingLudwig-Maximilians-Universitat Munchen,

[0038] in Vertretung des Freistaates Bayern 10 LMU17275PC continued up to at least one hour before sample taking, more preferably up to at least 30 min before sample taking, still more preferable up to at least 15 min before sample taking, most preferably until the sample was taken. Preferably the sample is known or suspected to comprise biomarkers referred to herein. Also preferably, the concentration of CAI 9.9 in the sample is known or suspected to be less than 37 U / ml, preferably less than 10 U / ml.

[0039] The term “biomarker”, as used herein, refers to a molecular species which serves as an indicator for a disease or physiological state as referred to herein. Said molecular species can be a chemical compound which is detectable in a sample of a subject, in particular a metabolite of the subject's metabolism. Moreover, the biomarker may also be a molecular species which is derived from said metabolite. In such a case, the actual metabolite will be chemically modified in the sample or during the determination process and, as a result of said modification, a chemically different molecular species, i.e. the analyte, will be the determined molecular species. Also, in case the biomarker has an activity, e.g. a catalytic activity and / or an activating activity, e.g. on target cells, the biomarker may also be determined via said activity, e.g. in an enzymatic assay. It is to be understood that in the aforesaid cases, the analyte may represent the actual biomarker and has the same potential as an indicator for the respective medical condition as the biomarker would have. Preferred modes of determination and analytes for the biomarkers of the present description are described in the context of the respective biomarkers herein below. Moreover, as is understood by the skilled person, a biomarker according to the present invention need not necessarily correspond to one molecular species. Rather, the biomarker may comprise stereoisomers or enantiomers of a compound and / or, e.g. in case the biomarker is a polypeptide, may comprise variant molecular species, e.g. translated from splice variants, glycosylation variants, peptidase processing variants, and the like. In an embodiment, the variants share at least one determinable feature, e.g. an epitope or an activity.

[0040] The biomarkers referred to herein and methods for their determination are in principle known in the art: Ceramide (dl8:l;C24:0): CAS NO: 102917-80-6, Lysophosphatidylethanolamine (C18:0): CAS NO: 899443-67-5, Phosphatidylethanolamine (C18:0, C22:6): CAS NO: 202647-82-3, Sphingomyelin (dl7:l; C16:0): CAS NO: 123065-40-7, and CA19.9: CAS NO: 92448-22-1. Further preferred biomarkers are Histidine (CAS NO: 71-00-1), Proline (CAS NO: 147-85-3), Tryptophan (CAS NO: 73-22-3), Ceramide (dl8:2,C24:0): CAS NO: 135941-18-3, Lysophosphatidylethanolamine (C18:2): CAS NO: 85046-18-0, Sphingomyelin (35:1): CHEBI: 133629, Sphingomyelin (41:2): CHEBL85762, and Sphingomyelin (dl8:2,C17:0):Ludwig-Maximilians-Universitat Munchen,

[0041] in Vertretung des Freistaates Bayern 11 LMU17275PC SpectraBase Compound ID: D4yNgl5DQf3. Preferably, at least one additional biomarker is determined in addition to the aforesaid biomarkers. Preferably, said additional biomarker is selected from the list consisting of aspartate aminotransferase (EC 2.6.1.1), alanine aminotransferase (EC 2.6.1.2), platelet count, haptoglobin (e.g. Genbank Acc No. NP_001119574.1), alpha2-macroglobulin (e.g. Genbank Acc No. NP_000005.3), apolipoprotein Al (e.g. Genbank Acc No. NP_000030.1), bilirubin (CAS NO: 635-65-4), cholesterol (CAS NO: 57-88-5), hyaluronan (CAS NO: 9004-61-9), prothrombin index, hepatocyte growth factor (HGF, e.g. Genbank Acc No. NP_000592.3), and urea (CAS NO: 57-13-6).

[0042] The term “determining” as used herein refers to semi quantitative or quantitative determination of a biomarker referred to herein. Determining the amount of a biomarker may be carried out by any technique which allows for establishing a measure of quantity of a biomarker in a semi quantitative or quantitative manner. Suitable techniques depend on the molecular nature and the properties of the biomarkers and are discussed elsewhere herein in more detail.

[0043] In principle, the amount of a biomarker can be determined by determining a complex of the analyte with a detection compound, in particular an antibody or fragment thereof, i.e. in an immunoassay. Said determining of a complex of the analyte may be performed in any format deemed appropriate by the skilled person, in particular a sandwich, competition, or other assay format. Said assays will develop a signal which is indicative for the amount of a biomarker. The amount of a biomarker may in an embodiment be determined in an activity assay, in particular in case the biomarker has catalytic, e.g. enzymatic, or signaling activity.

[0044] Preferably, the amount of a biomarker may be determined by detecting the amount of molecular species of the biomarker, or of fragments thereof. E.g., small molecule biomarkers may be detected as such or as their ions in mass spectrometry (MS). For polypeptide biomarkers detection of fragments thereof may be technically easier to put into practice. However, also other methods for detecting the amount of molecular species of the analyte are available, including chromatographic separation techniques such as liquid chromatography (LC), high performance liquid chromatography (HPLC), gas chromatography (GC), thin layer chromatography, and / or size exclusion or affinity chromatography, coupled to appropriate detection devices. Such a detection device may e.g. be a photometer, e.g. an UV / VIS-photometer or an MS device. Appropriate devices and methods are known in the art. FurtherLudwig-Maximilians-Universitat Munchen,

[0045] in Vertretung des Freistaates Bayern 12 LMU17275PC suitable methods comprise measuring a physical or chemical property specific for the biomarker such as its precise molecular mass or an NMR spectrum. Said methods comprise, preferably, biosensors, optical devices coupled to immunoassays, biochips, analytical devices such as mass- spectrometers, NMR-analyzers, surface plasmon resonance measurement equipment or chromatography devices.

[0046] The biomarkers to be determined in accordance with the present invention are as such known in the art. Moreover, methods for the determination of the amount of the biomarkers are known to the skilled person as well. For example, the biomarkers can be measured as described in the Examples section.

[0047] More preferably, determining of at least one biomarker comprises mass spectrometry (MS). For mass spectrometry, the analytes in the sample are ionized in order to generate charged molecules or molecule fragments. Afterwards, the mass-to-charge of the ionized analyte, in particular of the ionized biomarkers, or fragments thereof is measured. Thus, the mass spectrometry step preferably comprises an ionization step in which the biomarkers to be determined are ionized. Of course, other compounds present in the sample / eluate are ionized as well. Ionization of the biomarkers can be carried out by any method deemed appropriate, in particular by electron impact ionization, fast atom bombardment, electrospray ionization (ESI), atmospheric pressure chemical ionization (APCI), matrix assisted laser desorption ionization (MALDI). More preferably, the ionization step (for mass spectrometry) is carried out by electrospray ionization (ESI). Accordingly, the mass spectrometry is preferably ESI-MS (or if tandem MS is carried out, ESI-MS / MS).

[0048] Mass spectrometry as used herein encompasses all techniques which allow for the determination of the molecular weight (i.e. the mass) or a mass variable corresponding to a compound, i.e. a biomarker, to be determined in accordance with the methods proposed herein. Preferably, a combination of mass spectrometry with a separation and / or enrichment method is used, in particular gas chromatography mass spectrometry (GC-MS), liquid chromatography mass spectrometry (LC-MS), direct infusion mass spectrometry or Fourier transform ion-cyclotrone-resonance mass spectrometry (FT-ICR-MS), capillary electrophoresis mass spectrometry (CE-MS), high-performance liquid chromatography coupled mass spectrometry (HPLC-MS), quadrupole mass spectrometry, any sequentially coupled mass spectrometry, such as MS-MS or MS-MS-MS, inductively coupled plasma mass spectrometry (ICP-MS), pyrolysisLudwig-Maximilians-Universitat Munchen,

[0049] in Vertretung des Freistaates Bayern 13 LMU17275PC mass spectrometry (Py-MS), ion mobility mass spectrometry or time of flight mass spectrometry (TOF). How to apply these techniques is well known to the person skilled in the art. Moreover, suitable devices are commercially available. More preferably, mass spectrometry as used herein relates to LC-MS and / or GC-MS, i.e. to mass spectrometry being operatively linked to a prior chromatographic separation step. More preferably, mass spectrometry as used herein encompasses quadrupole MS.

[0050] The term “reference”, as used herein, relates to a value, e.g. an amount or any value derived therefrom, e.g. a score, which can be correlated to a medical condition and, preferably, which allows for the assessment as referred to herein to be made, more preferably enables allocation of a subject into either a group of subjects suffering from a disease or condition or being at risk for developing it, or a group of subjects which do not suffer from said disease or condition or which are not at risk for developing it. Such a reference can be a threshold value, e.g. a threshold amount, which separates these groups from each other. Accordingly, the reference may be a value which allows for allocation of a subject into a group of subjects suffering from a disease or condition or being at risk for developing it, or not. For example, the reference may be a value which allows for allocation of a subject into a group of subjects suffering from pancreatic cancer, or being at risk of developing pancreatic cancer. The reference may, however, also be a reference range, e.g., in an embodiment, a range of values for which pancreatic cancer can be excluded. Furthermore, the reference may be a value calculated from the aforesaid values, e.g. from the amounts of two or more biomarkers, in an embodiment to provide a score, e.g. a predictor score as specified elsewhere herein. A suitable reference separating the two groups can be provided without further ado e.g. by the statistical tests referred to herein elsewhere based on values of biomarkers from suitable reference groups as specified herein below. As the skilled person understands, it may not always be possible, although particularly envisaged, to provide a reference unambiguously allocating each and every possible value of a biomarker to one of the aforesaid groups; thus, there may be a range of values for which a clear assessment cannot be provided. Preferably, however, as indicated above, a reference enables the assessment to be made for each and every value of a biomarker or set of biomarkers which may be measured. As the skilled person understands, the specific value of a reference may depend on the assessment intended and on parameters thereof; thus, the reference value for assessing pancreatic cancer may typically be different from the reference value for assessing e.g. chronic pancreatitis. Relevant parameters having an influence on the reference may in particular be sensitivity and specificity of assessment.Ludwig-Maximilians-Universitat Munchen,

[0051] in Vertretung des Freistaates Bayern 14 LMU17275PC

[0052] A reference may in particular be derived from at least one reference group, the term "reference group" relating to a group of subjects with known status with regard to the assessment. Thus the reference group may e.g. be a group of subjects for which it is known whether they suffer from pancreatic cancer. The population of subjects in a reference group preferably comprises a plurality of subjects, e.g. at least 5, 10, 50, 100, 1,000, or 10,000 subjects. Typically, the subject to be diagnosed and the subjects of the said reference group are of the same species. The reference applicable for an individual subject may vary depending on various physiological parameters such as age, gender, or subpopulation. As is understood by the skilled person, prevalence of pancreatic cancer in the population is low; thus, a reference may be derived also from the average population. Assuming that contribution of actually afflicted subjects is low, such an average population reference group may be treated as a reference group known not to suffer from pancreatic cancer; preferably, in such case, the size of the reference group is sufficiently high, e.g. at least 100, more preferably at least 1000, even more preferably at least 10000 subjects. In view of the description herein, the skilled person understands that a reference group may, in principle, also be a mixed population of subjects with regard to pancreatic cancer, provided that the status of each member of said mixed population with regards to pancreatic cancer is or becomes known before deriving a reference from such group.

[0053] Reference amounts can, in principle, be calculated for a cohort of subjects based on the average or mean values for a given parameter such as biomarker amount by applying standard statistically methods. In particular, accuracy of a test such as a method aiming to diagnose an event, or not, is best described by its receiver-operating characteristics (ROC) (see especially Zweig 1993, Clin. Chem. 39:561-577). The ROC graph is a plot of all of the sensitivity / specificity pairs resulting from continuously varying the decision threshold over the entire range of data observed. The clinical performance of a diagnostic method depends on its accuracy, i.e. its ability to correctly allocate subjects to a certain prognosis or diagnosis. The ROC plot indicates the overlap between the two distributions by plotting the sensitivity versus 1 -specificity for the complete range of thresholds suitable for making a distinction. On the y-axis is sensitivity, or the true-positive fraction, which is defined as the ratio of number of truepositive test results to the product of number of true-positive and number of false-negative test results. This has also been referred to as positivity in the presence of a disease or condition. It is calculated solely from the affected subgroup. On the x-axis is the false-positive fraction, or 1 -specificity, which is defined as the ratio of number of false-positive results to the product ofLudwig-Maximilians-Universitat Munchen,

[0054] in Vertretung des Freistaates Bayern 15 LMU17275PC number of true-negative and number of false-positive results. It is an index of specificity and is calculated entirely from the unaffected subgroup. Because the true- and false-positive fractions are calculated entirely separately, by using the test results from two different subgroups, the ROC plot is independent of the prevalence of the event in the cohort. Each point on the ROC plot represents a sensitivity / -specificity pair corresponding to a particular decision threshold (i.e. reference). A test with perfect discrimination (no overlap in the two distributions of results) has an ROC plot that passes through the upper left corner, where the true-positive fraction is 1.0, or 100% (perfect sensitivity), and the false-positive fraction is 0 (perfect specificity). The theoretical plot for a test with no discrimination (identical distributions of results for the two groups) is a 45° diagonal line from the lower left corner to the upper right corner. Most plots fall in between these two extremes. If the ROC plot falls completely below the 45° diagonal, this is easily remedied by reversing the criterion for "positivity" from "greater than" to "less than" or vice versa. Qualitatively, the closer the plot is to the upper left corner, the higher the overall accuracy of the test. Dependent on a desired confidence interval, a threshold can be derived from the ROC curve allowing for the diagnosis or prediction for a given event with a proper balance of sensitivity and specificity, respectively. Accordingly, the reference to be used for the aforementioned method of the present invention, i.e. a threshold which allows to discriminate between subjects being at risk and not being at risk can be generated, usually, by establishing a ROC for said cohort as described above and deriving a threshold amount therefrom. Dependent on a desired sensitivity and specificity for a diagnostic method, the ROC plot allows deriving suitable thresholds. It will be understood that an optimal sensitivity may be desired for excluding a subject for being at increased risk (i.e. a rule-out), whereas an optimal specificity may be envisaged for a subject to be assessed as being at an increased risk (i.e. a rule-in).

[0055] The term “comparing” as used herein encompasses comparing the determined amount for a biomarker as referred to herein to a reference. It is to be understood that comparing as used herein refers to any kind of comparison made between the value for the amount with the reference. However, it is to be understood that preferably identical types of values are compared with each other, e.g., if an absolute amount is determined, the reference shall also be an absolute amount, if a relative amount is determined, the reference shall also be a relative amount, etc. The aforesaid comparison of identical types of values is also referred to a comparing to a "corresponding" value, e.g. a corresponding reference, herein. The term comparing also encompasses comparing a calculated score with a suitable reference core. Thus, preferably, theLudwig-Maximilians-Universitat Munchen,

[0056] in Vertretung des Freistaates Bayern 16 LMU17275PC corresponding reference is a value of the quantitative parameter allowing the determination of step (c) to be made. Also preferably, the corresponding reference comprises references for each biomarker derived from at least one subject known to suffer from pancreatic cancer. More preferably, the at least one corresponding reference comprises references for each biomarker derived from at least one subject known to not suffer from pancreatic cancer. Most preferably, the reference is a reference predictor score obtained using a trainable model trained as described elsewhere herein.

[0057] The comparison may be carried out manually or computer assisted. The value of the amount and the reference can be, e.g., compared to each other and the said comparison can be automatically carried out by a computer program executing an algorithm for the comparison. The computer program carrying out the said evaluation will provide the desired assessment in a suitable output format. As set forth above, it is also envisaged to calculate a score based on the amounts of the biomarkers, which may also be referred to as "predictor score", in particular a single score, and to compare this score to a reference score. The calculated score in an embodiment combines information on the amounts of the biomarkers. Moreover, in the score, the biomarkers may be weighted in accordance with their contribution to the establishment of the differentiation, wherein the weighting factor of the individual biomarkers may be different. The score can be regarded as a classifier parameter for the assessing as set forth herein. In particular, it enables providing the assessment based on a single score. Thus, the skilled person does not have to interpret the entire information on the amounts of the individual biomarkers. Using a scoring system as described herein, values of different dimensions or units for the biomarkers may be used since the values will be mathematically transformed into the score. Accordingly, e.g. values for absolute concentrations may be combined in a score with peak area ratios and / or enzymatic activity values. The reference score to be applied may be elected based on the desired sensitivity and / or the desired specificity. Preferably, the reference, in particular the reference score, is obtained by training at least one trainable model for assessing pancreas cancer in a subject, as described herein below.

[0058] Preferably, comparing comprises calculating a predictor score according to formula (I):

[0059] 1

[0060] P~ 1 + e~z

[0061] wherein P= predictor score; and z = B0+B1X1+B2X2+ — i-BnXn

[0062] wherein coefficients Bo, Bi, B2, ..., Bnare preferably selected from Table 2 or from Table 3,Ludwig-Maximilians-Universitat Munchen,

[0063] in Vertretung des Freistaates Bayern 17 LMU17275PC and wherein variables Xi, X2, ..., Xnare measurement values of the biomarkers, wherein said measurement values are preferably measurement values indicated as pg / ml and / or are loglO transformed and z-score normalized according to standard methods known to the skilled person. Also preferably, to address biases related to inter-platform variations and internal standards in the quantification of biomarkers, covariate shift correction using linear regression on matched samples is preferably applied. Thus, preferably, for each biomarker, a linear regression model based on the matched training and test samples is provided and correction slope and intercept are provided, followed by adjusting the test dataset according to the derived linear model to maintain consistency across different platforms. Effectiveness of the aforesaid correction may be evaluate e.g. by calculating a Concordance Correlation Coefficient (CCC), which measures the agreement between the corrected test data and the training data.

[0064] Preferably, the aforesaid coefficients are used with at least one significant digit, more preferably with at least two significant digits, still more preferably with at least three significant digits, most preferably with all significant digits as provided in Table 2 and / or Table 3.

[0065] The method comprises step (a) determining the biomarkers Ceramide (dl8:l;C24:0), Lysophosphatidylethanolamine (Cl 8:0), Phosphatidylethanolamine (Cl 8:0, C22:6), Sphingomyelin (dl7: 1; C16:0), and CA19.9 in a sample from a subject. The biomarkers and methods for determining them have been described herein above.

[0066] Preferably, the method comprises further determining at least one, preferably at least two, more preferably at least three, even more preferably at least four, most preferably all, marker(s) selected from the list consisting of Histidine, Proline, Tryptophan, Ceramide (dl8:2,C24:0), Lysophosphatidylethanolamine (C18:2), Sphingomyelin (35:1), Sphingomyelin (41:2), and Sphingomyelin (dl8:2,C17:0), preferably in step (a). In a preferred embodiment, the method comprises further determining at least five, preferably at least six, more preferably at least seven, most preferably all, marker(s) selected from aforesaid list, preferably in step (a). Also preferably, the method further comprises determining at least one further biomarker, preferably selected from the list consisting of aspartate aminotransferase, alanine aminotransferase, platelet count, haptoglobin, alpha2-macroglobulin, apolipoprotein Al, bilirubin, cholesterol, hyaluronan, prothrombin index, hepatocyte growth factor (HGF), Tissue inhibitor of metalloproteinases (TIMP), and urea, preferably in step (a). Also preferably, the method comprises further diagnostic steps, preferably sonography, magnetic resonance imaging,Ludwig-Maximilians-Universitat Munchen,

[0067] in Vertretung des Freistaates Bayern 18 LMU17275PC radiography, transient elastography, and / or determining subject age and / or gender, preferably in, preceding, or following step (a).

[0068] The method further comprises step (b) comparing the biomarkers determined in step (a) to at least one corresponding reference. Suitable references and methods of comparing have described herein above. Preferably, step (b) comprises calculating a predictor score for assessing pancreatic cancer from at least two, at least three, more preferably at least four, most preferably all, of the biomarkers referred to herein. Thus, step (b) may in particular comprise calculating a predictor score from the biomarkers Ceramide (dl8:l;C24:0), Lysophosphatidylethanolamine (Cl 8:0), Phosphatidylethanolamine (Cl 8:0, C22:6), Sphingomyelin (dl7: 1 ; C16:0), and CA19.9, preferably as specified herein above.

[0069] The term "comparing" has been specified herein above. In certain embodiments, comparing may comprise a comparison to an implied reference; e.g. in case a predictor score is calculated as a probability of a subject suffering from pancreatic cancer, an express reference value may not be required, since the meaning of a probability of e.g. 0.9, and / or a %probability of e.g.

[0070] 90% of a subject to suffer from pancreatic cancer are self-evident for the skilled person.

[0071] Preferably, the method is at least partially computer-implemented and at least said comparing in step (b) is performed by a trained automated machine learning derived generalized logistic regression model, wherein said model preferably was trained as specified herein below.

[0072] The method further comprises step (c) assessing pancreatic cancer in said subject based on said comparing in step (b). The term "assessing" has been specified herein above.

[0073] Advantageously, it was found in the work underlying the present invention that the method of assessing pancreas cancer as described herein allows for improved assessment of pancreatic cancer. In particular, the biomarker combinations (signatures) presented herein are a cost-effective tools designed to achieve a very high NP V, with very high specificity to safely exclude PDAC in patients with, CP or pancreatic lesions which require further diagnostic assessment. The signatures significantly outperform CAI 9.9 alone and thus can serve as an adjunct screening tool in higher-risk cohorts and is planned for the exclusion of PDAC in new-onset diabetes patients.Ludwig-Maximilians-Universitat Munchen,

[0074] in Vertretung des Freistaates Bayern 19 LMU17275PC The definitions made above apply mutatis mutandis to the following. Additional definitions and explanations made further below also apply for all embodiments described in this specification mutatis mutandis.

[0075] The present invention also relates to a method for assessing and treating pancreatic cancer, said method comprising the steps of the method for assessing pancreatic cancer as specified herein above and the further step of treating said pancreatic cancer in a subject identified to suffer therefrom.

[0076] Methods for treating pancreatic cancer are, in principle, known in the art and include in particular surgery and chemotherapy.

[0077] The present invention further relates to a computer-implemented training method of training at least one trainable model for assessing pancreas cancer in a subject, the method comprising: (i) providing the trainable model;

[0078] (ii) retrieving labeled training subject data comprising biomarker data of subjects having known pancreas cancer states, wherein said biomarker data are data of the biomarkers Ceramide (dl8:l;C24:0), Lysophosphatidylethanolamine (C18:0), Phosphatidylethanolamine (C18:0, C22:6), Sphingomyelin (dl7: 1 ; C16:0), and CA19.9; and

[0079] (iii) training the trainable model on the labeled training subject data,

[0080] wherein said trainable model is a Generalized Linear Model (GLM).

[0081] The term “training”, as used herein, relates to a process of determining parameters of at least one machine learning and / or deep learning model, specifically of the algorithm of the machine learning and / or deep learning model, specifically on at least one training data set or set of training data. The training specifically may comprise at least one optimization or tuning process, wherein a best parameter combination, e.g. according to at least one optimization procedure, is determined.

[0082] The term “trainable model”, as used herein, relates to at least one mathematical model configured for transforming one or more input values into one or more output values by using one or more parameters which may be adjusted in order to enable the model to be trained. The trainable model specifically may be a trainable mathematical model which is trainable on at least one training data set using one or more of machine learning, deep learning, neuralLudwig-Maximilians-Universitat Munchen,

[0083] in Vertretung des Freistaates Bayern 20 LMU17275PC networks, or other form of artificial intelligence. The term specifically may refer, without limitation, to the fact that the trainable model can be further trained, optimized or updated based on additional training data. Specifically, the trainable model is trained on a training dataset. The trainable model may be trained by using machine learning. The trainable model may be at least partially data-driven by being trained on data from historical training data.

[0084] The trainable model specifically may comprise at least one trainable model selected from the group consisting of a decision tree model, specifically at least one of an XGBoost model and a Random Forest model; a Support Vector Machine (SVM) model, specifically, a linear kernel SVM; a nearest neighbors model, specifically a KNeighbors Classifier; a Bayes model, specifically a GaussianNB() model; a regression model, specifically at least one of a LogisticRegression; a linear and a nonlinear regression model, e.g. a regression model comprising transformed features, such as log-transformed or polynomial; an Artificial Neural Network (ANN), specifically a non-linear Artificial Neural Network (ANN), in particular a deep learning architecture such as Convolutional NN, Recurrent NN, Long Short Term Memory NN, and the like; a kernel-based method; a tree regression model; a distributed gradientboosting framework model, specifically a light gradient-boosting machine model (LightGBM classifier). As referred to herein, the trainable model is a Generalized Linear Model (GLM), preferaly a binomial GLM.

[0085] Consequently, the term “trained model”, as used herein, relates to a trainable model which has gone through at least one training process as defined above, by applying, at least once, a set of training data to the trainable model. Specifically, one or more parameters of the trainable model might have been adapted, on the basis of the training data, in order to transform the trainable model into a trained model. Specifically, the trained model may be a trainable model which was trained on at least one training dataset, also denoted training data. The trained model may be or may comprise a classifier, configured for classifying an object, on the basis of one or more input variables, describing the object. The trained model may comprise at least one trained model selected from the group of trainable models as specified herein above.

[0086] The trainable model and / or the trained model specifically may be a model taken from a software library, providing a plurality of trainable models. As an example, the Scikit-learn open-source software library may be mentioned, for trainable models and / or trained models being programmed in the Python programming language. Additionally or alternatively, also, as anLudwig-Maximilians-Universitat Munchen,

[0087] in Vertretung des Freistaates Bayern 21 LMU17275PC example, reference may be made to the extreme Gradient Boosting (XGBoost) open-source software library, specifically providing gradient boosting trainable models, being programmed in one or more of the programming languages C++, Java, Python, R, Julia, Perl and Scala. Other libraries, however, are also feasible, as well as customized trainable models not retrieved from software libraries.

[0088] The term “classify”, as used herein, relates to the process of assigning an object to one or more predetermined or determinable classes. The classifying specifically may comprise a prediction or assignment providing an indication to which class out of a plurality of classes an object belongs, e.g. in accordance with one or more predetermined properties of the object.

[0089] The term “pancreas cancer state”, as used herein, relates to at least one item of information indicating at least one physical state or health state of a subject with respect to pancreas cancer. Thus, a pancreas cancer state may e.g. be "pancreas cancer: yes" or pancreas cancermo", may be "afflicted by pancreas cancer", or the like. As referred to herein, indication of a non-cancer state, e.g. "chronic pancreatitis:yes", may also be a pancreas cancer state, in as far as it implies that the subject does not suffer from pancreas cancer.

[0090] As outlined above, the trainable model is configured for classifying a subject’s health condition into a pancreas cancer state. Thus, as discussed above, the trainable model specifically may assign a subject's health condition to a state of suffering from pancreatic cancer, or not.

[0091] The term “retrieve”, as used herein, relates to the process of obtaining data from a data source. The data source may vary, in accordance with the specific application. Thus, in the context of the present invention, the retrieving of the labeled training subject data may take place by at least one of the following: downloading the labeled training subject data from at least one data source, such as from at least one data storage device, from a web- or cloud-based data storage device; obtaining the labeled training subject data via at least one computer network, such as the Internet; obtaining the labeled training subject data via at least one wire-based and / or wireless interface. The retrieving of the training data may fully or partially take place automatically, such as by automatic download of the labeled training subject data from publications of data on clinical studies, and / or may fully or partially take manually. Semiautomatic retrieving processes are also possible. Thus, as an example, data of clinical studies may be downloaded, e.g. automatically, and the data may be processed in order to generate theLudwig-Maximilians-Universitat Munchen,

[0092] in Vertretung des Freistaates Bayern 22 LMU17275PC labelled training subject data, e.g. by manual labeling of the training subject data downloaded from a web-based data source and / or by other data cleaning or data processing steps. Additionally or alternatively, the retrieving of the labeled training subject data may also comprise the actual measurement of the data, e.g., by evaluating subject samples, preferably as specified herein elsewhere.

[0093] The term “labeled”, as used herein, relates to the property of training data having, besides the actual biomarker data, additional information indicating a classification of the data. Thus, as an example, the labeled training subject data may contain, specifically for each subject or each data set of a subject, besides the biomarker data, an indication of the known pancreas cancer state of the subject. Thus, as an example, for each subject or each data set of the subject, besides the biomarker data of their respective subject, information may be comprised, indicating the pancreas cancer state of the subject. Additionally, the following state may also be comprised: the subject being in an apparently healthy state. The pancreas cancer state of the respective subjects may be obtained by other diagnostic means, such as by detecting one or more of the above-mentioned symptoms of the respective subjects contributing to the training subject data.

[0094] Consequently, the labeled training subject data may be obtained by using data from a plurality of subjects, such as by using data from one or more clinical studies. As an example, data from at least 100 subjects, specifically data from at least 500 subjects or more, may be used for generating the labeled training subject data. For, each subject, firstly, biomarker data may be generated. Further, for each subject, information on the pancreas cancer state of the respective subject may be obtained. As outlined above, diagnostic indicators other than biomarker data are preferably used for obtaining the pancreas cancer state of the respective subject, such as by detecting the presence or absence of the characteristic symptoms for each of the pancreas cancer states. Then, for each subject, the biomarker data are labeled at least with the respective information on the pancreas cancer state of the respective patient, in order to obtain a labeled training data set for the respective subject. In addition, and optionally, additional information on the respective subject may be comprised in the labeling, such as information on the age of the subject, information on the gender of the subject, medical information on the subject and the like, preferably as specified herein above. The aggregate of the labeled training data set of the subjects may then form part of the labeled training subject data.Ludwig-Maximilians-Universitat Munchen,

[0095] in Vertretung des Freistaates Bayern 23 LMU17275PC Preferably, the labeled training data may also be obtained by using data from one or more medical studies. Also, preferably, the training data are those of Table 4a, optionally further including the training data of Table 4b herein below.

[0096] The term “biomarker data”, as used herein, relates to all data pertaining to amount or concentrations of biomarkers in a sample of a subject. Preferably, the biomarker data are obtained as specified herein elsewhere. The biomarker data preferably comprise biomarker data from a sample as specified herein above.

[0097] One or more than one trainable model may be used. Even though the method generally relates to trainable models suited for classifying the patient’s health condition, the trainable model specifically may be or may comprise at least one trainable SVM model and / or at least one trainable XGBoost model. These models, for the present application, out of the available trainable model is in standard software libraries, have shown to provide excellent classification results for classifying a patient’s health condition according to pancreas cancer states.

[0098] As outlined above, the label of the training data specifically may contain information on the respective pancreas cancer state. Thus, generally, the labelling of the training subject data to obtain the labeled training subject data may contain information on the respective pancreas cancer state, as shown e.g. in Table 4a. The labeling may be part of the method or may be part of a preceding step of generating the training data, e.g. during one or more clinical studies.

[0099] The predetermined pancreas cancer states may comprise or may consist of: the patient suffering from pancreas cancer; the patient not suffering from pancreas cancer and / or the patient suffering from a benign pancreatic disease. Thus, the trainable model, when being in a trained state, may be configured for assessing patients suspected to suffer from a pancreas cancer state. Besides the categories listed above, one or more additional states may be comprised by the predetermined group of the pancreas cancer states, in particular as specified herein above; e.g. the subject suffering from pancreatic adenocarcinoma and / or the subject suffering from pancreatic ductal adenocarcinoma (PDAC) and / or the subject being in an apparently healthy state and / or the subject suffering from pancreatitis.Ludwig-Maximilians-Universitat Munchen,

[0100] in Vertretung des Freistaates Bayern 24 LMU17275PC The biomarker data of the labeled training subject data in step (ii) preferably are derived from a blood or blood-derived sample and are, preferably, obtained by MS analysis, more preferably by LC-MS.

[0101] The training method further may comprise at least one hyperparameter tuning step for tuning hyperparameters of the trainable model. The term “hyperparameter tuning step”, as used herein, relates to the process of adjusting, specifically optimizing, one or more hyperparameters of the trainable model. Thus, trainable models, e.g. trainable models from standard software libraries such as Scikit-learn, typically require one or more hyperparameters to be predetermined before starting the training process. The hyperparameter tuning step may comprise the pre-adjusting of these hyperparameters, e.g. in order to provide a suitable convergence for the training process. As an example, the labeled training data or at least a part thereof may be used for an initial test for training the trainable model, trying different sets of hyperparameters, wherein the performance of the training process is evaluated on the basis of measurable training performance values such as the speed of convergence, and, wherein, the set of hyperparameters providing a superior performance may finally be chosen for the trainable model. The trainable model may then be trained with this specific set of hyperparameters.

[0102] The hyperparameter tuning step, specifically, may be performed at least once before performing step (iii). It is, however, also possible to implement a plurality of tuning steps, intermitting with one or more partial training steps, e.g. in order to improve or even optimize the tuning of the hyperparameters during the process.

[0103] As outlined above, the trainable model specifically may comprise or specifically may consist of a trainable GLM model. As also outlined in further detail below, the trainable GLM model specifically may be trained using the following hyperparameters:

[0104] Regularization: Ridge; Logistic regression link: logit; Number of iterations: 40, Nlambad = 30, Lambda.max = 32.331; Sort metric = “Fl”, i.e. preferably the Fl score is calculated from the harmonic mean of the precision and recall; Stopping tolerance = 3, i.e. preferably if the performance of iteration is not improved after 3 subsequent rounds.

[0105] As will be outlined in further detail below, this set of hyperparameters is suited to provide training results leading to a trained GLM model on the basis of the labeled training data, the trained GLM model being capable of classifying the patient’s health condition with highLudwig-Maximilians-Universitat Munchen,

[0106] in Vertretung des Freistaates Bayern 25 LMU17275PC reliability. For further details on these hyperparameters, as an example, reference may be made to the Scikit-learn library.

[0107] The training subject data, as discussed above, specifically may be obtained from one or more patient studies. Specifically, the labeled training subject data may be patient data obtained from a plurality of different studies. More specifically, the labeled training subject data may contain patient data from differing age groups. The label of the training subject data, as outlined above, may contain information on the study. Additionally or alternatively, the label of the training subject data may contain information on the patient, such as on the group the respective patient belongs to. More specifically, the label may contain age information, such as information on the respective age group of the patients. Thus, the labelling of the training subject data to obtain the labeled training data may additionally contain information on the respective age groups of the patients.

[0108] The training method specifically may comprise splitting the labeled training subject data into a training data set and a test data set. Thus, more specifically, the splitting into a training data set and a test data set may comprise x%-y% splitting of the labeled training subject data, with x% being in the range 60 to 90, specifically in the range 65 to 75, and more specifically x%=70, and with y%=100-x. Therein, x% denotes the training data set and y% denotes the test data set.

[0109] The training method may further compromise a cross-validation in the training data set. Thus, the data of the training data set may be further split into k subsets or folds of equal size, with k being at least 2 or more folds, specifically in the range 3 to 10, more specifically k=5. The training and validation may be repeated k times, and each time the data of a different subset or fold may be used for validation. Additionally or alternatively, the training method may further comprise repeating the splitting of the labeled training data into the training data set and the test data set one or more times, specifically for a number of times in the range of 10 to 40 times, more specifically 20 times. Thus, a validation of the model may be performed with an xl%-yl% splitting of the labeled training subject data, and at least one additional x2%-y2% splitting of the labeled training subject data, with xl and x2 describing different partitions of the labeled training subject data, wherein, specifically, xl%=x2%. As an example, a degree of discrepancy between the results obtained for the multiple iterations may be used for qualifying and / or quantifying the training result.Ludwig-Maximilians-Universitat Munchen,

[0110] in Vertretung des Freistaates Bayern 26 LMU17275PC Additionally or alternatively, the training method may further comprise repeatedly performing step (iii) with multiple splittings. The multiple splittings may comprise an xl%-yl% splitting of the labeled training subject data and at least one additional x2%-y2% splitting of the labeled training subject data, with xl and x2 describing different partitions of the labeled training subject data, wherein, specifically, xl%=x2%. Alternatively or additionally, the multiple splittings may comprise differing splittings of the labeled training data, specifically with xl% x2%. Again, optionally, the training results, e.g. at least one numeric value qualifying and / or quantifying the training results, e.g. qualifying and / or quantifying the accuracy of the training, may be compared.

[0111] As outlined above, the training method specifically may comprise splitting the labeled training subject data into a training data set and a test data set. The training method may, then, further comprises the following step:

[0112] at least one validation step, the validation step comprising testing the accuracy of the training by using the test data set.

[0113] The testing of the accuracy specifically may comprise using at least one model accuracy metrics. Various model accuracy metrics are generally known in the field of machine learning. Specifically, one or more or even all of the following model accuracy metrics may be used: overall accuracy, balanced accuracy, weighted fl score. Other options, however, are also feasible.

[0114] The training method may further comprise the following step:

[0115] determining at least one target set of coefficients Bo, Bi, ..., Bnas specified herein above, by using at least one feature extraction method. Again, various feature extraction methods are generally known to the skilled person in the field of machine learning.

[0116] In a further aspect of the present invention, a computer-implemented method of assessing pancreas cancer in a subject is disclosed. The computer-implemented method comprises the following method steps, which preferably are performed in the given order. However, a different order is also feasible. Further, it is possible to perform two or more of the method steps simultaneously or in a fashion overlapping in time. Further, it is also possible to perform one, more than one or even all of the method steps repeatedly.Ludwig-Maximilians-Universitat Munchen,

[0117] in Vertretung des Freistaates Bayern 27 LMU17275PC The computer-implemented method for assessing pancreas cancer comprises the following steps:

[0118] (A) retrieving at least one trained model, preferably as specified herein above;

[0119] (B) retrieving biomarker data of the subject; and

[0120] (C) applying the trained model to the biomarker data, thereby assessing pancreas cancer in the subject.

[0121] The trained model specifically may be trained or have been trained by using the training method according to the present invention, such as according to any one of the embodiments described above and / or according to any one of the embodiments described in further detail below. Thus, for most of the terms, definitions and options as described above or as described in further detail below apply mutatis mutandis.

[0122] As outlined in further detail above, the computer-implemented method specifically may assign the patient or the patient’s health condition to one of the pancreas cancer states of the predetermined group of pancreas cancer states. The subject preferably is a subject as specified herein above. The group of predetermined pancreas cancer states may be the same as discussed above in the context of the training method. The biomarker data preferably comprise data for biomarkers also determined during the training of the trained model. Thus, as discussed above, the trained model may be trained by using the training method described herein.

[0123] The present invention also relates to a system comprising (i) a measuring device for determining the biomarkers Ceramide (dl8:l;C24:0), Lysophosphatidylethanolamine (C18:0), Phosphatidylethanolamine (C18:0, C22:6), Sphingomyelin (dl7: 1 ; C16:0), and CA19.9; and (ii) an evaluation device operably linked to the measuring device, said evaluation device comprising a data processor comprising instructions for carrying out a comparison of the biomarkers determined by the measuring device to at least one corresponding reference.

[0124] The term “system”, as used herein, relates to a collection of the indicated means which are operatively linked to each other. Said means may be implemented in a single system or may be physically separated devices which are operatively linked to each other. Preferably, the system further comprises an output device outputting the result of the assessment performed by the evaluation device. Preferably, the system further comprises a data storage device configured for storing, in particular in electronic form, data provided by the devices of the system, dataLudwig-Maximilians-Universitat Munchen,

[0125] in Vertretung des Freistaates Bayern 28 LMU17275PC provided by user input, in particular one or more reference, and / or instructions for performing the assessment, in particular machine-readable instructions for performing the assessment as specified herein above. As is understood by the skilled person, said data storage device(s) may be part of the measuring device and / or the evaluation device, and / or may be standalone data storage device(s) operatively linked to the other devices of the system.

[0126] The term "measuring device", as referred to herein, includes each and every device allowing determining at least one, preferably at least two, more preferably at least three, even more preferably at least four, most preferably all biomarkers referred to herein. Preferably, the measuring device comprises a receptacle for a sample. The receptacle may directly contact the sample, or may be a receptacle for a further means receiving the sample, wherein the further means may be e.g. a multi-well plate, to which a sample or a multiplicity of samples may be applied. Preferably, the measuring device comprises at least one detector for determining at least one biomarker, e.g., an MS device, a biosensor, a solid support coupled to a ligand specifically recognizing a biomarker, a Plasmon surface resonance devices, an NMR spectrometer, and the like).

[0127] The term "evaluation device", as referred to herein, may be any means capable of providing the analysis as specified; preferably, the evaluation device is or comprises a data processing means, in particular at least one processor, such as a microprocessor, a handheld device such as a mobile phone, or a computer. How to link the means in an operating manner will depend on the type of means included into the device. Preferably, the means are comprised by a single device. Preferably, the instructions and interpretations are comprised in an executable program code comprised in the device, such that, as a result of determination, an assessment of pancreatic cancer may be output to a user. Typical devices are those which can be applied without the particular knowledge of a specialized technician.

[0128] Preferably, the system comprises at least one processor, the processor being configured, e.g. by software programming, for performing the training method and / or the computer-implemented method for assessing pancreas cancer according to the present invention. More preferably, the system is system for assessing pancreas cancer and comprises at least one processor, the processor being configured, e.g. by software programming, for performing the method for assessing pancreas cancer, preferably the computer-implemented method for assessing pancreas cancer, according to the present invention.Ludwig-Maximilians-Universitat Munchen,

[0129] in Vertretung des Freistaates Bayern 29 LMU17275PC

[0130] The term “processor”, as used herein, relates to an arbitrary logic circuitry configured for performing basic operations of a computer or system, and / or, generally, to a device which is configured for performing calculations or logic operations. In particular, the processing unit may be configured for processing basic instructions that drive the computer or system. As an example, the processor may comprise at least one arithmetic logic unit (ALU), at least one floating-point unit (FPU), such as a math co-processor or a numeric coprocessor, a plurality of registers, specifically registers configured for supplying operands to the ALU and storing results of operations, and a memory, such as an LI and L2 cache memory. In particular, the processor may be a multi-core processor. Specifically, the processing unit may be or may comprise a central processing unit (CPU). Additionally or alternatively, the processor may be or may comprise a microprocessor, thus specifically the processing unit’s elements may be contained in one single integrated circuitry (IC) chip. Additionally or alternatively, the processing unit may be or may comprise one or more application-specific integrated circuits (ASICs) and / or one or more field-programmable gate arrays (FPGAs) or the like. The processing unit specifically may be configured, such as by software programming, for performing one or more evaluation operations.

[0131] The present invention also relates to an automated machine learning model obtained or obtainable according to the method according to the method of the present invention, preferably tangibly embedded on a data storage means.

[0132] The present invention also relates to a database for assessing pancreas cancer comprising at least one set of biomarker coefficients for calculating a predictor score from quantitative data of the biomarkers Ceramide (dl8:l;C24:0), Lysophosphatidylethanolamine (C18:0), Phosphatidylethanolamine (C18:0, C22:6), Sphingomyelin (d!7: 1 ; C16:0), and CA19.9 in a sample, and optionally at least one reference predictor score and / or at least one reference predictor range.

[0133] In a further aspect of the present invention, a computer program is proposed, the computer program comprising instructions which, when the program is executed by at least one processor of a computer or a computer network cause the processor to perform the training method according to the present invention, such as according to any one of the embodiments of theLudwig-Maximilians-Universitat Munchen,

[0134] in Vertretung des Freistaates Bayern 30 LMU17275PC training method described above and / or according to any one of the embodiments of the training method described in further detail below.

[0135] In a further aspect of the present invention, a computer program is proposed, the computer program comprising instructions which, when the program is executed by at least one processor of a computer or a computer network cause the processor to perform the assessment method according to the present invention, such as according to any one of the embodiments of the assessment method described above and / or according to any one of the embodiments of the assessment method described in further detail below.

[0136] In a further aspect of the present invention, a computer-readable storage medium is proposed, specifically a non-transient computer-readable storage medium, the computer readable storage medium comprising instructions which, when the instructions are executed by at least one processor of a computer or a computer network cause the processor to perform the training method according to the present invention, such as such as according to any one of the embodiments of the training method described above and / or according to any one of the embodiments of the training method described in further detail below.

[0137] In a further aspect of the present invention, a computer-readable storage medium is pro-posed, specifically a non-transient computer-readable storage medium, comprising instructions which, when the instructions are executed by at least one processor of a computer or a computer network cause the processor to perform the assessment method according to the present invention, such as according to any one of the embodiments of the assessment method described above and / or according to any one of the embodiments of the assessment method described in further detail below.

[0138] As outlined above, the methods as proposed herein are computer-implemented. As used herein, the terms “computer-readable data carrier” and “computer-readable storage medium” specifically may refer to non-transitory data storage means, such as a hardware storage medium having stored thereon computer-executable instructions. The computer-readable data carrier or storage medium specifically may be or may comprise a storage medium such as a randomaccess memory (RAM) and / or a read-only memory (ROM).Ludwig-Maximilians-Universitat Munchen,

[0139] in Vertretung des Freistaates Bayern 31 LMU17275PC Thus, specifically, the method steps of the methods as indicated above may be performed by using a computer or a computer network, preferably by using a computer program.

[0140] Further disclosed and proposed herein is a computer program product having program code means, in order to perform any one of the methods according to the present invention in one or more of the embodiments enclosed herein when the program is executed on a computer or computer network. Specifically, the program code means may be stored on a computer-readable data carrier and / or on a computer-readable storage medium.

[0141] In view of the above, the following embodiments are particularly envisaged:

[0142] Embodiment 1: A method for assessing pancreatic cancer in a subject, said method comprising

[0143] (a) determining the biomarkers Ceramide (dl8:l;C24:0), Lysophosphatidylethanolamine (Cl 8:0), Phosphatidylethanolamine (Cl 8:0, C22:6), Sphingomyelin (dl7:l; Cl 6:0), and CA19.9 in a sample from said subject;

[0144] (b) comparing the biomarkers determined in step (a) to at least one corresponding reference; and

[0145] (c) assessing pancreatic cancer in said subject based on said comparing in step (b).

[0146] Embodiment 2: The method of embodiment 1, wherein said pancreatic cancer is pancreatic adenocarcinoma, preferably pancreatic ductal adenocarcinoma (PDAC).

[0147] Embodiment s: The method of embodiment 1 or 2, wherein assessing comprises excluding pancreatic cancer, preferably comprises identifying a subject not suffering from pancreatic cancer.

[0148] Embodiment 4: The method of any one of embodiments 1 to 3, wherein said assessing comprises diagnosing pancreatic cancer, preferably comprises diagnosing PDAC.

[0149] Embodiment 5: The method of any one of embodiments 1 to 4, wherein said assessing comprises staging pancreatic cancer.

[0150] Embodiment 6: The method of any one of embodiments 2 to 5, wherein said assessing comprises prognosticating pancreatic cancer.

[0151] Embodiment 7: The method of any one of embodiments 1 to 6, wherein said assessing comprises differentiating between pancreatic cancer and a non-pancreatic cancer disorder, preferably pancreatitis, more preferably chronic pancreatitis.Ludwig-Maximilians-Universitat Munchen,

[0152] in Vertretung des Freistaates Bayern 32 LMU17275PC Embodiment 8: The method of any one of embodiments 1 to 7, wherein said subject is a human.

[0153] Embodiment 9: The method of any one of embodiments 1 to 8, wherein said subject is a subject at risk of suffering from pancreatic cancer.

[0154] Embodiment 10: The method of any one of embodiments 1 to 9, wherein said subject with an increased risk of suffering from pancreatic cancer is known or suspected to suffer from new onset type 2 diabetes, chronic pancreatitis, pancreatic cystic lesions, and / or is from a familial PDAC kinship.

[0155] Embodiment 11 : The method of any one of embodiments 1 to 10, wherein said method is comprised in a method of monitoring a subject at risk of suffering from pancreatic cancer. Embodiment 12: The method of any one of embodiments 1 to 11, wherein said method comprises further determining at least one, preferably at least two, more preferably at least three, even more preferably at least four, most preferably all, marker(s) selected from the list consisting of Histidine, Proline, Tryptophan, Ceramide (dl8:2,C24:0), Lysophosphatidylethanolamine (C18:2), Sphingomyelin (35:1), Sphingomyelin (41:2), and Sphingomyelin (dl8:2,C17:0), preferably in step (a).

[0156] Embodiment 13: The method of any one of embodiments 1 to 12, wherein at least one, preferably each, biomarker is determined by a method comprising independently selected a mass spectrometry (MS), preferably LC / MS, more preferably LC / MS-MS, an enzymatic assay, and / or an immunoassay.

[0157] Embodiment 14: The method of any one of embodiments 1 to 13, wherein said sample is a bodily fluid sample, preferably a blood sample or a blood-derived sample, more preferably a blood, plasma, or serum sample, even more preferably a plasma or serum sample.

[0158] Embodiment 15: The method of any one of embodiments 1 to 14, wherein the concentration of CA19.9 in said sample is known or suspected to be less than 37 U / ml, preferably less than 10 U / ml.

[0159] Embodiment 16: The method of any one of embodiments 1 to 15, wherein said determining said biomarkers comprises determining for each of said biomarkers a value of a quantitative parameter correlating with the amount of said biomarker in said sample.

[0160] Embodiment 17: The method of any one of embodiments 1 to 16, wherein step (b) comprises calculating a predictor score for assessing pancreatic cancer from all of said biomarkers.Ludwig-Maximilians-Universitat Munchen,

[0161] in Vertretung des Freistaates Bayern 33 LMU17275PC Embodiment 18: The method of embodiment 17, wherein step (b) comprises determining all biomarkers of embodiment 1 and wherein said predictor score is calculated according to formula (I)

[0162] 1

[0163] P~ 1 + e~z

[0164] wherein P= predictor score; and z = B0+B1X1+B2X2+ ••• +BnXn

[0165] and wherein variables Xi, X2, ... , Xnare measurement values of said biomarkers.

[0166] Embodiment 19: The method of embodiment 18, wherein coefficients Bo, Bi, B2, ..., Bnare selected from Table 2 or from Table 3.

[0167] Embodiment 20: The method of any one of embodiments 1 to 19, wherein said at least one corresponding reference is a value of said quantitative parameter allowing the determination of step (c) to be made.

[0168] Embodiment 21 : The method of any one of embodiments 1 to 20, wherein said at least one corresponding reference comprises references for each biomarker derived from at least one subject known to suffer from pancreatic cancer.

[0169] Embodiment 22: The method of embodiment 21, wherein amounts for each of the biomarkers being essentially identical or similar to the at least one corresponding reference are indicative for a subject suffering from pancreatic cancer; and / or amounts for each of the biomarkers being different from the corresponding references are indicative for a subject not suffering from pancreatic cancer.

[0170] Embodiment 23 : The method of any one of embodiments 1 to 22, wherein said at least one corresponding reference comprises references for each biomarker derived from at least one subject known to not suffer from pancreatic cancer.

[0171] Embodiment 24: The method of embodiment 23, wherein amounts for each of the biomarkers being essentially identical or similar to the corresponding references are indicative for a subject not suffering from pancreatic cancer; and / or amounts for each of the biomarkers being different from the corresponding references are indicative for a subject suffering from pancreatic cancer.

[0172] Embodiment 25: The method of any one of embodiments 1 to 24, wherein said method comprises determining at least one further biomarker, preferably selected from the list consisting of aspartate aminotransferase, alanine aminotransferase, platelet count, haptoglobin, alpha2-macroglobulin, apolipoprotein Al, bilirubin, cholesterol, hyaluronan, prothrombin index, hepatocyte growth factor (HGF), Tissue inhibitor of metalloproteinases (TIMP), and urea.Ludwig-Maximilians-Universitat Munchen,

[0173] in Vertretung des Freistaates Bayern 34 LMU17275PC Embodiment 26: The method of any one of embodiments 1 to 25, wherein said method comprises further diagnostic steps, preferably sonography, magnetic resonance imaging, radiography, transient elastography, and / or determining subject age and / or gender.

[0174] Embodiment 27: The method of embodiment 26, wherein said method is computer-implemented.

[0175] Embodiment 28: The method of any one of embodiments 1 to 27, wherein said method is computer implemented and wherein said comparing in step (b) is performed by a trained automated machine learning derived generalized logistic regression model.

[0176] Embodiment 29: The method of embodiment 28, wherein said model was trained as specified in any one of embodiments 31 to 47.

[0177] Embodiment 30: The method of any one of embodiments 1 to 29, wherein said reference is a reference predictor score obtained using a trainable model trained according to the method according to any one of embodiments 31 to 47.

[0178] Embodiment 31: A computer-implemented training method of training at least one trainable model for assessing pancreas cancer in a subject, the method comprising:

[0179] (i) providing the trainable model;

[0180] (ii) retrieving labeled training subject data comprising biomarker data of subjects having known pancreas cancer states, wherein said biomarker data are data of the biomarkers Ceramide (dl8:l;C24:0), Lysophosphatidylethanolamine (C18:0), Phosphatidylethanolamine (C18:0, C22:6), Sphingomyelin (dl7: 1 ; C16:0), and CA19.9; and

[0181] (iii) training the trainable model on the labeled training subject data,

[0182] wherein said trainable model is a Generalized Linear Model (GLM).

[0183] Embodiment 32: The method of embodiment 31, wherein said assessing is classifying pancreas cancer.

[0184] Embodiment 33: The method of embodiment 31 or 32, wherein the trainable model is a binomial GLM.

[0185] Embodiment 34: The method of any one of embodiments 31 to 33, wherein the labelling of the training subject data to obtain the labeled training subject data contains information on the respective pancreas cancer state.

[0186] Embodiment 35: The method of any one of embodiments 31 to 34, wherein the group of predetermined pancreas cancer states comprises, preferably consists of: the subject suffering from pancreatic adenocarcinoma and / or the subject suffering from pancreatic ductal adenocarcinoma (PDAC).Ludwig-Maximilians-Universitat Munchen,

[0187] in Vertretung des Freistaates Bayern 35 LMU17275PC Embodiment 36: The method of any one of embodiments 31 to 35, wherein the group of predetermined pancreas cancer states further comprises: the subject being in an apparently healthy state and / or the subject suffering from pancreatitis.

[0188] Embodiment 37: The method of any one of embodiments 31 to 36, wherein the biomarker data of the labeled training subject data in step (ii) are derived from a bodily fluid sample, preferably a blood sample or a blood-derived sample, more preferably a blood, plasma, or serum sample, even more preferably a plasma or serum sample and are, preferably, obtained by a method comprising independently selected a mass spectrometry (MS), preferably LC / MS, more preferably LC / MS-MS, an enzymatic assay, and / or an immunoassay.

[0189] Embodiment 38: The method of any one of embodiments 31 to 37, wherein the trainable GLM model is trained using the following hyperparameters:

[0190] Regularization: Ridge;

[0191] Logistic regression link: logit;

[0192] Number of iterations: 40,

[0193] Nlambad—30,

[0194] Lambda.max = 32.331

[0195] Sort_metric = “Fl”, and / or

[0196] Stopping tolerance = 3.

[0197] Embodiment 39: The method of any one of embodiments 31 to 38, wherein the labeled training subject data contain subject data from differing age groups.

[0198] Embodiment 40: The method of any one of embodiments 31 to 39, wherein the labelling of the training subject data to obtain the labeled training data additionally contains information on the respective age groups of the subjects.

[0199] Embodiment 41: The method of any one of embodiments 31 to 40, wherein the training comprises splitting the labeled training subject data into a training data set and a test data set. Embodiment 42: The method of embodiment 41, wherein the splitting into a training data set and a test data set comprises a x%-y% splitting of the labeled training subject data, with x% being in the range 60 to 90, specifically in the range 65 to 75, and more specifically x%= 70, and with y%=100-x.

[0200] Embodiment 43 : The method of any one of embodiments 31 to 42, wherein the method comprises a cross-validation in the training data set.

[0201] Embodiment 44: The method of any one of embodiments 31 to 43, wherein the method further comprises at least one validation step, the validation step comprising testing the accuracy of the training by using the test data set.Ludwig-Maximilians-Universitat Munchen,

[0202] in Vertretung des Freistaates Bayern 36 LMU17275PC Embodiment 45: The method of any one of embodiments 31 to 44, wherein testing of the accuracy comprises using at least one model accuracy metrics, specifically at least one model accuracy metrics selected from the group consisting of area under the curve (AUC), overall accuracy, balanced accuracy, positive predictive value, negative predictive value, true positive rate, true negative rate, false positive rate, false negative rate, and diagnostic odds ratio, preferably AUC.

[0203] Embodiment 46: The method of any one of embodiments 31 to 45, wherein said training data comprise the data of Table 4a and, optionally of Table 4b.

[0204] Embodiment 47: The method of any one of embodiments 31 to 46, further comprising at least one feature of any one of embodiments 1 to 30.

[0205] Embodiment 48: An automated machine learning model obtained or obtainable according to the method according to any one of embodiments 31 to 47, preferably tangibly embedded on a data storage means.

[0206] Embodiment 49: A database for assessing pancreas cancer comprising at least one set of biomarker coefficients for calculating a predictor score from quantitative data of the biomarkers Ceramide (dl8:l;C24:0), Lysophosphatidylethanolamine (C18:0), Phosphatidylethanolamine (C18:0, C22:6), Sphingomyelin (dl7:l; C16:0), and CA19.9 in a sample, and at least one reference predictor score and / or at least one reference predictor range.

[0207] Embodiment 50: The database of embodiment 49, wherein said database is tangibly embedded on a data storage means.

[0208] Embodiment 51: A computer-implemented method of assessing pancreas cancer in a subject, the method comprising

[0209] (A) retrieving at least one trained model, preferably as specified herein above;

[0210] (B) retrieving biomarker data of the subject; and

[0211] (C) applying the trained model to the biomarker data, thereby assessing pancreas cancer in the subject.

[0212] Embodiment 52: A system comprising

[0213] (i) a measuring device for determining the biomarkers Ceramide (dl8:l;C24:0), Lysophosphatidylethanolamine (Cl 8:0), Phosphatidylethanolamine (Cl 8:0, C22:6), Sphingomyelin (dl7: 1 ; C16:0), and CA19.9; and

[0214] (ii) an evaluation device operably linked to the measuring device, said evaluation device comprising a data processor comprising instructions for carrying out a comparison of the biomarkers determined by the measuring device to at least one corresponding reference.Ludwig-Maximilians-Universitat Munchen,

[0215] in Vertretung des Freistaates Bayern 37 LMU17275PC Embodiment 53: The system of embodiment 52, wherein the instructions for carrying out a comparison comprise instructions to carry out said comparison according to step (b) of the method according to any one of embodiments 1 to 30.

[0216] Embodiment 54: The system of embodiment 52 or 53, adapted to perform a method according to any one of embodiments 1 to 38, preferably comprising tangibly embedded instructions which, when carried out by the data processor, cause the system to perform a method according to any one of embodiments 1 to 38.

[0217] Embodiment 55: The system of any one of embodiments 52 to 54, wherein said system further comprises a database according to embodiment 49 or 50 operably coupled to the data processor.

[0218] Embodiment 56: The system of any one of embodiments 52 to 55, wherein said detection means is capable of specifically detecting said biomarkers.

[0219] Embodiment 57: The system of any one of embodiments 52 to 56, wherein said system is a system for assessing pancreatic cancer in a subject, and wherein said evaluation device further comprises means for assessing pancreatic cancer in said subject based on the comparison. Embodiment 58: The system of any one of embodiments 52 to 57, wherein said evaluation device is capable of automatically receiving values for the amounts of the biomarkers from the measuring device.

[0220] Embodiment 59: A method for assessing and treating pancreatic cancer, said method comprising the steps of the method according to any one of embodiments 1 to 30 and the further step of treating said pancreatic cancer in a subject identified to suffer therefrom.

[0221] Embodiment 60: Use of the biomarkers Ceramide (dl8:l;C24:0), Lysophosphatidylethanolamine (Cl 8:0), Phosphatidylethanolamine (Cl 8:0, C22:6), Sphingomyelin (dl7: 1 ; C16:0), and CA19.9 for assessing pancreatic cancer.

[0222] All references cited in this specification are herewith incorporated by reference with respect to their entire disclosure content and the disclosure content specifically mentioned in this specification.

[0223] Figure Legends

[0224] Figure 1. Trial enrollment and analysis populations. (A) Graphical explanation of cohort enrichment design: The incidence of PDAC in the general population (0.00014%) increases to 1% among individuals with a genetic predisposition, such as those with familial PDAC (>3Ludwig-Maximilians-Universitat Munchen,

[0225] in Vertretung des Freistaates Bayern 38 LMU17275PC FDR), chronic pancreatitis, or lean NOD patients. A sample size calculation for a prospective diagnostic study within this cohort warrants the recruitment of 29,000 patients. By targeting patients with an undefined pancreatic lesion detected through diagnostic CT-scan increases the effect size, hereby enriching the endpoint of PDAC prevalence in this population to 20%. (B) METAP AC study: A total of 1353 participants were enrolled, and 1129 participants were included in the primary analysis. (C) Enrollment of number of patients per center across 23 German centers. (D) Scatter plot illustrating correlation between population prevalence and cohort prevalence for risk factors (orange) and at-risk groups (green). The encircled points illustrate the comparison of PDAC prevalence within the general population versus the prevalence in an enriched cohort. Of note, Incidence equals prevalence in PDAC, Age is depicted as median values and gender is depicted as male to female ratio. R2: Pearson’s correlation coefficient. (E) Tabular representation of metabolites involved in z- or m-Metabolic signature, with X indicating being part of the signature.

[0226] Figure 2. Prediction score of i- and m-Metabolic signature. (A-B) Scatter plot illustrating the i-Metabolic signature (A) and m-Metabolic signature (B) performance for differential diagnosis of PDAC compared to CAI 9.9 alone. The y-axis depicts the predictive scores, whereas the x-axis depicts CAI 9.9 levels. The orange encircled points denote subjects that benefit from the z- o m-Metabolic signature in predicting PDAC (False Negatives with CA19.9 alone), and blue encircled points represent subjects that benefit by avoiding misdiagnosis (False Positives with CAI 9.9 alone). Asterisk depicts controls (IPMNs, Non-IPMNs Cystic lesions, Chronic pancreatitis, Acute pancreatitis and Metastases of extra-pancreatic origins) in the cohort. Adjoining boxplot illustrates the distribution of prediction score of z- or m-Metabolic signature in PDAC vs controls. Student’s t-test for PDAC vs controls in boxplot. (C) Comparison of specificity of z- and m-Metabolite signature with CAI 9.9 alone. Error bars indicate 95% confidence interval. Student’s t-test for z- or m-Metabolite signature vs CA19.9 alone in each subgroup. (D) Random Bootstrapping: A cohort prevalence ranging from 1% to 20% was mimicked by randomly selecting PDAC cases, which were compared to controls to assess the diagnostic performance of CA19.9, the z- (blue) and m-Metabolic signature (red). Here, the AUC, specificity, and sensitivity for both signatures operated independently of the cohort prevalence, while the specificity of CA19.9 increased, and its sensitivity decreased with increased prevalence. <0.05 considered statistically significant.Ludwig-Maximilians-Universitat Munchen,

[0227] in Vertretung des Freistaates Bayern 39 LMU17275PC Figure 3. Prediction score of m-Metabolic signature in SHIP- TREND-1. (A) Patient selection from SHIP-TREND-1. Total 243 out of 2507 participants with NOD were identified and 242 participants were included for further analysis. (B) Boxplot illustrating distribution of prediction score of m-Metabolite signature with and without simulated CA19.9 levels. Student’s t-test for NOD (without PDAC) vs NOD-PDAC in boxplot. P<0.05 considered statistically significant.

[0228] The following Examples shall merely illustrate the invention. They shall not be construed, whatsoever, to limit the scope of the invention.

[0229] Example 1: Study design

[0230] METAP AC (Plasma-Metabolome signature for the diagnosis of pancreatic cancer) is a prospective, consecutive, multi-disciplinary, investigator blinded multicentric study, designed to validate the performance of a previously identified plasma metabolite signatures(Mayerle et al. 2018; Mahajan et al. 2022) for early-stage PDAC exclusion in a population at risk. This phase IV study adhered to the EDRN guidelines (Srivastava and Wagner 2020) and was conducted at 23 German academic and community -based institutions. The study adhered to the STARD (Cohen et al. 2016) and TRIPOD guidelines (Collins et al. 2015). The study protocol was approved by the local institutional review board under the reference number BB 079 / 16 at the University Medicine Greifswald and 736-16 at LMU University. All participants or their legal representatives provided their written informed consent. Clinische Studien Gesellschaft (CSG) mbH, Berlin, Germany collected and monitored the data. The study statistician (first author) analyzed the data and vouches for the completeness and precision of the data and for the fidelity of the study to the protocol, along with the corresponding author.

[0231] Example 2:Study population

[0232] Sample size was calculated based on an expected sensitivity of 88% and specificity of 85%, with a margin of error of 5% and statistical power of 80% (a = 0.05). The PDAC incidence equaling prevalence was expected with 20% in a population with an undefined lesion of the pancreas detected on imaging. Minimum required sample size was determined with 1375 participants. Patients with a undefined pancreatic lesions identified through CT imaging who required further diagnostic evaluation were included. The exclusion criteria were set to be inability to give informed consent, current pregnancy and conditions precluding surgery, biopsy or clinical follow-up.Ludwig-Maximilians-Universitat Munchen,

[0233] in Vertretung des Freistaates Bayern 40 LMU17275PC

[0234] Example 3: Clinical procedures

[0235] Eligible participants with informed consents were recruited and blood samples were collected (EDTA plasma and serum), following overnight fasting. Final diagnosis of PDAC (either resection specimen, fine needle aspiration or biopsies of metastases) was confirmed by histopathology or ruled out in patients with non-pancreatic cancer disorder by a minimum clinical follow-up of 24 months. In case of controls, (IPMNs, non-IPMNs cystic lesions, chronic pancreatitis, acute pancreatitis and metastases of extra-pancreatic origins), final discharged diagnosis for 24 months follow-up considered as final diagnosis. A safety call concluded the study.

[0236] Example 4: Primary outcomes

[0237] The primary efficacy endpoint for MET APAC was set to be the diagnosis of PDAC with high specificity in a population at risk, employing two previously identified metabolite signatures (z-and m-Metabolic signature).

[0238] Example 5: Laboratory procedures

[0239] Whole-blood samples were processed to plasma and stored at -80°C before shipment to the central storage facility (University Medicine Greifswald). Serum CA19.9 levels were measured in a central laboratory (University Medicine Greifswald). Following collection of all biospecimens, 12 metabolites were analyzed for absolute concentration (pg / ml) using LC-MS / MS according to Metabolon Method TAM217 at the Bioanalytical testing site of Metabolon Inc. Morrisville, NC, USA.

[0240] Example 6: Detail targeted metabolite analysis

[0241] The targeted metabolite panel measures 12 analytes: Histidine, Proline, Tryptophan, Ceramide (dl8:l,C24:0), Ceramide (dl8:2,C24:0), Lysophosphatidylethanolamine (C18:0), Lysophosphatidylethanolamine (C18:2), Phosphatidylethanolamine (C18:0), Sphingomyelin (dl7:l,C16:0), Sphingomyelin (35:1), Sphingomyelin (41:2), and Sphingomyelin (dl8:2,C17:0). Analyte concentrations are analyzed by LC-MS / MS (Metabolon method, TAM217).

[0242] Calibration samples are prepared at eight different concentration levels by spiking phosphate buffered saline with corresponding calibration spiking solutions. Calibration samples, studyLudwig-Maximilians-Universitat Munchen,

[0243] in Vertretung des Freistaates Bayern 41 LMU17275PC samples, and quality control samples are spiked with a solution of isotopically labeled internal standards and subjected to protein precipitation with an organic solution (2:1 methanol: dichloromethane). Following centrifugation, an aliquot of the organic supernatant were derivatized using dansyl chloride. The derivatized extract is diluted and injected onto an Agilent 1290 Infinity / SCIEX QTRAP 5500 LC-MS / MS system equipped with an Ascentis Express 90 A Cl 8 column. The mass spectrometer is operated in positive mode using electrospray ionization (ESI). The peak areas of the respective product ions are measured against the peak area of the parent ions of the corresponding internal standard. Quantitation is performed using a weighted linear least squares regression analysis generated from fortified calibration standards prepared immediately prior to each run. Calculated concentrations are based on the analysis of 20.0 pL of plasma.

[0244] LC-MS / MS raw data was collected and processed using SCIEX software Analyst 1.7.3 and processed using SCIEX OS-MQ software v3.1.6. Data reduction was performed using Microsoft Excel for Microsoft 365 MSO.

[0245] Sample analysis was carried out in a 96-well plate format containing two calibration curves and six QC samples (2 replicates per QC level) per plate to monitor assay performance. Three levels of QC (Endogenous, Mx and ClinChek®) were prepared for TAM217. The Mx QC was prepared by reconstituting a lot of lyophilized plasma, with water and aliquoted into 1.5 ml vials. The Endogenous QC was prepared by aliquoting a pooled lot of human plasma procured from BioIVT. Endogenous QC was run at endogenous levels, no analyte spike was performed. The ClinChek® QCs were procured from Iris Technology (“ClinChek® Plasma Control for Amino Acids - Level I”). Accuracy was evaluated for the Endogenous and Mx QCs. Clinchek® QCs were evaluated using precision as no established concentrations were available. A note has been added for any samples with analytes that did not meet acceptance criteria. Analyte concentrations that fell below the limit of quantitation are extrapolated and given a comment of BLOQ. Analyte concentrations that fell above the limit of quantitation are extrapolated and given a comment of ALOQ. Analytes that cannot be extrapolated below the limit of quantitation and are considered not quantifiable and are recorded as N / Q as the concentration. The central storage facility, central laboratory, Metabolon Inc. as well as the statistician were blinded to clinical findings until metabolome and clinical database freezing.

[0246] Example 7: Statistical analysisLudwig-Maximilians-Universitat Munchen,

[0247] in Vertretung des Freistaates Bayern 42 LMU17275PC The study was designed to have >80% power to assess all the primary and secondary analyses to be achieved by enrollment of at least 250 PDAC patients.

[0248] The primary and secondary analyses were based on all available data without imputation for CA19.9 levels. To successfully transfer the technology, targeted metabolite analysis on a randomly selected plasma samples (N = 346) was performed using previously established LC-MS / MS methodology at Metanomics Health GmbH, Berlin, Germany (Mahajan et al. 2022) and quantification was correlated with LC-MS / MS analysis at Metabolon Inc. To reduce interplatform and internal standard associated bias for absolute quantification of metabolites, covariate shift correction was performed utilizing linear regression on matched samples. A linear regression model was constructed for each metabolite, which allowed the determination of the correction slope and intercept. Bias was corrected for all samples, which were then subject to further analysis. The effectiveness correction was assessed using the Concordance Correlation Coefficient (CCC). All metabolite profiling data were loglO-transformed to achieve an approximate normal distribution. The loglO-transformed, auto-scaled and median imputed data was independently evaluated with previously constructed 12 metabolites plus CA19.9 signature (i-Metabolite signature) and 4 metabolites plus CAI 9.9 signature (m-Metabolite signature) (Mahajan et al. 2022). Prediction scores for each patient were evaluated for its potential to differentiate PDAC from controls (IPMNs, non-IPMNs cystic lesions, chronic pancreatitis, acute pancreatitis and metastases of extra-pancreatic origins). The optimal cutoffs for z- and m-Metabolite signature were retained from previous studies (Mahajan et al. 2022). Subsequently, the prediction scores of z- and m-Metabolite signatures were compared to CAI 9.9 alone. The diagnostic performance of z- and m-Metabolite signature were evaluated using 10-fold cross validation. To mimic a population cohort with a 1-20% incidence of PDAC, PDAC patients were randomly sampled 1000 times to achieve the target prevalence (bootstrapping analysis). Briefly, the dataset was stratified by disease status, and CAI 9.9 values in the PDAC group were divided into quartiles for balanced sampling. For each prevalence level, 1000 bootstrap iterations were conducted, with samples drawn from each quartile according to the required prevalence. When quartiles lacked sufficient samples, sampling with replacement was utilized. The resulting cohort was then evaluated for the diagnostic performance of z- and m-Metabolite signatures. To account for the enrichment design, inverse probability weighting (IPW) was used to refine the analysis without requiring model retraining. Sampling weights were determined by comparing the original prevalence of pancreatic ductal adenocarcinoma (PDAC) at roughly 1% to the enriched sample prevalence of 20%. ThisLudwig-Maximilians-Universitat Munchen,

[0249] in Vertretung des Freistaates Bayern 43 LMU17275PC adjustment enables the calculation of model performance metrics that better represent expected outcomes in the target population, while still capitalizing on the statistical efficiency of the enriched design. To assess model performance, 1,000 bootstrap resampling iterations were conducted to estimate confidence intervals for key metrics, including the weighted area under the curve (AUC), ensuring reliable results.

[0250] The DeLong test was employed to compare the AUC of the two signatures. Performance metrics (AUC, sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) and balanced accuracy are presented as means and 95% CI. Planned subgroup analyses of the primary outcome were conducted for the performance of z- and m-Metabolite signatures in resectable PDAC (stage I-IIB), detectable CA19.9 (> 10 U / ml), CA19.9 (< 37 U / ml) and patients with < 10 U / ml CA19.9 status. A subgroup analysis was conducted to evaluate the differential diagnosis performance of the z- and m-Metabolic signatures in distinguishing PDAC from individual controls. Student’s t-test was utilized to assess the differences in predictive scores, while the false discovery rate was adjusted using the Bonferroni correction for multiple comparisons. For the primary outcome, a two-sided P-value of <0.05 was considered to indicate statistical significance. All data processing, modeling and assessment of performances was performed using R (version 4.4.0).

[0251] Example 8 (Comparative Example): Comparison to Mahajan et al. (2022)

[0252] In Mahajn et al. (2022), analysis was performed using a logistic regression model with elastic net regularization (glmnet). The model was trained using 10-fold cross-validation with deviance as the performance measure.

[0253] Model Specifications

[0254] • Family: Binomial (Logistic Regression)

[0255] • Standardization: Yes

[0256] • Cross-validation: 10-fold

[0257] • Performance Measure: Deviance

[0258] Model Performance

[0259] • Degrees of Freedom (Df): 11

[0260] • Percent Deviance Explained: 64.3%

[0261] • Lambda (Regularization Parameter): 0.0329Ludwig-Maximilians-Universitat Munchen,

[0262] in Vertretung des Freistaates Bayern 44 LMU17275PC

[0263] Coefficient Analysis

[0264] The model identified the following significant features with their corresponding coefficients are shown in Table 5.

[0265] Linear Equation

[0266] For the standardized variables, the model of Mahajan et al (2022) can be expressed as:

[0267] log(p / (l-p)) = B1(CA19_9) + B2(Proline) + B3 (Tryptophan) + B4(Histidine) + B5(Lysophosphatidylethanolamine (C18:0)) + B6(Lysophosphatidylethanolamine (C18:2)) + B7(Sphingomyelin (dl7: l,C16:0)) + B8(Sphingomyelin (dl8:2,C17:0)) + B9(Sphingomyelin (35:1)) + B10(Sphingomyelin (41:2)) + Bl 1 (Phosphatidylethanolamine (dl8:0,C22:6)) + B12(Ceramide (dl8:2,C24:0)) + B13(Ceramide (dl 8: l,C24:0)), where p is the probability of the positive class.

[0268] The aforesaid model of Mahajan et al. 2022, however omitting pre-classification according to CAI 9.9 expression, was compared to the models of the present invention; resulting performance parameters are summarized in Table 6. Further, performance parameters for CA19.9 alone are included as well.

[0269] Clearly, the models of the present invention have a better performance; also, contrary to Mahajan et al., 2022, pre-classification based on CA19.9 expression is not required.

[0270] Literature

[0271] Ben-Ami, Roni, Qiao-Li Wang, Jinming Zhang, Julianna G Supplee, Johannes F Fahrmann, Roni Lehmann- Werman, Lauren K Brais, et al. 2023. “Protein Biomarkers and Alternatively Methylated Cell-Free DNA Detect Early Stage Pancreatic Cancer.” Gut, December, gutjnl-2023-331074. doi.org / 10.1136 / gutjnl-2023-331074.

[0272] Boyd, Lenka N. C., Mahsoem Ali, Annalisa Comandatore, Ingrid Garajova, Laura Kam, Jisce R. Puik, Stephanie M. Fraga Rodrigues, et al. 2023. “Prediction Model for Early-Stage Pancreatic Cancer Using Routinely Measured Blood Biomarkers.” JAMA Network Open 6 (8): e2331197. doi.org / 10.1001 / jamanetworkopen.2023.31197.

[0273] Brand, Randall E., Jan Persson, Svein Olav Bratlie, Daniel C. Chung, Bryson W. Katona, Alfredo Carrato, Marien Castillo, et al. 2022. “Detection of Early-Stage Pancreatic Ductal Adenocarcinoma From Blood Samples: Results of a Multiplex BiomarkerLudwig-Maximilians-Universitat Munchen,

[0274] in Vertretung des Freistaates Bayern 45 LMU17275PC Signature Validation Study.” Clinical and Translational Gastroenterology 13 (3): e00468. doi.org / 10.14309 / ctg.0000000000000468.

[0275] Brand et al., Gut. 2007;56:1460-9

[0276] Byeon, Sooin, Matthew J. McKay, Mark P. Molloy, Anthony J. Gill, Jaswinder S. Samra, Anubhav Mittal, and Sumit Sahni. 2024. “Novel Serum Protein Biomarker Panel for Early Diagnosis of Pancreatic Cancer.” International Journal of Cancer 155 (2): 365- 71. doi.org / 10.1002 / ijc.34928.

[0277] Cao, Yingying, Rui Zhao, Kai Guo, Shuai Ren, Yaping Zhang, Zipeng Lu, Lei Tian, Tao Li, Xiao Chen, and Zhongqiu Wang. 2022. “Potential Metabolite Biomarkers for Early Detection of Stage-I Pancreatic Ductal Adenocarcinoma.” Frontiers in Oncology 11 (January):744667. doi.org / 10.3389 / fonc.2021.744667.

[0278] Cohen, Jeremie F., Daniel A. Korevaar, Douglas G. Altman, David E. Bruns, Constantine A.

[0279] Gatsonis, Lotty Hooft, Les Irwig, et al. 2016. “STARD 2015 Guidelines for Reporting Diagnostic Accuracy Studies: Explanation and Elaboration.” BMJ Open 6 (11): e012799. doi.org / 10.1136 / bmj open-2016-012799.

[0280] Collins, Gary S, Johannes B Reitsma, Douglas G Altman, and Karel Moons. 2015.

[0281] “Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD): The TRIPOD Statement.” BMC Medicine 13 (1): 1. doi. org / 10.1186 / s 12916-014-0241 -z.

[0282] Del Chiaro et al., World J Gastroenterol 2014; 20:12118-12131

[0283] Digiacomo, Luca, Erica Quagliarini, Daniela Pozzi, Roberto Coppola, Giulio Caracciolo, and Damiano Caputo. 2023. “Stratifying Risk for Pancreatic Cancer by Multiplexed Blood Test.” Cancers 15 (11): 2983. doi.org / 10.3390 / cancersl5112983.

[0284] Fahrmann, Johannes F., C. Max Schmidt, Xiangying Mao, Ehsan Irajizad, Maureen Loftus, Jinming Zhang, Nikul Patel, et al. 2021. “Lead-Time Trajectory of CA19-9 as an Anchor Marker for Pancreatic Cancer Early Detection.” Gastroenterology 160 (4): 1373-1383. e6. doi.org / 10.1053 / j.gastro.2020.11.052.

[0285] Firpo, Matthew A., Kenneth M. Boucher, Josh Bleicher, Gayatri D. Khanderao, Alessandra Rosati, Katherine E. Poruk, Sama Kamal, et al. 2023. “Multianalyte Serum Biomarker Panel for Early Detection of Pancreatic Adenocarcinoma.” JCO Clinical Cancer Informatics, no. 7 (March), e2200160. doi.org / 10.1200 / CCI.22.00160.

[0286] Haab, Brian, Lu Qian, Ben Staal, Maneesh Jain, Johannes Fahrmann, Christine Worthington, Denise Prosser, et al. 2024. “A Rigorous Multi-Laboratory Study of Known PDAC Biomarkers Identifies Increased Sensitivity and Specificity over CAI 9-9 Alone.” Cancer Letters 604 (November):217245. doi.org / 10.1016 / j.canlet.2024.217245. Iwano, Tomohiko, Kentaro Yoshimura, Genki Watanabe, Ryo Saito, Sho Kiritani, Hiromichi Kawaida, Takeshi Moriguchi, et al. 2021. “High-Performance Collective Biomarker from Liquid Biopsy for Diagnosis of Pancreatic Cancer Based on Mass Spectrometry and Machine Learning.” Journal of Cancer 12 (24): 7477-87.

[0287] doi. org / 10.7150 / jca.63244.

[0288] Katona, Bryson W ., Christine Worthington, Daniel Clay, Hannah Cincotta, Nuzhat A.

[0289] Ahmad, Gregory G. Ginsberg, Michael L. Kochman, and Randall E. Brand. 2023. “Outcomes of the IMMray PanCan-d Test in High-Risk Individuals Undergoing Pancreatic Surveillance: Pragmatic Data and Lessons Learned.” JCO Precision Oncology, no. 7 (September), e2300445. doi.org / 10.1200 / PO.23.00445.

[0290] Kim, Yoseop, Injoon Yeo, Iksoo Huh, Jaenyeon Kim, Dohyun Han, Jin-Young Jang, and Youngsoo Kim. 2021. “Development and Multiple Validation of the Protein MultiMarker Panel for Diagnosis of Pancreatic Cancer.” Clinical Cancer Research 27 (8): 2236-45. doi.org / 10.1158 / 1078-0432. CCR-20-3929.

[0291] Kim et al., Sci Rep 2023; 13(1): 106

[0292] Llach et al., Cancer Manag Res 2020;12:743-58Ludwig-Maximilians-Universitat Munchen,

[0293] in Vertretung des Freistaates Bayern 46 LMU17275PC Li, Tiandong, Junfen Xia, Huan Yun, Guiying Sun, Yajing Shen, Peng Wang, Jianxiang Shi, Keyan Wang, Hongwei Yang, and Hua Ye. 2023. “A Novel Autoantibody Signatures for Enhanced Clinical Diagnosis of Pancreatic Ductal Adenocarcinoma.” Cancer Cell International 23 (1): 273. doi.org / 10.1186 / sl2935-023-03107-l.

[0294] Mahajan, Ujjwal M., Bettina Oehrle, Simon Sirtl, Ahmed Alnatsha, Elisabetta Goni, Ivonne Regel, Georg Beyer, et al. 2022. “Independent Validation and Assay Standardization of Improved Metabolic Biomarker Signature to Differentiate Pancreatic Ductal Adenocarcinoma From Chronic Pancreatitis.” Gastroenterology 163 (5): 1407-22. doi.org / 10.1053 / j.gastro.2022.07.047.

[0295] Mayerle, Julia, Holger Kalthoff, Regina Reszka, Beate Kamlage, Erik Peter, Bodo

[0296] Schni ewind, Sandra Gonzalez Maldonado, et al. 2018. “Metabolic Biomarker Signature to Differentiate Pancreatic Ductal Adenocarcinoma from Chronic Pancreatitis.” Gut 67 (1): 128-37. doi.org / 10.1136 / gutjnl-2016-312432.

[0297] Nakamura, Kota, Zhongxu Zhu, Souvick Roy, Eunsung Jun, Haiyong Han, Ruben M. Munoz, Satoshi Nishiwada, et al. 2022. “An Exosome-Based Transcriptomic Signature for Noninvasive, Early Detection of Patients With Pancreatic Ductal Adenocarcinoma: A Multicenter Cohort Study.” Gastroenterology 163 (5): 1252-1266. e2. doi.org / 10.1053 / j.gastro.2022.06.090.

[0298] Pepe et al., J Natl Cancer Inst 2008; 100(20): 1432-8

[0299] Sahni, Sumit, AdvaitR. Pandya, William J. Hadden, Christopher B. Nahm, Sarah Maloney, Victoria Cook, James A. Toft, et al. 2021. “A Unique Urinary Metabolomic Signature for the Detection of Pancreatic Ductal Adenocarcinoma.” International Journal of Cancer 148 (6): 1508-18. doi.org / 10.1002 / ijc.33368.

[0300] Seeger, Nico, Stefan Gutknecht, Irin Zschokke, Isabella Fleischmann, Nadja Roth, Jiirg Metzger, Markus Weber, Stefan Breitenstein, and Lukasz Filip Grochola. 2024. “A Predictive Noninvasive Single-Nucleotide Variation-Based Biomarker Signature for Resectable Pancreatic Cancer: Protocol for a Prospective Validation Study.” JMIR Research Protocols 13 (May):e54042. doi.org / 10.2196 / 54042.

[0301] Shen, Xiaojing, Xiaolin Zhu, Hairong Liu, Rongtao Yuan, Qingyuan Guo, and Peng Zhao.

[0302] 2024. “Leveraging Genomic Signatures of Oral Microbiome- Associated Antibiotic Resistance Genes for Diagnosing Pancreatic Cancer.” Edited by Yash Gupta. PROS ONE 19 (4): e0302361. doi.org / 10.1371 / joumal. pone.0302361.

[0303] Singhi et al., Gastroenterology 2019;156(7):2024-40

[0304] Song, Jin, Lori J. Sokoll, Daniel W. Chan, and Zhen Zhang. 2021. “Validation of Serum Biomarkers That Complement CAI 9-9 in Detecting Early Pancreatic Cancer Using Electrochemiluminescent-Based Multiplex Immunoassays.” Biomedicines 9 (12): 1897. doi.org / 10.3390 / biomedicines9121897.

[0305] Srivastava, Sudhir, and Paul D. Wagner. 2020. “The Early Detection Research Network: A National Infrastructure to Support the Discovery, Development, and Validation of Cancer Biomarkers.” Cancer Epidemiology, Biomarkers & Prevention: A Publication of the American Association for Cancer Research, Cosponsored by the American Society of Preventive Oncology 29 (12): 2401-10. doi.org / 10.1158 / 1055-9965.EPI-20- 0237.

[0306] Wen, Yan-Rong, Xia-Wen Lin, Yu-Wen Zhou, Lei Xu, Jun-Li Zhang, Cui- Ying Chen, and Jian He. 2024. “N-Glycan Biosignatures as a Potential Diagnostic Biomarker for Early-Stage Pancreatic Cancer.” World Journal of Gastrointestinal Oncology 16 (3): 659-69. doi.org / 10.4251 / wjgo.vl6.i3.659.

[0307] Wu, Huanwen, Shiwei Guo, Xiaoding Liu, Yatong Li, Zhixi Su, Qiye He, Xiaoqian Liu, et al.

[0308] 2022. “Noninvasive Detection of Pancreatic Ductal Adenocarcinoma Using the Methylation Signature of Circulating Tumour DNA.” BMC Medicine 20 (1): 458. doi.org / 10.1186 / s 12916-022-02647-z.Ludwig-Maximilians-Universitat Munchen,

[0309] in Vertretung des Freistaates Bayern 47 LMU17275PC Xu, Caiming, Eunsung Jun, Yoshinaga Okugawa, Yuji Toiyama, Erkut Borazanci, John Bolton, Akinobu Taketomi, et al. 2024. “A Circulating Panel of circRNA Biomarkers for the Noninvasive and Early Detection of Pancreatic Ductal Adenocarcinoma.” Gastroenterology 166 (1): 178-190. el6. doi.org / 10.1053 / j.gastro.2023.09.050.

[0310] Yu, Jingru, Alexander Ploner, Maximilian Kordes, Matthias Lohr, Magnus Nilsson, Maria Evangelina Lopez De Maturana, Lidia Estudillo, et al. 2021. “Plasma Protein Biomarkers for Early Detection of Pancreatic Ductal Adenocarcinoma.” International Journal of Cancer 148 (8): 2048-58. doi.org / 10.1002 / ijc.33464.

[0311] Zweig 1993, Clin. Chem. 39:561-577udwig-Maximilians-Universitat Munchen,

[0312] n Vertretung des Freistaates Bayern 48 LMU17275PC able 2: Preferred coefficient values for formula (I) when using the m-metabolic signature

[0313]

[0314] able 3 : Preferred coefficient values for formula (I) when using the i-metabolic signature

[0315]

[0316] udwig-Maximilians-Universitat Munchen,

[0317] n Vertretung des Freistaates Bayern 49 LMU17275PC

[0318]

[0319] able 4a: Training data fortraining the m-metabolic signature; NO: sample number; PDAC: pancreatic ductal adenocarcinoma; CP: chronic ancreatitis; LPE: Lysophosphatidylethanolamine; SM: Sphingomyelin; PE: Phosphatidylethanolamine; Cer: Ceramide, re.: resectable.

[0320]

[0321] udwig-Maximilians-Universitat Munchen,

[0322] n Vertretung des Freistaates Bayern 50 LMU17275PC

[0323]

[0324] udwig-Maximilians-Universitat Munchen,

[0325] n Vertretung des Freistaates Bayern 51 LMU17275PC

[0326]

[0327] udwig-Maximilians-Universitat Munchen,

[0328] n Vertretung des Freistaates Bayern 52 LMU17275PC

[0329]

[0330] udwig-Maximilians-Universitat Munchen,

[0331] n Vertretung des Freistaates Bayern 53 LMU17275PC

[0332]

[0333] udwig-Maximilians-Universitat Munchen,

[0334] n Vertretung des Freistaates Bayern 54 LMU17275PC

[0335]

[0336] able 4b: Additional training data fortraining the i-metabolic signature; NO: sample number; Pro: proline; Trp: tryptophan; His: histidine; LPE: ysophosphatidylethanolamine; SM: Sphingomyelin; PE: Phosphatidylethanolamine; Cer: Ceramide.

[0337]

[0338] udwig-Maximilians-Universitat Munchen,

[0339] n Vertretung des Freistaates Bayern 55 LMU17275PC

[0340]

[0341] udwig-Maximilians-Universitat Munchen,

[0342] n Vertretung des Freistaates Bayern 56 LMU17275PC

[0343]

[0344] udwig-Maximilians-Universitat Munchen,

[0345] n Vertretung des Freistaates Bayern 57 LMU17275PC

[0346]

[0347] udwig-Maximilians-Universitat Munchen,

[0348] n Vertretung des Freistaates Bayern 58 LMU17275PC

[0349]

[0350] udwig-Maximilians-Universitat Munchen,

[0351] n Vertretung des Freistaates Bayern 59 LMU17275PC

[0352]

[0353] udwig-Maximilians-Universitat Munchen,

[0354] n Vertretung des Freistaates Bayern 60 LMU17275PC

[0355]

[0356] udwig-Maximilians-Universitat Munchen,

[0357] n Vertretung des Freistaates Bayern 61 LMU17275PC

[0358]

[0359] able 5: Coefficient values for the i-metabolic signature in th emodel according to Mahajan et al. 2022; two biomarkers (Sphingomyelin (41:2) and eramide (d 18 : l,C24:0)) were excluded from the model due to regularization.

[0360]

[0361] udwig-Maximilians-Universitat Munchen,

[0362] n Vertretung des Freistaates Bayern 62 LMU17275PC able 6: Performance marker comparison

[0363] N = (PDAC / Total) Model AUC Sensitivity Specificity Balanced accuracy

[0364]

[0365] i-Metabolic signature 0.846 Z = 21.417 0.675 0.904 0.805

[0366] (present invention) (0.842 - 0.849) P <0.001 (0.669 - 0.680) (0.898 - 0.911) (0.802- 0.808) All Stages m-Metabolic signature 0.846 Z = 21.417 0.599 0.936 0.790

[0367] (N = 489) (present invention) (0.842 - 0.849) P <0.001 (0.593 - 0.604) (0.931 - 0.940) (0.788 - 0.792)

[0368] CA19.9 0.799 - 0.818 0.791 0.806

[0369] Alone . (0.797 - 0.802) . (0.815 - 0.820) (0.787 - 0.794) (0.804 - 0.809) i-Metabolic signature 0.481 Z = -160.000 0.561 0.428 0.498

[0370]

[0371] <

Claims

1. Ludwig-Maximilians-Universitat Munchen,in Vertretung des Freistaates Bayern 63 LMU17275PC Claims1. A method for assessing pancreatic cancer in a subject, said method comprising(a) determining the biomarkers Ceramide (dl8:l;C24:0), Lysophosphatidylethanolamine (C18:0), Phosphatidylethanolamine (C18:0, C22:6), Sphingomyelin (dl7: 1; C16:0), and CA19.9 in a sample from said subject;(b) comparing the biomarkers determined in step (a) to at least one corresponding reference; and(c) assessing pancreatic cancer in said subject based on said comparing in step (b).

2. The method of claim 1, wherein assessing comprises excluding pancreatic cancer, preferably identifying a subject not suffering from pancreatic cancer; diagnosing pancreatic cancer, preferably diagnosing PDAC; staging pancreatic cancer; prognosticating pancreatic cancer; and / or differentiating between pancreatic cancer and a non-pancreatic cancer disorder, preferably pancreatitis, more preferably chronic pancreatitis.

3. The method of claim 1 or 2, wherein said subject is a subject at risk of suffering from pancreatic cancer; preferably is known or suspected to suffer from new onset type 2 diabetes, chronic pancreatitis, pancreatic cystic lesions, and / or is from a familial PDAC kinship.

4. The method of any one of claims 1 to 3, wherein said method comprises further determining at least one, preferably at least two, more preferably at least three, even more preferably at least four, most preferably all, marker(s) selected from the list consisting of Histidine, Proline, Tryptophan, Ceramide (dl8:2,C24:0), Lysophosphatidylethanolamine (C18:2), Sphingomyelin (35:1), Sphingomyelin (41:2), and Sphingomyelin (dl8:2,C17:0), preferably in step (a).

5. The method of any one of claims 1 to 4, wherein said sample is a bodily fluid sample, preferably a blood sample or a blood-derived sample, more preferably a blood, plasma, or serum sample, even more preferably a plasma or serum sample.Ludwig-Maximilians-Universitat Munchen,in Vertretung des Freistaates Bayern 64 LMU17275PC 6. The method of any one of claims 1 to 5, wherein step (b) comprises calculating a predictor score for assessing pancreatic cancer from all of said biomarkers and wherein said predictor score is calculated according to formula (I)1P~ 1 + e~zwherein P= predictor score; and z = B0+B1X1+B2X2+ — 1-BnXnand wherein variables Xi, X2, ... , Xnare measurement values of said biomarkers.

7. The method of claim 6, wherein coefficients Bo, Bi, B2, .. Bnare selected from Table 1 or from Table 2.

8. The method of any one of claims 1 to 7, wherein said method comprises determining at least one further biomarker, preferably selected from the list consisting of aspartate aminotransferase, alanine aminotransferase, platelet count, haptoglobin, alpha2- macroglobulin, apolipoprotein Al, bilirubin, cholesterol, hyaluronan, prothrombin index, hepatocyte growth factor (HGF), Tissue inhibitor of metalloproteinases (TIMP), and urea; and / or wherein said method comprises further diagnostic steps, preferably sonography, magnetic resonance imaging, radiography, transient elastography, and / or determining subject age and / or gender.

9. The method of any one of claims 1 to 8, wherein said method is computer- implemented, preferably wherein said method is computer implemented and wherein said comparing in step (b) is performed by a trained automated machine learning derived generalized logistic regression model, more preferably wherein said logistic regression model was trained as specified in any one of claims 10 to 13.

10. A computer-implemented training method of training at least one trainable model for assessing pancreas cancer in a subject, the method comprising:(i) providing the trainable model;(ii) retrieving labeled training subject data comprising biomarker data of subjects having known pancreas cancer states, wherein said biomarker data are data of the biomarkers Ceramide (dl8:l;C24:0), Lysophosphatidylethanolamine (C18:0), Phosphatidylethanolamine (Cl 8:0, C22:6), Sphingomyelin (dl7: 1; Cl 6:0), and CA19.9; and(iii) training the trainable model on the labeled training subject data,Ludwig-Maximilians-Universitat Munchen,in Vertretung des Freistaates Bayern 65 LMU17275PC wherein said trainable model is a Generalized Linear Model (GLM).

11. The method of claim 10, wherein said assessing is classifying pancreas cancer.

12. The method of claim 10 or 11, wherein the trainable model is a binomial GLM.

13. The method of any one of claims 10 to 12, wherein the trainable GLM model is trained using the following hyperparameters:Regularization: Ridge;Logistic regression link: logit;Number of iterations: 40,Nlambad—30,Lambda.max = 32.331Sort_metric = “Fl”, and / orStopping tolerance = 3.

14. An automated machine learning model obtained or obtainable according to the method according to any one of claims 10 to 13, preferably tangibly embedded on a data storage means.

15. A system comprising(i) a measuring device for determining the biomarkers Ceramide (dl8:l;C24:0), Lysophosphatidylethanolamine (C18:0), Phosphatidylethanolamine (C18:0, C22:6), Sphingomyelin (dl7: 1 ; C16:0), and CA19.9; and(ii) an evaluation device operably linked to the measuring device, said evaluation device comprising a data processor comprising instructions for carrying out a comparison of the biomarkers determined by the measuring device to at least one corresponding reference, wherein said system is adapted to perform the method according to any one of claims 1 to 9, preferably comprising tangibly embedded instructions which, when carried out by the data processor, cause the system to perform a method according to any one of claims 1 to 9.