Determination of a cancer risk score based on NMR spectroscopy data of a biofluid sample

The method employs 1H-NMR spectroscopy and machine learning to analyze biofluid samples, addressing the limitations of current cancer detection methods by enabling early and accurate differentiation between healthy and cancerous states, and providing a cancer risk score.

WO2025133230A1PCT designated stage expired Publication Date: 2025-06-26STRASSER PATRICK +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/088074
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-20
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Current AI- and ML-based cancer detection methods are limited in detecting specific types of cancer and often fail to detect cancerous diseases or pre-cancerous conditions early, as they focus on single signals and struggle to differentiate between malignant and benign conditions.

Method used

A computer-implemented method using 1H-NMR spectroscopy data from biofluid samples, combined with machine learning models, to evaluate molecular profiles and classify them into healthy or cancerous/pre-cancerous categories, thereby determining a cancer risk score.

Benefits of technology

This method enables early detection of cancerous and pre-cancerous diseases by analyzing the entire molecular profile of biofluid samples, differentiating between malignant and benign conditions, and providing a cancer risk score indicative of the probability of cancer occurrence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024088074_26062025_PF_FP_ABST
    Figure EP2024088074_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented method of determining a cancer risk score for an individual is described. The method comprises obtaining, at a computing device (100), Hydrogen-1 Nuclear Magnetic Resonance, 1H-NMR, spectroscopy data (300a-c) for a biofluid sample of an individual, the 1H-NMR spectroscopy data (300a-c) being indicative of an NMR signal intensity as a function of the chemical shift. The method further comprises evaluating, with at least one trained machine learning model (114) of the computing device (100), the 1H-NMR spectroscopy data (300a-c) in terms of a molecular profile (310a-c) of the biofluid sample, the molecular profile containing all hydrogen peaks in the 1H-NMR spectroscopy data associable to one or more of a metabolite, a protein, an amino acid, a micro molecule and a macromolecule contained in the biofluid sample. The method further comprises classifying, based on the evaluating with the at least one trained machine learning model (114), the molecular profile (310a-c) into at least a first class and a second class of molecular profiles, the first class being representative of molecular profiles (310a) associated with healthy individuals and the second class being representative of molecular profiles (310b, c) associated with individuals having a cancerous and / or pre-cancerous disease. Further, the method comprises determining, based on the classifying of the molecular profile, a cancer risk score indicative of a probability for cancer occurring at the individual.
Need to check novelty before this filing date? Find Prior Art

Description

DETERMINATION OF A CANCER RISK SCORE BASED ON NMR SPECTROSCOPY DATA OF A BIOFLUID SAMPLETECHNICAL FIELD

[0001] The present invention generally relates to the field of computer-aided oncology. More specifically, the present invention relates to a computer-implemented method of determining a cancer risk score for an individual. The present invention further relates to a computing device configured to perform said method, to a computer program instructing the computing device to perform said method, and to a computer-readable medium storing such computer program.BACKGROUND

[0002] Cancerous diseases, or generally cancer, affect millions of individuals or humans, and are a leading cause of death worldwide. The most common or widely spread cancer types are breast, lung, colon, and rectum / prostate cancers, but nearly any organ of the human body can be affected by cancer, such as liver, pancreas, kidneys, and others. As a general rule, cancer mortality can be reduced when detected and treated early. When detected early, cancer is more likely to respond to treatment and less likely to spread to other tissue, which can result in a greater probability of survival with less morbidity, as well as less expensive treatment.

[0003] In many cases, cancer, cancerous diseases or pre-cancerous diseases can be detected or diagnosed based on laboratory test results of a biofluid sample of an individual or patient, such as blood, urine, and Cerebrospinal Fluid (CSF). Currently established and widely used analysis methods of biofluid samples include, for example, blood chemistry tests that measure the amounts of certain substances in a blood sample, and complete blood count tests that measure the number of red blood cells, white blood cells, and platelets in a blood sample. Other cancer detection or analysis methods include tumor marker tests that measure substances that are produced by cancer cells or other cells of the body in response to cancer, and urinalysis that describes the color and content of a urine sample. Other approaches use liquid biopsy in combination with next-generation sequencing technologies for cancer detection.

[0004] Over the past years, also many developments have been made towards computer- implemented or automated cancer detection or assessment methods, for example to allow for a data-driven and / or individualized patient care based on laboratory test results of a biofluid sample of an individual, as described above. Generally, Artificial intelligence (Al) and Machine learning (ML), a subset of Al that enables computing devices to learn from training data, have emerged as useful tools in oncology and have shown to improve healthcare accuracy and patient outcomes.

[0005] However, the currently used Al- and ML-based methods for detecting cancer based on measurements of biofluid samples typically focus on specific substances present in a biofluid sample and are therefore usually limited to detecting particular types of cancer. Also, at least for some of the currently used cancer detection methods, an early detection of cancerous diseases or even pre-cancerous diseases is hardly possible, for example as the corresponding cancer signal may not be detectable at early stage of the cancer nor can it differentiate malignant from benign conditions.SUMMARY

[0006] It may, therefore, be desirable to provide for an improved computer-implemented method and corresponding computing device for determining a cancer risk score indicative of a probability for cancer occurring at an individual. The method and computing device described herein may particularly allow for or enable an early detection of cancerous disease and / or pre-cancerous disease and the differentiation of malignant from non-malignant or benign tumors.

[0007] This is achieved by the subject matter of the independent claims, wherein further embodiments are incorporated in the dependent claims and the following description.

[0008] Aspects of the present disclosure relate to a computer-implemented method of determining a cancer risk score for an individual, to a computing device configured to perform a method of determining a cancer risk score for an individual, to a corresponding computer program, and to a computer-readable medium storing such computer program. Any disclosure presented herein with reference to one or an aspect of the present disclosure equally applies to any other aspect of the present disclosure.

[0009] According to an aspect of the present disclosure, there is provided a computer-implemented method of determining, computing and / or calculating a cancer risk score for an individual. The method comprises: obtaining, at a computing device, Hydrogen-1 Nuclear Magnetic Resonance, 1 H-NMR, spectroscopy data for a biofluid sample of an individual, the 1 H-NMR spectroscopy data being indicative of an NMR signal intensity as a function of the chemical shift; evaluating, with at least one trained machine learning model of the computing device, the 1 H-NMR spectroscopy data in terms of a molecular profile of the biofluid sample, the molecular profile containing all hydrogen peaks in the 1 H-NMR spectroscopy data associable to and / or associated with one or more of a metabolite, a protein, an amino acid, a micro molecule and a macromolecule contained in the biofluid sample; classifying, based on the evaluating with the at least one trained machine learning model, the molecular profile into at least a first class and a second class of molecular profiles, the first class being representative of molecular profiles associated with healthy individuals and the second class being representative of molecular profiles associated with individuals having a cancerous and / or pre-cancerous disease; and determining, based on the classifying of the molecular profile, a cancer risk score indicative of a probability for cancer occurring at the individual.

[0010] Cancer can influence the overall molecular composition of one or more biofluids of an individual, such as for example the composition of blood, urine, cerebrospinal fluid (CSF) and other biofluids. This influence can include changes in the proportion of specific proteins, metabolites, micro- and / or macromolecules that are either directly released by the cancer cells or whose release is directly or indirectly induced by them.

[0011] The method described herein utilizes highly detailed 1 H-NMR spectroscopy data, also referred to as high-frequency NMR spectroscopy data or proton NMR spectroscopy data, obtained from a biofluid sample of a patient or individual in combination with machine learning (ML) to analyze and / or evaluate the molecular profile of the biofluid sample, for example to detect the aforementioned changes induced by cancer in the biofluid sample. The 1 H-NMR spectroscopy data and / or the molecular profile contained therein can provide information about the chemical environment of each hydrogen (1 H) atom in the biofluid sample. This means that depending on the structure of the molecule that contains the hydrogen atom, a different chemical environment can be present, thus changing the chemical shift of this specific hydrogen atom. For these reasons, high- frequency NMR or 1 H-NMR spectroscopy data of a biofluid sample can provide extremely high information density that can be used to differentiate between healthy and diseased patients having a benign, cancerous and / or pre-cancerous disease.

[0012] Generally, precancerous diseases can refer to health conditions or lesions that have an increased risk of developing into cancer but are not yet classified as cancer. The accompanying cell changes are often abnormal but not entirely uncontrolled, as is the case with cancer. An example of a precancerous condition or disease is dysplasia in the cervix, known as cervical intraepithelial neoplasia (CIN), which can lead to cervical cancer. An example of a cancerous disease is invasive cervical cancer, where cancer cells can divide uncontrollably and infiltrate surrounding tissue.

[0013] Accordingly, in the context of the present disclosure, precancerous diseases can refer to diseases that are either direct precursors of cancer, such as e.g. IPMN, a tumor (growth) of the pancreas that often later becomes cancer, or diseases that increase the risk of cancer, such as e.g. chronic pancreatitis.

[0014] While currently used cancer detection methods are usually limited to single signals that can be assigned to a particular tumor marker or specific signals in the respective measurement, such as a specific molecule, metabolite or lipoprotein, the method according to the present disclosure considers all signatures in the 1 H-NMR spectroscopy data. Thus, the entire molecular profile of the individual’s biofluid sample can be taken into account, analysed and evaluated by means of the method disclosed herein, which allows to reliably discern between healthy, precancerous, benign tumor and cancer patients or individuals.

[0015] Also, due to the high informational content of the 1 H-NMR spectroscopy data and since cancer can influence the molecular profile and metabolism of the individual at very early stage, various different changes in the molecular profile induced by or related to cancer can be detected at early stage of a cancerous disease or even pre-cancerous disease, thereby enabling early detection and treatment.

[0016] In particular, the method described herein may be configured to identify cancers of various origins, specifically by detecting a broad range of molecular signals or profiles that universally characterize malignant transformation. Unlike traditional diagnostic methods that typically focus heavily on genetic mutations or rely on the presence of cell-derived vesicles, the method or generally the approach described in this disclosure may be able and may be configured to detect molecular features on a (more) functional level such es e.g. in proteins, metabolites, lipids, lipid particles, andtheir molecular modifications. By capturing the subtle, yet consistently disrupted profiles of not only metabolites and lipids, but also or alternatively post-translational modifications, PTMs, like but not limited to any one of methylation, phosphorylation, glycosylation (glycan patterns), lipidation, acetylation or other modifications within the tumor’s biochemical environment, the method of this disclosure may transcend organ- or mutation-specific markers to detect the fundamental shifts that occur in virtually all cancers. This allows the method to not only identify tumors with a strong hypermetabolic effect but also cancers in general, since certain modifications, traits, and functionalities are necessary for a cell or group of cells to become malignant.

[0017] There may be different origins and / or types of molecular signals or profiles that may be used in the method. One of such origins and / or types may be altered glycosylation and / or other post- translational modifications (PTMs). Glycans and PTMs matter because essentially cancer progression disrupts the cellular machinery responsible for modifying proteins and lipids. Cancerous cells and their surrounding stroma for example change the enzymatic machinery responsible for attaching sugar chains (glycans) to proteins and lipids, altering how cells communicate, adhere, and interact with their environment. Similarly, methylation, phosphorylation, lipidation, acetylation, and other PTMs shift in predictable ways as cells transition from normal to malignant states.

[0018] A universal application of the method to all cancers or several cancers is thereby possible: Unlike markers tied to a single cancer type, PTM alterations are a common language of malignancy. They emerge regardless of the tumor’s location or histology, providing a universal signature that the method can detect, making it more versatile, and sensitive for malignant changes in the patient's serum.

[0019] The method may also employ broad-spectrum metabolic and lipid dysregulation, specifically but not only metabolic shifts: Malignant cells fundamentally rewire their metabolism to support continuous growth and survival, producing unique metabolite profiles. This altered metabolic landscape is not dependent on specific genetic mutations; it is a universal trait of cancer cells that can be picked up by analyzing the abundance and structure of key metabolites. Lipid remodeling is an alternative or additional option in the method: Beyond simple metabolite changes, cancers modulate lipid composition and the structure of lipid particles. Changes in lipid saturation, oxidation, and the distribution of lipoprotein-like particles form a distinct biochemical fingerprint that signals malignancy across various cancer types.

[0020] The method may further employ functional, chemical-level insights without reliance on genetic data. Specifically, one may ask why to focus on proteins, metabolites, and lipids as suggested herein? The answer is that traditional tests often focus on genetic abnormalities or single tumor markers like proteins or sugar residues. The herein disclosed method can transcend these limitations by focusing on the downstream consequences of malignancy - how proteins are modified, how metabolites are shifted, and how lipid structures change. The method may in particular use stable and universal indicators: Since PTMs and metabolite / lipid signatures reflect the cell’s functional state rather than a single mutation, they are more robust and broadly applicable. This functional approach allows for more accurate early detection and classification of a wide range of tumors.

[0021] The method of this disclosure is fundamentally new. On the one hand it is a shift from genecentric to chemistry-centric diagnostics: Historically, many diagnostic approaches centered on identifying a handful of genetic mutations or on capturing certain cell-secreted proteins or vesicles. The present method discards this narrow viewpoint. Instead, it can focus on the emergent chemical properties of proteins, metabolites, lipids and their sugar moieties and other modifications - properties that reliably change as normal cells become cancerous, regardless of their genetic background. On the other hand, the method can employ integration of previously overlooked universal markers: Prior methods rarely leveraged altered glycosylation or non-genetic methylation patterns as broad-spectrum cancer identifiers. These modifications were considered secondary or specialized. By placing these signals at the forefront of the analysis of the method of this disclosure, a new dimension of diagnostic information that has not been harnessed in a comprehensive, universal manner before is unlocked. Moreover, the method can employ a non-lnvasive, versatile application: Because these molecular patterns can be detected in easily accessible samples (e.g., bodily fluids or surface-accessible tissues) and do not depend on analyzing vesicles or extracting genetic material, the approach of this disclosure simplifies screening. It can be applied broadly and repeatedly, making early detection of a wide range of cancers more practical and cost-effective. Also, the method can use advanced data processing and machine learning integration: The conceptual leap from single-marker tests to analyzing global alterations in glycan structures, PTMs, and complex metabolite profiles was not feasible with older technologies. Recent advances in analytical chemistry, nuclear magnetic resonance spectrometry, and machine learning have enabled the herein disclosed integration and interpretation of these complex signals into a robust, reliable diagnostic signature.

[0022] In summary, the present method can identify and interpret a wide spectrum of non-genomic chemical modifications — abnormal glycan patterns, PTMs including methylation on proteins and metabolites, and distinctive metabolic rearrangements — that are universally linked to cancerous transformation. By building on these universal biochemical shifts rather than genetic or vesiclebased markers, a truly new diagnostic paradigm is provided that can detect cancers of different origins without relying on traditional, narrowly focused methods and which was not available before.

[0023] Moreover, the method described herein is not limited to a particular type of cancer or cancerous disease, but can be used to advantage to detect different types of cancer and / or even discern between different types of cancer, cancerous diseases and / or pre-cancerous diseases and / or benign diseases.

[0024] Furthermore, the method described herein can not only differentiate between healthy and diseased patients (patients with precancerous diseases or cancerous diseases) but also between malignant and non-malignant diseases (benign tumors). In particular, the method may be further comprising determining, in particular but not only based on the evaluating of the 1 H-NMR spectroscopy data in terms of a molecular profile of the biofluid sample and / or based on the classifying of the molecular profile, whether the cancer risk score indicative of the probability for cancer occurring at the individual is associated with a malignant disease or a non-malignant disease. Specifically, the method may comprise determining a malignant risk score indicative of a probabilityof the cancer occurring at the individual according to the probability of the cancer risk score being either malignant or non-malignant. Specifically, the determining of such malignant risk score may be based on molecular differences as provided by their molecular profiles and as described in the following. Generally, this is currently a very challenging task in the clinic since benign tumors and malignant tumors are similar in its molecular functionality, however the benign tumors stay within their borders (for example within the epithelia) and are mostly not deadly.

[0025] For benign tumor differentiation, the main difference between benign and malignant tumors is that benign tumors are noncancerous and don't spread, while malignant tumors are cancerous and can spread to other parts of the body. Benign tumors are usually not life threatening and don't spread to other parts of the body. They tend to grow slowly, have distinct borders, and are unlikely to recur once removed. Examples include fibroids in the uterus, lipomas in the skin, and moles on the skin. Malignant tumors are cancerous and can spread to other parts of the body through the bloodstream or lymphatic system. This process is called metastasis and can occur anywhere in the body, but is most common in the liver, lungs, brain, and bone. Malignant tumors grow uncontrollably and require treatment to prevent further spread.

[0026] The method of this disclosure is capable of differentiating malignant cancers from benign tumors by looking into particular molecular markers that are secreted, absorbed or changed by the cancer cells themselves or by particular cells in the tumor-microenvironment, the cancerous niche that is necessary to provoke tumor growth, allow the cancer to spread and allow cancer cells to metastasize. For example, the method may comprise detection of molecular changes that are induced by the different cell types in the tumor-microenvironment, for example by macrophages, stromal cells, like fibroblasts, or other immune cells, like dendritic cells, T cells, B cells, and others.

[0027] Below is an integrated summary highlighting the key molecular, cellular, and tumor- microenvironmental differences between benign and malignant tumors. This encapsulation emphasizes how specific molecular features, immune cell dynamics, and the tumor microenvironment distinguish malignant tumors from their benign counterparts.

[0028] For molecular differences, there are for once genetic alterations: Benign tumors generally harbor fewer genetic mutations. Their cells closely resemble their tissue of origin (well-differentiated) and show limited molecular deviations from normal cellular physiology. Malignant tumors exhibit extensive genetic and epigenetic alterations, such as mutation or loss of tumor suppressor genes, overexpression or activation of oncogenes, and widespread chromosomal instability. These alterations allow for unchecked proliferation, evasion of growth inhibition, and resistance to apoptosis.

[0029] Within the molecular differences, there are also differences in cellular metabolism and molecular signaling: Benign tumors tend to maintain relatively normal metabolic profiles with limited metabolic reprogramming. They usually do not induce significant systemic metabolic disturbances. Malignant tumors undergo marked metabolic reprogramming to support rapid growth, including increased glycolysis (Warburg effect) and altered amino acid (e.g., tryptophan) metabolism. They secrete and respond to a variety of growth factors and proteases that support local invasion and can produce signals that shape a pro-tumorigenic microenvironment.

[0030] Also, there are molecular markers and secreted factors: Benign tumors typically lack the robust molecular signatures associated with aggressive behavior. They produce fewer angiogenic factors and immunomodulatory molecules. Malignant tumors often express or secrete distinctive molecular markers, including elevated levels of certain metabolites, growth factors, cytokines, and checkpoint ligands (e.g., PD-L1) that promote angiogenesis, tissue invasion, immune evasion, and potential metastatic spread.

[0031] Further, there are immune cell differences, e.g., immune surveillance and cytotoxicity: In benign tumors, immune cells within or around benign lesions generally retain a more balanced, nonimmunosuppressive profile. Cytotoxic T lymphocytes and natural killer cells, if present, can recognize abnormal cells without being functionally inhibited. For malignant tumors, cancer cells actively modulate the immune landscape. They often recruit and reprogram immune cells into immunosuppressive phenotypes, such as regulatory T cells (Tregs) and M2-polarized macrophages. By expressing immune checkpoint molecules and inhibitory cytokines, malignant tumors suppress effective anti-tumor immunity, limiting cytotoxic T-cell activity and promoting immune escape. Another example is inflammation and immunomodulation. Benign tumors typically do not induce extensive chronic inflammation or manipulate immune infiltrates to their advantage. Malignant tumors frequently generate a pro-inflammatory yet immunosuppressive environment. This includes altered secretion of chemokines and cytokines that recruit immune cells but subdue their anti-tumor functions.

[0032] Moreover, there are tumor microenvironmental differences. Specifically, there are differences in angiogenesis and nutrient supply. Benign tumors usually have limited angiogenesis and maintain defined, often encapsulated borders. Nutrient and oxygen supplies are relatively constrained, preventing aggressive expansion. Malignant tumors actively induce angiogenesis to supply nutrients and oxygen, enabling rapid growth and increasing the risk of metastasis. This is facilitated through the secretion of VEGF and other angiogenic factors. Also, there are differences in Extracellular Matrix (ECM) and stromal interactions: In benign tumors, the ECM remains relatively intact, with less remodeling and fewer fibrotic or reactive stromal changes. In malignant tumors, the ECM is extensively remodeled, often by cancer-associated fibroblasts. This remodeling creates permissive pathways for invasion and metastasis. Stromal cells in malignant tumors can be co-opted to promote tumor growth and support a metastatic niche. Also, there are differences in the overall microenvironmental tone. In benign tumors, the local environment is comparatively stable, without widespread immunosuppression or invasive growth patterns. In malignant tumors, the tumor microenvironment is a complex, evolving niche characterized by disrupted architecture, high cellular diversity (cancer cells, fibroblasts, endothelial cells, and various immune cell subsets), and a distinctly immunosuppressive, pro-tumor milieu.

[0033] In essence, benign and malignant tumors differ substantially at the molecular, cellular, and environmental levels. While benign tumors remain relatively genetically stable, metabolically normal, and immunologically quiescent, malignant tumors leverage extensive genetic changes, metabolic rewiring, immune evasion strategies, and microenvironmental remodeling to facilitate aggressive growth, invasion, and metastasis. By interrogating these molecular and microenvironmentalsignatures - through the types of cells present, their metabolic states, and the factors they secrete - the innovative method disclosed herein can more accurately distinguish malignant cancers from benign lesions early, ultimately guiding better clinical decision-making.

[0034] The approach according to the method disclosed herein represents a significant advancement over existing methods for distinguishing between malignant and benign tumors by leveraging a comprehensive, molecularly-driven profiling strategy.

[0035] The key reasons the method of this disclosure stands out is, first, a holistic molecular signature rather than single markers. In other words, the method may use several markers combined, specifically in a molecular signature or holistic molecular signature, rather than a single marker. Typical approaches often rely on single biomarkers, imaging features, or limited immunohistochemical staining to infer whether a tumor is malignant or benign. The present method instead may employ a broad-spectrum analysis, simultaneously examining a wide array of molecules - such as metabolites, proteins, and lipid molecules - secreted by or altered within both the tumor and its microenvironment. By integrating these multiple molecular layers, the method can capture a far more nuanced, accurate signature of malignancy than any single-marker strategy.

[0036] Another key reason is that the method may employ direct assessment of tumormicroenvironment interactions. Conventional assessments, such as imaging or isolated tumor biopsies, primarily focus on tumor cells themselves and often overlook the dynamic cellular ecosystem around them. The method may however uniquely characterize the complex interplay between cancer cells and neighboring immune cells, fibroblasts, endothelial cells, and stromal components. This granular view enables the method to detect the subtle immunomodulatory and metabolic changes that accompany malignant progression, changes that benign lesions typically lack.

[0037] Yet another key reason is that the method may use enhanced specificity and sensitivity in early-stage differentiation. Typical clinical methods often struggle to distinguish early malignant changes from benign growths, leading to unnecessary follow-up scans, invasive procedures, or biopsies. The method of this disclosure however may identify metabolic and immune-regulatory shifts that are specific to malignant processes, allowing for more accurate, earlier differentiation. By detecting these molecular alterations before they become radiologically apparent, the method may significantly reduce false positives and improve early-stage diagnostic confidence.

[0038] Another key reason is that the method may reduce of reliance on imaging alone. While imaging techniques such as CT or MRI scans have improved tumor detection, they frequently cannot reliably distinguish malignant tumors from benign lesions solely based on structural characteristics. The molecular profiling suggested herein goes beyond size, shape, or radiographic density, offering a functional readout of tumor biology. This reduces the burden of repeated imaging studies and the associated costs, radiation exposure, and patient anxiety.

[0039] Another key reason is the non-invasive or minimally invasive sampling that the method employs. Many existing tests require invasive tissue biopsies, which carry risks and may not always be feasible. The method of this disclosure can however be applied to bodily fluids (e.g., blood or serum), collecting molecular signatures released by the tumor and its environment. This non- orminimally invasive sampling allows for safer, more frequent monitoring and better patient comfort, enabling clinicians to track changes over time without multiple, potentially risky procedures.

[0040] Finally, a key reason is the facilitating of personalized treatment decisions by the method. The comprehensive molecular profile not only classifies tumors accurately but may also provide insights into which signaling pathways are activated. This knowledge can inform clinicians about possible therapeutic targets, offering a stepping stone toward truly personalized oncology care. Early and accurate differentiation between benign and malignant tumors ensures patients receive the most appropriate interventions at the right time.

[0041] In summary, the approach of the method of this disclosure surpasses older methods by being able to deliver a multidimensional, metabolically and immunologically informed molecular portrait of tumors. This leads to more accurate differentiation between malignant and benign lesions, earlier diagnosis, fewer unnecessary procedures, and improved patient outcomes. It is a nextgeneration diagnostic tool that replaces limited, one-dimensional assessments with a rich, integrated molecular perspective.

[0042] Below is a comprehensive differentiation between precancerous and cancerous conditions, emphasizing how the method of this disclosure may be used for identifying and characterizing these states. Specifically, the method may, as described herein, use broad, integrative molecular profiling, combined with an in-depth examination of the tumor microenvironment (TME), which transcends current limitations in precision and clinical utility.

[0043] Fundamental distinctions between precancerous and cancerous conditions can be attributed to molecular and genetic alterations. Precancerous cells have begun to accumulate genetic and epigenetic changes, but these alterations are often partial or incomplete. The affected cells may express early oncogenic signals and show initial disruption of normal growth control but retain some regulatory checkpoints. Key tumor suppressor pathways may be partially compromised, yet not fully disabled. Cancerous malignant cells harbor widespread, well-established mutations, resulting in fully deregulated growth, robust growth factor autonomy, and diminished DNA repair capacity. Complete inactivation of tumor suppressors and / or constitutive activation of oncogenes provides a growth advantage that is no longer constrained by normal homeostatic controls. The cell population is genetically heterogeneous, exhibiting the capacity to adapt and evolve rapidly.

[0044] Regarding cellular behavior and phenotype, precancerous cells show atypical morphology and increased proliferation compared to their normal counterparts, but remain contained within their tissue of origin. They lack the invasive and metastatic capabilities characteristic of malignant cancers, and their growth rates, while elevated, are often slower and more orderly than those of fully malignant cells. Cancerous cells exhibit high proliferative rates, are poorly differentiated, and are capable of breaking through basement membranes and infiltrating surrounding tissues. Malignant cells can thrive in unfavorable conditions, show increased genomic instability, and are poised to metastasize to distant organs.

[0045] Regarding tumor microenvironment (TME) interactions, in precancerous stage, early changes in the local microenvironment occur, such as mild inflammatory infiltrates and initial stromal remodeling. Immune cells may still recognize and eliminate some aberrant cells. The TME is not yetfully co-opted to promote tumor progression, and angiogenesis is often limited or absent. In cancerous stage, malignant tumors hijack and reshape the TME extensively. There is robust angiogenesis, immune suppression, fibroblast activation, and ECM remodeling. The TME becomes immunosuppressive, nurturing cancerous cells, enabling invasion, and permitting metastasis.

[0046] Regarding clinical and therapeutic implications, in precancerous stage, these lesions represent early warning signs and can be intercepted or reversed before transitioning into full-blown cancer. Interventions may include localized removal, ablation, or less aggressive molecular treatments to restore normal cell regulation. In cancerous stage, treatment typically requires more intensive therapies (surgery, chemotherapy, radiation, immunotherapy), as the tumor is invasive and often resistant to standard interventions if detected late.

[0047] The method of this disclosure may employ a comprehensive molecular profiling at the earliest stages: Existing diagnostics often rely on late-stage biomarkers or broad imaging features that cannot reliably distinguish between truly precancerous lesions and either benign anomalies or very early-stage cancers. The method of this disclosure can utilize a wide-ranging multi-omics analysis to detect subtle molecular imprints that define the boundary between precancerous and malignant states. This is far more granular and actionable than conventional single-biomarker or imaging-only methods.

[0048] The method of this disclosure may also employ a decision based on a wide array of molecules - such as metabolites, proteins, and lipid molecules - secreted by or altered within both the tumor and its microenvironment. Unlike standard assessments that focus solely on the lesion itself, the method of this disclosure can capture and interpret (critical) changes in the TME, including immune cell composition, stromal cell activation, and extracellular matrix alterations. By delineating the supportive or suppressive niches around emerging lesions, early TME patterns that predict malignant progression can be identified. This TME-centric profiling is a breakthrough that current tools typically lack.

[0049] The method of this disclosure can also provide early-stage intervention guidance. Many existing strategies identify lesions only after malignant characteristics are well established, limiting the window of opportunity for prevention. The method of this disclosure enables clinicians and researchers to pinpoint lesions at a stage where relatively mild interventions may suffice to halt progression. By applying profile-driven insights, one can choose targeted strategies - such as immune modulation or metabolic correction - long before invasive cancer sets in.

[0050] The method may also discriminate between indolent and high-risk precancerous lesions. Standard protocols often over-diagnose and over-treat patients, as they cannot distinguish indolent lesions from those with a high probability of becoming malignant. The present method or, in other words, the method of this disclosure may identify the molecular hallmarks that correlate strongly with progression risk. This precision spares patients unnecessary invasive procedures and reduces healthcare costs by focusing interventions on lesions that truly need it.

[0051] Finally, the method is scalable and can be easily integrated with current screening protocols. The present method can integrate seamlessly into existing screening frameworks (e.g., imaging, cytology) by adding a molecular layer of analysis. The result is a refined decision-making process,enabling improved patient stratification, personalized follow-up intervals, and more informed patient counseling.

[0052] Accordingly, the present method represents a paradigm shift. By uniting a deep, integrative molecular approach using a wide array of molecules - such as metabolites, proteins, and lipid molecules - secreted by or altered within both the tumor and its microenvironment, and by focusing on the transition zone between precancerous and malignant states, a powerful, early-stage diagnostic and prognostic tool is provided. This comprehensive strategy ensures that interventions can be more targeted, effective, and implemented well before a lesion becomes a life-threatening malignancy - making our approach genuinely transformative in the field of cancer prevention and early detection.

[0053] Further, the computer-implemented method described herein can be considered or be utilized as screening method, which may be incorporated in a normal or routine checkup of patients or individuals, for example when a biofluid sample is drawn from the individual during a routine checkup at a health care provider (HCP) or hospital. Furthermore, 1 H-NMR is also well suited for high-throughput screening approaches, relatively inexpensive, and has high reproducibility. Hence, the method described herein can assist HCPs with an early detection or identification of individuals at risk for cancerous and / or pre-cancerous disease. As a consequence, the method described herein can allow to reliably and accurately risk stratify an individual as having a cancerous disease, a precancerous disease and / or as being healthy or having no cancer. In turn, this may allow for improved and more efficient utilization of healthcare resources as well as cost savings.

[0054] In the context of the present disclosure, the term “individual” may generally refer to a vertebrate, including animals and human beings. The term individual may be synonymously or interchangeably used herein with subject or patient.

[0055] As used herein, the 1 H-NMR spectroscopy data can refer to data obtained by and / or generated with a high-frequency NMR spectrometer, for example at a frequency equal to or above about 500 MHz, preferably equal to or above 600 MHz. For instance, the 1 H-NMR spectroscopy data, also referred to as 1 H-NMR spectrum, may be obtained with a pulse sequence similar to CPMGPR1 D.

[0056] Generally, the 1 H-NMR spectroscopy data can contain or be indicative of the free induction decay (FID), which may be measured as raw 1 H-NMR spectroscopy data in a 1 H-NMR spectroscopy of the biofluid sample. The measured FID may optionally be converted to a spectrum via Fourier Transformation, and further optionally referenced and / or converted to chemical shift in parts per million (ppm), for example based on normalization to a reference peak in the spectrum, such as e.g. the peak of lactate or anomere D-Glucose duplet. Accordingly, the phrase “1 H-NMR spectroscopy data indicative of an NMR signal intensity as a function of the chemical shift” includes 1 H-NMR spectroscopy data given as function of chemical shift, as well as 1 H-NMR spectroscopy data given as function of FID or another quantity, from which the NMR signal intensity as a function of chemical shift can be derived.

[0057] For example, normalization techniques can be employed within the method, including Probabilistic Quotient Normalization (PQN). PQN accounts for variations in the dilution of samplesby standardizing the data relative to a reference. This ensures that differences in metabolite concentrations are not skewed by sample preparation inconsistencies.

[0058] Further, or alternatively, the method may involve one or more alignment step(s) to address peak shift variability in spectral data. Techniques such as Recursive Segment-Wise Peak Alignment (RSPA) are particularly effective, as they iteratively adjust and align the peaks within segments of the spectra, improving the consistency and comparability of the dataset across samples.

[0059] As used herein, obtaining the 1 H-NMR spectroscopy data may comprise accessing said 1 H- NMR spectroscopy data and / or retrieving the 1 H-NMR spectroscopy data. For instance, the 1 H- NMR spectroscopy data may be accessed at and / or received from at least one memory or data storage of the computing device that carries out the method of the present disclosure, from a memory or data storage of another computing device, and / or from a remote data storage (a database, a further memory, a cloud storage or the like). Accordingly, retrieving the 1 H-NMR spectroscopy data may optionally comprise downloading said 1 H-NMR spectroscopy data from an external computing device or data storage.

[0060] Additionally or alternatively, obtaining the 1 H-NMR spectroscopy data may comprise receiving the 1 H-NMR spectroscopy data, e.g. from a computing device different than the computing device accessing the data. Accordingly, obtaining the 1 H-NMR spectroscopy data may comprise one or more of receiving the 1 H-NMR spectroscopy data, storing said 1 H-NMR spectroscopy data in the memory or data storage of the computing device, and retrieving the 1 H-NMR spectroscopy data by the computing device.

[0061] As used herein, evaluating, with the at least one trained machine learning model of the computing device, the 1 H-NMR spectroscopy data in terms of the molecular profile of the biofluid sample, may include processing and / or analyzing the 1 H-NMR spectroscopy data by means of the at least one trained ML model with respect to all hydrogen peaks contained in the 1 H-NMR spectroscopy data. Such evaluation of the 1 H-NMR spectroscopy data in terms of the molecular profile may include evaluating and / or analyzing the molecular profile of the biofluid sample with respect to one or more reference molecular profiles of one or more reference individuals. For example, the at least one ML model may be trained based on reference I HjNMR spectroscopy data of one or more reference individuals.

[0062] Based on the evaluation of the molecular profile and / or 1 H-NMR spectroscopy data of the biofluid sample, the at least one trained machine learning model can perform a binary or multiclass classification into at least two classes of molecular profiles, namely into at least the first class of molecular profiles associated with healthy individuals and the second class of molecular profiles associated with a cancerous and / or pre-cancerous disease. Accordingly, the at least one machine learning model may refer to or include a classification model or algorithm, which has been trained in or in accordance with a machine learning approach.

[0063] Further, as used herein, determining the cancer risk score may include computing and / or assessing the risk, likelihood and / or probability for a cancerous and / or pre-cancerous disease occurring and / or being present at the individual. This may include assessing, computing and / or determining the probability and / or likelihood for cancer, respectively a cancerous disease and / or pre-cancerous disease, occurring and / or being present at the individual. In other words, determining the cancer risk score may, for example, include computing and / or determining the probability and / or likelihood for cancer occurring at the individual based on or using the at least one trained machine learning model of the computing device.

[0064] The cancer risk score, as used herein, may refer to a numerical measure indicative of the determined risk, likelihood and / or probability for a cancerous and / or pre-cancerous disease occurring and / or being present at the individual. Therein, the cancer risk score may be provided on an arbitrary scale ranging from a minimum value, for example zero or 0, to a maximum value, for example one or 100. Any other scale, including relative and absolute scales, can be used to represent the cancer risk score. It should be noted that a plurality of cancer risk scores may be computed, for example using a plurality of different machine learning models or algorithms, as described in more detail hereinbelow.

[0065] Optionally, the cancer risk score may be provided as output by the computing device, for example at a user interface of the computing device. Further optionally, additional or contextual information associated with the cancer risk score may be determined by the computing device, which may optionally be provided as output by the computing device. Such contextual information may include a risk level or risk tier, such as low risk, medium risk, high risk, for a cancerous and / or pre-cancerous disease occurring and / or being present at the individual. Alternatively or additionally, contextual information may include an indication about a type of cancer and / or an estimated stage of cancer.

[0066] The computing device described herein may refer to any data processing device having one or more processors for data processing. The computing device may be embodied as a standalone computing device, as a server, and / or as computing network with a plurality of inter-operating computing devices, such as a cloud computing system or server system. Alternatively or additionally, the computing device may be embodied, at least in part, as mobile device, such as a smart phone a tablet computer, a notebook or the like.

[0067] According to an embodiment, classifying the molecular profile includes determining and / or computing a classification result indicative of a probability and / or likelihood for at least one of the first class and the second class, wherein the cancer risk score is determined based on the classification result. Accordingly, the cancer risk score may comprise or correspond to the classification result computed or generated by the at least one ML model. For instance, classifying the molecular profile may include coming to a conclusion as to whether the molecular profile of the biofluid sample is a member of the first class, the second class, and optionally one or more further classes. Alternatively or additionally, classifying the molecular profile may include computing a likelihood for the molecular profile of the biofluid sample to belong to the first class of molecular profiles, to the second class of molecular profiles and optionally one or more further classes of molecular profiles. In other words, the at least one machine learning model may be trained to determine a probability and / or likelihood for binary and / or multiclass classification of the molecular profile of the biofluid sample. It should be noted that in case of binary classification only a classification result of one of the first and second class may be computed and the classification result of the other one of the first and second classmay be computed based thereon, e.g. based on subtraction of the determined classification result from one or 100%.

[0068] According to an embodiment, the biofluid sample is selected from the group consisting of a blood serum sample, a blood plasma sample, a blood sample, a urine sample, and a cerebrospinal fluid sample. In other words, the biofluid sample may comprise one or more of a blood serum sample, a blood plasma sample, a blood sample, a urine sample, and a cerebrospinal fluid sample.

[0069] Depending on the type of cancer or cancerous disease, the composition of one or more biofluids of an individual may be altered, modified and / or changed, wherein such changes may optionally be induced at different stages of the cancer. A major part of cancer types and cancerous diseases expectedly change the composition or molecular composition of the blood of an individual, for example due to their altered metabolism and protein composition, which can change the proportion of one or more proteins, metabolites, amino acids, micro- and / or macromolecules that are either directly released by the cancer cells into the blood or whose release is directly or indirectly induced by the cancer, e.g. via biological networks. In these cases, a blood sample, a blood serum sample and / or a blood plasma sample may be used as biofluid sample to generate the 1 H-NMR spectroscopy data and determine the cancer risk score. Other cancer types or cancerous diseases, such as tumors of the bladder, kidney or urine ducts, on the other hand, may alter the molecular composition of urine at early cancer stage, and a urine sample may be utilized as biofluid sample to generate the 1 H-NMR spectroscopy data and determine the cancer risk score, which can allow for an early cancer detection. Yet other cancer types or cancerous diseases, such as brain tumors, may alter the molecular composition of CSF, and a CSF sample may be utilized as biofluid sample to generate the 1 H-NMR spectroscopy data and determine the cancer risk score. Accordingly, using one or more of a blood serum sample, a blood plasma sample, a blood sample, a urine sample, and a cerebrospinal fluid sample can allow for an earliest possible diagnosis or cancer detection.

[0070] The molecular differences in an individual’s bloodstream described herein are typically not attributable solely to the cancer cells themselves. While malignant cells directly alter their metabolic and protein composition - releasing distinct proteins, metabolites, amino acids, and various micro- and macromolecules into circulation - the tumor’s influence extends well beyond these direct contributions. Through complex biological networks and signaling cascades, the tumor indirectly prompts changes in other cells and tissues, resulting in systemic metabolic and lipidomic shifts that can be detected early and comprehensively. For example, the release of certain cytokines, such as interleukin-6 (IL-6), is one of many tumor-driven signals that do not originate exclusively from malignant cells. When secreted into the bloodstream, IL-6 can induce the liver and other tissues to remodel their metabolism, adjusting the levels of key lipid-binding proteins like apolipoproteins A1 and A2. Such changes represent systemic metabolic realignments, illustrating how the cancer and its microenvironment can reshape an individual’s biochemical landscape. Within the tumor microenvironment (TME), immune cells, stromal cells, and other supportive cell populations also undergo metabolic reprogramming. These non-malignant but transformed neighbors modify their consumption of nutrients, production of waste products, and overall cellular outputs in response to the tumor’s presence. For instance, fibroblasts may alter the extracellular matrix, influencing howproteins are processed and released. Immune cells might adapt their amino acid usage or lipid metabolism under the chronic stress of tumor antigens and immunosuppressive cues. Together, these shifts create a distinctive pattern of altered proteins, lipids, and metabolites and modifications of those in the bloodstream - patterns that do not rely solely on the tumor’s own secretions but also on the TME’s induced changes. The approach of the herein disclosed method captures the entire network of direct and indirect alterations - from malignant cells, the TME, and the host’s systemic response - it can identify the earliest subtle biochemical hints of cancer development. Instead of seeking a single marker from a particular cell type, the method may integrate multiple molecular signatures that arise from the interplay of cancer cells, immune cells, stromal components, and systemic metabolic regulators. This holistic perspective enables the detection of malignancies at very early stages, often before traditional diagnostics can pick up any structural or symptomatic evidence of disease.

[0071] According to an embodiment, the 1 H-NMR spectroscopy data is indicative of a NMR signal intensity in binned increments of the chemical shift, each increment having a width of less than or equal to about 0.02 ppm, preferably less than or equal to about 0.01 , even more preferably less than or equal to about 0.006 ppm, for example about 0.00016 ppm. It should be noted that the aforementioned ppm values and ranges apply to any incremental bin of the chemical shift discussed or disclosed herein. Generally, utilizing such a fine-structured binning of the 1 H-NMR spectrum can increase the informational content of the 1 H-NMR spectroscopy data, and thus allow to identify fine- structured changes in the molecular profile induced by cancer or cancerous disease at very early stage of the cancer.

[0072] According to an embodiment, the binning may be conducted at least over a spectrum from -1 ppm to 10 ppm of the chemical shift. In particular, the binning may be conducted over the whole spectrum, which may be from -1 ppm to 10 ppm. This may allow for fine grained information to be extracted over not only metabolically relevant regions but also regions in which more complex, overlapping protein signals are arising. Furthermore, by using a binning approach and not only extracting metabolites and lipids by deconvolution we are extracting overlapping regions of molecules or molecular substructures, like post translational modifications (PTMs), i.e. methylation, sugar-residues, protein substructures or proteins, that are either increased or decreased by the enzymes that are more or less strongly expressed in cancer cells, immune cells or stromal cells in the tumor microenvironment or the organ the tumor is located in. This approach also offers a higher resolution since not only peaks of abundant molecules, like metabolites or lipids are considered but also peaks of less abundant molecules that might be in the left or right shoulder of a peak of a more abundant protein. By not only looking into the area under the peak, but looking into the bins along the peak, these subtle changes are detected and can be considered when identifying a cancer risk score.

[0073] According to an embodiment, the method further comprises converting the 1 H-NMR spectroscopy data into a binned data structure based on mapping and / or associating the NMR signal intensity to binned increments of the chemical shift, each increment having a width of less than or equal to about 0.02 ppm, preferably less than or equal to about 0.01 , even more preferably less thanor equal to about 0.006 ppm, for example about 0.00016 ppm. Alternatively or additionally, the method may further comprise binning the 1 H-NMR spectroscopy data into a plurality of incremental bins of the chemical shift, each bin having a width of less than or equal to about 0.02 ppm, preferably less than or equal to about 0.01 , even more preferably less than or equal to about 0.006 ppm, for example about 0.00016 ppm. Optionally, the binned 1 H-NMR spectroscopy data may be normalized using probabilistic quotient normalisation. An advantage of such normalizing can be a reduction of the variability between different pre-analytical approaches.

[0074] In an exemplary implementation, a biofluid sample may be analyzed with a high-frequency NMR spectrometer, e.g. at or above about 500 MHz with a pulse sequence similar to CPMGPR1 D. The free induction decay (FID) may be measured and converted to a spectrum via Fourier Transformation, and the Fourier transformed spectrum may be normalized to a reference peak, such as e.g. lactate, anomere D-Glucose duplet, or another peak. Further, the baseline may be corrected and the spectrum may then be converted to binned increments of smaller or equal to 0.02 ppm, preferably less than or equal to about 0.01 , even more preferably less than or equal to about 0.006 ppm, for example about 0.00016 ppm. This allows to get a significant amount of data for the deciphering of the molecular profile of the biofluid sample and gives information about all hydrogen atoms present, wherein each bin may represent a specific position in the 1 H-NMR spectrum or spectroscopy data.

[0075] According to an embodiment, the at least one trained machine learning model is a machine- learned classifier configured to process, as input data, the 1 H-NMR spectroscopy data in a data structure of binned increments of the chemical shift. In other words, the at least one ML model may be trained or machine-learned to receive and / or process the 1 H-NMR spectroscopy data in a data structure of binned increments of the chemical shift, and provide as an output a classification result and / or the at least one cancer risk score, which may represent, include and / or correspond to the classification result.

[0076] According to an embodiment, the method may be focusing on significant changes in different bins or bin regions that allow a concrete and highly accurate differentiation between cancer and noncancer patient samples. In particular, the method described herein may be using only a couple hundred features to achieve a very high differentiation power (AUC of around 0.98), e.g. in the range of 100 to 900 features, in particular in the range of 200 to 700 features. This allows to keep the ML algorithm lean and allows to avoid biases like overfitting or underfitting. Hence, the method may have a significant small number in selected features (e.g., 400 vs. for example 3000 features) together with the significant improvement in accuracy (AUC of 0.98, specificity above 97.5%, sensitivity above 90%)

[0077] According to an embodiment, the at least one machine learning model is trained and / or configured to identify at least one molecular signature of cancerous and / or pre-cancerous disease in the molecular profile of the biofluid sample and / or in the 1 H-NMR spectroscopy data. Generally, the ML model may be trained and / or configured to identify and / or determine one or a plurality of molecular signatures of cancerous and / or pre-cancerous disease in the 1 H-NMR spectroscopy data. Therein, the one or more molecular signatures may be identified at any position or location in thespectrum, for example in one or more bins of the chemical shift. Further, a molecular signature of cancerous and / or pre-cancerous disease may include any alteration and / or change of the molecular profile of the biofluid sample with respect to one or more molecular profiles of healthy individuals. Such molecular signature, respectively such alteration and / or change in the molecular profile with respect to molecular profiles of healthy individuals, may, for example, include an additional hydrogen peak in the 1 H-NMR spectroscopy data induced by a cancerous and / or pre-cancerous disease, a hydrogen peak of increased height or size, a hydrogen peak of reduced height or size, a change in a width of a hydrogen peak, an overlap of a plurality of hydrogen peaks, removal of a hydrogen peak, a shift in ppm of one or more hydrogen peaks, or a combination thereof.

[0078] According to an embodiment, the determined cancer risk score is usable as computational biomarker indicative of a pathogenic molecular signature and / or molecular signatures of cancerous and / or pre-cancerous disease in the molecular profile of the biofluid sample. In other words, the determined cancer risk score may constitute a computational biomarker to inform about pathogenic molecular signatures in the individual’s biofluid sample. As described above, one or a plurality of different molecular signatures can be taken into consideration to compute the cancer risk score. Hence, also different computational biomarkers in the form of different cancer risk scores for different types of cancer, cancerous disease and / or pre-cancerous disease may be computed in accordance with the method described herein. Accordingly, the method described herein may be applied to or used to detect many different types of cancerous and / or pre-cancerous disease, thereby providing a versatile approach or method for cancer detection and / or assessing the risk of cancer.

[0079] According to an embodiment, identifying the at least one molecular signature includes identifying one or more bins in the 1 H-NMR spectroscopy data associated with and / or containing one or more hydrogen peaks indicative of a cancer-induced change in the molecular profile with respect to molecular profiles associated with or of healthy individuals and / or reference individuals. Based on the identified one or more bins in the 1 H-NMR spectroscopy data, the cancer risk score may be computed. Accordingly, the at least one ML model may be trained to identify one or more cancer-induced changes in one or more hydrogen peaks with respect to or relative to one or more molecular profiles of healthy individuals.

[0080] According to an embodiment, identifying the at least one molecular signature includes identifying one or more bins and / or a range of chemical shift in the 1 H-NMR spectroscopy data associated with an overlap of hydrogen peaks assignable to a plurality of metabolites, proteins, amino acids, micro-molecules and macromolecules contained in the biofluid sample. In other words, the ML model may be trained to identify an overlap of hydrogen peaks, which may be induced by a cancerous and / or pre-cancerous disease and which may comprise one or more bins of the chemical shift in the 1 H-NMR spectroscopy data. This may allow to identify complex patterns of molecular signatures induced by the cancerous and / or pre-cancerous disease, and / or to take complex patterns of molecular signatures into account for computing the cancer risk score. Thereby, overall versatility of the method may be increased, for example to allow for a determination of various different types of cancerous and / or pre-cancerous diseases.

[0081] According to an embodiment, identifying the at least one molecular signature includes identifying one or more bins and / or a range of chemical shift in the 1 H-NMR spectroscopy data non- assignable to single metabolites, proteins, amino acids, micro molecules, and macromolecules contained in the biofluid sample. Accordingly, also molecular signatures which may not be assignable to and / or associated with single metabolites, proteins, amino acids, micro molecules, and macromolecules contained in the biofluid sample may be taken into account to compute the cancer risk score. Also this may allow to increase an overall versatility of the method, for example to allow for a determination of various different types of cancerous and / or pre-cancerous diseases.

[0082] According to an embodiment, the at least one molecular signature is associated with a cancer-induced change in the proportion of one or more hydrogen peaks in the 1 H-NMR spectroscopy data associated with one or more metabolites, proteins, amino acids, micro molecules and macromolecules contained in the biofluid sample with respect to a biofluid sample of one or more healthy individuals. Therein, a change in the proportion of a hydrogen peak may refer to one or more of a height of the peak, a size of the peak, an integral of the peak, an area of the peak, a mean of the peak, a center of the peak, a location or position of the peak in the 1 H-NMR spectrum, and a combination thereof. Generally, by identifying changes in the proportion of one or more hydrogen peaks for classifying the molecular profile into at least the first and second class, the cancer risk score may be accurately and reliably be detected.

[0083] According to an embodiment, evaluating the 1 H-NMR spectroscopy data in terms of the molecular profile includes analyzing all hydrogen peaks in the 1 H-NMR spectroscopy data associated with one or more metabolites, proteins, amino acids, micro molecules and macromolecules contained in the biofluid sample. By analyzing all hydrogen peaks in the molecular profile, a plurality of or all molecular signatures induced by the cancerous and / or pre-cancerous disease in a molecular profile can be taken into account in the determination of the cancer risk score, thereby improving accuracy and robustness of the determination as well as increasing versatility of the determination.

[0084] According to an embodiment, evaluating the 1 H-NMR spectroscopy data in terms of the molecular profile includes detecting a pattern of hydrogen peaks associated with a cancer-induced change in the proportion of one or more hydrogen peaks in the 1 H-NMR spectroscopy data with respect to a biofluid sample of one or more healthy individuals. Alternatively or additionally, the at least one machine learning model may be trained to detect a pattern of hydrogen peaks associated with a cancer-induced change in the proportion of one or more hydrogen peaks in the 1 H-NMR spectroscopy data with respect to a biofluid sample of one or more healthy individuals. For instance, complex correlations between cancerous diseases and molecular signatures in the molecular profile of an individual may occur, which may lead to patterns of hydrogen peaks in the 1 H-NMR spectroscopy data that may be detected for determining the cancer risk score. Hence, by detecting such patterns of hydrogen peaks, the cancer risk score may be computed with high accuracy, reliability, and robustness.

[0085] According to an embodiment, obtaining the 1 H-NMR spectroscopy data includes acquiring raw 1 H-NMR spectroscopy data with an NMR spectrometer at a frequency above about 500 MHz.The raw 1 H-NMR spectroscopy data may can contain or be indicative of the free induction decay, as measured in a 1 H-NMR spectroscopy. Accordingly, acquiring raw 1 H-NMR spectroscopy data with an NMR spectrometer at a frequency above about 500 MHz may include measuring the free induction decay using an NMR spectrometer.

[0086] In an exemplary implementation, the method further spectral processing of the raw 1 H-NMR spectroscopy data and / or peak aligning the raw 1 H-NMR spectroscopy data. This may allow to correct for shifts in the position or chemical shift of hydrogen peaks of different biofluid samples, for example such that hydrogen peaks associated with a particular metabolite, but measured at different biofluid samples, is at the same position in the corresponding 1 H-NMR spectroscopy data. Hence, inter-comparability of 1 H-NMR spectroscopy data of different biofluid samples may be improved.

[0087] According to an embodiment, spectral processing of the raw 1 H-NMR spectroscopy data includes one or more of chemical shift referencing, phasing, and baseline correction of the raw 1 H- NMR spectroscopy data. Accordingly, one or more of these pre-processing steps may be applied to the raw 1 H-NMR spectroscopy data, which may increase inter-comparability of 1 H-NMR spectroscopy data of different biofluid samples, and improve robustness and reproducibility of the determination of the cancer risk score with respect to different biofluid samples.

[0088] According to an embodiment, the method further comprises one or more of normalizing, scaling, binning and filtering of the raw 1 H-NMR spectroscopy data and / or the 1 H-NMR spectroscopy data. Also this can increase inter-comparability of 1 H-NMR spectroscopy data of different biofluid samples, and improve robustness and reproducibility of the determination of the cancer risk score with respect to different biofluid samples.

[0089] According to an embodiment, the at least one machine learning model is trained for classification using one or more statistical machine learning algorithms. Generally, statistical machine learning algorithms may utilize statistical techniques to develop models that can learn from data and make predictions or decisions based on the learning. Exemplary and non-limiting statistical machine learning algorithms include a voting classifier, an ensemble method, and logistic regression.

[0090] A voting classifier is a machine learning model that trains on an ensemble of numerous models and predicts an output or at least one class based on the highest probability of the chosen class as the output. Ensemble methods, on the other hand, are techniques that aim at improving the accuracy of results in models by combining multiple models instead of using a single model. The combined models may thus increase the accuracy of the classification results obtained. Further, logistic regression relates to a statistical method that can be used for building machine learning models where the dependent variable is dichotomous and / or binary. Accordingly, logistic regression can particularly be utilized for binary classification into two classes, such as the first and second classes of molecular profiles.

[0091] According to an embodiment, the classifying of the molecular profile into at least the first class and the second class of molecular profiles comprises classifying, based on evaluating the 1 H- NMR spectroscopy data in terms of the molecular profile of the biofluid sample with a first trained machine learning model, the molecular profile into at least a healthy class of molecular profilesassociated with healthy individuals and a non-healthy or diseased class of molecular profiles associated with non-healthy or diseased individuals. Further, the method comprises, upon determining that the molecular profile of the biofluid sample is associated with or is in the non- healthy or diseased class of molecular profiles, classifying, based on evaluating the 1 H-NMR spectroscopy data in terms of the molecular profile of the biofluid sample with a second trained machine learning model, the molecular profile into a non-cancer class of molecular profiles associated with non-cancerous individuals and a cancer class of molecular profiles associated with cancerous and / or pre-cancerous individuals. Therein, the first trained machine learning model and the second machine learning model differ from one another. In particular, the first and second machine learning models may differ from one another in terms of the training of the respective models. For example, the first ML model may be trained to provide a binary classification of the molecular profile of the biofluid sample into the healthy class of molecular profiles associated with healthy individuals and the non-healthy or diseased class of molecular profiles associated with non- healthy or diseased individuals, whereas the second ML model may be trained to a binary classification of the molecular profile of the biofluid sample into the non-cancer class of molecular profiles associated with non-cancerous individuals and the cancer class of molecular profiles associated with cancerous and / or pre-cancerous individuals. Alternatively or additionally, different types of ML models may be utilized for the first and second ML models, such as e.g. different statistical ML models or algorithms. Alternatively or additionally, different classifications may be used for the first and second ML models, such as for example a classification with respect to cancer staging, e.g. metastatic and non-metastatic, and / or a classification with respect to sub-disease. Alternatively or additionally, the first and second ML models can differ in one or more of the number of parameters, the selected algorithm and feature engineering methods. For example, principal component analysis (PCA) and / or t-distributed stochastic neighbor embedding (tSNE) can be used to generate features based on unsupervised learning for the first and / or second ML models.

[0092] Accordingly, different ML models may be sequentially used or invoked to determine the cancer risk score. Applying a multi-step approach for determining the cancer risk score can generally increase accuracy and robustness of the overall method of determining the cancer risk score. Also, versatility of the overall method of computing the cancer risk score can be increased.

[0093] According to an embodiment, the method further comprises determining, based on the classifying of the molecular profile of the biofluid sample with the first trained machine learning model, a first cancer risk score indicative of a probability for the molecular profile being associated with or in the healthy class and / or the non-healthy or diseased class. The method further comprises determining, based on the classifying of the molecular profile of the biofluid sample with the second trained machine learning model, a second cancer risk score indicative of a probability for the molecular profile being associated with or in the non-cancer class and / or the cancer class of molecular profile, and determining the cancer risk score based on the first and second cancer risk scores. For instance, the classification result and / or second cancer risk score of the second ML model may be used as or constitute the cancer risk score. Alternatively or additionally, both the first and second cancer risk scores may be combined, for example in accordance with a predefinedmetrics, to provide the cancer risk score. Alternatively or additionally, a calibration classifier may be utilized to provide refined probabilities for the classification results, respectively the first and / or second cancer risk scores.

[0094] According to an embodiment, the method further comprises, upon determining that the molecular profile of the biofluid sample is associated with or is in the cancerous class of molecular profiles, classifying, based on the evaluating 1 H-NMR spectroscopy data in terms of the molecular profile of the biofluid sample with a third trained machine learning model, the molecular profile into a plurality of classes of molecular profiles, each class being associated with a particular type of cancerous disease and / or pre-cancerous disease. Optionally, a third cancer risk score indicative of a probability for the molecular profile being associated with a particular type of cancerous and / or pre- cancerous disease may be determined. Accordingly, a type of cancer, cancerous disease and / or pre-cancerous disease may be determined by means of the method described herein.

[0095] The third ML model may differ from one or both the first and second ML models. In particular, the third ML model may differ from the first and / or second ML model in terms of training. Specifically, the third ML model may be trained to provide a multiclass classification result for classifying the molecular profile into the plurality of classes of molecular profiles, each class being associated with a particular type of cancerous disease and / or pre-cancerous disease, whereas the first and second ML models may be trained to provide a binary classification result, as described above. Optionally, also a different type of ML model may be used for the third ML model in comparison to the first and second ML models. Alternatively or additionally, different classifications may be used for the first, second and third ML models, such as for example a classification with respect to cancer staging, e.g. metastatic and non-metastatic, and / or a classification with respect to sub-disease. Alternatively or additionally, the first, second and third ML models can differ in one or more of the number of parameters, the selected algorithm and feature engineering methods. For example, principal component analysis (PCA) and / or t-distributed stochastic neighbor embedding (tSNE) can be used to generate features based on unsupervised learning for at least one of the first, second and third ML models.

[0096] According to an embodiment, the method may further comprise providing and / or adjusting a cut-off value of the cancer risk score to produce a desired ratio of false-positives and false-negatives (for the probabilities for cancer occurring at the individual). The cut-off value may in particular but not only be provided and / or adjusted to meet the desired ratio. By adjusting one or more cancer risk score cut-off values, the at least one ML model can be adjusted for patient or individual groups and desired false-positive and / or false-negative ratios, for example based on an Area-Under-The-Curve- AUC-receiver operator curve (AUC-ROC). In other words, the method may include a process of setting desired false positive and / or false negative ratios that may be catered to a given patient or group of patients. The at least one ML model can take into account these ratios.

[0097] According to an embodiment, classifying the molecular profile may include using a three- stage network approach, wherein the first stage determines whether the molecular profile falls within the first class or the second class, the second stage determines whether the molecular profile is malignant or non-malignant and classifies the molecular profile / sample into a malignant class ornon-malignant class, and the third stage determines a type of cancer of the molecular profile and classifies the molecular profile into a class corresponding to the type of cancer. Accordingly, instead of categorizing malignant vs non-malignant, the method categorizes by the type of the cancer. That is, the method may determine healthy vs unhealthy, cancerous vs noncancerous, and then determines what disease or cancer is present.

[0098] According to an embodiment, the 1 H-NMR spectroscopy data, prior to evaluation, may be normalized using probabilistic quotient normalization.

[0099]

[0100] A further aspect of the present disclosure relates to a computing device configured to perform steps of the method, as described hereinabove and hereinbelow. It is emphasized that any feature, function, element and / or step described herein with reference to the method can be a feature, function and / or element of the computing device, and vice versa.

[0101] The computing device includes one or more processors for data processing. Further, the computing device may comprise a data storage and / or memory for storing data, such as for example the 1 H-NMR spectroscopy data, one or more cancer risk scores, and / or other data. Further, the data storage and / or memory may store software instructions and / or a computer program, which, when executed by a computing device, instructs the computing device to perform steps of the method described hereinabove and hereinbelow.

[0102] Further, the at least one machine learning model, for example one or more of the first ML model, the second ML model, and the third ML model may be implemented in software and / or hardware at the computing device.

[0103] In an exemplary implementation, the computing device may comprise one or more communication interfaces configured to communicate with one or more remote devices, for example an external data storage or database, and / or one or more remote computing devices. At least a part of the 1 H-NMR spectroscopy data may be received via the communication interface from an external data storage or database. Alternatively or additionally, the computed cancer risk score or information related thereto may be transmitted to one or more remote devices or stored at the external data storage or database.

[0104] A further aspect of the present disclosure relates to computer program, which, when executed by a computing device, instructs the computing device to perform steps of the method described hereinabove and hereinbelow.

[0105] Yet another aspect of the present disclosure relates to a non-transitory computer-readable medium storing a computer program, which, when executed by a computing device, instructs the computing device to perform steps of the method described hereinabove and hereinbelow.BRIEF DESCRIPTION OF THE DRAWINGS

[0106] Exemplary embodiments will be further described with reference to Figures, wherein:

[0107] Figure 1 shows a computing device according to an exemplary embodiment;

[0108] Figure 2 shows 1 H-NMR spectroscopy data to illustrate steps of a computer-implemented method of determining a cancer risk score;

[0109] Figure 3A shows a flow chart illustrating steps of a computer-implemented method of determining a cancer risk score according to an exemplary embodiment;

[0110] Figures 3B to 3D each show 1 H-NMR spectroscopy data usable as input in the method illustrated of Figure 3A; and

[0111] Figures 4 to 6 each show a flow chart illustrating steps of a computer-implemented method of determining a cancer risk score according to exemplary embodiments.

[0112] The Figures are schematic only and not true to scale. In principle, identical or like parts, elements and / or steps are provided with identical or like reference numerals in the Figures.DETAILED DESCRIPTION OF EXEMPLARY EMBODIMENTS

[0113] Figure 1 shows a computing device or system 100 configured to determine a cancer risk score according to an exemplary embodiment.

[0114] The computing device 100 comprises a processing circuitry 110 with one or more processors 112 for data processing.

[0115] The computing device 100 further comprises at least one machine learning model 114. The ML model 114 may be implemented in software and / or hardware, and may be part of the processing circuitry 110 of the computing device 100. The computing device 100 may optionally include a plurality of machine learning models 114, such as e.g. a first, a second, and a third ML model, as e.g. described in more detail with reference to Figure 6 hereinbelow. The optional plurality of ML models are jointly shown in Figure 1 with reference numeral 114.

[0116] The computing device 100 further comprises at least one data storage 116 and / or memory 116 for storing data and / or software instructions, for example in the form of a computer program for instructing the computing device 100 to carry out steps of the method of determining the cancer risk score, as described hereinabove and hereinbelow.

[0117] The computing device shown in Figure 1 further comprises a user interface 118 for outputting data and / or information, such as the cancer risk score, to a user of the computing device 100, and / or for controlling the computing system 100 by the user.

[0118] Further, the computing system 100 comprises a communication interface 120 for communicatively and / or operatively coupling the computing device 100 to an external computing device 500, and / or for coupling the computing system 100 to one or more external data sources 500, for example to receive or obtain 1 H-NMR spectroscopy data therefrom and / or to transmit data thereto, such as e.g. the cancer risk score.

[0119] Figure 2 shows 1 H-NMR spectroscopy data in 200 to illustrate steps of a computer- implemented method of determining a cancer risk score. Specifically, Figure 2 shows the NMR signal intensity in arbitrary units as a function of the chemical shift in parts per million, ppm.

[0120] Figure 2 shows an exemplary molecular profile 210a of a healthy individual in solid line in comparison to a molecular profile 210b of an individual having a cancerous and / or a pre-cancerous disease in dashed line. Therein, the molecular profiles 210a, b comprise or consist of all hydrogen peaks of the 1 H-NMR spectroscopy data 200.

[0121] Further, the molecular profiles 210a, b shown in Figure 2 comprise groups or peaks 220 assignable to metabolites, proteins, amino acids, micro molecules and macromolecules assignable to and / or associated with the hydrogen peaks shown in the respective molecular profiles 210a, b.

[0122] As can be seen by comparing molecular profile 210a with molecular profile 210b, the molecular profiles 210a, b of healthy individuals and individuals having a cancerous and / or pre- cancerous disease differ in several regions of the chemical shift. Such difference may be referred to herein as molecular signatures associated with cancerous and / or pre-cancerous disease. Based on these differences and / or molecular signatures, the ML model 114 and / or computing device 100 may determine the cancer risk score, as described in more detail hereinabove and hereinbelow.

[0123] For instance, a peak 222 assignable to choline phospholipide is present in the molecular profile 210b of the cancerous individual, whereas the profile 210a of the healthy individual lacks this peak. Also, proportions of the taurine peak 220, the glutamine peak 224, the glutamine and glutamate peak 226, the amino acids peaks 228, the lipid and fatty acid peaks 230, 232 and others differ between the molecular profile 210a of the healthy individual and the molecular profile 210b of the cancerous and / or pre-cancerous individual. One or more of such differences may be used by the ML model 114 and / or computing device 100 to determine the cancer risk score, as described in more detail hereinabove and hereinbelow.

[0124] Figure 3A shows a flow chart illustrating steps of a computer-implemented method of determining a cancer risk score according to an exemplary embodiment. Figures 3B to 3D each show 1 H-NMR spectroscopy data 300a, 300b, 300c usable as input in the method of Figure 3A. Specifically, Figures 3B to 3D each show the NMR signal intensity in arbitrary units as a function of the chemical shift in parts per million, ppm. Therein, the 1 H-NMR spectroscopy data 300a of Figure 3B correspond to or show a molecular profile 310a associated with a healthy individual. The 1 H- NMR spectroscopy data 300b of Figure 3C correspond to or show a molecular profile 310b associated with an individual having a pre-cancerous disease, such as e.g. pancreatitis, and the 1 H- NMR spectroscopy data 300c of Figure 3D correspond to or show a molecular profile 310c associated with an individual having a cancerous disease, such as e.g. pancreatic cancer. In the following, it may be commonly referred to Figures 1 to 3D.

[0125] The method comprises a step S1 of obtaining, at a computing device 100, Hydrogen-1 Nuclear Magnetic Resonance, 1 H-NMR, spectroscopy data 300a-c for a biofluid sample of an individual, the 1 H-NMR spectroscopy data 31 OOa-c being indicative of an NMR signal intensity as a function of the chemical shift. Optionally, step S1 may comprise accessing the 1 H-NMR spectroscopy data at and / or retrieving the 1 H-NMR spectroscopy data 300a-c from a data storage 116 of the computing device 100 and / or from an external computing device 500.

[0126] In a further step S2, the method comprises evaluating and / or analyzing, with at least one trained machine learning model 114 of the computing device, the 1 H-NMR spectroscopy data 300a- c in terms of a molecular profile 31 Oa-c of the biofluid sample, the molecular profile 31 Oa-c containing all hydrogen peaks in the 1 H-NMR spectroscopy data 300a-c associable to one or more of a metabolite, a protein, an amino acid, a micro molecule and a macromolecule contained in the biofluid sample.

[0127] The method further comprises a step S3 of classifying, based on the evaluating with the at least one trained machine learning model 114, the molecular profile 310a-c into at least a first class and a second class of molecular profiles 310a-c, the first class being representative of molecular profiles 310a associated with healthy individuals and the second class being representative of molecular profiles 310b, c associated with individuals having a cancerous and / or pre-cancerous disease.

[0128] In a further step S4, the method comprises determining, based on the classifying of the molecular profile 310a-c, a cancer risk score indicative of a probability for cancer occurring at the individual. Optionally, step S4 may include computing a probability of cancer occurring at the individual. Further optionally, the cancer risk score and / or contextual information may be outputted by the computing device 100, e.g. at the user interface 118.

[0129] As mentioned above, each of the 1 H-NMR spectroscopy data 300a-c shown in Figures 3B- 3D may be used as input to determine a corresponding cancer risk score for the corresponding individual and / or biofluid sample. When using the spectrum 300a of a healthy individual, the cancer risk score may indicate a high probability above about 50% that the molecular profile 310a is associated with the first class of molecular profiles of a healthy or non-cancerous individual. On the other hand, when using one of the 1 H-NMR spectroscopy data 300b, c of Figures 3C and 3D, the cancer risk score may indicate a high probability above about 50% that the molecular profile 310b, c is associated with the second class of molecular profiles of an individual having a cancerous and / or pre-cancerous disease. Optionally, different classes for cancerous and pre-cancerous disease may be utilized and the cancer risk score may indicate whether the individual has a cancerous disease, such as pancreatic cancer illustrated in Figure 3D, or a pre-cancerous disease, such as pancreatitis illustrated in Figure 3C. For example, a cancer risk score above 0.5 may indicate cancerous disease, whereas a cancer risk score below 0.5 may indicate a healthy individual and / or an individual having a pre-cancerous disease.

[0130] By means of the method disclosed herein, for example, pancreatic cancer Stage l-IV can be differentiated from chronic pancreatitis in a binary classification with over 95% accuracy in 10-fold cross validation.

[0131] In the following, the method of determining the cancer risk score, aspects related thereto and advantages thereof are summarized. The method disclosed herein can accurately detect cancers, cancerous diseases and / or pre-cancerous diseases that are normally diagnosed when it is too late, and survival rates are low. The method involves a combination of high-frequency 1 H-NMR (nuclear magnetic resonance) spectroscopy and machine learning applied on a biofluid sample, e.g. a blood sample. A cancer risk score is generated as output for the different cancerous and / or pre-cancerous diseases to allow for earlier detection when compared to other known techniques. In contrast to currently used or known techniques, the method disclosed herein may not only allow to infer information about the health status of the individual from known cancer-induced signals or molecular signatures, but also from molecular signatures that have not been identified yet. Furthermore, the method described herein is fast, scalable, and inexpensive, therefore it is very suitable for industrial use.

[0132] Cancer cells can influence the overall molecular composition of biofluids, such as e.g. blood, due to their altered metabolism and protein composition. This influence can include changes in the proportion of specific proteins, metabolites, micro- and / or macromolecules that are either directly released by the cancer cells or whose release is directly or indirectly induced by them, e.g. via biological networks, as e.g. described with reference to Figure 2. To detect this change, high- frequency 1 H-NMR spectroscopy and / or 1 H-NMR spectroscopy 200, 300a-c data is used, which allows to deliver information about the chemical environment of each hydrogen atom in the biofluid sample, thus indirectly giving information about every single molecule and its concentration. Depending on the structure of the molecule that contains the corresponding hydrogen atom, a different chemical environment may be present, thus changing the chemical shift of this specific hydrogen atom. For these reasons, 1 H-NMR spectroscopy data 200, 300a-c of the biofluid sample can give an extremely high information density that can be used to differentiate between healthy individuals and individuals having a cancerous and / or pre-cancerous disease.

[0133] Currently, analysis methods are limited to signals that can be assigned to a specific molecule, for example a metabolite or a lipoprotein. The great advantage of the method described herein may be that all hydrogen peaks in the 1 H-NMR spectrum are considered and thus the entire molecular profile 31 Oa-c of the individual’s biofluid sample can be taken into account. Already at early stages, the cancer cells can influence the metabolism and molecular profile, which can be reliably detected by means of the method of the present disclosure.

[0134] Early cancer detection is one of oncology's biggest challenges and therefore a dynamic, fastgrowing field. With the method described herein high-risk patients or individuals can be screened with predispositions for traits of cancer. Therefore, patients or individuals may provide a biofluid sample, e.g. blood sample, at the doctor, HCP or hospital that then can be analyzed with the method described herein.

[0135] For the NMR analysis, the biofluid samples can be analyzed with a high-frequency NMR spectrometer with a pulse sequence similar to CPMGPR1 D. The free induction decay (FID) can be converted to a spectrum via Fourier Transformation, normalized to a reference peak (lactate, anomere D-Glucose duplet, etc.), the baseline can be corrected and the spectrum can be converted to binned increments of smaller or equal to 0.02 ppm, preferably less than or equal to about 0.01 , even more preferably less than or equal to about 0.006 ppm, for example about 0.00016 ppm, as shown in Figures 3B to 3D. This allows to get a significant amount of data for the deciphering of the molecular composition or profile 31 Oa-c of the biofluid sample and can give information about all hydrogen atoms present, wherein each bin may represents a specific position in the spectrum.

[0136] Whilst most other cancer or multi-cancer detection approaches use liquid biopsy in combination with next-generation sequencing technologies to detect cancer, the method according to the present disclosure is not focusing on nucleotide sequences but rather on the molecular profile 31 Oa-c of the individual’s biofluid sample using 1 H-NMR spectroscopy, machine learning, and optionally biological reasoning. In particular, hydrogen peaks that cannot be assigned to specific metabolites, proteins, amino acids, micro- and / or macromolecules, but to an overlap of multiple molecules, proteins, nucleic acids, and / or other components or constituents of the biofluid sample,can be utilized to determine the cancer risk score. Due to this, the method according to the present disclosure can provide improved sensitivities and specificities of over 95% (in both preliminary and pilot study data), compared with around 60-80 % for other multi-cancer detection approaches, or tumor markers with even lower accuracy.

[0137] The method according to the present disclosure can allow for the differentiation of healthy and sick patients or individuals, respectively individuals having a cancerous and / or pre-cancerous disease. For training the at least one ML model 114, 1-HNMR spectroscopy data 200, 300a-c of a training data set of preferably a plurality of individuals can be normalized. Afterwards, feature selection and engineering for features of the at least one ML model 114 can be conducted. For example, one or more regions in and / or bins of chemical shift may be selected, based on the normalized 1 H-NMR spectroscopy data 300a-c, as features of the ML model 114, which may provide a difference between healthy and cancer / pre-cancer samples and thus may represent the molecules, proteins, amino acids, micro- and / or macromolecules that reflect or indicate changes between healthy or non-cancerous individuals and cancerous / pre-cancerous individuals.

[0138] Optionally, methods such as principal components (PCA) may be used to create additional engineered and / or extracted features for the at least one ML model 114. The identified features may contain information from metabolites, amino acids, proteins, micro- and / or macromolecules. After feature selection, for example by a p- value based filter, the at least one ML model 114 can be trained on the selected, engineered and / or extracted features.

[0139] A machine learning approach is used to train the at least one ML model 114 that can predict the molecular profile 31 Oa-c of the healthy / non-cancerous and the diseased / cancerous / pre- cancerous patients or individuals. Using the molecular profile 31 Oa-c can enable the detection of cancer signals or signatures from the individual’s biofluid sample. The at least one ML model 114 can be validated through cross-validation and optionally tested on a separate test dataset of 1 H- NMR spectroscopy data 200, 300a-c of one or more further individuals, e.g. different than those considered in the training data set.

[0140] Accordingly, the ability of high-frequency NMR spectroscopy can be combined with a wellthought through machine-learning approach and feature identification to detect highly complex correlations in the 1 H-NMR spectroscopy data 200, 300a-c, which may be emblematic of the individual’s biofluid sample and / or the corresponding molecular profile 31 Oa-c, which can allow for the earliest possible detection of a cancerous and / or pre-cancerous disease.

[0141] For example, by outputting and / or computing one or more classification results and / or probabilities for the respective classes, a cancer risk score can be determined that can inform about the extent to which the molecular profile 31 Oa-c resembles a certain health state. Based on the cancer risk score, information on how likely the predicted health state is can be provided.

[0142] Optionally, by adjusting cancer risk score cut-off values, the at least one ML model 114 can be adjusted for patient or individual groups and desired false-positive and / or false-negative ratios, for example based on an Area-Under-The-Curve-AUC-receiver operator curve (AUC-ROC).

[0143] Figure 4 shows a flow chart illustrating steps of a computer-implemented method of determining a cancer risk score according to an exemplary embodiment. The method of Figure 4may be performed by the computing device 100 described with reference to Figure 1. Specifically, Figure 4 illustrates training of the ML model 114 of the computing device 100.

[0144] In a first step 400, raw 1 H-NMR spectroscopy data is acquired at an NMR frequency at or above about 500 MHz, e.g. with a pulse sequence similar to CPMG and / or CPMGPR1 D.

[0145] In step 402, the 1 H-NMR spectroscopy data may be aligned and / or spectral processing may be applied, such as e.g. one or more of chemical shift referencing, phasing, and baseline correction of the raw 1 H-NMR spectroscopy data may be applied.

[0146] In a further step 404, the raw 1 H-NMR spectroscopy data may be binned to incremental bins of the chemical shift, each bin having a width of less than or equal to about 0.02 ppm, preferably less than or equal to about 0.01 , even more preferably less than or equal to about 0.006 ppm, for example about 0.00016 ppm.

[0147] At step 406, the binned 1 H-NMR spectroscopy data may be normalized using probabilistic quotient normalization. At step 408, labels may be assigned to the binned 1 H-NMR spectroscopy data, such as 0 and 1 for binary classification.

[0148] At step 410, feature selection, identification and / or extraction may be performed for features of the ML model 114 as described above. For instance, KBest may be utilized for feature selection.

[0149] At optional step 412, bin features may be dropped based on ppm range, unsupervised clustering features may be performed, undersampling and / or oversampling may be performed.

[0150] At step 414, scaling may be applied, e.g. with StandardScaler.

[0151] At step 416, the ML model 114 may be trained for binary or multiclass classification using a statistical machine learning algorithm, such as e.g. VotingClassifier, Ensemble methods, Logistic Regression, or others.

[0152] At step 418, the trained ML model can be used in inference to generate a cancer risk score, e.g. indicative of a probability between 0 and 1 for binary classification. Therein, the cancer risk score can act as computational biomarker to inform about pathogenic molecular signatures in an individual’s biofluid sample.

[0153] Figure 5 shows a flow chart illustrating steps of a computer-implemented method of determining a cancer risk score according to an exemplary embodiment. The method of Figure 5 may be performed by the computing device 100 described with reference to Figure 1. Specifically, Figure 5 illustrates inference of the ML model 114 of the computing device 100.

[0154] In a first step 500, raw 1 H-NMR spectroscopy data is acquired at an NMR frequency at or above about 500 MHz, e.g. with a pulse sequence similar to CPMG and / or CPMGPR1 D.

[0155] In step 502, the 1 H-NMR spectroscopy data may be aligned and / or spectral processing may be applied, such as e.g. one or more of chemical shift referencing, phasing, and baseline correction of the raw 1 H-NMR spectroscopy data.

[0156] In a further step 504, the raw 1 H-NMR spectroscopy data may be binned to incremental bins of the chemical shift, each bin having a width of less than or equal to about 0.02 ppm, preferably less than or equal to about 0.01 , even more preferably less than or equal to about 0.006 ppm, for example about 0.00016 ppm.

[0157] At step 506, the binned 1 H-NMR spectroscopy data may be normalized using probabilistic quotient normalisation At step 508, scaling may be applied, e.g. with StandardScaler.

[0158] At step 510, the trained ML model 114, e.g. trained according to the method of Figure 4, can be used to generate a cancer risk score, e.g. indicative of a probability between 0 and 1 for binary classification. Therein, the cancer risk can act as computational biomarker to inform about pathogenic molecular signatures in an individual’s biofluid sample. For instance, one of the 1 H-NMR spectroscopy data 310a-c shown in Figures 3B to 3D may used as input at step 510 to compute or determine the corresponding cancer risk score.

[0159] Figure 6 shows a flow chart illustrating steps of a computer-implemented method of determining a cancer risk score according to an exemplary embodiment. The method of Figure 6 may be performed by the computing device 100 described with reference to Figure 1. Also, the method shown in Figure 6 can be utilized to supplement step S3 of the method shown in Figure 3. Accordingly, the method of Figure 6 may include steps S1 to S4 of the method of Figure 3, wherein step S3 may be supplemented with the detailed steps shown in and described with reference to Figure 6.

[0160] At step 600 of the method of Figure 6, step S3 of classifying the molecular profile 31 Oa-c into at least the first class and the second class of molecular profiles 31 Oa-c comprises classifying, based on evaluating the 1 H-NMR spectroscopy data 300a-c in terms of the molecular profile 31 Oa-c of the biofluid sample with a first trained machine learning model, the molecular profile 31 Oa-c into at least a healthy class of molecular profiles 310a associated with healthy individuals and a non- healthy or diseased class of molecular profiles 310b,c associated with non-healthy or diseased individuals. For reasons of clarity and simplicity, a classification into the healthy class is shown with reference numeral 602 in Figure 6, and a classification into the non-healthy or diseased class is shown with reference numeral 604 in Figure 6.

[0161] Further, the method comprises at step 606, upon determining at step 604 that the molecular profile 310b, c of the biofluid sample is associated with or is in the non-healthy or diseased class of molecular profiles 310b, c, classifying, based on evaluating the 1 H-NMR spectroscopy data in terms of the molecular profile of the biofluid sample with a second trained machine learning model, the molecular profile into a non-cancer class 310a of molecular profiles associated with non-cancerous individuals and a cancer class of molecular profiles 310b, c associated with cancerous and / or pre- cancerous individuals. For reasons of clarity and simplicity, a classification into the non-cancer class is shown with reference numeral 608 in Figure 6, and a classification into the cancer class is shown with reference numeral 610 in Figure 6.

[0162] Optionally, a first cancer risk score indicative of a probability for the molecular profile 310a being associated with or in the healthy class and / or the non-healthy or diseased class may be determined at steps 600-604. Further optionally, a second cancer risk score indicative of a probability for the molecular profile 310b, c being associated with or in the non-cancer class and / or the cancer class of molecular profiles 310b, c may be determined at steps 606-610. Further optionally, the cancer risk score, respectively a final cancer risk score may be determined at steps 606-610, e.g. based on the first and second cancer risk scores.

[0163] Further, the method comprises at step 612, upon determining that the molecular profile of the biofluid sample is associated with or is in the cancerous class of molecular profiles, classifying, based on the evaluating 1 H-NMR spectroscopy data in terms of the molecular profile of the biofluid sample with a third trained machine learning model, the molecular profile into a plurality of classes of molecular profiles, each class being associated with a particular type of cancerous disease and / or pre-cancerous disease, as indicated by reference numerals 614a, 614b, 614c in Figure 6.

[0164] Optionally, a third cancer risk score indicative of a probability for the molecular profile being associated with a particular type of cancerous, such as pancreatic cancer, and / or pre-cancerous disease, such as pancreatitis, may be determined. Accordingly, a type of cancer, cancerous disease and / or pre-cancerous disease may be determined by means of the method described herein.

[0165] While the invention has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive; the invention is not limited to the disclosed embodiments. Other variations to the disclosed embodiments can be understood and effected by those skilled in the art and practicing the claimed invention, from a study of the drawings, the disclosure, and the claims.

[0166] As used herein, the word “comprising” does not exclude other elements or steps, and the indefinite article “a” or “an” does not exclude a plurality. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage. Any reference signs in the claims should not be construed as limiting the scope.

[0167] As used herein, the phrase “being indicative of’ may mean “reflecting” and / or “comprising”. Accordingly, an entity / element referred to herein as “being indicative of [...]” can be synonymously or interchangeably used herein with one or both said entity / element “comprising [...]” and said entity / element “reflecting [...]”.

[0168] Furthermore, the terms first, second, third or (a), (b), (c) and the like in the description and in the claims are used for distinguishing between similar elements and not necessarily for describing a sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances and that the embodiments of the invention described herein are capable of operation in other sequences than described or illustrated herein.

[0169] In the context of the present invention any numerical value indicated is typically associated with an interval of accuracy that the person skilled in the art will understand to still ensure the technical effect of the feature in question. As used herein, the deviation from the indicated numerical value is in the range of ± 10%, and preferably of ± 5%. The aforementioned deviation from the indicated numerical interval of ± 10%, and preferably of ± 5% is also indicated by the terms “about” and “approximately” used herein with respect to a numerical value.

Claims

CLAIMS1 . A computer-implemented method of determining a cancer risk score for an individual, the method comprising: obtaining, at a computing device (100), Hydrogen-1 Nuclear Magnetic Resonance, 1 H-NMR, spectroscopy data (300a-c) for a biofluid sample of an individual, the 1 H-NMR spectroscopy data (300a-c) being indicative of an NMR signal intensity as a function of the chemical shift; evaluating, with at least one trained machine learning model (114) of the computing device (100), the 1 H-NMR spectroscopy data (300a-c) in terms of a molecular profile (310a-c) of the biofluid sample, the molecular profile containing all hydrogen peaks in the 1 H-NMR spectroscopy data associable to one or more of a metabolite, a protein, an amino acid, a micro molecule and a macromolecule contained in the biofluid sample; classifying, based on the evaluating with the at least one trained machine learning model (114), the molecular profile (310a-c) into at least a first class and a second class of molecular profiles, the first class being representative of molecular profiles (310a) associated with healthy individuals and the second class being representative of molecular profiles (310b, c) associated with individuals having a cancerous and / or pre-cancerous disease; and determining, based on the classifying of the molecular profile, a cancer risk score indicative of a probability for cancer occurring at the individual.

2. The method according to the preceding claim, wherein determining the cancer risk score includes computing and / or determining the probability for cancer occurring at the individual based on the at least one trained machine learning model (114) of the computing device (100).

3. The method according to any one of the preceding claims, wherein classifying the molecular profile (310a-c) includes determining and / or computing a classification result indicative of a probability for the first class and / or the second class; and wherein the cancer risk score is determined based on the classification result.

4. The method according to any one of the preceding claims, wherein the at least one machine learning model (114) is trained to determine a probability for binary and / or multiclass classification of the molecular profile (310a-c) of the biofluid sample.

5. The method according to any one of the preceding claims, wherein the biofluid sample is selected from the group consisting of a blood serum sample, a blood plasma sample, a blood sample, a urine sample, and a cerebrospinal fluid sample.

6. The method according to any one of the preceding claims, wherein the 1 H-NMR spectroscopy data (300a-c) is indicative of a NMR signal intensity in binned increments of the chemical shift, each increment having a width of less than or equal to about 0.02 ppm, preferablyless than or equal to about 0.01 , even more preferably less than or equal to about 0.006 ppm, for example about 0.00016 ppm.

7. The method according to any one of the preceding claims, further comprising: converting the 1 H-NMR spectroscopy data (300a-c) into a binned data structure based on mapping the NMR signal intensity to binned increments of the chemical shift, each increment having a width of less than or equal to about 0.02 ppm, preferably less than or equal to about 0.01 , even more preferably less than or equal to about 0.006 ppm, for example about 0.00016 ppm.

8. The method according to any one of the preceding claims, further comprising: binning the 1 H-NMR spectroscopy data (300a-c) into a plurality of incremental bins of the chemical shift, each bin having a width of less than or equal to about 0.02 ppm, preferably less than or equal to about 0.01 , even more preferably less than or equal to about 0.006 ppm, for example about 0.00016 ppm; and optionally normalizing the binned 1 H-NMR spectroscopy data (300a-c) using probabilistic quotient normalization.

9. The method according to claim 8, wherein the binning is conducted over a spectrum from -1 ppm to 10 ppm of the chemical shift.

10. The method according to any one of the preceding claims, wherein the at least one trained machine learning model (114) is a machine-learned classifier configured to process, as input data, the 1 H-NMR spectroscopy data (300a-c) in a data structure of binned increments of the chemical shift.11 . The method according to any one of the preceding claims, wherein the at least one machine learning model (114) is trained and / or configured to identify at least one molecular signature of cancerous and / or pre-cancerous disease in the molecular profile (310a-c) of the biofluid sample.

12. The method according to any one of the preceding claims, wherein the determined cancer risk score is usable as computational biomarker indicative of a pathogenic molecular signature and / or a molecular signature of cancerous and / or pre-cancerous disease in the molecular profile (310a-c) of the biofluid sample.

13. The method according to any one of claims 11 and 12, wherein identifying the at least one molecular signature includes identifying one or more bins in the 1 H-NMR spectroscopy data (300a-c) associated with and / or containing one or more hydrogen peaks indicative of a cancer-induced change in the molecular profile (310b, c) with respect to molecular profiles (310a) associated with healthy individuals.

14. The method according to any one of claims 11-13, wherein identifying the at least onemolecular signature includes identifying one or more bins and / or a range of chemical shift in the 1 H- NMR spectroscopy data (300a-c) associated with an overlap of hydrogen peaks assignable to a plurality of metabolites, proteins, amino acids, micro-molecules and macromolecules contained in the biofluid sample; and / or wherein identifying the at least one molecular signature includes identifying one or more bins and / or a range of chemical shift in the 1 H-NMR spectroscopy data (300a-c) non-assignable to single metabolites, proteins, amino acids, micro molecules, and macromolecules contained in the biofluid sample.

15. The method according to any one of claims 11-14, wherein the at least one molecular signature is associated with a cancer-induced change in the proportion of one or more hydrogen peaks in the 1 H-NMR spectroscopy data (300a-c) associated with one or more metabolites, proteins, amino acids, micro molecules and macromolecules contained in the biofluid sample with respect to a biofluid sample of one or more healthy individuals.

16. The method according to any one of the preceding claims, wherein evaluating the 1 H-NMR spectroscopy data (300a-c) in terms of the molecular profile (31 Oa-c) includes analyzing all hydrogen peaks in the 1 H-NMR spectroscopy data (300a-c) associated with one or more metabolites, proteins, amino acids, micro molecules and macromolecules contained in the biofluid sample.

17. The method according to any one of the preceding claims, wherein evaluating the 1 H-NMR spectroscopy data (300a-c) in terms of the molecular profile (31 Oa-c) includes detecting a pattern of hydrogen peaks associated with a cancer-induced change in the proportion of one or more hydrogen peaks in the 1 H-NMR spectroscopy data (300a-c) with respect to a biofluid sample of one or more healthy individuals; and / or wherein the at least one machine learning model is trained to detect a pattern of hydrogen peaks associated with a cancer-induced change in the proportion of one or more hydrogen peaks in the 1 H-NMR spectroscopy data (300a-c) with respect to a biofluid sample of one or more healthy individuals.

18. The method according to any one of the preceding claims, wherein obtaining the 1 H-NMR spectroscopy data (300a-c) includes acquiring raw 1 H-NMR spectroscopy data with an NMR spectrometer at a frequency above about 500 MHz.

19. The method according to the preceding claim, further comprising spectral processing of the raw 1 H-NMR spectroscopy data and / or peak aligning the raw 1 H-NMR spectroscopy data.

20. The method according to the preceding claim, wherein spectral processing includes one or more of chemical shift referencing, phasing, and baseline correction of the raw 1 H-NMR spectroscopy data.21 . The method according to any one of the preceding claims, further comprising one or more of normalizing, scaling, binning and filtering of the raw 1 H-NMR spectroscopy data and / or the 1 H-NMR spectroscopy data.

22. The method according to any one of the preceding claims, wherein the at least one machine learning model (114) is trained for classification using one or more statistical machine learning algorithms, preferably based on one or more of a voting classifier, an ensemble method, and logistic regression.

23. The method according to any one of the preceding claims, wherein the classifying of the molecular profile (310a-c) into at least the first class and the second class of molecular profiles comprises: classifying, based on evaluating the 1 H-NMR spectroscopy data in terms of the molecular profile of the biofluid sample with a first trained machine learning model, the molecular profile into at least a healthy class of molecular profiles (310a) associated with healthy individuals and a non- healthy class of molecular profiles (310b, c) associated with non-healthy individuals; and upon determining that the molecular profile of the biofluid sample is associated with or in the non- healthy class of molecular profiles (310b, c), classifying, based on evaluating the 1 H-NMR spectroscopy data in terms of the molecular profile of the biofluid sample with a second trained machine learning model, the molecular profile into a non-cancer class of molecular profiles (310a) associated with non-cancerous individuals and a cancer class of molecular profiles associated with cancerous and / or pre-cancerous individuals (310b, c). wherein the first trained machine learning model and the second machine learning model differ from one another.

24. The method according to the preceding claim, further comprising: determining, based on the classifying of the molecular profile (310a-c) of the biofluid sample with the first trained machine learning model, a first cancer risk score indicative of a probability for the molecular profile (310a-c) being associated with or in the healthy class and / or the non-healthy class; determining, based on the classifying of the molecular profile of the biofluid sample with the second trained machine learning model, a second cancer risk score indicative of a probability for the molecular profile (31 Oa-c) being associated with or in the non-cancer class and / or the cancer class of molecular profiles; and determining the cancer risk score based on the first and second cancer risk scores.

25. The method according to any one of claims 23 and 24, further comprising: upon determining that the molecular profile of the biofluid sample is associated with or in the cancerous class of molecular profiles, classifying, based on the evaluating 1 H-NMR spectroscopydata in terms of the molecular profile of the biofluid sample with a third trained machine learning model, the molecular profile (310a-c) into a plurality of classes of molecular profiles, each class being associated with a particular type of cancerous disease and / or pre-cancerous disease; and optionally determining a third cancer risk score indicative of a probability for the molecular profile being associated with or in a particular type of cancerous and / or pre-cancerous disease.

26. The method according to any one of the preceding claims, further comprising: providing and / or adjusting a cut-off value of the cancer risk score to produce a desired ratio of false-positives and false-negatives.

27. The method according to any one of the preceding claims, further comprising: classifying the molecular profile (310a-c) using a three-stage network approach, wherein the first stage determines whether the molecular profile (31 Oa-c) falls within the first class or the second class, the second stage determines whether the molecular profile (31 Oa-c) is malignant or non- malignant and classifies the molecular profile (31 Oa-c) into a malignant class or non-malignant class, and the third stage determines a type of cancer of the molecular profile (31 Oa-c) and classifies the molecular profile (31 Oa-c) into a class corresponding to the type of cancer.

28. The method according to any one of the preceding claims, wherein the 1 H-NMR spectroscopy data (300a-c), prior to evaluation, is normalized using probabilistic quotient normalization.

29. The method according to any one of the preceding claims, wherein the method further comprises: determining, based on the evaluating of the 1 H-NMR spectroscopy data (300a-c) in terms of a molecular profile (31 Oa-c) of the biofluid sample and / or based on the classifying of the molecular profile (31 Oa-c), whether the cancer risk score indicative of the probability for cancer occurring at the individual is associated with a malignant disease or a non-malignant disease.

30. A computer program, which, when executed by a computing device (100), instructs the computing device (100) to perform the method according to any one of the preceding claims.31 . A non-transitory computer-readable medium storing a computer program according to the preceding claim.

32. A computing device (100) configured to perform the method according to any one of claims 1 to 29.

Citation Information

Patent Citations

  • Machine learning detection of hypermetabolic cancer based on nuclear magnetic resonance spectra

    WO2023196571A1