Liquid biopsy

A novel PCR-based liquid biopsy method targeting specific leukocyte RNA biomarkers addresses the limitations of current diagnostic tools for adenocarcinomas, offering sensitive, cost-effective, and rapid early-stage detection.

WO2025109299A1PCT designated stage expired Publication Date: 2025-05-30BIOMAVERICKS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/GB2024/052623
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-23
Filing Date
2024-10-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Current diagnostic methods for adenocarcinomas, particularly liquid biopsies, face challenges such as low sensitivity, high cost, and long turnaround times, leading to missed early detections and ineffective disease monitoring.

Method used

A novel liquid biopsy method utilizing PCR to quantify leukocyte RNA biomarkers, specifically targeting PTPRC, IGF2R, ITGAX, TNFRSF1A, and PAICS, to differentiate healthy individuals from patients with adenocarcinomas, offering a cost-effective and rapid diagnostic tool.

Benefits of technology

The method achieves sensitive and specific detection of adenocarcinomas, enabling early-stage diagnosis and potentially reducing mortality rates by providing a non-invasive and cost-effective diagnostic solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2024052623_30052025_PF_FP_ABST
    Figure GB2024052623_30052025_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein are methods for diagnosing cancer in an RNA sample obtained from a subject through determining the expression levels of a panel of biomarkers, as well as sets of primers and kits for performing the method. The panel of biomarkers comprises sequences within each of the leukocyte mRNA biomarkers BIOM01 (PTPRC; CD45), BIOM17 (IGF2R; CD222), BIOM19 (ITGAX; CD11C), BIOM24 (TNFRSF1A; CD120A) and BIOM75 (PAICS) and the cancer is selected from the group consisting of: Bile Duct, Breast, Cervix, Colorectum, Oesophagus, Lung, Mouth, Ovary, Pancreas, Prostate, and Stomach.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] LIQUID BIOPSY Field of Invention The present invention pertains to the field of medical diagnostics, particular to molecular diagnostics for early detection of adenocarcinomas. It is especially relevant to non-invasive diagnostic tools utilising leukocyte-derived RNA biomarkers for the detection and differentiation of various early-stage cancers, such as bile duct, breast, cervical, colorectal, oesophageal, lung, mouth, ovarian, pancreatic, prostate, and stomach cancers. Furthermore, the invention pertains to the genomics field, specifically concerning the design of distinctive primer sets targeting the identified RNA biomarkers. The identified biomarkers originate from transcripts produced by leukocytes present in blood samples. The designed primer sets, and their application enhance the early detection of these cancers. Reference to Sequence Listing This application contains a Sequence Listing in computer readable ST26 form, which is incorporated herein by reference. Background The increasing global burden of cancer, which is projected to rise from 19 million cases in 2020 to 29 million by 2040, brings with it an urgent call to action in the healthcare sector (Gordon-Dseagu & Vlad 2023). With cancer positioned as the second leading cause of mortality worldwide after cardiovascular diseases, attention to the various types and subtypes of cancer contributing to these alarming statistics is crucial. Adenocarcinomas, a broad category of cancers arising from glandular cells, account for 52% of the total cancer cases (Sung et al., 2021; PMID: 33538338). These cancers predominantly affect the bile duct, breast, cervix, colorectum, oesophagus, lung, mouth, ovary, pancreas, prostate, and stomach. The high prevalence and diversity of adenocarcinomas underscore the necessity for targeted research and tailored interventions. The development of these cancers often spans 8-10 years and is often influenced by multiple factors. On the one hand, extended exposure to modifiable risk factors such as smoking, obesity, and alcohol abuse are known precursors to adenocarcinoma. On the other hand, inherited genetic risk factors, such as BRCA1 / 2 mutations (breast, prostate) and Lynch (colorectal), Peutz-Jeghers (pancreatic), and Plummer-Vinson (oesophageal) syndromes can predispose individuals to these cancers. This intricate interplay between lifestyle and genetics underlies the transition from normal cells to malignant tumours. As we explore the diagnostic landscape of adenocarcinomas, it becomes apparent that early detection plays a vital role in prognosis. Some types of adenocarcinomas, such as breast cancer, often exhibit noticeable symptoms in early stages, enabling prompt surgical or therapeutic intervention that leads to favourable outcomes. Unfortunately, most others, like pancreatic cancer, remain asymptomatic or present nonspecific symptoms until they reach an advanced stage. These cancers' silent nature poses a significant challenge to early detection and effective treatment. The ability to detect such cancers in their early stages would confer numerous benefits. According to Blackford et al. (2020; PMID: 31958122), the mortality rates for adenocarcinomas such as pancreatic cancer vary significantly based on the stage of detection, ranging from 17% in stage 1 to 97% in advanced stage IV. Therefore, enhancing the development and implementation of early-stage diagnostic tools is key to reducing mortality rates, improving prognosis, and alleviating the burden on global healthcare systems. Diagnostic methodologies are a cornerstone in the management of adenocarcinomas, with liquid biopsy emerging as a promising non-invasive approach for early cancer detection. This technique allows molecular profiling of diseases, offering a cost-effective and less invasive alternative to imaging approaches such as MRI, CT, PET scans, and endoscopic ultrasound. Laboratory tests in liquid biopsy can detect various components in the blood like tumour antigens, circulating tumour DNA, circulating tumour cells (CTCs), and tumour-derived exosomes, which carry tumour-specific proteins and genetic material (Grunvald et al., 2020, PMID: 33081107; Lin et al., 2021, PMID: 34803167; Dai et al., 2020, PMID: 32759948). These components provide critical insight into the tumour’s characteristics, laying the groundwork for personalised therapy. The working principle of these technologies capitalises on the biological processes associated with tumour growth and metastasis. As tumours grow and undergo epithelial-mesenchymal transition, cells are sloughed off into the bloodstream, becoming CTCs. Concurrently, abnormal cell growth causes cell apoptosis, releasing genomic DNA into the blood. Both CTCs and cell-free DNA bear tumour-specific signatures, setting them apart from their normal counterparts. The same principle extends to exosomes, extracellular vesicles containing tumour-specific protein and genetic materials. However, these components exist in low concentrations in the bloodstream, necessitating the need for enrichment procedures. Heterogeneity in tumour cells and individual patient differences pose challenges to canonical enrichment approaches, such as antibodies targeting EPCAM+ cells or 5-hydroxymethylation DNA. These limitations become especially evident in late-stage patients. Even in optimal scenarios, the average sensitivity of these tests sits around 51.5% (Klein et al., 2021, PMID: 34176681), leaving a large margin for missed diagnoses. In addition, the high cost and extended turnaround time of up to 90 days for certain tests, such as those detecting differentially methylated tumour DNA, can serve as potential roadblocks to widespread adoption. When cancers progress to late stages due to missed early detection or ineffective disease monitoring, patients face a grim scenario: average morbidity rates exceeding 80%, limited treatment options, amplified financial burdens, and diminished quality of life. These severe outcomes underline the critical need for innovative, precise, cost-effective, and rapid diagnostic tools for adenocarcinomas to amplify early detection rates, enhance patient outcomes, and reduce the worldwide cancer burden. While tumour heterogeneity and individual variances present barriers to the clinical application of tumour-origin components in the blood, leukocytes provide a promising diagnostic avenue by their response to tumour growth. Through various interactions with tissue-resident immune cells, either direct, indirect via the extracellular matrix, or mediated by soluble molecules like cytokines and chemokines, leukocytes undergo alterations in their gene expression profiles (Baghban et al., 2020; PMID: 32264958). These changes, in turn, initiate the secretion of molecules that recruit more immune cells from peripheral blood, a process prevalent in both inflammation and tumour progression (Masopust & Soerens, 2019; PMID: 30726153; Cotechini et al., 2021; PMID: 33924237). Consequently, the resulting transcriptomic changes provide an opportunity to differentiate early cancer from healthy states, underscoring the potential of immune response in diagnostics irrespective of tumour heterogeneity (Munn & Bronte, 2016; PMID: 26609943). Several studies have found the gene expression profiles in blood samples from cancer patients to be significantly different from those in non-cancerous individuals. Investigations into the feasibility of using RNA-seq of peripheral blood samples, both at single-cell resolution and in bulk measurements, have shown promise in differentiating healthy individuals from cancer cases in the context of many adenocarcinomas. However, the overlap of differentially expressed genes among these cancers limits the effectiveness of statistical analyses of individual cancers without cross-comparison with other types. Despite these challenges, the potential of leukocyte biomarkers for cancer detection remains substantial. They offer the ability to capture the systemic immune response to cancer, potentially increasing the accuracy of the diagnostic test, while providing a less invasive and potentially cost-effective diagnostic method. Yet, the identification of tumour-specific leukocyte biomarkers, especially those quantifiable by commercially viable methods such as qPCR, remains elusive. Our invention directly addresses these challenges by offering a unique method utilising PCR to quantify leukocyte RNA biomarkers. This novel approach effectively differentiates healthy individuals from patients with four different types of adenocarcinomas, paving the way for an in-depth discussion of our approach in the subsequent sections of this patent application. Summary of Invention This invention presents a novel method and its accompanying diagnostic kit, useful for the early detection of cancer, by leveraging distinct biomarker profiles in leukocytes extracted from biofluid samples. Central to this method is the identification and quantification of five novel biomarkers (BIOM01, 17, 19, 24, 75, corresponding respectively to PTPRC (CD45), IGF2R (CD222), ITGAX (CD11C), TNFRSF1A (CD120A), and PAICS). Together, these biomarkers enable the simultaneous detection of various cancers, for example bile duct, breast, cervical, colorectal, oesophageal, lung, mouth, ovarian, pancreatic, prostate, and stomach cancers. The invention provides several advantages: 1. Biomarker Specialisation: The method uniquely harnesses a combination of five distinct biomarkers. When profiled collectively, they enable a comprehensive detection capability spanning eight specific cancers. This multiplexed approach presents a novel paradigm in cancer screening. 2. Biomarker Detection: The method incorporates primers specifically designed to amplify target sequences 1kb upstream from the 5' and 1kb downstream from the 3' of selected biomarker mRNA molecules. These primer pairs uniquely target exon junctions of BIOM01, 17, 19, 24, 75, and specific regions in the housekeeping gene GAPDH. 3. Optimised qPCR Procedure: The method is fine-tuned to enhance the sensitivity and specificity in detecting cancer biomarkers from biofluid samples containing leukocytes. This involves the optimization of various qPCR parameters and considerations of specific isoforms of the mentioned biomarkers. 4. Diagnostic Kit: This invention further provides a diagnostic kit to facilitate the entire process, containing: a. Custom-designed primers for the target biomarkers and reference gene GAPDH; b. Reagents containing reverse transcriptase, DNA polymerase, dNTPs, buffer solutions, and fluorescent dye molecules that bind to double-stranded DNA; c. Guidelines for the usage and interpretation of qualitative and quantitative data. 5. Scaling Up: The method has provisions for scaling up the optimized qPCR procedure, catering to mass testing and larger sample sizes. This incorporates techniques for optimal sample storage, alternative sample collection, cell enrichment, large-scale RNA extraction, premix assembly, and advanced data analysis using neural networks. 6. Predictive Modelling: The method utilises the biomarker profiles not only for diagnosis but also for predicting the disease's progression and aggressiveness by comparing patient profiles with references. Accordingly, in a first aspect, the invention provides a method for diagnosing cancer in a subject, comprising: a) providing an RNA sample obtained from said subject; b) determining the expression levels of each of a panel of biomarkers in the sample, c) comparing determined expression levels against predetermined criteria, and d) diagnosing the subject with cancer based on said comparison, wherein the panel of biomarkers comprises sequences within each of PTPRC (CD45), IGF2R (CD222), ITGAX (CD11C), TNFRSF1A (CD120A), and PAICS. The method may provide a diagnosis of cancer or non-cancerous. Alternatively, the diagnosis may provide a diagnosis of cancer type wherein the cancer is selected from the group consisting of: bile duct, breast, cervical, colorectal, oesophageal, lung, mouth, ovarian, pancreatic, prostate, and stomach. In this way, the biomarkers allow stratification or characterisation of patients by cancer type. In a second aspect, the invention provides a set of primers for the amplification of the selected biomarker sequences in mRNA molecules, comprising primers specific sequences within BIOM01, BIOM17, BIOM19, BIOM24, and BIOM75. This set of primers may find utility in the method of the first aspect. In a third aspect, the invention provides a cancer diagnostic kit, comprising the set of primers, which finds use in these methods. The kit may comprise the PCR environment described below. Also described herein is a PCR environment comprising one or more wells, wherein one or more qPCR reactants are affixed to a surface of the well. In some embodiments, Taq polymerase and / or reverse transcriptase enzymes are affixed to the surface of the well. Additionally or alternatively, affixed to the surface of the well may be a primer or primer pair which targets a biomarker sequence within a biomarker selected from PTPRC (CD45), IGF2R (CD222), ITGAX (CD11C), TNFRSF1A (CD120A), or PAICS mRNA, as described herein. In some embodiments, the PCR environment comprises at least five wells, wherein each of the at least five wells in turn has affixed to a surface thereof a primer or primer pair which targets a biomarker sequence within PTPRC (CD45), IGF2R (CD222), ITGAX (CD11C), TNFRSF1A (CD120A), and PAICS mRNA respectively. The PCR environment may additionally comprise a well wherein, affixed to the surface thereof, is a primer or primer pair which targets a biomarker sequence within a housekeeping or reference gene e.g. GAPDH as described herein. In some embodiments, the PCR environment comprises multiple “repeats” of wells containing primers for each of the biomarker panel and GAPDH (i.e. a PCR environment comprising 6 or more wells total would have one well comprising primers of each of PTPRC (CD45), IGF2R (CD222), ITGAX (CD11C), TNFRSF1A (CD120A), PAICS, and GAPDH mRNA; a PCR environment comprising 12 or more wells total would have two wells comprising primers of each of the aforementioned panel; a PCR environment comprising 18 or more wells total would have three wells comprising primers of each of the aforementioned panel; etc). The surface is preferably the bottom of the wells. This environment can be used in the methods herein, and simplifies the process by avoiding the need for primer design or aliquoting of multiple reagents. The environment may be used with a single premix comprising free nucleotides and buffer, in which can be combined with sample mRNA and loaded into the PCR environment. Also provided is a method for characterising or determining cancer type in a subject having, or known or suspected to have, cancer, comprising: a) providing an RNA sample obtained from said subject; b) determining the expression levels of each of a panel of biomarkers in the sample, c) comparing determined expression levels against predetermined criteria, and d) determining the cancer type in the subject based on said comparison, wherein the panel of biomarkers comprises sequences within each of PTPRC (CD45), IGF2R (CD222), ITGAX (CD11C), TNFRSF1A (CD120A), and PAICS. In some embodiments, the cancer type is selected from the group consisting of: bile duct, breast, cervical, colorectal, oesophageal, lung, mouth, ovarian, pancreatic, prostate, and stomach. Also provided is a method for determining cancer type in a population of subjects having, or known or suspected to have, cancer, comprising: providing a population of RNA samples obtained from said population of subjects, and, for each subject; determining the expression levels of each of a panel of biomarkers in the sample, comparing determined expression levels against predetermined criteria, and determining the cancer type in the subject based on said comparison, wherein the panel of biomarkers comprises sequences within each of PTPRC (CD45), IGF2R (CD222), ITGAX (CD11C), TNFRSF1A (CD120A), and PAICS; and wherein the cancer type is selected from the group consisting of: bile duct, breast, cervical, colorectal, oesophageal, lung, mouth, ovarian, pancreatic, prostate, and stomach. Also provided is a method for predicting or determining cancer progression in a subject, comprising: a) providing an RNA sample obtained from said subject; b) determining the expression levels of each of a panel of biomarkers in the sample, c) comparing determined expression levels against predetermined criteria, and d) predicting that the subject will undergo cancer progression, or determining that the subject has undergone cancer progression, based on said comparison, wherein the panel of biomarkers comprises sequences within each of PTPRC (CD45), IGF2R (CD222), ITGAX (CD11C), TNFRSF1A (CD120A), and PAICS; and wherein the cancer is selected from the group consisting of: bile duct, breast, cervical, colorectal, oesophageal, lung, mouth, ovarian, pancreatic, prostate, and stomach. Also provided is a method for determining whether a subject previously diagnosed with cancer has undergone relapse or recurrence of their cancer, comprising: a) providing an RNA sample obtained from said subject; b) determining the expression levels of each of a panel of biomarkers in the sample, c) comparing determined expression levels against predetermined criteria, and d) determining that the subject has undergone relapse or recurrence of their cancer based on said comparison, wherein the panel of biomarkers comprises sequences within each of PTPRC (CD45), IGF2R (CD222), ITGAX (CD11C), TNFRSF1A (CD120A), and PAICS, and wherein the cancer is selected from the group consisting of: breast, cervical, colorectal, lung, mouth, ovarian, pancreatic, prostate, and stomach. The subject may have been previously diagnosed as in remission or as substantially or completely cancer-free. Description of FiguresFigure 1 , and75. Individual data points within each group represent different samples, categorised broadly into healthy individuals and various cancer cases. The Y- as a metric for expression variations among the isoforms. Figure 2. Boxplot representing the Coefficient of Variation (CV) of biomarker levels within individuals, stratified by control status (CTRL) or cancer type (breast, cervix, colon, lung, mouth, ovary, pancreas, prostate, stomach). Each point on the plot corresponds to a patient. Figure 3. Boxplot illustrating the Coefficient of Variation (CV) of biomarker levels within individuals, grouped by individual biomarkers (BIOM01, 17, 19, 24, and 75). Each point represents a specific patient's CV for that biomarker. -G. Individual data points within each group represent different samples, categorized broadly into healthyindividuals and various cancer cases. The Y- ing as a metricfor expression variations among the isoforms. The distinct distributions highlight the pronounced differences between BIOM01 and BIOM24, pivotal biomarkers for our early cancer detection methodology. Detailed Description Provided is a method of diagnosing cancer in a subject. As used herein, “diagnosing” encompasses determining the disease state of a subject. It includes determining whether a subject has or lacks a disease, i.e. cancer, or a specific cancer as outlined herein, or whether a subject is likely to have (i.e. has an increased probability of having) said disease. Also contemplated is the diagnosis of a specific disease in a subject known to have a different, related, and / or general disease, e.g. diagnosing a specific cancer type (bile duct, breast, cervical, colorectal, oesophageal, lung, mouth, ovarian, pancreatic, prostate, and stomach, etc) in a patient known to have a (non-specific, or different) cancer. In this sense, the method may be used to determine cancer type in a subject. Diagnosis may be prognostic, and may determine or predict the likely course or outcome of a medical condition or treatment. Diagnosis may also include determining disease progression, for example determining the progression, relapse, or remission of cancer in a subject. The methods may be employed in the context of disease surveillance, optionally for patients previously diagnosed as having a cancer (who may subsequently be in remission). The methods described herein may be utilised to guide treatment. For example, a method of treating cancer in a subject may comprise diagnosing cancer in the subject as described herein, followed by administering a treatment. Diagnosis may include selecting a subject for treatment or prophylaxis. The methods relate to the diagnosis of cancer. Cancer may be metastatic. Cancer may be stage I, stage II, stage III, or stage IV. Advantageously, the methods of the invention allow early detection of cancers, for example, at stage III or earlier, stage II or earlier, or stage I. In some cases, the cancer may be at a “precancerous” or “non-invasive” stage, also referred to as stage 0, dysplasia or “carcinoma in situ”. Preferably, cancer is an adenocarcinoma. Specific cancer types, or specific cancers, of relevance to the invention are bile duct, breast, cervical, colorectal, oesophageal, lung, mouth, ovarian, pancreatic, prostate, and stomach. The biomarker panel allows the diagnosis of these specific cancers. A subject in accordance with the present disclosure may be any animal. In some embodiments a subject may be mammalian. In some embodiments a subject may be human. In some embodiments a subject may be a non-human animal, e.g. a non-human mammal. The subject may be male or female. The subject may be a patient. The patient may have a disease / condition described herein. A subject may have been diagnosed with a disease / condition described herein, i.e. cancer or a specific cancer described herein (bile duct, breast, cervical, colorectal, oesophageal, lung, mouth, ovarian, pancreatic, prostate, and stomach), may be suspected of having said disease / condition, or may be at risk from developing said disease / condition. The subject may have been diagnosed with a cancer, e.g. a specific cancer described herein, and have subsequently been diagnosed as in remission, as cancer-free, or as having no evidence of disease (NED). In these instances, the methods of the invention may be employed as cancer surveillance, e.g. to determine whether the subject has relapsed. The subject may have been diagnosed with one specific cancer described herein (e.g. breast) and subsequently be diagnosed with a different cancer described herein (e.g. cervical). The method involves providing an RNA sample obtained from the subject. The RNA sample may be previously obtained from the subject, or may be actively obtained from the subject, e.g. in a previous step. In some embodiments, the method involves providing or obtaining an RNA-containing sample obtained from a subject, and extracting the RNA from it, so as to produce an RNA sample. An RNA sample may be extracted from any source within the subject, and RNA extraction may be performed through any method known in the art. The methods may use an RNA sample derived from leukocytes and, as such, the RNA sample may be obtained from a source containing leukocytes. RNA samples may be extracted one or more biofluids, for example a biofluid selected from blood (e.g. peripheral blood), urine, saliva, cerebrospinal fluid, pleural effusion, ascites, and / or any other diagnostic biofluid. Additionally or alternatively, RNA samples may be derived from solid tumour tissues, for example solid tumour tissues that have been processed to enrich tumour-infiltrating leukocytes. Suitable processing methods include both enzymatic digestion and non-enzymatic methods. RNA samples for use in the method may be pooled from more than one source. Preferably, the RNA sample is obtained from peripheral blood. Peripheral blood may be enriched for peripheral blood mononuclear cells (PBMCs), e.g. one or more of lymphocytes (T cells, B cells, NK cells) and monocytes, and / or other leukocyte subtypes prior to RNA extraction. Enrichment may be achieved through centrifugation and isolation of the PBMC and / or other leukocyte subtype fraction. Similar enrichment steps may be performed on samples obtained from other sources, in order to increase the number of leukocytes, PBMCs and / or other leukocyte subtypes in the sample. In some alternatives, reverse transcription is conducted on the RNA sample to produce cDNA. This cDNA sample may be used in place of the RNA sample in the methods. Following RNA sample provision, the method involves determining the expression levels of each of a panel of biomarkers in the sample. To this end, the present invention identified a series of biomarkers which are capable of diagnosing and categorising cancer in a subject. These biomarkers are sequences within PTPRC (CD45), IGF2R (CD222), ITGAX (CD11C), TNFRSF1A (CD120A), and PAICS. Exemplary sequences within these genes for use as biomarkers include SEQ ID NO:3, 6, 9, 12, and 15, which are referred to as BIOM01 (or M1), BIOM17 (or M2), BIOM19 or (M3), BIOM24 (or M4), and BIOM75 (or M5) respectively. In some embodiments, the panel of biomarkers comprises sequences within one or more, two or more, three or more, four or more, or all five of PTPRC (CD45), IGF2R (CD222), ITGAX (CD11C), TNFRSF1A (CD120A), and PAICS. Additional biomarkers may be included. In some embodiments, the panel of biomarkers comprises sequences within PTPRC (CD45), IGF2R (CD222), ITGAX (CD11C), TNFRSF1A (CD120A), and PAICS. In further embodiments, the sequence length within each biomarker can range from at least 10 codons to 900 codons. For example, the sequence length can be at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, or 850, codons, to 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, or 900 codons. Accordingly, the sequence length of each biomarker can be no more than about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, or 900 codons. Accordingly, the sequence length of each biomarker can be about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, or 900 codons. To determine appropriate codon lengths, the skilled person would review pre-mRNA and mature mRNA transcript lengths for of the selected leukocyte genes. For example, in the case of PTPRC (CD45), IGF2R (CD222), ITGAX (CD11C), TNFRSF1A (CD120A), and PAICS known mRNA transcript sizes can be found in Table 1 below: Table 1: Leukocyte biomarker transcript sizes: Leukocyte biomarker Known transcript sizes (bp) 1074, 5333, 5135, 4991, 4850, 5357, 5213, 5159, 5015, 4874, 1395, 3427, 3536, 5066, PTPRC (CD45) 5377, 5179, 5035, 5213, 5159, 5015, 3427, 3536 IGF2R (CD222) 14061, 10459, 10459 4663, 4092, 3147, 2560, 2052, 2006, 4036, ITGAX (CD11C) 3147, 2560, 2052, 2006, 4036 TNFRSF1A (CD120A) 2097, 2017, 2171, 2289 898, 814, 786, 634, 511, 425, 398, 425, PAICS 432, 425, 898, 511, 398, 425 The biomarker panel may comprise one or more biomarker sequences SEQ ID NO: 3, 6, 9, 12, and / or 15, or derivatives or variants thereof. A derivative or variant may comprise or consist of a sequence with at least 70%, at least 80%, and least 90% or at least 95% homology or identity to a reference sequence e.g. SEQ ID NO:3, 6, 9, 12, and / or 15. A derivative or variant may comprise a sequence with 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more deletions, insertions, and / or substitutions relative to a reference sequence e.g. SEQ ID NO:3, 6, 9, 12, and / or 15. The expression levels may be determined through any known means, including for example quantitative PCR (qPCR), RNASeq, and / or RNA microarrays. A preferred method is qPCR, as it offers a combination of specificity with affordability. In this method, the reverse transcription is conducted on the extracted RNA to produce cDNA, and a qPCR reaction is performed on the cDNA using primer pairs specific to each of the panel of biomarkers. The expression levels may be normalised to a reference control. The reference control used is a “housekeeping gene” which does not change its expression levels across sample or treatment groups. This allows normalisation to account for different quantities of RNA in tested samples, and provides a comparator for expression levels in order to identify relative up or down regulation (collectively “changes in expression”). Exemplary housekeeping genes in mammals and / or humans include -2-microglobulin (B2M), TATA box binding protein (TBP), Glyceraldehyde-3-phosphate dehydrogenase (GAPDH), Glucuronidase, beta glycosidase (GUSb), Ribosomal protein, large p2 (RPLP2), Actin, beta (ACTB), Ribosomal RNA 18S (18S), Ubiquitin c (UBC), Phospholipase A2 (YWHAZ), ATP synthase subunit (ATP5B), Calnexin (CANX), Cytochrome c-1 (CYC1) and succinate dehydrogenase complex subunit A (SDHA). GAPDH is preferred, but the skilled person will readily appreciate that the identity of the housekeeping gene is less important than the fact that it possesses stable expression, and the present methods may be adapted to use any housekeeping gene known to the skilled person as a reference control. In particular, qPCR data may be expressed as values. This method directly uses the threshold cycle (CT, i.e. the cycle at which the fluorescence level reaches a certain threshold) information generated from a qPCR system to calculate relative gene expression in target and reference samples, using a reference control as the normalizer. This method is well known in the art (see, for example, Livak KJ, Schmittgen TD. Methods. Vol.25. San Diego, CA: 2001. Analysis of relative gene expression data using real-time quantitative PCR and the 2(-Delta Delta C(T)) Method; pp.402–408). In some embodiments, such as in qPCR, the biomarkers are detected through the use of specific primers. Primers amplify a target sequence within the mRNA transcript or cDNA thereof. The primer sequences may be selected to amplify a target sequence spanning around 2kb, 1.9kb, 1.8kb, 1.7kb, 1.6kb, 1.5kb, 1.4kb, 1.3kb, 1.2kb, 1.1kb, 1kb, 0.9kb, 0.8kb, 0.7kb, 0.6kb, or 0.5kb upstream from the 5' and around 2kb, 1.9kb, 1.8kb, 1.7kb, 1.6kb, 1.5kb, 1.4kb, 1.3kb, 1.2kb, 1.1kb, 1kb, 0.9kb, 0.8kb, 0.7kb, 0.6kb, or 0.5kb downstream from the 3' of the target sequence. Preferably, the primer sequences may be selected to amplify a target sequence spanning around 1kb upstream from the 5' and around 1kb downstream from the 3' of the target sequence. Exemplary target sequences are outlined in Table 3. Primer sequences may be selected to target exon junctions within the target transcript. Advantageously, this avoids amplification of gDNA, and may allow resolution of different splice variants. Primer sequences may be selected to yield amplicons from about 50 to about 500 base pairs (bp) in length. Amplicons may be about 50 to about 450 bp in length, about 50 to about 400 bp in length, about 60 to about 350 bp in length, about 70 to about 330 bp in length, about 80 to about 320 bp in length, about 80 to about 317 base pairs (bp) in length. Preferred amplicons may be no smaller than about 50, 60, 70 or 80 bp. Preferred amplicons may be no larger than about 500, 450, 400, 350, 340, 330, 320, 317, 310, or 300 bp. In some embodiments, primer sequences may be selected to bind to one or more regions within each mRNA leukocyte biomarker. In some embodiments, primer sequences bind to a least one region. In another embodiment, primer sequences bind to a least two regions. In some instances, wherein primers bind to two or more regions, the regions may be distally located and non-overlapping. In other instances, the regions may be adjacent and non- overlapping. In further instances, the regions may partially or entirely overlap. Primers sequences are preferably selected so as to have a Guanine-Cytosine (GC) content between 40-65%, and melting temperature (Tm) values ranging from about 58 to about 62°C. Furthermore, primers that possess self-complementarity or exhibit 3' end complementarity less than 6 nucleotide bases may be especially preferred. Exemplary target sequences and primer pairs are outlined in Table 3. BIOM01 may be detected through use of a forward primer comprising a sequence selected from SEQ ID NO:1, 19, 20, and 21 and / or a reverse primer comprising a sequence selected from SEQ ID NO:2, 22, 23, and 24, or any primers overlapping / having homology therewith. In particular, different BIOM01 isoforms may be detected by primer pairs comprising SEQ ID NO:1 and 2, SEQ ID NO:20 and 24, SEQ ID NO:19 and 22, SEQ ID NO:19 and 23 or SEQ ID NO:19 and 23. BIOM17 may be detected through use of a forward primer comprising a sequence of SEQ ID NO:4 and / or a reverse primer comprising a sequence of SEQ ID NO:5. BIOM19 may be detected through use of a forward primer comprising a sequence of SEQ ID NO:7 and / or a reverse primer comprising a sequence of SEQ ID NO:8. BIOM24 may be detected through use of a forward primer comprising a sequence selected from SEQ ID NO:10, 25, 26 and 27, and / or a reverse primer comprising a sequence selected from SEQ ID NO:11, 28, 29, 30, 31 and 32. In particular, different BIOM24 isoforms may be detected by primer pairs comprising SEQ ID NO:25 and 28, SEQ ID NO:25 and 29, SEQ ID NO:26 and 28, SEQ ID NO:26 and 29, and SEQ ID NO:26 and 30. BIOM75 may be detected through use of a forward primer comprising a sequence of SEQ ID NO:13 and / or a reverse primer comprising a sequence of SEQ ID NO:14. Further primers may be used, including derivatives of those outlined above which comprise 1, 2, 3, 4, 5 or more substitutions, additional residues, or deletions relative to the sequences listed above, whilst retaining specificity. The biomarkers may relate to specific isoforms of PTPRC (CD45), IGF2R (CD222), ITGAX (CD11C), TNFRSF1A (CD120A), and / or PAICS. Specific isoforms may be splice variants, SNPs, mutations, or similar. Detection, e.g. primers, may be able to distinguish between expression of two or more isoforms. For example, primers may have a target sequence present in one isoform as the result of mutation splice variants at exon bridges but absent in another. For example, the biomarker panel may include a sequence specific to the RA, RB, RC and / or RO isoform of BIOM01. In particular, the RA and RC isoforms, and especially the RC isoform, is contemplated. Alternatively or additionally, the biomarker panel may include a sequence specific to the isoforms 1, 2, 3, 4 and / or 5 of BIOM75. The methods may employ primers specific to any of these isoforms as described herein. In some embodiments, the biomarker panel distinguishes between the RA and RC isoforms of BIOM01. The methods may employ primers specific to any of these isoforms as described herein. The determined expression levels are compared against predetermined criteria, so as to diagnose the subject with cancer based on this comparison. The “predetermined criteria” are a signature of determined expression levels of the biomarkers that have previously been determined, or predicted, to correspond with a specific diagnosis. These may be experimentally or algorithmically determined. Predetermined criteria may be used to stratify subjects into one or more groups, with a corresponding disease categorisation. GAPDH as a reference control, are as follows: H1) Samples where BIOM01 < 2.4, BIOM17 < 2.4, BIOM19 < 2.4, BIOM24 lies between 1.2-3.2, and BIOM75 < 1.2 are categorised as healthy. Samples which do not meet this criterion (i.e. where BIOM01 2.4, BIOM17 , BIOM19 2.4, BIOM24 ) can be categorised as cancerous. H2) pre-categorised healthy samples where BIOM01 < 1, BIOM17 < 1, BIOM19 < 1, BIOM01 + BIOM17 + BIOM19 > 0.2, and BIOM24 > 1 or the standard deviation (SD) of BIOM01, 17, 19, 24, and 75 < 0.9, BIOM01 + BIOM17 + BIOM19 + BIOM24 + BIOM75 > 2 are re-categorised as cancerous. Pre-categorised cancerous samples where BIOM17 < 1, BIOM19 < 1, BIOM24 < 1, BIOM75 < 1 and BIOM17 + BIOM19 + BIOM24 + BIOM75 < -1 are re-categorised as healthy. Predetermined criteria H1 and H2 may provide useful diagnosis alone, stratifying patients into cancer or healthy groups without further subcategorization. Additional predetermined criteria may be used to subcategorise cancer by type as follows: C1) samples where BIOM01 lies between 5-7, BIOM17 lies between 4-6, and BIOM19 lies between 3-5, and BIOM75 < 1 are categorised as Pancreatic Cancer; C2) samples which do not fall into any of criteria H1-H2 or C1 with BIOM01 and BIOM24 values between 2-6 are categorised as Oesophageal Cancer if SD of BIOM01, 17, 19, 24, and 75 < 1 and BIOM75 lies between 1.2-3.2; and categorised as Lung Cancer if SD of BIOM01, 17, 19, 24, and 75 1, BIOM75 1.2, and / or BIOM75 3.2; C3) samples which do not fall into any of criteria H1-H2 or C1-C2 are categorised as Prostate Cancer if BIOM19 < 0 and BIOM24 < 0; categorised as Stomach Cancer if BIOM01 < 1, BIOM17 > 0.5, BIOM19 > 0.5, BIOM24 > 1; categorised as Ovarian Cancer if BIOM01 > 5, BIOM17 < 4, BIOM19 < 4, BIOM24 > 5; categorised as Lung Cancer if BIOM01, BIOM17, BIOM19, and BIOM24 > 2 and BIOM75 < -1.75; categorised as Mouth Cancer if BIOM01, BIOM17, and BIOM19 < 1, BIOM24 < 2, and BIOM01 + BIOM17 + BIOM19 + BIOM24 > 2.5; categorised as Oesophageal Cancer if BIOM01, BIOM17, and BIOM24 > 6; categorised as Bile Duct Cancer if BIOM01 > 6, BIOM17 > 6, BIOM19 > 2, and BIOM24 < 4; and categorised as Cervical Cancer if BIOM01 < 2, BIOM19 > 1; C4) samples which do not fall into any of criteria H1-H2 or C1-C3 are categorised as Breast Cancer if BIOM75 > 1.5 and BIOM24 < 10, or if BIOM75 lies in 0-1, BIOM24 < 1, and BIOM01 > 0, or if BIOM75 lies in -1-0 and BIOM01 < 0 or BIOM17 < 0; and categorised as Colon Cancer (also referred to as colorectal) if BIOM75 > 1.5 and BIOM24 10, or if BIOM75 lies in 1-1.5, or if BIOM75 lies in 0-1, BIOM24 < 1, and BIOM01 0, or if BIOM75 lies in 0-1 and BIOM24 > 1, or if BIOM75 lies in -1-0, BIOM01 0 or BIOM17 0. These are laid out in table 2 below: Table 2 – Predetermined Criteria Additional criteria Cat. Diagnosis M1 M2 M3 M4 M5 BIOM01 BIOM17 BIOM19 BIOM24 BIOM75 Healthy <2.4 <2.4 <2.4 1.2-3.2 <1.2 - H1 Cancerous 2.4 2.4 2.4 <1.2, >3.2 1.2 - Healthy - <1 <1 <1 <1 M2 + M3 + M4 + M5 < -1 M1 + M2 + M3 > 0.2, SD H2 Cancerous <1 <1 <1 >1 - of M1-M5 < 0.9, or M1 + M2 + M3 + M4 + M5 > 2 C1 Pancreatic 5-7 4-6 3-5 - <1 Not in H1, H2 or C1, SD of C2 Lung 2-6 - - 2-6 <1.2, >3.2 M1-M5 1 Not in H1, H2 or C1, SD of Oesophageal >6 - - >6 1.2-3.2 M1-M5 < 1 Prostate - - <0 <0 - Not in H1, H2 or C1-C2 Stomach <1 >0.5 >0.5 >1 - Not in H1, H2 or C1-C2 Ovary >5 <4 <4 >5 - Not in H1, H2 or C1-C2 Lung >2 >2 >2 >2 <-1.75 Not in H1, H2 or C1-C2 Not in H1, H2 or C1-C2, Mouth >1 >1 >1 <2 - M1 + M2 + M3 + M4 > 2.5 Oesophageal >6 >6 <2 >6 - Not in H1, H2 or C1-C2 Bile Duct >6 >6 >2 <4 - Not in H1, H2 or C1-C2 Cervical <2 - >1 - - Not in H1, H2 or C1-C2 Breast - - - <10 >1.5 Not in H1, H2 or C1-C3 Colon - - - 10 >1.5 Not in H1, H2 or C1-C3 Colon - - - - 1-1.5 Not in H1, H2 or C1-C3 Breast >0 - - <1 0-1 Not in H1, H2 or C1-C3 Colon 0 - - <1 0-1 Not in H1, H2 or C1-C3 Colon - - - >1 0-1 Not in H1, H2 or C1-C3 Colon >0 >0 - - -1-0 Not in H1, H2 or C1-C3 Breast <0 - - - -1-0 Not in H1, H2 or C1-C3 Breast - <0 - - -1-0 Not in H1, H2 or C1-C3 Following diagnosis, a patient may be selected for treatment, prescribed a treatment, and / or administered a treatment. Suitable treatments will be anti-cancer therapies and may be specific to the diagnosis. Sets of primers as used in the method are also provided. These allow the amplification of biomarker sequences and are suitable for use in the detection methods. These sets comprise primers capable of amplifying the biomarker panel described herein. The biomarker sequences are the amplicons of the primers. Sequences for these biomarkers. Especially BIOM01, BIOM17, BIOM19, BIOM24 and BIOM75, are as described herein. The sets of primers include primers (or primer pairs of forward and reverse primers) specific for sequences within one, two, three, four or all five of PTPRC (CD45), IGF2R (CD222), ITGAX (CD11C), TNFRSF1A (CD120A), and / or PAICS. The set may additionally include primers for a reference control or housekeeping gene, such as GAPDH. Specific primer sequences are outlined above. Also provided are cancer diagnostic kits, comprising a set of primers suitable for use in the method. This may include additional reagents for performing the detection method described above, for example agents for conducting reverse transcription (e.g. a reverse transcriptase enzyme, a DNA polymerase, and / or fluorescent dye molecules capable of binding to double- stranded DNA), or quantitative polymerase chain reaction (qPCR) reagents (e.g. dyes, polymerases, nucleotides, reaction premix). The kit may also comprise instructions for use. In some embodiments, the diagnostic kit comprises a well containing one or more reagents selected from one or both of a set of primers relating to a biomarker as described herein, a DNA polymerase (e.g. a Taq polymerase), and / or a reverse transcriptase enzyme. One or more of the reagents may be immobilised to the well, for example at the bottom of the well. The kit may comprise multiple wells, for example comprised within a plate. In this embodiment, it may be desirable for a group of wells to contain different sets of primers relating to different biomarkers, such that a single RNA or cDNA sample may be loaded into the group of different wells and the reaction performed, such that each well amplifies a single biomarker and so that the group as a whole can be used to detect all the biomarkers of the invention. Preferably, a group comprises a well for each biomarker of the invention along with a well containing primers for a reference gene. The kit may comprise multiple groups of wells so that repeat measurements of the same sample and / or multiple samples may be amplified and analysed in parallel. This kit is useful in the methods of the invention. Also provided is use of the cancer diagnostic kit to perform a method of diagnosing cancer in a subject as outlined herein. Accordingly, in one embodiment, the cancer diagnostic kit comprises a method of detecting biomarkers according to the first aspect of the invention; comprising one or more of the following: a. a set of primers according to the second aspect of the invention; b. one or more reagents for conducting reverse transcription, selected from a reverse transcriptase enzyme, a DNA polymerase, and / or fluorescent dye molecules capable of binding to double-stranded DNA; c. one or more quantitative polymerase chain reaction (qPCR) reagents; and / or d. instructions to carry out the methods described above. For convenience, the meaning of certain terms and phrases used in the specification, examples, and appended claims, are provided herein. If there is an apparent discrepancy between the usage of a term in other parts of this specification and its definition provided in this section, the definition in this section shall prevail. The term “about” when referring to a number or a numerical range means that the number or numerical range referred to is an approximation within experimental variability (or within statistical experimental error), and thus the number or numerical range may vary from, for example, between 1% and 15% of the stated number or numerical range. The term “at least” prior to a number or series of numbers is understood to include the number adjacent to the term “at least”, and all subsequent numbers or integers that could logically be included, as clear from context. When at least is present before a series of numbers or a range, it is understood that “at least” can modify each of the numbers in the series or range. Methods described herein may preferably be performed in vitro or ex vivo. The term ‘in vitro’ is intended to encompass procedures performed with cells in culture, the term ‘ex vivo’ is intended to encompass procedures performed outside or separate from intact multi-cellular organisms, whereas the term ‘in vivo’ is intended to encompass procedures with / on intact multi-cellular organisms. The present disclosure includes the combination of the aspects and preferred features described except where such a combination is clearly impermissible or expressly avoided. Throughout this specification, including the claims which follow, unless the context requires otherwise, the word ‘comprise,’ and variations such as ‘comprises’ and ‘comprising,’ will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps. It must be noted that, as used in the specification and the appended claims, the singular forms ‘a,’ ‘an,’ and ‘the’ include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from ‘about’ one particular value, and / or to ‘about’ another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by the use of the antecedent ‘about,’ it will be understood that the particular value forms another embodiment. Examples embodying certain aspects of the invention shall now be described, with reference to the following figures. Further aspects and embodiments will be apparent to those skilled in the art. All documents mentioned in this text are incorporated herein by reference. Examples The foregoing invention is described with reference to the following examples. The current challenge faced by primary care practitioners, clinicians, oncologists, and laboratory professionals is that cancer is often detected at an advanced stage. Traditional diagnostic measures such as endoscopic ultrasound and CT / MRI scans, while useful for advanced cases, are neither cost-efficient nor practical for early screening. Simultaneously, blood tests that detect tumour antigens or circulating DNA, cells, and exosomes show limited sensitivity or specificity, rendering them unsuitable for widespread use. These constraints emphasise an urgent need for reliable, cost-effective, and non-invasive early detection methods for adenocarcinomas such as pancreatic, breast, and colorectal cancer. In response to this need, our invention introduces a novel liquid biopsy test focusing on leukocytes, predominantly found in the peripheral blood. The rationale behind this approach lies in the sophisticated interplay between the tumour microenvironment and the immune system (Baghban et al., 2020; PMID: 32264958). Within this complex network, tumours manipulate tissue-resident immune cells, causing alterations in their gene expression profiles (Munn & Bronte, 2016; PMID: 26609943). These cells then release soluble molecules to recruit immune cells from peripheral blood (Masopust & Soerens, 2019; PMID: 30726153; Cotechini et al., 2021; PMID: 33924237). As a result, these recruited leukocytes exhibit gene expression changes that reflect the presence and progression of the tumour, regardless of the cancer’s origin or stage. The opportunity to capture these early immune responses through our liquid biopsy test offers a promising avenue for early detection and monitoring of adenocarcinomas, paving the way for more effective treatments and improved patient outcomes. Our novel diagnostic test is based on unique oligonucleotides designed to amplify selected regions within distinct transcripts. The process commences with the extraction and reverse transcription of leukocyte RNA into complementary DNA (cDNA). Our uniquely designed primer sets then target and amplify regions of interest within each of the four selected biomarker genes. This leads to the quantitative polymerase chain reaction (qPCR) yielding distinct cycle threshold (Ct) values for each biomarker. After normalising these values to the internal reference gene GAPDH, the normalised Ct values are compared with reference ranges derived from known healthy or cancerous cases. It is the combination of the expression patterns of these biomarkers that allows for the identification of either a healthy status or one of the selected types of cancer referred to above. The unique combination strategy provides the potential for future expansion of the biomarker panel to cover additional cancers. Our diagnostic approach, supported by a statistical interpretation of Ct values, offers a robust alternative to more complex techniques such as next-generation sequencing or intricate machine learning algorithms. This reduction in complexity decreases the dependence on sophisticated, costly equipment, potentially increasing its accessibility across various healthcare settings. Beyond effective early diagnosis, the test can serve as a non-invasive, cost-effective tool for monitoring disease progression in diagnosed patients and for evaluating treatment outcomes. By providing critical and interpretable information, this method aids healthcare providers in making informed decisions regarding further medical interventions, thus facilitating timely adjustments in treatment, and potentially improving patient outcomes. Biomarker Identification Data Acquisition Our quest to discover potential biomarkers for early cancer detection began with an exploration of leukocyte-derived RNA-sequencing (RNA-seq) data from pancreatic cancer patients. The motivation behind this choice was manifold. Firstly, pancreatic cancer presents an exigent need for early detection, as the mortality rate for stage 4 patients exceeds 97%, the highest amongst adenocarcinomas. Furthermore, the immune response to tumour progression is largely conserved across cancer types, implying a shared set of differentially expressed genes (DEGs) in leukocytes of cancer patients versus healthy individuals. Hence, by studying pancreatic cancer, which epitomizes the archetype of an adenocarcinoma, we hypothesized that our findings could apply to a broader cancer landscape. Although our immediate focus has been on bile duct, breast, cervical, colorectal, oesophageal, lung, mouth, ovarian, pancreatic, prostate, and stomach cancers, owing to our sample collection at the time of this patent application, the underlying methodology should apply to additional cancer types as well. To comprehensively capture the transcriptomic landscape, we leveraged both bulk and single- cell RNA-seq datasets. Bulk RNA-seq offers a comprehensive snapshot of the transcriptome by capturing global gene expression profiles, which may be subject to dropouts in single-cell RNA-seq due to its inherently lower sequencing depth, reaching only a few million reads per cell as compared to the extensive coverage of over a hundred million reads in bulk RNA-seq. However, single-cell RNA-seq adds a crucial dimension of resolution, allowing us to identify cell type-specific biomarkers and infer critical cell-cell interactions. Furthermore, single-cell data permits the creation of numerous mini-bulk samples, essential for the robust training of our statistical models. Pre-processing and Quality Control Upon data acquisition, we curated count matrices for each patient. The data were merged and normalised using the Seurat (v.4.0) package in R (v.4.0.2) programming environment. We filtered out cells with a low number of counts and genes, then identified and eliminated doublets. We removed patient variance, which was followed by Uniform Manifold Approximation and Projection (UMAP) visualisation, to classify the data and infer cell type information through an unsupervised approach. Parallelly, we also curated bulk RNA-seq data for validation purposes from previous studies. Using these expression matrices, we merged and normalised the samples, prepared patient phenotype data, and created an ExpressionSet object to store both the count matrix and phenotype data. Unsupervised Clustering For unsupervised clustering, we analysed the top features from each patient’s single-cell data.. This process enabled us to classify the data and infer cell-type information. This set the stage for differential expression analysis and subsequent biomarker selection. Differential Analysis The differential analysis was conducted in a two-pronged manner, first on the single-cell data, followed by a comparison with bulk RNA-seq data, to account for the technological differences inherent in each method. For single-cell data, the ‘FindMarkers’ function embedded in Seurat was utilised, applying the Wilcoxon test algorithm. These leukocyte subtypes were determined by canonical marker expression, including but not limited to CD4.Tcm (IL7R), CD4.Tem (CCR7), CD8.T (GZMK), CD8.Tex (CD8A), and NK.Act. (GNLY). This single-cell data-based analysis identified 3,654 gene candidates. Subsequently, the bulk RNA-seq data were assessed using the limma package. This process began with the normalised matrices, curated from previous studies, and presented as transcripts per million. The next step involved conducting Empirical Bayes Statistics on the linear model fit object, followed by extraction of the top-ranked genes. Interestingly, all 3,654 genes identified in the single-cell dataset were also found to be differentially expressed in the bulk datasets. Equipped with these differentially expressed genes, we prepared to meticulously shortlist the candidates, recognising that statistical significance does not always equate to biological relevance. Shortlisting Candidates To identify key players in tumour progression from the initial pool of 3,654 differentially expressed genes, we targeted and analysed genes implicated in immune-tumour cell interactions. From this analysis we condensed our list of immune genes with substantial involvement in immune-tumour cell interactions. We cross-checked our gene candidates for protein-level differential expression in normal and tumour pancreatic tissues. This rigorous approach narrowed down our list to 32 promising biomarkers, which showed differential expression across immune cell types and were implicated in tumour-immune interactions, consistent at both RNA and protein levels. With this curated list of biomarkers, we further validated in silico using bulk RNA-seq data, which allowed us to assess their predictive potential in a broader context. In Silico Validation of Biomarkers Neural Network Model Creation and Training We trained a neural network model using the R Torch package, with the training data derived from resampled single-cell data of the discovery cohort. Our primary source of training data was derived from single-cell data that had been resampled from the discovery cohort. This approach was chosen to ensure the model had a robust and diverse dataset to learn from. A noteworthy aspect of our model was its specific focus on a set of 32 genes. These genes were selected as variable features, meaning they played a pivotal role during both the training and validation phases of the model. By concentrating on these genes, we aimed to enhance the model’s precision and relevance in its predictions and interpretations. The resampling process resulted in a substantial number of mini-bulk samples, totalling 18,205. Such a significant volume was beneficial as it provided a rich dataset, ensuring the model was exposed to a wide variety of scenarios and patterns. This, in turn, bolstered the training process, making the model more robust and reliable. To ensure a balanced approach and to evaluate the model’s performance accurately, we divided the mini-bulk samples into two equal parts. Half of these samples, amounting to 9,102, were dedicated to training the model. The remaining half was reserved for validation purposes. This division allowed us to train the model effectively and then test its performance on unseen data, ensuring its reliability and accuracy in real-world scenarios. Further, parameters were introduced while training models such as training rate, training cycle, batch size, loss function, validation, and model optimizers to increase the accuracy. The neural network algorithm was introduced with both categorical and continuous variables to find a unique pattern, the data frame was used to separate continuous (x_cont) and categorical (x_cat) input data and the target data (y) where the predicted values are given by the model. Various methods were used to optimise the retrieval of individual samples and the total number of samples. The embedding module with compatible parameters was set which converts categorical data into dense vectors (embeddings). This is useful for handling categorical data in neural networks. The net is the main neural network module. It consists of an embedding layer for categorical data, followed by several fully connected layers. The network ends with a unique function which processes the classification tasks. Model Testing and Validation We developed a neural network model with the aim of achieving high accuracy in predictions. When we evaluated the model’s performance against the validation subset, the results were highly promising. Various strategies were used to evaluate the model performance where the data were PDAC specific samples were split equally with control samples, data were normalised, un-normalised and only using specific data such PDAC only and control only were used to evaluate the model. The model was able to correctly identify all the positive cases, showcasing a sensitivity of 100%. Sensitivity, often referred to as the true positive rate, measures the proportion of actual positives that are correctly identified. This means that our model did not miss out on any positive cases in the validation subset, which is crucial in many medical and diagnostic applications. On the other hand, the model’s specificity was 99.2%. Specificity, sometimes termed the true negative rate, gauges the proportion of actual negatives that are correctly identified. A specificity of 99.2% implies that our model made very few false positive errors, further attesting to its reliability. To ensure that our model’s performance was not just a result of overfitting to a specific dataset, we conducted further tests using an entirely separate cohort. This cohort comprised 101 patients, distinct from our initial dataset. In this external validation, the model continued to perform admirably, achieving a sensitivity of 95%. This indicates that it correctly identified 95% of the positive cases in this new cohort. However, the specificity was slightly lower at 79%, suggesting that while it correctly identified 79% of the negative cases, there were some false positives. This might be due to the low cancer samples and high control samples. These findings formed the basis for our qPCR assay and in-vitro validation on patient samples were performed to validate the findings. Assay Development Biomarker Selection In our pursuit to develop an In Vitro Diagnostic product, the RNA-seq approach enabled the initial identification of biomarkers that could differentiate between healthy and cancerous cases. The qPCR approach was then employed to validate these findings, with primers specifically designed to target the 32 biomarker candidates. The objective was to identify GAPDH and non-cancerous control samples, as these patterns enable precise categorisation of each case into specific cancer types or non-cancer categories. Through rigorous testing of all 32 biomarkers, we pinpointed 5 — BIOM01, BIOM17, BIOM19, BIOM24, BIOM75 — that, when considered together, can robustly distinguish healthy individuals from those with one of four primary adenocarcinomas: breast, colon, lung, and pancreatic cancer. The unique expression patterns of these biomarkers allow for targeted differentiation of these primary cancers. Additionally, preliminary results suggest potential for differentiating other cancers such as cervix, gastric, oral, ovary, and prostate, albeit these findings are based on limited samples and warrant further investigation. BIOM01, also known as CD45, is a surface antigen found ubiquitously on leukocytes and plays a vital role in the regulation of T cell and B cell activation. Functionally, it operates as a type 1 transmembrane protein tyrosine phosphatase, which significantly impacts leukocyte differentiation. This is mainly facilitated through distinct isoforms created by alternatively spliced transcripts in its exons 4, 5, and 6 (Dornan et al., 2002; PMID: 11694532). While upregulation of this biomarker has been associated with conditions like leukaemia and lymphoma (Rheinländer et al., 2018; PMID: 29366662), the downregulation in peripheral blood leukocytes, specifically pertaining to the isoform RC, is a relatively unexplored area. This is particularly interesting because it contrasts the findings of a previous study by Tang et al. (2019; PMID:31169017) that reported higher expression levels in tumour-infiltrating leukocytes. BIOM17 and BIOM24 act as receptors for Insulin Like Growth Factor 2 Receptor (IGF2R) and Tumour Necrosis Factor Receptor Superfamily Member 1A (TNFRSF1A; also known as Tumour Necrosis Factor Receptor 1, TNFR1), respectively. While the roles these receptors play in modulating leukocyte activity during cancer progression remain unclear, insights might be gained from studies involving the loss of function in bladder cancer cells (Liu et al., 2020; PMID: 31847523), or gain of function in renal cancer and melanoma (Takeda et al., 2019; PMID: 31748500). These studies may shed light on the potential roles of these receptors in regulating cell differentiation and proliferation. recognition in leukocytes. This integrin is pivotal in recruiting leukocytes to inflamed tissues (von Andrian et al., 1991; PMID: 1715568), and in modulating phagocytosis. Higher expression levels of these molecules are linked to leukocyte activation (Wen et al., 2022; PMID: 35167661); however, the relationship between downregulation of these molecules and immune suppression, as well as cancer progression, is still an emerging area of research (Fagerholm et al., 2019; PMID: 30837997). BIOM75, identified as phosphoribosyl aminoimidazole carboxylase (PAICS), is pivotal in de novo purine biosynthesis, which is vital for DNA / RNA synthesis. Its role in gastric carcinogenesis, especially in conditions like colon cancer, has been established (Huang et al., 2020, PMID: 32632107; Agarwal et al., 2020, PMID: 32422575). Specifically, it interacts with histone deacetylase (HDAC) 1 / 2, acting as a safeguard against DNA damage response. Furthermore, studies have suggested its involvement in the progression of other cancers, including those of the breast (Gallenne et al., 2017; PMID: 28411283), lung (Goswami et al., 2015; PMID: 26140362), prostate (Chakravarthi et al., 2017; PMID: 27550065), and bladder (Chakravarthi et al., 2018; PMID: 30121007). However, the role of BIOM75 dysregulation in immune cells relating to cancer progression remains ambiguous, as does the correlation between specific expression levels and distinct cancer types. Primer Design We designed primers to target the exon junctions of 32 biomarker candidates. The five biomarkers demonstrating significant diagnostic potential are highlighted at the top of Table 3. The designed primers yielded amplicons ranging from 80 to 317 base pairs (bp) in length, maintaining a Guanine-Cytosine (GC) content between 40-65% and melting temperature I values from 58 to 62°C. Design Process 1) The protein coding region of each selected gene was taken, and their nucleotide sequences were processed using software tools such as IDT Primer Quest, Primer3, and Primer Blast. 2) Specific exon regions for each gene were identified. To improve primer specificity, custom conditions — including amplicon size, GC content, and overlapping exon regions — were set, producing multiple primer sets. 3) Primers spanning multiple exons with the smallest product size were prioritised. 4) Validation of these primers was carried out with NCBI Primer Blast and Primer3, using the FASTA sequence of the specific gene and the sequences of the forward and reverse primers. 5) Primers targeting the intended regions and minimizing potential non-specific amplification from other genes were shortlisted for oligonucleotide synthesis. Final Selection Unspecific amplification can arise due to primers partially matching and binding to unintended regions in the genome. This results in undesired amplicons, which could be less than 500 bp and / or approximate the size of our expected amplicons, potentially confounding our results. Moreover, primers that possess self-complementarity or exhibit 3' end complementarity greater than 6.00 can lead to primer-dimer formations or other artifacts. To ensure accuracy and specificity, primer sets that might produce such unintended results or exhibit these problematic characteristics were not selected for synthesis. Table 3. Primer Design and Amplicon Sequence ID Type Size / bp Sequence (Including primer sequence; F: Forward; R: Reverse) Primers 17; 17 F: CGGACCAATGGCTCTGC (SEQ ID NO:1); R: ACGGTGACAGAGCTGGA (SEQ ID NO:2) BIO Product 114 CGGACCAATGGCTCTGCCCTGCACCTTAGGGCTCGGGATGCTGCTGG M01 CCCTGCCAGGGGCCTTGGGCTCGGGTGGCAGCGCGGAGGACAGCGT GGGCTCCAGCTCTGTCACCGT (SEQ ID NO:3) Primers 20; 21 F: ATGCCTGCCACAGAGATTAC (SEQ ID NO:4); R: TATAGGATGAACCTCCGCTCT (SEQ ID NO:5) BIO Product 111 ATGCCTGCCACAGAGATTACCTGGAAAGTAAAACTTGTTCTCTGAGCG M17 GCGAGCAGCAGGATGTCTCCATAGACCTCACACCACTTGCCCAGAGC GGAGGTTCATCCTATA (SEQ ID NO:6) Primers 20; 22 F: CCGCCAGATATTGCAGAAGA (SEQ ID NO:7); R: CTCTCATAAATGCCTCCTGTCC (SEQ ID NO:8) BIO Product 104 CCGCCAGATATTGCAGAAGAAGGTGTCGGTCGTGAGTGTGGCTGAAA M19 TTACGTTCGACACATCCGTGTACTCCCAGCTTCCAGGACAGGAGGCAT TTATGAGAG (SEQ ID NO:9) Primers 22; 22 F: CTCCAAATGCCGAAAGGAAATG (SEQ ID NO:10); R: ATAATGCCGGTACTGGTTCTTC (SEQ ID NO:11) BIO Product 100 CTCCAAATGCCGAAAGGAAATGGGTCAGGTGGAGATCTCTTCTTGCAC M24 AGTGGACCGGGACACCGTGTGTGGCTGCAGGAAGAACCAGTACCGGC ATTAT (SEQ ID NO:12) Primers 22; 20 F: CAAACAGTCTTATCGGGACCTC (SEQ ID NO:13); R: BIO CTCTCTGCAACCCACTCAAA (SEQ ID NO:14) M75 Product 84 CAAACAGTCTTATCGGGACCTCAAAGAAGTAACTCCTGAAGGGCTCCA AATGGTAAAGAAAAACTTTGAGTGGGTTGCAGAGAG (SEQ ID NO:15) Primers 21; 21 F: GTGGTCTCCTCTGACTTCAAC (SEQ ID NO:16); R: CCTGTTGCTGTAGCCAAATTC (SEQ ID NO:17) HK_ Product 129 GTGGTCTCCTCTGACTTCAACAGCGACACCCACTCCTCCACCTTTGAC GH GCTGGGGCTGGCATTGCCCTCAACGACCACTTTGTCAAGCTCATTTCC TGGTATGACAACGAATTTGGCTACAGCAACAGG (SEQ ID NO:18) Sample Preparation Patient Recruitment and Ethical Approval Patients aged between 18 and 75 diagnosed with a variety of cancers through conventional methods, such as CT / MRI scans and other medical examinations, were recruited for the study. Patients who had undergone surgical and / or other treatments such as chemo / radiotherapy, those with multiple cancers due to metastasis, and other outlier cases were excluded from statistical analysis. All patients provided informed consent for blood sampling. Ethical approval for this study was obtained from our host institute Basavatarakam Indo American Cancer Hospital. Blood Sample Processing and RNA Extraction For each patient, 0.5ml of blood was processed using the QIAamp®RNA Blood Mini Kit (Cat. No.52304 from QIAGEN®). In brief: Sample Preparation: Blood (0.5ml) was mixed with 2.5ml of Buffer BL and incubated on ice for 10-15 minutes. Centrifugation: The sample was centrifuged at 400xg for 10 minutes at 4°C. The supernatant was discarded, retaining the pellet. Pellet Resuspension: The pellet was resuspended in 1ml of Buffer EL and centrifuged again under the same conditions. Lysis: resuspension. Homogenisation: The lysate was transferred to a QIAshredder®spin column, followed by centrifugation at 20,000xg for 2 minutes. Ethanol Treatment: was then transferred to a new QIAamp®spin column for a brief centrifugation. Washing Steps: RNA Elution: of RNase-free water to the QIAamp®membrane and centrifuging at 8,000xg for 1 minute. RNA Quality Assessment The quality of extracted RNA was assessed using NanoDrop®measurements to ensure optimal absorbance ratios for A260 / A280 (around 2.0) and A260 / A230 (between 2.0-2.2). Those with high quality were then subjected to cDNA synthesis. cDNA Synthesis We utilised the Prime Script™1st Strand cDNA Synthesis Kit from Takara Bio®(Cat. No. 6110A) for the synthesis of complementary DNA (cDNA). The procedure is described below: Premix Solution Preparation: o Using the manufacturer's guidelines, a premix was prepared for each RNA sample: diluted using nuclease-free water) Denaturation of Template RNA: o The premix was incubated at 65°C for 5 minutes in a heat block. Immediately following incubation, samples were transferred to ice. First Strand cDNA Synthesis: o To the denatured RNA, the following were added: 0.5 l RNase inhibitorpt RTase -free water o The mixture was subjected to the following thermal cycler program: 30°C for 10 minutes 42°C for 60 minutes 95°C for 5 minutes o Samples were immediately transferred to ice post-incubation. cDNA Quantification and Dilution: o cDNA quantity was assessed using NanoDrop®. o cDNA samples were diluted to their respective concentrations, typically a 1:10 dilution. The resulting cDNA samples were then subjected to the qPCR procedure. qPCR Procedure The qPCR procedure was undertaken using cDNA samples derived from 0.5ml of whole blood. These assays were conducted on the Thermo Scientific®QuantStudio™ 5 instrument, utilising SYBR Green-based chemistry to relatively quantify the targeted biomarkers. We employed the commercial kit from Takara Bio®, specifically the TB Green Premix Ex Taq®(Tli RNase H Plus, Cat. No. RR42WR). For validation purposes, every assay incorporated a reference gene to gauge a relative CT value across both healthy and cancer patient samples. The primers underwent standardisation through gradient PCR and Agarose gel electrophoresis to determine the annealing temperature and optimal primer concentration. The selected annealing temperature, 60°C, proved suitable for simultaneous biomarker studies in either a single tube or a 96-well PCR chosen for the assays. Given that both SYBR Green®and ROX®dye are sensitive to light, reagent preparation was diligently done in dimly lit or dark conditions to maintain fluorescence quality during qPCR. Reaction Mixture Assembly 2O All reactions were assembled on 96-Well Semi-Skirted Plates (Thermo Fisher®, Cat. No. AB0900), kept on ice blocks to mitigate potential priming errors and degradation of reagents. Plate Preparation and Cycling Conditions Once the reaction mixture was set on the 96-well plate, it underwent a bubble check, followed by sealing using Tarsons optical PCR plate sealer. After sealing, the plates were vortexed and centrifuged at 3,000xg for 1 minute. They were then loaded onto the Applied Biosystems®QuantStudio™ 5 Real-Time PCR System (Thermo Fisher®, Cat. No. A34322) with the following qPCR profile: Step 1. Initial Denaturation: 95°C for 3 minutes Step 2. Denaturation: 95°C for 34 seconds Step 3. Annealing: 60°C for 30 seconds Step 4. Extension: 72°C for 30 seconds Step 5. The denaturation to extension steps (steps 2-4) were repeated for 39 times. To inspect for nonspecific amplification, a melt curve analysis was included: 95°C for 15 seconds Ramp down at -0.2°C / second to 60°C and hold for 1 minute. Cooling to ambient room temperature Data Collection and Normalisation Post-amplification, Ct values of selected biomarkers were recorded for each sample. GAPDH served as the reference gene for normalisation. A negative control was also maintained by replacing cDNA with 2ul of ddH2O. Data Interpretation Following the qPCR procedure, data interpretation becomes the next crucial step in deriving meaningful results. Initial Data Pre-processing The raw qPCR output, which is presented as Ct values, undergoes a quality control (QC) step. Samples without Ct values, indicative of potential technical errors, are discarded. Additionally, samples where the Ct value for the endogenous control gene, GAPDH, falls outside the range of 16.0-22.0 are omitted. This stipulation is because the average Ct value for GAPDH is 18.9 (n=12), and a more than 10-fold deviation from this value might suggest cell type enrichment or RNA degradation. Post QC, Ct values of each biomarker are normalised against GAPDH. This normalisation sample’s GAPDH Ct value from the mean Ct value of GAPDH. Then, we calculate the Log2 fold change for each biomarker in comparison to healthy controls. To streamline our description, the biomarkers BIOM01, BIOM17, BIOM19, BIOM24, and BIOM75 will henceforth be referred to as M1 to M5 respectively. Categorisation To elucidate the diagnostic potential of our data, we employ the following procedure to categorise samples: Healthy Subject Classification o H1. Initial Classification – Healthy versus Cancerous Subjects: Criteria for Healthy Classification: A subject is preliminarily classified as healthy if the absolute values of biomarkers M1, M2, and M3 are each less than 2.4, the absolute value of biomarker M4 is between 1.2 and 3.2, and the absolute value of biomarker M5 is less than 1.2. Criteria for Cancerous Classification: Subjects that do not satisfy the healthy classification criteria are preliminarily identified as cancerous. o H2. Reclassification Criteria for Initially Classified Subjects: For Subjects Initially Classified as Healthy: A subject is reclassified as cancerous if the absolute values of biomarkers M1, M2, and M3 are each less than 1, with their sum exceeding 0.2, and biomarker M4 is greater than 1; or if the standard deviation of biomarkers M1 through M5 is less than 0.9, with the sum of their values exceeding 2. For Subjects Initially Classified as Cancerous: A subject is reclassified as healthy if the absolute values of biomarkers M2, M3, M4, and M5 are each less than 1, and their sum is less than -1. Cancer Subtype Classification o C1. Pancreatic Cancer Classification Criteria: A subject is classified as having pancreatic cancer if biomarker M1 has a value between 5 and 7, biomarker M2 has a value between 4 and 6, biomarker M3 has a value between 3 and 5, and the absolute value of biomarker M5 is less than 1. o C2. Initial Classification Criteria for Lung Cancer and Oesophageal Cancer: Lung Cancer: For remaining subjects, they are classified as having lung cancer if M1 and M4 are between 2 and 6, SD is greater than or equal to 1, and the absolute value of M5 is less than or equal to 1.2, or greater than or equal to 3.2. Oesophageal Cancer: For remaining subjects, they are classified as having oesophageal cancer if M1 and M4 are between 2 and 6, SD is less than 1, and the absolute value of M5 is between 1.2 and 3.2. C3. Classification Criteria for Prostate Cancer, Stomach Cancer, Ovarian Cancer, Lung Cancer, Mouth Cancer, Oesophageal Cancer, Bile Duct Cancer, Cervical Cancer: Prostate Cancer: For remaining subjects, they are classified as having Prostate Cancer if M3 is less than 0, M4 is less than 0. Stomach Cancer: For remaining subjects, they are classified as having Stomach Cancer if the absolute value of M1 is less than 1, M2 is greater than 0.5, M3 is greater than 0.5, and M4 is greater than 1. Ovarian Cancer: For remaining subjects, they are classified as having Ovary Cancer if the absolute value of M1 is greater than 5, M2 is less than 4, M3 is less than 4, and M4 is greater than 5. Lung Cancer: For remaining subjects, they are classified as having Lung Cancer if the absolute values of M1, M2, M3, and M4 are all greater than 2, and M5 is less than -1.75. Mouth Cancer: For remaining subjects, they are classified as having Mouth Cancer if the absolute values of M1, M2, and M3 are less than 1, M4 is less than 2, and the sum of M1, M2, M3, and M4 is greater than 2.5. Oesophageal Cancer: For remaining subjects, they are classified as having Oesophageal Cancer if the absolute values of M1, M2, and M4 are greater than 6, and M3 is less than 2. Bile Duct Cancer: For remaining subjects, they are classified as having Bile Duct Cancer if M1 and M2 are greater than 6, M3 is greater than 2, and the absolute value of M4 is less than 4. Cervical Cancer: For remaining subjects, they are classified as having Cervical Cancer if the absolute value of M1 is less than 2, and M3 is greater than 1. o C4. Classification Criteria for Breast Cancer and Colorectal Cancer: Breast Cancer: For any remaining subjects, they are classified as having Breast Cancer if: The absolute value of M5 is greater than 1.5 and M4 is less than 10, or M5 is between 0 and 1, the absolute value of M4 is less than 1, and M1 is greater than 0, or M5 is between -1 and 0 and either M1 is less than 0 or M2 is less than 0. Colorectal Cancer: For any remaining subjects, they are classified as having Colorectal Cancer if The absolute value of M5 is greater than 1.5 and M4 is greater than or equal to 10, or The absolute value of M5 is between 1 and 1.5, or M5 is between 0 and 1 and either the absolute value of M4 is less than 1 and M1 is less than or equal to 0, or the absolute value of M4 is greater than 1, or M5 is between -1 and 0 and either M1 is greater than or equal to 0 or M2 is greater than or equal to 0. These categorisation criteria are summarised in Table 2 With these steps, we have successfully categorised all samples, thus transforming our raw data into actionable diagnostic insights. In Vitro Validation The validity of this method is anchored on its performance during in vitro validation. Through this phase, the technique’s reliability and applicability in a real-world setting are tested, ensuring that the results are both consistent and relevant. 1. Sample Population We undertook a prospective study, where a total of 98 samples were meticulously collected from patients when they were diagnosed with cancers or noncancerous diseases at the Basavatarakam Indo American Cancer Hospital. The rationale behind this approach was to harness real-time diagnostic data, ensuring the accuracy and relevance of the samples in relation to the presence or absence of specific cancers. The collection and use of these samples were carried out with full ethical approval. Table 4. Sample Population Sample Type Label Age Sex Stage Sample Type Label Age Sex StageSAM39 CTRL CTRL 28 Female 0SAM53CASEBreast 52 Female 3SAM40 CTRL CTRL 30 Female 0SAM88CASEBreast 53 Female 3SAM38 CTRL CTRL 50 Female 0SAM68CASEBreast 54 Female 3SAM06 CTRL CTRL 51 Female 0 SAM21 CASE Breast 57 Female 3SAM37 CTRL CTRL 37 Male 0SAM87CASEBreast 57 Female 3SAM05 CTRL CTRL 53 Male 0SAM82CASEBreast 63 Female 3SAM41 CTRL CTRL 57 Male 0 SAM36 CASE Cervix 45 Female 2 SAM07 CTRL CTRL 61 Male 0 SAM14 CASE Colorectum 48 Male 1 SAM42 CTRL CTRL 66 Male 0 SAM23 CASE Colorectum 66 Male 1 SAM15 CASE Bile Duct 60 Male 2 SAM86 CASE Colorectum 34 Female 2 SAM09 CASE Breast 30 Female 1 SAM78 CASE Colorectum 48 Female 2PAT03CASEBreast 45 Female 1SAM16 CASE Colorectum 67 Female 2PAT02CASEBreast 48 Female 1SAM85 CASE Colorectum 80 Female 2SAM64CASEBreast 50 Female 1SAM28 CASE Colorectum 52 Male 2SAM20 CASE Breast 52 Female 1 SAM79 CASE Colorectum 41 Male 2PAT01CASEBreast 52 Female 1SAM66 CASE Colorectum 55 Male 2SAM52CASEBreast 57 Female 1SAM29 CASE Colorectum 35 Female 3SAM72CASEBreast 57 Female 1SAM31 CASE Colorectum 55 Male 3SAM94CASEBreast 59 Female 1SAM10 CASE Colorectum 57 Male 3SAM51CASEBreast 62 Female 1SAM45 CASE Colorectum 60 Male 3SAM71CASEBreast 62 Female 1 PAT05CASEEsophagus 52 Female 1SAM84CASEBreast 62 Female 1 PAT06CASEEsophagus 67 Female 1PAT04CASEBreast 62 Female 1SAM44 CASE Esophagus 54 Male 1SAM93CASEBreast 65 Female 1SAM43 CASE Esophagus 63 Male 2SAM24 CASE Breast 66 Female 1 SAM49 CASE Esophagus 60 Male 3SAM58CASEBreast 31 Female 2 PAT08CASELung 47 Male 1SAM90CASEBreast 32 Female 2 PAT10CASELung 59 Male 1SAM55CASEBreast 33 Female 2 PAT11CASELung 63 Male 1SAM50 CASE Breast 35 Female 2PAT09CASELung 67 Male 1SAM12 CASE Breast 43 Female 2 SAM32 CASE Lung 72 Male 1SAM70CASEBreast 44 Female 2SAM26 CASE Lung 56 Female 2SAM27 CASE Breast 46 Female 2SAM61CASELung 57 Female 2SAM91CASEBreast 46 Female 2 SAM63CASELung 70 Female 2SAM60CASEBreast 48 Female 2SAM18 CASE Lung 72 Female 2SAM89CASEBreast 48 Female 2SAM48 CASE Lung 31 Female 3SAM65CASEBreast 50 Female 2SAM25 CASE Lung 50 Female 3SAM95CASEBreast 55 Female 2SAM30 CASE Lung 57 Female 3SAM33 CASE Breast 56 Female 2 SAM22 CASE Lung 70 Male 3SAM54CASEBreast 60 Female 2SAM11 CASE Mouth 41 Male 1SAM59CASEBreast 64 Female 2SAM19 CASE Ovary 47 Female 3SAM83 CASE Breast 64 Female 2 PAT07 CASE Pancreas 57 Male 1SAM56CASEBreast 66 Female 2 SAM62CASEPancreas 38 Female 3SAM57 CASE Breast 70 Female 2 SAM46 CASE Pancreas 42 Female 3 SAM13 CASE Breast 71 Female 2 SAM17 CASE Pancreas 67 Female 3 SAM74 CASE Breast 80 Female 2 SAM47 CASE Pancreas 70 Female 3SAM96CASEBreast 82 Female 2SAM34 CASE Prostate 69 Male 1SAM73 CASE Breast 46 Female 3 SAM08 CASE Stomach 44 Female 1SAM67CASEBreast 48 Female 3 PAT12CASEStomach 69 Male 1SAM92 CASE Breast 50 Female 3 SAM35 CASE Stomach 38 Female 3 2. Categorisation Outcomes Our categorisation approach took a multi-tiered and iterative path. Each step aimed to ensure more specific and refined classifications: Step 1. Healthy Subject Classification (H1): o Accurately identified controls SAM06, 07, 37, 38, 39, 40, 41, 42 as healthy. o Misclassifications: SAM09, 12 (Breast), SAM10, 29, 31 (Colorectum), SAM11 (Mouth), SAM35 (Stomach), SAM36 (Cervix) were incorrectly identified as healthy; SAM05 (CTRL) was incorrectly identified as cancerous. Step 2. Further Refinement for Healthy Subject Classification (H2): o SAM05 was correctly reclassified as healthy. o SAM09, 12 (Breast), SAM10, 29, 31 (Colorectal), SAM11 (Mouth), SAM35 (Stomach), SAM36 (Cervix) were correctly reclassified as cancerous. Step 3. Determine Pancreatic Cancer (C1): o Identified SAM17, 46, 47, 62, PAT07 as pancreatic cancer (PDAC). Step 4. Pre-determine Lung or Oesophageal Cancer (C2): o Identified SAM18, 22, 25, 30, 32, as lung cancer. o Identified SAM43, 44, PAT05, 06 as oesophageal cancer. Step 5. Determine Prostate, Stomach, Ovarian, Lung, Mouth, Oesophageal, Bile Duct, or Cervical Cancer (C3): o Identified SAM34 as prostate cancer. o Identified SAM08, 35, PAT12 as stomach cancer. o Identified SAM19 as ovarian cancer. o Identified SAM26, 48, 61, 63, PAT08, 09, 10, 11 as lung cancer. o Identified SAM11 as mouth cancer. o Identified SAM49 as oesophageal cancer. o Identified SAM15 as bile duct cancer. o Identified SAM36 as cervical cancer. Step 6. Determine Breast or Colorectal Cancer (C4): o Identified SAM09, 12, 13, 20, 21, 24, 27, 33, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 64, 65, 67, 68, 70, 71, 72, 73, 74, 82, 83, 84, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, PAT01, 02, 03, 04 as breast cancer. o Identified SAM10, 14, 16, 23, 28, 29, 31, 45 , 66, 78, 79, 85, 86 as colorectal cancer.

[0002] 6yphy y y y y y y y ttcelhtlhtlhtlhtlhtlhtlhtlhtl u tstastastastastastastastastststststststststststststst a a a a a a a a a D e e e e e e e eaeaeaeaeaeaeaeaeaeaeaeaeaeaS e e e e e e e e eelirBrBrBrBrBrBrBrBrBr r r r r r r r r r r rerH H H H H H H H HBB B B B B B B B B B B B B 5y tphy y y y y y y y r r r r r r r r r r r r r r r r r r r r r rththth h h h h h cue e e e e e e e e e e e e e e e e e e e e eel l l tl tl tl tl tl tlDcta a a a a a a a a ncancancancancancanc c c c c c c c c c c c c c canananananananananananananananSe e e e e e e e eeliC C C C C C C C CaH H H H H H H H H B C C C C C C C C C C C C C 4y y y y y y y y y rererererererererererererer r r r r r r r r rphetlht ht ht ht ht ht ht htt alalalalalalalalacncncncncncncncncncncncncencencencencence e e e encncncncncnSe e e e e e e e e aCaCaCaCaCaCaCaCaCaCaCa a a a a a a a a a a aH H H H H H H H H C C C C C C C C C C C C 3y y y r r r r r r r r r r r r r r r r r r r r r r rphy y y y y yte lhtlhtlhtlhtlhtlhtlhtlhtlece e e e e e e e e e e e e e e e e e e e e et a ncncncncncncncncncncncnc c c c c c c c c c cS ea a a a a a a aa a a a a a a a a a a anananananananananananaHeHeHeHeHeHeHeHeH C C C C C C C C C C C C C C C C C C C C C C C 2yphy y y y y y y y r r r r rte lhtlhtlhtlhtlhthththtececececercercercercercercercercercercercercercer r r r rcecececececta a a a alalalala nananan n n n n n n n n n n n n n n n n n n nSe e e e e e e e eC C CaCaCaCaCaCaCaCaCaCaCaCaCaCa a a a a a aH H H H H H H H H C C C C C C C 1yphy y y y?y r?r r r r r r r r r r r r r r r r r r?r rththththtree hy yththtec yh e e e e e e e e e e e e e e e e e eyh e etlSal l l l c l l leaHeaHea a n a a a nta lacncancancancancancancancancancanc c c c c c c tananananananana lacncanaHeHeHa e e eC H H H CeH C C C C C C C C C C C C C C C C C CeH C C M6 5 78U 33583 6 8488 367 50639 256575320503900585637585555 690 0S 21.036.51. 50.682.94.28.6.70-1-3-0-1-. 6.71191. 6.81751. 1.596.7 1 7 8 0 5 3 4 5 3 6 9 3 3 1 2 02241..97.23.33.31.20.21.28.14.32.25.29. . 223244.. 8012.2D0S23759409383789250353477554627770418976754 4 9 0 2 8 1 6 8 5 9 4.6 9 9 5 1 3 9 5 6 7 7 0 1 3 4 3 7 9 35013147575206213183816690.0.0.0.1.1.1.0.1.4.0.1.2.5.2.1.5.6.6.4.6.6.2.7.3.5.7.4.7.0.6.157470 3887 73 5717 4 2889 3 2 6 1 3 3 7 3 0 7 0 3014 0 72611 122 0 6 42 471 94 97451675 0 131 0 0 9 6 43726 4M2.00.0 .0. 2-0-.70.0 .0. 2 . 0 5 1 . 1 0 . 5 . .6.4. 24.4.5.8.2. . 66.3.-0-.12-.0.2.32-.1.22-.01-2-2-3-.41-1-2-3-2-22.03-2-4846398 488 433599 3 83 236 5 391 83384 31 3 41559914 544 204 1 13133301560 60597 860 51508 0 4 073 8M4.3.5.. . . 4.. 5. 8. 5.1.9..8. 2.. . . 11.. . 01 1.. 6 . . 2 6 2 . 11 1 11-2-0- 11- 12- 1 4 1010- 21171783 3481.50151.8.7.041.1313 282801237 952 68 0 606 54 3 8 9 1 7 0 5 0 0 4 2 0 2 4 3 7 0 7 2 0 0 060 2024662021034008248763978 5 5 7 8 5 7 4 04 2 667 7M.0 .0.-0.-0-.0 .0-.1 .0-.1.3.0.3.0.3.3.0.3.1.2.22.22.22.92.71.03.63.22.33.2 .06-.61.0no9 7 22 3 6 3 7 4 0 1it2064013 3 32.47 539 7 8 4.2.0.0 02 1 5507 3 1236812384330565677342509869 5 0 7 5 6 5086786822460700 7 0 194781aM.s1.0.01 .01 1 1.2 .6 .0 .1 .0 .6.4. . . . . . . . . . . . . ..0 . .i - - - - - - - 0 7 8 6 6 5 6 0 5 6 7 7 6 5 - 5 0rog1 10 3536 86 608863 37 90 580776539115306029609754747049 36880686062 5 1etaM1. 32.90.09. 50.30.22.19. 60.00.2 2 6 6 9 9 9 2 1 2 5 6 10. 4 3 7 4 00080017.0.6.4.8.4.3.7.5.9.6.1. . 0 . . . . . . . .C- - - - - - 2 7 1 7 7 7 7 7 1 3 3elpml tcaeL L L L L L L L L utst t t t t t t t t t t t t t t t t t t t tasasasasasasasasasasas s s s s s s s s s sSbR. aTRTR R R R R R R De e e e e e e e e e eaeaeaeaeaeaeaeaeaeaeae5 LC CTCTCTCTCTCTCTCeli rBBrBrBrBrBrBrBrBrBrBrBrBrBrBrBrBrBrBrBrBrBrBelbaTep L L L L L L L L L E E E E E E E E E E EyR R R R R R R R R S S S S S S S S S S SESESESESESESESESESESESES TTCTCTCTCTCTCTCTCTCACACACACACACACACACACACACACACACACACACACACACACAC elp9304836073501470245190 3020 4602 10 252749151748 40 394285095505210772mM AM AM AM AM AM AM AM AM AM AM ATATAM AM ATAM AM AMMMMTAMMMMMMMMM a SS S S S S S S S S S S P P S S P S SASASASASPASASASASASASASASAStststststststststststststststst t t t t t t tm xi um tum tum tum tum tum tum tum tum tum tuta a a a as s s s s s s svc c c c c c c c c c cereBreBreBreaBreaBreaBreaBreaBreaBreaBreaBreaBreaBreaBreaBreaBreaBrea a a aBreBreBreBreBrre ere e e e e e e e e eB Corlorororororororororool l l l l l l l l lCoCoCoCoCoCoCoCoCoCoCrererererererer r r r r r r r r r r r r r r r x r r r r r r r r r r rc c c c c c cecececececececececececececece e i e e e e e e e e e e enan n n n n n n n n n n n n n n n n n n ncncnvrc c c c c c c c c c cen n n n n n n n n n nCaCaCaCaCaCaCaCaCaCaCaCaCaCaCaCaCaCaCaCaCaCaC CaCaCaCaCaCaCaCaCaCaCaCrererererererererererererererererererer r r r r r r r r r r r r r r rc c c c c c c c ce e e e e e e e e e e e e e e en n n n n n n n ncncncncncncncncncncncncncncncnc c c c c c c c c c caCaCaCaCaCaCaCaCaCaCaCaCaCaCaCaCaCaCaCaCaCaCaCanCanCanCanCanCanCanCanCanCanCanCaCrerererererererererererererererererererererererererererererererer r rc c c c c c c c c c c c c c c c c c c c c ce e en n n n n n n n n n n nc c c c c c c c c c c c ca a a a a a a a a a a ananananananananananananananan n n n n n n n nC C C C C C C C C C C C C C C C C C C C C C C C C CaCaCaCaCaCaCaCaCaCrerererererererererererererererererererer r r r r r r r r r r r r r rcnc c c c c c c c c c c c c c c c c c cecececececececececececececececanananananananananananananan n n n n n n n n n n n n n n n n n n n nC C C C C C C C C C C C C CaCaCaCaCaCaCaCaCaCaCaCaCaCaCaCaCaCaCaCaCaCrercerncercer r r r r r r r r r r r r r r r r r r?cececececececececececececececece e e eyhrerererererer r r?ce e ey?hyhananananananananananananCanCanCanCanCanCancCancCancCancntlacncncncncn ncncncntlatlaC C C C C C C C C C C CaCaCe a a a a a a a a aH C C C C C C C C CeHeH54640049531 9534826602 937588074589 4 9 3 8 459 2 0 0 904 7 3410.7660 0 3 194 526 8 2 3 65076164 761.733.829.527.11 1.734.6-6.811.728.227.2482. 7.1.0.9.2.6.3.2. .4.4. 10 3 3 4 1 935203152263216212226342.. . . . . 024162028171.. . . 86114202.1 .1-29056899677451900608328442109296354263845191592541861261092294316 4 73 7 8 3 8 8 5 3 8 4 4 8 9 2 2 6 3 8 9 9 57 4 4. . . . . . . . . . .6 3 3 8 6 1 6 6 1 7 5 8 8 77 4 7 4 6 2 3 4 7 5 6.1.6.7.6.4.6.2.4.6.1.8.7.1.5.7.5.5.3.1.1.4.2.0.037749261371 4 3 1 675772373 8 876 376 7 4 3 573 8 3 371.8 5 806 25325872618 585 6 968 7972867 03127 782990 4 549 6 8 873. . . . 1-2-2-2-1-.0 .2. . . . 3-2-2-2-2-.0 .20-..04.122-.3 .1. . 2-0-3-.2 .2.-39-.1 .4.7-2.3-2.1-3.-25-. 01.40.8-1.2-0.-02-.038516.88793288041570 9 0634177752 3 371119025362088 96550268 00651541 3 1 0374807 0 84525343 668564681.. 4881.. 4761.11.06.. . . 58811141.. . . 50617181.. 4861.. . 342151.. . 159161.. .4.1. 6011711121.0 77. 4 6 35.1.101.5.1.0014371 1 08.1118318 4 7 5 5 6 1 6 0 0 7 0 0 1 2 7 9 8 800 8297200241005481787007910 5 1 6 54 1 7 0 5 9 6 8 5 9 092 8 3 0 6 6 1 3 9 7 4322.3.3.1.2 .0-.2.2.3.3.2.0.2.1.2.3.0.31.83.01.14.93.1 .11-.21.22.92.01.52.80.22.13.14.0 .0-545 2 8 5 14 1 5 6 738.853729630176155386633325 7 996.6.6.4.5.0.5.6.6.7.5 .650395645955365605763860148 933.35 6 5 0 0 9 5 1 8 1 038 7.4730666081208446053.1.-.6.3.6.6.4.3.6.5.521.50-.4.8.4.5.4.0.2.4.50-1-084 0 2 012 9 1 7 2 5 7 0 5 008 7 9 1 8 1 9 8 9 3 9895.688.927.653.7 17. 87.0554 7 2 0 0 2 3 5 22-.95.46.13.68.94.43.33.07.88. 6.300 7 4 6 7 1 5 2 9 32214626610546711.45.74.02.46.1 6 9 0 6 6 4 9 7 9 2 713.4.1.0.2.4.2.7.2.4.6.6.0 .1-t t t t t t t t t t t t t t t t t tmmmmmmst t t t tas xi n n n n n ututututututeasreasreasreasreaseaseaseaseaseaseaseaseaseas s s s s s s seaeaeaeaeaeaeaeae vrolo o o oc c c c c col l l leBrBrBrBrBrBrBrBrBrBrBrBrBrBrBrBB B B B e o o o oB B Br r r r rerererererC C C C C ColoolololololCoCoCoCoCoCESESESESESESESESESESESESESESESESESESESESESESESESESESESESESESESESE E EA A AA A A A A A A A A ACACACACACACACACACAS S SC C C C C C C C C C C C CACACACACACACACACACACACACAC1906985659334595386575314769377629358886127828636887589766413261829213MAM AM AM AM AM AM AM AM AM AM AM AM AM AM AM AM AM AM AM AMMMMMMMMMMMMMMMS S S S S S S S S S S S S S S S S S S SASASASASASASASASASASASASASASAS mumstuctusecgusususu s s s s sagre hag g ghahahahgngngngng g g g g g g g g htyr aea a a aet hrererereratchachacaor p p p p p u u u unununununununununuloolo oavc c c c c s mmmosososososL L L L L L L L L L L LuLM OnananananaoroPtoStoStS C Ce e e e eO O O O OP P P P Psrususususu s s s s sercegagagagaga g g g ghggg g ggg g g hncnh h h h h n n n n n n n n n n n n n tuyr a a a a aetaere a chchccrecrecrecrct a a aaCapCopsopopopouLuLuLuLuLuLuLuLuLuLuLuLuL o v n n n n nsomomomoesesesOesM Oa a a a a rPt t tO O Oe P P P P P S S SO sr rusususur r r rs s s s se eg g g g r r r r r r r a a a a a r r r rc ca a a a ececececec gececec gec g g gecece e e e e ece e enanhphphphp n n n n nnun n nnnn n nn nrcrcrcrcrc ncncncnCaCoso o o aeseseseCaCaCaCaCL aCaCauCL auCLuLuL aCaCnanPanPanPanPaa a a aPC C C C O O O Or r r r r r r rs s s s se e e e e e e ererererererererererererer r a a a a a r r r rc c c c c c c c c c ce e e e e e e e e e en nc c c c c c c c c c c r r r r r c c c ca ananananananananananananan n n n n n n n c c c c c n n n nC C C C C C C C C C C C C CaCaCaCaCaCaCaCaCnananananaa a a aP P P P PC C C Crererer r r r r r r r r r r r r r r r r r r r r r r r r r r rc c cececececececececececececece e e e e e e e e e e e e e en n n n n n n n n n n n n n n ncncncncncncncncncncncncnc c caCaCa a a a a a a a a a aCaCaCaCaCaCan n nC C C C C C C C C C CaCaCaCaCaCaCaCaCaCaCaCaC?y r r r r r r r r r r r r r r r r r r?r?hrtle ece e e e e e e e e e e e e e e e eyh erererererererereyhacenancancancancancancancancancanc c c c c c c c tanananananananana l canc c c c c c c c tanananan n n n nlaH C C C C C C C C C C C C C C C C C C Ce a a a a aH C C C C C C C C CeH3810 2382241599866311034217986665976 237 8169 7 6 8 9 4 3112.468.9 7 1 9 9 9 3 7 1 9 6 2 5 7 8 9 712 550431338932572 9.471.84.14.16.12.59.17.14.11.25.13. . . . . 228191712521.. 3921.. 2202.. . . .48. . . 011261615-3171.147225409922743845947763 5 6 9 2 6 3 4 7 3 2 9 8 5 4 3 1 5 7 48 3 8 9 6 4 1 2 0 5 99931592460415123477644882 0 0 1 6 9 6 6. . . . . .1.8.2.2.2.1.2.4.0 5 1 3 5 7 5 50 2 1 1 0 3.4.2.8.2.1.1.0.3.1.4.2.3.2.2.1.1.0757080405767768 3244063876 399985 977736 74 303594 43 02 77 6736 7472038.4 7 0 6 36. 3 8 1 7 041.1.2.3.2.122.4.3.3.3.8 .32.9-2.-22-. 6.03 9.205-. 10.10.8-0.-04-. 61.20-. 60.40.-03-.280.21. 13.0-9083139302 736 3 6 9 5 9 0 2 3 3 3 7 8 9 24.0.9.69.034.7. 1.546 5 7 8 5 6 0 7 8 4 6 9 1 504511874799 6 3 972 9 51.0.5.6.7.3.3.7.0.2 9 9 0 4 9 8 4 7 6 6 . 0 5 11 3 1 3 3 344 3 3 4 2 7 6 7 5.5.4.2.4.1.7.1.7.5.0.44-.5.3.1297687.727115503260 3 0 6 3 7 2 3 5 2 0 2 3 4 6 930589236566414055483438688751 7 74 0 7 1 405 1 8 5 74 1 601 7 4520.2.0.2.2.3.1.2.2.1.2.0.4.2.2.4.2.2.21.52.90.3 .01-.42.04.45.3 .15-.02.2 .0-70 45 53 007419511892 90 85344033925208 196048 69 11 28 771 6 506 8 1 96.043. 1 0 0 9 5 42. 9 1 1 9 9 57. 0 7 11. 65. 66447458.394332-.20-.0.3.3.6.1.10-.2.1.7.4.4.211.3.0.40-.21-.4.4.4.31-.3.2.02626713111650063000583034109780 0 4 9 0 2 8 7 7 1 8 5 8 2 3 66 4 1 6 9 47 7 1 8 8 3 1 0 7 4 4 8 4 6 1 2. . . . . .6.3.0.7.4.3.6.5.5.0.1.8.4.5 7 5 1 4 5 5 5 0 7 0 11 7 4 5 3 5 6 7 7 6 7 3 6 7 7 5 9 2 3.2.0.6.3.6.6.6.5.2.0.6.0mumstc utusgusususu s s s s sec agagag gh ya a a a aet hchchcreol rhophphaphaph gpn gun gun gung gg gg g g g gt reunununununununununu uoav recrecrecrecractsamamamoloso o o o L L L L L L L L L L L L LM OnananananaorotototCoCesesesese P P P P P P S S SO O O O OESESESESESESESESESESESESESESESESESESESESESESESESESESESESESE EA A AA A A A A A A A ACACACACACACACACAS SC C C C C C C C C C C CACACACACACACACACACACAC0154 5064 3 9 2 6 1 3 8 8 5 0 2 1 9 2 6 7 7 4 8 5T0T4 4 4 800 1 9T1T1T03 2 6 6 1 4 2 3 2 1 1 706 4 1 4 3 0 213MAM AA AM AM AM AA AATAM AM AM AM AM AMMMMMMTAMMMMMMTAMS S P P S S S P P P P S S S S SASASASASASASPASASASASASASPAS 3. Sensitivity and Specificity Sensitivity and specificity are fundamental measures in evaluating the accuracy of a diagnostic test. In the context of our categorisation method, we have computed these values based on our ability to correctly identify cancer cases and control (healthy) samples. Sensitivity Sensitivity is the ability of the test to correctly identify those with the disease. It is calculated using the formula: =+Given the results of our categorisation: True Positive (TP): Number of cancer cases correctly categorised as cancerous. In our case, it is 89. False Negative (FN): Number of cancer cases incorrectly categorised as non- cancerous (healthy). In our case, it is 0. The sensitivity of our categorisation method is approximately 100%: Specificity Specificity is the ability of the test to correctly identify those without the disease. It is calculated using the formula: =+Given the results of our categorisation: True Negative (TN): Number of control samples correctly categorised as non- cancerous. In our case, it is 9. False Positive (FP): Number of control samples incorrectly categorised as cancerous. In our case, it is 0 since all controls were correctly identified. The specificity of our categorisation method is approximately 100%: Repeatability and Replicability Repeatability emphasises the consistency within the same conditions, whereas Replicability refers to the consistency of a measure across different test conditions. In the absence of multiple assays on the same samples, we evaluated the repeatability of our assay by examining the consistency of biomarker measurements within each sample. The coefficient of variation (CV) serves as an indirect measure of the assay's repeatability. Lower CV values signify greater repeatability. For some cancers, such as Cervix, Mouth, Ovary, Pancreas, and Prostate, only one sample was available, limiting the scope for calculating variability within those categories. 1. Intra-assay Variability (Repeatability) Biomarker Variation Within Individuals: Across all samples, the CV ranges from 0.211 to 1.636, with a median CV of 0.677. This metric captures the variability in biomarker expression levels within individuals. Controls (CTRL): The CV for controls ranges from 0.283 to 1.608, with a median CV of 0.720. Breast Cancer: CV for Breast Cancer samples ranges from 0.331 to 1.636, with a median of 0.767. Colorectal Cancer: CV for Colorectal Cancer samples ranges from 0.410 to 0.991, with a median of 0.721. Oesophageal Cancer: CV for Oesophageal Cancer samples ranges from 0.211 to 0.899, with a median of 0.668. Lung Cancer: CV for Lung Cancer samples ranges from 0.370 to 0.951, with a median of 0.583. Pancreatic Cancer: CV for Pancreatic Cancer samples ranges from 0.498 to 0.767, with a median of 0.545. Stomach Cancer: CV for Stomach Cancer samples ranges from 0.454 to 1.104, with a median of 0.664. Other Cancer: CV for bile duct, cervical, mouth, ovarian, prostate cancer samples range from 0.497 to 0.833, with a median of 0.745. Biomarker Variation Across Individuals: ’BIOM01: This biomarker has demonstrated the highest variability in stomach cancer patients, with a CV of 1.405, while pancreatic cancer patients exhibited the least CV of 0.262. BIOM17: Oesophagus cancer patients showed the maximum variability with a CV of 1.022, while pancreatic cancer patients exhibited the least CV of 0.350. BIOM19: Oesophagus cancer patients had the highest variability with a CV of 0.570, while lung cancer patients showed the least at 0.427. BIOM24: Oesophageal cancer patients exhibited the maximum variability with a CV of 0.938, and lung cancer patients demonstrated the least variability with a CV of 0.332. BIOM75: Stomach cancer patients showed the maximum variability with a CV of 1.008, while oesophagus cancer patients exhibited the least CV of 0.277. 2. Inter-assay Variability (Replicability) Our assay design and procedures prioritise replicability, even though we did not perform multiple assays on the same sample. Each stage of our assay is delineated with precision based on established standards, ensuring consistent results across different runs. This is complemented by our meticulous documentation of every step, from sample collection to final measurement, facilitating transparency and ease of replication for other researchers. Regular equipment calibration and adherence to rigorous protocols further ensure that our findings can be reliably replicated under similar conditions. 3. Control Samples In our study, we selected control samples from individuals diagnosed with type II diabetes. This choice was deliberate: we primarily employed these samples as controls given their relevance in differentiating biomarker levels specific to pancreatic cancer. Diabetes can be both a symptom and a risk factor for pancreatic cancer, making it a pertinent reference group. In future investigations, we aim to expand our control groups, incorporating completely healthy individuals and those with other conditions, to further validate and refine our biomarker specificity and robustness. 4. Isoform-Specific Results The intricate interplay between cancer progression and the isoforms of specific biomolecules has been increasingly recognized. For instance, various isoforms of BIOM01 have been identified to be pivotal in the transition of naive T cells to memory T cells, underscoring their potential involvement in cancer-related processes (Barashdi et al., 2021; PMID: 34039664). Furthermore, the multifaceted roles of the five known isoforms of BIOM24 in controlling myeloid cell activity through the modulation of the TNF signaling pathway are well-documented (Wajant & Siegmund, 2019; PMID: 31192209). Nevertheless, the exact representation and function of these isoforms within leukocytes during cancer progression remain elusive. In our quest to elucidate the roles of these isoforms, we devised specific primer sets to target them. These are catalogued as follows: BIOM01 Isoforms: RA: Forward Primer BM01V-F2, Reverse Primer BM01V-R3 RB: Forward Primer BM01V-F1, Reverse Primer BM01V-R1 RC: Forward Primer BM01V-F1, Reverse Primer BM01V-R2 RO: Forward Primer BM01V-F1, Reverse Primer BM01V-R3 BIOM24 Isoforms: Isoform 1: Forward Primer BM24V-F1, Reverse Primer BM24V-R1 Isoform 2: Forward Primer BM24V-F1, Reverse Primer BM24V-R2 Isoform 3: Forward Primer BM24V-F2, Reverse Primer BM24V-R1 Isoform 4: Forward Primer BM24V-F2, Reverse Primer BM24V-R2 Isoform 5: Forward Primer BM24V-F2, Reverse Primer BM24V-R3 Table 6. Primer Design for Amplification of BIOM01 and BIOM24 Isoforms Ref ID. Type Sequence SEQ ID NO: BM01V-F1 Forward Primer AGTATTTGTGACAGGGCAAAGC 19. BM01V-F2 Forward Primer ATCCCCGGACTCTTTGGATAAT 20. BM01V-F3 Forward Primer AATGCAAAACTCAACCCTACCC 21. BM01V-R1 Reverse Primer AGGTGAGGCGTCTGTACTGA 22. BM01V-R2 Reverse Primer AACTGGGTCTGTAGGAAAGGT 23. BM01V-R3 Reverse Primer CTCAGAGTGGTTGTTTCAGAGG 24. BM24V-F1 Forward Primer CACTGCCGCTGCCACA 25. BM24V-F2 Forward Primer TGCCATGCAGGTTTCTTTCTAAG 26. BM24V-F3 Forward Primer AAAAAGAGGGGGAGCTTGAAG 27. BM24V-R1 Reverse Primer CGGTCCACTGTGCAAGAAGA 28. BM24V-R2 Reverse Primer AAACAATGGAGTAGAGCTTGGACT 29. BM24V-R3 Reverse Primer GAGATCGCGCCACTGCATT 30. BM24V-R4 Reverse Primer AAAGAAAATGACCAGGGGCAA 31. BM24V-R5 Reverse Primer TGAAGCCTGGAGTGGGACT 32. values (refer to Figure 4), revealed significant variations across different sample types. Specifically, data points from healthy individuals (n=3) were contrasted against cancer cases (n=9), which comprised 4 breast cancers, 2 colon cancers, 1 lung cance differences observed for BIOM01 can be attributed to the differential expression levels of its Isoforms RA and RC. Similarly, the variations in BIOM24 values can be attributed to differential le and BIOM24, and the distribution of the patient samples, BIOM01 and BIOM24 emerge as the focal biomarkers for this invention, underscoring their potential utility in early cancer detection. Discussion The qPCR method was employed to analyse 46 samples, representing a limited dataset. It should be noted that while the sample size for experimental validation was limited, we have performed in silico validation of these selected biomarkers using over 2,000 RNA- seq datasets, further bolstering our confidence in the results. The 46 samples analysed originated exclusively from an Indian cohort. Although this focused cohort may limit the generalisability, it is worth noting that our preliminary discovery datasets encompass diverse ethnic groups, offering a broader foundation for the biomarkers chosen. For the purposes of this study, we amalgamated stages 1 to 3 as early-stage cancer cases, contrasting them with stage 4 which is categorised as late-stage cancer cases. Our thorough trials have determined that samples stored at 4°C for no more than 24 hours deliver the best results. Although samples stored at 4°C within a week can still yield reliable results, samples kept at room temperature beyond 4 hours, immediately stored at -20°C, or treated with 10% DMSO volume prior to storage at -20°C or -80°C without isolating PBMCs either produce no RNA or yield less than 10ng of RNA. Such samples did not meet our quality standard and were thus excluded from the study. Applications Early Detection: Given the non-invasive nature of the blood test and the specificity of the identified biomarkers, our invention holds promise in screening populations at high risk for specific cancers. Earlier detection generally translates to more successful intervention and improved patient outcomes. Differential Diagnosis: With an array of identified biomarkers, our test can aid in differentiating between multiple types of cancers and even potentially between malignant and benign conditions, thus providing more precise diagnostic insights. Reduced Dependency on Imaging: While imaging techniques like CT or MRI are indispensable, they can be expensive, time-consuming, and sometimes inconclusive. Our test, being faster and potentially more affordable, could serve as a preliminary diagnostic tool, directing patients towards specific imaging only when necessary. Disease Surveillance: Patients already diagnosed with cancer could benefit from regular testing using our method to monitor the progression of the disease. Variations in biomarker levels might provide insights into disease dynamics, especially in cases where imaging may not capture subtle changes. Recurrence Detection: Post-therapeutic interventions, it is imperative to monitor patients for potential recurrence. Our test could serve as a routine monitoring tool, offering a quicker indication of relapse compared to waiting for clinical symptoms or periodic imaging. Treatment Efficacy: Changes in biomarker levels can be indicative of how a patient is responding to a given therapy. For instance, a declining trend might signify effective treatment, while a stagnant or increasing trend could warrant a change in therapeutic strategy. Minimising Side Effects: By providing rapid feedback on treatment efficacy, our test might help in adjusting drug dosages or treatment modalities, thereby minimizing potential side effects or ineffective treatments. Stratifying Patient Risk: By gauging the levels and combinations of specific biomarkers, our invention could help classify patients into different risk categories, facilitating personalized treatment strategies. Predicting Disease Evolution: Certain biomarker profiles might correlate with aggressive disease forms or with a propensity for metastasis. Recognizing such patterns early can guide therapeutic decisions and patient counselling. Guidance for Follow-Up: Based on biomarker profiles, clinicians can make informed decisions about the frequency and nature of follow-up appointments, ensuring that high- risk patients receive more intensive monitoring.

[0003] References Agarwal, S., et al. (2020). PAICS, a De Novo Purine Biosynthetic Enzyme, Is Overexpressed in Pancreatic Cancer and Is Involved in Its Progression. Transl Oncol, 13, 100776. Baghban, R., et al. (2020). Tumor microenvironment complexity and therapeutic implications at a glance. Cell Commun Signal, 18, 59. Barashdi, M.A.A., et al. Protein tyrosine phosphatase receptor type C (PTPRC or CD45). J Clin Pathol, 74, 548-552. Blackford, A.L., et al. (2020). Recent Trends in the Incidence and Survival of Stage 1A Pancreatic Cancer: A Surveillance, Epidemiology, and End Results Analysis. J Natl Cancer Inst, 112, 1162-1169. Chakravarthi, B.V.S.K., et al. (2017). Expression and Role of PAICS, a De Novo Purine Biosynthetic Gene in Prostate Cancer. Prostate.77, 10-21. Chakravarthi, B.V.S.K., et al. (2018). A Role for De Novo Purine Metabolic Enzyme PAICS in Bladder Cancer Progression. Neoplasia, 20, 894-904. Cotechini, T., et al. (2021). Tissue-Resident and Recruited Macrophages in Primary Tumor and Metastatic Microenvironments: Potential Targets in Cancer Therapy. Cells, 10, 960. Dai, J., et al. (2020). Exosomes: key players in cancer and potential therapeutic strategy. Signal Transduct Target Ther, 5, 145. Dornan, S., et al. (2002). Differential association of CD45 isoforms with CD4 and CD8 regulates the actions of specific pools of p56lck tyrosine kinase in T cell antigen receptor signal transduction. J Biol Chem, 277, 1912-1918. Fagerholm, S.C., et al. (2019). Beta2-Integrins and Interacting Proteins in Leukocyte Trafficking, Immune Suppression, and Immunodeficiency Disease. Front Immunol, 10, 254. Gallenne, T., et al. (2017). Systematic functional perturbations uncover a prognostic genetic network driving human breast cancer. Oncotarget, 8, 20572-20587. Gordon-Dseagu, V. and Vlad, I. (2023). Differences in cancer incidence and mortality across the globe. Retrieved on 21 August 2023 from https: / / www.wcrf.org / differences-in- Goswami, M.T., et al. (2015). Role and regulation of coordinately expressed de novo purine biosynthetic enzymes PPAT and PAICS in lung cancer. Oncotarget, 6, 23445- 23461. Grunvald, M.W., et al. (2020). Current Status of Circulating Tumor DNA Liquid Biopsy in Pancreatic Cancer. Int J Mol Sci, 21, 7651. Huang, N., et al. (2020). PAICS contributes to gastric carcinogenesis and participates in DNA damage response by interacting with histone deacetylase 1 / 2. Cell Death Dis, 11, 507. Klein, E.A., et al. (2021). Clinical validation of a targeted methylation-based multi-cancer early detection test using an independent validation set. Ann Oncol, 32, 1167-1177. Korsunsky, I., et al. (2019). Fast, sensitive and accurate integration of single-cell data with Harmony. Nat Methods, 16, 1289-1296. Lee, J.J., et al. (2021). Elucidation of Tumor-Stromal Heterogeneity and the Ligand- Receptor Interactome by Single-Cell Transcriptomics in Real-world Pancreatic Cancer Biopsies. Clin Cancer Res, 27, 5912-5921. Lin, D., et al. (2021). Circulating tumor cells: biology and clinical significance. Signal Transduct Target Ther, 6, 404. Liu, S.B., et al. (2020). Loss of IGF2R indicates a poor prognosis and promotes cell proliferation and tumorigenesis in bladder cancer via AKT signaling pathway. Neoplasma, 67, 129-136. Lorenson, M.Y., et al. (2019). Enzyme-linked oligonucleotide hybridization assay for direct oligo measurement in blood. Biol Methods Protoc, 4, bpy014. Masopust, D., & Soerens, A.G. (2019). Tissue-Resident T Cells and Other Resident Leukocytes. Annu Rev Immunol, 37, 521-546. Munn, D.H., & Bronte, V. (2016). Immune suppressive mechanisms in the tumor microenvironment Curr Opin Immunol, 39, 1-6. Opitz, L., et al. (2010). Impact of RNA degradation on gene expression profiling. BMC Med Genomics, 3, 36. Rheinländer, A., et al. (2018). CD45 in human physiology and clinical medicine. Immunol Lett, 196, 22-32. Shen, Y., et al. (2018). Impact of RNA integrity and blood sample storage conditions on the gene expression analysis. Onco Targets Ther, 11, 3573-3581. Steele, N.G., et al. (2020). Multimodal Mapping of the Tumor and Peripheral Blood Immune Landscape in Human Pancreatic Cancer. Nat Cancer, 1, 1097-1112. Sung, H., et al. (2021). Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA Cancer J Clin, 71, 209-249. Takeda, T., et al. (2019). Upregulation of IGF2R evades lysosomal dysfunction-induced apoptosis of cervical cancer cells via transport of cathepsins. Cell Death Dis, 10, 876. Tang, D., et al. (2019). Identification of key pathways and genes changes in pancreatic cancer cells (BXPC-3) after cross-talk with primary pancreatic stellate cells using bioinformatics analysis. Neoplasma, 66, 681-693. von Andrian, U.H., et al. (1991). Two-step model of leukocyte-endothelial cell interaction in inflammation: distinct roles for LECAM-1 and the leukocyte beta 2 integrins in vivo. Proc Natl Acad Sci U S A, 88, 7538-7542. Wajant, H., & Siegmund, D. (2019). TNFR1 and TNFR2 in the Control of the Life and Death Balance of Macrophages. Front Cell Dev Biol, 7, 91. 139, 3480-3492.

Claims

CLAIMS 1. A method for diagnosing cancer in a subject, comprising: a) providing an RNA sample obtained from said subject; b) determining the expression levels of each of a panel of biomarkers in the sample, c) comparing determined expression levels against predetermined criteria, d) diagnosing the subject with cancer based on said comparison, wherein the panel of biomarkers comprises sequences within each of the leukocyte mRNA biomarkers BIOM01 (PTPRC; CD45), BIOM17 (IGF2R; CD222), BIOM19 (ITGAX; CD11C), BIOM24 (TNFRSF1A; CD120A), and BIOM75 (PAICS); and wherein the cancer is selected from the group consisting of: Bile Duct, Breast, Cervix, Colorectum, Oesophagus, Lung, Mouth, Ovary, Pancreas, Prostate, and Stomach.

2. The method of claim 1, wherein step b comprises: conducting reverse transcription on the extracted RNA to produce cDNA; and conducting a quantitative polymerase chain reaction (qPCR) on the cDNA and calculating s in the panel.

3. The method of claim 1 or 2, wherein the sequence length within each of the leukocyte mRNA biomarkers is: from at least 10 codons up to 900 codons, optionally from at least 20 codons up to 800 codons; at least 30 codons up to 700 codons; at least 40 codons up to 600 codons; at least 50 codons up to 500 codons; preferably, from at least 50 codons up to 500 codons.

4. The method of claim 3, wherein the calculated expression levels are normalised relative to GAPDH.

5. The method of claim 4, wherein the predetermined criteria are: i) subjects where BIOM01 < 2.4, BIOM17 < 2.4, BIOM19 < 2.4, BIOM24 lies between 1.2-3.2, and BIOM75 < 1.2 are categorised as healthy, optionally unless BIOM01 < 1, BIOM17 < 1, BIOM19 < 1, BIOM01 + BIOM17 + BIOM19 > 0.2, and BIOM24 > 1 orthe standard deviation (SD) of BIOM01, 17, 19, 24, and 75 < 0.9, BIOM01 + BIOM17 + BIOM19 + BIOM24 + BIOM75 > 2 wherein the samples are categorised as cancerous. ii) subjects where BIOM BIOM BIOM BIOMBIOM BIOM BIOM17 < 1,BIOM19 < 1, BIOM24 < 1, BIOM75 < 1 and BIOM17 + BIOM19 + BIOM24 + BIOM75 < -1 wherein the samples are categorised as healthy.

6. The method according to claim 5 where the predetermined criteria are used to subcategorise cancer by type, and additionally include: iii) subjects where BIOM01 lies between 5-7, BIOM17 lies between 4-6, and BIOM19 lies between 3-5, and BIOM75 < 1 are categorised as Pancreatic Cancer; iv) subjects which do not fall into any of criteria i) to iii) with BIOM01 and BIOM24 values between 2-6 are categorised as Oesophageal Cancer if SD of BIOM01, 17, 19, 24, and 75 < 1 and BIOM75 lies between 1.2-3.2; and categorised as Lung Cancer if SD of BIOM01, 17, 19, 24, and 75 1, BIOM75 1.2, and / or BIOM75 3.2; v) subjects which do not fall into any of criteria i) to iv) are categorised as Prostate Cancer if BIOM19 < 0 and BIOM24 < 0; categorised as Stomach Cancer if BIOM01 < 1, BIOM17 > 0.5, BIOM19 > 0.5, BIOM24 > 1; categorised as Ovarian Cancer if BIOM01 > 5, BIOM17 < 4, BIOM19 < 4, BIOM24 > 5; categorised as Lung Cancer if BIOM01, BIOM17, BIOM19, and BIOM24 > 2 and BIOM75 < -1.75; categorised as Mouth Cancer if BIOM01, BIOM17, and BIOM19 < 1, BIOM24 < 2, and BIOM01 + BIOM17 + BIOM19 + BIOM24 > 2.5; categorised as Oesophageal Cancer if BIOM01, BIOM17, and BIOM24 > 6; categorised as Bile Duct Cancer if BIOM01 > 6, BIOM17 > 6, BIOM19 > 2, and BIOM24 < 4; and categorised as Cervical Cancer if BIOM01 < 2, BIOM19 > 1; vi) subjects which do not fall into any of criteria i) to v) are categorised as Breast Cancer if BIOM75 > 1.5 and BIOM24 < 10, or if BIOM75 lies in 0-1, BIOM24 < 1, and BIOM01 > 0, or if BIOM75 lies in -1-0 and BIOM01 < 0 or BIOM17 < 0; and categorised as Colon Cancer if BIOM75 > 1.5 and BIOM24 10, or if BIOM75 lies in 1-1.5, or if BIOM75 lies in 0-1, BIOM24 < 1, and BIOM01 0, or if BIOM75 lies in 0-1 and BIOM24 > 1, or if BIOM75 lies in -1-0, BIOM01 0 or BIOM17 0.

7. The method of any previous claim, wherein one or more of the biomarker sequences spans an exon junction.

8. The method of claim 7, wherein one or more of the biomarkers comprises one or more sequences selected from SEQ ID NO:3, 6, 9, 12, and / or 15.

9. The method of any previous claim, wherein the sample is selected from peripheral blood, urine, saliva, cerebrospinal fluid, pleural effusion, and ascites.

10. A set of primers for the amplification of a panel of biomarkers, comprising primer pairs which target biomarker sequences within each of PTPRC (CD45), IGF2R (CD222), ITGAX (CD11C), TNFRSF1A (CD120A), and PAICS mRNA molecules.

11. The set of primers according to claim 10, wherein one or more sequence to which the primers are specific and / or of the biomarker sequences comprises an exon junction.

12. The set of primers according to claim 10 or claim 11, wherein: i) one or more primers binds to a site 1kb upstream from the 5' and 1kb downstream from the 3' of the target biomarker sequence in the mRNA molecule; ii) one or more primers yields an amplicon ranging from 80 to 317 base pairs (bp) in length; iii) one or more primers has a Guanine-Cytosine (GC) content of 40-65%; iv) one or more primers has a melting temperature (Tm) ranging from 58 to 62°C; and / or vi) one or more of the primers possess self-complementarity; and / orvii) one or more of the primers exhibit 3' end complementarity less than 6 nucleotide bases.

13. The set of primers according to any one or claims 10 to 12, wherein the primer sequences are selected from: i) a PTPRC (CD45) forward primer comprising a sequence selected from SEQ ID NO:1, 19, 20, and 21 and / or a reverse primer comprising a sequence selected from SEQ ID NO:2, 22, 23 and 24; ii) a IGF2R (CD222) forward primer comprising a sequence of SEQ ID NO:4 and / or a reverse primer comprising a sequence of SEQ ID NO:5; iii) a ITGAX (CD11C) forward primer comprising a sequence of SEQ ID NO:7 and / or a reverse primer comprising a sequence of SEQ ID NO:8; iv) a TNFRSF1A (CD120A) forward primer comprising a sequence selected from SEQ ID NO:10, 25, 26 and 27, and / or a reverse primer comprising a sequence selected from SEQ ID NO:11, 28, 29, 30, 31 and 32; and / or v) a PAICS forward primer comprising a sequence of SEQ ID NO:13 and / or a reverse primer comprising a sequence of SEQ ID NO:

14.

14. The set of primers according to any one or claims 10 to 13, wherein the biomarker sequences have 70% homology to SEQ ID NO:3, 6, 9, 12, and / or 15.

15. The set of primers according to any one or claims 10 to 14, additionally comprising primers specific for a sequence within GAPDH.

16. The set of primers according to claim 15, wherein the sequence within GAPDH is SEQ ID NO:18.

17. The set of primers according to claim 15 or 16, wherein the primers specific for a sequence within GAPDH comprise SEQ ID NO:16 and / or 17.

18. A cancer diagnostic kit, comprising a set of primers according to any one of claims 10 to 17; and optionally one or more of: a) one or more reagents for conducting reverse transcription, optionally selected from a reverse transcriptase enzyme, a DNA polymerase, and / or fluorescent dye molecules capable of binding to double-stranded DNA; b) one or more quantitative polymerase chain reaction (qPCR) reagents; and / or c) Instructions for use according to the method of claims 1 to 9.