Non-coding RNAS for diagnosing pancreatic ductal adenocarcinoma (PDAC) in liquid biopsies

A method using non-coding RNA biomarkers and machine learning models in liquid biopsies effectively addresses the limitations of current PDAC diagnostics, achieving early detection with high accuracy.

WO2026159707A1PCT designated stage Publication Date: 2026-07-30RAMOT AT TEL AVIV UNIVERSITY LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
RAMOT AT TEL AVIV UNIVERSITY LTD
Filing Date
2026-01-18
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Current diagnostic methods for pancreatic ductal adenocarcinoma (PDAC) are limited by low sensitivity and specificity, often detecting the disease at advanced stages, and existing biomarkers like CA19-9 have limitations, necessitating a shift towards non-coding RNA (ncRNA) biomarkers for early detection.

Method used

A method utilizing a combination of non-coding RNA biomarkers, including hsa-miR-4446-3p, hsa-miR-6073, NAALADL2-AS1, and others, quantified through next-generation sequencing and machine learning models, to diagnose PDAC in liquid biopsies.

Benefits of technology

The method achieves an accuracy of 82% balanced accuracy and 91% area under the ROC curve, enabling early detection of PDAC with a high predictive power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IL2026050055_30072026_PF_FP_ABST
    Figure IL2026050055_30072026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to methods for determining in a bodily fluid sample of a subject whether the subject is afflicted with pancreatic ductal adenocarcinoma (PDAC), based on levels of non-coding RNA (ncRNA) biomarkers in the sample. In some embodiments of the invention, the biomarkers are identified by training a machine learning model on levels of ncRNA from subjects afflicted with PDAC and controls, and determining whether the subject is afflicted with PDAC by applying the machine learning model to levels of the biomarkers in a sample from the subject. Specific biomarkers identified by the methods of the invention include hsa-miR-4446-3p, hsa-miR-6073, and NAALADL2-AS1, and optionally Y RNA ENST00000365068.1. Further provided are kits for detection of the biomarkers.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] NON-CODING RNAS FOR DIAGNOSING PANCREATIC DUCTAL ADENOCARCINOMA (PDAC) IN LIQUID BIOPSIES FIELD OF THE INVENTION

[0002] The present invention is generally directed to methods for diagnosing cancer by evaluating non-coding RNA levels in blood. More specifically, the invention relates to using machine learning models for diagnosing pancreatic ductal adenocarcinoma (PDAC) based on levels of specific non-coding cell-free RNAs in plasma.

[0003] BACKGROUND OF THE INVENTION

[0004] Pancreatic ductal adenocarcinoma (PDAC) is the most common type of pancreatic cancer, accounting for over 90% of all pancreatic neoplasms. Despite advancements in treatment and diagnosis, PDAC continues to have a high morbidity rate and poor prognosis. It is the fourth leading cause of cancer-related deaths globally, contributing to 22% of gastrointestinal cancer fatalities and 5% of all cancer deaths. The disease primarily affects individuals over 60, with more than 80% of cases diagnosed in this age group. Known risk factors include smoking and obesity. Additionally, genetics play a significant role in PDAC development, with 5% of all cases carrying a germline mutation in oncogenes like KRAS, or tumor suppressors such as CDKN2A and TP53.

[0005] Potential causes contributing to the high morbidity rate of the disease are the treatment options available for PDAC which are notably limited and often ineffective. One of the primary challenges in treating PDAC is its inherent resistance to conventional cancer treatments. Chemotherapy, a standard treatment for many types of cancer, often proves ineffective against PDAC. This resistance can be attributed to the dense stromal environment of the tumor, which can create a physical barrier that impedes the delivery of chemotherapeutic agents to the cancer cells. Surgical resection, while potentially curative, is only an option for a small percentage of patients, as most are diagnosed when the disease is already at an advanced stage. Even when surgery is possible, it is a high-risk procedure given the pancreas's deep location within the abdominal cavity, and it is often associated with significant morbidity and mortality.

[0006] The current treatment options for PDAC are limited and often ineffective, the difficulties encountered in PDAC treatment have prompted a shift towards more individualized therapeutic approaches. Personalized medicine, which customizes treatment plans based on the unique attributes of a patient's malignancy, may offer a fresh path for patient care, but a more critical challenge and a higher reward lies in the early detection of the disease. Traditional diagnostic methods, such as imaging techniques, are often unable to detect PDAC until it has advanced to alate stage. The most commonly used biomarker for PDAC, Cancer Antigen 19-9 (CAI 9-9), has limitations in sensitivity and specificity, and its levels can be influenced by non-cancerous conditions, leading to false positives. Moreover, not all PDAC patients exhibit elevated CA19-9 signals, limiting its utility as a universal diagnostic marker.

[0007] However, the landscape of PDAC diagnosis is evolving with the advent of novel biomarkers, particularly non-coding RNAs. These biomarkers can be detected in bodily fluids such as blood or saliva, offering a less invasive and more accessible testing process. The use of advanced technologies like Next-Generation Sequencing (NGS; Deep Sequencing or High-Throughput Sequencing), bioinformatics analysis and machine learning (ML) have enabled the identification of unique RNA signatures associated with PDAC. These RNA signatures could potentially serve as reliable diagnostic markers, enabling early detection and intervention, which is crucial for improving patient outcomes.

[0008] The search for indicative biomarkers in liquid biopsies has been extensive with methods aiming to detected early signs of cancer for general and specific malignancies with diverse and potentially exotic non-coding genomic biotypes. Non-coding RNAs (ncRNAs) are a class of RNA molecules that do not code for proteins but play crucial roles in various biological processes. They are typically less than 200 nucleotides in length and include microRNAs (miRNAs), small interfering RNAs (siRNAs), PlWI-interacting RNAs (piRNAs), long non-coding RNAs (IncRNAs; as their name implies, they could be much longer) and others. These ncRNAs are involved in gene regulation, both at the transcriptional and post-transcriptional level. For instance, miRNAs bind to messenger RNAs (mRNAs) and inhibit their translation into proteins. Similarly, siRNAs are involved in the RNA interference (RNAi) pathway, leading to targeted degradation of specific mRNAs. piRNAs are primarily found in the germline and are essential for maintaining genome stability. The study of ncRNAs has opened up new avenues for understanding disease mechanisms, including cancer. PDAC has been targeted as a specific cancer type due to the shortage in high impact biomarkers in early stage. The majority of approaches aim at identifying changes in regulation of miRNAs associated with the tumor progression and proliferation via high-throughput sequencing of known and specific short sequences identified in solid tumor assays and their trace in blood plasma.

[0009] While some miRNAs have been shown to be associated with PDAC, casting a broader net has a potential for catching other RNAs and combinations that may better and more accurately predict the cancer.SUMMARY OF INVENTION

[0010] The following embodiments are described and illustrated in conjunction with compositions and methods which are meant to be exemplary and illustrative, not limiting in scope. In various embodiments, one or more of the above-described problems have been reduced or eliminated, while other embodiments are directed to other advantages or improvements.

[0011] In some embodiments, there is provided a method for determining in a bodily fluid sample of a subject whether the subject is likely afflicted with pancreatic ductal adenocarcinoma (PDAC), the method including:

[0012] a. quantifying in the bodily fluid sample levels of at least 3 non-coding RNA (ncRNA) biomarkers including hsa-miR-4446-3p, hsa-miR-6073 , and NAAI DL2-ASR and b. determining, based on the ncRNA biomarkers levels, whether the subject is likely afflicted with PDAC.

[0013] In some embodiments, the at least 3 ncRNA biomarkers further include at least one additional ncRNA biomarker selected from Y RNA ENST00000365068.1, DLEU2, RP11-874J12.4, hsa-miR-432-5p, IncRNA C15orf54, hsa-miR-6721-5p, CTD-3092A11.2, Y RNA ENST00000516950.1, DPYD-AS2, hsa-miR-6131, misc RNA 7SK, KCTD21-AS1, IncRNA RP13-216E22.4, MIR223HG-201, hsa-miR-4433b-3p, IncRNA RP11-930011.2, VIM-AS1, and combinations thereof.

[0014] In some embodiments, the determining includes applying to the levels of the ncRNA biomarkers in the sample a machine learning classification model optimized based on levels of the ncRNA biomarkers in PDAC patients and controls, and / or comparing the levels of the ncRNA biomarkers in the sample to control levels of the ncRNA biomarkers.

[0015] In some embodiments, when hsa-miR-4446-3p and hsa-miR-6073 levels are lower and NAALADL2-AS1 level is higher in the sample of the subject compared to control levels, the subject is considered likely to be afflicted with PDAC.

[0016] In some embodiments, the at least 3 ncRNA biomarkers further include at least one additional ncRNA biomarker selected from Y RNA ENST00000365068.1, DELU2, RP11-874J12.4, and hsa-miR-432-5p.

[0017] In some embodiments, the at least one additional ncRNA biomarker includes Y RNA ENST00000365068.1, DELU2, RP11-874J12.4, and hsa-miR-432-5p.

[0018] In some embodiments, the at least 3 ncRNA biomarkers are four ncRNA biomarkers, further including Y RNA ENST00000365068.1.

[0019] In some embodiments, when Y RNA ENST00000365068.1, DELU2, RP11-874J12.4, and / or hsa-miR-432-5p levels are lower in the sample of the subject compared to control levels,the subject is considered likely to be afflicted with PDAC.

[0020] In some embodiments, the at least 3 ncRNA biomarkers further include MIR223HG-201. In some embodiments, when MIR223HG-201 levels are lower in the sample of the subject compared to the control levels, the subject is considered likely to be afflicted with PDAC.

[0021] In some embodiments, quantifying the levels of the ncRNA biomarkers includes isolating RNA from the sample and quantifying levels of the ncRNA biomarkers in the isolated RNA.

[0022] In some embodiments, the RNA is cell-free RNA.

[0023] In some embodiments, the quantifying includes sequencing the RNA.

[0024] In some embodiments, the quantifying further includes in silico filtering out contaminants and / or in silico enriching for ncRNAs, by performing sequence alignments.

[0025] In some embodiments, the method further includes normalizing the ncRNA biomarkers levels by using Trimmed Mean of M-values (TMM) to account for library size and composition bias.

[0026] In some embodiments, the machine learning model is trained on levels of the ncRNA biomarkers obtained from subjects diagnosed with PDAC and from control subject not afflicted with PDAC.

[0027] In some embodiments, the machine learning model is selected from a Light Gradient Boosting Machine (LGBM) classifier, an ExtraTrees classifier, a Random Forest classifier, a Support Vector Machine (SVM) classifier, and an XGBoost classifier.

[0028] In some embodiments, the machine learning model is an LGBM classifier optimized based on the ncRNA biomarkers.

[0029] In some embodiments, the control levels of the ncRNA biomarkers are obtained by quantifying the levels of the ncRNA biomarkers in a bodily fluid sample of a subject not afflicted with PDAC.

[0030] In some embodiments, the bodily fluid sample is selected from blood, plasma, serum, cerebral fluid, urine, and combinations thereof.

[0031] In some embodiments, the bodily fluid sample is cell-free.

[0032] In some embodiments, the subject is a smoker.

[0033] In some embodiments, the subject has a body mass index (BMI) of above 20, 30, or 35. In some embodiments, the subject is at least about 55, 60, or 65 years old.

[0034] In some embodiments, the method further includes treating the subject with an anticancer treatment, when the subject is determined as likely to be afflicted with PDAC.

[0035] In some embodiments, there is provided a method for training a machine learning model for determining whether a subject is likely afflicted with a cancer, including the steps of:a. performing a differential gene expression analysis on ncRNA levels from subjects diagnosed with the cancer and controls, to identify a plurality of ncRNAs that are differentially expressed between the subjects diagnosed with the cancer and the controls;

[0036] b. selecting most informative ncRNAs by applying a machine learning classifier to the plurality of differentially expressed ncRNAs;

[0037] c. evaluating randomized combinations of the most informative ncRNAs to identify a group of ncRNA biomarkers; and

[0038] d. training a classification model based on the group of ncRNA biomarkers.

[0039] In some embodiments, the ncRNA levels are obtained by steps including:

[0040] a. sequencing RNA isolated from bodily fluid samples of subjects diagnosed with PDAC and controls;

[0041] b. filtering out contaminants and enriching for ncRNAs by performing sequence alignments; and

[0042] c. quantifying levels of the ncRNAs.

[0043] In some embodiments, the differential gene expression analysis includes applying minimum covariance determinant (MCD) to the ncRNA levels.

[0044] In some embodiments, the differential gene expression analysis includes using supervised gradient boosting.

[0045] In some embodiments, selecting the most informative ncRNAs includes using an ExtraTrees classifier.

[0046] In some embodiments, selecting the most informative ncRNAs includes using a gradient boosting algorithm.

[0047] In some embodiments, selecting the most informative ncRNAs includes mapping genes to known cellular pathways.

[0048] In some embodiments, there is provided a method of treating a subject suspected of being afflicted with PDAC, including:

[0049] a. quantifying in a bodily fluid sample of the subject levels of at least 3 ncRNA biomarkers including hsa-miR-4446-3p, hsa-miR-6073 , and NAAI DL2-ASR b. determining, based on the ncRNA biomarkers levels, whether the subject is likely afflicted with PDAC; and

[0050] c. treating the subject with an anti cancer treatment, when the subject is determined as likely to be afflicted with PDAC.

[0051] In some embodiments, the at least 3 ncRNA biomarkers further include at least oneadditional ncRNA biomarker selected from Y RNA ENST00000365068.1, DLEU2, RP11-874J12.4, hsa-miR-432-5p, IncRNA C15orf54, hsa-miR-6721-5p, CTD-3092A11.2, Y RNA ENST00000516950.1, DPYD-AS2, hsa-miR-6131, misc RNA 7SK, KCTD21-AS1, IncRNA RP13-216E22.4, MIR223HG-201, hsa-miR-4433b-3p, IncRNA RP11-930011.2, VIM-AS1, and combinations thereof.

[0052] In some embodiments, the determining is done by applying to the levels of the ncRNA biomarkers in the sample a machine learning classification model optimized based on levels of the ncRNA biomarkers in PDAC patients and controls, or by comparing the levels of the ncRNA biomarkers in the sample to control levels of the ncRNA biomarkers.

[0053] In some embodiments, there is provided a kit for determining in a bodily fluid sample of a subj ect whether the subj ect is likely afflicted with PDAC, the kit including reagents for quantifying levels of at least 3 ncRNA biomarkers including hsa-miR-4446-3p, hsa-miR-6073, and NAALADL2-AS1, in the bodily fluid sample.

[0054] In some embodiments, the reagents include: (a) a probe specific to the mature sequence of hsa-miR-4446-3p; (b) a probe specific to the mature sequence of hsa-miR-6073; and (c) a probe and / or a primer pair specific for NAALADL2- AS 1.

[0055] In addition to the exemplary embodiments described above, further embodiments will become apparent by reference to the figures and by studying the following detailed descriptions.

[0056] BRIEF DESCRIPTION OF DRAWINGS

[0057] The invention will now be described in relation to certain examples and embodiments with reference to the following illustrative figures.

[0058] Fig. 1 shows a schematic illustration of the integrated bioinformatics and ML workflow. The process commences with a stepwise alignment procedure, an iterative method ensuring optimal sequence alignment, critical for accurate downstream analysis. Subsequently, the workflow identifies differentially expressed genes, pivotal in deciphering gene expression patterns and their biological significance. The workflow then prioritizes genes using the ExtraTrees algorithm, robust ML technique that aids in the selection of salient genes for further investigation. The final stage employs gradient boosting classification, a potent ML algorithm, for data classification and predictive modeling.

[0059] Figs. 2A-2H show a detailed analysis of non-coding RNA biotypes and classification performance. A visual representation of expression analysis of various ncRNA biotypes and the performance of the gradient-boosting classification model. Figs.2A-2F present violin plots for six distinct ncRNA biotypes: mitochondrial tRNA (Fig 2A), miRNA (Fig 2B), piRNA (Fig 2C),miscRNA (Fig 2D), mitochondrial rRNA (Fig 2E), IncRNA (Fig 2F). These plots offer a detailed view of the distribution of expression levels for each biotype in both PDAC and control samples. The violin plots illustrate the density and spread of the data. Upper panels: left: PDAC, right: control. Fig. 2G presents a Receiver Operating Characteristic (ROC) curve with cross-validation, showcasing the performance of the model when classifying based on the top 7 genetic features (ncRNAs). The area under the ROC curve (AUC) is also indicated, offering a single summary measure of the model's predictive power. Each colored line represents a single evaluation in a fivefold cross validation routine, whereas the bolded dark blue line represents the cross validation average with standard deviation as shaded grey region. Balanced accuracy = 0.82 + / - 0.07; Accuracy = 0.87 + / - 0.04; Fl = 0.86 + / - 0.05; Precision = 0.89 + / - 0.03; Recall = 0.87 + / - 0.04; ROC AUC = 0.91 + / - 0.07; MCC (Matthews Correlation Coefficient^ 0.72 + / - 0.07; Score = 0.85.

[0060] Fig. 2H shows a heatmap of the top 20 features, representation of the expression levels of these features in PDAC and control samples. The heatmap uses the Ward metric for hierarchical clustering, which groups samples based on the similarity of their expression patterns for these top features. This figure underscores the effectiveness in distinguishing between PDAC and healthy plasma samples using plasma-derived ncRNAs and highlights the potential of these RNAs as biomarkers for PDAC.

[0061] Figs. 3A-3E show differences in distribution of data between control and PDAC groups.

[0062] Fig. 3A shows a PCA of the TMM (trimmed mean of the log expression ratios) normalized count matrix. The PCA reduces the dimensionality of the data, allowing for visualization in a 2D space. The plot shows the first two principal components. The data points are color-coded based on their labels, Control and PDAC. The decision boundaries of the MCD (Minimum Covariance Determinant) estimator are also shown, identifying potential outliers in the data. The plot provides an overview of the data distribution and the difficulty of separation in an unsupervised manner between the two classes. Figs.3B-3E present a distribution of normalized counts for four top genes in both 'Control' and 'PDAC groups. Each subplot consists of a boxplot (left) and a histogram with a kernel density estimate (right). The boxplots provide a summary of the central tendency, dispersion, and skewness of the gene expression data, while the histograms offer a visual interpretation of data distribution. Y-axis box plot and X-axis histograms units: Trimmed Mean of M-values (TMM)-normalized values. The genes are labeled in each subplot. These plots visualize the differences in gene expression between Control (blue) and PDAC (red) groups for the top most important genes. Fig.3B hsa-miR-4446-3p is downregulated in PDAC compared to controls; Fig.

[0063] 3C hsa-miR-6073 is downregulated in PDAC compared to controls; Fig. 3D Antisense NAALADL2-AS1 is upregulated in PDAC compared to controls; Fig. 3E Y_RNAENST00000365068.1 is downregulated in PDAC compared to controls.

[0064] Fig- 4 show results of a gene set enrichment and pathway analysis of the top 20 genes identified from the DGE (differential gene expression) analysis. Dot plot of enriched terms is shown, with each dot representing a specific term. The size of the dot corresponds to the count of the term. The adjusted P value for pancreatic cancer was the most significant (3.76* 10-12). FDR: False Discovery Rate.

[0065] Fig- 5 shows a clustered heatmap of correlation coefficients between the normalized read counts of the top 20 non-coding genes and the clinical attributes (age, BMI, weight, height, gender) of the PDAC samples. Each cell in the heatmap represents the correlation coefficient between a specific gene and a clinical feature. The color scale, ranging from dark blue (negative correlation) to light yellow (positive correlation), indicates the strength and direction of the correlation. The dendrogram on the y-axis groups genes is based on the similarity of their correlation patterns with the clinical features. Notably, BMI and weight display correlations peaking at -0.39 with certain genes, including the IncRNA antisense MIR223HG-201. Conversely, gender and height demonstrate correlations approximating zero with the selected genes. The heatmap underscores the intricate interplay between gene expression and clinical features in PDAC.

[0066] DETAILED DESCRIPTION OF THE INVENTION

[0067] In the following description, various aspects of the disclosure will be described. For the purpose of explanation, specific configurations and details are set forth in order to provide a thorough understanding of the different aspects of the disclosure. However, it will also be apparent to one skilled in the art that the disclosure may be practiced without specific details being presented herein. Furthermore, well-known features may be omitted or simplified in order not to obscure the disclosure.

[0068] Pancreatic ductal adenocarcinoma (PDAC) presents a formidable challenge in oncology, with survival rates that have remained stagnant over the past decades. The majority of patients are diagnosed at advanced stages, often with metastases, making them eligible only for palliative care. Even for those with localized tumors, radical surgery often fails to ensure long-term survival, with median survival post-diagnosis ranging from eight to ten months and most patients experiencing early tumor relapse.

[0069] While survival is predominantly determined by the stage of the tumor, both long-term and short-term survival have been observed in patients diagnosed with early-stage tumors. Circulating miRNAs have been found to have prognostic potential in various cancers, which may largely be due to their altered expression during tumorigenesis and their partial stability in bodily fluids.The present invention is directed to early diagnosis of PDAC, by evaluating levels of noncoding RNA (ncRNA) biomarkers in a liquid sample, such as a blood sample. The inventors combined next generation sequencing (NGS), bioinformatics analysis, and machine learning (ML) to identify unique RNA signatures associated with PDAC (Fig. 1). The step-wise sequence mapping procedure developed for this purpose facilitated accurately retrieving signals from plasma samples. A differential gene expression (DGE) analysis identified genes that were significantly differentially expressed between PDAC and control samples. The ExtraTrees algorithm was then used to rank these genes based on their importance. The top-ranking 20 genes were then subjected to gene set enrichment and pathway analysis. Notably, the pancreatic cancer pathway emerged as the top-ranking pathway, significantly ahead of subsequent pathways (Fig.

[0070] 4). This finding underscores the potential of these genes as biomarkers for PDAC specifically in liquid biopsies.

[0071] The results provided in the present invention demonstrate that the top seven biomarkers identified through the ExtraTrees classifier prioritization process yield a classification balanced accuracy of 82% and an ALIC of 91% (Fig.2G), while partial combinations with fewer biomarkers also provided a high predictive ability. The top seven ncRNA biomarkers include three miRNAs: hsa-miR-4446-3p, hsa-miR-6073, and hsa-miR-432-5p a Y RNA (component of the Ro ribonucleoprotein (RNP) complex in humans) transcript Y RNA ENST00000365068.1,' and three IncRNAs: NAALADL2-AS1, DELU2, and RP11-874J12.4. As shown in Figs. 3B-3E, hsa-miR-4446-3p, hsa-miR-6073, and Y RNA ENST00000365068.1 were downregulated in PDAC samples, while NAALADL2-AS1 was upregulated in PDAC samples.

[0072] Methods of predicting PDAC based on ncRNA biomarkers

[0073] In some embodiments, there is provided a method for determining in a bodily fluid sample of a subject whether the subject is likely afflicted with PDAC, the method including:

[0074] a. quantifying in the bodily fluid sample levels of at least 3 ncRNA biomarkers comprising hsa-miR-4446-3p, hsa-miR-6073, and NAALADL2-AS1,' and

[0075] b. determining, based on the ncRNA biomarkers levels, whether the subject is likely afflicted with PDAC.

[0076] In some embodiments, the at least 3 ncRNA biomarkers include at least one microRNA (miRNA). In some embodiments, the ncRNA biomarkers include at least one small interfering RNA (siRNA). In some embodiments, the ncRNA biomarkers include at least one PIWI-interacting RNA (piRNA). In some embodiments, the ncRNA biomarkers include at least one long non-coding RNA (IncRNA). In some embodiments, the ncRNA include at least one hairpin RNA.In some embodiments, the at least 3 ncRNA biomarkers are at least four ncRNA biomarkers, including hsa-miR-4446-3p, hsa-miR-6073, NAALADL2-AS1 and Y_RNA ENST00000365068.1.

[0077] In some embodiments, the at least 3 ncRNA biomarkers further include at least one additional ncRNA biomarker selected from ncRNAs in Table 2, namely: Y_RNA ENST00000365068.1, DLEU2, RP11-874J12.4, hsa-miR-432-5p, IncRNA C15orf54, hsa-miR-6721-5p, CTD-3092A11.2, Y RNA ENST00000516950.1, DPYD-AS2, hsa-miR-6131, misc RNA 7SK, KCTD21-AS1, IncRNA RP13-216E22.4, MIR223HG-201, hsa-miR-4433b-3p, IncRNA RP11-930011.2, VIM-AS1, and combinations thereof.

[0078] In some embodiments, the ncRNA biomarkers include up to 5, 4, 3, or to 2, nucleotide differences (mismatches in a sequence alignment) compared to the sequence of the corresponding accession number indicated in Table 2. In some embodiments, the ncRNA biomarkers include up to 10%, 5%, or 1%, nucleotide differences (mismatches in a sequence alignment) compared to the sequence of the corresponding accession number indicated in Table 2.

[0079] In some embodiments, the at least one additional ncRNA biomarker is selected from Y RNA ENST00000365068.1 , DELU2, RP11-874J12.4, and hsa-miR-432-5p. In some embodiments, the at least one additional ncRNA biomarker includes Y RNA ENST00000365068.1 , DELU2, RP11-874J12.4, and hsa-miR-432-5p.

[0080] In some embodiments, the at least one additional ncRNA biomarker includes MIR223HG-201.

[0081] In some embodiments, the at least one additional ncRNA biomarker includes Y RNA ENST00000365068.1 and DLEU2; Y RNA ENST00000365068.1 and RP11-874J12.4; Y RNA ENST00000365068.1 and hsa-miR-432-5p; DLEU2 and RP11-874J12.4; DLEU2 and hsa-miR-432-5p; or RP11-874J12.4 and hsa-miR-432-5p.

[0082] In some embodiments, the at least one additional ncRNA biomarker includes Y RNA ENST00000365068.1, DLEU2, and RP11-874J12.4; Y_RNAENST00000365068.1, DLEU2, and hsa-miR-432-5p; or DLEU2, RP11-874J12.4, and hsa-miR-432-5p.

[0083] In some embodiments, when the ncRNA is miRNA, the miRNA is a mature miRNA.

[0084] In some embodiments, at least one of the adenine nucleotides of the miRNA biomarkers are replaces with inosine.

[0085] The present invention is directed to PDAC diagnostics based on a liquid biopsy, i.e., a non-invasive or minimally invasive diagnostic test that analyzes bodily fluids to detect and analyze biomarkers shed by a tumor or other tissues.In some embodiments, the bodily fluid is selected from blood, plasma, serum, cerebral fluid, urine, and combinations thereof.

[0086] In some embodiments, the bodily fluid sample is cell-free. The term “cell-free” with respect to the sample means that cells are not present in the sample, or that cells are present in the sample at a negligible amount such that RNA extracted from the cells is present at a negligible amount compared to the total RNA amount extracted from the sample. In some embodiments, a “negligible amount” is less than 1%

[0087] If the bodily fluid includes cells, the method may further include a step including separating the cells from the fluid before quantifying ncRNA biomarkers levels in the cell-free fluid.

[0088] In some embodiments, the quantifying includes isolating RNA from the sample and quantifying levels of the ncRNA biomarkers in the isolated RNA.

[0089] In some embodiments, the isolated RNA is cell-free RNA.

[0090] The term “cell-free” with respect to RNA means that the extracted RNA is not cellular RNA (RNA extracted from cells) but rather RNA that is present outside cells (e.g. in exosomes or freely circulating in blood plasma), or that any cellular RNA is present at a negligible amount (i.e., less than 1%) compared to the total RNA amount extracted from the sample.

[0091] Methods for isolating RNA are well known in the field and include, for example, organic extraction (phenol-chloroform / trizol), silica column-based (spin-columns), and magnetic beadbased isolation. Additionally, isolating extracellular RNA from exosomes may be done by ultracentrifugation, polymer precipitation or size-exclusion chromatography, lysis of the vesicle, and RNA extraction using a column or beads. A nonlimiting example is provided in the experimental part below, including using the QIAamp® Circulating Nucleic Acid Kit (Qiagen).

[0092] The quantifying of levels of the ncRNA biomarkers may be done by any suitable method such as hybridization with probes specific for the biomarkers, sequencing, such as next generation sequencing (NGS), or sequence-specific quantitative real-time PCR (qRT-PCR).

[0093] In some embodiments, the quantifying includes sequencing the RNA.

[0094] In some embodiments, the quantifying includes in silico filtering out contaminants. Nonlimiting examples for contaminants which may be filtered out include rRNA, tRNA, and coding RNA.

[0095] In some embodiments, the quantifying includes in silico enriching for ncRNAs.

[0096] The filtering and / or the enriching may be performed by sequence alignments to known sequences. For example, contaminant filtering may be conducted by alignment to rRNA, tRNA, and / or coding RNA in the Ensembl database. The enrichment for ncRNAs may be performed bymapping the RNA sequences to specific databases, e.g., piRNAdb32, Ensembl ncRNA database, miRBase hairpin sequences database, or miRBase mature sequences.

[0097] In some embodiments, the in silico filtering out contaminants is performed by sequence alignments.

[0098] In some embodiments, the in silico enriching for ncRNAs is performed by sequence alignments.

[0099] In some embodiments, the quantifying includes aligning the ncRNA sequences to at least one sequence database and counting hits in the alignment results for quantifying each ncRNA level.

[0100] In some embodiments, the method further includes normalizing the ncRNA biomarkers levels by using Trimmed Mean of M-values (TMM) to account for library size and composition bias.

[0101] In some embodiments, determining whether the subject is likely afflicted with PDAC includes applying to the levels of biomarkers in the sample a machine learning classification model optimized based on levels of the ncRNA biomarkers in PDAC patients and controls.

[0102] The term “controls” or “control subjects”, as used herein, relates to subjects who are not diagnosed with PDAC at the time of obtaining their ncRNA biomarkers levels. In some embodiments, the controls have never been diagnosed with PDAC. In some embodiments, the controls are healthy subjects.

[0103] In some embodiments, the machine learning model is selected from a Light Gradient Boosting Machine (LGBM) classifier, an ExtraTrees classifier, a Random Forest classifier, a Support Vector Machine (SVM) classifier, and an XGBoost classifier.

[0104] In some embodiments, the machine learning model is an LGBM classifier optimized based on the ncRNA biomarkers. In some embodiments, the optimization comprises a grid search. In some embodiments, the grid search comprises a five-fold cross-validation routine to evaluate parameter combinations. In some embodiments, the model is further optimized by incorporating clinical parameters, such as hospital sample origin, subject age, subject weight, and / or subject gender.

[0105] In some embodiments, the machine learning classification model provides a diagnostic score. In some embodiments, the diagnostic score is a probability value ranging from 0 to 1. In some embodiments, the diagnostic score is calculated as a function of the weighted levels of the ncRNA biomarkers and the clinical parameters. In some embodiments, when the diagnostic score is above a predetermined threshold the subject is determined as likely to be afflicted with PDAC.In some embodiments, the threshold is determined based on a Receiver Operating Characteristic (ROC) curve analysis to optimize sensitivity and specificity.

[0106] In some embodiments, the predetermined threshold is at least about 0.5. In some embodiments, the predetermined threshold is about 0.5, 0.6, or 0.7.

[0107] In some embodiments, determining whether the subject is likely afflicted with PDAC is based on the diagnostic score.

[0108] In some embodiments, determining whether the subject is likely afflicted with PDAC includes comparing the levels of the ncRNA biomarkers in the sample to control levels of the ncRNA biomarkers.

[0109] In some embodiments, the control levels of the ncRNA biomarkers are obtained by quantifying the levels of the ncRNA biomarkers in a bodily fluid sample of a control subject. In some embodiments, the control levels of the ncRNA biomarkers are precalculated reference values representing normal levels, not associated with PDAC.

[0110] In some embodiments, determining whether the subject is likely afflicted with PDAC further includes comparing the levels of the ncRNA biomarkers in the sample to levels of the ncRNA biomarkers typical to PDAC patients.

[0111] In some embodiments, the levels of the ncRNA biomarkers typical to PDAC patients are obtained by quantifying the levels of the ncRNA biomarkers in a bodily fluid sample of a PDAC patient. In some embodiments, the levels of the ncRNA biomarkers typical to PDAC patients are precalculated reference values representing levels in PDAC patients.

[0112] In some embodiments, when the levels of hsa-miR-4446-3p, hsa-miR-6073, Y RNA ENST00000365068.1, DLEU2, RP11-874J12.4, MIR223HG-201, and / or hsa-miR-432-5p in the sample are lower in the sample of the subject than the control levels, than the subject is likely afflicted with the cancer.

[0113] In some embodiments, when the levels of NAALADL2-AS1 in the sample of the subject are higher than the control levels, then the subject is likely afflicted with the cancer.

[0114] In some embodiments, when the levels of hsa-miR-4446-3p and hsa-miR-6073 are lower and / or the levels of NAALADL2-AS1 is higher in the sample of the subject compared to control levels, the subject is considered likely to be afflicted with PDAC.

[0115] In some embodiments, when the levels of Y RNA ENST00000365068.1, DELU2, RP11-874J12.4, and / or hsa-miR-432-5p are lower in the sample of the subj ect compared to control levels, the subject is considered likely to be afflicted with PDAC.

[0116] In some embodiments, when the levels of MIR223HG-201 are lower in the sample of thesubject compared to the control levels, the subject is considered likely to be afflicted with PDAC. In some embodiments, the term “lower” relates to sample ncRNA biomarkers levels lower by at least 10%, 20%, 30%, 40%, or 50% compared to the control levels. In some embodiments, the term “higher” relates to sample ncRNA biomarkers levels higher by at least 10%, 20%, 30%, 40%, or 50% compared to the control levels.

[0117] In some embodiments, the at least 3 biomarkers can predict that a patient is likely to be afflicted with PDAC with an accuracy of at least about 80%, 85%, or 87%, and / or an area under the receiver operating characteristic (ROC) curve of at least 0.8, 0.85, or 0.9.

[0118] In some embodiments, the subject is at risk for having cancer. In some embodiments, the subject is at risk for having PDAC. The term “subject being at risk for having cancer / PDAC” means that the subject has a risk factor which is considered to be associated with the cancer or specifically with PDAC. Such a factor may be, e.g., a family history, genetic biomarker, environmental exposure (including e.g., smoking), or physiological factor (including age, and body mass index (BMI)).

[0119] In some embodiments, the subject does not present any symptoms related to cancer. In some embodiments, the subject does not present symptoms related to PDAC (i.e., PDAC is at an early stage). In some embodiments, the subject presents at least some symptoms of PDAC.

[0120] Nonlimiting examples for symptoms of cancer include unexplained weight loss, fatigue, persistent pain, skin changes (jaundice, darkening, or redness), lumps or thickening under the skin, fever, night sweats, persistent cough, changes in bowel or bladder habits, unusual bleeding or bruising, and combinations thereof.

[0121] Nonlimiting examples for symptoms of PDAC include jaundice (yellowing of the eyes and skin), abdominal pain radiating to the back, unexplained weight loss, loss of appetite, steatorrhea (oily or foul-smelling stools), dark urine, pale stools, new-onset diabetes, nausea and vomiting, blood clots (deep vein thrombosis), and combinations thereof.

[0122] In some embodiments, the subject is a mammal. In some embodiments, the subject is a human subject.

[0123] In some embodiments, the subject has a history of smoking.

[0124] In some embodiments, the subject has a BMI or at least about 25, 30, 35, or 40.

[0125] In some embodiments, the subject is at least about 55, 60, or 65 years old.

[0126] In some embodiments, the method further includes treating the subject with an anticancer treatment and / or with an anticancer agent, when the subject is determined as likely to be afflictedwith PDAC.

[0127] Nonlimiting examples for anticancer treatments suitable for treating PDAC include chemotherapy and targeted therapies, such as drugs targeting BRCA1 / 2 / PALB2, KRAS G12C, NRG1 fusion, NTRK fusion, or HER2.

[0128] Nonlimiting examples for anticancer drugs suitable for treating PDAC include chemotherapeutic drugs such as FOLFIRINOX (Folinic acid, Fluorouracil (5-FU), Irinotecan, and Oxaliplatin), Gemcitabine + Nab-paclitaxel, and NALIRIFOX (liposomal irinotecan, 5-fluorouracil, leucovorin, and oxaliplatin); and targeted therapy drugs such as Olaparib, Sotorasib, Adagrasib, Zenocutuzumab, Larotrectinib, Entrectinib, Trastuzumab, and deruxtecan.

[0129] Methods for training the machine learning model

[0130] In some embodiments, there is provided a method for training a machine learning model for determining whether a subject is likely afflicted with a cancer, including the steps of:

[0131] a. performing a differential gene expression analysis on ncRNA levels from subjects diagnosed with the cancer and controls, to identify a plurality of ncRNAs that are differentially expressed between the subjects diagnosed with the cancer and the controls;

[0132] b. selecting most informative ncRNAs by applying a machine learning classifier to the plurality of differentially expressed ncRNAs;

[0133] c. evaluating randomized combinations of the most informative ncRNAs to identify a group of ncRNA biomarkers; and

[0134] d. training a classification model based on the group of ncRNA biomarkers.

[0135] Definitions and embodiments mentioned above and which may be relevant to embodiments in this section also apply here, and vice versa. Some particularly relevant embodiments may be pointed out or explicitly repeated. For terms used herein, unless stated otherwise, their definition and embodiments are intended to be the same as above (mutatis mutandis).

[0136] The term “differentially expressed”, as used herein, relates to differences in ncRNA levels between PDAC samples and control samples. This terminology is used since it is likely that differences in ncRNA levels between samples result from differences in expression of the ncRNA in cells of the subject (e.g., the cancer cells). The method considers only statistically significant differentially expressed genes as part of the gene selection process. Additionally, for the same reasons, the term “expression level” is sometimes used herein to describe ncRNA levels in blood samples.

[0137] In some embodiments, the cancer is selected from pancreatic cancer, liver cancer, kidneycancer, stomach cancer, colorectal cancer, esophageal cancer, prostate cancer, bladder cancer, uterine cancer, gallbladder cancer, bile duct cancer, adrenal cancer, spleen cancer, small intestine cancer, and thymus cancer. In some embodiments, the cancer is pancreatic ductal adenocarcinoma (PDAC).

[0138] In some embodiments, the ncRNA levels are obtained by steps including:

[0139] a. sequencing RNA isolated from bodily fluid samples of subjects diagnosed with PDAC and controls;

[0140] b. filtering out contaminants and enriching for ncRNAs by performing sequence alignments; and

[0141] c. quantifying levels of the ncRNAs.

[0142] The levels of ncRNA may be obtained from the subjects diagnosed with the cancer and the controls by the same methods described above with reference to the subject being diagnosed, i.e., by isolating cell-free RNA and quantifying levels of ncRNA in the isolated RNA as described above.

[0143] In some embodiments, the quantifying includes aligning the ncRNA sequences to sequence databases and counting hits in the alignment results.

[0144] In some embodiments, the quantifying includes in silico filtering out contaminants. Nonlimiting examples for contaminants which may be filtered out include rRNA, tRNA, and coding RNA.

[0145] In some embodiments, the quantifying includes in silico enriching for ncRNAs.

[0146] The filtering and / or the enriching may be performed by sequence alignments to known sequences. For example, contaminant filtering may be conducted by alignment to rRNA, tRNA, and / or coding RNA in the Ensembl database. The enrichment for ncRNAs may be performed by mapping the RNA sequences to specific databases, e.g., piRNAdb32, Ensembl ncRNA database, miRBase hairpin sequences database, or miRBase mature sequences. The stepwise routine coupled with the stringency of sequence mapping in each step ensures enrichment of high classifying biotypes which subsequently contributes to the high accuracy of the machine learning model.

[0147] The differentially expressed ncRNAs may be upregulated in cancer samples compared to control samples, or may be downregulated in cancer samples compared to control samples. In some embodiments, some of the ncRNA biomarkers are upregulated in cancer samples compared to control samples and some of the ncRNA biomarkers are downregulated in cancer samples compared to control samples.

[0148] In some embodiments, the differential gene expression analysis includes applying minimum covariance determinant (MCD) to the ncRNA levels.In some embodiments, the differential gene expression analysis includes using supervised gradient boosting.

[0149] In some embodiments, the plurality of differentially expressed ncRNAs are defined by having a P value of < 0.01 in the differential expression analysis. A nonlimiting example for the analysis of differential expression may be found in the experimental section.

[0150] As used herein, the term “plurality” relates to more than 1. In some embodiments, the plurality of differentially expressed ncRNAs is 2, 3, 4, 5, 6, 7 or more ncRNAs.

[0151] Selecting most informative ncRNAs by the machine learning classifier may be done by using various suitable classifiers. In some embodiments, the classifier is based on randomized decision trees. A nonlimiting example is the ExtraTrees classifier, as described in the experimental section.

[0152] In some embodiments, selecting most informative ncRNAs includes using an ExtraTrees classifier.

[0153] In some embodiments, selecting most informative ncRNAs by the machine learning classifier includes running multiple iterations of the algorithm, and ranking at each iteration the ncRNAs based on their importance, determined by the frequency of its appearance in the top 30 features across all iterations.

[0154] In some embodiments, selecting most informative ncRNAs includes mapping genes to known cellular pathways. In some embodiments, the most informative ncRNAs are highly associated with cancer-specific pathways and / or PD AC-specific pathways.

[0155] In some embodiments, selecting most informative ncRNAs includes using differential expression analysis coupled with a gradient boosting algorithm.

[0156] Methods of treatment and methods of diagnosis by using a machine learning model

[0157] In some embodiments, there is provided a method of treating a subject suspected of being afflicted with PDAC, including:

[0158] a. quantifying in a bodily fluid sample of the subject levels of at least 3 ncRNA biomarkers including hsa-miR-4446-3p, hsa-miR-6073 , and NAAI DL2-ASR b. determining, based on the ncRNA biomarkers levels, whether the subject is likely afflicted with PDAC; and

[0159] c. treating the subject with an anti cancer treatment, when the subject is determined as likely to be afflicted with PDAC.

[0160] In some embodiments, the determining is done by applying to the ncRNA biomarkers levels a machine learning classification model optimized based on levels of the ncRNA biomarkers in PDAC patients and controls, or by comparing the levels of the ncRNA biomarkers in the sampleto control levels of the ncRNA biomarkers.

[0161] In some embodiments, there is provided a method for determining in a bodily fluid sample of a subject whether the subject is likely afflicted with PDAC, the method including:

[0162] a. quantifying in the bodily fluid sample levels of at least 3 ncRNA biomarkers including hsa-miR-4446-3p, hsa-miR-6073 , and NAALADL2-AS1,'

[0163] b. applying to the levels of the ncRNA biomarkers in the sample a machine learning classification model optimized based on levels of the ncRNA biomarkers in PDAC patients and controls to provide a diagnostic score; and

[0164] c. determining, based on the diagnostic score, whether the subject is likely afflicted with PDAC.

[0165] Definitions and embodiments mentioned above and which may be relevant to embodiments in this section also apply here, and vice versa. Some particularly relevant embodiments may be pointed out or explicitly repeated. For terms used herein, unless stated otherwise, their definition and embodiments are intended to be the same as above (mutatis mutandis).

[0166] The term “treating” or “treatment”, as used herein, refers to means of obtaining a desired physiological effect. The effect may be therapeutic in terms of partially or completely curing the cancer and / or symptoms thereof. The term encompasses both inhibiting the disease, i.e. arresting its development, and ameliorating the disease, i.e. causing regression of the disease, e.g., by eliminating or ameliorating its symptoms.

[0167] In some embodiments, the at least 3 ncRNA biomarkers are at least 4 ncRNA biomarkers, including hsa-miR-4446-3p, hsa-miR-6073, NAALADL2-AS1, and Y_RNA ENST00000365068.1.

[0168] In some embodiments, the at least 3 ncRNA biomarkers further include at least one additional ncRNA biomarker selected from Y RNA ENST00000365068.1, DLEU2, RP11-874J12.4, hsa-miR-432-5p, IncRNA C15orf54, hsa-miR-6721-5p, CTD-3092A11.2, Y RNA ENST00000516950.1, DPYD-AS2, hsa-miR-6131, misc RNA 7SK, KCTD21-AS1, IncRNA RP13-216E22.4, MIR223HG-201, hsa-miR-4433b-3p, IncRNA RP11-930011.2, VIM-AS1, and combinations thereof.

[0169] In some embodiments, the bodily fluid is cell-free.

[0170] In some embodiments, the method further includes normalizing the ncRNA biomarkers levels by using Trimmed Mean of M-values (TMM) to account for library size and composition bias.

[0171] In some embodiments, the quantifying includes sequencing the RNA, and aligning the RNA sequences to sequence databases for in silica filtering out contaminants, for in silica enriching forncRNAs, and / or for counting hits to quantify each ncRNA level.

[0172] In some embodiments, the machine learning classification model is an LGBM classifier optimized based on the ncRNA biomarkers.

[0173] In some embodiments, the diagnostic score is a probability value ranging from 0 to 1. In some embodiments, when the diagnostic score is above a predetermined threshold the subject is determined as likely to be afflicted with PDAC. In some embodiments, the predetermined threshold is at least about 0.5. In some embodiments, the predetermined threshold is about 0.5, 0.6, or 0.7.

[0174] Kits for diagnosing PDAC based on ncRNA biomarkers

[0175] In some embodiments, there is provided a kit for determining in a bodily fluid sample of a subj ect whether the subj ect is likely afflicted with PDAC, the kit including reagents for quantifying levels of at least 3 ncRNA biomarkers including hsa-miR-4446-3p. hsa-miR-6073 , and NAALADL2-AS1, in the bodily fluid sample.

[0176] Definitions and embodiments mentioned above and which may be relevant to the embodiments in this section also apply here, and vice versa. Some particularly relevant embodiments may be pointed out or explicitly repeated. For terms used herein, unless stated otherwise, their definition and embodiments are intended to be the same as above (mutatis mutandis).

[0177] In some embodiments, the at least 3 ncRNA biomarkers are at least 4 ncRNA biomarkers, including hsa-miR-4446-3p, hsa-miR-6073, NAALADL2-AS1 and Y_RNA ENST00000365068.1.

[0178] Reagents for quantifying levels of the at least 3 ncRNA biomarkers include at least agents which specifically recognize the 3 ncRNA biomarkers.

[0179] As noted above, methods involved in determining in a bodily fluid sample of a subject whether the subject is likely afflicted with PDAC based on levels of ncRNA biomarkers include isolating RNA from the sample, and quantifying the levels of the ncRNA biomarkers in the sample by well-known methods such as PCR, probe hybridization, and sequencing.

[0180] Accordingly, in some embodiments, the reagents further include non-specific reagents such as reagents for isolating RNA from bodily fluids, and reagents needed for quantifying the ncRNA biomarkers levels in the sample, such as PCR reagents, hybridization reagents, or sequencing reagents.

[0181] In some embodiments, the reagents for quantifying the ncRNA biomarkers levels include: (a) a probe specific to the mature sequence of hsa-miR-4446-3p; (b) a probe specific to the maturesequence of hsa-miR-6073; and (c) a probe and / or a primer pair specific for NAALADL2- AS 1. In some embodiments, the reagents further include (d) a probe and / or a primer pair specific for Y RNA ENST00000365068.1.

[0182] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the invention pertains.

[0183] The term "a" and "an" refers to one or to more than one (i.e., to at least one, or to one or more) of the grammatical object of the article. By way of example, “an element” means one element or more than one element.

[0184] The term "about", when referring to a measurable value such as an amount, a ratio, and the like, is meant to encompass variations of ±10% of the indicated value, as such variations are also suitable to perform the disclosed invention. Any numerical values appearing in the application are intended to be construed as if preceded by “about”, unless indicated otherwise.

[0185] While certain embodiments of the invention have been illustrated and described, it will be clear that the invention is not limited to the embodiments described herein. Numerous modifications, changes, variations, substitutions, and equivalents will be apparent to those skilled in the art without departing from the spirit and scope of the present invention as described by the claims, which follow.

[0186] The following examples are presented in order to more fully illustrate some embodiments of the invention. They should in no way be construed, however, as limiting the broad scope of the invention. One skilled in the art can readily devise many variations and modifications of the principles disclosed herein without departing from the scope of the invention.

[0187] EXAMPLES

[0188] Materials and Methods

[0189] Cohort aggregation and sample acquisition

[0190] A cohort of 122 viable samples was assembled (n = 122). This cohort included 79 control subjects (ncontroi = 79) and 43 PDAC patients (TIPDAC = 43), where five of the samples were sequenced twice in different stages of the disease (ntotai = 122). Out of the initial 123 samples procured, one was excluded due to unsuccessful RNA extraction, which resulted in unreadable outputs. Post-collection, and plasma isolation, the samples were preserved at -80°C. Each sample was accompanied by demographic data, including age, weight, height, and gender.

[0191] Some biases in the data were later partially mitigated during specification of covariates inthe differential gene expression linear model as described below.

[0192] Table 1. Summary of Demographic and Clinical Characteristics

[0193] Total Control PDAC

[0194]

[0195] BMI (Mean : SD) 25.62 : 54)0 2601 • 5. 12 2476 M64 Hadassah hospital 1 10 (9046%) 79 (100%) 31 (724)9%)

[0196] Sheba hospital 12 (9.84%) 0 (0%) 12 (2791%)

[0197] SD: standard deviation; BMI: body mass index;

[0198] Cell-free small RNA isolation from blood plasma

[0199] For the purification of cell-free nucleic acids from plasma samples, the QIAamp® Circulating Nucleic Acid Kit (Qiagen) was used, which provides a structured method for this purpose. The process begins with the lysis of plasma samples using proteinase K and Buffer ACL under denaturing conditions to release nucleic acids from their protein or vesicle-bound states. After lysis, Buffer ACB is added to adjust the binding conditions, allowing the nucleic acids to adhere to a silica membrane within the QIAamp® Mini columns. The bound nucleic acids are then washed with Buffer ACW1, Buffer ACW2, and ethanol to remove residual contaminants. Finally, the nucleic acids are eluted in Buffer AVE. This method is designed to handle large sample volumes and facilitate the recovery of nucleic acids, making it applicable for the sequencing of small cell-free RNA fragments. The samples’ output was assessed using random sampling using Qubit RNA High Sensitivity (Invitrogen) and Bioanalyzer RNA Small RNA chip (Agilent) to ensure minimal loss of material.

[0200] Library preparation for cell-free RNA-seq

[0201] For the purpose of turning the purified cell-free nucleic acids into a library suitable for sequencing, the SMARTer® smRNA-Seq Kit for Illumina (Takara) was used. This kit is specifically designed for the preparation of sequencing libraries from small RNA fragments. The process involves polyadenylation of the RNA, followed by cDNA synthesis using template switching technology, which incorporates adapters required for sequencing. After amplification and size selection to ensure the correct library size, the samples are prepared for sequencing. The library preparation was carried out using standardized protocols suitable for the sequencing of small cell-free RNA fragments. The cDNA libraries prepared were assessed with Bioanalyzer DNA High Sensitivity chip (Agilent) and samples that did not meet minimal criteria were excluded from analysis (1 sample).RNA-seq quality control, stepwise alignment and quantification

[0202] RNA-seq reads (FASTQ) were trimmed with fastp (v.0.23.4), quality assessed using FastQC (v.0.12.1) and visualized using MultiQC (v.1.21). Initial contamination filtering of ribosomal and transfer RNAs, as well as coding DNA (ENSEMBL release 86) was performed using Bowtie2 (v.2.4.5) allowing complete matches only. Identical sequential mapping procedure using Bowtie2 was performed to retrieve read mapping to piRNAdb and ENSEMBL’ s ncRNA database. Resulting non mapped reads were mapped to miRBase mature and hairpin sequences (release 22.1) using Bowtie. Quantification of mature miRNA sequences was facilitated using miRTOP (v.0.4.25). Quantification of piRNA and ncRNA was made via manual examination of the alignment files and SAMtools (v.1.19.2) commands.

[0203] Differential Gene Expression (DGE) analysis and count normalization

[0204] Aggregated quantifications were loaded into R and converted to edgeR (v.4.0) objects where counts were aggregated to the gene level from transcripts. Count filtration was performed using the filterByExpr function under default settings and resulting in 1,388 filtered genes. Count normalization was performed with edgeR weighted trimmed mean of the log expression ratios (TMM). Outlying samples were discarded from downstream Differential Gene Expression analysis (DGE). Discovery of the outlying samples was facilitated using The Minimum Covariance Determinant (MCD) method - a robust statistical technique used to estimate the parameters of a multivariate dataset. It is particularly useful for detecting outliers in highdimensional datasets, such as normalized count data. 20% contamination parameter offered sufficient balance between number of discarded samples and overall stability.

[0205] Subsequently, DGE comparisons were calculated with the following model:

[0206] -age + gender + input volume + hospital + label (Control r.s' PDAC) Significant differential expression was considered at adjusted P values of <0.01 and Log Fold Change values > 0.5 or < -0.5 resulting in 397 differentially expressed genes.

[0207] Feature importance and ranking

[0208] In the analysis process, the power of machine learning (ML) was leveraged to identify the most significant genes that could potentially serve as key features for downstream classification. To achieve this, the differentially expressed genes and the normalized count matrix were utilized as inputs for the ExtraTrees routine, a meta estimator that fits a number of randomized decision trees on various sub-samples of the dataset and uses averaging to improve the predictive accuracy and control over-fitting. The selection of optimal hyperparameters is a critical step in ML algorithms to ensure the best performance. In the present analysis, cross-validation was employed, a robust statistical method that provides an honest assessment of the model's predictiveperformance. To ensure the robustness of the results, the algorithm was iteratively run 1,000 times. In each iteration, the genes were ranked based on their importance. The importance of a gene was determined by the frequency of its appearance in the top 30 features across all iterations. After the completion of the iterative process, the results were compiled and the top 20 ranked genes were selected. These genes, which consistently appeared as the most important features across multiple iterations, were then used as the key features for downstream classification tasks. This approach ensures that the present model is built on the most informative and relevant features, thereby enhancing its predictive power and reliability.

[0209] Gene set enrichment and pathway analysis

[0210] In order to gain a deeper understanding of the biological significance of the top 20 genes identified from the feature importance routine, a comprehensive gene set enrichment and pathway analysis was conducted using NcPath enrichment analysis of human ncRNA and KEGG (release 111.0) signaling pathways. Mapping these genes onto KEGG pathways allowed to uncover the key biological pathways that these genes are involved in and to understand their functional context and validate their known role in PDAC related pathways.

[0211] Gradient boosting classification

[0212] To offer a classification procedure, a gradient boosting algorithm was utilize, specifically the Light Gradient Boosting Machine (LGBM) classifier. The training of the model was conducted using repeated cross-validation. This method is particularly effective in preventing overfitting and providing an honest assessment of the model's predictive performance. To further enhance the stability and accuracy of the present model, a 5-fold bagging approach was employed. This technique involves training the model on different subsets of the data and then averaging the predictions. This not only reduces the variance of the model but also helps in avoiding overfitting. The identification of the optimal number of features that best perform classification was achieved using a custom scoring equation. This equation calculates the mean and standard deviation of seven key metrics: balanced accuracy, accuracy, weighted Fl score, precision, and recall. The final score is the average of these seven metrics, providing a comprehensive evaluation of the model's performance.

[0213] Example 1: Enabling non-coding profiling of the cell-free RNA transcriptome

[0214] To quantify cell-free RNA-seq data from human plasma, multi-aligner stepwise sequence mapping procedure was developed, influenced by existing methods but unique to the usage intention of accurately and specifically retrieving signals that are previously known to bear high potential distinction between PDAC and healthy plasma samples. The importance of miRNAs asa potential high specificity feature compared to other known RNAs is emphasized by allowing up to two mismatches during mapping to miRBase sequence database, compared to full matches only while mapping to ENSEMBL non-coding sequences and piRNAdb. Subsequently read mappings from these diverse sources were aggregate and expression levels for tens of thousands of gene counts were quantified, while excluding specific and high-quality alignments to transfer and ribosomal RNA sequences (Fig. 1).

[0215] Following this, a TMM counts normalization and DGE analysis was performed to identify differentially expressed genes between PDAC and control samples (data not shown). The most differentially expressed gene for the mitochondrial tRNA (Mt tRNA) group is MT-TQ_ENST00000387372.1; The most differentially expressed genes for the microRNA (MiRNA) group are: hsa-miR-134-5p, hsa-miR-320c, hsa-let-7b-5p, hsa-let-7i-5p, hsa-miR-93-5p, hsa-miR-320a-3p, hsa-miR-432-5p, hsa-miR-342-5p, hsa-let-7d-5p, and hsa-let-7e-5p; The most differentially expressed genes for the PlWI-interacting RNA (piRNA) group are: hsa-piR-14786 and hsa-piR-11120; The most differentially expressed gene for the miscellaneous tRNA (misc RNA) group is DLEU2, 7SK; The most differentially expressed genes for the mitochondrial rRNA (Mt rRNA) group are: MT-RNR1 and MT-RNR2; The most differentially expressed genes for the long non-coding (Inc) RNA group are: C15orf54, ZNF667-AS1, RP11-930011.2, MIR4435-2HG, MIR4435-2HG, RP11-20D14.6, LINC00152, ZNF667-AS1, LINC00989, MALAT1, and LINC00989.

[0216] A robust ExtraTrees algorithm was employed to rank the importance of genetic features (ncRNAs) and train a gradient-boosted classification model on the top genetic features. This blend of bioinformatics and ML methodologies allowed achieving exceptional results that outperform existing approaches in classification tasks using plasma derived non-coding RNAs (ncRNA)s. The top 20 genes are presented in Table 2.

[0217] Table 2. Top 20 Ranked ncRNAs Identified by the ExtraTrees Routine.

[0218] Accession No. Mean Std Count Importance Importance Importance hsa-miR-4446-3p MIMAT0018965 0.014894 0.005917 954 hsa-miR-6073 MIMAT0023698 0.014332 0.005727 954 NAALADL2-AS1 ENST00000426315.1 0.011694 0.004016 948 Y RNA ENST00000365068.1 0.012318 0.004933 891 ENST00000365068.1

[0219] DLEU2 ENST00000614667.1 0.011427 0.004455 835

[0220] RP11-874J12.4 ENST00000577848.1 0.011055 0.004207 830 hsa-miR-432-5p MIMAT0002814 0.010819 0.004178 827 IncRNA C15orf54 ENST00000625107.1 0.009861 0.003446 813hsa-miR-6721-5p MIMAT0025852 0.011645 0.004544 805

[0221] CTD-3092A11.2 ENST00000602663.1 0.011123 0.004365 801 Y RNA ENST00000516950.1 0.011071 0.004198 794 ENST00000516950.1

[0222] DPYD-AS2 ENST00000422259.1 0.009608 0.003261 769 hsa-miR-6131 MIMAT0024615 0.011233 0.004398 760 misc RNA 7SK ENST00000365328.1 0.010671 0.004205 752 KCTD21-AS1 ENST00000600795.1 0.00875 0.00281 744 IncRNA RP13- ENSG00000271199 0.008391 0.002512 717 216E22.4

[0223] MIR223HG-201 ENST00000621933.1 0.010194 0.003838 712 hsa-miR-4433b-3p MIMAT0030414 0.010528 0.004083 702 IncRNARPll- ENST00000444482.1 0.009267 0.003152 698 930011.2

[0224] VIM-AS1 ENST00000592146.2 0.010141 0.003905 678

[0225] In Table 2: Accessions: ENST - Ensemble transcript; ENSG = Ensemble gene; MIM -miRBase; MIR: antisense IncRNA gene; miR: mature microRNA. For miRNA - hits may have up to 2 nucleotide (about 10%) difference compared to the accession number sequence. 'Mean Importance' represents the average importance score of each gene across 1,000 iterations. ' Std Importance' denotes the standard deviation of the importance score, providing an insight into the variability of each gene's importance. 'Count Importance' indicates the number of times each gene appeared in the top 30 features across all iterations, reflecting its consistency in being identified as a significant feature.

[0226] Table 3 presents performance metrics for the 1, 2, 3..20 top genes. The top-performing model includes the seven highest-ranking genes selected by the ExtraTrees algorithm (the first seven ncRNAs shown in Table 2), achieving 82% balanced accuracy (likelihood of correct diagnosis), 91% area under the receiver operating characteristic curve (AUC) (at its peak), and an Fl score of 86%, which is the weighted harmonic mean of precision and recall (Fig. 2G, Table 3). However, models including fewer ncRNAs also obtained very high results, e.g.: a model including the three highest-ranking genes in Table 2 achieved 84% balanced accuracy and an AUC of 89% (Table 3).

[0227] Table 3. Model Performance Metrics by Number of Selected Genes

[0228] Number of Balanced Fl Precision Recall ROCAUC MCC Top Genes Accuracy Weighted

[0229] 1 0.75 0.796 0.828 0.811 0.813 0.582 2 0.744 0.789 0.819 0.803 0.798 0.564 3 0.744 0.789 0.819 0.803 0.784 0.564 4 0.835 0.874 0.905 0.885 0.869 0.752 5 0.806 0.849 0.882 0.86 0.872 0.698 6 0.806 0.849 0.882 0.86 0.896 0.6987 0.817 0.858 0.888 0.869 0.909 0.716 8 0.811 0.85 0.878 0.86 0.897 0.696 9 0.811 0.85 0.878 0.86 0.908 0.696 10 0.811 0.85 0.878 0.86 0.899 0.696 11 0.799 0.841 0.87 0.852 0.902 0.676 12 0.799 0.841 0.87 0.852 0.902 0.676 13 0.799 0.841 0.87 0.852 0.906 0.676 14 0.81 0.85 0.876 0.86 0.906 0.692 15 0.805 0.849 0.882 0.86 0.902 0.698 16 0.776 0.822 0.859 0.836 0.909 0.642 17 0.787 0.831 0.864 0.844 0.906 0.659 18 0.783 0.83 0.871 0.844 0.905 0.665 19 0.776 0.822 0.859 0.836 0.911 0.642 20 0.781 0.823 0.853 0.836 0.902 0.638

[0230] In Table 3: ROC AUC: Receiver Operating Characteristic - Area Under the Curve; MCC: Matthews Correlation Coefficient

[0231] Figs. 3B-3E show a distribution of normalized counts for the four top genes in the 'Control' and 'PDAC groups. As can be seen, hsa-miR-4446-3p, hsa-miR-6073, and Y RNA ENST00000365068.1 are downregulated in PDAC compared to controls (Figs. 3B, 3C, and 3E, respectively), and Antisense NAALADL2-AS1 is upregulated in PDAC compared to controls (Fig. 3D). Additionally, the ncRNAs DELU2, RP11-874J12.4, and / or hsa-miR-432-5p were also found to be downregulated in PDAC compared to controls (data not shown

[0232] Example 2: The Differential Expression Analysis in more detail

[0233] The initial phase of the DGE analysis involved the application of Principal Component Analysis (PCA). PCA, a statistical procedure that employs an orthogonal transformation to convert a set of observations to a lower dimension was utilized. Despite its utility, the PCA plot did not exhibit a distinct unsupervised separation between the control and PDAC groups (Fig. 3A). This outcome highlighted the intricate and high-dimensional nature of the gene expression data, necessitating a more robust and flexible method such as supervised gradient boosting. Furthermore, the MCD method was employed for outlier detection and removal, with a majority of these outliers originating from the control group. MCD, a robust estimator of multivariate location and scatter, is particularly effective for outlier detection in high-dimensional datasets. The removal of these outliers aimed to enhance the robustness of subsequent analyses and ensure that the findings were not disproportionately influenced by these extreme observations.

[0234] A comprehensive DGE analysis was conducted with the objective of identifying statistically significant differentially expressed genes between PDAC and control samples. Following this, afeature selection process was implemented to extract the top-ranking candidates from these genes. Subsequently, the top 20 genes were subjected to gene set enrichment and pathway analysis. This served as an orthogonal validation method, assessing the genes' association with pancreatic and other cancer-related pathways (Fig.4). Notably, KEGG pancreatic cancer pathway emerged as the top-ranking pathway, with an adjusted P-value of 3.75e-12, significantly ahead of the subsequent pathways. In addition, several signaling pathways closely associated with cancer, aging, and cell proliferation also ranked highly. For instance, cellular senescence, a state of stable cell cycle arrest linked to aging and cancer, is highly correlated with PDAC, which is predominantly diagnosed in patients in their 70s. The MAPK signaling pathway, despite encompassing 1,025 genes, is a common pathway in cancer due to its strong association with cell proliferation.

[0235] Among the top differentially expressed genes are IncRNAs C15orf54, ZNF667-AS1, and DPYD-AS2 (Table 2). These genes have been previously described in the literature as potential prognostic biomarkers for gastrointestinal cancers and PDAC. However, in the ExtraTrees prioritization model, these IncRNA genes did not rank as the top three differentiators. Interestingly, ZNF667-AS1 did not even appear in the top 20, signaling its limited ability to act as a strong biomarker for PDAC. These findings demonstrate the increased precision and improved predictive potential for PDAC diagnostics of the biomarkers identified by the method of the invention.

[0236] Example 3: Correlation Analysis of Non-Coding Gene Expression and Clinical Features in PDAC Patients

[0237] The next aim was to assess correlation between the normalized read counts of the top 20 non-coding genes discovered in the ExtraTrees selection and the clinical attributes of the PDAC samples. This analysis was executed to investigate potential associations between gene expression and clinical parameters, which could elucidate the biological underpinnings of PDAC. Fig. 5 depicts a clustered heatmap that reveals a spectrum of correlation coefficients between the selected genes and the clinical features. Notably, BMI and weight display correlations peaking at -0.39 with certain genes, indicating a moderate inverse relationship. This inverse correlation suggests that an increase in BMI and weight is associated with a decrease in the expression of these genes, and vice versa. This observation could imply a potential involvement of these genes in metabolic processes associated with PDAC.

[0238] Interestingly, the IncRNA antisense MIR223HG-201 gene exhibits a correlation coefficient of -0.39 with BMI among the PDAC patients This moderate inverse relationship suggests that as the levels of antisense transcript MIR223HG-201 increase, the incidence or severity of PDAC tends to decrease, and vice versa. The role of the microRNA MIR223HG-201 in cancer is complexand varies across different types of cancer. It is typically repressed in hepatocellular carcinoma and leukemia, although higher expression levels of MIR223HG-201 are linked to colorectal and recurrent ovarian cancers, and pancreatic cancers. In some cases, MIR223HG-201 downregulation correlates with high tumor burden, disease aggressiveness, and poor patient prognosis.

[0239] Another gene of interest in the analysis is the small nuclear RNA (snRNA) 7SK, which exhibits a correlation coefficient of -0.34 with BMI. The 7SK snRNA is a ncRNA consisting of 331-333 base pairs. Previous studies have indicated that over-expression of 7SK snRNA promotes apoptosis in cancerous cells, suggesting its potential role as an endogenous anti -cancer agent. In the present study, the inverse correlation with BMI could imply a potential involvement of 7SK snRNA in metabolic processes associated with PDAC.

[0240] Conversely, gender and height demonstrate correlations approximating zero with the selected genes, indicating an absence of a linear relationship. Despite the late-onset nature of PDAC, the selected top 20 genes do not exhibit a significant correlation with age. This observation could suggest that the expression of these genes is not directly influenced by aging processes, but rather by other yet unidentified factors or mechanisms.

[0241] Example 4. Investigation of Predicted miRNA Target Proteins

[0242] The ExtraTrees routine to prioritize top candidate genes for classification, described above, resulted in 20 ncRNA candidates, from which six are miRNA downregulated in PDAC patients. Since miRNAs target multiple mRNAs and proteins, it was hypothesized that these miRNAs might lead to individual proteins of functional, diagnostic, or therapeutic significance. The predicted microRNA target protein list was generated by NcPath, which was filtered to only include validated interactions with a strong support entry on miRTarbase. The proteins that were found in the miRTarbase are CDKN1A, E2F1, E2F2, PIK3R1, BAK1, RAD51, and BCL2L1. These proteins are known to play significant roles in various biological processes and pathways. CDKN1 A, a cyclin-dependent kinase inhibitor, plays a crucial role in cell cycle regulation and is often found to be dysregulated in various cancers. E2F1 and E2F2, members of the E2F family of transcription factors, are involved in cell cycle control and apoptosis, and their aberrant expression is associated with tumorigenesis. PIK3R1 is a regulatory subunit of the PI3Ks, which are part of the PI3K / AKT / mT0R pathway, a critical pathway in cancer that regulates cell survival and growth. BAK1 and BCL2L1 are key players in the regulation of apoptosis, a process often disrupted cancer cells to evade cell death. RAD51 is involved in the repair of DNA double-strand breaks, and its overexpression has been linked to resistance to radiation therapy in several cancers.Example 5: diagnosing PDAC by ncRNA levels in a liquid sample

[0243] A peripheral blood sample is obtained from a subject suspected of being afflicted with PDAC, and plasma is separated via centrifugation at 3000 x g for 10 minutes at 4°C. Total cell-free RNA is isolated from the plasma using the QIAamp® Circulating Nucleic Acid Kit (Qiagen) according to the manufacturer’s instructions, including a proteinase K lysis step and elution in Buffer AVE (Qiagen). The isolated RNA is quantified for quality control using a Qubit™ RNA High Sensitivity assay (Thermo Fisher Scientific) and a Bioanalyzer RNA Small RNA chip (Agilent) to ensure the presence of suitable small RNA fragments.

[0244] The levels of a specific ncRNA signature primarily including hsa-miR-4446-3p, hsa-miR-6073, NAALADL2-AS1, and optionally additional ncRNAs from Table 2 are then measured via next generation sequencing (NGS) using the SMART er® smRNA-Seq Kit (Takara Bio) or by sequence-specific quantitative real-time PCR (qRT-PCR) using TaqMan® probes.

[0245] To ensure quantitative accuracy and allow for comparison across subjects, the resulting expression data are normalized using the TMM method to account for library size and composition biases. These normalized values are then processed by a diagnostic machine learning framework. Specifically, the expression data are analyzed using a classification algorithm, such as an LGBM classifier, that is trained to distinguish PDAC from non-PDAC controls based on the expression profiles of the identified ncRNA biomarkers. The model utilizes a 5 -fold bagging approach and repeated cross-validation to evaluate the multi-marker signature, ultimately generating a diagnostic risk score.

[0246] Based on the performance of this model in clinical cohorts, where it achieves a classification accuracy of 82% and an AUC of 91%, a score exceeding a predetermined threshold (such as 0.5) indicates a high probability of PDAC and warrants further clinical intervention.

Claims

CLAIMSWhat is claimed is:

1. A method for determining in a bodily fluid sample of a subject whether the subject is likely afflicted with pancreatic ductal adenocarcinoma (PDAC), the method comprising:a. quantifying in the bodily fluid sample levels of at least 3 non-coding RNA (ncRNA) biomarkers comprising hsa-miR-4446-3p, hsa-miR-6073 , and NAALADL2-AS1,' and b. determining, based on the ncRNA biomarkers levels, whether the subject is likely afflicted with PDAC.

2. The method of claim 1, wherein the at least 3 ncRNA biomarkers further comprise at least one additional ncRNA biomarker selected from Y RNA ENST00000365068.1, DLEU2, RP11- 874J12.4, hsa-miR-432-5p, IncRNA C15orf54, hsa-miR-6721-5p, CTD-3092A11.2, Y RNA ENST00000516950.1, DPYD-AS2, hsa-miR-6131, misc RNA 7SK, KCTD21-AS1, IncRNA RP13-216E22.4, MIR223HG-201, hsa-miR-4433b-3p, IncRNA RP11-930011.2, VIM-AS1, and combinations thereof.

3. The method of claim 1 or 2, wherein the determining comprises applying to the levels of the ncRNA biomarkers in the sample a machine learning classification model optimized based on levels of the ncRNA biomarkers in PDAC patients and controls, and / or comparing the levels of the ncRNA biomarkers in the sample to control levels of the ncRNA biomarkers.

4. The method of claim 3, wherein when hsa-miR-4446-3p and hsa-miR-6073 levels are lower and NAALADL2-AS1 level is higher in the sample of the subject compared to control levels, the subject is considered likely to be afflicted with PDAC.

5. The method of claim 3 or 4, wherein the at least 3 ncRNA biomarkers further comprise at least one additional ncRNA biomarker selected from Y RNA ENST00000365068.1, DELU2, RP11- 874J 12.4, and hsa-miR-432-5p.

6. The method of claim 5, wherein the at least one additional ncRNA biomarker comprises Y RNA ENST00000365068.1, DELU2, RP11-874J12.4, and hsa-miR-432-5p.

7. The method of claim 6, wherein when Y RNA ENST00000365068.1, DELU2, RP 11-874J12.4, and / or hsa-miR-432-5p levels are lower in the sample of the subject compared to control levels, the subject is considered likely to be afflicted with PDAC.

8. The method of any one of claims 3-7 , wherein the at least 3 ncRNA biomarkers further comprise MIR223HG-201.

9. The method of claim 8, wherein when MIR223HG-201 levels are lower in the sample of the subject compared to the control levels, the subject is considered likely to be afflicted with PDAC.

10. The method of any one of the preceding claims, wherein quantifying the levels of the ncRNA biomarkers comprises isolating RNA from the sample and quantifying levels of the ncRNA biomarkers in the isolated RNA.

11. The method of claim 10, wherein the RNA is cell-free RNA.

12. The method of claim 10 or 11, wherein the quantifying comprises sequencing the RNA.

13. The method of claim 12, wherein the quantifying further comprises in silico filtering out contaminants and / or in silico enriching for ncRNAs, by performing sequence alignments.

14. The method of any one of claims 3-13, further comprising normalizing the ncRNA biomarkers levels by using Trimmed Mean of M-values (TMM) to account for library size and composition bias.

15. The method of any one of claims 3-14, wherein the machine learning classification model is trained on levels of the ncRNA biomarkers obtained from subjects diagnosed with PDAC and from control subject not afflicted with PDAC.

16. The method of any one of claims 3-15, wherein the machine learning classification model is selected from a Light Gradient Boosting Machine (LGBM) classifier, an ExtraTrees classifier, a Random Forest classifier, a Support Vector Machine (SVM) classifier, and an XGBoost classifier.

17. The method of claim 16, wherein the machine learning machine learning classification model model is an LGBM classifier optimized based on the ncRNA biomarkers.

18. The method of any one of claims 3-14, wherein the control levels of the ncRNA biomarkers are obtained by quantifying the levels of the ncRNA biomarkers in a bodily fluid sample of a subject not afflicted with PDAC.

19. The method of any one of the preceding claims, wherein the bodily fluid sample is selected from blood, plasma, serum, cerebral fluid, urine, and combinations thereof.

20. The method of any one of the preceding claims, wherein the bodily fluid sample is cell-free.

21. The method of any one of the preceding claims, wherein the subject is a smoker.

22. The method of any one of the preceding claims, wherein the subject has a body mass index (BMI) of above 20, 30, or 35.

23. The method of any one of the preceding claims, wherein the subject is at least about 55, 60, or 65 years old.

24. The method of any one of the preceding claims, further comprising treating the subject with an anticancer treatment, when the subject is determined as likely to be afflicted with PDAC.

25. A method for training a machine learning model for determining whether a subject is likely afflicted with a cancer, comprising the steps of:a. performing a differential gene expression analysis on ncRNA levels from subjects diagnosed with the cancer and controls, to identify a plurality of ncRNAs that are differentially expressed between the subjects diagnosed with the cancer and the controls;b. selecting most informative ncRNAs by applying a machine learning classifier to the plurality of differentially expressed ncRNAs;c. evaluating randomized combinations of the most informative ncRNAs to identify a group of ncRNA biomarkers; andd. training a classification model based on the group of ncRNA biomarkers.

26. The method of claim 25, wherein the ncRNA levels are obtained by steps comprising:a. sequencing RNA isolated from bodily fluid samples of subjects diagnosed with PDAC and controls;b. filtering out contaminants and enriching for ncRNAs by performing sequence alignments; andc. quantifying levels of the ncRNAs.

27. The method of claim 25 or 26, wherein the differential gene expression analysis comprises applying minimum covariance determinant (MCD) to the ncRNA levels.

28. The method of any one of claims 25-27, wherein the differential gene expression analysis includes using supervised gradient boosting.

29. The method of any one of claims 25-28, wherein selecting the most informative ncRNAs includes using an ExtraTrees classifier.

30. The method of any one of claims 25-29, wherein selecting the most informative ncRNAs includes using a gradient boosting algorithm.

31. The method of any one of claims 25-30, wherein selecting the most informative ncRNAs includes mapping genes to known cellular pathways.

32. A method of treating a subject suspected of being afflicted with pancreatic ductal adenocarcinoma (PDAC), comprising:a. quantifying in a bodily fluid sample of the subject levels of at least 3 non-coding RNA (ncRNA) biomarkers comprising hsa-miR-4446-3p. hsa-miR-6073 , and NAALADL2- AS1- b. determining, based on the ncRNA biomarkers levels, whether the subject is likely afflicted with PDAC; andc. treating the subject with an anticancer treatment, when the subject is determined as likely to be afflicted with PDAC.

33. The method of claim 32, wherein the at least 3 ncRNA biomarkers further comprise at least one additional ncRNA biomarker selected from Y_RNA ENST00000365068.1, DLEU2, RP11-874J12.4, hsa-miR-432-5p, IncRNA C15orf54, hsa-miR-6721-5p, CTD-3092A11.2, Y RNA ENST00000516950.1, DPYD-AS2, hsa-miR-6131, misc RNA 7SK, KCTD21-AS1, IncRNA RP13-216E22.4, MIR223HG-201, hsa-miR-4433b-3p, IncRNA RP11-930O11.2, VIM-AS1, and combinations thereof.

34. The method of claim 32 or 33, wherein the determining is done by applying to the levels of the ncRNA biomarkers in the sample a machine learning classification model optimized based on levels of the ncRNA biomarkers in PDAC patients and controls, or by comparing the levels of the ncRNA biomarkers in the sample to control levels of the ncRNA biomarkers.

35. A kit for determining in a bodily fluid sample of a subj ect whether the subj ect is likely afflicted with PDAC, the kit comprising reagents for quantifying levels of at least 3 ncRNA biomarkers comprising hsa-miR-4446-3p. hsa-miR-6073 , and NAALADL2-AS1, in the bodily fluid sample.

36. The kit of claim 35, wherein the reagents comprise: (a) a probe specific to the mature sequence of hsa-miR-4446-3p; (b) a probe specific to the mature sequence of hsa-miR-6073; and (c) a probe and / or a primer pair specific for NAALADL2-AS1.