A miRNA-BASES METHOD FOR LUNG CANCER DETECTION

A 9-miRNA signature addresses the limitations of LDCT by providing a sensitive and specific blood test for lung cancer screening, enhancing diagnostic accuracy and reducing false positives.

WO2026114962A1PCT designated stage Publication Date: 2026-06-04FONDAZIONE DI RELIGIONE E CULTO CASA SOLLIEVO DELLA SOFFERENZA IRCCS - OPERA DI SAN PIO DA PIETRELCINA

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
FONDAZIONE DI RELIGIONE E CULTO CASA SOLLIEVO DELLA SOFFERENZA IRCCS - OPERA DI SAN PIO DA PIETRELCINA
Filing Date
2025-11-26
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Current lung cancer screening methods, such as low-dose computed tomography (LDCT), suffer from high false positive rates, radiation exposure concerns, and limited sensitivity, particularly in early-stage tumors, necessitating a more specific and sensitive diagnostic biomarker for early detection and differentiation between benign and malignant nodules.

Method used

A 9-circulating miRNA signature (c-miR signature) is developed through a multi-platform workflow and validated in large cohorts, identifying specific miRNAs (miR-29a-3p, miR-328-3p, miR-30c-5p, miR-200c-3p, miR-450b-5p, miR-218-5p, miR-184, miR-190b-5p, and miR-1233-3p) for early-stage lung cancer detection, using quantitative PCR or digital PCR for risk assessment.

Benefits of technology

The c-miR signature achieves 70-82% sensitivity and 67% specificity, accurately identifying individuals at high risk of lung cancer and distinguishing between benign and malignant nodules, reducing unnecessary LDCT screenings and improving diagnostic accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000015_0001
    Figure IMGF000015_0001
  • Figure IMGF000016_0001
    Figure IMGF000016_0001
  • Figure IMGF000016_0002
    Figure IMGF000016_0002
Patent Text Reader

Abstract

The present invention describes a method for diagnosing lung cancer in a subject by detecting specific circulating miRNAs ("c-miR signature") in a biological sample obtained from that subject. Such c-miR signature provides an indication and a prognosis of lung cancer in place of or in addition to alternative known methods, including, but not limited to, low-dose computed tomography (LDCT).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] “A miRNA-BASES METHOD FOR LUNG CANCER DETECTION”

[0002] The present invention describes a method for diagnosing lung cancer in a subject by detecting specific circulating miRNAs (“c-miR signature”) in a biological sample obtained from that subject. Such c-miR signature provides an indication and a prognosis of lung cancer in place of or in addition to alternative known methods, including, but not limited to, low-dose computed tomography (LDCT).

[0003] STATE OF THE ART

[0004] Lung cancer (LC) is the most common cancer worldwide, accounting for 2.5 million diagnoses per year. The high prevalence of advanced-stage disease at diagnosis (-80%), leads to relatively poor 5-year survival rates (<32%, ) for LC patients, and a global burden of 1.8 million of deaths per year (1). Tobacco smoking is the major risk factor for LC, with individuals who smoke cigarettes having a 15- to 30-fold increased risk of developing or dying from LC compared to lifetime never-smokers. Nonetheless, other risk factors exist, including a history of respiratory diseases / infections, exposure to asbestos, radon gas, and other carcinogens, living in heavily polluted areas, and maintaining an unhealthy diet (1). Therefore, LC primary prevention strategies rely mainly on smoking cessation and non-initiation campaigns, together with other highly recommended measures, such as establishing new regulations to ban asbestos and enforce radon disclosure, certification, and mitigation, as well as reducing air pollution.

[0005] On the other hand, LC secondary prevention strategies involve screening high-risk individuals annually with low-dose computed tomography (LD-CT). This approach has been shown to be effective resulting in approximately 20-30% reduction in LC mortality by increasing the chance of early-stage diagnosis and effective treatment (2). However, in addition to concerns about cost and radiation exposure, a significant burden of false positive findings (up to 28%) and overdiagnosis is associated with LD-CT screening (3).

[0006] This complicates the interpretation of LDCT results and ensuing decisions about the screening time interval: a difficulty that might be alleviated by a first-line screening test, such as the “c-miR signature”, that significantly reduces unnecessary LDCTs for individuals without lung cancer. Therefore, the identification of tumor biomarkers is highly desirable to complement the LD-CT screening program, by enhancing its specificity and reducing the rate of false-positives.

[0007] Recently, several types of circulating biomarkers in body fluids such as circulating tumor cells (CTC), circulating tumor DNA (ctDNA), circulating proteins, and tumor-derived extracellular vesicles and microRNAs have been proposed (4,5).

[0008] W02016 / 038119 and Montani et al. [7] discloses a method for diagnosing lung cancer in a subject by detecting a decrease and an increased abundance of different miRNAs in a blood sample obtained from that patient, the presence of which provides an earlier indication of cancer than alternative art-recognized methods, including, but not limited to, low-dose computed tomography (LDCT).

[0009] However, the lack of validation in large multi-center cohorts, lower sensitivity in early stage tumors, and technological limitation affect their clinical application (6). Additionally, there are very few on-going trials to validate these signatures (NCT02247453, NCT01248806, and NCT03452514).

[0010] Furthermore although, several studies have proposed different c-miRs, comprising of more than 100 different miRNA species included, with minor overlap among identified candidates (7, 8-12), the lack of consistency in the identified signatures, and limited application in clinical settings have been frequently debated. This is mainly due to poor study design, uncontrolled pre-analytic and analytic variability in c-miRs detection, lack of samples collected from real screening cohorts, and application of technologies hardly transferred to the routine clinical setting (6).

[0011] There is therefore a long-felt yet unmet need for diagnostic biomarkers and for a method to identify subjects and patients at high risk of developing lung cancer, that can be used both as a first line screening procedure to pre-select patients who require further diagnostic investigation by LDCT and as a second line screening procedure to distinguish between benign and malign nodules, in alternative or complementary to known methods.

[0012] DETAILED DESCRIPTION OF THE INVENTION

[0013] The inventors of the present invention have now surprisingly found that a 9 c- miRNA signature (i.e. the ‘c-miR signature’) can be used to diagnose subjects with lung cancer. Based on said signature, they have thus developed a new diagnostic in vitro or ex-vivo method that is a minimally invasive and relatively cheap blood test for use as a first-line screening procedure to pre-select patients who require further diagnostic investigation by LDCT and as a second line screening procedure to distinguish between benign and malign nodules.

[0014] Said method reduces the size of the target screening population, and it is undoubtedly advantageous in terms of costs, screening uptake rates and reduced medicalization of participants.

[0015] To tackle this, the inventors explored the serum / plasma circulating miRNome of lung cancer patients and high-risk individuals, by using a specifically designed multi-platform workflow with a multi-center design to identify a robust panel of circulating miRNAs (c-miRs) biomarkers diagnostics for early-stage Lung Cancer (LC), which they extensively validated in two large groups of high-risk individuals (smokers, >30 packs / year; aged, >50 years / old) from two distinct European LDCT screening cohorts. Assessing microRNAs in serum or plasma presents a promising approach to complement LC screening based on LD-CT scans, by potentially refining the eligibility criteria and improving the workup of nodules. In the present study, we identified a signature of 9 c-miRs with high diagnostic performance within the overall profiling of 276 lung cancer and 451 non-cancer controls. With -70% accuracy, 76%-82% sensitivity and -67% specificity, this molecular test demonstrates an high clinical utility (13). The test threshold of the method is currently based on the maximization of both sensitivity and specificity, but it can be further customized to achieve desired target of either sensitivity and or specificity.

[0016] Remarkably, our c-miR signature assigns LC risk to asymptomatic individuals, independently of other well-known predictors, including age, gender, smoking, nodule size and density. Moreover, the classification ability is maintained throughout different subset, with 63-73% of stage I LC correctly predicted as positive by the 9-c-miRs signature. These results, together with the ability of the test to distinguish malignant lesions from benign nodules, is particularly attractive in LC screening context.

[0017] The identified panel of c-miRs comprises miR-29a-3p, and miR-328-3p composing the previously proposed miR-Test (7), and miR-30c-5p, member of both the miR- Test and the MSC classifier (11). We also confirmed the presence of miR-200c-3p, miR-450b-5p, and miR-218-5p as diagnostic c-miRs, as previously proposed by Wozniak et al. (14) in their signature of 24 c-miRs. Additionally, we identified miR- 184, that was found in plasma extracellular vesicles (EVs), and differentially expressed between non-small cell lung cancer (NSCLC) and high-risk screening controls (15). Another marker that we identified is miR-190b-5p, which was previously included in a plasma panel of six microRNAs with high accuracy in discriminating LC from asymptomatic high-risk samples (16). Finally, we identified miR-1233-3p, which has not been fully explored in the context of LC diagnosis, but recognized as a potential circulating diagnostic marker in other cancers (i.e. ovarian cancer, renal cell carcinoma) and involved in immune- regulation.

[0018] The proposed new signature of said 9 c-miRs was therefore identified through stringent and powerful analytical approaches, comprising meta-analytic and machine learning techniques, careful evaluation of batch-bias of clustered data, application of different feature-selection methods, and evaluation of potential overfitting, thus conferring high generalizability and robustness to the identified predictors.

[0019] Another strength of our study lies in its comprehensive design and analytical approach, including both non-screening populations (meta-signature step), and screening populations (signature validation step). This approach accounts for different baseline LC risks, nodules sizes, and densities, tumor stages and histologies. This inclusive design facilitates the identification of c-miRs representative of high heterogeneity at cellular, histological, molecular and genetic levels peculiar to LC, by capturing information across many biological phenotypes, influenced by the tumor microenvironment and the immune system of the individual.

[0020] Moreover, the top-down strategy employed, which involved selecting the best performing diagnostic c-miRs from two screening cohorts, confers our signature the advantage of being fully arranged for a practical LC screening framework. Notably, the limited length of the signature, comprising of only 9 c-miRs and 3 normalizers, could facilitate its cost-effective implementation in clinical settings.

[0021] An embodiment of the present invention is therefore a method in-vitro or ex-vivo for identifying subjects affected by lung cancer, comprising the steps of: a) detecting the amount of each of the 9 miRNAs having sequence hsa-miR- 29a-3p (SEQ ID NO. 1), hsa-miR-30c-5p (SEQ ID NO. 2), hsa-miR-184 (SEQ ID NO. 3), hsa-miR-190b-5p (SEQ ID NO. 4), hsa-miR-200c-3p (SEQ ID NO. 5), hsa-miR-218-5p (SEQ ID NO. 6), hsa-miR-328-3p (SEQ ID NO. 7), hsa-miR-450b-5p (SEQ ID NO. 8) and hsa-miR-1233-3p (SEQ ID NO. 9) in a biological sample from a subject; b) normalizing the data obtained in step a).

[0022] According to further a preferred embodiment said method further comprises step c) of calculating a risk score.

[0023] Preferably when the miRNAs are detected by standard qPCR said risk score is calculated according to the following formula:

[0024] RS = -2.7465+hsa-miR-1233-3p*0.0622+hsa-miR-184*0.1151+hsa-miR- 190b-5p*(-0.0335)+hsa-miR-200c-3p*0.2769+hsa-miR-218-5p* (-0.0265)+hsa-miR-29a-3p*(-0.3578)+hsa-miR-30c-5p*1.5194 +hsa-miR-328-3p*(-0.9112)+hsa-miR-450b-5p*0.0709

[0025] Preferably when the miRNAs are detected by digital PCR (ddPCR) said risk score is calculated according to the following formula:

[0026] RS=-0.1725+hsa-miR-1233-3p*(-0.00128)+hsa-miR-184*(-0.00001)+hsa- miR-190b-5p*(-0.00008)+hsa-miR-200c-3p*0.000032+hsa-miR-218- 5p*(-0.00152)+hsa-miR-29a-3p*0.000002745+hsa-miR-30c- 5p*0.00000231+hsa-miR-328-3p*(-0.00002)+hsa-miR-450b-5p*(- 0.00012)

[0027] In a further preferred embodiment the method of the present invention further comprises step d) wherein the subject is classified as a subject at high risk, at intermediate risk or at low risk of developing lung cancer, based on the value of the tumor probability (i.e. ‘PROB TUM’).

[0028] Said tumor probability is derived from the risk score (RS).

[0029] Said tumor probability is calculated using the following formula: PROB_TUM=100*EXP(RS) / (1+(EXP(RS)).

[0030] Preferably, when the tumor probability is greater than or equal to 5% but lower than or equal to 60% it indicates that the subject has an intermediate risk of developing lung cancer, when the risk score is more than 60% the subject has an high risk of developing lung cancer and when the risk score is less than 5% the subject has a low risk of developing lung cancer.

[0031] According to a preferred embodiment said biological sample is a biological fluid or a tissue sample.

[0032] Preferably said biological fluid is whole blood, serum, plasma, saliva, urine or lymph fluid, more preferably said biological fluid is serum or plasma.

[0033] Preferably said tissue sample is a fresh tissue sample or a frozen tissue sample. According to a further preferred embodiment, in the method of the present invention the miRNAs in step a) is detected by quantitative qPCR, digital PCR (ddPCR), RNA sequencing (RNA-seq), Affymetrix microarray or custom microarray, or digital detection through molecular barcoding selected from NanoString technology.

[0034] Preferaby, the miRNAs in step a) is detected by quantitative PCR (qPCR) or digital PCR (ddPCR).

[0035] According to a preferred embodiment when the miRNAs are detected with qPCR the normalization of the data reported in step b) is made by using three miRNAs normalizer selected from miR-16-5p (SEQ ID NO: 10), miR-19a-3p (SEQ ID NO: 11) and miR-19b-3p (SEQ ID NO: 12). According to a preferred embodiment when the miRNAs are detected with ddPCR the normalization of the data reported in step b) is made by using two miRNAs normalizer selected from miR-16-5p (SEQ ID NO: 10) and ath-MIR159a (SEQ ID NO: 13).

[0036] Preferably said subject is a asymptomatic individual, high risk individual or a patient affected by lung cancer.

[0037] More preferably said patient affected by lung cancer is at an early-stage lung cancer, at a stage I lung cancer, at a stage II lung cancer, or with locally-advanced lung cancer i.e. Stage IIIA.

[0038] Said risk individual has a 20 pack-year or more smoking history, and smoke now or have quit within the past 15 years and are between 50 and 60 years old.

[0039] According to a preferred embodiment, the method of the present invention is a first- line screening procedure to pre-select patients or subjects who require further diagnostic investigation by LDCT.

[0040] According to a further preferred embodiment, said method is a second-line screening procedure to distinguish between benign and malign nodules in a subject. In a further preferred embodiment, the method according to the present invention has an accuracy between 65% and 80%, preferably about 70%.

[0041] Preferably, said method has a sensitivity between 70% and 90%, more preferably between 76% and 82%.

[0042] Preferably, said method has a specificity between 60% and 80%, more preferably about 67%.

[0043] Preferably, the subjects of the method according to the present invention may present one or more risk factors for developing lung cancer. Exemplary risk factors for developing lung cancer include, but are not limited to, a personal or family history of cancer, a history of smoking and / or exposure to second-hand smoke, and / or having limited access to preventative or curative medical care.

[0044] Alternatively, or in addition, a subject who is exposed to fine particulates from his or her environment is at an increased risk of developing lung cancer. For example, an individual who is exposed to smoke (including secondhand smoke), radon gas, asbestos and other chemicals (e.g. arsenic, chromium, and / or nickel), and / or nanosize particulates (e.g. dust and particulates from a manufacturing facility and / or motor vehicle exhaust) are at an increased risk of developing lung cancer. The combination of a family history of cancer and any one or more of the risk factors described herein may further increase an individual's risk of developing lung cancer.

[0045] In a further preferred embodiment, the method of the present invention further comprises a step of performing low-dose computed tomography (LDCT) or referring the subject for LDCT if the subject is diagnosed with lung cancer.

[0046] Preferably the method of the present invention further comprises a step of providing a treatment or referring the subject for treatment if the subject is diagnosed with lung cancer.

[0047] According to a further preferred embodiment the method of the present invention correctly predicted between 60% to 75% stage I LC as positive in a group of patients, preferably between 63% to 73% stage I LC as positive in a group of patients (at the cut-off maximizing the Youden index).

[0048] According to a preferred embodiment, the method of the present invention identifies patients included in the group of aggressive lung cancer with a poor prognosis, selected from patients with shorter overall survival and and / or patients with shorter disease-free survival, and / or patients responsive to a treatment, and / or with patients with metastatic disease which can include, but not limited to, patients with early-stage disease (stage I). According to a further preferred embodiment, the method of the present invention identifies patients included in the group of non aggressive lung cancer with a goodprognosis, selected from patients with longer overall survival, and / or patients with longer disease-free survival, and / or patients responsive to treatment, patients and / or without metastatic disease which can include, but not limited to, patients with early- stage disease (stage I).

[0049] According to a further preferred embodiment, the method of the present invention is used for the prognostic risk stratification of patients with lung cancer and / or to identify alternative therapeutic options after surgery, selected from systemic adjuvant chemotherapy, selected from platinum-based combinations, preferably cisplatin, carboplatin plus a third generation agents such as gemcitabine, vinorelbine, a taxane or camptothecin, molecular targeted therapeutics, immunotherapeutic, radiotherapy, or a combination thereof.

[0050] According to a further preferred embodiment, the method of the present invention it is a first-line screening procedure to pre-select subjects at risk of developing lung cancer.

[0051] FIGURES

[0052] Figure 1A-E. Fig. 1. A) Schematic representation of the study. B) Hierarchical clustering analysis of the 321 c-miRs commonly identified in all 3 c-miR expression datasets (GSE64591, GSE46729, GSE68951). Data (arrays) were median centered. Colors are as per the legend. C) Volcano plot for the 321 c-miRs common to the 3 datasets GSE64591, GSE46729, GSE68951. Log2 fold-change and -loglO proportion of false positive (pfp) are reported, as per RankProd non-parametric method. Each dot represents one miRNA. Selected 45 c-miRs differentially expressed (pfp<0.05) are highlighted in red (upregulated) or in blue (downregulated) in tumor vs. normal samples; D) ROC curves and AUC, for the model including 45 c-miRs and 36 c-miRs selected for the validation study, applied to the IARC dataset (GSE64591); E) Hierarchical clustering analysis of the 45 c- miRs expression (data were median-array-centered) in the pools (N=6) of samples (N=108) collected at IRCCS Casa Sollievo della Sofferenza Hospital (CSS) and Humanitas Research Hospital (HUM). Colors are as per the legend. On the right, bubbles represent the different criteria applied (as per the legend) to identify the 36 c-miRs. Highlighted in green, the 29 c-miRs selected (reliable detection in Step 2 analysis), and in yellow, the remaining 7 c-miRs selected in Step 3 and GSE64591 analysis.

[0053] Figure 2A-F. Fig. 2. A) Hierarchical clustering analysis of the median 36 c-miRs and 13 c-miRs (external signature) expression on real -word multi-center LD-CT screening cohorts of high-risk subjects (Step 4 analysis). Colors are as per the legend. B) ROC curves, AUC, and optimism-adjusted AUC (200 bootstrap) for the 9-c-miRs model in the following cohorts: MUG screen-detected lung cancer (LC) and normal controls (N), HUM screen-detected lung cancer (LC) and normal controls (N), and IARC GSE64591 lung cancer (LC) and normal controls (N). C) ROC curve and AUC for the 9-c-miRs model applied to the HUM cohort: screen- detected lung cancer (LC) and benign (BEN). D) Distribution of the probability of having lung cancer (tumor predicted probability) using the 9-c-miRs model in MUG and HUM cohorts; green lines represent the mean values. E) Distribution of tumor predicted probabilities using the 9-c-miRs model in the various LC subtypes (adenocarcinoma, AC; squamous cell carcinoma, SCC; and other subtypes, Other); green lines represent the mean values. F) ROC curves and AUC for the 9-c-miRs model applied to stage I disease only in MUG and HUM cohorts.

[0054] Figure 3. Flow chart of the study design.

[0055] Figure 4. AUC in the MUG and HUM cohorts when the model based on 13 c-miRs composing our previously derived miR-Test was applied.

[0056] Figure 5. Table S4. List of the 36 c-miRs and 45 c-miRs.

[0057] Figure 6. Table 1. Patients and tumors characteristics of the multi -center study. MUG=Medical University of Gdansk; HUM=Humanitas Research Hospital; LC=Lung cancer; N=normal; BEN=benign.

[0058] DEFINITIONS

[0059] Unless otherwise defined, all terms of art, notations and other scientific terminology used herein are intended to have the meanings commonly understood by those persons skilled in the art to which this disclosure pertains. In some cases, terms with commonly understood meanings are defined herein for clarity and / or for ready reference; thus, the inclusion of such definitions herein should not be construed to represent a substantial difference over what is generally understood in the art.

[0060] The term “microRNA" or “miRNA” or “c-miRs” used herein refers to a small noncoding RNA molecule of about 22 nucleotides found in plants, animals and some viruses, that has role in RNA silencing and post-transcriptional regulation of gene expression. miRNA exerts its functions via base-pairing with complementary sequences within mRNA molecules.

[0061] The term “signature" herein refers to an expression pattern derived from combination of several miRNA (i.e. transcripts) used as biomarkers.

[0062] The term “A UC” herein refers to the area under the ROC curve (Receiver Operating Characteristic curve), that is a graph showing the performance of a binary classification model at various classification thresholds.

[0063] The term “aggressive” herein refers to a cancer diagnosed in patients with an adverse prognosis (i.e. with shorter overall survival, and / or with shorter disease- free survival, and / or responsive to a treatment, and / or with metastatic disease).

[0064] The term “prognostic / prognosis” herein refers to the ability to discriminate patients with good / poor prognosis.

[0065] The term “first-line screening procedure" herein refers to the initial test or examination used to identify individuals who may have a particular disease or condition before any symptoms appear, and before more specific or invasive diagnostic tests are performed.

[0066] The term “second-line screening procedure" herein refers to a follow-up test such as LD-CT, performed after an initial (first-line) screening procedure yields a positive or suspicious result, in order to confirm, refine, or exclude the presence of disease.

[0067] The term “LD-CT" herein refers to low-dose computed tomography.

[0068] The term “biomarkers" (short for biological markers) herein refers to biological indicators (for example a transcript, i.e. miRNA) and / or measures of some biological state or condition.

[0069] The terms "comprising", "having", "including" and "containing" should be understood as 'open' terms (i.e. meaning "including, but not limited to") and should also be deemed a support for terms such as "consist essentially of, "consisting essentially of, " consist of ', or "consisting of.

[0070] EXAMPLES

[0071] 1. Meta-signature identification

[0072] To develop a robust circulating miRNA (c-miRs) diagnostic signature, we devised a strategy based on four main steps (Figure 1 A; Figure 3). In Step 1 (meta-signature identification), we performed an in-silico analysis of publicly available datasets of c-miRs expression profile (GSE64591, GSE46729, GSE68951) downloaded from Gene Expression Omnibus (GEO) database, including 150 LC patients and 136 cancer-free subjects (Table SI), whose plasma / serum samples were analyzed with different screening platforms (i.e., low-density qRT-PCR arrays, microarray). A total of 321 c-miRs were commonly identified in all three datasets (FigurelB; Table S2) including 5 c-miRs (miR-15b-5p, miR-19a-3p, miR-19b-3p, miR-24-3p, miR- 197-3p) we previously used as housekeeping c-miRs (7), and then analyzed by the non-parametric RankProd method (see methods), to combine data sets from different origins (meta-analysis) and increase the power for differentially expressed c-miRs identification (Table S2; see methods). Hence, we found 45 c-miRs differentially expressed (percentage of false predictions, pfp <0.05) between lung cancer and normal controls (Figure 1C; Table S2). We then verified the performance of the 45-c-miRs signature by applying unconditional logistic regression to the largest dataset (GSE64591, N=100 lung cancer, 100 normal controls), thus modeling the odds of LC as a function of the 45 c-miRs, which led to the result of AUC=0.87 (95%CI: 0.83-0.92) (Figure ID).

[0073] Table SI. Clinico-pathological characteristics of the 3 datasets GSE64591, GSE46729, GSE68951.

[0074] Table S2. List of 321 c-miRs common to the 3 datasets GSE64591, GSE46729, GSE68951. Fold change (FC) and proportion of false positive (PFP) are reported as per RankProd non-parametric method. List of selected 45 c-miRs (pfp<0.05) is also indicated.

[0075]

[0076] Next (Step 2 - pilot study), we further analyzed the 45-c-miRs signature in plasma samples from a total of 54 LC patients and 54 non-cancer controls (Table S3).

[0077] Table S3. Clinico-pathological characteristics of the plasma cohorts used in the pilot study.

[0078] Notably, among these 45 c-miRs, miR-197-3p was then used as housekeeping c- miR as previously explained. This cohort of LC patients and relative controls were retrospectively collected at IRCCS Casa Sollievo della Sofferenza Hospital (San Giovanni Rotondo, Italy) and Humanitas Research Hospital (Milan, Italy). Plasma samples were pooled (LC pools, N=6; CTRL pools, N=6) and profiled by a custom TaqMan low-density array (see methods) to rank c-miRs based on reliable qRT- PCR data (i.e., <35 Ct in at least 50% of pools) which resulted in 29 markers (15) (Figure IE; Figure 5: Table S4).

[0079] Finally (Step 3), in order to avoid any possible selection bias due to underrepresentation of some c-miRs in the 12 pools of samples analyzed, we measured the 45 c-miRs intracellular expression profile in a total of 173 LC cell lines available in the Cancer Cell Lines Encyclopedia (CCLE), together with the evaluation of the effect of 45 c-miRs on LC risk (univariate and multivariable logistic regression) in the largest dataset of plasma samples used in Step 1 (GSE64591). After scoring c-miRs based on detection in pools of samples, expression in LC cell lines, and association with LC risk (see criteria highlighted in Figure 5: Table S4 and Figure IE), we added 7 c-miRs to the 29 c-miRs list, resulting in a total of 36 c-miRs. Notably, the refined 36-c-miRs signature showed an analogous diagnostic performance (AUC=0.86, 95%CI: 0.81-0.91) of the 45-c- miR signature (AUC=0.86 vs. AUC=0.87, respectively; Figure ID).

[0080] 2. Analysis of c-miRs diagnostic for LC in a multi-center large cohort of LD- CT screened high-risk individuals

[0081] We then explored the diagnostic performance of such 36-c-miR signature in real- word multi-center LD-CT lung cancer screening cohort of high-risk subjects (smokers, >30 packs / year; aged, >50 years / old) (Step 4-signature reduction and validation). This cohort, collected at Medical University of Gdansk (MUG, Poland; MOLTEST-BIS trial (20)) and Humanitas Research Hospital in Milan (HUM, Italy; SMAC-1 trial (21)), included 72 screen-detected LC cases and 261 normal and benign controls (Figure 5: Table S4).

[0082] We used custom OpenArray™ qRT-PCR technology to profile plasma concentration of 36 c-miRs, which allowed us to profile all samples (N=333) of this large multi -center cohort in a single batch of 7 arrays, and in a short timeframe (< 3 days), thus limiting technical variability (batch effect) during the analysis of samples. We also included in the analysis other 13 c-miRs (aka miR-Test) that we had previously showed to accurately diagnose LD-CT detected LC by analyzing serum samples (7). We selected three normalizers (miR-16-5p, miR-19a-3p and miR-19b-3p) by correlating the expression profile analysis of six housekeeping c- miRs previously identified (7) with the reference miR-16-5p (Table S5).

[0083] Table S5. Correlation matrix (Pearson's correlation) and significance for reference c-miRs in the multicentric study.

[0084] After the normalization process, we analyzed the c-miRs expression distribution across centers (Figure 2A. Although differences in c-miRs expression across centers and group (cases / normal) were not consistently significant, we proceeded with analyses stratified by center or adjusted by center to account for possible center-effect. We then applied logistic regression (see methods) using the full screening cohort (both MUG and HUM individuals) with center correction, or restricting the analysis to the larger Polish cohort (MUG), in order to assemble a c- miRs diagnostic signature for LD-CT detected LC (Table S6).

[0085] Table S6. Diagnostic performance of the c-miRs signatures in different subsets of the multicentric study. Unconditional logistic model was fitted to each subset. Area under the curve (AUC), optimism-adjusted AUC (200 bootstrap), sensitivity (SE), specificity (SP), positive predictive value (PPV) and negative predictive value (NPV) are reported. The cut-off corresponding to the maximum Youden index

[0086] (with SE>SP) was considered. MUG=Medical University of Gdansk; HUM=Humanitas Research Hospital; LC=Lung cancer; N=normal.

[0087] Using a stepwise approach in the MUG cohort, we identified a 9-c-miRs signature

[0088] (hsa-miR-29a-3p, hsa-miR-30c-5p, hsa-miR-184, hsa-miR-190b-5p, hsa-miR- 200c-3p, hsa-miR-218-5p, hsa-miR-328-3p, hsa-miR-450b-5p, hsa-miR-1233-3p), with an AUC of 0.78 (SE, 76%; SP, 67%; ACC=70%) (Table 2; Table S7, Figure 2B).

[0089] Table 2. Diagnostic performance of the 9-c-miRs signature in different subsets of the multi-center study. Unconditional logistic model was fitted to each subset. Area under the curve (AUC), optimism-adjusted AUC (200 bootstrap), accuracy (ACC) sensitivity (SE), specificity (SP), positive predictive value (PPV) and negative predictive value (NPV) are reported. The cut-off corresponding to the maximum Youden index (with SE>SP) was considered. MUG=Medical University of Gdansk; HUM=Humanitas Research Hospital; LC=Lung cancer; N=normal.

[0090] Table S7. The 9 miRNA signature

[0091] Next, we validated this signature in the independent Italian cohort (HUM), where the 9-c-miRs signature demonstrated an AUC of 0.75 (SE, 82%; SP, 68%; ACC=71%) (Table 2; Table S8; Figure 2B). Table S8. Unconditional logistic model regression coefficients for the 9 c-miRs signature. Odds ratio (OR) and 95% Confidence Interval (95% CI) are reported. MUG=Medical University of Gdansk; HUM=Humanitas Research Hospital; LC=Lung cancer; N=normal.

[0092]

[0093] A stepwise approach with center correction or penalized approach did not allow to obtain signatures with better performance. Notably, the model based on 13 c-miRs composing our previously derived miR-Test (7) showed an AUC of 0.68 (SE, 66%; SP, 59%; ACC=61%) and AUC of 0.80 (SE, 86%; SP, 59%; ACC=66%) in the MUG and HUM cohorts, respectively (Figure 4). Therefore, while in the HUM cohort the performance of the previously developed miR-Test was similar to that of the newly identified 9-c-miR signature (AUC 0.80 vs. 0.75, respectively), in the larger MUG cohort the 9-c-miR signature demonstrated superior overall diagnostic performance compared with the miR-Test (AUC 0.78 vs. 0.68).

[0094] In conclusion, the 9-c-miR test was the only one to achieve better performance when validated across large, multi-center cohorts.

[0095] To further validate robustness of this 9-c-miRs signature across different miRNA profiling platforms and cohorts, we applied it in the IARC dataset (GSE64591), obtaining an AUC=0.78 (0.71-0.84) comparable to the results of the multi-center cohort (Figure 2B). Moreover, the 9-c-miRs showed a good discriminatory power even when comparing LC to benign nodules (AUC=0.71, 0.58-0.84, Figure 2C) and a remarkable separation of tumor predicted probability distributions between LC and controls in all the evaluated cohorts (Figure 2D), while no significant difference was found across LC subtypes (Figure 2E).

[0096] 3. Multivariate and subgroup analysis of 9-c-miRs LC signature in the Polish and Italian multi-center cohort We then investigated if the diagnostic accuracy of 9 c-miRs was independent from other individual- and nodule-related risk factors. Multivariable models showed significant odds ratios (OR) for 5% increase in 9-c-miRs LC predicted probability, after adjustment for age, gender, smoking (status and intensity), nodule size and density (OR=1.27, 95%CI: 1.07-1.50 in MUG cohort; OR=1.51, 95%CI: 1.16-1.95, in HUM cohort, Table 3). Subgroup analyses revealed that the 9-c-miRs risk model worked well in all the subsets considered, with a few non-significant associations due to small sample size (Table 4).

[0097] Table 3. Univariate and multivariable logistic model for the 9-c-miRs signature in the multi-center study. MUG=Medical University of Gdansk; HUM=Humanitas Research Hospital; LC=Lung cancer; N=normal; SCR=screening.

[0098] Table 4. Subgroup analysis for 9-c-miRs logistic model in the multi-center study.

[0099] MUG=Medical University of Gdansk; HUM=Humanitas Research Hospital;

[0100] LC=Lung cancer; N=normal; SCR=screening.

[0101] Next, we assessed the diagnostic power of the 9-c-miRs for detecting stage I LC in both centers. Despite, the small sample size of the HUM cohort (only 8 stage I LC) which limited the generalizability of the findings, multivariable analysis revealed that every 5% increase in predicted probability of the 9-c-miRs signature was associated with a 29% increase in the odds of having stage I LC in the Polish cohort (p<0.0001) and 25% increase in the Italian cohort (p=0.0224) (Table 4). Consistent with these findings, ROC curves restricted to stage I LC showed AUC=0.76 (0.68- 0.84), and 0.69 (0.49-0.89) in MUG and HUM, respectively (Figure 2F, Table 4). Additionally, the 9-c-miRs signature correctly predicted 27 of 37 (73%) stage I LC as positive in the MUG cohort, while in the HUM cohort, it identified 5 of 8 (62.5%) stage I LC as positive (at the cut-off maximizing the Youden index) (Table S9).

[0102] Table S9. 9 c-miRs prediction in stage I tumors and benign. The cut-off maximizing the Youden index was applied to define positive and negative calls (MUG cut- off=25%; HUM cut-off=22.5%). MUG=Medical University of Gdansk; HUM=Humanitas Research Hospital; LC=Lung cancer; BEN=benign.

[0103] We further challenged our 9-c-miRs signature in independent cohort of 277 patients (residents in US; FUTCH-cohort) with lung cancer (N=115) or with benign lung nodules (N=162). In addition, we employed an alternative technology, digital quantitative PCR (ddPCR), to further assess the reproducibility of the 9-c-miR signature across different analytical platforms. Data were normalized using a reduced set of reference miRNAs (has-miR-16-5p, and ath-MIR159a (SEQ ID NO: 13)), with the aim of minimizing both the complexity and the cost of the 9-c-miR assay. ROC analysis yielded an AUC of 0.66 (Figure 6) in this additional U.S. cohort of patients with lung cancer versus those with benign disease, which is consistent with the results previously obtained in the MUG cohort when comparing lung cancer patients with individuals presenting benign lung nodules (AUC = 0.71; Figure 2C).

[0104] MATERIALS AND METHODS

[0105] Study design

[0106] We structured the identification of the diagnostic c-miRs signature into four main Steps, as depicted in Figure 1A: Step 1 - meta-signature identification; Step 2 - pilot study; Step 3 - cancer cell lines; Step 4 - signature reduction and validation (multi-center European screening study).

[0107] Meta-signature identification

[0108] C-miRs expression analysis was performed by using the following publicly available datasets: GSE64591, GSE46729, and GSE68951 (Gene Expression Omnibus database, https: / / www.ncbi.nlm.nih.gov / geo / query / acc.cgi). Plasma or serum samples for a total of 150 lung cancer and 136 normal controls were profiled through TaqMan Human MicroRNA Array A + B Card Set v3.0, Applied Biosystems (GSE64591), Affymetrix Multispecies miRNA-1 Array, Thermo Fisher Scientific (GSE46729), and Agilent-031181 Unrestricted Human miRNA V16.0 Microarray, Agilent Technologies, Inc. (GSE68951). Non-parametric RankProd method (17) was applied to combine datasets from different origins (meta-analysis) and identify differentially expressed c-miRs (pfp<0.05) between lung cancer and normal controls (meta-signature). Pilot and validation cohorts selection

[0109] After approval from the Institutional Review Board (Medical University of Gdansk approval numbers NKEBN / 42 / 2009 and NKBBN / 376 / 2014; and Humanitas Clinical and Research Center approval number CE Humanitas ex DM 390 / 18; Fondazione IRCCS Casa Sollievo della Sofferenza approval number BIO- POLMONE - V1.0 08 Giu 16), informed consent was obtained from all the participants. Additionally, our study was conducted in accordance with the Declaration of Helsinki.

[0110] Pilot study cohort: A preliminary detection analysis of c-miRs composing the metasignature was performed in a clinical cohort, including retrospectively collected plasma samples from Casa Sollievo della Sofferenza Hospital, San Giovanni Rotondo, Italy (N=24 lung cancer, N=24 controls) and Humanitas Research Hospital, Milan, Italy (N=30 lung cancer, N=30 controls). We randomly split the cohort into 6 subsets of lung cancer and 6 subsets of controls, with proportional allocation of samples for each center. Specifically, each subset included 8 samples from Casa Sollievo della Sofferenza Hospital and 10 samples from Humanitas Research Hospital, respectively.

[0111] Validation cohort - Multi-center European LD-CT screening study. Two LD-CT screening cohorts of high-risk subjects were enrolled at Humanitas Research Hospital, Milan (HUM, Italy; SMAC-1 trial, NCT04315766) and Medical University of Gdansk (MUG, Poland; MOLTEST-BIS trial), and contributed to the CLEARLY project (https: / / transcan.eu / output-results / funded-projects / clearly.kl) funded by TRANSCAN-2 (JTC 2016). The screening inclusion criteria and protocols were described here (20, 21). A total of 72 lung cancers, 221 normal controls, and 40 benign lung lesions were analyzed.

[0112] Plasma collection MUG (Poland) protocol: a total of 10 ml of blood was collected in EDTA- containing vacutainer tubes. Of this, 1 ml (2 x 500 pl) of blood was then ali quoted into cryovials and stored at -80°C. The remaining 9 ml of blood was centrifuged at 600 x g for 20 minutes at 4°C to separate it into three layers: i) plasma (top); ii) white blood cells (buffy coat, middle); iii) and erythrocytes (bottom). The plasma supernatant was carefully aspirated and pooled into a centrifuge tube, followed by a second centrifugation at 1500 x g for 15 minutes at 4°C. After this step, 6 x 500 pl of plasma was aliquoted into labeled cryovials and stored at -80°C. The white blood cells (buffy coat) were collected from the initial centrifuge tube and also stored at -80°C.

[0113] HUM (Italy) protocol: the first 3 ml of blood drawn from the patient was discarded to prevent potential skin contamination, and was not used for plasma preparation. The remaining blood sample (minimum 4.5 ml) was collected in tubes with 0.129M Na Citrate anticoagulant (BD Vacutainer Sodium Citrate tubes - 363079 - Light blue top). The samples were centrifuged within 2 hours and 30 minutes at 1300g for 10 minutes at room temperature. The plasma was then carefully transferred into a 1.5 ml microtube, ensuring the interface was not touched and then subjected to a second centrifugation at 1300g for 10 minutes at room temperature. After this, the plasma was transferred into a 15 ml Falcon tube by gentle pipetting and then 3 x ~0.4 ml aliquots of plasma immediately dispensed into 0.5 ml Cryobank 2D coded tubes (Thermo Fisher Scientific, Cryobank vials, 2D coded, racked, blue cap; Cat. No. 374025) and placed on dry ice. The aliquots were kept on dry ice at all times before being transferred to a dedicated -80°C freezer for storage. The samples were also shipped on dry ice.

[0114] To identify and eventually exclude hemolytic samples that could negatively affect c-miRs profiling (18), we analyzed the hemolysis index as previously described (19). The absorbance peaks at 414 nm, and 385 nm, along with the hemolysis index were recorded in the database. Samples with an H I. < 0.2 were flagged as low- hemolytic.

[0115] Qiacube RNA extraction procedure

[0116] Plasma samples were thawed on ice, and RNA was extracted with the commercial miRNeasy mini-Kit (QIAgen™, Hilden, Germany) using the QIAcube™ robot, following the manufacturer's instructions.

[0117] TaqMan Low-Density Array™ c-miRs analysis

[0118] Pools plasma aliquots for each subset of the pilot study cohort were created and profiled for c-miRs (Step 2) expression using custom TaqMan Low-Density Array™. RNA was reverse-transcribed using the TaqMan Advanced miRNA cDNA Synthesis Kit (Thermo Fisher Scientific). Poly(A) tailing, adapter ligation, RT reaction, and miR-Amp were performed following the manufacturer’s instructions. qRT-PCR was performed following the manufacturer’s instructions (i.e., 95C for 30s, 45 cycles of 95C for 5s, and 60C for 30s) using a Card Custom Advance (Thermo Fisher Scientific) in a QuantStudio 12k Flex (Thermo Fisher Scientific). Raw Ct values were normalized on six housekeeping c-miRs, as previously described 7).

[0119] OpenArray™ qRT-PCR technology for high-throughput c-miRs analysis

[0120] Total RNA was reverse transcribed using the custom RT Primers and components supplied with the TaqMan MicroRNA Reverse Transcription Kit. The preamplified samples were assembled with TaqMan OpenArray Real-Time PCR Master Mix and loaded to OpenArray plates using the OpenArray Accufill System and compatible accessories (OpenArray 384-well Sample Plates, OpenArray 384- Well Plate Seals, OpenArray AccuFill System Tips, and QuantStudio 12K Flex OpenArray Accessories Kit). The OpenArray plates used in the study contained custom TaqMan OpenArray Human MicroRNA, allowing us to investigate c-miRs expression in each sample. Real-time PCR reactions were performed using the QuantStudio 12 K Flex Real-Time PCR System and the default parameters of the amplification cycles. Data were quantified by relative threshold cycles (Crt) methods and normalized on miR-16-5p as per the manufacturer’s instructions (Thermo Fisher Scientific). To further compensate for variability, we added additional normalizers by selecting the two miRNAs most correlating with miR-16- 5p, from the list of six miRNAs previously described as housekeeping (7). Normalization was performed using the following procedure: a scaling factor (SF) was calculated for each sample, by subtracting the average of the three normalizers to a constant value (K=19.079)(7 . Data were then normalized using the formula: Crt normalized=miRNA Crt raw - SF. Normalization was skipped for Crt=40. Batch bias was investigated through analysis of the miRNAs expression distribution across centers (FC analysis, Wilcoxon Rank Sum test).

[0121] Digital PCR (ddPCR) miRNA profiling

[0122] RNA was reverse-transcribed using the TaqMan Advanced miRNA cDNA Synthesis Kit (Thermo Fisher Scientific). Poly(A) tailing, adapter ligation, RT reaction and miR-Amp were performed following manufacturer’s instructions. MiRNAs expression was also determined using the ddPCR system from Bio-Rad Laboratories. In brief, 1 pl of diluted cDNA (see Results for dilution for each specific miRNA investigated) was combined with a 20 pl PCR reaction mixture that included 10 pl of Bio-Rad’s 2x ddPCR Supermix for probes (#186-3010), 1 pl of TaqMan primer-probe mix (Applied BioSystems), and nuclease-free water. Droplets were formed by dispensing the mixture into a plastic cartridge with 70 pl of QX100 Droplet Generation oil (# 1863005, Bio-Rad Laboratories). The cartridges were then processed in the QX200 Droplet Generator (Bio-Rad Laboratories). The resulting droplets were transferred to a 96-well PCR plate (Thermo Scientific) and underwent PCR amplification on the Cl 000 Touch Thermal Cycler (Bio-Rad Laboratories), using a modified protocol with an annealing / extension temperature of 58°C over 45 cycles. After amplification, the plate was analyzed using the QX200 Droplet Reader (Bio-Rad Laboratories) and the fraction of PCR-positive droplets was calculated using a Poisson distribution. Concentrations were determined using QuantaSoft software and expressed as copies per microliter (copies / pl). Copies / pl counts of the 9 c-miRs were normalized on the copies / pl of miR-16-5p and ath-MIR159a using the following procedure: a scaling factor (SF) was calculated for each sample, by dividing the average of the two normalizers to a constant value (K=185.41). Data were then normalized using the formula: Copies / pl normalized=miRNA Copies / pl raw / SF.

[0123] Signature reduction

[0124] To reduce the size of the signature and select the best set of diagnostic c-miRs, we applied different feature selection methods modeling the odds of lung cancer as a function of miRNAs expression. In detail: A) unconditional logistic regression with stepwise selection (selection parameters: significance^.2 to enter into the model, significance^.25 to stay in the model) was applied to the full screening cohort, including lung cancer and normal controls from both HUM and MUG; batch-effect correction was introduced, including the center as an adjustment covariate in the model; B) penalized unconditional logistic regression was applied to the full screening cohort, with Lasso regularization. Cross-validated (10-fold) loglikelihood with optimization (100 simulations) of the tuning penalty parameter was used to control for potential overfitting. The center was introduced in the model as an unpenalized covariate to correct for batch-effect; C) unconditional logistic regression with stepwise selection (selection parameter: significance^.3 to enter into the model, significance^.35 to stay in the model) was applied to the MUG cohort; D) penalized unconditional logistic regression was applied to the MUG cohort, with Lasso regularization (10-fold cross-validation, on 100 simulations). Feature selection with other techniques (unconditional logistic regression with elastic-net regularization, diagonal linear discriminant analysis) were applied, without reaching better performance (data not shown).

[0125] Statistical analysis

[0126] Patients and tumors characteristics were presented as number and percentage for categorical variables, and as median with first and third quartiles (QI; Q3) for continuous variables; we used Fisher’s exact test and Wilcoxon rank sum test to compare differences in distribution of categorical and continuous variables, respectively. Differentially expressed miRNAs across different datasets (metasignature identification) were identified through the RankProd method, with a nonparametric permutation approach to calculate fold-change and associated proportion of false positive (pfp). Logistic regression was used to model the odds of lung cancer as a function of miRNAs expression. Feature selection approaches based on stepwise or Lasso regularization were used to select panels of miRNAs discriminating between lung cancer and normal controls. Internal validation of models was conducted using bootstrapping techniques (200 bootstraps). The diagnostic performance of predictive models was evaluated by calculating the area under the curve (AUC), accuracy (ACC), sensitivity (SE), specificity (SP), positive predictive value (PPV), and negative predictive value (NPV). Youden index was used to identify the optimal threshold. A value of p less than 0.05 was considered statistically significant. All statistical analyses were performed using SAS software, version 9.4 (SAS Institute, Inc., Cary, NC), and R 3.3.1 (R Core Team, 2016).

[0127] REFERENCES 1. M. B. Schabath, M. L. Cote, Cancer Progress and Priorities: Lung Cancer. Cancer Epidemiol Biomarkers Prev 28, 1563-1579 (2019).

[0128] 2. H. J. De Koning, C. M. Van Der Aalst, P. A. De Jong, E. T. Scholten, K. Nackaerts, M. A. Heuvelmans, J.-W. J. Lammers, C. Weenink, U. Yousaf-Khan, N. Horeweg, S. Van ’T Westeinde, M. Prokop, W. P. Mali, F. A. A. Mohamed Hoesein, P. M. A. Van Ooijen, J. G. J. V. Aerts, M. A. Den Bakker, E. Thunnissen, J. Verschakelen, R. Vliegenthart, J. E. Walter, K. Ten Haaf, H. J. M. Groen, M. Oudkerk, Reduced Lung-Cancer Mortality with Volume CT Screening in a Randomized Trial. N Engl J Med 382, 503-513 (2020).

[0129] 3. D. E. Jonas, D. S. Reuland, S. M. Reddy, M. Nagle, S. D. Clark, R. P. Weber, C. Enyioha, T. L. Malo, A. T. Brenner, C. Armstrong, M. Coker-Schwimmer, J. C. Middleton, C. Voisin, R. P. Harris, Screening for Lung Cancer With Low-Dose Computed Tomography: Updated Evidence Report and Systematic Review for the US Preventive Services Task Force. JAMA 325, 971 (2021).

[0130] 4. C.-H. Marquette, J. Boutros, J. Benzaquen, M. Ferreira, J. Pastre, C. Pison, B. Padovani, F. Bettayeb, V. Fallet, N. Guibert, D. Basille, M. Hie, V. Hofrnan, P. Hofman, AIR project Study Group, Circulating tumour cells as a potential biomarker for lung cancer screening: a prospective cohort study. Lancet Re spir Med 8, 709-716 (2020).

[0131] 5. L. M. Seijo, N. Peled, D. Ajona, M. Boeri, J. K. Field, G. Sozzi, R. Pio, J. J. Zulueta, A. Spira, P. P. Massion, P. J. Mazzone, L. M. Montuenga, Biomarkers in Lung Cancer Screening: Achievements, Promises, and Challenges. Journal of Thoracic Oncology 14, 343-357 (2019).

[0132] 6. E. Dama, T. Colangelo, E. Fina, M. Cremonesi, M. Kallikourdis, G. Veronesi, F. Bianchi, Biomarkers and Lung Cancer Early Detection: State of the Art. Cancers (Basel) 13, 3919 (2021).

[0133] 7. F. Montani, M. J. Marzi, F. Dezi, E. Dama, R. M. Carletti, G. Bonizzi, R. Bertolotti, M. Bellomi, C. Rampinelli, P. Maisonneuve, L. Spaggiari, G. Veronesi, F. Nicassio, P. P. Di Fiore, F. Bianchi, miR-Test: A Blood Test for Lung Cancer Early Detection. JNCI: Journal of the National Cancer Institute 107 (2015), doi: 10.1093 / jnci / djv063.

[0134] 8. J. Zyla, R. Dziadziuszko, M. Marczyk, M. Sitkiewicz, M. Szczepanowska, E. Bottom, G. Veronesi, W. Rzyman, J. Polanska, P. Widlak, miR-122 and miR-21 are Stable Components of miRNA Signatures of Early Lung Cancer after Validation in Three Independent Cohorts. The Journal of Molecular Diagnostics , SI 525157823002453 (2023).

[0135] 9. M. Smolarz, P. Widlak, Serum Exosomes and Their miRNA Load — A Potential Biomarker of Lung Cancer. Cancers 13, 1373 (2021).

[0136] 10. M. Boeri, C. Verri, D. Conte, L. Roz, P. Modena, F. Facchinetti, E. Calabro, C. M. Croce, U. Pastorino, G. Sozzi, MicroRNA signatures in tissues and plasma predict development and prognosis of computed tomography detected lung cancer. Proc. Natl. Acad. Sci. U.S.A. 108, 3713-3718 (2011).

[0137] 11. G. Sozzi, M. Boeri, M. Rossi, C. Verri, P. Suatoni, F. Bravi, L. Roz, D. Conte, M. Grassi, N. Sverzellati, A. Marchiano, E. Negri, C. La Vecchia, U. Pastorino, Clinical Utility of a Plasma-Based miRNA Signature Classifier Within Computed Tomography Lung Cancer Screening: A Correlative MILD Trial Study. JCO 32, 768-773 (2014).

[0138] 12. F. Bianchi, F. Nicassio, M. Marzi, E. Belloni, V. Dall’Olio, L. Bernard, G. Pelosi, P. Maisonneuve, G. Veronesi, P. P. Di Fiore, A serum circulating miRNA diagnostic test to identify asymptomatic high-risk individuals with early stage lung cancer. EMBO Mol Med 3, 495-503 (2011).

[0139] 13. M. Power, G. Fell, M. Wright, Principles for high-quality, high-value testing. Evid Based Med 18, 5-10 (2013).

[0140] 14. M. B. Wozniak, G. Scelo, D. C. Muller, A. Mukeria, D. Zaridze, P. Brennan, J. D. Hoheisel, Ed. Circulating MicroRNAs as Non-Invasive Biomarkers for Early Detection of Non-Small-Cell Lung Cancer. PLoS ONE 10, e0125026 (2015).

[0141] 15. G. P. Vadla, B. Daghat, N. Patterson, V. Ahmad, G. Perez, A. Garcia, Y. Manjunath, J. T. Kaifi, G. Li, C. Y. Chabu, Combining plasma extracellular vesicle Let-7b-5p, miR-184 and circulating miR-22-3p levels for NSCLC diagnosis and drug resistance prediction. Sci Rep 12, 6693 (2022).

[0142] 16. S. Lu, H. Kong, Y. Hou, D. Ge, W. Huang, J. Ou, D. Yang, L. Zhang, G. Wu, Y. Song, X. Zhang, C. Zhai, Q. Wang, H. Zhu, Y. Wu, C. Bai, Two plasma microRNA panels for diagnosis and subtype discrimination of lung cancer. Lung Cancer 123, 44-51 (2018).

[0143] 17. F. Hong, R. Breitling, C. W. McEntee, B. S. Wittner, J. L. Nemhauser, J. Chory, RankProd: a bioconductor package for detecting differentially expressed genes in meta-analysis. Bioinformatics 22, 2825-2827 (2006).

[0144] 18. M. J. Marzi, F. Montani, R. M. Carletti, F. Dezi, E. Dama, G. Bonizzi, M. T. Sandri, C. Rampinelli, M. Bellomi, P. Maisonneuve, L. Spaggiari, G. Veronesi, F. Bianchi, P. P. Di Fiore, F. Nicassio, Optimization and Standardization of Circulating MicroRNA Detection for Clinical Application: The miR-Test Case. Clinical Chemistry 62, 743-754 (2016).

[0145] 19. V. Appierto, M. Callari, E. Cavadini, D. Morelli, M. G. Daidone, P. Tiberio, A Lipemia-Independent Nanodrop ® -Based Score to Identify Hemolysis in Plasma and Serum Samples. Bioanalysis 6, 1215-1226 (2014).

[0146] 20. M. Ostrowski, T. Marjanski, R. Dziedzic, M. Jelitto-Gorska, K. Dziadziuszko, E. Szurowska, R. Dziadziuszko, W. Rzyman, Ten years of experience in lung cancer screening in Gdansk, Poland: a comparative study of the evaluation and surgical treatment of 14 200 participants of 2 lung cancer screening programmes. Interactive Cardiovascular and Thoracic Surgery 29, 266-274 (2019).

[0147] 21. A. Antonicelli, P. Muriana, G. Favaro, G. Mangiameli, E. Lanza, M. Profili, F. Bianchi, E. Fina, G. Ferrante, S. Ghislandi, D. Pistillo, G. Finocchiaro, G.

[0148] Condorelli, R. Lembo, P. Novellis, E. Dieci, S. De Santis, G. Veronesi, The Smokers Health Multiple Actions (SMAC-1) Trial: Study Design and Results of the Baseline Round. Cancers 16, 417 (2024).

Claims

Claims1. A method in-vitro or ex-vivo for identifying subjects affected by lung cancer, comprising the steps of: a) detecting the amount of each of the 9 miRNAs having sequence hsa-miR- 29a-3p (SEQ ID NO. 1), hsa-miR-30c-5p (SEQ ID NO. 2), hsa-miR-184 (SEQ ID NO. 3), hsa-miR-190b-5p (SEQ ID NO. 4), hsa-miR-200c-3p (SEQ ID NO. 5), hsa-miR-218-5p (SEQ ID NO. 6), hsa-miR-328-3p (SEQ ID NO. 7), hsa- miR-450b-5p (SEQ ID NO. 8) and hsa-miR-1233-3p (SEQ ID NO. 9) in a biological sample from a subject; b) normalizing the data obtained in step a).

2. Method according to claim 1 further comprising step c) of calculating a risk score.

3. Method according to claim 1, wherein when the miRNs are detected by standard qPCR, said risk score is calculated according to the following formula: RS = -2.7465+hsa-miR-1233-3p*0.0622+hsa-miR-184*0.1151+hsa-miR-190b- 5p*(-0.0335)+hsa-miR-200c-3p*0.2769+hsa-miR-218-5p*(-0.0265)+hsa-miR- 29a-3p*(-0.3578)+hsa-miR-30c-5p*1.5194+hsa-miR-328-3p*(-0.9112)+hsa-miR- 450b-5p*0.07094. Method according to claim 1, wherein when the miRNAs are detected by digital PCR said risk score is calculated according to the following formula: RS=-0.1725+hsa-miR-1233-3p*(-0.00128)+hsa-miR-184*(-0.00001)+hsa-miR- 190b-5p*(-0.00008)+hsa-miR-200c-3p*0.000032+hsa-miR-218-5p* (-0.00152)+hsa-miR-29a-3p*0.000002745+hsa-miR-30c-5p*0.00000231+hsa- miR-328-3p*(-0.00002)+hsa-miR-450b-5p*(-0.00012)5. Method according to claim 1 further comprising step d) wherein the subject is classified as subjects at high risk, at intermediate risk or at low risk of developinglung cancer based on the value of the tumor probability.

6. Method according to claim 5, wherein said tumor probability is calculated using the following formula:PROB_TUM=100*EXP(RS) / (1+(EXP(RS)).

7. Method according to claim 5 or 6, wherein when the tumor probability is greater than or equal to 5% but lower than or equal to 60% it indicates that the subject has an intermediate risk of developing lung cancer, when the risk score is more than 60% the subject has an high risk of developing lung cancer and where the risk score is less than 5% the subject has a low risk of developing lung cancer.

8. The method according to any of the previous claims, wherein the biological sample is a biological fluid or a tissue sample.

9. Method according to claim 8, wherein the biological fluid is whole blood, serum, plasma, saliva, urine or lymph fluid, preferably said biological fluid is serum or plasma.

10. Method according to claim 8, wherein the tissue sample is a fresh tissue sample or a frozen tissue sample.

11. The method according to any of the previous claims, wherein the miRNAs in step a) is detected by quantitative PCR (qPCR), digital PCR (ddPCR), RNA sequencing, Affymetrix microarray or custom microarray, or digital detection through molecular barcoding selected from NanoString technology, preferably qPCR or digital PCR.

12. The method according to any of the previous claims, wherein when the miRNAs are detected with qPCR the normalization of the data reported in step b) is made by using three miRNAs normalizer selected from miR-16-5p (SEQ ID NO: 10), miR-19a-3p (SEQ ID NO: 11) and miR-19b-3p (SEQ ID NO: 12).

13. Method according to any of the previous claims, wherein the subject is anasymptomatic subject, a high risk individual or a patient affected by lung cancer.

14. Method according to any of the previous claims, wherein the subject affected by lung cancer is at an early-stage lung cancer, at a stage I lung cancer, at a stage II lung cancer or with locally-advanced lung cancer Stage III.

15. Method according to any of the previous claims, characterized in that it is a first-line screening procedure to pre-select patients or subjects who require further diagnostic investigation by LDCT.

16. Method according to any of the previous claims, characterized in that it is a second-line screening procedure to distinguish between benign and malign nodules in a subject.

17. Method according to any of the previous claims, characterized in that it has an accuracy between 65% to 80%, preferably about 70%.

18. Method according to any of previous claims characterized in that it has a sensitivity between 70% to 90%, preferably between 76% to 82%19. Method according to any of the previous claims, characterized in that it has specificity between 60% and 80%, preferably about 67%.

20. Method according to any of the previous claims, characterized in that it identifies patients included in the group of aggressive lung cancer with a poor prognosis, selected from patients with shorter overall survival and and / or patients with shorter disease-free survival, and / or patients responsive to a treatment, and / or with patients with metastatic disease which can include patients with early-stage disease (stage I).

21. Method according to any of the previous claims, characterized in that it identifies patients included in the group of non aggressive lung cancer with a goodprognosis, selected from patients with longer overall survival, and / or patients with longer disease-free survival, and / or patients responsive to treatment, patients and / orwithout metastatic disease which can include patients with early-stage disease (stage I).

22. Method according to any of the previous claims, characterized in that it is a first-line screening procedure to pre-select subjects at risk of developing lung cancer.