Method for identifying if a subject is at risk of developing lung cancer

A noninvasive method for identifying lung cancer risk involves analyzing nasal samples for cfDNA fragments from specific genes and detecting genetic variants, offering a reliable and convenient screening solution for early lung cancer detection.

WO2025109033A1PCT designated stage expired Publication Date: 2025-05-30ONCOSWAB GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/083042
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-21
Filing Date
2024-11-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

There is a need for a noninvasive, reliable, and convenient method for lung cancer screening that can identify individuals at risk early, as early lung cancer often has no symptoms and current methods are invasive or require extensive tissue sampling.

Method used

The method involves detecting cell-free DNA (cfDNA) fragments from target genes such as ALK, EGFR, MAP2K1, PIK3CA, and TP53 in nasal samples, and analyzing for genetic variants using PCR and next-generation sequencing (NGS) techniques.

Benefits of technology

This method is effective in identifying lung cancer risk with a high sensitivity and specificity, even in the absence of direct tumor connection, and can be performed with minimal nasal fluid samples, making it noninvasive and cost-effective.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000020_0001
    Figure IMGF000020_0001
  • Figure IMGF000021_0001
    Figure IMGF000021_0001
  • Figure IMGF000021_0002
    Figure IMGF000021_0002
Patent Text Reader

Abstract

The present invention relates to a method for identifying if a subject is at risk of developing lung cancer, the method comprising detecting in at least one nasal sample from said subject at least one or more cell free DNA (cfDNA) being fragments of target genes selected from the group comprising ALK, EGFR, MAP2K1, PIK3CA and TP53; and detecting the presence or absence of genetic variants of the said at least one or more cfDNA.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD FOR IDENTIFYING IF A SUBJECT IS AT RISK OF DEVELOPING

[0002] LUNG CANCER

[0003] TECHNICAL FIELD

[0004] The present invention relates to a method for identifying if a subject is at risk of developing lung cancer.

[0005] BACKGROUND OF THE INVENTION

[0006] Lung cancer is by far the leading cause of cancer death in the US, accounting for about 1 in 5 of all cancer deaths. There are 2 main types of lung cancer including non-small cell lung cancer (NSCLC) and small cell lung cancer (SCLC). Most people diagnosed with lung cancer are 65 or older. Early lung cancer often has no symptoms and can only be detected by medical imaging such as low dose computed tomography (LDCT).

[0007] In addition to diagnosis by CT, in vitro diagnostic tests have been developed based on tumor tissue samples and next generation sequencing (NGS) to detect biomarkers associated with lung cancer allowing stratification of patients who may benefit from treatment with targeted therapies. In these tests, DNA can be extracted from tumor tissue sample, or cell-free DNA (cfDNA) can be extracted from plasma sample for which real time PCR can be performed to detect variants of biomarkers.

[0008] Cancer patients usually have a high level of cfDNA in their serum or plasma because of cellular necrosis or apoptosis. The fraction of cfDNA that derived from tumor cells is named circulating tumor DNA (ctDNA) and molecular profiling of ctDNA has been suggested as a prognostic tool (Yan-yan Yan, Front. Cell Dev. BioL, 22 February 2021 Volume 9 - 2021 ).

[0009] Multiple biomarkers have been detected in lung cancer including genes coding for the expression of epithelial growth factor receptor (EGFR), anaplastic lymphoma kinase (ALK), B-Raf proto-oncogene (BRAF), mesenchymal epithelial transition (MET) and Besides, in lung cancer, it has been shown that biomarkers detected through nasal swabs can be effective for detecting aberrant methylation of certain genes, such as CDH133 and Septin 94 in nasopharyngeal carcinoma (US2011 / 0217717).

[0010] As early lung cancer often has no symptoms, there is an urgent need for a noninvasive procedure that provides easier and more convenient access to lung cancer screening to increase patient’s adherence, and that has a reliable positive predictive value.

[0011] SUMMARY OF THE INVENTION

[0012] The present invention provides a method for identifying if a subject is at risk of developing lung cancer, the method comprising detecting in at least one nasal sample from said subject at least one or more cell free DNA (cfDNA) being fragments of target genes selected from the group comprising ALK, EGFR, MAP2K1 , PIK3CA and TP53; and detecting the presence or absence of genetic variants of the said at least one or more cfDNA.

[0013] The invention also provides a computer program containing instructions and / or a computer-readable media containing means for carrying out said method.

[0014] The invention also provides a kit directed to said method for identifying if a subject is at risk of developing lung cancer, said kit comprising: i) probes for detecting genetic variants of cfDNA being fragments of target genes selected from the group comprising ALK, EGFR, MAP2K1 , PIK3CA, TP53, KRAS, MET, NRAS, CDH13, CHD1 , DAPK, RASSF1 , SEPT9, BRAF, ERBB2, and / or ROS1 ; and optionally ii) reagents and instructions for use.

[0015] BRIEF DESCRIPTION OF THE FIGURES

[0016] Figure 1 illustrates the method steps of the invention, wherein the sample is a nasal sample.

[0017] Figure 2 shows the amount of total nucleic acids ng / pl (number in bold) and the amount of cfDNA, expressed in % of total nucleic acids, as a fraction of total nucleic acids, in nasal samples obtained from 12 lung cancer patients. Figure 3 shows the heatmap of target genes from nasal samples of 7 lung cancer patients in comparison to control samples from 3 healthy patients.

[0018] Figure 4 shows the results of a ddPCR carried out with nasal samples of lung cancer patients (patient) and healthy patients (control). In A: the total number of droplets. In B: the number of positive droplets. In C: the variant allele fraction VAF.

[0019] Figure 5 shows the microbiome in the nasal of lung cancer patients in A: the microorganisms and in B: the bacteria species found in nasal samples from lung cancer patients (n=20) compared to control samples (n=50).

[0020] Figure 6 shows data obtained by Fast Aneuploidy Screening Test-Sequencing System (Fast-Seq) using nasal samples from lung cancer patients compared to control samples.

[0021] DETAILED DESCRIPTION OF THE INVENTION

[0022] The present invention relates to a method for identifying if a subject is at risk of developing lung cancer, the method comprising detecting in at least one sample from said subject at least one or more cell free DNA (cfDNA) being fragments of target genes selected from the group comprising CDH13, CHD1 , DAPK, RASSF1 , SEPT9, ALK, BRAF, EGFR, ERBB2, KRAS, MAP2K1 , MET, NRAS, PIK3CA, ROS1 and / or TP53.

[0023] As used herewith, a subject at risk of developing lung cancer is a subject at risk of developing and / or having a lung cancer. Early-stage lung cancer may develop or progress further over time. Lung cancer can be detected at all stages of development, at early stage or at other stages of development by the method of the invention. Thus, the invention relates also to a method for identifying if a subject is at risk of having lung cancer.

[0024] Circulating free nucleic acids (cfNA) such as circulating free DNA (cfDNA) or RNA (cfRNA) are present in biological fluids independent of the cells.

[0025] As used herewith, cell free DNA (cfDNA) are circulating free DNA found in different body fluids, both in healthy and not healthy subjects. cfDNA is released from cells into the circulatory system throughout the body and can be found for example in plasma, cerebral spinal fluid (CSF), pleural fluid, urine, nasal fluid, and saliva. The fraction of cfDNA that may derived from tumor cells is named circulating tumor DNA (ctDNA).

[0026] In the context of the present invention, cfDNA may also include all types of cfNA such as cfRNA.

[0027] As shown in the example 6, cfDNA have been found in nasal samples from lung cancer patients at an average of 20 ng / pL of cfDNA from 1 mL of nasal fluids (Figure 2). Advantageously, the amount of cfDNA required in the method of the invention is very small, and the method can be performed using at least one nasal sample from a subject using a minimal amount of nasal fluid such as 0.5 mL.

[0028] As used herewith, target genes are biomarkers that have been found in body fluids of patients with lung cancer such as those described in tables 4-6. Preferably, the target genes are selected from the group comprising CDH13, CHD1 , DAPK, RASSF1 , SEPT9, ALK, BRAF, EGFR, ERBB2, KRAS, MAP2K1 , MET, NRAS, PIK3CA, ROS1 and / or TP53.

[0029] In one embodiment, the present invention relates to a method for identifying if a subject is at risk of developing lung cancer, the method comprising detecting in at least one nasal sample from said subject at least one or more cell free DNA (cfDNA) being fragments of target genes selected from the group comprising ALK, EGFR, MAP2K1 , PIK3CA and TP53; and detecting the presence or absence of genetic variants of the said at least one or more cfDNA.

[0030] In the present invention, the method comprises detecting in at least one nasal sample from said subject at least one or more cell free DNA (cfDNA) that are fragments of target genes comprising ALK, EGFR, MAP2K1 , PIK3CA and TP53.

[0031] Thus, the method may further comprise the analysis of the target genes comprising ALK, EGFR, MAP2K1 , PIK3CA and TP53. For example, the analysis can be done following the results of detection of the presence of cfDNA in the at least one nasal sample from a subject. Preferably, the method comprises the analysis of the target genes comprising ALK, EGFR, MAP2K1 , PIK3CA, TP53, KRAS, MET, and NRAS. More preferably, the method comprises the analysis of the target genes comprising ALK, EGFR, MAP2K1 , PIK3CA, TP53, KRAS, MET, NRAS, CDH13, CHD1 , DAPK, RASSF1 , SEPT9, BRAF, ERBB2, and ROS1 . The analysis can be carried out for example by PCR analysis or NGS analysis.

[0032] As shown in the example, the analysis of the target genes ALK, EGFR, MAP2K1 , PIK3CA and TP53 has been made using at least one nasal sample per subject. The method is very reliable as at least one or more cfDNA / ctDNA being fragments of the target genes ALK, EGFR, MAP2K1 , PIK3CA and TP53 were found in 6 out of 7 lung cancer patients using nasal samples, whereas no ctDNA were found in healthy patients (Figure 3 and example 8). As shown in this example, by testing a minimum set of target genes, the efficacy of the method is at least 85%. This result is unexpected as the nasal cavity lacks a direct connection to tumor lung tissue. The nasal environment is also subject to DNA degradation from mucosal enzymes, bacteria, and contaminants diluting any possible trace of ctDNA. These data demonstrate that the method of the invention is effective for detecting tumor genetic variants ctDNA of target genes directly in at least one nasal sample from a lung cancer subject.

[0033] Advantageously, the method of the present invention through the detection of cfDNA / ctDNA can be performed using samples obtained from subjects through a non- invasive procedure using a swab. The collection of the sample may be done by rotating a swab on the nasal mucosa, or by rubbing the nasopharyngeal cavity and / or oropharynx. Preferably, the collection of nasal samples may be done by gently and slowly rotating a swab on the nasal mucosa of the anterior nare zone. Thus, the nasal sample collection is easy, safety and cost-effective.

[0034] Preferably, the target genes are selected from the group comprising ALK, EGFR, MAP2K1 , PIK3CA, TP53, KRAS, MET, and NRAS.

[0035] As shown in the example 7, ctDNA have been found in nasal samples from lung cancer patients. A large number of lung cancer mutations (68 mutations) were found in cfDNA / ctDNA being fragments of target genes selected from ALK, EGFR, KRAS, MAP2K1 , MET, NRAS, PIK3CA and TP53 (Table 8).

[0036] More preferably, the target genes are selected from the group comprising ALK, EGFR, MAP2K1 , PIK3CA, TP53, KRAS, MET, NRAS, CDH13, CHD1 , DAPK, RASSF1 , SEPT9, BRAF, ERBB2, and ROS1.

[0037] Even more preferably, target genes are selected from the group consisting of CDH13, CHD1 , DAPK, RASSF1 , SEPT9, ALK, BRAF, EGFR, ERBB2, KRAS, MAP2K1 , MET, NRAS, PIK3CA, ROS1 and / or TP53. Again even more preferably, target genes are selected from the group consisting of CDH13, CHD1 , DAPK, RASSF1 , SEPT9, ALK, BRAF, EGFR, ERBB2, KRAS, MAP2K1 , MET, NRAS, PIK3CA, ROS1 and TP53. The method of the invention may further comprises analyzing the presence or absence of genetic variants of the at least one or more cfDNA when cfDNA being fragments of target genes selected from the group comprising CDH13, CHD1 , DAPK, RASSF1 , SEPT9, ALK, BRAF, EGFR, ERBB2, KRAS, MAP2K1 , MET, NRAS, PIK3CA, ROS1 and / or TP53 have been detected in a sample obtained from the subject. Preferably, the method of the invention may further comprises analyzing the presence or absence of genetic variants of the at least one or more cfDNA when cfDNA being fragments of target genes selected from the group comprising ALK, EGFR, MAP2K1 , PIK3CA and TP53 have been detected in a sample obtained from the subject.

[0038] Preferably, the method comprises the detection of least one or more cfDNA being fragments of target genes selected from the group comprising CDH13, CHD1 , DAPK, RASSF1 , SEPT9, ALK, BRAF, EGFR, ERBB2, KRAS, MAP2K1 , MET, NRAS, PIK3CA, ROS1 and / or TP53 in at least one sample selected from a nasal sample, and / or a salivary sample. More preferably, the sample is a nasal sample.

[0039] In the present invention, combination of samples obtained from said subject provides more cfDNA / ctDNA, thereby increasing the sensitivity and specificity of the method. For example, a nasal sample can be combined with any other type of samples from body fluids such as saliva, plasma, cerebral spinal fluid (CSF), pleural fluid, urine, sputum, stool, and / or seminal fluid.

[0040] In one embodiment, the method may comprise the step of combining at least two samples independently selected from nasal samples and / or salivary samples obtained from the subject. Preferably, the method comprises the step of combining at least two, three, four, five, six, seven, eight, nine, ten or more samples from nasal samples and / or a salivary sample. More preferably, the method comprises the step of combining at least two, three, four, five, six, seven, eight, nine, ten or more samples from nasal samples.

[0041] In another embodiment, the method may comprise the step of combining at least two samples obtained from the subject at different time points during at least one day and detecting in said combined sample at least one or more cfDNA being fragments of target genes selected from the group comprising CDH13, CHD1 , DAPK, RASSF1 , SEPT9, ALK, BRAF, EGFR, ERBB2, KRAS, MAP2K1 , MET, NRAS, PIK3CA, R0S1 and / or TP53. Preferably, the method comprises the step of combining at least two samples obtained from the subject at different time points during one, two, three, four, five, six, seven, eight, nine, ten or more weeks.

[0042] For example, one nasal sample can be combined with at least one further nasal sample obtained from the subject at different time points during at least one week.

[0043] Thus, the method can be used to monitor over time the presence of cfDNA / ctDNA in at least one nasal sample from a subject at different time points following surgery of said subject.

[0044] In another embodiment, the method may comprise detecting in at least two nasal samples obtained from the subject at different time points at least one or more cfDNA being fragments of target genes selected from the group comprising CDH13, CHD1 , DAPK, RASSF1 , SEPT9, ALK, BRAF, EGFR, ERBB2, KRAS, MAP2K1 , MET, NRAS, PIK3CA, ROS1 and / or TP53. Preferably, the target genes are selected from the group comprising ALK, EGFR, MAP2K1 , PIK3CA and TP53. Thus, the method comprises the step of monitoring over time the presence of cfDNA / ctDNA in at least two nasal samples from said subject at different time points during one, two, three, four, five, six, seven, eight, nine, ten or more weeks.

[0045] In one embodiment of the method of the present invention, the detection of the at least one or more cfDNA may be determined by polymerase chain reaction (PCR) analysis.

[0046] PCR analysis can be performed by digital droplet PCR (ddPCR), and / or multiplex PCR.

[0047] As used herein, multiplex PCR analysis refers to the amplification of several cfDNA sequences simultaneously using multiple primers in one reaction, in which, preferably the multiplex PCR is carried out using primers for amplification only of the genetic variants of cfDNA (i.e. the fraction of cfDNA being ctDNA) comprising targeted mutations. ddPCR analysis can be used to identify and quantify a targeted mutation (a hotspot) in the target genes such as those determined in the present invention (table 8 or table 1 ) with very high sensitivity.

[0048] Preferably, the PCR analysis is a digital droplet PCR (ddPCR) analysis.

[0049] Thus, the PCR analysis comprises using probes hybridizing the genetic variants of the at least one or more cfDNA being fragments of the target genes ALK, EGFR, MAP2K1 , PIK3CA, TP53, KRAS, MET, NRAS, CDH13, CHD1 , DAPK, RASSF1 , SEPT9, BRAF, ERBB2, and / or ROS1.

[0050] For example, the genetic variants of the said at least one or more cfDNA are selected from the genetic variants described in table 1 and / or table 8.

[0051] Following positive result by PCR such as ddPCR, i.e. amplification of the genetic variants of the at least one or more cfDNA (i.e. the fraction of cfDNA being ctDNA) being fragments of target genes selected from the group comprising CDH13, CHD1 , DAPK, RASSF1 , SEPT9, ALK, BRAF, EGFR, ERBB2, KRAS, MAP2K1 , MET, NRAS, PIK3CA, ROS1 and / or TP53, the method of the invention further comprises analyzing the presence of genetic variants of the at least one or more cfDNA, for example by nucleic acid sequencing analysis.

[0052] As used herewith, nucleic acid sequencing analysis includes methods for determining the sequence of nucleic acids selected from DNA and / or RNA.

[0053] Advantageously, in the method of the invention, the polymerase chain reaction (PCR) analysis and nucleic acid sequencing analysis can be carried out sequentially.

[0054] Thus, in the case of amplification by PCR such as ddPCR, the method further comprises a nucleic acid sequencing analysis, preferably the nucleic acid sequencing analysis is next-generation sequencing analysis (NGS).

[0055] As used herein, genetic variants of cfDNA include substitutions, insertions, deletions, gene rearrangements, exon skipping, fusions and copy number alterations in cfDNA and / or cfRNA. For example, genetic variants of cfDNA are ctDNA and the presence of at least one or more ctDNA in a sample from a subject indicates that said subject has an increased risk of having and / or developing a cancer.

[0056] The method of the invention may be carried out in combination with low dose computed tomography (LDCT), prior or after LDCT. For example, a low dose computed tomography (LDCT) may be carried out to confirm the prognostic made by the method of the invention and establish a diagnosis.

[0057] In the absence of amplification by PCR such as ddPCR, a nucleic acid sequencing analysis is not necessary allowing a simple and low costs method.

[0058] In the method of the invention, the presence of genetic variants of at least one or more cfDNA being fragments of target genes independently selected from the group comprising CDH13, CHD1 , DAPK, RASSF1 , SEPT9, ALK, BRAF, EGFR, ERBB2, KRAS, MAP2K1 , MET, NRAS, PIK3CA, ROS1 and / or TP53, indicates that said subject has an increased risk of developing or having lung cancer.

[0059] Thus, the presence of genetic variants of one, two, three, four, five, six, seven, eight, nine, ten or more cfDNA being fragments of target genes independently selected from the group comprising CDH13, CHD1 , DAPK, RASSF1 , SEPT9, ALK, BRAF, EGFR, ERBB2, KRAS, MAP2K1 , MET, NRAS, PIK3CA, ROS1 and / or TP53, indicates that said subject has an increased risk of developing lung cancer.

[0060] Preferably, in the present invention, the presence of variants of at least one, two, three, four, five, six, seven, eight, nine, ten or more cfDNA being fragments of target genes independently selected from the group consisting of CDH13, CHD1 , DAPK, RASSF1 , SEPT9, ALK, BRAF, EGFR, ERBB2, KRAS, MAP2K1 , MET, NRAS, PIK3CA, ROS1 and TP53 indicates that said subject has an increased risk of developing a lung cancer.

[0061] As shown in example 8, each lung cancer patient has a specific profile (figure 3). The analysis of nasal samples from said patients shows the presence of at least one, two, or three cfDNA being fragments of target genes independently selected from the group comprising ALK, BRAF, EGFR, ERBB2, KRAS, MAP2K1 , MET, NRAS, PIK3CA, ROS1 and TP53 and genetic variants of these cfDNA (ctDNA) have been found in said nasal samples.

[0062] In another embodiment, the method may further comprise the step of detecting the presence or absence of 16S rRNA genes in said sample, preferably, in a nasal sample. For example, the method may further comprise the step of detecting and / or analyzing the presence or absence of genetic variants of 16S rRNA genes.

[0063] Most bacterial species contain more than one ribosomal RNA operon copy in their genomes, with some species containing up to 15 such copies.

[0064] As used herewith, the presence of genetic variants of 16S rRNA (rrs) genes indicates the presence of microorganisms in a sample. As shown in the example 10, the analysis of the nasal microbiome profile of patients with lung cancer in comparison with the microbiome profile of healthy subjects show the presence of multiple bacteria selected from the phylum Actinobacteriota, Bacteroidota, Firmicutes, and / or Proteobacteria. In this example, it has been shown that the relative abundance of some bacteria such as Corynebacterium, Dolosigranulum or Moraxella was found higher than in samples from healthy subject (figure 5A and 5B).

[0065] Thus, in the method of the invention, the presence of genetic variants of the at least one or more cfDNA and / or 16S rRNA genes indicates that said subject has an increased risk of developing and / or having lung cancer.

[0066] In another embodiment, the invention relates to a method for identifying if a subject is at risk of developing lung cancer, the method comprising detecting in at least one sample from said subject at least one or more cell free DNA (cfDNA) being fragments of target genes selected from the group comprising CDH13, CHD1 , DAPK, RASSF1 , SEPT9, ALK, BRAF, EGFR, ERBB2, KRAS, MAP2K1 , MET, NRAS, PIK3CA, ROS1 and / or TP53, wherein the detection of at least one or more cfDNA may be determined by fast aneuploidy screening test sequencing system (Fast-seq).

[0067] Fast-seq may be used to analyze somatic copy number alterations (SCNA) which are present in cfDNA of lung cancer subjects and estimate ctDNA level. Following positive detection by Fast-seq of cfDNA / ctDNA, samples are analyzed subsequently by carrying out a nucleic acid sequencing analysis to detect variants in said cfDNA.

[0068] The method may further comprise the step of detecting aneuploidy levels in circulating free DNA (cfDNA) through genome sequencing, either before, during, or after treatment. Thus, in the method of the invention, the detection of tumor-derived aneuploidy in the at least one or more cfDNA / ctDNA indicates that said subject has an increased risk of developing lung cancer.

[0069] Preferably, the invention relates to a method for identifying if a subject is at risk of developing lung cancer, the method comprising detecting in at least one sample from said subject at least one or more cell free DNA (cfDNA) being fragments of target genes selected from the group comprising CDH13, CHD1 , DAPK, RASSF1 , SEPT9, ALK, BRAF, EGFR, ERBB2, KRAS, MAP2K1 , MET, NRAS, PIK3CA, ROS1 and / or TP53, wherein the detection of at least one or more cfDNA is determined by PCR analysis and / or fast aneuploidy screening test sequencing system (Fast-seq).

[0070] In another embodiment, the method of the invention comprises the step of i) detecting the presence of genetic variants of the at least one or more cfDNA; ii) detecting at least one or more 16S rRNA gene; and / or iii) detecting tumor-derived aneuploidy in the at least one or more cfDNA. The presence of genetic variants of the at least one or more cfDNA, of one or more 16S rRNA gene, and / or tumor-derived aneuploidy in a sample, preferably a nasal sample, from a subject indicates that said subject has an increased risk of developing lung cancer (figure 1 ).

[0071] The method may also comprise a step of detecting in at least one sample from said subject at least one or more cfRNA of target genes selected from the group comprising ALK, ROS1 , RET, NTRK, and / or MET.

[0072] The present invention also relates to a method for identifying if a subject is at risk of developing lung cancer disease, the method comprising i) detecting in at least one sample from said subject at least one or more cell free DNA (cfDNA) being fragments of target genes independently selected from the group comprising CDH13, CHD1 , DAPK, RASSF1 , SEPT9, ALK, BRAF, EGFR, ERBB2, KRAS, MAP2K1 , MET, NRAS, PIK3CA, ROS1 and / or TP53; and ii) determining clinical risk factors of said subject.

[0073] In case of positive result in step i), the method further comprises analyzing the presence or absence of genetic variants of the at least one or more cfDNA.

[0074] Clinical risk factors are for example selected from the group comprising age (years), gender, race or ethnic group, education, bmi (kg / m2), copd, emphysema, personal history of cancer, family history of lung cancer, personal history of pneumonia, smoking status (current smoker, former smoker), smoking duration (years), smoking intensity (cigarettes per day), pack-years of smoking, years since cessation, breathing upon exertion, cough intensity, changes in voice, flu or pneumonia within the past 2 year, exposure to asbestos or other chemicals, current residence, physical activity, exposures (coronary heart disease, myocardial infarction, stroke, diabetes mellitus, chronic kidney disease, arthritis, asthma), personal history of cancer.

[0075] Preferably, the method further comprises the determination of clinical risk factors of said subject, said clinical risk factors are selected from the group comprising smoking status, age, gender, BMI, personal history of cancer, family history of lung cancer, personal history of pneumonia, and / or physical activity.

[0076] Advantageously, the method of the invention provides a comprehensive lung cancer management including in addition to the detection of the presence or absence of genetic variants of the at least one or more cfDNA, a risk assessment of high-risk population based on clinical risk factors.

[0077] In the present invention, lung cancer may be selected from non-small cell lung cancer (NSCLC), small cell lung cancer (SCLC), bronchogenic carcinoma, alveolar carcinoma, bronchial adenoma, sarcoma, lymphoma, chondromatous hamartoma, incidental pulmonary nodules related lung cancer and / or mesothelioma. Furthermore, NSCLC may be adenocarcinoma, squamous cell carcinoma, adenosquamous carcinoma, large cell carcinoma and sarcomatoid. The method of the invention can be used for example, to improve the management of incidental pulmonary nodules related lung cancer by reducing the high rate of false positives, unnecessary biopsies and long follow-up periods associated with the current standard of care.

[0078] The method of the invention may further comprise the step of treating the subject with at least one anticancer agent or a combination thereof based on the detected presence of genetic variants of the at least one or more cfDNA.

[0079] Advantageously, the method of the invention makes it possible to identify mutations relevant to a specific treatment or therapy, which improves the precision and effectiveness of therapeutic options. In particular, the treatment is tailored to a lung cancer patient in correlation with the type of lung cancer (NSCLC, SCLC, bronchogenic carcinoma, alveolar carcinoma, bronchial adenoma, sarcoma, lymphoma, chondromatous hamartoma, and mesothelioma), and genetic variants of target genes / cfDNA detected in that patient.

[0080] Anticancer agents are for example Gefitinib, Erlotinib, Afatinib, Osimertinib, Amivantamab, or Mobocertinib for patients having EGFR mutations; Sotorasib, or Adagrasib for patients having KRAS mutations; Crizotinib, Capmatinib, Tepotinib, or Glumetinib for patients having MET mutations; Vemurafenib, or Dabrafenib for patients having BRAF V600E mutations; Crizotinib, or Lorlatinib for patients having ALK rearrangements; Crizotinib, Repotrectinib, or Entrectinib for patients having ROS1 rearrangements; Trastuzumab or deruxtecan for patient having HER2 (ERBB2) mutations; Selpercatinib or Pralsetinib for patients having RET rearrangements; Larotrectinib, Entrectinib, or Repotrectinib for patients having NTRK gene fusions.

[0081] Further, patients under treatment with a known mutation in the target genes of the invention can be monitored with the method of the invention using nasal fluid samples, for example by carrying out a PCR reaction such as ddPCR to follow the patient over time for their response to treatment. Thus, the method of the invention provides a comprehensive lung cancer management of lung cancer patients at all stages of the disease to follow the efficacy of a treatment and remission. Thus, the method may further comprise the step of monitoring over time the response of the subject to the treatment with at least one anticancer agent.

[0082] The method of the invention also provides the step of monitoring the drug resistance of the subject over time, allowing for adjustment to treatment.

[0083] The method may further comprises using a computer program containing instructions and / or a computer-readable media having stored thereon the computer program containing means for carrying out the method of the invention.

[0084] The computer program model may include risk values correlated with cfDNA detection results including sequencing data, aneuploidy data, with microbiome analysis data, with clinical risk factors, with CT scans data and / or with data from any other imaging techniques.

[0085] Such a radiogenomic assessment guarantees a highly effective method for determining the relative risk (RR) of a subject of having a lung cancer.

[0086] The relative risk (RR) in this model may be calculated by the formula:

[0087] RR = exp(Pi * Smoking Status + j32* Pack-years of Smoking + p3* Environmental Carcinogen Exposure + j34* Family History of Lung Cancer + (35* Weighted Genetic Risk Score + p6* Biomarker 1 + (37* Biomarker 2 + p8* Biomarker 3 + ...); wherein to f3s are the coefficients or log odds ratios (log OR) estimated from the statistical modeling process; and wherein variables such as smoking status and the biomarkers represent an individual's risk factor values and are included as covariates in the model.

[0088] The specific values of these coefficients |3n (e.g. Pi to (38) are estimated during the model-building process, through logistic regression. These coefficients represent the contribution of each variable to the overall risk of lung cancer when all other variables are held constant. A machine learning algorithm has been developed.

[0089] Thus, in another aspect, the present invention relates to a computer program containing instructions and / or a computer-readable media or a data processing apparatus containing means for carrying out the method of the invention by combining data from the different techniques selected from PCR (ddPCR), NGS, Fast-seq, microbiome analysis of 16S rRNA sequences with the patient history, thereby improving the method accuracy of detecting subjects having lung cancer at all stages, in particular at early- stage.

[0090] In another aspect, the present invention relates to a kit directed to the method for identifying if a subject is at risk of developing lung cancer diseases, said kit comprising: i) specific probes for detecting genetic variants of cfDNA being fragments of target genes selected from the group comprising ALK, EGFR, MAP2K1 , PIK3CA, TP53, KRAS, MET, NRAS, CDH13, CHD1 , DAPK, RASSF1 , SEPT9, BRAF, ERBB2, and / or ROS1 ; and optionally ii) reagents and instructions for use.

[0091] For example, specific probes can be used for detecting genetic variants of cfDNA being fragments of target genes selected from the group comprising ALK, EGFR, MAP2K1 , PIK3CA and TP53.

[0092] Specific probes can also be used for detecting genetic variants of cfDNA being fragments of target genes selected from the group comprising ALK, EGFR, MAP2K1 , PIK3CA, TP53, KRAS, MET, and NRAS.

[0093] Specific probes can also be used for detecting genetic variants of cfDNA being fragments of target genes selected from the group comprising ALK, EGFR, MAP2K1 , PIK3CA, TP53, KRAS, MET, NRAS, CDH13, CHD1 , DAPK, RASSF1 , SEPT9, BRAF, ERBB2, and ROS1.

[0094] Specific probes can be prepared to detect genetic variants of cfDNA being fragments of target genes selected from the group comprising CDH13, CHD1 , DAPK, RASSF1 , SEPT9, ALK, BRAF, EGFR, ERBB2, KRAS, MAP2K1 , MET, NRAS, PIK3CA, ROS1 and / or TP53, such as those described in table 1 and / or table 8 targeting genetic variants of the target genes of the invention. Specific probes for analyzing cfDNA and / or ctDNA are described but not limited to those described in table 2 (SEQ ID NO: 1 to 56), or in table 7 (SEQ ID NO: 61 -87).

[0095] The kit may further comprise multiple swabs for collecting cfDNA during a prescribed period of time or re-usable swabs with interchangeable or washable heads, and a nucleic acid (DNA / RNA) preservation liquid for sample preservation.

[0096] EXAMPLES

[0097] Example 1 : High throughput screening of biomarkers in nasal swabs and saliva samples The study enrolled patients over 18 years of age, diagnosed with lung cancer independent of the stage. Legally incapable as well as patients that had recent nasal trauma or surgery, a markedly deviated nasal septum, a history of chronically blocked nasal passages or severe coagulopathy are discarded for this study. Patients that signed the informed consent, are subjected to nasal swab collection procedure which consists in gently and slowly rotating the swab against the nasal mucosa (anterior nare zone). The method does not require further introduction of the swab like nasopharyngeal testing or brushing of the nasal epithelium.

[0098] Nasal swab sample

[0099] The nasal swab collections is done by inserting the swab 1 -2 cm into one of the anterior nares. The swab is gently and slowly rotated against the nasal mucosa for 5-10 seconds, withdraw and repeated with the anterior nare using the same swab. The procedure is different from the nasopharyngeal swabbing method which requires a sample from the nasopharynx, on the upper part of the throat behind the nose and does not require further introduction of the swabs. The amount of sample collected is determined by stage of the patient. Diagnosed patients receive 1 swab and for screening 5-10 samples are taken to pull more DNA into the analysis. The resulting samples are stored at 4°C or at -20°C until further analyzed with DNA extraction kit method. DNA extraction from nasal swabs samples is performed using standardized kit methods (e.g. MagMAX™ Cell-Free DNA Isolation Kit with QIAamp DNA Midi Kit (Qiagen)).

[0100] Saliva sample

[0101] Donor will rinse their mouth 10 min before saliva collection. Saliva samples are collected using a saliva sample collection device (CIDA device) in combination with a Swab which is pressed 2 for two minutes bellow the tongue, above the tongue and in the right and left cheeks. The resulting samples are stored at -20°C until further analyzed with DNA extraction kit method such as using standardized kit methods.

[0102] Libraries panels preparation To evaluate the genomic profiling present in nasal swab samples, libraries can be prepared using Oncomine™ Pan-Cancer Cell-Free Assay (whole blood) and Oncomine™ Lung cfDNA Assay (Thermofisher) library panels for high throughput screening followed by next-generation sequencing (NGS). Together both library panels cover 63 genes, >150 lung cancer hotspots as well as short indels, MET exon 14 skipping, gene fusions, copy number genes and tumor suppressor genes (these panels are intended for plasma samples).

[0103] The procedure of the samples involves several steps, including reverse transcription of cfRNA, PCR amplification of target sequences, purification of amplicons, size selection, and analysis using next-generation sequencing (NGS).

[0104] Samples libraries preparation

[0105] DNA samples (max 40 ng) are pipeted in AB-1400 plates with superscript VILO™ Master Mix and place in thermal cycler to create cDNA. The plates are then centrifuged and 2uL of each library panel as mentioned above is added per sample in combination to the provided cfDNA master mix. Samples are then vortex, spin and place in the thermocycler for amplification, followed by purification using AMPurecourt beads according to the manufacturer protocol. The purified targets are then mixed with the Tag-sequencing barcode, library primer and cfDNA library PCR master mix and placed in the thermocycler for amplification.

[0106] To prepare the libraries for sequencing, the samples are centrifuged and mixed AMPurecourt beads for purification and analyzed by NGS.

[0107] Digital droplet PCR (ddPCR) ddPCR was performed using the ddPCRTM Supermix for Probes (NO dUTP, Bio-Rad) as described in the manufacturers protocol and the following probes.

[0108] Table 1

[0109] Kits: ddPCRTM Supermix for Probes (NO dUTP) van Bio-Rad (20x) 1863023

[0110] Multiplex PCR

[0111] A single multiplex PCR amplification is carried out to amplify using primer sequences as follows:

[0112] Table 2

[0113] PCR amplifications consist of 9 pl DNA, 10 pl KAPA2G Robust DNA Polymerase (KapaBiosystems) and 1 pl of primers (0.2 pM final for each primer). The PCR program is 95°C for 3 minutes, 35 x 95°C for 15s, 60°C for 15s, 72°C for 15s followed by 3 min at 72°C.

[0114] Example 2: Data analysis by next-generation sequencing (NGS) NGS protocol

[0115] Unaligned bam files generated by the Ion Torrent sequencer are mapped against the human reference genome (GRCh37 / hg19) using the TMAP 5.0.7 software with default parameters (https: / / github.com / iontorrent / TS). Subsequently, variant calling was done using the Ion Torrent specific caller, Torrent Variant Caller (TVC)-5.0.2, using the recommended Variant Caller Parameter for each cancer panel.

[0116] DNA sequencing data, identifying genetic variants, and infer disease progression, is generated by the Ion Torrent sequencer or equivalent software. Unaligned bam files, which contain the raw sequencing data, are processed to align the sequences to the human reference genome (GRCh37 / hg19), which provides a framework for identifying differences between the sample DNA and the reference genome. Once the sequencing data is aligned, variant calling is performed using the Ion Torrent-specific software (Torrent Variant Caller (TVC)-5.0.2.) which analyses the sequences to identify genetic variants. Variant interpretation is performed using Genetic Assistant software, which assigns functional prediction, conservation scores, and disease-associated information to each variant based on relevant databases and literature. Integrative Genomics Viewer is used to visually inspect the variants and validate their presence in the samples. Quantitative trait loci analyses, graphical Gaussian models, and Mendelian randomization are used to create directed links relating genetic variation in DNA mutations and transcriptional activity of pathways in cancer progression and recurrence. The genomic profiling of the microbiome present in the nose is also explored using rRNA gene sequencing targeting the V3-V4 hypervariable regions of the 16S rRNA gene (primers 341 F and 785R). The 5’ ends of the primers for each sample will be tagged with specific barcodes, followed by purification of the amplicons and quantification. Library preparation and sequencing will be performed using the standard instructions of the 16S Metagenomic Sequencing Library Preparation protocol (HluminaTM, Inc., San Diego, CA, United States). The measurements show the relative / absolute abundance of bacteria species in a sample in correlation with lung cancer. The presence of 16S rRNA gene, and thus genetic variants of 16S rRNA gene indicates that said subject has an increased risk of developing a lung cancer.

[0117] Table 3: Primer SEQ ID NO: 57-60

[0118] Example 3: Biomarkers for lung cancer In the present invention, over 200 potential cancer biomarkers have been detected in nasal fluids such as those described in Table 4.

[0119]

[0120] Biomarkers found in saliva are shown in Table 5 below.

[0121] Table 5: Biomarkers for lung cancer from saliva

[0122] Example 4: Biomarkers for lung cancer - Blood sample.

[0123] The biomarkers from blood can be measured in the nasal fluids. Table 6: Biomarkers for lung cancer from blood sample

[0124] Example 5: Screening process to identify if a person at risk of developing a lung cancer Figure 1 illustrates the process of identifying a person at risk of developing a lung cancer.

[0125] In a first protocol, a nasal swab sample is obtained from a person to identify whether the person is at risk of developing a lung cancer. DNA extraction from nasal swab samples is performed using the method as described in example 1 to prepare a library. A PCR reaction using the protocol as described herein is performed with the sample containing cell free DNA (cfDNA).

[0126] When amplification of cfDNA is observed following PCR reaction (SEQ ID NO: 1 -56), sequencing analysis such as Next Generation Sequencing is carried out to identify the presence of biomarkers specific for lung cancer or somatic copy number alterations (SCNA), which are present in lung cancer.

[0127] The absence of amplification of cfDNA following PCR reaction is indicative that the person is not at risk of developing a lung cancer.

[0128] In addition to the first protocol, or as an alternative, 16S-RNA seq is performed to analyze the nasal microbiome. The microbiome can identify smokers from non-smokers as well as correlate with the risk of lung cancer.

[0129] In addition to the first protocol, or as an alternative, DNA samples extracted from nasal swab sample are submitted to Fast Aneuploidy Screening Test-Sequencing System (Fast-seq). It is a faster and cost-effective method that simplifies the analysis of the genome by targeting repetitive fragments from different locations in the genome using a single primer pair (SEQ ID NO: 61 -87). With this alternative method, high throughput and decreased cost can be achieved by replacing laborious sequencing library preparation steps with PCR employing a single primer pair designed to amplify a discrete subset of repeated regions, which are present in lung cancer and involve alterations in a large portion of the cancer genome.

[0130] Fast-seq protocol

[0131] Amplicon libraries targeting LINE-1 (L1 ) elements are prepared from 1 ng of cell-free DNA (cfDNA). Target-specific L1 primers and Phusion Hot Start II Polymerase are utilized following the method outlined in a previous publication (https: / / doi.Org / 10.1002 / 1878-0261 .13196). The resulting PCR products undergo purification using AMPure Beads (Beckman Coulter) and serve as templates for a second PCR. In this step, sequencing adaptors and sample-specific indexes are incorporated into the amplicons (see table below). Sequencing libraries are quantified using the NEBNext Library Quant Kit for Illumina (New England Biolabs). Subsequently, libraries from 20 samples are pooled in equimolar amounts. The pooled libraries are then subjected to sequencing on the MiSeq platform (Illumina), generating a minimum of 90,000 single reads of 150 base pairs. Fast-Seq sequencing results of LINE-1 elements across the genome are mapped to the human reference genome hg19 using Burrows-Wheeler alignment (vO.7.17). For each chromosomal arm, a Z-score is computed as a measure of deviation from a reference panel of healthy / d iploid subjects. The Z-score is calculated by subtracting the mean and dividing by the standard deviation of normalized read-counts for the respective chromosome arm, enabling the assessment of over- and under-representation. Z-scores per chromosome arm are squared and then summed to derive a genome-wide aneuploidy score for each patient. Based on the cutoff established by Belie et aL, a genome-wide aneuploidy score is categorized as either high (> 5) or low (< 5).

[0132] Table 7

[0133] Multiplex PCR: The most promising biomarkers (3-10) are tested and ranked by Area Under Curve (AUC) using Logistic Regression models (python). The final model with the optimised combination is determined by the highest AUC of the cross-validations.

[0134] A PCR assay is then developed based on a cocktail mix used to amplify amplicons in the target regions simultaneously.

[0135] Achieving high specificity and sensibility: To achieve higher reduction in health costs we will use Receiver Operating Characteristic (ROC) Curve Analysis to plot and adjust trade-offs between sensitivity and specificity. During the screening process, biomarkers are selected that promote a higher sensitivity to lung cancer which will reduce the false negative rate.

[0136] The PCR test focused on sensitivity is combined with a highly specific NGS (Next- Generation Sequencing) which is performed only in positive patients (occurring in less than 0.5% of samples). This method allows to achieve highly specific and sensitive results, while maintaining cost-effectiveness. In this model an initial PCR assay allows to establish the initial risk of lung cancer by using the biomarkers with highest sensitivity. Repetitive nasal samples are taken to further increase such values on early stage and screening patients.

[0137] Example 6: cfDNA in nasal sample of patients

[0138] Nasal fluids from patients having lung cancer are obtained non-invasively using the nasal swab sample procedure as detailed in example 1 . A control sample is obtained from a healthy subject, having no cancer. Figure 2 shows the amount of total nucleic acids (ng / pL) and cfDNA (% of total nucleic acids) for each patient identified by number 1 to 12.

[0139] Both control and patient samples have shown cfDNA values ranging from 5 - 100 ng / pL at an average of 20 ng / pL from just 1 mL of nasal fluids. These values are at least 20 times higher than the ones obtained from 1 mL of plasma of lung cancer patients reported to be in the average of 0.838 ng / pL.

[0140] Advantageously, the amount of cfDNA required in the method of the invention is very low, meaning that all measurements can be done using at least one nasal sample from a subject with as little as 0.5 mL of a nasal fluid.

[0141] Example 7: Target genes and mutations with lung cancer patients Tag-sequencing analysis was performed by tagging the cfDNA / RNA from nasal fluid samples of lung cancer patients and healthy individuals. The panel used for this analysis can identify more than 150 lung cancer hotspots in 20 genes. These hotspots correspond to specific regions of the target genes that are associated with lung cancer.

[0142] The procedure of the sample preparation involves several steps including reverse transcription of cfRNA, PCR amplification of target sequences, purification of amplicons, size selection and NGS analysis as described in example 1 and 2.

[0143] Genetic variants in at least 8 genes have been identified in 15 patients and 68 lung cancer hotspots but not in control samples (not shown) of healthy persons, in which no genetic variant has been found (Table 8, below). The 68 lung cancer hotspots are mutations found in the target genes ALK, EGFR, KRAS, MAP2K1 , MET, NRAS, PIK3CA and TP53. They have been found using nasal fluid samples from lung cancer patients.

[0144] Table 8:

[0145] All these mutations are relevant for decision making of effective treatments for the corresponding patients. For example, anticancer agents such as Alectinib is used as a treatment of patients having mutations in ALK and Mobocertinib is used as a treatment of patients having EGFR mutation.

[0146] Further, patients under treatment with a known mutation in the target genes of the invention can be monitored carrying out a PCR reaction according to the method of the invention using nasal fluid samples to follow patient over time for their response to treatment.

[0147] Example 8: Performance of the method of the invention with lung cancer patients Nasal samples from 7 lung cancer patients have been tested for the target genes ALK, BRAF, EGFR, ERBB2, KRAS, MAP2K1 , MET, NRAS, PIK3CA, ROS1 and TP53. The method of the invention was performed using ddPCR and NGS as described in example 1 and 2.

[0148] As shown on figure 3, sample 1 is positive for ALK, EGFR and TP53, sample 2 is positive for TP53, sample 3 is positive for ALK, EGFR, sample 4 is positive for EGFR, TP53, sample 5 is negative for the target genes tested, sample 6 is positive for PIK3CA, and sample 7 is positive for MAP2K1 and TP53. The 3 control samples from healthy subjects are negative.

[0149] Thus, each lung cancer patient has a specific profile of the target genes selected from the group comprising ALK, EGFR, MAP2K1 , PIK3CA and TP53.

[0150] Example 9: Analysis of samples from patients with a KRAS mutation

[0151] Digital droplet PCR was performed using nasal fluids from 6 patients having lung cancer with a known KRAS mutation according to the protocol described in example 1 .

[0152] As shown on figure 4A, all samples have total droplet counts in the tens of thousands, sufficient for accurate quantification of DNA copies.

[0153] Data in figure 4B show the number of positives droplets. The presence of a significant number of positive droplets in each sample for lung cancer patient in comparison to the control sample indicates that the KRAS mutation is detectable in all samples. In figure 4C, the variant allele frequency (VAF) is the proportion of DNA sequencing reads that contain a specific genetic variant (eg. a mutation in KRAS) relative to the total reads covering that position in the gene.

[0154] For example, a VAF of 0.5 means that 0.5% of the analyzed molecules (cfDNA) contained that mutation indicating that there is ctDNA in the sample pool. In figure 4C, all nasal samples from lung cancer patients have a positive VAF indicating that all nasal samples contain ctDNA corresponding to the target gene KRAS.

[0155] Overall, these data indicate that nasal fluids are an effective way to both screen and diagnose for lung cancer non-invasively.

[0156] Example 10: Nasal microbiome analysis of patients with lung cancer

[0157] For the analysis of bacterial metagenomes, the fourth hypervariable domain (V4) of the microbial 16S ribosomal RNA (rRNA) gene has been amplified according to the method described in example 2. Nasal samples from 20 lung cancer patients have been analyzed and microbial classification was subsequently performed based on 16S rRNA amplicon NGS (Figure 5A).

[0158] Figure 5B shows data from the twelve most abundant bacteria genera found in lung cancer patients (n=20) in comparison to control samples (n=50 from a public available database) but more than 40 bacteria genera were found in samples.

[0159] Data were compared with the nasal microbiome of 50 control samples from healthy persons having no lung cancer showing marked differences between several bacteria types and indicating that the nasal microbiome can be used to predict the risk of lung cancer. In particular, the relative abundance of corynebacterium and moraxella have been found higher, whereas, the relative abundance of staphylococcus, streptococcus has been found lower in lung cancer patients. Thus, these comparative data show a correlation between the microbiome from nasal samples and the presence of lung cancer.

[0160] These data correlate with microbiome changes such as corynebacterium and streptococcus observed in lung tissues from lung cancer patients.

[0161] Since obtaining direct tissue from lung can be challenging, the method of the present invention based on nasal samples is an easier way to screen the lung microbiome. The lung microbiome has also been related with clinical factors such as smoking history and other lung diseases such as chronic obstructive pulmonary disease, cystic fibrosis, asthma, idiopathic pulmonary fibrosis.

[0162] Example 1 1 : Analysis of somatic copy number alterations of patients with lung cancer Nucleic acid sequencing analysis has been performed to analyze somatic copy number alterations (SCNA) which are present in nasal samples from 6 lung cancer patients using Fast Aneuploidy Screening Test-Sequencing System (Fast-Seq) as described in example 5.

[0163] Nasal fluid samples were submitted to Fast-Seq. Sequencing results of LINE-1 elements across the genome are mapped to the human reference genome and the aneuploidy score is calculated by subtracting the mean and dividing by the standard deviation of normalized read-counts for the respective chromosome arm of a control healthy patient having no cancer, enabling the assessment of over- and underrepresentation.

[0164] As shown in figure 6, the increase in aneuploidy scores in patients with lung cancer indicates the presence of ctDNA.

[0165] In blood samples, the aneuploidy score has been shown to correlate with the amount of circulating tumor cells (CTCs).

[0166] These data show that aneuploidy score obtained by using nasal samples comprising at least one or more cfDNA / ctDNA can be used as a primary screening step to evaluate tumor load without prior mutation knowledge. The method is cost-effective and suitable for various metastatic contexts, relying on the common presence of aneuploidy in tumors and requiring only 1 nanogram of cfDNA.

[0167] Example 12: Relative risk determination

[0168] This model includes the biomarkers as additional covariates alongside clinical and genetic factors. The biomarkers incorporated into this model are described in tables 4 and 5. These biomarkers are considered potential risk factors for lung cancer.

[0169] The relative risk (RR) in this model would include these biomarkers alongside clinical and genetic variables.

[0170] The following risk factors are considered for lifestyle risk: Age (years), Gender, Race or Ethnic Group, Education, BMI (kg / m2), COPD, Emphysema, Personal history of cancer, Family history of lung cancer, Personal history of pneumonia, Smoking status (Current smoker, Former smoker), Smoking duration (years), Smoking intensity (cigarettes per day), Pack-years of smoking, Years since cessation, Breathing upon exertion, Cough intensity, changes in voice, flu or pneumonia within the past 2 year, exposure to asbestos or other chemicals, current residence, Physical activity, Exposures (Coronary heart disease, Myocardial infarction, Stroke, Diabetes mellitus, Chronic kidney disease, Arthritis, Asthma), Personal history of cancer

[0171] The formula for RR is:

[0172] RR = exp(Pi * Smoking Status + (32* Pack-years of Smoking + p3* Environmental Carcinogen Exposure + (34* Family History of Lung Cancer + (35* Weighted Genetic Risk Score + p6* Biomarker 1 + (37* Biomarker 2 + p8* Biomarker 3 + ...)

[0173] Where, |3n (e.g. Pi to (3s) are the coefficients or log odds ratios (log OR) estimated from the statistical modeling process. Variables such as smoking status and the biomarkers represent an individual's risk factor values and are included as covariates in the model. The specific values of the coefficients |3n (e.g. Pi to f3s) are estimated during the modelbuilding process, through logistic regression. These coefficients represent the contribution of each variable to the overall risk of lung cancer when all other variables are held constant.

Claims

CLAIMS1 . A method for identifying if a subject is at risk of developing lung cancer, the method comprising detecting in at least one nasal sample from said subject at least one or more cell free DNA (cfDNA) being fragments of target genes selected from the group comprising ALK, EGFR, MAP2K1 , PIK3CA and TP53; and detecting the presence or absence of genetic variants of the said at least one or more cfDNA.

2. The method according to claim 1 , further comprising the target genes KRAS, MET, and NRAS.

3. The method according to any one of claims 1 -2, further comprising the target genes CDH13, CHD1 , DAPK, RASSF1 , SEPT9, BRAF, ERBB2, and ROS1 .

4. The method according to any one of the preceding claims, further comprising the step of detecting the presence or absence of 16S rRNA genes.

5. The method according to any one of the preceding claims, wherein the presence of genetic variants of the at least one or more cfDNA and / or of 16S rRNA genes indicates that said subject has an increased risk of developing and / or having lung cancer.

6. The method according to any one of the preceding claims, further comprising the step of treating the subject with at least one anticancer agent and / or a combination thereof based on the detected presence of genetic variants of the said at least one or more cfDNA.

7. The method according to claim 6, further comprising the step of monitoring over time the response of the subject to the treatment with the at least one anticancer agent and / or a combination thereof.

8. The method according to any of the preceding claims, wherein the lung cancer is selected from the group comprising non-small cell lung cancer (NSCLC), small cell lung cancer (SCLC), bronchogenic carcinoma, alveolar carcinoma, bronchialadenoma, sarcoma, lymphoma, chondromatous hamartoma, incidental pulmonary nodules related lung cancer, and mesothelioma.

9. The method according to any one of the preceding claims, wherein the at least one nasal sample is combined with at least one sample selected from salivary sample, plasma sample, cerebral spinal fluid (CSF) sample, pleural fluid sample, and / or urine sample.

10. The method according to any one of the preceding claims, wherein the at least one nasal sample is combined with at least one further nasal sample obtained from the subject at different time points during at least one week.11 . The method according to any one of the preceding claims, wherein the at least one or more cfDNA are detected in at least two nasal samples obtained from the subject at different time points.

12. The method according to any of the preceding claims, further comprising determining the clinical risk factors of said subject, said clinical risk factors are selected from the group comprising smoking status, age, gender, BMI, personal history of cancer, family history of lung cancer, personal history of pneumonia, and / or physical activity.

13. The method according to any one of the preceding claims, further comprising the analysis of the target genes comprising ALK, EGFR, MAP2K1 , PIK3CA and TP53.

14. The method according to any one of the preceding claims, wherein the detection of the genetic variants of the at least one or more cfDNA is determined by polymerase chain reaction (PCR) analysis selected from the group comprising digital droplet PCR (ddPCR) analysis and / or multiplex PCR analysis.

15. The method according to claim 14, wherein the PCR analysis comprises using probes hybridizing the genetic variants of the at least one or more cfDNA beingfragments of the target genes comprising ALK, EGFR, MAP2K1 , PIK3CA, TP53, KRAS, MET, NRAS, CDH13, CHD1 , DAPK, RASSF1 , SEPT9, BRAF, ERBB2, and / or ROS1.

16. The method according to any one of the preceding claims, further comprising analyzing the presence of genetic variants of the said at least one or more cfDNA by a nucleic acid sequencing analysis, preferably the nucleic acid sequencing analysis is next-generation sequencing analysis (NGS).

17. The method according to any one of the preceding claims, wherein the detection of at least one or more cfDNA is determined by fast aneuploidy screening test sequencing system (Fast-seq) and / or nucleic acid sequencing analysis, preferably the nucleic acid sequencing analysis is next-generation sequencing analysis (NGS).

18. The method according to any one of the preceding claims, wherein the genetic variants of the said at least one or more cfDNA are selected from the genetic variants as described in table 4.

19. A computer program containing instructions and / or a computer-readable media containing means for carrying out the method according to any one of claims 1 - 18.

20. A kit directed to the method for identifying if a subject is at risk of developing lung cancer according to any one of claims 1 -18, said kit comprising: i probes for detecting genetic variants of cfDNA being fragments of target genes selected from the group comprising ALK, EGFR, MAP2K1 , PIK3CA, TP53, KRAS, MET, NRAS, CDH13, CHD1 , DAPK, RASSF1 , SEPT9, BRAF, ERBB2, and / or ROS1 ; and optionally ii reagents and instructions for use.

Citation Information

Patent Citations

  • Diagnostic and prognostic methods for lung disorders using gene expression profiles from nose epithelial cells

    US20110217717A1

  • Kits and methods for testing for lung cancer risks

    WO2021046502A2