Highly sensitive and specific determination of DNA methylation profiles

A novel method for analyzing DNA methylation profiles using repetitive sequences and a classifier improves cancer detection sensitivity and universality, addressing limitations of existing liquid biopsy methods by targeting LINE-1 retrotransposons and providing comprehensive cancer monitoring.

JP2025537530APending Publication Date: 2025-11-18アンスティテュキュリ +5
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025525063
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-21
Filing Date
2023-09-01
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Current methods for detecting cancer through liquid biopsies are limited by low sensitivity and lack of universality, failing to capture tumor heterogeneity and requiring invasive tissue biopsies, and existing DNA methylation analysis methods are biased towards preselected mutations, making it difficult to establish a pan-cancer test.

Method used

A method involving clustering DNA subsequences, aligning at CpG dinucleotide positions, and using a classifier to distinguish between healthy and cancerous CpG methylation profiles, particularly targeting repetitive sequences like LINE-1 retrotransposons for comprehensive methylation analysis without relying on a reference genome.

Benefits of technology

Enhances sensitivity and universality in detecting cancer-specific signatures, enabling early-stage cancer detection and monitoring, and predicting treatment response by analyzing circulating DNA methylation patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025537530000001
    Figure 2025537530000001
  • Figure 2025537530000002
    Figure 2025537530000002
  • Figure 2025537530000003
    Figure 2025537530000003
Patent Text Reader

Abstract

The present invention relates to methods for determining the methylation profile of a DNA sequence of interest and for accurately distinguishing between healthy and cancerous methylation profiles, as well as kits for carrying them out.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD OF THE INVENTION The present invention relates to the field of medicine. In particular, the present invention relates to a method for determining the methylation profile of a DNA sequence of interest and a method for accurately distinguishing between healthy and cancerous methylation profiles, as well as a kit for carrying out the same.

[0002] Background of the Invention The discovery of mutations that contribute to the development and progression of human cancers has led to the emergence of a new generation of biomarkers. These genetic biomarkers are now recognized as useful in determining the most appropriate treatment, thereby enabling the development of targeted treatment plans and improving clinical outcomes. Current methods for identifying specific mutations in personalized medicine involve the use of tissue biopsies. However, biopsy procedures are invasive, painful, risky, and sometimes impractical. Furthermore, they fail to capture tumor heterogeneity due to spatial bias, and biopsies are also temporally biased because they are not typically performed repeatedly.

[0003] Liquid biopsies are an attractive method to complement or replace tissue biopsies. Extensive research has demonstrated that tumor genetic alterations can be detected in the plasma of patients with cancer. This paved the way for the implementation of molecular analysis performed on liquid biopsies to noninvasively determine tumor genotypes and demonstrated the potential of circulating tumor DNA (ctDNA) as a marker of cancer progression. Recent advances have demonstrated the advantages of ctDNA quantification, including the feasibility of large-scale tumor genotyping, since ctDNA originates from all tumor sites in the body. Furthermore, ctDNA is a powerful prognostic factor, capable of detecting clinically imperceptible tumor masses after surgery or during treatment. These approaches promise optimal management of cancer patients and currently play an important role in oncology.

[0004] However, several technical hurdles still limit their widespread use. Samples collected during early stages of tumor progression or during and after treatment may contain fewer than one copy of a mutation per milliliter of plasma. This is below the detection limit of most commonly used techniques, even when simultaneously testing for multiple gene mutations. Furthermore, most methods are biased toward preselected recurrent mutations and do not cover all tumors. In our study, we observed that approximately 25% of breast cancer patients, even in advanced stages, had no traceable common mutations in their plasma DNA. The sensitivity of methods targeting genetic alterations remains limited by the low number of recurrent mutations detectable per genome. Therefore, there is a need to develop novel detection tools that are more sensitive and informative.

[0005] Multiple studies have demonstrated that epigenetic processes play a central role in the development, progression, and treatment of cancer. Epigenetic changes (i.e., chromatin modification patterns, e.g., changes in DNA methylation and histone modifications) are promising candidates for cancer detection, diagnosis, and prognosis. These markers provide an additional level of information that is ignored by methods that examine genetic alterations alone. Unlike point mutations, which affect only a single base pair per genome, epigenetic changes are distributed throughout the genome and affect multiple residues per region. Therefore, novel diagnostic strategies integrating epigenetic biomarkers will not only achieve improved sensitivity but also cover cases without detectable mutations. Because the epigenetic landscape is highly cell type-specific, epigenetic markers may also provide information about the tumor's tissue of origin. This is different from the case of oncogenic mutations, which are often common across multiple cancer types. Epigenetic markers may play a crucial role in detecting early stages of cancer when the chances of recovery are greatest, residual disease, early relapse, or acquired resistance during treatment. This will allow for better monitoring of cancerous disease and provide a new therapeutic window for treatment and cure.

[0006] Aberrant DNA methylation is a hallmark of neoplastic cells, which exhibit global genomic hypomethylation combined with widespread hypermethylation of tumor suppressor genes. DNA methylation is a stable modification that affects numerous CpG sites per region and genome. Furthermore, concordance of methylation status between multiple CpGs from the same region may help detect low-frequency abnormalities among heterogeneous molecular populations. Finally, combining multiple genomic regions allows for the capture of a broad range of tumor alleles and coverage of heterogeneous cancer patient profiles.

[0007] To date, most studies investigating methylation patterns in plasma DNA have either used PCR-based methods to target a small number of regions at high depth or high-throughput sequencing to probe the entire genome at low depth. Both approaches have limited sensitivity. More recent studies have investigated the performance of more regions at high depth based on capturing regions of interest combined with deep sequencing. These methods have enabled highly sensitive detection and classification of cancers from plasma DNA. However, these methods primarily focus on hypermethylated and unique sequences in cancers, necessitating the targeting of specific regions for each cancer subtype. This makes it difficult to establish a universal pan-cancer test.

[0008] Therefore, there is a strong need to develop novel methods capable of assessing the potential of DNA methylation, especially circulating DNA methylation, as a universal tumor biomarker, and novel sensitive strategies for detecting cancer-specific signatures, preferably in blood samples.

[0009] Summary of the Invention The present invention relates to a method for determining the CpG methylation profile of at least one DNA sequence of interest or any fragment thereof, comprising: a) clustering a set of sub-sequences obtained from a DNA sequence of interest into clusters of sub-sequences; b) for each cluster, selecting one subsequence from among the subsequences of the cluster as a reference sequence; c) aligning the reference sequences of said clusters by aligning at CpG dinucleotide positions; d) for each cluster, aligning the remaining subsequences to the selected reference sequence; e) determining the CpG methylation status of each subsequence by determining whether the CpG dinucleotide at each CpG site of the subsequence is methylated or not, thereby determining a CpG methylation profile of the subsequence, including the CpG methylation level and / or the proportion of CpG methylated haplotypes; Including, wherein the DNA sequence of interest is or consists of repetitive sequences, said repetitive sequences being distributed throughout the genome of the subject and preferably containing a high density of CpG dinucleotides; wherein the method optionally comprises a first step of obtaining or providing a set of subsequences of said DNA sequence of interest, The present invention relates to a method, wherein the method optionally comprises repeating some or each of steps a) through e) for another set of subsequences from the DNA sequence of interest.

[0010] In particular, the repeat sequence is a retrotransposon, such as a LINE, HERV, SINE, SVA or a subfamily thereof, such as, in particular, LINE-1, L1PA, HERV-K and Alu, or a satellite repeat, such as a Sat2 or Sat3 element, preferably a LINE-1 retrotransposon or any fragment or variant thereof, more preferably a LINE-1 retrotransposon as set forth in SEQ ID NO: 2 or 29 or any fragment or variant thereof.

[0011] The present invention also provides a computer-implemented method for training a classifier to accurately distinguish between healthy and cancerous CpG methylation profiles, comprising: a) providing a training set of CpG methylation profiles of DNA sequences of interest or subsequences thereof, wherein the DNA sequences of interest are repeated and distributed throughout the genome and contain a high density of CpG dinucleotides, or preprocessed information obtained from the training set of CpG methylation profiles of DNA sequences of interest or subsequences thereof, as input to a classifier, wherein the training set of CpG methylation profiles includes CpG methylation profiles of DNA sequences or subsequences thereof from subjects identified as healthy subjects and subjects identified as cancer subjects; b) generating an output of the classifier for each CpG methylation profile input of the DNA sequence of interest or a subsequence thereof, wherein the output classifies the CpG methylation profile input of the DNA sequence of interest or a subsequence thereof as a healthy CpG methylation profile or a cancerous CpG methylation profile. Including, The present invention relates to a method wherein the CpG methylation profile comprises the CpG methylation level and / or the proportion of CpG methylated haplotypes of the DNA sequence or a subsequence thereof.

[0012] In particular, the CpG methylation profile of a DNA sequence of interest or a subsequence thereof is determined by the methods described herein.

[0013] The present invention also provides an in vitro or in silico method for determining the health status of a subject, in particular for determining whether a subject is a healthy subject or a subject suffering from cancer or cancer recurrence, comprising: (a) providing a DNA sequence of interest or a subsequence thereof from a subject, or preprocessed information obtained from said DNA sequence or subsequence, said DNA sequence of interest being a DNA sequence that encodes a repetitive sequence containing a high density of CpG dinucleotides distributed throughout the subject's genome, as input to a classifier trained to distinguish between healthy and cancerous CpG methylation profiles; (b) using the classifier to identify the CpG methylation profile of the subject's DNA sequence of interest or a subsequence thereof as an output of the classifier as a healthy CpG methylation profile or a cancerous CpG methylation profile.

[0014] The present invention also provides an in vitro or in silico method for determining the origin of a tumor in a subject, comprising: (a) providing a DNA sequence of interest or a subsequence thereof from a subject, or preprocessed information obtained from said DNA sequence or subsequence, said DNA sequence of interest being a DNA sequence that is distributed throughout the subject's genome and encodes a repetitive sequence containing a high density of CpG dinucleotides, as input to a classifier trained to distinguish between healthy CpG methylation profiles and cancerous CpG methylation profiles from various tumor origins; (b) using the classifier to identify a CpG methylation profile of the subject's DNA sequence of interest or a subsequence thereof as an output of the classifier as a healthy CpG methylation profile or a cancerous CpG methylation profile from a particular tumor origin, thereby determining the origin of a tumor from the subject.

[0015] The present invention also provides an in vitro or in silico method for determining the stage of a tumor in a subject, comprising: (a) providing a DNA sequence of interest or a subsequence thereof from a subject, or preprocessed information obtained from said DNA sequence or subsequence, said DNA sequence of interest being a DNA sequence that is distributed throughout the subject's genome and encodes a repeat sequence containing a high density of CpG dinucleotides, as input to a classifier trained to distinguish between healthy CpG methylation profiles and various stages of cancerous CpG methylation profiles; (b) using the classifier to identify a CpG methylation profile of the subject's DNA sequence of interest or a subsequence thereof as an output of the classifier as a healthy CpG methylation profile or a cancerous CpG methylation profile of a particular stage, thereby determining the stage of a tumor from the subject.

[0016] The present invention also provides an in vitro or in silico method for monitoring the response of a subject suffering from cancer to an anti-cancer treatment, comprising: (a) providing a classifier trained to distinguish between DNA sequences having a healthy CpG methylation profile and DNA sequences having a cancerous CpG methylation profile, as a first input, at least one DNA sequence of interest or a subsequence thereof from a first liquid biopsy from a subject suffering from cancer prior to administration of an anti-cancer treatment to the subject, said DNA sequence of interest being repeated throughout the subject's genome and comprising a high density of CpG sites or fragments thereof, or preprocessed information obtained from said first liquid biopsy, and as a second input, a second liquid biopsy from said subject after administration of an anti-cancer treatment, said second liquid biopsy comprising at least one DNA sequence of interest or a subsequence thereof, or preprocessed information obtained from said second liquid biopsy; (b) using the classifier to identify each CpG methylation profile of each DNA sequence of the first liquid biopsy as having a healthy CpG methylation profile or a cancerous CpG methylation profile as a first output of the classifier, and to identify each CpG methylation profile of each DNA sequence of the second liquid biopsy as having a healthy CpG methylation profile or a cancerous CpG methylation profile as a second output of the classifier; wherein the number of DNA sequences of interest classified as having a healthy CpG methylation profile in the second output of the classifier exceeds the number of DNA sequences of interest classified as having a healthy CpG methylation profile in the first output of the classifier, indicating that the subject will be responsive to the anti-cancer treatment; and wherein the number of DNA sequences of interest classified as having a healthy CpG methylation profile in the second output of the classifier is equal to or less than the number of DNA sequences of interest classified as having a healthy CpG methylation profile in the first output of the classifier, indicating that the subject will not be responsive to the anti-cancer treatment.

[0017] The present invention also provides an in vitro or in silico method for assessing the efficacy of a compound to revert a cancerous CpG methylation profile of a DNA sequence of interest from a subject suffering from cancer to a healthy CpG methylation profile, comprising: (a) providing a DNA sequence of interest or a subsequence thereof from a subject treated with a compound, said DNA sequence of interest being repeated and distributed throughout the subject's genome and containing a high density of CpG dinucleotides or any fragment thereof, or preprocessed information obtained from said at least one DNA sequence of interest or a subsequence thereof, as input to a classifier trained to distinguish between DNA sequences having a healthy CpG methylation profile and DNA sequences having a cancerous CpG methylation profile; (b) using the classifier to detect DNA sequences with healthy CpG methylation profiles and / or DNA sequences with cancerous CpG methylation profiles as an output of the classifier; The present invention relates to a method wherein the amount of DNA sequences having a healthy methylation profile exceeds a reference amount of DNA sequences having a healthy methylation profile obtained from the subject before any treatment with the compound, indicating that the compound is capable of reverting the cancerous CpG methylation profile to a healthy CpG methylation profile.

[0018] The present invention also provides an in vitro or in silico method for predicting the ability of a compound to treat cancer, comprising: assessing the efficacy of the compound to revert a cancerous CpG methylation profile of a DNA sequence of interest from the subject to a healthy CpG methylation profile; The present invention relates to a method, wherein an amount of DNA sequences classified as having a healthy CpG methylation profile above a reference amount indicates that the compound is useful for treating the cancer.

[0019] In particular, in the testing methods disclosed herein, the CpG methylation profile of a DNA sequence of interest or a subsequence thereof is determined by the methods for determining a CpG methylation profile disclosed herein.

[0020] In some embodiments, the classifier is trained according to the training methods disclosed herein.

[0021] In some embodiments, the DNA sequence of interest is a circulating cell-free DNA (cfDNA) sequence.

[0022] The present invention also provides - a memory storing at least one instruction for a classifier trained according to the training method disclosed herein; a processor accessing said memory to read the instructions and to perform any of the methods of the invention.

[0023] The present invention also provides a kit of primers or probes targeting a partial sequence of a DNA sequence encoding a LINE-1 retrotransposon, preferably a LINE-1 retrotransposon as set forth in SEQ ID NO: 2 or 29, The kit comprises at least four primers or probes selected from the group of primers or probes having the sequence shown in SEQ ID NO: 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 or 26, respectively, or a sequence having at least 85% identity thereto.

[0024] The present invention also relates to the use of said kit for amplifying a partial sequence of a DNA sequence encoding a LINE-1 retrotransposon, preferably a LINE-1 retrotransposon as set forth in SEQ ID NO: 2 or 29, for the diagnosis of cancer, The use relates to such cancers being preferably selected from the group consisting of colon cancer, breast cancer, lung cancer, uveal melanoma cancer, ovarian cancer and gastric cancer.

[0025] Detailed Description of the Invention Improved methods for detecting multiple types of cancer are needed, including the need for efficient cancer screening at an early stage.

[0026] Notably, cancer-associated hypomethylation has been reported in almost all classes of repetitive sequences, from dispersed retrotransposons to clustered satellite repeat DNA, and in multiple forms of cancer.

[0027] To gain a comprehensive picture of hypomethylation during carcinogenesis and improve sensitivity, we choose to target retrotransposons, particularly those of the long interspersed nucleobase-1 family (L1). Indeed, these elements, present in thousands of copies per genome, are hypomethylated in multiple cancers.

[0028] Two studies have investigated the global methylation profile of L1 in lung and colon cancer plasma using qPCR-based methods, but reported low detection sensitivities of less than 70%. Indeed, detecting methylation profiles at single-base-pair resolution in repeats requires sophisticated downstream analysis due to the inherent difficulty of mapping these sequences. To overcome this, we developed a computational tool for accurately aligning sequencing data without a reference genome. We implemented a predictive model trained by a machine learning algorithm to integrate methylation patterns globally and at the single-molecule level.

[0029] The present invention dramatically improves the sensitivity of DNA detection in a cost-effective manner and provides an optimal trade-off between the number of targeted regions and sequencing depth.

[0030] This specification also relates to a novel method that uses multiple cancer hypomethylation markers to capture a wide range of tumor alleles and cover the heterogeneous profiles of cancer patients in a single test. This method investigates selected regions that provide genome-wide information. This is because repetitive elements, such as retrotransposons, retain half of the CpG sites present in the human genome. This makes it possible to generate highly accurate and comprehensive methylation profiles from as little as a few nanograms of DNA using available sequencing depth. This method can be widely used to develop routine clinical tests.

[0031] The greatest originality and competitive advantage of the present invention is its ability to investigate DNA methylation at the level of repeat sequences. Repeat hypomethylation is common in many, if not all, cancer types and is an interesting marker for pan-cancer detection. Previous studies have set these regions aside because mapping these regions is inherently challenging and differentially methylated region (DMR) analysis is typically performed on mapped data. The present inventors have developed a novel solution for detecting methylation profiles in repeats at single-base-pair resolution without relying on mapping to a reference genome. This allows for high sensitivity and preservation of most of the data, essential for handling minute amounts of DNA. As shown herein, the method disclosed by the present inventors demonstrated high performance in detecting cancer samples and was evaluated in six cancer types, including three early-stage cancers of interest for pan-cancer diagnosis.

[0032] definition In order that the present invention may be more readily understood, certain terms are defined herein. Additional definitions are set forth throughout the detailed description.

[0033] Unless otherwise specified, all technical terms, notations, and other scientific terms used herein are intended to have the meaning commonly understood by one of ordinary skill in the art to which this invention belongs. In some cases, terms having a commonly understood meaning are defined herein for clarity and / or ready reference, and the inclusion of such definitions herein should not necessarily be construed as representing a deviation from what is commonly understood in the art. The techniques and procedures described or referenced herein are generally well understood and commonly employed using conventional methodology by those skilled in the art.

[0034] As used herein, "CpG" or "CG" are used interchangeably and refer to cytosine ("C") and guanine ("G") nucleotides linked by a phosphodiester bond, and in particular to the specific CG dinucleotide located at a "CpG site." In mammals, DNA methylation occurs predominantly at CpG dinucleotides. In certain embodiments, the cytosine residue of a CpG dinucleotide is methylated to 5-methylcytosine.

[0035] As used herein, "CpG island" refers to a stretch of DNA, particularly circulating cell-free DNA (cfDNA, also identified herein as "cell-free DNA"), in which the frequency of CpG sites is high compared to other regions of DNA. To be recognized as a CpG island, a sequence must meet the following criteria: a (G+C) content of 0.50 or greater, a CpG dinucleotide ratio of 0.60 or greater, and both of these occurring within a 200 bp or 500 bp sequence window.

[0036] As used herein, the term "methylation" refers to the methylation of a cytosine residue, particularly the methylation of the C5 position of cytosine and / or the N4 position of cytosine, preferably the methylation of the C5 position of cytosine. A cytosine contained in a CpG site that may be methylated is called a "methylation-prone cytosine." A methylated cytosine is called a "methylated cytosine." In some examples, methylation specifically refers to the methylation of a cytosine residue present in a CpG site.

[0037] As used herein, the term "differentially methylated" refers to a CpG methylation site whose methylation profile differs between a first state and a second state, e.g., a healthy state and a cancerous state.

[0038] As used herein, the term "hypomethylation" refers to a lower level of methylation that can be reported at the level of nuclear regions or CpG island CG dinucleotides in a state of interest compared to a reference state (e.g., at least one less methylated cytosine in a cancer state than in a healthy control).

[0039] As used herein, the term "hypermethylation" refers to a high level of methylation that can be reported at the level of a CG dinucleotide, nuclear region, or CpG island in a state of interest compared to a reference state (e.g., at least one more methylated cytosine in a cancer state than in a healthy control).

[0040] As used herein, the term "subsequence of a DNA sequence" refers to a portion or fragment of an original DNA sequence. A subsequence specifically consists of contiguous nucleic acids from the original DNA sequence. A subsequence of a DNA sequence is shorter (i.e., contains fewer nucleic acids) than the original DNA sequence.

[0041] As used herein, the term "amplicon" or "amplicon molecule" refers to a nucleic acid molecule generated by amplification of a template nucleic acid molecule, such as cfDNA or a nucleic acid molecule having a sequence complementary thereto, or a double-stranded nucleic acid comprising any such nucleic acid molecule. As used herein, the term "oligonucleotide primer" or "primer" refers to a nucleic acid molecule used, available, or intended for use in generating an amplicon from a template nucleic acid molecule. Under amplification-permissive conditions (e.g., in the presence of nucleotides and DNA polymerase and at an appropriate temperature and pH), an oligonucleotide primer can provide a starting point for amplification from the template to which it hybridizes. Typically, an oligonucleotide primer is a single-stranded nucleic acid between 5 and 200 nucleotides in length. Those skilled in the art will understand that the optimal primer length for generating an amplicon from a template nucleic acid molecule may vary depending on conditions including temperature parameters, primer composition, and transcription or amplification method. An oligonucleotide primer pair, as used herein, refers to a set of two oligonucleotide primers complementary to the first and second strands, respectively, of a template double-stranded nucleic acid molecule. The first and second members of an oligonucleotide primer pair may be referred to as the "forward" and "reverse" oligonucleotide primers, respectively, relative to a template nucleic acid strand. The forward oligonucleotide primer is capable of hybridizing to a nucleic acid strand complementary to the template nucleic acid strand, and the reverse oligonucleotide primer is capable of hybridizing to the template nucleic acid strand, with the position of the forward oligonucleotide primer relative to the template nucleic acid strand being 5' to the position of the reverse oligonucleotide primer sequence relative to the template nucleic acid strand.Those skilled in the art will understand that the identification of the first and second oligonucleotide primers as forward and reverse oligonucleotide primers, respectively, is arbitrary insofar as this identification depends on whether a given nucleic acid strand or its complementary strand is utilized as the template nucleic acid molecule.

[0042] As used herein, the term " probe " refers to a single-stranded or double-stranded nucleic acid molecule that can hybridize with complementary target (e.g., DNA, cfDNA or amplicon), and comprises detectable moiety.In some instances, for example, as defined herein, probe is a capture probe that is useful for detecting, identifying and / or isolating target sequence (e.g., gene sequence).

[0043] As used herein, "sequence identity" between two sequences is described by the parameters "sequence identity," "sequence similarity," or "sequence homology." In the context of the present invention, "sequence identity" between two sequences (A) and (B) is determined by comparing the two sequences aligned in an optimal manner over a comparison window. The sequence alignment can be performed by methods well known in the art, for example, using the Needleman-Wunsch global alignment algorithm or the Smith-Waterman local alignment algorithm. Analysis software matches similar sequences using similarity measures that account for various deletions and other modifications. Once all alignments are obtained, the % identity can be obtained by dividing the total number of aligned identical nucleic acid residues by the total number of nucleic acid residues contained in the longest sequence between sequences (A) and (B). To compare two nucleic acid sequences, for example, BLAST or EMBOSS Needle tools can be used. EMBOSS Needle creates an optimal global alignment of two sequences using the Needleman-Wunsch algorithm.

[0044] As used herein, the term "diagnosis" refers to determining whether a subject has or will develop a disease, disorder, condition, or state, and / or the qualitative or quantitative probability / likelihood thereof. For example, when diagnosing cancer, diagnosis may include determining the risk, type, stage, grade, or other classification of the cancer. In some instances, for example, as described herein, diagnosis may be or include determining the prognosis and / or likely response to one or more general or specific therapeutic agents or regimens.

[0045] The term "treatment" refers to any action intended to improve a patient's health status, for example, the cure, prevention, prophylaxis, and delay of a disease or disease symptoms. Treatment designates both curative and / or prophylactic treatment of a disease. Curative treatment is defined as treatment that results in the cure of, or that alleviates, improves, and / or eliminates, relieves, and / or stabilizes, a disease or disease symptoms or the suffering caused directly or indirectly by a disease. Preventative treatment includes both treatment that results in the prevention of a disease and treatment that reduces and / or delays the progression and / or onset of, or the risk of developing, a disease. In certain embodiments, such terms refer to the amelioration or eradication of a disease, disorder, infection, or symptoms associated therewith. In other embodiments, the term refers to minimizing the spread or worsening of cancer. Treatment according to the present invention does not necessarily imply 100% or complete treatment. Rather, there are various degrees of treatment that are recognized by those skilled in the art as having potential benefit or therapeutic effect. Preferably, the term "treatment" refers to the application or administration of a composition comprising one or more active agents to a subject with a disorder / disease.

[0046] As used herein, the term "classifier performance" refers to the predictive ability of a machine learning model. To measure the performance of a classifier, various types of classification performance metrics, such as accuracy, sensitivity, specificity, or area under the ROC curve, are used.

[0047] As used herein, the term "computer-implemented method" refers to a method involving a programmable apparatus / device, particularly a computer, a computer network, or a computer-readable medium carrying a computer program, wherein at least one step of the method is performed using at least one computer program. The computer-implemented method may further include at least one step that is not performed using a computer program.

[0048] As used herein, the term "and / or" means that each of the two specified features or components is deemed to be specifically disclosed, with or without the other. For example, "A and / or B" means that (i) A, (ii) B, and (iii) A and B are each deemed to be specifically disclosed, as if each were individually stated.

[0049] The articles "a" and "an" are used herein to refer to one or to more than one (i.e., to at least one) of the grammatical object of the article. By way of example, "an element" refers to one element or to more than one element.

[0050] The term "about," when used herein with respect to a value, refers to a value similar to the referenced value, given the context. Generally, a person of ordinary skill in the art familiar with the context will understand the relative degree of variation encompassed by the term "about" given the context. For example, in some embodiments, e.g., as described herein, the term "about" may encompass a range of values ​​within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or fractions thereof, of the referenced value.

[0051] Methods of the Invention Determination of CpG methylation profile In a first aspect, the present invention provides a method, in particular a computer-implemented method, for determining the CpG methylation profile of a DNA sequence of interest, comprising: a) clustering a set of subsequences obtained from a DNA sequence of interest into clusters of subsequences; b) for each cluster, selecting one subsequence from among the subsequences of the cluster as a reference sequence; c) aligning the reference sequences of said clusters by aligning at CpG dinucleotide positions; d) for each cluster, aligning the remaining subsequences to the selected reference sequence; e) determining the CpG methylation status of each subsequence by determining whether the CpG dinucleotide at each CpG site of the subsequence is methylated or not, thereby determining a CpG methylation profile of the subsequence, including the CpG methylation level and / or the proportion of CpG methylated haplotypes; Including, wherein the DNA sequence of interest is distributed throughout the genome of the subject, and is preferably a DNA sequence from the subject encoding a repetitive sequence containing a high density of CpG dinucleotides or any fragment of said repetitive sequence; wherein the method optionally comprises a first step of obtaining or providing a set of subsequences of said DNA sequence of interest, The present invention relates to a method, wherein the method optionally comprises performing / repeat (one or more times) some or each of steps a) to e) for another set of subsequences.

[0052] The method may also include an additional final step of comparing the CpG methylation profiles (determined in step e)) to optimize the CpG methylation profile.

[0053] In certain embodiments, the DNA is obtained / isolated / extracted from or comprised / included in a biological sample, particularly a biological sample from a subject, such biological samples being particularly described below in the section "Subjects and Biological Samples."

[0054] Preferably, the DNA sequence of interest is retrieved from cfDNA, where the methylation patterns of cellular DNA are preserved in cell-free DNA (cfDNA).

[0055] As used herein, "circulating cell-free DNA" and "cfDNA" refer to DNA fragments released from cells into body fluids, such as plasma. This term includes normal circulating cell-free DNA, circulating tumor DNA (ctDNA), cell-free mitochondrial DNA (cf-mtDNA), and cell-free fetal DNA (cf-fDNA). The terms "circulating tumor DNA" and "ctDNA" are used interchangeably and refer to a portion of circulating cell-free DNA released from cancer cells.

[0056] Preferably, the subsequence of the DNA of interest is a sequence containing a high density of CpG sites.

[0057] As used herein, "high density of CpG dinucleotides" or "high CpG dinucleotide density" refers to the density of CpG dinucleotides normalized by the density of G and C nucleotides in a DNA sequence or refers to the number of CpG dinucleotides in a DNA sequence. A CpG dinucleotide density is considered "high" if the ratio of observed to expected CpG dinucleotides (observed CpG / expected CpG) is 0.6 or greater.

[0058] In one embodiment, the DNA of interest contains at least 10, 15, 20, 25, 30, 35, 40, 45, or 50 CpG sites. In particular, the DNA of interest contains 20 to 40 CpG sites, and more particularly 30 to 35 CpG sites. Such CpG sites are particularly distributed over one or more DNA subsequences.

[0059] In a particular embodiment, the DNA sequence of interest is a DNA sequence that contains differentially methylated CpG dinucleotides, particularly those DNA sequences that are known to be hypermethylated in healthy subjects and hypomethylated in subjects suffering from cancer.

[0060] In certain embodiments, the DNA sequence and / or subsequence of interest comprises at least one differentially methylated region, preferably containing a high density of CpG sites.

[0061] As used herein, the term "differentially methylated region" or "DMR" refers to a DNA sequence that contains one or more differentially methylated CpG sites between a first state and a second state, e.g., a healthy state and a cancer state. A DMR that contains a greater or higher number or frequency of methylated CpG sites in a selected state of interest, e.g., a cancer state, may be referred to as a "hypermethylated DMR." A DMR that contains a lower number or frequency of methylated sites in a selected state of interest, e.g., a cancer state, may be referred to as a "hypomethylated DMR." The term "DMR" also designates a DNA sequence that has a different methylation profile, i.e., different methylation levels and / or methylation patterns, between a first state and a second state, e.g., a healthy state and a cancer state. In certain embodiments described herein, a DMR is a subsequence of a DNA sequence of interest. In certain embodiments described herein, a DMR is an amplicon produced by amplification using, for example, oligonucleotide primers, e.g., an oligonucleotide primer pair selected for amplification of the DMR or for amplification of a DNA subsequence of interest. In certain aspects described herein, a DMR is a region of DNA amplified by an oligonucleotide primer pair, e.g., a region having the sequence of or complementary to the oligonucleotide primer. Preferably, the DMR has a high density of CpG sites.

[0062] The DNA contemplated by the present invention is a DNA sequence containing repetitive sequences, preferably cfDNA or any fragment thereof. Preferably, such repetitive sequences contain one or more DMRs.

[0063] "Repetitive sequence," "repeated sequence," or "repeat" are used interchangeably herein to refer to multiple copies of a nucleotide sequence in a genome. They are abundantly distributed in the genomes of eukaryotic organisms. Repetitive sequences can be easily recognized as two large families: "tandem repeats" and "interspersed repeats."

[0064] Preferably, there are at least 100, at least 1000, at least 10,000, at least 100,000 or at least 1,000,000 copies of the repeat sequence in the genome of the subject.

[0065] In particular embodiments, the repeat sequence is selected from the group consisting of LINE, HERV, SINE, SVA, Sat2 and Sat3 elements (subfamilies thereof, e.g., L1PA, HERV-K and Alu) or any variant (i.e., analogous version) thereof. A variant of a particular (reference) sequence is a sequence that has at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 98% sequence identity with said particular (reference) sequence.

[0066] In one embodiment, the repeat sequence of the present invention is a tandem repeat sequence. Tandem repeats are composed of one or more nucleotides repeated head-to-tail in blocks or arrays, and are usually non-coding sequences. Depending on the size and total length of the repeat unit, they can be further classified into satellites (sat1, sat2, sat3, centromeric alpha-satellite, telomere), minisatellites (variable number of tandem repeats (VNTR)) and microsatellites (simple sequence repeats, SSR).

[0067] Preferably, the tandem repeat sequence is Sat2 or Sat3 satellite, in particular human satellite 2 or 3, or a fragment or variant thereof. Sat2 and Sat3 satellites are particularly described in Altermose et al. PLOS Computational Biology 2014 Volume 10 Issue 5 e1003628. SAT2 / 3 is rich in tandem repeats of pentameric GGAAT and divergent sequences containing CGGAT.

[0068] In another embodiment, the repetitive sequences are interspersed repeats, also called "interspersed repeats." Transposable elements (transposons), such as DNA transposons and retrotransposons, are interspersed repeats.

[0069] In certain embodiments, the repeat sequence is a transposon or retrotransposon. A transposable element or transposon is a small DNA segment that can replicate and insert DNA copies into random sites on the same or different chromosomes. In eukaryotes, such as humans, transposons can be classified into class I or class II. In particular, class I elements (so-called copy-and-paste retrotransposons) use reverse-transcribed RNA intermediates to produce their own copies, while class II elements (so-called cut-and-paste DNA transposons) excise from the donor site and reintegrate into another location in the genome (Wicker, T. et al. Nat. Rev. Genet. 8, 973-982 (2007)). Thus, the repeat sequence may be a class I or II transposon.

[0070] Preferably, the repetitive sequences are class I transposons, i.e. retrotransposons, in particular, repetitive sequences are evolutionarily young retrotransposons, preferably specific to the human or primate genome.

[0071] In an embodiment, the repeat sequence is a SINE or SINE-VNTR-Alu (SVA) element or any fragment or variant thereof.

[0072] The term "short interspersed element repeats" (SINEs) refers to retrotransposons, which constitute one of the major components of the genomic repeat fraction. More than one million copies of one class of retrotransposons, short interspersed element repeats (SINEs), are present in mammalian genomes, particularly in gene-rich regions of the genome.

[0073] The term "SINE-VNTR-Alu" (SVA) is used herein to refer to a nonautonomous, hominidae-specific retrotransposon known to be associated with human disease. SVAs are evolutionarily young and likely translocated by LINE-1 reverse transcriptases. SVA elements affect the host through a variety of mechanisms, including insertional mutagenesis, exon shuffling, alternative splicing, and generation of differentially methylated regions (DMRs). A typical SVA averages approximately 2 kilobases (kb), but SVA insertions can range in size from 700 to 4,000 base pairs (bp). SVA retrotransposons have been described, inter alia, by Hanks and Kazazian (Semin Cancer Biol. 2010 Aug; 20(4):234-245) and Gianfrancesco et al. (Neuropeptides. 2017 Aug; 64: 3-7).

[0074] In an embodiment, the repeat sequence is an Alu element or any fragment or variant thereof.

[0075] The term "Alu transposon" or "Alu element" refers to short stretches of DNA originally characterized by the action of the Arthrobacter luteus (Alu) restriction endonuclease. Alu elements are the most abundant transposable elements, with over one million copies distributed throughout the human genome. Alu elements are approximately 300 base pairs in length and are therefore classified as short interspersed nucleotide elements (SINEs), within the class of repetitive DNA elements. The typical structure of an Alu element is 5'-part A-A5TACA6-part B-polyA tail-3', where parts A and B are similar nucleotide sequences.

[0076] In a preferred embodiment, the Alu retrotransposon has a nucleotide sequence as set forth in SEQ ID NO:1 or a similar sequence, i.e., a sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 98% sequence identity thereto.

[0077] In one embodiment, the repeat sequence is a HERV element or any fragment or variant thereof. Preferably, the repeat sequence is a HERV-K element, which belongs to the HERV element subfamily. The terms "human endogenous retrovirus K" (HERV-K), "human teratocarcinoma-derived virus" (HDTV), or "human mouse mammary tumor virus-like-2" (HML-2) are used interchangeably and refer to a family of human endogenous retroviruses. HERV-K expression in humans is associated with various cancers. The human genome contains hundreds of copies of HERV-K. HERV-K elements are described, inter alia, in Agoni et al. (Front Oncol. 2013; 3: 180) and Garcia-Montojo M et al. (Crit Rev Microbiol. 2018 Nov; 44(6): 715-738).

[0078] In another embodiment, the repeat sequence is a LINE element, preferably a LINE-1 retrotransposon or any fragment or variant thereof. Preferably, the repeat sequence is a primate-specific copy of a LINE-1 retrotransposon (e.g., L1PA or L1HS).

[0079] The terms "LINE-1," "LINE1," or "L1" are used interchangeably herein and refer to the reverse-transcription transposon LINE-1 (also known as long spreading element-1 or long distribution element-1). LINE1 is a class I transposable element and belongs to the group of long interspersed repeat elements (LINEs). LINE-1 retrotransposons comprise approximately 17% of the human genome. A typical LINE-1 element is approximately 6,000 base pairs (bp) in length and consists of two non-overlapping open reading frames (ORFs) flanked by overlapping untranslated regions (UTRs) and target sites.

[0080] In embodiments, the LINE-1 retrotransposon has a nucleotide sequence as set forth in SEQ ID NO:2 or SEQ ID NO:29 or a similar sequence, i.e., a sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 98% sequence identity thereto.

[0081] In a preferred embodiment, the LINE-1 retrotransposon has a nucleotide sequence as set forth in SEQ ID NO:29 or a similar sequence, i.e., a sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 98% sequence identity thereto.

[0082] In an embodiment, the DNA repeat sequence of interest comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 60 or 90 CpG dinucleotides, preferably at least 30 or at least 90 CpG dinucleotides.

[0083] The methods of the present invention may include a step of pre-treating the DNA or a subsequence thereof.

[0084] In one embodiment, the DNA sequence of interest, for example, the cfDNA of interest, is treated to deaminate unmethylated cytosine before determining the CpG methylation profile of said DNA sequence of interest.This treatment of DNA can be used to deaminate unmethylated cytosine and produce uracil in DNA.In this way, when using specific probe and / or primer to sequence and / or amplify, the methylation status of DNA can be detected based on identifying the base change from cytosine to uracil or thymine.

[0085] Deamination of cytosine can be carried out by any method known in the art, for example, by treatment with a bisulfite reagent, e.g., sodium bisulfite, or an enzyme, e.g., Tet methylcytosine dioxygenase 2 (TET2), T4-phage beta-glucosyltransferase (T4-BGT), and apolipoprotein B mRNA editing enzyme catalytic subunit 3A (APOBEC3A) enzyme, any one or any combination of enzymes such as those described in Vaisvila et al., Genome Res. 2021. 31: 1280-1289.

[0086] Other methods for analyzing methylation status do not necessarily use cytosine deamination (e.g., TET-assisted pyridine borane sequencing (TAPS) and immunoprecipitation-based methods combined with deep sequencing (MeDIP)).

[0087] For example, DNA methylation-sensitive or DNA methylation-specific restriction enzymes that distinguish between methylated and unmethylated molecules can be used, and quantitative PCR (qPCR) and droplet digital PCR (ddPCR) methods that target methylated and unmethylated molecules can also be used.

[0088] Another method to distinguish cytosine methylation status without conversion is direct sequencing using Oxford Nanopore Technologies (ONT), which accurately detects 5mC changes even in plasma DNA (Cheng et al. Clin. Chem. 61, 1305-1306 (2015)).

[0089] Preferably, the DNA sequence of interest is treated with a bisulfite reagent, preferably sodium bisulfite.

[0090] Bisulfite reagents that can be used in the context of the present invention include, for example, bisulfite, disulfite, hydrogen sulfite, or any combination thereof, among others. These reagents may be useful for distinguishing between methylated and unmethylated nucleic acids. Bisulfite interacts differently with cytosine and 5-methylcytosine. In typical bisulfite-based methods, contacting DNA with bisulfite deaminates unmethylated cytosine to uracil, while leaving methylated cytosine unaffected, resulting in selective retention of methylated cytosine but not unmethylated cytosine. The same applies to EM-seq (enzymatic deamination of unmethylated cytosine). Thus, uracil or thymine residues are present in place of unmethylated cytosine residues, providing a discrimination signal for unmethylated cytosine residues, while the remaining (methylated) cytosine residues provide a discrimination signal for methylated cytosine residues. The processed samples can be analyzed, for example, by next-generation sequencing (NGS) or targeted bisulfite NGS / deep sequencing, specifically to identify cytosine to thymine base changes when uracil (U) is copied as thymine (T) during an amplification step (e.g., PCR amplification).

[0091] A variety of methylation assay techniques can be used in combination with bisulfite treatment to determine the methylation profile of a DNA sequence of interest. Such assays may include, for example, sequencing of bisulfite-treated nucleic acids, PCR (e.g., involving sequence-specific amplification), methylation-sensitive high-resolution melting (MS-HRM) PCR (see, e.g., Hussmann 2018 Methods Mol Biol. 1708:551-571), quantitative multiplex methylation-specific PCR (QM-MSP) (see, e.g., Fackler 2018 Methods Mol Biol. 1708:473-496), Methylation Specific Nuclease-assisted Minor-allele Enrichment (MS-NaME) (see, e.g., Liu 2017 Nucleic Acids Res. 45(6):e39), pyrosequencing, and Methylation-sensitive Single Nucleotide Primer Extension (Ms-SNuPE™) (see, e.g., Gonzalgo 2007 Nat Protoc. 2(8):1931-6), among others.

[0092] In certain embodiments, DNA sequences are processed by TET-assisted pyridine borane sequencing (TAPS). TAPS specifically converts only methylated cytosines, preserving DNA integrity and allowing for analysis of only a small fraction of DNA. This also allows for improved downstream analysis, as the resulting reads retain their full complexity. In certain embodiments, subsequences of DNA of interest are amplified from bisulfite-treated DNA samples. In certain embodiments, high-throughput and / or next-generation sequencing technologies are used to achieve base-pair level / scale resolution of DNA sequences and allow for analysis of methylation profiles.

[0093] In one embodiment, the target DNA sequence is not treated to deaminate unmethylated cytosine before determining the CpG methylation profile.Those skilled in the art are aware of techniques that can identify CpG methylation without requiring deamination treatment.Single molecule real-time (SMRT) sequencing theoretically provides the opportunity to directly evaluate specific base modifications of native DNA molecules using the kinetic signal of DNA polymerase without any prior chemical / enzymatic conversion and PCR amplification.For example, in Nanopore sequencing devices (e.g., Oxford Nanopore Technologies) and single molecule real-time (SMRT) sequencing (e.g., Pacific BioSciences PacBio), electrolytic current signals are highly sensitive to base modifications, such as 5-methylcytosine (5-mC), which allows the detection of native CpG methylation sites (Cheng et al., Clin. Chem.61, 1305-1306 (2015)).

[0094] In an embodiment, the method of the present invention comprises a step of amplifying the DNA sequences of interest or subsequences thereof, in particular, the DNA sequences or subsequences are amplified, for example, by polymerase chain reaction (PCR), preferably multiplex PCR, prior to clustering.

[0095] As used herein, the term "amplification" refers to the use of a template nucleic acid molecule, e.g., cfDNA, in combination with various reagents to generate additional nucleic acid molecules, e.g., "amplicons," from the template nucleic acid molecule. The additional nucleic acid molecules can be identical or similar (e.g., at least 80% identical, e.g., at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical) to the template nucleic acid molecule, its complementary sequence, and / or a segment thereof.

[0096] The process of DNA amplification can be carried out by any method known to those skilled in the art, for example, polymerase chain reaction (PCR). Preferably, the method for determining the CpG methylation profile is a PCR-based method.

[0097] Thus, in an embodiment, the subsequence of DNA of interest is an amplicon.

[0098] A DNA subsequence can be amplified, for example, using a specific primer pair complementary to the DNA sequence of interest, e.g., LINE-1, etc. A person skilled in the art knows how to design suitable primer pairs, for example, using alignment tools.

[0099] In embodiments, the amplicon is selected from the group consisting of amplicon #1, #2, #3, #4, #5, #6, #7, or #8, and any combination thereof, such as those set forth in Figure 1 herein and / or as set forth in SEQ ID NOs: 3, 4, 5, 6, 7, 8, 9, and 10, respectively. In particular, the method includes studying amplicon #1, #2, #4, #5, #6, #7, and #8, and optionally, amplicon #3.

[0100] Preferably, the amplicon studied in the method of the present invention comprises or consists essentially of an amplicon having a sequence selected from the group consisting of SEQ ID NOs: 3, 4, 5, 6, 7, 8, 9, 10 and any sequence having at least 85, 90, 95, 98 or 99% sequence identity thereto.

[0101] Alternatively, DNA subsequences can be amplified with universal or degenerate primers.

[0102] The primers used to amplify the DNA subsequence may include an adapter, such as a unique molecular identifier (UMI) or a unique dual index (UDI).

[0103] UMIs can be of any length suitable to produce a sufficiently large number of unique UMIs. In certain aspects, UMIs can be 5-20 nucleotides in length. Thus, each UMI can be about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length. In one embodiment, the UMI is a nucleotide sequence that is 16 nucleotides in length.

[0104] In an embodiment, the primer contains no CpG sites, up to two CpG sites, is methylation independent, and preferably contains no CpG sites towards the 5' end of the primer.

[0105] In this embodiment, the DNA sequence targeted by the primers is in the range of 100 to 200 bp, preferably 101 to 150 bp.

[0106] In an embodiment, the method further comprises removing PCR replicates after PCR amplification of the DNA sequence or subsequence of interest. A common practice for removing PCR duplicates is to remove all but one read of the same sequence, assuming that such reads were generated from the same DNA molecule by PCR. Preferably, the method of the present invention relies on the use of primers containing unique molecular identifiers (UMIs) to accurately detect PCR duplicates. The primers may further contain common or universal sequences CS1 and / or CS2. The common sequences may consist of or comprise, for example, common sequence 1 (CS1) (5'-ACACTGACGACATGGTTCTACA-3' SEQ ID NO:27) and / or common sequence 2 (CS2) (5'-TACGGTAGCAGAGACTTGGTCT-3' SEQ ID NO:28) universal primer sequences.

[0107] In an embodiment, the method of the present invention preferably comprises a step of capturing a DNA sequence, preferably a cfDNA sequence. In particular, the DNA sequence or DNA subsequence is captured before amplification.

[0108] In certain embodiments, one or more probes are used to capture a DNA sequence or subsequence of interest.

[0109] In an embodiment, the method comprises using at least 20, 50, 100, 200, 250 or 300 probes, particularly 200-250 probes, preferably 210-230 probes, that target different regions of the DNA sequence of interest.

[0110] In a very particular embodiment, the method comprises using probes of 100-150 bp, preferably 120 bp, starting every 20-30 bp, preferably every 24 bp.

[0111] In another embodiment, the method comprises using 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 probes that target different regions of the DNA sequence of interest.

[0112] In one embodiment, the methods of the invention involve a capture-based approach (NGS Methylation Detection System) such as that developed by TWIST Bioscience for DNA methylation analysis.

[0113] In certain embodiments, the method includes screening / capturing, aligning, and / or clustering DNA sequences and / or subsequences thereof. Such steps can be performed using DNA sequencing by any method known in the art, such as massively parallel sequencing (e.g., next-generation sequencing (NGS)), sequencing-by-synthesis, real-time (e.g., single-molecule) sequencing, bead emulsion sequencing, or nanopore sequencing. Quantitative polymerase chain reaction (qPCR) (e.g., methylation-sensitive restriction enzyme quantitative polymerase chain reaction or MSRE-qPCR) can also be used.

[0114] Preferably, the clustering of DNA subsequences is based on DNA nucleotide sequence similarity or identity and / or similar sequence length.

[0115] In some embodiments, DNA subsequence clustering is performed using algorithmic support, particularly vsearch. If the pairwise identity with the centroid is greater than 0, the subsequence is added to the cluster. Pairwise identity is defined as the number of (matching sequences) / (alignment length).

[0116] In one embodiment, the total number of clusters is dynamically selected, resulting in DNA subsequences being clustered into n clusters, where n is dynamically defined according to the size of the cluster, and a cluster is defined as a reference sequence that accounts for at least 15% of all sequences.

[0117] In certain embodiments, the DNA subsequences are clustered into at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 clusters, such as 5-15 clusters or 5-10 clusters, preferably 10 clusters. The number of clusters considered may be based, inter alia, on the proportion of total reads represented by a given cluster.

[0118] In certain embodiments, the DNA subsequences are clustered based on sequence identity, and the resulting clusters include subsequences that share at least about 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity.

[0119] The comparison of sequences and determination of percent identity between two sequences can be accomplished using any method known in the art and / or may be by computational algorithms, such as BLAST (basic local alignment search tool).

[0120] The method of the invention preferably comprises the step of selecting a reference sequence for (and in) each cluster, which can be one or more sequences selected from the cluster.

[0121] Preferably, the reference sequence is the most representative or representative sequence of the cluster under consideration, meaning the sequence representing the largest number of sequences having at least 60%, at least 70% or at least 80% sequence identity with the consensus size / sequence found within the cluster.

[0122] In particular, the reference sequence is not a reference genome, and thus the method involves aligning sequencing data from repetitive sequences without using a reference genome.

[0123] Preferably, the reference sequence is the sequence with the longest nucleic acid length in the cluster. In an embodiment where the method is a PCR-based method, the longest nucleic acid length is the length of the insert sequence between the forward primer and the reverse primer, optionally with an error of + / - 15 nucleic acids, preferably + / - 10 nucleic acids, more preferably + / - 5 nucleic acids.

[0124] In a preferred embodiment, the total number of clusters is dynamically selected so that the DNA subsequences are clustered into n clusters, where n is dynamically defined depending on the size of the clusters, and the reference sequence is the centroid of the n largest clusters (i.e., the cluster containing the largest number of sequences). The centroid of a sequence is the central sequence that minimizes the sum of the distances to all sequences in the cluster.

[0125] Preferably, the method comprises selecting the largest clusters (ie the clusters that contain more DNA sequences compared to the total number of DNA sequences).

[0126] In particular, such a cluster comprises a number of DNA subsequences representing a minimum of at least 10%, 15%, 20%, 25% or 30% of the total number of DNA subsequences.

[0127] In the context of the methods described herein, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 reference sequences are identified for each cluster.

[0128] Preferably, 5 to 20, 5 to 15 or 5 to 10 reference sequences, 1 to 20, 1 to 15 or 1 to 10 reference sequences are defined for each cluster, more preferably 10 reference sequences are selected / selected / identified / determined.

[0129] In an embodiment, one reference sequence is defined per cluster, and a total of 10 sequences from the 10 largest clusters are used as reference sequences.

[0130] The method of the invention preferably comprises the step of aligning the(respective) reference sequence(s) of the(respective) cluster(s) by aligning at the positions of the CpG dinucleotides.

[0131] In this way, the reference sequences are aligned to obtain a pool of reference sequences, which constitutes a reference for aligning all subsequences.

[0132] Preferably, the alignment of the remaining subsequences on the selected reference sequence is based on the score of the most favorable of the G / G alignment, followed by the T / T, C / C, C / T and T / C alignments, and finally the alignment of the remaining nucleotides of the sequence. This score can be calculated, for example, using the mafft program (preferably with the following parameters: --textmatrix<custom_score_matrix.txt> --retree 2). Reference sequences are preferably pairwise aligned using a custom scoring matrix to prioritize alignment of the dinucleotides CG / TG over other possible dinucleotide combinations (e.g., AG / AC / TC).

[0133] This step aims to identify the location of CpG dinucleotides in a reference sequence, in particular, alignment of DNA sequences or subsequences thereof from healthy subjects to identify the location of CpG dinucleotides.

[0134] The method preferably includes aligning the remaining subsequences of the cluster to the selected reference sequence.

[0135] In a particular embodiment, the method comprises a step of aligning all subsequences, in particular all amplified subsequences, to a reference sequence, which allows for an additional check on the quality of the alignment of the DNA subsequences.

[0136] In particular, reads are aligned to the identified reference sequence using an algorithm with a time complexity of O(n), where n is the number of reads to align. This allows for the alignment of millions of reads. An algorithm is said to have linear time complexity when its execution time increases linearly with the length of the input.

[0137] In certain optional embodiments, the method comprises an additional step of checking the presence of methylation-prone (enough) CpG dinucleotide sites in each reference sequence. This step is carried out either before or after, preferably after, the step of aligning the reference sequences of the cluster. In particular, at least 20% of CG+TG dinucleotides within the dinucleotide sites are selected, and these sites preferably have at least 5% CG dinucleotides (especially after bisulfite sequencing, thus representing methylated CG). This step may be referred to herein as "CG calling." Preferably, the method comprises a step of identifying methylated or methylation-prone CpG sites.

[0138] Preferably, the reference sequence comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 60 or 90 CpG dinucleotides, preferably at least 30 or 90 CpG dinucleotides.

[0139] The determination of the presence of methylation-prone CpG dinucleotide sites in each reference sequence is preferably carried out between steps d) and e) of the above-mentioned method or during step e) of the above-mentioned method.

[0140] Alternatively, the determination of the presence of methylation-prone CpG dinucleotide sites in each reference sequence can be carried out between steps b) and c) of the above-mentioned method.

[0141] In certain embodiments, the determination of methylation-prone CpG dinucleotide sites is performed on DNA from a healthy subject or a population of healthy subjects, in particular to avoid biases associated with hypomethylation in cancer.

[0142] Once the reference sequences are aligned, it is possible to easily identify CpG sites and check (to confirm) whether a particular CpG site may be methylated by comparing all aligned subsequences. In particular, the dinucleotides CG and TG are identified in each subsequence.

[0143] At a particular dinucleotide position, the proportion of the dinucleotide CG and the dinucleotide TG, respectively, at that particular dinucleotide position in the set of subsequences is calculated (the number of CGs and the number of TGs in all subsequences).

[0144] For a particular dinucleotide position in the sequence, a CG percentage of greater than 5% combined with a CG+TG percentage of greater than 20% indicates a CpG site that is particularly susceptible to methylation after bisulfite sequencing, and thus indicates a methylated CG.

[0145] A TG ratio of more than 95% indicates a dinucleotide site that is poorly methylated, especially one that is repeatedly poorly methylated.

[0146] A CG+TG ratio of less than 20% indicates a dinucleotide site that is unlikely to be a repeating template suitable for methylation.

[0147] The method of the present invention preferably includes a step of determining the CpG methylation status of each subsequence by determining whether the CpG dinucleotide is methylated or not at each CpG site of the subsequence, thereby determining a CpG methylation profile comprising the CpG methylation level and / or proportion of CpG methylated haplotypes of the subsequence.

[0148] Methylation status and methylation profiles can be assessed by various methods known in the art and / or provided herein. Methods for measuring methylation status can be, for example, but are not limited to, whole genome sequencing, targeted enzymatic methylation sequencing, methylation status-specific polymerase chain reaction (PCR), mass spectroscopy, methylation arrays, methylation-specific nucleases, mass-based separation, target-specific capture, and / or methylation-specific oligonucleotide primers. Certain methods for assessing methylation utilize bisulfite reagents (e.g., sodium bisulfite) or enzymatic conversion reagents (e.g., Tet methylcytosine dioxygenase 2).

[0149] As used herein, "methylation status" or "methylation state" refers to the fact that a CpG dinucleotide is methylated or not.

[0150] As used herein, "methylation profile" refers to the number, frequency, or pattern of methylation at CpG methylation sites in a sequence of interest, particularly a sequence of interest in a DNA sequence. Thus, a change in the methylation profile between a first state and a second state may be or include an increase in the number, frequency, or pattern of methylated CpG sites, or may be or include a decrease in the number, frequency, or pattern of methylated sites. In various examples, the change in methylation state is a change in methylation level and / or methylation pattern.

[0151] Preferably, the CpG methylation profile comprises the CpG methylation level of the subsequence, in particular the methylation level at each CpG site and / or the proportion of CpG methylation haplotypes.

[0152] As used herein, the term "methylation level" or "methylation value" refers to a numerical representation of methylation status, e.g., a number representing the frequency or ratio of methylation of CpG sites in a subset of sequences of interest. In certain embodiments, the methylation level is the proportion of CpGs that are methylated in the subset of DNA sequences. This means that, for a particular CpG dinucleotide in a cluster of subsequences, the methylation status is determined for such position in each subsequence of the cluster. The ratio can then be established as follows: (number of methylated CpG dinucleotides at a particular dinucleotide position) / (total number of subsequences in the cluster). Preferably, the methylation level at a particular CpG site in a cluster of subsequences is the ratio of CG dinucleotides to the number of CG+TG+TA dinucleotides (CG / CG+TG / CG+TG+TA).

[0153] As used interchangeably herein, the terms "methylation pattern," "methylation motif," "methylation signature," or "methylation haplotype" refer to a numerical representation of a "methylation profile," e.g., the number of unique sequences of strings / strands of methylated cytosines at single molecule resolution, thereby generating unique motifs. This refers, in particular, to the methylation status of consecutive CpG sites within each sequence (or subsequence or amplicon).

[0154] In certain embodiments, if a cytosine is methylated, the attribute number is "1." Meanwhile, an unmethylated cytosine has the attribute number "0." Thus, in a sequence of interest, a run of methylated or unmethylated cytosines results in a run of attribute numbers "1" or "0." This alternating sequence of 0s and 1s provides a particular methylation pattern for the DNA subsequence under study.

[0155] "Methylation haplotype proportion" refers to the proportion or percentage of a particular methylation pattern in a population of sequences, particularly a cluster of subsequences.

[0156] For example, if a sequence of interest contains two CpG sites, the methylation haplotypes are 0-0 (both CpG sites are unmethylated), 1-1 (both sites are methylated), 1-0 (the first CpG site is methylated), or 0-1 (the second CpG site is methylated). The proportions of methylation haplotypes in a cluster of subsequences are the proportions of haplotypes 0-0, 1-1, 1-0, and 0-1 in all subsequences of the cluster.

[0157] In certain embodiments, some or each of steps a) through f) of the methods described herein are performed on another set of subsequences of interest, which may be completely distinct / different or at least partially distinct from the set originally (or any previously) used.

[0158] In certain embodiments and for a particular DNA sequence, one subsequence or at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30 or 35 subsequences can be analyzed for its / their methylation profile, separately or simultaneously.

[0159] These partial sequences may be overlapping partial sequences in the target DNA. In particular, when aligned to the complete DNA sequence, the partial sequences may overlap by 5 to 100 bp, particularly 5 to 25 bp, at the 3' or 5' ends.

[0160] Preferably, the subsequence can be defined to cover the entire repeat sequence of interest. Alternatively, the subsequence can be defined to cover only a portion of the repeat sequence of interest. Preferably, the subsequence includes a DMR, particularly a region that is differentially methylated in cancer compared to healthy conditions.

[0161] In embodiments, the repeat sequence is LINE-1 and the method includes at least 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 200, 250, or 300 subsequences of interest.

[0162] In an embodiment, the repetitive sequence is LINE-1, the method includes 5 to 10 subsequences of interest, more preferably 7 to 9 subsequences of interest, and even more preferably 8 subsequences of interest, and the method for determining the CpG methylation profile is performed for each of the subsequences, for example, for 8 subsequences.

[0163] In an embodiment, the repetitive sequence is LINE-1, the method includes 100 to 300 subsequences of interest, more preferably 200 to 300 subsequences of interest, and even more preferably 250 subsequences of interest, and the method of determining the CpG methylation profile is performed for each of the subsequences, for example, for 250 subsequences.

[0164] Preferably, the subsequence of interest comprises or consists essentially of a sequence selected from SEQ ID NOs: 3, 4, 5, 6, 7, 8, 9, 10, and any sequence having at least 85, 90, 95, 98, or 99% sequence identity thereto.

[0165] Each subsequence was synthesized using a primer pair, specifically: i) a forward primer having the sequence shown in SEQ ID NO: 11 and a reverse primer having the sequence shown in SEQ ID NO: 12, which in particular preferably targets amplicon #1 as set forth in SEQ ID NO: 3; ii) a forward primer having the sequence shown in SEQ ID NO: 13 and a reverse primer having the sequence shown in SEQ ID NO: 14, which in particular preferably targets amplicon #2 as set forth in SEQ ID NO: 4; iii) a forward primer having the sequence shown in SEQ ID NO: 15 and a reverse primer having the sequence shown in SEQ ID NO: 16, particularly targeting amplicon #3 as set forth in SEQ ID NO: 5; iv) a forward primer having the sequence shown in SEQ ID NO: 17 and a reverse primer having the sequence shown in SEQ ID NO: 18, which in particular preferably targets amplicon #4 as set forth in SEQ ID NO: 6; v) a forward primer having the sequence shown in SEQ ID NO: 19 and a reverse primer having the sequence shown in SEQ ID NO: 20, which in particular preferably targets amplicon #5 as set forth in SEQ ID NO: 7; vi) a forward primer having the sequence shown in SEQ ID NO: 21 and a reverse primer having the sequence shown in SEQ ID NO: 22, particularly targeting amplicon #6 as set forth in SEQ ID NO: 8; vii) a forward primer having the sequence shown in SEQ ID NO:23 and a reverse primer having the sequence shown in SEQ ID NO:24, particularly targeting amplicon #7 as set forth in SEQ ID NO:9; viii) In particular, a forward primer having the sequence shown in SEQ ID NO:25 and a reverse primer having the sequence shown in SEQ ID NO:26, which preferably targets amplicon #8 as set forth in SEQ ID NO:10. and any combination thereof, particularly, it can be amplified with a primer pair consisting of these.

[0166] The primers described herein may include an adapter, such as a unique molecular identifier (UMI) or a unique dual index (UDI). The primers may further include consensus or universal sequences CS1 and / or CS2. The consensus sequences may be, for example, consensus sequence 1 (CS1) (5'-ACACTGACGACATGGTTCTACA-3', SEQ ID NO:27) and / or consensus sequence 2 (CS2) (5'-TACGGTAGCAGAGACTTGGTCT-3', SEQ ID NO:28) universal primer sequences.

[0167] In a particular embodiment, the method comprises a step of identifying the amplicons prior to the clustering step. Such identification can be carried out in particular by the sequences of the primers and probes used.

[0168] In certain embodiments, the method comprises repeating some or each of steps a) to e) once or several times, optionally with the additional step of comparing the CpG methylation profiles to provide an optimized CpG methylation profile.

[0169] Determining the CpG methylation profile of a DNA sequence of interest provides insight into the state of that DNA sequence (e.g., healthy or cancerous) and therefore the state of the subject from which the DNA sequence is derived.

[0170] In certain preferred embodiments, the present invention makes it possible to distinguish a "healthy CpG methylation profile" from a "cancerous CpG methylation profile".

[0171] As used herein, the term "healthy CpG methylation profile" refers to a CpG methylation profile that correlates with or is indicative of a healthy subject, particularly a subject that is not suffering from cancer.

[0172] As used herein, the term "cancerous CpG methylation profile" refers to a CpG methylation profile that correlates with cancer or indicates that a subject is suffering from cancer.

[0173] In some embodiments, the term "cancerous CpG methylation profile" refers to: - a CpG methylation profile that correlates with a particular cancer origin / type or indicates that the subject has a particular cancer origin / type (cancer origin is, for example, breast cancer, colon cancer or brain cancer); - a CpG methylation profile that correlates with a particular stage of cancer or indicates that a subject has a particular stage of cancer (e.g., the stage of cancer is stage I, II, or III); and / or - a CpG methylation profile that correlates with cancer metastasis or indicates that a subject is suffering from cancer metastasis Refers to...

[0174] In some embodiments, the term "cancerous CpG methylation profile" refers to: - a CpG methylation profile that correlates with a particular cancer origin / type and metastasis or indicates that a subject suffers from a particular cancer origin / type and metastasis (cancer origin is, for example, breast cancer, colon cancer or brain cancer); - a CpG methylation profile that correlates with a particular stage and metastasis of cancer or indicates that a subject has a particular stage of cancer (e.g., stage I, II, or III) and metastasis; or - CpG methylation profiles that correlate with a particular cancer origin / type and stage or indicate that a subject has cancer of a particular origin / type and stage. Refers to...

[0175] In certain preferred aspects, the methods described herein are partially or fully computer-implemented methods.

[0176] training method In another aspect, the present invention relates to methods, in particular computer-implemented methods, for training a classifier to accurately distinguish between healthy and cancerous CpG methylation profiles or to determine the health status of a subject, in particular to accurately distinguish between healthy and cancerous subjects or between different types of cancer. These training methods rely in particular on determining the CpG methylation profile of a DNA sequence of interest or a subsequence thereof, and preferably use the methods for determining the CpG methylation profile described herein, for example in the paragraph "Determining the CpG methylation profile" appearing herein above. Thus, the aspects and embodiments described herein above with respect to determining the CpG profile can be applied to any of the training methods described herein below.

[0177] Examples of cancers that may be mentioned in the methods of the present invention are described herein, particularly in the section "Subjects and Biological Samples."

[0178] As used herein, the term "classifier" refers to a computer-implemented algorithm that performs classification, i.e., determines a likelihood score or probability that an object will be classified within one group of objects (e.g., a group of healthy CpG methylation profiles) as opposed to one or more other groups of objects (e.g., a group of cancerous CpG methylation profiles), and maps the input objects to categories (e.g., healthy CpG methylation profiles or cancerous CpG methylation profiles). The term "classifier" may refer to one or more classifiers. For example, multiple classifiers can be trained and data can be processed by these classifiers in parallel and / or as a pipeline. For example, the output of one type of classifier (e.g., the output from an intermediate layer of a neural network) may be provided as input to another type of classifier.

[0179] Examples of classifiers that can be used in the context of the present invention include, but are not limited to, artificial neural networks of various architectures (e.g., deep, convolutional, fully connected) and supervised machine learning classifiers, such as support vector machine (SVM) classifiers, random forest classifiers, decision tree classifiers, K-nearest neighbor classifiers (KNN), logistic regression classifiers, nearest neighbor classifiers, Gaussian mixture models (GMM), and nearest centroid classifiers. This is not an exhaustive list, and those skilled in the art will be able to identify similar algorithms that can be used as well, even though they are not specifically mentioned here. The details and rules of the functions of the mentioned algorithms have already been widely described in the literature. The key contribution is the set of input data provided to the classifier (i.e., a set of CpG methylation profiles or preprocessed information obtained from a set of CpG methylation profiles). Based on this input data, any suitable supervised machine learning technique can be used to generate a suitable model. Therefore, the selection of a suitable algorithm is secondary in nature and can be made in many different ways and in various combinations, as will be apparent to those skilled in the art.

[0180] Preferably, the classifier is selected from a support vector machine (SVM) classifier, a random forest (RF) classifier, a decision tree classifier, a K-nearest neighbor classifier (KNN), a logistic regression classifier, a nearest neighbor classifier, a Gaussian mixture model (GMM), a nearest centroid classifier, and an artificial neural network, such as a deep neural network, a convolutional neural network, or a fully connected neural network. More preferably, the classifier is selected from a support vector machine (SVM) classifier, a random forest (RF) classifier, and a neural network, in particular a convolutional neural network (CNN). Even more preferably, the classifier is a random forest classifier.

[0181] A classifier utilizes some training data to understand whether a given input object belongs to one category / class or another. The classifier may be provided with a training set of biological samples from subjects, e.g., healthy subjects and / or cancer subjects, which training set includes DNA sequences, particularly cfDNA sequences, that exhibit healthy or cancerous CpG methylation profiles. Alternatively, the classifier may be provided with preprocessed information obtained from such a training set of DNA sequences.

[0182] In an aspect, the present invention provides a method, typically a computer-implemented method, for training a classifier to accurately distinguish between healthy and cancerous CpG methylation profiles, comprising: a) providing a training set of CpG methylation profiles of DNA sequences of interest or subsequences thereof, wherein the DNA sequences of interest are repeated and distributed throughout the genome and contain a high density of CpG dinucleotides, or preprocessed information obtained from the training set of CpG methylation profiles of DNA sequences of interest or subsequences thereof, as input to a classifier, wherein the training set of CpG methylation profiles includes CpG methylation profiles of DNA sequences or subsequences thereof from subjects identified as healthy subjects and subjects identified as cancer subjects; b) generating an output of the classifier for each CpG methylation profile input of the DNA sequence of interest or a subsequence thereof, wherein the output classifies the CpG methylation profile input of the DNA sequence of interest or a subsequence thereof as a healthy CpG methylation profile or a cancerous CpG methylation profile; c) To evaluate the performance of the classifier to distinguish between healthy and cancerous CpG methylation profiles; Including, The present invention relates to a method wherein the CpG methylation profile comprises the CpG methylation level and / or the proportion of CpG methylated haplotypes of a DNA sequence or a subsequence thereof.

[0183] The present invention also provides a method, typically a computer-implemented method, for training a classifier to determine the health status of a subject, in particular to accurately distinguish between healthy subjects and subjects suffering from cancer, the method comprising: a) providing a training set of CpG methylation profiles of DNA sequences of interest or subsequences thereof, wherein the DNA sequences of interest are repeated and distributed throughout the genome and contain a high density of CpG dinucleotides, or preprocessed information obtained from the training set of CpG methylation profiles of DNA sequences of interest or subsequences thereof, as input to a classifier, wherein the training set of CpG methylation profiles includes CpG methylation profiles of DNA sequences or subsequences thereof from subjects identified as healthy subjects and subjects identified as cancer subjects; b) generating an output of the classifier for each CpG methylation profile input of the DNA sequence of interest or a subsequence thereof, wherein the output classifies the CpG methylation profile input of the DNA sequence of interest or a subsequence thereof as a healthy CpG methylation profile or a cancerous CpG methylation profile; c) classifying subjects with a healthy CpG methylation profile as healthy subjects and subjects with a cancerous CpG methylation profile as subjects suffering from cancer; d) assessing the performance of the classifier to distinguish between healthy subjects and subjects with cancer; Including, The present invention relates to a method wherein the CpG methylation profile comprises the CpG methylation level and / or the proportion of CpG methylated haplotypes of a DNA sequence or a subsequence thereof.

[0184] Step c) or step d) of the training method described herein above relates to assessing the performance of the classifier to distinguish between i) healthy CpG methylation profiles or subjects and ii) cancerous CpG methylation profiles or subjects.

[0185] In certain preferred embodiments, the CpG methylation profile of a DNA sequence of interest or a subsequence thereof is determined using the methods for determining a CpG methylation profile described herein, in particular in the section "Determining a CpG methylation profile".

[0186] In one embodiment, the method for determining the CpG methylation profile comprises aligning the DNA sequences of interest of a subset of healthy subjects to identify the locations of CpG dinucleotides that are prone to methylation.Preferably, highly methylated CpG sites are determined.

[0187] Alternatively or additionally, such methods may further comprise aligning the DNA sequences of interest of the subset from cancer subjects to identify the locations of CpG dinucleotides that are prone to methylation. Preferably, unmethylated CpG sites are determined.

[0188] In a particular embodiment, the evaluation of the performance of a classifier in the context of the present invention is based on the classification of the CpG methylation profile of a DNA sequence or subsequence thereof of interest using a test set comprising CpG methylation profiles of DNA sequences or subsequences thereof from healthy and cancer subjects, said test set being distinct from the training set, wherein the healthy or cancer state of each subject in the test and training sets is known, and wherein the CpG methylation profiles of the DNA sequences or subsequences thereof of the test set are obtained and processed using the same methods used to obtain and process the CpG methylation profiles of the DNA sequences or subsequences thereof using the training set.

[0189] Biological and clinical data from healthy individuals and cancer patients can be easily retrieved from clinical trials. Collaborative research, such as the NIH Collaboratory Distributed Research Network, provides mediated or shared access to clinical data repositories by qualified researchers. Additionally, the UW De-identified Clinical Data Repository (DCDR) and the Stanford Center for Clinical Informatics enable initial cohort identification. The National Program of Cancer Registries (CDC) provides support for states and localities to maintain registries that provide high-quality data. Data collected by local cancer registries enables public health professionals to more effectively understand and address the burden of cancer. Clinical data is also published by the European Medicines Agency (EMA).

[0190] The performance of a classifier can be evaluated using any method known to those skilled in the art, for example, the performance of a classifier can be evaluated by precision, recall, or F1 score.

[0191] This classifier is considered to be a good enough classifier to distinguish between healthy and cancerous states, provided that false positives are minimized.

[0192] To improve the performance of the classifier, the training method, i.e., steps a) to c) or steps a) to d) (depending on the method considered herein above), can be repeated with some modifications, for example, by increasing the number of healthy and cancerous CpG methylation profiles in the training set of DNA sequences, using separate training sets of DNA sequences. Another possibility to improve the performance of the classifier is to increase the number of sets of cfDNA subsequences of interest used.

[0193] In certain embodiments, the performance of a classifier is the accuracy of the classifier. The accuracy of a classifier is a measure of the correct predictions of the classifier compared to its total data points. It is particularly the ratio of correct prediction units to the total number of predictions made by the classifier. It is preferably the proportion of correct classifications using either an independent test set or cross-validation.

[0194] In certain embodiments, the performance of a classifier is expressed / calculated as an AUROC (area under the receiver operating characteristic curve), AUC (area under the curve), or ROC (receiver operating characteristic) curve. The ROC curve can be used to select a classifier threshold that maximizes true positives and then minimizes false positives. It is hypothesized that the higher the AUC, the better the model's performance in distinguishing between positive and negative classes.

[0195] For example, to evaluate the performance of a classifier, the true positive rate and false positive rate are evaluated in each run by interpolating to generate all points on the ROC curve. At the end of the run, an average ROC curve is generated and a 95% confidence interval is calculated based on the results of all runs. In a multi-class context, an ROC curve for each class can be generated by taking the class under consideration as the positive class.

[0196] Preferably, the CpG methylation profile of the DNA sequence of interest or a subsequence thereof is determined using the methods for determining a CpG methylation profile disclosed above.

[0197] Test Method In another particular aspect, the present invention relates to a method, particularly a computer-implemented method, for determining the health status of a subject, for determining whether a subject is healthy or suffering from cancer, for determining the origin or stage of a tumor, or for monitoring the effect / response of an anti-cancer treatment / drug, for evaluating the efficacy of a compound for reverting a cancerous CpG methylation profile of a cfDNA sequence of interest from a subject suffering from cancer to a healthy CpG methylation profile, or for predicting the ability of a compound to treat cancer. These testing methods rely, in particular, on determining the CpG methylation profile of a cfDNA sequence of interest or a subsequence thereof using the method for determining the CpG methylation profile disclosed herein above, particularly in the section "Determining the CpG methylation profile." Thus, each and every aspect and embodiment described in connection with determining the CpG profile can be applied to any of the testing methods described herein below.

[0198] Examples of cancers contemplated by the test methods of the present invention are described below, particularly in the section "Subjects and Biological Samples," and apply to any of the test methods described below.

[0199] Preferably, in the testing methods described herein, the classifier is trained using the training methods described herein, in particular the training methods described herein above in the "Training Method" section. As such, each and every aspect and embodiment described in relation to the training methods can be applied to any of the testing methods described herein below.

[0200] In a particular embodiment, the present invention provides an in vitro or in silico method for determining the health status of a subject, in particular for determining whether a subject is a healthy subject or a subject suffering from cancer or cancer recurrence, comprising: i) providing a DNA sequence of interest or a subsequence thereof from a subject, or preprocessed information obtained from said DNA sequence or subsequence, said DNA sequence of interest being a DNA sequence that is distributed throughout the subject's genome and encodes a repeat sequence containing a high density of CpG dinucleotides, as input to a classifier trained to distinguish between healthy and cancerous CpG methylation profiles; ii) using the classifier to identify the CpG methylation profile of the subject's DNA sequence of interest or a subsequence thereof as an output of the classifier as a healthy CpG methylation profile or a cancerous CpG methylation profile.

[0201] In particular, determining the health status of a subject includes identifying the CpG methylation profile of a DNA sequence of interest or a subsequence thereof from the subject as a healthy CpG methylation profile or a cancerous CpG methylation profile, wherein the number of CpG methylation profiles classified as cancerous CpG methylation profiles exceeding the number of CpG methylation profiles classified as healthy CpG methylation profiles indicates that the subject is suffering from cancer and / or the number of CpG methylation profiles classified as healthy CpG methylation profiles exceeding the number of CpG methylation profiles classified as cancerous CpG methylation profiles indicates that the subject is not suffering from cancer.

[0202] In another embodiment, the number of CpG methylation profiles classified as healthy CpG methylation profiles being below a statistically significant threshold indicates that the subject has cancer and / or the number of CpG methylation profiles classified as healthy CpG methylation profiles being equal to or greater than a statistically significant threshold indicates that the subject does not have cancer.

[0203] Alternatively, a number of CpG methylation profiles classified as cancerous CpG methylation profiles above the predicted level threshold indicates that the subject has cancer, and / or a number of CpG methylation profiles classified as cancerous CpG methylation profiles below the predicted level threshold indicates that the subject does not have cancer.

[0204] In embodiments, the prediction level threshold is at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, or 60%. Preferably, the prediction level threshold is at least 50%. In particular, when the algorithm is a random forest, the prediction level threshold is at least 50%, e.g., if at least 50% of the decision trees classify the methylation profile as "cancerous", the methylation profile is considered (classified) as "cancerous".

[0205] Thus, in certain embodiments, if the CpG methylation profile (e.g., methylation level and / or methylation haplotype) of a DNA sequence of interest or a subsequence thereof for a particular subject is considered "healthy" in at least 51% of cases (i.e., the number of runs performed on the statistical model, e.g., at least 1000, 5000, or 10000 runs), the subject is identified as being cancer-free.

[0206] In another embodiment, a subject is identified as having cancer if the CpG methylation profile of a DNA sequence of interest or a subsequence thereof from the subject is deemed "cancerous" in at least 50% of cases (i.e., the number of runs performed in the statistical model).

[0207] In certain embodiments, the prediction level threshold is dynamically determined or set. Varying the prediction score threshold used for classification allows for emphasis on sensitivity. This also allows for more fine-tuning of each individual model by selecting a threshold different from the default 0.5 typically used for classification in order to accurately classify more samples of a given class A without necessarily increasing the misclassification rate of samples of a given class B.

[0208] The prediction level threshold can be easily set by one skilled in the art depending on the data set, cancer type, or response type (eg, healthy subjects vs. cancer subjects, tumor origin, etc.).

[0209] In certain embodiments, the method of determining the health status of a subject comprises determining the presence of a primary tumor, preferably an early stage tumor (particularly a stage I, II or III primary tumor), the presence of cancer recurrence or the presence of metastasis in the subject if the classifier identifies the CpG methylation profile as a cancerous CpG methylation profile.

[0210] The methods of the present invention are particularly useful for early diagnosis of cancer or early stage (ie, stage I or II or III) cancer.

[0211] Thus, in certain aspects, the present invention provides an in vitro or in silico method for determining the stage of a tumor from a subject, comprising: (a) providing a DNA sequence of interest or a subsequence thereof from a subject, or preprocessed information obtained from said DNA sequence or subsequence, said DNA sequence of interest being a DNA sequence that encodes a repetitive sequence containing a high density of CpG dinucleotides distributed throughout the subject's genome, as input to a classifier trained to distinguish between healthy and cancerous CpG methylation profiles; (b) using the classifier, i) identifying a CpG methylation profile of a DNA sequence of interest or a subsequence thereof from said subject as a healthy CpG methylation profile or a cancerous CpG methylation profile; ii) comparing the cancerous CpG methylation profile of the DNA sequence of interest or a subsequence thereof from the subject with the CpG methylation profiles of tumors of different known stages, in particular the CpG methylation profiles of tumors from a population of cancer subjects with tumors of different known stages; iii) correlating the cancerous CpG methylation profile with a particular stage to determine the stage of the tumor from the subject as an output of the classifier.

[0212] The present invention also provides an in vitro or in silico method for determining the origin of a tumor in a subject, comprising: (a) providing a DNA sequence of interest or a subsequence thereof from a subject, or preprocessed information obtained from said DNA sequence or subsequence, said DNA sequence of interest being a DNA sequence that is distributed throughout the subject's genome and encodes a repetitive sequence containing a high density of CpG dinucleotides, as input to a classifier trained to distinguish between healthy CpG methylation profiles and cancerous CpG methylation profiles from various tumor origins; (b) using the classifier to identify a CpG methylation profile of the subject's DNA sequence of interest or a subsequence thereof as an output of the classifier as a healthy CpG methylation profile or a cancerous CpG methylation profile from a particular tumor origin, thereby determining the origin of a tumor from the subject.

[0213] In embodiments, the tumor origin is the origin of a metastasis.

[0214] In another aspect, the methods of the present invention can be advantageously used for the early detection of cancer recurrence.

[0215] In addition to or instead of determining the stage of a tumor, the methods of the present invention can be used to determine the origin of the tumor, i.e., the type of primary tumor, eg, colon cancer, breast cancer, lung cancer, etc.

[0216] Therefore, the present invention also provides an in vitro or in silico method for determining the origin of a tumor from a subject, comprising: (a) providing a DNA sequence of interest or a subsequence thereof from a subject, or preprocessed information obtained from said DNA sequence or subsequence, said DNA sequence of interest being a DNA sequence that encodes a repetitive sequence containing a high density of CpG dinucleotides distributed throughout the subject's genome, as input to a classifier trained to distinguish between healthy and cancerous CpG methylation profiles; (b) using the classifier, i) identifying a CpG methylation profile of a DNA sequence of interest or a subsequence thereof from said subject as a healthy CpG methylation profile or a cancerous CpG methylation profile; ii) comparing the cancerous CpG methylation profile of the DNA sequence of interest or a subsequence thereof from the subject with the CpG methylation profiles of tumors of different known origins, in particular the CpG methylation profiles of tumors from a population of cancer subjects with tumors of different known origins; iii) correlating the cancerous CpG methylation profile with a particular origin to determine the origin of the tumor from the subject as an output of the classifier.

[0217] The present invention also provides an in vitro or in silico method for determining the stage of a tumor in a subject, comprising: (a) providing a DNA sequence of interest or a subsequence thereof from a subject, or preprocessed information obtained from said DNA sequence or subsequence, said DNA sequence of interest being a DNA sequence that is distributed throughout the subject's genome and encodes a repeat sequence containing a high density of CpG dinucleotides, as input to a classifier trained to distinguish between healthy CpG methylation profiles and various stages of cancerous CpG methylation profiles; (b) using the classifier to identify a CpG methylation profile of the subject's DNA sequence of interest or a subsequence thereof as an output of the classifier as a healthy CpG methylation profile or a cancerous CpG methylation profile of a particular stage, thereby determining the stage of a tumor from the subject.

[0218] The present invention also provides an in vitro or in silico method for determining the origin and stage of a tumor in a subject, comprising: (a) providing a DNA sequence of interest or a subsequence thereof from a subject, or preprocessed information obtained from said DNA sequence or subsequence, said DNA sequence of interest being a DNA sequence that is distributed throughout the subject's genome and encodes a repeat sequence containing a high density of CpG dinucleotides, as input to a classifier trained to distinguish between healthy CpG methylation profiles and cancerous CpG methylation profiles of various tumor origins and stages; (b) using the classifier to identify a CpG methylation profile of the subject's DNA sequence of interest or a subsequence thereof as an output of the classifier as a healthy CpG methylation profile or a cancerous CpG methylation profile of a particular tumor origin and stage, thereby determining the origin and stage of a tumor from the subject.

[0219] The methods of the present invention are useful for diagnosing primary tumors and metastases. Typically, metastases spread from a primary tumor, but more than 10% of patients presenting to an oncology department have metastases of unknown primary tumor.

[0220] To this end, the present invention also provides an in vitro or in silico method for determining whether a subject is suffering from cancer metastasis, comprising: (a) providing a DNA sequence of interest or a subsequence thereof from a subject, or preprocessed information obtained from said DNA sequence or subsequence, said DNA sequence of interest being a DNA sequence that encodes a repetitive sequence containing a high density of CpG dinucleotides distributed throughout the subject's genome, as input to a classifier trained to distinguish between healthy and cancerous CpG methylation profiles; (b) using the classifier, i) identifying a CpG methylation profile of a DNA sequence of interest or a subsequence thereof from said subject as a healthy CpG methylation profile or a cancerous CpG methylation profile; ii) comparing the cancerous CpG methylation profile of the DNA sequence of interest or a subsequence thereof from the subject with known CpG methylation profiles of metastases, in particular CpG methylation profiles of tumors determined / obtained from a population of cancer subjects known to have metastases (e.g., by the methods described herein); iii) correlating the cancerous CpG methylation profile with known cancer metastasis CpG methylation profiles to determine as an output of the classifier whether the subject is suffering from cancer metastasis.

[0221] The present invention also provides an in vitro or in silico method for determining whether a subject is suffering from cancer metastasis, comprising: (a) providing a DNA sequence of interest or a subsequence thereof from a subject, or preprocessed information obtained from said DNA sequence or subsequence, said DNA sequence of interest being a DNA sequence that is distributed throughout the subject's genome and encodes a repetitive sequence containing a high density of CpG dinucleotides, as input to a classifier trained to distinguish between a healthy CpG methylation profile and a cancerous CpG methylation profile associated with metastasis; (b) using the classifier to identify a CpG methylation profile of the subject's DNA sequence of interest or a subsequence thereof as an output of the classifier as a healthy CpG methylation profile or a cancer metastatic CpG methylation profile, thereby determining whether the subject is suffering from cancer metastasis.

[0222] The present invention also provides an in vitro or in silico method for determining the stage of a tumor in a subject, wherein determining the stage of the tumor includes determining the presence of metastases, the method comprising: (a) providing a DNA sequence of interest or a subsequence thereof from a subject, or preprocessed information obtained from said DNA sequence or subsequence, said DNA sequence of interest being a DNA sequence that is distributed throughout the subject's genome and encodes a repeat sequence containing a high density of CpG dinucleotides, as input to a classifier trained to distinguish between healthy CpG methylation profiles and cancerous CpG methylation profiles of various stages, including metastasis; (b) using the classifier to identify the CpG methylation profile of the subject's DNA sequence of interest or a subsequence thereof as an output of the classifier as a healthy CpG methylation profile or a cancerous CpG methylation profile of a particular stage, including metastasis, to determine the stage of the tumor and the presence of metastasis.

[0223] The present invention also provides an in vitro or in silico method for determining the origin of a tumor from a subject and whether the subject is suffering from metastasis, comprising: (a) providing a DNA sequence of interest or a subsequence thereof from a subject, or preprocessed information obtained from said DNA sequence or subsequence, said DNA sequence of interest being a DNA sequence that is distributed throughout the subject's genome and encodes a repeat sequence containing a high density of CpG dinucleotides, as input to a classifier trained to distinguish between healthy CpG methylation profiles and metastatic cancerous CpG methylation profiles from various tumor origins; (b) using the classifier to identify the CpG methylation profile of the DNA sequence of interest or a subsequence thereof in the subject as an output of the classifier as a healthy CpG methylation profile or a metastatic cancerous CpG methylation profile from a particular tumor origin, thereby determining the origin of the tumor in the subject and determining whether the patient is suffering from metastasis.

[0224] The methods of the invention for determining whether a subject is suffering from cancer (i.e., whether the patient has cancer or is healthy) can be followed by methods for determining the origin of the cancer, for determining the stage of the cancer and / or for determining whether the patient is suffering from metastasis, particularly if the subject has been classified as a cancer subject.

[0225] The test method of the present invention for determining whether a subject has cancer can be performed once or multiple times during the subject's lifetime, thereby making it possible to monitor the occurrence of cancer, the progression of cancer, or the occurrence of cancer recurrence or metastatic recurrence.

[0226] In certain embodiments, the subject suffering from cancer is a subject who has undergone / been exposed to anti-cancer treatment, such as resective surgery, chemotherapy, radiation therapy, or immunotherapy.

[0227] In certain embodiments, DNA from a subject is provided once to determine whether the subject has cancer, and then, if the subject has cancer, one or more times during first-line treatment, particularly to determine whether the treatment is effective, e.g., whether the subject is cured of cancer or whether symptoms associated with cancer are claimed.

[0228] The methods of the invention for determining whether a subject has cancer may be performed after first-line treatment or after the patient is considered cured of cancer, for example, 6 months, 1 year, 2 years, 3 years, 4 years, 5 years or 10 years after first-line treatment or after the patient is identified / deemed cured.

[0229] This advantageously allows for monitoring the occurrence of cancer or metastatic recurrence. If such an event occurs, a second line of treatment may be administered to the subject. DNA from the subject can be provided one or more times during the second line of treatment.

[0230] The efficacy of first-line and / or second-line treatments can be assessed by methods of monitoring response to anti-cancer treatments, particularly therapeutic compounds, of the invention, which can be performed one or more times during first-line and / or second-line treatment.

[0231] Thus, the present invention provides an in vitro or in silico method for monitoring the response of a subject suffering from cancer to an anti-cancer treatment, in particular to a therapeutic compound / agent, comprising: (i) providing a classifier trained to distinguish between DNA sequences having a healthy CpG methylation profile and DNA sequences having a cancerous CpG methylation profile, as a first input, at least one DNA sequence of interest or a subsequence thereof from a first liquid biopsy from a subject suffering from cancer prior to administration of a therapeutic compound to the subject, said DNA sequence of interest being repeated throughout the subject's genome and comprising a high density of CpG sites or fragments thereof, or preprocessed information obtained from said first liquid biopsy, and as a second input, a second liquid biopsy from said subject after administration of a therapeutic compound, said second liquid biopsy comprising at least one DNA sequence of interest or a subsequence thereof, or preprocessed information obtained from said second liquid biopsy; (ii) using the classifier to identify each CpG methylation profile of each DNA sequence of the first liquid biopsy as having a healthy CpG methylation profile or a cancerous CpG methylation profile as a first output of the classifier, and to identify each CpG methylation profile of each DNA sequence of the second liquid biopsy as having a healthy CpG methylation profile or a cancerous CpG methylation profile as a second output of the classifier; wherein the number of DNA sequences of interest classified as having a healthy CpG methylation profile in the second output of the classifier is greater than the number of DNA sequences of interest classified as having a healthy CpG methylation profile in the first output of the classifier, indicating that the subject will be responsive to the therapeutic compound; and wherein the number of DNA sequences of interest classified as having a healthy CpG methylation profile in the second output of the classifier is less than or equal to the number of DNA sequences of interest classified as having a healthy CpG methylation profile in the first output of the classifier, indicating that the subject will not be responsive to the therapeutic compound.

[0232] Alternatively, a number of DNA sequences of interest classified as having a cancerous CpG methylation profile in the second output of the classifier that is less than the number of DNA sequences of interest classified as having a cancerous CpG methylation profile in the first output of the classifier indicates that the subject is responsive to the therapeutic compound, whereas a number of DNA sequences of interest classified as having a cancerous CpG methylation profile in the second output of the classifier that is greater than or equal to the number of DNA sequences of interest classified as having a cancerous CpG methylation profile in the first output of the classifier indicates that the subject is not responsive (resistant) to the therapeutic compound.

[0233] Without the aid of machine learning approaches, it is possible to monitor the response of patients suffering from cancer to anti-cancer treatment, in particular to therapeutic compounds / agents, by determining the CpG methylation profile of the patient before and after treatment, in particular according to the methods for determining the CpG methylation profile disclosed herein.

[0234] Preferably, the anti-cancer treatment is selected from the group consisting of resective surgery, chemotherapy, radiotherapy or immunotherapy.

[0235] Preferably, the therapeutic compound is a chemotherapeutic or immunotherapeutic compound. Chemotherapeutic compounds may be, but are not limited to, alkylating agents, antimetabolites, plant alkaloids, topoisomerase inhibitors, and antitumor antibiotics. Immunotherapeutic compounds may be, but are not limited to, antibodies, cytokines, or interferons.

[0236] The method for monitoring the response of a subject suffering from cancer to anti-cancer treatment, particularly to a therapeutic compound, relies on the alteration of a cancerous CpG methylation profile of a DNA sequence of interest from a subject suffering from cancer to a healthy CpG methylation profile, which alteration to a healthy CpG methylation profile indicates the efficacy of the anti-cancer treatment in the subject under study.

[0237] In another embodiment, a method for monitoring the response of a subject suffering from cancer to anti-cancer treatment, particularly to a therapeutic compound, relies on changes in CpG methylation levels. For example, if a DNA sequence contains a CpG site known to be hypomethylated in cancer, a decrease in hypomethylation indicates the efficacy of the anti-cancer treatment in the subject under study.

[0238] For this reason, the present invention also provides an in vitro or in silico method for assessing the efficacy of a compound to revert a cancerous CpG methylation profile of a DNA sequence of interest from a subject suffering from cancer to a healthy CpG methylation profile, comprising: (i) providing a DNA sequence of interest or a subsequence thereof from a subject treated with a compound, said DNA sequence of interest being repeated and distributed throughout the subject's genome and containing a high density of CpG dinucleotides or any fragment thereof, or preprocessed information obtained from said at least one DNA sequence of interest or a subsequence thereof, as input to a classifier trained to distinguish between DNA sequences having a healthy CpG methylation profile and DNA sequences having a cancerous CpG methylation profile; (ii) using the classifier to detect DNA sequences with healthy CpG methylation profiles and / or DNA sequences with cancerous CpG methylation profiles as an output of the classifier; The present invention relates to a method wherein the amount of DNA sequences having a healthy methylation profile exceeds a reference amount of DNA sequences having a healthy methylation profile obtained from the subject before any treatment with a compound, indicating that the compound is capable of reverting the cancerous CpG methylation profile to a healthy CpG methylation profile.

[0239] The ability of a compound to revert a cancerous CpG methylation profile of a DNA sequence of interest from a subject suffering from cancer to a healthy CpG methylation profile may indicate the ability of the compound to treat cancer.

[0240] Thus, the present invention also relates to in vitro or in silico methods of predicting or testing the ability of a compound to treat cancer, which methods typically comprise using the methods disclosed herein to assess the efficacy of a compound to revert a cancerous CpG methylation profile of a DNA sequence in a subject to a healthy CpG methylation profile, wherein an amount of DNA sequences classified as having a healthy CpG methylation profile in a biological sample of a subject treated with the compound that is greater than a reference amount obtained from the subject's biological sample prior to any treatment of the subject with the compound indicates that the compound is useful for treating the cancer.

[0241] In certain embodiments, the testing methods disclosed herein rely on preprocessed information obtained from a DNA sequence of interest or a subsequence thereof.

[0242] Preferably, in the testing method disclosed herein, samples are randomly drawn from the training dataset without replacement. This means that the samples used to train the classifier are different (i.e., not the same) from the samples used in the testing method of the present invention. For example, the sample population is divided into 60% for training the classifier and 40% for the testing method.

[0243] Preferably, in the methods disclosed herein that use a classifier, the classifier is selected from a support vector machine (SVM) classifier, a random forest (RF) classifier, a decision tree classifier, a K-nearest neighbor classifier (KNN), a logistic regression classifier, a nearest neighbor classifier, a Gaussian mixture model (GMM), a nearest centroid classifier and an artificial neural network, such as a deep neural network, a convolutional neural network or a fully connected neural network, more preferably selected from a support vector machine (SVM) classifier, a random forest (RF) classifier and a convolutional neural network (CNN), and even more preferably an RF classifier.

[0244] In certain embodiments, the methods of the present invention further comprise determining the presence of a mutation or genetic alteration in at least one gene that is deregulated in cancer, e.g., one of the most commonly used mutations for detecting ctDNA among the 299 recurrent oncogenic mutations identified from The Cancer Genome Atlas (TCGA) described in Bailey et al., Cell. 2018;173(2):371-385.e18, in particular EGFR, TPp53, AKT1, BRAF, ERBB2, KIT, MET, RB1, FGFR, JAK, TSC, PDGFR, PIK3CA, ESR1, NRAS, CTNNB1, FBXW7, APC, CDKN2A, PTEN, FGFR2, HRAS, KRAS, PPP2R1A, GNAS (see, e.g., Cohen, Science 1, eaar3247-10 (2018)), or any combination thereof.

[0245] Mutations in such genes are well documented in the art to be associated with cancer.

[0246] In further particular aspects, the methods described herein, in particular the training and testing methods, are computer-implemented methods.

[0247] In another particular aspect, the present invention provides a method for producing a pharmaceutical composition comprising: - a memory storing at least one instruction for a classifier trained according to any of the training methods described herein, in particular the method for training a classifier to accurately distinguish between healthy and cancerous CpG methylation profiles or between healthy and cancerous subjects, a processor accessing said memory to read the aforementioned instructions and to perform the testing method of the invention, in particular a method for determining the health status of a subject, in particular a method for determining whether a subject is a healthy subject or a subject suffering from cancer or a recurrence of cancer, a method for determining the origin of a tumor in a subject, a method for determining the stage of a tumor, a method for monitoring the response of a subject suffering from cancer to a therapeutic compound or a method for evaluating the efficacy of a compound for reverting a cancerous CpG methylation profile of a DNA sequence of interest from a subject suffering from cancer to a healthy CpG methylation profile or a method for evaluating the efficacy of a compound for treating cancer.

[0248] kit In another aspect, the present invention relates to a kit of primers or probes targeting DNA sequences, preferably DNA sequences from a subject that encode repetitive sequences distributed throughout the genome of the subject, preferably containing a high density of CpG dinucleotides, more preferably retrotransposons.

[0249] In particular, the primers or probes target DNA encoding a LINE-1 retrotransposon, preferably a LINE-1 retrotransposon as set forth in SEQ ID NO: 2 or 29 or having at least 85%, 90%, 95%, 97%, 98% or 99% sequence identity thereto.

[0250] Such primers or probes may be complementary to any region of the LINE-1 retrotransposon, such as the 5'UTR, ORFI, ORFII, or 3'UTR regions.

[0251] In certain embodiments, the primers targeting LINE-1 comprise an adapter, such as a unique molecular identifier (UMI) or a unique dual index (UDI).

[0252] In certain embodiments, the primers comprise common or universal sequences CS1 and / or CS2. The common sequences may consist of or comprise, for example, common sequence 1 (CS1) (5'-ACACTGACGACATGGTTCTACA-3' SEQ ID NO:27) and / or common sequence 2 (CS2) (5'-TACGGTAGCAGAGACTTGGTCT-3' SEQ ID NO:28) universal primer sequences.

[0253] In very particular embodiments, the kits of the invention may comprise probes of 100-150 bp, preferably 120 bp, starting every 20-30 bp, preferably every 24 bp. Preferably, such probes are designed to cover at least 75%, 80%, 85%, 90% or 95% of the DNA sequence of interest.

[0254] In another embodiment, the kit of the present invention may contain primers of 20 to 100 bp. Such probes are designed to cover at least 5% of the target DNA sequence.

[0255] Particular kits of the invention comprise at least four primers or probes selected from the group of primers or probes having the sequence set forth in SEQ ID NO: 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 or 26 respectively or a sequence having at least 80%, 85%, 90%, 95%, 97%, 98% or 99% identity thereto.

[0256] In a particular embodiment, the present invention relates to a kit of primer pairs targeting a subsequence of a DNA sequence encoding a LINE-1 retrotransposon, preferably a LINE-1 retrotransposon as set forth in SEQ ID NO: 1, said kit comprising: i) a forward primer having the sequence shown in SEQ ID NO: 11 and a reverse primer having the sequence shown in SEQ ID NO: 12; ii) a forward primer having the sequence shown in SEQ ID NO: 13 and a reverse primer having the sequence shown in SEQ ID NO: 14; iii) a forward primer having the sequence shown in SEQ ID NO: 15 and a reverse primer having the sequence shown in SEQ ID NO: 16; iv) a forward primer having the sequence shown in SEQ ID NO: 17 and a reverse primer having the sequence shown in SEQ ID NO: 18; v) a forward primer having the sequence shown in SEQ ID NO: 19 and a reverse primer having the sequence shown in SEQ ID NO: 20; vi) a forward primer having the sequence shown in SEQ ID NO: 21 and a reverse primer having the sequence shown in SEQ ID NO: 22; vii) a forward primer having the sequence shown in SEQ ID NO: 23 and a reverse primer having the sequence shown in SEQ ID NO: 24; viii) a forward primer having the sequence shown in SEQ ID NO: 25 and a reverse primer having the sequence shown in SEQ ID NO: 26 and any combination thereof.

[0257] Preferably, the kit comprises at least 5, 6, 7 or 8 primer pairs as disclosed herein above, more preferably 8 primer pairs.

[0258] The present invention also relates to the use of the kit of the present invention for amplifying a partial sequence of a DNA sequence encoding a LINE-1 retrotransposon, preferably a LINE-1 retrotransposon as set forth in SEQ ID NO: 2 or 29, particularly for cancer diagnosis by PCR multiplex.

[0259] Subjects and biological samples The sample analyzed using the methods and kits provided herein can be any biological sample and / or any sample containing DNA, typically obtained from a subject.

[0260] In certain embodiments, the DNA sequence of interest is cfDNA. Methylation patterns of cellular DNA are preserved in cell-free DNA (cfDNA).

[0261] cfDNA can be provided in a biological sample, such as a fluid sample obtained from a subject. In fact, cfDNA can be found in bodily fluids, such as plasma, serum, or urine. The concentration of cfDNA is typically low in biological samples, but can increase significantly under certain conditions, including (but not limited to) pregnancy, autoimmune disorders, myocardial infarction, and cancer. Circulating tumor DNA (ctDNA) is a component of circulating DNA, particularly from cancer cells. cfDNA and ctDNA provide real-time or near-real-time measures of the methylation status of source tissues. cfDNA and ctDNA have a blood half-life of approximately 2 hours, resulting in a sample collected at a given time providing a relatively timely reflection of the status of the source tissue.

[0262] In another specific embodiment, the DNA sequence of interest is circulating tumor DNA (ctDNA), which is tumor-derived fragmented DNA present in the bloodstream and is not associated with cells.

[0263] As used herein, the term "subject" refers to an organism, typically a mammal (e.g., a human). In particular, the subject suffers from a disease, disorder, or (abnormal) condition. In a particular embodiment, the subject is susceptible / prone to developing a disease, disorder, or (abnormal) condition. In another particular embodiment, the subject exhibits one or more symptoms or characteristics of a disease, disorder, or condition. In another particular embodiment, the subject does not suffer from a disease, disorder, or (abnormal) condition or does not exhibit any symptoms or characteristics of a disease, disorder, or condition, i.e., the subject is a healthy subject. In another embodiment, the subject is a subject that exhibits one or more features characteristic of a susceptibility to or risk of developing a disease, disorder, or (abnormal) condition. A particular subject is a patient. In another embodiment, the subject is an individual for whom a diagnosis has been established and / or who has been exposed to a therapeutic treatment or to whom a therapeutic compound / agent has been administered.

[0264] In certain embodiments, the subject is a human, particularly a child, infant, adolescent, or adult, particularly an adult at least 18 years of age, preferably an adult at least 40 years of age, and more preferably an adult at least 50 years of age. When the subject is a human subject, the "human subject" may also be identified herein as an "individual."

[0265] As used herein, the term "biological sample" typically refers to a sample obtained or derived from a biological source of interest (e.g., a tissue, organism, or cell culture) as described herein. In certain aspects, the biological source is or includes an organism, e.g., an animal or a human. A biological sample may include a biological tissue or bodily fluid. A biological sample may be or include cells, tissue, or bodily fluid. A biological sample may consist of or include blood, blood cells, free-floating nucleic acids, e.g., DNA, a biopsy sample, ascites, surgical specimen, cell-containing bodily fluid, sputum, saliva, stool, urine, cerebrospinal fluid, peritoneal fluid, pleural effusion, lymphatic fluid, gynecological fluid, synovial fluid, secretions, excretions, skin swab, vaginal swab, oral swab, nasal swab, washing or lavage, e.g., ductal lavage, bronchoalveolar lavage, aspirate, scraping, and / or bone marrow.

[0266] In certain aspects, the biological sample consists of or comprises samples obtained from a single subject or multiple subjects.

[0267] In an embodiment, the biological sample is a biopsy, in particular a solid biopsy or a liquid biopsy.

[0268] A tissue biopsy requires the removal of solid tissue from a subject's body. This biopsy is typically taken from a tissue or organ suspected of containing a solid tumor or tumor cells. Tissue biopsies are typically used when the location of a known tumor is suspected or confirmed and can be obtained.

[0269] The biological sample is preferably a liquid biopsy sample, such as a blood, plasma, serum, sputum, bronchial fluid, or pleural effusion sample.

[0270] The biological sample is preferably derived from blood, such as serum (also identified herein as "serum") or plasma (also identified herein as "plasma"), preferably plasma.

[0271] In a preferred embodiment, the DNA sequence of interest, typically the cfDNA sequence of interest, is obtained by drawing blood from the patient, followed by separating the plasma and extracting the DNA from the plasma.

[0272] Nucleic acid can be isolated from sample (for example, isolate cfDNA from blood or plasma) in various ways known in the art.Nucleic acid can be isolated, for example, by standard DNA purification techniques (but not limited to), for example, by direct gene capture (for example, by clarifying sample to remove assay inhibitors, and if present, by capturing target nucleic acid from the clarified sample with a capture agent to produce capture complex, and then isolating the capture complex to recover target nucleic acid).

[0273] In certain embodiments, the subject is a human subject diagnosed with or suffering from cancer, diagnosed with or suffering from cancer at risk of having cancer, and / or diagnosed with or suffering from cancer at immediate risk of having cancer.

[0274] As used herein, the terms "cancer," "malignancy," "neoplasm," "tumor," and "carcinoma" are used interchangeably to refer to a disease, disorder, or condition in which cells exhibit or relatively exhibit abnormal, uncontrolled, and / or autonomous growth, such that cells exhibit or exhibit an abnormally elevated proliferation rate and / or an abnormal growth phenotype. In certain aspects, cancer comprises one or more tumors. In certain aspects, cancer consists of or comprises pre-cancerous (e.g., benign), malignant, pre-metastatic, metastatic, and / or non-metastatic cells. In certain aspects, cancer consists of or comprises a solid tumor. In certain aspects, cancer consists of or comprises a hematological tumor. In the context of the present invention, cancer may be, for example, colon cancer, hematopoietic cancer such as leukemia, lymphoma (Hodgkin's and non-Hodgkin's), myeloma or myeloproliferative disorders, such as sarcoma, melanoma, adenoma, carcinoma of solid tissue, in particular squamous cell carcinoma of the mouth, throat, larynx and / or lung, liver / bile duct cancer, genitourinary cancer such as prostate cancer, cervical cancer, bladder cancer, urothelial cancer, ovarian cancer, uterine cancer and / or endometrial cancer, renal cell cancer, bone cancer, pancreatic cancer, skin cancer, cutaneous cancer, intraocular melanoma, uveal melanoma, cancer of the endocrine system, thyroid cancer, parathyroid cancer, head and neck cancer, breast cancer, gastrointestinal cancer, cancer of the nervous system.

[0275] In another embodiment, the subject has a benign tumor / lesion, such as a papilloma-induced tumor / lesion.

[0276] Preferably, the cancer detected by the method or kit of the present invention is selected from the group consisting of colon cancer, breast cancer, lung cancer, uveal melanoma, ovarian cancer and gastric cancer.

[0277] In a particular embodiment, the cancer detected by the method or kit of the invention is an early stage cancer, in particular a stage I, II or III cancer.

[0278] As used herein, the term "cancer stage" refers to a qualitative or quantitative assessment of the level of progression of cancer. In certain embodiments, criteria used to determine the stage of cancer include, for example, one or more of the following: the location of the cancer in the body, the size of the tumor, whether the cancer has spread to lymph nodes, whether the cancer has spread to one or more different parts of the body, etc. In certain embodiments, the cancer is staged using the so-called TNM system, where T refers to the size and spread of the main tumor, usually called the primary tumor, N refers to the number of nearby lymph nodes that have cancer, and M refers to whether the cancer has metastasized. Early-stage cancer is a term used to describe cancer in which the cancer is in its early stages of growth and may not have spread to other parts of the body. This term particularly refers to stages I, II, and sometimes III. In certain embodiments, the cancer is stage I-III (cancer is present, and the higher the number, the larger the tumor and the more it has spread to nearby tissues) or stage IV (cancer has spread to distant parts of the body). In certain embodiments, the cancer is assigned a stage selected from the group consisting of in situ (abnormal cells are present but have not spread to nearby tissues), localized (cancer is limited to where it began and there are no signs of spreading), regional (cancer has spread to nearby lymph nodes, tissues, or organs), distant (cancer has spread to distant parts of the body), or unknown (there is not enough information to identify the stage of the cancer).

[0279] Purpose The methods and kits of the present invention can be used in a variety of applications, for example, the disclosed methods and kits can be used to screen, detect or diagnose cancer, or to assist in the screening of cancer.

[0280] In certain embodiments, screening will provide a diagnosis of a condition, eg, type, origin, or stage of cancer, using the methods and / or kits disclosed herein.

[0281] Those skilled in the art will appreciate that regular, preventative, and / or prophylactic screening for cancer improves the efficiency of diagnosis and, if necessary, treatment. As noted above, early-stage cancer includes cancer stages I-III according to at least one system of cancer staging. Thus, the present disclosure provides methods and kits that are particularly useful for early cancer diagnosis and treatment.

[0282] In certain embodiments, cancer screening of the present invention is performed one or more times for a given subject, hi certain embodiments, cancer screening is performed periodically, for example, every six months, every year, every two years, every three years, every four years, every five years, or every ten years.

[0283] The methods of the invention may be performed on subjects who are asymptomatic at the time of screening, so that the methods and kits of the invention are particularly likely to detect early stage cancers.

[0284] In certain embodiments, screening using the methods and / or kits of the present disclosure can be followed by a further diagnostic confirmation assay, which can be performed to confirm the diagnosis obtained from the methods disclosed herein, e.g., solid biopsy or imaging methods.

[0285] Any of the methods described herein can be used to determine the initiation of treatment or to select appropriate / optimized treatment in a subject in whom a tumor is present.

[0286] Also, in certain aspects, the present invention relates to methods for treating tumors, which may optionally include performing a diagnostic method of the present invention, wherein if a tumor is present, e.g., if the subject is identified as a cancer subject, the tumor is treated.

[0287] The present invention also relates to administering an appropriate anti-cancer treatment, in particular a drug / medication or drug / medication combination, optionally in conjunction with surgery and / or radiation therapy, if the method of the present invention indicates that the subject is suffering from cancer.

[0288] Alternatively, the diagnostic method can be used to perform additional diagnostic steps when a tumor is detected, for example, by analysis of a solid biopsy and / or tumor imaging. [Brief explanation of the drawings]

[0289] [Figure 1]Figure 1. Targeting DNA methylation patterns of primate-specific LINE-1 elements from plasma DNA. A. CpG density along the structure of a 95-CpG-containing human-specific LINE-1 (L1HS) element. The L1PA_cfDNAme assay targets 34 CpGs (approximately 30%). Each target amplicon is highlighted with a black bar below the structure. The number of CpG sites detected per amplicon is displayed in blue. B-C. Circos plots showing genome-wide hits obtained with the L1PA_cfDNAme assay. B. This panel displays overlapping hits obtained with deep sequencing in healthy donor plasma versus ovarian cancer tissue. C. This panel displays overlapping hits obtained with deep sequencing in healthy donor plasma versus uveal melanoma tissue. D. Histogram summarizing the most representative L1 subfamilies targeted by the L1PA_cfDNAme assay across three deep sequencing samples. Each bar corresponds to a sample and is ordered as follows: Healthy = healthy plasma, 54M reads; OVC = ovarian cancer tissue, 44M reads; UVM = uveal melanoma tissue, 46M reads. L1 subfamilies are ordered by their most representative (total copies targeted by the assay across the three deep sequencing samples used, descending order). Colors highlight the relative contribution of L1P A copies that are uniquely mapped (black), randomly mapped (gray), and hit by both (crosshatched). L1 copies are obtained by finding overlaps between all hits and annotated copies of L1P / L1HS in repeatmasker (hg38 genome version). Each copy is counted once. [Figure 2]Figure 2. L1PA hypomethylation is detectable in plasma DNA in multiple forms of cancer. A. Mean levels of methylation in healthy plasma, healthy tissue collected adjacent to ovarian tumors, ovarian tumors, and breast tumors. The mean level of methylation for each sample corresponds to the % CG dinucleotides at each CpG site averaged by the number of CpG sites. p-values ​​were calculated using Student's t-test (pHP=0.026, pOvcTumors=1.5e-05, pBrcTumors=1.4e-05). B. Mean levels of methylation in six cancer types, including four metastatic stage cohorts and three non-metastatic stage cohorts. The mean level of methylation for each sample corresponds to the % CG dinucleotides at each CpG site averaged by the number of CpG sites. p-values ​​were calculated using Student's t-test (pcolon_M+<1e-4, pbreast_M+<1e-4, plung_M+=0.0377, puvea_M+=0.0017, povary_M0<1e-4, pstomach_M0<1e-4, pbreast_M0=0.0227). The dashed black line represents the median. C. Methylation levels at each target CpG site (x-axis) for each sample (y-axis) are shown as a heatmap. No clustering was performed on the data, with the target CpG site (amplicon # indicated) on the x-axis and sample type (healthy human donor, colon cancer, breast cancer, and ovarian cancer plasma) on the y-axis. The metaplot represents the mean levels for healthy vs. cancer samples at each CpG site. D. Differential methylation levels between healthy samples and patients for each cancer type. The CG sites are arranged in 5' to 3' order along the L1 structure. P values ​​were calculated using Student's t test (corrected for multiple testing using the p.adjust method "BH", an R function corresponding to Benjamini & Hochberg (1995)). [Figure 3]Figure 3. L1PA methylation changes distinguish cancer samples from healthy donors. A-B. Receiver operating characteristic (ROC) curves obtained for classification of plasma samples using the methylation percentages at the 33 targeted CpGs in all cancer samples (A) or by cancer subtype (B). C-D. ROC curves obtained for classification of plasma samples using the haplotype percentages within each amplicon target in all cancer samples (C) or by cancer subtype (D). All classifications involve 5,000 stratified random replicates trained on 60% of samples and tested by bootstrap on the remaining 40%. CRC_M+ n=75, BRC_M+ n=97, NSCLC_M+ n=52, UVM n=70, OVC_M0 n=23, GAC_M0 n=27, and BRC_M0 n=40 were tested on 123 healthy donors. E. Sensitivity at 99% specificity by cancer class for the two models (ordered by highest sensitivity, bars indicate 95% CI). [Figure 4]Figure 4. Specific performance of cancer samples is reproducible across independent cohorts. A. Table showing the number of patients by cancer type and stage (non-metastatic: M0 vs. metastatic: M+) in Cohort 1 and Cohort 2. B. Heatmap showing the methylation levels of each target CpG site (x-axis) for each healthy sample (y-axis) from Cohort 1 vs. Cohort 2. No clustering was performed on the data, with the target CpG site (shown as amplicon #) aligned on the x-axis. The metaplot represents the average level for donors from Cohort 1 vs. Cohort 2 at each CpG site (gray intensity corresponds to the legend highlighted on the side of the heatmap). C. Comparison of the average methylation levels, excluding amplicon 3, between healthy donors from Cohorts 1 and 2 and five cancer types common to the two cohorts. Methylation levels were calculated as previously described for Figure 2. p-values ​​are calculated using Student's t-test (povary_M0_C1 vs C2 < 1e-4, pstomach_C1M0 vs C2M += 2e-4). Dashed lines represent medians. D-E. ROC curves obtained with classifiers trained on cohort 1 and tested on cohort 2 using the methylation percentage at each CpG target excluding the CpG in amplicon 3 for all cancer samples combined (D) or by cancer type (E). No bootstrapping step. F. Sensitivity at 99% specificity for different cancer classes using classifiers based on the methylation level of a single CpG, excluding (dark gray) or including (light gray) amplicon 3. Bars indicate 95% CI. [Figure 5]Figure 5. Cancer origin can be directly estimated from L1PA methylation status. A-D. ROC curves obtained with a "multiclass" classifier trained on Cohort 1 and tested on Cohort 2 using single-CpG methylation levels with (A) or without (B) amplicon #3 and bootstrapping, or with (C) or without (D) haplotype proportions and bootstrapping with (C) amplicon #3 and bootstrapping. Only classes that were homogeneous between Cohorts 1 and 2, i.e., healthy donors, CRC_M+, BRC_M+, and OVC_M0M+, were included in this study. E. AUC by cancer class for two models (single-CpG methylation or haplotype proportions) with and without amplicon #3 and bootstrapping. Bars indicate 95% CI. [Figure 6] Figure 6. Scheme of the targeted bisulfite sequencing strategy used to construct the L1PA-cfDNAme library. The protocol begins with one cycle of linear PCR to incorporate unique molecular identifiers (UMIs) to identify each starting molecule present in the sample. We also incorporated a second set of molecular identifiers (UIDs) during a second PCR to generate a library with sufficient nucleotide diversity, critical for successful downstream sequencing (see the Methods section for more details). [Figure 7] Figure 7. Overview flowchart showing the pipeline developed for reference-free alignment of sequencing data. [Figure 8] Figure 8. A. Comparison of cancer detection rates using L1PA_cfDNAme and common recurrent mutations evaluated in previous studies. B. Comparison of cancer detection rates using L1PA_cfDNAme and the non-invasive Galleri test developed by Grail. [Figure 9] Figure 9. ROC curve obtained with a classifier trained on cohort 1 and tested on cohort 2 using the proportion of haplotypes excluding CpGs in amplicon #3.

[0290] Example The following examples are presented for illustrative and non-limiting purposes and serve to illustrate the present invention.

[0291] Materials and Methods method Preparation of plasma DNA Whole blood was collected into BD Vacutainer EDTA tubes (BD Biosciences). Plasma was separated within 3 hours of collection to ensure high-quality cell-free DNA. Blood was centrifuged at 820 g for 10 minutes at room temperature. The supernatant was then transferred to a new 2 ml tube and centrifuged at 16,000 g for 10 minutes at 15°C to remove any remaining cell debris. Plasma was collected, transferred to a new 2 ml tube, and stored at -80°C for further processing. DNA was extracted from 2 ml of plasma using the automated QIAsymphony Circulating DNA Kit (Qiagen) or the manual QIAamp Circulating Nucleic Acid Kit (Qiagen) according to the manufacturer's instructions, and the isolated DNA was eluted with 60 μl or 36 μl of elution buffer, respectively. Plasma DNA was quantified using the Qubit® dsDNA HS Assay Kit (Thermo Fisher Scientific) on a Qubit® 2.0 Fluorometer according to the manufacturer's instructions and stored at -20°C until use.

[0292] Preparation of DNA from cell lines and tissues DNA isolation from cell lines and peripheral blood mononuclear cells (PBMCs) was performed using the QIAamp DNA Mini Kit or QIAamp DNA Blood Mini Kit (Qiagen) according to the manufacturer's instructions. DNA from cryopreserved and formalin-fixed, paraffin-embedded (FFPE) tumor tissues was extracted using the classic phenol-chloroform protocol and the NucleoSpin® FFPE DNA Kit (Macherey-Nagel), respectively. Isolated DNA was quantified using the Qubit® dsDNA BR Assay Kit on a Qubit® 2.0 Fluorometer.

[0293] Bisulfite conversion Bisulfite treatment of genomic DNA isolated from cancer tissues, cancer cell lines, and PBMCs was performed using the EZ DNA Methylation-Gold Kit™ (Zymo Research, CA, USA) according to the manufacturer's instructions. Up to 200 ng of genomic DNA was treated with Zymo CT conversion reagent in a thermal cycler at 98°C for 10 minutes, followed by 64°C for 2.5 hours. The bisulfite-treated DNA was purified using the spin column provided with the kit. Bisulfite treatment of plasma DNA was performed using the Zymo EZ DNA Methylation-Lightning Kit™ (Zymo Research, CA, USA) according to the manufacturer's instructions. DNA (up to 200 ng) isolated from 2 ml of plasma was treated with the Zymo Lightning conversion reagent under the following cycling conditions: 98°C for 8 minutes and 54°C for 60 minutes. The bisulfite-treated DNA was purified using the spin column provided with the kit. The bisulfite-treated DNA was stored at −70°C and further used to construct a sequencing library.

[0294] Primer design Eight primer pairs were designed using the LINE-1 human-specific consensus sequence from Repbase (Figure 1A). The 5' UTR (promoter region) is CpG-rich and a common target for methylation quantification, but L1PA transcripts are often 5'-truncated. Therefore, to target more L1PA elements and improve assay sensitivity, primers for ORFI and ORFII as well as the 3' UTR were also designed. All primers were designed on the plus strand of bisulfite-converted DNA using MethPrimer or PyroMark software. Target regions ranged from 101 bp to 150 bp, with an average size of 170 bp, to better capture plasma DNA fragments containing two to seven CpG targets. The primers were methylation-independent and contained no CpG sites and ~2 CpG sites, with no CpG sites toward the 5' end of the primer. To avoid methylation-biased amplification, degenerate primers targeting both methylated and unmethylated states were used for primers containing CpG sites. Both target-specific primers contained Fluidigm's universal CS (consensus sequence) tag at their 5' ends for the latter amplification step. Sixteen Ns (random nucleotide sequences) were incorporated as unique molecular identifier (UMI) sequences between the target-specific sequence and S2 in the reverse primer, enabling identification of unique individual molecules and accurate scoring of DNA methylation rates. These primers allowed the generation of 4.294 billion distinct UMIs. Because LINE-1 covers thousands of copies per genome, a large number of distinct UMIs were used to uniquely barcode each target molecule. One strand of each template molecule was encoded with a UMI using a single-cycle linear target-specific PCR.

[0295] A 16-nucleotide stretch was incorporated between the target-specific sequence and CS1 in the forward primer to increase the diversity of the target sequencing library and improve sequencing quality. All primers were obtained from Eurogentec (RP-cartridge purification method). Seven amplicons (#1, #3, #4, #5, #6, #7, #8) were multiplexed in a single reaction. Amplicon #2 was treated separately because it overlaps with other primers. The designed primers were evaluated by in silico PCR using converted human genomic DNA as a reference.

[0296] Preparation of targeted bisulfite sequencing library (Figure 6) Sequencing libraries were prepared using three PCR steps: target-specific linear amplification (UMI assignment), target-specific exponential amplification, and barcode PCR.

[0297] Each library was prepared in two individual reactions containing (due to overlap between primers): 1. Multiplex PCR amplification of seven probes (L1HS-2, 4, 6, 7, 8, 9, 15) using Platinum™ Multiplex PCR Master Mix, Thermofisher, Life Technologies SAS 2. Single PCR amplification of one probe (L1HS-14) using Hot Star Taq Plus DNA polymerase (Qiagen) Includes.

[0298] In multiplex reactions, each primer was used at a final concentration of 0.01–0.06 μM (lower concentrations were used to avoid primer dimers), while in single reactions, 0.1 μM primer was used. Up to 5 ng and 4 ng of bisulfite-converted DNA (plasma and tissue DNA) were used for multiplex and single reactions, respectively.

[0299] UMI assignment for multiplex reactions was performed using the Platinum™ Multiplex PCR kit Master Mix (Thermofisher, Life Technologies SAS) in 25 μL reactions containing 1× Platinum™ Multiplex PCR Master Mix, 0.01–0.06 μM L1HS reverse primer mix (containing 16 Ns), and up to 5 ng bisulfite-converted DNA under the following thermocycling conditions: 95°C for 5 minutes, followed by one cycle of 95°C for 30 seconds, 58°C for 90 seconds, and 72°C for 40 seconds. UMI assignments for single reactions were generated using Hot Star Taq Plus DNA polymerase (Qiagen) in 25 μL reactions containing 1× Taq PCR buffer, 0.65 U Hot Star Taq (5 U / μL), 0.2 μM dNTPs, 1.5 mM MgCl2, 0.1 μM L1HS-14 reverse primer (containing 16 Ns), and up to 4 ng bisulfite-converted DNA under the following thermocycling conditions: 95°C for 10 minutes, followed by one cycle of 94°C for 60 seconds, 58°C for 30 seconds, and 72°C for 40 seconds.

[0300] To completely remove the reverse primer and dNTPs, 25 μL of each reaction mixture was treated with 50 U Exonuclease I (Thermo Fisher Scientific) and 10 U FastAP Thermosensitive Alkaline Phosphatase (Thermo Fisher Scientific) at 37°C for 1 hour, followed by heat inactivation at 80°C for 15 minutes.

[0301] Target-specific exponential amplification for multiplex reactions was performed using the Platinum™ Multiplex PCR kit Master Mix in 50 μL reactions containing 1× Platinum™ Multiplex PCR Master Mix, 0.01–0.06 μM L1HS forward primer mix (containing 16 Ns), 0.2 μM CS2 reverse primer, and 20 μL of purified PCR product under the following thermocycling conditions: 95°C for 5 minutes, followed by 28 cycles of 95°C for 30 seconds, 58°C for 90 seconds, and 72°C for 30 seconds, followed by a 10-minute incubation at 72°C. Target-specific exponential amplification for a single reaction was performed using a Hot Star Taq Plus DNA polymerizer in a 50 μL reaction containing 1× Taq PCR buffer, 0.65 U Hot Star Taq (5 U / μL), 0.2 μM dNTPs, 1.5 mM MgCl, 0.2 μM L1HS-14 forward primer (containing 16 Ns), 0.2 μM CS2 reverse primer, and 16 μL of purified PCR product under the following thermocycling conditions: 95°C for 10 min, followed by 28 cycles of 94°C for 60 s, 58°C for 30 s, and 72°C for 30 s, followed by a 10-min incubation at 72°C.

[0302] PCR products from multiplex and single reactions were pooled together after quantification by qPCR. The pooled products were purified at a 1.2x ratio using Agencourt AMPure XP (Beckman Coulter) according to the manufacturer's protocol. The purified DNA was eluted in 30 μl of water.

[0303] Barcode PCR was performed using universal fluidic primers and complete sequencing adapters to introduce sample-specific barcodes. 25 μL of purified pooled PCR product, 1x Phusion HF buffer, 1 U Phusion Hot Start II DNA polymerase (Thermo Fisher Scientific), 0.2 μM fluidic primers, and 0.2 mM dNTPs were mixed in a final volume of 50 μL and amplified under the following thermocycling conditions: 98°C for 2 minutes, followed by 20–25 cycles of 98°C for 10 seconds, 62°C for 30 seconds, and 72°C for 30 seconds, followed by a 5-minute incubation at 72°C.

[0304] Finally, the amplified products were purified by double (upper and lower) size selection using two consecutive AMPure XP steps. In the first step, a low concentration of AMPure XP beads (0.6-0.7x ratio) was used. In this step, beads containing large fragments were discarded, and the supernatant was collected for the next step (reverse purification). In the second step, a higher concentration of beads (1.1-1.2x ratio) was used. In this step, beads containing the desired fragments were collected and purified according to the manufacturer's protocol. The size-selected library was eluted with 15 µL of low-EDTA TE buffer.

[0305] Libraries were quantified by fluorescent assay using the Qubit HS DNA kit (Thermo Fisher Scientific). Libraries were then quantified by electrophoretic assay using the Caliper LabChip HS DNA (PerkinElmer) or BioAnalyzer HS DNA kit (Agilent), qualified, and equimolar pooled for sequencing. Sequencing was performed on an Illumina HiSeq-rapid run mode or NovaSeq (PE30,170).

[0306] Lead pretreatment The sequencing facility provides multiple files in FASTQ format. Each sample is an individual FASTQ file containing its corresponding raw sequence, consisting of a number of sequence parts: CS1, forward UMI, forward primer, insert, reverse primer, reverse UMI, and CS2. For each sample, the reads are demultiplexed (i.e., cut using the atropos program) using the forward and reverse primer sequences to generate a FASTA file for each primer set containing the insert sequence and a reverse UMI for deduplication (the reverse UMI is unique for each input DNA molecule).

[0307] The insert sequences and reverse UMIs are then filtered by expected size, preferably with a tolerance of ±5 bases for inserts, and for example, no tolerance for UMIs consisting of at least 16 bases. The sequences of the inserts and reverse UMIs are then concatenated to form a single sequence, which is then de-duplicated using, for example, the vsearch program. The reverse UMIs are then trimmed, and the inserts are separated into separate FASTA files. All inserts obtained for all samples are aggregated on a primer set basis into one FASTA file, resulting in one file for each primer set.

[0308] Clustering, reference sequence extraction, and global alignment (Figure 7) For example, using the vsearch program (preferably with the following parameter: --cluster_fast <inputfasta>--notrunclabels--fasta_width 0--iddef 4--id 0--qmask none--clusterout_sort--consout <referencefasta>(using ) and apply clustering based on minimum sequence identity to each file (clustering of DNA subsequences is performed with the help of vsearch). If the pairwise identity with the centroid is greater than 0, this subsequence is added to a cluster. Pairwise identity is defined as the number of (matching sequences) / (alignment length), or a randomly selected subsample of 20 million reads if a given file contains more. For example, an awk program is used to separate the reference sequences of the "n", e.g., 10, largest clusters into separate files: one FASTA file per primer set each containing n (in this example, 10) reference sequences from the sequences of the n largest clusters. For example, use the mafft program (preferably with the following parameters: --textmatrix<custom_score_matrix.txt> --retree 2) for each of the files to pairwise align the n (e.g., 10) reference sequences using a custom scoring matrix that prioritizes dinucleotide CG / TG alignments over other possible dinucleotide combinations (e.g., AG / AC / TC...), resulting in n (in this example, 10) reference sequence database FASTA files for each primer set. Finally, the files are aligned using, for example, the mothur program [preferably with the following parameters: #align.seqs(candidate= <inputfasta>、template= <referencefasta>, align=needleman, match=1, mismatch=-1, gapopen=-1, gapextend=0) for each primer set FASTA file to align all sequences from all samples to the corresponding reference sequence database.

[0309] CG Calling To call CG dinucleotide sites of interest, a 2-bp sliding window was used across all aligned sequences to determine the distribution of dinucleotides along each subsequence of the DNA sequence of interest (here, each amplicon target). The percentage of CG dinucleotides was calculated at each position along the sequence (subsequence of interest), as well as the percentage of any other dinucleotide. A first threshold value for the CG / TG dinucleotide ratio (e.g., 20% or higher) was used to determine whether a considered position in the sequence (subsequence of interest) qualifies as a CG site. After bisulfite treatment, information about whether actual TG dinucleotides or TG dinucleotides resulting from conversion of unmethylated CG dinucleotides is lost. To mitigate this, a second threshold value for the TG ratio (e.g., 95% or higher) is preferably applied to these preselected sites to potentially eliminate from selection CG sites that are actual TG sites (e.g., not the result of bisulfite conversion).

[0310] Extraction of methylation levels and haplotypes From the aligned sequences, methylation patterns / profiles are extracted, e.g., using dedicated python scripts, and compiled for each sample into either the average level of methylation at each identified CG dinucleotide site or the percentage of methylation haplotypes (the methylation status of consecutive CpG sites within each subsequence of interest, e.g., an amplicon).

[0311] Machine learning trained classification model The resulting data (expressed as the average level of methylation per CG site or the proportion of methylated haplotypes) were specifically subjected to random forests (Breiman, L. Random Forests. Machine Learning 45, 5-32 (2001)) from the Python package Scikit-Learn (Pedregosa et al., JMLR 12, pp. 2825-2830, 2011). https: / / doi.org / 10.1023 / A:1010933404324) is used to perform supervised training of a statistical model using a classifier algorithm (with the following hyperparameters: n_estimators=300, criterion='gini', max_depth=None, min_samples_split=2, min_samples_leaf=1, min_weight_fraction_leaf=0.0, max_features='sqrt', max_leaf_nodes=None, min_impurity_decrease=0.0, bootstrap=True, oob_score=False, warm_start=False, class_weight=None, ccp_alpha=0.0, max_samples=None).

[0312] There are three main rationales for choosing Random Forest over other training methods: Random Forest is less prone to overfitting, performs better even when the quantitative relationship between features and observations is biased in favor of the former, as is the case when using a methylation haplotype data representation (WIREs Data Mining Knowl Discov 2012, 2: 493-507 doi: 10.1002 / widm.1072), and because Random Forest inherently returns a measure of variable importance, it greatly facilitates the interpretability of model decisions.

[0313] It is also worth noting that during the development of this test, most classifier learning algorithms from Scikit-Learn were tested, with Random Forest being one of the top performers.

[0314] * The features used to train the models were either the average level of methylation per CG site (n=33) or the proportion of methylated haplotypes (i.e., the combination of all possible methylation states of CG sites within a given amplicon) (n=274). No additional transformations or feature selection were performed on the data.

[0315] * We evaluate the model as follows: To estimate variances and confidence intervals, we perform 5000 classification runs. In each run, an equal number of samples from each class are randomly drawn without replacement from the training dataset. The samples from these draws are stratified by class and split into 60% for training and 40% for evaluation.

[0316] The true positive rate and false positive rate are evaluated for each run, with interpolation performed to generate all points on the ROC curve. At the end of the 5000 runs, an average ROC curve is generated and a 95% confidence interval is calculated based on the results of all runs. In the multi-class case, an ROC curve for each class is generated by treating the class under consideration as the positive class (thus not giving any special weighting to the control class, i.e., healthy plasma).

[0317] result By targeting primate-specific LINE-1 elements, plasma DNA methylation patterns can be determined genome-wide (Figure 1).

[0318] We have developed a method for detecting methylation patterns in DNA, particularly circulating cell-free DNA, specifically a PCR-based targeted bisulfite method combined with computer-implemented sequencing (also identified herein as "deep sequencing"). We used sodium bisulfite-based chemical conversion to achieve base-pair resolution analysis. Base-pair resolution analysis is preferred for addressing methylation levels at single CpG dinucleotides and co-methylation of multiple CpG sites to determine methylation haplotypes (methylation states of consecutive CpG sites). We specifically designed eight amplicons (target subsequences of interest) targeting the primate-specific L1 element (L1PA) for use in multiplex PCR (Figure 1A). Primers are equipped with unique molecular identifiers (UMIs), which aid in signal convolution and detection of true low-frequency alterations, helping to reduce errors. We detected thousands of L1PA elements scattered throughout the genome, as observed with genomic hits from deep-sequenced healthy plasma (data not shown). We observed very similar patterns in deep-sequenced ovarian cancer and uveal melanoma tissues (Figure 1B-C) as well as in healthy and standard-range cancer plasma, demonstrating the robustness of this approach. Overall, the estimated number of targeted L1PA elements is approximately 30,000–40,000 per genome, including half of the human-specific copies (L1HS) and many copies of other L1PA subfamilies. This represents 82–125,000 CpG sites (Figure 1D).

[0319] After multiplex PCR deep sequencing, reads are traditionally mapped back to the genome. However, the majority of sequencing reads from repetitive sequences (approximately 80%) are lost or randomly assigned during the mapping process and subsequently lost through classical differentially methylated region (DMR) calling. Therefore, we developed a computational tool to optimize (improve accuracy / performance) aligned repetitive sequence data without using a reference genome. To do this, we clustered all obtained high-quality reads based on their similarity and extracted a reference sequence from each cluster. Next, we aligned all reads to the identified reference sequence in a computational time of O(n) (n is the number of aligned reads, and O is Landau's Big O). This allows for alignment of millions of reads. This genome-free alignment allows for agnostic extraction of informative CpG sites. We selected sites (target subsequences of interest) with a minimum CG / TG content, particularly 20% or more, and preferably at least 5% CG, to increase, preferably ensure, the likelihood that the CG position or site of interest exhibits methylation. This selection was performed on healthy samples to avoid bias associated with cancer hypomethylation. We searched 33 of the 34 CpG positions covered by the patient panel for the L1HS consensus sequence. The 34th CpG was absent in copies belonging to the L1PA6-8 family, which accounted for 32% of the hits obtained. Therefore, in certain instances, CG / TG signals were obtained below the 20% threshold used to identify CpG sites.

[0320] Overall, we performed an unbiased method to search for methylation sites contained within the youngest LINE-1 elements present in the human genome without relying on the current genomic annotation, which remains scarce for the repeat sequences of interest. The method disclosed herein allowed us to study the levels and patterns of DNA methylation, particularly cfDNA methylation, both globally and at each CpG site.

[0321] L1PA hypomethylation is detectable in plasma DNA in multiple forms of cancer (Figure 2). We validated our method on cancer cell lines, healthy tissues, and tumor tissues and observed statistically significant L1PA hypomethylation in ovarian and breast tumors compared with healthy plasma samples and healthy tissue samples collected adjacent to ovarian tumors (Figure 2A). Next, we tested plasma samples from colon and ovarian cancers, where significant proportions of L1 hypomethylation have previously been reported. We detected highly statistically significant L1PA hypomethylation in cfDNA from metastatic stage colon cancer (CRC_M+) and stage III / IV ovarian cancer (OVC_M0 / M+, 80% of which are stage II). We also detected highly statistically significant L1PA hypomethylation in metastatic stage breast cancer (BRC_M+) and metastatic stage uveal melanoma (UVM_M+), as well as early stage gastric cancer (GAC_M0) (Figure 2B). Hypomethylation was less pronounced in metastatic non-small cell lung cancer (NSCLC_M+) and early-stage breast cancer (BRC_M0). Indeed, focusing on global methylation levels provides only partial information. We further calculated the methylation level at each CpG site of the target L1HS sequence (n = 33). We observed specific patterns of methylation along the L1 structure across the various amplicons (i.e., subsequences of interest) (Figure 2C). Overall, these patterns were robust across the 123 healthy plasma samples tested. As expected, the 5' portion of L1 was highly methylated, particularly within the second amplicon, but portions of the final CG target, amplicon #8, were also highly methylated (62% in healthy samples). We observed that the most highly methylated sites were the central five CpGs of amplicon #2, which were flanked by two CpG sites with low methylation levels. However, this was not necessarily the site that showed the most significant difference in methylation levels between healthy and cancer samples (Figure 2D). We observed that all CpG sites showed different patterns in one or more cancer subtypes, and different CpGs may be informative for different groups.For example, the first two CpGs of amplicon #3 were highly significant in metastatic colorectal cancer (CRC_M+) but not in metastatic breast cancer (BRC_M+), suggesting that the hypomethylation patterns of L1PA differ between cancer types.

[0322] L1PA hypomethylation-based classifier distinguishes cancer samples from healthy donors across multiple forms of cancer (Figure 3) We trained a classification model using the random forest algorithm based on these 33 CpG sites corresponding to the methylation levels of each CpG target and evaluated its ability to distinguish between healthy and tumor plasma. Methylation of the L1PA element demonstrated excellent ability to distinguish between healthy and tumor plasma from six types of cancer, with an area under the curve (AUC) of 0.95 (95% confidence interval (CI) = 0.92-0.98, Figure 3A). This model performed extremely well for metastatic stages of colon and breast cancer (AUC_CRC_M+ = 0.99; 95% CI = 0.99-1.00; AUC_BRC_M+ = 0.99; 95% CI = 0.99-1.00), as well as for stage III / IV ovarian cancer and non-metastatic gastric cancer (AUC_OVC_M0 = 0.99; 95% CI = 0.98-1.00; AUC_GAC_M0 = 0.98; 95% CI = 0.93-1.00) (Figure 3B). Further superior performance is observed for metastatic lung cancer and uveal melanoma (AUC_NSCLC_M+=0.97; 95% CI=0.92-1.00; AUC_UVM_M+=0.96; 95% CI=0.92-0.99), and more importantly, for early stages of breast cancer (AUC_BRC_M0=0.95; 95% CI=0.89-1.00).

[0323] We also developed an approach to integrate methylation haplotypes at the single-molecule level, which corresponds to the true methylation pattern of neighboring CpGs detected for each amplified molecule. Based on the combination of 33 target CpGs, we were able to extract a total of 274 unique haplotypes from the 33 CpG sites. Here, we wanted to see whether classification could be improved by more detailed signals. We observed similar results for the methylation levels at each CpG site (Figures 3C-3D). However, in the case of ovarian cancer, the haplotype-based model was even better and more robust (Figure 3E).

[0324] Specific performance in cancer samples is reproducible across independent cohorts (Figure 4) To validate our cfDNA L1 targeted bisulfite sequencing approach, we tested a second, independent cohort consisting of 30 healthy donors and 160 patients with the same cancer types as Cohort 1, excluding uveal melanoma (Figure 4A). First, we compared the methylation patterns at each CpG site along the L1 structure and observed very good overall reproducibility (Figure 4B). We compared global methylation levels, excluding amplicon #3, and observed very similar distributions for each cancer type between the two cohorts (Figure 4C). Next, we tested the classifier trained on the first cohort on this new, independent sample set. When testing all cancers without annotating different histologies, we observed very good classification results, with an overall AUC of 0.99 (95% CI = 0.99-1.00, Figure 4D). We observed that excluding CpG sites in amplicon 3 improved performance. This was true across all classification criteria, across all cancers, and across cancer types, particularly lung and ovarian cancers (Figure 9). Indeed, in adjusted analyses, we achieved very high performance for each cancer type (AUC_CRC_M+ = 0.99; 95% CI = 0.99-1.00; AUC_BRC_M+ = 0.99; 95% CI = 0.99-1.00; AUC_NSCLC_M+ = 0.96; 95% CI = 0.94-0.98; AUC_OVC_M0M+ = 0.99; 95% CI = 0.99-1.00; AUC_GAC_M+ = 0.99; 95% CI = 0.99-1.00) (Figure 4E). We further observed excellent sensitivity with 99% specificity, again improved by excluding CpG sites in amplicon 3 (Figure 4F), demonstrating the robustness of detecting hypomethylation in L1PA elements from cfDNA for cancer detection.

[0325] The origin of cancer can be directly inferred from the methylation status of L1PA (Fig. 5). We further analyzed whether we could identify the origin of cancer, in addition to disease status, using a "multi-class" learning model. For this classification, we provided cancer type annotations in the training set and tested multiple possible cancer classes corresponding to each cancer type. We first tested this approach using the methylation levels of a single CpG and only on classes that were homogeneous between Cohort 1 and Cohort 2, namely, healthy donors, CRC_M+, BRC_M+, and OVC_M0M+. Notably, we observed that the multiclass model detected the correct cancer type at performance far above chance (AUC_HD = 0.98, 95% CI = 0.94-1.00; AUC_CRC_M+ = 0.83, 95% CI = 0.80-0.85; AUC_BRC_M+ = 0.72, 95% CI = 0.67-0.76; AUC_OVC_M0M+ = 0.73, 95% CI = 0.68-0.78) (Figure 5A). We also observed that in this case, using bootstrapping and including amplicon 3 tended to improve performance for breast and ovarian cancers (Figure 5A vs. Figures 5B and 5E), and we also observed very similar results for the prototype data (Figures 5C-E). We also tested a multiclass model that also included lung and stomach cancer groups (data not shown). Overall, the methylation profile of L1PA elements provides information about the origin of cancer.

[0326] Comparison with existing methods (Figure 8) Surprisingly, determining L1PA methylation profiles with the aid of the present invention allows for the identification of tumor plasma with substantially higher accuracy or performance than methods based on mutation detection. In comparison, the identification of the same tumor samples via clinically used methods that detect frequently occurring mutations does not exceed 59% for ovarian cancer (compared to 95% in the context of the present invention), 38% for colon cancer (compared to 98% in the context of the present invention), or 52% for metastatic breast cancer (compared to 95% in the context of the present invention) (Figure 8A). The inventors also achieved a remarkable performance of 94% detection rate in a cohort of 27 early gastric cancers, compared to 12.5% ​​for mutation screening.

[0327] Additionally, the method of the present invention achieves similar or higher levels of sensitivity for five cancer types compared to the Galleri® test (Klein, EA et al. Annals of Oncology 32, 1167-1177 (2021)) (Figure 8B). In contrast to the method of the present invention, which targets repetitive sequences, e.g., retrotransposons, the Galleri® method is a capture-based method that targets 100,000 uniquely mappable regions (as opposed to repetitive targets, which are elements dispersed throughout the genome and can be mapped to specific regions in the genome). Additionally, the method developed by the present inventors only requires targeting approximately 82-125,000 CpG sites, 10-fold fewer than the 1,100,000 CpG sites targeted by the Galleri® test. The most original and competitive advantage of the method disclosed herein by the present inventors is its ability to investigate DNA methylation, particularly cfDNA methylation.

[0328] Discussion In this study, we established robust evidence that targeting transposon hypomethylation from cell-free DNA is a sensitive and specific biomarker for noninvasively detecting multiple forms of cancer. We developed a "turnkey" assay capable of identifying tumor plasma and quantifying tumor burden. We investigated selected repetitive regions that provide genome-wide information because half of the CpG sites present in the human genome are recurrently conserved. This novel method targets LINE-1 retrotransposon hypomethylation, a common feature across multiple forms of cancer, to capture a wide range of tumor alleles and cover the heterogeneous profiles of cancer patients in a single test. A cfDNA L1-targeted bisulfite sequencing assay provided a surrogate indicator of genome-wide DNA methylation levels. This enabled us to generate highly accurate and comprehensive methylation profiles from minute amounts of cfDNA (down to a few nanograms) using affordable sequencing depth. Therefore, we expect this method to be widely applicable for the development of routine clinical tests. The most original and competitive feature of this study is its investigation of cfDNA methylation in repetitive sequences. Repeat hypomethylation is common in many, if not all, cancer types and is a promising marker for pan-cancer detection. Previous studies have neglected these regions because they are inherently difficult to map, and differentially methylated region (DMR) analysis has typically been performed on mapped data. We developed a novel method to detect methylation profiles in repeats at single-base-pair resolution without mapping to a reference genome. This allows us to retain most of the data produced, achieving high sensitivity and helping us handle trace amounts of cfDNA. The results disclosed herein demonstrate high performance in detecting cancer samples, and we have established its feasibility in six cancer types, including three early-stage cancers.

[0329] Overall, this assay targets approximately 82-125,000 CpG sites, which is 10-fold less than the 1,100,000 CpG sites targeted by the existing Galleri test. Meanwhile, we achieved similar or higher levels of sensitivity in four of the five cancers tested across both studies (Figure 8B). For example, we achieved 94% sensitivity (99% specificity) for early-stage gastric cancer, compared with 47% for the Galleri test. We also achieved 73% sensitivity (99% specificity) for early-stage breast cancer, compared with only 28% for the Galleri test (Figure 8B). This was also in the range of sensitivity achieved by the CancerSeek test (for BRC_M0, sensitivity was less than 40%) (Cohen, JD et al. Science 1, eaar3247-10 (2018)), which investigates a panel of mutations in 16 genes linked to 8 proteins. This highlights that breast cancer is one of the most difficult cancers to detect by liquid biopsy and that our approach could be a game changer for non-invasive early detection of breast cancer.

[0330] The inventors were also able to demonstrate that the methylation status of L1PA detected in plasma DNA allows the origin of the detected cancer to be inferred.< / referencefasta> < / inputfasta> < / referencefasta> < / inputfasta>

Claims

1. 1. A method for determining the CpG methylation profile of at least one DNA sequence of interest or any fragment thereof, comprising: a) clustering a set of subsequences obtained from a DNA sequence of interest into clusters of subsequences; b) for each cluster and from each cluster, selecting one subsequence from among the subsequences of said cluster as a reference sequence; c) aligning the reference sequences of said clusters by aligning at CpG dinucleotide positions; d) aligning the remaining subsequences to the selected reference sequence; e) determining the CpG methylation status of each subsequence by determining whether the CpG dinucleotide at each CpG site of the subsequence is methylated or not, thereby determining a CpG methylation profile of the subsequence, including the CpG methylation level and / or the proportion of CpG methylated haplotypes; Including, wherein the DNA sequence of interest is or consists of repetitive sequences, said repetitive sequences being distributed throughout the genome of the subject and preferably containing a high density of CpG dinucleotides; wherein the method optionally comprises a first step of obtaining or providing a set of subsequences of said DNA sequence of interest, wherein the method optionally includes repeating some or each of steps a) through e) for another set of subsequences from the DNA sequence of interest; method.

2. 2. The method of claim 1, wherein the repetitive sequence is a retrotransposon, such as LINE, HERV, SINE, SVA or a subfamily thereof, such as, in particular, LINE-1, L1PA, HERV-K and Alu, or a satellite repeat, such as a Sat2 or Sat3 element, preferably a LINE-1 retrotransposon or any fragment or variant thereof, more preferably a LINE-1 retrotransposon as set forth in SEQ ID NO: 2 or 29 or any fragment or variant thereof.

3. 1. A computer-implemented method for training a classifier to accurately distinguish between healthy and cancerous CpG methylation profiles, comprising: a) providing as input to a classifier a training set of CpG methylation profiles of DNA sequences of interest or subsequences thereof, wherein the DNA sequences of interest are repeated and distributed throughout the genome and contain a high density of CpG dinucleotides, or preprocessed information obtained from the training set of CpG methylation profiles of DNA sequences of interest or subsequences thereof, wherein the training set of CpG methylation profiles includes CpG methylation profiles of DNA sequences or subsequences thereof from subjects identified as healthy subjects and subjects identified as cancer subjects; b) generating an output of the classifier for each CpG methylation profile input of the DNA sequence of interest or a subsequence thereof, said output classifying the CpG methylation profile input of the DNA sequence of interest or a subsequence thereof as a healthy CpG methylation profile or a cancerous CpG methylation profile. Including, wherein the CpG methylation profile comprises the CpG methylation level and / or the proportion of CpG methylated haplotypes of a DNA sequence or a subsequence thereof; method.

4. 4. The method of claim 3, wherein the CpG methylation profile of a DNA sequence of interest or a subsequence thereof is determined by the method of claim 1 or 2.

5. 1. An in vitro or in silico method for determining the health status of a subject, in particular determining whether a subject is a healthy subject or a subject suffering from cancer or cancer recurrence, comprising: a) providing a DNA sequence of interest or a subsequence thereof from a subject, or pre-processed information obtained from said DNA sequence or subsequence, said DNA sequence of interest being a DNA sequence that is distributed throughout the subject's genome and encodes a repetitive sequence containing a high density of CpG dinucleotides, as input to a classifier trained to distinguish between healthy and cancerous CpG methylation profiles; b) using the classifier to identify the CpG methylation profile of the subject's DNA sequence of interest or a subsequence thereof as an output of the classifier as a healthy CpG methylation profile or a cancerous CpG methylation profile. method.

6. 1. An in vitro or in silico method for determining the origin of a tumor from a subject, comprising: a) providing a DNA sequence of interest or a subsequence thereof from a subject, or pre-processed information obtained from said DNA sequence or subsequence, said DNA sequence of interest being a DNA sequence that is distributed throughout the subject's genome and encodes a repetitive sequence containing a high density of CpG dinucleotides, as input to a classifier trained to distinguish between healthy CpG methylation profiles and cancerous CpG methylation profiles from various tumor origins; b) using the classifier to identify the CpG methylation profile of the subject's DNA sequence of interest or a subsequence thereof as an output of the classifier as a healthy CpG methylation profile or a cancerous CpG methylation profile from a particular tumor origin, thereby determining the origin of the tumor from the subject. method.

7. 1. An in vitro or in silico method for determining the stage of a tumor from a subject, comprising: a) providing a DNA sequence of interest or a subsequence thereof from a subject, or preprocessed information obtained from said DNA sequence or subsequence, said DNA sequence of interest being a DNA sequence that is distributed throughout the subject's genome and encodes a repeat sequence containing a high density of CpG dinucleotides, as input to a classifier trained to distinguish between healthy CpG methylation profiles and various stages of cancerous CpG methylation profiles; b) using the classifier to identify the CpG methylation profile of the subject's DNA sequence of interest or a subsequence thereof as an output of the classifier as a healthy CpG methylation profile or a cancerous CpG methylation profile of a particular stage, thereby determining the stage of a tumor from the subject. method.

8. 1. An in vitro or in silico method for monitoring the response of a subject suffering from cancer to an anti-cancer treatment, comprising: a) providing a classifier trained to distinguish between DNA sequences having a healthy CpG methylation profile and DNA sequences having a cancerous CpG methylation profile, as a first input, at least one DNA sequence of interest or a subsequence thereof from a first liquid biopsy from a subject suffering from cancer prior to administration of an anti-cancer treatment to the subject, said DNA sequence of interest being repeated throughout the subject's genome and comprising a high density of CpG sites or fragments thereof, or pre-processed information obtained from said first liquid biopsy, and as a second input, a second liquid biopsy comprising at least one DNA sequence of interest or a subsequence thereof from said subject after administration of an anti-cancer treatment, or pre-processed information obtained from said second liquid biopsy; b) using the classifier to identify each CpG methylation profile of each DNA sequence of the first liquid biopsy as having a healthy CpG methylation profile or a cancerous CpG methylation profile as a first output of the classifier, and to identify each CpG methylation profile of each DNA sequence of the second liquid biopsy as having a healthy CpG methylation profile or a cancerous CpG methylation profile as a second output of the classifier; wherein the number of DNA sequences of interest classified as having a healthy CpG methylation profile in the second output of the classifier exceeds the number of DNA sequences of interest classified as having a healthy CpG methylation profile in the first output of the classifier, indicating that the subject will be responsive to the anti-cancer treatment; and wherein the number of DNA sequences of interest classified as having a healthy CpG methylation profile in the second output of the classifier is less than or equal to the number of DNA sequences of interest classified as having a healthy CpG methylation profile in the first output of the classifier, indicating that the subject will not be responsive to the anti-cancer treatment. method.

9. 1. An in vitro or in silico method for assessing the efficacy of a compound to revert a cancerous CpG methylation profile of a DNA sequence of interest from a subject suffering from cancer to a healthy CpG methylation profile, comprising: a) providing a DNA sequence of interest or a subsequence thereof from a subject treated with a compound, said DNA sequence of interest being repeated and distributed throughout the subject's genome and containing a high density of CpG dinucleotides or any fragment thereof, or pre-processed information obtained from said at least one DNA sequence of interest or a subsequence thereof, as input to a classifier trained to distinguish between DNA sequences with a healthy CpG methylation profile and DNA sequences with a cancerous CpG methylation profile; b) using the classifier to detect DNA sequences with healthy CpG methylation profiles and / or DNA sequences with cancerous CpG methylation profiles as an output of the classifier; wherein the amount of DNA sequences with a healthy methylation profile exceeds a reference amount of DNA sequences with a healthy methylation profile obtained from the subject before any treatment with the compound, indicating that the compound is capable of reverting the cancerous CpG methylation profile to a healthy CpG methylation profile. method.

10. 1. An in vitro or in silico method for predicting the ability of a compound to treat cancer, comprising: assessing the efficacy of the compound to revert a cancerous CpG methylation profile of a DNA sequence of interest from a subject of claim 9 to a healthy CpG methylation profile; wherein an amount of DNA sequences classified as having a healthy CpG methylation profile above a reference amount indicates that the compound is useful for treating the cancer. method.

11. The method according to any one of claims 5 to 10, wherein the CpG methylation profile of the DNA sequence of interest or a subsequence thereof is determined by the method according to any one of claims 1 to 4.

12. The method according to any one of claims 5 to 11, wherein the classifier is trained according to the method according to claim 3 or 4.

13. 13. The method of any one of claims 1 to 12, wherein the DNA sequence of interest is a circulating cell-free DNA (cfDNA) sequence.

14. a memory storing at least one instruction of a classifier trained according to the method of any one of claims 3, 4 and 12; a processor accessing said memory to read said instructions and to execute the method of any one of claims 5 to 13, Computer system.

15. A kit of primers or probes targeting a partial sequence of a DNA sequence encoding a LINE-1 retrotransposon, preferably a LINE-1 retrotransposon as set forth in SEQ ID NO: 2 or 29, The kit comprises at least four primers or probes selected from the group of primers or probes having the sequences shown in SEQ ID NOs: 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or 26, respectively, or sequences having at least 85% identity thereto; kit.

16. 16. Use of the kit according to claim 15 for amplifying a partial sequence of a DNA sequence encoding a LINE-1 retrotransposon, preferably a LINE-1 retrotransposon as set forth in SEQ ID NO: 2 or 29, for the diagnosis of cancer, Preferably, the cancer is selected from the group consisting of colon cancer, breast cancer, lung cancer, uveal melanoma cancer, ovarian cancer and gastric cancer. use.