Whole genome repeat landscape in cancer and cell-free DNA

The ARTEMIS method, by analyzing k-mers in nucleic acid sequences, identifies and assesses various repetitive sequence types across the entire genome, addressing the shortcomings of existing technologies in identifying cancer-related genomic changes and enabling precise cancer diagnosis and personalized treatment.

CN122074154APending Publication Date: 2026-05-22JOHNS HOPKINS UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JOHNS HOPKINS UNIVERSITY
Filing Date
2024-08-14
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing technologies are insufficient for effectively analyzing and monitoring complex repetitive sequences in the human genome, especially in cancer-related genomic changes, leading to inadequate research on these key regions.

Method used

A method called ARTEMIS is provided that identifies and evaluates various repetitive sequence types across the entire genome, including satellite DNA, RNA elements, LINE, SINE, LTR, etc., by analyzing k-mers in nucleic acid sequences. It then uses machine learning and epigenetic analysis to generate scores for cancer diagnosis and treatment.

Benefits of technology

It enables precise diagnosis and monitoring of cancer-related genomic changes, provides new methods for cancer diagnosis and treatment, improves the ability to identify repetitive sequence changes, and supports personalized treatment decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122074154A_ABST
    Figure CN122074154A_ABST
Patent Text Reader

Abstract

Compositions and methods for analyzing repetitive sequences in the genome are provided, which are used in methods for large scale analysis of these regions for characterization, detection, and monitoring of human cancer.
Need to check novelty before this filing date? Find Prior Art

Description

This application claims the benefit of U.S. Provisional Application No. 63 / 532,642, filed August 14, 2023, the entire contents of which are incorporated herein by reference.

[0001] Statement on Federally Funded Research This invention was developed with government support from grants granted by the National Institutes of Health under the names CA006973, CA062924, CA121113, CA233259, CA271896, and GM136577. The government holds certain rights to this invention. Technical Field

[0002] Compositions for analyzing repetitive sequences in the genome are used in methods for large-scale analysis of these regions to characterize, detect, and monitor human cancer. Background Technology

[0003] Repetitive sequences make up more than half of the human genome and contain a wide variety of elements that vary greatly among individuals and have a crucial impact on genome structure and function. 1,2 Due to technical limitations of short read alignment and reliance on incomplete genome assembly, repetitive sequences have historically been overlooked. 3 These repetitive sequences include tandem repeats, such as human satellite DNA sequences, and sporadic repeats, including structural RNA, long sporadic nuclear elements (LINEs), short sporadic nuclear elements (SINEs), long terminal repeats (LTRs), and other transposon elements. The recently completed telomere-to-telomere genome, adding nearly 200 Mb to the previous reference genome, revealed the genomic and epigenomic status of repetitive sequences and revitalized research on these essential genomic regions. 4-6 . Summary of the Invention

[0004] A genome-wide pathway for analyzing repetitive sequence landscapes in next-generation sequencing is provided. This pathway is referred to in this paper as ARTEMIS. A nalysis of R epea T El e M The analysis of duplication elements in dlSease (disease-related repetitive elements) can assess thousands of unique duplication types that occur across the entire genome and span multiple subfamilies, including multiple families (satellite DNA, RNA elements, transposable elements, LINE, SINE, LTR).

[0005] In some respects, read lengths can be, for example, less than 800, 700, 600, 500, 400, or 300 bp. In other respects, read lengths can be, for example, 50 to 400 bp, or 50 to 300 or 250 bp, or 75 to 300 bp, or 75 to 250 bp.

[0006] In some respects, the aforementioned approaches and methods may be independent of comparison.

[0007] In one aspect, a method for identifying a type of repeating element is provided, comprising: extracting repeating sequences and coordinates from known repeating element types; identifying nucleic acid sequences (k-mers) in nucleic acid sequences including genomic or cell-free DNA; selecting k-mers appearing in a single repeating type and identifying unique k-mers of the repeating element type; wherein the unique k-mers identify one or more repeating element types.

[0008] In one aspect, a method for identifying repeat element types is provided, comprising: extracting repeat sequences and coordinates from known repeat element types; identifying nucleic acid sequences (k-mers) in an RNA sequence; selecting k-mers appearing in a single repeat type and identifying unique k-mers of the repeat element type; wherein the unique k-mers identify one or more repeat element types.

[0009] In one aspect, a method for identifying repeat element types is provided, comprising: a) extracting repeat sequences and genomic coordinates from known repeat element types; b) selecting nucleic acid sequences (k-mers) appearing in a single repeat type and identifying unique k-mers for the repeat element type, wherein the unique k-mers identify one or more repeat element types; and c) identifying k-mers in nucleic acid sequences comprising genomic or cell-free DNA. In some aspects, the type and / or frequency of k-mers in nucleic acid sequences comprising genomic or cell-free DNA can be identified.

[0010] In one aspect, a method for identifying repeat element types is provided, comprising: a) extracting repeat sequences and genomic coordinates from known repeat element types; b) selecting nucleic acid sequences (k-mers) appearing in a single repeat type and identifying unique k-mers for the repeat element type, wherein the unique k-mers identify one or more repeat element types; and c) identifying k-mers in genomic or cell-free DNA. In some aspects, the type and / or frequency of k-mers in genomic or cell-free DNA can be identified.

[0011] In some implementations, repeating element types are excluded from families containing low-complexity, unknown, simple repeating sequences or combinations thereof.

[0012] In some implementations, elements from each family are aggregated, said families including tRNA, srpRNA, snRNA, scRNA, rRNA, RNA elements, DNA, retrotran, retrotransposon, or combinations thereof.

[0013] In some embodiments, the family includes long interspersed nuclear elements (LINE), short interspersed nuclear elements (SINE), long terminal repeat sequences (LTR), satellite DNA, transposon elements, RNA elements, or combinations thereof. In some embodiments, k-mers that appear in a single repeat type element but not in a non-repeat region are selected.

[0014] In some embodiments, the k-mer comprises up to or about 5, 8, 10 to 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, or 200 nucleotides. In some embodiments, the k-mer comprises 10 to 40, 50, 60, 70, 80, 90, or 100 nucleotides. In some embodiments, the k-mer comprises 15 to 35 nucleotides. In some embodiments, the k-mer comprises 20 to 30 nucleotides. In some embodiments, the k-mer comprises 22 to 28 nucleotides. In some embodiments, the k-mer comprises 23 to 26 nucleotides. In some embodiments, the k-mer comprises about 24 nucleotides.

[0015] In some implementations, a landscape of k-mer repetitive sequences is generated, each unique k-mer and its inverse complementary sequence are counted, and the k-mer counts for each type of repetition element are summed. In some implementations, gene regions containing multiple transcripts are counted to identify k-mers present in each gene.

[0016] In some embodiments, genes are sorted by total k-mer count and corrected k-mer density. In some embodiments, the corrected k-mer density comprises the number of types appearing in a gene divided by the k-mer density. In some embodiments, the k-mer density is the total number of k-mers per Mb.

[0017] In some implementations, the types of repetitive elements in cancer diagnosis-related genes are increased compared to non-cancer controls. In some implementations, the families of repetitive element types defined by k-mers are altered across the entire genome compared to normal controls. In some implementations, changes in the genome compared to a non-cancer genome include focal amplifications, deletions, copy number changes, rearrangements, repetitive element alterations, or combinations thereof.

[0018] In another aspect, a method for diagnosing and treating cancer in a subject is provided, suitably comprising: generating a k-mer repeat sequence landscape from a biological sample of the subject, analyzing changes in repeat elements, said analysis including selecting normal samples, calculating the ratio of k-mer repeat sequence landscapes between samples to diagnose the subject having cancer, and treating the subject diagnosed with cancer.

[0019] In some implementations, a subject is diagnosed with cancer if the k-mer repeat sequence landscape between two samples contains a tumor / normal ratio below the 1st percentile or above its 99th percentile of the normal / normal ratio, wherein the normal / normal ratio is determined by computer simulation (…). in silico The samples were obtained through secondary sampling from normal samples, with each subsample containing approximately half the coverage of the original sample. In some implementations, changes in the tumor / normal and normal / normal ratios are correlated with one or more indicators of genomic instability.

[0020] In some embodiments, the one or more genomic instability indicators include entropy, nonmodal ploidy fraction, nondiploidy fraction, loss of heterozygosity fraction, number of breakpoints, tumor mutational burden, ploidy, modal ploidy, or a combination thereof.

[0021] In some embodiments, the method further includes repeating element types with overlapping k-mers and detecting the difference in the total k-mer count of repeating element types between tumors with and without focal amplification.

[0022] In some implementations, the k-mer repetitive sequence landscape is fed into a machine learning or artificial intelligence program to generate cross-validation or external validation scores for distinguishing tumor samples from normal samples.

[0023] In some embodiments, the biological sample comprises genomic DNA, cell-free DNA (cfDNA), or a combination thereof. In some embodiments, the epigenetic characteristics, fragment length, and fragment coverage differences in histone marker regions of the cfDNA fragments are assessed. In some embodiments, the localization of histone markers within repeat element types and the density of each histone marker are assessed. In some embodiments, based on histone marker density, the ratio of observed k-mer counts to expected k-mer count variations between repeat element types is assessed.

[0024] In some implementations, cancer treatments or therapies include surgery, adjuvant chemotherapy, neoadjuvant chemotherapy, radiation therapy, hormone therapy, cytotoxic therapy, immunotherapy, adoptive T-cell therapy, targeted therapy, and combinations thereof.

[0025] In another aspect, a method for detecting and monitoring cancer progression in subjects is provided, and may suitably include: detecting multiple features of a subject sample to generate a k-mer repetitive sequence landscape; further classifying the features into families comprising long interspersed nuclear elements (LINEs), short interspersed nuclear elements (SINEs), long terminal repeats (LTRs), satellite DNA, transposable elements, RNA elements, or combinations thereof, defining multiple megabase (Mb) bins via epigenetic analysis; calculating the alignment fragment coverage for each bin, wherein bins with a mappability less than 0.9 or a GC content less than 0.3 are excluded; training a regression model using the repetitive sequence landscape and epigenetic profile, and integrating the scores together using the regression model to generate a comprehensive score. In some embodiments, the method further includes incorporating the fragmentation and the scores obtained through one or more software models, machine learning models, artificial intelligence, or combinations thereof.

[0026] In some respects, the methods and materials described in this paper may also include machine learning.

[0027] In some aspects, a method is provided for determining whether a subject has responded to treatment, and may suitably include any one or more methods embodied herein. In some embodiments, the treatment includes surgery, adjuvant chemotherapy, neoadjuvant chemotherapy, radiation therapy, hormone therapy, cytotoxic therapy, immunotherapy, adoptive T-cell therapy, targeted therapy, and combinations thereof.

[0028] definition Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. It will be further understood that terms such as those defined in common dictionaries should be interpreted as having meanings consistent with their meanings in the context of the relevant art, and should not be interpreted in an idealized or overly formal sense unless specifically defined herein.

[0029] As used herein, the singular forms “a,” “an,” and “the” are intended to also include the plural forms, unless the context clearly indicates otherwise. Furthermore, with regard to the use of the terms “including,” “includes,” “having,” “has,” “with,” or variations thereof in the specification and / or claims, such terms are intended to be inclusive in a manner similar to the term “comprising.”

[0030] The terms “about” or “approximately” mean within an acceptable margin of error for a particular value, as determined by one person skilled in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, according to practice in the art, “about” may mean within one standard deviation or more than one standard deviation. Alternatively, “about” may mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value or range. Or, particularly with respect to biological systems or processes, the term may mean within five orders of magnitude or even two orders of magnitude of the numerical value. When a particular value is described in an application and claims, unless otherwise stated, the term “about” should be assumed to mean within an acceptable margin of error for that particular value.

[0031] As used herein, the term "cancer" means any disease, condition, trait, genotype, or phenotype known in the art as characterized by uncontrolled cell growth or replication; including liver cancer (including hepatocellular carcinoma (HCC)), lung cancer (including non-small cell lung cancer), gastric cancer, colorectal cancer, and leukemias such as acute myeloid leukemia (AML), chronic myeloid leukemia (CML), acute lymphoblastic leukemia (ALL), and chronic lymphocytic leukemia; AID-related cancers such as Kaposi's sarcoma; breast cancer; bone cancers such as osteosarcoma, chondrosarcoma, Ewing's sarcoma, fibrosarcoma, giant cell tumor, ameloblastoma, and chordoma; and brain cancers such as meningioma, glioblastoma, and low-grade astrocytoma. Oligodendroglioma, pituitary adenoma, schwannoma, and metastatic brain cancer; head and neck cancers including various lymphomas such as mantle cell lymphoma, non-Hodgkin lymphoma, adenoma, squamous cell carcinoma, laryngeal cancer, gallbladder cancer and bile duct cancer; retinal cancers such as retinoblastoma, esophageal cancer, gastric cancer, multiple myeloma, ovarian cancer, uterine cancer, thyroid cancer, testicular cancer, endometrial cancer, melanoma, bladder cancer, prostate cancer, pancreatic cancer, sarcoma, nephroblastoma, cervical cancer, head and neck cancer, skin cancer, nasopharyngeal carcinoma, liposarcoma, epithelial carcinoma, renal cell carcinoma, gallbladder adenocarcinoma, parotid gland adenocarcinoma, endometrial sarcoma, multidrug-resistant cancers; and proliferative diseases and conditions such as neovascularization associated with tumor angiogenesis.

[0032] The terms "cell-free nucleic acid," "cell-free DNA," or "cfDNA" refer to nucleic acid fragments that circulate within an individual's body (e.g., in the bloodstream) and originate from one or more healthy cells and / or one or more cancer cells. Furthermore, cfDNA may originate from other sources, such as viruses, fetuses, etc.

[0033] The term "cfDNA sequence coverage" refers to the average number of cfDNA molecules that overlap with a specific site.

[0034] The term "circulating tumor DNA" or "ctDNA" refers to nucleic acid fragments originating from tumor cells or other types of cancer cells that may enter an individual's bloodstream due to biological processes such as apoptosis or necrosis of dying cells or by being actively released by living tumor cells.

[0035] As used herein, when referring to the defined or described elements of articles, compositions, apparatus, methods, processes, systems, etc., the terms “comprising,” “comprise,” or “comprised,” and their variations, mean inclusive or open-ended, allowing for additional elements, thereby indicating that the defined or described articles, compositions, apparatus, methods, processes, systems, etc., include those specified elements—or, where appropriate, their equivalents—and may include other elements, still falling within the scope / definition of the defined articles, compositions, apparatus, methods, processes, systems, etc.

[0036] "Diagnostic" or "diagnosed" means identifying the presence or nature of a pathological condition. Diagnostic methods differ in their sensitivity and specificity. The "sensitivity" of a diagnostic test is the percentage of diseased individuals who test positive ("true positive" percentage). Diseased individuals not detected by the test are "false negatives." Subjects who are not diseased and test negative in the test are called "true negatives." The "specificity" of a diagnostic test is 1 minus the false positive rate, where the "false positive" rate is defined as the percentage of diseased individuals who test positive. While a particular diagnostic method may not provide a definitive diagnosis of the condition, it is sufficient if the method provides positive indications that aid in diagnosis.

[0037] As used in this article, "effective amount" means the amount that provides therapeutic or preventative benefits.

[0038] As used herein, the terms “fragment profile,” “position-dependent differences in fragment patterns,” and “differences in fragment size and coverage in a position-dependent manner across the entire genome” are equivalent and can be used interchangeably. In some embodiments, determining the cfDNA fragment profile in a mammal can be used to identify mammals with cancer. For example, low-coverage whole-genome sequencing can be performed on cfDNA fragments obtained from a mammal (e.g., samples obtained from a mammal), and the sequenced fragments can be mapped to the genome (e.g., within a non-overlapping window) and evaluated to determine the cfDNA fragment profile. As described herein, the cfDNA fragment profile of mammals with cancer is more heterogeneous (e.g., in terms of fragment length) than that of healthy mammals (e.g., mammals without cancer). Therefore, this disclosure also provides methods and materials for evaluating, monitoring, and / or treating mammals (e.g., humans) with or suspected of having cancer. In some embodiments, this document provides methods and materials for identifying mammals with cancer. For example, samples obtained from a mammal (e.g., blood samples) can be evaluated to determine the presence of cancer in the mammal based at least in part on the mammal's cfDNA fragment profile, and optionally to determine the primary tissue of the cancer. In some embodiments, methods and materials are provided for monitoring mammals with cancer. For example, samples obtained from mammals (e.g., blood samples) can be evaluated to determine the presence of cancer in the mammal based at least in part on the mammal's cfDNA fragment profile. In some embodiments, methods and materials are provided for identifying mammals with cancer and administering one or more cancer treatments to the mammal to treat the mammal. For example, samples obtained from mammals (e.g., blood samples) can be evaluated to determine whether the mammal has cancer based at least in part on the mammal's cfDNA fragment profile, and one or more cancer treatments can be administered to the mammal.

[0039] The term "ensemble learning" refers to an algorithm that combines predictions from two or more models. An "ensemble method" is a machine learning technique that combines multiple base models to produce an optimal predictive model. Widely used ensemble learning strategies include bagging, stacking, and boosting. Bootstrap aggregation, or simply bagging, is an ensemble learning method that seeks diverse ensemble members by varying the training data. The name bagging comes from the abbreviation of Bootstrap AGGregatING. As the name suggests, the two key elements of bagging are bootstrap and aggregation.

[0040] Bootstrap aggregation (bagging) involves training an ensemble on a bootstrap dataset. The bootstrap set is created by selecting instances with replacement from the original training dataset. Therefore, a bootstrap set may contain instances given zero, one, or more times. Ensemble members can also impose constraints on features (e.g., nodes in a decision tree) to encourage exploration of different features. Variance and feature considerations of local information in the bootstrap set contribute to the diversity of the ensemble and can enhance its performance. To reduce overfitting, members can be validated using an out-of-bag set (instances not in their bootstrap sets).

[0041] Stacking is a general process in which a single learner is trained to combine individual learners. Here, the individual learners are called the first-layer learners, and the combiner is called the second-layer learner or meta-learner. Stacking (sometimes called stacked generalization) involves training a model to combine the predictions of several other learning algorithms. First, all the other algorithms are trained using available data. Then, a combined algorithm (the final estimator) is trained to make the final prediction using the predictions of all the other algorithms (base estimators) as additional input, or using cross-validation predictions from the base estimators (which prevents overfitting). Although logistic regression models are commonly used as combiners in practice, stacking can theoretically represent any ensemble technique if any combination algorithm is used. Stacking typically produces better performance than any single trained model. It has been successfully used for supervised learning tasks (regression, classification, and distance learning) and unsupervised learning (density estimation). The key elements of stacking are: the training dataset remains constant, the machine learning model learns how to optimally combine predictions for each different machine learning algorithm used in the ensemble.

[0042] Boosting ensemble learning is an ensemble approach that seeks to modify the training data to focus attention on instances where previously fitted models misclassified the training dataset. In boosting, the training dataset for each subsequent classifier becomes increasingly focused on instances misclassified by the previously generated classifier. A key feature of boosting ensembles is the idea of ​​correcting prediction errors. Models are fitted and added sequentially to the ensemble, such that the second model attempts to correct the predictions of the first, the third corrects the second, and so on. This typically involves using very simple decision trees that make only one or a few decisions, called weak learners in boosting. The predictions of weak learners are combined through simple voting or averaging, although the contributions are weighted proportionally to their performance or ability. The goal is to develop a so-called "strong learner" from multiple specially constructed "weak learners." Typically, the training dataset remains unchanged, and the learning algorithm is modified to give more or less attention to a particular sample (data row) based on whether it has been correctly or incorrectly predicted by a previously added ensemble member. For example, data rows can be weighted to indicate the degree of attention the learning algorithm must give when learning the model. The key elements of boosting are as follows: biasing the training data towards instances that are difficult to predict, repeatedly adding ensemble members to correct the predictions of previous models, and combining the predictions using a weighted average of the models.

[0043] The term "genomic nucleic acid" or "genomic DNA" refers to nucleic acids, including chromosomal DNA, derived from one or more healthy (e.g., non-tumor) or tumor cells. In various embodiments, genomic DNA can be extracted from cells derived from blood cell lineages, such as white blood cells (WBCs).

[0044] As used herein, a “k-mer” or “k-polymer” refers to a DNA sequence consisting of k consecutive nucleotides (located in the forward or reverse strand of a DNA molecule), where k is a natural integer greater than 1. Any sequence of length L will contain L-k+1 k-mers. In some respects, a k-mer can also be described as a polynucleotide subsequence of length k, including short sequences such as fewer than 200, 150, 100, 90, 80, 70, 60, or 40 bases.

[0045] As used in this article, a “LINE” (long interspersed nuclear element) is a relatively long non-LTR retrotransposon. They are widely distributed in eukaryotic genomes, typically comprising 21.1% of the human genome. Each LINE is approximately 7000 base pairs long. LINEs can be transcribed into mRNA and translated into proteins with reverse transcriptase function. This reverse transcriptase generates DNA copies of the LINE RNA. These DNA copies can integrate into new sites in the genome. The human genome has only one abundant LINE, called LINE-1. LINE-1 elements are approximately 6000 base pairs long. There are approximately 100,000 truncated LINE-1 elements in the human genome. Random mutations can occur in LINEs. Due to random mutations, LINEs degenerate. They can no longer be transcribed or translated. Furthermore, LINEs are divided into five main groups, such as L1, RTE, R2, I, and jockey. These five groups are further subdivided into another 28 clades. LINEs are typically propagated through a mechanism called targeted priming reverse transcription (TPRT). LINE insertions can lead to human diseases such as hemophilia A, cancer, and Mendelian genetic disorders. Hypomethylation of LINE can also trigger certain types of cancer.

[0046] "Optional" or "optionally" means that the event or situation described below may or may not occur, and the description includes examples of where the event or situation occurs and examples of where it does not occur.

[0047] As used in this specification and the appended claims, the term "or" is generally used in its meaning as including "and / or" unless the content clearly indicates otherwise.

[0048] "Parenteral" administration of immunogenic compositions includes techniques such as subcutaneous (sc), intravenous (iv), intramuscular (im), or intrasternal injection or infusion.

[0049] The terms “patient,” “individual,” or “subject” are used interchangeably herein and refer to a mammalian subject to be treated, preferably a human. In some embodiments, the methods of the present invention can be used in laboratory animals, veterinary applications, and the development of animal models of disease, including but not limited to rodents (including mice, rats, and hamsters) and primates.

[0050] As used herein, the term "reference genome" can refer to a digital or previously identified database of nucleic acid sequences assembled into a representative instance of a species or subject. A reference genome can be assembled from nucleic acid sequences from multiple subjects, samples, or organisms and does not necessarily represent the nucleic acid composition of a single person. Reference genomes can be used to map sequencing reads from a sample to chromosomal locations. For example, reference genomes for human subjects, as well as many other organisms, can be found at the National Center for Biotechnology Information (NCBI) website, ncbi.nlm.nih.gov.

[0051] The term “read length” or “read length” refers to any nucleotide sequence, including sequence reads obtained from an individual and / or nucleotide sequences derived from initial sequence reads obtained from samples obtained from an individual.

[0052] The terms “sample,” “patient sample,” “biological sample,” etc., encompass a wide range of sample types obtained from patients, individuals, or subjects and used in diagnostic, prognostic, and / or monitoring assays. Patient samples can be obtained from healthy subjects, patients with illness, or patients with lung cancer. In some embodiments, the “provided” sample may be obtained by the person (or machine) performing the assay, or it may be obtained by another person and transferred to the person (or machine) performing the assay. Furthermore, samples obtained from patients may be aliquoted, and only a portion may be used for diagnosis. Additionally, the sample, or a portion thereof, may be stored under conditions suitable for maintaining the sample for subsequent analysis. This definition specifically covers blood and other liquid samples of biological origin (including, but not limited to, peripheral blood, serum, plasma, cord blood, amniotic fluid, cerebrospinal fluid, urine, saliva, feces, and synovial fluid), and solid tissue samples (such as biopsy specimens or tissue cultures or their derived cells and progeny). In some embodiments, the sample comprises cerebrospinal fluid. In a particular embodiment, the sample comprises a blood sample. In another embodiment, the sample comprises a plasma sample. In yet another embodiment, a serum sample is used. The definition of "sample" also includes samples that have been processed in any way after acquisition, such as by centrifugation, filtration, precipitation, dialysis, chromatography, treatment with reagents, washing, or enrichment of certain cell populations. The term further encompasses clinical samples and also includes cells in cultures, cell supernatants, tissue samples, organs, etc. Samples may also contain freshly frozen and / or formalin-fixed, paraffin-embedded tissue blocks, such as tissue blocks prepared from clinical or pathological biopsies for pathological analysis or immunohistochemical studies.

[0053] The term "sequence read" refers to the nucleotide sequence read from a sample obtained from an individual. Sequencing reads can be obtained using a variety of methods known in the art.

[0054] As used in this article, "SINE" (short, scattered nuclear elements) are a class of much shorter non-LTR retrotransposons. They are approximately 100 to 700 base pairs long. SINEs are also DNA elements that amplify themselves in eukaryotic genomes via RNA intermediates. SINEs comprise approximately 13% of mammalian genomes. The internal regions of SINEs originate from tRNA. They remain highly conserved. They are commonly found in many vertebrate and invertebrate species. Copy number variations and mutations in SINEs can be incorporated to construct phylogenetic-based species classifications. SINEs can be classified into three main types: CORE-SINE, V-SINE, and AmnSINE. Alu elements are the most common SINEs in primates. In addition, more than 50 human diseases are associated with SINE insertion. When they insert into or near exons, they can cause missplicing or alter reading frames. This leads to disease phenotypes such as breast cancer, colon cancer, leukemia, hemophilia, cystic fibrosis, colon cancer, Dent disease, neurofibromatosis, etc.

[0055] As used herein, a “therapeuticly effective” amount (i.e., effective dose) of a compound or reagent means an amount sufficient to produce a therapeutically (e.g., clinically) desired outcome. The composition may be administered once or more daily to once or more weekly, including every other day. Those skilled in the art will understand that certain factors will affect the dosage and timing required for effective treatment of a subject, including, but not limited to, the severity of the disease or condition, prior treatment, the subject’s overall health and / or age, and any other pre-existing conditions. Furthermore, treatment of a subject with a therapeutically effective amount of the disclosed compound may comprise a single treatment or a series of treatments.

[0056] As used herein, the terms “treat,” “treating,” and “treatment” refer to the reduction or improvement of a disease and / or its associated symptoms. It should be understood that, while not excluding, treating a condition does not require the complete elimination of the disease, condition, or its associated symptoms.

[0057] Genes: All genes, gene names, and gene products disclosed herein are intended to correspond to homologs of any species to which the compositions and methods disclosed herein are applicable. It should be understood that when genes or gene products from a particular species are disclosed, this disclosure is intended to be exemplary only and should not be construed as limiting unless the context in which it appears clearly indicates otherwise. Thus, for example, the disclosure of genes or gene products herein is intended to cover homologous and / or orthologous genes and gene products from other species.

[0058] Scope: Throughout this disclosure, various aspects of this disclosure may be presented in the form of scope. It should be understood that the description in scope form is merely for convenience and brevity and should not be construed as a rigid limitation on the scope of this disclosure. Accordingly, a description of a scope should be considered as specifically disclosing all possible sub-scopes within that scope and the individual numerical values ​​thereof. A description of a scope such as 1 to 6 should be considered as specifically disclosing sub-scopes such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., and the individual numbers within that scope, such as 1, 2, 2.7, 3, 4, 5, 5.3, and 6. This applies regardless of the breadth of the scope.

[0059] Any composition or method provided herein may be combined with one or more of any other compositions and methods provided herein. Attached Figure Description

[0060] This patent or application document contains at least one color drawing. Upon request and payment of the necessary fees, the Patent Office will provide a copy of this patent or application publication with color drawings.

[0061] Figure 1 This is a schematic diagram of one implementation of the ARTEMIS method. De novo identification of k-mers revealed approximately 1.1 billion unique k-mers, spanning 6 families and 1280 different repeat elements. The k-mer repetitive sequence landscape was defined as the sum of all k-mer counts for each repetitive sequence type identified across all sequencing reads, normalized for coverage. These landscapes were used in machine learning to generate ARTEMIS scores for disease prediction.

[0062] Figure 2 (including) Figures 2A-2EFigure 2A shows the k-mer repetitive sequence landscape in human cancers, revealing extensive differences from normal tissues. The heatmap (Figure 2A) shows the k-mer repetitive sequence landscape ratios for each PCAWG tumor compared to its matched normal tissue, revealing a large number of tumor-specific changes that can be correlated with genomic instability indicators (n=469 tumor / normal pairs, representing all PCAWG samples for which genomic instability indicators were available). Each PCAWG tumor is listed along the y-axis, and each individual repetitive element type along the x-axis. A ratio greater than 1 (red) indicates an increase in that element in the tumor, while a ratio less than 1 (blue) indicates a decrease. As indicated by the yellow bars of evidence on the x-axis, most of the changes thus identified occurred in elements for which there was no prior evidence of change in cancer (820 out of 1280). Figure 2B shows elements derived from all six families of repetitive elements, sorted by Benjamini-Hochberg corrected p-values ​​according to the Wilcoxon signed-rank test, which compares the overlap of repetitive elements with tumor-specific structural breakpoints with their overlap with randomly selected genomic regions. Solid circles represent newly discovered cancer-related elements identified in this study, while hollow circles represent elements with prior evidence of cancer association. Red circles represent elements with breakpoint exhaustion, and yellow circles represent elements with breakpoint enrichment. Figure 2C: Box plot showing the relationship between PCAWG breast cancer (n=91 tumors / normal pair) and... ERBB2 Tumors with overlapping (1 Mb) regions of repeating element types: normal k-mer count ratio distribution. The ratio of all elements with >0.5% k-mers found in this region is shown for each patient (left), with Benjamini-Hochberg corrected p-values ​​for each comparison plotted using the Wilcoxon signed-rank test; red dots indicate p<0.05 (right). Bold element names indicate those newly identified as cancer-related in this study. Figure 2D: Box plot showing the ratio of tumors with k-mers occurring within LINE-1-mediated deletions in PCAWG lung tumors containing at least one LINE-1-mediated deletion: normal k-mer count ratio (n=5, Supplementary Data File 2). Figure 2E: Kaplan-Meir curves for overall survival and progression-free survival in AJCC stage III or IV PCAWG tumors (n=167), stratified into two groups based on predicted ARTEMIS scores. The group shown in blue has ARTEMIS scores below the median, and the group shown in red has ARTEMIS scores above the median.

[0063] Figures 3A-3B: K-mer repetitive sequence landscape capturing tumor-specific changes in plasma. Figure 3A: Top plot, each bar shows the percentage (dark blue) of k-mers found on chrY in the chm13 reference genome and the percentage (light blue) of k-mers found on all other chromosomes for a given human satellite sequence 2 or 3 element type. Bottom plot, in non-cancer individuals (n=158), the distribution of coverage-normalized k-mer counts in cfDNA for these satellite sequence types in males (n=87) and females (n=71). The p-value of the Wilcoxon signed-rank test is shown at the top of each plot. Figure 3B: k-mer counts in PCAWG tissues (top, n=54 hepatocellular carcinoma, n=48 squamous cell lung carcinoma, and n=38 adenocarcinoma tumor / normal pairs) and plasma cfDNA (bottom; n=75 hepatocellular carcinoma patients and n=133 non-cancer patients; n=29 squamous cell lung carcinoma patients and n=158 non-cancer patients; n=62 adenocarcinoma patients and n=158 non-cancer patients). For each cancer type, the top five features that were significantly different in both tissues and plasma and had at least 1000 expected k-mers per million aligned reads are shown as separate plots. P-values ​​are shown at the top of each plot and were calculated using the Wilcoxon signed-rank test.

[0064] Figure 4 This shows the effect of epigenetic status on the representativeness of repetitive elements in cfDNA. (Top image) Figure 1 The peak values ​​per Mb for each chromatin state of each histone type are summarized. Peak density was scaled in each CHIP-Seq experiment to eliminate the effect of differences in peak number across experiments. Subfigure 2 shows the percentage of histone peaks for each type in each of the 1280 repeat elements divided into six families, illustrating the variation in epigenetic state among repeat sequence types. Subfigure 3 shows the distribution of aligned fragment lengths for fragments overlapping with each histone marker and for all fragments in plasma from non-cancer patients derived from the LUCAS cohort, showing an increase in shorter fragments within regions associated with activation-related histone markers. Lines are medians, and shading indicates + / - 1 standard deviation for plotting distributional differences. Figure 4 In plasma from non-cancer patients derived from the LUCAS cohort, a plot of genome-wide coverage relative to each histone marker region shows decreased coverage in activation-related histone marker regions. The x-axis represents the log-mean coverage, and the y-axis represents the logarithmic difference in counts. Subplot 5, in plasma from non-cancer patients derived from the LUCAS cohort, shows the reduced representativeness of elements with higher numbers of activated histone markers by the ratio of observed mean k-mer counts to expected k-mer counts for features located at the top and bottom decimal places of histone marker density.

[0065] Figures 5A-5C The ARTEMIS and ARTEMIS-DELFI methods for early lung cancer detection using cfDNA are shown. The distribution of ARTEMIS and combined ARTEMIS-DELFI scores in the cross-validated LUCAS cohort for both cancer and non-cancer patients shows that scores are lower in non-cancer patients than in cancer patients (Figure 5A). ROC analysis of ARTEMIS and ARTEMIS-DELFI scores demonstrates high performance in classifying individuals with and without lung cancer in the full LUCAS cohort and in subgroups classified by cancer stage (Figure 5B). The sensitivity and specificity achieved by ARTEMIS and ARTEMIS-DELFI in the external validation cohort, at a scoring threshold selected in the cross-validation cohort to achieve 50%–80% specificity, elucidates the general applicability of these methods for early detection in high-risk populations (Figure 5C).

[0066] Figure 6 This is a detailed schematic diagram of the ARTEMIS method.

[0067] Figure 7 The relationship between genome-wide repetitive element size and the number of unique k-mers identified is shown. The x-axis shows the number of unique k-mers identified in the reference genome, and the y-axis shows the size of genome-wide repetitive elements. Each point represents one of 1280 repetitive element types, grouped into 6 families.

[0068] Figures 8A-8D This is a characterization of genome-wide ARTEMIS k-mers and their enrichment in cancer-related genes. The number of chromosomes containing at least one k-mer for each of the 1280 repetitive element features is shown (Fig. 8A). Each box plot represents this distribution for an element subfamily. The color bar indicates the number of unique k-mers defining this element (log10). The distribution of the 1.1 billion k-mers defining these elements across the genome shows that the vast majority of k-mers appear only once in the genome (Fig. 8B). The number of repetitive element features (1280) found on each chromosome shows their genome-wide distribution (Fig. 8C). Gene set enrichment frontier analysis (GSEA) for the COSMIC Cancer Gene Census, such as sorting by total k-mers per gene (top left) and sorting by k-mer density adjusted for the number of unique types found per gene (top right), shows the enrichment of repetitive sequence k-mers in cancer-related genes (Fig. 8D). As a control, the same number of randomly selected genes are shown in the bottom subplot. In each subplot, the gray line represents 100 permutations of gene rankings, simulating a zero distribution, while the blue line represents the set of genes being analyzed.

[0069] Figure 9 This is a gene set enrichment analysis based on the total number and density of k-mers. It shows all KEGG gene sets enriched with an FDR < 0.05, including many gene sets related to cancer cell processes or pathways.

[0070] Figure 10 This is a simulation of k-mer counting in short-read next-generation sequencing. 50 million 100bp paired-end reads were simulated from the chm13 reference genome, incorporating the actual sequencing error rate. Reads truly originating from repetitive sequence regions were selected, and k-mers of repetitive elements were counted. For each true-source element (i.e., rows with a sum of 1), the percentage of k-mers counted from a given element type was plotted, showing that 98% of the counted k-mers appeared in reads originating from the correct repetitive elements. Insets show only SINE element types.

[0071] Figure 11 This shows the localization of k-mers from human satellite sequences 2 and 3 in CHM13. Altemose et al. described... 1 Subgroups of human satellite sequences 2 and 3 were not annotated in the RepeatMasker orbit. In the simulation, k-mers derived from these element types largely overlapped with the broader satellite sequence regions defined in the RepeatMasker orbit (left panel) and were counted across the entire genome (right panel).

[0072] Figure 12 This demonstrates the correlation between k-mer count and copy number. The ratio of tumor to normal k-mer counts per chromosome arm for PCAWG samples shows a correlation between k-mer count and increased tumor copy number. The gray dashed line represents the expected k-mer count ratio for copy-neutral regions (diploid). The observed fitted line slope is less than 1, indicating that fewer k-mers were found per chromosome arm than expected based on copy number, consistent with the view that such repetitive sequences may experience deletion when they contribute to increases in nearby genomic content.

[0073] Figures 13A-13C Showing PCAWG tumor ERBB2 and SOX2 / PIK3CADistribution of gene copy numbers and changes associated with repetitive sequences. Copy number of ERBB2 in breast cancer (Fig. 13A), copy number of SOX2 in lung cancer (Fig. 13B, left panel), copy number of PIK3CA in lung cancer (Fig. 13B, middle panel), and the number of samples with or without SOX2 / PIK3CA amplification, showing copy number >5 along chr3, with regions of SOX2 and PIK3CA marked by black boxes (Fig. 13B, right panel). Tumors with repetitive element types overlapping the SOX2 / PIK3CA region (31 Mb) in lung cancer: the normal k-mer count ratio distribution shows an increased count for tumors carrying focal amplification (Fig. 13C). All elements with >10.0% k-mers found in this region are plotted. For each comparison, p-values ​​corrected for Benjamini-Hochberg from the Wilcoxon signed-rank test are plotted. Several of these elements represent novel changes not previously associated with cancer (bold names).

[0074] Figure 14 This indicates that the k-mer repetitive sequence landscape in human cancers reveals numerous differences from normal tissues, exceeding the range that could be expected from chromosome arm gains and deletions alone. The k-mer repetitive sequence landscape ratios for each PCAWG tumor compared to its matched normal tissue are shown, with all insignificant changes or changes in scale expected solely by chromosome arm gains and deletions colored white. Even after copy number normalization, significant variations in the k-mer repetitive sequence landscape remain apparent, including changes in novel repetitive elements. The heatmap shows 329 of 333 tumors, including those for which copy number data were available from TCGA.

[0075] Figure 15 This demonstrates the impact of LINE-1 mediated deletions on k-mer counts in the PCAWG lung cancer cohort. The left panel shows that tumors with deletions have lower k-mer counts regardless of flanking copy number. The right panel shows the families of repeating elements represented by k-mers in each deleted region. As expected, most k-mers originate from the LINE repeating element type.

[0076] Figures 16A-16C This demonstrates the impact of SINE element changes on PCAWG tumor survival. To adjust for tumor type, tumors were stratified according to the type-specific median of SINE changes, ensuring that each group contained half of each tumor type (Figures 16A, 16B). This differentiated tumors based on overall survival and showed a trend toward significant progression-free survival. The distribution of SINE element changes in PCAWG tumors exhibited a tissue-type-specific pattern of change (Figure 16C).

[0077] Figures 17A and 17B illustrate the impact of genomic stability indices on PCAWG tumor survival. Overall survival (Figure 17A) and progression-free survival (Figure 17B) did not show a significant association with genomic stability indices.

[0078] Figure 18 This demonstrates the phylogenetic variability of repeating elements in normal PCAWG tissue samples. The variability of the observed count versus expected count ratio for each repeating element type is shown in all normal PCAWG samples. The top figure shows the coefficient of variation for each of the 1275 repeating types (whose k-mers appear in <75% of chrY) (we excluded five repeating types with high proportions of k-mers in chrY because sex-related variation could affect the coefficient of variation for these repeating types). The bottom figure shows the distribution of k-mer counts, with the inset highlighting the 10 features with the largest coefficients of variation, most of which are satellite elements.

[0079] Figure 19 This demonstrates that ARTEMIS can effectively distinguish PCAWG tumors from normal samples across all cancer types. The ARTEMIS model was trained using penalized logistic regression on a landscape of k-merged repeat sequences from 333 tumor samples and 333 normal samples, and evaluated using the average score from 5-fold cross-validation (repeated 10 times). ARTEMIS shows high overall performance in distinguishing PCAWG tumors from normal tissue across all tumor types and within individual tumor types.

[0080] Figure 20 This indicates that the k-merchant repeat sequence landscape in the PCAWG samples exhibits consistency under secondary sampling coverage. The correlation coefficients of 42 PCAWG lung cancer samples between the original 40-80X samples and the secondary-sampled 30X and 1-2X versions (top), and the Bland-Altman plot of five random samples (bottom) show consistency. In the Bland-Altman plot, the x-axis is the logarithm of the k-merchant count mean, and the y-axis is the logarithm of the count difference.

[0081] Figure 21 The k-mer repetitive sequence landscape in PCAWG secondary sampling samples of normal tissue showed minimal difference. The heatmap (left) and ROC analysis (right) showing the k-mer repetitive sequence landscape ratio between two half-coverage secondary sampling samples of 100 random normal samples produced minimal difference, indicating that the observed variation between tumor and normal tissue is due to biological changes rather than technological variations.

[0082] Figure 22The comparison of the k-mer repetitive sequence landscape in plasma sequenced at 1-2x coverage shows consistency across platforms and sequencing batches. The top figure, showing PCA analysis of healthy samples from the LUCAS cohort by genomic library batch and sex, reveals that while sex differences in k-mer counts are significant, technical differences between library batches are not. The correlation (bottom left) and Bland-Altman plots (bottom right) for the k-mer repetitive sequence landscape of the LUCAS cohort sequenced on HiSeq and NovaSeq platforms, along with the Bland-Altman plots for five randomized cancer and healthy patients, show high consistency among sequencing repetitions. On the Bland-Altman plots, the x-axis represents the logarithmic value of the count mean, and the y-axis represents the difference in the logarithmic value of the count. The dashed line distinguishes repetitive element types based on whether the expected count is less than or more than 1000 k-mers per million aligned reads—elements to the right of this line show higher consistency across platforms in low-coverage sequencing.

[0083] Figure 23 This diagram shows the queues and datasets used for plasma analysis.

[0084] Figure 24 This shows the localization of 1 Mb bins with high-density epigenetic markers used for plasma ARTEMIS analysis.

[0085] Figure 25 (including Figures 25A-25C): Importance and stability of features in locked ARTEMIS and ARTEMIS-DELFI models for lung cancer. Figure 25A: Standardized coefficients of individual features preserved by the LASSO regression model for lung cancer, scaled according to their integration weights in the ARTEMIS-DELFI model. Figure 25B: Stability analysis of scores generated by these models for individual patients using different cross-validation fold sizes. Figure 25C: Stability analysis of inter-group scores generated using different cross-validation fold sizes.

[0086] Figure 26 This section showcases ARTEMIS for early detection of liver cancer using cfDNA. The top figure shows the distribution of ARTEMIS and combined ARTEMIS-DELFI scores in a cross-validated liver cancer cohort for both cancer and non-cancer patients, indicating that non-cancer patients scored lower than cancer patients. The bottom figure shows ROC analysis of the ARTEMIS score and the combined ARTEMIS-DELFI model, demonstrating high performance in classifying individuals with and without liver cancer across the entire cohort and within subgroups based on cancer stage.

[0087] Figure 27This study demonstrates the application of the locked ARTEMIS cfDNA model in a cohort of patients receiving lung cancer treatment. The top-to-bottom plots show similar dynamic changes in the maximum mutational allele score, ARTEMIS-DELFI score, DELFI score, and ARTEMIS score at each obtained time point. Patients are plotted from left to right in ascending order of progression-free survival, demonstrating that the MAF trajectory associated with tumor burden can be represented by scores derived from the k-mer repetitive sequence landscape and fragmentation profile.

[0088] Figure 28 Germplasmic variability of repeating elements in normal PCAWG tissue samples. The variability of the observed versus expected count ratios for each repeating element type in all normal PCAWG samples. The top figure shows the coefficient of variation for each of the 1275 repeating types in which <75% k-mers appear on chrY (we excluded five repeating types with high proportions of k-mers on chrY because sex-related variability could affect the coefficient of variation for these repeating types). The bottom figure shows the distribution of k-mer counts, with the inset highlighting the 10 features with the highest coefficients of variation, many of which are satellite elements.

[0089] Figure 28 ARTEMIS can effectively distinguish PCAWG tumors from normal samples across all cancer types. The ROC curves of the ARTEMIS model were trained using penalized logistic regression on a landscape of k-merged repeat sequences from 525 tumor and 525 normal samples, and evaluated using the average score from 10 repeated 5-fold cross-validations. The overall ROC curve and ROC curves differentiated by ethnicity or tumor type are shown.

[0090] Figure 29 (include Figure 29 (A and 29B): The impact of ARTEMIS score on PCAWG tumor survival. Figure 29 A: ARTEMIS scores of all tumor tissues compared to normal samples. Showing all stage III or IV tumors and their matched normal tissues (n=167). Figure 29 B: Kaplan-Meier plots of overall survival and progression-free survival for these cancer patients stratified by high or low ARTEMIS scores. To adjust for tumor type, tumors were stratified by type-specific median ARTEMIS score, ensuring that each group contained half of each tumor type. P-values ​​were calculated using a log-rank test. Figure 29 In the meantime, values ​​with high ARTEMIS scores are lower on the y-axis compared to values ​​with low ARTEMIS scores.

[0091] Figure 30This section simulates the combined effects of tumor-specific structural and epigenetic changes on cfDNA representativeness. The bar charts show the proportions of structural and epigenetic changes in repetitive elements in cfDNA from cancer patients that were simulated, with tumor-derived changes exhibiting consistent or inconsistent directionality in plasma. Red bars indicate simulations where plasma coverage is affected by both tumor-derived structural and epigenetic changes, while gray bars indicate simulations considering only structural changes. In paired bar charts, the left side depicts the impact of epigenetic changes on plasma coverage, and the right side depicts the absence of such effects.

[0092] Figure 31 (include Figure 31 A-31B): ARTEMIS uses cfDNA for liver cancer detection. Figure 31 A: Distribution of ARTEMIS and combined ARTEMIS-DELFI scores in cancer and non-cancer patients in a cross-validated hepatocellular carcinoma cohort. Figure 31 B: ROC analysis of ARTEMIS score and combined ARTEMIS-DELFI model to classify individuals with and without liver cancer in the whole cohort and in subgroups by cancer stage.

[0093] Figure 32 External validation of locked ARTEMIS and ARTEMIS-DELFI models for lung cancer detection. Period-wise ROC analysis of external validation of locked ARTEMIS and ARTEMIS-DELFI models for lung cancer detection.

[0094] Figure 33 Validation of locked ARTEMIS and ARTEMIS-DELFI models for detecting lung cancer recurrence. The locked ARTEMIS and ARTEMIS-DELFI models for lung cancer detection were compared with scores from samples from individuals with a history of cancer recurrence to samples from individuals who had not experienced recurrence. P-values ​​were calculated using the Wilcoxon signed-rank test.

[0095] Figure 34 (including Figures 34A-34B) shows the validation of the locked ARTEMIS and ARTEMIS-DELFI models for detecting lung cancer recurrence. The locked ARTEMIS and ARTEMIS-DELFI models for lung cancer detection were compared with scores from samples from individuals with a history of cancer recurrence to samples from individuals who did not experience recurrence. P-values ​​were calculated using the Wilcoxon signed-rank test. Detailed Implementation

[0096] In one aspect, an alignment-independent de novo k-mer discovery method is provided for identifying repetitive elements from whole-genome sequencing. Using this method, it was demonstrated in recently characterized long-read reference genomes that repetitive elements are enriched in genes and pathways that are typically altered in human cancers. In a subsequent example section analyzing 2428 samples from 1758 individuals, tumor-specific changes in representative repetitive elements were revealed to be detectable in tissue and circulating cell-free DNA (cfDNA) from cancer patients. A machine learning model trained using whole-genome repetitive sequence landscapes and fragmentation profiles in cfDNA was used to detect early-stage lung cancer patients and validated in an independent diagnostic cohort. The k-mer repetitive sequence landscape approach allows for the reconstruction of repetitive sequence landscapes using low-coverage whole-genome sequencing and facilitates large-scale analysis of these regions for the characterization and detection of human cancers. ARTEMIS

[0097] ARTEMIS (Disease Repetitive Element Analysis) is a genome-wide, alignment-independent, whole-genome approach developed for analyzing repetitive sequence landscapes in short-read sequencing. ARTEMIS assesses over 1200 individual repetitive sequence types, occurring genome-wide and spanning 57 subfamilies across 6 families (satellite sequences, RNA elements, transposable elements, LINE, SINE, LTR). In this study, we used ARTEMIS to show that repetitive sequence landscapes are enriched in genes that are universally altered in human cancers, and that tumor-specific changes in repetitive sequences reflect a combination of structural and epigenetic changes in the cancer genome. Genome-wide repetitive sequence landscape analysis using ARTEMIS can be implemented using low-coverage whole-genome sequencing, allowing for the analysis of repetitive sequence landscapes in cfDNA for the detection of human cancers.

[0098] Therefore, in some aspects, a method for identifying repeat element types includes: extracting repeat sequences and coordinates from known repeat element types; identifying short nucleic acid sequences (k-mers) in genomic or cell-free DNA; selecting k-mers appearing in a single repeat sequence type and identifying unique k-mers of the repeat element type; wherein the unique k-mers identify one or more repeat element types. In some embodiments, repeat element types are excluded from families containing low-complexity, unknown, simple repeats, or combinations thereof. In some embodiments, elements from each family, including tRNA, srpRNA, snRNA, scRNA, rRNA, RNA elements, DNA, reverse transcripts, or combinations thereof, are aggregated.

[0099] In some embodiments, one or more k-mers are generated from one or more nucleic acid sequences. The nucleic acid sequences can be any sequence, such as genomic sequences or cfDNA. In one embodiment, k-mer generation is performed by an automated algorithm that receives a nucleic acid sequence and generates multiple k-mers from that sequence. K-mers can be generated from a single organism or species, or from multiple organisms or species. For example, k-mers can be generated from the genomic sequences of common, possible, or known contaminants to detect these contaminants during quality control analysis. Furthermore, k-mers can be generated from the genomic sequences of different organisms that may be present in a composite sample to detect and analyze sequence data from multiple different organisms during quality control analysis.

[0100] k-mers extracted from one or more nucleic acid sequences can have the same length or a variety of different lengths. As just one example, the generated k-mers can be approximately 20, 24, 28, or 32 bases, but longer and shorter k-mers are also suitable. Larger k-mers are less likely to find multiple matches in the reference genome, which can lead to multiple types of k-mer annotations. For example, small k-mers map to many parts of the genome and do not provide useful and unique annotation information. However, large k-mers also increase the amount of memory required to store the generated k-mers. Shorter k-mers reduce the amount of memory required to store the generated k-mers. Therefore, determining the optimal k-mer depends in particular on various factors such as computing power and memory size.

[0101] In some embodiments, k-mers are selected that appear in a single repeating element and not in non-repetitive regions. In some embodiments, the k-mer comprises 10 to 40, 50, 60, or 70 nucleotides. In some embodiments, the k-mer comprises 10 to 40 nucleotides. In some embodiments, the k-mer comprises 15 to 35 nucleotides. In some embodiments, the k-mer comprises 20 to 30 nucleotides. In some embodiments, the k-mer comprises 22 to 28 nucleotides. In some embodiments, the k-mer comprises 23 to 26 nucleotides. In some embodiments, the k-mer comprises about 24 nucleotides. system

[0102] In some instances, this disclosure provides systems, methods, or kits that may include data analysis implemented in measurement devices (e.g., laboratory instruments such as sequencers), software code executed on computing hardware. The software may be stored in memory and executed on one or more hardware processors. The software may be organized into routines or packages that can communicate with each other. Modules may include one or more devices / computers, and potentially one or more software routines / packages executed on said one or more devices / computers. For example, an analytical application or system may include at least a data receiving module, a data preprocessing module, a data analysis module (capable of operating on one or more types of genomic data), a data interpretation module, or a data visualization module.

[0103] A data receiving module can connect laboratory hardware or instruments to a computer system that processes laboratory data. A data preprocessing module can manipulate the data to prepare it for analysis. Examples of data manipulation that can be applied in the preprocessing module include affine transformation, denoising, data cleaning, reformatting, or secondary sampling. A data analysis module can be specifically designed to analyze genomic data from one or more genomic materials. It can, for example, acquire assembled genomic sequences and perform probabilistic and statistical analyses to identify anomalous patterns associated with disease, pathology, state, risk, condition, or phenotype. A data interpretation module can use analytical methods, for example, derived from statistics, mathematics, or biology, to support the understanding of the relationship between identified anomalous patterns and health status, functional state, prognosis, or risk. The data analysis and / or data interpretation modules can include one or more machine learning models, which can be implemented in hardware, for example, software that embodies the machine learning model. A data visualization module can use mathematical modeling, computer graphics, or rendering methods to create visual representations of the data that facilitate the understanding or interpretation of the results. This disclosure provides computer systems programmed to perform the methods of this disclosure.

[0104] In some implementations, the methods disclosed herein may include computational analysis of nucleic acid sequencing data from samples from one or more individuals. The analysis may identify variants inferred from the sequence data to identify sequence variants based on probabilistic modeling, statistical modeling, mechanistic modeling, network modeling, or statistical inference. Non-limiting examples of analytical methods include principal component analysis, autoencoders, singular value decomposition, Fourier basis functions, wavelets, discriminant analysis, regression, support vector machines, tree-based methods, networks, matrix factorization, and clustering. Non-limiting examples of variants include germline variants or somatic mutations. In some instances, a variant may refer to a known variant. A known variant may be scientifically confirmed or reported in the literature. In some instances, a variant may refer to a presumptive variant associated with a biological change. The biological change may be known or unknown. In some instances, a presumptive variant may have been reported in the literature but has not yet been biologically confirmed. Alternatively, a presumptive variant may never have been reported in the literature but can be inferred based on the computational analysis disclosed herein. In some instances, a germline variant may refer to a nucleic acid that induces natural or normal variation.

[0105] In some implementations, the computer system includes a central processing unit (CPU, also referred to herein as a "processor" and "computer processor"), which may be a single-core or multi-core processor, or multiple processors for parallel processing; memory (e.g., cache, random access memory, read-only memory, flash memory, or other memory); electronic storage units (e.g., hard disks); communication interfaces (e.g., network adapters) for communicating with one or more other systems; and peripheral devices, such as adapters for caches, other memories, data storage, and / or electronic displays. The memory, storage units, interfaces, and peripheral devices may communicate with the CPU via a communication bus (solid line) such as a motherboard. Storage units may be data storage units (or databases) for storing data. One or more analyte characteristic input values ​​may be input from the one or more measuring devices. Exemplary analytes and measuring devices are as described herein.

[0106] Computer systems can be operatively coupled to computer networks (“networks”) with the aid of communication interfaces. The network may be the Internet, an intranet and / or an extranet, or an intranet and / or extranet communicating with the Internet. In some cases, the network is a telecommunications and / or data network. The network may include one or more computer servers capable of distributed computing, such as cloud computing via a network (“cloud”), to perform various aspects of the analysis, calculation, and generation of this disclosure, such as, for example, activating valves or pumps to transfer reagents or samples from one chamber to another, or applying heat to samples (e.g., during amplification reactions), processing and / or determining other aspects of samples, performing sequencing analysis, measuring sets of values ​​representing molecular categories, identifying feature sets and eigenvectors from determination data, processing eigenvectors using machine learning models to obtain output classifications, and training machine learning models (e.g., iteratively searching for optimal values ​​for machine learning model parameters). Such cloud computing may be provided by cloud computing platforms such as Amazon Web Services (AWS), Microsoft Azure, Google Cloud Platform, and IBM Cloud. In some cases, with the help of computer systems, networks can achieve peer-to-peer networking, which allows devices coupled to the computer system to act as clients or servers.

[0107] A CPU can execute a series of machine-readable instructions, which can be embodied in a program or software. These instructions can be stored in a memory location, such as memory. The instructions can be directed to the CPU, which can then be programmed or otherwise configured to implement the methods disclosed herein. The CPU can be part of a circuit, such as an integrated circuit. One or more other components of the system can be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).

[0108] Storage units can store files, such as drivers, libraries, and saved programs. Storage units can also store user data, such as user preferences and user programs. In some cases, a computer system may include one or more additional data storage units located outside the computer system, such as those located on a remote server communicating with the computer system via an intranet or the Internet.

[0109] A computer system can communicate with one or more remote computer systems via a network. For example, a computer system can communicate with a user's remote computer system. Examples of remote computer systems include personal computers (e.g., portable PCs), slates or tablet PCs (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, smartphones (e.g., Apple® iPhone, Android-enabled devices, Blackberry®), or personal digital assistants. Users can access computer systems via a network.

[0110] The methods described herein can be implemented by machine-executable code (e.g., a computer processor) stored in an electronic storage location within a computer system, such as, for example, in a memory or electronic storage unit. The machine-executable or machine-readable code can be provided in software form. During use, the code can be executed by the CPU. In some cases, the code can be retrieved from a storage unit and stored in memory so that the CPU can access it at any time. In some cases, an electronic storage unit can be excluded, and the machine-executable instructions are stored in memory.

[0111] Code can be pre-compiled and configured for use by machines with processors suitable for executing the code, or it can be compiled during runtime. The code can be provided in a programming language, which can be selected to enable the code to be executed either pre-compiled or compiled.

[0112] Various aspects of the systems and methods provided herein, such as computer systems, can be embodied in programming. These aspects of the technology can be considered as “products” or “artifacts” typically existing in the form of machine (or processor) executable code and / or associated data carried or embodied in a type of machine-readable medium. Machine-executable code can be stored on electronic storage units, such as memory (e.g., read-only memory, random access memory, flash memory) or hard disks. Medium of the “storage” type can include any or all tangible memory, such as various semiconductor memories, magnetic tape drives, and disk drives, of computers, processors, or their associated modules, which can provide non-transitory storage for software programming at any time. All or part of the software can sometimes be communicated via the Internet or various other telecommunications networks. Such communication, for example, can load software from one computer or processor to another, for example, from a management server or host computer to a computer platform for an application server. Therefore, another type of medium that can carry software elements includes light waves, radio waves, and electromagnetic waves, such as physical interfaces between local devices, via wired and fiber optic fixed-line networks, and over various air links. Physical elements carrying such waves, such as wired or wireless links, optical links, etc., can also be considered as media carrying software. As used herein, unless limited to non-transitory, tangible "storage" media, the term "readable medium" for a computer or machine refers to any medium that participates in providing instructions to a processor for execution.

[0113] Therefore, machine-readable media, such as computer-executable code, can take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media include, for example, optical discs or disks, any storage device such as any computer, such as those used to implement the database shown in the figure. Volatile storage media include dynamic memory, such as the main memory of a computer platform. Tangible transmission media include coaxial cables; copper wires and optical fibers, including wires that form the bus within a computer system.

[0114] Carrier transmission media can take the form of electrical or electromagnetic signals, or sound or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Therefore, common forms of computer-readable media include, for example: floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punched cardboard tape, any other physical storage media with a perforated pattern, RAM, ROM, PROM and EPROM, FLASH-EPROM, any other memory chips or cassette tapes, carrier waves for transmitting data or instructions, cables or links for transmitting such carrier waves, or any other media from which a computer can read programming code and / or data. Many of these forms of computer-readable media can participate in delivering one or more sequences of one or more instructions to a processor for execution.

[0115] The computer system may include or communicate with an electronic display, which includes a user interface (UI) for providing information such as the current stage of sample processing or assay (e.g., a specific step being performed, such as a lysis step or sequencing step). Input is received by the computer system from one or more measurements. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces. Algorithms may, for example, process and / or assay samples, perform sequencing analyses, identify k-mers containing each type of repetitive sequence, measure sets of values ​​representing molecular categories, identify feature sets and eigenvectors from assay data, process eigenvectors using a machine learning model to obtain an output classification, and train machine learning models (e.g., iteratively searching for optimal values ​​for machine learning model parameters).

[0116] In some implementations, systems capable of executing one or more algorithms to determine changes in cfDNA mutation profiles, mutation frequencies, and / or fragmentation profiles, such as laptops, desktops, iPads, and mobile devices, classify subjects as cancer patients based on their cfDNA mutation profiles, mutation frequencies, and / or fragmentation. 54 These systems further execute machine learning algorithms that can be used to generate models, such as, for example, high-risk groups and low-risk general groups (using Mathios et al.). 27 Features and coverage from transcription factor binding sites 45(Punished logistic regression). These models can be trained on a cohort of subjects using 5-fold cross-validation with 10 replicates, and a score for each sample is calculated from the average of all replicates and evaluated using AUC-ROC. For example, the first model uses high-risk non-cancer and HCC patients, while the second model uses non-cancer individuals without liver pathology. The locked high-risk model trained on the cohort is applied to a second, different cohort to generate cancer predictions on an external validation set. "Class labels" can be applied to each sample, indicating the sample classification for any number of input features. For example, the class labels for the cohort set can identify k-merged repeat sequence landscapes, features indicating cfDNA mutation profiles, mutation frequencies, and / or fragmentation profiles based on genomic localization, etc. The resulting training set is fed to a machine learning unit, such as a neural network or support vector machine. Using the training set, the machine learning unit can generate models to classify samples according to the k-merged repeat sequence landscape to generate ARTEMIS scores, cfDNA mutation profiles, mutation frequencies, and / or fragmentation profiles for disease prediction.

[0117] In some implementations, a method for creating a trained classifier is provided, comprising the steps of: extracting repetitive sequences and coordinates from known repetitive element types; identifying short nucleic acid sequences (k-mers) in the genome or cell-free DNA; selecting k-mers appearing in a single repetitive type and identifying unique k-mers of the repetitive element type; wherein the unique k-mers identify one or more repetitive element types. For the analysis of a single sample, the k-mer repetitive sequence landscape is defined as the count of all k-mers in the sequencing sample that match each of 1280 repetitive sequence types divided by the number of aligned sequence reads. Since changes in repetitive sequences can occur during the initiation of cancer and other diseases, this comprehensive compilation of repetitive sequence features can be used to train machine learning models to distinguish genomes in normal and disease states.

[0118] As an example, a trained classifier can use a learning algorithm selected from the following groups: random forest, neural network, support vector machine, and linear classifier. Each of the multiple different categories can be selected from the following groups: health, breast cancer, colon cancer, lung cancer, pancreatic cancer, prostate cancer, ovarian cancer, melanoma, and liver cancer.

[0119] A trained classifier can be applied to methods for classifying samples from subjects. Such classification methods may include: (a) providing a multi-parameter model of k-merged repeat sequences representing test samples from subjects; and (b) classifying the test samples using the trained classifier. After the test samples are classified into one or more categories, therapeutic interventions can be administered to subjects based on the sample classification.

[0120] In some implementations, a training set is provided to the machine learning unit (e.g., a neural network or support vector machine). Using the training set, the machine learning unit can generate a model to classify samples based on responses to one or more treatments. This is also known as "calling." The developed model can incorporate information from any part of the test vector.

[0121] Typically, machine learning can be used to reduce a dataset generated from all combinations of (primary samples / analytes / tests) to the best predictive feature set, such as a feature set that meets specified criteria. Statistical learning and / or regression analysis can be applied in various instances. Models generating various model hypotheses, ranging from simple to complex and from small to large, can be applied to the data in a cross-validation paradigm. From simple to complex, this includes considering feature representations from linear to nonlinear and from non-hierarchical to hierarchical. From small to large, this includes considering the size of the basis vector space to which the data is projected and the number of interactions between features involved in the modeling process.

[0122] Machine learning techniques can be used to evaluate commercially viable detection modalities that optimize cost / performance / commercial coverage, as defined in the initial problem. Threshold checks can be performed: if a method applied to a reserved dataset not used in cross-validation exceeds the initial constraints, the assay is locked and production is initiated. For example, thresholds for assay performance could include the expected minimum accuracy, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), or a combination thereof. For example, the expected minimum accuracy, PPV, NPV, clinical sensitivity, clinical specificity, or a combination thereof may be at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. As another example, the expected minimum AUC can be at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.81, at least about 0.82, at least about 0.83, at least about 0.84, at least about 0.85, at least about 0.86, at least about 0.87, at least about 0.88, at least about 0.89, at least about 0.90, at least about 0.91, at least about 0.92, at least about 0.93, at least about 0.94, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99. A subset of assays can be selected from a set of assays to be performed on a given sample. This selection is based on the total cost of performing that subset of assays and is subject to thresholds for assay performance, such as the expected minimum accuracy, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), and combinations thereof. If the thresholds are not met, the assay engineering process can loop back to constraint settings for possible easing, or loop back to a wet laboratory to change the parameters in which data are acquired. In the case of a clinical problem, biological constraints, budget, laboratory equipment, etc., can limit the scope of the problem.

[0123] In some implementations, the computer processing of machine learning techniques may include statistical, mathematical, biological methods, or any combination thereof. In various instances, any computer processing method may include dimensionality reduction methods, logistic regression, principal component analysis, autoencoders, singular value decomposition, Fourier basis functions, wavelets, discriminant analysis, support vector machines, tree-based methods, random forests, gradient boosting trees, logistic regression, matrix factorization, network clustering, statistical tests, and neural networks.

[0124] In some implementations, the computer processing of machine learning techniques may include logistic regression, multiple linear regression (MLR), dimensionality reduction, partial least squares (PLS) regression, principal component regression, autoencoders, variational autoencoders, singular value decomposition, Fourier basis methods, wavelets, discriminant analysis, support vector machines, decision trees, classification and regression trees (CART), tree-based methods, random forests, gradient boosting trees, logistic regression, matrix factorization, multidimensional scaling analysis (MDS), dimensionality reduction methods, t-distributed random neighborhood embeddings (t-SNE), multilayer perceptrons (MLP), network clustering, neural fuzzy logic, neural networks (shallow and deep), artificial neural networks, Pearson product-moment correlation coefficient, Spearman rank correlation coefficient, Kendall rank correlation coefficient, or any combination thereof. In some instances, the computer processing method is a supervised machine learning method, including, for example, regression, support vector machines, tree-based methods, and neural networks. In some instances, the computer processing method is an unsupervised machine learning method, including, for example, clustering, networks, principal component analysis, and matrix factorization.

[0125] For supervised learning, training samples (e.g., thousands) can include measurement data (e.g., various analytes) and known labels, which can be determined via other time-consuming processes, such as imaging of subjects and analysis performed by trained practitioners. Example labels can include subject classifications, such as a discrete classification of whether a subject has cancer, or a continuous classification providing discrete probabilities (e.g., risk or rating). The learning module can optimize the model's parameters to achieve a quality metric (e.g., prediction accuracy against known labels) using one or more specified criteria. A quality metric can be determined for any arbitrary function, including the entire set of risk, loss, utility, and decision functions. Gradients can be used in conjunction with learning steps (e.g., a measure of how much the model's parameters should be updated for a given time step in the optimization process).

[0126] As described above, examples can be used for a variety of purposes. For instance, plasma (or other samples) can be collected from subjects with symptoms of a condition (e.g., those known to have the condition) and healthy subjects. Genetic data (e.g., k-mer repeat sequences, cfDNA) can be acquired and analyzed to obtain a variety of different features, which may include features based on genome-wide analysis. These features can form a feature space, which can be searched, stretched, rotated, translated, and subjected to linear or nonlinear transformations to generate an accurate machine learning model that can distinguish between healthy subjects and those with the condition (e.g., identifying a subject's disease or non-disease state). The outputs derived from these data and models (which may include the probability of the condition, the stage (level) of the condition, or other values) can be used to generate another model that can be used to recommend further procedures, such as recommending a biopsy or continuous monitoring of the subject's condition.

[0127] To invoke a treatment response, the invocation algorithm can obtain genetic information and treatment responses from multiple individuals suffering from the disease or condition. The data can first be standardized (using the same procedures used for clustering algorithms). The invocation operation (classification) can be performed using, for example, a Bayesian model. The "invocation score" for each invocation can be the product of the training score and the data-model fit score. After scoring all treatment responses, the application can calculate a comprehensive score.

[0128] In some implementations, the training dataset includes clinical data selected from cancer stage, surgical procedure type, age, tumor grade, tumor invasion depth, postoperative complication occurrence, and presence of venous invasion. In some implementations, the training dataset is preprocessed, including converting the provided data into class-conditional probabilities.

[0129] Another implementation uses machine learning techniques to train a statistical classifier, specifically a support vector machine, for each cancer stage category based on the frequency of word occurrences in a corpus of histological reports for each patient. New reports can then be classified according to the most probable stage, thereby facilitating the collection and analysis of population staging data.

[0130] In some implementations, the machine learning algorithm is selected from the group consisting of: supervised or unsupervised learning algorithms selected from support vector machines, random forests, nearest neighbor analysis, linear regression, binary decision trees, discriminant analysis, logistic classifiers, and cluster analysis.

[0131] Typically, the system may include a report generator for reporting cancer detection results and treatment plans. This report generator system may be a central data processing system configured to communicate directly via a communication link with: a remote data site or laboratory, a healthcare institution / healthcare provider (treatment professional), and / or the patient / subject. The laboratory may be a medical laboratory, diagnostic laboratory, medical facility, healthcare institution, point-of-care testing equipment, or any other remote data site capable of generating clinical information for the subject. The subject's clinical information includes, but is not limited to, laboratory test data, X-ray data, examinations, and diagnoses. Healthcare providers or institutions... 26 This includes healthcare providers such as doctors, nurses, home health aides, technicians, and physician assistants, and the institution is any healthcare facility equipped with a healthcare provider. In some cases, the healthcare provider / institution also serves as a remote data site. In cancer treatment implementation plans, the subject may have cancer, etc.

[0132] Other clinical information for cancer subjects includes the results of laboratory tests, imaging, or medical procedures specific to that cancer, which are readily identifiable by a person skilled in the art. Appropriate sources of clinical information for cancer include, but are not limited to: CT scans, MRI scans, ultrasound scans, bone scans, PET scans, bone marrow tests, barium meal X-rays, endoscopy, lymphangiography, IVU (intravenous urography) or IVP (intravenous pyelography), lumbar puncture, cystoscopy, immunological tests (anti-malignant antibody screening), and cancer marker tests.

[0133] Subject clinical information can be obtained manually or automatically from the laboratory. To simplify the system, information is automatically obtained at predetermined or regular time intervals. Regular time intervals refer to the intervals at which laboratory data are automatically collected using the methods and systems described herein, based on time measurements such as hours, days, weeks, months, years, etc. In one implementation, data collection and processing are performed at least once daily. In another implementation, data transmission and collection are performed monthly, bi-weekly, weekly, or every few days. Alternatively, information retrieval can be performed at predetermined but irregular time intervals. For example, the first retrieval step may occur after one week, and the second retrieval step may occur after one month. Data transmission and collection can be tailored to the nature of the condition being managed and the frequency of tests and medical examinations required by the subject.

[0134] In some embodiments, a genetic reporter is generated from a subject's sample (e.g., cfDNA). Polynucleotides in the sample can be sequenced, for example, by whole-genome sequencing or NGS sequencing, producing multiple sequence reads. In some embodiments, the genetic information includes variables defining the genomic organization of cancer cells or a single disseminated cancer cell genome. In some embodiments, the genetic information includes k-mer repetitive sequences or k-mer landscapes generated from the genomes of one or more subjects. In some embodiments, the genetic information includes sequence or abundance data of one or more loci from an individual's cell-free DNA.

[0135] Genetic variants can also be identified. Genetic variants include sequence variants, copy number variants, and nucleotide modification variants. Sequence variants are variations in the nucleotide sequence of a gene. Copy number variants are deviations in the copy number of a portion of the genome from the wild type. Genetic variants include, for example, single nucleotide variants (SNPs), insertions, deletions, inversions, translocations, translocations, gene fusions, chromosome fusions, gene truncations, copy number variations (e.g., aneuploidy, partial aneuploidy, polyploidy, gene amplification), aberrant changes in nucleic acid chemical modifications, aberrant changes in epigenetic patterns, and aberrant changes in nucleic acid methylation. This process then determines the frequency of genetic variants in a sample containing genetic material. Because this process involves noise, it separates information from noise. The sensitivity of detecting genetic variants can be improved by increasing the read depth of polynucleotides (e.g., by sequencing samples from subjects at two or more time points with higher read depth).

[0136] To improve diagnostic confidence, multiple measurements can be performed. Alternatively, measurements can be performed at multiple time points (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more time points) to determine whether the cancer is progressing, in remission, or stable. Diagnostic confidence can be used to identify disease states. For example, cell-free polynucleotides taken from a subject can include polynucleotides derived from normal cells as well as polynucleotides derived from disease cells, such as cancer cells. Polynucleotides from cancer cells may carry genetic variants, such as somatic mutations and copy number variants. When cell-free polynucleotides from a subject sample are sequenced, cfDNA mutation profiles, mutation frequencies, and / or fragmentation profiles can be generated, as described in the Examples section below.

[0137] In some implementations, diagnostic confidence is based on the ARTEMIS score, as detailed in the Examples section below. Briefly, an ensemble penalized logistic regression model is trained using a k-mer repetitive sequence landscape, and the predictions are defined as the ARTEMIS score. First, a k-mer repetitive sequence landscape is obtained for each sample using 786 features expected to have more than 1000 k-mers per million aligned reads. This filtering is employed because low-abundance features exhibit greater technical variability at low coverage. The remaining features are divided into six families (LINE, SINE, satellite sequences, LTR, RNA elements, DNA elements) and centered and scaled within their respective families in each sample. RNA and DNA elements are merged because only two RNA element features remain after filtering against a threshold of 1000 k-mers / million aligned reads. 561 1 Mb bins are further defined using the epigenetic analysis described above, where >90% of the bases in these bins are covered by a peak from one of the aforementioned histone CHIP-Seq experiments, or >30% of the bases are covered by one of the three chromatin states defined above. The alignment fragment coverage within these bins was calculated. Bins with a mappability <0.9 or GC content <0.3 were excluded from downstream analysis. We then trained six penalized logistic regression (PLR) models using the repetitive sequence landscape for the five repetitive families and the stated epigenetic profiles, and integrated their scores using leave-one-out cross-validation with nested 5-fold cross-validation to train each learner. The combined score was defined as the ARTEMIS score.

[0138] The methods and systems described herein can be used to detect a variety of cancers. Like most cells, cancer cells can be characterized by their turnover rate, where old cells die and are replaced by new ones. Typically, dying cells release DNA or fragments of DNA into the bloodstream upon contact with the vascular system of a given subject. This is also true for cancer cells at different stages of disease. Cancer cells can also be characterized by various genetic aberrations, such as copy number variations and mutations, depending on the stage of the disease. This phenomenon can be used to detect the presence or absence of cancer in an individual using the methods and systems described herein.

[0139] In the early detection of cancer, any of the systems or methods described herein (including mutation detection or copy number variation detection) can be used to detect cancer. These systems and methods can be used to detect any number of genetic aberrations that may lead to or be caused by cancer. These may include, but are not limited to, k-mer landscape, k-mer repetitive sequences, cfDNA mutation profiles, mutation frequencies, cfDNA fragmentation profiles, mutations, insertions, deletions, copy number variations, transversions, translocations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structural alterations, gene fusions, gene truncations, gene amplifications, gene duplications, chromosomal damage, DNA damage, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, abnormal changes in nucleic acid methylation, infection, and cancer.

[0140] Furthermore, the systems and methods described herein can also be used to aid in the characterization of certain cancers. Genetic data generated by the systems and methods of this disclosure, such as k-mer landscapes, can enable practitioners to better characterize specific forms of cancer. Typically, cancers are heterogeneous in both composition and stage. Genetic profiling data can allow for the characterization of specific subtypes of cancer, which can be important for the diagnosis or treatment of that specific subtype. This information can also provide clues to the prognosis of subjects or practitioners regarding specific types of cancer.

[0141] The systems and methods provided herein can be used to monitor known cancers or other diseases in specific subjects. This may enable subjects or practitioners to adjust treatment regimens based on disease progression. In this instance, the systems and methods described herein can be used to construct a subject-specific k-mer repeat sequence landscape of disease process, a combination of the k-mer repeat sequence landscape and cfDNA mutation profile, cfDNA mutation profile, mutation frequency, and / or fragmentation profile. In some cases, cancer progresses, becoming more aggressive and genetically unstable. In other instances, cancer may remain benign, inactive, or dormant. The systems and methods disclosed herein can be used to determine disease progression.

[0142] Furthermore, the systems and methods described herein can be used to determine the effectiveness of specific treatment regimens. In one instance, certain treatment regimens may vary over time and correlate with the cancer's k-mer repeat sequence landscape, the combination of the k-mer repeat sequence landscape with the cfDNA mutation profile, the cfDNA mutation profile, mutation frequency, and / or fragmentation profile. This correlation can be used to select therapies. Additionally, if cancer remission is observed after treatment, the systems and methods described herein can be used to monitor residual disease or disease recurrence.

[0143] Furthermore, the methods disclosed herein can be used to characterize the heterogeneity of abnormal conditions in subjects. These methods include generating a landscape of k-mer repeat sequences of extracellular polynucleotides from the subject, a combination of the k-mer repeat sequence landscape and a cfDNA mutation profile, a cfDNA mutation profile, mutation frequencies, and / or fragmentation profiles, wherein the cfDNA mutation profile contains a large amount of data derived from spectral variation and mutation analysis. In some cases, including but not limited to cancer, diseases can be heterogeneous. Disease cells can be dissimilar. In the case of cancer, some tumors are known to contain different types of tumor cells, some of which are at different stages of cancer. In other instances, heterogeneity can include multiple disease lesions. Again, in the case of cancer, multiple tumor lesions may be present, possibly one or more of which are the result of metastasis from the primary site (also known as distant metastasis).

[0144] The methods disclosed herein can be used to generate datasets that are sums of k-mer repetitive sequence landscapes, combinations of k-mer repetitive sequence landscapes and cfDNA mutation profiles, profiles, fingerprints, or genetic information derived from different cells in heterogeneous diseases. These datasets can contain individual or combined copy number variation and mutation analyses.

[0145] Furthermore, these reports are submitted and accessed electronically via the internet. Data analysis takes place at sites outside the participants' locations. Reports are generated and transmitted to the participants' locations. Participants can access reports reflecting their tumor burden via a network-connected computer.

[0146] Healthcare providers can use the annotation information to select alternative drug treatment options and / or to provide information about drug treatment options to insurance companies. This method may include adding annotations to drug treatment options for conditions described in, for example, the NCCN Clinical Practice Guidelines in Oncology™ or the American Society of Clinical Oncology (ASCO) Clinical Practice Guidelines.

[0147] Reports are generated for subjects with cancer regarding their k-mer repeat landscape, the combination of the k-mer repeat landscape and cfDNA mutation profile, genomic location mapping, and cfDNA mutation profile variations. These reports, compared to other profiles of subjects with known outcomes, can indicate that a specific cancer is aggressive and resistant to treatment. Subjects are monitored and retested after a period of time. If, at the end of the cycle, there are no changes in the k-mer repeat landscape, the combination of the k-mer repeat landscape and cfDNA mutation profile, the cfDNA mutation profile, mutation frequency, and / or fragmentation variation profile, this may indicate that the current treatment is ineffective. Comparisons are made with the k-mer repeat landscape, the combination of the k-mer repeat landscape and cfDNA mutation profile, and the cfDNA mutation profile of other subjects. For example, if changes in the k-mer repeat landscape or k-mer repeat sequences indicate cancer progression, the previously prescribed treatment regimen is no longer effective, and a new treatment regimen needs to be prescribed.

[0148] In some implementations, the system receives genetic information from a DNA sequencer. The program then determines specific k-mer repetitive sequence landscapes, combinations of k-mer repetitive sequence landscapes with cfDNA mutation profiles, or cfDNA alterations and their frequencies. These reports are submitted and accessed electronically via the Internet. Data analysis occurs at a site outside the participant's location. Reports are generated and transmitted to the participant's location. Participants can access reports reflecting their tumor burden via a networked computer.

[0149] While temporal information can be used to enhance information on the k-mer repetitive sequence landscape, the combination of the k-mer repetitive sequence landscape and cfDNA mutation profile, the cfDNA mutation profile, or mutation frequency, other consistent analytical methods can also be applied. In other implementations, historical comparisons can be used in conjunction with other consistent k-mer repetitive sequence landscapes, the combination of the k-mer repetitive sequence landscape and cfDNA mutation profile, the cfDNA mutation profile, mutation frequency, and / or fragmentation profile. Consistent k-mer repetitive sequence landscapes, the combination of the k-mer repetitive sequence landscape and cfDNA mutation profile, the cfDNA mutation profile, and mutation frequency can be normalized for control samples. Molecular measurements mapped to reference sequences can also be compared across the genome to identify regions in the genome where the k-mer repetitive sequence landscape, the combination of the k-mer repetitive sequence landscape and cfDNA mutation profile, the cfDNA mutation profile, and mutation frequency have changed or remained unchanged. Consistent analytical approaches include, for example, linear or nonlinear methods derived from digital communication theory, information theory, or bioinformatics for constructing consistent k-mer repeat sequences, k-mer repeat sequence landscapes, cfDNA mutation profiles, and mutation frequencies (e.g., voting, averaging, statistical, maximum a posteriori or maximum likelihood detection, dynamic programming, Bayesian, hidden Markov, or support vector machine methods, etc.). After determining the sequence read coverage, a stochastic modeling algorithm is applied to convert the normalized nucleic acid sequence read coverage for each window region into discrete copy number states. In some cases, this algorithm may include one or more of the following: hidden Markov models, dynamic programming, support vector machines, Bayesian networks, grid decoding, Viterbi decoding, expectation maximization, Kalman filtering methods, and neural networks.

[0150] Artificial neural networks (NNets) mimic the neural structure of the brain, designing networks that resemble "neurons." They process records one at a time, or in batch mode, and "learn" by comparing the record's classification (which is largely arbitrary at the beginning) with the record's known actual classification. In MLP-NNet, the error from the initial classification of the first record is fed back into the network and used to modify the network's algorithm a second time, and so on for many iterations. Neural networks use an iterative learning process where data samples (rows) are presented to the network one at a time, and the weights associated with the input values ​​are adjusted each time.

[0151] After all samples have been presented, the process typically restarts. During this learning phase, the network learns by adjusting the weights to predict the correct class label for the input samples. Due to the connections between units, neural network learning is also known as "connectionist learning." Advantages of neural networks include high tolerance for noisy data and the ability to classify patterns that have not yet been trained on. One neural network algorithm is the backpropagation algorithm, such as Levenberg-Marquadt. Once the network is built for a specific application, it is ready to be trained. To begin this process, the initial weights are randomly chosen. Then training or learning begins.

[0152] The network uses weights and functions in the hidden layers to process one record from the training data at a time, then compares the result with the expected output. The error is then propagated back through the system, causing adjustments to the weights applied to the next record. This process repeats over and over as the weights are continuously adjusted. During network training, the same set of data is processed multiple times as the connection weights are continuously optimized.

[0153] In one implementation, the training step of the machine learning unit on the training dataset can generate one or more classification models for application to test samples. These classification models can be applied to the test samples to predict a subject's response to a treatment intervention.

[0154] Comparing sequence coverage to control samples or reference sequences may aid in normalization across windows. In this embodiment, genomic DNA and / or cell-free DNA are extracted and isolated from readily available bodily fluids, such as blood. For example, cell-free DNA can be extracted using various methods known in the art, including but not limited to isopropanol precipitation and / or silica-based purification. DNA can be extracted from any number of subjects, such as subjects without cancer, subjects at risk of cancer, or subjects known to have cancer (e.g., by other means).

[0155] Following the isolation / extraction steps, a variety of different sequencing operations can be performed on cell-free polynucleotide samples. Before sequencing, the sample can be treated with one or more reagents (e.g., enzymes, unique identifiers such as barcodes, probes, etc.). In some cases, if the sample is treated with a unique identifier (such as a barcode), the sample or sample fragment can be labeled individually or subgrouped using that unique identifier. The labeled sample can then be used for downstream applications, such as sequencing reactions that trace individual molecules back to the parent molecule.

[0156] K-mer repeat sequences and / or cell-free polynucleotides can be labeled or tracked to allow for subsequent identification and tracing of specific polynucleotides. Assigning identifiers (e.g., barcodes) to each polynucleotide or polynucleotide subgroup allows for the assignment of unique identities to each sequence or sequence fragment. This allows for data acquisition from individual samples, rather than being limited to sample averages. In some instances, nucleic acids or other molecules derived from a single strand can share a common tag or identifier and can therefore be subsequently identified as being derived from that strand. Similarly, all fragments from a single strand of nucleic acid can be labeled with the same identifier or tag, allowing for subsequent identification of fragments from said parent strand. In other cases, gene expression products (e.g., mRNA) can be labeled to quantify expression, thereby enabling the counting of barcodes or combinations of barcodes with the sequences they are linked to. In still other cases, systems and methods can be used as PCR amplification controls. In such cases, multiple amplification products from a PCR reaction can be labeled with the same tag or identifier. If subsequent sequencing of the products reveals sequence differences, the differences between products with the same identifier can be attributed to PCR errors. Furthermore, individual sequences can be identified based on the sequence data characteristics of the read length itself. For example, the detection of unique sequence data at the start (initiation) and end (termination) portions of each sequencing read can be used individually, or in combination with the length or number of base pairs of the unique sequence for each sequencing read, to assign unique identities to individual molecules. Fragments from a single strand of nucleic acid that have been assigned unique identities can thus allow for subsequent identification of fragments from the parent strand. This can be combined with bottlenecking of the initial starting genetic material to limit diversity.

[0157] Typically, the methods and systems described in this article can be used to prepare genomic and / or cell-free polynucleotide sequences for downstream sequencing reactions. Common sequencing methods include next-generation sequencing (NGS), classic Sanger sequencing, whole-genome bisulfite sequencing (WGSB), small RNA sequencing, and low-coverage whole-genome sequencing (lcWGS).

[0158] As used herein, the term "sequencing" refers to any of a variety of techniques used to determine the sequence of a biomolecule, such as nucleic acids, like DNA or RNA. Exemplary sequencing methods include, but are not limited to: targeted sequencing, single-molecule real-time sequencing, exome sequencing, RNA sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy termination sequencing, whole-genome sequencing, hybridization sequencing, pyrosequencing, capillary electrophoresis, gel electrophoresis, duplex sequencing, cyclic sequencing, single-base extension sequencing, solid-phase sequencing, high-throughput sequencing, massively parallel signal sequencing, emulsion PCR, low-temperature denaturing co-amplification PCR (COLD-PCR), multiplex PCR, reversible dye terminator sequencing, paired-end sequencing, near-term sequencing, exonuclease sequencing, ligation sequencing, short-read sequencing, single-molecule sequencing, synthetic sequencing, real-time sequencing, reversible terminator sequencing, nanopore sequencing, 454 sequencing, Solexa genome analysis sequencing, SOLiD™ sequencing, MS-PET sequencing, and combinations thereof. In some implementations, sequencing can be performed using a gene analyzer, such as those commercially available from Illumina or Applied Biosystems. In some implementations, the sequencing method can be massively parallel sequencing, i.e., simultaneously (or rapidly sequentially) sequencing any one of at least 100, 1,000, 10,000, 100,000, 1 million, 10 million, 100 million, or 1 billion polynucleotide molecules.

[0159] After sequencing, quality scores are assigned to reads. A quality score can be a representation of the read length indicating whether it is usable in subsequent threshold-based analyses. In some cases, the quality or length of some reads is insufficient for subsequent mapping steps. Sequencing reads with quality scores of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered from the dataset. In other cases, sequencing reads assigned quality scores of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered from the dataset. Genomic fragment reads that meet the specified quality score thresholds are mapped to a reference genome, or a reference sequence known to be free of mutations. After mapping alignment, mapping scores are assigned to the sequence reads. A mapping score can be a representation of the read mapped back to the reference sequence, indicating whether each position is uniquely mappable. In some cases, reads can be sequences unrelated to mutation analysis. For example, some sequence reads may originate from contaminant polynucleotides. Sequencing reads with mapping scores of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered from the dataset. In other cases, sequencing reads with assigned mapping scores below 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered from the dataset. For each mappable base, bases that do not meet the minimum mappability threshold or are of low quality can be replaced with the corresponding base found in the reference sequence.

[0160] A variety of cancers can be detected using the methods and systems described herein. Like most cells, cancer cells can be characterized by their turnover rate, i.e., the death of old cells and their replacement by new cells. Typically, dead cells in contact with the vascular system in a given subject can release DNA or DNA fragments into the bloodstream. This is also true for cancer cells at different stages of disease. Depending on the stage of disease, cancer cells can also be characterized by various genetic abnormalities such as copy number variations and mutations. This phenomenon can be used to detect the presence or absence of cancer in an individual using the methods and systems described herein.

[0161] The types and number of cancers that can be detected may include, but are not limited to: leukemia, brain cancer, lung cancer, skin cancer, nasal cancer, laryngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, colorectal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, solid tumors, heterogeneous tumors, homogeneous tumors, etc.

[0162] Furthermore, the systems and methods described herein can also be used to help characterize certain cancers. The genetic data generated by the systems and methods disclosed herein can help practitioners better characterize specific forms of cancer. Typically, cancers are heterogeneous in both composition and stage. Genetic profiling data can allow for the characterization of specific cancer subtypes, which can be important for the diagnosis or treatment of that specific subtype. This information can also provide subjects or practitioners with clues about the prognosis of a particular type of cancer.

[0163] The systems and methods provided herein can be used to monitor known cancers or other diseases in a specific subject. This allows subjects or practitioners to adjust treatment plans based on disease progression. In this instance, the systems and methods described herein can be used to construct a genetic profile of a specific subject's disease process. In some cases, cancer can progress, becoming more aggressive and genetically unstable. In other instances, cancer may remain benign, inactive, or dormant. The systems and methods disclosed herein may help determine disease progression.

[0164] Furthermore, the systems and methods described in this article may help determine the efficacy of specific treatment regimens. In one instance, if treatment is successful, a successful treatment regimen actually increases the number of copy number variations or mutations detected in the subject's blood because more cancer cells die and release DNA. In other instances, this may not occur. In another instance, certain treatment regimens may be correlated with the genetic profile of cancer over time. This correlation may help in selecting therapies. Additionally, if cancer remission is observed after treatment, the systems and methods described in this article may help monitor residual disease or disease recurrence.

[0165] Data is transmitted to a computer for processing via a direct connection or the Internet. The data processing aspects of the system can be implemented using digital electronic circuits, or in computer hardware, firmware, software, or a combination thereof. The data processing apparatus of this disclosure can be implemented in a computer program product tangibly contained in a machine-readable storage device for execution by a programmable processor; and the data processing method steps of this disclosure can be executed by a programmable processor that executes instruction programs to achieve the functions of this disclosure by manipulating input data and generating output. The data processing aspects of this disclosure can advantageously be implemented in one or more computer programs executable on a programmable system including at least one programmable processor, at least one input device, and at least one output device, the at least one programmable processor being connected to receive data and instructions from and send data and instructions to the data storage system. Each computer program can be implemented in a high-level programming or object-oriented programming language, or in assembly language or machine language, if desired; and in any case, the language can be a compiled language or an interpreted language. For example, suitable processors include both general-purpose and special-purpose microprocessors. Typically, the processor receives instructions and data from read-only memory and / or random access memory. Storage devices suitable for tangibly representing computer program instructions and data include all forms of non-volatile memory, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM disks. Any of the foregoing can be supplemented by or integrated into an ASIC (Application-Specific Integrated Circuit).

[0166] To provide interaction with the user, this method can be implemented using a computer system having a display device for displaying information to the user, such as a monitor or LCD (liquid crystal display) screen, and an input device through which the user can provide input to the computer system, such as a keyboard, a two-dimensional pointing device (such as a mouse or trackball), or a three-dimensional pointing device (such as a data glove or gyro mouse). The computer system can be programmed to provide a graphical user interface through which computer programs interact with the user. The computer system can also be programmed to provide virtual reality or three-dimensional display interfaces.

[0167] Example Example 1: Landscape of whole-genome repetitive sequences in cancer and cell-free DNA Changes in repetitive sequences have long been considered to be associated with cancer development. Transposable elements are thought to regulate gene expression, and their movement in cancer may be driven by the loss of their silencing mechanisms through global hypomethylation. 7 This movement can lead to oncogene activation and genomic instability.8 The type of repeating sequence shows differential enrichment at structural breakpoints. For example, tandem repeats are enriched at copy number variation (CNV) breakpoints, while... Alu Repeated sequences are enriched in missing and repeat breakpoints. 9 The aforementioned variations in repetitive elements are a widespread feature of cancer genomes, but variations in individual element types have been observed in different cancer types. 10,11 Transposable elements act as enhancers of dysregulated tissue-specific transcription factors in cancer, while typical tandem repeat amplification is associated with gene regulation and varies substantially between primary tumor sites. 12,13 Instability and amplification of repetitive sequences around and in the centromere region of cancer patients may drive erroneous chromosome segregation and other structural changes. 14-18 These changes are associated with lower overall survival. 19 .

[0168] With the development of liquid biopsy technologies for human cancer detection and whole-genome characterization, the analysis of repetitive sequences in cfDNA has begun. Retrotransposon elements and non-telomeric satellite DNA have been shown to be highly present in cfDNA. 20 However, liquid biopsies aimed at incorporating repetitive sequences are limited to determining overall cfDNA levels or assessing aneuploidy. 21-24 Despite the aforementioned progress, a systematic analysis of the overall summaries of repetitive sequences in any human cancer has not yet been conducted, primarily due to the inability to identify and quantify repetitive sequences on a genome-wide basis.

[0169] To address these challenges, ARTEMIS (Analysis of Repetitive Elements in Disease) was developed as an alignment-independent, genome-wide method for analyzing repetitive sequence landscapes in short-read sequencing. ARTEMIS evaluates over 1200 individual repetitive sequence types, which occur genome-wide and span 57 subfamilies across 6 families (satellite sequences, RNA elements, transposable elements, LINE, SINE, LTR). In this study, ARTEMIS was used to show that repetitive sequence landscapes are enriched in genes that are pervasively variable in human cancers, and that tumor-specific changes in repetitive sequences reflect a combination of structural and epigenetic changes in the cancer genome. Genome-wide repetitive sequence landscape analysis using ARTEMIS can be performed using low-coverage whole-genome sequencing, allowing for the analysis of repetitive sequence landscapes in cfDNA for the detection of human cancers.

[0170] result K-mers were retrieved from whole-genome repetitive elements. To develop ARTEMIS, we performed a de novo search of short sequences (k-mers) because we assumed these sequences had sufficient complexity to identify different types of repetitive elements in the genome. Figure 1 and Figure 6 For example, a 24-bp k-mer sequence can theoretically distinguish 281 trillion (4) 24 (Number of sequences). The telomere-to-telomere reference genome (chm13), recently assembled from long-read sequencing, was used. 5,25 We found 4.73 billion 24-bp k-mer sequences in the genome, and a total of 4.17 billion of these 24-bp k-mers are unique to repetitive elements. As associated repetitive elements have diversified in their sequence composition over time, we identified 1.1 billion 24-bp k-mers that uniquely define each of the 1266 recently identified repetitive types. To be included in this set, a k-mer cannot appear in a non-repetitive region of the genome, nor in multiple repetitive types. Each of the 1266 repetitive sequence types analyzed is defined by a 24-bp k-mer with a median of 43,297 and spanning an average of 2.6 Mb of genome sequence. Figure 7 We further incorporated 58k 24-bp k-polymers of enhanced annotations from 14 human satellite sequence subtypes. 26 These 1.1 billion k-mers, representing 1280 repetitive sequence types, were found on all chromosomes, and 98% of the k-mers used were observed only once in the T2T reference genome (Figures 8A to 8C). These k-mers also represent genomic regions that cannot be aligned with high quality in typical short-read next-generation sequencing, such as human satellite sequences. This allows ARTEMIS to consider the entire genome, not just the approximately 60-85% of reads that can be aligned with high quality in next-generation sequencing. 27, 28 To verify that these repetitive sequence landscape k-mers are not confused with human-related microbial genomes. 29,30 We examined 1545 reference genomes representing common microorganisms and found a median of 190 ARTEMIS k-mers (range 0-2550) per microbial genome, representing <0.0003% of the 1.1 billion possible k-mers counted in ARTEMIS across all cases. For individual sample analysis, we defined the k-mer repetitive sequence landscape as the count of all k-mers in the sequencing sample that match each of the 1280 repetitive sequence types divided by the number of aligned sequence reads. Since variations in repetitive sequences can occur during the initiation of cancer and other diseases, this comprehensive aggregation of repetitive features can be used to train machine learning models to distinguish between normal and disease states in a genome.

[0171] Genome-wide enrichment of repeat k-mers in cancer-related genes and pathways We first examined the genome-wide distribution of 1.1 billion k-mers that define unique repetitive sequence types and found that repetitive elements are enriched in genes that are commonly altered in human cancers. Among the 736 genes in the COSMIC Cancer Driver Gene Survey... 31 We identified 474 repetitive k-mer sequences with a higher-than-expected number in their exons or introns (normalized enrichment score = 9.12, false discovery rate q = 0.00), including genes amplified, deleted, and rearranged in cancer (normalized enrichment scores of 1.92, 4.13, and 6.45, respectively, and false discovery rates q = 0.01, 0.00, and 0.00, respectively). This enrichment remained significant even after adjusting for the size of these genes (Fig. 8D) and reflected an average 15-fold increase in repetitive k-mers in these regions (p < 2.2e-16, Wilcoxon signed-rank test). In contrast, analysis of the same number of randomly selected genes in the genome did not show enrichment of repetitive k-mer sequences (normalized enrichment score = -1.0, false discovery rate q = 0.97). Repetitive k-mer sequences were also significantly increased in pathways that are generally dysregulated in cancer, including cell adhesion, growth, and signal transduction, as well as in cancer type-specific gene sets. In summary, these observations on the localization of k-mers of repetitive sequences suggest that, during tumorigenesis, alterations in key genes affecting human cancer oncogenic pathways can be selectively influenced by genomic changes associated with repetitive sequences.

[0172] The landscape of k-mer repetitive sequences is altered in cancer genomes. Given the vast number of genomic changes occurring during tumorigenesis, we used short-read next-generation sequencing to assess whether the k-mer repetitive sequence landscape is altered in cancer. Considering the challenge of distinguishing highly correlated repetitive sequences, we integrated short-read whole-genome sequencing (WGS) data with typical sequence error rates and analyzed them in an alignment-independent manner. We found that despite potential sequencing errors, the high complexity of k-mer sequences allows them to remain specific to their defined repetitive sequence families (98% of the counted k-mers were found in reads derived from their true repetitive sequence types). Figure 10 and Figure 11 ).

[0173] We analyzed data from the whole-genome pan-cancer analysis (PCAWG). 32Matched tumors and normal tissues from 333 cancer patients, including those with lung cancer (n=86), colorectal cancer (n=60), breast cancer (n=91), liver cancer (n=54), and ovarian cancer (n=42), were used to determine whether genome-wide k-mer counts for specific repetitive element types were altered in tumors. An average of 24.2 billion total k-mers, representing 1280 repetitive elements, were identified in each sample sequenced at 30–60X. Compared to their matched normal tissues, a median of 827 repetitive elements (range 249–1246) showed significantly increased or decreased k-mer counts in cancer (Figure 2A, Methods). Nearly two-thirds (820 of 1280) of the altered elements had not been previously observed to be altered in human cancers (Figure 2A, Table 1). Although alterations were frequently observed in elements within LTRs, transposable elements, and RNA elements, the highest rates of alteration were observed in elements derived from satellite sequences, LINEs, and SINEs. Nearly a quarter of the elements studied originated from the maximal repeat subfamily ERV1 of the LTR, which is hypothesized to aberrantly activate transcription in cancer cells through tumor expansion adaptation. 33 On average, although individual altered elements varied across tissue types, more than half of the 300 LTR ERV1 elements were altered in all five tumor types studied. While 21 ERV1 alterations have been previously described... 34 We observed changes in an additional 279 novel ERV1 variants across the cancer types analyzed (Tables S6 and S7). These changes are similar to other genomic variations in cancer genomes. 35,36 The changes in the k-mer repeat sequence landscape are highly complex, and no two patients studied had exactly the same set of changes.

[0174] We analyzed matched tumors and normal tissues from 525 cancer patients from the genome-wide pan-cancer analysis (PCAWG) (Table S4) Nature578, 82-93 (2020) included patients with breast cancer (n=91), lung cancer (n=86), colorectal cancer (n=60), liver cancer (n=54), thyroid cancer (n=48), head and neck squamous cell carcinoma (n=44), ovarian cancer (n=42), gastric cancer (n=38), bladder cancer (n=23), cervical cancer (n=20), and prostate cancer (n=19), and determined whether genome-wide k-mer counts for specific repeat element types were altered in tumors. An average of 22.4 billion total k-mers, representing 1280 repeat elements, were identified in each sample sequenced from 30 to 60X. Compared to their matched normal tissues, a median of 807 repeat elements (range 246–1280) showed increased or decreased k-mer counts in tumors (Figure 2A, Tables S5 and S6). Nearly two-thirds (820 out of 1280) of the altered elements had not been previously observed to change in human cancer (Figure 2A, Table 1). Although changes were also frequently observed in elements within LTRs, transposable elements, and RNA elements, the highest rates of alteration were observed in elements derived from satellite sequences, LINEs, and SINEs. Nearly a quarter of the elements studied originated from the maximally repetitive subfamily ERV1 of the LTR (Table 1), which is hypothesized to aberrantly activate transcription in cancer cells through tumor expansion adaptation (i.e., the process by which reactivated transposable elements drive oncogene expression). 33 On average, although individual altered elements varied across tissue types, over 40% of the 300 LTR ERV1 elements were altered across all 12 tumor types studied. While variations in 21 ERV1 elements have been previously described... 34 However, we observed an additional 279 new ERV1 changes across the cancer types analyzed. This is similar to other large-scale changes in the cancer genome. 35,36 The changes in the k-mer repeat sequence landscape are highly complex, and no two patients studied had exactly the same set of changes.

[0175] We hypothesize that changes in the k-mer repetitive sequence landscape are partly related to structural changes occurring during tumorigenesis, such as chromosome copy number changes, rearrangements, or focal amplification or deletion. Therefore, we found that k-mer counts reflect genome-wide chromosome arm gains and deletions in the analyzed tumors (r=0.81; p<2.2e-16, Spearman correlation). Figure 12Furthermore, tumors with greater variations in the k-mer repetitive sequence landscape exhibited higher chromosomal instability, as reflected by other measures of overall genomic entropy, loss of heterozygosity (LOH), genomic nonmodal ploidy fraction, and whole-genome structural changes (r=0.34, p=1.7e-9; r=0.2, p=3.2e-4; r=0.32, p=2.9e-8, Spearman correlation, respectively) (Figure 2A). In contrast, tumor mutational burden (TMB), a measure of single-base sequence changes in a single cancer, was not correlated with whole-genome k-mer repetitive sequence landscape changes (r=0.081, p=0.166, Spearman correlation).

[0176] Rearrangements caused by copy number neutral translocations, inversions, duplications, or deletions may be facilitated by homologous sequence swapping. 37 Analysis of the locations of repetitive elements and tumor-specific sequence breakpoints in 333 analyzed samples identified 170 elements enriched at breakpoint sites, including LINE, SINE, LTR, TE, and RNA elements, and included 92 elements previously not shown to be altered in cancer, suggesting these elements may play a role in promoting these structural changes (Figure 2B). Analysis of focal amplifications of five or more copies showed that the content of repetitive elements in all subfamilies was associated with increased amplicon copy number (r=0.91, p<2.2e-16, Spearman correlation analysis). As an example, in breast tumors... ERBB2 Analysis of the surrounding 1 Mb region (where gains are known to occur) revealed a significant increase in 14 repeat elements, including eight elements previously undocumented in cancer (Fig. 2C, 13A). Similarly, driver genes are found in squamous cell lung cancer. PIK3CA and SOX2 Gain in a region of approximately 30 Mb on chromosome 3q 38,39 The results show an increase in k-mers for repeating elements that overlap with these regions, including nine previously unknown elements that are altered in cancer (Figures 13B and 13C).

[0177] Changes in the repetitive sequence landscape cannot be fully explained by chromosomal or focal copy number changes or genomic rearrangements. After comparing changes in repetitive elements observed in the analyzed cancer genomes with segmental copy number alterations, we determined that 88% of repetitive sequence changes (median 706 changes per tumor, ranging from 232 to 1246) were on the order of magnitude larger than changes expected solely due to copy number gains and deletions. Figure 14A set of 242 elements exhibited changes that could not be explained by copy number variations in at least 75% of the tumors studied. These changes included a reduction in LINE-1-mediated deletions of k-mer elements in squamous cell lung cancer, and lower-than-expected levels of repetitive sequences in copy number regions, consistent with the concept that such repetitive sequences may undergo deletion when they contribute to increases in nearby genomic content (Figure 2D). Figure 12 , Figure 15 ) 37, 40 Overall, these analyses highlight the ability of k-mer repetitive sequence landscapes to detect and characterize a wide range of structural changes in human cancers, including large chromosomal alterations, pervasive amplification or deletion of driver gene regions, and changes that directly target repetitive sequences.

[0178] To assess the potential clinical significance of alterations in repetitive elements in the cancer genome, we investigated whether tumor alterations in any of these families were associated with overall survival or progression-free survival in patients on the PCAWG dataset. We found that alterations in SINE elements were associated with longer overall survival (p=0.005) and progression-free survival (p=0.003) (Figure 2E), and remained significant for overall survival even after adjusting for tumor type (Figures 16A and 16B). While the overall SINE alteration load appears to be a pan-cancer biomarker, the distribution of individual element alterations differed across cancer types (Figure 16C). Interestingly, this variation in patient outcomes was not observed in other non-repetitive sequence whole-genome indicators, including genomic entropy, LOH, or nonmodal ploidy fraction (Figures 17A and 17B). Our observations, along with previous analyses suggesting that reactivation and increase of repetitive elements in the cancer genome may lead to an immune response, are consistent with findings that suggest reactivation and increase of repetitive elements in the cancer genome may contribute to an immune response. 41-43 Or increased genomic instability 44 Both mechanisms can reduce tumor cell adaptability and lead to improved patient outcomes.

[0179] Although repeating elements exhibit phylogenetic variability among different individuals ( Figure 18 However, the ARTEMIS score generated from the machine learning model of k-merged repeat sequence landscapes of mismatched tumor and normal samples was able to distinguish 333 PCAWG tumors from normal tissues with high performance across all cancer types analyzed (AUC range: = 0.93 to >0.99). Figure 19 ).

[0180] landscape of k-mer repeat sequences in cfDNA We sought to determine whether the approach we used to characterize the repetitive sequence landscape could be used to assess circulating cfDNA. Theoretically, detecting repetitive sequences using low-coverage whole-genome sequencing is feasible because ARTEMIS aggregates a large number of k-mer-defined examples of repetitive elements across the entire genome while maintaining sufficient granularity to identify disease-specific genomic features. As a first step in this analysis, we determined that the repetitive sequence landscape in PCAWG remained highly consistent even when these were resampled to different sequencing depths ranging from >60X to 1X coverage. Figure 20 and 21 We further found that the k-mer repetitive sequence landscape in cfDNA was consistent across different sequencing platforms and experimental batches. Figure 22 ).

[0181] To determine whether low-coverage cfDNA sequencing could quantify the repetitive sequence landscape in plasma, we first examined satellite sequence families known to be distributed on the Y chromosome (chrY) in male and female ensembles (n=158). In the plasma of males (n=87), the k-mer count of human satellite sequence types known to be found only or primarily on chrY was significantly higher than in females (n=71) (all types p<2.2e-16) (Figure 3A), while there was no significant difference between males and females in satellite sequences not found on chrY (all types p>0.1).

[0182] We then identified the most variable repetitive elements in PCAWG tumors and evaluated the presence of these elements in the cfDNA of prospectively collected individuals from a diagnostic cohort of patients at risk for lung or liver cancer (lung cancer cohort n=287; liver cancer cohort n=208) who had previously undergone sequencing. 27, 45 In each cohort, the increase or decrease of many repetitive sequence k-mers observed in tumors compared to plasma from cancer-free individuals was evident in the plasma of patients with squamous or adenocarcinoma subtypes of lung or liver cancer (Figure 3B). These include changes in elements previously documented to play a role in cancer, such as LINE L1 elements, as well as changes in newly discovered elements currently revealed to be altered in cancer, including elements from subfamilies such as DNA-hAT-Charlie and LTR ERV1, ERVL-MaLR, and ERVL.

[0183] Because genome-wide chromatin and epigenetic changes can alter the representativeness of circulating cfDNA fragments. (27, 28, 45-49)We hypothesize that the repetitive sequence landscape in cfDNA may differ from the expected repetitive sequence content in genomic DNA. We have previously shown that cfDNA fragmentation profiles reflect the open and closed chromatin states across the entire genome. (28, 45) Here, we analyzed cfDNA from 158 cancer-free individuals and showed differential density of repeat element types in regions with different histone markers. Figure 4 (A and B), and individual cfDNA fragments originating from regions of chromatin with active transcription or activated histone markers are shorter and exhibit lower coverage in plasma ( Figure 4 (C and D). Overall, for regions with high density of activating chromatin histone markers, the count of repetitive sequence landscape k-mers in cfDNA was lower than for regions with low density of these markers, while the opposite was observed for repressive histone markers. Figure 4 E). Whole-genome simulations show that the repetitive sequence landscape in cfDNA can be influenced by both tumor-specific epigenome and genomic variations. Figure 31 ).

[0184] ARTEMIS k-merged repeat sequence landscape analysis for cancer detection and monitoring Given its ability to identify repetitive sequence landscape variations in cfDNA, we evaluated the potential of the ARTEMIS method for non-invasive cancer detection. Figure 23 We previously described the use of sensitive and readily available whole-genome cfDNA fragmentation assay (DELFI) for lung and liver cancer screening in high-risk populations. 27, 45 Here, we used high-density histone markers (whose differential expression affects the representativeness of repetitive sequences), Figure 24 The k-mer repetitive sequence landscape and epigenetic profile of the cfDNA region were analyzed to generate cross-validated ensemble machine learning models for detecting lung cancer in a prospective diagnostic cohort (n=287) or liver cancer in a high-risk group (n=208), respectively. 27, 45 ARTEMIS classified lung cancer patients with an AUC of 0.82 (95% CI 0.77–0.87) and when compared with DELFI genome-wide fragmentation features... 28When combined, the ARTEMIS-DELFI model classified lung cancer patients with an AUC of 0.91 (95% CI 0.88–0.95) (Figures 5A and 5B, Figure 25). Similar performance was observed in a cohort of individuals at risk for hepatocellular carcinoma, where ARTEMIS detected individuals with hepatocellular carcinoma in patients with cirrhosis or viral hepatitis with an AUC of 0.87 (95% CI 0.82–0.92), and the AUC increased to 0.91 (95% CI 0.87–0.95) when combined with DELFI (Figures 25 and 5B). Figure 26 We validated the locked ARTEMIS and ARTEMIS-DELFI models in an external cohort consisting of non-cancer individuals with high and average risk of lung cancer, as well as patients with lung cancer at all stages, and observed similar performance to that in the cross-validated training cohort (Figure 5C). We further applied these models to an independent cohort of patients with advanced lung cancer receiving tyrosine kinase inhibitor therapy. 50 Furthermore, it was demonstrated that ARTEMIS and combined ARTEMIS-DELFI scores were associated with the observed circulating tumor DNA (ctDNA) mutation allele fraction (MAF) during treatment (ARTEMIS: r=0.70, p=2.67e-12; ARTEMIS-DELFI: r=0.80, p<2.2e-16, Spearman correlation). Figure 27 ).

[0185] Finally, given the observed tumor-specific variations in the repetitive sequence landscape, we evaluated whether ARTEMIS could help identify the primary tissue in tumors or cfDNA samples from cancer patients. We first examined whether the k-mer repetitive sequence landscape could capture tissue-specific signals. We trained a machine learning model using the k-mer repetitive sequence landscape to distinguish tissue types and found that, despite relying solely on genomic features (which are generally considered to show fewer tissue-specific differences than transcriptomic and epigenetic features), it achieved a mean accuracy of 92% in classifying PCAWG tumors by primary tissue among the tumor types studied. This is consistent with the observation that while variations in the repetitive sequence landscape are a pan-cancer feature of the cancer genome, the specific repetitive elements altered vary from tumor type to tumor (Figure 2A). We then extended this approach to cfDNA and trained a cross-validated classification model on cfDNA from a multi-cancer cohort comprising 226 individuals with tumors of the breast, ovary, lung, colorectal, bile duct, stomach, or pancreas using ARTEMIS-DELFI. 28Here, despite the small number of samples used to train such classifiers, we found that ARTEMIS-DELFI correctly classified the detected patients among different cancer types, with average accuracy of 72% or 84% for the highest or top two predictions, respectively (Table 2).

[0186] ARTEMIS k-merged repeat sequence landscape analysis for cancer detection and monitoring Given its ability to identify changes in the repetitive sequence landscape in cfDNA, we evaluated the potential of the ARTEMIS method for non-invasive cancer detection. Figure 23 We previously described the use of sensitive and readily available whole-genome cfDNA fragmentation assay (DELFI) for lung and liver cancer screening in high-risk populations. 27, 45 Here, we use high-density histone markers (which differentially affect the representativeness of repetitive sequences), Figure 24 The k-mer repetitive sequence landscape and epigenetic profile in the cfDNA region of the lung cancer were used as features in a machine learning model to detect lung cancer in a diagnostic cohort (n=287) prospectively collected by the Danish Lung Cancer Screening Study (LUCAS) and liver cancer in a high-risk group (n=208). 27, 45 ARTEMIS classified lung cancer patients with an AUC of 0.82 (95% CI 0.78–0.87), and when compared with DELFI genome-wide fragmentation features... 28 When combined, the ARTEMIS-DELFI model classified lung cancer patients with an AUC of 0.91 (95% CI 0.88–0.94) (Figures 5A and 5B, Figure 25). Similar performance was observed in a cohort of individuals at risk for hepatocellular carcinoma, where ARTEMIS detected individuals with hepatocellular carcinoma in patients with cirrhosis or viral hepatitis with an AUC of 0.87 (95% CI 0.82–0.93), and the AUC increased to 0.90 (95% CI 0.86–0.94) when combined with DELFI. Figure 30 We validated the locked ARTEMIS and ARTEMIS-DELFI models in an external cohort consisting of non-cancer individuals with high and average risk of lung cancer (n=400) and patients with lung cancer at all stages (n=88), and observed similar performance to that in the cross-validated training cohort (Figure 5C). Figure 32 Analysis of an independent set of reserved patients (n=25) with a history of cancer from the LUCAS cohort using locked ARTEMIS and ARTEMIS-DELFI models showed that patients who had experienced cancer recurrence had higher scores compared to those who had not. Figure 33 We further applied these models to an independent cohort (n=19) of patients with advanced lung cancer who received tyrosine kinase inhibitor therapy. 50 Furthermore, it was demonstrated that ARTEMIS and combined ARTEMIS-DELFI scores were associated with the observed circulating tumor DNA (ctDNA) mutation allele fraction (MAF) during treatment (ARTEMIS: r=0.70, p=2.67x10). -12 ; ARTEMIS-DELFI: r=0.80, p<2.2e -16 (Spearman correlation). Analysis of ARTEMIS-DELFI scores at the first time point after treatment initiation (median = 6 days) identified patients with scores higher or lower than the pre-treatment median who had shorter or longer progression-free survival, respectively (median = 1.4 months for the high-score group vs. 8.9 months for the low-score group, p < 0.001, two-sided log-rank test) (Figure 34). discuss We have specifically demonstrated that ARTEMIS can reconstruct a genome-wide landscape of repetitive sequences reflecting potential changes in human cancer. These alterations reflect structural changes in the cancer genome, including focal amplification, deletion, copy number changes, and rearrangements, as well as direct changes in repetitive elements. Through this analysis, we found that repetitive elements are enriched in genes that are universally altered in human cancer, including at tumor-specific rearrangement breakpoints. Cancer-specific changes in the repetitive sequence landscape were observed across the entire genome, including previously unknown elements altered in human cancer. These elements may provide a potential basis for the widespread alteration of genes, pathways, and chromosomes, as well as genomic instability, in human cancer. Furthermore, the expansion or contraction of repetitive elements, which can now be fully identified, provides new methods for detecting and examining mechanisms influencing cancer and other diseases.

[0187] We found that changes in the repetitive sequence landscape can be detected in circulation, and that plasma signals are further altered by epigenetic changes in repetitive elements (affecting their susceptibility to fragmentation). We and others have previously shown that changes in chromatin accessibility, transcription factor binding, and methylation can alter the representativeness of cfDNA in the blood. 28, 45-48 In this study, we demonstrate that epigenetic states influenced by histone acetylation and methylation (leading to altered gene expression levels) have a profound impact on the size and coverage of cfDNA across different regions of the genome, including repetitive regions. These analyses suggest that the landscape of k-mer repetitive sequences in plasma can reveal both structural and epigenetic changes within the genome.

[0188] The use of repetitive sequence landscapes in cross-validation and external validation settings for cfDNA-based detection of lung cancer, liver cancer, and other cancers demonstrates that ARTEMIS, alone or in combination with other genome-wide features, can provide a pathway for non-invasive detection, surveillance, and primary tissue identification of cancer. Future work will be important in validating the ARTEMIS method for non-invasive detection of other cancers. One limitation of ARTEMIS is its reliance on assessing changes in the repetitive sequence landscape, which are phylogenetically inherent in individuals. 3, 14, 51-53 However, ARTEMIS improves early diagnosis by identifying genome-wide variations that may not be apparent in other liquid biopsy methods when tumor features such as mutations or arm-level changes go undetected. Future research will be valuable in characterizing the k-mer repetitive sequence landscape across diverse individuals, as the chm13 reference genome is derived from a single individual and comparisons with representative groups of healthy genotypes from different germline backgrounds can improve performance. Furthermore, the functional significance of variations within repetitive sequence families remains poorly understood and can be improved through further analysis in cancer and other disease states. Only about 43% of the k-mers used in the ARTEMIS method, occurring genome-wide, are located within approximately 28,000 known genes, and many repetitive sequence types in our landscape have not yet been studied in human cancers. Considering the size, diversity, and potential clinical significance of these genomic regions, our study provides unique insights into cancer genomics and a proof-of-concept for the practicality of genome-wide k-mer repetitive sequence landscapes as tissue- and blood-based biomarkers.

[0189] Table 1: Subfamilies of repetitive elements altered in cancer as identified by ARTEMIS

[0190]

[0191]

[0192] *The test is based on the ARTEMIS-DELFI test with 90% specificity. Lung cancer patients include those with additional lung cancer who have received prior treatment.

[0193] References (used for the disclosures shown by superscript numbers in Embodiment 1 above, corresponding to the reference numbers shown in the list below) 1. M.R. Vollger, X. Guitart, P.C. Dishuck, L. Mercuri, W.T. Harvey,A. Gershman, M. Diekhans, A. Sulovari, K.M. Munson, A.P. Lewis, K. Hoekzema,D. Porubsky, R. Li, S. Nurk, S. Koren, K.H. Miga, A.M. Phillippy, W. Timp, M.Ventura, E.E. Eichler, Segmental duplications and their variation in acomplete human genome. Science. 376, eabj6965 (2022). 2. S. Aganezov, S.M. Yan, D.C. Soto, M. Kirsche, S. Zarate, P.Avdeyev, D.J. Taylor, K. Shafin, A. Shumate, C. Xiao, J. Wagner, J. McDaniel,N.D. Olson, M.E.G. Sauria, M.R. Vollger, A. Rhie, M. Meredith, S. Martin, J.Lee, S. Koren, J.A. Rosenfeld, B. Paten, R. Layer, C.-S. Chin, F.J.Sedlazeck, N.F. Hansen, D.E. Miller, A.M. Phillippy, K.H. Miga, R.C. McCoy,M.Y. Dennis, J.M. Zook, M.C. Schatz, A complete reference genome improvesanalysis of human genetic variation. Science. 376, eabl3533 (2022). 3. S.J. Hoyt, J.M. Storer, G.A. Hartley, P.G.S. Grady, A. Gershman,L.G. de Lima, C. Limouse, R. Halabian, L. Wojenski, M. Rodriguez, N.Altemose, A. Rhie, L.J. Core, J.L. Gerton, W. Makalowski, D. Olson, J. Rosen,A.F.A. Smit, A.F. Straight, M.R. Vollger, T.J. Wheeler, M.C. Schatz, E.E.Eichler, A.M. Phillippy, W. Timp, K.H. Miga, R.J. O’Neill, From telomere totelomere: The transcriptional and epigenetic state of human repeat elements.Science. 376, eabk3112 (2022). 4. A. Gershman, M.E.G. Sauria, X. Guitart, M.R. Vollger, P.W. Hook,S.J. Hoyt, M. Jain, A. Shumate, R. Razaghi, S. Koren, N. Altemose, G.V.Caldas, G.A. Logsdon, A. Rhie, E.E. Eichler, M.C. Schatz, R.J. O’Neill, A.M.Phillippy, K.H. Miga, W. Timp, Epigenetic patterns in a complete humangenome. Science. 376, eabj5089 (2022). 5. S. Nurk, S. Koren, A. Rhie, M. Rautiainen, AV Bzikadze, A.Mikheenko, MR Vollger, N. Altemose, L. Uralsky, A. Gershman, S. Aganezov,SJ Hoyt, M. Diekhans, GA Logsdon, M. Alonge, SE Anton, Bouffard, M. G. Bouffard. SY Brooks, GV Caldas, N.-C. Chen, H. Cheng, C.-S. Chin, W. Chow, LG de Lima, PC Dishuck, R. Durbin, T. Dvorkina, ITFiddes, G. Formenti, RS Fulton, A. Fungtammasan, E. Garrison, PGS Grady,TA Graves-Lindsay, IM Hall, NF Hansen, GA Hartley, M. Haukness, K. Howe, M. W. Hunka, C. Jain, M. Jainka, M. Jain. ED Jarvis, P. Kerpedjiev, M.Kirsche, M. Kolmogorov, J. Korlach, M. Kremitzki, H. Li, VV Maduro, T.Marschall, AM McCartney, J. McDaniel, DE Miller, JC Mullikin, EWMyers, ND Olson, B. Paten, P. Pevsky, PA, D. Pevsky, D. Pevzner. T.Potapova, EI Rogaev, JA Rosenfeld, SL Salzberg, VA Schneider, FJSedlazeck, K. Shafin, CJ Shew, A. Shumate, Y. Sims, AFA Smith, DC Soto,I. Sovi, JM Storer, A. Streets, BA Sullivan,F.Thibaud-Nissen, J.Torrance, J. Wagner, B.P. Walenz, A. Wenger, J.M.D. Wood, C. Xiao, S.M. Yan,A.C. Young, S. Zarate, U. Surti, R.C. McCoy, M.Y. Dennis, I.A. Alexandrov,J.L. Gerton, R.J. O’Neill, W. Timp, J.M. Zook, M.C. Schatz, E.E. Eichler,K.H. Miga, A.M. Phillippy, The complete sequence of a human genome. Sci NewYork N Y. 376, 44-53 (2022). 6. N. Altemose, GA Logsdon, AV Bzikadze, P. Sidhwani, SALangley, GV Caldas, SJ Hoyt, L. Uralsky, FD Ryabov, CJ Shew, MEGSauria, M. Borchers, A. Gershman, A. Mikheenko, VA Shepelev, T. Dvorkina,O. Kunyavskaya, MR Vollger, A. Rhie, AM McCartney, M. Asri, R. Lorig-Roach, K. Shafin, JK Lucas, S. Aganezov, D. Olson, LG de Lima, T.Potapova, GA Hartley, M. Haukness, P. Kerpedjiev, F. Gusev, K. Tigyokyis, S. B. Yourongyis, A.B. S. Nurk, S. Koren, SR Salama, B. Paten, EI Rogaev, A.Streets, GH Karpen, AF Dernburg, BA Sullivan, AF Straight, TJWheeler, JL Gerton, EE Eichler, AM Phillippy, W. Timp, MY Dennis,RJ O'Neill, JM Zook, Zook, PA, M.M. Diekhans, CHLangley, IA Alexandrov, KH Miga, Complete genomic and epigenetic maps ofhuman centromeres. Sci New York N Y. 376, eabl4178 (2022). 7. L. Grégoire, A. Haudry, E. Lerat, The transposable elementenvironment of human genes is associated with histone and expression changesin cancer. Bmc Genomics. 17, 588 (2016). 8. S.L. Anwar, W. Wulaningsih, U. Lehmann, Transposable Elements inHuman Cancer: Causes and Consequences of Deregulation. Int J Mol Sci. 18, 974(2017). 9. P. Bose, K.E. Hermetz, K.N. Conneely, M.K. Rudd, Tandem Repeatsand G-Rich Sequences Are Enriched at Human CNV Breakpoints. Plos One. 9,e101607 (2014). 10. S. Sato, M. Gillette, P. R. de Santiago, E. Kuhn, M. Burgess, K.Doucette, Y. Feng, C. Mendez-Dorantes, P. J. Ippoliti, S. Hobday, M. A.Mitchell, K. Doberstein, S. M. Gysler, M. S. Hirsch, L. Schwartz, M. J.Birrer, S. J. Skates, K. H. Burns, S. A. Carr, R. Drapkin, LINE-1 ORF1p as acandidate biomarker in high grade serous ovarian carcinoma. Sci Rep-uk .13,1537(2023). [ PMC free article ] [ PubMed ] 11. Martinez JG, Perez-Escuredo P, Castro-Santos C, CAMarcos, Pendas JLL, Fraga MF, Hermsen MA. Cell Oncol .35, 259–267(2012). 12. K. Karttunen, D. Patel, J. Xia, L. Fei, K. Palin, L. Aaltonen, B. Sahu, Biorxiv, in press, doi:10.1101 / 2022.12.16.520732. 13. Erwin GS, Gursoy G, Al-Abri R, Suriyaprakash A, Dolzhenko E, Zhu K, Hoerner CR, White SM, Ramirez L, Vadlakonda A, von Kraut K, Park J, Brannon CM, Sumano DA, Kirtikar RA, Erwin AA,TJ Metzner, RKC Yuen, AC Fan, JT Leppert, MA Eberle, M.Gerstein, MP Snyder, Recurrent repeat expansions in human cancer genomes. Nature .613, 96-102(2023). 14. L. G. de Lima, E. Howe, V. P. Singh, T. Potapova, H. Li, B. Xu,J. Castle, S. Crozier, C. J. Harrison, S. C. Clifford, K. H. Miga, S. L.Ryan, J. L. Gerton, PCR amplicons identify widespread copy number variationin human centromeric arrays and instability in cancer. Cell Genom .1, 100064(2021). 15. A. K. Saha, M. Mourad, M. H. Kaplan, I. Chefetz, S. N. Malek, R.Buckanovich, D. M. Markovitz, R. Contreras-Galindo, The Genomic Landscape ofCentromeres in Cancers. Sci Rep-uk .9, 11259(2019). 16. S. Decombe, F. Loll, L. Caccianini, K. Affannoukoué, I. Izeddin,J. Mozziconacci, C. Escudé, J. Lopes, Epigenetic rewriting at centromeric DNArepeats leads to increased chromatin accessibility and chromosomalinstability. Epigenet Chromatin .14, 35(2021). 17. P. Ly, S. F. Brunner, O. Shoshani, D. H. Kim, W. Lan, T.Pyntikova, A. M. Flanagan, S. Behjati, D. C. Page, P. J. Campbell, D. W.Cleveland, Chromosome segregation errors generate a diverse spectrum ofsimple and complex genomic rearrangements. Nat Genet .51, 705-715(2019). 18. K. Ichida, K. Suzuki, T. Fukui, Y. Takayama, N. Kakizawa, F.Watanabe, H. Ishikawa, Y. Muto, T. Kato, M. Saito, K. Futsuhara, Y. Miyakura,H. Noda, T. Ohmori, F. Konishi, T. Rikiyama, Overexpression of satellitealpha transcripts leads to chromosomal instability via segregation errors atspecific chromosomes. Int J Oncol .52, 1685-1693(2018). 19. F. Bersani, E. Lee, P. V. Kharchenko, A. W. Xu, M. Liu, K. Xega,O. C. MacKenzie, B. W. Brannigan, B. S. Wittner, H. Jung, S. Ramaswamy, P. J.Park, S. Maheswaran, D. T. Ting, D. A. Haber, Pericentromeric satelliterepeat expansions through RNA-derived DNA intermediates in cancer. Proc National Acad Sci .112, 15148-15153(2015). 20. S. Grabuschnig, J. Soh, P. Heidinger, T. Bachler, E. Hirschböck,I. R. Rodriguez, D. Schwendenwein, C. W. Sensen, Circulating cell-free DNA ispredominantly composed of retrotransposable elements and non-telomericsatellite DNA. J Biotechnol .313, 48-56(2020). 21. U. Gezer, A. J. Bronkhorst, S. Holdenrieder, The Utility ofRepetitive Cell-Free DNA in Cancer Liquid Biopsies. Diagnostics .12, 1363(2022). 22. C. Douville, J. D. Cohen, M. Ptak, M. Popoli, J. Schaefer, N.Silliman, L. Dobbyn, R. E. Schoen, J. Tie, P. Gibbs, M. Goggins, C. L.Wolfgang, T.-L. Wang, I.-M. Shih, R. Karchin, A. M. Lennon, R. H. Hruban, C.Tomasetti, C. Bettegowda, K. W. Kinzler, N. Papadopoulos, B. Vogelstein,Assessing aneuploidy with repetitive element sequencing. Proc National Acad Sci .117, 4858-4863(2020). 23. C. Douville, S. Springer, I. Kinde, J. D. Cohen, R. H. Hruban, A.M. Lennon, N. Papadopoulos, K. W. Kinzler, B. Vogelstein, R. Karchin,Detection of aneuploidy in patients with cancer through amplification of longinterspersed nucleotide elements (LINEs). Proc National Acad Sci .115, 1871-1876(2018). 24. C. Rago, D. L. Huso, F. Diehl, B. Karim, G. Liu, N. Papadopoulos,Y. Samuels, V. E. Velculescu, B. Vogelstein, K. W. Kinzler, L. A. Diaz,Serial Assessment of Human Tumor Burdens in Mice by the Analysis ofCirculating DNA. Cancer Res .67, 9364-9370(2007). 25. A. Rhie, S. Nurk, M. Cechova, SJ Hoyt, DJ Taylor, N.Altemose, PW Hook, S. Koren, M. Rautiainen, IA Alexandrov, J. Allen, M.Asri, AV Bzikadze, N.-J. Chen, C.-S. Chin, M. Diekhans, P. Flicek, G.Formenti, A. Fungtammasan, CG Girón, E. Garrison, A. Gershman, J. Gerton,PGS Grady, A. Guarracino, L. Haggerty, R. Halabian, NF Hansen, R.Harris, GA Hartley, WT Harvey, M. Heinness, J. Heinness, T. Hourlier, RM Hubley, SE Hunt, M. Jain, RK Kesharwani, AP Lewis, H. Li, GALogsdon, JK Lucas, W. Makalowski, C. Markovic, FJ Martín, AMMcCartney, RC McCoy, J. McDaniel, BM McNulty, P. Medev, A. Murvedko, KMT Murphyson, HE Munphyson. Olsen, ND Olson, LFPaulin, D. Porubsky, T. Potapova, F. Ryabov, SL Salzberg, MEG Sauria,FJ Sedlazeck, K. Shafin, VA Shepelev, A. Shumate, JM Storer, L.Surapaneni, AMT O'Neill, F. Thibaud-Ni, Timpcz, M. Timpcz, Timpcz, M. Vollger, BP Walenz, AC Watwood, MHWeissensteiner , AMWenger , MA Wilson , Zarate S , Zhu Y , Zook JM , Eichler EE , O'Neill R , Schatz MC , Miga KH , Makova KD , Phillippy AM , Biorxiv , inpress , doi:10.1101 / 2022.12.01.518724 . [ PMC free article ] [ PubMed ] 26. Altemose N, Miga KH, Maggioni M, Willard HF. Plos Comput Biol .10, e1003628(2014). 27. Mathios D, Johansen JS, Christian S, Medina JE, Phallen J, Larsen KR, Bruhm DC, Niknafs N, Ferreira L, Adleff V, Adleff JYChiao, Leal A, Noe M, JR White, Arun AS, Hruban C, AVannapragada, SD Jensen, M-B. Wørntoft, AH Madsen, B Carvalho, M deWit, J Carey, NC Dracopoli, T Maddala, KC Fang, A.-R. [ PubMed ] Hartman, PMForde, V. Anagnostou, JR Brahmer, RJA Fijneman, HJ Nielsen, GAMeijer, CL Andersen, A. Mellemgaard, SE Bojesen, RB Scharpf, VEVelculescu. Nat Commun .12, 5060(2021). 28. S. Cristiano, A. Leal, J. Phallen, J. Fiksel, V. Adleff, DCBruhm, S. Ø. Jensen, JE Medina, J. Hruban, JR White, DN Palsgrove,N. Niknafs, V. Anagnostou, P. Forde, J. Naidoo, K. Marrone, J. Brahmer, BDWoodward, H. Husain, KL van Rooijen, M.-B. Wørntoft, AH Madsen, CJH van de Velde, M. Verheij, A. Cats, CJA Punt, GR Vink, NCT vanGrieken, M. Koopman, RJA Finjneman, JS Johansen, HJ Nielsen, GAMeijer, CL Andersen, RB Scharpf, VE Velculesculescug, genome-wide fragmentation with patient-wide DNA fragmentation cancer. Nature .570, 385-389(2019). 29. BA Methé, KE Nelson, M. Pop, ..., RK Wilson, O. White, Aframework for human microbiome research. Nature .486, 215-21(2011). 30. C. Huttenhower, D. Gevers, R. Knight, ..., RK Wilson, O.White, Structure, function and diversity of the healthy humanmicrobiome. Nature .486, 207-214(2012). 31. J. G. Tate, S. Bamford, H. C. Jubb, Z. Sondka, D. M. Beare, N.Bindal, H. Boutselakis, C. G. Cole, C. Creatore, E. Dawson, P. Fish, B.Harsha, C. Hathaway, S. C. Jupe, C. Y. Kok, K. Noble, L. Ponting, C. C.Ramshaw, C. E. Rye, H. E. Speedy, R. Stefancsik, S. L. Thompson, S. Wang, S.Ward, P. J. Campbell, S. A. Forbes, COSMIC: the Catalogue Of SomaticMutations In Cancer. Nucleic Acids Res .47, gky1015-(2018). 32. The ICGC / TCGA Pan-Cancer Analysis of Whole Genomes Consortium,Pan-cancer analysis of whole genomes. Nature .578, 82-93(2020). 33. A. Babaian, D. L. Mager, Endogenous retroviral promoterexaptation in human cancer. Mob. DNA .7, 24(2016). 34. H. S. Jang, N. M. Shah, A. Y. Du, Z. Z. Dailey, E. C. Pehrsson,P. M. Godoy, D. Zhang, D. Li, X. Xing, S. Kim, D. O'Donnell, J. I. Gordon, T.Wang, Transposable elements drive widespread expression of oncogenes in humancancers. Nat. Genet .51, 611-617(2019). 35. T. Sjoblom, S. Jones, L. D. Wood, D. W. Parsons, J. Lin, T. D.Barber, D. Mandelker, R. J. Leary, J. Ptak, N. Silliman, S. Szabo, P.Buckhaults, C. Farrell, P. Meeh, S. D. Markowitz, J. Willis, D. Dawson, J. K.V. Willson, A. F. Gazdar, J. Hartigan, L. Wu, C. Liu, G. Parmigiani, B. H.Park, K. E. Bachman, N. Papadopoulos, B. Vogelstein, K. W. Kinzler, V. E.Velculescu, The Consensus Coding Sequences of Human Breast and ColorectalCancers. Science .314, 268-274(2006). 36. LD Wood, DW Parsons, S Jones, J Lin, T Sjoblom, RJLeary, D Shen, SM Boca, T Barber, J Ptak, N Silliman, Szabo S, Dezso Z, Ustyanksky V, Nikolskaya T, Nikolsky Y, Karchin R, Wilson PA, Kaminker JS, Zhang, Croshaw R, Willis J, Dawson D, Shipitsin M, Willson JKV, Sukumar S, Polyak K, Park BH, Pethiyagoda CL, PVKPant, Ballinger DG, Sparks AB, Hartigan J, Smith DR, Suh E, Buckhaults, N Papadopoulos, Buckhaults, P, Markowitz, SD, Parmigiani, KW Kinzler,VE Velculescu, B. Vogelstein, The Genomic Landscapes of Human Breast andColorectal Cancers. Science .318, 1108–1113(2007). [ PMC free article ] [ PubMed ] 37. Burssed B, Zamariolli M, Bellucco FT, Melaragno MI,Mechanisms of structural chromosomal rearrangement formation. Mol. Cytogenet .15, 23(2022). 38. J. Qian, PP Massion, Role of Chromosome 3q Amplification inLung Cancer. J Thorac Oncol .3, 212–215(2008). 39. T. Hussenet, S. Dali, J. Exinger, B. Monga, B. Jost, D. Dembelé,N. Martinet, C. Thibault, J. Huelsken, E. Brambilla, S. du Manoir, SOX2 Is anOncogene Activated by Recurrent 3q26.3 Amplifications in Human Lung SquamousCell Carcinomas. Plos One .5, e8960(2010). 40. E. Rheinbay, M. M. Nielsen, F. Abascal, ..., M. J. van de Vijver,L. van’t Veer, C. von Mering, Analyses of non-coding somatic drivers in 2,658cancer whole genomes. Nature .578, 102-111(2020). 41. K. B. Chiappinelli, P. L. Strissel, A. Desrichard, H. Li, C.Henke, B. Akman, A. Hein, N. S. Rote, L. M. Cope, A. Snyder, V. Makarov, S.Budhu, D. J. Slamon, J. D. Wolchok, D. M. Pardoll, M. W. Beckmann, C. A.Zahnow, T. Merghoub, T. A. Chan, S. B. Baylin, R. Strick, Inhibiting DNAMethylation Causes an Interferon Response in Cancer via dsRNA IncludingEndogenous Retroviruses. Cell .162, 974-86(2014). 42. M. Russo, S. Morelli, G. Capranico, Expression of down-regulatedERV LTR elements associates with immune activation in human small-cell lungcancers. Mob. DNA .14, 2(2023). 43. M. Onishi-Seebacher, G. Erikson, Z. Sawitzki, D. Ryan, G. Greve,M. Libbert, T. Jenuwein, Repeat to gene expression ratios in leukemic blastcells can stratify risk prediction in acute myeloid leukemia. BMC Med. Genom .14, 166(2021). 44. J. R. Kemp, M. S. Longworth, Crossing the LINE Toward GenomicInstability: LINE-1 Retrotransposition in Cancer. Front. Chem .3, 68(2015). 45. Z. H. Foda, A. V. Annapragada, K. Boyapati, D. C. Bruhm, N. A.Vulpescu, J. E. Medina, D. Mathios, S. Cristiano, N. Niknafs, H. T. Luu, M.G. Goggins, R. A. Anders, J. Sun, S. H. Mehta, D. L. Thomas, G. D. Kirk, V.Adleff, J. Phallen, R. B. Scharpf, A. K. Kim, V. E. Velculescu, Detectingliver cancer using cell-free DNA fragmentomes. Cancer Discov (2022), doi:10.1158 / 2159-8290.cd-22-0659. 46. M. W. Snyder, M. Kircher, A. J. Hill, R. M. Daza, J. Shendure,Cell-free DNA Comprises an In Vivo Nucleosome Footprint that Informs ItsTissues-Of-Origin. Cell .164, 57-68(2016). 47. P. Ulz, G. G. Thallinger, M. Auer, R. Graf, K. Kashofer, S. W.Jahn, L. Abete, G. Pristauz, E. Petru, J. B. Geigl, E. Heitzer, M. R.Speicher, Inferring expressed genes by whole-genome sequencing of plasmaDNA. Nat Genet .48, 1273-1278(2016). 48. Q. Zhou, G. Kang, P. Jiang, R. Qiao, W. K. J. Lam, S. C. Y. Yu,M.-J. L. Ma, L. Ji, S. H. Cheng, W. Gai, W. Peng, H. Shang, R. W. Y. Chan, S.L. Chan, G. L. H. Wong, L. T. Hiraki, S. Volpi, V. W. S. Wong, J. Wong, R. W.K. Chiu, K. C. A. Chan, Y. M. D. Lo, Epigenetic analysis of cell-free DNA byfragmentomic profiling. Proc National Acad Sci .119, e2209852119(2022). 49. SY Shen, R. Singhania, G. Fehringer, A. Chakravarthy, MHARoehrl, D. Chadwick, PC Zuzarte, A. Borgida, TT Wang, T. Li, O. Kees, Z.Zhao, A. Spreafico, T. and S. Medina, Y. Wang, D. I. Rottafico, S. E. Rottabi, S.E. Chow, T. Murphy, A. Arruda, GM O'Kane, J. Liu, M. Mansour, JDMcPherson, C. O'Brien, N. Leighl, PL Bedard, N. Fleshner, G. Liu, MDMinden, S. Gallinger, A. Goldenberg, TJ Pugh, MM Hoffman, D. J. Hungman, D. J. Hungman, R. J. Hung. Sensitive tumor detection andclassification using plasma cell-free DNA methylomes. Nature .563, 579-583(2018). 50. J. Phallen, A. Leal, BD Woodward, PM Forde, J. Naidoo, KA Marrone, JR Brahmer, J. Fiksel, JE Medina, S. Cristiano, DNPalsgrove, CD Gocke, DC Bruhm, P. Keshavarzian, V. Adleff, E. Weihe, V. ARB, V.A. Velculescu, H. Husain, Early NoninvasiveDetection of Response to Targeted Therapy in Non-Small Cell LungCancer. Cancer Res .79, 1204-1213(2019). 51. N. Altemose, A classical revival: Human satellite DNAs enter the genomics era. Semin Cell Dev Biol .128, 2-14(2022). 52. JR Iben, RJ Maraia, tRNA gene copy number variation in humans. Gene .536, 376-384 (2014). 53. W.-W. Liao, M. Asri, J. Ebler, ..., H. Li, B. Paten, A drafthuman pangenome reference. Nature .617, 312-324(2023). 54. DC Bruhm, D. Mathios, ZH Foda, AV Annapragada, JEMedina, V. Adleff, EJ Chiao, L. Ferreira, S. Cristiano, JR White, SAMazzilli, E. Billatos, A. Spira, AH Zaidi, J. Mueller, AK Kim, V.Anagnostou, J. Phallen, RB Scharpf, VE Velculescu, Single-molecule genome-wide mutation profiles of cell-free DNA for non-invasive detection of cancer. Nat. Genet. , 1-10(2023). Example 2 This embodiment describes the method used in Embodiment 1 and additional results.

[0194] research group PCAWG available in our protected data cloud 2We obtained matched tumor and normal BAM files from 343 patients with lung cancer, liver cancer, ovarian cancer, colorectal cancer, and breast cancer, and excluded 10 patients whose tumor or normal samples were on the PCAWG blacklist (Table S12). We then used Samtools to convert the BAM files to FASTQ for use in the ARTEMIS workflow.

[0195] We analyzed whole-genome sequencing data (1-2x coverage) from cfDNA derived from 819 individuals with and without lung cancer, 208 individuals with and without liver cancer, and data from our previous publications. 3-6 The described 423 patient individuals from a multi-cancer cohort, who were either cancer-free or had tumors of the breast, ovary, lung, bile duct, colorectal, stomach, duodenum, or pancreas. Figure 23 In short, the LUCAS cohort is a prospective group of 365 patients examined at Bispebjerg Hospital in Copenhagen, Denmark, who showed positive imaging findings on chest X-ray or CT. 287 of these patients were used by Mathios et al. 3 The described early detection analysis. The liver cancer cohort consisted of 208 individuals with hepatocellular carcinoma or at high risk for liver cancer (cirrhosis, hepatitis B, or hepatitis C). 170 of these individuals were part of the Johns Hopkins University School of Medicine HCC biomarker registry and the IntraVenous Experience-Related AIDS Study (ALIVE). 6 This was a prospective collection. Thirty-eight patients with HBV or cirrhosis were retrospectively collected by BioIVT (Westbury, NY). External validation cohorts included colorectal cancer screening cohorts from Denmark (Endoscopy III) and the Netherlands (COCOS, Dutch trial registration number NTR1829). 7 385 non-cancer individuals from Alleghany Health Network, Boston University, and the Detection of Early Lung Cancer Among Military Personnel (DECAMP) consortium. 8 The study included lung cancer patients and individuals at high risk of lung cancer, as well as lung cancer patients from BioIVT (n=57) (Westbury, New York). The lung cancer surveillance cohort included 19 patients who received tyrosine kinase inhibitor therapy at the University of California, San Diego or Johns Hopkins University. As previously described, targeted sequencing and whole-genome sequencing data at 75 time points were available. 4,5The multi-cancer cohort consisted of samples from ILS Bio / Bioreclamation, Aarhus University, Herlev Hospital of the University of Copenhagen, Hvidovre Hospital, Utrecht University Medical Center, Amsterdam University Academic Medical Center, Netherlands Cancer Institute, and the University of California, San Diego. Patient sample collection used in this study complied with all relevant ethical guidelines. The collection protocol was approved by the Danish Regional Ethics Committee and the Danish Data Protection Authority (LUCAS cohort, Endoscopy III sample), the Dutch Health Council (COCOS sample), the U.S. Department of Defense Office for the Protection of Human Research (DECAMP sample), and the institutional review committees of the aforementioned institutions (AHN sample, liver cancer cohort, lung cancer surveillance cohort, and multi-cancer cohort). Written informed consent was provided by all patients, and the study was conducted in accordance with the Declaration of Helsinki.

[0196] The blood collection, cfDNA extraction, and sequencing protocols have been previously described. 3,4,6,9 In short, whole blood was collected into EDTA or Streck tubes, and plasma was separated by centrifugation and aliquoted into EDTA tubes. cfDNA was isolated from 2–4 mL of plasma, and next-generation sequencing libraries were prepared using 15 ng (if available) or the full purified amount (if less than 15 ng was available). Except for lung cancer surveillance samples and multi-cancer cohort samples, which underwent 12 cycles, all libraries were subjected to four cycles of PCR amplification. All cfDNA data used for modeling were sequenced on an Illumina HiSeq2500 platform with 1–2x coverage using 100 bp paired ends. For the LUCAS cohort, we analyzed cohorts sequenced on both HiSeq2500 and NovaSeq6000 platforms to allow for technical comparison of sequencing duplicates.

[0197] K-mers discovered from scratch We first extracted all repeat sequences and their coordinates of known repeat element types from the RepeatMasker orbitals of chm13 (T2T-CHM13v2.0). We excluded repeat sequences from low-complexity, unknown, and simple repeat sequence families, leaving 1287 repeat sequence types spanning 57 subfamilies (including 13 families). For simplicity, we grouped all elements from the tRNA, srpRNA, snRNA, scRNA, and rRNA families into RNA elements, and all elements from the DNA, DNA®, retrotransposon, and RC families into transposon elements, ultimately resulting in six families (LINE, SINE, LTR, satellite sequences, transposon elements, and RNA elements).

[0198] Then, we were influenced by Altemose et al.1 This inspired a de novo discovery process for k-mergers. We used Jellyfish. 10 We counted all unique 24-mers occurring in each of the 1287 repeat types, as well as 24-mers occurring in parts of the genome other than all repeat regions. We then selected k-mers that appeared only in a single repeat type and not in non-repeatable regions of the genome. K-mers of a sequence and its inverse complement were counted together because the reference genome represents one strand, but we expected half of the paired-end reads to originate from the inverse complement. We identified at least one unique k-mer in 1266 of the 1287 repeat types. We also additionally included k-mers from Altemose et al. 1 The supplemental material includes 58,426 k-mers from 14 HSATII and HSATIII subfamilies defined in the RepeatMasker satellite sequence annotations. These k-mers overlap with a wider range of satellite sequence types in the RepeatMasker orbits, but to maintain consistency with previous publications, we allowed these k-mers to be counted across multiple repetitive sequence types. In total, we identified 1,140,845,806 unique k-mers and defined 1,280 repetitive element types. To validate the low co-occurrence of these k-mers in common human-associated microbial genomes, we counted 1,545 microbial genomes from the Human Microbiome Project, downloadable from NCBI Entrez. 11,12 The number of k-mers in the sample.

[0199] Generation of k-merged repeating sequence landscapes We obtained all sequencing reads for each sample, counted each unique k-mer and its reverse complement, and summed the k-mer counts for each repeat type. We normalized the summed counts to the number of reads aligned with MAPQ>=30 (samtools view -c -q 30 -F 3844). Our method considers all reads, including those from genomic portions not provided in hg19 and / or unaligned repeat types.

[0200] ARTEMIS machine learning model in the organization We centered and scaled the coverage-normalized k-merged repeat sequence landscape counts for each tumor and normal tissue sample and trained a penalized logistic regression model to generate a cross-validation ARTEMIS score to distinguish tumor from normal tissue samples (the ARTEMIS score for each sample was calculated as the average of 5-fold cross-validations with 10 repetitions). We further trained a multi-class gradient boosting model (GBM) using the k-merged repeat sequence landscape to generate cross-validation (5-fold cross-validation) predictions for primary tumor tissues (the model generates a multinomial probability vector for each sample, where each element corresponds to a possible primary tumor tissue, and the predicted class is selected based on the element with the maximum value).

[0201] ARTEMIS is used for early detection, primary tissue, and monitoring of cancer in cfDNA. We obtained the k-cluster repeat sequence landscape for each sample using 786 features with an expected number of over 1000 k-clusters per million alignment reads. This filtering was employed because features with low abundance exhibit greater technical variability at low coverage. Figure 23 To allow for the integration of different feature categories, we used nested cross-validation to generate the ARTEMIS score. The inner cross-validation loop trained six penalized logistic regression (PLR) models (Lasso regression, α=1, with a penalty parameter selected in the range of 0.00001–0.1 by resampling within each cross-validation fold) featuring repetitive sequence landscapes (PLR models for each of the five repetitive sequence families, and a PLR model for the epigenetic profile, Supplementary Materials and Methods). The outer cross-validation loop was trained using a leave-one-individual-out architecture; we integrated six scores for each of N-1 individuals using the PLR ​​models. The score obtained by applying this PLR model to the six scores of the reserved patient features is the ARTEMIS score for cancer detection in cfDNA.

[0202] To incorporate the DELFI fragmentation profile, we integrated the ARTEMIS score with three other models: a PLR model using principal component analysis of the ratio of short to long fragments within a 5 Mb bin across the entire genome (D. Mathios et al., Nat Commun 12, 5060 (2021)), a PLR model based on aneuploidy z-scores of 39 chromosome arms (D. Mathios et al., Nat Commun 12, 5060 (2021) and a gradient boosting model based on coverage in 5 Mb bins across the whole genome ( 28The ensemble of this combination yielded a joint ARTEMIS-DELFI score. We retrained the lung cancer model on the full LUCAS cohort and then applied the locked model to four external validation cohorts: the JHU validation set (n=431) from Mathios et al. (2021) (D. Mathios et al., Nat Commun 12, 5060 (2021)), a subset of patients with prior cancer, with or without cancer recurrence (n=25) from the LUCAS cohort, and a validation set from the AHN / DECAMP cohort of Bruhm et al. (2023). 56 ), and the lung cancer surveillance cohort (n=19) from Phallen et al. (2019) 50 .

[0203] Finally, we trained a multi-class GBM using the features described above to generate ARTEMIS and ARTEMIS-DELFI scores for primary tissue classification. The final ARTEMIS model for primary tissues used the ensemble components described above as features, while the ARTEMIS-DELFI model used these components along with additional fragmentation features (aneuploidy z-scores for 39 chromosome arms and the ratio of short to long fragments within a 5 Mb bin across the whole genome). Each ensemble component produced a multinomial probability vector, each corresponding to a possible tumor site. We determined the classification based on the vector element with the maximum value. When ensembled multiple GBM classifiers, all elements of the vector were used as feature inputs to the ensemble model. Nested cross-validation was used on all cancer samples (n=423) from Cristiano et al. (2019) ( 28 We trained the models on [the platform] and reported the performance of the ARTEMIS and ARTEMIS-DELFI detection models in detecting all cancers at a 90% specificity threshold when trained on a full cohort including cancer-free patients. Consistent with previous publications, for primary tissue analysis, we included baseline time points from the aforementioned lung cancer surveillance cohort to increase the number of lung cancers available for classification analysis.

[0204] Gene set enrichment analysis We downloaded the coordinates of all genes from the chm13 version of the compiled ncbiRefSeq from the UCSC Genome Explorer. For genes with multiple transcripts, we defined regions containing all transcripts and then counted which of the 1.1 billion possible k-mers appeared in each gene. We then categorized the genes based on two metrics: (1) the total number of k-mers and (2) a corrected k-mer density metric—the number of types appearing in the gene divided by the k-mer density (total k-mers / Mb). We used this correction because our requirement that a k-mer appear only in a single repetitive element type meant that for genes with a greater number of repetitive sequence types occurring therein, the chance of finding k-mers was actually lower, despite these genes having a higher overall repetitive sequence density. We used the KEGG gene set and the COSMIC cancer gene census. 13-15 Both ordinal lists are used in gene set enrichment analyses. For KEGG, the default minimum gene set size of 25 and maximum gene set size of 500 are used. For COSMIC analysis, a minimum of 15 and a maximum of 1000 are used to ensure that all gene sets are included.

[0205] Simulating the k-merchant repeat landscape in short-read sequencing We simulated 50 million paired-end reads from the chm13 reference genome with an error rate of 0.1% (corresponding to Q30), segmented the reads according to the primary repeat element type (or reads appearing in non-repetitive regions), and counted 1.1 billion possible k-mers for each group. Then, for each repeat element type, we calculated the percentage of all counted k-mer occurrences in reads originating from that repeat element type.

[0206] Landscape analysis of k-mer repeat sequences in PCAWG tumors We generated a k-merged repeat sequence landscape for all 686 PCAWG samples (343 matched tumor / normal pairs) and excluded 10 pairs from the PCAWG blacklist, leaving 666 samples. To define significant changes in repeat elements, we selected 100 normal samples and subsampled two computer-simulated normal samples, each with approximately half the coverage of the original samples. We then calculated the ratio of k-merged repeat sequence landscapes between the two samples and set the threshold for determining significant changes in tumors to be either below the 1st percentile or above the 99th percentile of the normal / normal ratio. This ensured that changes observed in tumors were greater than those expected to be attributable solely to chance variants. We compared the number of such changes with several genomic instability indicators (entropy, nonmodal ploidy fraction, nondiploidy fraction, loss of heterozygosity fraction, breakpoint count, tumor mutational burden, ploidy, and modal ploidy). 16Related. The visualized sample set (n=293, Figure 2A) can be obtained from Anagnostou et al. 16 A tumor definition was established to obtain genomic stability metrics. In short, this required tumors to have TCGA copy number data (for ploidy correlation metrics) and to belong to the mc3 (multicenter, multicancer mutation detection) mutation detection set used for TMB calculation. Correlation analysis was performed based on data from Anagnostou et al. (2020). 16 The analysis was performed on this subset (293 out of the original 333), while the remaining analyses used all 333 non-blacklisted samples.

[0207] To determine whether these observed changes were greater than those expected attributable solely to chromosome arm gains and deletions, we used copy number data from the TCGA array to calculate the mean arm-level ploidy for 329 out of 333 tumors with available data. We counted all k-mers on each chromosome arm and determined the theoretical count of diploid k-mers. We then adjusted these counts at the chromosome arm level for the copy number profile of each individual sample. This adjustment could not be performed for acrocentromere chromosome arms or the Y chromosome, which were not available in the TCGA data, so we assumed they were diploid. We then calculated the predicted tumor:normal ratio (if all changes in the k-mer repetitive sequence landscape were attributable solely to individual chromosome arm gains and deletions) and defined confidence intervals around these for each element in each sample using the range of values ​​observed in the random downsampling experiments described above. We then further filtered for changes that exceeded the range predicted by chromosome arm gains and deletions.

[0208] To extend these copy number analyses, we identified repetitive k-mers found on each chromosome arm and calculated the tumor-to-normal k-mer count ratio at the arm level for each sample and for each chromosome arm from which copy number data were available from TCGA. We compared the expected ratio of k-mer counts with the ratios observed across the arm-level ploidy range. We repeated this analysis for amplicon regions with copy numbers greater than 4, up to 10 of the most amplified regions per tumor. We also randomly sampled 10 diploid regions from each tumor containing one of the aforementioned amplified regions as controls. After filtering amplicons ranging in size from 10 kb to 5 Mb, we obtained 1976 regions, of which 1972 contained repetitive sequence k-mers. We compared the expected ratio of k-mer counts with the ratios observed across the amplicon copy number range.

[0209] Next, we identified two common copy number gain patterns in tumors—those found in breast cancer tumors. ERBB2 The 1Mb region and lung squamous cell tumors contain SOX2 and PIK3CAThe 31 Mb region was identified. We extracted the types of repeat elements that overlapped with k-mers in these regions and assessed the difference in genome-wide k-mer counts of these repeat elements between tumors with and without focal amplification.

[0210] To assess the impact of small-kb-scale repetitive sequence-related structural changes on the k-merchant repetitive sequence landscape, we identified 19 LINE-1-mediated deletions occurring in lung squamous cell tumors from PCAWG. 17 Of the 19 LINE-1 mediated deletions ranging in size from <1 kb to >50 Mb, 11 contained at least one ARTEMIS k-mer. We further restricted these k-mer sets to those appearing only within the genomic regions defined by the deletions and compared the counts of these k-mers between tumor and normal samples with and without the deletions.

[0211] We divided patients into two groups whose number of repeat element variations in the six major repeat element families (LINE, SINE, LTR, satellite sequences, DNA elements, and RNA elements) was above or below the median, and used the R package... survival and survminer Overall survival and progression-free survival were compared between the two groups. Only the number of SINE changes was significantly associated with survival. To correct for the influence of tumor type in this analysis, we also performed an analysis stratified patients according to the median of tumor type-specific changes, effectively ensuring that half of the patients from each tumor type were in each group. We also compared Anagnostou et al. (2020). 16 Survival analysis was performed using the genome stability metrics defined in the [reference].

[0212] Finally, we centered and scaled the coverage-normalized counts and trained a penalized logistic regression model using these k-merged repeat sequence landscapes to generate a cross-validated ARTEMIS score (the average score of 5-fold cross-validation across 10 replicates) to distinguish tumor from normal tissue samples. We further trained a multi-class gradient boosting model (GBM) using the k-merged repeat sequence landscape to generate a cross-validated ARTEMIS score (5-fold cross-validation) to distinguish primary tumor tissue. We conservatively excluded breast tumors from this analysis because they were sequenced at different sites than other PCAWG samples, which could overstate performance due to technical differences. To assess the extent to which sequencing coverage affects the repeat sequence landscape, we resampled sequencing data from PCAWG lung cancer (with matched blood from normal samples, n=42) to 30x and 1–2x coverage and evaluated the consistency between the resampled normalized k-merged counts and the original normalized k-merged counts.

[0213] Observe new changes in repeating elements We quantified the variation in each repeating element using the analysis for single tumor-normal pairs as described above. Furthermore, we compared the k-mer counts for each feature between all tumor and normal samples using the Wilcoxon rank-sum test, employed Benjamini-Hochberg correction for multiple tests, and used Cohen's D for effect size. To assess the novelty of these variations for 57 subfamilies comprising 1280 repeating element types, we first performed a systematic search of the descriptions of each of the 57 subfamilies as follows: We first used a manually compiled pan-cancer paper on repetitive elements, using PCAWG and TCGA, to analyze paired tumor / normal studies. 17,18 We retrieved prior evidence from PCAWG / TCGA. We categorized subfamilies containing elements described in these studies as having "prior evidence in TCGA / PCAWG". For elements within these subfamilies, we considered them to have prior evidence for the change if they were explicitly identified as involving cancer-related changes; otherwise, we categorized the changes we observed as "novel". For the LINE-L1 subfamily, we conservatively considered all 132 elements to have prior evidence, as Rodriguez-Martin et al. (2020) identified nearly 20,000 LINE-L1-related changes in the PCAWG sample. 17 However, the element naming conventions prior to CHM13 were quite different, making differentiation difficult. This approach yielded descriptions of 18 subfamilies comprising 1008 elements. Of these, 275 elements were described as altering in cancer, and 733 were novel alterations that we identified.

[0214] For subfamilies not included in the aforementioned studies, we searched PubMed for "subfamilial name and cancer," or for subfamilies with few elements and few search results, we searched for "subfamilial name" or "single element name." For subfamilies where evidence was found using this method, we categorized the elements as having "previous evidence of the subfamilial." If a search for another subfamilial on PubMed yielded evidence of a subfamilial and its elements, we grouped them into the same category. This method yielded descriptions of 12 subfamilies comprising 185 elements. 19-31 .

[0215] If neither of the above two search attempts yields evidence of a subfamily or its internal elements, we classify it as "novel." This includes the final 27 subfamilies comprising 87 elements.

[0216] Analysis of the k-mer repeat sequence landscape in cfDNA We provide the LUCAS queue and its corresponding external validation set. 3 And for the AHN / DECAMP queue 9 Lung Cancer Surveillance Cohort 4 and liver cancer cohort 6 We generated a landscape of k-mer repetitive sequences. In the LUCAS cohort, we compared k-mer counts in lung and non-cancerous samples for repetitive element types that frequently changed in PCAWG lung tumors and satellite repetitive element types with different proportions of k-mers appearing on chrY.

[0217] We investigated the impact of epigenetic traits on the representativeness of circulating cfDNA fragments. In non-cancer patients in the LUCAS cohort, we examined differences in fragment length and coverage of fragments appearing in different histone marker regions. We defined histone chip-seq assays from the lymphoblastoid cell line GM12878 based on the ENCODE histone markers. 13 These regions (ENCFF001SUG, ENCFF001SUI, ENCFF001SUJ, ENCFF001SUE, ENCFF001SUL, ENCFF001SUF, ENCFF001SUN, ENCFF001SUO, ENCFF001SUP, ENCFF001SUQ). We further define chromatin states using ENCODE. 33 We categorized states 1-5 (promoter, enhancer 1, enhancer 2, 5' end of transcription 1, 5' end of transcription 2) as activated states, states 7-9 (3' end of transcription 1, 3' end of transcription 2, 3' end of transcription 3) as 3' end transcription, and states 10-13 (PC repression 1, PC repression 2, heterochromatin 1, heterochromatin 2) as repressed states. We then examined the localization of these histone markers within repeat element types in chm13 and identified repeat element types with higher and lower histone marker densities for each. Finally, we analyzed how the ratio of observed k-mer counts to expected k-mer counts (based on the chm13 reference genome) varied across different repeat element types based on histone marker densities.

[0218] Early detection, tissue origin, and monitoring of cancer using ARTEMIS We trained an ensemble penalized logistic regression model using k-cluster repeat sequence landscapes and defined the predictions as ARTEMIS scores. First, we extracted the k-cluster repeat sequence landscape for each sample using 786 features expected to exceed 1000 k-clusters per million alignment reads. This filtering was employed because low-abundance features exhibit greater technical variability at low coverage. Figure 22We divided the remaining features into six families (LINE, SINE, satellite sequences, LTR, RNA elements, DNA elements) and centered and scaled the features within each family in each sample. RNA and DNA elements were merged because only two RNA element features remained after filtering at a threshold of 1000 k-mers / million alignment reads. Using the epigenetic analysis described above, we further defined 561 1 Mb bins in which >90% of the bases were covered by a peak from one of the histone CHIP-Seq experiments described above, or >30% of the bases were covered by one of the three chromatin states defined above. We calculated the alignment fragment coverage in these bins, as in our previous fragmentation-based analysis. 3,6 Bins with a mappability <0.9 or GC content <0.3 were excluded from downstream analysis. We then trained six penalized logistic regression (PLR) models using the repetitive sequence landscape for the five repetitive families described above and the epigenetic profiles as described above. We then integrated the scores of these PLR ​​models using leave-one-out cross-validation and nested 5-fold cross-validation to train each learner. The combined scores were defined as ARTEMIS scores.

[0219] To incorporate DELFI fragmentation profiles, we integrated ARTEMIS scores with three other models: a PLR model using principal component analysis of the ratio of short to long fragments in 5 Mb bins across the whole genome. 3 PLR model based on aneuploidy z-score of 39 chromosome arms 3 And a gradient boosting model based on coverage in 5 Mb bins of the whole genome. 5 The integration of this combination yields a joint ARTEMIS-DELFI score. We compare the ARTEMIS score and the joint ARTEMIS-DELFI score with those from Mathios et al. (2021). 3 The scores of the described DELFI models were compared. We retrained the lung cancer model on the full LUCAS cohort and then applied the locked model to three external validation cohorts: from Mathios et al. (2021). 3 The JHU validation set (n=431) is from Bruhm et al. (2023). 9 The AHN / DECAMP validation set, and from Phallen et al. (2019). 4 A lung cancer surveillance cohort.

[0220] Finally, we trained a multi-class GBM using the ensemble architecture described above to generate ARTEMIS and ARTEMIS-DELFI scores for primary tissue classification. Cross-validation was then performed using data from Cristiano et al. (2019). 5 The model was trained on all cancer samples, and when trained on a full cohort including non-cancer patients, the performance of ARTEMIS and ARTEMIS-DELFI in detecting all cancers at a 90% specificity threshold was reported. Consistent with previous publications, for primary tissue analysis, we included baseline time points from the aforementioned lung cancer surveillance cohort to increase the number of lung cancers available for classification analysis.

[0221] Bioinformatics and Statistical Software All statistical analyses were performed using R version > 4.0.5. GSEA analysis was performed using Broad Institute software GSEAv4.3.2. 13,15 conduct 13,15 We used Jellyfish. 10 De novo k-mer discovery was performed in the chm13 reference genome, and these k-mers were then counted in patient samples. For alignment read counting and fragmentation analysis, adapter sequences were pruned using fastp (0.20.0), paired-end reads were aligned to the hg19 reference genome using Bowtie2 (2.3.0), and read counts were performed using samtools (1.17). Classification was performed using the R package Caret. R packages used for visualization included regioneR. 34 , fgsea (github.com / ctlab / fgsea / ), ComplexHeatmap 35 and ggplot2 36 .

[0222] Statistics and repeatability The computer code, software version, and computing environment used to reproduce the results in this study will be publicly available on GitHub after publication. As noted in the paper, the p-values ​​for comparisons between the two groups were determined using the Wilcoxon rank-sum test. As noted in the paper, the correlations of continuous variables were determined using the Pearson product-moment correlation coefficient or the Spearman rank correlation coefficient. The DeLong test was used to compare the ROC curves. All confidence intervals for the area under the ROC curve indicate a 95% confidence level and are based on the DeLong method.

[0223] References (used in Embodiment 2 above, as shown by superscript numbers in Embodiment 2, corresponding to the reference numbers in the list immediately below) https: / / doi.org / 10.1016 / S0022-5376(01)80010-0 , Google Scholar Crossref , CAS 1. Altemose N, Miga KH, Maggioni M, Willard HF, GenomicCharacterization of Large Heterochromatic Gaps in the Human Genome Assembly.Plos Comput Biol. 10, e1003628 (2014). 2. The ICGC / TCGA Pan-Cancer Analysis of Whole Genomes Consortium,Pan-cancer analysis of whole genomes. Nature. Rev. 578, 82–93 (2020). 3. Mathios D, Johansen JS, Christian S, Medina JE, Phallen J, Larsen KR, DC Bruhm, Niknafs N, Ferreira L, Adleff V, Chiao JY, Leal A, Noe M, White JR, Arun AS, Hruban C, AVAnnapragada S, Ø. Jensen, M.-BW Ørntoft, AH Madsen, B. Carvalho, M. deWit, J. Carey, NC Dracopoli, T. Maddala, KC Fang, A.-R. [ PubMed ] Hartman, PMForde, V. Anagnostou, JR Brahmer, RJA Fijneman, HJ Nielsen, GAMeijer, CL Andersen, A. Mellemgaard, SE Bojesen, RB Scharpf, VEVelculescu. Nat Commun. Rev. 12, 5060 (2021). 4. J. Phallen, A. Leal, BD Woodward, PM Forde, J. Naidoo, KAMarrone, JR Brahmer, J. Fiksel, JE Medina, S. Cristiano, DNPalsgrove, CD Gocke, DC Bruhm, P. Keshavarzian, V. Adleff, E. Weihe, V. ARB, Scharstop, V. ARB. Velculescu, H. Husain, Early NoninvasiveDetection of Response to Targeted Therapy in Non–Small Cell Lung Cancer.Cancer Res. 79, 1204–1213 (2019). 5. S. Cristiano, A. Leal, J. Phallen, J. Fiksel, V. Adleff, DCBruhm, S. Ø. Jensen, JE Medina, J. Hruban, JR White, DN Palsgrove,N. Niknafs, V. Anagnostou, P. Forde, J. Naidoo, K. Marrone, J. Brahmer, BDWoodward, H. Husain, KL van Rooijen, M.-BW Ørntoft, AH Madsen, CJH van de Velde, M. Verheij, A. Cats, CJA Punt, NCT Vink, M. Koopken, M. Kopken. RJA Fijneman, JS Johansen, HJ Nielsen, GAMeijer, CL Andersen, RB Scharpf, VE Velculescu, Genome-wide cell-free DNA fragmentation in patients with cancer. Nature. 570, 385–389 (2019). 6. ZH Foda, AV Annapragada, K. Boyapati, DC Bruhm, NAVulpescu, JE Medina, D. Mathios, S. Cristiano, N. Niknafs, HT Luu, MG Goggins, RA Anders, J. Sun, SH Mehta, DL Thomas, GD Kirk, V.Ad Phallen, JRB, Phallen, RB. Kim, VE Velculescu, Detectingliver cancer using cell-free DNA fragmentomes. Cancer Discov (2022), doi:10.1158 / 2159-8290.cd-22-0659. 7. EM Stoop, MC de Haan, TR de Wijkerslooth, PM Bossuyt,M. van Ballegooijen, CY Nio, MJ van de Vijver, K. Biermann, M. Thomeer,ME van Leerdam, P. Fockens, J. Stoker, EJ Kuipers, E. Dekker,Participation and yield of colonoscopy versus non-cathartic CT colonography in population-based screening for randomized controlled colorectal cancer. Lancet Oncol. 13, 55–64 (2012). 8. E. Billatos, F. Duan, E. Moses, H. Marques, I. Mahon, L. Dymond,C. Apgar, D. Aberle, G. Washko, A. Spira, D. investigators, Detection ofearly lung cancer among military personnel (DECAMP) consortium: studyprotocols. Bmc Pulm Med. 19, 59 (2019). 9. D. Bruhm, D. Mathios, Z. H. FOda, A. V. Annapragada, J. E. Medina,Vi. Adleff, E. J. Chiao, L. Ferreira, S. Cristiano, J. R. White, S. A.Mazzilli, E. Billatos, A. Spira, A. H. Zaidi, J. Mueller, A. K. Kim, V.Anagnostou, J. Phallen, R. B. Scharpf, V. E. Velculescu, Single moleculegenome-wide mutation profiles of cell-free DNA for non-invasive detection ofcancer. In Press. Nature Genetics. 10. G. Marçais, C. Kingsford, A fast, lock-free approach forefficient parallel counting of occurrences of k-mers. Bioinformatics. 27,764–770 (2011). 11. B. A. Methé, K. E. Nelson, M. Pop, …, R. K. Wilson, O. White, Aframework for human microbiome research. Nature. 486, 215–21 (2011). 12. C. Huttenhower, D. Gevers, R. Knight, …, R. K. Wilson, O. White,Structure, function and diversity of the healthy human microbiome. Nature.486, 207–214 (2012). 13. A. Subramanian, P. Tamayo, V. K. Mootha, S. Mukherjee, B. L.Ebert, M. A. Gillette, A. Paulovich, S. L. Pomeroy, T. R. Golub, E. S.Lander, J. P. Mesirov, Gene set enrichment analysis: A knowledge-basedapproach for interpreting genome-wide expression profiles. Proc National AcadSci. 102, 15545–15550 (2005). 14. J. G. Tate, S. Bamford, H. C. Jubb, Z. Sondka, D. M. Beare, N.Bindal, H. Boutselakis, C. G. Cole, C. Creatore, E. Dawson, P. Fish, B.Harsha, C. Hathaway, S. C. Jupe, C. Y. Kok, K. Noble, L. Ponting, C. C.Ramshaw, C. E. Rye, H. E. Speedy, R. Stefancsik, S. L. Thompson, S. Wang, S.Ward, P. J. Campbell, S. A. Forbes, COSMIC: the Catalogue Of SomaticMutations In Cancer. Nucleic Acids Res. 47, gky1015- (2018). 15. VK Mootha, CM Lindgren, K.-F. Eriksson, A. Subramanian, S.Sihag, J. Lehar, P. Puigserver, E. Carlsson, M. Ridderstråle, E. Laurila, N.Houstis, MJ Daly, N. Patterson, JP Mesirov, TR Golub, P. Tamayo, B.Spiegelman, ES Lander, JN Hirs, Alchhortshun, DLC Grotshun. PGC-1α-responsive genes involved in oxidative phosphorylation are coordinatelydownregulated in human diabetes. Nat Genet. 34, 267–273 (2003). 16. V. Anagnostou, N. Niknafs, K. Marrone, DC Bruhm, JR White,J. Naidoo, K. Hummelink, K. Monkhorst, F. Lalezari, M. Lanis, S. Rosner, JE Reuss, KN ​​Smith, V. Adleff, K. Rodgers, Z. Belcaid, L. Rhymee, B. Levy,J. Feliciano, CL Hann, DS Ettinger, C. Georgiades, F. Verde, P. Illei,QK Li, AS Baras, E. Gabrielson, MV Brock, R. Karchin, DM Pardoll,SB Baylin, JR Brahmer, RB Scharpf, PM Forde, VE Velculescu,Multimodal genomic prediction of immune checkpoint blockade feature non-small-cell lung cancer. Nat Cancer. 1, 99–111 (2020). 17. B. Rodriguez-Martin, EG Alvarez, A. Baez-Ortega, …, L. van'tVeer, C. von Mering, Pan-cancer analysis of whole genomes identifies driverrearrangements promoted by LINE-1 retrotransposition. Nat Genet. 52, 306–319(2020). 18. H. S. Jang, N. M. Shah, A. Y. Du, Z. Z. Dailey, E. C. Pehrsson,P. M. Godoy, D. Zhang, D. Li, X. Xing, S. Kim, D. O’Donnell, J. I. Gordon, T.Wang, Transposable elements drive widespread expression of oncogenes in humancancers. Nat. Genet. 51, 611–617 (2019). 19. C. Yandım, G. Karakülah, Dysregulated expression of repetitiveDNA in ER+ / HER2- breast cancer. Cancer Genet. 239, 36–45 (2019). 20. K. Zillner, J. Komatsu, K. Filarsky, R. Kalepu, A. Bensimon, A.Nmeth, Active human nucleolar organizer regions are interspersed withinactive rDNA repeats in normal and tumor cells. Epigenomics. 7, 363–378(2015). 21. M. Uemura, Q. Zheng, C. M. Koh, W. G. Nelson, S.Yegnasubramanian, A. M. D. Marzo, Overexpression of ribosomal RNA in prostatecancer is common but not linked to rDNA promoter hypomethylation. Oncogene.31, 1254–1263 (2012). 22. V. Valori, K. Tus, C. Laukaitis, D. T. Harris, L. LeBeau, K. A.Maggert, Human rDNA copy number is unstable in metastatic breast cancers.Epigenetics. 15, 85–106 (2020). 23. S. Shuai, H. Suzuki, A. Diaz-Navarro, F. Nadeu, S. A. Kumar, A.Gutierrez-Fernandez, J. Delgado, M. Pinyol, C. López-Otín, X. S. Puente, M.D. Taylor, E. Campo, L. D. Stein, The U1 spliceosomal RNA is recurrentlymutated in multiple cancers. Nature. 574, 712–716 (2019). 24. H. Suzuki, S. A. Kumar, S. Shuai, A. Diaz-Navarro, A. Gutierrez-Fernandez, P. D. Antonellis, F. M. G. Cavalli, K. Juraschka, H. Farooq, I. Shibahara, M. C. Vladoiu, J. Zhang, N. Abeysundara, D. Przelicki, P. Skowron, N. Gauer, B. Luu, C. Daniels, X. Wu, A. Forget, A. Momin, J. Wang, W. Dong, S.-K. Kim, W. A. ​​Grajkowska, A. Jouvet, M. Fèvre-Montange, M. L. Garrè, A. A. N. Rao, C. Giannini, J. M. Kros, P. J. French, N. Jabado, H.-K. Ng, WSPoon, C. G. Eberhart, I. F. Pollack, J. M. Olson, W. A. ​​Weiss, T. Kumabe, E. López-Aguilar, B. Lach, M. Massimino, E. G. V. Meir, J. B. Rubin, R. Vibhakar, L. B. Chambless, N. Kijima, A. Klekner, L. Bognár, J. A. Chan, CC Faria, J. Ragoussis, S. M. Pfister, A. Goldenberg, RJ Wechsler-Reya, SD Bailey, L. Garzia, AS Morrissy, MA Marra, X. Huang, D. Malkin, O. Ayrault, V. Ramaswamy, Nature.574, 707–711 (2019). 25. C. Tan, J. Cao, L. Chen, X. Xi, S. Wang, Y. Zhu, L. Yang, L. Ma,D. Wang, J. Yin, T. Zhang, Z. J. Lu, Noncoding RNAs Serve as Diagnosis andPrognosis Biomarkers for Hepatocellular Carcinoma. Clin. Chem. 65, 905–915(2019). 26. S. Ganesan, Breaking satellite silence: human satellite II RNAexpression in ovarian cancer. J. Clin. Investig. 132, e161981 (2022). 27. D. T. Ting, D. Lipson, S. Paul, B. W. Brannigan, S. Akhavanfard,E. J. Coffman, G. Contino, V. Deshpande, A. J. Iafrate, S. Letovsky, M. N.Rivera, N. Bardeesy, S. Maheswaran, D. A. Haber, Aberrant Overexpression ofSatellite Repeats in Pancreatic and Other Epithelial Cancers. Science. 331,593–596 (2011). 28. W. Arancio, C. Coronnello, Repetitive Sequence Transcription inBreast Cancer. Cells. 11, 2522 (2022). 29. A. K. Saha, M. Mourad, M. H. Kaplan, I. Chefetz, S. N. Malek, R.Buckanovich, D. M. Markovitz, R. Contreras-Galindo, The Genomic Landscape ofCentromeres in Cancers. Sci Rep-uk. 9, 11259 (2019). 30. K. Tsumagari, L. Qi, K. Jackson, C. Shao, M. Lacey, J. Sowden, R.Tawil, V. Vedanarayanan, M. Ehrlich, Epigenetics of a tandem DNA repeat:chromatin DNaseI sensitivity and opposite methylation changes in cancers.Nucleic Acids Res. 36, 2196–2207 (2008). 31. C. S. Herrington, M. Worsham, S. A. Southern, P. Mackowiak, S. R.Wolman, Loss of sequences on the short arm of chromosome 17 is a late eventin squamous carcinoma of the cervix. Mol. Pathol. 54, 160 (2001). 32. The ENCODE Project Consortium, An Integrated Encyclopedia of DNAElements in the Human Genome. Nature. 489, 57–74 (2012). 33. JWK Ho, YL Jung, T. Liu, BH Alver, S. Lee, K. Ikegami,K.-A. Sohn, A. Minoda, MY Tolstorukov, A. Appert, SCJ Parker, T. Gu,A. Kundaje, NC Riddle, E. Bishop, TA Egelhofer, SS Hu, AAAlekseyenko, A. Rechtsteiner, D. Asker, JA Belsky, SK Bowman, QBChen, RA-J. Chen, D.S. Day, Y. Dong, A.C. Dose, X. Duan, C.B. Epstein,S. Ercan, EA Feingold, F. Ferrari, JM Garrigues, N. Gehlenborg, PJGood, P. Haseley, D. He, M. Herrmann, MM Hoffman, TE Jeffers, PVKharchenko, P. Kolasinska-Zwierz, CV Kotwaliwale, N. Kumar, SA Langley, Langley, IEN Larch, IEN, Labbre, M. Labbre, M. Litorrecht. R. Park, MJ Pazin, HN Pham, A. Plachetka, B. Qin, YB Schwartz, N. Shoresh, P. Stempor, A.Vielle, C. Wang, CM Whittle, H. Xue, RE Kingston, JH Kim, BEBernstein, AF Dernburg, V. Pirrotta, TD Nobleda, M. W. No. Kellis, DM MacAlpine, S. Strome, SCR Elgin, XS Liu, JD Lieb, J. Ahringer, GH Karpen, PJPark, Comparative analysis ofmetazoan chromatin organization. Nature. 512, 449–452 (2014). 34. B. Gel, A. Díez-Villanueva, E. Serra, M. Buschbeck, MAPeinado, R. Malinverni, regioneR: an R / Bioconductor package for the association analysis of genomic regions based on permutation tests. Bioinformatics. 32, 289–291 (2016). 35. Z. Gu, R. Eils, M. Schlesner, Complex heatmaps reveal patterns and correlations in multidimensional genomic data. Bioinformatics. 32, 2847–2849 (2016). 36. H. Wickham, ggplot2, Elegant Graphics for Data Analysis (2009), doi:10.1007 / 978-0-387-98141-3 Other embodiments Based on the foregoing description, it is evident that the disclosures herein can be varied and modified to suit various uses and conditions. Such embodiments are also within the scope of the following claims.

[0224] All references to sequences, patents, and publications in this specification are incorporated herein by reference as if each individual patent and publication were specifically and individually indicated to be incorporated by reference. The various references cited by the applicant in this document do not constitute “prior art” to the disclosures of any particular reference.

Claims

1. A method for identifying the type of repeating elements, comprising: Extract repetitive sequences and genomic coordinates from known repetitive element types; select nucleic acid sequences (k-mers) appearing in a single repetitive type and identify unique k-mers of the repetitive element type; wherein the unique k-mers identify one or more repetitive element types, and identify the type and / or frequency of k-mers in the genome or cell-free nucleic acid sequences.

2. The method of claim 1, wherein the type and / or frequency of k-mers in the genome or cell-free DNA sequence are identified.

3. The method of claim 1 or 2, wherein the type and / or frequency of k-mers in the genome or cell-free RNA are identified.

4. The method of any one of claims 1 to 3, wherein the type and frequency of k-mers in the genome or cell-free nucleic acid sequence are identified.

5. The method of any one of claims 1 to 4, wherein repeating element types are excluded from families comprising low-complexity, unknown, simple repeating sequences or combinations thereof.

6. The method of any one of claims 1 to 5, wherein the aggregated elements are from each family, said family including tRNA, srpRNA, snRNA, scRNA, rRNA, RNA elements, DNA, retrotran, retrotransposon, or combinations thereof.

7. The method of any one of claims 1 to 6, wherein the family comprises long interspersed nuclear elements (LINE), short interspersed nuclear elements (SINE), long terminal repeat sequences (LTR), satellite sequences, transposable elements, RNA elements, or combinations thereof.

8. The method of any one of claims 1 to 7, wherein the k-mer comprises a maximum or less than 200 nucleotides.

9. The method of any one of claims 1 to 8, wherein the k-mer comprises a maximum or less than 100 nucleotides.

10. The method of any one of claims 1 to 8, wherein the k-mer comprises 10 to 40 nucleotides.

11. The method of any one of claims 1 to 8, wherein the k-polymer comprises 15 to 35 nucleotides.

12. The method of any one of claims 1 to 8, wherein the k-mer comprises 20 to 30 nucleotides.

13. The method of any one of claims 1 to 8, wherein the k-polymer comprises 22 to 28 nucleotides.

14. The method of any one of claims 1 to 8, wherein the k-mer comprises 23 to 26 nucleotides.

15. The method of any one of claims 1 to 5, wherein the k-mer comprises about 24 nucleotides.

16. The method of any one of claims 1 to 15, wherein a landscape of k-mer repetitive sequences is generated.

17. The method of claim 16, wherein each unique k-mer and its inverse complementary sequence are counted, and the k-mer counts for each repeating type element are summed.

18. The method of any one of claims 1 to 17, wherein gene regions comprising a plurality of transcripts are counted to identify k-mers present in each gene.

19. The method of claim 18, wherein the genes are sorted according to the total number of k-mers and the corrected k-mer density.

20. The method of claim 19, wherein the corrected k-mer density comprises the number of types appearing in the gene divided by the k-mer density.

21. The method of claim 20, wherein the k-polymer density is the total number of k-polymers per Mb.

22. The method of any one of claims 1 to 21, wherein the types of repetitive elements in cancer diagnosis-related genes are increased compared with non-cancer controls.

23. The method of any one of claims 1 to 22, wherein the family of repeating element types defined by the k-mer is altered across the entire genome compared to a normal control.

24. The method of claim 23, wherein the alterations in the genome compared to a non-cancer genome include focal amplification, deletion, copy number changes, rearrangements, changes in repetitive elements, or combinations thereof.

25. The method of any one of claims 1 to 24, wherein repetitive element sequences are identified from RNA transcripts.

26. A method for diagnosing and treating cancer in a subject, comprising: Generating k-mer repeat sequence landscapes from biological samples of subjects, generating k-mer repeat sequence landscapes from normal sample populations, comparing k-mer repeat sequence landscapes between subject samples and normal samples to diagnose subjects having cancer, and treating subjects diagnosed with cancer using cancer treatment methods.

27. The method of claim 26, wherein the changes in the k-mer repetitive sequence landscape are associated with one or more indicators of genomic instability.

28. The method of claim 27, wherein the one or more genomic instability indicators include entropy, nonmodal ploidy fraction, nondiploidy fraction, loss of heterozygosity fraction, breakpoint count, tumor mutational burden, ploidy, modal ploidy, or a combination thereof.

29. The method of any one of claims 26 to 28, further comprising repeating element types having overlapping k-mers, and analyzing the difference in the total k-mer count of said repeating element types between tumors with and without focal amplification.

30. The method of any one of claims 26 to 29, wherein the k-mer repeat sequence landscape is input into a machine learning or artificial intelligence program to generate a cross-validation score for distinguishing samples containing tumor material from normal samples.

31. The method of any one of claims 26 to 30, wherein the biological sample comprises genomic DNA, cell-free DNA (cfDNA), RNA, or a combination thereof.

32. The method of any one of claims 26 to 31, wherein the biological sample comprises cell-free DNA (cfDNA), and the cfDNA characteristics are evaluated, the cfDNA characteristics including fragment length and coverage of fragments in regions identified by histone marker frequencies.

33. The method of claim 32, wherein the location of histone markers within the repeat element type and the density of each histone marker are evaluated.

34. The method of claim 33, wherein the change in the ratio of observed k-mer count to expected k-mer count between repeat element types is assessed based on histone marker density.

35. The method of any one of claims 26 to 34, wherein the treatment is selected from surgery, adjuvant chemotherapy, neoadjuvant chemotherapy, radiotherapy, hormone therapy, cytotoxic therapy, immunotherapy, adoptive T-cell therapy, targeted therapy, and combinations thereof.

36. A method for detecting and monitoring cancer progression in a subject, comprising: Multiple features of the subject sample were detected to generate a k-mer repetitive sequence landscape; one or more features were classified into families, wherein the families include long interspersed nuclear elements (LINE), short interspersed nuclear elements (SINE), long terminal repeats (LTR), satellite sequences, transposable elements, RNA elements, or combinations thereof; multiple features were defined, including fragment length and fragment coverage in regions identified by histone marker frequencies; a regression model was trained using the repetitive sequence landscape and epigenetic profile; and these scores were integrated together using the regression model to generate a comprehensive score.

37. The method of claim 36, further comprising incorporating whole-genome fragmentation and scores obtained through one or more software models, machine learning models, artificial intelligence, or a combination thereof.

Citation Information

Patent Citations

  • Spray nozzle

    CA121113A

  • Tip for billiard cues

    CA233259A

  • Plug attachment

    CA271896A

  • Medicinal compound

    CA62924A