Method and reagent for detecting circular DNA molecules in biological samples
Error-corrected sequencing methods effectively detect and quantify eccDNA molecules by aligning DNA fragments and enriching them, addressing the challenge of eccDNA detection in biological samples and offering disease markers.
Patent Information
- Application Number
- JP2025512631
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-16
- Filing Date
- 2023-08-29
- Publication Date
- 2025-09-04
AI Technical Summary
Existing methods are inadequate for detecting and quantifying extrachromosomal circular DNA (eccDNA) molecules in biological samples, which are crucial markers for various human diseases, including cancer and chronic kidney disease, due to their non-linear nature and uneven distribution during cell division.
A method involving error-corrected sequencing, such as double-stranded sequencing, is employed to identify eccDNA by aligning DNA fragments with a reference genome, detecting candidate breakpoints, and enriching or preserving eccDNA molecules through selective removal or isolation of linear DNA, followed by adapter ligation to generate sequencing libraries.
This approach enhances the detection and quantification of eccDNA, providing a reliable marker for disease states by accurately identifying and characterizing these molecules, even in low abundance, and enabling disease treatment strategies.
Smart Images

Figure 2025529130000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 384,066, filed November 16, 2022, and U.S. Provisional Application No. 63 / 373,851, filed August 29, 2022, each of which is incorporated by reference in its entirety.
[0002] Reference to electronically submitted sequence listings This application contains a Sequence Listing that has been submitted electronically in XML file format, which is incorporated herein by reference in its entirety. The XML copy was created on August 29, 2023, is named "TSB-018SeqList.XML", and is 4,047 bytes in size.
[0003] The present technology generally relates to methods for preparing and analyzing nucleic acid libraries, such as DNA libraries for next-generation sequencing (NGS) applications, such as double-stranded sequencing. In particular, some embodiments of the present technology are directed to the detection and / or quantification of extrachromosomal circular DNA (eccDNA) molecules in biological samples. [Background technology]
[0004] Extrachromosomal circular DNA (eccDNA) molecules reside within the nuclei of eukaryotic cells and vary in size from less than 100 bp to several megabases. eccDNA molecules can contain any element present in the human genome, ranging from small non-coding regions to entire genes. During cell division, eccDNA molecules may be maintained, but because they lack centromeres, they are not evenly distributed during cell division.
[0005] eccDNA molecules affect human health. For example, eccDNA molecules that promote tumorigenesis are commonly present in cancer. Furthermore, elevated levels of eccDNA molecules are found in cell-free urine samples from individuals with chronic kidney disease. Therefore, detection of eccDNA could be a marker for human disease.
[0006] Thus, there is a need for new methods for detecting and quantifying non-linear molecules, such as eccDNA, in sequencing libraries obtained from biological samples, as well as methods for preparing libraries from biological samples in a manner that maximizes the preservation of such molecules. The present disclosure addresses this need and also provides other advantages. Summary of the Invention
[0007] In one aspect, the disclosure provides a method for detecting candidate extrachromosomal circular DNA (eccDNA) molecules in a biological sample, the method comprising: (a) providing a sequencing library comprising a plurality of double-stranded DNA fragments obtained from the sample; (b) obtaining error-corrected sequences for the double-stranded DNA fragments in the library; (c) detecting possible insertions present in the double-stranded DNA fragments by aligning a plurality of the error-corrected sequences with a reference genome; and (d) detecting candidate eccDNA breakpoints present in one or more of the fragments in which the possible insertions are detected.the candidate eccDNA breakpoint comprises sequence B located upstream of sequence A, wherein (i) in the reference genome, sequence A is present upstream of sequence B, (ii) in the reference genome, the first nucleotide of sequence A is located Y nucleotides upstream from the last nucleotide of sequence B, and (iii) in the candidate eccDNA breakpoint, the last nucleotide of sequence B is located approximately immediately upstream of the first nucleotide of sequence A; (e) detecting a candidate eccDNA molecule from among the fragments containing the candidate eccDNA breakpoint, wherein the candidate eccDNA molecule is a fragment containing a candidate eccDNA breakpoint that is not excluded by any of (i), (ii), or (iii), the step comprising: (i) comparing the length of the fragment with a distance of Y nucleotides, wherein if the fragment is determined to be longer than Y nucleotides, the fragment is not a candidate eccDNA molecule. (ii) determining whether the error-corrected sequence of the fragment contains a duplication of any subsequence contained within the region from sequence A to sequence B in the reference genome, wherein if a duplication is detected, it indicates that the fragment is not a candidate eccDNA molecule; and / or (iii) determining whether a sequence located upstream of sequence A in the reference genome is present upstream of sequence B in the error-corrected sequence of the fragment, or a sequence located downstream of sequence B in the reference genome is present downstream of sequence A in the error-corrected sequence of the fragment, wherein if it is detected that a sequence located upstream of sequence A in the reference genome is present upstream of sequence B in the error-corrected sequence, or a sequence located downstream of sequence B in the reference genome is present downstream of sequence A in the error-corrected sequence, it indicates that the fragment is not a candidate eccDNA molecule.
[0008] In some embodiments, the error-corrected sequence of the method is obtained in (b) by consensus sequencing.In some embodiments, the consensus sequencing is double-stranded sequencing (DS).In some embodiments, the consensus sequencing is single-stranded consensus sequencing (SSCS) or the combination of DS and SSCS. In some embodiments, the biological sample is selected from the group consisting of a sperm sample, a semen sample, a prostatic fluid sample, a testicular biopsy sample, a spermatogonial sample, a germ cell sample, a gamete sample, a swab sample, a lavage sample, an aspirate sample, a biopsy sample, a tissue sample, a tumor sample, a precancerous lesion sample, a liquid biopsy sample, a hyperplasia sample, a hypertrophy sample, a dysplasia sample, a urine sample, a cerebrospinal fluid (CSF) sample, other body fluid samples, an autopsy sample, an autopsy sample, a surgical sample, a model organism sample, a plasma sample, a serum sample, a gastric juice sample, a bone marrow sample, a stool sample, a brushing sample, a bile sample, a pancreatic juice sample, a synovial fluid sample, a sputum sample, a mucus sample, a vitreous humor sample, a forensic sample, an environmental sample, a bacterial sample, a fungal sample, a mammalian sample, a human sample, and a diagnostic sample.
[0009] In some embodiments, the biological sample contains potential cancer cells or nucleic acids potentially derived from cancer. In some embodiments, the biological sample contains cell-free DNA. In some embodiments, the biological sample contains cells exposed to a potentially toxic substance. In some embodiments, the potentially toxic substance is a potential clastogen, aneugen, mutagen, and / or teratogen, an aneugen, mutagen, and / or teratogen. In some embodiments, the presence and / or characteristics of eccDNA molecules in the sample are used to identify a disease state or physiological condition. In some embodiments, the disease state or physiological condition is selected from the group consisting of inflammation, autoimmunity, infection, organ transplant rejection, stem cell transplant rejection, therapeutic cell rejection, therapeutic cell response, immunotherapy response, pregnancy, preeclampsia, radiation exposure, sun exposure, drug exposure, and hypersensitivity. In some embodiments, the double-stranded DNA fragments are obtained by enzymatic fragmentation. In some embodiments, the average length of the double-stranded DNA fragments in the library is between about 100 bp and 1000 bp, or greater than about 1000 bp, 2000 bp, 3000 bp, 4000 bp, 5000 bp, 6000 bp, 7000 bp, 8000 bp, 9000 bp, or 10,000 bp.
[0010] In some embodiments, the length of the candidate eccDNA molecule is between about 100 and 1000 nucleotides. In some embodiments, the length of the candidate eccDNA molecule is less than about 500 nucleotides. In some embodiments, the length of the candidate eccDNA molecule is greater than about 1000 bp, 2000 bp, 3000 bp, 4000 bp, 5000 bp, 6000 bp, 7000 bp, 8000 bp, 9000 bp, or 10,000 bp. In some embodiments, the length of the candidate eccDNA molecule is greater than about 100 kb, 200 kb, 300 kb, 400 kb, 500 kb, 600 kb, 700 kb, 800 kb, 900 kb, 1 Mb, 2 Mb, or 3 Mb. In some embodiments, the candidate eccDNA molecule comprises a gene. In some embodiments, the candidate eccDNA molecule comprises an origin of replication. In some embodiments, the length of the candidate eccDNA molecule is approximately equal to the distance Y nucleotides. In some embodiments, the length of the candidate eccDNA molecule is exactly equal to the distance Y nucleotides. In some embodiments, the length of the candidate eccDNA molecule is less than about 50%, 60%, 70%, 80%, 90%, or more of the average length of the DNA fragments in the library. In some embodiments, the error-corrected sequences obtained in (b) are specific to a single genomic region. In some embodiments, the error-corrected sequences obtained in (b) are specific to about 1 to about 30 distinct genomic loci.
[0011] In some embodiments, the method is performed with or without an enrichment step to increase the proportion of double-stranded circular DNA molecules among all double-stranded nucleic acids in the sample, and further includes a step of comparing the frequencies in the library of possible insertions detected in step (c), candidate eccDNA breakpoints detected in step (d), and / or candidate eccDNA molecules detected in step (e) obtained by the method performed with or without the enrichment step. In some embodiments, the enrichment step comprises selectively removing double-stranded linear DNA molecules from the sample. In some embodiments, the double-stranded linear DNA molecules are selectively removed by treating the sample with one or more exonucleases. In some embodiments, the enrichment step comprises selectively isolating double-stranded circular DNA molecules from the sample. In some embodiments, the double-stranded circular DNA molecules are selectively isolated by electrophoresis, column filtration, density gradient centrifugation, selective extraction, and / or using a DNA-binding protein that selectively binds to or maintains binding to double-stranded circular DNA molecules relative to double-stranded linear DNA molecules. In some embodiments, the DNA binding protein is a helicase.
[0012] In some embodiments, the ratio of the number of candidate eccDNA molecules detected in step (e) to the number of error-corrected sequences obtained in step (b), the number of possible insertions detected in step (c), or the number of candidate eccDNA breakpoints detected in step (d) is higher when the method is performed with an enrichment step than when the method is performed without an enrichment step. In some embodiments, the frequency of possible insertions detected in step (c) with the enrichment step is less than about 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, or 10% of the frequency of possible insertions detected in step (c) without the enrichment step. In some embodiments, the frequency of candidate eccDNA breakpoints detected in step (d) with the enrichment step is less than about 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, or 10% of the frequency of candidate eccDNA breakpoints detected in step (d) without the enrichment step. In some embodiments, the frequency of candidate eccDNA molecules detected in step (e) with the enrichment step is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or more of the frequency of candidate eccDNA molecules detected in step (e) without the enrichment step. In some embodiments, enrichment of circular DNA molecules may not significantly decrease the frequency of different categories, e.g., when substantially all of the insertions present in a given sample correspond to eccDNA molecules. In some embodiments (e.g., when substantially all of the insertions present in a given sample correspond to eccDNA molecules), enrichment of circular DNA molecules may increase the frequency of different categories.
[0013] In some embodiments, the method further comprises calculating the probability that the candidate eccDNA molecule identified in step (e) is a bona fide eccDNA molecule. In some embodiments, the calculation is based in part on any one or more of the frequencies, ratios, or percentages disclosed herein. In some embodiments, the calculation is based in part on the relationship between the length of the candidate eccDNA molecule and the average length of double-stranded DNA fragments in the library, where a shorter length of the candidate eccDNA molecule relative to the average length of double-stranded DNA fragments in the library indicates a high probability that the candidate eccDNA molecule is a bona fide eccDNA molecule. In some embodiments, the calculation is based in part on the observation that the length of the candidate eccDNA molecule is nearly or exactly equal to the distance Y nucleotides, where the observation indicates a high probability that the candidate eccDNA molecule is a bona fide eccDNA molecule.
[0014] In some embodiments, any one or more of steps (a)-(e), determining any one or more of the frequencies, ratios, or percentages disclosed herein, or any of the calculations disclosed herein, are performed on a computer. In some embodiments, any one or more of steps (a)-(e), determining any one or more of the frequencies, ratios, or percentages disclosed herein, or any of the calculations disclosed herein, are performed on a cloud.
[0015] In another aspect, the present disclosure provides a computer-based system for performing any one or more of the methods disclosed herein.
[0016] In another aspect, the present disclosure provides a method of treating a disease or other medical condition in a mammalian subject, the method comprising the steps of (i) performing any one of the methods disclosed herein on a biological sample obtained from the subject, (ii) identifying in the sample one or more candidate eccDNA molecules indicative of a physiological condition associated with the disease or medical condition, and (iii) administering to the subject a treatment for the disease or medical condition.
[0017] In another aspect, the present disclosure provides a method for preparing a sequencing library for detecting candidate extrachromosomal circular DNA (eccDNA) molecules in a biological sample, the method comprising: (a) providing a biological sample containing double-stranded DNA; (b) preparing a first, non-enriched portion of the biological sample and a second, enriched portion enriched in double-stranded circular DNA molecules, wherein the second portion is prepared by selectively removing linear double-stranded DNA molecules and / or selectively isolating double-stranded circular DNA molecules from the sample; (c) fragmenting double-stranded DNA molecules in the first portion of the biological sample to generate a population of non-enriched double-stranded DNA fragments and enzymatically fragmenting double-stranded DNA molecules in the second portion of the biological sample to generate a population of enriched double-stranded DNA fragments; (d) ligating sequencing adaptors to a plurality of the non-enriched double-stranded DNA fragments to generate a non-enriched sequencing library; and (e) ligating sequencing adaptors to a plurality of the enriched double-stranded DNA fragments to generate an enriched sequencing library.
[0018] In some embodiments, the biological sample is selected from the group consisting of a sperm sample, a semen sample, a prostatic fluid sample, a testicular biopsy sample, a spermatogonial sample, a germ cell sample, a gamete sample, a swab sample, a lavage sample, an aspirate sample, a biopsy sample, a tissue sample, a tumor sample, a precancerous lesion sample, a liquid biopsy sample, a hyperplasia sample, a hypertrophy sample, a dysplasia sample, a urine sample, a cerebrospinal fluid (CSF) sample, other body fluid samples, an autopsy sample, an autopsy sample, a surgical sample, a model organism sample, a plasma sample, a serum sample, a gastric juice sample, a bone marrow sample, a stool sample, a brushing sample, a bile sample, a pancreatic juice sample, a synovial fluid sample, a sputum sample, a mucus sample, a vitreous humor sample, a forensic sample, an environmental sample, a bacterial sample, a fungal sample, a mammalian sample, a human sample, and a diagnostic sample.
[0019] In some embodiments, the biological sample contains nucleic acids from potential cancer cells or potential cancers. In some embodiments, the biological sample contains cell-free DNA. In some embodiments, the biological sample contains cells exposed to a potentially toxic substance. In some embodiments, the potentially toxic substance is a potential clastogen, aneugen, mutagen, and / or teratogen, an aneugen, mutagen, and / or teratogen. In some embodiments, the fragmentation is enzymatic fragmentation. In some embodiments, the second sample is enriched by treating the sample with one or more exonucleases. In some embodiments, the one or more exonucleases include exonuclease I, exonuclease T, exonuclease VII, exonuclease III, T7 exonuclease, exonuclease V (RecBCD), exonuclease VIII, lambda exonuclease, or T5 exonuclease.
[0020] In some embodiments, the method further comprises treating a portion of the second sample with one or more endonucleases before treating it with one or more exonucleases, and comparing the candidate eccDNA obtained with and without the one or more endonucleases. In some embodiments, the second sample is enriched by selectively isolating double-stranded circular DNA molecules from the sample. In some embodiments, the double-stranded circular DNA molecules are selectively isolated by electrophoresis, column filtration, density gradient centrifugation, selective extraction, and / or using a DNA-binding protein that selectively binds to or maintains binding to double-stranded circular DNA molecules compared to double-stranded linear DNA molecules. In some embodiments, the DNA-binding protein is a helicase. In some embodiments, the biological sample is treated with DTT before (b), (c), or step (d).
[0021] In some embodiments, the method further includes preparing third and fourth portions of the biological sample by removing subportions from the first non-enriched portion and the second enriched portion prepared in step (b), respectively; treating portions of the third and fourth portions with a reagent that induces cleavage in double-stranded circular DNA molecules at DNA damage sites, and leaving another portion of the third and fourth portions untreated; and ligating sequencing adapters to the treated and untreated portions of the third and fourth portions.
[0022] In some embodiments, the reagent is a combination of FPG (formamidopyrimidine [fapy]-DNA glycosylase) or UDG (uracil-DNA glycosylase) and endonuclease VIII. In some embodiments, the sequencing adapter is a double-stranded sequencing adapter. In some embodiments, the sequencing adapter comprises a Y-shape. In some embodiments, the sequencing adapter is a hairpin adapter.
[0023] In other aspects, the present disclosure provides a sequencing library prepared using any one of the methods disclosed herein.
[0024] In other aspects, the present disclosure provides kits for carrying out any one or more of the methods disclosed herein.
[0025] Many aspects of the present disclosure can be better understood with reference to the following drawings. The components in the drawings are not necessarily to scale, emphasis instead being placed upon clearly illustrating the principles of the present disclosure. [Brief explanation of the drawings]
[0026] [Figure 1A] FIG. 1 shows a schematic diagram of a system for identifying eccDNA molecules according to one embodiment. [Figure 1B] 1 shows a flowchart including steps that can be performed to identify candidate eccDNA molecules according to embodiments of the present disclosure. [Figure 1C] 1 shows a flowchart including steps for identifying candidate eccDNA molecules according to one embodiment. [Figure 2] Paired blood and sperm samples from six patients were analyzed by double-strand sequencing using the TwinStrand DuplexSeq mutagenicity assay. All sperm samples had fewer point mutations but more indels than their corresponding blood samples. [Figure 3A]Pooled blood and sperm samples (unpaired) were subjected to double-strand sequencing using the TwinStrand DuplexSeq mutagenicity assay and analyzed using either mechanical fragmentation (MF) or enzymatic fragmentation (EF) methods. Sperm samples had a lower frequency of point mutations and a higher frequency of indel calls compared to blood. The excess indels in sperm DNA, primarily insertions (not shown), were detected in libraries prepared using both fragmentation methods, but indel frequencies accounted for a higher proportion of total variant calls in sperm EF libraries (bottom). [Figure 3B] Pooled blood and unpaired sperm samples were subjected to double-strand sequencing using the TwinStrand DuplexSeq mutagenicity assay and analyzed using either mechanical fragmentation (MF) or enzymatic fragmentation (EF).The size distribution of indels in sperm DNA shows a periodicity similar to that reported for eccDNA, particularly in the microDNA subtype. [Figure 4] The indel calls are plotted by indel length for control DNA and sperm DNA libraries prepared by mechanical fragmentation (MF) or enzymatic fragmentation (EF). All apparent insertions greater than 20 bp in length were analyzed for the origin of the inserted DNA. If the called insert sequence matched the sequence immediately downstream of the called insertion position, forming a BA junction as shown in Figure 1B, the event was scored as an apparent tandem duplication (IsTandemDup = TRUE, orange). If the inserted DNA mapped to another site on the same chromosome or did not map to any site on the same chromosome, the event was scored as FALSE or NA, respectively. The left panel shows the number of indel calls by indel length for control DNA (also referred to as "devDNA" in the examples). The right panel shows the number of indel calls by indel length for sperm DNA samples. The excess insertions present in sperm cells were almost exclusively apparent tandem duplications containing BA junctions. All insertion calls greater than 100 bp in length were apparent tandem duplications. [Figure 5]The sequence reads and primary alignment for a specific indel call are shown. A shows an example of a variant call with the sequence color-coded to correspond to the subsequent panels. B shows two consensus "reads" supporting the variant call, with different portions of the sequence color-coded to correspond to the other panels. Asterisks indicate novel junctions, and italics indicate soft-clipped portions of the reads. C shows the read alignment displayed in IGV, with arrows added to correspond to the color coding in the other panels. D shows a schematic diagram of a circular DNA molecule derived from genomic DNA with a novel junction at the circularization site (shown as a gray, clear interface). The scissors mark indicates a single cleavage event that generates a fragment of the exact same length as the indel call. [Figure 6] A shows a schematic diagram of (i) the reference allele ABCD, (ii) the duplex consensus read pair, (iii) a split alignment of the R1 consensus read supporting a novel junction between D and A (DA junction), and two possible alternative alleles containing a DA junction. (iv) Excision and circularization of ABCD to form a DNA circle, and (v) a chromosomal tandem duplication (TD) of ABCD all generate a DA junction. B shows a histogram of allele lengths for all insertion calls identified in sperm DNA, color-coded by whether they contained a DA junction. The dashed vertical line indicates the median fragment size of the library (approximately the median insert size, 233 bp). C shows the subset of events with allele lengths less than the median fragment size and containing a DA junction (gray events to the left of the vertical dashed line in B, n = 63). (i) Chromosomal TD should show a fragment size distribution similar to that of the entire library (gray distribution), but DNA circular molecules should show fragment sizes smaller than the allele size, i.e., less than the median fragment size (diagonal hatching). (ii) All observed fragments were smaller than the median fragment size, which was significantly different from the distribution expected for chromosomal TD (binomial test p-value = 1.08 × 10-19). [Figure 7]Schematic diagram showing the relationship between allele length and fragment size in DA junction-containing molecules derived from circular DNA or chromosomal tandem duplications (TDs). For a reference allele ABCD (A) with an allele length of 400 bp, if an indel variant mapping to the reference is derived from circular DNA (B, left), a single fragmentation cleavage event (indicated by scissors) generates a linear fragment of length equal to the 400 bp allele length (C, left). Two or more cleavages of the circular DNA molecule generate fragments shorter than the allele length (not shown). In contrast, if an indel variant mapping to the reference is derived from the ABCD chromosomal TD (B, right), fragmentation of the indel variant can generate fragments shorter than the reference allele length (not shown), fragments equal to the allele length (C, left), or fragments longer than the 400 bp reference allele length (C, right). Note that fragments longer than the allele length (C, right) contain sequences in which at least a portion of the reference allele is repeated multiple times. [Figure 8] Exonuclease V treatment increased the relative frequency (number of hypothetical circular molecules per duplex base pair) of candidate DNA circles in HeLa cell and human sperm DNA as detected by DS. Error bars indicate Wilson's binomial confidence intervals. [Figure 9] A shows that the frequency of putative DNA circles per duplex base pair detected by DS is higher in tumor DNA compared to matched normal DNA. The numbers in the inset indicate the number of unique candidate DNA circles detected in each sample. B shows that the common somatic mutation frequency calculated from the same DS dataset shows no clear correlation with the DNA circle frequency, suggesting that these two measures are independent. In both panels, error bars indicate Wilson's binomial confidence interval. [Figure 10]Panel A shows the frequency of putative DNA circles per duplex base pair detected by DS in human TK6 cells treated with different genotoxic compounds. For each compound, the frequency of candidate circles increased in a dose-dependent manner, with at least one dose group being significantly higher than the untreated group. Panel B shows that mutation frequencies calculated from the same DS dataset did not show a clear correlation with DNA circle frequency, suggesting that these two measures are independent. In both panels, individual points represent measurements from replicate cultures, and error bars indicate group-based confidence intervals calculated using a t-distribution. Asterisks indicate p-values calculated from a quasi-Poisson generalized linear model for between-group comparisons: *p<0.05, **p<0.01, ***p<0.001. [Figure 11] Panel A shows the frequency of hypothetical DNA circular molecules per duplex base pair detected by DS in human TK6 cells cocultured with HepaRG cells and treated with water or cyclophosphamide, compared with the control group. The frequency of DNA circular molecules was significantly higher in the cyclophosphamide-treated samples than in the control group. Panel B shows that the mutation frequency in the cyclophosphamide-treated samples was significantly higher than in the control group. In both panels, individual points represent measurements from replicate cultures, and error bars indicate group-based confidence intervals calculated using a t-distribution. Asterisks indicate p-values calculated from a quasi-Poisson generalized linear model for between-group comparisons: *p<0.05, **p<0.01, ***p<0.001. DETAILED DESCRIPTION OF THE INVENTION
[0027] As used herein, unless otherwise clear from the context, the term "a" is understood to mean "at least one." As used herein, the term "or" is understood to mean "and / or." As used herein, the terms "comprising" and "including" are understood to encompass the listed element or step both when presented alone and when presented with one or more additional elements or steps. Ranges provided herein are intended to be inclusive. As used herein, the term "comprise" and variations thereof (e.g., "comprising," "comprises," etc.) are not intended to exclude other additives, elements, integers, or steps.
[0028] The terms "about" or "approximately," as used herein in reference to a value, refer to a value that is contextually similar to a reference value. Generally, one of ordinary skill in the art will understand the relevant degree of variation encompassed by "about" or "approximately," given the context. For example, in some embodiments, the terms "about" or "approximately" can encompass values within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less of the reference value.
[0029] Unless the word "or" in a list of two or more items is expressly limited to mean only one item exclusive of the other items, the use of "or" in that list should be interpreted as including (a) any single item in the list, (b) all of the items in the list, or (c) any combination of the items in the list. Furthermore, the term "comprising" is used throughout to mean including at least the recited features, and does not exclude greater numbers of the same features and / or additional types of other features. It will also be understood that, while specific embodiments have been described herein for illustrative purposes, various modifications can be made without departing from the art. Furthermore, while advantages associated with particular embodiments of the present disclosure have been described in the context of those embodiments, other embodiments may also exhibit those advantages, and not all embodiments necessarily exhibit those advantages, to be within the scope of the present disclosure. Accordingly, the present disclosure and related art may encompass other embodiments not expressly shown or described herein.
[0030] The term "subject" includes a cell, tissue, or organism, human or non-human, in vivo, ex vivo, or in vitro, male or female.
[0031] The term "mammal" includes humans and non-humans, including, but not limited to, humans, non-human primates, canines, felines, murines, bovines, equines, and porcines.
[0032] The term "sample" can include a single cell or multiple cells, cell fragments, or a bodily fluid fraction (e.g., a blood sample) obtained from a subject by venipuncture, excretion, ejaculation, massage, biopsy, needle aspiration, lavage, scraping, surgical incision or intervention, or other means known to those of skill in the art. Examples of bodily fluid fractions include amniotic fluid, aqueous humor, bile, lymph, breast milk, interstitial fluid, blood, plasma, earwax, Cowper's fluid (pre-ejaculatory fluid), chyle, chyme, female ejaculate, menstrual blood, mucus, saliva, urine, vomit, tears, vaginal fluid, sweat, serum, semen, sebum, pus, pleural effusion, cerebrospinal fluid, synovial fluid, intracellular fluid, and vitreous humor.
[0033] As used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. The following definitions and descriptions of various subordinate features may be used in combination with any of the methods described. When described in connection with a particular method, this is an example to aid in the explanation and does not limit the feature to that particular method.
[0034] I. Overview Disclosed herein is a method for identifying extrachromosomal circular DNA (eccDNA) from a sample. See FIG. 1A for an overview of an embodiment of a system for identifying eccDNA molecules. FIG. 1A shows a sample 110, a sequencing assay 120, and an extrachromosomal circular DNA (eccDNA) detection system 130.
[0035] In various embodiments, the sample 110 is obtained from a subject. In various embodiments, the sample may be obtained by an individual or a third party, such as a medical professional. Examples of medical professionals include doctors, paramedics, nurses, first responders, psychologists, phlebotomists, medical physicists, nurse practitioners, surgeons, dentists, and other obvious medical professionals recognized by those of skill in the art. In various embodiments, the sample 110 comprises a cell or cell population. For example, the cell or cell population may have previously been exposed to a compound or foreign substance. In such embodiments, the system diagram shown in FIG. 1A is useful for evaluating a compound or foreign substance provided to the cell or cell population.
[0036] Generally, the sample 110 is analyzed by performing a sequencing assay 120 to determine a plurality of sequence reads of nucleic acids present in the sample 110. In certain embodiments, the sequencing assay 120 includes performing an error-corrected sequencing method, one example of which is double-stranded sequencing (DS), which is described in further detail herein.
[0037] The eccDNA detection system 130 analyzes multiple sequence reads generated by the sequencing assay 120 to identify the presence or absence of eccDNA. In various embodiments, the eccDNA detection system 130 identifies candidate eccDNA sequence reads that contain reference allele junctions. The eccDNA detection system 130 can distinguish between eccDNA sequence reads and non-eccDNA sequence reads (e.g., sequence reads derived from chromosomal tandem duplications) and determine an eccDNA profile 140. In various embodiments, the eccDNA profile 140 refers to at least the amount of eccDNA molecules. In various embodiments, the eccDNA profile 140 refers to at least the frequency of eccDNA molecules.
[0038] In various embodiments, the eccDNA profile 140 is useful for various purposes. For example, the eccDNA profile 140 may be useful for assessing the clastogenicity of a potential clastogen previously provided to cells in the sample 110. As another example, the eccDNA profile 140 may be useful for assessing the genotoxicity of a xenobiotic previously provided to cells in the sample 110. In various embodiments, the eccDNA profile 140 is useful for assessing cancer risk in the sample 110. For example, a sample having a higher amount or frequency of eccDNA may be assessed as having a higher cancer risk compared to another sample having a lower amount or frequency of eccDNA.
[0039] II. Detection of eccDNA in biological samples The present disclosure relates to methods and related reagents, kits, and systems for detecting, quantifying, and characterizing circular DNA molecules in biological samples using error-corrected sequencing methods, such as double-stranded sequencing (DS). Such error-corrected sequencing offers the unexpected advantage of being able to detect low-volume sequences in a sample derived from extrachromosomal circular DNA (eccDNA). The present disclosure provides methods for preparing sequencing libraries from any biological sample that may contain eccDNA, thereby preserving and / or enriching the eccDNA present in the sample. The present disclosure also provides methods for sequencing and analyzing libraries prepared from such biological samples, thereby enabling the detection, characterization, and / or quantification of eccDNA. Furthermore, methods are provided for identifying additional conditions and features for sequencing library preparation and / or analysis that can improve the identification of eccDNA in a sample. The present methods and reagents can be used in any application, including the preparation of DNA libraries by fragmenting double-stranded DNA molecules and ligating adapters to the resulting double-stranded DNA fragments. The method can be used to generate libraries prepared from double-stranded DNA from cell samples, tissue samples, blood samples, biopsy samples, liquid biopsies, cell-free samples, forensic samples, environmental samples, or other sources that may contain eccDNA.
[0040] Extrachromosomal circular DNA (eccDNA) molecules are extrachromosomal circular DNA derived from endogenous chromosomes present in the nucleus, cytoplasm, or outside the cell, and their size can vary from less than 100 base pairs (bp) to several megabases. eccDNA includes, or may be referred to by, terms such as ecDNA, covalently closed circular DNA (generally referring to circular viral DNA molecules), microDNA, telomeric circles (a group of eccDNA involved in the immortalization of telomerase-negative cancers through selective lengthening of telomeres), and episomes (autonomously replicating circular DNA generally referring to bacterial DNA). As used herein, eccDNA generally refers to "simple" eccDNA molecules, i.e., eccDNA formed by circularization of a single contiguous segment of a genome, as opposed to hybrid or chimeric eccDNA, which may contain segments from other sites within the same genome or from other sources. Without wishing to be bound by theory, the term "microDNA" can refer to eccDNA molecules of less than 1,000, 2,000, 3,000, 4,000, 5,000, or more base pairs (bp). Due to their small size, microDNAs are less likely to carry full-length genes than larger eccDNA molecules, but may carry partial genes or microRNAs.
[0041] Small eccDNAs (commonly referred to as microDNAs) have been identified in many cell and tissue types in diverse eukaryotes, including plants, birds, rodents, and humans. The origin of these small, circular DNA molecules is unknown, but most proposed mechanisms involve the aberrant repair of DNA breaks. The potential functions of microDNAs are similarly poorly understood. Despite these unknowns, there is growing interest in using microDNAs as biomarkers for various disease states.
[0042] As used herein, the terms "eccDNA," "DNA circle," "circle," "circular DNA," and "microDNA" can be used interchangeably and refer to circular DNA molecules.
[0043] The method provided herein can be used to detect candidate microDNAs without exonuclease enrichment, allowing for simultaneous detection of chromosomal mutations and hypothesized microDNAs. Because double-stranded sequencing (DS) labels original DNA molecules with unique molecular identifiers (UMIs) before amplification, it allows for more quantitative assessment of microDNAs relative to each other and chromosomal DNA compared to current methods involving enrichment or rolling circle amplification. It has been shown that the abundance of microDNAs increases when cells are exposed to various clastogenic agents and apoptosis-inducing compounds. Therefore, microDNA detection by DS-based assays may provide a valuable indicator of genomic instability.
[0044] In certain embodiments, the method is directed to sequencing-based detection, quantification, and / or characterization of candidate eccDNA in a biological sample. In certain embodiments of this process, double-stranded DNA in the biological sample is fragmented, and sequencing adapters are ligated to one or both ends of the resulting double-stranded DNA fragments. The DNA fragments are then sequenced and analyzed to obtain a consensus sequence for multiple fragments. Upon alignment to a reference genome, certain fragments are identified as apparent indels (i.e., corresponding to small insertions or deletions), subcategories of apparent indels corresponding to apparent insertions, or apparent structural variants. In some embodiments, the apparent indels are 1,000 bp or less in length. In some embodiments, the apparent indels are 900 bp or less in length. In some embodiments, the apparent indels are 800 bp or less in length. In some embodiments, the apparent indels are 700 bp or less in length. In some embodiments, the apparent indels are 600 bp or less in length. In some embodiments, the apparent indel is 500 bp or less in length. In some embodiments, the apparent indel is 400 bp or less in length. In some embodiments, the apparent indel is 300 bp or less in length.In some embodiments, the apparent indel length is 20bp, 21bp, 22bp, 23bp, 24bp, 25bp, 26bp, 27bp, 28bp, 29bp, 30bp, 31bp, 32bp, 33bp, 34bp, 35bp, 36bp, 37bp, 38bp, 39bp, 40bp, 41bp, 42bp, 43bp, 44bp, 45bp, 46bp, 47bp, 48bp, 49bp, 50bp, 51bp, 52bp, 53bp, 54bp, 55bp, 56bp, 57bp, 58bp, 59bp, 60bp, 61bp, 62bp, 63bp, 64bp, 65bp, 66bp, 67bp, 68bp, 69bp , 70bp, 71bp, 72bp, 73bp, 74bp, 75bp, 76bp, 77bp, 78bp, 79bp, 80bp, 81bp, 82bp, 83bp, 84bp, 85bp, 86bp, 87bp, 88bp, 89bp, 90bp, 91bp, 92bp, 93bp, 94bp, 95bp, 96bp, 97bp, 98bp, 99bp, 100bp, 150bp, 200bp, 250bp, 300bp, 350bp, 400bp, 450bp, 500bp, 550bp, 600bp, 650bp, 700bp, 750bp, 800bp, 850bp, 900bp, 950bp, 1000bp, or more.In some embodiments, the apparent indel length is about 20 bp, about 21 bp, about 22 bp, about 23 bp, about 24 bp, about 25 bp, about 26 bp, about 27 bp, about 28 bp, about 29 bp, about 30 bp, about 31 bp, about 32 bp, about 33 bp, about 34 bp, about 35 bp, about 36 bp, about 37 bp, about 38 bp, about 39 bp, about 40 bp, about 41 bp, about 42 bp p, about 43bp, about 44bp, about 45bp, about 46bp, about 47bp, about 48bp, about 49bp, about 50bp, about 51bp, about 52bp, about 53bp, about 54bp, about 55bp, about 56bp, about 57bp, about 58bp, about 59bp, about 60bp, about 61bp, about 62bp, about 63bp, about 64bp, about 65bp, about 66bp, about 67bp, about 68bp, about 69bp , about 70bp, about 71bp, about 72bp, about 73bp, about 74bp, about 75bp, about 76bp, about 77bp, about 78bp, about 79bp, about 80bp, about 81bp, about 82bp, about 8 3bp, about 84bp, about 85bp, about 86bp, about 87bp, about 88bp, about 89bp, about 90bp, about 91bp, about 92bp, about 93bp, about 94bp, about 95bp, about 96bp, about 97 bp, about 98 bp, about 99 bp, about 100 bp, about 150 bp, about 200 bp, about 250 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, about 500 bp, about 550 bp, about 600 bp, about 650 bp, about 700 bp, about 750 bp, about 800 bp, about 850 bp, about 900 bp, about 950 bp, about 1000 bp, or more.In some embodiments, the apparent indel length is 20 bp or more, 21 bp or more, 22 bp or more, 23 bp or more, 24 bp or more, 25 bp or more, 26 bp or more, 27 bp or more, 28 bp or more, 29 bp or more, 30 bp or more, 31 bp or more, 32 bp or more, 33 bp or more, 34 bp or more, 35 bp or more, 36 bp or more, 37 bp or more, 38 bp or more, 39 bp or more, 40 bp or more, 41 bp or more, 42 bp or more , 43bp or more, 44bp or more, 45bp or more, 46bp or more, 47bp or more, 48bp or more, 49bp or more, 50bp or more, 51bp or more, 52bp or more, 53bp or more, 54bp or more, 55bp or more, 56 bp or more, 57bp or more, 58bp or more, 59bp or more, 60bp or more, 61bp or more, 62bp or more, 63bp or more, 64bp or more, 65bp or more, 66bp or more, 67bp or more, 68bp or more, 69bp or more Above, 70bp or more, 71bp or more, 72bp or more, 73bp or more, 74bp or more, 75bp or more, 76bp or more, 77bp or more, 78bp or more, 79bp or more, 80bp or more, 81bp or more, 82bp or more, 83bp or more, 84bp or more, 85bp or more, 86bp or more, 87bp or more, 88bp or more, 89bp or more, 90bp or more, 91bp or more, 92bp or more, 93bp or more, 94bp or more, 95bp or more, 96bp or more, 97bp or more, 98bp or more, 99bp or more, 100bp or more, 150bp or more, 200bp or more, 250bp or more, 300bp or more, 350bp or more, 400bp or more, 450bp or more, 500bp or more, 550bp or more, 600bp or more, 650bp or more, 700bp or more, 750bp or more, 800bp or more, 850bp or more, 900bp or more, 950bp or more, 1000bp or more, or more.In some embodiments, the apparent indel length is 20bp, 21bp, 22bp, 23bp, 24bp, 25bp, 26bp, 27bp, 28bp, 29bp, 30bp, 31bp, 32bp, 33bp, 34bp, 35bp, 36bp, 37bp, 38bp, 39bp, 40bp, 41bp, 42bp, 43bp, 44bp, 45bp, 46bp, 47bp, 48bp, 49bp, 50bp, 51bp, 52bp, 53bp, 54bp, 55bp, 56bp, 57bp, 58bp, 59bp, 60bp, 61bp, 62bp, 63bp, 64bp, 65bp, 66bp, 67bp, 68bp, 69bp , 70bp, 71bp, 72bp, 73bp, 74bp, 75bp, 76bp, 77bp, 78bp, 79bp, 80bp, 81bp, 82bp, 83bp, 84bp, 85bp, 86bp, 87bp, 88bp, 89bp, 90bp, 91bp, 92bp, 93bp, 94bp, 95bp, 96bp, 97bp, 98bp, 99bp, 100bp, 150bp, 200bp, 250bp, 300bp, 350bp, 400bp, 450bp, 500bp, 550bp, 600bp, 650bp, 700bp, 750bp, 800bp, 850bp, 900bp, 950bp, 1000bp, or more.
[0045] In some embodiments, the apparent indel length is greater than 1 bp, greater than 2 bp, greater than 3 bp, greater than 4 bp, greater than 5 bp, greater than 6 bp, greater than 7 bp, greater than 8 bp, greater than 9 bp, greater than 10 bp, greater than 11 bp, greater than 12 bp, greater than 13 bp, greater than 14 bp, greater than 15 bp, greater than 16 bp, greater than 17 bp, greater than 18 bp, greater than 19 bp, or greater than 20 bp. In some embodiments, the apparent indel length is greater than 20 bp.
[0046] In some embodiments, the apparent insertion length is 20 bp or more, 21 bp or more, 22 bp or more, 23 bp or more, 24 bp or more, 25 bp or more, 26 bp or more, 27 bp or more, 28 bp or more, 29 bp or more, 30 bp or more, 31 bp or more, 32 bp or more, 33 bp or more, 34 bp or more, 35 bp or more, 36 bp or more, 37 bp or more, 38 bp or more, 39 bp or more, 40 bp or more, 41 bp or more, 42 bp or more, 43 bp or more, 44 bp or more, 45 bp or more, 46 bp or more, 47 bp or more, 48 bp or more, 49 bp or more, 50 bp or more, 51 bp or more, 52 bp or more, 53 bp or more, 54 bp or more, 55 bp or more, 56 bp or more, 57 bp or more, 58 bp or more, 59 bp or more, 60 bp or more, 61 bp or more, 62 bp or more, 63 bp or more, 64 bp or more, 65 bp or more, 66 bp or more, 67 bp or more, 68 bp or more, 69 bp or more, 70 bp or more, 71 bp or more, 72 bp or more, 73 bp or more, 74 bp or more, 75 bp or more, 76 bp or more, 77 bp or more, 78 bp or more, 79 bp or more, 80 bp or more, 81 bp or more, 82 bp or 3bp or more, 44bp or more, 45bp or more, 46bp or more, 47bp or more, 48bp or more, 49bp or more, 50bp or more, 51bp or more, 52bp or more, 53bp or more, 54bp or more, 55bp or more, 56bp or more, 57bp or more, 58bp or more, 59bp or more, 60bp or more, 61bp or more, 62bp or more, 63bp or more, 64bp or more, 65bp or more, 66bp or more, 67bp or more, 68bp or more, 69bp or more , 70bp or more, 71bp or more, 72bp or more, 73bp or more, 74bp or more, 75bp or more, 76bp or more, 77bp or more, 78bp or more, 79bp or more, 80bp or more, 81bp or more, 82bp or more, 83bp or more, 84bp or more, 85bp or more, 86bp or more, 87bp or more, 88bp or more, 89bp or more, 90bp or more, 91bp or more, 92bp or more, 93bp or more, 94bp or more, 95bp or more, 96bp or more, 97bp or more, 98bp or more, 99bp or more, 100bp or more, 150bp or more, 200bp or more, 250bp or more, 300bp or more, 350bp or more, 400bp or more, 450bp or more, 500bp or more, 550bp or more, 600bp or more, 650bp or more, 700bp or more, 750bp or more, 800bp or more, 850bp or more, 900bp or more, 950bp or more, 1000bp or more, or more.In some embodiments, the apparent insertion length is about 20 bp, about 21 bp, about 22 bp, about 23 bp, about 24 bp, about 25 bp, about 26 bp, about 27 bp, about 28 bp, about 29 bp, about 30 bp, about 31 bp, about 32 bp, about 33 bp, about 34 bp, about 35 bp, about 36 bp, about 37 bp, about 38 bp, about 39 bp, about 40 bp, about 41 bp, about 42 bp , about 43bp, about 44bp, about 45bp, about 46bp, about 47bp, about 48bp, about 49bp, about 50bp, about 51bp, about 52bp, about 53bp, about 54bp, about 55bp, about 56bp, about 57bp, about 58bp, about 59bp, about 60bp, about 61bp, about 62bp, about 63bp, about 64bp, about 65bp, about 66bp, about 67bp, about 68bp, about 69bp, About 70bp, about 71bp, about 72bp, about 73bp, about 74bp, about 75bp, about 76bp, about 77bp, about 78bp, about 79bp, about 80bp, about 81bp, about 82bp, about 83 bp, about 84bp, about 85bp, about 86bp, about 87bp, about 88bp, about 89bp, about 90bp, about 91bp, about 92bp, about 93bp, about 94bp, about 95bp, about 96bp, about 97 bp, about 98 bp, about 99 bp, about 100 bp, about 150 bp, about 200 bp, about 250 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, about 500 bp, about 550 bp, about 600 bp, about 650 bp, about 700 bp, about 750 bp, about 800 bp, about 850 bp, about 900 bp, about 950 bp, about 1000 bp, or more.In some embodiments, the apparent insertion length is 20 bp or more, 21 bp or more, 22 bp or more, 23 bp or more, 24 bp or more, 25 bp or more, 26 bp or more, 27 bp or more, 28 bp or more, 29 bp or more, 30 bp or more, 31 bp or more, 32 bp or more, 33 bp or more, 34 bp or more, 35 bp or more, 36 bp or more, 37 bp or more, 38 bp or more, 39 bp or more, 40 bp or more, 41 bp or more, 42 bp or more, 43 bp or more, 44 bp or more, 45 bp or more, 46 bp or more, 47 bp or more, 48 bp or more, 49 bp or more, 50 bp or more, 51 bp or more, 52 bp or more, 53 bp or more, 54 bp or more, 55 bp or more, 56 bp or more, 57 bp or more, 58 bp or more, 59 bp or more, 60 bp or more, 61 bp or more, 62 bp or more, 63 bp or more, 64 bp or more, 65 bp or more, 66 bp or more, 67 bp or more, 68 bp or more, 69 bp or more, 70 bp or more, 71 bp or more, 72 bp or more, 73 bp or more, 74 bp or more, 75 bp or more, 76 bp or more, 77 bp or more, 78 bp or more, 79 bp or more, 80 bp or more, 81 bp or more, 82 bp or 3bp or more, 44bp or more, 45bp or more, 46bp or more, 47bp or more, 48bp or more, 49bp or more, 50bp or more, 51bp or more, 52bp or more, 53bp or more, 54bp or more, 55bp or more, 56bp or more, 57bp or more, 58bp or more, 59bp or more, 60bp or more, 61bp or more, 62bp or more, 63bp or more, 64bp or more, 65bp or more, 66bp or more, 67bp or more, 68bp or more, 69bp or more , 70bp or more, 71bp or more, 72bp or more, 73bp or more, 74bp or more, 75bp or more, 76bp or more, 77bp or more, 78bp or more, 79bp or more, 80bp or more, 81bp or more, 82bp or more, 83bp or more, 84bp or more, 85bp or more, 86bp or more, 87bp or more, 88bp or more, 89bp or more, 90bp or more, 91bp or more, 92bp or more, 93bp or more, 94bp or more, 95bp or more, 96bp or more, 97bp or more, 98bp or more, 99bp or more, 100bp or more, 150bp or more, 200bp or more, 250bp or more, 300bp or more, 350bp or more, 400bp or more, 450bp or more, 500bp or more, 550bp or more, 600bp or more, 650bp or more, 700bp or more, 750bp or more, 800bp or more, 850bp or more, 900bp or more, 950bp or more, 1000bp or more, or more.In some embodiments, the apparent insertion length is 20bp, 21bp, 22bp, 23bp, 24bp, 25bp, 26bp, 27bp, 28bp, 29bp, 30bp, 31bp, 32bp, 33bp, 34bp, 35bp, 36bp, 37bp, 38bp, 39bp, 40bp, 41bp, 42bp, 43bp, 44bp, 45bp, 46bp, 47bp, 48bp, 49bp, 50bp, 51bp, 52bp, 53bp, 54bp, 55bp, 56bp, 57bp, 58bp, 59bp, 60bp, 61bp, 62bp, 63bp, 64bp, 65bp, 66bp, 67bp, 68bp, 69bp, 70bp, 71bp, 72bp, 73bp, 74bp, 75bp, 76bp, 77bp, 78bp, 79bp, 80bp, 81bp, 82bp, 83bp, 84bp, 85bp, 86bp, 87bp, 88bp, 89bp, 90bp, 91bp, 92bp, 93bp, 94bp, 95bp, 96bp, 97bp, 98bp, 99bp, 100bp, 150bp, 200bp, 250bp, 300bp, 350bp, 400bp, 450bp, 500bp, 550bp, 600bp, 650bp, 700bp, 750bp, 800bp, 850bp, 900bp, 950bp, 1000bp, or more.
[0047] In some embodiments, the apparent insertion length is greater than 1 bp, greater than 2 bp, greater than 3 bp, greater than 4 bp, greater than 5 bp, greater than 6 bp, greater than 7 bp, greater than 8 bp, greater than 9 bp, greater than 10 bp, greater than 11 bp, greater than 12 bp, greater than 13 bp, greater than 14 bp, greater than 15 bp, greater than 16 bp, greater than 17 bp, greater than 18 bp, greater than 19 bp, or greater than 20 bp. In some embodiments, the apparent insertion length is greater than 20 bp.
[0048] In some embodiments, the apparent structural variant is due to the apparent insertion or the apparent duplication.
[0049] After the apparent indels and / or apparent structural variants are identified, the sequences can be evaluated to determine whether they are derived from a candidate eccDNA molecule, and in some embodiments, the probability that a given candidate eccDNA molecule is actually an eccDNA molecule is calculated or estimated. For example, FIG. 1B shows a flowchart including steps that can be performed to identify candidate eccDNA or candidate eccDNA in a sample according to embodiments of the present disclosure. In the first step of the flowchart shown in FIG. 1B, these indels or insertions are identified or provided. Categories of the insertion sequences include, for example, tandem duplications (TDs), insertions that are not tandem duplications (non-TD insertions), and eccDNA molecules. In the next step of the flowchart shown in FIG. 1B, non-TD insertions are excluded from the category by detecting candidate eccDNA breakpoint sequences (denoted as "BA" breakpoints in FIG. 1B). The breakpoints can be present in both eccDNA and tandem duplications, and include sequences A and B that are present on the same chromosome in the reference genome and are separated by a distance of Y nucleotides. For example, the distance Y starts from the first nucleotide of sequence A and continues to the last nucleotide of sequence B (as shown in Figure 1B, α refers to the entire sequence from the first nucleotide of A to the last nucleotide of B, and the length of α in the genome is equal to Y nucleotides).
[0050] In contrast to the arrangement of A and B in the reference genome, at the candidate breakpoint, sequence A and sequence B are approximately adjacent, with their order reversed on the chromosome. Figure 1B illustrates the possibility that a sequence starting at A (i.e., the first nucleotide of A) and ending at B (i.e., the last nucleotide of B) in the reference genome could result in the generation of a candidate breakpoint BA sequence, either through a tandem duplication event or the formation of eccDNA. In some embodiments, the breakpoint includes the last nucleotide of sequence B, immediately upstream of the first nucleotide of sequence A. However, due to potential imprecision in the closure of the eccDNA circular structure, it should be understood that in some embodiments, the last nucleotide of B and the first nucleotide of A may be separated by one or more nucleotides, or the last nucleotide of B and / or the first nucleotide of A may be deleted or mutated.
[0051] Next, as shown in Figure 1B, sequences containing candidate eccDNA breakpoints are evaluated using any one or more of a variety of methods to assess the likelihood that they correspond to bona fide eccDNA molecules. These evaluations are based on several properties of eccDNA (i.e., "simple" eccDNA formed by circularization of a single contiguous segment of the genome, as opposed to hybrid or chimeric eccDNA, which may contain segments from other sites within the same genome or from any other source) that allow them to be distinguished from tandem duplications. For example, since the simple eccDNA formed as shown in Figure 1B essentially consists of sequence α from the first nucleotide of sequence A to the last nucleotide of sequence B (having length Y), sequences derived from bona fide eccDNA would not extend beyond distance Y, would not contain sequences shown to the left of A or to the right of B in Figure 1B, and would not contain multiple copies of any portion of α.
[0052] Thus, a sequence having a BA breakpoint may include any sequence shown to the left of A or to the right of B in Figure 1B, Figure 6A, or Figures 7A-7C that is longer than Y (e.g., 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more nucleotides than Y) (e.g., a contiguous sequence of 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, or more nucleotides to the left of A or to the right of B). Multiple copies of a subsequence within α (e.g., a contiguous sequence of 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, or more nucleotides) indicate that the sequence is not derived from eccDNA. It should be understood that it is not necessary to perform all three of these assessments; any one assessment may be sufficient. Thus, in some embodiments, only one of the three assessments is used.
[0053] In some embodiments, the length of the fragment containing the candidate eccDNA (or the length of the consensus sequence obtained from the fragment) is compared to distance Y to determine whether the length of the fragment or consensus sequence is approximately equal to or exactly equal to Y. As described above, the eccDNA formed as shown in Figure 1B, 6A, or 7A-7C has a length approximately equal to distance Y, and therefore, if the eccDNA is cleaved (i.e., linearized) only once during the library preparation process, the fragments in the sequencing library derived from the eccDNA will also have a length approximately equal to Y. Alternatively, if the eccDNA is cleaved multiple times during the fragmentation process, and the resulting fragments containing BA sequences are ligated to sequencing adaptors and included in the library, the length of the insert fragment will be shorter than Y.
[0054] Thus, the observation that the length of a fragment or consensus sequence is equal to or nearly equal to the distance Y can be evidence that the fragment is derived from authentic eccDNA (in the case of tandem duplications, it is generally unlikely that a fragment length equal to Y will occur by chance, especially if the average length of the fragments in the library is significantly greater than the distance Y). Such evidence can be used, for example, to calculate or estimate the probability that a given fragment or consensus sequence is derived from authentic eccDNA.
[0055] Further, see Figure 1C, which shows a flowchart including steps for identifying candidate eccDNA molecules according to an embodiment. In certain embodiments, identifying at least one eccDNA in a sample includes identifying the presence or absence of eccDNA in the sample.
[0056] Step 160 involves obtaining a plurality of sequence reads of the sequenced double-stranded DNA using double-strand sequencing. In some embodiments, obtaining a plurality of sequence reads of the sequenced double-stranded DNA using double-strand sequencing includes performing or having performed double-strand sequencing on the sample. In some embodiments, the double-strand sequencing includes ligating adapters to the ends of the double-stranded DNA, at least one of the adapters including a sequence that labels one strand of the double-stranded DNA with a unique nucleotide sequence that distinguishes it from the complementary strand. In some embodiments, the double-strand sequencing includes amplifying a strand of the double-stranded DNA using the ligated adapters to generate at least a first-strand amplicon and a second-strand amplicon. In some embodiments, the double-strand sequencing includes sequencing at least the first-strand amplicon and the second-strand amplicon to generate a plurality of sequence reads including a first-strand sequence read and a second-strand sequence read. In some embodiments, obtaining a plurality of sequence reads of double-stranded DNA sequenced using double-stranded sequencing comprises identifying or having identified eccDNA using the plurality of sequence reads of the double-stranded sequencing.
[0057] In step 165, a subset of sequence reads are identified, each of which independently includes a reference allele junction (e.g., DA), where the nucleic acid sequence at the end of the reference allele (e.g., D in reference allele ABCD) joins the nucleic acid sequence at the start of the reference allele (e.g., A in reference allele ABCD).
[0058] Step 170 distinguishes between candidate eccDNA sequence reads and chromosomal tandem duplication sequence reads. In various embodiments, in addition to distinguishing between candidate eccDNA sequence reads and chromosomal tandem duplication sequence reads, step 175 includes selecting a subset of DA junction-containing consensus sequencing reads whose allele lengths are less than a threshold apparent insert size, estimating the fragment size of each read pair in the subset and comparing the fragment size to a threshold apparent insert size one by one, or comparing the distribution of estimated fragment sizes to the insert size distribution of the entire library, and comparing the estimated fragment size of any consensus read pair to the allele size of that read pair.
[0059] In step 180, eccDNA is identified based on the candidate eccDNA sequence reads identified in step 170. In some embodiments, identifying eccDNA based on the identified candidate eccDNA sequence reads comprises determining the amount of eccDNA according to the candidate eccDNA sequence reads. In some embodiments, identifying eccDNA based on the identified candidate eccDNA sequence reads comprises determining the frequency of eccDNA according to the candidate eccDNA sequence reads. In some embodiments, identifying eccDNA based on the identified candidate eccDNA sequence reads comprises determining the quality of eccDNA according to the candidate eccDNA sequence reads. In some embodiments, identifying eccDNA based on the identified candidate eccDNA sequence reads comprises determining any characteristic of the eccDNA according to the candidate eccDNA sequence reads. In some embodiments, any characteristic of the eccDNA includes, but is not limited to, size, chromosomal location of origin, chromosomal origin annotation (e.g., gene region, exon region, intron region, regulatory element, repetitive sequence), nucleotide sequence content (e.g., GC content, mononucleotide, dinucleotide, trinucleotide repeat, junction microhomology, etc.).
[0060] Thus, in some embodiments, the terms "candidate eccDNA molecule" or "candidate eccDNA" are used interchangeably and correspond to double-stranded fragments obtained from a biological sample, or consensus sequences derived from such fragments, that are identified as apparent indels or apparent structural variants and contain candidate eccDNA breakpoints (the "BA" or "DA" breakpoints shown in FIG. 1B, FIG. 6A, or FIG. 7A-7C, which may be used interchangeably). In some embodiments, a sequence is determined to correspond to a candidate eccDNA molecule if it satisfies all of the following conditions: (i) its length does not exceed distance Y; (ii) it does not include sequences outside the genomic region defined by A and B, or A, B, C, and D in the reference genome (the sequences to the left of A or right of B in FIG. 1B, or the sequences to the left of A and right of D in FIG. 6A or FIG. 7A-7C); and (iii) it does not include multiple copies of a subsequence from region α from the first nucleotide of sequence A to the last nucleotide of sequence B in the reference genome.
[0061] As mentioned above, in some embodiments, such "candidate eccDNA" may include bona fide eccDNA molecules as well as short fragments derived from tandem duplications (i.e., fragments shorter than distance Y). Accordingly, the present disclosure provides additional steps to help identify bona fide eccDNA molecules from among this category of candidates, and / or to help estimate or calculate the probability that a given candidate is in fact an eccDNA molecule.
[0062] In some embodiments, information about the relative size of fragments corresponding to a candidate eccDNA molecule to the sizes of fragments in a sequencing library can be used to calculate or estimate the probability that the fragment is derived from a bona fide eccDNA molecule. For example, as described elsewhere herein, the observation that a fragment has a size exactly or approximately the same as distance Y can be an indication that the fragment is likely derived from a bona fide eccDNA molecule, particularly if the average length of the fragments in the library is greater than Y.
[0063] Furthermore, the greater the average length of the fragments in the library is, and the smaller the variation in fragment size (e.g., standard deviation from the average fragment length), the higher the probability that the fragments are true eccDNA and not small fragments of tandem duplication that arose by chance. In some embodiments of the method, as described elsewhere herein, the library is prepared with a relatively high average fragment length to reduce the likelihood that a given fragment in the library will have a size less than or equal to Y nucleotides by chance. In some embodiments, the average length (e.g., mean or median) of the fragments in the library is between about 100 bp and 1000 bp. In some embodiments, the average length (e.g., mean or median) of the fragments in the library is about 1000 bp, 2000 bp, 3000 bp, 4000 bp, 5000 bp, 6000 bp, 7000 bp, 8000 bp, 9000 bp, or 10,000 bp.
[0064] In some embodiments, the length of the candidate eccDNA molecule is between about 100 and 1000 nucleotides, or less than about 500 nucleotides. In some embodiments, the length of the candidate eccDNA molecule is greater than about 1000 bp, 2000 bp, 3000 bp, 4000 bp, 5000 bp, 6000 bp, 7000 bp, 8000 bp, 9000 bp, or 10,000 bp. In some embodiments, the length of the candidate eccDNA molecule is greater than about 100 kb, 200 kb, 300 kb, 400 kb, 500 kb, 600 kb, 700 kb, 800 kb, 900 kb, 1 Mb, 2 Mb, or 3 Mb.
[0065] In some embodiments, the candidate eccDNA molecule comprises a gene (e.g., a coding sequence, a promoter, a regulatory element). In some embodiments, the candidate eccDNA molecule comprises an origin of replication. In some such embodiments, multiple copies of the eccDNA are predicted or observed to be present in one or more cells of a sample.
[0066] In certain embodiments, the method is performed multiple times with a given biological sample, one or more times including an enrichment step for circular DNA molecules and one or more times not including the enrichment step. The enrichment step for eccDNA molecules can be performed by any of a variety of methods. For example, in some embodiments, one or more exonucleases are used to remove linear (i.e., non-circular) DNA molecules from the sample. In such embodiments, some or all of the linear DNA in the sample is digested, leaving only circular DNA molecules in the sample. In some embodiments, circular DNA molecules are selectively isolated or purified from the sample. Any method capable of separating linear double-stranded DNA molecules from covalently closed circular double-stranded DNA molecules can be used in the method.
[0067] By performing the method with and without the enrichment step and comparing the candidate eccDNA molecules detected in both conditions, it is possible to determine or estimate the number of bona fide eccDNA molecules present in the original sample, i.e., the proportion of candidate eccDNA molecules detected in the absence of exonuclease treatment that are bona fide eccDNA molecules.
[0068] A non-limiting list of exonucleases that can be used in the present methods includes, for example, Exonuclease I, Exonuclease T, Exonuclease VII, Exonuclease III, T7 Exonuclease, Exonuclease V (RecBCD), Exonuclease VIII, Lambda Exonuclease, and / or T5 Exonuclease. In some embodiments, one or more exonucleases are used. In some embodiments, the one or more exonucleases include Exonuclease I, Exonuclease T, Exonuclease VII, Exonuclease III, T7 Exonuclease, Exonuclease V (RecBCD), Exonuclease VIII, Lambda Exonuclease, or T5 Exonuclease. In some embodiments, the one or more exonucleases include Exonuclease I. In some embodiments, the one or more exonucleases comprise exonuclease T. In some embodiments, the one or more exonucleases comprise exonuclease VII. In some embodiments, the one or more exonucleases comprise exonuclease III. In some embodiments, the one or more exonucleases comprise T7 exonuclease. In some embodiments, the one or more exonucleases comprise exonuclease V (RecBCD). In some embodiments, the one or more exonucleases comprise exonuclease VIII. In some embodiments, the one or more exonucleases comprise lambda exonuclease. In some embodiments, the one or more exonucleases comprise T5 exonuclease.
[0069] In some embodiments, a portion of the sample is also subjected to an endonuclease treatment prior to exonuclease treatment, which linearizes any circular DNA molecules in the sample and removes all double-stranded DNA molecules in the sample, providing additional evidence that the DNA molecules remaining after exonuclease treatment are truly circular.
[0070] In some embodiments, the enrichment step comprises performing size selection. In some embodiments, the size selection comprises the use of paramagnetic beads, electrophoresis, column filtration, density gradient centrifugation, or selective extraction. In some embodiments, the size selection is performed at a size threshold of about 10,000 bp. In some embodiments, the size selection comprises the use of paramagnetic beads, electrophoresis, column filtration, density gradient centrifugation, or selective extraction, and is performed at a size threshold of at least 10,000 bp. In some embodiments, the size selection uses a size threshold of about 10,000 bp. In some embodiments, the size selection uses a size threshold of at least 10,000 bp. In some embodiments, the size selection comprises the use of paramagnetic beads. In some embodiments, the size selection comprises the use of electrophoresis. In some embodiments, the size selection comprises the use of column filtration. In some embodiments, the size selection comprises the use of density gradient centrifugation. In some embodiments, the size selection comprises the use of selective extraction. In some embodiments, the size selection comprises the use of paramagnetic beads at a size threshold of about 10,000 bp. In some embodiments, the size selection comprises using electrophoresis at a size threshold of about 10,000 bp. In some embodiments, the size selection comprises using column filtration at a size threshold of about 10,000 bp. In some embodiments, the size selection comprises using density gradient centrifugation at a size threshold of about 10,000 bp. In some embodiments, the size selection comprises using paramagnetic beads at a size threshold of at least 10,000 bp. In some embodiments, the size selection comprises using electrophoresis at a size threshold of at least 10,000 bp. In some embodiments, the size selection comprises using column filtration at a size threshold of at least 10,000 bp. In some embodiments, the size selection comprises using density gradient centrifugation at a size threshold of at least 10,000 bp.In some embodiments, the size selection comprises using selective extraction with a size threshold of at least 10,000 bp.
[0071] In some embodiments, the enrichment step involves electrophoresis, column filtration (e.g., using silica columns designed or used for plasmid isolation), density gradient centrifugation, selective extraction, and / or the use of a DNA-binding protein that differentially binds or retains double-stranded circular DNA molecules relative to double-stranded linear DNA molecules. In some embodiments, the DNA-binding protein is a helicase or other protein that binds to and translocates along double-stranded DNA, thus dissociating from linear DNA with ends but not from circular DNA with no ends. According to such embodiments, circular DNA molecules can be isolated by binding the protein to DNA in a sample and purifying the protein-DNA complex by affinity-based methods (e.g., using a specific antibody against the protein, biotinylating the protein, or other methods known to those of skill in the art).
[0072] In some embodiments, for example, where it is important to ensure that each detected eccDNA is eccDNA, systematic enrichment steps, such as size selection or exonuclease treatment, can be performed to ensure that all detected eccDNA is in fact eccDNA and not tandem duplicate fragments, etc. However, in some embodiments, it may be sufficient to be able to calculate or estimate the probability that a given eccDNA molecule is in fact eccDNA based on information obtained during a separate performance of the method, including the enrichment step. Furthermore, even when an enrichment step is performed, it will be understood that the performance of the method, including the enrichment step, and the performance of the method, excluding the enrichment step, do not have to be performed simultaneously. For example, after collection of a biological sample, an initial analysis, including the enrichment step, can be performed first, followed by one or more follow-up analyses, without the enrichment step, using the same biological sample. The follow-up analyses can be performed at any time, such as months or years after the initial performance of the method, including the enrichment step. Similarly, the method can be performed initially without the enrichment step, followed by analyses including the enrichment step. For example, when candidate eccDNA molecules are detected in a sample, it may be desirable to repeat the method, including exonuclease treatment or other enrichment steps, to confirm, quantify, and / or characterize the detected candidate eccDNA molecules. Furthermore, in some embodiments, no enrichment step may be performed on a given biological sample, such as when sufficient previous analysis, including enrichment steps, has been performed on similar or related samples to allow new samples to be reliably analyzed without enrichment.
[0073] In some embodiments, one or more conditions during library preparation can be altered and the candidate and / or eccDNA molecules detected before and after the alteration can be compared to identify conditions that can improve the yield of total eccDNA molecules in a sample or the yield of specific types of eccDNA molecules (e.g., eccDNA of a particular size, the presence or absence of an origin of replication, the presence or absence of a gene region, etc.). Library preparation conditions that can be altered in such studies include, but are not limited to, steps related to nucleic acid isolation from the sample, nucleic acid washing or pretreatment (e.g., DTT treatment), fragmentation conditions (mechanical fragmentation, sonication, enzymatic fragmentation such as Covaris, including the type and concentration of enzyme used), duration of the fragmentation step, type and concentration of sequencing adapters used, conditions for the ligation step, amount of DNA used, etc.
[0074] The inclusion of enrichment steps such as size selection and exonuclease treatment has different effects on different nucleic acids in a biological sample. For example, because exonucleases act on nucleic acid ends, they are expected to digest linear nucleic acids in a sample but not circular nucleic acids (especially circular nucleic acids without damage or nicks). Therefore, because the different steps shown in the flowcharts of Figures 1B and 1C may detect different types of linear DNA (e.g., tandem duplications (TD) and non-TD insertions), the number of fragments detected at each step is expected to decrease as circular DNA molecules are enriched. For example, for fragments containing indels or insertions provided or received in the first step, without enrichment, the category may include TD, non-TD insertions (or other indels), and eccDNA; however, with enrichment, the same category will only include eccDNA (provided that TD and non-TD insertions are completely removed by the enrichment step). Similarly, for fragments containing BA breakpoints detected in subsequent steps, without enrichment, the category may include TD and eccDNA; however, with enrichment, only eccDNA will be present. Furthermore, regarding the "candidate eccDNA" molecules remaining after the third step, without enrichment, the category could in principle contain both eccDNA and TD-derived fragments shorter than Y, whereas with enrichment, only eccDNA should remain. Therefore, if non-TD insertions and / or TDs exist in the fragments along with eccDNA, the number of fragments detected in the first two steps may be reduced more significantly by enrichment than the number of candidate eccDNAs detected in the final step.
[0075] Thus, in some embodiments, the ratio of the number of candidate eccDNA molecules detected in the final step shown in FIG. 1B to the total number of error-corrected sequences obtained from the library, or to the number of insertions or indels detected in the sequences, or to the number of candidate eccDNA breakpoints detected in insertion- or indel-containing sequences, is higher when the method includes an enrichment step than when the method does not include an enrichment step.
[0076] In some embodiments, the number of error-corrected sequences obtained with a method that includes an enrichment step is less than about 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, or 10% of the number of error-corrected sequences obtained with a method that does not include an enrichment step.
[0077] In some embodiments, the frequency of apparent indels or apparent structural variants in the error-corrected sequences detected by a method that includes an enrichment step is less than about 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, or 10% of the frequency of possible insertions or indels detected by a method that does not include an enrichment step.
[0078] In some embodiments, the frequency of candidate eccDNA breakpoints detected in fragments containing apparent indels or apparent structural variants detected by a method that includes an enrichment step is less than about 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, or 10% of the frequency of candidate eccDNA breakpoints detected by a method that does not include an enrichment step.
[0079] In some embodiments, the frequency of candidate eccDNA molecules detected in fragments containing eccDNA breakpoints detected by a method that includes an enrichment step is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or more of the frequency of candidate eccDNA molecules detected by a method that does not include an enrichment step.
[0080] In some embodiments, enrichment for circular DNA molecules may not significantly decrease the frequency of different categories, e.g., when substantially all of the insertions present in a given sample correspond to eccDNA molecules. In some embodiments (e.g., when substantially all of the insertions present in a given sample correspond to eccDNA molecules), enrichment for circular DNA molecules may increase the frequency of different categories.
[0081] In some embodiments, any one or more quantitative elements of the information obtained by the method (such as any frequency, ratio, or percentage described herein) are used in a calculation to determine or estimate the probability that the candidate eccDNA molecule identified in step (e) is a bona fide eccDNA molecule. In some embodiments, the calculation is performed partially or completely on a computer and / or cloud.
[0082] In some embodiments, the calculation is based in part on the relationship between the length of the candidate eccDNA molecule and the average length of the double-stranded DNA fragments in the library, where a shorter length of the candidate eccDNA molecule relative to the average length of the double-stranded DNA fragments in the library indicates a higher probability that the candidate eccDNA molecule is a bona fide eccDNA molecule.
[0083] In some embodiments, the calculation is based in part on the observation that the length of the candidate eccDNA molecule is nearly or exactly equal to the distance Y nucleotides, which observation indicates a high probability that the candidate eccDNA molecule is a bona fide eccDNA molecule.
[0084] The present disclosure also provides computer-based systems for performing any one or more of the methods disclosed herein.
[0085] In some embodiments, the present disclosure provides for the treatment of a disease or other medical condition in a mammalian subject. In some such embodiments, the method comprises the steps of (i) performing any of the methods disclosed herein on a biological sample obtained from the subject, (ii) identifying one or more candidate eccDNA molecules in the sample that are indicative of a physiological condition associated with the disease or medical condition, and (iii) administering to the subject a treatment for the disease or medical condition.
[0086] In some embodiments, the disease state or physiological condition is selected from the group consisting of cancer, inflammation, autoimmune disease, infectious disease, organ transplant rejection, stem cell transplant rejection, therapeutic cell rejection, therapeutic cell response, immunotherapy response, pregnancy, pre-eclampsia, radiation exposure, sun exposure, drug exposure, and hypersensitivity.
[0087] II. Sequencing Library Preparation The method is based on the analysis of sequence information obtained from double-stranded nucleic acid obtained from a biological sample and can be used to detect and / or quantify eccDNA in a biological sample.
[0088] Biological samples The present methods can be used to detect, quantify, and / or characterize candidate extrachromosomal circular DNA (eccDNA) molecules in any type of biological sample. In some embodiments, the biological sample contains cells. In some embodiments, the biological sample contains cell-free DNA. In some embodiments, the biological sample contains both cells and cell-free DNA. In some embodiments, the biological sample is obtained from a subject, such as a blood sample, tissue sample, tumor biopsy, liquid biopsy, swab, lavage fluid, urine sample, saliva sample, or other sample containing cells and / or cell-free DNA that can be analyzed using the present methods. In some embodiments, the biological sample contains sperm cells, such as a sperm sample or semen sample. In some embodiments, the sample is a prostatic fluid sample, a testicular biopsy sample, a spermatogonial sample, a germ cell sample, a gamete sample, a swab, a lavage fluid, an aspirate, a biopsy tissue, a tissue sample, a tumor sample, a precancerous lesion sample, a liquid biopsy, a hyperplasia sample, a hypertrophy sample, a dysplasia sample, a urine sample, a cerebrospinal fluid (CSF) sample, other body fluid sample, an autopsy sample, a postmortem sample, a surgical sample, a model organism sample, a plasma sample, a serum sample, a gastric juice sample, a bone marrow sample, a stool sample, a brushing sample, a bile sample, a pancreatic juice sample, a synovial fluid sample, a sputum sample, a mucus sample, a vitreous humor sample, a forensic sample, an environmental sample, a bacterial sample, a fungal sample, a mammalian sample, a human sample, and a diagnostic sample.
[0089] In some embodiments, the biological sample contains cancer cells or nucleic acids derived from cancer cells, such as a tumor sample, blood sample, or other liquid biopsy. In some embodiments, the presence and / or characteristics of eccDNA molecules in a sample can indicate the presence of a particular genetic event in an individual, such as cancer-associated genomic instability or apoptotic DNA degradation. Furthermore, in some embodiments, the presence and / or characteristics of eccDNA molecules in a sample can be used to identify disease states or physiological conditions, such as inflammation, autoimmune disease, infection, organ transplant rejection, stem cell transplant rejection, therapeutic cell rejection, therapeutic cell response, immunotherapy response, pregnancy, preeclampsia, radiation exposure, sun exposure, drug exposure, hypersensitivity, etc.
[0090] In some embodiments, the biological sample comprises cells that have been exposed to a potentially toxic agent (e.g., a potentially clastogenic, aneugenic, mutagenic, and / or teratogenic agent). In some embodiments, the biological sample is obtained from an individual that has been exposed to the agent, or comprises cells that have been intentionally exposed to the agent in order to assess the genotoxic potential of the agent.
[0091] As used herein, the terms "biological sample" or "sample" generally refer to a sample obtained or derived from a biological source of interest (e.g., a tissue, organism, or cell culture) as described herein. In some embodiments, the source of interest includes an organism, such as an animal or a human. In other embodiments, the source of interest includes a microorganism, such as a bacterium, a virus, a protozoan, or a fungus. In still other embodiments, the source of interest may be a synthetic tissue, a synthetic organism, a cell culture, a nucleic acid, or other material. In still other embodiments, the source of interest may be a plant-based organism. In still other embodiments, the sample may be an environmental sample (e.g., a water sample, a soil sample, an archaeological sample, or other sample collected from a non-living source). In other embodiments, the sample may be a multi-organism sample (e.g., a mixed-organism sample). In some embodiments, the biological sample is or includes a biological tissue or a bodily fluid. In some embodiments, the biological sample may consist of or comprise bone marrow, blood, blood cells, ascites, tissue or fine needle biopsy sample, cell-containing bodily fluid, free nucleic acid, sputum, saliva, urine, cerebrospinal fluid, peritoneal fluid, pleural fluid, feces, lymphatic fluid, gynecological fluid, skin swab, vaginal swab, Pap smear, oral swab, nasal swab, lavage or irrigation fluid (e.g., ductal lavage or bronchoalveolar lavage), vaginal fluid, aspirate, scraping, bone marrow specimen, tissue biopsy specimen, fetal tissue or fluid, surgical specimen, feces, other bodily fluid, secretion, and / or excretion, and / or cells obtained therefrom. In some embodiments, the biological sample is or comprises cells obtained from an individual. In some embodiments, the obtained cells are or comprise cells obtained from the individual from whom the sample was taken. In certain embodiments, the biological sample is a liquid biopsy obtained from a subject. In some embodiments, the sample is a "primary sample" taken directly from the source of interest. The primary sample is taken by any suitable means.For example, in some embodiments, a primary biological sample is obtained by a method selected from the group consisting of biopsy (e.g., fine needle aspiration or tissue biopsy), surgery, or bodily fluid collection (e.g., blood, lymph, stool, etc.). In some embodiments, as the context will dictate, the term "sample" refers to a preparation obtained by processing the primary sample (e.g., by removing one or more components of the primary sample and / or adding one or more reagents), such as by filtration through a semipermeable membrane. Such a "processed sample" may include, for example, nucleic acids or proteins extracted from the sample, or materials obtained by processing the primary sample using techniques such as mRNA amplification, reverse transcription, or isolation and purification of specific components.
[0092] In some embodiments, the sample is a forensic sample (e.g., blood, tissue, sperm, hair, saliva, or other sample containing cells or cell-free DNA collected from a known or unknown source), and the method can be used to identify, for example, the individual from whom the sample originated, or the tissue or cell type from which the cell-free DNA originated.
[0093] As used herein, the term "subject" generally refers to a mammal (e.g., a human, including in some embodiments a fetus). In some embodiments, the subject is suffering from the relevant disease, disorder, or condition. In some embodiments, the subject is susceptible to the disease, disorder, or condition. In some embodiments, the subject exhibits one or more symptoms or characteristics of the disease, disorder, or condition. In some embodiments, the subject does not exhibit any symptoms or characteristics of the disease, disorder, or condition. In some embodiments, the subject possesses one or more characteristics indicative of susceptibility to or risk for the disease, disorder, or condition. In some embodiments, the subject is a patient. In some embodiments, the subject is an individual for whom diagnosis and / or treatment is or has been performed.
[0094] fragmentation Double-stranded DNA obtained from a biological sample can be fragmented in a variety of ways. For example, fragmentation can be achieved by physical shearing (sonication, Covaris fragmentation, etc.) or enzymatic methods using enzyme cocktails that cleave DNA phosphodiester bonds. As a result of either of these methods, the intact nucleic acid material (e.g., genomic DNA (gDNA)) is reduced to a mixture of randomly or semi-randomly sized nucleic acid fragments. In certain embodiments of this method, enzymatic fragmentation is used.
[0095] Adapter ligation In certain embodiments, the method includes ligating one or more sequencing adapters to the fragmented double-stranded nucleic acid molecules to generate double-stranded adapter-fragment complexes. The adapter molecules may contain one or more of a variety of features suitable for MPS or next-generation sequencing (NGS) platforms. For example, they may include a sequencing primer recognition site, an amplification primer recognition site, a barcode (e.g., a single molecule identifier (SMI) sequence, an indexing sequence), a single-stranded portion, a double-stranded portion, a strand discrimination element (SDE), or the like. In DS or any next-generation sequencing technology, using highly pure sequencing adapters is important for obtaining high-quality, reproducible data and maximizing the sequencing yield of a sample (i.e., the relative proportion of input molecules converted into independent sequence reads). This is particularly important in DS technology, as it is necessary to ensure the recovery of both strands of the original double-stranded molecule.
[0096] In some embodiments, the adapter has a Y-shape. In some embodiments, the adapter has a loop or hairpin shape. In some embodiments, the adapter is adapted to be as described in U.S. Patent Nos. 11,332,784, 11,479,807, 10,287,631, 9,752,188, 11,155,869, 11,098,359, 11,242,562, 11,198,907, 10,570,451, 10,385,393, 10,370,713, 11,130,996, 10, In one embodiment, one or more of the adapters disclosed in U.S. Pat. Nos. 10,689,699, 10,604,804, 10,689,700, 10,711,304, 10,760,127, 10,752,951, 11,047,006, 11,118,225, 11,608,529, 11,555,220, or 11,549,144 (all of which are incorporated herein by reference in their entirety) may be used.
[0097] IV. Sequencing and Analysis As described herein, the method includes performing a sequencing assay (e.g., sequencing assay 120 described in connection with FIG. 1A) to analyze a sample. In certain embodiments, the sequencing assay includes an error-corrected sequencing method, such as double-stranded sequencing (DS). Double-stranded sequencing (DS) is a method for generating error-corrected nucleic acid sequence reads from double-stranded nucleic acid molecules. In certain aspects of the technology, DS can be used to independently sequence both strands of individual nucleic acid molecules, such that sequence reads derived during massively parallel sequencing are recognizable as originating from the same double-stranded nucleic acid parent molecule and, after sequencing, distinguishable as distinct, separate entities. By comparing the sequence reads obtained from each strand, it is possible to obtain an error-corrected sequence of the original double-stranded nucleic acid molecule (known as a double-stranded consensus sequence). The DS process allows for confirmation of whether one or both strands of the original double-stranded nucleic acid molecule are represented in the generated sequencing data used to form the double-stranded consensus sequence. Double-stranded sequencing methods are described, for example, in U.S. Pat. Nos. 9,752,188, 11,479,807, 10,287,631, 9,752,188, 11,155,869, 11,098,359, 11,242,562, 11,198,907, 10,570,451, 10,385,393, 10,370,713, 11,130,996, Nos. 10,689,699, 10,604,804, 10,689,700, 10,711,304, 10,760,127, 10,752,951, 11,047,006, 11,118,225, 11,608,529, 11,555,220, and 11,549,144, all of which are incorporated herein by reference in their entirety.
[0098] In addition to double-stranded sequencing, any sequencing method capable of generating error-corrected sequencing reads is encompassed within the scope of the present disclosure.For example, many embodiments of single-consensus sequencing and / or a combination of single-consensus sequencing and double-stranded consensus sequencing are envisioned.Furthermore, other embodiments of the present technology may have different configurations, components, or procedures than those described herein.Therefore, those skilled in the art will understand that the present technology may have other embodiments with additional elements, and may have other embodiments that do not have some of the features shown and described herein.
[0099] After generating double-stranded adapter-fragment DNA complexes containing at least one SMI (single molecule identifier) and at least one SDE (strand discrimination element), the complexes are subjected to DNA amplification, such as PCR, or other biochemical DNA amplification methods, to generate one or more copies of the first-strand target nucleic acid sequence and one or more copies of the second-strand target nucleic acid sequence. The one or more amplified copies of the first-strand target nucleic acid molecule and the one or more amplified copies of the second-strand target nucleic acid molecule can then be subjected to DNA sequencing, preferably performed using a "next-generation" massively parallel DNA sequencing platform.
[0100] Sequence reads generated from first-strand and second-strand target nucleic acid molecules derived from an original double-stranded target nucleic acid molecule are distinguishable based on sharing a related, substantially unique SMI, and in some embodiments (e.g., double-stranded sequencing embodiments), can be distinguished from complementary strand target nucleic acid molecules by SDE. After identification, an error-corrected sequence can be generated by comparing one or more sequence reads generated from the first-strand target nucleic acid molecule with one or more sequence reads generated from the second-strand target nucleic acid molecule. For example, nucleotide positions where bases match between the first-strand and second-strand target nucleic acid sequences are considered to be true sequences, while mismatched nucleotide positions between the two strands can be identified as potential sites of technical errors and excluded. This allows for the generation of an error-corrected sequence of the original double-stranded target nucleic acid molecule. In some embodiments, a single-stranded consensus sequence (SSCS) is generated by comparing one or more sequence reads generated from the first and / or second strands with each other. In some embodiments, a double-stranded consensus sequence is obtained by comparing the SSCS of both strands derived from the same original double-stranded molecule.
[0101] In some embodiments, identifying eccDNA using multiple sequence reads from double-stranded sequencing or identifying eccDNA includes fragment length analysis. In some embodiments, the fragment length analysis is achieved by reviewing the genome alignment of the double-stranded consensus reads that support the variant call, which is visualized in Integrated Genomics Viewer (IGV) or a similar genome browser. In some embodiments, reviewing the genome alignment of the double-stranded consensus reads that support the variant call is performed manually or using software.
[0102] In some embodiments, analysis is performed to assess the presence of DA junctions. Without wishing to be bound by theory, junctions where the end of an allele merges with the initiation site (referred to herein as "ABCD" alleles, and subsequently referred to as "DA" junctions) are expected features of circular DNA. To identify DA junctions, the alternative allele sequences of apparent insertions or structural variants (identified by variant calling) can be written in any format, such as FASTA format, and used as input to any pattern matching or sequence alignment algorithm, such as the matchPattern function in Biostrings (a Bioconductor package). The algorithm is used to find the optimal single and / or best match or alignment between the alternative allele sequence and the associated reference sequence and report the reference genomic coordinates of the preferred match. In some embodiments, the software allows for approximately one mismatch every 50 base pairs. In some embodiments, the algorithm allows for up to one mismatch every 50 base pairs. In some embodiments, the algorithm allows for at least one mismatch every 50 base pairs. In some embodiments, the algorithm allows for one mismatch approximately every 50 base pairs. The genomic coordinates of the single or best alignment of the alternative allele sequence with the associated reference sequence are compared to the start coordinate of the apparent insertion or apparent structural variant call. If the coordinates of the variant call and the best alignment coordinate of the alternative allele to the reference genome are identical or nearly identical, a DA junction is determined to exist. In some embodiments, the existence of a DA junction can be further confirmed by inspection of auxiliary alignments and / or BLAT searches of soft-clipped sequences.
[0103] In some embodiments, methods and reagents are used to enrich target nucleic acid material, e.g., to limit detection of eccDNA molecules to one or more genomic regions or loci of interest. For example, in some embodiments, the error-corrected sequences obtained in (b) are specific to a single genomic region. In some embodiments, the error-corrected sequences are specific to 2, 3, 4, 5, or more distinct genomic regions. In some embodiments, the error-corrected sequences obtained in (b) are specific to about 1 to about 30 distinct genomic loci. In some embodiments, the error-corrected sequences are specific to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more distinct loci. In some embodiments, the error-corrected sequences obtained in (b) are specific to the whole exome.
[0104] In some embodiments, according to aspects of the present technology, sequencing reads generated from the double-stranded sequencing step discussed herein can be further filtered to remove reads from DNA-damaging molecules (e.g., damaged after tissue or blood collection). In some embodiments, DNA-damaging molecules (e.g., damaged after tissue or blood collection) can be removed before double-stranded sequencing. For example, DNA repair enzymes such as uracil-DNA glycosylase (UDG), formamidopyrimidine DNA glycosylase (FPG), and 8-oxoguanine DNA glycosylase (OGG1) can be used to repair DNA damage (e.g., in vitro DNA damage). These DNA repair enzymes are glycosylases that remove damaged bases from DNA. For example, UDG removes uracil generated by cytosine deamination (resulting from spontaneous hydrolysis of cytosine), and FPG removes 8-oxoguanine (e.g., the most common DNA damage caused by reactive oxygen species). FPG also has leasing activity and can generate single-base gaps at abasic sites. Such abasic sites subsequently fail to amplify by PCR, for example, because the polymerase cannot replicate the template. In some embodiments, single-stranded DNA gaps formed by the leasing activity may prevent complete amplification of the strand during PCR. Therefore, by using such DNA damage repair enzymes, damaged DNA that does not contain true mutations but may cause artifactual / erroneous mutation calls after sequencing and double-stranded sequence analysis can be effectively removed.
[0105] In some embodiments, the glycosylase (FPG, UDG, OGG1, etc.) is used to linearize circular DNA molecules, allowing the biological sample to contain cell-free DNA, allowing its removal by exonuclease treatment, thereby providing a means for detecting DNA damage in eccDNA.
[0106] V. Method The present disclosure provides methods for identifying at least one eccDNA in a sample.
[0107] In some embodiments, a method for identifying at least one extrachromosomal circular DNA (eccDNA) in a sample containing double-stranded DNA includes performing or having performed double-stranded sequencing on the sample, as described herein. In some embodiments, the double-stranded sequencing includes ligating an adapter to an end of the double-stranded DNA, as described herein. In some embodiments, at least one adapter includes a nucleotide sequence that labels a strand of the double-stranded DNA, such that the strand of the double-stranded DNA has a nucleotide sequence that is distinct from its complementary strand. In some embodiments, the double-stranded sequencing includes amplifying the strand of the double-stranded DNA using the ligated adapter to generate at least a first-strand amplicon and a second-strand amplicon. In some embodiments, the double-stranded sequencing includes sequencing at least the first-strand amplicon and the second-strand amplicon to generate a plurality of sequence reads, including a first-strand sequence read and a second-strand sequence read. In some embodiments, the double-stranded sequencing comprises generating an error-corrected sequence read by comparing the first-strand sequence read with the second-strand sequence read and excluding mismatched nucleotide positions.
[0108] In some embodiments, a method for identifying at least one extrachromosomal circular DNA (eccDNA) in a sample containing double-stranded DNA includes identifying or having identified the eccDNA using multiple sequence reads from the double-stranded sequencing.
[0109] In some embodiments, a method for identifying at least one extrachromosomal circular DNA (eccDNA) in a sample comprising double-stranded DNA comprises: performing or causing to be performed double-stranded sequencing on the sample; the double-stranded sequencing comprising: ligating adapters to ends of the double-stranded DNA, wherein at least one adapter comprises a nucleotide sequence that tags a strand of the double-stranded DNA such that the strand of the double-stranded DNA has a nucleotide sequence that is distinct from its complementary strand; amplifying the strand of the double-stranded DNA using the ligated adapters to generate at least first-strand amplicons and second-strand amplicons; and sequencing at least the first-strand amplicons and the second-strand amplicons to generate a plurality of sequence reads comprising first-strand sequence reads and second-strand sequence reads; and identifying or causing to be identified the eccDNA using the plurality of sequence reads from the double-stranded sequencing.
[0110] In some embodiments, a method for identifying at least one extrachromosomal circular DNA (eccDNA) in a sample comprising double-stranded DNA includes the steps of performing or causing double-stranded sequencing on the sample, ligating adapters to ends of the double-stranded DNA, wherein at least one adapter comprises a nucleotide sequence that tags a strand of the double-stranded DNA such that the strand of the double-stranded DNA has a nucleotide sequence that is distinct from its complementary strand, and amplifying the strands of the double-stranded DNA using the ligated adapters to identify at least a first strand of the double-stranded DNA. performing or causing to be performed double-stranded sequencing, the double-stranded sequencing comprising: generating an amplicon and a second-strand amplicon; sequencing at least the first-strand amplicon and the second-strand amplicon to generate a plurality of sequence reads including first-strand sequence reads and second-strand sequence reads; and generating error-corrected sequence reads by comparing the first-strand sequence reads with the second-strand sequence reads and excluding mismatched nucleotide positions; and identifying or causing to be identified the eccDNA using the plurality of sequence reads from the double-stranded sequencing.
[0111] In some embodiments, a method for identifying at least one extrachromosomal circular DNA (eccDNA) in a sample comprising double-stranded DNA comprises: performing or causing double-stranded sequencing on the sample; ligating adapters to ends of the double-stranded DNA, wherein at least one adapter comprises a nucleotide sequence that tags a strand of the double-stranded DNA such that the strand of the double-stranded DNA has a nucleotide sequence that is distinct from its complementary strand; amplifying the strand of the double-stranded DNA using the ligated adapters to generate at least first-strand amplicons and second-strand amplicons; sequencing at least the first-strand amplicons and the second-strand amplicons to generate a plurality of sequence reads, the plurality of sequence reads comprising first-strand sequence reads and second-strand sequence reads; and generating error-corrected sequence reads by comparing the first-strand sequence reads and the second-strand sequence reads and excluding mismatched nucleotide positions. and identifying or causing to be performed the eccDNA using the plurality of sequence reads of the double-stranded sequencing, wherein the step of identifying or causing to be performed the eccDNA comprises the steps of: identifying a subset of sequence reads from the plurality of sequence reads, each of which independently contains a reference allele junction (e.g., DA); and distinguishing, from the subset of sequence reads, hypothesized eccDNA sequence reads from chromosomal tandem duplication sequence reads, wherein the step of distinguishing the hypothesized eccDNA sequence reads comprises the steps of selecting a subset of DA junction-containing consensus sequencing reads whose allele length is less than a threshold apparent insert size; estimating the fragment size of each read pair in the subset and comparing it one by one with the threshold apparent insert size, or comparing the distribution of estimated fragment sizes with the distribution of insert sizes in the entire library; and comparing the estimated fragment size of any consensus read pair to the allele.and determining the amount of eccDNA according to the differentiated hypothesized eccDNA sequence reads.
[0112] In some embodiments, identifying or having identified the eccDNA using the plurality of sequence reads from double-stranded sequencing includes identifying a subset of sequence reads from the plurality of sequence reads, wherein each sequence read of the subset of sequence reads independently comprises a reference allele junction. As used herein, the terms "reference allele junction," "junction," "DA junction," "BA," and "BA junction" can be used interchangeably and refer to a nucleic acid sequence in which a nucleic acid sequence at the end of a reference allele is linked to a nucleic acid sequence at the start of a reference allele.
[0113] In some embodiments, identifying or causing to be identified the eccDNA using the plurality of sequence reads from double-stranded sequencing comprises distinguishing, from the subset of sequence reads, hypothesized eccDNA sequence reads and chromosomal tandem duplication sequence reads. In some embodiments, identifying or causing to be identified the eccDNA using the plurality of sequence reads from double-stranded sequencing comprises determining the amount of eccDNA according to the distinguished hypothesized eccDNA sequence reads.
[0114] In some embodiments, identifying or having identified the eccDNA using the plurality of sequence reads from double-stranded sequencing includes identifying a subset of sequence reads from the plurality of sequence reads, each of which independently includes a reference allele junction; distinguishing hypothesized eccDNA sequence reads and chromosomal tandem duplication sequence reads from the subset of sequence reads; and determining the amount of eccDNA according to the distinguished hypothesized eccDNA sequence reads.
[0115] 1. A method for identifying at least one extrachromosomal circular DNA (eccDNA) in a sample comprising double-stranded DNA, the method comprising: performing or causing double-strand sequencing on the sample; ligating adapters to ends of the double-stranded DNA, at least one of the adapters comprising a nucleotide sequence that tags a strand of the double-stranded DNA such that the strand of the double-stranded DNA has a nucleotide sequence that is distinct from its complementary strand; amplifying the strand of the double-stranded DNA using the ligated adapters to generate at least a first-strand amplicon and a second-strand amplicon; sequencing at least the first-strand amplicon and the second-strand amplicon to generate a plurality of sequence reads, the plurality of sequence reads comprising a first-strand sequence read and a second-strand sequence read; and generating error-corrected sequence reads by comparing the first-strand sequence reads with the second-strand sequence reads and excluding mismatched nucleotide positions. performing or causing to be performed double-stranded sequencing, the step including: identifying or causing to be performed the eccDNA using the plurality of sequence reads of the double-stranded sequencing, wherein the step of identifying or causing to be identified the eccDNA includes the steps of: identifying, from the plurality of sequence reads, a subset of sequence reads that each independently include a reference allele junction; distinguishing, from the subset of sequence reads, hypothesized eccDNA sequence reads and chromosomal tandem duplication sequence reads; and determining the amount of eccDNA according to the distinguished hypothesized eccDNA sequence reads.
[0116] In some embodiments, a method for identifying extrachromosomal circular DNA (eccDNA) includes obtaining a plurality of sequence reads of double-stranded DNA sequenced using double-stranded sequencing; identifying a subset of sequence reads from the plurality of sequence reads, each of the subsets independently including a reference allele junction; distinguishing hypothesized eccDNA sequence reads from chromosomal tandem duplication sequence reads from the subset of sequence reads; and identifying eccDNA according to the distinguished hypothesized eccDNA sequence reads.
[0117] In some embodiments, a method for identifying extrachromosomal circular DNA (eccDNA) comprises obtaining a plurality of sequence reads of double-stranded DNA sequenced using double-stranded sequencing; identifying a subset of sequence reads from the plurality of sequence reads, each of which independently contains a reference allele junction; and distinguishing hypothesized eccDNA sequence reads from the subset of sequence reads into chromosomal tandem duplication sequence reads, wherein distinguishing the hypothesized eccDNA sequence reads comprises selecting a subset of DA junction-containing consensus sequencing reads whose allele length is less than a threshold apparent insert size; estimating the fragment size of each read pair in the subset and comparing it one by one with the threshold apparent insert size, or comparing the distribution of estimated fragment sizes with the distribution of insert sizes in the entire library; and comparing the estimated fragment size of any consensus read pair with the allele; and identifying eccDNA according to the distinguished hypothesized eccDNA sequence reads.
[0118] In some embodiments, the at least one adapter sequence is or contains at least one non-standard nucleotide, in some embodiments, the non-standard nucleotide is selected from uracil, methylated nucleotides, RNA nucleotides, ribose nucleotides, 8-oxoguanine, biotinylated nucleotides, dethiobiotin nucleotides, thiol-modified nucleotides, acrydite-modified nucleotides, iso-dC, iso-dG, 2'-O-methyl nucleotides, inosine nucleotides, locked nucleic acids, peptide nucleic acids, 5-methyl dC, 5-bromodeoxyuridine, 2,6-diaminopurine, 2-aminopurine nucleotides, abasic nucleotides, 5-nitroindole nucleotides, adenylated nucleotides, azido nucleotides, digoxigenin nucleotides, I-linkers, 5'-hexynyl-modified nucleotides, 5-octadiynyl dU, photocleavable spacers, non-photocleavable spacers, click chemistry-enabled modified nucleotides, fluorescent dyes, biotin, furan, BrdU, Fluoro-dU, and any combination thereof.
[0119] In some embodiments, the reference allele comprises a nucleic acid sequence having the formula ABCD. In some embodiments, the reference allele junction comprises a nucleic acid sequence comprising a nucleic acid sequence at the end of the reference allele and a nucleic acid sequence at the beginning of the reference allele. In some embodiments, the reference allele junction comprises a nucleic acid sequence having the formula DA. In some embodiments, the reference allele junction comprises a nucleic acid sequence having the formula DA, where B and C are absent. In some embodiments, the reference allele junction comprises a nucleic acid sequence having the formula DA, where B or C is absent. In some embodiments, the reference allele junction comprises a nucleic acid sequence having the formula DA, where D comprises the nucleic acid sequence at the end of the reference allele and A comprises the nucleic acid sequence at the beginning of the reference allele. In some embodiments, the nucleic acid sequence DA comprises, in the 5' to 3' direction, nucleic acid sequence D operably linked to nucleic acid sequence A. In some embodiments, nucleic acid sequence D is located downstream of nucleic acid sequence A at the reference genomic locus of the reference allele. In some embodiments, nucleic acid sequence A is located upstream of nucleic acid sequence D at the reference genomic locus of the reference allele.
[0120] In some embodiments, the reference allele junction comprises at least 1 base pair (bp) of nucleic acid sequence. In some embodiments, the reference allele junction comprises about 1 bp, 2 bp, 3 bp, 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 15 bp, 20 bp, about 30 bp, about 40 bp, about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 160 bp, about 170 bp, about 180 bp, about 190 bp, about 200 bp, about 210 bp, about 220 bp, about 230 bp, about 240 bp, about 250 bp, about 260 bp, about 270 bp, about 280 bp, about 290 bp, about 310 bp, about 320 bp, about 330 bp, about 340 bp, about 350 bp, about 360 bp, about 370 bp, about 380 bp, about 390 bp, about 400 bp, about 410 bp, about 420 bp, about 430 bp, about 440 bp, about 450 bp, about 460 bp, about 470 bp, about 480 bp, about 490 bp, about 500 bp, about 510 bp, about 520 bp, about 530 bp, about 540 bp, about 550 bp, 00bp, about 210bp, about 220bp, about 230bp, about 240bp, about 250bp, about 260bp, about 270bp, about 280bp, about 290bp, about 300bp, about 310bp, about 320bp, about 330bp, about 340bp, about 350bp, about 360bp, about 370bp, about 380bp, about 390bp, about 400bp, about 410bp, about 420bp, about 430bp, about 440bp, about 450bp, about 460bp, about 470bp, about 480bp, approximately 490bp, approximately 500bp, approximately 510bp, approximately 520bp, approximately 530bp, approximately 540bp, approximately 550bp, approximately 560bp, approximately 570bp, approximately 580bp, approximately 590bp, approximately 600bp, approximately 610bp, Approximately 620bp, approximately 630bp, approximately 640bp, approximately 650bp, approximately 660bp, approximately 670bp, approximately 680bp, approximately 690bp, approximately 700bp, approximately 710bp, approximately 720bp, approximately 730bp, approximately 740bp, approximately 750bp, These include nucleic acid sequences of about 760 bp, about 770 bp, about 780 bp, about 790 bp, about 800 bp, about 810 bp, about 820 bp, about 830 bp, about 840 bp, about 850 bp, about 860 bp, about 870 bp, about 880 bp, about 890 bp, about 900 bp, about 910 bp, about 920 bp, about 930 bp, about 940 bp, about 950 bp, about 960 bp, about 970 bp, about 980 bp, about 990 bp, about 1000 bp, or more in length. In some embodiments, the reference allele junction is at least 1 bp, 2 bp, 3 bp, 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 15 bp, 20 bp, at least 30 bp, at least 40 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp,At least 110bp, at least 120bp, at least 130bp, at least 140bp, at least 150bp, at least 160bp, at least 170bp, at least 180bp, at least 190bp, at least 200bp, at least 210bp, at least 220bp, at least 230bp, at least 240bp, at least 250bp, at least 260bp, at least 270bp, at least 280bp, at least 290bp, at least 300bp, at least 310bp, at least 320bp, at least 330bp, at least 340bp, at least 350bp, at least 360bp, at least 370bp, at least 380bp, at least 390bp, at least 400bp, at least 410bp, at least 420bp, at least 430bp, at least 440bp, at least 450bp, at least 460bp, at least 470bp, at least 480bp, at least 490bp, at least 500bp, at least 510bp, at least 520bp, at least 530bp, at least 540bp, at least 550bp, At least 560bp, at least 570bp, at least 580bp, at least 590bp, at least 600bp, at least 610bp, at least 620bp, at least 630bp, at least 640bp, at least 650bp, at least 660bp, at least 670bp, at least 680bp, at least 690bp, at least 700bp, at least 710bp, at least 720bp, at least 730bp, at least 740bp, at least 750bp, at least 760bp, at least 770bp, at least 7 80bp, at least 790bp, at least 800bp, at least 810bp, at least 820bp, at least 830bp, at least 840bp, at least 850bp, at least 860bp, at least 870bp, at least 880bp, at least 890bp, at least 900bp, at least 910bp, at least 920bp, at least 930bp, at least 940bp, at least 950bp, at least 960bp, at least 970bp, at least 980bp, at least 990bp, at least 1000bp,In some embodiments, the reference allele junction comprises a nucleic acid sequence of 1 bp, 2 bp, 3 bp, 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 15 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, 110 bp, 120 bp, 130 bp, 140 bp, 150 bp, 160 bp, 170 bp, 180 bp, 190 bp, 210 bp, 220 bp, 230 bp, 240 bp, 250 bp, 260 bp, 270 bp, 280 bp, 290 bp, 300 bp, 310 bp, 320 bp, 330 bp, 340 bp, 350 bp, 360 bp, 370 bp, 380 bp, 390 bp, 400 bp, 410 bp, 420 bp, 430 bp, 440 bp, 450 bp, 460 bp, 470 bp, 480 bp, 490 bp, 500 bp, 510 bp, 520 bp, 530 bp, 540 bp, 550 bp, 560 bp, 570 bp, 580 bp, 590 bp, 610 bp, 620 bp, 630 bp, 640 bp, 650 bp, 660 bp, 670 0bp, 200bp, 210bp, 220bp, 230bp, 240bp, 250bp, 260bp, 270bp, 280bp, 290bp, 300bp, 310bp, 320bp, 330 bp, 340bp, 350bp, 360bp, 370bp, 380bp, 390bp, 400bp, 410bp, 420bp, 430bp, 440bp, 450bp, 460bp, 470b p, 480bp, 490bp, 500bp, 510bp, 520bp, 530bp, 540bp, 550bp, 560bp, 570bp, 580bp, 590bp, 600bp, 610bp , 620bp, 630bp, 640bp, 650bp, 660bp, 670bp, 680bp, 690bp, 700bp, 710bp, 720bp, 730bp, 740bp, 750bp, Nucleic acid sequences of 760bp, 770bp, 780bp, 790bp, 800bp, 810bp, 820bp, 830bp, 840bp, 850bp, 860bp, 870bp, 880bp, 890bp, 900bp, 910bp, 920bp, 930bp, 940bp, 950bp, 960bp, 970bp, 980bp, 990bp, 1000bp, or longer in length.
[0121] In some embodiments, the reference allele junction is at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, at least 15%, at least 16%, at least 17%, at least 18%, at least 19%, at least 20%, at least 21%, at least 22%, at least 23%, at least 24%, at least 25%, at least 26%, at least 27%, at least 28%, at least 30%, at least 31%, at least 32%, at least 33%, at least 34%, at least 35%, at least 36%, at least 37%, at least 38%, at least 39%, at least 40%, at least 41%, at least 42%, at least 43%, at least 44%, at least 45%, at least 46%, at least 47%, at least 48%, at least 49%, at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 8 %, at least 24%, at least 25%, at least 26%, at least 27%, at least 28%, at least 29%, at least 30%, at least 31%, at least 32%, at least 33%, at least 34%, at least 35%, at least 36%, at least 37%, at least 38%, at least 39%, at least 40%, at least 41%, at least 42%, at least 43%, at least 44%, at least 45%, at least 46%, at least 47%, at least 48%, at least 49 %, at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75 %, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the nucleic acid sequence.In some embodiments, the reference allele junction is about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 11%, about 12%, about 13%, about 14%, about 15%, about 16%, about 17%, about 18%, about 19%, about 20%, about 21%, about 22%, about 23%, about 24%, about 25%, about 26%, about 27%, about 28%, about 29%, about 30%, about 31%, about 32%, about 33%, about 34%, about 35%, about 36%, about 37%, about 38%, about 39%, about 40%, about 41%, about 42%, about 43%, about 44%, about 45%, about 46%, about 47%, about 48%, about 49%, about 50%, about 51%, about 52%, about 53%, about 54%, about 55%, about 56%, about 57%, about 58%, about 59%, about 60%, about 61%, about 62%, about 63%, about 64%, about 65%, about 66%, about 67%, about 68%, about 69%, about 70%, about 71%, about 72%, about 73%, about 74%, about 75%, about 76%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, about 100%, about 101%, about 2%, approximately 23%, approximately 24%, approximately 25%, approximately 26%, approximately 27%, approximately 28%, approximately 29%, approximately 30%, approximately 31%, approximately 32%, approximately 33%, approximately 34%, approximately 35%, approximately 36%, approximately 37%, approximately 38%, approximately 39%, approximately 40%, approximately 41%, approximately 42%, approximately 43%, approximately 44%, approximately 45%, approximately 46%, approximately 47%, approximately 48%, approximately 49% ,approximately 50%,approximately 51%,approximately 52%,approximately 53%,approximately 54%,approximately 55%,approximately 56%,approximately 57%,approximately 58%,approximately 59%,approximately 60%,approximately 61%,approximately 62%,approximately 63%,approximately 64%,approximately 65%,approximately 66%,approximately 67%,approximately 68%,approximately 69%,approximately 70%,approximately 71%,approximately 72%,approximately 73%,approximately 74%,approximately 75%,approximately 76%,approximately The nucleic acid sequence may have a length of 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100%. In some embodiments, the reference allele junction is 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 1 9%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the length of the nucleic acid sequence.
[0122] In some embodiments, the nucleic acid sequence DA has at least 1 base pair (bp). In some embodiments, the nucleic acid sequence DA has about 1 bp, 2 bp, 3 bp, 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 15 bp, 20 bp, about 30 bp, about 40 bp, about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 160 bp, about 170 bp, about 180 bp, about 190 bp, or about 200 bp. , about 210bp, about 220bp, about 230bp, about 240bp, about 250bp, about 260bp, about 270bp, about 280bp, about 290bp, about 300bp, about 310bp, about 320bp, about 330bp, about 340 bp, about 350bp, about 360bp, about 370bp, about 380bp, about 390bp, about 400bp, about 410bp, about 420bp, about 430bp, about 440bp, about 450bp, about 460bp, about 470bp, about 4 80bp, approximately 490bp, approximately 500bp, approximately 510bp, approximately 520bp, approximately 530bp, approximately 540bp, approximately 550bp, approximately 560bp, approximately 570bp, approximately 580bp, approximately 590bp, approximately 600bp, approximately 610bp, Approximately 620bp, approximately 630bp, approximately 640bp, approximately 650bp, approximately 660bp, approximately 670bp, approximately 680bp, approximately 690bp, approximately 700bp, approximately 710bp, approximately 720bp, approximately 730bp, approximately 740bp, approximately 750b p, about 760 bp, about 770 bp, about 780 bp, about 790 bp, about 800 bp, about 810 bp, about 820 bp, about 830 bp, about 840 bp, about 850 bp, about 860 bp, about 870 bp, about 880 bp, about 890 bp, about 900 bp, about 910 bp, about 920 bp, about 930 bp, about 940 bp, about 950 bp, about 960 bp, about 970 bp, about 980 bp, about 990 bp, about 1000 bp, or longer. In some embodiments, the nucleic acid sequence DA is at least 1 bp, 2 bp, 3 bp, 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 15 bp, 20 bp, at least 30 bp, at least 40 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, at least 110 bp, at least 120 bp,At least 130bp, at least 140bp, at least 150bp, at least 160bp, at least 170bp, at least 180bp, at least 190bp, at least 200bp, at least 210bp, at least 220bp, at least 230bp, at least 240bp, at least 250bp, at least 260bp, at least 270bp, at least 280bp, at least 290bp, at least 300bp, at least 310bp, at least 320bp, at least 330bp, at least 340bp, at least Also 350bp, at least 360bp, at least 370bp, at least 380bp, at least 390bp, at least 400bp, at least 410bp, at least 420bp, at least 430bp, at least 440bp, at least 450bp, at least 460bp, at least 470bp, at least 480bp, at least 490bp, at least 500bp, at least 510bp, at least 520bp, at least 530bp, at least 540bp, at least 550bp, at least 560bp, at least 570 bp, at least 580bp, at least 590bp, at least 600bp, at least 610bp, at least 620bp, at least 630bp, at least 640bp, at least 650bp, at least 660bp, at least 670bp, at least 680bp, at least 690bp, at least 700bp, at least 710bp, at least 720bp, at least 730bp, at least 740bp, at least 750bp, at least 760bp, at least 770bp, at least 780bp, at least 790bp, In some embodiments, the nucleic acid has a length of at least 800 bp, at least 810 bp, at least 820 bp, at least 830 bp, at least 840 bp, at least 850 bp, at least 860 bp, at least 870 bp, at least 880 bp, at least 890 bp, at least 900 bp, at least 910 bp, at least 920 bp, at least 930 bp, at least 940 bp, at least 950 bp, at least 960 bp, at least 970 bp, at least 980 bp, at least 990 bp, at least 1000 bp, or more.The nucleic acid sequence DA can be 1 bp, 2 bp, 3 bp, 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 15 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, 110 bp, 120 bp, 130 bp, 140 bp, 150 bp, 160 bp, 170 bp, 180 bp, 190 bp, 200 bp, 210 bp p, 220bp, 230bp, 240bp, 250bp, 260bp, 270bp, 280bp, 290bp, 300bp, 310bp, 320bp, 330bp, 340bp, 35 0bp, 360bp, 370bp, 380bp, 390bp, 400bp, 410bp, 420bp, 430bp, 440bp, 450bp, 460bp, 470bp, 480bp, 490bp, 500bp, 510bp, 520bp, 530bp, 540bp, 550bp, 560bp, 570bp, 580bp, 590bp, 600bp, 610bp, 620b p, 630bp, 640bp, 650bp, 660bp, 670bp, 680bp, 690bp, 700bp, 710bp, 720bp, 730bp, 740bp, 750bp, 76 The length may be 0 bp, 770 bp, 780 bp, 790 bp, 800 bp, 810 bp, 820 bp, 830 bp, 840 bp, 850 bp, 860 bp, 870 bp, 880 bp, 890 bp, 900 bp, 910 bp, 920 bp, 930 bp, 940 bp, 950 bp, 960 bp, 970 bp, 980 bp, 990 bp, 1000 bp, or more.
[0123] In some embodiments, the nucleic acid sequence DA is at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, at least 15%, at least 16%, at least 17%, at least 18%, at least 19%, at least 20%, at least 21%, at least 22%, at least 23%, at least 24%, at least 25%, at least 26%, at least 27%, at least 28%, at least 30%, at least 31%, at least 32%, at least 33%, at least 34%, at least 35%, at least 36%, at least 37%, at least 38%, at least 39%, at least 40%, at least 41%, at least 42%, at least 43%, at least 44%, at least 45%, at least 46%, at least 47%, at least 48%, at least 49%, at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86 at least 24%, at least 25%, at least 26%, at least 27%, at least 28%, at least 29%, at least 30%, at least 31%, at least 32%, at least 33%, at least 34%, at least 35%, at least 36%, at least 37%, at least 38%, at least 39%, at least 40%, at least 41%, at least 42%, at least 43%, at least 44%, at least 45%, at least 46%, at least 47%, at least 48%, at least 49 %, at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least or at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length.In some embodiments, the nucleic acid sequence DA is about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 11%, about 12%, about 13%, about 14%, about 15%, about 16%, about 17%, about 18%, about 19%, about 20%, about 21%, about 22%, about 23%, about 24%, about 25%, about 26%, about 27%, about 28%, about 29%, about 30%, about 31%, about 32%, about 33%, about 34%, about 35%, about 36%, about 37%, about 38%, about 39%, about 40%, about 41%, about 42%, about 43%, about 44%, about 45%, about 46%, about 47%, about 48%, about 49%, about 50%, about 51%, about 52%, about 53%, about 54%, about 55%, about 56%, about 57%, about 58%, about 59%, about 60%, about 61%, about 62%, about 63%, about 64%, about 65%, about 66%, about 67%, about 68%, about 69%, about 70%, about 71%, about 72%, about 73%, about 74%, about 75%, about 76%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91% Approximately 23%, approximately 24%, approximately 25%, approximately 26%, approximately 27%, approximately 28%, approximately 29%, approximately 30%, approximately 31%, approximately 32%, approximately 33%, approximately 34%, approximately 35%, approximately 36%, approximately 37%, approximately 38%, approximately 39%, approximately 40%, approximately 41%, approximately 42%, approximately 43%, approximately 44%, approximately 45%, approximately 46%, approximately 47%, approximately 48%, approximately 49% ,approximately 50%,approximately 51%,approximately 52%,approximately 53%,approximately 54%,approximately 55%,approximately 56%,approximately 57%,approximately 58%,approximately 59%,approximately 60%,approximately 61%,approximately 62%,approximately 63%,approximately 64%,approximately 65%,approximately 66%,approximately 67%,approximately 68%,approximately 69%,approximately 70%,approximately 71%,approximately 72%,approximately 73%,approximately 74%,approximately 75%,approximately 76%,approximately 77%,approximately 78%,approximately 79%,approximately 80%,approximately 81%,approximately 82%,approximately 83%,approximately 84%,approximately 85%,approximately 86%,approximately 87%,approximately 88%,approximately 89%,approximately 90%,approximately 91%,approximately 92%,approximately 93%,approximately 94%,approximately 95%,approximately 96%,approximately 97%,approximately 98%,approximately 99%,approximately 100%,approximately 101%,approximately 102%,approximately 103%,approximately 104%,approximately 105%,approximately 106%,approximately 107%,approximately 108%,approximately 109%,approximately 110%,approximately 111%,approx 6%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100% of the length. In some embodiments, the nucleic acid sequence DA is 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 12 having a length of 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.
[0124] In some embodiments, the nucleic acid sequence D has at least 1 base pair (bp). In some embodiments, the nucleic acid sequence D has about 1 bp, 2 bp, 3 bp, 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 15 bp, 20 bp, about 30 bp, about 40 bp, about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 160 bp, about 170 bp, about 180 bp, about 190 bp, about 200 bp, Approximately 210bp, approximately 220bp, approximately 230bp, approximately 240bp, approximately 250bp, approximately 260bp, approximately 270bp, approximately 280bp, approximately 290bp, approximately 300bp, approximately 310bp, approximately 320bp, approximately 330bp, approximately 340b p, about 350bp, about 360bp, about 370bp, about 380bp, about 390bp, about 400bp, about 410bp, about 420bp, about 430bp, about 440bp, about 450bp, about 460bp, about 470bp, about 48 0bp, about 490bp, about 500bp, about 510bp, about 520bp, about 530bp, about 540bp, about 550bp, about 560bp, about 570bp, about 580bp, about 590bp, about 600bp, about 610bp, about 620bp, about 630bp, about 640bp, about 650bp, about 660bp, about 670bp, about 680bp, about 690bp, about 700bp, about 710bp, about 720bp, about 730bp, about 740bp, about 750bp , about 760 bp, about 770 bp, about 780 bp, about 790 bp, about 800 bp, about 810 bp, about 820 bp, about 830 bp, about 840 bp, about 850 bp, about 860 bp, about 870 bp, about 880 bp, about 890 bp, about 900 bp, about 910 bp, about 920 bp, about 930 bp, about 940 bp, about 950 bp, about 960 bp, about 970 bp, about 980 bp, about 990 bp, about 1000 bp, or more. In some embodiments, the nucleic acid sequence D is at least 1 bp, 2 bp, 3 bp, 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 15 bp, 20 bp, at least 30 bp, at least 40 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, at least 110 bp, at least 120 bp, at least 130 bp,At least 140bp, at least 150bp, at least 160bp, at least 170bp, at least 180bp, at least 190bp, at least 200bp, at least 210bp, at least 220bp, at least 230bp, at least 240bp, at least 250bp, at least 260bp, at least 270bp, at least 280bp, at least 290bp, at least 300bp, at least 310bp, at least 320bp, at least 330bp, at least 340bp, at least 350bp, at least 360bp, at least 370bp, at least 380bp, at least 390bp, at least 400bp, at least 410bp, at least 420bp, at least 430bp, at least 440bp, at least 450bp, at least 460bp, at least 470bp, at least 480bp, at least 490bp, at least 500bp, at least 510bp, at least 520bp, at least 530bp, at least 540bp, at least 550bp, at least 560bp, at least 570bp, at least 580bp, at least 590bp, at least 600bp, at least 610bp, at least 620bp, at least 630bp, at least 640bp, at least 650bp, at least 660bp, at least 670bp, at least 680bp, at least 690bp, at least 700bp, at least 710bp, at least 720bp, at least 730bp, at least 740bp, at least 750bp, at least 760bp, at least 770bp, at least 780bp, at least 790bp, In some embodiments, the nucleic acid sequence D has a length of at least 800 bp, at least 810 bp, at least 820 bp, at least 830 bp, at least 840 bp, at least 850 bp, at least 860 bp, at least 870 bp, at least 880 bp, at least 890 bp, at least 900 bp, at least 910 bp, at least 920 bp, at least 930 bp, at least 940 bp, at least 950 bp, at least 960 bp, at least 970 bp, at least 980 bp, at least 990 bp, at least 1000 bp, or more.2bp, 3bp, 4bp, 5bp, 6bp, 7bp, 8bp, 9bp, 10bp, 15bp, 20bp, 30bp, 40bp, 50bp, 60bp, 70bp, 80bp, 90b p, 100bp, 110bp, 120bp, 130bp, 140bp, 150bp, 160bp, 170bp, 180bp, 190bp, 200bp, 210bp, 220bp, 2 30bp, 240bp, 250bp, 260bp, 270bp, 280bp, 290bp, 300bp, 310bp, 320bp, 330bp, 340bp, 350bp, 360 bp, 370bp, 380bp, 390bp, 400bp, 410bp, 420bp, 430bp, 440bp, 450bp, 460bp, 470bp, 480bp, 490bp, 500bp, 510bp, 520bp, 530bp, 540bp, 550bp, 560bp, 570bp, 580bp, 590bp, 600bp, 610bp, 620bp, 63 0bp, 640bp, 650bp, 660bp, 670bp, 680bp, 690bp, 700bp, 710bp, 720bp, 730bp, 740bp, 750bp, 760bp , 770bp, 780bp, 790bp, 800bp, 810bp, 820bp, 830bp, 840bp, 850bp, 860bp, 870bp, 880bp, 890bp, 900bp, 910bp, 920bp, 930bp, 940bp, 950bp, 960bp, 970bp, 980bp, 990bp, 1000bp, or longer.
[0125] In some embodiments, the nucleic acid sequence D is at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, at least 15%, at least 16%, at least 17%, at least 18%, at least 19%, at least 20%, at least 21%, at least 22%, at least 23%, at least 24%, at least 25%, at least 26%, at least 27%, at least 28%, at least 29%, at least 30%, at least 31%, at least 32%, at least 33%, at least 34%, at least 35%, at least 36%, at least 37%, at least 38%, at least 39%, at least 40%, at least 41%, at least 42%, at least 43%, at least 44%, at least 45%, at least 46%, at least 47%, at least 48%, at least 49%, at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, , at least 24%, at least 25%, at least 26%, at least 27%, at least 28%, at least 29%, at least 30%, at least 31%, at least 32%, at least 33%, at least 34%, at least 35%, at least 36%, at least 37%, at least 38%, at least 39%, at least 40%, at least 41%, at least 42%, at least 43%, at least 44%, at least 45%, at least 46%, at least 47%, at least 48%, at least 49%, at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 10 9%, at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least At least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length.In some embodiments, nucleic acid sequence D is about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 11%, about 12%, about 13%, about 14%, about 15%, about 16%, about 17%, about 18%, about 19%, about 20%, about 21%, about 22%, about 23%, about 24%, about 25%, about 26%, about 27%, about 28%, about 29%, about 30%, about 31%, about 32%, about 33%, about 34%, about 35%, about 36%, about 37%, about 38%, about 39%, about 40%, about 41%, about 42%, about 43%, about 44%, about 45%, about 46%, about 47%, about 48%, about 49%, about 50%, about 51%, about 52%, about 53%, about 54%, about 55%, about 56%, about 57%, about 58%, about 59%, about 60%, about 61%, about 62%, about 63%, about 64%, about 65%, about 66%, about 67%, about 68%, about 69%, about 70%, about 71%, about 72%, about 73%, about 74%, about 75%, about 76%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, about 100%, about 101%, about 10 23%, approximately 24%, approximately 25%, approximately 26%, approximately 27%, approximately 28%, approximately 29%, approximately 30%, approximately 31%, approximately 32%, approximately 33%, approximately 34%, approximately 35%, approximately 36%, approximately 37%, approximately 38%, approximately 39%, approximately 40%, approximately 41%, approximately 42%, approximately 43%, approximately 44%, approximately 45%, approximately 46%, approximately 47%, approximately 48%, approximately 49% ,approximately 50%,approximately 51%,approximately 52%,approximately 53%,approximately 54%,approximately 55%,approximately 56%,approximately 57%,approximately 58%,approximately 59%,approximately 60%,approximately 61%,approximately 62%,approximately 63%,approximately 64%,approximately 65%,approximately 66%,approximately 67%,approximately 68%,approximately 69%,approximately 70%,approximately 71%,approximately 72%,approximately 73%,approximately 74%,approximately 75%,approximately 76%,approximately 77%,approximately 78%,approximately 79%,approximately 80%,approximately 81%,approximately 82%,approximately 83%,approximately 84%,approximately 85%,approximately 86%,approximately 87%,approximately 88%,approximately 89%,approximately 90%,approximately 91%,approximately 92%,approximately 93%,approximately 94%,approximately 95%,approximately 96%,approximately 97%,approximately 98%,approximately 99%,approximately 100%,approximately 101%,approximately 102%,approximately 103%,approximately 104%,approximately 105%,approximately 106%,approximately 107%,approximately 108%,approximately 109%,approximately 110%,approximately 111%,approx 6%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100% of the length. In some embodiments, the nucleic acid sequence D is 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122 9%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the length.
[0126] In some embodiments, the nucleic acid sequence A has at least 1 base pair (bp). In some embodiments, the nucleic acid sequence A has about 1 bp, 2 bp, 3 bp, 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 15 bp, 20 bp, about 30 bp, about 40 bp, about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 160 bp, about 170 bp, about 180 bp, about 190 bp, about 200 bp, Approximately 210bp, approximately 220bp, approximately 230bp, approximately 240bp, approximately 250bp, approximately 260bp, approximately 270bp, approximately 280bp, approximately 290bp, approximately 300bp, approximately 310bp, approximately 320bp, approximately 330bp, approximately 340b p, about 350bp, about 360bp, about 370bp, about 380bp, about 390bp, about 400bp, about 410bp, about 420bp, about 430bp, about 440bp, about 450bp, about 460bp, about 470bp, about 48 0bp, about 490bp, about 500bp, about 510bp, about 520bp, about 530bp, about 540bp, about 550bp, about 560bp, about 570bp, about 580bp, about 590bp, about 600bp, about 610bp, about 620bp, about 630bp, about 640bp, about 650bp, about 660bp, about 670bp, about 680bp, about 690bp, about 700bp, about 710bp, about 720bp, about 730bp, about 740bp, about 750bp , about 760 bp, about 770 bp, about 780 bp, about 790 bp, about 800 bp, about 810 bp, about 820 bp, about 830 bp, about 840 bp, about 850 bp, about 860 bp, about 870 bp, about 880 bp, about 890 bp, about 900 bp, about 910 bp, about 920 bp, about 930 bp, about 940 bp, about 950 bp, about 960 bp, about 970 bp, about 980 bp, about 990 bp, about 1000 bp, or more. In some embodiments, the nucleic acid sequence A is at least 1 bp, 2 bp, 3 bp, 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 15 bp, 20 bp, at least 30 bp, at least 40 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, at least 110 bp, at least 120 bp, at least 130 bp,At least 140bp, at least 150bp, at least 160bp, at least 170bp, at least 180bp, at least 190bp, at least 200bp, at least 210bp, at least 220bp, at least 230bp, at least 240bp, at least 250bp, at least 260bp, at least 270bp, at least 280bp, at least 290bp, at least 300bp, at least 310bp, at least 320bp, at least 330bp, at least 340bp, at least 350bp, at least 360bp, at least 370bp, at least 380bp, at least 390bp, at least 400bp, at least 410bp, at least 420bp, at least 430bp, at least 440bp, at least 450bp, at least 460bp, at least 470bp, at least 480bp, at least 490bp, at least 500bp, at least 510bp, at least 520bp, at least 530bp, at least 540bp, at least 550bp, at least 560bp, at least 570bp, at least 580bp, at least 590bp, at least 600bp, at least 610bp, at least 620bp, at least 630bp, at least 640bp, at least 650bp, at least 660bp, at least 670bp, at least 680bp, at least 690bp, at least 700bp, at least 710bp, at least 720bp, at least 730bp, at least 740bp, at least 750bp, at least 760bp, at least 770bp, at least 780bp, at least 790bp, In some embodiments, nucleic acid sequence A has a length of at least 800 bp, at least 810 bp, at least 820 bp, at least 830 bp, at least 840 bp, at least 850 bp, at least 860 bp, at least 870 bp, at least 880 bp, at least 890 bp, at least 900 bp, at least 910 bp, at least 920 bp, at least 930 bp, at least 940 bp, at least 950 bp, at least 960 bp, at least 970 bp, at least 980 bp, at least 990 bp, at least 1000 bp, or more.2bp, 3bp, 4bp, 5bp, 6bp, 7bp, 8bp, 9bp, 10bp, 15bp, 20bp, 30bp, 40bp, 50bp, 60bp, 70bp, 80bp, 90b p, 100bp, 110bp, 120bp, 130bp, 140bp, 150bp, 160bp, 170bp, 180bp, 190bp, 200bp, 210bp, 220bp, 2 30bp, 240bp, 250bp, 260bp, 270bp, 280bp, 290bp, 300bp, 310bp, 320bp, 330bp, 340bp, 350bp, 360 bp, 370bp, 380bp, 390bp, 400bp, 410bp, 420bp, 430bp, 440bp, 450bp, 460bp, 470bp, 480bp, 490bp, 500bp, 510bp, 520bp, 530bp, 540bp, 550bp, 560bp, 570bp, 580bp, 590bp, 600bp, 610bp, 620bp, 63 0bp, 640bp, 650bp, 660bp, 670bp, 680bp, 690bp, 700bp, 710bp, 720bp, 730bp, 740bp, 750bp, 760bp , 770bp, 780bp, 790bp, 800bp, 810bp, 820bp, 830bp, 840bp, 850bp, 860bp, 870bp, 880bp, 890bp, 900bp, 910bp, 920bp, 930bp, 940bp, 950bp, 960bp, 970bp, 980bp, 990bp, 1000bp, or longer.
[0127] In some embodiments, nucleic acid sequence A is at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, at least 15%, at least 16%, at least 17%, at least 18%, at least 19%, at least 20%, at least 21%, at least 22%, at least 23%, at least 24%, at least 25%, at least 26%, at least 27%, at least 28%, at least 29%, at least 30%, at least 31%, at least 32%, at least 33%, at least 34%, at least 35%, at least 36%, at least 37%, at least 38%, at least 39%, at least 40%, at least 41%, at least 42%, at least 43%, at least 44%, at least 45%, at least 46%, at least 47%, at least 48%, at least 49%, at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at , at least 24%, at least 25%, at least 26%, at least 27%, at least 28%, at least 29%, at least 30%, at least 31%, at least 32%, at least 33%, at least 34%, at least 35%, at least 36%, at least 37%, at least 38%, at least 39%, at least 40%, at least 41%, at least 42%, at least 43%, at least 44%, at least 45%, at least 46%, at least 47%, at least 48%, at least 49%, at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 10 9%, at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least At least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length.In some embodiments, nucleic acid sequence A is about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 11%, about 12%, about 13%, about 14%, about 15%, about 16%, about 17%, about 18%, about 19%, about 20%, about 21%, about 22%, about 23%, about 24%, about 25%, about 26%, about 27%, about 28%, about 29%, about 30%, about 31%, about 32%, about 33%, about 34%, about 35%, about 36%, about 37%, about 38%, about 39%, about 40%, about 41%, about 42%, about 43%, about 44%, about 45%, about 46%, about 47%, about 48%, about 49%, about 50%, about 51%, about 52%, about 53%, about 54%, about 55%, about 56%, about 57%, about 58%, about 59%, about 60%, about 61%, about 62%, about 63%, about 64%, about 65%, about 66%, about 67%, about 68%, about 69%, about 70%, about 71%, about 72%, about 73%, about 74%, about 75%, about 76%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91% of the length of the reference allele. 23%, approximately 24%, approximately 25%, approximately 26%, approximately 27%, approximately 28%, approximately 29%, approximately 30%, approximately 31%, approximately 32%, approximately 33%, approximately 34%, approximately 35%, approximately 36%, approximately 37%, approximately 38%, approximately 39%, approximately 40%, approximately 41%, approximately 42%, approximately 43%, approximately 44%, approximately 45%, approximately 46%, approximately 47%, approximately 48%, approximately 49% ,approximately 50%,approximately 51%,approximately 52%,approximately 53%,approximately 54%,approximately 55%,approximately 56%,approximately 57%,approximately 58%,approximately 59%,approximately 60%,approximately 61%,approximately 62%,approximately 63%,approximately 64%,approximately 65%,approximately 66%,approximately 67%,approximately 68%,approximately 69%,approximately 70%,approximately 71%,approximately 72%,approximately 73%,approximately 74%,approximately 75%,approximately 76%,approximately 77%,approximately 78%,approximately 79%,approximately 80%,approximately 81%,approximately 82%,approximately 83%,approximately 84%,approximately 85%,approximately 86%,approximately 87%,approximately 88%,approximately 89%,approximately 90%,approximately 91%,approximately 92%,approximately 93%,approximately 94%,approximately 95%,approximately 96%,approximately 97%,approximately 98%,approximately 99%,approximately 100%,approximately 101%,approximately 102%,approximately 103%,approximately 104%,approximately 105%,approximately 106%,approximately 107%,approximately 108%,approximately 109%,approximately 110%,approximately 111%,approx 6%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100% of the length. In some embodiments, the nucleic acid sequence A is 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122 9%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the length.
[0128] In some embodiments, distinguishing between hypothesized eccDNA sequence reads and chromosomal tandem duplication sequence reads involves comparing the estimated fragment size of the subset of sequence reads to a threshold derived from the insert size distribution of all sequence reads in the library. Insert size is a metric generated when aligning reads / consensus reads to a reference genome. In most sequencing libraries, the majority of read pairs (obtained from paired-end sequencing) are coordinately mapped (R1 maps upstream of R2), and the calculated insert size corresponds to the size of the DNA fragment before adaptor ligation. Therefore, the insert size distribution of a sequencing library closely approximates the fragment size distribution of the original DNA after fragmentation and before adaptor ligation. However, the majority of (consensus) read pairs resulting from molecules containing DA junctions are disconcertingly aligned (R2 maps upstream relative to R1, as shown in Figure 5C). In the case of disconcerting read pairs, the insert size calculated by the alignment software does not accurately represent the length of the original DNA fragment. The fragment length must be estimated by manually or computationally reconstructing the structure and sequence of the original DNA fragment.
[0129] In some embodiments, distinguishing between hypothesized eccDNA sequence reads and chromosomal tandem duplication sequence reads further comprises selecting a subset of DA junction-containing consensus sequencing reads whose allele length is less than a threshold apparent insert size. In some embodiments, distinguishing between hypothesized eccDNA sequence reads and chromosomal tandem duplication sequence reads further comprises estimating the fragment size of each read pair in the subset and comparing it one-by-one to a threshold apparent insert size, or comparing the estimated fragment size distribution to the insert size distribution in the entire library. In some embodiments, distinguishing between hypothesized eccDNA sequence reads and chromosomal tandem duplication sequence reads further comprises selecting a subset of DA junction-containing consensus sequencing reads whose allele length is less than a threshold apparent insert size, and estimating the fragment size of each read pair in the subset and comparing it one-by-one to the threshold apparent insert size, or comparing the estimated fragment size distribution to the insert size distribution in the entire library.
[0130] In some embodiments, distinguishing between hypothesized eccDNA sequence reads and chromosomal tandem duplication sequence reads further comprises comparing the estimated fragment size of any consensus read pair to the allele size of that read pair.
[0131] In some embodiments, distinguishing between hypothesized eccDNA sequence reads and chromosomal tandem duplication sequence reads further includes selecting a subset of DA junction-containing consensus sequencing reads whose allele length is less than a threshold apparent insert size, estimating the fragment size of each read pair in the subset and comparing it one-by-one to the threshold apparent insert size, or comparing the distribution of estimated fragment sizes to the distribution of insert sizes in the entire library, and comparing the estimated fragment size of any consensus read pair to the allele size of that read pair.
[0132] In some embodiments, the threshold apparent insert size is about 20 base pairs (bp). In some embodiments, the threshold apparent insert size is at least 20 base pairs (bp). In some embodiments, the threshold apparent insert size is about 20 bp, about 30 bp, about 40 bp, about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 160 bp, about 170 bp, about 180 bp, about 190 bp, about 200 bp, about 210 bp, about 220 bp, about 230 bp, or about 240 bp. , about 250bp, about 260bp, about 270bp, about 280bp, about 290bp, about 300bp, about 310bp, about 320bp, about 330bp, about 340bp, about 350bp, about 360bp, about 370bp , about 380bp, about 390bp, about 400bp, about 410bp, about 420bp, about 430bp, about 440bp, about 450bp, about 460bp, about 470bp, about 480bp, about 490bp, about 500bp , about 510bp, about 520bp, about 530bp, about 540bp, about 550bp, about 560bp, about 570bp, about 580bp, about 590bp, about 600bp, about 610bp, about 620bp, about 630b p, about 640bp, about 650bp, about 660bp, about 670bp, about 680bp, about 690bp, about 700bp, about 710bp, about 720bp, about 730bp, about 740bp, about 750bp, about 760b p, about 770 bp, about 780 bp, about 790 bp, about 800 bp, about 810 bp, about 820 bp, about 830 bp, about 840 bp, about 850 bp, about 860 bp, about 870 bp, about 880 bp, about 890 bp, about 900 bp, about 910 bp, about 920 bp, about 930 bp, about 940 bp, about 950 bp, about 960 bp, about 970 bp, about 980 bp, about 990 bp, about 1000 bp, or more. In some embodiments, the threshold apparent insert size is at least 20 bp, at least 30 bp, at least 40 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, at least 110 bp, at least 120 bp, at least 130 bp, at least 140 bp, at least 150 bp,at least 160bp, at least 170bp, at least 180bp, at least 190bp, at least 200bp, at least 210bp, at least 220bp, at least 230bp, at least 240bp, at least 250bp, at least 260bp, at least 270bp, at least 280bp, at least 290bp, at least 300bp, at least 310bp, at least 320bp, at least 330bp, at least 340bp, at least 350bp, at least 360bp, at least 370bp, at least 380bp, at least 390bp, at least 400bp, at least 410bp, at least 420bp, at least 430bp, at least 440bp, at least 450bp, at least 460bp, at least 470bp, at least 480bp, at least 490bp, at least 500bp, at least 510bp, at least 520bp, at least 530bp, at least 540bp, at least 550bp, at least 560bp, at least 570bp, at least 580bp, At least 590bp, at least 600bp, at least 610bp, at least 620bp, at least 630bp, at least 640bp, at least 650bp, at least 660bp, at least 670bp, at least 680bp, at least 690bp, at least 700bp, at least 710bp, at least 720bp, at least 730bp, at least 740bp, at least 750bp, at least 760bp, at least 770bp, at least 780bp, at least 790bp, at least 8 In some embodiments, the threshold apparent insert size is 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 810 bp, 820 bp, 930 bp, 940 bp, 950 bp, 960 bp, 970 bp, 980 bp, 990 bp, 1000 bp, or more.70bp, 80bp, 90bp, 100bp, 110bp, 120bp, 130bp, 140bp, 150bp, 160bp, 170bp, 180bp, 19 0bp, 200bp, 210bp, 220bp, 230bp, 240bp, 250bp, 260bp, 270bp, 280bp, 290bp, 300bp, 31 0bp, 320bp, 330bp, 340bp, 350bp, 360bp, 370bp, 380bp, 390bp, 400bp, 410bp, 420bp, 4 30bp, 440bp, 450bp, 460bp, 470bp, 480bp, 490bp, 500bp, 510bp, 520bp, 530bp, 540bp, 5 50bp, 560bp, 570bp, 580bp, 590bp, 600bp, 610bp, 620bp, 630bp, 640bp, 650bp, 660bp, 670bp, 680bp, 690bp, 700bp, 710bp, 720bp, 730bp, 740bp, 750bp, 760bp, 770bp, 780bp, 790bp, 800bp, 810bp, 820bp, 830bp, 840bp, 850bp, 860bp, 870bp, 880bp, 890bp, 900bp, 910bp, 920bp, 930bp, 940bp, 950bp, 960bp, 970bp, 980bp, 990bp, 1000bp, or more.
[0133] In some embodiments, the estimated fragment size is about 20 base pairs (bp). In some embodiments, the threshold apparent insert size is at least 20 base pairs (bp). In some embodiments, the threshold apparent insert size is about 20 bp, about 30 bp, about 40 bp, about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 160 bp, about 170 bp, about 180 bp, about 190 bp, about 200 bp, about 210 bp, about 220 bp, about 230 bp, or about 240 bp. , about 250bp, about 260bp, about 270bp, about 280bp, about 290bp, about 300bp, about 310bp, about 320bp, about 330bp, about 340bp, about 350bp, about 360bp, about 370bp , about 380bp, about 390bp, about 400bp, about 410bp, about 420bp, about 430bp, about 440bp, about 450bp, about 460bp, about 470bp, about 480bp, about 490bp, about 500bp , about 510bp, about 520bp, about 530bp, about 540bp, about 550bp, about 560bp, about 570bp, about 580bp, about 590bp, about 600bp, about 610bp, about 620bp, about 630b p, about 640bp, about 650bp, about 660bp, about 670bp, about 680bp, about 690bp, about 700bp, about 710bp, about 720bp, about 730bp, about 740bp, about 750bp, about 760b p, about 770bp, about 780bp, about 790bp, about 800bp, about 810bp, about 820bp, about 830bp, about 840bp, about 850bp, about 860bp, about 870bp, about 880bp, about 890bp, about 900bp, about 910bp, about 920bp, about 930bp, about 940bp, about 950bp, about 960bp, about 970bp, about 980bp, about 990bp, about 1000bp, or more. In some embodiments, the predicted fragment size is at least 20bp, at least 30bp, at least 40bp, at least 50bp, at least 60bp, at least 70bp, at least 80bp, at least 90bp, at least 100bp, at least 110bp, at least 120bp, at least 130bp, at least 140bp, at least 150bp, at least 160bp,At least 170bp, at least 180bp, at least 190bp, at least 200bp, at least 210bp, at least 220bp, at least 230bp, at least 240bp, at least 250bp, at least 260bp, at least 270bp, at least 280bp, at least 290bp, at least 300bp, at least 310bp, at least 320bp, at least 330bp, at least 340bp, at least 350bp, at least 360bp, at least 370bp, 380bp, at least 390bp, at least 400bp, at least 410bp, at least 420bp, at least 430bp, at least 440bp, at least 450bp, at least 460bp, at least 470bp, at least 480bp, at least 490bp, at least 500bp, at least 510bp, at least 520bp, at least 530bp, at least 540bp, at least 550bp, at least 560bp, at least 570bp, at least 580bp, at least 590bp, at least 600bp, at least 610bp, at least 620bp, at least 630bp, at least 640bp, at least 650bp, at least 660bp, at least 670bp, at least 680bp, at least 690bp, at least 700bp, at least 710bp, at least 720bp, at least 730bp, at least 740bp, at least 750bp, at least 760bp, at least 770bp, at least 780bp, at least 790bp, at least 800bp, at least 810bp, at least 820bp, at least 830bp, at least 840bp, at least 850bp, at least 860bp, at least 870bp, at least 880bp, at least 890bp, at least 900bp, at least 910bp, at least 920bp, at least 930bp, at least 940bp, at least 950bp, at least 960bp, at least 970bp, at least 980bp, at least 990bp, at least 1000bp, at least 1010bp 90bp, at least 600bp, at least 610bp, at least 620bp, at least 630bp, at least 640bp, at least 650bp, at least 660bp, at least 670bp, at least 680bp, at least 690bp, at least 700bp, at least 710bp, at least 720bp, at least 730bp, at least 740bp, at least 750bp, at least 760bp, at least 770bp, at least 780bp, at least 790bp, at least 800bp p, at least 810 bp, at least 820 bp, at least 830 bp, at least 840 bp, at least 850 bp, at least 860 bp, at least 870 bp, at least 880 bp, at least 890 bp, at least 900 bp, at least 910 bp, at least 920 bp, at least 930 bp, at least 940 bp, at least 950 bp, at least 960 bp, at least 970 bp, at least 980 bp, at least 990 bp, at least 1000 bp, or more. In some embodiments, the predicted fragment size is 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp,90bp, 100bp, 110bp, 120bp, 130bp, 140bp, 150bp, 160bp, 170bp, 180bp, 190bp, 200bp , 210bp, 220bp, 230bp, 240bp, 250bp, 260bp, 270bp, 280bp, 290bp, 300bp, 310bp, 320 bp, 330bp, 340bp, 350bp, 360bp, 370bp, 380bp, 390bp, 400bp, 410bp, 420bp, 430bp, 4 40bp, 450bp, 460bp, 470bp, 480bp, 490bp, 500bp, 510bp, 520bp, 530bp, 540bp, 550bp, 560bp, 570bp, 580bp, 590bp, 600bp, 610bp, 620bp, 630bp, 640bp, 650bp, 660bp, 670bp, 680bp, 690bp, 700bp, 710bp, 720bp, 730bp, 740bp, 750bp, 760bp, 770bp, 780bp, 790bp, 800bp, 810bp, 820bp, 830bp, 840bp, 850bp, 860bp, 870bp, 880bp, 890bp, 900bp, 910bp, 920bp, 930bp, 940bp, 950bp, 960bp, 970bp, 980bp, 990bp, 1000bp, or more.
[0134] In some embodiments, the hypothesized eccDNA is identified as one in which the predicted fragment size is equal to or less than the allele length. In some embodiments, the predicted fragment size is about 20 base pairs (bp). In some embodiments, the allele length is at least 20 base pairs (bp). In some embodiments, the allele length is about 20 bp, about 30 bp, about 40 bp, about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 160 bp, about 170 bp, about 180 bp, about 190 bp, about 200 bp, about 210 bp, about 220 bp, about 230 bp, about 240 bp, or about 250 bp. , about 260bp, about 270bp, about 280bp, about 290bp, about 300bp, about 310bp, about 320bp, about 330bp, about 340bp, about 350bp, about 360bp, about 370bp, about 380 bp, about 390bp, about 400bp, about 410bp, about 420bp, about 430bp, about 440bp, about 450bp, about 460bp, about 470bp, about 480bp, about 490bp, about 500bp, about 51 0bp, about 520bp, about 530bp, about 540bp, about 550bp, about 560bp, about 570bp, about 580bp, about 590bp, about 600bp, about 610bp, about 620bp, about 630bp, about 640bp, about 650bp, about 660bp, about 670bp, about 680bp, about 690bp, about 700bp, about 710bp, about 720bp, about 730bp, about 740bp, about 750bp, about 760bp, In some embodiments, the allele length is at least 20 bp, at least 30 bp, at least 40 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, at least 110 bp, at least 120 bp, at least 130 bp,At least 140bp, at least 150bp, at least 160bp, at least 170bp, at least 180bp, at least 190bp, at least 200bp, at least 210bp, at least 220bp, at least 230bp, at least 240bp, at least 250bp, at least 260bp, at least 270bp, at least 280bp, at least 290bp, at least 300bp, at least 310bp, at least 320bp, at least 330bp, at least 340bp, at least 350bp, At least 360bp, at least 370bp, at least 380bp, at least 390bp, at least 400bp, at least 410bp, at least 420bp, at least 430bp, at least 440bp, at least 450bp, at least 460bp, at least 470bp, at least 480bp, at least 490bp, at least 500bp, at least 510bp, at least 520bp, at least 530bp, at least 540bp, at least 550bp, at least 560bp, at least 570bp, At least 580bp, at least 590bp, at least 600bp, at least 610bp, at least 620bp, at least 630bp, at least 640bp, at least 650bp, at least 660bp, at least 670bp, at least 680bp, at least 690bp, at least 700bp, at least 710bp, at least 720bp, at least 730bp, at least 740bp, at least 750bp, at least 760bp, at least 770bp, at least 780bp, at least 790bp, In some embodiments, the allele length is at least 800 bp, at least 810 bp, at least 820 bp, at least 830 bp, at least 840 bp, at least 850 bp, at least 860 bp, at least 870 bp, at least 880 bp, at least 890 bp, at least 900 bp, at least 910 bp, at least 920 bp, at least 930 bp, at least 940 bp, at least 950 bp, at least 960 bp, at least 970 bp, at least 980 bp, at least 990 bp, at least 1000 bp, or more.40bp, 50bp, 60bp, 70bp, 80bp, 90bp, 100bp, 110bp, 120bp, 130bp, 140bp, 150bp, 160bp, 1 70bp, 180bp, 190bp, 200bp, 210bp, 220bp, 230bp, 240bp, 250bp, 260bp, 270bp, 280bp, 290 bp, 300bp, 310bp, 320bp, 330bp, 340bp, 350bp, 360bp, 370bp, 380bp, 390bp, 400bp, 410b p, 420bp, 430bp, 440bp, 450bp, 460bp, 470bp, 480bp, 490bp, 500bp, 510bp, 520bp, 530bp, 540bp, 550bp, 560bp, 570bp, 580bp, 590bp, 600bp, 610bp, 620bp, 630bp, 640bp, 650bp, 660bp, 670bp, 680bp, 690bp, 700bp, 710bp, 720bp, 730bp, 740bp, 750bp, 760bp, 770bp, 780bp, 790bp, 800bp, 810bp, 820bp, 830bp, 840bp, 850bp, 860bp, 870bp, 880bp, 890bp, 900bp, 910bp, 920bp, 930bp, 940bp, 950bp, 960bp, 970bp, 980bp, 990bp, 1000bp, or more.
[0135] The method (as described herein) for identifying at least one extrachromosomal circular DNA (eccDNA) in a sample containing double-stranded DNA further comprises a step of eccDNA enrichment, which may include size selection as described herein and / or exonuclease treatment as described herein.
[0136] The methods described herein are used to assess the clastogenicity of potential clastogenic agents, such as, but not limited to, chemical compounds, physical exposures, biological agents, complex mixtures, and / or environmental exposures. In some embodiments, a method for assessing the clastogenicity of a potential clastogenic agent includes obtaining double-stranded DNA from one or more cells exposed to the clastogenic agent, performing double-stranded sequencing on the double-stranded DNA to obtain a plurality of sequence reads of the double-stranded DNA, identifying a subset of sequence reads from the plurality of sequence reads that each independently include a reference allele junction, distinguishing, from the subset of sequence reads, hypothesized eccDNA sequence reads and chromosomal tandem duplication sequence reads, determining an eccDNA profile according to the distinguished hypothesized eccDNA sequence reads, and assessing the clastogenicity of the potential clastogenic agent according to the determined eccDNA profile. In some embodiments, a method for assessing the clastogenicity of a potential clastogenic agent includes the steps of obtaining double-stranded DNA comprising a hypothesized extrachromosomal circular DNA (eccDNA) from one or more cells exposed to the clastogenic agent, performing double-stranded sequencing on the double-stranded DNA to obtain a plurality of sequence reads of the double-stranded DNA, identifying a subset of sequence reads from the plurality of sequence reads, each of which independently comprises a reference allele junction, distinguishing, from the subset of sequence reads, hypothesized eccDNA sequence reads and chromosomal tandem duplication sequence reads, determining an eccDNA profile according to the distinguished hypothesized eccDNA sequence reads, and assessing the clastogenicity of the potential clastogenic agent according to the determined eccDNA profile.
[0137] In some embodiments, a method for assessing the clastogenicity of a potential clastogenic agent includes obtaining double-stranded DNA comprising a hypothesized extrachromosomal circular DNA (eccDNA) from one or more cells exposed to the clastogenic agent; performing double-stranded sequencing on the double-stranded DNA to obtain a plurality of sequence reads of the double-stranded DNA; identifying, from the plurality of sequence reads, a subset of sequence reads that each independently contain a reference allele junction; and distinguishing, from the subset of sequence reads, hypothesized eccDNA sequence reads from chromosomal tandem duplication sequence reads, wherein distinguishing the hypothesized eccDNA sequence reads comprises selecting a subset of consensus sequencing reads that contain a DA junction and have an allele length less than a threshold apparent insert size. and estimating the fragment size of each read pair in the subset and comparing it one by one with the threshold apparent insert size, or comparing the estimated fragment size distribution with the insert size distribution in the entire library, and comparing the estimated fragment size of any consensus read pair to the allele; determining an eccDNA profile according to the differentiated hypothesized eccDNA sequence reads; and assessing the clastogenicity of the potential clastogenic agent according to the determined eccDNA profile, wherein assessing the clastogenicity of the potential clastogenic agent further comprises comparing the eccDNA profile from one or more cells exposed to the clastogenic agent with a control or untreated sample from the same cohort.
[0138] In some embodiments, a method for assessing clastogenicity of a potential clastogenic agent includes obtaining double-stranded DNA from one or more cells; performing double-stranded sequencing on the double-stranded DNA to obtain a plurality of sequence reads of the double-stranded DNA; identifying, from the plurality of sequence reads, a subset of sequence reads that each independently comprise a reference allele junction; and distinguishing, from the subset of sequence reads, hypothesized eccDNA sequence reads and chromosomal tandem duplication sequence reads, wherein distinguishing the hypothesized eccDNA sequence reads. The method includes the steps of: selecting a subset of DA junction-containing consensus sequencing reads whose allele length is less than a threshold apparent insert size; estimating the fragment size of each read pair in the subset and comparing it one by one with the threshold apparent insert size, or comparing the distribution of estimated fragment sizes with the distribution of insert sizes in the entire library; and comparing the estimated fragment size of any consensus read pair with the allele; and evaluating the clastogenicity of the potential clastogenic agent according to the differentiated assumed eccDNA. In some embodiments, the one or more cells do not contain eccDNA.
[0139] In some embodiments, the method described herein is used to evaluate the genotoxicity of a compound. In some embodiments, the compound is a xenobiotic as described herein, a clastogenic agent as described herein, or any mutagen. In some embodiments, a method for evaluating genotoxicity includes the following steps: obtaining double-stranded DNA containing a hypothesized extrachromosomal circular DNA (eccDNA) from one or more cells exposed to a xenobiotic; performing double-stranded sequencing on the double-stranded DNA to obtain a plurality of sequence reads of the double-stranded DNA; identifying a subset of sequence reads that each independently contain a reference allele junction from the plurality of sequence reads; distinguishing hypothesized eccDNA sequence reads from the subset of sequence reads and chromosomal tandem duplication sequence reads; determining an eccDNA profile according to the distinguished hypothesized eccDNA sequence reads; and evaluating the clastogenicity of the xenobiotic according to the determined eccDNA profile.
[0140] In some embodiments, a method for assessing genotoxicity of a compound or exposure includes: a) assessing clastogenicity, the method comprising the steps of: obtaining double-stranded DNA from one or more cells exposed to a potential genotoxicant; performing double-stranded sequencing on the double-stranded DNA to obtain a plurality of sequence reads of the double-stranded DNA; identifying, from the plurality of sequence reads, a subset of sequence reads that each independently include a reference allele junction; distinguishing, from the subset of sequence reads, hypothesized eccDNA sequence reads and chromosomal tandem duplication sequence reads; and identifying the distinguished hypothesized eccDNA sequence reads. a) determining a profile of eccDNA according to sequence reads, and assessing the clastogenicity of the potential genotoxin according to the determined profile of eccDNA; and b) assessing mutagenicity, comprising obtaining double-stranded DNA from one or more cells exposed to a mutagen, performing double-stranded sequencing on the double-stranded DNA to obtain a plurality of sequence reads of the double-stranded DNA, determining a mutational profile of the double-stranded DNA, and assessing the mutagenicity of the potential genotoxin according to the determined mutational profile of the double-stranded DNA.
[0141] In some embodiments, a method for assessing genotoxicity of a compound or exposure includes: a) assessing clastogenicity, the method comprising: obtaining double-stranded DNA from one or more cells exposed to a potential genotoxicant; performing double-stranded sequencing on the double-stranded DNA to obtain a plurality of sequence reads of the double-stranded DNA; identifying, from the plurality of sequence reads, a subset of sequence reads that each independently contain a reference allele junction; and distinguishing, from the subset of sequence reads, hypothesized eccDNA sequence reads from chromosomal tandem duplication sequence reads, wherein distinguishing the hypothesized eccDNA sequence reads comprises selecting a subset of DA junction-containing consensus sequencing reads whose allele length is less than a threshold insert size; and estimating the fragment size of each read pair in the subset and comparing it to the threshold insert size one by one. a) distinguishing the potential genotoxin from the alleles, or comparing the estimated fragment size distribution with the insert size distribution in the entire library and comparing the estimated fragment sizes of any consensus read pairs with the alleles, determining an eccDNA profile according to the distinguished hypothesized eccDNA sequence reads, and evaluating the clastogenicity of the potential genotoxin according to the determined eccDNA profile; and b) evaluating mutagenicity, the evaluating step comprising obtaining double-stranded DNA from one or more cells exposed to a mutagen, performing double-stranded sequencing on the double-stranded DNA to obtain a plurality of sequence reads of the double-stranded DNA, determining a mutational profile of the double-stranded DNA, and evaluating the mutagenicity of the potential genotoxin according to the determined mutational profile of the double-stranded DNA. In some embodiments, the potential genotoxin is any compound, physical exposure, environmental exposure, biological agent, or other factor capable of damaging DNA.
[0142] In some embodiments, the method described herein is used to assess cancer risk in a sample.In some embodiments, the method for assessing cancer risk in a sample includes the following steps: obtaining double-stranded DNA from the sample, comprising hypothesized extrachromosomal circular DNA (eccDNA); performing double-stranded sequencing on the double-stranded DNA to obtain a plurality of sequence reads of the double-stranded DNA; identifying a subset of sequence reads from the plurality of sequence reads, each of which independently comprises a reference allele junction; distinguishing hypothesized eccDNA sequence reads and chromosomal tandem duplication sequence reads from the subset of sequence reads; determining an eccDNA profile according to the distinguished hypothesized eccDNA sequence reads; and assessing the cancer risk of the sample according to the determined eccDNA profile.
[0143] In some embodiments, a method for assessing cancer risk in a sample includes obtaining double-stranded DNA from the sample, the double-stranded DNA comprising a hypothesized extrachromosomal circular DNA (eccDNA); performing double-stranded sequencing on the double-stranded DNA to obtain a plurality of sequence reads of the double-stranded DNA; identifying, from the plurality of sequence reads, a subset of sequence reads each independently comprising a reference allele junction; and distinguishing, from the subset of sequence reads, hypothesized eccDNA sequence reads and chromosomal tandem duplication sequence reads, wherein distinguishing the hypothesized eccDNA sequence reads comprises determining whether the allele length is greater than or equal to a threshold intercept. The method includes the steps of: selecting a subset of consensus sequencing reads containing DA junctions that are less than the threshold insert size; estimating the fragment size of each read pair in the subset and comparing it to the threshold insert size one by one, or comparing the distribution of estimated fragment sizes with the distribution of insert sizes in the entire library; and comparing the estimated fragment size of any consensus read pair to the allele; determining an eccDNA profile according to the differentiated hypothesized eccDNA sequence reads; and assessing the cancer risk of the sample according to the determined eccDNA profile. In some embodiments, the sample is a cancerous sample or a normal sample. In some embodiments, assessing the cancer risk of the sample based on the eccDNA profile further comprises comparing the eccDNA profile with a known eccDNA profile.
[0144] In some embodiments, the cell, tissue, or organoid is exposed to one or more compounds. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is an animal cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is any animal cell. In some embodiments, the cell is any human cell. In some embodiments, the cell is any cell. In some embodiments, the cell is a muscle cell. In some embodiments, the cell is a nerve cell. In some embodiments, the cell is a blood cell. In some embodiments, the cell is a connective tissue cell. In some embodiments, the cell is an epithelial cell. In some embodiments, the cell is a germ cell. In some embodiments, the cell is an endocrine cell. In some embodiments, the cell is an immune system cell. In some embodiments, the cell is a stem cell. In some embodiments, the cell is a normal cell. In some embodiments, the cell is a cancer cell. In some embodiments, the tissue is an animal tissue. In some embodiments, the tissue is a human tissue. In some embodiments, the tissue is any animal tissue. In some embodiments, the tissue is any human tissue. In some embodiments, the tissue is any tissue. In some embodiments, the tissue is epithelial tissue. In some embodiments, the tissue is connective tissue. In some embodiments, the tissue is muscle tissue. In some embodiments, the tissue is nervous tissue. In some embodiments, the organoid is animal organoid. In some embodiments, the organoid is human organoid. In some embodiments, the organoid is any organoid.
[0145] As described herein, eccDNA is a DNA fragment present outside of linear chromosomes and is often a by-product of DNA breakage and repair. eccDNA can contain both coding and non-coding sequences and is found in both normal and cancer cells. As used herein, "clastogenicity" refers to the ability of a particular substance to induce chromosomal breakage. Without wishing to be bound by any particular theory, chromosomal breakage events may contribute to the formation of eccDNA. Chromosomal breaks, characteristic of chromosomal breakage events, may result in these circular DNAs, adding further complexity to the genome. Both eccDNA and clastogenicity serve as markers of genomic instability and are associated with diseases such as cancer. For example, elevated levels of eccDNA and chromosomal breakage events are frequently observed in malignant tumors, suggesting their potential usefulness as diagnostic or prognostic markers. In some embodiments, a clastogenic agent or potential clastogenic agent is a substance, such as a compound, its mutant, or its derivative, that can induce chromosomal breakage and cause mutations. Non-limiting examples include the chemical compounds ethyl methanesulfonate (EMS), methyl methanesulfonate (MMS), ethylene oxide, acetaldehyde, formaldehyde, benzene, vinyl chloride, cadmium, nickel compounds, chromium (VI) compounds, lead compounds, arsenic compounds, cyclophosphamide, nitrogen mustard, melphalan, chlorambucil, colchicine, 5-fluorouracil, hydroquinone, adriamycin, actinomycin D, camptothecin, etoposide, cisplatin, azathioprine, 6-mercaptopurine, methotrexate, aflatoxin, thalidomide, hydroxyurea, bleomycin, naphthalene, 2-acetylaminofluorene (2-AAF), variants thereof, or derivatives thereof. Non-limiting examples of physical clastogens include X-rays, gamma rays, and ultraviolet radiation. Non-limiting examples of biological agents include certain oncoviruses, such as HPV and Epstein-Barr virus, and bacteria, such as Helicobacter pylori.Non-limiting examples of complex mixtures and environmental exposures include tobacco smoke, polychlorinated biphenyls (PCBs), asbestos, diesel exhaust, crude oil, and pesticides such as atrazine and paraquat.
[0146] In some embodiments, the compound is a foreign substance, hi some embodiments, the foreign substance is selected from environmental pollutants, hydrocarbons, food additives, oil mixtures, pesticides, other foreign substances, synthetic polymers, carcinogens, pharmaceuticals, antioxidants, and any combination thereof.
[0147] In some embodiments, the method comprises exposing one or more cells, tissues, or organoids to a compound and obtaining eccDNA from the one or more cells, tissues, or organoids before performing or allowing double-strand sequencing as described herein.In some embodiments, the method further comprises evaluating the clastogenicity of the compound based on the determined eccDNA profile.In some embodiments, the eccDNA profile comprises any one or any combination of eccDNA quantity, eccDNA frequency, eccDNA quality, eccDNA or its fragment size, eccDNA genomic location, or other eccDNA characteristics.In some embodiments, the eccDNA profile comprises eccDNA quantity.In some embodiments, the eccDNA profile comprises eccDNA frequency.In some embodiments, the eccDNA profile comprises eccDNA quality.In some embodiments, the eccDNA profile comprises eccDNA or its fragment size.In some embodiments, the eccDNA profile comprises eccDNA genomic location.In some embodiments, the eccDNA profile comprises any eccDNA characteristic. In some embodiments, the potential clastogen is a direct clastogen. In some embodiments, the potential clastogen is an indirect clastogen.
[0148] VI. Computer Implementation In some embodiments, the methods disclosed herein are implemented on one or more computers. In various embodiments, the eccDNA detection system 130 shown in FIG. 1A can be embodied as one or more computers. For example, the identification or execution of eccDNA using multiple sequence reads from double-stranded sequencing, and database storage, can be implemented in hardware, software, or a combination of both. In one embodiment, a machine-readable storage medium is provided, the medium containing data storage material encoded with machine-readable data, which, when used with a machine programmed to use the data, can display the execution and results of any of the disclosed datasets and identification models. The data can be used for various purposes, such as patient monitoring and treatment planning. The methods can be implemented in a computer program executed on a programmable computer, including a processor, a data storage system (including volatile and non-volatile memory and / or storage elements), a graphics adapter, a pointing device, a network adapter, at least one input device, and at least one output device. The program code is applied to the input data to perform the functions described above and generate output information, which is applied to one or more output devices in a known manner. The computer may be, for example, a personal computer, microcomputer, or workstation of conventional design.
[0149] Each program may be implemented in a high-level procedural or object-oriented programming language to communicate with a computer system. However, if desired, the program may also be implemented in assembly or machine language. In either case, the language may be a compiled or interpreted language. Each computer program is preferably stored on a general-purpose or special-purpose programmable computer-readable storage medium or device (e.g., ROM or magnetic disk) that, when read by a computer, configures and operates the computer to carry out the procedures described herein. The system may also be considered to be implemented as a computer-readable storage medium configured with a computer program, which causes the computer to operate in a specific, predetermined manner and perform the functions described herein.
[0150] The data and its database can be provided in a variety of media for ease of use. "Media" refers to a product containing the signature pattern information. The database can be recorded on a computer-readable medium, i.e., any medium that can be directly read and accessed by a computer. Such media include, but are not limited to, magnetic storage media such as floppy disks, hard disk storage media, and magnetic tape; optical storage media such as CD-ROMs; electrical storage media such as RAM and ROM; and hybrids of these categories, such as magnetic / optical storage media. Those skilled in the art can readily understand how to create a product containing a record of the database information using any currently known computer-readable medium. "Recording" refers to the process of storing information on a computer-readable medium using any method known in the art. Any convenient data storage structure can be selected based on the means for accessing the stored information. A variety of data processor programs and formats can be used for storage, such as word processing text files, database formats, etc.
[0151] VII. Kit The present disclosure also provides kits for practicing the methods disclosed herein, including methods for detecting a hypothesized eccDNA molecule in a biological sample, methods for treating a disease or other medical condition in a mammalian (e.g., human) subject, and methods for preparing a sequencing library for detecting a hypothesized eccDNA molecule in a biological sample. The kits can include, for example, one or more reagents (e.g., sequencing adapters, exonucleases, ligases) for practicing any of the methods disclosed herein, one or more physical devices (e.g., reaction vessels, columns, supports, or containers) for practicing the methods, and / or instructions (e.g., printed instructions and / or instructions provided in electronic media) for practicing any of the methods disclosed herein.
[0152] VIII. Additional Embodiments A method for detecting candidate extrachromosomal circular DNA (eccDNA) molecules in a biological sample, the method comprising the steps of: (a) providing a sequencing library containing a plurality of double-stranded DNA fragments obtained from the sample; (b) obtaining error-corrected sequences for the double-stranded DNA fragments in the library; (c) detecting possible insertions present in the double-stranded DNA fragments by aligning a plurality of the error-corrected sequences with a reference genome; and (d) detecting candidate eccDNA breakpoints present in one or more of the fragments in which the possible insertions are detected.the candidate eccDNA breakpoint comprises sequence B located upstream of sequence A, wherein (i) sequence A is present upstream of sequence B in the reference genome, (ii) the first nucleotide of sequence A is located Y nucleotides upstream from the last nucleotide of sequence B in the reference genome, and (iii) in the candidate eccDNA breakpoint, the last nucleotide of sequence B is located approximately immediately upstream of the first nucleotide of sequence A; (e) detecting a candidate eccDNA molecule from among fragments containing the candidate eccDNA breakpoint, wherein the candidate eccDNA molecule is a fragment containing a candidate eccDNA breakpoint that is not excluded by any of (i), (ii), or (iii) below: (i) comparing the length of the fragment with a distance of Y nucleotides, wherein if the fragment is determined to be longer than Y nucleotides, it indicates that the fragment is not a candidate eccDNA molecule; and (ii) determining whether the error-corrected sequence of the fragment contains a duplication of any subsequence contained within the region from sequence A to sequence B in the reference genome, wherein if a duplication is detected, it indicates that the fragment is not a candidate eccDNA molecule; and / or (iii) determining whether a sequence located upstream of sequence A in the reference genome is present upstream of sequence B in the error-corrected sequence of the fragment, or whether a sequence located downstream of sequence B in the reference genome is present downstream of sequence A in the error-corrected sequence of the fragment, wherein if it is detected that a sequence located upstream of sequence A in the reference genome is present upstream of sequence B in the error-corrected sequence, or a sequence located downstream of sequence B in the reference genome is present downstream of sequence A in the error-corrected sequence, this indicates that the fragment is not a candidate eccDNA molecule.
[0153] In some embodiments, the error-corrected sequence obtained in (b) is obtained by consensus sequencing.In some embodiments, the consensus sequencing is double-stranded sequencing (DS).In some embodiments, the consensus sequencing is single-stranded consensus sequencing (SSCS) or the combination of DS and SSCS. In some embodiments, the biological sample is selected from the group consisting of sperm samples, semen samples, prostatic fluid samples, testicular biopsy samples, spermatogonial samples, germ cell samples, gamete samples, swab samples, lavage samples, aspirate samples, biopsy samples, tissue samples, tumor samples, precancerous lesion samples, liquid biopsy samples, hyperplasia samples, hypertrophy samples, dysplasia samples, urine samples, cerebrospinal fluid (CSF) samples, other bodily fluid samples, autopsy samples, postmortem samples, surgical samples, model organism samples, plasma samples, serum samples, gastric juice samples, bone marrow samples, stool samples, brushing samples, bile samples, pancreatic juice samples, synovial fluid samples, sputum samples, mucus samples, vitreous humor samples, forensic samples, environmental samples, bacterial samples, fungal samples, mammalian samples, human samples, and diagnostic samples. In some embodiments, the biological sample contains potentially cancerous cells or nucleic acids potentially derived from cancer. In some embodiments, the biological sample contains cell-free DNA. In some embodiments, the biological sample contains cells exposed to a potentially toxic substance. In some embodiments, the potentially toxic agent is a potential clastogen, aneugen, mutagen, and / or teratogen, aneugen, mutagen, and / or teratogen. In some embodiments, the presence and / or characteristics of eccDNA molecules in the sample are used to identify a disease state or physiological condition. In some embodiments, the disease state or physiological condition is selected from the group consisting of cancer, inflammation, autoimmune disease, infectious disease, organ transplant rejection, stem cell transplant rejection, therapeutic cell rejection, therapeutic cell response, immunotherapy response, pregnancy, pre-eclampsia, radiation exposure, sun exposure, drug exposure, and hypersensitivity.In some embodiments, the double-stranded DNA fragments are obtained by enzymatic fragmentation. In some embodiments, the average length of the double-stranded DNA fragments in the library is between about 100 bp and 1000 bp. In some embodiments, the average length of the double-stranded DNA fragments in the library is greater than about 1000 bp, 2000 bp, 3000 bp, 4000 bp, 5000 bp, 6000 bp, 7000 bp, 8000 bp, 9000 bp, or 10,000 bp. In some embodiments, the length of the candidate eccDNA molecules is between about 100 and 1000 nucleotides. In some embodiments, the length of the candidate eccDNA molecules is less than about 500 nucleotides. In some embodiments, the length of the candidate eccDNA molecules is greater than about 1000 bp, 2000 bp, 3000 bp, 4000 bp, 5000 bp, 6000 bp, 7000 bp, 8000 bp, 9000 bp, or 10,000 bp. In some embodiments, the length of the candidate eccDNA molecule is greater than about 100 kb, 200 kb, 300 kb, 400 kb, 500 kb, 600 kb, 700 kb, 800 kb, 900 kb, 1 Mb, 2 Mb, or 3 Mb. In some embodiments, the candidate eccDNA molecule comprises a gene. In some embodiments, the candidate eccDNA molecule comprises an origin of replication. In some embodiments, the length of the candidate eccDNA molecule is approximately equal to a distance of Y nucleotides. In some embodiments, the length of the candidate eccDNA molecule is exactly equal to a distance of Y nucleotides. The length of the candidate eccDNA molecule is less than about 50%, 60%, 70%, 80%, 90%, or more of the average length of the A fragments. In some embodiments, the error-corrected sequences obtained in (b) are specific to a single genomic region. In some embodiments, the error-corrected sequences obtained in (b) are specific to about 1 to about 30 distinct genomic loci.In some embodiments, the method is performed with or without an enrichment step to increase the proportion of double-stranded circular DNA molecules among all double-stranded nucleic acids in the sample, and further includes a step of comparing the frequencies in the library of possible insertions detected in step (c), candidate eccDNA breakpoints detected in step (d), and / or candidate eccDNA molecules detected in step (e) obtained by the method performed with or without the enrichment step. In some embodiments, the enrichment step comprises selectively removing double-stranded linear DNA molecules from the sample. In some embodiments, the double-stranded linear DNA molecules are selectively removed by treating the sample with one or more exonucleases. In some embodiments, the enrichment step comprises selectively isolating double-stranded circular DNA molecules from the sample. In some embodiments, the double-stranded circular DNA molecules are selectively isolated by electrophoresis, column filtration, density gradient centrifugation, selective extraction, and / or using a DNA-binding protein that selectively binds to or maintains binding to double-stranded circular DNA molecules relative to double-stranded linear DNA molecules. In some embodiments, the DNA-binding protein is a helicase. In some embodiments, the ratio of the number of candidate eccDNA molecules detected in step (e) to the number of error-corrected sequences obtained in step (b), the number of possible insertions detected in step (c), or the number of candidate eccDNA breakpoints detected in step (d) is higher when the method is performed with the enrichment step than when the method is performed without the enrichment step. In some embodiments, the frequency of possible insertions detected in step (c) with the enrichment step is less than about 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, or 10% of the frequency of possible insertions detected in step (c) without the enrichment step. In some embodiments, the frequency of candidate eccDNA breakpoints detected in step (d) with the enrichment step is less than about 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, or 10% of the frequency of candidate eccDNA breakpoints detected in step (d) without the enrichment step.In some embodiments, the frequency of candidate eccDNA molecules detected in step (e) with the enrichment step is at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or more of the frequency of candidate eccDNA molecules detected in step (e) without the enrichment step. In some embodiments, the method further includes calculating the probability that the candidate eccDNA molecule identified in step (e) is a bona fide eccDNA molecule. In some embodiments, the calculation is based in part on any one or more of the frequencies, ratios, or percentages described herein. In some embodiments, the calculation is based in part on the relationship between the length of the candidate eccDNA molecule and the average length of double-stranded DNA fragments in the library, where a shorter length of the candidate eccDNA molecule relative to the average length of double-stranded DNA fragments in the library indicates a high probability that the candidate eccDNA molecule is a bona fide eccDNA molecule. In some embodiments, the calculation is based in part on observing that the length of the candidate eccDNA molecule is approximately or exactly equal to the distance Y nucleotides, indicating a high probability that the candidate eccDNA molecule is a bona fide eccDNA molecule. In some embodiments, any one or more of steps (a)-(e), determining any one or more of the frequencies, ratios, or percentages, or any calculations disclosed herein are performed on a computer. In some embodiments, any one or more of steps (a)-(e), determining any one or more of the frequencies, ratios, or percentages, or any calculations disclosed herein are performed on a cloud. In some embodiments, a computer-based system for performing any of the methods provided herein is provided.
[0154] In some embodiments, a method of treating a disease or other medical condition in a mammalian subject comprises (i) performing a method disclosed herein on a biological sample obtained from the subject, (ii) identifying in the sample one or more candidate eccDNA molecules indicative of a physiological condition associated with the disease or medical condition, and (iii) administering to the subject a treatment for the disease or medical condition.
[0155] In some embodiments, a method for preparing a sequencing library for detecting candidate extrachromosomal circular DNA (eccDNA) molecules in a biological sample includes: (a) providing a biological sample containing double-stranded DNA; (b) preparing a first, non-enriched portion of the biological sample and a second, enriched portion enriched in double-stranded circular DNA molecules, wherein the second portion is prepared by selectively removing linear double-stranded DNA molecules and / or selectively isolating double-stranded circular DNA molecules from the sample; (c) fragmenting double-stranded DNA molecules in the first portion of the biological sample to generate a population of non-enriched double-stranded DNA fragments and enzymatically fragmenting double-stranded DNA molecules in the second portion of the biological sample to generate a population of enriched double-stranded DNA fragments; (d) ligating sequencing adaptors to a plurality of the non-enriched double-stranded DNA fragments to generate a non-enriched sequencing library; and (e) ligating sequencing adaptors to a plurality of the enriched double-stranded DNA fragments to generate an enriched sequencing library. In some embodiments, the biological sample is selected from the group consisting of sperm samples, semen samples, prostatic fluid samples, testicular biopsy samples, spermatogonial samples, germ cell samples, gamete samples, swab samples, lavage samples, aspirate samples, biopsy samples, tissue samples, tumor samples, precancerous lesion samples, liquid biopsy samples, hyperplasia samples, hypertrophy samples, dysplasia samples, urine samples, cerebrospinal fluid (CSF) samples, other bodily fluid samples, autopsy samples, autopsy samples, surgical samples, model organism samples, plasma samples, serum samples, gastric juice samples, bone marrow samples, stool samples, brushing samples, bile samples, pancreatic juice samples, synovial fluid samples, sputum samples, mucus samples, vitreous humor samples, forensic samples, environmental samples, bacterial samples, fungal samples, mammalian samples, human samples, and diagnostic samples. In some embodiments, the biological sample contains nucleic acids from potentially cancerous cells or potentially cancerous cells. In some embodiments, the biological sample contains cell-free DNA.In some embodiments, the biological sample contains cells exposed to a potentially toxic substance. In some embodiments, the potentially toxic substance is a potential clastogen, aneugen, mutagen, and / or teratogen, aneugen, mutagen, and / or teratogen. In some embodiments, the fragmentation is enzymatic fragmentation. In some embodiments, the second sample is enriched by treating the sample with one or more exonucleases. In some embodiments, the one or more exonucleases include exonuclease I, exonuclease T, exonuclease VII, exonuclease III, T7 exonuclease, exonuclease V (RecBCD), exonuclease VIII, lambda exonuclease, or T5 exonuclease. In some embodiments, the method further comprises treating a portion of the second sample with one or more endonucleases prior to treatment with one or more exonucleases, and comparing the candidate eccDNA obtained with and without treatment with the one or more endonucleases. In some embodiments, the second sample is enriched by selectively isolating double-stranded circular DNA molecules from the sample. In some embodiments, the double-stranded circular DNA molecules are selectively isolated by electrophoresis, column filtration, density gradient centrifugation, selective extraction, and / or using a DNA-binding protein that selectively binds to or maintains binding to double-stranded circular DNA molecules relative to double-stranded linear DNA molecules. In some embodiments, the DNA-binding protein is a helicase. In some embodiments, the biological sample is treated with DTT prior to (b), (c), or step (d).In some embodiments, the method further includes (b) preparing third and fourth portions of the biological sample by removing subportions from the first unenriched portion and the second enriched portion, respectively; treating a portion of the third and fourth portions with a reagent that induces cleavage in double-stranded circular DNA molecules at DNA damage sites, leaving another portion of the third and fourth portions untreated; and ligating sequencing adapters to the treated and untreated portions of the third and fourth portions. In some embodiments, the reagent is a combination of FPG (formamidopyrimidine [fapy]-DNA glycosylase) or UDG (uracil-DNA glycosylase) and endonuclease VIII. In some embodiments, the sequencing adapters are double-stranded sequencing adapters. In some embodiments, the sequencing adapters comprise a Y-shape. In some embodiments, the sequencing adapters are hairpin adapters. In some embodiments, a sequencing library prepared using any of the methods provided herein is provided. In some embodiments, a kit for performing any of the methods provided herein is provided.
[0156] Further aspects of the invention are described in the following numbered paragraphs. 1. A method for identifying at least one extrachromosomal circular DNA (eccDNA) in a sample containing double-stranded DNA, comprising: performing or causing double-stranded sequencing on the sample; Identifying or allowing the eccDNA to be identified from the multiple sequence reads of the double-stranded sequencing; The method includes: 2. The use of double-stranded sequencing to identify at least one extrachromosomal circular DNA (eccDNA) in a sample containing double-stranded DNA. The definitions, explanations and other features that apply to the methods of this application apply equally to the above terms of use. 3. In the method of item 1 or the use of item 2, the double-stranded sequencing a) tagging double-stranded DNA fragments in the sample by ligating adapters to the ends of the DNA, where each strand within each DNA fragment is i) tagged with a unique molecular identifier (UMI) to identify the strand as being derived from a single DNA molecule; and ii) each strand is labeled with a strand identifier (SDE) to distinguish between a first strand and a second strand within the single DNA molecule; b) amplifying the tagged DNA; c) sequencing the tagged amplification products, wherein reads from each strand are identifiable as originating from one DNA molecule by the UMI label, and reads from a first strand of the DNA molecule are distinguishable from reads from a second strand of the same DNA molecule by the SDE label; Includes. UMI labels are also referred to in the art as SMI (Single Molecular Identifier) labels. 4. In the method or use of paragraph 3, further: d) comparing the first strand sequence read with the second strand sequence read for a read from a single DNA molecule; Includes. During comparison, nucleotides at positions where the first strand read and the second strand read do not match are excluded (or can be excluded). 5. In the method or use of any of the above paragraphs, the double-stranded sequencing comprises single-stranded consensus sequencing (SSCS) and / or duplex consensus sequencing (DCS). Clauses 3-5 apply equally to other aspects of the invention (e.g., those described in claim 38 and paragraph 7). In addition, claims 2-37 may depend from either clause 1 or clauses 3-5.
[0157] IX. Working Example The following sections provide some non-limiting examples of how to prepare a sequencing library and detect candidate eccDNA using double-stranded sequencing.
[0158] Example 1 Preparation of sequencing libraries from sperm and blood samples and analysis of sequencing data.
[0159] PRJ00150 preparation Samples: Paired blood and sperm samples from six young men. DNA from the blood was isolated using a Qiagen kit according to the manufacturer's instructions. The sperm were isolated using standard protocols (including high-concentration DTT and bead beating). Otherwise, the samples were processed using Covaris and double-stranded sequencing standard protocols for enzymatic fragmentation, respectively.
[0160] The data was processed through the TwinStrand standard internal pipeline.
[0161] Methods for identifying candidate eccDNA: The mut files were preprocessed with dsreporter and then filtered by variation_type==”indel”&length(alt)>length(ref) to identify all variant calls consistent with insertions. From here on, the alt allele is referred to as α.
[0162] Histograms of length(α) were generated for various combinations of the samples and the treatments, and significant periodicity was observed in most samples / groups.
[0163] PRJ00178 preparation The sample was sperm DNA (pooled across individuals) from PRJ00150. All other samples were TS human devDNA (DNA extracted from blood collected from young, healthy human blood donors).
[0164] Sample preparation: None (processed according to standard protocol).
[0165] The devDNA was exposed to 10 mM, 150 mM, or 600 mM DTT (including IDTE and water controls) overnight at 37°C to mimic DTT-associated damage during sperm DNA extraction, and then purified with 1.6xSPRI beads.
[0166] To repair existing DNA damage that could cause artifactual variant calling, the sperm DNA and the devDNA were treated with 1x LCM (FPG and UDG) at 37°C for 1 hour, and then purified with 1.8x SPRI beads.
[0167] The samples were otherwise processed according to Covaris and Duplex-sequencing standard protocols for enzymatic fragmentation, respectively.
[0168] Data processing,
[0169] Methods for identifying candidate eccDNA:
[0170] The mut files were preprocessed with dsreporter and then filtered by variation_type==”indel”&length(alt)>length(ref) to identify all variant calls consistent with insertions. From here on, the alt allele is referred to as α.
[0171] Histograms of length(α) were generated for the various combinations of the samples and the treatments, and the histograms showed significant periodicity in most samples / groups.
[0172] The matchedPattern function in the BSgenome.Hsapiens.UCSC.hg38 package was used to identify the location of α within the chromosome containing the variant call (allowing 1 mismatch per 50 bases).
[0173] If the alt allele (α) matches the chromosome and location of the variant call, this is considered to indicate the presence of a BA junction, which matches either a chromosomal tandem duplication or eccDNA. These were considered candidate eccDNAs for the reasons listed below.
[0174] Reads supporting representative candidate eccDNAs were identified and validated by IGV.
[0175] Customer mouse tumor-normal sample For all sample types, the DNA was extracted using a Qiagen kit, the duplex sequencing was performed according to the kit protocol, and the data were analyzed using the TS pipeline on DNANexus.
[0176] Additional data processing: The mut files were preprocessed with dsreporter and then filtered by variation_type==”indel”&length(alt)>length(ref) to identify all variant calls consistent with insertions. From here on, the alt allele is referred to as α.
[0177] Histograms of length(α) were generated for various combinations of the samples and the treatments, and significant periodicity was observed in most samples / groups.
[0178] Evidence that this is not an artifact caused by DNA damage during sperm DNA isolation Pretreatment of the sperm DNA with LCM to remove damaged molecules did not reduce the number of identified eccDNA candidates (compared to the corresponding control group).Furthermore, pretreatment of the devDNA with high concentrations of DTT did not increase the number of identified eccDNA candidates (compared to the corresponding control group).
[0179] evidence that the candidate eccDNA is not (primarily) a chromosomal tandem duplication (TD); The size distribution of these alphas is similar to that previously reported for eccDNA (especially microDNA), with the periodicity corresponding to the DNA length surrounding histones. Chromosomal TDs in this size range (<1 kb) likely arise during DNA replication or DNA repair, where normal chromatin structure is disrupted and therefore are not expected to affect the results. The size distribution pattern is consistent with that of known eccDNAs (see, e.g., Dillon, L. et al., Cell Reports, 2015; Mehana, P. et al., PLoS One, 2017). Furthermore, the size distribution is independent of the shearing method (a similar trend is observed in mechanically sheared libraries, see Figure 4), further suggesting the biological relevance of these indel calls.
[0180] Furthermore, no examples were found in which the insert size of the supporting consensus read was larger than α, even when α was smaller than the median insert size of the library, a situation that would be expected to occur by chance in chromosomal TD.
[0181] Additionally, no examples were found in which the consensus read supporting the indel call contained a portion of the α multiple times or contained sequences derived from outside the α in the reference genome.
[0182] Furthermore, in the specific example shown in Figures 5A-5D, some of the indel calls showed lengths that perfectly matched the lengths of the two sections of apparent duplication joined by the new junction. This strongly supports the possibility that the indel calls originated from circular genome-derived DNA present in the original sample. The circular DNA was likely cleaved in a single cleavage event during sample preparation. The enzymatic fragmentation method cleaves DNA randomly, as evidenced by the fact that many other fragments in the library align to the same reference region but have various different cleavage sites / fragment ends. Therefore, in a tandem duplication event, it is highly unlikely that the enzyme cleaves at the exact same breakpoint in both copies of the duplication.
[0183] Example 2. Methods for verifying the identification of candidate eccDNAs. Major demonstration experiments Identify sample types with high observed or predicted candidate eccDNA counts, including:
[0184] Sample types for which candidate eccDNA has already been identified (e.g., human sperm, mouse gastrointestinal tumors),
[0185] Reanalyzing existing DS data to identify other samples / sample types with high candidate eccDNA counts;
[0186] This involves identifying sample types that have traditionally been shown to have high eccDNA loads by other methods.
[0187] Ideally, genomic DNA is isolated from all samples of interest using a single, gentle DNA isolation method (e.g., Qiagen Blood & Tissue Kit). While specialized extraction methods may be required for some sample types, such as sperm, care is taken to isolate high-quality DNA with minimal extraction-related damage.
[0188] For each sample, prepare two corresponding DS libraries as follows:
[0189] First library preparation is prepared by hybrid capture using the species-specific mutagenic panel according to the standard protocol, with approximately 500 ng of total genomic DNA as input.
[0190] Second library preparation (exonuclease V (RecBCD) pretreatment to enrich for circular DNA),
[0191] The genomic DNA is spiked with bacterial plasmids at a 1:1 molar ratio (1 genome equivalent per plasmid) as a control to monitor the enrichment of circular DNA.
[0192] 50 units of exonuclease V (NEB) is added to 5 μg of the genomic DNA spiked with the plasmid, and the mixture is incubated at 37° C. for 2 hours.
[0193] To inactivate the Exo V, EDTA is added to a final concentration of 11 mM, and the mixture is incubated at 65°C for 30 minutes.
[0194] The remaining DNA is purified by 1.5x SPRI purification and the samples are quantified as follows:
[0195] The concentration is measured using the Qubit HS DNA kit, and the target is a 10-fold or greater reduction in total DNA yield compared to before digestion.
[0196] If the DNA concentration is greater than 0.5 ng / µL, run the Agilent Genomic DNA ScreenTape to confirm removal of high molecular weight chromosomal DNA.
[0197] To confirm the enrichment of circular DNA, qPCR was performed using primers targeting three randomly selected chromosomal regions and one plasmid region, aiming for a ≥100-fold reduction in at least one of the chromosomal regions relative to the plasmid region (compared to the pre-digested sample).
[0198] If the initial reduction target (≥10-fold by mass, ≥100-fold by qPCR) is not achieved, adjust the Exo V treatment protocol (e.g., double the enzyme units and incubation time). If reduction is still insufficient, record the amount of reduction and prepare libraries as follows:
[0199] Library preparation method using ExoV-digested DNA,
[0200] If more than 1 ng of DNA remains after the digestion, prepare a DS library using 1 ng of input DNA according to the standard protocol, increasing the number of cycles in each PCR step (determined empirically).
[0201] If less than 1 ng of DNA remains after the digestion, a known amount of Exo V-digested DNA is mixed with bacterial plasmid DNA (to serve as carrier DNA during library construction) to prepare the library (up to the first PCR step), noting that the bacterial plasmid DNA is substantially removed in the hybrid capture step of the DS protocol.
[0202] The DS library will be sequenced and analyzed according to the standard protocol, with the addition of a proposed eccDNA candidate identification pipeline.
[0203] Additional details, If necessary, one or both of the DS libraries in the sample pair (+ / - Exo V) are downsampled to have identical PTFS before proceeding to identify the eccDNA candidates.
[0204] The frequency of the eccDNA candidates is calculated as the number of the eccDNA candidates per informative duplex base in each library.
[0205] Expected results (evident for the eccDNA candidate identification pipeline);
[0206] The frequency of the eccDNA candidates (defined above) should be significantly higher in the Exo V-treated library compared to the control library, with the magnitude of the change depending on the amount of depletion due to the Exo V treatment (estimated by DNA yield and qPCR quantification as above).
[0207] If the same eccDNA candidate is identified in both libraries of the sample pair (+ / - Exo V), this is strong evidence that they are true circular DNA. However, even if the same eccDNA candidate is not identified between the library pair, this does not negate the possibility that they are the eccDNA. If only a single copy of the eccDNA is present in the biological sample, it will be randomly distributed to either of the libraries and will be identified in only one (or none) of the libraries. Similarly, eccDNA present at extremely low copy numbers in the biological sample may not be detected in both libraries due to simple sampling statistics.
[0208] The similar size distribution profiles of eccDNA candidates in both library types (with and without Exo V treatment) support the validity of the eccDNA calls in the unmodified libraries.
[0209] If no substantial reduction in eccDNA frequency is observed in the Exo V-treated library compared to the control (suggesting that most candidates are not true eccDNA), or if a reduction is observed that is significantly lower than expected based on the amount of chromosomal DNA reduction in the Exo V-treated library (suggesting a mixture of true-positive and false-positive eccDNA calls),
[0210] Individual eccDNA candidates supported by 10 or more double-stranded consensuses (alt allele counts >10) in the non-Exo-treated library are not detected in the Exo-treated library.
[0211] As an alternative to, or in addition to, Exo V-dependent removal of linear chromosomal DNA, DNA extracted from a single sample can be used to prepare a pair of libraries using two different methods: one for total chromosomal DNA (e.g., Qiagen Blood and Tissue kit) and one for plasmid isolation.
[0212] Example 3. Discovery of eccDNA in human sperm The following analysis was performed using the samples and data provided in Example 1. Briefly, human sperm DNA was obtained from six healthy young donors. The double-strand sequencing (DS) libraries were prepared using 500 ng of DNA as input. The TwinStrand® DuplexSeq™ Human Mutagenesis Assay panel (20 x 2.4 kb target regions spanning the genome) was used for hybrid capture. Because the allele length distribution of large insertions was reminiscent of small extrachromosomal circular DNA (eccDNA) or microDNA, these variant calls were analyzed in more detail. First, the variant location and the alternative allele sequence were analyzed to determine whether they supported the expected junction type for either circular DNA or chromosomal tandem duplication (a junction fusing the ends of alleles ABCD, hereafter referred to as a "DA junction"). To identify DA junctions (Figures 6A-6C, 7A-7C), the alternative allele sequences were exported to a FASTA file and used as input for the matchPattern function in Biostrings (Bioconductor package). The reference genome sequence used was BSgenome.Hsapiens.UCSC.hg38 (Bioconductor package), allowing one mismatch per 50 bp. The search was performed chromosome-by-chromosome, detecting matches only on the chromosome for the variant call. Zero or one match was obtained for each sequence. If the matchPattern function returned a match located exactly 1 bp downstream of the variant's location in the original mut file, the variant was determined to support a DA junction.
[0213] For fragment length analysis, duplex consensus reads supporting variant calls were visualized in IGV. First, a BLAT search of auxiliary alignments and / or soft-clipped sequences was performed to confirm the presence of DA junctions. For all manually inspected events (read pairs with DA junctions), we confirmed the following: all 5' soft-clipped bases aligned to the other end of the allele, supporting the DA junction; 3' soft-clipped bases represented readthrough to the DS adapter (present only in fragments shorter than 142 bp); and the majority of consensus read pairs containing DA junctions aligned as mismatched pairs. One or both of the consensus reads contained substantial soft-clipped regions, indicating auxiliary alignment to the other end of the allele (across the DA junction). For each read pair, the physical DNA fragment length was estimated using the following procedure: a pseudo-read pair was constructed consisting of the primary alignment of one read in the pair and the auxiliary alignment of the other read. The difference between the rightmost end of the reverse-aligned read and the leftmost end of the forward-aligned read (including the 5' soft-clipped base) was calculated. A small number of DA junction-containing consensus read pairs were aligned cooperatively, and in this case, the fragment length was estimated as the calculated insert size plus the 5' soft-clipped base. The distance between the 5' ends of outgoing read pairs (primary-primary or primary-supplementary) was recorded based on visualization and confirmed to be equal to the allele length minus the estimated fragment length, except for one confirmed chromosomal TD (tandem duplication). To ensure a reasonable probability of detecting duplicated sequences, if any, a subset of apparent insertion variant calls with allele lengths less than the median insert size (233 bp) across the sperm DS library was selected for systematic manual inspection. Given the null hypothesis that all DA junctions originate from chromosomal TDs, the length distribution of the fragments supporting this subset of variant calls was expected to be similar to the insert size distribution across the entire library.Specifically, it was estimated that approximately half of the fragments were less than 233 bp, the remaining half were greater than 233 bp, and the majority of the fragments contained overlapping ABCD sequences (fragments with lengths exceeding the allele length). A binomial test p-value was calculated to test whether the observed results were significantly different from the null hypothesis.
[0214] Most apparent large insertions were called relative to the end and start of the reference allele based on a split alignment of double-stranded consensus reads (Figure 6A i-iii). Given the reference allele (ABCD), this type of alignment reveals a junction connecting the end of the reference allele (D) to the start of the reference allele (A), termed a DA junction. DA junctions are formed by chromosomal tandem duplication (TD) of the ABCD allele (Figure 6A v) or when the ABCD allele is excised and circularized (Figure 6A iv). Figure 6B shows the periodic allele length distribution of apparent insertions, color-coded by the presence or absence of a DA junction. Across the six sperm samples, 96% (307 of 320) of apparent insertions larger than 20 bp and 100% (n = 307) of apparent insertions larger than 125 bp contained a DA junction (Figure 6B). In contrast, no DA junctions were detected in inserts longer than 20 bp (n = 7) in the corresponding blood samples. The length distribution of apparent inserts containing DA junctions is remarkably similar to the distribution of small extrachromosomal circular DNA (eccDNA, also known as microDNA) reported in various cell types, including sperm.
[0215] To determine whether the DA junctions detected by DS are derived from microDNA, we investigated the relationship between the length of the DNA fragments containing the DA junction and the length of the ABCD allele. Genomic DNA was enzymatically fragmented during DS library preparation, and the size distribution of DNA fragments in the final DS library was approximated by the insert size distribution of the sequenced library (insert size was calculated as the distance between the 5' ends of read 1 and read 2 in a read pair aligned to the reference genome). If the DA junction is derived from a chromosomal tandem duplication (TD), the size distribution of fragments containing the junction should be independent of the size of the duplicated ABCD allele and similar to the distribution of the entire library (Figures 7A-7C). Furthermore, fragments larger than the ABCD allele contain at least two copies of a portion of the ABCD sequence (e.g., CDABCD) and may also contain flanking regions of non-ABCD sequences (Figures 7A-7C). Conversely, circular DNA molecules must undergo one or more cleavages to generate linear DNA fragments in the final library that are less than or equal to the length of the DNA circle (defined as the length of the ABCD allele (Figures 7A-7C)).
[0216] To ensure a reasonable probability of detecting duplicated sequences, we selected all DA-containing apparent insertions (n = 63) whose ABCD allele length was less than or equal to the median insert size (233 bp) of the six sperm libraries, and manually confirmed the aligned duplex consensus reads supporting each call. If most or all of the DA-containing fragments were derived from chromosomal TD, approximately half of the fragments would be greater than 233 bp and the other half would be less (Figure 6C; Thiel distribution, Figures 7A-7C). Furthermore, at least half of the events would have fragment sizes exceeding the ABCD allele length, thus containing the duplicated ABCD sequence and possibly the non-ABCD wing sequence (Figures 7A-7C). Conversely, all of the microDNA-derived fragments would be less than 233 bp, the length of the ABCD allele, and would not contain the duplicated or wing sequence (Figure 6C; shaded distribution, Figures 7A-7C). All 63 confirmed apparent insertions had fragment lengths less than 233 bp, which was highly inconsistent with the DA junctions being primarily derived from the chromosomal TD (binomial test p-value = 1.08 × 10-19), indicating that the majority of the DA junctions were derived from the microDNA. Furthermore, in 62 of the 63 apparent insertions, the fragment size was shorter than the ABCD allele length, consistent with microDNA origin. Only one chromosomal duplication (TD) was definitively identified: a 75-bp allele located on the 180-bp fragment, containing duplicated and flanking sequences. Notably, this was the smallest DA junction-containing variant call in the dataset, much smaller than all other DA junction-containing alleles (Figure 6B).
[0217] All DA junction-containing events with an allele length greater than 125 bp were considered candidate eccDNAs. Thus, 307 candidate eccDNAs were detected across the six sperm samples (average 51 per sample, standard deviation 18), yielding an average frequency of 41 unique DNA circles per billion duplex bases in sperm (standard deviation 14). This was higher than the combined frequency of all other large indels (>20 bp) and structural variant (SV) types. No candidate eccDNAs were identified in the six corresponding blood samples.
[0218] Example 4. Detection of eccDNA using exonuclease enrichment The DNA samples (HeLa (BioChain), DNA from HeLa cells purchased from BioChain; Blood (TS1), DNA from human whole blood extracted using an Agilent DNA extraction kit; Sperm (TS2), DNA from human sperm cells or tissues suspended in Qiagen RLT lysis buffer containing 10% TCEP and homogenized by bead beating, then extracted using a modified version of the Qiagen DNeasy Mini protocol; and Sperm (TS3), DNA from human sperm cells or tissues suspended in Qiagen RLT containing 10% TCEP, supplemented with Proteinase K to a final concentration of 200 μg / ml, and incubated at 56°C for 2 hours, then extracted using a modified version of the Qiagen DNeasy Mini protocol) were subjected to exonuclease treatment to remove linear DNA and enrich for circular DNA species present in the samples. If the DA junction-containing fragments detected by DS are primarily or exclusively derived from DNA circles, exonuclease treatment will increase the frequency of detected events (frequency = (number of circles detected) / (number of duplex bases sequenced)). If the DA junction-containing fragments are derived from chromosomal tandem duplications, exonuclease treatment will not increase their frequency.
[0219] A 1 μg aliquot of the DNA was treated with Escherichia coli exonuclease V (RecBCD, NEB M0345L) according to the manufacturer's recommended protocol, but with varying incubation times. For each DNA sample and incubation time, three parallel reactions were pooled to ensure sufficient yield. A no-enzyme control reaction ("Control") was performed for the 30-minute time point. After exonuclease treatment, the DNA was purified with SPRI beads (1.2x ratio). The exonuclease treatment reduced total DNA by 85% to >99% in all sample types compared to the no-enzyme control. The enrichment of circular DNA was further confirmed by qPCR quantification of the ratio of mitochondrial DNA to nuclear DNA in the HeLa cells and the blood samples.
[0220] DS libraries were prepared using 300 ng of DNA input for the control group and 2-150 ng of DNA input for the exonuclease-treated samples. The human mutagenic panel (20 sites across the genome x 2.4 kb target regions) was used for hybrid capture. Candidate circular DNA was defined as an "indel" variant call (obtained from standard DS variant calling) with the following additional features: the alt allele is longer than the reference allele and is greater than 125 bp; and The variant call is consistent with the occurrence of a DA junction in the consensus read pair.
[0221] Figure 8 shows the frequency of candidate circles (frequency = (number of detected circles) / (number of sequenced double-stranded bases)). Exonuclease treatment resulted in a time-dependent increase in the frequency of candidate circles in three of the four sample types tested. In the fourth sample type, blood from young healthy donors, no candidate circles were detected in either the control or exonuclease-treated groups. This finding indicates that the DA junction-containing fragments detected by DS (Duplex-Sequencing) are primarily or exclusively derived from DNA circles.
[0222] Example 5 Candidate eccDNA in tumor and normal samples Without wishing to be bound by any particular theory, circular DNA profiles (quantity and characteristics of eccDNA and microDNA) may differ between normal and cancer tissues, and circular DNA profiles may be used as biomarkers for monitoring cancer progression and post-treatment outcomes. If DS detects true biological differences in the frequency of circular DNA, differences will be observable between paired tumor-normal samples.
[0223] Triplicate sets of paired tumor and normal DNA representing three common cancer types (breast, colon, and lung) were purchased from BioChain. DS libraries were prepared using 750 ng of DNA as input. Hybrid capture was performed using the DuplexSeq Human Mutagenesis Assay panel (20 x 2.4 kb genome-wide target regions). Hypothetical DNA circles were defined as "indel" variant calls (from standard DS variant calling) with the following additional features: the alt allele is longer than the reference allele and is greater than 125 bp; and The variant call is consistent with the occurrence of a DA junction in the consensus read pair.
[0224] Figures 9A-9B show the frequency of the hypothetical DNA circles (calculated as frequency = (number of circles detected) / (number of duplex bases sequenced)). The DS showed a moderate increase in hypothetical DNA circles in all three tumors compared to the corresponding normal samples (Figure 9A). The increase in hypothetical DNA circles did not correspond to an increase in the mutation frequency in two of the three tumor types (Figure 9B).
[0225] Example 6. eccDNA as an indicator of clastogenicity To assess whether DNA circle frequency could be a surrogate indicator of clastogenicity, the DS data for various genotoxicants were reanalyzed. ENU is a well-known potent mutagen and clastogen. The DS data showed that compound #1 was nonmutagenic (at the doses tested), and compound #2 increased the mutation frequency in a dose-dependent manner.
[0226] Human TK6 cells were treated with different potentially genotoxic compounds. DS libraries were prepared using 500 ng of DNA as input. The human mutagenic panel (20 sites across the genome × 2.4 kb target region) was used for hybrid capture. Candidate circles were defined as "indel" variant calls (from standard DS variant calling) in which the alt allele was longer than the reference allele and had a length greater than 125 bp. In other datasets, all variants meeting this criterion were also shown to be associated with DA junctions. Figures 10A-10B show the frequency of candidate circles, calculated as frequency = (number of circles detected) / (number of duplex bases sequenced).
[0227] Treatment with ENU, a known clastogen, increased the frequency of candidate DNA circles in a dose-dependent manner, reaching statistical significance at the highest dose. Treatment with compounds #1 and #2 also increased the frequency of candidate DNA circles in a dose-dependent manner, despite the lack of mutagenicity of compound #1 at the doses tested.
[0228] In a separate experiment, TK6 and HepaRG cells were co-cultured in a system where the cells were physically separated but shared culture medium. The co-cultures were treated with ultrapure water (control) or cyclophosphamide, with each group run in triplicate. DS libraries were prepared using variable input amounts of DNA. Hybrid capture was performed using the DuplexSeq Human Mutagenesis Assay panel (20 x 2.4 kb target regions spanning the genome).
[0229] DNA circles were defined as "indel" variant calls (from standard DS variant calling) in which the alt allele was longer than the reference allele and had a length greater than 125 bp. In other datasets, all variants meeting this criterion were shown to be associated with DA junctions. Figures 11A-11B show the frequency of DNA circles, calculated as frequency = (number of circles detected) / (number of duplex bases sequenced). Cyclophosphamide, a known chromosome breaker, significantly increased both DNA circle and mutation frequencies in the TK6-HepaRG coculture system.
[0230] X. Conclusion From the foregoing description, it will be understood that specific embodiments of the present technology have been described for illustrative purposes, although well-known structures and functions have not been described or illustrated in detail to avoid unnecessarily obscuring the description of the embodiments of the present technology. Where the context permits, terms described in the singular or plural shall include the respective plural or singular forms.
[0231] The above detailed description of embodiments of the present technology is not intended to be exhaustive or to limit the technology to the precise form disclosed above. While specific embodiments and examples of the present technology have been described above for illustrative purposes, those skilled in the art will recognize that various equivalent modifications are possible within the scope of the present technology. For example, even if steps are presented in a particular order, alternative embodiments may perform the steps in a different order. The various embodiments described herein can also be combined to provide yet other embodiments. All references cited herein are incorporated by reference in their entirety as if fully set forth herein. [Sequence table] <?xml version="1.0" encoding="UTF-8"?> <!DOCTYPE ST26SequenceListing PUBLIC "- / / WIPO / / DTD Sequence Listing 1.3 / / EN" "ST26SequenceListing_V1_3.dtd"> <st26sequencelisting dtdversion="V1_3" filename="TSB-018SeqList.xml" softwarename="WIPO Sequence" softwareversion="2.3.0" productiondate="2023-08-29"> <applicationidentification> <ipofficecode> WO< / ipofficecode> <applicationnumbertext>< / applicationnumbertext> <filingdate>< / filingdate> < / applicationidentification> <applicantfilereference> TSB-018WO< / applicantfilereference> <earliestpriorityapplicationidentification> <ipofficecode> US< / ipofficecode> <applicationnumbertext> 63 / 373,851< / applicationnumbertext> <filingdate> 2022-08-29< / filingdate> < / earliestpriorityapplicationidentification> <applicantname languagecode="en"> TwinStrand Biosciences, Inc.< / applicantname> <inventiontitle languagecode="en"> METHODS AND REAGENTS FOR DETECTION OF CIRCULAR DNA MOLECULES IN BIOLOGICAL SAMPLES< / inventiontitle> <sequencetotalquantity> 3< / sequencetotalquantity> <sequencedata sequenceidnumber="1"> <insdseq> <INSDSeq_length>131< / INSDSeq_length> <INSDSeq_moltype>DNA< / INSDSeq_moltype> <INSDSeq_division>PAT< / INSDSeq_division> <INSDSeq_feature-table> <insdfeature> <INSDFeature_key>source< / INSDFeature_key> <INSDFeature_location>1..131< / INSDFeature_location> <INSDFeature_quals> <insdqualifier> <INSDQualifier_name>mol_type< / INSDQualifier_name> <INSDQualifier_value>genomic DNA< / INSDQualifier_value> < / insdqualifier> <insdqualifier id="q2"> <INSDQualifier_name>organism< / INSDQualifier_name> <INSDQualifier_value>Homo sapiens< / INSDQualifier_value> < / insdqualifier> < / INSDFeature_quals> < / insdfeature> < / INSDSeq_feature-table> <INSDSeq_sequence>tgtttaggatgccccaggaatgtgcagatctgcgttggaagggtagaaacgggaatgaaagcaaatggtaaggatccatggaacgaaatatggggtggagatagcagcgtttccaccaaacttctatcgaa< / INSDSeq_sequence> < / insdseq> < / sequencedata> <sequencedata sequenceidnumber="2"> <insdseq> <INSDSeq_length>142< / INSDSeq_length> <INSDSeq_moltype>DNA< / INSDSeq_moltype> <INSDSeq_division>PAT< / INSDSeq_division> <INSDSeq_feature-table> <insdfeature> <INSDFeature_key>source< / INSDFeature_key> <INSDFeature_location>1..142< / INSDFeature_location> <INSDFeature_quals> <insdqualifier> <INSDQualifier_name>mol_type< / INSDQualifier_name> <INSDQualifier_value>other DNA< / INSDQualifier_value> < / insdqualifier> <insdqualifier id="q4"> <INSDQualifier_name>organism< / INSDQualifier_name> <INSDQualifier_value>synthetic construct< / INSDQualifier_value> < / insdqualifier> < / INSDFeature_quals> < / insdfeature> < / INSDSeq_feature-table> <INSDSeq_sequence>atgaaagcaaatggtnnggatccatggaacgaaatatggggtggagatagcagcgtttccaccaaacttctatcgaactttaggatgccccaggantgtgcagatctgcgttggaagggtagaaacgggnnnnnnnnnnnnn< / INSDSeq_sequence> < / insdseq> < / sequencedata> <sequencedata sequenceidnumber="3"> <insdseq> <INSDSeq_length> 142< / INSDSeq_length> <INSDSeq_moltype> DNA< / INSDSeq_moltype> <INSDSeq_division> PAT< / INSDSeq_division> <INSDSeq_feature-table> <insdfeature> <INSDFeature_key>source< / INSDFeature_key> <INSDFeature_location>1..142< / INSDFeature_location> <INSDFeature_quals> <insdqualifier> <INSDQualifier_name>mol_type< / INSDQualifier_name> <INSDQualifier_value>other DNA< / INSDQualifier_value> < / insdqualifier> <insdqualifier id="q6"> <INSDQualifier_name>organism< / INSDQualifier_name> <INSDQualifier_value>synthetic construct< / INSDQualifier_value> < / insdqualifier> < / INSDFeature_quals> < / insdfeature> < / INSDSeq_feature-table> <INSDSeq_sequence> nnnnnnnnnnnnatgaaagcaaatggtaaggatccatggaacgaaatatggggtggaganagcgtntccaccaaacttctatcgaagtttaggatgccccaggaatgcagatctgcgttggaagggtagaaacggg< / INSDSeq_sequence> < / insdseq> < / sequencedata> < / st26sequencelisting>
Claims
1. A method for identifying at least one extrachromosomal circular DNA (eccDNA) in a sample containing double-stranded DNA, the method comprising performing or having performed double-sequencing on the sample, the double-sequencing comprising the steps of: ligating adapters to the ends of the double-stranded DNA, wherein the at least one adapter comprises a nucleotide sequence that tags a strand of the double-stranded DNA such that the strand of the double-stranded DNA has a nucleotide sequence that is distinct from its complementary strand; amplifying the strand of the double-stranded DNA using the ligated adapters to generate at least a first strand amplicon and a second strand amplicon; sequencing the at least the first strand amplicon and the second strand amplicon to generate a plurality of sequence reads comprising a first strand sequence read and a second strand sequence read; and identifying or identifying the eccDNA using the plurality of sequence reads of the double-sequencing.
2. 2. The method of claim 1, wherein the step of performing or having performed double sequencing comprises generating error-corrected sequence reads by comparing the first strand sequence reads and the second strand sequence reads by discounting mismatched nucleotide positions.
3. 2. The method of claim 1, wherein identifying or identifying the eccDNA comprises identifying a subset of sequence reads from the plurality of sequence reads, wherein each sequence read of the subset of sequence reads independently comprises a reference allele junction.
4. The method of claim 3, comprising: distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads from among the subset of sequence reads; and determining the amount of eccDNA according to the distinguished putative eccDNA sequence reads.
4. The method of claim 3, wherein the reference allele junction comprises the nucleic acid sequence DA.
5. The method of claim 4, wherein nucleic acid sequence DA comprises nucleic acid sequence D operably linked to nucleic acid sequence A in the 5' to 3' direction.
6. The method of claim 4, wherein nucleic acid sequence D is located downstream of nucleic acid sequence A at the reference genomic locus of the reference allele.
7. 8. The method of any one of claims 4 to 7, wherein the nucleic acid sequence DA is at least 1 base pair (bp) in length.
8. 4. The method of claim 3, wherein the reference allele junction is due to an apparent indel or an apparent structural variant.
9. 9. The method of claim 8, wherein the apparent indel is at least 1 bp in length.
10. The method of claim 8, wherein the apparent structural variant is due to an apparent insertion or an apparent duplication.
11. The method of claim 10, wherein the apparent structural variant due to the apparent insertion or apparent duplication is at least 20 bp in length.
12. 12. The method of any one of claims 1 to 11, wherein distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads comprises selecting a subset of consensus sequence reads that contain DA junctions with allele lengths less than a threshold apparent insert size.
13. 13. The method of any one of claims 1 to 12, wherein distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads comprises inferring the fragment size of each read pair in the subset and comparing it one-by-one to a threshold apparent insert size, or comparing the distribution of inferred fragment sizes to the distribution of insert sizes in the entire library.
14. 14. The method of any one of claims 1 to 13, wherein distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads comprises comparing the inferred fragment size of any one consensus read pair to the allele size of that read pair.
15. 16. The method of any one of claims 12 to 15, wherein distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads comprises: selecting a subset of consensus sequencing reads containing DA junctions with allele lengths less than a threshold apparent insert size; estimating the fragment size of each read pair in the subset and comparing it one by one to a threshold apparent insert size, or comparing the distribution of estimated fragment sizes to the distribution of insert sizes in the entire library; and comparing the estimated fragment size of any one consensus read pair to the allele.
16. 16. The method of any one of claims 12 to 15, wherein the threshold apparent insert size is at least 20 bp in length.
17. The method of any one of claims 12 to 15, wherein the predicted fragment size is at least 20 bp in length.
18. The method according to any one of claims 12 to 15, wherein the putative eccDNA is identified as having the putative fragment size equal to or less than the allele length.
19. 19. The method of claim 18, wherein the allele length is at least 20 bp.
20. 20. The method of any one of claims 1 to 19, wherein the method comprises performing or having performed double sequencing on the sample, wherein the double sequencing ligates adapters to the ends of the double-stranded DNA, at least one adapter comprising a nucleotide sequence that tags the strand of the double-stranded DNA such that the strand of the double-stranded DNA has a nucleotide sequence that is distinct from its complementary strand; amplifying the strand of the double-stranded DNA using the ligated adapters to generate at least first-strand amplicons and second-strand amplicons; sequencing the at least first-strand amplicons and second-strand amplicons to generate first-strand sequence amplicons. generating a plurality of sequence reads including a first strand sequence read and a second strand sequence read; generating error-corrected sequence reads by comparing the first strand sequence read and the second strand sequence read by discounting mismatched nucleotide positions; and identifying or having identified eccDNA using the plurality of sequence reads of duplex sequencing, wherein the identification or having identified eccDNA includes the following steps: identifying a subset of sequence reads from the plurality of sequence reads, wherein the sequence reads of the subset of sequence reads each independently include a reference allele junction. Among the subset of sequence reads, putative eccDNA sequence reads are distinguished from chromosomal tandem duplication sequence reads, wherein distinguishing the putative eccDNA sequence reads includes the following steps: selecting a subset of consensus sequencing reads including DA junctions with allele lengths less than a threshold apparent insert size; estimating the fragment size of each read pair in the subset and comparing it one by one with a threshold apparent insert size, or comparing the distribution of estimated fragment sizes with the distribution of insert sizes in the entire library; and comparing the estimated fragment size of any one consensus read pair with the allele; determining the amount of eccDNA based on the identified putative eccDNA sequence.
21. Before double sequencing is performed or has been performed: the method of claim 1 or 20, further comprising exposing one or more cells to a potential clastogen; and obtaining eccDNA from the one or more cells.
22. 22. The method of claim 21, further comprising assessing the clastogenicity of the potential clastogen based on the determined eccDNA profile.
23. 23. The method of claim 22, wherein the profile comprises any one or any combination of the quantity, frequency, quality, size, genomic location, or any other characteristic of the eccDNA.
24. 24. The method of claim 23, wherein the profile comprises the frequency of eccDNA.
25. 25. The method of any one of claims 21 to 24, wherein the potential clastogen is a chemical compound, a physical exposure, a biological agent, or a complex mixture and / or an environmental exposure.
26. 26. The method of any one of claims 1 to 25, wherein the sample is selected from the group consisting of a sperm sample, a semen sample, a prostatic fluid sample, a testicular biopsy sample, a spermatogonial sample, a germ cell sample, a gamete sample, a swab, a lavage, an aspirate, a biopsy, a tissue sample, a tumor sample, a pre-neoplastic sample, a liquid biopsy, a hyperplasia sample, a hypertrophy sample, a dysplasia sample, a urine sample, a CSF sample, any other body fluid sample, an autopsy sample, a post-mortem sample, a surgical sample, a model organism sample, a plasma sample, a serum sample, a stomach sample, a bone marrow sample, a stool sample, a brushing sample, a bile sample, a pancreatic juice sample, a synovial fluid sample, a sputum sample, a mucus sample, a vitreous sample, a forensic sample, an environmental sample, a bacterial sample, a fungal sample, a mammalian sample, a human sample, and a diagnostic sample.
27. The method according to any one of claims 1 to 26, wherein the sample contains cancer cells or cancer-derived nucleic acids.
28. The method according to any one of claims 1 to 27, wherein the biological sample comprises cell-free DNA.
29. 29. The method of any one of claims 1 to 28, wherein the at least one adapter sequence is or comprises at least one non-standard nucleotide.
30. 30. The method of claim 29, wherein the non-standard nucleotide is selected from uracil, methylated nucleotides, RNA nucleotides, ribose nucleotides, 8-oxo-guanine, biotinylated nucleotides, desthiobiotin nucleotides, thiol-modified nucleotides, acrydite-modified nucleotides, iso-dC, iso-dG, 2'-0-methyl nucleotides, inosine nucleotide-locked nucleic acids, peptide nucleic acids, 5-methyl-dC, 5-bromodeoxyuridine, 2,6-diaminopurine, 2-aminopurine nucleotides, basic nucleotides, 5-nitroindole nucleotides, adenylated nucleotides, azido nucleotides, digoxigenin nucleotides, I-linkers, 5'-hexynyl-modified nucleotides, 5-octadinyl-dU, photocleavable spacers, non-photocleavable spacers, click chemistry-compatible modified nucleotides, fluorescent dyes, biotin, furan, BrdU, fluoro-dU, roto-dU, and any combination thereof.
31. 10. The method of claim 1, wherein the method further comprises performing eccDNA enrichment.
32. 32. The method of claim 31 , wherein said performing eccDNA enrichment comprises performing size selection. performing exonuclease treatment; and / or using a DNA binding protein that differentially binds to or remains bound to double-stranded circular DNA molecules compared to double-stranded linear DNA molecules.
33. 33. The method of claim 32, wherein the size selection comprises the use of paramagnetic beads at a size threshold of about 10,000 bp, electrophoresis, column filtration, density gradient centrifugation, or selective extraction.
34. 33. The method of claim 32, wherein the exonuclease is selected from the group consisting of exonuclease I, exonuclease T, exonuclease VII, exonuclease III, T7 exonuclease, exonuclease V (RecBCD), exonuclease VIII, lambda exonuclease, T5 exonuclease, and any combination thereof.
35. A method for identifying extrachromosomal circular DNA (eccDNA), comprising: obtaining a plurality of sequence reads of sequenced double-stranded DNA using duplex sequencing; identifying a subset of sequence reads from the plurality of sequence reads, wherein the sequence reads of the subset of sequence reads each independently contain a reference allele junction; distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads among the subset of sequence reads; and identifying eccDNA according to the distinguished putative eccDNA sequence reads.
36. 36. The method of claim 35, wherein the reference allele junction comprises the nucleic acid sequence DA.
37. 37. The method of claim 36, wherein the nucleic acid sequence DA comprises nucleic acid sequence D operably linked to nucleic acid sequence A in the 5' to 3' direction.
38. 38. The method of claim 37, wherein the nucleic acid sequence D is located downstream of the nucleic acid sequence A at the reference genomic locus of the reference allele.
39. 39. The method of any one of claims 36 to 38, wherein the nucleic acid sequence DA is at least 1 base pair (bp) in length.
40. 36. The method of claim 35, wherein the reference allele junction is due to an apparent indel or an apparent structural variant.
41. 41. The method of claim 40, wherein the apparent indel is at least 1 bp in length.
42. 41. The method of claim 40, wherein the apparent structural variant is due to an apparent insertion or an apparent duplication.
43. 43. The method of claim 42, wherein the apparent structural variant due to the apparent insertion or apparent duplication is at least 20 bp in length.
44. 44. The method of any one of claims 35 to 43, wherein distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads comprises selecting a subset of consensus sequence reads that include DA junctions with allele lengths less than a threshold apparent insert size.
45. 45. The method of any one of claims 35 to 44, wherein distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads comprises inferring the fragment size of each read pair in the subset and comparing it one-by-one to a threshold apparent insert size, or comparing the distribution of inferred fragment sizes to the distribution of insert sizes in the entire library.
46. 46. The method of any one of claims 35 to 45, wherein distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads comprises comparing the inferred fragment size of any one consensus read pair to the allele size of that read pair.
47. 47. The method of any one of claims 35 to 46, wherein distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads comprises: selecting a subset of consensus sequencing reads containing DA junctions with allele lengths less than a threshold apparent insert size; inferring the fragment size of each read pair in the subset and comparing it to a threshold apparent insert size one by one, or comparing the distribution of inferred fragment sizes to the distribution of insert sizes in the entire library; and comparing the inferred fragment size of any one consensus read pair to the allele.
48. 48. The method of any one of claims 44 to 47, wherein the threshold apparent insert size is at least 20 bp in length.
49. 48. The method of any one of claims 44 to 47, wherein the deduced fragment size is at least 20 bp in length.
50. 48. The method of any one of claims 44 to 47, wherein the putative eccDNA is identified as having the putative fragment size less than or equal to the allele length.
51. 51. The method of claim 50, wherein the allele length is at least 20 bp.
52. 36. The method of claim 35, comprising obtaining multiple sequence reads of the sequenced double-stranded DNA using duplex sequencing. identifying a subset of sequence reads from a plurality of sequence reads, wherein the sequence reads of the subset of sequence reads each independently comprise a reference allele junction; selecting a subset of consensus sequence reads comprising a DA junction from the subset of sequence reads, comprising distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads, wherein distinguishing the putative eccDNA sequence reads includes selecting a subset of consensus sequence reads comprising a DA junction, wherein the allele length is smaller than a threshold apparent insert size; estimating the fragment size of each read pair in the subset and comparing it one by one with a threshold apparent insert size, or comparing the distribution of the estimated fragment sizes with the distribution of insert sizes of the entire library; and comparing the estimated fragment size of any one consensus read pair with the allele; and identifying eccDNA according to the distinguished putative eccDNA sequence reads.
53. A method for assessing potential clastogenicity, comprising the steps of: obtaining double-stranded DNA comprising putative extrachromosomal circular DNA (eccDNA) from one or more cells; performing duplex sequencing of the double-stranded DNA to obtain a plurality of sequence reads of the double-stranded DNA; identifying a subset of sequence reads from the plurality of sequence reads, wherein each sequence read of the subset of sequence reads independently comprises a reference allele junction; distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads among the subset of sequence reads; determining an eccDNA profile according to the identified putative eccDNA sequence reads; and assessing potential clastogenicity according to the determined eccDNA profile.
54. 54. The method of claim 53, wherein the one or more cells are one of one or more cells exposed to a potentially clastogenic, control, or untreated cell.
55. 54. The method of claim 53, wherein the reference allele junction comprises the nucleic acid sequence DA.
56. 56. The method of claim 55, wherein the nucleic acid sequence DA comprises nucleic acid sequence D operably linked to nucleic acid sequence A in the 5' to 3' direction.
57. 57. The method of Claim 56, wherein nucleic acid sequence D is located downstream of nucleic acid sequence A at the reference genomic locus of the reference allele.
58. 58. The method of any one of claims 53 to 57, wherein the nucleic acid sequence DA is at least 1 bp in length.
59. 54. The method of claim 53, wherein the reference allele junction is due to an apparent indel or an apparent structural variant.
60. 60. The method of Claim 59, wherein the apparent indel is at least 1 bp in length.
61. 60. The method of claim 59, wherein the apparent structural variant is due to an apparent insertion or an apparent duplication.
62. 62. The method of claim 61, wherein the apparent structural variant due to the apparent insertion or apparent duplication is at least 20 bp in length.
63. 63. The method of any one of claims 53 to 62, wherein distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads comprises selecting a subset of consensus sequence reads that contain DA junctions with allele lengths smaller than a threshold apparent insert size.
64. 64. The method of any one of claims 53 to 63, wherein distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads comprises inferring the fragment size of each read pair in the subset and comparing it one-by-one to a threshold apparent insert size, or comparing the distribution of inferred fragment sizes to the distribution of insert sizes in the entire library.
65. 65. The method of any one of claims 53 to 64, wherein distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads comprises comparing the putative fragment size of any one consensus read pair to the allele size of that read pair.
66. The method of any one of claims 53 to 65, wherein distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads comprises the steps of: selecting a subset of consensus sequence reads containing DA junctions and having allele lengths smaller than a threshold apparent insert size; estimating the fragment size of each read pair in the subset and comparing it to a threshold apparent insert size one by one or comparing the distribution of estimated fragment sizes to the distribution of insert sizes in the entire library; and comparing the estimated fragment size of any one consensus read pair to the allele.
67. 67. The method of any one of claims 62 to 66, wherein the threshold apparent insert size is at least 20 bp in length.
68. 67. The method of any one of claims 62 to 66, wherein the predicted fragment size is at least 20 bp in length.
69. 67. The method of any one of claims 62 to 66, wherein the putative eccDNA is identified as having a putative fragment size equal to or less than the allele length.
70. 70. The method of claim 69, wherein the allele length is at least 20 bp.
71. 54. The method of Claim 53, wherein assessing the potential clastogen further comprises comparing the eccDNA profile from one or more cells exposed to the potential clastogen with a control or untreated sample from the same cohort.
72. 54. The method of claim 53, wherein the method comprises obtaining double-stranded DNA comprising putative extrachromosomal circular DNA (eccDNA) from one or more cells exposed to the potential clastogen. performing duplex sequencing of the double-stranded DNA to obtain a plurality of sequence reads of the double-stranded DNA; identifying a subset of sequence reads from the plurality of sequence reads, wherein each sequence read of the subset of sequence reads independently comprises a reference allele junction; discriminating putative eccDNA sequence reads from among the subset of sequence reads from chromosomal tandem duplication sequence reads, wherein discriminating the putative eccDNA sequence reads comprises selecting a subset of consensus sequence reads comprising a DA junction having an allele length less than a threshold apparent insert size; and comparing the estimated fragment size of any one consensus read pair with an allele; determining an eccDNA profile according to the discriminative predicted eccDNA sequence reads; and evaluating the clastogenicity of a potential clastogen according to the determined eccDNA profile, wherein the method further comprises comparing the eccDNA profile from one or more cells exposed to the potential clastogen with a control or untreated sample from the same cohort.
73. 73. The method of any one of Claims 53 or 72, wherein said profile comprises any one or any combination of the quantity, frequency, quality, size, genomic location, or any other characteristic of said eccDNA.
74. 74. The method of Claim 73, wherein said profile comprises a frequency of said eccDNA.
75. 73. The method of any one of Claims 53 or 72, wherein the method comprises distinguishing tumor DNA from paired normal DNA based on the frequency of eccDNA in the tumor.
76. 73. The method of any one of Claims 53 or 72, wherein the method further comprises performing eccDNA enrichment.
77. 77. The method of Claim 76, wherein said step of performing eccDNA enrichment comprises performing size selection. performing an exonuclease treatment; and / or using a DNA binding protein that differentially binds to or remains bound to double-stranded circular DNA molecules compared to double-stranded linear DNA molecules.
78. 78. The method of claim 77, wherein said size selection comprises the use of paramagnetic beads at a size threshold of about 10,000 bp, electrophoresis, column filtration, density gradient centrifugation, or selective extraction.
79. 78. The method of Claim 77, wherein the exonuclease is selected from any one of exonuclease I, exonuclease T, exonuclease VII, exonuclease III, T7 exonuclease, exonuclease V (RecBCD), exonuclease VIII, lambda exonuclease, T5 exonuclease, and any combination thereof.
80. 80. The method of any one of claims 53-79, wherein the potential clastogen is a chemical compound, a physical exposure, a biological agent, or a complex mixture and / or an environmental exposure.
81. A method for assessing the clastogenicity of a potential clastogen, comprising obtaining double-stranded DNA from one or more cells; performing duplex sequencing of the double-stranded DNA to obtain a plurality of sequence reads of the double-stranded DNA; identifying a subset of sequence reads from the plurality of sequence reads, wherein the sequence reads of the subset of sequence reads each independently comprise a reference allele junction; distinguishing, from the subset of sequence reads, putative eccDNA sequence reads from chromosomal tandem duplication sequence reads; and assessing the clastogenicity potential according to the distinguished putative eccDNA.
82. 82. The method of Claim 81, wherein said one or more cells do not contain eccDNA.
83. 82. The method of claim 81 , wherein the one or more cells are one of one or more cells exposed to a clastogen, a control cell, or an untreated cell.
84. 82. The method of Claim 81, wherein the reference allele junction comprises the nucleic acid sequence DA.
85. 85. The method of claim 84, wherein nucleic acid sequence DA comprises nucleic acid sequence D operably linked to nucleic acid sequence A in the 5' to 3' direction.
86. 86. The method of claim 85, wherein nucleic acid sequence D is located downstream of nucleic acid sequence A at the reference genomic locus of the reference allele.
87. 87. The method of any one of claims 84 to 86, wherein the length of the nucleic acid sequence DA is at least 1 bp.
88. 88. The method of claim 87, wherein the reference allele junction is due to an apparent indel or an apparent structural variant.
89. 89. The method of claim 88, wherein the apparent indel is at least 1 bp in length.
90. 88. The method of claim 87, wherein the apparent structural variant is due to an apparent insertion or an apparent duplication.
91. 89. The method of Claim 88, wherein the apparent insertion or apparent duplication apparent structural variant is at least 20 bp in length.
92. 92. The method of any one of claims 81 to 91, wherein distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads comprises selecting a subset of consensus sequence reads containing DA junctions with allele lengths less than a threshold apparent insert size; estimating the fragment size of each read pair in the subset and comparing it one-by-one to a threshold apparent insert size, or comparing the distribution of estimated fragment sizes to the distribution of insert sizes in the entire library; and comparing the estimated fragment size of any one consensus read pair to the allele.
93. 93. The method of any one of claims 81 to 92, wherein the threshold apparent insert size is at least 20 bp in length.
94. 94. The method of any one of claims 81 to 93, wherein the predicted fragment size is at least 20 bp in length.
95. 82. The method of Claim 81, wherein assessing the clastogenicity of the potential clastogen further comprises comparing the eccDNA profile from one or more cells exposed to the potential clastogen with a control sample or an untreated sample from the same cohort.
96. 82. The method of claim 81, wherein the method comprises obtaining double-stranded DNA from one or more cells. Performing duplex sequencing of the double-stranded DNA to obtain a plurality of sequence reads of the double-stranded DNA; identifying a subset of sequence reads from the plurality of sequence reads, wherein each sequence read of the subset of sequence reads independently comprises a reference allele junction; distinguishing putative eccDNA sequence reads from the subset of sequence reads from chromosomal tandem duplication sequence reads, wherein distinguishing the putative eccDNA sequence reads comprises selecting a subset of consensus sequence reads comprising a DA junction with an allele length shorter than a threshold apparent insert size; estimating the fragment size of each read pair in the subset and comparing it one by one with a threshold apparent insert size or comparing the distribution of the estimated fragment sizes with the distribution of insert sizes of the entire library; and comparing the estimated fragment size of any one consensus read pair with the allele; evaluating the potential clastogenicity according to the identified putative eccDNA.
97. 82. The method of Claim 81, wherein said method further comprises performing eccDNA enrichment.
98. 98. The method of Claim 97, wherein said enriching for eccDNA comprises performing size selection; performing exonuclease treatment; and / or using a DNA-binding protein that differentially binds to or remains bound to double-stranded circular DNA molecules compared to double-stranded linear DNA molecules.
99. 99. The method of claim 98, wherein the size selection comprises the use of paramagnetic beads, electrophoresis, column filtration, density gradient centrifugation, or selective extraction at a size threshold of about 10,000 bp.
100. 99. The method of Claim 98, wherein the exonuclease is selected from any one of exonuclease I, exonuclease T, exonuclease VII, exonuclease III, T7 exonuclease, exonuclease V (RecBCD), exonuclease VIII, lambda exonuclease, T5 exonuclease, and any combination thereof.
101. 101. The method of any one of claims 81-100, wherein the potential clastogen is a chemical compound, a physical exposure, a biological agent, or a complex mixture and / or an environmental exposure.
102. A method for evaluating genotoxicity, comprising the steps of: a) evaluating clastogenicity: obtaining double-stranded DNA from one or more cells exposed to a potential genotoxin; performing duplex sequencing of the double-stranded DNA to obtain a plurality of sequence reads of the double-stranded DNA; identifying a subset of sequence reads from the plurality of sequence reads, wherein the sequence reads of the subset of sequence reads each independently include a reference allele junction; discriminating putative eccDNA sequence reads from the subset of sequence reads from chromosomal tandem duplicated sequence reads; determining an eccDNA profile according to the identified putative eccDNA sequence reads; and evaluating the clastogenicity of the potential genotoxin according to the determined eccDNA profile; and b) evaluating mutagenicity: obtaining double-stranded DNA from one or more cells exposed to the potential genotoxin; performing duplex sequencing of the double-stranded DNA to obtain a plurality of sequence reads of the double-stranded DNA; determining a mutational profile of the double-stranded DNA; and evaluating the mutagenicity of the potential genotoxin according to the determined mutational profile of the double-stranded DNA.
103. 103. The method of claim 102, wherein the reference allele junction comprises the nucleic acid sequence DA.
104. 104. The method of claim 103, wherein nucleic acid sequence DA comprises nucleic acid sequence D operably linked to nucleic acid sequence A in the 5' to 3' direction.
105. 105. The method of claim 104, wherein nucleic acid sequence D is located downstream of nucleic acid sequence A at the reference genomic locus of the reference allele.
106. 106. The method of any one of claims 103 to 105, wherein the nucleic acid sequence DA is at least 1 base pair (bp) in length.
107. 103. The method of claim 102, wherein the reference allele junction is due to an apparent indel or an apparent structural variant.
108. 108. The method of claim 107, wherein the apparent indel is at least 1 bp in length.
109. 108. The method of claim 107, wherein the apparent structural variant is due to an apparent insertion or an apparent duplication.
110. 110. The method of claim 109, wherein the apparent insertion or the apparent structural variant due to the apparent duplication is at least 20 bp in length.
111. 111. The method of any one of claims 102-110, wherein distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads comprises selecting a subset of consensus sequence reads that include DA junctions with allele lengths less than a threshold insert size.
112. 113. The method of any one of claims 102-112, wherein distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads comprises inferring a fragment size for each read pair in the subset and comparing it pairwise to a threshold insert size, or comparing the distribution of inferred fragment sizes to the distribution of insert sizes for the entire library.
113. 114. The method of any one of claims 102-113, wherein distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads comprises comparing the inferred fragment size of any one consensus read pair to the allele size of that read pair.
114. 114. The method of any one of claims 102-113, wherein distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads comprises selecting a subset of consensus sequence reads containing DA junctions with allele lengths less than a threshold insert size; inferring the fragment size of each read pair in the subset and comparing it pairwise to a threshold insert size, or comparing the distribution of inferred fragment sizes to the distribution of insert sizes across the entire library; and comparing the inferred fragment size of any one consensus read pair to the allele.
115. 115. The method of any one of claims 111 to 114, wherein the threshold apparent insert size is at least 20 bp in length.
116. 116. The method of any one of claims 111 to 115, wherein the predicted fragment size is at least 20 bp in length.
117. 117. The method of any one of claims 111 to 116, wherein the putative eccDNA is identified as having a putative fragment size equal to or less than the allele length.
118. 118. The method of claim 117, wherein the allele length is at least 20 bp.
119. 103. The method of Claim 102, wherein said profile comprises any one or any combination of the quantity, frequency, quality, size, genomic location, or any other characteristic of said eccDNA.
120. 120. The method of Claim 119, wherein said profile comprises a frequency of said eccDNA.
121. 121. The method of Claim 120, wherein said method further comprises performing eccDNA enrichment.
122. 103. The method of claim 102, wherein the method comprises the step of: a) assessing clastogenicity, the method comprising the step of obtaining double-stranded DNA from one or more cells exposed to a potential genotoxin. performing duplex sequencing of the double-stranded DNA to obtain a plurality of sequence reads of the double-stranded DNA; identifying a subset of sequence reads from the plurality of sequence reads, wherein the sequence reads of the subset of sequence reads each independently include a reference allele junction; distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads among the subset of sequence reads; determining an eccDNA profile according to the identified putative eccDNA sequence reads, wherein distinguishing the putative eccDNA sequence reads includes the steps of: selecting a subset of consensus sequence reads including DA junctions having an allele length less than a threshold insert size; inferring the fragment size of each read pair in the subset and comparing it one by one with a threshold insert size, or comparing the distribution of the inferred fragment sizes with the distribution of insert sizes of the entire library; and comparing the inferred fragment size of any one consensus read pair with the allele; and evaluating the clastogenicity of the potential genotoxin according to the determined eccDNA profile; and Steps including assessing mutagenicity: obtaining double-stranded DNA from one or more cells exposed to a potential genotoxin; performing duplex sequencing of the double-stranded DNA to obtain multiple sequence reads of the double-stranded DNA; determining a mutation profile of the double-stranded DNA; and assessing the mutagenicity of the potential genotoxin according to the determined mutation profile of the double-stranded DNA.
123. The method of any one of claims 102 or 122, wherein the eccDNA enrichment performed comprises performing size selection; performing exonuclease treatment; and / or using a DNA binding protein that differentially binds to or remains bound to double-stranded circular DNA molecules compared to double-stranded linear DNA molecules.
124. 124. The method of claim 123, wherein size selection comprises the use of paramagnetic beads at a size threshold of about 10,000 bp, electrophoresis, column filtration, density gradient centrifugation, or selective extraction.
125. 124. The method of Claim 123, wherein the exonuclease is selected from any one of exonuclease I, exonuclease T, exonuclease VII, exonuclease III, T7 exonuclease, exonuclease V (RecBCD), exonuclease VIII, lambda exonuclease, T5 exonuclease, and any combination thereof.
126. 126. The method of any one of claims 102-125, wherein the xenobiotic is selected from any one of environmental pollutants, hydrocarbons, food additives, oil mixtures, pesticides, other xenobiotics, synthetic polymers, carcinogens, drugs, antioxidants, and combinations thereof.
127. A method for assessing cancer risk in a sample, the method comprising: obtaining double-stranded DNA containing putative extrachromosomal circular DNA (eccDNA) from the sample; performing duplex sequencing of the double-stranded DNA to obtain a plurality of sequence reads of the double-stranded DNA; identifying a subset of sequence reads from the plurality of sequence reads, wherein the sequence reads of the subset of sequence reads each independently comprise a reference allele junction; distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads among the subset of sequence reads; determining an eccDNA profile according to the distinguished putative eccDNA sequence reads; and assessing the cancer risk of the sample according to the determined eccDNA profile.
128. 128. The method of claim 127, wherein the reference allele junction comprises the nucleic acid sequence DA.
129. 129. The method of claim 128, wherein nucleic acid sequence DA comprises nucleic acid sequence D operably linked to nucleic acid sequence A in the 5' to 3' direction.
130. 130. The method of claim 129, wherein nucleic acid sequence D is located downstream of nucleic acid sequence A at the reference genomic locus of the reference allele.
131. 131. The method of any one of claims 128 to 130, wherein the nucleic acid sequence DA is at least 1 base pair (bp) in length.
132. 128. The method of claim 127, wherein the reference allele junction is due to an apparent indel or an apparent structural variant.
133. 133. The method of claim 132, wherein the apparent indel is at least 1 bp in length.
134. 133. The method of claim 132, wherein the apparent structural variant is due to an apparent insertion or an apparent duplication.
135. 135. The method of Claim 134, wherein the apparent structural variant due to the apparent insertion or apparent duplication is at least 20 bp in length.
136. 136. The method of any one of claims 127-135, wherein distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads comprises selecting a subset of consensus sequence reads that include DA junctions with allele lengths less than a threshold insert size.
137. 138. The method of any one of claims 127-137, wherein distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads comprises inferring a fragment size for each read pair in the subset and comparing it pairwise to a threshold insert size, or comparing the distribution of inferred fragment sizes to the distribution of insert sizes for the entire library.
138. 139. The method of any one of claims 127-138, wherein distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads comprises comparing the inferred fragment size of any one consensus read pair to the allele size of that read pair.
139. 139. The method of any one of claims 127-138, wherein distinguishing putative eccDNA sequence reads from chromosomal tandem duplication sequence reads comprises selecting a subset of consensus sequence reads containing DA junctions with allele lengths less than a threshold insert size; inferring the fragment size of each read pair in the subset and comparing it pairwise to a threshold insert size, or comparing the distribution of inferred fragment sizes to the distribution of insert sizes across the entire library; and comparing the inferred fragment size of any one consensus read pair to the allele.
140. 140. The method of any one of claims 139 to 139, wherein the threshold apparent insert size is at least 20 bp in length.
141. 140. The method of any one of claims 136 to 139, wherein the predicted fragment size is at least 20 bp in length.
142. 142. The method of any one of claims 127 to 141, wherein the putative eccDNA is identified as having a putative fragment size equal to or less than the allele length.
143. 143. The method of claim 142, wherein the allele length is at least 20 bp.
144. 128. The method of Claim 127, wherein said profile comprises any one or any combination of the quantity, frequency, quality, size, genomic location, or any other characteristic of said eccDNA.
145. 145. The method of Claim 144, wherein said profile comprises a frequency of said eccDNA.
146. The method of claim 127, wherein the method comprises the steps of obtaining double-stranded DNA containing putative extrachromosomal circular DNA (eccDNA) from the sample, performing duplex sequencing of the double-stranded DNA to obtain a plurality of sequence reads of the double-stranded DNA; identifying a subset of sequence reads from the plurality of sequence reads, wherein each sequence read of the subset of sequence reads independently comprises a reference allele junction; selecting a subset of consensus sequence reads from the subset of sequence reads that comprises a DA junction having an allele length less than a threshold insert size, comprising distinguishing putative eccDNA sequence reads from chromosomal tandem duplicated sequence reads, wherein the putative eccDNA sequence reads are identified; estimating the fragment size of each read pair in the subset and comparing it one by one with a threshold insert size, or comparing the distribution of the estimated fragment sizes with the distribution of insert sizes of the entire library; and comparing the estimated fragment size of any one consensus read pair with an allele; determining an eccDNA profile according to the identified putative eccDNA sequence reads; and assessing the cancer risk of the sample according to the measured eccDNA profile.
147. 147. The method of any one of Claims 127 or 146, wherein said method further comprises performing eccDNA enrichment.
148. 148. The method of Claim 147, wherein said eccDNA enrichment comprises performing size selection; performing exonuclease treatment; and / or using a DNA binding protein that differentially binds to or remains bound to double-stranded circular DNA molecules compared to double-stranded linear DNA molecules.
149. 149. The method of claim 148, wherein said size selection comprises the use of paramagnetic beads at a size threshold of about 10,000 bp, electrophoresis, column filtration, density gradient centrifugation, or selective extraction.
150. 149. The method of Claim 148, wherein the exonuclease is selected from any one of exonuclease I, exonuclease T, exonuclease VII, exonuclease III, T7 exonuclease, exonuclease V (RecBCD), exonuclease VIII, lambda exonuclease, T5 exonuclease, and any combination thereof.
151. 128. The method of claim 127, wherein the sample is selected from the group consisting of a sperm sample, a semen sample, a prostatic fluid sample, a testicular biopsy sample, a spermatogonial sample, a germ cell sample, a gamete sample, a swab, a lavage, an aspirate, a biopsy, a tissue sample, a tumor sample, a pre-neoplastic sample, a liquid biopsy, a hyperplasia sample, a hypertrophy sample, a dysplasia sample, a urine sample, a CSF sample, any other body fluid sample, an autopsy sample, a post-mortem sample, a surgical sample, a model organism sample, a plasma sample, a serum sample, a stomach sample, a bone marrow sample, a stool sample, a brushing sample, a bile sample, a pancreatic juice sample, a synovial fluid sample, a sputum sample, a mucus sample, a vitreous sample, a forensic sample, an environmental sample, a bacterial sample, a fungal sample, a mammalian sample, a human sample, and a diagnostic sample.
152. 128. The method of Claim 127, wherein the sample is a cancer sample or a healthy sample.
153. 128. The method of Claim 127, wherein assessing the cancer risk of the sample according to the determined profile of eccDNA further comprises comparing the eccDNA profile with known eccDNA profiles.