Preparation and compositions of cot-1 nucleic acid

WO2025123005A3PCT designated stage expired Publication Date: 2025-07-31ILLUMINA INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/059144
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-07
Filing Date
2024-12-09
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing nucleic acid analysis techniques face challenges with non-specific binding of hybridization probes, leading to off-target capture and compromised sequencing results.

Method used

The method involves generating Cot-1 DNA by contacting genomic DNA fragments with a hybridization probe panel, separating bound and unbound portions, and providing the unbound portion as Cot-1 DNA, which can be further purified using solid-phase reversible immobilization beads.

Benefits of technology

This approach effectively reduces non-specific binding and off-target capture, enhancing the specificity and accuracy of nucleic acid sequencing by providing a hybridization blocker enriched with repetitive elements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024059144_31072025_PF_FP_ABST
    Figure US2024059144_31072025_PF_FP_ABST
Patent Text Reader

Abstract

Presented herein are Cot-1 methods and compositions that may be used for reducing non-specific binding, such as enhancing specific enrichment of target sequences in a nucleic acid library. In an embodiment, a Cot-1 composition may be a modified Cot-1 composition that has sequences with similarity to coding sequences removed. In an embodiment, a Cot-1 composition may be a synthetic Cot-1 composition. In an embodiment, a Cot-1 composition may be a minimal or reduced Cot-1 composition.
Need to check novelty before this filing date? Find Prior Art

Description

PREPARATION AND COMPOSITIONS OF COT-1 NUCLEIC ACIDCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] The present application claims priority to and the benefit of U.S. Provisional Application No. 63 / 607,488, filed on December 7, 2024, which is incorporated by reference herein in its entirety.BACKGROUND

[0002] The present disclosure relates generally to Cot-1 hybridization blockers, such as those for reducing off-target capture of nucleic acids in various assays.

[0003] The subject matter discussed in this section should not be assumed to be prior art merely as a result of its mention in this section. Similarly, a problem mentioned in this section or associated with the subject matter provided as background should not be assumed to have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which in and of themselves can also correspond to implementations of the claimed technology.

[0004] Certain analysis techniques rely on hybridization of nucleic acids to their complements. For example, targeted sequencing methods use a panel or set of probes that hybridize to target sequences in the nucleic acid library, which may also contain off-target sequences that are not of interest. Hybridization of the probes to the target sequences allows these sequences to be separated from the rest of the fragments in the library for sequencing. By targeting only a portion of the nucleic acid library, hybrid capture methods avoid sequencing of the off-target nucleic acid fragments that do not contain sequences of interest. Off-target reads not only waste sequencing yield, but also potentially compromise variant calling for somatic mutations of low frequency. However, certain probes may be associated with non-specific binding. In such cases, hybridization blockers, such as Cot-1 DNA, may be used to reduce non-specific binding.BRIEF SUMMARY

[0005] Provided herein is a method of generating Cot-1 DNA that includes contacting genomic DNA fragments with a hybridization probe panel under conditions that permit hybridization of probes of the hybridization probe panel to the genomic DNA fragments, wherein the probes of hybridization probe panel are configured to hybridize to coding regions of the genomic DNA fragments; separating a first portion of the genomic DNA fragments from a second portion of the genomic DNA fragments based on differential binding to the hybridization probe panel, wherein the first portion is not bound to the probes and the second portion is bound to the probes; and providing the first portion of the genomic DNA fragments as Cot-1 DNA.

[0006] Provided herein is also a method of generating Cot-1 DNA that includes providing genomic DNA fragments; exposing the genomic DNA fragments to a denaturing temperature to yield denatured DNA; allowing the denatured DNA to reanneal to yield reannealed DNA; and purifying the reannealed DNA using solid-phase reversible immobilization beads; and providing the purified DNA as Cot-1 DNA.

[0007] Provided herein is also a method of providing a minimal Cot-1 composition that includes providing probe sequence information of a panel of probes configured to hybridize to targeted regions of genomic DNA; identifying, in the probe sequence information, sequences that overlap with repetitive element sequences in the genomic DNA; and providing a plurality of oligonucleotides having a length of 60-100 nucleotides that correspond to the repetitive elements sequences as the minimal Cot-1 composition.

[0008] Provided herein is also a Cot DNA composition prepared by a method comprising: contacting genomic DNA fragments with a hybridization probe panel under conditions that permit hybridization of probes of the hybridization probe panel to the genomic DNA fragments, wherein the probes of hybridization probe panel are configured to hybridize to coding regions of the genomic DNA fragments; separating a first portion of the genomic DNA fragments from a second portion of the genomic DNAfragments based on differential binding to the hybridization probe panel, wherein the first portion is not bound to the probes and the second portion is bound to the probes; and providing the first portion of the genomic DNA fragments as the Cot DNA composition.

[0009] Provided herein is also a synthetic Cot-1 DNA composition comprising any combination of SEQ ID NOS: 1-1364.

[0010] Provided herein is also a synthetic Cot-1 DNA composition comprising a plurality of oligonucleotides having a length of 60-100 nucleotides and comprising SEQ ID NOS: 1-1364.

[0011] The details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0001] FIG. 1 is a schematic illustration of an example target enrichment sequencing technique with or without a Cot-1 hybridization blocker, according to an embodiment of the present disclosure;

[0002] FIG. 2 is a flow diagram of a Cot-1 preparation technique, according to an embodiment of the present disclosure;

[0003] FIG. 3 is a schematic illustration of a Cot-1 preparation technique, according to an embodiment of the present disclosure;

[0004] FIG. 4 is a flow diagram of a Cot-1 preparation technique, according to an embodiment of the present disclosure;

[0005] FIG. 5 is a schematic illustration of a Cot-1 preparation technique, according to an embodiment of the present disclosure;

[0006] FIG. 6 is a flow diagram of a Cot-1 preparation technique, according to an embodiment of the present disclosure;

[0007] FIG. 7 is a schematic illustration of a Cot-1 preparation technique, according to an embodiment of the present disclosure e;

[0008] FIG. 8 shows padded read enrichment data for a minimal Cot-1 composition as discuss herein; and

[0009] FIG. 9 is a block diagram of a sequencing device configured to acquire sequencing data, according to an embodiment.DETAILED DESCRIPTION

[0010] Hybridization probes may have imperfect specificity for their nucleic acid targets. For example, in an exome sequencing reaction, certain hybridization probes may pull down intronic or intergenic sequences from a nucleic acid library along with target sequences. These off-target fragments, once pulled down, are then present in the pool of nucleic acid fragments that are sequenced, which results in incorrect sequences being present in the sequence data. The use of hybridization blockers, such as Cot-1 DNA, may block certain types of off-target binding.

[0011] FIG. 1 is a schematic representation of a target enrichment capture workflow with or without Cot-1 to reduce off-target binding. The workflow includes preparation of a nucleic acid library 30 formed from a plurality of nucleic acid fragments 32 from the sample, such as a sample including genomic DNA (e.g., a human genome, an animal genome, a bacterial genome) or other nucleic acids. The nucleic library includes fragments having sequences from regions of interest (e g., fragment 32a) that include target sequences (e.g., target sequence 14) as well as fragments that are off-target (e.g., fragment 32b) having only off-target sequences. Among the off-target fragments 32b may be certain repetitive sequence fragments 32c that may be subject to off-target binding. It should be understood that fragments 32a that include target sequences maybe formed entirely from regions of interest or may include other regions that are not of interest such that the library 30 may include mixed target fragments 32d.

[0012] The target-specific hybridization probes 20 are designed to be complementary to one or more target sequences 14 on the fragments 32s. Accordingly, under hybridization conditions, one or more target-specific hybridization probes 20 will bind to the complementary target sequences 14. This facilitates separation, using probe- linked affinity tags 40, of fragments 32a that have the target sequences 14 from the fragments 32b that do not have the target sequences 14 to create a target-enriched sample for sequencing. However, among the probes 20, there may be a subset of probes 20a with high specificity for target sequences 14 and that do not bind to off-target sequences in the fragments 32b. Another subset of the probes 20b may be a nonspecific subset that has some non-specific binding to sequences present in the repetitive sequence fragments 32c and / or the mixed fragments 32d.

[0013] Accordingly, as shown in panel 50, when the probes 20 are contacted with strands of the fragments 32, the specific subset 20a binds to on-target fragments 32a with the non-specific subset has at least some binding to the repetitive sequence fragments 32c. Thus, at a separation step (e.g., using beads 56 with affinity tag binders 58), not only are the target fragments 32a captured, but there is also some capture of off-target repetitive sequence fragments 32c. In addition, in some cases, the beads 56 may have some non-specific binding ability to off-target fragments 32b. Accordingly, after separation, the captured strands include strands from on-target fragments 32a, which may include mixed fragments 32d as well as from off-target fragments 32b, including repetitive sequence fragments 32c. These undesired strands from off-target fragments 32b can then enter the subsequent sequencing steps of the workflow to generate undesired sequence data.

[0014] In contrast, when a Cot-1 hybridization blocker is provided in the workflow, as shown in panel 52, the Cot-1 DNA interacts with the non-specific subset of probes 20b to reduce off-target fragment capture. In addition, the Cot-1 DNA may interact directly with the beads 56 to reduce another source of non-specific binding. Thus, in theillustrated example, the beads 56 provide captured strands from on-target fragments 32a with a reduction in a presence of off-target fragment strands from off-target fragments 32b.

[0015] Cot-1 is a useful hybridization blocker that reduces non-specific binding in target enrichment and other probe-based workflows, such as comparative genomic hybridization (CGH), microarray assays, hybridization pulldowns, in situ hybridization (e.g., FISH), . Cot-1 is prepared by denaturing and renaturing large amounts of genomic DNA fragments, followed by purification of the initial duplex DNA fraction. Cot-1 is thus the fraction of the genome that is generally rich in repetitive elements and the like. Cot-1 DNA can be prepared from a eukaryotic (e.g., mammalian, human) genomic DNA source, such as placental DNA, by shearing, denaturing, and reannealing under conditions that enrich repetitive elements. The Cot-1 fraction of human genomic DNA consists largely of rapidly annealing repetitive elements. Repetitive genomic DNA sequences can make up a large fraction of a source genome. For example, more than 40% of the human genome is estimated to be these repetitive sequences. While the fraction may vary in other eukaryotes, many genomes also contain large fractions of these repetitive sequences.

[0016] Repetitive sequences may be characterized by their reannealing kinetics after denaturation, which are typically faster than those of low copy sequences. This may be a result of the lower complexity of the repetitive sequences. Cot or COt analysis is a technique that measures how much repetitive DNA is in a DNA sample such as a genome. Cot analysis involves heating a sample of genomic DNA to a denaturation temperature following by a temperature decrease to permit reannealing of the denatured strands and monitoring reannealing of the DNA. A sequence of single-stranded repetitive sequence DNA is more likely to find a complementary strand because of their relative abundance compared to lower copy sequences, such that common sequences renature more rapidly than rare sequences. The rate of reannealing is proportional to the number of copies of that sequence in the DNA sample, and a sample with a high fraction of repetitive sequences will renature rapidly, while complex sequences willrenature more slowly relative to one another. The amount of renaturation or reannealing is measured relative to a Cot or COt value.

[0017] For a concentration C of single-stranded (e.g., denatured) DNA at time t (e.g., time in seconds Ts), the loss of single-stranded DNA may be expressed as -dC / dt=kC2, where k is the rate constant in in liters (mole nt)-l sec-1. A Cot-1 value may be, in embodiments, expressed as Cot-1 = 1 = mol / Lx Ts. Thus, a particular sample of DNA (or nucleic acids) may have a Cot-1 value based on the initial concentration of DNA (calculated in moles of nucleotide per liter and assuming an average molecular weight of 339 g / mol). In one example starting at a DNA concentration of 300ng / L, for 1 mole of dNTP equal to an average of 339 g / mol, then moles of sheared DNA is equal to (0.300 g / L) / (339 g / L) = 8.85 x 10’4mol / L. Thus, for this example, the reannealing reaction must proceed for 1130 s in order for Cot to equal 1. However, the reannealing time may vary depending on the initial nucleic acid concentration and, in some cases, the GC composition of the source genomic DNA. The reannealing reaction can be stopped by alcohol precipitation. While certain embodiments of the disclosure relate to Cot-1 DNA having a COT value of 1, it should be understood that other Cot values may also be encompassed in the disclosed techniques (e.g., Cot-1.2, Cot-1.5, Cot-2). Further, Cot DNA may refer to a reannealed DNA composition as discussed herein with a Cot value associated with at least 50%, at least 80%, at least 95%, or at least 99% reannealing of the source DNA. In certain embodiments, the Cot DNA may be prepared from a genomic DNA source and / or may be synthetic. In an embodiment, the Cot DNA source is human placental DNA or mammalian placental DNA, such as male placental DNA.

[0018] Cot-1 prepared from genomic DNA may have variable quality and effectiveness as a hybridization blocker due to variability of the source genomic DNA, the reannealed fragment size, and / or the preparation workflow. Provided herein are Cot DNA (e.g., Cot-1 DNA) composition and preparation techniques that provide higher quality and / or more consistent Cot DNA. Cot DNA or Cot-1 DNA may refer to a heterogenous DNA composition. In an embodiment, Cot DNA is double-stranded or mostly double-stranded. The DNA composition may include mostly repetitive elements as generally discussed herein. Cot DNA may, in embodiments, refer to a particular preparation technique and the composition yielded from the preparation technique. Cot DNA may be non-naturally occurring. In one example, a Cot composition that includes DNA fragments within a particular size band (e.g., lOObp-lOOObp) and that is enriched for repetitive elements does not occur in nature and is generally from modification and / or manipulation of a genomic DNA source.

[0019] In one approach, Cot DNA may be improved by removing low copy sequences, such as coding or exome sequences, such that a reannealed fraction is artificially enriched for repetitive sequences. As discussed herein, repetitive elements in DNA fragments tend to reanneal more quickly, because they are more likely to find complements due to their relative abundance in the genome. Thus, by selecting appropriate annealing and reaction times, Cot DNA may be enriched for repetitive elements, while low copy sequences are less likely to be part of the reannealed fragment population. Nonetheless, some coding regions may reanneal to one another and form part of heterogeneous Cot DNA compositions. When Cot DNA is provided as a hybridization blocker, the Cot DNA is formed from an exogenous source. Therefore, it is not desirable to have coding regions in the Cot DNA, because these regions may bind to or associate with hybridization probes in a variety of assay contexts, decreasing the available probes to bind to target DNA or RNA of interest. Because hybridization blockers may be provided in excess concentrations relative to target nucleic acids, the depletion of probes may impact the signal able to be generated from coding regions of the target DNA or RNA, particularly for low concentration sequences. This is turn may impact the accuracy of sequencing results.

[0020] FIG. 2 is a flow diagram 100 of a Cot-1 preparation technique that includes an enrichment step to remove low copy sequences from reannealed genomic DNA fragments, such as coding sequences, during preparation. At step 104, genomic fragments, e.g., genomic DNA fragments, are provided. The fragmentation may be via shearing, sonication, acoustic shearing, enzymatic fragmentation, chemicalfragmentation, etc. The fragments may be of a suitable size, such as 5Obp to lOOObp. In an embodiment, the fragments may undergo a size selection step to exclude fragments outside of a desired size range.

[0021] At step 106, the DNA fragments are contacted with a set of probes for a removal step under conditions that permit hybridization. The conditions that permit hybridization may include an initial denaturing step to separate strands of doublestranded fragments (e.g., temperatures above 75°C, above 80°C, in a range of 75°C- 98°C) and contact with the probes at a subsequent hybridization temperature that is lower than the denaturing temperature. In an embodiment, the hybridization temperature is high enough to prevent non-specific binding. In an embodiment, the hybridization temperature is 45-60°C.

[0022] A variety of hybridization or washing conditions include variations in temperature or buffer composition and can be varied according to the specificity of the reaction needed. A range of stringency includes, for example, high, moderate or low stringency conditions. In an emobiment, stringent conditions are selected to be about 5- 10° C. lower than the thermal melting point (Tm) for the specific sequence at a defined ionic strength and pH. The Tmis the temperature, under defined ionic strength, pH and nucleic acid concentration, at which 50% of the probes complementary to the target hybridize to the target sequence at equilibrium. Differences in the number of hydrogen bonds as a function of base pairing between perfect matches and mismatches can be exploited as a result of their different Tms. Accordingly, a hybrid comprising perfect complementarity will melt at a higher temperature than one comprising at least one mismatch, all other parameters being equal. Stringent hybridization conditions may also include those in which the salt concentration is less than about 1.0 M sodium ion, generally about 0.01 to 1.0 M sodium ion concentration or other salts at pH 7.0 to 8.3 and the temperature is at least about 30° C. for short probes such as 10 to 50 nucleotides and at least about 60° C. for long probes such as greater than 50 nucleotides. Low stringency conditions include NaCl concentrations of about 1.0 M. Furthermore, low stringency conditions can include MgCh concentrations of about 10 mM, moderatestringency of about 1-10 mM, and high stringency conditions include concentrations of about 1 mM. Stringent conditions also can be achieved with the addition of helix destabilizing agents such as formamide. For example, low stringency conditions include formamide concentrations of about 0 to 10%, while high stringency conditions utilize formamide concentrations of about 40%.

[0023] In an embodiment, the probes are a panel of probes representative of coding regions in the genome, e.g., human coding regions, mammalian coding regions. For example, the probes may be a whole exome sequencing panel. In an embodiment, the probes may be a targeted panel of probes. In some cases, the panel may be a commercially available panel, such as whole exome or targeted sequencing panels available from Illumina.

[0024] After binding of the probes to coding sequencing still present in the DNA fragments, the probe-coding sequence complexes can be separated using probe separation techniques at step 110. For example, the probes may carry affinity tags (e.g., biotin), and the separation may be mediated by pulling the probe-coding sequence complexes out using affinity tag binders (e.g., avidin, streptavidin) linked to magnetic beads. This permits a separated of a probe-bound portion of the reaction mixture that includes the probe-coding sequence complexes from an unbound portion that includes coding sequence-depleted reannealed fragments that are not bound to probes. This unbound portion can be provided as Cot DA, e.g., Cot-1 DNA.

[0025] FIG. 3 is a schematic illustration of an embodiment of the process 100 showing fragment strands 150 (e.g., single strands of the fragments) generated from DNA fragments from a genomic DNA source. The strands 150 include a large portion of repetitive elements 156 and some presence of coding sequences 152. In an embodiment, the DNA fragments may be from a Cot DNA preparation. That is, the strands 150 may be generated from denaturing a Cot-1 DNA composition formed using conventional techniques, and the process 100 may be used to deplete coding sequences from prepared Cot-1 DNA to homogenize the composition and increase a fraction of repetitive elements 156 in the Cot-1 DNA after the process 100 is complete. In otherembodiments, the strands 150 may be generated as part of an initial Cot-1 preparation technique using genomic DNA to generate fragments and after the fragmentation step is performed. The strands may be generated using denaturing temperatures as discussed herein.

[0026] The strands 150 are exposed to reannealing temperatures, e.g., 65°C, in the presence of an appropriate buffer to promote hybridization of probes 162 to any individual strands 150 including coding sequences 152 as well as reannealing of repetitive elements 156 in the strands 150 to one another based on complementarity. The coding sequences 152 that are bound to the probes 162 are not available for reannealing. The probes 162 may include probes to coding sequences 152 as well as their complements. However, in an embodiment, any unbound strands 150 may be digested using a single-stranded nuclease (e.g., SI, Pl). The bound coding sequences 152 can be removed from the reannealed fragments 168 using probe affinity tags 164 coupled to beads (e.g., streptavidin magnetic beads, SMB), and the reannealed fragments 168 can be provided as the Cot-1 product 168. The reannealed fragments 168 may be precipitated as part of purification. Precipitation as discussed herein may refer to alcohol precipitation. In an embodiment, the ethanol precipitation is EtOH (0°C) / 3M NaOAC Precipitation, -20°C overnight or -80°C 1 hour.

[0027] FIG. 4 is directed to a process 200 for improving Cot DNA quality. In embodiments, the process 200 may be used alone or in combination with other techniques discussed herein. For example, steps of the process 200 may be incorporated into the process 100 (FIG. 2). The process 200 may be used as part of a Cot-1 preparation to improve Cot-1 purity. The process 200 includes providing genomic DNA fragments at step 204. The fragments may be provided as generating discussed herein, via fragmentation of a genomic DNA source. The fragments are denatured at denaturing temperatures (e.g., above 90°C, 92°C-100°C for 10 minutes), and then allowed to reanneal under conditions that promote hybridization of repetitive elements (e.g., 60°C-65°C, 60-120 minutes, reaction mixture concentration of 0.3-0.5M NaCl) at step 210. The reannealed DNA can be purified from single-stranded DNA inthe reaction at step 212 using solid phase reversible immobilization beads (SPRI beads). SPRI beads function to preferentially bind to double-stranded or reannealed DNA relative to single-stranded DNA. Additionally, the single-stranded DNA may be digested using SI or Pl nuclease that targets single- stranded nucleic acids. Both SI and Pl nucleases cleave RNA and single stranded DNA with no base specificity . SI nuclease is more active on DNA than RNA. Its reaction products are oligonucleotides or single nucleotides with 5' phosphoryl groups. Pl nuclease cleaves its substrate at every position yielding nucleoside 5' monophosphates, and it does not recognize or act on double-stranded DNA . Thus, using bead-based separation (e.g., magnetic pulldown and washing of supernatant), the purified reannealed DNA can be provided as Cot-1 DNA at step 216. For example, the reannealed DNA can be separated and precipitated. In an embodiment, the SPRI beads can be tuned to bind fragments of a particular size that corresponds to the genomic DNA fragments.

[0028] In another example, a Cot-1 preparation technique may include process steps that amplify desired fractions of Cot-1 to enrich the prepared Cot-1 for sequences or repetitive elements of interest. FIG. 5 is a schematic illustration of an embodiment of an enrichment workflow to generate Cot-1 DNA from genomic DNA 250. The genomic DNA 250 is fragmented and denatured to generate fragment strands 256 (e.g., single strands of the fragments). The strands 256 include a large portion of repetitive elements that reanneal to generate reannealed fragments 258 and with some non-annealed single fragment strands 260 remaining after the strands 256 are exposed to reannealing temperatures, e.g., 65°C, in the presence of an appropriate buffer to promote renaturing. The non-annealed single fragment strands 260 may be digested using a single-stranded nuclease (e.g., SI, Pl). The reannealed fragments 258 can be amplified to generate amplified Cot-1 264 prior to precipitation. The amplification may be via random hexamer primers. In such an embodiment, the disparity between the repetitive element fraction and less abundant fractions in the reannealed fragments 258 can be exponentially increased such that the amplified Cot-1 264 is artificially enriched for the repetitive elements after amplification. In an embodiment the reannealed fragments 258 can be adapterized via library preparation techniques to attach specific adapters (e.g.,Illumina sequencing adapters) via ligation and / or tagmentation as generally discussed in US Patent No. 11788139, which is hereby incorporated by reference herein. The adapters may include universal primer regions that can be used to amplify adapterized reannealed fragments 258 to generate the amplified Cot-1 264. In an embodiment the reannealed fragments 258 can be amplified using targeted primers designed in silico using known bioinformatic sequences for repetitive elements. The amplified Cot-1 264 may be processed and / or precipitated as generally discussed herein to remove components of the amplification reactions to generated purified amplified Cot-1 264. The amplified Cot-1 264 is a non-naturally occurring composition with a fractional distribution of repetitive elements that is modified or skewed outside of naturally occurring distributions.

[0029] A proposed function of Cot-1 in next generation sequencing (NGS) target enrichment assays is to prevent off-target capture of non-desirable repetitive elements. In some enrichment panels, probes with sequence identity to repetitive elements are unintentionally present. Cot-1 DNA, being composed entirely of repetitive elements, will hybridize to these unintentional probes and reduce or eliminate their ability to hybridize to and capture library fragments composed of repetitive sequence.

[0030] If enrichment probes targeting repetitive sequences are identified, there are various strategies to resolve their function. In one example, the probes having the unintentional repetitive element probe(s) can be removed from the panel. However, this is not applicable to panels that have already been established. Provided herein is a Cot- 1 improvement technique in which common repetitive elements unintentionally probed in enrichment panels are identified bioinformatically. Based on this identification, sense and antisense oligos across these elements can be provided as a proxy for a synthetic Cot-1 and / or as an additive to Cot-1 as generally discussed herein. This approach can be applied to multiple panels that have already been physically manufactured that share common design strategies. This will result in a “minimal synthetic Cot-1” useful for target enrichment.

[0031] FIG. 6 is a flow diagram of a bioinformatic process 300 for designing and generating a minimal synthetic Cot-1. In an embodiment, the process including a step 306 of providing (e.g., accessing, transmitting, and / or receiving) information representative of probe sequences for a panel of probes of interest. For example, the probes may be configured to bind to targeted regions of a sample, which may be a genomic DNA sample or another nucleic acid sample source. The probes represent regions complementary to sequences in the sample DNA. In an embodiment, the probes may be 20bp-120bp in length. Using the sequencing information, sequences that overlap with repetitive elements in genomic DNA are identified at step 310. In an embodiment, the repetitive elements may be those of Table 1 or Table 2, by way of example. The genomic DNA may be a reference genome or a reference genome representation. In an embodiment, the overlap is identified using a database of characterized repetitive elements from a population. In an embodiment, the identification may be via a sequence analysis tool, such as a sequence analysis tool of a computer (see FIG. 9). The overlap may be based on one or more of sequence identity, sequence similarity, or sequence homology. In an embodiment, the sequences are scored based on homology, and sequences are designated as overlapping based on a homology score or value above a preset threshold. In an embodiment, the sequences are scored based on homology, and sequences are designated as overlapping based on a preset number of highest-homology sequences (e.g., highest 100 or 500).

[0032] After identification, the process 300 proceeds to a step 312 of providing a plurality of oligonucleotides having a length of 60-100 nucleotides that correspond to the sequences as the minimal Cot-1 composition. The corresponding sequences may include sense and antisense oligonucleotides. Providing the plurality of oligonucleotides may include synthesizing the oligonucleotides of the minimal synthetic Cot-1 composition. Providing the plurality of oligonucleotides may include transmitting sequence of the oligonucleotides of the minimal synthetic Cot-1 composition to a DNA synthesis facility and generating synthesis instructions to synthesize the oligonucleotides of the minimal synthetiCot-1 composition.

[0033] FIG. 7 is a schematic illustration of a universal synthetic Cot-1 composition that may be used as part of a reaction mixture for various probe panels (illustrated as probe panels 350a, 350b, 350c by way of example) that may have probes that target repetitive elements. For sequences of each of the probes of the various panels, the overlapping sequences 354a, 354b, 354c may be identified and used to design respective oligonucleotides of minimal Cot-1 compositions 356a, 356b, 356c. Pooling these oligonucleotides yields a universal synthetic Cot-1 composition 360. In an embodiment, the overlapping sequences 354a, 354b, 354c may be subject to additional processing and analysis by pooling the overlapping sequences 354a, 354b, 354c in silico and removing duplicates or generating a representative composition of the pool before progressing to synthesis of the universal synthetic Cot-1 composition 360.

[0034] The minimal synthetic Cot-1 composition and / or the universal synthetic Cot-1 composition as discussed herein may include single-stranded or double-stranded oligonucleotides. The minimal synthetic Cot-1 composition and / or the universal synthetic Cot-1 composition may be provided as a reagent or as part of a kit for targeted enrichment NGS techniques. In an embodiment, the minimal synthetic Cot-1 composition and / or the universal synthetic Cot-1 composition may be used in conjunction with conventional Cot-1 and / or the improved Cot-1 compositions as discussed herein.EXAMPLES

[0035] Sequences of a panel of interest (composed of 39,759 target enrichment probes) were evaluated with RepeatMasker (version 4.1.2-pl Smit, A, R Hubley, and P Green. “RepeatMasker.” https: / / repeatmasker.org / .) using the NCBI / RMBLAST search engine (version 2.10.0+) using a curated database of repetitive elements (Dfam version 3.3 Storer, Jessica, Robert Hubley, Jeb Rosen, Travis J. Wheeler, and Arian F. Smit. “The Dfam Community Resource of Transposable Element Families, Sequence Models, and Genome Annotations.” Mobile DNA 12, no. 1 (January 12, 2021): 2. https: / / doi.org / 10.1186 / s 13100-020-00230-y). Resulting output from RepeatMasker was parsed to obtain the number of probes in the enrichment panel with homology torepetitive element classes (Error! Reference source not found.). Individual repeat elements within repeat classes were matched to the corresponding entries in Dfam. Only repeat elements with sequence identity to > 30 probes from the panel of interest were considered (Error! Reference source not found.). In Table 2, Alpha / ALR elements were not targeted in panel of interest, but manually added as they are common repeat elements in the human genome.

[0036] The disclosed repetitive elements as discussed herein may be, for example, those referred to in Table 1 or Table 2. The disclosed minimal synthetic Cot-1 may include oligonucleotides in any combination of SEQ ID NOS: 1-1354.Table 1 Repeat class and number of enrichment probes with sequence homology within the panel of interest

[0037] The identification of repetitive element sequences from Dfam and selection of k- mer oligonucleotides (80-mer oligonucleotides in this example) representing minimal synthetic Cot-1 were as follows. Sequences for each of the Dfam accessions (Table 2)were retrieved from the curated Dfam database version 3.6 (file: Dfam_curatedonly.embl.gz). Dfam sequences were divided into 80-mer, nonoverlapping segments using a custom Python script. Resulting 80-mers were binned into clusters by sequence similarity and a best representative sequence was chosen using BBTools dedupe. sh with the following parameters: minoverlap=60, k=17, findoverlap=t, minidentity=80, cluster=t, pbr=t, absorbrc=t, pto=tThe resulting 682 representative sequences were processed to replace any observed ambiguous base (“N”, which represents either A, T, G, or C) were replaced with A (adenosine). The reverse-complement of the 682 sequences were added to the set, resulting in 1364 80-mer sequences (SEQ ID NOS: 1-1364, Table 3) representing both strands of DNA to comprise a minimal synthetic Cot-1 to test with the panel of interest. Oligonucleotides of (SEQ ID NOS: 1-1364) were synthesized as a concentrated pool for full-functional testing along with the panel of interest.

[0038] FIG. 8 shows the use of the minimal synthetic Cot-1 had a positive impact on padded read enrichment, a measure of on-target performance in NGS target enrichment assays. Padded read enrichment was measured using 0, 15, and 150 uM concentration of minimal synthetic Cot-1 oligonucleotides provided as a hybridization blocker in a target enrichment reaction using a probe panel of interest.

[0039] The techniques disclosed herein may be implemented in conjunction with a sequencing device and / or a sequence analysis device. For example, the steps of FIG. 6, including the identification of overlapping sequences, may be performed using a sequence analysis tool of a sequence analysis device. It should be understood that the disclosed sequence analysis may include analysis of thousands or millions of nucleotide bases and / or thousands or millions of sequence combinations. FIG. 9 is a schematic diagram of a sequencing device 500 that may be used in conjunction with the disclosed embodiments for acquiring sequencing data of overlapping as generally discussed herein. The sequence device 500 may be implemented according to any sequencing technique, such as those incorporating sequencing-by-synthesis methods described inU.S. Patent Publication Nos. 2007 / 0166705; 2006 / 0188901; 2006 / 0240439; 2006 / 0281109; 2005 / 0100900; U.S. Pat. No. 7,057,026; WO 05 / 065814; WO 06 / 064199; WO 07 / 010,251, the disclosures of which are incorporated herein by reference in their entireties. In an embodiment, the sequencing device 500 includes a separate sample processing device 502 and an associated computer 504. However, as noted, these may be implemented as a single device. Further, the associated computer 504 may be local to or networked or otherwise in communication with the sample processing device 502. In the depicted embodiment, the biological sample may be loaded into the sample processing device 502 on a sample substrate 510, e.g., a flow cell or slide, that is imaged to generate sequence data. For example, reagents that interact with the biological sample fluoresce at particular wavelengths in response to an excitation beam generated by an imager 512 and thereby return radiation for imaging. For instance, the fluorescent components may be generated by fluorescently tagged nucleic acids that hybridize to complementary molecules of the components or to fluorescently tagged nucleotides that are incorporated into an oligonucleotide using a polymerase. As will be appreciated by those skilled in the art, the wavelength at which the dyes of the sample are excited and the wavelength at which they fluoresce will depend upon the absorption and emission spectra of the specific dyes. Such returned radiation may propagate back through the directing optics. This retrobeam may generally be directed toward detection optics of the imager 512.

[0040] The imager detection optics may be based upon any suitable technology, and may be, for example, a charged coupled device (CCD) sensor that generates pixilated image data based upon photons impacting locations in the device. However, it will be understood that any of a variety of other detectors may also be used including, but not limited to, a detector array configured for time delay integration (TDI) operation, a complementary metal oxide semiconductor (CMOS) detector, an avalanche photodiode (APD) detector, a Geiger-mode photon counter, or any other suitable detector. TDI mode detection can be coupled with line scanning as described in U.S. Patent No. 7,329,860, which is incorporated herein by reference. Other useful detectors aredescribed, for example, in the references provided previously herein in the context of various nucleic acid sequencing methodologies.

[0041] The imager 512 may be under processor control, e.g., via a processor 514, and the sample receiving device 502 may also include I / O controls 516, an internal bus 518, non-volatile memory 520, RAM 522 and any other memory structure such that the memory is capable of storing executable instructions, and other suitable hardware components that may be similar to those described with regard to FIG. 31. Further, the associated computer 504 may also include a processor 524, I / O controls 526, communications circuity 527, and a memory architecture including RAM 528 and nonvolatile memory 530, such that the memory architecture is capable of storing executable instructions 532. The hardware components may be linked by an internal bus, which may also link to the display 534. In embodiments in which the sequencing device 500 is implemented as an all-in-one device, certain redundant hardware elements may be eliminated.

[0042] The processor 514, 524 may be programmed to assign individual sequencing reads to a sample based on the associated index sequence or sequences according to the techniques provided herein. In particular embodiments, based on the image data acquired by the imager 512, the sequencing device 500 may be configured to generate sequencing data that includes base calls for each base of a sequencing read. Further, based on the image data, even for sequencing reads that are performed in series, the individual reads may be linked to the same location via the image data and, therefore, to the same template strand. In this manner, index sequencing reads may be associated with a sequencing read of an insert sequence before being assigned to a sample of origin. The processor 514, 524 may also be programmed to perform downstream analysis on the sequences corresponding to the inserts for a particular sample subsequent to assignment of sequencing reads to the sample.

[0043] Throughout this application various publications, patents and / or patent applications have been referenced. The disclosure of these publications in their entireties is hereby incorporated by reference in this application. The term comprising isintended herein to be open-ended, including not only the recited elements, but further encompassing any additional elements. While only certain features of the invention have been illustrated and described herein, many modifications and changes will occur to those skilled in the art. It is, therefore, to be understood that the appended claims are intended to cover all such modifications and changes as fall within the true spirit of the invention. Further, elements of the disclosed embodiments may be combined or exchanged. Accordingly, other embodiments are within the scope of the following claims.Table 1 : Sequences of 80-mer oligos representing a minimal synthetic Cot-1

Claims

CLAIMS1. A method of generating Cot-1 DNA, comprising: contacting genomic DNA fragments with a hybridization probe panel under conditions that permit hybridization of probes of the hybridization probe panel to the genomic DNA fragments, wherein the probes of hybridization probe panel are configured to hybridize to coding regions of the genomic DNA fragments; separating a first portion of the genomic DNA fragments from a second portion of the genomic DNA fragments based on differential binding to the hybridization probe panel, wherein the first portion is not bound to the probes and the second portion is bound to the probes; and providing the first portion of the genomic DNA fragments as Cot-1 DNA.

2. The method of claim 1, comprising fragmenting genomic DNA by sonicating or shearing the genomic DNA to generate the genomic DNA fragments.

3. The method of claim 1, wherein the genomic DNA comprises human placental DNA.

4. The method of claim 1, wherein the conditions that permit hybridization comprise exposing the genomic DNA fragments to one or more cycles of denaturing temperatures and reannealing temperatures.

5. The method of claim 1, wherein the conditions that permit hybridization comprise exposing the genomic DNA fragments to a range of denaturing temperatures and subsequently holding a reannealing temperature.

6. The method of claim 1, comprising exposing the first portion to a denaturing temperature and a subsequent reannealing temperature before providing the first portion of the genomic DNA fragments as Cot-1 DNA.

7. The method of claim 1, comprising contacting the first portion with a singlestranded nuclease before providing the first portion of the genomic DNA fragments as Cot-1 DNA.

8. The method of claim 1, wherein separating the first portion from the second portion comprises using affinity tags coupled to the individual probes.

9. The method of claim 8, wherein separating the first portion from the second portion comprises using affinity tag binders coupled to magnetic beads, wherein the affinity tag binders are configured to bind to the affinity tags.

10. The method of claim 9, comprising washing the affinity tag binder coupled to the magnetic beads for reuse.

11. The method of claim 1, comprising precipitating the first portion before providing the first portion of the genomic DNA fragments as Cot-1 DNA.

12. The method of claim 1, wherein the Cot-1 DNA is at least 95% double-stranded DNA13. The method of claim 1, wherein the hybridization probe panel comprises an exome panel or a whole exome panel.

14. A method of generating Cot-1 DNA, comprising: providing genomic DNA fragments;exposing the genomic DNA fragments to a denaturing temperature to yield denatured DNA; allowing the denatured DNA to reanneal to yield reannealed DNA; and purifying the reannealed DNA using solid-phase reversible immobilization beads; and providing the purified DNA as Cot-1 DNA.

16. The method of claim 15, wherein the beads are magnetic.

17. The method of claim 15, wherein the beads bind to reannealed DNA in a specific size ranging between 150bp to 800bp.

18. The method of claim 15, wherein purifying the reannealed DNA using solidphase reversible immobilization beads comprises: allowing the reannealed DNA to bind to the solid-phase reversible immobilization beads; separating the bound reannealed DNA from unbound components in a reaction mixture; eluting the reannealed DNA from the solid-phase reversible immobilization beads to yield the purified DNA.

19. A method of providing a minimal Cot-1 composition, comprising: providing probe sequence information of a panel of probes configured to hybridize to targeted regions of genomic DNA; identifying, in the probe sequence information, sequences that overlap with repetitive element sequences in the genomic DNA; and providing a plurality of oligonucleotides having a length of 60-100 nucleotides that correspond to the repetitive elements sequences as the minimal Cot-1 composition.

20. The method of claim 19, wherein the oligonucleotides are single-stranded or double-stranded.

21. The method of claim 19, wherein the plurality of oligonucleotides comprises at least 1000 oligonucleotides.

22. The method of claim 19, comprising: generating k-mers of the repetitive element sequences having the overlap with the sequences; clustering the k-mers by sequence similarity into a plurality of clusters; selecting, for each cluster, a best representative sequence; and generating the plurality of oligonucleotides using best representative sequences of the plurality of clusters.

23. The method of claim 22, wherein the best representative sequence comprises a synthetic or modified sequence in which ambiguous bases are replaced by adenosine.

24. The method of claim 22, wherein generating the plurality of oligonucleotides comprises generating first oligonucleotides of the best representative sequences and second oligonucleotides that are complements of the best representative sequences.

25. A Cot DNA composition prepared by a method comprising: contacting genomic DNA fragments with a hybridization probe panel under conditions that permit hybridization of probes of the hybridization probe panel to the genomic DNA fragments, wherein the probes of hybridization probe panel are configured to hybridize to coding regions of the genomic DNA fragments; separating a first portion of the genomic DNA fragments from a second portion of the genomic DNA fragments based on differential binding to thehybridization probe panel, wherein the first portion is not bound to the probes and the second portion is bound to the probes; and providing the first portion of the genomic DNA fragments as the Cot DNA composition.

26. A synthetic Cot-1 DNA composition comprising any combination of SEQ ID NOS: 1-1364.

27. A synthetic Cot-1 DNA composition comprising a plurality of oligonucleotides having a length of 60-100 nucleotides and comprising SEQ ID NOS: 1-1364.

Citation Information

Patent Citations

  • Mitigation of Cot-1 DNA distortion in nucleic acid hybridization

    CN101405407A

  • C0t-1DNA, its preparation method and application

    CN102465178A

  • Method for specific enrichment of nucleic acid sequences

    US20130130917A1

  • Hybridization methods and reagents

    US20220106590A1