Sequence capture technology and probe design
The k-mer based strategy for producing DNA probe libraries addresses the inefficiencies of existing methods by creating chimeric probes that enrich target sequences effectively, enhancing data quality and reducing costs, suitable for genomics and genetic engineering applications.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- MONSANTO TECHNOLOGY LLC
- Filing Date
- 2025-11-06
- Publication Date
- 2026-05-15
AI Technical Summary
Existing methods for designing DNA probe libraries are labor-intensive and costly, and individual probes or libraries often fail to effectively enrich target sequences due to saturating sequences present in multiple copies in the genome, leading to poor data quality and inefficiency in genome engineering projects.
A k-mer based strategy is employed to identify unique k-mer polynucleotide sequences, form contigs, and produce chimeric double-stranded DNA sequences, which are then fragmented to create a library of DNA probes capable of enriching target sequences, avoiding saturating probes and enabling scalable, cost-effective production.
The method allows for high-fidelity enrichment of target sequences, enabling efficient analysis of up to 2,000 regions per assay, reducing production time and cost, and facilitating applications in genomics and genetic engineering.
Smart Images

Figure IMGF000025_0001 
Figure IMGF000026_0001 
Figure IMGF000027_0001
Abstract
Description
TITLE OF THE INVENTIONSEQUENCE CAPTURE TECHNOLOGY AND PROBE DESIGNCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the priority of U.S. Provisional Application Serial No. 63 / 718.247, filed November 8, 2024, the entire disclosure of which is incorporated herein by reference.INCORPORATION OF SEQUENCE LISTING
[0002] A sequence listing containing the file named “MONS577WO_ST26” which is 110 kilobytes (measured in MS-Windows®) and created on November 6, 2025, and comprises 72 sequences, is incorporated herein by reference in its entirety.FIELD OF THE INVENTION
[0003] This present disclosure relates to the field of producing a DNA probe library for enriching and sequencing one or more DNA sequences of interest, and more specifically to methods of producing a DNA probe library using a k-mer analysis strategy for efficient and effective enrichment of DNA sequences of interest.BACKGROUND OF THE INVENTION
[0004] Sequence capture can be described as a method to selectively isolate sequences of interest for sequencing and analysis. Rather than sequencing the entire genome, which can be costly and time-consuming, sequence capture allows researchers to efficiently isolate and analyze specific genes or genomic regions, including those of transgenic origin. In general, sequence capture requires designing a library of short DNA or RNA probes, hybridization of such probes to fragmented genomic DNA, capture of the fragmented genomic DNA of interest through hybridization (i.e. genomic DNA hybridizing the designed probes), isolation of the fragmented genomic DNA of interest bound to the probes, and sequence analysis of the enriched DNA sequences of interest. Designing probe libraries is a critical step toward effective enrichment of sequences of interest as well as further analysis of such sequences.1US_ACTIVE\131553842.V1
[0005] Many characteristics must be evaluated in designing a library of probes in addition to complementarity to the target sequence or region of interest. For example, probe length, melting temperature, GC content, the presence of repetitive sequences, and the propensity for secondary structures, are characteristics known to affect the efficacy of sequence capture probes. However, even when carefully evaluated and optimized with respect to characteristics known to affect probe efficacy, not all individual probes or probe libraries are capable of enriching a target sequence of interest. The inability to properly enrich sequences of interest leads to poor data quality (e.g., failed sequencing reactions or reduced depth of sequencing at intended targets). This remains a significant bottle neck towards the development of high-quality probe libraries. At the same time, it is expensive and labor-intensive to test every probe design individually, and it is undesirably expensive to carefully design and order single- sequence individual biotinylated probes. These limitations impair the ability to apply library enrichment strategies to increasingly complex genome engineering projects. Therefore, there remains a significant need in the art for a comprehensive, low-cost, flexible, and highly scalable approach to generate effective libraries of DNA probes capable of enriching target sequences of interest across the genome.SUMMARY OF THE INVENTION
[0006] In one aspect the present disclosure provides, a method of producing a library of DNA probes, the method comprising identifying all or substantially all k-mer polynucleotide sequences between about 18 to about 24 nucleotides in length within a DNA sequence of interest within a genome; determining the number of times the k-mer polynucleotide sequences are found within the genome; identifying two or more contigs of at least about 50 nucleotides in length, wherein each k-mer polynucleotide sequence comprised in the two or more contigs is found less than or equal to about five times within the genome; producing a chimeric double-stranded DNA sequence comprising the two or more contigs; and fragmenting the chimeric double- stranded DNA sequence to produce a library of DNA probes. In one embodiment, the method comprises identifying all or substantially all k-mers of about 20 nucleotides in length within the DNA sequence of interest. In another embodiment, each k-mer polynucleotide sequence comprised in the two or more contigs is found less than or equal to four times, less than or equal to three times, less than or equal to two times, or less than or equal to one time, within the genome. In yet another embodiment, the contigs are at least about 55, at least about 65, at least about 75, at least about 85, at least about 95, or at2US_ACTIVE\131553842.V1least about 100 nucleotides, in length. The contigs, in specific embodiments, are at least about 80 nucleotides in length. In another embodiment, the method comprises producing a chimeric doublestranded DNA sequence comprising three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, twenty or more, twenty-five or more, or thirty or more, contigs. Non-limiting examples of fragmenting the chimeric doublestranded DNA sequence include shearing, endonuclease digestion, or a combination thereof. In further embodiments, the method comprises labeling the fragmented double- stranded DNA sequences. In still further embodiments, labeling comprises biotinylating the fragmented doublestranded DNA sequences. In certain embodiments, the k-mer polynucleotide sequences found less than about five times within the genome do not comprise mononucleotide repeats or dinucleotide repeats longer than about ten nucleotides in length, trinucleotide repeats longer than twelve nucleotides in length, or palindromic repeats capable of forming a hairpin longer than 6 nucleotides in length. In specific embodiments, the k-mer polynucleotide sequences do not comprise repeated sequence elements, including tandem repeats or inverted repeats (e.g., at any distance from each other in the sequence greater than 10 bp in length). In further embodiments, the genome is a plant genome. A library of DNA probes produced by the methods disclosed herein is also provided. In some embodiments, the library of probes comprises double-stranded DNA sequences capable of hybridizing to at least one, at least two, at least ten, at least twenty, at least fifty, at least one hundred, at least two hundred, at least five hundred, at least one thousand, or at least two thousand DNA sequences of interest, within the genome.
[0007] In other embodiments, a method of enriching the presence of a DNA sequence of interest in a sample is provided, the method comprising contacting the sample with a library of DNA probes produced by the methods disclosed herein, wherein the sample comprises fragmented genomic DNA; subjecting the sample and the library of DNA probes to hybridization conditions; and enriching the fragmented genomic DNA hybridized to the library of DNA probes in the sample. In certain embodiments, the hybridization conditions comprise high stringency conditions.
[0008] In another aspect, a method of producing a library of DNA probes is provided, the method comprising identifying all or substantially all k-mer polynucleotide sequences between about 18 to about 24 nucleotides in length within a plasmid DNA sequence of interest; determining the number of times the k-mer polynucleotide sequences are found within a genome; identifying two or more contigs of at least about 50 nucleotides in length, wherein each k-mer polynucleotide3US_ACTIVE\131553842.V1sequence comprised in the two or more contigs is found less than or equal to about five times within the genome; producing a chimeric double- stranded DNA sequence comprising the two or more contigs; and fragmenting the chimeric double- stranded DNA sequence to produce a library of DNA probes. In some embodiments, the chimeric double-stranded DNA sequence further comprises one or more additional contigs, wherein each k-mer polynucleotide sequence comprised within the one or more additional contigs is found less than or equal to about one time within the genome. In other embodiments, the chimeric double- stranded DNA sequence further comprises at least one additional polynucleotide sequence, wherein the at least one additional polynucleotide sequence is located: between the two or more contigs; at the 5’ end of the chimeric double-stranded DNA sequence; or at the 3’ end of the chimeric double-stranded DNA sequence. In still other embodiments, the method further comprises labeling the fragmented double-stranded DNA sequences. In particular embodiments, the plasmid DNA comprises a transformation vector. In other embodiments, the method comprises identifying all or substantially all k-mers of about 20 nucleotides in length within the plasmid DNA sequence of interest. In another embodiment, each k-mer polynucleotide sequence comprised in the two or more contigs is found less than or equal to four times, less than or equal to three times, less than or equal to two times, or less than or equal to one time, within the genome. A library of DNA probes produced by the methods disclosed herein is also provided, for example, wherein the library of probes comprises double-stranded DNA sequences derived from at least one, at least two, at least three, at least four, at least five, or at least six plasmid DNA sequences of interest. In specific embodiments, the library of probes is capable of identifying a transgene or a transgene insertion site.
[0009] In yet another aspect provided herein is a method of identifying a saturating sequence within a DNA sequence of interest, the method comprising: identifying all or substantially all k- mer polynucleotide sequences between about 18 to about 24 nucleotides in length within a DNA sequence of interest within a genome; determining the number of times the k-mer polynucleotide sequences are found within the genome; and identifying the k-mer polynucleotide sequences found more than one time within the genome. In some embodiments, the method comprises identifying k-mer polynucleotide sequences found one or more times, more than two times, more than three times, more than five times, or more than ten times, within the genome.4US_ACTIVE\131553842.V1BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The following drawings form part of the present specification and are included to further demonstrate certain aspects of the present invention. The invention may be better understood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein.
[0011] FIG. 1 shows an exemplary workflow of the k-mer based probe design.
[0012] FIG. 2 shows schematic results of sequencing a region of interest in multiple strains using a library of probes produced by the methods described herein. Twelve probes produced from three separate chimeric double- stranded DNA sequences were used to identify two alleles from 4 different soybean germplasms. Each germplasm was homozygous for one of two alleles; twelve polymorphisms were identified that distinguish the two alleles
[0013] FIG. 3A shows a graphical representation of k-mer based analysis (using 20-mers) to identify regions suitable for probe design for a region of interest. FIG. 3B shows a region with maximum range of probe possibilities. FIG. 3C shows a region with restricted range of probe possibilities.
[0014] FIG. 4A shows schematic sequencing results of ten regions of interest used to distinguish between corn line 01DK2 and 80IDM2 using a library of probes produced by the methods described herein. The results shown relate to the region of interest on Chromosome 10. FIG. 4B shows schematic sequencing results of ten regions of interest used to distinguish between corn line 01DK2 and 80IDM2 using a library of probes produced by the methods described herein. The results shown relate to the region of interest on Chromosome 1.
[0015] FIG. 5 shows k-mer counts across a validated sequence capture probe predicted by 15-17 k-mers as compared to using a 20 k-mer. The dashed grey line indicates the threshold for unacceptable k-mer counts. Any k-mers at or above this line would indicate that the region is not acceptable for a sequence capture probe. All k-mers with counts >10 were adjusted to 10 to improve graph readability.
[0016] FIG. 6 shows k-mer counts for a sequence capture probe predicted by 25-27 k-mers as compared to using a 20 k-mer. The dashed grey line indicates the threshold for unacceptable k- mer counts. Any k-mers at or above this line would indicate that the region is not acceptable for5US_ACTIVE\131553842.V1a sequence capture probe. The current algorithm for k-mer counting has a maximum threshold of 255 occurrences in the genome, after which, for bioinformatic efficiency, any additional occurrences in the genome are ignored. All k-mers with counts >10 were adjusted to 10 to improve graph readability.
[0017] FIG. 7 shows an exemplary sequence capture and analysis workflow using k-mer based probes.BRIEF DESCRIPTION OF THE SEQUENCES
[0018] SEQ ID NO:1 is an exemplary genomic Zea mays sequence on chromosome 1.
[0019] SEQ ID NO:2 is an exemplary primary forward primer sequence; SP6F.
[0020] SEQ ID NO:3 is an exemplary primary reverse primer sequence; SP6R.
[0021] SEQ ID NO:4 is an exemplary primary forward primer sequence; T7F.
[0022] SEQ ID NO:5 is an exemplary primary reverse primer sequence; T7R.
[0023] SEQ ID NO:6 is an exemplary secondary primer sequence; SP6.
[0024] SEQ ID NO:7 is an exemplary secondary primer sequence; T7.
[0025] SEQ ID NOs. 8-20 are exemplary corn target sequences.
[0026] SEQ ID NOs. 21-30 are exemplary K-mer probes targeting regions of interest in the soy genome.
[0027] SEQ ID NOs. 31 and 32 are exemplary contig designs for the soy genome.
[0028] SEQ ID NOs. 33 and 34 are exemplary contig designs containing both com and soy sequences.
[0029] SEQ ID NO: 35 is the polynucleotide sequence of the soy lectin gene.
[0030] SEQ ID NOs. 36-45 are exemplary K-mer probes used for enrichment of canola sequences.
[0031] SEQ ID NO: 46 is an exemplary contig designs containing both canola and cotton sequences.6US_ACTIVE\131553842.V1
[0032] SEQ ID NO: 47 is the polynucleotide sequence of the canola FAT4 gene.
[0033] SEQ ID NOs. 48-57 are exemplary K-mer probes designed to enrich regions of interest in the cotton genome.
[0034] SEQ ID NO: 58 is the polynucleotide sequence of the cotton ACP gene.
[0035] SEQ ID NOs. 59-68 are exemplary 90-mer probes for enriching genomic locations of interest in corn.
[0036] SEQ ID NOs. 69-71 are exemplary designs for 30-mer, 50-mer, and 90-mer contig probes embedded in nonsense sequence to produce chimeric probe designs.
[0037] SEQ ID NO: 72 is the polynucleotide sequence of the cotton PDC-3 gene.DETAILED DESCRIPTION OF THE INVENTION
[0038] The present disclosure provides novel k-mer based methods allowing production of high- quality probe libraries capable of enriching a target sequence of interest. The methods provided herein avoid the current limitations associated with sequence capture in a robust and cost-effective manner, without compromising data quality. In particular, the present inventors found that one of the most significant problems associated with designing probe libraries is that individual probes designed to a region of interest can be “saturating” probes, whose sequences (or highly similar sequences) are present in many iterations in the genome (potentially up to thousands of nearly identical copies). These probes enrich many regions of the genome and thus dilute the ability to enrich target sequences of interest during sequence capture. Furthermore, the present inventors found that when a saturating probe is included in a probe library, all probes enrich poorly, and data quality (depth of sequencing at intended targets) is comprised. In contrast, a single probe that fails to enrich the intended target(s) or fails to enrich any DNA sequence has no discernible effect on the efficacy of the probe library overall.
[0039] Thus, the present disclosure provides for the first time a k-mer based strategy to avoid potential issues related to saturating probes; and effectively produce probe libraries capable of enriching a target sequence of interest. As used herein, the term “saturating,” such as a saturating probe or saturating sequence, refers to a DNA sequence that is found multiple times within a genome and reduces the ability to enrich other target sequence(s) of interest during sequence7US_ACTIVE\131553842.V1capture when included in a probe library or mixture of probe libraries. In some embodiments, the methods provided herein comprise identifying the number of times a short k-mer (e.g., a 20-mer) sequence appears in the genome. This analysis can be performed to identify the longest sequence near or overlapping a target region of interest for which every k-mer polynucleotide sequence is unique or absent (a “clear region”). In other embodiments, it can be used to identify acceptable regions in which not every 20-mer is unique or absent, but for which there is an expectation that the probe library derived from said sequence will not be a saturating sequence. For example, closely related gene family members are generally enrichable (and generally are bioinformatically distinguishable) with regions of interest for which some, many, or every 20-mer is not unique. In such embodiments, a saturating sequence may refer to a k-mer polynucleotide sequence found more than two times, more than three times, more than four times, more than five times, more than six times, more than seven times, more than eight times, more than nine times, more than ten times, more than eleven times, more than twelve times, more than thirteen times, more than fourteen times, more than fifteen times, or more than twenty times, within the genome. After determining the number of times the k-mer polynucleotide sequences are found within the genome, the methods provided herein may comprise identifying one or more contigs of at least about 50 nucleotides in length, wherein each k-mer polynucleotide sequence comprised in the contig is found less than or equal to about five times within the genome. As used herein, a “contig” (also referred to as a contiguous sequence) is a continuous sequence of nucleotides found within a reference sequence, e.g., a genome of interest. In further embodiments, said one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, twenty or more, twenty-five or more, or thirty or more contigs can be concatenated to produce a chimeric sequence; and fragmented to produce a library of DNA probes.
[0040] The disclosed methods, as further demonstrated herein, are robust and avoid the need to carefully design and order “perfect” individual biotinylated probes of exact sequence. The methods described herein have successfully been used for scaled production of probe libraries, e.g., production by a single person of up to 48 probe libraries per day (with potential to produce probes against 480 or more individual target regions), whereas previous methods struggle to produce 10 probes (producing probes against 100 or more individual target regions) with greater than two days of effort. Moreover, the high-fidelity of the probe libraries produced using the methods described8US_ACTIVE\131553842.V1have successfully been used to analyze as many as 200 target regions per single assay, and may be used to analyze as many as 500, 1,000, 1,500, or 2,000 target regions per single assay.
[0041] Therefore, the methods provided by the present disclosure have broad applications in the study of genomics and genetic engineering, including but not limited to, transgene insertion, genetic modifications, zygosity, and variant analysis. The k-mer based methods described herein can be used to confidently investigate almost any sequence-based query.A. K-mer Based Analysis for Probe Design
[0042] The methods provided herein comprise identifying all or substantially all k-mer polynucleotide sequences within a DNA sequence of interest. Tn some embodiments, the method may comprise identifying all or substantially all k-mer polynucleotide sequences between about 12 to about 18 nucleotides in length, about 13 to about 19 nucleotides in length, about 14 to about 20 nucleotides in length, about 15 to about 21 nucleotides in length, about 16 to about 22 nucleotides in length, about 17 to about 23 nucleotides in length, about 18 to about 24 nucleotides in length, about 19 to about 25 nucleotides in length, about 20 to about 26 nucleotides in length, about 18 to about 24 nucleotides in length, about 19 to about 24 nucleotides in length, about 20 to about 24 nucleotides in length, about 21 to about 24 nucleotides in length, about 22 to about 24 nucleotides in length, about 18 to about 23 nucleotides in length, about 18 to about 22 nucleotides in length, about 18 to about 21 nucleotides in length, about 18 to about 20 nucleotides in length, about 19 to about 23 nucleotides in length, or about 20 to about 22 nucleotides in length, including all ranges derivable therebetween. In other embodiments, the method may comprise identifying all or substantially all k-mer polynucleotide sequences about 12 in length, about 13 in length, about 14 in length, about 15 in length, about 16 in length, about 17 in length, about 18 in length, about 19 in length, about 20 in length, about 21 in length, about 22 in length, about 23 in length, about 24 in length, about 25 in length, or about 26 in length, including all ranges derivable therebetween.
[0043] The methods provided herein further comprise identifying the number of times the k-mer polynucleotide sequences are found within a genome. Such method steps identify the longest possible regions that are unique in the genome assembly of interest; or identify regions that are strongly predicted to cause poor enrichment during sequence capture. For example, the methods described here may comprise identifying k-mer polynucleotide sequences (e.g., 20-mers) that are present at high copy number in the genome of interest and removing them from the set of k-mers9US_ACTIVE\131553842.V1used to design probe elements. Tn some embodiments, the methods comprise identifying each k- mer polynucleotide sequence found less than or equal to about ten times within the genome, about nine times within the genome, about eight times within the genome, about seven times within the genome, about six times within the genome, about five times within the genome, about four times within the genome, about three times within the genome, about two times within the genome, or about one time within the genome. In further embodiments, the methods comprise identifying each k-mer polynucleotide sequence found at least about one time within the genome, at least about two times within the genome, at least about three times within the genome, at least about four times within the genome, at least about five times within the genome, at least about six times within the genome, at least about seven times within the genome, at least about eight times within the genome, at least about nine times within the genome, or at least about ten or more times within the genome.
[0044] As demonstrated herein, short regions of near identity to a region of interest are sufficient to enrich selected DNA sequence(s) of interest. Regions as short as about 50 nucleotides are demonstrated to be sufficient for enrichment, particularly at higher levels of stringency. Conventional stringency conditions are described by Sambrook et al., 1989, and by Haymes et al.. In: Nucleic Acid Hybridization, A Practical Approach, IRL Press, Washington, DC (1985). Contiguous sequences, wherein each k-mer polynucleotide sequence comprised in the contig is found less than or equal to a number of times within the genome may be identified. In view of the description provided herein, those skilled in the art can determine the appropriate number of times within the genome each k-mer polynucleotide sequence may be comprised in the contig, as well as the desired length of the contig. For example, the nature and number of the DNA sequence(s) of interest and the stringency conditions for enrichment are each non-limiting considerations that may be evaluated. As such, in some embodiments the methods provided herein comprise identifying contigs of at least about 50 nucleotides in length, about 55 nucleotides in length, about 60 nucleotides in length, about 65 nucleotides in length, about 70 nucleotides in length, about 75 nucleotides in length, about 80 nucleotides in length, about 85 nucleotides in length, about 90 nucleotides in length, about 95 nucleotides in length, about 100 nucleotides in length, about 105 nucleotides in length, about 1 10 nucleotides in length, about 1 15 nucleotides in length, about 120 nucleotides in length, about 125 nucleotides in length, about 130 nucleotides in length, about 135 nucleotides in length, about 140 nucleotides in length, about 145 nucleotides in length, about 15010US_ACTIVE\131553842.V1nucleotides in length, about 160 nucleotides in length, about 170 nucleotides in length, about 180 nucleotides in length, about 190 nucleotides in length, or about 200 nucleotides or more in length, including all ranges derivable therebetween. In related embodiments, each k-mer polynucleotide sequence comprised in the contig is found less than or equal to about ten times within the genome, about nine times within the genome, about eight times within the genome, about seven times within the genome, about six times within the genome, about five times within the genome, about four times within the genome, about three times within the genome, about two times within the genome, or about one time within the genome.
[0045] The methods described herein further comprise producing a chimeric double- stranded DNA sequence comprising one or more contigs. As demonstrated herein, individual contigs can be concatenated into a chimeric sequence. Moreover, fragmenting such a chimeric doublestranded DNA sequence can produce a library of DNA probes capable of enriching all concatenated DNA sequences of interest represented in the chimeric sequence. For example, in certain embodiments, fragmenting of the chimeric sequence produces a library of DNA probes. Although fragmentation results in both DNA probes that are entirely within a single contig as well as others that comprise partial or complete sequences of adjacent contigs of the chimeric probe design, the presence of some fraction of fragments that are non-functioning does not negatively impact the ability of other functioning fragments to enrich: such methods have proven to be robust, while avoiding the need to carefully design perfect individual biotinylated probes. As such, production of probe libraries from chimeric double stranded DNA as described herein allows for scaled production of DNA probe libraries at significantly reduced costs and reduced time scales. Furthermore, the length of each double stranded DNA (dsDNA) probe is independent of all other designed probes, i.e., they may be of different lengths as long as fit the criteria for selection described herein. The probe tiles may be concatenated as is convenient for ease of dsDNA production; changing the order and / or direction of individual probe contigs in the chimeric dsDNA design does not materially affect the efficacy of the resulting probe libraries. Also, additional intercalating sequences and / or extending 5’ and / or 3’ sequences may be added to chimeric probe designs without lowering the efficacy of the resulting probe library. In some embodiments, it is necessary or desirable to add additional polynucleotide sequence(s) between contigs or even internally within a contig design, as long as each individual resulting contig fragment is 50 bp or longer, so that the dsDNA may be more robustly produced. These additional polynucleotide11US_ACTIVE\131553842.V1sequences are only required to pass the same bioinformatic cut-off for k-mers that the concatenated contigs were also subjected to. Such additional polynucleotide sequences can comprise at least about 1 nucleotide, at least about 2 nucleotides, at least about 3 nucleotides, at least about 4 nucleotides, at least about 5 nucleotides, at least about 6 nucleotides, at least about 7 nucleotides, at least about 8 nucleotides, at least about 9 nucleotides, at least about 10 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 35 nucleotides, at least about 40 nucleotides, at least about 45 nucleotides, or at least about 50 nucleotides, including all ranges derivable therebetween. Additional sequences that may be used include, but are not limited to, sequences derived from plasmid DNA; sequences derived from genomic DNA from a distinct organism, such as a plant pathogen; sequences derived from the reverse or complement of any sequence in the target species or from any other species; sequences derived from genomic DNA from a non-isogenic germplasm; sequences derived from genomic DNA from the target organism, or sequences derived from any artificially produced DNA sequence. In most cases the added sequence may have a k-mer score of 0 for each of the possible k-mers analyzed against the target organism, and is strongly predicted to not affect the depth of sequencing coverage. For a gDNA sequence taken from the target organism, the added sequence represents the addition of extra target region(s), but does not materially affect the depth of coverage of the regions of interest. Added sequences may also ensure that the projected synthesis of DNA for probe production is more robust. Additionally, because enrichment does not require perfect sequence matches across the entirety of each probe design, artificial polymorphisms (most often SNPs, or 5’- or 3 truncation of an individual contig design) may be introduced anywhere in the chimeric-double stranded DNA sequence, for instance to disrupt undesired long mono-, di- an trinucleotide repeats or to disrupt an undesired palindromic region or repeats (either in same orientation or in reverse complement orientation) at any distance from each other in the chimeric sequence. In specific embodiments, “nonsense” DNA sequences of any useful length can be intercalated between contigs in a chimeric dsDNA design, or even can be intercalated into the middle of contigs, as long as the resultant split contigs each have length of about 50 bp or greater. Nonsense DNA, such as the reverse (not reverse complement) of any chosen sequence (e.g. the reverse of a gene, intergenic region, or a plasmid sequence) can be used in chimeric probe designs in order to make the design longer without enriching additional regions; for improving the GC content of regions in the design; and to make the design fit the requirements for synthesis. Also,12US_ACTIVE\131553842.V1because enrichment does not require perfect annealing between probe and gDNA library sequences, artificial SNPs may be introduced to fit the requirements for synthesis, for instance by breaking up a direct or inverted repeats.
[0046] In some embodiments, the chimeric double- stranded DNA sequence comprises at least about 100 nucleotides, at least about 150 nucleotides, at least about 200 nucleotides, at least about 250 nucleotides, at least about 300 nucleotides, at least about 350 nucleotides, at least about 400 nucleotides, at least about 450 nucleotides, at least about 500 nucleotides, at least about 550 nucleotides, at least about 600 nucleotides, at least about 650 nucleotides, at least about 700 nucleotides, at least about 750 nucleotides, at least about 800 nucleotides, at least about 850 nucleotides, at least about 900 nucleotides, at least about 1000 nucleotides, at least about 1100 nucleotides, at least about 1250 nucleotides, at least about 1500 nucleotides, at least about 1750 nucleotides, at least about 2000 nucleotides, at least about 3000 nucleotides, or at least about 5000 or more nucleotides in length, including all ranges derivable therebetween. Chimeric doublestranded DNA sequences provided herein may also be described as comprising two of more contigs, three or more contigs, four or more contigs, five or more contigs, six or more contigs, seven or more contigs, eight or more contigs, nine or more contigs, ten or more contigs, fifteen or more contigs, twenty or more contigs, twenty-five or more contigs, or thirty or more contigs. The methods described herein have been successfully used to produce probe sets capable of enriching as many as 200 DNA sequences of interest in a single assay. Also provided herein are chimeric double-stranded DNA sequences comprising DNA sequences from more than one organism or source. For example, DNA sequences derived from transgenes, endogenous genes, plasmids, and other sources, which may all be comprised within a single double- stranded DNA sequence.B. Applications of K-mer Based Probe Design
[0047] The methods of the present disclosure have broad applications for the investigation of almost any sequence-based query. For example, the methods described herein may be advantageous in the study of genomic modification, zygosity analysis, SNP analysis, variant analysis, identification of chromosome rearrangements, strain identification, transgene characterization, and in-class genetic contamination. In particular, the methods described herein are capable of producing probe libraries for characterizing random insertions, site-directed integrations, gene edits (e.g., INDELs, chromosome rearrangements), or any other element of13US_ACTIVE\131553842.V1interest. The methods described herein provide significant improvements over the prior art. For example, the methods provided herein can produce a library of probes to target any DNA sequence of interest, including those that are difficult to assess with amplicon-based strategies. Furthermore, using mixtures of probe libraries (pooled probes) produced by the methods described herein do not require re-validation of the pooled probe sets. Multiplexing probe libraries to enrich multiple target regions does not require additional laboratory costs, since enrichment by hybridization and capture is accomplished by a single pool of probe or probes against a single pool of many sequenceable and demultiplexable gDNA librariesC. Sequence Capture and Enrichment
[0048] “Sequence capture” refers to the process of enriching specific DNA fragments from a DNA library (e.g., a genomic DNA library) using probes that are complementary to the target sequence of interest. For example, using a library of DNA probes produced by the methods described herein. Sequence capture methodologies enable enrichment of DNA sequence(s) of interest from material derived from single cells, tens of cells, hundreds of cells, thousands of cells, or millions of cells. For example, DNA may be enriched from material derived from about 1 cell to about 1,000,000,000 cells, about 1 cell to about 500,000,000 cells, about 1 cell to about 100,000,000 cells, about 1 cell to about 50,000,000 cells, about 1 cell to about 10,000,000 cells, about 1 cell to about 5,000.000 cells, about 1 cell to about 1,000.000 cells, about 1 cell to about 500,000 cells, about 1 cell to about 100,000 cells, about 1 cell to about 50,000 cells, about 1 cell to about 10,000 cells, about 1 cell to about 5,000 cells, about 1 cell to about 1,000 cells, about 1 cell to about 500 cells, about 1 cell to about 100 cells, about 1 cell to about 50 cells, about 1 cell to about 25 cells, or about 1 cell to about 10 cells, including all ranges derivable therebetween. The cell or cells from which DNA may be enriched can be derived from any organism. Nonlimiting examples of such organisms include eukaryotic single-celled organisms, eukaryotic multicellular organisms, fungi, plants, animals, reptiles, amphibians, insects, mammals, and humans. In particular embodiments, samples for use in the methods provided by the present disclosure may be derived from any plant, plant part, cell, or tissue thereof (including transgenic and non-transgenic plants), non-limiting examples of which include seeds, leaves, stems, roots, root tips, anthers, pistils, seed, grain, embryo, pollen, ovules, flower, shoot, tissue, petiole, cells, meristematic cells, and the like.14US_ACTIVE\131553842.V1
[0049] Provided herein are methods of enriching the presence of a DNA sequence of interest in a sample. The libraries of DNA probes produced by the methods described herein can be used in combination with standard sequence capture procedures known in the art. In some embodiments, such methods comprise contacting the sample with a library of DNA probes produced by the methods described herein, wherein the sample comprises fragmented genomic DNA; subjecting the sample and the library of DNA probes to hybridization conditions; and enriching the fragmented genomic DNA hybridized to the library of DNA probes in the sample. In some embodiments, the fragmented genomic DNA is modified at the 5’ and / or the 3’ ends to allow for sequencing and demultiplexing of data when employing a standard sequencing platform. In further embodiments, the hybridization conditions comprise high stringency conditions. Using a library of DNA probes produced by the methods described herein allows for enrichment of DNA comprising a DNA sequence of interest or portion thereof. In some embodiments, a DNA sequence of interest may be defined as a sequence comprising at least 1, at least 5, at least 10, at least 25, at least 50, at least 100, at least 500, at least 1,000, at least 2,000. at least 3,000. at least 4,000, at least 5,000, at least 10,00, at least 15,000, at least 20,000, at least 25,000, at least 30,000, at least 40,000, at least 50,000, at least 100,000, at least 150,000, at least 200,000, at least 250,000, at least 300,000, at least 400,000, or at least 500,000 nucleotides, including all ranges derivable therebetween. Furthermore, such libraries may be used to enrich DNA comprising the DNA sequence of interest and at least 1, at least 5, at least 10. at least 25. at least 50, at least 100, at least 500, at least 1,000, at least 2,000, at least 3,000, at least 4,000, at least 5,000, at least 10,00, at least 15,000, at least 20,000, at least 25,000, at least 30,000, at least 40,000, at least 50,000, or at least 100,000 further nucleotides, including all ranges derivable therebetween. In other embodiments, a DNA sequence of interest may be within a genome, a B chromosome, a sex chromosome, a supernumerary chromosome, an accessory chromosome, a plasmid DNA sequence, a linear or circular fragment of a plasmid-derived DNA sequence or genomic DNA sequence, a plastid DNA sequence, or within a mitochondrial DNA sequence. In still further embodiments, a DNA sequence of interest may be a transgene, a transgene insertion site, a genomic modification site, a foreign DNA sequence, or a gene of interest (e.g., an endogenous gene of interest; including noncoding and / or coding regions). In some embodiments, the library of DNA probes can be produced from sequences derived from one or more distinct organisms and may contain one or more types of DNA sequences (including, e.g., a transgene DNA sequence, a plastid DNA sequence, etc.)15US_ACTIVE\131553842.V1D. Probe Labeling
[0050] As described herein, in some embodiments, the methods comprise labeling the fragmented double-stranded DNA sequences, such as a fragmented chimeric double- stranded DNA sequence. As used herein, the term “label” refers to a directly or indirectly detectable oligonucleotide or nucleotide modification that is conjugated directly or indirectly to the composition to be detected. Non-limiting examples of a nucleotide modification that may be present, in certain embodiments, in a label as described herein include a biotin label, a fluorescent label, or a chemical modification. Methods to introduce oligonucleotide labels are known in the art and any such method may be used according to the methods provided by the present disclosure. For example, after the chimeric double-stranded DNA is fragmented, it may be subjected to end-repair involving making the ends blunt, and adding a single adenine (A) residue to the 3’ end of each strand (also known as A- tailing). To facilitate the ligation of adapters to the fragmented DNA, the adapters are designed with a single thymine (T) overhang on their 3’ end, creating a complementary 1 bp overhang (T pairs with A), which allows for efficient ligation between the adapter and the fragmented probe DNA. Modification can also be used to increase stability and efficiency, e.g., phosphorothioate linkages, 5’ phosphorylation of the reverse primer (e.g. “5’-Phos”), or 3’ modifications such as dideoxynucleotide linkages (e.g. “ddC”). The designed primers must be compatible with a second set of primers used during PCR amplification. Any primer sequence(s) may be used with the methods disclosed herein, provided such sequences do not interfere with subsequent sequencing. Exemplary primer sequences provided herein include, but are not limited to, a SP6F Primer (SP6F) 5’- CACGACTATTTAGGTGACACTATAGT-3’; a SP6R primer (SP6R) 5’-Phos- CTATAGTGTCACCTAAATAGTCGTGddC-3’; a T7F primer (T7F) 5’-CTCCGATAATACGACTCACTATAGGGT-3’; a T7R primer (T7R) 5’-Phos- CCCTATAGTGAGTCGTATTATCGGAGddC-3’; a Biotin-SP6 Primer 5’ Biotin- CACGACTATTTAGGTGACAC-3’; and a Biotin-T7 Primer 5’Biotin- CTCCGATAATACGACTCACTA-3’ (SEQ ID NOs: 2-7).
[0051] As used herein the terms “polynucleotide,” “oligonucleotide,” and “nucleic acid” may be used interchangeably and include linear oligomers of natural or modified monomers or linkages. A polynucleotide may include, for example, deoxyribonucleosides, ribonucleosides, a-anomeric forms thereof, peptide nucleic acids, and the like, capable of specifically binding to a target polynucleotide by way of a regular pattern of monomer-to-monomer interactions, such as16US_ACTIVE\131553842.V1Watson-Crick type of base pairing, base stacking, Hoogsteen, or reverse Hoogsteen type base pairing. Monomers may be linked, in some embodiments, by a phosphodiester bond or an analog thereof to form polynucleotides. Whenever a polynucleotide is represented by a sequence of letters herein, a person of ordinary skill in the art would understand that the nucleotides are in 5' to 3' orientation from left to right. A person of ordinary skill in the art would further understand that if a polynucleotide is presented as a sequence of letters that “A” denotes adenine, “C” denotes cytosine, “G” denotes guanine, and “T” denotes thymine, and “U” denotes uracil, unless otherwise noted. Analogs of phosphodiester linkages include, but are not limited to, phosphorothioate, phosphorodithioate, phosphoranilidate, and phosphoramidate linkages. It is clear to those skilled in the art when polynucleotides having natural or non-natural nucleotides may be employed. For example, a person of ordinary skill in the art would understand when processing by enzymes may be employed or when polynucleotides consisting of natural nucleotides are required.
[0052] In certain embodiments, polynucleotides of the present disclosure may comprise one or more modified or substituted sugar moieties. Such modified or substituted sugars may, in some embodiments, improve stability in the presence of nucleases or binding affinity. Non-limiting examples of modified or substituted sugars include carbocyclic or acyclic sugars, sugars having substitute groups at one or more of their 2', 3' or 4' positions, sugars having substitutes in place of one or more hydrogen atoms of the sugar, and sugars having a linkage between any two other atoms in the sugar. A large number of sugar modifications are known in the art and any such sugar modification may be used according to the present disclosure.
[0053] In some embodiments, polynucleotides may include one or more nucleobase modifications or substitutions which are structurally distinguishable from, yet functionally interchangeable with, naturally occurring or synthetic unmodified nucleobases. Modified nucleobases may include in certain embodiments, synthetic or natural nucleobases such as 5- methylcytosine (5-me-C), 5 -hydroxymethyl cytosine, 7-deaza-guanine, or 7-deaza-adenine, 2- aminopyridine, 2-pyridone, 5-substituted pyrimidines, 6-azapyrimidines and N-2 substituted purines, N-6 substituted purines, 0-6 substituted purines, 2 aminopropyladenine, 5- propynyluracil, and 5-propynylcytosine.17US_ACTIVE\131553842.V1E. Definitions
[0054] The following definitions are provided to define and clarify the meaning of these terms in reference to the relevant embodiments of the present disclosure as used herein and to guide those of ordinary skill in the art in understanding the present disclosure. Unless otherwise noted, terms are to be understood according to their conventional meaning and usage in the relevant art, particularly in the field of molecular biology and plant genomics.
[0055] When introducing elements of the present disclosure or the embodiment(s) thereof, the articles “a,” “an,” “the,” and “said” are intended to mean that there are one or more of the elements.
[0056] The term “and / or,” when used in a list of two or more items, means any one of the items, any combination of the items, or all of the items with which this term is associated.
[0057] The terms “comprising,” “including,” and “having” are intended to be inclusive and mean that there may be additional elements other than the listed elements. For example, any method that “comprises,” “has” or “includes” one or more steps is not limited to possessing only those one or more steps and can also cover other unlisted steps. Similarly, any composition or device that “comprises,” “has” or “includes” one or more features is not limited to possessing only those one or more features and can cover other unlisted features.
[0058] As used herein, a “plant” includes a whole plant, explant, plant part, somatic or germline tissue, seed, seedling, or plantlet at any stage of regeneration or development.
[0059] As used herein, the term “transgene” refers to a DNA molecule artificially incorporated into the genome of an organism as a result of human intervention, such as by plant transformation methods. The term “transgenic” means comprising a transgene, for example a “transgenic plant” refers to a plant comprising a transgene in its genome.
[0060] As used herein, “genome,” “genomic DNA,” or “gDNA,” refers to chromosomal DNA of an organism.
[0061] As used herein, a “genomic modification” (also referred to as “modification”) or “genomic edit” (also referred to as “edit”) refers to any modification to a genomic nucleotide sequence as compared to a wild-type or control plant. A genomic modification or genomic edit comprises a deletion, an insertion, a substitution, an inversion, a duplication, or any combination thereof. In some embodiments, a genomic modification may comprise a known structural variant present in18US_ACTIVE\131553842.V1an organism as compared to a wild-type or control organism, or may represent a novel structural variant as compared to a wild-type or control organism.
[0062] All methods described herein can be performed in any suitable order unless otherwise indicated herein or clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., “such as”) provided with respect to certain embodiments herein is intended merely to illuminate the present disclosure and does not pose a limitation on the scope of the present disclosure otherwise claimed.
[0063] Other objects, features, and advantages of the present disclosure are apparent from detailed description provided herein. It should be understood, however, that the detailed description and any specific examples provided, while indicating specific embodiments of the disclosure, are given by way of illustration only, since various changes and modifications within the spirit and scope of the disclosure will become apparent to those skilled in the art from this detailed description. Any embodiment of the present disclosure may be used in combination with any other embodiment described herein.
[0064] All references herein are incorporated herein by reference in their entirety.EXAMPLES
[0065] The following examples are included to illustrate embodiments of the present disclosure. It should be appreciated by those of skill in the art that the techniques disclosed in the examples that follow represent techniques discovered by the inventor to function well in the practice of the disclosure. However, those of skill in the art should, in light of the present disclosure, appreciate that many changes can be made in the specific embodiments which are disclosed and still obtain a like or similar result without departing from the concept, spirit and scope of the disclosure. More specifically, it will be apparent that certain agents which are both chemically and physiologically related may be substituted for the agents described herein while the same or similar results would be achieved. All such similar substitutes and modifications apparent to those skilled in the art are deemed to be within the spirit, scope and concept of the disclosure as defined by the appended claims.19US_ACTIVE\131553842.V1Example 1 : K-mer Based Probe Design
[0066] The following example describes the design of probes for sequence capture that selectively enrich a region of interest while avoiding enrichment of excessive copies of genomic regions that will cause significant reductions in sequencing depth and data quality.
[0067] The sequence capture probes were designed using a computational algorithm to identify and evaluate potential sequence capture probes. Briefly, the sequence of a genomic region of interest for enrichment was entered into the algorithm. The algorithm used a sliding window to identify each 20-mer of the sequence of interest (k-mer) and calculate how many times the 20-mer sequence appeared in the entire genome of the respective plant. K-mers were selected for probe design based on appearing only once or 5 or less times in the genome.
[0068] If a number of consecutive 20-mers pass the selection criteria (e.g., appearing between 0 and 5 times in the genome), then the original sequence from which the consecutive 20-mers were derived is identified as a potential region for probe design. Ideal probe regions for assessing genomic editing are at least 50 bp and preferably greater than 100 bp in length, of which each 20- mer comprising the larger sequence is unique or absent in the genome assembly, and is within 200 bp of the genomic sequence of interest. In situations where an ideal probe is not designable in a given target region, non-ideal probes are useable, and are designed by relaxing one or more of the “ideal” restrictions. For alternate analyses, the details of an ideal probe can be different. For instance, if significant allelic variation is not present (or expected not to be present, such as in the case of large foreign DNA insertions) then probes that cover the entirety of the regions that fit the k-mer analysis standard are preferred.Example 2: Preparing K-mer Probes for Sequence Capture
[0069] To generate synthetic k-mer probes, the DNA probes designed using the methodology described in Example 1 were concatenated together to create a chimeric double- stranded DNA sequence design. The double- stranded DNA sequence comprises one or greater distinct probe sequences. The DNA sequence can be synthesized using methods known in the art and as described herein.
[0070] Following synthesis of the designed dsDNA, probes were produced using components of the Roche Kapa Hyperplus Kit in a modification of the method used to produce gDNA libraries.20US_ACTIVE\131553842.V1The modifications result in short (<300 bp) biotinylated probe sets that are suitable for enrichment strategies. The chimeric double- stranded DNA sequences (0.25 to 1 pg of DNA) were semirandomly sheared using fragmentation enzymes in a total volume of 50 pl at 37° C for 20 minutes, resulting in many short fragments in the 50 to 250 bp range. The sheared DNA was size selected using SizeSelector-1 Beads (Aline). 190 pl of undiluted beads were added to the 50 pl fragmentation mix (“3.8x beads”). The mixture was incubated at room temperature for 5 mins, placed on the magnet for approximately 3 mins. 220 pl of the supernatant was moved to a fresh tube that contained 2.5 ml of undiluted beads (resulting in a “4.0x” bead solution. The mixture was incubated at room temperature for 5 mins, placed on the magnet for approximately 3 mins. The supernatant was removed, and the bead pellet was washed twice with 180 pL of 80% ethanol. Beads were dried at room or elevated (37° C to 45° C) temperature for at least 5 minutes. Sheared and sized DNA was eluted from the beads in 55 pl of water for at least 3 minutes, then placed on a magnet for at least 2 minutes. The solution was removed from the beads.
[0071] A probe library was prepared using 50 pL of the sheared and sized DNA solution by adding 7 pL of End-Repair and A-Tailing (ERAT) buffer and 3 pL of ERAT Enzyme. The solution was incubated at 65°C for 30 minutes, then incubated at 4° C or 10°C until ready to continue. Next, 2.5 pL of annealed SP6_F / SPC6_R adapter mix, 2.5 pL of annealed T7_F / T7_R adapter mix, 5 pL of water, 30 pL of Ligation Buffer, 10 pL of DNA ligase was added and incubated at 20° C for 15 minutes. The reaction was then purified with SeqPure PCR Purification Beads (Biochain) by adding 198 of undiluted beads to the 110 pl reaction (“1.8x” Biochain beads). Following 5 minutes of incubation at room temperature, samples were moved to a magnet for at least 3 minutes. Supernatant was removed and discarded. Pellets were dried and the DNA was eluted (as described above) in 25 pL of water.
[0072] For amplification, 20 pL of the reaction was added to a fresh tube along with 100 pL of KAPA HiFi HotStart Reaction Mix, 20 pL of 5 uM Biotin-T7 primer, 20 pL of 5 uM Biotin-SP6 primer, and 40 pL water. The resulting 200 pL reaction mixture was split into four 50 pL reactions for PCR reaction. The PCR reaction was incubated at 98°C for 45 seconds; then at least 12 cycles of: 98°C for 15 seconds, 60°C for 30 seconds, 72°C for 30 seconds; then a final 72°C incubation for up to 5 minutes. Following amplification, the reaction was cleaned using DNA SizeSelector-1 Beads using the same protocol described above, except for volumes that result in a “2.5x” and “4.0x” bead ratio: 125 pl of undiluted beads are added to each 50 pl PCR reaction21US_ACTIVE\131553842.V1(2.5X); 150 pl of supernatant is transferred to a fresh tube containing 64.2 pl of undiluted beads (resulting in “4.0x” ratio). Following washing and drying steps, each reaction is eluted in approximately 35 to 50 pl of sterile water; after placement on a magnet, the supernatants of the four pools are combined (140 to 200 pl total yield of designed probe library). Probe library sizes of approximately 150-300 bp were confirmed using an Agilent High Sensitivity DNA chip. The final volume of up to 200 pl is convenient for present uses, but larger volumes from the same preparation can be used, as is convenient for workflows. Probe libraries can be stored at -20° C for at least two years with no apparent loss of efficacy. Data obtained further demonstrated that different probe lengths that can be attained by varying bead sizing parameters. Probe length does not appear to have material effect on quality of enrichment attainable; 3.8x / 4.0x bead sizing conditions give -100 to 150 bp probe sizes. Larger sized probes (e.g., >300 bp average, and with wider range of sizes around the average) may also be readily purified and used to effectively enrich DNA sequences of interest.Example 3: Identifying SNPs in a Genome using K-mer Probes
[0073] The library of DNA probes designed using the methodologies in the above examples was used to enrich highly repetitive regions in the corn genome. To test whether the library of probes could be used to enrich DNA sequences near repetitive regions of interest and to identify SNPs varying among the 01 DKD2, 80IDM2, and LH244 com lines within the regions of interest, the probes were designed against 10 target sequences in the corn genome. The probes were designed using the methodologies described in Example 1 and Example 2. Briefly, probes were designed to 10 regions, each on a different chromosome. Each region was approximately 2000 bp in length, and 100 bp probes were designed to enrich a region that included a potential SNP at or near the center of the abstracted sequence. This example demonstrates that the library of DNA probes can enrich a standard set of regions of interest; and provides confirmation that enrichment of regions containing structural variants (e.g., SNPs) is robust.
[0074] DNA was extracted from wild-type corn plants from corn lines 01DKD2, 80IDM2, and LH244. The DNA was purified using Macherey-Nagel Plant 24 extraction kit according to the manufacturer’s instructions. Genomic DNA libraries suitable for sequencing Illumina NovaSeq6000 were prepared in 96-well format using the Kapa HyperPlus Kit (Roche) as per manufacturer’s instructions. Pooled libraries were stored at 4°C for up to a week before22US_ACTIVE\131553842.V1hybridization and capture. Longer storage of gDNA libraries at -20° C or -80° C is also possible. Hybridization and capture of sequences of interest were accomplished using 7 pl of biotinylated synthetic probe library and 0.5 to 1.5 ug of pooled gDNA libraries. The KAPA Hyper Capture Reagent Kit (Roche) was used for hybridization and capture according to the manufacturer’s instructions, including substitution of KAPA Hybrid Enhancer Reagent (“KHE;” Roche) for the kit-provided enhancer that is optimized for non-plant use. Briefly: gDNA library pool is mixed with 20 pl KHE and 2 volumes of undiluted A-line beads for 5 minutes at room temperature. After a three-minute subjection to a magnet the supernatant is removed, the on-magnet pellet washed once with 80% ethanol and then dried for at least 5 minutes. The gDNA library was eluted by sequential additions of 13.4 pl Universal Enhancing Oligos, 28 pl Hybridization buffer and 12 pl Component H. Following room temperature incubation for at least three minutes, and subjecting the mixture to a magnet for at least 3 minutes, the entire supernatant (approximately 53 pl and containing the gDNA libraries) is added to a tube or 96-well plate to which 7 pl of prepared and mixed probe has been pre-aliquoted. The mixture is incubated at 95°C for 10 mins and cooled to 51° C for at least 10 hours for hybridization. After incubation, 60 pL of the hybridized solution was added to beads equivalent to 50 pl of mixed Dynabeads (Invitrogen), which had been prepared according to manufacturer’s specifications and from which the entire final wash solution had been removed. The hybridized DNA was incubated with the Dynabeads for approximately 15 mins at 47 °C.
[0075] The Dynabeads were washed using 100 pL of wash buffer that had been heated to 47°C, mixed, and then placed on the magnet to enable removal of wash buffer with pipette. The Dynabeads were then washed 5 more times using a series of buffers, and the liquid removed while the beads were on a magnet. Following final wash, the beads were suspended in 40 pL of water.
[0076] DNA was amplified from the suspension by using 20 pL of the capture library (including beads), 25 pL of KAPA HiFi HotStart Ready Mix, 2.5 pL each of two PCR oligo mixtures. The resulting solution was then amplified using a PCR cycle with 14 cycles at 98°C for 15 seconds, 60°C for 30 seconds, 72°C for 45 seconds; then incubated at 72° C for 5 minutes and then held at 4°C until use. The amplification solution was purified using Aline Size Selection Beads according to manufacturer’s instructions for a “1.6X” bead ratio. The resulting DNA pool’s concentration and average size were characterized using an Agilent High Sensitivity DNA Chip. Samples were23US_ACTIVE\131553842.V1submitted to sequencing on a NovaSeq 6000 instrument, from which an average of 1 million reads per sample was requested.
[0077] Following sequencing, the DNA sequences were compiled and aligned. 10 target sequences and a control sequence comprising an approximately 1 kb region of the lectin gene (as a control for library quality) were analyzed across the 10 chromosome sequences. Of the 10 regions, 8 were found to have SNPs among the corn lines in the target sequences as detailed in Table 1. The identified SNPs can be used as a single point confirmation for the presence of one or more of the three germplasms. For instance, the identified and enriched Chromosome 2 SNP is different in each of the three germplasms, thus is able to distinguish the presence of any single parental strain, if only 1 version of the SNP is found, as well as any offspring of a cross between the germplasms, where 2 of the 3 SNP variants are expected to be found and can identify the parent strains.Table 124US_ACTIVE\131553842.V1Example 4: Analysis of K-mer Length for Probe Design
[0078] The following example describes the influence of k-mer length on the design of sequence capture probes. In particular, k-mers of 15-27 nucleotides in length were calculated and compared for both a validated sequence capture probe and a region that only met the uniqueness criteria for longer k-mer sizes. The validated sequence capture probe comprised nucleotides 606 to 736 of SEQ ID NO:1. This probe uniquely matched the genome when using an 80% identity cutoff and generated robust and specific sequencing data. In FIG. 6, the counts of 15-mers, 16-mers, 17- mers, and 20-mers are shown across the validated probe region. K-mers of 15-17 nucleotides had multiple k-mers with >5 counts, which would have led to excluding this region. Although k-mers of this size could be used, it is very stringent and would exclude many regions that could be used as sequence capture probes.
[0079] The region comprising nucleotides 1556-1607 of SEQ ID NO:1 was predicted to be a valid sequence capture probe for k-mers in the 25-27 size range, but not for k-mers less than 25 nucleotides (FIG. 7). When the 50 nucleotide probe regions were matched to the genome using BLAST, >1000 genomic locations are reported with >80% identity. This suggests that this region is unacceptable for use as a sequence capture probe. These two regions (nucleotides 606 to 736 and nucleotides 1556-1607 of SEQ ID NO:1) indicate that although k-mers of various sizes can be used, using k-mers less than or equal to 17 nucleotides in length discarded regions that make good sequence capture probes, whereas using k-mers greater than or equal to 25 nucleotides identify regions that are unspecific and would require additional steps to discard unspecific probes.25US_ACTIVE\131553842.V1Example 5: K-mer Probes Genomic DNA Enrichment from Corn, Soy, Cotton, and Canola
[0080] To demonstrate that the probe design methodologies described in Example 1 using corn are applicable to multiple crops species of interest, probes against regions of interest in the genomes of soy. cotton, and canola were design, synthesized, and libraries produced. Genomic DNA (gDNA) libraries of each species were enriched using probes designed and prepared as described in Examples 1 and 2. Enriched gDNA library samples were then sequenced and mapped to the target regions to demonstrate the utility of the k-mer design methodology as described in Example 3.Soy Probes and Combined Soy / Corn Probes
[0081] The methodology described in the above examples was used to enrich regions of interest in the Glycine max (soybean) genome. Briefly, 10 k-mer probes of 50 bp length were designed to unique regions on 10 separate chromosomes of the soy genome (Table 2). The k-mer probes were concatenated to produced contigs for library preparation. Soy genomic assembly sequences including the probe target and approximately 1000 bp 5’ and 3’ of the probe were used to map the enriched reads.Table 2: K-mer probes targeting regions of interest in the soy genome.26US_ACTIVE\131553842.V1
[0082] The full contig sequences were designed by concatenating the ten 50bp contigs derived from soy with sequences derived either from corn or from nonsense sequences that were not targeted to any region of the genome. Two designs of 50-mers derived from soy separated by 50bp nonsense sequences were produced (Table 3) and two additional sequences derived from soy and corn sequences were produced (Table 4); all four were tested and results are shown in Table 5. Three technical replicates were carried out for enrichment and sequencing.Table 3: Contig designs for the soy genome with the sequences that the k-mer probes were concatenated with to create the contigs for library preparation. 50-mer soy probes are in bold; underlined sequences highlight the change in order of the 50-mer contigs, resulting in placement of 3’ and 5’-most elements of Design 1 into the middle of Design 2; plain text indicates nonsense DNA sequence.27US_ACTIVE\131553842.V1Table 4: Contig designs containing both corn and soy sequences. Soy sequences from Table 2 are underlined and alternate between plain text and bolded to highlight differences in order of the two designs. Com sequences are not underlined and alternate between plain text and italicized text to highlight differences in order of the two designs. Design 3 alternates between com and soy derived sequences; Design 4 groups all 50-mers grouped by species.28US_ACTIVE\131553842.V1Table 5: K-mer probes enrich soy genome sequences of interest when k-mer probes are part of contigs containing both com and non-sense sequences with the soy probes.29US_ACTIVE\131553842.V1
[0083] The k-mer probes enriched all 10 of the regions of interest in the soy genome along with the control soy lectin gene (SEQ ID NO: 35). The non-deduplicated depth of the endogenous control gene (definition of 100%) was greater than 10,000 per analysis of 1 million raw reads per sample. The full probe sequences were synthesized as double stranded (dsDNA) and prepared as a probe library as described in Example 2. Enrichment of each of the 10 soy targets was comparable for each of the four contig designs (Table 5). For each contig design, all 10 probes enriched the sequence of interest. Probes derived from Contig Design 1 enriched the sequences between 15-123.4% of the control gene. Probes derived from Contig Design 2 enriched the sequences between 17.1-132.4% of the control gene. Probes derived from Contig Design 3 enriched the sequences between 20.2-128.7% of the control gene. Probes derived from Contig Design 4 enriched sequences between 17.4-127.8% of the control gene.
[0084] This example demonstrates that DNA probes designed using the described methodologies can enrich sequences in the soy genome and that the depth of enrichment is independent of the position of the designed probe contigs within a chimeric probe design, and independent of the presence or absence of adjoining probes relevant to targeted sequences. It also demonstrates that adjacent placement of sequences from alternate organisms or nonsense sequences does not interfere with the ability of probe contigs to enrich regions of interest.Canola Probes
[0085] The methodology described in the above examples was used to enrich regions of interest in the Brassica napus (canola) genome. Briefly, 50 bp k-mer probes were designed to unique regions on separate chromosomes. The probes were concatenated to form a contig, and a probe library was produced as described in previous examples. Canola genomic assembly sequences including the probe target and approximately 1000 bp 5’ and 3’ of the probe were used to map the enriched reads. The 10 canola probe contigs and their chromosome are provided in Table 6. A30US_ACTIVE\131553842.V1single probe design was created for both canola and cotton enrichment, with alternating 50-mer contig designs from the two species (Table 7).Table 6: K-mer probes used for enrichment of canola sequences. The chromosome location, probe sequence, and Peak Read Depth as a Percentage of the Control are provided.Table 7: Alternating 50-mer contig design containing sequences from cotton and canola. Canola- derived sequences are bold and underlined. Cotton-derived sequences are in plain text.31US_ACTIVE\131553842.V1
[0086] gDNA was extracted from wild-type (WT) plants from canola line 65037. As described in Example 3, the gDNA was purified and processed into gDNA libraries. The gDNA libraries were enriched using the designed k-mer probes and an endogenous canola gene, FAT4 (SEQ ID NO: 47). and enriched libraries suitable for Illumina sequencing were prepared for sequencing. Individual read-pairs from sequencing were mapped to the 10 targeted sequences (SEQ ID Nos: 36-45) and the endogenous control gene, canola FAT4. The peak read depths of the 10 target sequences were compared to the read depth of FAT4 of the same sample to calculate the peak read depth as a percentage of the control. Two technical replicates were carried out for the enrichment and sequencing. Results are shown in Table 6.
[0087] The k-mer probes enriched all ten of the regions of interest in the canola genome. Read depths compared to the canola FAT4 control ranged from 76.8% to 131.0% with an average enrichment rate of 96.9%.Cotton Probes
[0088] The methodology described in the above examples was used to enrich regions of interest in the Gossypium hirsutum (cotton) genome. Briefly, 50 bp k-mer probes were designed to unique regions on 10 separate chromosomes. Cotton genomic assembly sequences including the probe target and approximately 1000 bp 5’ and 3’ of the concatenated k-mer probes were used to map the enriched reads. The 10 cotton k-mer probes and their chromosome locations of their target sequences are provided in Table 8.32US_ACTIVE\131553842.V1Table 8: K-mer probes enrich cotton sequences. 10 k-mer probes were designed to enrich regions of interest in the cotton chromosomes. The peak reads are provided as a percentage of the control gene ACP.
[0089] gDNA was extracted from WT plants from cotton line DP393. As described in Example 3, the gDNA was purified and processed into gDNA libraries. The gDNA libraries were enriched using the dual-use canola plus cotton design from above and an endogenous cotton probe; enriched libraries suitable for Illumina sequencing were prepared for sequencing. Individual read-pairs from sequencing were mapped to the 10 targeted sequences and the endogenous control gene, cotton33US_ACTIVE\131553842.V1ACP (SEQ TD NO: 58). The peak read depths of the 10 target genes were compared to the read depth of ACP of the same sample to calculate the peak read depth as a percentage of the control. Two technical replicates were carried out for the enrichment and sequencing. Results are shown in Table 8.
[0090] The k-mer probes enriched all ten of the regions of interest in the cotton genome along with the control gene. Read depths compared to the cotton ACP control ranged from 67.5% to 87.0% with an average enrichment rate of 96.9%. These results and the previous examples demonstrate that the k-mer probe design methodology can be used to design probes capable of enriching sequences of multiple plant species.Example 6: 20-mer through 90-mer probe design and testing
[0091] To demonstrate that the k-mer probe methodology is applicable for design of probes of variable lengths, a series of probes were designed using the methodology described above with lengths of 20bp, 30bp, 40bp, 50bp, 60bp, 70bp, 80bp, and 90bp. Each of the probes were designed using the previously described methodology consistent with the k-mer analysis. The 10 probes, for each of the 8 sets of probes, targeted the middle region of approximately 2000 bp regions of interest in the 10 corn chromosomes. The sequences of the 90-mer probes are provided in Table 9. Alternate probes of lengths 80, 70, 60, 50, 40, 30 and 20 bp were designed centered on the 90 bp probe sequence designs; chimeric designs of approximately 1000 bp were made from each set of 10 probe designs by surrounding each shorter sequence with nonsense sequence on either side to produce a 100 bp probe plus nonsense sequence; ten 100 bp designs were concatenated to produce -1000 bp probe designs with differing lengths of probe contig designs. As examples, the 30 bp, 50 bp and 90 bp designs are shown in Table 10.Table 9: 90-mer probes enrich genomic locations of interest in com34US_ACTIVE\131553842.V1Table 10: Illustrative designs: 30-mer, 50-mer or 90-mer contig probes embedded in nonsense sequence to produce -1000 bp chimeric probe designs. Bolded and underlined sequences denote the k-mer probes. Plain text sequences are non-sense DNA.35US_ACTIVE\131553842.V136US_ACTIVE\131553842.V1
[0092] gDNA was extracted from WT plants from com line 01DKD2. As described in Example 3, the gDNA was purified and processed into gDNA libraries. The genomic DNA libraries were enriched using the 20-mer through 90-mer designs from above and an endogenous com probe; enriched libraries suitable for Illumina sequencing were prepared for sequencing. Individual readpairs from sequencing were mapped to the 10 targeted sequences and the endogenous control gene, cotton PDC-3 (SEQ ID NO: 72). The peak read depths of the 10 target genes were compared to the read depth of PDC-3 of the same sample to calculate the peak read depth as a percentage of the control. Two technical replicates were carried out for the enrichment and sequencing. Results are shown in Table 11.
[0093] The relative read depths of the sequences enriched by the 90 bp probes normalized to the depth of the PDC-3 control ranged from 80.8% to 172.2% with an average read depth of 132.4%. This indicates that k-mer lengths longer than 50bp are capable of enriching genomic DNA sites of interest.Table 11. Probe of normalized enrichment rates for probes of 20-mer, 30-mer, 40-mer, 50-mer, 60-mer, 70-mer, 80-mer, and 90-mer lengths targeted to the same regions on Chromosomes 1-10 of the corn genome.37US_ACTIVE\131553842.V1
[0094] For all 10 targets, 20-mer probe contigs failed to measurably enrich the targets of interest at the standard stringency of enrichment (Table 11). For all ten targets, every probe 30 bp or longer was shown to enrich the targets to usable peak depths, and the probes enriched at an average normalized depth of 67.7%. K-mer informed designs of 30 bp or greater can enrich target sequences using the standard enrichment conditions, and the 40 to 90-mer probes enriching the sequences at a higher normalized read depth compared to the shorter length k-mers. Shorter regions that fit the k-mer criteria can be used in regions of high repetitiveness with an unrepetitive sequence as short as about 30 bp while longer reads also enrich the gDNA.* *
[0095] All of the methods disclosed and claimed herein can be made and executed without undue experimentation in light of the present disclosure. While the compositions and methods of this invention have been described in terms of preferred embodiments or aspects, it will be apparent to those of skill in the art that variations may be applied to the methods and in the steps or in the sequence of steps of the method described herein without departing from the concept, spirit, and scope of the invention. More specifically, it will be apparent that certain agents which are both chemically and physiologically related may be substituted for the agents described herein while the same or similar results would be achieved. All such similar substitutes and modifications apparent to those skilled in the art are deemed to be within the spirit, scope and concept of the invention as defined by the appended claims.38US_ACTIVE\131553842.V1
Claims
CLAIMS1. A method of producing a library of DNA probes, the method comprising: a) identifying all or substantially all k-mer polynucleotide sequences between about 18 to about 24 nucleotides in length within a DNA sequence of interest within a genome; b) determining the number of times the k-mer polynucleotide sequences are found within the genome; c) identifying two or more contigs of at least about 50 nucleotides in length, wherein each k-mer polynucleotide sequence comprised in the two or more contigs is found less than or equal to about five times within the genome; d) producing a chimeric double- stranded DNA sequence comprising the two or more contigs; and e) fragmenting the chimeric double- stranded DNA sequence to produce a library of DNA probes.
2. The method of claim 1, wherein the method comprises identifying all or substantially all k-mers of about 20 nucleotides in length within the DNA sequence of interest.
3. The method of claim 1, wherein each k-mer polynucleotide sequence comprised in the two or more contigs is found less than or equal to four times, less than or equal to three times, less than or equal to two times, or less than or equal to one time within the genome.
4. The method of claim 1, wherein the contigs are at least about 55, at least about 65, at least about 75, at least about 85, at least about 95, or at least about 100 nucleotides, in length.
5. The method of claim 4, wherein the contigs are at least about 80 nucleotides in length.
6. The method of claim 1, wherein the method comprises producing a chimeric doublestranded DNA sequence comprising three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, twenty or more, twenty-five or more, or thirty or more, contigs.39US_ACTIVE\131553842.V17. The method of claim 1 , wherein fragmenting the chimeric double-stranded DNA sequence comprises endonuclease digestion.
8. The method of claim 1, wherein fragmenting the chimeric double- stranded DNA sequence comprises shearing.
9. The method of claim 1, wherein the method comprises labeling the fragmented doublestranded DNA sequences.
10. The method of claim 9, wherein labeling comprises biotinylating the fragmented doublestranded DNA sequences.
11. The method of claim 1, wherein the k-mer polynucleotide sequences found less than about five times within the genome do not comprise mononucleotide repeats or dinucleotide repeats longer than about ten nucleotides in length, trinucleotide repeats longer than twelve nucleotides in length, or palindromic repeats capable of forming a hairpin longer than 6 nucleotides in length.
12. The method of claim 1, wherein the genome is a plant genome.
13. A library of DNA probes produced by the method of claim 1.
14. The library of DNA probes of claim 13, wherein the library of probes comprises doublestranded DNA sequences capable of hybridizing to at least two, at least ten, at least twenty, at least fifty, at least one hundred, or at least two hundred DNA sequences of interest, within the genome.
15. A method of enriching the presence of a DNA sequence of interest in a sample, the method comprising: a) contacting the sample with a library of DNA probes produced by the method of claim 1, wherein the sample comprises fragmented genomic DNA; b) subjecting the sample and the library of DNA probes to hybridization conditions; and c) enriching the fragmented genomic DNA hybridized to the library of DNA probes in the sample.40US_ACTIVE\131553842.V116. The method of claim 15, wherein the hybridization conditions comprise high stringency conditions.
17. A method of producing a library of DNA probes, the method comprising: a) identifying all or substantially all k-mer polynucleotide sequences between about 18 to about 24 nucleotides in length within a plasmid DNA sequence of interest; b) determining the number of times the k-mer polynucleotide sequences are found within a genome; c) identifying two or more contigs of at least about 50 nucleotides in length, wherein each k-mer polynucleotide sequence comprised in the two or more contigs is found less than or equal to about five times within the genome; d) producing a chimeric double-stranded DNA sequence comprising the two or more contigs; and e) fragmenting the chimeric double- stranded DNA sequence to produce a library of DNA probes.
18. The method of claim 17, wherein the chimeric double-stranded DNA sequence further comprises one or more additional contigs, wherein each k-mer polynucleotide sequence comprised within the one or more additional contigs is found less than or equal to about one time within the genome.
19. The method of claim 17, wherein the chimeric double-stranded DNA sequence further comprises at least one additional polynucleotide sequence, wherein the at least one additional polynucleotide sequence is located: between the two or more contigs; at the 5’ end of the chimeric double- stranded DNA sequence; or at the 3’ end of the chimeric double-stranded DNA sequence.
20. The method of claim 17, wherein the method further comprises labeling the fragmented double-stranded DNA sequences.
21. The method of claim 17, wherein the plasmid DNA comprises a transformation vector.41US_ACTIVE\131553842.V122. The method of claim 17, wherein the method comprises identifying all or substantially all k-mers of about 20 nucleotides in length within the plasmid DNA sequence of interest.
23. The method of claim 17, wherein each k-mer polynucleotide sequence comprised in the two or more contigs is found less than or equal to four times, less than or equal to three times, less than or equal to two times, or less than or equal to one time within the genome.
24. A library of DNA probes produced by the method of claim 17.
25. The library of DNA probes of claim 24, wherein the library of probes comprises doublestranded DNA sequences derived from at least one, at least two, at least three, at least four, at least five, or at least six plasmid DNA sequences of interest.
26. The library of DNA probes of claim 24, wherein the library of probes is capable of identifying a transgene insertion site.
27. A method of identifying a saturating sequence within a DNA sequence of interest, the method comprising: a) identifying all or substantially all k-mer polynucleotide sequences between about 18 to about 24 nucleotides in length within a DNA sequence of interest within a genome; b) determining the number of times the k-mer polynucleotide sequences are found within the genome; and c) identifying the k-mer polynucleotide sequences found more than one time within the genome.
28. The method of claim 27, wherein the method comprises identifying k-mer polynucleotide sequences found more than two times, more than three times, more than five times, or more than ten times, within the genome.42US_ACTIVE\131553842.V1