Depletion probes
Patent Information
- Application Number
- JP2024551954
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-03-01
- Filing Date
- 2023-03-01
- Publication Date
- 2026-03-02
AI Technical Summary
【0011】 本発明の局面は、RNA枯渇プローブのセットを提供する。例示的なセットは、ヘテロ二重鎖を形成するために枯渇されるべきRNAにハイブリダイズする複数のDNAオリゴを含み、ここで上記DNAオリゴは、(i)上記ヘテロ二重鎖に所定の範囲内の融解温度を与える、および(ii)オフターゲットRNAに対するマッチングを最小にするように選択された配列を有する。好ましくは、上記ヘテロ二重鎖は、上記RNA分子に沿って不均一な分布で形成される。上記セットにおけるヘテロ二重鎖間の間隔は、高度に変動性であり得る、および/または数学的に無作為であり得る。上記DNAオリゴは、隣接するヘテロ二重鎖間の間隔を最小にするように選択され得る。上記DNAオリゴは、低GCオリゴおよび高GCオリゴ(これは、低GCオリゴより高い融解温度を有する、および/または上記低GCオリゴより短い)の両方を含み得る。上記プローブセットは、所定の範囲内の推定される融解温度を有するRNA分子に相補的な候補配列のセットを生成する工程;参照オフターゲットRNAに対するマッチスコアに少なくとも部分的に基づいて、上記候補配列の各々にコスト関数を割り当てる工程;および上記RNA分子に沿った位置にマッピングする上記候補配列のセットを選択する工程であって、ここで上記セットは蓄積するコスト関数、介在ギャップ、および/または重複を最小にするアルゴリズムによって選択される工程によって、設計され得る。上記プローブセットは、リボソームRNA、ミトコンドリアRNA、グロビン転写物などのうちのいずれかを枯渇するように設計され得る。注目すべきことには、上記ヘテロ二重鎖は、任意の規則的な、反復する、または均一なパターンまたは間隔を形成しないRNA分子に沿った位置において形成される。
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] Technical Field The present disclosure relates to depletion probes. [Background technology]
[0002] background The transcriptome refers to the RNA transcripts of a cell or organism at a given time. Knowledge of the transcriptome can reveal active cellular processes and provide information about cellular regulation, proliferation and dysfunction.
[0003] Transcriptome analysis typically involves microarray technology, and more commonly next-generation sequencing technology (e.g., RNA-Seq). The general goal of RNA-Seq is to detect all of the diverse transcripts present (including mRNA and non-coding RNA), and to detect splice variants, mutations, mobile genetic elements, and expression levels during or at various developmental stages. Transcriptomics and RNA-Seq have applications in diagnostics, disease profiling, pathogen detection, evolutionary biology, and other research areas. For example, RNA-Seq can potentially identify genes that involve resistance to environmental stress (e.g., drought resistance in crops). In another example, transcriptomic profiling can provide information about the mechanisms of drug resistance, potentially revealing strategies to combat hospital-acquired antibiotic-resistant infections. Summary of the Invention [Means for solving the problem]
[0004] Abstract The present invention provides RNA depletion probes, i.e., sets of short DNA oligos that hybridize along the length of the RNA to be depleted from a sample. The DNA oligos mediate digestion of the target RNA by RNase H. The present invention is useful for depleting non-target RNA from a sample to facilitate capture of low abundance RNA or simply achieve higher yields of target RNA. Depletion probes according to the present invention are designed based on biochemistry and biophysical properties of the probe / target interaction, so that all of the depletion probes of a set exhibit uniform, consistent behavior. For example, the depletion probes can be designed to have melting temperatures (Tm) within a narrow predefined range. Probes according to the present invention are designed to achieve improved performance, often resulting in probe sets with irregular, even seemingly random, spacing along the RNA to be removed from a sample.
[0005] Despite the irregular spacing, the probe sets of the present invention outperform prior art probe sets due to a rational biology-based design that models Tm and screens reference RNA sequence data to omit candidate probe sequences with off-target matches. Prior art probe sets sacrifice complete deletion to achieve uniform probes that bind along the RNA to be deleted. Due to the limiting assumption that probe sets must be tiled along the length of the RNA with uniform spacing and melting temperatures, many probes bind poorly due to factors such as GC content, secondary structure, or length. The present invention addresses these issues by using a design methodology that works within a selected range of melting temperatures and has reduced off-target binding.
[0006] The probes of the present invention form heteroduplexes along one or more RNA molecules with non-predictable, seemingly random spacing. In fact, the probe sets of the present invention may leave long gaps of RNA, in some cases up to about 100 bases or more, but are still superior to prior art probe sets designed to follow some simple pattern of short, regular gaps along the target RNA. Furthermore, the probes of the present invention show improved depletion even in highly degraded samples. Another advantage of the probes designed according to the present invention is that it allows probe hybridization and RNA digestion to occur in one rapid step, as opposed to using heating / cooling cycles followed by lengthy and suboptimal digestion temperatures. Because the probe sets of the present invention use probes designed to exploit the actual sequence content and biophysical properties of heteroduplex formation, RNA removal using the probe sets of the present invention reliably removes highly abundant RNA from a sample. Because such tools can be used to reliably remove RNA (e.g., ribosomal RNA and / or globin transcripts), RNA-Seq assays using the probe sets of the present invention better detect their residual, and sometimes rare, mRNA transcripts. Because rare genetic events are markers for critical conditions (eg, cancer or pathogenic infections), the probe sets of the present invention are useful for early and accurate detection of critical biological information.
[0007] In a particular aspect, the present invention provides a method for depleting RNA. A preferred method comprises hybridizing a plurality of DNA oligos to a target RNA molecule in a sample to form a heteroduplex, wherein the DNA oligos have sequences selected to (i) provide the heteroduplex with a melting temperature within a predetermined range, and (ii) minimize the match to a reference off-target RNA. The method of the present invention further comprises digesting the RNA in the heteroduplex. The heteroduplex can be formed with a non-uniform distribution along the RNA molecule. The DNA oligos can include low GC oligos and high GC oligos that have a higher melting temperature than the low GC oligos. The DNA oligos can include low GC oligos and high GC oligos that are shorter than the low GC oligos. The DNA oligos can be selected to minimize the spacing between adjacent heteroduplexes and can be present in non-uniform concentrations.
[0008] In some embodiments, DNA oligos are selected by generating a set of candidate sequences that are complementary to the target RNA (in this case, the RNA to be removed) with predicted melting temperatures within a predetermined range; assigning a cost function to each of the candidate sequences based at least in part on the match score to a reference off-target RNA; and selecting a set of the candidate sequences that map to a position along the RNA molecule, where the set is selected by an algorithm that minimizes the cumulative cost function, intervening gaps, and / or overlaps. The DNA sequences can be selected by a computer system to provide DNA oligos that each uniquely anneal to an RNA molecule with a similar annealing temperature, with minimal or no overlaps, and with minimal or no annealing to off-target RNA. Optionally, the computer system further selects the sequences to provide the DNA oligos with a minimum length and / or to minimize gaps between heteroduplexes.
[0009] In some embodiments, DNA oligos are tested and shown not to stably hybridize to protein-coding transcripts in a standardized human reference RNA extract. The target RNA molecules may include any of ribosomal RNA, globin transcripts, or mitochondrial RNA. A preferred method includes performing the steps described to digest ribosomal RNA and globin transcripts from the sample. The digesting step may include treating the sample with RNAse H. Additionally, a preferred method may include reverse transcribing non-target RNA molecules from the sample after the digesting step. In a preferred embodiment, the heteroduplexes are formed at positions along the RNA molecule that do not form any regular, repeating, or uniform pattern or spacing. For example, the spacing between the heteroduplexes may be highly variable and / or mathematically random. The positioning of the heteroduplexes may leave one or more uncovered regions of the RNA molecule of at least about 10 to about 100 bases. Furthermore, the coverage does not have to be end-to-end, ie, the 5' and 3' terminal sequences need not be covered.
[0010] The method of the present invention can increase the ratio of poly-A-tailed RNA to non-coding RNA in the sample. At least two of the sequences of the DNA oligos are selected to hybridize with the RNA and be depleted at positions overlapping or adjacent to each other. The DNA oligos used in the present invention can have lengths ranging from about 20 bases to about 40 bases, such as between 18 and 44, inclusive. The DNA oligos can be present in the mixture at non-equimolar concentrations. For example, based on empirical observations, some of the probes can be boosted, e.g., doubled in concentration.
[0011] An aspect of the present invention provides a set of RNA depletion probes. An exemplary set includes a plurality of DNA oligos that hybridize to the RNA to be depleted to form a heteroduplex, where the DNA oligos have sequences selected to (i) provide the heteroduplex with a melting temperature within a predetermined range, and (ii) minimize matching to off-target RNA. Preferably, the heteroduplexes are formed with a non-uniform distribution along the RNA molecule. The spacing between heteroduplexes in the set can be highly variable and / or mathematically random. The DNA oligos can be selected to minimize the spacing between adjacent heteroduplexes. The DNA oligos can include both low GC oligos and high GC oligos (which have a higher melting temperature and / or are shorter than the low GC oligos). The probe set can be designed by generating a set of candidate sequences that are complementary to an RNA molecule with a predicted melting temperature within a predetermined range; assigning a cost function to each of the candidate sequences based at least in part on the match score to a reference off-target RNA; and selecting a set of the candidate sequences that map to a position along the RNA molecule, where the set is selected by an algorithm that minimizes the cumulative cost function, intervening gaps, and / or overlaps. The probe set can be designed to deplete any of ribosomal RNA, mitochondrial RNA, globin transcripts, etc. Of note, the heteroduplexes are formed at positions along the RNA molecule that do not form any regular, repeating, or uniform pattern or interval. [Brief description of the drawings]
[0012] [Figure 1] FIG. 1 illustrates diagrammatically a method for probe design according to certain embodiments.
[0013] [Diagram 2] FIG. 2 shows the steps of the method for probe selection.
[0014] [Diagram 3] FIG. 3 shows the distribution of inter-probe gap sizes.
[0015] [Figure 4] FIG. 4 is a histogram of gap sizes.
[0016] [Diagram 5] FIG. 5 is a histogram of probe lengths.
[0017] [Figure 6] FIG. 6 illustrates a diagram of the Bloom filter design methodology.
[0018] [Figure 7] FIG. 7 describes an example of a method for probe design.
[0019] [Figure 8] FIG. 8 shows ribosomal RNA depletion from whole blood.
[0020] [Figure 9] FIG. 9 shows the results of hemoglobin RNA depletion from whole blood.
[0021] [Figure 10] FIG. 10 shows rRNA depletion from formalin-fixed paraffin-embedded (FFPE) tissue with 10 ng RNA.
[0022] [Figure 11] FIG. 11 shows rRNA depletion from FFPE tissue with 400 ng RNA.
[0023] [Figure 12] FIG. 12 shows rRNA depletion from universal mouse and rat references. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0024] Detailed Description The present disclosure relates to methods for designing and using RNA depletion probes, as well as probe sets designed according to such methods. The methods and materials of the present disclosure are useful for removing species of RNA molecules from a sample. The methods of the present invention have particular application in transcriptomics and certain RNA-based workflows intended to identify or quantify a certain population of RNA molecules. For example, it is often the case that a particular RNA molecule (e.g., mRNA) in a sample needs to be isolated. The depletion probes of the present invention are useful, for example, for depleting ribosomal catalytic RNAs (which do not themselves code for proteins) that may interfere with the detection of target mRNAs in a sample. As an example, one approach to isolate and evaluate mRNAs is to use polyA-specific priming with oligo-dT as an approach to avoid rRNA molecules. However, since the rRNAs are still present and interfere, there are various workflows that benefit from using target capture other than polyT primer hybridization to the mRNA polyA tail. Furthermore, partially degraded samples (e.g., formalin-fixed paraffin-embedded (FFPE) samples) frequently have 3' polyA tails separated from the 5' ends of transcripts, rendering information from the 5' ends inaccessible to techniques that rely on polyA enrichment. Furthermore, certain RNA species (e.g., long non-coding RNAs (lncRNAs)) do not contain polyA tails. In such instances, the lack of ability to deplete non-target RNAs makes target detection more challenging.
[0025] The methods and compositions of the present invention provide tools for depleting any RNA species from a sample. RNA depletion according to embodiments of the present invention works by using a "depletion probe" or DNA oligo designed to hybridize to the RNA molecule to be depleted. The DNA oligo anneals to the RNA to form a heteroduplex. RNase H recognizes the DNA-RNA heteroduplex and cleaves the RNA in the heteroduplex. RNA depletion using RNase H is attractive because it does not require that the RNA to be preserved contain a polyA tail, and the presence of a polyA tail on the RNA molecule targeted for depletion is not critical. Thus, even any mRNA transcript (e.g., globin) can be depleted by RNase H depletion. Moreover, the methods of the present invention allow lncRNA and mRNA enrichment even in highly degraded samples.
[0026] According to the present invention, RNase H depletion utilizes depletion probes, i.e., short DNA oligos that hybridize to the RNA to be depleted and form a heteroduplex. Some conventional approaches use a set of depletion probes that are tiled along the RNA molecule. The use of tiled probes is attractive due to the possibility that the entire target molecule will be covered. Other approaches have suggested allowing a certain spacing between the probes, for example leaving an inter-probe gap of a certain number of bases between each adjacent pair of probe targets. Furthermore, some depletion probe sets may include probes that overlap with each other and are apparently described as super-coverage (e.g., coverage >1). In such descriptions, "overlap" is a recognized term, but the choice of phrase is misleading to the extent that it describes two probes hybridizing to one target molecule with one of them leaving a single-stranded end unbound, where the other probe is fully bound. In practical use, each of the so-called overlapping probes will find independent copies of the same type of target RNA molecule, such that each is likely to bind fully to one molecule, where the term overlapping (and as used herein) simply acknowledges that the DNA oligos are designed to complementary regions where, for two oligos, it is their cognate targets that overlap.
[0027] Regardless of the probe tiling strategy, whether adjacent, uniformly spaced, or overlapping, existing probe designs reveal a design paradigm of determining which coverage pattern is most likely to be suitable and applying that coverage pattern to cover the molecule to be depleted. The coverage pattern may require that the probe targets are overlapping, adjacent, or regularly spaced. Once the pattern is selected, it is applied to the molecule to design the probe, and the probe is synthesized and used for RNA depletion.
[0028] The present invention recognizes that conventional probe coverage strategies create laboratory reagents that impose limitations on their use in analytical assays. For example, existing depletion probe sets can only be used under very limited conditions due to the diversity of GC content or due to RNA secondary structure. Also, optimizing the reaction temperature for the majority of probes means that the reaction is carried out at a temperature that is too high or too low for some of the probes. Conventional probe coverage strategies are not well suited to potential regions of sequence similarity between the RNA to be depleted and the RNA of interest.
[0029] The present invention addresses such limitations. The present invention provides a method of probe design that first relies on the biophysical properties of the molecules to be depleted and preserved. The method of the present invention uses sequence data on RNA molecules, also uses reference RNA data, and applies modeling to predict probes that hybridize to target RNA without marking for destruction the mRNA to be preserved.
[0030] Probes are designed based on biophysical properties. The methods and compositions of the present invention use sequence information, GC richness, reference sequences, and modeled / predicted biochemical behavior to generate and select well-functioning probe sets with uniform, consistent hybridization throughout the probe set and no off-target hybridization. The key insight of the present invention is that probe design does not require spatial patterns to be indiscriminately projected onto the length of the RNA molecule. Instead, desired properties (e.g., melting temperature range) are established, and then a set of probes with those desired properties is designed. For example, a metaset containing all possible probes with desired properties can be generated by an algorithm, for example, from sequence data, and then the metaset can be reduced based on mapping, filtering, and / or scoring to select a smaller but useful probe set. The selected set of useful probes is synthesized as DNA oligos and used to deplete selected RNAs from the sample, thus enriching the RNA of interest. Importantly, a reference dataset of sequences of non-target RNAs can be used, for example, in a filtering or scoring step to deplete probes from the metaset that potentially deplete rare mRNAs of interest from the sample.
[0031] In the initial step of generating all or a substantial number of possible probes, some embodiments begin by establishing a near uniform Tm (or Tm range) between any probe and its target, and also aim to minimize the Tm between each probe and any potential off-targets.
[0032] Many approaches can be used to arrive at the file set. The selected probes preferably bind across the target RNA. An algorithm can be used to find a probe set that binds across all regions of the target (e.g., 45S, mitochondrial rRNA, or globin transcript). Certain embodiments use a graph-based algorithm that seeks to minimize (without specific constraints) the spacing between probes and the potential Tm with off-targets (e.g., perhaps predicted by percent homology with the strongest non-target BLAST hit). Other target-spanning algorithms can be used to arrive at the set of probes. The probe set can be partially supplemented with one or even more rounds of similar design and / or manual work based on empirical data. For example, some probes can be "boosted" in concentration to improve performance, for example, based on results in initial test assays.
[0033] Using such methodology, the present invention provides a method for making depletion probes, a set of depletion probes, and a method for RNaseH-based depletion of rRNA and other highly expressed RNAs, for example, from an RNA-Seq library, using the depletion probes. Any suitable analysis package or algorithm can be used in probe design. A preferred embodiment generates candidate probes in a manner that is informed by the sequence information of the target and a set of reference RNA sequence information for a sample or organism (as an interesting aside, it does not matter if the algorithm uses a reference RNA sequence data set that also includes the target). The present disclosure provides a method for designing probes, and probe sets that are designed by a method that is systematic and rational, and allows for rapid design iteration to optimize the performance of the probe set.
[0034] At least three basic approaches to designing probes are contemplated. The first approach utilizes Bloom filters to identify unique k-mers (Bloom filtered). The second approach generates a set of probes with balanced melting temperatures (Tm) that are filtered by a basic local alignment search tool (e.g., BLAST or related algorithms) (BLAST filtered) to avoid creating probes that deplete non-target RNAs. The third approach generates a set of probes with balanced melting temperatures (Tm) that score the probes using BLAST and then selects a probe set that traverses the target RNA while minimizing the score (BLAST scored).
[0035] In these approaches, potential probes can be screened for uniqueness to RefSeq human RNA (manually masked for targeted sequences). In certain exemplary embodiments, RNA28SNx and RNA18SNx were used as development models. Probes for other RNAs (e.g., 5S and mitochondrial rRNA probes) can also be designed similarly. After generating an initial set of probes, a final candidate set that covers the target can be selected, for example, using bioinformatics programming techniques known to those skilled in the art. A preferred embodiment uses BLAST scoring methodology.
[0036] FIG. 1 diagrammatically illustrates a BLAST scoring method 101 for probe design according to certain embodiments. The method begins by selecting a target and obtaining sequence data. This can be done by accessing a publicly available database of genetic information (e.g., GenBank or Ensemble). For example, the method 101 can be used to deplete Homo sapiens 18S ribosomal N3 RNA from a biological sample. The computer system of the present invention can be used to retrieve the NCBI reference sequence NR_146152.1. A software package (e.g., BioPerl or BioRuby) can be used to import sequence data (e.g., the RNA18SN3 FASTA file for 18S ribosomal RNA N2, stored in GenBank under NCBI reference sequence NR_146152.1). This file can be brought into the system and set as a target. The system can be operated to bring in multiple targets (e.g., essentially simultaneously, e.g., in one "run"). Preferred embodiments target at least two or three RNA molecules each (eg, the longest ribosomal RNA and a globin transcript).
[0037] Once the target is set, method 101 includes setting boundaries for the melting temperature. This can be done by setting a target temperature and then allowing a range such as 50°C ±3. The system then goes through the target file and generates all possible strings that, when synthesized as DNA oligos, will have a Tm within the boundaries. In one simple embodiment, a software package on the computer system starts at the 5' end of the target file, reads each possible k-mer, and calculates the tm for each. Any suitable calculation can be performed. In some embodiments, the system uses [4(G+C) + 2(A+T)]. To illustrate with a trivial example, NR_146152.1 starts with a T. Thus, the first possible k-mer is a 1-mer consisting of a T. The calculation returns 2, which is outside the boundaries (not within 3°C of 50), so the system does not write the k-mer to the file. The system then reads the 5' 2-mer, which is a TA, and so on. All strings that result in a Tm within the boundaries are written to the file, and then the system moves to position 2 in the target file.
[0038] For this step, it may be preferable to establish a minimum (and / or maximum) length. The system iterates over the target file and writes all possible k-mers with Tm that fall within the boundaries to the target file. For example, all probes that fall within the boundaries may be written to a FASTA file.
[0039] The method may proceed to BLAST each probe against reference RNA sequence data. In one simple example, the BLAST score is used to simply omit probes that match non-target RNA above a certain threshold. In certain embodiments, the BLAST result is used to score probes (e.g., [0..1]), where the score is used as the cost in a cost function. The computer system that performs the operation may preferably use blastn-short.
[0040] The BLAST scores can be used to filter out probes or to assign scores to the probes that are used by an algorithm to select a set of probes to cover a target RNA.
[0041] The algorithmic task of the computer system is to cover the target RNA (conveniently from the 5' to the 3' end), a task that may be represented to the computer as traversing a graph where the probes are nodes and any inter-probe gaps are edges (or, trivially, vice versa). The system attempts to traverse the graph using an arbitrary cost as a cost function for using that probe. This may be done using a version of the Dijsktra algorithm where the edge weight is the "distance" between two nodes plus the "cost" of visiting the target node. In some embodiments, the edge weight includes a gap opening value and a separate gap extension value. In some other embodiments, the cost of visiting a node may be based on its similarity to potential non-target RNAs.
[0042] FIG. 2 illustrates the steps of the method of probe selection. In the top panel, all probes are generated. For example, the GenBank file is read for all strings of a certain length that are calculated to have a Tm within the boundaries, and the strings are written to a FASTA file. In the figure, each square is a candidate probe, a sequence from the GenBank file. The shaded blocks are probes that are eliminated by BLAST against the reference RNA sequence data. Probes that match the reference are eliminated from consideration (directly exhausted in the BLAST filter, and effectively avoided in the BLAST score), and the algorithm tries to build a directed graph through the probes. Any suitable algorithm can be used. One particular embodiment uses a graph traversal algorithm that essentially explores a directed acyclic graph with an optimized cost, where the nodes and edges are assigned certain scores. In a BLAST scoring embodiment, probes that match the reference (shaded in FIG. 2) will have a relatively high cost. Since edges may have a cost proportional to their length, the algorithm may attempt to minimize gaps (as depicted, the rectangular "probes" are nodes and the arrows are edges). Nodes along the paths found are the probe sets that are selected. In this algorithm, the selected probe sets provide non-overlapping coverage of the target RNA molecule.
[0043] Graph algorithms are known in the art and have contributed to sequence analysis due to the natural fit between the linear nature of genetic information carried in nucleic acid sequences and the linear nature of paths through directed graphs. Graph platforms (e.g., Neo4J) are attractive here because they cross the software / hardware barrier. Software packages can read, write, and query (e.g., traverse) data stored as graphs, but graph storage makes non-standard use of computer hardware. For example, some graph platforms use index-free proximity where a graph query reads a node, but that node contains a reference to the physical location in physical memory where the adjacent node is written, so the traversal moves to the referenced location in the hardware. This is in comparison to relational databases that use lookup index tables written in familiar software to record relationships, and have index-free proximity using spatial addresses on the hardware in a way that is much faster than index tables. Because the algorithms of the present disclosure lend themselves to implementation in a graph platform, the methods of the present disclosure can potentially run very quickly and be scaled up to query any arbitrarily sized reference target and BLAST FASTA probes against any arbitrarily sized reference set without any difficulty.
[0044] In an exemplary embodiment, the cost function for traversing the graph was set to k*score^r (the k,r parameters can be specified from the command line, but the defaults are k=256, r=2). Effectively, k is equal to the distance penalty paid by leaving a gap of sqrt(k) bases when score=1 (the worst possible scenario). In the cost function, r defines the rate at which the score penalty is reduced, so that larger values of r mean that small amounts of small differences / mismatches are dealt with more liberally. The graph traversal minimizes the cost function, and the visited nodes become the probe set. The probe associated with the graph traversal with the optimal score (not necessarily the absolute minimum) is used.
[0045] Figure 3 shows the distribution of inter-probe gap sizes for different values of k in the cost function. Increasing the value of k appears to decrease the overall BLAST score of the probes, with a potential trade-off in overall coverage. The probe set can be manually considered. For example, a biologist may choose k=1000 since any match to an off-target RNA will get a very expensive score. Thus, the resulting probe set will not deplete rare mRNAs, even though a 28-base gap in ribosomal RNA (the final value in the k=1000 row in the figure) may remain undepleted after RNase H digestion.
[0046] Thus, the software package first analyzes the target sequence to generate a set of probes (candidate sequences complementary to the RNA molecule within the estimated melting temperature boundary), BLASTs the probes to assign a filter function to each of the candidate sequences based at least in part on the match score to the reference off-target RNA, and finds a cost-minimizing directed graph through the probes (candidate sequences) to select a set of probes that map to positions along the RNA molecule. As discussed, the set is selected by an algorithm that minimizes the cumulative cost function, intervening gaps, and overlaps. The algorithm described does not literally minimize overlaps, but simply avoids them. Other algorithms are within the scope of this disclosure. In fact, overlaps do not necessarily need to be avoided, since they are not expected to actually compete for binding to the same copy of the molecule. Using method 101, the probes are selected by a computer system to provide DNA oligos that each anneal uniquely to the RNA molecule, preferably with minimal or no overlaps, and with minimal or no annealing of off-target RNA.
[0047] A second method of probe design, the BLAST-filter, is also within the scope of the present invention. In the BLAST filter embodiment, a computer system can be used to first generate all probes within a given Tm range, for example, using an input target RNA sequence file and a Perl script. In any embodiment, the probe length can be set to have a default (e.g., 21), but it can also be varied, for example, a different value can be used, the value can be switched on-the-fly, or a range of values can be set. Similarly, the temperature parameter can be varied. For example, some embodiments use a script that allows a 3.0 degree variation around the initial value (i.e., if the minimum and maximum lengths do not allow a probe to be generated within the Tm initial value, the script will generate the next closest candidate within 3.0 degrees - this can be adjusted or eliminated). To generate all probes, they are designed by finding the first probe of the desired Tm starting from every base from 5'. An additional strategy useful in any embodiment to ensure that nothing is lost at the 3' end is to have the script generate probes from the last 250 bases of the 3' end. In various preferred embodiments, the script literally uniquifies the generated probe sets.
[0048] In this BLAST-filter embodiment, the system BLASTs probes against the RefSeq - target sequence with blastn-short. Probes with percent identity above a given threshold are filtered out. Results show that this strategy gives results with some similarity to the Bloom filter methodology. These BLAST-filtering methodologies can be used to generate probes with the most stringent identity criteria possible.
[0049] Any suitable approach can be used to generate probe sets. The insight of the present invention is to first use information about the biochemical properties of target RNA and off-target RNA to generate probe sets that actually perform the intended task of depleting target RNA and leaving off-target RNA. After generating probes according to biophysical properties, those skilled in the art can apply other criteria (e.g., cost function of Dijsktra algorithm that minimizes gaps) that promote results similar to the main objective (no gaps) of the prior art probe design strategy. Any suitable approach that uses biophysical properties can be used (e.g., Bloom-filter, BLAST-filter, or BLAST-score method is presented here). The method of the present invention provides probe sets that tend to have certain characteristic features.
[0050] Certain features of the probe sets designed according to the method of the present invention are characteristic and noteworthy. For example, the probe spacing appears to be highly variable and random (sequence of gaps below). Some probes are adjacent to one or more other probes (i.e., the 5' end abuts the 3' end of another probe and / or the 3' end abuts the 5' end of another probe). Some probes overlap with other probes (note the comments elsewhere about overlap). Many gaps are more than 30 nucleotides (max=111, see FIG. 4). The probes have various lengths (min=18, max=44, see FIG. 5). The Tm of the probes is fairly uniform and mostly falls into two discontinuous groups (generally, we found that probes targeting GC-rich segments are required to have higher Tm than those targeting more balanced segments). Coverage of all targets is not end-to-end. In practice, coverage of the 45S rRNA is missing 5 nucleotides at the 5' end and 47 nucleotides at the 3' end. The probes are not present in the mixture at equimolar concentrations - some have been boosted up to 10x (based on empirical data) to improve performance.
[0051] Some of these features are contrary to previous thinking about depletion probes. For example, in prior art probe designs, 30-base and even 111-base gaps are considered rare and should be avoided. Furthermore, random mixtures of gaps, flanking, and overlapping were treated as being avoided. Importantly, length uniformity is sacrificed, and the probes of the present invention have various lengths and substantially uniform melting temperatures. Because the probes of the present disclosure are designed first from biophysical properties, they are superior to prior art probes (e.g., substantially uniform Tm ensures consistent digestion by RNase H, because the intended heteroduplexes are all stable under similar conditions, whereas in the prior art's rigid tiling strategy, the probe and target in the AT-rich region may melt apart). As a consequence of the design paradigm, probe spacing is highly variable and appears random.
[0052] The disclosed strategy has distinct advantages. For example, the need to denature the target and / or probe is reduced - allowing for a reduction in workflow. The present invention eliminates the need to add enzymes while the reaction mix is heated - likely due to more efficient denaturation of off-target probe binding. Taken together, these reductions cut workflow time in half compared to competitors, with no apparent loss in data quality.
[0053] Figure 4 is a histogram of gap sizes determined for one particular RNA. As can be seen, the probe spacing is highly variable and appears random. A set of probes was designed for a particular RNA.The following are the interprobe gap size sequences where the probes form heteroduplexes with that particular RNA (negative values indicate overlap): 7, 10, 10, 10, 1, 9, 29, 18, 14, 15, 3, 21, 0, 0, 0, 7, 12, 17, 2, 1, 2, 3, 2, 33, 6, 18, 16, 34, 6, 9, 11, 2, 19, 3, 20, 11, 1, 6, 18, 12, 10, 13, 11, 13, 17, 6, 6, 8, 16, 14, 19, 10, 0, 11, 8, 2, 14, 5, 7, 8, 3, 9, 11, 21, 0, 3, 17, 12, 7, 13, 14, 15, 18, 1, 1, 0 ,2,3,11,4,13,18,19,6,5,5,14,1,0,0,18,19,13,15,6,12,3,111,1,1,46,21,12,12,8,11,12,9,0,1,7,3,2,0,14,9,0,7,7,20,0,7,1,3,1,1,0,0 ,1,0,10,3,3,26,4,11,9,10,1,2,1,5,0,0,5,19,7,17,10,5,6,0,3,8,5,0,6,6,6,7,13,5,29,0,21,29,3,7,5,12,7,8,12,0,9,4,7,9,5,34,3,13,1 0, 9, 9, 4, 22, 36, 15, 26, 29, 8, 14, 17, 4, 44, 12, 11, 47, 9, 10, 9, 28, 25, 15, 0, 19, 22, 15, 37, 14, 12, 33, 5, 29, 39, 45, 7, 41, 25, 13, 32, 33, 10, 16, 39, 0, 0 ,3,18,24,7,20,7,4,3,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,25,0,0,0,0,0,0,0,0,0,-28,19,10,0,0,0,16,0,4,14,10,10,22,-9,0,18,4,0,0,3,8,2,4,1 4, 8, 2, 2, 6, 3, 5, 8, 8, 25, 25, 14, 4, 5, -8, 2, 13, 12, 16, 2, 2, 5, 9, 6, 2, 9, 6, 10, 11, 3, 8, 4, 4, 4, 5, 3, 3, 0, 0, 0, 3, 0, 10, 0, 4, 0, 0, 0, 5, 10, 0, 0, 0, 0, 9, 6, 13, 0, 1, -9, 8, 12, 13, 26, 11, 9, 45, -16, -5, -25, 4, 0, 18, 11, 0, 34, 14, 1, 0, 0, 0, 0, 1, 4, 8, 1, 0, 3, 5, 22, 14, 37, 11, 16, 9, 7, 0, 45, 10, 6, 16, 45, 22, 0, 28.There is no apparent uniformity or pattern: the probe sets were not designed to form a pattern along the molecule.
[0054] Similar to the gap size, the probe length is variable.
[0055] Figure 5 shows a histogram of probe lengths for a set designed for a particular RNA, where d$Length is the probe length. The gap histogram and the probe length histogram reveal an important property of the probe sets of the invention: the gaps are not uniform. In fact, the distribution of gaps is neither regular nor standard.
[0056] Note from the histogram that there is at least one gap above 100. The mode is likely zero. There are also negative values. Similarly, the probe length includes at least 1 at 44, and an apparent mode just above 20. The histogram shows such a pattern because the probes were not designed to gap or length specifications. In fact, biophysical properties may affect not only sequence selection, but also probe length. For example, high GC oligos may be allowed to be made at higher ends within the Tm boundary, but shorter than low GC oligos. Depending on the algorithm, the system may try to minimize the spacing between adjacent heteroduplexes, but may nevertheless give such a histogram (with values >100).
[0057] Using the probe of the present disclosure, RNA depletion can be performed by hybridizing multiple DNA oligos to target RNA molecules in a sample to form heteroduplexes, where the DNA oligos have sequences selected to (i) provide the heteroduplex with a melting temperature within a predetermined range and (ii) minimize match to reference off-target RNA; and digesting the RNA in the heteroduplex. These steps can be performed within a single-cell RNA-Seq (scRNA-Seq) workflow, for example, the method can include separating single cells into fluid partitions, releasing nucleic acid from cells in the partitions (e.g., by chemical or thermal cell lysis), and depleting the target RNA molecules. The fluid partitions can be, for example, emulsions or microfluidic devices, or droplets in wells, for example, plates such as multi-well plates or picotiter plates. The method may include introducing reaction reagents (e.g., any of the lysis reagents, depletion probes of the invention, RNase H, capture oligos, primers, template switching oligos, tagmentation reagents, reverse transcriptase, dNTPs, etc.) into the wells upon droplet formation by pipetting or by a droplet merger. Any of the primers or oligos may include barcodes, sequencing adapters, universal priming sites, etc. For example, in some embodiments, every RNA transcript is captured by an oligo containing a unique or near-unique sequence that functions as a unique molecular identifier, and every cDNA is given a partition barcode or a cell barcode.
[0058] Using the depletion probes of the present invention, heteroduplexes may form and show uneven distribution along the RNA molecule. RNase H digests the target RNA. Preferably, the DNA oligos (i.e., the depletion probes) have been tested and shown not to stably hybridize to protein-coding transcripts in standardized human reference RNA extracts. After digestion, overabundant RNA (e.g., ribosomal or globin) is substantially depleted. The depletion of transcripts is copied into cDNA (optionally barcoded by molecules and / or cells). Downstream method steps may further copy or amplify the cDNA molecules. For example, the cDNA molecules may be amplified with, for example, sequencing adaptors, such as Illumina Y adaptors, at the ends of library members to create a sequencing library.
[0059] The method described herein is useful for providing a set of RNA depletion probes comprising a plurality of DNA oligos that can hybridize to a target RNA molecule to form a heteroduplex. The DNA oligos have sequences selected to (i) provide the heteroduplex with a melting temperature within a predetermined range, and (ii) minimize matches to off-target RNA. Preferably, the heteroduplexes are formed with a non-uniform distribution along the RNA molecule, e.g., the spacing between the heteroduplexes can be highly variable and / or mathematically random. The probes can be algorithmically or even automatically designed by a computer system operable to generate a set of candidate sequences that are complementary to an RNA molecule with an estimated melting temperature within a predetermined range; assign a filtering function to each of the candidate sequences based at least in part on the match score to the reference off-target RNA; and select a set of the candidate sequences that map to a position along the RNA molecule, where the set is selected by an algorithm that minimizes the accumulated cost function, intervening gaps, and overlaps. The computer system may output the selected set as any suitable file, such as a FASTA file, a file formatted for input to an oligonucleotide synthesis instrument, or a format accepted by a commercial oligonucleotide supplier, such as Integrated DNA Technologies (IDT). Furthermore, in some embodiments, the computer system may include an application programming interface that selects the probe set and transmits a well-formed IDT order via Internet protocols. Once synthesized, the probe set may be provided in packaging for use in a laboratory or for shipping or sale. For example, the probe set may be lyophilized in a tube or in a solution in a tube, where the tube may be any suitable tube, such as a 1.5 mL microcentrifuge tube sold under the trademark EPPENDORF. The probe set may then be used in a method of performing RNA depletion, as described above.
[0060] As described above, the probe set can be selected by a methodology using Bloom filter, BLAST filter, or BLAST scoring. A preferred BLAST scoring embodiment is detailed above. A Bloom filter embodiment is described in the examples. Besides Bloom filter, BLAST filter or BLAST scoring methodology, or any other suitable methodology, can be used to provide depletion probes. The present invention provides RNA depletion probes, i.e., a set of short DNA oligos that hybridize along the length of target RNA and mediate digestion of target RNA by RNase H to deplete overabundant RNA molecules from a sample. Since the depletion probes according to the present invention are first designed based on the biochemistry and biophysical properties of the probes, all of the depletion probes of the set exhibit substantially uniform and consistent behavior in binding to target RNA in a sample. The probes are primarily designed to specific performance goals and biophysical properties, resulting in a probe set with irregular, even essentially random, spacing along the target RNA molecule. EXAMPLES
[0061] Working Example Example 1. Bloom filter Figure 6 illustrates the workflow for the Bloom filter design methodology. The workflow begins with setting the target RNA and having a RefSeq dataset. For example, a package such as BioPerl can be used to open a gateway to the latest version of RefSeq, as described in Pruitt, 2007, NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins, Nucleic Acids Res 35(Database):D61-5 (incorporated by reference). The target RNA is removed (or masked) from RefSeq.
[0062] Further, the targets are k-merized or parsed into k-mers, such as all possible k-mers. It is trivial for one of skill in the art to write a script, such as a Perl script, that writes all possible k-mers into a FASTA file. Depending on the platform, it may be preferable to convert the FASTA file into a binary alignment map (BAM) for downstream processing, e.g., by the PathSeq pipeline.
[0063] Use Bloom filters to search for k-mers from the RefSeq in the target. This can be done using the PathSeq GATK pipeline to build a Bloom filter package that identifies k-mers present in the target sequence that are not present elsewhere in the RefSeq.
[0064] A Bloom filter is a space-efficient probabilistic data structure useful for testing whether an element is a member of a set. There can be false positive matches, but no false negatives. An empty Bloom filter is a bit array with m bits all set to 0. Also, k different hash functions must be defined, each of which maps or hashes some set element to one of the m array locations, generating a uniform random distribution. Typically, k is a small constant that depends on the desired false error rate ε, while m is proportional to k and the number of elements to be added. To add an element, the element is given to each of the k hash functions to obtain k array locations. The bits in all these locations are set to 1. To query an element (test whether it is in the set), the element is given to each of the k hash functions to obtain k array locations. If any of the bits in these locations are 0, the element is definitely not in the set; if so, all of the bits are set to 1 when the element is inserted. If all are 1, then either the element is present in the set or the bit was accidentally set to 1 during the insertion of another element, resulting in a false positive. With a simple Bloom filter, there is no way to distinguish between the two cases, but more advanced techniques can address this issue. See Najam, 2019, Pattern matching for DNA sequencing data using multiple Bloom filters, Biomed Res Int 14:7074387 (incorporated by reference).
[0065] Bloom filters are applied to k-mers. For example, Bloom filters can be used to simply eliminate k-mers found in RefSeq. Alternatively, Bloom filters can be used to find k-mers in RefSeq and assign them a score (e.g., some constant multiple of % match or some other such arbitrary score).
[0066] With the set of k-mers in a file or data structure, the system tiles the set of k-mers across the target RNA. This can be done as a graph traversal, which is well suited for computational implementation even on large (including very, very large) data volumes due to the unique way that graph platforms such as Neo4J access hardware, e.g., with index-free proximity. The system tiles the RNA sequence. Tiling an (RNA) sequence can be stated as the problem of traversing a directed graph where edges have costs. In some embodiments, the cost of an edge is initially the distance (in bases) between the end of a probe and the beginning of the next probe. The lowest cost path from the beginning to the end of the RNA can be most efficiently (and optimally) calculated using Dijkstra's algorithm (see, e.g., Rodriguez-Puente, 2013, Algorithm for shortest path search in geographic information systems by using reduced graphs, Springerplus 2:291, incorporated by reference). Dijkstra runs in O[|V| + |E|* log(|V|)]. Using the simple heuristic of only considering probes within a fixed distance of each other (as opposed to all-to-all), |V| ∈ O(|E|) - so the runtime complexity is O[|E|*log(|E|)].
[0067] Dijkstra's algorithm was implemented in a general manner that allows it to operate on any graph object whose nodes may hold any data type. It is noted that in the original implementation of the tiling algorithm, the graph was constructed such that the distance from probe x to probe y was start(y)-stop(x). This works, but Dijkstra is a greedy algorithm and may leave gaps at the end of its tiling path (i.e., if there are multiple optima, it chooses those that incur costs later in the optimization, and typically accumulates these in longer stretches). This is relatively easily defined by setting the distance between nodes to be [start(y)-stop(x)]^2, so that the cost grows more quickly when a single gap is widened than when opening many short gaps. Additionally or alternatively, it may be appropriate to assign a gap extension penalty separate from the gap formation penalty.
[0068] It is noted that the graph algorithm can read from the initial probe candidates, which cover every base in the target (e.g., RNA18SN3), but long stretches remain uncovered. For RNA18SN3, all filtered 21-mers cover all but 4 bases, but by tiling twice Dijkstra, 223 and 222 bases remain uncovered, respectively. Since the probes are not overlapping and there is only one probe starting at each position, it is very unlikely that a continuous coverage path can be found when the probes are filtered out. Interestingly, this does not require any restrictions on applicability, and the resulting probe sets can be used for RNA depletion. It is noted that in existing algorithms, when using an unfiltered set of probes (i.e., when every possible probe is presented to the tiler), the return is an optimal path with a small coverage gap at the end (if the length of the tiled sequence is not a multiple of the length of the probe, the sequence length is all that is determined in this case). Any input that includes a suitable tiling returns the same input. Manual curation of the results can be helpful, but manual curation of the results is likely to ensure that suitable results are obtained.
[0069] Example 2. Design workflow FIG. 7 describes an example of a method for human / mouse / rat (H / M / R) probe design. The method includes algorithmic design of depletion probes for human transcript targets; removing probes with high homology to undesired human transcript targets; algorithmic design of depletion probes for mouse / rat transcript targets; removing probes with high homology to undesired mouse / rat transcript targets; cross-checking human probes against the mouse / rat transcriptome and vice versa; and removing probes with high homology to undesired transcript targets in H / M / R. After such design methodology, the final probe set can be synthesized, used to remove unwanted RNA from the sample, or packaged for storage, shipping, distribution, or later use. For example, the probes can be placed in tubes or vials in an appropriate solution or lyophilized for later reconstitution. Such vials or tubes can contain other reagents, such as RNase H, as needed.
[0070] Example 3. Results Figure 8 shows ribosomal RNA depletion of total RNA from whole blood. The bar on the left, "Before RNA depletion," indicates that the sample was >80% ribosomal RNA. After rRNA depletion, as shown by the bar on the right, the sample was <5% rRNA and approximately 75% mRNA. Thus, these results indicate that the probe sets and methods of the present disclosure are effective for depleting ribosomal RNA from samples.
[0071] 9 shows the results of hemoglobin RNA depletion from whole blood using the probe set and method of the present disclosure. Before hemoglobin depletion, the sample contained about 35.75% residual hemoglobin RNA, as shown by the bar on the left. After depletion, the sample contained about 0.01% residual hemoglobin RNA. Thus, these results show that the probe set and method of the present disclosure are effective for depleting hemoglobin RNA from samples.
[0072] FIG. 10 shows that the methods and probe sets of the present invention are useful for robust ribosomal RNA depletion from total RNA samples derived from formalin-fixed paraffin-embedded (FFPE) tissues with 10 ng RNA.
[0073] FIG. 11 shows the results of a FFPE blank with 400 ng RNA.
[0074] For each of the four blocks with 10ng and 400ng total RNA per block obtained, in all cases, the % ribosomal RNA after depletion was less than or equal to about 7%, and in most cases, significantly less than 5%.Thus, these results show that the probe set and method of the present disclosure are effective for depleting RNA from FFPE samples.
[0075] Figure 12 shows quantitative results of ribosomal RNA depletion from total RNA derived from universal mouse and rat references. Using the method and probe set of the present invention, the mouse sample was depleted from about 85% ribosomal RNA to less than about 5% ribosomal RNA. The rat sample originally had about 80% ribosomal RNA, and the method and probe set of the present invention depleted the rat sample to less than about 5% ribosomal RNA. Thus, these results show that the probe set and method of the present disclosure are effective for depleting RNA from rat or mouse samples.
Claims
1. A method for depleting target RNA molecules in a sample, the method comprising: combining the sample with (i) a plurality of DNA oligos complementary to the target RNA molecules, and (ii) a nuclease; and after said combining, hybridizing DNA oligos from said plurality to said target RNA molecule to form a DNA / RNA heteroduplex; using said nuclease to digest RNA of said formed DNA / RNA heteroduplex to deplete said target RNA molecule; A method that encompasses (a) a DNA oligo of the plurality of DNA oligos has minimized annealing to off-target RNA molecules; (b) the plurality of DNA oligos (i) each oligo of the plurality of DNA oligos has a melting temperature within a predetermined range; (ii) the plurality of DNA oligos comprises oligos complementary to different regions of the target RNA molecule, and the DNA / RNA heteroduplex is formed in a non-uniform distribution along the target RNA molecule. having a sequence such as The method of claim 1.
3. A method for depleting target RNA molecules in a sample, the method comprising: (a) forming a DNA / RNA heteroduplex by contacting a plurality of DNA oligos with the target RNA molecule; wherein a DNA oligo of the plurality of DNA oligos is complementary to the target RNA molecule and has minimized annealing to off-target RNA molecules; The plurality of DNA oligos (i) the DNA oligos of the plurality of DNA oligos have melting temperatures within a predetermined range; (ii) the plurality of DNA oligos comprises DNA oligos complementary to different regions of the target RNA molecule, and the DNA / RNA heteroduplex is formed in a non-uniform distribution along the target RNA molecule. and (b) digesting the RNA in the formed DNA / RNA heteroduplex; A method that encompasses
4. The method described in claim 2 or 3, wherein the predetermined range is such that each of the plurality of DNA oligos has a melting temperature within 3.0 degrees of the initial melting temperature value.
5. The method described in claim 3, wherein the DNA oligos of the plurality of DNA oligos have minimized annealing to off-target RNA molecules.
6. A method described in any one of claims 1 to 3, wherein the target RNA molecule comprises ribosomal RNA, globin transcript, or mitochondrial RNA. (i) the nucleotide distance between one or more of the different regions is 10 to 100 nucleotides; and / or (ii) the mode of the nucleotide distance between the DNA / RNA heteroduplexes on the target RNA molecule is zero, and / or (iii) the spacing between the DNA / RNA heteroduplexes on the target RNA molecule is mathematically random; The method according to claim 2 or claim 3.
8. A method described in any one of claims 1 to 3, wherein the multiple DNA oligos have sequences that leave a gap of a predetermined number of bases between each adjacent pair of regions of the target RNA molecule, overlap each other by a predetermined number of bases, and / or leave one or more uncovered regions of the target RNA molecule.
9. A method described in any one of claims 1 to 3, wherein the DNA oligos of the multiple DNA oligos are present in the sample at non-equimolar concentrations.
10. A method according to any one of claims 1 to 3, wherein the plurality of DNA oligos include oligos having different lengths.
11. The method according to any one of claims 1 to 3, wherein each of the plurality of DNA oligos is 18 to 44 bases in length.
12. A method described in any one of claims 1 to 3, wherein each of the plurality of DNA oligos anneals to the target RNA molecule.
13. The method according to claim 1, further comprising a step of selecting the sequence of the DNA oligo, wherein the step of selecting comprises: generating a set of candidate sequences complementary to the target RNA molecule, wherein the candidate sequences have predicted melting temperatures within the predetermined range; assigning a cost function to each of the candidate sequences based at least in part on a match score to a reference off-target RNA; and selecting a sequence of said oligos from said candidate sequences of said plurality of DNA oligos that minimizes said cost function, intervening gaps, and overlaps; A method comprising:
14. The method of any one of claims 1 to 3, comprising the step of depleting additional target RNA molecules in the sample, the method further comprising the step of contacting the sample with an additional plurality of DNA oligos, each having a sequence complementary to the additional target RNA molecule and having minimized annealing to off-target RNA molecules, the method comprising: (a) forming additional DNA / RNA heteroduplexes by contacting the additional plurality of DNA oligos with the additional target RNA molecules; wherein the additional plurality of DNA oligos are (i) each oligo of the plurality of DNA oligos and each oligo of the additional plurality of DNA oligos has a melting temperature within a predetermined range; (ii) the additional plurality of DNA oligos comprises oligos complementary to different regions of the additional target RNA molecule, and the additional DNA / RNA heteroduplexes are formed in a non-uniform distribution along the additional target RNA molecule. and (b) digesting the RNA in the additional heteroduplex formed; The method includes:
15. A method for depleting target RNA molecules in a sample, said method comprising: (i) RNAse H, and (ii) a plurality of DNA oligos, each having a sequence selected to be complementary to the target RNA molecule and to minimize annealing to off-target molecules, the plurality of DNA oligos comprising: (a) the DNA oligo forms a DNA / RNA heteroduplex with the target RNA molecule; (b) each oligo of the plurality of DNA oligos has a melting temperature within a predetermined range; (c) the plurality of DNA oligos comprises oligos complementary to different regions of the target RNA molecule, and the DNA / RNA heteroduplex is formed in a non-uniform distribution along the target RNA molecule. A plurality of DNA oligos having sequences selected to contacting the method.
16. The method described in claim 15, wherein the multiple DNA oligos hybridize to and digest the target RNA molecule in a single step.
17. A computer-implemented method for designing a set of RNA depletion probes for depleting a target RNA molecule, wherein the RNA depletion probes are DNA oligos, the method comprising: obtaining a set of candidate DNA oligo sequences that contain sequences complementary to the target RNA molecule and have melting temperatures within a predetermined range; selecting from the set of candidate DNA oligo sequences a plurality of sequences that are complementary to different regions of the target RNA molecule and minimize annealing to off-target molecule sequences; It encompasses wherein the distribution of nucleotide distances between the different regions is non-uniform; method.
18. The method described in claim 17, wherein the set of RNA depletion probes is selected to provide the DNA oligos with a minimum length and / or to minimize gaps between the sequences of the target RNA molecule.
19. A method for depleting target RNA molecules in a sample, the method comprising: contacting the sample with a plurality of DNA oligos, each complementary to the target RNA molecule and having minimized annealing to off-target molecules, wherein the plurality of DNA oligos comprises: (i) the DNA oligo forms a DNA / RNA heteroduplex with the target RNA molecule; (ii) each oligo of the plurality of DNA oligos has a melting temperature within a predetermined range; (iii) the plurality of DNA oligos comprises oligos complementary to different regions of the target RNA molecule, and the DNA / RNA heteroduplexes formed at positions along the target RNA molecule have a non-uniform distribution. and digesting the RNA in the formed DNA / RNA heteroduplex; A method that encompasses