Method for Preparing a Pool of Personalized Off-Target Guide RNAs for CRISPR
By purifying mRNA from subject samples, preparing cDNA, amplifying target sequences, fragmenting DNA, ligating tags, hybridizing initiating oligonucleotide pools, removing unrelated sequences, preparing sgRNA pools, and using sgRNA-guided nucleic acid binding proteins to cleave the polynucleotide mixture and selecting uncut target polynucleotides, the problem of difficult to reduce the complexity and depth of sequencing libraries in the prior art is solved, and the cost of sequencing and data management is reduced.
Patent Information
- Application Number
- CN202080076212.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-31
- Filing Date
- 2020-10-19
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-10-19
AI Technical Summary
The prior art is difficult to effectively reduce the complexity and sequencing depth of sequencing libraries, resulting in high sequencing costs and data management costs.
The uncleaved target polynucleotide pool was prepared by purifying mRNA from subject samples, preparing cDNA, amplifying target sequences, fragmenting DNA, ligating tags, hybridizing initiating oligonucleotide pools, removing unrelated sequences, preparing sgRNA pools, and using sgRNA-guided nucleic acid binding proteins to cleave the polynucleotide mixture, uncleaved target polynucleotides were selected.
The enrichment of target polynucleotides is achieved, reducing the complexity and sequencing depth of the sequencing library, and significantly reducing the sequencing cost and data management cost.
Smart Images

Figure CN114616342B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to methods for obtaining an enriched population of target polynucleotides using synthetic single guide RNAs (sgRNAs) for an sgRNA-guided nucleic acid binding protein, and to methods for obtaining a pool of personalized, target-agnostic synthetic single guide RNAs (sgRNAs) for an sgRNA-guided nucleic acid binding protein. Also provided are kits comprising a pool of sgRNAs obtainable by the methods of the present invention, the use of a pool of sgRNAs obtainable by the methods of the present invention, and methods of monitoring a disease condition. Background Art
[0002] Next-generation sequencing (NGS) is a major driver of genetic and molecular research, including modern diagnostics especially in the field of cancer medicine. This technology provides a powerful method for studying DNA or RNA samples. New and improved methods and protocols have been developed to support various applications, including the analysis of genetic variation and sample-specific differences. To improve such methods, methods have been developed that aim to target-enrich sequencing libraries by focusing on specific sequences, transcripts, genes, or genomic subregions, or by eliminating unwanted sequences.
[0003] Targeted enrichment can be useful in some cases, for example, where it is desired to analyze a specific part of the entire genome. Efficient sequencing of the complete exome (all transcribed sequences) is a typical example of this method. Other examples include enrichment of specific transcripts, enrichment of mutation hotspots, or exclusion of interfering nucleic acid material. Specifically, in the context of personalized sequence determination or patient-based molecular monitoring, these targeted enrichment strategies are important, where (i) the transcriptome library of a patient is to be analyzed and monitored and (ii) only a subset of sequences in the library is diagnostically relevant.
[0004] Current techniques for targeted enrichment include (i) hybridization capture, where nucleic acid strands from an input sample are specifically hybridized in solution or on a solid support to pre-prepared DNA fragments complementary to the target region of interest, such that the sequences of interest can be physically captured and separated; (ii) selective circularization or molecular inversion probes (MIPs), where single-stranded DNA circles comprising the target region sequence are formed in a highly specific manner by gap filling and ligation chemistry, resulting in a structure with common DNA elements, which are then used to selectively amplify the target region of interest; and (iii) polymerase chain reaction (PCR) amplification, where PCR is directed to the target region of interest by performing multiplex long-range PCR in parallel, a limited number of standard multiplex PCRs, or highly multiplex PCR methods that amplify a very large number of short fragments (Mertes et al., 2011, Briefings in functional Genomics, 10, 6, 374 - 386).
[0005] Recently, the CRISPR / Cas technology has been used for targeted enrichment purposes, specifically for removing unwanted sequences from sequencing libraries.
[0006] The CRISPR / Cas technology is a new and highly versatile genome editing and epigenome editing tool based on the repurposing of the CRISPR / Cas (clustered regularly interspaced short palindromic repeats / Cas) bacterial immune system (Cong et al., 2013, Science, 339, 819 - 824). When complexed with a short RNA oligonucleotide called single-guide RNA (sgRNA), the Cas nuclease can introduce double-strand breaks (DSBs) at specific sgRNA complementary positions.
[0007] The CRISPR / Cas system has also been repurposed as a programmable restriction endonuclease to direct cleavage in a very precise and customized manner (Lee et al., 2015, Nucleic Acids Res., 43, 1–9).
[0008] Accordingly, methods using CRISPR / Cas have been developed that selectively deplete excess sequences in a process called depletion of abundant sequences by hybridization (DASH). DASH is used to remove targets, such as ribosomal RNA (rRNA) from mRNA-seq and wild-type KRAS background sequences from cancer samples, by directing targeted cleavage of the targets and preventing further amplification and sequencing of the targets (Gu et al., 2016. Genome Biol., 17, 41). According to Gu et al., using DASH after transposon-mediated fragmentation but before subsequent amplification steps (which depend on the presence of adapter sequences at both ends of the fragments) can prevent amplification of target sequences (mitochondrial rRNA), thus ensuring that they do not appear in the final sequencing library.
[0009] However, this enrichment method is suitable for removing only specific substances from sequencing libraries, while all other sequences remain in the library. Therefore, there is a need for a general personalized transcriptome-based enrichment and analysis method that enables effective reduction of the complexity of sequencing libraries, and thereby involves reduced sequencing depth and a manageable amount of data to be processed and thereby enables repeated implementation (e.g.) in monitoring methods. Summary of the Invention
[0010] The present invention addresses this need and provides a method for obtaining an enriched population of target polynucleotides, the method comprising: (i) purifying a population of mRNA molecules from a sample obtained from a subject; (ii) preparing cDNA from the mRNA molecules of step (i); (iii) amplifying one or more target sequences from the cDNA obtained in step (ii) to obtain a pool of DNA molecules; (iv) fragmenting the amplified DNA molecules, preferably to a size of 20 to 30 bp; (v) ligating the fragments of step (iv) to a tag capable of binding to a cognate interactant to obtain a pool of tagged catcher oligonucleotides; (vi) providing a pool of starting oligonucleotides for preparing a pool of synthetic single-guide RNAs (sgRNAs) for an sgRNA-guided nucleic acid-binding protein, wherein the starting oligonucleotides comprise a promoter segment, a random segment that is a potential complementary sequence to the catcher oligonucleotide, and a binding segment that is complementary to at least a portion of a scaffold sequence for interacting with the sgRNA-guided nucleic acid-binding protein; (vii) hybridizing the pool of starting oligonucleotides with the tagged catcher oligonucleotides; (viii) removing complexes of starting oligonucleotides and tagged catcher oligonucleotides from the pool of starting oligonucleotides by binding the tag to a cognate interactant (preferably located on magnetic beads or a suitable surface), thereby obtaining a reduced pool of starting oligonucleotides; (ix) preparing a pool of sgRNAs with the reduced pool of starting oligonucleotides obtained in step (viii); (x) using the pool of sgRNAs obtained in step (ix) to cleave a mixture of polynucleotides obtained from a test sample by an sgRNA-guided nucleic acid-binding protein; and (xi) size-selecting one or more uncleaved target polynucleotides from the mixture of polynucleotides obtained in step (x). By providing a pool of sgRNAs capable of binding to the sequences, the method advantageously allows for the removal of all sequences that are not relevant to the personalized target polynucleotide target, i.e., a panel of genes of the patient transcriptome. Thereby, the complexity of the resulting personalized sequencing library is greatly reduced, and sequencing operations are performed on the library, specifically next-generation sequencing (NGS), including a much lower sequencing depth, which significantly reduces the sequencing cost as well as the costs of subsequent data management and data processing.
[0011] In another aspect, the present invention relates to a method for obtaining a pool of personalized target-agnostic synthetic single-guide RNAs (sgRNAs) for an sgRNA-guided nucleic acid-binding protein, the method comprising: (i) purifying a population of mRNA molecules from a sample obtained from a subject; (ii) preparing cDNA from the mRNA molecules of step (i); (iii) amplifying one or more target sequences from the cDNA obtained in step (ii) to obtain a pool of DNA molecules; (iv) fragmenting the amplified DNA molecules, preferably to a size of 20 to 30 bp; (v) ligating the fragments of step (iv) to a tag capable of binding to a cognate interactant to obtain a pool of tagged catcher oligonucleotides; (vi) providing a pool of starting oligonucleotides for preparing a pool of synthetic single-guide RNAs (sgRNAs) for an sgRNA-guided nucleic acid-binding protein, wherein the starting oligonucleotides comprise a promoter segment, a random segment that is a potential complementary sequence to the catcher oligonucleotide, and a binding segment that is complementary to at least a portion of a scaffold sequence for interacting with the sgRNA-guided nucleic acid-binding protein; (vii) hybridizing the pool of starting oligonucleotides with the tagged catcher oligonucleotides; (viii) removing complexes of starting oligonucleotides and tagged catcher oligonucleotides from the pool of starting oligonucleotides by binding the tag to a cognate interactant (preferably located on magnetic beads or a suitable surface), thereby obtaining a reduced pool of starting oligonucleotides; and (ix) preparing a pool of sgRNAs with the reduced pool of starting oligonucleotides obtained in step (viii).
[0012] In a preferred embodiment of the present invention, the amplification (iii) is carried out as a polymerase chain reaction (PCR).
[0013] In a particularly preferred embodiment, the tag capable of binding to a cognate interactant is biotin and wherein the cognate interactant is streptavidin.
[0014] In other embodiments, the step of ligating the fragment to a biotin tag comprises tailing with activated biotin, a ligation reaction with biotin, or ligation to biotin by click chemistry.
[0015] In another embodiment, the sgRNA-guided nucleic acid-binding protein is a DNA-binding Cas protein, preferably Cas9 protein or a derivative thereof.
[0016] In other embodiments of the present invention, the random segment comprises about 10 to 30 random nucleotides, preferably, the random segment comprises about 20 random nucleotides.
[0017] In another embodiment of the present invention, steps (vii) to (viii) as mentioned above are repeated 1, 2, 3, 4, 5 times or more.
[0018] In another embodiment of the present invention, the one or more target polynucleotides or target sequences represent a gene, one or more exons of a gene and / or its open reading frame or sub-part; or a group of different genes, a group of one or more exons of different genes and / or a group of its open reading frame or sub-part, or any combination of any of the previously mentioned elements.
[0019] In another preferred embodiment of the present invention, the method for obtaining an enriched personalized population of target polynucleotides as described above further comprises, as step (xii), the step of sequencing the size-selected uncut target polynucleotides.
[0020] In another aspect, the present invention relates to a kit comprising a pool of sgRNAs obtainable by a method for obtaining a personalized pool of target-agnostic synthetic single-guide RNAs (sgRNAs) for an sgRNA-guided nucleic acid-binding protein as described above, and an sgRNA-guided nucleic acid-binding protein. Preferably, the sgRNA-guided nucleic acid-binding protein is a Cas9 protein or a derivative thereof.
[0021] In another aspect, the present invention relates to the use of a pool of sgRNAs obtainable by a method for obtaining a personalized pool of target-agnostic synthetic single-guide RNAs (sgRNAs) for an sgRNA-guided nucleic acid-binding protein as described above for removing target-agnostic polynucleotides from a mixture of polynucleotides in a Cas9-based endonuclease assay.
[0022] In another aspect, the present invention relates to a method for monitoring a disease condition, which comprises performing at predetermined time intervals the method for obtaining an enriched personalized population of target polynucleotides as described above. Preferably, the method is performed according to the treatment requirements of the disease.
[0023] It should be understood that the above features and those to be explained below can be used not only in the various combinations specified, but also in other combinations or alone without departing from the scope of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 A schematic diagram showing the steps for obtaining a personalized pool of target-agnostic synthetic single-guide RNAs (sgRNAs) according to an embodiment of the present invention is shown.
[0025] Figure 2Shows the steps for preparing a targeting-irrelevant sgRNA used in accordance with an embodiment of the present invention. Detailed Embodiments
[0026] Although the present invention will be described with respect to specific embodiments, such description should not be construed in a limiting sense.
[0027] Before describing in detail the exemplary embodiments of the present invention, definitions important for the understanding of the present invention are provided.
[0028] Unless the context clearly dictates otherwise, as used in this specification and in the appended claims, the singular form "a" also includes the respective plural forms.
[0029] In the context of the present invention, the terms "about" and "approximately" denote an interval of accuracy that a person skilled in the art will understand still ensures the technical effect of the feature being discussed. The term generally represents a deviation from the indicated numerical value of ±20%, preferably ±15%, more preferably ±10%, and even more preferably ±5%.
[0030] It should be understood that the term "comprising" is not restrictive. For the purposes of the present invention, the terms "consisting of" or "consisting essentially of" are considered to be preferred embodiments of the term "comprising". If a group is defined hereinafter as comprising at least certain embodiments, this means that a group consisting preferably only of these embodiments is also covered.
[0031] Furthermore, the terms "(i)", "(ii)", "(iii)", or "(a)", "(b)", "(c)", "(d)", or "first", "second", "third", etc. in a description or claim are used to distinguish similar elements and are not necessarily used to describe an order or time sequence.
[0032] It should be understood that, where appropriate, the terms used are interchangeable and that the embodiments of the present invention described herein can be operated in a sequence different from that described or shown herein. Unless otherwise stated, where the terms relate to a method, procedure or use step, there is no time or time interval coherence between the steps, i.e., the steps can be performed simultaneously or there can be a time interval of seconds, minutes, hours, days, weeks, etc. between these steps.
[0033] It should be understood that the present invention is not limited to the specific methods, procedures, etc. described herein, as these can vary. It should also be understood that the terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the scope of the present invention, which will be limited only by the appended claims.
[0034] The drawings are considered to be schematic, and the elements shown in the drawings need not be shown to scale. However, a variety of elements are shown so that their function and general purpose will be apparent to those skilled in the art.
[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.
[0036] As described above, in one aspect the invention relates to a method for obtaining an enriched population of personalized target polynucleotides, the method comprising: (i) purifying a population of mRNA molecules from a sample obtained from a subject; (ii) preparing cDNA from the mRNA molecules of step (i); (iii) amplifying one or more target sequences from the cDNA obtained in step (ii) to obtain a pool of DNA molecules; (iv) fragmenting the amplified DNA molecules, preferably to a size of 20 to 30 bp; (v) ligating the fragments of step (iv) to a tag capable of binding to a cognate interactant to obtain a pool of tagged catcher oligonucleotides; (vi) providing a pool of starting oligonucleotides for preparing a pool of synthetic single guide RNAs (sgRNAs) for an sgRNA-guided nucleic acid binding protein, wherein the starting oligonucleotides comprise a promoter segment, a random segment that is a potential complementary sequence to the catcher oligonucleotide, and a binding segment that is complementary to at least a portion of a scaffold sequence for interacting with the sgRNA-guided nucleic acid binding protein; (vii) hybridizing the pool of starting oligonucleotides with the tagged catcher oligonucleotides; (viii) removing complexes of the starting oligonucleotides and the tagged catcher oligonucleotides from the pool of starting oligonucleotides by binding the tag to a cognate interactant (preferably located on magnetic beads or a suitable surface), thereby obtaining a reduced pool of starting oligonucleotides; (ix) preparing a pool of sgRNAs with the reduced pool of starting oligonucleotides obtained in step (viii); (x) using the pool of sgRNAs obtained in step (ix) to cleave a mixture of polynucleotides obtained from a test sample by an sgRNA-guided nucleic acid binding protein; and (xi) size selecting one or more uncleaved target polynucleotides from the mixture of polynucleotides obtained in step (x).
[0037] As used herein, the term "target polynucleotide" refers to any target nucleic acid molecule suitable for molecular analysis. Preferably, the target polynucleotide is a DNA or cDNA molecule. In a typical embodiment, the target polynucleotide is derived from a test sample of a subject, or in a specific embodiment, from a group of subjects. As used herein, the term "target sequence" refers to a nucleic acid derived from the transcriptome of a subject, or in a specific embodiment, a group of subjects. Thus, the target sequence can be provided as mRNA, or preferably, as a cDNA molecule. In a specific embodiment of the present invention, the target polynucleotide or the target sequence represents a gene, one or more exons of a gene, and / or its open reading frame or subpart. In other embodiments, the target polynucleotide or the target sequence can also be a group of different genes, a group of one or more exons of different genes, and / or a group of its open reading frame or subparts. Combinations of the previously mentioned elements are also preferred. The form and content of the target polynucleotide are typically reflected by the content and sequence of the target sequence derived from the subject's transcriptome, and in turn, the target sequence is reflected by the sequence and form of the catcher oligonucleotide used in the method of the present invention.
[0038] In the first step of the method of the present invention, a population of mRNA molecules is purified from a sample obtained from a subject.
[0039] As used herein, the term "sample" or "test sample" refers to any biological material obtained from a subject by a suitable method known to those skilled in the art. Samples used in the context of the present invention should preferably be collected in a clinically acceptable manner, more preferably in a manner that preserves nucleic acids, specifically RNA, and more preferably mRNA molecules. The biological sample may include body tissues and / or fluids such as blood, or blood components such as serum or plasma, sweat, sputum or saliva, semen and urine, and excretory or fecal samples. Additionally, the biological sample may contain cell extracts derived from a cell population, the cell population including epithelial cells, preferably tumor epithelial cells or epithelial cells derived from a tissue suspected of being a tumor. Cancerous tissues or samples containing cancer cells are particularly preferred. Other disease-related tissue samples or other biological samples are also preferred. Liquid biopsy samples are particularly preferred. As an alternative, the biological sample may be of animal origin. In certain embodiments, cells may be used as the original source of polynucleotides. In certain embodiments, samples, specifically samples after initial processing, may be mixed. The present invention preferably contemplates the use of non-mixed samples. In specific embodiments of the present invention, specific pre-enrichment steps may also be performed on the content of the biological sample. For example, the sample may be contacted with (e.g.) magnetic particles functionalized with ligands specific for the cell membranes or organelles of certain cell types. Subsequently, the material concentrated by the magnetic particles may be used to prepare polynucleotides. In other embodiments of the present invention, biopsy or excised samples may be obtained and / or used. These samples may contain cells or cell lysates. Additionally, cells, such as tumor cells, may be enriched by filtration of fluid or liquid samples, such as blood, urine, sweat, etc. These filtration treatments may also be combined with pre-enrichment steps based on ligand-specific interactions as described above.
[0040] As used herein, the term "purification" refers to the preparation and purification of nucleic acids, preferably in a manner that prevents degradation of nucleic acids such as DNA or RNA. As used herein, the term "purification" refers to the preparation and purification of nucleic acids, preferably in a manner that prevents degradation of nucleic acids such as DNA or RNA. In other steps, if necessary, cells may be purified from body tissues and body fluids and then further processed to obtain polynucleotides. Purification may be performed according to any suitable method known to those skilled in the art. These methods may include the use of nucleic acid extraction and washing protocols, the use of column-based purification protocols, or preferably, the use of bead-based protocols. The protocols may start with a pretreated sample or a crude sample. In a preferred embodiment, purification is suitable for the preparation of RNA, specifically mRNA. As an alternative, purification may be suitable for the preparation of DNA.
[0041] Thus, a "polynucleotide obtained from a test sample" as used herein is generally a polynucleotide that has been prepared from a test sample, e.g., a DNA and / or RNA molecule, preferably a DNA molecule. Additionally, in certain embodiments, the polynucleotide can also be purified, e.g., purified according to the above-described protocols.
[0042] In other steps of the method of the present invention, cDNA is prepared from previously obtained and purified mRNA molecules. The preparation of cDNA is carried out according to any suitable method or protocol known to those skilled in the art. Generally, a reverse transcription method using poly-T and random oligonucleotides can be used. For example, a poly-T oligonucleotide can be used, which is complementary to the 3'-polyA tail of the mRNA molecule. The poly-T oligonucleotide can have any suitable length, e.g., 6 to 15 nucleotides.
[0043] As a counter-oligonucleotide of the 3' region of the mRNA molecule, for example, a random oligonucleotide can be used. These random oligonucleotides can include 5 to 15 nucleotides, which have a random base sequence, i.e., no pre-determined sequence. The random base sequence generally covers all sequence possibilities in the covered extension, including single nucleotide extensions such as poly-T, poly-A, poly-G, poly-C. Thus, the oligonucleotides are used as a set of different polynucleotides to cover all possibilities. The number of different oligonucleotides used for the annealing step depends on the size of the sequence covered by the longer sequence, and longer sequences require more different oligonucleotides to cover all possible sequence variations compared to shorter sequences. Preferably, the random oligonucleotides contain 6, 7, 8, 9, or 10 nucleotides, i.e., the random oligonucleotides are hexamers, heptamers, octamers, nonamers, or decamers. Hexamers (6 nucleotides) are particularly preferably used.
[0044] In other specific embodiments, the oligonucleotides used for the preparation of cDNA can be sequence-specific, or a combination of a poly-T oligonucleotide and a sequence-specific, i.e., non-random, oligonucleotide can be used. The use of sequence-specific oligonucleotides enables the complexity of the obtained cDNA set to be reduced to a pre-determined set. The sequence-specific oligonucleotides can advantageously bind to a specific gene or a group of genes, preferably a gene or a part thereof as defined herein, or a group of genes, or any target sequence as defined herein, i.e., for copying.
[0045] In a specific embodiment, a set of specific oligonucleotides can be used, wherein the set covers a gene, an exon, its open reading frame or a sub - part thereof in a continuous manner, such that copying is enabled. The term "continuous" means that when used for copying or amplification, the specific oligonucleotides cover or amplify the entire sequence of the gene, exon, its open reading frame or sub - part thereof in the form of the obtained product without overlap, or in other embodiments, have an overlap of 2, 3, 4, 5, 6, 7, 8, 9, 10, 15 or 20 or more nucleotides between each splice.
[0046] Reverse transcription can be carried out using any suitable reverse transcriptase known to those skilled in the art. Examples of such suitable reverse transcriptases are reverse transcriptases that do not contain RNase H activity. Specific examples include MMLV reverse transcriptase without RNase H activity or commercially available reverse transcriptases such as SuperScript, SuperScript II, SuperScript III, StrataScript, etc.
[0047] Preferably, the reverse transcriptase reaction is carried out using a suitable buffer or in the presence of a suitable buffer. Such a buffer can (for example) contain TrisHCL, for example, at a concentration of 250 mM, KCl, for example, at a concentration of 375 mM, MgCl2, for example, at a concentration of 15 mM and DTT, for example, at a concentration of 0.1 M, preferably at pH 8.3. Additionally, a suitable amount of dNTPs, for example, dATP, dCTP, dGTP and dTTP, for example, are used at a suitable concentration, such as 10 mM.
[0048] After the preparation of cDNA, the cDNA molecules are amplified. The amplification can preferably be carried out as a polymerase chain reaction (PCR). Suitable methods and protocols will be known to the person skilled in the art. Examples of PCR methods or PCR-based methods suitable for nucleic acid amplification contemplated by the present invention include basic PCR, hot start PCR, long PCR, quantitative end point PCR, rapid amplified polymorphic DNA analysis, rapid amplification of cDNA ends, differential display PCR, and high-fidelity PCR. The amplification can preferably be carried out using a suitable polymerase, for example, Pfu DNA polymerase or other thermostable DNA polymerases with high fidelity and corresponding systems. Other details will be known to the person skilled in the art and can be sourced from suitable literature sources, such as: PCR, Methods and Protocols, 2017, Lucilia Domingues, Springer, New York, McClelland and Welsh, 1994, PCR Methods Appl., 4:S59-65. Alternative amplification methods also contemplated in the context of the present invention include NASBA, loop-mediated isothermal amplification using Bst-DNA polymerase, isothermal amplification using phi29 DNA polymerase, recombinase-polymerase amplification, and helicase-dependent isothermal amplification.
[0049] Typically, one or more target sequences as defined herein, for example, a specific gene or a group of genes, etc., preferably, a gene or a part thereof or a group of genes as defined herein, or those containing one or more of these target sequences, are amplified according to the present invention. Thus, the target sequences can be amplified by sequence-specific oligonucleotides, for example, sequence-specific oligonucleotides complementary to regions or segments of the target sequences. The amplification yields a pool of double-stranded DNA molecules. The number of molecules obtained and the complexity of the pool can be adjusted by the parameters of the PCR method, such as the number of PCR cycles, the hybridization temperature for oligonucleotide annealing, the buffer composition, the length of the oligonucleotides, the temperature profile and variations during PCR, etc.
[0050] In other specific embodiments of the present invention, the amplified target cDNA molecules can be separated, for example, according to length, sequence, association with genes, gene groups, gene families, etc. The amplified cDNA molecules can be used in the form of a pool as described herein or in a separated form.
[0051] In a subsequent step of the method for obtaining an enriched personalized population of target polynucleotides according to the invention, the amplified DNA molecules are fragmented. Fragmentation can be carried out according to any suitable method known to the person skilled in the art. For example, fragmentation can be achieved by restriction digestion or any suitable shearing protocol, such as, for example, adaptive focused acoustic shearing (AFA) or Covaris shearing, the use of hydrodynamic forces, sonication, point-sink shearing or using a French press shearing procedure. Acoustic shearing is preferably used, specifically Covaris shearing. Also preferably, the size of the resulting polynucleotides is similar to or within a predetermined range. The range of the resulting fragments can be between about 20 and 100 bp, more preferably between about 20 and 50 bp. In a particularly preferred embodiment, the resulting size of the fragmented DNA molecules is between about 20 and 30 bp.
[0052] The fragmented DNA molecules can be obtained as blunt-ended or sticky-ended fragments. Fragments with sticky ends should be provided, and they need to undergo a blunting reaction to prepare for the subsequent step of ligation to the tags. The blunting activity can be, for example, a terminal repair step. The terminal repair step can be carried out by using any suitable terminal repair enzyme activity, such as, for example, DNA polymerase I, preferably its Klenow fragment, T4 DNA polymerase or T4 polynucleotide kinase. Preferably, terminal repair is carried out at 20 °C using T4 DNA polymerase, T4PNK and Klenow.
[0053] In subsequent steps, as mentioned above, the fragments of step (iv), preferably further modified by end repair if necessary, are ligated to a tag capable of binding to a cognate interactant. This step yields a pool of tagged catcher oligonucleotides. As used herein, the term "tag" refers to a molecule capable of binding to a cognate interactant and thereby picking out from a liquid solution or mixture of molecules. Examples of suitable tags and interactants are biotin tags, e.g., DNA fragments obtained by the steps described above to obtain "catcher oligonucleotides", and streptavidin interactants provided on a suitable surface such as in a reaction vessel or on magnetic beads, etc. Derivatives of biotin and streptavidin known to the person skilled in the art are also contemplated. Other examples include magnetic beads bound to DNA fragments obtained by the steps described above, and a magnetic separator that attracts said magnetic beads, e.g., magnetic beads attached to a surface. Embodiments are also contemplated in which the DNA fragments obtained by the steps described above are immobilized on a surface or solid phase. Such immobilization can be, for example, in the form of a column. The solid phase can have any suitable form or structure, e.g., consisting of agarose. This enables potential hybridizing molecules to pass through or move through the surface or solid phase or on the surface or solid phase and thereby bind them to the catcher oligonucleotides and thus remove them from the pool. In certain embodiments, this activity can be repeated one or more times. Thus, interactant-tag binding can advantageously be used to pick out the catcher oligonucleotides together with any bound or linked binding partner containing a complementary target sequence from a mixture of a liquid solution.
[0054] Typically, the catcher oligonucleotides thus correspond to a part of a polynucleotide that should be enriched or should be included in an enriched personalized population of target polynucleotides. For example, the catcher oligonucleotides are complementary to a target polynucleotide as defined herein or a polynucleotide containing a target sequence, e.g., covering a gene, one or more exons of a gene and / or its open reading frame or subpart, or in other embodiments, covering a group of different genes, a group of one or more exons of different genes and / or a group of their open reading frames or subparts, as well as combinations of the elements mentioned previously. Preferably, a group of catcher oligonucleotides covers a group of genes, etc. Specificity, subject-dependence and thus personalization of the gene or group of genes, etc. are achieved by using RNA molecules from a subject sample.
[0055] As used herein, the term "ligated" refers to any suitable step of joining a DNA fragment to a tag as described above. Ligation of the tag is performed on both strands of the double-stranded amplified cDNA fragment. Preferably, ligation can be performed at the 3' end of each strand. Such ligation procedures can (for example) include tailing with an activated tag, such as biotinylation. Preferably, blunt-ended cDNA fragments are tailed with A, where in a preferred embodiment, the A-tailing step involves the use of biotinylated-dATP. A-tailing activity can be carried out by any suitable A-tailing enzyme activity, such as Taq polymerase or Klenow fragment. Preferably, Taq DNA polymerase is used for A-tailing at 65°C. Further details can be sourced from suitable literature sources, such as Nucleic Acids Research, 2010, 38, 13, e137.
[0056] Alternatively, the ligation reaction can be based on a ligation reaction with a tag, such as biotin, or a ligation with a tag, such as biotin, by click chemistry.
[0057] The ligation can be chemical or enzymatic. Enzymatic ligation is preferred. Chemical ligation typically requires the presence of a condensing agent. Examples of chemical ligation contemplated by the present invention utilize electrophilic phosphorothioate groups. Other examples include the use of cyanogen bromide as a condensing agent. Enzymatic ligation can be performed using any suitable ligase known to those skilled in the art. Examples of suitable ligases include T4 DNA ligase, Escherichia coli (E. coli) DNA ligase, T3 DNA ligase, and T7 DNA ligase. Alternatively, ligases such as Taq DNA ligase, Tma DNA ligase, 9°N DNA ligase, T4 polymerase 1, T4 polymerase 2, or thermostable 5'App DNA / RNA ligase can be used.
[0058] Ligation of the tag to the cDNA fragment can be performed by click chemistry. The term "click chemistry" refers to the reaction between an azide and an alkyne to obtain a covalently 1,5-disubstituted 1,2,3-triazole product and is essentially based on Cu catalysis. Generally, the catalyst can be introduced as a Cu-TBTA complex. Preferably, the ligation between the cDNA fragment and the tag, such as biotin, is performed by introducing an alkynyl residue into the DNA molecule. This is typically performed using alkynyl triphosphate by a termination reaction or a nick translation reaction. Alternatively, an alkynyl group can also be introduced during PCR. This reaction yields alkyne-modified DNA, which can react with a homologous azide-activated ester, preferably an azide-activated tag, more preferably an azide-activated biotin.
[0059] Subsequently, the tagged catcher oligonucleotide is provided as a single-stranded molecule. Thus, the double-stranded molecule in which both strands contain the tag is denatured to obtain single strands. In a specific embodiment, the single-stranded nature of the catcher oligonucleotide is maintained or improved by adding a suitable buffer or using suitable reaction parameters.
[0060] In other steps of the method for obtaining an enriched, personalized population of target polynucleotides according to the invention, a pool of starting oligonucleotides for preparing synthetic single guide RNAs (sgRNAs) for an sgRNA-guided nucleic acid-binding protein is provided.
[0061] Typically, the corresponding part of the method is based on the use of the CRISPR / Cas system. As used herein, the term "CRISPR / Cas system" refers to a biochemical method for specifically cleaving and modifying nucleic acids, also known as genome editing. For example, genes in the genome can generally be inserted, removed, or turned off through the CRISPR / Cas system, and nucleotides in genes or nucleic acid molecules can also be altered. The concept and the effect of the steps of the CRISPR / Cas system have multiple similarities with RNA interference, because short RNA fragments of about 18 to 20 nucleotides mediate the binding to the target in the bacterial defense mechanism. In the CRISPR / Cas system, generally an RNA-guided nucleic acid-binding protein, such as a Cas protein, binds to certain RNA sequences as a ribonucleoprotein. For example, a Cas endonuclease (e.g., Cas9, Cas5, Csn1, or Csx12, or derivatives thereof) can bind to certain RNA sequences called crRNA repeats and cleave the DNA adjacent to these sequences. Without wishing to be bound by theory, it is believed that the crRNA repeats form an RNA secondary structure and are subsequently bound by a nucleic acid-binding protein (e.g., Cas) that changes its protein folding, enabling the target DNA to bind to the RNA. In addition, the presence of a PAM motif, i.e., a protospacer adjacent motif, in the target DNA is necessary for activating the nucleic acid-binding protein (e.g., Cas). The DNA is generally cleaved at 3 nucleotides before the PAM motif. The sequence binding to the target DNA, i.e., the crRNA spacer, usually follows the crRNA repeats; generally, the two sequences, i.e., the crRNA repeat motif and the target-binding segment, are labeled as "crRNA". This second part (the target-binding segment) of the crRNA is the crRNA-spacer sequence with a variable adaptor function. It is complementary to the target DNA and binds to the target DNA. Another RNA, tracrRNA or trans-acting CRISPR RNA, is also required. The tracrRNA is partially complementary to the crRNA part, so that they bind to each other. The tracrRNA generally binds to the precursor crRNA, forms an RNA double helix, and is converted into an active form by RNase III. These properties enable binding to the DNA and cleavage by the endonuclease function of a nucleic acid-binding protein (e.g., Cas) near the binding site.
[0062] As used herein, the term "initiating oligonucleotide" refers to a short nucleic acid molecule or nucleic acid oligomer. Its length can vary depending on the specific application, targeting method, genetic background of the organism involved, etc. Generally, the length of the initiating oligonucleotide is between about 40 and 250 nucleotides, for example, 40, 45, 50, 55, 60, 65, 100, 150, 200 or 250 nucleotides or any value between the recited values. Advantageously, the length of the oligonucleotide is 55 nucleotides. Preferably, the initiating oligonucleotide is a single-stranded DNA molecule. In a specific alternative embodiment, RNA, PNA, CNA, HNA, LNA or ANA molecules or mixtures thereof are also contemplated as initiating oligonucleotides. As used herein, the term "PNA" refers to peptide nucleic acid, i.e., an artificially synthesized polymer similar to DNA or RNA. The PNA backbone generally consists of repeating N-(2-aminoethyl)-glycine units linked by peptide bonds. A variety of purine and pyrimidine bases are linked to the backbone via methylene carbonyl bonds. As used herein, the term "CNA" refers to cyclopentane nucleic acid, i.e., a nucleic acid molecule containing, for example, 2'-deoxycarbovir. The term "HNA" refers to hexitol nucleic acid, i.e., a DNA analogue constructed from standard nucleobases and a phosphorylated 1,5-anhydrohexitol backbone. As used herein, the term "LNA" refers to locked nucleic acid. Generally, locked nucleic acid is modified and thus is an inaccessible RNA nucleotide. The ribose moiety of an LNA nucleotide can be modified with an additional bridge connecting the 2' and 4' carbons. Such a bridge locks the ribose in the 3'-endo conformational conformation. The locked ribose conformation enhances base stacking and backbone pre-organization. Presumably, this increases the thermal stability, i.e., the melting temperature of the oligonucleotide. As used herein, the term "ANA" refers to arabinonucleic acid or a derivative thereof. In the context of the present invention, a preferred ANA derivative is 2'-deoxy-2'-fluoro-β-D-arabinonucleoside (2'F-ANA).
[0063] The oligonucleotide is typically provided in a liquid, for example, an aqueous solution. The solution may contain or consist of a suitable buffer, such as a hybridization buffer, for example, which contains SSC, NaCl, sodium phosphate, SDS, TE and / or MgCl2.
[0064] As used herein, "pool of starting oligonucleotides for preparing synthetic single guide RNA (sgRNA)" refers to a group of oligonucleotides that provide the features necessary for a pool for preparing synthetic single guide RNA (sgRNA). In this context, as used herein, the term "synthetic single guide RNA (sgRNA)" or "single guide RNA (sgRNA)" refers to an artificial or synthetic combination of the crRNA and tracrRNA sequences of the CRISPR / Cas system as described above. Generally, the sgRNA contains a sequence segment that can be used to direct a DNA-binding protein to a binding site. As described by Jinek et al., 2012, Science, 337, 816-821, the crRNA and tracrRNA can be combined into a functional entity (sgRNA), which satisfies the two activities (crRNA and tracrRNA) mentioned above. For example, nucleotides 1-42 of crRNA-sp2, nucleotides 1-36 of crRNA-sp2, or nucleotides 1-32 of crRNA-sp2 can be combined with nucleotides 4-89 of tracrRNA. Other options for obtaining sgRNA can be derived from Nowak et al., 2016, Nucleic Acids Research, 44, 20, 9555–9564. For example, sgRNAs can be provided that contain different forms of the upper stem structure, or sgRNAs in which the spacer sequence is differentially truncated from the standard 20 nucleotides to 14 or 15 nucleotides. Further envisioned variants include those in which the putative RNAP III terminator sequence is removed from the lower stem. Also envisioned are variants in which the upper stem is extended to increase sgRNA stability and increase its assembly with a nucleic acid-binding protein for sgRNA guidance, e.g., a Cas protein. According to other embodiments of the invention, depending on the form or properties of the nucleic acid-binding protein for sgRNA guidance, e.g., a Cas protein, the sequence and form of the sgRNA can be different. Thus, different combinations of sequence elements can be used based on the source of the nucleic acid-binding protein for sgRNA guidance. The invention also contemplates any future developments in this context and includes any modifications or improvements to the sgRNA-nucleic acid-binding protein interaction that go beyond the information that can be derived from Jinek et al., 2012 or Nowak et al., 2016. In a specific embodiment, the sgRNA to be used can have a sequence as shown in any one of SEQ ID NOs: 1 to 3. Particular preference is given to using Streptococcus pyogenes sgRNA, e.g., as used in commercially available kits such as the EnGen sgRNA Synthesis Kit provided by NewEngland Biolabs Inc.Similar sgRNA formats from other commercially available suppliers, or separately prepared sgRNAs, are also envisioned. Such sgRNAs can be derived from the sequences shown in SEQ ID NO: 1 if used in conjunction with a homologous nucleic acid binding protein from Streptococcus pyogenes. As an alternative, if used in conjunction with a homologous nucleic acid binding protein from Staphylococcus aureus, such sgRNAs can be derived from the sequences shown in SEQ ID NO: 2. In other alternatives, if used in conjunction with a homologous nucleic acid binding protein from Streptococcus thermophilus, the sgRNA can be derived from the sequences shown in SEQ ID NO: 3.
[0065] Generally, the features required to prepare synthetic single-guide RNA (sgRNA) include all elements necessary to generate an sgRNA molecule suitable for use in the CRISPR / Cas system as described above herein. Thus, these features include the presence of a promoter segment; the presence of a random segment that serves as or contains a potential complementary sequence to a catcher oligonucleotide as described herein, which serves as a complementary sequence to a potential binding or hybridization interaction partner with a matching sequence; and the presence of a binding element that is complementary to at least a portion of a scaffold sequence, which is used to interact with an sgRNA-guided nucleic acid binding protein. The features mentioned can be provided in any suitable order. Preferably, the order is from 5' to 3': (i) the promoter segment, (ii) the random segment, and (iii) the binding element for the scaffold sequence. The members of the oligonucleotide set differ according to the sequence of the random segment.
[0066] As used herein, a "promoter segment" refers to any suitable promoter structure capable of initiating RNA transcription. Preferably, the promoter is a promoter that functions under in vitro conditions. In other embodiments, the promoter can be a constitutive promoter or a regulatable promoter. Examples of suitable promoters are the T7 RNA polymerase promoter. In alternative embodiments, the promoter can be the U6 RNA polymerase III promoter, the type III RNA polymerase III promoter H1, or the cytomegalovirus promoter (CMV), preferably, the minimal CMV promoter. The promoter segment can be accompanied by or further include other elements, such as spacer elements, guiding elements, etc. For example, upstream of the spacer sequence of the promoter, the segment can preferably include 1 or 2 guanine residues. Other details of the promoter segment can be sourced from suitable literature sources, such as Milligan et al., 1987, Nucleic Acids Research, 15, 21, 8783-8798 or Nowak et al., 2016, Nucleic Acids Research, 44, 20, 9555–9564.
[0067] As used herein, a "random segment" refers to a nucleic acid extension containing a random base sequence that typically covers all sequence possibilities in the covered extension, including single nucleotide extensions, such as poly-T, poly-A, poly-G, poly-C. In certain embodiments, most or a certain amount of possible nucleotide combinations or sequences can be represented by the random segment, for example, 99%, 95%, 90%, 85%, 80%, 75%, 70% or less, or any value between the above values. In a preferred embodiment, the random segment contains about 10 to 30 random nucleotides, for example, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 nucleotides. Particularly preferably, the random segment contains about 20 nucleotides.
[0068] A "pool" of starting oligonucleotides can have any suitable size. Generally, the size of the pool depends on the length of the random segment, so a longer sequence means a larger oligonucleotide pool, and a longer segment requires more different oligonucleotides to cover all possible sequence variations. Thus, the pool of oligonucleotides can include segments with random base sequences, i.e., no pre-determined sequence, and thus contains all possible nucleotide combinations or sequences in the segment.
[0069] In a specific embodiment, some random segments may occur in excess, while others may occur in deficiency. This over-expression or under-expression can be deliberately controlled, for example, according to need, the known or expected sequence composition in the sample, or the addition of separation or elimination techniques used during sample preparation, such as size exclusion, etc. For example, it may be advantageous to provide a higher sequence occurrence for nucleic acid substances that are more frequently present in the test sample, such as rDNA sequences, repetitive sequences, etc.
[0070] In other embodiments, the pool of oligonucleotides may contain one or several occurrences of a single specific nucleotide combination or sequence. For example, each possible nucleotide combination or sequence in the random segment defined above may occur 1, 2, 3, 4, 5, 10, 50, 100, 1000 times or more frequently in the pool of oligonucleotides.
[0071] The term "binding element complementary to at least a portion of the scaffold sequence for interaction with an sgRNA-guided nucleic acid-binding protein" refers to a nucleic acid segment that contains a sequence complementary to an oligonucleotide that combines the functions of crRNA and tracrRNA as in an sgRNA, preferably as described above, and contains a crRNA repeat motif and an RNA double helix-forming element. The term "scaffold sequence" as used herein refers to the structural motif that is generally required for binding and interaction with an sgRNA-guided nucleic acid-binding protein as defined above, for example, a Cas protein. By providing the scaffold function in a separate oligonucleotide, a hybridization step between the pool of starting oligonucleotides (preferably, the reduced pool of starting oligonucleotides according to the present invention) and the scaffold sequence, followed by DNA extension, results in the generation of a double-stranded template molecule in one entity that contains a promoter segment, a random segment that is potentially complementary to a catcher oligonucleotide, and the sgRNA scaffold function. Figure 2 Schematically shows a method for obtaining the double-stranded template molecule that contains a promoter segment, a random segment that is potentially complementary to a catcher oligonucleotide, and the sgRNA scaffold function in one entity, as well as subsequent steps for generating an sgRNA molecule based on the template molecule.
[0072] As used herein, the term "complementary" refers to the presence of matching base pairs in opposing nucleic acid strands. For example, for the nucleotide or base A in the sense strand, the complementary or antisense strand binds through the nucleotide or base T, or vice versa; similarly, for the nucleotide or base G in the sense strand, the complementary or antisense strand binds through the nucleotide or base C, or vice versa. In certain embodiments of the present invention, this scheme of full or complete complementarity can be altered by the possibility of the presence of single or multiple non-complementary bases or nucleotide extensions within the sense and / or antisense strands. Thus, within the concept of a pair of sense and antisense strands, the two strands can be fully complementary or can be only partially complementary, e.g., showing about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% or 100% complementarity between all nucleotides of the two strands or between all nucleotides in a specific segment as defined herein. Non-complementary bases can include one of the nucleotides A, T, G, C, i.e., showing a mismatch between A and G or between T and C, or can include any modified nucleobase, which includes (e.g.) modified bases as described in WIPO Standard ST.25. In addition, the present invention also contemplates complementarity between different nucleic acid molecules, e.g., between a DNA strand and an RNA strand, between a DNA strand and a PNA strand, between a DNA strand and a CNA strand, etc. Preferably, the complementarity between the strands or segments as defined herein is complete or 100% complementarity.
[0073] As used herein, the term "complementary to at least a portion of the scaffold sequence" refers to the binding segment having a complementary overlap with the oligonucleotide containing the scaffold sequence. The overlap can be (e.g.) 5, 7, 10, 12, 15, 18, 20, 22, 25, 28 or 30 nucleotides in length, or an overlap of any value between the above values. Longer overlaps are also contemplated. Short overlaps in the range of 5 to 20 nucleotides are preferred. The length of the overlap can be further adjusted according to hybridization efficiency. The overlap is typically at the 3' end of the starting oligonucleotide and the 5' end of the oligonucleotide containing the scaffold sequence. Within the overlap, the match or complementarity between complementary bases is preferably 100%. In alternative embodiments, the match is less than 100%, e.g., 99%, 95%, 90%, 85% or less than 85%.
[0074] The term "potential complementary sequence" as used in the context of the random segments refers to the sequence of the random segments which, based on specific probabilities related to size, nucleotide combinations, etc., can be complementary to one or more catcher oligonucleotides as defined above herein and which are derived from the cDNA / mRNA molecules of a subject. Thus, the random segments can contain sequences that are complementary to the catcher oligonucleotides and sequences that are not complementary to the catcher oligonucleotides. The gist of the present invention is to identify or obtain those starting oligonucleotides for the preparation of synthetic single guide RNA (sgRNA) that are indeed complementary to the personalized catcher oligonucleotides, thereby enabling the regulatable removal of these sequences from a pool of starting oligonucleotides.
[0075] In the case of at least partial complementarity between the catcher oligonucleotide and the starting oligonucleotide, hybridization is preferably carried out between the catcher oligonucleotide and at least a portion of the potential complementary sequence provided in the starting oligonucleotide as defined herein. The binding between these molecules enables the effective removal of the starting oligonucleotides containing the complementary sequence from the pool of starting oligonucleotides.
[0076] Preferably, the catcher oligonucleotide contains about 20 to 30 nucleotides, for example, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 nucleotides. The catcher oligonucleotide can also contain other elements, such as spacer elements, barcode sequences, etc. These elements can be added to the catcher oligonucleotide before or after ligation to the tags as described above.
[0077] In a subsequent step, the pool of starting oligonucleotides and the catcher oligonucleotide are hybridized. The hybridization is typically carried out in a liquid solution, for example, an aqueous solution containing a suitable buffer as defined above. The hybridization can be carried out according to any suitable temperature, ionic concentration, and / or pH parameters known to the person skilled in the art. For example, the hybridization can be carried out at a temperature and / or pH and / or ionic concentration at which most, preferably all, of the complementary bases in the starting oligonucleotides undergo complementary base pairing with the catcher oligonucleotide in solution. For example, non-specific binding or hybridization reactions can be avoided by setting the temperature to a value that allows only complete, i.e., 100% complementary binding. In an alternative embodiment, the temperature can be set to a value that allows about 99%, 98%, 95%, 90%, 85% or 80% complementary base complementary binding.
[0078] In other steps of the method for obtaining an enriched population of target polynucleotides according to the invention, complexes of the starting oligonucleotides and the catcher oligonucleotides are removed from the pool of starting oligonucleotides. Preferably, removal is initiated by binding the tag to a cognate interactant, which is preferably located on magnetic beads or a suitable surface. For example, by introducing magnetic beads containing streptavidin molecules, the biotin tag can be bound to the tag. In addition, the hybridized starting oligonucleotides, i.e., those oligonucleotides containing sequences that match the catcher oligonucleotide sequence, are also indirectly bound to the magnetic beads. Subsequently, the magnetic beads with the bound nucleic acids can be removed from the solution. For example, the magnetic beads can be magnetic beads that can be removed by magnetic force. Different removal options are also contemplated, such as centrifugation or filtration. As an alternative, removal can be effected by magnetic bead-magnetic force interaction. In this case, the catcher oligonucleotide is linked to the magnetic bead. After hybridization with the matching starting oligonucleotide, a magnetic force is applied and the complex between the catcher oligonucleotide and the starting oligonucleotide can be removed from the solution to, for example, a magnetic zone.
[0079] After performing the above steps, the pool of starting oligonucleotides is reduced, i.e., starting oligonucleotides that are complementary to the sequence of the catcher oligonucleotide are no longer present in the pool or their presence is reduced. As used herein, the term "reduced presence" means that the number of starting oligonucleotides that are complementary to the sequence of the catcher oligonucleotide is reduced by 5-fold, 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 100-fold.
[0080] In a specific embodiment, the previous steps (vii) and (viii), namely the step of hybridizing the starting oligonucleotide pool with the tagged catcher oligonucleotides and the step of removing the complex of the starting oligonucleotides and the tagged catcher oligonucleotides as defined above, can be repeated one or several times. For example, these steps can be repeated 1, 2, 3, 4, 5 or more times. In certain embodiments, the repetition can be associated with the amplification step (e.g., by PCR) of the reduced starting oligonucleotide pool. Thus, suitable primer binding sites can be present on each starting oligonucleotide and used for amplification. Pfu DNA polymerase or other thermostable DNA polymerases with high fidelity and corresponding systems can be preferably used for amplification. Each repetition can be combined with a washing step and / or a quality control step. For example, the presence of specific elements in the reduced pool can be determined by real-time PCR. In a specific embodiment, a new catcher oligonucleotide or a new set of catcher oligonucleotides is used for each repetition. In other embodiments, different catcher oligonucleotides or different sets of catcher oligonucleotides are used. For example, one or more parts of the cDNA molecules as described above herein are used for each repetition. For example, the difference can be different parts of the entire target sequence, such as adjacent sequence parts of a gene or genomic sequence (if the gene is covered by several adjacent or consecutive or partially overlapping catcher oligonucleotides).
[0081] In other steps, a pool of sgRNAs with the reduced starting oligonucleotide pool obtained in step (viii) is generated. Generally, this step can be implemented as Figure 2 shown. Generally, a single-stranded oligonucleotide, preferably a DNA molecule, containing the crRNA and tracrRNA functions combined in the sgRNA, preferably containing the crRNA repeat motif and the RNA double helix-forming element as described above herein, is hybridized via the binding element present as defined above herein to the starting oligonucleotides obtained in step (viii), and the binding element is present in the starting oligonucleotides. Subsequently, the single-stranded part of the hybrid complex can be filled by a DNA extension reaction. This reaction is preferably carried out with a DNA polymerase, such as T4 DNA polymerase or Klenow enzyme.
[0082] In certain embodiments, the resulting double-stranded template molecule containing a promoter segment, a random segment that is unrelated to binding to the catcher oligonucleotide (i.e., does not show complementarity to the catcher oligonucleotide sequence), and the sgRNA scaffold function in one entity can be amplified (e.g., by PCR). Subsequently, the template can be transcribed into an RNA molecule via the promoter segment as defined above herein, thereby generating an sgRNA that can be used for CRISPR / Cas activity.
[0083] In a specific embodiment, a pool of sgRNAs was obtained according to a commercialization protocol and based on a commercial kit (e.g., a commercial kit provided by New England Biolabs, such as the EnGen sgRNA Synthesis Kit). Similar forms of sgRNAs from other commercial suppliers are also contemplated.
[0084] Thus, in another aspect, the present invention relates to a method for obtaining a pool of personalized target-agnostic synthetic single-guide RNAs (sgRNAs) for an sgRNA-guided nucleic acid-binding protein. The method essentially comprises steps (i) to (ix) as defined herein. In an alternative aspect, the present invention contemplates the provision of a pool of reduced starting oligonucleotides, which essentially comprises steps (i) to (iv) as defined above herein.
[0085] In certain embodiments, the sgRNAs obtained in step (ix) can be stored, modified, and / or purified to allow for appropriate further use. For example, any 5' triphosphate residues that may be present can be removed, e.g., by using alkaline phosphatase. Purification of the sgRNAs can be performed according to any suitable protocol known to the person skilled in the art, e.g., using a spin column to remove proteins, salts, and nucleotides. In certain embodiments, a quality check of the sgRNAs is also contemplated before further use, e.g., by UV light absorbance at 260 nm.
[0086] In other steps of the method for obtaining an enriched personalized population of target polynucleotides according to the present invention, a pool of sgRNAs obtained in step (ix) as defined above is used to cleave a mixture of polynucleotides obtained from a test sample as defined herein by an sgRNA-guided nucleic acid-binding protein.
[0087] As used herein, the term "mixture of polynucleotides" refers to nucleic acids derived from a sample as mentioned above. The polynucleotides used in this step of the method according to the present invention are preferably DNA molecules or cDNA molecules. The DNA molecules can be genomic DNA or derivatives thereof. The use of DNA libraries is also contemplated.
[0088] In a preferred embodiment, the mixture of polynucleotides obtained from a test sample, particularly from a liquid biopsy sample, comprises genomic DNA and / or cDNA molecular fragments referred to as cell-free circulating DNA (ccfDNA). The size of these DNA species typically ranges between 70 - 300 base pairs. CcfDNA generally comprises degraded DNA fragments released into the plasma. CcfDNA can be used to describe various forms of DNA freely circulating in the bloodstream, including circulating tumor DNA (ctDNA) and cell-free fetal DNA (cffDNA). In cancer, elevated levels of cfDNA are observed, particularly in advanced disease. There is evidence that cfDNA becomes increasingly frequent in circulation with increasing age. CcfDNA has been shown to be a useful biomarker for a large number of diseases other than cancer and fetal medicine. This includes (but is not limited to) trauma, sepsis, aseptic inflammation, myocardial infarction, stroke, transplantation, diabetes, and sickle cell disease. Most of the ccfDNA is double-stranded extracellular molecules of DNA.
[0089] In a preferred embodiment, the mixture of polynucleotides obtained from a test sample comprises genomic DNA and / or cDNA molecules, optionally with their size adjusted to a predetermined value. For example, the mixture of polynucleotides can comprise sheared or fragmented genomic DNA or cDNA molecules as defined above. In an optional embodiment, the size of the obtained polynucleotides can be adjusted to a predetermined range. Exemplary ranges are from about 2 kb to 2.5 kb, 2.5 kb to 3 kb, or 3 kb to 3.5 kb, etc., 5 kb to 6 kb, 10 kb to 12 kb, etc. The size of the polynucleotides and any optional adjustment can be based on the length of the target sequence as mentioned herein.
[0090] According to the CRSIPR / Cas method as described above, cleavage is performed by a sgRNA-guided nucleic acid-binding protein, e.g., a nuclease such as Cas, using the pool of sgRNAs obtained in step (ix). Generally, the mixture of polynucleotides as defined above is added to a reaction solution comprising the sgRNA as mentioned above at a suitable concentration, a suitable reaction buffer, and a sgRNA-guided nucleic acid-binding protein at a suitable concentration. The reaction can be incubated at a suitable temperature. Subsequently, in certain embodiments, the reaction can be terminated by adding a protease, e.g., proteinase K, or preferably, by performing a heat denaturation step, e.g., at 65 °C.
[0091] In a specific embodiment, the cleavage can also be carried out according to a commercialization protocol and based on a commercialization kit (e.g., a commercialization kit provided by New England Biolabs, such as EnGen Cas9NLS, Streptococcus pyogenes in vitro digestion kit). Similar digestion protocols from other commercial suppliers are also contemplated.
[0092] In the final step of the method for obtaining an enriched personalized population of target polynucleotides of the present invention, size selection is performed, which enables the separation of uncleaved target polynucleotides from the cleaved polynucleotides obtained in step (x). Generally, since random target sequences are used for sgRNA preparation, polynucleotide molecules containing matching random sequences are cleaved with the CRSIPR / Cas system. Polynucleotide molecules containing target sequences (which are not recognized by the sgRNAs in the pool of sgRNAs because these molecules have been removed by hybridization with catcher oligonucleotides as described above), i.e., the target polynucleotides according to the present invention, will not be cleaved and thus have a larger size. Size selection can be carried out by any suitable method. For example, agarose gel- or polyacrylamide gel-based methods or bead-based methods can be used. In a preferred alternative embodiment, magnetic beads can be used to remove short fragments.
[0093] Subsequently, the obtained target polynucleotides can be purified, stored, and / or used for additional or subsequent activities.
[0094] In a specific embodiment of the present invention, another activity carried out using the polynucleotide is sequencing the obtained target polynucleotides. As used herein, the term "sequencing" refers to any suitable sequencing method known to those skilled in the art. Preferably, next-generation sequencing (NGS) or second-generation sequencing techniques can be used, which are generally large-scale parallel sequencing methods carried out in a highly parallel manner. For example, sequencing can be carried out according to a parallel sequencing method on platforms such as Roche 454, GS FLX Titanium, Illumina, Life Technologies Ion Proton, Oxford Nanopore Technologies, Solexa, Solid, or Helicos Biosciences Heliscope systems. In certain embodiments, sequencing can also include additional preparation, sequencing, and subsequent imaging and initial data analysis steps of the polynucleotide.
[0095] For example, the preparation step can include randomly fragmenting the polynucleotide into smaller sizes and generating sequencing templates, such as fragment templates. The spatially separated templates (for example) can be attached or immobilized on a solid surface, thereby allowing simultaneous sequencing reactions. In a typical example, a library of nucleic acid fragments is generated and adapters containing universal primer sites are ligated to the ends of the fragments. Subsequently, the fragments are denatured into single strands and captured by magnetic beads. After amplification, a large number of templates can be attached or immobilized in a polyacrylamide gel, or chemically cross-linked to an amino-coated glass surface, or deposited on individual microtiter plates. As an alternative, solid-phase amplification can be used. In this method, forward and reverse primers are typically attached to a solid support. The surface density of the amplified fragments is defined by the ratio of primers to templates on the support. This method can generate millions of spatially separated template clusters that can hybridize with universal sequencing primers for large-scale parallel sequencing reactions. Other suitable alternatives include multiple displacement amplification methods. Suitable sequencing methods include (but are not limited to) Illumina's cyclic reversible termination (CRT) or sequencing by synthesis (SBS), sequencing by ligation (SBL), single molecule addition (pyrosequencing), or real-time sequencing. An example platform using the CRT method is Illumina / Solexa and HelicoScope. Exemplary SBL platforms include Life / APG / SOLiD carrier oligonucleotide ligation detection. An exemplary pyrosequencing platform is Roche / 454. Exemplary real-time sequencing platforms include the Pacific Biosciences platform and the Life / Visi-Gen platform. Other sequencing methods for obtaining large-scale parallel nucleic acid sequence data include nanopore sequencing, sequencing by hybridization, sequencing based on a nanotransistor array, sequencing based on a scanning tunneling microscope (STM), or sequencing based on a nanowire-molecular sensor. More detailed information about sequencing methods is known to those skilled in the art or can be sourced from suitable literature sources, such as Goodwin et al., 2016, Nature Reviews Genetics, 17, 333-351, van Dijk et al., 2014, Trends in Genetics, 9, 418-426, or Feng et al., 2015, Genomics Proteomics Bioinformatics, 13, 4-16.
[0096] In another aspect, the present invention relates to target polynucleotides obtainable by a method of obtaining an enriched personalized population of target polynucleotides as defined above herein. Thus, the target polynucleotides may be present, for example, in a mixture with non-target polynucleotides in a size-separable state. Accordingly, a purified portion of the target polynucleotides can be obtained by separating the target polynucleotides from the non-target polynucleotides as described herein. Similarly, the target polynucleotides may be present in a separated form in a gel, such as an agarose gel or a polyacrylamide gel, and can thus be extracted therefrom using methods known to those skilled in the art. The obtained target polynucleotides can be purified, stored, or modified according to any suitable method. The target polynucleotides can be provided in any suitable buffer or liquid, or can be provided in a dry or lyophilized form.
[0097] In another aspect, the present invention relates to a pool of synthetic single-guide RNAs (sgRNAs) that are target-irrelevant for an sgRNA-guided nucleic acid-binding protein, such as Cas9, obtainable by the method of the present invention. Thus, the pool can be provided as RNA molecules. Preferably, the RNA is purified and / or clarified. It can be provided in any suitable buffer or liquid, or can be provided in a dry or lyophilized form. In an alternative embodiment, the pool of sgRNAs is provided as a pool of reduced starting oligonucleotides according to the present invention. In another example, it can be provided as a mixture of starting oligonucleotides and oligonucleotides containing a scaffold sequence as described above.
[0098] In another aspect, the present invention relates to a kit that comprises a pool of sgRNAs obtainable by a method of obtaining a personalized pool of target-irrelevant synthetic single-guide RNAs (sgRNAs) for an sgRNA-guided nucleic acid-binding protein and an sgRNA-guided nucleic acid-binding protein. The kit is preferably for enriching a personalized population of target polynucleotides. The features of the method as defined above herein also apply to the kit of the present invention. The kit can, for example, comprise reagents and components as defined in one or more steps of the method of the present invention. For example, the kit can comprise reagents or components for cleaving a mixture of polynucleotides obtained from a test sample by an sgRNA-guided nucleic acid-binding protein. In different embodiments, the kit can comprise or can additionally comprise reagents or components for performing size selection. Generally, the kit can comprise a suitable buffer solution, a marker, or a washing solution, etc. In addition, the kit can comprise a certain amount of known nucleic acid molecules or proteins, which can be used for calibration of the kit or as an internal control. The corresponding components will be known to those skilled in the art.
[0099] In addition, the kit may include a leaflet of instructions and / or may provide information regarding its use, etc.
[0100] An apparatus for carrying out the above method steps is also contemplated. The apparatus may, for example, consist of different modules, which may carry out one or more steps of the method according to the invention. These modules may be combined in any suitable manner, for example, they may be present in a single location or be separate. The implementation of the method at different time points and / or different locations is also contemplated. After some of the steps of the method as defined herein, there may be a break or pause, during which the reagents or products, etc. are appropriately stored, for example, stored in a refrigerator or a cooling device. In the case where these steps are carried out in a specific module of the apparatus as defined herein, the module may serve as a storage vehicle. The module may also be used to transport the reaction products or reagents to different locations, for example, different laboratories, etc.
[0101] In another aspect, the present invention relates to the use of a pool of sgRNAs obtainable by a method of obtaining a pool of synthetic single guide RNAs (sgRNAs) that are target-irrelevant for an sgRNA-guided nucleic acid-binding protein in an sgRNA-guided nucleic acid-binding protein-based assay for removing target-irrelevant polynucleotides from a mixture of polynucleotides. The assay may include the step of cleaving a mixture of polynucleotides obtained from a test sample by an sgRNA-guided nucleic acid-binding protein, wherein the nucleic acid-binding protein is directed to sequences that will be cleaved by the sgRNAs obtained by the method according to the invention. The features of the method as defined above herein also apply to the use or assay as mentioned above.
[0102] In a preferred embodiment of the above method, kit or application, the nucleic acid-binding protein for sgRNA guidance is a DNA-binding Cas protein. Examples of these DNA-binding Cas proteins are Cas2, Cas3, Cas5, Csn1 or Csx12 or Cas9. Derivatives or mutants thereof are also contemplated. In a particularly preferred embodiment, the nucleic acid-binding protein for sgRNA guidance is derived from the Cas9 protein family or a derivative thereof. More preferably, the nucleic acid-binding protein for sgRNA guidance is Cas9 or a derivative thereof. Preferably, the derivative is a functional derivative having nuclease activity. The present invention also contemplates the use of Cas9 from different bacterial sources. For example, the Cas9 protein can be derived from Streptococcus pyogenes, Staphylococcus aureus or Streptococcus thermophiles. Preferably, the Cas9 is Streptococcus pyogenes Cas9 protein. Other detailed information regarding the forms and uses of Cas proteins can be obtained from suitable literature sources such as Jiang and Doudna, 2017, Annu. Rev. Biophys., 46, 505-529, Makarova et al., 2011, Biology Direct, 6, 38 or Wang et al., 2016, Annu. Rev. Biochem., 85, 22.1-22.38.
[0103] In another aspect, the present invention relates to a method of monitoring a disease condition, which comprises performing a method of obtaining an enriched personalized population of target polynucleotides at a predetermined time interval. Thus, the aim of the method is to provide personalized cDNA molecules of a patient for preparing sgRNAs unrelated to the target, and thus to provide a selection of one or more uncut target polynucleotides at specific intervals. Subsequently, the target polynucleotides are analyzed according to their sequence and / or length etc., and they can be compared (for example) with previously obtained target polynucleotides or with reference polynucleotides with respect to their sequence, length or other parameters. The repeated provision of the corresponding cDNA and ultimately the target polynucleotides enables the tracking of the target polynucleotides, for example, the molecular changes of a gene or a group of genes. These molecular changes can indicate the onset or presence of a disease or the absence or cure of a disease, or can have any other suitable diagnostic value. The molecular changes can also enable (for example) the adoption of a treatment method by defining the length of treatment, selecting a suitable medicament, defining potential co-therapies, etc.
[0104] In a preferred embodiment, the intervals between the enriched personalized populations of target polynucleotides that can be used relative to a population of a subject's mRNA molecules are monitored according to any suitable considerations. For example, any suitable interval can be implemented, e.g., between hours to years. Preferably, the interval can be defined according to the requirements of the treatment strategy. For example, if the disease to be treated is a disease that requires therapy readjustment monthly or bi-monthly, the interval can be set accordingly. Thus, subjects can be stratified according to the results of the monitoring method.
[0105] As used herein, the term "stratifying a subject" means classifying a subject by factors other than the treatment itself. In the context of the present invention, such a factor can be a molecular change as defined above herein. Stratification can (for example) help control confounding variables, or assist in detecting and interpreting between variables. Generally, a patient can be analyzed relative to their molecular changes after certain intervals. In the event of encountering or suspecting a particular molecular condition, a particular form of therapy or a specifically adjusted form of therapy can be used.
[0106] As used herein, the term disease can be any disease suitable for molecular analysis. Preferably, the disease is cancer.
[0107] In a specific embodiment, the cancer can be breast cancer, prostate cancer, ovarian cancer, kidney cancer, lung cancer, pancreatic cancer, bladder cancer, uterine cancer, kidney cancer, brain cancer, gastric cancer, colon cancer, melanoma or fibrosarcoma, glioblastoma or hematological leukemia or lymphoma, e.g., spinal or lymphatic lymphoma.
[0108] Turning now to Figure 1, a schematic diagram of steps for obtaining a pool of target - irrelevant synthetic single - guide RNAs (sgRNAs) according to an embodiment of the present invention is provided. In the first step, a sample 1 of a subject is purified 2. This step obtains mRNA 3, which is reverse - transcribed 4, resulting in the production of cDNA copies 5 of the mRNA. The cDNA is further amplified 6 by polymerase chain reaction 7. The amplification obtains a PCR product 8, which is fragmented 9, resulting in the production of short fragments 10, which particularly contain sticky ends. Subsequently, the short fragments 10 are blunt - ended 11 to obtain blunt - ended fragments 12. The blunt - ended fragments 12 are ligated 14 to a tag 13, preferably biotin. This obtains a pool 15 of tagged catcher oligonucleotides. In addition, a pool 16 of starting oligonucleotides is provided, which contains a promoter segment 17, a random segment 18 as a potential complementary sequence to the catcher oligonucleotide, and a binding segment 19, which is complementary to at least a part of the scaffold sequence used for interaction with the sgRNA - guiding nucleic acid - binding protein. The pool 16 of starting oligonucleotides is hybridized 22 with single - stranded catcher oligonucleotides 20 derived from the pool 15 of tagged catcher oligonucleotides, wherein the catcher oligonucleotides contain a tag 13 and a segment 21 complementary to the random segment 18. Subsequently, the complex of the starting oligonucleotides 16 and the catcher oligonucleotides 20 is removed 23 from the solution or reaction mixture. In the solution or reaction mixture, a reduced pool 24 of starting oligonucleotides is retained.
[0109] Figure 2 A schematic diagram of steps for preparing a target - irrelevant sgRNA 107 used according to an embodiment of the present invention is shown. The method starts with a pool 16 of starting oligonucleotides, which contains a promoter segment 17, a random segment 18 as a potential complementary sequence to the catcher oligonucleotide, and a binding segment 19, which is complementary to at least a part of the scaffold sequence used for interaction with the sgRNA - guiding nucleic acid - binding protein, and hybridizes 102 with a scaffold oligonucleotide 100. Subsequent reaction steps 109 occur in a single tube 101. After hybridization, single - stranded regions are filled in a DNA extension reaction 103, providing a double - stranded DNA molecule 104. In the next step, transcription 106 of the dsDNA molecule is initiated 105 by promoter activity. This obtains an sgRNA molecule 107, which contains a sequence 18 that is not specific to the target polynucleotide according to the present invention and an sgRNA scaffold segment 108.
[0110] The accompanying drawings are provided for illustrative purposes. Accordingly, it should be understood that the drawings should not be regarded as restrictive. Those skilled in the art will clearly be able to envision further modifications to the principles described herein.
[0111] List of reference numerals
[0112] 1 Sample of a subject
[0113] 2 Purification of the sample
[0114] 3 mRNA
[0115] 4 Reverse transcription
[0116] 5 cDNA
[0117] 6 Amplification
[0118] 7 Polymerase chain reaction (PCR)
[0119] 8 PCR product
[0120] 9 Fragmentation
[0121] 10 Short fragment
[0122] 11 Passivation
[0123] 12 Passivated fragment
[0124] 13 Tag
[0125] 14 Ligation to the tag
[0126] 15 Pool of tagged catcher oligonucleotides
[0127] 16 Pool of starting oligonucleotides
[0128] 17 Promoter segment
[0129] 18 Random segment as potential complementary sequence to the catcher oligonucleotide
[0130] 19 Binding segment
[0131] 20 Single-stranded catcher oligonucleotide
[0132] 21 Segment complementary to the random segment
[0133] 22 Hybridization reaction
[0134] 23 Removal of the complex from the solution
[0135] 24 Continue with the pool of starting oligonucleotides reduced
[0136] 100 Scaffold oligonucleotide
[0137] 101 Single-tube reaction
[0138] 102 Hybridization reaction with the scaffold oligonucleotide
[0139] 103 DNA extension reaction
[0140] 104 Double-stranded DNA molecule
[0141] 105 Promoter activity
[0142] 106 Transcription reaction
[0143] 107 Target-unrelated sgRNA molecule
[0144] 108 sgRNA scaffold segment
[0145] 109 Reaction steps in a single tube
Claims
1. A method for obtaining an enriched population of target polynucleotides, the method comprising: (i) purifying a population of mRNA molecules from a sample obtained from a subject; (ii) preparing cDNA from the mRNA molecules of step (i); (iii) amplifying one or more target sequences from the cDNA obtained in step (ii) to obtain a pool of DNA molecules; (iv) fragmenting the amplified DNA molecules to obtain fragments; (v) ligating the fragments of step (iv) to a tag capable of binding to a cognate interactant to obtain a pool of tagged catcher oligonucleotides; (vi) providing a pool of starting oligonucleotides for preparing a pool of synthetic single-guide RNAs (sgRNAs) for an sgRNA-guided nucleic acid-binding protein, wherein the starting oligonucleotides comprise a promoter segment, a random segment that is a potential complementary sequence to the catcher oligonucleotide, and a binding segment that is complementary to at least a portion of a scaffold sequence for interacting with the sgRNA-guided nucleic acid-binding protein; (vii) hybridizing the pool of starting oligonucleotides with the tagged catcher oligonucleotides; (viii) removing complexes of starting oligonucleotides and tagged catcher oligonucleotides from the pool of starting oligonucleotides by binding the tag to a cognate interactant, thereby obtaining a reduced pool of starting oligonucleotides; (ix) preparing a pool of sgRNAs using the reduced pool of starting oligonucleotides obtained in step (viii); (x) using the pool of sgRNAs obtained in step (ix) to cleave a mixture of polynucleotides obtained from a test sample by an sgRNA-guided nucleic acid-binding protein; and (xi) size-selecting one or more uncleaved target polynucleotides from the mixture of polynucleotides obtained in step (x); wherein step (ix) of preparing a pool of sgRNAs using the reduced pool of starting oligonucleotides obtained in step (viii) comprises: hybridizing a single-stranded oligonucleotide to the starting oligonucleotides obtained in step (viii) via a binding element present in the starting oligonucleotides; filling in the single-stranded portions of the hybridized complex by a DNA extension reaction; amplifying the resulting double-stranded template molecules; transcribing the template into an RNA molecule via the promoter segment; and wherein the sgRNA comprises a sequence not specific to the target polynucleotide and an sgRNA scaffold segment.
2. The method according to claim 1, wherein, In step (iv), the amplified DNA molecules are fragmented to a size of 20 to 30 bp.
3. The method according to claim 1, wherein, In step (viii), the tag is bound to a cognate interactant located on magnetic beads or a suitable surface.
4. A method for obtaining a pool of personalized target-agnostic synthetic single-guide RNAs (sgRNAs) for an sgRNA-guided nucleic acid-binding protein, the method comprising: (i) purifying a population of mRNA molecules from a sample obtained from a subject; (ii) preparing cDNA from the mRNA molecules of step (i); (iii) Amplify one or more target sequences from the cDNA obtained in step (ii) to obtain a pool of DNA molecules; (iv) Fragment the amplified DNA molecules to obtain fragments; (v) Link the fragments of step (iv) to a tag capable of binding to a cognate interactant to obtain a pool of tagged catcher oligonucleotides; (vi) Provide a pool of starting oligonucleotides for preparing a pool of synthetic single-guide RNAs (sgRNAs) for an sgRNA-guided nucleic acid-binding protein, wherein the starting oligonucleotides comprise a promoter segment, a random segment that serves as a potential complementary sequence to the catcher oligonucleotides, and a binding segment that is complementary to at least a portion of a scaffold sequence for interacting with the sgRNA-guided nucleic acid-binding protein; (vii) Hybridize the pool of starting oligonucleotides with the tagged catcher oligonucleotides; (viii) Remove complexes of starting oligonucleotides and tagged catcher oligonucleotides from the pool of starting oligonucleotides by binding the tag to a cognate interactant, thereby obtaining a reduced pool of starting oligonucleotides; and (ix) Prepare a pool of sgRNAs with the reduced pool of starting oligonucleotides obtained in step (viii); wherein, step (ix) of preparing a pool of sgRNAs with the reduced pool of starting oligonucleotides obtained in step (viii) comprises: Hybridizing a single-stranded oligonucleotide to the starting oligonucleotides obtained in step (viii) via a binding element present in the starting oligonucleotides; Filling in the single-stranded portions of the hybridization complexes by a DNA extension reaction; Amplifying the resulting double-stranded template molecules; Transcribing the template into an RNA molecule via the promoter segment; and wherein, the sgRNA comprises a sequence not specific to the target polynucleotide and an sgRNA scaffold segment.
5. The method according to claim 4, wherein In step (iv), fragment the amplified DNA molecules to a size of 20 to 30 bp.
6. The method according to claim 4, wherein In step (viii), bind the tag to a cognate interactant located on magnetic beads or a suitable surface.
7. The method according to any one of claims 1 to 6, wherein the amplification (iii) is carried out as a polymerase chain reaction (PCR).
8. The method according to any one of claims 1 to 6, wherein the tag capable of binding to a cognate interactant is biotin and wherein the cognate interactant is streptavidin.
9. The method according to claim 8, wherein the step of linking the fragment to the biotin tag comprises tailing with activated biotin, a ligation reaction with biotin, or linking to biotin via click chemistry.
10. The method according to any one of claims 1 to 6, wherein the sgRNA-guided nucleic acid-binding protein is a DNA-binding Cas protein.
11. The method according to claim 10, wherein the DNA-binding Cas protein is a member of the Cas9 protein family.
12. The method according to claim 11, wherein, The DNA-binding Cas protein is Cas9 protein or a derivative thereof.
13. The method according to any one of claims 1 to 6, wherein the random segment comprises 10 to 30 random nucleotides.
14. The method according to claim 13, wherein, The random segment comprises 20 random nucleotides.
15. The method according to any one of claims 1 to 6, wherein steps (vii) and (viii) are repeated 1, 2, 3, 4, 5 times or more.
16. The method according to any one of claims 1 to 6, wherein the one or more target polynucleotides or target sequences represent a gene, one or more exons of a gene and / or its open reading frame or subpart; or a group of different genes, a group of one or more exons of different genes and / or a group of its open reading frame or subpart, or any combination of any of the previously mentioned elements.
17. The method according to any one of claims 1 to 3, further comprising, as step (xii), sequencing the size-selected uncut target polynucleotides.
Citation Information
Patent Citations
Methods of depleting a target molecule from an initial collection of nucleic acids, and compositions and kits for practicing the same
US20150225773A1
Depletion of abundant sequences by hybridization (DASH)
US20180051320A1
Compositions and methods for targeted depletion, enrichment, and partitioning of nucleic acids using crispr / cas system proteins
US20180298421A1