DNA targeted capture sequencing methods, systems, and devices based on CRISPR-dCas9

By employing a CRISPR-dCas9-based DNA targeting capture method, the dCas9-sgRNA complex is used to directly capture ultra-long and ultra-short nucleic acids, solving the problems of complex operation, long time consumption, and high cost in existing technologies. This method achieves effective enrichment of ultra-long and ultra-short nucleic acids and is suitable for various clinical tests.

CN118006746BActive Publication Date: 2026-07-17BEIJING CAPITALBIO MEDLAB CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING CAPITALBIO MEDLAB CO LTD
Filing Date
2024-02-08
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing targeted sequencing technologies are complex, time-consuming, and costly, and cannot directly capture ultra-long and ultra-short nucleic acids. Furthermore, the application of the CRISPR/Cas9 system in the field of nucleic acid capture is limited.

Method used

A complex is formed by inactivated Cas9 (dCas9) and sgRNA. Biotin-labeled dCas9 specifically binds to the target sequence, and the target sequence is enriched using methods such as streptavidin, thereby directly capturing ultra-long and ultra-short nucleic acids.

Benefits of technology

It simplifies the operation process, achieves effective enrichment of ultra-long and ultra-short nucleic acids, overcomes the shortcomings of existing technologies, and is applicable to fields such as tumor detection, pathogen detection, and drug resistance gene detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118006746B_ABST
    Figure CN118006746B_ABST
Patent Text Reader

Abstract

This invention provides a CRISPR-dCas9-based DNA targeted capture sequencing method, system, device, medium, and program product, relating to the field of precision medicine. The method includes: obtaining a DNA sequence containing a target sequence; the target sequence includes: sequences of 200-1000 bp, sequences greater than or equal to 1000 bp, and sequences less than or equal to 200 bp; obtaining a dCas9-sgRNA complex; mixing the DNA sequence and the dCas9-sgRNA complex, wherein the dCas9-sgRNA complex captures the target sequence in the DNA sequence, resulting in a dCas9-sgRNA-DNA complex formed by the binding of the dCas9-sgRNA complex to the target sequence, thereby achieving the separation and enrichment of the target DNA. This invention solves the problems of complex, time-consuming, and costly existing targeted sequencing operations, and is applicable not only to the capture of conventional length nucleic acid fragments but also to the direct capture of ultra-long and ultra-short nucleic acids.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomedical technology, and more specifically, to a CRISPR-dCas9-based DNA targeted capture sequencing method, system, device, medium, and program product. Background Technology

[0002] Modern medicine, especially personalized or precision medicine, increasingly relies on DNA analysis. DNA in clinical samples is increasingly used to find biomarkers for disease diagnosis, prognosis, and prediction. With the development of sequencing technology, second- and third-generation sequencing play a crucial role in disease research, clinical diagnosis, and personalized medicine. Genetic testing has become an indispensable tool in scientific research and medical practice. Compared to whole-genome sequencing (WGS), targeted sequencing technology aims to perform rapid and accurate sequencing analysis of specific gene regions or specific sequences within the genome. By introducing specific primers or probes, targeted sequencing can selectively amplify and sequence gene regions of interest. This precise approach plays a vital role in studying genetic variations, discovering pathogenic genes, analyzing tumor gene mutations, and detecting pathogenic microorganisms and their drug resistance genes.

[0003] Currently, the most commonly used gene capture methods include probe hybridization capture and multiplex PCR amplification. Probe hybridization capture utilizes the principle of complementary base pairing in nucleic acids. Modified probes targeting specific genomic regions are designed according to research needs and hybridized with a nucleic acid library containing sequencing adapters, allowing the probe to specifically bind to the target DNA sequence. The complex formed by the probe and the target region can be recovered using magnetic beads or other methods, capturing the target region from the entire DNA sample. Although probe hybridization capture sequencing is generally less expensive than whole-genome sequencing, the overall capture process remains costly due to the costs of probe design and synthesis. Furthermore, probe hybridization capture technology has drawbacks such as operational complexity and time consumption. Multiplex PCR amplification is a PCR technique that simultaneously amplifies multiple target sequences in a single reaction. It amplifies multiple target sequences by introducing multiple primer pairs simultaneously in the same reaction. Compared to probe hybridization capture, multiplex PCR amplification is less expensive, simpler to operate and analyze, and has high specificity; however, primer design for this technique is more challenging. Primer design must ensure that multiple primers do not interfere with each other in the same reaction and maintain specificity and relative consistency. Therefore, multiplex PCR amplification typically has low throughput and limited application flexibility.

[0004] Furthermore, in recent years, with the development of the CRISPR / Cas (Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) / CRISPR-Associated Protein) system, its applications are no longer limited to gene editing. The CRISPR / Cas9 system can target and cleave sequences via sgRNA, and then use post-cleavage modifications to separate the target fragments. However, this cleavage strategy typically requires pretreatment such as end passivation of the sample nucleic acids, which is cumbersome; and there are often many limitations on the location of the sgRNA, such as if the sgRNA (single guide RNA) spacing is too small, leading to excessive fragmentation of the target sequence by Cas9 and affecting nucleic acid recovery efficiency. These drawbacks limit the wider application of the CRISPR / Cas9 system in nucleic acid capture detection. Modified dCas9, i.e., inactivated Cas9 (dead Cas9), although losing nuclease activity, retains DNA binding ability, bringing new possibilities for the application of the CRISPR / Cas system in nucleic acid capture. Typically, dCas9 is used in studies such as gene expression regulation and epigenetic modification. By fusing various transcriptional regulatory domains to the C-terminus of dCas9, transcription factors can be recruited or target regions modified, thereby enabling the study of transcriptional regulation of target genes. Currently, dCas9 has been widely used in transcriptional regulation, but its potential applications in nucleic acid capture still require further exploration. Summary of the Invention

[0005] This invention aims to address at least one of the technical problems existing in the prior art. To this end, this invention provides a CRISPR-dCas9-based DNA targeted capture sequencing method and system. The method utilizes a complex formed by a dead Cas9 (dCas9), which retains DNA binding ability but lacks nuclease activity, and sgRNA to directly capture the target sequence of extracted nucleic acids. The complex formed by biotin-labeled dCas9 and sgRNA can specifically bind to the target sequence, and then capture methods such as streptavidin are used to achieve effective enrichment of the target sequence. This invention solves the problems of existing targeted sequencing operations being complex, time-consuming, costly, and unable to directly capture ultra-long and ultra-short nucleic acids.

[0006] The first aspect of this application discloses a DNA targeted capture sequencing method based on CRISPR-dCas9, the method comprising:

[0007] Obtain a DNA sequence containing a target sequence; the target sequence includes: a sequence of 200-1000 bp, a sequence of 1000 bp or more, and a sequence of 200 bp or less.

[0008] Obtain the dCas9-sgRNA complex;

[0009] The DNA sequence and the dCas9-sgRNA complex are mixed, and the dCas9-sgRNA complex captures the target sequence in the DNA sequence to obtain a dCas9-sgRNA-DNA complex in which the dCas9-sgRNA complex binds to the target sequence.

[0010] In some embodiments, the target sequence includes a sequence of 10 kb or more;

[0011] Optionally, the target sequence includes a sequence of 60-120 bp; preferably, it includes a sequence of any of the following lengths: 60 bp, 80 bp, 100 bp, or 120 bp.

[0012] In some embodiments, the method for obtaining the dCas9-sgRNA complex includes:

[0013] Obtain sgRNA;

[0014] The target sgRNA is determined from the sgRNA according to the specified interval length;

[0015] Obtain the forward and reverse primers for the target sgRNA;

[0016] Based on the aforementioned forward and reverse primers, template DNA for in vitro transcription was synthesized using PCR technology to obtain PCR products;

[0017] The PCR product was purified to obtain purified sgRNA;

[0018] The purified sgRNA was assembled according to the system mixing standard to obtain the dCas9-sgRNA complex.

[0019] Optionally, the specified interval length includes any one or more of the following: 20bp, 100bp, 200bp;

[0020] Optionally, the sgRNA is obtained by the following method: determining the target sequence; determining the sgRNA on the target sequence based on the PAM sequence; and screening out the sgRNA that meets the requirements from the sgRNA on the target sequence, which is the sgRNA.

[0021] Optionally, the screening criteria include any one or more of the following: GC content within the range, homopolymer content less than or equal to the first threshold, dinucleotide repeats less than or equal to the second threshold, no hairpin structure, and no off-target effects from human genomes.

[0022] Optionally, the purification process includes: purifying the PCR product using a solid-phase medium to obtain purified sgRNA template DNA; performing in vitro transcription using an in vitro transcription kit to remove the template DNA and obtain in vitro transcribed sgRNA; and purifying the in vitro transcribed sgRNA using an RNA purification kit to obtain the purified sgRNA.

[0023] Optionally, the system mixing standard includes the following components: sgRNA, dCas9-Biotin, reaction buffer, and nuclease-free water.

[0024] In some embodiments, the dCas9-sgRNA complex comprises a complex formed by the binding of a nuclease-free Cas9 protein to sgRNA.

[0025] Optionally, the dCas9 protein includes conventional dCas9 protein as well as various dCas9 proteins formed through other modification processes.

[0026] In some embodiments, the DNA sequence containing the target sequence is derived from one or more of the following samples: human cells, Acinetobacter baumannii ATCC 19606, Klebsiella pneumoniae ATCC 43816, Escherichia coli ATCC 11775, Pseudomonas aeruginosa ATCC 27853, and Staphylococcus aureus ATCC 43300.

[0027] In some embodiments, the method further includes: capturing the dCas9-sgRNA-DNA complex using a solid-phase medium to obtain a captured solid-phase medium; washing and purifying the solid-phase medium to obtain purified target DNA;

[0028] The purified target DNA was used for library construction, sequencing, and data analysis to calculate the enrichment fold.

[0029] Optionally, the enrichment factor can be calculated as follows: enrichment factor = percentage of target sequence reads after capture / percentage of target sequence reads not captured;

[0030] Optionally, the solid medium is a magnetic bead; and the surface of the magnetic bead is fixed with streptavidin.

[0031] A second aspect of this application discloses a CRISPR-dCas9-based DNA targeted capture sequencing system, the system comprising:

[0032] The first acquisition unit is used to acquire a DNA sequence containing a target sequence; the target sequence includes: a sequence of 200-1000 bp, a sequence of 1000 bp or more, and a sequence of 200 bp or less.

[0033] The second acquisition unit is used to acquire the dCas9-sgRNA complex;

[0034] A capture unit is used to mix the DNA sequence and the dCas9-sgRNA complex, wherein the dCas9-sgRNA complex captures the target sequence in the DNA sequence to obtain a dCas9-sgRNA-DNA complex in which the dCas9-sgRNA complex binds to the target sequence.

[0035] A third aspect of this application discloses a computer device, the device comprising: a memory and a processor; the memory being used to store program instructions; the processor being used to invoke the program instructions, which, when executed, are used to perform the steps of the method disclosed in the first aspect.

[0036] The fourth aspect of this application discloses a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method disclosed in the first aspect.

[0037] The fifth aspect of this application discloses a computer program product, including a computer program that, when executed by a processor, implements the steps of the method disclosed in the first aspect.

[0038] The sixth aspect of this application discloses the application of the CRISPR-dCas9-based DNA targeted capture sequencing method disclosed in the first aspect in the preparation of DNA detection, diagnostic and therapeutic reagents.

[0039] This application has the following beneficial effects:

[0040] 1. This application innovatively discloses a CRISPR-dCas9-based DNA targeted capture sequencing method. It utilizes a complex formed by a dead Cas9 (dCas9), lacking nuclease activity but retaining DNA binding ability, and sgRNA to directly capture the target sequence of extracted nucleic acids. The complex formed by biotin-labeled dCas9 and sgRNA can specifically bind to the target sequence, and then capture methods such as streptavidin are used to achieve effective enrichment of the target sequence. This invention solves the problems of existing targeted sequencing methods, such as complexity, time consumption, high cost, and inability to directly capture ultra-long and ultra-short nucleic acids.

[0041] 2. This application directly captures ultra-long and ultra-short DNA, demonstrating that this technical approach can effectively enrich ultra-long (>10kb) and ultra-short DNA (60bp-120bp), improving upon the current situation where reported CRISPR-dCas9 capture systems only capture nucleic acids of conventional size (200-1000bp). Furthermore, it overcomes the problem that when the CRISPR-dCas9 complex binds to double-stranded DNA, the double-stranded DNA unwinds, forming a hybrid strand with the sgRNA in the CRISPR-dCas9 complex. Therefore, whether ultra-short DNA can effectively bind to the CRISPR-dCas9 complex still requires further research. Additionally, there are currently no reports on whether the CRISPR-dCas9 system can effectively extract ultra-long DNA (>10kb) for effective enrichment due to its large molecular weight.

[0042] 3. This application can effectively overcome the cumbersome operation of the previously reported CRISPR-dCas9 capture system, which requires the preparation of nucleic acid libraries in advance and the modification of sgRNA used for capture. This process directly uses biotin-labeled dCas9, and the sgRNA preparation and capture process are simple. Moreover, there is no need to construct a library before capture, which simplifies the operation of the entire system.

[0043] Based on the above characteristics, this technical approach has a wide range of applications, including fusion gene detection in tumor detection and pathogen sequence species annotation, drug resistance gene detection, and cell-free DNA detection in targeted pathogen detection. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 This is a schematic diagram of the method flow provided in the first aspect of the present invention;

[0046] Figure 2 This is a schematic flowchart of the system provided in the second aspect of the present invention;

[0047] Figure 3 This is a schematic diagram of a computer device provided in an embodiment of the present invention;

[0048] Figure 4 This is a schematic diagram of the technical principle of DNA targeting and capture based on CRISPR-dCas9 provided in the embodiments of the present invention;

[0049] Figure 5 This is an agarose gel electrophoresis image of the extracted nucleic acid from the simulated sample provided in this embodiment of the invention;

[0050] Figure 6 This is the enrichment fold of each target sequence provided in the embodiments of the present invention (with sgRNA and dCas9 added alone as a control);

[0051] Figure 7 This refers to the enrichment fold of the targeted capture experiment with different sgRNA intervals provided in the embodiments of the present invention. Detailed Implementation

[0052] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0053] In some of the processes described in the specification, claims, and accompanying drawings of this invention, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0054] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0055] Figure 1This is a schematic diagram of a DNA targeted capture sequencing method based on CRISPR-dCas9 provided in an embodiment of the present invention. Specifically, the method includes the following steps:

[0056] 101: Obtain a DNA sequence containing the target sequence; the target sequence includes: a sequence of 200-1000 bp, a sequence of 1000 bp or more, and a sequence of 200 bp or less;

[0057] In some embodiments, the target sequence includes a sequence of 10 kb or more;

[0058] Optionally, the target sequence includes a sequence of 60-120 bp; preferably, it includes any of the following lengths: 60 bp, 80 bp, 100 bp, and 120 bp. It should be noted that this embodiment only lists the above-mentioned lengths, but in actual operation, it is not limited to the above-mentioned sequences.

[0059] In some embodiments, the DNA sequence containing the target sequence is derived from one or more of the following samples: human cells, Acinetobacter baumannii ATCC 19606, Klebsiella pneumoniae ATCC 43816, Escherichia coli ATCC 11775, Pseudomonas aeruginosa ATCC 27853, and Staphylococcus aureus ATCC 43300.

[0060] 102: Obtain the dCas9-sgRNA complex;

[0061] In some embodiments, the method for obtaining the dCas9-sgRNA complex includes:

[0062] Obtain sgRNA; determine the target sgRNA from the sgRNA according to a specified interval length; obtain the forward and reverse primers for the target sgRNA; based on the forward and reverse primers, synthesize template DNA for in vitro transcription using PCR technology to obtain PCR products; purify the PCR products to obtain purified sgRNA; assemble the purified sgRNA according to the system mixing standard to obtain the dCas9-sgRNA complex;

[0063] Optionally, the specified interval length includes any one or more of the following: 20bp, 100bp, 200bp;

[0064] Optionally, the sgRNA is obtained by the following method: determining the target sequence; determining the sgRNA on the target sequence based on the PAM sequence; and screening out the sgRNA that meets the requirements from the sgRNA on the target sequence, which is the sgRNA.

[0065] Optionally, the screening criteria include any one or more of the following: GC content within the range (25-75%), homopolymer less than or equal to the first threshold (5), binucleonucleotide repeat less than or equal to the second threshold (3), no hairpin structure, and no off-target effects from human genome; wherein, the preferred range for GC content is 25-75%, the preferred first threshold is 5, and the preferred second threshold is 3;

[0066] Optionally, the purification process includes: purifying the PCR product using a solid-phase medium to obtain purified sgRNA template DNA; performing in vitro transcription using an in vitro transcription kit to remove the template DNA and obtain in vitro transcribed sgRNA; and purifying the in vitro transcribed sgRNA using an RNA purification kit to obtain the purified sgRNA.

[0067] Optionally, the system mix standard includes the following components: sgRNA, dCas9-Biotin, reaction buffer, and nuclease-free water. Assembly is completed by incubation at room temperature (25°C) for 30 min.

[0068] In some embodiments, the dCas9-sgRNA complex comprises a complex formed by the binding of a nuclease-free Cas9 protein to sgRNA.

[0069] Optionally, the dCas9 protein includes conventional dCas9 protein as well as various dCas9 proteins formed through other modification processes.

[0070] The primer sequences corresponding to sgRNA are as follows:

[0071]

[0072]

[0073]

[0074] 103: Mix the DNA sequence and the dCas9-sgRNA complex, wherein the dCas9-sgRNA complex captures the target sequence in the DNA sequence to obtain a dCas9-sgRNA-DNA complex in which the dCas9-sgRNA complex binds to the target sequence.

[0075] In some embodiments, the method further includes: capturing the dCas9-sgRNA-DNA complex using a solid-phase medium to obtain a captured solid-phase medium; washing and purifying the solid-phase medium to obtain purified target DNA, i.e., the dCas9-sgRNA-DNA complex.

[0076] The purified target DNA was used for library construction, sequencing, and data analysis to calculate the enrichment fold.

[0077] Optionally, the enrichment factor can be calculated as follows: enrichment factor = percentage of target sequence reads after capture / percentage of target sequence reads not captured;

[0078] Optionally, the solid medium is a magnetic bead; streptavidin is fixed on the surface of the magnetic bead.

[0079] The DNA-dCas9-sgRNA complex captured on the surface of magnetic beads can be easily and quickly separated from the DNA library or mixture by magnetic separation technology.

[0080] The DNA in the DNAdCas9-sgRNA complex captured by magnetic beads can be purified using various DNA purification techniques. The purified DNA can then be analyzed using sequencing technology to interpret its sequence information.

[0081] Figure 2 This invention provides a CRISPR-dCas9-based DNA targeted capture sequencing system, comprising:

[0082] The first acquisition unit 201 is used to acquire a DNA sequence containing a target sequence; the target sequence includes: a sequence of 200-1000 bp, a sequence of 1000 bp or more, and a sequence of 200 bp or less.

[0083] The second acquisition unit 202 is used to acquire the dCas9-sgRNA complex;

[0084] The capture unit 203 is used to mix the DNA sequence and the dCas9-sgRNA complex, wherein the dCas9-sgRNA complex captures the target sequence in the DNA sequence to obtain a dCas9-sgRNA-DNA complex in which the dCas9-sgRNA complex binds to the target sequence.

[0085] Figure 3 This invention provides a computer device comprising: a memory and a processor; the memory for storing program instructions; and the processor for calling the program instructions, which, when executed, perform the steps of the method described above.

[0086] This invention also includes a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method described above.

[0087] This invention also discloses a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.

[0088] Specifically, Example 1 illustrates the entire process using a targeted capture sequencing experiment of ultra-long nucleic acid fragments as an example:

[0089] 1. Experimental materials:

[0090] Samples: Acinetobacter baumannii ATCC 19606, Klebsiella pneumoniae ATCC 43816, Escherichia coli ATCC 11775, Pseudomonas aeruginosa ATCC 27853, Staphylococcus aureus ATCC43300, and human cells.

[0091] Reagents: Microbial genome extraction kit, dCas9 protein, streptavidin magnetic beads, PCR mix, T7 in vitro transcription kit, RNA purification kit, transposase library preparation kit, etc.

[0092] 2. Experimental methods:

[0093] Step 1: Design sgRNA sequences for the 5 target sequences in the experimental system.

[0094] First, all possible sgRNAs on the target sequence are identified based on the Protospacer Adjacent Motif (PAM) sequence (NGG). Low-quality sgRNAs are excluded (considering factors such as GC content, homopolymers, binucleonucleotide repeats, hairpin structures, and off-target effects in the human genome). Then, a set of sgRNAs is selected according to specific intervals (intervals include 20bp, 100bp, 200bp, etc.). All the sgRNAs from the target sequences together constitute the sgRNA library.

[0095] Step 2: Preparation of sgRNA template strand for in vitro transcription.

[0096] Based on the above sgRNA sequence, sgRNA primers for in vitro transcription were designed, and all forward primers were mixed together in equal quantities. Template DNA for in vitro transcription was synthesized by PCR. The template sequence is: AAAAGCACCGACTCGGTGCCACTTTTTCAAGTTGATAACGGACTAGCCTTATTTTAACTTGCTATTTCTAGCTCTAAAAC. The forward primer sequence is: TTCTAATACGACTCACTATAGNNNNNNNNNNNNNNNNNNNNGTTTTAGAGCTAGA, where N represents the sequence in the sgRNA that is complementary to the target DNA. The reverse primer sequence is: AAAAGCACCGACTCGGTGCC.

[0097] The amplification system is as follows:

[0098] Element 50 μl reaction system PCR Mix 12.5μl 10μM forward primer 2.5μl 10μM reverse primer 2.5μl 1μM template DNA 2μl Nuclease-free water 18μl

[0099] The amplification conditions are as follows:

[0100]

[0101] Step 3: Purify the PCR product using magnetic beads.

[0102] After the reaction, the PCR product was purified using magnetic beads. The purification steps were as follows: 90 μl of AMPure XP magnetic beads were added to the PCR product, mixed thoroughly, and allowed to stand for 5 min. The PCR tube was placed in a magnetic rack to separate the magnetic beads and liquid. After the solution became clear, the supernatant was carefully removed. The PCR tube remained in the magnetic rack, and 200 μl of nuclease-free water was added to freshly prepared 80% ethanol to rinse the magnetic beads. After incubation at room temperature for 30 sec, the supernatant was carefully removed. This rinsing process was repeated once. 10 μl of pipette was used to remove any remaining liquid. The PCR tube remained in the magnetic rack, and the magnetic beads were left to dry at room temperature. 22 μl of nuclease-free water was added, and the mixture was pipetted until thoroughly mixed. After standing at room temperature for 5 min, the PCR tube was briefly centrifuged and placed in a magnetic rack to stand. After the solution became clear, 20 μl of the supernatant was carefully transferred to a new PCR tube. The concentration of the recovered product was determined using a Qubit analyzer.

[0103] Step 4: In vitro transcription of sgRNA.

[0104] In vitro transcription of sgRNA was performed using the T7 in vitro transcription kit. The steps were as follows: Clean the lab bench to prevent ribonuclease contamination. Add the following reagents to the PCR tube in sequence: 10 μl NTP Buffer Mix, 1 μg sgRNA template DNA purified in the previous step, 2 μl T7 RNA polymerase Mix, and water to a final volume of 30 μl. The reaction conditions were: 37℃, 16 h.

[0105] Remove template DNA after the reaction: Add 20 μl of nuclease-free water and 2 μl of DNase to every 30 μl of reaction, mix, and incubate at 37°C for 15 min.

[0106] Step 5: RNA purification.

[0107] RNA was purified using an RNA purification kit, and the concentration of sgRNA was determined using Qubit.

[0108] Step 6: Assembly of the dCas9-sgRNA complex.

[0109] Mix the components according to the system shown in the table below:

[0110] Components Dosage Nuclease-free water Make up to 20μl reaction buffer 2μl sgRNA 321.7ng dCas9-Biotin 2μl

[0111] The above system was assembled after incubation at room temperature (25°C) for 30 minutes.

[0112] Step 7: Prepare simulated samples and extract genomes.

[0113] Equal amounts of strains A. baumannii ATCC 19606, K. pneumoniae ATCC 43816, E. coli ATCC 11775, P. aeruginosa ATCC 27853, and S. aureus ATCC 43300 were mixed together and then mixed with human cells to prepare a simulated sample. Genomic nucleic acids were obtained using a microbial genome extraction kit.

[0114] Step 8: Incubate the dCas9-sgRNA complex with the simulated sample genome.

[0115] Prepare the reaction mixture in the PCR tubes according to the table below:

[0116] Components or operations Dosage reaction buffer 2μl Simulated sample DNA 16μl dCas9-sgRNA complex 2μl Total volume 20μl

[0117] Gently tap to mix and briefly incubate on a PCR instrument as follows: 37°C, 45 min, to achieve binding of the dCas9-sgRNA complex to the target sequence.

[0118] Step 9: Affinity magnetic beads specifically capture dCas9-sgRNA-DNA complexes.

[0119] Take 10 μl of streptavidin magnetic beads, wash the beads twice with 1×dCas9 binding buffer, then suspend them in 5 μl of 1×dCas9 binding buffer. Add this to the 20 μl dCas9-sgRNA-DNA incubation mixture and incubate at room temperature for 10 min by rotation. Place the PCR tube on a magnetic separator to separate the magnetic beads and supernatant.

[0120] Step 10: Elution, Library Construction and Sequencing

[0121] Wash the magnetic beads three times with 1× binding buffer. Resuspend the magnetic beads in 30 μl of 0.2% SDS and incubate at room temperature for 5 min. Purify the DNA in the supernatant. The purification steps are as follows: Add 30 μl of AMPure XP Beads to the PCR product, mix thoroughly, and let stand for 5 min. Place the PCR tube in a magnetic rack to separate the magnetic beads and liquid. After the solution is clear, carefully remove the supernatant. Keep the PCR tube in the magnetic rack at all times, add 200 μl of nuclease-free water and freshly prepared 80% ethanol to wash the magnetic beads, incubate at room temperature for 30 sec, and carefully remove the supernatant. Repeat the rinsing once. Keep the PCR tube in the magnetic rack at all times, open the cap and dry the magnetic beads at room temperature. Add 22 μl of nuclease-free water, mix thoroughly by pipetting, and let stand at room temperature for 5 min. Briefly centrifuge the PCR tube and place it in the magnetic rack to stand. After the solution is clear, carefully transfer 20 μl of supernatant to a new PCR tube. Measure the concentration of the recovered product using Qubit.

[0122] Library construction was performed using a transposase library preparation kit, and sequencing was performed using the Illumina sequencing platform.

[0123] Step 11: Data Analysis After Disconnection

[0124] After the data was processed, adapters and low-quality sequences were removed using the FASTP software. Then, sequence alignment was performed using BWA, comparing the data with the human reference genome, the microbial reference genome of the simulated sample, and the target sequence. The alignment results were then statistically analyzed, and the enrichment factor was calculated (enrichment factor = percentage of target sequence reads after capture / percentage of target sequence reads not captured).

[0125] 2. Experimental Results

[0126] As attached Figure 5 The figure shows the agarose gel electrophoresis of the extracted simulated sample nucleic acid (M: Marker; 1-3: simulated sample nucleic acid). As can be seen from the figure, the simulated sample nucleic acid used for dCas9 capture has a main peak of >10kb after extraction.

[0127] As attached Figure 6 As shown, the enrichment fold for each target sequence is displayed (with sgRNA and dCas9 added separately as controls); specifically, it shows the enrichment fold for the five target sequences in the experimental system. As can be seen from the figure, the Cas9 enrichment process can achieve an average enrichment of 6.6-fold for the target sequences, with a maximum enrichment of 25.9-fold.

[0128] Specifically, Example 2 illustrates the entire process using a targeted capture sequencing experiment of ultrashort nucleic acid fragments as an example:

[0129] 1. Experimental materials

[0130] Samples: Primers were designed for PCR amplification to obtain target sequences of different lengths, sul2, including 60bp, 80bp, 100bp, and 120bp. A 100bp non-target sequence was amplified as a background sequence. The target and non-target sequences were mixed at a ratio of 1:99 for the capture experiment.

[0131] 2. Experimental Methods

[0132] In this embodiment, four sgRNAs were designed to capture target sequences of different lengths. The specific implementation method is the same as the dCas9 capture procedure in Example 1.

[0133] 3. Experimental Results

[0134] The table below shows the percentage of target sequences that were not captured and those that were captured, as well as the enrichment fold for different fragment lengths. As can be seen from the table, the dCas9 capture process can enrich nucleic acid sequences from 60bp to 120bp, with an average enrichment fold of 248.9-fold and a maximum enrichment fold of 552-fold. This result demonstrates that this technique can effectively enrich ultrashort nucleic acid sequences.

[0135]

[0136] In addition, in this embodiment, we compared the capture effects of different sgRNA intervals (i.e., the targeted capture experiment of flat sgRNA), selecting two target sequences, catB7 and sul2, as test genes. The sgRNA sets were designed with intervals of 200 bp, 100 bp, and 20 bp between the selected sgRNAs. The specific implementation method is the same as the dCas9 capture procedure in Example 1.

[0137] The results are as follows Figure 7 As shown, dCas9 capture procedures with different spacings all achieved target sequence enrichment greater than 10-fold, and the capture effect increased as the sgRNA spacing decreased. This experiment demonstrates that dCas9 capture procedures can achieve better capture results with a denser sgRNA design.

[0138] This implementation demonstrates that the dCas9 capture process can improve capture performance by utilizing a planar sgRNA design. Compared to CRISPR-Cas9 cleavage, the dCas9 capture process can utilize as much sgRNA as possible on the target sequence, thereby increasing capture sites and improving capture performance.

[0139] The significance of targeted capture of ultra-long or ultra-short nucleic acid fragments lies in:

[0140] Current gene capture technologies primarily target standard-sized nucleic acids, such as those in libraries matched to next-generation sequencing (200-1000 bp). However, in practical clinical research and applications, ultrashort and ultralong nucleic acid fragments also play an irreplaceable role, providing crucial genetic information for disease research and pathogen detection. Ultrashort nucleic acid fragments, such as circulating tumor DNA (ctDNA), enable liquid biopsies, making early cancer detection and disease monitoring more convenient and accurate. Furthermore, the detection of circulating free DNA (cfDNA) can be used to diagnose various infectious diseases, including viral, bacterial, fungal, and parasitic infections. By analyzing cfDNA in blood, specific gene sequences of pathogens can be detected, helping to determine the type and extent of infection. The analysis of these short fragments allows clinical research to gain a more comprehensive understanding of the molecular basis of diseases, providing new insights for precision medicine and treatment planning.

[0141] Meanwhile, research on ultralong nucleic acid fragments plays a crucial role in genomics, structural variation, and genetic disease research. Obtaining ultralong nucleic acid fragment sequences helps to delve deeper into complex structural variations in the genome, such as gene fusions, revealing the underlying mechanisms of diseases and laying the foundation for the development of novel treatments. For pathogen detection, ultralong nucleic acid sequence analysis can contribute to more accurate pathogen identification and drug resistance gene annotation. Therefore, capture sequencing of both ultrashort and ultralong nucleic acid fragments can inject new vitality into clinical research, bringing new opportunities for a deeper understanding of diseases and the development of treatment strategies.

[0142] This invention presents a novel targeted enrichment technique, distinct from probe capture and multiplex amplification. It utilizes biotin-labeled, inactive Cas9 (dCas9) lacking nuclease activity but retaining DNA-binding capacity to capture target nucleic acids. Nucleic acids extracted from the sample DNA can be captured directly without fragmentation or library construction. Unlike Cas9 capture, because dCas9 lacks nuclease activity, sgRNA can be laid flat on the target sequence at certain intervals, providing more capture sites. This invention verifies that smaller sgRNA intervals result in better capture performance. This invention can capture and enrich ultra-long (>10kb) and ultra-short (60bp) nucleic acid fragments, applicable to different types of nucleic acid capture, such as relatively complete genomic nucleic acid DNA and short cell-free DNA. Combined with third-generation and second-generation sequencing, it can obtain more target sequence information and can be used for fusion gene detection in tumor detection and sequence species annotation, drug resistance gene detection, and cell-free DNA detection in targeted pathogen detection. Furthermore, this invention directly captures extracted nucleic acids without prior library construction, and the sgRNA used requires no special modification, simplifying the overall capture process.

[0143] The verification results of this verification embodiment show that assigning inherent weights to indications can moderately improve the performance of this method compared to the default settings.

[0144] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0145] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0146] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0147] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0148] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0149] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0150] The computer device provided by the present invention has been described in detail above. For those skilled in the art, there will be changes in the specific implementation and application scope based on the ideas of the embodiments of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A DNA targeted capture sequencing method based on CRISPR-dCas9, characterized in that, The method includes: Obtain a DNA sequence containing a target sequence; the specific length of the target sequence is 60bp, 80bp, 100bp, or 120bp. Obtain sgRNA; identify the target sgRNA from the sgRNA according to a specified interval length; obtain the forward and reverse primers for the target sgRNA; synthesize template DNA for in vitro transcription using PCR technology based on the forward and reverse primers to obtain PCR products; purify the PCR products to obtain purified sgRNA; assemble the purified sgRNA according to the system mixing standard to obtain the dCas9-sgRNA complex; the specified interval length is 20 bp; The DNA sequence and the dCas9-sgRNA complex are mixed, and the dCas9-sgRNA complex captures the target sequence in the DNA sequence to obtain a dCas9-sgRNA-DNA complex in which the dCas9-sgRNA complex binds to the target sequence.

2. The DNA targeted capture sequencing method based on CRISPR-dCas9 according to claim 1, characterized in that, The sgRNA is obtained by the following method: determining the target sequence; determining the sgRNA on the target sequence based on the PAM sequence; and selecting the sgRNA that meets the requirements from the sgRNA on the target sequence.

3. The DNA targeted capture sequencing method based on CRISPR-dCas9 according to claim 2, characterized in that, The screening criteria include any one or more of the following: GC content within the range, homopolymer content less than or equal to the first threshold, binucleonucleotide repeats less than or equal to the second threshold, no hairpin structure, and no off-target effects from human genomes.

4. The DNA targeted capture sequencing method based on CRISPR-dCas9 according to claim 1, characterized in that, The purification process includes: purifying the PCR product using a solid-phase medium to obtain purified sgRNA template DNA; performing in vitro transcription using an in vitro transcription kit to remove the template DNA and obtain in vitro transcribed sgRNA; and purifying the in vitro transcribed sgRNA using an RNA purification kit to obtain the purified sgRNA.

5. The DNA targeted capture sequencing method based on CRISPR-dCas9 according to claim 1, characterized in that, The system mix standard includes the following components: sgRNA, dCas9-Biotin, reaction buffer, and nuclease-free water.

6. The DNA targeted capture sequencing method based on CRISPR-dCas9 according to claim 1, characterized in that, The dCas9-sgRNA complex comprises a complex formed by the binding of dCas9 protein, which has no nuclease activity, and sgRNA.

7. The DNA targeted capture sequencing method based on CRISPR-dCas9 according to claim 6, characterized in that, The dCas9 protein includes conventional dCas9 protein as well as various dCas9 proteins formed through other modification processes.

8. The DNA targeted capture sequencing method based on CRISPR-dCas9 according to any one of claims 1-7, characterized in that, The DNA sequence containing the target sequence is derived from one or more of the following samples: human cells, Acinetobacter baumannii ATCC 19606, Klebsiella pneumoniae ATCC 43816, Escherichia coli ATCC 11775, Pseudomonas aeruginosa ATCC 27853, and Staphylococcus aureus ATCC 43300.

9. The DNA targeted capture sequencing method based on CRISPR-dCas9 according to any one of claims 1-7, characterized in that, The method further includes: capturing the dCas9-sgRNA-DNA complex using a solid-phase medium to obtain a captured solid-phase medium; washing and purifying the solid-phase medium to obtain purified target DNA. The purified target DNA was used for library construction, sequencing, and data analysis to calculate the enrichment fold.

10. The CRISPR-dCas9-based DNA targeted capture sequencing method according to claim 9, characterized in that, The enrichment factor is calculated as follows: Enrichment factor = Percentage of target sequence reads after capture / Percentage of target sequence reads not captured.

11. The CRISPR-dCas9-based DNA targeted capture sequencing method according to claim 9, characterized in that, The solid medium is a magnetic bead; the surface of the magnetic bead is fixed with streptavidin.