A library construction method for low-concentration nucleic acids and application thereof in preparation of detection and diagnosis reagents

By adding specific DNA fragments to low-concentration nucleic acid samples and cutting them using the CRISPR/Cas9 system, the problems of high failure rate and adapter self-ligation rate in the construction of low-concentration nucleic acid libraries are solved, improving library quality and the detection rate of target sequences, making it suitable for pathogen metagenomic detection.

CN116716666BActive Publication Date: 2026-07-21HANGZHOU MATRIDX BIOTECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU MATRIDX BIOTECH CO LTD
Filing Date
2023-03-14
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies suffer from high failure rates, high adapter self-ligation rates, and high background nucleic acid sequence proportions in library construction of low-concentration nucleic acid samples, resulting in poor sequencing data quality and increased costs, especially with limited effectiveness in pathogen metagenomic detection.

Method used

After adding specific DNA fragments to low-concentration DNA samples, sgRNA designed using the CRISPR/Cas9 system is used to cut and remove the specific DNA fragments, thereby improving the success rate and quality of library construction.

Benefits of technology

It significantly reduces the proportion of adapter self-ligation, improves the detection rate of target sequences, simplifies operation, reduces costs, and is suitable for library construction of low-concentration nucleic acid samples, especially significantly improving library quality in pathogen metagenomic detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116716666B_ABST
    Figure CN116716666B_ABST
Patent Text Reader

Abstract

The application relates to a library construction method for low-concentration nucleic acid and application in preparation of a detection diagnostic reagent, and belongs to the technical field of biomolecules. The method is characterized in that a specific DNA fragment is added to low-concentration DNA, then library construction is carried out, and then the specific DNA fragment in the library is removed by using sgRNA and a CRISPR / Cas9 system designed for the specific DNA fragment. The application solves the problem of low starting concentration in library construction by first adding a specific DNA fragment, and then removes the added sequence by combining the CRISPR / Cas system after library construction, so that the data proportion of the target sequence is improved. The operation is simple, and the problems of high failure rate, high self-ligation rate of the connector, high proportion of background nucleic acid sequence and the like in the library construction of low-concentration DNA are effectively solved. The application is particularly suitable for low-concentration nucleic acid samples in pathogen metagenome sequencing, and the detection of the target sequence is improved under limited sequencing data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the application of a method for constructing libraries for low-concentration nucleic acids in the preparation of diagnostic reagents, belonging to the field of biomolecular technology. Technical Background

[0002] With the development of high-throughput DNA sequencing technology, also known as next-generation sequencing (NGS), its clinical diagnostic applications are becoming increasingly widespread. By analyzing DNA sequences, it enables precise diagnosis and treatment of diseases. Currently, from the diagnosis of hereditary diseases and the detection of tumor driver gene mutation sites to the application of pathogen metagenomic sequencing, NGS technology plays a crucial role in the diagnosis of various diseases.

[0003] Next-generation sequencing (NGS) technology is used in the diagnosis of hereditary diseases. Analysis of sequencing results can reveal chromosomal abnormalities and various genetic variations in the genome. Examples include the widely used and mature non-invasive prenatal testing (NIPT) and CNV (Catalytic Vulnerability Analysis). In oncology, the detection of mutations in known specific target genes facilitates selective cancer therapy. Furthermore, the detection of tumor driver gene large panels and tumor tumor mesenchymal stem cell (TMB) mutations are among the most promising applications in cancer treatment and companion diagnostics. In recent years, metagenomic next-generation sequencing (mNGS), based on NGS technology, has seen rapid development and application in the clinical diagnosis of infectious diseases caused by pathogenic microorganisms. mNGS technology plays a crucial role in viral sequence tracing and global pandemic control. For instance, the SARS-CoV-2 virus was first discovered using this technology, and its complete genome sequence was obtained, providing the most important foundational information for rapid understanding and knowledge of the SARS-CoV-2 virus. However, in pathogen metagenomic detection, due to the complexity of sample types and sources, there are often a large number of samples in clinical practice that are difficult to replicate and have extremely low nucleic acid content. After library construction, these samples often produce a high proportion of adapter self-ligation and irrelevant environmental DNA sequences, which poses a greater challenge to the quality of library construction, the amount of sequencing data, and the validity of sequencing analysis results.

[0004] The CRISPR / Cas system is an acquired immune system developed by bacteria over a long period of evolution. It destroys invading bacteriophages and exogenous DNA through enzymatic cleavage. Currently, the CRISPR / Cas system has become the most efficient, simplest, and most widely used gene editing system. The CRISPR / Cas system can be divided into three different types: Type I, Type II, and Type III. The most widely used gene editing technologies are primarily Type II, such as Cas9 and Cas12. The CRISPR / Cas system mainly consists of three parts: the Cas protein and two RNA components: a crRNA and a tracrRNA. The crRNA contains a nucleotide spacer sequence transcribed from the CRISPR locus, while the tracrRNA is essential for the maturation of the crRNA and the formation of the Type II monitoring complex. Through artificial design, these two RNAs are linked to form a single guide RNA (sgRNA). Guided by the sgRNA, the Cas9 protein recognizes the PAM region of the target gene and specifically cleaves the target sequence, producing a DNA double-strand break. Currently, CRISPR / Cas9 technology has become the fastest, most efficient, and most widely used gene editing technology. In recent years, the CRISPR / Cas system has also seen increasing applications in in vitro molecular studies, such as utilizing its specific cleavage activity.

[0005] The success of metagenomic next-generation sequencing (mNGS) technology in pathogen sequencing applications hinges on the quality of the sequencing library, including appropriate library size, purity, and a high percentage of effective target sequences. When constructing and sequencing libraries from samples with extremely low nucleic acid concentrations, the resulting libraries often contain numerous adapter self-ligations and background sequences—all invalid library sequences—leading to increased sequencing data and costs, while still failing to achieve satisfactory results. Extremely low DNA concentrations in clinical samples also frequently result in library construction failures. Currently, there are two main solutions to these problems associated with low-concentration nucleic acid library construction:

[0006] 1) Before library construction, whole genome amplification (WGA) is performed on low-concentration DNA nucleic acids to obtain sufficient DNA for library construction. Currently, this method is mainly used in single-cell DNA sequencing for scientific research. Other scientific research and special legal needs also use this method for amplification before library construction. However, the WGA method amplifies the entire low-concentration DNA, which often amplifies background nucleic acids in the sample and does not effectively improve the proportion of target data. In addition, the WGA process is relatively complex for clinical operation, the reagent cost is high, and different WGA methods have different effects. For these reasons, this method is not currently used for samples used in clinical diagnosis, especially for pathogen metagenomic detection.

[0007] 2) The main focus is on optimizing reagents and processes from sample extraction to nucleic acid and then to library construction. For example, selecting suitable micro-nucleic acid extraction kits to improve sample nucleic acid yield; and choosing library construction kits with better performance for low-concentration nucleic acid library construction to improve the success rate and quality of library construction. However, due to the large number of manufacturers and types of extraction and library construction kits currently available, it is difficult to establish unified experimental expectations, and this method inherently has very limited effectiveness in improving low-concentration nucleic acid library construction. Overall, there is still a lack of suitable solutions for low-concentration DNA library construction in clinical practice; there is no simple, economical, and effective solution. Summary of the Invention

[0008] This invention aims to solve the aforementioned technical problems by providing a method for constructing libraries for low-concentration nucleic acids. This invention improves the quality and success rate of library construction by adding specific synthetic DNA sequence fragments to low-concentration DNA, and then removing the synthetic sequences using a CRISPR / Cas9 system. This invention can effectively improve the success rate and quality of low-concentration DNA library construction.

[0009] The technical solution of the present invention is as follows:

[0010] A method for constructing a library for low-concentration nucleic acids involves adding a specific DNA fragment to a low-concentration DNA sample, constructing a library, and then using an sgRNA and CRISPR / Cas9 system designed for the specific DNA fragment to remove the specific DNA fragment from the library.

[0011] In the above-described technical solution of this invention, the specific DNA fragment can be an artificially synthesized fragment sequence; or a fragment product derived from known plasmids, other microorganisms, or the genomes of animals and plants, after PCR; its size ranges from 50bp to 500bp, and the amount added ranges from 0.1 to 10ng. Furthermore, the sequence of steps for adding the specific DNA fragment includes not only adding it after DNA fragmentation followed by library construction, but also adding the specific DNA fragment first, followed by DNA fragmentation and library construction.

[0012] This invention increases the initial nucleic acid concentration of the library by artificially adding specific nucleic acid fragments before library construction. This avoids problems such as high library construction failure rate, high adapter self-ligation rate, and high background nucleic acid sequence ratio caused by low concentration. The artificially added sequence is then cut using a CRISPR / Cas system to disrupt the library structure, remove the added sequence, and increase the detection rate of the target sequence.

[0013] This invention can improve the success rate and quality of library construction, and significantly reduce the proportion of linker self-ligation. It is an effective solution to the problems of poor quality and high failure rate in low-concentration nucleic acid library construction. Furthermore, this method is simple, fast, economical, and can achieve high throughput.

[0014] The method of this invention (SpikeCas library construction) can be applied to the construction of libraries starting with all low-concentration nucleic acids. In particular, it significantly improves the library quality and success rate for pathogen metagenomics of special sample types, and enhances the detection of target sequences.

[0015] The CRISPR / Cas system of this invention includes conventional Cas9 protein as well as various Cas9 proteins formed through other modification processes, including Cas12a, Cas12b and their various modified variants.

[0016] The CRISPR system in this invention is a complex formed by the combination of CRISPR / Cas protein and a library; the library is constructed based on a specific DNA fragment and an sgRNA designed for that specific DNA fragment.

[0017] The sgRNA also includes a combination of crRNA and tracrRNA corresponding to the Cas protein. The sgRNA can be used as a single or multiple sgRNAs depending on the Spike sequence.

[0018] As a preferred embodiment of the above technical solution, the specific DNA fragment is a synthetically produced fragment (Spike sequence), which is as follows:

[0019] GTATGATTTGATCGTCACAATGACATAATAGAGAGATTGATTTAGTGACTCGGACAATAAAATGCGTTGTGAGAGGTTAAGCAAGCA.

[0020] Design two primers for spCas9 sg RNA targeting this Spike DNA:

[0021] Primers for sgRNA1: taatacgactcactataggGAGAGATTGATTTAGTGACTgttttagagctagaaatagc. Primers for sgRNA2: taatacgactcactataggACAATAAAATGCGTTGTGAGgttttagagctagaaatagc.

[0022] The sgRNA was amplified by PCR using the Cas9-scaffold RV primer sequence AGCACCGACTCGGTGCCACT and the pSGKP vector as a template. After product recovery, sgRNA was obtained by in vitro transcription using T7 RNA transcriptase.

[0023] Another objective of this invention is to provide the application of the above-described library construction method in the preparation of diagnostic reagents.

[0024] In summary, the present invention has the following beneficial effects:

[0025] 1) This invention solves the problem of low initial concentration of library construction by first introducing specific DNA fragments, and then removes the introduced sequence using the CRISPR / Cas system after library construction, thereby increasing the proportion of target sequence data.

[0026] 2) The CRISPR / Cas system, excluding the input sequences, does not introduce additional background library sequences after library construction.

[0027] 3) The method of the present invention is simple to operate and effectively solves the problems of high failure rate, high adapter self-ligation rate and high background nucleic acid sequence ratio when constructing low-concentration DNA libraries; it is particularly suitable for low-concentration nucleic acid samples in pathogen metagenomic sequencing, and improves the detection of target sequences with limited sequencing data. Attached Figure Description

[0028] Figure 1 A flowchart illustrating the principle of low-concentration nucleic acid SpikeCas library construction;

[0029] Figure 2 This is a bar chart illustrating the sequencing results of low-concentration nucleic acid libraries constructed using the SpikeCas method.

[0030] Figure 3 Table of results for SpikeCas library construction experiments using low-concentration nucleic acids;

[0031] Figure 4 Table showing the results of low-concentration nucleic acid SpikeCas library construction for clinical samples. Detailed Implementation

[0032] The method of the present invention will be described in detail below with reference to specific embodiments. These embodiments should be understood as specific descriptions and expositions of the present invention, and should not be construed as limiting the scope of the present invention. Modifications or alterations made by those skilled in the art based on the present invention, and these equivalent modifications, also fall within the scope of the claims set forth in the claims of the present invention.

[0033] Example 1

[0034] Example of low-concentration nucleic acid SpikeCas library construction, the process is as follows: Figure 1 As shown.

[0035] In this embodiment, the artificially added DNA sequence is an 89bp Spike DNA sequence (referred to as Spike DNA throughout this document), which is as follows:

[0036] GTATGATTTGATCGTCACAATGACATAATAGAGAGATTGATTTAGTGACTCGGACAATAAAATGCGTTGTGAGAGGTTAAGCAAGCA.

[0037] Two spCas9 sg RNA primers were designed targeting this Spike DNA:

[0038] sgRNA1 primers: taatacgactcactataggGAGAGATTGATTTAGTGACTgttttagagctagaaatagc

[0039] sgRNA2 primers: taatacgactcactataggACAATAAAATGCGTTGTGAGgttttagagctagaaatagc.

[0040] The sgRNA was amplified by PCR using the pSGKP vector as a template and the Cas9-scaffold RV primer sequence AGCACCGACTCGGTGCCACT, respectively. After product recovery, the sgRNA was obtained by in vitro transcription using T7 RNA transcriptase.

[0041] sgRNA1:

[0042] GAGAGAUUGAUUUAGUGACUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAA GGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCCGGGCU.

[0043] sgRNA2:

[0044] ACAAUAAAAUGCGUUGUGAGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAA GGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCCGGGCU.

[0045] In this embodiment, we selected cultured Jurakat cells for nucleic acid extraction. The DNA nucleic acid was diluted to 0.01 ng as an extremely low concentration sample nucleic acid, then fragmented. Next, 0.5 ng of a synthetically produced 89 bp spike DNA sequence was added, followed by filler addition, adapter ligation, and normal library construction. The constructed library was then amplified by library PCR. Then, according to the following reaction system, the library was cut using sgRNA1, sgRNA2, or both sgRNA1 and sgRNA2. In this embodiment, we preferred to cut the library using both sgRNA1 and sgRNA2 simultaneously (i.e., SpikeCas library construction).

[0046] Cas9 cleavage library reaction system:

[0047]

[0048] Procedure: Incubate at 37℃ for 1 hour, then at 70℃ for 10 minutes, followed by library PCR amplification and library purification. Sequencing is performed using the Illumina sequencing platform.

[0049] The experimental results are shown in Figure 2 and Figure 3 .

[0050] Figure 1 and Figure 2 This indicates that direct library construction using DNA at extremely low concentrations leads to significant adapter self-ligation and increases in background microbial nucleic acid levels, while the number of specific target sequence reads remains unchanged. After library construction using the SpikeCas method, adapter self-ligation is largely eliminated, and the number of target pathogen reads is significantly increased.

[0051] Example 2

[0052] The Spike DNA sequence artificially added in this embodiment is the same as that in Example 1, and the sgRNA1 and sgRNA2 sequences used are also the same as those in Example 1. Furthermore, in this embodiment, a scheme that uses sgRNA1 and sgRNA2 simultaneously is preferred.

[0053] In this embodiment, two clinically derived cerebrospinal fluid samples showed extremely low nucleic acid concentrations after extraction. Sample 1-01D2135250N had a concentration of 0.011 ng / μL, and sample 2-01C2132042N had a concentration of 0.012 ng / μL, almost the lowest nucleic acid concentration detectable by qubit. Conventional library construction and SpikeCas library construction were performed on these two samples, respectively.

[0054] The SpikeCas library construction process involves fragmenting two samples, adding 1 ng of synthetically produced spike DNA sequence, followed by completion, A addition, and adapter ligation for normal library construction. The constructed library is then amplified by library PCR. Finally, the library is digested using sgRNA1 and sgRNA2 in the following reaction system.

[0055] Cas9 cleavage library reaction system:

[0056]

[0057] Procedure: Incubate at 37℃ for 1 hour, keep warm at 70℃ for 10 minutes, then perform library PCR amplification, followed by library purification, and sequencing using the Illumina sequencing platform.

[0058] The experimental results are shown in Figure 4 .

[0059] Figure 4 The results show that after SpikeCas library construction using low-concentration clinical sample DNA, compared to the normal direct library construction sample 1-2135250N cerebrospinal fluid sample which did not detect Mycobacterium tuberculosis complex (negative sample), the adapter self-ligation rate decreased from 77.3% to 6.3% after SpikeCas library construction, and Mycobacterium tuberculosis complex was successfully detected (RPM 3.81). For sample 2-2132042N, the adapter self-ligation rate decreased from 70.8% to 5.8%, and the RPM for detecting Mycobacterium tuberculosis complex increased from 3.12 to 81.33. This indicates that even with low nucleic acid concentrations in clinical samples, SpikeCas library construction, compared to the normal direct library construction method, can significantly reduce adapter self-ligation and, with limited data, significantly improve the detection of target sequences or target pathogen sequences.

Claims

1. A method for constructing a library for low-concentration nucleic acids, characterized in that: Low-concentration DNA is fragmented, and then a specific DNA fragment is added to it. A library is then constructed, and the specific DNA fragment is removed from the library using an sgRNA and CRISPR / Cas system designed for the specific DNA fragment. The specific DNA fragment has a size range of 50bp to 500bp and is an artificially synthesized fragment sequence or a fragment product derived from known plasmids, microorganisms, or plant and animal genomes after PCR. The Cas protein in the CRISPR / Cas system is selected from Cas9 protein or Cas12b protein.

2. The method for constructing a library for low-concentration nucleic acids according to claim 1, characterized in that: The Cas9 protein is a modified Cas9 variant protein, and the Cas12b protein is a modified Cas12b variant protein.

3. The application of the library construction method as described in any one of claims 1 to 2 in the preparation of diagnostic reagents.