Synthetic gene high-throughput library construction and sequencing verification method

The preparation of DNB nanospheres by stepwise PCR amplification and single-strand circularization using high-throughput sequencing technology solves the problem of high detection and screening costs in DNA chip synthesis, enabling low-cost detection and screening of DNA chip synthesis products and improving the success rate of accurate detection.

CN121992069APending Publication Date: 2026-05-08BGI TECH (CHANGZHOU) CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BGI TECH (CHANGZHOU) CO LTD
Filing Date
2024-11-05
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

In existing DNA chip synthesis technologies, the accuracy detection and screening of synthesized DNA fragments are costly, and traditional Sanger sequencing methods result in high error rates and high costs, making it difficult to achieve low-cost commercial applications.

Method used

High-throughput sequencing technology was used to prepare DNB nanospheres through stepwise PCR amplification and single-strand circularization. Library construction and sequencing verification were then performed using gene-specific primers, shortening the process and reducing costs.

Benefits of technology

This technology enables efficient detection and accurate screening of DNA chip synthesis products, reducing detection and screening costs, shortening the time required, and improving the success rate of accurate detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121992069A_ABST
    Figure CN121992069A_ABST
Patent Text Reader

Abstract

The invention provides an amplification library building method for high-throughput sequencing verification of synthetic genes. The amplification library building method comprises the following steps: obtaining a joint connection product of which two ends are connected with double-chain joints after a to-be-detected synthetic gene sequence, and carrying out first PCR (Polymerase Chain Reaction) amplification by adopting a first gene specific primer by taking the joint connection product as a template to obtain a first amplification product; taking the first amplification product as a template, and carrying out second PCR amplification by adopting a second gene specific primer to obtain a second amplification product; performing thermal denaturation single-chain separation and single-chain cyclization on the second amplification product to obtain single-chain circular DNA; the preparation method comprises the following steps: carrying out DNB preparation on single-stranded circular DNA to obtain DNB nanospheres; and sequencing the DNB nanospheres to obtain synthetic gene sequence information. The method verifies that the synthesized gene is high in flux, and the flux requirement of large-scale gene synthesis can be met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of biotechnology, specifically, to the field of high-throughput sequencing of synthetic genes, and more specifically, to a method for constructing and sequencing-verifying a high-throughput library of synthetic genes. Background Technology

[0002] With the development of genetic science and genetic engineering, DNA synthesis technology plays an increasingly important role in the life sciences. De novo synthesis of DNA elements, including regulatory sequences, whole genomes, artificial metabolic pathways, and even complete artificial genomes, will bring about tremendous changes to human life science research. Since the first synthesis of oligonucleotide chains in 1961 (Nirenberg et al. (1961) Proc. Natl Acad. Sci. USA 54:1588), DNA synthesis and assembly technologies have made great strides.

[0003] In 2004, Tian et al. successfully synthesized 292 different oligonucleotide chains on a DNA chip and assembled them into a 14.6 kb DNA fragment using oligonucleotide microarray chips, making it possible to synthesize and assemble large-scale, efficient, and low-cost DNA long fragments based on chip technology (Tian (2004) Nature 432:1050). In 2010, Kosuri et al. used Agilent's commercial DNA microarray chips to synthesize hundreds of genes (Kosuri et al. (2010) Nat. Biotech. 28:1295) and provided a complete and comprehensive technical route (Eroshenko et al. (2012) Curr. Protoc. in Chem. Biol. 4:1). Based on this related technology, GEN9 was established in the United States in 2012, becoming the first commercial company to provide DNA chip synthesis services globally. Its DNA synthesis products are priced at approximately US$0.26 per bp, lower than the market price of traditional DNA synthesis.

[0004] Two major challenges facing DNA chip synthesis technology are the high error rate of oligonucleotide chains and the impact of the high complexity of oligonucleotide libraries on assembly. Improving the accuracy of DNA chip synthesis and reducing assembly and screening costs are crucial for its commercial application. In existing DNA chip synthesis technologies, the correctness detection and screening of synthesized DNA fragments are often accomplished using first-generation sequencing technology—Sanger sequencing. Due to the inherent limitations of DNA chip synthesis, the assembled products are often complex. Compared to the designed sequence, each assembled DNA molecule may contain multiple highly complex variations (including single-base variations, nucleotide insertions, and deletions). The probability of single-base variations varies from 1‰ to 1% depending on the platform and design. Therefore, when the single-base error rate is 5‰, the probability of obtaining a completely correct DNA fragment in one attempt when assembling a 750bp DNA fragment is only 2.33%, the probability of one or fewer errors is 11.11%, and the probability of two or fewer errors is 27.63%. To select the correct molecules from complex DNA assembly products whose actual sequences perfectly match the designed sequences, and to ensure a 90% success rate, 98 single clones need to be selected for Sanger sequencing. When errors exceed three, the costs of error correction materials and time become prohibitive, and DNA chip synthesis technology loses its cost advantage. Therefore, traditional Sanger sequencing-based methods for detecting the correctness of DNA chip synthesis products are expensive, representing the main source of cost for DNA chip synthesis technology and a major bottleneck for further cost reduction. Summary of the Invention

[0005] This invention provides a method for amplification and library preparation for high-throughput sequencing verification of synthetic genes. This method successfully applies high-throughput sequencing technology to the detection of DNA chip synthetic products through a library construction strategy, which greatly reduces the cost of product detection and correct product screening, thereby significantly reducing the cost of DNA chip synthesis.

[0006] This invention is achieved through the following technical solution:

[0007] A high-throughput detection method for DNA synthesis products includes the following steps:

[0008] After obtaining the synthetic gene sequence to be tested, a linker ligation product with double linkers attached to both ends is obtained. The linker ligation product includes optional target and non-target regions.

[0009] Using the adapter ligation product as a template, a first PCR amplification is performed using a first gene-specific primer to obtain a first amplification product, wherein the first gene-specific primer binds to the target region; and

[0010] Using the first amplification product as a template, a second PCR amplification is performed using a second gene-specific primer to obtain a second amplification product, wherein the second gene-specific primer binds to the target region.

[0011] The second amplification product was subjected to thermal denaturation to separate the single strands and then circularize them to obtain single-stranded circular DNA.

[0012] Single-stranded circular DNA was processed into DNB nanospheres.

[0013] The DNB nanospheres were sequenced to obtain the synthetic gene sequence information.

[0014] As a preferred embodiment of the present invention, the target region includes multiple target genes, the first gene-specific primer includes multiple primers that bind to the multiple target genes respectively; the second gene-specific primer includes multiple primers, which are nested primers inside the first gene-specific primers, and bind to the multiple target genes respectively.

[0015] As a preferred embodiment of the present invention, the second gene-specific primer includes a portion that binds to the target region and a portion located at the 5' end that is the same as or partially the same as the second strand of the double-linked head.

[0016] As a preferred embodiment of the present invention, the number of cycles for the first PCR amplification is 10 to 30 cycles, preferably 20 cycles; the number of cycles for the second PCR amplification is 30 to 40 cycles, preferably 35 cycles.

[0017] As a preferred embodiment of the present invention, the target region includes one or more of the following: a vector homology arm base sequence, a base sequence with an index tag, and a single-stranded circular primer-matched base sequence.

[0018] As a preferred embodiment of the present invention, the first PCR amplification is one or more of forward amplification and reverse amplification.

[0019] As a preferred embodiment of the present invention, the first gene-specific primer and the second gene-specific primer are shown in SEQ ID NO:1-6.

[0020] As a preferred embodiment of the present invention, the first gene-specific primer is shown in SEQ ID NO:1-4.

[0021] As a preferred embodiment of the present invention, the second gene-specific primer is shown in SEQ ID NO:5-6.

[0022] The method of this invention solves the problem of batch synthesizing genes, cloning them together, obtaining sequencing libraries through amplification library construction technology, and performing high-throughput sequencing verification. The amplification library construction and sequencing process is nearly 40% faster than the existing high-throughput library construction process, and the cost of amplification library construction can be reduced by nearly 30% compared with the existing methods. Attached Figure Description

[0023] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0024] Figure 1 This is a schematic diagram of the process for forward amplification and detection of DNA chip synthesis products using a high-throughput sequencing strategy according to the present invention.

[0025] Figure 2 This is a schematic diagram of the process for detecting DNA chip synthesis products using a high-throughput sequencing strategy for forward and reverse amplification according to the present invention.

[0026] Figure 3 This is a gel electrophoresis image of the first PCR amplification result in an embodiment of the present invention.

[0027] Figure 4 This is a gel electrophoresis image of the second PCR amplification result in an embodiment of the present invention. Detailed Implementation

[0028] The present invention will be further explained and described below with reference to specific embodiments. Unless otherwise specified, the experimental methods used in the following embodiments are conventional methods. Unless otherwise specified, the materials and reagents used in the following embodiments are commercially available.

[0029] 1. Gene-specific primer design

[0030] 1) Based on the same vector used for cloning and ligating the synthesized genes in batches, gene-specific primers are designed according to the vector homologous sequence to amplify the synthesized genes in batches carrying the vector homologous sequence. The amplification product is a mixture of PCR products of all synthesized genes.

[0031] 2) The primers for amplification also include index sequences to distinguish different clone numbers of the synthesized gene; the index sequence is represented by "NNNNNNNNNN".

[0032] 3) The amplification primers also include base sequences that match the single-stranded circularization primers, which facilitates binding with the single-stranded circularization primers and promotes the transformation of single-stranded linear DNA into single-stranded circular DNA.

[0033] The primer design information is as follows:

[0034] primer_id sequence length serial number 1_Ad153_PCR_1F GAACGACATGGCTACGATCCGACTTGTGCCAATTGTCAAGCTAAGTTCAG 50 SEQ ID NO:1 1_Ad153_PCR_1R TTGTCTTCCTAAGACCGCTTGGCCTCCGACTGATCAGGTGCCAATGTTCAGTCTAC 56 SEQ ID NO:2 2_Ad153_PCR_1F GAACGACATGGCTACGATCCGACTTGATCAGGTGCCAATGTTCAGTCTAC 50 SEQ ID NO:3 2_Ad153_PCR_1R TTGTCTTCCTAAGACCGCTTGGCCTCCGACTGTGCCAATTGTCAAGCTAAGTTCAG 56 SEQ ID NO:4 Ad153_PCR_2F / 5Phos / GAACGACATGGCTACGATCCGAC 23 SEQ ID NO:5 Ad153_PCR_x_2R TGTGAGCCAAGGAGTTGNNNNNNNNNNTTGTCTTCCTAAGACCGCTTGGC 50 SEQ ID NO:6

[0035] 2. Library construction using amplification method

[0036] The gene was amplified and synthesized using a stepwise PCR amplification method.

[0037] 1) First PCR amplification system and reaction conditions

[0038]

[0039] The PCR product was not purified after the first step and was detected by 1.5% agarose gel electrophoresis. The results are as follows: Figure 3 As shown, the first step of amplification successfully amplified the synthetic gene.

[0040] 2) Second PCR amplification system and reaction conditions

[0041] 2 μL of the first PCR amplification product was used as a template for the second PCR amplification.

[0042]

[0043] After the second step of PCR products was completed, they were detected by 1.5% agarose gel electrophoresis. The results are as follows: Figure 4 The results showed that the second PCR amplification products were successfully amplified, and the amplification products were also recovered by gel cutting.

[0044] 3. PCR product cyclization and DNB preparation

[0045] Circulation was performed using the MGI circularization kit, and the obtained single-circular DNA (ssDNA) passed qubit concentration quantification. Simultaneously, ssDNA was used to prepare DNB, and the prepared DNB also passed qubit concentration quantification. The data results for single-strand circularization and DNB preparation are as follows:

[0046]

[0047] 4. Sequencing

[0048] High-throughput sequencing was performed using the MGI-G99 sequencer, sequencing type SE400. The sequencing data volume met the analysis requirements. The SE400 sequencing data report is shown in Table 1.

[0049] Table 1: SE400 Sequencing Data Download Report

[0050]

[0051]

[0052] Results of this high-throughput sequencing data: Total Reads (M) is 114.56 Mb, Q30 (%) is 71.93%, SplitRate (%) is 98.28%; after sequencing analysis, the gene accuracy is 97.8% (980 / 1002). Note: Gene accuracy = number of genes with at least one correct clone / total number of genes validated in this batch.

[0053] 5. High-throughput sequencing analysis of forward and reverse amplification products

[0054] A subset of genes was sampled and re-validated using Sanger sequencing. The results of the high-throughput sequencing analysis and the Sanger sequencing analysis are shown in Tables 2 and 3.

[0055] The high-throughput sequencing analysis results of the forward amplification products were consistent with those of the reverse amplification products. A subset of genes were sampled for further Sanger sequencing validation, and the results showed consistency between the high-throughput sequencing analysis and the Sanger sequencing analysis. This demonstrates that the amplification-based high-throughput sequencing technology workflow is entirely feasible for gene synthesis validation.

[0056] Table 2: High-throughput sequencing analysis results of forward amplification products

[0057]

[0058]

[0059] Table 3: High-throughput sequencing analysis results of reverse amplification products

[0060]

[0061]

[0062]

[0063] Tables 2 and 3 present the high-throughput sequencing analysis results of the forward and reverse amplification products. It can be seen that the results are consistent with those of the reverse amplification products. A sample of genes was sampled for further Sanger sequencing validation, and the analysis showed that the high-throughput sequencing results were consistent with the Sanger sequencing results. This demonstrates that the high-throughput sequencing technology based on amplification is entirely feasible for gene synthesis validation.

[0064] In summary, based on the experimental results above, it can be concluded that the amplification method for library construction and sequencing of synthetic genes solves the problem that some application scenarios cannot accommodate the mixing of all synthetic genes for library construction and sequencing; further shortens the overall high-throughput sequencing verification time; and reduces the overall cost of high-throughput sequencing verification.

[0065] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0066] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for amplification and library construction of synthetic genes validated by high-throughput sequencing, characterized in that, include: After obtaining the synthetic gene sequence to be tested, a linker ligation product with double linkers attached to both ends is obtained. The linker ligation product includes optional target and non-target regions. Using the adapter ligation product as a template, a first PCR amplification is performed using a first gene-specific primer to obtain a first amplification product, wherein the first gene-specific primer binds to the target region. and Using the first amplification product as a template, a second PCR amplification is performed using a second gene-specific primer to obtain a second amplification product, wherein the second gene-specific primer binds to the target region. The second amplification product was subjected to thermal denaturation to separate the single strands and then circularize them to obtain single-stranded circular DNA. Single-stranded circular DNA was processed into DNB nanospheres. The DNB nanospheres were sequenced to obtain the synthetic gene sequence information.

2. The method according to claim 1, characterized in that, The target region includes multiple target genes. The first gene-specific primer includes multiple primers that bind to the multiple target genes respectively. The second gene-specific primer includes multiple primers, which are nested primers inside the first gene-specific primer, and bind to the multiple target genes respectively.

3. The method according to claim 1, characterized in that, The second gene-specific primer includes a portion that binds to the target region and a portion located at the 5' end that is identical or partially identical to the second strand of the double-linked head.

4. The method according to claim 1, characterized in that, The number of cycles for the first PCR amplification is 10 to 30, preferably 20; the number of cycles for the second PCR amplification is 30 to 40, preferably 35.

5. The method according to claim 1, characterized in that, The target region includes one or more of the following: a vector homology arm base sequence, a base sequence with an index tag, and a single-stranded circular primer-matched base sequence.

6. The method according to claim 1, characterized in that, The first PCR amplification is one or more of forward amplification and reverse amplification.

7. The method according to claim 1, characterized in that, The first gene-specific primer and the second gene-specific primer are shown in SEQ ID NO:1-6.

8. The method according to claim 7, characterized in that, The first gene-specific primer is shown in SEQ ID NO:1-4.

9. The method according to claim 7, characterized in that, The second gene-specific primer is shown in SEQ ID NO:5-6.