Gene sequencing method and kit
By optimizing sequencing reagents and sequencing protocols, and employing fluorescent labeling and reversible blocking groups, the problems of short read lengths and high error rates in high-throughput sequencing have been solved, resulting in longer read lengths and higher quality gene sequencing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN SALUS BIOMED CO LTD
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-12
Smart Images

Figure CN122012683A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of gene sequencing technology, and in particular to a gene sequencing method and kit. Background Technology
[0002] High-throughput sequencing (HTS), also known as next-generation sequencing (NGS), is a technology capable of rapidly sequencing millions to billions of DNA molecules in parallel. Compared to traditional Sanger sequencing, its core advantages lie in its high throughput, low cost, and high speed, fundamentally changing genomics research and transforming large-scale sequencing projects, such as the Human Genome Project, from time-consuming and costly endeavors into tasks that can be accomplished in regular laboratories. The entire high-throughput sequencing workflow includes library preparation, cluster generation or DNA nanosphere (DNB) preparation and loading, and sequencing reaction.
[0003] Currently, the mainstream platform technologies in the market include sequencing-by-synthesis technology from companies such as Illumina and Silergy, combined probe-anchored polymerization sequencing technology from BGI Genomics, and semiconductor sequencing technology from Thermo Fisher Scientific. However, current mainstream next-generation sequencing platforms still suffer from relatively short read lengths, and most still have error rates above 0.1%.
[0004] Therefore, how to reduce the error rate of high-throughput sequencing and effectively improve sequencing quality remains a key research focus and challenge in the field of high-throughput sequencing technology. Summary of the Invention
[0005] The purpose of this application is to provide an improved gene sequencing method and a kit for use with this sequencing method.
[0006] To achieve the above objectives, this application adopts the following technical solution:
[0007] The first aspect of this application discloses a gene sequencing method, comprising sequencing a nucleic acid molecule to be tested using sequencing reagent 1 and sequencing reagent 2; wherein, sequencing reagent 1 comprises 4 second nucleotide analogs and 4 first nucleotide analogs, and sequencing reagent 2 comprises 4 first nucleotide analogs; the first nucleotide analogs have the structure shown in Formula 1, and the second nucleotide analogs have the structure shown in Formula 2;
[0008] Formula 1 ;
[0009] Formula 2 ;
[0010] In Formulas 1 and 2, the "Base" of the four second nucleotide analogs and the four first nucleotide analogs are adenine A, cytosine C, guanine G, and thymine T, respectively; in Formula 1, R1 is a reversible blocking group; in Formula 2, "Dye" is a fluorescent group, "Linker" is a linker, and "X" is independently selected from any one of CH2, NH, S, CF2, CBr2, and Se, and the four second nucleotide analogs are labeled with four different fluorescent groups; sequencing of the nucleic acid molecule to be tested includes sequencing after the sequencing primers are hybridized to the nucleic acid molecule to be tested, using at least one of the following schemes:
[0011] Scheme 1: (1) Contact with sequencing reagent 1, and under the action of metal ions and polymerase, the second nucleotide analog forms a complex with the nucleic acid molecule to be tested, the 3' end of the sequencing primer, or the 3' end of the extended sequencing primer; (2) Contact with imaging buffer to wash away unbound second nucleotide analog; (3) Signal acquisition; (4) After signal acquisition, remove the bound second nucleotide analog; (5) Contact with sequencing reagent 2, and under the action of polymerase, the first nucleotide analog polymerizes to the 3' end of the sequencing primer or the 3' end of the extended sequencing primer; (6) Remove the reversible blocking group of the first nucleotide analog; Repeat steps (1) to (6) to complete the sequencing of the nucleic acid molecule to be tested;
[0012] Scheme 2: (1) Contact with sequencing reagent 2, and under the action of polymerase, the first nucleotide analog polymerizes to the 3' end of the sequencing primer or the 3' end of the sequencing primer extension; (2) Contact with sequencing reagent 1, and under the action of metal ions and polymerase, the second nucleotide analog forms a complex with the nucleic acid molecule to be tested and the 3' end of the sequencing primer extension; (3) Contact with imaging buffer to wash away unbound second nucleotide analog; (4) Signal acquisition; (5) After signal acquisition, remove the bound second nucleotide analog; (6) Remove the reversible blocking group of the first nucleotide analog; Repeat steps (1) to (6) to complete the sequencing of the nucleic acid molecule to be tested.
[0013] In both Scheme 1 and Scheme 2, signal acquisition mainly refers to acquiring the signal of the fluorescent group of the second nucleotide analog in the complex, so that the base type of the second nucleotide analog can be determined based on the fluorescent group signal during subsequent analysis, and then the base type at the corresponding position of the nucleic acid molecule to be tested can be determined.
[0014] Because the excitation efficiency of the laser on the dye varies, and the light-gathering range of the filter may deviate from the wavelength of the dye, there may be a large difference in the signal of the four bases. Therefore, this application uses four second nucleotide analogs and four first nucleotide analogs in sequencing reagent 1. By adding the first type of nucleotides, the signal values of the four bases are adjusted to obtain better imaging results, which facilitates the base recognition of the subsequent algorithm and improves the quality.
[0015] It should be noted that this application uses optimized sequencing reagent 1 and sequencing reagent 2 for sequencing to avoid residual scarring after excision. Simultaneously, the optimized sequencing protocol enables longer read lengths with higher quality and a lower error rate. Compared to existing sequencing-by-synthesis (SBS) routes, both schemes in this application begin signal acquisition immediately upon contact with sequencing reagent 1. No synthesis reaction occurs at this step; that is, fluorescent dNTPs do not extend to form complexes. After imaging, excision is not required; elution can be performed directly with high-salt buffer.
[0016] It should also be noted that the key to this application lies in the optimized sequencing reagent 1, sequencing reagent 2, and sequencing scheme; as for reversible blocking groups, fluorescent groups, linkers, etc., existing technologies can be referenced, and no specific limitations are made here.
[0017] In one implementation of this application, step (1) of Scheme 1 further includes contacting with a cleaning reagent to provide a reaction environment for the formation of the complex in step (1).
[0018] In one implementation of this application, step (5) of Scheme 1 further includes contacting with a cleaning reagent to provide a reaction environment for the polymerization of the first nucleotide analog in step (5).
[0019] In one implementation of this application, step (6) of Scheme 1 further includes contacting with a cleaning reagent to wash away the unpolymerized first nucleotide analog and providing a reaction environment for removing the reversible blocking group of the first nucleotide analog.
[0020] In one implementation of this application, step (1) of Scheme 2 further includes contacting with a cleaning reagent to provide a reaction environment for the polymerization of the first nucleotide analog in step (1).
[0021] In one implementation of this application, step (2) of scheme two further includes contacting with a cleaning reagent to wash away the unpolymerized first nucleotide analog and to provide a reaction environment for the formation of the complex in step (2).
[0022] In one implementation of this application, step (6) of scheme two further includes contacting the cleaning reagent first to provide a reaction environment for removing the reversible blocking group of the first nucleotide analog.
[0023] The second aspect of this application discloses a gene sequencing kit, which includes sequencing reagent 1 and sequencing reagent 2; sequencing reagent 1 includes 4 second nucleotide analogs and 4 first nucleotide analogs, and sequencing reagent 2 includes 4 first nucleotide analogs; the first nucleotide analogs have the structure shown in Formula 1, and the second nucleotide analogs have the structure shown in Formula 2.
[0024] It should be noted that the gene sequencing kit of this application is actually a combination of sequencing reagent 1 and sequencing reagent 2 used in the gene sequencing method of this application, which is convenient to use; therefore, the specific limitations of sequencing reagent 1 and sequencing reagent 2 in the kit of this application, such as the limitations of each group in Formula 1 and Formula 2, can be referred to the gene sequencing method of this application, and will not be repeated here.
[0025] In one implementation of this application, the gene sequencing kit further includes at least one of imaging buffer, excision reagent, cleaning reagent 1, and cleaning reagent 2; the imaging buffer includes 10-100 mM Tris-HCl, 10-50 mM NaCl, 1-2 M Betaine, 0.02-0.5% Tween-20, 0.1-1 mM EDTA, 1-70 mM non-catalytic metal ions, and 5-70 mM antioxidant; the excision reagent includes 10-50 mM tris(3-hydroxypropyl)phosphine, 0.2-2 M NaCl, 10-100 mM Tris-HCl, and 0.02-0.5% Tween-20; cleaning reagent 1 includes 20-100 mM NaCl and 25-100 mM C6H5Na3O7·2H2O; and cleaning reagent 2 includes 10-100 mM Tris-HCl, 1-5 mM NaCl, and 0.1-1 mM EDTA.
[0026] Among them, the excision reagent is used to excise the reversible blocking group; the cleaning reagent 1 and the cleaning reagent 2 are used for cleaning in different steps. For example, the cleaning reagent 2 is used in steps (1) and (5) of scheme 1, and steps (1) and (2) of scheme 2. The cleaning reagent 1 is used in the remaining steps. After the signal acquisition is completed, the removal of the bound second nucleotide analog is also done using the cleaning reagent 1.
[0027] In one implementation of this application, the non-catalytic metal ion includes at least one of strontium chloride, calcium chloride, and barium chloride.
[0028] In one implementation of this application, the antioxidant includes at least one of sodium ascorbate, Trolox, and glutathione.
[0029] In one implementation of this application, sequencing reagent 1 is prepared by removing the antioxidant from the imaging buffer and adding polymerase, four second nucleotide analogs and four first nucleotide analogs.
[0030] In one implementation of this application, the four second nucleotide analogs, or "Base," of the sequencing reagent 1 are second nucleotide analogs of adenine A, cytosine C, guanine G, and thymine T, respectively, and are labeled as second nucleotide analog A, second nucleotide analog T, second nucleotide analog C, and second nucleotide analog G, wherein the concentration of second nucleotide analog A is 0.02-0.05 μM, the concentration of second nucleotide analog T is 0.5-1 μM, the concentration of second nucleotide analog C is 0.5-1 μM, and the concentration of second nucleotide analog G is 0.02-0.05 μM; the four first nucleotide analogs, or "Base," are first nucleotide analogs of adenine A, cytosine C, guanine G, and thymine T, respectively, and are labeled as first nucleotide analog A, first nucleotide analog T, first nucleotide analog C, and first nucleotide analog G, wherein the concentration of first nucleotide analog A is 0.01-0.02 μM, and the concentration of first nucleotide analog T is 0.003-0.005 μM. μM, the concentration of the first nucleotide analog C is 0.002-0.005 μM, and the concentration of the first nucleotide analog G is 0.008-0.01 μM.
[0031] In one implementation of this application, the sequencing reagent 2 includes 10-50 mM Tris-HCl, 5-50 mM NaCl, 2-20 mM (NH4)2SO4, 0.01-0.1 mg / mL polymerase, 2-20 mM MgSO4, 0.2-1 mM EDTA, and four first nucleotide analogs, wherein each of the first nucleotide analogs A, T, C, and G is 0.2-1 μM.
[0032] It should be noted that in the nucleotide analogs shown in Formulas 1 and 2 of this application, the reversible blocking group, linker, and fluorescent group can all refer to existing technologies. For example, the optional reversible blocking group R1 in Formula 1 includes, but is not limited to, at least one of methylene azide, allyl, ester, phosphate, hydroxylamine, disulfide bond, photocleavable group, 2-nitrobenzyl, and azo compound. The linker can be a cleavable linker or a non-cleavable linker. The fluorescent group labels of the four second nucleotide analogs are non-repeatingly selected from existing known fluorescent groups.
[0033] In one implementation of this application, the reversible blocking group includes at least one of the following groups.
[0034] .
[0035] In one implementation of this application, the cleavable linker includes at least one of the following: electrophilic cleavage linker group, nucleophilic cleavage linker group, photolytic linker group, cleavage group under reducing conditions, cleavage group under oxidizing conditions, safety handle type linker group, and group cleaved by elimination mechanism.
[0036] In one implementation of this application, the cleavable linker is at least one selected from alkyl, allyl, azidomethylene, 2-nitrobenzyl, and dithio.
[0037] In one implementation of this application, the non-cleavable linker includes at least one of a polyethylene glycol chain and a polypeptide chain.
[0038] In one implementation of this application, the fluorescent groups of the four second nucleotide analogs are non-repeatingly selected from AF488, AF532, AF633, AF680, AF660, AF700, AF647, AF594, AF555, AF568, CY3, CY5, CY5.5, CY7, CY7.5, ROX, R6G, ATTO495, ATTO532, ATTO700, ATTO680, ATTO655, ATTO647N, ATTO594, ATTO Rho101, ATTO590, ATTO Thio12, FAM, VIC, TET, JOE, HEX, CAL Fluor Orange560, TAMRA, CAL Fluor Red610, TEXAS RED, and CAL Fluor Red635, iFluor488, iFluor514, iFluor532, iFluor546, iFluor555, iFluor568, iFluor590, i Fluor610, iFluor633, iFluor647, iFluor680, iFluor700, iFluor710, Quasar705, Quasar670.
[0039] Due to the adoption of the above technical solutions, the beneficial effects of this application are as follows:
[0040] The gene sequencing method of this application, using optimized sequencing reagent 1 and sequencing reagent 2, as well as an optimized sequencing scheme, can achieve sequencing with longer read lengths, higher quality, and lower error rate, providing a new method and approach for high-throughput sequencing. Attached Figure Description
[0041] Figure 1 This is a technical roadmap of the gene sequencing method scheme one in the embodiments of this application;
[0042] Figure 2This is a technical roadmap of the second gene sequencing method in the embodiments of this application;
[0043] Figure 3 This is the ATCG image of dNTP-C2-H, i.e., sequencing reagent 1, in the embodiments of this application;
[0044] Figure 4 These are the ATCG images of dNTP-N3-H, i.e., sequencing reagents 1-2, in the embodiments of this application;
[0045] Figure 5 This is a comparison of the Q30 curves of sequencing reagents 1 and 1-2 using dNTPs in the embodiments of this application.
[0046] Figure 6 This is a Binding image of metal ion group 3 used in the embodiments of this application;
[0047] Figure 7 This is a Binding image of metal ion group 4 used in the embodiments of this application;
[0048] Figure 8 This is a Binding image of metal ion group 5 used in the embodiments of this application;
[0049] Figure 9 This is a graph showing the variation in the number of SE50 clusters of some metal ions in the embodiments of this application;
[0050] Figure 10 This is a graph showing the changes in SE50 signal intensity of some metal ions in the embodiments of this application;
[0051] Figure 11 This is the result of comparing the image quality (C bases) of four groups of buffer solution experiments in the embodiments of this application;
[0052] Figure 12 These are the sequencing signal value curves of four sets of buffer solutions in the embodiments of this application;
[0053] Figure 13 These are the sequencing Q-value curves of four sets of buffer solutions in the embodiments of this application;
[0054] Figure 14 These are the result images of polymerase group 1 and group 2 used in the embodiments of this application;
[0055] Figure 15 This is the result of comparing the signal intensity of polymerase groups 3, 4 and 5 in the embodiments of this application;
[0056] Figure 16 The results are the Q30 curve comparison results of polymerase groups 3, 4 and 5 in the embodiments of this application. Detailed Implementation
[0057] Current mainstream next-generation sequencing platforms still suffer from short read lengths and error rates mostly exceeding 0.1%. This application addresses this issue by synthesizing fluorescently modified dNTPs to avoid scarring after excision, and employing an optimized sequencing protocol system to achieve longer read lengths, higher quality, and lower error rates.
[0058] This application employs an improved edge-to-edge sequencing method, the main technical route of which is as follows: Figure 1 and Figure 2 As shown. Figure 1 As the sequencing method of this application, after the sequencing primers hybridize to the nucleic acid molecule to be tested, the main cyclic steps of the first method include: (1) contacting sequencing reagent 1, under the action of metal ions and polymerase, the second nucleotide analog forms a complex with the nucleic acid molecule to be tested, the 3' end of the sequencing primer or the 3' end of the extended sequencing primer; (2) contacting imaging buffer to wash away unbound second nucleotide analogs; (3) signal acquisition; (4) after signal acquisition, removing the bound second nucleotide analogs; (5) contacting sequencing reagent 2, under the action of polymerase, the first nucleotide analog polymerizes to the 3' end of the sequencing primer or the 3' end of the extended sequencing primer; (6) removing the reversible blocking group of the first nucleotide analog; repeating steps (1) to (6) to complete the sequencing of the nucleic acid molecule to be tested.
[0059] Furthermore, step (1) of Scheme 1 also includes contacting the cleaning reagent first to provide a reaction environment for the formation of the complex in step (1). For example, cleaning reagent 2 is used first for cleaning, and then sequencing reagent 1 is added. Furthermore, step (5) of Scheme 1 also includes contacting the cleaning reagent first to provide a reaction environment for the polymerization of the first nucleotide analog in step (5). For example, after signal acquisition is completed, cleaning reagent 1 is used first for cleaning, then cleaning reagent 2 is used, and then sequencing reagent 2 is added. Furthermore, step (6) of Scheme 1 also includes contacting the cleaning reagent first to wash away the unpolymerized first nucleotide analog and to provide a reaction environment for the removal of the reversible blocking group of the first nucleotide analog. For example, after the reaction of sequencing reagent 2 is completed, cleaning reagent 1 is added for cleaning, and then the reversible blocking group removal reaction is performed.
[0060] Figure 2Scheme 2 of the sequencing method of this application, after the sequencing primers hybridize to the nucleic acid molecule to be tested, the main cyclic steps of Scheme 2 include: (1) contacting sequencing reagent 2, under the action of polymerase, the first nucleotide analog polymerizes to the 3' end of the sequencing primer or the 3' end of the extended sequencing primer; (2) contacting sequencing reagent 1, under the action of metal ions and polymerase, the second nucleotide analog forms a complex with the nucleic acid molecule to be tested and the 3' end of the extended sequencing primer; (3) contacting imaging buffer to wash away the unbound second nucleotide analog; (4) signal acquisition; (5) after the signal acquisition is completed, the bound second nucleotide analog is removed; (6) the reversible blocking group of the first nucleotide analog is removed; repeat steps (1) to (6) to complete the sequencing of the nucleic acid molecule to be tested.
[0061] Furthermore, step (1) of scheme two also includes contacting the cleaning reagent first to provide a reaction environment for the polymerization of the first nucleotide analog in step (1). For example, cleaning reagent 2 is used first for cleaning, and then sequencing reagent 2 is added for reaction. Furthermore, step (2) of scheme two also includes contacting the cleaning reagent first to wash away the unpolymerized first nucleotide analog and to provide a reaction environment for the formation of the complex in step (2). For example, after the reaction of sequencing reagent 2 is completed, cleaning reagent 2 is used first for cleaning, and then sequencing reagent 1 is added. Furthermore, step (6) of scheme two also includes contacting the cleaning reagent first to provide a reaction environment for removing the reversible blocking group of the first nucleotide analog. For example, after signal acquisition is completed, cleaning reagent 2 is used first for cleaning, then cleaning reagent 1 is used for cleaning, and then the reversible blocking group removal reaction is performed.
[0062] The present application will be further described in detail below through specific embodiments. The following embodiments are only for further illustration of the present application and should not be construed as limiting the present application.
[0063] Unless otherwise specified, the technical means used in the embodiments are conventional means well known to those skilled in the art.
[0064] Example 1
[0065] This example demonstrates the design of different nucleotide analogs and compares their quality in binding sequencing. Specifically, it includes:
[0066] 1. Library: Ecoli standard library (item number SL-Q00100), with inserted fragments in the range of 450-550 bp.
[0067] 2. Sequencing was performed using the Saluseq Nimbo platform.
[0068] 1) Sequencing reagent preparation:
[0069] This example designs two sets of first nucleotide analogs and second nucleotide analogs. The first set includes four first nucleotide analogs and four second nucleotide analogs. The bases of the four first nucleotide analogs are adenine (A), cytosine (C), guanine (G), and thymine (T), respectively. The reversible blocking groups are all azides, referred to as dNTP-N3, and the structures are shown in Formula 1. The base labels of the four second nucleotide analogs are linked to the bases of the four second nucleotide analogs via linkers: dye-A is linked to adenine (A), dye-B to cytosine (C), dye-C to guanine (G), and dye-D to thymine (T). The four second nucleotide analogs have some modifications at the R position, referred to as dNTP-C2-H, and the structures are shown in Formula 2. The second set contains four primary nucleotide analogs and four secondary nucleotide analogs. The bases of the four primary nucleotide analogs are adenine (A), cytosine (C), guanine (G), and thymine (T), respectively. All four contain azide reversible blocking groups, referred to as dNTP-N3, with structures shown in Formula 1. The base labels for the four secondary nucleotide analogs are linked to the bases of the four secondary nucleotide analogs via linkers: dye-A linked to adenine (A), dye-B linked to cytosine (C), dye-C linked to guanine (G), and dye-D linked to thymine (T), referred to as dNTP-N3-H, with structures shown in Formula 3. It should be noted that the primary nucleotide analogs in the first and second sets have identical structures, being dNTP-N3 with azide reversible blocking groups.
[0070] Formula 1 ;
[0071] Formula 2 ;
[0072] Formula 3
[0073] In Formulas 1, 2, and 3, the "Base" of the four second nucleotide analogs and the four first nucleotide analogs are adenine A, cytosine C, guanine G, and thymine T, respectively; in Formula 1, R1 is azide; in Formulas 2 and 3, "Dye" is a fluorescent group, "Linker" is a linker, and the four second nucleotide analogs are labeled with four different fluorescent groups; in Formula 2, "X" is independently selected from any one of CH2, NH, S, CF2, CBr2, and Se, specifically CH2 in this example. The linker can be a cleavable linker or a non-cleavable linker. A cleavable linker includes at least one of the following: electrophilic cleavable linker group, nucleophilic cleavable linker group, photolytic linker group, cleavable linker group under reducing conditions, cleavable linker group under oxidizing conditions, safety handle type linker group, and group cleaved by elimination mechanism. Specifically, the cleavable linker is selected from at least one of alkyl, allyl, azidomethylene, 2-nitrobenzyl, and dithio. A non-cleavable linker includes at least one of polyethylene glycol chain and polypeptide chain. In this example, it is specifically azidomethylene.
[0074] Compound 1 was purchased from Saluseq, catalog numbers SL-C088, SL-C089, SL-C090, and SL-C091; compound 2 was purchased from Saluseq, catalog numbers SL-C150, SL-C151, SL-C152, and SL-C153; compound 3 was purchased from Saluseq, from the Saluseq Nimbo sequencing kit, catalog number SL-S00333.
[0075] Imaging buffer: 50 mM Tris-HCl (25℃, pH 8.8), 25 mM NaCl, 2 M Betaine, 0.05% Tween-20, 0.5 mM EDTA, 25 mM Strontium chloride, 20 mM sodium ascorbate.
[0076] Sequencing reagent 1 containing dNTP-C2-H and sequencing reagents 1-2 containing dNTP-N3-H were prepared separately.
[0077] Sequencing Reagent 1: Based on the imaging buffer, remove sodium ascorbate and add 0.04 mg / mL 9°N polymerase (produced by Shenzhen Sailu Medical) and the above-mentioned first set of 4 second nucleotide analogs and 4 first nucleotide analogs mixture, wherein the concentration of second nucleotide analog A is 0.05 μM, the concentration of second nucleotide analog T is 1 μM, the concentration of second nucleotide analog C is 1 μM, the concentration of second nucleotide analog G is 0.025 μM, the concentration of first nucleotide analog A is 0.0167 μM, the concentration of first nucleotide analog T is 0.003 μM, the concentration of first nucleotide analog C is 0.005 μM, and the concentration of first nucleotide analog G is 0.0083 μM.
[0078] Sequencing reagents 1-2: Based on the imaging buffer, sodium ascorbate was removed, and 0.04 mg / mL 9°N polymerase (produced by Shenzhen Sailu Medical) and the second set of 4 second nucleotide analogs and 4 first nucleotide analogs prepared above were added. The concentrations of second nucleotide analog A, T, C and G were 0.025 μM, and the concentrations of first nucleotide analog A, T, C and G were 0.003 μM, 0.005 μM, and 0.0083 μM, respectively.
[0079] Sequencing reagent 2: 25 mM Tris-HCl, 25 mM NaCl, 10 mM (NH4)2SO4, 0.02 mg / mL 9°N polymerase (Sailu Medical), 10 mM MgSO4, 1 mM EDTA, and 1 μM of each of the four first nucleotide analogs (dNTP-N3) prepared above.
[0080] Excision reagents: 25 mM tris(3-hydroxypropyl)phosphine (abbreviated THPP), 0.5 M NaCl, 50 mM Tris-HCl, pH 9.0, 0.05% Tween-20.
[0081] Cleaning reagent 1: 750 mM NaCl, 75 mM C6H5Na3O7·2H2O, pH 7.0.
[0082] Cleaning reagent 2: 20 mM Tris-HCl, 2.5 mM NaCl, 0.5 mM EDTA, pH 9.0.
[0083] 2) First, use sequencing reagent 1 and perform 100 cycles of sequencing according to protocol 1. The technical route is as follows: Figure 1 As shown, the details are as follows:
[0084] Amplification and sequencing primer hybridization: Library amplification and sequencing primer hybridization were performed using the Ecoli standard library from Saluseq Medical (catalog number SL-Q00100) and the Saluseq Nimbo sequencing kit (catalog number SL-S00333).
[0085] Binding: Add 120 μL of washing reagent 2 to provide a similar environment for the upcoming binding; add 100 μL of sequencing reagent 1, set the temperature to 60℃, and react for 10 s. Under the action of 9°N polymerase (produced by Shenzhen Sailu Medical) and strontium chloride metal ions, the second nucleotide analog forms a stable complex with the nucleic acid molecule to be tested (i.e., the template strand), the -OH end of the primer, or the -OH end of the first nucleotide analog bound in the previous cycle, so that the second nucleotide analog is stably chelated at the position of the next nucleotide of the template strand; then add 250 μL of imaging reaction solution to wash away the unbound second nucleotide analog and stabilize it for signal acquisition; after acquisition, add 120 μL of washing reagent 1 to react and remove the metal ions in sequencing reagent 1, destroy the structure of the polymerase, and the complex disperses and is washed away.
[0086] Polymerization: Add 120 μL of washing reagent 2 to provide a similar environment for the upcoming polymerization; add 120 μL of sequencing reagent 2, set the temperature to 60 ℃, and react for 25 s to allow the four first nucleotide analogs to polymerize to the primer chain under the action of polymerase. Then add 120 μL of washing reagent 1 to wash away the unreacted first nucleotide analogs. Add 120 μL of excision reagent and react at 65 ℃ for 33 s to expose the -OH ends of the polymerized first nucleotide analogs.
[0087] The binding-polymerization-excision process is repeated to proceed to the next cycle of sequencing.
[0088] The primers were re-hybridized using the same chip, and sequencing reagents 1-2 were used to perform 100 cycles of sequencing according to the same protocol 1 described above. The two sets of data were then compared. The results are as follows: Figures 3 to 5 , and as shown in Table 1. Figure 3 This is the ATCG image of dNTP-C2-H, i.e., sequencing reagent 1. Figure 4 The images are ATCG images of dNTP-N3-H, i.e., sequencing reagents 1-2. Figure 5 The figure shows the comparison results of the Q30 curves of sequencing reagent 1 and sequencing reagent 1-2 for dNTPs. Table 1 shows the comparison of various indicators in the sequencing reports of sequencing reagent 1 and sequencing reagent 1-2 for dNTPs.
[0089] Table 1
[0090] Sequencing reagent 1 Sequencing reagents 1-2 Software Version 1.2.0.415 1.2.0.415 CycleNumber 100 / 100 100 / 100 Read1 Length 100 100 FovNumber 232 / 232 232 / 232 Total Reads (M) 72.91 18.48 Q30(%) 93.22 79.21 Purity (%) 96.27 87.05 ValidRatio(%) 83.01 21.22 NFilterRatio(%) 11.05 40.66 PurityFilterRatio(%) 5.94 38.12 Density(um2) 0.67 0.67
[0091] Figures 3 to 5 As shown in Table 1, the images of dNTP-N3-H (sequencing reagents 1-2) are poor, with more non-specific bindings, such as... Figure 3 and Figure 4 As shown in Table 1, sequencing reagent 1 is significantly better than sequencing reagents 1-2 in terms of sequencing quality over 100 cycles. Figure 5 As shown.
[0092] Based on the above experiments, this example uses the same raw materials and follows Scheme 2 to perform 100 cycles of sequencing, with the technical route as follows: Figure 2 As shown, the details are as follows:
[0093] Amplification and sequencing primer hybridization: Library amplification and sequencing primer hybridization were performed using the Ecoli standard library from Saluseq Medical (catalog number SL-Q00100) and the Saluseq Nimbo sequencing kit (catalog number SL-S00333).
[0094] The cyclic steps include: adding 120 μL of washing reagent 2 to provide a similar environment for the upcoming binding; adding 120 μL of sequencing reagent 2, setting the temperature to 60 °C, and reacting for 25 s to allow the four first nucleotide analogs to polymerize to the primer strand under the action of polymerase; subsequently adding 120 μL of washing reagent 2 to wash away unreacted first nucleotide analogs and simultaneously provide a similar environment for the upcoming binding; adding 100 μL of sequencing reagent 1, setting the temperature to 60 °C, and reacting for 10 s, under the action of 9°N polymerase (Sailu Medical) and the metal ion strontium chloride, the second nucleotide analog forms a stable complex with the target nucleic acid molecule (i.e., the template strand) and the -OH end of the bound first nucleotide analog, allowing the second nucleotide analog to stably chelate at the position of the next nucleotide on the template strand; then adding 250 μL of imaging reaction solution to wash away unbound second nucleotide analogs and stabilize the signal acquisition; after acquisition, adding 120 μL of washing reagent 2 for washing, and then adding 120 μL of... The washing reagent 1 is added in μL to react with the metal ions in the sequencing reagent 1, which destroys the structure of the polymerase. The complex disperses and is washed away. Then, 120 μL of excision reagent is added and the reaction is carried out at 65 °C for 33 s to expose the -OH end of the first nucleotide analog on the polymerase.
[0095] Repeat the above cyclical steps to complete the sequencing.
[0096] The results showed that Scheme 2 and Scheme 1 had similar imaging effects, and the sequencing quality of 100 cycles was comparable for both; therefore, Scheme 1 and Scheme 2 can be used interchangeably as needed.
[0097] Example 2
[0098] 1. Library construction: Use the Cyclo Ecoli standard library (catalog number SL-Q00100), with insert fragments of 450-550 bp. The standard library and sequencing kit are the same as in Example 1.
[0099] 2. Sequencing reagent preparation:
[0100] The four first nucleotide analogs and four second nucleotide analogs are the same as in Example 1. The bases of the four first nucleotide analogs are adenine A, cytosine C, guanine G, and thymine T, respectively. The reversible blocking group is azide, referred to as dNTP-N3, and the structure is shown in Formula 1. The base labels of the four second nucleotide analogs are all linked to the bases of the four second nucleotide analogs through linkers, namely dye-A linked to adenine A, dye-B linked to cytosine C, dye-C linked to guanine G, and dye-D linked to thymine T, referred to as dNTP-C2-H, and the structure is shown in Formula 2.
[0101] Imaging buffer: 10 mM Tris-HCl (25℃, pH 8.8), 50 mM NaCl, 2 M Betaine, 0.05% Tween-20, 0.5 mM EDTA, 5 mM metal ions, 20 mM sodium ascorbate; In this example, groups with different metal ions were designed for the experiment, and the different metal ions are shown in Table 2.
[0102] Sequencing reagent 1: Based on imaging buffers designed with different metal ions, sodium ascorbate was removed, and 0.04 mg / mL 9°N polymerase (Shenzhen Sailu Medical) and the above-prepared mixture of 4 second nucleotide analogs and 4 first nucleotide analogs were added respectively. The concentrations of second nucleotide analog A, T, C and G were 0.025 μM, and the concentrations of first nucleotide analogs A, T, C and G were 0.003 μM, 0.005 μM, and 0.0083 μM respectively.
[0103] Sequencing reagent 2: 50 mM Tris-HCl, 50 mM NaCl, 10 mM (NH4)2SO4, 0.02 mg / mL 9°N polymerase (Sailu Medical), 3 mM MgSO4, 1 mM EDTA, and 1 μM of each of the four first nucleotide analogs prepared above.
[0104] Excision reagents: 20 mM tris(3-hydroxypropyl)phosphine (abbreviated THPP), 0.5 M NaCl, 50 mM Tris-HCl, pH 9.0, 0.05% Tween-20.
[0105] Cleaning reagent 1: 50 mM NaCl, 50 mM C6H5Na3O7·2H2O, pH 7.0.
[0106] Cleaning reagent 2: 20 mM Tris-HCl, 1-5 mM NaCl, 0.5 mM EDTA, pH 9.0.
[0107] 3. Sequencing reagents containing different metal ions were prepared to evaluate their effect on the formation of multi-component complexes. The main metal ions included MgCl2, MnCl2, CaCl2, SrCl2, BaCl2, and NiCl2, all at a concentration of 5 mM. Each metal ion was tested for 50 cycles, and the specific sequencing procedure was the same as in Example 1. The groups corresponding to different metal ions are shown in Table 2.
[0108] Table 2
[0109] metal ions <![CDATA[MgCl2]]> <![CDATA[MnCl2]]> <![CDATA[CaCl2]]> <![CDATA[SrCl2]]> <![CDATA[BaCl2]]> <![CDATA[NiCl2]]> Group 1 2 3 4 5 6
[0110] 4. Experimental Results:
[0111] Partial metal ion binding images (200% view) such as 6 to Figure 8 As shown, Figures 6 to 8 The results for groups 3, 4, and 5 are shown in sequence. The changes in the number of SE50 clusters for some metal ions are shown below. Figure 9 As shown, the SE50 signal intensity changes of some metal ions are as follows: Figure 10 As shown.
[0112] The results showed that: 1) Groups 1 and 6 had no signal; Group 2 had a weak signal and the image focal plane was lost during signal acquisition, making it impossible to complete the SE50 read length test; Groups 3, 4, and 5 were analyzed, and the image quality was as follows: Figures 6 to 8 As shown, under the same test conditions, CaCl2 and SrCl2 have better image quality, while BaCl2 has poor image quality. 2) Analyze the number of cluster points in different test groups of 50 cycles, such as... Figure 9 As shown, it is evident that the number of cluster points in group 5 is significantly insufficient, indicating that the binding efficiency of group 5 is low, while group 4 is more stable. 3) Analyze the changes in signal intensity of different channels in different test groups over 50 cycles, such as... Figure 10 As shown, it is clear that the signal value of group 5 is lower than that of groups 3 and 4, while the difference between groups 3 and 4 is not significant.
[0113] In summary, based on the overall results, SrCl2 showed the best binding effect, followed by CaCl2.
[0114] Example 3
[0115] This example tests the effect of different buffer systems on the formation of multi-component complexes.
[0116] 1. Library construction: Use the Cyclo Ecoli standard library (catalog number SL-Q00100), with insert fragments of 450-550 bp. The standard library and sequencing kit are the same as in Example 1.
[0117] 2. Sequencing reagent 1 was prepared separately using MOPS, HEPES, and Tris buffer systems. Other reagents, including imaging buffer, were the same as in Example 1. MOPS, HEPES, and MOPS were purchased from Aladdin Reagent Co., Ltd.
[0118] Group 1: 50 mM Tris-HCl (25 ℃ pH 9.0), 50 mM NaCl, 2 M Betaine, 0.05% Tween-20, 0.5 mM EDTA, 5 mM SrCl2, 9°N polymerase.
[0119] Group 2: 50 mM MOPS (25℃, pH 7.8), 50 mM NaCl, 2 M Betaine, 0.05% Tween-20, 0.5 mM EDTA, 5 mM SrCl2, 9°N polymerase.
[0120] Group 3: 50 mM HEPES (25°C pH8.0), 50 mM NaCl, 2 M Betaine, 0.05% Tween-20, 0.5mM EDTA, 5 mM SrCl2 9°N polymerase.
[0121] Group 4: 50 mM Tricine (25 ℃ pH 8.5), 50 mM NaCl, 2 M Betaine, 0.05% Tween-20, 0.5 mM EDTA, 5 mM SrCl2, 9°N polymerase.
[0122] Based on the above groups, four second nucleotide analogs and four first nucleotide analogs are added to obtain the corresponding four groups of sequencing reagents 1. The four second nucleotide analogs and four first nucleotide analogs and their dosages are the same as in Example 1.
[0123] 3. Following technical route 2, i.e. scheme 2, SE150 sequencing was performed on the 4 sets of data, and the differences in images, signal values, quality values, and error rates were analyzed.
[0124] The results are as follows Figures 11 to 13 , and as shown in Table 3. Figure 11 This is a comparison of the image quality (C bases) of four sets of experiments. Figure 12 These are the sequencing signal value curves from four experimental groups. Figure 13Table 3 shows the sequencing Q-value curves for the four experiments, and the results of the alignment rate and error rate analysis for the four experiments.
[0125] Table 3
[0126] test Group 1 Group 2 Group 3 Group 4 sequencing length 150 150 150 150 Phasing 0.07 0.07 0.07 0.08 Prephasing 0.07 0.07 0.07 0.07 Raw Q30 (%) 96.39 92.02 94.39 95.77 Density / um^2 0.69 0.65 0.68 0.7 Available data volume / M 74.18 68.86 72.86 73.70 Mapping_rate(%) 99.999 99.999 99.999 99.999 Mismatch_rate (%) 0.101 0.462 0.203 0.183
[0127] Figures 11 to 13 And the results in Table 3 show that: 1) In terms of original image quality and signal, group 1 has a higher signal, the other three groups of conditional images are not significantly different, and group 2 has a slightly lower signal. Figure 11 As shown; 2) Judging from the signal changes and Q30 trends, group 1 is better than group 4, while group 2 has poorer signal and Q30, as shown. Figure 12 and Figure 13 As shown in Table 3; 3) Comparing the error rates of the four groups, group 4 is significantly better than the other two groups. Overall, group 1 has the best conditions.
[0128] Example 4
[0129] This example tested the effect of different polymerases on dNTP binding efficiency.
[0130] 1. Preparation of DNA polymerase and its mutants
[0131] 1) Preparation of recombinant bacteria
[0132] The encoding gene for 9°N DNA polymerase was inserted into the XbaI and NotI molecules of pET-28a (Novagen) to obtain a recombinant vector, which was named SL-V. The SL fusion protein is a fusion protein obtained by adding a his tag consisting of six histidine residues after the first methionine residue of SL. The sequence of the SL fusion protein (labeled SL1) is shown in SEQ ID NO.1. The mutants of 9°N DNA polymerase were labeled SL2 and SL3, respectively. The sequence of SL2 is shown in SEQ ID NO.2, and the sequence of SL3 is shown in SEQ ID NO.3.
[0133] SL-V was introduced into Escherichia coli BL21(DE3) (Tiangen Biotech (Beijing) Co., Ltd.) to obtain a recombinant bacterium, which was named BL21-SL-V. BL21-SL-V expresses the SL fusion protein.
[0134] 2) Preparation of DG fusion protein
[0135] Induced expression
[0136] The recombinant BL21-SL-V obtained in step 1 was inoculated into 100 mL of Kan-LB medium (Kan-LB medium is a liquid medium in which kanamycin is added to LB medium to obtain a final kanamycin concentration of 50 µg / mL) and cultured overnight at 37 ℃ and 200 rpm (approximately 16 h). The resulting bacterial solution was inoculated into 1000 mL of Kan-LB medium at a ratio of 1:100 and cultured at 37 ℃ and 200 rpm for 4 h. IPTG was added to the resulting bacterial solution until its final concentration in the bacterial solution was 0.5 mM, and then cultured overnight at 25 ℃ and 200 rpm (approximately 16 h). BL21-SL-V cells were collected at 10000 g for 20 min.
[0137] Bacterial cell disruption and purification
[0138] Take the BL21-SL-V cells from step 2.1 and add 20 ml of lysis buffer (50 mM Tris, 200 mM NaCl, 5% Glycerol, pH 7.5) to 1 g of cells. Resuspend the cells in the lysis buffer to obtain a bacterial resuspension. Place the bacterial resuspension in an ultrasonic disruptor and sonicate at 40% power (approximately 400W) for 5 seconds, pause for 5 seconds, for a total of 20 minutes. Centrifuge the ultrasonic product at 20000 g at 4 ℃ for 30 minutes and collect the supernatant.
[0139] After filtering the supernatant through a 0.45 μm syringe filter (Life Sciences), the supernatant was loaded at an appropriate flow rate for Ni column affinity chromatography (HisTrap HP pre-packed column, 5 mL, 17-5248-02, GE Healthcare). After loading, the supernatant was equilibrated with affinity buffer A (75 mM Tris, 500 mM NaCl, 20 mM imidazole, 5% Glycerol, pH 7.4) for 10 CV; eluted with 50% affinity buffer 2 for 5 CV, and the Ni column affinity chromatography eluent corresponding to peaks greater than or equal to 100 mAU was collected.
[0140] The eluent corresponding to peak values greater than or equal to 100 mAU was dialyzed overnight using ion buffer A (50 mM Tris, 50 mM NaCl, 5% Glycerol, pH 7.4). The dialyzed solution was then loaded at a controlled flow rate for cation exchange chromatography (HiTrap SP HP pre-packed ion exchange column, 5 mL, 17-1152-01, GE Healthcare). After loading, the column was equilibrated with ion buffer A for 10 CV; eluted with 50% ion buffer B for 5 CV. The eluent corresponding to peak values greater than or equal to 100 mAU was collected. The eluent was then used in a dialysis bag (SpectrumLabs, 131267) and dialyzed for 24 hours in dialysate (50 mM Tris, 400 mM KCl, 0.2 mM EDTA, 5% Glycerol). Finally, glycerol was added to a final concentration of 50% and quantified to 0.8 mg / mL to obtain the purified DG fusion protein.
[0141] SEQ ID NO.1 (9°N -SL1)
[0142] MHHHHHHHILDTDYITENGKPVIRVFKKENGEFKIEYDRTFEPYFYALLKDDSAIEDVKKVTAKRHGTVVKVKRAEKVQKKFLGRPIEVWKLYFNHPQDVPAIRDRIRAHPAVVDIYEYDIPFAKRYLIDPEGDEELTMLAFAIATLYHEGEFGTGPILMISYADGSEARVITWKKIDLPYVDVVSTEKEMIKRFLRVVREKDPDVLITYNGDNFDFAYLKKRCEELGIKFTLGRDGSEPKIQRMGDRFAVEVKGRIHFDLYPVIRRTINLPTYTLEAVYEAVFGKPKEKVYAEEIAQAWESGEGLERVARYSMEDAKVTYELGREFFPMEAQLSRLIGQSLWDVSRSSTGNLVEWFLLRKAYKRNELAPNKPDERELARRRGGY AGGYVKEPERGLWDNIVYLDFRSLSASIITHNVSPDTLNREGCKEYDVAPEVGHKFCKDFPGFIPSLLGDLLEERQKIKRKMKATVDPLEKKLLDYRQRLIKILANSFYGYYGYAKARWYCKECAESVTAWGREYIEMVIRELEEKFGFKVLYADTDGLHATIPGADAETVKKKAKEFLKYINPKLPGLLELEYEGFYVRGFFVTKKKYAVIDEEGKITTRGLEIVRRDWSEIAKETQARVLEAILKHGDVEEAVRIVKEVTEKLSKYEVPPEKLVIHEQITRDLRDYKATGPHVAKRLAARGVKIRPGTVISYIVLKGSGRIGDRAIPADEFDPTKHRYDAEYYYIENQVLPAVERILKAFGYRKEDLRYQKTKQVGLGAWLKVKGKK
[0143] SEQ ID NO.2(9°N -SL2)
[0144] MHHHHHHHILDTDYITENGKPVIRVFKKENGEFKIEYDRTFEPYFYALLKDDSAIEDVKKVTAKRHGTVVKVKRAEKVQKKFLGRPIEVWKLYFNHPQDVPAIRDRIRAHPAVVDIYEYDIPFAKRYLIDPEGDEELTMLAFAIATLYHEGEFGTGPILMISYADGSEARVITWKKIDLPYVDVVSTEKEMIKRFLRVVREKDPDVLITYNGDNFDFAYLKKRCEELGIKFTLGRDGSEPKIQRMGDRFAVEVKGRIHFDLYPVIRRTINLPTYTLEAVYEAVFGKPKEKVYAEEIAQAWESGEGLERVARYSMEDAKVTYELGREFFPMEAQLSRLIGQSLWDVSRSSTGNLVEWFLLRKAYKRNELAPNKPDERELARRRGGY AGGYVKEPERGLWDNIVYLDFRSLSASIITHNVSPDTLNREGCKEYDVAPEVGHKFCKDFPGFIPSLLGDLLEERQKIKRKMKATVDPLEKKLLDYRQRLIKILANSFYGYYGYAKARWYCKECAESVTAWGREYIEMVIRELEEKFGFKVLYADPDGLHATIPGADAETVKKKAKEFLKYINPKLPGLLELEYEGFYVRGFFVTKKKYAVIDEEGKITTRGLEIVRRDWSEIAKETQARVLEAILKHGDVEEAVRIVKEVTEKLSKYEVPPEKLVIHEQITRDLRDYKATGPHVAKRLAARGVKIRPGTVISYIVLKGSGRIGDRAIPADEFDPTKHRYDAEYYYIENQVLPAVERILKAFGYRKEDLRYQKTKQVGLGAWLKVKGKK
[0145] SEQ ID NO.3(9°N -SL3)
[0146] MHHHHHHILDTDYITENGKPVIRVFKKENGEFKIEYDRTFEPYFYALLKDDSAIEDVKKVTAKRHGTVVKVKRAEKVQKKFLGRPIEVWKLYFNHPQDVPAIRDRIRAHPAVVDIYEYDIPFAKRYLIDKGLIPMEGDEELTMLAFAIATLYHEGEEFGTGPILMISYADGSEARVITWKKIDLPYVDVVSTEKE MIKRFLRVVREKDPDVLITYNGDNFDFAYLKKRCEELGIKFTLGRDGSEPKIQRMGDRFAVEVKGRIHFDLYPVIRRTINLPTYTLEAVYEAVFGKPKEKVYAEEIAQAWESGEGLERVARYSMEDAKVTYELGREFFPMEAQLSRLIGQSLWDVSRSSTGNLVEWFLLRKAYKRNELAPNKPDERELARRRGGY AGGYVKEPERGLWDNIVYLDFSSLSASIIITHNVSPDTLNREGCKEYDVAPEVGHKFCKDFPGFIPSLLGDLLEERQKIKRKMKATVDPLEKKLLDYRQRLIKILANSFYGYYGYAKARWYCKECAESVTAWGREYIEMVIRELEEKFGFKVLYADTDGLHATIPGADAETVKKKAKEFLKYINPKLPGLLELEY EGFYVRGFFVTKKKYAVIDEEGKITTRGLEIVRRDWSEIAKETQARVLEAILKHGDVEEAVRIVKEVTEKLSKYEVPPPEKLVIHEQITRDLRDYKATGPHVAVAKRLAARGVKIRPGTVISYIVLKGSGRIGDRAIPADEFDPTKHRYDAEYYIENQVLPAVERILKAFGYRKEDLRYQKTKQVGLGAWLKVKGKK
[0147] 2. Library construction: Use the Cyclo Ecoli standard library (catalog number SL-Q00100), with insert fragments ranging from 450 to 550 bp. The standard library and sequencing kit are the same as in Example 1.
[0148] 3. Select different types of polymerases to prepare sequencing reagent 1. Other reagents are the same as in Example 1.
[0149] The different polymerases and their amounts are shown in Table 4. The remaining components of sequencing reagent 1 are the same as in Example 1. Other reagents used for sequencing, such as sequencing reagent 2, imaging buffer, excision reagent, washing reagent 1, and washing reagent 2, are the same as in Example 1.
[0150] Table 4
[0151] Group polymerase Enzyme dosage brand Group 1 BST 0.08 mg / mL NEB Group 2 Klenow 0.08 mg / mL NEB Group 3 9°N -SL1 0.04 mg / mL Salus Group 4 9°N -SL2 0.04 mg / mL Salus Group 5 9°N -SL3 0.04 mg / mL Salus
[0152] 4. Following technical route 2, i.e. scheme 2, perform SE100 sequencing, analyze the images, and compare the results.
[0153] The results are as follows Figures 14 to 16 , and as shown in Table 5. Figure 14 These are images from group 1 and group 2. Figure 15 It compares the signal strength of groups 3, 4, and 5. Figure 16 Table 5 compares the Q30 curves of Groups 3, 4, and 5. It also compares the data of the three groups.
[0154] Table 5
[0155] Metric Group 3 Group 4 Group 5 Raw_total_reads(M) 58.2693 82.1994 73.7408 Raw_GC_content(%) 50.6248 50.5389 50.4292 Clean_total_reads(M) 52.4597 79.8895 70.8102 Clean_Q30(%) 93.0317 96.21 94.8425 All_reads 52459679 79889503 70810186 Mapped_reads 52436475 79885779 70778457 Mapping_rate(%) 99.955768 99.995339 99.955191 Unique_mapping_rate(%) 96.853301 96.896386 96.819057 Mapped_bases(cigar) 5222305808 7961292844 7051842377 All_mismatch_bases 29075517 11753107 25418826 Mismatch_rate (%) 0.556 0.147 0.360
[0156] Figures 14 to 16 And the results in Table 5 show:
[0157] i. Groups 1 and 2 have virtually no signal, such as Figure 14 As shown.
[0158] ii. Groups 3 and 4 showed better image quality. Under the same conditions, sequencing at 100 bp, groups 4 and 5 were superior to group 3 in terms of signal intensity and Q30, especially in Q30, where group 3 performed poorly. Figure 15 and Figure 16 As shown.
[0159] iii. Comparing the error rates of the three groups, group 4 was significantly better than the other two groups, as shown in Table 5.
[0160] Overall, the 9°N-SL2 enzyme in group 4 is the best.
[0161] Example 5
[0162] The WES standard library (NA12878) was sequenced, and the variant results were analyzed.
[0163] 1. Library Construction: A WES NA12878 standard library was constructed using the Agitec Capture Kit (catalog number PH2006465) and the Novizan Library Construction Kit (catalog number ND627), with insert fragments ranging from 400-500 bp. The NA12878 human genomic DNA standard was purchased from Beijing Bio-Innovation Technology Co., Ltd., under the brand name Coriell.
[0164] 2. PE150 sequencing was performed using the methods and optimal reagents from Examples 1 to 4 to evaluate the accuracy and precision of variant detection.
[0165] Sequencing reagent 1 formulation: 50 mM Tris-HCl (25 ℃ pH9.0), 50 mM NaCl, 2 M Betaine, 0.05% Tween-20, 0.5 mM EDTA, 5 mM SrCl2, 0.04 mg / mL 9°N -SL2, 4 second nucleotide analogs and 4 first nucleotide analogs and their dosages are the same as in Example 1.
[0166] The results are shown in Table 6. Table 6 compares the mutation results of the two sequencing methods.
[0167] Table 6
[0168] SBS_WES SBB1_WES SBB2_WES Total_reads(M) data volume 163.599 160.9665 164.1198 Q30 (%) quality 97.6795 98.9587 98.9884 All_reads 79049800 77664746 77692535 Mapping_rate(%) alignment rate 99.957009 99.960418 99.959093 Mismatch_rate error rate 3.02E-03 7.76E-04 7.20E-04 Capture rate (%) 74.7115 76.2201 76.0201 Mean_coverage_depth 130.3529 130.4613 130.4625 Recall: SNP (%) - call SNP sensitivity 99.31 99.552 99.571 Precision: SNP (%) - call SNP accuracy 99.781 99.791 99.789 F-score: SNP (%) 99.545 99.491 99.525 Recall: InDel (%) - call InDel sensitivity 96.027 96.913 97.108 Precision: InDel (%) call InDel accuracy 98.538 98.568 98.743 F-score:InDel(%) 97.777 97.822 97.959
[0169] In Table 6, "SBS_WES" refers to the mainstream SBS sequencing method currently on the market, "SBB1_WES" refers to the sequencing method according to Scheme 1 of Example 1, and "SBB 2_WES" refers to the sequencing method according to Scheme 2 of Example 1. The sequencing data were analyzed for variants using DeepVariant analysis software.
[0170] The results in Table 6 show that Scheme 1, Scheme 2 and the matching reagents of this application have better Q30, lower error rate and better variation analysis results.
[0171] The above description, in conjunction with specific embodiments, provides a further detailed explanation of this application and should not be construed as limiting the specific implementation of this application to these descriptions. Those skilled in the art to which this application pertains can make several simple deductions or substitutions without departing from the concept of this application.
Claims
1. A gene sequencing method, characterized in that: This includes using sequencing reagent 1 and sequencing reagent 2 to sequence the nucleic acid molecules to be tested; The sequencing reagent 1 includes four second nucleotide analogs and four first nucleotide analogs, and the sequencing reagent 2 includes four first nucleotide analogs; The first nucleotide analog has the structure shown in Formula 1, and the second nucleotide analog has the structure shown in Formula 2; Set 1 ; Formula 2 ; In Formulas 1 and 2, the "Base" of the four second nucleotide analogs and the four first nucleotide analogs are adenine A, cytosine C, guanine G, and thymine T, respectively; in Formula 1, R1 is a reversible blocking group; in Formula 2, "Dye" is a fluorescent group, "Linker" is a linker, and "X" is independently selected from CH2, NH, S, CF2, CBr2, and Se. The four second nucleotide analogs are labeled with four different fluorescent groups. The sequencing of the nucleic acid molecule to be tested includes sequencing using at least one of the following schemes after the sequencing primers are hybridized to the nucleic acid molecule to be tested. Scheme 1: (1) Contact with sequencing reagent 1, and under the action of metal ions and polymerase, the second nucleotide analog forms a complex with the nucleic acid molecule to be tested, the 3' end of the sequencing primer, or the 3' end of the extended sequencing primer; (2) Contact with imaging buffer to wash away unbound second nucleotide analog; (3) Signal acquisition; (4) After signal acquisition, remove the bound second nucleotide analog; (5) Contact with sequencing reagent 2, and under the action of polymerase, the first nucleotide analog polymerizes to the 3' end of the sequencing primer or the 3' end of the extended sequencing primer; (6) Remove the reversible blocking group of the first nucleotide analog; Repeat steps (1) to (6) to complete the sequencing of the nucleic acid molecule to be tested; Scheme 2: (1) Contact with sequencing reagent 2, and under the action of polymerase, the first nucleotide analog polymerizes to the 3' end of the sequencing primer or the 3' end of the sequencing primer extension; (2) Contact with sequencing reagent 1, and under the action of metal ions and polymerase, the second nucleotide analog forms a complex with the nucleic acid molecule to be tested and the 3' end of the sequencing primer extension; (3) Contact with imaging buffer to wash away unbound second nucleotide analog; (4) Signal acquisition; (5) After signal acquisition, remove the bound second nucleotide analog; (6) Remove the reversible blocking group of the first nucleotide analog; Repeat steps (1) to (6) to complete the sequencing of the nucleic acid molecule to be tested.
2. The gene sequencing method according to claim 1, characterized in that: In the first scheme, step (1) further includes contacting the cleaning agent first to provide a reaction environment for the formation of the complex in step (1).
3. The gene sequencing method according to claim 1, characterized in that: In the first scheme, step (5) further includes contacting the cleaning reagent first to provide a reaction environment for the polymerization of the first nucleotide analog in step (5).
4. The gene sequencing method according to claim 1, characterized in that: In the first scheme, step (6) further includes contacting the first nucleotide analog with a cleaning agent to wash away the unpolymerized first nucleotide analog and to provide a reaction environment for removing the reversible blocking group of the first nucleotide analog.
5. The gene sequencing method according to claim 1, characterized in that: In the second scheme, step (1) further includes contacting the cleaning reagent first to provide a reaction environment for the polymerization of the first nucleotide analog in step (1).
6. The gene sequencing method according to claim 1, characterized in that: In the second scheme, step (2) further includes contacting with a cleaning reagent to wash away the unpolymerized first nucleotide analog and to provide a reaction environment for the formation of the complex in step (2).
7. The gene sequencing method according to claim 1, characterized in that: In the second scheme, step (6) further includes contacting the cleaning reagent first to provide a reaction environment for removing the reversible blocking group of the first nucleotide analog.
8. A gene sequencing kit, characterized in that: Includes sequencing reagent 1 and sequencing reagent 2; The sequencing reagent 1 includes four second nucleotide analogs and four first nucleotide analogs, and the sequencing reagent 2 includes four first nucleotide analogs; The first nucleotide analog has the structure shown in Formula 1, and the second nucleotide analog has the structure shown in Formula 2; Set 1 ; Formula 2 ; In Formula 1 and Formula 2, the "Base" of the four second nucleotide analogs and the four first nucleotide analogs are adenine A, cytosine C, guanine G, and thymine T, respectively; in Formula 1, R1 is a reversible blocking group; in Formula 2, "Dye" is a fluorescent group, "Linker" is a linker, and "X" is independently selected from any one of CH2, NH, S, CF2, CBr2, and Se. The four second nucleotide analogs are labeled with four different fluorescent groups.
9. The gene sequencing kit according to claim 8, characterized in that: It also includes at least one of imaging buffer, excision reagent, cleaning reagent 1, and cleaning reagent 2; The imaging buffer solution comprises 10-100 mM Tris-HCl, 10-50 mM NaCl, 1-2 M Betaine, 0.02-0.5% Tween-20, 0.1-1 mM EDTA, 1-70 mM non-catalytic metal ions, and 5-70 mM antioxidant. The excision reagent includes 10-50 mM tris(3-hydroxypropyl)phosphine, 0.2-2 M NaCl, 10-100 mM Tris-HCl, and 0.02-0.5% Tween-20; The cleaning reagent 1 includes 20-100 mM NaCl and 25-100 mM C6H5Na3O7·2H2O; The cleaning reagent 2 includes 10-100 mM Tris-HCl, 1-5 mM NaCl, and 0.1-1 mM EDTA; Optionally, the non-catalytic metal ion includes at least one of strontium chloride, calcium chloride, and barium chloride; Optionally, the antioxidant includes at least one of sodium ascorbate, Trolox, and glutathione; Optionally, the sequencing reagent 1 is prepared by removing the antioxidant from the imaging buffer and adding polymerase, four second nucleotide analogs and four first nucleotide analogs; Optionally, in the sequencing reagent 1, the four second nucleotide analogs, or "Base," are second nucleotide analogs of adenine A, cytosine C, guanine G, and thymine T, respectively, and are labeled as second nucleotide analog A, second nucleotide analog T, second nucleotide analog C, and second nucleotide analog G, respectively. The concentration of second nucleotide analog A is 0.02-0.05 μM, the concentration of second nucleotide analog T is 0.5-1 μM, the concentration of second nucleotide analog C is 0.5-1 μM, and the concentration of second nucleotide analog G is 0.02-0.05 μM. Similarly, the four first nucleotide analogs, or "Base," are first nucleotide analogs of adenine A, cytosine C, guanine G, and thymine T, respectively, and are labeled as first nucleotide analog A, first nucleotide analog T, first nucleotide analog C, and first nucleotide analog G, respectively. The concentration of first nucleotide analog A is 0.01-0.02 μM, and the concentration of first nucleotide analog T is 0.003-0.005 μM. μM, the concentration of the first nucleotide analog C is 0.002-0.005 μM, and the concentration of the first nucleotide analog G is 0.008-0.01 μM; Optionally, the sequencing reagent 2 includes 10-100 mM Tris-HCl, 5-50 mM NaCl, 2-20 mM (NH4)2SO4, 0.01-0.1 mg / mL polymerase, 2-20 mM MgSO4, 0.2-1 mM EDTA, and four first nucleotide analogs, wherein each of the first nucleotide analogs A, T, C, and G is 0.2-1 μM.
10. The gene sequencing kit according to claim 8 or 9, characterized in that: The reversible blocking group includes at least one of methylene azide, allyl, ester, phosphoric acid, hydroxylamine, disulfide bond, photolytic group, 2-nitrobenzyl, and azo compound; Optionally, the reversible blocking group includes at least one of the following groups: ; And / or, the connector is a shardable connector or a non-shardable connector; Optionally, the cleavable linker includes at least one of the following: electrophilic cleavage linker group, nucleophilic cleavage linker group, photolytic linker group, cleavage group under reducing conditions, cleavage group under oxidizing conditions, safety handle type linker group, and group cleaved by elimination mechanism. Optionally, the cleavable linker is at least one selected from alkyl, allyl, azidomethylene, 2-nitrobenzyl, and dithio. Optionally, the non-cleavable linker includes at least one of a polyethylene glycol chain and a polypeptide chain; And / or, the fluorescent labels of the four second nucleotide analogs are non-repeatingly selected from AF488, AF532, AF633, AF680, AF660, AF700, AF647, AF594, AF555, AF568, CY3, CY5, CY5.5, CY7, CY7.5, ROX, R6G, ATTO495, ATTO532, ATTO700, ATTO680, ATTO655, ATTO647N, ATTO594, ATTO Rho101, ATTO590, ATTO Thio12, FAM, VIC, TET, JOE, HEX, CAL Fluor Orange560, TAMRA, CAL Fluor Red610, TEXAS RED, CAL Fluor Red635, iFluor488, iFluor514, iFluor532, iFluor546, iFluor555, iFluor568, iFluor590, i Fluor610, iFluor633, iFluor647, iFluor680, iFluor700, iFluor710, Quasar705, Quasar670.