Joint fixed on chip and sequencing method of sample library connected on joint
By designing special base linker sequences and enzyme denaturation treatment on the chip and adjusting the sequencing order, the problem of prolonged analysis result output time in the existing technology is solved, and the effect of rapid node-by-node output of analysis results is achieved.
Patent Information
- Application Number
- CN202510866208.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-12
AI Technical Summary
Existing sequencing technology cannot perform first-end and second-end tag sequencing first, and then first-end and second-end target fragment sequencing, which results in prolonged analysis result output time and inability to output analysis results at the fastest speed.
A connector fixed on the chip is designed, including P5 and P7 connector sequences, which introduce special bases, such as U, I, and G*, at specific positions. Through enzyme digestion and denaturation treatment, the sequencing order is adjusted. Tag sequencing is performed first, and then target fragment sequencing is performed. Sequencing-while-splitting and multi-node batch analysis are supported.
It reduces the output time of short-read sequencing sample results under different read length modes, supports splitting while sequencing, improves the timeliness and immediacy of analysis, and can output analysis results in batches on multiple nodes.
Smart Images

Figure CN120624432A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sequencing, and in particular to a method for sequencing a linker fixed on a chip and a sample library connected to the linker. Background Art
[0002] With the popularization of sequencing technology and the development of its application directions, sequencing modes with a single application direction or read length can no longer meet the needs of users. One of them is to simultaneously sequence mixed libraries with different read lengths and output analysis results in batches and multiple nodes. In addition, there are also high requirements for the timeliness or immediacy of the output of analysis results. How to improve the timeliness or immediacy of analysis has become a key factor.
[0003] The conventional sequencing model first sequences the target fragment and tag, then sequences the tag and target fragment at the second end. Data analysis, however, involves data splitting (splitting sequencing data belonging to the same sample from multiple sample data) based on the instructions of the tag within the sample. This requires completing the first-end target fragment sequencing and obtaining complete dual-tag sequencing. In other words, data analysis can only be performed after the longest sequencing read at the first end is complete. This significantly impacts the time it takes to produce analysis results for dual-tag short-read samples. Therefore, to improve the timeliness or immediacy of analysis, it is necessary to adjust the sequencing order, sequencing the first- and second-end tags first, followed by sequencing the first- and second-end target fragments. This allows single-end dual-tag samples to be split and analyzed for different read lengths as needed during first-end target fragment sequencing. After second-end target fragment sequencing is complete, data from the sample with the longest sequencing read is split and analyzed. This final result is then used to correct the split results for the long reads required. This sequencing model allows for the fastest node-by-node analysis results. According to this approach, the P5 adapter on the sequencing chip of current mainstream gene sequencers only has a special U base. After completing step 2, reading the Index2 (second tag) sequence, there is no special base recognition site, making it impossible to remove the oligonucleotides bound to the P5 adapter during Index2 sequencing. The presence of these oligonucleotides hinders the binding of the first target fragment sequencing primer in step 3. Therefore, it is impossible to complete the task of first sequencing the index and then sequencing the target fragment, and subsequently it is impossible to output the analysis results in the fastest node-by-node manner. Summary of the Invention
[0004] In a first aspect, the present invention provides a connector fixed on a chip, comprising a P5 connector sequence and a P7 connector sequence, wherein the P5 connector sequence comprises a first identifier base and a second identifier base, the first identifier base is located at the position of the T base, and the second identifier base is located at the position of the C base; the P7 connector sequence comprises a third identifier base, and the third identifier base is located at the position of the G base; wherein T is thymine, C is cytosine, and G is guanine.
[0005] Preferably, the first marker base is U, the second marker base is I, and the third marker base is G*, wherein U is uracil, I is hypoxanthine, and G* is 8-oxoguanine.
[0006] Preferably, the P5 linker sequence is 5'-AATGATACGGCGACCACCGAGAUCTAIAC-3'; and the P7 linker sequence is 5'-CAAGCAGAAGACGGCATACGAG*AT-3'.
[0007] Preferably, one of the first identifier base and the second identifier base is close to the 5' end of the P5 linker sequence.
[0008] Preferably, when the first identifier base is close to the 5' end of the P5 adapter sequence, the distance between the first identifier base and the 5' end of the P5 adapter sequence is a first distance, and the distance between the tag primer binding site at the corresponding position on the complementary chain of the P5 adapter sequence and the 5' end of the P5 adapter sequence is a second distance, and the first distance is less than or equal to the second distance.
[0009] Preferably, if the first distance is smaller than the second distance, the difference between the first distance and the second distance is greater than or equal to 2 bp.
[0010] Preferably, the length of the tag primer sequence bound to the tag primer binding site at the corresponding position on the complementary strand of the P5 adapter sequence is greater than or equal to 15 bp and does not overlap with the tag sequence.
[0011] Preferably, the third identifier base is close to the 5' end or 3' end of the P7 linker sequence.
[0012] A second aspect of the present invention provides a method for sequencing a sample library, wherein the sample library is connected to the adapter fixed on the chip as described in the first aspect, and when the first identifier base is close to the 5' end of the P5 adapter sequence, the method comprises the following steps:
[0013] The second marker base is removed, and the chain connected to the P5 adapter sequence is denatured and removed, leaving the chain connected to the P7 adapter sequence;
[0014] For the remaining chain, the first tag close to the 5' end of the P7 adapter sequence is first sequenced, and the first tag primer and the complementary sequence of the first tag sequence are removed by denaturation during sequencing; then the P5 adapter sequence fixed on the chip is used as a primer to sequence the second tag away from the 5' end of the P7 adapter sequence, and the first identification base is removed, and then the second tag primer and the complementary sequence of the second tag sequence are removed by denaturation during sequencing; then one end of the nucleic acid fragment is sequenced, and the primer at one end of the nucleic acid fragment and the complementary sequence at one end of the nucleic acid fragment are removed by denaturation.
[0015] A third aspect of the present invention provides a method for sequencing a sample library, wherein the sample library is connected to the adapter fixed on the chip as described in the first aspect, and when the second identifier base is close to the 5' end of the P5 adapter sequence, the method comprises the following steps:
[0016] The first marker base is removed, and the chain connected to the P5 adapter sequence is denatured and removed, leaving the chain connected to the P7 adapter sequence including the sample library;
[0017] For the remaining chain, the first tag close to the 5' end of the P7 adapter sequence is first sequenced, and the first tag primer and the complementary sequence of the first tag sequence are removed by denaturation during sequencing; then the P5 adapter sequence fixed on the chip is used as a primer to sequence the second tag away from the 5' end of the P7 adapter sequence, and the second identification base is removed, and then the second tag primer and the complementary sequence of the second tag sequence are removed by denaturation during sequencing; then one end of the nucleic acid fragment is sequenced, and the primer at one end of the nucleic acid fragment and the complementary sequence at one end of the nucleic acid fragment are removed by denaturation.
[0018] Preferably, the method further comprises the steps of:
[0019] Reconnect the sample library starting from the P5 adapter sequence fixed on the chip;
[0020] The third marker base is removed, and the chain connected to the P7 linker sequence is denatured and removed, leaving the chain connected to the P5 linker sequence;
[0021] The remaining strand is then sequenced at the other end of the nucleic acid fragment.
[0022] Preferably, if the first marker base is U, the second marker base is I, and the third marker base is G*, human alkyladenine DNA glycosylase or 3-methyladenine DNA glycosylase II is used to remove the I base, uracil DNA glycosylase and nuclease VIII are used to remove the U base, and 8-oxoguanine DNA glycosylase is used to remove the G* base.
[0023] Compared with the prior art, the present invention has the following beneficial effects:
[0024] 1. The present invention changes the traditional sequencing mode and can support first-end and second-end tag sequencing, and then first-end and second-end target fragment sequencing.
[0025] 2. When different read length modes are mixed for sequencing in the present invention, the result output time of short read length sequencing samples can be shortened.
[0026] 3. The present invention can support the sequencing-while-splitting mode, and support multi-node batch output of analysis results. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 A schematic diagram of a connector fixed on a chip provided in Example 1;
[0028] Figure 2 This is a schematic diagram of the sample library provided in Example 1 after bridge PCR amplification;
[0029] Figure 3 A schematic diagram of the sequencing process provided in Example 1;
[0030] Figure 4 A schematic diagram of a connector fixed on a chip provided in Example 2;
[0031] Figure 5 Schematic diagram of the sample library provided in Example 2 after bridge PCR amplification;
[0032] Figure 6 A schematic diagram of the sequencing process provided in Example 2;
[0033] In the figure, P5 is the P5 adapter sequence, P7 is the P7 adapter sequence, DNA insert is the nucleic acid fragment, Index1 primer site is the first index primer binding site, Index2 primer site is the second index primer binding site, Read1 primer site is the primer binding site at one end of the nucleic acid fragment, Index1 primer is the first index primer, Index2primer is the second index primer, and Read1 primer is the primer at one end of the nucleic acid fragment. DETAILED DESCRIPTION
[0034] The following is a clear and complete description of the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present invention.
[0035] Example 1
[0036] like Figure 1 As shown, this embodiment provides a connector fixed on a chip, including a P5 connector sequence and a P7 connector sequence, wherein the P5 connector sequence includes two identifier bases, namely an I base and a U base, wherein the U base is close to the 5' end of the P5 connector sequence and the I base is far from the 5' end of the P5 connector sequence. The I base in the P5 connector sequence can replace any C base in the P5 connector sequence in the prior art, and the U base in the P5 connector sequence can replace any T base in the P5 connector sequence in the prior art. The P7 connector sequence includes an identifier base G*, and the G* base in the P7 connector sequence can replace any G base in the P7 connector sequence in the prior art.
[0037] The P5 and P7 adapter sequences immobilized on the chip in this example are compatible with, and completely complementary to, the adapter sequences of currently available mainstream libraries. Therefore, sequencing can be completed using the adapters provided in this example without modification or optimization of the existing mainstream library structures.
[0038] The sample library is connected to the connector fixed on the chip provided in this embodiment, and can be Figure 2 As shown ( Figure 2 Taking only one pair of complementary chains on a pair of P7 linker sequences and a P5 linker sequence as an example, there are actually tens of millions or even hundreds of millions of such complementary chains on the chip).
[0039] When tag sequencing is required first and then target fragment sequencing, the sequencing process can be as follows: Figure 3 Specifically, the following steps may be included:
[0040] In step 1, human alkyladenine DNA glycosylase or 3-methyladenine DNA glycosylase II is used to remove the marker base I base away from the 5' end of the chain connected to the P5 linker sequence, denature and remove the chain connected to the P5 linker sequence, and leave the chain connected to the P7 linker sequence.
[0041] In step 2, the first tag primer is bound to the first tag primer binding site, and the complementary sequence of the first tag sequence is synthesized with the first tag primer as the synthesis starting point, and the reading (sequencing) of the first tag sequence is completed through the complementary sequence, and then the first tag primer and the complementary sequence of the first tag sequence are removed by denaturation.
[0042] Step three: Use the P5 linker sequence on the chip as the second tag primer, and use the second tag primer as the synthesis starting point to synthesize the complementary sequence of the second tag sequence. The second tag sequence is read (sequencing) through the complementary sequence. Then, use uracil DNA glycosylase and nuclease endonuclease VIII to remove the marker base U base near the 5' end of the chain connected to the P5 linker sequence. Finally, denaturation is performed to remove the second tag primer and the complementary sequence of the second tag sequence.
[0043] This completes tag sequencing.
[0044] Step 4: The primer at one end of the nucleic acid fragment is combined with the primer binding site at one end of the nucleic acid fragment, and the complementary sequence of the sequence at one end of the nucleic acid fragment is synthesized with the primer at one end of the nucleic acid fragment as the starting point of synthesis, and the reading (sequencing) of the sequence at one end of the nucleic acid fragment is completed through the complementary sequence, and then the primer at one end of the nucleic acid fragment and the complementary sequence of the sequence at one end of the nucleic acid fragment are removed by denaturation.
[0045] This completes the sequencing of one end of the nucleic acid fragment.
[0046] Step 5: Reconnect the sample library starting from the P5 adapter sequence on the chip to synthesize a new chain (ie, restore to the original complementary chain state).
[0047] Step 6: Use 8-oxoguanine DNA glycosylase to remove the marker base G* on the P7 linker sequence, denature and remove the chain connected to the P7 linker sequence, and leave the chain connected to the P5 linker sequence.
[0048] Step seven, the primer at the other end of the nucleic acid fragment is combined with the primer binding site at the other end of the nucleic acid fragment, and the complementary sequence of the sequence at the other end of the nucleic acid fragment is synthesized using the primer at the other end of the nucleic acid fragment as the starting point of synthesis, and the reading (sequencing) of the sequence at the other end of the nucleic acid fragment is completed through the complementary sequence.
[0049] At this point, the sequencing of the other end of the nucleic acid fragment is completed.
[0050] Example 2
[0051] like Figure 4 As shown, this embodiment provides a connector fixed on a chip, including a P5 connector sequence and a P7 connector sequence, wherein the P5 connector sequence includes two identification bases, namely an I base and a U base, wherein the I base is close to the 5' end of the P5 connector sequence and the U base is far from the 5' end of the P5 connector sequence. The I base in the P5 connector sequence can replace any C base in the P5 connector sequence in the prior art, and the U base in the P5 connector sequence can replace any T base in the P5 connector sequence in the prior art. The P7 connector sequence includes an identification base G*, and the G* base in the P7 connector sequence can replace any G base in the P7 connector sequence in the prior art.
[0052] The P5 and P7 adapter sequences immobilized on the chip in this example are compatible with, and completely complementary to, the adapter sequences of currently available mainstream libraries. Therefore, sequencing can be completed using the adapters provided in this example without modification or optimization of the existing mainstream library structures.
[0053] The sample library is connected to the connector fixed on the chip provided in this embodiment, and can be Figure 5 As shown ( Figure 5 Taking only one pair of complementary chains on a pair of P7 linker sequences and a P5 linker sequence as an example, there are actually tens of millions or even hundreds of millions of such complementary chains on the chip).
[0054] When tag sequencing is required first and then target fragment sequencing, the sequencing process can be as follows: Figure 6 Specifically, the following steps may be included:
[0055] Step 1: Use uracil DNA glycosylase and endonuclease VIII to remove the marker base U base away from the 5' end of the chain connected to the P5 linker sequence, denature and remove the chain connected to the P5 linker sequence, and leave the chain connected to the P7 linker sequence.
[0056] In step 2, the first tag primer is bound to the first tag primer binding site, and the complementary sequence of the first tag sequence is synthesized with the first tag primer as the synthesis starting point, and the reading (sequencing) of the first tag sequence is completed through the complementary sequence, and then the first tag primer and the complementary sequence of the first tag sequence are removed by denaturation.
[0057] Step 3: Use the P5 adapter sequence on the chip as the second tag primer, and use the second tag primer as the synthesis starting point to synthesize the complementary sequence of the second tag sequence. The second tag sequence is read (sequencing) through the complementary sequence. Then, human alkyladenine DNA glycosylase or 3-methyladenine DNA glycosylase II is used to remove the marker base I base near the 5' end of the chain connected to the P5 adapter sequence. Finally, the second tag primer and the complementary sequence of the second tag sequence are denatured and removed.
[0058] This completes tag sequencing.
[0059] Step 4: The primer at one end of the nucleic acid fragment is combined with the primer binding site at one end of the nucleic acid fragment, and the complementary sequence of the sequence at one end of the nucleic acid fragment is synthesized with the primer at one end of the nucleic acid fragment as the starting point of synthesis, and the reading (sequencing) of the sequence at one end of the nucleic acid fragment is completed through the complementary sequence, and then the primer at one end of the nucleic acid fragment and the complementary sequence of the sequence at one end of the nucleic acid fragment are removed by denaturation.
[0060] This completes the sequencing of one end of the nucleic acid fragment.
[0061] Step 5: Reconnect the sample library starting from the P5 adapter sequence on the chip to synthesize a new chain (ie, restore to the original complementary chain state).
[0062] Step 6: Use 8-oxoguanine DNA glycosylase to remove the marker base G* on the P7 linker sequence, denature and remove the chain connected to the P7 linker sequence, and leave the chain connected to the P5 linker sequence.
[0063] Step seven, the primer at the other end of the nucleic acid fragment is combined with the primer binding site at the other end of the nucleic acid fragment, and the complementary sequence of the sequence at the other end of the nucleic acid fragment is synthesized using the primer at the other end of the nucleic acid fragment as the starting point of synthesis, and the reading (sequencing) of the sequence at the other end of the nucleic acid fragment is completed through the complementary sequence.
[0064] At this point, the sequencing of the other end of the nucleic acid fragment is completed.
[0065] Example 3
[0066] In the joint provided in this embodiment,
[0067] The P5 linker sequence is: 5′-AATGATACGGCGACCACCGAGAUCTAIAC-3′;
[0068] The sequence of the P7 linker is: 5'-CAAGCAGAAGACGGCATACGAG*AT-3'.
[0069] After the nucleic acid fragment is amplified, sequencing is performed using the following steps: ① First, use AlkA enzyme (human alkyladenine DNA glycosylase (hAAG) or 3-methyladenine DNA glycosylase II) to cut the I base and sequence the first and second tags. ② Then use UDG (uracil DNA glycosylase) and Endo VIII enzyme (nuclease VIII) to cut the U base. ③ Perform 75-cycle read length (SE75) sequencing and generate a result report. ④ After waiting for 1.5 hours, complete the 100-cycle read length (SE100) sequencing and generate a result report. And so on, finally complete the 300-cycle read length (PE150) sequencing and generate a result report.
[0070] In this example, different sample libraries were mixed and loaded onto the sequencing machine, and split according to the SE75, SE100, SE150, PE100, and PE150 modes, respectively. Six experiments were conducted to test the performance parameters of different split modes, including Q20, Q30, and GC content. The sequencing results are shown in Table 1.
[0071] Table 1
[0072]
[0073] In the six experiments, the time to generate results for different read lengths was 7 hours and 30 minutes, 9 hours, 12 hours and 5 minutes, 20 hours and 30 minutes, and 23 hours and 50 minutes, respectively, and the average Q30 value was above 90%. The results are shown in Table 2.
[0074] Table 2
[0075]
[0076]
[0077] By mixing different libraries onto the machine, result reports can be output at each node, solving the current pain point of requiring the entire sequencing run to generate all reports, greatly improving sequencing efficiency. The statistical results of the difference in sequencing time between the present invention and conventional sequencing technologies are shown in Table 3.
[0078] Table 3
[0079] Read length Duration of the invention Regular duration Save time SE75 7h30min 23h30min 16h SE100 9h 23h30min 14h30min SE150 12h05min 23h30min 11h25min PE100 20h30min 23h30min 3h PE150 23h50min 23h30min -20min
[0080] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A connector fixed on a chip, characterized in that: It includes a P5 linker sequence and a P7 linker sequence, wherein the P5 linker sequence includes a first identifier base and a second identifier base, the first identifier base is located at the position of the T base, and the second identifier base is located at the position of the C base; the P7 linker sequence includes a third identifier base, and the third identifier base is located at the position of the G base; wherein T is thymine, C is cytosine, and G is guanine.
2. The connector fixed on a chip according to claim 1, wherein: The first marker base is U, the second marker base is I, and the third marker base is G*, wherein U is uracil, I is hypoxanthine, and G* is 8-oxoguanine.
3. The connector fixed on a chip according to claim 1, wherein: The P5 linker sequence is 5'-AATGATACGGCGACCACCGAGAUCTAIAC-3'; the P7 linker sequence is 5'-CAAGCAGAAGACGGCATACGAG*AT-3'.
4. The connector fixed on a chip according to claim 1, wherein: One of the first identifier base and the second identifier base is close to the 5' end of the P5 linker sequence.
5. The connector fixed on a chip according to claim 4, wherein: When the first identifier base is close to the 5' end of the P5 adapter sequence, the distance between the first identifier base and the 5' end of the P5 adapter sequence is the first distance, and the distance between the tag primer binding site at the corresponding position on the complementary chain of the P5 adapter sequence and the 5' end of the P5 adapter sequence is the second distance, and the first distance is less than or equal to the second distance.
6. The connector fixed on a chip according to claim 5, wherein: If the first distance is smaller than the second distance, the difference between the first distance and the second distance is greater than or equal to 2 bp.
7. The connector fixed on a chip according to claim 6, wherein: The length of the tag primer sequence bound to the tag primer binding site at the corresponding position on the complementary strand of the P5 adapter sequence is greater than or equal to 15 bp and does not overlap with the tag sequence.
8. The connector fixed on a chip according to claim 1, wherein: The third identification base is close to the 5' end or 3' end of the P7 linker sequence.
9. A method for sequencing a sample library, characterized in that: The sample library is connected to the linker fixed on the chip according to any one of claims 1 to 8, and when the first marker base is close to the 5' end of the P5 linker sequence, the method comprises the following steps: The second marker base is removed, and the chain connected to the P5 adapter sequence is denatured and removed, leaving the chain connected to the P7 adapter sequence; For the remaining chain, the first tag close to the 5' end of the P7 adapter sequence is first sequenced, and the first tag primer and the complementary sequence of the first tag sequence are removed by denaturation during sequencing; then the P5 adapter sequence fixed on the chip is used as a primer to sequence the second tag away from the 5' end of the P7 adapter sequence, and the first identification base is removed, and then the second tag primer and the complementary sequence of the second tag sequence are removed by denaturation during sequencing; then one end of the nucleic acid fragment is sequenced, and the primer at one end of the nucleic acid fragment and the complementary sequence at one end of the nucleic acid fragment are removed by denaturation.
10. A method for sequencing a sample library, characterized in that: The sample library is connected to the linker fixed on the chip according to any one of claims 1 to 8, and when the second marker base is close to the 5' end of the P5 linker sequence, the method comprises the following steps: The first marker base is removed, and the chain connected to the P5 adapter sequence is denatured and removed, leaving the chain connected to the P7 adapter sequence including the sample library; For the remaining chain, the first tag close to the 5' end of the P7 adapter sequence is first sequenced, and the first tag primer and the complementary sequence of the first tag sequence are removed by denaturation during sequencing; then the P5 adapter sequence fixed on the chip is used as a primer to sequence the second tag away from the 5' end of the P7 adapter sequence, and the second identification base is removed, and then the second tag primer and the complementary sequence of the second tag sequence are removed by denaturation during sequencing; then one end of the nucleic acid fragment is sequenced, and the primer at one end of the nucleic acid fragment and the complementary sequence at one end of the nucleic acid fragment are removed by denaturation.
11. The method for sequencing a sample library according to claim 9 or 10, wherein: The method further comprises the steps of: Reconnect the sample library starting from the P5 adapter sequence fixed on the chip; The third marker base is removed, and the chain connected to the P7 linker sequence is denatured and removed, leaving the chain connected to the P5 linker sequence; The remaining strand is then sequenced at the other end of the nucleic acid fragment.
12. The method for sequencing a sample library according to claim 11, wherein: If the first marker base is U, the second marker base is I, and the third marker base is G*, human alkyladenine DNA glycosylase or 3-methyladenine DNA glycosylase II is used to remove the I base, uracil DNA glycosylase and nuclease VIII are used to remove the U base, and 8-oxoguanine DNA glycosylase is used to remove the G* base.