Method for constructing a library of tagged sequence vectors by short oligonucleotides and use thereof

By using primer dimer technology with short oligonucleotide tag primers and universal primers, the problems of high cost and complex operation in vector library construction in existing technologies have been solved, achieving efficient, low-cost and high-throughput vector library construction, supporting gene editing and barcode tagging.

CN116254258BActive Publication Date: 2026-03-20FUJIAN AGRI & FORESTRY UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-03
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing technologies require the synthesis of long single-stranded oligonucleotides when constructing tag sequence vector libraries, which is costly, prone to introducing mutations, and cumbersome, making it difficult to achieve efficient, low-cost, and high-throughput vector library construction.

Method used

Using single-stranded oligonucleotide tag primers with a length of less than 60 nt and universal primers containing modified bases, primer dimers are formed by primer annealing and linked to a linearized vector. The tag sequence is inserted using the base excision repair mechanism of E. coli, simplifying the operation process.

Benefits of technology

It significantly reduces primer synthesis costs, improves operational fidelity and efficiency, enables high-throughput, low-cost vector library construction, and allows tag sequences to be used for gene editing and barcode labeling, supporting efficient genetic engineering research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116254258B_ABST
    Figure CN116254258B_ABST
Patent Text Reader

Abstract

The application discloses a method for constructing a tag sequence-containing vector library by short oligonucleotides and application thereof. The method needs two oligonucleotides, i.e. a customized tag primer and a universal primer which can be used for all library constructions. The middle of the tag primer is a tag sequence, and both sides are annealing sequences. The universal primer has sequences complementary to the annealing sequences of the tag primer on both sides, and a modified base which can be recognized as a base damage by a host microorganism in the middle. The tag primer library is annealed with an equal amount of the universal primer to obtain primer dimers. The primer dimers have base-paired DNA double strands on both sides, and are cohesive ends. The middle of the primer dimers is an unpaired omega loop structure, and the double strands of the omega loop contain the tag information carried by the tag primer and the modified base carried by the universal primer respectively. The primer dimers are connected with linearized vectors, and then are transformed into E. coli, so that the region of the modified base of the omega loop is replaced by the tag sequence through the base excision repair mechanism of the endogenous cells.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of genetic engineering, and particularly relates to a method for constructing a tag sequence-containing vector library by short oligonucleotides and application thereof. BACKGROUND

[0002] Cas protein binds to target DNA depending on the base complementarity of gRNA, and only needs to simply change the gRNA sequence obtained by transcription to customize new target points. By using CRISPR / Cas gene editing technology, a gRNA library targeting each member of a gene family or a signal pathway is customized, and a high-throughput, large-scale directed mutant library is created, which is an important means of functional genomics and biological breeding. Through modification of Cas protein, change of host DNA damage mode and repair pathway, high-throughput CRISPR technology can also realize the creation of mutant libraries such as base editing, transcription activation and transcription inhibition. The fundamental prerequisite for realizing this technology is to be able to simply, efficiently and low-costly customize gRNA vector library.

[0003] DNA barcode refers to a short nucleotide sequence on DNA, which can distinguish different sample sources. DNA barcode can be used for high-throughput mixed sequencing, which can reduce the cost of sequencing while ensuring that different molecules do not interfere with each other. By using DNA barcode to label different cells, the cell lineage can be tracked for developmental biology research.

[0004] In the prior art, a tag sequence is introduced into a vector, and first, a long single-stranded oligonucleotide (~100 nt) is synthesized, with the tag sequence in the middle and the two ends being conservative sequences. Then, by using PCR technology, a double-stranded DNA fragment is amplified by using universal primers designed on the conservative sequences. Finally, by enzyme digestion / ligation or homologous recombination technology, the double-stranded DNA fragment is inserted into the vector. The length of the single-stranded oligonucleotide to be synthesized is positively correlated with the cost and error rate, and the commercial synthesis cost of each base of oligonucleotide above 60 nt (including 60 nt) is several times higher than that of oligonucleotide below 59 nt (including 59 nt). Some technologies can use shorter primers, but still need PCR to obtain double-stranded DNA fragments. PCR technology can introduce mutations, and the recovery rate of short fragment amplification products is low. The prior art also uses positive and negative primer annealing to obtain double-stranded DNA fragments, but the positive and negative primers of different tags need to be independently annealed, and the operation of constructing a tag vector library is complicated, and low-cost degenerate primers cannot be used as tag primers. SUMMARY

[0005] In order to solve the above problems, the application provides a new method for constructing a vector library containing a tag sequence and application, wherein the tag sequence can be used as a target of a gRNA library, so that the application can be used for creating a high-throughput gene editing vector library. The tag sequence can be used as a barcode of vector DNA, so that the application can be used for creating a barcode vector library, and the barcode can be given to a genetically engineered cell obtained by using the vector library. If the tag sequence is located in a protein coding region, a vector library for site-directed saturation mutation of a protein can be obtained.

[0006] In order to solve the above problems, the application adopts the technical scheme of:

[0007] A method for constructing a vector library containing a tag sequence by using a short oligonucleotide, comprising:

[0008] 1) preparing a universal primer that can be used for all library construction; that is, under the premise of a vector skeleton and linearization mode, preparing a single-stranded oligonucleotide that can be used for any type of vector library (the same linearized vector).

[0009] 2) customizing a tag primer containing a preset tag sequence (about 20 nt) in the middle;

[0010] The tag primer is a single-stranded oligonucleotide with a length of less than 60 nt, the middle of which is a tag sequence to be introduced into a vector, and the two sides of which are annealing sequences (about 15 nt).

[0011] The universal primer is a single-stranded oligonucleotide (about 40 nt), the middle of which is one or more modified bases that can be recognized by a host microorganism as base damage, and the two sides of which are homologous sequences of the annealing sequences of the tag primer; the two sides of the tag primer and the universal primer can be complementary to each other, so that a primer dimer, that is, a double-stranded DNA short fragment, can be formed.

[0012] 3) annealing the tag primer and an equal amount of the universal primer through the homologous sequences to obtain a primer dimer;

[0013] The two sides of the primer dimer are base-paired DNA double strands, and the two ends of the primer dimer are cohesive ends; the middle of the primer dimer is an unpaired loop structure, which is in the shape of “Ω”; the double strands of the “Ω” loop structure respectively contain the tag sequence carried by the tag primer and the modified base carried by the universal primer.

[0014] 4) preparing a linearized vector with two cohesive ends compatible with the cohesive ends of the primer dimer;

[0015] 5) connecting the primer dimer and the linearized vector skeleton into a recombination molecule by using T4 DNA ligase; wherein the primer dimer can be connected to the linearized vector through the matching of the cohesive ends.

[0016] 6) Transform E. coli, replace the region containing the modified base of the omega loop with the tag sequence by the cell's endogenous base excision repair mechanism.

[0017] 7) Pick single clone, shake bacteria to extract plasmid, sequence verification.

[0018] Further, the sticky ends at both ends of the primer dimer are incompatible; by designing the tag primer and the universal primer, the two ends of the primer dimer will form incompatible sticky ends (if the sticky ends are compatible, the efficiency will be greatly reduced).

[0019] Preferably, the tag sequence of the tag primer is the targeting sequence of the genomic DNA of the organism; it can be used to construct a: vector library that causes double-strand breaks at a specified site in the genomic DNA of the cell; b: vector library that causes single-strand breaks at a specified site in the genomic DNA of the cell; c: vector library that causes base damage at a specified site in the genomic DNA of the cell; d: vector library that recruits transcriptional regulators (including activators and inhibitors) at a specified site in the genomic DNA of the cell.

[0020] Preferably, the tag sequence of the tag primer is a degenerate primer containing degenerate bases; the tag sequence is located in the protein coding region, and the degenerate primer is used to introduce mutations at a specific site of the protein, which can be used for the construction of protein / peptide library such as site-directed saturation mutation library and phage display library.

[0021] Further, the linearized vector is usually obtained by double digestion of a circular plasmid with Type II restriction endonuclease or single digestion with Type IIS restriction endonuclease, and its two ends are incompatible (if the sticky ends are compatible, the efficiency will be greatly reduced) and compatible with the sticky ends of the primer dimer.

[0022] The application also provides a method for identifying the identity of a transgenic organism, comprising: the vector library obtained by the above method, identifying each member in the library through the tag sequence.

[0023] Further, the specific steps are: after stably transforming the cells with the vector library obtained by the above method, extracting the DNA of the transformants, and identifying the tag sequence integrated on its genome by sequencing, thereby giving each independent transformant a unique identity and obtaining information such as the copy number of the transgene.

[0024] The application has the following beneficial effects:

[0025] Compared with most of the existing technologies which need to customize more than 100 nt, the present application only needs to customize a single-stranded tag primer with very short length (<60 nt) for each tag sequence, which can significantly reduce the cost of primer synthesis. Compared with most of the existing technologies which need PCR, the present application uses primer annealing to obtain double-stranded DNA carrying tag sequence information, which is more faithful and will not affect the library construction efficiency due to the difficulty in recovering small fragments. Compared with the existing technologies which need to anneal each pair of primers separately, the present application uses universal primers containing modified bases, which simplifies the operation method and also realizes the demand for low-cost barcode preparation using degenerate primers. BRIEF DESCRIPTION OF DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0027] Figure 1 The principle diagram of constructing a tag sequence-containing vector library by short oligonucleotides according to the present application.

[0028] Figure 2 The secondary structure of primer duplex in the embodiments;

[0029] In the formula, the degenerate base and the modified base are underlined; the lowercase letters represent the annealing sequence, which is the downstream sequence of the rice U3 promoter and the upstream sequence of the sgRNA scaffold in this embodiment, to realize seamless insertion of the tag sequence. The annealing sequence needs to be redesigned for different linearized vectors; the bold letters represent the cohesive end.

[0030] Figure 3 The vector map of plasmid pGN-Cas9-U3’=Amp-SacB=scaffold’ in the embodiments;

[0031] Wherein, the white font represents the essential elements for ensuring the propagation of the plasmid in E. coli and Agrobacterium, and the black font represents the T-DNA region of the binary vector. LB and RB are the left border and the right border of the binary vector, respectively, Pe35S, HygR and T35S are the elements for conferring the resistance of plants to hygromycin, PZmUbi, Cas9 and TNOS are the elements for conferring the gene editing function of plant cells, POsU3' and Scaffold' are the rice U3 promoter and sgRNA scaffold respectively with the downstream sequence and the upstream sequence missing, and the missing sequences will be subsequently filled by the primer duplex. AmpR and SacB are the positive selection marker and the negative selection marker, respectively, and are used to reduce the mutation rate of vector propagation and to improve the positive rate of subsequent molecular cloning. The BsaI enzyme digestion site is the position for linearizing the plasmid.

[0032] Figure 4 DNA sequence of the cohesive end region of the linearized vector pGN-Cas9-U3' = Amp-SacB = scaffold' in the example;

[0033] Wherein, the BsaI enzyme digestion recognition site is located on the replaced fragment.

[0034] Figure 5 Growth of E. coli colonies transformed with the vector library in the example on kanamycin LB plates;

[0035] The results show that the efficiency of the application is very high and has use value.

[0036] Figure 6 Vector map of each member of the CRISPR / Cas9 genome editing vector library constructed using mixed tag primers in the example;

[0037] Wherein, the sequences introduced by the primer duplex, the rice U3 promoter and the sgRNA are both restored to be complete.

[0038] Figure 7 Sequencing verification results of 7 randomly selected clones in the example. DETAILED DESCRIPTION

[0039] In order to make the objects, technical solutions and advantages of the embodiments of the application clearer, the technical solutions in the embodiments of the application will be described clearly and completely below with reference to the embodiments of the application. Obviously, the described embodiments are part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the application.

[0040] Refer to the drawings Figure 1As shown, the method for constructing a tag sequence-containing vector library by short oligonucleotides according to the present application comprises the following steps:

[0041] 1) Prepare a universal primer in advance. The middle of the universal primer contains a modified base (such as dU) that can be recognized by the host microorganism as a base lesion. The two sides of the universal primer are each ~15 nt annealing sequences. The length of the annealing sequence is not fixed, but too long will increase the cost, and too short may affect the subsequent annealing, thereby affecting the efficiency. According to experimental determination, the Tm value of the annealing sequence at each end should be above 45℃ to ensure the efficiency of the method. The Tm values of the annealing sequences at both ends should be as close as possible. The specific sequence of the annealing sequence is determined by the vector sequence and the selected restriction enzyme for circularizing the vector. When the linearized vector is certain, the universal primer can be reused. Dilute the universal primer with sterile double distilled water to a concentration of 100 μM.

[0042] 2) According to the different needs of different vector libraries, customize single-stranded oligonucleotides (<60 nt) containing a preset tag sequence (~20 nt) in the middle, named tag primer. The two ends of the tag primer are annealing sequences that can anneal with the universal primer, so the annealing sequences on the two primers are complementary to each other. Mix all the tag primers equally to obtain a tag primer mixture. The concentration of the tag primer mixture also needs to be diluted to 100 μM.

[0043] 3) Configure a 10 μL reaction system containing 1 μL of tag primer (100 μM), 1 μL of universal primer, 1 μL of 10×T4 DNA ligase Buffer (containing 10 mM ATP or using 10 mM ATP solution instead), and 0.5 μL of T4 polynucleotide kinase (NEB, M0201), and finally 5.5 μL of sterile double distilled water to make up. Set the reaction program on the PCR instrument as follows: 37℃ for 30 min (T4 polynucleotide kinase transfers the phosphate group of ATP to the 5' end of the primer), 65℃ for 20 min (inactivate T4 polynucleotide kinase), 4℃ for 1 min, 95℃ for 5 min, gradually cool to 25℃ (5℃ / min), and finally dilute 200 times with sterile double distilled water. At this time, the product is 0.05 μM of 5' phosphorylated primer dimer. The 5' and 3' ends of conventional commercial oligonucleotides are both hydroxyl groups. The 5' hydroxyl group of the primer dimer formed by direct annealing will introduce one nick on each DNA single strand after being connected with the linear vector, affecting the subsequent expected base excision repair. Therefore, T4 polynucleotide kinase must be used in this step to obtain 5' phosphorylated primer dimer, thereby improving the efficiency of subsequent ligation and the efficiency of introducing information carried by the tag primer into the vector. The middle of the primer dimer is a non-paired loop structure, showing an "Ω" shape; the double strands of the "Ω" loop structure contain the tag sequence carried by the tag primer and the modified base carried by the universal primer, respectively; the two ends of the primer dimer form incompatible cohesive ends.

[0044] 4) Preparation of linearized vector. It is recommended to use double digestion of Type II restriction enzyme or single digestion of Type IIS restriction enzyme, which has sticky ends incompatible with each other (if the cohesive ends are compatible, the efficiency will be greatly reduced) and compatible with the cohesive ends of the primer duplex. Take the Bsal enzyme digestion reaction as an example, configure a 50 μL reaction system: containing 1 μg of vector DNA, 5 μL of rCutSmart Buffer and 1 μL of Bsal (NEB, R3733), and finally add sterile double distilled water to make up. Reaction at 37°C for 15 min to overnight. It is recommended to recover the linearized vector after digestion. If not recovered, Type IIS restriction enzyme or the replaced fragment on the vector as a lethal sequence needs to be selected, otherwise it will affect the positive rate due to the high self-connection rate; and after the completion of the enzyme digestion reaction, the restriction enzyme needs to be inactivated by high temperature or proteinase K to avoid affecting the subsequent connection. When the linearized vector is constant, the universal primer is fixed and reusable, and the annealing sequence of the tag primer is fixed.

[0045] 5) Ligation reaction. Configure a 20 μL reaction system: containing 1 μL of linearized vector (~20 ng / μL), 1 μL of primer duplex (0.05 μM), 1 μL of 10 x T4 DNA ligase Buffer and 1 μL of T4 DNA ligase, and finally add sterile double distilled water to make up. Reaction at 25°C for 1 h or at 16°C overnight.

[0046] 6) Transform the ligation product into E. coli competent cells. After thawing the E. coli competent cells in ice bath, add the reaction solution (not more than 1 / 10 proportion) and mix gently, ice bath for 30 min; 42°C water bath heat shock for 60 s, quickly put back into ice and stand for 2 min, this step cannot shake the centrifuge tube; add 500-900 μL of LB medium to the centrifuge tube and mix, recover the bacteria for 45-60 min at 37°C, 180 rpm; take 100-200 μL of bacterial solution and spread on solid LB plates with corresponding resistance; after the plate is dry, seal it, invert it in a 37°C incubator, and culture overnight (~15 h).

[0047] 7) Select single clones, shake bacteria to extract plasmids, and verify by sequencing.

[0048] Example 1

[0049] ​Tag primers (Tag-f) and universal primers (Uni-r) were synthesized separately. The tag primers in this embodiment are degenerate primers, where W represents A / T, K represents T / G, Y represents T / C, D represents A / T / G, R represents A / G, S represents G / C, N represents A / T / G / C, V represents A / G / C, H represents A / T / C, M represents A / C, and B represents T / G / C. The primer sequences are shown in Table 1. Therefore, tag primer Tag-f is actually an oligonucleotide library containing 20736 members. The universal primer in this embodiment selected U as the modifying base, which can be recognized by the *E. coli* host as a deamination injury of base C, indicated by underline.

[0050] Table 1. Primer sequences used to construct the vector library

[0051] Name Sequence Tag-f <![CDATA[TGTGCAGATGATCCGTGGCA GWCKYGDARTSNVAHAMCBT GTTTCAGAGCTAGAAATA]]> Uni-r TTGCTATTTCTAGCTCTGAAAC U TGCCACGGATCATCTG]]>

[0052] Prepare a 10 μL reaction mixture containing 1 μL Tag-f (100 μM), 1 μL Uni-r (100 μM), 1 μL 10×T4 DNA ligase buffer, and 0.5 μL T4 polynucleotide kinase (NEB, M0201), with 5.5 μL of sterile double-distilled water added to make up the difference. Set the reaction program on the PCR instrument as follows: 37℃ for 30 min, 65℃ for 20 min, 4℃ for 1 min, 95℃ for 5 min, then gradually cool to 25℃ (5℃ / min). Finally, dilute 200-fold with sterile double-distilled water; the product at this point is 0.05 μM of 5' phosphorylated primer dimer. Figure 2 ).

[0053] The plasmid pGN-Cas9-U3'=Amp-SacB=scaffold' was digested with the Type IIS restriction enzyme BsaI. Figure 3 Prepare a 50 μL reaction system containing 1 μg pGN-Cas9-U3'=Amp-SacB=scaffold', 5 μL rCutSmartBuffer, and 1 μL... (NEB, R3733), and finally made up the difference with sterile double-distilled water. The PCR instrument was set to 37℃ for 1 hour, and no recovery was required to obtain a linearized vector concentration of 20 ng / μL. Figure 4 ).

[0054] Prepare a 20 μL reaction mixture containing 1 μL linearized vector (~20 ng / μL), 1 μL primer dimer (0.05 μM), 1 μL 10×T4 DNA ligase buffer, and 1 μL T4 DNA ligase. Top up with sterile double-distilled water. Incubate at 25°C for 1 h.

[0055] 100 μL E. coli competent cells were thawed on ice, 10 μL reaction solution was added and mixed gently, and then the mixture was placed on ice for 30 min; the mixture was heated at 42℃ for 60 s, quickly placed on ice and stood for 2 min; 900 μL LB medium was added to the centrifugal tube and mixed, and then the bacteria were recovered at 37℃ and 180 rpm for 45 min; 150 μL of the bacterial solution was spread on an LB plate containing kanamycin (50 mg / L) and sucrose (5%); after the plate was dried, it was sealed, inverted in a 37℃ incubator, and cultured overnight. Figure 5

[0056] The expected map of each member of the vector library is shown in Figure 6 Seven single clones were randomly selected, and Sanger sequencing was used to verify that they met the expectations. Figure 7

[0057] The present application can establish a vector library containing different tag sequences with high fidelity, high efficiency and high throughput. Since the universal primer can be reused, only the single-stranded tag primer needs to be customized for new library construction, and the length is very short. The present application can also significantly reduce the cost of primer synthesis for library construction. The tag sequence can be a gene editing target to achieve high-throughput gene editing. The tag sequence can be used as a DNA barcode to customize the polymorphism of the vector and achieve the identification of transgenic organisms. The tag sequence can also be used to construct a site-directed saturation mutation library.

[0058] The above examples are only used to illustrate the technical solutions of the present application, but not to limit it; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing examples, or make equivalent substitutions for part of the technical features; and these modifications or substitutions do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.​​

Claims

1. A method for constructing a tag sequence vector library using short oligonucleotides, characterized in that, include: 1) Prepare universal primers that can be used for all library constructions; 2) Customize the tag primer containing a preset tag sequence in the middle; The tag primer is a single-stranded oligonucleotide less than 60 nt in length, with the tag sequence to be introduced into the vector in the middle and annealing sequences on both sides. The universal primer is a single-stranded oligonucleotide, with one or more uracil or other modified bases that can be recognized as base damage by the host microorganism in the middle, and homologous sequences of the tag primer annealing sequence on both sides; the sequences on both sides of the tag primer and the universal primer can complement each other; 3) Obtain primer dimers by annealing the tag primers and an equal amount of universal primers between homologous sequences; The primer dimer consists of two base-paired DNA double strands on both sides, with sticky ends at both ends; the middle of the primer dimer is an unpaired circular structure in the shape of an "Ω"; the double strands of the "Ω" circular structure contain the tag sequence carried by the tag primer and the modified base carried by the universal primer, respectively; the sticky ends at both ends of the primer dimer are incompatible; 4) Prepare linearized vectors with two sticky ends compatible with the sticky ends of primer dimers; 5) The primer dimer and the linearized vector backbone were ligated into a recombinant molecule using T4 DNA ligase; 6) Transform E. coli and replace the region containing the modified bases of the Ω-loop with the tag sequence through the endogenous base excision repair mechanism; 7) Select single clones, extract plasmids by shaking, and verify by sequencing.

2. The method for constructing a tag sequence vector library using short oligonucleotides according to claim 1, characterized in that, The tag sequence of the tag primer is the target sequence of the organism's genomic DNA.

3. The method for constructing a tag sequence vector library using short oligonucleotides according to claim 1, characterized in that, The tag sequence of the tag primer is a degenerate primer containing degenerate bases.

4. The method for constructing a tag sequence vector library using short oligonucleotides according to claim 1, characterized in that, The linearized vector is obtained by double digestion of a circular plasmid with a Type II restriction endonuclease or single digestion with a Type IIS restriction endonuclease, and its two ends are incompatible sticky ends that are compatible with the sticky ends of the primer dimer.

5. An application of the method as described in claim 2 in constructing a, b, c, and d; in, a: A vector library that creates double-strand breaks at specified sites in the cellular genomic DNA; b: Vector libraries that induce single-strand breaks at specified sites in the cellular genomic DNA; c: Vector libraries that cause base damage at specified sites in the cell's genomic DNA; d: A vector library that recruits transcriptional regulatory factors at designated sites on the cellular genomic DNA.

6. The application according to claim 5, characterized in that, The regulatory factors include activators and inhibitors.

7. The application of the method as described in claim 3 in the construction of protein / peptide libraries.

8. A method for identifying genetically modified organisms, characterized in that, The vector library obtained by any one of claims 1-4 is used to identify each member in the library by tag sequence.

9. The method for identifying transgenic organisms according to claim 8, characterized in that, The specific steps are as follows: After the vector library obtained by any one of claims 1-4 is stably transformed into cells, the DNA of the transformants is extracted, and the tag sequence integrated into their genome is identified by sequencing.

Citation Information

Patent Citations

  • Method for generating higher order genome editing libraries

    CN110249049A