De novo synthesis method of genome DNA (Deoxyribose Nucleic Acid) of oversized fragment human highly repetitive

Through the assembly of yeast lithium acetate and CRISPR-mediated in vivo yeast assembly technology, the assembly problem of highly repeatable ultra-large fragments of DNA is solved, and efficient and accurate gene composition is achieved, which is suitable for the fields of synthetic biology and biopharmaceuticals.

CN120555480APending Publication Date: 2025-08-29TIANJIN UNIV SYNTHETIC BIOLOGY FRONTIER RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510847756.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

In the prior art, when synthesizing super-large fragments of DNA containing highly repeat sequences, there are problems such as difficult assembly, high error rate and poor stability. Especially in Saccharomyces cerevisiae, DNA breakage and fragment loss are prone to occur, making it difficult to achieve accurate genome assembly.

Method used

The yeast lithium acetate assembly method is used to assemble small fragments by adding homologous arms, and the in vivo assembly technology mediated by yeast sexual reproduction and CRISPR-mediated in vivo assembly is used to realize the synchronous assembly of multiple small fragments and the splicing of large fragments, reduce misassembly assembly, and use the eukaryotic characteristics of yeast to improve stability.

Benefits of technology

It significantly improves the synthesis accuracy and stability of highly repeat genomic DNA, reduces the probability of mismatch, expands the scope of application of gene combinatorial, and is suitable for the fields of synthetic biology, gene therapy and biopharmaceuticals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention relates to the technical field of biology, in particular to a de novo synthesis method of genome DNA of a super-large-fragment human highly repetitive sequence. The invention provides a strategy for gradually synthesizing and assembling the Mb-level genome DNA fragment with the human highly repetitive sequence from a small fragment, a medium-length fragment, a large fragment to a super-large fragment, the strategy is simple, convenient, efficient and accurate, and the problem of synthesizing the super-large fragment DNA containing the highly repetitive sequence gene in the prior art is solved; accurate assembly of complete chromosomes of saccharomyces cerevisiae is achieved, research on genomes containing complex repetitive sequences is promoted, and the method has wide application prospects in the fields of synthetic biology, gene therapy, biological pharmacy and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of biotechnology, and in particular to a de novo assembly and synthesis strategy for Mb-level genomic DNA fragments with highly repetitive human sequences. Background Art

[0002] As an emerging interdisciplinary subject, synthetic biology has broken the limitations of natural evolution. With DNA assembly technology as its core, it has shown great application potential in the fields of biology, medicine, new materials, etc. DNA assembly technology is a key technology in the field of synthetic biology in recent years. It plays a decisive role in the construction of large metabolic pathways and genomes. It is mainly divided into two methods: in vivo assembly and in vitro assembly. Small fragments of DNA are usually assembled in vitro due to their easy operation and high stability. However, the existing various in vitro DNA assembly methods are difficult to assemble and have cumbersome operation procedures when faced with DNA molecules of several hundred kb. In addition, most methods still need to be transferred into the host for cloning and expression after the in vitro assembly is completed, which greatly limits the application of in vitro assembly technology. Therefore, the assembly of large fragments of DNA relies more on the in vivo assembly of microorganisms with efficient homologous recombination capabilities.

[0003] Currently, yeast Saccharomyces cerevisiae and Escherichia coli are commonly used host bacteria for in vivo DNA assembly. Saccharomyces cerevisiae possesses efficient homologous recombination capabilities, offering significant advantages when assembling very large DNA fragments, such as microbial genomes. However, when assembling DNA fragments as large as 300 kb, DNA breakage is prone to occur. When cloning fragments rich in repetitive sequences, fragment loss often occurs, leading to increased synthesis errors. While E. coli boasts easy cultivation, simple molecular manipulation, and a low homologous sequence requirement, resulting in a short DNA assembly cycle and ease of transfer, its genetic material is free DNA, making stable inheritance impossible.

[0004] The artificial synthesis of the Saccharomyces cerevisiae genome project (Sc2.0) is dedicated to the redesign and chemical reconstruction of the entire genome of Saccharomyces cerevisiae. Its core goal is to achieve the artificial full synthesis of the 12 Mb Saccharomyces cerevisiae genome. At present, 8.5 complete artificial synthetic chromosomes have been successfully designed and synthesized. In this process, genome assembly is a crucial link. However, because Saccharomyces cerevisiae is a eukaryotic organism, the genome fragments are long and assembly is extremely difficult. At present, for the assembly of Saccharomyces cerevisiae chromosomes, short fragments are mostly constructed in vitro and replicated and enriched by Escherichia coli, and slightly longer fragments rely on the yeast endogenous homologous recombination mechanism and adopt a hierarchical assembly method. However, these methods have problems such as complex assembly process and easy loss of fragments. Especially for regions containing highly repetitive sequences, mismatches occur frequently, which makes it difficult to meet the needs of accurately synthesizing complete chromosomes. Specifically:

[0005] (1) Mismatches caused by highly repetitive sequences: In the study of de novo synthesis of human genomic DNA, the synthesis of very large DNA fragments is a key step. Currently, conventional synthesis of very large DNA fragments often uses the eSwAP-IN homology arm connection method (see Figure 1 However, when faced with genes containing highly repetitive sequences in the genome, existing technologies have exposed obvious flaws. For example, the 276 kb SynB region in the human genome contains more than 68 similar gene fragments. When synthesized using traditional homology arm ligation, these highly similar repetitive sequences are very likely to mismatch, resulting in errors in the entire gene synthesis (see Figure 2 、 Figure 3 ).

[0006] (2) There are defects in the selection of chassis organisms: Existing in vivo assembly technologies mainly rely on model organism chassis such as Escherichia coli, Bacillus subtilis, and Saccharomyces cerevisiae. In the E. coli system, although researchers have developed a series of systems for constructing large DNA fragments through the Lambda-red system, such as the REXER and CONEXER systems developed by Jason Chin's team combined with CRISPR / Cas9 technology to achieve the assembly of a 1.1 Mb human genome fragment in E. coli, E. coli is a prokaryotic organism with exposed circular DNA as its genetic material and lacks a nucleus. When processing complex genomes, it cannot well preserve high-level structures, which is not conducive to the stable synthesis of genes containing highly repetitive sequences.

[0007] Therefore, a simple, efficient and accurate method is urgently needed to solve the difficulties encountered by existing technologies in synthesizing very large fragments of DNA containing highly repetitive sequence genes, so as to achieve the precise assembly of the complete chromosome of Saccharomyces cerevisiae. Summary of the Invention

[0008] In view of this, the present application provides a de novo assembly synthesis strategy for Mb-level genomic DNA fragments with highly repetitive human sequences. The yeast lithium acetate assembly method is used to assemble small fragments, increase the homology arms, and compress the number of fragments by >10 times (>10 fragments per assembly), thereby achieving synchronous assembly of multiple small fragments and reducing the occurrence of incorrect assembly; protoplast transformation is used to achieve splicing and assembly of multiple large fragments (40~100 kb); yeast sexual reproduction and CRIPSR-mediated in vivo assembly are used to avoid the difficulties of easy breakage during in vitro extraction and low efficiency of re-introduction into cells.

[0009] In order to achieve the above-mentioned invention objectives, this application provides the following technical solutions:

[0010] The present application provides a method for de novo synthesis of genomic DNA of ultra-large fragments of highly repetitive human sequences, including: large-fragment DNA assembly and ultra-large-fragment DNA assembly;

[0011] The steps of assembling large DNA fragments include:

[0012] Step (i): Synthesize a length of 2~8 kb (can be 2,000 bp, 2,500 bp, 3,000 bp, 3,500 bp, 4,000 bp, 4,500 bp, 5,000 bp, 5,500 bp, 5,600 bp, 5,700 bp, 5,800 bp, 5,900 bp, 6,000 bp, 6,100 bp, 6,200 bp, 6,300 bp, 6,400 bp, 6,500 bp, 6,600 bp, 6,700 bp, 6,800 bp, 6,900 bp, 7,000 bp, 7,100 bp, 7,200 bp, 7,300 bp, 7,400 bp, 7,500 bp, 7,600 bp, bp, 7,700 bp, 7,800 bp, 7,900 bp or 8,000 bp), wherein the 3' end of each primary DNA fragment contains a 40-3,000 bp (which can be 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, 200 bp, 300 bp, 400 bp, 410 bp, 420 bp, 430 bp, 440 bp, 450 bp, 460 bp, 470 bp, 480 bp, 490 bp, 500 bp, 510 bp, 520 bp, 530 bp, 540 bp, 550 bp, 560 bp, 570 bp, 580 bp, 590 bp, 600 bp, 700 bp) sequence with the 5' end of the next primary DNA fragment. bp, 800 bp, 900 bp, 1,000 bp, 2,000 bp, or 3,000 bp) homology arms with identical base sequences;

[0013] Step (ii): the adjacent 8 to 15 primary DNA fragments in step (i) are respectively transferred into N chromosomes of Saccharomyces cerevisiae cells, and the primary DNA fragments are connected through homology arms in the Saccharomyces cerevisiae cells to form multi-level DNA fragments;

[0014] Step (iii): extracting the multi-level DNA fragments from step (ii);

[0015] Step (iv): using the multi-level DNA fragments from step (iii) as new primary DNA fragments, and repeating steps (ii) and (iii) until a large DNA fragment with a length of 300 to 800 kb is formed;

[0016] The steps of assembling the super-large fragment DNA include:

[0017] Step (v): transferring the large DNA fragments from step (iv) into a pair of Saccharomyces cerevisiae chromosomes in groups of two, wherein the pair of Saccharomyces cerevisiae chromosomes contain a Cas9 expression plasmid and an sgRNA expression plasmid, respectively;

[0018] Step (vi): mating the pair of Saccharomyces cerevisiae in step (v) to obtain a diploid, wherein the sgRNA in the diploid binds to a specific site on the Saccharomyces cerevisiae chromosome, mediating Cas9 to perform cutting, and homologous recombination occurs on the Saccharomyces cerevisiae chromosome in the Saccharomyces cerevisiae to obtain an ultra-large DNA fragment;

[0019] Step (vii): extracting the oversized DNA fragments from step (vi);

[0020] Step (viii): Using the ultra-large DNA fragment as a new large DNA fragment, repeat steps (v) to (vii) until a complete ultra-large human highly repetitive sequence genomic DNA is obtained.

[0021] In some specific embodiments of the present application, step (ii) of the de novo synthesis method is as follows: the adjacent 8 to 15 primary DNA fragments are grouped together, and each group is transferred together with linearized pCC1BAC into Saccharomyces cerevisiae for homologous recombination to form an artificial chromosome;

[0022] Step (iii) is: enzymatically cutting the artificial chromosome to remove the skeleton and separate the multi-level DNA fragments.

[0023] In some specific embodiments of the present application, the transfer in step (ii) of the above-mentioned de novo synthesis method includes transfer via a lithium acetate-mediated transformation method.

[0024] In some specific embodiments of the present application, the Saccharomyces cerevisiae cell of step (ii) of the above de novo synthesis method includes BY4741.

[0025] In some specific embodiments of the present application, the pair of Saccharomyces cerevisiae in step (v) of the above de novo synthesis method are VL6-48 and VL6-48A, respectively.

[0026] In some specific embodiments of the present application, in step (v) of the above-mentioned de novo synthesis method, one of the pair of Saccharomyces cerevisiae contains only a Cas9 expression plasmid, and the other contains only a sgRNA expression plasmid.

[0027] In some specific embodiments of the present application, step (v) of the above-mentioned de novo synthesis method is: transferring the large DNA fragment and the artificial chromosome backbone (PRS414 or PRS416) into Saccharomyces cerevisiae by protoplast transformation method to form an artificial chromosome by homologous recombination;

[0028] Step (vi) is: mating the pair of Saccharomyces cerevisiae in step (v) to obtain a diploid, wherein the sgRNA in the diploid binds to a specific site of the artificial chromosome, mediating Cas9 to perform cutting, and the artificial chromosome undergoes homologous recombination in the Saccharomyces cerevisiae to obtain an ultra-large fragment of DNA.

[0029] In some specific embodiments of the present application, the sequence of the sgRNA that recognizes the specific site in the above-mentioned de novo synthesis method is shown as SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7 or SEQ ID NO: 8.

[0030] In some specific embodiments of the present application, the sequence of the site recognized by the Cas9 in the above-mentioned de novo synthesis method is shown as SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3 or SEQ ID NO: 4.

[0031] The present application also provides a method for assembling an ultra-large fragment of human highly repetitive sequence genomic DNA (which can be 0.8-1.2 Mb, 0.9-1.1 Mb or 1 Mb), comprising:

[0032] Saccharomyces cerevisiae A and Saccharomyces cerevisiae B of different mating types are mated to obtain a diploid. The sgRNA in the diploid binds to specific sites on artificial chromosomes A and B, mediating Cas9 cutting. The artificial chromosomes A and B undergo homologous recombination in Saccharomyces cerevisiae to obtain an ultra-large fragment of genomic DNA with highly repetitive human sequences.

[0033] In some specific embodiments of the present application, the above-mentioned assembly method includes:

[0034] Saccharomyces cerevisiae A and Saccharomyces cerevisiae B of different mating types are mated. The mating type of Saccharomyces cerevisiae A or Saccharomyces cerevisiae B is not specifically limited, as long as the mating types are different and can mate. After mating, a diploid is obtained. The sgRNA in the diploid binds to specific sites (i.e., the target sites of the sgRNA) on the artificial chromosomes A and B, mediating Cas9 cleavage. The artificial chromosomes A and B undergo homologous recombination in the Saccharomyces cerevisiae, thereby obtaining an ultra-large fragment of genomic DNA containing highly repetitive human sequences.

[0035] The artificial chromosome A and sgRNA are from Saccharomyces cerevisiae A, and the artificial chromosome B and Cas9 are from Saccharomyces cerevisiae B;

[0036] The sgRNA is expressed by an sgRNA expression plasmid, and the Cas9 is expressed by a Cas9 expression plasmid;

[0037] The artificial chromosome A has a synthetic segment A of 500,000 to 600,000 bp (which may be 500,000 bp, 510,000 bp, 520,000 bp, 530,000 bp, 540,000 bp, 550,000 bp, 560,000 bp, 570,000 bp, 580,000 bp, 590,000 bp or 600,000 bp), and the artificial chromosome B has a synthetic segment A of 500,000 to 600,000 bp (which may be 500,000 bp, 510,000 bp, 520,000 bp, 530,000 bp, 540,000 bp, 550,000 bp, 560,000 bp, 570,000 bp, bp, 580,000 bp, 590,000 bp or 600,000 bp), wherein the synthetic fragment A and the synthetic fragment B have homology arms;

[0038] The specific site is not particularly limited, as long as it allows the artificial chromosome A and artificial chromosome B to undergo controllable and expected homologous recombination in Saccharomyces cerevisiae;

[0039] The homology arm and the target site may include, overlap or be 0-100 bp apart, and may be 2 bp, 3 bp, 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 15 bp, 20 bp, 25 bp, 50 bp or 75 bp;

[0040] The method for assembling the artificial chromosome A comprises: mating Saccharomyces cerevisiae A1 and Saccharomyces cerevisiae A2 of different mating types to obtain a diploid, wherein sgRNA in the diploid binds to specific sites of the artificial chromosome A1 and the artificial chromosome A2, mediating Cas9 to perform cutting, and the artificial chromosome A1 and the artificial chromosome A2 undergo homologous recombination in the Saccharomyces cerevisiae to obtain the artificial chromosome A;

[0041] The artificial chromosome A1 and sgRNA are from Saccharomyces cerevisiae A1, and the artificial chromosome A2 and Cas9 are from Saccharomyces cerevisiae A2;

[0042] The artificial chromosome A1 has a length of 200,000 to 400,000 bp (can be 200,000 bp, 210,000 bp, 220,000 bp, 230,000 bp, 240,000 bp, 250,000 bp, 260,000 bp, 270,000 bp, 280,000 bp, 290,000 bp, 300,000 bp, 310,000 bp, 320,000 bp, 330,000 bp, 340,000 bp, 350,000 bp, 360,000 bp, 370,000 bp, 380,000 bp, 390,000 bp or 400,000 bp). bp), the artificial chromosome A2 has a length of 200,000 to 400,000 bp (which may be 200,000 bp, 210,000 bp, 220,000 bp, 230,000 bp, 240,000 bp, 250,000 bp, 260,000 bp, 270,000 bp, 280,000 bp, 290,000 bp, 300,000 bp, 310,000 bp, 320,000 bp, 330,000 bp, 340,000 bp, 350,000 bp, 360,000 bp, 370,000 bp, 380,000 bp, 390,000 bp or 400,000 bp). bp), wherein the synthetic fragment A1 and the synthetic fragment A2 have homology arms;

[0043] The method for assembling the artificial chromosome B comprises: mating Saccharomyces cerevisiae B1 and Saccharomyces cerevisiae B2 of different mating types to obtain a diploid, wherein sgRNA in the diploid binds to specific sites of the artificial chromosome B1 and the artificial chromosome B2, mediating Cas9 to perform cutting, and homologous recombination of the artificial chromosome B1 and the artificial chromosome B2 occurs in the Saccharomyces cerevisiae to obtain the artificial chromosome B;

[0044] The artificial chromosome B1 and sgRNA are from Saccharomyces cerevisiae B1, and the artificial chromosome B2 and Cas9 are from Saccharomyces cerevisiae B2;

[0045] The artificial chromosome B1 has 200,000 to 400,000 bp (can be 200,000 bp, 210,000 bp, 220,000 bp, 230,000 bp, 240,000 bp, 250,000 bp, 260,000 bp, 270,000 bp, 280,000 bp, 290,000 bp, 300,000 bp, 310,000 bp, 320,000 bp, 330,000 bp, 340,000 bp, 350,000 bp, 360,000 bp, 370,000 bp, 380,000 bp, 390,000 bp or 400,000 bp). bp), the artificial chromosome B2 has a length of 200,000 to 400,000 bp (which may be 200,000 bp, 210,000 bp, 220,000 bp, 230,000 bp, 240,000 bp, 250,000 bp, 260,000 bp, 270,000 bp, 280,000 bp, 290,000 bp, 300,000 bp, 310,000 bp, 320,000 bp, 330,000 bp, 340,000 bp, 350,000 bp, 360,000 bp, 370,000 bp, 380,000 bp, 390,000 bp or 400,000 bp). bp), wherein the synthetic fragment B1 and the synthetic fragment B2 have homology arms;

[0046] The length of the homology arm is 40 to 3,000 bp (which can be 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, 150 bp, 200 bp, 250 bp, 300 bp, 350 bp, 400 bp, 410 bp, 420 bp, 430 bp, 440 bp, 450 bp, 460 bp, 470 bp, 480 bp, 490 bp, 500 bp, 510 bp, 520 bp, 530 bp, 540 bp, 550 bp, 560 bp, 570 bp, 580 bp, 590 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 2,000 bp, or 3,000 bp). bp).

[0047] In some specific embodiments of the present application, the assembling method of the artificial chromosome A1, the artificial chromosome A2, the artificial chromosome B1 or the artificial chromosome B2 of the above-mentioned assembly method comprises:

[0048] introducing the artificial chromosome vector backbone A and the medium-length fragment into Saccharomyces cerevisiae C by a protoplast transformation method, and causing homologous recombination in the Saccharomyces cerevisiae to obtain the artificial chromosome A1, the artificial chromosome A2, the artificial chromosome B1 or the artificial chromosome B2;

[0049] The medium-length fragments are 4 to 7 segments, each of which is 40,000 to 71,000 bp in length (can be 40,000 bp, 45,000 bp, 50,000 bp, 55,000 bp, 60,000 bp, 65,000 bp, 71,000 bp, or 75,000 bp), and have homology arms at both ends, enabling the end-to-end connection of the medium-length fragments by homologous recombination to form a synthetic fragment C;

[0050] The synthetic fragment C and the artificial chromosome vector backbone A have homology arms, which can connect them to form a loop;

[0051] The length of the homology arm is 40 to 3,000 bp (can be 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, 150 bp, 200 bp, 250 bp, 300 bp, 350 bp, 400 bp, 450 bp, 500 bp, 550 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 2,000 bp or 3,000 bp).

[0052] In some specific embodiments of the present application, the method for assembling the medium-length fragments of the above-mentioned assembly method comprises:

[0053] The artificial chromosome vector backbone B and the small fragment are introduced into Saccharomyces cerevisiae D by a lithium acetate-mediated transformation method, and homologous recombination occurs in the Saccharomyces cerevisiae to obtain the medium-length fragment;

[0054] The small fragments are 8 to 12 segments, each of which is 2,000 to 6,000 bp in length (can be 2,000 bp, 2,500 bp, 3,000 bp, 3,500 bp, 4,000 bp, 4,500 bp, 5,000 bp, 5,500 bp, 5,600 bp, 5,700 bp, 5,800 bp, 5,900 bp or 6,000 bp), and have homology arms at both ends, so that the small fragments can be connected end to end by homologous recombination to form a synthetic fragment D;

[0055] The synthetic fragment D has an enzyme cleavage site, which can be cleaved by enzymes to form a sticky end that can be connected to the linearized artificial chromosome vector backbone B to form a ring;

[0056] The length of the homology arm is 40 to 3,000 bp (can be 40 bp, 50 bp, 60 bp, 70 bp, (can be 80 bp, 90 bp, 100 bp, 150 bp, 200 bp, 250 bp, 300 bp, 350 bp, 400 bp, 450 bp, 500 bp, 550bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 2,000 bp or 3,000 bp).

[0057] In some specific embodiments of the present application, the Saccharomyces cerevisiae A, Saccharomyces cerevisiae B, Saccharomyces cerevisiae A1, Saccharomyces cerevisiae A2, Saccharomyces cerevisiae B1, Saccharomyces cerevisiae B2 and / or Saccharomyces cerevisiae C in the above assembly method are BY4741 strains.

[0058] In some specific embodiments of the present application, the Saccharomyces cerevisiae D in the above assembly method is BY4741.

[0059] In some specific embodiments of the present application, the artificial chromosome vector backbone A in the above assembly method is a yeast artificial chromosome, which can be PRS414 and / or PRS416.

[0060] In some specific embodiments of the present application, the artificial chromosome vector backbone B in the above assembly method is pCC1BAC.

[0061] In some specific embodiments of the present application, the enzyme cleavage site in the above-mentioned assembly method is NotI.

[0062] In some specific embodiments of the present application, the artificial chromosome A in the above assembly method has a Ura3 marker, and the artificial chromosome B has a Trp1 marker.

[0063] In some specific embodiments of the present application, the artificial chromosome A1 of the above assembly method has a Ura3 marker, the artificial chromosome A2 has a Trp1 marker, the artificial chromosome B1 has a Ura3 marker, and the artificial chromosome B2 has a Trp1 marker.

[0064] In some specific embodiments of the present application, the Saccharomyces cerevisiae used in the above assembly method can be replaced by other eukaryotic single cells.

[0065] The present invention has the following beneficial effects:

[0066] (1) Improved synthesis accuracy: Compared with the SwAP-IN method, the present invention adopts the strategy of synthesizing medium-length fragments first and then gradually assembling them into ultra-long fragments, which effectively reduces the mismatch probability of highly repetitive sequences during the synthesis process and significantly improves the accuracy of ultra-large fragment DNA synthesis containing highly repetitive sequence genes;

[0067] (2) Compared with the existing E. coli synthesis system, the present invention takes advantage of the yeast system, giving full play to the advantages of yeast as a eukaryotic organism in stable inheritance, nuclear structure, and retention of DNA high-level structure, providing a good environment for the accurate assembly of DNA fragments, and further ensuring the reliability of the synthesis process;

[0068] (3) Expanding the scope of application: The method of the present invention provides an efficient and accurate technical means for the field of de novo genome synthesis, which helps to promote the research on genomes containing complex repetitive sequences and has broad application prospects in many fields such as synthetic biology, gene therapy, and biopharmaceuticals. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art.

[0070] Figure 1 Shows the connection mode of eSwAP-IN homology arms;

[0071] Figure 2 indicates mismatches of highly similar repeat sequences;

[0072] Figure 3 Shows similar gene fragments in the SynB region of the human genome;

[0073] Figure 4 The technical route for assembling a 500 kb large fragment is shown;

[0074] Figure 5 The electropherogram of the 500 kb large fragment assembly is shown;

[0075] Figure 6 Shown hAZFa assembly technology route;

[0076] Figure 7 Shows the electrophoresis diagram of SynA, SynB, SynG, SynC, SynAG, and SynBC;

[0077] Figure 8 5 shows the composition of hAZFa;

[0078] Figure 9 Shows the comparison between Saccharomyces cerevisiae and Escherichia coli;

[0079] Figure 10 Shows the element composition of gene fragments from different sources;

[0080] Figure 11 The fragments were present in independent chromosomes with higher-order structures;

[0081] Figure 12 The genetic stability of the fragments was shown. DETAILED DESCRIPTION

[0082] The application discloses a de novo assembly synthesis strategy for Mb-level genomic DNA fragments with human highly repetitive sequences. Those skilled in the art can learn from the contents of this article and appropriately improve the process parameters to achieve it. It is particularly important to point out that all similar replacements and modifications are obvious to those skilled in the art and are considered to be included in the present invention. The methods and applications of the present invention have been described by preferred embodiments, and relevant personnel can obviously change or appropriately modify and combine the methods and applications described herein without departing from the content, spirit and scope of this application to achieve and apply the technology of the present invention.

[0083] It should be understood that the expression "one or more of" includes individually each of the items recited after the expression and various combinations of two or more of the recited items, unless otherwise apparent from the context and usage. The expression "and / or" in conjunction with three or more recited items should be understood to have the same meaning, unless otherwise apparent from the context.

[0084] The terms "comprising", "having" or "containing", including their grammatical synonyms, should generally be understood as open and non-restrictive, e.g., not excluding other unrecited elements or steps, unless otherwise specifically stated or understood from the context.

[0085] It should be understood that the order of steps or the order in which certain actions are performed is not important as long as the application remains operable. Additionally, two or more steps or actions may be performed simultaneously.

[0086] The use of any and all examples or exemplary language such as "for example" or "including" herein is intended only to better illustrate the present application and does not limit the scope of the present application. No language in this specification should be construed as indicating any non-claimed element is essential to the practice of the present application.

[0087] In addition, the numerical ranges and parameters used to define this application are approximate values. The relevant numerical values ​​in the specific examples have been presented as accurately as possible. However, any numerical value inherently inevitably contains standard deviations due to individual testing methods. Therefore, unless otherwise expressly stated, it should be understood that all ranges, amounts, values, and percentages used in this disclosure are modified by the word "about." As used herein, "about" generally means that the actual value is within plus or minus 10%, 5%, 1%, or 0.5% of a particular value or range.

[0088] In some specific embodiments of the present application, the hierarchical assembly is divided into three steps: synthesis of medium-length fragments, assembly of large fragments, and construction of ultra-large-length fragments.

[0089] Medium-length fragment synthesis: Multiple medium-length DNA fragments are synthesized, preferably around 5 kb in length. This length ensures fragment accuracy during synthesis and facilitates subsequent assembly. Precise chemical synthesis ensures the sequence accuracy of each medium-length fragment, reducing the error rate during the initial synthesis phase.

[0090] Large fragment assembly: Multiple synthesized medium-length fragments of approximately 5 kb are further assembled into large fragments ranging in length from 40 to 71 kb. During this assembly process, specific assembly techniques, such as homologous recombination in yeast, are used to connect the medium-length fragments in the correct order and orientation. Because medium-length fragments are relatively short, the repetitive sequences contained within them are more easily identified and connected locally, reducing assembly errors caused by mismatches of long-range repetitive sequences.

[0091] Ultra-Long Fragment Construction: Multiple 40-71 kb fragments are assembled into ultra-long fragments to complete the synthesis of specific regions across the entire genome. This process also leverages the yeast assembly system, leveraging its advantages in processing complex DNA sequences to ensure accurate construction of ultra-long fragments.

[0092] In some specific implementations, relying on efficient assembly of large fragments in vivo, a 500 kb fragment can be assembled de novo within 1 month. For technical routes, see Figure 4 , electrophoresis results see Figure 5 .

[0093] The PAM site and sgRNA sequence information involved in this application are as follows:

[0094] S1: cggccaacgcgaacccttagtgg (SEQ ID NO: 1);

[0095] S2:aacggtgttcacgccgcatccgg(SEQ ID NO: 2);

[0096] S3:atagtgtcacctaaatagcttgg(SEQ ID NO: 3);

[0097] S4:tgatgaacctgaatcgccagagg(SEQ ID NO: 4);

[0098] sgRNAS1:tctttgaaaagataatgtatgattatgctttcactcatatttatacagaaacttgatgttttctttcgagtatatacaaggtgattacatgtacgtttgaagtacaactctagattttgtagtgccctcttgggctagcggtaaaggtgcgcattttttcacaccctacaatgttctgttcaaaagattttggtcaaacgctgtagaagtgaaagttggtgcgcatgtttcggcgttcgaaacttctccGCAGTGAAAGATAAATGATCcggccaacgcgaacccttagGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGGTGCTTTTTTTGTTTTTTATGTCTtcgagtcatgtaattagttatgtcacgcttacGttcacgccctccccccacatccgctctaaccgaaaaggaaggagttagacaacctgaagtctaggtccctatttatttttttatagttatgttagtattaagaacgttatttatatttcaaatttttcttttttttctgtacagacgcgtgtacgcatgtaacattatactgaaaaccttgcttgagaaggttttgggacgctcgaaggcttt(SEQ ID NO: 5);

[0099] sgRNAS2:(SEQ ID NO: 6);

[0100] sgRNAS3:(SEQ ID NO: 7);

[0101] sgRNA S4: tctttgaaaagataatgtatgattatgctttcactcatatttatacagaaacttgatgttttctttcgagtatatacaaggtgattacatgtacgtttgaagtacaactctagattttgtagtgccctcttgggctagcggtaaaggtgcgcattttttcacaccctacaatgttctgttcaaaagattttggtcaaacgctgtagaagtgaaagttggtgcgcatgtttcggcgttcgaaacttctccGCAGTGAAAGATAAATGATCtgatgaacctgaatcgccagGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGGTGCTTTTTTTGTTTTTTATGTCTtcgagtcatgtaattagttatgtcacgcttacGttcacgccctccccccacatccgctctaaccgaaaaggaaggagttagacaacctgaagtctaggtccctatttatttttttatagttatgttagtattaagaacgttatttatatttcaaatttttcttttttttctgtacagacgcgtgtacgcatgtaacattatactgaaaaccttgcttgagaaggttttgggacgctcgaaggcttt (SEQ ID NO: 8).

[0102] The fragment information involved in this application is shown in Table 1.

[0103] Table 1

[0104]

[0105]

[0106]

[0107]

[0108]

[0109]

[0110] DNA source:a GenScript, b Genewiz.

[0111] In this application, "ultra-large fragments", "Mb level" and "megabases" can be used interchangeably to indicate genomic DNA with a length of not less than 0.5 Mb, 0.5~2 Mb, 0.8~1.5 Mb, 0.8~1.2 Mb, or 0.9~1.1 Mb.

[0112] In this application, “highly repetitive sequence” refers to a DNA sequence that has hundreds or even millions of copies in a genome, with an average length of 200-300 bp for the repeat units.

[0113] Unless otherwise specified, the raw materials, reagents, consumables and instruments involved in this application are all common commercial products and can be purchased from the market.

[0114] The present invention will be further described below with reference to the embodiments.

[0115] Example

[0116] The specific implementation process is divided into three main steps: Figure 6 (showing technology roadmap).

[0117] Step 1: Synthesis and assembly of medium-length fragments

[0118] Using lithium acetate-mediated transformation, a total of 233 commercially synthesized DNA fragments of approximately 5 kb (see Table 1) were assembled into 23 medium-length DNA fragments (40-71 kb) in the Saccharomyces cerevisiae BY4741 strain. Specifically:

[0119] (1) The DNA fragments were mixed and digested with NotI and purified using the TIANgel purification kit (DP209-03, Tiangen). The purified fragments were mixed with the PCR linearized pCC1BAC vector at a molar ratio of 1:1 (this is the ratio of each fragment to vector). 8 to 12 5 kb DNA fragments were co-transformed into yeast using the standard LiAc / SS-DNA / PEG transformation protocol (Gietz, RD & Schiestl, RH High-efficiency yeast transformation using the LiAc / SS carrier DNA / PEG method. Nature Protocols 2, 31-34 (2007)). Yeast colonies were selected on SC-His agar plates and cultured at 30°C for 3 days. The assembled DNA fragments were digested with NotI and purified by ethanol precipitation to obtain purified enzyme-digested DNA fragments (a total of 23 fragments, see Table 1 for details).

[0120] Step 2: Assembly of SynA, SynG, SynB, and SynC

[0121] The medium-length fragments were assembled into Saccharomyces cerevisiae VL6-48α or VL6-48a by yeast protoplast transformation (Kouprina, N. & Larionov, V. Selective isolation of genomic loci from complex genomes by transformation-associated recombination cloning in the yeast Saccharomyces cerevisiae. Nature Protocols 3, 371-377 (2008)) to obtain SynA, SynG, SynB and SynC.

[0122] To prepare yeast protoplasts, transfer overnight culture of VL6-48α or VL6-48a to 50 mL of YPAD medium (initial OD 600 = 0.2), cultured at 30°C for approximately 5 hours, and then harvested by centrifugation, washed with ddH2O, and then rinsed with 1 M sorbitol. The cell pellet was resuspended in 20 mL of SPE solution (containing 20 μL of 10 mg / mL lyticase-20T solution and 40 μL of β-mercaptoethanol), mixed, and incubated at 30°C for 40 minutes to digest the cell walls. The pellet was washed twice with 1 M sorbitol and resuspended in 200 μL of STC solution. The purified enzyme-digested DNA fragments (see Table 1 for details) were mixed and added to the protoplasts in STC solution with 50 ng of PCR-amplified linearized vector (PRS414 for SynA and SynB, and PRS416 for SynG and SynC) and incubated for 10 minutes. Then add 800 μL of PEG 8000 solution (20% [w / v] PEG 8000, 10 mM CaCl2, 10 mMTris-HCl, pH 7.5), mix gently by inversion, and let it stand at room temperature for 10 minutes. Collect the protoplasts by centrifugation, resuspend them in 800 μL SOS solution, and incubate them at 30°C for 40 minutes. Transfer the solution to melted SORB-TOP selection medium (SC medium containing 1 M sorbitol and 3% bacterial agar), quickly spread it on an SC agar plate containing 1 M sorbitol, and culture at 30°C for 5-7 days until transformant colonies appear. Electrophoresis results are shown in the table. Figure 7 .

[0123] Step 3: Assembly of SynAG, SynBC, and hAZFa

[0124] The assembly of SynAG (599 kb), SynBC (549 kb) and hAZFa (1.14 Mb) was completed by combining CRISPR-Cas with yeast mating technology. Figure 8 1 shows the composition of hAZFa. Specifically:

[0125] The VL6-48α and VL6-48a strains carrying the preassembled fragments (SynA, SynB, SynG, and SynC) were cultured in 3 mL SC medium of the corresponding auxotrophic type at 30°C overnight and then transferred to a tube containing 3 mL YPAD liquid at a 1:1 ratio (initial OD 600 =0.2). After 8 hours of incubation, serial 10-fold dilutions were performed, and 100 μL of 10 0 , 10 -1 , 10 -2 The dilution was spread on triple-deficient SC medium (SC-His-Lys-Ura for SynAG assembly, SC-His-Lys-Trp for SynBC assembly). MATα yeast carries a Cas9 expression plasmid (His3 tag) and SynA (Trp1 tag), and MATa yeast carries an sgRNA expression plasmid (Lys2 tag) and SynG (Ura3 tag). After the two yeast strains mate to form diploids, the CRISPR system is activated to induce specific cleavage at the S1 / S2 sites of SynA and the S2 site of SynG, releasing the SynA fragment and the SynG fragment with the Ura3 vector. With the help of efficient homologous recombination in yeast, the SynA fragment is assembled into the SynG fragment with the Ura3 tag, and the same is true for SynBC. For electrophoresis results, see Figure 7 For the final assembly of the 1.14 Mb hAZFa construct, a backbone vector containing the Saccharomyces cerevisiae CEN6 centromere was used, and strains with TRP1 marker loss were screened by replica plating to increase the positive rate. Transformants grown on the plates were initially verified by PCR.

[0126] The present invention successfully assembled a 1.14 Mb super-large fragment using Saccharomyces cerevisiae. Compared with Escherichia coli, Saccharomyces cerevisiae is superior in terms of efficient assembly, stable inheritance, nuclear structure, higher-order structure and genetic material existence form ( Figure 9 ), capable of assembling structurally complex gene fragments ( Figure 10 ), and integrated into the chromosome ( Figure 11 ), can be stably inherited for more than 200 generations ( Figure 12 ), which solves the problem faced by existing technologies in synthesizing very large fragments of DNA containing highly repetitive sequence genes. This has broad application prospects in many fields such as synthetic biology, gene therapy, and biopharmaceuticals.

[0127] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of this application.

Claims

1. A method for de novo synthesis of genomic DNA containing ultra-large fragments of highly repetitive human sequences, characterized in that The method comprises: large-fragment DNA assembly and ultra-large-fragment DNA assembly; The steps of assembling large DNA fragments include: Step (i): synthesizing continuous primary DNA fragments of 2 to 8 kb in length, wherein the 3' end of each primary DNA fragment contains a homology arm having a base sequence of 40 to 3,000 bp identical to the 5' end of the next primary DNA fragment; Step (ii): the adjacent 8 to 15 primary DNA fragments in step (i) are respectively transferred into N chromosomes of Saccharomyces cerevisiae cells, and the primary DNA fragments are connected through homology arms in the Saccharomyces cerevisiae cells to form multi-level DNA fragments; Step (iii): extracting the multi-level DNA fragments from step (ii); Step (iv): using the multi-level DNA fragments from step (iii) as new primary DNA fragments, and repeating steps (ii) and (iii) until a large DNA fragment with a length of 300 to 800 kb is formed; The steps of assembling the super-large fragment DNA include: Step (v): transferring the large DNA fragments from step (iv) into a pair of Saccharomyces cerevisiae chromosomes in groups of two, wherein the pair of Saccharomyces cerevisiae chromosomes contain a Cas9 expression plasmid and an sgRNA expression plasmid, respectively; Step (vi): mating the pair of Saccharomyces cerevisiae in step (v) to obtain a diploid, wherein the sgRNA in the diploid binds to a specific site on the Saccharomyces cerevisiae chromosome, mediating Cas9 to perform cutting, and homologous recombination occurs on the Saccharomyces cerevisiae chromosome in the Saccharomyces cerevisiae to obtain an ultra-large DNA fragment; Step (vii): extracting the oversized DNA fragments from step (vi); Step (viii): Using the ultra-large DNA fragment as a new large DNA fragment, repeat steps (v) to (vii) until a complete ultra-large human highly repetitive sequence genomic DNA is obtained.

2. The de novo synthesis method according to claim 1, wherein The transfer into N Saccharomyces cerevisiae cell chromosomes in step (ii) refers to: the adjacent 8 to 15 primary DNA fragments are grouped together, and each group is transferred into Saccharomyces cerevisiae together with the linearized pCC1BAC.

3. The de novo synthesis method according to claim 1 or 2, wherein: The transfer in step (ii) includes transfer via a lithium acetate-mediated conversion method.

4. The de novo synthesis method according to claim 1, wherein The Saccharomyces cerevisiae cell of step (ii) includes BY4741.

5. The de novo synthesis method according to claim 1, wherein The pair of Saccharomyces cerevisiae in step (v) are VL6-48 and VL6-48A.

6. The de novo synthesis method according to claim 1, wherein In step (v), one of the pair of Saccharomyces cerevisiae contains only the Cas9 expression plasmid, and the other contains only the sgRNA expression plasmid.

7. The de novo synthesis method according to claim 1, wherein The step (v) of separately transferring the DNA fragments into a pair of Saccharomyces cerevisiae chromosomes refers to transferring the large DNA fragment and the artificial chromosome backbone into Saccharomyces cerevisiae by a protoplast transformation method, wherein the artificial chromosome backbone is PRS414 or PRS416.

8. The de novo synthesis method according to claim 1, wherein The sequence of the sgRNA that recognizes the specific site is shown in SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7 or SEQ ID NO:

8.

9. The de novo synthesis method according to claim 1, wherein The sequence of the site recognized by the Cas9 is shown in SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3 or SEQ ID NO:

4.

10. A method for assembling genomic DNA of ultra-large fragments of highly repetitive human sequences, characterized in that: include: Saccharomyces cerevisiae A and Saccharomyces cerevisiae B of different mating types are mated to obtain a diploid. The sgRNA in the diploid binds to specific sites on artificial chromosomes A and B, mediating Cas9 cutting. The artificial chromosomes A and B undergo homologous recombination in Saccharomyces cerevisiae to obtain an ultra-large fragment of genomic DNA with highly repetitive human sequences.