Vectors and methods for constructing polycistronic and / or multigenic transgenes

JP2026527460APending Publication Date: 2026-08-14BIOSOLUTION DESIGNS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-03
Publication Date
2026-08-14

Smart Images

  • Figure 2026527460000005
    Figure 2026527460000005
  • Figure 2026527460000006
    Figure 2026527460000006
  • Figure 2026527460000007
    Figure 2026527460000007
Patent Text Reader

Abstract

Cloning vectors and methods useful for constructing multigenic and / or polycistronic constructs for delivery to cells and / or organisms are disclosed. The vectors are genetically engineered to include unique acceptor domains that independently allow for the sequential or repetitive addition of compatible DNA and / or RNA inserts. By inserting a first nucleotide insert into the acceptor domain, the complementary restriction site at the 3' end of the first insert segment is disrupted, while the 3' restriction site at the insertion site of the cloning vector is regenerated, forming a circular vector containing the first insert. Subsequently, cleavage at the 3' end of the first insert forms an insertion site for a second insert segment, and thereafter, each cleavage at the 3' end of a growing linear strand consisting of a desired number of inserts proceeds to the repetitive assembly of additional insert segments.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Cross - reference to related applications This application claims the benefit of U.S. Provisional Application No. 63 / 511,725, filed on July 3, 2023, the entire content of which is incorporated herein by reference in its entirety.

[0002] Sequence listing This application was created on July 1, 2024 and includes as a sequence listing the complete content of the accompanying text file "Sequence.txt" which contains 25,000 bytes, the content of which is incorporated herein by reference.

[0003] There are many types of cloning vectors, but the most commonly used ones are genetically engineered plasmids. A plasmid is a naturally occurring small circular DNA molecule that can be stably maintained within organisms such as bacteria. Plasmids have been modified as useful tools for inserting foreign DNA fragments for cloning purposes. Commercially available vectors have features that allow DNA fragments to be easily inserted into or removed from the vector, such as restriction enzyme recognition sites. The vector and the foreign DNA fragment are treated with a restriction enzyme that cleaves the DNA, generating DNA fragments with protruding ends called blunt ends or sticky ends. Vector DNA and DNA insertion fragments with compatible ends can then be ligated by molecular ligation. Usually, a series of unique restriction enzyme sites are aggregated within a multiple cloning site or polylinker. After the DNA fragment is ligated to the cloning vector, it can be subcloned into another vector designed for a more specific application.

[0004] All cloning vectors commonly used in molecular biology possess key features necessary for their function, such as selection markers. However, some also possess additional features specific to their intended use. For simplicity and convenience, cloning is often performed using Escherichia coli (E. coli). Therefore, the cloning vectors used often possess elements necessary for growth and maintenance within E. coli, such as functional origins of replication (ori). The ColE1 origin is found in many plasmids. Furthermore, some vectors contain elements that allow them to be maintained in organisms other than E. coli.

[0005] With the advancement of gene therapy, the need to deliver multiple genes to treat complex diseases has become recognized in the scientific community. It is advantageous to deliver multigenic therapies that enable the expression of all genes of interest in all target cells. However, when multiple types of vectors are delivered, uptake by target cells can be uneven, potentially resulting in an uneven distribution of transgene expression. Many conventional vectors, including those described in US10036026 and US20080050808, were designed in an era when de novo DNA synthesis was unreliable and cloning of genomic or cDNA fragments was considered the best method to ensure nucleotide sequence fidelity.

[0006] Despite the widespread commercialization and use of plasmid cloning vectors, no vectors specifically designed to enable the rapid, repetitive, or sequential construction of multigenic and / or polycistronic constructs within a single vector exist. As construct sizes increase, the overlap of restriction enzyme sites present in multiple cloning sites and DNA insertion fragments makes the assembly and excision of multigenic and / or polycistronic constructs potentially difficult or impossible. Therefore, there remains a demand for vectors capable of delivering multiple gene products within a single payload.

[0007] Furthermore, there is a need for standardization of gene modules to logically link useful gene components and enable their exchange between multiple classes of plasmid DNA vectors. More specifically, there is a need to develop standardized methods for constructing single-gene and multigenic vectors useful for in vitro transcription, transient transfection, and stable integration. [Overview of the project]

[0008] One aspect of the present invention relates to a method for constructing a polycistronic or multigenic transgene, comprising the steps of providing a set of interrelated plasmid DNA vectors genetically engineered to have one or more unique acceptor sites, and a set of genetically engineered nucleotide fragments having multiple compatibilitys. Each compatible genetically engineered nucleotide fragment has a 5' end and a 3' end that are compatible with any of the one or more unique acceptor sites.

[0009] First, a first plasmid DNA vector is selected and linearized at a unique restriction enzyme site. This forms a first acceptor site consisting of a first restriction enzyme site having 5' and 3' ends, where the first nucleotide fragment has 5' and 3' ends complementary to the 5' and 3' ends of the first restriction enzyme site. Next, the first nucleotide fragment is inserted into the first acceptor site. By ligating the end of the first inserted nucleotide fragment with the end of the acceptor site, the 3' end of the first restriction enzyme site is regenerated, while the 5' end of the first restriction enzyme site is destroyed, and a new 5' end of the first restriction enzyme site is generated. The linearization, insertion, and ligation steps are repeated for one or more additional nucleotide fragments different from the first nucleotide fragment. This iterative or sequential process yields a multigenic and / or polycistronic transgene in which the first inserted nucleotide fragment segment and the one or more additional nucleotide fragments are linked in series.

[0010] These various plasmid DNA vectors are suitable for a variety of applications, including transient transfection, in vitro transcription, gene delivery, and shuttle to viral or nonviral delivery vectors. In one embodiment, a method is provided for transferring effector elements from at least one plasmid DNA cloning vector to an in vitro transcription vector in order to express a target protein or polypeptide in eukaryotic cells. Thus, in another embodiment, a multigenic and / or polycistronic transgene is constructed in a first plasmid DNA vector, excised, and then inserted into a second plasmid DNA vector.

[0011] In another embodiment, the plasmid DNA cloning vector comprises one or more genetically engineered unique acceptor sites and is configured to repeatedly or sequentially generate multigenic and / or polycistronic constructs. The plasmid DNA cloning vector has nucleotide sequence identity selected from the group consisting of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, and SEQ ID NO:7.

[0012] In yet another embodiment, a combination of nucleotide molecules for assembling a multigenic and / or polycistronic construct is provided, comprising a plasmid DNA cloning vector genetically engineered to have one or more unique acceptor sites, and one or more compatible genetically engineered fragments that can be inserted into at least one of the one or more unique acceptor sites. Each compatible genetically engineered fragment has a 5' end and a 3' end that fit into a single unique acceptor site, and the plasmid DNA cloning vector and each of the genetically engineered fragments are designed such that ligating one of the fragments into one of the unique acceptor sites disrupts the 5' end of the acceptor site and regenerates the 3' end.

[0013] The method using the vector integration system of the present invention is based on the following three types of acceptor vector / insertion fragment ligation events.

[0014] 1. A single restriction enzyme site forms an "acceptor site" into which a fragment bounded by 5'-SalI-fragment-XhoI-3' can be inserted. This event is the only and intentionally created instance in the entire BOP structure, and depending on the orientation of the inserted fragment, one of the overhangs defined by the inserted fragment is generated or destroyed.

[0015] 2. In the primary case, when an acceptor vector is constructed by cleaving at two non-complementary restriction enzyme sites, a unique 5' overhang and a unique 3' overhang are generated, and a fragment bounded by the same 5' and 3' overhangs is ligated to these, and the 5' and 3' restriction enzyme sites are regenerated.

[0016] 3. When used only for the EL and TUL modules, an acceptor vector is constructed by cleaving at two non-complementary restriction enzyme sites, generating a unique 5' overhang and a unique 3' overhang (e.g., a TUL acceptor domain with XhoI and NotI). A fragment bounded by a non-identical but compatible (complementary) 5' sticky end (e.g., SalI) and an identical 3' overhang (NotI) is then ligated, disrupting the 5' acceptor restriction enzyme site and regenerating the 3' restriction enzyme site.

[0017] In one embodiment, the one or more unique acceptor sites comprise a plurality of acceptor sites, and the plasmid DNA cloning vector further comprises a plurality of spacers having spacers between each of the plurality of acceptor sites, provided that each spacer is distinct from all other spacers in the plasmid DNA cloning vector. The plurality of spacers are selected from the group consisting of CAGAGTCCC, GGGAGGTTT, ACTCAAGG, GCAGAAGTC, AGCCAACCT, TGCCGAGTC, CCAGCCGCC, GAAGAGGT, CACTTCCTG, CTCTGAGCC, AGCTCCAGT, and ATATCACGC.

[0018] In one embodiment, one or more compatible genetically engineered fragments inserted into the vector are genetically engineered not to encode restriction enzyme recognition sites for SalI, PvuI, KasI, AgeI, XhoI, NotI, and AseI. In another embodiment, the fragment is further designed not to encode a restriction enzyme recognition site for ScaI. In yet another embodiment, the fragment is further designed not to encode a restriction enzyme recognition site for BsgI.

[0019] In yet another embodiment, a method is provided for removing and substituting at least one restriction enzyme site or splice element from a nucleotide sequence for insertion into a cloning vector. This method includes the steps of: identifying at least one restriction enzyme site or splice element to be removed; determining a substitution desired by the user according to a set of rules used by an algorithm; generating an output DNA sequence; comparing whether the amino acid sequence translated from the output DNA sequence is identical to the amino acid sequence translated from the original DNA sequence; and validating the output DNA sequence as having passed a validation step if both amino acid sequences are identical. Furthermore, nucleotides are added to the 5' and 3' ends of the output DNA sequence to enable de novo synthesis, site-directed mutagenesis, and / or vector construction. Specifically, an 8-nucleotide spacer sequence (GAGAGAGA) is added upstream of the translation initiation site (ATG) at the 5' end, followed by a PvuI restriction endonuclease recognition site (CGATCG). Furthermore, a recognition site for ScaI restriction endonuclease (AGTACT) is added downstream of the 3' translation termination site (TAA, TAG, or TGA), followed by an 8-nucleotide spacer sequence (GAGAGAGA).

[0020] In another embodiment, a combination of nucleotide molecules and nucleotide molecular segments for assembling a multigenic and / or polycistronic construct is provided, comprising a first nucleotide molecule genetically engineered to have one or more unique acceptor sites, and one or more compatible genetically engineered nucleotide molecular segments that can be inserted into at least one of the one or more unique acceptor sites. Each compatible genetically engineered nucleotide molecular segment has a 5' end and a 3' end that fit a single unique acceptor site, and the first nucleotide molecule and each of the genetically engineered nucleotide molecular segments are designed such that ligating one of the segments to one of the unique acceptor sites results in the disruption of the 5' end of the acceptor site and the regeneration of the 3' end.

[0021] In yet another embodiment, a kit is provided for constructing a polycistronic or multigenic transgene, which comprises a set of interrelated plasmid DNA vectors genetically engineered to have one or more unique acceptor sites. Each of the interrelated plasmid DNA vectors is designed to allow insertion and ligation of multiple compatible genetically engineered nucleotide fragments, each compatible genetically engineered nucleotide fragment having 5' and 3' ends that fit with one or more unique acceptor sites of any of the interrelated plasmid vectors in the set. Here, the first plasmid DNA vector is one vector in the set, and one or more of the interrelated plasmid DNA vectors are insertable into at least one other plasmid DNA vector in the set. Other features and advantages of the present invention are described in the following detailed description of the invention, some of which will become apparent from that description or will be understood through the practice of the invention. The present invention is carried out and achieved by compositions and methods particularly specified herein and in the claims. [Brief explanation of the drawing]

[0022] The accompanying drawings constitute part of this specification and illustrate embodiments of the invention, and together with the above-mentioned summary of the invention and the detailed description of the invention below, they illustrate the invention. This patent or patent application file contains at least one drawing drawn in color. A copy of this patent or patent application publication containing the color drawing will be provided by the Patent Office upon request, subject to payment of the necessary fees. [Figure 1]Figure 1 illustrates the central dogma of molecular biology (DNA is transcribed into RNA, and RNA is translated into protein) in relation to the vector of the present invention. The figure details the modular structure of eukaryotic genes at each level: a) double-stranded DNA, b) primary RNA transcript showing the arrangement of exons (rectangles) and introns (single lines between rectangles), and candidate splicing domains (dashed lines), c) spliced ​​messenger RNA (mRNA) with ligated exons (continuous rectangles), where the 5' untranslated region (5'UTR) is defined by the leftmost rectangle shown in lighter color, the open reading frame (ORF) is defined by the region shown in darker color, the 3' untranslated region (3'UTR) is defined by the rightmost rectangle shown in lighter color, and the poly-A tail, and d) translated protein (rectangle). The figure further illustrates that a particular region of DNA contains multiple different modules that exert biological functions via DNA, RNA, or protein-encoded units (DNA-coding modules). Furthermore, this figure shows that the DNA encoding module can be abstracted into three distinct domains (vector-defining gene modules) that can encode plasmid DNA-based vectors. These three vector-defining gene modules consist of a transcriptional regulatory (R) domain or element, a biological effector (E) domain or element, and a transcriptional processing (P) domain or element. [Figure 2A] Figures 2A and 2B show an overview of a typical pBOP-30 vector. Figure 2A shows pBOP-30 at the top level, illustrating the order and identity of restriction enzyme recognition sites and spacers in the polyMOD framework (top row), and the modular configuration within the circular vector shown below. The modular structure of pBOP-30 can be used to design multigenic and / or polycistronic constructs using Sal / Xho → Xho or Sal / Not strategies. [Figure 2B] Figure 2B shows the polyMOD framework introduced within the E element, which is an ORF, and is used to construct substructures of the ORF. [Figure 3]Figure 3 shows an example of modular elements inserted into the pBOP-30 vector to form a complete transcription unit (TU) within the vector. An example of this completed construct is shown as pBOP-30-TU-CMEGBH. [Figure 4] Figure 4 is a schematic diagram showing an example of a method for constructing a multi-gene construct using the pBOP-30 vector structure by subcloning a transcription unit (TU) into the XhoI site. A single XhoI restriction enzyme site forms an "acceptor site" into which a fragment bounded by 5'-SalI-fragment-XhoI-3' can be inserted. This event is the only and intentionally provided single example throughout the BOP structure, and depending on the orientation of the inserted fragment, one of the overhangs defined by the inserted fragment is generated or destroyed. [Figure 5] Figure 5 is a schematic diagram showing an example of another method for constructing a multi-gene construct using the pBOP-30 vector structure by directionally subcloning a transcription unit into a transcription unit linker (TUL). [Figure 6] Figure 6 shows a method for constructing a larger multi-mer type plasmid DNA vector in a three-step process by maintaining the TUL in the 3'-terminal transcription unit. [Figure 7AB] Figure 7 shows the polyMOD framework in the upper part and the module arrangement within the multi-effector / polycistronic pBOP-30-IRES vector in the lower part. The pBOP-30-IRES vector is useful for constructing a 2-ORF polycistron in two steps and for constructing larger polycistronic constructs. [Figure 8A] Figures 8A - 8B show the multi-effector / polycistronic IRES shuttle system of pBOP-30-IRES, which can be used to construct a 3-ORF polycistron in two steps and for constructing larger polycistronic constructs. Figure 8A shows the process of constructing a bi-ORF construct. [Figure 8B] Figure 8B shows the process of constructing a tri-ORF construct. [Figure 9A] Figures 9A-9E show vector structures useful for incorporating bioactive RNA species (bioRNAs) into the effector domains of transcription units, and methods for constructing multi-effector E domains combined with one or more bioRNAs or ORFs by combining these with pBOP-30 or pBOP-30-IRES vectors. Figure 9A shows the modular structure of a BOP-compliant intron-coding bioactive RNA (bioRNA) element. The upper panel shows a method for placing one or more bioRNAs within a synthetic splicing unit, and the lower panel shows a method for placing one or more bioRNAs within a wild-type splicing unit. [Figure 9B] Figure 9B shows the modular structure of the pBOP-30-bioRNA-1 vector, which is designed to either place a single bioRNA within a single effector domain or place it upstream of a multi-effector containing one or more ORFs downstream of the bioRNA. [Figure 9C] Figure 9C shows the modular structure of the pBOP-30-bioRNA-2 vector, which is designed to position bioRNA downstream of one or more ORFs. [Figure 9D] Figure 9D shows the process of positioning the bioRNA within the EL domain. [Figure 9E] Figure 9E shows how to position the bioRNA element upstream of the second ORF controlled by IRES and downstream of the first ORF. [Figure 10A] Figures 10A to 10C show various configurations of the pBOP-40 vector. Figure 10A shows the polyMOD framework in the upper section and the module structure of the pBOP-40 vector in the lower section. [Figure 10B] Figure 10B shows an example of a pre-assembled construct built within the pBOP-40 vector. [Figure 10C] Figure 10C shows a three-step method for constructing a multigenic construct within the pBOP-40 vector through sequential steps. [Figure 11A] Figure 11A shows the modular structure of the genome integration negative selector backbone (GINSB). [Figure 11B] Figure 11B shows the modular structure of the polyMOD framework and the pBOP-40-NS vector. [Figure 12] Figure 12 shows the polyMOD framework and modular structure of the pBOP-30-IVT vector, which includes an inversely complementary BsgI site for cleaving the polyA site within the polyA sequence. [Figure 13] Figure 13 shows a flowchart of the algorithm for fitting protein-coding effector sequences to vectors in the BOP system, divided into three stages: Sanitize Sequence Input, BOPize Sequence, and Verify Sequence. [Figure 14] Figure 14 shows a flowchart of the Sanitize Sequence Input process. [Figure 15] Figure 15 shows a flowchart of the BOPize Sequence process. [Figure 16] Figure 16 shows a flowchart of the Verify Sequence process. [Modes for carrying out the invention]

[0023] This invention provides an object-linking architecture designed for constructing multigenic therapies that can be delivered by delivery as synthetic RNA or by delivery to ex vivo cells or in vivo using viral or nonviral DNA delivery methods. The significance of the object-linking system lies in the rapid exchange of one or more pre-validated and standardized genetic elements within or between existing DNA vectors. In the absence of a standardized object-linking platform, researchers are forced to construct DNA molecules by assembling DNA fragments from multiple methods or rely on modern de novo synthesis techniques to construct entire DNA molecules from scratch. This invention provides an architecture suitable for the low-cost and rapid construction of therapeutically relevant complex DNA molecules by enabling the rapid linking of specified genetic elements, which can be produced by various molecular biological methods, using established genetic engineering techniques. Most importantly, the final products and / or intermediates obtained using the vectors and methods of this invention enable the low-cost and rapid production of vector variant libraries suitable for rapid screening of desired therapeutic functions without requiring de novo synthesis of different species.

[0024] DNA vectors for constructing DNA molecules having one or more transcription units and methods for using them are provided. These vectors leverage the availability of de novo synthesis (DNS) for researchers. The methods of use guide the generation of DNA sequences in which specific restriction sites are removed and / or replaced by nucleotide alterations or deletions, allowing a given genetic function to be bounded by those restriction sites.

[0025] In one embodiment, a bilayered, interconnected plasmid series is provided, useful for the repetitive or sequential assembly of multiple DNA segments. The basic vector structure includes two unique acceptor domains that allow for the independent, repetitive or sequential addition of biological effectors, or entire transcription units. These two acceptor domains have a common molecular architecture, with their 5' ends defined by complementary restriction site products that are destroyed upon ligation, and their 3' ends defined by a common site that is regenerated by ligation. When a first DNA segment is inserted into the acceptor domains, the complementary restriction site at the 5' end of the first DNA segment is destroyed, while the 3' restriction site is regenerated at the insertion site, forming a circular vector having the first insert. Next, a cut is made at the 3' end of the first insert to form the insertion site for the second DNA segment. Thereafter, each subsequent cut at the 3' end of the elongating linear strand, consisting of the desired number of inserts, allows for the repetitive or sequential assembly of additional DNA segments.

[0026] In another embodiment, a method of using the vector is disclosed. This method of use includes simple guidelines for ensuring that de novo synthesis products, nucleotide fragments derived from genomic DNA, cDNA, and / or clones derived from other plasmid vectors are compatible with the vector of the present invention, i.e., that the nucleotide sequences inserted into any of the vectors do not encode specific restriction enzyme sites. In one embodiment, the user may need to introduce one or more mutations into the inserted sequence to ensure that SalI, PvuI, KasI, AgeI, XhoI, NotI, and AseI are not present in the inserted sequence. In another embodiment, any ScaI sites must also be mutated to make the components compatible with a separate polycistronic internal ribosome entry site (IRES) shuttle system or a bioactive RNA shuttle system. In yet another embodiment, any BsgI sites are removed from all sequences contained within the insert into the in vitro transcription vector.

[0027] In this specification, "Bird of Prey TM The terms “BOP” and “BoP” are used interchangeably to refer to the vectors of the present invention, which form an interconnected vector system for various applications in biological systems. The pBOP-30 vector is the foundation of the BOP system and was used to construct all subsequent vectors disclosed herein. The terms “BOPization” and “BOPizer” refer to the processes and / or algorithms used to prepare nucleotide sequences for insertion into the BOP vectors of the present invention.

[0028] In this specification, the term “mutate” refers to any process used to alter the nucleotide identity of a target DNA sequence. Various techniques are well known to those skilled in the art and can be used to modify or remove specific nucleotides. Site-directed mutagenesis of DNA sequences is also well known to those skilled in the art as gene editing. The use of de novo synthesis may also be referred to as the creation of a designed mutant nucleotide sequence using the target sequence as a template. For the purposes of this invention, mutation or editing is performed on a target DNA sequence to remove or introduce one or more restriction enzyme recognition sites. Generally, mutations are understood to be conserved, that is, the mutant nucleotide sequence translates into a protein that is indistinguishable from (i.e., identical to) the amino acid sequence translated from the non-mutant sequence. In other words, the desired mutation to be introduced is a conserved substitution.

[0029] Some of the abbreviations used herein include TU (transcription unit), R (transcription regulator), E (biological effector), P (transcription processing), IRES (internal ribosome entry site), ORF (open reading frame), UTR (untranslated region), EL (effector linker), CCEs (compatible cohesive ends), TUV (transcription unit vector), TUL (transcription unit linker), CIP (calf intestinal phosphatase), CMV (cytomegalovirus), GIC (genome integration controller), CMD (chromatin modification domain), and PSE (positive selector). element (positive selector element), BE (bioindicator element), TUAD (transcription unit acceptor domain), GFP (green fluorescent protein), RFP (red fluorescent protein), BDNF (brain-derived neurotrophic factor), EMCV (encephalomyocarditis virus), IVT (in vitro transcription), IVTP (in vitro transcription promoter)This includes in vitro transcription promoters, genomic integration negative selector backbones (GINSB), synthetic splice donors (SD), 5' portion synthetic introns (SI-5), 3' portion synthetic introns (SI-3), synthetic splice acceptors (SA), wildtype splice donors (WD), wildtype introns (WI-5), wildtype introns (WI-3), wildtype introns (WI-3), and wildtype splice acceptors (WA).

[0030] In this specification, the term “transcription unit (TU)” is used to describe the genetic material necessary to control the initiation and termination of transcription, as well as intervening sequences related to the regulation of post-transcriptional modifications and the control of protein translation, and to protein products and / or bioactive RNA species. A vector may contain a single transcription unit consisting of defined “domains” for arranging specific biological “elements.” Nucleotide sequences suitable for each domain are classified as “elements” and inserted into the corresponding domains. Thus, after assembly, a single TU consists of three functional elements: a) a transcriptional regulatory (R) element, b) a biological effector (E) element, and c) a transcriptional processing (P) element, each of which is inserted into the R domain, E domain, and P domain, respectively (see Figure 2). The present invention may be used to construct vectors having multiple effectors within the E domain, and vectors having multiple TUs in the context of the entire vector. To prevent intravector recombination during and / or after pDNA production, the genetic sequences of each TU element (R, E, and P) must be sufficiently different from each other.

[0031] In this specification, the term "multigenic" refers to a gene construct that codes for multiple mutually distinct transcription units. Multigenic constructs may also be referred to as bigenic, trigenic, quadgenic, etc.

[0032] In this specification, the term "polycistronic" is interchangeable with the term "multi-effector transcript" and refers to an RNA transcript encoding multiple open frames (ORFs) and / or one ORF and one or more bioactive RNA (bioRNA) species. Polycistronic constructs may also be referred to as bicistronic, tricistronic, quadcistronic, etc. Bioactive RNAs are non-coding RNAs that include, but are not limited to, microRNAs (miRNAs), short hairpin RNAs (shRNAs), miRNA antagonists (anti-miRs), and long non-coding RNAs (lncRNAs).

[0033] The R element can consist of DNA, RNA, or a combination of DNA and RNA sequences. DNA sequences can promote transcription, while RNA sequences can provide post-transcriptional and pre-translational regulation. For example, the 5' untranslated sequence (5'UTR) can affect RNA stability, intracellular localization, and / or translation efficiency. RD refers to a subdomain of a transcriptional regulatory domain that primarily encodes DNA-based regulatory factors, while RR refers to a subdomain of a transcriptional regulatory domain that primarily encodes RNA-based regulatory factors.

[0034] The E element may encode a bioactive RNA species (e.g., microRNA (miRNA)) and / or polyamino acids (peptides and / or proteins). In this specification, the term “biological effector” refers to a DNA unit capable of producing an ORF that can be transcribed and translated into a bioactive RNA molecule, long non-coding RNA (lncRNA), or protein. An effector domain may contain multiple ORFs and / or bioactive RNA species. EN refers to a subdomain of an ORF-coding effector that places a protein subelement at its amino (N) terminus. EI refers to a subdomain of an ORF-coding effector that places a protein subelement internally and is flanked by an amino-terminal and a carboxyl-terminal element. EC refers to a subdomain of an ORF-coding effector that places a protein subelement at its carboxyl (C) terminus.

[0035] The P element may encode RNA-based 3'UTR sequences and polyadenylation signals, as well as RNA / DNA hybrid coding sequences that control transcription termination, and DNA coding elements that can promote upstream transcription, DNA coding elements that can regulate downstream transcription, DNA coding elements that can prevent read-through of transcripts from downstream genes encoded on the complementary strand, or DNA coding elements that can prevent gene silencing (e.g., enhancers, repressors, insulators, and / or chromatin remodeling sequences).

[0036] In this specification, the term "polyMOD framework" refers to a nucleotide sequence encoding a series of restriction enzyme recognition sites, with spacer sequences inserted between the restriction enzyme recognition sites. The polyMOD framework is similar to multiple cloning sites found in cloning vectors, but differs in that the selection and sequence order of restriction enzyme sites are carefully designed to assemble inserts or elements (i.e., R elements, E elements, P elements, etc.) into a multigenic or polycistronic construct in a controlled manner. The polyMOD framework allows for the ordering of elements that replicates the order of genomic elements such as promoters, intron / exon coding sequences, polyA, and other regulatory sequences, even when various genetically engineered inserts are selected from arbitrary biological sources or when they are de novo synthesized with desired nucleotide sequences. The polyMOD framework also differs from conventional multiple cloning sites in that spacer sequences are inserted between restriction enzyme recognition sites to ensure sufficient space for each enzyme to bind to its corresponding site and cleave the DNA sequence. In conventional multiple cloning sites, spacers are usually not needed because the user selects and cuts only one or two locations within the site. However, in the polyMOD framework, almost all (and sometimes all) of the sites are used to assemble functional transcription units.

[0037] In this specification, the term "rapid assembly" refers to the ability to directly, repeatedly, or sequentially attach fragments or inserts to an elongating linear insert chain without the need to use another restriction enzyme to cleave other regions of the polyMOD framework or cloning vector. In some cases, the assembly process must be carried out sequentially, such as when joining individual modules or elements to form transcription units. In other cases, the assembly process is carried out iteratively, such as when combining transcription units to create multigene or polycistron constructs.

[0038] In this specification, the terms “internal ribosome entry site” or “IRES” refer to a DNA sequence that, when transcribed, produces an RNA species having a tertiary structure that binds to ribosomes and facilitates protein translation. The vectors and methods of the present invention enable the construction of IRES libraries, and individual IRES modules may exhibit preferential activity in a given cell type and / or cellular state. Furthermore, since the function of IRESs may be influenced by the tertiary structure of upstream and downstream RNAs, an IRES library provides the user with the ability to construct and test which IRES module functions best with a given ORF in a given cell type. In this specification, the term “Ori” is a term used by those skilled in the art to identify the DNA sequence encoding the origin of replication in a plasmid vector. The Ori in the vectors of the present invention include, but are not limited to, those used in widely available plasmids such as pUC, pBR322, pET, pGEX, pColE1, R6K*, pACYC, pSC101, pBluescript, pGEM and / or PCDF. Alternatively, it may include Ori identified as pMB1, ColE1, p15A, pAC101, F1, CloDF13, and / or CDF. The bacterial replication system may include elements in cosmids (COSMID), fosmids (FOSMID), and bacterial artificial chromosomes (BAC).

[0039] In this specification, the term "spacers" refers to nucleotide sequences that interpose between restriction enzyme sites encoded in a vector. Any element inserted into the vector replaces a spacer. For example, in the vector shown in Figure 2, the spacer between the PvuI and ScaI sites is replaced when the E element is successfully inserted into the vector. For example, the E element spacer is a "placeholder" at the position where the spacer is removed and the E element module is inserted into the vector. Other spacers include, but are not limited to, positive selector element spacers, E element spacers, TUL element spacers, P element spacers, and 5' chromatin modification domain spacers. Therefore, the nucleotide sequences encoding the polyMOD framework of a vector can be represented as a string of restriction enzyme sites in parentheses flanking the spacers, as in the following example: [(PstI)+(spacer)+(SalI)+(R element spacer)+(PvuI)+(E element spacer)+(ScaI)+(EL element spacer)+(KasI)+(P element spacer)+(AgeI)+(XhoI)+(TUL element spacer)+(NotI)+(spacer)+(XbaI)+((pBOP-30 vector backbone (VB))+(PstI)].

[0040] Spacers are preferably 6 to 15 base pairs long and exist as trimmers to avoid translational frameshifts. Ideally, spacers should not encode a) restriction enzyme binding sites, b) transcription factor binding sites, c) Kozak sequences, d) RNA splice donor or splice acceptor sequences, e) translation start (ATG) codons, or f) polyadenylation signals (AATAAA), either within their own sequence or when ligated to upstream or downstream sequences. Furthermore, to avoid homologous recombination events, each spacer may be used only once per vector. Candidate spacer sequences include, but are not limited to, the exemplary nucleotide sequences shown in Table 1.

[0041] Table 1 Candidate Spacer Arrays [Table 1]

[0042] Table 2 provides two examples of how the spacer arrays listed in Table 1 can be used in a vector. Each spacer is used only once per vector, but can be distributed in any order within the vector. The notation N / A indicates that the element is not part of the vector in question. Table 2. Candidate use examples of spacer arrays in two vectors. [Table 2]

[0043] In each vector, positive and / or negative selection factors are used. Therefore, antibiotic resistance genes are used to select cells or bacteria that have successfully undergone insertion cloning. Antibiotics useful for this purpose include, but are not limited to, kanamycin, spectinomycin, streptomycin, ampicillin, carbenicillin, bleomycin, erythromycin, polymyxin B, tetracycline, and chloramphenicol.

[0044] Figure 1 shows the modular structure of eukaryotic genes and the molecular components associated with the pre-transcriptional, transcriptional, post-transcriptional, pre-translational, translational, and post-translational stages. The pre-transcriptional stage includes chromatin modification domains, silencers, etc., while the transcriptional stage includes promoters, enhancers, repressors, etc. The post-transcriptional stage includes RNA splicing, RNA stabilizing or destabilizing elements, and intracellular RNA transport. The pre-translational stage includes intracellular localization-specific translation initiation, cell or state-specific translation initiation, etc. The translational stage includes translational stuttering or translational stalling for controlled protein folding. The post-translational stage includes phosphorylation, glycosylation, etc. Furthermore, Figure 1 deliberately abstracts gene regulatory elements with the aim of developing a modular DNA vector manufacturing system in which each of these transcription units can be used to produce DNA molecules encoding one or more transcription units that encode one or more biological effectors.

[0045] The BOP transcription unit consists of four main elements, namely TR, ED, EL, and TP, which are also identified as R, E, EL, and P. Each of these elements is intentionally designed to have a defined substructure to enable efficient modular exchange of defined functional elements. These subdomains are flanked by one or more restriction enzyme sites specific to the defined element or module class. Figure 2 shows how the four main elements of the transcription unit are defined by five rare restriction enzyme sites, and the junction of each functional element is defined by a shared restriction site as follows: 1) the transcriptional regulatory element or domain R is flanked by SalI and PvuI; 2) the biological effector element or domain E is flanked by PvuI and KasI; 3) the effector linker element or domain EL is defined by ScaI and KasI; and 4) the transcriptional processing element or domain P is flanked by KasI and AgeI. Furthermore, Figure 2 shows that the R domain can be subdivided into two sub-elements RD and RR by the introduction of the AfeI site. Within the E domain, the ORF-based effector can be subdivided into three distinct subdomains, EN, EI, and EC, by the fixed placement of the XmaI and MfeI sites. The P domain can be subdivided into two sub-elements, P-5 and P-3, by the introduction of the FspI site.

[0046] The R subdomain allows for efficient differentiation between DNA-encoded functional elements (RD) and RNA-encoded functional elements (RR). RD subdomains are designed to contain, but are not limited to, DNA-based elements encoding promoters, enhancers, repressors, chromatin modification domains, and insulator sequences. RR subdomains are designed to contain, but are not limited to, RNA-based elements encoding splicing units and 5' untranslated regions. However, distinguishing these R sub-elements as DNA-based or RNA-based does not preclude the intentional placement of DNA-based motifs within RR domains. For example, important DNA-encoded transcription factor binding sites are often found downstream of transcription start sites.

[0047] ORF-based effectors often consist of distinct protein-coding sequences without synthetic substructures, although some ORFs compliant with BOPs may consist of one or more defined subdomains. In some cases, users may wish to affix one or more functional motifs to either the amino-terminus or carboxy-terminus of the target protein. These functional fusion protein motifs often confer one or more of the following properties, but are not limited to these examples: 1) intracellular localization signals, 2) visualization tags, and 3) protein purification elements. The ORF substructure of BOPs allows for the distinct placement of motifs EN located at the amino-terminus of the protein and motifs EC located at the carboxy-terminus. The EI domain is located between the N-terminal and C-terminal subdomains.

[0048] The substructures of P allow for the distinction between RNA-encoded and DNA-encoded motifs. Unlike the substructures of R, the substructures of P consist of a mixture of DNA-encoded and RNA-encoded elements. Therefore, the nomenclature of P substructures uses positional terms such as the 5' or 3' side of the FspI splitting site. The P-5 domain is designed to contain RNA-based splicing elements, a 3' untranslated region sequence, and a polyadenylation signal. The P-3 domain is designed to contain RNA / DNA hybrid motifs responsible for transcription termination, as well as DNA-based motifs that can encode enhancer, repressor, chromatin modification domains, and insulator elements.

[0049] For the purpose of cloning DNA into the vector of the present invention, the insertion DNA to be cloned must be modified to remove a specific nucleotide sequence. Mutations to the sequence of interest are introduced to remove restriction enzyme sites that interfere with the design and use of polycistronic and / or multigenic constructs. In this specification, the term “compliant” means a nucleotide sequence that does not contain restriction enzyme sites that interfere with the design or use of the construct and is therefore ready for insertion into the vector of the present invention.

[0050] It is generally understood in the art that when modifying promoter sequences, it is advantageous to make the smallest possible changes, given the importance of individual trans-factor binding sites and their relative distance from the corresponding transcription start sites. To obtain the full utility of this system, it is necessary to remove specific restriction enzyme sites from the components. Therefore, while restriction enzymes can be used within the vector, they must be removed by editing from any newly synthesized DNA sequences intended for insertion into the vector's cloning site. Alternative methods for introducing mutations into sequences include, but are not limited to, site-directed mutagenesis performed directly from cDNA, genomic clones, or genomic DNA or reverse-transcribed mRNA. The restriction enzyme sites shown in Table 3 are located within the indicated vector domains and must be removed from the DNA sequences of elements to be inserted into the corresponding domains. Novel sites useful for transcription unit subdomains are shown in bold.

[0051] Table 3 Restriction enzyme recognition sites within the vector domain [Table 3]

[0052] Certain R and P elements may contain functional domains affected by transcription factor binding sites and / or hairpin structures, making it impossible to introduce mutations without adversely impacting their function. Therefore, within R and P domains, only major architectural sites are mutated. The major sites in transcription unit multigenic vectors are SalI, PvuI, ScaI, KasI, AgeI, XhoI, NotI, and AseI. On the other hand, when encoding amino acids within an ORF, it is possible to select from multiple different codons, thus offering greater freedom in mutating restriction enzyme sites within an ORF. Given this, more restriction enzyme sites are selected for specific functional purposes, such as pivoting sub-ORF components to surround a defined effector ORF. For example, XmaI and MfeI may be used to establish the amino-terminal and carboxy-terminal modules, respectively.

[0053] Table 4 provides the name and sequence list number for each plasmid DNA vector. The table also lists the restriction enzyme recognition sites present within the polyMOD framework for each vector.

[0054] Table 4 Restriction enzyme sites of vectors, sequence numbers, and polyMOD frameworks [Table 4]

[0055] The pBOP-30 vector contains a nucleotide sequence encoding [(PstI)+(spacer)+(SalI)+(R element spacer)+(PvuI)+(E element spacer)+(ScaI)+(EL element spacer)+(KasI)+(P element spacer)+(AgeI)+(XhoI)+(TUL element spacer)+(NotI)+(spacer)+(XbaI)+((pBOP-30 vector backbone (VB)))] (Note: A spacer is not required between AgeI and XhoI).

[0056] The pBOP-IRES vector contains [(PstI)+(spacer)+(SalI)+(R element spacer)+(PvuI)+(reporter ORF-1 element (e.g., red fluorescent protein) spacer)+(EcoRV)+(IRES element spacer)+(PacI)+(reporter ORF-2 element (e.g., green fluorescent protein) spacer)+(ScaI)+(EL element spacer)+(KasI)+(P element spacer)+(AgeI)+(XhoI)+(TUL element spacer)+(NotI)+(spacer)+(XbaI)+((pBOP-30 vector backbone (VB)))].

[0057] The pBOP-30-bioRNA-1 vector contains [(PstI)+(spacer)+(SalI)+(R element spacer)+(PvuI)+(EcoRV)+(bioRNA element spacer)+(ScaI)+(EL element spacer)+(KasI)+(P element spacer)+(AgeI)+(XhoI)+(TUL element spacer)+(NotI)+(spacer)+(XbaI)+((pBOP-30 vector backbone (VB)))].

[0058] The pBOP-30-bioRNA-2 vector contains [(PstI)+(spacer)+(SalI)+(R element spacer)+(PvuI)+(E element spacer)+(EcoRV)+(bioRNA element spacer)+(ScaI)+(EL element spacer)+(KasI)+(P element spacer)+(AgeI)+(XhoI)+(TUL element spacer)+(NotI)+(spacer)+(XbaI)+((pBOP-30 vector backbone (VB)))].

[0059] The pBOP-40 vector contains a nucleotide sequence encoding [(PstI)+(GID element spacer)+(AscI)+(5' chromatin modification domain (CMD) element spacer)+(SphI)+(Positive selection factor element (PSE) spacer)+(SpeI)+(Bioindicator element (BE) spacer)+(PspOMI)+(Spacer)+(SalI)+(R element spacer)+(PvuI)+(E element spacer)+(ScaI)+(EL element spacer)+(KasI)+(P element spacer)+(AgeI)+(XhoI)+(Spacer)+(NsiI)+(3' chromatin modification domain (CMD) element spacer)+(NotI)+(3' GIC element spacer)+(pBOP-30 vector backbone (VB))].

[0060] The pBOP-40-NS vector is a variant of pBOP-40 that further contains one or two negative selectors flanking the GIC element. Its structure is defined as [(FspI)+(NegSel-1 element spacer)+(PstI)+(5'GIC element spacer)+(AscI)+(spacer)+(NotI)+(3'GIC element spacer)+(XbaI)+(NegSel-2 element spacer)+(AfeI)+(pBOP-30 vector backbone (VB))]. The entire NegSel-1 polyMOD can be isolated by FspI and PstI, and the NegSel-2 polyMOD can be isolated by XbaI and AfeI.

[0061] The pBOP-30-IVT vector contains a nucleotide sequence encoding [(PstI)+(spacer)+(SalI)+(IVTP element)+(AfeI)+(RR ACC)+(PvuI)+(E element)+(KasI)+(P-5 ACC)+(FspI)+(polyA element)+(reverse BsgI)+(AgeI)+(spacer)+(XbaI)+((pBOP-30 vector backbone (VB))+(PstI))].

[0062] In certain applications, additional editing of the DNA may be required. For example, in DNA sequences used to synthesize mRNA products by in vitro transcription, BsgI sites present in the 5' regulatory motif, ORF, and 3' regulatory motif must be removed. This is because the enzyme is used to linearize the polyA sequence in the in vitro transcription process of the present invention.

[0063] Each vector can be described as containing a nucleotide sequence encoding a polyMOD framework having a column of restriction enzyme sites separated by spacer sequences. While these vectors are interrelated, each is designed for a different application. The pBOP-30 vector is designed for transient transfection leading to non-persistent episomal delivery of pDNA. The pBOP-30-IRES vector is designed to construct a multi-effector or polycistronic transcription unit having effector domains capable of encoding two or more ORFs and / or two or more bioRNAs. The pBOP-30-IVT vector is designed for in vitro transcription. The pBOP-40 and pBOP-40-NS vectors have structures that allow for the incorporation of a single effector domain, a single transcription unit, or a multigen into a virus production shuttle vector or a non-viral genome integration shuttle vector. The pBOP-30-bioRNA vector is designed to encode bioactive RNA species, either as a single effector, positioned upstream or downstream of a single ORF, or within a mixed ORF / bioRNA multi-effector transcription unit using the pBOP-30-IRES vector.

[0064] These seven vectors are "interrelated" in that they are designed to allow the transfer of a completed multigenic or polycistrone constructed within one vector to another for use in different applications. For example, a multigenic construct constructed within the pBOP-30-IVT vector can be excised as a whole and inserted into a pBOP-40 vector for transfer into a gene delivery or viral delivery system. Thus, the interrelationships between these vectors enable rapid testing of constructs and subsequent application to therapeutic uses. Furthermore, those skilled in the art will understand that some of the selected components can be reacted in a single reaction vessel. While this is not the case between the pBOP-30 and pBOP-40 systems, each system contains a subset of elements suitable for dynamic vector assembly.

[0065] The overall structure of the pBOP-40 vector contains restriction enzyme-separated gene modules that encode important biological functions for controlling and evaluating the delivery and stable integration of foreign DNA into eukaryotic genomes. The pBOP-40 vector structure includes gene modules that allow for the inclusion of a) genomic integration regulators, b) chromatin modification domains, c) positive selection factor elements, and d) bioindicator elements. Because these domains reside within the target receptor shuttle vector, the components of the pBOP-40 vector only need to have mutations in restriction enzyme sites that enable the ordered arrangement of GIC, CMD, PSE, and BE, and allow for the stepwise insertion of single effector, single transcription unit, or multigenic transcription unit into transcription unit receptor domains 1, 2, or 3. The deliberate selection of restriction enzyme sites enables the stepwise addition of these components to the TUAD.

[0066] The TUAD-1 position is always used first and can be used to place single effectors prepared in PvuI and KasI into TUAD-1 in PvuI and KasI, or to place single transfer units without TULs prepared in SalI and AgeI into TUAD-1 in SalI and AgeI. After TUAD-1 is filled, single or multigenic transfer units, along with their respective TULs, can be prepared in SalI and NotI and placed into the TUAD-2 position in XhoI and NotI. Finally, single or multigenic transfer units, along with their respective TULs, can be prepared in SalI and NotI and inserted in reverse into the TUAD-3 position in PspOMI and SalI, which have adhesive ends compatible with NotI. For multigenic vectors requiring a CMD at the 3' end, the user may need to construct the 3'-side final transfer unit within the 3'CMD-TU shuttle vector.

[0067] The GIC modules flank the 5' and 3' ends of the pBOP-40 genome integration target receptor shuttle vector, respectively, with the 5' GIC defined by PstI and AscI, and the 3' GIC by NotI and XbaI. Depending on the biochemical activity of the GICs, one or both GIC modules may be occupied. The GIC sequences may encode homologous recombination arms usable with or without site-specific nucleases, such as transposase integration donor sites like the Sleeping Beauty transposon system or piggyBac transposons, small tyrosine recombinase sites like the Cre-related lox site or the Frp-related Frt site, large serine recombinases like SF370, PhiC31, or Bxb1, or ROSA26 or HPRT loci.

[0068] The chromatin modification domain (CMD) module flanks the 5' and 3' ends of the pBOP-40 genome integration target receptor shuttle vector, respectively. Based on the arrangement of restriction enzyme sites on the 5' side, the 5' CMD is defined by AscI and SphI, and the 3' CMD by NsiI and NotI. The chromatin modification domain consists of DNA sequences that can influence pre-transcriptional and transcriptional mechanisms. These include, but are not limited to, sequences that affect the epigenetic structure of genes, as well as sequences that can promote or inhibit transcription. Examples include long-range cis-regulatory elements known as insulators, such as the CTCF insulator and the β-globin locus. Insulators exhibit two functions: a) enhancer-blocking insulators prevent distal enhancers from acting on the promoters of adjacent genes, and b) barrier insulators prevent euchromatin silencing due to the diffusion of adjacent heterochromatin.

[0069] Positive selection factor elements (PSEs) encode genes that protect cells from toxins. PSE modules are designed to encode transcription units with a different structure from pBOP-30. A PSE module is defined by the sequence SphI-spacer-HindIII-spacer-EcoRI-spacer-SpeI, with the entire PSE module delimited by a 5' SphI site and a 3' SpeI site. In the context of pBOP-40, PSE modules are positioned in antisense orientation to prevent the possibility of transcriptional read-over to downstream modules. PSE modules can be constructed within pBOP-40 or in a separate shuttle vector, in which case the R domain is delimited by SpeI and EcoRI, the E domain by EcoRI and HindIII, and the P domain by HindIII and SphI. Examples of PSEs include, but are not limited to, genes encoding enzymes that metabolize zeosin, kanamycin, neomycin, puromycin, and hygromycin.

[0070] The bioindicator element (BE) module encodes a gene that can provide a useful bioassay signal, such as a reporter gene. The BE module is designed to encode a transcription unit with a different structure from pBOP-30. The BE module is defined by the sequence SpeI-spacer-HindIII-spacer-EcoRI-spacer-PspOMI, and the entire BE module is delimited by the 5' SpeI site and the 3' PspOMI site. In the context of pBOP-40, the BE module is positioned in antisense orientation to prevent the possibility of transcriptional read-over to downstream modules. The BE module can be constructed within pBOP-40 or in a separate shuttle vector, in which case the R domain is delimited by PspOMI and EcoRI, the E domain by EcoRI and HindIII, and the P domain by HindIII and SpeI. Examples of BEs include, but are not limited to, a) fluorescent proteins, b) reporter gene enzymes that can generate optical or colorimetric signals by acting on defined substrates, and c) genes encoding reporter gene channels or pumps that enable the uptake of imaging agents into cells or tissues.

[0071] Before describing embodiments of the present invention in more detail, it should be understood that the present invention is not limited to the specific embodiments described herein and can be modified. Furthermore, the terms used herein are used solely for the purpose of describing specific embodiments and should not be constrained, as the scope of the present invention is limited only by the appended claims.

[0072] Where a numerical range is provided, each intermediate value between the upper and lower limits of that range is understood to be included in that range and encompassed by the invention, unless otherwise explicitly indicated by the context or description. Furthermore, any narrower range between any two values ​​within that range is also encompassed by the invention, unless otherwise explicitly indicated by the context or description.

[0073] Unless otherwise defined herein, all technical and scientific terms relating to the present invention have the same meanings as those generally understood by those skilled in the art to which the present invention pertains. While representative and illustrative methods and materials are described herein, identical or equivalent methods and materials may also be used in carrying out or testing the present invention.

[0074] All publications and patents cited herein are incorporated herein by reference, each publication or patent being individually and expressly incorporated by reference, and are used to disclose and explain the methods and / or materials cited in such publications. Any reference to a publication is intended to disclose that publication prior to the filing date and should not be construed as acknowledging that the present invention does not have any priority over that publication by prior art. It should also be noted that the stated publication date may differ from the actual date of prior art and should be independently verified.

[0075] In this specification and the attached claims, the singular forms "a," "an," and "the" are to be interpreted as including the plural unless the context clearly indicates otherwise. It should also be noted that claims may be constructed in a manner that excludes any element. Therefore, this statement is intended to provide a basis for including exclusive terms such as "solely" and "only," or negative limitations such as "a particular feature or element is absent," "a particular feature or element is excluded," or "a particular feature or element is not included" in the claims.

[0076] A reader of this disclosure will see that each embodiment described and illustrated herein has separable components and features and can be readily combined with features of other embodiments without departing from the scope or spirit of the invention. Any method may be carried out in the order described or in any other logically possible order. [Examples]

[0077] (Example 1) Figure 2 shows the vector identified as pBOP-30, which has the nucleotide sequence SEQ ID NO:1. Like other vectors of the present invention, pBOP-30 contains mutations made to the vector backbone, which are necessary to a) maintain the function of Ori, b) maintain the function of the kanamycin resistance gene, and c) remove restriction sites within the polyMOD framework. This is because these restriction sites must be specific to the polyMOD framework at the locations where the gene module needs to be inserted. The polyMOD framework of pBOP-30 is shown in the line drawing above the circular plasmid diagram and provides the order and identity of the restriction enzyme recognition sites separated by spacer sequences, which are inserted into the pBOP-30 backbone at the PstI site. The combination of the polyMOD framework and the backbone represents an "empty" vector and has the nucleotide sequence identity SEQ ID NO:1. The circular plasmid diagram of pBOP-30 shows the locations of the various domains and subdomains into which the gene elements, i.e., nucleotide segments, are inserted into the vector. As with all vectors disclosed herein, these domains represent the positions in which the indicated elements are placed within the vector by excising and replacing the spacers shown in the polyMOD framework diagram.

[0078] The pBOP-30 vector was used to construct the pBOP-30 system construct, identified as pBOP-38-CMV promoter-chiron splice unit-EGFP-bGHpA (pBOP38-CMEGBH), as shown in Figure 3. The pBOP-38-CMEGBH single-gene construct is designed for transient transfection in eukaryotic cells. The pBOP38-CMEGBH vector allows for the "substitution" of defined regulatory sequences and / or biological effector coding sequences. A key feature of pBOP38-CMEGBH is its ability to be used for iterative addition of transcription units by utilizing its property of having compatible sticky ends when cleaved with restriction enzymes SalI and XhoI.

[0079] (Example 2) There are at least two methods that can be used to construct multi-transcription unit vectors or multigenic constructs by taking advantage of the fact that the SalI and XhoI restriction sites have compatible sticky ends, and that when they are linked by ligation, the two separate sites are destroyed. All vectors disclosed herein are suitable for constructing multigenic constructs because they both include SalI and XhoI architectures for iteratively assembling multiple transcription units.

[0080] As shown in Figure 4, when using protocol #1, a method of transcription unit linking, a transcription unit (e.g., TU-1) from one reference vector can be inserted into the 3' end of another transcription unit vector (TU-2) by subcloning the SalI / XhoI-delimited fragment of TU-1 into the XhoI site of the other transcription unit vector (TU-2). Since SalI and XhoI are compatible sticky ends, there is a 50% probability that the TU-1 fragment will be inserted into the TU-2 vector in either the forward or reverse direction. In other words, a single XhoI restriction enzyme site forms an "acceptor site" into which a fragment beginning with SalI at the 5' end and ending with XhoI at the 3' end can be inserted. This is the sole and intentional single exemplary case within the entire BOP structure, and depending on the orientation of the insert, an overhang delimited by the insert is generated or destroyed.

[0081] Those skilled in the art will understand the possibility of constructing larger multigenic constructs in a single reaction vessel. While this is thermodynamically disadvantageous and does not allow control over the order of transcription unit components, it is possible to construct multigenic constructs in a single reaction vessel and then "find" the desired results using next-generation sequencing or PCR.

[0082] Given that protocol #1 does not allow control over the orientation of transcription units, protocol #2 enables control over left-to-right sense chain orientation, i.e., 5' to 3'. Protocol #2 utilizes the transcription unit linker TUL, defined by the ordering of the XhoI site, a spacer sequence designed for efficient restriction enzyme binding, and the NotI site. Since each transcription unit has a TUL at its 3' end, any transcription unit vector can function as either a) a transcription unit donor or b) a transcription unit vector receptor.

[0083] As shown in Figure 5, the TU-1 vector can function as a vector receptor for transcription units within the TU-2 vector. By cleaving the TU-1 vector with XhoI and NotI, and then dephosphorylating it with a phosphatase (e.g., CIP) to prevent autoligation, the TU-1 vector becomes a transcription unit receptor vector. The TU-2 donor sequence is prepared by cleaving the TU-2 vector with SalI and NotI. The resulting restriction enzyme-digested TU-1 receptor vector and TU-2 donor vector are subjected to agarose gel electrophoresis, and suitable fragments are gel-purified. Some transcription unit donor fragments may be similar in size to the backbone of the transcription unit donor vector and therefore may not be efficiently separated on the agarose gel. In such cases, the transcription unit donor vector can be cleaved with SalI, NotI, and AseI. AseI cleaves the vector backbone into two smaller fragments, allowing separation from the transcription unit donor fragment. A multigenic vector with two transcription units can be fabricated by ligating a TU-2 donor fragment with a TU-1 receptor vector fragment. During this process, the compatible sticky ends of SalI and XhoI are disrupted, resulting in the multigenic vector having a SalI site at the 5' end of the TU-1 module and a complete TUL at the 3' end of the TU-2 module, allowing for the addition of subsequent transcription unit modules. This is shown in Figure 5, where two dual-effector multigenic vectors are initially constructed and then linked using transcription unit linker modules. By repeating these steps, additional transcription units can be iteratively added.

[0084] Those skilled in the art will understand the possibility of integrating two different transcription unit linking protocols to perform a multigenic assembly reaction in a single reaction vessel. When using the integrated protocol, the TU-1 receptor vector is prepared by digestion with XhoI and NotI, one or more transcription unit donor fragments are prepared by digestion with SalI and XhoI, and the final transcription unit donor fragment is prepared by digestion with SalI and NotI.

[0085] (Example 3) Figure 6 shows a method for constructing a larger multimer, i.e., a multigenic construct, by maintaining TUL at the 3' end transcription unit. In this embodiment, four distinct transcription units TU-1, TU-2, TU-3, and TU-4 are ligated from left to right, i.e., from 5' to 3', in the order [(TU-1)+(TU-2)+(TU-3)+(TU-4)]. In step 1, two multigenic vectors, [(TU-1)+(TU-2)] and [(TU-3)+(TU-4)], are constructed. This is achieved by subcloning TU-2, separated by SalI / NotI, to TU-1 via XhoI / NotI, and separately by subcloning TU-4, separated by SalI / NotI, to TU-3 via XhoI / NotI. In step 2, the [(TU-3)+(TU-4)] multigen is excised from the vector by SalI / NotI and subcloned into the [(TU-1)+(TU-2)] multigen by XhoI / NotI, thereby obtaining a quad-multigen defined by [(TU-1)+(TU-2)] and [(TU-3)+(TU-4)].

[0086] (Example 4) A multigenic construct useful for vaccine development is constructed within the transcription unit structure of the pBOP-30 vector (not shown). To prevent intravector recombination during pDNA production and / or after transfection, the gene sequences of each transcription unit element R, E, and P must be sufficiently different from each other. Three different transcription units, TU-1, TU-2, and TU-3, are linked together according to the molecular structure [(TU-1)+(TU-2)+(TU-3)]. Each transcription unit contains a sequence encoding either an antigen or an immunoactivator as the E element. TU-1 is defined by the structure [(EF1alpha promoter)+(COVID-19 spike protein)+(bovine growth hormone polyA)]. TU-2 is defined by the structure [(CMV promoter)+(COVID-19 envelope protein)+(SV40 polyA)]. TU-3 is defined by the structure [(PGK promoter)+(IL12 multieffector)+(HSVTK polyA)]. First, the TU-2 fragment, delimited by SalI / NotI, is subcloned to TU-1 via XhoI / NotI to obtain the [(TU-1)+(TU-2)] multigenic construct. Next, the TU-3 fragment, delimited by SalI / NotI, is subcloned to the [(TU-1)+(TU-2)] multigenic construct via XhoI / NotI to obtain the [(TU-1)+(TU-2)+(TU-3)] multigenic construct. The resulting vector (data not shown) can be used as a pDNA-based vaccine that can be delivered by either electroporation or a nonviral transfection reagent.

[0087] (Example 5) Figure 7A shows pBOP-30-IRES, a multi-effector or polycistronic IRES vector system useful for constructing a polycistrone consisting of two ORFs. The polyMOD framework of pBOP-30-IRES is shown in the line drawing above the circular plasmid diagram and provides the order and identity of restriction enzyme recognition sites separated by spacer sequences, inserted into the pBOP-30 backbone at the PstI and XbaI sites. The base IRES vector is defined by the order of gene modules [(PvuI)+(ORF-1 element spacer)+(EcoRV)+(IRES element spacer)+(PacI)+(ORF-2 element spacer)+(ScaI)+(effector linker (EL))+(KasI)]. pBOP-30-IRES has the nucleotide sequence SEQ ID NO:2. In useful embodiments, red fluorescent protein RFP is pre-positioned at the ORF-1 element receptor and green fluorescent protein GFP is pre-positioned at the ORF-2 element receptor. This functions as an IRES function assay system. IRES function can be tested by confirming the presence of both RFP and GFP expression in transfected mammalian cells. Non-functional IRESs will show only RFP expression.

[0088] (Example 6) Figure 8A shows a method for constructing a multi-effector 2-cistron transcription unit within pBOP-30-IRES following a two-step process.

[0089] Step 1: Subcloning ORF-1 into the RFP module. To do this, the IRES shuttle is cleaved at the PvuI sticky end and EcoRV blunt end and dephosphorylated with bovine enteral phosphatase CIP. Next, ORF-1 is isolated by cleaving at the PvuI sticky end and ScaI blunt end. Finally, ORF-1 is ligated into the IRES shuttle. Ligation destroys the two blunt end sites, EcoRV and ScaI, and regenerates the sticky end PvuI site. The resulting polycistrone is [(PvuI)+(ORF-1)+(IRES-1)+(PacI)+(GFP)+(ScaI)+(EL)+(KasI)]. Those skilled in the art will understand the usefulness of this first product [(ORF-1)+(IRES)+(GFP)], as downstream GFP can serve as confirmation of transfection efficiency to test the function of ORF-1 in the target cell type. Cells exhibiting green fluorescence are expected to also express the ORF-1 protein, functioning as gene tracers, i.e., as reporter genes or expression indicators.

[0090] Step 2: Subcloning ORF-2 into the GFP module of the polycistrone constructed in Step 1. First, the [(ORF-1)(IRES-1)(RFP)] polycistrone is cleaved at the PacI and KasI sticky ends and dephosphorylated. Next, ORF-2 is cleaved at the PvuI and KasI sticky ends. PvuI and PacI have compatible sticky ends, which are destroyed by ligation. The resulting dual effector [(ORF-1)+(IRES-1)+(ORF-2)] can be a) used as a substrate for constructing a larger polycistrone, or b) can function as an effector module for subcloning into either pBOP-30 transcription unit in PvuI / KasI.

[0091] Figure 8B shows a multi-effector or polycistronic IRES shuttle system useful for constructing a polycistron consisting of three ORFs, i.e., a 3-cistron construct. Therefore, the final product includes a triple-effector polycistron. The method for constructing a triple effector polycistronic transfer unit follows a two-step process.

[0092] Step 1: Subcloning ORF-3 into the GFP module of the IRES-2 shuttle. To do this, the IRES-2 shuttle is cleaved at the PacI and KasI sticky ends and dephosphorylated by CIP. Next, ORF-3 is isolated by cleaving at the PvuI and KasI sticky ends. Finally, ORF-3 is ligated to the IRES-2 shuttle. PvuI and PacI have compatible sticky ends, which are destroyed by ligation, respectively, and the sticky end KasI site is regenerated. The resulting polycistrone is [(PvuI)+(RFP)(EcoRV)+(IRES2)+(PacI)+(ORF-3)+(ScaI)+(EL)+(KasI)].

[0093] Step 2: Subcloning [(IRES-2)+(ORF-3)] into the IRES linker IL module of the pBOP-30-IRES polycistrone. First, the [(ORF-1)(IRES-1)(ORF-2)] polycistrone is cleaved at the ScaI blunt end and KasI sticky end and dephosphorylated. Next, [(IRES-2)+(ORF-3)] is cleaved at the EcoRV blunt end and KasI sticky end. The two blunt end sites, ScaI and EcoRV, are destroyed by ligation. The resulting triple effector [(ORF-1)+(IRES-1)+(ORF-2)+(IRES-2)+(ORF-3)] can be a) used as a substrate for constructing a larger polycistrone, or b) can function as an effector module for subcloning into any pBOP-30 transcription unit in PvuI / KasI.

[0094] (Example 7) The polycistronic construct is constructed within the structure of the pBOP-30-IRES cloning vector (not shown). This cloning vector has a multi-effector polycistrone defined by [(PvuI)+(RFP ORF)+(EcoRV)+(IRES-1)+(PacI)+(GFP ORF)+(ScaI)+(effector linker (EL))+KasI]. Two distinct effector ORF-1 and ORF-2 are defined by the molecular structure [(PvuI)+(ORF-X)+(ScaI)+(EL)+(KasI)], respectively. Both ORF-1 and ORF-2 can be prepared by 1) de novo synthesis, 2) performing a prepared PCR reaction using a pDNA vector encoding a reference ORF as a template, or 3) isolating an ORF fragment from a reference transcription unit. Regardless of how ORF-1 and ORF-2 are obtained, it is essential to first define the desired order of ORF-1 and ORF-2 in the desired multi-effects polysystron.

[0095] The multi-effector transcription unit is constructed with the terminal structure [(ORF-1)+(IRES-1)+(ORF-2)]. ORF-1 encodes brain-derived neurotrophic factor BDNF-1, ORF-2 encodes apolipoprotein E-epsilon-2, APOE ε2, and IRES-1 encodes EMCV IRES, the internal ribosome entry site of encephalomyocarditis virus. The construction method involves the following two steps to construct the [(BDNF-1)+(EMCV IRES)+(APOE ε2)] multi-effector element.

[0096] Step 1: Subcloning ORF-1 into the GFP module. To do this, the IRES shuttle is cleaved at the PvuI sticky end and EcoRV blunt end, and dephosphorylated by CIP. Next, ORF-1 is isolated by cleaving at the PvuI sticky end and ScaI blunt end. Finally, ORF-1 is ligated into the IRES shuttle. Ligation destroys the two blunt end sites, EcoRV and ScaI, and regenerates the sticky end PvuI site. The resulting polycistrone is [(PvuI)+(ORF-1)+(IRES-1)+(PacI)+(RFP)+(ScaI)+(EL)+(KasI)].

[0097] Step 2: Subcloning ORF-2 into the RFP module of the polycistrone constructed in Step 1. First, the [(ORF-1)(IRES-1)(RFP)] polycistrone is cleaved at the PacI and KasI sticky ends and dephosphorylated. Next, ORF-2 is cleaved at the PvuI and KasI sticky ends. PvuI and PacI have compatible sticky ends and are destroyed by ligation, respectively. The resulting dual effector [(ORF-1)+(IRES-1)+(ORF-2)] can be used as a substrate for constructing a larger polycistrone, or it can function as an effector module for subcloning into either pBOP-30 transcription units in PvuI / KasI.

[0098] (Example 8) The construct containing both protein-coding effectors and intron-coding bioRNAs is constructed within the structure of a cloning vector identified as pBOP-30-bioRNA-2. The pBOP-30-bioRNA-2 vector has the nucleotide sequence SEQ ID NO:4. This cloning vector has a multi-effector polycistrone defined by [(PvuI)+(ORF)+(EcoRV)+(bioRNA-1)+(ScaI)+(effector linker (EL)+KasI)]. The protein-coding effector (ORF-1) is defined by the molecular structure [(PvuI)+(ORF-X)+(ScaI)+(EL)+(KasI)]. Both ORF-1 and bioRNA-1 can be prepared by 1) de novo synthesis, 2) performing a prepared PCR reaction using a pDNA vector encoding a reference ORF as a template, or 3) isolating an ORF fragment from a reference transcription unit. Regardless of the method of obtaining ORF-1 and bioRNA-1, it is essential to first define the desired order of ORF-1 and bioRNA-1 in the desired multi-effector construct.

[0099] In this example, as shown in Figure 9D, the multi-effector transcription unit is constructed with the terminal structure [(ORF-1)+(bioRNA-1)]. ORF-1 encodes brain-derived neurotrophic factor BDNF-1, and bioRNA-1 encodes intron 16 of the human heterogeneous nuclear ribonucleoprotein K gene HNRNPK, which contains a copy of the human microRNA miR7-1. The method for constructing the [(BDNF-1)+(HNRNPK intron 16)] multi-effector element includes the following steps.

[0100] First, bioRNA-1 is subcloned into the EL domain. To do this, the ORF-1-containing vector is cleaved at the ScaI blunt end and KasI sticky end, and then dephosphorylated with bovine enteric phosphatase CIP. Next, bioRNA-1 is isolated by cleaving at the EcoRV blunt end and KasI sticky end. Finally, bioRNA-1 is ligated into the ORF-1-containing vector. Ligation disrupts the two blunt-end sites, EcoRV and ScaI, introduces a new ScaI site and IL domain, and regenerates the sticky-end KasI site. The resulting construct is [(PvuI)+(ORF-1)+(bioRNA-1)+(ScaI)+(IL)+(KasI)]. The resulting dual effector [(ORF-1)+(bioRNA-1)] can be used as a substrate for constructing larger polycistrons, or it can function as an effector module for subcloning into any pBOP-30 transcription unit in PvuI / KasI.

[0101] (Example 9) Figure 10A shows the pBOP-40 vector, which can be used for gene delivery, virus production, and nonviral genome integration. The pBOP-40 vector has the nucleotide sequence SEQ ID NO: 5. The polyMOD framework of pBOP-40 is shown in the line drawing above the circular plasmid diagram and provides the order and identity of restriction enzyme recognition sites separated by spacer sequences, which are inserted into the pBOP-40 backbone at the PstI and XbaI sites.

[0102] The detailed structure of pBOP-40 when all elements are assembled is defined by [(PstI)+(5' Genome Integration Regulator (GIC))+(AscI)+(5' Chromatin Modification Domain)+(SphI)+(Positive Selector Element (PSE))+(SpeI)+(BE)+PspOMI+(Spacer)+(SalI)+(R Element)+(PvuI)+(E Element)+(ScaI)+(EL Element)+(KasI)+(P Element)+(AgeI)+(XhoI)+(TUL Element)+(NsiI)+(3'CMD)+(NotI)+(3'GIC)+(Vector Backbone (VB))+(PstI)]. The PstI region shown in the figure is depicted as a single region, which indicates the location where the closed circular plasmid is joined, representing that the two ends are joined at one point. The right-to-left arrows shown below the PSE and BE elements indicate that these transcription units are positioned in reverse orientation, i.e., antisense orientation.

[0103] Figure 11A further illustrates the structures of the two transcription unit receptor domains, TUAD-1 and TUAD-2. A single pBOP-30 system transcription unit can be placed in TUAD-1 at SalI / AgeI. A pBOP-30 system multigenic construct separated by SalI / XhoI can be placed in pBOP-40 at SalI / XhoI or XhoI, although this is a thermodynamically difficult event because the sticky ends of SalI / XhoI or XhoI attempt to self-anneal. A pBOP-30 system multigenic construct separated by SalI / NotI can be subcloned into pBOP-40 at XhoI / NotI. Unlike the TUL of the pBOP-30 system, the TUAD of pBOP-40 contains a CMD element separated by NsiI / NotI. If a 3'CMD is required in a pBOP-40 system multigenic vector, it may be necessary to construct a transcription unit based on TUAD-2.

[0104] Figure 10B shows an example of a construct built within the pBOP-40 vector. Here, the identity of the DNA module assembled with the pBOP-40 vector is shown, rather than the general reception sites of each element shown in Figure 10A. Figure 10C shows a typical step for constructing a construct within the pBOP-40 vector. The TUAD-1 position is always used first and can be used to position single effectors prepared in PvuI and KasI into the transcription unit effector domain of TUAD-1 in PvuI and KasI, or to position single transcription units without TULs prepared in SalI and AgeI into TUAD-1 in SalI and AgeI. After TUAD-1 is filled, single transcription units or multigenic transcription units with their respective TULs, prepared in SalI and NotI, can be positioned at the TUAD-2 position in XhoI and NotI. Finally, single or multigenic transcription units with TULs prepared in SalI and NotI can be inserted in reverse into the TUAD-3 position in PspOMI and SalI, which have sticky ends compatible with NotI. For multigenic vectors requiring a CMD at the 3' end, the user may need to construct the 3' final transcription unit within a 3'CMD-TU shuttle vector.

[0105] (Example 10) This embodiment provides a method for constructing a vector for integration into a CHO cell line at a well-characterized genomic locus containing the cassette exchange receptor domain of a large serine recombinase. The final product encodes constitutive expression of a single-chain antibody Fc fusion protein and expression of a tetracycline-inducible glycosyltransferase.

[0106] The construct is assembled within the pBOP-40 vector, which is designed to facilitate genome integration by matching recombinase donor sites located within the 5' and 3' genome integration regulatory domains with corresponding recombinase acceptor sites pre-positioned at suitable CHO genome loci.

[0107] The transcription unit TU-1, which encodes a single-chain antibody and an Fc domain fusion protein, is assembled within the pBOP-30 vector. The assembled TU element is then subcloned into the TUAD-1 receptor domain of pBOP-40 in SalI / AgeI to obtain pBOP-40-TU1. The pBOP-30 multigenic vector encoding tetracycline-inducible glycosyltransferase is also constructed with the structure [(TU-2)+(TU-3)], where TU-2 encodes a tetracycline-inducible promoter R element that drives glycosyltransferase expression, and TU-3 encodes a ubiquitous and constitutive promoter R element that drives the expression of tetracycline regulatory proteins. Next, the [(TU-2)+(TU-3)] multigenic element, separated by SalI / NotI, is subcloned into the TUAD-2 receptor domain of the pBOP-40-TU1 vector at XhoI / NotI to obtain the final vector pBOP-40[(TU-1)+(TU-2)+(TU-3)]. This vector is inserted into the large serine recombinase receptor domain of CHO cells via recombinase-mediated cassette exchange reaction (RMCE). This allows CHO cells to express single-chain antibody Fc fusion proteins and tetracycline-inducible glycosyltransferases.

[0108] (Example 11) Bird of Prey TMThe modularity of the system allows for the efficient design and construction of a genomic integration negative selector backbone (GINSB), into which libraries of different GID modules can be inserted. Figure 11A shows the modular structure of the GINSB, and Figure 11B shows a map of pBOP-40-NS, a pBOP-40 vector containing the GINSB. The nucleotide sequence of pBOP-40-NS is SEQ ID NO: 6. The polyMOD framework of pBOP-40 is shown as a line drawing above the circular plasmid diagram, providing the order and identity of restriction enzyme recognition sites separated by spacer sequences. This polyMOD was inserted into the pBOP-30 backbone using NsiI at the 5' end and NheI at the 3' end. NsiI and PstI have compatible sticky ends, and NheI and XbaI also have compatible sticky ends; therefore, the PstI and XbaI sites are maintained for the subsequent placement of the CID domain, thereby creating pBOP-40. The pBOP-40-NS vector has the nucleotide sequence SEQ ID NO:6. While it is not essential for both negative selector domains to be occupied for GINSB function, the overall selectivity is improved by placing two different negative selector genes at the NegSel-1 and NegSel-2 positions.

[0109] The vectors of the present invention enable the integration of foreign genes into the genome by either an integrase-based method or a homologous recombination-based method. The efficiency of these processes can be influenced by numerous factors, including the promotion of genome site selection. For example, using a nuclease that induces site-specific genome breaks can increase the probability of homologous recombination-mediated genome repair events. Regardless of the method used to integrate the desired foreign gene into the genome, there is always a possibility of random integration events. Confirmation of genome integration is often evaluated by killing cells that have not incorporated the positive selection factor gene with a toxic chemical. Examples of positive selection factors include, but are not limited to, the NeoR, PuroR, HygroR, and ZeoR genes. Historically, to eliminate random genome integration events, researchers have flanked the desired "genome integration regulator" with negative selection factor genes. Historical examples of negative selection factors used in the creation of knockout or knock-in transgenic animals include, but are not limited to, chemoselection genes (e.g., HSVTK, CDA, etc.) and / or optical selection genes (e.g., fluorescent proteins, chromophores, etc.).

[0110] One example of using GINSB is a configuration in which a short-lived red fluorescent protein is placed at the NegSel-1 position and nothing is placed at the NegSel-2 position. In this example, the expression of the red fluorescent protein degRFP, flanked by degron, is driven by the strong promoter CMV, and GID encodes the Sleeping Beauty transposon donor site. When co-introduced with the Sleeping Beauty transposase, the transposase can be supplied as a plasmid expression vector, mRNA, or recombinant protein, but GINSB provides short-term red fluorescent protein expression until the transposer donor vector is degraded or diluted by successive cell divisions. Since fluorescent proteins can have long half-lives, the use of degron ensures that RFP expression is driven by the strong CMV promoter rather than protein stability. If RFP expression persists after multiple cell divisions, it suggests that abnormal random genomic integration has occurred. The use of GINSB is particularly useful when pursuing site-directed genomic integration by homologous recombination.

[0111] (Example 12) Figure 12 shows the pBOP-30-IVT in vitro transcription vector, which includes a reverse-complementary BsgI site for cleaving the polyA site within the polyA sequence. The pBOP-30-IVT vector can accommodate the reverse-complementary BsgI site to linearize the polyA site within the polyA sequence. The polyMOD framework of pBOP-30-IVT is shown as a line drawing above the circular plasmid diagram, providing the order and identity of restriction enzyme recognition sites separated by spacer sequences, inserted into the pBOP-30 backbone at the PstI and XbaI sites. The nucleotide sequence of pBOP-30-IVT is SEQ ID NO:7. The molecular structure of pBOP-30-IVT is defined by [(PstI)+(SalI)+(in vitro transcription promoter (IVTP))+(AfeI)+(RR receptor domain)+(PvuI)+(ORF element)+(KasI)+(P-5 receptor domain)+(FspI)+(polyA element)+(reverse complement BsgI)+(AgeI)+(XbaI)+(pBOP-30 vector backbone (VB))].

[0112] (Example 13) This example provides a representative method for constructing a multi-effector construct encoding BDNF and APOE ε2 effector proteins within a pBOP-30-IVT vector (not shown). This method provides a pBOP-30 system in vitro transcription vector driven by a T7 in vitro transcription promoter and transcribes multi-effector polycistronic mRNA. First, the pBOP-30-IVT vector is cleaved with PvuI and KasI and dephosphorylated. Then, the PvuI / KasI-separated multi-effector [(BDNF-1)+(EMCV IRES)+(APOE ε2)] polycistrone is subcloned into the PvuI / KasI site of the pBOP-30-IVT vector to obtain a final vector having a sequence encoding [(T7 promoter)+([(BDNF-1)+(EMCV IRES)+(APOE ε2)])+(polyA)]. The obtained vector is amplified and purified using a standard plasmid preparation and purification protocol. Next, this DNA is linearized with BsgI and purified on an agarose gel. The resulting linear DNA is then used as a template for RNA synthesis using T7 RNA polymerase.

[0113] (Example 14) This embodiment provides a representative "BOPization rule" for preparing nucleotide sequences for insertion into a BoP vector and offers a high-level explanation of the ORF-focused BOPization process.

[0114] A transcriptional unit (TU) is the functional unit of BoP, also known as a "gene," and consists of all elements from the transcriptional regulatory domain (TR) to the transcriptional processing domain (TP), flanked by the upstream SalI (GTCGAC) site and the downstream AgeI (ACCGGT) site. A TU is divided into four domains: the transcriptional regulatory domain (TR), the effector domain (E), the effector linker domain (EL), and the transcriptional processing domain (TP). To reduce the possibility of gene substitution when introduced into a living organism, it is desirable that the TR and TP domains do not originate from the same gene or genomic region.

[0115] The TR domain is flanked by an upstream SalI(GTCGAC) site and a downstream PvuI(CGATCG) site, and contains upstream regulatory elements necessary for the precise temporal and spatial expression of a gene. The TR domain may contain enhancers, chromatin modification domains, promoters, 5'UTR, and splice elements that can regulate the time, location, and duration of transcription, as well as post-transcriptional and pre-translational events. The TR domain can be divided into two subdomains: an RD subdomain flanked by SalI(GTCGAC) and AfeI(AGCGCT) restriction sites, and an RR subdomain flanked by AfeI(AGCGCT) and PvuI(CGATCG) restriction sites. In a common embodiment, DNA-encoded regulatory elements (e.g., enhancers, promoters, etc.) and transcription start sites (TSS, including up to 50 bp downstream of the TSS) are placed in the RD domain, and RNA-encoded regulatory elements (e.g., 5'UTR, introns, etc.) are placed in the RR subdomain. Furthermore, bioRNAs in the form of intron-coding microRNAs can also be placed in the RR subdomain. A common method for removing restriction sites from innate DNA sequences that do not fit into the TR domain is to introduce nucleic acid sequence changes that do not alter the predicted transcription factor binding sites.

[0116] The E domain is flanked by an upstream PvuI(CGATCG) site and a downstream ScaI(AGTACT) site and contains a "biological effect" element. The effector may be a protein-coding ORF or a non-coding, biologically active RNA (non-coding RNA, long non-coding RNA, microRNA, etc.). The E domain can be divided into three subdomains: the EN subdomain flanked by PvuI(CGATCG) and XmaI(CCCGGG) restriction sites, the EI subdomain flanked by XmaI(CCCGGG) and MfeI(CAATTG) restriction sites, and the EC subdomain flanked by MfeI(CAATTG) and ScaI(AGTACT) sites. A common method for removing restriction sites from innate protein-coding DNA sequences that do not fit the E domain is to introduce nucleic acid sequence changes by altering the protein-coding codon (as described in the BOPizer algorithm). On the other hand, a common method for removing restriction sites from natural RNA-coding DNA sequences that do not fit the E domain (e.g., microRNAs, lncRNAs) is to introduce nucleic acid sequence changes that do not affect the predicted folding structure of the corresponding RNA sequence.

[0117] The effector linker domain (EL) is flanked by an upstream ScaI (AGTACT) site and a downstream KasI (GGCGCC) site, maintaining the ability to add additional effectors to the transcription unit using IRES. Furthermore, bioRNAs in the form of intron-coding microRNAs can also be placed in the EL domain. The EL domain is not divided into subdomains. A common method for removing restriction sites from native DNA sequences that do not fit into the EL domain is to introduce nucleic acid sequence changes that do not alter the predicted folding structure of the corresponding RNA sequence.

[0118] The transcriptional processing domain (TP) is flanked by an upstream KasI (GGCGCC) site and a downstream AgeI (ACCGGT) site, and may contain splicing elements, a 3'UTR, a DNA / RNA transcription termination factor crossover region, a polyadenylation (polyA) site (RNA-based), DNA-based elements, enhancers, and chromatin modification domains. All of these can regulate the time, location, and duration of transcription, as well as post-transcriptional and pre-translational events. The TP domain can be divided into two subdomains: the P-5 subdomain flanked by KasI (GGCGCC) and FspI (TGCGCA) restriction sites, and the P-3 subdomain flanked by FspI (TGCGCA) and AgeI (ACCGGT) restriction sites. A common embodiment involves placing RNA-encoded regulatory elements (e.g., 3'UTR, polyadenylation signals, etc.) in the P-5 domain and DNA-encoded regulatory elements (e.g., enhancers, silencers, insulators, etc.) in the P-3 subdomain. Furthermore, combinations of elements encoded in the TR and TP domains control where TU effectors are expressed. These important compartmentalizers are located upstream (i.e., in the TR domain) and downstream (i.e., in the TP domain) of native genes due to the three-dimensional structure of DNA / chromatin. A common technique for removing restriction sites from native DNA sequences that do not fit into the TP domain is to introduce nucleic acid sequence changes that do not alter the expected transcription factor binding sites.

[0119] Similar to the EL domain, the transcription unit linker (TUL) is a conserved domain that allows for the addition of more transcription units within the same vector backbone. The added transcription units should have distinct TR, E, and TP domains to prevent intravector recombination. The TUL is flanked by an upstream XhoI (CTCGAG) site and a downstream NotI (GCGGCCGC) site.

[0120] The algorithm for fitting protein-coding effector sequences to the BoP system is divided into three stages: Sanitize Sequence Input, BOPize Sequence, and Verify Sequence, as shown in Figure 14. In the first stage of Sanitize Sequence Input, shown in Figure 15, the input sequence is quality-checked and standardized. The algorithm checks whether the length of the protein-coding sequence is divisible by 3 to ensure that the correct protein reading frame is used. Next, all whitespace characters are removed and character cases are normalized.

[0121] In the second stage of BOPize Sequence, sequences corresponding to specific restriction sites are removed, as shown in Figure 16. As the first step in Figure 16, the algorithm scans the sequence to detect one of the following restriction sites: AarI (GCAGGTG or CACCTGC), AfeI (AGCGCT), AgeI (ACCGGT), AscI (GGCGCGCC), AseI (ATTAAT), AsiSI (GCGATCGC), BsgI (CTGCAC or GTGCAG), EcoRI (GAATTC), EcoRV (GATATC), FspI (TGCGCA), HindIII (AAGCTT), I-CeuI (CGTAACTATAACGGTCCTAAGGTAGCGAA; SEQ ID NO: 8), I-SceI (TAGGGATAACAGGGTAAT; SEQ ID NO: 8) NO:9), KasI (GGCGCC), MfeI (CAATTG), MluI (ACGCGT), NotI (GCGGCCGC), NsiI (ATGCAT), PacI (TTAATTAA), PmeI (GTTTAAAC), PspOMI (GGGCCC), PstI (CTGCAG), PvuI (CGATCG), SalI (GTCGAC), Sap I (GCTCTTC or GAAGAGC), SbfI (CCTGCAGG), ScaI (AGTACT), SmaI (CCCGGG), SpeI (ACTAGT), SphI (GCATGC), SrfI (GCCCGGGC), SwaI (ATTTAAAT), XbaI (TCTAGA), XhoI (CTCGAG), and XmaI (CCCGGG).

[0122] In the second step, the algorithm scans the sequence to detect the following elements necessary for efficient splicing of RNA transcripts: splice donor sites (AGGT), branch point sites (CTAAC, CTAAT, CTGAC, or CTGAT), and splice acceptor sites (CAGG).

[0123] To remove restriction sites or splice elements, the algorithm utilizes the fact that many amino acids are encoded by multiple sequences (for example, leucine is encoded by six codons: CTA, CTC, CTG, CTT, TTA, or TTG). Replacing the original codon with another codon encoding the same amino acid does not affect the amino acid sequence, but it can significantly alter the DNA sequence of the CDS. However, since the usage frequency of each codon is not constant, the algorithm systematically changes from the most commonly used codons to the less frequently used codons, and avoids using rare codons.

[0124] If a restriction site or splice element overlaps with a codon encoding alanine (A or Ala, GCA, GCC, GCG, or GCT), the original codon sequence is replaced with GCC. If this change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with GCT. If this change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element still remains, the original codon sequence is replaced with GCA. If this change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element still remains after all these changes, a wider window is explored by considering adjacent amino acid coding sequences. If the restriction site or splice element still cannot be removed, the algorithm provides the user with an error message.

[0125] If a restriction site or splice element overlaps with a cysteine-coding codon (C or Cys, TGC or TGT), the original codon sequence is replaced with TGC. If this change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with TGT. If this change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element still remains after all these changes, a wider window is explored by examining adjacent amino acid-coding sequences. If the restriction site or splice element still cannot be removed, the algorithm provides the user with an error message.

[0126] If a restriction site or splice element overlaps with a codon encoding aspartic acid (D or Asp, GAC or GAT), the original codon sequence is replaced with GAC. If this change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with GAT. If this change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element still remains after all these changes, a wider window is explored by examining adjacent amino acid coding sequences. If the restriction site or splice element still cannot be removed, the algorithm provides the user with an error message.

[0127] If a restriction site or splice element overlaps with a codon encoding glutamate (E or Glu, GAA or GAG), the original codon sequence is replaced with GAG. If this change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with GAA. If this change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element still remains after all these changes, a wider window is explored by examining adjacent amino acid coding sequences. If the restriction site or splice element still cannot be removed, the algorithm provides the user with an error message.

[0128] If a restriction site or splice element overlaps with a codon encoding phenylalanine (F or Phe, TTC or TTT), the original codon sequence is replaced with TTC. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with TTT. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element still remains after all these changes, a wider window is explored by considering adjacent amino acid coding sequences. If the restriction site or splice element cannot be removed, the algorithm provides the user with an error message.

[0129] If a restriction site or splice element overlaps with a codon encoding glycine (G or Glyc, GGA, GGC, GGG, or GGT), the original codon sequence is replaced with GGC. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with GGG. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with GGA. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element still remains after all these changes, a wider window is investigated by considering adjacent amino acid coding sequences. If the restriction site or splice element cannot be removed, the algorithm provides the user with an error message.

[0130] If a restriction site or splice element overlaps with a codon encoding histidine (H or His, CAC or CAT), the original codon sequence is replaced with CAC. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with CAT. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element still remains after all these changes, a wider window is investigated by considering adjacent amino acid coding sequences. If the restriction site or splice element cannot be removed, the algorithm provides the user with an error message.

[0131] If a restriction site or splice element overlaps with a codon encoding isoleucine (I or Ile, ATA, ATC or ATT), the original codon sequence is replaced with ATC. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with ATT. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element still remains after all these changes, a wider window is investigated by considering adjacent amino acid coding sequences. If the restriction site or splice element cannot be removed, the algorithm provides the user with an error message.

[0132] If a restriction site or splice element overlaps with a codon encoding lysine (K or Lys, AAA or AAG), the original codon sequence is replaced with AAG. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with AAA. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element still remains after all these changes, a wider window is investigated by considering adjacent amino acid coding sequences. If the restriction site or splice element cannot be removed, the algorithm provides the user with an error message.

[0133] If a restriction site or splice element overlaps with a codon encoding leucine (L or Leu, CTA, CTC, CTG, CTT, TTA or TTG), the original codon sequence is replaced with CTG. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with CTC. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with TTG. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with CTT. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains after all these changes, a wider window is investigated by considering adjacent amino acid coding sequences. If the restriction site or splice element cannot be removed, the algorithm provides the user with an error message.

[0134] If a restriction enzyme site or splice element overlaps with a codon encoding methionine (M or Met, the only ATG codon), the algorithm explores a wider window by considering adjacent amino acid coding sequences to remove the site. If the restriction site or splice element cannot be removed, the algorithm provides the user with an error message.

[0135] If a restriction site or splice element overlaps with a codon encoding asparagine (N or Asn, AAC or AAT), the original codon sequence is replaced with AAC. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with AAT. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element still remains after all these changes, a wider window is investigated by considering adjacent amino acid coding sequences. If the restriction site or splice element cannot be removed, the algorithm provides the user with an error message.

[0136] If a restriction site or splice element overlaps with a codon encoding proline (P or Pro, CCA, CCC, CCG, or CCT), the original codon sequence is replaced with CCC. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with CCT. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with CCA. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains after all these changes, a wider window is investigated by considering adjacent amino acid coding sequences. If the restriction site or splice element cannot be removed, the algorithm provides the user with an error message.

[0137] If a restriction site or splice element overlaps with a codon encoding glutamine (Q or Gln, CAA or CAG), the original codon sequence is replaced with CAG. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with CAA. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element still remains after all these changes, a wider window is explored by considering adjacent amino acid coding sequences. If the restriction site or splice element cannot be removed, the algorithm provides the user with an error message.

[0138] If a restriction site or splice element overlaps with a codon encoding arginine (R or Arg, AGA, AGG, CGA, CGC, CGG, or CGT), the original codon sequence is replaced with AGA. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with AGG. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with CGG. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains after all these changes, a wider window is investigated by considering adjacent amino acid coding sequences. If the restriction site or splice element cannot be removed, the algorithm provides the user with an error message.

[0139] If a restriction site or splice element overlaps with a codon encoding serine (S or Ser, AGC, AGT, TCA, TCC, TCG, or TCT), the original codon sequence is replaced with AGC. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with TCC. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with TCT. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains after all these changes, a wider window is investigated by considering adjacent amino acid coding sequences. If the restriction site or splice element cannot be removed, the algorithm provides the user with an error message.

[0140] If a restriction site or splice element overlaps with a codon encoding threonine (T or Thr, ACC, ACA, ACG, or ACT), the original codon sequence is replaced with ACC. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with ACA. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with ACT. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element still remains after all these changes, a wider window is investigated by considering adjacent amino acid coding sequences. If the restriction site or splice element cannot be removed, the algorithm provides the user with an error message.

[0141] If a restriction site or splice element overlaps with a codon encoding valine (V or Val, GTA, GTC, GTG or GTT), the original codon sequence is replaced with GTG. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with GTC. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with GTT. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element still remains after all these changes, a wider window is investigated by considering adjacent amino acid coding sequences. If the restriction site or splice element cannot be removed, the algorithm provides the user with an error message.

[0142] If a restriction enzyme site or splice element overlaps with a codon encoding tryptophan (W or Trp, the only TGG codon), the algorithm explores a wider window by considering adjacent amino acid coding sequences to remove the site. If the restriction site or splice element cannot be removed, the algorithm provides the user with an error message.

[0143] If a restriction site or splice element overlaps with a tyrosine-coding codon (Y or Tyr, TAC or TAT), the original codon sequence is replaced with TAC. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element remains, the original codon sequence is replaced with TAT. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element still remains after all these changes, a wider window is investigated by considering adjacent amino acid-coding sequences. If the restriction site or splice element cannot be removed, the algorithm provides the user with an error message.

[0144] If a restriction site or splice element overlaps with a stop codon (Ochre or Och, Amber or Amb, Opal or Opa, TAA, TAG or TGA), the original stop codon sequence is replaced with TAG. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site remains, the original stop codon sequence is replaced with TAA. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site remains, the original stop codon sequence is replaced with TGA. If this codon change removes the restriction site, the algorithm proceeds to the next restriction site. If the site or element still remains after all these changes, a wider window is investigated by considering the adjacent amino acid coding sequence. If the restriction site or splice element cannot be removed, the algorithm provides the user with an error message.

[0145] To confirm that this process does not introduce any amino acid substitutions, the most important part of the final post-processing step is sequence verification. In this step, as shown in Figure 17, the initial sequence is translated from DNA to amino acid sequence, and the BOP-modified sequence is translated from DNA to amino acid sequence, and then these are compared.

[0146] If the two amino acid sequences are identical, this QC step is passed; otherwise, an error message is sent to the user. After it is verified that any modifications to remove restriction enzyme sites or splice elements do not affect the amino acid sequence encoding the peptide, sequences are added to both the 5' and 3' ends of the CDS to enable de novo synthesis and vector construction. At the 5' end, located upstream of the translation start site ATG, an 8-nucleotide spacer sequence GAGAGAGA is placed, followed by the PvuI restriction endonuclease recognition site CGATCG. On the other hand, at the 3' end, located downstream of the translation termination sites TAA, TAG, or TGA, the ScaI restriction endonuclease recognition site AGTACT is placed, followed by an 8-nucleotide spacer sequence GAGAGAGA.

[0147] While the present invention has been described in relation to several representative embodiments, those skilled in the art will understand that the invention can be practiced with modifications within the spirit and scope of the appended claims. Therefore, the present invention should not be limited to the embodiments described herein, but should also include all modifications and their equivalents within the spirit and scope of the description herein.

Claims

1. A method for constructing a polycistronic or multigenic transgene, A step of linearizing a first plasmid DNA vector to form a first acceptor region including a first restriction region having a 5' end and a 3' end, A step of providing a first nucleotide fragment having 5' and 3' ends complementary to the 5' and 3' ends of the first restriction site, A step of inserting the first nucleotide fragment into the first acceptor site, A step of ligating the end of the first inserted nucleotide fragment to the end of the acceptor site, The ligation process involves the regeneration of the 3' end of the first restricting portion, the destruction of the 5' end of the first restricting portion, and the generation of a new 5' end of the first restricting portion. A step of performing the linearization step, the insertion step and the ligation step for one or more additional nucleotide fragments, Each of the one or more additional nucleotide fragments is produced in a process different from that of the first nucleotide fragment. The first insertion fragment and the one or more additional DNA fragments generate a multigenic and / or polycistronic transgene in which the ends are linearly linked together, A method that includes this.

2. In the method according to claim 1, A method wherein the selected plasmid DNA vector has nucleotide sequence identity selected from the group consisting of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, and SEQ ID NO:

7.

3. The method according to claim 1, A step of providing a set of interconnected plasmid DNA vectors having one or more unique acceptor sites and comprising a plurality of compatible genetically modified nucleotide fragments, wherein each of the plurality of compatible genetically modified nucleotide fragments has a 5' end and a 3' end that fit one or more of the one or more unique acceptor sites, and the first plasmid DNA vector is one of the vectors in the set. The steps include: excising the multigenic and / or polycistronic transgene from the first plasmid DNA vector, and inserting the first plasmid DNA vector into the second plasmid DNA vector included in the set; A method that further includes this.

4. A plasmid DNA cloning vector comprising one or more genetically modified unique acceptor sites, The vector is configured to sequentially and / or iteratively generate multigenic and / or polycistronic constructs. Plasmid DNA cloning vector.

5. In the plasmid DNA cloning vector according to claim 4, The plasmid DNA cloning vector has nucleotide sequence identity selected from the group consisting of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, and SEQ ID NO:

7. Plasmid DNA cloning vector.

6. In the plasmid DNA cloning vector according to claim 5, The aforementioned nucleotide sequence identity is SEQ ID NO:

1. Plasmid DNA cloning vector.

7. In the plasmid DNA cloning vector according to claim 5, The aforementioned nucleotide sequence identity is SEQ ID NO:

2. Plasmid DNA cloning vector.

8. In the plasmid DNA cloning vector according to claim 5, The aforementioned nucleotide sequence identity is SEQ ID NO:

3. Plasmid DNA cloning vector.

9. In the plasmid DNA cloning vector according to claim 5, The aforementioned nucleotide sequence identity is SEQ ID NO:

4. Plasmid DNA cloning vector.

10. In the plasmid DNA cloning vector according to claim 5, The aforementioned nucleotide sequence identity is SEQ ID NO:

5. Plasmid DNA cloning vector.

11. In the plasmid DNA cloning vector according to claim 5, The aforementioned nucleotide sequence identity is SEQ ID NO:

6. Plasmid DNA cloning vector.

12. In the plasmid DNA cloning vector according to claim 5, The aforementioned nucleotide sequence identity is SEQ ID NO:

7. Plasmid DNA cloning vector.

13. A combination of nucleotide molecules for assembling multigenic and / or polycistronic constructs, A plasmid DNA cloning vector genetically modified to have one or more unique acceptor sites, One or more fitted genetically modified fragments that can be inserted into at least one of the one or more unique acceptor sites, Includes, Each of the one or more fitted genetically modified fragments has a 5' end and a 3' end that fit to one of the one or more unique acceptor sites, Each of the plasmid DNA cloning vector and the one or more fitted genetically modified fragments is designed such that, by ligating one of the one or more fitted genetically modified fragments to one of the one or more unique acceptor sites, the 5' end of the acceptor site is destroyed and the 3' end of the acceptor site is regenerated. A combination of nucleotide molecules.

14. In the combination described in claim 13, The above-mentioned one or more unique acceptor sites include multiple acceptor sites, The plasmid DNA cloning vector further comprises a plurality of spacers, each having a spacer between each of the plurality of acceptor sites. However, this is subject to the condition that each spacer in the plasmid DNA cloning vector is different from any other spacer. combination.

15. In the combination described in claim 13, The plurality of spacers are selected from the group consisting of CAGAGTCCC, GGGAGGTTT, ACTCAAGG, GCAGAAGTC, AGCCAACCT, TGCCGAGTC, CCAGCCGCC, GAAGAGGT, CACTTCCTG, CTCTGAGCC, AGCTCAGT, and ATATCACGC. combination.

16. In the combination described in claim 13, Each of the one or more compatible genetically modified fragments is genetically modified so as not to encode restriction enzyme recognition sites for Sall, Pvul, KasI, Agel, Xhol, Notl, and Asel. combination.

17. In the combination described in claim 16, Each of the one or more compatible genetically modified fragments is further genetically modified so as not to encode a restriction enzyme recognition site of Seal. combination.

18. In the combination described in claim 16, Each of the one or more compatible genetically modified fragments is further genetically modified so as not to encode a restriction enzyme recognition site for Bsgl. combination.

19. A combination of nucleotide molecules and nucleotide molecular segments for assembling multigenic and / or polycistronic constructs, A first nucleotide molecule genetically modified to have one or more unique acceptor sites, One or more fitted genetically modified nucleotide molecular segments that can be inserted into at least one of the one or more unique acceptor sites, Includes, Each of the one or more fitted genetically modified nucleotide molecular segments has a 5' end and a 3' end that fit to one of the one or more unique acceptor sites, Each of the first nucleotide molecule and the one or more adapted genetically modified nucleotide molecular segments is designed such that, by ligating one of the one or more adapted genetically modified nucleotide molecular segments to one of the one or more unique acceptor sites, the 5' end of the acceptor site is destroyed and the 3' end of the acceptor site is regenerated. combination.

20. In the combination described in claim 19, The above-mentioned one or more unique acceptor sites include multiple acceptor sites, The first nucleotide molecule further comprises a plurality of spacers having spacers between each of the plurality of acceptor sites, However, this is subject to the condition that each spacer within the first nucleotide molecule is different from any of the other spacers. combination.

21. In the combination described in claim 19, The plurality of spacers are selected from the group consisting of CAGAGTCCC, GGGAGGTTT, ACTCAAGG, GCAGAAGTC, AGCCAACCT, TGCCGAGTC, CCAGCCGCC, GAAGAGGT, CACTTCCTG, CTCTGAGCC, AGCTCAGT, and ATATCACGC. combination.

22. In the combination described in claim 19, Each of the one or more compatible genetically modified nucleotide molecular segments is genetically modified so as not to encode restriction enzyme recognition sites for Sall, Pvul, KasI, Agel, Xhol, Notl, and Asel. combination.

23. In the combination described in claim 22, Each of the one or more compatible genetically modified nucleotide molecular segments is further genetically modified so as not to encode a restriction enzyme recognition site of Seal. combination.

24. In the combination described in claim 22, Each of the one or more compatible genetically modified nucleotide molecular segments is further genetically modified so as not to encode a restriction enzyme recognition site for Bsgl. combination.

25. A method for removing and replacing at least one restriction site or splice element from a nucleotide sequence for insertion into a cloning vector, A step of identifying at least one restriction site or splice element to be removed from the nucleotide sequence, A step of determining a desired substitution to be performed by a user according to a set of rules used by an algorithm, wherein the rules include the following: a) If at least one restriction enzyme site or splice element overlaps with a codon encoding the amino acid alanine (abbreviated as A or Ala, encoded by the codon sequence GCA, GCC, GCG, or GCT), The original codon sequence is replaced with GCC, and if the restriction region is removed by the codon change, the algorithm moves on to the next restriction region. If the relevant region or element remains, the original codon sequence is replaced with a GCT, and if the restriction region is removed by the codon change, the algorithm moves on to the next restriction region. If the relevant region or element remains, the original codon sequence is replaced with a GCA, and if the restriction region is removed by the codon change, the algorithm moves on to the next restriction region. If the aforementioned site or element remains after all these changes, a wider window may be considered by examining the sequences encoding adjacent amino acids. If the aforementioned restricted portion or splice element cannot be removed, the algorithm provides the user with an error message; b) If the restriction enzyme site or splice element overlaps with a codon encoding the amino acid cysteine ​​(abbreviated as C or Cys, and encoded by the codon sequence of TGC or TGT), The original codon sequence is replaced with TGC, and if the restriction region is removed by the codon change, the algorithm moves on to the next restriction region. If the relevant region or element remains, the original codon sequence is replaced with the TGT, and if the restriction region is removed by the codon change, the algorithm moves on to the next restriction region. If the aforementioned site or element remains after all these changes, a wider window may be considered by examining the sequences encoding adjacent amino acids. If the aforementioned restricted portion or splice element cannot be removed, the algorithm provides the user with an error message; c) If the restriction enzyme site or splice element overlaps with a codon encoding the amino acid aspartic acid (abbreviated as D or Asp, and encoded by the codon sequence GAC or GAT), The original codon sequence is replaced with the GAC, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the relevant region or element remains, the original codon sequence is replaced with the GAT. If the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned site or element remains after all these changes, a wider window may be considered by examining the sequences encoding adjacent amino acids. If the aforementioned restricted portion or splice element cannot be removed, the algorithm will provide the user with an error message; d) If the restriction enzyme site or splice element overlaps with a codon encoding the amino acid glutamic acid (abbreviated as E or Glu, encoded by the codon sequence GAA or GAG), The original codon sequence is replaced with a GAG, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the relevant region or element remains, the original codon sequence is replaced with the GAA; if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned site or element remains after all these changes, a wider window may be considered by examining the sequences encoding adjacent amino acids. If the aforementioned restricted portion or splice element cannot be removed, the algorithm will provide the user with an error message; e) If the restriction enzyme site or splice element overlaps with a codon encoding the amino acid phenylalanine (abbreviated as F or Phe, encoded by the codon sequence TTC or TTT), The original codon sequence is replaced with the TTC, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the relevant region or element remains, the original codon sequence is replaced with TTT, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned site or element remains after all these changes, a wider window may be considered by examining the sequences encoding adjacent amino acids. If the aforementioned restricted portion or splice element cannot be removed, the algorithm will provide the user with an error message; f) If the restriction enzyme site or splice element overlaps with a codon encoding the amino acid glycine (abbreviated as G or Glyc, and encoded by the codon sequence GGA, GGC, GGG, or GGT), The original codon sequence is replaced with GGC, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned region or element remains, the original codon sequence is replaced with GGG, and if the restriction region is removed by the codon change, the algorithm moves on to the next restriction region. If the aforementioned region or element remains, the original codon sequence is replaced with GGA, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned site or element remains after all these changes, a wider window may be considered by examining the sequences encoding adjacent amino acids. If the aforementioned restricted portion or splice element cannot be removed, the algorithm will provide the user with an error message; g) If the restriction enzyme site or splice element overlaps with a codon encoding the amino acid histidine (abbreviated as H or His, encoded by a CAC or CAT codon sequence), The original codon sequence is replaced with CAC, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned region or element remains, the original codon sequence is replaced with CAT, and if the restriction region is removed by the codon change, the algorithm moves on to the next restriction region. If the aforementioned site or element remains after all these changes, a wider window may be considered by examining the sequences encoding adjacent amino acids. If the aforementioned restricted portion or splice element cannot be removed, the algorithm will provide the user with an error message; h) If the restriction enzyme site or splice element overlaps with a codon encoding the amino acid isoleucine (abbreviated as I or Ile, encoded by the codon sequence ATA, ATC, or ATT), The original codon sequence is replaced with an ATC, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned region or element remains, the original codon sequence is replaced with ATT, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned site or element remains after all these changes, a wider window may be considered by examining the sequences encoding adjacent amino acids. If the aforementioned restricted portion or splice element cannot be removed, the algorithm will provide the user with an error message; i) If the restriction enzyme site or splice element overlaps with a codon encoding the amino acid lysine (abbreviated as K or Lys, encoded by the codon sequence AAA or AAG), The original codon sequence is replaced with AAG, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned region or element remains, the original codon sequence is replaced with AAA, and if the restriction region is removed by the codon change, the algorithm moves on to the next restriction region. If the aforementioned site or element remains after all these changes, a wider window may be considered by examining the sequences encoding adjacent amino acids. If the aforementioned restricted portion or splice element cannot be removed, the algorithm will provide the user with an error message; j) If the restriction enzyme site or splice element overlaps with a codon encoding the amino acid leucine (abbreviated as L or Leu, and encoded by the codon sequence CTA, CTC, CTG, CTT, TTA, or TTG), The original codon sequence is replaced with a CTG, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned region or element remains, the original codon sequence is replaced with a CTC, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned region or element remains, the original codon sequence is replaced with TTG, and if the restriction region is removed by the codon change, the algorithm moves on to the next restriction region. If the aforementioned region or element remains, the original codon sequence is replaced with CTT, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned site or element remains after all these changes, a wider window may be considered by examining the sequences encoding adjacent amino acids. If the aforementioned restricted portion or splice element cannot be removed, the algorithm will provide the user with an error message; k) If the restriction enzyme site or splice element overlaps with a codon encoding the amino acid methionine (abbreviated as M or Met, and encoded by a codon sequence consisting only of ATG), The algorithm examines a wider window by looking at sequences encoding adjacent amino acids in order to remove the aforementioned site. If the aforementioned restricted portion or splice element cannot be removed, the algorithm will provide the user with an error message; l) If the restriction enzyme site or splice element overlaps with a codon encoding the amino acid asparagine (abbreviated as N or Asn, and encoded by the codon sequence AAC or AAT), The original codon sequence is replaced with AAC, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned region or element remains, the original codon sequence is replaced with AAT, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned site or element remains after all these changes, a wider window may be considered by examining the sequences encoding adjacent amino acids. If the aforementioned restricted portion or splice element cannot be removed, the algorithm will provide the user with an error message; m) If the restriction enzyme site or splice element overlaps with a codon encoding the amino acid proline (abbreviated as P or Pro, and encoded by the codon sequence CCA, CCC, CCG, or CCT), The original codon sequence is replaced with CCC, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned region or element remains, the original codon sequence is replaced with a CCT, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned region or element remains, the original codon sequence is replaced with CCA, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned site or element remains after all these changes, a wider window may be considered by examining the sequences encoding adjacent amino acids. If the aforementioned restricted portion or splice element cannot be removed, the algorithm will provide the user with an error message; n) If the restriction enzyme site or splice element overlaps with a codon encoding the amino acid glutamine (abbreviated as Q or Gin, encoded by the codon sequence CAA or CAG), The original codon sequence is replaced with CAG, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned region or element remains, the original codon sequence is replaced with CAA, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned site or element remains after all these changes, a wider window may be considered by examining the sequences encoding adjacent amino acids. If the aforementioned restricted portion or splice element cannot be removed, the algorithm will provide the user with an error message; o) If the restriction enzyme site or splice element overlaps with a codon encoding the amino acid arginine (abbreviated as R or Arg, and encoded by the codon sequence AGA, AGG, CGA, CGC, CGG, or CGT), The original codon sequence is replaced with AGA, and if the restriction site is removed by the codon change, the algorithm moves on to the next restriction site. If the aforementioned region or element remains, the original codon sequence is replaced with AGG, and if the restriction region is removed by the codon change, the algorithm moves on to the next restriction region. If the aforementioned region or element remains, the original codon sequence is replaced with CGG, and if the restriction region is removed by the codon change, the algorithm moves on to the next restriction region. If the aforementioned site or element remains after all these changes, a wider window may be considered by examining the sequences encoding adjacent amino acids. If the aforementioned restricted portion or splice element cannot be removed, the algorithm will provide the user with an error message; p) If the restriction enzyme site or splice element overlaps with a codon encoding the amino acid serine (abbreviated as S or Ser, and encoded by the codon sequence AGC, AGT, TCA, TCC, TCG, or TCT), The original codon sequence is replaced with AGC, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned region or element remains, the original codon sequence is replaced with a TCC, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned region or element remains, the original codon sequence is replaced with TCT, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned site or element remains after all these changes, a wider window may be considered by examining the sequences encoding adjacent amino acids. If the aforementioned restricted portion or splice element cannot be removed, the algorithm will provide the user with an error message; q) If the restriction enzyme site or splice element overlaps with a codon encoding the amino acid threonine (abbreviated as T or Thr, encoded by the codon sequence ACC, ACA, ACG, or ACT), The original codon sequence is replaced with ACC, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned region or element remains, the original codon sequence is replaced with ACA, and if the restriction region is removed by the codon change, the algorithm moves on to the next restriction region. If the aforementioned region or element remains, the original codon sequence is replaced with ACT, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned site or element remains after all these changes, a wider window may be considered by examining the sequences encoding adjacent amino acids. If the aforementioned restricted portion or splice element cannot be removed, the algorithm will provide the user with an error message; r) If the restriction enzyme site or splice element overlaps with a codon encoding the amino acid valine (abbreviated as V or Vai, encoded by the codon sequence GTA, GTC, GTG, or GTT), The original codon sequence is replaced with GTG, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned region or element remains, the original codon sequence is replaced with a GTC, and if the restriction region is removed by the codon change, the algorithm moves on to the next restriction region. If the aforementioned region or element remains, the original codon sequence is replaced with a GTT, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned site or element remains after all these changes, a wider window may be considered by examining the sequences encoding adjacent amino acids. If the aforementioned restricted portion or splice element cannot be removed, the algorithm will provide the user with an error message; s) If the restriction enzyme site or splice element overlaps with a codon encoding the amino acid tryptophan (abbreviated as W or Trp, encoded by a single TGG codon sequence), The algorithm examines a wider window by looking at sequences encoding adjacent amino acids in order to remove the site in question. If the aforementioned restricted portion or splice element cannot be removed, the algorithm will provide the user with an error message; t) If the restriction enzyme site or splice element overlaps with a codon encoding the amino acid tyrosine (abbreviated as Y or Tyr, encoded by the codon sequence TAC or TAT), The original codon sequence is replaced with a TAC, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned region or element remains, the original codon sequence is replaced with TAT, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned site or element remains after all these changes, a wider window may be considered by examining the sequences encoding adjacent amino acids. If the aforementioned restricted portion or splice element cannot be removed, the algorithm will provide the user with an error message; u) If the restriction enzyme site or splice element overlaps with a codon encoding a stop codon (abbreviated as Ochre or Och, Amber or Amb, Opal or Opa, and encoded by the codon sequence TAA, TAG, or TGA), The original stop codon sequence is replaced with a TAG, and if the restriction region is removed by this codon change, the algorithm moves on to the next restriction region. If the region remains, the original stop codon sequence is replaced with TAA, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the region remains, the original stop codon sequence is replaced with TGA, and if the restriction region is removed by the codon change, the algorithm moves to the next restriction region. If the aforementioned site or element remains after all these changes, a wider window may be considered by examining the sequences encoding adjacent amino acids. If the aforementioned restricted portion or splice element cannot be removed, the algorithm provides the user with an error message; A process to generate an output DNA sequence identical to the amino acid sequence translated from the original DNA sequence, The process of verifying the output DNA sequence, A step of passing the verification step when the two amino acid sequences are identical, A step of adding nucleotides to the 5' and 3' ends of the output DNA sequence in order to enable de novo synthesis, site-directed mutagenesis and / or vector construction, At the 5' end, an 8-nucleotide spacer sequence (GAGAGAGA) is added upstream of the translation initiation site (ATG), followed by the PvuI restriction endonuclease recognition site (CGATCG). At the 3' end, a step is taken to add a recognition site for ScaI restriction endonuclease (AGTACT) downstream of the translation termination site (TAA, TAG, or TGA), followed by the addition of an 8-nucleotide spacer sequence (GAGAGAGA). A method that includes this.

26. A kit for constructing polycistronic or multigenic transgenes, It comprises a set of interrelated plasmid DNA vectors genetically engineered to have one or more unique acceptor sites, Each of the aforementioned interrelated plasmid DNA vectors is designed to allow the insertion and ligation of multiple compatible genetically engineered nucleotide fragments. Each compatible genetically engineered nucleotide fragment has a 5' end and a 3' end that are compatible with one of the one or more unique acceptor sites of any one of the related plasmid DNA vectors in the set. The first plasmid DNA vector is one vector in the set, Furthermore, one or more of the related plasmid DNA vectors are insertable into at least one other plasmid DNA vector in the set of related plasmid DNA vectors. kit.

27. In the kit according to claim 26, The aforementioned set of interrelated plasmid DNA cloning vectors includes at least several different vectors having nucleotide sequence identity selected from the group consisting of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, and SEQ ID NO:

7. kit.

28. In the kit according to claim 26, The aforementioned set of interrelated plasmid DNA cloning vectors includes vectors having nucleotide sequence identity of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6 and SEQ ID NO:

7. kit.