Targeted integration methods
The SDI system addresses the inefficiencies of random integration in recombinant protein production by using orthogonal DNA enzymes for targeted integration and selective removal of unwanted sequences, enhancing the reliability and efficiency of protein expression in eukaryotic cells.
Patent Information
- Application Number
- JP2022562134
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-04-08
- Filing Date
- 2021-04-06
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2041-04-06
AI Technical Summary
Current methods for producing recombinant proteins in eukaryotic cells, such as CHO cells, suffer from high biological noise and inefficiency due to random integration of DNA, leading to heterogeneous cell populations and unreliable protein expression, making it difficult to optimize gene cassette designs and compare variants.
A novel site-directed integration (SDI) system using orthogonal DNA enzymes for targeted integration of a donor vector into predefined genomic locations, followed by positive and negative selection to ensure integration at the correct site and remove unwanted sequences, thereby reducing biological noise and improving expression efficiency.
The SDI system allows for rapid and specific selection of cells with targeted integration, reducing the need for extensive screening and improving the consistency and efficiency of recombinant protein production by minimizing random integration events.
Smart Images

Figure 0007819113000016 
Figure 0007819113000017 
Figure 0007819113000018
Abstract
Description
[Technical Field]
[0001] The present invention relates to the field of optimized expression systems for the production of recombinant proteins. More specifically, it relates to cell-based methods utilizing targeted integration of donor vectors into specific predefined genomic locations in the genome of eukaryotic host cells, wherein said vectors and host cells contain nucleic acid components that allow for selective selection of those cells that have integrated the donor vector into the predefined genomic location in the host cell genome, and for detection and elimination of cells that have undergone any additional random integration events into other parts of the genome. [Background technology]
[0002] Over the past 30 years, recombinant protein therapeutics have evolved from a novelty to a dominant position among marketed drugs. Recombinant production of therapeutic proteins represents a market worth over $100 billion annually and plays a vital role in the global economy and advanced medical care. Therapeutic protein classes include recombinant proteins (insulin, growth factors, cytokines, and blood factors), vaccines (antigens, VLPs), and monoclonal antibodies. By far, the most dominant form is the monoclonal antibody. While some recombinant proteins can be produced in single-celled microorganisms such as Escherichia coli (E. coli), for more complex proteins, including the monoclonal antibody class, Chinese hamster ovary (CHO) cells are the exclusive host for production [1, 2].
[0003] The primary approach to generating high-performance therapeutic protein-producing cell lines in industry today is to introduce recombinant protein genes into the genome of host CHO cell lines via a random integration approach and then select / screen for individual cells with the integrated gene at an active genomic site in sufficient copy number to produce sufficiently high transcription and, at the same time, possess a phenotype capable of supporting high protein translation and secretion. This is a highly intensive and time-consuming process with a high inherent uncertainty and biological variation. The typical process duration ranges from 3 to 12 months, depending on the host cell growth, the level of automation performed, and the end point (e.g., if evaluation of long-term clonal stability is included).
[0004] One fundamental limitation of the random integration approach is low sampling, which is the cellular diversity in the transfected cell pool. Approximately only 0.1-1% of transfected cells integrate the recombinant DNA. Furthermore, this subpopulation is highly heterogeneous with respect to the integration location, copy number, and integrity of the integrated DNA. Adding to the inherent overall phenotypic diversity of CHO cells, which is inherent to CHO cells due to their high genomic and epi-genomic plasticity, finding high-producing clones is somewhat challenging. This is also why high variability in protein production from non-clonal stable pools is commonly observed (stochastic sampling of phenotypic diversity).
[0005] This lack of sampling and high biological noise also make it difficult to compare different gene cassette designs of therapeutic protein candidates for expression optimization. Comparing multiple variants through parallel generation of stable pools is a very intensive task, and high biological noise would make the results unreliable. The use of co-transfection of variant libraries is hampered by the fact that random integration typically results in the integration of multiple copies of the expression vector, and therefore, any cells generated through such a workflow will typically contain integrated copies from more than one gene cassette design. Improving protein expression through cell line engineering strategies based on random integration of effector genes is hindered for the same reasons.
[0006] One potential major improvement to all of the above limitations is the use of targeted integration (site-directed integration; SDI) of the gene of interest (GOI). In this scenario, a pre-identified genomic location known to support high and stable transcription is used as the target of the GOI. The use of a smart combination of pre-transfected sequences and vector design, including the use of co-transfected nucleic acid enzymes such as nucleases or recombinases, facilitates targeted integration and ensures that all cells in culture contain the correctly inserted GOI and therefore have high transcription rates. This would significantly reduce the number of clones required for cell line development (CLD) screening operations and reduce biological noise in comparisons of gene cassette design or cell line engineering efforts. Several technical solutions for targeted integration have been described in the art [3-6]. Nevertheless, challenges remain.
[0007] A common challenge for all strategies utilizing targeted integration is generating high and sufficient expression of the GOI, since typical solutions result in the integration of a single GOI copy. Other challenges and limitations of available solutions described in the art are outlined below.
[0008] The Flp-In system (based on Flp / flippase recombinase, also known as flippase recombinase) for targeted integration [7] is an example of a solution that utilizes a single recombinase recognition sequence in combination with a recombinase to enable targeted integration at a predefined genomic location. After recombinase activation, the entire expression vector is integrated at the recombinase recognition sequence. Integration at the recombinase recognition site inactivates one selectable marker and activates a second selectable marker, allowing cells with correct integration events to be selected. The major drawbacks of this solution are: (i) there is no mechanism to detect or remove cells with integrated additional copies of the expression vector due to random integration events; (ii) there is no mechanism to remove sequence regions, such as the plasmid backbone sequence and active selectable marker gene, that may negatively affect expression of the GOI; and (iii) the method has questionable flexibility in the selection of selectable markers, since activation of the selectable marker during integration results in the fusion of extra amino acids at the N-terminus, which may affect its functionality.
[0009] To avoid the presence of sequences that potentially have a negative effect on GOI expression after targeted integration, different solutions for cassette exchange reactions at predefined genomic locations have been described in the art [3-6]. An example of such a solution is disclosed by Rentschler [8]. The predefined genomic location utilizes an active selectable marker gene (GFP) flanked by two orthogonal recombinase recognition sequences, both of which are targets of the same recombinase. Similarly, the GOI in the expression vector is flanked by two recombinase recognition sequences that match the two present in the genome. Upon recombinase activation, cassette exchange can occur between the selectable marker cassette and the GOI cassette. Cells undergoing cassette exchange can be selected by the absence of GFP expression. The disadvantages of this type of solution are: (i) there is no mechanism to detect or remove cells with additional copies of the expression vector integrated by random integration events; and (ii) because selection of cells undergoing cassette exchange is initially based on the absence of an active gene product, the selection time point must be delayed to allow for GFP degradation / dilution.
[0010] Haghighat-Khah RE et al. disclose a two-step site-specific cassette exchange system in insects, namely, the Aedes aegypti mosquito and the Plutella xylostella moth. [9] This exchange system utilizes the phiC31 recombinase for integration of an expression vector at a predefined genomic location, followed by the use of a second recombinase (Cre or Flp) for expression of the plasmid backbone sequence. However, the exchange system of Haghighat-Khah RE et al. does not provide a means to distinguish between targeted and random integration events. Additionally, no means is provided for removing the selectable marker gene.
[0011] Yuan et al. described a recombinase-based method for generating transgenic cells free of selectable markers and vector backbones, utilizing PhiC31-mediated gene delivery to a pseudo-attP sequence naturally present in the genome of target cells
[10] . Selection of cells in which integration occurred was achieved via the presence of an active eGFP expression cassette in the expression vector, and an att-B-TK fusion gene, which is inactivated upon targeted integration, was used as a negative selection marker to eliminate random integration events in a second selection step. The selection system and the plasmid bacterial backbone were then excised using two other recombinases, Cre and Dre. Significant drawbacks of the method disclosed by Yuan, Y. et al. for adoption in recombinant protein production applications based on integration into a single predefined genomic location are that (i) the method does not provide a means to distinguish cells that have integrated only at the predefined location from cells that have integrated at both the predefined site and a random pseudo attP site, because an inactive TK gene results from both scenarios; (ii) the first selection step cannot be performed until the transient expression of the selection marker has disappeared, which requires time; and (iii) the first selection step does not distinguish between the desired integration, integration at the pseudo attP site, or a random integration event.
[0012] Thus, there remains a need in the art to identify improved expression systems for the production of recombinant proteins. In this regard, the following desirable properties are required: (a) Ability to support GOI expression levels versus random integration resolution; (b) rapid and specific selection of cells that have undergone integration at a predefined genomic location; (c) further means for detecting and removing cells that have undergone an unwanted integration event; (d) Means to avoid the presence of sequences that may have a negative effect on GOI expression in isolated cells There is a need in the art to improve existing SDI systems to combine the above into a single solution. Summary of the Invention [Means for solving the problem]
[0013] The above mentioned problems are now solved or at least alleviated by the provision of the methods and means further presented herein.
[0014] The present disclosure provides a novel solution for recombinant protein production utilizing site-directed integration (SDI) of a single-copy donor vector into a predefined genomic location in an isolated eukaryotic host cell. The disclosed SDI-based system is based on a unique and inventive combination of well-established nucleic acid components for efficient integration of the donor vector into a dedicated target site in the host cell. The method provides specific positive selection of host cells that have integrated the donor vector into the dedicated predefined genomic location. The method also provides a step for detecting and optionally removing, by negative selection, any cells in which an unwanted integration event has occurred at another location in the host cell genome. This two-step selection method is unique and will be highly useful in the field of recombinant protein production.
[0015] As previously mentioned herein, an essential element to enable advanced applications such as improved cell line development, flexible cell line engineering, and simultaneous probing of gene construct libraries is increased control of recombinant gene integration into host cell lines and better control of their copy number, which is now provided by the present disclosure.
[0016] Although putative hotspot locations have been identified, the method presented herein was initially set up using Chinese hamster ovary (CHO) cells, as the SDI system should be applicable to any eukaryotic cell system, including mammalian cells such as human cells.
[0017] Thus, in a first aspect, the present disclosure relates to a method for targeted integration of a donor vector into a predefined genomic location in a eukaryotic cell, said method comprising: i) a. a nucleic acid sequence I1 comprising a recognition site for a first DNA enzyme; b. a nucleic acid sequence E1 containing a recognition site for a second DNA enzyme; and c. promoter nucleic acid sequence P1; providing a eukaryotic cell comprising a predefined genomic location comprising: ii) a. nucleic acid sequence I2; b. a nucleic acid sequence of interest; c. a nucleic acid sequence E2 comprising a recognition site for the second DNA enzyme; d. a nucleic acid sequence encoding a first selectable marker; and e. Optionally, an expression cassette encoding a second selectable marker; providing a donor vector; iii) contacting the donor vector with the cell in the presence of a first DNA enzyme, wherein the presence of the first DNA enzyme allows for recombination between the nucleic acid sequence I2 of the donor vector and the nucleic acid sequence I1 present at a predefined genomic location of the cell; iv) selecting cells having the donor vector integrated at a predefined genomic location by detecting expression of a first selection marker in the cells, wherein expression of the first selection marker is activated by the promoter nucleic acid sequence P1 at the predefined genomic location; and v) isolating the cells selected in the previous step Includes:
[0018] In a further aspect, the present disclosure relates to an isolated eukaryotic cell obtained by the methods described herein.
[0019] In yet a further aspect, the present disclosure relates to the use of an isolated eukaryotic cell obtained by the methods described herein for producing a recombinant protein.
[0020] In yet a further aspect, the present disclosure relates to a method of producing a recombinant protein, the method comprising: i) obtaining an isolated eukaryotic cell comprising a donor vector comprising one or more nucleic acid sequences of interest integrated into a predefined genomic location by performing the method disclosed herein, wherein at least one nucleic acid sequence of interest comprises at least one expression cassette comprising a gene encoding a protein of interest; ii) producing in the cells of step i) the protein encoded by the gene of interest; and iii) isolating the protein of step ii). Includes: [Brief explanation of the drawings]
[0021] [Figure 1-1] Figure 1. Schematic representation of the general concept of the method for targeted integration of a donor vector into a predefined genomic location in a host cell through the use of at least two DNA enzymes with orthogonal specificity. [Figure 1-2] Continued from Figure 1. [Figure 2] FIG. 1 shows a schematic representation of the landing pad design (nucleic acid sequences present at predefined genomic locations in the host cell genome) and matching donor vectors for LP1P1 and HyClone LP2P2 cell lines. [Figure 3]Figure 1 illustrates an example of a flow cytometry plot 7 days post-transfection compared to a non-transfected control (NC). The density of cells in the plot with enriched major populations is visualized by alternating black and white regions (20% of the total cells in each region). The top row shows FACS data for a non-transfected control (NC) culture of HyClone CHO cells. The middle row shows FACS data for a random integration control (RI) based on HyClone CHO cells (lacking LP) transfected with donor vector B only (no PhiC31). The bottom row shows FACS data for a HyClone CHO LP2P2 cell line transfected with PhiC31 and donor vector B (SDI). Gates (B, D, F) in the middle plots of each row were set based on the non-transfected control and report the percentage of cells activating the selectable marker above background. [Figure 4] FIG. 1 shows a schematic representation of the landing pad cell line and donor vector used and the changes in the landing pad predicted to occur via the activity of PhiC31(1) and Cre(2). [Figure 5] Figure 1 shows flow cytometry plots of SDI populations 7 days after transfection of Cre recombinase variants compared to negative mock-transfected controls. The density of cells in the plots with enriched major populations is visualized by alternating black and white regions (20% of the total cells in each region). The top panel shows a plot of a mock transfection lacking a nucleic acid molecule encoding Cre recombinase, the middle panel shows a plot of a population transfected with a Cre recombinase expression plasmid, and the bottom panel shows a plot of a population transfected with synthetic Cre recombinase mRNA. [Figure 6] FIG. 1 shows a schematic representation of the landing pad cell line and donor vector used and the changes in the landing pad predicted to occur through the activity of PhiC31 recombinase (1) and Cre recombinase (2). [Figure 7]Flow cytometry plots from steps performed according to Figure 6. The density of cells in the plot with enriched major populations is visualized by alternating black and white areas (20% of the total cells in each area). The left plot in the top panel shows the population after a second round of eGFP (green fluorescent protein)-positive sorting. The middle plot in the top panel shows the population 7 days after Cre recombinase transfection. The right plot in the top panel shows the population after eGFP-negative sorting performed with gate E after Cre recombinase transfection. The bottom panel shows the population plot 7 days after transfection of step 2 sorted cells with DNA donor vector B. [Figure 8] Figure 1 shows a schematic representation of the landing pad cell lines and donor vectors used. GSx = glutamine synthetase gene variant. [Figure 9] Figure 1 shows flow cytometry plots of cell population generation using SDI of donor vectors. The density of cells in the plot with enriched major populations is visualized by alternating black and white regions (20% of the total cells in each region). The top panel shows plots of populations that underwent G418 selection, RFP (red fluorescent protein)-positive FACS sorting, and transfection with synthetic Cre recombinase mRNA. eGFP histograms are shown for both the RFP-negative subpopulation (corresponding to integration at the landing pad using Cre recombinase-mediated excision of TagRFP-T) and the RFP-positive subpopulation (corresponding to failed Cre recombinase-mediated excision of TagRFP-T, which may be caused by off-target or truncated integration at the landing pad). The bottom panel shows plots of the final SDI pool generated using FACS sorting of the RFP-negative / GFP-positive cells from the top panel. [Figure 10-1] Figure 10. The donor vector's first selectable marker linked via an IRES element to a gene encoding the second DNA enzyme, both of which are activated upon integration at a predefined genomic location. [Figure 10-2] Continued from Figure 10. [Figure 11-1] Figure 11. A diagram showing the case where the predefined genomic region also contains an expression cassette for the first DNA enzyme, positioned such that upon integration of the donor vector at the predefined genomic location, the recognition sites for the second DNA enzyme are flanked and therefore can be removed in the presence of the second DNA enzyme. [Figure 11-2] Continued from Figure 11. [Figure 12-1] Figure 12. Diagram illustrating a variant of targeted integration in which a gene editing enzyme is used to catalyze the integration of a donor vector into a predefined genomic location in the host cell genome. [Figure 12-2] Continued from Figure 12. [Figure 13-1] Figure 13. A diagram showing a variant of targeted integration that uses recombinase-mediated cassette exchange (RMCE) to catalyze integration at a predefined genomic location in the host cell genome. [Figure 13-2] Continued from Figure 13. [Figure 14-1] Figure 14. A variant of targeted integration using a single recombinase recognition site pair to catalyze integration of a donor vector at a predefined genomic location in a host cell. [Figure 14-2] Continued from Figure 14. [Figure 15-1] Figure 15. Diagram showing the use of a single recombinase recognition site pair to catalyze integration at a predefined genomic location, operably fusing promoter P1, present at the predefined genomic location, to the 5' portion of a split intron. [Figure 15-2] Continued from Figure 15. DETAILED DESCRIPTION OF THE INVENTION
[0022] The present disclosure will now be described in more detail in connection with the accompanying drawings and several non-limiting examples.
[0023] definition The details of this disclosure are described below. Although any materials and methods similar or equivalent to those described herein can be used in the practice or testing of the present invention, the preferred materials and methods are described. All words and terms used herein shall be deemed to have the same meaning as commonly ascribed to them by one of ordinary skill in the art, unless a different meaning is apparent from the context.
[0024] A composition that "comprises" one or more recited components may also include other components not specifically recited.
[0025] The singular forms "a" and "an" will be construed as including the plural.
[0026] "Expression" is used to mean the production of a protein from a gene and refers to and includes the successful operation of the steps of the "central dogma", namely, transcription, translation, and protein folding to arrive at an active protein.
[0027] As defined herein, an "expression vector" is a vector that contains nucleic acid sequences for achieving protein expression from the vector when present in a host cell. As used herein, an expression vector is used, for example, to introduce a particular gene of interest into a cell and then direct the cellular machinery for protein synthesis to produce the protein of interest encoded by the gene of interest. An expression vector can contain an "expression cassette," which contains nucleic acid sequences for promoting protein expression. In addition, a vector may contain other nucleic acid sequence elements or components.
[0028] As referred to herein, a "donor vector" is a vector, preferably a DNA vector, that contains nucleic acid elements or components for facilitating integration of the vector into a predefined genomic location of an isolated eukaryotic host cell. The donor vector contains a nucleic acid sequence that facilitates a recombination event, including a nucleic acid sequence present in a predefined genomic location of the host cell, a nucleic acid sequence of interest that optionally encodes a protein of interest, a recognition site for a second DNA enzyme, and a nucleic acid sequence encoding a first selectable marker. Optionally, it may also contain an expression cassette for a second selectable marker. A "donor vector" may also sometimes be referred to herein as a "vector." A "donor vector" may sometimes be in the form of an expression vector, such as when the donor vector contains an expression cassette encoding a second selectable marker. More specifically, the donor vector described herein contains at least a nucleic acid sequence I2 for recombination with I1 present in a predefined genomic location of the eukaryotic cell. In addition, if the nucleic acid of interest encodes a protein of interest, it contains the nucleic acid sequence of interest, also referred to herein as a gene of interest ("GOI"). It also contains a nucleic acid sequence E2, which contains a recognition site for a second DNA enzyme that allows for excision of a portion of the vector backbone once stable integration of the donor vector occurs at a predefined genomic location in the host cell. It also contains a nucleic acid sequence encoding a first selectable marker (SM1), whose expression is achieved only if the donor vector is correctly integrated into the predefined genomic location in the host cell. Finally, the donor vector optionally contains an expression cassette encoding a second selectable marker (SM2). After the action of the second DNA enzyme, the second selectable marker will be expressed in the cell and be able to direct only if a random integration event of the vector occurs and be used in the second round of selection in the method. The donor vector is preferably, but not limited to, a DNA donor vector. DNA donor vectors are sometimes abbreviated as "DDV."
[0029] An "expression cassette" is a nucleic acid component that forms part of an expression vector that contains all the elements required for the initiation of transcription and translation of a protein of interest. The gene of interest that encodes the protein of interest also forms part of the expression cassette. The expression cassette contains, for example, a promoter, which is essential for the initiation of transcription, and other sequences that facilitate transcription, such as enhancer sequences. Sometimes, the term "integration cassette" is used herein, which corresponds to a nucleic acid sequence from a donor vector that remains in a predefined genomic location after the action of a second DNA enzyme. An "integration cassette" may include an "expression cassette."
[0030] As used herein, a "gene" of interest refers to the nucleic acid components required to produce a protein of interest, and can also refer to multiple genes of interest present in the same expression cassette, since a protein of interest can comprise multiple polypeptide chains. Expression cassettes containing multiple genes of interest can utilize individual promoters to achieve transcription of each gene, or two or more genes can be transcribed as a common mRNA, i.e., with each gene separated by an IRES element. This is in keeping with the idea that whenever "a" is used herein, it can refer to multiple. An example of when an expression cassette contains more than one gene of interest is when an antibody is to be expressed from the gene of interest, e.g., the light and heavy chain antibodies are present as separate genes in the expression cassette.
[0031] An "intron" is a nucleic acid sequence that is a gene once transcribed and removed by RNA splicing during production of the final RNA product. Introns are non-coding regions of an RNA transcript, or the DNA that encodes it, that are removed by pre-translational splicing.
[0032] As used herein, a promoter operably fused to the 5' portion of a split intron means that transcription of the 5' portion of the split intron is driven by the promoter. The 5' portion of a split intron is defined herein as comprising a splice donor site sequence (e.g., GT). The 3' portion of a split intron may be defined herein as comprising (i) a splice branch site sequence, (ii) a Py-rich sequence region, and (iii) a splice acceptor site sequence (e.g., AG).
[0033] Transcription involves the conversion of DNA into RNA by the cellular machinery. A "transcriptional control sequence" is a segment of nucleic acid sequence that has the ability to increase or decrease the ultimate expression of a particular gene, i.e., the sequence is capable of regulating the transcription of the gene. Examples of transcriptional control sequences are promoters, enhancers, and similar elements.
[0034] The untranslated region ("UTR") refers to either of the two sections on each side of the coding sequence on the mRNA strand: on the 5' side, it is called the 5'UTR, and on the 3' side, it is called the 3'UTR.
[0035] As referred to herein, an upstream open reading frame (uORF) is an open reading frame (ORF) in the 5' untranslated region (5'UTR) of an mRNA molecule. uORFs are generally involved in the regulation of eukaryotic gene expression. The translation of uORFs typically inhibits the downstream expression of the primary ORF (open reading frame), thus causing reduced protein expression when present. Approximately half of human genes contain these regions.
[0036] Internal ribosome entry sites ("IRES") are RNA elements that allow for translation initiation in a cap-dependent manner. They are often referred to as distinct regions of an RNA molecule that can recruit eukaryotic ribosomes to mRNA. The location of IRES elements is often in the 5'UTR region, but they can occur elsewhere in the mRNA.
[0037] A "plasmid" is a small, circular, extrachromosomal DNA molecule that can replicate independently of a cell and is found in bacteria. Plasmids are often used as vectors for molecular cloning, i.e., transferring and introducing selected DNA into a host cell. Plasmids are constructed from specific and essential elements and can contain genes that can be homologous or heterologous to the bacterial host cell. For example, plasmids always contain a bacterial origin of replication and most often contain specific antibiotic resistance genes.
[0038] As referred to herein, a "nucleic acid sequence of interest" may be defined as a nucleic acid sequence that one desires to incorporate into a cell to affect the functionality of said cell, which may include a gene of interest ("GOI") that encodes a protein of interest.
[0039] As used herein, a "recombinant" protein refers to a protein produced from an expression cassette introduced into a cell by an expression vector. Techniques for producing recombinant proteins are well known to those skilled in the art.
[0040] A "promoter" is a region of DNA that initiates transcription of a gene upon binding of RNA polymerase to it. A promoter is located near the transcription start site of a gene.
[0041] As referred to herein, a "host cell" relates to a eukaryotic cell that is intended to be transformed or has been transformed by a donor vector disclosed herein.
[0042] An "isolated cell," "isolated host cell," or "isolated eukaryotic host cell" refers to a cell that has been isolated from its natural environment, meaning that it does not contain any additional components that may occur in nature and is no longer part of that natural environment.
[0043] As used herein, a "predefined genomic location," sometimes referred to as a "landing pad" (abbreviated as "LP") or, rather, a predefined genomic location comprising a landing pad sequence, is intended to refer to a location, or nucleic acid location, characterized by a specific nucleic acid sequence in a host cell genome. A predefined genomic location may also be referred to herein as a "safe harbor site" and / or a "recombination site." At the predefined genomic location of the host cell, a recombination event between nucleic acid sequences 11 and 12, facilitated by the presence of a first DNA enzyme, occurs, thereby initiating expression of a first selectable marker, indicating a successful integration event. Essentially, the predefined genomic location comprises a nucleic acid sequence comprising a recognition site for a first DNA enzyme, a nucleic acid sequence comprising a recognition site for a second DNA enzyme, and a promoter nucleic acid sequence.
[0044] As used herein, "targeted integration" refers to the incorporation or introduction of a nucleic acid sequence element or component into another nucleic acid element or component, facilitating a recombination event between such sequences, thereby generating a hybrid sequence from the original sequence. Such an integration event is triggered by the presence of an enzyme-recognition nucleic acid sequence in any one or some of the nucleic acid sequence elements or components that form the basis for the recombination.
[0045] An "enzyme recognition site" refers to a particular combination of nucleotides in a nucleic acid sequence that is recognized by a specific enzyme, facilitating the binding of the enzyme to it, which then initiates an action at the recognition site, such as a recombination event between two sequences.
[0046] As referred to herein, the term "DNA enzyme" is defined as an enzyme that acts on DNA, such as by cutting pieces of DNA or by cutting DNA and incorporating it into another DNA sequence. This term includes enzymes such as Crisps / Cas9, recombinases, integrases, nucleases, etc., but the present disclosure is not limited thereto.
[0047] As referred to herein, a "first DNA enzyme" can be functionally defined as an enzyme that can participate in the integration of a donor vector at a predefined genomic location in a host cell in the methods disclosed herein. The function of the first DNA enzyme is to introduce, rather than remove, a nucleic acid sequence into a predefined genomic region. The first DNA enzyme can be one specific enzyme, or it can be different enzymes when used in the methods disclosed herein. This is the case, for example, when the integration of the donor vector is sequential, thereby repeating multiple times to introduce multiple copies / variants of the nucleic acid sequence of interest / donor vector into a predefined genomic location in the host cell, or when reversible integration of the nucleic acid sequence of interest is performed. Examples of "first DNA enzymes" for use in the context of the present methods are provided elsewhere herein.
[0048] As referred to herein, a "second DNA enzyme" can be functionally defined as an enzyme that can participate in the methods disclosed herein in excising a nucleic acid sequence region from a predefined genomic location into which the donor vector has been integrated, said nucleic acid sequence region being flanked by specific sequences recognized by the second DNA enzyme. When the second DNA enzyme recognizes a sequence, it will cleave the nucleic acid sequence component between these sequences. Examples of "second DNA enzymes" for use in the context of the present methods are provided elsewhere herein.
[0049] "In the presence of a first DNA enzyme" and / or "in the presence of a second DNA enzyme" means that the first and / or second DNA enzyme is provided in any form described herein, e.g., as a protein expressed from a donor vector, a separate expression vector, an expression cassette present in the genome of the cell, synthetic mRNA, etc. "In the presence" is intended to indicate that the function of the first DNA enzyme and / or the second DNA enzyme is provided in any suitable manner disclosed herein.
[0050] As referred to herein, a "selection marker" is, for example, in the present context, a marker (first selection marker) that can indicate that a specific event has occurred, that is, that the integration of the donor vector has occurred at a predefined genomic location in the host cell. The selection marker is often a fluorescent protein that is expressed by the host cell once the donor vector has been integrated at the correct site in the host cell genome. The expression of the fluorescent protein can be detected, for example, by FACS (fluorescence-activated cell sorting). Other possible selection markers are described elsewhere herein.
[0051] A "first selection marker," also abbreviated herein as "SM1," can be defined as a silent, inactive, or promoterless selection marker when present in a donor vector. The first selection marker contains a non-coding section that is compatible with a promoter present at a predefined genomic location. When the donor vector is integrated into the correct position at the predefined genomic location, the first selection marker can be expressed because it now has a promoter to initiate transcription. Once the selection marker is expressed, a cell population expressing the first selection marker can be selected as positive for stable integration of the donor vector at the predefined genomic location. The first selection marker can also be referred to herein as a "reporter." Examples of suitable first selection markers are provided elsewhere herein.
[0052] A "second selectable marker," also abbreviated herein as "SM2," is an optimal characteristic of the donor vectors disclosed herein and can be defined as a non-silent, active and / or functional selectable marker when present in the donor vector. The selectable marker is encoded as part of an expression cassette; i.e., the selectable marker is transiently expressed upon entry into a cell and will then promote stable expression independent of where in the genome it is introduced. The second selectable marker, in most embodiments of the methods presented herein, is a negative selectable marker, meaning that cells expressing this marker are preferably not used for recombinant protein production, since these cells will also integrate the donor vector at locations other than the predefined genomic location.
[0053] Detailed Description Provided herein are methods utilizing a specific site-directed integration (SDI) system for targeted and detectable integration of a donor vector into a predefined genomic location in an isolated eukaryotic cell. In addition, the methods allow for the identification of random integration events of the donor vector into other parts of the eukaryotic host cell genome other than the predefined location.
[0054] In combination, this provides a "double" selection of a population of cells that have positively integrated the donor vector at a target site (a predefined genomic location) in the host cell genome, preferably in the absence of additional random integration of the donor vector at other locations in the host cell genome, thereby providing an optimized system for subsequent recombinant protein expression. The method uses a combined selection strategy based on positive (integration at a predefined genomic location) and subsequent negative (absence of random integration events) selection of a population of cells.
[0055] The overall solution is based on the integration of so-called "landing pad" (LP) sequences at predefined genomic locations selected for their ability to support high transcription and their long-term stability. Landing pads are designed together with matching donor vectors that allow for controlled integration into predefined sites and direct selection of cells in which only the desired integration has occurred. Herein, predefined genomic locations and landing pads / landing pad sequences may be used interchangeably.
[0056] The basic design of the SDI system involves (i) integration of a donor vector containing a nucleic acid sequence of interest into a pre-targeted genomic location in a eukaryotic host cell, (ii) selection of cells that have integrated a single copy of the donor vector into the pre-defined genomic location using at least one, or alternatively, two orthogonal selection steps, and (iii) optionally, the use of a combination of two types of DNA enzyme recognition sequences along with two different DNA enzymes, e.g., specific recombinases, to enable removal of undesired sequences from the donor vector at the pre-defined genomic location.
[0057] A general implementation of the method is outlined in Figure 1. The isolated eukaryotic cell contains a predefined genomic location (i) containing a nucleic acid sequence I1 containing a recognition site for a first DNA enzyme, (ii) a nucleic acid sequence E1 containing a recognition site for a second DNA enzyme, and (iii) a promoter nucleic acid sequence P1 containing a transcription initiation site. I1, E2, and P1 are formed in two symmetrical 5'-3' sequence orientations: O1 = [I1, P1 with 3'-5' orientation, E1] or O2 = [E1, P1 with 5'-3' orientation, I1]. The donor vector contains (i) a nucleic acid sequence I2 that promotes recombination with I1 in the presence of the first DNA enzyme, (ii) a promoter-less first selectable marker gene (SM1), (iii) a recognition site E2 for the second DNA enzyme, (iv) an integration cassette IC, and optionally (v) an active expression cassette for a second selectable marker gene (SM2). SM1, SM2 (when present), E2, and IC are formed in one of two symmetric clockwise orientations: O3 = [I2, IC, E2, SM2, SM1 with counterclockwise orientation] or O4 = [I2, SM1, SM2, E2, IC with clockwise orientation]. Nucleic acid sequence elements present in the predefined genomic locations and donor vectors are always formed in one of two matching orientations: (a) O1 / O3 or (b) O2 / O4.
[0058] Integration of the entire donor vector or a portion of the donor vector into a predefined genomic location of the isolated eukaryotic cell is achieved by introducing the donor vector into the cell in the presence of a first DNA enzyme, the presence of which allows recombination between nucleic acid sequence I2 of the donor vector and nucleic acid sequence I1 present at the predefined genomic location of the cell.
[0059] Integration at a predefined genomic location is located in the SM1 gene so that P1 can effect transcription of the SM1 gene and, therefore, expression of the SM1 gene product. Thus, cells that have integrated the entire donor vector or a portion of the donor vector at a predefined genomic location can be selected and isolated by using expression of SM1 as a positive selection criterion.
[0060] Optionally, undesired sequences that may potentially negatively impact the intended functionality of the isolated cells can be specifically removed from predefined genomic locations in a supplementation step leaving only the integration cassette (IC) and remaining sequences from I1, I2, E1, and E2.
[0061] Upon integration of the entire donor vector or a portion of the donor vector at a predefined genomic location, the plasmid backbone sequence (i.e., the sequence for plasmid propagation in bacteria) and the expression cassettes for SM1 and SM2 (if present) become flanked by two nucleic acid sequences, E1 and E2 (see FIG. 1 ). In the presence of the second DNA enzyme, this region of sequence flanked by E1 and E2 is excised from the predefined genomic location via the second DNA enzyme acting on E1 and E2. Cells in which the region flanked by E1 and E2 has been excised can be selected and isolated in a negative selection step based on the absence of SM1 expression (if SM2 is not present in the original donor vector) and / or the absence of SM2 expression (if SM2 is present in the original donor vector). In addition to achieving removal of undesired sequences, this supplementary selection step consistently increases the specificity of isolating cells that have integrated the entire donor vector or part of the donor vector at a predefined genomic location, since any cells that achieve SM1 activation (via a non-specific mechanism) after integration outside the predefined genomic location will not have SM1 flanked by E1 and E2 and therefore will not be selected in the negative selection step based on SM1 expression.
[0062] Using SM2 present in the donor vector, a selection step with improved functionality can be performed after the action of the second DNA enzyme. Because SM2 is provided as an active expression cassette, any copy of the donor vector integrated into an undesired genomic location will result in expression of SM2. Importantly, however, because E1 is present only at the predefined genomic location, such an integration event will not lead to an SM2 expression cassette flanked by E1 and E2. Thus, after the action of the second DNA enzyme, which leads to excision of the sequence region flanked by E1 and E2, cells that have integrated a single copy of the integration cassette (IC) at, and only at, the predefined genomic location can be selected and isolated in a negative selection step based on the absence of SM2 expression.
[0063] The integration cassette (IC) typically contains an expression cassette for a gene of interest (GOI), although application of the method is not limited thereto.
[0064] Further examples of specific implementations and general methods are generally illustrated, but not limited to, using only one of two possible symmetric orientations of key sequence elements present in predefined genomic locations and present in the donor vector.
[0065] One particular implementation of the design concept is outlined in Figure 4, focusing on the landing pad (LP1P1) and DNA donor vector. Results of experiments performed based on this implementation are also described and discussed in the experimental section of Example 2. This implementation is merely an example of one way of practicing the invention, and is not intended to be limiting.
[0066] Thus, in one implementation illustrated in Figure 4, a eukaryotic host cell line contains, at a predefined genomic location, a first recombinase recognition sequence (attP1) for the recombinase PhiC31 recombinase, a promoter in a 3' to 5' orientation, and a second recombinase recognition sequence (loxP) for the recombinase Cre recombinase.
[0067] PhiC31 recombinase is a DNA recombinase derived from the Streptomyces phage φC31. This enzyme can mediate recombination between the two nucleic acid sequences attB and attP. Cre recombinase is also a site-specific recombinase used in this system to subsequently excise the selection system and the plasmid bacterial backbone. Thus, once the initial selection has been made, Cre recombinase can be described as "cleaning" the vector backbone from non-useful sequences. Both PhiC31 recombinase and Cre recombinase are well-known enzymes used in site-specific recombination (
[10] ).
[0068] The matching DNA donor vector comprises a promoter-less first selection marker (exemplified here by RFP, red fluorescent protein) encoded in a counterclockwise orientation, a matching PhiC31 recombinase recognition sequence (attB1), an expression cassette comprising a nucleic acid sequence encoding a protein of interest, a complementary recombinase recognition sequence for Cre recombinase (loxP), a complete functional expression cassette of a second selection marker (optionally exemplified here by FC-eGFP), and a plasmid backbone (including sequences for bacterial propagation, etc.).
[0069] Cotransfection of a DNA donor vector and a vector for expression of PhiC31 containing a predefined genomic location of a landing pad (LP) sequence into eukaryotic host cells will result in integration of the donor vector at the LP via PhiC31-mediated recombination of attP1 and attB2 in a fraction of the transfected cells. Upon integration at the predefined genomic location, a promoterless selectable marker will be positioned so that it is activated by the promoter at the predefined genomic location. Activity of the first selectable marker can then be used to select for cells that have integrated at the LP (using FACS in the case of RFP). Appropriate selection should generate a pool of cells in which the majority of cells have a single copy integrated at the LP. However, a fraction of cells are expected to have additional copies integrated via off-target integration mechanisms, such as DNA repair-mediated random integration at genomic pseudo-attP sequences and PhiC31-mediated integration. To select against such events and simultaneously allow for the removal (i.e., "wiping out") of the selectable marker cassette and plasmid backbone at a predefined genomic location, a second recombinase-mediated step was designed.
[0070] Because the predefined genomic location contains a loxP sequence and the DNA donor vector also contains a strategically placed loxP sequence, an integration event at the predefined genomic location will contain both selectable markers (as well as other unwanted sequence elements, such as the plasmid backbone) flanked by two loxP sequences. In contrast, most off-target events should not lead to a loxP-flanked selectable marker (although some random integration events of the linked donor vector may lead to an adjacent second selectable marker gene, this should be very rare). By using a second transfection of a vector encoding Cre recombinase, the region flanked by loxP sequences can be excised from the genome of the corresponding cell. Cells with a single copy integrated at the predefined genomic location (lacking off-target integration) and the unwanted sequence elements to be removed can be selected via the absence of selectable marker activity (absence of eGFP activity using FACS). This is also referred to as selection by negative selection.
[0071] Some key general theoretical advantages of the general SDI system disclosed herein are: (1) It allows for a selection process that minimizes the likelihood that isolated cells differ from the desired outcome of having a single copy, i.e., gene of interest (GOI), integrated at a predefined genomic location, and only at that location. This is important in CLD operations because it reduces biological variation and therefore the amount of screening required. It also improves the likelihood that cells isolated from a CLD operation will perform well in the platform culture process. This is a critical property because optimization of expression cassette designs, which is based on transfection of a donor vector mixture containing a library of expression cassette designs, requires a one-to-one correlation between cell phenotype and a single corresponding gene cassette design.
[0072] (2) Only sequences responsible for the productivity of the cell line are retained. Cell sources are not consumed for the expression of selectable marker proteins or truncated GOI versions, as may be the case in random integration (RI)-based cell line development (CLD). The presence of sequences of bacterial origin, which potentially have a negative impact on the long-term expression stability of the GOI, can be avoided.
[0073] (3) Because the selectable marker is not part of a predefined genomic location sequence, there is flexibility in the selection of the selectable marker. The optimal selectable marker can be selected based on the application.
[0074] (4) The desired integration event activates the expression of a first selection marker, which allows for the positive selection of cells with integration at a predefined genomic location with high specificity. Using a selection marker such as a fluorescent protein or cell surface marker, this allows for a very short period between transfection and selection of positive integrants (e.g., using FACS or MACS). Results should be obtained in 2-3 days. This shortens the time required for the CLD procedure. In addition, early isolation of cells that have undergone integration at the desired location from cells that have undergone undesired integration events or have not undergone integration may have an additional advantage, as it minimizes the risk that undesired cells will outgrow the desired cells. Thus, the efficiency and performance of the method may be improved compared to methods lacking this feature.
[0075] (5) The method allows for sequential integration at the same genomic location without the construction of unwanted sequences. This can be achieved by placing the novel sequence required for the second integration event at a predefined genomic location downstream of the first GOI (or nucleic acid sequence of interest), as illustrated for PhiC31 used as the first DNA enzyme in Figure 6. This property would also enable the generation of host eukaryotic cell lines with multiple landing pads that can be individually enhanced. This allows multiple copies of a GOI to be integrated at one or several predefined genomic locations to achieve increased expression of the corresponding protein of interest (POI). Alternatively, it can be used to enable protein- and clone-specific cell line engineering through the regulated integration of cellular effector proteins that improve expression.
[0076] Preferred implementations of the method, utilizing a serine recombinase such as PhiC31 or Bxb1 as the first DNA enzyme in combination with a single matching recombinase recognition sequence pair (i.e., attP / attB) further have the potential for superior integration efficiency. Specifically, it is thus operably fused to the 5' portion of the split intron in combination with the promoter P1 or P2. PhiC31- or Bxb1-mediated recombination of their corresponding attP / attB pairs is an irreversible reaction; therefore, in theory, integration should be limited only by transfection efficiency and plasmid stability. This contrasts with Cre-based or CRISPR / Cas9-based integration, in which competing, non-productive reaction pathways may exist.
[0077] Thus, the present disclosure provides new and improved methods for the efficient and selective targeted integration of a nucleic acid sequence of interest (e.g., encoding a protein of interest) into a host cell. Isolated host cells that have selectively integrated a single copy donor vector containing the nucleic acid sequence of interest represent an excellent system for recombinant protein production that will find use in many different applications.
[0078] Thus, in a first aspect, the present disclosure relates to a method for targeted integration of a donor vector into a predefined genomic location in an isolated eukaryotic cell, said method comprising: i) a. a nucleic acid sequence I1 comprising a recognition site for a first DNA enzyme; b. a nucleic acid sequence E1 containing a recognition site for a second DNA enzyme; and c. promoter nucleic acid sequence P1; providing an isolated eukaryotic cell comprising a predefined genomic location; ii) a. nucleic acid sequence I2; b. a nucleic acid sequence of interest; c. a nucleic acid sequence E2 comprising a recognition site for the second DNA enzyme; d. a nucleic acid sequence encoding a first selectable marker; e. Optionally, an expression cassette encoding a second selectable marker; providing a donor vector; iii) contacting the donor vector with the cell in the presence of a first DNA enzyme, wherein the presence of the first DNA enzyme allows for recombination between the nucleic acid sequence I2 of the donor vector and the nucleic acid sequence I1 present at a predefined genomic location of the cell; iv) selecting cells having the donor vector integrated at a predefined genomic location by detecting expression of a first selection marker in the cells, wherein expression of the first selection marker is activated by the promoter nucleic acid sequence P1 at the predefined genomic location; and v) isolating the cells selected in the previous step Includes:
[0079] The first selectable marker may also be abbreviated and referred to herein as "SM1."
[0080] The second selectable marker may also be abbreviated and referred to herein as "SM2."
[0081] Non-limiting examples of first DNA enzymes are DNA recombinases described elsewhere herein, such as PhiC31 or Bxb1 recombinase. When used as a first DNA enzyme, a characterizing property of a recombinase is that it introduces, rather than removes, nucleic acid sequence regions into predefined genomic regions.
[0082] Non-limiting examples of second DNA enzymes are the DNA recombinases described elsewhere herein, such as PhiC31 recombinase, Bxb1 recombinase, Cre recombinase, and Dre recombinase. When used as a second DNA enzyme, a characterizing property of a recombinase is that it removes, rather than introduces, nucleic acid sequence regions from a predefined genomic region.
[0083] The nucleic acid sequence I1 comprising the recognition site for the first DNA enzyme may be an attP or attB site for PhiC31 or Bxb1 recombinase present at said predefined genomic location, or as otherwise exemplified herein, depending on which first DNA enzyme is used in the present context. As an example, it may also be a loxP site for Cre recombinase, or as otherwise exemplified herein.
[0084] Nucleic acid sequence I2 may be an attB or attP site (recognition site) for PhiC31 or Bxb1 recombinase present in the donor vector, or as otherwise exemplified herein, depending on which first DNA enzyme is used in the present context. As an example, it may also be a loxP site for Cre recombinase, or as otherwise exemplified herein.
[0085] The nucleic acid sequence E1 can be a loxP site for Cre recombinase or a roxP site for Dre recombinase. It can also be an attP or attB site for PhiC31 or Bxb1 recombinase, or as otherwise exemplified herein.
[0086] The nucleic acid sequence E2 may be a loxP site for Cre recombinase or a roxP site for Dre recombinase. It may also be an attP or attB site for PhiC31 or Bxb1 recombinase, or as otherwise exemplified herein.
[0087] If the first DNA enzyme is a PhiC31 recombinase, the second DNA enzyme is not a PhiC31 recombinase. The same applies to any other first and second DNA enzymes, i.e., the first and second DNA enzymes are never identical in the same SDI system.
[0088] Herein, the first selectable marker (SM1) of the donor vector may be linked to a gene encoding the second DNA enzyme via an IRES element or the amino acid sequence of SM1, and the second DNA enzyme is fused with a self-cleaving peptide, so that both the first selectable marker and the second DNA enzyme are activated upon integration at a predefined genomic location. This is illustrated in Figure 10. Once the donor vector is integrated into the predefined genomic location, no further introduction of the nucleic acid vector is required using the method steps, ensuring the presence of the second DNA enzyme. SM1 expression can be increased until the intracellular concentration of the second DNA enzyme reaches a value high enough to promote nuclear localization and excision of the sequence region flanked by E1 and E2. With a properly timed positive selection step, cells that have undergone integration at a predefined genomic location will contain levels of SM1 that allow for positive selection.
[0089] As used herein, the predefined genomic location may also include an expression cassette for the first DNA enzyme positioned such that upon integration of the donor vector at the predefined genomic location, the expression cassette is flanked by a recognition site for the second DNA enzyme and removed from the predefined genomic region through the action of the second DNA enzyme. This is illustrated in Figure 11. This should further simplify the method and improve the likelihood of high integration efficiency. Because the expression cassette is removed during a later step of the method, no cellular resources are consumed in the expression of the first DNA enzyme in the final isolated cells, avoiding any negative consequences of the long-term presence of the first DNA enzyme.
[0090] Thus, the first DNA enzyme may be provided by expression from a predefined genomic location or by introduction into the cell in any form that results in the transient presence of the first DNA enzyme in the cell, including introduction of the isolated protein itself, introduction of a separate expression plasmid containing an expression cassette for the first DNA enzyme, the presence of an active expression cassette for the first DNA enzyme in the donor vector, or introduction of a synthetic mRNA encoding the first DNA enzyme.
[0091] As previously described herein, all aspects of the present disclosure allow flexibility in the selection of selectable markers without making any changes to predefined genomic locations. SM1 can be selected from the group of (i) antibiotic resistance genes, (ii) metabolic enzyme genes such as GS or DHFR, (iii) fluorescent protein genes, or (iv) cell surface markers such as CD4 or CD10. SM2 can be selected from the group of (i) toxic product-producing enzymes such as TK, (ii) fluorescent protein genes, or (iii) cell surface markers such as CD4 or CD10.
[0092] Preferably, both selection markers are chosen from the group of (i) fluorescent protein genes or (ii) cell surface markers, which allow for a fast selection process via methods such as FACS or MACS.
[0093] If the selection marker is a fluorescent protein, expression of the first or second selection marker can be detected, for example, by using FACS. If the selection marker is an antibiotic resistance gene, integration can be detected by culturing cells in the presence of the corresponding antibiotic. If the cells survive in medium supplemented with the antibiotic, the donor vector has been successfully integrated.
[0094] Also provided herein is a method further comprising step vi) comprising excising a nucleic acid sequence flanked by nucleic acid sequences E1 and E2 from a predefined genomic location of a cell isolated in step v) of the method herein in the presence of a second DNA enzyme, wherein the presence of the second DNA enzyme enables recombination between nucleic acid sequences E1 and E2, and the presence of the nucleic acid sequence flanked by nucleic acid sequences E1 and E2 in the cell is indicative of stable integration of the donor vector into the predefined genomic location of the cell.
[0095] As described herein, recombinases are useful for excising nucleic acid sequences flanked by appropriate nucleic acid regions (E1 and E2) in the host cell genome. This is simply a "cleaning up" process in the host cell genome, since some portions of the nucleic acid sequence introduced into a predefined genomic location become redundant after integration and selection. Their presence can also consume cellular energy. Excision of a nucleic acid sequence refers to the ability of a second DNA enzyme to excise and remove portions of the nucleic acid sequence from the host cell genome by binding to a specific combination of nucleotides, i.e., E1 and E2. The presence of the nucleic acid sequences E1 and E2 at a predefined genomic location is, in principle, evidence that stable integration of the donor vector has occurred.
[0096] Herein, step vi) may form part of step iii) or may be performed after step iii), such as after step iv) or after step v) of the method. Step vi) may also be performed before step v).
[0097] Also provided is a method wherein the donor vector of step ii) further comprises e) an expression cassette encoding a second selection marker, and the cells isolated in step v) are additionally selected based on their non-expression of the second selection marker, expression of the second selection marker indicating that the donor vector has integrated at a location different from the predefined genomic location of the cell.
[0098] Also provided herein is a method, wherein the donor vector of step ii) further comprises e) an expression cassette encoding a second selection marker, and the cells isolated in step vii) are selected based on their non-expression of the second selection marker, wherein expression of the second selection marker indicates that the donor vector has integrated at a location different from the predefined genomic location of the cell.
[0099] The expression cassette encoding a second selectable marker is positioned on the donor vector such that upon integration at the predefined genomic location, it will be flanked by E1 and E2 sequences. However, if the donor vector integrates outside of the predefined genomic location, the expression cassette encoding a second selectable marker will not be flanked by E1 and E2 sequences.
[0100] Thus, if expression of the second selection marker (SM2) among cells in the cell population can be detected, for example, by FACS, after the action of the second DNA enzyme, this means that undesired integration events of the donor vector have occurred at other locations in the cells. Such cells can be removed, allowing for selection (by negative selection) for cells in which integration of the donor vector has occurred only at the predefined genomic location.
[0101] Also provided is a method in which step vi) is performed after step v), the method further comprising step vii) performed after step vi), which comprises isolating cells in which the nucleic acid sequence flanked by nucleic acid sequences E1 and E2 has been excised from the predefined genomic location of the cells isolated in step vi).
[0102] The second DNA enzyme may be provided as an isolated protein itself, it may be expressed from an expression cassette on a separate expression vector or plasmid, or it may be expressed from a synthetic mRNA encoding the second DNA enzyme. It may also be expressed from a donor vector once integrated into a predefined genomic location, as previously described.
[0103] Also provided herein is a method wherein the nucleic acid sequence of interest in the donor vector of step ii) comprises at least one expression cassette comprising a gene encoding a protein of interest. The protein of interest can be any type of recombinant protein that the user desires to express, such as an antibody or other therapeutic protein.
[0104] Also provided herein are methods wherein the excised nucleic acid sequence lacks at least one expression cassette containing a gene encoding a protein of interest.
[0105] Cell line development based on random integration (RI) typically results in top clones with multiple integrated copies of the target gene. It would be advantageous for the disclosed method to also have the ability to increase target gene copy number in a controlled manner to ensure that the transcription levels (mRNA copies per cell) required for competitive protein expression levels can be achieved. Integration of multiple copies at a single time can be problematic because the size of the expression plasmid poses challenges for bacterial expansion efficiency, transfection efficiency, plasmid stability in CHO cells, and integration efficiency.
[0106] Therefore, methods for sequentially integrating multiple copies of a nucleic acid of interest into a host cell genome are also provided herein. In addition to reducing the size of the expression plasmid, inserting copies in a sequential manner offers the additional benefit of potentially gradually increasing recombinant expression imposed on the cell. This, in turn, may improve the likelihood of isolating a highly productive phenotype due to the possibility of gradually adapting to new stressor environments. As already mentioned, recursive integration at predefined genomic locations may also be used to enable protein- and clone-specific cell line engineering by first introducing an expression cassette for a protein of interest, followed by the introduction of a cellular effector gene that can improve expression of the protein of interest in a subsequent integration step.
[0107] The technology presented herein, based on a key property of PhiC31 / att, which holds the key to simple, controlled, sequential integration of multiple copies, is the presence of an orthogonal attP / attB pair. The orthogonal attP / attB pair differs from the natural sequence only in the central nucleotide pair. Examples of repeated integration at the same genomic location using the orthogonal recognition site of a first DNA enzyme are shown in the experimental section of Example 3 and in Figures 6 and 8. Of course, the same approach as the PhiC31 / att-based technology described herein can be used for other DNA enzymes, such as Cre, Dre, Flp, or CRISPR / Cas9, for which an orthogonal DNA enzyme recognition pair / sequence exists.
[0108] Therefore, also provided herein is a method comprising the steps defined elsewhere herein to result in the sequential integration of multiple nucleic acid sequences of interest, wherein the donor vector of step ii) of the method provided herein comprises: f. a nucleic acid sequence I3 comprising a recognition site for a first DNA enzyme; and g. promoter nucleic acid sequence P2 Further includes:
[0109] The presence of an additional recognition site I3 in the donor vector results in the recurrent targeted integration of two or more donor vectors at the same predefined genetic location, as described above. When a first recombination event occurs within nucleic acid sequence pair I1 / I2 (e.g., attB1 / attP1), thereby generating a hybrid sequence (attR1), an additional recognition site I3 (e.g., attP2) is maintained, which can facilitate a next round of integration using a second donor vector containing an additional recognition site I4 (e.g., attB2). Figure 6 illustrates how such a sequential integration approach can be performed. The method involves rounds of excision of nucleic acid sequence components from the predefined genomic location where the first or additional donor vector integrates to make way for the integration of additional nucleic acid sequences of interest.
[0110] Thus, provided herein is a method for sequential targeted integration of n additional donor vectors into predefined genomic locations in a eukaryotic cell, wherein a first donor vector further comprises, in addition to the components already described herein, f. a nucleic acid sequence I3 comprising a recognition site for a first DNA enzyme; and g. a promoter nucleic acid sequence P2.
[0111] A method for sequential targeted integration of n additional donor vectors into predefined genomic locations of a eukaryotic cell, comprising performing at least steps i) to iv), and optionally v), vi) and / or vii), in any suitable order, for sequential targeted integration of n additional donor vectors into predefined genomic locations of an isolated eukaryotic cell; n is an integer greater than or equal to 1, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, or any other number; The method is: (A) I. Providing the cells isolated in step v) or the cells isolated in step vii); II. A. a nucleic acid sequence I4 capable of recombining in the presence of said first DNA enzyme with the corresponding nucleic acid sequence I3 present in the cell provided in the previous step; B. a nucleic acid sequence of interest; C. Nucleic acid sequence E2; D. a nucleic acid sequence encoding a first selectable marker; and E. Optionally, an expression cassette encoding a second selectable marker; F. Optionally, a nucleic acid sequence I1 containing a recognition site for a first DNA enzyme and a promoter nucleic acid sequence P1 providing a donor vector comprising: III. introducing the donor vector of step II) into a cell in the presence of a first DNA enzyme, wherein the presence of the first DNA enzyme allows recombination between the nucleic acid sequence I4 of the donor vector and the nucleic acid sequence I3 present at a predefined genomic location of the cell; IV. selecting cells having the donor vector integrated at a predefined genomic location by detecting expression of a first selection marker in the cells, wherein expression of the first selection marker is activated by the promoter nucleic acid sequence P2 at the predefined genomic location of the cells; V. isolating the cells selected in the preceding step; VI. excising a nucleic acid sequence flanked by nucleic acid sequences E1 and E2 from a predefined genomic location of the cell isolated in step V in the presence of a second DNA enzyme, wherein the presence of the second DNA enzyme permits recombination between the nucleic acid sequences E1 and E2, and the presence in the cell of the nucleic acid sequence flanked by nucleic acid sequences E1 and E2 is indicative of stable integration of the donor vector into the predefined genomic location of the cell; VII. Isolating cells, wherein the nucleic acid sequence flanked by sequences E1 and E2 has been excised from the predefined genomic location of the cells isolated in step V. integrating a first additional donor vector into a predefined genomic location of the cell; (B) where n is greater than the number of additional donor vectors integrated at the predefined genomic location of the cells isolated in the preceding step; I. providing a cell isolated in the preceding step, which contains the nucleic acid sequence I1 integrated at a predefined genomic location; II. Steps ii), iii) to iv), v), vi), and vii) integrating an additional donor vector comprising the compound into a predefined genomic location in the cell; (C) where n is greater than the number of additional donor vectors integrated at the predefined genomic location of the cells isolated in the preceding step; I. providing a cell containing the nucleic acid sequence I3 integrated at a predefined genomic location, obtained by carrying out the aforementioned step (B); II. Repeating steps (A) II through (A) IX and step (B) until the cell has n additional donor vectors integrated into predefined genomic locations. integrating an additional donor vector into a predefined genomic location of the cell, comprising: Includes:
[0112] As previously described herein, recognition site I4 in the donor vector results in recombination with recognition site I3, which is already present in a predefined genomic location in the host cell (i.e., incorporated in a previous round of integration). This allows second and further copies / multiple copies of the nucleic acid of interest to be introduced into the predefined genomic location by subsequent introduction of recognition site variants via the donor vector. The particular recognition site variant (pair) (i.e., the first DNA enzyme recognition site) used for recombination between the introduced donor vector and a nucleic acid sequence present in a predefined genomic region can be reused through rounds of recombination events, as long as a different pair of recognition sites is present during each round of integration. The same is true for selectable markers.
[0113] Specifically, there is provided an iterative method for integrating any desired number of donor vector copies into a predefined genomic location of said eukaryotic cell based solely on two orthogonal pairs of recognition sequences for a first recombinase enzyme and two variants of said first selection marker, (i) the first recombinase enzyme is selected from the group of serine recombinases, such as PhiC31 or Bxb1 or mutated variants thereof; (ii) the two selectable marker variants are selected from the group of (a) fluorescent proteins or (b) heterologous cell surface markers; (iii) The donor vector used in the odd-numbered integration step is (a) a first version of the first selectable marker; (b) a first recombinase recognition sequence from a first pair of recognition sequences; (c) a first recombinase recognition sequence derived from the recognition sequence of a second orthogonal pair; Includes; (iv) The donor vector used in the even-numbered integration step is (a) a second version of the first selectable marker; (b) a second recombinase recognition sequence derived from the recognition sequence of the second orthogonal pair; (c) a second recombinase recognition sequence derived from the first pair of recognition sequences; Includes; (v) integration at odd integration steps is facilitated by recombination between recombinase recognition sequences from said first pair of recombinase recognition sequences; (vi) integration in even-numbered integration steps is facilitated by recombination between recombinase recognition sequences from the orthogonal second pair of recombinase recognition sequences; (vii) prior to the interaction of the odd number of integration steps, the first version of the first selectable marker is excised from the predefined genomic location by the presence of the second recombinase enzyme acting at recombinase recognition sequences E1 and E2; (viii) prior to the interaction of even number of integration steps, the second version of the first selectable marker is excised from the predefined genomic location by the presence of a second recombinase enzyme that acts at recombinase recognition sequences E1 and E2; (ix) the second recombinase enzyme is selected from the group of tyrosine recombinases, such as Cre, Dre, or Flp; (x)E1=E2.
[0114] - the cell of step i) comprises n predefined genomic locations, each of which is a nucleic acid sequence I, I11 to I1 n Including I11 to I1 n are different from each other; - step ii) comprises providing 1 to n donor vectors, each of the 1 to n donor vectors comprising the nucleic acid sequences I21 to I22, respectively; n In the presence of the first DNA enzyme, the corresponding I11 to I1 n Each of the 1 to n donor vectors is capable of recombining nucleic acid sequences and contains a first selection marker, SM1 to SM2. n Including SM1 to SM n but differ from each other; - step iv) comprises introducing 1 to n donor vectors into the cell; - step v) is the step of selecting different first selection markers SM1 to SM2 in the cells; n selecting cells having each of the 1 to n donor vectors integrated at its corresponding predefined genomic location by detecting each of the 1 to n donor vectors; Further provided is a method wherein n is an integer greater than or equal to 2.
[0115] The donor vector of step ii) is f. a nucleic acid sequence I3 comprising a recognition site for a first DNA enzyme; and g. promoter nucleic acid sequence P2 Further comprising: Also provided is a method wherein the sequence to be excised comprises an expression cassette containing a gene encoding a protein of interest.
[0116] The fact that the excised sequence includes an expression cassette containing a gene encoding a protein of interest means that there is a reversible integration whereby a second round of recombination can introduce a new and "first" nucleic acid sequence of interest encoding the protein of interest at a predefined genomic location. Thus, there are no multiple copies of the same gene of interest at a predefined genomic location, which is the purpose of sequential integration as previously mentioned herein. This can be exploited for the reuse of high performance clones for the expression of another protein of interest.
[0117] Figure 8 and Example 4 illustrate the generation of SDI cell pools using two sequential selection steps: adding a second DNA enzyme after integration to remove nucleic acid sequences that no longer serve a purpose in the cells. In this example, an antibiotic resistance gene was used as the first selection marker (SM1). A second round of selection was performed using Cre recombinase to excise the nucleic acid sequence flanked by loxP nucleic acid regions at each end. The presence of random integration events was detected by double-positive GFP / RFP (green / red fluorescent protein) signals using FACS. Cells were sorted based on positive / negative GFP / RFP signals. This additional step results in the elimination of cells that can integrate one donor vector at a predetermined genomic location but can randomly integrate a second or additional donor vector at random, non-target locations in the host cell genome.
[0118] As previously mentioned herein, the first DNA enzyme may be a recombinase. The first DNA enzyme may also be a mixture of different DNA enzymes, such as recombinases, as long as none of the DNA enzymes in the first DNA enzyme is the same as the second DNA enzyme.
[0119] As used herein, there may be more than one recognition site for the first DNA enzyme present in the donor vector, such as two or more recognition sites. This means that there will also be more than one recognition site for the first DNA enzyme present at a predefined genomic location, such as two or more recognition sites. An example of such a system utilizing a recombinase is shown in FIG. 13, which illustrates recombinase-mediated cassette exchange (RMCE) to catalyze integration at a predefined genomic location. The modified variants thereof described in FIGS. 1-2, 9-11, and 15 are also encompassed by the present disclosure.
[0120] Thus, in the example in Figure 13, the predefined genomic location includes, in 5' to 3' sequence order, (i) a first recognition site for a first recombinase enzyme (I1a); (ii) a second recognition site for the first recombinase enzyme (I1b), (iii) a promoter P1 with a 3' to -5' orientation, and (v) a recognition site for a second recombinase enzyme.
[0121] In this example, the donor vector comprises, in 5' to 3' sequence order: (i) a third recognition site for the first recombinase enzyme (I2a); (ii) an integration cassette (IC), exemplified here by an expression cassette for a gene of interest (GOI); (iii) a recognition site E2 for the second recombinase enzyme; (iv) an expression cassette for a second selection marker (SM2); (v) a gene for a first selection marker (SM1) encoded in a 3' to 5' orientation; and (vi) a fourth recognition site for the first recombinase enzyme (I2b).
[0122] Introduction of the donor vector and the first recombinase into a population of cells results in (a) integration of the integration cassette in the donor vector, i.e., the sequence region flanked by the third and fourth recombinase recognition sites, at a predefined genomic location for a fraction of the cells (see Figure 13b, panel (ii)), and (b) off-target genomic integration (outside the predefined genomic location) of the donor vector for a fraction of the cells (see Figure 13b, panel (iii)).
[0123] Integration at the predefined genomic location results in the formation of an active expression cassette for SM1 (see Figure 13b, panel (ii)). Furthermore, after integration at the predefined genomic location, both SM1 and SM2 are flanked by two recognition sites for the second recombinase enzyme.
[0124] Integration via an off-target event (see Figure 13b, panel (iii)) typically does not lead to activation of SM1, but rather to the integration of an active SM2 that is not flanked by the two recognition sites for the second recombinase enzyme.
[0125] Cells that have undergone integration at a predefined genomic location differ from cells that have no integration event (see Figure 13b, panel (i)) and cells that have undergone only off-target integration events via the activity of SM1. Thus, the activity of SM1 can be used to select for cells that have undergone integration at the LP.
[0126] In order to eliminate cells that have undergone off-target integration events in addition to integration at a predefined genomic location, the recombinase activity of the second recombinase is introduced into cells selected for SM1 activity. For integration at the LP, this results in the excision of both SM1 and SM2, thus resulting in their corresponding activities. For off-target integration events, this reaction cannot occur, and SM2 activity remains. As a result, cells that have undergone only the desired targeted integration event at a predefined genomic location can be selected from cells that have undergone multiple integration events through the absence of SM2 activity.
[0127] The predefined genomic location for the final selected cells (see Figure 13c) does not contain the expression cassette for SM2, the activated expression cassette for SM1, or any remaining sequences from the donor vector, except for sequences generated via recombination of E1 and E2 (E).
[0128] The first recombinase enzyme can be selected from the group of (i) serine recombinases or (ii) tyrosine recombinases.
[0129] The first to fourth recombinase recognition sites are (a) I1a=I1b and I2a=I2b, using one matching recognition site for a serine recombinase, such as PhiC31 [I1a=I1b=attP or attB and I2a=I2b=attB or attP] or Bxb1 [I1a=I1b=Bxb1 attP or Bxb1 attB and I2a=I2b=Bxb1 attB or Bxb1 attP]; (b) using mutated recognition pairs of serine recombinases such as PhiC31 (I1a and I1b = different attP or attB variants; I2a and I2b = different attB or attP variants), to select two different matching recognition site pairs; and (c) I1a=I2a and I1b=I2b for tyrosine recombinase are the present mutated recognition site variant pairs, such as Cre (LoxP1 and LoxP2 selected from available mutated loxP pairs), Dre (rox1 and rox2 selected from available mutated rox pairs), or FLP (FRT1 and FRT2 selected from available mutated FRT pairs). can be selected by:
[0130] The second recombinase enzyme is different from the first recombinase enzyme and can be selected from the group of (i) serine recombinases or (ii) tyrosine recombinases.
[0131] The recognition sites E1 and E2 for the second recombinase enzyme may be identical in sequence, as exemplified by (i) E1=E2=loxP or a mutated variant thereof for use with Cre recombinase, (ii) E1=E2=rox or a mutated variant thereof for use with Dre recombinase, or (iii) E1=E2=FRT or a mutated variant thereof for use with FLP (flippase) recombinase.
[0132] The recognition sites E1 and E2 of the second recombinase enzyme may have different sequences, as exemplified by (i) E1=attP and E2=attB or a mutated variant thereof for use with PhiC31 recombinase, or (ii) E1=Bxb1 attP and E2=Bxb1 attB or a mutated variant thereof for use with Bxb1 recombinase.
[0133] Additional recognition sites for the first DNA enzyme may be referred to herein as variants of I1, ie, I1a and I1b, and variants of I2, ie, I2a and I2b, etc.
[0134] therefore, (a) I1 contains two recombinase recognition site variants I1a and I1b; and (b) I2 comprises two recombinase recognition site variants I2a and I2b; and (c) I1a is capable of recombining with I2a and I1b is capable of recombining with I2b in the presence of the first DNA enzyme; Methods are also provided herein.
[0135] Sometimes, I1a is identical to I2a and I1b is identical to I2b. I1a, I1b, I2a and I2b may each be selected from loxP, rox or FRT or variants thereof, and the first DNA enzyme may each be selected from the group consisting of Cre recombinase, Dre recombinase and FLP recombinase [3].
[0136] In this specification, (a) I1 contains a single recombinase recognition site; and (b) I2 contains a single recombinase recognition site; and (c) I1 and I2 are capable of recombining in the presence of said first DNA enzyme. A method is also provided.
[0137] The recombinase recognition site formed by I1 may also differ in sequence from the recombinase recognition site formed by I2. The recombinase recognition site provided herein may be selected from attB, attP, Bxb1 attP, Bxb1 attB, or variants thereof. The recombinase may be PhiC31 or Bxb1 recombinase or a mutant thereof.
[0138] Any variant or mutant of a recognition site / DNA enzyme may be a functionally equivalent variant or mutant thereof, and one skilled in the art may construct and produce such functionally equivalent variants or mutants.
[0139] An example of using a single recombinase recognition site pair to catalyze integration at a predefined genomic location is shown in Figure 14. The modified variants described in Figures 1, 9-11, and 15 are also encompassed by the present disclosure.
[0140] In this example, the predefined genomic location includes, in 5' to 3' sequence order, (i) a first recognition site for a first recombinase enzyme (I1); (ii) a promoter P1 with a 3' to -5' orientation, and (iii) a recognition site for a second recombinase enzyme.
[0141] In this example, the donor vector contains, in 5' to 3' sequence order: (i) a second recognition site for the first recombinase enzyme (I2); (ii) an integration cassette (IC), exemplified here by an expression cassette for a gene of interest (GOI); (iii) a second recognition site E2 for the second recombinase enzyme; (iv) an expression cassette for a second selection marker (SM2); and (v) a gene for a first selection marker (SM1) encoded in a 3' to 5' direction.
[0142] Introduction of the donor vector and first recombinase into a population of LP cells results in (a) integration of the donor vector at the LP for a fraction of the LP cells (see Figure 14b, panel (ii)) and (b) off-target genomic integration (outside the predefined genomic location) of the donor vector for a fraction of the cells (see Figure 14b, panel (iii)).
[0143] Integration at the predefined genomic location results in the formation of an active expression cassette for SM1 (see Figure 14b, panel (ii)). Furthermore, after integration at the predefined genomic location, both SM1 and SM2 are flanked by two recognition sites, E1 and E2, for the second recombinase enzyme.
[0144] Integration via an off-target event (see Figure 14b, panel (iii)) typically does not lead to activation of SM1 but incorporates an active SM2 that is not flanked by recognition sites for the second recombinase enzyme.
[0145] Cells that have undergone integration at a predefined genomic location are distinct from cells that have no integration event (see Figure 14b, panel (i)) and cells that have undergone only off-target integration events via the activity of SM1. Thus, the activity of SM1 can be used to select for cells that have undergone integration at a predefined genomic location.
[0146] In order to eliminate cells that have undergone off-target integration events in addition to integration at a predefined genomic location, the recombinase activity of the second recombinase is introduced into cells selected for SM1 activity. Due to integration at a predefined genomic location, this results in the excision of both SM1 and SM2, thus resulting in their corresponding activities. Due to off-target integration events, this reaction cannot occur, and SM2 activity remains. As a result, cells that have undergone only the desired targeted integration events in LP can be selected from LP cells that have undergone multiple integration events through the absence of SM2 activity.
[0147] The predefined genomic location for the final selected cells (see Figure 14c) does not contain the expression cassette for SM2, the activated expression cassette for SM1, or any remaining sequences from the donor vector, except for sequences generated via recombination of I1 and I2 (I12) and E1 and E2 (E).
[0148] The first recombinase enzyme can be selected from the group of serine recombinases, such as PhiC31 and Bxb1. The first and second recombinase recognition sites (I1 and I2) can be selected with matching recognition sites for the selected recombinase according to: (a) I1 = attP variant and I2 = attB variant or (b) I1 = attB variant and I2 = attP variant.
[0149] The second recombinase enzyme is different from the first recombinase enzyme and can be selected from the group of (i) serine recombinases or (ii) tyrosine recombinases.
[0150] The recognition sites E1 and E2 for the second recombinase enzyme may be identical in sequence, as exemplified by (i) E1=E2=loxP or a mutated variant thereof for use with Cre recombinase, (ii) E1=E2=rox or a mutated variant thereof for use with Dre recombinase, or (iii) E1=E2=FRT or a mutated variant thereof for use with FLP recombinase.
[0151] The recognition sites E1 and E2 of the second recombinase enzyme may have different sequences, as exemplified by (i) E1=attP and E2=attB or mutated variants thereof for use with PhiC31 recombinase, or (ii) E1=Bxb1 attP and E2=Bxb1 attB.
[0152] Also provided is the method illustrated in Figure 15, which illustrates a single recombinase recognition site pair for catalyzing integration at a predefined genomic location, where a promoter P1 present at the predefined genomic location is operably fused to the 5' portion of a split intron. Methods previously described but modified according to Figure 15 (i.e., using a split intron design) are also encompassed by the present disclosure.
[0153] In this example, the predefined genomic location further comprises a 5' portion of an intron having a 3' to 5' orientation and a functional sequence region F1 having a 3' to 5' orientation between the first recognition site of the first recombinase enzyme and the promoter P1 having a 3' to 5' orientation.
[0154] In this example, the donor vector further comprises a sequence region located between the first selection marker SM1 and the second recognition site for the first recombinase enzyme, the sequence region comprising, in 5' to 3' sequence order, (a) a functional sequence region F3 having a 3' to 5' orientation and (b) the 3' portion of an intron having a 3' to 5' orientation, further comprising a functional sequence region F2 downstream of a splice acceptor site sequence.
[0155] Upon integration at a predefined genomic location (see Figure 15b, top panel), a complete expression cassette, including a functional intron, for the first selectable marker SM1 is formed, and thus expression of SM1 is activated.
[0156] After an off-target integration event (see Figure 15b, bottom panel), a truncated version of the SM1 expression cassette is integrated.
[0157] By chance, transcription of the truncated SM1 cassette can occur, which can result from (a) promoter rescue, where the donor vector is integrated in such a way that the truncated SM1 cassette is located in frame with the native promoter present in the cell genome, or (b) cleavage of the donor vector and concatemer formation, such that the promoter present in the donor vector is reoriented in frame with the truncated SM1 cassette, followed by integration of the resulting concatemer.
[0158] Such random events can reduce the specificity of SM1, which is based on the selection of cells that integrate the donor vector at a predefined genomic location. Improved specificity can be achieved by using specific combinations of functional sequence regions F1-F3 (see Figure 15b).
[0159] In the first design of F1-F3, (a) SM1 (when present in the donor vector) lacks an ATG start codon and is directly fused to the 3' intron, and (b) F1 consists of a transcription start site (TSS), the first 5'-UTR region, a Kozak / translation start site, and an ATG start codon, all in a 3'-to-5' orientation (from 3' to 5'). After an off-target integration event, this means that any integrated SM1 gene lacks a start codon and therefore does not result in expression of a functional SM1 protein. However, upon integration at a predefined genomic location, a functional expression cassette is formed. Upon splicing of the intron, the ATG start codon is fused directly to SM1, thereby leading to proper expression of the SM1 protein.
[0160] In the second design of F1-F3, (a) SM1 contains an ATG start codon, (b) F3 consists of a second 5'-UTR region and a Kozak / translation initiation site (KOS) with a 3'-5' orientation (from 3' to 5'), (c) F2 contains at least one short upstream open reading frame (uORF) with a 3'-5' orientation, and (d) F1 consists of a transcription start site (TSS) and a first 5'-UTR region. After an off-target integration event, the truncated SM1 cassette will typically retain one or more uORFs. These uORFs reduce initiation at the intended SM1 start codon, thereby improving discrimination between off-target integration-based SM1 activation and SM1 activation based on integration at a predefined genomic location. Preferably, a series of multiple uORFs is used, positioned at the shortest distance from the SM1 start codon (immediately downstream of the intron splice branch site).
[0161] The use of a split intron design also improves activated SM1 expression, because the optimal 5'-UTR sequence can be used for SM1. In designs lacking a split intron, the sequence generated through recombination of I1 and I2 (see Figure 15) would be encompassed by the SM1 5'-UTR. This leads to an extended 5'-UTR with a potentially suboptimal sequence composition, which can reduce the obtainable expression level of SM1 (affecting the specificity of the SM1-based positive selection step). Through the use of the split intron described herein, the I1 / I2 recombination product is incorporated into the fully formed intron upon integration at a predefined genomic location (see Figure 15). Upon production of mature SM1 mRNA by the cell, the intron is spliced out, and the corresponding SM1 5'-UTR is fully defined by F1 and F2. Thus, the SM1 5'-UTR can be designed with complete control to optimize SM1 expression for the intended purpose. Variations in the design of F2–F3 further confers flexibility to the expression level of SM1 upon integration at the LP. Expanding the length of the 5′-UTR region of F3 reduces SM1 expression, and adding a transcriptional enhancer element in F2 can increase SM1 expression above that achieved with the optimal 5′-UTR alone.
[0162] Finally, the use of a split intron design can improve the recombination efficiency between I1 and I2 at a predefined genomic location, as shown in the experimental section of Example 1. One possible explanation for the observed improved integration efficiency is that the 5' portion of the split intron at the predefined genomic location serves as a critical spacer that can avoid / reduce steric interference between the RNA polymerase initiation complex and the copy of the first DNA enzyme (e.g., PhiC31) that binds around the start transcription site, and accomplishes this function through the binding and manipulation of I1. To increase integration efficiency, the 5' portion of the split intron can be designed to have a length of at least 50 bp, at least 100 bp, or at least 300 bp.
[0163] Figure 12 illustrates a method in which a gene editing enzyme is used to catalyze the integration of a donor vector at a predefined genomic location (wherein the first DNA enzyme is a gene editing enzyme). Modifications thereof described in Figures 1, 9-11, and 15 are also encompassed by the present disclosure.
[0164] In this example, the predefined genomic location includes, in 5' to 3' sequence order, (i) a left homology arm (LHA), (ii) a recognition / cleavage site (CS) for the gene editing enzyme, (iii) a right homology arm (RHA) with a 3' to 5' orientation (i.e., with a splice donor site at the end closest to the promoter) that also functions as the 5' portion of the intron, (iv) a promoter P1 with a 3' to 5' orientation, and (v) a recognition site E1 for the second DNA enzyme.
[0165] In this example, the donor vector comprises, in 5' to 3' sequence order, (i) the left homologous arm (LHA), (ii) an integration cassette (IC), exemplified here by an expression cassette for a gene of interest (GOI), (iii) a recognition site E2 for the second DNA enzyme, (iv) an expression cassette for a second selectable marker (SM2), (v) a gene for a first selectable marker (SM1) encoded in a 3' to 5' orientation, (vi) a 3'-portion of an intron having a 3' to 5' orientation (i.e., with a splice branch site and a splice acceptor site at the end proximal to SM1), and (vii) the right homologous arm (RHA), which also functions as a 5'-portion of an intron having a 3' to 5' orientation.
[0166] Introduction of a gene editing enzyme with cleavage specificity for the donor vector and the CS into a population of eukaryotic cells results in (a) a double-strand break at the CS at a predefined genomic location for a fraction of the eukaryotic cells, (b) integration of the donor vector region flanked by the LHA and RHA by homology-directed DNA repair for a fraction of the eukaryotic cells with a double-strand break at the CS, and (c) off-target genomic integration of the donor vector (outside the predefined genomic region) for a fraction of the LP cells.
[0167] Integration at a predefined genomic location results in the formation of an active expression cassette for SM1 (see Figure 12b, panel (ii)). The integration event also generates a complete functional intron between promoter P1 and SM1, so that the mature mRNA of SM1 does not contain RHA. Furthermore, after integration at LP, both SM1 and SM2 are flanked by two recognition sites for the second DNA enzyme.
[0168] Off-target integration (see Figure 12b, panel (iii)) typically does not lead to activation of SM1 but incorporates an active SM2 in which the two recognition sites for the second DNA enzyme are not adjacent.
[0169] Eukaryotic cells that have undergone integration at a predefined genomic location are distinct from cells that have no integration events (see Figure 12b, panel (i)) and cells that have undergone only off-target integration events via the activity of SM1. Thus, the activity of SM1 can be used to select cells that have undergone integration at a predefined genomic location.
[0170] To eliminate cells undergoing off-target integration events in addition to integration at a predefined genomic location, a recombinase activity (a second DNA enzyme) capable of recombining E1 and E2 is introduced into cells selected for SM1 activity. Due to integration at the LP, this results in the excision of both SM1 and SM2, and thus their corresponding activities. Due to off-target integration events, this reaction cannot occur, and SM2 activity remains. As a result, cells undergoing only the desired targeted integration event at a predefined genomic location can be selected from cells undergoing multiple integration events through the absence of SM2 activity.
[0171] The predefined genomic location for the final selected cells (see Figure 12c) does not contain the expression cassette for SM2, the activated expression cassette for SM1, or any remaining sequences from the donor vector, except for sequences generated via recombination of E1 and E2 (E).
[0172] The gene editing enzyme may be selected from the group of, but is not limited to, (i) Zn-finger nucleases (ZFNs); homing endonucleases such as meganucleases; (iii) TALENs; or (iv) DNA- or RNA-guided nucleases such as CRISPR / Cas9.
[0173] The second DNA enzyme has recombinase activity and may be selected from the group of (i) serine recombinases or (ii) tyrosine recombinases.
[0174] E1 and E2 may be identical in sequence, as exemplified by (i) E1=E2=loxP or a mutated variant thereof for use with Cre recombinase, (ii) E1=E2=rox or a mutated variant thereof for use with Dre recombinase, or (iii) E1=E2=FRT or a mutated variant thereof for use with FLP recombinase.
[0175] E1 and E2 can have different sequences, as exemplified by (i) E1=attP and E2=attB or (ii) E1=Bxb1 attP and E2=Bxb1 attB for use with PhiC31 recombinase.
[0176] Thus, provided herein is a method in which the first DNA enzyme is a gene-editing enzyme, such as a gene-editing nuclease, whereby (a) I1 comprises a cleavage site of the gene-editing nuclease and two sequence regions, LHA1 and RHA1; (b) I2 comprises two sequence regions, LHA2 and RHA2, that are homologous to LHA1 and LHA2; and (c) I1 and I2 are capable of recombining in the presence of the first DNA enzyme.
[0177] As previously mentioned, methods are provided wherein the gene editing enzyme is selected from the group consisting of, but not limited to, (i) zinc finger nucleases (ZFNs); (ii) homing endonucleases, such as meganucleases; (iii) TALENS, and (iv) DNA- or RNA-guided nucleases, such as CRISPR / Cas9.
[0178] Nucleic acid sequences E1 and E2 may each be the same recombinase recognition site, such as loxP, rox, or FRT, or a variant thereof, with the proviso that E1 and E2 are different from I1 and I2.
[0179] The second DNA enzyme may be selected from the group consisting of Cre recombinase, Dre recombinase, and FLP recombinase, with the proviso that the first DNA enzyme is not Cre recombinase, Dre recombinase, or FLP recombinase.
[0180] In the methods provided herein, promoter nucleic acid sequences P1 and / or P2, when integrated at the predefined genomic location, may be operably fused to the 5' portion of a split intron. This is illustrated in Figure 15, previously discussed herein. Introduction of a split intron between promoters P1 (or P2) and I1 (or a variant thereof) at the predefined genomic location provides a "spacer" that minimizes steric hindrance that may arise from polymerase blockage to the promoter. The presence of this spacer results in improved expression of the first selectable marker (SM1), as shown in the experimental section of Example 1.
[0181] Thus, a method is provided herein in which the predefined genomic location further comprises a functional sequence region F1 having a 3' to 5' orientation between the 5' portion of an intron having a 3' to 5' orientation and the first recognition site for the first recombinase enzyme and the promoter P1 having a 3' to 5' orientation, and the donor vector further comprises a sequence region located between the first selection marker SM1 having a 3' to 5' orientation and the second recognition site for the first recombinase enzyme, the sequence region comprising, in 5' to 3' sequence order, (a) a functional sequence region F3 having a 3' to 5' orientation and (b) a 3' portion of an intron having a 3' to 5' orientation, further comprising a functional sequence region F2 downstream of a splice acceptor site.
[0182] As previously discussed herein, the excised nucleic acid sequence may be: (a) a nucleic acid sequence encoding a first selectable marker; (b) a promoter nucleic acid sequence P1 or P2; and / or (c) an expression cassette encoding a second selectable marker.
[0183] The above-described design of the excised nucleic acid sequence results in the selection of cells in which the donor vector has randomly integrated at a location other than the predefined genomic location based on the expression of a second selection marker (SM2). This means that, using such a design, expression of SM2 can be positive only for cells that integrate the donor vector outside the predefined genomic location after the action of the second DNA enzyme. Therefore, a second round of selection can be performed using a negative selection step based on the expression of SM2 to catch cells in which the donor vector has integrated outside the predefined genomic location. Removal of the expression cassette encoding the second selection marker is also an improvement to the method, as it conserves energy for the cells, which can instead be used to produce the protein of interest.
[0184] The first selection marker may be selected from the group of (i) fluorescent proteins and (ii) heterologous cell surface markers, in addition to those described elsewhere herein. The use of fluorescent proteins or cell surface markers as selection markers offers particular advantages because selection can be performed using rapid and direct isolation methods (i.e., FACS or MACS-based) as soon as the concentration of the first selection marker is increased above a certain limit (allowing detection of fluorescence above background in FACS and allowing efficient binding to magnetic beads in MACS). This is in contrast to selection markers based on metabolic enzymes or antibiotic resistance genes, which require lengthy and indirect isolation strategies, and are based on cells with an active selection marker, which grow too slowly while cells lacking the active selection marker grow too slowly.
[0185] The first DNA enzyme may be provided in the form of a plasmid, mRNA or purified protein, and optionally, the first DNA enzyme may be encoded by and expressed from the donor vector. The first DNA enzyme may also be expressed from an expression cassette encoding the first DNA enzyme present at a predefined genomic location in the cell of step i) of the methods disclosed herein.
[0186] As already mentioned herein, the donor vector of step ii) may further comprise an expression cassette encoding a second DNA enzyme, the expression of which is activated when the donor vector is integrated into the predefined genomic location of the cell of step i) in the method disclosed herein.
[0187] The second DNA enzyme may also be provided in the form of a plasmid, mRNA or purified protein.
[0188] The eukaryotic cells for use in the methods provided herein may be selected from the group consisting of yeast cells, filamentous fungal cells, plant cells, insect cells, or mammalian cells. Mammalian cells may be, but are not limited to, human, monkey, rodent, or mouse cells. The eukaryotic cells are the isolated eukaryotic cells previously described herein. An isolated cell is a cell that has been isolated or removed from its natural environment.
[0189] Eukaryotic cells for use in the methods presented herein may be specifically selected based on their suitability for the production of recombinant proteins in bioreactors. Suitable cells can be selected from the group of CHO or HEK cell lines.
[0190] Eukaryotic cells for use in the methods presented herein may be specifically selected based on their similarity to cell types present in mammalian species, such as humans.
[0191] Eukaryotic cells for use in the methods presented herein may be selected from a group of cell lines capable of growing in suspension culture.
[0192] In another aspect, an isolated eukaryotic cell obtainable by the method described herein is also provided. The isolated eukaryotic cell obtainable by the method disclosed herein contains one or more nucleic acid sequences of interest integrated at a predefined genomic location. Preferably, the isolated cell obtained does not contain any donor vector integrated at a location other than the predefined genomic location in the host cell genome.
[0193] In yet another aspect, there is also provided the use of an isolated eukaryotic cell obtainable by the methods described herein for producing a recombinant protein.
[0194] Thus, in another aspect, there is also provided a method of producing a recombinant protein, said method comprising: i) obtaining an isolated eukaryotic cell comprising one or more nucleic acid sequences of interest integrated into a predefined genomic location by carrying out the method disclosed herein, wherein at least one nucleic acid sequence of interest comprises at least one expression cassette comprising a gene encoding a protein of interest; ii) producing in the cells of step i) the protein encoded by the gene of interest; and iii) isolating the protein of step ii). Includes:
[0195] In a further aspect, a. a nucleic acid sequence comprising a recognition site for a first DNA enzyme; b. a nucleic acid sequence of interest; c. a nucleic acid sequence containing a recognition site for a second DNA enzyme A donor vector comprising:
[0196] In yet a further aspect, an isolated eukaryotic cell is provided comprising a predefined genomic location, the predefined location comprising: a. a nucleic acid sequence I1 containing a recognition site for a first DNA enzyme; b. a nucleic acid sequence E1 containing a recognition site for a second DNA enzyme; and c. promoter nucleic acid sequence P1; d. a nucleic acid sequence encoding a first selectable marker; e. Optionally, an expression cassette encoding a second selectable marker Includes:
[0197] In yet a further aspect, a recombinant expression system is provided for targeted integration of a nucleic acid sequence of interest into a host cell, said expression system comprising: a. a nucleic acid sequence comprising a recognition site for a first DNA enzyme; b. a nucleic acid sequence of interest; c. a nucleic acid sequence containing a recognition site for a second DNA enzyme a donor vector comprising: a. a nucleic acid sequence I1 containing a recognition site for a first DNA enzyme; b. a nucleic acid sequence E1 containing a recognition site for a second DNA enzyme; and c. promoter nucleic acid sequence P1; d. a nucleic acid sequence encoding a first selectable marker; e. Optionally, an expression cassette encoding a second selectable marker an isolated eukaryotic cell comprising a predefined genomic location, Includes:
[0198] The present disclosure will now be illustrated by the following Examples section, without being limited to the examples provided therein, which merely illustrate different ways of practicing the present invention.
[0199] Experimental Section Abbreviation LP = Landing Pad LP1P1=landing pad including attP1 LP2P2 = landing pad 2 containing attP2 and split intron CHO = Chinese hamster ovary FC-eGFP = enhanced green fluorescent protein fused to IgG1-derived FC TagBFP2 = blue fluorescent protein variant TagRFP-T = red fluorescent protein variant G418 = Also known as geneticin, a broad-spectrum antibiotic that selects for mammalian cells expressing the neomycin resistance gene (NeoR).
[0200] Sequence Listing The following sequences are used in the experimental section, but the present disclosure is not limited to these sequences. Therefore, variants of said sequences are also envisaged, whose function remains essentially the same as that of the original sequence.
[0201] attP1 (SEQ ID NO: 1)
[0202] [ka]
[0203] attB1 (SEQ ID NO: 2)
[0204] [ka]
[0205] attP2 (SEQ ID NO: 3)
[0206] [ka]
[0207] attB2 (SEQ ID NO: 4)
[0208] [ka]
[0209] PhiC31 gene (SEQ ID NO: 5)
[0210] [ka]
[0211] loxP (SEQ ID NO: 6)
[0212] [ka]
[0213] Cre gene (SEQ ID NO: 7)
[0214] [ka]
[0215] FC-eGFP gene (SEQ ID NO: 8)
[0216] [ka]
[0217] FC-TagBFP2 gene (SEQ ID NO: 9)
[0218] [ka]
[0219] TagRFP-T gene (SEQ ID NO: 10)
[0220] [ka]
[0221] eGFP gene (SEQ ID NO: 11)
[0222] [ka]
[0223] TagBFP2 gene (SEQ ID NO: 12)
[0224] [ka]
[0225] NeoR gene (SEQ ID NO: 13)
[0226] [ka] [Example]
[0227] Efficiency of phiC31 recombinase-mediated integration and selectable marker activation To investigate integration efficiency, HyClone CHO LP cells and non-LP HyClone CHO control cells were transfected with a combination of PhiC31 recombinase expression plasmids and either donor vector A or B (Figure 2). The donor vector contains expression cassettes for FC-eGFP and FC-TagBFP2, as well as the promoterless TagRFP-T gene, positioned to be activated upon integration at the LP in LP cells.
[0228] Two HyClone-CHO LP variants and matching donor vectors were investigated (see Figure 2). In HyClone-CHO LP1P1, the promoter in LP is located immediately downstream of attP1. In donor vector A, the TagRFP-T gene is located immediately upstream of attB1. In HyClone-CHO LP2P2, the 5' portion of the split intron is located between attP2 of LP and the downstream promoter. In donor vector B, the 3' portion of the split intron is located between attB2 and the upstream TagRFP-T gene. Integration efficiency was assessed by flow cytometry 7 days after transfection by measuring the percentage of cells displaying RFP signals above background (defined relative to non-transfected controls, see Figure 3).
[0229] In Figure 3, an example of the resulting flow cytometry data is shown for HyClone CHO LP2P2 in comparison to a control. The complete set of results is summarized in Table 1. For LP1P1, only a non-transfected control was used, while for LP2P2, both a random integration control (RI control, donor vector only) and a mock att integration control (donor vector + PhiC31 in a CHO cell line lacking LP) were performed. The data show that both LP variants are functional, but the LP2P2 variant utilizing a split-intron design provides superior integration efficiency.
[0230] [Table 1] [Example]
[0231] Efficiency of Cre recombinase-mediated excision of the vector backbone at the integration site HyClone CHO LP1P1 cells were transfected using a donor vector containing a PhiC31 expression plasmid and expression cassettes for FC-eGFP and FC-TagBFP2, as well as a promoterless TagRFP-T gene, positioned to be activated upon integration at the LP in LP cells (Figure 4). Cells that integrated the donor vector at the landing pad (LP) were enriched by several FACS sorting steps gating on Tag-RFP-T signals above background and the average expression of both FC-eGFP and FC-TagBFP2. The resulting sorted and expanded pool of cells was then transfected a second time using either (a) a Cre recombinase-expressing plasmid, (b) synthetic mRNA encoding Cre, or (c) a mock transfection solution lacking the Cre recombinase-encoding nucleic acid molecule. Seven days after the second transfection, all cell populations were analyzed by flow cytometry to assess the efficiency of excision of the region flanked by two loxP sites.
[0232] A plot of the flow cytometry analysis after Cre recombinase transfection can be seen in Figure 5. The data show an increase in cells that do not express FC-TagBFP2 for the two Cre recombinase-treated pools compared to the mock control. This, in turn, clearly indicates correct integration of the donor vector at the LP, such that FC-TagBFP2 is flanked by two loxP sites at which the Cre recombinase enzyme can act. The data show that the Cre recombinase-catalyzed excision reaction is highly efficient, with yields up to at least 80%. [Example]
[0233] Recurrent integration at the same genomic location by use of orthogonal attP / attB pairs HyClone CHO LP1P1 cells were transfected using a donor vector containing the PhiC31 expression plasmid and an attB1 sequence followed by an attP2 sequence, the 5' portion of a split intron, a promoter, a loxP sequence, and an expression cassette for eGFP (Figure 6, Donor Vector A). Seven days after transfection, eGFP-positive cells were sorted by FACS, followed by cell expansion. A second, more stringent round of selection for eGFP-positive cells was then performed. Cells expanded after the second round of positive selection (Figure 7, Step 1 selection) were transfected using synthetic mRNA encoding Cre recombinase. Seven days after Cre recombinase transfection, eGFP-negative cells were sorted using FACS. After expansion, the eGFP-negative cell pool was analyzed by flow cytometry (Figure 7, Step 2 selection). Data from the sorting and analysis steps are shown in Figure 7, top panel. During these steps, it is assumed that the landing pad in the CHO genome is altered, as shown by steps (1) and (2) in FIG.
[0234] To verify the functionality of the altered landing pad, the eGFP-negative pool obtained after the final selection (Figure 7, selection in step 2) was transfected with DNA donor vector B (Figure 6, step 3) and analyzed by flow cytometry 7 days after transfection (Figure 7, bottom panel). The data demonstrate functionality of the novel landing pad. Finally, cells from the eGFP-negative pool were cloned using single cell sorting by FACS, and the landing pad region of their genome was amplified by PCR and sequenced. Correct alteration of the landing pad was confirmed for multiple clones by sequencing (complete coverage of the novel landing pad region), indicating successful implementation of the alterations outlined in Figure 7. [Example]
[0235] Generation of SDI cell pools using two successive selection steps HyClone CHO LP2P2 cells (the clone generated in Example 3) were transfected with the PhiC31 recombinase expression plasmid and donor vector constructed according to FIG.
[0236] Starting two days after transfection, cells were cultured in the presence of G418 to select for cells that integrated the donor vector at the landing pad, thereby expressing the neomycin resistance gene (Neomycin-resistant gene). R ) was activated. After returning to high viability (>98%) after G418 selection, cells were sorted by FACS based on GFP / RFP double-positive signals. After expansion, sorted cells were transfected using synthetic mRNA encoding Cre recombinase. Seven days after Cre recombinase transfection, cells were FACS sorted by GFP-positive / RFP-negative signals (Figure 9, gate E). The final SDI pool was analyzed by flow cytometry after expansion.
[0237] FACS / flow cytometry data can be seen in Figure 9. The additional selection step after Cre recombinase transfection reduces heterogeneity in the pools, as shown by the mean and CV of eGFP signals for TagRFP-T positive (incorrect integration) and TagRFP-T negative (correct integration) cells.
[0238] [References] TIFF0007819113000015.tif189170
Claims
1. 1. A method for targeted integration of a donor vector into a predefined genomic location in an isolated eukaryotic cell, comprising: i) a. a nucleic acid sequence I1 containing a recognition site for a first DNA enzyme; b. a nucleic acid sequence E1 containing a recognition site for a second DNA enzyme; and c. promoter nucleic acid sequence P1 providing an isolated eukaryotic cell comprising a predefined genomic location; ii) a. nucleic acid sequence I2; b. a nucleic acid sequence of interest; c. a nucleic acid sequence E2 comprising a recognition site for the second DNA enzyme; d. a nucleic acid sequence encoding a first selectable marker; e. Optionally, an expression cassette encoding a second selectable marker providing a donor vector; iii) contacting the donor vector with the cell in the presence of a first DNA enzyme, wherein the presence of the first DNA enzyme allows for recombination between nucleic acid sequence I2 of the donor vector and nucleic acid sequence I1 present at the predefined genomic location of the cell; iv) selecting cells having the donor vector integrated at the predefined genomic location by detecting expression of the first selection marker in the cells, wherein the expression of the first selection marker is activated by the promoter nucleic acid sequence P1 at the predefined genomic location; and v) isolating the cells selected in the previous step. Including, the method further comprising step vi) forming part of or performed after step iii), wherein step vi) comprises excising, in the presence of a second DNA enzyme, a nucleic acid sequence flanked by nucleic acid sequences E1 and E2 from the predefined genomic location of the cell, wherein the presence of the second DNA enzyme permits recombination between the nucleic acid sequences E1 and E2, and the presence in the cell of the nucleic acid sequence flanked by nucleic acid sequences E1 and E2 is indicative of stable integration of the donor vector into the predefined genomic location of the cell; step vi) is performed before step v), and the donor vector of step ii) further comprises e) an expression cassette encoding a second selection marker, and the cells isolated in step v) are additionally selected based on their non-expression of the second selection marker, expression of the second selection marker indicating that the donor vector has integrated at a location different from the predefined genomic location of the cell; or the method further comprises step vii) performed after step vi), wherein step vi) comprises isolating cells, wherein the nucleic acid sequence flanked by the nucleic acid sequences E1 and E2 has been excised from the predefined genomic location of the cells obtained in step vi), wherein the donor vector of step ii) further comprises e) an expression cassette encoding a second selection marker, and the cells isolated in step vii) are selected based on their non-expression of the second selection marker, wherein expression of the second selection marker indicates that the donor vector has integrated at a location different from the predefined genomic location of the cell. method.
2. 2. The method of claim 1, wherein the nucleic acid sequence of interest of the donor vector in step ii) comprises at least one expression cassette comprising a gene encoding a protein of interest.
3. 3. The method of claim 2, wherein the excised nucleic acid sequence lacks the at least one expression cassette containing a gene encoding a protein of interest.
4. The donor vector of step ii) f. a nucleic acid sequence I3 comprising a recognition site for a first DNA enzyme; and g. promoter nucleic acid sequence P2 4. The method of claim 1, further comprising:
5. further comprising sequential targeted integration of n additional donor vectors into said predefined genomic locations of said eukaryotic cell; n is an integer greater than or equal to 1; (A) I. Providing the cells isolated in step v) of claim 1 or isolated in step vii) of claim 1; II. A. a nucleic acid sequence I4 capable of recombining with the nucleic acid sequence I3 present in the cell provided in the preceding step in the presence of the first DNA enzyme; B. a nucleic acid sequence of interest; C. Nucleic acid sequence E2; D. a nucleic acid sequence encoding a first selectable marker; and E. Optionally, an expression cassette encoding a second selectable marker; F. Optionally, a nucleic acid sequence I1 containing a recognition site for a first DNA enzyme and a promoter nucleic acid sequence P1 providing a donor vector comprising: III. introducing the donor vector of step II) into the cell in the presence of a first DNA enzyme, wherein the presence of the first DNA enzyme allows recombination between the nucleic acid sequence I4 of the donor vector and the nucleic acid sequence I3 present at the predefined genomic location of the cell; IV. selecting cells having the donor vector integrated at the predefined genomic location by detecting expression of the first selection marker in the cells, wherein the expression of the first selection marker is activated by the promoter nucleic acid sequence P2 at the predefined genomic location of the cells; V. isolating the cells selected in the preceding step; VI. excising, from the predefined genomic location of the cell isolated in step V, a nucleic acid sequence flanked by nucleic acid sequences E1 and E2 in the presence of a second DNA enzyme, wherein the presence of the second DNA enzyme permits recombination between nucleic acid sequences E1 and E2, and the presence in the cell of a nucleic acid sequence flanked by nucleic acid sequences E1 and E2 is indicative of stable integration of the donor vector into the predefined genomic location of the cell; VII. Isolating cells in which the nucleic acid sequence flanked by sequences E1 and E2 has been excised from the predefined genomic location of the cells isolated in step V. integrating a first additional donor vector into said predefined genomic location of said cell; (B) provided that n is greater than the number of additional donor vectors integrated at the predefined genomic location of the cell isolated in the preceding step; I. providing a cell isolated in the preceding step, comprising said nucleic acid sequence I1 integrated at said predefined genomic location; II. Step of carrying out step ii) according to claim 1 or 4, steps iii) to iv) according to claim 1, step v) according to claim 1, step vi) according to claim 1, and step vii) according to claim 1 integrating an additional donor vector into said cell at said predefined genomic location; (C) provided that n is greater than the number of additional donor vectors integrated at the predefined genomic location of the cell isolated in the preceding step; I. providing the cell obtained by carrying out the aforementioned step (B) and containing the nucleic acid sequence I3 integrated at the predefined genomic location; II. Repeating steps (A)II through (A)VII and step (B) until the cell has n additional donor vectors integrated at the predefined genomic locations. integrating an additional donor vector into said cell at said predefined genomic location, The method of claim 4 when dependent on claim 3, comprising:
6. - the cell of step i) according to claim 1 comprises n predefined genomic locations, each of which is a nucleic acid sequence I, I1, respectively 1 ~I1 n and I1 1 ~I1 n but differ from each other; - step ii) of claim 1 comprises providing 1 to n donor vectors, each of the 1 to n donor vectors comprising the nucleic acid sequence I2 1 ~I2 n and in the presence of the first DNA enzyme, 1 ~I1 n and each of said 1 to n donor vectors is capable of recombining with a nucleic acid sequence comprising a first selectable marker SM 1 ~SM n Including SM 1 ~SM n but differ from each other; - step iv) according to claim 1 comprises introducing the 1 to n donor vectors into the cells; - step v) of claim 1, wherein the different first selection markers SM in the cells 1 ~SM n selecting cells having each of the 1 to n donor vectors integrated at its corresponding predefined genomic location by detecting each of the 1 to n donor vectors; 3. The method according to claim 1, wherein n is an integer of 2 or greater.
7. The donor vector of step ii) according to claim 1, f. a nucleic acid sequence I3 comprising a recognition site for a first DNA enzyme; and g. promoter nucleic acid sequence P2 Further comprising:
3. The method of claim 2, wherein the excised sequence comprises the expression cassette containing a gene encoding a protein of interest.
8. the first DNA enzyme is a recombinase; (a) I1 comprises two recombinase recognition site variants I1a and I1b; (b) I2 comprises two recombinase recognition site variants I2a and I2b; (c) in the presence of the first DNA enzyme, I1a is capable of recombining with I2a and I1b is capable of recombining with I2b.
9. 9. The method of claim 8, wherein I1a is identical to I2a, I1b is identical to I2b, I1a, I1b, I2a, and I2b are each selected from loxP, rox, or FRT, and the first DNA enzyme is each selected from the group consisting of Cre recombinase, Dre recombinase, and FLP recombinase.
10. the first DNA enzyme is a recombinase; (a) I1 contains a single recombinase recognition site; (b) I2 contains a single recombinase recognition site; 8. The method of claim 1, wherein (c) I1 and I2 are capable of recombining in the presence of the first DNA enzyme.
11. 11. The method of claim 10, wherein the recombinase recognition site contained in I1 differs in sequence from the recombinase recognition site contained in I2, and the recombinase recognition site is selected from attB or attP.
12. 12. The method of claim 10 or 11, wherein the recombinase is PhiC31 or Bxb1 recombinase.
13. the first DNA enzyme is a gene-editing nuclease; (a) I1 comprises a cleavage site of the gene-editing nuclease and two sequence regions LHA1 and RHA1; (b) I2 contains two sequence regions LHA2 and RHA2 that are homologous to LHA1 and LHA2; 8. The method of claim 1, wherein (c) I1 and I2 are capable of recombining in the presence of the first DNA enzyme.
14. 14. The method of claim 1, wherein the promoter nucleic acid sequences P1 and / or P2 are operably fused to the 5' portion of a split intron when integrated at the predefined genomic location.
15. the nucleic acid sequence to be excised is (a) said nucleic acid sequence encoding a first selectable marker; (b) the promoter nucleic acid sequence P1 or P2; and / or (c) the expression cassette encoding a second selectable marker 15. The method of any one of claims 1 to 14, comprising:
16. 16. The method of any one of claims 1 to 15, wherein the first selection marker is selected from the group of (i) a fluorescent protein and (ii) a heterologous cell surface marker.
17. 17. The method of any one of claims 1 to 16, wherein the donor vector of step ii) further comprises an expression cassette encoding a second DNA enzyme, the expression of which is activated when the donor vector is integrated into the predefined genomic location of the cell of step i).
18. 18. The method of any one of claims 1 to 17, wherein the first DNA enzyme is expressed from an expression cassette encoding the first DNA enzyme present in the predefined genomic location of the cell of step i).
19. 19. An isolated eukaryotic cell obtainable by the method according to any one of claims 1 to 18.
20. 1. A method for producing a recombinant protein, comprising: i) obtaining an isolated eukaryotic cell comprising one or more nucleic acid sequences of interest integrated at a predefined genomic location by carrying out the method of any one of claims 1 to 18, wherein at least one nucleic acid sequence of interest comprises at least one expression cassette comprising a gene encoding a protein of interest; ii) producing in the cell of step i) the protein encoded by the gene of interest; and iii) isolating the protein of step ii). A method comprising: