Cell production method, cells, and protein production method

The method of integrating a target gene into host cells using a single recombinase and donor vector, with specific recognition sites, addresses the inefficiencies of current protein production methods by achieving high-expression and cost-effective production of therapeutic proteins.

WO2025115876A1PCT designated stage expired Publication Date: 2025-06-05FUJIFILM CORP
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/041881
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-11-26
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Current methods for producing cells that stably express therapeutic proteins, such as humanized monoclonal antibodies, are costly and inefficient due to the need for multiple types of donor vectors and enzymes for integration into the host genome.

Method used

A method involving the integration of a target gene into the genome of a host cell using a single type of recombinase and a single type of donor vector, with specific recognition sites for recombinase (RRS1 to RRS6) arranged to facilitate efficient and high-expression integration of multiple copies of the target gene into a highly expressed region of the host genome.

Benefits of technology

This method enables the creation of cells that highly express a target gene, resulting in enhanced productivity of therapeutic proteins with reduced costs and complexity compared to existing techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024041881_05062025_PF_FP_ABST
    Figure JP2024041881_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention involves introducing a donor vector of a target gene into a host cell, reacting a recombinase, and selecting a cell expressing the target gene from the host cell. The genome of the host cell has a region R containing RRS1, RRS2, RRS3, and RRS4, which are recognition sites of the recombinase, one by one in this order, and the donor vector has RRS5 and RRS6, which are recognition sites of the recombinase, and the target gene disposed between RRS5 and the RRS6. RRS1 and RRS4 can be recombined with RRS5 and cannot be recombined with RRS6, and RRS2 and RRS3 can be recombined with RRS6 and cannot be recombined with RRS5.
Need to check novelty before this filing date? Find Prior Art

Description

Method for producing cells, and method for producing cells and proteins

[0001] The present disclosure relates to methods for producing cells, cells and proteins.

[0002] Patent document 1 discloses a site-specific integration host cell containing an endogenous Fer1L4 gene, wherein an exogenous nucleotide sequence has been integrated into the Fer1L4 gene. Patent document 2 discloses a cell comprising an exogenous nucleic acid integrated at a specific site within an expression-enhancing locus, wherein the exogenous nucleic acid sequence encodes a bispecific antigen-binding protein. Patent document 3 discloses a cell comprising a first exogenous nucleic acid integrated into a first expression-enhancing locus and a second exogenous nucleic acid integrated into a second expression-enhancing locus, wherein the first and second exogenous nucleic acids together encode an antigen-binding protein. Patent Document 4 discloses a mammalian cell comprising a first recombination target site (RTS) chromosomally integrated at a first high integration (HI) locus, wherein the first HI locus is within an active genomic compartment of accessible chromatin and within approximately 30,000 base pairs of a TAD boundary, and the first HI locus overlaps with a region of the cellular genome that interacts with at least one enhancer element.

[0003] Non-Patent Document 1 discloses that the directionality of DNA integration by Bxb1 integrase depends only on the central dinucleotides of attP and attB. Non-Patent Document 2 discloses that among 15 serine recombinase candidates that integrate DNA into the human genome, Bxb1 integrase is the most accurate and efficient. Non-Patent Document 3 discloses that among four serine integrases φBT1, TG1, φRv1, and Bxb1, Bxb1 integrase is the most efficient.

[0004] European Patent Application Publication No. 2711428 International Publication No. WO 2017 / 184831 International Publication No. WO 2017 / 184832 International Publication No. WO 2020 / 072480

[0005] Molecular Cell, 2003, Vol. 12, 1101-1111BMC Biotechnology, 2013, 13:87Acta Biochim Biophys Sin, 2017, 49(1), 44-50

[0006] There is a technology for integrating a gene of interest into the genome of a host cell in order to create cells that stably produce a medical protein such as a humanized monoclonal antibody. From a cost perspective, it is preferable to have a small number of donor vectors for the gene of interest and enzymes that recombine the donor vector with the host genome. Furthermore, from the perspective of target protein production, it is preferable to place multiple copies of the gene of interest in a highly expressed region of the host genome. For example, when the target protein is an antibody, it is desirable to introduce one type of donor vector carrying a gene of interest containing an H-chain coding sequence and an L-chain coding sequence into the host cell, and then insert multiple copies of the gene of interest into a highly expressed region of the host genome through a recombination reaction using one type of recombinase.

[0007] The present disclosure has been made under the above circumstances. An object of the present disclosure is to provide a method for producing cells that highly express a gene of interest. An object of the present disclosure is to provide cells that highly express a gene of interest. An object of the present disclosure is to provide a method for producing a protein with excellent productivity.

[0008] Specific means for solving the problems include the following aspects. <1> A method for producing a cell by incorporating a target gene into the genome of a host cell using one type of recombinase and one type of donor vector, the method comprising: introducing the donor vector for the target gene into the host cell; reacting the recombinase in the host cell into which the donor vector has been introduced; and selecting cells that express the target gene from the host cells that have been reacted with the recombinase, wherein the genome of the host cell and the donor vector are as follows (1) to (4): (1) the genome of the host cell has a region R containing one each of RRS1, RRS2, RRS3, and RRS4, which are recognition sites for the recombinase, in this order; (2) the donor vector has RRS5 and RRS6, which are recognition sites for the recombinase, and the target gene located between RRS5 and RRS6; (3) RRS1 and RRS4 are capable of recombining with RRS5 but are incapable of recombining with RRS6; (4) RRS2 and RRS3 are recombinable with RRS6 but not with RRS5. <2> The method for producing a cell according to <1>, wherein the genome of the host cell further meets the following (5): (5) RRS1 and RRS4 have the same sequence, and RRS2 and RRS3 have the same sequence. <3> The method for producing a cell according to <1> or <2>, wherein the donor vector further meets the following (6): (6) The transcription direction of the target gene placed between RRS5 and RRS6 is from RRS6 to RRS5. <4> The method for producing a cell according to any one of <1> to <3>, wherein the genome of the host cell further meets the following (7): (7) The region R has a first selection marker gene placed between RRS1 and RRS2 and a second selection marker gene placed between RRS3 and RRS4. <5> The method for producing the cell according to any one of <1> to <4>, wherein the donor vector further comprises the following (8): (8) a third selection marker gene located between RRS5 and RRS6. <6> The method for producing the cell according to any one of <1> to <5>, further comprising introducing an expression vector for a recombinase into the host cell. <7> The method for producing the cell according to any one of <1> to <6>, wherein the recombinase is serine recombinase.<8> The method for producing a cell according to any one of <1> to <7>, wherein the host cell is a mammalian cell. <9> The method for producing a cell according to any one of <1> to <7>, wherein the host cell is a CHO cell. <10> The method for producing a cell according to any one of <1> to <9>, wherein the target gene is a gene encoding at least one selected from the group consisting of an enzyme, an antibody, an interleukin, a cytokine, a chemokine, a hormone, a growth factor, a transcription factor, a receptor, a viral preparation, a vaccine, a medical protein, a subunit thereof, and a fragment thereof.

[0009] <11> A cell having a gene of interest integrated into its genome, the cell satisfying the following conditions (A) to (C): (A) the genome has a region G containing one each of Site 1, Site 2, Site 3, and Site 4, in this order, which are sites formed by recombination of recombinase recognition sites; (B) Site 1 and Site 4 have sequence identity, and Site 2 and Site 3 have sequence identity; and (C) Region G has a gene of interest arranged between Site 1 and Site 2 and a gene of interest arranged between Site 3 and Site 4. <12> The cell according to <11>, further satisfying the following condition (D): (D) the transcription direction of the gene of interest arranged between Site 1 and Site 2 is from Site 2 to Site 1, and the transcription direction of the gene of interest arranged between Site 3 and Site 4 is from Site 3 to Site 4. <13> The cell according to <11> or <12>, wherein the recombinase is a serine recombinase. <14> The cell according to any one of <11> to <13>, wherein the cell is a mammalian cell. <15> The cell according to any one of <11> to <13>, wherein the cell is a CHO cell. <16> The cell according to any one of <11> to <15>, wherein the target gene is a gene encoding at least one selected from the group consisting of enzymes, antibodies, interleukins, cytokines, chemokines, hormones, growth factors, transcription factors, receptors, viral preparations, vaccines, medical proteins, subunits thereof, and fragments thereof. <17> A method for producing a protein, comprising culturing the cell according to any one of <11> to <16>, and expressing a protein encoded by the target gene.

[0010] According to the present disclosure, a method for producing cells that highly express a gene of interest is provided. According to the present disclosure, a cell that highly expresses a gene of interest is provided. According to the present disclosure, a method for producing a protein with excellent productivity is provided.

[0011] Fig. 1 is a conceptual diagram showing the form of recombination between region R of the host genome and a donor vector. Fig. 2 is a schematic diagram of the host genome construction vector used in the examples. Fig. 3 is a schematic diagram of the donor vector used in the examples. Fig. 4 is a schematic diagram of the recombinase expression vector used in the examples. Fig. 5 is a schematic diagram of region G of the clone created in the examples.

[0012] Hereinafter, embodiments of the present disclosure will be described. These descriptions and examples are intended to illustrate the embodiments and do not limit the scope of the embodiments. The mechanisms of action described in this disclosure include assumptions, and the validity of these assumptions does not limit the scope of the embodiments.

[0013] When describing embodiments of the present disclosure with reference to the drawings, the configuration of the embodiments of the present disclosure is not limited to the configuration shown in the drawings. The sizes of elements in the drawings are conceptual, and the relative relationships between the sizes of elements are not limited thereto.

[0014] In the present disclosure, the term "step" includes not only an independent step but also a step that cannot be clearly distinguished from other steps as long as the purpose of the step is achieved.

[0015] In the present disclosure, a numerical range indicated using "to" indicates a range that includes the numerical values ​​before and after "to" as the minimum and maximum values, respectively. In the numerical ranges described in stages in the present disclosure, the upper or lower limit value described in one numerical range may be replaced with the upper or lower limit value of another numerical range described in stages. Furthermore, in the numerical ranges described in the present disclosure, the upper or lower limit value of that numerical range may be replaced with a value shown in an example.

[0016] In the present disclosure, each component may contain multiple corresponding substances. When referring to the amount of each component in a composition in the present disclosure, if multiple substances corresponding to each component are present in the composition, the total amount of the multiple substances present in the composition is meant unless otherwise specified.

[0017] In this disclosure, the term "nucleic acid" includes all nucleic acids (e.g., deoxyribonucleic acid (DNA), ribonucleic acid (RNA), their analogs, natural products, and artificial products), as well as nucleic acids to which small molecules, groups (e.g., methyl groups), non-nucleic acid molecules, structures, etc. are linked. Nucleic acids may be single-stranded or double-stranded.

[0018] In the present disclosure, a donor vector is a substance that has the function of delivering exogenous nucleic acid into a cell and its genome, and is itself a nucleic acid. There are no limitations on the origin, form, or base sequence of the donor vector. The donor vector may be a circular nucleic acid or a linear nucleic acid. The donor vector may be a single-stranded nucleic acid or a double-stranded nucleic acid. The donor vector is preferably double-stranded DNA.

[0019] In the present disclosure, there is no limitation on the number of amino acid residues in a protein. Proteins include proteins in which amino acids have been post-translationally modified. Post-translational modifications of amino acids include phosphorylation, methylation, acetylation, glycosylation, lipidation, and the like. In the present disclosure, amino acids are represented by the three-letter and one-letter codes defined by IUPAC-IUBMB JCBN (IUPAC-IUBMB Joint Commission on Biochemical Nomenclature). Unless otherwise specified, amino acids referred to in the present disclosure are L-amino acids.

[0020] In the present disclosure, nucleotide sequence identity and amino acid sequence identity are calculated using BLAST (Basic Local Alignment Search Tool) (https: / / blast.ncbi.nlm.nih.gov / Blast.cgi).

[0021] In the present disclosure, recombinase is a general term for enzymes that recombine nucleic acids, and is a term that includes integrase. RRS is an abbreviation for recombinase recognition site.

[0022] In the present disclosure, when referring to the orientation or base sequence of an RRS (recombinase recognition site), of the two DNA strands constituting the double-stranded DNA, the DNA strand displaying the recombinase recognition sequence is referred to as the sense strand, and the strand complementary to the sense strand is referred to as the antisense strand. In the present disclosure, the identity of the base sequence of an RRS refers to the identity of the base sequence read in the 5'→3' direction of the sense strand (i.e., the DNA strand displaying the recombinase recognition sequence).

[0023] <Method for Producing Cells> The present disclosure provides a method for producing cells that highly express a gene of interest. The method for producing cells of the present disclosure is a method for producing cells by integrating a gene of interest into the genome of a host cell using one type of recombinase and one type of donor vector.

[0024] The method for producing cells disclosed herein includes introducing a donor vector for a gene of interest into a host cell, reacting a recombinase in the host cell into which the donor vector has been introduced, and selecting cells that express the gene of interest from the host cells that have been reacted with the recombinase.

[0025] The origin, size, and nucleotide sequence of the target gene are not limited. Examples of target genes include genes encoding at least one selected from the group consisting of enzymes, antibodies, interleukins, cytokines, chemokines, hormones, growth factors, transcription factors, receptors, viral preparations, vaccines, medical proteins, subunits thereof, and fragments thereof. That is, examples of proteins encoded by target genes (referred to as "target proteins" in the present disclosure) include at least one selected from the group consisting of enzymes, antibodies, interleukins, cytokines, chemokines, hormones, growth factors, transcription factors, receptors, proteins constituting viral preparations, vaccines, medical proteins, subunits thereof, and fragments thereof.

[0026] In the present disclosure, an antibody is not limited to an immunoglobulin, but may be any molecule that binds to an antigen. In the present disclosure, the term antibody includes antibody fragments and antigen-binding molecules. In the present disclosure, the heavy chain of an antibody is also referred to as an H chain, and the light chain of an antibody is also referred to as an L chain.

[0027] A gene of interest has all sequences necessary for the expression of a protein of interest. That is, a gene of interest includes a coding sequence for the protein of interest and all nucleic acids necessary for the transcription and translation of the coding sequence in a host cell (e.g., a promoter, a transcription terminator, a polyadenylation sequence). A gene of interest may include one copy of the coding sequence for the protein of interest, or two or more copies. For example, a gene of interest may include at least one copy of the coding sequence for each subunit to express all subunits of a heteromultimeric protein. For example, a gene of interest may include at least one copy of a sequence encoding an antibody heavy chain and a sequence encoding an antibody light chain.

[0028] The gene of interest may further include a sequence encoding at least one selected from the group consisting of nucleic acids constituting the viral formulation, transcriptional regulatory nucleic acids, and non-coding RNAs, such as microRNA (miRNA), short hairpin RNA (shRNA), small interfering RNA (siRNA), small nuclear RNA (snRNA), ribosomal RNA (rRNA), and transfer RNA (tRNA).

[0029] Host cells may be prokaryotic or eukaryotic. Examples of prokaryotic cells include bacterial cells. Examples of eukaryotic cells include fungi, yeast, insect cells, and mammalian cells.

[0030] Examples of bacterial cells include gram-negative bacterial cells such as Escherichia coli, Salmonella typhimurium, Serratia marcescens, Pseudomonas putida, Pseudomonas aeruginosa, etc.; and gram-positive bacterial cells such as Bacillus subtilis, etc. Preferred bacterial cells are those of the Enterobacteriaceae family, and a more preferred example is Escherichia coli, particularly strain B or K12.

[0031] An example of a fungus is Aspergillus oryzae.

[0032] Examples of yeast include Saccharomyces cerevisiae, Pichia pastoris, and Hansenula polymorpha.

[0033] Examples of insect cells include BmN cells derived from the silkworm (Bombyx mori), Sf9 cells and Sf21 cells derived from the armyworm (Spodoptera frugiperda), S2 cells derived from the fruit fly (Drosophila melanogaster), and Pv11 cells derived from the sleeping chironomid (Polypedilum vanderplanki).

[0034] Examples of mammalian cells include Chinese hamster ovary cells (CHO cells), baby hamster kidney cells (BHK cells), human embryonic kidney cell lines (e.g., HEK293 cells), human retinoblastoma-derived cell lines (e.g., PER.C6 cells), mouse myeloma cell lines (e.g., NS0 cells and SP2 / 0 cells), and established cell lines derived from these cells.

[0035] Examples of CHO cells include CHO-DG44 cells, CHO-K1 cells, CHO-DXB11 cells, and CHOpro3 cells. - Cells and cell lines derived from these cells.

[0036] Examples of mammalian cells include cells that have the ability to differentiate into other cells, such as pluripotent stem cells, such as embryonic stem cells (ES cells) and induced pluripotent stem cells (iPS cells), and multipotent stem cells, such as mesenchymal stem cells, tissue stem cells, and somatic stem cells.

[0037] Examples of means for introducing a donor vector into a host cell include electroporation, lipofection, microinjection, and cell infection with a viral vector. Electroporation is preferred from the viewpoints of high safety, high introduction efficiency, and low cytotoxicity.

[0038] Recombinase can be reacted in host cells into which a donor vector has been introduced by, for example, maintaining the culture environment of the host cells at the optimal temperature for the recombinase.

[0039] The recombinase may be an enzyme endogenous to the host cell, an enzyme introduced into the host cell by an expression vector, or an enzyme added to the host cell in the form of a protein or RNA. The recombinase expression vector may be integrated into the host genome or may exist in the host cell as an extrachromosomal element.

[0040] An example of an embodiment of the method for producing a cell of the present disclosure includes introducing an expression vector for a recombinase into a host cell. From the viewpoint of ensuring that the recombinase acts inside the host cell at the desired time, it is preferable that the recombinase be introduced into the host cell using an expression vector.

[0041] There is no limitation on the order in which the recombinase expression vector and the target gene donor vector are introduced into the host cell. The recombinase expression vector and the target gene donor vector may be introduced into the host cell together or separately. From the viewpoint of not increasing the number of steps and time required for producing the target cell, it is preferable to introduce the recombinase expression vector and the target gene donor vector into the host cell together.

[0042] The base nucleic acid and base sequence for constructing a recombinase expression vector are not limited. Examples of the base nucleic acid include viral vectors, non-viral vectors, and artificial nucleic acids. The base nucleic acid may be a circular nucleic acid or a linear nucleic acid. Examples of viral vectors include nucleic acids derived from adenovirus, adeno-associated virus, retrovirus, vaccinia virus, poxvirus, lentivirus, herpesvirus, baculovirus, or bacteriophage. Examples of non-viral vectors include artificial plasmids and bacterial vectors obtained by modifying bacterial genes.

[0043] There are no limitations on the origin, type, or form of recombinase. Recombinases commonly used in genetic engineering include serine recombinases (a type with a serine residue at the active site) and tyrosine recombinases (a type with a tyrosine residue at the active site), which are enzymes derived from bacteriophages. These were discovered as enzymes that integrate the phage genome into the bacterial genome when the bacteriophage infects the bacteria. Some serine recombinases and tyrosine recombinases have been confirmed to function in mammalian cells.

[0044] Preferred properties of the recombinase used in the cell production method of the present disclosure include high specificity for the base sequence of the recognition site, no other factors other than the recombinase are required for the recombination reaction, and the recombination reaction is irreversible.

[0045] As the recombinase used in the cell production method of the present disclosure, serine recombinase is preferred from the viewpoint of having all of the above properties. As the serine recombinase, one selected from the group consisting of Bxb1, φC31, TP901, A118, SPβc, TG1, φBT1, φRv1, φ370.1, Wβ, Pa01, and Pa03 is preferred from the viewpoint of being able to recombine mammalian genomes. Among these, Bxb1 recombinase (also known as Bxb1 integrase) is preferred from the viewpoint of excellent accuracy and efficiency of the recombination reaction.

[0046] When constructing an expression vector for a bacteriophage-derived recombinase, the codons of the recombinase gene are optimized to enable expression in a host cell, and a coding sequence for a nuclear localization signal is preferably added to the recombinase gene.

[0047] Cells expressing a target gene are selected from host cells reacted with a recombinase, for example, based on the concentration and / or purity of the target protein. For example, this can be done by setting a standard value for the concentration and / or purity of the target protein and selecting cells that reach the standard value, or by selecting cells with a relatively high concentration and / or purity of the target protein. Specifically, the following steps (S1) to (S4) are performed.

[0048] (S1) Adding a selection agent to the culture medium of the host cells. (S2) Dividing the host cells into single cells. (S3) Collecting a portion of the culture medium of the single-celled cells and measuring the concentration and / or purity of the target protein. (S4) Selecting cells whose concentration and / or purity of the target protein is equal to or higher than a standard value or is relatively high.

[0049] The purity of a target protein refers to the proportion (by mass or number) of the target protein in the total amount of multiple proteins. When the target protein is a multimeric protein, proteins that are not in their native form (e.g., proteins lacking some subunits, or proteins in which one subunit is replaced by another) may be produced. Therefore, it is desirable to have a low proportion of proteins that are not in their native form, i.e., a high purity of proteins in their native form.

[0050] The genome of the host cell (also referred to as the "host genome" in the present disclosure) and the donor vector used in the cell production method of the present disclosure have the following features (1) to (4).

[0051] (1) The genome of the host cell has a region R containing one each of recombinase recognition sites RRS1, RRS2, RRS3, and RRS4 in this order. (2) The donor vector has recombinase recognition sites RRS5 and RRS6 and a target gene located between RRS5 and RRS6. (3) RRS1 and RRS4 can recombine with RRS5 but cannot recombine with RRS6. (4) RRS2 and RRS3 can recombine with RRS6 but cannot recombine with RRS5.

[0052] Figure 1 shows the form of recombination between region R of the host genome and a donor vector. Figure 1 shows the form in which region R has a selection marker gene, but region R does not have to have a selection marker gene. The abbreviations in Figure 1 have the following meanings: GoI: gene of interest 1st MG: first selection marker gene 2nd MG: second selection marker gene

[0053] The cell production method of the present disclosure uses a host genome and a donor vector having any of the features (1) to (4), and thereby integrates two genes of interest into region R of the host genome. By providing region R in a highly expressed region of the host genome, integration of two genes of interest into the highly expressed region can be efficiently achieved.

[0054] An example of an embodiment of the genome of the host cell is further the following form (5): (5) RRS1 and RRS4 have the same sequence, and RRS2 and RRS3 have the same sequence.

[0055] The host genome being in form (5) increases the feasibility of integrating two genes of interest into region R of the host genome.

[0056] An example of an embodiment of the donor vector is the following embodiment (6): (6) The transcription direction of the target gene placed between RRS5 and RRS6 is from RRS6 to RRS5.

[0057] The transcription direction of the target gene is the potential transcription direction in the donor vector. Configuration (6) means that the elements constituting the target gene, i.e., the coding sequence of the target protein and all nucleic acids necessary for the transcription and translation of the coding sequence in a host cell (e.g., promoter, transcription terminator, polyadenylation sequence), are arranged from RRS6 to RRS5 in an order and direction that allows transcription and translation.

[0058] The donor vector has the configuration (6), which results in recombination such that the transcription directions of two adjacent target genes in the host genome move away from each other (←→). This configuration in which the transcription directions of two adjacent target genes move away from each other (←→) is expected to result in higher expression of the target genes than a configuration in which the transcription directions of two adjacent target genes move toward each other (→←).

[0059] The donor vector shown in Figure 1 is in form (6). In the form after recombination shown in Figure 1, the transcription directions of the two target genes adjacent to each other on the host genome are directed away from each other (←→).

[0060] An example of an embodiment of the genome of the host cell is further described below as form (7): (7) In region R, the host cell has a first selection marker gene located between RRS1 and RRS2 and a second selection marker gene located between RRS3 and RRS4.

[0061] The host genome having the form (7) facilitates the selection and enrichment of host cells in which recombination has occurred within region R after the application of the recombinase, or in which region R is present in the genome before the application of the recombinase.

[0062] Region R of the host genome shown in Figure 1 is in the form (7). The transcription directions of the first selection marker gene and the second selection marker gene may be the same (→→ or ←←), approaching each other (→←), or moving away from each other (←→).

[0063] An example of an embodiment of the donor vector is further described below as (8): (8) A third selectable marker gene is located between RRS5 and RRS6.

[0064] The donor vector's morphology (8) facilitates the selection and enrichment of host cells in which the gene of interest has been integrated into the genome.

[0065] RRS1 to RRS6, the genome of the host cell, and the donor vector will be described in detail below.

[0066] [RRS1 to RRS6] First, the characteristics of the RRSs of serine recombinases will be described. The RRSs of serine recombinases are generally called attP (phage attachment site) and attB (bacterial attachment site), since serine recombinases are derived from bacteriophages. Serine recombinases recombine DNA between attP and attB. Sequences that have a base sequence similar to that of native attP or attB and that are recognized by serine recombinases are called pseudo attP and pseudo attB. The number of bases in attP and pseudo attP may range from 1 bp to 1000 bp, generally ranges from 10 bp to 300 bp, and more generally ranges from 20 bp to 200 bp. The number of bases in attB and pseudo attB may range from 1 bp to 1000 bp, typically from 10 bp to 300 bp, and more typically from 20 bp to 200 bp.

[0067] Below, examples of the RRS of serine recombinase are shown, including the native attP and native attB of Bxb1 recombinase (also known as Bxb1 integrase), and examples of pseudo attP and pseudo attB. The RRS of serine recombinase may determine whether or not attP and attB can recombine depending on the difference between two bases in the center or vicinity of the sequence (referred to as the "center portion" in this disclosure). For Bxb1 recombinase, the difference between two bases in the center usually determines whether or not attP and attB can recombine. In each of the sequences below, the two bases in the center that determine whether or not attP and attB can recombine are underlined.

[0068] Native attP SEQ ID NO: 1: 5'-GTCGTGGTTTGTCTGGTCAACCACCGCGGTCTCAGTGGTGTACGGTACAAACCCCGAC-3' Native attB SEQ ID NO: 2: 5'-TCGGCCGGCTTGTCGACGACGGCGGTCTCCGTCGTCAGGATCATCCGGGC-3'

[0069] An example of pseudo attP: SEQ ID NO: 3: 5'-GTCGTGGTTTGTCTGGTCAACCACCGCGCTCTCAGTGGTGTACGGTACAAACCCCGAC-3' An example of pseudo attB: SEQ ID NO: 4: 5'-TCGGCCGGCTTGTCGACGACGGCGCTCTCCGTCGTCAGGATCATCCGGGC-3'

[0070] SEQ ID NO: 3 is a sequence in which the first base of the central portion "GT" of SEQ ID NO: 1 is modified to "CT" in the central portion. SEQ ID NO: 4 is a sequence in which the first base of the central portion "GT" of SEQ ID NO: 2 is modified to "CT" in the central portion.

[0071] In the RRS of serine recombinase, whether or not attP and attB can recombine may depend on the difference between the two bases in the center. In the case of SEQ ID NOs: 1 to 4, whether or not recombination can occur is as follows: SEQ ID NOs: 1 and 2, which have the same two bases in the center, can recombine. SEQ ID NOs: 3 and 4, which have the same two bases in the center, can recombine. SEQ ID NOs: 1 and 4, which have different two bases in the center, cannot recombine. SEQ ID NOs: 3 and 2, which have different two bases in the center, cannot recombine. In the present disclosure, "incapable of recombination" includes forms in which recombination is impossible and forms in which the probability of recombination occurring is lower than expected.

[0072] RRS1 to RRS6 are preferably designed as base sequences recognized by the same type of serine recombinase, taking into consideration the sequence of the two bases in the center. Specifically, the following forms (a) to (h) are preferred. Forms (a) to (h) make it easy to achieve forms (3) and (4).

[0073] (a) RRS1 and RRS4 have two identical bases in the center, and the entire base sequence is identical. The identity of the entire base sequence is preferably 80% or more, more preferably 90% or more, even more preferably 95% or more, and most preferably 100%. (b) RRS2 and RRS3 have two identical bases in the center, and the entire base sequence is identical. The identity of the entire base sequence is preferably 80% or more, more preferably 90% or more, even more preferably 95% or more, and most preferably 100%. (c) RRS1 (and RRS4) and RRS2 (and RRS3) have one or both of the two bases in the center that are different, and the entire base sequence is identical. The identity of the entire base sequence is preferably 80% or more, more preferably 90% or more, and even more preferably 95% or more. It is preferable that RRS1 (and RRS4) and RRS2 (and RRS3) have the same sequence except for one or both of the two bases in the central portion. (d) The number of bases in each of RRS1 to RRS4 is preferably 1 bp to 1000 bp, more preferably 10 bp to 300 bp, and even more preferably 20 bp to 200 bp. The difference in the number of bases between RRS1 to RRS4 is preferably 30% or less, more preferably 20% or less, and even more preferably 15% or less. It is most preferable that the number of bases in RRS1 to RRS4 is the same.

[0074] (e) The two bases in the center of RRS5 are identical to the two bases in the center of RRS1 and RRS4. (f) The two bases in the center of RRS6 are identical to the two bases in the center of RRS2 and RRS3. (g) RRS5 and RRS6 differ in one or both of the two bases in their center, and have identity over their entire base sequences. The identity over the entire base sequence is preferably 80% or more, more preferably 90% or more, and even more preferably 95% or more. RRS5 and RRS6 preferably have the same sequence except for one or both of the two bases in their center. (h) The number of bases in each of RRS5 and RRS6 is preferably 1 bp to 1000 bp, more preferably 10 bp to 300 bp, and even more preferably 20 bp to 200 bp. The difference in the number of bases between RRS5 and RRS6 is preferably 30% or less, more preferably 20% or less, and even more preferably 15% or less. Most preferably, RS5 and RRS6 have the same number of bases.

[0075] When serine recombinase is used to create cells, examples of embodiments of RRS1 to RRS6 include the following forms (i) and (j).

[0076] (i) RRS1 to RRS4 are the native attP and pseudo attP of the serine recombinase, and RRS5 and RRS6 are the native attB and pseudo attB of the serine recombinase. RRS1, RRS4, and RRS5 (or RRS2, RRS3, and RRS6) are the native att of the serine recombinase.

[0077] (j) RRS1 to RRS4 are the native attB and pseudo attB of the serine recombinase, and RRS5 and RRS6 are the native attP and pseudo attP of the serine recombinase. RRS1, RRS4, and RRS5 (or RRS2, RRS3, and RRS6) are the native att of the serine recombinase.

[0078] The native recognition sequence of serine recombinase can be found in academic papers, technical literature, etc. There are 16 possible sequences that can be applied to the two central bases of attP and attB. That is, in the 5' to 3' direction, they are "GT," "CT," "AT," "TT," "GA," "CA," "AA," "TA," "GC," "CC," "AC," "TC," "GG," "CG," "AG," and "TG." From these, the two central bases of the pseudo-att are selected.

[0079] When Bxb1 recombinase is used to create cells, an example of an embodiment of RRS1 to RRS6 is as follows. This embodiment is referred to as Ex(1). RRS1 and RRS4 are SEQ ID NO: 1, RRS2 and RRS3 are SEQ ID NO: 3, RRS5 is SEQ ID NO: 2, and RRS6 is SEQ ID NO: 4.

[0080] When Bxb1 recombinase is used to create cells, an example of an embodiment of RRS1 to RRS6 is as follows. This embodiment is referred to as Ex(2). RRS1 and RRS4 are SEQ ID NO: 3, RRS2 and RRS3 are SEQ ID NO: 1, RRS5 is SEQ ID NO: 4, and RRS6 is SEQ ID NO: 2.

[0081] When Bxb1 recombinase is used to create cells, an example of an embodiment of RRS1 to RRS6 is as follows. This embodiment is referred to as Ex(3). RRS1 and RRS4 are SEQ ID NO: 2, RRS2 and RRS3 are SEQ ID NO: 4, RRS5 is SEQ ID NO: 1, and RRS6 is SEQ ID NO: 3.

[0082] When Bxb1 recombinase is used to create cells, an example of an embodiment of RRS1 to RRS6 is as follows. This embodiment is referred to as Ex(4). RRS1 and RRS4 are SEQ ID NO: 4, RRS2 and RRS3 are SEQ ID NO: 2, RRS5 is SEQ ID NO: 3, and RRS6 is SEQ ID NO: 1.

[0083] Another example of an embodiment of RRS1 to RRS6 is a form in which two bases in the center are modified based on Ex(1) to Ex(4). Two of the 16 possible two-base sequences are selected, and one is used for the two bases in the center of RRS1, RRS4, and RRS5, and the other is used for the two bases in the center of RRS2, RRS3, and RRS6.

[0084] An example of an embodiment of Ex(1) is shown in Table 1. Table 1 shows only the bases of the sense strand (the DNA strand displaying the recognition sequence of the recombinase), and omits the bases of the antisense strand. The arrow indicates the 5' to 3' direction of the sense strand. In each sequence shown in Table 1, the two bases in the center that determine whether recombination between RRSs occurs are underlined.

[0085]

[0086] Ex(1) has other forms depending on the orientation of each RRS in addition to the forms shown in Table 1. The orientation of each RRS is not limited to the forms shown in Table 1, as long as it is an orientation that realizes the transfer of the target gene from the donor vector to two locations within region R.

[0087] When serine recombinase is used to generate cells, the orientations of RRS1 to RRS4 in region R and the orientations of RRS5 and RRS6 in the donor vector are preferably as follows to efficiently transfer the target gene from the donor vector to two locations within region R. In the following explanation, the orientation of RRS is indicated as the 5' to 3' direction of the sense strand (the DNA strand displaying the recognition sequence of the recombinase).

[0088] ・When the direction of RRS5 and RRS6 is "→ target gene ←", the direction of RRS1, RRS2, RRS3 and RRS4 is "→ ← → ←". (Form shown in Table 1) ・When the direction of RRS5 and RRS6 is "← target gene →", the direction of RRS1, RRS2, RRS3 and RRS4 is "← → ← →". ・When the direction of RRS5 and RRS6 is "→ target gene →", the direction of RRS1, RRS2, RRS3 and RRS4 is "→ → ← ←". ・When the direction of RRS5 and RRS6 is "← target gene ←", the direction of RRS1, RRS2, RRS3 and RRS4 is "← ← → →". The orientations of RRS1 and RRS4 are opposite. That is, the sense strand of RRS1 and the sense strand of RRS4 are different DNA strands. The orientations of RRS2 and RRS3 are opposite. That is, the sense strand of RRS2 and the sense strand of RRS3 are different DNA strands.

[0089] [Host Cell Genome] The genome of the host cell (also referred to as "host genome" in the present disclosure) has region R. Region R is a region that includes one each of RRS1, RRS2, RRS3, and RRS4, which are recognition sites for recombinase. The order of the RRSs in region R is RRS1, RRS2, RRS3, and RRS4. Region R is a continuous region. The host genome may have one region R throughout the genome, or may have two or more regions R.

[0090] By recombination with a donor vector, the R region allows insertion of a target gene between RRS1 and RRS2, and also between RRS3 and RRS4, so that two target genes can be inserted per R region.

[0091] An example of an embodiment of the host genome has a first selection marker gene located between RRS1 and RRS2 and a second selection marker gene located between RRS3 and RRS4 in region R. The first selection marker gene and the second selection marker gene each contain all nucleic acids necessary for gene expression. The size and base sequence of the first selection marker gene and the second selection marker gene are not limited.

[0092] The first and second selection marker genes may be the same gene or different genes, but from the viewpoint of not increasing the number of steps and time required for cell selection and enrichment, it is preferable that the first and second selection marker genes be the same gene.

[0093] One embodiment of the first selection marker gene and the second selection marker gene is a negative selection gene used to select and enrich for host cells in which recombination has occurred within region R. An example of a negative selection gene is a suicide gene that induces cell death in response to a specific drug. Examples of suicide genes include the herpes simplex virus-derived thymidine kinase gene (selection drug: ganciclovir), the inducible caspase 9 gene (selection drug: AP1903), and the cytosine deaminase gene (selection drug: 5-fluorocytosine).

[0094] An example of an embodiment of the first selection marker gene and the second selection marker gene is a gene that expresses a positive selection marker used to select and enrich host cells having region R in their genome. An example of a positive selection marker is a fluorescent protein. Any known fluorescent protein can be used as the fluorescent protein. The fluorescent protein is preferably a monomeric, highly bright fluorescent protein.

[0095] In one embodiment, one of a negative selection gene and a positive selection gene is placed between RRS1 and RRS2. In another embodiment, both a negative selection gene and a positive selection gene are placed between RRS1 and RRS2. In one embodiment, one of a negative selection gene and a positive selection gene is placed between RRS3 and RRS4. In another embodiment, both a negative selection gene and a positive selection gene are placed between RRS3 and RRS4.

[0096] In region R, the number of bases between the outer end of RRS1 and the outer end of RRS4, which is the recognition site farthest from RRS1, is, for example, 100 kbp or less, 70 kbp or less, 50 kbp or less, 30 kbp or less, or 10 kbp or less. The number of bases between the outer end of RRS1 and the outer end of RRS4 is, for example, 100 bp or more, 1 kbp or more, or 2 kbp or more.

[0097] In region R, the number of bases between the outer end of RRS2 (the end closer to RRS1) and the outer end of RRS3 (the end closer to RRS4) is preferably 50 bp or more, more preferably 100 bp or more, and even more preferably 200 bp or more. According to this embodiment, the two target genes inserted into region R by recombination are closely spaced at an appropriate distance, and high expression of the target genes can be expected.

[0098] Region R may be a region already present in the host genome, or may be a region newly formed in the host genome.

[0099] Formation of region R in the host genome is carried out, for example, by incorporating region R into the host genome using a vector carrying region R (referred to as a "host genome construction vector" in this disclosure).

[0100] The vector for constructing a host genome has at least RRS1, RRS2, RRS3, and RRS4 in this order. One embodiment of the vector for constructing a host genome has a first selection marker gene located between RRS1 and RRS2, and a second selection marker gene located between RRS3 and RRS4.

[0101] The base nucleic acid and base sequence for constructing a vector for constructing a host genome construction are not limited. Examples of the base nucleic acid include viral vectors, non-viral vectors, and artificial nucleic acids. The base nucleic acid may be a circular nucleic acid or a linear nucleic acid. Examples of viral vectors include nucleic acids derived from adenovirus, adeno-associated virus, retrovirus, vaccinia virus, poxvirus, lentivirus, herpesvirus, baculovirus, or bacteriophage. Examples of non-viral vectors include artificial plasmids and bacterial vectors obtained by modifying bacterial genes.

[0102] An example embodiment of a host genome is a host genome having region R inserted into at least one safe harbor in the host genome, with region R within the safe harbor.

[0103] A safe harbor within a genome is a region where a host cell survives even if a gene is inserted, and where the inserted gene is expressed. A safe harbor within a genome is identified by a chromosome number or an accession number and a base number in a public base sequence database. Examples of public base sequence databases include the International Nucleotide Sequence Databases (INSD) and the NCBI Reference Sequence Database (RefSeq). A safe harbor within a genome may be referred to by the name of a known gene present within or near that region.

[0104] The safe harbor in the genome where the region R is inserted may be a known safe harbor or a newly discovered safe harbor. Known safe harbors can be found in publicly available databases, academic papers, technical literature, etc.

[0105] When there are multiple safe harbors, at least one may be selected as the insertion region of region R. Methods for selecting a safe harbor include, for example, selecting a safe harbor in which the expression level (pg / cell / copy) of the protein encoded by the inserted gene is relatively high; or selecting a safe harbor in which the expression level (pg / cell / copy) of the protein encoded by the inserted gene exceeds a predetermined standard. The protein expression level of the safe harbor may be data obtained from publicly available databases, academic papers, technical literature, etc., or may be data obtained by actually inserting a gene into the safe harbor and measuring the protein expression level.

[0106] Insertion of region R into a target safe harbor within the genome is possible using known genome editing techniques.

[0107] [Donor Vector] The donor vector has recombinase recognition sites RRS5 and RRS6, and a target gene located between RRS5 and RRS6.

[0108] A gene of interest has all sequences necessary for the expression of a protein of interest. That is, a gene of interest includes a coding sequence for the protein of interest and all nucleic acids necessary for the transcription and translation of the coding sequence in a host cell (e.g., a promoter, a transcription terminator, a polyadenylation sequence). A gene of interest may include one copy of the coding sequence for the protein of interest, or two or more copies. For example, a gene of interest may include at least one copy of the coding sequence for each subunit to express all subunits of a heteromultimeric protein. For example, a gene of interest may include at least one copy of a sequence encoding an antibody heavy chain and a sequence encoding an antibody light chain.

[0109] Promoters that can be used in prokaryotic cells include those disclosed in J. Mol. Biol. 1986; 189(1): 113-30, phage polymerase promoters, and E. coli polymerase promoters. Specific examples include T7A1, T7A2, T7A3, λpL, λpR, lac, lacUV5, trp, tac, trc, phoA, and rrnB.

[0110] Examples of promoters that can be used in yeast cells include the gal promoter, the AOX1 promoter, the AOX2 promoter, the GAP promoter, the GAL1 promoter, and the GAL10 promoter.

[0111] Examples of promoters that can be used in insect cells include the polyhedrin promoter, P10 promoter, viral infection early expression protein (IE-1) promoter, MT promoter, COPIA promoter, CMV promoter, RSV promoter, SV40 promoter, heat shock protein promoter, OPIE2 promoter, and actin5C promoter.

[0112] Examples of promoters that can be used in mammalian cells include virus-derived promoters and housekeeping gene-derived promoters. Examples of virus-derived promoters include the human CMV promoter, rat CMV promoter, SV40 promoter, RSR-LTR promoter, and HSK-TK promoter. Examples of housekeeping gene-derived promoters include the hEF-1α promoter, Chinese hamster EF-1α promoter, β-actin promoter, and mouse phosphoglycerate kinase (mPGK) promoter. A preferred example of a promoter that can be used in mammalian cells is the EF-1α promoter, more preferably the hEF-1α promoter.

[0113] The gene of interest may contain a coding sequence for a secretory leader to facilitate extracellular transport or secretion of the protein of interest. A secretory leader is a type of signal peptide that induces extracellular transport or secretion of a polypeptide.

[0114] When a gene of interest contains a coding sequence for a secretory leader, the coding sequence for the secretory leader and the coding sequence for the protein of interest are arranged in the same reading frame. Here, "arranged in the same reading frame" means that the coding sequences for the secretory leader and the protein of interest are arranged so that they can be expressed as a single polypeptide. A linker or spacer coding sequence may or may not be present between the coding sequence for the secretory leader and the coding sequence for the protein of interest in the gene of interest. A preferred form is one in which the coding sequence for the protein of interest is arranged in the same reading frame downstream of the coding sequence for the secretory leader. From this form of gene of interest, a fusion protein is expressed in which the secretory leader is arranged at the N-terminus of the protein of interest. A more preferred form is one in which the coding sequence for the protein of interest is arranged contiguously in the same reading frame downstream of the coding sequence for the secretory leader. From this form of gene of interest, a fusion protein is expressed in which the secretory leader is linked to the N-terminus of the protein of interest. Here, "downstream" refers to the order of the two coding sequences. When both coding sequences are arranged so that coding sequence B is transcribed after coding sequence A, coding sequence B is said to be arranged downstream of coding sequence A. The secretory leader of the fusion protein is generally cleaved from the fusion protein during transport or secretion of the fusion protein.

[0115] Examples of secretory leaders include fibronectin secretory leaders, collagen secretory leaders, and albumin secretory leaders. From the viewpoint of high extracellular secretion rate of the fusion protein, the fibronectin secretory leader is preferred.

[0116] Examples of fibronectin secretory leaders include amphibian fibronectin secretory leaders and mammalian fibronectin secretory leaders. Examples of amphibian fibronectin secretory leaders include Xenopus fibronectin secretory leaders. Examples of mammalian fibronectin secretory leaders include human, rat, mouse, bovine, porcine, canine, feline, and Chinese hamster fibronectin secretory leaders and functional equivalents thereof.

[0117] The origin of the fibronectin secretory leader is preferably selected depending on the type of host cell. When the host cell is a human cell, it is preferable to use a human fibronectin secretory leader for the gene of interest. When the host cell is a rat cell, it is preferable to use a rat fibronectin secretory leader for the gene of interest. When the host cell is a CHO cell, it is preferable to use a Chinese hamster fibronectin secretory leader for the gene of interest.

[0118] An example embodiment of a gene of interest comprises a hEF-1α promoter, a coding sequence for a fibronectin secretory leader, a coding sequence for a protein of interest, and a polyA sequence operably linked to each other.

[0119] The transcription direction of the target gene located between RRS5 and RRS6 is preferably from RRS6 to RRS5. In this embodiment, the transcription directions of the two target genes inserted into region R of the host genome are directed away from each other (←→).

[0120] An example of an embodiment of the donor vector has a third selection marker gene located between RRS5 and RRS6. The third selection marker gene contains all nucleic acids necessary for gene expression. The size and base sequence of the third selection marker gene are not limited. The third selection marker gene is a gene that expresses a positive selection marker used to select and enrich host cells in which a gene of interest has been integrated into the genome.

[0121] A third example of a selectable marker gene is a gene that confers resistance to a selective drug. Examples of the selective drug include antibiotics and enzyme inhibitors.

[0122] When the selective drug is an antibiotic, the selectable marker gene is an antibiotic resistance gene, which is a gene for an enzyme that degrades the antibiotic, such as a puromycin resistance gene, a hygromycin resistance gene, a neomycin resistance gene, a chloramphenicol resistance gene, a tetracycline resistance gene, an erythromycin resistance gene, a spectinomycin resistance gene, a kanamycin resistance gene, a G418 resistance gene, a blasticidin resistance gene, a zeocin resistance gene, a phleomycin resistance gene, and an ampicillin resistance gene.

[0123] An example of a system in which the selective drug is an enzyme inhibitor is the DHFR-MTX system. In the DHFR-MTX system, the selective drug is methotrexate (MTX) and the selectable marker gene is the dihydrofolate reductase (DHFR) gene. The DHFR-MTX system is effective in host cells (e.g., CHO-DG44 cells) that lack the DHFR gene.

[0124] An example of a system in which the selective drug is an enzyme inhibitor is the GS-MSX system. In the GS-MSX system, the selective drug is methionine sulfoximine (MSX) and the selective marker gene is the glutamine synthetase (GS) gene. The GS-MSX system is effective in host cells lacking the GS gene (e.g., GS knockout CHO cells).

[0125] A third example of a selectable marker gene is a fluorescent protein gene. Any known fluorescent protein can be used as the fluorescent protein. The fluorescent protein is preferably a monomeric, highly fluorescent protein. When the host genome contains a fluorescent protein gene, it is preferable to avoid overlapping of excitation wavelengths and emission wavelengths between fluorescent proteins.

[0126] As the third selection marker gene, a combination of the above-mentioned genes may be used. For example, a drug resistance gene and a fluorescent protein gene may be placed between RRS5 and RRS6.

[0127] The base nucleic acid and base sequence for constructing the donor vector are not limited. Examples of the base nucleic acid include viral vectors, non-viral vectors, and artificial nucleic acids. The base nucleic acid may be a circular nucleic acid or a linear nucleic acid. Examples of viral vectors include nucleic acids derived from adenovirus, adeno-associated virus, retrovirus, vaccinia virus, poxvirus, lentivirus, herpesvirus, baculovirus, or bacteriophage. Examples of non-viral vectors include artificial plasmids and bacterial vectors obtained by modifying bacterial genes.

[0128] <Cells> The present disclosure provides cells that highly express a gene of interest. The cell of the present disclosure is a cell in which an exogenous gene of interest has been integrated into the genome.

[0129] The origin, size, and nucleotide sequence of the target gene are not limited. Examples of target genes include genes encoding at least one selected from the group consisting of enzymes, antibodies, interleukins, cytokines, chemokines, hormones, growth factors, transcription factors, receptors, viral preparations, vaccines, medical proteins, subunits thereof, and fragments thereof. That is, examples of target proteins include at least one selected from the group consisting of enzymes, antibodies, interleukins, cytokines, chemokines, hormones, growth factors, transcription factors, receptors, proteins constituting viral preparations, vaccines, medical proteins, subunits thereof, and fragments thereof.

[0130] A gene of interest has all sequences necessary for the expression of a protein of interest. That is, a gene of interest includes a coding sequence for the protein of interest and all nucleic acids necessary for the transcription and translation of the coding sequence in a cell (e.g., a promoter, a transcription terminator, a polyadenylation sequence). A gene of interest may include one copy of the coding sequence for the protein of interest, or two or more copies. For example, a gene of interest may include at least one copy of the coding sequence for each subunit to express all subunits of a heteromultimeric protein. For example, a gene of interest may include at least one copy of a sequence encoding an antibody heavy chain and a sequence encoding an antibody light chain.

[0131] The gene of interest may further include a sequence encoding at least one selected from the group consisting of nucleic acids constituting the viral formulation, transcriptional regulatory nucleic acids, and non-coding RNAs, such as microRNA (miRNA), short hairpin RNA (shRNA), small interfering RNA (siRNA), small nuclear RNA (snRNA), ribosomal RNA (rRNA), and transfer RNA (tRNA).

[0132] The cells of the present disclosure may be prokaryotic or eukaryotic cells. Examples of prokaryotic cells include bacterial cells. Examples of eukaryotic cells include fungi, yeast, insect cells, and mammalian cells. Specific examples of bacterial cells, fungi, yeast, and insect cells are the same as those given in the description of the method for producing the cells.

[0133] Examples of mammalian cells include Chinese hamster ovary cells (CHO cells), baby hamster kidney cells (BHK cells), human embryonic kidney cell lines (e.g., HEK293 cells), human retinoblastoma-derived cell lines (e.g., PER.C6 cells), mouse myeloma cell lines (e.g., NS0 cells and SP2 / 0 cells), and established cell lines derived from these cells.

[0134] Examples of CHO cells include CHO-DG44 cells, CHO-K1 cells, CHO-DXB11 cells, and CHOpro3 cells.- Cells and cell lines derived from these cells.

[0135] Examples of mammalian cells include cells differentiated from mammalian cells having differentiation potential, such as cells differentiated by introducing a gene of interest into pluripotent stem cells (ES cells, iPS cells, etc.) or multipotent stem cells (mesenchymal stem cells, tissue stem cells, somatic stem cells, etc.).

[0136] The cells of the present disclosure have the following forms (A) to (C).

[0137] (A) The genome has a region G containing one each of Site 1, Site 2, Site 3, and Site 4, in that order, which are sites formed by recombination of recombinase recognition sites. (B) Site 1 and Site 4 have sequence identity, and Site 2 and Site 3 have sequence identity. (C) Region G has a gene of interest located between Site 1 and Site 2, and a gene of interest located between Site 3 and Site 4.

[0138] In the present disclosure, the identity of the base sequences of Sites 1 to 4 refers to the identity of the base sequences read in the 5' to 3' direction on the DNA strand adjacent to each site, in the 5' to 3' direction toward the target gene. The read strand of Site 1 is different from the read strand of Site 4, and the read strand of Site 2 is different from the read strand of Site 3.

[0139] The sequence identity between Site 1 and Site 4 is, for example, 80% or more, 90% or more, 95% or more, or 100%. The sequence identity between Site 2 and Site 3 is, for example, 80% or more, 90% or more, 95% or more, or 100%.

[0140] The cells of the present disclosure can be produced using one type of recombinase and a host genome and a donor vector having any of the forms (1) to (4). When the host genome and donor vector have any of the forms (1) to (4), the cells of the present disclosure can have any of the forms (A) to (C). Figure 1 shows examples of the form of region G produced by a host genome and a donor vector having any of the forms (1) to (4).

[0141] When the cell of the present disclosure is a cell produced by one type of recombinase and a host genome and donor vector having forms (1) to (4), Site 1 is a site formed by recombination between RRS1 and RRS5, Site 2 is a site formed by recombination between RRS2 and RRS6, Site 3 is a site formed by recombination between RRS3 and RRS6, and Site 4 is a site formed by recombination between RRS4 and RRS5. The number of bases in each of Sites 1 to 4 may be in the range of 1 bp to 1000 bp, generally in the range of 10 bp to 300 bp, and more generally in the range of 20 bp to 200 bp.

[0142] An example of an embodiment of Sites 1 to 4 is a site formed by recombination of recognition sites for a serine recombinase, such as one selected from the group consisting of Bxb1, φC31, TP901, A118, SPβc, TG1, φBT1, φRv1, φ370.1, Wβ, Pa01, and Pa03.

[0143] Examples of embodiments of Site 1 to Site 4 include the following forms (a) to (d).

[0144] (a) Site 1 and site 4 have identical two bases in the center and have identical entire base sequences. The identity of the entire base sequences is, for example, 80% or more, 90% or more, 95% or more, or 100%. (b) Site 2 and site 3 have identical two bases in the center and have identical entire base sequences. The identity of the entire base sequences is, for example, 80% or more, 90% or more, 95% or more, or 100%. (c) Site 1 (and site 4) and site 2 (and site 3) have different one or both of the two bases in the center and have identical entire base sequences. The identity of the entire base sequences is, for example, 80% or more, 90% or more, or 95% or more. Site 1 (and site 4) and site 2 (and site 3) may have identical sequences except for one or both of the two bases in the center. (d) Sites 1 to 4 may each have a number of bases ranging from 1 bp to 1000 bp, typically ranging from 10 bp to 300 bp, and more typically ranging from 20 bp to 200 bp.

[0145] In one embodiment of Sites 1 to 4, Site 1 and Site 4 have the same sequence, and Site 2 and Site 3 have the same sequence. Cells having such morphology can be produced by a host genome and a donor vector having morphology (1) to (4) and morphology (5).

[0146] An example of an embodiment of the cell of the present disclosure is further the following form (D): (D) The transcription direction of the target gene located between Site 1 and Site 2 is from Site 2 to Site 1, and the transcription direction of the target gene located between Site 3 and Site 4 is from Site 3 to Site 4.

[0147] Configuration (D) means that the transcription directions of two target genes arranged in region G are directions that move away from each other (←→). A configuration in which the transcription directions of two adjacent target genes are directions that move away from each other (←→) is expected to result in higher expression of the target genes than a configuration in which the transcription directions of two adjacent target genes are directions that move toward each other (→←).

[0148] Cells comprising morphology (D) can be produced by host genomes and donor vectors comprising morphology (1) to (4) and morphology (6).

[0149] The post-recombination form shown in Figure 1 is form (D), in which the transcription directions of the two target genes aligned within region G are directed away from each other (←→).

[0150] An example of an embodiment of the cell of the present disclosure has, in region G, a selection marker gene (1) located between site 1 and site 2, and a selection marker gene (2) located between site 3 and site 4. The selection marker genes (1) and (2) are genes that express positive selection markers used for selecting and enriching the cells of the present disclosure. Cells having this morphology can be produced using a host genome and donor vector having morphology (1) to (4) and morphology (8). Specific examples of the selection marker genes (1) and (2) are the same as the third selection marker gene listed in the explanation of the donor vector.

[0151] Region G is a continuous region. The cell of the present disclosure may have one region G or two or more regions G throughout the genome.

[0152] In region G, the number of bases between the outer end of site 1 and the outer end of site 4, which is the site farthest from site 1, is, for example, 100 kbp or less, 70 kbp or less, 50 kbp or less, 30 kbp or less, or 10 kbp or less. The number of bases between the outer end of site 1 and the outer end of site 4 is, for example, 100 bp or more, 1 kbp or more, or 2 kbp or more.

[0153] In region G, the number of bases between the outer end of site 2 (the end closer to site 1) and the outer end of site 3 (the end closer to site 4) is preferably 50 bp or more, more preferably 100 bp or more, and even more preferably 200 bp or more. According to this embodiment, the two target genes present in region G are closely aligned at an appropriate distance, and high expression of the target genes can be expected.

[0154] <Protein Production Method> The present disclosure provides a protein production method with excellent productivity. The protein production method of the present disclosure uses cells that highly express a target gene, thereby achieving excellent target protein productivity.

[0155] The method for producing a protein of the present disclosure involves culturing the cells of the present disclosure to express the protein encoded by the gene of interest. By culturing the cells, the protein of interest is produced within the cells, and the protein of interest accumulates in the culture medium and / or the cells.

[0156] The cell culture method and medium composition may be selected depending on the type of cell. Culture conditions (e.g., culture scale, cell density, temperature, and CO 2 The concentration may also be selected depending on the type of cells.

[0157] An example of an embodiment of the protein production method of the present disclosure includes recovering a target protein from a culture medium. Methods for recovering a target protein from a culture medium include, for example, centrifugation, filtration, diafiltration, ion exchange chromatography, affinity chromatography, hydrophobic interaction chromatography, gel filtration chromatography, and high-performance liquid chromatography (HPLC). The recovered target protein is used, for example, in the production of a pharmaceutical composition.

[0158] An example of an embodiment of the protein production method of the present disclosure includes recovering cells in which a target protein has accumulated from a culture medium. Methods for recovering cells from a culture medium include, for example, centrifugation and filtration. The target protein accumulates inside or on the surface of the cells depending on its properties. The recovered cells are, for example, administered, transfused, or transplanted into a mammal.

[0159] The cell production method of the present disclosure will be described in more detail below using specific examples. The materials, processing procedures, etc. shown in the following specific examples can be changed as appropriate without departing from the spirit of the present disclosure. The scope of the cell production method of the present disclosure should not be construed as being limited by the specific examples shown below.

[0160] The base sequences of RRS1 to RRS6 in the following examples are as follows: In each of the sequences below, the two bases in the center that determine whether or not recombination between RRSs occurs are underlined.

[0161] RRS1 and RRS4 SEQ ID NO: 1: 5'-GTCGTGGTTTGTCTGGTCAACCACCGCGGTCTCAGTGGTGTACGGTACAAACCCCGAC-3' RRS5 SEQ ID NO: 2: 5'-TCGGCCGGCTTGTCGACGACGGCGGTCTCCGTCGTCAGGATCATCCGGGC-3'

[0162] RRS2 and RRS3 SEQ ID NO: 3: 5'-GTCGTGGTTTGTCTGGTCAACCACCGCGCTCTCAGTGGTGTACGGTACAAACCCCGAC-3' RRS6 SEQ ID NO: 4: 5'-TCGGCCGGCTTGTCGACGACGGCGCTCTCCGTCGTCAGGATCATCCGGGC-3'

[0163] RRS1 and RRS4 are the native attP of Bxb1 recombinase (also known as Bxb1 integrase). RRS5 is the native attB of Bxb1 recombinase. RRS2 and RRS3 are sequences in which the two central bases "GT" of the native attP of Bxb1 recombinase have been modified to "CT". RRS6 is a sequence in which the two central bases "GT" of the native attB of Bxb1 recombinase have been modified to "CT".

[0164] <Construction of a vector for constructing a host genome> A vector for constructing a host genome was prepared using a custom synthesis service for artificial genes. Hereinafter, this vector will be referred to as "Vector A." Vector A has RRS1 to RRS4, a first negative selection gene between RRS1 and RRS2, and a second negative selection gene between RRS3 and RRS4. The first and second negative selection genes are herpes simplex virus-derived thymidine kinase genes. Vector A has a replication origin for amplification using Escherichia coli and an ampicillin resistance gene as a selection marker.

[0165] A schematic diagram of the structure of Vector A is shown in Figure 2. The gene arrangement and transcription direction are as shown in Figure 2. The orientations of RRS1 to RRS4 are as shown in Table 1. The total length of Vector A is approximately 8 kbp, the number of bases between the outer end of RRS1 and the outer end of RRS4 is approximately 4.5 kbp, and the number of bases between the outer end of RRS2 (the end closest to RRS1) and the outer end of RRS3 (the end closest to RRS4) is approximately 300 bp.

[0166] <Construction of donor vector> The following DNA fragment (1) and DNA fragment (2) were synthesized using a custom synthesis service for artificial genes. DNA fragment (1): red fluorescent protein mCherry gene-puromycin resistance gene. Contains all nucleic acids necessary for gene expression, with the coding sequence for the 2A self-cleaving peptide present between the two genes. DNA fragment (2): antibody L chain gene-H chain gene. Contains all nucleic acids necessary for expression of each chain gene.

[0167] Using In-Fusion HD Cloning Kit (Takara Bio, product code 639648), a vector was constructed by ligating DNA fragment (1) and DNA fragment (2). DNA fragment (2) was further ligated to the constructed vector to obtain the following DNA fragment (3). DNA fragment (3): red fluorescent protein mCherry gene-puromycin resistance gene-L chain gene-H chain gene-L chain gene-H chain gene

[0168] A DNA fragment was prepared by PCR in which RRS5 was added to one end of DNA fragment (3) and RRS6 was added to the other end, and this DNA fragment was ligated to a base vector to prepare a donor vector. Hereinafter, this vector will be referred to as "donor vector B."

[0169] Donor vector B has an antibody gene (light chain gene-heavy chain gene-light chain gene-heavy chain gene) as the gene of interest between RRS5 and RRS6. Donor vector B has a red fluorescent protein mCherry gene-puromycin resistance gene as selection marker genes between RRS5 and RRS6. Donor vector B has a replication origin for amplification using E. coli and an ampicillin resistance gene as a selection marker. Hereinafter, the entire group of genes placed between RRS5 and RRS6 will be referred to as "GoI-MG."

[0170] Figure 3 shows a schematic diagram of the donor vector B. The gene order and transcription direction are as shown in Figure 3. The orientations of RRS5 and RRS6 are as shown in Table 1.

[0171] <Construction of Bxb1 Expression Vector> An expression vector for Bxb1 recombinase was prepared using a custom synthesis service for artificial genes. Hereinafter, this expression vector will be referred to as "Vector C." Vector C has a replication origin for amplification using E. coli, an ampicillin resistance gene as a selection marker, and the Bxb1 gene. The Bxb1 gene is a gene whose codons have been optimized for expression in mammalian cells, and to which an SV40-derived nuclear localization signal sequence has been added at the 5' end.

[0172] Figure 4 shows a schematic diagram of the structure of vector C. The gene arrangement and transcription direction are as shown in Figure 4.

[0173] <Cell culture> CHO-DG44 cells were used as host cells. For the maintenance and passage of CHO-DG44 cells, a liquid medium prepared by adding hypoxanthine / thymidine (HT Supplement (100X)) to a serum-free basal medium (Thermo Fisher Scientific, CD OptiCHO Medium) was used. For single cell cloning experiments, a liquid medium prepared by adding 10% (v / v) fetal bovine serum to an IMDM basal medium was used.

[0174] <Establishment of host cells> Vector A was introduced into CHO-DG44 cells by electroporation. This treatment was performed using a 4D-Nucleofector device and SF Cell Line 4D-NucleofectorX Kit L (Lonza; "Nucleofector" is a registered trademark). 11 μg of Vector A was used for the treatment.

[0175] After introducing Vector A, the cells were maintained and passaged using a passage medium. On day 6 of culture, one cell was seeded per well in a 96-well plate, and the cells were single cloned. Genomes were extracted from the 24 established clones, and one clone with one copy of Vector A inserted into the genome was obtained using a digital PCR system (Bio-Rad Laboratories, ddPCR Supermix for Probes (No dUTP) #1863024). Hereinafter, this clone is referred to as "CHO-159B3 cells."

[0176] Region R, the region into which Vector A was inserted in the genome of CHO-159B3 cells, was amplified by PCR and subjected to Sanger sequencing analysis (using the contract analysis service of FASMAC Corporation). The results of the sequence analysis confirmed that RRS1 to RRS4, the first TK gene, and the second TK gene were present in region R as designed. That is, the order of RRS1 to RRS4, the first TK gene, and the second TK gene in region R was as shown in the schematic diagram of FIG. 1, and the orientation of RRS1 to RRS4 in region R was as shown in Table 1. The number of bases between the outer end of RRS1 and the outer end of RRS4 was approximately 4.5 kbp, and the number of bases between the outer end of RRS2 (the end closest to RRS1) and the outer end of RRS3 (the end closest to RRS4) was approximately 300 bp.

[0177] When the copy numbers of the first TK gene and the second TK gene (relative to the Txnip gene copy number) in the genome of CHO-159B3 cells were measured, it was found that each of the first TK gene and the second TK gene had about 1 copy.

[0178] <Integration of target gene into host genome> Donor vector B and vector C were introduced into CHO-159B3 cells by electroporation. This process was performed using a 4D-Nucleofector device and SF Cell Line 4D-NucleofectorX Kit L (Lonza). 12 μg of donor vector B and 6 μg of vector C were used for the process.

[0179] After introducing donor vector B and vector C, the cells were passaged in a subculture medium to monitor the expression and response of Bxb1 recombinase. On day 11 of culture, 30 cells were seeded per well in a 96-well plate, and selection was performed using ganciclovir and puromycin, as well as visual selection using red fluorescence.

[0180] Genomes were extracted from the 21 established clones, and the region corresponding to region R was amplified by PCR and subjected to Sanger sequencing. Clones in which sites formed by recombination of Bxb1 recombinase were present at the positions of RRS1 and RRS4, which were present in region R, were initially selected.

[0181] The clones selected from the primary selection were subjected to full-length sequence analysis of region G formed by recombination between region R and donor vector B, and one clone was selected in which sites 1 to 4 and GoI-MG were present as designed. Sequence analysis was performed using a long-read sequencer MinION Mk1C (Oxford Nanopore Technologies).

[0182] Region G of one selected clone contained sites 1 to 4, a GoI-MG located between sites 1 and 2, and a GoI-MG located between sites 3 and 4. The transcription directions of the two GoI-MGs were pointing away from each other (←→). The base sequence of region G shared over 99% identity with the designed base sequence.

[0183] Figure 5 shows a schematic diagram of region G of the above clone. The gene order and transcription direction are as shown in Figure 5. Of the two GoI-MGs, Figure 5 details the GoI-MG located between site 1 and site 2. The GoI-MG located between site 3 and site 4 contains the same gene group as the GoI-MG located between site 1 and site 2, but in the opposite direction.

[0184] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[0185] The disclosure of Japanese Application No. 2023-202229, filed on November 29, 2023, is incorporated herein by reference in its entirety.

Claims

1. A method for producing a cell by incorporating a gene of interest into the genome of a host cell using one type of recombinase and one type of donor vector, comprising: introducing the donor vector for the gene of interest into the host cell; reacting the recombinase in the host cell into which the donor vector has been introduced; and selecting a cell expressing the gene of interest from the host cell into which the recombinase has been reacted, wherein the genome of the host cell and the donor vector are as follows: (1) the genome of the host cell has a region R including one each of RRS1, RRS2, RRS3, and RRS4 in this order as recognition sites for the recombinase; (2) the donor vector has RRS5 and RRS6 as recognition sites for the recombinase, and the gene of interest arranged between RRS5 and RRS6; (3) the RRS1 and RRS4 are capable of recombining with RRS5 and incapable of recombining with RRS6; and (4) the RRS2 and RRS3 are capable of recombining with RRS6 and incapable of recombining with RRS5.

2. The method for producing a cell according to claim 1, wherein the genome of the host cell further comprises the following (5): (5) RRS1 and RRS4 have the same sequence, and RRS2 and RRS3 have the same sequence.

3. The method for producing a cell according to claim 1, wherein the donor vector is further defined as (6) below: (6) The transcription direction of the target gene located between RRS5 and RRS6 is from RRS6 to RRS5.

4. A method for producing a cell according to claim 1, wherein the genome of the host cell further comprises the following (7): (7) a first selection marker gene located between RRS1 and RRS2, and a second selection marker gene located between RRS3 and RRS4 in the region R.

5. The method for producing a cell according to claim 4, wherein the donor vector further comprises the following (8): (8) a third selection marker gene located between RRS5 and RRS6.

6. A method for producing the cell of claim 1, further comprising introducing an expression vector for said recombinase into said host cell.

7. The method of producing the cell of claim 1, wherein the recombinase is a serine recombinase.

8. The method of producing a cell according to claim 1, wherein the host cell is a mammalian cell.

9. The method of producing a cell according to claim 1, wherein the host cell is a CHO cell.

10. A method for producing a cell according to any one of claims 1 to 9, wherein the target gene is a gene encoding at least one selected from the group consisting of enzymes, antibodies, interleukins, cytokines, chemokines, hormones, growth factors, transcription factors, receptors, viral preparations, vaccines, medical proteins, subunits thereof, and fragments thereof.

11. A cell having a target gene incorporated into its genome, which is one of (A) to (C) below; (A) the genome has a region G including one each of site 1, site 2, site 3, and site 4, in that order, which are sites formed by recombination of recombinase recognition sites, (B) site 1 and site 4 have sequence identity, and site 2 and site 3 have sequence identity, and (C) region G has the target gene arranged between site 1 and site 2, and the target gene arranged between site 3 and site 4.

12. The cell of claim 11, further comprising the following (D): (D) the transcription direction of the target gene located between site 1 and site 2 is from site 2 to site 1, and the transcription direction of the target gene located between site 3 and site 4 is from site 3 to site 4.

13. The cell of claim 11, wherein the recombinase is a serine recombinase.

14. The cell of claim 11, wherein the cell is a mammalian cell.

15. The cell of claim 11, wherein the cell is a CHO cell.

16. The cell according to claim 11, wherein the target gene is a gene encoding at least one selected from the group consisting of an enzyme, an antibody, an interleukin, a cytokine, a chemokine, a hormone, a growth factor, a transcription factor, a receptor, a viral preparation, a vaccine, a medical protein, a subunit thereof, and a fragment thereof.

17. A method for producing a protein, comprising culturing the cell according to any one of claims 11 to 16 and expressing a protein encoded by the target gene.

Citation Information

Patent Citations

  • Site-specific integration

    EP2711428A1

  • Compositions and methods for making antibodies based on use of an expression-enhancing locus

    WO2017184831A1

  • Compositions and methods for making antibodies based on use of expression-enhancing loci

    WO2017184832A1

  • SSI cells with predictable and stable transgene expression and methods of formation

    WO2020072480A1

  • Methods and means for generating efficient silencing constructs using recombination cloning

    JP2004516852A