Search method, cell production method, cell selection method, cell, and method for producing cell product

By employing a search method to identify and target high-expression TADs in the genome, the efficiency of producing cell lines that stably express therapeutic proteins is enhanced, addressing the limitations of current methods.

WO2025115875A1PCT designated stage expired Publication Date: 2025-06-05FUJIFILM CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/041880
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-11-26
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Current methods for producing cells that stably express therapeutic proteins, such as humanized monoclonal antibodies, are inefficient due to the lack of precise targeting of high-expression regions in the genome.

Method used

A search method is developed to identify Topologically Associating Domains (TADs) with high transcriptional activity in the genome, allowing for the targeted insertion of a target gene into these regions to enhance gene expression.

Benefits of technology

This approach significantly increases the likelihood of producing cell lines that highly express the target gene, leading to improved productivity of therapeutic proteins.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024041880_05062025_PF_FP_ABST
    Figure JP2024041880_05062025_PF_FP_ABST
Patent Text Reader

Abstract

An embodiment of this search method comprises: calculating a TAD score indicative of the transcription activity for each TAD included in a genome; and selecting a TAD on the basis of the TAD score. Another embodiment of this search method comprises: analyzing a genome obtained from a cell expressing a foreign gene; and selecting a boundary of a region including the foreign gene. This cell production method comprises inserting a gene of interest into a region found by this search method. This cell selection method comprises: calculating a TAD score indicative of the transcription activity for each TAD included in a genome; selecting a TAD on the basis of the TAD score; and selecting a cell in which a target gene is present in the selected TAD. This cell is derived from a Chinese hamster, and has a target gene inserted in at least one prescribed region. This method for producing a cell product comprises culturing cells and expressing a target gene.
Need to check novelty before this filing date? Find Prior Art

Description

Search method, cell production method, cell selection method, and method for producing cells and cell products

[0001] The present disclosure relates to methods of discovery, methods of producing cells, methods of selecting cells, and methods of producing cells and cell products.

[0002] Patent document 1 discloses a site-specific integration host cell containing an endogenous Fer1L4 gene, wherein an exogenous nucleotide sequence has been integrated into the Fer1L4 gene. Patent document 2 discloses a cell comprising an exogenous nucleic acid integrated at a specific site within an expression-enhancing locus, wherein the exogenous nucleic acid sequence encodes a bispecific antigen-binding protein. Patent document 3 discloses a cell comprising a first exogenous nucleic acid integrated into a first expression-enhancing locus and a second exogenous nucleic acid integrated into a second expression-enhancing locus, wherein the first and second exogenous nucleic acids together encode an antigen-binding protein. Patent Document 4 discloses a mammalian cell comprising a first recombination target site (RTS) chromosomally integrated at a first high integration (HI) locus, wherein the first HI locus is within an active genomic compartment of accessible chromatin and within approximately 30,000 base pairs of a TAD boundary, and the first HI locus overlaps with a region of the cellular genome that interacts with at least one enhancer element.

[0003] European Patent Application Publication No. 2711428 International Publication No. WO 2017 / 184831 International Publication No. WO 2017 / 184832 International Publication No. WO 2020 / 072480

[0004] There is a technology for integrating a gene of interest into the genome of a host cell in order to create cells that stably produce therapeutic proteins such as humanized monoclonal antibodies. If the region of the genome where a gene is highly or stably expressed is known in advance, then by inserting the gene of interest into that region, it is possible to create a cell line that highly or stably expresses the gene of interest with a high probability.

[0005] The present disclosure has been made under the above circumstances. An objective of the present disclosure is to provide a search method for finding a target region on a genome into which a gene of interest is inserted and which highly expresses the gene of interest. An objective of the present disclosure is to provide a method for producing cells which highly express a gene of interest. An objective of the present disclosure is to provide a selection method for selecting cells which highly express a gene of interest. An objective of the present disclosure is to provide cells which highly express a gene of interest. An objective of the present disclosure is to provide a method for producing a cell product which is highly productive of a substance encoded by a gene of interest.

[0006] Specific means for solving the problems include the following aspects: <1> A search method for finding a target region on a genome into which a gene of interest is to be inserted, the search method including the following (1) for finding a TAD as the target region: (1) calculating a TAD score indicating transcriptional activity for each TAD in the genome, and selecting a TAD based on the TAD score; <2> The search method according to <1>, further including the following (2): (2) identifying a TAD score-SH, which is the TAD score of the TAD belonging to at least one known safe harbor in the genome, determining a threshold based on at least one TAD score-SH, and selecting a TAD based on the threshold and the TAD score. <3> A search method for finding a target region on a genome into which a gene of interest is to be inserted, the search method comprising the steps of (a) to (d) below, and finding a region specified by the following boundary pair P as the target region: (a) preparing a cell having a genome F in which a foreign gene has been incorporated and expressing the foreign gene; (b) obtaining genome F from the cell and analyzing genome F to identify region F, which is a region containing the foreign gene, and boundary F, which is the boundary between region F and the genome; (c) calculating a TAD score indicating transcriptional activity for each TAD in the genome, and specifying, for each boundary F, TAD score-F, which is the TAD score of the TAD to which it belongs, and selecting boundary F based on TAD score-F; and (d) finding boundary pair P that sandwiches region F from the selected boundaries F. <4> The searching method according to <3>, further comprising the following (p), and performing (c) for the boundary F selected by (p) below: (p) measuring the methylation rate of the foreign gene present in region F, and selecting a boundary F for region F with a low methylation rate. <5> The searching method according to <3> or <4>, further comprising the following (q), and performing (d) for the boundary F selected by (c) and (q) below: (q) identifying, for at least one known safe harbor in the genome, a TAD score-SH, which is the TAD score of the TAD belonging to the safe harbor, determining a threshold based on at least one TAD score-SH, and selecting a boundary F based on the threshold and the TAD score-F.<6> The searching method according to any one of <3> to <5>, further comprising the following (r): (r) selecting a boundary pair P such that a region F between the boundary pair P in the genome F is not a region formed by a chromosomal translocation. <7> The searching method according to any one of <1> to <6>, wherein the TAD score is a value obtained by multiplying the density of genes present in the TAD by their average expression levels. <8> The searching method according to any one of <1> to <7>, wherein the genome is a mammalian cell genome. <9> The searching method according to any one of <1> to <7>, wherein the genome is a CHO cell genome.

[0007] <10> A method for producing a cell, comprising: finding a TAD by the searching method according to <1> or <2>; and inserting a gene of interest into the TAD of a genome having the TAD. <11> A method for producing a cell, comprising: finding a region specified by boundary pair P by the searching method according to any one of <3> to <9>; and inserting a gene of interest within ±10 kbp before and after the region of a genome having the region. <12> The method for producing a cell according to <10> or <11>, wherein the gene of interest is a gene encoding at least one selected from the group consisting of an enzyme, an antibody, an interleukin, a cytokine, a chemokine, a hormone, a growth factor, a transcription factor, a receptor, a transcription regulatory nucleic acid, a non-coding RNA, a viral formulation, a vaccine, a medical protein, a subunit thereof, and a fragment thereof.

[0008] <13> A method for selecting cells expressing a target gene, the method comprising (11) below: (11) calculating a TAD score indicating transcriptional activity for each TAD in the genome of a cell, selecting TADs based on the TAD score, and selecting cells in which the target gene is present in the selected TAD. <14> The method for selecting cells according to <13>, further comprising (12) below: (12) identifying a TAD score-SH, which is the TAD score of a TAD belonging to at least one known safe harbor in the genome, determining a threshold based on at least one TAD score-SH, selecting TADs based on the threshold and the TAD score, and selecting cells in which the target gene is present in the selected TAD. <15> The method for selecting cells according to <13> or <14>, further comprising (13) below: (13) measuring the methylation rate of a target gene present in a selected TAD, and selecting cells with a low methylation rate. <16> The method for selecting cells according to any one of <13> to <15>, wherein the TAD score is a value obtained by multiplying the density of genes present in the TAD by their average expression levels. <17> The method for selecting cells according to any one of <13> to <16>, wherein the cells are mammalian cells. <18> The method for selecting cells according to any one of <13> to <16>, wherein the cells are CHO cells. <19> The method for selecting cells according to any one of <13> to <18>, wherein the target gene is a gene encoding at least one selected from the group consisting of enzymes, antibodies, interleukins, cytokines, chemokines, hormones, growth factors, transcription factors, receptors, transcription regulatory nucleic acids, non-coding RNAs, viral preparations, vaccines, medical proteins, subunits thereof, and fragments thereof.

[0009] <20> A cell derived from a Chinese hamster, wherein a gene of interest has been inserted into at least one region selected from the 40 regions shown in Table 1. <21> A cell derived from a Chinese hamster, wherein a gene of interest has been inserted into at least one region selected from the 40 regions shown in Table 2. <22> A cell derived from a Chinese hamster, wherein a gene of interest has been inserted into at least one region selected from the 40 regions shown in Table 3. <23> The cell according to any one of <20> to <22>, wherein the cell derived from a Chinese hamster is a CHO cell. <24> The cell according to any one of <20> to <23>, wherein the gene of interest is a gene encoding at least one selected from the group consisting of an enzyme, an antibody, an interleukin, a cytokine, a chemokine, a hormone, a growth factor, a transcription factor, a receptor, a transcription regulatory nucleic acid, a non-coding RNA, a viral preparation, a vaccine, a medical protein, a subunit thereof, and a fragment thereof.

[0010] <25> A method for producing a cell, comprising inserting a gene of interest into at least one region selected from the 40 regions shown in Table 1 of the genome of a cell derived from a Chinese hamster. <26> A method for producing a cell, comprising inserting a gene of interest into at least one region selected from the 40 regions shown in Table 2 of the genome of a cell derived from a Chinese hamster. <27> A method for producing a cell, comprising inserting a gene of interest into at least one region selected from the 40 regions shown in Table 3 of the genome of a cell derived from a Chinese hamster. <28> The method for producing a cell according to any one of <25> to <27>, wherein the cell derived from a Chinese hamster is a CHO cell. <29> The method for producing a cell according to any one of <25> to <28>, wherein the gene of interest is a gene encoding at least one selected from the group consisting of an enzyme, an antibody, an interleukin, a cytokine, a chemokine, a hormone, a growth factor, a transcription factor, a receptor, a transcription regulatory nucleic acid, a non-coding RNA, a viral preparation, a vaccine, a protein for medical use, a subunit thereof, and a fragment thereof.

[0011] <30> A method for producing a cell product, comprising culturing cells produced by the cell production method according to any one of <10> to <12> and <25> to <29>, and expressing the gene of interest. <31> A method for producing a cell product, comprising culturing cells selected by the cell selection method according to any one of <13> to <19>, and expressing the gene of interest. <32> A method for producing a cell product, comprising culturing the cells according to any one of <20> to <24>, and expressing the gene of interest. <33> The method for producing a cell product according to any one of <30> to <32>, wherein the gene of interest is a gene encoding at least one selected from the group consisting of an enzyme, an antibody, an interleukin, a cytokine, a chemokine, a hormone, a growth factor, a transcription factor, a receptor, a transcription regulatory nucleic acid, a non-coding RNA, a viral formulation, a vaccine, a medical protein, a subunit thereof, and a fragment thereof.

[0012] According to the present disclosure, there is provided a search method for finding a target region on a genome into which a gene of interest is inserted and which highly expresses the gene of interest. According to the present disclosure, there is provided a method for producing cells which highly express a gene of interest. According to the present disclosure, there is provided a selection method for selecting cells which highly express a gene of interest. According to the present disclosure, there is provided cells which highly express a gene of interest. According to the present disclosure, there is provided a method for producing a cell product which has excellent productivity for a substance encoded by a gene of interest.

[0013] 1 is a conceptual diagram of a second screening method. FIG. 2 is a scatter plot showing the antibody production performance of 60 antibody-producing strains created in the examples. FIG. 3 is a histogram of the copy number of foreign genes present within paired boundaries obtained from genome analysis of 31 antibody-producing strains. FIG. 4 is a histogram of the average methylation rates of promoter regions and antibody subunit regions obtained from genome analysis of 31 antibody-producing strains. FIG. 5 is an example of a contact map of CHO cells. FIG. 6 is a distribution diagram of TAD scores for 2,502 TADs detected in the genome of a CHO cell. FIG. 7 is a scatter plot showing the performance of high-performance regions discovered by this embodiment. FIG. 8 is a graph showing the performance of high-performance regions discovered by this embodiment. FIG. 9 is a graph showing that high-performance regions discovered by this embodiment can insert multiple copies of a coding sequence. FIG. 10 is a graph showing that high-performance regions discovered by this embodiment can stably express multiple copies of a coding sequence.

[0014] Hereinafter, embodiments of the present disclosure will be described. These descriptions and examples are intended to illustrate the embodiments and do not limit the scope of the embodiments. The mechanisms of action described in this disclosure include assumptions, and the validity of these assumptions does not limit the scope of the embodiments.

[0015] When describing embodiments of the present disclosure with reference to the drawings, the configuration of the embodiments of the present disclosure is not limited to the configuration shown in the drawings. The sizes of elements in the drawings are conceptual, and the relative relationships between the sizes of elements are not limited thereto.

[0016] In the present disclosure, the term "step" includes not only an independent step but also a step that cannot be clearly distinguished from other steps as long as the purpose of the step is achieved.

[0017] In the present disclosure, a numerical range indicated using "to" indicates a range that includes the numerical values ​​before and after "to" as the minimum and maximum values, respectively. In the numerical ranges described in stages in the present disclosure, the upper or lower limit value described in one numerical range may be replaced with the upper or lower limit value of another numerical range described in stages. Furthermore, in the numerical ranges described in the present disclosure, the upper or lower limit value of that numerical range may be replaced with a value shown in an example.

[0018] In the present disclosure, each component may contain multiple corresponding substances. When referring to the amount of each component in a composition in the present disclosure, if multiple substances corresponding to each component are present in the composition, the total amount of the multiple substances present in the composition is meant unless otherwise specified.

[0019] In this disclosure, the term "nucleic acid" includes all nucleic acids (e.g., deoxyribonucleic acid (DNA), ribonucleic acid (RNA), their analogs, natural products, and artificial products), as well as nucleic acids to which small molecules, groups (e.g., methyl groups), non-nucleic acid molecules, structures, etc. are linked. Nucleic acids may be single-stranded or double-stranded.

[0020] In the present disclosure, a donor vector is a substance that has the function of delivering exogenous nucleic acid into a cell and its genome, and is itself a nucleic acid. There are no limitations on the origin, form, or base sequence of the donor vector. The donor vector may be a circular nucleic acid or a linear nucleic acid. The donor vector may be a single-stranded nucleic acid or a double-stranded nucleic acid. The donor vector is preferably double-stranded DNA.

[0021] In the present disclosure, there is no limitation on the number of amino acid residues in a protein. Proteins include proteins in which amino acids have been post-translationally modified. Post-translational modifications of amino acids include phosphorylation, methylation, acetylation, glycosylation, lipidation, and the like. In the present disclosure, amino acids are represented by the three-letter and one-letter codes defined by IUPAC-IUBMB JCBN (IUPAC-IUBMB Joint Commission on Biochemical Nomenclature). Unless otherwise specified, amino acids referred to in the present disclosure are L-amino acids.

[0022] In the present disclosure, nucleotide sequence identity and amino acid sequence identity are calculated using BLAST (Basic Local Alignment Search Tool) (https: / / blast.ncbi.nlm.nih.gov / Blast.cgi).

[0023] A TAD (topologically associating domain, topology associating domain) is a structural unit in a genome detected by three-dimensional genome structure analysis, and is a region with a relatively high probability of spatial contact. The size of a TAD is generally several hundred kbp to several Mbp. The algorithm for three-dimensional genome structure analysis may be a publicly available algorithm, an improved version of a publicly available algorithm, or a newly developed algorithm. The genome structure data used for three-dimensional genome structure analysis may be data obtained from publicly available databases, academic papers, technical literature, etc., or may be data obtained by analyzing an actual intracellular genome. Examples of methods for three-dimensional genome structure analysis include the Hi-C method; the in situ Hi-C method, low Hi-C method, SAFE Hi-C method, Hi-CO method, and Micro-C method, which are based on the Hi-C method.

[0024] The Hi-C method is a three-dimensional genome structure analysis method that comprehensively detects spatially adjacent regions within a genome across the entire genome. Hi-C stands for high-throughput chromosome conformation capture. From the calculation results of Hi-C analysis, regions with a relatively high spatial contact probability, i.e., TADs, are detected, and the genome is partitioned into multiple TADs. The size of TADs detected by Hi-C analysis is generally several hundred kbp to several Mbps. Mammalian genomes are partitioned into thousands of TADs by Hi-C analysis. The Hi-C analysis algorithm used to detect TADs may be a publicly available algorithm, an improved version of a publicly available algorithm, or a newly developed algorithm. The genome structure data used in Hi-C analysis may be data obtained from publicly available databases, academic papers, technical literature, etc., or data obtained by analyzing actual intracellular genomes.

[0025] A safe harbor within a genome is a region where a host cell survives even if a gene is inserted, and where the inserted gene is expressed. A safe harbor within a genome is identified by a chromosome number or an accession number and a base number in a public base sequence database. Examples of public base sequence databases include the International Nucleotide Sequence Databases (INSD) and the NCBI Reference Sequence Database (RefSeq). A safe harbor within a genome may be referred to by the name of a known gene present within or near that region.

[0026] <Search Method> The present disclosure provides a search method for finding a target region on a genome for inserting a gene of interest. The present disclosure provides a first search method and a second search method.

[0027] <First Search Method> The target region found by the first search method is a TAD, which is a structural unit in the genome, and which has relatively high transcriptional activity.

[0028] There are TADs with relatively high transcriptional activity and TADs with relatively low transcriptional activity. The first screening method aims to find TADs with relatively high transcriptional activity and includes the following (1).

[0029] (1) Calculate a TAD score, which indicates transcriptional activity, for each TAD in the genome and select TADs based on the TAD score.

[0030] In the present disclosure, the TAD score is an evaluation criterion that indicates the level of transcriptional activity. A higher TAD score is usually a preferable evaluation criterion. Selecting a TAD based on the TAD score means, for example, selecting a TAD with a relatively high TAD score; or selecting a TAD with a TAD score that exceeds a predetermined standard.

[0031] The TAD score may be any index that identifies TADs with relatively high transcriptional activity from among the numerous TADs found in a genome (e.g., a mammalian genome is divided into several thousand TADs). The TAD score may be, for example, the total expression level, average value, or median value of the gene expression levels present in the TAD, the density of the genes present in the TAD, the reciprocal of the methylation rate or unmethylation rate of CpG sites within the TAD, a value obtained by multiplying two or more of these values, or an index based on two or more of these values.

[0032] The genes present in TADs can be known from publicly available databases, academic papers, technical literature, etc. The number of genes present in TADs involved in calculating the TAD score is at least one gene, preferably two or more genes, the more the better, and it is preferable to cover as many genes present in TADs as possible. The gene expression levels may be data obtained from publicly available databases, academic papers, technical literature, etc., or may be data obtained by actually quantifying the gene expression levels. The gene expression levels can be quantified by known mRNA quantification methods.

[0033] An example of a TAD score is the product of the density of genes present in a TAD and the average expression level. A higher value is preferable. The density of genes present in a TAD is the number of endogenous genes present in a TAD divided by the number of base pairs in the TAD (genes / Mb).

[0034] An example of an embodiment of the first search method further includes the following (2): (2) identifying a TAD score-SH, which is the TAD score of the TAD belonging to at least one known safe harbor in the genome, determining a threshold based on at least one TAD score-SH, and selecting a TAD based on the threshold and the TAD score.

[0035] Known safe harbors within a genome can be found in public databases, academic papers, technical literature, etc. The number of safe harbors identifying a TAD score-SH is at least one, preferably two or more, e.g., eight or less.

[0036] When there are many known safe harbors, at least one may be selected. Methods for selecting a safe harbor include, for example, selecting a safe harbor in which the expression level (pg / cell / copy) of the protein encoded by the inserted gene is relatively high; or selecting a safe harbor in which the expression level (pg / cell / copy) of the protein encoded by the inserted gene exceeds a predetermined standard. The protein expression level of the safe harbor may be data obtained from publicly available databases, academic papers, technical literature, etc., or may be data obtained by actually inserting a gene into the safe harbor and measuring the protein expression level.

[0037] The coordinates of the safe harbor (i.e., the chromosome number or the accession number and base number in a public base sequence database) are applied to the TAD section (the output of a three-dimensional genome structure analysis (e.g., Hi-C analysis)), the TAD to which the safe harbor belongs is identified, and then the TAD score of the TAD is identified. Here, the TAD score is the TAD score calculated for each TAD in (1). In the present disclosure, the TAD score of the TAD to which the safe harbor belongs is referred to as the "TAD score-SH." At least one TAD score-SH is identified.

[0038] The threshold based on the TAD score-SH is, for example, the minimum, maximum, mean or median value of at least one TAD score-SH.

[0039] A higher TAD score is usually a preferable evaluation criterion. Selecting a TAD based on a threshold and a TAD score means selecting a TAD with a TAD score above the threshold.

[0040] TADs selected based on a threshold based on TAD score-SH are predicted to exhibit gene expression levels that are comparable to or exceed known safe harbors.

[0041] Even if (1) and (2) cannot be clearly distinguished from each other, the first search method includes (1) and (2) as long as the objectives of (1) and (2) are achieved.

[0042] <Second Search Method> The target regions found by the second search method are regions within the genome that are safe harbors and have relatively high transcriptional activity.

[0043] The second search method includes the following steps (a) to (d) for the purpose of finding regions in the genome that are safe harbors and have relatively high transcriptional activity, and finds regions specified by boundary pair P. A region specified by boundary pair P is a region whose start and end points are boundary pair P.

[0044] (a) Preparing a cell having a genome F in which a foreign gene has been integrated and expressing the foreign gene. (b) Obtaining genome F from the cell and analyzing genome F to identify region F, which is a region containing the foreign gene, and boundary F, which is the boundary between region F and the genome. (c) Calculating a TAD score indicating transcriptional activity for each TAD in the genome, and for each boundary F, identifying TAD score-F, which is the TAD score of the TAD to which it belongs, and selecting boundary F based on TAD score-F. (d) Finding a boundary pair P that sandwiches region F from among the selected boundaries F.

[0045] Fig. 1 is a conceptual diagram of the second search method, which shows the relationship between the genome, foreign gene, cell, genome F, region F, boundary F, TAD, and boundary pair P, which are the targets of the second search method.

[0046] The foreign gene in (a) and (b) is a gene that is not originally present in the genome that is the target of the second search method, and is a gene that is focused on during genome analysis in order to find a safe harbor. The origin, type, size, and base sequence of the foreign gene are not limited. One example of an embodiment of the foreign gene is a structural gene (i.e., a gene that encodes a protein).

[0047] In the present disclosure, a foreign gene refers to a foreign gene that can be expressed in a cell. The foreign gene has all sequences necessary for the expression of the foreign gene. When the foreign gene is a structural gene, the foreign gene has all sequences necessary for the expression of the protein encoded by the foreign gene, including the protein coding sequence and all nucleic acids necessary for the transcription and translation of the coding sequence in the cell (e.g., promoter, transcription terminator, polyadenylation sequence). The foreign gene may contain one copy of the protein coding sequence, or two or more copies. For example, the foreign gene may contain at least one copy of the coding sequence for each subunit to express all subunits of a heteromultimeric protein. For example, the foreign gene may contain at least one copy each of the sequence encoding an antibody heavy chain and the sequence encoding an antibody light chain.

[0048] The cells in (a) may be either an established cell line or newly created cells. The cells may be either a polyclonal cell population or monoclonal cells. The cells in (a) do not die even when an exogenous gene is integrated into their genome and express the exogenous gene. Therefore, a safe harbor can be found by analyzing genome F obtained from the cells.

[0049] A preferred example of the cell in (a) is a monoclonal cell that stably and highly expresses a foreign gene. The region of the genome F of the cell into which the foreign gene is inserted is predicted to be a high-performance safe harbor (i.e., a region in which the host cell survives even after the gene is inserted and the inserted gene is stably and highly expressed). By analyzing the genome F obtained from the cell, a high-performance safe harbor can be efficiently identified.

[0050] Specifically, (b) includes, for example, genome extraction from cells, preparation of a sequencing library, library sequencing, mapping of reads onto the genome, mapping of the foreign gene onto the reads, identification of the boundary between the genome and the foreign gene, extraction of reads containing the foreign gene, measurement of the copy number of the foreign gene contained in the reads, and identification of boundary pairs sandwiching the foreign gene. From the viewpoint of obtaining long reads and DNA methylation data, single-molecule real-time sequencing or nanopore sequencing is preferred for library sequencing.

[0051] Region F is a region within genome F that is distinct from the genome and contains at least one copy of a foreign gene. Boundary F is the boundary between region F and the genome in genome F. The coordinates of boundary F are specified by the chromosome number of the genome to be searched or the accession number and base number of a public base sequence database. Examples of public base sequence databases include INSD (the International Nucleotide Sequence Databases) and RefSeq (NCBI Reference Sequence Database).

[0052] An example of the embodiment of (b) includes narrowing down region F to a region containing multiple copies (i.e., two or more copies) of the foreign gene. In this case, boundary F of region F containing multiple copies of the foreign gene is subject to (c). The number of copies of the foreign gene contained in region F may be, for example, two to six copies, two to five copies, or two to four copies. Region F containing multiple copies of the foreign gene is multicopy-tolerant, allowing stable production of the target product even when multiple copies of a target gene with a relatively long base length are inserted. In other words, a high production yield of the target product is maintained even after long-term culture. Conventionally, when multiple copies of a foreign gene are inserted at a single site, the production yield of the substance encoded by the foreign gene tends to be unstable, but this problem is improved according to the embodiments of the present disclosure. Examples of target genes with relatively long base lengths and multiple copies include nucleic acids in which two or more copies of a coding sequence are linked, and nucleic acids in which at least one copy of the coding sequence for each subunit of a heteromultimeric protein is linked. A region into which a nucleic acid containing two or more linked copies of a coding sequence can be inserted is a region that can enable high production of a substance encoded by a gene of interest.A region into which a nucleic acid containing at least one linked copy of the coding sequences for all subunits of a heteromultimeric protein can be inserted is a region that can enable stable production of a heteromultimeric protein.

[0053] The TAD and TAD score in (c) are synonymous with the TAD and TAD score in (1) of the first search method, and the specific form and calculation method are also the same. An example of an embodiment of the TAD score is the value obtained by multiplying the density of genes present in the TAD by the average expression level. The higher this value, the more preferable it is as an evaluation criterion.

[0054] The coordinates of boundary F (i.e., the chromosome number or the accession number and base number in a public base sequence database) are applied to the TAD section (which is the output of a three-dimensional genome structure analysis (e.g., Hi-C analysis)) to identify the TAD to which boundary F belongs, and then the TAD score of that TAD is determined. Here, the TAD score is the TAD score calculated for each TAD. In the present disclosure, the TAD score of the TAD to which boundary F belongs is referred to as "TAD score-F."

[0055] A higher TAD score is usually a preferable evaluation criterion. Selecting the boundary F based on the TAD score-F means, for example, selecting the boundary F where the TAD score-F is relatively high, or selecting the boundary F where the TAD score-F exceeds a predetermined standard.

[0056] (d) is to find a pair of boundaries P that sandwich a region F from among the boundaries F selected by (c). The pair of boundaries P is one end and the other end of the same region F in the genome F.

[0057] Since boundary pair P is a pair of boundary F selected by (c), it is a coordinate pair that exists in a TAD with relatively high transcriptional activity. Therefore, the region identified by boundary pair P (i.e., the region with boundary pair P as its start and end points) is expected to be a region with relatively high transcriptional activity. The coordinate pair of boundary pair P is identified by the chromosome number of the genome to be searched or the accession number and base number of a public base sequence database. Examples of public base sequence databases include the International Nucleotide Sequence Databases (INSD) and NCBI Reference Sequence Database (RefSeq).

[0058] In (b), when region F is narrowed down to a region containing multiple copies of a foreign gene, the region specified by boundary pair P can stably produce the target product even when a relatively long target gene (e.g., a nucleic acid in which two or more copies of a coding sequence are linked, or a nucleic acid in which at least one copy of the coding sequence for all subunits of a heteromultimeric protein is linked) is inserted. In other words, a high production amount of the target product can be maintained even after long-term culture.

[0059] An example of an embodiment of the second searching method further includes the following (p), and performs (c) for the boundary F selected by (p): (p) measuring the methylation rate of the foreign gene present in region F, and selecting boundary F of region F with a low methylation rate.

[0060] DNA methylation generally represses gene expression. Therefore, (p) is to select a region F where gene expression is not repressed and its boundary F. By performing (c) on the boundary F selected by (p) and then performing (d), a boundary pair P relating to a region where high gene expression can be expected is found.

[0061] The methylation rate of the exogenous gene can be obtained from data obtained by performing single-molecule real-time sequencing or nanopore sequencing as the library sequencing in (b). The methylation rate value is, for example, methylated cytosines at CpG sites / total cytosines at CpG sites × 100. The methylation rate value may be the value for the entire exogenous gene, the value for a portion of the exogenous gene (e.g., the value for the protein-coding sequence), or the value for each range obtained by dividing the exogenous gene by function (e.g., the value for each of the promoter region and the protein-coding sequence). The minimum, maximum, average, or median of at least one methylation rate obtained from the target range is used as the representative value to select region F and its boundary F.

[0062] Selecting a region F with a low methylation rate means, for example, selecting a region F with a relatively low methylation rate; or selecting a region F with a methylation rate below a predetermined standard (e.g., 30%, 20%, 10%).

[0063] An example of an embodiment of the second search method further includes the following (q), and performs (d) for the boundary F selected by (c) and (q): (q) identifying a TAD score-SH, which is the TAD score of the TAD belonging to at least one known safe harbor in the genome, determining a threshold based on at least one TAD score-SH, and selecting a boundary F based on the threshold and the TAD score-F.

[0064] The known safe harbors, TAD score-SH, and thresholds in the genome in (q) are synonymous with the known safe harbors, TAD score-SH, and thresholds in the genome in (2) of the first search method, and the specific forms and methods of identification are also the same.

[0065] A higher TAD score is usually a preferable evaluation criterion. Selecting a boundary F based on a threshold and a TAD score-F means selecting a boundary F that has a TAD score-F that is greater than the threshold.

[0066] (c) and (q) are performed to select a boundary F, and then (d) is performed to find a boundary pair P for regions predicted to show gene expression levels equivalent to or exceeding known safe harbors.

[0067] Even if (c) and (q) cannot be clearly distinguished from each other, the second search method includes (c) and (q) as long as the objectives of (c) and (q) are achieved.

[0068] An example of an embodiment of the second searching method further includes (r) selecting a boundary pair P such that a region F between the boundary pair P in the genome F is not a region formed by a chromosomal translocation.

[0069] To confirm that region F is not a region formed by chromosomal translocation, for example, one coordinate and the other coordinate of a boundary pair P sandwiching region F belong to the same chromosome, and the base length of region F (in other words, the distance between boundary pair P on genome F) is not too long. For example, if the base length of region F is 100 kbp or less, region F is determined to not be a region formed by chromosomal translocation, and a boundary pair P sandwiching this region F is selected.

[0070] Even if (d) and (r) cannot be clearly distinguished from each other, the second search method includes (d) and (r) as long as the objectives of (d) and (r) are achieved.

[0071] The genomes targeted by the first search method are genomes of any cells, as long as they contain TADs, which are structural units within the genome. Examples of cells include fungi, yeast, insect cells, mammalian cells, and plant cells.

[0072] The genomes targeted by the second search method are genomes of any cell, as long as they contain TADs, which are structural units within the genome, and it is possible to integrate a gene into the genome and create cells that express the gene. Examples of cells include fungi, yeast, insect cells, mammalian cells, and plant cells.

[0073] An example of a fungus is Aspergillus oryzae.

[0074] Examples of yeast include Saccharomyces cerevisiae, Pichia pastoris, and Hansenula polymorpha.

[0075] Examples of insect cells include BmN cells derived from the silkworm (Bombyx mori), Sf9 cells and Sf21 cells derived from the armyworm (Spodoptera frugiperda), S2 cells derived from the fruit fly (Drosophila melanogaster), and Pv11 cells derived from the sleeping chironomid (Polypedilum vanderplanki).

[0076] Examples of mammalian cells include Chinese hamster ovary cells (CHO cells), baby hamster kidney cells (BHK cells), human embryonic kidney cell lines (e.g., HEK293 cells), human retinoblastoma-derived cell lines (e.g., PER.C6 cells), mouse myeloma cell lines (e.g., NS0 cells and SP2 / 0 cells), and established cell lines derived from these cells.

[0077] Examples of CHO cells include CHO-DG44 cells, CHO-K1 cells, CHO-DXB11 cells, and CHOpro3 cells. - Cells and cell lines derived from these cells.

[0078] Examples of mammalian cells include cells that have the ability to differentiate into other cells, such as pluripotent stem cells, such as embryonic stem cells (ES cells) and induced pluripotent stem cells (iPS cells), and multipotent stem cells, such as mesenchymal stem cells, tissue stem cells, and somatic stem cells.

[0079] <Cell Production Method> The present disclosure provides a method for producing a cell that highly expresses a gene of interest. The present disclosure provides a cell production method that includes a first screening method and a cell production method that includes a second screening method.

[0080] The cell production method including the first screening method includes finding a TAD by the first screening method and inserting a target gene into the TAD of a genome having the TAD.

[0081] In the cell production method including the first screening method, the insertion region of the target gene is located within a TAD selected based on the TAD score indicating transcriptional activity, i.e., within a TAD with relatively high transcriptional activity. Therefore, the produced cells are likely to highly express the target gene.

[0082] A cell production method including the second search method includes finding a region identified by boundary pair P using the second search method, and inserting a target gene within ±10 kbp before and after the region of a genome having the region.

[0083] In the cell production method including the second search method, the insertion region of the target gene may be within ±10 kbp before and after the region specified by boundary pair P (i.e., the region starting and ending at boundary pair P), or it may be outside the region starting and ending at boundary pair P, or it may straddle the inside and outside of the region starting and ending at boundary pair P. Because the insertion region of the target gene is within or near the safe harbor found by the second search method, the produced cells are likely to highly express the target gene without dying.

[0084] The cell production method including the second screening method may limit the insertion region of the gene of interest to a narrower range. Examples of the insertion region of the gene of interest include within ±8 kbp before and after the region specified by boundary pair P; within ±5 kbp before and after the region specified by boundary pair P; within ±3 kbp before and after the region specified by boundary pair P; within ±1 kbp before and after the region specified by boundary pair P; and within the region specified by boundary pair P.

[0085] The genome into which the target gene of interest is inserted can be the genome of any cell, as long as it has TADs, which are structural units within the genome. Examples of cells include fungi, yeast, insect cells, mammalian cells, and plant cells. Specific examples of cells are the same as those given in the description of the search method of the present disclosure.

[0086] In one embodiment of the cell production method of the present disclosure, a gene of interest is inserted into the genome of a mammalian cell, such as Chinese hamster ovary cells (CHO cells), baby hamster kidney cells (BHK cells), human embryonic kidney cell lines (e.g., HEK293 cells), human retinoblastoma-derived cell lines (e.g., PER.C6 cells), mouse myeloma cell lines (e.g., NS0 cells and SP2 / 0 cells), and established cell lines derived from these cells.

[0087] In one embodiment of the cell production method of the present disclosure, a gene of interest is inserted into the genome of a CHO cell. Examples of CHO cells include CHO-DG44 cells, CHO-K1 cells, CHO-DXB11 cells, and CHOpro3 cells. -Cells and cell lines derived from these cells.

[0088] In one embodiment of the cell production method of the present disclosure, a gene of interest is inserted into the genome of a cell that has the ability to differentiate into other cells, such as pluripotent stem cells (ES cells, iPS cells, etc.) and multipotent stem cells (mesenchymal stem cells, tissue stem cells, somatic stem cells, etc.).

[0089] Insertion of a gene of interest into a target region of the genome is possible using known genome editing techniques.

[0090] The origin, size, and base sequence of the target gene are not limited. Target genes include nucleic acids that encode proteins and nucleic acids that do not encode proteins. Examples of target genes include genes that encode at least one selected from the group consisting of enzymes, antibodies, interleukins, cytokines, chemokines, hormones, growth factors, transcription factors, receptors, transcription regulatory nucleic acids, non-coding RNAs, viral preparations, vaccines, medical proteins, subunits thereof, and fragments thereof.

[0091] When the target gene is a nucleic acid encoding a protein, examples of the protein encoded by the target gene (referred to as "target protein" in the present disclosure) include at least one selected from the group consisting of enzymes, antibodies, interleukins, cytokines, chemokines, hormones, growth factors, transcription factors, receptors, proteins constituting viral preparations, vaccines, medical proteins, subunits thereof, and fragments thereof.

[0092] In the present disclosure, an antibody is not limited to an immunoglobulin, but may be any molecule that binds to an antigen. In the present disclosure, the term antibody includes antibody fragments and antigen-binding molecules. In the present disclosure, the heavy chain of an antibody is also referred to as an H chain, and the light chain of an antibody is also referred to as an L chain.

[0093] Examples of nucleic acids that do not encode proteins include transcriptional regulatory nucleic acids, non-coding RNA, and nucleic acids that constitute viral formulations. Non-coding RNA (ncRNA) includes microRNA (miRNA), short hairpin RNA (shRNA), small interfering RNA (siRNA), small nuclear RNA (snRNA), ribosomal RNA (rRNA), and transfer RNA (tRNA).

[0094] In the present disclosure, a gene of interest refers to a gene of interest that can be expressed in a cell. The gene of interest has all sequences necessary for the expression of the gene of interest. When the gene of interest is a structural gene, the gene of interest has all sequences necessary for the expression of the protein encoded by the gene of interest, including the protein coding sequence and all nucleic acids necessary for the transcription and translation of the coding sequence in the cell (e.g., promoter, transcription terminator, polyadenylation sequence). The gene of interest may contain one copy of the protein coding sequence, or two or more copies. For example, the gene of interest may contain at least one copy of the coding sequence for each subunit to express all subunits of a heteromultimeric protein. For example, the gene of interest may contain at least one copy each of a sequence encoding an antibody heavy chain and a sequence encoding an antibody light chain.

[0095] In one embodiment of the cell production method disclosed herein, a nucleic acid comprising two or more copies of a coding sequence linked together is inserted as a gene of interest into a target region of the genome. This embodiment produces cells capable of high production of a substance encoded by the gene of interest. From the viewpoint of long-term cell passage stability, the number of copies of the coding sequence contained in the gene of interest is preferably not too high, preferably 2 to 6 copies, more preferably 2 to 5 copies, and even more preferably 2 to 4 copies.

[0096] In one embodiment of the cell production method of the present disclosure, a nucleic acid comprising at least one copy of the coding sequences for all subunits of a heteromultimeric protein linked together is inserted as a gene of interest into a target region of the genome. This embodiment produces cells capable of stably producing a heteromultimeric protein. In this embodiment, the gene of interest is preferably a nucleic acid comprising two or more copies of the set of coding sequences for all subunits linked together, from the viewpoint of increasing the expression level of the heteromultimeric protein. The copy number of the set contained in the gene of interest is preferably not too high, from the viewpoint of long-term cell passage stability, and is preferably 2 to 6 copies, more preferably 2 to 5 copies, and even more preferably 2 to 4 copies.

[0097] <Cell Selection Method> The present disclosure provides a selection method for selecting cells that highly express a target gene.

[0098] The origin, size, and base sequence of the target gene are not limited. Target genes include nucleic acids that encode proteins and nucleic acids that do not encode proteins. Examples of target genes include genes that encode at least one selected from the group consisting of enzymes, antibodies, interleukins, cytokines, chemokines, hormones, growth factors, transcription factors, receptors, transcription regulatory nucleic acids, non-coding RNA, viral preparations, vaccines, medical proteins, subunits thereof, and fragments thereof. The meaning and specific form of the target gene are the same as those of the target gene described in the cell production method of the present disclosure.

[0099] The cell sorting method of the present disclosure includes the following (11): (11) calculating a TAD score indicating transcriptional activity for each TAD in the genome of a cell, selecting TADs based on the TAD score, and selecting cells in which a target gene is present in the selected TAD.

[0100] The TAD and TAD score in (11) are synonymous with the TAD and TAD score in (1) of the first search method, and the specific form and calculation method are also the same. An example of an embodiment of the TAD score is the value obtained by multiplying the density of genes present in the TAD by the average expression level. The higher this value, the more preferable it is as an evaluation criterion.

[0101] The presence of the target gene in the TAD can be confirmed, for example, by extracting the genome from the cell, preparing a sequencing library, performing library sequencing, mapping the reads onto the genome, mapping the target gene onto the reads, extracting reads containing the target gene, and measuring the copy number of the target gene contained in the reads. From the viewpoints of obtaining long reads and DNA methylation data, single-molecule real-time sequencing or nanopore sequencing is preferred for library sequencing.

[0102] The cells selected by (11) are expected to be cells that highly express the target gene because the target gene is present in a TAD with relatively high transcriptional activity.

[0103] An example of the embodiment of (11) includes selecting cells in which multiple copies of the target gene (i.e., two or more copies of the coding sequence) are present in the selected TAD. These cells are likely to express the target gene at a higher level. From the viewpoint of long-term subculture stability of the cells, it is preferable that the number of copies of the target gene present in the selected TAD is not too high, preferably 2 to 6 copies, more preferably 2 to 5 copies, and even more preferably 2 to 4 copies.

[0104] An example of an embodiment of the cell sorting method of the present disclosure further includes the following (12): (12) identifying a TAD score-SH, which is the TAD score of the TAD belonging to at least one known safe harbor in the genome, determining a threshold based on at least one TAD score-SH, selecting a TAD based on the threshold and the TAD score, and selecting cells in which the target gene is present in the selected TAD.

[0105] The known safe harbors, TAD score-SH, and thresholds in the genome in (12) are synonymous with the known safe harbors, TAD score-SH, and thresholds in the genome in (2) of the first search method, and the specific forms and methods of identification are also the same.

[0106] Cells selected by (12) are expected to exhibit gene expression levels comparable to or exceeding those of cells in which the gene of interest has been inserted into a known safe harbor.

[0107] Even if (11) and (12) cannot be clearly distinguished from each other, the cell sorting method of the present disclosure includes (11) and (12) as long as the objectives of (11) and (12) are achieved.

[0108] An example of an embodiment of the cell sorting method of the present disclosure further includes the following (13): (13) measuring the methylation rate of a target gene present in the selected TAD, and selecting cells with a low methylation rate.

[0109] DNA methylation generally suppresses gene expression, so by performing (13), cells that highly express the target gene are selected.

[0110] The methylation rate of the target gene can be obtained from data obtained by performing single-molecule real-time sequencing or nanopore sequencing as the library sequencing in (11). The methylation rate value is, for example, methylated cytosines at CpG sites / total cytosines at CpG sites × 100. The methylation rate value may be the value for the entire target gene, the value for a portion of the target gene (e.g., the value for the protein-coding sequence), or the value for each range obtained by dividing the target gene by function (e.g., the values ​​for the promoter region and the protein-coding sequence). The minimum, maximum, average, or median of at least one methylation rate obtained from the target range is used as the representative value to be used for cell selection.

[0111] Selecting cells with a low methylation rate of a target gene means, for example, selecting cells with a relatively low methylation rate; or selecting cells with a methylation rate below a predetermined standard (e.g., 30%, 20%, 10%).

[0112] The cells that can be subjected to the cell sorting method of the present disclosure are any cells as long as they have a genome comprising the structural unit TAD. Examples of cells include fungi, yeast, insect cells, mammalian cells, and plant cells. Specific examples of cells are the same as those listed in the description of the screening method of the present disclosure.

[0113] An example of an embodiment of the cell sorting method of the present disclosure targets mammalian cells, including Chinese hamster ovary cells (CHO cells), baby hamster kidney cells (BHK cells), human embryonic kidney cell lines (e.g., HEK293 cells), human retinoblastoma-derived cell lines (e.g., PER.C6 cells), mouse myeloma cell lines (e.g., NS0 cells and SP2 / 0 cells), and established cell lines derived from these cells.

[0114] An embodiment of the cell sorting method of the present disclosure targets CHO cells. Examples of CHO cells include CHO-DG44 cells, CHO-K1 cells, CHO-DXB11 cells, and CHOpro3 cells. - Cells and cell lines derived from these cells.

[0115] An example of an embodiment of the cell selection method of the present disclosure targets cells differentiated from mammalian cells having differentiation potential, for example, cells differentiated by introducing a target gene into pluripotent stem cells (ES cells, iPS cells, etc.) or multipotent stem cells (mesenchymal stem cells, tissue stem cells, somatic stem cells, etc.).

[0116] <Cells> The present disclosure provides cells that have an exogenous gene of interest integrated into their genome and that highly express the gene of interest.

[0117] The cells of the present disclosure are derived from Chinese hamsters. Examples of cells derived from Chinese hamsters include fibroblasts, adipocytes, adipose-derived stem cells, bone marrow-derived stem cells, ovarian cells, and cell lines derived from these cells.

[0118] An example of an embodiment of a cell derived from a Chinese hamster is a Chinese hamster ovary cell (CHO cell). Examples of CHO cells include CHO-DG44 cells, CHO-K1 cells, CHO-DXB11 cells, and CHOpro3 cells. - Cells and cell lines derived from these cells.

[0119] The cells of the present disclosure are cells in which a gene of interest has been inserted into at least one region selected from the 40 regions shown in Table 1 below in the genome of a cell derived from a Chinese hamster. The 40 regions shown in Table 1 are regions identified by the Chinese hamster RefSeq accession number and base number, i.e., "RefSeq ID," "start," and "end" in Table 1.

[0120]

[0121] Table 1 also shows the TAD to which each region belongs. The TAD here refers to the TAD detected in the Examples described below. "TAD start" and "TAD end" in Table 1 are the boundary coordinates of the TAD obtained in the Examples described below. The meanings of "TAD start" and "TAD end" in Tables 2 and 3 are the same.

[0122] The 40 regions shown in Table 1 are regions consisting of regions discovered by the second screening method of the present disclosure and their neighboring regions, and are regions where genes can be highly expressed. Therefore, a cell having a genome in which a gene of interest has been inserted into at least one region selected from the 40 regions shown in Table 1 is a cell that can highly express the gene of interest.

[0123] A preferred example of the cell of the present disclosure is a cell in which a gene of interest has been inserted into at least one region selected from the group consisting of 40 regions in the genome of a cell derived from a Chinese hamster shown in Table 2 below and regions having 90% or more sequence identity with any of the 40 regions. The 40 regions shown in Table 2 are regions identified by the Chinese hamster RefSeq accession number and base number, i.e., "RefSeq ID," "start," and "end" in Table 2.

[0124]

[0125] The 40 regions shown in Table 2 are regions within the 40 regions shown in Table 1. The 40 regions shown in Table 2 are regions consisting of regions discovered by the second screening method of the present disclosure and their neighboring regions, and are regions in which genes can be highly expressed. Therefore, a cell having a genome in which a gene of interest has been inserted into at least one region selected from the 40 regions shown in Table 2 is a cell that can highly express the gene of interest.

[0126] A more preferred example of the cell of the present disclosure is a cell in which a gene of interest has been inserted into at least one region selected from the group consisting of 40 regions in the genome of a cell derived from a Chinese hamster shown in Table 3 below and regions having 90% or more sequence identity with any of the 40 regions. The 40 regions shown in Table 3 are regions identified by the Chinese hamster RefSeq accession number and base number, i.e., "RefSeq ID," "start," and "end" in Table 3.

[0127]

[0128] The 40 regions shown in Table 3 are regions within the 40 regions shown in Table 2. The 40 regions shown in Table 3 are regions discovered by the second screening method of the present disclosure, and are safe harbor regions where genes are highly expressed. Therefore, a cell having a genome in which a gene of interest has been inserted into at least one region selected from the 40 regions shown in Table 3 is a cell that can stably and highly express the gene of interest.

[0129] The origin, size, and base sequence of the target gene are not limited. Target genes include nucleic acids that encode proteins and nucleic acids that do not encode proteins. Examples of target genes include genes that encode at least one selected from the group consisting of enzymes, antibodies, interleukins, cytokines, chemokines, hormones, growth factors, transcription factors, receptors, transcription regulatory nucleic acids, non-coding RNAs, viral preparations, vaccines, medical proteins, subunits thereof, and fragments thereof.

[0130] When the target gene is a nucleic acid encoding a protein, examples of the target protein include at least one selected from the group consisting of enzymes, antibodies, interleukins, cytokines, chemokines, hormones, growth factors, transcription factors, receptors, proteins constituting viral preparations, vaccines, medical proteins, subunits thereof, and fragments thereof.

[0131] Examples of nucleic acids that do not encode proteins include transcriptional regulatory nucleic acids, non-coding RNA, and nucleic acids that constitute viral formulations. Non-coding RNA (ncRNA) includes microRNA (miRNA), short hairpin RNA (shRNA), small interfering RNA (siRNA), small nuclear RNA (snRNA), ribosomal RNA (rRNA), and transfer RNA (tRNA).

[0132] In the present disclosure, a gene of interest refers to a gene of interest that can be expressed in a cell. The gene of interest has all sequences necessary for the expression of the gene of interest. When the gene of interest is a structural gene, the gene of interest has all sequences necessary for the expression of the protein encoded by the gene of interest, including the protein coding sequence and all nucleic acids necessary for the transcription and translation of the coding sequence in the cell (e.g., promoter, transcription terminator, polyadenylation sequence). The gene of interest may contain one copy of the protein coding sequence, or two or more copies. For example, the gene of interest may contain at least one copy of the coding sequence for each subunit to express all subunits of a heteromultimeric protein. For example, the gene of interest may contain at least one copy each of a sequence encoding an antibody heavy chain and a sequence encoding an antibody light chain.

[0133] One embodiment of the cell of the present disclosure has a genome in which a nucleic acid comprising two or more copies of a coding sequence linked together is inserted as a gene of interest. The cell of this embodiment can produce a high level of a substance encoded by the gene of interest. From the viewpoint of long-term cell passage stability, it is preferable that the number of copies of the coding sequence contained in the gene of interest is not too high, preferably 2 to 6 copies, more preferably 2 to 5 copies, and even more preferably 2 to 4 copies.

[0134] An example of an embodiment of the cell of the present disclosure has a genome in which a nucleic acid in which at least one copy of the coding sequences for all subunits of a heteromultimeric protein are linked together is inserted as a gene of interest. The cell of this embodiment can stably produce a heteromultimeric protein. In this embodiment, the gene of interest is preferably a nucleic acid in which two or more copies of the coding sequence set for all subunits are linked together, from the viewpoint of increasing the expression level of the heteromultimeric protein. The copy number of the set contained in the gene of interest is preferably not too high, from the viewpoint of long-term passage stability of the cell, and is preferably 2 to 6 copies, more preferably 2 to 5 copies, and even more preferably 2 to 4 copies.

[0135] The present disclosure provides a method for producing a Chinese hamster-derived cell that highly expresses a gene of interest. The cell production method includes inserting the gene of interest into at least one region selected from the 40 regions listed in Table 1 of the genome of the Chinese hamster-derived cell. A preferred example of the cell production method includes inserting the gene of interest into at least one region selected from the 40 regions listed in Table 2 of the genome of the Chinese hamster-derived cell. A more preferred example of the cell production method includes inserting the gene of interest into at least one region selected from the 40 regions listed in Table 3 of the genome of the Chinese hamster-derived cell. Inserting the gene of interest into a target region of the genome can be achieved using known genome editing techniques.

[0136] <Method for producing a cell product> The present disclosure provides a method for producing a cell product with excellent productivity. The method for producing a cell product of the present disclosure uses cells that highly express a gene of interest, and thereby achieves excellent productivity of a substance encoded by the gene of interest (referred to as a "substance of interest" in the present disclosure). The cells that highly express a gene of interest are at least one type selected from the group consisting of cells produced by the cell production method of the present disclosure, cells selected by the cell selection method of the present disclosure, and cells of the present disclosure.

[0137] The method for producing a cell product disclosed herein involves culturing cells, expressing a gene of interest, and producing a substance of interest. By culturing the cells, the substance of interest is produced within the cells, and the substance of interest accumulates in the culture medium and / or the cells.

[0138] The cell culture method and medium composition may be selected depending on the type of cell. Culture conditions (e.g., culture scale, cell density, temperature, and CO 2 The concentration may also be selected depending on the type of cells.

[0139] An example of an embodiment of the method for producing a cell product of the present disclosure includes recovering a target substance from the culture medium. Methods for recovering the target substance from the culture medium include, for example, centrifugation, filtration, diafiltration, ion exchange chromatography, affinity chromatography, hydrophobic interaction chromatography, gel filtration chromatography, and high performance liquid chromatography (HPLC). The recovered target substance can be used, for example, in the production of a pharmaceutical composition.

[0140] An example of an embodiment of the method for producing a cell product of the present disclosure includes recovering cells in which a target substance has accumulated from a culture medium. Methods for recovering cells from a culture medium include, for example, centrifugation and filtration. The target substance accumulates inside or on the surface of the cells depending on its properties. The recovered cells are, for example, administered, transfused, or transplanted into a mammal.

[0141] The search method and the like of the present disclosure will be explained in more detail below using specific examples. The materials, processing procedures, and the like shown in the following specific examples can be changed as appropriate without departing from the spirit of the present disclosure. The scope of the search method and the like of the present disclosure should not be construed as being limited by the specific examples shown below.

[0142] <Creation of antibody-producing strains> [Method] A plasmid carrying genes encoding known IgG (one H-chain gene and one L-chain gene, for a total of two antibody subunit genes) was constructed based on an artificial plasmid containing the dihydrofolate reductase (DHFR) gene. The H-chain gene and L-chain gene each contain the hEF-1α promoter, a coding sequence for a fibronectin secretory leader, a coding sequence for the antibody subunit, and a polyA sequence. The fibronectin secretory leader is a signal peptide that induces extracellular secretion of polypeptides. The order of the two antibody subunit genes on the plasmid is L-chain-H-chain. Hereinafter, this gene group (L-chain-H-chain) will be referred to as "Gol."

[0143] The plasmid was linearized and introduced into CHO-DG44 cells by electroporation. The cells were statically cultured in a medium containing methotrexate (MTX) for 14 to 21 days to establish an MTX-resistant cell pool. One cell was seeded per well in a 96-well plate and incubated at 37°C and CO 2 The cells were statically cultured in an atmosphere containing 10% (v / v) antibody. On day 14 of culture, the culture supernatant was collected, and the antibody concentration was measured using an Octet Qke molecular interaction analyzer (Sartorius). Clones with the highest antibody concentrations were selected and cultured in 24-well plates, and then cultured in bioreactor tubes for scale-up. 60 strains of cells with the highest antibody concentrations were selected. To confirm the long-term passage stability of the cells, the cells were suspended in a passage medium and passaged every three days. Passage was terminated when the total number of cell divisions exceeded 60.

[0144] <Fed-batch culture test> [Method] Sixty antibody-producing strains before and after long-term passage were each suspended in 40 mL of basal medium, transferred to a 125 mL flask for shaking culture, and incubated at 37°C and CO 2Shaking culture and feeding were performed at 140 rpm in a 5% (v / v) atmosphere. A fixed amount of feed medium was added daily from day 3 to day 13 of culture. Sampling was performed every 1 to 3 days to measure cell density, culture medium components, and antibody concentration. Measurements were performed using a Vi-CELL XR cell counter (Beckman Coulter), a FLEX2 cell culture environment analyzer (Nova Biomedical), and a CedexBio product analyzer (Roche Diagnostics). On day 14 of culture, the culture medium was collected, and cells and cell debris were removed using a depth filter (pore size 0.22 μm) to obtain the culture supernatant. The antibody concentration in the culture supernatant was measured by liquid chromatography using a protein A column.

[0145] [Results] Figure 2 shows the antibody production performance of 60 antibody-producing strains. In the scatter plot in Figure 2, the horizontal axis represents the change in specific production rate before and after long-term passaging, and the vertical axis represents the relative specific production rate. From the 60 antibody-producing strains, 31 strains were selected that showed a change on the horizontal axis of -20% or more.

[0146] <Measurement of GoI insertion position> [Method] Thirty-one antibody-producing strains were cultured with shaking for two weeks. The medium used was a liquid medium containing serum-free basal medium (CD OptiCHO Medium, model number 12681-011, Thermo Fisher Scientific) supplemented with L-glutamic acid and methotrexate. 4 × 10 cells were collected from each cell culture. 6Cells were harvested and genomes were extracted using a long-chain genome extraction kit. The genome length was adjusted and enriched. A library was created from the enriched genome using a long-read sequencing kit (Ligation Sequencing Kit, model number SQK-LSK109, Oxford Nanopore Technologies). The DNA sequences in the library were read using a nanopore sequencer (model number M1CCapEx, Oxford Nanopore Technologies). The fast5 file output from the sequencer was base-called to obtain a fastq file and corresponding fasta file. From all reads, reads containing H-chain and L-chain genes were extracted using BLAST (Basic Local Alignment Search Tool), and these were mapped to the CHO genome (GCF_003668045.3) to obtain mapping data (bam files). The minimap2 mapping tool was used. The .bam file was visualized using Integrative Genomics Viewer (IGV) (Broad Institute) to comprehensively identify the boundary between the GoI and the genome. Hereinafter, the boundary between the GoI and the genome is referred to as a "junction," and its coordinates are referred to as "junction coordinates."

[0147] [Results] A total of 226 junctions were found in 31 antibody-producing strains.

[0148] <Measurement of GoI insertion pattern> [Method] A fastq file was obtained by base calling the fast5 file output from the sequencer. The reads were mapped to the CHO genome (CriGri-PICRH). Minimap2 was used as the mapping tool. Only reads that mapped within 6500 bp from the junction coordinates were extracted. Using the self-developed software "TaulVis," regions of homology with the genome and with GoI within the extracted reads were visualized. Reads in which the two junctions and GoI could be read as a single sequence (i.e., "junction-GoI-junction") were extracted.

[0149] [Results] The above series of reads contained a total of 176 junctions.

[0150] <Identifying junctions flanking GoI> [Method] The H chain gene and L chain gene were mapped onto the above-mentioned series of reads. Minimap2 was used as the mapping tool. When the full length of the H chain gene could be mapped, it was counted as the H chain copy number. Similarly, when the full length of the L chain gene could be mapped, it was counted as the L chain copy number. Reads in which one or more copies of the H chain gene and L chain gene were counted were extracted, and the junctions contained in the extracted reads were identified.

[0151] [Results] Figure 3 shows a histogram of the copy numbers of the H-chain gene and the L-chain gene. In the histogram in Figure 3, the horizontal axis represents the copy number and the vertical axis represents the number of junctions. A total of 165 junctions were contained in the reads in which one or more copies of the H-chain gene and the L-chain gene were counted.

[0152] <Detection of Hypomethylated Regions> [Method] For reads in which one or more copies of the H-chain gene and L-chain gene were counted, methylated cytosines at CpG sites were detected from the fast5 file described above using a base modification analysis tool (Megalodon, Oxford Nanopore Technologies). The methylation rate (methylated cytosines at CpG sites / total cytosines at CpG sites × 100) was measured for each promoter region, H-chain coding sequence, and L-chain coding sequence, and the arithmetic mean of the methylation rates was calculated.

[0153] [Results] Figure 4 shows the average methylation rates for the promoter region, H-chain coding sequence, and L-chain coding sequence. In the histogram in Figure 4, the horizontal axis represents the average methylation rate, and the vertical axis represents the number of junctions. A total of 140 junctions were contained in the reads in which the average methylation rate for all promoter regions and antibody subunit regions was 10% or less.

[0154] <Obtaining TAD boundary coordinates> [Method] The genome structure data (fastq file, SRRID: SRR12194154) to be used for Hi-C analysis was obtained from a paper analyzing the genome structure of CHO cells (William Hilliard, Kelvin H. Lee. Systematic identification of safe harbor regions in the CHO genome through a comprehensive epigenome analysis. Biotechnology and Bioengineering, 2020. https: / / doi.org / 10.1002 / bit.27599). Using the mapping program bwa, the fastq file was mapped to the CHO genome (CriGri-PICRH, https: / / www.ncbi.nlm.nih.gov / datasets / genome / GCF_003668045.3 / ) (alignment parameters: -A1, -B4, -E50, -L0) to obtain mapping data (bam file). Next, a contact map was obtained from the mapping data (.bam file) using the hicBuildMatrix command (binSize: 100000) of HiCExplorer, a HiC data analysis and visualization tool. Next, the boundary coordinates of TADs (a set of pairs of start and end points) were obtained from the contact map using the hiCFindTADs command (minDepth: 300000, maxDepth: 600000, Step: 100000, minBoundaryDistance 400000).

[0155] [Results] The genome structure data obtained from the above paper contained 281,721,369 paired-end reads. All reads were mapped to the CHO genome, resulting in 150 million paired reads, which were used to obtain a contact map for each chromosome. Figure 5 shows the contact map (expressed as a heat map) for RefSeq accession number NW_023276806.1. 2,502 TADs were detected across the entire CHO genome.

[0156] <Calculation of TAD Score> [Method] Based on the coordinate information of the endogenous genes in CHO cells (GCF_003668045.3_CriGri-PICRH-1.0_genomic.gtf), the endogenous genes contained in each of the 2,502 TADs were detected. RNA-Seq (Takara Bio) was performed on CHO-DG44 cells to quantify the mRNA levels of the endogenous genes. The mRNA levels were corrected to TPM (transcripts per million) to obtain the expression levels of the endogenous genes. For each of the 2,502 TADs, the density of the endogenous genes was multiplied by the average expression level of the endogenous genes (the average of the TPM values ​​after common logarithmic transformation), and this value was used as the TAD score.

[0157] [Results] The distribution of TAD scores is shown in Figure 6. The "x" in Figure 6 represents the 2502 TAD scores in descending order.

[0158] <Identification of TAD scores of known safe harbors> [Method] The Fer114 locus, Hprt locus, and C12orf35 locus are known as safe harbors for CHO cells. The TADs to which these three loci belong were identified, and their TAD scores were determined. The minimum of the three TAD scores was determined as the threshold.

[0159] [Results] The TAD scores of the TADs to which the three loci belonged were as follows. These are shown as three horizontal lines in the distribution diagram in Figure 6: Fer114 locus = 34.3, Hprt locus = 12.3, and C12orf35 locus = 9.0. The threshold was determined to be 9.0.

[0160] <Selection of junctions with high TAD scores> [Method] The TADs to which 140 junctions belonged and their TAD scores were identified. From among the 140 junctions, junctions with a TAD score above a threshold of 9.0 (i.e., the TAD score of the C12orf35 locus) were selected.

[0161] [Results] The "◯" in Figure 6 indicates 140 junctions selected based on the methylation rate. Of the 140 junctions, 28 had a TAD score of more than 9.0.

[0162] <Detection of junction pairs without chromosomal translocation> [Method] When junction pairs sandwiching a region containing one or more copies of the H-chain gene and the L-chain gene could be determined to be on the same chromosome based on the junction coordinates and the distance between the junction pairs on the reads was 100 kbp or less, it was determined that there was no chromosomal translocation between the junction pairs. Based on this determination criterion, junction pairs without chromosomal translocation between the junctions were detected.

[0163] [Results] Of the 28 junctions with a TAD score of over 9.0, 14 were junction pairs without chromosomal translocations between them. The seven regions between the seven junction pairs were estimated by this series of processes to be safe harbors for the inserted gene and regions where the inserted gene is stably and highly expressed. Table 4 shows the coordinates of the seven regions and 14 junctions (the "start" and "end" of the seven regions).

[0164]

[0165] <Verification of gene expression levels in seven regions> [Method] Plasmids were prepared in which homology arms for the seven regions were added to both ends of the inserted gene for each of the seven regions shown in Table 4. The inserted genes were the two antibody subunit genes (L chain-H chain) described above.

[0166] The prepared plasmid was introduced into CHO cells by electroporation (4D-Nucleofector X unit system, Lonza). Using PCR to amplify the boundary between the genome and the inserted gene and a digital PCR system (QX-200, Bio-Rad Laboratories), strains in which only one copy of the inserted gene had been introduced into the intended region were selected. For comparison, a strain in which only one copy of the inserted gene had been introduced into the Fer114 locus, a known safe harbor for CHO cells, was produced in the same manner as above. For comparison, a linearized inserted gene was introduced into CHO cells by electroporation, allowing random integration into the genome. A strain in which only one copy of the inserted gene had been introduced per genome was selected using a digital PCR system. The cell lines were cultured, and the mRNA levels of the H and L chains were quantified by real-time PCR (CFX96, Bio-Rad Laboratories).

[0167] [Results] Figure 7 shows the amount of mRNA for the heavy chain. The vertical axis represents the ratio of mRNA amount to the average amount of mRNA amount for 47 strains produced by random integration into which only one copy of the inserted gene was introduced. "RI" represents the 47 strains produced by random integration. "TI" represents, from top to bottom, the strain with insertion into the Fer114 locus, the strain with insertion into ID029, and the strain with insertion into ID040. The results shown in Figure 7 show that gene insertion into the region discovered by this embodiment has a higher probability of producing a strain that highly expresses the target gene than random integration.

[0168] 8 shows the specific production rates of H chains for the strains with insertion at ID040, ID029, ID038, ID032, and ID024. The specific production rates here refer to the ratio to the production rate of the strain with insertion at the Fer114 locus. The five strains above had specific production rates of over 0.8, which were comparable to the strain with insertion at the Fer114 locus.

[0169] <Creation of a multicopy insertion strain> [Method] An attempt was made to insert multiple copies of the coding sequence into the seven regions shown in Table 4. For each of the seven regions, a plasmid was created in which homology arms for the seven regions were added to both ends of the inserted gene. The inserted gene was a nucleic acid in which two or four of the two antibody subunit genes (L chain - H chain) described above were linked together.

[0170] The prepared plasmid was introduced into CHO cells by electroporation (4D-Nucleofector X unit system, Lonza). Strains in which only two or four copies of the inserted gene were introduced into the intended region were selected using PCR and a digital PCR system (QX-200, Bio-Rad Laboratories), which amplify the boundary between the genome and the inserted gene. To confirm the long-term passage stability of the cells, the cells were suspended in a passage medium and passaged every three days. Passage was terminated when the total number of cell divisions exceeded 60.

[0171] [Results] Figure 9 shows the copy numbers of the L chain and H chain determined by ddPCR. The horizontal axis represents the name of the produced strain, with the strain identification number added after the region ID. Strains ID040-1 and ID040-2 had inserted therein a nucleic acid in which two L chains and two H chains were linked. ID040 was a region in which at least two copies of an L chain and an H chain could be inserted. Strains ID029-1 and ID029-2 had inserted therein a nucleic acid in which four L chains and four H chains were linked. ID029 was a region in which at least four copies of an L chain and an H chain could be inserted. Strains ID032-1 and ID032-2 had inserted therein a nucleic acid in which four L chains and four H chains were linked. ID032 was a region in which at least four copies of an L chain and an H chain could be inserted.

[0172] <Fed-batch culture test> [Method] The ID040-1 strain, ID040-2 strain, ID029-1 strain, ID029-2 strain, ID032-1 strain, and ID032-2 strain were subjected to a culture test.

[0173] Six antibody-producing strains before and after long-term passage were each suspended in 40 mL of basal medium, transferred to a 125 mL flask for shaking culture, and incubated at 37°C and CO 2Shaking culture and feeding were performed at 140 rpm in a 5% (v / v) atmosphere. A fixed amount of feed medium was added daily from day 3 to day 13 of culture. Sampling was performed every 1 to 3 days to measure cell density, culture medium components, and antibody concentration. Measurements were performed using a Vi-CELL XR cell counter (Beckman Coulter), a FLEX2 cell culture environment analyzer (Nova Biomedical), and a CedexBio product analyzer (Roche Diagnostics). On day 14 of culture, the culture medium was collected, and cells and cell debris were removed using a depth filter (pore size 0.22 μm) to obtain the culture supernatant. The antibody concentration in the culture supernatant was measured by liquid chromatography using a protein A column.

[0174] [Results] Figure 10 shows the antibody production performance of six antibody-producing strains before and after long-term culture. "Pre" indicates the antibody-producing strain before long-term culture, and "Post" indicates the antibody-producing strain after long-term culture. The vertical axis represents the relative specific production rate. The results shown in Figure 10 demonstrate that antibody-producing strains in which two or four copies of the L-chain-H-chain have been inserted into ID040, ID029, or ID032 do not experience a decrease in antibody production even after long-term culture, and are capable of stable antibody production. Compared to the strain in which two copies of the L-chain-H-chain have been inserted into ID040, the strain in which four copies of the L-chain-H-chain have been inserted into ID029 or ID032 has higher antibody productivity, demonstrating that the antibody production amount increases as the copy number of the coding sequence increases.

[0175] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[0176] The disclosure of Japanese Application No. 2023-202228, filed on November 29, 2023, is incorporated herein by reference in its entirety.

Claims

1. A search method for finding a target region on a genome into which a target gene is to be inserted, comprising the steps of: (1) finding a TAD as the target region; (1) calculating a TAD score indicating transcriptional activity for each TAD in the genome, and selecting a TAD based on the TAD score.

2. The searching method of claim 1, further comprising the following (2): (2) identifying a TAD score-SH, which is the TAD score of a TAD belonging to at least one known safe harbor in the genome; determining a threshold based on at least one of the TAD score-SH; and selecting a TAD based on the threshold and the TAD score.

3. A search method for finding a target region on a genome into which a gene of interest is to be inserted, comprising the steps of (a) to (d) below, and finding a region identified by the following boundary pair P as the target region; (a) preparing a cell having a genome F in which an exogenous gene has been incorporated and expressing the exogenous gene; (b) obtaining the genome F from the cell, and analyzing the genome F to identify a region F which is a region containing the exogenous gene and a boundary F which is the boundary between the region F and the genome; (c) calculating a TAD score indicating transcriptional activity for each TAD in the genome, and for each boundary F, identifying a TAD score-F which is the TAD score of the TAD to which it belongs, and selecting the boundary F based on the TAD score-F; and (d) finding a boundary pair P which sandwiches the region F from among the selected boundaries F.

4. The searching method of claim 3, further comprising the following (p), performing (c) on the boundary F selected by (p): (p) measuring the methylation rate of the foreign gene present in the region F, and selecting the boundary F of the region F in which the methylation rate is low.

5. The searching method of claim 3 or 4, further comprising the following (q), and performing (d) for the boundary F selected by (c) and (q) below; (q) identifying a TAD score-SH, which is the TAD score of the TAD belonging to at least one of the known safe harbors in the genome, determining a threshold based on at least one of the TAD score-SH, and selecting the boundary F based on the threshold and the TAD score-F.

6. A searching method according to claim 3 or claim 4, further comprising the following (r): (r) selecting a boundary pair P such that the region F between the boundary pair P in the genome F is not a region formed by a chromosomal translocation.

7. A searching method described in any one of claims 1 to 4, wherein the TAD score is a value obtained by multiplying the density of genes present in the TAD by their average expression level.

8. The searching method according to claim 1 or 3, wherein the genome is a mammalian cell genome.

9. The method of claim 1 or 3, wherein the genome is a CHO cell genome.

10. A method for producing a cell, comprising: finding the TAD by the searching method described in claim 1; and inserting a target gene into the TAD of a genome having the TAD.

11. A method for producing a cell, comprising: finding a region identified by the boundary pair P using the search method described in claim 3; and inserting a target gene within ±10 kbp before and after the region of a genome having the region.

12. A method for producing a cell according to claim 10 or 11, wherein the target gene is a gene encoding at least one selected from the group consisting of an enzyme, an antibody, an interleukin, a cytokine, a chemokine, a hormone, a growth factor, a transcription factor, a receptor, a transcriptional regulatory nucleic acid, a non-coding RNA, a viral preparation, a vaccine, a medical protein, a subunit thereof, and a fragment thereof.

13. A method for selecting cells expressing a target gene, comprising the following (11): (11) calculating a TAD score indicating transcriptional activity for each TAD in the genome of a cell, selecting a TAD based on the TAD score, and selecting a cell in which the target gene is present in the selected TAD.

14. The method for selecting cells according to claim 13, further comprising the following (12): (12) identifying a TAD score-SH, which is the TAD score of a TAD belonging to at least one known safe harbor in the genome; determining a threshold based on at least one of the TAD score-SH; selecting a TAD based on the threshold and the TAD score; and selecting a cell in which the target gene is present in the selected TAD.

15. A method for selecting cells according to claim 13 or 14, further comprising the following (13): (13) measuring the methylation rate of the target gene present in the selected TAD, and selecting cells having a low methylation rate.

16. A method for selecting cells described in claim 13 or claim 14, wherein the TAD score is a value obtained by multiplying the density of genes present in the TAD by their average expression level.

17. The method for selecting cells according to claim 13, wherein the cells are mammalian cells.

18. The method for selecting cells according to claim 13, wherein the cells are CHO cells.

19. The method for selecting cells described in claim 13, wherein the target gene is a gene encoding at least one selected from the group consisting of enzymes, antibodies, interleukins, cytokines, chemokines, hormones, growth factors, transcription factors, receptors, transcription regulatory nucleic acids, non-coding RNA, viral preparations, vaccines, medical proteins, subunits thereof, and fragments thereof.

20. A cell derived from a Chinese hamster, in which a target gene has been inserted into at least one region selected from the 40 regions shown in Table 1 below.

21. The cell of claim 20, wherein the cell derived from Chinese hamster is a CHO cell.

22. The cell according to claim 20 or 21, wherein the target gene is a gene encoding at least one selected from the group consisting of an enzyme, an antibody, an interleukin, a cytokine, a chemokine, a hormone, a growth factor, a transcription factor, a receptor, a transcriptional regulatory nucleic acid, a non-coding RNA, a viral preparation, a vaccine, a medical protein, a subunit thereof, and a fragment thereof.

23. A method for producing a cell product, comprising culturing cells produced by the method for producing cells according to claim 10 or 11, and expressing the target gene.

24. A method for producing a cell product, comprising culturing cells selected by the method for selecting cells according to claim 13 or 14, and expressing the target gene.

25. A method for producing a cell product, comprising culturing the cell according to claim 20 or 21 to express the target gene.

Citation Information

Patent Citations

  • Site-specific integration

    EP2711428A1

  • Compositions and methods for making antibodies based on use of an expression-enhancing locus

    WO2017184831A1

  • Compositions and methods for making antibodies based on use of expression-enhancing loci

    WO2017184832A1

  • SSI cells with predictable and stable transgene expression and methods of formation

    WO2020072480A1

  • SSI cells with predictable and stable transgene expression and methods of formation

    JP2022513319A