Method for integrating genes in CHO (Chinese hamster ovary) cells and application of method

By screening essential genes in CHO cells and using the CRISPR/Cas system for homologous recombination repair, efficient and stable gene integration in CHO cells is achieved, solving the problem of insufficient gene integration targets in CHO cells and improving the efficiency of recombinant protein production.

CN120249403APending Publication Date: 2025-07-04TIANJIN INST OF IND BIOTECH CHINESE ACADEMY OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410013162.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-04
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The lack of efficient and stable gene integration targets in CHO cells leads to the problems of long screening cycles and low construction efficiency during recombinant protein production.

Method used

By screening the essential genes in CHO cells, using the CRISPR/Cas system to introduce breaks at the essential genes, and using homologous recombination repair (HDR) to integrate the target genes, combining homologous arms and regulatory elements, efficient and stable gene integration is achieved.

Benefits of technology

It improves the accuracy and efficiency of gene integration, shortens the screening cycle, and improves the expression level of recombinant proteins and the construction speed of CHO cell lines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120249403A_ABST
    Figure CN120249403A_ABST
Patent Text Reader

Abstract

The invention relates to a method for integrating genes in CHO cells and application of the method, and belongs to the field of molecular biology. According to the method for integrating the genes in the CHO cells, provided by the invention, the target genes are integrated at the screened CHO essential genes, and the cells obtained through homologous repair (HDR) are screened by utilizing whether the essential genes can be normally expressed or not, so that the screening efficiency is improved, and the target genes are accurately, efficiently and stably integrated in the CHO cell genome; meanwhile, the CHO essential gene screened by the invention is more beneficial to integration and screening of CHO cells.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of molecular biology, and relates to a method for integrating genes in CHO cells and its application, specifically to a method for efficiently mediating gene integration in CHO cells and its application. Background Art

[0002] Recombinant proteins are proteins obtained by applying recombinant DNA or recombinant RNA technology, which refers to obtaining a target gene by applying gene cloning or chemical synthesis technology, ligating it to a suitable expression vector, introducing it into a specific host cell, and expressing a protein molecule. Among them, recombinant drug proteins are an important part of biopharmaceuticals, including recombinant proteins, recombinant drug protein engineering antibodies, therapeutic vaccines, etc. Since the structure is complex or glycosylation is very important for the activity of proteins, mammalian cells have become an important platform for the production of recombinant drug proteins at present because they have post-translational modifications (PTMs) similar to human cells. Among the mammalian cells used, nearly 70% are Chinese hamster ovary (CHO) cells.

[0003] The expression form of foreign genes in host cells can be divided into transient expression and stable expression. Stable expression refers to: (1) the expression after foreign genes are transfected into eukaryotic cells and integrated into the genome. The stable expression level of recombinant genes is generally 1-2 orders of magnitude lower than that of transient expression; (2) although the host cells have been passaged many times or the conditions have changed, the expression level still remains stable. To achieve stable expression of foreign genes in host cells, the commonly used vector systems include: 1. Retrovirus system: It can effectively infect host cells and mediate the efficient integration of the foreign gene expression cassette into the genome, but its loading capacity is limited, and the preparation process of recombinant virus particles is complex; 2. Eukaryotic expression plasmid system: The preparation process is relatively simple, but it inserts into the host genome by random DNA recombination, and the integration efficiency is extremely low; 3. Transposon system: It uses a plasmid system, the preparation process is relatively simple, and the foreign gene is integrated into the genome by a transposase, and the integration efficiency is relatively low.

[0004] At present, the eukaryotic expression plasmid system is the most commonly used in industry. However, due to the random integration of foreign genes into the host genome, silencing or low expression levels often occur, and stable and highly expressing clone cell lines can only be isolated from a large number of cell clones. Among them, the production of recombinant drug proteins is restricted by factors such as low expression levels of target proteins in CHO cells, long cell cycle for screening high-yield and stable cells, and high cell culture costs (Chusainow J, et al. A study of monoclonal antibody producing CHO cell lines: what makes a stable high producer? Biotechnol. Bioeng. 2009).

[0005] CRISPR / Cas (Clustered Regularly Interspaced Short Palindromic Repeats / CRISPR-Associated systems), whose full name is clustered regularly interspaced short palindromic repeats / CRISPR-associated proteins, is an acquired immune system of bacteria and archaea. It can recognize foreign double-stranded DNA through single-stranded RNA and cut foreign DNA with the help of Cas nucleases to resist the invasion of foreign DNA. The CRISPR-Cas system has been widely studied and applied in mammalian genome editing.

[0006] Non-homologous end joining (NHEJ) and homologous recombination repair (HDR) are common DNA repair mechanisms when the body's DNA is damaged. Among them, NHEJ is the main one and HDR is the auxiliary one. At present, the gene seamless knock-in method mainly combines the CRISPR-Cas system with the HDR mechanism to achieve precise gene editing. However, due to the interference of the powerful NHEJ repair system in the host and the limitation of HDR efficiency, the gene integration method based on HDR generally has low efficiency. Therefore, efficient screening strategies are usually required to enrich cell clones with successful gene integration. For example, by co-integrating reporter genes such as fluorescence and antibiotic resistance. However, these selection strategies are time-consuming, inefficient and not suitable for use in fields such as biomedicine.

[0007] Complementary expression of disrupted essential genes has gradually been adopted as an effective screening strategy. Currently, plasmid-based stable transfection of CHO cells mainly utilizes two selection markers, GS and DHFR, which is a typical screening strategy based on essential genes. However, existing selection markers are difficult to meet the requirements in various situations. However, due to the large mammalian cell genome and complex signaling networks, the identification of essential genes is not as easy as in most model microorganisms. In 2015, using CRISPR technology through mutagenesis or the establishment of sgRNA libraries, researchers screened out approximately 2,000 essential genes related to proliferation and survival in human cells respectively (TIM WANG, et al. Identification and characterization of essential genes in the human genome, Science, 2015; VINCENTA. BLOMEN, et al. Gene essentiality and synthetic lethality in haploid human cells, Science, 2015). In 2022, based on multiple literatures including the above-mentioned ones, researchers comprehensively screened out 5,072 essential genes in human cells (Luke Funk, et al. The phenotypic landscape of essential human genes, Cell, 2022). However, there has been no comprehensive literature report on essential genes in CHO so far, and it is urgent to identify essential genes in CHO cells to broaden the application scope of the screening strategy of complementary expression of disrupted essential genes.

[0008] In addition, in the field of recombinant protein expression in CHO cells, there is a lack of research on high-level and stable expression targets in the CHO cell genome. Therefore, it is necessary to propose a gene integration system that can achieve efficient and precise targeting in CHO cells, including gene integration methods and high-expression integration targets, to overcome the drawbacks of random integration in traditional expression plasmid methods and achieve the rapid construction of high-yield CHO cell lines. Summary of the Invention

[0009] Problems to be Solved by the Invention

[0010] There are few reports on essential genes in CHO and they cannot meet the requirements in various situations. At the same time, CHO cells are the most widely used host cells for recombinant protein production in biopharmaceuticals. However, the construction of current engineering cells usually relies on the traditional stable transfection strategy of plasmid random integration, which has problems such as long screening cycles and low construction efficiency.

[0011] Solutions for Solving the Problems

[0012] In the present invention, essential genes in CHO cells are screened, and a programmable gene integration strategy is used to target these highly expressed and stable sites (essential genes) in CHO cells, so that CHO engineering cell lines can be rapidly constructed and the recombinant protein expression level can be improved.

[0013] The technical solution of the present invention is as follows:

[0014] [1]. A method for integrating a target gene into the genome of a host cell, the method comprising:

[0015] a. Contacting one or more host cells, which are CHO cells, with the following (i) and (ii):

[0016] (i) A nuclease capable of generating a break at an essential gene in the host cell, wherein the essential gene encodes a gene product required for cell survival and / or proliferation; and

[0017] (ii) A donor template comprising a knock-in cassette, the knock-in cassette comprising a coding sequence region of a target gene product and a coding sequence or partial coding sequence region of an essential gene; the coding sequence or partial coding sequence region of the essential gene is upstream of the coding sequence region of the target gene product;

[0018] b. Selecting host cells expressing (iii) and (iv):

[0019] (iii) The product of the target gene, and

[0020] (iv) The gene product encoded by the essential gene required for cell survival and / or proliferation, or a functional variant thereof.

[0021] [2]. The method according to [1], wherein the essential gene is selected from at least one of the genes shown in Table 1;

[0022] Preferably, the essential gene is selected from at least one of the following genes:

[0023] GAPDH, Eef1a1, Rplp0, PKM, Tubb, Aldoa, Ybx3, Fkbp10, Cdc42, Cs, Cap1, Nop14, Fdx2 and Ccnd3.

[0024] [3]. The method according to [1] or [2], wherein the break is a single-strand break or a double-strand break, preferably a double-strand break;

[0025] Optionally, the break occurs within any exon of the coding sequence of the essential gene.

[0026] [4]. According to the method of any one of [1] to [3], the 5' and 3' ends of the knock-in cassette respectively comprise a first homologous region and a second homologous region, which are capable of recombining with a third homologous region and a fourth homologous region respectively, wherein the third homologous region and the fourth homologous region are respectively located on both sides of the break generated in the essential gene.

[0027] [5]. According to the method of any one of [1] to [4], wherein the coding sequence region of the target gene product comprises the coding sequence of at least one target gene product;

[0028] Optionally, the coding sequence of a signal peptide is included between the coding sequences of the target gene products.

[0029] [6]. According to the method of [5], wherein the knock-in cassette comprises a gene product expression regulatory element capable of separating the gene products encoded by a plurality of the essential genes, and a regulatory element for expressing the gene products encoded by the essential genes and the target gene product as separate gene products; Optionally, at least one of the gene products is a protein and the regulatory element enables the protein to be expressed separately from other gene products;

[0030] Optionally, the regulatory element comprises an IRES, a 2A element or a promoter;

[0031] Optionally, the 2A element is selected from at least one of T2A, P2A, E2A and F2A elements;

[0032] Optionally, the promoter is selected from at least one of CMV promoter, EF-1α promoter, SV40 promoter, CHEF-1α promoter and GAC promoter.

[0033] [7]. According to the method of [6], wherein the partial coding sequence of the essential gene in the knock-in cassette encodes the C-terminal fragment of the protein encoded by the essential gene; the C-terminal fragment includes the amino acid sequence encoded by the coding sequence of the essential gene across the break.

[0034] [8]. According to the method of any one of [1] to [7], wherein the nuclease is selected from the group consisting of CRISPR / Cas nuclease, zinc finger nuclease, TAL-effector DNA binding domain-nuclease fusion protein, transposase and site-specific recombinase; Optionally, the nuclease includes Cas9 endonuclease.

[0035] [9]. According to the method described in any one of [1] to [8], the host cell is contacted with at least one nuclease and at least one donor template to achieve integration of a target gene at at least one site in the host cell genome.

[0036]

[10] . According to the method described in any one of [1] to [9], the target gene comprises any gene that can be expressed in CHO cells.

[0037]

[11] . According to the method described in any one of [1] to

[10] , it further comprises the step of recovering the host cell, wherein the donor template has undergone homologous recombination at the site where the essential gene is cleaved.

[0038]

[12] . A host cell, comprising:

[0039] An insertion cassette is inserted into the coding sequence of an essential gene in the genome of the host cell, wherein the essential gene encodes a gene product required for the survival and / or proliferation of the cell, and the insertion cassette comprises a coding sequence region of the target gene product and a coding sequence or partial coding sequence region of the essential gene; the coding sequence or partial coding sequence region of the essential gene is downstream of the coding sequence region of the target gene product;

[0040] Optionally, the expression of the target gene product and the gene product encoded by the essential gene is initiated by the endogenous promoter of the essential gene;

[0041] Preferably, the host cell is a CHO cell;

[0042] Preferably, the host cell expresses: the product of the target gene, and the gene product encoded by the essential gene required for the survival and / or proliferation of the cell, or a functional variant thereof;

[0043] Optionally, the essential gene comprises the genes shown in Table 1;

[0044] Optionally, the target gene comprises any gene that can be expressed in CHO cells.

[0045]

[13] . According to the host cell described in

[12] , wherein the host cell comprises regulatory elements capable of expressing the gene product encoded by the essential gene and the target gene product as separate gene products; optionally, wherein at least one of the gene products is a protein and the regulatory elements enable the protein to be expressed separately from other gene products;

[0046] Optionally, the regulatory elements comprise IRES, 2A element, and / or promoter;

[0047] Optionally, the 2A element is selected from at least one of T2A, P2A, E2A, and F2A elements;

[0048] Optionally, the promoter is selected from at least one of CMV promoter, EF-1α promoter, SV40 promoter, CHEF-1α promoter, and GAC promoter.

[0049]

[14] . A pharmaceutical composition comprising: a host cell prepared by the method according to any one of [1] to

[11] , or a host cell as described in any one of

[12] or

[13] ; and a pharmaceutically acceptable carrier or excipient.

[0050]

[15] . A kit comprising: a host cell prepared by the method according to any one of [1] to

[11] , or a host cell as described in

[12] or

[13] , or a pharmaceutical composition as described in

[14] .

[0051] Effects of the Invention

[0052] The method for integrating a target gene into the CHO cell genome provided by the present invention finds an essential gene with high transcriptional intensity in CHO cells, integrates the target gene at the essential gene, and utilizes the survival pressure brought by the normal expression of the essential gene to quickly eliminate cells repaired by NHEJ mismatch, greatly improving the screening efficiency, and achieving precise, efficient, and stable integration of the target gene into the CHO cell genome. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 Shows an exemplary integration strategy targeting an essential gene in an embodiment of the present invention.

[0054] Figure 2 Shows an exemplary integration strategy targeting the GAPDH gene in an embodiment of the present invention.

[0055] Figure 3 Shows the GFP intensity corresponding to two DNA donors (cleavable linker and promoter) in an embodiment of the present invention.

[0056] Figure 4 Shows the yield of cell monoclonal antibodies screened in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0057] The following will detail various exemplary embodiments, features, and aspects of the present invention. The special word "exemplary" here means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" here does not have to be construed as superior to or better than other embodiments.

[0058] In addition, for a better illustration of the present invention, numerous specific details are given in the following detailed embodiments. Those skilled in the art should understand that the present invention can also be implemented without certain specific details. In other instances, methods, means, apparatuses, and steps well-known to those skilled in the art are not described in detail so as to highlight the gist of the present invention.

[0059] Unless otherwise specified, the units used in this specification are all international standard units, and the numerical values and numerical ranges appearing in the present invention should be understood to include the inevitable systematic errors in industrial production.

[0060] In this specification, the meaning expressed by "may" includes both the meaning of performing a certain process and the meaning of not performing a certain process.

[0061] In this specification, the "some specific / preferred embodiments", "other specific / preferred embodiments", "embodiments", etc. mentioned refer to the specific elements (e.g., features, structures, properties, and / or characteristics) related to the embodiments described, which are included in at least one of the embodiments described herein, and may or may not exist in other embodiments. Additionally, it should be understood that the elements can be combined in various embodiments in any suitable manner.

[0062] As used herein, the term "CRISPR / Cas nuclease" refers to any CRISPR / Cas protein having DNA nuclease activity, such as Cas9 or Cas12 proteins that exhibit specific association (or "targeting") with a DNA target site (e.g., within a genomic sequence in a cell) in the presence of a guide molecule. The strategies, systems, and methods disclosed herein can use any combination of CRISPR / Cas nucleases disclosed herein or known to those of ordinary skill in the art. Those of ordinary skill in the art will know additional CRISPR / Cas nucleases and variants suitable for the context of this disclosure, and it should be understood that the disclosure is not limited in this regard.

[0063] As used herein, the term "nuclease" refers to any protein that catalyzes the cleavage of phosphodiester bonds. In some embodiments, the nuclease is a DNA nuclease. In some embodiments, the nuclease is a "nickase" that, when it cleaves double-stranded DNA, such as genomic DNA in a cell, results in a single-strand break. In some embodiments, the nuclease cleaves double-stranded DNA, such as genomic DNA in a cell, resulting in a double-strand break. In some embodiments, the nuclease binds to a specific target site within double-stranded DNA, and the target site overlaps or is adjacent to the location of the resulting break. In some embodiments, the nuclease causes a double-strand break that includes overhangs from 0 (blunt ends) to 22 nucleotides in both the 3' and 5' directions. As discussed herein, CRISPR / Cas nucleases, zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), and meganucleases are exemplary nucleases that can be used according to the strategies, systems, and methods of the present disclosure.

[0064] As used herein, the term "ribonucleoprotein (RNP)" refers to a nucleoprotein that contains RNA, which is a form of binding nucleic acids and proteins together. Ribonucleoproteins include ribosomes, telomerase, and small nuclear ribonucleoproteins (snRNPs).

[0065] As used herein, the term "essential gene" with respect to a cell refers to a gene that encodes at least one gene product required for cell survival and / or proliferation. Essential genes can be housekeeping genes that are essential for the survival of all cell types, or genes that need to be expressed in a particular cell type for survival and / or proliferation under specific culture conditions. In some embodiments, the loss of essential gene function results in a significant reduction in cell survival, e.g., a significantly reduced survival time of cells characterized by the loss of essential gene function compared to cells of the same cell type but without the loss of the same essential gene function. In some embodiments, the loss of essential gene function results in the death of the affected cells. In some embodiments, the loss of essential gene function results in a significant reduction in cell proliferation, e.g., a significant reduction in the ability of cells to divide, which can be manifested during an important time period required for cells to complete the cell cycle, or, in some preferred embodiments, cells completely lose the ability to complete the cell cycle and thus proliferate.

[0066] As used herein, the terms "polypeptide", "protein", and "peptide" are used interchangeably herein to refer to a polymeric form of amino acids of any length, which may include encoded and non-encoded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides with similar peptide backbones.

[0067] As used herein, "cargo gene" is the same as "gene of interest". It should be understood that the methods and cells of the present invention are not limited to any specific gene product of interest, and the choice of the gene product of interest will depend on the type of cell and the ultimate use of the cell.

[0068] As used herein, the term "IgG1" is a subtype of IgG antibody, which is an immunoglobulin widely present in human blood and tissues. It is a dimer composed of four single-chain antibody molecules, and each single chain contains a variable region and a constant region. Each single-chain antibody is composed of a heavy chain and a light chain. The heavy chain of IgG1 is the γ1 chain, and the light chain can be either the κ chain or the λ chain.

[0069] As used herein, the term "GAPDH" refers to glyceraldehyde-3-phosphate dehydrogenase, which is an essential protein that catalyzes the oxidative phosphorylation of glyceraldehyde-3-phosphate in the presence of inorganic phosphate and nicotinamide adenine dinucleotide (NAD), which is an important energy-producing step in carbohydrate metabolism.

[0070] As used herein, the term "ribosome entry site" or "IRES (internal ribosome entry site)" is a nucleic acid sequence, the presence of which enables protein translation initiation to be independent of the 5' cap structure, thus making it possible to initiate translation directly from the middle of messenger RNA (mRNA). IRES is commonly constructed in the middle of eukaryotic bicistronic mRNA. In this way, the first protein is usually initiated by the 5' cap structure, while the second protein is initiated by the IRES. The expression of the two proteins before and after the IRES is usually proportional, so the expression of one protein can be reflected according to the expression of one of the reporter genes.

[0071] As used herein, the term "2A peptide" refers to a class of peptide fragments that are 18-22 amino acid residues in length and can induce the self-cleavage of recombinant proteins containing the 2A peptide in cells. These peptides all have a certain sequence motif, which often causes the ribosome to be unable to connect at the junction of the last glycine (G) and proline (P), thus resulting in the "cleavage" effect. Four commonly used 2A peptides include P2A, T2A, E2A, and F2A, among which P2A usually has the highest cleavage efficiency (close to 100% in some cases). Although it is not a necessary condition for the function of the 2A peptide, adding a GSG (Gly-Ser-Gly) sequence to the N-terminus of the 2A peptide sequence can improve the cleavage efficiency induced by the 2A peptide.

[0072] In the present invention, the term "transfection" refers to the entry of recombinant plasmid vectors or free nucleotides into eukaryotic cells mediated by liposomes and the like.

[0073] In some embodiments of the present invention, the term "codon optimization" may refer to that the nucleotide sequence encoding a polypeptide has been configured to contain codons preferred by a host cell or organism, so as to improve gene expression in the host cell or organism and enhance translation efficiency. In other embodiments of the present invention, the term "codon optimization" refers to the "codon optimization" of the coding sequence or partial coding sequence region of an essential gene in the knock-in cassette, so as to prevent further recognition and cleavage by nucleases (such as repeated binding of the gRNA targeting domain sequence), and to reduce the possibility of recombination after integrating the knock-in cassette into the genome of the cell.

[0074] In the present invention, the terms "signal peptide", "secretory peptide", "signal sequence" or "signal peptide sequence" refer to such short peptides that, when fused with a protein of interest (such as an antibody of the present disclosure), can promote the secretion of the protein of interest expressed by a cell to the cell membrane or outside the cell. The signal peptide is usually located at the N-terminus of the protein of interest, and various signal peptides are known to those skilled in the art, such as but not limited to hemagglutinin signal sequence, human insulin signal sequence, human interleukin-2 signal sequence, albumin signal sequence, etc.

[0075] In the present invention, the term "polycistronic" or "multicistronic" when used herein with reference to a knock-in cassette refers to the fact that the knock-in cassette can express two or more proteins from the same mRNA transcript. Similarly, a "bicistronic" knock-in cassette is a knock-in cassette that can express two proteins from the same mRNA transcript.

[0076] As used herein, the knock-in efficiency refers to the integration efficiency of the gene of interest detected by real-time fluorescence quantitative PCR (RT-), and the specific calculation method is -ΔΔCt the 2-ΔΔCt method, where ΔΔCt = ΔCt (experimental group) - ΔCt (control group); ΔCt (experimental group) = Ct (gene of interest) - Ct (housekeeping gene), and the same applies to ΔCt (control group); the experimental group is a mixed cell population subjected to gene editing using the method disclosed in the present invention, and the control group is a homozygous cell population obtained previously with one copy of the gene of interest integrated into the genome; the housekeeping gene is GAPDH.

[0077] In the present invention, the term "pharmaceutical composition" means a composition containing one or more antibodies described herein, as well as other components such as a physiological / pharmaceutically acceptable carrier and excipient. The purpose of the pharmaceutical composition is to facilitate the administration to an organism, promote the absorption of the active ingredient and thus exert its biological activity.

[0078] In the present invention, the term "pharmaceutically acceptable" (or "pharmacologically acceptable") refers to molecular entities and compositions that, when appropriately administered to animals or humans, do not produce adverse reactions, allergic reactions, or other untoward effects. As used herein, the term "pharmaceutically acceptable carrier" includes any and all solvents, dispersion media, coatings, antibacterial agents, isotonic agents, and absorption delaying agents, buffers, excipients, binders, lubricants, gels, surfactants, etc. that can be used as a medium for pharmaceutically acceptable substances.

[0079] The technical solution of the present invention will be described in detail as follows:

[0080] In the present invention, the nuclease is used to target the coding frame of an essential gene in the CHO cell genome to break the coding sequence of the essential gene, and then a repair template with upstream and downstream homologous arms, a partial coding frame of the essential gene, and the target gene is transferred into CHO cells. At this time, the cells will have two repair methods: NHEJ (non-homologous end joining) with higher efficiency and HDR (homologous recombination repair) with lower efficiency in the cells; cells repaired by NHEJ mismatch repair will die and be eliminated because they cannot synthesize functional essential proteins; cells correctly knocked in by HDR can normally express the essential gene, the cells can survive normally, and at the same time, the target gene product can also be expressed. Therefore, the method of the present invention solves the problem that NHEJ interference causes low HDR efficiency during cell gene integration and establishes an efficient genome precise integration system.

[0081] Further, there is a partial coding frame of the essential gene replaced by synonymous codons between the upstream homologous arm and the target gene in the repair template capable of causing HDR in CHO cells. This partial coding frame can seamlessly connect with the essential gene sequence in the upstream homologous arm and translate the complete protein corresponding to the essential gene, enabling the cells undergoing HDR to grow normally (see Figure 1 ). Thus, it can be seen that the present invention also greatly improves the screening efficiency by using the survival pressure brought about by the normal expression of the essential gene.

[0082] Furthermore, for the CHO cells, which are commonly used as recombinant protein expression hosts in biomedicine, the inventors have screened essential genes with relatively high transcriptional levels from the genome of CHO cells. Therefore, the present invention discloses essential gene loci capable of enabling high-level and stable expression of target genes, which can be used as preselected targets for gene integration (see Table 1 specifically), providing more choices for the efficient integration and screening of target genes.

[0083] Table 1. List of essential genes in CHO cells that can be used for targeted integration and have relatively high transcriptional levels

[0084]

[0085]

[0086]

[0087]

[0088]

[0089] <Method for integrating a target gene into the genome of a host cell>

[0090] According to some embodiments of the present invention, provided is a method for integrating a target gene into the genome of a host cell, the method comprising:

[0091] a. contacting one or more host cells, which are CHO cells, with the following (i) and (ii):

[0092] (i) a nuclease capable of generating a break at an essential gene in the host cell, wherein the essential gene encodes a gene product required for cell survival and / or proliferation; and

[0093] (ii) a donor template, the donor template comprising a knock-in cassette, the knock-in cassette comprising a coding sequence region of the target gene product, and a coding sequence or a partial coding sequence region of the essential gene; the coding sequence or the partial coding sequence region of the essential gene is upstream of the coding sequence region of the target gene product;

[0094] b. selecting host cells expressing (iii) and (iv):

[0095] (iii) the product of the target gene, and

[0096] (iv) the gene product encoded by the essential gene required for cell survival and / or proliferation, or a functional variant thereof.

[0097] In some embodiments, the CHO cells are Chinese hamster ovary cells, such as cell lines CHO-K1, CHO-S, CHO-K1SV, CHO-GS, CHO-DXB11 or CHO-DG44.

[0098] In some embodiments, the essential genes control the basic functions required for cell survival, including DNA replication, transcription, mRNA splicing, translation, vesicle trafficking, cell division, etc. The deletion of essential genes will cause the cell to lose the physiological and biochemical functions necessary for cell survival and / or proliferation, resulting in a significant reduction in cell survival time and cell death. Optionally, the essential genes in CHO cells are listed in Table 1. In some specific embodiments, the essential gene is selected from at least one of the following genes: GAPDH, Eef1a1, Rplp0, PKM, Tubb, Aldoa, Ybx3, Fkbp10, Cdc42, Cs, Cap1, Nop14, Fdx2, and Ccnd3.

[0099] In some embodiments, one or more host cells are contacted with two or more nucleases capable of creating a break at the essential gene in the host cell, wherein the essential gene encodes a gene product required for cell survival and / or proliferation; and one or more host cells are contacted with two or more donor templates. Specifically, the nuclease creates a break at the essential gene through a guide RNA (gRNA) targeting the essential gene. In some embodiments, the break is a single-strand break or a double-strand break, preferably a double-strand break. In some optional embodiments, the break occurs within any exon of the coding sequence of the essential gene. In some preferred embodiments, the break occurs within any one of the last first, second last, third last, or fourth last exons of the coding sequence of the essential gene, more preferably within the second last or third last exon of the coding sequence of the essential gene. The break occurring within the above exons is more conducive to completely inactivating the essential gene, so that the cells undergoing NHEJ mismatch repair cannot survive, and the coding sequence or partial coding sequence region of the codon-optimized essential gene in the knock-in cassette can be as short as possible to improve the efficiency of integrating the target gene into the host cell genome and reduce costs. Exemplarily, some essential genes, their targeting sequences, and break positions are listed in Table 2.

[0100] In some exemplary embodiments, the essential genes include the GAPDH gene and the PKM gene. The second last exon of the GAPDH gene is targeted by a nuclease, and the corresponding gRNA targeting sequence is as shown in SEQ ID NO:1; the second last exon of the PKM gene is targeted by a nuclease, and the corresponding gRNA targeting sequence is as shown in SEQ ID NO:4.

[0101] Table 2. Targeting Sequences of Some Essential Genes

[0102]

[0103] In some embodiments, the donor template is used to integrate a gene of interest into the host cell genome by homology-directed repair (HDR). Whether single-stranded or double-stranded, the donor template generally includes regions homologous to DNA regions within or near (e.g., flanking or adjacent to) the target sequence to be cleaved, i.e., the knock-in cassette includes a first homologous region (i.e., 5' homologous arm) and a second homologous region (i.e., 3' homologous arm) at the 5' and 3' ends, respectively, which are capable of recombining with a third homologous region and a fourth homologous region, respectively, where the third homologous region and the fourth homologous region are located on either side of the break generated in the essential gene. These homologous regions are referred to herein as "homologous arms" and are schematically illustrated below with respect to the knock-in cassette (which may be separated from one or both of the homologous arms by additional spacer sequences not shown):

[0104] [5' homologous arm]-[knock-in cassette]-[3' homologous arm].

[0105] In some embodiments, the homologous arms can have any suitable length (including 0 nucleotides if only one homologous arm is used), and the 5' and 3' homologous arms can have the same length or can have different lengths. The selection of an appropriate homologous arm length may be affected by various factors, such as the desire to avoid homology or microhomology with certain sequences (e.g., Alu repeats or other extremely common elements). For example, the 5' homologous arm can be shortened to avoid sequence repeat elements. In other embodiments, the 3' homologous arm can be shortened to avoid sequence repeat elements. In some embodiments, both the 5' and 3' homologous arms can be shortened to avoid including certain sequence repeat elements.

[0106] In some alternative embodiments, the knock-in cassette includes a polycistronic (e.g., bicistronic) knock-in cassette that includes coding sequences for two or more gene products (i.e., the coding sequence region of the gene product of interest includes at least one coding sequence for a gene product of interest). Further, the coding sequence of a signal peptide is included between the coding sequences of the gene products of interest; exemplarily, the knock-in cassette includes a first coding sequence for a first gene product of interest, a linker (e.g., 2A peptide, IRES), and a second coding sequence for a second gene product of interest.

[0107] In some alternative embodiments, the knock-in cassette comprises gene product expression regulatory elements capable of separating the gene products encoded by multiple of the essential genes, and regulatory elements for expressing the gene products encoded by the essential genes and the gene product of interest as separate gene products; optionally, wherein at least one of the gene products is a protein and the regulatory elements enable the protein to be expressed separately from other gene products. In some alternative embodiments, the regulatory elements comprise IRES, 2A elements, and the 2A elements are selected from T2A (EGRGSLLTCGDVEENPGP, SEQ ID NO: 33), P2A (ATNFSLLKQAGDVEENPGP, SEQ ID NO: 34), E2A (QCTNYALLKLAGDVESNPGP, SEQ ID NO: 35), and / or F2A (VKQTLNFDLLKLAGDVESNPGP, SEQ ID NO: 36) elements. In these embodiments, the gene product of interest and the gene products encoded by the essential genes are expressed by the endogenous promoters of the essential genes.

[0108] In other alternative embodiments, the knock-in cassette comprises regulatory elements capable of enhancing the expression of the gene of interest, and the regulatory elements may be promoters, and any promoter capable of enhancing the expression of the gene of interest is included in the present invention; exemplarily, the promoter is one or more of the CMV promoter, EF-1α promoter, SV40 promoter, CHEF-1α promoter, and GAC promoter. If the promoter contains a stop codon, the expression of the upstream essential gene can be terminated, or, preferably, a stop codon may also be included downstream of the coding sequence or a partial coding sequence region of the essential gene to terminate the expression of the essential gene. The examples of the present application have confirmed that, compared with using IRES and 2A elements as regulatory elements for expressing the gene products encoded by the essential genes and the gene product of interest as separate gene products, when using a promoter as the regulatory element, the expression intensity of the gene product of interest is higher.

[0109] In some alternative embodiments, a terminator and / or a polyA signal sequence are further included downstream of the coding sequence of the target gene. As used herein, "polyA signal sequence" or "polyadenylation signal sequence" is a sequence that triggers endonuclease cleavage of mRNA and adds a series of adenosines to the 3'-end of the cleaved mRNA. The polyA signal sequence can be AATAAA. The AATAAA sequence can be replaced with other hexanucleotide sequences that are homologous to AATAAA and capable of emitting a polyadenylation signal, including ATTAAA, AGTAAA, CATAAA, TATAAA, GATAAA, ACTAAA, AATATA, AAGAAA, AATAAT, AAAAAA, AATGAA, AATCAA, AACAAA, AATCAA, AATAAC, AATAGA, AATTAA, or AATAAG.

[0110] In some embodiments, the partial coding sequence of the essential gene in the knock-in cassette encodes a C-terminal fragment of the protein encoded by the essential gene; the C-terminal fragment includes the amino acid sequence encoded by the disrupted coding sequence of the essential gene, so as to enable the expression of the gene product of the essential gene and / or the target gene product after integrating the knock-in cassette into the genome of the cell.

[0111] In some exemplary embodiments, if the break occurs within the penultimate exon of the essential gene, the partial coding sequence of the essential gene encodes the penultimate exon and the last exon of the essential gene. Specifically, the partial coding sequence of the essential gene includes a part of the 3'-end of the third-to-last exon and the sequence between the 5'-end of the penultimate exon and the break site, as well as the sequences of the penultimate exon and the last exon downstream of the break site.

[0112] In some embodiments, the nuclease is selected from the group consisting of CRISPR / Cas nucleases, zinc finger nucleases (ZFNs), TAL-effector DNA-binding domain-nuclease fusion proteins (TALENs), transposases, and site-specific recombinases. In some alternative embodiments, the nuclease includes a Cas9 endonuclease.

[0113] In some embodiments, the target gene includes any gene that can be expressed in CHO cells. Exemplarily, the product of the target gene can include GFP fluorescent protein, monoclonal antibody, human serum albumin, human interferon α2b, human interferon β, coagulation factor VIII, coagulation factor IX, and Fc fusion protein.

[0114] <Host cell>

[0115] According to some embodiments of the present invention, a host cell is provided, which is prepared by the method of integrating a gene of interest into the genome of the host cell. It comprises: inserting a knock-in cassette into the coding sequence of an essential gene in the genome of the host cell, wherein the essential gene encodes a gene product required for the survival and / or proliferation of the cell, the knock-in cassette comprises a coding sequence region of the gene product of interest, and a coding sequence or a partial coding sequence region of the essential gene; the coding sequence or the partial coding sequence region of the essential gene is downstream of the coding sequence region of the gene product of interest.

[0116] In some optional embodiments, the gene product of the gene of interest and the gene product encoded by the essential gene are expressed under the initiation of the endogenous promoter of the essential gene;

[0117] In some preferred embodiments, the host cell is a CHO cell;

[0118] In some preferred embodiments, the host cell expresses: the product of the gene of interest, and the gene product encoded by the essential gene required for the survival and / or proliferation of the cell, or a functional variant thereof.

[0119] In some embodiments, the host cell comprises regulatory elements capable of expressing the gene product encoded by the essential gene and the gene product of the gene of interest as separate gene products; optionally, at least one of the gene products is a protein and the regulatory elements enable the protein to be expressed separately from other gene products. In some alternative embodiments, the regulatory elements comprise an IRES, a 2A element, and the 2A element is selected from T2A (EGRGSLLTCGDVEENPGP, SEQ ID NO:33), P2A (ATNFSLLKQAGDVEENPGP, SEQ ID NO:34), E2A (QCTNYALLKLAGDVESNPGP, SEQ ID NO:35), and / or F2A (VKQTLNFDLLKLAGDVESNPGP, SEQ ID NO:36) elements.

[0120] In some other alternative embodiments, the knock-in cassette comprises a regulatory element capable of enhancing the expression of the gene of interest, and the regulatory element may be a promoter, and any promoter capable of enhancing the expression of the gene of interest is included in the present invention; exemplarily, the promoter is one or more of a CMV promoter, an EF-1α promoter, an SV40 promoter, a CHEF-1α promoter, and a GAC promoter.

[0121] In some other alternative embodiments, a terminator and / or a polyA signal sequence are further included downstream of the coding sequence of the gene of interest.

[0122] In some embodiments, the target gene comprises any gene that can be expressed in CHO cells. Exemplarily, the products of the target gene may include GFP fluorescent protein, monoclonal antibody, human serum albumin, human interferon α2b, human interferon β, coagulation factor VIII, coagulation factor IX, and Fc fusion protein.

[0123] <Drug Compositions and Kits>

[0124] According to some embodiments of the present invention, there is provided a drug composition comprising the host cell described above, and optionally, a pharmaceutically acceptable carrier or excipient.

[0125] The term "pharmaceutically acceptable" means that when the molecular entity and the composition are appropriately administered to an animal or a human, they do not produce adverse, allergic, or other untoward reactions.

[0126] Exemplarily, some substances that can be used as pharmaceutically acceptable carriers or their components are sugars such as lactose, glucose, and sucrose; starches such as corn starch and potato starch; cellulose and its derivatives such as sodium carboxymethyl cellulose, ethyl cellulose, and methyl cellulose; tragacanth powder; malt; gelatin; talc; solid lubricants such as stearic acid and magnesium stearate; calcium sulfate; vegetable oils such as peanut oil, cottonseed oil, sesame oil, olive oil, corn oil, and cocoa butter; polyols such as propylene glycol, glycerol, sorbitol, mannitol, and polyethylene glycol; alginic acid; emulsifiers such as wetting agents such as sodium lauryl sulfate; coloring agents; flavoring agents; tabletting agents, stabilizers; antioxidants; preservatives; pyrogen-free water; isotonic saline solutions, and phosphate buffer solutions, etc.

[0127] In some embodiments, the composition of the present invention can be made into various dosage forms as needed, and the physician can determine the dose beneficial to the patient according to factors such as the type of patient, age, weight, general disease condition, and administration route. The administration route can be by injection or other treatment methods.

[0128] According to some embodiments of the present invention, there is provided a kit containing the host cell according to the present invention, or the drug composition as described above.

[0129] Examples

[0130] The embodiments of the present invention will be described in detail below in conjunction with examples. However, those skilled in the art will understand that the following examples are only for illustrating the present invention and should not be construed as limiting the scope of the present invention. For those not specified in the examples, the conventional conditions or the conditions recommended by the manufacturer are followed. Those reagents or instruments not indicated by the manufacturer can be obtained as conventional products through commercial purchase.

[0131] Example 1: Integrating Cargo Genes by Targeting the Essential Gene GAPDH

[0132] In previous studies, it has been experimentally demonstrated that GAPDH (GeneID: 100736557) is an essential gene in CHO-K1 cells (purchased from ATCC CCL-61). To test the feasibility of the system, Cas9 and gRNA targeting the GAPDH gene (the gRNA targeting sequence is shown as SEQ ID NO:1 in Table 2, and the gRNA scaffold sequence is shown in SEQ ID NO:15) were used in CHO-K1 cells to introduce a DNA double-strand break at a position approximately 230 bp from the stop codon in the 6th exon (i.e., the second-to-last exon) of GADPDH.

[0133] According to the known method (for specific steps, see A ribonucleoprotein-based decaplex CRISPR / Cas9 knockout strategy for CHO host engineering. Biotechnol Prog. 2022 Jan;38(1):e3212), the CRISPR / Cas nuclease (Takara, catalog number 632678) and the gRNA targeting the GAPDH gene were introduced into CHO-K1 cells by nucleofection (electroporation) of ribonucleoprotein (RNP). At the same time, a double-stranded DNA donor template fragment was also electroporated. This donor template included a knock-in cassette, which contained, in the 5' to 3' order:

[0134] A 5' homologous arm with a length of approximately 150 bp, which contained a part of the 3' end of exon 5 and the sequence between the 5' end of exon 6 and the cleavage site (there is no intron between exons 5 and 6);

[0135] The codon-optimized sequence of exons 6 and 7 downstream of the cleavage site (optimized to prevent repeated binding of the gRNA targeting domain sequence);

[0136] The sequence encoding the T2A self-cleaving peptide (T2A), or the coding sequence of the stop codon of exon 7 and the CMV promoter;

[0137] The sequence of the cargo gene, that is, the gene of interest (in this example, the sequence encoding the GFP gene, and it can also be other various genes);

[0138] The polyA signal sequence (it can also be not set, and the natural polyA signal sequence of the target gene is used), and a 3' homologous arm with a length of approximately 150 bp (which contains the coding part of exon 6).

[0139] The 5' and 3' homologous arms flanking the knock-in cassette were designed to correspond to the sequences around the RNP cleavage site (such as Figure 2As shown, the sequences of the knock-in cassette of the T2A version and the homologous arms on both sides are shown in SEQ ID NO: 16, and the sequences of the knock-in cassette of the CMV promoter version and the homologous arms on both sides are shown in SEQ ID NO: 39).

[0140] The concentration of RNP during electroporation was 4 μM, and the donor dsDNA fragment encoding GFP was 2.5 μg. The obtained CHO-K1 cells after electroporation were statically cultured in an incubator at 37 °C and 5% CO2.

[0141] As Figure 2 shown, in cells edited by DNA nuclease but not successfully targeted by the DNA donor template, NHEJ-mediated synthesis of the essential GAPDH protein was not possible, leading to cell death. In cells successfully targeted by the DNA donor template, the GAPDH coding region was restored by correct integration of the knock-in cassette, allowing these cells to survive and continue to proliferate; moreover, the cargo gene downstream (3') of the GAPDH coding sequence was expressed, resulting in the production of a functional gene product.

[0142] On the 4th day after electroporation, the proportion of cells expressing green fluorescent protein was detected by flow cytometry and the knock-in (KI, Knockin) efficiency of the cargo gene encoding GFP integrated into GAPDH was detected by RT-qPCR (the primers used are shown in Table 3). The statistically determined knock-in efficiencies of three parallel groups of the T2A version were 62.5 ± 6.6%; the statistically determined knock-in efficiencies of three parallel groups of the CMV promoter version were 70.5 ± 4.9%. Compared with 4 days after electroporation, the knock-in efficiency of the cargo gene (the gene encoding GFP) was significantly higher at 9 days, reaching 81.2 ± 5.7% in the T2A version group; and reaching 89.6 ± 6.1% in the CMV promoter version group. This is consistent with the expectation that there is a large amount of cell death in RNP-induced GAPDH knockout cells, indicating that at this time, more GAPDH knockout cells that failed to successfully transfer the knock-in cassette died due to the lack of normal expression of GAPDH.

[0143] Table 3. Primers required for RT-qPCR verification

[0144] Targeting Position Sequence SEQ ID NOs: GAPDH (Reference Forward Primer) GTGGAATCTACTGGCGTCTT SEQ ID NO:27 GAPDH (Reference Forward Primer) GTTGTCATACTTGTCTTGGTTCA SEQ ID NO:28 GFP (Forward Primer) CGACCACTACCAGCAGAA SEQ ID NO:29 GFP (Reverse Primer) GAACTCCAGCAGGACCAT SEQ ID NO:30 IgG1 Heavy Chain (Forward Primer) TGCTGACTCTGTGGAAGGAA SEQ ID NO:31 IgG1 Heavy Chain (Reverse Primer) GAGGCGGTAGACAGGTAGG SEQ ID NO:32

[0145] Example 2: Integrating Cargo Genes by Targeting the Gene PKM

[0146] The knock-in integration and selection method described in Example 1 was used to target the PKM gene (GeneID: 100751347) in CHO-S. The PKM gene encodes pyruvate kinase, which is a key enzyme in the glycolysis pathway. The gRNA targeting sequence of the 11th exon of the PKM gene is shown as SEQ ID NO: 4 in Table 2.

[0147] The same knock-in integration and selection methods described in Example 1 were also used to target the Hsp90b1 gene (Gene ID: 100773331) in CHO-S. This gene encodes a member of the adenosine triphosphate (ATP) - metabolizing chaperone family, which plays a role in stabilizing and folding other proteins.

[0148] Although CHO-S was tested for the purposes of this experiment, the methods described can also be applied to other CHO cell types.

[0149] The cargo gene (gene of interest) integrated was GFP. The donor template sequence of the PKM gene is shown as SEQ ID NO:17, and the donor template sequence of the Hsp90b1 gene is shown as SEQ ID NO:38.

[0150] The RNP concentration during electroporation was 4 μM, and the donor dsDNA fragment encoding GFP was 2.5 μg. The CHO-K1 cells obtained after electroporation were statically cultured in an incubator at 37 °C and 5% CO2.

[0151] Using no RNP as a control: Only the donor template sequence (SEQ ID NO:17 or SEQ ID NO:38) fragment was electroporated into CHO-S cells.

[0152] RT-qPCR was performed 4 days after electroporation (the primers used are shown in Table 3) to determine the extent of successful integration of each knock-in cassette at its corresponding target site. Approximately 76.2% of the cells in the experimental group transfected with RNP containing the guide RNA with the targeting sequence (SEQ ID NO:4 in Table 2) and the donor fragment containing the donor template sequence (SEQ ID NO:17) integrated GFP. In contrast, only approximately 1% of the cells transfected with the donor fragment containing only the donor template sequence (SEQ ID NO:17) (no RNP control group) integrated GFP. Culturing was continued until day 9, and the knock-in efficiency of the experimental group continued to increase, with the integration rate of the GFP gene being approximately 85.6%. Approximately 26.6% of the cells in the experimental group transfected with RNP containing the guide RNA with the targeting sequence (SEQ ID NO:37 in Table 2) and the donor fragment containing the donor template sequence (SEQ ID NO:38) integrated GFP. In contrast, only approximately 0.5% of the cells transfected with the donor fragment containing only the donor template sequence (SEQ ID NO:17) (no RNP control group) expressed GFP. Culturing was continued until day 9, and the knock-in efficiency of the experimental group continued to increase, with the integration rate of the GFP gene being approximately 31.1%.

[0153] Example 3: Integrating Promoters and Cargo Genes by Targeting the Essential Gene Fkbp10

[0154] To determine the feasibility and expression effect of using a single promoter to initiate cargo genes, gene integration was performed in CHO-K1 using gRNA targeting the Fkbp10 (Gene ID: 100771014) gene (the gRNA targeting sequence is shown as SEQ ID NO: 8 in Table 2) and two donor templates (Template 1: the version of T2A peptide linked to GFP, SEQ ID NO: 18; Template 2: the version of promoter linked to GFP, SEQ ID NO: 19). The sequence order in each of the two donor fragments is as follows: 5'-end homologous arm, i.e., partial sequence of the 9th exon (i.e., the penultimate exon) upstream of the cleavage site and partial sequence of the intron upstream thereof; partial sequence of the 9th exon and the 10th exon after codon optimization; exogenous sequence (i.e., cargo gene sequence); 3'-end homologous arm, i.e., partial sequence of the 9th exon and the adjacent intron downstream of the cleavage site. Among them, the version of T2A peptide linked to GFP in the exogenous sequence is the fusion expression of the T2A peptide sequence and the GFP sequence; the version of promoter linked to GFP in the exogenous sequence is the stop codon of Fkbp10, CMV enhancer, CMV promoter, GFP gene, and CaMV polyA signal sequence.

[0155] The RNP concentration during electroporation was 4 μM, and the donor dsDNA fragment encoding GFP was 2.5 μg. The CHO-K1 cells obtained after electroporation were statically cultured in an incubator at 37°C and 5% CO2.

[0156] On the 4th day after electroporation, the proportion of cells expressing green fluorescent protein was detected by flow cytometry and the knock-in (KI) efficiency of the cargo gene encoding GFP integrated into Fkbp10 was detected by RT-qPCR (the primers used are shown in Table 3). The RT-qPCR data showed that the statistically determined knock-in efficiency of the three parallels in the Template 1 experimental group was 53.1 ± 4.6%; the statistically determined knock-in efficiency of the three parallels in the Template 2 experimental group was 56.2 ± 5.1%; the comparison of the green fluorescent protein expression intensity under a fluorescence inverted microscope is shown in Figure 3 , and the fluorescence intensity after integration of Template 2 was higher than that of Template 1, indicating that inserting an additional strong promoter in front of the cargo gene can further promote the expression of the cargo gene.

[0157] Example 4: Integrating Multiple Cargo Genes by Targeting GAPDH

[0158] Use the gRNA targeting the GAPDH gene (the gRNA targeting sequence is shown as SEQ ID NO:1 in Table 2), and replace the cargo gene in the donor template with the gene sequences encoding the heavy chain protein of IgG1 and the gene sequence encoding the light chain protein of IgG1. Specifically, a stop codon is added to the optimized sequence of the GAPDH codon, and the downstream sequences are, in order, the CMV enhancer and the CMV promoter, the gene sequence of the heavy chain protein of IgG1 fused with the secretion signal peptide, the T2A peptide, the gene sequence of the light chain protein of IgG1 fused with the secretion signal peptide, and the CaMV polyA signal sequence. The specific sequence and annotation of this donor template are shown as SEQ ID NO:20.

[0159] The above corresponding RNP and donor fragment were electroporated into CHO-K1 cells cultured in serum-free suspension. The concentration of RNP during electroporation was 4 μM, and the donor dsDNA fragment encoding GFP was 2.5 μg. The CHO-K1 cells obtained after electroporation were statically cultured in an incubator at 37 °C and 5% CO2.

[0160] Four days after electroporation, the IgG integration efficiency was detected by RT-qPCR (the primers used are shown in Table 3), and the targeted integration efficiency was approximately 55%. The monoclonal cell lines with targeted integration were isolated by the limiting dilution method and named GP1 to GP5 respectively. Cells GI1 to GI5 were cultured in shake flasks containing CD CHO medium (Gibco, catalog number 10743002) (37 °C, 100 rpm, 8% CO2). After 48 h of culture, the supernatants were collected, and the monoclonal antibody yields of the above 5 cell lines were as Figure 4 shown. The data indicate that the CHO engineering cells constructed using the present invention have good protein expression and production performance.

[0161] Example 5: Verification of the Integration Efficiency of More Essential Genes

[0162] Use the targeting sequences of the genes shown in Table 2 to target the essential genes Eef1a1, Rplp0, Tubb, Aldoa, Ybx3, and Cdc42 of CHO-K1 respectively. The donor fragments used contain the upstream homologous arm of the cleavage site, the optimized sequence of partial essential gene codons, the T2A peptide, the GFP coding gene, and the downstream homologous arm of the cleavage site. The sequences of the donor fragments corresponding to each site are as follows: Eef1a1 is shown in SEQ ID NO:21; Rplp0 is shown in SEQ ID NO:22; Tubb is shown in SEQ ID NO:23; Aldoa is shown in SEQ ID NO:24; Ybx3 is shown in SEQ ID NO:25; Cdc42 is shown in SEQ ID NO:26.

[0163] According to the known method (A ribonucleoprotein-based decaplex CRISPR / Cas9 knockout strategy for CHO host engineering. Biotechnol Prog. 2022 Jan; 38(1):e3212), the CRISPR / Cas nuclease (Takara, catalog number 632678), the gRNA targeting each gene, and the corresponding DNA donor fragment were transfected into CHO-K1 cells by nucleofection (electroporation) of ribonucleoprotein (RNP). The concentration of RNP during electroporation was 4 μM, and the donor dsDNA fragment encoding GFP was 2.5 μg. The CHO-K1 cells obtained after electroporation were statically cultured in an incubator at 37 °C and 5% CO2. On the 4th day after electroporation, the knock-in (KI) efficiency of the cargo gene encoding GFP integrated into GAPDH was detected. The knock-in efficiencies of Eef1a1, Rplp0, Tubb, Aldoa, Ybx3, and Cdc42 were 56.4 ± 5.7%, 62.9 ± 4.9%, 75.5 ± 6.3%, 46.3 ± 3.9%, 59.1 ± 3.8%, and 51.6 ± 5.7%, respectively. On the 9th day after electroporation, the knock-in efficiency of the cargo gene (the gene encoding GFP) increased, reaching 69.3 ± 5.9%, 78.6 ± 4.7%, 86.9 ± 5.1%, 59.5 ± 4.2%, 69.7 ± 6.8%, and 62.3 ± 4.6%, respectively. The above data indicate that the CHO gene integration method based on essential genes in this patent has a certain generality at multiple essential gene loci.

[0164] SEQ ID NO:15: gRNA scaffold sequence

[0165] GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTT

[0166] SEQ ID NO:16:

[0167] Note: Indicates the partial sequence of the upstream homologous arm of the GAPDH gene cleavage site, exons 5 and 6; Indicates the codon-optimized exons 6 and 7; Indicates the T2A sequence; Indicates the GFP coding sequence; Indicates the partial sequence of the downstream homologous arm of the GAPDH gene cleavage site, exon 6.

[0168]

[0169] SEQ ID NO:17:

[0170] Note: Indicates the upstream homologous arm of the PKM gene break site, and a partial sequence of the upstream intron-exon 11; Indicates codon-optimized exons 11 and 12; Indicates the T2A sequence; Indicates the GFP coding sequence; Indicates the downstream homologous arm of the PKM gene break site, a partial sequence of exon 11 and a partial sequence of the downstream intron.

[0171]

[0172] SEQ ID NO:18: Indicates the upstream homologous arm of the Fkbp10 gene break site; Indicates the codon-optimized sequence; Indicates the T2A sequence; Indicates the GFP coding sequence; Indicates the downstream homologous arm of the Fkbp10 gene break site.

[0173]

[0174] SEQ ID NO:19: Indicates the upstream homologous arm of the Fkbp10 gene break site; Indicates the codon-optimized sequence; the part not underlined indicates the CMV enhancer and promoter sequences; Indicates the GFP coding sequence; Indicates the CaMV polyA signal sequence; Indicates the downstream homologous arm of the Fkbp10 gene

[0175]

[0176] SEQ ID NO:20: Indicates the upstream homologous arm of the GAPDH gene break site; Indicates the codon-optimized sequence; the part not underlined indicates the CMV enhancer and promoter sequences; Indicates the coding sequence of the secretion signal peptide-IgG1 heavy chain-T2A peptide-secretion signal peptide-IgG1 light chain; Indicates the CaMV polyA signal sequence; Indicates the downstream homologous arm of the GAPDH gene.

[0177]

[0178]

[0179] SEQ ID NO:21: Represents the homologous arm upstream of the cleavage site of the Eef1a1 gene; Represents the codon-optimized sequence; Represents the T2A sequence; Represents the GFP coding sequence; Represents the homologous arm downstream of the cleavage site of the Eef1a1 gene.

[0180]

[0181] SEQ ID NO:22: Represents the homologous arm upstream of the cleavage site of the Rplp0 gene; Represents the codon-optimized sequence; Represents the T2A sequence; Represents the GFP coding sequence; Represents the homologous arm downstream of the cleavage site of the Rplp0 gene

[0182]

[0183] SEQ ID NO:23: Represents the homologous arm upstream of the cleavage site of the Tubb gene; Represents the codon-optimized sequence; Represents the T2A sequence; Represents the GFP coding sequence; Represents the homologous arm downstream of the cleavage site of the Tubb gene

[0184]

[0185] SEQ ID NO:24: Represents the homologous arm upstream of the cleavage site of the Aldoa gene; Represents the codon-optimized sequence; Represents the T2A sequence; Represents the GFP coding sequence; Represents the homologous arm downstream of the cleavage site of the Aldoa gene

[0186]

[0187] SEQ ID NO:25: Represents the homologous arm upstream of the cleavage site of the Ybx3 gene; Represents the codon-optimized sequence; Represents the T2A sequence; Represents the GFP coding sequence; Indicates the homologous arm downstream of the Ybx3 gene cleavage site

[0188]

[0189] SEQ ID NO:26: Indicates the homologous arm upstream of the Fkbp10 gene cleavage site; Indicates the codon-optimized sequence; Indicates the T2A sequence; Indicates the GFP coding sequence; Indicates the homologous arm downstream of the Cdc42 gene cleavage site

[0190]

[0191] SEQ ID NO:38: Indicates the homologous arm upstream of the Hsp90b1 gene cleavage site; Indicates the codon-optimized sequence; Indicates the T2A sequence; Indicates the GFP coding sequence; Indicates the homologous arm downstream of the Hsp90b1 gene cleavage site

[0192]

[0193] SEQ ID NO:39: Indicates the homologous arm upstream of the GAPDH gene cleavage site; Indicates the codon-optimized sequence; the un-underlined part indicates the CMV enhancer and promoter sequences; Indicates the GFP coding sequence; Indicates the CaMV polyA signal sequence; Indicates the homologous arm downstream of the GAPDH gene cleavage site

[0194]

[0195] It should be noted that although the technical solutions of the present invention are introduced by specific examples, those skilled in the art can understand that the present invention should not be limited thereto. The above has described the embodiments of the present invention. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and changes are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other ordinary skilled in the art in the technical field to understand the embodiments disclosed herein.

Claims

1. A method for integrating a target gene into the genome of a host cell, the method comprising: a. contacting one or more host cells, which are CHO cells, with the following (i) and (ii): (i) a nuclease capable of generating a break at an essential gene in the host cell, wherein the essential gene encodes a gene product required for cell survival and / or proliferation; and (ii) a donor template, the donor template comprising a knock-in cassette, the knock-in cassette comprising a coding sequence region of a target gene product, and a coding sequence or a partial coding sequence region of an essential gene; the coding sequence or the partial coding sequence region of the essential gene is upstream of the coding sequence region of the target gene product; b. selecting host cells that express (iii) and (iv): (iii) the product of the target gene, and (iv) the gene product encoded by the essential gene required for cell survival and / or proliferation, or a functional variant thereof.

2. The method according to claim 1, wherein, The essential gene is selected from at least one of the genes shown in Table 1; Preferably, the essential gene is selected from at least one of the following genes: GAPDH, Eef1a1, Rplp0, PKM, Tubb, Aldoa, Ybx3, Fkbp10, Cdc42, Cs, Cap1, Nop14, Fdx2, and Ccnd3.

3. The method according to claim 1 or 2, wherein the break is a single-strand break or a double-strand break, preferably a double-strand break; Optionally, the break occurs within any exon of the coding sequence of the essential gene.

4. The method according to any one of claims 1 to 3, wherein the 5' and 3' ends of the knock-in cassette respectively comprise a first homologous region and a second homologous region, which are capable of recombining with a third homologous region and a fourth homologous region respectively, wherein the third homologous region and the fourth homologous region are respectively located on both sides of the break generated in the essential gene.

5. The method according to any one of claims 1 to 4, wherein The coding sequence region of the target gene product comprises the coding sequence of at least one target gene product; Optionally, the coding sequence between the coding sequences of the target gene products comprises the coding sequence of a signal peptide.

6. The method according to claim 5, wherein The knock-in cassette comprises gene product expression regulatory elements capable of separating the gene products encoded by multiple essential genes, and regulatory elements for expressing the gene products encoded by the essential genes and the target gene product as separate gene products; optionally, at least one of the gene products is a protein and the regulatory element enables the protein to be expressed separately from other gene products; Optionally, the regulatory element comprises an IRES, a 2A element, or a promoter; Optionally, the 2A element is selected from at least one of the T2A, P2A, E2A, and F2A elements; Optionally, the promoter is selected from at least one of the CMV promoter, EF-1α promoter, SV40 promoter, CHEF-1α promoter, and GAC promoter.

7. The method according to claim 6, wherein, The partial coding sequence of the essential gene in the knock-in cassette encodes the C-terminal fragment of the protein encoded by the essential gene; the C-terminal fragment comprises the amino acid sequence encoded by the coding sequence of the essential gene spanning the break.

8. The method according to any one of claims 1 to 7, wherein, The nuclease is selected from the group consisting of CRISPR / Cas nucleases, zinc finger nucleases, TAL-effector DNA binding domain-nuclease fusion proteins, transposases, and site-specific recombinases; optionally, the nuclease comprises a Cas9 endonuclease.

9. The method according to any one of claims 1 to 8, wherein the host cell is contacted with at least one nuclease and at least one donor template so that integration of a target gene is achieved at at least one site in the host cell genome.

10. The method according to any one of claims 1 to 9, wherein the target gene comprises any gene that can be expressed in CHO cells.

11. The method according to any one of claims 1 to 10, further comprising the step of recovering the host cell, wherein, The donor template has undergone homologous recombination at the break generated in the essential gene.

12. A host cell, comprising: An insertion cassette is inserted into the coding sequence of an essential gene in the genome of the host cell, wherein the essential gene encodes a gene product required for the survival and / or proliferation of the cell, and the insertion cassette comprises a coding sequence region of a target gene product and a coding sequence or a partial coding sequence region of the essential gene; the coding sequence or the partial coding sequence region of the essential gene is downstream of the coding sequence region of the target gene product; Optionally, the expression of the target gene product and the gene product encoded by the essential gene is initiated by the endogenous promoter of the essential gene; Preferably, the host cell is a CHO cell; Preferably, the host cell expresses: the product of the target gene, and the gene product encoded by the essential gene required for the survival and / or proliferation of the cell, or a functional variant thereof; Optionally, the essential gene comprises the genes shown in Table 1; Optionally, the target gene comprises any gene that can be expressed in CHO cells.

13. The host cell according to claim 12, wherein, The host cell comprises regulatory elements capable of expressing the gene product encoded by the essential gene and the target gene product as separate gene products; optionally, at least one of the gene products is a protein and the regulatory elements enable the protein to be expressed separately from other gene products; Optionally, the regulatory elements comprise an IRES, a 2A element, and / or a promoter; Optionally, the 2A element is selected from at least one of T2A, P2A, E2A, and F2A elements; Optionally, the promoter is selected from at least one of CMV promoter, EF-1α promoter, SV40 promoter, CHEF-1α promoter, and GAC promoter.

14. A pharmaceutical composition, comprising: a host cell prepared by the method according to any one of claims 1 to 11, or a host cell according to any one of claims 12 or 13; and a pharmaceutically acceptable carrier or excipient.

15. A kit, comprising: a host cell prepared by the method according to any one of claims 1 to 11, or a host cell according to any one of claims 12 or 13, or the pharmaceutical composition according to claim 14.