Target integration into mammalian sequences to enhance gene expression

Site-specific integration of transgenes into ERVs or LTR-RTs in mammalian cells stabilizes expression and reduces variability, addressing the limitations of random integration and epigenetic effects, thereby improving therapeutic protein production.

JP7843231B2Active Publication Date: 2026-04-09SELEXIS SA
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2020-12-24
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Current methods for integrating transgenes into mammalian cells suffer from low and unstable expression due to random integration and epigenetic effects, leading to clonal variability and the need for extensive screening to identify high-expression cell clones.

Method used

Site-specific targeting of transgenes into endogenous retroviral sequences (ERVs) or LTR-retrotransposons (LTR-RTs) in the mammalian genome, modulating the host's DNA repair pathway to achieve stable and high-level expression, using engineered cells with insertion sites at ERV or LTR-RT loci.

Benefits of technology

Ensures high and stable transgene expression rates with identical genomic configurations, reducing the need for clone screening and minimizing viral particle release, while enhancing therapeutic protein production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007843231000024
    Figure 0007843231000024
  • Figure 0007843231000025
    Figure 0007843231000025
  • Figure 0007843231000026
    Figure 0007843231000026
Patent Text Reader

Abstract

Disclosed herein is a cell in which exogenous nucleic acid sequence (for example, transgene) is stably integrated into genome, and the method for producing and using such cell, in or near the integration site of the sequence that comprises at least a part of endogenous retrovirus (ERV) or LTR-retrotransposon (LTR-RT), or alternatively, the sequence that comprises ERV or LTR-RT that is or was part of the genome of cell.Advantageously, high level and / or stable production of transgene expression product can be achieved.The integration and expression of transgene can be promoted by adjusting the DNA repair pathway of cell, for example, by transiently expressing the gene that codes for the protein that forms part of DNA repair pathway during the integration of transgene.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Expression of recombinant proteins in mammalian cells is very important for the biotechnological production of recombinant proteins and / or for therapeutic uses such as gene therapy and cell therapy. Generation of each cell line requires successful integration of the transgene into the host genome and its expression in the cells. Currently, the mainstream strategies for cell line development are: i) random integration of the transgene into the chromosomes of the cells, ii) selection of cells into which the transgene has been integrated, and iii) selection of specific cells exhibiting optimal production capacity characteristics. However, this approach is limited by the number of transgene copies integrated and by the epigenetic effects due to the genomic environment of the transgene, often causing low and unstable transcription and / or high clonal variability.

[0002] In general, to overcome these problems associated with cell line development, epigenetic regulators can be used to protect the transgene from negative position effects (Bell and Felsenfeld, 1999). These epigenetic regulators include boundary or insulator elements, locus control regions (LCRs), stabilizing and anti-repressor (STAR) elements, ubiquitous chromatin opening elements (UCOEs), and matrix attachment regions (MARs). All of these epigenetic regulators have been used in recombinant protein production (Zahn-Zabal et al., 2001, Kim et al., 2004) and gene therapy (Agarwal et al., 1998, Castilla et al., 1998) in mammalian cell lines.

[0003] Publications and other materials, including patents and patent applications, used herein to illustrate the present invention and, more specifically, to provide additional details relating to the implementation of the present invention, are incorporated herein by reference in their entirety. For convenience, publications are referred below either by number, by author and year of publication, or by patent / publication number, with reference to the attached bibliography. [Overview of the project] [Problems that the invention aims to solve]

[0004] In particular, site-specific targeting of transgenes is needed, which is suitable for increasing and stabilizing transgene expression in mammalian cells. Site-specific targeting of transgenes is also needed because it advantageously results in cells with identical genomic configurations, eliminating the need to screen many cell clones to identify and select those with high levels of transgene expression. Suitable insertion sites, referred to as "landing pads," for specific subclasses of transgenes are also needed. These insertion sites advantageously ensure stability of the cell line into which the transgene is incorporated and high long-term expression rates, even from a single transgene or a low copy number. Therefore, identifying and validating suitable insertion sites for transgenes in mammalian cells is particularly needed in the art for efficient and reliable transgene expression. Cell clones used for therapeutic protein production that lack expression of endogenous retroviral sequences (ERVs) and / or do not release viral particles into the cell supernatant, or release them to a low degree, along with the therapeutic protein produced, are also needed. One or more of the above needs, as well as other needs, are addressed herein. [Means for solving the problem]

[0005] In particular, the stable integration of exogenous nucleic acid sequences, such as transgenes, within or near at least a portion of the insertion sequences of endogenous retroviral sequences (ERVs) or LTR-retrotransposons (LTR-RTs) in the mammalian genome is disclosed. In certain embodiments, this results in the production of high levels and / or stable transgene expression products. In certain embodiments, this is achieved and / or facilitated by modulating the host cell's DNA repair pathway.

[0006] At least one gene locus containing an insertion site for an ERV sequence or an LTR-RT sequence, Engineered cells, preferably mammalian cell lines, are disclosed, including engineered CHO cells, engineered pig cells, or engineered human cells, which include a transgene incorporated into at least one locus, and which include the genome of the cell.

[0007] At least one locus into which the transgene is incorporated, and by optional selection, the corresponding allelic locus, may be an ERV sequence or an LTR-RT sequence insertion locus.

[0008] At least one locus into which the transgene is incorporated may be a locus corresponding to the wild-type allele (e.g., ERV-deficient or LTR-RT-deficient) of an ERV sequence insertion locus or an LTR-RT sequence insertion locus (e.g., an ERV-integrated or LTR-RT-integrated genomic sequence). The transgene may also be incorporated adjacent to or replacing the corresponding ERV sequence or LTR-RT sequence at the insertion locus. The transgene may also be incorporated into either or both loci in more than 20%, more than 30%, or even more than 40% of the transgene-containing cells in a cell population. The locus may be homozygous and may contain, for example, at least two copies of SEQ ID NO: 1 or SEQ ID NO: 2, or parts thereof. The locus may be heterozygous and may contain, for example, both SEQ ID NO: 1 and SEQ ID NO: 2, or parts thereof.

[0009] In particular embodiments, Within the cell's genome, At least one gene locus containing an insertion site for an endogenous retrovirus (ERV) sequence or an LTR-retrotransposon (LTR-RT) sequence, (i) a) ERV sequence or LTR-RT sequence, or b) insertion site, and optionally a portion of the sequence in a), and / or (ii)(i) Allele wild-type corresponding sequence including at least one gene locus, At least one transgene encoding at least one transgene expression product incorporated into at least one gene locus and The subject is engineered cells, preferably mammalian cell lines, including engineered CHO cells, engineered pig cells, or engineered human cells, which include engineered CHO-K1 cells. The cells may contain (i) and (ii) on different chromosomes, for example, chromosome 15 and chromosome 9 of CHO cells. The cells may contain (i) and (ii), and at least one transgene may be incorporated into (ii) and not into (i), or may be incorporated into (i) as well.

[0010] Certain embodiments relate to a population of cells including the engineered cells described herein. At least one transgene may be incorporated into more than 20%, more than 30%, or even more than 40% of the cells (i) and / or (ii) in the cell population. The engineered cells of the cell population may include (i) and (ii) as described above. (i) may include at least nucleotides 29021-40247 (or 29521-39747) of SEQ ID NO: 1, or sequences having 95%, 98%, or 99% sequence identity with nucleotides 29021-40247 (or 29521-39747) of SEQ ID NO: 1, and (ii) may include at least nucleotides 29020-31020 (or 29520-30520) of SEQ ID NO: 2, or sequences having 95%, 98%, or 99% sequence identity with nucleotides 29020-31020 (or 29520-30520) of SEQ ID NO: 2. The manipulated cells / cell population lacks expression of endogenous retroviral sequences (ERVs) in certain preferred embodiments. In certain preferred embodiments, no detectable viral particles are present in the culture supernatant of the cell population.

[0011] At least one transgene expression product may be the product of interest, e.g., the protein of interest. A cell / cell population may, by arbitrary selection, express the product of interest / protein of interest at a rate (e.g., picograms per cell per day, μg / l, or mg / l) at least 1.5 times, 2 times, 2.5 times, 3 times, or more than the amount of the product of interest / protein of interest per unit time that would be produced if at least one transgene were integrated into the genome outside at least one locus.

[0012] ERV or LTR-RT includes endogenous retroviral elements of type C (ERV C), MLV (mouse leukemia virus), XMRV (heterotropic mouse leukemia virus-associated virus), MMTV (mouse mammary tumor virus), MERV-L (mouse ERV with L-tRNA PBS), VL30 (virus-like 30), IAP (intraciscan A particle), MusD (Mus type D-associated retrovirus), PERV (porcine endogenous retrovirus), KoRV (koala retrovirus), enJSRV (Yagsikte sheep retrovirus), MaLR (mammalian apparent LTR retrotransposon), HERV (human endogenous retrovirus), for example, HERV-E (human ERV with E-tRNA PBS), HERV-H (human ERV with H-tRNA PBS), HERV-K (human ERV with K-tRNA PBS), HERV-L (L-tRNA The group may be selected from human ERVs with PBS, HERV-W (human ERVs with W-tRNA PBS), and combinations thereof.

[0013] An ERV or LTR-RT sequence may include at least one ERV subsequence selected from the group consisting of a gag (group-specific antigen) gene, a pol (polymerase) gene, an env (envelope) gene, a MA (matrix), a CA (capsid), an NC (nucleocapsid), a SP1 (spacer peptide 1), a SP2 (spacer peptide 2), or further domains encoding proteins such as pp12 or p6, and a long terminal repeat (LTR) of the ERV, as well as combinations thereof, and the transgene is optionally incorporated into one of the subsequences.

[0014] Cells may be transfected with one or more vectors containing one or more genes from Table 2, or sequences having at least 80%, 85%, 90%, 95%, 98%, or 99% sequence identity with sequence numbers 25-28, 38-58, and / or 59, and the cytoplasm of the cells may be optionally transfected with one or more exogenous chemical inhibitors and / or stimulants of DNA repair pathways (DRPs), such as NU7441, olaparib, a DNA ligase IV inhibitor, Scr7, KU-0060648, anti-EGFR antibody C225 (cetuximab), compound 401(2 The following may further be included: NHEJ inhibitors, myrin, myrin derivatives, PolQ inhibitors, CtIP inhibitors, and combinations thereof, selected from the group consisting of (4-morpholinyl)-4H-pyrimide[2,la]isoquinoline-4-one, vanillin, waltmannin, DMNB, IC87361, LY294002, OK-1035, CO15, NK314, PI103 hydrochloride, and combinations thereof; MMEJ inhibitors, HR inhibitors, e.g., RI-1 and BO2, HR stimulants, e.g., RS-1, NHEJ stimulants, e.g., IP6, and any combination of any one of the above inhibitors and / or stimulants. In any of the manipulated cells, the locus may have at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with a sequence selected from SEQ ID NO: 1 and / or SEQ ID NO: 2. The introduced gene may be a landing pad. The manipulated cells are, in certain preferred embodiments, Chinese hamster ovary (CHO) cells, or human or porcine cells.

[0015] One embodiment is a method for incorporating a transgene into the genome of a cell, preferably a mammalian cell line, (a) Providing at least one transgene as part of a vector, for example, a plasmid or viral vector, wherein the vector integrates the transgene into at least one locus in a cell that includes an insertion site for an endogenous retrovirus (ERV) sequence or an LTR-retrotransposon (LTR-RT) sequence, or (b) to provide, optionally, at least one transgene and / or nickase as part of a vector, wherein the nuclease and / or nickase is preferably encoded by at least one vector, and the nuclease and / or nickase introduces double-strand and / or single-strand breaks into at least one locus of a cell containing an insertion site for an endogenous retrovirus (ERV) sequence or an LTR-retrotransposon (LTR-RT) sequence in order to incorporate the transgene, and further optionally, to provide at least one vector encoding at least one target element that guides the at least one nuclease and / or nickase. Optionally, upmodulate, in particular stimulate, at least one primary DNA repair pathway (DRP) in a cell, and optionally, downmodulate, in particular stimulate, at least one secondary DRP in a cell, or vice versa. Transfecting cells with at least one transgene, By arbitrary selection, the engineered cells containing the transgene incorporated into the gene locus are isolated, Includes methods.

[0016] Furthermore, cells / cell lines may be transfected with one or more genes from Table 2, or sequence numbers 25-28, 38-58, and / or 59, or sequences having at least 80%, 85%, 90%, 95%, 98%, or 99% sequence identity with sequence numbers 25-28, 38-58, and / or 59, preferably as part of one or more further vectors, and / or cells / cell lines may be exposed to chemicals that affect the cell's DNA repair pathway (DRP). Cells may also be transfected with one or more further vectors that include and express sequence numbers 25-28, 38-58, and / or 59, preferably sequence numbers 25-28.

[0017] At least one nuclease and / or nickas is a transposase, integrase, recombinase, e.g., site-specific recombinase, nickas, or nuclease, e.g., site-specific nuclease, a fusion protein containing a programmable DNA-binding domain and a nuclease domain, or any combination thereof, or a homing end nuclease, restriction enzyme, zinc finger nuclease or zinc finger nickas, meganuclease or mega nickas, transcription activator-like effector nuclease or transcription activator-like effector nickas, RNA guide nuclease or RNA guide Recombinases may include donicasse, DNA guide nuclease or DNA guide nickases, megaTAL nuclease, BurrH nuclease, ARCUS nuclease, modified or chimeric versions or variants thereof, and any combination thereof, in particular zinc finger nuclease or zinc finger nickases, transcription activator-like effector nuclease or transcription activator-like effector nickases, RNA guide nuclease or RNA guide nickases, and RNA guide nuclease or RNA guide nickases may, optionally, be part of CRISPR-based systems, restriction enzymes, and combinations thereof. Recombinases may include Cre recombinase, FLP recombinase, lambda integrase, PhiC31 integrase, Dre recombinase, xb1 integrase, gamma delta resolverase, R4 integrase, Tn3 resolverase, or TP901-1 recombinase. In certain embodiments, the nuclease is a transcription activator-like effector nuclease or an RNA guide nuclease.

[0018] The vectors used herein may be plasmids or viral vectors, such as AAV vectors.

[0019] The first and / or second DRPs include excision, canonical homology-directed repair (canonical HDR), homologous recombination (HR), alternative homology-directed repair (Alt-HDR), double-strand break repair (DSBR), single-strand annealing (SSA), synthesis-dependent strand annealing (SDSA), break-induced replication (BIR), alternative end joining (Alt-EJ), microhomology-mediated end joining (MMEJ), and DNA synthesis-dependent microhomology-mediated end joining. The following can be selected from the group consisting of end joining (SD-MMEJ), canonical non-homologous end joining repair (C-NHEJ), alternative non-homologous end joining (A-NHEJ), trans-damage DNA synthesis repair (TLS), base excision repair (BER), nucleotide excision repair (NER), mismatch repair (MMR), DNA damage response (DDR), blunt end joining, single-strand break repair (SSBR), interstrand crosslink repair (ICL), Fanconi anemia (FA) pathway, and combinations thereof.

[0020] In certain embodiments, at least one first DRP is homologous recombination (HR) and at least one second DRP is one or more non-homologous end-joining (NHEJ) DNA repair pathways; at least one first DRP may be an Alt-EJ pathway, e.g., MMEJ, and at least one second DRP may be one or more non-homologous end-joining (NHEJ) DNA repair pathways; at least one first DRP may be an Alt-EJ pathway, e.g., MMEJ, and at least one second DRP may be a homologous recombination (HR) DNA repair pathway; or at least one first DRP may be an Alt-EJ pathway, e.g., MMEJ, and at least one second DRP may be one or more alternative DNA repair pathways.

[0021] The interference / alteration of DRP may be an upregulation thereof and may take the form of a) causing and expressing an overexpression of at least one component of DRP in the cell, including, b) introducing at least one component of the DRP into the cell, and / or c) contacting the cell with at least one stimulant of a component of DRP, such as a chemical stimulant, such as an HR stimulant, such as RS-1 and / or an NHEJ stimulant, such as IP6.

[0022] The interference / alteration may also be a downregulation and may take the form of a) contacting the cell with at least one inhibitor of a component of DRP, such as a chemical inhibitor, such as NU7441, olaparib, DNA ligase IV inhibitor, Scr7, KU-0060648, anti-EGFR antibody C225 (cetuximab), compound 401 (2-(4-morpholinyl)-4H-pyrimido[2,1-a]isoquinolin-4-one), vanillin, wortmannin, DMNB, IC87361, LY294002, OK-1035, CO15, NK314, PI103 hydrochloride, and combinations thereof, an NHEJ inhibitor, mirin, a derivative of mirin, an inhibitor of PolQ, an inhibitor of CtIP, and combinations thereof, an MMEJ inhibitor, an HR inhibitor, such as RI-1 and / or BO2, b) inactivating or downregulating at least one component of DRP by contacting the cell with at least one inhibitory nucleic acid, such as miRNA, siRNA, shRNA, or expressing it in the cell, and / or c) expressing a protein that inhibits the DRP in the cell, or any combination thereof may take the form of.

[0023] One embodiment includes engineered cells produced by one of the methods disclosed herein.

[0024] Another embodiment is a kit for introducing at least one transgene into a cell, wherein One container contains at least one locus including an endogenous retrovirus (ERV) sequence or an LTR-retrotransposon (LTR-RT) sequence, for example, the insertion sites of SEQ ID NOs: 1 and 2, preferably (i) a sequence having 95%, 98%, or 99% sequence identity with nucleotides 29021-40247 of SEQ ID NO: 1, or (ii) a locus containing at least nucleotides 29020-31020 of SEQ ID NO: 2, or a sequence having 95%, 98%, or 99% sequence identity with nucleotides 29020-31020 of SEQ ID NO: 2, or nucleotides 29521-39747 of SEQ ID NO: 1. (ii) a vector encoding a nuclease and / or nickase that targets a locus, for example, an ERV sequence or LTR-RT sequence, for example, SEQ ID NO: 3, incorporated into an insertion site, which includes (ii) a sequence having 95%, 98%, or 99% sequence identity with nucleotides 29520 to 30520 of SEQ ID NO: 2, or nucleotides 29520 to 30520 of SEQ ID NO: 2, and optionally, at least one vector encoding at least one target element that guides the at least one nuclease and / or nickase. Optionally, in a separate container, at least one vector encoding at least one target element that guides the at least one nuclease and / or nickase, In a separate container, at least one stimulant and / or inhibitor of the DNA repair pathway (DRP), and / or One or more genes encoding one or more of the DRP proteins in Table 2, or one or more vectors containing sequences having at least 80%, 85%, 90%, 95%, 98%, or 99% sequence identity with SEQ ID NOs. 25-28, 38-58, and / or 59, Furthermore, instructions for a method of transfecting cells with an introduced gene using at least one nuclease and / or nicasse, and at least one stimulant and / or inhibitor. Includes, includes kit. [Brief explanation of the drawing]

[0025] [Figure 1] Figure 1 is a schematic diagram of a gene locus containing an integration site, showing one integration site containing the ERV sequence, in this case the allele containing ERV C 109F (SEQ ID NO: 3), and its wild-type counterpart allele. [Figure 2A] Figure 2A shows a dual CRISPR-based approach to target the ERV C 109F locus and the wild-type allele for transgene integration. [Figure 2B] Figure 2B shows a dual CRISPR-based approach to target the ERV C 109F locus and the ERV C 109F allele for transgene integration. [Figure 3A] Figure 3A shows vectors used for targeted integration into the ERV C 109F allele with or without stimulation of the Alt-EJ repair pathway, using vectors that do not contain homologous sequences. [Figure 3B] Figure 3B shows vectors used for targeted integration into the wild-type allele (Figure 3B) with or without stimulation of the Alt-EJ repair pathway, using vectors that do not contain homologous sequences. [Figure 4A] Figure 4A shows vectors used for targeted integration into the ERV C 109F allele with or without stimulation of the HR repair pathway, using vectors containing homologous sequences. [Figure 4B] Figure 4B shows vectors used for targeted integration into wild-type alleles with or without stimulation of the HR repair pathway, using vectors containing homologous sequences. [Figure 5A]Figure 5A shows the development of a TaqMan qPCR assay for the wild-type allele. The thick horizontal lines enclosed in the box indicate the amplicon locations in the TaqMan qPCR assay. [Figure 5B] Figure 5B shows the development of the TaqMan qPCR assay for the ERV C 109F allele. The thick horizontal lines enclosed in the box indicate the amplicon locations in the TaqMan qPCR assay. [Figure 6] Figure 6 shows the percentage of clones with or without DPR stimulation, with or without DNA homology sequences, including both alleles, the wild-type allele, the ERV C 109F allele, or the integration of the transgene at a random locus. [Figure 7A] Figures 7A, 7B, and 7C show fluorescence in situ hybridization using transgene probes for the detection of integration events at the ERV locus. Notably, due to chromosomal rearrangements, a single "ERV locus" can be found on two different chromosomes (chromosome 15 and chromosome 9) in CHO cells. [Figure 7B] Figures 7A, 7B, and 7C show fluorescence in situ hybridization using transgene probes for the detection of integration events at the ERV locus. Notably, due to chromosomal rearrangements, a single "ERV locus" can be found on two different chromosomes (chromosome 15 and chromosome 9) in CHO cells. [Figure 7C] Figures 7A, 7B, and 7C show fluorescence in situ hybridization using transgene probes for the detection of integration events at the ERV locus. Notably, due to chromosomal rearrangements, a single "ERV locus" can be found on two different chromosomes (chromosome 15 and chromosome 9) in CHO cells. [Figure 8A] Figure 8A shows GFP fluorescence measurements for each type of cell clone: ​​no DNA homology / Alt-EJ stimulation (compared to Figure 6). [Figure 8B]Figure 8B shows GFP fluorescence measurements for each type of cell clone: ​​no DNA homology / no Alt-EJ stimulation (compared to Figure 6). [Figure 8C] Figure 8C shows GFP fluorescence measurements for each type of cell clone: ​​DNA homology / HR stimulation (Figure 8C) (compared to Figure 6). [Figure 8D] Figure 8D shows GFP fluorescence measurements for each type of cell clone: ​​DNA homology / no HR stimulation (Figure 8D) (compared to Figure 6). [Figure 9] Figure 9 shows a comparison of GFP fluorescence results obtained for clones using targeted integration with DNA sequence homology in transgene vectors, comparing HR mechanism regulation with the absence of HR mechanism regulation. [Figure 10A] Figures 10A, 10B, and 10C illustrate the integration sites of the type C ERV 109F sequence on chromosome 15 (Figure 10A) and its wild-type allele locus on chromosome 9 (Figure 10B). CRISPR break sites are indicated by triangles below the DNA sequence in the upper panel, or by boxes at the bottom of Figure 10C. [Figure 10B] Figures 10A, 10B, and 10C illustrate the integration sites of the type C ERV 109F sequence on chromosome 15 (Figure 10A) and its wild-type allele locus on chromosome 9 (Figure 10B). CRISPR break sites are indicated by triangles below the DNA sequence in the upper panel, or by boxes at the bottom of Figure 10C. [Figure 10C] Figures 10A, 10B, and 10C illustrate the integration sites of the type C ERV 109F sequence on chromosome 15 (Figure 10A) and its wild-type allele locus on chromosome 9 (Figure 10B). CRISPR break sites are indicated by triangles below the DNA sequence in the upper panel, or by boxes at the bottom of Figure 10C. [Figure 11A] Figure 11A shows the vectors used for transfection to generate trastuzumab-producing CHO cell clones. On the left side of the figure, the co-transfected vectors for transient expression are shown. [Figure 11B] Figure 11B shows the vectors used for transfection to generate trastuzumab-producing CHO cell clones. In the figure, the immunoglobulin (Ig) expression vectors, i.e., vectors containing the Tras_Hc (heavy chain) and Tras_Lc (light chain) sequences, and their integration sites on chromosome 15 and chromosome 9, respectively, are shown. [Figure 12A] Figure 12A is a schematic diagram of four types of clones characterized after targeted integration of the transgene at the ERV109F genomic locus in CHO-M cells. The transgene is integrated into both locus alleles (referred to as “integration into both loci” or simply “both” in the figure) (ERV109F allele on chromosome 15 and wild-type allele on chromosome 9). [Figure 12B] Figure 12B is a schematic diagram of four types of clones characterized after targeted integration of the transgene at the ERV109F genomic locus in CHO-M cells. The transgene is integrated into the ERV109F allele on chromosome 15, but not into the wild-type allele on chromosome 9 (referred to as "integration into the ERV locus" or simply "ERV" in the figure). [Figure 12C] Figure 12C is a schematic diagram of four types of clones characterized after targeted integration of the transgene at the ERV109F genomic locus in CHO-M cells. The transgene is integrated into the wild-type allele on chromosome 9, but not into the ERV109F allele on chromosome 15 (referred to in the figure as "integration into a locus without ERV" or simply "ERV deficiency"). [Figure 12D] Figure 12D is a schematic diagram of four types of clones characterized after targeted integration of the transgene at the ERV109F genomic locus in CHO-M cells. In these clones, the transgene is not integrated into any allele at the locus but is randomly integrated into the host chromosome. The gray arrows represent the PCR primers used to characterize the genomic integration sites of the cloned transgenes. [Figure 13A]Figure 13A shows the multiplier of decrease in ERV C 109F expression after transgene integration, compared to parental CHO-M cells. Total RNA was extracted from the cell clones shown, viral RNA levels were determined by RT-qPCR, and processed using the delta-delta Ct (cycle threshold) calculation method. The hatching corresponds to the hatching shown in Figures 12A-12C. The figures show cultures with low titers. [Figure 13B] Figure 13B shows the multiplier of decrease in ERV C 109F expression after transgene integration, compared to parental CHO-M cells. Total RNA was extracted from the cell clones shown, viral RNA levels were determined by RT-qPCR, and processed using the delta-delta Ct (cycle threshold) calculation method. The hatching corresponds to the hatching shown in Figures 12A-12C. The figures show cultures with moderate titers. [Figure 13C] Figure 13C shows the multiplier of decrease in ERV C 109F expression after transgene integration, compared to parental CHO-M cells. Total RNA was extracted from the cell clones shown, viral RNA levels were determined by RT-qPCR, and processed using the delta-delta Ct (cycle threshold) calculation method. The hatching corresponds to the hatching shown in Figures 12A-12C. The figure shows cultures with high titers. [Figure 14A] Figure 14A shows the determination of trastuzumab production levels in the supernatants of various CHO cell clone types. For clone types possessing the Tras expression construct integrated into the indicated genomic "locus" (allele), ELISA assays were performed on cell-free supernatants of cultures prepared in 96-well plates for 3 days. The panel shows the titers of trastuzumab antibodies obtained from cultures exhibiting low levels of production for each clone type. [Figure 14B]Figure 14B shows the determination of trastuzumab production levels in the supernatants of various CHO cell clone types. For clone types possessing the Tras expression construct integrated into the indicated genomic "locus" (allele), ELISA assays were performed on cell-free supernatants of cultures prepared in 96-well plates for 3 days. The panel shows the titers of trastuzumab antibodies obtained from cultures exhibiting moderate levels of production for each clone type. [Figure 14C] Figure 14C shows the determination of trastuzumab production levels in the supernatants of various CHO cell clone types. For clone types possessing the Tras expression construct integrated into the indicated genomic "locus" (allele), ELISA assays were performed on cell-free supernatants of cultures prepared in 96-well plates for 3 days. The panel shows the titers of trastuzumab antibodies obtained from cultures exhibiting high levels of production for each clone type. [Figure 15A]Figures 15A-F provide measurements of clonal trastuzumab production capacity in 10-day fed-batch cultures in 24-deep-well plates (Figures 15A (low titer), 15B (moderate titer), and 15C (high titer)) and 3 ml of culture medium with 300,000 cells / ml of starting material, or in cultures performed in 96-well plates (Figures 15D (low titer), 15E (moderate titer), and 15F (high titer)). Tras protein titer measurement is performed using the LabChip LCGXII system (registered trademark) (PERKIN ELMER, The work was performed using (Inc.). Hatching corresponds to the hatching notation shown in Figures 13A-13C and 14A-14C: In each graph, from left to right: the transgene is integrated into both alleles of the locus ("both") (ERV109F allele on chromosome 15 and wild-type allele on chromosome 9); the transgene is integrated into the ERV109F allele on chromosome 15 but not into the wild-type allele on chromosome 9 ("ERV"); the transgene is not integrated into either allele of the locus but is randomly integrated into the host chromosome, or the transgene is integrated into the wild-type allele on chromosome 9 but not into the ERV109F allele on chromosome 15 ("ERV-deficient"). [Figure 15B]Figures 15A-F provide measurements of clonal trastuzumab production capacity in 10-day fed-batch cultures in 24-deep-well plates (Figures 15A (low titer), 15B (moderate titer), and 15C (high titer)) and 3 ml of culture medium with 300,000 cells / ml of starting material, or in cultures performed in 96-well plates (Figures 15D (low titer), 15E (moderate titer), and 15F (high titer)). Tras protein titer measurement is performed using the LabChip LCGXII system (registered trademark) (PERKIN ELMER, The work was performed using (Inc.). Hatching corresponds to the hatching notation shown in Figures 13A-13C and 14A-14C: In each graph, from left to right: the transgene is integrated into both alleles of the locus ("both") (ERV109F allele on chromosome 15 and wild-type allele on chromosome 9); the transgene is integrated into the ERV109F allele on chromosome 15 but not into the wild-type allele on chromosome 9 ("ERV"); the transgene is not integrated into either allele of the locus but is randomly integrated into the host chromosome, or the transgene is integrated into the wild-type allele on chromosome 9 but not into the ERV109F allele on chromosome 15 ("ERV-deficient"). [Figure 15C]Figures 15A-F provide measurements of clonal trastuzumab production capacity in 10-day fed-batch cultures in 24-deep-well plates (Figures 15A (low titer), 15B (moderate titer), and 15C (high titer)) and 3 ml of culture medium with 300,000 cells / ml of starting material, or in cultures performed in 96-well plates (Figures 15D (low titer), 15E (moderate titer), and 15F (high titer)). Tras protein titer measurement is performed using the LabChip LCGXII system (registered trademark) (PERKIN ELMER, The work was performed using (Inc.). Hatching corresponds to the hatching notation shown in Figures 13A-13C and 14A-14C: In each graph, from left to right: the transgene is integrated into both alleles of the locus ("both") (ERV109F allele on chromosome 15 and wild-type allele on chromosome 9); the transgene is integrated into the ERV109F allele on chromosome 15 but not into the wild-type allele on chromosome 9 ("ERV"); the transgene is not integrated into either allele of the locus but is randomly integrated into the host chromosome, or the transgene is integrated into the wild-type allele on chromosome 9 but not into the ERV109F allele on chromosome 15 ("ERV-deficient"). [Figure 15D]Figures 15A-F provide measurements of clonal trastuzumab production capacity in 10-day fed-batch cultures in 24-deep-well plates (Figures 15A (low titer), 15B (moderate titer), and 15C (high titer)) and 3 ml of culture medium with 300,000 cells / ml of starting material, or in cultures performed in 96-well plates (Figures 15D (low titer), 15E (moderate titer), and 15F (high titer)). Tras protein titer measurement is performed using the LabChip LCGXII system (registered trademark) (PERKIN ELMER, The work was performed using (Inc.). Hatching corresponds to the hatching notation shown in Figures 13A-13C and 14A-14C: In each graph, from left to right: the transgene is integrated into both alleles of the locus ("both") (ERV109F allele on chromosome 15 and wild-type allele on chromosome 9); the transgene is integrated into the ERV109F allele on chromosome 15 but not into the wild-type allele on chromosome 9 ("ERV"); the transgene is not integrated into either allele of the locus but is randomly integrated into the host chromosome, or the transgene is integrated into the wild-type allele on chromosome 9 but not into the ERV109F allele on chromosome 15 ("ERV-deficient"). [Figure 15E]Figures 15A-F provide measurements of clonal trastuzumab production capacity in 10-day fed-batch cultures in 24-deep-well plates (Figures 15A (low titer), 15B (moderate titer), and 15C (high titer)) and 3 ml of culture medium with 300,000 cells / ml of starting material, or in cultures performed in 96-well plates (Figures 15D (low titer), 15E (moderate titer), and 15F (high titer)). Tras protein titer measurement is performed using the LabChip LCGXII system (registered trademark) (PERKIN ELMER, The work was performed using (Inc.). Hatching corresponds to the hatching notation shown in Figures 13A-13C and 14A-14C: In each graph, from left to right: the transgene is integrated into both alleles of the locus ("both") (ERV109F allele on chromosome 15 and wild-type allele on chromosome 9); the transgene is integrated into the ERV109F allele on chromosome 15 but not into the wild-type allele on chromosome 9 ("ERV"); the transgene is not integrated into either allele of the locus but is randomly integrated into the host chromosome, or the transgene is integrated into the wild-type allele on chromosome 9 but not into the ERV109F allele on chromosome 15 ("ERV-deficient"). [Figure 15F]Figures 15A-F provide measurements of clonal trastuzumab production capacity in 10-day fed-batch cultures in 24-deep-well plates (Figures 15A (low titer), 15B (moderate titer), and 15C (high titer)) and 3 ml of culture medium with 300,000 cells / ml of starting material, or in cultures performed in 96-well plates (Figures 15D (low titer), 15E (moderate titer), and 15F (high titer)). Tras protein titer measurement is performed using the LabChip LCGXII system (registered trademark) (PERKIN ELMER, The work was performed using (Inc.). Hatching corresponds to the hatching notation shown in Figures 13A-13C and 14A-14C: In each graph, from left to right: the transgene is integrated into both alleles of the locus ("both") (ERV109F allele on chromosome 15 and wild-type allele on chromosome 9); the transgene is integrated into the ERV109F allele on chromosome 15 but not into the wild-type allele on chromosome 9 ("ERV"); the transgene is not integrated into either allele of the locus but is randomly integrated into the host chromosome, or the transgene is integrated into the wild-type allele on chromosome 9 but not into the ERV109F allele on chromosome 15 ("ERV-deficient"). [Figure 16A] Figure 16A shows a 14-day test of four highly productive clones obtained from Figure 15C in an Ambr® 15 automated microscale bioreactor system (SARTORIUS Stedim, Germany) to evaluate their production capacity on a larger scale. [Figure 16B] Figure 16B shows a 14-day test of four highly productive clones obtained from Figure 15F in an Ambr® 15 automated microscale bioreactor system (SARTORIUS Stedim, Germany) to evaluate their production capacity on a larger scale. [Modes for carrying out the invention]

[0026] Transgene-producing cells and cell lines, as well as methods for producing and using them, are disclosed herein. For the production of a desired transgene, one or more ERV or LTR-RT loci in the cell's genome are targeted for transgene integration and expression. ERV sequences capable of forming viral particles may, in certain embodiments, be excluded or at least made non-functional with respect to viral particle production. The targeted ERV or LTR-RT locus may actually contain one allele that contains (or contained before removal) an ERV or LTR-RT sequence, while the other allele does not contain it and is a so-called wild-type allele that has never contained an ERV or LTR-RT sequence. The transgene may be introduced into a cell, preferably at an ERV or LTR-RT locus, for integration into an allele containing an ERV or LTR-RT sequence, into an allele that does not contain an ERV or LTR-RT sequence, or into both.

[0027] In one example, the transgene encodes an antibiotic selection gene, as well as a gene encoding the target protein, such as the heavy and light chains of immunoglobulin or human erythropoietin. The transgene is inserted into a vector containing a promoter upstream of the transgene and a Selexis Genetic Element (SGE) downstream of the transgene. The transcriptional activator-like effector (TALE) nickase is engineered to recognize a specific sequence of DNA and cleave it at the ERV locus located 5 bp upstream and 5 bp downstream of the ERV integration site. The selected ERV is integrated into only one of two alleles at the locus, while the other allele is a so-called wild-type allele that has never contained an ERV. CHO-K1 cells are transfected with vectors containing the gene encoding the TALE nickase, vectors containing the transgene, and vectors designed for transient expression of MRE11. Cells that show integration of the introduced gene into the wild-type allele locus / alleus of the ERV sequence, but do not show integration into the allele containing ERV, are selected for the production of the target protein.

[0028] In another example, the kit is used to create CHO cells that produce a target transgene. The kit's CHO cells are engineered to remove any incorporated ERV sequences that produce viral particles or virus-like particles. The cells are also engineered to insert landing pads into the allele wild-type corresponding locus / alleus of the ERV sequence. The landing pads encode green fluorescent protein (GFP). The kit also includes a vector encoding nickase for the sequence in the landing pad, as well as at least one vector encoding at least one target element that guides the at least one nickase to the landing pad. The kit also includes a vector designed for transient expression of CIRBP, as well as a vector into which the target transgene can be incorporated. After incorporating the target transgene, which is a single-domain antibody, all vectors are co-transfected into the engineered CHO cells. CHO cells that do not express GFP are selected. An expression vector for RS-1, a RAD51 stimulant, is also part of the kit and is added during co-transfection to stimulate homologous recombination (HR).

[0029] The cells / cell population according to the present invention (the latter is often used to refer to cells of a cell line that exhibits homogeneous properties of cells within a cell population) can be maintained under cell culture conditions. Eukaryotic cells / cell populations, preferably mammalian cells, e.g., human or non-human mammalian cells. Non-limiting examples of this type of cell include human cells, e.g., HEK cells (human fetal kidney), Chinese hamster ovary (CHO) cells, mouse myeloma cells including NS0 and Sp2 / 0 cells, and porcine cells, e.g., LLCPK (porcine kidney epithelium) cells. Modified versions of CHO cells include CHO DG44, CHO-K1, and CHO pro-3. In one preferred embodiment, the SURE CHO-M cell line (SELEXIS SA, Switzerland) is used.

[0030] The insertion site of an endogenous retrovirus (ERV) sequence or LTR-retrotransposon (LTR-RT) sequence in the cell genome is a nucleic acid sequence having a length of 100 nucleotides or less, preferably 90, 80, 70, 60, 50, 40, 30, 20, 10, 5, 4, 3, or 2 nucleotides or less, which is the locus corresponding to the wild-type allele of (i) that may be referred to herein as ERV deletion, and which includes an ERV or LTR-RT sequence, i.e., an ERV or LTR-RT sequence incorporated into the cell genome, and which includes (i)b) an ERV or LTR-RT sequence that was included before the complete or partial removal of the ERV or LTR-RT sequence, or (ii) a wild-type allele locus of (i) that may be referred to herein as ERV deletion. (i)a) and b) are referred herein to as “ERV sequence or LTR-RT sequence insertion locus / allele” or “ERV or LTR-RT insertion locus / allele” and (ii) is referred herein to as “allele wild-type corresponding locus / allele of ERV sequence insertion locus / allele” or “allele wild-type corresponding locus / allele of LTR-RT sequence insertion locus / allele” or simply the “allele wild-type (wt) corresponding sequence” as described above. As will be readily apparent to those skilled in the art, - One allele, (i)a) containing an ERV or LTR-RT sequence, (i)b) Prior to the complete or partial removal of the ERV or LTR-RT sequence, the system contained an ERV or LTR-RT sequence, and / or (ii)(i) is the wild-type allele counterpart, A gene locus likely exists.

[0031] Cells that combine at least two different alleles of a single locus, for example, (i)a) and (ii), or (i)b) and (ii), may be referred to herein as heterozygous with respect to that locus.

[0032] Cells combining (i)a) and (ii) are referred to as hemizygous with respect to single-copy ERV sequences.

[0033] A cell that has two identical alleles at a single locus, for example, (i)a) and (i)a), is said to be homozygous with respect to that locus.

[0034] - One allele and the corresponding allele locus, therefore both alleles, (i)a) contain an ERV or LTR-RT sequence, or (i)b) contain an ERV or LTR-RT sequence before complete or partial removal of the ERV LTR-RT sequence, - Cells in which one allele and the corresponding allele locus, and therefore both alleles are (ii) wild-type allele loci of an ERV sequence or LTR-RT sequence insertion locus, are also within the scope of the present invention.

[0035] One non-restrictive example of such insertion sites for ERV sequences is found in SEQ ID NOs: 1 and 2, at the 3' end of SEQ ID NO: 4, and at the 5' end of SEQ ID NO: 5. The ERV sequence is shown in SEQ ID NO: 3.

[0036] The allele wild-type corresponding loci of (i), which contain the corresponding 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, 5, 4, 3, and 2 nucleotides of the insertion site, however, have no evidence of current or past insertion of ERV or LTR-RT sequences. A non-restrictive example of such an allele correspondence to SEQ ID NO: 1 is SEQ ID NO: 2, where the insertion site is a nucleotide around nucleotide 30020. SEQ ID NO: 1 is the "ERV sequence or LTR-RT sequence insertion locus," while SEQ ID NO: 2 is the allele wild-type corresponding locus of the corresponding "ERV sequence or LTR-RT sequence insertion locus."

[0037] In this context, a locus is generally a location on a chromosome in a eukaryotic cell where a specific genomic sequence is located, having a maximum of 60,000 nucleotides in the genome (see Figures 5A and B), but with a length of 1,000, 900, 800, 700, or 600 nucleotides. However, as is known to those skilled in the art, loci can be identified, for example, in the human genome, with a length of 612 to 4,767,747 base pairs (Taher and Ovcharenko, 2009). An allele is a specific form of a locus, distinguished from other forms by its specific nucleotide sequence. In cells subjected to vigorous genomic rearrangement, which are many cells that can be maintained under cell culture conditions and used for the production of transgene products, such loci may be found on different chromosomes as a result of this genomic rearrangement. In general, to capture these configurations that share a genomic sequence consisting of 1,000 to 60,000 nucleotides defining a locus, the different alleles of such a locus may be referred to, for example, as discussed elsewhere herein, as the “wild-type allele locus of an ERV or LTR-RT sequence insertion locus” and the “ERV or LTR-RT sequence insertion locus.” An example of such a locus is the locus with the insertion site of ERV-C 109F. This locus is located on chromosomes 15 and 9 due to rearrangement (see Figures 7A–7C). On chromosome 15, the ERV-C 109F sequence is incorporated into the genome, while on chromosome 7, the “wild-type allele locus” is located. The fact that this is actually a single locus present on two chromosomes can be inferred from the corresponding surrounding sequences at the 5' and / or 3' of the insertion site, for example, for ERV-C 109F, from sequence numbers 1 and 2, for example, nucleotides 1-30020 of sequence number 1 and nucleotides 1-30020 of sequence number 2.Therefore, gene loci on multiple chromosomes often exhibit high sequence identity, for example, 100% sequence identity for sequences upstream of the ERV-C 109F insertion site (nucleotides 1-30020 of SEQ ID NO: 1 and nucleotides 1-30020 of SEQ ID NO: 2). However, slight mismatches that reduce sequence identity to, for example, 99%, 98%, 97%, 96%, or 95% are possible. Additionally, deletions may exist around the ERV or LTR-RT sequence insertion sites.

[0038] Loci containing insertion sites for endogenous retroviral (ERV) sequences or LTR-retrotransposon (LTR-RT) sequences generally number up to 60,000, 50,000, 40,000, or 30,000, but are generally identifiable by sequences of fewer than 20,000 or 10,000 nucleotides, or in certain specific cases, fewer than 9,000, 8,000, 7,000, 6,000, 5,000, 4,000, 3,000, or 2,000 nucleotides, or are identifiable by sequences of 1,000 to 600 nucleotides, including approximately 900, 800, or 700 nucleotides. As mentioned above, one such locus is shown in SEQ ID NO: 1, and its corresponding wild-type allele locus is shown in SEQ ID NO: 2. As those skilled in the art will understand, for example, the integration of a transgene may occur within any portion of an endogenous retrovirus (ERV) sequence (see, e.g., SEQ ID NO: 3) or an LTR-retrotransposon (LTR-RT) sequence, or within a chromosomal sequence of a locus adjacent to an ERV or LTR-RT sequence (an ERV or LTR-RT adjacent sequence), or replace an ERV or LTR-RT sequence or an ERV or LTR-RT adjacent sequence that is completely or partially deleted at a locus (see, e.g., SEQ ID NO: 4 and / or 5), which is within the scope of the present invention. A preferred locus containing an insertion site of an ERV or LTR-RT sequence is a locus in which the ERV sequence contains, preferably, the complete gag gene, pol gene, or env gene, but at least a portion thereof, and / or at least one, preferably two LTRs, most preferably all of these subsequences of the ERV, or each of the subsequences of the LTR-RT. In a particular embodiment, a more preferred locus containing an insertion site for an endogenous retrovirus (ERV) sequence or an LTR-retrotransposon (LTR-RT) sequence is a corresponding wild-type allele locus lacking any ERV sequence or LTR-RT sequence.

[0039] As those skilled in the art will understand, ERV sequences other than those shown in SEQ ID NOs: 1-5, as well as LTR-RT sequences and their respective loci, are also within the scope of the present invention. Some non-limiting examples of such ERV sequences are listed in Table 1 as SEQ ID NOs: 12-24.

[0040] [Table 1]

[0041] As shown, a locus containing an ERV sequence or LTR-RT sequence and its insertion site have an allele-corresponding locus that includes the insertion site but may or may not contain an ERV sequence or LTR-RT sequence. Furthermore, as described above, in certain embodiments, neither locus may contain an ERV sequence nor an LTR-RT sequence: the cell may have been manipulated to remove an ERV sequence or an LTR-RT sequence or a portion thereof, and in certain embodiments, a portion of the locus, i.e., an ERV or LTR-RT insertion site / a sequence adjacent to the ERV or LTR-RT sequence inserted therein. The cell may have been manipulated in this way at the time of transgene integration, or alternatively, the transgene may be integrated into a cell that has already been manipulated to remove or alter a portion of an ERV sequence or LTR-RT sequence. Each of these changes and the resulting cells are disclosed, for example, in U.S. Patent Application No. 62 / 784,566 and U.S.-designated International Patent Application Publication No. 2020 / 136149, which are incorporated herein by reference in their entirety.

[0042] Hemizygosity with respect to ERV / LTR-RT sequences refers to the fact that there is only one copy of a given ERV / LTR-RT sequence at a particular locus in a diploid cell. This means that there is an "ERV or LTR-RT insertion locus" and an "allele wild-type corresponding locus" for the ERV or LTR-RT insertion locus. Homozygosity with respect to ERV / LTR-RT sequences means that the alleles at a locus are corresponding and that at a particular locus in a diploid cell, both have an "ERV or LTR-RT insertion locus" or both have an "allele wild-type corresponding."

[0043] The introduced gene can be incorporated into a locus that is hemizygous or homozygous with respect to the ERV sequence / LTR-RT sequence.

[0044] LTR-retrotransposon (LTR-RT) sequences, also known as mammalian LTR-retrotransposon sequences or MaLR sequences, contain at least two LTR sequences adjacent to regions encoding two enzymes: at least two enzymes, integrase and reverse transcriptase (RT), and these include the gag and pol genes. In contrast to ERVs, LTR-RT sequences do not contain the env gene encoding the envelope protein (ENV) (Havecker et al., 2004).

[0045] Endogenous retrovirus (ERV) sequences constitute the remnants of retroviral integration into the cell's genome and include at least a portion of the gag, pol, and env genes, as well as / or at least one, preferably two, long terminal repeats (LTRs). The functional units or portions that make up an ERV sequence are also referred to as ERV subsequences. Thus, the gag, pol, env, and LTRs are considered ERV subsequences. In a preferred embodiment, at least one, preferably two, of the gag, pol, or env genes express their respective proteins. In a more preferred embodiment of ERV selection, the ERV sequence releases VPs (viral particles) or VLPs (viral-like particles). The size of a complete endogenous retrovirus is, on average, 6–12 kb, and it contains the gag, pol, and env genes, which always occur in the same order. The coding sequence is flanked by two LTRs (long terminal repeat sequences). Most ERVs are deficient due to having numerous inactivating mutations. In addition, they can be inactivated (i.e., cease to be transcribed) by epigenetic silencing. However, some ERVs still have open reading frames in their genome and / or they can be transcriptionally active. Mammalian ERVs have potent similarities and can originate from genera of gamma-retroviruses and beta-retroviruses, including intracisional A particles (IAP or IAPS), feline leukemia virus (FeLV), mouse leukemia virus (MLV), koala infectious disease virus (KoRV), and mouse mammary tumor virus (MMTV). ERVs maintained within the genome may have certain advantages for cells into which they are incorporated, including providing a source of genetic diversity and protection against other viral pathogens. However, they can be infectious and pose risks, particularly as a result of ERV awakening due to cancer, cellular stress, and / or epigenetic alteration, in the context of expression of transgenes, i.e., proteins, as described elsewhere herein.

[0046] The three main proteins encoded within the retroviral genome are Gag, Pol, and Env. Gag (antigen group), encoded by the gag gene, is a polyprotein that processes into the matrix and other core proteins, including the nucleoprotein core particle that determines the retroviral core. Pol, encoded by the pol gene, is a reverse transcriptase with RNase H and integrase functions. Its activity results in the pre-integration form of the double-stranded DNA of the virus, integration into the host genome via integrase function, and reverse transcription after integration into the host genome via RNase function. Env, encoded by the env gene, is an envelope protein present in the viral lipid layer that determines viral affinity.

[0047] The gag gene produces a Gag precursor protein, which is expressed from unspliced ​​viral mRNA. During the viral maturation process, the Gag precursor protein is cleaved by a viral-encoded protease (the product of the pol gene) into four small proteins, generally denoted as MA (matrix), CA (capsid), NC (nucleocapsid), and a further protein domain (e.g., pp12 in mouse leukemia virus (MLV) or p6 in HIV).

[0048] Viral proteases (Pro), integrases (IN), RNase H, and reverse transcriptase (RT) are expressed within the context of Gag-Pol fusion proteins. Gag-Pol precursors are generally produced by ribosome frameshift events, which are triggered by specific cis-acting RNA motifs (in which a short stem-loop follows a heptanucleotide sequence in the distal region of Gag RNA). When ribosomes encounter this motif, they shift to the pol reading frame for approximately 5% of their time without interrupting translation. The frequency of ribosome frameshifts explains why Gag and Gag-Pol precursors are produced in a ratio of approximately 20:1.

[0049] During viral maturation, the virus-encoded protease cleaves the Pol polypeptide from Gag and further digests it to separate the protease, RT, RNase H, and integrase activity. However, not all of these cleavage processes are efficient; for example, approximately 50% of the RT protein remains linked to RNase H as a single polypeptide (p65) (Hope & Trono, 2000).

[0050] The pol gene encodes reverse transcriptase. During the reverse transcription process, polymerase creates a double-stranded DNA copy of the single-stranded genomic RNA present in the virion. RNase H removes the original RNA template from the first DNA strand, enabling the synthesis of the complementary DNA strand. The primary functional species of polymerase is the heterodimer. All pol gene products can be found within the capsid of the released virion.

[0051] The IN protein mediates the insertion of proviral DNA into the genomic DNA of infected cells. This process is mediated by three different functions of IN.

[0052] The Env protein is expressed from single-splicing mRNA. First, it is synthesized in the endoplasmic reticulum, then Env migrates to the Golgi complex, where it undergoes glycosylation. Env glycosylation is generally required for infectivity. Cellular proteases cleave the protein into its transmembrane and surface domains.

[0053] Viral genomic RNA expressed from certain ERVs in the genome can be released from cells in the form of VPs. Other expressed ERVs may cause the formation of VLPs, such as RVLPs (retrovirus-like particles), but not VP formation, and therefore may not result in the release of particles containing viral genomic RNA. However, generally speaking, what is released is likely to be infectious.

[0054] In the context of this application, VP refers to a viral particle containing at least a portion of the viral genome. In some cases, VP may contain full-length viral genome RNA and therefore may be a functional VP. As used in the context of this invention, VLP is a particle that appears to be a VP but lacks any portion of the viral genome.

[0055] The vector according to the present invention is a nucleic acid molecule capable of transporting other nucleic acids to which it is linked. Plasmids are, for example, one type of vector. Viral vectors, such as lentivirus or adeno-associated virus (AAV) vectors, are another type of vector.

[0056] In a particular embodiment of the present invention, the vector is used to transport exogenous nucleic acids into cells or a population of cells.

[0057] As used herein, exogenous nucleic acids mean that the referenced nucleic acid is introduced into a host cell. The source of an exogenous nucleic acid may be, for example, a homogeneous or heterogeneous nucleic acid that expresses the protein of interest. Correspondingly, the term endogenous refers to a nucleic acid molecule already present in the host cell. The term heterogeneous nucleic acid refers to a nucleic acid molecule derived from a source that is not of the host cell species, while homogeneous nucleic acid refers to a nucleic acid molecule derived from the same species as the host cell. Thus, the exogenous nucleic acid according to the present invention may utilize either or both heterogeneous and / or homogeneous nucleic acids.

[0058] For example, the cDNA of the human interferon gene is a heterologous exogenous nucleic acid when introduced into CHO cells, but an allologous exogenous nucleic acid in HeLa cells. When introduced into cells, the exogenous gene may be part of the vector, or it may be introduced together with additional endogenous or exogenous nucleic acid sequences.

[0059] A transgene is an exogenous nucleic acid that codes for a product of interest, such as a protein, also referred to as a "transgene expression product." In certain embodiments, one or more transgenes are required to generate a cell line that produces a product of interest, specifically a protein of interest, such as an antibody. This may require transgenes that code for the light chain and heavy chain for producing the antibody, i.e., the protein of interest, as well as an antibiotic selection transgene used to select cells that are stably transfected. The transgene expression product may also be a simple marker protein, such as an antibiotic selection gene, highly sensitive green fluorescent protein (GFP), or β-galactosidase (lacZ). In this case, the transgene may be incorporated under the control of a specific gene promoter, replace a completely or partially removed ERV or LTR-RT sequence, or be incorporated into an allele wild-type counterpart. Such a transgene may, together with another transgene, such as a transgene that codes for a protein of interest, or another transgene expression product, function as a landing pad for the incorporation of a transgene expression product that yields a protein of interest, such as a therapeutic protein. For example, the transgene expression product may be the light or heavy chain of an antibody; however, the “protein of interest” is an immunoglobulin composed of four chains. However, the product of interest, for example, the protein of interest, is generally a protein (including fusion proteins), not limited to signaling proteins such as α-IFN, β-IFN, γ-IFN, τ-IFN, ω-IFN, cytokines such as erythropoietin, or antibodies such as monoclonal antibodies, or fusion proteins, as well as regulatory RNA such as siRNA, shRNA, or mRNA. The “protein of interest” is a therapeutic protein recovered from the cell supernatant and measured in picograms per cell per day, μg / l, or mg / l.

[0060] As used herein, transfection refers to the introduction of nucleic acids, including naked nucleic acids, purified nucleic acids, or vectors containing specific nucleic acids, into cells, specifically eukaryotic cells, including mammalian cells. Any known transfection method can be used in the context of the present invention. Some of these methods involve enhancing the permeability of biological membranes to introduce nucleic acids into cells. Notable examples are electroporation or microporation. The methods may be used on their own or assisted by ultrasound, electromagnetic and thermal energy, chemical permeability enhancers, pressure, etc., to selectively enhance the flow rate of nucleic acids into host cells. Other transfection methods, such as transfection based on carriers including lipofection or viruses (also called transduction), and transfection based on chemicals, are also within the scope of the present invention. However, any method for introducing nucleic acids into cells can be used. Transiently transfected cells will retain / express the transfected RNA / DNA for a short period and will not be passaged. Stably transfected cells continuously express the transfected DNA and pass it through generations, and the exogenous nucleic acid is incorporated into the cell's genome.

[0061] When used herein, “DNA repair pathway” or “DRP” refers to cellular mechanisms that enable a cell to maintain its genomic integrity and function in response to the detection of DNA damage, such as single-strand or double-strand breaks. Depending on several parameters, such as the type and length of the DNA damage, or the cell cycle of the cell at the time of the damage, DRPs include excision, canonical homology-directed repair (canonical HDR), homologous recombination (HR), alternative homology-directed repair (Alt-HDR), double-strand break repair (DSBR), single-strand annealing (SSA), synthesis-dependent strand annealing (SDSA), break-induced replication (BIR), alternative end joining (Alt-EJ), microhomology-mediated end joining (MMEJ), and DNA synthesis-dependent This includes, but is not limited to, sex microhomology-mediated end joining (SD-MMEJ), non-homologous end joining (NHEJ) pathways, such as canonical non-homologous end joining (C-NHEJ) repair, alternative non-homologous end joining (A-NHEJ) pathway, trans-transparency DNA synthesis (TLS) repair, base excision repair (BER), nucleotide excision repair (NER), mismatch repair (MMR), DNA damage response (DDR), blunt end joining, single-strand break repair (SSBR), inter-strand crosslink repair (ICL), and Fanconi anemia pathway (FA). However, the DRP of the present invention is preferably selected from the above group.

[0062] DNA repair pathways can be inhibited or, rather, preferred / enhanced. Genes, mRNAs, or corresponding proteins involved in such pathways can be modified to inhibit or preferentially / enhance the pathway (see examples in Table 2).

[0063] [Table 2-1]

[0064] [Table 2-2]

[0065] [Table 2-3]

[0066] [Table 2-4]

[0067] [Table 2-5]

[0068] Examples of NHEJ inhibitors (= inhibitors of PARP1, Ku70 / 80, DNA-PKcs, XRCC4 / XLF, ligase IV, ligase III, XRCCl, Artemis, and PNK) include, without limitation, NU7441 (Leahy et al., Identification of a highly potent and selective DNA-dependent protein kinase (DNA-PK) inhibitor (NU7441) by screening of chromenone libraries. (Leahy et al., (2004), NU7026 (Willmore et al., 2004), olaparib, DNA ligase IV inhibitor, Scr7 (Maruyama et al., 2015)), KU-0060648 (Robert et al., 2015), anti-EGFR antibody C225 (cetuximab) (Dittmann et al., Examples include compound 401 (2-(4-morpholinyl)-4H-pyrimido[2,la]isoquinoline-4-one), vanillin, waltomannin, DMNB, IC87361, LY294002, OK-1035, CO15, NK314, and PI103 hydrochloride.

[0069] Examples of MMEJ inhibitors include, but are not limited to, MRE11 inhibitors, such as myrin and its derivatives (Shibata et al, 2014), PolQ inhibitors, and CtIP inhibitors (Sfeir and Symington, 2015).

[0070] Examples of HR inhibitors include, but are not limited to, RI-1 and BO-2.

[0071] Examples of HR stimulants include RS-1 (RAD51 stimulant), but are not limited to these.

[0072] Examples of NHEJ stimulants include, but are not limited to, IP6 (inositol hexakisphosphate, a DNA-PK enhancer, Hanakahi 2000, Ma 2002, Cheung 2008).

[0073] Downregulation of DRP reduces the activity of such DRP in cells or cell populations. Downregulation of DRP may be 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% of the repair activity (hereinafter, "activity") without downregulation. Downregulation can be achieved by a number of means, including, but is not limited to, contacting the cells or cell population with one or more inhibitors, e.g., chemical inhibitors of DRP / its components; inactivating DRP / its components; downregulating DRP / its components (e.g., by contacting the cells or cell population with one or more inhibitory nucleic acids, e.g., miRNA, siRNA, shRNA, or any combination thereof, or by expressing them in the cells or cell population); and / or mutating one or more genes of DRP / its components.

[0074] In a preferred embodiment, a DRP is downregulated, which is either non-productive or competes with another DRP, and is therefore referred to as a competing pathway or a non-productive pathway.

[0075] For example, the NHEJ pathway can be inhibited, for instance, by MMEJ and related mechanisms, to prioritize the productive integration of exogenous DNA. In the context of the present invention, any active DRP may compete with another active DRP in the cell and is therefore a competitive DRP pathway. Non-productive DRPs in the context of the present invention are pathways that do not mediate or only inefficiently mediate the integration of exogenous DNA into the cellular genome. For example, synthesis-dependent strand annealing (SDSA), break-induced replication (BIR), base excision repair (BER), nucleotide excision repair (NER), mismatch repair (MMR), DNA damage response (DDR), blunt end joining, single-strand break repair (SSBR), and interstrand crosslink repair (ICL) are generally inefficient in mediating the integration of exogenous DNA.

[0076] Downregulation of one DRP generally results in one or more other DNA repair pathways taking over the repair function of the downregulated DRP. The one or more DRPs that take over the repair function are generally upregulated. Upregulation of one or more DRPs may be 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% of their activity without downregulation. A DRP upregulated as a result of downregulation of another competing DRP is considered "preferred" (or enhanced) compared to the downregulated DRP. The degree of preferential / enhancement may be proportional to the degree of downregulation, for example, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% higher activity compared to the unregulated activity of the downregulated DRP. The downregulated activity of DRP may shift to one pathway, or it may shift to two or more pathways that inherit the DNA repair function of the downregulated DRP.

[0077] Apart from downregulating other DRPs, DRPs can also be expressed, for example, by causing overexpression of one or more components of the DRPs in the cells or cell populations; by heterologously introducing the components of the DRPs into the cells or cell populations; or by contacting the cells or cell populations with one or more modifiers of one or more components of the DRPs, preferably stimulants, such as chemical stimulants, to mutate one or more genes of the DRPs, where the mutation enhances the expression or activity of one or more components of the DRPs.

[0078] Adjustment, particularly up-adjustment, can be achieved by transfecting cells / cell populations with one or more genes listed in Table 2, and / or sequences having sequence identity (e.g., 99%, 98%, 97%, 96%, or 95%) with the genes listed in SEQ ID NOs. 25-28 and 38-59 and these sequences as described elsewhere herein, simultaneously with, or transiently less than 24 hours, less than 18 hours, less than 12 hours, less than 8 hours, less than 4 hours, less than 2 hours, less than 1 hour before, and in certain embodiments less than, after, the embedded vector shown in Figure 11B, for example, by co-transfecting (i.e., simultaneously or within 1 hour). Several representative vectors are shown in Figure 11A. However, as will be understood by those skilled in the art, any gene encoding a protein involved in DRP can be used in such a vector. Some of these genes are listed in Table 2, and certain preferred genes are listed in SEQ ID NOs. 25-28 and 38-59. As can be seen in Figure 11A, certain preferred vectors insert the DRP gene between ITRs (terminal inverted repeat sequences), i.e., the DRP gene is "adjacent to two ITRs on both sides", and in certain preferred embodiments, also include gene elements, e.g., MARs (e.g., SGEs). Transfection with one or more vectors having MRE11, POLQ, CIRBP, and / or RAD51, e.g., one or more vectors having SEQ ID NO: 25 (hMRE11), SEQ ID NO: 26 (cgPOLQ), SEQ ID NO: 27 (hCIRBP), and / or SEQ ID NO: 28 (hRAD51) is preferred.Other preferred DRP genes are those whose expression products are (i) used by both MMEJ and HR (Mre11, Rad50, Nbs1, CtIP, Exo1, BLM, ATM, ERCC1, Srs2, Xpf, Pol δ, Ligase I, Ligand III), (ii) used by MMEJ but not by HR (e.g., LigIV, XRCC1, PARP1, POLQ, WRN, POLB, Pol4), and (iii) used by HR but not by MMEJ proteins (e.g., BRCA1, 53BP1, MDC1).

[0079] When used herein, a chemical stimulant refers to a chemical compound that can be used to enhance gene expression or protein activity. As will be readily apparent to those skilled in the art, the chemical stimulant will depend on which component of which DNA repair pathway (DPR) it stimulates. For example, RS-1, a RAD51 stimulant, stimulates HR. IP6 (inositol hexakisphosphate) and other DNA-PK enhancers are NHEJ stimulants (see, e.g., Hanakahi 2000, Ma 2002, Cheung 2008).

[0080] When used herein, a chemical inhibitor refers to a chemical compound that can be used to inhibit gene expression or protein activity. Similarly, as will be readily apparent to those skilled in the art, the chemical inhibitor will depend on which component of which DPR is stimulated. Examples of chemical inhibitors of MMEJ include, but are not limited to, MRE11 inhibitors, e.g., myrin and its derivatives (Shibata et al, Molec. Cell (2014) 53:7-18), PolQ inhibitors, and CtIP inhibitors (Sfeir and Symington, “Microhomology-Mediated End Joining: A Back-up Survival Mechanism or Dedicated Pathway?” Trends Biochem Sci (2015) 40:701-714). Examples of HR inhibitors include RI-1 (RAD51 inhibitor 1) and BO2 (3-(phenylmethyl)-2-[(1E)-2-(3-pyridinyl)ethenyl]-4(3H)-quinazolinone). See also U.S. Patent Publications 2019 / 0194694A1 and 2015 / 0361451A1.

[0081] Chemical stimulants and inhibitors are generally exogenous, i.e., added to the cell supernatant and taken up by the cells. Such inhibitors may be added to cells / cell populations, for example, simultaneously, or within 1 hour, or within 24 hours, 18 hours, 12 hours, 8 hours, 4 hours, or 2 hours, along with the various vectors described herein. Nucleases and / or nicasses: Induction of double-strand / single-strand breaks

[0082] Different molecules can introduce double-strand and / or single-strand breaks into genomic nucleic acids. Examples of nucleases or nickases of the present invention include, but are not limited to, homing end nucleases, restriction enzymes, zinc finger nucleases or zinc finger nickases, meganucleases or mega nickases, transcription activator-like effector (TALE) nucleases or TALE nickases, guided, in particular, nucleic acid guide nucleases or nickases, e.g., RNA guide nucleases or RNA guide nickases, DNA guide nucleases, e.g., Argonaut (NgAgo) or DNA guide nickase of Natronobacterium gregoryi, mega TAL nucleases, BurrH nucleases, ARCUS nucleases, modified or chimeric versions or variants thereof, and combinations thereof. RNA guide nucleases or RNA guide nickases are optionally part of a CRISPR-based system.

[0083] In a preferred embodiment, these double-strand and / or single-strand breaks are introduced by one or more nucleases or nickases. Nucleases can introduce double-strand and / or single-strand breaks. The term nickase is used for molecules that introduce single-strand breaks and may be nucleases having a partially inactive DNA cleavage domain. For example, the nuclease domains of nucleases can be mutated independently of each other to create DNA "nickases" that can introduce single-strand breaks with the same specificity as each nuclease. Subject to the limitations described herein, the following considerations regarding nucleases apply equally to nickases.

[0084] Nucleases can cleave phosphodiester bonds between monomers in nucleic acids. Many nucleases are involved in DNA repair by recognizing damaged sites and cleaving them from the surrounding DNA. These enzymes may be part of a complex. Exonucleases are nucleases that digest nucleic acids from the ends. In this context, preferred endonucleases are those that act on the central region of the target molecule. Deoxyribonucleases act on DNA, and ribonucleases act on RNA. Many nucleases involved in DNA repair are not sequence-specific. In this context, however, sequence-specific nucleases are preferred. In one preferred embodiment, the sequence-specific nuclease is specific to larger nucleotide fragments in the target genome, for example, five or more nucleotides, or 10, 15, 20, 25, 30, 35, 40, 45, or even 50 or more nucleotides, in the ranges of 5-50, 10-50, 15-50, 15-40, and 15-30, which in a particular embodiment are preferred as target sequences in the target genome. The larger such “recognition sequences” are, the fewer target sites there are in the genome, and the more specific the cleavage performed by the nuclease or nickase on the genome, the more site-specific the resulting cleavage becomes. Site-specific nucleases generally have fewer than 10, fewer than 5, fewer than 4, fewer than 3, fewer than 2, or a single (only one) target site in the genome. Nucleases that have been manipulated to alter genomic nucleic acids, including by cleaving specific genomic target sequences, are referred to herein as manipulated nucleases. CRISPR-based systems are one type of manipulated nuclease. However, such manipulated nucleases may be based on any nuclease described herein. In one preferred embodiment, the codons of each nuclease are optimized for expression in eukaryotic cells, e.g., mammalian cells.The nuclease / system of the present invention may also include one or more linkers and / or additional functional domains, for example, a 5-3' exonuclease or a 3-5' exonuclease, or other non-nuclease domains, for example, an endoprocessing enzyme domain exhibiting a helicase domain.

[0085] Restriction enzymes are sequence-specific nucleases that are often specific to smaller nucleotide fragments and, as a result, have short recognition sequences. The first letter of the name is derived from the genus, and the second two letters are derived from the species of prokaryotic cell from which it was isolated. For example, EcoRI is derived from the bacterium Escherichia coli RY13. Many restriction enzymes are restriction endonucleases, which, for example, introduce blunt or alternating cuts in the middle of nucleic acids. Many restriction enzymes are sensitive to the methylation status of the DNA they target. Cuts can be blocked or damaged if certain bases at the enzyme's recognition site are altered.

[0086] Examples of methylation-sensitive restriction enzymes that are important in epigenetics include DpnI and DpnII, which are sensitive to the detection of N6-methyladenine in the GATC recognition site, and HpaII and MspI, which are sensitive to the detection of C5-methylcytosine in the CCGG recognition site.

[0087] Table 3 lists some exemplary restriction enzymes used in the examples, along with their recognition sites found in the reference CHO genome, their CpG methylation sensitivity, and the number of target sites.

[0088] [Table 3]

[0089] Endonucleases that recognize sequences larger than 12 base pairs are called meganucleases. Meganucleases / nickases are endodeoxyribonucleases characterized by large recognition sites (e.g., 12-40 base pairs, e.g., 20-40 or 30-40 base pair double-stranded DNA sequences), and as a result, these sites can occur only once in any given genome.

[0090] Homing endonucleases are a meganuclease form of double-stranded DNase possessing a large, asymmetric recognition site and a coding sequence typically embedded in either an intron or intein. Because homing endonuclease recognition sites are extremely rare within the genome, they cleave at very few locations, sometimes even a single location within the genome (see also International Publication No. 2004067736 and U.S. Patent No. 8697395B2).

[0091] Zinc finger nucleases / nickases (ZFNs) are artificial restriction enzymes produced by fusing a zinc finger DNA-binding domain to a DNA-cleaving domain. The zinc finger domain can be manipulated to target specific desired DNA sequences. ZFNs are described, for example, by Urnov F., et al. (Highly efficient endogenous human gene correction using designed zinc-finger nucleases (2005) Nature 435:646-651).

[0092] Transcription activator-like effector (TALE) nucleases / nickases are restriction enzymes that can be manipulated to cleave specific sequences of DNA. Because TALEs can be manipulated to bind to virtually any desired DNA sequence, when combined with a DNA cleavage domain, they can cleave DNA at specific locations. TALE nucleases are described, for example, by Mussolino et al. (A Novel TALE nuclease scaffold enables high genome editing activity in combination with low toxicity (2011) Nucl.Acids Res. 39 (21):9283-9293).

[0093] RNA guide nucleases / nickases, particularly endonucleases, include, for example, Cas9 or Cpf1. CRISPR systems are described in detail. Any CRISPR-based system is part of the present invention. When a different RNA guide endonuclease is used, a suitable guide RNA, sgRNA, or crRNA, or other suitable RNA sequences that interact with the RNA guide endonuclease and target genomic target sites in genomic nucleic acids, may be used.

[0094] In a particular preferred embodiment, the nuclease is an RNA guide nuclease. Non-limiting examples of RNA guide nucleases, including nucleic acid guide nucleases for use in this disclosure, include Casl, CaslB, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csnl and Csxl2), Cas10, CasX, CasY, Cpf1, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Examples include, but are not limited to, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, CsxlO, Csxl6, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4, Cms1, Cpf1, their homologs, their orthologues, or modified versions thereof, MAD7, e.g., MADzyme (INSCRIPTA), C2cl, C2c2, C2c3.

[0095] In a particular preferred embodiment, the nuclease is a DNA guide nuclease. “DNA guide nuclease” refers to a system comprising a DNA guide (gDNA) and an endonuclease. The DNA guide, e.g., 5'-phosphorylated single-stranded DNA (ssDNA), guides the endonuclease to cleave a double-stranded DNA target within the DNA guide nickas. “Argonaut-based system” refers to a DNA guide nuclease based on a single-stranded DNA guide (gDNA) and an endonuclease derived from the Argonaut (Ago) protein family. The gDNA directs the endonuclease's target to a specific DNA sequence, resulting in sequence-specific DNA cleavage. The Ago protein can be modified by mutagenesis to have improved activity at 37°C. Several Argonaute proteins are characterized from Natronobacterium gregoryi (NgAgo, see, e.g., Gao et al., DNA-guided genome editing using the Natronobacterium gregoryi Argonaute, Nature Biotechnology, published online May 2, 2016), Rhodobacter sphaeroides (RsAgo, see, e.g., Olivnikov et al.), Thermo thermophiles (TaAgo, see, e.g., Swarts et al (2014), Nature 507(7491): 258-261), and Pyrococcus furiosus Argonaute (PfAgo).

[0096] The use of Argonaut-based systems enables targeted cleavage of genomic DNA within cells.

[0097] TtAgo is a prokaryotic Argonaut protein thought to be involved in gene silencing. TtAgo originates from the bacterium Thermus thermophilus. (See, for example, Swarts et al, ibid, G. Sheng et al, (2013) Proc. Natl. Acad. Sci. USA III, 652).

[0098] One of the best-known prokaryotic Ago proteins is derived from T. thermophilus (TtAgo, Swarts et al., ibid.). This "guide DNA" to which TtAgo is bound induces a protein-DNA complex to bind to a Watson-Crick complementary DNA sequence on a third DNA molecule. Once the target DNA can be identified using the sequence information in these guide DNAs, the TtAgo-guide DNA complex cleaves the target DNA. This mechanism is also supported by the structure of the TtAgo-guide DNA complex while bound to the target DNA (G. Sheng et al., ibid.). Ago derived from Rhodobacter sphaeroides (RsAgo) has similar properties (ibid.).

[0099] An exogenous guide DNA of any DNA sequence can be loaded onto the TtAgo protein (Swarts et al., ibid.). Since the specificity of TtAgo cleavage is induced by the guide DNA, a TtAgo-DNA complex formed using guide DNA specified by an exogenous investigator will therefore direct TtAgo target DNA cleavage to target DNA specified by a complementary investigator. In this manner, targeted double-strand breaks can be created in DNA. The use of a TtAgo-guide DNA system (or an orthologous Ago-guide DNA system derived from another organism) enables targeted cleavage of genomic DNA within cells. Such cleavage may be single-stranded or double-stranded. For cleavage of mammalian genomic DNA, it may be preferable to use a codon-optimized version of TtAgo for expression in mammalian cells. Furthermore, it may be preferable to treat cells with a TtAgo-DNA complex formed in vitro, in which the TtAgo protein is fused to a cell-permeable peptide. Ago-RNA-mediated DNA cleavage can be used to achieve a number of results, including gene knockout, targeted gene addition, gene modification, and targeted gene deletion, using standard techniques in the field of DNA cleavage utilization.

[0100] Exemplary examples of Argonaut-based systems and gDNA designs are disclosed in International Publication No. 2017 / 107898, Chinese Publication No. 105483118, International Publication No. 2017 / 139264, U.S. Patent Applications No. 2017367280 and 20180201921, and the references cited herein, all of which are incorporated herein by reference in their entirety. Argonaut-based systems optionally include one or more linkers and / or additional functional domains, e.g., 5-3' exonuclease or 3-5' exonuclease, or other non-nuclease domains, e.g., endoprocessing enzyme domains exhibiting a helicase domain.

[0101] "MegaTAL nuclease / nickase" refers to an engineered nuclease comprising an engineered TALE DNA-binding domain and an engineered meganuclease or engineered homing endonuclease. The TALE DNA-binding domain can be designed to bind to DNA at virtually any locus of nucleic acid sequences in the genome, and when such a DNA-binding domain is fused to an engineered meganuclease, it can cleave the target sequence. Exemplary examples of megaTAL nuclease and TALE DNA-binding domain designs are disclosed, for example, in Boissel et al. (MegaTALs: a rare-cleaving nuclease architecture for therapeutic genome engineering (2013), Nucleic Acids Research 42 (4):2591-2601) and the references cited therein, all of which are incorporated herein by reference in their entirety. MegaTAL nucleases optionally include one or more linkers and / or additional functional domains, such as a C-terminal domain (CTD) polypeptide, an N-terminal domain (NTD) polypeptide, a 5-3' exonuclease or 3-5' exonuclease, or other non-nuclease domains, such as an endoprocessing enzyme domain exhibiting a helicase domain.

[0102] A "TALE DNA-binding domain" is the DNA-binding portion of a transcription activator-like effector (TALE or TAL effector) that mimics plant transcription activators that manipulate the plant transcriptome (see, for example, Kay et al., 2007. Science 318:648-651). TALE DNA-binding domains intended in specific embodiments may be de novo-engineered or engineered from naturally occurring TALEs, including, but not limited to, AvrBs3 derived from Xanthomonas campestris pv. vesicatoria, Xanthomonas gardneri, Xanthomonas translucens, Xanthomonas axonopodis, Xanthomonas perforans, Xanthomonas alfalfa, Xanthomonas citri, Xanthomonas euvesicatoria, and Xanthomonas oryzae, as well as brgl 1 and hpxl7 derived from Ralstonia solanacearum. Exemplary examples of TALE proteins for inducing and designing DNA-binding domains are disclosed in U.S. Patent No. 9,017,967 and the references cited herein, all of which are incorporated herein by reference in their entirety.

[0103] "BurrH nuclease" refers to a nuclease-active fusion protein containing modular base-specific nucleic acid-binding domains (MBBBDs). These domains are derived from proteins from the bacterial intracellular symbiont Burkholderia Rhizoxinica or from other similar proteins identified from marine organisms. By combining different modules of these binding domains, modular base-specific binding domains can be manipulated to have binding properties to specific nucleic acid sequences, such as DNA-binding domains. Such manipulated MBBDs can then be fused to a nuclease-catalyzing domain to cleave DNA at virtually any locus of nucleic acid sequences in the genome. Exemplary examples of BurrH nuclease and MBBD designs are disclosed in International Publication 2014 / 018601 and U.S. Publication 2015225465A1, as well as the references cited herein, all of which are incorporated herein by reference in their entirety. The BurrH nuclease optionally includes one or more linkers and / or additional functional domains, e.g., a 5-3' exonuclease or a 3-5' exonuclease, or other non-nuclease domains, e.g., an endoprocessing enzyme domain exhibiting a helicase domain.

[0104] Enzymes such as transposases or integrases may also be used as niccasses / nucleases in the context of the disclosed methods and cells.

[0105] Target elements for targeting at least one locus in a cell's genome containing an insertion site for an endogenous retrovirus (ERV) sequence or an LTR-retrotransposon (LTR-RT) sequence are generally sequences that promote and / or guide the activity of nickases and / or nucleases. Such target elements include, for example, a guide RNA containing a single-stranded guide RNA (sgRNA) or crRNA (CRISPR RNA), and are encoded by Cas9, Cpf1, or Cms1 nuclease expression vectors that target, for example, the ERV C 109F 5' genome sequence (SEQ ID NO: 8 and SEQ ID NO: 10) and the ERV C 109F 3' genome sequence (SEQ ID NO: 9 and SEQ ID NO: 11). Upregulation and / or downregulation of DRPs can be adjusted depending on the type of element used. For example, in certain embodiments, the homologous recombination (HR) pathway is upregulated for DSBs produced by CRISPR cleavage sites 16 and 17 (see Figures 5A and 5B). In addition, vectors containing the transgene may include 5' and 3' homology arms (SEQ ID NOs. 6 and 7), which are present in the vector and loci, including the insertion site.

[0106] The sequence specificity of the CRISPR (clustered, regularly spaced, short, palindromic repeat sequences) system is determined by small RNAs. CRISPR loci consist of a series of repeat sequences separated by "spacer" sequences that match the genomes of bacteriophages and other mobile gene elements. The repeat-sequence-spacer array is transcribed as a long precursor and processed within the repeat sequences to produce small crRNAs that specify target sequences (also known as protospacers) to be cleaved by the CRISPR system. For cleavage, the presence of a sequence motif immediately downstream of the target sequence, also known as a protospacer flanking motif (PAM), is often required. CRISPR-related (cas) genes typically encode enzymatic mechanisms responsible for the development and targeting of crRNAs (CRISPR RNAs) flanked on both sides by repeat-sequence-spacer arrays. For example, Cas9 is a dsDNA endonuclease that uses crRNA guidance to specify cleavage sites. The loading of crRNA guides into Cas9 occurs during the processing of the crRNA precursor and requires a small antisense RNA, tracrRNA, and RNAse Ill relative to the precursor. In contrast to genome editing with ZFNs or TALENs, altering the target specificity of Cas9 does not require protein manipulation and only requires the design of a short crRNA guide, also called sgRNA when the crRNA is fused to tracrRNA (trans-activated CRISPR RNA).

[0107] To date, three different types of Cas9 nucleases (e.g., Cas9) have been employed in genome editing protocols. The first is wild-type Cas9, which can site-specifically cleave double-stranded DNA, leading to the activation of double-strand break (DSB) repair mechanisms. DSBs are repaired by the intracellular non-homologous end joining (NHEJ) pathway, which can result in insertions and / or deletions (indels) that disrupt the target locus. Alternatively, if a donor template homologous to the target locus is provided, DSBs can be repaired by the homology-directed repair (HDR) pathway, enabling precise substitution mutations.

[0108] Furthermore, we developed mutant forms (nCas9) possessing only nickase activity (e.g., Cas9D10A) to improve the precision of the Cas9 system. This means that it cleaves only one DNA strand and does not activate the NHEJ. Instead, when a homologous repair template is provided, DNA repair occurs solely via the high-fidelity HDR pathway, resulting in a reduction of indel mutations. Cas9D10A is therefore more attractive in terms of target specificity for many applications when a locus is targeted by a paired Cas9 complex designed to generate adjacent DNA nicks. Such Cas nickases can also be fused with other functional or catalytic domains, such as domains providing deamination activity (e.g., for base editing purposes).

[0109] The third type is based on enzymatically inactive Cas9 (eiCas9), also known as dead Cas9 or dCas9. This system includes Cas9 mutants lacking endonuclease activity due to mutations in their endonuclease domains (e.g., RuvC and HNH domains). dCas9 is still capable of binding to its guide RNA and DNA strands and can be fused to functional or catalytic domains that provide DNA modification activity, selected from, but not limited to, nuclease activity (e.g., Fok1), Clo51, methyltransferase activity, demethylase activity, deamination activity, depurination activity, integrase activity, transposase activity, and recombinase activity. Other domains that provide protein modification activity include, but are not limited to, repressive domains (e.g., KRAB domain), activating domains (e.g., VP16), methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity, and deglycosylation activity.

[0110] The term sequence identity refers to a measure of the identity of a nucleotide or amino acid sequence. Generally, sequences are aligned to achieve the highest degree of match. "Identity" itself has a recognized meaning within the relevant art and can be calculated using publicly available techniques. (See, for example, Computational Molecular Biology, Lesk, AM, ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, DW, ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, AM, and Griffin, HG, eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991). While several methods exist for determining the identity between two polynucleotide or polypeptide sequences, the term "identity" is well known to those skilled in the art to define identical nucleotides or amino acids at a given position within the sequence (Carillo, H. & Lipton, D., SIAM J Applied Math 48:1073 (1988)).

[0111] Whether any particular nucleic acid molecule is at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to, for example, sequence numbers 1, 2, 3, 4, 5, or any part thereof, of gamma retrovirus-like sequences can conventionally be determined by using a known computer program, such as DNAsis software (Hitachi Software, San Bruno, Calif.), for initial sequence alignment, followed by ESEE version 3.0 DNA / protein sequencing software for multiple sequence alignments.

[0112] Whether an amino acid sequence is at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to a protein expressed by, for example, SEQ ID NO: 1 or 3 or a part thereof, can conventionally be determined using a computer program, such as the BESTFIT program (Wisconsin Sequence Analysis Package®, version 8 for Unix, Genetics Computer Group, University Research Park, 575 Science Drive, Madison, Wis. 53711). BESTFIT uses the local homology algorithm of Smith and Waterman, Advances in Applied Mathematics 2:482-489 (1981) to find the segment with the highest homology between two sequences.

[0113] When using DNAsis, ESEE, BESTFIT, or any other sequence alignment program to determine whether a particular sequence is, for example, 95% identical to a reference sequence according to the present invention, the percentage of identity is calculated over the entire length of the reference nucleic acid or amino acid sequence, and the parameters are set so that a homology gap of up to 5% of the total number of nucleotides in the reference sequence is acceptable.

[0114] Another preferred method for determining the highest overall agreement between a query sequence (the sequence of the present invention) and a target sequence, also known as global sequence alignment, can be determined using a FASTDB computer program based on the algorithm of Brutlag et al. (Comp. App. Biosci. (1990) 6:237-245). In sequence alignment, both the query sequence and the target sequence are DNA sequences. RNA sequences can be compared by converting U to T. The result of the global sequence alignment is in percent identity units. Preferred parameters used in FASTDB alignment of DNA sequences to calculate percent identity are matrix=unitary, k-tuple=4, mismatch penalty=1, binding penalty=30, randomization group length=0, cutoff score=1, gap penalty=5, gap size penalty=0.05, window size=500 or the shorter of the length of the target nucleotide sequence.

[0115] For example, a polynucleotide having 95% "identity" to the reference sequence of the present invention is identical to the reference sequence, except that the polynucleotide sequence may contain up to 5 point mutations on average for every 100 nucleotides of the reference nucleotide encoding the polypeptide. In other words, to obtain a polynucleotide having a nucleotide sequence that is at least 95% identical to the reference nucleotide sequence, up to 5% of the nucleotides in the reference sequence may be deleted or substituted with other nucleotides, or up to 5% of the total nucleotides in the reference sequence may be inserted into the reference sequence. The query sequence may be an entire sequence, an ORF (open reading frame), or any fragment as specified herein.

[0116] The NCBI Basic Local Alignment Search Tool (BLAST) (Altschul et al. J. Mol. Biol. 215:403-410, 1990) is available for use with the sequence analysis programs blastp, blastn, blastx, tblastn, and tblastx from several suppliers, including the National Center for Biotechnology Information (NCBI, Bethesda, Md.) and the Internet. It can be accessed on the NCBI website along with instructions on how to use this program to determine sequence identity and sequence similarity.

[0117] The present invention also covers not only sequences having a particular sequence identity with the sequences disclosed herein, but also any sequence variant of any of the sequences disclosed herein. The present invention therefore also covers sequence variants in any context in which a particular sequence identity is referred to, and vice versa. A “sequence variant” refers to a polynucleotide or polypeptide that is different from the sequences disclosed herein (polynucleotide or polypeptide sequences) but retains their essential properties. Generally, variants are closely similar to the sequences disclosed herein and are identical in many regions.

[0118] Variants may include changes in the coating region, the non-coding region, or both. Sequence variants that involve changes resulting in silent substitutions, additions, or deletions, but which do not alter, for example, the properties or activity of the encoded polypeptide, are particularly preferred. Nucleotide variants produced by silent substitutions due to degeneracy of the gene code are preferred. Furthermore, variants in which 5 to 10, 1 to 5, or 1 to 2 amino acids are substituted, deleted, or added in any combination are also preferred.

[0119] The amino acid sequence of a variant polypeptide may differ from the amino acid sequence shown in Sequence ID No. 3 by the insertion or deletion of one or more amino acid residues, and / or the substitution of one or more amino acid residues with different amino acid residues. Preferably, the amino acid changes are minor in nature, such as conservative amino acid substitutions that do not significantly affect protein folding and / or activity, small deletions typically of 1 to about 30 amino acids, small amino-terminal or carboxyl-terminal extensions, such as an amino-terminal methionine residue, small linker peptides of up to about 20 to 25 residues, or small extensions that enhance purification by altering net charge or another function, such as a polyhistidine tract, antigenic epitope, or binding domain. Examples of conservative substitutions are found in the groups of basic amino acids (arginine, lysine, and histidine), acidic amino acids (glutamic acid and aspartic acid), polar amino acids (glutamine and asparagine), hydrophobic amino acids (leucine, isoleucine, and valine), aromatic amino acids (phenylalanine, tryptophan, and tyrosine), and small amino acids (glycine, alanine, serine, threonine, and methionine). Amino acid substitutions that do not generally alter specific activity are well known in the art and are described, for example, by H. Neurath and RLHill, 1979, In, The Proteins, Academic Press, New York. The most common exchanges are Ala / Ser, Val / Ile, Asp / Glu, Thr / Ser, Ala / Gly, Ala / Thr, Ser / Asn, Ala / Val, Ser / Gly, Tyr / Phe, Ala / Pro, Lys / Arg, Asp / Asn, Leu / Ile, Leu / Val, and their reverses. A "consecutive nucleotide" at a particular percentile means that the nucleotides are directly following each other. Therefore, 10% of the nucleotides in Sequence ID No. 2, which contains 60,000 nucleotides, could be nucleotides 1-6000, or nucleotides 2-6001, and so on.

[0120] For example, gene silencing using siRNA is described elsewhere, for example, in U.S. Patent Publication No. 20180016583, which is incorporated herein by reference in its entirety, particularly with respect to its disclosure and gene silencing. [Examples]

[0121] The objectives of the following examples are, firstly, to confirm the insertion of a transgene into a locus via CRISPR-mediated cleavage (Figure 2), and secondly, to compare the effects of homology on a transgene sequence with an integrated locus with those without homology to the transgene (Figures 3A, 3B, 4A, 4B, 5A, 5B). In addition, the upregulation / downregulation and stimulating / inhibiting effects on DNA repair pathways, such as homologous recombination (HR) or alternative end joining (Alt-EJ), were compared (Figure 6).

[0122] [Table 4]

[0123] Example 1: Target insertion into the ERV C 109F gene locus This example illustrates the integration of a transgene into the ERV C 109F locus (SEQ ID NOs: 1, 2).

[0124] CHO-M cells were transfected with vectors targeting each of the genomic ERV C 109F loci (Figures 3A, 3B, 4A, 4B). Such vectors were either homologous sequence-free (Figures 3A, 3B) or homologous sequence-containing (Figures 4A, 4B). For both the "non-homologous" and "homologous" approaches, the non-homologous end-joining (NHEJ) DNA repair pathway was inhibited using the chemical inhibitor Nu7441. Furthermore, each approach included experiments aimed at evaluating the presence or absence of homologous recombination (HR) or alternative end-joining (Alt-EJ) stimulation.

[0125] In either approach, there were four possibilities for transgene integration at the defined ERV C 109F locus: i) integration without a transgene, or ii) integration in the wild-type allele of the ERV C 109F locus, or iii) integration in the ERV-C 109F allele, or finally iv) integration in both alleles of (ii) and (iii), which may be referred to herein as “both loci.”

[0126] These results were obtained using three TaqMan qPCR assays developed to determine the category of each clone. These assays are described in relation to Figures 5A and 5B, as well as Figure 12.

[0127] Figure 6 shows the percentage of clones falling into categories (i), (ii), (iii), and (iv) as defined above. The figure particularly shows the targeting efficiency of CRISPR targets compared to random integration. Clones obtained by non-targeted integration correspond to clones that show a fluorescence signal similar to the signal obtained for CHO-M wild-type cells, suggesting that the transgene integration occurred in random genomic sequences because it did not target the ERV locus.

[0128] In this example, it was observed that HR stimulation and NHEJ inhibition resulted in a higher frequency of transgene integration in the ERV allele when DNA homology-based transgene integration was performed compared to when DNA homology was not used. This may reflect the fact that homology assists targeted transgene integration by homologous recombination. However, the highest frequency of targeted integration in both alleles occurred when expression vectors without homology were used, as well as during Alt-EJ stimulation and NHEJ inhibition (Figure 6).

[0129] In conclusion, a highly efficient process for transgene integration using CRISPR and gRNA is designed here, resulting in 70%–90% higher targeted integration compared to untargeted integration. Furthermore, stimulation of either the HR or Alt-EJ pathway increased the efficiency of targeted integration in this example, for integration in alleles including ERV or both alleles, depending on the presence or absence of DNA homology.

[0130] Figures 7A and 7B show histograms of chromosome distribution frequencies of DNA-FISH signals in Pool 1 and Pool 3, respectively (see Table 4 for pool composition): Cells from the Pool 1 and Pool 3 populations were blocked at metaphase using colesemide and spread on glass slides. DNA-FISH experiments were then performed on each sample using probes targeting promoters that activate GFP (green fluorescent protein) expression, and images were collected using a confocal microscope, Zeiss LSM800. Finally, the images were analyzed using a karyotype analyzer to obtain karyotypes. Targeted integration on chromosome 15 (Chr15) and chromosome 9 (Chr9) was enriched by up to 60% and 19.5% in Pool 1, and up to 45% and 32% in Pool 3. The n value indicates the number of karyotypes analyzed.

[0131] In each pool, multiple transgene integration sites were observed on other chromosomes, but at a much lower frequency than the total number of random transgene integration events, which accounted for approximately 20%.

[0132] Finally, these results confirm that CRISPR cleavage, combined with NHEJ inhibition and activation of HR, or in particular the Alt-EJ MMEJ mechanism, enables highly efficient integration of target transgenes on chromosomes 15 and 9.

[0133] These results indicate that targeted integration can be highly efficient upon Alt-EJ activation, with up to approximately 80% of the total integration occurring in the portions of chromosome 9 and / or chromosome 15 containing the two alleles of the ERV-109F locus, namely the ERV-deficient wild-type and the ERV C 109F-containing allele. Since all CHO cells undergo extensive chromosomal rearrangement after isolation from native Chinese hamster cells and selection for optimal growth characteristics in vitro, CHO-M cells do not possess homologous sequences for all chromosomal loci of homologous chromosomes. This explains why homologous sequences in the genome may be located on different chromosomes, as observed at the loci examined.

[0134] Figure 7C shows representative karyotypes of cells obtained from Pool 1, with transgene insertions occurring on chromosomes 9 and 15, respectively. DNA is stained gray, and DNA-FISH probes for telomere repeats (sequence TTAGGGTTAGGGTTAGGG [SEQ ID NO: 60]) and promoters that activate GFP expression are stained white. White arrows indicate the location of FISH signals for ERV and wild-type alleles.

[0135] Example 2: Increased transgene expression by targeting the ERV C 109F gene locus. This example illustrates that the ERV C 109F locus enables enhanced and stable expression of exogenous transgenes. Figure 8 shows the analysis of GFP-expressing CHO clones based on FITC fluorescence analysis. 340,000 cells were transfected under different conditions and seeded in semi-solid medium in the presence of antibiotic selector two days after transfection. Forty-two clones were selected (ClonePix®) based on fluorescence intensity and cultured. Nine days after selection, FITC was measured (Cytoflex®) in 2000 cells per clone, and the type of transgene integration event was determined for each clone by qPCR analysis (TaqMan®). The 42 clones were categorized into four groups, as described above, based on whether the transgene was located as a wild-type allele, an allele containing ERV C 109F, both alleles, or an untargeted integration corresponding to an out-of-target genomic integration event.

[0136] Figures 8A, B, C, and D show the fluorescence intensity (FITC) according to the type of transgene locus integration. For each type of integration, the fluorescence indication allowed us to investigate whether one locus provides higher transgene expression compared to others, and whether the transgene expression level changes due to modification of the repair pathway.

[0137] The results, overall, suggest that targeting the ERV C 109F locus mediates higher FITC fluorescence compared to random genomic integration represented by untargeted integration. Furthermore, it can be understood that the modification of the DNA repair pathway resulted in higher fluorescence levels for integration of the ERV C 109F locus into the wild-type allele compared to integration into the ERV allele (e.g., FITC fluorescence of clones isolated from pools 1 and 3, Figures 8A and 8C). Integration of the ERV C 109F locus into the wild-type allele also resulted in higher fluorescence compared to integration into both alleles. Therefore, in certain embodiments of the invention, integration of the ERV C 109F locus into the wild-type allele without integration into the allele occupied by the corresponding ERV is preferred. This embodiment is particularly preferred when potential negative effects may occur after integration into the ERV-containing allele, and / or when the wild-type ERV-deficient allele is desirable for higher and more stable transgene expression than the ERV-containing allele. Those skilled in the art will understand that, even in embodiments without transgene integration in the ERV-containing allele, the ERV sequence has been modified in certain embodiments to eliminate or reduce the release of viral particles from cells, so that the release of viral particles is no longer detectable (see Duroy, 2020).

[0138] Furthermore, the modification of the DNA repair pathway is also thought to enhance the expression of the transgene at this locus, particularly when this modification is applied, resulting in significantly higher expression than without Alt-EJ modification. Therefore, in certain embodiments, it is particularly preferable (see, for example, the dashed line showing the median fluorescence intensity in Figures 8A and 8B).

[0139] Figure 9 shows the results obtained from different transfections and clones from different batches obtained from pools 3 and 4. These results show fluorescence obtained from clones generated using DNA homology-based targeted integration, compared to cells generated with HR adjustment (pool 3, Figure 8C) compared to cells without HR adjustment (pool 4, Figure 8D).

[0140] These results validate previous findings, particularly by demonstrating that modification of the HR repair pathway can also be used to generate cell clones exhibiting increased transgene expression when the transgene is integrated into the wild-type allele or both alleles, compared to integration into only the ERV-containing allele or non-targeted integration. These results further verify that modification of DNA repair mechanisms is advantageous for obtaining increased transgene expression.

[0141] Example 3: Expression of the complex protein This example demonstrates that by configuring the protein to enable GFP expression when incorporated into a wild-type allele lacking ERV109F on chromosome 9 and / or a homologous ERV109F-containing allele on chromosome 15 (Figure 10), it is also possible to express the trastuzumab therapeutic protein.

[0142] First, CHO cells were transfected with an expression vector for CRISPR-mediated cleavage along with a trastuzumab expression vector (Figure 11B). Some transfections also included vectors mediating the expression of alternative end-joining mechanisms, such as POLQ, MRE11, and CIRBP proteins that enhance the microhomology-mediated end-joining (MMEJ) pathway (Figure 11B). This was done with the aim of promoting the integration of the trastuzumab expression vector at chromosomal CRISPR cleavage sites, as the trastuzumab expression vector does not exhibit significant DNA sequence homology to chromosomal loci due to the absence of homologous sequences. Cells containing the genomically integrated transgene sequence were selected for puromycin antibiotic resistance, and then clonal populations were isolated using a ClonePix® FL Imager from Molecular Devices, LLC.

[0143] The derived clones were then analyzed by PCR assay using primers indicated by arrows in Figure 12 to determine the genomic locus in which the transgene was incorporated: Amplification using the primers indicated by arrows in Figure 12B indicates that no transgene insertion occurred in the ERV-deficient wild-type allele (referred to as the ERV-free locus in the figure). The primers shown in Figure 12C evaluate whether the ERV is still inserted into its normal locus without the transgene insertion. Amplification using all primers shown in Figure 12D indicates that both alleles are intact and that the transgene integration must have occurred elsewhere in the genome. In Figure 12A, no amplification with any primer pair indicates that there is an insertion in both alleles. Thus, four types of clones were identified, including those with a random location, an allele on chromosome 15 containing ERV109F, an ERV109F-deficient wild-type allele on chromosome 9, or a transgene incorporated in either the alleles on chromosomes 15 or 9. As predicted, clones in which the Tras transgene was introduced into the allele containing ERV or both alleles, and therefore the ERV109F sequence was replaced with the Tras coding sequence, exhibited the most potent reduction in ERV expression, up to a reduction of more than 17-fold, reaching undetectable levels within the PCR background (Figure 13A-C).

[0144] Clones were then screened for trastuzumab selection, and representative clones expressing trastuzumab at low (Figure 14A), moderate (Figure 14B), and high (Figure 14C) levels were selected for each clone type. As observed with respect to GFP expression, the highest levels of trastuzumab production were obtained with the integration of the Tras coding sequence into the ERV-deficient chromosome 9 locus, followed by clones with the Tras sequence integrated into both loci. Production levels obtained by targeted integration in ERV109F-containing and / or-deficient wild-type alleles were significantly higher than those obtained by random genomic integration. Overall, these findings indicate that the genomic environment of the ERV locus is highly favorable for gene expression, and that the most productive clones are obtained with the integration of the transgene into the ERV-deficient chromosome 9 wild-type allele.

[0145] Next, we evaluated whether the high Tras titers obtained in the supernatant of small-scale, short-term non-feed cultures could be transferred to therapeutic protein-producing cultures. Specific production capacity levels obtained at 6-10 day intervals in small 96-well plates (Figures 15D-F), as well as titers obtained from the supernatant of fed-batch cultures placed in shaking 24-well plates (Figures 15A-C), indicated that optimal Tras expression was achieved by incorporating the transgene into the ERV-deficient wild-type allele. Upscaling

[0146] The ability to produce the clone that mediates the highest titer for each type of genomic integration (Figures 15C and 15F) was evaluated on a larger scale using fed-batch cultures produced in an Ambr® 15 bioreactor. A specific production capacity exceeding 20 picograms of Tras IgG secreted per cell per day was obtained for a clone with the transgene integrated into the ERV-deficient wild-type allele (Figure 16B). This clone also mediates the highest titer of antibody released into the cell culture supernatant (Figure 16A).

[0147] In summary, surprisingly, the optimal target integration locus for high transgene expression and optimal production of therapeutic protein was found to be a chromosomal locus containing a highly expressed ERV, more preferably the chromosome 15 genomic allele containing ERV109F, and even more preferably the ERV109F-deficient wild-type allele on chromosome 9. Expression vectors of alternative end-binding factors such as MRE11 and PolQ MMEJ proteins (see Figure 11A) could be co-transfected with the therapeutic protein vector to prioritize high-frequency target integration of non-homologous plasmid sequences at this preferred genomic locus, and this co-transfecting was performed. Dual integration of the therapeutic protein expression vector in both ERV-containing and ERV-deficient wild-type alleles is also highly preferred, allowing for high levels of therapeutic protein expression, as well as cessation of expression of potentially harmful retroviral sequences that lead to the release of viral particles into the cell supernatant, along with the therapeutic protein of interest, as a result of a single transfection. material and method cell culture

[0148] Chinese hamster ovary (CHO-M) cells adapted to a suspension were maintained in serum-free BalanCD CHO medium (Irvine Scientific) supplemented with L-glutamine (GE Healthcare). The CHO-M viable cell density and fluorescence signal of green fluorescently transfected cells were evaluated using a Cytoflex flow cytometer (Beckman Coulter). Cells were cultured in 50 ml C50 bioreactor tubes (TPP, Switzerland) at 37°C, 5% CO2, and in a humidified incubator with a stirring speed of 180 rpm, and subcultured every 3-4 days. Plasmid construction

[0149] Two EGFP (highly sensitive green fluorescent protein) expression vectors were used in this study. Both vectors possess the same eukaryotic expression cassette, consisting of an antibiotic resistance cassette followed by an EGFP expression cassette, along with a downstream SV40 enhancer and SELEXIS Genetic Element (SGE). SELEXIS SGEs are intrinsic epigenetic DNA-based elements that control the dynamic organization of chromatin across all mammalian cells. They enable transcriptional enhancement by isolating the incorporated transgene from the silencing effect of the surrounding chromatin.

[0150] Duroy et al. 2019 described that of the 173 type C ERVs identified in the CHO genome, only one can be transcribed and produce viral particles present in the CHO culture supernatant. Because this ERV sequence exists only in a hemizygous state in the CHO cell genome, the integration locus is specific (Figure 1). This means that one allele exists in a CHO genome without ERV integration (wild-type (WT) or ERV-deficient wild-type allele), while the other allele exists in a homologous DNA sequence within the CHO genome that can find the integrated ERV (referred to as the ERV-C 109F-containing allele, or ERV-C 109F allele). Only the last allele will express the corresponding viral sequence in the cell, resulting in the presence of ERV-C 109F mRNA in the cytoplasm.

[0151] One of the EGP-containing vectors used in this application further includes two 750 bp-long homology sequences located on either side of the allele breakpoint and CRISPR break site (5' and 3' homology arms) shown in Figure 2, corresponding to DNA sequences derived from genomic loci surrounding the ERV C 109F-containing and ERV-deficient wild-type alleles (SEQ ID NO: 6 and SEQ ID NO: 7).

[0152] Two sets of CRISPR vectors were used in addition to the EGFP vector to introduce site-specific DSBs. CRISP 16 and CRISP 17 DSBs (SEQ ID NOs. 8 and 9) are preferably repaired by the homologous recombination pathway using 5' and 3' homology arms present as homologous sequences in the vector and the wild-type allele and / or ERV-containing allele of the locus. CRISP 50 and CRISP 51 DSBs (SEQ ID NOs. 10 and 11) are preferably repaired by the Alt-EJ pathway using microhomology sequences present in the vector and the wild-type allele. Transfection and single-cell isolation

[0153] To inhibit the NHEJ DNA repair pathway, CHO-M cells growing in suspension were pre-treated with 0.5 μM Nu7441 to inhibit DNA-PKc. Cells stimulated via the homologous recombination (HR) pathway were further treated with 1 μM RS-1. Pre-treated cells were transfected with two expression vectors containing highly sensitive green fluorescent protein (EGFP) coding sequences, as shown in Figures 3 and 4, for ease of detection (340,000 cells / transfection). To stimulate the Alt-EJ repair pathway, i.e., the MMEJ pathway, cells were also transfected with the hMRE11 (SEQ ID NO: 25), cgPOLQ (SEQ ID NO: 26), and hCIRBP (SEQ ID NO: 27) genes, cloned into separate expression vectors. Meanwhile, to stimulate the HR repair pathway, cells were also transfected with the hMRE11 and hRAD51 (SEQ ID NO: 28) genes, cloned into separate expression vectors.

[0154] One day after transfection, cells were centrifuged and the medium was changed to remove Nu7441 and RS-1. Two days after transfection, cells were seeded in a semi-solid medium containing 3 μg / ml of puromycin at a cell density of 5000 cells / ml. After growing in the semi-solid medium for 10 days, 42 EGFP-expressing clones per transfection were selected based on fluorescence intensity (ClonePix, Molecular Devices) and cultured in BalanCD CHO medium. EGFP expression levels (FITC) were measured in 2000 cells per clone 9 days after selection (experience with DNA repair pathway stimulation) and 6 days after selection (experience without DNA repair pathway stimulation) (Cytoflex®). Results are presented categorized by qPCR analysis (TaqMan®). TaqMan(registered trademark) qPCR assay DNA extraction

[0155] Genomic DNA (gDNA) was extracted from 2 × 10⁶ cells using the CellsDirect One-Step qRT-PCR Kit (Thermo Fisher Scientific®) according to the manufacturer's instructions. gDNA quantification was performed using a NanoDrop® spectrophotometer (Thermo Fisher Scientific®). qPCR assay

[0156] Three Taqman® qPCR assays were designed (Figures 5A and 5B; probes and primers are shown in SEQ ID NOs. 29–37). The linearity and efficiency of all TaqMan qPCR assays were validated using a standard curve approach. All assays showed high linearity (=0.999) and very good efficiency (≧0.97).

[0157] qPCR was performed using the Rotor-Gene Multiplex PCR Kit® and FAM or HEX-labeled TaqMan® qPCR assay at QIAGEN's Rotor-Gene. Data analysis was performed using Rotor-Gene Q Series® software (v2.3.1).

[0158] To verify the absence or presence of a reference locus, TaqMan® specificity qPCR assays were performed for three loci using a standard curve approach with appropriate negative controls.

[0159] The absence or presence of a single amplicon in the CT range corresponding to the control allowed us to determine whether the target locus for the integration of the wild-type allele of the ERV-containing allele or the ERV integration locus was similar to that of untransfected CHO-M cells. The "yes or no" results of three different TaqMan assays allowed us to determine which category each clone belonged to. Fluorescence In-Situ Hybridization Experiment

[0160] Cells were blocked at metaphase using colsemid and spread on glass slides. DNA-FISH experiments were performed on each sample using probes targeting promoters that activate EGFP (green fluorescent protein) vector expression. Images were collected using a Zeiss LSM800 confocal microscope. Finally, the images were analyzed using a karyotype analyzer to obtain the karyotype. Generation and characterization of trastuzumab-producing cell clones

[0161] To enable targeted integration of an IgG-encoding vector into the ERV109F integrated genomic locus, CHO-M cells were co-transfected with PuroBT+_Tras_Hc and PuroBT+_Tras_Lc trastuzumab (Tras) immunoglobulin (IgG) expression plasmids (Figure 11B) and a CRISPR-sgRNA vector. Cells were selected for resistance to puromycin, and single-cell clones were isolated using a ClonePix® FL Imager from MOLECULAR DEVICES, LLC. CHO-M cell line and fed batch culture

[0162] Parental Selexis CHO-M cells stably expressing human monoclonal IgG1 antibody, and derived clonal cell lines, were cultured as follows: Seed train cultures were subculturified every 3-4 days before N-1 seeding. Four days before microbioreactor inoculation, CHO-M cultures were subculturified in 0.30 × 10⁶ volumes according to process requirements. 6 Cells were subcultured in a shaking flask at a seeding cell density of 1 cells / ml(N-1). The cells were cultured in chemically defined BalanCD Growth A® culture medium (IRVINE SCIENTIFIC, USA) supplemented with 6 mM L-glutamine (HyClone, USA) at 37.0°C, 5% CO2, and 120 rpm in an incubator (KUHNER, Germany). Total protein quantification assay

[0163] The LabChip LCGXII system (PERKIN ELMER, Inc.), an automated microfluidic capillary gel electrophoresis system, was used for the total protein assay. Samples containing proteins were mixed with amine-reactive fluorescent dyes that nonspecifically label proteins, and the proteins were detected by laser-induced fluorescence at the outlet of the separation channel. Characterization of clones regarding the effectiveness of targeted integration

[0164] The genomic integration sites of the Tras expression vector in IgG-producing cell clones were analyzed by q-PCR assays performed on genomic DNA. Quantitative PCR (q-PCR) was performed multiplexed using three Taqman probes. Two Taqman assays were designed to determine the presence of the ERV109F junction sequence between ERV and genomic DNA on both sides of the ERV integration locus on chromosome 15 (Chr). Lack of amplification indicated that one or more transgene copies were integrated into the ERV109F allele and that the ERV sequence was deleted (Figures 12A and 12B). A third Taqman assay was designed to assess the presence of the wild-type allele, i.e., the ERV-deficient chromosome 9 genomic sequence, to evaluate transgene integration at this locus (Figures 12A and 12C).

[0165] If no product was obtained from the three q-PCR assays, it could be inferred that the transgene copy was integrated into both alleles (Figure 12A). If all three Taqman assays yielded positive results, this indicated that both alleles were intact and that the sequence encoding Tras was therefore integrated elsewhere in the genome (Figure 12D). Thus, such clones were classified as those in which transgene integration occurred randomly. ERV109F expression

[0166] Total cellular RNA was extracted using QIAGEN's RNeasy® kit according to the manufacturer's protocol. Two DNAse treatments were performed during and after extraction. RNA was reverse transcribed to DNA using PROMEGA's GoScript® reverse transcriptase (RT) kit.

[0167] RT-qPCR assays of ERV109F RNA from Tras-producing clones and parental CHO-M cells were performed using a Taqman assay designed to detect the ERV109F long terminal repeat sequence, with cellular GAPDH housekeeping gene mRNA as the reference. The multiplier of ERV109F expression reduction was determined according to the delta-delta CT calculation method (Livak and Schmittgen, Analysis of Relative Gene Expression Data Using Real-Time Quantitative PCR and the 2-ΔΔCT Method, Methods 25, 402-408, 2001). This assay allowed for the determination of the decrease in ERV expression after transgene integration at the ERV locus, thereby further verifying transgene integration and ERV sequence detection at the ERV109F locus. Assays for cell clone culture and trastuzumab antibody production capacity

[0168] The cell culture process used to determine Tras production in fed-batch culture was as follows: Cell growth and production performance were evaluated using classical fed-batch static culture under agitation in 96-deep-well or 24-deep-well plates. Fed-batch culture was also performed using the Ambr15® automated microscale bioreactor system (SARTORIUS Stadium, Germany) equipped with a cooling system that allows for temperature shifts. All cultures were performed with 40% dissolved oxygen (DO), agitation speed of 1000–1400 rpm, maintaining a temperature of 36.5°C, then shifting to 33.0°C (time shift depending on seeding density), controlling the pH to 6.90±0.10, then shifting to 7.00±0.20 using CO2 and 1M carbonate (time shift depending on seeding density).

[0169] Fed batch cultures in 24-deep wells or 96-deep wells were seeded at a target cell density of 300,000 cells / ml using culture volumes of 3 ml and 250 ml, respectively. In a microbioreactor, according to the Ambr15 seeding density process, 1.00 × 10⁶ cells were seeded at an initial working volume of 13 mL. 6 Cells were seeded at a target cell density of 100 cells / mL. Feed supplements, Cell Culture Supplement 1 and Cell Culture Supplement 2, were added to the culture at various time points depending on the seeding density process. Glucose solution (SIGMA ALDRICH, USA) was added daily based on the glucose concentration, as needed to maintain good cell viability and high levels of production. Microbioreactor samples were collected daily for cell counting and determination of viable cell density (VCD). Cell viability was measured using Bioprofile® FLEX2 (NOVA BIOMEDICAL, USA). Cells were grown for up to 14 days.

[0170] Those skilled in the art will understand that the above description is not limiting, but rather provides examples of certain embodiments of the present invention. Those skilled in the art may, by the guidance provided above, conceive of a wide range of alternative embodiments not specifically described herein. Bibliography Duroy et al., Characterization and mutagenesis of Chinese hamster ovary cells endogenous retroviruses to inactivate viral particle release. Biotechnol Bioeng. (published online October 2019) US Patent Publication 2020-0109421A1 Havecker et al., The diversity of LTR retrotransposons, Genome Biology 2004, 5:225 (2004) Bell and Felsenfeld, Stopped at the border: boundaries and insulators. Curr Opin Genet Dev 9, 191-198 (1999). Zahn-Zabal et al., Development of stable cell lines for production or regulated expression using matrix attachment regions , J. Biotechnol. 87: 29-42 (2001) Kim et al.,Regulation of Swi6 / HP1-dependent heterochromatin assembly by cooperation of components of the mitogen-activated protein kinase pathway and a histone deacetylase Clr6.J. Biol. Chem.; 279: 42850-42859 (2004) Agarwal et al., Scaffold attachment region-mediated enhancement of retroviral vector expression in primary T cells. J Virol 72, 3720-3728 ((1998) Castilla et al., Engineering passive immunity in transgenic mice secreting virus-neutralizing antibodies in milk, Nature Biotech. Vol. 16, 349-354 (1998) Taher and Ovcharenko, Variable locus length in the human genome leads to ascertainment bias in functional inference for non-coding elements, Bioinformatics, Vol. 25 no. 5 2009, pages 578-584 (2009) Leahy et al., Identification of a highly potent and selective DNA-dependent protein kinase (DNA-PK) inhibitor (NU7441) by screening of chromenone libraries. (Bioorg. Med. Chem. Lett. 14:6083-6087 (2004) Willmore et al., A novel DNA-dependent protein kinase inhibitor, NU7026, potentiates the cytotoxicity of topoisomerase II poisons used in the treatment of leukemia. Blood 103 (12):4659-65 (2004) Maruyama et al., Increasing the efficiency of precise genome editing with CRISPR-Cas9 by inhibition of nonhomologous end joining, Nat. Biotechnol. 33 :538- 542 (2015) Robert et al., Pharmacological inhibition of DNA- PK stimulates Cas9-mediated genome editing, Genome Med 7:93 (2015) Dittmann et al., Inhibition of radiation- induced EGFR nuclear import by C225 (Cetuximab) suppresses DNA-PK activit, Radiother and Oncol 76: 157 (2005) Shibata et al, inhibitors of PolQ, inhibitors of CtI, Molec. Cell 53:7-18 (2014) Sfeir and Symington, "Microhomology-Mediated End Joining: A Back-up Survival Mechanism or Dedicated Pathway?" Trends Biochem Sci 40:701-714 (2015) Urnov F., et al., Highly efficient endogenous human gene correction using designed zinc-finger nucleases Nature 435:646-651 (2005) Mussolino et al., A novel TALE nuclease scaffold enables high genome editing activity in combination with low toxicity Nucl. Acids Res. 39(21):9283-9293 (2011) Gao et al., DNA-guided genome editing using the Natronobacterium gregoryi Argonaute, Nature Biotechnology, (published online May 2, 2016) Swarts et al., DNA-guided DNA interference by a prokaryotic Argonaute, Nature 507(7491): 258-261 (2014) Sheng et al,Structure-based cleavage mechanism of Thermus thermophilus Argonaute DNA guide strand-mediated DNA target cleavage, Proc. Natl. Acad. Sci. U.S.A. III, 652 (2013). Boissel et al., MegaTALs: a rare-cleaving nuclease architecture for therapeutic genome engineering, Nucleic Acids Research 42 (4):2591 -2601 (2013) Kay etal., A bacterial effector acts as a plant transcription factor and induces a cell size regulator, Science 318:648-651 (2007) Computational Molecular Biology, Lesk, A. M., ed., Oxford University Press, New York (1988) Biocomputing: Informatics and Genome Projects, Smith, D. W., ed., Academic Press, New York (1993) Computer Analysis of Sequence Data, Part I, Griffin, A. M., and Griffin, H. G., eds., Humana Press, New Jersey (1994) Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press (1987) Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York (1991) Carillo, H. & Lipton, D., The Multiple Sequence Alignment Problem in Biology, SIAM J Applied Math 48:1073 (1988) Smith and Waterman, Advances in Applied Mathematics 2:482-489 (1981) Brutlag et al., Comp. App. Biosci. 6:237-245 (1990) Altschul et al., J. Mol. Biol. 215:403-410 (1990) H. Neurath and R. L. Hill, In: The Proteins, Academic Press, New York (1979)

Claims

1. Manipulated CHO cells, wherein within the genome of the manipulated CHO cells, At least one gene locus containing an insertion site for an endogenous retrovirus (ERV) sequence, wherein the insertion site is (i) a) the ERV C 109F sequence, or b) a portion of the sequence in a), and / or (ii)(i) Allele wild-type corresponding sequence Includes, (i) comprises at least nucleotides 29021-40247 of SEQ ID NO: 1, or a sequence having 95%, 98%, or 99% sequence identity with nucleotides 29021-40247 of SEQ ID NO: 1, and (ii) comprises at least nucleotides 29020-31020 of SEQ ID NO: 2, or a sequence having 95%, 98%, or 99% sequence identity with nucleotides 29020-31020 of SEQ ID NO: 2, or (i) However, the locus is characterized in that it includes at least a sequence having 95%, 98%, or 99% sequence identity with nucleotides 29521 to 39747 of SEQ ID NO: 1, or nucleotides 29521 to 40247 of SEQ ID NO: 1, and (ii) includes at least a sequence having 95%, 98%, or 99% sequence identity with nucleotides 29520 to 30520 of SEQ ID NO: 2, or nucleotides 29520 to 30520 of SEQ ID NO: 2, and At least one transgene encoding at least one transgene expression product incorporated into the at least one gene locus and Includes, The manipulated CHO cells, comprising (i) and (ii) on chromosome 15 and chromosome 9, respectively.

2. The engineered CHO cell according to claim 1, comprising (i) and (ii), wherein at least one transgene is incorporated into (ii).

3. The manipulated CHO cell according to claim 2, wherein at least one of the introduced genes is not incorporated into (i).

4. The manipulated CHO cell according to claim 2, wherein at least one of the introduced genes is also incorporated into (i).

5. A cell population comprising the manipulated CHO cells according to any one of claims 1 to 4, wherein the at least one transgene is incorporated into more than 20%, more than 30%, or more than 40% of (i) and / or (ii) of the cells in the cell population.

6. The manipulated CHO cells or cell population according to any one of claims 1 to 4, or the cell population according to claim 5, wherein the manipulated CHO cells or cell population lack expression of endogenous retroviral sequences (ERVs) and / or do not contain detectable viral particles in the culture supernatant of the cell population.

7. The manipulated CHO cells according to any one of claims 1 to 4, or the cell population according to claim 5, wherein the at least one transgene expression product is the protein of the choice.

8. The manipulated CHO cells or cell population according to any one of claims 1 to 4, or the cell population according to claim 5, wherein the manipulated CHO cells or cell population express the target protein at an amount of picograms, μg / l, or mg / l per cell per day that is at least 1.5 times, 2 times, 2.5 times, 3 times, or more than the amount of the target protein that would be expressed if the at least one transgene were incorporated into the genome outside the at least one locus, per unit time.

9. The manipulated CHO cells or cell population according to any one of claims 1 to 8, wherein the manipulated CHO cells are transfected with one or more vectors containing one or more genes from Table 1-1 to 1-5 or SEQ ID NOs. 25-28, 38-58, and / or 59, or genes having a sequence identity of at least 90% with SEQ ID NOs. 25-28, 38-58, and / or 59, and which encode proteins involved in the DNA repair pathway (DRP). Table 1-1 Table 1-2 Table 1-3 Table 1-4 Table 1-5

10. The cytoplasm of the manipulated CHO cells is reacted to one or more exogenous chemical inhibitors and / or stimulants of DRP, or NU7441, olaparib, DNA ligase IV inhibitor, Scr7, KU-0060648, anti-EGFR antibody C225 (cetuximab), compound 401 (2-(4-morpholinyl)-4H-pyrimido[2,l-a]isoquinoline-4-one), vanillin, waltmannin, DMNB, IC87361, LY294002, OK-1035, CO15, NK314, PI103 hydrochloride, etc. The engineered CHO cells or cell population according to any one of claims 1 to 8, further comprising an NHEJ inhibitor selected from the group consisting of combinations thereof, or an MMEJ inhibitor selected from the group consisting of myrin, a derivative of myrin, a PolQ inhibitor, a CtIP inhibitor, or combinations thereof, or an HR inhibitor, RI-1 or BO2, or an HR stimulant or RS-1, or an NHEJ stimulant or IP6, or any combination of any one of the above-mentioned inhibitors and / or stimulants.

11. The engineered CHO cells or cell population according to any one of claims 1 to 10, wherein the introduced gene is incorporated into a landing pad sequence containing a selectable marker gene.

12. A method for incorporating a transgene into the genome of CHO cells, (a) To provide at least one transgene as part of a vector, plasmid or viral vector containing the at least one transgene, wherein the vector contains the transgene at at least one locus of the CHO cell which includes an insertion site for an endogenous retrovirus (ERV) sequence, and the insertion site is (i) a) the ERV C 109F sequence, or b) a portion of the sequence in a), and / or (ii)(i) Allele wild-type corresponding sequences, Includes, (i) comprises at least nucleotides 29021-40247 of SEQ ID NO: 1, or a sequence having 95%, 98%, or 99% sequence identity with nucleotides 29021-40247 of SEQ ID NO: 1, and (ii) comprises at least nucleotides 29020-31020 of SEQ ID NO: 2, or a sequence having 95%, 98%, or 99% sequence identity with nucleotides 29020-31020 of SEQ ID NO: 2, or, (i) comprises a sequence having 95%, 98%, or 99% sequence identity with at least nucleotides 29521 to 39747 of SEQ ID NO: 1, or nucleotides 29521 to 40247 of SEQ ID NO: 1, and (ii) comprises a sequence having 95%, 98%, or 99% sequence identity with at least nucleotides 29520 to 30520 of SEQ ID NO: 2, or nucleotides 29520 to 30520 of SEQ ID NO: 2, characterized in that it is incorporated into a gene locus, provided, or (b) To provide at least one transgene and at least one nuclease and / or nickase, wherein the nuclease and / or nickase incorporates a double-strand and / or single-strand break in order to incorporate the transgene, at at least one locus of the CHO cell which includes an insertion site for an endogenous retrovirus (ERV) sequence, (i) a) the ERV C 109F sequence, or b) a portion of the sequence in a), and / or (ii)(i) Allele wild-type corresponding sequences, Includes, (i) comprises at least nucleotides 29021-40247 of SEQ ID NO: 1, or a sequence having 95%, 98%, or 99% sequence identity with nucleotides 29021-40247 of SEQ ID NO: 1, and (ii) comprises at least nucleotides 29020-31020 of SEQ ID NO: 2, or a sequence having 95%, 98%, or 99% sequence identity with nucleotides 29020-31020 of SEQ ID NO: 2, or, (i) comprises a sequence having 95%, 98%, or 99% sequence identity with at least nucleotides 29521-39747 of SEQ ID NO: 1, or nucleotides 29521-40247 of SEQ ID NO: 1, and (ii) comprises a sequence having 95%, 98%, or 99% sequence identity with at least nucleotides 29520-30520 of SEQ ID NO: 2, or nucleotides 29520-30520 of SEQ ID NO: 2, characterized in that it is introduced into a gene locus and provided. Transfecting the CHO cells with at least one of the introduced genes, Methods that include...

13. The method according to claim 12, wherein at least one of the transgenes in (b) is provided as part of the vector.

14. The method according to claim 12, wherein the nuclease and / or nickase is encoded by at least one vector.

15. The method according to claim 12, further comprising providing at least one vector encoding at least one target element that guides the at least one nuclease and / or nickase.

16. The method according to claim 12, further comprising upregulating or stimulating at least one first DNA repair pathway (DRP) of the CHO cells, or similarly doing the reverse.

17. The method according to claim 12, further comprising downmodulating or stimulating at least one second DNA repair pathway (DRP) of the CHO cells, or vice versa.

18. The method according to claim 12, further comprising isolating an engineered CHO cell containing the at least one transgene incorporated into the gene locus.

19. The aforementioned CHO cells further, - One or more genes from Table 2-1 to 2-5 or SEQ ID NOs. 25 to 28, 38 to 58, and / or 59, or genes with sequences having at least 90% sequence identity with SEQ ID NOs. 25 to 28, 38 to 58, and / or 59, which are transfected with genes encoding proteins involved in the DNA repair pathway (DRP), and / or - The method according to claim 12, wherein the CHO cells are brought into contact with a chemical substance that affects the DRP. Table 2-1 Table 2-2 Table 2-3 Table 2-4 Table 2-5

20. The method according to claim 19, wherein the CHO cells are transfected with the gene as part of one or more further vectors.

21. The method according to claim 12, wherein the CHO cells are transfected with one or more vectors that include and express SEQ ID NOs. 25-28, 38-58, and / or 59, or SEQ ID NOs. 25-28.

22. The at least one nuclease and / or nicasse described above, Transposases, integrases, recombinases or site-specific recombinases, nickases, or nucleases or site-specific nucleases, fusion proteins containing a programmable DNA-binding domain and a nuclease domain, or Any combination of these, or The method according to any one of claims 12 to 21, wherein the homing endonuclease, restriction enzyme, zinc finger nuclease or zinc finger nickase, meganuclease or mega nickase, transcription activator-like effector nuclease or transcription activator-like effector nickase, RNA guide nuclease or RNA guide nickase, DNA guide nuclease or DNA guide nickase, mega TAL nuclease, BurrH nuclease, ARCUS nuclease, modified or chimeric versions or variants thereof, and any combination thereof.

23. The method according to any one of claims 12 to 21, wherein the at least one nuclease and / or nickase is a zinc finger nuclease or zinc finger nickase, a transcription activator-like effector nuclease or transcription activator-like effector nickase, an RNA guide nuclease or RNA guide nickase.

24. The method according to claim 22 or 23, wherein the RNA guide nuclease or RNA guide niccas is part of a CRISPR-based system, restriction enzyme, or combination thereof.

25. A kit for introducing at least one transgene into CHO cells, One container contains at least one locus containing an insertion site for an endogenous retrovirus (ERV) sequence selected from SEQ ID NOs: 1 and 2, or at least one gene containing (i) nucleotides 29021-40247 of SEQ ID NO: 1, or a sequence having 95%, 98%, or 99% sequence identity with nucleotides 29021-40247 of SEQ ID NO: 1, and (ii) at least one gene containing nucleotides 29020-31020 of SEQ ID NO: 2, or a sequence having 95%, 98%, or 99% sequence identity with nucleotides 29020-31020 of SEQ ID NO:

2. A vector encoding a nuclease and / or nickase that targets a stromatous locus, or at least one gene locus comprising (i) a sequence having 95%, 98%, or 99% sequence identity with nucleotides 29521-39747 of SEQ ID NO: 1, or nucleotides 29521-40247 of SEQ ID NO: 1, and (ii) a sequence having 95%, 98%, or 99% sequence identity with nucleotides 29520-30520 of SEQ ID NO: 2, or nucleotides 29520-30520 of SEQ ID NO: 2, or an ERV sequence or SEQ ID NO: 3 incorporated into the insertion site. In a separate container, at least one stimulant and / or inhibitor of the DNA repair pathway (DRP), and / or The following describes one or more vectors comprising one or more genes encoding one or more DRP proteins from Tables 3-1 to 3-5 or SEQ ID NOs. 25-28, 38-58, and / or 59, or genes having a sequence with at least 90% sequence identity with SEQ ID NOs. 25-28, 38-58, and / or 59, and encoding genes that encode proteins involved in DRP; and a method for transfecting CHO cells with the at least one transgene using the at least one nuclease and / or nicasse, and the at least one stimulant and / or inhibitor. A kit that includes this. Table 3-1 Table 3-2 Table 3-3 Table 3-4 Table 3-5

26. The kit according to claim 25, further comprising at least one vector encoding at least one target element that guides the at least one nuclease and / or nickase.

27. The kit according to claim 25, comprising in a separate container at least one vector encoding at least one target element for guiding the at least one nuclease and / or nickase.

Citation Information

Patent Citations

  • Improved eukaryotic cells for protein production and methods for their production

    JP2018537986A