Glutamine synthase markers for recombinant expression

By using GS derived from koalas, ostriches, snakes, pigeons, junglefowl, gerbils, or night bats as selectable markers, the problem of low selection efficiency in existing technologies is solved, enabling efficient recombinant protein production and identification of high-yield cell clones.

CN121335982APending Publication Date: 2026-01-13SHANGHAI ZHENGE BIOTECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202480022874.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-29
Filing Date
2024-01-26
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing glutamine synthase (GS) derived from Chinese hamsters has low selection efficiency and yield as a selectable marker in recombinant protein production, and cannot effectively support the development of high-yield cell lines.

Method used

Using GS derived from koalas, ostriches, snakes, pigeons, junglefowl, gerbils, or night bats as selectable markers, and encoding their nucleotide sequences for use in an expression system for recombinant protein production, combined with mRNA destabilizing elements and degradation determinants, we can identify cells with high transcriptional activity and produce high-yield recombinant proteins.

Benefits of technology

It significantly increased the proportion of positive cell clones expressing the target protein and the POI expression level of host cells, achieving efficient recombinant protein production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121335982A_ABST
    Figure CN121335982A_ABST
Patent Text Reader

Abstract

Disclosed herein are novel selectable markers and uses thereof. Specifically, provided herein is a nucleotide sequence encoding glutamine synthetase (GS) derived from Taokuo bear, ostrich, snake, pigeon, chicken, gerbil, or bat as a selectable marker for identifying genomic loci with high transcriptional activity or host cells with high productivity of a protein of interest, and / or a use for accelerating the identification process. Related screening methods, production methods and expression systems are also included.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference to related applications

[0001] This application claims priority to PCT patent application No. PCT / CN2023 / 073653, filed on January 29, 2023, the entire contents of which are incorporated herein by reference. 1. Reference to the electronically submitted sequence list

[0002] This application incorporates a sequence of XML files for reference, the file being titled “817A004WO01P_SL.XML”, created on January 13, 2023, and having a size of 42,082 bytes. Technical Field

[0003] This invention relates to molecular biology and cell biology. The present invention includes selectable markers for recombinant protein production and related methods of use. Background Technology

[0004] Selectable markers are commonly used in recombinant protein production. When producing clones expressing recombinant proteins, host cells are typically transfected with a DNA vector encoding both the target protein and the selectable marker. Selectable markers allow for the selection of cell clones with expression vectors, as well as high-yield clones. Glutamine synthase (GS) has been widely used as a selectable marker in recombinant protein production in eukaryotic cells, such as Chinese hamster ovary (CHO) cells. Currently, almost all selectable markers used in GS systems are derived from Chinese hamsters (Cricetulus griseus), which have relatively low selection efficiency and yield. More effective and efficient production methods are essential to support the development of innovative biopharmaceuticals and biosimilars, requiring high-yielding cell lines with the desired quality properties. Therefore, the need for expression systems with improved effectiveness and efficiency in cell line production, particularly in the selection step of high-yield clonal cell lines, remains unmet and urgent. The vectors, cells, and expression systems presented in this article address this need and offer relevant advantages. Summary of the Invention

[0005] This article provides the use of nucleotide sequences encoding glutamine synthase (GS) as selectable markers, wherein the GS is koala GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS.

[0006] In some embodiments of the uses provided herein, GS is derived from the koala GS of the Phascolarctidae family. In some embodiments, GS has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 1. In some embodiments, GS exhibits reduced activity compared to wild-type koala GS. In some embodiments, GS has the amino acid sequence shown in SEQ ID NO: 1.

[0007] In some embodiments of the uses provided herein, GS is derived from ostrich GS of the family Struthionidae. In some embodiments, GS has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 2. In some embodiments, GS has reduced activity compared to wild-type ostrich GS. In some embodiments, GS has the amino acid sequence shown in SEQ ID NO: 2.

[0008] In some embodiments of the uses provided herein, GS is derived from snakes of the family Elapidae. In some embodiments, GS has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3. In some embodiments, GS exhibits reduced activity compared to wild-type snake GS. In some embodiments, GS has the amino acid sequence shown in SEQ ID NO: 3.

[0009] In some embodiments of the uses provided herein, GS is derived from pigeon GS of the Columbidae family. In some embodiments, GS has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 5. In some embodiments, GS has reduced activity compared to wild-type pigeon GS. In some embodiments, GS has the amino acid sequence shown in SEQ ID NO: 5.

[0010] In some embodiments of the uses provided herein, GS is derived from wild-type red junglefowl GS. In some embodiments, GS has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 6. In some embodiments, GS has reduced activity compared to wild-type red junglefowl GS. In some embodiments, GS has the amino acid sequence shown in SEQ ID NO: 6.

[0011] In some embodiments of the uses provided herein, GS is derived from the gerbil (Muridae). In some embodiments, GS has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 7. In some embodiments, GS activity is reduced compared to wild-type gerbil GS. In some embodiments, GS has the amino acid sequence shown in SEQ ID NO: 7.

[0012] In some embodiments of the uses provided herein, GS is derived from the nocturnal bat GS of the family Vespertilionidae. In some embodiments, GS has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 8. In some embodiments, GS activity is reduced compared to wild-type nocturnal bat GS. In some embodiments, GS has the amino acid sequence shown in SEQ ID NO: 8.

[0013] In some embodiments, this document provides for the use of nucleotide sequences encoding GS as optional markers, wherein the GS comprises a catalytic domain from koala GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS. In some embodiments, the GS comprises a catalytic domain from a koala GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 1. In some embodiments, the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 1. In some embodiments, the GS comprises a catalytic domain from an ostrich GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 2. In some embodiments, the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 2. In some embodiments, the GS comprises a catalytic domain from a snake GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 3. In some embodiments, the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 3. In some embodiments, the GS comprises a catalytic domain from a pigeon GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 5. In some embodiments, the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 163-412 shown in SEQ ID NO: 5. In some embodiments, the GS comprises a catalytic domain from a wild red junglefowl GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 150-399 shown in SEQ ID NO: 6. In some embodiments, the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 150-399 shown in SEQ ID NO: 6.In some embodiments, the GS comprises a catalytic domain from a gerbil GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 7. In some embodiments, the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 7. In some embodiments, the GS comprises a catalytic domain from a night bat GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 8. In some embodiments, the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 8.

[0014] In some embodiments of the uses provided herein, the nucleotide sequence encoding the GS is operatively linked to an mRNA destabilizing element. In some embodiments, the GS comprises a degradation determinant. In some embodiments, the degradation determinant has an amino acid sequence selected from SEQ ID NO: 23-25.

[0015] In some embodiments, the uses provided herein are for identifying genomic loci with high transcriptional activity. In some embodiments, the uses provided herein are for identifying host cells capable of producing a target protein (POI). In some embodiments, the uses provided herein are for the recombinant production of POIs. In some embodiments, the uses provided herein are for the production of recombinant proteins in mammalian cells. In some embodiments, the mammalian cells are Chinese hamster ovary (CHO) cells. In some embodiments, the POI is selected from antibodies, enzymes, soluble proteins, secretory proteins, membrane proteins, and fusion proteins.

[0016] This article also provides deoxyribonucleic acid (DNA) vectors suitable for recombinant protein production or genome integration, which contain nucleotide sequences encoding GS (GS coding sequences), wherein the GS is koala GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, gerbil GS or night bat GS.

[0017] In some embodiments of the vector provided herein, the encoded GS is derived from the koala GS of the Phascolarctidae family. In some embodiments, the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 1. In some embodiments, the GS activity is reduced compared to wild-type koala GS. In some embodiments, the GS has the amino acid sequence shown in SEQ ID NO: 1. In some embodiments, the GS encoding sequence is at least 80% identical to SEQ ID NO: 10.

[0018] In some embodiments of the vector provided herein, the encoded GS is derived from ostrich GS of the Struthionidae family. In some embodiments, the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 2. In some embodiments, the GS activity is reduced compared to wild-type ostrich GS. In some embodiments, the GS has the amino acid sequence shown in SEQ ID NO: 2. In some embodiments, the GS encoding sequence has at least 80% identity with SEQ ID NO: 11.

[0019] In some embodiments of the vector provided herein, the encoded GS is derived from snake GS of the Elapidae family. In some embodiments, the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 3. In some embodiments, the GS activity is reduced compared to wild-type snake GS. In some embodiments, the GS has the amino acid sequence shown in SEQ ID NO: 3. In some embodiments, the GS encoding sequence has at least 80% identity with SEQ ID NO: 12.

[0020] In some embodiments of the vector provided herein, the encoded GS is derived from pigeon GS of the Columbidae family. In some embodiments, the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 5. In some embodiments, the GS activity is reduced compared to wild-type pigeon GS. In some embodiments, the GS has the amino acid sequence shown in SEQ ID NO: 5. In some embodiments, the GS encoding sequence is at least 80% identical to SEQ ID NO: 14.

[0021] In some embodiments of the vector provided herein, the encoded GS is derived from the red junglefowl GS of the Phasianidae family. In some embodiments, the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 6. In some embodiments, the GS activity is reduced compared to wild-type red junglefowl GS. In some embodiments, the GS has the amino acid sequence shown in SEQ ID NO: 6. In some embodiments, the GS encoding sequence is at least 80% identical to SEQ ID NO: 15.

[0022] In some embodiments of the vector provided herein, the encoded GS is derived from the gerbil GS of the Muridae family. In some embodiments, the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 7. In some embodiments, the GS activity is reduced compared to wild-type gerbil GS. In some embodiments, the GS has the amino acid sequence shown in SEQ ID NO: 7. In some embodiments, the GS encoding sequence is at least 80% identical to SEQ ID NO: 16.

[0023] In some embodiments of the vector provided herein, the encoded GS is derived from the nocturnal bat GS of the Vespertilionidae family. In some embodiments, the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 8. In some embodiments, the GS activity is reduced compared to the wild-type nocturnal bat GS. In some embodiments, the GS has the amino acid sequence shown in SEQ ID NO: 8. In some embodiments, the GS encoding sequence has at least 80% identity with SEQ ID NO: 17.

[0024] This document also provides DNA vectors suitable for recombinant protein production or genome integration, comprising a nucleotide sequence encoding a GS (GS-coding sequence), wherein the GS comprises a catalytic domain derived from koala GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS. In some embodiments, the GS comprises a catalytic domain derived from a koala GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 1. In some embodiments, the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 1. In some embodiments, the GS comprises a catalytic domain derived from an ostrich GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 2. In some embodiments, the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 2. In some embodiments, the GS comprises a catalytic domain from a snake GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 3. In some embodiments, the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 3. In some embodiments, the GS comprises a catalytic domain from a pigeon GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 5. In some embodiments, the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 163-412 shown in SEQ ID NO: 5. In some embodiments, the GS comprises a catalytic domain from a wild red junglefowl GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 150-399 shown in SEQ ID NO: 6. In some embodiments, the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 150-399 shown in SEQ ID NO: 6.In some embodiments, the GS comprises a catalytic domain from a gerbil GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 7. In some embodiments, the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 7. In some embodiments, the GS comprises a catalytic domain from a night bat GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 8. In some embodiments, the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 8.

[0025] In some embodiments of the vectors provided herein, the GS coding sequence is operatively linked to an mRNA destabilizing element. In some embodiments of the vectors provided herein, the encoded GS comprises a degradation determinant. In some embodiments, the degradation determinant has an amino acid sequence selected from SEQ ID NO: 23-25.

[0026] In some embodiments of the vectors provided herein, the GS coding sequence is operatively linked to the simian cavitation virus 40 (SV40) promoter. In some embodiments of the vectors provided herein, the GS coding sequence is operatively linked to a poly(A) tail.

[0027] In some embodiments, the vectors provided herein are suitable for recombinant protein production and also include expression cassettes. In some embodiments, the vectors provided herein include two or more expression cassettes. In some embodiments, the expression cassettes include a nucleotide sequence encoding a POI (POI coding sequence). In some embodiments, the POI is selected from antibodies, enzymes, soluble proteins, secreted proteins, membrane proteins, or fusion proteins. In some embodiments, the POI is an antibody selected from: IgG1 antibody, IgG2 antibody, IgG3 antibody, IgG4 antibody, IgA antibody, IgM antibody, Fab, Fab', F(ab')2, Fv, scFv, (scFv)2, single-domain antibody (sdAb), single-chain antibody (scAb), and heavy-chain antibody (HCAb). In some embodiments, the POI is an antibody selected from: monoclonal antibody, bispecific antibody, multispecific antibody, bivalent antibody, and multivalent antibody. In some embodiments, the POI consists of one or more copies of the same polypeptide. In some embodiments, the POI comprises two different polypeptides. In some implementations, the POI is an antibody comprising a light chain and a heavy chain, each encoded by a separate nucleotide sequence on the vector.

[0028] In some implementations, the vectors provided herein are suitable for genome integration.

[0029] In some embodiments, this document provides the use of the vectors provided herein for identifying host cells capable of producing POIs.

[0030] In some embodiments, this document provides the use of the vectors provided herein for identifying genomic loci with high transcriptional activity.

[0031] In some embodiments, this document provides a method for identifying host cells capable of producing POIs, comprising introducing a vector described herein into a population of host cells, culturing the population of host cells in a glutamine-free medium, wherein host cells capable of growing in the medium are identified as host cells capable of producing POIs.

[0032] In some embodiments, this document provides a method for identifying genomic loci with high transcriptional activity, comprising introducing a vector described herein into a population of host cells, culturing the host cell population in a glutamine-free medium, wherein host cells capable of growing in the medium are identified as host cells having a GS coding sequence inserted at a genomic locus with high transcriptional activity. In some embodiments, the method provided herein further includes sequencing the genome of the identified host cells to locate the genomic locus with high transcriptional activity.

[0033] In some embodiments of the methods provided herein, the host cell population is cultured in the presence of a GS inhibitor. In some embodiments, the GS inhibitor is methionine sulfoximine (MSX).

[0034] In some embodiments, this document provides host cells comprising the vectors disclosed herein. In some embodiments, the host cells have wild-type endogenous GS. In some embodiments, the endogenous GS of the host cells has reduced activity or has been knocked out. In some embodiments, The host cell is a mammalian cell. In some embodiments, the host cell is a CHO cell.

[0035] In some embodiments, this document provides the use of the host cells disclosed herein for the in vitro production of POIs. In some embodiments, this document provides a method for the in vitro production of POIs, which includes culturing the host cells disclosed herein for a sufficient time under POI-producing conditions.

[0036] In some embodiments, this document provides a method for in vitro production of POIs, which includes replacing the GS coding sequence with the POI coding sequence in host cells identified in the methods described herein, and culturing the host cells for a sufficient time under POI-producing conditions. In some embodiments, the method further includes separating the POI from other components in culture. In some embodiments, the separation includes extraction, continuous liquid-liquid extraction, pervaporation, membrane filtration, membrane separation, reverse osmosis, electrodialysis, distillation, crystallization, centrifugation, extractive filtration, ion exchange chromatography, adsorption chromatography, or ultrafiltration.

[0037] In some embodiments, this document provides an expression system for the in vitro production of POIs, comprising the DNA vector disclosed herein and a host cell. In some embodiments, the host cell is a CHO cell. In some embodiments, the expression system provided herein further includes a glutamine-free culture medium. In some embodiments, the expression system provided herein further includes a GS inhibitor. In some embodiments, the expression system provided herein further includes a method for introducing the vector into the host cell. In some embodiments, the expression system provided herein is also included in a kit. Attached Figure Description

[0038] Figure 1A scatter plot of antibody expression from the pools is provided. Cells were electroporated with a denosumab expression vector containing a GS-selectable biomarker from a specified species and then divided into 300 pools in 96-well plates. All pools were cultured for 20 days in the presence of 50 μM MSX. Antibody levels in the supernatant from each pool were measured using the Octet label-free system. Expression from all positive pools (left) or the top 30 pools (right) is plotted. Horizontal lines represent medians.

[0039] Figure 2 The titer distribution of antibody expression from positive pools is provided. Cells were electroporated with a denosumab expression vector containing a GS-selectable biomarker from a specified species, and then divided into 300 pools in 96-well plates. All pools were cultured for 20 days in the presence of 50 μM MSX. Antibody levels in the supernatant from each pool were measured by Octet.

[0040] Figures 3A-3C Destabilized GS selectable markers further improved selection efficiency. Flow cytometry analysis was performed on CHO cells cultured in the absence or presence of 50 μM MSX. Cells were electroporated using a denosumab expression vector containing a GS selectable marker from the specified species. Figure 3A The results are shown using Chinese Hamster GS, Koala GS, Ostrich GS, or Snake GS. Figure 3B The results are shown using Pigeon GS, Red Junglefowl GS, Gerbil GS, or Night Bat GS. Figure 3C The results of flow cytometry analysis of CHO cells cultured in the presence of 50 μM MSX are shown.

[0041] Figure 4 The amino acid sequences of GS from Chinese hamster (Cricetulus griseus) , koala (Phascolarctos cinerecus) , night bat (Pipistrellus kuhlii) , gerbil (Meriones unguiculatus) , ostrich (Struthio camelus australis) , red junglefowl (Gallus gallus) , pigeon (Columbia livia) , and snake (Pseudonajatextilis) are provided. Detailed Implementation

[0042] This paper provides an expression system using GS derived from koalas, ostriches, snakes, pigeons, junglefowl, gerbils, or night bats as selectable markers in (e.g.) recombinant protein production, and a method for screening cell clones with high yields. Unbound by theory, the invention presented herein is based on the unexpected discovery that GS from species phylogenetically distant from CHO host cells, specifically GS from koalas, ostriches, snakes, pigeons, junglefowl, gerbils, or night bats, can be used as highly efficient selectable markers, and their use not only leads to a significant increase in the proportion of positive cell clones expressing the target protein (“POI”), but also a significant increase in the POI expression level in the host cells (e.g., CHO cells). As disclosed herein, when using GS derived from koalas, ostriches, snakes, pigeons, junglefowl, gerbils, or night bats as selectable markers, efficient identification of cell clones with high recombinant POI yields is achieved. Therefore, in some embodiments, this document provides the use of nucleotide sequences encoding GS (GS-coding sequences) as optional markers, wherein the GS is koala GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS; this document also provides expression vectors containing GS-coding sequences, wherein the GS is koala GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS, and methods of using them; this document also provides expression systems and kits containing these vectors.

[0043] Before further describing the present invention, it should be understood that the present invention is not limited to the specific embodiments described herein, and that the terminology used herein is for the purpose of describing specific embodiments and is not intended to be limiting.

[0044] Unless otherwise defined herein, the scientific and technical terms used in this disclosure shall have the meanings commonly understood by those skilled in the art. Furthermore, unless the context otherwise requires, singular terms shall include plural terms and plural terms shall include singular terms. Generally, the terminology used in conjunction with the cell and tissue culture, molecular biology, immunology, microbiology, genetics, and protein and nucleic acid chemistry and hybridization described herein, as well as the techniques used in these fields, are those well known and commonly used in the art.

[0045] The term “one” refers to one or more of the same entity; for example, “one carrier” should be understood to mean one or more carriers.

[0046] When used herein, the term “and / or” will be considered as a specific disclosure of each of the two specified features or components, with or without the other. Thus, as used herein in phrases such as “A and / or B”, the term “and / or” is intended to include “A and B”, “A or B”, “A alone”, and “B alone”. Similarly, as used in phrases such as “A, B and / or C”, the term “and / or” is intended to cover each of the following: A, B and C; A, B or C; A or C; A or B; B or C; A and C; A and B; B and C; A alone; B alone; and C alone.

[0047] As used herein in the context of two or more polynucleotides or polypeptides, the terms "identical," "percentage of identity," and their syntactic equivalents refer to two or more sequences or subsequences that are identical or have a specified percentage of identical nucleotide or amino acid residues when compared and aligned for maximum correspondence (with gaps introduced where necessary) and without considering any conserved amino acid substitutions as part of sequence identity. Percentage of identity can be measured using sequence comparison software or algorithms or by visual inspection. A variety of algorithms and software that can be used to obtain amino acid or nucleotide sequence alignments are well known in the art. These include (but are not limited to) the BLAST, ALIGN, Megalign, BestFit, GCG Wisconsin software packages, and their variations. In some embodiments, the two polynucleotides or polypeptides provided herein are substantially identical, meaning that when compared and aligned for maximum correspondence, they have at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, and in some embodiments, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% nucleotide or amino acid residue identity, as measured by sequence comparison algorithms or by visual inspection. In some embodiments, there is identity within amino acid sequence regions of at least about 10 residues, at least about 20 residues, at least about 40-60 residues, at least about 60-80 residues, or any integer value between them. In some embodiments, there is identity within regions of more than 60-80 residues, such as at least about 80-100 residues, and in some embodiments, the sequences are substantially identical across the full length of the compared sequences (e.g., the coding region of a target protein or antibody). In some embodiments, identity exists within nucleotide sequence regions of at least about 10 bases, at least about 20 bases, at least about 40-60 bases, at least about 60-80 bases, or any integer value between them. In some embodiments, identity exists within regions greater than 60-80 bases, such as at least about 80-100 bases or more, and in some embodiments, the sequences are substantially identical over the full length of the compared sequences (e.g., the nucleotide sequence encoding a POI). Unless otherwise stated, the percentage of identity as used herein is calculated using global alignment (i.e., comparing two sequences over their entire length).

[0048] As used herein, the term "about" refers to a conventional range of error for various values ​​readily known to those skilled in the art. References to "about" values ​​or parameters herein include and describe embodiments involving that value or parameter itself. For example, a description of "about X" includes a description of "X". In some embodiments, "about" represents a value up to ±10% of the listed values, such as ±1%, ±2%, ±3%, ±4%, ±5%, ±6%, ±7%, ±8%, ±9%, or ±10%. In many embodiments, the term "about" covers a variation of ±5%, ±2%, ±1%, or ±0.5% of the numerical value of a number. In some embodiments, the term "about" covers a variation of ±5% of the numerical value of a number. In some embodiments, the term "about" covers a variation of ±2% of the numerical value of a number. In some embodiments, the term "about" covers a variation of ±1% of the numerical value of a number.

[0049] Scope: Throughout this disclosure, various aspects of the invention may be presented in a scope format. It should be understood that the scope format is for convenience and brevity only and should not be considered a rigid limitation on the scope of the invention. Therefore, a description of the scope should be considered to have all possible sub-scopes specifically disclosed, as well as the individual values ​​within those scopes. For example, a description of a scope such as 1 to 6 should be considered to have specifically disclosed sub-scopes such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., and the individual values ​​within those scopes, such as 1, 2, 2.7, 3, 4, 5, 5.3, and 6. This is applicable regardless of the width of the scope.

[0050] This document describes exemplary genes and peptides with reference to GenBank numbers, GI numbers, Uniprot numbers, and / or SEQ ID NO. It should be understood that those skilled in the art can readily identify homologous sequences from reference sequence sources, including (but not limited to) Uniprot (uniprot.org / ), GenBank (ncbi.nlm.nih.gov / genbank / ), and / or EMBL (embl.org / ). 6.1 Selectable markers

[0051] To produce recombinant proteins on an industrial scale, identifying cell clones capable of producing large quantities of recombinant proteins is crucial. The optional markers provided herein allow for the efficient identification of cell clones with high POI yields. High cellular yields are achieved by integrating expression vectors into highly transcriptionally active sites or by having multiple copies of the expression vector present in the cell. Accordingly, this document provides the use of nucleotide sequences encoding glutamine synthase (GS) as optional markers, wherein the GS is koala GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS. In some embodiments, this document provides the use for recombinant protein production. In some embodiments, this document provides the use for identifying cell clones capable of producing POIs. In some embodiments, this document provides the use for identifying genomic loci that integrate transgenes for efficient expression.

[0052] As used herein and as commonly understood in the art, a “selectable marker” is a gene that confers a characteristic to its vector that allows for artificial selection. This characteristic is typically a positive characteristic, such as resistance to antibiotics or key enzyme activity. Selectable markers are often an integral part of a vector (such as an expression vector) and are commonly used in molecular biology and genetic engineering to indicate the success of a procedure for introducing a vector into cells. Once a vector containing a selectable marker is introduced into a population of host cells, the cells can be cultured in a medium in which the expression of the selectable marker is required for cell survival and growth. In this way, the selectable marker can select cells that have successfully accepted and expressed the vector. Selection conditions can be tuned for different levels of stringency. Generally, the more stringent the culture conditions, the higher the expression of the selectable marker required for the growth of host cells under those conditions.

[0053] Glutamine synthase, or GS, is an enzyme classified under the Enzyme Committee (EC) number 6.3.1.2. GS catalyzes the ATP-dependent conversion of glutamate and ammonia to glutamine and plays a crucial role in nitrogen metabolism. The biochemical reaction catalyzed by GS can be represented as: ATP + L-glutamate + NH3 <=> ADP + phosphate + L-glutamine. The enzyme activity catalyzing the above reaction is referred to herein as "GS activity".

[0054] Wild-type GS typically has two main domains: a β-grasping domain (e.g., amino acid residues from position 30 to position 104 in wild-type hamster GS, SEQ ID NO: 9) and a catalytic domain (amino acid residues from position 110 to position 359 in wild-type hamster GS, SEQ ID NO: 9).

[0055] Glutamine cassettes (GS) are commonly used selectable markers in mammalian cell lines. In some cell lines, endogenous GS can be inactivated or inhibited. Consequently, these cells cannot synthesize glutamine themselves and can only grow when glutamine is added to the culture medium, or only in glutamine-free medium when they have been introduced with an expression vector containing the GS gene. Positive selection can be maintained by using glutamine-free medium. Furthermore, different concentrations of GS activity inhibitors, including methionine sulfoximine (MSX) and its derivatives, phosphorus-containing analogs of glutamate, and bisphosphates, can be used to produce different levels of selection strictness. Increased GS inhibitors select for cells with higher GS expression because GS activity is inhibited, and only cells with higher GS levels can grow under this treatment. Therefore, cells with amplified copy numbers of expression cassettes in the chromosome, or cells with expression cassettes inserted at highly transcriptionally active sites, can be identified.

[0056] In some embodiments, this document provides for the use of nucleotide sequences encoding GS (GS-encoding sequences) as optional markers, wherein the GS is a koala GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS. As used herein and understood in the art, the term “encoding” or its syntactic equivalent refers to the inherent property of a particular nucleotide sequence to serve as a template for the synthesis of other polymers and macromolecules having defined nucleotide sequences (i.e., rRNA, tRNA, and mRNA) or defined amino acid sequences and the biological properties produced therefrom. Thus, if transcription and translation of the mRNA corresponding to the gene produces a protein, then the gene encodes the protein. Unless otherwise stated, a “nucleotide sequence” “encoding” an amino acid sequence can be any nucleotide sequence that is a degenerate form of each other and encodes the same amino acid sequence. Nucleotide sequences encoding proteins and RNA may include introns. Table 1: Exemplary GS amino acid and nucleotide sequences.

[0057] In some embodiments, this document provides the use of the nucleotide sequence encoding the koala GS (koala GS coding sequence) as an optional marker. The koala GS is a GS derived from the koala. In some embodiments, the koala GS is derived from the family Phascolarctidae. In some embodiments, the koala GS is derived from the genus Phascolarctos. In some embodiments, the koala GS is derived from the koala (Phascolarctoscinerecus). Exemplary amino acid sequences of the koala GS can be found in the table above and in public databases, such as Uniprot#A0A6P5K4T4.

[0058] This document provides the use of the koala GS encoding sequence as an alternative marker. In some embodiments, the nucleotide sequence encodes a koala GS having an amino acid sequence comprising or consisting of SEQ ID NO: 1. In some embodiments, the koala GS is a functional variant of a GS having the amino acid sequence shown in SEQ ID NO: 1. Functional variants of a GS having the amino acid sequence shown in SEQ ID NO: 1 maintain the basic structure and GS activity of the reference GS. The functional variant may have, for example, amino acid substitutions, deletions, and / or additions of about 1 to about 25, about 1 to about 20, about 1 to about 15, about 1 to about 10, or about 1 to about 5 amino acid substitutions, deletions, and / or additions compared to SEQ ID NO: 1. In some embodiments, the amino acid sequence of the functional variant differs from SEQ ID NO: 1 by only up to 22, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 amino acid variation, including the presence of substitutions, insertions, and deletions. Variations in the amino acid sequence can be amino acid substitutions. Variations in the amino acid sequence can be conserved amino acid substitutions. The variants may be naturally present in the koala species (e.g., allelic variants or splice variants). Alternatively, the variants may be obtained through genetic engineering. In some embodiments, the functional variants exhibit reduced GS activity compared to wild-type koala GS. In some embodiments, the functional variants with reduced GS activity contain one or more mutations in the glutamate-binding domain, ATP-binding domain, amino-binding domain, or any combination thereof. In some embodiments, the functional variants with reduced GS activity contain one or more mutations in the β-grasping domain, catalytic domain, or any combination thereof.

[0059] This document provides the use of the koala GS coding sequence as an optional marker. In some embodiments, the nucleotide sequence encodes a koala GS having an amino acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 1. In some embodiments, GS has an amino acid sequence having at least 90% identity with SEQ ID NO: 1. In some embodiments, GS has an amino acid sequence having at least 92% identity with SEQ ID NO: 1. In some embodiments, GS has an amino acid sequence having at least 95% identity with SEQ ID NO: 1. In some embodiments, GS has an amino acid sequence having at least 98% identity with SEQ ID NO: 1. In some embodiments, GS has an amino acid sequence having at least 99% identity with SEQ ID NO: 1. In some embodiments, GS has the amino acid sequence shown in SEQ ID NO: 1. In some embodiments, the koala GS is derived from the Phascolarctidae family and has an amino acid sequence that is at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to that in SEQ ID NO: 1. In some embodiments, the koala GS is derived from the genus *Phascolarctos* and has an amino acid sequence with at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 1. In some embodiments, the koala GS is derived from the family *Phascolarctidae* and has an amino acid sequence with at least 95% identity with SEQ ID NO: 1. In some embodiments, the GS provided herein as an alternative marker has reduced activity compared to wild-type koala GS. In some embodiments, the GS with reduced activity contains one or more mutations in the glutamate-binding region, ATP-binding region, amino-binding region, or any combination thereof. In some embodiments, the functional variants with reduced GS activity contain one or more mutations in the β-grasping domain, catalytic domain, or any combination thereof.

[0060] This document provides the use of koala GS encoding sequences as alternative biomarkers. In some embodiments, the nucleotide sequence encodes a functional fragment of koala GS having the amino acid sequence shown in SEQ ID NO: 1. The functional fragment of koala GS having the amino acid sequence shown in SEQ ID NO: 1 maintains the GS activity of a reference GS. In some embodiments, the functional fragment comprises at least 100 consecutive amino acids of SEQ ID NO: 1. In some embodiments, the functional fragment comprises at least 150 consecutive amino acids of SEQ ID NO: 1. In some embodiments, the functional fragment comprises at least 200 consecutive amino acids of SEQ ID NO: 1. In some embodiments, the functional fragment comprises at least 250 consecutive amino acids of SEQ ID NO: 1. In some embodiments, the functional fragment comprises at least 300 consecutive amino acids of SEQ ID NO: 1. In some embodiments, the functional fragment provided herein as an alternative biomarker exhibits reduced GS activity compared to wild-type koala GS.

[0061] Since the catalytic domain itself is sufficient to provide enhanced GS marker performance, in some embodiments, GS comprising a catalytic domain from koala GS are provided herein. In some embodiments, the catalytic domain of koala GS (e.g., GS from koala (Phascolarctos cinerecus)) may consist of amino acids 134-351 of the protein. In some embodiments, the catalytic domain of koala GS (e.g., GS from koala (Phascolarctoscinerecus)) may consist of amino acids 110-359 of the protein. In some embodiments, the GS comprises a catalytic domain from koala GS having an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 1. In some embodiments, the catalytic domain has an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 1. In some embodiments, the catalytic domain has amino acids 110-359 shown in SEQ ID NO: 1. The GS may further include an N-terminal region from any other GS. In some embodiments, the GS may also include a beta-grasp domain from any other GS. The GS may include an N-terminal region (e.g., a beta-grasp domain) from, for example, hamster GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS. In some embodiments, the GS provided herein includes a catalytic domain from a koala GS and a beta-grasp domain from a hamster GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS.

[0062] This document provides the use of the koala GS coding sequence as an alternative marker. In some embodiments, the koala GS exhibits reduced stability at the mRNA or protein level. In some embodiments, the koala GS coding sequence provided herein is operatively linked to an mRNA destabilizing element. In some embodiments, the koala GS provided herein contains a degradation determinant. The degradation determinant may be any degradation determinant disclosed herein or otherwise known in the art. In some embodiments, the koala GS has an N-terminal degradation determinant. In some embodiments, the koala GS has a C-terminal degradation determinant. In some embodiments, the degradation determinant is a PEST sequence (e.g., SEQ ID NO: 23). In some embodiments, the degradation determinant is an ODD sequence (e.g., SEQ ID NO: 24). In some embodiments, the degradation determinant is an IκBα sequence (e.g., SEQ ID NO: 25). In some embodiments, the degradation determinant has an amino acid sequence selected from SEQ ID NO: 23-25.

[0063] In some embodiments, this document provides the use of nucleotide sequences encoding ostrich GS as optional markers. Ostrich GS is a GS derived from ostriches. In some embodiments, ostrich GS is derived from the family Struthionidae. In some embodiments, ostrich GS is derived from the genus Struthio. In some embodiments, ostrich GS is derived from the South African ostrich (Struthio camelus australis). Exemplary amino acid sequences of ostrich GS can be found in the table above and in public databases, such as Uniprot#A0A093JMX8.

[0064] This document provides the use of ostrich GS coding sequences as alternative markers. In some embodiments, the nucleotide sequence encodes an ostrich GS having an amino acid sequence comprising or consisting of SEQ ID NO:2. In some embodiments, the ostrich GS is a functional variant of a GS having the amino acid sequence shown in SEQ ID NO:2. Functional variants of GS having the amino acid sequence shown in SEQ ID NO:2 maintain the basic structure and GS activity of the reference GS. The functional variant may have, for example, amino acid substitutions, deletions, and / or additions of about 1 to about 25, about 1 to about 20, about 1 to about 15, about 1 to about 10, or about 1 to about 5 amino acid substitutions compared to SEQ ID NO:2. In some embodiments, the amino acid sequence of the functional variant differs from SEQ ID NO:2 by only up to 22, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 amino acid variation, including the presence of substitutions, insertions, and deletions. Variations in the amino acid sequence can be amino acid substitutions. Variations in the amino acid sequence can be conserved amino acid substitutions. The variant may be naturally present in ostrich species (e.g., allelic variants or splice variants). Alternatively, the variant can be obtained through genetic engineering. In some embodiments, the functional variant exhibits reduced GS activity compared to wild-type ostrich GS. In some embodiments, the functional variant with reduced GS activity contains one or more mutations in the glutamate-binding domain, ATP-binding domain, ammonia-binding domain, or any combination thereof. In some embodiments, the functional variant with reduced GS activity contains one or more mutations in the β-grasping domain, catalytic domain, or a combination thereof.

[0065] This document provides the use of the ostrich GS encoding sequence as an optional marker. In some embodiments, the nucleotide sequence encodes an ostrich GS having an amino acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 2. In some embodiments, the GS has an amino acid sequence having at least 90% identity with SEQ ID NO: 2. In some embodiments, the GS has an amino acid sequence having at least 92% identity with SEQ ID NO: 2. In some embodiments, the GS has an amino acid sequence having at least 95% identity with SEQ ID NO: 2. In some embodiments, the GS has an amino acid sequence having at least 98% identity with SEQ ID NO: 2. In some embodiments, the GS has an amino acid sequence having at least 99% identity with SEQ ID NO: 2. In some embodiments, GS has the amino acid sequence shown in SEQ ID NO: 2. In some embodiments, the ostrich GS is derived from the family Struthionidae and has an amino acid sequence that is at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to that of SEQ ID NO: 2. In some embodiments, the ostrich GS is derived from the genus Struthio and has an amino acid sequence that is at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to that of SEQ ID NO: 2. In some embodiments, the ostrich GS is derived from the family Struthionidae and has an amino acid sequence with at least 95% identity to SEQ ID NO: 2. In some embodiments, the GS provided herein as an alternative marker has reduced activity compared to wild-type ostrich GS. In some embodiments, the GS with reduced activity contains one or more mutations in the glutamate-binding region, ATP-binding region, amino-binding region, or any combination thereof. In some embodiments, functional variants with reduced GS activity contain one or more mutations in the β-grasping domain, catalytic domain, or any combination thereof.

[0066] This document provides the use of ostrich GS encoding sequences as selectable markers. In some embodiments, the nucleotide sequence encodes a functional fragment of ostrich GS having the amino acid sequence shown in SEQ ID NO: 2. The functional fragment of ostrich GS having the amino acid sequence shown in SEQ ID NO: 2 maintains the GS activity of a reference GS. In some embodiments, the functional fragment comprises at least 100 consecutive amino acids of SEQ ID NO: 2. In some embodiments, the functional fragment comprises at least 150 consecutive amino acids of SEQ ID NO: 2. In some embodiments, the functional fragment comprises at least 200 consecutive amino acids of SEQ ID NO: 2. In some embodiments, the functional fragment comprises at least 250 consecutive amino acids of SEQ ID NO: 2. In some embodiments, the functional fragment comprises at least 300 consecutive amino acids of SEQ ID NO: 2. In some embodiments, the functional fragment provided herein as a selectable marker exhibits reduced GS activity compared to wild-type ostrich GS.

[0067] Since the catalytic domain itself is sufficient to provide enhanced GS marker performance, in some embodiments, GS comprising a catalytic domain derived from ostrich GS are provided herein. In some embodiments, the catalytic domain of ostrich GS (e.g., GS from the South African ostrich (Struthio camelus australis)) may consist of amino acids 134-351 of the protein. In some embodiments, the catalytic domain of ostrich GS (e.g., GS from the South African ostrich (Struthio camelus australis)) may consist of amino acids 110-359 of the protein. In some embodiments, the GS comprises a catalytic domain from an ostrich GS having an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 2. In some embodiments, the catalytic domain has an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 2. In some embodiments, the catalytic domain has amino acids 110-359 shown in SEQ ID NO: 2. The GS may further include an N-terminal region from any other GS. In some embodiments, the GS may also include a β-grasping domain from any other GS. The GS may include an N-terminal region (e.g., a β-grasping domain) from, for example, hamster GS, koala GS, snake GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS. In some embodiments, the GS provided herein includes a catalytic domain from an ostrich GS and a β-grasping domain from a hamster GS, koala GS, snake GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS.

[0068] This document provides the use of ostrich GS coding sequences as selectable markers. In some embodiments, ostrich GS exhibits reduced stability at the mRNA or protein level. In some embodiments, the ostrich GS coding sequences provided herein are operatively linked to mRNA destabilizing elements. In some embodiments, the ostrich GS provided herein contains a degradation determinant. The degradation determinant may be any degradation determinant disclosed herein or otherwise known in the art. In some embodiments, ostrich GS has an N-terminal degradation determinant. In some embodiments, ostrich GS has a C-terminal degradation determinant. In some embodiments, the degradation determinant is a PEST sequence (e.g., SEQ ID NO: 23). In some embodiments, the degradation determinant is an ODD sequence (e.g., SEQ ID NO: 24). In some embodiments, the degradation determinant is an IκBα sequence (e.g., SEQ ID NO: 25). In some embodiments, the degradation determinant has an amino acid sequence selected from SEQ ID NO: 23-25.

[0069] In some embodiments, this document provides the use of nucleotide sequences encoding snake GS as optional markers. Snake GS is a GS derived from snakes. In some embodiments, snake GS is derived from the family Elapidae. In some embodiments, snake GS is derived from the genus *Pseudonaja*. In some embodiments, snake GS is derived from the eastern brown snake *Pseudonaja textilis*. Exemplary amino acid sequences of snake GS can be found in the table above and in public databases, such as Uniprot#A0A670XZ19.

[0070] This document provides the use of snake GS coding sequences as alternative markers. In some embodiments, the nucleotide sequence encodes a snake GS having an amino acid sequence comprising or consisting of SEQ ID NO:3. In some embodiments, the snake GS is a functional variant of a GS having the amino acid sequence shown in SEQ ID NO:3. Functional variants of GS having the amino acid sequence shown in SEQ ID NO:3 maintain the basic structure and GS activity of the reference GS. The functional variant may have, for example, amino acid substitutions, deletions, and / or additions of about 1 to about 25, about 1 to about 20, about 1 to about 15, about 1 to about 10, or about 1 to about 5 amino acid substitutions compared to SEQ ID NO:3. In some embodiments, the amino acid sequence of the functional variant differs from SEQ ID NO:3 by only up to 22, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 amino acid variation, including the presence of substitutions, insertions, and deletions. The variation in the amino acid sequence can be an amino acid substitution. The variation in the amino acid sequence can be a conserved amino acid substitution. The variant may be naturally present in snake species (e.g., allelic variants or splice variants). Alternatively, the variant can be obtained through genetic engineering. In some embodiments, the functional variant exhibits reduced GS activity compared to wild-type snake GS. In some embodiments, the functional variant with reduced GS activity contains one or more mutations in the glutamate-binding domain, ATP-binding domain, ammonia-binding domain, or any combination thereof. In some embodiments, the functional variant with reduced GS activity contains one or more mutations in the β-grasping domain, catalytic domain, or a combination thereof.

[0071] This document provides the use of the snake GS encoding sequence as an optional marker. In some embodiments, the nucleotide sequence encodes a snake GS having an amino acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3. In some embodiments, the GS has an amino acid sequence having at least 90% identity with SEQ ID NO: 3. In some embodiments, the GS has an amino acid sequence having at least 92% identity with SEQ ID NO: 3. In some embodiments, the GS has an amino acid sequence having at least 95% identity with SEQ ID NO: 3. In some embodiments, the GS has an amino acid sequence having at least 98% identity with SEQ ID NO: 3. In some embodiments, the GS has an amino acid sequence having at least 99% identity with SEQ ID NO: 3. In some embodiments, GS has the amino acid sequence shown in SEQ ID NO: 3. In some embodiments, the snake GS is derived from the family Elapidae and has an amino acid sequence that is at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to that of SEQ ID NO: 3. In some embodiments, the snake GS is derived from the genus Pseudonaja and has an amino acid sequence that is at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to that of SEQ ID NO: 3. In some embodiments, the snake GS is derived from the Elapidae family and has an amino acid sequence with at least 95% identity to SEQ ID NO: 3. In some embodiments, the GS provided herein as an alternative marker has reduced activity compared to wild-type snake GS. In some embodiments, the GS with reduced activity contains one or more mutations in the glutamate-binding region, ATP-binding region, amino-binding region, or any combination thereof. In some embodiments, functional variants with reduced GS activity contain one or more mutations in the β-grasping domain, catalytic domain, or a combination thereof.

[0072] Since the catalytic domain itself is sufficient to provide enhanced GS marker performance, in some embodiments, GS comprising a catalytic domain derived from snake GS are provided herein. In some embodiments, the catalytic domain of snake GS (e.g., GS from the eastern brown snake (Pseudonaja textilis)) may consist of amino acids 134-351 of the protein. In some embodiments, the catalytic domain of snake GS (e.g., GS from the eastern brown snake (Pseudonaja textilis)) may consist of amino acids 110-359 of the protein. In some embodiments, the GS comprises a catalytic domain from snake GS having an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3. In some embodiments, the catalytic domain has an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 3. In some embodiments, the catalytic domain has amino acids 110-359 shown in SEQ ID NO: 3. The GS may further include an N-terminal region from any other GS. In some embodiments, the GS may also include a β-grasping domain from any other GS. The GS may include an N-terminal region (e.g., a β-grasping domain) from, for example, hamster GS, koala GS, ostrich GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS. In some embodiments, the GS provided herein includes a catalytic domain from a snake GS and a β-grasping domain from a hamster GS, koala GS, ostrich GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS.

[0073] This document provides the use of snake GS coding sequences as selectable markers. In some embodiments, the nucleotide sequence encodes a functional fragment of snake GS having the amino acid sequence shown in SEQ ID NO: 3. The functional fragment of snake GS having the amino acid sequence shown in SEQ ID NO: 3 maintains the GS activity of a reference GS. In some embodiments, the functional fragment comprises at least 100 consecutive amino acids of SEQ ID NO: 3. In some embodiments, the functional fragment comprises at least 150 consecutive amino acids of SEQ ID NO: 3. In some embodiments, the functional fragment comprises at least 200 consecutive amino acids of SEQ ID NO: 3. In some embodiments, the functional fragment comprises at least 250 consecutive amino acids of SEQ ID NO: 3. In some embodiments, the functional fragment comprises at least 300 consecutive amino acids of SEQ ID NO: 3. In some embodiments, the functional fragment provided herein as a selectable marker exhibits reduced GS activity compared to wild-type snake GS.

[0074] This document provides the use of snake GS coding sequences as selectable markers. In some embodiments, snake GS exhibits reduced stability at the mRNA or protein level. In some embodiments, the snake GS coding sequences provided herein are operatively linked to mRNA destabilizing elements. In some embodiments, the snake GS provided herein contains a degradation determinant. The degradation determinant may be any degradation determinant disclosed herein or otherwise known in the art. In some embodiments, snake GS has an N-terminal degradation determinant. In some embodiments, snake GS has a C-terminal degradation determinant. In some embodiments, the degradation determinant is a PEST sequence (e.g., SEQ ID NO: 23). In some embodiments, the degradation determinant is an ODD sequence (e.g., SEQ ID NO: 24). In some embodiments, the degradation determinant is an IκBα sequence (e.g., SEQ ID NO: 25). In some embodiments, the degradation determinant has an amino acid sequence selected from SEQ ID NO: 23-25.

[0075] In some embodiments, this document provides the use of nucleotide sequences encoding pigeon GS as optional markers. Pigeon GS is a GS derived from pigeons. In some embodiments, pigeon GS originates from the family Columbidae. In some embodiments, pigeon GS originates from the genus *Columbia*. In some embodiments, pigeon GS originates from the rock pigeon (*Columbia livia*). Exemplary amino acid sequences of pigeon GS can be found in the table above and in public databases, such as NCBI#XP_005501161.

[0076] This document provides the use of pigeon GS encoding sequences as alternative markers. In some embodiments, the nucleotide sequence encodes a pigeon GS having an amino acid sequence comprising or consisting of SEQ ID NO: 5. In some embodiments, the pigeon GS is a functional variant of a GS having the amino acid sequence shown in SEQ ID NO: 5. Functional variants of GS having the amino acid sequence shown in SEQ ID NO: 5 maintain the basic structure and GS activity of the reference GS. The functional variant may have, for example, amino acid substitutions, deletions, and / or additions of about 1 to about 25, about 1 to about 20, about 1 to about 15, about 1 to about 10, or about 1 to about 5 amino acid substitutions compared to SEQ ID NO: 5. In some embodiments, the amino acid sequence of the functional variant differs from SEQ ID NO: 5 by only up to 22, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 amino acid variation, including the presence of substitutions, insertions, and deletions. The variation in amino acid sequence can be an amino acid substitution. The variation in amino acid sequence can be a conserved amino acid substitution. The variant may be naturally present in pigeon species (e.g., allelic variants or splice variants). Alternatively, the variant can be obtained through genetic engineering. In some embodiments, the functional variant exhibits reduced GS activity compared to wild-type pigeon GS. In some embodiments, the functional variant with reduced GS activity contains one or more mutations in the glutamate-binding domain, ATP-binding domain, ammonia-binding domain, or any combination thereof. In some embodiments, the functional variant with reduced GS activity contains one or more mutations in the β-grasping domain, catalytic domain, or a combination thereof.

[0077] This document provides the use of the pigeon GS encoding sequence as an optional marker. In some embodiments, the nucleotide sequence encodes a pigeon GS having an amino acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 5. In some embodiments, the GS has an amino acid sequence having at least 90% identity with SEQ ID NO: 5. In some embodiments, the GS has an amino acid sequence having at least 92% identity with SEQ ID NO: 5. In some embodiments, the GS has an amino acid sequence having at least 95% identity with SEQ ID NO: 5. In some embodiments, the GS has an amino acid sequence having at least 98% identity with SEQ ID NO: 5. In some embodiments, the GS has an amino acid sequence having at least 99% identity with SEQ ID NO: 5. In some embodiments, GS has the amino acid sequence shown in SEQ ID NO: 5. In some embodiments, the pigeon GS is derived from the family Columbidae and has an amino acid sequence that is at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to that of SEQ ID NO: 5. In some embodiments, the pigeon GS is derived from the genus Columbia and has an amino acid sequence that is at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to that of SEQ ID NO: 5. In some embodiments, the pigeon GS is derived from the Columbidae family and has an amino acid sequence with at least 95% identity to SEQ ID NO: 5. In some embodiments, the GS provided herein as an alternative marker has reduced activity compared to wild-type pigeon GS. In some embodiments, the GS with reduced activity contains one or more mutations in the glutamate-binding region, ATP-binding region, ammonia-binding region, or any combination thereof. In some embodiments, functional variants with reduced GS activity contain one or more mutations in the β-grasping domain, catalytic domain, or a combination thereof.

[0078] Since the catalytic domain itself is sufficient to provide enhanced GS marker performance, in some embodiments, GS comprising a catalytic domain derived from pigeon GS are provided herein. In some embodiments, the catalytic domain of pigeon GS (e.g., GS from Columbia livia) may consist of amino acids 187-404 of the protein. In some embodiments, the catalytic domain of pigeon GS (e.g., GS from Columbia livia) may consist of amino acids 163-412 of the protein. In some embodiments, the GS comprises a catalytic domain from pigeon GS having an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 5. In some embodiments, the catalytic domain has an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 163-412 shown in SEQ ID NO: 5. In some embodiments, the catalytic domain has amino acids 163-412 shown in SEQ ID NO: 5. The GS may further include an N-terminal region from any other GS. In some embodiments, the GS may also include a β-grasping domain from any other GS. The GS may include an N-terminal region (e.g., a β-grasping domain) from, for example, hamster GS, koala GS, ostrich GS, snake GS, junglefowl GS, gerbil GS, or night bat GS. In some embodiments, the GS provided herein includes a catalytic domain from a pigeon GS and a β-grasping domain from a hamster GS, koala GS, ostrich GS, snake GS, junglefowl GS, gerbil GS, or night bat GS.

[0079] This document provides the use of pigeon GS encoding sequences as selectable markers. In some embodiments, the nucleotide sequence encodes a functional fragment of pigeon GS having the amino acid sequence shown in SEQ ID NO: 5. The functional fragment of pigeon GS having the amino acid sequence shown in SEQ ID NO: 5 maintains the GS activity of a reference GS. In some embodiments, the functional fragment contains at least 100 consecutive amino acids of SEQ ID NO: 5. In some embodiments, the functional fragment contains at least 150 consecutive amino acids of SEQ ID NO: 5. In some embodiments, the functional fragment contains at least 200 consecutive amino acids of SEQ ID NO: 5. In some embodiments, the functional fragment contains at least 250 consecutive amino acids of SEQ ID NO: 5. In some embodiments, the functional fragment contains at least 300 consecutive amino acids of SEQ ID NO: 5. In some embodiments, the functional fragment provided herein as a selectable marker exhibits reduced GS activity compared to wild-type pigeon GS.

[0080] This document provides the use of pigeon GS coding sequences as selectable markers. In some embodiments, pigeon GS exhibits reduced stability at the mRNA or protein level. In some embodiments, the pigeon GS coding sequences provided herein are operatively linked to mRNA destabilizing elements. In some embodiments, the pigeon GS provided herein contains a degradation determinant. The degradation determinant may be any degradation determinant disclosed herein or otherwise known in the art. In some embodiments, the pigeon GS has an N-terminal degradation determinant. In some embodiments, the pigeon GS has a C-terminal degradation determinant. In some embodiments, the degradation determinant is a PEST sequence (e.g., SEQ ID NO: 23). In some embodiments, the degradation determinant is an ODD sequence (e.g., SEQ ID NO: 24). In some embodiments, the degradation determinant is an IκBα sequence (e.g., SEQ ID NO: 25). In some embodiments, the degradation determinant has an amino acid sequence selected from SEQ ID NO: 23-25.

[0081] In some embodiments, this document provides the use of nucleotide sequences encoding *Gallus gallus* GS as optional markers. *Gallus gallus* GS is a GS derived from *Gallus gallus*. In some embodiments, *Gallus gallus* GS is derived from the family Phasianidae. In some embodiments, *Gallus gallus* GS is derived from the genus *Gallus*. In some embodiments, *Gallus gallus* GS is derived from *Gallus gallus*. Exemplary amino acid sequences of *Gallus gallus* GS can be found in the table above and in public databases, such as NCBI#XP_040560456.

[0082] This document provides the use of the Red Junglefowl GS encoding sequence as an alternative marker. In some embodiments, the nucleotide sequence encodes a Red Junglefowl GS having an amino acid sequence comprising or consisting of SEQ ID NO: 6. In some embodiments, the Red Junglefowl GS is a functional variant of the GS having the amino acid sequence shown in SEQ ID NO: 6. Functional variants of the GS having the amino acid sequence shown in SEQ ID NO: 6 maintain the basic structure and GS activity of the reference GS. The functional variant may have, for example, amino acid substitutions, deletions, and / or additions of about 1 to about 25, about 1 to about 20, about 1 to about 15, about 1 to about 10, or about 1 to about 5 amino acid substitutions, deletions, and / or additions compared to SEQ ID NO: 6. In some embodiments, the amino acid sequence of the functional variant differs from SEQ ID NO: 6 by only up to 22, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 amino acid variation, including the presence of substitutions, insertions, and deletions. The variation in the amino acid sequence can be an amino acid substitution. The variation in the amino acid sequence can be a conserved amino acid substitution. The variant may be naturally present in the Red Junglefowl species (e.g., allelic variants or splice variants). Alternatively, the variant can be obtained through genetic engineering. In some embodiments, the functional variant exhibits reduced GS activity compared to wild-type junglefowl GS. In some embodiments, the functional variant with reduced GS activity contains one or more mutations in the glutamate-binding domain, ATP-binding domain, ammonia-binding domain, or any combination thereof. In some embodiments, the functional variant with reduced GS activity contains one or more mutations in the β-grasping domain, catalytic domain, or a combination thereof.

[0083] This document provides the use of the wild red junglefowl GS coding sequence as an optional marker. In some embodiments, the nucleotide sequence encodes a wild red junglefowl GS having an amino acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 6. In some embodiments, the GS has an amino acid sequence having at least 90% identity with SEQ ID NO: 6. In some embodiments, the GS has an amino acid sequence having at least 92% identity with SEQ ID NO: 6. In some embodiments, the GS has an amino acid sequence having at least 95% identity with SEQ ID NO: 6. In some embodiments, the GS has an amino acid sequence having at least 98% identity with SEQ ID NO: 6. In some embodiments, the GS has an amino acid sequence having at least 99% identity with SEQ ID NO: 6. In some embodiments, the GS has the amino acid sequence shown in SEQ ID NO: 6. In some embodiments, the red junglefowl GS is derived from the Phasianidae family and has an amino acid sequence that is at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the amino acid sequence in SEQ ID NO: 6. In some embodiments, the red junglefowl GS is derived from the Gallus genus and has an amino acid sequence that is at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the amino acid sequence in SEQ ID NO: 6. In some embodiments, the wild-type red junglefowl GS is derived from the Phasianidae family and has an amino acid sequence with at least 95% identity to SEQ ID NO: 6. In some embodiments, the GS provided herein as an alternative marker has reduced activity compared to wild-type red junglefowl GS. In some embodiments, the GS with reduced activity contains one or more mutations in the glutamate-binding region, ATP-binding region, amino-binding region, or any combination thereof. In some embodiments, functional variants with reduced GS activity contain one or more mutations in the β-grasping domain, catalytic domain, or any combination thereof.

[0084] Since the catalytic domain itself is sufficient to provide enhanced GS marker performance, in some embodiments, GS comprising a catalytic domain from a wild red junglefowl GS are provided herein. In some embodiments, the catalytic domain of a wild red junglefowl GS (e.g., a GS from *Gallus gallus*) may consist of amino acids 174-391 of the protein. In some embodiments, the catalytic domain of a wild red junglefowl GS (e.g., a GS from *Gallus gallus*) may consist of amino acids 150-399 of the protein. In some embodiments, the GS comprises a catalytic domain from a wild red junglefowl GS having an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 150-399 shown in SEQ ID NO: 6. In some embodiments, the catalytic domain has an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 150-399 shown in SEQ ID NO: 6. In some embodiments, the catalytic domain has amino acids 150-399 as shown in SEQ ID NO: 6. The GS may further include an N-terminal region from any other GS. In some embodiments, the GS may also include a β-grasping domain from any other GS. The GS may include an N-terminal region (e.g., a β-grasping domain) from, for example, hamster GS, koala GS, ostrich GS, snake GS, pigeon GS, gerbil GS, or night bat GS. In some embodiments, the GS provided herein includes a catalytic domain from a red junglefowl GS and a β-grasping domain from a hamster GS, koala GS, ostrich GS, snake GS, pigeon GS, gerbil GS, or night bat GS.

[0085] This document provides the use of a wild-type jungle germ cell (RG) GS encoding sequence as an alternative biomarker. In some embodiments, the nucleotide sequence encodes a functional fragment of a jungle germ cell GS having the amino acid sequence shown in SEQ ID NO: 6. The functional fragment of a jungle germ cell GS having the amino acid sequence shown in SEQ ID NO: 6 maintains the GS activity of a reference GS. In some embodiments, the functional fragment comprises at least 100 consecutive amino acids of SEQ ID NO: 6. In some embodiments, the functional fragment comprises at least 150 consecutive amino acids of SEQ ID NO: 6. In some embodiments, the functional fragment comprises at least 200 consecutive amino acids of SEQ ID NO: 6. In some embodiments, the functional fragment comprises at least 250 consecutive amino acids of SEQ ID NO: 6. In some embodiments, the functional fragment comprises at least 300 consecutive amino acids of SEQ ID NO: 6. In some embodiments, the functional fragment provided herein as an alternative biomarker exhibits reduced GS activity compared to wild-type jungle germ cell GS.

[0086] This document provides the use of the *Gynostemma pentaphyllum* GS coding sequence as an alternative marker. In some embodiments, *Gynostemma pentaphyllum* GS exhibits reduced stability at the mRNA or protein level. In some embodiments, the *Gynostemma pentaphyllum* GS coding sequence provided herein is operatively linked to an mRNA destabilizing element. In some embodiments, the *Gynostemma pentaphyllum* GS provided herein contains a degradation determinant. The degradation determinant may be any degradation determinant disclosed herein or otherwise known in the art. In some embodiments, the *Gynostemma pentaphyllum* GS has an N-terminal degradation determinant. In some embodiments, the *Gynostemma pentaphyllum* GS has a C-terminal degradation determinant. In some embodiments, the degradation determinant is a PEST sequence (e.g., SEQ ID NO: 23). In some embodiments, the degradation determinant is an ODD sequence (e.g., SEQ ID NO: 24). In some embodiments, the degradation determinant is an IκBα sequence (e.g., SEQ ID NO: 25). In some embodiments, the degradation determinant has an amino acid sequence selected from SEQ ID NO: 23-25.

[0087] In some embodiments, this document provides the use of nucleotide sequences encoding gerbil GS as optional markers. Gerbil GS is a GS derived from gerbils. In some embodiments, gerbil GS is derived from the family Muridae. In some embodiments, gerbil GS is derived from the genus Meriones. In some embodiments, gerbil GS is derived from the long-clawed gerbil (Meriones unguiculatus). Exemplary amino acid sequences of gerbil GS can be found in the table above and in public databases, such as NCBI#XP_021488531.

[0088] This document provides the use of the gerbil GS encoding sequence as an alternative marker. In some embodiments, the nucleotide sequence encodes a gerbil GS having an amino acid sequence comprising or consisting of SEQ ID NO: 7. In some embodiments, the gerbil GS is a functional variant of the GS having the amino acid sequence shown in SEQ ID NO: 7. Functional variants of the GS having the amino acid sequence shown in SEQ ID NO: 7 maintain the basic structure and GS activity of the reference GS. The functional variant may have, for example, amino acid substitutions, deletions, and / or additions of about 1 to about 25, about 1 to about 20, about 1 to about 15, about 1 to about 10, or about 1 to about 5 amino acid substitutions, deletions, and / or additions compared to SEQ ID NO: 7. In some embodiments, the amino acid sequence of the functional variant differs from SEQ ID NO: 7 by only up to 22, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 amino acid variation, including the presence of substitutions, insertions, and deletions. The variation in amino acid sequence can be an amino acid substitution. The variation in amino acid sequence can be a conserved amino acid substitution. The variant may be naturally present in gerbil species (e.g., allelic variants or splice variants). Alternatively, the variant can be obtained through genetic engineering. In some embodiments, the functional variant exhibits reduced GS activity compared to the wild-type gerbil GS. In some embodiments, the functional variant with reduced GS activity contains one or more mutations in the glutamate-binding domain, ATP-binding domain, ammonia-binding domain, or any combination thereof. In some embodiments, the functional variant with reduced GS activity contains one or more mutations in the β-grasping domain, catalytic domain, or any combination thereof.

[0089] This document provides the use of the gerbil GS coding sequence as an optional marker. In some embodiments, the nucleotide sequence encodes a gerbil GS having an amino acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 7. In some embodiments, GS has an amino acid sequence having at least 90% identity with SEQ ID NO: 7. In some embodiments, GS has an amino acid sequence having at least 92% identity with SEQ ID NO: 7. In some embodiments, GS has an amino acid sequence having at least 95% identity with SEQ ID NO: 7. In some embodiments, GS has an amino acid sequence having at least 98% identity with SEQ ID NO: 7. In some embodiments, GS has an amino acid sequence having at least 99% identity with SEQ ID NO: 7. In some embodiments, the GS has the amino acid sequence shown in SEQ ID NO: 7. In some embodiments, the gerbil GS is derived from the family Muridae and has an amino acid sequence that is at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the amino acid sequence in SEQ ID NO: 7. In some embodiments, the gerbil GS is derived from the genus Meriones and has an amino acid sequence that is at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the amino acid sequence in SEQ ID NO: 7. In some embodiments, the gerbil GS is derived from the Muridae family and has an amino acid sequence with at least 95% identity to SEQ ID NO: 7. In some embodiments, the GS provided herein as an alternative marker has reduced activity compared to wild-type gerbil GS. In some embodiments, the GS with reduced activity contains one or more mutations in the glutamate-binding region, ATP-binding region, amino-binding region, or any combination thereof. In some embodiments, functional variants with reduced GS activity contain one or more mutations in the β-grasping domain, catalytic domain, or any combination thereof.

[0090] Since the catalytic domain itself is sufficient to provide enhanced GS marker performance, in some embodiments, GS comprising a catalytic domain derived from gerbil GS are provided herein. In some embodiments, the catalytic domain of gerbil GS (e.g., GS from Meriones unguiculatus) may consist of amino acids 134-351 of the protein. In some embodiments, the catalytic domain of gerbil GS (e.g., GS from Meriones unguiculatus) may consist of amino acids 110-359 of the protein. In some embodiments, the GS comprises a catalytic domain from a gerbil GS having an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 7. In some embodiments, the catalytic domain has an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 7. In some embodiments, the catalytic domain has amino acids 110-359 shown in SEQ ID NO: 7. The GS may further include an N-terminal region from any other GS. In some embodiments, the GS may also include a β-grasping domain from any other GS. The GS may include an N-terminal region (e.g., a β-grasping domain) from, for example, hamster GS, koala GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, or night bat GS. In some embodiments, the GS provided herein includes a catalytic domain from a gerbil GS and a β-grasping domain from a hamster GS, koala GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, or night bat GS.

[0091] This document provides the use of gerbil GS-encoding sequences as selectable biomarkers. In some embodiments, the nucleotide sequence encodes a functional fragment of gerbil GS having the amino acid sequence shown in SEQ ID NO: 7. The functional fragment of gerbil GS having the amino acid sequence shown in SEQ ID NO: 7 maintains the GS activity of a reference GS. In some embodiments, the functional fragment comprises at least 100 consecutive amino acids of SEQ ID NO: 7. In some embodiments, the functional fragment comprises at least 150 consecutive amino acids of SEQ ID NO: 7. In some embodiments, the functional fragment comprises at least 200 consecutive amino acids of SEQ ID NO: 7. In some embodiments, the functional fragment comprises at least 250 consecutive amino acids of SEQ ID NO: 7. In some embodiments, the functional fragment comprises at least 300 consecutive amino acids of SEQ ID NO: 7. In some embodiments, the functional fragment provided herein as a selectable biomarker exhibits reduced GS activity compared to wild-type gerbil GS.

[0092] This document provides the use of the gerbil GS coding sequence as an alternative marker. In some embodiments, the gerbil GS exhibits reduced stability at the mRNA or protein level. In some embodiments, the gerbil GS coding sequence provided herein is operatively linked to an mRNA destabilizing element. In some embodiments, the gerbil GS provided herein contains a degradation determinant. The degradation determinant may be any degradation determinant disclosed herein or otherwise known in the art. In some embodiments, the gerbil GS has an N-terminal degradation determinant. In some embodiments, the gerbil GS has a C-terminal degradation determinant. In some embodiments, the degradation determinant is a PEST sequence (e.g., SEQ ID NO: 23). In some embodiments, the degradation determinant is an ODD sequence (e.g., SEQ ID NO: 24). In some embodiments, the degradation determinant is an IκBα sequence (e.g., SEQ ID NO: 25). In some embodiments, the degradation determinant has an amino acid sequence selected from SEQ ID NO: 23-25.

[0093] In some embodiments, this document provides the use of nucleotide sequences encoding the night bat GS as optional markers. The night bat GS is a GS derived from the night bat. In some embodiments, the night bat GS is derived from the family Vespertilionidae. In some embodiments, the night bat GS is derived from the genus Pipistrellus. In some embodiments, the night bat GS is derived from Pipistrellus kuhlii. Exemplary amino acid sequences of the night bat GS are available in the table above and in public databases, such as NCBI#XP_036298993.

[0094] This document provides the use of the *Noctiluca spp.* GS encoding sequence as an alternative marker. In some embodiments, the nucleotide sequence encodes a *Noctiluca spp.* GS having an amino acid sequence comprising or consisting of SEQ ID NO: 8. In some embodiments, the *Noctiluca spp.* GS is a functional variant of a GS having the amino acid sequence shown in SEQ ID NO: 8. Functional variants of a GS having the amino acid sequence shown in SEQ ID NO: 8 maintain the basic structure and GS activity of the reference GS. The functional variant may have, for example, amino acid substitutions, deletions, and / or additions of about 1 to about 25, about 1 to about 20, about 1 to about 15, about 1 to about 10, or about 1 to about 5 amino acid substitutions compared to SEQ ID NO: 8. In some embodiments, the amino acid sequence of the functional variant differs from SEQ ID NO: 8 by only up to 22, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 amino acid variation, including the presence of substitutions, insertions, and deletions. The variation in amino acid sequence can be an amino acid substitution. The variation in amino acid sequence can be a conserved amino acid substitution. The variant may be naturally present in *Noctiluca spp.* species (e.g., allelic variants or splice variants). Alternatively, the variant can be obtained through genetic engineering. In some embodiments, the functional variant exhibits reduced GS activity compared to the wild-type night bat GS. In some embodiments, the functional variant with reduced GS activity contains one or more mutations in the glutamate-binding domain, ATP-binding domain, ammonia-binding domain, or any combination thereof. In some embodiments, the functional variant with reduced GS activity contains one or more mutations in the β-grasping domain, catalytic domain, or a combination thereof.

[0095] This document provides the use of the *Noctiluca nocturnalis* GS encoding sequence as an optional marker. In some embodiments, the nucleotide sequence encodes a *Noctiluca nocturnalis* GS having an amino acid sequence having at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 8. In some embodiments, the GS has an amino acid sequence having at least 90% identity with SEQ ID NO: 8. In some embodiments, the GS has an amino acid sequence having at least 92% identity with SEQ ID NO: 8. In some embodiments, the GS has an amino acid sequence having at least 95% identity with SEQ ID NO: 8. In some embodiments, the GS has an amino acid sequence having at least 98% identity with SEQ ID NO: 8. In some embodiments, the GS has an amino acid sequence having at least 99% identity with SEQ ID NO: 8. In some embodiments, GS has the amino acid sequence shown in SEQ ID NO: 8. In some embodiments, the night bat GS is derived from the family Vespertilionidae and has an amino acid sequence that is at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 8. In some embodiments, the night bat GS is derived from the genus Pipistrellus and has an amino acid sequence that is at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 8. In some embodiments, the nocturnal bat GS is derived from the family Vespertilionidae and has an amino acid sequence with at least 95% identity to SEQ ID NO: 8. In some embodiments, the GS provided herein as an alternative marker has reduced activity compared to wild-type nocturnal bat GS. In some embodiments, the GS with reduced activity contains one or more mutations in the glutamate-binding region, ATP-binding region, ammonia-binding region, or any combination thereof. In some embodiments, functional variants with reduced GS activity contain one or more mutations in the β-grasping domain, catalytic domain, or a combination thereof.

[0096] Since the catalytic domain itself is sufficient to provide enhanced GS marker performance, in some embodiments, GS comprising a catalytic domain from a night-bat GS are provided herein. In some embodiments, the catalytic domain of a night-bat GS (e.g., a GS from *Pipistrellus kuhlii*) may consist of amino acids 134-351 of the protein. In some embodiments, the catalytic domain of a night-bat GS (e.g., a GS from *Pipistrellus kuhlii*) may consist of amino acids 110-359 of the protein. In some embodiments, the GS comprises a catalytic domain from a night-bat GS having an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 8. In some embodiments, the catalytic domain has an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 8. In some embodiments, the catalytic domain has amino acids 110-359 shown in SEQ ID NO: 8. The GS may further include an N-terminal region from any other GS. In some embodiments, the GS may also include a β-grasping domain from any other GS. The GS may include an N-terminal region (e.g., a β-grasping domain) from, for example, a hamster GS, a koala GS, an ostrich GS, a snake GS, a pigeon GS, a junglefowl GS, or a gerbil GS. In some embodiments, the GS provided herein includes a catalytic domain from a night bat GS and a β-grasping domain from a hamster GS, a koala GS, an ostrich GS, a snake GS, a pigeon GS, a junglefowl GS, or a gerbil GS.

[0097] This document provides the use of the *Noctua noctua* GS encoding sequence as an alternative biomarker. In some embodiments, the nucleotide sequence encodes a functional fragment of *Noctua noctua* GS having the amino acid sequence shown in SEQ ID NO: 8. The functional fragment of *Noctua noctua* GS having the amino acid sequence shown in SEQ ID NO: 8 maintains the GS activity of a reference GS. In some embodiments, the functional fragment comprises at least 100 consecutive amino acids of SEQ ID NO: 8. In some embodiments, the functional fragment comprises at least 150 consecutive amino acids of SEQ ID NO: 8. In some embodiments, the functional fragment comprises at least 200 consecutive amino acids of SEQ ID NO: 8. In some embodiments, the functional fragment comprises at least 250 consecutive amino acids of SEQ ID NO: 8. In some embodiments, the functional fragment comprises at least 300 consecutive amino acids of SEQ ID NO: 8. In some embodiments, the functional fragment provided herein as an alternative biomarker exhibits reduced GS activity compared to wild-type *Noctua noctua* GS.

[0098] This document provides the use of the Night Bat GS coding sequence as an optional marker. In some embodiments, the Night Bat GS exhibits reduced stability at the mRNA or protein level. In some embodiments, the Night Bat GS coding sequence provided herein is operatively linked to an mRNA destabilizing element. In some embodiments, the Night Bat GS provided herein contains a degradation determinant. The degradation determinant may be any degradation determinant disclosed herein or otherwise known in the art. In some embodiments, the Night Bat GS has an N-terminal degradation determinant. In some embodiments, the Night Bat GS has a C-terminal degradation determinant. In some embodiments, the degradation determinant is a PEST sequence (e.g., SEQ ID NO: 23). In some embodiments, the degradation determinant is an ODD sequence (e.g., SEQ ID NO: 24). In some embodiments, the degradation determinant is an IκBα sequence (e.g., SEQ ID NO: 25). In some embodiments, the degradation determinant has an amino acid sequence selected from SEQ ID NO: 23-25.

[0099] In some embodiments, this document provides the use of nucleotide sequences encoding chimeric GS (chimeric GS coding sequences) as optional markers, wherein the chimeric GS comprises fragments derived from different species. For example, the chimeric GS may have a first fragment and a second fragment, each independently derived from koala GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS. In some embodiments, the first and second fragments are derived from two different species. For illustrative purposes, in some embodiments, the chimeric GS may have a fragment derived from koala GS and a fragment derived from ostrich GS. In some embodiments, the chimeric GS may have a fragment derived from koala GS and a fragment derived from snake GS. In some embodiments, the chimeric GS may have a fragment derived from ostrich GS and a fragment derived from snake GS. For illustrative purposes, in some embodiments, the chimeric GS may have a fragment derived from koala GS, a fragment derived from ostrich GS, and a fragment derived from snake. Variations and arrangements of combinations of fragments derived from different species are explicitly considered herein. Those skilled in the art will be able to identify specific fusion structures and confirm their GS activity using the methods disclosed herein through routine experiments.

[0100] like Figure 4As shown, the koala GS may contain the amino acid sequence shown in SEQ ID NO: 9 (wild-type hamster GS) selected from the following positions with one or more amino acid substitutions: G12, S19, E24, Q27, V33, C49, C53, V54, E55, E56, F68, S72, S80, V82, M84, E92, Q106, S125, H128, L140, D152, L160, R172, M176, T191, V197, K1 98, H199, A200, R213, K230, V234, A236, S240, T260, H269, K271, K276, R282, H304, K305, D311, D318, S320, T328, E332, A339, C341, F349, I355, V356, D366, and Q367, as specified in SEQ ID NO: 9. In some embodiments, the amino acid substitution may be selected from G12N, S19N, E24N, Q27L, V33I, C49S, C53S, V54I, E55D, E56D, F68Y, S72G, S80V, V82A, M84L, E92D, Q106R, S125T, H128K, L140M, D152N, L160P, R172G, M176A, T191V, V197A, K19 8M, H199P, A200S, R213E, K230E, V234M, A236V, S240P, T260A, H269F, K271E, K276R, R282K, H304N, K305E, D311E, D318N, S320G, T328M, E332D, A339D, C341R, F349Y, I355L, V356I, D366E and Q367E, according to SEQ ID NO: 9.

[0101] Ostrich GS can contain the amino acid sequence shown in SEQ ID NO: 9 (wild-type hamster GS) selected from the following positions with one or more amino acid substitutions: N10, G12, Q15, S19, V33, G39, C49, C53, V54, E56, S70, S72, S80, V82, E92, F98, Q106, K107, P108, K118, D152, L160, R172, M176, T191, Y194, K198 H199, I206, R213, V220, K230, A236, T237, T260, E264, N265, K271, S278, R282, K305, N308, N310, D311, D318, S320, T328, Q331, A339, C341, F349, I355, Q367 and Q370, numbered according to SEQ ID NO: 9. In some embodiments, the amino acid substitution may be selected from: N10S, G12A, Q15H, S19K, V33I, G39H, C49H, C53S, V54L, E56D, S70A, S72G, S80R, V82A, E92D, F98L, Q106R, K107Q, P108S, K118R, D152N, L160P, R172G, M176V, T191G, Y194N, K198M, H19 9P, I206V, R213E, V220I, K230E, A236V, T237S, S240P, T260S, E264D, N265G, K271E, S278G, R282Q, K305E, N308S, N310H, D311E, D318N, S320G, T328N, Q331H, A339D, C341R, F349Y, I355L, Q367E and Q370E, as specified in SEQ ID NO: 9.

[0102] Snake GS can contain the amino acid sequence shown in SEQ ID NO: 9 (wild-type hamster GS) selected from the following positions with one or more amino acid substitutions: N10, G12, S19, E24, V33, G39, D48, C49, C53, E56, S72, S80, V82, E92, F98, F102, Q106, H115, H128, D152, L160, R172, I175, M176, K189, T191, Y194, K198, H 199, I206, I212, R213, V220, K230, A236, T237, S240, T260, N265, H269, K271, E272, S278, R282, K305, N310, D311, D318, S320, T328, E332, A339, C341, F349, I355, Q367 and Q370, numbered according to SEQ ID NO: 9. In some embodiments, the amino acid substitution may be selected from: N10S, G12T, S19K, E24D, V33I, G39F, D48E, C49R, C53N, E56D, S72G, S80V, V82S, E92D, F98L, F102L, Q106R, H115S, H128N, D152N, L160P, R172G, I175V, M176V, K189N, T191G, Y194N, K198M, H1 99P, I206V, I212V, R213D, V220I, K230E, A236V, T237S, S240P, T260N, N265G, H269Q, K271E, E272D, S278G, R282Q, K305E, N310H, D311E, D318N, S320G, T328N, E332V, A339D, C341R, F349Y, I355L, Q367E and Q370E, as specified in SEQ ID NO: 9.

[0103] Pigeon GS can contain the amino acid sequence shown in SEQ ID NO: 9 (wild-type hamster GS) selected from the following positions with one or more amino acid substitutions: N10, G12, Q15, S19, E24, V33, G39, C49, C53, V54, E56, S72, S80, V82, E92, F98, Q106, K107, P108, K118, D152, L160, R172, M176, T191, Y194, K198 H199, I206, R213, V220, K230, A236, T237, S240, T260, E264, N265, K271, R282, K305, N310, D311, D318, S320, T328, Q331, K334, A339, C341, F349, I355, Q367 and Q370, numbered according to SEQ ID NO: 9. In some embodiments, the amino acid substitution may be selected from: N10S, G12A, Q15H, S19K, E24D, V33I, G39H, C49H, C53S, V54L, E56D, S72G, S80R, V82A, E92D, F98L, Q106R, K107Q, P108S, K118R, D152N, L160P, R172G, M176V, T191G, Y194N, K198M, H19 9P, I206V, R213E, V220I, K230E, S240P, A236V, T237S, S240P, T260E, E264D, N265G, K271E, R282Q, K305E, N310H, D311E, D318N, S320G, T328S, Q331H, K334R, A339D, C341R, F349Y, I355L, Q367E and Q370E, as specified in SEQ ID NO: 9.

[0104] Wild-type hamster GS may contain the amino acid sequence shown in SEQ ID NO: 9 (wild-type hamster GS) selected from the following positions with one or more amino acid substitutions: N10, G12, Q15, S19, V33, G39, C49, C53, V54, E56, S70, S72, S80, V82, E92, F98, Q106, K107, P108, E110, K118, D152, L160, R172, M176, T191, Y194, K1 98, H199, I206, R213, V220, K230, A236, T237, S240, T260, E264, N265, K271, R282, K305, N308, N310, D311, D318, S320, T328, Q331, A339, C341, F349, I355, Q367 and Q370, as specified in SEQ ID NO: 9. In some embodiments, the amino acid substitution may be selected from: N10S, G12A, Q15H, S19K, V33I, G39H, C49H, C53S, V54L, E56D, S70A, S72G, S80R, V82A, E92D, F98L, Q106R, K107Q, P108S, E110D, K118R, D152N, L160P, R172G, M176V, T191G, Y194N, K19 8M, H199P, I206V, R213E, V220I, K230E, A236V, T237S, S240P, T260N, E264D, N265G, K271E, R282Q, K305E, N308S, N310H, D311E, D318N, S320G, T328N, Q331H, A339D, C341R, F349Y, I355L, Q367E and Q370E, according to SEQ ID NO: 9.

[0105] Gerbil GS may contain the amino acid sequence shown in SEQ ID NO: 9 (wild-type hamster GS) selected from the following positions with one or more amino acid substitutions: H8, M16, M18, S19, C53, S72, S80, Q106, T116, D122, H128, L140, D152, L160, R172, M176, T191, Y194, K198, H199, I212, K230, S240, T260, H269, K271, R282, K305, D318, F337, A339, C341, F349, I355, and Q367, as numbered according to SEQ ID NO: 9. In some embodiments, the amino acid substitution may be selected from H8Q, M16T, M18L, S19A, C53S, S72G, S80T, Q106R, T116S, D122E, H128R, L140M, D152N, L160P, R172G, M176V, T191A, Y194N, K198M, H199P, I212V, K230E, S240P, T260A, H269F, K271E, R282Q, K305E, D318N, F337L, A339D, C341R, F349Y, I355L, and Q367E.

[0106] The nocturnal bat GS may contain the amino acid sequence shown in SEQ ID NO: 9 (wild-type hamster GS) selected from the following positions with one or more amino acid substitutions: A2, H8, E24, Q27, V33, C49, C53, V54, E55, F68, S72, S80, E92, V97, Q106, H115, L140, D152, L160, K169, R172, M176, V188, T191, Y194, K198, H199, R213, K230, S240, T260, H269, K271, A273, K276, R282, L294, K305, D318, T328, A339, C341, I355, V356, and Q367, according to SEQ ID NO: 9. NO: 9. In some embodiments, the amino acid substitution may be selected from A2S, H8Q, E24D, Q27L, V33I, C49S, C53S, V54I, E55D, F68M, S72G, S80I, E92D, V97I, Q106H, H115Y, L140M, D152N, L160P, K169R, R172G, M176V, V18 8I, T191G, Y194N, K198M, H199P, R213E, K230E, S240P, T260A, H269Y, K271E, A273S, K276R, R282Q, L294Q, K305E, D318N, T328A, A339D, C341R, I355L, V356I and Q367E, numbered according to SEQ ID NO: 9.

[0107] In some embodiments, the amino acid sequence of the GS provided herein, which can be used as a valid selectable marker, may be included in a SEQ ID NOx with one or more amino acid substitutions at positions selected from the following. NO: 9: A2, H8, N10, G12, Q15, M16, M18, S19, E24, Q27, V33, G39, D48, C49, C53, V54, E55, E56, F68, S70, S72, S80, V82, M84, E92, V97, F98, F102, Q106, K107, P108, E110, H115, T116, K118, D122, S125, H128, L140, D152, L160, K169, R172, I175, M176, V188, K189, T191, Y194, V1 97, K198, H199, A200, I206, I212, R213, V220, K230, V234, A236, T237, S240, T260, E264, N265, H269, K271, E272, A273, K276, S278, R282, L294, H304, K305, N308, N310, D311, D318, S320, T328, Q331, E332, K334, F337, A339, C341, F349, I355, V356, D366, Q367, and Q370, as specified in SEQ ID NO: 9.

[0108] In some embodiments, the amino acid sequence of the GS provided herein may include SEQ ID NO: 9 having at least one amino acid substitution in the N-terminal region. The amino acid sequence of the GS provided herein may include SEQ ID NO: 9 having an amino acid substitution in A2. The amino acid substitution may be A2S. The amino acid sequence of the GS provided herein may include SEQ ID NO: 9 having an amino acid substitution in H8. The amino acid substitution may be H8Q. The amino acid sequence of the GS provided herein may include SEQ ID NO: 9 having an amino acid substitution in N10. The amino acid substitution may be N10S. The amino acid sequence of the GS provided herein may include SEQ ID NO: 9 having an amino acid substitution in G12. The amino acid substitution may be G12A, G12N, or G12T. The amino acid sequence of the GS provided herein may include SEQ ID NO: 9 having an amino acid substitution in Q15. The amino acid substitution may be Q15H. The amino acid sequence of the GS provided herein may include SEQ ID NO: 9 having an amino acid substitution in M16T. The amino acid sequence of the GS provided herein may include SEQ ID NO: 9 having an amino acid substitution in M18. The amino acid substitution may be M18L. The amino acid sequence of GS provided herein may be contained in SEQ ID NO: 9 with an amino acid substitution in S19. The amino acid substitution may be S19A, S19K, or S19N. The amino acid sequence of GS provided herein may be contained in SEQ ID NO: 9 with an amino acid substitution in E24. The amino acid substitution may be E24D or E24N. The amino acid sequence of GS provided herein may be contained in SEQ ID NO: 9 with an amino acid substitution in Q27. The amino acid substitution may be Q27L.

[0109] In some embodiments, the amino acid sequence of the GS provided herein may include SEQ ID NO: 9 having at least one amino acid substitution in the β-grasping domain. The amino acid sequence of the GS provided herein may include SEQ ID NO: 9 having an amino acid substitution in V33. The amino acid substitution may be V33I. The amino acid sequence of the GS provided herein may include SEQ ID NO: 9 having an amino acid substitution in G39. The amino acid substitution may be G39H. The amino acid substitution may be G39F or G39H. The amino acid sequence of the GS provided herein may include SEQ ID NO: 9 having an amino acid substitution in D48. The amino acid substitution may be D48E. The amino acid sequence of the GS provided herein may include SEQ ID NO: 9 having an amino acid substitution in C49. The amino acid substitution may be C49S. The amino acid substitution may be C49H, C49S, or C49R. The amino acid sequence of the GS provided herein may include SEQ ID NO: 9 having an amino acid substitution in C53. The amino acid substitution may be C53S or C53N. The amino acid sequence of the GS provided herein may include SEQ ID NO: 9 having an amino acid substitution in V54. The amino acid substitution can be V54I or V54L. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in E55. The amino acid substitution can be E55D. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in E56. The amino acid substitution can be E56D. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in F68. The amino acid substitution can be F68Y or F68M. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in S72. The amino acid substitution can be S70A or S72G. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in S80. The amino acid substitution can be S80V, S80R, S80T, or S80I. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in V82. The amino acid substitution can be V82A or V82S. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in M84. The amino acid substitution may be M84L. The amino acid sequence of GS provided herein may be contained in SEQ ID NO: 9 with an amino acid substitution in E92. The amino acid substitution may be E92D. The amino acid sequence of GS provided herein may be contained in SEQ ID NO: 9 with an amino acid substitution in V97. The amino acid substitution may be V97I.The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at F98. The amino acid substitution can be F98L. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at F102. The amino acid substitution can be F102L. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at Q106. The amino acid substitution can be Q106R or Q106H. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at K107. The amino acid substitution can be K107Q. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at P108. The amino acid substitution can be P108S.

[0110] In some embodiments, the amino acid sequence of the GS provided herein may include SEQ ID NO: 9 with at least one amino acid substitution in the catalytic domain. The amino acid sequence of the GS provided herein may include SEQ ID NO: 9 with an amino acid substitution in E110. The amino acid substitution may be E110D. The amino acid sequence of the GS provided herein may include SEQ ID NO: 9 with an amino acid substitution in H115. The amino acid substitution may be H115S or H115Y. The amino acid sequence of the GS provided herein may include SEQ ID NO: 9 with an amino acid substitution in T116. The amino acid substitution may be T116S. The amino acid sequence of the GS provided herein may include SEQ ID NO: 9 with an amino acid substitution in K118. The amino acid substitution may be K118R. The amino acid sequence of the GS provided herein may include SEQ ID NO: 9 with an amino acid substitution in D122. The amino acid substitution may be D122E. The amino acid sequence of the GS provided herein may include SEQ ID NO: 9 with an amino acid substitution in S125. The amino acid substitution may be S125T. The amino acid sequence of the GS provided herein may include SEQ ID NO: 9 with an amino acid substitution in H128. The amino acid substitution may be H128L. The amino acid substitution may be H128K, H128N, or H128R. The amino acid sequence of the GS provided herein may be contained in SEQ ID NO: 9 with an amino acid substitution in L140. The amino acid substitution may be L140M. The amino acid sequence of the GS provided herein may be contained in SEQ ID NO: 9 with an amino acid substitution in D152. The amino acid substitution may be D152N. The amino acid sequence of the GS provided herein may be contained in SEQ ID NO: 9 with an amino acid substitution in L160. The amino acid substitution may be L160P. The amino acid sequence of the GS provided herein may be contained in SEQ ID NO: 9 with an amino acid substitution in K169. The amino acid substitution may be K169R. The amino acid sequence of the GS provided herein may be contained in SEQ ID NO: 9 with an amino acid substitution in R172. The amino acid substitution may be R172G. The amino acid sequence of the GS provided herein may be contained in SEQ ID NO: 9 with an amino acid substitution in I175. The amino acid substitution may be I175V. The amino acid sequence of the GS provided herein may be contained in SEQ ID NO: 9 with an amino acid substitution in M176. The amino acid substitution can be M176A or M176V. The amino acid sequence of GS provided herein can be contained in SEQ ID NO: 9 of V188 with an amino acid substitution. The amino acid substitution can be V188I.The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in K189. The amino acid substitution can be K189N. The amino acid substitution can be K189Q. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in T191. The amino acid substitution can be T191A, T191V, or T191G. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in Y194. The amino acid substitution can be Y194N. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in V197. The amino acid substitution can be V197A. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in K198. The amino acid substitution can be K198M. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in H199. The amino acid substitution can be H199P.

[0111] The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at A200. The amino acid substitution can be A200S. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at I206. The amino acid substitution can be I206V. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at I212. The amino acid substitution can be I212V. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at R213. The amino acid substitution can be R213E. The amino acid substitution can be R213E or R213D. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at V220. The amino acid substitution can be V220I. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at K230. The amino acid substitution can be K230E. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at V234. The amino acid substitution can be V234M. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at A236. The amino acid substitution can be A236V. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at T237. The amino acid substitution can be T237S. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at S240. The amino acid substitution can be S240P. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at T260. The amino acid substitution can be T260A, T260S, T260N, or T260E. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at E264. The amino acid substitution can be E264D. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at N265. The amino acid substitution can be N265G. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at H269. The amino acid substitution can be H269F, H269Q, or H269Y. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at K271. The amino acid substitution can be K271E. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at E272. The amino acid substitution can be E272D.The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at A273. The amino acid substitution can be A273S. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at K276. The amino acid substitution can be K276R. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at S278. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at R282. The amino acid substitution can be R282K or R282Q. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at L294. The amino acid substitution can be L294Q. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at H304. The amino acid substitution can be H304N. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at K305. The amino acid substitution can be K305E. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution at N308. The amino acid substitution can be N308S. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in N310. The amino acid substitution can be N310H. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in D311. The amino acid substitution can be D311E. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in D318. The amino acid substitution can be D318N. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in S320. The amino acid substitution can be S320G. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in T328. The amino acid substitution can be T328N, T328M, T328S, or T328A. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in Q331. The amino acid substitution can be Q331H. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in E332. The amino acid substitution may be E332D or E332V. The amino acid sequence of the GS provided herein may be contained in SEQ ID NO: 9 with an amino acid substitution in K334. The amino acid substitution may be K334R. The amino acid sequence of the GS provided herein may be contained in SEQ ID NO: 9 with an amino acid substitution in F337.The amino acid substitution can be F337L. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in A339. The amino acid substitution can be A339D. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in C341. The amino acid substitution can be C341R. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in F349. The amino acid substitution can be F349Y. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in I355. The amino acid substitution can be I355L. The amino acid sequence of the GS provided herein can be contained in SEQ ID NO: 9 with an amino acid substitution in V356. The amino acid substitution can be V356I.

[0112] In some embodiments, the amino acid sequence of the GS provided herein may include SEQ ID NO: 9 having at least one amino acid substitution in the C-terminal region. The amino acid sequence of the GS provided herein may include SEQ ID NO: 9 having an amino acid substitution in D366. The amino acid substitution may be D366E. The amino acid sequence of the GS provided herein may include SEQ ID NO: 9 having an amino acid substitution in Q367. The amino acid substitution may be Q367E. The amino acid sequence of the GS provided herein may include SEQ ID NO: 9 having an amino acid substitution in Q370. The amino acid substitution may be Q370E.

[0113] In some embodiments, the GS provided herein, which can be used as an effective selectable marker, may have an amino acid sequence that is at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, or 98% identical to SEQ ID NO: 9, having at least one amino acid at a position selected from the following positions. Replacements: A2, H8, N10, G12, Q15, M16, M18, S19, E24, Q27, V33, G39, D48, C49, C53, V54, E55, E56, F68, S70, S72, S80, V82, M84, E92, V97, F98, F102, Q106, K107, P108, E110, H115, T116, K118 , D122, S125, H128, L140, D152, L160, K169, R172, I175, M176, V188, K189, T191, Y194, V19 7. K198, H199, A200, I206, I212, R213, V220, K230, V234, A236, T237, S240, T260, E264, N2 65, H269, K271, E272, A273, K276, S278, R282, L294, H304, K305, N308, N310, D311, D318, S320, T328, Q331, E332, K334, F337, A339, C341, F349, I355, V356, D366, Q367, and Q370, as designated by SEQ ID NO: 9. In some embodiments, the amino acid sequence of the GS provided herein may comprise SEQ ID NO: 9 having about 1, about 3, about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, or about 80 amino acid substitutions.In some embodiments, the amino acid substitutions are selected from A2S, H8Q, N10S, G12A, G12N, G12T, Q15H, M16T, M18L, S19A, S19K, S19N, E24N, E24D, Q27L, V33I, G39H, G39F, D48E, C49H, C49S, C49R, C53S, C53N, V54I, V54L, E55D, E56D, F68Y, F68M, S70A, S72G, S80V, S8 0R, S80T, S80I, V82A, V82S, M84L, E92D, V97I, F98L, F102L, Q106R, Q106H, K107Q, P108S, E110D, H115S, H115Y, T 116S, K118R, D122E, S125T, H128K, H128N, H128R, L140M, D152N, L160P, K169R, R172G, I175V, M176A, M176V, V18 8I, K189N, T191A, T191V, T191G, Y194N, V197A, K198M, H199P, A200S, I206V, I212V, R213E, R213D, V220I, K230E , V234M, A236V, T237S, S240P, T260A, T260S, T260N, T260E, E264D, N265G, H269F, H269Q, H269Y, K271E, E272D, A 273S, K276R, S278G, R282K, R282Q, L294Q, H304N, K305E, N308S, N310H, D311E, D318N, S320G, T328N, T328M, T328S, T328A, Q331H, E332D, E332V, K334R, F337L, A339D, C341R, F349Y, I355L, V356I, D366E, Q367E and Q370E, numbered according to SEQ ID NO: 9.

[0114] In some embodiments, the GS provided herein may have an amino acid sequence that is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, or at least 98% identical to SEQ ID NO: 9, having at least one amino acid substitution at the following positions: N10, G12, Q15, S19, E24, V33, G39, C49, C53, V54, E56, S72, S80, V82. E92, F98, Q106, K107, P108, K118, L140, D152, L160, R172, M176, T191, Y194, K198, H199, I206, R213, V220, K230, A236, T237, S240, T260, E264, N265, H269, K271, R282, K305, N310, D311, D318, S320, T328, Q331, A339, C341, F349, I355, Q367 and Q370. In some embodiments, the amino acid sequence of the GS provided herein may include SEQ ID NO(s) having about 3, about 5, about 10, about 15, about 20, about 25, about 30, about 35, or about 40 amino acid substitutions at the following positions. NO: 9: N10S, G12A, G12N, G12T, Q15H, S19A, S19K, S19N, E24N, E24D, V33I , G39H, G39F, C49H, C49S, C49R, C53S, C53N, V54I, V54L, E56D, S72G, S80 V, S80R, S80T, S80I, V82A, V82S, E92D, F98L, Q106R, Q106H, K107Q, P108 S, K118R, L140M, D152N, L160P, R172G, M176A, M176V, T191A, T191V, T19 1G, Y194N, K198M, H199P, I206V, R213E, R213D, V220I, K230E, A236V, T237S, S240P, T260A, T260S, T260N, T260E, E264D, N265G, H269F, H269Q, H269Y, K271E, R282K, R282Q, K305E, N310H, D311E, D318N, S320G, T328N, T328M, T328S, T328A, Q331H, A339D, C341R, F349Y, I355L, Q367E and Q370E.In some embodiments, the amino acid sequence of the GS provided herein may include SEQ ID NO: 9 with about one amino acid substitution. In some embodiments, the amino acid sequence of the GS provided herein may include SEQ ID NO: 9 with about three amino acid substitutions. In some embodiments, the amino acid sequence of the GS provided herein may include SEQ ID NO: 9 with about five amino acid substitutions. In some embodiments, the amino acid sequence of the GS provided herein may include SEQ ID NO: 9 with about ten amino acid substitutions. In some embodiments, the amino acid sequence of the GS provided herein may include SEQ ID NO: 9 with about fifteen amino acid substitutions. In some embodiments, the amino acid sequence of the GS provided herein may include SEQ ID NO: 9 with about twenty amino acid substitutions.

[0115] Some of the GS disclosed herein as alternative markers exhibit reduced activity or stability (at the mRNA or protein level). Reduced activity or stability in alternative markers leads to more stringent selection, as higher transcriptional activity or higher copy numbers of the expression cassette are required for growth under these conditions. In some embodiments, functional variants or fragments of GS from koalas, ostriches, snakes, pigeons, junglefowl, gerbils, or night bats are provided herein, exhibiting reduced GS activity compared to their wild-type counterparts. As used herein, the term "reduced activity" refers to a decrease in the ability of a variant or fragment of an enzyme (such as GS) to exercise its enzymatic activity (such as GS activity) compared to its wild-type counterpart. The activity of the reduced enzyme is approximately 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% of the activity of the wild-type enzyme. In other words, if the activity of the wild-type enzyme is considered to be 100%, then compared to the activity of its wild-type counterpart, the activity of the mutant enzyme (or the enzyme with reduced activity) is reduced by approximately 10%, or approximately 20%, or approximately 30%, or approximately 40%, or approximately 50%, or approximately 60%, or approximately 70%, or approximately 80%, or approximately 90%, or approximately 91%, or approximately 92%, or approximately 93%, or approximately 94%, or approximately 95%, or approximately 96%, or approximately 97%, or approximately 98%, or approximately 99%. For example, in some embodiments, the “reduced activity” GS may have approximately 20%, approximately 30%, approximately 40%, approximately 50%, approximately 60%, approximately 70%, approximately 80%, or approximately 90% of the GS activity compared to its wild-type counterpart.

[0116] In some embodiments, the GS disclosed herein, used as an optional marker, exhibits reduced stability at the mRNA level. In some embodiments, the GS coding sequence disclosed herein is operatively linked to an mRNA destabilizing element. In some embodiments, the GS coding sequence disclosed herein is operatively linked to two or more mRNA destabilizing elements. A key element of mRNA stability is the composition of ribonucleoproteins (RNPs)—for example, the binding of poly-A-binding protein (PABP) at the 3' end and the binding of cap-binding protein eIF4E at the 5' end are essential for cytoplasmic mRNA stability. Several other RNA-binding proteins (RBPs) have been found to bind within the 5'- or 3'-untranslated region (UTR) to dynamically regulate mRNA stability under various cellular conditions, particularly at instability-promoting sites such as adenosine monophosphate-uridine monophosphate enrichment elements (AU-enriched elements; AREs). The 3'-UTR also typically contains microRNA binding sites, which generally accelerate mRNA degradation by recruiting degradation mechanisms and stripping stable mRNP components. (Koh et al., SciRep 9, 5976 (2019); Forrest et al., (2020) PLoS ONE 15(2):e0228730). Additionally, stem-loop destabilizing elements (SLDEs) were found to enhance mRNA degradation independently of nearby AREs (Putland et al., Molecular and Cellular Biology, 22(6):1664-73 (2002)). In some embodiments, the disclosed GS coding sequence is operatively linked to an ARE in the 3'-UTR. In some embodiments, the disclosed GS coding sequence is operatively linked to two or more ARE copies in the 3'-UTR. In some embodiments, the disclosed GS coding sequence is operatively linked to an SLDE in the 3'-UTR. In some embodiments, the disclosed GS coding sequence is operatively linked to both an ARE and an SLDE in the 3'-UTR.

[0117] In some embodiments, the GS disclosed herein, used as an optional marker, exhibits reduced stability at the protein level. In some embodiments, the GS provided herein contains degradation determinants. In some embodiments, the GS provided herein contains two or more degradation determinants. Intracellular protein degradation is primarily mediated by the ubiquitin (Ub)-proteasome system (UPS). Ub ligases recognize substrate proteins through their degradation signals (i.e., degradation determinants) and conjugate Ub, a 9-kDa protein (typically in the form of a poly-Ub chain), to amino acid residues (typically an inner lysine) of the target substrate, thereby targeting it for degradation via the proteosome. (Varshavsky, PNAS (2019) 116(2):358-366; Timms; BiochemSoc Trans. 2020 48(4):1557–1567). Degradation determinants can be located at the N-terminus (N-degradation determinant), C-terminus (C-degradation determinant), or internal sites of the target protein. Therefore, in some embodiments, the GS containing a degradation determinant may have a degradation determinant fused to its N-terminus or C-terminus. In some embodiments, the nucleotide sequence encoding a platypus GS provided herein is linked at its 5' or 3' end to a nucleotide sequence encoding a degradation determinant, thereby causing the degradation determinant to be fused to the N-terminus or C-terminus of the GS. In some embodiments, the GS provided herein contains an Arg / N-degradation determinant. In some embodiments, the GS provided herein contains an Ac / N-degradation determinant. In some embodiments, the GS provided herein contains an fMet / N-degradation determinant. In some embodiments, the GS provided herein contains a Pro / N-degradation determinant. In some embodiments, the GS provided herein contains a degradation determinant having a synthetic degradation determinant, which is typically a non-natural short peptide having 5-30 amino acids. In some embodiments, the GS provided herein contains a PEST degradation determinant (a peptide sequence rich in proline (P), glutamic acid (E), serine (S), and threonine (T)). In some embodiments, the GS provided herein contains a PEST degradation determinant from ornithine decarboxylase: SHGFPPEVEEQAAGTLPMSCAQESGMDRHPAACASARINV (SEQ ID NO: 23).In some embodiments, the GS provided herein comprises an ODD (oxygen-dependent degradation) domain from transcription factor HIF1a aa530-652: EFKLELVEKLFAEDTEAKNPFSTQDTDLDLEMLAPYIPMDDDFQLRSFDQ LSPLESSSASPESASPQSTVTVFQQTQIQEPTANATTTTATTDELKTVTKD RMEDIKILIASPSPTHIHKETT (SEQ ID NO: 24). In some embodiments, the GS provided herein comprises an IkappaBalpha (IκBα) degradation determinant: IQQQLGQLTLENLQMLPESEDEESYDTESEFTEFTEDELPYDDCVFGGQR (SEQ ID NO: 25).

[0118] Accordingly, this document provides the use of nucleotide sequences encoding the GS of koalas, ostriches, snakes, pigeons, junglefowl, gerbils, or night bats disclosed herein, or functional variants or fragments thereof, as selectable markers. In some embodiments, the selectable marker can be used to identify genomic loci for integration of expression cassettes. Genomic loci are selected due to their high transcriptional activity. In some embodiments, the selectable marker can be used to identify cell clones capable of producing POIs. Cell clones are selected due to their high POI yield. In some embodiments, the selectable marker is used for recombinant production of POIs. POIs are further described in the following sections. For example, POIs may be selected from antibodies, enzymes, soluble proteins, secretory proteins, membrane proteins, and fusion proteins. In some embodiments, POIs are recombinantly produced in mammalian cells. In some embodiments, POIs are recombinantly produced via CHO cells. 6.2 Carrier

[0119] This document also provides deoxyribonucleic acid (DNA) vectors containing a nucleotide sequence encoding a GS (GS-coding sequence), wherein the GS is a koala GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS. In some embodiments, the vectors provided herein are suitable for recombinant production and contain an expression cassette. In some embodiments, the vectors provided herein are suitable for genome integration.

[0120] As used herein and as understood in the art, the term "vector" refers to a medium for carrying genetic material (e.g., nucleotide sequences) that can be introduced into a host cell, where it can be replicated and / or expressed. Suitable vectors include, for example, expression vectors, plasmids, phage vectors, viral vectors, episomes, and artificial chromosomes. A vector may include one or more selectable marker genes and appropriate expression control sequences. Selectable marker genes may be included to provide antibiotic or toxin resistance, supplement nutrient deficiencies, or provide critical nutrients absent in the culture medium. Expression control sequences may include constitutive and inducible promoters, transcription enhancers, transcription terminators, etc., well known in the art. DNA regions (such as control elements and protein-coding sequences) may be "operably linked" when they are functionally related to each other. For example, if a promoter controls transcription, the promoter is operably linked to the coding sequence; or if the location of the ribosome binding site allows translation, it is operably linked to the coding sequence.

[0121] As used herein and as understood in the art, an "expression cassette" is a distinct and contiguous component of a vector DNA that includes regulatory sequences that control the expression of nucleotide sequences that the expression cassette may carry. Regulatory sequences include, for example, transcription initiation (promoter) and termination sequences, enhancers, introns, origin of replication sites, polyadenylation sequences, peptide signaling, and chromatin isolating elements. Regulatory sequences are described in, for example, Goeddel, G ENE E XPRESSION T ECHNOLOGY :M ETHODS IN E NZYMOLOGY 185, Academic Press, San Diego, Calif. (1990). Simply put, expression cassettes guide host cell mechanisms to prepare RNA and proteins encoded by the nucleotide sequences contained within the cassette. Therefore, expression in cells from different organisms or species (e.g., bacteria, yeast, plant, and mammalian cells) requires different regulatory sequences. The vectors described herein can have one or more expression cassettes. An expression cassette can be “empty” and contain a multiple cloning site (MCS) for inserting a nucleotide sequence encoding a POI (POI-coding sequence). Expression cassettes containing a POI-coding sequence can be loaded. As used herein and understood in the art, a “multiple cloning site” or “MCS” refers to a short DNA segment on the vector containing multiple restriction sites to allow insertion of a POI-coding sequence.

[0122] In some embodiments, the expression cassette may include a POI-encoding nucleotide sequence. In some embodiments, the expression may have more than one POI-encoding nucleotide sequence, i.e., polycistronic. A polycistronic expression cassette contains more than one cistron, which can be transcribed into mRNA simultaneously expressing two or more individual polypeptides. In some embodiments, the expression cassette may be bicistronic, i.e., containing two cistrons. mRNA transcribed from a bicistronic expression cassette can simultaneously express two individual polypeptides. In some embodiments, a tricistronic expression cassette may be tricistronic, i.e. containing three cistrons. mRNA transcribed from a tricistronic expression vector can simultaneously express three individual polypeptides.

[0123] Cistrons within an expression cassette can be separated by, for example, internal ribosome entry sites (IRES) or 2A elements. As understood in the art, an IRES refers to a nucleotide sequence within the expression cassette that, when transcribed into mRNA, can directly recruit ribosomes without requiring prior scanning of the untranslated regions of the mRNA by the ribosomes. As understood in the art, 2A elements encoding self-cleaving short peptides (approximately 20 amino acids) provide a mechanism for the subsequent separation of equimolarly generated target peptides. Illustrative 2A self-cleaving peptides include P2A, E2A, F2A, and T2A (see table below).

[0124] The DNA vector provided herein contains a nucleotide sequence encoding a GS, which is any koala GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS described herein. The nucleotide sequence encoding such a GS can be a naturally occurring nucleotide sequence. Alternatively, the triplet codons encoding such a GS can be optimized for expression in specific host cells, such as CHO cells. Software and algorithms for codon optimization are known in the art, including, for example, the algorithms described in Raab et al. (2010, SystSynth Biol. 4:215-25).

[0125] In some embodiments, this document provides a DNA vector having a nucleotide sequence encoding a GS as a koala GS (koala GS coding sequence). The koala GS can be any koala GS described herein. In some embodiments, the koala GS originates from the family Phascolarctidae. In some embodiments, the koala GS originates from the genus Phascolarctos. In some embodiments, the koala GS originates from the koala (Phascolarctos cinerecus). In some embodiments, the koala GS has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 1. In some embodiments, the GS has the amino acid sequence shown in SEQ ID NO: 1. In some embodiments, the koala GS is a functional variant of the GS having the amino acid sequence shown in SEQ ID NO: 1. In some embodiments, the koala GS has an amino acid sequence having at least 95% identity with SEQ ID NO: 1. In some embodiments, the koala GS is a functional fragment of GS having the amino acid sequence shown in SEQ ID NO: 1. In some embodiments, the koala GS has an amino acid sequence comprising at least 100 consecutive amino acids of SEQ ID NO: 1.

[0126] In some embodiments, this document provides a DNA vector having a nucleotide sequence encoding a GS comprising a catalytic domain from a koala GS. In some embodiments, the koala GS (e.g., the GS from *Phascolarctos cinerecus*) may consist of amino acids 134-351 of the protein. In some embodiments, the catalytic domain of the koala GS (e.g., the GS from *Phascolarctos cinerecus*) may consist of amino acids 110-359 of the protein. In some embodiments, the GS comprises a catalytic domain from a koala GS having an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 1. In some embodiments, the catalytic domain has an amino acid sequence that is at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to amino acids 110-359 shown in SEQ ID NO: 1. In some embodiments, the catalytic domain has amino acids 110-359 shown in SEQ ID NO: 1.

[0127] In some embodiments, the koala GS has reduced activity compared to the wild-type koala GS. In some embodiments, the koala GS exhibits reduced stability at the mRNA or protein level. In some embodiments, the koala GS coding sequence is operatively linked to an mRNA destabilizing element. In some embodiments, the koala GS contains a degradation determinant. The degradation determinant may be any degradation determinant disclosed herein or otherwise known in the art. In some embodiments, the koala GS has an N-terminal degradation determinant. In some embodiments, the koala GS has a C-terminal degradation determinant. In some embodiments, the degradation determinant is a PEST sequence (e.g., SEQ ID NO: 23). In some embodiments, the degradation determinant is an ODD sequence (e.g., SEQ ID NO: 24). In some embodiments, the degradation determinant is an IκBα sequence (e.g., SEQ ID NO: 25). In some embodiments, the degradation determinant has an amino acid sequence selected from SEQ ID NO: 23-25.

[0128] In some embodiments, a DNA vector is provided having a nucleotide sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 10. In some embodiments, the nucleotide sequence encoding the koala GS has at least 80% identity with SEQ ID NO: 10. In some embodiments, the nucleotide sequence encoding the koala GS has at least 85% identity with SEQ ID NO: 10. In some embodiments, the nucleotide sequence encoding the koala GS has at least 90% identity with SEQ ID NO: 10. In some embodiments, the nucleotide sequence encoding the koala GS has at least 95% identity with SEQ ID NO: 10. In some embodiments, the nucleotide sequence encoding the koala GS is identical to that of SEQ ID NO: 10. In some embodiments, the nucleotide sequence encoding the koala GS is optimized for expression via a specific host cell, such as a CHO cell.

[0129] In some embodiments, this document provides a DNA vector having a nucleotide sequence encoding a GS as an ostrich GS (ostrich GS coding sequence). The ostrich GS can be any ostrich GS described herein. In some embodiments, the ostrich GS is derived from the family Struthionidae. In some embodiments, the ostrich GS is derived from the genus Struthio. In some embodiments, the ostrich GS is derived from the South African ostrich (Struthio camelus australis). In some embodiments, the ostrich GS has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 2. In some embodiments, the GS has the amino acid sequence shown in SEQ ID NO: 2. In some embodiments, the ostrich GS is a functional variant of the GS having the amino acid sequence shown in SEQ ID NO: 2. In some embodiments, the ostrich GS has an amino acid sequence having at least 95% identity with SEQ ID NO: 2. In some embodiments, the ostrich GS is a functional fragment of the GS having the amino acid sequence shown in SEQ ID NO: 2. In some embodiments, the ostrich GS has an amino acid sequence comprising at least 100 consecutive amino acids of SEQ ID NO: 2.

[0130] In some embodiments, this document provides a DNA vector having a nucleotide sequence encoding a GS comprising a catalytic domain from an ostrich GS. In some embodiments, the catalytic domain of an ostrich GS (e.g., a GS from the South African ostrich (Struthio camelus australis)) may consist of amino acids 134-351 of the protein. In some embodiments, the catalytic domain of an ostrich GS (e.g., a GS from the South African ostrich (Struthio camelus australis)) may consist of amino acids 110-359 of the protein. In some embodiments, the GS comprises a catalytic domain from an ostrich GS having an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 2. In some embodiments, the catalytic domain has an amino acid sequence that is at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to amino acids 110-359 shown in SEQ ID NO: 2. In some embodiments, the catalytic domain has amino acids 110-359 shown in SEQ ID NO: 2.

[0131] In some embodiments, ostrich GS has reduced activity compared to wild-type ostrich GS. In some embodiments, ostrich GS exhibits reduced stability at the mRNA or protein level. In some embodiments, the coding sequence of ostrich GS is operatively linked to an mRNA destabilizing element. In some embodiments, ostrich GS contains a degradation determinant. The degradation determinant may be any degradation determinant disclosed herein or otherwise known in the art. In some embodiments, ostrich GS has an N-terminal degradation determinant. In some embodiments, ostrich GS has a C-terminal degradation determinant. In some embodiments, the degradation determinant is a PEST sequence (e.g., SEQ ID NO: 23). In some embodiments, the degradation determinant is an ODD sequence (e.g., SEQ ID NO: 24). In some embodiments, the degradation determinant is an IκBα sequence (e.g., SEQ ID NO: 25). In some embodiments, the degradation determinant has an amino acid sequence selected from SEQ ID NO: 23-25.

[0132] In some embodiments, a DNA vector is provided having a nucleotide sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 11. In some embodiments, the nucleotide sequence encoding ostrich GS has at least 80% identity with SEQ ID NO: 11. In some embodiments, the nucleotide sequence encoding ostrich GS has at least 85% identity with SEQ ID NO: 11. In some embodiments, the nucleotide sequence encoding ostrich GS has at least 90% identity with SEQ ID NO: 11. In some embodiments, the nucleotide sequence encoding ostrich GS has at least 95% identity with SEQ ID NO: 11. In some embodiments, the nucleotide sequence encoding ostrich GS is identical to SEQ ID NO: 11. In some embodiments, the nucleotide sequence encoding ostrich GS is optimized for expression via specific host cells, such as CHO cells.

[0133] In some embodiments, this document provides a DNA vector having a nucleotide sequence encoding a snake GS (snake GS coding sequence). The snake GS can be any snake GS described herein. In some embodiments, the snake GS originates from the family Elapidae. In some embodiments, the snake GS originates from the genus *Pseudonaja*. In some embodiments, the snake GS originates from the eastern brown snake *Pseudonaja textilis*. In some embodiments, the snake GS has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3. In some embodiments, the GS has the amino acid sequence shown in SEQ ID NO: 3. In some embodiments, the snake GS is a functional variant of the GS having the amino acid sequence shown in SEQ ID NO: 3. In some embodiments, the snake GS has an amino acid sequence having at least 95% identity with SEQ ID NO: 3. In some embodiments, the snake GS is a functional fragment of the GS having the amino acid sequence shown in SEQ ID NO: 3. In some embodiments, the snake GS has an amino acid sequence comprising at least 100 consecutive amino acids of SEQ ID NO: 3.

[0134] In some embodiments, this document provides a DNA vector having a nucleotide sequence encoding a GS comprising a catalytic domain from a snake GS. In some embodiments, the catalytic domain of a snake GS (e.g., a GS from the eastern brown snake (Pseudonaja textilis)) may consist of amino acids 134-351 of the protein. In some embodiments, the catalytic domain of a snake GS (e.g., a GS from the eastern brown snake (Pseudonaja textilis)) may consist of amino acids 110-359 of the protein. In some embodiments, the GS comprises a catalytic domain from a snake GS having an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with the amino acid sequence shown in SEQ ID NO: 3. In some embodiments, the catalytic domain has an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with the amino acid sequence shown in SEQ ID NO: 3. In some embodiments, the catalytic domain has amino acids 110-359 as shown in SEQ ID NO: 3.

[0135] In some embodiments, snake GS has reduced activity compared to wild-type snake GS. In some embodiments, snake GS exhibits reduced stability at the mRNA or protein level. In some embodiments, the coding sequence of snake GS is operatively linked to an mRNA destabilizing element. In some embodiments, snake GS contains a degradation determinant. The degradation determinant may be any degradation determinant disclosed herein or otherwise known in the art. In some embodiments, snake GS has an N-terminal degradation determinant. In some embodiments, snake GS has a C-terminal degradation determinant. In some embodiments, the degradation determinant is a PEST sequence (e.g., SEQ ID NO: 23). In some embodiments, the degradation determinant is an ODD sequence (e.g., SEQ ID NO: 24). In some embodiments, the degradation determinant is an IκBα sequence (e.g., SEQ ID NO: 25). In some embodiments, the degradation determinant has an amino acid sequence selected from SEQ ID NO: 23-25.

[0136] In some embodiments, a DNA vector is provided having a nucleotide sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 12. In some embodiments, the nucleotide sequence encoding snake GS has at least 80% identity with SEQ ID NO: 12. In some embodiments, the nucleotide sequence encoding snake GS has at least 85% identity with SEQ ID NO: 12. In some embodiments, the nucleotide sequence encoding snake GS has at least 90% identity with SEQ ID NO: 12. In some embodiments, the nucleotide sequence encoding snake GS has at least 95% identity with SEQ ID NO: 12. In some embodiments, the nucleotide sequence encoding snake GS is identical to SEQ ID NO: 12. In some embodiments, the nucleotide sequence encoding snake GS is optimized for expression via specific host cells, such as CHO cells.

[0137] In some embodiments, this document provides a DNA vector having a nucleotide sequence encoding a GS as a pigeon GS (pigeon GS coding sequence). The pigeon GS can be any pigeon GS described herein. In some embodiments, the pigeon GS originates from the family Columbidae. In some embodiments, the pigeon GS originates from the genus *Columbia*. In some embodiments, the pigeon GS originates from *Columbia livia*. In some embodiments, the pigeon GS has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 5. In some embodiments, the GS has the amino acid sequence shown in SEQ ID NO: 5. In some embodiments, the pigeon GS is a functional variant of the GS having the amino acid sequence shown in SEQ ID NO: 5. In some embodiments, the pigeon GS has an amino acid sequence having at least 95% identity with SEQ ID NO: 5. In some embodiments, the pigeon GS is a functional fragment of the GS having the amino acid sequence shown in SEQ ID NO: 5. In some embodiments, the pigeon GS has an amino acid sequence comprising at least 100 consecutive amino acids of SEQ ID NO: 5.

[0138] In some embodiments, this document provides a DNA vector having a nucleotide sequence encoding a GS comprising a catalytic domain from a pigeon GS. In some embodiments, the catalytic domain of a pigeon GS (e.g., a GS from Columbialivia) may consist of amino acids 187-404 of the protein. In some embodiments, the catalytic domain of a pigeon GS (e.g., a GS from Columbialivia) may consist of amino acids 163-412 of the protein. In some embodiments, the GS comprises a catalytic domain from a pigeon GS having an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with the amino acid sequence shown in SEQ ID NO: 5. In some embodiments, the catalytic domain has an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with the amino acid sequence shown in SEQ ID NO: 5. In some embodiments, the catalytic domain has amino acids 163-412 as shown in SEQ ID NO: 5.

[0139] In some embodiments, pigeon GS has reduced activity compared to wild-type pigeon GS. In some embodiments, pigeon GS exhibits reduced stability at the mRNA or protein level. In some embodiments, the coding sequence of pigeon GS is operatively linked to an mRNA destabilizing element. In some embodiments, pigeon GS contains a degradation determinant. The degradation determinant may be any degradation determinant disclosed herein or otherwise known in the art. In some embodiments, pigeon GS has an N-terminal degradation determinant. In some embodiments, pigeon GS has a C-terminal degradation determinant. In some embodiments, the degradation determinant is a PEST sequence (e.g., SEQ ID NO: 23). In some embodiments, the degradation determinant is an ODD sequence (e.g., SEQ ID NO: 24). In some embodiments, the degradation determinant is an IκBα sequence (e.g., SEQ ID NO: 25). In some embodiments, the degradation determinant has an amino acid sequence selected from SEQ ID NO: 23-25.

[0140] In some embodiments, a DNA vector is provided having a nucleotide sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 14. In some embodiments, the nucleotide sequence encoding pigeon GS has at least 80% identity with SEQ ID NO: 14. In some embodiments, the nucleotide sequence encoding pigeon GS has at least 85% identity with SEQ ID NO: 14. In some embodiments, the nucleotide sequence encoding pigeon GS has at least 90% identity with SEQ ID NO: 14. In some embodiments, the nucleotide sequence encoding pigeon GS has at least 95% identity with SEQ ID NO: 14. In some embodiments, the nucleotide sequence encoding pigeon GS is identical to SEQ ID NO: 14. In some embodiments, the nucleotide sequence encoding pigeon GS is optimized for expression via specific host cells, such as CHO cells.

[0141] In some embodiments, this document provides a DNA vector having a nucleotide sequence encoding a GS as a red junglefowl GS (red junglefowl GS coding sequence). The red junglefowl GS can be any red junglefowl GS described herein. In some embodiments, the red junglefowl GS is derived from the family Phasianidae. In some embodiments, the red junglefowl GS is derived from the genus *Gallus*. In some embodiments, the red junglefowl GS is derived from *Gallus gallus*. In some embodiments, the red junglefowl GS has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 6. In some embodiments, the GS has the amino acid sequence shown in SEQ ID NO: 6. In some embodiments, the red junglefowl GS is a functional variant of the GS having the amino acid sequence shown in SEQ ID NO: 6. In some embodiments, the red junglefowl GS has an amino acid sequence having at least 95% identity with SEQ ID NO: 6. In some embodiments, the red junglefowl GS is a functional fragment of the GS having the amino acid sequence shown in SEQ ID NO: 6. In some embodiments, the chicken GS has an amino acid sequence comprising at least 100 consecutive amino acids of SEQ ID NO: 6.

[0142] In some embodiments, this document provides a DNA vector having a nucleotide sequence encoding a GS comprising a catalytic domain from a wild junglefowl GS. In some embodiments, the catalytic domain of a wild junglefowl GS (e.g., a GS from *Gallus gallus*) may consist of amino acids 174-391 of the protein. In some embodiments, the catalytic domain of a wild junglefowl GS (e.g., a GS from *Gallus gallus*) may consist of amino acids 150-399 of the protein. In some embodiments, the GS comprises a catalytic domain from a wild junglefowl GS having an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 150-399 shown in SEQ ID NO: 6. In some embodiments, the catalytic domain has an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 150-399 shown in SEQ ID NO: 6. In some embodiments, the catalytic domain has amino acids 150-399 as shown in SEQ ID NO: 6.

[0143] In some embodiments, the red junglefowl GS has reduced activity compared to wild-type red junglefowl GS. In some embodiments, the stability of the red junglefowl GS is reduced at the mRNA or protein level. In some embodiments, the coding sequence of the red junglefowl GS is operatively linked to an mRNA destabilizing element. In some embodiments, the red junglefowl GS contains a degradation determinant. The degradation determinant may be any degradation determinant disclosed herein or otherwise known in the art. In some embodiments, the red junglefowl GS has an N-terminal degradation determinant. In some embodiments, the red junglefowl GS has a C-terminal degradation determinant. In some embodiments, the degradation determinant is a PEST sequence (e.g., SEQ ID NO: 23). In some embodiments, the degradation determinant is an ODD sequence (e.g., SEQ ID NO: 24). In some embodiments, the degradation determinant is an IκBα sequence (e.g., SEQ ID NO: 25). In some embodiments, the degradation determinant has an amino acid sequence selected from SEQ ID NO: 23-25.

[0144] In some embodiments, a DNA vector is provided herein having a nucleotide sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 15. In some embodiments, the nucleotide sequence encoding the red junglefowl GS has at least 80% identity with SEQ ID NO: 15. In some embodiments, the nucleotide sequence encoding the red junglefowl GS has at least 85% identity with SEQ ID NO: 15. In some embodiments, the nucleotide sequence encoding the red junglefowl GS has at least 90% identity with SEQ ID NO: 15. In some embodiments, the nucleotide sequence encoding the red junglefowl GS has at least 95% identity with SEQ ID NO: 15. In some embodiments, the nucleotide sequence encoding the red junglefowl GS is identical to SEQ ID NO: 15. In some embodiments, the nucleotide sequence encoding the red junglefowl GS is optimized for expression via specific host cells, such as CHO cells.

[0145] In some embodiments, this document provides a DNA vector having a nucleotide sequence encoding a GS as a gerbil GS (gerbil GS coding sequence). The gerbil GS can be any gerbil GS described herein. In some embodiments, the gerbil GS originates from the family Muridae. In some embodiments, the gerbil GS originates from the genus Meriones. In some embodiments, the gerbil GS originates from the long-clawed gerbil (Meriones unguiculatus). In some embodiments, the gerbil GS has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 7. In some embodiments, the GS has the amino acid sequence shown in SEQ ID NO: 7. In some embodiments, the gerbil GS is a functional variant of the GS having the amino acid sequence shown in SEQ ID NO: 7. In some embodiments, the gerbil GS has an amino acid sequence having at least 95% identity with SEQ ID NO: 7. In some embodiments, the gerbil GS is a functional fragment of the GS having the amino acid sequence shown in SEQ ID NO: 7. In some embodiments, the gerbil GS has an amino acid sequence comprising at least 100 consecutive amino acids of SEQ ID NO: 7.

[0146] In some embodiments, this document provides a DNA vector having a nucleotide sequence encoding a GS comprising a catalytic domain from a gerbil GS. In some embodiments, the catalytic domain of a gerbil GS (e.g., a GS from the long-clawed gerbil (Meriones unguiculatus)) may consist of amino acids 134-351 of the protein. In some embodiments, the catalytic domain of a gerbil GS (e.g., a GS from the long-clawed gerbil (Meriones unguiculatus)) may consist of amino acids 110-359 of the protein. In some embodiments, the GS comprises a catalytic domain from a gerbil GS having an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 7. In some embodiments, the catalytic domain has an amino acid sequence that is at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to amino acids 110-359 shown in SEQ ID NO: 7. In some embodiments, the catalytic domain has amino acids 110-359 shown in SEQ ID NO: 7.

[0147] In some embodiments, the gerbil GS has reduced activity compared to wild-type gerbil GS. In some embodiments, the gerbil GS exhibits reduced stability at the mRNA or protein level. In some embodiments, the gerbil GS coding sequence is operatively linked to an mRNA destabilizing element. In some embodiments, the gerbil GS contains a degradation determinant. The degradation determinant may be any degradation determinant disclosed herein or otherwise known in the art. In some embodiments, the gerbil GS has an N-terminal degradation determinant. In some embodiments, the gerbil GS has a C-terminal degradation determinant. In some embodiments, the degradation determinant is a PEST sequence (e.g., SEQ ID NO: 23). In some embodiments, the degradation determinant is an ODD sequence (e.g., SEQ ID NO: 24). In some embodiments, the degradation determinant is an IκBα sequence (e.g., SEQ ID NO: 25). In some embodiments, the degradation determinant has an amino acid sequence selected from SEQ ID NO: 23-25.

[0148] In some embodiments, a DNA vector is provided having a nucleotide sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 16. In some embodiments, the nucleotide sequence encoding gerbil GS has at least 80% identity with SEQ ID NO: 16. In some embodiments, the nucleotide sequence encoding gerbil GS has at least 85% identity with SEQ ID NO: 16. In some embodiments, the nucleotide sequence encoding gerbil GS has at least 90% identity with SEQ ID NO: 16. In some embodiments, the nucleotide sequence encoding gerbil GS has at least 95% identity with SEQ ID NO: 16. In some embodiments, the nucleotide sequence encoding gerbil GS is identical to SEQ ID NO: 16. In some embodiments, the nucleotide sequence encoding gerbil GS is optimized for expression via specific host cells, such as CHO cells.

[0149] In some embodiments, this document provides a DNA vector having a nucleotide sequence encoding a GS as a night bat GS (night bat GS coding sequence). The night bat GS can be any night bat GS described herein. In some embodiments, the night bat GS originates from the family Vespertilionidae. In some embodiments, the night bat GS originates from the genus Pipistrellus. In some embodiments, the night bat GS originates from Pipistrellus kuhlii. In some embodiments, the night bat GS has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 8. In some embodiments, the GS has the amino acid sequence shown in SEQ ID NO: 8. In some embodiments, the night bat GS is a functional variant of a GS having the amino acid sequence shown in SEQ ID NO: 8. In some embodiments, the night bat GS has an amino acid sequence having at least 95% identity with SEQ ID NO: 8. In some embodiments, the night bat GS is a functional fragment of a GS having the amino acid sequence shown in SEQ ID NO: 8. In some embodiments, the Night Bat GS has an amino acid sequence comprising at least 100 consecutive amino acids of SEQ ID NO: 8.

[0150] In some embodiments, this document provides a DNA vector having a nucleotide sequence encoding a GS comprising a catalytic domain from a night bat GS. In some embodiments, the catalytic domain of a night bat GS (e.g., a GS from *Pipistrellus kuhlii*) may consist of amino acids 134-351 of the protein. In some embodiments, the catalytic domain of a night bat GS (e.g., a GS from *Pipistrellus kuhlii*) may consist of amino acids 110-359 of the protein. In some embodiments, the GS comprises a catalytic domain from a night bat GS having an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 8. In some embodiments, the catalytic domain has an amino acid sequence that is at least 85%, at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to amino acids 110-359 shown in SEQ ID NO: 8. In some embodiments, the catalytic domain has amino acids 110-359 shown in SEQ ID NO: 8.

[0151] In some embodiments, the *Noctua spp.* GS has reduced activity compared to wild-type *Noctua spp.* In some embodiments, the *Noctua spp.* GS exhibits reduced stability at the mRNA or protein level. In some embodiments, the coding sequence of the *Noctua spp.* GS is operatively linked to an mRNA destabilizing element. In some embodiments, the *Noctua spp.* GS contains a degradation determinant. The degradation determinant may be any degradation determinant disclosed herein or otherwise known in the art. In some embodiments, the *Noctua spp.* GS has an N-terminal degradation determinant. In some embodiments, the *Noctua spp.* GS has a C-terminal degradation determinant. In some embodiments, the degradation determinant is a PEST sequence (e.g., SEQ ID NO: 23). In some embodiments, the degradation determinant is an ODD sequence (e.g., SEQ ID NO: 24). In some embodiments, the degradation determinant is an IκBα sequence (e.g., SEQ ID NO: 25). In some embodiments, the degradation determinant has an amino acid sequence selected from SEQ ID NO: 23-25.

[0152] In some embodiments, a DNA vector is provided having a nucleotide sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 17. In some embodiments, the nucleotide sequence encoding the night bat GS has at least 80% identity with SEQ ID NO: 17. In some embodiments, the nucleotide sequence encoding the night bat GS has at least 85% identity with SEQ ID NO: 17. In some embodiments, the nucleotide sequence encoding the night bat GS has at least 90% identity with SEQ ID NO: 17. In some embodiments, the nucleotide sequence encoding the night bat GS has at least 95% identity with SEQ ID NO: 17. In some embodiments, the nucleotide sequence encoding the night bat GS is identical to SEQ ID NO: 17. In some embodiments, the nucleotide sequence encoding the night bat GS is optimized for expression via specific host cells, such as CHO cells.

[0153] This document provides a DNA vector comprising a nucleotide sequence encoding a GS (GS-encoding sequence), wherein the GS may have an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, or at least 98% identity with SEQ ID NO: 9, having at least one amino acid at a position selected from the following positions. Replacements: A2, H8, N10, G12, Q15, M16, M18, S19, E24, Q27, V33, G39, D48, C49, C53, V54, E55, E56, F68, S70, S72, S80, V82, M84, E92, V97, F98, F102, Q106, K107, P108, E110, H115, T116, K118 , D122, S125, H128, L140, D152, L160, K169, R172, I175, M176, V188, K189, T191, Y194, V19 7. K198, H199, A200, I206, I212, R213, V220, K230, V234, A236, T237, S240, T260, E264, N2 65, H269, K271, E272, A273, K276, S278, R282, L294, H304, K305, N308, N310, D311, D318, S320, T328, Q331, E332, K334, F337, A339, C341, F349, I355, V356, D366, Q367, and Q370, as designated by SEQ ID NO: 9. In some embodiments, the amino acid sequence of the GS provided herein may comprise SEQ ID NO: 9 having about 1, about 3, about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, or about 80 amino acid substitutions.In some embodiments, the amino acid substitutions are selected from A2S, H8Q, N10S, G12A, G12N, G12T, Q15H, M16T, M18L, S19A, S19K, S19N, E24N, E24D, Q27L, V33I, G39H, G39F, D48E, C49H, C49S, C49R, C53S, C53N, V54I, V54L, E55D, E56D, F68Y, F68M, S70A, S72G, S80V, S8 0R, S80T, S80I, V82A, V82S, M84L, E92D, V97I, F98L, F102L, Q106R, Q106H, K107Q, P108S, E110D, H115S, H115Y, T 116S, K118R, D122E, S125T, H128K, H128N, H128R, L140M, D152N, L160P, K169R, R172G, I175V, M176A, M176V, V18 8I, K189N, T191A, T191V, T191G, Y194N, V197A, K198M, H199P, A200S, I206V, I212V, R213E, R213D, V220I, K230E , V234M, A236V, T237S, S240P, T260A, T260S, T260N, T260E, E264D, N265G, H269F, H269Q, H269Y, K271E, E272D, A 273S, K276R, S278G, R282K, R282Q, L294Q, H304N, K305E, N308S, N310H, D311E, D318N, S320G, T328N, T328M, T328S, T328A, Q331H, E332D, E332V, K334R, F337L, A339D, C341R, F349Y, I355L, V356I, D366E, Q367E and Q370E, numbered according to SEQ ID NO: 9.

[0154] In some embodiments, the DNA vector provided herein comprises a nucleotide sequence encoding GS having an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, or 98% identity with SEQ ID NO: 9, and having at least one amino acid substitution at the following positions: N10, G12, Q15, S19, E24, V33, G39, C49, C53, V54, E56, S72, S80, V82. E92, F98, Q106, K107, P108, K118, L140, D152, L160, R172, M176, T191, Y194, K198, H199, I206, R213, V220, K230, A236, T237, S240, T260, E264, N265, H269, K271, R282, K305, N310, D311, D318, S320, T328, Q331, A339, C341, F349, I355, Q367 and Q370.In some embodiments, the amino acid sequence of the GS provided herein may include SEQ ID NO(s) having about 3, about 5, about 10, about 15, about 20, about 25, about 30, about 35, or about 40 amino acid substitutions at the following positions. NO: 9: N10S, G12A, G12N, G12T, Q15H, S19A, S19K, S19N, E24N, E24D, V33I , G39H, G39F, C49H, C49S, C49R, C53S, C53N, V54I, V54L, E56D, S72G, S80 V, S80R, S80T, S80I, V82A, V82S, E92D, F98L, Q106R, Q106H, K107Q, P108 S, K118R, L140M, D152N, L160P, R172G, M176A, M176V, T191A, T191V, T19 1G, Y194N, K198M, H199P, I206V, R213E, R213D, V220I, K230E, A236V, T237S, S240P, T260A, T260S, T260N, T260E, E264D, N265G, H269F, H269Q, H269Y, K271E, R282K, R282Q, K305E, N310H, D311E, D318N, S320G, T328N, T328M, T328S, T328A, Q331H, A339D, C341R, F349Y, I355L, Q367E and Q370E. In some embodiments, the amino acid sequence of the GS provided herein may include SEQ ID NO: 9 with about one amino acid substitution. In some embodiments, the amino acid sequence of the GS provided herein may include SEQ ID NO: 9 with about three amino acid substitutions. In some embodiments, the amino acid sequence of the GS provided herein may include SEQ ID NO: 9 with about five amino acid substitutions. In some embodiments, the amino acid sequence of the GS provided herein may include SEQ ID NO: 9 with about ten amino acid substitutions. In some embodiments, the amino acid sequence of the GS provided herein may include SEQ ID NO: 9 with about fifteen amino acid substitutions. In some embodiments, the amino acid sequence of the GS provided herein may include SEQ ID NO: 9 with about twenty amino acid substitutions.

[0155] In some embodiments, the vectors provided herein are suitable for producing recombinant POIs. The vectors provided herein contain any GS coding sequence disclosed herein and an expression cassette. In some embodiments, the expression cassette contains an MCS, which can be used to insert a target gene encoding the POI for recombinant production. In some embodiments, the expression cassette on the vectors provided herein may contain a nucleotide sequence encoding a POI (POI coding sequence).

[0156] As used interchangeably herein and as understood in the art, “polypeptide” or “peptide” refers to a polymer of amino acids of any length linked together by peptide bonds. It may include non-natural or modified amino acids or may be broken down by non-amino acid components. A “protein” contains one or more polypeptides. In globular proteins, such as enzymes, the polypeptide chain of amino acids folds into a three-dimensional functional shape or tertiary structure at least partially through disulfide bonds (SS) with other amino acids in the same polypeptide. Other interactions, such as hydrogen bonds, ionic bonds, covalent bonds, and hydrophobic interactions, also contribute to the tertiary structure. In some proteins, such as antibody molecules and hemoglobin, several polypeptides are bonded together to form a quaternary structure. Polypeptides, peptides, or proteins can also be modified by, for example, glycosylation, lipolysis, acetylation, phosphorylation, or any other manipulation or modification.

[0157] The vectors provided herein include single-gene vectors, dual-gene vectors, and multi-gene vectors. In some embodiments, a POI expressed by a vector provided herein may contain one or more copies of the same polypeptide. In some embodiments, a single-gene vector containing a single nucleotide sequence encoding a polypeptide is provided herein. In some embodiments, a POI expressed by a vector provided herein may contain two or more different polypeptides. In some embodiments, the two or more different polypeptides are encoded by different nucleotide sequences on different vectors. Different vectors can be co-introduced into the same host cell for recombinant production. In some embodiments, the different polypeptides forming the POI are encoded by different nucleotide sequences on the same vector. In some embodiments, a dual-gene vector containing two nucleotide sequences, each encoding a polypeptide, is provided herein. In some embodiments, a multi-gene vector containing multiple nucleotide sequences, each encoding a target polypeptide, is provided herein.

[0158] In some embodiments, the vectors provided herein are dual-gene or multi-gene vectors containing two or more nucleotide sequences encoding two or more polypeptides that are part of the same POI. In some embodiments, the vectors provided herein are dual-gene or multi-gene vectors containing two or more nucleotide sequences encoding two or more polypeptides that form more than one POI. The two or more nucleotide sequences encoding the target polypeptide can be placed in one or more expression cassettes. Accordingly, in some embodiments, the vectors provided herein may have a single bicistronic, tricistronic, or polycistronic expression cassette, wherein all encoding nucleotide sequences are operatively linked to a common expression control sequence. In some embodiments, the vectors provided herein may have two or more expression cassettes, wherein the encoding nucleotide sequences are placed under the control of expression control sequences in different expression cassettes.

[0159] For example, in some embodiments, the vector provided herein comprises an expression cassette containing a first nucleotide sequence encoding a first polypeptide and a second nucleotide sequence encoding a second polypeptide. The nucleotide sequences may be linked by a separator element. In some embodiments, the vector provided herein comprises a first expression cassette containing a first nucleotide sequence and a second expression cassette containing a second nucleotide sequence. In some embodiments, the vector provided herein comprises a bicistronic expression cassette containing first and second nucleotide sequences. The nucleotide sequences may be linked by a separator element.

[0160] The separator element included in the vectors disclosed herein may be, for example, an IRES or a 2A element. In some embodiments, the vectors provided herein contain nucleotides encoding a 2A self-cleaving peptide. In some embodiments, the GS coding sequence and the POI coding sequence are linked by a 2A coding sequence. In some embodiments, the POI coding sequence is linked by a 2A coding sequence. Illustrative 2A self-cleaving peptides include P2A, E2A, F2A, and T2A. In some embodiments, the vectors provided herein contain an IRES. In some embodiments, the GS coding sequence and the POI coding sequence are linked by an IRES. In some embodiments, the POI coding sequence is linked by an IRES.

[0161] For example, in some embodiments, this document provides expression vectors for the recombinant production of antibodies having a light chain and a heavy chain, wherein the vector comprises a first nucleotide sequence encoding the light chain and a second nucleotide sequence encoding the heavy chain. In some embodiments, the first and second nucleotide sequences are placed in the same expression cassette, separated by a 2A coding sequence or an IRES. In some embodiments, the first and second nucleotide sequences are placed in first and second expression cassettes, respectively.

[0162] The following sections further describe POIs that can be expressed by the vectors disclosed herein. For example, POIs may be selected from antibodies, enzymes, soluble proteins, secreted proteins, membrane proteins, and fusion proteins. In some embodiments, the POI is an antibody selected from IgA antibodies, IgM antibodies, IgG1 antibodies, IgG2 antibodies, IgG3 antibodies, IgG4 antibodies, Fab, Fab', F(ab')2, Fv, scFv, (scFv)2, single-domain antibodies (sdAb), single-chain antibodies (scAb), and heavy-chain antibodies (HCAb). In some embodiments, the POI is an antibody selected from monoclonal antibodies, bispecific antibodies, multispecific antibodies, bivalent antibodies, and multivalent antibodies.

[0163] In some embodiments, the vectors provided herein are suitable for genome integration. In some embodiments, the vectors provided herein are suitable for recombinant protein production. In some embodiments, expression vectors comprising a GS coding sequence and an expression cassette are provided herein. A wide variety of expression vectors can be used. Examples of vectors are plasmids, autonomously replicating sequences, and transposons. Exemplary vectors also include, without limitation, plasmids, bacteriophages, phage particles, fosmids, artificial chromosomes such as yeast artificial chromosomes (YAC), bacterial artificial chromosomes (BAC), P1-derived artificial chromosomes (PAC), mammalian artificial chromosomes (MAC), bacteriophages such as λ phage or M13 phage, and animal viruses.

[0164] Examples of animal viruses useful as vectors include, without limitation, retroviruses (including lentiviruses), adenoviruses, adeno-associated viruses, cytomegaloviruses, herpesviruses (e.g., herpes simplex virus), poxviruses, baculoviruses, bovine papillomaviruses, papillomaviruses, and multifocal papillomaviruses (e.g., SV40). In some embodiments, the expression vectors provided herein are adeno-associated virus (AAV) vectors, lentiviral vectors, retroviral vectors, replicative adenovirus vectors, replication-defective adenovirus vectors, herpesvirus vectors, and baculovirus vectors. In some embodiments, the vector is an adenoviral vector. In some embodiments, the vector is a retroviral vector. In some embodiments, the vector is an adeno-associated virus vector. Examples of expression vectors are the pClneo vector (Promega) for expression in mammalian cells; and pLenti4 / V5-DEST for lentivirus-mediated gene transfer and expression in mammalian cells. TM pLenti6 / V5-DEST TMand pLenti6.2 / V5-GW / lacZ (Invitrogen). Exemplary transposable subsystems such as Sleeping Beauty and PiggyBac can also be used. (Ivics et al., Cell, 91(4):501–510(1997); et al., (2007) Nucleic Acids Research. 35(12):e87.

[0165] In some embodiments, the vector is a cell-free gene vector or an extrachromosomally maintained vector. As used herein, the term "cell-free gene" refers to a vector that can replicate but does not integrate into the host chromosomal DNA and is not gradually lost from dividing host cells, and also indicates extrachromosomal or cell-free gene replication of the vector.

[0166] In some embodiments, the vectors provided herein are engineered to have a sequence encoding a DNA replication origin, or "ori," from a lymphotropic herpesvirus or gamma herpesvirus, adenovirus, SV40, bovine papillomavirus, or yeast, specifically corresponding to the oriP of EBV as the replication origin of a lymphotropic herpesvirus or gamma herpesvirus. In some embodiments, the lymphotropic herpesvirus may be Epstein-Barr virus (EBV), Kaposi's sarcoma herpesvirus (KSHV), squirrel monkey herpesvirus (HS), or Marek's disease virus (MDV). Epstein-Barr virus (EBV) and Kaposi's sarcoma herpesvirus (KSHV) are also examples of gamma herpesviruses. Typically, the host cell contains a viral replication transactivator protein that activates replication.

[0167] The “expression control sequences,” “control elements,” or “regulatory sequences” present in expression vectors are those vector untranslated regions—origins of replication, selection cassettes, promoters, enhancers, translation initiation signal (ribosome binding site sequence or Kozak sequence) introns, polyadenylated sequences, and the 5' and 3' untranslated regions—that interact with host cell proteins to carry out transcription and translation. The strength and specificity of these elements can vary. Depending on the vector system and host used, any number of suitable transcriptional and translational elements can be used, including ubiquitous promoters and inducible promoters.

[0168] Mammalian expression vectors may contain non-transcriptional elements such as origin of replication, suitable promoters and enhancers linked to the gene to be expressed, and other 5' or 3' flanked non-transcriptional sequences, and 5' or 3' non-translational sequences such as essential ribosome binding sites, polyadenylation sites, splice donor and acceptor sites, and transcription termination sequences. Expression of recombinant proteins in insect cell culture systems (e.g., baculoviruses) also provides a robust method for producing correctly folded and biologically functional proteins. Baculovirus systems for producing heterologous proteins in insect cells are well known to those skilled in the art.

[0169] In some embodiments, the vectors provided herein include a promoter operatively linked to a GS coding sequence. As understood in the art, a promoter is a nucleotide sequence that defines where transcription of a gene by an RNA polymerase begins. Promoter sequences are typically located directly upstream of the transcription start site or at the 5' end. DNA regions are operatively linked when they are functionally related to each other. Structures in operatively linked nucleotide sequences are capable of performing or characterized by performing the desired operation. Those skilled in the art will recognize that elements or structures in a nucleic acid sequence do not necessarily have to be in a tandem or adjacent order for operative linking. For example, if a promoter controls transcription of a sequence, then that promoter is operatively linked to the coding sequence; or if the location of the ribosome binding site allows translation, then it is operatively linked to the coding sequence.

[0170] In some embodiments, this document provides an expression vector comprising a GS encoding sequence and an expression cassette, wherein the GS encoding sequence is operatively linked to a promoter. The promoter may be any promoter described herein or otherwise known in the art. In some embodiments, the promoter is a CMV promoter. In some embodiments, the promoter is an SV40 promoter.

[0171] As described above, the vectors described herein can be used to recombinantly produce POIs. In some embodiments, the vectors provided herein comprise a GS coding sequence and an expression cassette comprising a POI coding sequence. The expression cassette comprises a promoter operatively linked to the POI coding sequence. The promoter may be any promoter disclosed herein or otherwise known in the art. In some embodiments, the promoter is the same as the promoter operatively linked to the GS coding sequence. In some embodiments, the promoter is different from the promoter operatively linked to the GS coding sequence.

[0172] In some embodiments, the expression cassette comprises two or more nucleotide sequences encoding two or more polypeptides, wherein the two or more nucleotide sequences are operatively linked to the same promoter. In some embodiments, the vector provided herein comprises two or more expression cassettes, each expression cassette comprising a promoter operatively linked to a nucleotide sequence encoding a polypeptide. In some embodiments, the promoters in different expression cassettes are the same promoter. In some embodiments, the promoters in different expression cassettes are different. The promoter can be any promoter disclosed herein or otherwise known in the art.

[0173] The promoter can be a forward promoter or a reverse promoter. In some embodiments, the promoter is a mammalian promoter. In some embodiments, one or more promoters are natural promoters. In some embodiments, one or more promoters are non-natural promoters. In some embodiments, one or more promoters are non-mammal promoters.

[0174] Universally prevalent promoters that can be used in this illustrative disclosure include (but are not limited to) cytomegalovirus (CMV) promoters, simian virus 40 (SV40) promoters (e.g., early or late), Moloney murine leukemia virus (MoMLV) LTR promoters, Raul's sarcoma virus (RSV) LTR promoters, herpes simplex virus (HSV) (thymidine kinase) promoters, spleen lesion-forming virus (SFFV) promoters, U1 promoters, U6 promoters, H1 promoters, H5 promoters, P7.5 promoters, P11 promoters, and T7 promoters. Motion promoter, Sp6 promoter, elongation factor 1-α (EF1α) promoter, early growth response 1 (EGR1), ferritin H (FerH), ferritin L (FerL), glyceraldehyde-3-phosphate dehydrogenase (GAPDH), eukaryotic translation initiation factor 4A1 (EIF4A1), heat shock protein 70 kDa (HSPA5), heat shock protein 90 kDa β, member 1 (HSP90B1), heat shock protein 70 kDa (HSP70), β-kinin (β-KIN), lac promoter, human ROSA The following promoters are used: 26 loci (Irions et al., Nature Biotechnology 25, 1477-82 (2007)), upstream activation sequence (UAS) promoter, ubiquitin C promoter (UBC), phosphoglycerate kinase-1 (PGK) promoter, cytomegalovirus enhancer / chicken β-actin (CAG) promoter, tetracycline response element (TRE), araC promoter, araBAD promoter, tryptophan (trp) promoter, Ptac promoter, and β-actin promoter. In some embodiments, the CMV promoter is used. In some embodiments, the SV40 promoter is used. In some embodiments, the EF1α promoter is used.

[0175] In some embodiments, the promoter is constitutively driven expression; that is, the promoter is a constitutive promoter. In some embodiments, the promoter is an inducible promoter. Inducible promoters are not limited and can be any inducible promoter known in the art. In some embodiments, the presence of one or more environmental or chemical stimuli promotes the expression of the inducible promoter. For example, in some embodiments, the inducible promoter drives expression in the presence of: chemical molecules, such as tetracyclines and their derivatives (e.g., doxycycline), cumate and their derivatives; or environmental stimuli, such as heat or light.

[0176] In some embodiments, the inducible promoter is based on a tetracycline-controlled transcriptional activation system, a cumate repressor system, a lac repressor system, an arabinose-regulated pBad promoter system, an alcohol-regulated AlcA promoter system, a steroid-regulated LexA promoter system, a heat shock-induced Hsp70 or Hsp90 promoter system, or a blue light-induced pR promoter system. Therefore, in some embodiments, the inducible promoter comprises a nucleic acid sequence that binds to a tetracycline transactivator, such as a tetracycline-responsive element. In some embodiments, expression of the inducible promoter is turned on in the presence of tetracycline or its derivatives (Tet-on system), while in other embodiments, expression of the inducible promoter is turned off in the presence of tetracycline or its derivatives (Tet-off system). In some embodiments, the inducible promoter is based on a cumate repressor system. Therefore, in some embodiments, the inducible promoter comprises a nucleic acid sequence that binds to a CymR repressor, such as a cumate operon sequence.

[0177] In some embodiments, the expression of an inducible promoter is driven by the dimerization of a transcription factor. In some embodiments, transcription is performed by bacterial EL222, which dimers in the presence of blue light to drive expression from the C120 promoter or its regulatory elements. In some embodiments, the inducible promoter comprises a nucleic acid sequence derived from the C120 promoter or regulatory element. Illustrative examples of inducible promoters / systems also include (but are not limited to) steroid-inducible promoters, such as promoters for genes encoding glucocorticoids or estrogen receptors (inducible by treatment with the corresponding hormone), metallothionein promoters (inducible by treatment with various heavy metals), MX-1 promoters (inducible by interferon), the “GeneSwitch” mifepristone-regulated system (Sirin et al., 2003, Gene, 323:67), cumate-inducible gene switches (WO 2002 / 088346), tetracycline-dependent regulatory systems, etc.

[0178] In some embodiments, this document provides vectors suitable for the expression of recombinant proteins in mammalian cells (e.g., CHO cells), comprising: a nucleotide sequence operably linked to a promoter encoding the koala, ostrich, snake, pigeon, junglefowl, gerbil, or night bat GS disclosed herein, and an expression cassette comprising a nucleotide sequence encoding a POI operably linked to the promoter. The promoter may be any promoter disclosed herein or otherwise known in the art. In some embodiments, the promoter is the same. In some embodiments, the promoter is different. In some embodiments, the POI is an antibody comprising both a heavy chain and a light chain. Therefore, in some embodiments, this document provides vectors suitable for the expression of recombinant antibodies in mammalian cells (e.g., CHO cells), comprising: a nucleotide sequence operably linked to a promoter encoding the koala, ostrich, snake, pigeon, junglefowl, gerbil, or night bat GS disclosed herein, a first expression cassette comprising a nucleotide sequence operably linked to the promoter encoding a light chain, and a second expression cassette comprising a nucleotide sequence operably linked to the promoter encoding a heavy chain. In some embodiments, this document provides a vector suitable for the expression of recombinant antibodies in mammalian cells (e.g., CHO cells), comprising: a nucleotide sequence operably linked to a promoter encoding the koala, ostrich, snake, pigeon, junglefowl, gerbil, or night bat GS disclosed herein, and an expression cassette comprising a promoter operably linked to a first nucleotide sequence encoding a light chain and a second nucleotide sequence encoding a heavy chain. The promoter may be any promoter disclosed herein or otherwise known in the art, such as the CMV promoter and the SV40 promoter. In some embodiments, the promoter is the same. In some embodiments, the promoter is different.

[0179] For expression in mammalian cells, most mammalian mRNAs have a polyadenylated 3' end or are linked to multiple adenines forming a poly(A) tail. In some embodiments, the vectors provided herein contain one or more polyadenylated signals. When more than one polyadenylated signal is provided, the polyadenylated signal may be located at the 3' end of each set of nucleotide sequences encoding the polypeptide. Thus, in one example, the vectors provided herein contain a polyadenylated signal at the 3' end of each of at least one POI coding sequence and a polyadenylated signal at the 3' end of the GS coding sequence. In some embodiments, a polyadenylated signal may be provided at the 3' end of each expression cassette. For illustrative purposes, the vectors provided herein may contain nucleotide sequences encoding the heavy and light chains of an antibody, wherein the heavy chain coding sequence and the light chain coding sequence are placed in the same expression cassette. In some embodiments, these vectors may contain a polyadenylated signal at the 3' end of the expression cassette and a polyadenylated signal at the 3' end of the GS coding sequence. In some embodiments, these vectors may include a polyadenylation signal at the 3' end of the heavy chain coding sequence, a polyadenylation signal at the 3' end of the light chain coding sequence, and a polyadenylation signal at the 3' end of the GS coding sequence. 6.3 Host Cell

[0180] This document provides host cells comprising the vectors described herein. The vectors may be any vectors disclosed herein. In some embodiments, this document provides host cells comprising a vector having the GS coding sequence disclosed herein and an expression cassette having either an MCS or a POI coding sequence having a nucleotide sequence for inserting a POI (POI coding sequence). In some embodiments, this document provides host cells comprising a POI coding sequence inserted at a transcriptionally active locus generated using the optional markers described herein or the methods described herein. In some embodiments, this document provides stable cell lines comprising the vectors described herein.

[0181] In some embodiments, the host cells provided herein can be grown in a glutamine-free culture medium. In some embodiments, the host cells provided herein can be grown in the presence of a GS inhibitor. The GS inhibitor can be any GS inhibitor disclosed herein or otherwise known in the art, including MSX and its derivatives, phosphorus-containing analogs of glutamate, and bisphosphonates. In some embodiments, the GS inhibitor is MSX or a derivative thereof. In some embodiments, the GS inhibitor is a phosphorus-containing analog of glutamate. In some embodiments, the GS inhibitor is a bisphosphonate. GS inhibitors can be added at different concentrations to produce different levels of selectivity.

[0182] In some embodiments, the host cells provided herein can be grown in a culture medium supplemented with about 1-10 μM, about 10-50 μM, about 50-100 μM, or about 100-300 μM of MSX. In some embodiments, the host cells provided herein can be grown in a culture medium supplemented with about 1 μM, about 3 μM, about 5 μM, about 10 μM, about 25 μM, about 50 μM, about 75 μM, about 100 μM, about 150 μM, about 200 μM, about 250 μM, or about 300 μM of MSX. In some embodiments, the host cells provided herein can be grown in a culture medium supplemented with about 5 μM of MSX. In some embodiments, the host cells provided herein can be grown in a culture medium supplemented with about 10 μM of MSX. In some embodiments, the host cells provided herein can be grown in a culture medium supplemented with about 50 μM of MSX. In some embodiments, the host cells provided herein can be grown in a culture medium supplemented with about 100 μM of MSX.

[0183] The host cell described herein may contain a vector. In some embodiments, the host cell may contain a vector containing the GS coding sequence disclosed herein and an expression cassette. As described above, an expression cassette may contain one or more nucleotide sequences encoding one or more target polypeptides. In some embodiments, the host cell described herein may contain a vector containing the GS coding sequence disclosed herein and two or more expression cassettes. The host cell may contain multiple copies of the same vector. In some embodiments, the host cell may contain two different vectors, each containing the GS coding sequence disclosed herein and at least one expression cassette. In some embodiments, the host cell may contain multiple different vectors, each containing the GS coding sequence disclosed herein and at least one expression cassette.

[0184] In some embodiments, the host cells provided herein may possess wild-type endogenous GS. In some embodiments, the endogenous GS of the host cells has a mutation that results in reduced activity. In some embodiments, the endogenous GS of the host cells is inactivated. In some embodiments, the endogenous GS of the host cells is knocked out. Cells in which the GS gene is knocked out lose their endogenous GS activity and cannot grow in a glutamine-free environment without the insertion or incorporation of an exogenous GS gene from (for example) a vector described herein.

[0185] The host cells provided in this article may be eukaryotic cell lines, such as yeast cell lines (e.g., Saccharomyces cerevisiae or Yarrowia lipolytica cell lines), fungal cell lines (e.g., Aspergillus niger cell lines), insect cell lines (e.g., Spodoptera fugiperda cell lines, such as Sf9), or mammalian cell lines. Examples of suitable mammalian host cell lines include (but are not limited to) COS-7 (monkey kidney), L-929 (mouse fibroblast), C127 (mouse mammary tumor), NS0 (non-secreting mouse myeloma), SP2 / 0 (mouse myeloma), 3T3 (mouse fibroblast), CHO (Chinese hamster ovary), HeLa (human cervical cancer), BHK (hamster kidney fibroblast), HEK-293 (human embryonic kidney) cell lines (e.g., HEK293-F, HEK293-H, HEK293-T), PERC.6 (human embryonic retinal cell), HROC277 (human colorectal adenocarcinoma cell), VERO (African green monkey kidney), MDCK (canine kidney), WI38 (human lung fibroblast), V79 (Chinese hamster lung), BHK (young hamster kidney fibroblast), and their variants.

[0186] In some embodiments, the host cell provided herein is COS-7 cell. In some embodiments, the host cell provided herein is L-929 cell. In some embodiments, the host cell provided herein is C127 cell. In some embodiments, the host cell provided herein is NSO cell. In some embodiments, the host cell provided herein is SP2 / 0 cell. In some embodiments, the host cell provided herein is 3T3 cell. In some embodiments, the host cell provided herein is CHO cell. In some embodiments, the host cell provided herein is HeLa cell. In some embodiments, the host cell provided herein is BHK cell. In some embodiments, the host cell provided herein is HEK-293 cell. In some embodiments, the host cell provided herein is PERC.6 cell. In some embodiments, the host cell provided herein is HROC277 cell. In some embodiments, the host cell provided herein is VERO cell. In some embodiments, the host cell provided herein is MDCK cell. In some embodiments, the host cell provided herein is WI38 cell. In some embodiments, the host cell provided herein is V79 cell. In some embodiments, the host cell provided herein is BHK cell.

[0187] In some embodiments, the host cells provided herein are suitable for post-translational modifications (“PTMs”) in POIs, such as glycosylation, phosphorylation, and disulfide bonding.

[0188] In some embodiments, the cell line is a stable cell line. As understood in the art, a stable cell line refers to a cell clone containing a vector described herein capable of continuously expressing a POI. In some embodiments, any one or more vectors disclosed herein are transiently introduced into the cell. In some embodiments, a host cell for expressing a POI is provided herein, wherein the cell contains a vector disclosed herein, the vector containing a GS coding sequence and an expression cassette containing a POI coding sequence. In some embodiments, the vector is transiently introduced into the host cell without integration into the cell's genome. In some embodiments, the coding sequence of the vector is stably integrated into the cell's genome.

[0189] In some embodiments, the host cell is a eukaryotic cell. In some embodiments, the host cell is a mammalian cell. In some embodiments, the host cell is a Chinese hamster ovary (CHO) cell. More than half of the approved and currently marketed therapeutic proteins are produced using CHO cells, primarily due to the unique characteristics of CHO cells, such as human-like post-translational modifications of the product and their adaptability to bioprocess development and large-scale production. In some embodiments, this document provides CHO cells comprising the vectors provided herein. The vector can be any vector described herein.

[0190] In some embodiments, the host cell included in the methods described above and disclosed herein is a CHO cell. In some embodiments, the CHO cells have wild-type endogenous GS. Exemplary CHO cell lines having wild-type GS may be, for example, CHO-S, CHO-K1, CHOK1SV, CHOZN K1, or FreeStyle CHO-S. In some embodiments, CHO cells having wild-type GS may have mutations in genes other than GS that reduce or eliminate the enzymatic activity of proteins involved in, for example, glycosylation or metabolic pathways. The mutations may be naturally occurring mutations or genetically engineered mutations. Exemplary CHO cell lines having impaired or inactivated glycosylation enzymes (e.g., fucosyltransferase 8) include CHO FUT8 KO. Exemplary CHO cell lines having impaired or inactivated enzymes in metabolic pathways (e.g., dihydrofolate reductase, DHFR) include CHO-DG44, CHO-DUXB11, and CHO-DUKX. In some embodiments, CHO cells with wild-type GS may have enhanced enzymatic activity in genes other than GS, the products of which are involved in cell growth / survival, metabolism, and protein modification. In some embodiments, CHO cells have amplified genes encoding anti-apoptotic proteins. Gene amplification may be caused by naturally occurring mutations or genetic engineering alterations. In some embodiments, CHO cells have exogenous genes encoding anti-apoptotic proteins.

[0191] In some embodiments, the host cell provided herein is a CHO-S cell. In some embodiments, the host cell provided herein is a CHO-K1 cell. In some embodiments, the host cell provided herein is a CHOK1SV cell. In some embodiments, the host cell provided herein is a CHOZNK1 cell. In some embodiments, the host cell provided herein is a FreeStyle CHO-S cell. In some embodiments, the host cell provided herein is a CHO-DG44 cell. In some embodiments, the host cell provided herein is a CHO-DUXB11 cell. In some embodiments, the host cell provided herein is a CHO-DUKX cell.

[0192] In some embodiments, the endogenous GS of the CHO cells provided herein has reduced activity compared to wild-type hamster GS. In some embodiments, the endogenous GS of the CHO cells provided herein is inactivated. In some embodiments, the endogenous GS of the CHO cells provided herein is knocked out. Exemplary CHO cell lines with endogenous GS knockout may be, for example, CHOK1SV GS-KO, CHOZN GS- / -, or CHOZN GS KO. In some embodiments, the host cell provided herein is CHOK1SV GS-KO cells. In some embodiments, the host cell provided herein is CHOZN GS- / - cells. In some embodiments, the host cell provided herein is CHOZN GS KO cells.

[0193] The CHO cells described in this article are commercially available from manufacturers such as Lonza, ECACC, Sigma-Aldrich / Merck, Fisher, and Horizon. Most CHO lines used for recombinant protein production were originally derived from CHO K1 (Kao & Puck, 1968) and are therefore very genetically similar.

[0194] In some embodiments, the CHO cells provided herein can be grown in glutamine-free medium. In some embodiments, the CHO cells provided herein can be grown in the presence of a GS inhibitor. In some embodiments, the CHO cells provided herein can be grown in glutamine-free medium in the presence of a GS inhibitor. In some embodiments, the CHO cells provided herein can be selected based on their ability to grow in glutamine-free medium. In some embodiments, the CHO cells provided herein can be selected based on their ability to grow in the presence of a GS inhibitor. In some embodiments, the CHO cells provided herein can be selected based on their ability to grow in glutamine-free medium in the presence of a GS inhibitor. The GS inhibitor can be any GS inhibitor disclosed herein or otherwise known in the art, including MSX and its derivatives, phosphorus-containing analogs of glutamate, and bisphosphates. In some embodiments, the CHO cells provided herein can be grown in glutamine-free medium supplemented with MSX. In some embodiments, the CHO cells provided herein can be selected based on their ability to grow in glutamine-free medium supplemented with MSX. Different concentrations of GS inhibitors can be supplemented to produce different levels of selection strictness. In some embodiments, approximately 1-10 μM, approximately 10-50 μM, approximately 50-100 μM, or approximately 100-300 μM of MSX are supplemented. In some embodiments, approximately 1 μM, approximately 3 μM, approximately 5 μM, approximately 10 μM, approximately 25 μM, approximately 50 μM, approximately 75 μM, approximately 100 μM, approximately 150 μM, approximately 200 μM, approximately 250 μM, or approximately 300 μM of MSX are supplemented. In some embodiments, approximately 50 μM of MSX is supplemented. 6.4 Target Protein (POI)

[0195] This document provides the use of GS derived from koalas, ostriches, snakes, pigeons, junglefowl, gerbils, or night bats as alternative markers in the recombinant production of POIs, or in some embodiments, in the recombinant production of target mRNAs. A POI can be any protein to be expressed. In some embodiments, the POI is selected from antibodies, enzymes, soluble proteins, secreted proteins, membrane proteins, or fusion proteins. In some embodiments, the antibody can be a therapeutic protein. In some embodiments, the POI is an antibody. In some embodiments, the POI is an enzyme. In some embodiments, the enzyme can be used for enzyme replacement therapy. In some embodiments, the POI is a soluble protein. In some embodiments, the POI is a secreted protein. In some embodiments, the POI is a membrane protein. In some embodiments, the POI is a fusion protein.

[0196] In some embodiments, expression of POI can cause cytotoxicity when expressed in a reference expression system. In some embodiments, POI is a protein expressed in low yield in a conventional expression system. In some embodiments, protein expression or quality is significantly improved by expression according to the disclosed methods, for example, using koala, ostrich, snake, pigeon, junglefowl, gerbil, or night bat GS as optional markers.

[0197] In some embodiments, the POI is a human protein. In some embodiments, the POI is a mammalian protein.

[0198] In some embodiments, the POI is a monomer, i.e., composed of a single polypeptide. In some embodiments, the POI comprises two or more copies of the same polypeptide. In some embodiments, the POI is a polymer and comprises at least two different polypeptides. In some embodiments, the POI comprises non-natural amino acids. In some embodiments, the POIs provided herein include post-translational modifications (“PTMs”), such as glycosylation, phosphorylation, ubiquitination, nitrosation, methylation, acetylation, esterification, etc.

[0199] In some embodiments, the POI is an antibody. Exemplary antibodies that can be produced by the compositions and methods disclosed herein include intact monoclonal antibodies, single-domain antibodies (sdAbs; e.g., camel antibodies, alpaca antibodies), single-chain Fv (scFv) antibodies, heavy chain antibodies (HCAbs), light chain antibodies (LCAbs), multispecific antibodies, bispecific antibodies, monospecific antibodies, monovalent antibodies, and any other modified immunoglobulin molecules containing an antigen-binding site (e.g., dual variable domain immunoglobulin molecules), provided that the antibody exhibits the desired biological activity. Antibodies also include (but are not limited to) mouse antibodies, camel antibodies, chimeric antibodies, humanized antibodies, and human antibodies. Based on the characteristics of their heavy chain constant domains, referred to as α, δ, ε, γ, and μ, respectively, antibodies can be any of the five major classes of immunoglobulins: IgA, IgD, IgE, IgG, and IgM, or their subclasses (isotypes) (e.g., IgG1, IgG2, IgG3, IgG4, IgA1, and IgA2). Unless otherwise expressly stated, the term "antibody" as used herein includes the "antigen-binding fragment" of a complete antibody. The term "antigen-binding fragment" as used herein refers to a portion or fragment of a complete antibody that is the antigen-determining variable region of the complete antibody. Examples of antigen-binding fragments include (but are not limited to) Fab, Fab', F(ab')2, Fv, linear antibodies, single-chain antibody molecules (e.g., scFv), heavy chain antibodies (HCAb), light chain antibodies (LCAb), disulfide-linked scFv (dsscFv), double-chain antibodies, triple-chain antibodies, tetra-chain antibodies, microantibodies, dual variable domain antibodies (DVD), single variable domain antibodies (sdAb; e.g., camel antibodies, alpaca antibodies), and single variable domain (VHH) of heavy chain antibodies, as well as bispecific or multispecific antibodies formed from antigen fragments.

[0200] Some antibodies contain at least one heavy chain and one light chain. When used to refer to antibodies, the term "heavy chain" refers to a polypeptide chain of approximately 50-70 kDa, in which the amino-terminal portion includes a variable region of approximately 120 to 130 or more amino acids and the carboxyl-terminal portion includes a constant region. Based on the amino acid sequence of the heavy chain's constant region, the constant region can be one of five different types, referred to as α, δ, ε, γ, and μ. The different heavy chains have different sizes: α, δ, and γ contain approximately 450 amino acids, while μ and ε contain approximately 550 amino acids. When combined with a light chain, these different types of heavy chains produce five well-known antibody classes: IgA, IgD, IgE, IgG, and IgM, which include four IgG subclasses: IgG1, IgG2, IgG3, and IgG4. The heavy chain can be a human heavy chain. When used to refer to antibodies, the term "light chain" refers to a polypeptide chain of approximately 25 kDa, wherein the amino-terminal portion comprises a variable region of approximately 100 to approximately 110 or more amino acids and the carboxyl-terminal portion comprises a constant region. The approximate length of a light chain is 211 to 217 amino acids. Based on the amino acid sequence of the constant domain, two distinct types exist, referred to as κ or λ. The amino acid sequences of light chains are well known in the art. Light chains can be human light chains.

[0201] In some embodiments, the POI is a therapeutic protein. In some embodiments, the POI may be useful in the production of clinical test kits or other diagnostic assays.

[0202] In some embodiments, the compositions and methods of the present invention are used to produce therapeutic proteins. In some embodiments, the POI is selected from abariciac, abatacept, abciximab, adalimumab, aflibercept, galactosidase β, abiglutide, interleukin, afacilipes, alenzumab, vidarcephalin, aglucosidase alfa, alicurumab, aliskiren, α-1-protease inhibitor, alteplase, anapeptidase, acestatin, anipexase, human anthrax immunoglobulin, antihemophilic factor, antithrombin α, human antithrombin III, antithymocyte globulin, antithymocyte globulin (horse), antithymocyte globulin (rabbit), aprotinin, asimomab, asfotase alfa, asparaginase, Erwinia chrysogenum. Chrysanthemi (asparaginase), Atezolizumab, autologous chondrocyte culture, Baliximab, Becaprolem, Belapatine, Belimumab, Belactal, Bevacizumab, Bivalirudin, Bonatumab, Botulinum toxin type A, Botulinum toxin type B, Bentuximab, Brodatumab, Busereline, Cl-esterase inhibitor (human), Cl-esterase inhibitor, Canazine, Canazine, Carotenumab, Pegylated cetolizumab (Certolizumab pegol), Cetuximab, Human chorionic gonadotropin (A), Human chorionic gonadotropin (HCG), Human chorionic gonadotropin, Coagulation factor IX, Coagulation factor Vila, Human coagulation factor X, Coagulation factor XIIIA-subunit, Coagulation enzyme, Alfaconesta (conestat) Alfa, adrenocorticotropic hormone, ticokine, daclizumab, daptomycin, daratumumab, dabepoetin alpha, defibrin polynucleotide, denileukindiftitox, denosumab, disiludin, dartuximab, alfa chain enzyme, trexobin alpha, dulaglutide, eculizumab, efalizumab, emoxetine alpha, elemothrombin alpha, eleothrombin alpha, enfuviride, alfa edoretin, eportin ξ, epitubatide, etanercept, evolocumab, exenatide, factor IX complex (human), fibrinogen concentrate (human) ), plasmin (also known as fibrinolytic enzyme), filgrastim, filgrastim-sndz, follicle-stimulating hormone α, follicle-stimulating hormone β, sulfonase, gastric factor, gemtruzumab, oxazolidin, glatiramer acetate, recombinant glucagon, glutapeptide, golimumab, bacitracin D, hepatitis A vaccine, hepatitis B immunoglobulin, human calcitonin, human tetanus toxoid immunoglobulin, human rabies virus immunoglobulin, human p(D) immunoglobulin, human serum albumin, human varicella-zoster immunoglobulin, hyaluronidase, tiimumab, tiimumab (Ibritumomab)Tiuxetan), Idacillinumab, Iduthioplasma, Imiglucerase, Human Immunoglobulin, Infliximab, Aspart Insulin, Bovine Insulin, Degludec Insulin, Detemir Insulin, Glargine Insulin, Lisglut Insulin, Lispro Insulin, Porcine Insulin, Regular Insulin, Regular Insulin, Porcine Insulin, Protamine Insulin, Interferon α-2a, Interferon α-2b, Interferon Alfacon-1, Interferon α-nl, Interferon α-n9, Interferon β-1a, Interferon β-1b, Interferon γ-1b, Intravenous Immunoglobulin, etc. Lilimumab, Exexumab, Laronibase, Legasil, Lepirudine, Lepranylbutazone, Liraglutide, Lucisin, Luteinizing Hormone Alpha, Mecaserine, Urotropic Gonadotropin, Meporibumab, Beta-epotine, Metriptin, Moromab, Natazumab, Alpha-interferon, Necitumumab, Neciritide, Nivolumab, Oxtoxicumab, Obituzumab, Oxyclothromab ... Panitumumab, Pembrolizumab, Pertuzumab, Porcine lung phospholipid α, Pramlintide, Preotact, Human Protein S, Ramucirumab, Ranibizumab, Raburicase, Recibacubitumab, Reteplase, Linasicept, Rituximab, Romistachytin, Sacoloxetine, Salmon Calcitonin, Saxagstim, Satumuzumab-Pendetide, Lysosomal Acid Lipase α, Secretin, Secukinumab, Sermorelin, Serum Albumin, Iodized Serum Albumin, Secuximab, Cimodococcus α, Prevexil (Sipuleucel-T), recombinant growth hormone, recombinant growth hormone, streptokinase, sulodide, susotropin α, afataglitazone, teduglutide, teicoplanin, tenepase, teriparatide, temorelin, coagulation regulatory protein α, thymosin alpha, thyroglobulin, thyrotropin α, thyrotropin α, tocilizumab, tosimoumab, trastuzumab, purified protein derivative of tuberculin, turoctocog α, follicle-stimulating hormone, urokinase, uterotonic acid, vasopressin, vedotin, and veraglitazone α.

[0203] In some embodiments, the POI is a soluble protein, a secreted protein, or a membrane protein. In some embodiments, the POI is not limited to dopamine receptor 1 (DRD1), cystic fibrosis transmembrane transport regulator (CFTR), Cl-esterase inhibitor (Cl-Inh), IL2-inducible T-cell kinase (ITK), or NAD enzyme. In some embodiments, the NAD enzyme is SARM1. In some embodiments, SARM1 is a deletion variant representing a mature protein.

[0204] In some embodiments, the POI is a membrane protein. Illustrative membrane proteins include ion channels, gap junctions, ionotropic receptors, transport proteins, and integrated membrane proteins such as cell surface receptors (e.g., G protein-coupled receptors (GPCRs), tyrosine kinase receptors, integrins, etc.), and proteins that shuttle between the membrane and cytosol in response to signal transduction (e.g., Ras, Rac, Raf, Ga subunits, arrestin, Src, and other effector proteins). In some embodiments, the POI is a G protein-coupled receptor. In some embodiments, the POI is a 7-(sub)-transmembrane domain receptor, 7TM receptor, heptaspiral receptor, serpentine receptor, or G protein-linked receptor (GPLR). In some embodiments, the POI is a class A, class B, class C, class D, class E, or class F GPCR. In some embodiments, the POI is a class 1, class 2, class 3, class 4, class 5, or class 6 GPCR. In some implementations, the POI is a rhodopsin-like GPCR, a secretin receptor family GPCR, a metabolizing glutamate / pheromone GPCR, a fungal mating pheromone receptor, a cyclic AMP receptor, or a coiled / smooth GPCR.

[0205] POIs expressed using the compositions and methods of the present invention may also include proteins associated with enzyme substitution, such as β-galactosidase, α-galactosidase, imiglucerase, taligulcerase alfa, verasidase alfa, vidarcephalinase, lysosomal acid lipase alfa, laronidase, iduxitase, ileosulfate esterase alfa, thiosulfate esterase, alginate α, factor VIII, C3 inhibitors, and Hurler and Hunter corrective factors. In some embodiments, the POI is a nucleoside enzyme, NAD+ nucleoside enzyme, hydrolase, glycosylation enzyme, glycosylation enzyme that hydrolyzes N-glycosyl compounds, NAD+ glycohydrolase, NAD enzyme, DPN enzyme, DPN hydrolase, NAD hydrolase, pyridine diphosphate nucleoside enzyme, nicotinamide adenine dinucleotide dinucleotide enzyme, NAD glycohydrolase, NAD nucleoside enzyme, or nicotinamide adenine dinucleotide glycohydrolase. In some embodiments, the POI is an enzyme involved in nicotinic acid and nicotinamide metabolism and calcium signaling pathways. In some embodiments, the POI can be a secreted protein, such as Cl-Inh.

[0206] POIs expressed using the compositions and methods of the present invention can also be fusion proteins. In some embodiments, the POI is a fusion protein containing an Fc region. In some embodiments, the POI is a therapeutic protein containing an Fc region. In some embodiments, the POI is afacilipe, etanercept, abatacept, belacept, aflibercept, linalcept, romiplostim, antihemophilic factor-Fc fusion protein, or enothrombin α.

[0207] In some embodiments, this document also provides proteins expressed by introducing the vectors disclosed herein into mammalian cells (e.g., CHO cells). In some embodiments, this document also provides proteins produced by mammalian cells comprising the vectors disclosed herein. 6.5 Expression System and Reagent Kit

[0208] The optional biomarkers provided herein allow for the efficient identification of cell clones with high POI yields and improved POI production. Therefore, this document also provides expression systems for the in vitro production of POIs, comprising the vectors and / or cells provided herein. In some embodiments, the vectors, cells, and systems provided herein are used for the reliable production of POIs that are difficult to express. In some embodiments, the expression system efficiently identifies host cells with high yields and leads to improved POI production (e.g., higher expression rates, higher yields). For illustration, the expression systems provided herein and the POIs produced therefrom can exhibit one or more of the following improvements: reliable production; reduced need for expression optimization; suitability for therapeutic applications; minimal batch-to-batch variability; improved functional activity; and consistent activity.

[0209] In some embodiments, the expression system provided herein comprises the vector and host cell disclosed herein. The vector can be any vector disclosed herein. As provided, the vector in the expression system provided herein may comprise a nucleotide sequence encoding a koala GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS. In some embodiments, the vector is adapted for genome integration. In some embodiments, the vector is adapted for recombinant protein production and also comprises an expression cassette. The expression cassette may be empty (i.e., no POI coding sequence has been inserted). The expression cassette may also contain one or more POI coding sequences. In some embodiments, the vector provided herein may comprise two or more expression cassettes.

[0210] In some embodiments, the expression system provided herein comprises a vector disclosed herein. In some embodiments, the expression system provided herein may comprise a vector containing the GS coding sequence disclosed herein and an expression cassette. As described above, an expression cassette may comprise one or more nucleotide sequences encoding one or more target peptides. In some embodiments, the expression system provided herein may comprise a vector containing the GS coding sequence disclosed herein and two or more expression cassettes. In some embodiments, the expression system provided herein may have two different vectors, each containing the GS coding sequence disclosed herein and at least one expression cassette. In some embodiments, the expression system provided herein may have multiple different vectors, each containing the GS coding sequence disclosed herein and at least one expression cassette.

[0211] The expression system provided herein comprises the host cell described herein. In some embodiments, the host cell provided herein may have wild-type endogenous GS. In some embodiments, the endogenous GS of the host cell has a mutation that leads to reduced activity. In some embodiments, the endogenous GS of the host cell is inactivated. In some embodiments, the endogenous GS of the host cell is knocked out.

[0212] When the expression system provided herein is used to produce recombinant POIs, the vector is introduced into a host cell. In some embodiments, the vector is stably introduced into the host cell. In some embodiments, the vector is transiently introduced into the host cell. Therefore, in some embodiments of the expression system provided herein, the vector and the host cell are two separate components. In some embodiments of the expression system provided herein, the vector is present within the host cell.

[0213] Any host cell disclosed herein or otherwise known in the art that can be used for recombinant production of POIs can be used in the expression systems disclosed herein. In some embodiments, the host cell is a eukaryotic cell. In some embodiments, the host cell is a mammalian cell. Examples of suitable mammalian host cell lines include (but are not limited to) COS-7 (monkey kidney), L-929 (mouse fibroblast), C127 (mouse mammary tumor), NS0 (non-secreting mouse myeloma), SP2 / 0 (mouse myeloma), 3T3 (mouse fibroblast), CHO (Chinese hamster ovary), HeLa (human cervical cancer), BHK (hamster kidney fibroblast), HEK-293 (human embryonic kidney) cell lines (e.g., HEK293-F, HEK293-H, HEK293-T), PERC.6 (human embryonic retinal cell), HROC277 (human colorectal adenocarcinoma cell), VERO (African green monkey kidney), MDCK (canine kidney), WI38 (human lung fibroblast), V79 (Chinese hamster lung), BHK (young hamster kidney fibroblast), and their variants.

[0214] In some embodiments, the host cell included in the expression system disclosed herein is a CHO cell. In some embodiments, the CHO cells have wild-type endogenous GS. Exemplary CHO cell lines having wild-type GS may be, for example, CHO-S, CHO-K1, CHOK1SV, CHOZN K1, or FreeStyle CHO-S. In some embodiments, CHO cells having wild-type GS may have mutations in genes other than GS that reduce or eliminate the enzymatic activity of proteins involved in, for example, glycosylation or metabolic pathways. The mutations may be naturally occurring or genetically engineered. Exemplary CHO cell lines having impaired or inactivated glycosylation enzymes (e.g., fucosyltransferase 8) include CHO FUT8 KO. Exemplary CHO cell lines having impaired or inactivated enzymes in metabolic pathways (e.g., dihydrofolate reductase, DHFR) include CHO-DG44, CHO-DUXB11, and CHO-DUKX. In some embodiments, CHO cells with wild-type GS may have enhanced enzymatic activity in genes other than GS, the products of which are involved in cell growth / survival, metabolism, and protein modification. In some embodiments, CHO cells have amplified genes encoding anti-apoptotic proteins. Gene amplification may be caused by naturally occurring mutations or genetic engineering alterations. In some embodiments, CHO cells have exogenous genes encoding anti-apoptotic proteins.

[0215] In some embodiments, the endogenous GS of CHO cells may have reduced activity compared to wild-type hamster GS. In some embodiments, the endogenous GS of the CHO cells provided herein is inactivated. In some embodiments, the endogenous GS of the CHO cells provided herein is knocked out. Exemplary CHO cell lines with endogenous GS knockout may be, for example, CHOK1SV GS-KO, CHOZN GS- / -, or CHOZN GS KO.

[0216] In some embodiments, the expression system provided herein also includes a glutamine-free culture medium. Any suitable culture medium can be used, depending on the host cells included in the expression system. A variety of culture media are well known and commercially available in the art. For example, the expression system provided herein may include CHO cells and a culture medium suitable for CHO cells. Exemplary culture media for CHO cells may be, for example, serum-free medium (SFM), protein-free medium (PFM), and chemically defined medium (CDM) (Li et al., Front. Bioeng. Biotechnol., (2021) https: / / doi.org / 10.3389 / fbioe.2021.646363; Ritacco et al., Biotechnology Progress 34.6 (2018): 1407-1426).

[0217] Key components of the culture medium in the expression system described herein include water, carbon source, nitrogen source and phosphate, certain amino acids, fatty acids, vitamins, trace elements, and salts. The water should be uncontaminated and endotoxin-free. In some embodiments, water for injection (WFI) is used. Glucose can serve as the primary energy and carbon source. Although CHO cells can maintain high viability in glucose-limited media, glucose levels are typically kept high due to the rapid growth and nutrient consumption rate of CHO cells in recombinant protein production media. Alternatives to glucose include galactose, fructose, mannose, and other hexoses. The choice of carbon source can affect the glycosylation of recombinant proteins. For example, a medium containing a high concentration of mannose can inhibit intracellular α-mannosidase, thereby increasing the percentage of mannose glycosylation in the product, which can enhance both antibody-dependent cell-mediated cytotoxicity (ADCC) and antibody clearance in the human body.

[0218] Amino acids are also key components of cell culture media, and maintaining most amino acids within specific concentration ranges in the medium is crucial for CHO cell culture. Studies have shown that optimizing the amino acid composition of cell culture media can improve growth profiles and titers, and can also achieve desired product glycosylation patterns. In addition to increasing titers and peak cell density, selected amino acids at specific concentrations can have a protective effect on cell growth in bioreactors. Some amino acids can also eliminate or mitigate some of the adverse effects of ammonium and pCO2 accumulation and high osmotic pressure. Some amino acids also act as signaling molecules, thereby reducing apoptosis rates in mammalian cells. Essential amino acids include histidine, isoleucine, leucine, lysine, methionine, phenylalanine, threonine, tryptophan, and valine, which are provided in high concentrations in the medium. Specifically, tryptophan can be a limiting factor, and tryptophan supplementation has been shown to improve both titers and peak cell density. Although non-essential amino acids can be synthesized by cultured mammalian cells, cell culture media typically still contain most or all of these amino acids to support cell growth and protein production. In fact, most non-essential amino acids have a significant impact on the cell culture process. Furthermore, some amino acid substitutions can help achieve improved solubility and stability. For example, tyrosine, the least soluble amino acid, can be replaced by disodium phosphate tyrosine or a tyrosine-containing dipeptide to improve solubility. Cysteine ​​is one of the most unstable amino acids and can be oxidized to cystine, which has low solubility, at neutral pH. The highly soluble and stable cysteine ​​derivative S-thiocysteine ​​has been reported to be used as an alternative cysteine ​​source and antioxidant in CHO cell culture media.

[0219] Lipids are major components of biological membranes and also serve as energy and signal transduction molecules in mammalian cells. Normally, CHO cells can synthesize lipids themselves. However, lipid supplementation in serum-free culture media has been shown to benefit cell viability and product glycosylation. Exogenous supplementation of phospholipids, such as phosphatidic acid and lysophosphatidic acid, has been shown to stimulate CHO cell growth. As major components of phospholipids, choline and ethanolamine have shown cell-growth-enhancing effects comparable to those of mixed lipids.

[0220] Vitamins act as coenzymes, prosthetic groups, or cofactors in signaling cascades and in enzyme inhibition and activation. Although required in trace amounts, vitamins are a major component of cell culture media, especially in CDM. In CHO cell culture, the addition of vitamins has been shown to increase the volumetric yield of mAbs by up to 3 times.

[0221] The effective concentrations of trace elements in cell culture media are often very low, and can even be below detectable levels. However, their importance cannot be ignored. For example, in CHO cell culture, the concentration of copper should be carefully optimized relative to both culture performance and product quality. Additionally, iron is an essential component of CDM; and zinc supplementation has been shown to not only increase mAb yield by 1.2-fold but also reduce apoptosis. Other trace elements that may be included in the culture media disclosed herein include manganese, molybdenum, selenium, and vanadium, as well as germanium, rubidium, zirconium, cobalt, nickel, tin, and chromium for certain cell types.

[0222] Salts play important chemical and biological roles in CHO cell culture media, including maintaining cell membrane potential, osmotic pressure, and buffering. The main ions added to CHO culture media include sodium, potassium, magnesium, calcium, chloride, phosphate, carbonate (bicarbonate), sulfate, and nitrate.

[0223] Growth factors, typically peptides, small proteins, and hormones, act as signaling molecules influencing cell growth, proliferation, recovery, and differentiation. In many early culture media, growth factors are provided in serum form. In SFM (Self-Focused Growth Factor), only small amounts of specific growth factors are provided, thus minimizing the overall complexity of the culture medium formulation. Widely used growth factors include, for example, insulin and its analogues, as well as other autocrine growth factors such as brain-derived neurotrophic factor (BDNF), fibroblast growth factor 8 (FGF8), growth-regulating alpha protein (CXCL1), hepatocyte growth factor (HGF), hepatocellular carcinoma-derived growth factor (HDGF), leukemia inhibitory factor (LIF), macrophage colony-stimulating factor 1 (CSF1), and vascular endothelial growth factor C (VEGFC). As an alternative, the small molecule antioxidant chelate gold tricarboxylic acid (ATA) has shown insulin-like effects on promoting CHO cell growth.

[0224] Polyamines are ubiquitous molecules in mammalian cells, playing crucial roles in a variety of metabolic processes, including DNA synthesis and transcription, ribosome function, ion channel regulation, and cell signal transduction. Although mammalian cells synthesize several polyamines from ornithine, supplementing the culture medium with polyamines such as putrescine, spermidine, and spermine is essential for supporting and accelerating CHO cell growth.

[0225] The culture media described herein may include non-nutrient components to provide a more stable physical or chemical environment for cells. These components, which may include buffers, surfactants, and / or antifoamers, can have a significant impact on cell growth and yield. Exemplary buffers include bicarbonate buffer systems (CO2 / NaHCO3), organic zwitterionic buffers such as HEPES, or phosphate buffers. Exemplary surfactants include Pluronic F-68 and silicone-based antifoamers such as Antifoam C (Sigma-Aldrich).

[0226] Accordingly, the culture medium provided herein can be supplemented to achieve higher titers, specific glycosylation patterns, cell densities, etc. In some embodiments, the expression system provided herein may further include supplements such as amino acids, lipids, trace elements, salts, or any other supplements disclosed herein or otherwise known in the art (e.g., Ritacco et al., Biotechnology Progress 34.6 (2018):1407-26). Supplements may be included directly in the culture medium or stored separately.

[0227] In some embodiments, the expression system provided herein further includes a GS inhibitor. The GS inhibitor may be any GS inhibitor disclosed herein or otherwise known in the art. The GS inhibitor may be MSX and its derivatives, a phosphorus-containing analog of glutamate, or a bisphosphate. Different concentrations of GS inhibitor may be added to produce different levels of selectivity. In some embodiments, the expression system provided herein also includes MSX. In some embodiments, the expression system provided herein includes a culture medium and MSX. In some embodiments, the expression system provided herein includes a glutamine-free culture medium and MSX. In some embodiments, the culture medium and MSX are two separate components of the expression system. In some embodiments, the culture medium is supplemented with MSX. In some embodiments, the expression system provided herein includes a glutamine-free culture medium supplemented with about 1-10 μM, about 10-50 μM, about 50-100 μM, or about 100-300 μM MSX. In some embodiments, the expression system provided herein comprises a glutamine-free culture medium supplemented with about 1 μM, about 3 μM, about 5 μM, about 10 μM, about 25 μM, about 50 μM, about 75 μM, about 100 μM, about 150 μM, about 200 μM, about 250 μM, or about 300 μM of MSX.

[0228] In some embodiments, the expression systems provided herein also include methods for introducing the vector into host cells. The method of introducing the vector into cells can be any manner known in the art. A variety of methods can be used to introduce polynucleotides or vectors into host cells, which are well known in the art and are selected in part based on the host cell. For example, vectors can be introduced into cells using chemical, physical, biological, or viral methods. Methods for introducing polynucleotides or vectors into host cells include (but are not limited to) the use of calcium phosphate, dendritic polymers, cationic polymers, lipid transfection, liposomes, Fugene, peptide dendritic polymers, electroporation, cell extrusion, acoustic perforation, optical transfection, protoplast fusion, puncture transfection, hydrodynamic delivery, injection, gene gun, magnetic transfection, particle bombardment, nuclear transfection, and viral transduction.

[0229] In some embodiments, the expression systems provided herein may include methods of inserting a vector into the host cell genome to generate stable cell lines. Genome integration methods include, for example, lentiviral transfection, baculovirus gene transfer to mammalian cells (BacMam), retroviral transfection, CRISPR / Cas9, and / or transposons. In some embodiments, the expression systems provided herein may include methods of transiently introducing a vector into host cells. In some embodiments, transient transfection methods include using viral vectors, helper lipids, such as PEI, Lipofectamine, or Fectamine293.

[0230] The expression systems disclosed herein may be provided, for example, in the form of a kit. In some embodiments, the kit comprises a vial containing a DNA vector and another vial containing host cells. Thus, a kit for the in vitro production of POIs is provided herein, comprising a DNA vector having a GS coding sequence and an expression cassette, and host cells. In some embodiments, the host cells are CHO cells. In some embodiments, the kits provided herein also comprise glutamine-free culture medium, MSX, or both. In some embodiments, the kit also includes instructions for use. 6.6 How to use

[0231] This article also provides methods for using the optional markers, vectors, cells, and expression systems disclosed herein in, for example, the identification of genomic loci with high transcriptional activity, the identification of host cells with high POI expression, and the production of recombinant proteins in vitro. 6.6.1 Screening Method

[0232] The selectable biomarkers disclosed herein can be used to identify host cells with high POI expression. The selectable biomarkers disclosed herein can be used to identify genomic loci with high transcriptional activity. In some embodiments, the method includes (1) screening GS-expressing cells from a cell population in which the GS coding sequence disclosed herein has been integrated into its random genomic loci, and (2) identifying genomic loci in which the GS coding sequence is integrated into the GS-expressing cells, wherein the genomic loci represent loci with high transcriptional activity. In some embodiments, the screening step includes culturing the cell population under conditions in which only GS-expressing cells can grow. For example, cells can be cultured in a glutamine-free medium. Cells can also be cultured in a medium supplemented with a GS inhibitor (e.g., MSX). GS inhibitors can be added at different concentrations to produce different levels of selection strictness. In some embodiments, the screening step includes two or more rounds of screening. In some embodiments, two or more rounds of screening include different levels of strictness.

[0233] Accordingly, in some embodiments, the methods provided herein include introducing a vector containing the GS coding sequence described herein into a host cell population and culturing the host cell population in a glutamine-free medium. The GS coding sequence is integrated into a random locus of the host cell genome, and only those with the GS coding sequence inserted at a genomic locus having high transcriptional activity can express GS at sufficient levels to grow in a glutamine-free medium. In some embodiments, the medium also contains a GS inhibitor (e.g., MSX). Any method for introducing nucleotide sequences into cells disclosed herein or otherwise known in the art can be used.

[0234] In some embodiments, genomic loci can be located by sequencing. Any sequencing method known in the art can be used. In some embodiments, next-generation sequencing is used. Genomic loci in which the GS coding sequence is integrated into GS-expressing cells can be identified using any method known in the art. In some embodiments, the methods provided herein further include replacing the GS coding sequence with a POI coding sequence in the identified host cell. The replacement of the GS coding sequence with a POI coding sequence can be performed using any method known in the art. For example, in some embodiments, recombinases such as Cre or Flp can be used. In some embodiments, the replacement can be assisted by DNA breaks produced by enzymes such as zinc finger nucleases, transcription activator-like effector nucleases (TALENs), or CRISPR-Cas systems. The CRISPR-Cas system can be the CRISPR-Cas9 system (Cong et al., Science 2013, 339:819-823).

[0235] The optional markers and vectors disclosed herein can be used in the identification of host cells with high POI yields. The method includes (1) introducing a vector disclosed herein, having both a GS-coding sequence and a POI-coding sequence, into a population of host cells, and (2) identifying host cells expressing POI. Any method disclosed herein or otherwise known in the art for introducing a vector into host cells can be used in the methods disclosed herein. In some embodiments, the identification step includes culturing host cells under conditions in which only GS-expressing cells can grow. For example, cells can be cultured in a glutamine-free medium. Cells can also be cultured in a medium supplemented with a GS inhibitor (e.g., MSX). GS inhibitors can be added at different concentrations to produce different levels of selection strictness. In some embodiments, the screening step includes two or more rounds of screening. In some embodiments, two or more rounds of screening include different levels of strictness. In some embodiments, this document also provides a method for screening cell clones with high POI yield, which includes culturing a population of cell clones transfected with the vector provided herein in a glutamine-free medium supplemented with a GS inhibitor, wherein cell clones capable of growing in said medium are identified as cell clones with high POI yield.

[0236] In some embodiments, the methods provided herein identify host cells expressing POI at specific levels. For example, in some embodiments, the methods provided herein identify host cells with POI production titers in the range of 0.2–500 μg / mL at an early stage (e.g., in 96-well plates or culture tubes). In some embodiments, the methods provided herein identify host cells with POI production titers in the range of 0.05–20 g / L at a later stage (e.g., in flasks or bioreactors).

[0237] In some implementations, early POI production can achieve the following titers: at least 1 μg / mL, at least 2 μg / mL, at least 5 μg / mL, at least 8 μg / mL, at least 10 μg / mL, at least 20 μg / mL, at least 50 μg / mL, at least 80 μg / mL, at least 100 μg / mL, at least 150 μg / mL, at least 200 μg / mL, at least 250 μg / mL, at least 300 μg / mL, at least 350 μg / mL, at least 400 μg / mL, at least 450 μg / mL, or at least 500 μg / mL. In some implementations, early POI production can achieve the following titers: about 1 μg / mL, about 2 μg / mL, about 5 μg / mL, about 8 μg / mL, about 10 μg / mL, about 20 μg / mL, about 50 μg / mL, about 80 μg / mL, about 100 μg / mL, about 150 μg / mL, about 200 μg / mL, about 250 μg / mL, about 300 μg / mL, about 350 μg / mL, about 400 μg / mL, about 450 μg / mL, or about 500 μg / mL.

[0238] In some embodiments, late-stage POI production can achieve the following titers: at least 0.05 g / L, at least 0.1 g / L, at least 0.2 g / L, at least 0.5 g / L, at least 0.8 g / L, at least 1 g / L, at least 2 g / L, at least 5 g / L, at least 8 g / L, at least 10 g / L, at least 12 g / L, at least 15 g / L, at least 18 g / L, or at least 20 g / L. In some embodiments, late-stage POI production can achieve the following titers: about 0.05 g / L, about 0.1 g / L, about 0.2 g / L, about 0.5 g / L, about 0.8 g / L, about 1 g / L, about 2 g / L, about 5 g / L, about 8 g / L, about 10 g / L, about 12 g / L, about 15 g / L, about 18 g / L, or about 20 g / L.

[0239] In some implementations, early POI production can achieve the following titers: 1-10 μg / mL, 1-50 μg / mL, 1-100 μg / mL, 1-200 μg / mL, 1-500 μg / mL, 10-50 μg / mL, 10-100 μg / mL, 10-200 μg / mL, 10-500 μg / mL, 50-100 μg / mL, 50-200 μg / mL, 50-500 μg / mL, 100-200 μg / mL, 100-500 μg / mL, or 200-500 μg / mL. In some implementations, late-stage POI production can achieve the following titers: 0.05-0.1 g / L, 0.05-0.2 g / L, 0.05-0.5 g / L, 0.05-1 g / L, 0.05-2 g / L, 0.05-5 g / L, 0.05-10 g / L, 0.05-20 g / L, 0.1-0.2 g / L, 0.1-0.5 g / L, 0.1-1 g / L, 0.1-2 g / L, 0.1-5 g / L, 0.1-10 g / L, 0.1-20 g / L, 0.2-0.5 g / L. Between g / L, 0.2-1g / L, 0.2-2g / L, 0.2-5g / L, 0.2-10g / L, 0.2-20g / L, 0.5-1g / L, 0.5-2g / L, 0.5-5g / L, 0.5-10g / L, 0.5-20g / L, 1-2g / L, 1-5g / L, 1-10g / L, 1-20g / L, 2-5g / L, 2-10g / L, 2-20g / L, 5-10g / L, 5-20g / L, or 10-20g / L.

[0240] In some embodiments, the methods provided herein include separating a population of host cells into multiple pools by introducing a vector described herein, and measuring POI expression in each pool to determine the cell pool with the desired production capacity.

[0241] The expression level or production capacity of POIs in host cells or host cell pools can be measured using any method known in the art. Exemplary detection methods may include immunohistochemistry, immunocytochemistry, flow cytometry (e.g., FACS), magnetic beads conjugated with antibody molecules, ELISA assays, etc.

[0242] The host cell that can be used in the methods described above can be any cell that requires glutamine for survival and growth. In some embodiments, the host cell can be a eukaryotic cell, such as a yeast cell (e.g., a Saccharomyces cerevisiae or Yarrowia lipolytica cell line), a fungal cell line (e.g., an Aspergillus niger cell line), an insect cell line (e.g., a Spodoptera fugitiveda cell line, such as Sf9), or a mammalian cell. Examples of suitable mammalian host cell lines include (but are not limited to) COS-7 (monkey kidney), L-929 (mouse fibroblast), C127 (mouse mammary tumor), NS0 (non-secreting mouse myeloma), SP2 / 0 (mouse myeloma), 3T3 (mouse fibroblast), CHO (Chinese hamster ovary), HeLa (human cervical cancer), BHK (hamster kidney fibroblast), HEK-293 (human embryonic kidney) cell lines (e.g., HEK293-F, HEK293-H, HEK293-T), PERC.6 (human embryonic retinal cell), HROC277 (human colorectal adenocarcinoma cell), VERO (African green monkey kidney), MDCK (canine kidney), WI38 (human lung fibroblast), V79 (Chinese hamster lung), BHK (young hamster kidney fibroblast), and their variants.

[0243] In some embodiments, the host cell included in the methods described above and disclosed herein is a CHO cell. In some embodiments, the CHO cells have wild-type endogenous GS. Exemplary CHO cell lines having wild-type GS may be, for example, CHO-S, CHO-K1, CHOK1SV, CHOZN K1, or FreeStyle CHO-S. In some embodiments, CHO cells having wild-type GS may have mutations in genes other than GS that reduce or eliminate the enzymatic activity of proteins involved in, for example, glycosylation or metabolic pathways. The mutations may be naturally occurring mutations or genetically engineered mutations. Exemplary CHO cell lines having impaired or inactivated glycosylation enzymes (e.g., fucosyltransferase 8) include CHO FUT8 KO. Exemplary CHO cell lines having impaired or inactivated enzymes in metabolic pathways (e.g., dihydrofolate reductase, DHFR) include CHO-DG44, CHO-DUXB11, and CHO-DUKX. In some embodiments, CHO cells with wild-type GS may have enhanced enzymatic activity in genes other than GS, the products of which are involved in cell growth / survival, metabolism, and protein modification. In some embodiments, CHO cells have amplified genes encoding anti-apoptotic proteins. Gene amplification may be caused by naturally occurring mutations or genetic engineering alterations. In some embodiments, CHO cells have exogenous genes encoding anti-apoptotic proteins.

[0244] In some embodiments, the endogenous GS of CHO cells may have reduced activity compared to wild-type hamster GS. In some embodiments, the endogenous GS of the CHO cells provided herein is inactivated. In some embodiments, the endogenous GS of the CHO cells provided herein is knocked out. Exemplary CHO cell lines with endogenous GS knockout may be, for example, CHOK1SV GS-KO, CHOZN GS- / -, or CHOZN GS KO.

[0245] To select genomic loci with high transcriptional activity or host cells expressing high levels of POIs, GS inhibitors can be used to generate stringent selection conditions. This is because high transcriptional activity / expression will be required to produce sufficient GS activity levels to allow host cell survival and growth. Any GS inhibitor disclosed herein or otherwise known in the art can be used in the methods disclosed herein, including MSX and its derivatives, phosphorus-containing analogs of glutamate, and bisphosphates. Different concentrations of GS inhibitors can be added to produce different levels of stringency. In some embodiments, MSX is used. In some embodiments, the MSX in the selection conditions is at about 1-10 μM, about 10-50 μM, about 50-100 μM, or about 100-300 μM. In some embodiments, the host cells provided herein can be grown in a culture medium supplemented with about 1 μM, about 3 μM, about 5 μM, about 10 μM, about 25 μM, about 50 μM, about 75 μM, about 100 μM, about 150 μM, about 200 μM, about 250 μM, or about 300 μM of MSX.

[0246] For illustrative purposes, in some embodiments, this document provides a method for identifying CHO cells with high POI production capacity, comprising (1) introducing a vector having a koala GS coding sequence and a POI coding sequence disclosed herein into a population of CHO cells and (2) identifying POI-expressing cells by culturing the CHO cell population in a glutamine-free medium supplemented with approximately 50 μM MSX and measuring the POI expression of surviving cells. 6.6.2 Production Method

[0247] The optional biomarkers, vectors, cells, and expression systems provided herein can be used for the in vitro production of POIs. Accordingly, this document also provides uses of the optional biomarkers, vectors, cells, or expression systems disclosed herein for the in vitro production of POIs. In some embodiments, this document provides a method for the in vitro production of POIs, comprising culturing the host cells disclosed herein under certain conditions for a sufficient time to produce POIs. In some embodiments, the host cells comprise a vector disclosed herein having a GS coding sequence disclosed herein and a POI coding sequence disclosed herein. In some embodiments, the host cells comprise a POI coding sequence inserted at a transcriptionally active genomic locus identified using the methods disclosed herein. The host cells can be any host cells disclosed herein suitable for recombinant production. In some embodiments, the host cells are CHO cells. In some embodiments, the host cells are CHO cells with wild-type endogenous GS. In some embodiments, the host cells are CHO cells with endogenous GS knocked out.

[0248] The POI can be any POI disclosed herein or otherwise known in the art. In some embodiments, the POI is selected from antibodies, enzymes, soluble proteins, secretory proteins, membrane proteins, or fusion proteins.

[0249] In some embodiments, the methods provided herein include identifying host cells with high POI production capacity using selectable markers disclosed herein, and culturing the host cells under certain conditions for a sufficient time to produce POI.

[0250] Methods for culturing host cells, such as CHO cells for recombinant protein production, are well known in the art. (Kim JY et al., Appl Microbiol Biotechnol. 2012; 93(3):917-30; Ritacco FV et al., Biotechnol Prog. 2018; 34(6):1407-1426; Fischer S et al., Biotechnol Adv. 2015; 33(8):1878-96). Using the optional markers provided herein, clones with high production capacity can be identified. For example, in some embodiments, the methods provided herein identify host cells with POI production titers in the range of 0.2-500 μg / mL at an early stage (e.g., in 96-well plates or culture tubes). In some embodiments, the in vitro methods provided herein can produce POI titers of approximately 0.2-500 μg / mL. In some embodiments, the methods provided herein produce POIs at the following titers: at least 1 μg / mL, at least 2 μg / mL, at least 5 μg / mL, at least 8 μg / mL, at least 10 μg / mL, at least 20 μg / mL, at least 50 μg / mL, at least 80 μg / mL, at least 100 μg / mL, at least 150 μg / mL, at least 200 μg / mL, at least 250 μg / mL, at least 300 μg / mL, at least 350 μg / mL, at least 400 μg / mL, at least 450 μg / mL, or at least 500 μg / mL. In some embodiments, the methods provided herein produce POI at the following titers: about 1 μg / mL, about 2 μg / mL, about 5 μg / mL, about 8 μg / mL, about 10 μg / mL, about 20 μg / mL, about 50 μg / mL, about 80 μg / mL, about 100 μg / mL, about 150 μg / mL, about 200 μg / mL, about 250 μg / mL, about 300 μg / mL, about 350 μg / mL, about 400 μg / mL, about 450 μg / mL, or about 500 μg / mL. In some embodiments, the methods provided herein identify host cells with POI production titers in the range of 0.05–20 g / L at later stages (e.g., in flasks or bioreactors). In some embodiments, the methods provided herein produce POI at the following titers: at least 0.05 g / L, at least 0.1 g / L, at least 0.2 g / L, at least 0.5 g / L, at least 0.8 g / L, at least 1 g / L, at least 2 g / L, at least 5 g / L, at least 8 g / L, at least 10 g / L, at least 12 g / L, at least 15 g / L, at least 18 g / L, or at least 20 g / L.

[0251] In some embodiments, the methods provided herein produce POIs at the following titers: about 1 μg / mL, about 2 μg / mL, about 5 μg / mL, about 8 μg / mL, about 10 μg / mL, about 20 μg / mL, about 50 μg / mL, about 80 μg / mL, about 100 μg / mL, about 150 μg / mL, about 200 μg / mL, about 250 μg / mL, about 300 μg / mL, about 350 μg / mL, about 400 μg / mL, about 450 μg / mL, or about 500 μg / mL. In some embodiments, the methods provided herein produce POI at the following titers: about 0.05 g / L, about 0.1 g / L, about 0.2 g / L, about 0.5 g / L, about 0.8 g / L, about 1 g / L, about 2 g / L, about 5 g / L, about 8 g / L, about 10 g / L, about 12 g / L, about 15 g / L, about 18 g / L, or about 20 g / L.

[0252] In some embodiments, the methods provided herein produce POIs at titers between 1-10 μg / mL, 1-50 μg / mL, 1-100 μg / mL, 1-200 μg / mL, 1-500 μg / mL, 10-50 μg / mL, 10-100 μg / mL, 10-200 μg / mL, 10-500 μg / mL, 50-100 μg / mL, 50-200 μg / mL, 50-500 μg / mL, 100-200 μg / mL, 100-500 μg / mL, or 200-500 μg / mL. In some embodiments, the method provided herein uses concentrations between 0.05-0.1 g / L, 0.05-0.2 g / L, 0.05-0.5 g / L, 0.05-1 g / L, 0.05-2 g / L, 0.05-5 g / L, 0.05-10 g / L, 0.05-20 g / L, 0.1-0.2 g / L, 0.1-0.5 g / L, 0.1-1 g / L, 0.1-2 g / L, 0.1-5 g / L, 0.1-10 g / L, 0.1-20 g / L, 0.2-0.5 g / L, and 0. POI production at titers between 0.2-1g / L, 0.2-2g / L, 0.2-5g / L, 0.2-10g / L, 0.2-20g / L, 0.5-1g / L, 0.5-2g / L, 0.5-5g / L, 0.5-10g / L, 0.5-20g / L, 1-2g / L, 1-5g / L, 1-10g / L, 1-20g / L, 2-5g / L, 2-10g / L, 2-20g / L, 5-10g / L, 5-20g / L, or 10-20g / L.

[0253] The in vitro method described herein can be used at approximately 1-800 μg / 10 6 POIs are produced at a rate of cells / day. In some embodiments, the in vitro methods provided herein can produce at least 5 μg / 10 6 10 cells / day, at least 10 μg / 10 6 10 cells / day, at least 75 μg / 10 6 10 cells / day, at least 100 μg / 10 6 10 cells / day, at least 150 μg / 10 6 10 cells / day, at least 200 μg / 10 6 10 cells / day, at least 250 μg / 10 6 10 cells / day, at least 300 μg / 10 610 cells / day, at least 400 μg / 10 6 10 cells / day or at least 500 μg / 10 6 POI production per cell / day. In some embodiments, the in vitro methods provided herein can produce approximately 5 μg / 10 cells / day. 6 10 cells / day, approximately 10 μg / 10 6 10 cells / day, approximately 75 μg / 10 6 100 μg / day 10 cells / day 6 10 cells / day, approximately 150 μg / 10 6 10 cells / day, approximately 200 μg / 10 6 10 cells / day, approximately 250 μg / 10 6 10 cells / day, approximately 300 μg / 10 6 10 cells / day, approximately 400 μg / 10 6 1 cell / day or approximately 500 μg / 10 6 POI production per cell per day. The POI expression level or production capacity of a host cell or host cell pool can be measured by any method known in the art. Exemplary detection methods may include immunohistochemistry, immunocytochemistry, flow cytometry (e.g., FACS), magnetic beads conjugated with antibody molecules, ELISA assays, etc.

[0254] In some embodiments, the methods provided herein further include separating the protein from other components in the culture. In some embodiments, the separation includes extraction, continuous liquid-liquid extraction, pervaporation, membrane filtration, membrane separation, reverse osmosis, electrodialysis, free-flow electrophoresis, affinity chromatography, immunoaffinity chromatography, high-performance liquid chromatography, distillation, crystallization, centrifugation, extractive filtration, size exclusion chromatography, hydrophobic interaction chromatography, ion exchange chromatography, adsorption chromatography, or ultrafiltration.

[0255] All papers, publications, and patents cited in this specification are incorporated herein by reference as if each individual paper, publication, or patent were specifically and individually indicated as such, and these papers, publications, and patents are incorporated herein by reference to disclose and describe methods and / or materials relating to the cited publications. However, any reference to any reference, paper, publication, patent, patent publication, or patent application cited herein is not and should not be construed as an admission or implication of any kind that they constitute valid prior art or are part of general common sense in any country worldwide.

[0256] Unless the context otherwise indicates, it is specifically intended to indicate that the various features described herein can be used in any combination.

[0257] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. 6.7 Exemplary Implementation

[0258] Implementation Method 1: Use of a nucleotide sequence encoding glutamine synthase (GS) as a selectable marker, wherein the GS is koala GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS.

[0259] Implementation Method 2: According to the use described in Implementation Method 1, the GS is a koala GS derived from the Phascolarctidae family.

[0260] Implementation 3: According to the use described in Implementation 2, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 1.

[0261] Implementation 4: The use according to Implementation 2 or 3, wherein the GS has reduced activity compared to wild-type koala GS.

[0262] Embodiment 5: According to the use described in Embodiment 1, wherein the GS has the amino acid sequence shown in SEQ ID NO: 1.

[0263] Implementation method 6: According to the use described in implementation method 1, wherein the GS is an ostrich GS derived from the Struthionidae family.

[0264] Embodiment 7: According to the use described in Embodiment 6, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 2.

[0265] Implementation method 8: The use according to implementation method 6 or 7, wherein the GS has reduced activity compared to wild-type ostrich GS.

[0266] Embodiment 9: According to the use described in Embodiment 1, wherein the GS has the amino acid sequence shown in SEQ ID NO: 2.

[0267] Implementation 10: According to the use described in Implementation 1, wherein the GS is a snake GS derived from the Elapidae family.

[0268] Embodiment 11: According to the use described in Embodiment 10, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 3.

[0269] Embodiment 12: The use according to Embodiment 10 or 11, wherein the GS has reduced activity compared to wild-type snake GS.

[0270] Embodiment 13: According to the use described in Embodiment 1, wherein the GS has the amino acid sequence shown in SEQ ID NO: 3.

[0271] Implementation Method 14: According to the use described in Implementation Method 1, the GS is a pigeon GS derived from the Columbidae family.

[0272] Embodiment 15: According to the use described in Embodiment 14, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 5.

[0273] Implementation 16: The use according to Implementation 14 or 15, wherein the GS has reduced activity compared to wild-type pigeon GS.

[0274] Embodiment 17: According to the use described in Embodiment 1, wherein the GS has the amino acid sequence shown in SEQ ID NO: 5.

[0275] Implementation method 18: According to the use described in implementation method 1, the GS is a red junglefowl GS derived from the Pheasantidae family.

[0276] Embodiment 19: According to the use described in Embodiment 18, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 6.

[0277] Implementation 20: The use according to Implementation 18 or 19, wherein the GS has reduced activity compared to wild-type junglefowl GS.

[0278] Embodiment 21: According to the use described in Embodiment 1, wherein the GS has the amino acid sequence shown in SEQ ID NO: 6.

[0279] Implementation Method 22: According to the use described in Implementation Method 1, wherein the GS is a gerbil GS derived from the Muridae family.

[0280] Embodiment 23: According to the use described in Embodiment 22, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 7.

[0281] Implementation 24: The use according to Implementation 22 or 23, wherein the GS has reduced activity compared to wild-type gerbil GS.

[0282] Embodiment 25: The use according to Embodiment 1, wherein the GS has the amino acid sequence shown in SEQ ID NO: 7.

[0283] Implementation Method 26: According to the use described in Implementation Method 1, wherein the GS is a night bat GS derived from the Vespertilionidae family.

[0284] Embodiment 27: According to the use described in Embodiment 26, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 8.

[0285] Implementation 28: The use according to Implementation 26 or 27, wherein the GS has reduced activity compared to wild-type night bat GS.

[0286] Embodiment 29: According to the use described in Embodiment 1, wherein the GS has the amino acid sequence shown in SEQ ID NO: 8.

[0287] Embodiment 30: According to the use described in Embodiment 1, the amino acid sequence of said GS includes SEQ ID NO: 9 with at least one amino acid substitution at the following positions: A2, H8, N10, G12, Q15, M16, M18, S19, E24, Q27, V33, G39, D48, C49, C53, V54, E55, E56, F68, S70, S72, S80, V82, M84, E92, V97, F98, F102, Q106, K107, P108, E110, H115, T116, K118, D122, S125, H128, L140, D152, L160, K169, R172, I175, M176, V188, K189, T191, Y194 V197, K198, H199, A200, I206, I212, R213, V220, K230, V234, A236, T237, S240, T260, E264, N265, H269, K271, E272, A273, K276, S278, R282, L294, H304, K305, N308, N310, D311, D318, S320, T328, Q331, E332, K334, F337, A339, C341, F349, I355, V356, D366, Q367, and Q370.

[0288] Embodiment 31: According to the use described in Embodiment 1, the amino acid sequence of said GS includes SEQ ID NO: 9, selected from the following positions having at least one amino acid substitution: N10, G12, Q15, S19, E24, V33, G39, C49, C53, V54, E56, S72, S80, V82, E92, F98, Q106, K107, P108, K118, L140, D152, L160, R172, M176, T191, Y194, K198, H19 9. I206, R213, V220, K230, A236, T237, S240, T260, E264, N265, H269, K271, R282, K305, N310, D311, D318, S320, T328, Q331, A339, C341, F349, I355, Q367 and Q370.

[0289] Embodiment 32: According to the use described in Embodiment 31, the amino acid sequence of said GS includes SEQ ID NO: 9: N10, G12, Q15, S19, E24, V33, G39, C49, C53, V54, E56, S72, S80, V82, E92, F98, Q106, K107, P108, K118, L140, D152, L160, R172, M176, T191, Y194, K198, H19 at positions selected from the following positions having about 3, about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, or about 55 amino acid substitutions. 9. I206, R213, V220, K230, A236, T237, S240, T260, E264, N265, H269, K271, R282, K305, N310, D311, D318, S320, T328, Q331, A339, C341, F349, I355, Q367 and Q370.

[0290] Implementation 33: Use of a nucleotide sequence encoding a GS as an optional marker, wherein the GS comprises a catalytic domain derived from koala GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS.

[0291] Embodiment 34: According to the use described in Embodiment 33, wherein the GS comprises a catalytic domain from a koala GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 1.

[0292] Embodiment 35: The use according to Embodiment 34, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 1.

[0293] Embodiment 36: According to the use described in Embodiment 33, wherein the GS comprises a catalytic domain from an ostrich GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 2.

[0294] Embodiment 37: According to the use described in Embodiment 36, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 2.

[0295] Embodiment 38: According to the use described in Embodiment 33, wherein the GS comprises a catalytic domain from a snake GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3.

[0296] Embodiment 39: The use according to Embodiment 38, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 3.

[0297] Embodiment 40: According to the use described in Embodiment 33, wherein the GS comprises a catalytic domain from a pigeon GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 5.

[0298] Embodiment 41: According to the use described in Embodiment 40, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 163-412 shown in SEQ ID NO: 5.

[0299] Embodiment 42: According to the use described in Embodiment 33, wherein the GS comprises a catalytic domain from a chicken GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 6.

[0300] Embodiment 43: The use according to Embodiment 42, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 150-399 shown in SEQ ID NO: 6.

[0301] Embodiment 44: According to the use described in Embodiment 33, wherein the GS comprises a catalytic domain from a gerbil GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 7.

[0302] Embodiment 45: According to the use described in Embodiment 44, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 7.

[0303] Embodiment 46: According to the use described in Embodiment 33, wherein the GS comprises a catalytic domain from a night-bat GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 8.

[0304] Embodiment 47: According to the use described in Embodiment 46, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 8.

[0305] Embodiment 48: The use according to any one of Embodiments 1 to 47, wherein the GS-encoded nucleotide sequence is operatively linked to an mRNA destabilizing element.

[0306] Embodiment 49: The use according to any one of Embodiments 1 to 47, wherein the GS contains a degradation determinant.

[0307] Embodiment 50: The use according to Embodiment 49, wherein the degradation determinant has an amino acid sequence selected from SEQ ID NO: 23-25.

[0308] Embodiment 51: The use according to any one of Embodiments 1 to 50, which is used to identify genomic loci with high transcriptional activity.

[0309] Embodiment 52: The use according to any one of Embodiments 1 to 50, which is used to identify host cells capable of producing the target protein (POI).

[0310] Implementation method 53: The use according to any one of implementation methods 1 to 50, which is used for the recombinant production of POI.

[0311] Implementation 54: According to the use described in Implementation 53, it is used for the production of recombinant proteins in mammalian cells.

[0312] Embodiment 55: According to the use described in Embodiment 54, the mammalian cells are Chinese hamster ovary (CHO) cells.

[0313] Embodiment 56: The use according to any one of Embodiments 52 to 55, wherein the POI is selected from antibodies, enzymes, soluble proteins, secretory proteins, membrane proteins, and fusion proteins.

[0314] Implementation Method 57: A deoxyribonucleic acid (DNA) vector suitable for recombinant protein production or genome integration, comprising a nucleotide sequence encoding a GS (GS coding sequence), wherein the GS is a koala GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS.

[0315] Implementation 58: The carrier according to Implementation 57, wherein the GS is derived from the koala GS of the Phascolarctidae family.

[0316] Embodiment 59: The carrier according to Embodiment 58, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 1.

[0317] Implementation method 60: The carrier according to implementation method 58 or 59, wherein the GS has reduced activity compared to wild-type koala GS.

[0318] Embodiment 61: The carrier according to Embodiment 57, wherein the GS has the amino acid sequence shown in SEQ ID NO: 1.

[0319] Embodiment 62: The carrier according to any one of Embodiments 58 to 61, wherein the GS encoding sequence has at least 80% identity with SEQ ID NO: 10.

[0320] Implementation method 63: The carrier according to implementation method 57, wherein the GS is an ostrich GS derived from the Struthionidae family.

[0321] Embodiment 64: The carrier according to Embodiment 63, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 2.

[0322] Embodiment 65: The carrier according to Embodiment 63 or 64, wherein the GS has reduced activity compared to wild-type ostrich GS.

[0323] Embodiment 66: The carrier according to Embodiment 57, wherein the GS has the amino acid sequence shown in SEQ ID NO: 2.

[0324] Embodiment 67: The carrier according to any one of Embodiments 63 to 66, wherein the GS encoding sequence has at least 80% identity with SEQ ID NO: 11.

[0325] Implementation method 68: The carrier according to implementation method 57, wherein the GS is derived from a snake GS of the Elapidae family.

[0326] Embodiment 69: The carrier according to Embodiment 68, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 3.

[0327] Implementation 70: The carrier according to Implementation 68 or 69, wherein the GS has reduced activity compared to wild-type snake GS.

[0328] Embodiment 71: The carrier according to Embodiment 57, wherein the GS has the amino acid sequence shown in SEQ ID NO: 3.

[0329] Embodiment 72: The carrier according to any one of Embodiments 68 to 71, wherein the GS encoding sequence has at least 80% identity with SEQ ID NO: 12.

[0330] Implementation method 73: The carrier according to implementation method 57, wherein the GS is derived from pigeon GS of the Columbidae family.

[0331] Embodiment 74: The carrier according to Embodiment 73, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 5.

[0332] Implementation method 75: The carrier according to implementation method 73 or 74, wherein the GS has reduced activity compared to wild-type pigeon GS.

[0333] Embodiment 76: The carrier according to Embodiment 57, wherein the GS has the amino acid sequence shown in SEQ ID NO: 5.

[0334] Embodiment 77: The carrier according to any one of Embodiments 73 to 76, wherein the GS encoding sequence has at least 80% identity with SEQ ID NO: 14.

[0335] Implementation method 78: The carrier according to implementation method 57, wherein the GS is derived from the red junglefowl GS of the Pheasantidae family.

[0336] Embodiment 79: The carrier according to Embodiment 78, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 6.

[0337] Implementation method 80: The carrier according to implementation method 78 or 79, wherein the GS has reduced activity compared to wild-type junglefowl GS.

[0338] Embodiment 81: The carrier according to Embodiment 57, wherein the GS has the amino acid sequence shown in SEQ ID NO: 6.

[0339] Embodiment 82: The carrier according to any one of Embodiments 78 to 81, wherein the GS encoding sequence has at least 80% identity with SEQ ID NO: 15.

[0340] Implementation method 83: The carrier according to implementation method 57, wherein the GS is derived from the gerbil GS of the Muridae family.

[0341] Embodiment 84: The carrier according to Embodiment 83, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 7.

[0342] Implementation method 85: The carrier according to implementation method 83 or 84, wherein the GS has reduced activity compared to wild-type gerbil GS.

[0343] Embodiment 86: The carrier according to Embodiment 57, wherein the GS has the amino acid sequence shown in SEQ ID NO: 7.

[0344] Embodiment 87: The carrier according to any one of Embodiments 83 to 86, wherein the GS encoding sequence has at least 80% identity with SEQ ID NO: 16.

[0345] Implementation method 88: The carrier according to implementation method 57, wherein the GS is a nocturnal bat GS derived from the Vespertilionidae family.

[0346] Embodiment 89: The carrier according to Embodiment 88, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 8.

[0347] Implementation 90: The carrier according to Implementation 88 or 89, wherein the GS has reduced activity compared to wild-type night bat GS.

[0348] Embodiment 91: The carrier according to Embodiment 57, wherein the GS has the amino acid sequence shown in SEQ ID NO: 8.

[0349] Embodiment 92: The carrier according to any one of Embodiments 88 to 91, wherein the GS encoding sequence has at least 80% identity with SEQ ID NO: 17.

[0350] Embodiment 93: The carrier according to Embodiment 57, wherein the amino acid sequence of the GS includes a SEQ ID having at least one amino acid substitution at a position selected from the following positions. NO: 9: A2, H8, N10, G12, Q15, M16, M18, S19, E24, Q27, V33, G39, D48, C49, C53, V54, E55, E56, F68, S70, S72, S80, V82, M84, E92, V97, F98 , F102, Q106, K107, P108, E110, H115, T116, K118, D122, S125, H128, L140, D152, L160, K169, R172, I175, M176, V188, K189, T191, Y194 V197, K198, H199, A200, I206, I212, R213, V220, K230, V234, A236, T237, S240, T260, E264, N265, H269, K271, E272, A273, K276, S278, R282, L294, H304, K305, N308, N310, D311, D318, S320, T328, Q331, E332, K334, F337, A339, C341, F349, I355, V356, D366, Q367, and Q370.

[0351] Embodiment 94: The carrier according to Embodiment 57, wherein the amino acid sequence of the GS comprises SEQ ID NO: 9 having at least one amino acid substitution at the following positions: N10, G12, Q15, S19, E24, V33, G39, C49, C53, V54, E56, S72, S80, V82, E92, F98, Q106, K107, P108, K118, L140, D152, L160, R172, M176, T191, Y194, K198, H19 9. I206, R213, V220, K230, A236, T237, S240, T260, E264, N265, H269, K271, R282, K305, N310, D311, D318, S320, T328, Q331, A339, C341, F349, I355, Q367 and Q370.

[0352] Embodiment 95: The carrier according to Embodiment 94, wherein the amino acid sequence of the GS comprises SEQ ID NO: 9: N10, G12, Q15, S19, E24, V33, G39, C49, C53, V54, E56, S72, S80, V82, E92, F98, Q106, K107, P108, K118, L140, D152, L160, R172, M176, T191, Y194, K198, H19 at positions having about 3, about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, or about 55 amino acid substitutions at the following positions. 9. I206, R213, V220, K230, A236, T237, S240, T260, E264, N265, H269, K271, R282, K305, N310, D311, D318, S320, T328, Q331, A339, C341, F349, I355, Q367 and Q370.

[0353] Implementation 96: A DNA vector suitable for recombinant protein production or genome integration, comprising a nucleotide sequence encoding a GS (GS coding sequence), wherein the GS comprises a catalytic domain derived from koala GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, gerbil GS or night bat GS.

[0354] Embodiment 97: The carrier according to Embodiment 96, wherein the GS comprises a catalytic domain from a koala GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 1.

[0355] Embodiment 98: The carrier according to Embodiment 97, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 1.

[0356] Embodiment 99: The carrier according to Embodiment 96, wherein the GS comprises a catalytic domain from an ostrich GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 2.

[0357] Embodiment 100: The carrier according to Embodiment 99, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 2.

[0358] Embodiment 101: The carrier according to Embodiment 96, wherein the GS comprises a catalytic domain from a snake GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 3.

[0359] Embodiment 102: The carrier according to Embodiment 101, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 3.

[0360] Embodiment 103: The carrier according to Embodiment 96, wherein the GS comprises a catalytic domain from a pigeon GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 5.

[0361] Embodiment 104: The carrier according to Embodiment 103, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 163-412 shown in SEQ ID NO: 5.

[0362] Embodiment 105: The carrier according to Embodiment 96, wherein the GS comprises a catalytic domain from a chicken GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 6.

[0363] Embodiment 106: The carrier according to Embodiment 105, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 150-399 shown in SEQ ID NO: 6.

[0364] Embodiment 107: The carrier according to Embodiment 96, wherein the GS comprises a catalytic domain from a gerbil GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 7.

[0365] Embodiment 108: The carrier according to Embodiment 107, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 7.

[0366] Embodiment 109: The carrier according to Embodiment 96, wherein the GS comprises a catalytic domain from a night-bat GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO: 8.

[0367] Embodiment 110: The carrier according to Embodiment 109, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO: 8.

[0368] Embodiment 111: The vector according to any one of Embodiments 57 to 110, wherein the GS coding sequence is operatively linked to an mRNA destabilizing element.

[0369] Embodiment 112: The carrier according to any one of Embodiments 57 to 110, wherein the GS contains a degradation determinant.

[0370] Embodiment 113: The carrier according to Embodiment 112, wherein the degradation determinant has an amino acid sequence selected from SEQ ID NO: 23-25.

[0371] Embodiment 114: A vector according to any one of Embodiments 57 to 113, wherein the vector is suitable for recombinant protein production and further comprises an expression cassette.

[0372] Implementation 115: The vector according to Implementation 114, wherein the GS encoding sequence is operatively linked to the simian cavitation virus 40 (SV40) promoter.

[0373] Implementation 116: The carrier according to Implementation 114 or 115, wherein the GS encoded sequence is operatively linked to a poly(A) tail.

[0374] Embodiment 117: The carrier according to any one of Embodiments 114 to 116 comprises two or more expression boxes.

[0375] Embodiment 118: The vector according to any one of Embodiments 114 to 117, wherein the expression cassette contains a nucleotide sequence encoding a POI (POI coding sequence).

[0376] Embodiment 119: The carrier according to Embodiment 118, wherein the POI is an antibody, enzyme, soluble protein, secretory protein, membrane protein or fusion protein.

[0377] Implementation Method 120: The vector according to Implementation Method 119, wherein the POI is an antibody selected from IgG1 antibody, IgG2 antibody, IgG3 antibody, IgG4 antibody, IgA antibody, IgM antibody, Fab, Fab', F(ab')2, Fv, scFv, (scFv)2, single-domain antibody (sdAb), single-chain antibody (scAb) and heavy-chain antibody (HCAb).

[0378] Implementation Method 121: The vector according to Implementation Method 119, wherein the POI is an antibody selected from the following: monoclonal antibody, bispecific antibody, multispecific antibody, bivalent antibody, and multivalent antibody.

[0379] Embodiment 122: The vector according to any one of Embodiments 118 to 121, wherein the POI consists of one or more copies of the same polypeptide.

[0380] Embodiment 123: The carrier according to any one of Embodiments 118 to 121, wherein the POI comprises two different polypeptides.

[0381] Embodiment 124: The vector according to Embodiment 123, wherein the POI is an antibody comprising a light chain and a heavy chain, each encoded by a separate nucleotide sequence on the vector.

[0382] Embodiment 125: A vector according to any one of Embodiments 57 to 124, wherein the vector is suitable for genome integration.

[0383] Embodiment 126: Use of the vector according to any one of Embodiments 118 to 124 for identifying host cells capable of producing POI.

[0384] Implementation 127: Use of the vector according to Implementation 125 for identifying genomic loci with high transcriptional activity.

[0385] Embodiment 128: A method for identifying host cells capable of producing POI, comprising introducing a vector according to any one of Embodiments 118 to 124 into a population of host cells, culturing the population of host cells in a glutamine-free culture medium, wherein host cells capable of growing in the culture medium are identified as host cells capable of producing POI.

[0386] Implementation 129: A method for identifying genomic loci with high transcriptional activity, comprising introducing a vector according to Implementation 125 into a population of host cells, culturing the population of host cells in a glutamine-free medium, wherein host cells capable of growing in the medium are identified as host cells having a GS coding sequence inserted at a genomic locus with high transcriptional activity.

[0387] Implementation 130: The method according to implementation 129 further includes sequencing the genome of the identified host cell to locate genomic loci with high transcriptional activity.

[0388] Implementation method 131: The method according to any one of implementation methods 128 to 130, wherein the host cell population is cultured in the presence of a GS inhibitor.

[0389] Implementation 132: The method according to Implementation 131, wherein the GS inhibitor is methionine sulfoxide (MSX).

[0390] Embodiment 133: Host cell comprising a vector according to any one of Embodiments 118 to 124.

[0391] Implementation method 134: The host cell according to implementation method 133 has wild-type endogenous GS.

[0392] Implementation 135: The host cell according to Implementation 133, wherein the endogenous GS of the host cell has reduced activity or is knocked out.

[0393] Embodiment 136: A host cell according to any one of Embodiments 133 to 135, wherein the host cell is a mammalian cell.

[0394] Implementation method 137: The host cell according to implementation method 136 is a CHO cell.

[0395] Embodiment 138: Use of the host cell according to any one of Embodiments 133 to 137 for in vitro production of POI.

[0396] Implementation Method 139: A method for producing POI in vitro, comprising culturing a host cell according to any one of Implementation Methods 133 to 137 for a sufficient time under certain conditions to produce the POI.

[0397] Implementation 140: A method for producing POI in vitro, comprising replacing the GS coding sequence with a POI coding sequence in a host cell identified in the method according to Implementation 129, and culturing the host cell under certain conditions for a time sufficient to produce POI.

[0398] Implementation 141: The method according to Implementation 139 or 140 further includes separating the POI from other components in the culture.

[0399] Implementation Method 142: The method according to Implementation Method 141, wherein the separation includes extraction, continuous liquid-liquid extraction, pervaporation, membrane filtration, membrane separation, reverse osmosis, electrodialysis, distillation, crystallization, centrifugation, extractive filtration, ion exchange chromatography, adsorption chromatography, or ultrafiltration.

[0400] Embodiment 143: An expression system for in vitro production of POI, comprising a DNA vector and a host cell according to any one of Embodiments 57 to 125.

[0401] Implementation 144: The expression system according to Implementation 143, wherein the host cell is a CHO cell.

[0402] Embodiment 145: The expression system according to Embodiment 143 or 144 further includes a glutamine-free culture medium.

[0403] Embodiment 146: The expression system according to any one of Embodiments 143 to 145 further includes a GS inhibitor.

[0404] Embodiment 147: The expression system according to any one of Embodiments 143 to 146 further includes a method for introducing the vector into the host cell.

[0405] Embodiment 148: The expression system according to any one of Embodiments 143 to 147 is included in the kit. 6.8 Experiment

[0406] Unless otherwise stated, the following embodiments are provided for illustrative purposes only and are not intended to be limiting. Therefore, the invention should not be considered as limited to the following embodiments, but should be considered as covering any and all changes that become apparent from the teachings provided herein. 6.8.1 Example 1: Identification of GS markers that significantly improve yield and screening efficiency

[0407] Carrier design: Currently, almost all selectable markers used in the GS system are GS derived from the Chinese hamster (Cricetulus griseus). To investigate whether GS from other species further improve titers or screening efficiency, a set of vectors was designed as follows: - Sequences encoding the eel green fluorescent protein UnaG, 2A autocleavage peptide, and glutamine synthase, located under the control of the SV40 promoter; - Two expression cassettes, in which the heavy and light chain sequences encoding the human anti-RANKL antibody denosumab are respectively controlled by the CMV promoter; - The sequence used in prokaryotic cells that encodes the protein conferring ampicillin resistance; and - The origin of prokaryotic replication.

[0408] The 2A self-cleaving peptide allows for the equimolar expression of two products isolated from it: green fluorescent protein UnaG and glutamine synthase. After random integration, GS expression in each individual cell was evaluated by monitoring UnaG levels using flow cytometry. In this study, denosumab was used as a secretory reporter molecule for titer evaluation using selectable glutamine synthase markers from different species. These vectors have distinct GS-coding sequences from each other, derived from the Chinese hamster (Cricetulus griseus), koala (Phascolarctos cinerecus), night bat (Pipistrellus kuhlii), gerbil (Meriones unguiculatus), ostrich (Struthio camelus australis), red junglefowl (Gallus gallus), pigeon (Columbia livia), and snake (Pseudonaja textilis).

[0409] Electroporation and cell screening: First, the expression vector described above was linearized using the restriction endonuclease PvuI (there is a PvuI site in the ampicillin resistance gene). Then, the linearized vector was introduced into host cells using a cell electroporator (LonzaNuleofector 2b). The host cells were derived from ATCC wild-type Chinese hamster ovaries CHO-K1 (CCL-61) and adapted for serum-free suspension culture at Shanghai ZhenGe Biotech Co., Ltd. Twenty-four hours after electroporation, the cell pools were plated into culture flasks or 96-well plates (approximately 1,200 cells / well, 300 wells per condition) with or without MSX (50 μM). 6.8.2 Example 2: Recombinant Protein Expression Analysis

[0410] Biomembrane Interference (BLI) Analysis: Twenty days post-inoculation, the supernatant from the cell pool was diluted 5-fold and analyzed using the Octet Qke label-free system and Octet ProA biosensor (Sartorius, catalog number No. 18-5010) according to the manufacturer's instructions. Figure 2 As shown, denosumab expression from all positive pools (denosumab expression >0.125 μg / mL) or the first 30 pools using a specified GS is plotted. When conventional Chinese hamster GS was used as an optional biomarker, only 4.3% of the 300 inoculated pools (13) expressed detectable denosumab (>0.125 μg / mL) on day 20, with a median expression of 1.07 μg / mL. Conversely, when GS from koalas, ostriches, snakes, pigeons, junglefowl, gerbils, or night bats were used as alternative biomarkers, 85.6% (257), 65.3% (196), 38.0% (114), 52.0% (156), 73.6% (221), 63.6% (191), or 94.0% (282) of the pools expressed denosumab (>0.125 μg / mL), with median expression values ​​of 2.86 μg / mL, 2.57 μg / mL, 1.75 μg / mL, 2.27 μg / mL, 2.51 μg / mL, 2.67 μg / mL, and 3.39 μg / mL, respectively.

[0411] Accordingly, Octet-based recombinant protein expression analysis confirmed that using koala, ostrich, snake, pigeon, junglefowl, gerbil, or night bat GS as selectable markers significantly improved the screening efficiency and yield of recombinant proteins.

[0412] Figure 2The table lists the number of positive pools with specified ranges of denosumab expression. Compared to using Chinese hamster GS as a candidate biomarker, using koala, ostrich, snake, pigeon, junglefowl, gerbil, or night bat GS as candidate biomarkers identified more positive pools with higher titers. Therefore, by screening a small number of starting pools using these newly identified candidate biomarkers, candidates with higher titers can be obtained.

[0413] Flow cytometry analysis of GS expression: Approximately 3 weeks post-inoculation, cell pools electroporated with the GS gene from the specified species were analyzed using an Attune NxT flow cytometer. Figures 3A-3C The 2A self-cleaving peptide between the green fluorescent proteins UnaG and GS leads to equimolar expression of both proteins, allowing direct comparison of GS expression by monitoring UnaG fluorescence signals. As shown, when using GS from Chinese hamsters as a selectable marker, the majority of test cells (>98%) remained UnaG “negative” in the presence of MSX, indicating low selection efficiency. Unexpectedly, when using GS from koalas, ostriches, snakes, pigeons, junglefowl, gerbils, or night bats… Figures 3A-3B When used as a selectable biomarker, a large number of cells showed positive UnaG expression, thus confirming that GS biomarkers from these species significantly improved MSX selection efficiency compared to traditional Chinese hamster GS. Furthermore, when GS biomarkers from these species were used as selectable biomarkers, the UnaG signal intensity was 10-100 times stronger. Figure 3C This indicates that GS expression (proportional to recombinant POI expression) is also 10-100 times higher.

Claims

1. Use of a nucleotide sequence encoding glutamine synthase (GS) as an optional marker, wherein the GS is koala GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS.

2. The use according to claim 1, wherein the GS is derived from the koala GS of the Phascolarctidae family.

3. The use according to claim 2, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO:

1.

4. The use according to claim 2 or 3, wherein the GS has reduced activity compared to wild-type koala GS.

5. The use according to claim 1, wherein the GS has the amino acid sequence shown in SEQ ID NO:

1.

6. The use according to claim 1, wherein the GS is an ostrich GS derived from the Struthionidae family.

7. The use according to claim 6, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO:

2.

8. The use according to claim 6 or 7, wherein the GS has reduced activity compared to wild-type ostrich GS.

9. The use according to claim 1, wherein the GS has the amino acid sequence shown in SEQ ID NO:

2.

10. The use according to claim 1, wherein the GS is a snake GS derived from the Elapidae family.

11. The use according to claim 10, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO:

3.

12. The use according to claim 10 or 11, wherein the GS has reduced activity compared to wild-type snake GS.

13. The use according to claim 1, wherein the GS has the amino acid sequence shown in SEQ ID NO:

3.

14. The use according to claim 1, wherein the GS is a pigeon GS derived from the Columbidae family.

15. The use according to claim 14, wherein the GS has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO:

5.

16. The use according to claim 14 or 15, wherein the GS has reduced activity compared to wild-type pigeon GS.

17. The use according to claim 1, wherein the GS has the amino acid sequence shown in SEQ ID NO:

5.

18. The use according to claim 1, wherein the GS is derived from the red junglefowl of the Pheasantidae family.

19. The use according to claim 18, wherein the GS has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO:

6.

20. The use according to claim 18 or 19, wherein the GS has reduced activity compared to wild-type junglefowl GS.

21. The use according to claim 1, wherein the GS has the amino acid sequence shown in SEQ ID NO:

6.

22. The use according to claim 1, wherein the GS is derived from the gerbil (Muridae).

23. The use according to claim 22, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO:

7.

24. The use according to claim 22 or 23, wherein the GS has reduced activity compared to wild-type gerbil GS.

25. The use according to claim 1, wherein the GS has the amino acid sequence shown in SEQ ID NO:

7.

26. The use according to claim 1, wherein the GS is a nocturnal bat GS derived from the Vespertilionidae family.

27. The use according to claim 26, wherein the GS has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO:

8.

28. The use according to claim 26 or 27, wherein the GS has reduced activity compared to wild-type night bat GS.

29. The use according to claim 1, wherein the GS has the amino acid sequence shown in SEQ ID NO:

8.

30. The use according to claim 1, wherein the amino acid sequence of said GS comprises SEQ ID NO: 9 having at least one amino acid substitution at a position selected from the following: A2, H8, N10, G12, Q15, M16, M18, S19, E24, Q27, V33, G39, D48, C49, C53, V54, E55, E56, F68, S70, S72, S80, V82, M84, E92, V97, F98, F102, Q106, K107, P108, E110, H115, T116, K118, D122, S125, H128, L140, D152, L160, K169, R172, I175, M176, V188, K189, T191, Y194 V197, K198, H199, A200, I206, I212, R213, V220, K230, V234, A236, T237, S240, T260, E264, N265, H269, K271, E272, A273, K276, S278, R282, L294, H304, K305, N308, N310, D311, D318, S320, T328, Q331, E332, K334, F337, A339, C341, F349, I355, V356, D366, Q367, and Q370.

31. The use according to claim 1, wherein the amino acid sequence of said GS comprises SEQ ID NO: 9 having at least one amino acid substitution at a position selected from the following: N10, G12, Q15, S19, E24, V33, G39, C49, C53, V54, E56, S72, S80, V82, E92, F98, Q106, K107, P108, K118, L140, D152, L160, R172, M176, T191, Y194, K198, H19 9. I206, R213, V220, K230, A236, T237, S240, T260, E264, N265, H269, K271, R282, K305, N310, D311, D318, S320, T328, Q331, A339, C341, F349, I355, Q367 and Q370.

32. The use according to claim 31, wherein the amino acid sequence of said GS comprises SEQ ID NO: 9: N10, G12, Q15, S19, E24, V33, G39, C49, C53, V54, E56, S72, S80, V82, E92, F98, Q106, K107, P108, K118, L140, D152, L160, R172, M176, T191, Y194, K198, H19 at positions selected from the group consisting of about 3, about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, or about 55 amino acid substitutions at the following positions.

9. I206, R213, V220, K230, A236, T237, S240, T260, E264, N265, H269, K271, R282, K305, N310, D311, D318, S320, T328, Q331, A339, C341, F349, I355, Q367 and Q370.

33. Use of a nucleotide sequence encoding a GS as an optional marker, wherein the GS comprises a catalytic domain derived from koala GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS.

34. The use according to claim 33, wherein the GS comprises a catalytic domain from a koala GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO:

1.

35. The use according to claim 34, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO:

1.

36. The use according to claim 33, wherein the GS comprises a catalytic domain from an ostrich GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO:

2.

37. The use according to claim 36, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO:

2.

38. The use according to claim 33, wherein the GS comprises a catalytic domain from a snake GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO:

3.

39. The use according to claim 38, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO:

3.

40. The use according to claim 33, wherein the GS comprises a catalytic domain from a pigeon GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO:

5.

41. The use according to claim 40, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 163-412 shown in SEQ ID NO:

5.

42. The use according to claim 33, wherein the GS comprises a catalytic domain from a wild chicken GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO:

6.

43. The use according to claim 42, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 150-399 shown in SEQ ID NO:

6.

44. The use according to claim 33, wherein the GS comprises a catalytic domain from a gerbil GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO:

7.

45. The use according to claim 44, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO:

7.

46. ​​The use according to claim 33, wherein the GS comprises a catalytic domain from a night-bat GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO:

8.

47. The use according to claim 46, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO:

8.

48. The use according to any one of claims 1 to 47, wherein the GS-encoded nucleotide sequence is operatively linked to an mRNA destabilizing element.

49. The use according to any one of claims 1 to 47, wherein the GS comprises a degradation determinant.

50. The use according to claim 49, wherein the degradation determinant has an amino acid sequence selected from SEQ ID NO: 23-25.

51. The use according to any one of claims 1 to 50, for identifying genomic loci with high transcriptional activity.

52. The use according to any one of claims 1 to 50, for identifying host cells capable of producing the target protein (POI).

53. The use according to any one of claims 1 to 50, for the recombinant production of POI.

54. The use according to claim 53, for the production of recombinant proteins in mammalian cells.

55. The use according to claim 54, wherein the mammalian cells are Chinese hamster ovary (CHO) cells.

56. The use according to any one of claims 52 to 55, wherein the POI is selected from antibodies, enzymes, soluble proteins, secretory proteins, membrane proteins, and fusion proteins.

57. A deoxyribonucleic acid (DNA) vector suitable for recombinant protein production or genome integration, comprising a nucleotide sequence encoding a GS (GS coding sequence), wherein the GS is a koala GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS.

58. The carrier according to claim 57, wherein the GS is derived from the koala GS of the Phascolarctidae family.

59. The vector according to claim 58, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO:

1.

60. The carrier according to claim 58 or 59, wherein the GS has reduced activity compared to wild-type koala GS.

61. The vector according to claim 57, wherein the GS has the amino acid sequence shown in SEQ ID NO:

1.

62. The carrier according to any one of claims 58 to 61, wherein the GS encoded sequence has at least 80% identity with SEQ ID NO:

10.

63. The carrier according to claim 57, wherein the GS is derived from ostrich GS of the Struthionidae family.

64. The vector according to claim 63, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO:

2.

65. The carrier according to claim 63 or 64, wherein the GS has reduced activity compared to wild-type ostrich GS.

66. The vector according to claim 57, wherein the GS has the amino acid sequence shown in SEQ ID NO:

2.

67. The vector according to any one of claims 63 to 66, wherein the GS encoded sequence has at least 80% identity with SEQ ID NO:

11.

68. The carrier according to claim 57, wherein the GS is derived from a snake GS of the Elapidae family.

69. The vector according to claim 68, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO:

3.

70. The carrier according to claim 68 or 69, wherein the GS has reduced activity compared to wild-type snake GS.

71. The vector according to claim 57, wherein the GS has the amino acid sequence shown in SEQ ID NO:

3.

72. The vector according to any one of claims 68 to 71, wherein the GS encoded sequence has at least 80% identity with SEQ ID NO:

12.

73. The carrier according to claim 57, wherein the GS is derived from pigeon GS of the Columbidae family.

74. The vector according to claim 73, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO:

5.

75. The carrier according to claim 73 or 74, wherein the GS has reduced activity compared to wild-type pigeon GS.

76. The vector according to claim 57, wherein the GS has the amino acid sequence shown in SEQ ID NO:

5.

77. The vector according to any one of claims 73 to 76, wherein the GS encoded sequence has at least 80% identity with SEQ ID NO:

14.

78. The carrier according to claim 57, wherein the GS is derived from the red junglefowl GS of the Pheasantidae family.

79. The vector according to claim 78, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO:

6.

80. The carrier according to claim 78 or 79, wherein the GS has reduced activity compared to wild-type junglefowl GS.

81. The vector according to claim 57, wherein the GS has the amino acid sequence shown in SEQ ID NO:

6.

82. The vector according to any one of claims 78 to 81, wherein the GS encoded sequence has at least 80% identity with SEQ ID NO:

15.

83. The carrier according to claim 57, wherein the GS is derived from the gerbil GS of the Muridae family.

84. The vector according to claim 83, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO:

7.

85. The carrier according to claim 83 or 84, wherein the GS has reduced activity compared to wild-type gerbil GS.

86. The vector according to claim 57, wherein the GS has the amino acid sequence shown in SEQ ID NO:

7.

87. The vector according to any one of claims 83 to 86, wherein the GS encoded sequence has at least 80% identity with SEQ ID NO:

16.

88. The carrier according to claim 57, wherein the GS is derived from the nocturnal bat GS of the Vespertilionidae family.

89. The vector according to claim 88, wherein the GS has an amino acid sequence that is at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO:

8.

90. The carrier according to claim 88 or 89, wherein the GS has reduced activity compared to wild-type night bat GS.

91. The vector according to claim 57, wherein the GS has the amino acid sequence shown in SEQ ID NO:

8.

92. The vector according to any one of claims 88 to 91, wherein the GS encoded sequence has at least 80% identity with SEQ ID NO:

17.

93. The vector according to claim 57, wherein the amino acid sequence of the GS comprises SEQ ID NO: 9 having at least one amino acid substitution at one of the following positions: A2, H8, N10, G12, Q15, M16, M18, S19, E24, Q27, V33, G39, D48, C49, C53, V54, E55, E56, F68, S70, S72, S80, V82, M84, E92, V97, F98, F102, Q106, K107, P108, E110, H115, T116, K118, D122, S125, H128, L140, D152, L160, K169, R172, I175, M176, V188, K189, T191, Y194 V197, K198, H199, A200, I206, I212, R213, V220, K230, V234, A236, T237, S240, T260, E264, N265, H269, K271, E272, A273, K276, S278, R282, L294, H304, K305, N308, N310, D311, D318, S320, T328, Q331, E332, K334, F337, A339, C341, F349, I355, V356, D366, Q367, and Q370.

94. The vector according to claim 57, wherein the amino acid sequence of the GS comprises SEQ ID NO: 9 having at least one amino acid substitution at the following positions: N10, G12, Q15, S19, E24, V33, G39, C49, C53, V54, E56, S72, S80, V82, E92, F98, Q106, K107, P108, K118, L140, D152, L160, R172, M176, T191, Y194, K198, H19 9. I206, R213, V220, K230, A236, T237, S240, T260, E264, N265, H269, K271, R282, K305, N310, D311, D318, S320, T328, Q331, A339, C341, F349, I355, Q367 and Q370.

95. The vector according to claim 94, wherein the amino acid sequence of said GS comprises SEQ ID NO: 9: N10, G12, Q15, S19, E24, V33, G39, C49, C53, V54, E56, S72, S80, V82, E92, F98, Q106, K107, P108, K118, L140, D152, L160, R172, M176, T191, Y194, K198, H19 at positions selected from the group consisting of about 3, about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, or about 55 amino acid substitutions at the following positions.

9. I206, R213, V220, K230, A236, T237, S240, T260, E264, N265, H269, K271, R282, K305, N310, D311, D318, S320, T328, Q331, A339, C341, F349, I355, Q367 and Q370.

96. A DNA vector suitable for recombinant protein production or genome integration, comprising a nucleotide sequence encoding a GS (GS coding sequence), wherein the GS comprises a catalytic domain derived from koala GS, ostrich GS, snake GS, pigeon GS, junglefowl GS, gerbil GS, or night bat GS.

97. The carrier according to claim 96, wherein the GS comprises a catalytic domain from a koala GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO:

1.

98. The carrier according to claim 97, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO:

1.

99. The carrier according to claim 96, wherein the GS comprises a catalytic domain from an ostrich GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO:

2.

100. The carrier according to claim 99, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO:

2.

101. The carrier according to claim 96, wherein the GS comprises a catalytic domain from a snake GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO:

3.

102. The carrier according to claim 101, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO:

3.

103. The carrier according to claim 96, wherein the GS comprises a catalytic domain from pigeon GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO:

5.

104. The carrier according to claim 103, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 163-412 shown in SEQ ID NO:

5.

105. The carrier according to claim 96, wherein the GS comprises a catalytic domain from a wild chicken GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO:

6.

106. The carrier according to claim 105, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 150-399 shown in SEQ ID NO:

6.

107. The carrier according to claim 96, wherein the GS comprises a catalytic domain from a gerbil GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO:

7.

108. The carrier according to claim 107, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO:

7.

109. The carrier according to claim 96, wherein the GS comprises a catalytic domain from a Night Bat GS having an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with SEQ ID NO:

8.

110. The carrier according to claim 109, wherein the catalytic domain has an amino acid sequence having at least 90%, at least 92%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acids 110-359 shown in SEQ ID NO:

8.

111. The vector according to any one of claims 57 to 110, wherein the GS coding sequence is operatively linked to an mRNA destabilizing element.

112. The carrier according to any one of claims 57 to 110, wherein the GS comprises a degradation determinant.

113. The carrier according to claim 112, wherein the degradation determinant has an amino acid sequence selected from SEQ ID NO: 23-25.

114. The vector according to any one of claims 57 to 113, wherein the vector is suitable for recombinant protein production and further comprises an expression cassette.

115. The vector of claim 114, wherein the GS encoding sequence is operatively linked to the simian cavitation virus 40 (SV40) promoter.

116. The carrier according to claim 114 or 115, wherein the GS encoded sequence is operatively linked to a poly(A) tail.

117. The carrier according to any one of claims 114 to 116, comprising two or more expression cassettes.

118. The vector according to any one of claims 114 to 117, wherein the expression cassette comprises a nucleotide sequence encoding a POI (POI coding sequence).

119. The vector according to claim 118, wherein the POI is an antibody, enzyme, soluble protein, secretory protein, membrane protein, or fusion protein.

120. The vector according to claim 119, wherein the POI is an antibody selected from IgG1 antibody, IgG2 antibody, IgG3 antibody, IgG4 antibody, IgA antibody, IgM antibody, Fab, Fab', F(ab')2, Fv, scFv, (scFv)2, single-domain antibody (sdAb), single-chain antibody (scAb) and heavy-chain antibody (HCAb).

121. The vector according to claim 119, wherein the POI is an antibody selected from the group consisting of monoclonal antibodies, bispecific antibodies, multispecific antibodies, bivalent antibodies, and multivalent antibodies.

122. The vector according to any one of claims 118 to 121, wherein the POI consists of one or more copies of the same polypeptide.

123. The carrier according to any one of claims 118 to 121, wherein the POI comprises two different polypeptides.

124. The vector according to claim 123, wherein the POI is an antibody comprising a light chain and a heavy chain, each encoded by a separate nucleotide sequence on the vector.

125. The vector according to any one of claims 57 to 124, wherein the vector is suitable for genome integration.

126. Use of the vector according to any one of claims 118 to 124 for identifying host cells capable of producing POI.

127. The use of the vector according to claim 125 for identifying genomic loci with high transcriptional activity.

128. A method for identifying host cells capable of producing POI, comprising introducing a vector according to any one of claims 118 to 124 into a population of host cells, culturing the population of host cells in a glutamine-free medium, wherein host cells capable of growing in the medium are identified as host cells capable of producing POI.

129. A method for identifying genomic loci with high transcriptional activity, comprising introducing the vector of claim 125 into a population of host cells, culturing the population of host cells in a glutamine-free medium, wherein host cells capable of growing in said medium are identified as host cells having a GS coding sequence inserted at a genomic locus with high transcriptional activity.

130. The method of claim 129 further comprises sequencing the genome of the identified host cell to locate genomic loci with high transcriptional activity.

131. The method according to any one of claims 128 to 130, wherein the host cell population is cultured in the presence of a GS inhibitor.

132. The method of claim 131, wherein the GS inhibitor is methionine sulfoxide (MSX).

133. A host cell comprising the vector according to any one of claims 118 to 124.

134. The host cell according to claim 133, having wild-type endogenous GS.

135. The host cell of claim 133, wherein the endogenous GS of the host cell has reduced activity or is knocked out.

136. The host cell according to any one of claims 133 to 135, wherein the host cell is a mammalian cell.

137. The host cell according to claim 136, wherein it is a CHO cell.

138. Use of the host cell according to any one of claims 133 to 137 for the in vitro production of POI.

139. A method for producing POI in vitro, comprising culturing a host cell according to any one of claims 133 to 137 for a period of time under certain conditions sufficient to produce POI.

140. A method for producing POI in vitro, comprising replacing the GS coding sequence with the POI coding sequence in a host cell identified in the method of claim 129, and culturing the host cell under certain conditions for a sufficient time to produce the POI.

141. The method of claim 139 or 140, further comprising separating the POI from other components in the culture.

142. The method of claim 141, wherein the separation comprises extraction, continuous liquid-liquid extraction, pervaporation, membrane filtration, membrane separation, reverse osmosis, electrodialysis, distillation, crystallization, centrifugation, extractive filtration, ion exchange chromatography, adsorption chromatography, or ultrafiltration.

143. An expression system for in vitro production of POIs, comprising a DNA vector and a host cell according to any one of claims 57 to 125.

144. The expression system of claim 143, wherein the host cell is a CHO cell.

145. The expression system according to claim 143 or 144 further comprises a glutamine-free culture medium.

146. The expression system according to any one of claims 143 to 145 further comprises a GS inhibitor.

147. The expression system according to any one of claims 143 to 146, further comprising a manner of introducing the vector into the host cell.

148. The expression system according to any one of claims 143 to 147, which is included in the kit.

Citation Information

Patent Citations

  • A system for inducible expression in eukaryotic cells

    WO2002088346A2