Compositions, constructs, cells and methods for increased recombinant protein expression by site-specific integration
Patent Information
- Application Number
- EP2022832293
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-10-12
- Filing Date
- 2022-06-28
- Publication Date
- 2025-07-16
AI Technical Summary
Current methods for recombinant protein expression in mammalian cells are time-consuming, labor-intensive, and inefficient due to random genomic integration, leading to low frequencies of high-level protein expression and unstable cell lines, with risks associated with viral vectors like lentivirus and retroviruses.
Site-specific integration using genomic amplification drop sites (GADS) in mammalian cells for targeted insertion and amplification of exogenous nucleotide sequences, including selection markers and protein-coding genes, to achieve stable and high-level recombinant protein production.
This approach results in stable, high-producing cell lines with consistent productivity over long periods, reducing the need for extensive screening and minimizing risks associated with viral vectors, while enabling rapid and efficient production of recombinant proteins.
Smart Images

Figure 1.1
Abstract
Description
COMPOSITIONS. CONSTRUCTS. CELLS AND METHODS FOR INCREASEDRECOMBINANT PROTEIN EXPRESSION BY SITE-SPECIFIC INTEGRATIONCROSS REFERENCE TO RELATED APPLICATIONS[1] This application claims priority to U.S. Provisional Patent Application Serial Number 63 / 215,876 filed June 28, 2021 and entitled “COMPOSITIONS, CONSTRUCTS, CELLS AND METHODS FOR INCREASED RECOMBINANT PROTEIN EXPRESSION BY SITE SPECIFIC INTEGRATION”, together with U.S. Provisional Patent Application Serial Number 63 / 254,997 filed October 12, 2021 and entitled “COMPOSITIONS, CONSTRUCTS, CELLS AND METHODS FOR CELL THERAPY”, the entire disclosure of each of which is hereby incorporated herein by reference.FIELD OF THE DISCLOSURE[2] The present disclosure relates to gene expression, and, in particular, to compositions, constructs, cells and methods for increased recombinant protein expression by site-specific integration.BACKGROUND[3] Approximately 20 to 30% of new drugs approved by the USFDA in recent years are biologies (in 2019 60 biologies were approved by CDER and CBER), up from 1 approval in 2000 (Batta et al., J Family Med Prim Care. 2020 Jan; 9(1): 105-114). There are currently more than 350 therapeutic biologies on the market and over 900 biologies in development. Development of expression systems for the efficient production of recombinant proteins is important for providing a source of a given protein for research or therapeutic use. The increased number of biologies in development has driven the need to develop simple and rapid high-output technologies for the development of recombinant protein expressing cell lines. The generation of commercial cell lines using conventional methods is a time-consuming, labor-intensive and repetitive process. The instant disclosure provides novel methods and materials that solve or ameliorate many of these significant problems.[4] Expression systems have been developed for both prokaryotic cells and for eukaryotic cells, which include yeast, Pichia pastoris, insect and mammalian cells. Expression in mammalian cells, for example Chinese hamster ovary (or "CHO") cells, is often preferred for themanufacture of therapeutic proteins, since post-translational modifications in such expression systems are more likely to resemble those found in human cells expressing proteins than the type of post-translational modifications that occur in microbial (prokaryotic) expression systems. Human cell lines like HEK293, HT1080, Per.C6, and other well-known cell lines are even more preferred. One skilled in the art can also develop means of using pluripotent, induced pluripotent, totipotent and adult stem cell lines in certain embodiments of the disclosure.[5] Recombinant expression plasmids comprising a gene of interest that encodes all or a portion of a desired protein are routinely used to generate CHO cells expressing the desired recombinant protein. These recombinant plasmids randomly integrate into the genome of the host producing recombinant proteins but the frequency of cell lines carrying the stably integrated recombinant gene that are capable of expressing a desired recombinant protein at high levels is extremely low. A large number of transfected mammalian cell lines must be screened to identify clones which express the recombinant proteins at high levels. During the construction and selection of protein-producing cells lines, cell lines with a large range of expression, growth and stability profiles are obtained. These variations can arise due to the inherent plasticity of the mammalian genome. They can also originate from stochastic gene regulation networks or in variation in the amount of recombinant protein produced resulting from random genomic integration of a transgene principally due to the "position variegation effect" or merely from the plasmid copy number, especially in light of the large size of the mammalian genome and the fact that only a small percentage of the genomic DNA contains transcriptionally active sequences.[6] As a consequence of these variations and the low (perhaps 1 in 10,000) frequency of genomic integration, resource-intensive and time-consuming efforts are required to screen many transfectants in the pool for these rare events, in order to isolate a commercially compatible production cell line (e. g., a combination of good growth, high productivity and stability of production, with desired product profile).[7] Expression augmenting sequences have been disclosed to increase expression of recombinant protein for eukaryotic expression systems (see for example WO 97 / 25420). An increase in the frequency of high-level recombinant gene expressing cell lines would provide a much greater pool of high protein expressing cell lines to choose from. This task can be accomplished by generating homologous recombinant plasmids targeted totranscriptionally active sites as disclosed herein and by devising a means to select for such cell lines.[8] There are a number of well-known and common amplification methods used to improve the yield of protein expression in mammalian cell systems. Amplification of the dihydrofolate reductase gene (Dhfr) by methotrexate (Mtx) exposure is commonly used for recombinant protein expression in Chinese hamster ovary (CHO) cells. However, this method is both time- and labor-intensive, and the cells that are generated are frequently unstable in culture. Further, the DHFR / MTX system has a long development time (taking up to 6 months), and the selection pressure is indirect.[9] Another common amplification system is the GS / MSX system. The glutamine synthetase (GS) expression system has been used for decades but has numerous issue. L-Methionine sulfoximine (MSX) inhibits the activity of glutamine synthetase, an enzyme essential for the production of glutamine. MSX is used as a media supplement to aid selection and amplification processes in recombinant mammalian cell lines that use GS as a selective marker. However, this system also has a long development time, provides indirect selection pressure and the amplified gene to be expressed is even more unstable than the cell-lines developed by the DHFR / MTX system.
[0010] The failings of existing amplification systems are discussed in Joseph J. Priolal, Nathan Calzadillal, Martina Baumann2, Nicole Borth3, Christopher G. Tate4 and Michael J. Betenbaughl High-throughput screening and selection of mammalian cells for enhanced protein production Biotechnol. J. 2016, 11. pl-13. DOI 10.1002 / biot.201500579.
[0011] One of the most common methods of recombinant protein production involves transfection with one of a variety of viruses including lentivirus and retroviruses. See for example, Tandon, et al., Bio Protoc. 2018 November 5; 8(21): . doi:10.21769 / BioProtoc.3073. This approach has many issues:1) Safety issues: A very high titer of infectious viruses carrying the gene of interest is required in a packaging cell line thus necessitating a high-level biosafety lab. Technicians must be highly trained to make and use these viruses, because the viruses can easily infect such technicians. The viruses may cause cancer or other problems because of virus integration triggered mutagenesis. Also, they have to wear special safety clothes, gloves, goggles etc. 17% of the human genome consists of long interspersed repeat element (“LINE”) sequences. These are hypothesized to be the remnants of previous virus infections. These viruses infected us and then some of themrandomly integrated into the human genome. These virus sequences are a permanent part of the human genome and may even constitutively produce virus proteins such as reverse transcriptase. Reverse-transcriptase transcribes an RNA sequence and makes the DNA sequence from it in a reversal of the usual protein production process. This DNA sequence might then integrate into the human genome and cause mutations. Even if defected lentivirus vectors are used, such vectors might recombine with these pre existing LINE sequences and result in an active virus. The use of virus vectors even for making cell lines poses potentially dangerous risks.2) Purification issues. Proteins purified from such viral systems require time consuming additional purification steps in order to ensure that the protein does not contain virus or viral toxin contaminants.3) Cell line stability. During the process of creating cell lines with a virus, the lentivirus integrates into the target genome randomly at hundreds or even thousands of sites. This random integration causes mutations in the cell line which can be advantageous for protein production in the short term but in the long term these mutations are harmful for the cell line. These methods are most useful for short term protein production. Over a longer term these randomly integrated viruses cease to produce proteins due to silencing. There is an evolutionary mechanism to silence integrated viruses. That is thought to be a natural mechanism our cells use to keep the LINE sequences at bay.
[0012] To achieve a long-term supply of a recombinant protein, a random integration system is not optimal. Embodiments of the present disclosure provide more stable protein production in a safer, faster, and easier system.
[0013] The technical problem underlying the present disclosure is to overcome the above- identified disadvantages, in particular to provide, preferably in a safe, simple and efficient manner, high producing cell lines with a high stability and positive growth and productivity characteristics, in particular cell lines which provide a consistent productivity over a long cultivation and production period.
[0014] In particular, the present disclosure solves or ameliorates much of these technical problems by the provision of a site-specific integration (SSI) host cell comprising an endogenous genomic amplification drop site(s) (GADS, which may also be referred to as genomic super expression drop site(s) or genomic overexpression drop site(s)), wherein an exogenous nucleotide sequence is integrated at said GADS. In some embodiments, the exogenous nucleotide sequence comprises at least one gene coding sequence of interest. In someembodiments, the nucleotide sequence comprises at least one selection marker gene, one promoter sequence and one enhancer sequence.
[0015] This background information is provided to reveal information believed by the applicant to be of possible relevance. No admission is necessarily intended, nor should be construed, that any of the preceding information constitutes prior art or forms part of the general common knowledge in the relevant art.SUMMARY
[0016] The following presents a simplified summary of the general inventive concept(s) described herein to provide a basic understanding of some aspects of the disclosure. This summary is not an extensive overview of the disclosure. It is not intended to restrict key or critical elements of embodiments of the disclosure or to delineate their scope beyond that which is explicitly or implicitly described by the following description and claims.
[0017] In accordance with one broad aspect of the disclosure, a nucleic acid construct is provided. The nucleic acid construct comprises a DNA fragment capable of site-specific insertion into a region of open chromatin in mammalian cells on at least one chromosome at more than one site.
[0018] In one embodiment, the nucleic acid construct is amplified more than 5 times following selective pressure.
[0019] In one embodiment, the DNA fragment comprises the sequence of SEQ ID NO: 1 , SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO: 5 or a functional fragment thereof.
[0020] In one embodiment, the DNA fragment is insertable at more than 5 sites on two or more chromosomes.
[0021] In accordance with another broad aspect of the disclosure, an expression vector comprising the nucleic acid construct as described above is provided.
[0022] In one embodiment, the expression vector further comprises a selectable marker. In one particular embodiment, the selectable marker provides resistance to puromycin.
[0023] In one embodiment, the expression vector further comprises a nucleic acid sequence encoding all or a portion of a protein of interest.
[0024] In one embodiment, the protein of interest is a drug selected from the group consisting of adalimumab, atezolizumab, nivolumab, pembrolizumab, etanercept, trastuzumab, bevacizumab, rituximab, aflibercept, infliximab, ustekinumab, ranibizumab, proteins of tumor biology, proteins of food industry, proteins of animal health, proteins of ageing,proteins of genetic disorders, vaccines, viral-like particles (VLPs), single proteins, virus inhibitor proteins, or the like.
[0025] In accordance with another broad aspect of the disclosure, a mammalian host cell transformed with the expression vector as described above is provided.
[0026] In one embodiment, the expression vector further comprises a selectable marker.
[0027] In one embodiment, the mammalian host cell is a 9HEK293 cell, a HT1080 cell, a Per.C6 cell, a pluripotent cell, an induced pluripotent cell, a totipotent cell, an adult stem cell, or a primary cell.
[0028] In accordance with another broad aspect of the disclosure, there is provided a Chinese hamster ovary (CHO) cell transformed with the expression vector described above.
[0029] In accordance with another broad aspect of the disclosure, there is provided a mouse cell transformed with the expression vector described above.
[0030] In accordance with another broad aspect of the disclosure, there is provided a human cell transformed with the expression vector described above.
[0031] In accordance with another broad aspect of the disclosure, there is provided a method for obtaining a recombinant protein, which comprises culturing a transformed host cell under conditions promoting expression of the recombinant protein and recovering the recombinant protein. In particular, the transformed host cell may be any one of the CHO cell, the mouse cell or the human cell, as described above.
[0032] In accordance with various other broad aspect of the disclosure, there is provided various plasmids constructed in accordance with various plasmid maps, typically including at least the nucleic acid construct described above.
[0033] In some embodiments, the plasmid further comprises a transgene and a selectable marker that provides resistance to puromycin.
[0034] In accordance with a further broad aspect of the disclosure, there is provided a DNA vector capable of integration into a mammalian genome. The DNA vector comprises a DNA fragment capable of site-specific insertion into a region of open chromatin in mammalian cells, wherein the DNA fragment is the sequence of SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, or a functional fragment thereof; and at least one transgene encoding at least one protein of interest to be expressed by a mammalian cell after the DNA vector integrates into the mammalian genome of the mammalian cell.
[0035] In some embodiments, the at least one protein of interest is a drug selected from the group consisting of adalimumab, atezolizumab, nivolumab, pembrolizumab, etanercept, trastuzumab, bevacizumab, rituximab, aflibercept, infliximab, ustekinumab, ranibizumab, proteins of tumor biology, proteins of food industry, proteins of animal health, proteins of ageing, proteins of genetic disorders, vaccines, viral-like particles (VLPs), single proteins, virus inhibitor proteins or the like.
[0036] Other aspects, features and / or advantages will become more apparent upon reading of the following non-restrictive description of specific embodiments thereof, given by way of example only with reference to the accompanying drawings.BRIEF DESCRIPTION OF THE FIGURES
[0037] Several embodiments of the present disclosure will be provided, by way of examples only, with reference to the appended drawings, wherein:
[0038] FIGURE 1 shows a plasmid map of a recombinant insertion and expression vector designed to create cell lines that express trastuzumab (Herceptin, i.e. an insertion vector for trastuzumab heavy and light chain production).
[0039] FIGURE 2 shows a plasmid map of another recombinant insertion and expression vector designed to create cell lines that express trastuzumab (Herceptin, i.e. another insertion vector for trastuzumab heavy and light chain production).
[0040] FIGURE 3 shows a fluorescent in situ hybridization (FISH) image of a trastuzumab producing CHO-K1 cell line, which cell line was produced in less than two weeks, wherein bright dots show the presence of trastuzumab expressing gene copies and wherein the arrows show that amplification is clearly visible on two chromosomes.
[0041] FIGURE 4 shows an image of a Western blot of trastuzumab producing CHO-K1 cell lines and also includes actin as a control, wherein the signal at 55 kDa is the trastuzumab heavy chain and the signal at 42 kDa is actin (internal control for quantitation of protein expression with Licor Odyssey gel documentation system).
[0042] FIGURE 5 shows an image of a Western blot of HEK293 cells showing the comparison of trastuzumab heavy chain production (55kDa) and control actin production (42kDa). In addition to the controls, trastuzumab production is shown in lanes labelled 10 / 5, 10 / 6, 10 / 8, and 10 / 13, wherein each lane corresponds to particular clones.
[0043] FIGURE 6 shows an image of a Western blot showing purified trastuzumab production of both heavy (55 kDa) and light (25 kDa) chains from clone 10 / 13 of FIGURE 5, wherein purification was done by successive rounds of protein G (“SpinTrap”) purification.
[0044] FIGURE 7 shows an inverted image of a Western blot showing the results of subcloning of clone 10 / 13, wherein the signal at 55 kDa is the trastuzumab heavy chain and the signal at 42 kDa is actin (internal control for quantitation of protein expression with Licor Odyssey gel documentation system).
[0045] FIGURE 8 shows a FISH analysis image demonstrating amplification of clone 10 / 13 across a variety of human chromosomes, wherein stars show targeted sites and amplification of trastuzumab genes.
[0046] FIGURE 9 shows a plasmid map of a plasmid carrying GADS2 (SEQ ID NO:2) for transfecting, insertion and amplification.
[0047] FIGURE 10 shows a FISH analysis image of FISH-stained MEF cells transfected with a GADS2 carrying plasmid according to the disclosure and as shown in FIGURE 9, which in this embodiment carries the Influenza A virus Hemagglutinin MYMC X-181 California strain sequence as a useful transgene, wherein targeting exclusively happened into a large acrocentric chromosome into the upper part of the long chromosomal arm in all 34 clones.
[0048] FIGURE 11 shows a FISH analysis image of FISH-stained cells Hamster chromosomes stained with DAPI (blue signal), wherein the transgene was labeled with a green fluorescent dye. As shown, one hamster chromosome has an amplified chromosome arm with hundreds of transgene copies. As shown on the right side of the photo, a newly formed, autonomous mammalian artificial chromosome is visible, which almost entirely consists of transgene sequences in hundreds of copies (green signal).
[0049] FIGURE 12 shows a plasmid map of a plasmid construct according to the disclosure carrying the GADS1 sequence (CDC27 pseudogene) and the Influenza A virus Hemagglutinin MYMC X-181 California strain sequence as a useful transgene.
[0050] FIGURE 13 shows a FISH analysis image of FISH-stained cells, wherein mouse chromosomes were stained with DAPI (blue) and transgenes were stained with a green fluorescent dye, the figure specifically showing an autonomous mammalian chromosome close to the middle of the photo with hundreds of copies from the transgene, and more specifically showing two chromosomes with an amplified chromosome arm (30 copies and 60 copies from the transgenes are present). Cells were transfected with the construct shown in FIGURE 12.
[0051] FIGURE 14 shows a plasmid map of a pIKRBBP7 plasmid for insertion and amplification of CIRBBP.
[0052] FIGURE 15 shows Western Blots of CIRBBP cell-lines generated using the pIKRBBP7 plasmid of FIGURE 14, specifically showing expression of CIRBBP using an anti-AVI-tag antibody as a primary antibody and anti-mouse-HRP antibody as a secondary antibody in the Western blots. Western blots were developed by ECL and chemiluminescent signal was photographed.
[0053] Elements in the several figures are illustrated for simplicity and clarity and have not necessarily been drawn to scale. For example, the dimensions of some of the elements in the figures may be emphasized relative to other elements for facilitating understanding of the various presently disclosed embodiments. Also, common, but well-understood elements that are useful or necessary in commercially feasible embodiments are often not depicted in order to facilitate a less obstructed view of these various embodiments of the present disclosure.DETAILED DESCRIPTION
[0054] Various implementations and aspects of the specification will be described with reference to details discussed below. The following description and drawings are illustrative of the specification and are not to be construed as limiting the specification. Numerous specific details are described to provide a thorough understanding of various implementations of the present specification. However, in certain instances, well-known or conventional details are not described in order to provide a concise discussion of implementations of the present specification.
[0055] Various methods and processes will be described below to provide examples of implementations of the disclosure disclosed herein. No implementation described below limits any claimed implementation and any claimed implementations may cover processes or methods that differ from those described below. The claimed implementations are not limited to methods or processes having all of the features of any one method or process described below or to features common to multiple or all of the methods or processes described below. It is possible that a method or process described below is not an implementation of any claimed subject matter.
[0056] Furthermore, numerous specific details are set forth in order to provide a thorough understanding of the implementations described herein. However, it will be understood bythose skilled in the relevant arts that the implementations described herein may be practiced without these specific details. In other instances, well-known methods, procedures and components have not been described in detail so as not to obscure the implementations described herein.
[0057] It is understood that for the purpose of this specification, language of “at least one of X, Y, and Z” and “one or more of X, Y and Z” may be construed as X only, Y only, Z only, or any combination of two or more items X, Y, and Z (e.g., XYZ, XY, YZ, ZZ, and the like). Similar logic may be applied for two or more items in any occurrence of “at least one ...” and “one or more...” language.
[0058] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0059] As used herein, the term “transgene” is given its broadest possible meaning to include any gene that one wants to express regardless of the species or source of that gene. For example, a wild-type human gene or a mutant mouse gene could both be a transgene when used in the constructs and methods herein regardless of the species of the host cell, for that matter.
[0060] Throughout the specification and claims, the following terms take the meanings explicitly associated herein, unless the context clearly dictates otherwise. The phrase “in one of the embodiments” or “in at least one of the various embodiments” as used herein does not necessarily refer to the same embodiment, though it may. Furthermore, the phrase “in another embodiment” or “in some embodiments” as used herein does not necessarily refer to a different embodiment, although it may. Thus, as described below, various embodiments may be readily combined, without departing from the scope or spirit of the innovations disclosed herein. The same logic may apply to examples.
[0061] In addition, as used herein, the term “or” is an inclusive “or” operator, and is equivalent to the term “and / or,” unless the context clearly dictates otherwise. The term “based on” is not exclusive and allows for being based on additional factors not described, unless the context clearly dictates otherwise. In addition, throughout the specification, the meaning of singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. The meaning of "in" includes "in" and "on."
[0062] The term “comprising” as used herein will be understood to mean that the list following is non-exhaustive and may or may not include any other additional suitable items, for example one or more further feature(s), component(s) and / or element(s) as appropriate.
[0063] Novel insertion sequences, referred to herein as GADS (genomic amplification drop site(s)), that facilitate increased expression of recombinant proteins in mammalian host cells, are disclosed. A preferred embodiment of the disclosure is a GADS that was obtained from human cell genomic DNA. Vectors comprising such human derived GADS can be used for insertion into not only human but also mouse and hamster genomes as a result of the highly conserved nature of the GADS. Alternatively, one skilled in the art can easily obtain mouse or hamster GADS sequences and other mammalian cell lines useful for protein expression could be envisioned by one skilled in the art. Another preferred embodiment of the disclosure is a GADS that was obtained from hamster cell genomic DNA. Another preferred embodiment of the disclosure is a GADS that was obtained from mouse cell genomic DNA.
[0064] The present disclosure discloses a GADS sequence from a genomic locus in the human genome that is capable of high recombinant gene amplification and expression. In a most preferred embodiment of the disclosure, the GADS is selected from the group consisting of (a) DNAs comprising nucleotides of SEQ ID NO:l; (b) fragments of SEQ ID NO:l that are useful as insertion and expression sites; (c) nucleotide sequences complementary to (a) and / or (b); (d) nucleotide sequences that are at least about 80%, more preferably about 90%, and more preferably about 95% identical in nucleotide sequence to (a), (b) and / or (c) and that are useful for insertion and expression of exogenous proteins; and (e) combinations of the foregoing nucleic acid sequences that are useful for insertion and expression of exogenous proteins.
[0065] Expression vectors comprising the novel GADS sequences are able to transform CHO, HEK293, or other mammalian cells to increase expression of recombinant proteins through genomic insertion into open chromatin and amplification. Thus, another embodiment of the disclosure is an expression vector comprising a GADS sequence. In a preferred embodiment, the expression vector further comprises a eukaryotic promoter / enhancer driving the expression of all or a portion of a protein of interest. Two or more different nucleic acids expressing exogenous proteins of interest can be present in an expression vector used to transfect a cell (e.g., CHO or HEK293 cells), wherein each nucleic acid sequence encodes a different polypeptide that assemble (when expressed) to form a desired protein. In an additional preferred embodiment, the expression vector comprises a plasmid that encodes a gene of interest and also encodes an amplifiable dominant selectable marker.A preferred marker is puromycin; other amplifiable markers known in the art are also suitable for use in certain embodiments of the expression vectors of the instant disclosure.
[0066] Mammalian host cells can be transformed with an expression vector of the present disclosure to produce high levels of recombinant protein. Accordingly, another embodiment of the disclosure provides a mammalian host cell transformed with an expression vector of the present disclosure. Also within the scope of the present disclosure are mammalian host cells transformed with two expression vectors, wherein each of the two expression vectors encodes at least one polypeptide subunit that when co-expressed assembles into a desired protein with biological activity. In a most preferred embodiment, the host cells are CHO or HEK cells.
[0067] The disclosure also provides a method for obtaining a recombinant protein, comprising transforming a host cell with an expression vector of the present disclosure, culturing the transformed host cell under conditions promoting amplification of the inserted exogenous vector and expression of the protein, and recovering the protein. In a preferred application of the disclosure, transformed host cells are selected with multiple selection steps with increasing concentrations of a selection antibiotic such as puromycin. Embodiments of this method are useful for creating an amplification and expression system which is tunable to the specific properties of different transgenes that one desires to produce.
[0068] In certain embodiments of the disclosure, the selectable antibiotic (preferably puromycin) concentration is increased in a series of steps to achieve the desired optimal amplification and expression level in each cell type (preferably human, mouse or hamster). The methods, constructs and systems disclosed herein also provide a large number and variety of individual clones with various levels of amplification providing different expression levels. The particular clones having optimal amplification levels can be selected to provide a stable cell line with a desired protein production level. One skilled in the art will understand that different proteins require different levels of production and the methods disclosed herein (as well as the constructs and uses) provide an adjustable combination of integration, amplification and protein production. One skilled in the art can use the methods and constructs disclosed herein to adjust the system for the best results for each individual desired protein product.
[0069] The instant disclosure therefore comprises methods, compositions and constructs useful for production of large amounts of recombinant proteins in cell lines that are stable for long periods of time. Embodiments of the disclosure include sequences and constructs forachieving non-random insertion of exogenous DNA sequences into the genome of mammalian cell lines followed by the amplification of the inserted DNA into multiple sites across numerous chromosomes. In one surprising aspect of the instant disclosure, the non- random insertion sites of the disclosure are substantially composed of euchromatin that are open and not silenced. Thus, the constructs of the instant disclosure are uniquely useful for the stable, long-term, large scale production of therapeutic proteins.
[0070] Embodiments of the instant disclosure include methods for achieving rapid amplification of the inserted DNA across the genomes of the transformed cells.
[0071] Certain embodiments of the recombinant expression vectors of the instant disclosure include novel sequences for achieving homologous recombination at specific sites in the genomes of mammalian cells.
[0072] Recombinant expression vectors include synthetic or cDNA-derived DNA fragments encoding a protein, operably linked to suitable transcriptional or translational regulatory elements derived from mammalian, viral, fungi, insect or bacterial genes. Such regulatory elements may include a transcriptional promoter, a sequence encoding suitable mRNA ribosomal binding sites, and sequences which control the termination of transcription and translation. Mammalian expression vectors may also comprise non-transcribed elements such as an origin of replication, a suitable promoter and enhancer linked to the gene to be expressed, other 5' or 3' flanking non-transcribed sequences, 5' or 3' non-translated sequences such as ribosome binding sites, a polyadenylation site, and transcriptional termination sequences. An origin of replication that confers the ability to replicate in a host, and a selectable gene to facilitate recognition of transformants, may also be incorporated. A preferred expression vector is shown in Figure 1.
[0073] DNA regions are operatively linked when they are functionally related to each other. For example, a promoter is operatively linked to a coding sequence if it controls the transcription of the sequence; or a ribosome binding site is operatively linked to a coding sequence if it is positioned so as to permit translation. Transcriptional and translational control sequences in expression vectors used in transforming cells are known in the art.
[0074] Transformed host cells are cells which have been transformed or transfected with expression vectors constructed using recombinant DNA techniques and which contain sequences encoding all or a portion of recombinant proteins. Expressed proteins may be secreted into the cell culture supernatant, depending on the DNA selected, but may also be deposited inside the cell and / or in the cell membrane. Various mammalian cell culturesystems can be employed to express recombinant protein according to embodiments of the present disclosure, all well known in the art, for example COS lines of monkey kidney cells, CHO cells, HeLa cells, HEK293 cells, Per.C6 cells, LMTK- cells, pluripotent cells, induced pluripotent cells, totipotent cells, adult stem cells, primary cells and BHK cell lines.
[0075] Several transformation protocols are known in the art, and are reviewed, for example, in Kaufman et. al., (1988) Meth. Enzymology 185:537. The transformation protocol chosen will depend on the host cell type and the nature of the gene of interest and can be chosen based upon routine experimentation. The basic requirements of any such protocol are first to introduce DNA encoding a protein of interest into a suitable host cell, and then to identify and isolate host cells which have incorporated the DNA in stable, expressible manner. Examples of methods useful for introducing DNA encoding a protein of interest can be found in Wigler et.al., (1980) Proc. Natl. Acad. Sci. USA 77:3567; Schaffner (1980) Proc. Natl. Acad. Sci. USA 77:2163; Potter et.al, (1988) Proc. Natl. Acad. Sci. USA 81:7161; and Shigekawa (1988) BioTechniques 6:742.
[0076] A method of amplifying the gene of interest is also desirable for expression of the recombinant protein, and typically involves the use of a selection marker. The novel characteristics of the instant homologous recombination vectors are ideal for amplification using a selection marker. Resistance to cytotoxic drugs is the characteristic most frequently used as a selection marker and can be the result of either a dominant trait (i.e., can be used independent of host cell type) or a recessive trait (i.e., useful in particular host cell types that are deficient in whatever activity is being selected for). Many amplifiable markers are suitable for use in the present disclosure (for example, as described in Maniatis, Molecular Biology: A Laboratory Manual, Cold Spring Harbor Laboratory, NY (1989)). Useful selectable markers for gene amplification in drug-resistant mammalian cells include DHFR-MTX (methotrexate) resistance (Alt et.al., (1978) J. Biol. Chem. 253:1357; Wigler et. al, (1980) Proc. Natl. Acad. Sci. USA 77:3567), and other markers known in the art (as reviewed, for example, in Kaufman et.al., (1988) Meth. Enzymology 185:537). The most widespread method for amplifying a target gene in cell culture is the use of methotrexate (Mtx) treatment to amplify dihydrofolate reductase (Dhfr), however, surprisingly, embodiments of the present disclosure provide a significantly better amplification method using increasing concentrations of puromycin.
[0077] A preferred selection and amplification marker is the gene that encodes puromycin resistance. In certain embodiments of the present disclosure high levels of puromycin areused to apply selective pressure on the cells and the exogenous gene spreads along with the puromycin resistance gene throughout the GADS. In certain embodiments of the disclosure more than 5 copies of the exogenous gene are spread across different sites and chromosomes. In certain preferred embodiments, more than 10 copies of the exogenous gene are spread across different sites and chromosomes and in certain other preferred embodiments more than 25 copies of the exogenous gene are spread across different sites and chromosomes. In the most preferred embodiments, more than 50 copies of the exogenous gene are spread across different sites and chromosomes, leading to very high expression of the exogenous gene.
[0078] Previous methods of applying puromycin selection have been time consuming (see for example, Prieto et al, Prieto et al. BMC Proceedings 2011, 5(Suppl 8):P7 at www. biomedcentral. com / 1753-6561 / 5 / S 8 / P7] and methods of the present disclosure found surprising and unexpected results using a novel protocol of multi-step increases in puromycin concentration. Preferred methods of the disclosure involve step wise increases in puromycin from 10 ug / ml up to 250 ug / ml with each increase taking place after less than 7 days. In more preferred embodiments puromycin concentration was increased after about 3 days.
[0079] In certain embodiments of the present disclosure cell lines are created that express one recombinant protein or peptide of interest. In more preferred embodiments of the disclosure the cell lines of the disclosure produce two or more different proteins or peptides. In still further preferred embodiments, four different proteins or peptides may be produced in a single cell line. An example of the expression of two different peptides is provided herein as shown in Figure 6 which depicts the production of both the heavy and light chains of the trastuzumab antibody following with transformation of HEK293 cells by the vector depicted in Figure 2.
[0080] Thus, preferred embodiments of the expression vectors may encode one recombinant protein or peptide of interest. In more preferred embodiments of the expression vectors two or more different proteins or peptides are encoded. In still further preferred embodiments, four different proteins or peptides may be encoded in a single vector. In other embodiments, multiple vectors may be used to express yet more exogenous proteins or peptides.
[0081] In certain embodiments of the disclosure the recombinant proteins produced may be antibodies. In certain preferred embodiments, both light and heavy chains of antibodies may be produced by the same cell line.
[0082] Embodiments of the disclosure can be used to produce almost any protein, even complex biologies containing multiple polypeptide chains. In certain embodiments of the disclosure the cell lines stably produce large amounts of a drug selected from adalimumab, atezolizumab, nivolumab, pembrolizumab, etanercept, trastuzumab, bevacizumab, rituximab, aflibercept, infliximab, ustekinumab, ranibizumab, proteins of tumor biology, proteins of food industry, proteins of animal health, proteins of ageing, proteins of genetic disorders, vaccines (VLPs, single proteins, virus inhibitor proteins).
[0083] Applicants have isolated and identified novel sequence elements that can improve expression of recombinant proteins from at least two to ten-fold in stable cell lines when inserted in an expression vector. We refer to these novel sequence elements as GADS for genomic amplification drop sites.
[0084] GADS according to the instant disclosure have been found in a variety of mammalian genomes and cell lines. Some are found in many species and some are found in only human cells.
[0085] The novel insertion and amplification sites according to the disclosure comprise sequences that are found at sites that are adjacent to each other in wild type human cells, and some such sequences are found in the intergenic spacer regions of rDNA gene sequences. The useful GADS sequences were discovered in a research program to uncover the best possible insertion and amplification sites for recombinant protein production and mammalian artificial chromosome production.
[0086] Methods according to the instant disclosure involve the novel use of sequences hypothesized to be used by mammalian cells to further chromosomal evolution. Sequences according to the instant disclosure are responsible for the amplification processes that occur in these vast rDNA regions in the continuously ongoing chromosomal evolution. These regions can amplify themselves and new chromosomes can be formed. These new chromosomes may be inherited through generations without any apparent side effects in humans. This process may seem slow, but in the evolutionary sense it is very fast. Extra chromosomes are formed naturally in the human population with a 0.043% frequency. These chromosomes are called sSMC chromosomes (small supernumerary marker chromosomes) and these are inherited in human families through many generations often without any apparent side effects (Fu S, Fu H et al. (1992). Yi Chuan Xue Bao 19(4):294- 7.; Gravholt CH, Friedrich U. (1995). Am J Med Genet 56(1): 106-11.; Csonka E. (2008). Hungarian Medical Journal 2(3):365-380.).
[0087] An extensive research program was conducted to uncover the exact DNA sequences that are responsible for the amplification process described above and this research surprising lead to the development of the instant recombinant protein expression disclosure. During the examination of the vast amount of possible DNA it was found that no rDNA gene sequences are involved in the process and so the search began for sequences in the non coding intergenic spacers between the genes and upstream and downstream from the rDNA genes in the non-coding chromosomal regions. This massive search required the examination of hundreds of kilobases of DNA sequences.
[0088] The extensive research described above uncovered two human GADS sequences: the CDC27 pseudogene sequence and a 2993 base pair (bp) sequence. Further examination and utilization of the 2993 bp sequence determined that a smaller portion of the sequence (904 bp) works well in certain embodiments of the methods for targeting and amplification in the HEK293 cell line according to the present disclosure. Embodiments of the disclosure utilizing this smaller sequence are advantageous when the exogenous gene to be expressed is large, or when several genes are to be expressed at the same time in the same cell line. In certain embodiments up to 4 genes can be expressed from one plasmid of the disclosure. In other embodiments multiple expression vectors may be used to transform the cells, increasing the number of polypeptides that can be produced simultaneously.
[0089] The present disclosure relates to the identification of recombinant protein integration sites in a variety of host genomes, and the construction of homologous recombination vectors for achieving high, stable recombinant gene expression in mammalian cells.
[0090] A number of GADS sequences have been identified and characterized. These sequences are disclosed as SEQ ID NO:l, SEQ ID NO:2 and SEQ ID NO:3 herein and also referred to as GADS1, GADS2 and GADS3 respectively.GADS1: Human CDC27 pseudogene
[0091] Certain embodiments of the instant disclosure include sequences known as the Homo sapiens cell division cycle 27 pseudogene (CDC27 pseudogene). This pseudogene has no previously known function. There are many versions of CDC27 pseudogenes which can be used in embodiments of the instant disclosure. CDC27 pseudogenes are found on at least 10 chromosomes in wild type human cells including chromosomes 2, 7, 13, 14, 15, 16, 20, 21, 22, and Y. CDC27 pseudogene 11 is found on chromosomes 7, 15, 16, 20, 21, 22, and Y in wild type human cells. However, in HEK293 cells CDC27 pseudogene 11 (hereinafterGADS1) sequences are found on 11 to 13 chromosomes by FISH analysis. The more widespread distribution in HEK293 is hypothesized to be the result of the altered genome found in this transformed cell line. The increased distribution of the GADS1 sequence in HEK293 cells is useful in certain embodiments of the disclosure because there are more target sites in these cultured cells. The same is true of CHO and mouse cell lines.
[0092] A preferred sequence according to certain embodiments of the instant disclosure includes the 1978 base pair CDC27 pseudogene 11 sequence (GADS1). The CDC27 pseudogene sequences according to the disclosure are useful in embodiments of the instant disclosure where protein production in cell lines originating from a wide variety of different species are desired because the CDC27 pseudogene sequence is highly conserved in a variety of mammalian species. Nucleotide sequence identity to Cricetulus griseus (Chinese Hamster) CDC27 pseudogene DNA sequence is 85.87% compared to the whole human sequence with only 1% gaps were detected in a Blast search. Nucleotide sequence identity to Mus musculus (mouse) CDC27 pseudogene DNA sequence is 84.65% compared to the whole human sequence with only 1% gaps were detected in a Blast search. The CDC27 pseudogene is even more highly conserved across different primate species.
[0093] GADS1 sequences according to the instant disclosure have been shown by FISH analysis to be present on at least 11 chromosomes in HEK293 cells, on the acrocentric arms in several hundred to thousands of copies. Thus, embodiments of the disclosure including constructs and methods that use these sequences are preferred for very high amplification and production of exogenous proteins.
[0094] GADS1 sequences useful according to the present disclosure have been identified on chromosomes 2, 22 and Y. (See RefSeq, accessed May 2014: htps: / / www.alliancegenome.Org / gene / HGNC:1728T In addition, GADS1 sequences are found in Nucleolus organiser regions (NORs) which are chromosomal regions crucial for the formation of the nucleolus. In humans, the NORs are located on the short arms of the acrocentric chromosomes 13, 14, 15, 21 and 22, the genes RNRl, RNR2, RNR3, RNR4, and RNR5 respectively. Finally, wild type human cells carry GADS1 sequences on Chromosomes 7, 15, 16, 20, 21, 22, Y. These regions carry tens of copies of this sequence, each. So, the wild type genome carries hundreds of copies of this sequence.
[0095] GADS1 sequences were successfully used in methods according to the present disclosure utilizing CHO-K1 cells where insertion and amplification of exogenous DNA was demonstrated. Targeting and amplification was demonstrated in three different regions onlarge metacentric hamster chromosomes: at the very end of metacentric chromosomes; in the middle of one chromosomal arm of metacentric chromosomes; and in the middle (close to the centromere) of metacentric chromosomes.
[0096] In other embodiments mouse GAD SI homologues may be used for exogenous gene expression in mouse cell lines.
[0097] SEQ ID NO:l according to the present disclosure comprises the 5’-3’ DNA sequence disclosed in the Sequence Listing provided herewith, originating from Homo sapiens and having a length of 1978 bp ("SEQ ID NO:l” and “GADS1” are used interchangeably herein).
[0098] In preferred embodiments of the instant disclosure GADS1 (SEQ ID NO:l) sequences are found in euchromatin regions, suitable for expression and are not silenced. This makes these sites ideal for recombinant protein production. There are many sites in the target genome where the gene of interest can be integrated with the help of CDC27 pseudogenes such as the preferred sequence of GADS1.
[0099] As discussed above and while not wanting to be bound by theory, the natural CDC27 pseudogenes are hypothesized to have an evolutionary function, because they are so extremely conserved. This CDC27 pseudogene sequence is a natural amplificator. In certain embodiments the application of antibiotic selection pressure to this sequence causes the GADS to be highly amplified together with the gene of interest.
[0100] The chromatin around the CDC27 pseudogenes is always open for gene expression making the sites extremely useful in certain embodiments of the disclosure for recombinant gene expression. The inventors hypothesize that these and the other identified GADS sequences are involved in chromosomal evolution and have evolved to essentially be “untouchable” zones - remaining intact throughout evolutionary history. In addition, the natural amplification of these GADS, is accomplished without silencing, setting the systems of the present disclosure apart from all of the other amplification and expression systems previously developed.
[0101] Further, a significant advantage of embodiments of the present disclosure over the previous known DHFR and GS based amplification processes that are used in the industry presently, is that the amplified chromosome arm created in embodiments of the present disclosure are stable and that is one reason why recombinant proteins are expressing stably at high levels in embodiments of the present disclosure.
[0102] In addition to GADS1, the present disclosure encompasses fragments of SEQ ID NO:l that also exhibit GADS activity.
[0103] Expression vectors comprising the isolated 1978 bp sequence (SEQ ID NO:l) and shorter fragments thereof are useful to transform CHO cells and result in high levels of stable protein expression. The present novel GADS1 is useful to improve expression of a recombinant protein driven by a promoter / enhancer region to which it is linked.
[0104] Expression vectors comprising the isolated 1978 bp sequence (SEQ ID NO:l) and shorter fragments thereof are useful to transform HEK293 cells and result in high levels of stable protein expression. The present novel GADS1 is useful to improve expression of a recombinant protein driven by a promoter / enhancer region to which it is linked.
[0105] Moreover, additional fragments of SEQ ID NO:l exhibiting GADS activity can be identified, as well as similar GADS motifs from other types of cells or from other integration sites in transformed cells. In addition, it is known in the art that subsequent processing of fragments of DNA prepared by restriction enzyme digestion can result in the removal of additional nucleotides from the ends of the fragments.
[0106] A fragment (21 lbp - SEQ ID NO:4) of GADS1 (SEQ ID NO:l) is also a part of the 904 bp sequence of GADS3 (SEQ ID NO:3) and also the part of the 2993 bp sequence GADS2 (SEQ ID NO:2). One skilled in the art can thus devise various fragments of the sequences disclosed herein for use in additional embodiments of the present disclosure.
[0107] Other combinations of fragments of SEQ ID NO:l can also be developed, for example, sequences that include multiple copies of all or a part of SEQ ID NO:l. Such combinations can be contiguously linked or arranged to provide optimal spacing of the fragments. Additionally, within the scope of the present disclosure are expression vectors comprising the sequence of SEQ ID NO:l arranged with insertion sequences therein (e.g., insertion of a gene encoding a desired protein at a certain selected site in SEQ ID NO:l).GADS2: 2993 bp long sequence
[0108] In additional embodiments of the instant disclosure a 2993 bp sequence with no previously known function, hereinafter called “GADS2” is provided. GADS2 sequences are found in wild type human cells on chromosomes 7, 15, 16, 20, 21, 22 and Y. These regions each carry tens of copies of this sequence. Thus wild-type human genomes carry hundreds ofcopies from this sequence. A small part of this sequence overlaps with the sequence ofGADS1.
[0109] There is no known equivalent sequence in non-human species to GADS2 but this result might be merely because the non-coding portions of these genomes have been less well characterized.
[0110] In certain embodiments the GADS2 sequence was used to make targeting and amplification of plasmids in mouse and hamster cell lines. These sequences were used to create independent mammalian artificial chromosomes as well as amplified chromosome arms. Thus, GADS2 sequences are useful for making artificial chromosome arms and independent artificial chromosomes. An embodiment of the instant application showing an example of such artificial chromosomes and chromosome arms is shown in Figures 11 and 13
[0111] GADS2: SEQ ID NO:2 according to the present disclosure comprises the 5’-3’ DNA sequence disclosed in the Sequence Listing provided herewith, originating from Homo sapiens and having a length of 2993 bp ("SEQ ID NO:2” and “GADS2” are used interchangeably herein).GADS3: 904 bp long sequence (a smaller fragment of the 2993 bp sequence)
[0112] A smaller part of the GADS2 sequence is especially useful in developing human cell lines expressing recombinant proteins. The smaller sequence comprises 904 base pairs of the GADS2 sequence and is referred to as GADS3 or SEQ ID NO:3. Wild type human cells carry these GADS3 sequences on chromosomes 7, 15, 16, 20, 21, 22, and Y. These regions carry tens of copies of this GADS3 sequence, each. The wild-type human genome carries hundreds of copies of GADS3 sequences.
[0113] Cultured human cells (including but not limited to HEK293 cells) carry hundreds to thousands of copies from this GADS3 sequence based on FISH experiments. As discussed previously, cultured human cells have altered genomes, chromosome numbers and chromosome rearrangements and amplifications are very frequent events. Thus, GADS3 sequences are ideal for certain preferred embodiments of the methods and constructs of the instant disclosure.
[0114] GADS3 has a number of desirable characteristics for use in preferred methods and constructs of the instant disclosure. 211 bp of the GADS3 sequence overlaps with GADS1 at the 5’ end.
[0115] No GADS3 nucleotide sequence identity was found to Cricetulus griseus Blast sequences.
[0116] No GADS3 nucleotide sequence identity was found to Mus musculus Blast sequences. Thus, in certain embodiments of the disclosure, methods and constructs incorporating GADS3 sequences are useful for production and amplification in human cell lines.
[0117] One significant advantage of embodiments of the instant disclosure comprising the GADS3 sequence is that GADS3 is small which is advantageous when you have large exogenous genes to express, or several genes to be expressed at the same time. Certain embodiments comprising expression vectors including GADS3 can include up to 4 exogenous genes to be expressed on a single plasmid. As disclosed in the examples, GADS3 vectors are useful for simultaneous expression of both the heavy and light chains of antibodies.
[0118] GADS3: SEQ ID NO:3 according to the present disclosure comprises the 5’-3’ DNA sequence disclosed in the Sequence Listing provided herewith, originating from Homo sapiens and having a length of 904bp ("SEQ ID NO:3” and “GADS3” are used interchangeably herein).GADS4: SEQ ID NO:4 - A 211bp fragment of GADS2 and GAD S3
[0119] There is a 211 bp overlap between CDC27 pseudogene 11 (GADSl, SEQ ID NO:l), the 904 bp sequence (GADS3, SEQ ID NO:3) and the 2993 bp sequence (GADS2, SEQ ID NO:2). This sequence, hereinafter "GADS4” or “SEQ ID NO:4”, can be used in certain embodiments of the disclosure.
[0120] GADS4: SEQ ID NO:4 according to the present disclosure comprises the 5’-3’ DNA sequence disclosed in the Sequence Listing provided herewith, originating from Homo sapiens and having a length of 211bp ("SEQ ID NO:4” and “GADS4” are used interchangeably herein).GADS5: SEQ ID NO:5 - A 293 bp fragment of GADS2
[0121] A fragment of the 2993 bp human sequence (GADS2, SEQ ID NO:2) has a 238 bp identity with 2% gaps with mouse genomic sequence. This sequence does not overlap with the 904 bp sequence (GADS3, SEQ ID NO:3) or with the CDC27 pseudogene 11 sequence(GADS1, SEQ ID NO: 1)
[0122] GADS5: SEQ ID NO:5 according to the present disclosure comprises the 5’-3’ DNA sequence disclosed in the Sequence Listing provided herewith, originating from Homo sapiens and having a length of 293 bp ("SEQ ID NO:5” and “GADS5” are used interchangeably herein).
[0123] One skilled in the art will recognize that changes can be made in the nucleotide sequences set forth in SEQ ID NO:l, SEQ ID NO:2 and SEQ ID NO:3 by site directed or random mutagenesis techniques that are known in the art. The resulting GADS variants can then be tested for GADS homologous recombination activity as described herein. DNAs that are at least about 80%, more preferably about 85%, and more preferably about 90% identical in nucleotide sequence to SEQ ID NO: 1, 2, or 3 or fragments thereof, having GADS homologous recombination activity are isolatable by experimentation and hypothesized to have GADS recombination. Accordingly, homologues of the disclosed GADS sequences and variants thereof are also encompassed by the present disclosure.
[0124] The following examples are illustrative of embodiments of the present disclosure and do not limit the scope of the disclosure in any way. All references cited herein, whether supra or infra, are hereby incorporated by reference in their entirety.EXAMPLESEXAMPLE 1: GADS 3 CHO-K1 cell line production of trastuzumab
[0125] The optimal DNA sequence for expression of trastuzumab was determined and the DNA was synthesized in a pUC plasmid and transferred into an appropriate insertion vector according to the disclosure for insertion in CHO-K1 cells. A typical plasmid is shown inFigure 1.
[0126] This plasmid was used to establish stable CHO-K1 cell lines with amplified “trastuzumab” as follows: a. CHO-K1 cells were seeded into one well of a 24- well plate (approximately 100000 cells). The next day cells were transfected with 1 pg trastuzumab containing plasmid. For this cell line we use the Turbofect reagent (see https : / / www. thermofisher com / order / catalo g / product / R0532# / R0532. accessedMay 31, 2021). b. The insertion plasmid was diluted with 100 mΐ serum-free DMEM medium. c. The Turbofect reagent was mixed thoroughly by vortexing. d. 2 mΐ Turbofect reagent was added to the insertion plasmid solution. e. This reaction was mixed by pipetting and incubated at room temperature for 20 minutes. f. The transfection mix was added evenly and drop by drop to the 24 cells in the wells with each well containing serum containing 1 ml DMEM medium. g. After 24 hours, the cells were collected using TrypLE Select reagent and then distributed into the wells of 2x96-well plates. About 40 ml of DMEM was used in order to achieve volume of about 200 mΐ x 96 x 2 ml altogether. h. After 24 hours antibiotic selection is begun by the addition of 10 pg / ml puromycin. i. After 3 days the selection medium is exchanged for fresh medium. j. After an additional 3 days the selection medium is again exchanged. k. After an additional 3 days the growing cell clones are collected from the wells of the 96-well plate by TrypLE Select reagent, and individually transferred into one well of a 24-well plate. l. 24 hours later antibiotic selection with 50 ug / ml Puromycin is started. m. 3 days later the selection medium is exchanged for medium containing 100 pg / ml Puromycin containing medium. n. The growing cell lines were examined by lysing the cells and purifying the total protein. Western blot analysis (see Figure 4) was used to determine the production of trastuzumab. o. The highest producing cell lines are checked with FISH experiments for the presence of the insertion plasmid. See, for example, Figure 3. p. The best clones are further selected with increased antibiotic selection. The cells that are growing at 100 pg / ml puromycin are split into 3 wells of a 24-well plateand after 24 hours the cells are subjected to 150 pg / ml, 175 pg / ml and 200 pg / ml puromycin. q. Surviving clones are examined with western blotting for protein production and with FISH for the presence of amplification. r. The best producing clones are grown in serum-free medium. s. Trastuzumab is purified from the medium. t. The concentration of trastuzumab is determined by spectrophotometer, HPLC or similar procedures. u. Highest yielding cell lines will be selected and trastuzumab can be purified.EXAMPLE 2: GAD S3 HEK293 cell line production of trastuzumab
[0127] Plasmids containing the GADS3 sequence (SEQ ID NO:3) were used to transfect HEK293 cell line. The insertion vector is shown in Figure 2. The production of trastuzumab was confirmed by Western blotting as shown in Figure 5 using the 42kDa actin protein expression for quantitation.
[0128] Trastuzumab is purified using protein G columns. The Protein G HP SpinTrap column was used for small scale purification. Figure 6 shows a blot of the purified trastuzumab both heavy and light chains (HETPlO / 13 Herceptin; 0.9-3 g / L (Bradford); 3.5x106 cell / column: 111 pg / cell). (for SpinTrap see https: / / www.sigmaaldrich.com / catalog / product / si gma / ge28903134?lang=hu®ion=HU&gclid=CiwKCAiwtdeFBhBAEiwAKOIv5914n5Xlr4X4sllmGD Z46KOJtAlk-rlomuiFYESnStEUWOXp-aaThoCi-4QAvD BwE accessed May 31, 2021).
[0129] Based on the Bradford protein concentration measurement, we found that trastuzumab production was at least 111 pg / cell.
[0130] The highest producing cell line (10 / 13) was subjected to subcloning (see Figure 7) and the highest producing clone (#11) was subjected to FISH analysis (see Figure 8).EXAMPLE 3: GADS2 cell line production in mouse cell lines
[0131] Mouse cell lines: Mouse embryonic fibroblast (MEF) cells from 3.5 days old individual mouse embryos were isolated and used to establish stable cell lines. The cells wereimmortalized with well documented methods: basically, cells were passaged every 3 days until immortalized. These cells show several markers of mesenchymal stem cells by FACS experiments. This process and these cell lines are useful for adult mesenchymal stem cells from adult tissues (preferably from adipose tissue after liposuction), modify those cells with mammalian artificial chromosomes created using the instant disclosure and using them directly or after differentiation into certain cell types (fat, cartilage, bone etc.)) for gene therapy and / or cell therapy experiments. The first step, shown here, demonstrates a vector according to the disclosure that carried the 2993 bp sequence (GADS2: SEQ ID NO:2) into which was inserted a useful gene (Influenza A virus Hemagglutinin MYMC X-181 California strain) as shown in Figure 9. This embodiment of the disclosure was transfected into the immortalized MEFs, producing 34 stable cell lines with small-scale amplification that resulted in 30-50 copies of the exogenous sequence, as shown in Figure 10. Surprisingly, targeting exclusively happened into a large acrocentric chromosome in the upper part of the long chromosomal arm in all 34 clones.EXAMPLE 4: GADS2 cell line production in CHO cell lines
[0132] A plasmid according to the disclosure (shown in Figure 9) was used to transfect CHO- DG44 cells producing 79 clones. Autonomous mammalian artificial chromosomes were formed in 16 cell lines out of the 79. Figure 11 shows an example of an autonomous mammalian artificial chromosome and an amplified chromosome arm together. One hamster chromosome has an amplified chromosome arm with hundreds of transgene copies. The autonomous mammalian artificial chromosome consists almost entirely of transgene sequences, representing hundreds of transgene copies.EXAMPLE 5: GADS1 cell line production in mouse cell lines
[0133] LMTK- mouse cells were transfected with a plasmid construct according to the disclosure carrying GADS1 sequence (CDC27 pseudogene: SEQ ID NO:l) and the Influenza A virus Hemagglutinin MYMC X-181 California strain sequence was inserted as a useful exogenous gene (Figure 12). We produced 32 cell lines. Eight cell lines carried autonomous mammalian artificial chromosomes (Figure 13).EXAMPLE 6: GADS1 cell line production in mouse cell lines
[0134] CHO-K1 cell lines were made with a plasmid according to the disclosure carrying the GADS1 sequence (CDC27 pseudogene: SEQ ID NO:l) and expressing RBBP7 protein (pIKRBBP7, see Figure 14). Two types of cell lines were constructed in this embodiment. One type of cell line expresses and secretes the RBBP7 protein into the culture medium (CIRBBP cell lines). In this embodiment, the RBBP7 protein is expressed with a hamster IgK secretion signal on the N-terminal. After the secretion signal, there are two tags for labeling and purification (an AVI tag and a 6xHis tag). 53 cell lines were produced with this construct and 24 of them stably produce the RBBP7 protein (see Figure 15). This was shown by Western blotting experiments.EXAMPLE 7: GADS1 cell line production in mouse cell lines
[0135] Additional RBBP7 protein producing cell lines were produced without the hamster IgK secretion signal to cause the protein to remain inside the cells rather than be excreted. Proteins produced in this embodiment are purified from cell lysates.
[0136] While the present disclosure describes various embodiments for illustrative purposes, such description is not intended to be limited to such embodiments. On the contrary, the applicant's teachings described and illustrated herein encompass various alternatives, modifications, and equivalents, without departing from the embodiments, the general scope of which is defined in the appended claims. Except to the extent necessary or inherent in the methods or processes themselves, no particular order to steps or stages of methods or processes described in this disclosure is intended or implied. In many cases the order of method or process steps may be varied without changing the purpose, effect, or import of the methods described.
[0137] Information as herein shown and described in detail is fully capable of attaining the above- described object of the present disclosure, the presently preferred embodiment of the present disclosure, and is, thus, representative of the subject matter which is broadly contemplated by the present disclosure. The scope of the present disclosure fully encompasses other embodiments which may become apparent to those skilled in the art, and is to be limited, accordingly, by nothing other than the appended claims, wherein any reference to an element being made in the singular is not intended to mean "one and onlyone" unless explicitly so stated, but rather "one or more." All structural and functional equivalents to the elements of the above-described preferred embodiment and additional embodiments as regarded by those of ordinary skill in the art are intended to be encompassed by the present claims. Moreover, no requirement exists for a system or method to address each and every problem sought to be resolved or ameliorated by the present disclosure, for such to be encompassed by the present claims. Furthermore, no element, component, or method step in the present disclosure is intended to be dedicated to the public regardless of whether the element, component, or method step is explicitly recited in the claims. However, that various changes and modifications in form, material, work-piece, and fabrication material detail may be made, without departing from the spirit and scope of the present disclosure, as set forth in the appended claims, as may be apparent to those of ordinary skill in the art, are also encompassed by the disclosure.
Claims
CLAIMSWhat is claimed is:
1. A nucleic acid construct comprising a DNA fragment capable of site-specific insertion into a region of open chromatin in mammalian cells on at least one chromosome at more than one site.
2. The nucleic acid construct according to Claim 1, which is amplified more than 5 times following selective pressure.
3. The nucleic acid construct of either one of Claim 1 or Claim 2, wherein said DNA fragment comprises the sequence of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO: 5, or a functional fragment thereof.
4. The nucleic acid construct according to Claim 2, wherein the DNA fragment is insertable at more than 5 sites on two or more chromosomes.
5. An expression vector comprising the nucleic acid construct of any one of Claims 1 to 4.
6. The expression vector of Claim 5, further comprising a selectable marker.
7. The expression vector of Claim 6, wherein the selectable marker provides resistance to puromycin.
8. The expression vector of any one of Claims 5 to 7, further comprising a nucleic acid sequence encoding all or a portion of a protein of interest.
9. The expression vector of Claim 8, wherein said protein of interest is a drug selected from the group consisting of adalimumab, atezolizumab, nivolumab, pembrolizumab, etanercept, trastuzumab, bevacizumab, rituximab, aflibercept, infliximab, ustekinumab, ranibizumab, proteins of tumor biology, proteins of food industry, proteins of animal health, proteins of ageing, proteins of genetic disorders, vaccines, viral-like particles (VLPs), single proteins, and virus inhibitor proteins.
10. A mammalian host cell transformed with the expression vector of claim 5.
11. The mammalian host cell of Claim 10, wherein the expression vector further comprises a selectable marker.
12. The mammalian host cell transformed of either one of Claim 10 or Claim 11, wherein the mammalian host cell is a 9HEK293 cell, a HT1080 cell, a Per.C6 cell, a pluripotent cell, an induced pluripotent cell, a totipotent cell, an adult stem cell, or a primary cell.
13. A Chinese hamster ovary (CHO) cell transformed with the expression vector of Claim 5.
14. A mouse cell transformed with the expression vector of Claim 5.
15. A human cell transformed with the expression vector of Claim 5.
16. A method for obtaining a recombinant protein, comprising culturing a transformed host cell according to any one of Claims 10 to 15, under conditions promoting expression of said recombinant protein, and recovering the recombinant protein.
17. A plasmid constructed according to a plasmid map as shown in Figure 1.
18. A plasmid constructed according to a plasmid map as shown in Figure 2.
19. A plasmid constructed according to a plasmid map as shown in Figure 9.
20. A plasmid constructed according to a plasmid map as shown in Figure 12.
21. A plasmid constructed according to a plasmid map as shown in Figure 14.
22. A plasmid comprising a DNA fragment according to Claim 3.
23. The plasmid according to Claim 22, further comprising a transgene and a selectable marker that provides resistance to puromycin.
24. A DNA vector capable of integration into a mammalian genome, said DNA vector comprising: a. a DNA fragment capable of site-specific insertion into a region of open chromatin in mammalian cells, wherein said DNA fragment is the sequence of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, or a functional fragment thereof; and b. at least one transgene encoding at least one protein of interest to be expressed by a mammalian cell after the DNA vector integrates into the mammalian genome of the mammalian cell.
25. The DNA vector of Claim 24, wherein said at least one protein of interest is a drug selected from the group consisting of adalimumab, atezolizumab, nivolumab, pembrolizumab, etanercept, trastuzumab, bevacizumab, rituximab, afbbercept, infliximab, ustekinumab, ranibizumab, proteins of tumor biology, proteins of food industry, proteins of animal health, proteins of ageing, proteins of genetic disorders, vaccines, viral-like particles (VLPs), single proteins, and virus inhibitor proteins.
Citation Information
Patent Citations
Generation and use of pluripotent stem cells
WO2012021632A2
Method for integrating long foreign gene into safe region of human pluripotent stem cell and allowing same to normally function therein
WO2020171222A1