Methods for assessing protein production in CHO cells
DNA methylation profiling of CHO cells addresses the instability in protein production by identifying genetically and phenotypically identical clones, ensuring consistent high yield and quality in industrial-scale biologics production.
Patent Information
- Application Number
- JP2025511596
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-01
- Filing Date
- 2023-08-23
- Publication Date
- 2025-10-24
AI Technical Summary
Current methods for assessing CHO cell lines for protein production are time-consuming and inaccurate, leading to instability and inefficiency in industrial-scale biologics production due to genetic and epigenetic variations, resulting in reduced protein productivity over time.
A method utilizing DNA methylation profiling of CHO cells to identify genetically and phenotypically identical clones, ensuring stability and consistency in protein production by comparing test methylation profiles with reference profiles using DNA methylation bead-based arrays.
Enhances the speed, quality, and efficiency of heterologous protein production by selecting clones with consistent methylation patterns, maintaining high protein yield and quality over long-term culture.
Smart Images

Figure 2025535223000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an epigenetics-based method for quantitatively and qualitatively assessing target protein production in CHO cells and cellular stability before, during or after the actual production of the protein. In particular, measures of variable methylation of promoters and / or CpG sites in CHO cells can provide insight into the quantitative and qualitative production of target proteins by CHO cells. [Background technology]
[0002] Chinese hamster ovary (CHO) cells have been known to be the driving force behind industrial production of recombinant therapeutic proteins since 1987 and are therefore widely used for biologics production. Approximately 70% of all recombinant biopharmaceutical proteins and all monoclonal antibodies approved since 2016 have been produced in CHO cells. Some advantages of utilizing CHO cells for biologics production include their resistance to genetic manipulation, ease of adaptation to manufacturing process scale, rapid growth rate, and ability to perform human-compatible post-translational modifications. However, biologics production systems in CHO cells face obstacles due to loss of protein productivity over time.
[0003] Although initial protein expression from cell lines is high, production declines during long-term culture. This results in reduced process yields, impacting timelines and increasing costs. Changes in the cell culture environment can lead to changes in the cellular behavior and protein productivity of producer cell lines. Some reasons for loss of productivity in CHO cells include the accumulation of numerous genomic mutations over long-term culture, transgene loss, and epigenetic regulation of the transgene insertion site. In particular, viral promoter integration sites are susceptible to transcriptional regulation via epigenetic modulation such as histone modifications and DNA methylation. The DNA methylation status of viral promoters is a key factor in protein production or expression stability in producer CHO cells. Increased DNA methylation in promoters leads to transgene silencing at the transcriptional level. Protein production variability in CHO cells has been linked to DNA methylation-mediated regulation of the cytomegalovirus major immediate-early and early enhancer (CMV) promoter and the simian vacuolating virus 40 (SV40) promoter, which are the most frequently used promoters for recombinant protein production in CHO cells.
[0004] Current methods for determining the suitability of CHO clones for target protein production are not only time consuming but also not very accurate in selecting clones or cells for optimal protein production.
[0005] Furthermore, genetically identical CHO clones can still result in heterogeneous phenotypes, potentially causing instability, inefficiency, and economic losses during industrial-scale heterologous protein production. Methods for comparing and selecting CHO clones using phenotypic analysis alone cannot guarantee consistency over time. Genotypic comparison of CHO clones cannot define how genes are differentially expressed to adapt to environmental conditions. As shown by Wippermann A, et al., Appl Microbiol Biotechnol. 2014 Jan;98(2):579-89, supplementation with butyrate, which is known to enhance cell-specific productivity in CHO cells, also resulted in changes in epigenetic silencing events. [Prior art documents] [Non-patent literature]
[0006] [Non-Patent Document 1] Wippermann A, et al., Appl Microbiol Biotechnol.2014 Jan;98(2):579-89 Summary of the Invention [Problem to be solved by the invention]
[0007] Therefore, there is a need in the art for efficient and affordable tools to comprehensively evaluate and regulate CHO metabolism and protein production. There is also a need in the art for methods to select and maintain identical CHO populations to improve the speed, quality, efficiency, and consistency of production. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a plot showing the results of a principle component analysis (PCA) of the 122 identified variably methylated regions (DMRs). [Figure 2]FIG. 2 is a plot showing the results of a principle component analysis (PCA) of the 289 identified variably methylated regions (DMRs). [Figure 3] 3 is a graph showing the viable cell counts of control CHO Humira431 cells and hyperosmolarity-treated CHO Humira431 cells. Sodium chloride was added to hyperosmolarity-treated CHO Humira431 cells on day 3. From day 3 onwards, in contrast to control CHO Humira431, a plateau in viable cell count was observed in hyperosmolarity-treated CHO Humira431 cells, which continued to increase until day 10, when viable cell count began to plateau. [Figure 4] 4 is a graph showing heterologous protein productivity of hyperosmolarity-treated CHO Humira431 cells and control CHO Humira431 cells on day 7 of fed-batch culture. Hyperosmolarity-treated CHO Humira431 cells were found to produce 86.5 pg / cell to 90.4 pg / cell of heterologous protein, in contrast to control CHO Humira431 cells, which produced 40.4 pg / cell to 41.3 pg / cell of heterologous protein. Therefore, addition of sodium chloride to hyperosmolarity-treated CHO Humira431 cells was found to increase heterologous protein productivity on day 7 of fed-batch culture. [Figure 5] Figure 5 is a graph showing the classification of six clones based on heterologous protein productivity on days 9, 11, and 14 of fed-batch culture. Based on productivity, clones 2C9, 3D11, and 2H2 are classified as low producers, clone 10A8 is classified as an intermediate producer, and clones 7H9 and 8F8 are classified as high producers. [Figure 6] FIG. 6 is a PCA plot showing the clustering of groups based on protein productivity. DETAILED DESCRIPTION OF THE INVENTION
[0009] The present invention solves the above-mentioned problems by providing a means not only to identify genetically identical CHO clones or cell lines but also to confirm that these clones and / or cell lines are phenotypically homogeneous, thus ensuring stability, efficiency, and reduced economic losses during heterologous protein production, particularly on an industrial scale. In particular, methods according to any aspect of the present invention use methylation patterns in CHO clones and / or cell lines and the preservation of these methylation patterns for the selection and maintenance of identical CHO populations to improve the speed, quality, efficiency, and consistency of heterologous protein production. Because genotypic comparison of CHO clones cannot define how genes are differentially expressed to adapt to environmental conditions, and phenotypic analysis alone cannot guarantee consistency over time, epigenetic methods, particularly DNA methylation, provide state-of-the-art technology for selecting not only genetically identical but also epigenetically, and therefore phenotypically, identical CHO clones for improved heterologous protein production. Methods according to any aspect of the present invention enable the use of DNA methylation as a tool for quantitatively and qualitatively improving protein production from CHO cells. Altering DNA methylation patterns on the viral promoter driving transgene expression transcriptionally increases protein expression in CHO cells.
[0010] According to one aspect of the present invention, there is provided a method for determining the suitability of at least one Chinese Hamster Ovary (CHO) test cell line for optimal heterologous protein production, the method comprising: (a) determining a test methylation profile from genomic material obtained from a CHO test cell line; (b) comparing the test methylation profile obtained from (a) with a reference methylation profile, the reference methylation profile comprising the methylation status of two or more CpG sites from at least one CHO reference cell line exhibiting at least one phenotype of interest for optimal heterologous protein production; Including, A method is provided, wherein a significant similarity of the test methylation profile of (a) compared to the reference methylation profile indicates that the CHO test cell line is suitable for optimal heterologous protein production, and wherein the test methylation profile and the reference methylation profile are from CpG sites from the CHO cell genome and are determined using a DNA methylation bead-based array.
[0011] Therefore, epigenetic technology provides a solution for quantitative and qualitative analysis of protein production. In particular, the reference methylation profile may include environment-specific or dynamic CpG sites, i.e., sites that are likely to play important roles in certain environmental conditions; CpG sites of viral promoters (CMV and SV40 promoters) and / or CpG sites in regulatory regions of candidate genes from pathways important in certain important biological processes of CHO cells (e.g., metabolism-related genes, protein production-related genes, cell growth / division-related genes, and methylation-related genes). More specifically, the test and reference methylation profiles are from CpG sites where the test methylation profile and the reference methylation profile are from CpG sites from the CHO cell genome.
[0012] The term "CHO cell genome" herein refers to the genomic DNA of a CHO cell, excluding the DNA of viruses used to introduce foreign DNA into the cell, particularly CMV and SV40. In particular, a CHO cell genome can refer to a cell having a genomic organization in the form naturally found in the wild. The term can also include genes added to the CHO genome by genetic modification (i.e., for improved production of proteins, etc.), but does not necessarily include genes and promoters from viruses used to introduce genes into the CHO genome. Thus, the term "CHO cell genome" can exclude viral genes and promoters and / or include endogenous or homologous genes of the CHO cell and / or genetically modified endogenous or homologous genes and / or intergenic genes of the CHO cell, DNA found between genes of the CHO cell.
[0013] The method according to any aspect of the present invention can be used to quantify the methylation level at any one of these CpG sites for CHO cells, particularly test CHO cells. This information can then be used to phenotypically assess, evaluate, and enhance CHO cells in various cell culture conditions. More specifically, machine learning models can be used to analyze the generated quantitative and qualitative methylation data. Even more specifically, the method according to any aspect of the present invention can be used in a predictive and accurate manner to design optimal cell culture conditions, particularly with regard to the selection of suitable CHO cell lines, compared to the current trial-and-error methods used. This therefore allows for direct online control of the manufacturing process, improving the robustness and therefore the overall quality of the molecules produced by the CHO cells.
[0014] CHO cell line refers to an immortalized Chinese hamster ovary cell line (CHO) derived from Cricetulus griseus. In particular, the CHO cell line may be selected from the group consisting of CHO-K1 (ATCC), CHO-DG44 (Thermo Fisher Scientific), CHO-DXB11 (ATCC), ExpiCHO-S™ cells (Thermo Fisher Scientific), FreeStyle™ CHO-S™ cells (Thermo Fisher Scientific), CHO 1-15 [subscript 500] (ATCC), and Agarabi CHO (ATCC).
[0015] The term "suitability" as used herein refers to a CHO cell line that is suitable for optimal heterologous protein production. In one example, a CHO cell line can be considered suitable for optimal heterologous protein production before a transgene is introduced into the cells. In this case, the CHO cell line can have phenotypic parameters or characteristics that allow the cell line to grow well, allow for easy incorporation of a transgene of interest, and, after incorporation of the transgene, allow for optimal heterologous protein production, where the protein is the product of the transgene of interest. These characteristics or phenotypic parameters include at least optimal glucose consumption, growth rate, lactate production, ammonia accumulation, etc. If a CHO cell line is confirmed to exhibit at least one of these phenotypic parameters, the CHO cell line can be considered suitable for optimal heterologous protein production when a transgene of interest is introduced into the cells.
[0016] In another example, a CHO cell line may be considered suitable for optimal heterologous protein production after a transgene has been introduced into the cells. In this case, the CHO cell line is genetically modified using methods known in the art to introduce the transgene into the cells, and the genetically modified cells are capable of optimal heterologous protein production when the protein is the translation product of the transgene. In this example, the CHO cell line may have at least one phenotype of interest that allows the genetically modified cell line to have good viability and optimal target protein production. These phenotypes of interest may include cell viability (viability), protein productivity (in terms of protein quantity and quality), phenotypic uniformity, cell depletion, etc. Thus, the methods according to any aspect of the present invention may be used with genetically modified CHO cell lines (i.e., with a transgene introduced into the cell line) or with non-genetically modified CHO cell lines. In both cases, the CHO cell line is intended for use in heterologous protein production.
[0017] As used herein, the term "transgene" refers to a gene that is removed from the genome of one organism and inserted into the genome of another organism by artificial techniques used for genetic modification. For example, a human gene is artificially introduced into the genome of a CHO cell for the production of at least one protein of interest, particularly a therapeutic protein.
[0018] As used herein, the term "therapeutic protein" refers to a genetically engineered version of a naturally occurring human protein. Examples of therapeutic proteins include antibody-based drugs, anticoagulants, blood factors, bone morphogenetic proteins, engineered protein scaffolds, enzymes, growth factors, hormones, interferons, interleukins, etc.
[0019] As used herein, the term "cell viability" refers to the ability of cells to survive and undergo cell proliferation. Cell viability is a measure of the proportion of living cells in a population. Cell proliferation refers to an increase in cell number due to cell division. Assays commonly used to test cell viability include the BrdU cell proliferation assay, the MTT cell proliferation assay, the trypan blue cell count, and the ATP cell viability assay.
[0020] As used herein, the term "cell exhaustion" refers to a state of cells that loses the ability to carry out metabolic activities, including heterologous protein production. Cell exhaustion can be determined by metabolite detection assays.
[0021] As used herein, the term "phenotypic homogeneity" refers to a state in which all cells in a population exhibit the same phenotype under specific conditions.
[0022] As used herein, the term "heterologous protein production" refers to the production of a protein that is not endogenous to a cell. It refers to the expression of a gene or part of a gene, particularly a transgene, in a host CHO cell that does not naturally express this gene. Assays commonly used to quantitate heterologous protein production include enzyme-linked immunosorbent assay (ELISA), chromatography, and bioprocess analyzers. As used herein, the term "host cell" refers to a cell line for the expression of a heterologous protein. For example, CHO cells are a primary host for the production of various therapeutic proteins.
[0023] The term "optimal heterologous protein production" herein refers to CHO cells capable of high-level protein production, particularly during industrial or large-scale production of recombinant proteins, where the protein is typically a functional protein not naturally present in wild-type CHO cells. In particular, for optimal heterologous protein production, CHO cell lines minimize metabolic burden and toxic effects on the cells. More specifically, "optimal heterologous protein production" refers to high-level protein production in which protein production is consistently maintained over the production period (i.e., long-term culture) so that the CHO cell line not only produces a high yield of the protein of interest, but also consistently maintains the quality of the produced protein. In particular, for a CHO cell according to any embodiment of the present invention to be capable of "optimal heterologous protein production," the cell must exhibit at least one or more of the following desired phenotypes: phenotypic uniformity, protein productivity, and protein quality. More specifically, for "optimal heterologous protein production," the CHO cell may exhibit phenotypic uniformity and protein productivity, or phenotypic uniformity and protein quality, or protein productivity and protein quality, or phenotypic uniformity, protein productivity, and protein quality.
[0024] As used herein, the term "protein productivity" refers to a measure of the amount of protein made per viable cell at a single titer point. It is calculated by dividing the titer (mg) by the viable cell density (VCD or cells / ml), with the final measurement expressed as the amount of protein per cell (mg / cell).
[0025] The term "protein quality" refers to post-translational modifications of proteins that determine their efficacy and function. Modifications generally include phosphorylation, glycosylation, ubiquitination, methylation, acetylation, protein folding, and the like. For example, protein glycosylation is an important quality attribute that regulates the efficacy, stability, and half-life of therapeutic proteins. Protein quality can be determined using immunoprecipitation-based techniques, biochemical assays, mass spectrometry (MS), and the like.
[0026] The terms "methylation profile," "methylation pattern," "methylation state," or "methylation status" are used herein to describe the state, condition, or status of methylation of a genomic sequence, and refer to features of a DNA segment at a particular genomic locus that are related to methylation. Such features include, but are not limited to, the presence or absence of methylation of any of the cytosine (C) residues within this DNA sequence, the location of methylated C residues, the proportion of methylated C in any particular stretch of residues, and allelic differences in methylation due to, for example, differences in allelic origin.
[0027] The term "methylation state" refers to the state of a particular methylation site (i.e., methylated vs. unmethylated), meaning that the residue or methylation site is methylated or unmethylated. A methylation profile can then be determined based on the methylation state of one or more methylation sites. Thus, the term "methylation profile" or even "methylation pattern" refers to the relative or absolute concentration of methylated or unmethylated C residues in any particular stretch of residues in the genomic material of a biological sample. For example, if a typically unmethylated cytosine (C) residue in a DNA sequence is methylated, this may be referred to as "hypermethylation," whereas if a typically methylated cytosine (C) residue in a DNA sequence is unmethylated, this may be referred to as "hypomethylation." Similarly, if a cytosine (C) residue in a DNA sequence (e.g., DNA of a sample nucleic acid from a test subject) is methylated compared to another sequence in a different region or different individual (e.g., compared to a normal nucleic acid or a standard nucleic acid of a reference sequence), the sequence is considered to be hypermethylated compared to other sequences. Alternatively, if a cytosine (C) residue in a DNA sequence is unmethylated compared to another sequence in a different region or in a different individual, the sequence is considered to be hypomethylated compared to other sequences. These sequences are said to be "variably methylated." Measuring the level of variable methylation can be performed by various methods known to those skilled in the art. One method, as a non-limiting example, is to measure the methylation level of each matched CpG site as determined by bisulfite sequencing.
[0028] The term "hypermethylation" refers to an average methylation state corresponding to an increased presence of 5-mCyt at one or more CpG dinucleotides within a DNA sequence of a test DNA sample compared to the amount of 5-mCyt found at the corresponding CpG dinucleotide in a normal control DNA sample.
[0029] The term "hypomethylation" refers to an average methylation state corresponding to a decreased presence of 5-mCyt at one or more CpG dinucleotides within a DNA sequence of a test DNA sample compared to the amount of 5-mCyt found at the corresponding CpG dinucleotide in a normal control DNA sample.
[0030] As used herein, "methylated nucleotide" or "methylated nucleotide base" refers to the presence of a methyl moiety on a nucleotide base, which is not normally present in recognized typical nucleotide bases. For example, cytosine does not contain a methyl moiety in its pyrimidine ring in its normal form, but 5-methylcytosine contains a methyl moiety at the 5th position of its pyrimidine ring. Thus, cytosine may not be considered a methylated nucleotide in its normal form, and 5-methylcytosine may be considered a methylated nucleotide. In another example, thymine may contain a methyl moiety at the 5th position of its pyrimidine ring, but for purposes herein, thymine may not be considered a methylated nucleotide when present in DNA. Typical nucleotide bases in DNA are thymine, adenine, cytosine, and guanine. Typical bases in RNA are uracil, adenine, cytosine, and guanine. Correspondingly, a "methylation site" is a position in a target gene nucleic acid region where methylation may occur. For example, a position containing CpG is a methylation site where the cytosine may or may not be methylated. In particular, the term "methylated nucleotide" refers to a nucleotide bearing a methyl group attached to a nucleotide position available for methylation. These methylated nucleotides are usually found in nature, and to date, methylated cytosine, which occurs primarily in association with the dinucleotide CpG, but also in association with CpNpG and CpNpN sequences, can be considered the most common. In principle, other naturally occurring nucleotides can also be methylated, but these are not considered in relation to any aspect of the present invention.
[0031] As used herein, the term "significantly similar," particularly in the context of comparing methylation profiles (such as comparing a test profile (from one or more test subjects) to a reference profile), refers to a similarity observed by statistical means (i.e., by using bioinformatics) and / or visual observation. Significant similarity is observed, for example, when a test profile overlaps with a reference profile defined by multiple training samples by multivariate statistical methods such as principal component analysis or multidimensional scaling. In particular, a test profile is significantly similar to a given reference profile if more than 50, 55, 60, 65, 70, 75, 80, 85, 90, 95% of the methylation pattern / profile overlaps with the methylation pattern / profile of the reference profile. Similarity of a test profile to more than one reference profile, such as two, three, or even all of the reference profiles, reduces the significance of the similarity.
[0032] As used herein, the term "genomic material" refers to a nucleic acid molecule or fragment of the genome of a CHO cell or cell line. In particular, such a nucleic acid molecule or fragment is DNA or RNA or a hybrid thereof, most preferably a molecule of the DNA genome of a CHO cell or cell line.
[0033] As used herein, a "DNA sample" refers to DNA extracted from a cell according to any aspect of the present invention using methods known in the art.
[0034] "Bisulfite treatment" of genomic DNA, which is used interchangeably with the term "bisulfite modification," refers to the treatment of genomic DNA with a deaminating agent, such as bisulfite, which can be used to treat all DNA, regardless of whether it is methylated or not. Specifically, the term "bisulfite" as used herein includes any suitable type of bisulfite, such as sodium bisulfite, or other chemical agents that can chemically convert cytosine (C) to uracil (U) without chemically modifying methylated cytosine and can therefore be used to differentially modify DNA sequences based on the methylation status of the DNA, e.g., U.S. Patent Application Publication No. 2010 / 0112595. As used herein, a reagent that "differentially modifies" methylated or unmethylated DNA includes any reagent that modifies methylated and / or unmethylated DNA in a process that results in distinguishable products from methylated and unmethylated DNA, thereby enabling identification of DNA methylation status. Such processes can include, but are not limited to, chemical reactions (such as bisulfite conversion of C to U) and enzymatic treatments (such as cleavage by methylation-dependent endonucleases). Thus, an enzyme that can preferentially cleave or digest methylated DNA is one that can cleave or digest DNA molecules with much higher efficiency when the DNA is methylated, while an enzyme that preferentially cleaves or digests unmethylated DNA shows significantly higher efficiency when the DNA is unmethylated.
[0035] Therefore, before step (a) according to any aspect of the invention is carried out, the genomic DNA contained in / obtained or extracted from the cells is first bisulfite treated.
[0036] Alternative methods available in the art may be used instead of bisulfite treatment. Those skilled in the art will understand which other methods to use. In one example, TET-assisted pyridine borane sequencing (TAPS) may be used to detect 5mC and 5hmC (Yibin Liu et al., Nature Biotechnology, 37:424-429 (2019)).
[0037] The term "test" as used herein in conjunction with the term cell refers to a cell that has been subjected to a method according to any embodiment of the present invention and that is the basis for the analytical application of the present invention. A "test cell" is therefore a CHO cell or a group of CHO cells that is being tested according to any embodiment of the present invention, or a profile obtained or generated in such a situation. Conversely, the term "reference" or "control" is intended to refer primarily to a given entity that is used for comparison with a test entity. In particular, a "test cell" refers to a cell that is being tested for suitability for optimal homologous protein production whose methylation status has to be determined, whereas a "control" or "reference" refers to a cell that is known to exhibit optimal homologous protein production or its methylation profile.
[0038] As used herein, a "CpG site" or "methylation site" is a nucleotide within a nucleic acid (DNA or RNA) that is susceptible to methylation, either by naturally occurring events in vivo or by events that initiate chemical methylation of the nucleotide in vitro. Some of these sites may be hypermethylated and some may be hypomethylated in cells. In some cases, a CpG site may not be considered completely hypermethylated or hypomethylated, but may be given a value that is a measure of the methylation of the CpG site. Thus, methylation may be quantified and is not necessarily an absolute case of hypermethylation or hypomethylation.
[0039] As used herein, a "methylated nucleic acid molecule" refers to a nucleic acid molecule that contains one or more nucleotides that are methylated.
[0040] As used herein, "CpG island" describes a segment of a DNA sequence that contains a functionally or structurally deviant CpG density. For example, Yamada et al. describe a set of standards for determining CpG islands. CpG islands must be at least 400 nucleotides long, have a GC content of more than 50%, and an OCF / ECF ratio of more than 0.6 (Yamada et al., 2004, Genome Research, 14, pp. 247-266). Others have defined CpG islands less strictly as sequences at least 200 nucleotides long with a GC content of more than 50% and an OCF / ECF ratio of more than 0.6 (Takai et al., 2002, Proc. Natl. Acad. Sci. USA, 99, pp. 3740-3745).
[0041] In particular, if there is variable methylation detected in the test cells, i.e., if the cells show absolute hypermethylation or hypomethylation or at least quantitative variable methylation at at least one CpG site compared to a reference (i.e., from a CHO cell line having at least one phenotype of interest), the test cells may also contain the phenotype of interest and be capable of optimal heterologous protein production. More specifically, if a CpG site shows the same methylation state in the test cells compared to the corresponding CpG site in the reference cell or reference methylation profile, the test cells may express the phenotype of interest and be capable of optimal heterologous protein production. Overall, this platform gives us the opportunity to detect a wide range of DNA methylation states in CHO cells and correlate them with industrially relevant parameters important for the development of at least one biopharmaceutical.
[0042] In particular, in step (a) of the method according to any embodiment of the present invention, the methylation status of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100 CpG sites is determined. A person skilled in the art will be able to determine the number of CpG sites that need to be used in step (a) of any embodiment of the present invention. Even more particularly, the methylation status of at least two CpG sites is determined in step (a) of the method according to any embodiment of the present invention.
[0043] As used herein, the term "epigenetic changes" refers to chemical (e.g., methylation) or protein (e.g., histone) changes that occur to a gene body or its promoter. Through epigenetic changes, environmental factors such as diet, stress, and prenatal nutrition can affect genes that are passed on from one generation to the next.
[0044] In particular, the reference methylation profile according to any embodiment of the present invention is a compilation of two or more CpG sites from at least one CHO reference cell line that exhibits at least one phenotype of interest for optimal heterologous protein production. In one example, different CpG sites are collected from a single reference CHO cell line that exhibits at least one phenotype of interest for optimal heterologous protein production. In another example, different CpG sites are collected from two or more cell lines, each of which exhibits at least one phenotype of interest for optimal heterologous protein production. Thus, the reference methylation profile according to any embodiment of the present invention may not be a naturally occurring methylation profile from a single CHO cell line, but rather an artificial profile obtained by combining related CpG sites from different reference CHO cell lines, each of which has at least one phenotype of interest for optimal heterologous protein production.
[0045] The phenotype of interest for an optimal heterologous protein may be selected from the group consisting of phenotypic uniformity, protein productivity, and protein quality.
[0046] According to a further aspect of the present invention there is provided a method for selecting at least one CHO cell comprising a phenotype of interest from a population of CHO cells from a parent clone, the method comprising: (a) determining a test methylation profile from genomic material obtained from a CHO cell; (b) comparing the test methylation profile of (a) with a reference methylation profile from a parent clone exhibiting the phenotype of interest; Including, (b) a significant similarity between the test methylation profile and the reference methylation profile indicates that the cell has the desired phenotype of the parent clone; A method is provided in which the phenotype of interest is selected from the group consisting of phenotypic uniformity, protein productivity, and protein quality, and the test methylation profile and the reference methylation profile are from CpG sites from the CHO cell genome and are determined using a DNA methylation bead-based array.
[0047] As used herein, the term "parent clone" refers to a cell line (CHO cell line) derived from a host cell into whose genome a transgene has been integrated. The term "subclone," as used herein with respect to a parent clone, refers to a clonal cell line derived from the parent clone that has the same genotype but a different phenotype due to epigenetic changes.
[0048] The method used according to this aspect of the present invention involves selecting at least one CHO cell line in at least one bioreactor that is genetically and phenotypically identical or significantly similar to a parental clone. In particular, phenotypic multiplicity typically arises during cell replication in a bioreactor of a parental clone. As used herein, the term "phenotypic multiplicity" refers to phenotypic variation present in a cell population, particularly CHO cells, under certain conditions without genotypic change. The method according to this aspect of the present invention allows for the selection of at least one clone from a phenotypically heterogeneous CHO cell population that has at least variation from the original / established parental clone that may also exhibit a phenotype of interest (e.g., production of at least one human-like protein). In particular, CHO cells identical or significantly similar to the parental clone can be identified by comparing the distribution of CpG site methylation (e.g., beta value distribution) in the clonal population of the bioreactor. CHO cells identical or significantly similar to the parental CHO cell line may have the same methylation profile. A partially methylated clonal population may also exhibit cell-to-cell variation.
[0049] Similarly, the method used according to this aspect of the present invention is to select at least one CHO cell or clonal population having a selective and specific methylation profile for protein productivity. In this example, the selected CHO cells have the same methylation profile as the parent clone from which the parent clone exhibits protein productivity. Thus, the reference methylation profile in this context refers to the methylation profile of the parent clone with protein productivity.
[0050] In another example, the method used according to this aspect of the present invention is to select at least one CHO cell or clonal population having a selective and specific methylation profile for protein quality. Protein quality can be measured based on ideal glycosylation / sugar backbone, etc. In this example, the selected CHO cells have the same methylation profile as the parent clone, which parent clone exhibits protein quality. Thus, the reference methylation profile in this context refers to the methylation profile of the parent clone having protein quality.
[0051] According to a further aspect of the present invention there is provided a method for identifying at least one CHO test cell line capable of producing at least one biosimilar to a heterologous protein produced by a CHO reference cell line, the method comprising: (a) determining a test methylation profile from genomic material obtained from a CHO test cell line; (b) comparing the test methylation profile of (a) with a reference methylation profile of a CHO reference cell line; Including, A method is provided in which a significant similarity between the test methylation profile and the reference methylation profile of (a) indicates that the two cell lines produce biosimilars, and the test methylation profile and the reference methylation profile are from CpG sites from the CHO cell genome and are determined using a DNA methylation bead-based array.
[0052] The term "biosimilar," as used herein, refers to a recombinant protein produced by genetically modified CHO cells that is highly similar to the original biological therapeutic reference product and shares quality, safety, and efficacy with the reference product. In particular, the produced product is phenotypically / epigenetically similar to the reference product. The term "biosimilar" is more clearly explained at least in A. Ishii-Watabe et al., (2019) Drug Metab. Pharmacokinet. 34(1):64-70 and Wolff-Holz, E. et al., (2019) BioDrugs 33, 621-634.
[0053] Information on the DNA methylation patterns of cell lines can result in clearer specified profiles for product release in CHO cells, serve as "copyright" protection from biosimilar developers, and can be developed as a potential "gold standard" for the regulatory processes required for biosimilar development.
[0054] According to yet another aspect of the present invention there is provided a method for identifying at least one CHO test cell line capable of producing at least one bioidentical to a heterologous protein produced by a CHO reference cell line, the method comprising: (a) determining a test methylation profile from genomic material obtained from a CHO test cell line; (b) comparing the test methylation profile of (a) with a reference methylation profile of a CHO reference cell line; Including, If the test methylation profile and the reference methylation profile of (a) are identical, it indicates that the two cell lines produce the same biological product, and the test methylation profile and the reference methylation profile are from CpG sites from the CHO cell genome and are determined using a DNA methylation bead-based array.
[0055] As used herein, the term "bioidentical" refers to a recombinant protein produced by genetically modified CHO cells that has the same molecular structure as the original biological therapeutic reference product. The term "bioidentical" is more clearly explained at least in Stanczyk FZ, et al., Climacteric. 2021;24:38-45.
[0056] CHO cells capable of producing biosimilar or bioidentical proteins have a CpG methylation profile significantly similar or identical to a reference profile from CHO cells, particularly a parental clone capable of producing a wild-type protein, particularly a protein most similar to a therapeutic protein. In another example, CHO cells producing biosimilar or bioidentical proteins have a methylation profile of selected regions (e.g., but not limited to, hypomethylated regions (LMRs) / partially methylated domains (PMDs) / variably methylated regions (DMRs) / variably methylated points (DMPs)) significantly similar or identical to a reference profile from CHO cells, particularly a parental clone capable of producing a wild-type protein, particularly a protein most similar to a therapeutic protein. In another example, CHO cells producing biosimilar or bioidentical proteins have a significantly higher CpG methylation distribution (e.g., beta value distribution) compared to other CHO cells. In yet another example, CHO cells producing biosimilar or bioidentical proteins have no or minimal amounts of partial methylation at each site compared to other cells. In particular, the heterologous protein is a monoclonal antibody and / or a therapeutic protein.
[0057] A hypomethylated region (LMR) is a region of the genome in which less than 60% of the CpGs within the region are methylated. More specifically, less than 50%, 40%, 30%, 20%, or 10% of the CpGs in an LMR are methylated. Any method known in the art can be used to identify or detect LMRs in genomic DNA. Well-known methods include using programs such as MethylSeekR. In particular, LMRs in genomic DNA have at least three consecutive CpGs and no single nucleotide polymorphisms (SNPs) at any of the CpG positions. Even more specifically, LMRs in genomic DNA are identified based on the methods disclosed at least in Burger, L., (2013) Nucleic Acids Research, 41(16):e155 and / or Stadler, M., (2011) Nature 480, 490-495. LMRs are known to have average methylation ranging from 10% to 50%, are regions of low CG density that do not overlap with CpG islands, tend to be enriched for H3K4me1, DHS, and p300 / CBP, and / or are located primarily distal to promoters in intergenic or intronic regions. - have an average methylation in the range of 10% to 50%; - Areas of low CG density, - enriched in histone H3 monomethylated at lysine 4 (H3K4me1), DNase I hypersensitive sites (DHS), and the transcriptional coactivators CREB-binding protein (CPB) and p300; - located primarily distal to the promoter in intergenic or intronic regions, and / or - No single nucleotide polymorphisms (SNPs) at any of the CpG positions.
[0058] Hypomethylated regions (LMRs) represent a key feature of the dynamic methylome. LMRs are localized decreases in the DNA methylation landscape and represent CpG-poor distal regulatory regions that often reflect the binding of transcription factors and other DNA-binding proteins. LMRs were first described in mice (Stadler et al. (2011), Nature: 480, 490-95). The evolutionary conservation of LMRs across mammals remains unexplored.
[0059] Variably methylated regions (DMRs) are genomic regions that have different methylation states across multiple biological samples, such as tissues, cells, and individuals. These are genomic regions that differ between phenotypes. Statistical power is likely to be greater when neighboring DMRs are considered as a whole [Gu H et al. (2010) Nat Methods 2010;7:133-6]. DMR lengths can range from hundreds to thousands of bases [Rakyan et al. (2011) Nat Rev Genet 12:529-41, 2011, Bock C (2012) Nat Rev Genet 2012;13:705-19].
[0060] DMRs can occur throughout the genome, but have been identified particularly around gene promoter regions, within gene bodies, and in intergenic regulatory regions. There are two types of regions: predefined and user-defined. Regions with special biological significance, such as CpG islands, CpG shores, and UTRs, are predefined. Many traditional statistical tests, including t-tests and Wilcoxon rank-sum tests, can be performed at the region level. For user-defined regions, criteria include a fixed region length, a fixed number of significant and adjacent CpG sites, and a significant and smoothed estimated effect size.
[0061] Partially methylated domains (PMDs) are extended regions within the genome that exhibit reduced average DNA methylation levels. They cover gene-deleted and transcriptionally inactive regions and tend to be heterochromatic.
[0062] Variably methylated positions (DMPs) are CpG sites with different DNA methylation status across different biological samples and are considered as potential functional regions involved in gene transcription regulation.
[0063] According to a further aspect of the present invention there is provided a method for assessing one or more phenotypic parameters of at least one test CHO cell line, the method comprising: (a) determining the test methylation status of one or more preselected methylation sites from genomic material obtained from a test CHO cell line; (b) determining a test methylation profile of the test CHO cell line from the methylation status determined in (a); (c) comparing the test methylation profile determined in (b) with at least one predetermined reference methylation profile, each predetermined reference methylation profile being specific for a reference CHO cell line having at least one phenotypic parameter; Including, If the test methylation profile is significantly similar to one of the predetermined reference methylation profiles, then the test CHO cell line has similar, or preferably the same, phenotypic parameters as the reference CHO cell line having the predetermined reference methylation profile, and the test methylation profile and the reference methylation profile are from CpG sites from the CHO cell genome and are determined using a DNA methylation bead-based array.
[0064] In particular, the phenotypic parameters are - Optimal carbohydrate metabolism - Optimal amino acid metabolism - Optimal lipid metabolism - Optimal protein productivity, and - Optimal cell viability is selected from the group consisting of:
[0065] As used herein, the term "carbohydrate metabolism" refers to nearly all or all biochemical processes involved in the metabolic formation, degradation, and interconversion of carbohydrates within cells. It includes multiple pathways, such as glycolysis, gluconeogenesis, glycogenolysis, and gluconeogenesis. For example, glycolysis is one of the important metabolic pathways of CHO cells. Through glycolysis, CHO cells consume glucose as the main carbon source for energy production and produce lactate as the most common metabolic by-product. In particular, the term "optimal carbohydrate metabolism" refers to the ideal or best carbohydrate metabolism possible by CHO cells.
[0066] Similarly, the term "amino acid metabolism" as used herein refers to the entire biochemical process involved in the metabolic formation, degradation, and interconversion of amino acids within cells. Amino acids are the basic building blocks of proteins and constitute all proteinaceous materials of cells, including the cytoskeleton, protein components of enzymes, receptors, and signaling molecules. Furthermore, amino acids are utilized for cell growth and maintenance. For example, glutaminolysis is an important metabolic pathway for CHO cells. Glutaminolysis is the primary pathway by which CHO cells assimilate organic nitrogen for biomass synthesis while releasing ammonium as a major by-product. In particular, the term "optimal amino acid metabolism" refers to the ideal or best amino acid metabolism possible by CHO cells.
[0067] The term "lipid metabolism" as used herein refers to the synthesis and degradation of lipids in cells, including the breakdown or storage of fats for energy and the synthesis of structural and functional lipids. Lipids are major components of cell membranes and act as secondary messengers in cellular communication, involved in signal transduction, transport, and secretion. Lipids are also important sources of energy through β-oxidation and the tricarboxylic acid (TCA) cycle. Lipid metabolism can have a significant impact on cell growth. For example, the processes of triacylglycerol synthesis and degradation in CHO cells can significantly affect overall cell metabolism and viability. In particular, the term "optimal lipid metabolism" refers to the ideal or best amino acid metabolism possible by CHO cells.
[0068] Carbohydrate, amino acid, and lipid metabolism can be determined by metabolite detection assays, HPLC, and bioprocess analyzers. These methods are further disclosed at least in Coulet, M. et al., Cells (2022), 11, 1929; Fan Y et al., Biotechnol Bioeng (2015) 112(3):521-535, and Ali AS et al., Biotechnol J. (2018); 13(10):e1700745.
[0069] As used herein, the term "preselected methylation sites" refers to methylation sites selected from genes or regions that exhibit the highest degree of methylation variation during training of the method and meet certain quality criteria, for example, a minimum sequencing coverage of 5x or more for five or more eligible CpG sites. Furthermore, genes with an average methylation level of less than 0.1 or an average methylation level of more than 0.9 can be excluded due to their limited dynamic range. A "reference methylation profile" can be defined based on multiple training samples using multivariate statistical methods such as principal component analysis or multidimensional scaling.
[0070] As used herein, the term "significantly similar" refers to a similarity observed by statistical means (i.e., by using bioinformatics) and / or visual observation, particularly in the context of comparing methylation profiles (such as comparing a test profile (from one or more test subjects) to a reference profile). Significant similarity is observed, for example, when a test profile overlaps with a reference profile defined by multiple training samples by multivariate statistical methods such as principal component analysis or multidimensional scaling. In particular, a test profile is significantly similar to a given reference profile if more than 50, 55, 60, 65, 70, 75, 80, 85, 90, 95% of the methylation pattern / profile overlaps with the methylation pattern / profile of the reference profile. Similarity of a test profile to more than one reference profile, such as two, three, or even all of the reference profiles, reduces the significance of the similarity.
[0071] The term "predetermined reference profile" as used herein refers to a typical or standard methylation profile of the genomic material of a CHO cell line having specific characteristics that depend on the context in which the term is used. In one example, with respect to a method for determining a CHO cell line that exhibits at least one phenotypic parameter according to any embodiment of the present invention that confers optimal heterologous protein production potential on the cell line, the term "predetermined reference profile" refers to a typical or standard methylation profile of the genomic material of a CHO cell line that exhibits one or more phenotypic parameters selected from the group consisting of optimal glucose consumption, optimal growth rate, optimal lactate production, and optimal ammonia accumulation. The predetermined reference profile can be obtained from one or more reference CHO cell lines, each expressing one or more phenotypic parameters.
[0072] The method according to this aspect of the invention attempts to generate a methylation profile of a CHO cell line with optimal heterologous protein production potential, as the cell line may exhibit cell viability, adaptability, low cell exhaustion, and a good metabolic readout. In particular, the method according to this aspect of the invention provides a prognostic methylation profile of the ideal parental cell line before transgene introduction.
[0073] According to yet another aspect of the present invention there is provided a method for developing a test system for determining whether a test CHO cell line is capable of optimal heterologous protein production, the method comprising: (a) determining the test methylation status of one or more preselected methylation sites from genomic material obtained from a test CHO cell line; (b) selecting, from the preselected methylation sites, a reference panel of methylation sites characterized by a specific and distinct variable methylation profile for each phenotypic parameter or phenotype of interest; (c) obtaining a test system by assigning a reference methylation profile for each phenotypic parameter or phenotype of interest; Including, A comparison of the test methylation profile obtained from the test sample with the reference methylation profile obtained in (c) allows for confirmation of whether the test CHO cell line is capable of optimal heterologous protein production, wherein the test methylation profile and the reference methylation profile are from CpG sites from the CHO cell genome and are determined using a DNA methylation bead-based array.
[0074] The term "reference panel of methylation sites" refers to specific and distinct CpG sites or regions used to form a reference methylation profile.
[0075] According to yet another aspect of the present invention there is provided a method for determining whether a CHO cell line is robust, stable and capable of optimal heterologous protein production prior to introducing a transgene into the cells, the method comprising: (a) determining a methylation profile from genomic material obtained from a CHO cell line; (b) comparing the methylation profile of (a) with a reference methylation profile for a CHO cell line that is robust, stable, and capable of optimal heterologous protein production; Including, A method is provided in which a significant similarity between the test methylation profile and the reference methylation profile of (a) indicates that the CHO cell line is robust, stable, and capable of optimal heterologous protein production, and the test methylation profile and the reference methylation profile are from CpG sites from the CHO cell genome and are determined using a DNA methylation bead-based array.
[0076] The DNA methylation profile of step (a) according to any embodiment of the present invention is determined using a DNA methylation-based array, in particular a bead-based DNA methylation array. Arrays according to any embodiment of the present invention are advantageous because they allow for understanding of the genomic stability of CHO cell lines and allow for better control over the manufacturing / process development / product development / scale-up / validation processes, thereby aiding in the selection of better CHO cell lines for industrial applications.
[0077] DNA methylation-based arrays enable a high-throughput, robust method for determining semi-quantitative / quantitative DNA methylation information from small samples of extracted DNA of interest. These custom-designed arrays can use Illumina iScan and Infinium platform technology or equivalents, which allows for, for example, 100,000 different bead types covalently attached to DNA methylation probes on each chip. Each probe represents one CpG methylation site at the end of the probe sequence. The DNA sample undergoes bisulfite conversion, amplification, fragmentation, precipitation, and resuspension steps before hybridization on the array chip. Once on the chip, the DNA hybridizes to beads for each CpG site, allowing for specific detection of methylation changes at each site by single-nucleotide extension. This is particularly advantageous because the array-based method is simple and the results of methylation-based arrays are accurate and reproducible.
[0078] Furthermore, compared to conventional sequencing, which takes several weeks to generate data, array technology requires a much shorter turnaround time. The amount and complexity of data generated is less than that of sequencing, making it less computationally intensive. This allows for more rapid calculations to achieve interpretable results from experimental groups. Overall, microarray technology is approximately 10 times faster and 10 times cheaper than conventional sequencing, yet still allows for quantification of methylation levels at specific CpG sites.
[0079] The term "array" as used herein refers to an intentionally created collection of probe molecules, which can be prepared synthetically or biosynthetically. The probe molecules within an array can be identical or different from one another. Arrays can be in a variety of formats, such as libraries of soluble molecules; libraries of compounds tethered to resin beads, silica chips, or other solid supports.
[0080] In particular, DNA methylation-based arrays provide a convenient platform for the simultaneous analysis of a large number of CpG sites, e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 50, 100, 500, 1000, 5000, 10,000, 100,000, or more sites or loci. In particular, arrays contain a plurality of different probe molecules, which may be attached to a substrate or spatially separated within the array. Examples of arrays that may be used in accordance with any embodiment of the present invention include slide arrays, silicon wafer arrays, liquid arrays, bead-based arrays, and the like. In one example, the array technology used in accordance with any embodiment of the present invention combines a miniaturized array platform for sample handling and data processing, a high level of assay multiplexing, and scalable automation.
[0081] In particular, an array according to any embodiment of the present invention may be an array of arrays, also referred to as a composite array, having multiple individual arrays configured to enable the processing of multiple samples simultaneously. Examples of composite arrays and the technology behind them are disclosed at least in U.S. Pat. No. 6,429,027 and U.S. Patent Application Publication No. 2002 / 0102578. A composite array substrate can include multiple individual array locations, each having multiple probes and each physically separated from other assay locations on the same substrate, thereby preventing fluid contacting one array location from contacting another array location. Each array location can have multiple different probe molecules attached directly to the substrate or attached to the substrate via rigid particles in wells (also referred to herein as beads in wells).
[0082] In one example, the array substrate can be an optical fiber bundle or array of bundles, such as those described in U.S. Patent Nos. 6,023,540, 6,200,737, and / or 6,327,410. The optical fiber bundle or array of bundles can have probes attached directly to the fibers or via beads. One skilled in the art will be able to readily determine which substrate is most suitable for arrays according to any aspect of the invention. WO2004110246 further discloses other substrates that can be used for arrays according to any aspect of the invention and methods of attaching beads to substrates.
[0083] In one example, the surface of the substrate can have physical modifications to allow probe attachment or to generate array locations. For example, the surface of the substrate can be modified to include chemically modified sites useful for covalently or non-covalently attaching probe molecules or particles to which probe molecules are attached. The probes can be attached using any of a variety of methods known in the art, including inkjet printing, spotting techniques, photolithographic synthesis, or mask-based printing. WO2004110246 discloses these techniques in more detail.
[0084] In one example, a DNA methylation-based array according to any embodiment of the invention can be a bead-based array, where the beads are associated with a solid support, such as those commercially available from Illumina, Inc. (San Diego, Calif.). Arrays of beads useful according to any embodiment of the invention can be in a fluidic format, such as the fluid flow of a flow cytometer or similar device. Commercially available fluidic formats for distinguishing beads include, for example, those used in Luminex's XMAP™ technology or Lynx Therapeutics' MPSS™ method.
[0085] As used herein, the terms "solid support," "support," and "substrate" are used interchangeably and refer to a material or group of materials having one or more surfaces that are rigid or semi-rigid. In many instances, at least one surface of the solid support is substantially flat, although in some instances it may be desirable to physically separate synthesis regions of different compounds, e.g., with wells, raised areas, pins, etched trenches, etc.
[0086] DNA methylation arrays according to any embodiment of the invention can be very high density arrays, for example, having about 10,000,000 probes / cm to about 2,000,000,000 probes / cm or about 100,000,000 probes / cm to about 1,000,000,000 probes / cm. High density arrays are particularly useful according to any embodiment of the invention due to the large number of CpG sites contained on the array.
[0087] A DNA methylation array according to any aspect of the present invention can be used to analyze or assess multiple such loci simultaneously or sequentially, as desired. In one example, multiple different probe molecules can be attached to a substrate or spatially separated within the array. Each probe is typically specific to a particular locus and can be used to distinguish the methylation status of the locus.
[0088] As used herein, the term "probe molecule" refers to a surface-immobilized molecule that can be recognized by a specific target. The probes used in the array can be specific for the methylated allele of a CpG site, the unmethylated allele of a CpG site, or both.
[0089] The term "target" as used herein refers to a molecule that has affinity for a given probe molecule. Targets can be naturally occurring or artificial molecules. They can also be used in their unchanged state or as aggregates. Targets can be covalently or non-covalently bound to binding members, either directly or via specific binding substances. Examples of targets that can be used according to any embodiment of the present invention are methylated and unmethylated CpG sites. Targets are sometimes referred to in the art as anti-probes. No difference in meaning is intended when the term target is used herein.
[0090] As used herein, the term "complementary" refers to hybridization or base pairing between nucleotides or nucleic acids, for example, between the two strands of a double-stranded DNA molecule or between an oligonucleotide primer and a primer-binding site on a single-stranded nucleic acid to be sequenced or amplified. Complementary nucleotides are generally A and T (or A and U), or C and G. Two single-stranded RNA or DNA molecules are said to be complementary when the nucleotides of one strand are optimally aligned, compared, and, by appropriate nucleotide insertion or deletion, pair with at least about 80%, usually at least about 90%-95%, and more preferably about 98%-100% of the nucleotides of the other strand. Fully complementary refers to 100% complementarity over the length of the sequence. For example, a 25-base probe is fully complementary to a target if all 25 bases of the probe are complementary to the contiguous 25-base sequence of the target, with no mismatches between the probe and target over the length of the probe.
[0091] According to another aspect of the present invention, there is provided a method for determining the regulation of transgene expression in at least one CHO cell line genetically modified with a transgene, comprising the steps of: - measuring the methylation level of at least one CpG site in at least one promoter of the transgene, the promoter is a viral promoter, Methods are provided in which DNA methylation levels are determined using a DNA methylation bead-based array.
[0092] According to another aspect of the present invention there is provided a method for determining the regulation of transgene expression in at least one CHO cell line genetically modified with a transgene, the method comprising: - measuring the methylation level of at least one CpG site in at least one promoter of the transgene, the promoter is selected from cytomegalovirus (CMV) and simian vacuolating virus 40 (SV40); Methods are provided in which DNA methylation levels are determined using a DNA methylation bead-based array.
[0093] As used herein, the term "promoter" or "gene promoter," which is used interchangeably with the term "regulatory region" or "regulatory sequence," refers to a corresponding contiguous gene DNA sequence extending from 1.5 kb upstream to 1.5 kb downstream relative to the transcription start site (TSS), or a contiguous portion thereof. Specifically, a "regulatory region" refers to a corresponding contiguous gene DNA sequence extending from 1.5 kb upstream to 0.5 kb downstream relative to the TSS. In some instances, a "regulatory region" refers to a corresponding contiguous gene DNA sequence extending from 1.5 kb upstream to the downstream end of a CpG island that overlaps with the region 1.5 kb upstream to 1.5 kb downstream of the TSS (and thus, in such cases, may extend beyond 1.5 kb downstream), and a contiguous portion thereof. Altering DNA methylation on gene promoters responsible for protein glycosylation can lead to improved protein quality. Protein glycosylation is an important quality attribute that regulates the efficacy, stability, and half-life of therapeutic proteins. Due to regulatory concerns, it is desirable to obtain consistent glycoform profiles in protein production. Therefore, DNA methylation can act as a tool to globally regulate CHO metabolism and protein production.
[0094] According to a further aspect of the present invention, there is provided the use of DNA methylation profiling to identify at least one suitable site or region within the genome of a CHO cell line for the introduction of at least one transgene. In particular, information about the CHO epigenome can be used to identify suitable transgene insertion sites based on methylation patterns that are optimal "hotspots" for transgene expression. For example, specific LMRs can be identified in the genome of a CHO cell line for targeted insertion of at least one transgene, e.g., because highly methylated sites are lenticular and not productive for transgene expression (TIS analysis). In another example, the CMV promoter and surrounding repetitive elements can also be identified as hotspots for transgene insertion using methylation profiling.
[0095] Methylation profiling can also be used to screen and select suitable promoters for use in CHO cells that result in optimal transgene expression. In particular, methylation data from different promoters and transgene insertion sites can be obtained and compared to select the best-performing promoters that may result in improved transgene expression. In particular, arrays according to any embodiment of the present invention can be used to monitor transgene activity (expression or silencing / imprinting) by quantifying DNA methylation levels of transgene promoters.
[0096] According to yet another aspect of the present invention, there is provided a DNA methylation bead-based array comprising at least: - a plurality of distinct locations, each location having at least one probe molecule comprising a nucleic acid sequence complementary to a plurality of CpG sites of the CHO cell; A DNA methylation bead-based array is provided, wherein the CpG sites of the CHO cells are from the CHO genome and selected from at least one CpG of Tables 5a-5f.
[0097] These CpG sites are environment-specific CpG sites (i.e., dynamic CpG sites) and are CpG sites found in the promoters of metabolism-related genes, protein production-related genes, cell growth and division-related genes, and epigenetic-related genes and the genes themselves.
[0098] "Environment-specific CpG sites," also known as dynamic CpG sites in the context of CHO cells, refer to CpG sites that are variably methylated among different CHO cell lines. The cell lines used in this analysis include CHO-K1 (ATCC), CHO-DG44 (Thermo Fisher Scientific), CHO-DXB11 (ATCC), ExpiCHO-S™ cells (Thermo Fisher Scientific), FreeStyle™ CHO-S™ cells (Thermo Fisher Scientific), and CHO 1-15 500 (ATCC) and Agarabi CHO (ATCC).
[0099] "Metabolic-related genes" in the context of CHO cells herein refer to genes associated with several metabolic pathways, such as glycolysis, the TCA cycle, the pentose phosphate pathway, the malate-aspartate shuttle, amino acid metabolism, lactate metabolism, cholesterol biosynthesis, nucleotide biosynthesis, nucleotide sugar biosynthesis, etc. Some examples of such genes include Hk2, Pgk1, Idh3a, Pgm1, and Pdha1. One skilled in the art would readily determine the genes found in CHO cells that fall into this category.
[0100] "Protein production-related genes," as used in connection with CHO cells herein, refer to genes associated with cellular processes such as DNA replication and repair, mRNA transcription, mRNA translation, post-translational modification, and protein folding and transport. Some examples of such genes include Gatb, Sec61a2, Ube2e3, Exosc1, Dna2, Pold1, etc. One of skill in the art would be able to readily determine other genes found in CHO cells that fall into this category.
[0101] "Cell growth and division-related genes," as used herein in reference to CHO cells, refer to genes associated with cellular processes such as cell cycle regulation, cytoskeleton-related elements, cell signaling, nucleotide metabolism, and cell death. Some examples of such genes include Camk1, Cd82, Cdk4, Col1a1, and Ctsb. Again, one of skill in the art would be able to readily determine other genes found in CHO cells that fall into this category.
[0102] "Epigenetic-related genes," as used herein in connection with CHO cells, refer to genes associated with epigenetic modifications, such as DNA methylation pathways, DNA demethylation pathways, folate and methionine cycles, and histone modifications. Some examples of such genes include Hat1, Shmt1, Bhmt, Dnmt1, and Ehmt1. One skilled in the art would be able to readily determine other genes found in CHO cells that fall into this category.
[0103] The term "viral promoter" as used herein in relation to CHO cells refers to promoters and enhancers of at least cytomegalovirus (CMV) and simian vacuolating virus 40. Viral promoters are typically rich in CpG sites, making them susceptible to DNA methylation and thus repressing protein expression.
[0104] The methods according to any embodiment of the present invention may also be used to predict whether a CHO test cell is capable of optimal heterologous protein production. [Example]
[0105] The above describes preferred embodiments, which, as will be understood by those skilled in the art, may be subject to changes or modifications in design, construction, or operation without departing from the scope of the claims. These variations are intended to be covered, for example, by the claims. [Example 1] Oxidative stress in CHO cell cultures Wet-Lab methodology For this experiment, the transgenic CHO cell line Agarabi CHO (ATCC® CRL-3440™) was grown in CD FortiCHO medium supplemented with 8 mM L-glutamine at 37°C, 8% CO2, and a shaking speed of 130 RPM. Six flasks of batch culture were maintained for 7 days, with three flasks representing technical replicates for the control set and three flasks representing technical replicates for the treatment set. Flasks were seeded with 3E5 viable cells / mL on day 0, and hydrogen peroxide was added to the treatment set at a final concentration of 120 μM every 48 hours to induce oxidative stress. Cell count, cell viability, and heterologous protein production were measured every two days, and cell pellets were collected on day 7 for both the control and treatment sets. Induction of oxidative stress in CHO cells by treatment with hydrogen peroxide resulted in a decrease in growth rate and cell viability compared to the control set, and therefore, there was a slight increase in heterologous protein productivity in the treatment set.
[0106] Genomic DNA was purified from the collected cell pellets using the DNeasy Blood & Tissue Kit (Qiagen) and quantified using PicroGreen or NanoDrop™ 2000. Genomic DNA (500 ng) from the control and treatment sets was used to prepare libraries for whole genome bisulfite sequencing (WGBS). Libraries were sequenced by a third party on the NovaSeq platform, which generated 125 GB of data per sample.
[0107] calculation methodology The raw sequencing data were subjected to quality control (fastqc)1, trimming of sequencing adapters (TrimGalore)2, and alignment with Bismark3. The CMV promoter combined with the CHOK1-GS (Cricetulus griseus) genome was used as the reference genome. Bismark was also used to remove duplicate reads and extract methylation counts from the alignment output. SNPs were excluded, and only counts with a minimum coverage of 10x were used for downstream analysis, resulting in 3,711,013 CpG sites for the hydrogen peroxide-treated sample. Because regulated methylation targets are most commonly clustered in short regions, modified single-linked clustering of methylation sites was performed using DMRfinder4. With a maximum distance between CpG sites of 100 bp, 1,728,014 genomic regions were found for the hydrogen peroxide-treated sample.
[0108] Analysis of variable methylation Variable methylation analysis was performed between the control and treatment groups using MethylKit5. Logistic regression was used to determine variable methylation across all regions, and the sliding linear model (SLIM) method for FDR correction. Regions with an FDR-corrected p-value of less than 0.05 and a methylation change of more than 25% between groups were determined as variable methylated regions (DMRs), which were 122 in the hydrogen peroxide-treated samples, as shown in Table 1. Principal component analysis (PCA) is a dimensionality reduction technique that highlights variation in a dataset. PCA analysis of DMRs is shown in Figure 1.
[0109] Preliminary results indicate that DMRs play a role in the epigenetic changes of oxidative stress, which could potentially be used as markers for future studies.
[0110] [Table 1]
[0111] [Example 2] Adaptation of CHO cells with medium supplements Wet-Lab methodology For this experiment, the transgenic CHO cell line Agarabi CHO (ATCC® CRL-3440™) was adapted for two weeks at 37°C, 8% CO2, and 130 RPM shaking in CD FortiCHO medium supplemented with 8 mM L-glutamine and 1 mg / L human insulin-like growth factor 1 (IGF-1). Batch cultures of six flasks were maintained for seven days, with three flasks representing technical replicates for the control set (no IGF-1 adaptation) and three flasks representing technical replicates for the IGF-1 adaptation set. Flasks were seeded with 3E5 viable cells / mL on day 0, and 1 mg / L insulin growth factor was added to the adaptation set. Cell counts, cell viability, and protein production were measured every two days, and cell pellets were collected on day 7 for both the control and treatment sets. Adaptation of CHO cells with IGF-1 did not significantly affect growth rate and viability, but heterologous protein productivity was doubled compared to the control set.
[0112] Genomic DNA was purified from the collected cell pellets using the DNeasy Blood & Tissue Kit (Qiagen) and quantified using PicroGreen or NanoDrop™ 2000. Genomic DNA (500 ng) from the control and adaptive sets was used to prepare libraries for whole genome bisulfite sequencing (WGBS). Library sequencing was performed by a third party on the NovaSeq platform, which generated 125 GB of data per sample.
[0113] calculation methodology The raw sequencing data were subjected to quality control (fastqc)1, trimming of sequencing adapters (TrimGalore)2, and alignment with Bismark3. The CMV promoter combined with the CHOK1-GS (Cricetulus griseus) genome was used as the reference genome. Bismark was also used to remove duplicate reads and extract methylation counts from the alignment output. SNPs were excluded, and only counts with a minimum coverage of 10x were used for downstream analysis, resulting in 4,244,091 CpG sites for the IGF-1-adapted sample. Because regulated methylation targets are most commonly clustered in short regions, modified single-linked clustering of methylation sites was performed using DMRfinder4. When the maximum distance between CpG sites was 100 bp, 2,048,904 genomic regions were found for the IGF-1-adapted sample.
[0114] Analysis of variable methylation Variable methylation analysis was performed using MethylKit5 between the control and adaptive groups. Logistic regression was used to determine variable methylation across all regions, and the sliding linear model (SLIM) method for FDR correction. Regions with an FDR-corrected p-value of less than 0.05 and a methylation change of more than 25% between groups were determined as variable methylated regions (DMRs), of which there were 289 in the IGF-1 adaptive samples, as listed in Table 2. Principal component analysis (PCA) is a dimensionality reduction technique that highlights variation in a dataset. PCA analysis of DMRs is shown in Figure 2.
[0115] Preliminary results indicate that DMRs play a role in the epigenetic changes of IGF-1 adaptation, which could potentially be used as markers for future studies.
[0116] [Table 2a]
[0117] [Table 2b]
[0118] [Example 3] Quality of heterologous proteins from CHO cells Wet-Lab methodology For this experiment, a transgenic CHO cell line, Humira431 clone (obtained from A*Star BTI), was grown in EX-CELL Advanced fed-batch medium supplemented with 6 mM L-glutamine at 37°C, 8% CO2, and 150 RPM shaking. Six flasks of fed-batch culture were maintained for 11 days, with three flasks representing technical replicates for the control set (C1, C2, C3) and three flasks representing technical replicates for the treatment set (T1, T2, T3). Flasks were seeded with 3E5 viable cells / mL on day 0, and cultures were fed with EX-CELL® Advanced CHO Feed 1 on days 3, 5, 7, and 9, supplementing glucose to 6 g / L using 45% glucose if glucose dropped below 3 g / L. To induce hyperosmolality in the cell culture medium, concentrated sodium chloride solution was added to the treatment set on day 3, increasing the medium osmolality from 320 mOsm / kg to 480 mOsm / kg. Cell count, cell viability, and heterologous protein production were measured every two days, and cell pellets were collected on day 7 for both the control and treatment sets. Induction of hyperosmolality in the CHO cell medium with sodium chloride resulted in a decrease in growth rate (Figure 3a), an increase in heterologous protein productivity (Figure 3b), and a change in the relative abundance of each N-glycan modification in the treatment set compared to the control set (Table 3). The alteration in the relative abundance of each N-glycan represents a change in the quality of the heterologous protein.
[0119] DNA extraction DNA was extracted using the PureLink Genomic DNA Isolation Minikit (Invitrogen), including RNAase treatment according to the manufacturer's instructions. DNA quantity was measured by PicoGreen assay to ensure the A260 / 280 ratio was ≤1.8, and DNA quality was assessed by NanoDrop (Thermo Scientific). Small samples were then analyzed using automated electrophoresis on a TapeStation (Agilent) to ensure each sample contained high-molecular-weight DNA.
[0120] Sequencing analysis Genomic DNA (500 ng) from the samples was used to prepare libraries for whole-genome bisulfite sequencing (WGBS). Libraries were sequenced by a third party on the NovaSeq platform, which generated 125 GB of data per sample at 20x coverage.
[0121] Data Processing: Sequencing data processing and analysis: Raw sequencing data were quality controlled (fastqc)1, trimmed sequencing adapters (TrimGalore)2, and aligned with Bismark3.
[0122] Bismark was also used to remove duplicate reads and extract methylation counts from the alignment output. SNPs were excluded, and only counts with a minimum coverage of 10x were used for downstream analysis.
[0123] The methylation ratios of the control (C1) and treated (T1) samples were then extracted. Sites with a methylation difference of 30% were then filtered (Table 4). These methylation sites may indicate differences in protein quality between samples.
[0124] [Table 3]
[0125] [Table 4a]
[0126] [Table 4b]
[0127] [Table 4c]
[0128] [Example 4] Heterologous protein level detection from CHO cells Wet-Lab methodology For this experiment, five transgenic CHO clones (obtained from A*Star BTI) were grown in EX-Cell Advanced fed-batch medium supplemented with 6 mM L-glutamine at 37°C, 8% CO2, and 225 RPM shaking. The five transgenic CHO cell lines included low producers (3D11, 2C9, 2H2), intermediate producers (10A8), and high producers (8F8, 7H9). Flasks were seeded at 3E5 viable cells / mL on day 0, and cultures were fed with Cell Boost 7a on days 3, 5, 7, 9, and 11. When glucose dropped below 2 g / L, it was replenished to 6 g / L using 45% glucose. Fed-batch cultures of the six clones were maintained for 14 days. Cell count, cell viability, and heterologous protein production were measured every two days, and cell pellets were harvested on day 9. The specific productivity (pg / cell / day) of all six clones was calculated for days 9, 11 and 14, as shown in FIG.
[0129] DNA extraction DNA was extracted using the PureLink Genomic DNA Isolation Minikit (Invitrogen), including RNAase treatment according to the manufacturer's instructions. DNA quantity was measured by PicoGreen assay to ensure the A260 / 280 ratio was ≤1.8, and DNA quality was assessed by NanoDrop (Thermo Scientific). Small samples were then analyzed using automated electrophoresis on a TapeStation (Agilent) to ensure each sample contained high-molecular-weight DNA.
[0130] Bisulfite conversion and BeadChip analysis Genomic DNA samples were then subjected to bisulfite conversion using the EZ DNA Methylation-Gold™ Kit (Zymo Research). Methylation levels were then quantified using our customized methylation BeadChip kit (Illumina), allowing for quantitative analysis of over 50,000 methylation sites across the genome at single-nucleotide resolution. After bisulfite conversion, samples were processed through a three-day workflow, including sample amplification, fragmentation, precipitation, hybridization to BeadChips, and X-staining, according to the Infinium HD Methylation Assay (Illumina, ref. #15019519 v07), before being imaged on an iScan (Illumina), which generates intensity files for beta calculations.
[0131] Data Processing: Processing Beadchip data: Customized chip array data processing was performed in R version 4.1.2 using Sesame version 1.14.2. DNA methylation levels at each site were calculated as methylation β values. β values are defined as methylation signal / (methylation signal + unmethylation signal). This can be calculated using the getBetas function. Normalized β values were generated and quality control was performed using the SeSAMe pipeline (Zhou et al., 2018). Low-intensity detection calling and generation (based on p-values) were performed using pOOBAH. Background subtraction based on normal exponential deconvolution with extra bleed-through subtraction, if necessary, was also performed using the out-of-band probe noob (Triche et al., 2013).
[0132] After obtaining the beta values, the control probes were removed from the data frame. CpG sites with NA beta values were also removed from the data frame.
[0133] To obtain differentially methylated positions (DMPs) between high-protein-producing clones (7H9 and 8F8) and low-protein-producing clones (2C9, 2H2, and 3D11), sample 10A8 was excluded from the beta value data frame before extracting DMPs. After excluding 10A8, DMPs between high-protein-producing clones and low-protein-producing clones were extracted using the dml and dmr functions from the sesame package. The dmr function resulted in a data frame. To obtain more statistically significant DMPs, only DMPs with Pr(|t|) < 0.05 were retained, and the rest were removed from the data frame. This left 901 CpG sites (after removing probes with NAs). A PCA plot of these sites was plotted using prcomp followed by the autoplot function. These citations are shown in Table 5.
[0134] [Table 5a]
[0135] [Table 5b]
[0136]
Table 5c
[0137]
Table 5d
[0138]
Table 5e
[0139]
Table 5f
Claims
1. 1. A method for determining the suitability of at least one Chinese hamster ovary (CHO) test cell line for optimal heterologous protein production, said method comprising: (a) determining a test methylation profile from genomic material obtained from said CHO test cell line; (b) comparing the test methylation profile obtained from (a) with a reference methylation profile, wherein the reference methylation profile comprises the methylation status of two or more CpG sites from at least one CHO reference cell line exhibiting at least one phenotype of interest for optimal heterologous protein production; Including, a significant similarity of the test methylation profile of (a) compared to the reference methylation profile indicates that the CHO test cell line is suitable for optimal heterologous protein production; the test methylation profile and the reference methylation profile are from CpG sites from a CHO cell genome and are determined using a DNA methylation bead-based array; method.
2. 2. The method of claim 1, wherein the reference methylation profile is a compilation of two or more CpG sites from at least one CHO reference cell line exhibiting at least one phenotype of interest for optimal heterologous protein production.
3. 3. The method of claim 1 or 2, wherein the phenotype of interest for the optimal heterologous protein is selected from the group consisting of phenotypic uniformity, protein productivity, and protein quality.
4. 1. A method for selecting at least one CHO cell comprising a phenotype of interest from a population of CHO cells from a parent clone, the method comprising: (a) determining a test methylation profile from genomic material obtained from said CHO cells; (b) comparing the test methylation profile of (a) with a reference methylation profile from a parent clone exhibiting the phenotype of interest; Including, (b) a significant similarity between the test methylation profile and the reference methylation profile indicates that the cell has the phenotype of interest of the parent clone; the test methylation profile and the reference methylation profile are from CpG sites from the CHO cell genome and are determined using a DNA methylation bead-based array; The phenotype of interest is selected from the group consisting of phenotypic uniformity, protein productivity, and protein quality. method.
5. 1. A method for identifying at least one CHO test cell line capable of producing at least one biosimilar to a heterologous protein produced by a CHO reference cell line, the method comprising: (a) determining a test methylation profile from genomic material obtained from said CHO test cell line; (b) comparing the test methylation profile of (a) to the reference methylation profile of the CHO reference cell line; Including, a significant similarity between the test methylation profile and the reference methylation profile of (a) indicates that the two cell lines produce biosimilars; the test methylation profile and the reference methylation profile are from CpG sites from the CHO cell genome and are determined using a DNA methylation bead-based array. method.
6. 1. A method for identifying at least one CHO test cell line capable of producing at least one bioidentical to a heterologous protein produced by a CHO reference cell line, the method comprising: (a) determining a test methylation profile from genomic material obtained from said CHO test cell line; (b) comparing the test methylation profile of (a) to the reference methylation profile of the CHO reference cell line; Including, if the test methylation profile and the reference methylation profile of (a) are identical, it indicates that the two cell lines produce biologically identical products; the test methylation profile and the reference methylation profile are from CpG sites from the CHO cell genome and are determined using a DNA methylation bead-based array. method.
7. 1. A method for assessing one or more phenotypic parameters of at least one test CHO cell line, said method comprising: (a) determining the test methylation status of one or more preselected methylation sites from the genomic material obtained from the test CHO cell line; (b) determining a test methylation profile for the test CHO cell line from the methylation status determined in (a); (c) comparing the test methylation profile determined in (b) to at least one predetermined reference methylation profile, each of the predetermined reference methylation profiles being specific for a reference CHO cell line having at least one phenotypic parameter; Including, if the test methylation profile is significantly similar to one of the predetermined reference methylation profiles, then the test CHO cell line has similar, or preferably the same, phenotypic parameters as the reference CHO cell line having the predetermined reference methylation profile; the test methylation profile and the reference methylation profile are from CpG sites from the CHO cell genome and are determined using a DNA methylation bead-based array. method.
8. 8. The method of claim 7, wherein the phenotypic parameter is selected from the group consisting of optimal carbohydrate metabolism, optimal amino acid metabolism, optimal lipid metabolism, optimal protein productivity, and optimal cell viability.
9. 1. A method for developing a test system to determine whether a test CHO cell line is capable of optimal heterologous protein production, said method comprising: (a) determining the test methylation status of one or more preselected methylation sites from the genomic material obtained from the test CHO cell line; (b) selecting from said preselected methylation sites a reference panel of methylation sites characterized by a specific and distinct variable methylation profile for each phenotypic parameter or phenotype of interest; (c) obtaining a test system by assigning a reference methylation profile for each of said phenotypic parameters or phenotypes of interest; Including, Comparison of the test methylation profile obtained from the test sample with the reference methylation profile obtained in (c) allows to determine whether the test CHO cell line is capable of optimal heterologous protein production; the test methylation profile and the reference methylation profile are from CpG sites from the CHO cell genome and are determined using a DNA methylation bead-based array. method.
10. 1. A method for determining whether a CHO cell line is robust, stable, and capable of optimal heterologous protein production prior to introducing a transgene into the cells, said method comprising: (a) determining a methylation profile from genomic material obtained from said CHO cell line; (b) comparing the methylation profile of (a) with a reference methylation profile for a CHO cell line that is robust, stable, and capable of optimal heterologous protein production; Including, a significant similarity between the test methylation profile of (a) and the reference methylation profile indicates that the CHO cell line is robust, stable, and capable of optimal heterologous protein production; the test methylation profile and the reference methylation profile are from CpG sites from a CHO cell genome and are determined using a DNA methylation bead-based array; method.
11. 11. The method of any one of claims 1 to 10, wherein the CpG site comprises at least one of the CpG sites provided in Tables 5a to 5f.
12. 1. A method for determining the regulation of transgene expression in at least one CHO cell line genetically modified with a transgene, said method comprising: - measuring the methylation level of at least one CpG site of at least one viral promoter of said transgene; Including, wherein the DNA methylation level is determined using a bead-based DNA methylation array. method.
13. 1. A DNA bead-based methylation array comprising at least: a plurality of distinct locations, each location having at least one probe molecule comprising a nucleic acid sequence complementary to a plurality of CpG sites in the CHO cell; Including, the CpG sites of the CHO cells are at least selected from Tables 5a to 5f; DNA bead-based methylation array.