Prediction of quality of recombinant protein from CHO cells
DNA methylation analysis of CHO cells at specific CpG sites predicts protein quality, addressing inefficiencies in CHO protein production by ensuring consistent protein quality and reducing time and costs.
Patent Information
- Application Number
- PCT/EP2025/069085
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-15
- Filing Date
- 2025-07-04
- Publication Date
- 2026-01-22
AI Technical Summary
Current methods for determining the quality of recombinant proteins produced by CHO cells are time-consuming and not very accurate, leading to inefficiencies and financial losses due to heterogenous phenotypes and genetic instability in protein production.
A method using DNA methylation analysis of specific CpG sites in CHO cells to predict the quality of heterologous proteins before large-scale production, utilizing bead-based DNA methylation arrays and machine learning models to ensure phenotypic similarity with reference cells producing known quality proteins.
Enables efficient selection of CHO cells capable of producing high-quality proteins, reducing time and costs by ensuring consistent protein quality through epigenetic analysis, thereby improving production efficiency and consistency.
Smart Images

Figure IMGF000022_0001 
Figure IMGF000026_0001 
Figure IMGF000007_0001
Abstract
Description
[0001] PREDICTION OF QUALITY OF RECOMBINANT PROTEIN FROM CHO CELLS
[0002] FIELD OF THE INVENTION
[0003] The present invention relates to a method based on epigenetics for predicting the quality of target protein produced by a CHO cell and a method for selecting a CHO cell capable of producing heterologous proteins of known, good or high quality. In particular, the measure of differential methylation of promotors and / or CpG sites of CHO cells may provide an insight into the quality of protein produced by the CHO cells.
[0004] BACKGROUND OF THE INVENTION
[0005] Chinese Hamster Ovary (CHO) cells are known to be the workhorses for the industrial production of recombinant therapeutic proteins since 1987 and are hence widely used for biologies production. About 70% of all recombinant biopharmaceutical proteins and all monoclonal antibodies approved since 2016 are being manufactured in CHO cells. Several advantages of utilizing CHO for biologies production include tolerance to genetic manipulations, ease of adaptation to manufacturing process scales, rapid growth rates, and ability to perform human-compatible post-translational modifications. However, the biologies production system in CHO faces a problem as the proteins produced by the cells do not always mirror the native protein and particularly lack the post-translational modification.
[0006] In the current market of bioproduction, an efficient and robust analytical method is in high demand to monitor recombinant protein production before, during and after the process of production. Such a method can fulfil the multi-metabolic demands of different types of clones and cell lines. Unfortunately, most of the methods currently available can only determine the quality of the heterologous protein produced after the CHO cell producing the protein is cultured and the protein analysed. In particular, the current methods of determining the suitability of a CHO clone for target protein production are not only time-consuming, but also not very accurate for selection of clones or cells for optimal protein production. Furthermore, genetically identical CHO clones can still result in products with heterogenous phenotypes or characteristics, thus creating instability, inefficiency, and financial losses during heterologous protein production at an industrial scale. Accordingly, there is still a need in the art for such a robust analytical tool to optimise methods of producing recombinant proteins in CHO cells.
[0007] BRIEF DESCRIPTION OF THE FIGURES
[0008] Figure 1 is a PCA plot for 16447 CpG sites plotted using the preomp followed by ggplot2 functions in Example 1 .
[0009] Figure 2 is a PCA plot for the Differentially Methylated Positions (DMPs) obtained from the list of 16447 CpG sites extracted using the dml function from the sesame package. The 238 CpG sites remaining as shown in Table 1 were plotted using the preomp followed by ggplot2 functions.
[0010] Figure 3 is a PCA plot of 16447 CpG sites from Example 2 that are plotted using the preomp followed by ggplot2 functions. Figure 4 is a is a PCA plot for the Differentially Methylated Positions (DMPs) between charge variant groups obtained from the list of 16447 CpG sites extracted using the dml function from the sesame package. The 191 CpG sites remaining as shown in Table 2 were plotted using the prcomp followed by ggplot2 functions.
[0011] DESCRIPTION OF THE INVENTION
[0012] The present invention solves the problems above by providing a means of predicting the capability of a CHO test cell to produce heterologous proteins of a known, possibly good or high quality by measuring DNA methylation in specific CpG sites of the test CHO cell. In particular, the heterologous protein produced is considered to be of known, possibly good or high quality based on the number of protein quality metrics that is present or expressed in the protein produced by the CHO test cell. By providing a means of predicting the quality of the heterologous protein produced by the CHO test cell prior to culturing of the cell in a large scale to determine the quality of the protein produced, time, energy and money is saved. The method according to any aspect of the present invention improves speed, quality, efficiency and consistency of present heterologous protein production by CHO cells. Since genotype comparisons of CHO clones cannot define how genes are expressed differentially to adapt to environmental conditions, and phenotypic analyses alone are not able to guarantee consistency over time, epigenetic methods, specifically DNA methylation therefore provides a state-of-the-art technology to select not only genetically identical, but epigenetically and therefore phenotypically identical CHO clones for improved heterologous protein production. The method according to any aspect of the present invention allow the use of DNA methylation as a tool to predict heterologous protein production ensuring efficient production of mainly good quality or known quality proteins.
[0013] According to one aspect of the present invention, there is provided a method of predicting quality of a heterologous protein to be produced from a test Chinese Hamster Ovary (CHO) cell, the method comprising the steps of:
[0014] (a) determining a test methylation profile of one or more pre-selected CpG sites within the DNA of the test CHO cell;
[0015] (b) comparing the test methylation profile obtained from (a) with a reference methylation profile, wherein the reference methylation profile comprises the methylation status of the preselected CpG sites from at least one CHO reference cell line that is capable of producing a known, preferably good, quality heterologous protein; wherein a significant similarity in the test methylation profile of (a) compared to the reference methylation profile, is indicative of the heterologous protein produced by the test cell being of known, preferably good, quality; and wherein a significant difference in the test methylation profile of (a) compared to the reference methylation profile, is indicative of the heterologous protein produced by the test cell being of unknown or poor quality; wherein the test and reference methylation profiles are determined using a bead-based DNA methylation array; and wherein the quality of the heterologous protein is determined by the presence of at least one protein quality metric in the protein produced.
[0016] In particular, the protein quality metric in the heterologous protein produced is analogous to a reference protein produced from the CHO reference cell line(s).
[0017] According to a further aspect of the present invention, there is provided a method of selecting at least one CHO cell capable of producing at least one heterologous protein with a desired protein quality metric from a population of CHO cells from a parental clone, the method comprising the steps of:
[0018] (a) determining a test methylation profile from genomic material obtained from the CHO cell, and
[0019] (b) comparing the test methylation profile of (a) with a reference methylation profile, wherein the reference methylation profile comprises the methylation status of more than one CpG site from at least one CHO reference cell line that produces a heterologous protein with the desired or known protein quality metric; wherein a significant similarity in the methylation profile of (a) compared to the reference methylation profile, is indicative of the CHO cell being capable of producing heterologous proteins with the desired or known protein quality metric; and wherein a significant difference in the test methylation profile of (a) compared to the reference methylation profile, is indicative of the CHO cell capable of producing heterologous proteins without the desired or known protein quality metric; and wherein the methylation profile of (a) and the reference methylation profile are determined using beadbased DNA methylation array.
[0020] Epigenetic technologies thus provide a solution for the quantitative and qualitative analysis of heterologous proteins produced from CHO cells. In particular, the reference methylation profile may comprise environmental specific CpG sites or dynamic CpG sites i.e., sites which seem to have a crucial role in several environmental conditions; CpG sites in the viral promoters (CMV and SV40 promoters), CpG sites from regulatory regions of candidate genes from pathways which are significant in certain important biological processes for the CHO cell (e.g. metabolic linked genes, protein production linked genes, cell growth / division linked genes, and methylation linked genes) and / or CpG sites in low methylated regions (LMR) / partially methylated domains (PMD) / differentially methylated regions (DMR) / differentially methylated points (DMP). More in particular, the test and reference methylation profiles are from CpG sites from wherein the test methylation profile and reference methylation profile are from CpG sites from the CHO cell genome.
[0021] The term ‘CHO cell genome’ herein refers to the genomic DNA of the CHO cell that excludes the DNA of a virus, particularly CMV and SV40, that are used to introduce foreign DNA to the cell. In particular, the CHO cell genome may denote the cell with a genome make-up that is in a form as seen naturally in the wild. The term may also include genes which have been added to the CHO genome by genetic modification (i.e. with regard to improved production of protein etc.) but not necessarily or not genes and promoters of viruses that have been used to introduce the genes into the CHO genome. The term “CHO cell genome” therefore may exclude virus genes and promoters and / or may include endogenous or homologous genes of the CHO cell and / or genetically modified endogenous or homologous genes of the CHO cell and / or intergenic genes, DNA found between the genes of the CHO cell.
[0022] Machine learning models may be used to analyse the quantitative and qualitative methylation data generated. Even more in particular, the methods according to any aspect of the present invention may be used in a predictive and precise way to determine suitable CHO cells that can be used to produce good and / or known quality of heterologous proteins, compared to the current methods of trial and error that are used. This thus allows the online and direct control of manufacturing processes, increasing the robustness and thus overall, the quality of molecules produced by the CHO cells.
[0023] The CHO cell line refers to immortal Chinese Hamster Ovary cell line (CHO) derived from Cricetulus griseus. In particular, the CHO cell line may be selected from the group consisting of CHO-K1 (ATCC), CHO-DG44 (Thermo Fisher Scientific), CHO-DXB1 1 (ATCC), ExpiCHO-S™ cells (Thermo Fisher Scientific), Freestyle™ CHO-S™ cells (Thermo Fisher Scientific), CHO 1 -15 [subscript 500] (ATCC) and Agarabi CHO (ATCC). In particular, a CHO cell line also refers to a collection of cells originating from one cell. For example, a CHO reference cell line refers to a cell a population or group of cells originating from a single cell that is capable of producing known or good quality heterologous proteins.
[0024] The term ‘heterologous protein’ as used herein refers to a protein which is not endogenous to the cell. It means an expression of a gene or part of a gene, particularly a transgene in a host CHO cell which does not naturally express this gene. The assays that are commonly used to quantify heterologous protein production include enzyme-linked immunosorbent assay (ELISA), chromatography & bioprocess analyser. The term ‘host cell’ as used herein refers to a cellular system for the expression of heterologous protein. For example, CHO cells are the main hosts for the production of various therapeutic proteins.
[0025] As used herein, the term ‘therapeutic protein’ refers to genetically engineered versions of naturally occurring human proteins. Examples of therapeutic proteins include antibody-based drugs, anticoagulants, blood factors, bone morphogenetic proteins, engineered protein scaffolds, enzymes, growth factors, hormones, interferons, interleukins and the like.
[0026] As used herein, the term ‘transgene’ refers to a gene that is taken from the genome of one organism and inserted into the genome of another organism by artificial techniques used in genetic modification. For example, a human gene is artificially introduced into the genome of CHO cells for the production of at least one protein of interest, particularly therapeutic proteins.
[0027] The term ‘quality of heterologous protein’ refers to the posttranslational modification of the protein that determines the efficacy and function of the protein. The post-translational modifications generally include phosphorylation, glycosylation, ubiquitination, S-nitrosylation, methylation, N-acetylation, sumoylation, oxidation, deamidation, lipidation, and proteolysis (https: / / www.thermofisher.com / de / de / home / life- science / protein-biology / protein-biology-learning-center / protein-biology-resource-library / pierce-protein- methods / overview-post-translational-modification.html). In particular, the quality of the heterologous protein is determined by the presence of at least one protein quality metric in the protein produced that is analogous to a reference protein produced from the CHO reference cell line. More in particular, the protein quality metric is selected from the group consisting of glycosylation, propensity to form aggregates, charge variation, and amidation. Details of these terms and means of measurement are provided at least in Ha TK, et al. Biotechnol Adv. 2022 Jan-Feb;54:107831 . For example, the term ‘Charge variation’ in therapeutic proteins, including mAbs, refers to a phenomenon in which the acidic or basic species of the protein appear more frequently than the main protein isoform (Du et al., 2012). Protein charge depends on the number and type of ionizable amino acids that make up the protein. In particular, “charge variation” used interchangeably with the term “charge variant” is the consequence of post-translational modification of a protein. In particular, the charge variants refer to different isoforms of a protein that have slight variations in their charge due to modifications, namely post-translational modifications. The term “post-translational modification" is defined above. For example, phosphorylation refers to addition of phosphate groups (usually from ATP) to serine, threonine, or tyrosine residues introduces negative charges. This can alter the overall charge of the protein, potentially leading to a more negative isoelectric point (pl). Phosphorylation can also create binding sites for other proteins, affecting signaling pathways. Glycosylation refers to addition of carbohydrate moieties (N-linked or O-linked) which can affect the charge depending on the sugar composition. Glycosylation can introduce neutral or negatively charged groups, which can influence the protein's overall charge and solubility. Acetylation refers to addition of acetyl groups, particularly at lysine residues, which neutralizes the positive charge of the amino group. This can lead to a decrease in the overall positive charge of the protein, affecting its interactions and stability. Methylation does not change the charge of the protein, as it adds non-polar methyl groups to lysine or arginine residues. However, it can influence protein interactions and stability. In ubiquitination, the attachment of ubiquitin (a small protein) can alter the protein's charge due to the presence of additional lysine residues in ubiquitin. This modification is often involved in targeting proteins for degradation, impacting their functional charge state. Similarly, sumoylation, involves the addition of small ubiquitin-like modifier (SUMO) proteins. This can also alter the charge and influence the protein's localization and activity. In oxidation, the oxidation of specific amino acids (like cysteine) can introduce new functional groups, potentially altering the charge and reactivity of the protein and in deamidation, the conversion of asparagine or glutamine residues to aspartic acid or glutamic acid can introduce negative charges, affecting the protein's overall charge. The cumulative effect of these modifications can lead to different charge variants of the same protein, which can be separated and analyzed using techniques such as isoelectric focusing or chromatography. This diversity in charge variants can have significant implications for protein function, stability, and interactions in biological systems.
[0028] The term ‘Protein aggregation’, an important quality attribute, is commonly observed during the development of therapeutic proteins (V' azquez-Rey and Lang, 2011). Protein aggregation is considered undesirable because of its detrimental effect on drug efficacy and safety. Aggregation of mAbs, which is strongly related to the balance between expression of heavy and light chains of immunoglobulins, is also an important issue in the mAb manufacturing process (van der Kant et al., 2017). More in particular, the protein quality metric is glycosylation. Protein glycosylation is a critical quality attribute that modulates the efficacy, stability, and half-life of a therapeutic protein. Conversely, a protein with no or very few protein quality metrics may be considered an unknown or poor-quality protein.
[0029] The method according to any aspect of the present invention may be used to predict the quality of the heterologous protein produced by a test cell by determining the DNA methylation profile of the test cell, preferably the DNA methylation profile of the pre-selected CpG sites. These pre-selected CpG sites have been established based on reference CHO cell lines that are known to produce good or known quality of heterologous proteins. In particular, the method according to any aspect of the present invention is used to determine if the heterologous proteins produced by the test cell have the desired post translational modification(s) or protein quality metric(s) that are found in the reference protein that is produced by the reference CHO cell. In particular, the reference protein produced by the reference CHO cell is analogous to the native protein. The method according to any aspect of the present invention is therefore advantageous as firstly, the quality of the heterologous protein produced by the CHO cell can be determined before the cell goes into large scale production of the protein. This saves time and money. Further, the method according to any aspect of the present invention ensures that the proteins produced are analogous to the reference or desired protein. Thus improving yield of the desired product.
[0030] In one example, the pre-selected CpG sites may be selected from the CpG sites listed in Tables 4 and 5. In particular, according to any aspect of the present invention, the protein quality metric is glycosylation and the pre-selected CpG sites may be selected from the CpG sites listed in Table 4. In another example, the protein quality metric is charge variation of the protein produced and the pre-selected CpG sites may be selected from the CpG sites listed in Table 5.
[0031] In a further example, the pre-selected CpG sites may be selected from the CpG sites listed in Table 3 below.
[0032] Table 3. Overlapping CpG sites between Tables 1 and 2.
[0033] In yet a further example, the pre-selected CpG sites may be selected from the CpG sites listed in Table 6 below. Table 6. Overlapping CpG sites between Tables 4 and 5.
[0034] Protein quality can be determined using Immunoprecipitation based techniques, Biochemical Assays, Mass spectrometry (MS) and the like.
[0035] The terms “methylation profile”, “methylation pattern”, “methylation state” or “methylation status,” are used herein to describe the state, situation or condition of methylation of a genomic sequence, and such terms refer to the characteristics of a DNA segment at a particular genomic locus in relation to methylation. Such characteristics include, but are not limited to, whether any of the cytosine (C) residues within this DNA sequence are methylated, location of methylated C residue(s), percentage of methylated C at any particular stretch of residues, and allelic differences in methylation due to, e.g., difference in the origin of the alleles.
[0036] The term "methylation status" refers to the status of a specific methylation site (i.e. methylated vs. nonmethylated) which means a residue or methylation site is methylated or not methylated. Then, based on the methylation status of one or more methylation sites, a methylation profile may be determined. Accordingly, the term "methylation profile" or also “methylation pattern” refers to the relative or absolute concentration of methylated C residues or unmethylated C residues at any particular stretch of residues in the genomic material of a biological sample. For example, if cytosine (C) residue(s) not typically methylated within a DNA sequence are methylated, it may be referred to as "hypermethylated"; whereas if cytosine (C) residue(s) typically methylated within a DNA sequence are not methylated, it may be referred to as "hypomethylated". Likewise, if the cytosine (C) residue(s) within a DNA sequence (e.g., the DNA from a sample nucleic acid from a test subject) are methylated as compared to another sequence from a different region or from a different individual (e.g., relative to normal nucleic acid or to the standard nucleic acid of the reference sequence), that sequence is considered hypermethylated compared to the other sequence. Alternatively, if the cytosine (C) residue(s) within a DNA sequence are not methylated as compared to another sequence from a different region or from a different individual, that sequence is considered hypomethylated compared to the other sequence. These sequences are said to be "differentially methylated". Measurement of the levels of differential methylation may be done by a variety of ways known to those skilled in the art. One method is to measure the methylation level of individual interrogated CpG sites determined by the bisulfite sequencing method, as a non-limiting example.
[0037] As used herein, a “methylated nucleotide” or a “methylated nucleotide base” refers to the presence of a methyl moiety on a nucleotide base, where the methyl moiety is usually not present in a recognized typical nucleotide base. For example, cytosine in its usual form does not contain a methyl moiety on its pyrimidine ring, but 5-methylcytosine contains a methyl moiety at position 5 of its pyrimidine ring. Therefore, cytosine in its usual form may not be considered a methylated nucleotide and 5- methylcytosine may be considered a methylated nucleotide. In another example, thymine may contain a methyl moiety at position 5 of its pyrimidine ring, however, for purposes herein, thymine may not be considered a methylated nucleotide when present in DNA. Typical nucleotide bases for DNA are thymine, adenine, cytosine and guanine. Typical bases for RNA are uracil, adenine, cytosine and guanine. Correspondingly a "methylation site" is the location in the target gene nucleic acid region where methylation has the possibility of occurring. For example, a location containing CpG is a methylation site wherein the cytosine may or may not be methylated. In particular, the term “methylated nucleotide” refers to nucleotides that carry a methyl group attached to a position of a nucleotide that is accessible for methylation. These methylated nucleotides are usually found in nature and to date, methylated cytosine that occurs mostly in the context of the dinucleotide CpG, but also in the context of CpNpG- and CpNpN- sequences may be considered the most common. In principle, other naturally occurring nucleotides may also be methylated but they will not be taken into consideration with regard to any aspect of the present invention.
[0038] As used herein, the term “significantly similar” refers to in particular in context with the comparison of methylation profiles (such as the comparison between test profiles (from test subject(s) and reference profiles) a similarity observed by statistical means (i.e. by using bioinformatics) and / or also by observation using the eye. A significant similarity is observed for example if a test profile overlaps with a reference profile that is defined by multiple training samples through multivariate statistical methods, such as Principal Component analysis or Multi-Dimensional Scaling. In particular, a test profile is significantly similar to the pre-determined reference profile if more than 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 % of the methylation pattern / profile overlaps with that of the reference profile. A similarity of a test profile to more than one, such as two, three or even all reference profiles reduces the significance of the similarity.
[0039] As used herein, the term “significantly different” refers to in particular in context with the comparison of methylation profiles (such as the comparison between test profiles (from test subject(s) and reference profiles) a difference observed by statistical means (i.e. by using bioinformatics) and / or also by observation using the eye. A significant difference is observed for example if a test profile does not overlap with a reference profile that is defined by multiple training samples through multivariate statistical methods, such as Principal Component analysis or Multi-Dimensional Scaling. In particular, a test profile is significantly different to the pre-determined reference profile if less than 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 % of the methylation pattern / profile overlaps with that of the reference profile.
[0040] As used herein, the term “genomic material” refers to nucleic acid molecules or fragments of the genome of the CHO cells or cell lines. In particular, such nucleic acid molecules or fragments are DNA or RNA or hybrids thereof, and most preferably are molecules of the DNA genome of CHO cells or cell lines.
[0041] As used herein, the “DNA sample” or ‘DNA in step (a)’ refers to the DNA extracted from the cell according to any aspect of the present invention using known methods in the art.
[0042] ‘Bisulfite treatment’ of genomic DNA used interchangeably with the term ‘bisulfite modification’, refers to the treatment of the genomic DNA with a deaminating agent such as a bisulfite that may be used to treat all DNA, methylated or not. In particular, the term “bisulfite” as used herein encompasses any suitable type of bisulfite, such as sodium bisulfite, or other chemical agents that are capable of chemically converting a cytosine (C) to an uracil (U) without chemically modifying a methylated cytosine and therefore can be used to differentially modify a DNA sequence based on the methylation status of the DNA, e.g., U.S. Pat. Pub. US 2010 / 0112595. As used herein, a reagent that "differentially modifies" methylated or non-methylated DNA encompasses any reagent that modifies methylated and / or unmethylated DNA in a process through which distinguishable products result from methylated and nonmethylated DNA, thereby allowing the identification of the DNA methylation status. Such processes may include, but are not limited to, chemical reactions (such as a C to U conversion by bisulfite) and enzymatic treatment (such as cleavage by a methylation-dependent endonuclease). Thus, an enzyme that preferentially cleaves or digests methylated DNA is one capable of cleaving or digesting a DNA molecule at a much higher efficiency when the DNA is methylated, whereas an enzyme that preferentially cleaves or digests unmethylated DNA exhibits a significantly higher efficiency when the DNA is not methylated.
[0043] Accordingly, before step (a) according to any aspect of the present invention is carried out, the genomic DNA contained / obtained or extracted from the cell, is first bisulfite treated.
[0044] An alternative method available in the art may be used instead of bisulfite treatment. A skilled person will understand which other methods to use. In one example, TET-assisted pyridine borane sequencing (TAPS) may be used for detection of 5mC and 5hmC (Yibin Liu, et al., Nature Biotechnology, 37: 424- 429 (2019).
[0045] The term “test” used in conjunction with the term CHO cell herein refers to a CHO cell that is subjected to the method according to any aspect of the present invention and is the basis for an analysis application of the present invention. A ‘test cell’ is therefore a CHO cell or a group of CHO cells being tested according to any aspect of the present invention, or a profile being obtained or generated in this context. Conversely, the term “reference” or ‘control’ shall denote, mostly predetermined, entities which are used for a comparison with the test entity. In particular, a ‘test cell’ refers to a cell being tested for capable of production of known or good quality heterologous protein production where the methylation status has to be determined and a ‘control’ or ‘reference’ refers to a cell which is known to display known or good quality protein production or a methylation profile thereof. The term ‘known or good quality protein production' is further explained in Ha TK, et al. Biotechnol Adv. 2022 Jan-Feb;54:107831 and Sha S, et al. Trends Biotechnol. 2016; 34(10):835-846.
[0046] As used herein, a “CpG site” or “methylation site” is a nucleotide within a nucleic acid (DNA or RNA) that is susceptible to methylation either by natural occurring events in vivo or by an event instituted to chemically methylate the nucleotide in vitro. Some of these sites may be hypermethylated and some may be hypomethylated in a cell. In some cases a CpG site may not be considered fully hypermethylated or hypomethylated but a value may be given that is a measure of methylation of the CpG site. Accordingly, methylation may be quantified and may not always be an absolute case of hypermethylation or hypomethylation.
[0047] As used herein, a “methylated nucleic acid molecule” refers to a nucleic acid molecule that contains one or more nucleotides that is / are methylated.
[0048] A “CpG island” as used herein describes a segment of DNA sequence that comprises a functionally or structurally deviated CpG density. For example, Yamada et al. have described a set of standards for determining a CpG island: it must be at least 400 nucleotides in length, has a greater than 50% GC content, and an OCF / ECF ratio greater than 0.6 (Yamada et al., 2004, Genome Research, 14, 247-266). Others have defined a CpG island less stringently as a sequence at least 200 nucleotides in length, having a greater than 50% GC content, and an OCF / ECF ratio greater than 0.6 (Takai et al., 2002, Proc. Natl. Acad. Sci. USA, 99, 3740-3745).
[0049] In particular, when there is differential methylation detected in a test cell, that is to say that the cell displays absolute hypermethylation or hypomethylation or at least quantitative differential methylation at, at least one CpG site in comparison to the reference (i.e., from a CHO cell line producing known and / or good quality heterologous proteins), then the test cell is also capable of or produces known and / or good quality heterologous proteins. More in particular, when the CpG site displays the same methylation status in the test cell in comparison to the corresponding CpG site in the reference cell or reference methylation profile, the test cell produces known and / or good quality heterologous proteins (i.e. proteins with at least one protein quality metric or the same protein quality metric as the reference cell). Overall, this platform gives us an opportunity to detect wide-spread DNA methylation status in CHO cells and correlate it with industrially relevant parameters which are crucial for the development of at least biological pharmaceutical products.
[0050] In particular, in the method according to any aspect of the present invention, in step (a) the methylation status of at least 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100 CpG sites are determined. A skilled person would be capable of determining the number of CpG sites that need to be used in step (a) according to any aspect of the present invention. Even more in particular, the skilled person would be able to determine the number of CpG sites that make up the ‘pre-selected methylation sites’ in step (a) to obtain the test methylation profile.
[0051] As used herein, the term “pre-selected methylation sites” refers to methylation sites that were selected from genes or regions that showed the highest degree of methylation variation during the training of the method and fulfils certain quality criteria such as a minimum sequencing coverage of >5x were considered and for >5 qualified CpG sites. Additionally, genes that have an average methylation level <0.1 or an average methylation level >0.9 can be excluded due to their limited dynamic range. “Reference methylation profiles” may be defined on the basis of multiple training samples using multivariate statistical methods, such as such as Principal Component analysis or Multi-Dimensional Scaling.
[0052] In particular, the reference methylation profile according to any aspect of the present invention is a compilation of more than one CpG site from at least one CHO reference cell line that produces heterologous proteins of a known, preferably good quality. In one example, the different CpG sites are collected from a single reference CHO cell line that produces heterologous proteins of a known, preferably good quality. In another example, the different CpG sites are collected from more than one cell line where each cell line produces heterologous proteins of a known, preferably good quality. The reference methylation profile according to any aspect of the present invention may thus not be a naturally occurring methylation profile from a single CHO cell line but an artificial profile obtained from combining relevant CpG sites from different reference CHO cell lines, each produces heterologous proteins of a known, preferably good quality. Such a reference profile may be especially accurate and / or specific at identifying any CHO test cell that produces heterologous proteins of good or known quality. The method according to this aspect of the present invention attempts to create a methylation profile for a CHO cell line that has the potential for predicting CHO cells that are capable of producing heterologous proteins of a known or good quality. In particular, the method according to this aspect of the present invention provides a prognostic methylation profile for ideal parental cell lines prior to and post transgene introduction.
[0053] Low Methylated Region (LMR) is a region of the genome wherein less than 60% of CpGs in that region are methylated. More in particular, less than 50%, 40%, 30%, 20% or 10% of the CpGs in the LMRs are methylated. Any method known in the art may be used to identify or detect LMRs in the genomic DNA. Well known methods include using programmes such as MethylSeekR. In particular, LMRs in the genomic DNA have at least three consecutive CpGs and have no single nucleotide polymorphisms (SNPs) in any of the CpG positions. Even more in particular, LMRs in the genomic DNA are identified based on the method disclosed at least in Burger, L., (2013) Nucleic Acids Research, 41 (16): e155 and / or Stadler, M., (2011) Nature 480, 490-495. LMRs are known to have an average methylation ranging from 10% to 50%; are regions of low CG density which do not overlap with CpG islands; tend to be enriched for H3K4me1 , DHSs, and p300 / CBP; and / or are primarily located distal to promoters in intergenic or intronic regions. In particular, LMRs:
[0054] - have an average methylation ranging from 10% to 50%,
[0055] - are regions of low CG density;
[0056] - are enriched for Histone H3 monomethylated at lysine 4 (H3K4me1), DNase I hypersensitive sites (DHSs) and transcriptional coactivators CREB binding protein (CPB) and p300;
[0057] - are primarily located distal to promoters in intergenic or intronic regions; and / or
[0058] - have no single nucleotide polymorphisms (SNPs) in any of the CpG positions.
[0059] Low-methylated regions (LMRs) represent a key feature of the dynamic methylome. LMRs are local reductions in the DNA methylation landscape and represent CpG-poor distal regulatory regions that often reflect the binding of transcription factors and other DNA-binding proteins. LMRs were originally described in the mouse (Stadler et al. (2011) Nature: 480, 490-95). Evolutionary conservation of LMRs beyond mammals has remained unexplored.
[0060] Differentially methylated regions (DMRs) are genomic regions with different methylation statuses among multiple biological samples like tissues, cells, individuals, etc. These are genomic regions that differ between phenotypes. The statistical power is likely to be greater when adjacent DMPs are considered together as a whole [Gu H et al (2010) Nat Methods 2010; 7:133-6]. The lengths of the DMRs may range between a few hundred to a few thousand bases [Rakyan et al (2011) Nat Rev Genet 12:529-41 , 2011 , Bock C (2012) Nat Rev Genet 2012; 13:705-19],
[0061] DMRs may occur throughout the genome but have been identified particularly around the promoter regions of genes, within the body of genes, and at intergenic regulatory regions. There are two types of regions, predefined or user defined. Regions with special biological meaning, such as CpG islands, CpG shores, UTRs and so on, are predefined. Many traditional statistical testings, including t-test and Wilcoxon rank sum test, can be performed at a region level. For user-defined regions, criteria such as a fixed region length, fixed numbers of significant and adjacent CpG sites, significant and smoothed estimated effect sizes, etc.
[0062] Partially methylated domains (PMDs) are extended regions in the genome exhibiting a reduced average DNA methylation level. They cover gene-poor and transcriptionally inactive regions and tend to be heterochromatic.
[0063] Differentially methylated Positions (DMP) are CpG sites with different DNA methylation status across different biological samples and regarded as possible functional regions involved in gene transcriptional regulation.
[0064] As used herein, the term ‘parental clone’ refers to a cell line derived from a CHO cell line in which a transgene has been integrated into the genome. The term ‘subclone’ as used herein in relation to a parental clone refers to a clonal cell line derived from parental clone having the same genotype but a different phenotype due to epigenetic changes.
[0065] The DNA methylation profile of step (a) according to any aspect of the present invention is determined using DNA methylation-based array. In particular, a bead-based DNA methylation array. The array according to any aspect of the present invention is advantageous as it enables the understanding of genome stability of the CHO cell line, enables better control over the manufacturing / process development / product development / scaling up / validation process, thereby aiding in the selection of better CHO cell lines for industrial applications.
[0066] DNA-Methylation-based arrays allow for a high-throughput and robust method to determine semi- quantitative / quantitative DNA-methylation information through a small sample of extracted DNA of interest. These custom designed arrays may use Illumina iScan and Infinium platform technology or an equivalent thereof, which allows on each chip for example 100,000 different bead types that covalently bind DNA-methylation probes. Each probe represents one CpG Methylation site at the end of the probe sequence. DNA samples undergo bisulfite conversion, amplification, fragmentation, precipitation and resuspension steps before hybridization on an array chip. Once on the chip the DNA hybridizes to the beads for each CpG site so that methylation changes at each site can be detected specifically through single nucleotide extension. This is especially advantageous as the array-based method is simple and the results of the methylation-based array are accurate and reproducible.
[0067] Further, compared to traditional sequencing which can take weeks to generate data, the array technology has a much shorter turn-around time. The volume and complexity of data generated is lesser compared to sequencing making it computationally less intensive. This allows for quicker computation to achieve interpretable results from experimental groups. Overall microarray technology is roughly 10x faster and 10x cheaper than traditional sequencing while still quantifiable for the methylation level at specific CpG sites.
[0068] The term “array” as used herein refers to an intentionally created collection of probe molecules which can be prepared either synthetically or biosynthetically. The probe molecules in the array can be identical or different from each other. The array can assume a variety of formats, for example, libraries of soluble molecules; libraries of compounds tethered to resin beads, silica chips, or other solid supports.
[0069] In particular, a DNA methylation-based array provides a convenient platform for simultaneous analysis of large numbers of CpG sites, for example, at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 50, 100, 500, 1000, 5000, 10,000, 100,000 or more sites or loci. In particular, the array comprises a plurality of different probe molecules that can be attached to a substrate or otherwise spatially distinguished in an array. Examples of arrays that may be used according to any aspect of the present invention include slide arrays, silicon wafer arrays, liquid arrays, bead-based arrays and the like. In one example, array technology used according to any aspect of the present invention combines a miniaturized array platform, a high level of assay multiplexing, and scalable automation for sample handling and data processing.
[0070] In particular, the array according to any aspect of the present invention may be an array of arrays, also referred to as a composite array, having a plurality of individual arrays that is configured to allow processing of multiple samples simultaneously. Examples of composite arrays and the technology behind them are disclosed at least in US 6,429,027 and US 2002 / 0102578. A substrate of a composite array may include a plurality of individual array locations, each having a plurality of probes, and each physically separated from other assay locations on the same substrate such that a fluid contacting one array location is prevented from contacting another array location. Each array location can have a plurality of different probe molecules that are directly attached to the substrate or that are attached to the substrate via rigid particles in wells (also referred to herein as beads in wells).
[0071] In one example, an array substrate can be a fibre optical bundle or array of bundles as described in US6,023,540, US6,200,737 and / or US6,327,410. An optical fibre bundle or array of bundles can have probes attached directly to the fibres or via beads. A skilled person would be able to easily determine which substrate will be most suitable for the array according to any aspect of the present invention. W020041 10246 further discloses other substrates and methods of attaching beads to the substrates that may be used in the array according to any aspect of the present invention.
[0072] In one example, a surface of the substrate may have physical alterations to enable the attachment of probes or produce array locations. For example, the surface of a substrate can be modified to contain chemically modified sites that are useful for attaching, either-covalently or non-covalently, probe molecules or particles having attached probe molecules. Probes may be attached using any of a variety of methods known in the art including, an ink-jet printing method, a spotting technique, a photolithographic synthesis method, or printing method utilizing a mask. W02004110246 discloses these techniques in more detail. In one example, the DNA methylation-based array according to any aspect of the present invention may be a bead-based array, where the beads are associated with a solid support such as those commercially available from Illumina, Inc. (San Diego, Calif.). An array of beads useful according to any aspect of the present invention can also be in a fluid format such as a fluid stream of a flow cytometer or similar device. Commercially available fluid formats for distinguishing beads include, for example, those used in XMAP(TM) technologies from Luminex or MPSS(TM) methods from Lynx Therapeutics.
[0073] The term “solid support”, “support”, and “substrate” as used herein are used interchangeably and refer to a material or group of materials having a rigid or semi-rigid surface or surfaces. In many examples, at least one surface of the solid support will be substantially flat, although in some examples it may be desirable to physically separate synthesis regions for different compounds with, for example, wells, raised regions, pins, etched trenches, or the like.
[0074] The DNA methylation array according to any aspect of the present invention may be a very high-density array, for example, those having from about 10,000,000 probes / cm2to about 2,000,000,000 probes / cm2or from about 100,000,000 probes / cm2to about 1 ,000,000,000 probes / cm2. High density arrays are especially useful according to any aspect of the present invention for including the multitude of CpG sites on the array.
[0075] The DNA methylation array according to any aspect of the present invention may be used to analyse or evaluate such pluralities of loci simultaneously or sequentially as desired. In one example, a plurality of different probe molecules can be attached to a substrate or otherwise spatially distinguished in an array. Each probe is typically specific for a particular locus and can be used to distinguish methylation state of the locus.
[0076] The term “probe molecules” as used herein refers to a surface-immobilized molecule that can be recognized by a particular target. Probes used in the array can be specific for the methylated allele of a CpG site, the non-methylated allele of the CpG site or both.
[0077] The term “target” as used herein refers to a molecule that has an affinity for a given probe molecule. Targets may be naturally occurring or man-made molecules. Also, they can be employed in their unaltered state or as aggregates. Targets may be attached, covalently or noncovalently, to a binding member, either directly or via a specific binding substance. Examples of targets which can be employed according to any aspect of the present invention are methylated and non-methylated CpG sites. Targets are sometimes referred to in the art as anti-probes. As the term targets is used herein, no difference in meaning is intended.
[0078] The term “complementary” as used herein refers to the hybridization or base pairing between nucleotides or nucleic acids, such as, for instance, between the two strands of a double stranded DNA molecule or between an oligonucleotide primer and a primer binding site on a single stranded nucleic acid to be sequenced or amplified. Complementary nucleotides are, generally, A and T (or A and U), or C and G. Two single stranded RNA or DNA molecules are said to be complementary when the nucleotides of one strand, optimally aligned and compared and with appropriate nucleotide insertions or deletions, pair with at least about 80% of the nucleotides of the other strand, usually at least about 90% to 95%, and more preferably from about 98 to 100%. Perfectly complementary refers to 100% complementarity over the length of a sequence. For example, a 25-base probe is perfectly complementary to a target when all 25 bases of the probe are complementary to a contiguous 25 base sequence of the target with no mismatches between the probe and the target over the length of the probe.
[0079] According to yet another aspect of the present invention, there is provided a use of a DNA-methylation based array for selection of a CHO clone that is capable of producing a heterologous protein with at least one desired protein quality metric.
[0080] The protein quality metric selected from the group consisting of glycosylation, propensity to form aggregates, charge variant, and amidation to the heterologous protein, preferably the protein quality metric is glycosylation of the heterologous protein.
[0081] According to a further aspect of the present invention, there is provided a bead based DNA methylation array comprising at least: a plurality of distinct locations, each location having at least one probe molecule comprising a nucleic acid sequence complementary to a plurality of CpG sites of a CHO cell, wherein the bead-based DNA methylation array is used in the method according to any aspect of the present invention.
[0082] ‘Environmental specific CpG sites’ also known as dynamic CpG sites in the context of CHO cells refer to the CpG sites that are differentially methylated among different CHO cell lines. The cell lines that were used in this analysis include CHO-K1 (ATCC), CHO-DG44 (Thermo Fisher Scientific), CHO-DXB11 (ATCC), ExpiCHO-S™ cells (Thermo Fisher Scientific), Freestyle™ CHO-S™ cells (Thermo Fisher Scientific), CHO 1-15soo (ATCC) and Agarabi CHO (ATCC).
[0083] ‘Metabolic linked genes’ in the context of CHO cells herein refer to genes that are related to several metabolism pathways such as Glycolysis, TCA cycle, Pentose Phosphate pathway, Malate-aspartate shuttle, Amino acid metabolism, Lactate metabolism, Cholesterol biosynthesis, Nucleotide biosynthesis, Nucleotide sugar biosynthesis etc. A few examples of such genes include Hk2, Pgk1 , Idh3a, Pgm1 , and Pdhal . A skilled person would easily determine the genes that are found in CHO cells that fall within this category.
[0084] ‘Protein production linked genes’ used in the context of CHO cells herein refer to genes that are related to cellular processes such as DNA replication and repair, mRNA transcription, mRNA translation, post- translational modifications, and protein folding and export. A few examples of such genes include Gatb, Sec61a2, Ube2e3, Exosd , Dna2, Poldi and the like. A skilled person would easily be able to determine the other genes that are found in CHO cells that fall within this category.
[0085] ‘Cell growth and division linked genes’ used in the context of CHO cells herein refer to genes that are related to cellular processes such as cell cycle regulation, Cytoskeleton-related elements, cell signalling, nucleotide metabolism, and cell death. A few examples of such genes include Camkl , Cd82, Cdk4, Col1a1 , and Ctsb. Again, a skilled person would easily be able to determine the other genes that are found in CHO cells that fall within this category.
[0086] ‘Epigenetic linked genes’ used in the context of CHO cells herein refer to genes that are related to epigenetic modifications such as DNA methylation pathway, DNA demethylation pathway, Folate and Methionine cycle, and Histone modifications. A few examples of such genes include Hat1 , Shmtl , Bhmt, Dnmtl , and Ehmtl . A skilled person would easily be able to determine the other genes that are found in CHO cells that fall within this category.
[0087] The examples adduced hereinafter describe the present invention by way of example, without any intention that the invention, the scope of application of which is apparent from the entirety of the description and the claims, be restricted to the embodiments specified in the examples.
[0088] EXAMPLES:
[0089] Example 1 :
[0090] Predicting the glycosylation of heterologous protein from CHO cells
[0091] Wet-Lab methodology
[0092] For this experiment, 18 transgenic CHO clones (from A*Star BTI) expressing biosimilar were grown in EX-CELL Advance CHO Fed-Batch media + 6mM Glutamine + 250mM MTX at 37°C, 8% CO2, at a shaking speed of 225 RPM. The clones were subjected to 14days of fed-batch culture where the flasks were seeded with 3E5 viable cells / mL on day 0 and the culture was fed with appropriate feed on Day 3, 5, 7, 9, 11 and glucose was topped up to 6g / l using 45% glucose when it dropped below 2g / l. The fed batch culture of 18 clones was maintained for 14 days. Cell count & cell viability were measured using Vi-CELL BLU cell counter (Beckman Coulter), and specific productivity were measured Cedex Bio Analyzer (Roche) for day 14 and cell pellets were collected on day 9. Culture media was harvested on day 14 and antibody was purified using protein A SpinTrap (Cytiva) which was followed by HILIC-FLR-MS analysis to identify the glycans and determine the relative abundance of each N-glycan. Based on the combined glycosylation result (fucosylation, galactosylation, mannosylation and sialyation), the 18 transgenic CHO clones were categorized into three groups as per their closeness (measured by Euclidean distance) to the innovator protein: Close (closeness <14), Medium (closeness 14-18) and Distant (closeness >18).
[0093] DNA Extraction DNA was extracted using the PureLink Genomic DNA Isolation Minikit kit (Invitrogen), including RNAase treatment following the manufacturer's instructions. DNA quantity was measured by PicoGreen assay and DNA quality was assessed via NanoDrop (Thermo Scientific) to ensure the A260 / 280 ratio was < 1 .8. A small amount of sample was then also analysed using automated electrophoresis on TapeStation (Agilent) to ensure each sample contained high molecular weight.
[0094] Bisulfite Conversion and BeadChip Analysis
[0095] The genomic DNA samples were then subjected to bisulfite conversion using the EZ DNA Methylation- Gold™ Kit (Zymo Research). The methylation levels were then quantified using our customized methylation BeadChip kits (Illumina) which can analyze over 30,000 methylation sites quantitatively across the genome at single-nucleotide resolution.
[0096] Data processing:
[0097] The customized chip array data processing was performed in R version 4.1 .2 using sesame version 1 .14.2. DNA methylation level for each site was calculated as methylation p-value. Beta values were defined as methylated signal / (methylated signal + unmethylated signal). It can be computed using getBetas function. The SeSAMe pipeline (Zhou et al. 2018) was used to generate normalized p-values and for quality control. Low intensity- based detection calling and making (based on p-value) was done with pOOBAH. Background subtraction based on normal-exponential deconvolution using out-of-band probes noob (Triche et al. 2013) and optionally with extra bleed- through subtraction were also implemented.
[0098] The Differential methylation analysis on the sample groups was also performed using Sesame.
[0099] Methods to obtain PCA plots
[0100] After obtaining the beta values, control and SV40 probes were filtered out of the data frame. CpG sites with NA beta values were also removed from the data frame. This resulted in 16447 CpG sites remaining and the Principal Component Analysis (PCA) plot for these sites were plotted using the prcomp followed by ggplot2 functions. The resulting PCA plot is shown in Figure 1 .
[0101] Differentially Methylated Positions (DMPs) between the three groups: Close (closeness <14), Medium (closeness 14-18) and Distant (closeness >18) extracted using the dml function from the sesame package. The dml function will result in a data frame and to obtain the more statistically significant DMPs, only DMPs with P-Value < 0.05 and Effect Size > 0.1 were retained while the rest were removed from the data frame. This resulted in 238 CpG sites remaining as shown in Table 1 and the PCA plot for these sites were plotted using the prcomp followed by ggplot2 functions. The resulting PCA plot is shown in Figure 2.
[0102] Table 1 : List of DMPs between different glycosylation groups
[0103] Table 4. List of DMPs between different glycosylation groups that exclude DMPs listed in W02024 / 046840
[0104] Example 2:
[0105] Predicting the Charge variant of heterologous protein from CHO cells
[0106] Wet-Lab methodology
[0107] For this experiment, 18 transgenic CHO clones (from A*Star BTI) expressing biosimilar were grown in EX-CELL Advance CHO Fed-Batch media + 6mM Glutamine + 250mM MTX at 37°C, 8% CO2, at a shaking speed of 225 RPM. The clones were subjected to 14days of fed-batch culture where the flasks are seeded with 3E5 viable cells / mL on day 0 and the culture was fed with appropriate feed on Day 3, 5, 7, 9, 11 and glucose is topped up to 6g / l using 45% glucose when it drops below 2g / l. The fed-batch culture of 18 clones is maintained for 14 days. Cell count & cell viability were measured using Vi-CELL BLU cell counter (Beckman Coulter), and specific productivity was measured using Cedex Bio Analyzer (Roche) for day 14 and cell pellets were collected on day 9. Culture media was harvested on day 14 and antibody was purified using protein A SpinTrap (Cytiva) which was followed by HPLC-CEX analysis on the purified samples to determine charge variants (acidic, basic and main species). Based on the percentage of the main species ofthe purified heterologous protein, the 18 transgenic CHO clones were categorized into two groups as per their difference from the innovator protein: Close (difference < 5.8 ), and Distant (difference > 5.8).
[0108] DNA Extraction, Bisulfite Conversion and BeadChip Analysis and Data processing were carried out as provided in Example 1 .
[0109] Methods to obtain PCA plots
[0110] After obtaining the beta values, control and SV40 probes were filtered out of the data frame. CpG sites with NA beta values were also removed from the data frame. This resulted in 16447 CpG sites remaining and the Principal Component Analysis (PCA) plot for these sites were plotted using the prcomp followed by ggplot2 functions. The resulting PCA plot is shown in Figure 3.
[0111] Differentially Methylated Positions (DMPs) between charge variant groups: Close (difference < 5.8 ), and Distant (difference > 5.8) were extracted using the dml function from the sesame package. The dml function will result in a data frame and to obtain the more statistically significant DMPs, only DMPs with P- Value < 0.05 and Effect Size > 0.1 were retained while the rest were removed from the data frame. This resulted in 191 CpG sites remaining as shown in Table 2 and the PCA plot for these sites were plotted using the prcomp followed by ggplot2 functions. The resulting PCA plot is shown in Figure 4.
[0112]
[0113] Table 2: List of DMPs between Charge variant groups
[0114]
[0115] Table 5. List of DMPs between different charge variant groups that exclude DMPs listed in W02024 / 046840
Claims
CLAIMS1 . A method of predicting quality of a heterologous protein to be produced from a population of test Chinese Hamster Ovary (CHO) cells, the method comprising the steps of:(a) determining a test methylation profile of pre-selected methylation CpG sites within the DNA of the test CHO cells;(b) comparing the test methylation profile obtained from (a) with a reference methylation profile, wherein the reference methylation profile comprises the methylation status of the preselected methylation CpG sites from CHO reference cells that are capable of producing a known, preferably good, quality heterologous protein; wherein a significant similarity in the test methylation profile of (a) compared to the reference methylation profile, is indicative of the heterologous protein produced by the test cells being of known, preferably good, quality; and wherein a significant difference in the test methylation profile of (a) compared to the reference methylation profile, is indicative of the heterologous protein produced by the test cells being of unknown or poor quality; wherein the test and reference methylation profiles are determined using bead-based DNA methylation array; and wherein the quality of the heterologous protein is based on the presence of at least one protein quality metric; and the pre-selected methylation CpG sites are selected from Tables 4 and 5.
2. The method according to claim 1 , wherein the protein quality metric is selected from the group consisting of glycosylation, propensity to form aggregates, charge variation, and amidation of the heterologous protein.
3. The method according to either claim 1 or 2, wherein the protein quality metric is glycosylation.
4. The method according to any one of the preceding claims, wherein the protein quality metric is glycosylation and the pre-selected methylation CpG sites are selected from Table 4.
5. The method according to any one of the preceding claims, wherein the protein quality metric is charge variation and the pre-selected methylation CpG sites are selected from Table 5.
6. The method according to any one of the preceding claims, wherein the pre-selected methylation CpG sites from the test and reference methylation profiles are from Table 3.
7. The method according to any one of the preceding claims, wherein the reference methylation profile is a compilation of more than one CpG site from more than one CHO reference cell line that produce a known, preferably good, quality heterologous protein with at least one protein quality metric.
8. A method of selecting at least one CHO cell capable of producing at least one heterologous protein with a desired protein quality metric from a population of CHO cells from a parental clone, the method comprising the steps of:(a) determining a test methylation profile from genomic material obtained from the CHO cell, and(b) comparing the test methylation profile of (a) with a reference methylation profile, wherein the reference methylation profile comprises the methylation status of more than one CpG site from the CHO cell genome of at least one CHO reference cell line that produces a heterologous protein with the desired protein quality metric; wherein a significant similarity in the methylation profile of (a) compared to the reference methylation profile, is indicative of the CHO cell being capable of producing heterologous proteins with the desired protein quality metric; and wherein a significant difference in the test methylation profile of (a) compared to the reference methylation profile, is indicative of the CHO cell capable of producing heterologous proteins without the desired protein quality metric; wherein the methylation profile of (a) and the reference methylation profile are determined using bead-based DNA methylation array of pre-selected methylation CpG sites; and wherein the pre-selected methylation CpG sites are selected from Tables 4 and 5.
9. The method according to claim 8, wherein the protein quality metric is selected from the group consisting of glycosylation, propensity to form aggregates, charge variation, and amidation to the heterologous protein, preferably the protein quality metric is glycosylation.
10. The method according to any one of the preceding claims, wherein the reference methylation profile is a compilation of more than one CpG site from more than one CHO reference cell line that produces a heterologous protein with the desired known protein quality metric.
11. Use of a DNA-methylation based array for the method of selection of a CHO clone according to any one the claims 8 to 10..
12. Use according to claim 11 , wherein the protein quality metric is selected from the group consisting of glycosylation, propensity to form aggregates, charge variation, amidation and a chemical or biological modification to the heterologous protein, preferably the protein quality metric is glycosylation of the heterologous protein.
13. A bead-based DNA methylation array comprising at least: a plurality of distinct locations, each location having at least one probe molecule comprising a nucleic acid sequence complementary to a plurality of CpG sites of a CHO cell, wherein the bead-based DNA methylation array is used in the method according to any one of the claims1 to 10.
Citation Information
Patent Citations
Alternative substrates and formats for bead-based array of arrays TM
US20020102578A1
Bisulfite Conversion Reagent
US20100112595A1
Fiber optic sensor with encoded microspheres
US6023540A
Photodeposition method for fabricating a three-dimensional, patterned polymer microstructure
US6200737B1
Target analyte sensors utilizing Microspheres
US6327410B1