Predicting the fitness of a CHO cell

By measuring methylation levels of specific CpG sites in CHO cells and using statistical algorithms to assess phenotypes, the method addresses the inefficiencies in current CHO cell selection processes, providing a predictive tool for protein production and improving bioprocessing efficiency.

WO2025108794A1PCT designated stage expired Publication Date: 2025-05-30EVONIK OPERATIONS GMBH
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/082142
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-24
Filing Date
2024-11-13
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Current methods for determining the suitability of CHO cells for protein production are time-consuming, non-specific, and unable to guarantee consistency over time, leading to inefficiencies and financial losses in bioprocessing.

Method used

A method based on epigenetics that measures methylation levels of specific CpG sites in CHO cells, using a statistical algorithm to determine phenotypes of interest and assess cell fitness, thereby providing a predictive tool for protein production.

Benefits of technology

This method allows for accurate, simple, and reproducible identification of CHO cells suitable for growth and protein production, ensuring stability, efficiency, and reducing financial losses in bioprocessing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000018_0001
    Figure IMGF000018_0001
  • Figure IMGF000019_0001
    Figure IMGF000019_0001
  • Figure IMGF000019_0002
    Figure IMGF000019_0002
Patent Text Reader

Abstract

The present invention relates to a method of establishing a biomarker for a first phenotype of interest of a CHO cell, the method comprising the steps of: (a) determining methylation values of CpG sites within genomic DNA obtained from a population of CHO cells that are a representation of the first phenotype of interest and are part of the training samples, (b) defining a set of specific CpG sites having reliable methylation values in the training samples of step (a); and (c) performing a penalized regression using the methylation values of step (b) as input and phenotype of interest correlated to the training samples as dependent variable, by applying a penalized regression model; thereby obtaining the specific CpG sites with corresponding weighting factors and intercept of the linear model equation as parameters defining the biomarker for the first phenotype of interest.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] PREDICTING THE FITNESS OF A CHO CELL

[0002] FIELD OF THE INVENTION

[0003] The invention relates to a method based on epigenetics for quantitatively and qualitatively assessing the fitness of a CHO cell for use in bioprocessing. In particular, the method measures the methylation levels of specific CpG sites and using a specific statistical algorithm determines phenotypes of interest expressed by the CHO cell and based on the phenotypes of interest expressed by the cell, the cell fitness of the CHO cell is accessed.

[0004] BACKGROUND OF THE ART

[0005] Chinese Hamster Ovary (CHO) cells are known to be the workhorses for the industrial production of recombinant therapeutic proteins since 1987 and are hence widely used for biologies production. About 70% of all recombinant biopharmaceutical proteins and all monoclonal antibodies approved since 2016 are being manufactured in CHO cells. Several advantages of utilizing CHO for biologies production include tolerance to genetic manipulations, ease of adaptation to manufacturing process scales, rapid growth rates, and ability to perform human-compatible post-translational modifications. However, the biologies production system in CHO faces a bottleneck due to the loss of protein productivity over time.

[0006] Initial expression and / or production of proteins by a cell line is often high; however, the production reduces during prolonged culture. This results in decreased process yield, impacts timelines and increases costs. Changes in cell culture environment can result in an alteration of cell behaviour and protein productivity of the producer cell line.

[0007] The current methods of determining the suitability of a CHO clone for target protein production are not only time-consuming, but also not specific enough as readout for the interplay between the cell and its immediate environment, which is critical for selection of clones or cells for optimal protein production. For example, US2017 / 0081732 discloses a method for the selection of a long-term polypeptide expressing or secreting cell based on the determination of histone acylation which is measured using PCR. However, the method is complicated and not accurate. Further, PCR being a non-high throughput method does not allow for examination of genome wide CpGs, making the process incomplete.

[0008] Further, genetically identical CHO clones can also result in heterogenous phenotypes creating instability to an established process, inefficiency and financial loss during heterologous protein production at an industrial scale. Methods to compare and select CHO clones that use only phenotypic analyses are not able to guarantee consistency over time. Genotype comparisons of CHO clones cannot define how genes are expressed differentially to adapt to environmental conditions. As shown by Wippermann A, et al., Appl Microbiol Biotechnol. 2014 Jan;98(2):579-89, supplementation of butyrate which is known to enhance cell specific productivities in CHO cells also led to alterations of epigenetic silencing events. Accordingly, there is a need in the art for a tool that is efficient and affordable to globally evaluate and regulate CHO metabolism and protein production. There is also a need in the art for methods of selection and maintenance of identical CHO populations in order to improve speed, quality, efficiency and consistency of production. In particular, there is a need in the art for new descriptive and predictive markers for positive and valuable phenotypes in CHO cells that can be used to identify CHO cells early that are well suited for genetic modifications and / or heterologous protein production.

[0009] DESCRIPTION OF THE INVENTION

[0010] The present invention solves the problems above by providing an accurate, simple and reproducible means of identifying CHO cells that are fit for growth, and protein production at an early stage of bioprocessing. In particular, the present invention provides a means of ensuring, stability, efficiency and reduction of financial loss during heterologous protein production particularly at an industrial scale.

[0011] The inventors have identified a number of CpG (Cytosine-phosphate Guanine) sites in CHO cells’ genome for which the level of DNA methylation is correlated with phenotypes of interest expressed by the cells. That is, measuring DNA methylation at these locations (CpG sites) enables making accurate predictions of the expression of phenotypes of interest in any CHO cell tested. Since genotype comparisons of CHO clones cannot define how genes are expressed differentially to adapt to environmental conditions, and phenotypic analyses alone are not able to guarantee consistency over time, epigenetic methods, specifically DNA methylation therefore provides a state-of-the-art technology to determine at any stage of the bioprocessing of a CHO cell, particularly at an early stage of biologies production in the CHO cell, whether the CHO cell tested shows any, and if so which, phenotypes of interest to determine its suitability for genetic modification and other science research. The method according to any aspect of the present invention allow the use of DNA methylation as a tool to predict protein production quantitatively and qualitatively from CHO cells.

[0012] According to one aspect of the present invention, there is provided a method of establishing a biomarker for a first phenotype of interest of a CHO cell, the method comprising the steps of:

[0013] (a) measuring methylation values of all CpG sites within genomic DNA obtained from a population of CHO cells that are a representation of the first phenotype of interest and are part of the training samples,

[0014] (b) defining a set of specific CpG sites having consistent and reproducible methylation values in the training samples of step (a); and

[0015] (c) performing a penalized regression using the methylation values of step (a) as input and phenotype of interest correlated to the training samples as dependent variable, by applying a penalized regression model; thereby obtaining the specific CpG sites with corresponding weighting factors and intercept of the linear model equation as parameters defining the biomarker for the first phenotype of interest. The biomarker according to any aspect of the present invention may also be called a ‘CHO methylation clock’ which provides a correlation between phenotypes of interest expressed by a CHO cell and DNA methylation patterns. Using the method according to any aspect of the present invention, a CHO fitness score for an unknown sample may be quantitatively analysed using the set of specific CpG sites related to each phenotype of interest.

[0016] The term ‘biomarker’ as used herein refers to a naturally occurring characteristic by which a particular pathological or physiological process can be identified in a CHO cell. In particular, the biomarker refers to a feature, particularly a DNA methylation pattern in a CHO cell that may be used to identify a phenotype of interest in the CHO cell. The biomarker for identifying a phenotype of interest in the CHO cell thus refers to specific CpG sites with corresponding weighting factors and intercept of the linear model equation as parameters. The biomarker for a first phenotype of interest refers to a first set of specific CpG sites with corresponding weighting factors and intercept of the linear model equation. The biomarker for a second phenotype of interest refers to a second set of specific CpG sites with corresponding weighting factors and intercept of the linear model equation. The biomarker for a third, fourth and subsequent phenotypes of interest refers to a third, fourth and subsequent set of specific CpG sites with corresponding weighting factors and intercept of the linear model equation. The biomarkers may be complied to obtain a collection of biomarkers to determine a range of phenotypes of interest of a CHO cell. In particular, the set of specific CpG sites are determined from all the CpG sites of the CHO cell based on the methylation values measured in step (a) using any method known in the art for measuring DNA methylation. In particular, the specific CpG sites are selected based on expressing reliable methylation values (i.e. the methylation values of the specific CpG sites are consistent and reproducible using any method known in the art for measurement of methylation). The set of specific CpG sites result in the same or significantly similar methylation values when DNA methylation is measured using the same method over many more than one measurement. Accordingly, the CpG sites result in consisting and reproducible readings of methylation values.

[0017] Steps (b) and (c) according to any aspect of the present invention may be carried out by a skilled person using a computer.

[0018] The term ‘CHO cell genome’ herein refers to the genomic DNA of the CHO cell that excludes the DNA of a virus, particularly CMV and SV40, that are used to introduce foreign DNA to the cell. In particular, the CHO cell genome may denote the cell with a genome make-up that is in a form as seen naturally in the wild. The term may also include genes which have been added to the CHO genome by genetic modification (i.e. with regard to improved production of protein etc.) but not necessarily or not genes and promoters of viruses that have been used to introduce the genes into the CHO genome. The term “CHO cell genome” therefore may exclude virus genes and promoters and / or may include endogenous or homologous genes of the CHO cell and / or genetically modified endogenous or homologous genes of the CHO cell and / or intergenic genes, DNA found between the genes of the CHO cell. The CHO cell line refers to immortal Chinese Hamster Ovary cell line (CHO) derived from Cricetulus griseus. In particular, the CHO cell line may be selected from the group consisting of CHO-K1 (ATCC), CHO-DG44 (Thermo Fisher Scientific), CHO-DXB11 (ATCC), ExpiCHO-S™ cells (Thermo Fisher Scientific), Freestyle™ CHO-S™ cells (Thermo Fisher Scientific), CHO 1-15 [subscript 500] (ATCC) and Agarabi CHO (ATCC).

[0019] As used herein, the term ‘population of CHO cells’ refers to more than one CHO cell used as reference CHO cells that display at least one phenotype of interest. The different CHO cells expressing the phenotype of interest are then used to determine different CpG sites that are commonly methylated to determine a methylation pattern or reference methylation profile for the first phenotype of interest. In one example, the different CpG sites are collected from more than one CHO cell where each cell displays a first phenotype of interest to obtain a reference methylation profile for a first phenotype of interest or a biomarker for the first phenotype of interest. The reference methylation profile according to any aspect of the present invention may thus not be a naturally occurring methylation profile from a single CHO cell but an artificial profile obtained from combining relevant CpG sites from different reference CHO cell lines, each with at least one phenotype of interest. The population of different CHO cells are considered a representation of the first phenotype of interest and the methylation patterns from each of the CHO cells that form the population of different CHO cells for the first phenotype of interest form part of the training samples for producing the biomarker for the first phenotype of interest. In one example, the different CpG sites are collected from a single reference CHO cell that displays at least one phenotype of interest. In another example, the different CpG sites are collected from more than one CHO cell where each cell displays at least one phenotype of interest. The reference profile may be defined by multiple training samples through multivariate statistical methods, such as Principal Component analysis or Multi- Dimensional Scaling.

[0020] As used herein, the term ‘phenotype of interest’ in connection with a CHO cell refers to the cell displaying at least one the following characteristics selected from the group consisting of optimal heterologous protein production, phenotypic homogeneity, protein quality, optimal carbohydrate metabolism, optimal amino acid metabolism, optimal lipid metabolism, optimal cell survivability and combinations thereof. In particular, the phenotype of interest refers to a characteristic that the CHO cell according to any aspect of the present invention displays that is beneficial to the survival of the cell, suitability of the cell for protein production and the overall protein production of the cell. All these phenotypes together may be used to determine a CHO fitness score or the CHO cell fitness level. The more phenotypes of interest or optimal phenotypes of interest that the CHO cell expresses the higher the fitness level and vice versa.

[0021] The term ‘suitability’ as used herein, refers to a CHO cell line that is fit for optimal heterologous protein production. In one example, a CHO cell may be considered suitable for optimal heterologous protein production before a transgene is introduced into the cell. In this case, the CHO cell may have at least one phenotype of interest or characteristics that enable the cell to grow well and allow for easy uptake of the transgene of interest and following the uptake of the transgene, allow for optimal heterologous protein production, where the protein is a product of the transgene of interest. These characteristics or phenotype of interest include at least optimal glucose consumption, growth rate, lactic acid production, ammonia accumulation and the like. When a CHO cell is confirmed of displaying at least one of these phenotypes of interest, the CHO cell may be considered suitable for optimal heterologous protein production when the transgene of interest is introduced into the cell.

[0022] In another example, a CHO cell may be considered suitable for optimal heterologous protein production after the transgene has been introduced into the cell. In this case, a CHO cell is genetically modified using methods known in the art to introduce a transgene into the cell and the genetically modified cell is capable of optimal heterologous protein production where the protein is a product of translation of the transgene. The CHO cell in this example, may have a least one phenotype of interest that enables the genetically modified cell line to have good viability and optimal target protein production. These phenotypes of interest may include cell viability (survivability), protein productivity (in terms of protein quantity and quality), phenotypic homogeneity, cell exhaustion, and the like. Accordingly, the method according to any aspect of the present invention may be used on a CHO cell that has been genetically modified (i.e. with transgene introduced into the cell) or on a CHO cell that has not yet been genetically modified. In both cases, the CHO cell may be for use in heterologous protein production.

[0023] As used herein, the term ‘transgene’ refers to a gene that is taken from the genome of one organism and inserted into the genome of another organism by artificial techniques used in genetic modification. For example, a human gene is artificially introduced into the genome of CHO cells for the production of at least one protein of interest, particularly therapeutic proteins.

[0024] As used herein, the term ‘therapeutic protein’ refers to genetically engineered versions of naturally occurring human proteins. Examples of therapeutic proteins include antibody-based drugs, anticoagulants, blood factors, bone morphogenetic proteins, engineered protein scaffolds, enzymes, growth factors, hormones, interferons, interleukins and the like.

[0025] As used herein, the term ‘cell survivability’ refers to the capability of a cell to be viable and perform cell proliferation. Cell viability is a measure of the proportion of live cells within a population. Cell proliferation refers to an increase in cell number due to cell division. The assays that are commonly used to test cell survivability include BrdU Cell Proliferation Assay, MTT Cell Proliferation Assays, trypan blue cell counting, and ATP Cell Viability Assays.

[0026] As used herein, the term ‘cell exhaustion’ refers to the state of the cell where it loses its capability to perform metabolic activity including heterologous protein production. Cell exhaustion can be determined by Metabolite Detection Assays. As used herein, the term ‘phenotypic homogeneity’ refers to a state when all the cells in a population exhibit the same phenotype under a certain condition.

[0027] The term ‘heterologous protein production’ as used herein refers to the production of a protein which is not endogenous to the cell. It means an expression of a gene or part of a gene, particularly a transgene in a host CHO cell which does not naturally express this gene. The assays that are commonly used to quantify heterologous protein production include enzyme-linked immunosorbent assay (ELISA), chromatography & bioprocess analyser. The term ‘host cell’ as used herein refers to a cellular system for the expression of heterologous protein. For example, CHO cells are the main hosts for the production of various therapeutic proteins.

[0028] The term ‘optimal heterologous protein production’ herein refers to CHO cells that are capable of high- level protein production, particularly during industrial production or large-scale production of recombinant proteins, where the protein is usually a functional protein that is not naturally occurring in the wild-type CHO cell. In particular, for optimal heterologous protein production a CHO cell has minimized metabolic burdens and toxic effects to the cell. More in particular, ‘optimal heterologous protein production’ refers to high level protein production where the CHO cell not only produces a high yield of the protein of interest but also that the protein production is constantly maintained over the period of production (i.e., the prolonged period of culture) such that the quality of the protein produced is also consistent and maintained. In particular, for a CHO cell according to any aspect of the present invention to be capable of ‘optimal heterologous protein production’, the cell must at least display one of more of the following phenotypes of interest: phenotypic homogeneity, protein productivity, and protein quality. More in particular, for ‘optimal heterologous protein production’, the CHO cell may comprise phenotypic homogeneity and protein productivity, or phenotypic homogeneity, and protein quality, or protein productivity, and protein quality, or phenotypic homogeneity, protein productivity, and protein quality.

[0029] The term ‘protein productivity’ as used herein refers to a measure of the amount of protein made per viable cell at a single titre point. It is calculated by dividing the titre (mg) by the viable cell density (VCD or cells / ml), and the final measurement is represented as the amount of protein per cell (mg / cell).

[0030] The term ‘protein quality’ refers to the posttranslational modification of the protein that determines the efficacy and function of the protein. The modifications generally include phosphorylation, glycosylation, ubiquitination, methylation, acetylation, protein folding etc. For example, protein glycosylation is a critical quality attribute that modulates the efficacy, stability, and half-life of a therapeutic protein. Protein quality can be determined using Immunoprecipitation based techniques, Biochemical Assays, Mass spectrometry (MS) and the like.

[0031] The term ‘carbohydrate metabolism’, as used herein refers to almost all or all of the biochemical processes responsible for the metabolic formation, breakdown, and interconversion of carbohydrates in cells. It involves multiple pathways such as glycolysis, gluconeogenesis, glycogenolysis, and glycogenesis. For example, glycolysis is one of the key metabolic pathways of CHO cells. Through glycolysis, CHO cells consume glucose as the main carbon source for energy production and generate lactate as the most common metabolic by-product. Particularly, the term ‘optimal carbohydrate metabolism’ refers to the ideal or best carbohydrate metabolism possible by a CHO cell.

[0032] Similarly, the term ‘amino acid metabolism’ as used herein refer to the whole of the biochemical processes responsible for the metabolic formation, breakdown, and interconversion of amino acids in CHO cells. Amino acids are the basic building blocks of proteins and constitute all proteinaceous material of the cell including the cytoskeleton, protein component of enzymes, receptors, and signalling molecules. In addition, amino acids are utilized for the growth and maintenance of cells. For example, glutaminolysis is a key metabolic pathway of CHO cells. Glutaminolysis is the prevalent pathway through which CHO cells assimilate organic nitrogen for biomass synthesis while releasing ammonium as the main byproduct. Particularly, the term ‘optimal amino acid metabolism’ refers to the ideal or best amino acid metabolism possible by a CHO cell.

[0033] The term ‘lipid metabolism’ as used herein refers to the synthesis and degradation of lipids in CHO cells, involving the breakdown or storage of fats for energy and the synthesis of structural and functional lipids. Lipids are the major component of cellular membranes, act as secondary messengers in cell communication, involved in signalling, transport and secretion. Lipids are also an important source of energy through p-oxidation and the tricarboxylic acid (TCA) cycle. Lipid metabolism can have a significant impact on cell growth. For example, the process of triacylglycerol synthesis and degradation in CHO cells can greatly affect overall cellular metabolism and viability. Particularly, the term ‘optimal lipid metabolism’ refers to the ideal or best amino acid metabolism possible by a CHO cell.

[0034] Carbohydrate, amino acid and lipid metabolism can be determined by Metabolite Detection Assays, HPLC and bioprocess analyser. These methods are further disclosed at least in Coulet, M. et al., Cells (2022), 11 , 1929; Fan Y, et al., Biotechnol Bioeng (2015) 112(3):521 -535 and Ali AS, et al., Biotechnol J. (2018); 13(10):e1700745.

[0035] In particular, the phenotype of interest according to any aspect of the present invention is selected from the group consisting of phenotypic homogeneity, protein quality, optimal carbohydrate metabolism, optimal amino acid metabolism, optimal lipid metabolism, optimal heterologous protein production, optimal cell survivability and combinations thereof.

[0036] In context of the present invention, the term “methylation value’ refers to the average methylation value of at least one cytosine (C) residue within the genomic DNA sequence of a CHO cell. Both p-value and M- value may be used as metrics to measure (average) methylation levels. The M-value is more statistically valid for the differential analysis of methylation levels. However, the p-value is much more biologically interpretable and needs to be reported when M-value method is used for conducting differential methylation analysis. Any method known in the art may be used to determine the methylation value of a CpG site or a region in the DNA with many CpG sites. In one example, DNA array may be used.

[0037] The term “methylation ratio” refers to number of methylated cytosine(s) divided by the total number of cytosine(s) covered at the specific site(s). The methylation ratios for the CpG sites in step (b) may be advantageously determined using bisulfite sequencing.

[0038] The term “read coverage of the CpG site” is to be understood as the number of reads that align with the known CpG site in the reference sequence. The methylation ratio and the read coverage of the genomic CpG sites may be determined in step (a) using bisulfite sequencing.

[0039] Step b) according to the first aspect of the present invention may include a bisulfite conversion process. In step c) a regression analysis may be used determine the phenotypes of interest displayed by the CHO cell tested.

[0040] Bisulfite treatment’ of genomic DNA used interchangeably with the term ‘bisulfite modification’, refers to the treatment of the genomic DNA with a deaminating agent such as a bisulfite that may be used to treat all DNA, methylated or not. In particular, the term “bisulfite” as used herein encompasses any suitable type of bisulfite, such as sodium bisulfite, or other chemical agents that are capable of chemically converting a cytosine (C) to an uracil (U) without chemically modifying a methylated cytosine and therefore can be used to differentially modify a DNA sequence based on the methylation status of the DNA, e.g., U.S. Pat. Pub. US 2010 / 0112595. As used herein, a reagent that "differentially modifies" methylated or non-methylated DNA encompasses any reagent that modifies methylated and / or unmethylated DNA in a process through which distinguishable products result from methylated and nonmethylated DNA, thereby allowing the identification of the DNA methylation status. Such processes may include, but are not limited to, chemical reactions (such as a C to U conversion by bisulfite) and enzymatic treatment (such as cleavage by a methylation-dependent endonuclease). Thus, an enzyme that preferentially cleaves or digests methylated DNA is one capable of cleaving or digesting a DNA molecule at a much higher efficiency when the DNA is methylated, whereas an enzyme that preferentially cleaves or digests unmethylated DNA exhibits a significantly higher efficiency when the DNA is not methylated.

[0041] Accordingly, before step (a) according to any aspect of the present invention is carried out, the genomic DNA contained / obtained or extracted from the cell, is first bisulfite treated.

[0042] An alternative method available in the art may be used instead of bisulfite treatment. A skilled person will understand which other methods to use. In one example, TET-assisted pyridine borane sequencing (TAPS) may be used for detection of 5mC and 5hmC (Yibin Liu, et al., Nature Biotechnology, 37: 424- 429 (2019).

[0043] Whole genome bisulfite sequencing is a genome-wide analysis of DNA methylation based on the sodium bisulfite conversion of genomic DNA, which is then sequenced on a next-generation sequencing platform. The sequences are then re-aligned to the reference genome to determine methylation states of the CpG dinucleotides based on mismatches resulting from the conversion of unmethylated cytosines into uracil. In particular, in step (a) a methylation ratio and read coverage of the CpG sites are determined using bisulfite sequencing; and in step (b), the set of specific CpG sites are defined using a cutoff value using bisulfite sequencing.

[0044] The term “test” used in conjunction with the term cell herein refers to a cell that is subjected to the method according to any aspect of the present invention and is the basis for an analysis application of the present invention. A ‘test cell’ is therefore a CHO cell or a group of CHO cells being tested according to any aspect of the present invention, or a profile being obtained or generated in this context. Conversely, the term “reference” or ‘control’ shall denote, mostly predetermined, entities which are used for a comparison with the test entity. In particular, a ‘test cell’ refers to a cell being tested for suitability of optimal homologous protein production where the methylation status has to be determined and a ‘control’ or ‘reference’ refers to a cell which is known to display optimal homologous protein production or a methylation profile thereof.

[0045] As used herein, a “CpG site” or “methylation site” is a nucleotide within a nucleic acid (DNA or RNA) that is susceptible to methylation either by natural occurring events in vivo or by an event instituted to chemically methylate the nucleotide in vitro. Some of these sites may be hypermethylated and some may be hypomethylated in a cell. In some cases a CpG site may not be considered fully hypermethylated or hypomethylated but a value may be given that is a measure of methylation of the CpG site. Accordingly, methylation may be quantified and may not always be an absolute case of hypermethylation or hypomethylation.

[0046] As used herein, a “methylated nucleic acid molecule” refers to a nucleic acid molecule that contains one or more nucleotides that is / are methylated.

[0047] The phenotype of interest-correlated training sample of a CHO cell in step (a) may, for example, be taken at any stage of growth of the CHO cell. The sample is always the CHO genomic DNA.

[0048] In particular, the penalized regression model is elastic net linear regression model. Elastic net is used to select CpGs, differentially methylated regions (DMRs), lowly methylated regions (LMRs) and CpGs of a subset of that are most predictive of each phenotype of interest. These selected specific CpGs according to any aspect of the present invention can then be used to develop a model for predicting cell fitness based on DNA methylation signatures. More in particular, the Elastic net algorithm is a machine learning based method that combines traditional Lasso and ridge regression techniques, placing emphasis on model sparsity while appropriately balancing the contributions of correlated variables. It is particularly useful in the method according to any aspect of the present invention for constructing linear models where the number of variables is substantially greater than the number of samples. In particular, the specific CpG sites according to any aspect of the present invention are distributed within low methylated regions (LMRs), CpG islands, variably methylated sites and / or differentially methylated regions (DMRs) in the genome of the CHO cell.

[0049] Low Methylated Region (LMR) is a region of the genome wherein less than 60% of CpGs in that region are methylated. More in particular, less than 50%, 40%, 30%, 20% or 10% of the CpGs in the LMRs are methylated. Any method known in the art may be used to identify or detect LMRs in the genomic DNA. Well known methods include using programmes such as MethylSeekR. In particular, LMRs in the genomic DNA have at least three consecutive CpGs and have no single nucleotide polymorphisms (SNPs) in any of the CpG positions. Even more in particular, LMRs in the genomic DNA are identified based on the method disclosed at least in Burger, L., (2013) Nucleic Acids Research, 41 (16): e155 and / or Stadler, M., (2011) Nature 480, 490-495. LMRs are known to have an average methylation ranging from 10% to 50%; are regions of low CG density which do not overlap with CpG islands; tend to be enriched for H3K4me1 , DHSs, and p300 / CBP; and / or are primarily located distal to promoters in intergenic or intronic regions. In particular, LMRs: have an average methylation ranging from 10% to 50%, are regions of low CG density; are enriched for Histone H3 monomethylated at lysine 4 (H3K4me1), DNase I hypersensitive sites (DHSs) and transcriptional coactivators CREB binding protein (CPB) and p300; are primarily located distal to promoters in intergenic or intronic regions; and / or have no single nucleotide polymorphisms (SNPs) in any of the CpG positions.

[0050] Low-methylated regions (LMRs) represent a key feature of the dynamic methylome. LMRs are local reductions in the DNA methylation landscape and represent CpG-poor distal regulatory regions that often reflect the binding of transcription factors and other DNA-binding proteins. LMRs were originally described in the mouse (Stadler et al. (2011) Nature: 480, 490-95). Evolutionary conservation of LMRs beyond mammals has remained unexplored.

[0051] A “CpG island” as used herein describes a segment of DNA sequence that comprises a functionally or structurally deviated CpG density. For example, Yamada et al. have described a set of standards for determining a CpG island: it must be at least 400 nucleotides in length, has a greater than 50% GC content, and an OCF / ECF ratio greater than 0.6 (Yamada et al., 2004, Genome Research, 14, 247-266). Others have defined a CpG island less stringently as a sequence at least 200 nucleotides in length, having a greater than 50% GC content, and an OCF / ECF ratio greater than 0.6 (Takai et al., 2002, Proc. Natl. Acad. Sci. USA, 99, 3740-3745). In context of the present invention, the terms “methylation profile”, “methylation pattern”, “methylation state” or “methylation status,” are used herein to describe the state, situation or condition of methylation of a genomic sequence, and such terms refer to the characteristics of a DNA segment at a particular genomic locus in relation to methylation. Such characteristics include, but are not limited to, whether any of the cytosine (C) residues within this DNA sequence are methylated, location of methylated C residue(s), percentage of methylated C at any particular stretch of residues, and allelic differences in methylation due to, e.g., difference in the origin of the alleles.

[0052] Differentially methylated regions (DMRs) are genomic regions with different methylation statuses among multiple biological samples like tissues, cells, individuals, etc. These are genomic regions that differ between phenotypes. The statistical power is likely to be greater when adjacent differentially methylated points (DMPs) are considered together as a whole [Gu H et al (2010) Nat Methods 2010; 7:133-6]. The lengths of the DMRs may range between a few hundred to a few thousand bases [Rakyan et al (2011) Nat Rev Genet 12:529-41 , 2011 , Bock C (2012) Nat Rev Genet 2012; 13:705-19],

[0053] DMRs may occur throughout the genome but have been identified particularly around the promoter regions of genes, within the body of genes, and at intergenic regulatory regions. There are two types of regions, predefined or user defined. Regions with special biological meaning, such as CpG islands, CpG shores, UTRs and so on, are predefined. Many traditional statistical testings, including t-test and Wilcoxon rank sum test, can be performed at a region level. For user-defined regions, criteria such as a fixed region length, fixed numbers of significant and adjacent CpG sites, significant and smoothed estimated effect sizes, etc.

[0054] Partially methylated domains (PMDs) are extended regions in the genome exhibiting a reduced average DNA methylation level. They cover gene-poor and transcriptionally inactive regions and tend to be heterochromatic.

[0055] Differentially methylated Positions (DMP) are CpG sites with different DNA methylation status across different biological samples and regarded as possible functional regions involved in gene transcriptional regulation.

[0056] In particular, the steps (a)-(c) according to any aspect of the present invention are repeated for a second phenotype of interest and subsequent phenotypes of interests thereafter such that there is at least one biomarker for each phenotype of interest to produce a compilation of biomarkers for different phenotypes of interest.

[0057] In particular, the biomarker according to any aspect of the present invention is a set of specific CpG sites, corresponding weighting factors and intercept of the linear model equation that are relevant to a first phenotype of interest. A compilation of biomarkers therefore refers to more than one biomarker wherein each biomarker is correlated to at least one phenotype of interest displayed by the CHO cell.

[0058] The coverage cutoff value defined in step (b) may be a minimum of 10, 9, 8, 7, 6, 5, 4 or 3. In particular, the coverage cutoff value defined in step (b) may be a minimum of 3. The method according to any aspect of the present invention may further comprise a step of (d) optimizing the fit of the regression model such that the number of CpG sites is reduced to a minimum. For optimizing the fit of the regression model in step (d) or in general for optimizing the fit of the regression model, the algorithm preferably applies ridge regression in combination with lasso regression.

[0059] Advantageously, the methods of ridge regression and lasso regression are balanced using an alpha value of between 0 and 1 .

[0060] According to a further aspect of the present invention, there is provided a method of predicting cell fitness of a CHO test cell, the method comprising the steps of:

[0061] (a) measuring DNA methylation levels of specific CpG sites in extracted genomic DNA from the CHO test cell; and

[0062] (b) comparing the methylation levels of these specific CpG sites from step (a) with methylation levels of the same specific CpG sites from reference CHO cells known to express at least one phenotype of interest; and

[0063] (c) deducing therefrom the phenotypes of interest(s) expressed by the test CHO cell, and wherein the cell fitness of the CHO test cell is determined by the number of phenotypes of interests that the CHO test cell expresses, and wherein the specific CpG sites in steps (a) and (b) are determined using the method according to any aspect of the present invention.

[0064] The term ‘cell fitness’ of a CHO cell as used herein refers to the ability of the CHO cell to grow, produce protein, survive and display other positive characteristics that make the cell suitable for use in cell culture. In particular, cell fitness of a CHO cell may be determined by the phenotypes of interest that the cell displays.

[0065] The phenotype of interest-correlated reference sample in step (b) serves as a control and represents an average methylation level for a pre-determined and specific phenotype of interest in a CHO cell. The method for predicting cell fitness of a CHO cell may be used for testing individual CHO cell and for testing a complete population of CHO cells.

[0066] The biological sample material deriving from the CHO cell refers to genomic DNA. As used herein, the term “genomic material” refers to nucleic acid molecules or fragments of the genome of the CHO cells. In particular, such nucleic acid molecules or fragments are DNA or RNA or hybrids thereof, and most preferably are molecules of the DNA genome of CHO cells.

[0067] As used herein, the “DNA sample” refers to the DNA extracted from the CHO cell according to any aspect of the present invention using known methods in the art.

[0068] According to a further aspect of the present invention, there is provided an in vitro method for predicting cell fitness of a CHO test cell, the method comprising the steps of: a) measuring DNA methylation levels of specific CpG sites in extracted genomic DNA from the CHO test cell and multiplying same with their respective weighing factors to obtain the weighted methylation ratios of those CpG sites, b) computing the sum over the weighted methylation ratios obtained in step (a) and adding the respective intercept of linear model equation, thus predicting amount of expression of a phenotype of interest in the CHO test cell, wherein the specific CpG sites, the respective weighing factors and the respective intercept of linear model equation are the parameters of a biomarker for a phenotype of interest and the cell fitness of the CHO test cell is determined by the number of phenotypes of interests that the CHO test cell expresses.

[0069] In particular, the biomarkers according to any aspect of the present invention are determined using the method of the first aspect of the present invention.

[0070] The DNA methylation level according to any aspect of the present invention may be determined using any method known in the art. In one example, the methylation levels can be measured using the commercial Illumina™ platform. In another example, the methylation levels can be determined using a DNA methylation bead based array.

[0071] Arrays allow for a high-throughput and robust method to determine semi-quantitative / quantitative DNA- methylation information through a small sample of extracted DNA of interest. These custom designed arrays may use Illumina iScan and Infinium platform technology or an equivalent thereof, which allows on each chip for example 100,000 different bead types that covalently bind DNA-methylation probes. Each probe represents one CpG Methylation site at the end of the probe sequence. DNA samples undergo bisulfite conversion, amplification, fragmentation, precipitation and resuspension steps before hybridization on an array chip. Once on the chip the DNA hybridizes to the beads for each CpG site so that methylation changes at each site can be detected specifically through single nucleotide extension. This is especially advantageous as the array-based method is simple and the results of the array are accurate and reproducible.

[0072] Further, compared to traditional sequencing which can take weeks to generate data, the array technology has a much shorter turn-around time. The volume and complexity of data generated is lesser compared to sequencing making it computationally less intensive. This allows for quicker computation to achieve interpretable results from experimental groups. Overall microarray technology is roughly 10x faster and 10x cheaper than traditional sequencing while still quantifiable for the methylation level at specific CpG sites.

[0073] The term “array” as used herein refers to an intentionally created collection of probe molecules which can be prepared either synthetically or biosynthetically. The probe molecules in the array can be identical or different from each other. The array can assume a variety of formats, for example, libraries of soluble molecules; libraries of compounds tethered to resin beads, silica chips, or other solid supports.

[0074] In particular, an array provides a convenient platform for simultaneous analysis of large numbers of CpG sites, for example, at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 50, 100, 500, 1000, 5000, 10,000, 100,000 or more sites or loci. In particular, the array comprises a plurality of different probe molecules that can be attached to a substrate or otherwise spatially distinguished in an array. Examples of arrays that may be used according to any aspect of the present invention include slide arrays, silicon wafer arrays, liquid arrays, bead-based arrays and the like. In one example, array technology used according to any aspect of the present invention combines a miniaturized array platform, a high level of assay multiplexing, and scalable automation for sample handling and data processing.

[0075] In particular, the array according to any aspect of the present invention may be an array of arrays, also referred to as a composite array, having a plurality of individual arrays that is configured to allow processing of multiple samples simultaneously. Examples of composite arrays and the technology behind them are disclosed at least in US 6,429,027 and US 2002 / 0102578. A substrate of a composite array may include a plurality of individual array locations, each having a plurality of probes, and each physically separated from other assay locations on the same substrate such that a fluid contacting one array location is prevented from contacting another array location. Each array location can have a plurality of different probe molecules that are directly attached to the substrate or that are attached to the substrate via rigid particles in wells (also referred to herein as beads in wells).

[0076] In one example, an array substrate can be a fibre optical bundle or array of bundles as described in US6,023,540, US6,200,737 and / or US6,327,410. An optical fibre bundle or array of bundles can have probes attached directly to the fibres or via beads. A skilled person would be able to easily determine which substrate will be most suitable for the array according to any aspect of the present invention. W020041 10246 further discloses other substrates and methods of attaching beads to the substrates that may be used in the array according to any aspect of the present invention.

[0077] In one example, a surface of the substrate may have physical alterations to enable the attachment of probes or produce array locations. For example, the surface of a substrate can be modified to contain chemically modified sites that are useful for attaching, either-covalently or non-covalently, probe molecules or particles having attached probe molecules. Probes may be attached using any of a variety of methods known in the art including, an ink-jet printing method, a spotting technique, a photolithographic synthesis method, or printing method utilizing a mask. W02004110246 discloses these techniques in more detail.

[0078] In one example, the array according to any aspect of the present invention may be a bead-based array, where the beads are associated with a solid support such as those commercially available from Illumina, Inc. (San Diego, Calif.). An array of beads useful according to any aspect of the present invention can also be in a fluid format such as a fluid stream of a flow cytometer or similar device. Commercially available fluid formats for distinguishing beads include, for example, those used in XMAP(TM) technologies from Luminex or MPSS(TM) methods from Lynx Therapeutics.

[0079] The term “solid support”, “support”, and “substrate” as used herein are used interchangeably and refer to a material or group of materials having a rigid or semi-rigid surface or surfaces. In many examples, at least one surface of the solid support will be substantially flat, although in some examples it may be desirable to physically separate synthesis regions for different compounds with, for example, wells, raised regions, pins, etched trenches, or the like.

[0080] The array or microarray according to any aspect of the present invention may be a very high-density array, for example, those having from about 10,000,000 probes / cm2to about 2,000,000,000 probes / cm2or from about 100,000,000 probes / cm2to about 1 ,000,000,000 probes / cm2. High density arrays are especially useful according to any aspect of the present invention for including the multitude of CpG sites on the array.

[0081] The array according to any aspect of the present invention may be used to analyse or evaluate such pluralities of loci simultaneously or sequentially as desired. In one example, a plurality of different probe molecules can be attached to a substrate or otherwise spatially distinguished in an array. Each probe is typically specific for a particular locus and can be used to distinguish methylation state of the locus.

[0082] The term “probe molecules” or ‘probes’ as used interchangeably herein refers to a surface-immobilized molecule that can be recognized by a particular target. Probes used in the array can be specific for the methylated allele of a CpG site, the non-methylated allele of the CpG site or both or for the methylated allele of a non-CpG site, the non-methylated allele of the non-CpG site or both.

[0083] The term “target” as used herein refers to a molecule that has an affinity for a given probe molecule. Targets may be naturally occurring or man-made molecules. Also, they can be employed in their unaltered state or as aggregates. Targets may be attached, covalently or noncovalently, to a binding member, either directly or via a specific binding substance. Examples of targets which can be employed according to any aspect of the present invention are methylated and non-methylated CpG sites. Targets are sometimes referred to in the art as anti-probes. As the term targets is used herein, no difference in meaning is intended.

[0084] The term “complementary” as used herein refers to the hybridization or base pairing between nucleotides or nucleic acids, such as, for instance, between the two strands of a double stranded DNA molecule or between an oligonucleotide primer and a primer binding site on a single stranded nucleic acid to be sequenced or amplified. Complementary nucleotides are, generally, A and T (or A and U), or C and G. Two single stranded RNA or DNA molecules are said to be complementary when the nucleotides of one strand, optimally aligned and compared and with appropriate nucleotide insertions or deletions, pair with at least about 80% of the nucleotides of the other strand, usually at least about 90% to 95%, and more preferably from about 98 to 100%. Perfectly complementary refers to 100% complementarity over the length of a sequence. For example, a 25-base probe is perfectly complementary to a target when all 25 bases of the probe are complementary to a contiguous 25 base sequence of the target with no mismatches between the probe and the target over the length of the probe.

[0085] According to another aspect of the present invention, there is provided a computer-implemented method of establishing a biomarker for a first phenotype of interest of a population of CHO cells, the method comprising the steps of: a) inputting methylation values of all CpG sites within genomic DNA obtained from a population of CHO cells that express a specific phenotype of interest associated to cell fitness of the cell and are part of the training samples, b) identifying and determining specific CpG sites from all the CpG sites in step (a) showing consistent and reproducible methylation values, and c) correlating the CpG methylation levels of the CpG sites obtained in step (a) with the phenotype of interest using elastic Net linear regression model thereby obtaining the specific CpG sites with corresponding weighting factors and intercept of the linear model equation as parameters defining the biomarker for the first phenotype of interest.

[0086] According to another aspect of the present invention, there is provided a computer program loaded into a memory of a computer, implementing the method according to any aspect of the present invention.

[0087] According to yet another aspect of the present invention, there is provided a tangible computer-readable medium comprising computer-readable code that, when executed by a computer, causes the computer to perform the method according to any aspect of the present invention.

[0088] According to a further aspect of the present invention, there is provided a use of the medium according to any aspect of the present invention to predict the cell fitness of a CHO test cell, wherein the cell fitness of the cell is based on the expression of at least one phenotype of interest by the cell.

[0089] Unless stated otherwise, all percentages (%) given are percentages by mass.

[0090] The examples adduced hereinafter describe the present invention by way of example, without any intention that the invention, the scope of application of which is apparent from the entirety of the description and the claims, be restricted to the embodiments specified in the examples.

[0091] BRIEF DESCRIPTION OF FIGURES

[0092] Figure 1 is a scatter plot of predicted IgG vs actual IgG on the test datasets from Example 3.

[0093] EXAMPLES EXAMPLE 1 :

[0094] Predicting the specific productivity (pg / cell / day) from CHO cells

[0095] Wet-Lab methodology

[0096] For this experiment, sixty transgenic CHO clones are grown in basal at 37°C, 8% CO2, at a shaking speed of 225 RPM. The sixty transgenic CHO cell lines include low producers (specific productivity <10pg / cell / day), intermediate producer (specific productivity 10-20 pg / cell / day) and high producers (specific productivity >20pg / cell / day). The flasks are seeded with 3E5 viable cells / mL on day 0 and the culture is fed with appropriate feed on Day 3, 5, 7, 9, 11 and glucose is topped up to 6g / l using 45% glucose when it drops below 2g / l. The fed-Batch culture of 60 clones is maintained for 14 days. Cell count, cell viability, and specific productivity are measured for day 14 and cell pellets are collected on day 14.

[0097] DNA Extraction

[0098] DNA is extracted using the PureLink Genomic DNA Isolation Minikit kit (Invitrogen), including RNAase treatment following the manufacturer's instructions. DNA quantity is measured by PicoGreen assay and DNA quality is assessed via NanoDrop (Thermo Scientific) to ensure the A260 / 280 ratio is < 1 .8. A small amount of sample is then also analysed using automated electrophoresis on TapeStation (Agilent) to ensure each sample contains high molecular weight.

[0099] Whole-genome bisulfite seguencing (WGBS)

[0100] The genomic DNA (500ng) from the samples are used to prepare libraries for Whole Genome Bisulfite Sequencing (WGBS). The sequencing of the libraries is performed by a third party on a NovaSeq platform which generated 125GB data per sample with 20X coverage.

[0101] Bioinformatics methodology

[0102] Preprocessing of Whole-genome bisulfite sequencing (WGBS) data

[0103] Raw WGBS data undergoes quality control followed by adaptor trimming with trim galore (https: / / www.bioinformatics.babraham.ac.uk / projects / trim_galore / ). The paired reads are then mapped to the reference genome of [CriGri-PICRH-1 .0](https: / / www. ncbi.nlm.nih.gov / assembly / GCF_003668045.3 / ) by Bismark (https: / / www.bioinformatics.babraham.ac.uk / projects / bismark / ). Bismark is also utilized to perform deduplication and extract the number of reads having methylated CpGs and the number of total reads at this position. Methylation ratios are determined by dividing the number of methylated reads by total reads. data

[0104] The customized chip array data processing is performed in R version 4.1 .2 using sesame version 1 .14.2.

[0105] DNA methylation level for each site was calculated as methylation p-value. Beta values are defined as methylated signal / (methylated signal + unmethylated signal). It can be computed using getBetas function. The SeSAMe pipeline (Zhou et al. 2018) was used to generate normalized p-values and for quality control. Low intensity- based detection calling and making (based on p-value) was done with pOOBAH. Background subtraction based on normal-exponential deconvolution using out-of-band probes noob (Triche et al. 2013) and optionally with extra bleed-through subtraction were also implemented.

[0106] Model establishment

[0107] A penalized regression model (https: / / cran.r-project.org / web / packages / glmnet / index.html in R and https: / / scikit-learn.org / in Python) are used to regress the specific productivity to the methylation ratio of each CpG. Briefly, Hyperparameters, including the alpha values, which control the balance between L1 and L2 regularization, were tuned through cross-validation. Mean Squared Error (MSE), or Root Mean Squared Error (RMSE), or R-squared (R2) are used to evaluate the model performance for determining optimal hyperparameters. With fine-tuned hyperparameters, the elastic net model is fitted to the full dataset. Coefficients of the features are analyzed to understand the impact on productivity.

[0108] EXAMPLE 2.

[0109] Predicting the growth status in the form viable cell density (cells / ml) from CHO cells

[0110] Wet-Lab methodology

[0111] For this experiment, sixty transgenic CHO clones are grown in basal medium supplemented at 37°C, 8% CO2, at a shaking speed of 225 RPM. The sixty transgenic CHO cell lines include slow growers (viable cell density <1 E7cells / ml ), intermediate growers (viable cell density 1 E7-3E7 cells / ml) and fast growers (viable cell density >3E7 cells / ml). The flasks are seeded with 3E5 viable cells / mL on day 0 and the culture is fed with appropriate feed on Day 3, 5, 7, 9, 11 and glucose is topped up to 6g / l using 45% glucose when it drops below 2g / l. The fed-Batch culture of 60 clones is maintained for 14 days. Cell count and cell viability are measured for day 7 and cell pellets were collected on day 7.

[0112] DNA Extraction

[0113] DNA is extracted using the PureLink Genomic DNA Isolation Minikit kit (Invitrogen), including RNAase treatment following the manufacturer's instructions. DNA quantity is measured by PicoGreen assay and DNA quality is assessed via NanoDrop (Thermo Scientific) to ensure the A260 / 280 ratio is < 1 .8. A small amount of sample is then also analysed using automated electrophoresis on TapeStation (Agilent) to ensure each sample contains high molecular weight.

[0114] Sequencing Analysis

[0115] The genomic DNA (500ng) from the samples are used to prepare libraries for Whole Genome Bisulfite Sequencing (WGBS). The sequencing of the libraries are performed by a third party on a NovaSeq platform which generated 125GB data per sample with 20X coverage. bisulfite data Raw WGBS data undergoes quality control followed by adaptor trimming with trim galore (https: / / www.bioinformatics.babraham.ac.uk / projects / trim_galore / ). The paired reads are then mapped to the reference genome of [CriGri-PICRH-1 .0](https: / / www. ncbi.nlm.nih.gov / assembly / GCF_003668045.3 / ) by Bismark (https: / / www.bioinformatics.babraham.ac.uk / projects / bismark / ). Bismark is also utilized to perform deduplication and extract the number of reads having methylated CpGs and the number of total reads at this position. Methylation ratios are determined by dividing the number of methylated reads by total reads.

[0116] Preprocessing of methylation array data

[0117] The customized chip array data processing is performed in R version 4.1 .2 using sesame version 1 .14.2. DNA methylation level for each site is calculated as methylation p-value. Beta values are defined as methylated signal / (methylated signal + unmethylated signal). It can be computed using getBetas function. The SeSAMe pipeline (Zhou et al. 2018) is used to generate normalized p-values and for quality control. Low intensity- based detection calling and making (based on p-value) was done with pOOBAH.

[0118] Background subtraction based on normal-exponential deconvolution using out-of-band probes noob (Triche et al. 2013) and optionally with extra bleed-through subtraction are also implemented.

[0119] Model establishment

[0120] A penalized regression model (https: / / cran.r-project.org / web / packages / glmnet / index.html in R and https: / / scikit-learn.org / in Python) is used to regress the cell counts to the methylation ratio of each CpG. Briefly, Hyperparameters, including the alpha values, which control the balance between L1 and L2 regularization, were tuned through cross-validation. Mean Squared Error (MSE), or Root Mean Squared Error (RMSE), or R-squared (R2) were used to evaluate the model performance for determining optimal hyperparameters. With fine-tuned hyperparameters, the elastic net model was fitted to the full dataset. Coefficients of the features were analyzed to understand the impact on cell viability.

[0121] EXAMPLE 3:

[0122] C / 0 cell culture

[0123] The cells used in this study belong to the Humira431 CHO cell line (A*Star BTI), which has been derived from CHO DG44 cells modified to express a therapeutic antibody (Adalimumab biosimilar). Cells were seeded at 3 x 105cells / ml in 30 ml media in 125 ml shake flasks and culture conditions maintained at 37°C, 8% CO2, with shaking at 150 rpm. For fed-batch cultures, 10% v / v of EX-CELL Advanced CHO Feed 1 without glucose (24368C, Merck) was added every alternate day from days 3 to 14, and D-(+)-Glucose solution (G8769, Merck) added when glucose concentration fell to 2 g / L. Samples were taken from each flask every day from day 3 until the end of the culture for cell count and metabolite analysis. Viable cell density (VCD) was determined using the Vi-CELL BLU analyser (Beckman Coulter) and metabolite concentrations were measured using the Cedex Bio Analyser (Roche).

[0124] Whole genome bisulfite sequencing (WGBS): Humira CHO samples were cultured either by fed-batch culture or by batch culture, the IgG titre were measured, followed by whole genome bisulfite sequencing (WGBS).

[0125] The EZ DNA Methylation-Gold™ Kit was used for bisulfate conversion, while the VAHTS Universal Pro DNA Library Prep Kit was used for library preparation. The Illumina NovaSeq 6000 was used for whole genome bisulfate sequencing (WGBS).

[0126] Data processing:

[0127] The raw WGBS data underwent quality control with FastQC and adaptor trimming using TrimGalore. The aligned reads were mapped to the CriGri-PICRH-1 .0 reference genome using Bismark, which was also used for deduplication and extraction of beta values for each CpG site. Only samples with above 10 million CpG sites with coverage above 10 were retained for the following analysis.

[0128] To ensure the methylation data’s reliability and suitability for model building, all samples needed to share the same CpG sites with sufficient coverage. Various coverage thresholds were hence explored to determine the optimal level. The threshold of 10 was then selected based on its balance of stringency and the retention of a substantial number of CpG sites. CpGs that associated with culture types were removed to avoid confounding effects, refining the dataset to 543,613 CpG sites in 1 14 samples suitable for modeling.

[0129] Machine learning model development:

[0130] Elastic Net penalized regression model was used to predict IgG values based on methylation data from 543,613 CpG sites in 114 CHO Humira samples. The performance of the predictive models was assessed using Root Mean Squared Error (RMSE), a metric that measures the average magnitude of error between the predicted and actual values, with lower RMSE indicating more accurate predictions. Model implementation and hyperparameter optimization were performed using the SciKit-Learn library and GridSearchCV.

[0131] Results:

[0132] The Elastic Net model identified 850 CpG sites (Table 1) that potentially influence IgG productivity through their methylation patterns. The optimal value for the alpha parameter was determined to be 0.1 , resulting in a best cross-validation RMSE score of 12.89. The model was then applied to the test datasets to predict IgG values, and the correlation between the predicted and actual values is depicted in Figure 1 .

[0133] Table 1 850 CpG sites identified using Elastic Net model that potentially influence IgG productivity through their methylation patterns 202300097 Foreign Countries

[0134] 202300097 Foreign Countries

[0135] 202300097 Foreign Countries

[0136] 202300097 Foreign Countries

[0137] 202300097 Foreign Countries

Claims

CLAIMS1 . A method of establishing a biomarker for a first phenotype of interest of a CHO cell, the method comprising the steps of:(a) measuring methylation values of all CpG sites within genomic DNA obtained from a population of CHO cells that are a representation of the first phenotype of interest and are part of the training samples,(b) defining a set of specific CpG sites having consistent and reproducible methylation values in the training samples of step (a); and(c) performing a penalized regression using the methylation values of step (a) as input and phenotype of interest correlated to the training samples as dependent variable, by applying a penalized regression model; thereby obtaining the specific CpG sites with corresponding weighting factors and intercept of the linear model equation as parameters defining the biomarker for the first phenotype of interest.

2. The method according to claim 1 , wherein the penalized regression model is elastic net linear regression model.

3. The method according to either claim 1 or 2, wherein steps (b) and (c) are carried out on a computer.

4. The method according to any one of the preceding claims, wherein the steps (a)-(c) are repeated for a second phenotype of interest and subsequent phenotypes of interests thereafter such that there is at least one biomarker for each phenotype of interest to produce a compilation of biomarkers for different phenotypes of interest.

5. The method according to any one of the preceding claims, wherein for optimizing the fit of the regression model in step (c), the algorithm applies ridge regression in combination with lasso regression.

6. The method according to claim 5, wherein the methods of ridge regression and lasso regression are balanced using an alpha value of between 0 and 1 .

7. The method according to any one of the preceding claims, wherein the DNA methylation value is determined using a DNA methylation bead based array.

8. The method according to any one of the claims 1 to 6, wherein in step (a) a methylation ratio and read coverage of the CpG sites are determined using bisulfite sequencing; and in step (b), the set of specific CpG sites are defined using a cutoff value using bisulfite sequencing.

9. The method according to claim 8, wherein the read coverage cutoff value defined in step (b) is a minimum of 3.

10. An in vitro method for predicting cell fitness of a CHO test cell, the method comprising the steps of: a) measuring DNA methylation levels of specific CpG sites in extracted genomic DNA from the CHO test cell and multiplying same with their respective weighing factors to obtain the weighted methylation ratios of those CpG sites, b) computing the sum over the weighted methylation ratios obtained in step (a) and adding the respective intercept of linear model equation, thus predicting amount of expression of a phenotype of interest in the CHO test cell, wherein the specific CpG sites, the respective weighing factors and the respective intercept of linear model equation are the parameters of a biomarker for a phenotype of interest and the cell fitness of the CHO test cell is determined by the number of phenotypes of interests that the CHO test cell expresses wherein the biomarkers are determined using the method according to any one of the claims 1 to 9.

11. The method according to any one of the preceding claims, wherein the phenotype of interest is selected from the group consisting of phenotypic homogeneity, protein quality, optimal carbohydrate metabolism, optimal amino acid metabolism, optimal lipid metabolism, optimal heterologous protein production, optimal cell survivability and combinations thereof.

12. The method according to any one of the preceding claims, wherein the specific CpG sites are distributed within low methylated regions (LMRs), CpG islands, variably methylated sites and / or differentially methylated regions in the genome of the CHO cell.

13. The method according to any one of the claims 10 to 12, wherein the DNA methylation level is determined using a DNA methylation bead based array.

14. A computer-implemented method of establishing a biomarker for a first phenotype of interest of a population of CHO cells, the method comprising the steps of: a) inputting methylation values of all CpG sites within genomic DNA obtained from a population of CHO cells that express a specific phenotype of interest associated to cell fitness of the cell and are part of the training samples, b) identifying and determining specific CpG sites from all the CpG sites in step (a) showing consistent and reproducible methylation values, and c) correlating the CpG methylation levels of the CpG sites obtained in step (a) with the phenotype of interest using elastic Net linear regression model thereby obtaining the specific CpG sites with corresponding weighting factors and intercept of the linear model equation as parameters defining the biomarker for the first phenotype of interest.

Citation Information

Patent Citations

  • Alternative substrates and formats for bead-based array of arrays TM

    US20020102578A1

  • Bisulfite Conversion Reagent

    US20100112595A1

  • Method for the Selection of a Long-Term Producing Cell using Histone Acylation as Markers

    US20170081732A1

  • Fiber optic sensor with encoded microspheres

    US6023540A

  • Photodeposition method for fabricating a three-dimensional, patterned polymer microstructure

    US6200737B1