Predicting the quality of a bioprocess and product
The method predicts the quality of eukaryotic bioprocesses and products by analyzing CpG methylation indices and reference courses, using machine learning models, thereby addressing the need for advanced epigenetic monitoring and control in biopharmaceutical production.
Patent Information
- Application Number
- PCT/EP2024/087042
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-12-18
- Publication Date
- 2025-06-26
AI Technical Summary
There is a recognized need for advanced monitoring and control of epigenetic factors to refine bioprocesses and ensure the quality of biopharmaceuticals, particularly in predicting the quality of eukaryotic bioprocesses and products.
A method is developed to predict the quality characteristics of a eukaryotic bioprocess or product by determining a genomic CpG-methylation index for selected CpG-methylation sites and comparing it to a predetermined reference CpG methylation course, using statistical or machine learning models to estimate or predict product quality attributes or process parameters.
This method enables accurate prediction of bioprocess quality characteristics, allowing for real-time monitoring and control, thereby enhancing the efficiency and reliability of biopharmaceutical production and ensuring adherence to quality standards.
Smart Images

Figure EP2024087042_26062025_PF_FP_ABST
Abstract
Description
[0001] PREDICTING THE QUALITY OF A BIOPROCESS AND PRODUCT
[0002] FIELD OF THE INVENTION
[0003] The invention refers to a method of describing the effect of the environment of a eukaryotic cell and based on this understanding of the effect of the environment on the cells, the prediction of the quality of a bioprocess that is based on such eukaryotic cell or a product of said bioprocess. The invention refers also to a method of designing the quality of a bioprocess or product, based on changing and optimizing the environment of a eukaryotic cell. To achieve this understanding of the effect of the environment on the cells, a CpG methylation index and a CpG methylation course are analysed.
[0004] BACKGROUND OF THE INVENTION
[0005] Ensuring the quality of biopharmaceuticals is essential, especially when adhering to GMP standards that focus on the safety and effectiveness of therapeutic products. The evolution of the industry has been significantly shaped by the integration of advanced technologies, particularly in the realm of quality assurance for biologies, given their complexity as compared to chemical drugs. Epigenetic profiling and the predictive capabilities of artificial intelligence (Al) and machine learning (ML) are at the forefront of this transformation.
[0006] The biopharmaceutical sector's reliance on mammalian cell lines, particularly Chinese Hamster Ovary (CHO) cells, is rooted in their advanced capability for applying post-translational modifications to biologies, which is critical for the correct activity of therapeutic proteins. The optimisation of cell culture conditions, including nutrition strategies, is pivotal in achieving the desired yield, quality, purity, and functional properties of biologies. This optimisation is inherently linked to the comprehension and modulation of the cellular milieu, where epigenetic mechanisms such as DNA methylation have a profound impact. Wippermann et al. (2014) developed a CHO- specific CpG island microarray to analyse genome-wide DNA methylation and studied the effects of butyrate on differential DNA methylation in CHO cultures. Their research showed that butyrate supplementation led to dynamic and reversible changes in DNA methylation, which were particularly evident in genes related to cell growth, productivity, and signalling pathways. Recent advancements in data science, particularly the application of Al and ML algorithms, have shown promise in analysing epigenetic markers to enhance quality assurance in biomanufacturing. These cutting-edge methods, as outlined by Holder et al. (2017), harness the capabilities of Active Learning (ACL) and Imbalanced Class Learning (ICL) to effectively select features from epigenetic datasets, addressing the imbalance problem inherent in genomic data. Incorporating deep learning methodologies, especially reinforcement learning algorithms, has revolutionised the automation of intricate genomic feature creation, which is essential for the classification tasks at the heart of predicting the quality of biologies. These cutting-edge Machine Learning (ML) strategies, as evidenced by Olivecrona et al. (2017) are pivotal in ensuring earlier adherence to Quality Assurance (QA) standards, thereby enhancing the overall efficiency of the biomanufacturing process and marking a significant shift in the landscape of predictive bio-quality analytics.
[0007] The research conducted by Crowgey et al. (2018) illustrates the application of epigenetic machine learning for the identification of biomarkers for conditions such as spastic cerebral palsy (CP), utilising DNA methylation patterns. This approach, which employs machine learning algorithms to analyse epigenetic data, can potentially improve the predictive accuracy of disease diagnostics.
[0008] Marx et al. (2022) provide an overview about the current understanding of epigenetic regulation in CHO cells and discuss its significance for shaping the cell’s phenotype.
[0009] EP2558591A1 discloses a method for selecting of a long-term producing CHO cell clone by determining of the methylation frequency of a CpG site within an expression cassette that expresses a protein of interest.
[0010] EP4165213A1 discloses a method of determining CpG methylation status in a DNA sample for diagnosing a disease, the method comprising: (a) subjecting the DNA sample to bisulfite conversion; (b) amplifying said DNA sample following said (a) to obtain an amplified DNA sample; (c) labeling CpG sites in said amplified DNA sample with a label to obtain a labeled DNA sample; (d) contacting said labeled DNA sample on an array comprising a plurality of probes for said DNA under conditions which allow specific hybridization between said plurality of probes and said DNA; and (e) detecting said hybridization, wherein an amount of said label is indicative of the CpG methylation status in said DNA sample. US11566291 B2 discloses a differential methylation level of CpG loci that are determinative of a biochemical reoccurrence of prostate cancer.
[0011] EP1589118A2 discloses methods and compositions for assessing CpG methylation and provides an unstructured nucleic acid (UNA) oligonucleotide that base pairs with CpG islands. The subject oligonucleotide may be present in an array and find use in methods for evaluating methylation of CpG islands in cells.
[0012] WO201 6097319A1 discloses methods for detecting CpG methylation in marker DNA for diagnosing cancer.
[0013] Ou et al (2021 ) disclose epigenetic changes that may impact survival and function of human regulatory T cells.
[0014] Kim, M et al (2011) disclose a mechanistic understanding of production instability in CHO cell lines expressing recombinant monoclonal antibodies, namely the molecular mechanisms underpinning qP instability over long-term sub-culture in CHO cell lines producing recombinant lgG1 and lgG2 monoclonal antibodies.
[0015] Patel, N. A.et al (2018) discloses antibody expression stability in CHO clonally derived cell lines and their subclones; namely the role of methylation in phenotypic and epigenetic heterogeneity.
[0016] Feichtinger, J et al (2016) discloses a comprehensive genome and epigenome characterization of CHO cells in response to evolutionary pressures and over time.
[0017] WO201 1128337A1 discloses a low-dosed dosage form for hormone replacement therapy (HRT). More particularly, it concerns a solid oral dosage form comprising about 0.5 mg estradiol and about 0.5 mg drospirenone, and at least one pharmaceutically acceptable excipient.
[0018] Wippermann, A et al (2013) discloses a CpG island microarray for analyses of genome-wide DNA methylation in Chinese hamster ovary cells
[0019] Carillo- Avila, JA et al (2022) describes quality control protocols of cell lines whose target molecule is DNA, and details the scope or purpose and their corresponding functionality.
[0020] Osterlehner, A et al (2011 ) discloses promoter methylation and transgene copy numbers predict unstable protein production in recombinant Chinese hamster ovary cell lines
[0021] There is a recognised need for advanced monitoring and control of epigenetic factors to refine bioprocesses. SUMMARY OF THE INVENTION
[0022] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Other features, details, utilities, and advantages of the claimed subject matter will be apparent from the following written detailed description, including those aspects illustrated in the accompanying drawings and defined in the appended claims.
[0023] It is the objective of the invention to provide a new strategy or method of determining and / or controlling the quality of a eukaryotic bioprocess and a product produced by such bioprocess. It is a particular objective to ensure the quality of a eukaryotic bioprocess and / or product, by determining and / or controlling the quality of a eukaryotic bioprocess.
[0024] The objective is solved by the claimed subject matter, and as further described herein.
[0025] The invention provides for a method for predicting the quality characteristics of a eukaryotic bioprocess or of a product that is produced by a eukaryotic bioprocess, comprising: a) determining for a eukaryotic bioprocess cell, a genomic CpG-methylation index for a number of CpG-methylation sites, at an index timepoint of that bioprocess; b) comparing the index to a predetermined reference CpG methylation course to determine at least one product quality attribute (QA) or process parameter (PP) that is predictive to the quality of the bioprocess or product for a duration including at least said index timepoint and a time period of the bioprocess following or preceding said timepoint; wherein the reference CpG methylation course comprises the CpG-methylation changes at said CpG-methylation sites over said duration, which has been determined to be indicative of said at least one QA or PP by comparing with a reference bioprocess of comparable type, based on an estimation or prediction model.
[0026] Specifically, said predicting of the quality characteristics comprises estimating, describing, adjusting, or controlling said quality characteristics.
[0027] According to a specific aspect, the invention provides for a method for predicting the quality characteristics of a eukaryotic bioprocess or of a product that is produced by a eukaryotic bioprocess, comprising: a) determining for a eukaryotic bioprocess cell, a genomic CpG-methylation index for a number of CpG-methylation sites within more than one genomic regulatory elements, at an index timepoint of that bioprocess; b) comparing the genomic CpG-methylation index to a predetermined reference CpG methylation course to predict at least one product quality attribute (QA) or process parameter (PP) for a duration of the bioprocess which includes at least said index timepoint and a time period of the bioprocess following or preceding said timepoint; wherein the reference CpG methylation course comprises the CpG-methylation changes at said CpG-methylation sites over said duration, which has been predetermined to be indicative of said at least one QA or PP by comparing with a reference bioprocess of comparable type, preferably wherein said at least one product quality attribute (QA) or process parameter (PP) are predicted based on an estimation or prediction model.
[0028] Preferably an estimation of prediction model is employed which estimates or predicts at least one QA or PP across multiple timepoints of the bioprocess.
[0029] Specifically, said estimation or prediction is performed by a statistical or informed machine learning model, or by any other mathematical (e.g., statistical), deep learning, Al or foundation models, or by hybrid models.
[0030] Specifically, the estimation or prediction model uses the CpG methylation course and data of a specific CpG sample to predict the QA or PP. It can utilize the CpG information of a training set in relation to the respective QA or PP, and can then use its estimation and prediction power to utilize a sample CpG index to calculate the respective QA and PP.
[0031] Specifically, the genomic CpG-methylation index comprises at least one CpG pattern or motif. CpG sites which have a certain methylation pattern are herein referred to as CpG pattern. The CpG sites can be categorized into motifs. Specifically, one or more CpG patterns and / or motifs can be chosen for the purpose of defining the genomic CpG-methylation index. Preferably, CpG patterns and / or motifs are used where CpG methylation changes over time and across comparable experiments, in particular where the change of methylation can be correlated to said at least one product quality attribute (QA) or process parameter (PP). Preferably, selected CpG patterns and / or motifs can be used for the machine leaning (ML) models establishment.
[0032] Specifically, the reference CpG methylation course has been predetermined to be indicative of said at least one QA or PP by comparing with a reference bioprocess of comparable type, such as carried out in a prior experiment, or is predetermined by utilizing a general model generated from a comparable cell line, species or process.
[0033] Particularly, based on a set of candidate PPs and target QAs, an in silica model can be developed to predict the outcome of applying said PPs to a bioprocess on the QAs, before the bioprocess is physically performed.
[0034] Specifically, said genomic CpG-methylation index (herein also referred to as CpG-methylation index, CpG index, or index) comprises or is composed of a set of CpG sites that is predictive of said QA or PP, or directly links to the respective QA or PP.
[0035] Specifically, the CpG methylation course is composed of the change of methylation for a set of CpG sites, which change is predictive of said QA or PP (or directly links to the respective QA or PP) during the bioprocess, in particular at said index timepoint and a time period of the bioprocess following or preceding said timepoint.
[0036] Specifically, a CpG site that directly links to the respective QA or PP can be inferred from a larger group of CpG sites that directly influence said CpG site, therefore the CpG index used to estimate or predict said QA or PP is preferably for a number of CpG sites that indirectly affect said QA or PP via their effect on the directly acting CpG.
[0037] Specifically, the CpG index is determined by obtaining a sample of the bioprocess (or a fraction of the bioprocess cell or respective cell culture), herein referred to as sampling, and analysing the sample for CpG methylation.
[0038] Specifically, the CpG index determined at one timepoint during the bioprocess ( / .e., a sampling timepoint) estimates the QAs or PPs at the sampling timepoint.
[0039] Specifically, the CpG index determined at one sampling timepoint predicts the PPs and QAs during said bioprocess going forward, in particular during a time period following the sampling timepoint.
[0040] Specifically, the CpG index determined at one sampling timepoint predicts the PPs and QAs during said bioprocess retrospectively, in particular during a time period before the sampling timepoint.
[0041] Specifically, the CpG index is determined at more than one sampling timepoints during the bioprocess, e.g., by sampling and analysing the samples for CpG methylation, thereby obtaining the CpG index at a number of index timepoints, and optionally a respective CpG-methylation course. Specifically, such CpG-methylation course that is composed of more than one CpG indices can be compared to the reference CpG methylation course. Specifically, the number of sampling timepoints is at least 1 , 2, 3, 4, 5, 6, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, or 120, or up to 120.
[0042] Specifically, the timepoints are measured in time units (e.g., hours, or days) or in a cell-relevant unit (e.g., divisions).
[0043] According to a specific aspect, the index timepoint is at a time during an adaptation phase, seed culture, growth phase, production phase, or maturation phase of the bioprocess.
[0044] Specifically, the CpG methylation index and methylation course is determined at a predefined timepoint.
[0045] According to a specific aspect, the estimation or prediction model as used in the method described herein estimates or predicts at least one QA or PP across multiple timepoints of the bioprocess.
[0046] Specifically, the estimation model provides for estimation of the respective QA or PP based on a sample at one timepoint e.g., an index timepoint.
[0047] Specifically, the prediction model provides for calculating or predicting the respective QA or PP, back and forth in time during the bioprocess e.g., before and / or after an index timepoint. For example, based on a sample in the seed train or in the production process, the further bioprocess (or manufacturing run) can be predicted.
[0048] Specifically, the CpG-methylation index is dynamic, because it changes during the bioprocess. The estimation or prediction model as used herein can predict the changes during the bioprocess cased on determining the index at only one index timepoint, and further predict the change of the respective QA or PP during the bioprocess.
[0049] Specifically, said estimation or prediction is performed by a statistical or informed machine learning model, or by any other mathematical (e.g., statistical), deep learning, Al or foundation model, or by a hybrid model, such as using combination of the foregoing models.
[0050] Specifically, said estimation or prediction model is configured to model at least one of: i) direct effects of CpG methylation on QA; ii) direct effects of PP on CpG methylation; iii) direct effects between different CpG methylation sites; iv) direct effects between any of i), ii), iii), or iv). According to a specific aspect, estimation or prediction model is a mathematical, deep learning, Al, or foundation model, or a hybrid model, preferably a statistical or informed machine learning model.
[0051] Specifically, said estimation or prediction model is a neural network model or a Bayesian model or Foundation model.
[0052] Specifically, said estimation or prediction model employs a deep neural network or convolutional neural network, regression model or a Foundation model, that uses the CpG methylation levels to obtain estimates or predictions for the associated QAs and PPs.
[0053] Specifically, said estimation or prediction model calculates the possible methylation range for each CpG-methylation site of said genomic CpG-methylation index using Bayesian probability analysis.
[0054] Specifically, said estimation or prediction model determines a theoretical maximum of the bioprocess based on optimal methylation states for the selected CpG sites; and comparing the current bioprocess performance to said theoretical maximum.
[0055] Specifically, the bioprocess can be controlled by predicting (or determining) deviations in the bioprocess from the theoretical maximum.
[0056] Specifically, the estimation or prediction model generates a trajectory of the respective QA or PP based on the methylation course of the selected CpG-methylation sites.
[0057] Specifically, the CpG-methylation sites are selected according to their impact on the respective QA or PP, as determined by differential methylation analysis.
[0058] Specifically, the CpG-methylation sites are selected by analysing the methylation state of one or more genomic regions, preferably within more than one regulatory region of the genome of the cell, and identifying by CpG-methylation trajectory analysis the CpG sites which are activated (i.e. , methylated) and deactivated (i.e. not methylated) in relation to the respective QA or PP.
[0059] Specifically, those CpG-methylation sites are selected, where the respective QA or PP is a function of the methylation state of said CpG-methylation sites.
[0060] Specifically, the respective QA or PP is a function of the methylation state of said CpG-methylation sites.
[0061] Specifically, said estimation or prediction model allows determining the trajectory for said at least one QA or PP based the genomic CpG-methylation index for a number of CpG-methylation sites at a single index timepoint, wherein the trajectory is constructed by said estimation or prediction model.
[0062] Specifically, a trajectory is determined by i) fitting methylation data to a mathematical curve; and ii) defining function parameters based on the curve shape; and iii) applying said function parameters to construct a methylation trajectory from a single timepoint measurement.
[0063] Specifically, for predicting the quality characteristics of a eukaryotic bioprocess or of a product that is produced by a eukaryotic bioprocess, as described herein, the genomic CpG-methylation index for selected CpG-methylation sites is determined at an index timepoint, and the CpG methylation changes at the selected CpG-methylation sites are predicted for timepoints following or preceding said index timepoint, using said trajectory function.
[0064] According to a specific aspect, said CpG-methylation index comprises a system of numbers obtainable by comparing the methylation levels at said CpG-methylation sites, which change according to each other, or by comparing the methylation levels to a fixed standard.
[0065] Specifically, the CpG index includes the results of testing the CpG methylation for one or more genomic sites of the bioprocess cells, in particular CpG-methylation sites (herein also referred to as CpG sites or simply CpGs).
[0066] Specifically, any genomic CpG site can be comprised in the CpG-index, thereby contributing to said CpG index, which allows estimating or predicting one or more QAs and / or PPs.
[0067] According to a specific aspect, the number of CpG-methylation sites comprises at least two CpG-methylation sites.
[0068] Specifically, the number of CpG-methylation sites, is at least 2, 3, 4, 5, 6, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, or 5000. For foundation models, a selection or all of the CpG methylation sites of the whole genome can be utilized.
[0069] Specifically, said CpG methylation sites are genomic CpG sites.
[0070] The genomic CpG sites are preferably comprised in chromosomal DNA and / or mitochondria DNA. According to a specific aspect, the number of CpG-methylation sites are selected from CpG sites within a genomic region which is less than 500 bp, or less than 250 bp region (herein referred to as CpG region).
[0071] Typically, CpGs within a CpG region react similarly (Affinito et al, 2020; Taryma- Lesniak et al, 2024). Therefore, the determination of CpG-methylation at only one CpG site within one CpG region can be extrapolated to any other CpG site within said CpG region.
[0072] The models can be fine-tuned to increase predictability, using a selection of CpG sites within the same genomic region or CpG island. Should we add a sentence, that several CpG site of a given region can be utilized to fine-tune the models?
[0073] According to a specific aspect, CpG-methylation sites are selected from CpG sites within a genomic region which is less than 500 bp, or less than 250 bp region (herein referred to as CpG region).
[0074] Specifically, CpG-methylation sites are selected from CpG sites within a genomic region of less than 500 bp, or less than 250 bp region, which is located within a genomic regulatory element.
[0075] Specifically, the CpG site or CpG region is located within a non-coding genomic region e.g., within a regulatory element, such as a promoter, enhancer, activator, repressor, silencer, ribosomal binding site, a transcriptional or translational start or stop sequence, insulator, transposon element or RNA regulator.
[0076] Specifically, said number of CpG-methylation sites are within one or more genomic regulatory elements, preferably selected from a promoter, enhancer, silencer, insulator, transposon element, or RNA regulator.
[0077] Specifically, said genomic regulatory elements are within more than one genomic regulatory elements selected from a promoter, enhancer, silencer, insulator, transposon element, or RNA regulator.
[0078] Specifically, said CpG methylation sites are within more than one genomic regulatory elements, in particular more than one different regulatory elements such as independently selected from promoters, enhancers, silencers, insulators, transposons elements, or RNA regulators.
[0079] Specifically, said CpG methylation sites are not within a non-genomic region, such as an exogenous transgene or an artificial vector, which has been incorporated into the cell. In particular, the CpG methylation sites selected for the purpose described herein are not within an exogenous transgene or a heterologous promoter. In particular, one or more CpG methylation sites selected for the purpose described herein are not within a CMV promoter, in particular a CMV promoter comprised in an artificial vector. In particular, one or more CpG methylation sites selected for the purpose described herein are not within a transgene, in particular transgene that is exogenous to the genome of the cell i.e., not naturally-occurring within the genome of the cell, such as not occurring within the genome of the wild-type cell.
[0080] Specifically, the CpG index includes the results of testing the quality or quantity of CpG methylation at the CpG site, or within the CpG region. The quality determination of CpG methylation typically refers to the methylation of selected CpG sites, such as e.g., to obtain a CpG methylation pattern or motif. The quantity determination of CpG methylation is typically determined for a specific CpG site or CpG sites and refers to the percentage of CpG methylation at the specific CpG site(s), as compared to all possible methylation at said CpG site(s) within a sample population of cells.
[0081] The present disclosure particularly refers to the following specific aspects.
[0082] Specifically, the bioprocess is a process of culturing the cell such as e.g., in a cell culture. Specifically, the process is a bioprocess to produce a biological product, such as a cell-expressed (e.g., an expression product) or cell-derived product.
[0083] Specifically, the QA is predictive to the quality of the bioprocess and / or the product. Specifically, the reference CpG methylation course has been predetermined upon comparing with a respective reference bioprocess, particularly a validated bioprocess, or one of robust and expected outcome.
[0084] Specifically, said bioprocess is process of culturing a bioprocess cell, which is a batch, repeated batch, fed-batch, perfusion, continuous or steady-state (e.g., chemostat) cell culture, or a process of culturing an organoid or a tissue.
[0085] Specifically, the bioprocess is a culture of tissue or tissue culture. The tissue culture may comprise a culture of a single cell, a population of cells, or a whole or part of an organ or organoid. Specifically, cells can be primary, quiescent, immortalized or a hybrid cell line.
[0086] Specifically, the cell culture is a culture of a clone, or developing cells. Cells can be cultured during certain developmental stages of the cells such as in a culture of evolving or maturating cells.
[0087] Specifically, the cell culture and respective reference culture can be of the same or different types, where type refers to cells from different strains (e.g., CHO-K1 and CHO-S) and cells from a different class (mammal and avian). Specifically, the reference cell culture is a culture of comparable cells, in particular isogenic cells.
[0088] Specifically, the cell culture is a culture of adherent cells, such as e.g., cells adhered to a solid surface, or non-adherent cells, such as e.g., cells in suspension.
[0089] Specifically, the bioprocess cell is a mammalian, animal, insect, avian, plant or yeast cell, preferably a CHO or human cell such as HEK, T-cells, or stem cells.
[0090] Specifically, the bioprocess cell is yeast (Saccharomyces, Pichia, Kluyveromyces, Hansenula and Yarrowia), avian, insect (Sf9, Sf21 , S2, High Five) or mammalian cell, or, preferably a rodent or human cell, preferably a cell used for production or bioassays selected from the group consisting of Chinese hamster ovary (CHO)-cell lines, mouse myeloma (NSO)-cell lines, HT1080, H9, HepG2, MCF7, MDBK Jurkat, MDCK, NIH3T3, PC12, BHK (baby hamster kidney cell), VERO, SP2 / 0, YB2 / 0, Y0, C127, L cell, COS, e.g., COS1 and COS7, QC1-3, HEK-293, VERO, PER.C6, HeLA, EBI, EB2, EB3, oncolytic or hybridoma-cell lines, Sf9, Sf21 , BY4741 , BY4742, W303, S288C, GS115, KM71 , X-33, SMD1168H, NSO, KM71 H, High Five, Tn5, S2, Hi-5 BTI- Tn-4, NK; Se301 , embryo fibroblast cells, myoblasts, satellite cells, fibroblasts, adipocytes, endothelial cells, mesenchymal stem cells, pericytes, chondrocytes, osteoblasts, collagen-producing cells, neurons, oncolytic or hybridoma-cell lines, as well as patient-derived primary cells or tissue or animal-derived primary cells (used for biomass generation such as cultured meat or fortified foods) and stem cells (totipotent, pluripotent, multipotent, unipotent).
[0091] According to a specific aspect, the duration comprises at least one generation time of the eukaryotic cell, or the duration of the bioprocess, preferably at least up to the limit of in vitro cell age (LIVCA).
[0092] Specifically, the duration comprises at least any one of 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 223, 240, 250, or even more generation times e.g., up to the limit of in vitro cell age (LIVCA) of the bioprocess cells. The LIVCA is understood as the timepoint or generation up to which a bioprocess cell line can be safely used to obtain a bioprocess product of same or at least similar quality, and can e.g., be determined by gene copy number analysis, product expression rate or biomarker control analysis, or in-depth characterization of an expressed bioprocess product employing regular analytical analysis methods. Specifically, the CpG methylation index is estimative and predictive to the quality of the bioprocess for the whole lifecycle or lifetime of the bioprocess.
[0093] The invention further provides for the use of the method described herein for various purposes, such as, for: a) controlling the specific productivity qp of the bioprocess, preferably wherein the CpG-methylation index comprises CpG sites which regulate genes of protein formation folding, or transportation within the cells and excretion into the cell culture fluid, preferably wherein said genes encode ribosomes, mRNA expression, endoplasmic reticulum structures, Golgi apparatus structures, chaperones, or b) controlling the product quality, such as product degradation, preferably wherein the CpG-methylation index comprises CpG sites which regulate one or more intracellular protein degradation genes, preferably wherein said genes encode proteases compartmentalized in proteasomes or in degradative organelles, such as lysosomes; or c) controlling the bioprocess e.g., by process analytical technology (PAT) which ensures quality in biopharmaceutical manufacturing by monitoring and controlling processes in real-time, preferably during Good Manufacturing Production (GMP), upscaling or downscaling, technology transfer, cell culture media development, monoclonality assessment, parametric release, raw material release, cell culture media release, quality by design, growth performance or any other cell culture based test, cell line development, cell line monitoring, such as monitoring adaption, aging, or maturation, process development, clone selection, cell bank release or stability assessment, preferably wherein the bioprocess is controlled by setting the CpG-methylation index according to the predetermined reference CpG methylation course; d) comparing the quality of different bioprocesses, wherein a first bioprocess is compared with a second bioprocess being used as a comparator, preferably wherein the CpG-methylation index for said second bioprocess is compared with a reference CpG- methylation course that is predetermined from said first bioprocess; preferably wherein: i) a first bioprocess produces an originator product, and a second different one produces a comparator product preferably a biosimilar, or biobetter; ii) a first bioprocess is performed using a specific fermentation mode, and second different one uses a comparator fermentation mode; or iii) a first bioprocess is performed at one manufacturing site, and a second different one is performed at a comparator manufacturing site; or iv) a first bioprocess is using a cell culture parameter or medium, and a second different one uses a respective comparator cell culture parameter or medium; or v) a first bioprocess uses a scale of production, and a second different one uses a comparator scale of production; or vi) a first bioprocess is performed to produce the product employing a specific bioreactor or bioreactor system, and a second different one uses a comparator bioreactor or bioreactor system; e) for controlling, correcting or adjusting a QA or PP based on deviations of the CpG index from the predetermined reference CpG-methylation course.
[0094] Specifically, for controlling the product degradation, the CpG-methylation index may comprise one or more CpG sites which regulate the proteasome 20S subunit beta 10 (PsmblO) gene, in particular one or more CpG sites that are within the promoter of the PsmblO gene. An exemplary PsmblO gene is of Cricetulus griseus (Chinese hamster), NCBI Gene ID: 100751737, which is endogenous to the genome of the CHO cell that is used in the bioprocess, or the respective ortholog thereof that is endogenous to the genome of the cell as used in the bioprocess.
[0095] Likewise, for controlling the product degradation, the CpG-methylation index may comprise one or more CpG sites which regulate the (Ulk1) gene, in particular one or more CpG sites that are within the promoter of the Ulk1 gene. An exemplary Ulk1 gene is of Cricetulus griseus (Chinese hamster), Gene ID: NCBI Gene ID: 100772932, which is endogenous to the genome of the CHO cell that is used in the bioprocess, or the respective ortholog thereof that is endogenous to the genome of the cell as used in the bioprocess.
[0096] Specifically, for controlling the product degradation, the CpG-methylation index may comprise one or more CpG sites within the promoter of the PsmblO gene and within the promoter of the Ulk1 gene.
[0097] For the CpG sites comprised in the promoter of the PsmblO gene and / or Ulk1 gene, it has been determined that there is a dynamic activation and deactivation (e.g., a zig-zag pattern) in experiments yielding a degraded or truncated (undesired) product, while the same CpG sites remain inert in experiments with an intact (desired) product.
[0098] Specifically, for monitoring adaption of the cell to a certain cell culture medium or conditions, it can be ensured that the process of adaption is complete once the cells reach an expected CpG-methylation index or fingerprint. By such monitoring, the time to adapt a cell line can be shortened. According to a specific aspect, said QA is a quality attribute of a bioprocess product, such as a) an expression product, preferably a proteinaceous such as a protein, polypeptide, peptide or a self-assembly system such as a virus like particle, a virusbased or nucleic acid product, and / or b) a cellular product, preferably a product comprising or consisting of a whole cell, a population of cells, tissue, organ, or organoid, or a fraction of any one of the foregoing, and / or c) a food product e.g., a food product which comprises a cellular product and / or a protein, and / or d) an assembly of any of the foregoing.
[0099] Specifically, the nucleic acid product is a DNA or RNA product, such as a virus, bacteriophage or vector.
[0100] Specifically, a virus-based product is understood as a viral substance (e.g., a viral vector) such as used in gene-therapy, a vaccine comprising a viral substance (e.g., a viral vector, viral antigen, attenuated or inactivated virus), such as an oncolytic vaccine.
[0101] Specifically, said at least one QA comprises: i) for an expression product, at least one of protein stability, integrity, degradation, post-translational modification, purity, folding, isoform, activity, potency, absence or presence of contaminants; or ii) for a cellular product, at least one of stability, karyotype, viability, cell surface properties, maturation status, purity, activity, potency, strength, oncogene activation, physical or organoleptic parameters, absence or presence of one or more biomarkers or contaminants;
[0102] According to a specific aspect, said PP is a process parameter of a bioprocess, such as a) a parameter that is determined or measured in a culture of the bioprocess cell as used herein, or b) a parameter determined by the presence or absence of substances comprised in the bioprocess, such as one or more cell-derived or culture-derived substances e.g., cellular or bioprocess metabolites or cell culture media components bioprocess metabolites.
[0103] Specifically, said at least one PPs selected from physical, biological or chemical parameters. Specific physical parameters include e.g., measurable physical parameters or factors that influence a bioprocess. Specific biological parameters include e.g., measurable cell-related parameters or factors that influence a bioprocess. Specific chemical parameters include e.g., measurable substances that influence a bioprocess.
[0104] Specifically, said at least one PP comprises: i) a parameter determined in a culture of the bioprocess cell, preferably cell viability, viable cell density, product titre, bioprocess specific productivity, product yield, metabolite production, substrate consumption, growth, doubling time, pH, pressure, temperature, shear stress, shear resistance, mixing time, dissolved oxygen, pCO2 level, redox state, cell cycle; aggregate formation, cell size, cell surface characteristics, oncogene activation, reverse transcriptase and / or oncogene expression, host cell protein content, host cell DNA content, or plasmid DNA content (such as the plasmid DNA content for transient expression systems); or ii) a parameter determined by the presence or absence of substances comprised in the bioprocess.
[0105] Specific examples of such substances comprised in the bioprocess are one or more of the following: a) sugars and energy metabolites, such as glucose, fructose, galactose, mannose, maltose, fucose, raffinose, sucrose, xylose, pyruvate, ribose, lactate, ammonia, acetone, ethanol, putrescine, or citrate; b) amino acids, such as arginine, cysteine, cystine, histidine, isoleucine, glutamine, leucine, methionine, phenylalanine, threonine, tryptophan, tyrosine, valine, alanine, asparagine, aspartic acid, glutamic acid, glycine, proline, serine, ornithine, or taurine; c) vitamins or hormones, such as ascorbic acid, Vitamin A, Vitamin K, Vitamin E, Biotin, Vitamin D, niacin, thiamine, pyridoxal phosphate, cobalamin (B12), folic acid, choline, cyanocobalamin, nicotinamide, p-aminobenzoic acid, pantothenic acid, pyridoxal, pyridoxamine, pyridoxine, riboflavin, thiamine, retinol, NO2, niacinamide, nicotinamide, pantothenate, acetyl-Coa, or succinyl-CoA; d) organic acids, such as acetic acid, lactic acid, isovaleric acid, fumaric acid, formic acid, citric acid, 3-hydroxybutyric acid, maleic acid, propionic acid, pyruvic acid, succinic acid, isocitrate, alpha-ketoglutarate, succinate, fumarate, malate, oxalacetate, or pyroglutamine; e) elements, such as Al, B, Ba, Be, Ca, Cl, Co, Cr, Cu, Fe, K, I, Mg, Mn, Mo, Na, Ni, P, S, Se, Si, Sn, Sr, Ti, Tl, V, or Zn; f) lipids, such as cholesterol, fatty acids, phospholipids, ethanolamine, inositol, triglycerides, omega-3 fatty acids, or lipid carriers such as cyclodextrins, ethanol, polysorbate, polyoxyethylene-polyoxypropylen, or polyethylengycol; g) nucleic acid precursors, such as adenosine, hypoxanthine, thymidine, adenosine, cytidine, guanosine, uridine, cytosine, NAD, NADH, or NADPH; h) proteins, such as insulin, insulin-like growth factor, transferrin, endothelial growth factor (VEGF), cytokines, or granulocyte colony stimulating factor (G-CSF); i) support substances such as an anti-clumping agent (i.e. pluronic, dextran sulphate; PES, PEG, ferric citrate); j) peptones such as soy peptone, yeast extract, casein hydrolysate; k) dipeptides such as Ala-Gin, Gly-Tyr, Carnosine (Ala-His), Anserine, Cysteinyl- Histidine.
[0106] Specifically, any of said PPs can be controlled or modulated to adjust the quality or quantity of methylation e.g., the methylation level at one or more CpG sites.
[0107] Specifically, any of said PPs can be modulated to adjust one or more QAs.
[0108] Particularly, the PPs can be modulated if a CpG index change is predicted, that is, the CpG index predicts a deviation from the reference CpG methylation course.
[0109] Specifically, adjustment of said any PP may have a direct or indirect effect on the CpG index at one or more timepoints following a sampling timepoint.
[0110] Specifically, any of said QAs can be controlled or modulated to adjust the quality or quantity of the bioprocess product.
[0111] Controlling or modulating a QA may comprise setting or stabilizing a QA to a target value, preferably obtaining a QA that essentially unchanged (+ / - 10%), or slightly changed (+ / - 20%), or substantially changed (+ / - 50%), at one or more timepoints following a sampling timepoint, preferably over a duration of the bioprocess.
[0112] Controlling or modulating a QA may comprise increasing or decreasing a QA, in particular a value of a QA, preferably obtaining a QA that is higher or lower than a measured value, or to increase or decrease a measured value during a bioprocess, such as e.g., to adjust the value based on the predicted value using the CpG index. The increase or decrease is preferably at least 1 .2-fold, or 1 .3-fold, or 1 .4-fold , or 1 .5-fold, or 1.6-fold, or 1.7-fold, or 1.8-fold, or 1.9-fold, or at least 2-fold, but could also be 10- or 100-fold, depending on the type of QA and assay.
[0113] Specific QAs of a bioprocess or product directly or indirectly influence the quality or quantity of the bioprocess product. Specific QAs are functional QAs or safety QAs.
[0114] Exemplary QAs are: i) Functional QA: protein stability, integrity, degradation, post-translational modification, strength, purity, folding, isoform, activity, potency cellular stability, physical and organoleptic parameters, maturation status, cell surface properties, viability, biomass. ii) Safety QA: absence or presence of adventitious agents, karyotype, post- translational modification, cellular stability, oncogene activation, maturation status, cell surface properties, impurity, folding, isoform.
[0115] Specifically, said QA which is a post-translational modification of a product comprises any one or more of glycation, glycosylation, phosphorylation, ubiquitination, S-nitrosylation, methylation, N-acetylation, lipidation, deamidation, sialylation, afucosylation, disulfide bonds, gamma carboxylation, charge variants, or free thiols.
[0116] According to a specific aspect, said at least one QA or PP determines the stability of bioprocess cells or is based on the stability of bioprocess cells, in particular a QA or PP which determines the stability of the bioprocess cells to produce the bioprocess product, if essentially unchanged (+ / - 10%), or slightly changed (+ / - 20%), or substantially changed (+ / - 50%) over said duration of the bioprocess.
[0117] Stability of a bioprocess or bioprocess cells is herein understood as the quality or quantity of the bioprocess product that is produced by the bioprocess or bioprocess cells during the bioprocess, which is essentially unchanged (+ / - 10%), or slightly changed (+ / - 20%), or substantially changed (+ / - 50%) over the duration of the bioprocess.
[0118] Stability of the bioprocess product is herein understood as maintaining a stability QA for a duration of a bioprocess during which the bioprocess produces a product of same or at least similar quality (e.g., essentially unchanged (+ / - 10%), or slightly changed (+ / - 20%) product quality).
[0119] Specifically, one or more of the QAs are stability QAs, which are herein understood to indicate the stability of the bioprocess and bioprocess product.
[0120] Specifically, a stability QA can be determined by growth or metabolic rate, biomarker control analysis, or in-depth characterization of the expressed bioprocess product such as correct post translational modifications, employing regular analytical analysis methods. Specifically, one or more of the QAs are product quality QAs, which are herein understood to indicate the quality of the bioprocess product to meet a predefined product specification.
[0121] Specifically, one or more of the PPs are quality PPs, which are herein understood to indicate the quality of the bioprocess to produce a bioprocess product to meet a predefined process specification.
[0122] Specifically, a quality QA can be determined by the absence of contaminating stability QAs.
[0123] Specifically, where the bioprocess product possibly contains a contaminating substance e.g., an undesired product of the bioprocess, such as a host cell protein, host cell DNA, or a bioburden indicator (e.g., originating from mycoplasma, or endogenous retrovirus), which have a negative impact on the product quality, the quality of the bioprocess or product can be determined by a low amount of said contaminating or otherwise undesired product of the bioprocess.
[0124] Specifically, a quality QA or quality PP can be determined by a high amount of the bioprocess product, or by the product yield, such as by a yield of at least 1 mg / L (e.g., for a protein of interest, POI), preferably at least 10 mg / L, preferably at least 100 mg / L, most preferred at least 1 g / L.
[0125] According to a specific aspect, the bioprocess cell is a producer cell.
[0126] There are suitable methods to determine the bioprocess cell’s specific productivity (pg / g cell dry mass (CDM) per hour) and / or volumetric productivity (pg / L per hour) for the bioprocess product. Productivity and its fold change can be determined e.g., on a small scale or large scale. For example, an increase in a recombinant protein production might be determined at small-scale by measuring the concentration in the culture medium by a respective immunoassay, such as an ELISA. It can also be determined quantitatively by the ForteBio Octet method, or by chromatography (HPLC).
[0127] Specific methods for determining the amount of POI production described herein, can refer to the specific production rate (qp) of the bioprocess product, and / or to a time integral of a viable cell concentration (IVC). Specifically, the method may include the combination of determining qp and IVC. Production or productivity, being defined as concentration of the product in the bioprocess or fraction thereof (e.g., the supernatant), is typically understood as a function of these two parameters (qp and IVC).
[0128] Specifically, the bioprocess product is heterologous or endogenous to the bioprocess cell. Specifically, the bioprocess product is a proteinaceous, nucleic acid, or cellular product, preferably wherein: a) the proteinaceous product is a protein, polypeptide, peptide or a self-assembly system such as a virus like particle; b) the nucleic acid product is a DNA or RNA product, such as a virus, bacteriophage or vector; c) the cellular product comprises or consists of a whole cell, a population of cells, tissue, organ, or organoid, or a fraction of any one of the foregoing.
[0129] Specifically, the bioprocess product can be used in the fields of Pharmaceutical and Biotechnology Industry, Human and animal diagnostic industry, Food and Beverage Industry, Diagnostic and Research Laboratories, Agricultural and Crop Sciences, Animal Feed and Nutrition, Cosmetics and Personal Care Industry, Environmental and Bioremediation Industry, Textile and Leather Industry, Biofuel and Renewable Energy Industry, Chemical Industry, Paper and Pulp Industry, Detergent and Cleaning Industry, Paints and Coatings Industry, Water Treatment Industry, Polymer and Plastics Industry, Mining and Metal Extraction Industry, Construction and Building Materials Industry, Electronics and Semiconductor Industry, Automotive and Transportation Industry, Aerospace and Defense Industry, Sports and Fitness Industry or Art and Cultural Preservation Industry.
[0130] Specifically, proteinaceous products can be naturally occurring or synthetic, heterologous or endogenous.
[0131] Specifically, the bioprocess product can be intracellular or is secreted from the bioprocess cell into the bioprocess cellular supernatant.
[0132] Specifically, the bioprocess product is preferably a mammalian derived or related protein such as a human protein or a protein comprising a human protein sequence, or a bacterial protein or bacterial derived protein.
[0133] In specific cases, the bioprocess product is a fusion or multimeric protein, specifically a dimer or tetramer e.g., an antigen-binding molecule that comprises at least one heavy chain and at least one light chain of an antibody or a self-assembly virus-like particle consisting of many proteins.
[0134] Specifically, the bioprocess product is a therapeutic or diagnostic protein or product, such as functioning in mammals e.g., a human therapeutic or diagnostic.
[0135] Specifically, the bioprocess product is an enzyme that carries out an intermediate reaction within the bioprocess such as cleavage or binding. Specifically, the bioprocess product is a food, or cosmetic product.
[0136] According to a specific aspect, the proteinaceous product includes antigenbinding proteins, protein antibiotics, toxin fusion proteins, host cell proteins, structural proteins, regulatory proteins, vaccine antigens, growth factors, blood clotting or coagulation factors, hormones, cytokines, enzymes, process enzymes, and metabolic enzymes.
[0137] Specifically, the antigen-binding protein is an antibody molecule.
[0138] Specifically, the antibody molecule is a monoclonal antibody.
[0139] Specifically, the antibody molecule is a full-length antibody, an antibody comprising one or more epitope binding fragments of a full-length antibody, or a bispecific or multi-specific antibody comprising one or more of said fragments.
[0140] Specifically, said one or more epitope binding fragments of a full-length antibody are Fab, Fab', F(ab')2, Fv, or scFv fragments, or single domain antibodies.
[0141] The present disclosure includes certain features and embodiments of the present invention, which apply to all aspects of the invention, as appropriate.
[0142] Specifically, the method described herein comprises extracting genomic DNA from the sample, chemically modifying the DNA to differentially detect methylated cytosines and the methylated over unmethylated DNA ratio is measured at multiple CpG sites in the genome.
[0143] In specific instances, the method comprises thousands of quantitative measurements per sample. Each measurement measures the extent of methylation at a particular genomic location (e.g., at a CpG) and can describe the state of the entire population or of individual cells.
[0144] Specifically, any of the CpG sites identified by a method described herein can be used to set up a CpG methylation course for use in any of the methods described herein.
[0145] Exemplary selections of genomic sites identified for certain QAs and PPs are further described in the Examples section.
[0146] The present invention is also based on applied machine learning and epigenetic profiling to optimise eukaryotic producer cell cultures, such as CHO cell cultures. Deep and reinforcement learning have can predict growth and productivity and be used to optimise the bioprocess PPs. Overall, integrating epigenetic analysis with data science was found to have a strong potential to advance bioprocess monitoring and control. A computer system is preferably employed that includes a novel and innovative cell culture quality prediction system. The quality prediction system may generate accurate predictions regarding the bioprocess and / or product quality.
[0147] In an embodiment, the quality prediction system may be configured to analyze large volumes of biological data. For example, the computer system may include or link to a database including, for example, tens of millions of data points. Depending on various factors, such as the data sources, the application, among others, the number of data points may vary.
[0148] The quality prediction system disclosed herein may be configured to recognize a selection of CpG indices or CpG methylation courses and the CpG index changes among a larger selection of CpG sites within the whole or part of the genome of the cell.
[0149] The data representation can be encoded with the biological data in a way that enables the expression of various structural and mechanistic relationships of the bioprocess cell’s CpG methylated genomic sites or CpG index associated with one or more process parameters of the bioprocess or quality attributes of the bioprocess or the product. Deep learning methods can then be applied to the data encoded to the data representation, potentially enabling the generation of analysis results that reflect the quality of the bioprocess or the product.
[0150] In an embodiment, the quality prediction system may be configured to enable the analysis of significant volumes of data that determine the varied and complex PPs of the bioprocess and QAs of the bioprocess or product, that are involved in performing accurate quality predictions.
[0151] Moreover, the quality prediction system may provide an efficient and scalable mechanism for enabling the analysis of biological data in part on the respective QA.
[0152] The methods described herein can be adapted for quality assurance in biopharmaceutical manufacturing, offering a means to predict the quality of bioprocesses and biologies early in the production process and ensure adherence to quality standards.
[0153] Specific embodiments refer to a method for controlling the quality of the bioprocess and the product by profiling the methylation of CpG sites and establishing an index, methylation course and index change. This technique can contribute to the stabilisation of bioprocess and product quality, the robustness of a bioprocess and the predictability, simulation and control of bioprocesses. Specific embodiments refer to the application of epigenetic profiling in conjunction with Al and ML, such as in the context of GMP standards, aiming to analyse current methods and potential advancements in biopharmaceutical quality assurance and quality control.
[0154] The methods described herein conveniently employ a computer program. A computer program as used herein comprises computer program code. When a computer program is executed on a computer, the computer program can execute the method described herein. A computer as used herein is particularly equipped with appropriate computer programs and an operating system for executing computer-implementable steps, such as the calculation of the correlation functions, or the determination of one or more characteristic parameters as used for the purpose described herein. Said one or more characteristic parameters may be a QA or PP as used herein.
[0155] Specifically, a computer program can be used for screening data of CpG methylation within the cellular genome or a certain genomic region, to determine correct or desired signals which yield one or more QA or PP and establish an index or course, reference or other, described herein and apply estimative or predictive models.
[0156] Specifically, said predictive modelling employs an estimative and / or predictive model which is a deep neural network or convolutional neural network or regression model, that uses the CpG methylation levels to obtain estimates or predictions for the associated QAs and PPs.
[0157] Particularly, the neural network could be used to predict gaps in the existing model and generate an accurate prediction even if the CpG site was not measured.
[0158] FIGURES
[0159] Figure 1. Estimated and Observed values for Integral of Viable Cell Density (IVCD, Fig. 1a), Viable Cell Density (VCD, Fig. 1ab) and Growth Rate (Fig. 1c) over the duration of the experiment (days 4 to 14). The dots represent the cumulative cell density, the current cell density and the daily changes in cell density, as measured in the lab. The lines represent the estimated values for the same QA based on the machine learning algorithm and the selected CpG sites. The gray shade represents the expected acceptable range where that QA’s values should lie according to the reference CpG index. The R-squared and RMSE (root-mean-square error) are given to indicate the accuracy of the measurement, where higher values of the former and smaller of the latter indicate higher prediction accuracy. An R-square higher than 0.9 reflects an almost perfect prediction and is indicative of full understanding of the CpG sites influencing that QA.
[0160] Figure 2. Estimated and Observed values for Cumulative Glucose Consumption (Fig. 2a), Specific Glucose Consumption (Fig. 2b) and Specific Lactate Production (Fig. 2c) over the duration of the experiment (days 4 to 14). The dots represent how much glucose was used by all cells in culture, how much glucose was used per cell and day and how much lactate was produced per cell and day, as measured in the lab. The lines represent the estimated values for the same QA based on the machine learning algorithm and the selected CpG sites. The gray shade represents the expected acceptable range where that QA’s values should lie according to the reference CpG index. The R-squared and RMSE (root-mean-square error) are given to indicate the accuracy of the measurement, where higher values of the former and smaller of the latter indicate higher prediction accuracy. An R-square higher than 0.9 reflects an almost perfect prediction and is indicative of full understanding of the CpG sites influencing that QA.
[0161] Figure 3. Estimated and Observed values for Specific Glutamine Consumption (Fig. 3a), Specific Glutamate Production (Fig. 3b) and Specific Ammonia Production (Fig. 3c) over the duration of the experiment (days 4 to 14). The dots represent how much glutamine was used per cell and day, how much glutamate was produced per cell and day and how much ammonia was produced per cell and day, as measured in the lab. The lines represent the estimated values for the same QA based on the machine learning algorithm and the selected CpG sites. The gray shade represents the expected acceptable range where that QA’s values should lie according to the reference CpG index. The R-squared and RMSE (root-mean-square error) are given to indicate the accuracy of the measurement, where higher values of the former and smaller of the latter indicate higher prediction accuracy. An R-square higher than 0.9 reflects an almost perfect prediction and is indicative of full understanding of the CpG sites influencing that QA.
[0162] Figure 4. Estimated and Observed values for Integral of Viable Cell Density (IVCD) over the duration of the experiment (days 4 to 14), when using CpG sites in the neighbouring region (Fig. 4a: 50 bp apart; Fig. 4b: 150 bp apart; Fig. 4c: 2,000 bp apart used as a negative control) of the CpG sites selected for the estimative CpG index in Figure 1. The dots represent cumulative cell density as measured in the lab. The lines represent the estimated values for the same QA based on the machine learning algorithm and the selected CpG sites. The gray shade represents the expected acceptable range where that QA’s values should lie according to the reference CpG index. The R-squared and RMSE (root-mean-square error) are given to indicate the accuracy of the measurement, where higher values of the former and smaller of the latter indicate higher prediction accuracy. An R-square higher than 0.9 reflects an almost perfect prediction and is indicative of full understanding of the CpG sites influencing that QA.
[0163] Figure 5. Predicted and Observed cell culture time point and cell culture medium. Confusion matrices describing the actual medium (Fig. 5a) or timepoint (Fig. 5b) used and the estimated one based on the classification model.
[0164] Figure 6. Predicted and Observed values for Cumulative Glucose Consumption over the duration of the experiment (days 4 to 14) by sampling at only one said timepoint (Fig. 6a through Fig. 6j are sampling on days 4 through 13, respectively). The dots represent how much glucose was used by all cells in culture. The lines represent the forward and backward predicted values for the same QA based on the machine learning algorithm and the selected CpG sites. The shade represents the expected acceptable range where that QA’s values should lie according to the reference CpG index. The R- squared and RMSE (root-mean-square error) are given to indicate the accuracy of the measurement, where higher values of the former and smaller of the latter indicate higher prediction accuracy. An R-square higher than 0.9 reflects an almost perfect prediction and is indicative of full understanding of the CpG sites influencing that QA.
[0165] Figure 7. Estimated and Observed values for Integral of Viable Cell Density (IVCD) over the duration of the experiment (days 4 to 14), when using the reference CpG index from Figure 1 on a bioprocess with different PPs (basal medium and glucose concentration) that yields a lower cell growth. The dots represent cumulative cell density as measured in the lab. The dotted line represents the estimated values for the same QA based on the machine learning algorithm and the selected CpG sites. The gray dashed line represents the reference experiment (to which the current experiment is being compared). The R-squared and RMSE (root-mean-square error) are given to indicate the accuracy of the measurement, where higher values of the former and smaller of the latter indicate higher prediction accuracy. An R-square higher than 0.9 reflects an almost perfect prediction and is indicative of full understanding of the CpG sites influencing that QA. Figure 8. Estimated and Observed values for Integral of Viable Cell Density (IVCD) of cultures of two different manufacturing scales (for a vaccine, and thus 10L is sufficient for the world markets) and lab environments: 250 mL shake flask scale (top) and 10 L bioreactor (bottom). Both show comparable accuracy, indicating robustness of models to changes in manufacturing size, inoculation cell density and controlled vs uncontrolled fermentation strategies. The 10L bioreactor was harvested on day 11 while the shake flask continued to be cultured to day 13. The dots represent cumulative cell density as measured in the lab. The line represents the estimated values for the same QA based on the machine learning algorithm and the selected CpG sites. The R-squared and RMSE (root-mean-square error) are given to indicate the accuracy of the measurement, where higher values of the former and smaller of the latter indicate higher prediction accuracy. An R-square higher than 0.9 reflects an almost perfect prediction and is indicative of full understanding of the CpG sites influencing that QA.
[0166] Figure 9. Estimated and Observed values for Cumulative Glucose Consumption of cultures of two different manufacturing scales and lab environments: 250 mL shake flask scale (top) and 10 L bioreactor (bottom). Both show comparable accuracy, indicating robustness of models to changes in manufacturing size, inoculation cell density and controlled vs uncontrolled fermentation strategies. The 10L bioreactor was harvested on day 11 while the shake flask continued to be cultured to day 13. The dots represent how much glucose was used by all cells in culture. The line represents the estimated values for the same QA based on the machine learning algorithm and the selected CpG sites. The R-squared and RMSE (root-mean-square error) are given to indicate the accuracy of the measurement, where higher values of the former and smaller of the latter indicate higher prediction accuracy. An R-square higher than 0.9 reflects an almost perfect prediction and is indicative of full understanding of the CpG sites influencing that QA.
[0167] Figure 10. Estimated and Observed values for Viable Cell Density (VCD) of two independent cultures from a batch fermentation mode shows the CpG-index is translatable to other cell culture modalities and can predict the lack of cell growth at the end of a batch culture as compared to a fed-batch. The batch cultures were harvested on day 6. The dots represent current cell density as measured in the lab. The line represents the estimated values for the same QA based on the machine learning algorithm and the selected CpG sites. The R-squared and RMSE (root-mean-square error) are given to indicate the accuracy of the measurement, where higher values of the former and smaller of the latter indicate higher prediction accuracy. An R-square higher than 0.9 reflects an almost perfect prediction and is indicative of full understanding of the CpG sites influencing that QA.
[0168] Figure 11. Methylation-based models can capture bioprocess deviation, without training of the deviation in advance (this means, that the model never saw nor was it trained to predict the deviation), a) A bioprocess deviation was simulated by decreasing the bioreactor temperature on day 6, which resulted in a lower product titer (as measured by ELISA). Dots represent product titer as measured in the lab (light gray: standard process, dark gray: deviated process) b) Before the deviation occurs, forward predictive modeling on day 4 predict that both the reference culture and the deviating culture will have the same yield at harvest. Lines represent the predicted product titer (light gray: standard process, dark gray: deviated process) c) Once the deviation occurs, daily estimation of the yield based on methylation data reflects a drop in product which mirrors the wet lab measurements. Lines represent the estimated titer values based on the machine learning algorithm and the selected CpG sites (light gray: standard process, dark gray: deviated process).
[0169] Figure 12. SDS-PAGE gel of the bioprocess product ran with different PPs. The degradation (highlighted in white) is only observed when one of the basal media was used and identified via differential methylation analysis to be caused by a protease.
[0170] Figure 13. Methylation analysis of key proteostasic promoters (PsmblO, Ulk1) identified by CpG trajectory analysis reveals their dynamic activation and deactivation (zig-zag pattern) only in experiments yielding a truncated (undesired) product (dashed line), while the same CpG sites remain inert in experiments with an intact (desired) product (solid line).
[0171] Figure 14. Comparison of observed (actual I wet-lab) fed-batch titer values (circles and triangles without lines) to predicted titer values based on methylation data from cell culture passaging (lines), which resembles cell line stability data.
[0172] The model accurately predicts that the titer of an End-of-Production (EOP, similar to LIVCA) cell bank after 13 days of fed-batch will be higher than that of an MCB cell bank. The model bases the prediction on samples obtained during the passaging used to generate the EOP cell bank, not on samples obtained during the actual fed-batch cultures and are thus predicted
[0173] Figure 15. The age of samples from different cell culture passages can be distinguished based on a subset of CpG probes, a) PCA shows that samples are arranged based on passage, from the earliest (P06, leftmost) to the latest (P34, rightmost). Only methylation data was used for clustering, b) Negative control that results in an unordered clustering. PCA including a random subset of CpG sites to demonstrate that the age clustering is due to this specific set of CpG sites and not a global aging of the epigenome.
[0174] Figure 16. Different cell banks can be detected based on their methylation fingerprint. Each datapoint is a different sample (unique timepoint of an experiment). MCB and RCB experiments cluster separately from the limit of invitro cell age (LIVCA) experiments. The samples collected while generating the EOP cell bank for the LIVCA experiment show progressive aging, from P6 (identical to RCB) to P34 (identical to EOP / LIVCA cell bank).
[0175] Figure 17. A cell line originally cultivated in CD-CHO was adapted to 4 different basal media. Monitoring of the beta value (methylation percentage) of selected CpG sites throughout the adaptation period (Passages 1 to 5) reveals that those sites do not change when the cell is already adapted to that medium (CD-CHO) but quickly (in one passage) change methylation status when exposed to a foreign medium. This difference in methylation of particular CpG sites is conserved during the entire adaptation process.
[0176] Figure 18. Comparison of the VCD models on independent cell lines. Two independent cell lines (Null CL and CLC) were harvested at different timepoints and their cell densities (horizontal pattern) normalized to the reference. Application of the estimative models reveals that the prediction (square pattern) of cell lines growth based on another cell line’s CpG Index (Example 1) is mostly accurate.
[0177] Figure 19. Comparison of the age segmentation on independent cell lines. 4 cell banks of different ages were clustered via PCA using the aging CpG sites determined in the original cell line (Example 9) and results in a perfect segregation from earliest (MCB) to oldest (EOP).
[0178] Figure 20. Schematic of the Bayesian neural network approach to develop an in silica model, where direct and indirect interactions between QAs, PPs and CpG sites are calculated, uncovering unknown relations and missing measurements, as well as defining those methylation patterns indicative of an optimal or suboptimal bioprocess.
[0179] Figure 21. Proof of principle of the Bayesian network, where based on the probability distribution of each CpG site in a given index, the maximum capacity or potential for the QA or PP associated with that index and bioprocess can be calculated. Figure 22. Illustrative example of how intrinsic variable raw material can be qualified to reduce bioprocess variability, and how this could be standardized using an adequate CpG index. Lower and upper limits (red) are the acceptance criteria for a particular product quality attribute. Using methylation based raw material qualification and release, each batch (black dot) could be more in control and no batches would be discarded (outside the red boundaries).
[0180] Figure 23. Estimated and Observed values for Viable Cell Density (Figure 23a) and Viability (Figure 23b) of human cells grown in two different conditions. Both show high accuracy, indicating the CpG index methodology can be translated to human cells and used to predict PPs. Wet lab data was measured at three timepoints (seeding, feeding and harvest), and VCD is normalized to fold growth compared to seeding.
[0181] The dots represent the number of cells in culture (Figure 23a) or their viability (Figure 23b). The line represents the estimated values for the same QA based on the machine learning algorithm and the selected CpG sites. The R-squared and RMSE (rootmean-square error) are given to indicate the accuracy of the measurement, where higher values of the former and smaller of the latter indicate higher prediction accuracy.
[0182] Figure 24. The Viable Cell Density model trained on Days 4 to 14 can be used to backwards predict timepoints not previously seen, such as Day 3. Cell densities at harvest during passaging (Day 3, normalized to first passage) as measured in the lab (checkered pattern) or back-calculated (horizontal line pattern).
[0183] Figure 25: Table 1 . CpG ranges associated to each of the bioprocess QAs or PPs. The methylation of any of the CpG sites within a given positional range can be combined into an index used for the estimation or prediction of the QAs or PPs. The table describes the QA or PP calculated by the CpG index (column 1), the genomic range within which the CpG sites can be used to calculate said index (columns 2 and 3, in the form of chromosomal scaffold plus the starting and end position from that scaffold) as well as the gene and pathway (when in a promoter region) controlled by the CpG sites in that region (columns 4 and 5). The grey banding separates the different QAs and PPs. Note that some CpG sites contribute to more than one CpG index.
[0184] DETAILED DESCRIPTION
[0185] Unless indicated or defined otherwise, all terms used herein have their usual meaning in the art, which will be clear to the skilled person. Reference is for example made to the standard handbooks, such as Sambrook et aL, 2012, Molecular Cloning: A Laboratory Manual, volumes 1-4, Cold Spring Harbor Press, NY); Lewin, "Genes IV", Oxford University Press, New York, (1990), and Janeway et aL, "Immunobiology" (5th Ed., or more recent editions), Garland Science, New York, 2001 , Ausubel et aL, Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, Md. (1989), Vega et aL, Gene Targeting, CRC Press, Ann Arbor Mich. (1995), and Vectors: A Survey of Molecular Cloning Vectors and Their Uses, Butterworths, Boston Mass. (1988).
[0186] The terms “comprise”, “contain”, “have” and “include” as used herein can be used synonymously and shall be understood as an open definition, allowing further members or parts or elements. “Consisting” is considered as a closest definition without further elements of the consisting definition feature. Thus “comprising” is broader and contains the “consisting” definition.
[0187] The term “about” or “around” as used herein refers to the same value or a value differing by + / -10% or + / -5% of the given value.
[0188] Specific terms as used throughout the specification have the following meaning. The present invention refers to a bioprocess and bioprocess cell.
[0189] The term “cell” as used herein with respect to a bioprocess cell shall refer to a single cell, a single cell clone, or a cell line of a producer cell.
[0190] A cell type may refer to cells of a certain species, strain or cell line. When comparing comparable cell types, such comparable cell types are typically of the same species, strain or cell line, or of a cell line of a completely different species, however, with common features as relevant to the bioprocess and respective bioprocess product.
[0191] Specifically, the bioprocess can be a culture of a bioprocess cell, in particular a cell culture or a culture of tissue cells, or tissue, including e.g., organoids or organs.
[0192] Specifically, the bioprocess cell as used herein is cultured in a culture of a clone or cell line of eukaryotic cells.
[0193] The term “cell line” as used herein refers to an established clone of a particular cell type that has acquired the ability to proliferate over a prolonged period of time. A cell line is typically used for expressing an endogenous or recombinant nucleic acid molecule or gene, or for producing nucleic acid products, cellular products, or products of a metabolic pathway to produce polypeptides or cell metabolites mediated by such polypeptides.
[0194] A “bioprocess cell” is herein also referred to as a “producer cell. Specific producer cells are producer cell lines which can be cultured in a cell culture, such cell lines are commonly understood as cell lines ready-to-use for cell culture in a bioreactor to obtain a product of a bioprocess.
[0195] The term “host cell” as used herein shall particularly apply to any cell, which is suitably used for recombination purposes, such as to produce a POI or a host cell metabolite.
[0196] Specifically, a producer cell culture is an ex vivo (or in vitro) cell culture. Recombinant host cells as described herein are artificial organisms and derivatives of native (wild type) host cells. It is well understood that the term “producer cell” does not include human beings. Specifically, the producer cells, methods and uses described herein, e.g., specifically referring to those comprising one or more genetic modifications, heterologous expression cassettes or artificial expression constructs, said transfected or transformed host cells and recombinant proteins, are non-naturally occurring, are “man-made” or synthetic, and are therefore not considered as a result of “law of nature”. Genetic modifications described herein may employ tools, methods and techniques known in the art, such as described by Sambrook et al., 2012, Molecular Cloning: A Laboratory Manual, volumes 1-4, Cold Spring Harbor Press, NY).
[0197] Specifically, a eukaryotic cell as used for the bioprocess described herein, is a producer cell, such as suitably used in a production process to produce a bioprocess product e.g., at industrial scale.
[0198] Suitable producer or host cells are commercially available, and engineered for the specific purpose of bioprocess production, for example, obtained from culture collections such as the DSMZ (Deutsche Sammlung von Mikroorganismen and Zellkulturen GmbH, Braunschweig, Germany) or the American Type Culture Collection (ATCC).
[0199] Specifically, the producer cell is of mammalian, plant, insect, avian, fungi or yeast cells. Specific examples of producer cells are of human or non-human animal origin.
[0200] Specifically, the producer cell is a vertebrate producer cell, such as derived from a mouse, rat, Chinese hamster ovary cell (CHO) cell, Syrian hamster, monkey, ape, dog, horse, ferret, or cat cell, an insect, avian or yeast, such as a methylotrophic yeast.
[0201] Specifically, the mammalian cell is a human, rodent or bovine cell, cell line or cell strain. Examples of specific mammalian cells suitable as host cells described herein are mouse myeloma (NSO)-cell lines, Chinese hamster ovary (CHO)-cell lines, HT1080, H9, HepG2, MCF7, MDBK Jurkat, MDCK, NIH3T3, PC12, BHK (baby hamster kidney cell), VERO, SP2 / 0, YB2 / 0, Y0, C127, L cell, COS, e.g., COS1 and COS7, QC1-3, HEK-293, VERO, PER.C6, HeLA, EBI, EB2, EB3, oncolytic or hybridoma-cell lines. Preferably the mammalian cells are CHO-cells or human cells, or respective cell lines.
[0202] Preferably, the cell is a CHO cell, or a HEK cell, or a T cell.
[0203] Examples of CHO cells include, but are not limited to, CHOK1 , CHOK1SV, Potelligent CHOK1SV, CHO GS knockout, CHOK1SV GS-KO, CHOS, CHO DG44, CHO DXB11 , CHOZN, or a CHO-derived cell.
[0204] Specifically, a cell line of a certain species such as CHO can be used. It has been determined that there are cross-cell line applications of the subject described herein. Where a model or correlation between CpG-methylation of certain CpG sites and their change during the bioprocess as a function of a QA or PP of the bioprocess has been established for a certain cell line (or strain), such model or correlation can be conveniently transferred to another cell line (or strain), in particular another cell line (or strain) of the same species. Specific CpG methylation indices are universal indices e.g., universal indices that are applicable across cell lines (or strains) of different species origin, such as different mammalian species.
[0205] It is well-known in the art that in mammals, DNA methylation predominantly takes place in the context of CpG dinucleotides (CpGs) and affects most CpG sites in the genome.
[0206] Eukaryotic cells also include avian cells, cell lines or cell strains, such as for example, EBx® cells, EB14, EB24, EB26, EB66, or EBvl3.
[0207] Specifically, the eukaryotic cell is an insect cell (e.g., Sf9, Mimic™ Sf9, Sf21 , High FiveTM (BT 1 -TN-5B1 -4), or BT 1 -Ea88 cells), an algae cell (e.g., of the genus Amphora, Bacillariophyceae, Dunaliella, Chlorella, Chlamydomonas, Cyanophyta (cyanobacteria), Nannochloropsis, Spirulina, or Ochromonas), or a plant cell (e.g., cells from monocotyledonous plants (e.g., maize, rice, wheat, or Setaria), or from a dicotyledonous plants (e.g., cassava, potato, soybean, tomato, tobacco, alfalfa, Physcomitrella patens or Arabidopsis).
[0208] Specifically, the producer host cell is a differentiated form of cells. The production host cell can be derived from a primary cell in culture.
[0209] Specifically, the producer cell is a producer cell line, such as an immortalized producer cell line.
[0210] Specifically, the producer cell is a cell line of a primary, including patient- and animal-derived, a quiescent, an immortalized or a hybrid cell line, or a cell mixture of the same type or different types, or a tissue including e.g., an organ or organoid. Specifically, the producer cell is a host cell engineered to express a protein of interest (POI), such as from a recombinant expression cassette expressing a gene of interest (GOI) that encodes the POI.
[0211] Specifically, the producer cell is a cell line of a recombinant host cell.
[0212] Specifically, the producer cell is a cell line of mammalian host cells e.g., cell lines used as cell factories such as for example human primary cells, telomerase immortalized cell lines, or cell lines immortalized by viral oncogenes including adenoviral E1A, HPV derived E6, EBV derived oncogenes, SV40, or combinations of transcription factors. Specific cell lines include telomerase immortalized endothelial or mesenchymal stem cells, HEK293, CHO, Vero, HEK, or CAP.
[0213] Specifically, the producer cell is selected from the group consisting of normal or immortalized human cells, such as induced pluripotent or adult stem cells, epithelial cells and cancer cells.
[0214] Specifically, the producer cell is a mammalian stem cell or dendritic cell, preferably of human origin.
[0215] Specifically, the producer cell is a stem cell e.g., a totipotent, pluripotent, multipotent, unipotent cell.
[0216] Specifically, the stem cell is a mesenchymal stem cell (MSC), amniotic stem cell, or induced pluripotent stem (iPS) cell, a dendritic cell, a hematopoietic cell, epithelial cell, endothelial cell, nerve cell, blood cell or immune cell. Specifically, the source cells of EVs can be amnion-derived multipotent progenitor cell, chorion derived mesenchymal stem cell, induced pluripotent stem cell, keratinocyte, fibroblast, embryonic stem cell, ectodermal stromal cell, endodermal stromal cell, olfactory ensheathing cell, dental pulp stem cell, or immortalized mesenchymal stem cell.
[0217] Cells suitably employed in large-scale cell culture production include mesenchymal stem cells, dendritic cells, and HEK cells or 293T cells.
[0218] Specifically, the producer cell culture is a culture of a mammalian body fluid or tissue, preferably of blood, urine, amniotic fluid, ascites, cerebrospinal fluid, saliva, synovial fluid, or bone marrow.
[0219] Specifically, the producer cell is of a tissue culture, such as of body tissue, or originating from body fluid.
[0220] Specific cell culture processes may employ tissue culture. The culture of pluripotent stem cells, which include embryonic stem cells, embryonic germ cells and induced pluripotent cells, can contribute to tissues of a prenatal, postnatal or adult organism.
[0221] The subject cells may be from any mammal, including humans, primates, domestic and farm animals, and zoo, laboratory or pet animals, such as dogs, cats, cattle, horses, sheep, pigs, goats, rabbits, rats, mice etc. They may be established cell lines or they may be primary cells, where “primary cells”, “primary cell lines", and “primary cultures” are used interchangeably herein to refer to cells and cells cultures that have been derived from a subject and allowed to grow in vitro for a limited number of passages.
[0222] The subject cells may be isolated from fresh or frozen cells, which may be from a neonate, a juvenile or an adult, and from tissues including skin, muscle, bone marrow, peripheral blood, umbilical cord blood, spleen, liver, pancreas, lung, intestine, stomach, adipose, and other differentiated tissues. The tissue may be obtained by biopsy or apheresis from a live donor, or obtained from a dead or dying donor within about 48 hours of death, or freshly frozen tissue, tissue frozen within about 12 hours of death and maintained at below about -20°C, usually at about liquid nitrogen temperature (-190 °C) indefinitely. For isolation of cells from tissue, an appropriate solution may be used for dispersion or suspension. Such solution will generally be a balanced salt solution, e.g. normal saline, PBS, Hank’s balanced salt solution, etc., conveniently supplemented with fetal calf serum or other naturally occurring factors, in conjunction with an acceptable buffer at low concentration, generally from 5-25 mM. Convenient buffers include HEPES, phosphate buffers, lactate buffers, etc.
[0223] As described herein, the bioprocess may use bioprocess cells of tissue, in particular of organs, or organoids.
[0224] Specifically, the tissue is of an organ, such as a kidney, brain, or of placenta. Specifically, the tissue is a tumor or metastasis tissue, or a benign tissue.
[0225] The tissue culture can be a culture of an organoid. Specifically, the tissue is of an organoid. The term "organoid" is used herein to mean a 3-dimensional growth of mammalian cells in culture that retains characteristics of the tissue in vivo, e.g. prolonged tissue expansion with proliferation, multilineage differentiation, recapitulation of cellular and tissue ultrastructure, etc. A primary organoid is an organoid that is cultured from an explant, i.e. a cultured explant. A secondary organoid is an organoid that is cultured from a subset of cells of a primary organoid, i.e. the primary organoid is fragmented, e.g. by mechanical or chemical means, and the fragments are replated and cultured. A tertiary organoid is an organoid that is cultured from a secondary organoid, etc.
[0226] The term “cell culture” or “culturing” or “cultivation” as used herein with respect to a bioprocess cell refers to the maintenance of cells in an artificial, e.g., an in vitro environment, under conditions favoring growth, differentiation or continued viability, in an active or quiescent state, of the cells, specifically in a controlled bioreactor according to methods known in the industry. It is to be understood that the term "cell culture" is a generic term and may be used to encompass the cultivation not only of individual cells, but also of tissues, organs or organoids.
[0227] The in vitro environment refers to the conditions in which cells are grown and maintained outside their natural setting, typically within a laboratory or controlled setting. The cellular bioprocess environment encompasses various factors that can significantly influence the growth, behaviour, and characteristics of cells being cultured. These factors include physical factors, like temperature, pH level, humidity, and the type of substrate or surface on which the cells are cultured. Maintaining optimal physical conditions is crucial for the survival and growth of cells as well as to process a product; nutrient availability, cells require specific nutrients such as amino acids, vitamins, sugars, and minerals for growth and proliferation. The culture medium used for cell growth provides these essential nutrients and is customized based on the cell type(s) being cultured; oxygen levels, cells need oxygen for their metabolic processes. The oxygen concentration in the cell culture environment must be controlled and maintained at levels appropriate for the specific cell type; sterility and contamination Control, the environment in which cells are cultured must be kept sterile to prevent contamination that could affect cell growth and compromise experimental or bioprocessing results; cellcell interactions, in some cases, cells need specific interactions with other cells or a three-dimensional structure to mimic their natural environment accurately. This might involve co-culturing different cell types or using scaffolds to create a more realistic environment. The cell culture environment is critical in influencing cellular behaviour, differentiation, gene expression, and other cellular functions. Creating and maintaining an environment that mimics the natural conditions required by the specific cell type being cultured or bioprocessed is essential for accurate research and bioprocess outcomes and applications in fields like medicine, biotechnology, and drug development but also for achieving, maintaining and processing of any bioprocess product. When culturing a cell culture using appropriate culture media, the cells are brought into contact with the media in a culture vessel or with substrate under conditions suitable to support culturing the cells in the cell culture. Standard cell culture media and techniques are well-known in the art.
[0228] Cell culture processes may employ batch culture, semi fed-batch, perfusion culture, continuous culture, and fed-batch culture. The production phase specifically follows a growth phase. Specifically, a bioreactor is used which is suitable for any such cell culture or respective fermentation processes.
[0229] Batch culture is a culture process by which a small amount of a seed culture solution is added to a medium and cells are grown without adding an additional medium or discharging a culture solution during culture. Continuous culture is a culture process by which a medium is continuously added and discharged during culture. The continuous culture also includes perfusion culture. Fed-batch culture, which is an intermediate between the batch culture and the continuous culture and also referred to as semi-batch culture, is a culture process by which a medium is continuously or sequentially added during culture but, unlike the continuous culture, a culture solution is not continuously discharged.
[0230] Specifically preferred is a fed-batch process which is based on feeding of a growth limiting nutrient substrate to a culture. The fed-batch strategy, including single fed-batch or repeated fed-batch fermentation, is typically used in bio-industrial processes to reach a high cell density in the bioreactor.
[0231] Each bioprocess can be defined by a set of “characteristics”, which particularly include the process parameters (PP), the quality attributes (QA) and one or more CpG indices, as described herein.
[0232] The cell cultures as described herein particularly employ techniques which provide for the production of a bioprocess product, such as to obtain the bioprocess product in the cell culture medium, which is separable from the cellular biomass, herein referred to as “cell culture supernatant”, or in a cell culture fraction comprising cellular biomass, which product can be purified to obtain the bioprocess product at a higher degree of purity.
[0233] A bioprocess product can be produced using a bioprocess cell (a producer cell) or a respective cell line described herein, by culturing in an appropriate medium, isolating the bioprocess product from the culture, and optionally purifying it by a suitable method. Methods for recovering and / or purifying a bioprocess product are well established in the art. Specifically, a physical or chemical or physical-chemical method is used. The physical or chemical or physical-chemical method can be a filtering method, a centrifugation method, an ultracentrifugation method, an extraction method, a lyophilization method, a precipitation method, a chromatography method or a combination of two or more of any such methods. Specifically, the chromatography method comprises one or more of size-exclusion chromatography (or gel filtration), ion exchange chromatography, e.g., anion or cation exchange chromatography, affinity chromatography, hydrophobic interaction chromatography, and / or multimodal chromatography.
[0234] Cell culture media provide the nutrients necessary to maintain and grow cells in a controlled, artificial and in vitro environment. Characteristics and compositions of the cell culture media vary depending on the particular cellular requirements. Important parameters include osmolality, pH, and nutrient formulations. Feeding of nutrients may be done in a continuous or discontinuous mode according to methods known in the art.
[0235] The cell line described herein can be cultured under suitable batch, fed-batch or continuous culture conditions. The culture may be performed in microtiter plates, shakeflasks, or a bioreactor, and optionally starting with a batch phase as the first step, followed by a fed-batch phase or a continuous culture phase as the second step.
[0236] Whereas a batch process is a cell culture mode in which all the nutrients necessary for culturing the cells are contained in the initial culture medium, without additional supply of further nutrients during fermentation, in a fed-batch or continuous process, after a batch phase, a feeding phase takes place in which one or more nutrients are supplied to the culture by feeding. Although in most processes the mode of feeding is critical and important, the producer cell culture and methods described herein are not restricted with regard to a certain mode of cell culture.
[0237] In certain embodiments, the cell culture process is a fed-batch process. Specifically, a producer cell is cultured in a growth phase and transitioned to a production phase in order to produce a desired bioprocess product.
[0238] Specifically, the cell culture method described herein comprises a growing phase and a production phase.
[0239] Specifically, the producer cell culture may comprise the steps: a) culturing the producer cell under growing conditions (growing phase, or “growth phase”); and a further step b) culturing the producer cell under growth-limiting conditions (production phase), during which the bioprocess product is produced.
[0240] Specifically, the second step b) follows the first step a).
[0241] Specifically, the batch phase is performed until a basal carbon source that is initially added to the cell culture is consumed by the cell line. The dissolved oxygen (DO) spike method can be used to determine basal carbon source consumption during batch phase.
[0242] Specifically, the batch phase is performed for around 8 to 96h.
[0243] In a typical system of cell culture and bioprocess product production, wherein a batch phase is followed by a fed-batch phase, specifically, the cultivation in the fed-batch phase is performed for around 8 to 5,000h.
[0244] In another embodiment, producer cells described herein are cultured in a continuous mode e.g., employing a chemostat. A continuous fermentation process is characterized by a defined, constant and continuous rate of feeding of fresh culture medium into a bioreactor, whereby culture broth is at the same time removed from the bioreactor at the same defined, constant and continuous removal rate. By keeping culture medium, feeding rate and removal rate at the same constant level, the cell culture parameters and conditions in the bioreactor remain constant.
[0245] In another embodiment, producer cells are cultured in a perfusion mode e.g., culturing cells within a device while supplying fresh medium and removing the supernatant.
[0246] Specifically, the producer cell as used herein is suitable for a cell culture in a bioreactor or is capable of being cultured or grown in a bioreactor.
[0247] A stable cell culture as described herein is specifically understood to refer to a cell culture maintaining the genetic properties, specifically keeping the bioprocess product production level high e.g., at least at a pg level, even after about 20 generations of cultivation, preferably at least 30 generations, more preferably at least 40 generations, most preferred of at least 50 generations. Specifically, a stable production cell line is provided which is considered a great advantage when used for industrial scale production.
[0248] The devices, facilities and methods used for the purpose described herein are specifically suitable for use in and with culturing any desired cell line. Further, the devices, facilities and methods are suitable for culturing any eukaryotic host cell type and are particularly suitable for production operations configured for production of pharmaceutical and biopharmaceutical products, such as polypeptide or protein products (POI), nucleic acid products (for example DNA or RNA), or cells and / or viruses such as those used in cellular and / or viral therapies. Unless stated otherwise herein, the devices, facilities, and methods can include any desired volume or production capacity including but not limited to bench-scale, pilot-scale, and full production scale capacities.
[0249] Moreover, the devices, facilities, and methods can include any suitable reactor(s) including but not limited to stirred tank, airlift, fiber, microfiber, hollow fiber, ceramic matrix, fluidized bed, fixed bed, and / or spouted bed bioreactors. As used herein, “reactor” can include a fermenter or fermentation unit, or any other reaction vessel and the term “reactor” is used interchangeably with “fermenter.”
[0250] In embodiments and unless stated otherwise herein, the devices, facilities, and methods described herein can also include any suitable unit operation and / or equipment not otherwise mentioned, such as operations and / or equipment for separation, purification, and isolation of such products. Any suitable facility and environment can be used, such as traditional stick-built facilities, modular, mobile and temporary facilities, or any other suitable construction, facility, and / or layout. For example, in some embodiments modular clean rooms can be used. Additionally, and unless otherwise stated, the devices, systems, and methods described herein can be housed and / or performed in a single location or facility or alternatively be housed and / or performed at separate or multiple locations and / or facilities.
[0251] The industrial process scale would preferably employ volumes of at least 10 L, specifically at least 50 L, preferably at least 1 m3, preferably at least 10 m3, most preferably at least 100 m3.
[0252] Production conditions in industrial scale are preferred, which refer to e.g., fed batch culture in reactor volumes of 100 L to 10 m3or larger, employing typical process times of several days, or continuous processes in fermenter volumes of approximately 50 - 1000 L or larger, with dilution rates of approximately 0.001 - 0.15 IT1.
[0253] The cell culture described herein is particularly advantageous for use in a method of bioprocess product production on an industrial manufacturing scale e.g., with respect to both the volume and the technical system, in combination with a cultivation mode that is based on feeding of nutrients, in particular a fed-batch or batch process, or a continuous or semi-continuous process (e.g., chemostat).
[0254] The bioprocess cell described herein is typically tested for its capacity to produce the bioprocess product e.g., to express a GOI for POI production, and tested for the bioprocess product yield by a suitable assay. For testing a POI yield, any of the following tests can be used: ELISA, activity assay, capillary electrophoresis, HPLC, or other suitable tests, such as SDS-PAGE and Western Blotting techniques, or mass spectrometry. Alternatively, when the product is the cell, techniques to quantify the biomass, the activity and assess cellular markers are employed.
[0255] As described herein, an estimation or prediction model may employ correlation analysis or estimation analysis.
[0256] The term “correlation analysis” or “estimation analysis” as used herein shall refer to an analysis configured to generate correlation functions between measured characteristic parameters of a system (such as of a CpG methylation at one or more genomic sites, a CpG pattern or motif, or the changes of any one of the foregoing within a system (such as a producer cell culture) over a certain time period (such as the duration of the bioprocess), and one or more correlated values of at least one attribute of the system (such as one or more QAs that is predictive to the quality of the product, or one or more PPs that is predictive to the bioprocess and / or reactive to changes during the bioprocess).
[0257] Particularly, an “estimation model” is herein understood as a computational framework specifically designed to infer current values of a bioprocess, such as process parameters or quality attributes, from existing methylation and bioprocess data, particularly the CpG index. This model utilises advanced algorithms to analyse and correlate epigenetic data with bioprocess variables, offering real-time insights. Its primary function is to provide immediate understanding and assessment of the current bioprocessing conditions, facilitating accurate and timely decision-making and controls in biomanufacturing.
[0258] The term “predictive modelling” or “prediction model” as used herein shall refer to a computer implemented model that encompasses a variety of statistical techniques from data mining, predictive modeling and machine learning, that analyze current and historical facts to make predictions about future or other unknown events. A computational framework can be utilized to interpret historical bioprocess data as well as quality attributes and to predict future states or outcomes in bioprocessing and product quality attributes by leveraging both past and current data. This model typically employs advanced algorithms to analyse trends, CpG patterns or CpG motifs in epigenetic data, such as CpG methylation levels and other bioprocess parameters and quality attributes. Its purpose is to predict future and describes past changes in the bioprocessing environment, such as fluctuations in nutrient levels or cellular responses. By predicting these future and past conditions, the model aids in proactive decisionmaking and planning in biomanufacturing, enhancing the efficiency and effectiveness bioprocesses.
[0259] The term “predicting” with respect to a condition, attribute or parameter as used herein shall always include “estimating”, “measuring” and “determining”. The term “predicting” specifically includes forward prediction which allows assessing and predicting the outcome of a process for a defined period, such as e.g., for at least one or more days, such as at least 10, or 20 days, till the end of the bioprocess or the harvest time. Specifically, the system described herein can predict one or more QAs and / or PPs in a standard or changing environment
[0260] Using predictive modelling, a reference genomic CpG methylation index or reference CpG-methylation course (over time) can be determined for a certain bioprocess, which reference indicates one or more certain QAs or PPs.
[0261] Such predictive models can foresee Out of Specification events, where the bioprocess is outside the proven, defined and expected ranges between which it should lie or its product quality should be, which would result in an unknown, potentially unsafe outcome.
[0262] Systems and methods described herein generally use historical measurements and machine learning to build a dynamic model of a cell culture process. For example, historical data, e.g., certain concentrations of cell culture components from real-world cell culture processes, may be used to train a dynamic, data-driven predictive model. When applied to a real-world process, the model can predict future or past bioprocess QAs and PPs based on current (e.g., real-time) measurements of the cell culture. The model also makes use of measurements taken in one or more earlier time intervals (e.g., by using data of the current day and also on one or more days prior to the current day). The model may be a neural network or a regression model. The predictions output by the model may be input to a model-predictive controller, which can be used to stabilized, improve or maximize a desired QA or PP. The controller can then take the appropriate control action or actions (e.g., adjust a PP), or to manage and control the cell culture in a manner that guides the cell culture process to the desired objective. The methods taught herein also have broader applications and can be utilized to evaluate and control the effects of certain PPs on the QAs of the bioprocess, predict genotypic responses to certain PPs, and / or to predict other types of cell culture or cell phenotype development, including in silica modelling.
[0263] An informed machine learning model can be used for correlation analysis and / or predictive modeling. For example, one or more supervised or semi-supervised machine learning methods can be used, such as selected from neural network methods, decision trees, k-nearest neighbors method, carrier vector machines, algorithm based on a linear model, a generalized linear model discriminant, a factor regression model, a partial least square model, a factor analysis, a support vector machine, a support vector regression, a graphical model, a tree -based model, a random forest model, a random ferns model, a naive Bayes model, a linear discriminant analysis, a quadratic linear discriminant analysis, a perceptron model, a neural network model, nearest neighbor model, a nearest prototype model, an ensemble model, a prototype -based supervised algorithm, a bagged model, a Bayesian model, a regularized linear model, a polynomial model, a rule -based model, a Gaussian process model, a mixture discriminant model, a regression spline model, a rule induction method, a prototype model, a quantile regression model, a relevance vector machine, a soft independent modelling of class analogies model, a principal component-based model, an independent-based model, a self-organizing map model.
[0264] Semi-supervised learning method combines one or more unsupervised learning steps (with no labeled training data) with one or more supervised learning steps (with only labeled training data). Supervised or semi-supervised machine learning methods may include only one machine learning method with associated distinct candidate parameter sets and I or distinct machine learning methods with respective one or more candidate parameter sets.
[0265] Specific aspects described herein refer to measured characteristic parameters of a system, such as a parameter of a CpG-methylation at one or more genomic sites, or a genomic CpG-profile, an epigenetic profile, or the changes of any one of the foregoing within the system. Specifically, the system is a producer cell culture as described herein.
[0266] As described herein, correlation or estimation analysis can be used to screen, select, adjust and / or determine (or predetermine) the best reference CpG-methylation course for comparing with the data of genomic CpG-methylation of a bioprocess for which the method described herein is applied. For comparison, characteristic parameters can be correlated with calibration data such as obtained from a reference cell culture. The term “reference” as used herein with respect to a bioprocess or culture is used to establish the process model, meaning the optimized or validated version with a known and expected outcome, on which changed may be applied to predict and characterize the influence of the environment on the cells or the bioprocess or measurable biomarkers.
[0267] In contrast to the reference bioprocess, a “comparator” bioprocess is any bioprocess compared to the reference or the model.
[0268] The term “CpG-methylation” as used herein is understood as follows.
[0269] The term "methylation" as used herein refers to attachment of a methyl group to a base constituting DNA. Preferably, methylation refers to methylation occurring at cytosine of a specific CpG site of a specific gene. The result for a single cell is threefold, the site on each chromosome is either methylated or it is not; meaning both chromosomes are methylated, both are non-methylated, or one is and the other is not methylated. The result for many cells could be any methylation level between “0” and “1 ”; i.e., all cells are not methylated, or all cells are methylated or anything in-between.
[0270] For a population, the CpG methylation level instead refers to the proportion or percentage of a specific CpG site being methylated in a bioprocess. This metric’s values are continuous and range from 0 (no methylation) to 1 (complete methylation).
[0271] The "methylated status" as used herein refers to the presence (“methylationdetection”) or absence (“non-methylation-detection”) of 5-methyl cytosine of one or more CpG dinucleotides in a DNA base sequence.
[0272] The "methylation level" as used herein refers to, for example, the amount or extent of methylation, “a level”, of methylation present in the DNA base sequences of a DNA- methylated genomic site. The methylation level may be identified by a microarray. The microarray may be performed using a probe immobilized on a solid surface.
[0273] The term “genomic site” as used herein refers to any locus or region within the coding or non-coding genome of a cell, including e.g., the chromosomal or episomal genome, including the genomic elements that are endogenous or heterologous (or artificial) to the cell, such as e.g., expression cassettes or vectors.
[0274] The methylation status or level may be measured by any suitable method, such as by PCR, methylation-specific PCR, real-time methylation-specific PCR, MethyLight PCR, MethyLight digital PCR, EpiTYPER, PCR using a methylated DNA-specific binding protein, sanger sequencing, sequencing by hybridization, sequencing by synthesis, Singe Molecule Real Time sequencing, optical mapping, next generation sequencing, DNA shearing, modified electrophoresis, Enzymatic fragmentation, quantitative PCR, a DNA chip assay, electron microscopy DNA sequencing, nanopore sequencing, SeqStudio™, Genexus™, GeneScan™, 3500 Genetic Analyzer , AVITI™, Pacific Biosciences’ Sequel™ (PacBio), Bionano Genomics, Ion Torrent™’ Ion GeneStudio™ pyrosequencing and bisulfite sequencing. Determining the methylation status or level may comprise the use of a high-throughput methylation assay.
[0275] A “CpG” site refers to a CpG site present in DNA of the genome. A CpG site may be present in various regulatory elements. These are specific sequences within the DNA that control the activity of genes. These elements play a fundamental role in determining when and to what extent a gene is transcribed into messenger RNA (mRNA), which then guides protein synthesis. There are several types of DNA regulatory elements, including promoters, sequences located near the beginning of a gene where RNA polymerase and other transcription factors bind to initiate transcription. They determine the start site for transcription; enhancers, DNA sequences that can be located far from the gene they regulate. When specific proteins bind to enhancers, they can increase the rate of transcription of the associated gene; silencers, sequences that, when bound by certain proteins, decrease the rate of transcription of a gene; insulators, sequences help in organizing the genome by regulating the interactions between enhancers, silencers, and promoters. They act as boundaries, preventing the effects of regulatory elements from spreading to neighbouring genes; insulators, sequences that help in organizing the genome by regulating the interactions between enhancers, silencers, and promoters. They act as boundaries, preventing the effects of regulatory elements from spreading to neighbouring genes; transposons, DNA sequences that can move or "transpose" within a genome, often changing their positions. Methylation plays a role in silencing transposon activity. When transposons are heavily methylated, it can suppress their mobility and reduce the risk of their insertion into important genes, which could disrupt gene function; regulation by miRNA and / or IncRNA, MicroRNAs (miRNAs) are a class of small non-coding RNA molecules, typically around 22 nucleotides in length and long- non coding RNA (IncRNA) with a typical size up to 1000 nucleotides, that play a crucial role in post-transcriptional gene regulation. They function by binding to specific messenger RNA (mRNA) molecules, leading to their degradation or inhibiting their translation, thereby regulating the expression of target genes. These regulatory elements work in a highly coordinated manner, interacting with various proteins (transcription factors), other molecules, and DNA methylation to modulate gene expression. The combination and activity of these elements contribute to the precise and dynamic control of gene expression, allowing cells to respond to developmental cues, environmental signals, and internal physiological changes.
[0276] The CpG methylation status or level may be that of hypermethylation. Identifying one or more of the genomic sites having the respective characteristic methylation status or level comprises obtaining a sample of genomic DNA, and determining, by analyzing the genomic DNA using an assay, the methylation status or level of at least one CpG dinucleotide sequence within at a genomic site or region.
[0277] CpG methylation analysis for multiple genomic sites may result in the methylation status or levels that change according to each other, or change when comparing the methylation status or level to a fixed standard or predetermined reference value, thereby obtaining a CpG methylation pattern, signature or index, or a CpG motif.
[0278] CpG methylation analysis at different timepoints during a cell culture may result in the methylation status or levels at one or more CpG sites that change over time for the cell culture, wherein the change can be determined according to each other, or relative to each other, or can be determined when comparing the methylation status or levels to a fixed standard or predetermined reference value or CpG methylation course.
[0279] Some regions of genomic DNA have a frequency of CpG that is closer to that is expected by chance, and these sequences are known as CpG islands. As used herein, a "CpG island" refers to an area with CpG dinucleotides of clustering. Typically, the CpG island is understood as a sequence of DNA, of at least 200 bp, that has a GC content at least 50%. Specifically, a CpG island has an observed or expected CpG content ratio of at least 0.6 ( / .e., a CpG dinucleotide content of at least 60% of which would be expected by chance).
[0280] The methylation status or level of a CpG island can indicate that the CpG-island is methylation-free or methylated.
[0281] The term “index” in the context of “index of CpG-methylation” or “CpG-methylation index” as used herein, shall refer to a composite index comprising multiple variables that can be used as a quality assessment tool to predict QA of a cell culture over a prolonged period. An index of CpG-methylation is typically a combination of values of CpG- methylation at selected genomic sites, in particular CpG sites, into a single data model.
[0282] The index can be used as a classifier with classification power to predict the process or result of a cell culture at certain timepoints. The genomic sites useful in providing an index of CpG-methylation can be predetermined using machine learning algorithms such as support vector machine, naive bayes, neural network, etc. as known to those in the art, or discriminant analysis algorithms based on classical statistics.
[0283] As a rule, methylated CpG-sites are silenced whereas a lack of methylation is indicative of activity.
[0284] In one aspect, a CpG-methylation index, is indicative of the cell culture quality over a period of time. Specifically, the index is based on the relative changes in CpG- methylation at preselected genomic sites as compared to a reference.
[0285] Specifically, methylation of a CpG site can have an effect on the methylation status or level of other CpG sites within the same genomic region that is spanning e.g., 500 bp, or 250 bp. For example, a CpG site can make reference to any other CpG sites within a neighbouring 250 bp range in either direction, due to the high correlation between all sites in the region.
[0286] The term “timepoint” is herein particularly understood as the time of sampling thereby obtaining a sample that is used for determining CpG-methylation. A timepoint, for which a CpG index is determined, is herein also referred to as “index timepoint”. Typically, sampling is performed as a way of determining the variables comprised in the index, or the index itself. The index timepoint is typically during a specific phase of the cell culture which duration could be hours, days, generation, cell passages or any other unit over time.
[0287] The term “sampling” as used herein is particularly understood as the process of obtaining a sample, or any obtention of a material for subsequent analysis, be it cellular media for nutrition analysis, cellular product for quality analysis or cells for cell behaviour or methylome analysis.
[0288] The term "expression” or “expression cassette” is herein understood to refer to nucleic acid molecules (herein also referred to as polynucleotides), which contain a desired coding sequence (herein referred to as a gene), and control sequences in operable linkage, so that cells transformed or transfected with these molecules incorporate the respective sequences and are capable of producing the encoded proteins or other cellular products, such as cell metabolites. The term “expression” as used herein refers to either expression of a polynucleotide or gene, or to the expression of the respective polypeptide or protein.
[0289] One or more expression cassettes are herein also understood as “expression system”. The expression system may be included in an expression construct, such as a vector; however, the relevant DNA may also be integrated into a host cell chromosome. Expression may refer to secreted or non-secreted expression products, including polypeptides or metabolites.
[0290] Expression cassettes are conveniently provided as expression constructs e.g., in the form of “vectors” or “plasmids”, which are typically DNA sequences that are required for the transcription of cloned recombinant nucleotide sequences i.e., of recombinant genes and the translation of their mRNA in a suitable host organism. Expression vectors or plasmids usually comprise an origin for autonomous replication or a locus for genome integration in the host cells, selectable markers, a number of restriction enzyme cleavage sites, a suitable promoter sequence and a transcription terminator, which components are operably linked together. The terms “plasmid” and “vector” as used herein include autonomously replicating nucleotide sequences as well as genome integrating nucleotide sequences, such as artificial chromosomes e.g., a yeast artificial chromosome (YAC).
[0291] Expression vectors may include but are not limited to cloning vectors, modified cloning vectors and specifically designed plasmids. Preferred expression vectors described herein are expression vectors suitable for expressing a recombinant gene in a eukaryotic host cell and are selected depending on the host organism. Appropriate expression vectors typically comprise regulatory sequences suitable for expressing DNA encoding a POI in a eukaryotic host cell. Examples of regulatory sequences include promoters, operators, enhancers, ribosomal binding sites, and sequences that control transcription and translation initiation and termination. The regulatory sequences are typically operably linked to the DNA sequence to be expressed.
[0292] To allow expression of a nucleotide sequence in a producer cell, a promoter sequence is typically regulating and initiating transcription of the downstream nucleotide sequence, with which it is operably linked. An expression cassette or vector typically comprises a promoter nucleotide sequence which is adjacent to the 5’ end of a coding sequence, e.g., upstream from and adjacent to the coding sequence (e.g., encoding a helper factor) or gene of interest (GOI), or if a signal or leader sequence is used, upstream from and adjacent to said signal and leader sequence, respectively, to facilitate translation initiation and expression of coding sequences to obtain the expression product (e.g., the POI).
[0293] Specific expression constructs as used for the purpose described herein comprise a promoter operably linked to a nucleotide sequence encoding a POI under the transcriptional control of said promoter. Specifically, a promoter can be used which is not natively associated with said coding sequence, such as to allow expression from a recombinant nucleotide sequence.
[0294] According to a specific aspect of producing a bioprocess product from a recombinant host cell culture, a producer cell comprises a recombinant expression cassette that expresses a GOI, which expression cassette is heterologous to the host cell or artificial, thus not naturally-occurring in the respective wild-type host cell.
[0295] Specifically, the expression cassette comprises or consists of an artificial fusion of polynucleotides, including a promoter operably linked to the GOI, and optionally further sequences, such as a signal, leader, or a terminator sequence. Specifically, the expression cassette comprises or consists of an artificial fusion of a promoter, the GOI, and one or more additional regulatory sequences in operable linkage to allow expression of the GOI from said expression cassette for POI production.
[0296] Specifically, the expression cassette comprises one or more expression control sequences operably linked to said GOI, preferably said one or more expression control sequences comprise a promoter which is an inducible, de-repressible or otherwise regulatable promoter, or a constitutive promoter.
[0297] Specifically, within the expression cassette, an expression cassette promoter and GOI are heterologous to each other, thus, not occurring in such combination (or operable linkage) in nature e.g., wherein either one (or only one) of the promoter and GOI is artificial or heterologous to the other and / or to the host cell described herein; the promoter is an endogenous promoter and the GOI is a heterologous GOI; or the promoter is an artificial or heterologous promoter and the GOI is an endogenous GOI; wherein both, the promoter and GOI, are artificial, heterologous or from different origin, such as from a different species or type (strain) of cells compared to the host cell described herein. Specifically, the promoter of the expression cassette is not naturally associated with and / or not operably linked to said GOI in the cell which is used as a host cell described herein.
[0298] Specifically, the promoter can be regulatable, such as e.g., inducible or derepressible.
[0299] Specifically, the promoter can be a constitutive promoter.
[0300] Specifically, a GOI expression cassette is comprised in an autonomously replicating vector or plasmid, or integrated within a chromosome of said host cell.
[0301] The expression cassette may be introduced into the host cell and integrated into the host cell genome (or any of its chromosomes) as intrachromosomal element e.g., at a specific site of integration or randomly integrated, whereupon a high producer host cell line is selected. Alternatively, the expression cassette may be integrated within an extrachromosomal genetic element, such as a plasmid or an artificial chromosome e.g., a bacterial or yeast artificial chromosome (BAC or YAC). According to a specific example, the expression cassette is introduced into the host cell by a vector, in particular an expression vector, which is introduced into the host cell by a suitable transformation technique. For this purpose, the GOI may be ligated into an expression vector.
[0302] Techniques for transfecting or transforming host cells for introducing a vector or plasmid are well known in the art. Reference is for example made to the standard handbooks, such as Sambrook et al., 2012, Molecular Cloning: A Laboratory Manual, volumes 1-4, Cold Spring Harbor Press, NY), and in other molecular biology manuals. These can include electroporation, spheroplasting, lipid vesicle mediated uptake, heat shock mediated uptake, calcium phosphate mediated transfection (calcium phosphate / DNA co-precipitation), viral infection, and particularly using modified viruses such as, for example, modified adenoviruses, microinjection and electroporation.
[0303] The term "gene expression", or “expressing a polynucleotide” or “expressing a nucleic acid molecule” as used herein, is meant to encompass at least one step selected from the group consisting of DNA transcription into mRNA, mRNA translation and processing, mRNA maturation, mRNA export, protein folding and / or protein transport.
[0304] The term “endogenous” as used herein is meant to include those molecules and sequences, in particular endogenous genes or proteins, which are present in the wildtype (native) cell, prior to any genetic modification for producing a bioprocess product. In particular, an endogenous nucleic acid molecule (e.g., a gene) or protein that does occur in (and can be obtained from) a particular cell as it is found in nature, is understood to be “endogenous to host cell”. Moreover, a cell “endogenously expressing” a nucleic acid or protein expresses that nucleic acid or protein as does a cell of the same particular type as it is found in nature. Moreover, a cell “endogenously producing” or that “endogenously produces” a nucleic acid, protein, or other compound produces that nucleic acid, protein, or compound as does a host cell of the same particular type as it is found in nature.
[0305] According to specific aspects, the genomic site (in particular, the CpG site), where the CpG methylation is determined, is endogenous to the cell as used herein for the bioprocess. Specifically, CpG sites selected for determining an index, are within one or more endogenous regions of the genome of the cell. The term “heterologous” as used herein with respect to a nucleotide sequence, construct such as an expression cassette, amino acid sequence or protein, refers to a compound which is either foreign to a given host cell, i.e. “exogenous”, such as not found in nature in said host cell; or that is naturally found in a given host cell e.g., is “endogenous”, however, in the context of a heterologous construct or integrated in such heterologous construct e.g., employing a heterologous nucleic acid fused or in conjunction with an endogenous nucleic acid, thereby rendering the construct heterologous. The heterologous nucleotide sequence as found endogenously may also be produced in an unnatural e.g., greater than expected or greater than naturally found, amount in the cell. The heterologous nucleotide sequence, or a nucleic acid comprising the heterologous nucleotide sequence, possibly differs in sequence from the endogenous nucleotide sequence but encodes the same protein as found endogenously. Specifically, heterologous nucleotide sequences are those not found in the same relationship to a host cell in nature. Any recombinant or artificial nucleotide sequence is understood to be heterologous.
[0306] The term "operably linked" as used herein refers to the association of nucleotide sequences on a single nucleic acid molecule, e.g., a vector, or an expression cassette, in a way such that the function of one or more nucleotide sequences is affected by at least one other nucleotide sequence present on said nucleic acid molecule. By operably linking, a nucleic acid sequence is placed into a functional relationship with another nucleic acid sequence on the same nucleic acid molecule. For example, a promoter is operably linked with a coding sequence of a recombinant gene, when it is capable of effecting the expression of that coding sequence. As a further example, a nucleic acid encoding a signal peptide is operably linked to a nucleic acid sequence encoding a POI, when it is capable of expressing a protein in the secreted form, such as a preform of a mature protein or the mature protein. Specifically, such nucleic acids operably linked to each other may be immediately linked i.e., without further elements or nucleic acid sequences in between the nucleic acid encoding the signal peptide and the nucleic acid sequence encoding a POI. Alternatively, a suitable linking sequence can be used such as e.g., a cloning site positioned between the promoter and the GOL
[0307] The term “polynucleotide”, “nucleic acid molecule(s)” or “nucleic acid sequence(s)” as interchangeably used herein, refers to nucleotides, either ribonucleotides or deoxyribonucleotides or a combination of both, in a polymeric unbranched form of any length. Preferably, a polynucleotide refers to deoxyribonucleotides in a polymeric unbranched form of any length. Here, nucleotides consist of a pentose sugar (deoxyribose), a nitrogenous base (adenine, guanine, cytosine or thymine) and a phosphate group.
[0308] A “promoter” sequence is typically understood as a non-coding regulatory sequence which, when operably linked to a coding sequence, controls the transcription of the coding sequence. A promoter sequence may be natively associated with the coding sequence, such as in a native (wild-type) cell for endogenous protein expression.
[0309] A promoter is herein described to initiate, regulate, or otherwise mediate or control the expression of a protein coding polynucleotide (DNA), such as a POI coding DNA. Promoter DNA and coding DNA may be or be derived from the same gene or from different genes, and may be derived from the same or different organisms.
[0310] Either the promoter or the coding sequence, or both, can be heterologous to the cell. A promoter may or may not be natively associated with the coding sequence. Any one or both of the promoter and the coding sequence can be endogenous and are herein also understood to be not natively associated in a cell, if comprised in a heterologous expression cassette.
[0311] A heterologous promoter may be heterologous to the polynucleotide to be expressed and / or an artificial promoter, or a promoter that is originating from the wildtype host cell, but positioned in the host cell genome within a heterologous expression cassette or positioned at a location where it is not naturally occurring in the wild-type host cell.
[0312] The term "protein of interest (POI)" as used herein refers to a polypeptide or a protein that is produced by means of recombinant technology in a host cell. More specifically, the protein may either be a polypeptide not naturally-occurring in the host cell i.e., a heterologous protein, or else may be native to the host cell i.e., a homologous protein to the host cell, but is produced, for example, by transformation or transfection with a self-replicating vector containing the nucleic acid sequence encoding the POI, or upon integration by recombinant techniques of one or more copies of the nucleic acid sequence encoding the POI into the genome of the host cell, or by recombinant modification of one or more regulatory sequences controlling the expression of the gene encoding the POI, e.g., of the promoter sequence. In some cases, the term POI as used herein also refers to any metabolite product by the host cell as mediated by the recombinantly expressed protein. The term “isolated” or “isolation” as used herein with respect to a bioprocess product shall refer to such compound that has been sufficiently separated from the environment with which it would naturally be associated, in particular from a cell culture or cell culture fraction such as a cell culture supernatant, so as to exist in “purified” or “substantially pure” form. Yet, “isolated” does not necessarily mean the exclusion of artificial or synthetic mixtures with other compounds or materials, or the presence of impurities that do not interfere with the fundamental activity, and that may be present, for example, due to incomplete purification.
[0313] The term “purified” as used herein shall refer to a preparation comprising at least 50% (mol / mol), preferably at least 60%, 70%, 80%, 90% or 95% of a compound (e.gr, a bioprocess product or a POI). Purity is measured by methods appropriate for the compound (e.g., chromatographic methods, polyacrylamide gel electrophoresis, HPLC analysis, and the like).
[0314] Isolation and purification methods for obtaining a recombinant polypeptide or protein product may comprise methods utilizing difference in solubility, such as salting out and solvent precipitation, methods utilizing difference in molecular weight, such as ultrafiltration and gel electrophoresis, methods utilizing difference in electric charge, such as ion-exchange chromatography, methods utilizing specific affinity, such as affinity chromatography, methods utilizing difference in hydrophobicity, such as reverse phase high performance liquid chromatography, and methods utilizing difference in isoelectric point, such as isoelectric focusing may be used.
[0315] The following standard methods are preferred: cell (debris) separation and wash by Microfiltration or Tangential Flow Filter (TFF) or centrifugation, POI purification by precipitation or heat treatment, POI activation by enzymatic digest, POI purification by chromatography, such as ion exchange (IEX), hydrophobic interaction chromatography (HIC), affinity chromatography, size exclusion (SEC) or HPLC chromatography, POI precipitation, concentration and washing, such as by ultrafiltration steps.
[0316] The term “recombinant” as used herein shall mean “being prepared by or the result of genetic engineering. A “recombinant cell” or “recombinant host cell” is herein understood as a cell or host cell that has been genetically engineered or modified to comprise a nucleic acid sequence which was not native (or endogenous) to said cell. A recombinant host may be engineered to comprise an expression vector or cloning vector containing a recombinant nucleic acid sequence, in particular employing nucleotide sequence foreign to the host. A recombinant protein is produced by expressing a respective recombinant nucleic acid in a host.
[0317] The term “recombinant” with respect to a POI as used herein, includes a POI that is prepared, expressed, created or isolated by recombinant means, such as a POI isolated from a host cell transformed or transfected to express the POI. In accordance with the present invention conventional molecular biology, microbiology, and recombinant DNA techniques within the skill of the art may be employed. Such techniques are explained fully in the literature. See, e.g., Sambrook et al., 2012, Molecular Cloning: A Laboratory Manual, volumes 1-4, Cold Spring Harbor Press, NY).
[0318] Certain recombinant host cells are “engineered” host cells which are understood as host cells which have been manipulated using genetic engineering i.e., by human intervention. When a host cell is engineered to express, overexpress, underexpress, knockin or knockout a given gene or the respective protein, the host cell is manipulated such that the host cell has the respective capability to express such gene or protein, to a different extent compared to the host cell under the same condition prior to manipulation, or compared to the host cells which are not engineered. Cells that are not engineered by any recombinant means or techniques are generally understood as being naturally occurring or wild type.
[0319] Specific QAs are biomarkers which indicate the quality of a cell-based bioprocess product e.g., a bioprocess product which comprises cellular material (in particular whole cells, or fractions thereof), which may comprise a biomarker such as a cell surface marker. Often, biomarkers are employed to monitor cell behaviour or activity, or as readout of the bioprocess or the product. Specifically, a “biomarker” refers to a measurable indicator or characteristic that can be objectively evaluated and assessed as a sign of normal biological processes, pathogenic processes, or responses to the cell's environment.
[0320] Specifically, a biomarker can be a “cell surface marker”. Cell surface markers, also known as cell surface antigens or surface proteins, are specific proteins, glycoproteins, or other molecules present on the outer surface of a cell's plasma membrane. These markers serve various functions, including (i) Cell identification: They distinguish one cell type from another. Different cell types express distinct combinations of surface markers, allowing for their identification and classification; (ii) Cell communication: Surface markers play a role in cell-cell communication and signaling. They can interact with other cells, signaling molecules, or the extracellular environment, triggering various cellular responses; (iii) Immune response: Some cell surface markers act as antigens, playing a crucial role in the immune system by identifying cells as "self or "non-self." This helps the immune system recognize and respond to foreign invaders or abnormal cells; (iv) Cell adhesion: Certain surface markers facilitate cell adhesion to other cells or to the extracellular matrix, maintaining tissue structure and supporting various physiological processes.
[0321] Examples of cell surface markers include major histocompatibility complex (MHC) molecules involved in immune recognition, CD molecules used to classify immune cells (e.g., CD4 and CD8 on T cells), receptors for hormones or growth factors, and adhesion molecules like integrins. The CD nomenclature (Cluster of Differentiation) is an extensive system used to classify cell surface molecules based on monoclonal antibody studies and their function. A partial list as there are over 300 CD markers identified to date, each with distinct roles in cell identification, activation, signaling, and immune responses across various cell types, includes but is not limited to: CD1a to CD1e: involved in antigen presentation, found on dendritic cells, some B cells, and thymocytes;, CD2: facilitates adhesion between T cells and other cells; CD3 (CD3e, CD3g, CD3d): part of the T cell receptor complex, critical for T cell signaling; CD4: found on helper T cells, involved in recognizing antigens presented by MHC class II molecules; CD5: expressed on T cells and involved in T cell activation and signaling; CD8: found on cytotoxic T cells, interacts with MHC class I molecules for target cell killing; CD9: involved in cell adhesion, migration, and signal transduction; CD10: commonly used as a marker for certain types of leukemia and lymphoma; CD14: found on monocytes, macrophages, and neutrophils, involved in innate immune responses; CD19: typically found on B cells, involved in B cell signaling and activation; CD20: expressed on B cells, a therapeutic target in some lymphomas; CD21 : acts as a receptor for complement fragments, found on B cells and follicular dendritic cells; CD23: found on B cells, involved in IgE regulation; CD25: IL-2 receptor alpha chain, expressed on activated T cells; CD28: co-stimulatory molecule on T cells, necessary for activation; CD30: found on activated T cells and B cells, associated with lymphomas; CD31 (PECAM-1): involved in cell adhesion and angiogenesis, found on endothelial cells, platelets, and leukocytes; CD34: found on hematopoietic stem cells and endothelial cells; CD40: found on antigen-presenting cells, interacts with CD40 ligand on T cells for immune responses; CD44: cell adhesion molecule, involved in migration and signaling in various cell types; CD45: Tyrosine phosphatase found on leukocytes, involved in signaling; CD56 (NCAM): found on natural killer cells and some T cells, associated with cytotoxicity; CD80 (B7-1): co-stimulatory molecule on antigen- presenting cells; CD86 (B7-2): co-stimulatory molecule on antigen-presenting cells; CD95 (Fas / APO-1): involved in apoptosis, found on T cells and other cells;CD117 (c- Kit): receptor for stem cell factor, important in hematopoiesis; CD133 (Prominin-1): marker for hematopoietic stem cells and various stem cell populations.
[0322] Therefore, the present invention provides for improved methods of quality control of bioprocesses, in particular of producer cell cultures.
[0323] The invention is based on the finding that DNA methylation of CpG sites in regulatory regions is a key epigenetic mechanism regulating gene expression. Epigenetic profiling methods, which assess the CpG methylation status, elucidate the cell culture dynamics and their impact on bioprocess or bioproduct quality, such as the biopharmaceutical stability.
[0324] Specifically, the models described herein are based on the rate of change of different CpG sites and changes in the network of CqG sites. Changes to the environment of the cell can be accurately modeled which is novel and surprising compared to other existing bioprocessing models.
[0325] According to specific examples, it has been found that process parameters such as nutrient usage, and cell growth can be determined retrospectively, and predictively using a combination of estimative and predictive models, based on a CpG index composed of sites within demarcated DNA regions. Deviations from the reference specifications are detected at the methylome level providing invaluable information to take decisions about the outcome of the current bioprocess. Moreover, cellular stability can be confirmed with a quick and uncostly method and the identification of genes responsible for bioprocess relevant QAs or PPs can lead to a targeted approach to avoid an unwanted effect or prime a desired one.
[0326] QAs are often considered critical in determining the effectiveness, reliability, performance, and overall value of the bioprocess product, in particular the end product.
[0327] In manufacturing processes, quality PPs typically refer to characteristics like durability, precision, consistency, efficiency, or adherence to certain standards or specifications.
[0328] In Pharmaceuticals and Biotechnology, QAs can be critical in assessing the safety, efficacy, and consistency of drugs, vaccines, or biological products. Attributes might involve purity, potency, stability, sterility, or other specific criteria outlined by regulatory bodies. In Product Development, when designing any product, identifying and maintaining specific quality attributes can be essential to meet customer needs and expectations, ensuring that the final product functions as intended.
[0329] In Analytical Testing and Diagnostic, they refer to specific characteristics or properties of a bioprocess effected by a substance, material, or biological sample that are assessed and measured to ensure the quality, safety, efficacy, or proper functioning of a product or bioprocess.
[0330] It has been found that the methylation dynamics of each CpG site can be studied individually along the time course of an experiment. By comparing how each methylation pattern changes over time and across experiments, it is possible to categorize the CpG sites into motifs. Such categorization represents distinct types of methylation behaviour, and can be used for the machine leaning (ML) models establishment.
[0331] Below are examples of motifs and method to identify them:
[0332] • Inert I Indicator: Characterized using F-statistics, this motif shows no significant changes in methylation across the dataset (both experiments or timepoints), implying stable methylation states, or that the appropriate stimulus has not been provided to the cells. Inert indicates full methylation or unmethylation of the CpG site (i.e. a beta value of 0 or 1) and indicator an intermediate (yet stable) methylation state. While both inerts and indicators are used in the same way for ML model establishment, the distinction can be made as they may reflect different properties of the cellular population. An example of this is the CpG site associated with gene Capn2 (scaffold NW_003613581v1 , position 789418), a protease which may only be activated upon stimuli.
[0333] • Category: Utilizes F-statistics to compare variances among experiments but not among different timepoints. Unlike inerts and indicators, there is no change in methylation over time, but the base conditions of the culture have primed the cell population to different background levels of methylation. An example of this is the CpG site associated with gene SGK1 (scaffold NW_003615041v1 , position 359320), which distinguishes between different basal media and is involved in insulin sensing.
[0334] • Cycle: Characterized by a sinusoidal function, indicating periodic changes in methylation levels that repeat over time. An example of this is the CpG site associated with gene Mrps18c (scaffold NW_003621021v1 , position 9199), involved in protein translation. • Signal: Identified using Gaussian Mixture Models, this category captures sudden spikes in methylation, representing transient changes with no intermediate data points. An example of this is the CpG site associated with PsmblO as described in Example 7.
[0335] • Growth I Depletion: Modelled using sigmoidal and inverse sigmoidal functions, respectively, to model a constant increase or decrease in methylation. An example of this is the CpG site associated with G protein-coupled receptor Gpr119 (scaffold NW_003613712v1 , position 77065), involved in protein translation.
[0336] It can be concluded that usage of the identified CpG sites, together with the developed machine learning models, can accurately predict bioprocess performance and product quality, as well as model changes in said bioprocess, reducing time and cost.
[0337] The Examples which follow are set forth to aid in the understanding of the invention but are not intended to, and should not be construed to limit the scope of the invention in any way. The Examples do not include detailed descriptions of conventional methods and devices. Such methods and devices are well known to those of ordinary skill in the art.
[0338] EXAMPLES
[0339] Example 1 : Identifying CpG sites involved in the CpG indices for specific process metrics
[0340] Wet-lab methodology
[0341] For this experiment, CHO-K1 cells (CHO-K1 , ATCC CCL-61 derivative) expressing a non-glycosylated single chain antibody were grown in 3 different media (ActiPro (Cytiva), CDM4PERMAb (Cytiva), First CHOice (UGA Biopharma) under fed- batch conditions for a period of 14 days. Cells were fed every day starting on day 4 to maintain glucose and glutamine levels at 6.5 g / L and 8 mM, respectively. Every day, cell number and viability and nutrient usage and waste production were measured, and cell pellets and supernatant were collected for methylation analysis and protein and spent media analysis, respectively. All cultures had a viability of >90% at harvest, indicative of good culture conditions, and expressed the product, as detected by Western Blot and chromatography, among others.
[0342] Genomic DNA was extracted using the standard DNeasy Blood & Tissue Kit (Qiagen) workflow or an alternative automated method. All samples had a DIN (DNA Integrity number) greater than 7, indicative of good DNA quality. Libraries were prepared according to the Infinium protocol (Illumina) before measurement in the TALOS chip (Evonik), generating an .idat file.
[0343] Computational methodology
[0344] The .idat file contains fluorescence values representative of the methylated and unmethylated fraction of each CpG site assayed. These intensities are background- corrected and normalized with the SeSaMe package (Bioconductor), outputting beta values (percentage of methylation per CpG site). Beta values below the default detection confidence threshold are removed, and new values imputed (for example, using k- Nearest Neighbors) if at least a beta value for that particular CpG site exists in half of the experiments.
[0345] To reduce the complexity of the dataset before biomarker candidate discovery, the variability of each CpG site’s methylation across experiments and over the duration of each experiment was calculated. Then, a regression model identified those sites most important for the prediction of each of the PPs / QAs, which were combined into what we termed “CpG indices”, which based on the percentage of methylation at a given timepoint can accurately return the wet lab measurements for multiple PPs / QAs (a comprehensive list of these is as follows: cumulative glucose consumption, specific glucose consumption, specific lactate production, specific glutamine consumption, specific glutamate production, specific ammonia production, product titer, specific productivity, viable cell density, integral of viable cell density, growth rate; with representative examples illustrated in Figs. 1-3).
[0346] The CpG sites identified are shown in Table 1 (Fig.25). A total of 218 sites have been identified as predictors of 11 different quality attributes of the culture, with some of them being involved in more than one QA (grey shading). For every CpG site, the genomic location (as per the 2011 Cricetulus griseus reference assembly genome, CriGri_1 .0, INSDC Assembly GCA_000223135.1 , Aug 2011) is indicated, and when known, the gene and pathway that it regulates. The relevance of these CpG sites was confirmed on a subset of the dataset not used for training. IVCD, cumulative glucose consumption, specific glucose consumption, specific glutamine consumption and specific ammonia production were predicted with an R-squared > 0.9, while specific lactate production, VCD, growth rate, and specific glutamate production were predicted with an R-squared > 0.7 (Figures 1-3).
[0347] These results showcase the successful establishment of a CpG index consisting of as few as 5 CpG sites, which can estimate with accuracy the cell culture density, stage or metabolic profile at that same timepoint. We also created additional CpG indices by substituting the CpG sites thereby obtaining an index for an about 250 bp region that comprises said CpG sites, yielding comparable results, while CpG sites outside of the about 250 bp region were used as a negative control. This highlights the coherent regulation of a given QA / PP by any CpG in a given region (Figure 4).
[0348] Finally, we repeated the regression model using this time categorical values instead of continuous ones. The model predicted whether cells had been grown in any of the 3 basal media with no errors, irrespective of the timepoint. In addition, the model could also predict with high accuracy at which timepoint the sample had been taken, suggesting that with more sampling points we could also predict more specific cellular behaviours like cell cycle stage (Figure 5).
[0349] Based on the data available so far, the CpG sites responsible for the metabolism of glucose, glutamine, lactate, glutamate and ammonia have been selected. However, supporting literature (O'Neill et al. 2022, Mohammad et al. 2019) shows similar metabolic rates for other metabolites such as amino acids or trace elements. Moreover, full basal media can be distinguished with 100% accuracy, therefore, it is expected that the same machine learning approach can be applied to other medium components.
[0350] Example 2: Setting up a reference CpG methylation course for a bioprocess.
[0351] The CpG sites shown in Example 1 can accurately estimate cell culture QA / PP at the timepoint of sampling. To extend this prediction to timepoints before and after sampling, each of the QA / PP were fitted to a curve and the function needed to plot each curve based on its shape was defined. This shape is representative of the combined changes of each of the predictive CpG sites. Because the parameters to draw this function are known, it is possible with a single sampling point to construct the rest of the trajectory for that QA and / or PP and experiment. Figures 6 demonstrates how based on the methylation pattern of the selected CpG sites and their fitness to the functions described above, from a single timepoint the predictive model can construct the methylation state at the rest of the timepoints, and from there estimate the actual QA / PP, as described in Example 1. These results underpin the relevance of the selected sites and method not only as indicators of the current cellular state, but also as indices of cellular memory and trajectory.
[0352] Example 3: Comparing a CpG index of a bioprocess with the reference CpG methylation course.
[0353] In bioprocessing, a common strategy to achieve higher cell densities is to screen multiple basal media or perform a design of experiments (DoE) with different feeding concentrations. The CpG index can also estimate the performance of the bioprocess when testing changes to the reference process parameter. This was done so by repeating the experiment from Example 1 with the following changes. First, the cells were grown in a different basal medium. Second, the glucose concentration was doubled. The observed bioprocess data outcome was quite different, with almost a 50% decrease in total cell number in the new bioprocess strategy under consideration. Indeed, the experiment was stopped at day 11 due to a viability below the acceptance criteria.
[0354] Next, the models established during experiment 1 , which had not been retrained on this data, and therefore had never experienced this basal medium or nutrient composition were applied. Estimation of the cell growth (IVCD) based on the methylation readout and the CpG index was accurate (Figure 7), predicting as well a lower cell density as compared to the original process (the IVCD from Example 1 is shown here again as a dashed line for reference). These results confirm the validity of the models in cultures deviating from the standard process, and their applicability to Design of Experiment (DoE) or medium screening approaches. In addition, the results highlight that the CpG index is dynamic and the models work across a wide range of bioprocess outcomes (such as an IVCDs of between ~7 and 15 million). Example 4: Usability of CpG index to characterize scale-down models, develop scale-up models and confirm successful tech transfer
[0355] The original bioprocess (which was run in shake flasks) was adapted to a controlled manufacturing scale bioreactor operated in a different location. Therefore, there were changes to both the scale and the lab environment of the bioprocess. The same feeding strategy was maintained but the inoculation cell density was increased from 300,000 cells / mL to 500,000 cells / mL and the process was stopped earlier. The bioreactor culture and a reference shake flask culture were run in parallel using cells from the same seed train, and samples for methylation analysis were taken at the same timepoints.
[0356] Cell culture process parameters were calculated using the same models as in Example 1 , and the accuracy (based on R2 and RMSE) of both small and big scale cultures in the new facility is compared (Figures 8 and 9). These results highlight the robustness of the methylation-based models to changes in cell culture scales and geometry, as well as initial evidence of their use to validate technology transfer between labs.
[0357] Example 5: Translatability of the model to other fermentation modes.
[0358] While fed-batch is by far the most widespread bioprocess modality, others exist such as batch fermentation, where cells are grown for a shorter period of time as they do not receive nutrients once the culture starts. To test the accuracy of the CpG Index on this fermentation type, the cells for 6 days following a batch fermentation strategy were cultured and the models from Example 1 were applied.
[0359] As expected, viable cell density was lower than in fed-batch, and viability dropped faster, due to the lack of feeding. Importantly, the models accurately predict that the cells do not follow the exponential growth that normally starts on Day 4 during fed-batch, rather, it estimates that the cell number remains constant (Figure 10). This example shows that methylation changes and the CpG index are conserved when cells are cultured in different fermentation modes, which when combined with the previous use cases is a very powerful tool during bioprocess optimization.
[0360] Example 6: Monitoring and detecting deviations in a process under control. Examples 3, 4 and 5 describe how the CpG index can estimate the performance of a bioprocess based on daily analysis of the DNA methylation when parameters such as nutrition, manufacturing scale or mode and manufacturing plant are purposedly changed. However, the CpG index can also be used as part of a control strategy (i.e. in a GMP operation) where deviations from the expected outcome are reflected in said index.
[0361] Two identical fed-batch processes were run in parallel, starting from the same seed train. In GMP, all process parameters such as bioreactor temperature are defined within a narrow range and tightly monitored, as deviations can cause out of specification results and product rejection. To simulate a deviation of the process, the temperature in one of the bioreactors was decreased by 5 °C on day 6, which yielded 25% less product (Figure 11a). Figure 11 b shows that before the deviation occurs (on day 6), the models (forward prediction on Day 4) predict that both bioprocesses will yield a similar amount of product, as is expected from a reproducible manufacturing process. However, when looking at the CpG-index on a daily basis (Figure 11c), it can be seen that the estimated amount of product after day 6 (when the deviation occurs) starts to differ between both cultures, reflecting the actual observed behaviour in the culture. Indeed, by day 12 the model estimates a 35 % decrease in yield in the “uncontrolled” process, close to the real outcome.
[0362] This example illustrates how the herein disclosed DNA methylation-based models can detect deviations in the bioprocess almost immediately after they take place. This is an extremely powerful tool to control the bioprocess. Furthermore, in cases when there are no wet lab measurements around the time of the deviation, revisiting methylation samples can simplify investigations by pinpointing the time point when the deviation from the standard process (and standard CpG-index) occurred.
[0363] Example 7: Identification of CpG sites responsible for bioprocess product degradation.
[0364] In occasions, a process may be under control when looking at the defined critical process parameters and still yield a low-quality product. This occurs even in well-defined processes, where the trigger for the low quality cannot be pinpointed with the current techniques or knowledge, and the low-quality result can only be discovered later in the process, after the cell culture has been completed and the product purified. Here, we propose that analysis of selected CpG sites can reveal poor bioprocess outcomes or products with failing critical quality attributes, therefore allowing manufacturers to stop bioprocesses earlier (potentially correcting them) to save resources.
[0365] CHO cells were grown according to the bioprocess described in Example 1 and processed as above. In samples cultured with particular medium components, a degradation of the bioprocess product was observed (Figure 12, white box).
[0366] An initial differential methylation analysis between samples showing intact and fragmented products did not reveal clear changes in the CpG-index. However, the trajectories of certain CpGs (the daily change in methylation value) were “stable” in experiments with no degradation and shifted (or signalled) in those experiments with degradation. Further analysis via gene ontology identified these “degradation-signaling” CpGs as protease PsmblO and autophagy-related Ulk1 (Figure 13). Proteases and control of protein degradation have been shown to be regulated by shifts in methylation (Chen and Chai, 2002), while the autophagy-lysosomal pathway is also critical in proteostasis, or correct maintenance of the proteome (Korovila et al. 2017). Therefore, a link can be established between some medium components (or concentrations) in ActiPro or CDM4PERMAb, which are not present in FirstCHOice, that can activate or repress said enzymes via methylation.
[0367] These results suggest that identification of QA-impacting genes, such as protein integrity, can be performed via differential methylation analysis. This is of extreme novelty, as knowledge of how the activity of these CpG sites affects the product quality can help inform in advance whether the bioprocess will yield an unusable product, before the use of conventional analytical techniques becomes possible.
[0368] Example 8. Use of the established CpG indices to simulate fed-batch data from a cell line stability testing.
[0369] Cell line stability testing is key to define the limit of in vitro cell age (LIVCA). This limit is reached when the cell becomes unstable as evidenced by growth rates and productivities that deviate from those of the cell at earlier passages (younger age).
[0370] To assess whether the cells were still stable, a master cell bank (MCB) for 35 passages were cultured to generate an end of production cell bank (EOP) which was 108 days older than the MCB. Passaging was performed in shake flasks every three days with an inoculation cell density of 300,000 cells I mL. Cells were taken at the end of each passage.
[0371] Cells from the MCB and EOP were run in parallel in fed batch mode following the same strategy from Example 1. Titer was monitored daily to assess productivity as the most common strategy to understand and define the limit of in vitro cell age. During the fed batch, cultures were also sampled daily for methylation assessment.
[0372] Wet lab analysis (via HPLC) shows that the EOP cells produce more protein than MCB, although both titers are comparable (Fehler! Verweisquelle konnte nicht gefunden werden.4). Next, it was tested if instead of running the full fed-batch culture, one could simply predict the fed-batch performance based on the passaging data (that is, the cells that were passaged every three days for a hundred days to generate an EOP from the MCB).
[0373] For this, the methylation data was taken from seed train samples matching the MCB and EOP (passages P6 and P34, respectively) and simulated via the forward prediction models from Example 2 what the titer would have been had these cells being grown as fed batch cultures for an additional 10 days (Figure 14). These predictions align with the observed data and capture the higher titer of EOP as compared to MCB.
[0374] These results support two conclusions. First, Example 5 had already shown the adaptability of the developed and described estimative (daily) models to other fermentation modes (fed batch to batch). Here, this is expanded by demonstrating that the predictive models also translate well from batch to fed-batch fermentation. Second, they provide the unique use case of using methylation-based CpG index to predict the outcome of a stability-indicating study without the need to conduct a fed-batch experiment, proposing a resource-saving alternative to the current method.
[0375] Example 9. Prediction of cell age as indicator for cell line stability.
[0376] CHO cells were grown according to the process described in Example 8. Cell growth kinetics were comparable among all passages, from P6 to P34 (growth rate = 0.79 ± 0.02 days-1). Material was extracted and processed as in Example 1 , and the beta values for each sample generated as above.
[0377] Then, the methylation trend of every CpG site over time was considered. Random fluctuations are to be expected due to cellular events and therefore disregarded, but consistent gains or losses of methylation are indicative of cell aging. Time series analysis revealed 868 CpG sites that fit this criterium, and the relevance of the 15 with the highest loading was confirmed using unsupervised learning.
[0378] When plotting passage 6 to 34 using only the top 15 age-related CpG sites a longitudinal distribution of the cells was observed, with an increase in age shown along the x axis (PC1) (Figure 15a). To confirm that the “aging distribution” is due to a set of CpG sites and not due to a global epigenome shift on the old cells, the PCA was repeated using a subset of random CpG sites as a negative control, which yielded an unordered clustering (Figure 15b).
[0379] Finally, these data were integrated with the samples already described in Example 8, which include samples that are “younger” than P6 and older than P34. Moreover, these new samples have been grown as fed-batch. Figure 16 reveals how the new samples respect the original distribution, with MCB / RCB samples (P0 and P1) being to the left of P6, and LIVCA samples (P35) being to the right of P34. Importantly, the samples here are a combination of fed-batch and batch, highlighting how these “aging” CpG sites are conserved across fermentation modes.
[0380] These results reveal that comparison of the epigenetic methylation pattern is a reliable method of determining cellular age. Further, comparison of the epigenetic methylation pattern is a quick and effective method of comparing cellular stability and cell banks established at different cellular generations including LIVCA.
[0381] Example 10. A methodology to track and characterize cell line adaptation.
[0382] Example 3 demonstrated how changes in the basal medium impact fed-batch performance. However, before the start of the fed-batch cells need to be adapted to their environment, and this usually involves adaptation to the basal medium specific for said bioprocess. This adaptation phase is performed for an arbitrary number of passages that err on the side of caution, as there are no clear indicators of when cells have successfully adapted (other than no decline in growth rate). Here, it is shown that monitoring cell adaptation to the environment can be performed in a more rigorous and reproducible manner by following the methylation status of selected CpG sites.
[0383] Cells originally grown in CD-CHO (Gibco) were adapted to 4 additional different basal media over 5 passages. The reference medium was run in parallel as a control. The growth rate was identical in all passages and across all media, suggesting adaptation had been successful.
[0384] To confirm the adaptation of the cells, and the speed of this event, we looked at the methylation profile of those CpG sites which were differentially expressed at the end of the 5-passage adaptation period. The methylation percent in CD-CHO samples (to which the cells were already adapted) did not shift throughout the culture. However, the methylation status of those same basal medium “sensing” CpG sites shifted slightly in as little as one passage (Figure 17) and maintained that separation throughout the entire adaptation phase.
[0385] This result suggests that if cells are being adapted to new conditions, it can be ensured that the process is complete once the cells reach the expected methylation fingerprint, shortening production times.
[0386] Example 11. Prediction of cell growth and cell age on independent cell lines.
[0387] Two additional, independent cell lines were analyzed. The first one (Null CL, Host CHO-GS Cell Line transfected with a plasmid without gene of interest), derived from CHO-K1 , ATCC CCL-61) was cultured with and without methotrexate (MTX), a selection agent commonly used in cell culture recently reported to cause DNA methylation changes (Guderud et al. 2021). The second one (CLC, CHO-GS Cell line stably transfected with a plasmid encoding the gene of interest, derived from CHO-K1 , ATCC CCL-61) was cultured at different cell bank stages (MCB, RCB, two independent EOPs).
[0388] Cells were cultured for varying lengths of time and Viable Cell Density was only measured at harvest. Therefore, all cell densities were normalized to the reference (Null CL without MTX). We applied the CpG index established in Example 1 (without retraining the ML model, therefore using only the CpG sites determined for the original cell line) and accurately estimated the VCD (within the 10% range of the actual value) for 5 of the 6 experiments (Figure 18).
[0389] Next, we focused only on the cell line CLC, where we have multiple cell banks: MCB (Master Cell Bank = early passage), WCB (Working Cell Bank = mid-early passage), EOP1 and EOP2 (End of Production Cell Bank = late passage, generated twice independently). To confirm that the “cell aging” subset established in Example 9 is transferable to other cell lines too, we plotted a PCA using only the CpG sites from that Example. Indeed, here all cell banks also projected along the x axis based on their age: MCB > WCB > EOP1 = EOP2 and WCB clustered nearer to MCB than to EOP, as expected due to the fact that WCB is closer to MCB than to EOP in terms of passage number (Figure 19).
[0390] These results highlight how CpG indices established in one cell line can be applied to different cell lines, even if they are not of the same CHO strain. This is crucial as it evidences the conserved behavior of CpG sites and cements its use as a universal tool without the need for extensive retraining or data re-generation.
[0391] Further, the fact that commonly used cell culture reagents to increase productivity such as MTX do not impact the CpG indices is significant, as once more it corroborates the fact that the models could be applied to cell culture strategies relying on unstable clones, transient transfections or DNA modifying reagents such as VPA. Example 12. Establishment of an in silico model to predict bioprocess outcome and calculate bioprocess optima.
[0392] Based on the prediction accuracies observed in the previous examples, we noticed some of the QAs or PPs cannot be fully estimated due to missing measurements. To overcome this technical issue, and make the appropriate estimative and predictive models failproof, a neural network was developed (Figure 9) to model the (i) direct effect of a CpG on a QA, (ii) the direct effect of a PP on a CpG, (iii) the direct effect of a CpG on another CpG and (iv) the indirect effect of any of the above onto each other. This digitalization of the observed methylation patterns, how they are affected upon by the bioprocess PPs and their effect on the bioprocess QAs serves a three-fold purpose:
[0393] A) Ensure accurate predictions even in the absence of a sampling point;
[0394] B) Identify unknown cause-effect relationship, as well as CpG sites or regulatory regions that have not been measured due to lack of knowledge of their effect on gene activity;
[0395] C) Translate the CpG indices and methylation courses here described to a fully computational system, informing the user about the optimal or suboptimal outcome of the current bioprocess and avoiding the need to conduct process development or comparison experiments on the reference bioprocess in the first place.
[0396] As a first proof of concept, this network was used to calculate the maximum potential VCD given the current bioprocess conditions. Based on Bayesian probabilities, it is possible to estimate which is the possible range of methylation for each CpG site involved in a CpG index. There are states that are mutually exclusive, such as CpG site A with a 20% methylation and CpG site B with a 100% methylation. Calculating all the possible combinations of methylation for each CpG site in the index will give an optimum, which signifies the maximum capacity (or performance) of the bioprocess at its current state. Figure 21 shows how the current bioprocess (termed ‘actual’) is close to the optimum (termed ‘mode’) but some CpG sites (which process parameters influence those CpG sites is currently being determined) need tweaking to reach it. Example 13. Assessing quality of raw materials in upstream manufacturing.
[0397] Qualification of raw materials such as the powdered basal medium or processspecific growth factors are the first step in any manufacturing process and will determine its success. Even in controlled processes, differences in the raw materials undetected by the established analytical techniques can lead to variation or even out of specification results for product CQAs. Similar to Example 7, CpG index analyses could detect the variability of such raw materials and be used as a release test, as illustrated in the conceptual example in Figure 22.
[0398] Example 14. Establishment of predictive models on human cells and alternative species.
[0399] Human cells (T-cells, prepared from different donors) were grown on fed-batch modality on two culture conditions causing very distinct outcomes: optimal conditions, yielding high cell density, and suboptimal conditions where the basal medium lacked key growth factors, yielding low cell density. Wet lab measurements for VCD and viability, as well as samples for methylation analysis, were taken at 3 timepoints: seeding, feeding and harvest, Methylation samples were processed as in Example 1 although this time an Infinium MethylationEPIC v2.0 BeadChip (Illumina) was used to assay the methylation levels of each CpG site.
[0400] The computational analysis was also performed according to the methodology described in Example 1 . Shortly, feature selection based on experimental variability was employed to select those methylation sites that could be involved in QA / PP estimation, and later multiple machine learning approaches were compared and ranked (via cross validation) to determine which one yielded the most robust value estimations. Like in Example 1 , the CpG sites contributing the most to model establishment were incorporated into a CpG index that performed with the same accuracy on an independent, testing dataset. Each CpG index is unique to each QA or PP, and the coefficients of each CpG site inside that index are the output on the algorithm yielding the lowest RMSE (root mean squared error) and MAE (mean absolute error) across all cross-validation folds.
[0401] Figure 23 displays the accuracy of the methylation-based human models in estimating the values for Viable Cell Density (Figure 23a) and viability (Figure 23b). Using only the methylation values of those CpG sites within that particular’s PP or QA CpG index the cellular observations (cell count and viability) can be calculated. More importantly, the model is robust across a spectrum of culture conditions, being accurate even when the difference in VCD at harvest is >3-fold.
[0402] These results showcase three key points about the methylation-based predictive technology: i) bioprocesses using human cells are also amenable to be monitored and controlled based on DNA methylation; ii) given that in the cell therapy field control strategies are still being developed due to bioprocess outcome variability, and human cells are usually the product, this technology proposes a novel and robust approach to monitor both the bioprocess PPs but also the QAs such as the cells in this case; iii) validates the methodology behind our methylation-based approach as robust, implicitly also validating that these models could be extended to species other than these two,
[0403] Example 15. Prediction of cell growth based on stability time series.
[0404] The predictive models described in Example 2 cover timepoints in which the model had been trained (Days 4 to 14). However, during cell passaging (for example, to generate stability indicating data) cells are passaged (grown sequentially) for an extended period of time with harvesting every 3 days. Here, the application of the established CpG indices is shown to predict what the viable cell density is at earlier timepoints.
[0405] Cells were grown as per Example 8 and harvested at Day 3 (during the exponential growth) All data was normalized to the first passage (P6) and the methylation-based VCD model applied. Figure 24 shows how the model is able to predict the trends in growth (i.e. P6 and P9 grow more than P7) as well as yield precise estimation of the cell density.
[0406] These results highlight the applicability of the models to timepoints beyond those in which they were initially trained. REFERENCES
[0407] Affinito, O., Palumbo, D., Fierro, A., Cuomo, M., De Riso, G., Monticelli, A., Miele, G., Chiariotti, L. and Cocozza, S., 2020. Nucleotide distance influences co-methylation between nearby CpG sites. Genomics, 112(1 ), pp.144-150.
[0408] Chen, L. M., and K. X. Chai. 2002. 'Prostasin serine protease inhibits breast cancer invasiveness and is transcriptionally regulated by promoter DNA methylation', Int J Cancer, 97: 323-9.
[0409] Crowgey, Erin L., Adam G. Marsh, Karyn G. Robinson, Stephanie K. Yeager, and Robert E. Akins. 2018. 'Epigenetic machine learning: utilising DNA methylation patterns to predict spastic cerebral palsy', BMC bioinformatics, 19: 225-25.
[0410] Guderud, K., Sunde, L.H., Flam, S.T., Maehlen, M.T., Mjaavatten, M.D., Norli, E.S., Evenrod, I.M., Andreassen, B.K., Franzenburg, S., Franke, A. and Rayner, S., 2021. Methotrexate treatment of newly diagnosed RA patients is associated with DNA methylation differences at genes relevant for disease pathogenesis and pharmacological action. Frontiers in immunology, 12, p.713611 .
[0411] Holder, Lawrence B., M. Muksitul Haque, and Michael K. Skinner. 2017. 'Machine learning for epigenetics and future medical applications', Epigenetics, 12: 505-14.
[0412] Korovila, I., Hugo, M., Castro, J.P., Weber, D., Hohn, A., Grune, T. and Jung, T., 2017. Proteostasis, oxidative stress and aging. Redox biology, 13, pp.550-567. Marx Nicolas, Peter Eisenhut, Marcus Weinguny, Gerald Klanert, Nicole Borth. 2022. „How to train your cell - Towards controlling phenotypes by harnessing the epigenome of Chinese hamster ovary production cell lines”, Biotechnology Advances 56:107924.
[0413] Wippermann, Anna, Sandra Klausing, Oliver Rupp, Stefan P. Albaum, Heino Buntemeyer, Thomas Noll, and Raimund Hoffrogge. 2014. 'Establishment of a CpG island microarray for analyses of genome-wide DNA methylation in Chinese hamster ovary cells', Applied microbiology and biotechnology, 98: 579-89.
[0414] Olivecrona, Marcus, Thomas Blaschke, Ola Engkvist, and Hongming Chen. 2017. 'Molecular de-novo design through deep reinforcement learning', Journal of cheminformatics, 9: 48-14.
[0415] Ou, Kristy, Dania Hamo, Anne Schulze, Andy Roemhild, Daniel Kaiser, Gilles Gasparoni, Abdulrahman Salhab, Ghazaleh Zarrinrad, Leila Amini, Stephan Schlickeiser, Mathias Streitz, Jbrn Walter, Hans-Dieter Volk, Michael Schmueck- Henneresse, Petra Reinke, and Julia K. Polansky. 2021. 'Strong Expansion of Human Regulatory T Cells for Adoptive Cell Therapy Results in Epigenetic Changes Which May Impact Their Survival and Function', Frontiers in cell and developmental biology, 9: 751590-90.
[0416] Mohammad, A., C. Agarabi, S. Rogstad, E. DiCioccio, K. Brorson, M. Ashraf, P. J. Faustino, and C. N. Madhavarao. 2019. 'An ICP-MS platform for metal content assessment of cell culture media and evaluation of spikes in metal concentration on the quality of an lgG3:kappa monoclonal antibody during production', J Pharm Biomed Anal, 162: 91-100.
[0417] O'Neill, E. N., J. C. Ansel, G. A. Kwong, M. E. Plastino, J. Nelson, K. Baar, and D. E. Block. 2022. 'Spent media analysis suggests cultivated meat media will require species and cell type optimization', NPJ Sci Food, 6: 46.
[0418] Taryma-Lesniak, O., Bihkowski, J., Przybylowicz, P.K., Sokolowska, K.E., Borowski, K. and Wojdacz, T.K., 2024. Methylation patterns at the adjacent CpG sites within enhancers are a part of cell identity. Epigenetics & Chromatin, 17(1 ), p.30.
Claims
CLAIMS1. A method for predicting the quality characteristics of a eukaryotic bioprocess or of a product that is produced by a eukaryotic bioprocess, comprising: a) determining for a eukaryotic bioprocess cell, a genomic CpG-methylation index for a number of CpG-methylation sites within more than one genomic regulatory elements, at an index timepoint of that bioprocess; b) comparing the genomic CpG-methylation index to a predetermined reference CpG methylation course to predict at least one product quality attribute (QA) or process parameter (PP) for a duration of the bioprocess which includes at least said index timepoint and a time period of the bioprocess following or preceding said timepoint; wherein the reference CpG methylation course comprises the CpG-methylation changes at said CpG-methylation sites over said duration, which has been predetermined to be indicative of said at least one QA or PP by comparing with a reference bioprocess of comparable type, preferably wherein said at least one QA or PP is predicted based on an estimation or prediction model.
2. The method of claim 1 , wherein said bioprocess is process of culturing a bioprocess cell, which is a batch, repeated batch, fed-batch, perfusion, continuous, or steady-state cell culture, or a process of culturing an organoid or a tissue.
3. The method of claim 1 or 2, wherein the index timepoint is at a time during an adaptation phase, seed culture, growth phase, production phase, or maturation phase of the bioprocess.
4. The method of any one of claims 1 to 3, wherein said duration comprises at least one generation time of the eukaryotic cell, or the duration of the bioprocess, preferably at least up to the limit of in vitro cell age (LIVCA).
5. The method of any one of claims 1 to 4, wherein CpG-methylation sites are selected from CpG sites within a genomic region of less than 500 bp that is located within a genomic regulatory element.
6. The method of any one of claims 1 to 5, wherein said genomic regulatory elements are selected from a promoter, enhancer, silencer, insulator, transposon element, or RNA regulator.
7. The method of any one of claims 1 to 6, wherein: a) said at least one QA comprises: i) for an expression product, at least one of protein stability, integrity, degradation, post-translational modification, purity, folding, isoform, activity, potency, absence or presence of one or more biomarkers or contaminants; ii) for a cellular product, at least one of stability, karyotype, viability, cell surface properties, maturation status, purity, activity, potency, strength, oncogene activation, physical or organoleptic parameters, absence or presence of contaminants; b) said at least one PP comprises i) a parameter determined in a culture of the bioprocess cell, preferably cell viability, viable cell density, product titre, bioprocess specific productivity, product yield, metabolite production, substrate consumption, growth, doubling time, pH, pressure, temperature, shear stress, shear resistance, mixing time, dissolved oxygen, pCO2 level, redox state, cell cycle; aggregate formation, cell size, cell surface characteristics, oncogene activation, reverse transcriptase and / or oncogene expression, host cell protein content, host cell DNA content, or plasmid DNA content, or ii) a parameter determined by the presence or absence of substances comprised in the bioprocess, such as a cell-derived or culture-derived substance, metabolites, or cell culture media components.
8. The method of any one of claims 1 to 7 wherein said at least one QA or PP determines the stability of bioprocess cells or the bioprocess.
9. The method of any one of claims 1 to 8, wherein said estimation or prediction model estimates or predicts at least one QA or PP across multiple timepoints of the bioprocess.
10. The method of any one of claims 1 to 9, wherein said estimation or prediction model is configured to model at least one of: i) direct effects of CpG methylation on QA; ii) direct effects of PP on CpG methylation; iii) direct effects between different CpG methylation sites; iv) direct effects between any of i), ii), iii), or iv).11 . The method of any one of claims 1 to 10, wherein said estimation or prediction model is a mathematical, deep learning, Al, or foundation model, or a hybrid model, preferably a statistical or informed machine learning model.
12. The method of any one of claims 1 to 11 , wherein said estimation or prediction model employs a deep neural network or convolutional neural network or regression model, that uses the CpG methylation levels to obtain estimates or predictions for the associated QAs and PPs.
13. The method of any one of claims 1 to 12, wherein said CpG-methylation index comprises a system of numbers obtainable by comparing the methylation levels at said CpG-methylation sites, which change according to each other, or by comparing the methylation levels to a fixed standard.
14. The method of any one of claims 1 to 13, wherein the bioprocess cell is a mammalian, animal, insect, avian, plant or yeast cell, preferably a CHO or human cell.
15. The method of any one of claims 1 to 14, wherein the product is heterologous or endogenous to the bioprocess cell, preferably a proteinaceous, nucleic acid, or cellular product, preferably wherein: a) the proteinaceous product is a protein, polypeptide, peptide, virus, or virus-like particle, or an assembly of any of the foregoing; b) the nucleic acid product is a DNA or RNA product, such as a virus, bacteriophage, plasmid or vector; c) the cellular product comprises or consists of a whole cell, a population of cells, tissue, organ, or organoid, or a fraction of any one of the foregoing.
16. The method of any one of claims 1 to 15, wherein the product is selected from the group consisting of an antigen-binding protein, an enzyme, a peptide, a protein antibiotic, a toxin fusion protein, a structural protein, a regulatory protein, a protein vaccine antigen, a hormone, a growth factor, a cytokine, and a blood clotting or coagulation factor, preferably wherein: a) the antigen-binding protein is an antibody molecule, preferably a monoclonal antibody which is any of a full-length antibody, an antibody comprising one or more epitope binding fragments of a full-length antibody, or a bispecific or multi-specific antibody comprising one or more of said fragments, preferably wherein said one or more epitope binding fragments of a full-length antibody are Fab, Fab', F(ab')2, Fv, or scFv fragments, or single domain antibodies; b) the product is a process enzyme or metabolic enzyme.
17. Use of the method of any one of claims 1 to 16 in a method of a) controlling the specific productivity qp of the bioprocess, preferably wherein the CpG-methylation index comprises CpG sites which regulate genes of protein formation folding, or transportation within the cells and excretion into the cell culture fluid, preferably wherein said genes encode ribosomes, mRNA expression, endoplasmic reticulum structures, Golgi apparatus structures, chaperones, or b) controlling the product quality, such as product degradation, preferably wherein the CpG-methylation index comprises CpG sites which regulate one or more intracellular protein degradation genes, preferably wherein said genes encode proteases compartmentalized in proteasomes or in degradative organelles, such as lysosomes; or c) controlling the bioprocess, preferably during Good Manufacturing Production (GMP), upscaling or downscaling, technology transfer, cell culture media development, monoclonality assessment, raw material release, parametric release, cell culture media release, quality by design, growth performance or any other cell culture based test, cell line development, cell line monitoring, such as monitoring adaption, aging, or maturation, process development, clone selection, cell bank release or stability assessment, preferably wherein the bioprocess is controlled by setting the CpG-methylation index according to the predetermined reference CpG methylation course; d) comparing the quality of different bioprocesses, wherein a first bioprocess is compared with a second bioprocess being used as a comparator, preferably wherein theCpG-methylation index for said second bioprocess is compared with a reference CpG- methylation course that is predetermined from said first bioprocess.
18. The use according to claim 17, wherein: a) a first bioprocess produces an originator product, and a second different one produces a comparator product preferably a biosimilar, or biobetter; or b) a first bioprocess is performed using a specific fermentation mode, and second different one uses a comparator fermentation mode; or c) a first bioprocess is performed at one manufacturing site, and a second different one is performed at a comparator manufacturing site; or d) a first bioprocess is using a cell culture parameter or medium, and a second different one uses a respective comparator cell culture parameter or medium; or e) a first bioprocess uses a scale of production, and a second different one uses a comparator scale of production; or f) a first bioprocess is performed to produce the product employing a specific bioreactor or bioreactor system, and a second different one uses a comparator bioreactor or bioreactor system; or g) for controlling, correcting or adjusting a QA or PP based on deviations of the CpG index from the predetermined reference CpG-methylation course.
Citation Information
Patent Citations
Methods and compositions for assessing CpG methylation
EP1589118A2
Method for the selection of a long-term producing cell
EP2558591A1
Methods of detecting methylated cpg
EP4165213A1
Differential methylation level of CpG loci that are determinative of a biochemical reoccurrence of prostate cancer
US11566291B2
Low-dosed solid oral dosage forms for hrt
WO2011128337A2
Cited By
Cell type and proportion prediction method of methylation basic model based on image and related equipment
CN121191162A