Data-driven predictive modeling for cell line selection in biopharmaceutical production

JP7901127B2Active Publication Date: 2026-08-05AMGEN INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
AMGEN INC
Filing Date
2024-09-24
Publication Date
2026-08-05

AI Technical Summary

Benefits of technology

【0015】 別の態様では、上記のプロセスに加えて又はその代わりに、1つ以上の機械学習アルゴリズムを使用して、いずれのクローンがサブクローニングステージから小規模スクリーニング培養(例えば、図1のステージ11からステージ12)に進むべきかを選択し得る。典型的には、サブクローニングステージの終わりに高い細胞生産性スコア及び多くの細胞数の両方を有するクローンは、小規模スクリーニング培養(流加バッチ実験)において高い性能を達成する最良の候補であると考えられてきた。このアプローチは、典型的には、およそ30~100クローンの流加バッチステージへの前進をもたらす。しかしながら、本明細書に記載の機械学習アルゴリズムは、サブクローニングステージ及び先行する細胞プールステージの両方で候補クローンの種々の属性を分析し、仮想小規模(例えば、流加バッチ)培養実験から生じる特定の製品品質属性(例えば、力価、細胞増殖又は比生産性)を予測することにより、このプロセスを改善することができる。クローンの生成及び増殖のマイクロタイタープレートに基づく方法(すなわち図1のサブクローニングステージ11)は、例えば、Berkeley Lights Beacon(商標)光-電子細胞株生成及び分析システムなど、より効率的であり、高スループットあり、且つ高含有量のスクリーニングツールの使用で置換され得る。候補細胞株について製品品質属性値を予測した後、候補は、予測された値に従ってランク付けされ、それにより細胞株開発の次のステージに向けた候補クローンのより小さいサブセットの選択を容易にする。有利には、これらの値に従って作成されたランキングは、基礎となる予測値が比較的低い精度を示し、したがって表面上では不十分であるように見えても、特定の機械学習モデルでは高度に正確であり得る。実施形態に応じて、このプロセスは、小規模スクリーニング培養のための候補クローン/細胞株(すなわち小規模培養において最良の性能を示すものである可能性がより高いクローン)を選択する場合、より少ないリソース使用(例えば、時間、コスト、労力、設備などに関して)を必要とし、且つ/又はより良好な標準化を提供し得る。例えば、流加バッチステージに進められる細胞の数を減らすことは、他の薬物製品について他の細胞株を試験する能力を解放し得る。いくつかの実施形態では、小規模スクリーニングステージは、様々な細胞株のランキングに基づいて完全にスキップされ得る(例えば、プロセス10のステージ11からステージ14に直接進むことにより)。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007901127000013
    Figure 0007901127000013
  • Figure 0007901127000014
    Figure 0007901127000014
  • Figure 0007901127000015
    Figure 0007901127000015
Patent Text Reader

Abstract

To provide methods for predicting a relative rank of cell lines advanced from a clone generation and analysis process according to a certain product quality attribute.SOLUTION: A method for facilitating selection of master cell lines from among candidate cell lines that produce recombinant proteins comprises: receiving first attribute values for the candidate cell lines measured using an opto-electronic cell line generation and analysis system; acquiring second attribute values that include one or more attribute values measured at a cell pool screening stage of the candidate cell lines; determining ranking of the candidate cell lines according to a product quality attribute associated with hypothetical small-scale screening cultures, where the determination of the ranking comprises predicting, for each of the candidate cell lines, a value of the product quality attribute by analyzing the first and second multiple attribute values using a machine learning-based regression estimator, and comparing the predicted values; and presenting an indication of the ranking to a user via a user interface.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to related applications Priority is claimed to U.S. Provisional Patent Application No. 62 / 841,186, filed Apr. 30, 2019, and U.S. Provisional Patent Application No. 63 / 014,398, filed Apr. 23, 2020, the entire disclosures of which are incorporated herein by reference.

[0002] This application generally relates to cell line (clone) selection technology, and more specifically to a technology for predicting the relative rank of cell lines advanced from a clone generation and analysis process according to specific product quality attributes.

Background Art

[0003] Large complex molecules (e.g., proteins), known as biologics in the biopharmaceutical industry, are derived from living systems. A general workflow for the development of biologics starts with research and development. At this initial stage, diseases or indications representing unmet important medical needs are targeted. Researchers determine promising drug candidates based on, for example, the appropriate target product profile governing aspects such as safety, efficacy, and route of administration. Ultimately, a combination of in vitro studies and computational models selects a particular molecule as the top drug candidate for a particular disease and target population. After the top candidate is selected, the blueprint of that molecule is formalized into a gene, and the gene of interest is inserted into an expression vector. Then, the expression vector is inserted into a host cell in a process known as transfection. When transfection is successful, the cell can incorporate the gene of interest into its own production mechanism and ultimately acquire the ability to produce the desired pharmaceutical.

[0004] Because each cell possesses unique characteristics, the products produced by each cell vary slightly in terms of productivity (e.g., potency) and product quality. Generally, for economic and safety reasons, it is more desirable to produce drugs with consistently high potency and consistently high quality. Higher product concentrations or potencies help reduce the manufacturing footprint required to achieve the desired yield, thus saving both capital and operating costs. Higher product quality ensures that a larger proportion of the drug is safe, effective, and usable, which also saves costs. In relation to cell line development, product quality attributes are evaluated through assays performed on the product of interest. These assays often involve chromatographic analysis, which is used to determine attributes such as the degree of glycosylation and the proportion of unusable proteins due to cleavage (clipping) or aggregation (aggregates), as well as other factors.

[0005] Based on criteria for productivity and product quality, the "best" cell lines or clones are selected in a process known as "cell line selection," "clone selection," or "clone screening." The selected cell lines / clones are used for the master cell bank, which serves as a uniform starting point for all future manufacturing (e.g., clinical and commercial).

[0006] Ensuring a consistent product batch helps promote more uniform and predictable pharmacokinetic and pharmacodynamic responses in patients. However, when producing the desired product using a “pool” of heterogeneous cells obtained after gene transfer, many variants of the resulting product may exist. This is because, during gene transfer, the target gene is incorporated into candidate host cells in various ways. For example, there may be differences in copy number (i.e., the number of copies into which the target gene is incorporated) and other differentiation factors between the unique footprints of different cells. The production of the desired product can also vary due to slight differences in the internal mechanisms of individual cells, including the nature of post-translational modifications. These variations are undesirable, especially considering the need to ultimately control and guarantee a measurable safe response in patients. Therefore, typically, cell lines in a master cell bank are required to be “clonally induced,” meaning the master cell bank contains only cells derived from a common single cell ancestor. This theoretically facilitates ensuring a greater degree of uniformity in the drugs produced, although there will be slight but inevitable differences due to natural genetic variation from random mutations during cell division. Therefore, the clonal screening process is crucial not only for distributing productive and high-quality starting materials, but also for identifying the only cell line that meets the "clonal" requirement.

[0007] Figure 1 shows a typical clone screening process 10. The first stage 11 describes a conventional microtiter plate-based method for clone generation and proliferation, which may take 2-3 weeks. Hundreds of pooled heterogeneous cells are sorted into single-cell cultures by processes such as fluorescence-activated cell sorting (FACS) or limiting dilution. After restoring to a healthy and stable population, these clone-derived cells are analyzed, and the selected population is moved to stage 12. In stage 12, clone cells are cultured in “small cell cultures” (e.g., a 10-day fed-batch method) in small containers such as spin tubes, 24-well plates, or 96-depth plates. In this small-scale process, bolus nutrients are added periodically, and different measurements of cell proliferation and viability are obtained. Typically, hundreds or thousands of these small-scale cultures are carried out in parallel. At the end of the culture (e.g., day 10), the cells are collected for assay and analysis.

[0008] In Stage 12, the growth and productivity characteristics of clones in small cultures are analyzed to select the "top" or "best" clones (e.g., top 4) for scale-up culture to be carried out in the third Stage 14. The scale-up (or "large-scale") process is useful compared to the small-scale culture in Stage 12 to better represent the process that will ultimately be used in clinical and commercial production. The scale-up process may be carried out, for example, by culturing for 15 days in a 3-5 liter perfusion bioreactor. These perfusion bioreactors are adapted to the more efficient movement of waste products and nutrients, thereby increasing the overall productivity of the culture. Perfusion bioreactors typically allow for more precise control and monitoring, as well as a greater number of measurable variables, such as routine and continuous process conditions and metabolite concentrations.

[0009] Following the scale-up process in Stage 14, the medium and product are collected and analyzed. Finally, in the fourth Stage 16, the scale-up product that yields the highest titer and exhibits the best product quality attributes (PQA) is selected, typically as the "best" or "winning" clone. Lastly, in the fifth Stage 18, the winning clone is used as a master cell bank for future clinical and commercial manufacturing. [Overview of the Initiative] [Problems that the invention aims to solve]

[0010] The conventional clone screening processes of the types described above are extremely resource-intensive, typically taking several months and requiring hundreds or thousands of assays and cell cultures. However, as the pace of biotechnology accelerates and more emphasis is placed on further molecular processing in early-stage pipelines, the need for faster clone screening is increasing. Furthermore, conventional clone screening processes lack standardized criteria for selecting which clones to advance to the next stage / bioprocess and ultimately selecting the winning clone, partly because the unique combination of modality, composition, and sequence characteristics for each different drug candidate means that different factors may be more or less important. [Means for solving the problem]

[0011] The embodiments described herein relate to systems and methods for creating, evaluating, and / or applying performance predictive models for cell lines and bioprocesses in clonal selection. In particular, they are used to create robust machine learning models that improve performance while reducing development timelines and resource usage.

[0012] In one embodiment, one or more machine learning algorithms can be used to predict the performance of each clone and all clones in a hypothetical scale-up (bioreactor) culture based on measurements and other data regarding small-scale real-world cultures of these same clones. Large-scale culture performance can be predicted for a hypothetical / virtual culture duration (e.g., 15 days of culture), and each prediction can be made almost instantly. Depending on the embodiment, this process may result in the selection of better clones / cell lines for scale-up experiments (i.e., clones that are more likely to perform better in large-scale cultures), or it may result in the selection of a "winner" clone without any scale-up experiments (e.g., by selecting the clone with the best predicted bioreactor performance), which may shorten the critical path of a biopharmaceutical program by more than a month.

[0013] Using the predictive models described herein, clones with higher productivity and / or higher quality can be identified compared to the conventional "funnel" approach (i.e., from Stage 12 to Stage 14 and then proceeding to Stage 16 in Figure 1). This improvement arises because small-scale results, despite some similarities, do not fully represent the results of the scale-up. In other words, simply selecting the clone with the best productivity and / or best product quality according to some predetermined criteria in Stage 12 does not necessarily result in the best productivity and / or best product quality in Stage 14 (according to the same criteria).

[0014] Furthermore, interpretable machine learning algorithms can be used to identify the most important input features (e.g., measurements of small cultures) for achieving accurate predictions. This can be useful considering that a very large number of attributes (e.g., over 600) may be tracked in any given clone screening program. Therefore, it may be possible to make sufficiently accurate predictions using a relatively small number of input features (e.g., about 10 features), eliminating the need to measure many other attributes. Knowledge of the correlation between measurements and the desired predictive target can provide scientific insights and may also give rise to hypotheses for further research that could lead to improvements in future bioprocesses.

[0015] In another embodiment, in addition to or instead of the above process, one or more machine learning algorithms may be used to select which clones should proceed from the subcloning stage to a small-scale screening culture (e.g., stages 11 to 12 in Figure 1). Typically, clones that have both a high cell productivity score and a large number of cells at the end of the subcloning stage have been considered the best candidates to achieve high performance in a small-scale screening culture (fed-batch experiment). This approach typically results in approximately 30–100 clones advancing to the fed-batch stage. However, the machine learning algorithms described herein can improve this process by analyzing various attributes of candidate clones in both the subcloning stage and the preceding cell pool stage and predicting specific product quality attributes (e.g., titer, cell proliferation, or specific productivity) that will result from a virtual small-scale (e.g., fed-batch) culture experiment. The microtiter plate-based method for clone generation and proliferation (i.e., subcloning stage 11 in Figure 1) can be replaced by the use of more efficient, high-throughput, and high-content screening tools, such as the Berkeley Lights Beacon® photoelectron cell line generation and analysis system. After predicting product quality attribute values ​​for candidate cell lines, the candidates are ranked according to the predicted values, thereby facilitating the selection of a smaller subset of candidate clones for the next stage of cell line development. Advantageously, the ranking created according to these values ​​may be highly accurate in a particular machine learning model, even if the underlying predictions show relatively low accuracy and therefore appear insufficient on the surface. Depending on the embodiment, this process may require less resource use (e.g., in terms of time, cost, effort, equipment, etc.) and / or provide better standardization when selecting candidate clones / cell lines for small-scale screening cultures (i.e., clones that are more likely to exhibit the best performance in small-scale cultures). For example, reducing the number of cells advanced to the fed-batch stage may free up the ability to test other cell lines for other drug products.In some embodiments, the small-scale screening stage can be completely skipped based on the ranking of various cell lines (for example, by proceeding directly from stage 11 to stage 14 of process 10).

[0016] Those skilled in the art will understand that the drawings described herein are included for illustrative purposes only and do not limit the disclosure. The drawings are not necessarily to scale and instead focus on illustrating the principles of the disclosure. In some cases, various aspects of the embodiments described may be exaggerated or enlarged to facilitate understanding of the embodiments described. In the drawings, similar reference numerals appearing throughout the drawings refer to components that are generally functionally and / or structurally similar. [Brief explanation of the drawing]

[0017] [Figure 1] This illustrates the various stages of a typical clonal screening process. [Figure 2] This is a simplified block diagram of an exemplary system capable of carrying out the method of the first aspect of the present invention as described herein. [Figure 3] This is an illustrative process flowchart for generating machine learning models tailored to specific use cases. [Figure 4A] This demonstrates the exemplary performance of various models in a variety of different use cases. [Figure 4B] This demonstrates the exemplary performance of various models in a variety of different use cases. [Figure 5A] Exemplary feature importance metrics are presented for various different use cases and models. [Figure 5B] Exemplary feature importance metrics are presented for various different use cases and models. [Figure 5C] Exemplary feature importance metrics are presented for various different use cases and models. [Figure 5D] Exemplary feature importance metrics are presented for various different use cases and models. [Figure 6A] Screenshots provided by an exemplary user interface for setting parameters of each use case and analyzing predicted outputs are shown. [Figure 6B] Screenshots provided by an exemplary user interface for setting parameters of each use case and analyzing predicted outputs are shown. [Figure 7] It is a flowchart of an exemplary method for facilitating the selection of a master cell line from candidate cell lines that produce recombinant proteins. [Figure 8] It is a simplified block diagram of an exemplary system that can implement the method of the second aspect of the invention described herein. [Figure 9] It is an exemplary graphic output showing the relationship between cell number and cell productivity score for cell line selection. [Figure 10] An exemplary process for generating and evaluating a machine learning model is shown. [Figure 11A] An exemplary output from a regression estimator that can be used for feature reduction is shown. [Figure 11B] An exemplary output from a regression estimator that can be used for feature reduction is shown. [Figure 12A] Model performance and / or feature importance observed for various models and target product quality attributes are shown. [Figure 12B] Model performance and / or feature importance observed for various models and target product quality attributes are shown. [Figure 12C] Model performance and / or feature importance observed for various models and target product quality attributes are shown. [Figure 12D] Model performance and / or feature importance observed for various models and target product quality attributes are shown. [Figure 12E] Model performance and / or feature importance observed for various models and target product quality attributes are shown. [Figure 12F]This shows the observed model performance and / or feature importance for various models and target product quality attributes. [Figure 12G] This shows the observed model performance and / or feature importance for various models and target product quality attributes. [Figure 13A] This section compares rankings based on real-world fed-batch cultures with model-predicted rankings. [Figure 13B] This section compares rankings based on real-world fed-batch cultures with model-predicted rankings. [Figure 13C] This section compares rankings based on real-world fed-batch cultures with model-predicted rankings. [Figure 14] This is a flowchart illustrating an exemplary method for facilitating the selection of cell lines to proceed to the next cell line screening stage from among multiple candidate cell lines that produce recombinant proteins. [Modes for carrying out the invention]

[0018] The various concepts introduced above and discussed in more detail later can be implemented in any of many ways, and the concepts described are not limited to any particular mode of implementation. Examples of embodiments are provided for illustrative purposes only.

[0019] Figure 2 is a simplified block diagram of an exemplary system 100 capable of carrying out the method of the first embodiment described herein. The system 100 includes a computing system 102 that is communicably connected to a training server 104 via a network 106. Generally, the computing system 102 is configured to use one or more machine learning (ML) models 108 trained by the training server 104 to predict the large-scale (bioreactor) cell culture performance (e.g., productivity and / or product quality attributes) of a particular cell line based on small-scale culture measurements of that cell line and optionally on other parameters (e.g., modality).

[0020] Network 106 may be a single communication network or may include one or more types of communication networks (e.g., one or more wired and / or wireless local area networks (LANs) and / or one or more wired and / or wireless wide area networks (WANs) such as the Internet). In various embodiments, the training server 104 trains and / or uses the ML model 108 as a “cloud” service (e.g., Amazon Web Services), or the training server 104 may be a local server. However, in the illustrated embodiment, the ML model 108 is trained by the server 104 and, if necessary, transferred to the computing system 102 via Network 106. In other embodiments, one, some, or all of the ML models 108 may be trained on the computing system 102 and then uploaded to the server 104. In yet another embodiment, the computer system 102 trains and maintains / stores the model 108, in which case the system 100 may omit both Network 106 and the training server 104.

[0021] Figure 2 illustrates a scenario in which the computing system 102 makes predictions based on measurements of a specific small cell culture 110. The culture 110 may be a culture of a specific cell line (e.g., derived from Chinese hamster ovary (CHO) cells) in a single container such as a well or vial. The cell line of culture 110 may be any suitable cell line that produces recombinant proteins and may be of any particular modality. The cell line may be, for example, a monoclonal antibody (mAb) producing cell line or a cell line that produces bispecific or other multispecific antibodies. It will also be understood that the computing system 102 may make predictions based on measurements of cells cultured in a microfluidic environment such as an optoelectronic device as described herein.

[0022] One or more analytical instruments 112 are collectively configured to acquire physical measurements used by the computing system 102 to make predictions, as will be discussed further later. The analytical instruments 112 can acquire measurements directly and / or indirect or “soft” sensor measurements, or can facilitate their acquisition. As used herein, the term “measurement” may refer to a value directly measured / detected by an analytical instrument (e.g., one of the instruments 112), a value calculated by an analytical instrument based on one or more direct measurements, or a value calculated by another device (e.g., the computing system 102) based on one or more direct measurements. The analytical instruments 112 may include fully automated instruments and / or instruments requiring human assistance. As merely one example, the analytical instrument 112 may include one or more chromatographs (e.g., instruments configured to perform size exclusion chromatography (SEC), cation exchange chromatography (CEX), and / or hydrophilic interaction chromatography (HILIC)), one or more instruments configured to obtain measurements for determining the titer of a target product, and one or more devices configured to directly or indirectly measure the metabolite concentrations of a culture medium (e.g., glucose, glutamine, etc.).

[0023] The computer system 102 may be a general-purpose computer specifically programmed to perform the operations discussed herein, or it may be a dedicated computing device. As can be seen from Figure 2, the computing system 102 includes a processing unit 120, a network interface 122, a display 124, a user input device 126, and a memory unit 128. However, in some embodiments, the computing system 102 includes two or more computers that are located together or far apart from each other. In these distributed embodiments, the operations described herein relating to the processing unit 120, the network interface 122, and / or the memory unit 128 may be divided among multiple processing units, network interfaces, and / or memory units, respectively.

[0024] The processing unit 120 includes one or more processors, each of which may be a programmable microprocessor that executes software instructions stored in the memory unit 128 to perform some or all of the functions of the computing system 102 as described herein. The processing unit 120 may include, for example, one or more central processing units (CPUs) and / or one or more graphics processing units (GPUs). Alternatively or in addition, some of the processors in the processing unit 120 may be other types of processors (e.g., application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), etc.), and some of the functions of the computing system 102 as described herein may be implemented in hardware instead. The network interface 122 may include any suitable hardware (e.g., front-end transmitter and receiver hardware), firmware and / or software configured to communicate with the training server 104 over the network 106 using one or more communication protocols. For example, the network interface 122 may be or include an Ethernet interface that enables the computing system 102 to communicate with the training server 104 over the internet or an intranet, etc.

[0025] The display 124 may use any suitable display technology (e.g., LED, OLED, LCD, etc.) to present information to the user, and the user input device 126 may be a keyboard or other suitable input device. In some embodiments, the display 124 and the user input device 126 are integrated within a single device (e.g., a touchscreen display). Generally, the display 124 and the user input device 126 may be combined to enable the user to interact with a graphical user interface (GUI) (e.g., an interface as described later with reference to Figures 6A and 6B) provided by the computing system 102. However, in some embodiments, the computing system 102 does not include the display 124 and / or the user input device 126, or one or both of the display 124 and the user input device 126 are included in another computer or system (e.g., a customer's device) that is communicatively coupled to the computing system 102.

[0026] The memory unit 128 may include one or more volatile and / or non-volatile memories. It may include one or more of any suitable memory types, such as read-only memory (ROM), random access memory (RAM), flash memory, solid-state drives (SSDs), and hard disk drives (HDDs). Collectively, the memory unit 128 may store one or more software applications, data received / used by those applications, and data output / generated by those applications. These applications, when executed by the processing unit 120, include a large-scale predictive application 130 that predicts the performance (e.g., productivity and / or product quality attributes) of a particular cell line in a virtual / hypothetical large-scale culture based on small-scale measurements obtained by the analytical instrument 112 (and possibly on other information such as modality). The various modules of application 130 are described below, and it will be understood that these modules may be distributed among different software applications, and / or the functionality of any one such module may be divided among two or more software applications.

[0027] The data acquisition unit 132 of application 130 collects values ​​of various attributes related to small cell cultures, such as culture 110. For example, the data acquisition unit 132 can receive measurements directly from the analytical instrument 112. In addition or alternatively, the data acquisition unit 132 can receive information stored in a measurement database (not shown in Figure 2) and / or information entered by the user (e.g., via the user input device 126). For example, the data acquisition unit 132 may receive modality, target drug product, drug protein scaffold type and / or any other appropriate information entered by the user and / or stored in the database. In addition or alternatively, the data acquisition unit may receive measurements from an optoelectronic device as described herein.

[0028] For a given small cell culture corresponding to a specific cell line, the prediction unit 134 of application 130 operates based on attribute values ​​collected by the data acquisition unit 132 and uses a local machine learning model 136 to output one or more predicted attribute values ​​corresponding to a hypothetical large-scale culture. That is, the attribute values ​​collected by the data acquisition unit 132 are used as input / features for the machine learning model 136. The attributes whose values ​​are predicted may include one or more productivity metrics (e.g., titer) and / or one or more product quality metrics (e.g., SEC main peak, low molecular weight peak and / or high molecular weight peak percentage). In the illustrated embodiment, the machine learning model 136 is a local copy of model 108 trained by the training server 104 and can be stored, for example, in the RAM of the memory unit 128. However, as described above, the server 104 may utilize all models 108 in other embodiments, in which case the local copy does not need to reside in the memory unit 128.

[0029] The visualization unit 138 of application 130 generates a user interface that allows the user to input information indicating a use case (e.g., which large attribute values ​​to predict, modalities, etc.) via the user input device 126, and allows the user to observe a visual representation of the predictions made by the prediction unit 134 (and / or other information derived therefrom) via the display 124. Exemplary screenshots of the user interface that can be generated by the visualization unit 138 are described later with reference to Figures 6A and 6B.

[0030] The operation of System 100 according to one embodiment will be described in more detail here with respect to a specific scenario (in which Application 130 is used to predict large-scale performance for a number of different cell lines (clones) in a small cell culture, including a specific cell line in a small cell culture 110). By doing so, a better selection of cell lines can be identified for scale-up (for example, for stage 14 in process 10 in Figure 1), or the scale-up stage can be skipped entirely (for example, by directly passing from stage 12 to stage 16 of process 10 based on predictions for various cell lines).

[0031] First, the training server 104 trains the machine learning model 108 using data stored in the training database 140. The machine learning model 108 may include a number of different types of machine learning-based regression estimators (e.g., decision tree regressor models, random forest regressor models, linear support vector regression models, extreme gradient boosting (xgboost) regressor models, etc.) and, optionally, one or more non-regression-based models (e.g., neural networks). Furthermore, in some embodiments, the model 108 may include two or more models of any given type (e.g., two or more models of the same type trained on different historical datasets and / or using different feature sets). In addition, different models of the model 108 may be trained to predict different large-scale culture attribute values ​​(e.g., titer or chromatographic SEC values, etc.). As will be discussed further later with reference to Figures 4A and 4B, each of the machine learning models 108 may be optimized (trained and tuned) for a particular specification case or for a particular class of specification cases. Furthermore, as will be discussed later with reference to Figures 5A-5D, each of the machine learning models 108 may be used to identify which features (e.g., small culture attribute values) best predict a particular large culture attribute value, and / or may be trained or retrained using a feature set that includes only the features that best predict the particular large culture attribute value.

[0032] The training database 140 may include a single database stored in a single memory location (e.g., HDD, SSD, etc.) or multiple databases stored in one or more memory locations. For each different model within the machine learning model 108, the training database 140 may store corresponding sets of training data (e.g., input / feature data and corresponding labels), which may overlap between training datasets. For example, to train a model that predicts titer, the training database 140 may include numerous feature sets (each of which includes historical small-scale culture measurements and possibly other information (e.g., modality) performed by one or more analytical instruments (e.g., analytical instrument 112 and / or similar instruments)) along with labels for each feature set. In this example, the labels for each feature set represent the large-scale culture titer value (e.g., endpoint titer at day 15) measured when the small-scale culture cell line was scaled up in a bioreactor. In some embodiments, all features and labels are numerical, and non-numerical classifications or categories are mapped to numerical values ​​(for example, the tolerance values ​​for modality features / inputs [Bispecific Format 1, Bispecific Format 2, Bispecific Format 1 or 2] are mapped to the values ​​[10, 01, 00]).

[0033] In some embodiments, the training server 104 uses additional labeled datasets within the training database 140 to validate the trained machine learning models 108 (for example, to ensure that one given machine learning model 108 provides at least a minimum acceptable accuracy). The validation of the models 108 will be discussed further later with reference to Figure 3. In some embodiments, the training server 104 also continuously updates / improves one or more machine learning models 108. For example, after the machine learning models 108 have been trained to initially provide a sufficient level of accuracy, additional measurements, both small (feature) and large (label) measurements, may be used to improve the predictive accuracy.

[0034] Application 130 may retrieve a specific machine learning model 108 corresponding to a desired use case from the training server 104 via network 106 and network interface 122. The use case is, for example, indicated by the user via a user interface (see, for example, Figure 6A, as described later). Once the model is retrieved, the computing system 102 saves a local copy as a local machine learning model 136. In other embodiments, the model is not retrieved as described above; instead, the input / feature data is sent to the training server 104 (or another server) as needed to use the appropriate model of model 108.

[0035] The data acquisition unit 132 collects the necessary data according to the feature set used by model 136. For example, the data acquisition unit 132 may communicate with the analytical instrument 112 to collect measurements of titer, chromatographic values, metabolite concentrations, and / or other specific attributes of the small cell culture 110. In one such embodiment, the data acquisition unit 132 sends commands to one or more analytical instruments 112 to cause one or more instruments to automatically collect the desired measurements. In another embodiment, the data acquisition unit 132 collects measurements of the small cell culture 110 by communicating with a different computing system (not shown in Figure 2) coupled to (and optionally controlling) the analytical instrument 112. As described above, the data acquisition unit 132 may also receive information entered by the user (e.g., modality, targeted drug product, etc.). In some embodiments, some user input information collected by the data acquisition unit 132 is used to select a suitable one of the models 108, while other user input information collected by the data acquisition unit 132 is used (or used to obtain) one or more features / inputs to the selected model.

[0036] After the data acquisition unit 132 collects attribute values ​​related to small cell cultures 110 (and optionally other data such as targeted drug products) and uses them as input / features by the local machine learning model 136, the prediction unit 134 causes the model 136 to operate based on those inputs / features to output predictions of one or more large cell culture attribute values ​​for the same cell line. It should be understood that in some embodiments and / or scenarios, the local machine learning model 136 may include two or more models, each predicting / outputting different large cell culture attribute values.

[0037] The large-scale culture attribute values ​​output by Model 136 may include, for example, one or more productivity attributes such as titer or viable cell density (VCD), and / or one or more product quality attribute values ​​such as SEC main peak (MP) percentage, SEC low molecular weight (LMW) peak percentage, and / or SEC high molecular weight (HMW) peak percentage. The visualization unit 138 causes the user interface displayed on the display 124 to show the predicted attribute values ​​and / or other information derived from the predicted attribute values. For example, the visualization unit 138 may cause the user interface to show whether the predicted attribute values ​​meet one or more cell line selection criteria (for example, after application 130 has compared the attribute values ​​to one or more respective thresholds).

[0038] The above process can be repeated for a large number of different cell lines, each of which can be used for small cell cultures similar to small cell culture 110. For example, a computing system 102 (or another computing system not shown in Figure 2) may cause the analytical instrument 112 to continuously acquire measurements from hundreds or thousands of small cell cultures (each containing a different clone / cell line), and a prediction unit 134 may cause the model 136 to operate on each set of measurements (and possibly other data) to output a large-scale prediction or set of predictions.

[0039] The prediction unit 134 can store the predictions made by model 136 for each cell line and / or the information obtained from each set of predictions in the memory unit 128 or another suitable memory / location. After predictions have been made and saved for all cell lines under consideration, a “winning” cell line may be selected (e.g., similar to stage 16 in Figure 1). The selection of the winning cell line may be fully automated according to several criteria specific to the use case (e.g., by assigning specific weights to productivity and product quality attributes and then comparing scores), or it may involve human interaction (e.g., by simply displaying the predicted large attribute values ​​to the user via display 124). Alternatively, after predictions have been made and saved for all cell lines under consideration, a subset of cell lines may be selected for scale-up (e.g., similar to stage 14 in Figure 1). Again, this selection may be fully automated according to several criteria specific to the use case, or it may involve human interaction.

[0040] As described above, the training server 104 can train a number of different predictive models 108 that are particularly well-suited to specific specification cases or specific class specification cases. Furthermore, interpretable machine learning models can be used to avoid the time and cost of having to perform and collect a very large number of small analytical measurements (and possibly other information). For example, the training server 104 could train one of the models 108 for hundreds of features (e.g., about 600 features), and then the training server 104 (or a human reviewer) could analyze the trained model (e.g., the weights assigned to each feature) to determine the most predictive features (e.g., about 10 features). A new version of that particular model, or a version of that model trained using only the most predictive features, can then be used with a much smaller set of features. Identifying highly predictive features can also be useful for other purposes, such as providing new scientific insights that may give rise to new hypotheses (which could then lead to improvements in bioprocesses).

[0041] Various techniques for determining which model is best suited to a particular use case and for identifying the most predictive features for a given model or use case are described below with reference to Figures 3-5.

[0042] Generally, a model that performs well for a particular use case can be identified by training many different models using historical training data generated from previous clonal screening operations and comparing their results. Historical data may include small-scale cell line development data (e.g., small-scale fed-batch measurement data) and scaled-up bioreactor data (e.g., perfusion bioreactor measurements). Furthermore, historical data may include both categorical data such as culture medium type and modality, and numerical data such as metabolite concentrations and potency values. For small-scale cell line development data (also referred to herein simply as “cell line development data” or “CLD data”), growth factors such as viability, VCD, and glucose concentration can be collected periodically over time (e.g., on different days of a 10-day culture). For scaled-up bioreactor data (also referred herein as “bioprocess development data” or “BD data”), these attributes, along with additional attributes such as pH level and dissolved oxygen concentration, may be collected and recorded in relation to each feature set. Bioreactor data may also include data that serves as labels for various feature sets, such as product titers and other analytical results from assays (e.g., SEC and / or CEX analysis results). Various measures can be taken to ensure a robust training dataset (e.g., providing standardized heterogeneous data, removing outliers, assigning missing values, etc.).

[0043] In some embodiments, special feature engineering techniques are used to extract or derive useful features. For example, convolutional neural networks (or APIs that automatically extract summary statistics from time data, such as tsfresh) can be used to detect time dependencies between various attributes (e.g., a high correlation between the VCD on day 0 of a small culture and the VCD on day 6 of a small culture). These time dependencies can be used to extract / derive useful features for model training. Other feature engineering techniques may also be used, such as variance thresholding, principal component analysis (PCA), mutual information regression, analysis of variance (ANOVA), and removal of features with high covariance.

[0044] In any supervised machine learning regression model generated using historical training data, the task is to predict values ​​from input / feature data x.

number

number

number

[0045] Figure 3 shows a modular and flexible process 200 that can be used as a framework for identifying models that perform well for each of numerous different use cases. First, in stage 202, relevant data corresponding to a given use case is selected from the available historical data. A “use case” can be defined in various ways depending on how it is determined which data is relevant to that use case. For example, a use case may be defined as a specific target variable (y), a specific modality or set of modalities, and possibly one or more specific restrictions on a feature dataset. As a more specific example, a use case may correspond to (1) endpoint titer for a large culture (bioreactor) as the target variable, (2) all modalities (e.g., monoclonal antibodies and possibly bispecificity or multispecificity formats), and (3) using only cell line development history data as features for (and / or for derivation of) the training data. Conversely, other use cases may include (1) chromatographic analysis results for large cultures as target variables (e.g., SEC main peak), (2) a single modality (e.g., a specific monoclonal antibody or a bispecific or multispecific antibody format), and (3) the use of both cell line development history data and bioreactor history data as features of training data (and / or for derivation).

[0046] The model library for use cases is registered in Stage 204. Stage 204 involves selecting a number of candidate machine learning models / estimators that may or may not prove particularly suitable for predicting target attribute values ​​for use cases. To obtain accurate and interpretable results, some or all of the machine learning models selected in Stage 204 should meet two criteria. First, machine learning models that can assign weights to input features are preferred because they can explain the relative importance of each input feature in predicting the target output. Second, sparsity-induced machine learning models are preferred (e.g., models that initially accept many attribute values ​​as features but require only a small subset of these attribute values ​​as features to make accurate predictions). This characteristic improves interpretability while reducing overfitting by eliminating features that do not significantly affect the target outcome. Sparsity-induced models can also save time and cost because there is no need to measure the excluded attribute values. Regression models / estimators based on decision trees (e.g., decision / ID tree models, random forest models, xgboost models, gradient boosting models, etc.) or other machine learning algorithms (e.g., support vector machines (SVMs) with linear basis function kernels and / or radial basis function kernels, elastic networks, etc.) are particularly well-suited to satisfy both of the above criteria. In some embodiments, one or more neural networks can also be selected at stage 204, although this is not conventionally considered interpretable.

[0047] In Stage 206, the machine learning pipeline is designed to train each model considered to be for a use case (i.e., each model selected for the library in Stage 204). For example, Stage 206 may involve performing k-fold validation for each model (e.g., k=10 if the model is trained and evaluated 10 times across different 90 / 10 partitions of the dataset selected in Stage 202). Within the machine learning pipeline, the dataset selected in Stage 202 may first be transformed by standard scaling, such as normalizing the mean of each feature to zero (μ=0) and the standard deviation to 1 (σ=1). This allows each feature's importance to be considered on an equal standard without bias due to the unequal magnitudes of the raw values ​​corresponding to different features.

[0048] After normalization, the model's hyperparameters are tuned. For example, Bayesian search techniques can be used to tune the hyperparameters. This technique performs a computationally more efficient Bayesian-guided search than grid search or random search, while achieving a similar level of performance to random search. Relatively simple algorithms, such as non-boosting and non-neural network algorithms, may use a relatively small number of iterations of Bayesian search (e.g., 10), while more complex algorithms, such as gradient boosting, xgboost, and neural network algorithms, may use a relatively large number of iterations of Bayesian search (e.g., 30) for higher-dimensional search spaces. Hyperparameters can be selected through k-fold validation (e.g., k=5). Each trained model with tuned hyperparameters is then evaluated using the test dataset. The coefficient of determination (R) is calculated for each model. 2 Algorithmic performance metrics such as ) and root mean square error (RMSE) can be obtained. RMSE can be calculated as follows:

number

number

[0049] In Stage 208, the best model for a use case is selected according to several criteria. For example, the "best" model may be the one that has the lowest mean RMSE across 10 cross-validation splits after 90 / 10k split validation, among all the models used to register the model library in Stage 204 and trained in Stage 206 (by Equation 3 above). RMSE is used to avoid the tendency to compare model performance across use cases with a singular normalization metric, and is therefore used in R 2 It could be a better metric. Furthermore, R 2 The metric may, in some cases, take extremely negative values ​​in several cross-validation sets, which can distort the dynamics of model comparisons when averaged. RMSE may be used more than mean absolute error (MAE) to penalize larger errors between predictions and actuals.

[0050] Subsequently, in Stage 210, the final production model for the use case is output. The final production model may be of the same type as the model selected in Stage 208, but may be retrained on the entire dataset selected in Stage 202 to obtain better (e.g., optimal) hyperparameters. By training on the entire dataset, the final production model can generalize better and may exhibit a similar or higher level of average accuracy compared to that obtained during cross-partitioning validation. The final production model is then saved as a trained model and ready to make predictions for new experiments.

[0051] In one embodiment, process 200 is performed by the training server 104 in Figure 2 (with human input at various stages, including, optionally, defining use cases and / or registering candidate models in the model library). Process 200 can be repeated for each use case and for any appropriate number of use cases (e.g., 5, 10, 100). Once final production models for different use cases are output at each iteration of stage 210, the training server 104 may add these final production models to the machine learning model 108. Subsequently, and before making predictions for various clones / cell lines of small cell cultures (e.g., culture 110) in the manner discussed above with reference to Figure 2, the computing system 102 or the training server 104 may select an appropriate final production model from the model 108. This selection can be based on user input indicating use cases (e.g., as described later with reference to Figure 6A) and on an algorithm or mapping (e.g., performed by application 130) that matches user-specified use cases to final production models. Alternatively, if no exact match exists, such an algorithm may match the user-specified use case to the final production model of Model 108 that best fits the user-specified use case (for example, by mapping categorical parameters such as modality to numerical values ​​and calculating the vector distance between the numerical parameters that define the use case).

[0052] As mentioned above, reducing the number of features required for a particular model can be advantageous. Therefore, if the "best" model from stage 208 is retrained in stage 210, only the features that best predict the desired output (e.g., titer) may be utilized. To identify smaller feature sets, process 200 can perform recursive feature removal (RFE), which allows for the recursive reduction of explanatory features used in the final production model, discarding the least important features. The RFE algorithm trains the data by utilizing a subset of features to obtain optimal model performance with respect to the constraints on the number of features. Pairing RFE with sparsity-inducing models / estimators such as decision trees or elastic networks can further reduce the number of explanatory features, at the cost of increased interpretability. Throughout RFE, an elbow plot can be used to determine the "sweet spot" or inflection point between interpretability and accuracy.

[0053] In addition to determining the accuracy of each model in a model library, it can be important to know the prediction interval (also known as the "confidence" interval). For example, a slightly less accurate model may be preferred over a more accurate model if it has a much tighter prediction interval. However, complex machine learning algorithms can generate only point predictions without any intervals. Therefore, in some embodiments, a conformal prediction framework is used. The conformal prediction interval allows for the assignment of error limits to each new observation and can be used as a wrapper for any machine learning estimator. This framework is applicable when it is assumed that the training and test data originate from the same distribution. If this interchangeability condition is met, a subset of the training data can be used to construct a misfit function on which the underlying sample distribution is measured.

[0054] In one embodiment, the “misfit” API is used in conjunction with the induced conformal prediction framework, which allows the model to be trained only once, immediately before prediction intervals are generated in parallel for all new observations. The induced conformal prediction framework requires disparate calibration sets of the training set. This helps in building robust prediction intervals, but removing samples from the training set to build the misfit function reduces the statistical power of the model. A normalization process (e.g., by a KNN-based approach) can be used to generate specific decision boundaries for each prediction.

[0055] The prediction intervals generated by the conformal prediction framework include future observations at a rate equal to 1-α (where α is the significance level), but the width of the generated intervals depends heavily on the underlying function. Naturally, narrower intervals result in greater reliability in point predictions.

[0056] Figures 4A and 4B show exemplary model performance (RMSE across 10 cross-validation intervals) for several different use cases. In all use cases shown, the target variable (attribute value) is either large-scale (bioreactor) endpoint titer or large-scale SEC analysis metrics. Bioreactor endpoint titer may represent the product concentration yield from the cell culture medium (HCCF) collected on the final day of perfusion bioreactor culture (e.g., day 15). This is a weighted average composite titer from the culture supernatant and perfusion permeate. Productivity is assessed using endpoint titer. SEC analysis assesses the chromatographic peak profile of the product based on protein size. The three elution peaks are typically separated into three classifications: low molecular weight (LMW), main peak (MP), and high molecular weight (HMW). High-quality clones ideally have high SEC MP, low SEC LMW, and low SEC HMW. MP represents usable product, LMW represents cleavage clipping, and HMW represents aggregated masses. SEC is one of several core analyses typically used to assess product quality.

[0057] In Figures 4A and 4B, "CLD" refers to the development of a cell line that demonstrates the use of small-scale culture data to train the model, while "BD" refers to the development of a bioprocess that demonstrates the use of large-scale culture data to train the model. For example, the use case "Titer-All Modality-CLD" is one in which the target attribute value is the bioreactor endpoint titer, all modalities (e.g., mAbs and bispecific or multispecific antibodies) are included, and only small-scale culture data is used to train the model. For each model in each plot, the thin horizontal line (with short vertical lines at both ends) represents the total RMSE range over 10-fold cross-validation, the thick horizontal line represents the + / - standard deviation range relative to the RMSE, and the vertical line within the thick horizontal line represents the mean RMSE over all 10-fold cross-validation.

[0058] For example, as shown in Figure 4A, the random forest regressor model provides the lowest mean RMSE for use cases "Titer-All Modality-CLD" and "Titer-Dual Singularity-CLD", the xgboost model provides the lowest mean RMSE for use cases "Titer-mAb-CLD" and "Titer-All Modality-CLD+BD", the decision tree model provides the lowest mean RMSE for use case "Titer-Dual Singularity-CLD+BD", and the SVM (linear kernel) model provides the lowest mean RMSE for use case "Titer-mAb-CLD+BD". As shown in Figure 4B, the xgboost model provides the lowest mean RMSE for use cases "SEC MP-all modality-CLD", "SEC MP-bispecificity-CLD", "SEC MP-mAb-CLD", "SEC MP-all modality-CLD_BD", and "SEC MP-mAb-CLD+BD", while the SVM (linear kernel) model provides the lowest mean RMSE for use case "SEC MP-bispecificity-CLD+BD".

[0059] Although not shown in Figure 4B, similar results can be obtained for SEC HMW and SEC LMW. For SEC HMW target attribute values, the decision tree model provides the lowest mean RMSE for use cases "SEC HMW-all modality-CLD", "SEC LMW-all modality-CLD", "SEC LMW-bispecificity-CLD", and "SEC LMW-all modality-CLD+BD", the xgboost model provides the lowest RMSE for use cases "SEC HMW-bispecificity-CLD", "SEC HMW-mAb-CLD", "SEC HMW-bispecificity-CLD+BD", "SEC HMW-mAb-CLD+BD", and "SEC LMW-bispecificity-CLD+BD", the random forest model provides the lowest RMSE for use case "SEC HMW-all modality-CLD+BD", the elastic network provides the lowest RMSE for use case "SEC LMW-mAb-CLD", and the SVM (linear kernel) model provides the lowest RMSE for use case "SEC LMW-mAb-CLD+BD".

[0060] In some embodiments, the application 130 of the computing system 102 in Figure 2 determines a use case (target attribute value, modality, and dataset type) for a given collection of candidate clones / cell lines based on user input (e.g., input via display 124) and requests one of the corresponding models 108 from the training server 104. For example, model 108 may include all of the “lowest mean RMSE” ​​models shown above, and the server 104 or computing system 102 may store a database associating each of these models with the use case (or multiple use cases) in which the model provided the lowest mean RMSE. The server 104 or computing system 102 can then access the database to select the best model appropriate for the determined use case. In an alternative embodiment, the computing system 102 sends data indicating the use case to the training server 104, and in response, the training server 104 selects one of the corresponding models 108 and sends it to the computing system 102 to save as a local machine learning model 136. In yet another embodiment, as described above, the selected model may be available remotely from the computing system 102 (for example, on the server 104).

[0061] In some cases, the user may want to test two or more use cases to select a winning clone or a set of clones to be scaled up in a bioreactor for further screening. In these cases, application 130 (or a remote server such as server 104) may select and run multiple models, all of which will be used to make large-scale predictions for each clone / cell line. For example, when selecting a winning clone, the user may want to consider both titer and SEC main peak at large scale. Therefore, application 130 may select and / or run a first machine learning model (e.g., a random forest model) for the use case corresponding to the endpoint titer and a second machine learning model (e.g., an xgboost model) for the use case corresponding to the SEC main peak. As another example, when selecting a winning clone, the user may want to consider titer, SEC main peak, SEC low molecular weight, and SEC high molecular weight at large scale, and application 130 may select and / or run a random forest model for titer, an xgboost model for the SEC main peak, and a decision tree model for both SEC low molecular weight and SEC high molecular weight.

[0062] As mentioned above, an interpretable model may be preferable to identify which input / feature best predicts a particular target attribute value. For example, a tree-based learning method may output metrics indicating how important each feature is for the purpose of reducing the model's mean squared error when that feature is used as a node in the decision tree. Furthermore, a coefficient plot may represent normalized directional coefficients that weight each input / feature when predicting the target attribute value.

[0063] Figures 5A–5D show exemplary feature importance metrics for various different use cases and models. Figure 5A shows feature importance and coefficient plots for a model predicting endpoint titer in a large-scale (bioreactor) environment, and Figure 5B shows feature importance plots for titer prediction filtered by modality. From these two plots, it can be seen that "CLD-Titer × SEC Main Peak-Day 10" is consistently a high-importance feature for models derived using only CLD (cell line development) data. It can also be seen that VCD is a particularly important property when predicting titer, rather than specific productivity (indicated as "qp," with units of pg per cell per day). This indicates that having better cell proliferation in the culture is more important than having high specific productivity for the purpose of generating high titer. The term "iVCD" in Figure 5A refers to the integrated VCD, which describes the total amount (cells × days) in the reactor.

[0064] Figure 5C shows feature importance and coefficient plots for a model predicting the endpoint SEC main peak in a large-scale (bioreactor) environment, and Figure 5D shows feature importance plots for SEC main peak predictions filtered by modality. These plots show that modality and modifications to the protein scaffold are important determinants of the SEC main peak. For example, the CLD modality on day 0 (converted to numerical values) has a strong negative correlation with the SEC main peak, indicating that molecules corresponding to the bispecificity format generally have lower predicted SEC main peaks. The term "project" in Figure 5D refers to a specific project, and therefore a specific product.

[0065] In some embodiments, the training server 104 in Figure 2 trains any given model of the machine learning model 108 using N most important features for a particular use case and model (where N is a predetermined positive integer such as 10 or a number exceeding a threshold importance metric for all features), and only these N features are collected by the data collection unit 132 for processing by the local model 136. In some embodiments, N is determined using recursive feature removal (RFE) as described above. Through RFE, the training server 104 may perform multiple iterations of training to reduce the final number of inputs / features used to make predictions. As described above, the ideal number of features (i.e., the number of features used to train various models 108 used in production) can be selected by examining elbow plots that graph the number of features against model performance, with inflection points representing the "sweet spot" between accuracy and interpretability in each of such graphs.

[0066] For the features discussed above, any appropriate attribute may be used (for example, to initially train various models and, if the features are sufficiently important, to train the final production model). A non-exhaustive list of possible attributes / features for both cell line development (CLD) and bioprocess development (BD) datasets is shown in Table 1 below.

[0067] [Table 1]

[0068] [Table 2]

[0069] [Table 3]

[0070] Table 4

[0071] As described above, one or more machine learning models (e.g., Model 108) selected (e.g., by Application 130 or Server 104) to make predictions for large-scale cultures may depend on use cases or a set of use cases entered by the user via a graphical user interface. Figure 6A shows an exemplary screenshot of such a user interface 400, which Application 130 may have presented on, for example, Display 124. As seen in the exemplary embodiment of Figure 6A, the user interface may allow the user to (1) enter two target attributes (i.e., large-scale bioreactor attributes predicted by the corresponding machine learning model), (2) indicate whether the input / features should include only cell line development data or both cell line development and bioprocess development (bioreactor) data, (3) indicate one or more modalities being considered, and (4) indicate the desired prediction / confidence interval. Based on user input, application 130 or server 104 can select an appropriate model from model 108, i.e., the final production model obtained from stage 210 of process 200 for each of the user-instructed use cases, in order to make predictions. In the exemplary screenshot 400, it can be seen that a single set of user inputs may correspond to two use cases (i.e., one for each of two target attributes, and each of those use cases includes the same user-selected dataset and modality). The selected model may be downloaded as a local model (e.g., each similar to model 136) or remain on server 104 for use in a cloud service. User activation of the “Predict” control is detected by application 130 (or server 104), and in response, application 130 (or server 104) causes the model to act on each feature set to predict each large attribute value. In other embodiments, it will be understood that the user interface may provide different user controls than those shown in Figure 6A.

[0072] The predictions made by the selected / applied model can be presented to the user in any appropriate manner. An example of such presentation is shown in screenshot 410 of Figure 6B, which corresponds to an embodiment in which predictions for all clones / cell lines can be shown simultaneously. In Figure 6B, each clone / cell line is plotted as a dark circle on a two-dimensional graph. In the results shown in the exemplary scenario of Figure 6B, a user who desires a clone with a high SEC main peak and high titer would select one or both of the two clones in the upper right corner of the graph as the top clone (or the application 130 would automatically select them instead). In some embodiments, the application 130 also allows the user to toggle the display of the prediction interval for each prediction. Furthermore, in some embodiments, the application 130 allows the user to view feature importance and / or coefficient plots related to various models / predictions (e.g., plots similar to those shown in Figures 5A-5D).

[0073] Figure 7 is a flowchart of an exemplary method 500 that facilitates the selection of a master cell line from among candidate cell lines that produce recombinant proteins. When method 500 executes software instructions of application 130 stored in memory unit 128, it is executed by processing unit 120 of computing system 102 or by one or more processors of server 104, for example (e.g., in the execution of a cloud service).

[0074] In block 502, attribute values ​​related to a small cell culture are received for a specific cell line. At least some of the received attribute values ​​are measurements of the small cell culture (e.g., one or more media characteristics such as endpoint titer, SEC MP, SEC LMW, SEC HMW, VCD, viability, glucose or other metabolite concentration, and / or any other CLD measurements shown in Table 1 above). In some embodiments, attribute values ​​may be received from the optoelectronic equipment described herein. In some embodiments and / or scenarios, other data are also received in block 502, such as user input data (e.g., identifier of a specific cell line, modality of a drug produced using a specific cell line, indication of a drug product produced using a specific cell line, and / or protein scaffold type associated with a drug produced using a specific cell line). Furthermore, in some embodiments, one or more attribute values ​​related to a large cell culture can be received (e.g., in embodiments where a small culture is scaled up to perform a large measurement on day 0, to better predict large performance on day 15 without necessarily performing a large culture for the entire period).

[0075] In some embodiments, the small culture attribute values ​​received in block 502 include measurements obtained on different days of the small culture. For example, the first attribute value may be the potency of the small culture on day 10 (e.g., the endpoint potency of the 10-day culture), and the second attribute value may be the VCD value of the small culture on day 0. In a further example, the third attribute value may be the VCD value of the small culture on day 6, and so on. In other exemplary embodiments, the combination of small measurements may be the same as or similar to the one labeled "CLD" in any of the plots in Figures 5A–5D.

[0076] In block 504, for a given cell line, one or more attribute values ​​associated with a virtual large-scale cell culture are predicted by analyzing the attribute values ​​(and optionally user input data) received in block 502 using at least a machine learning-based regression estimator (e.g., a decision tree regression estimator, a random forest regression estimator, an xgboost regression estimator, a linear SVM regression estimator, etc.). The predicted attribute values ​​may include, for example, titer (e.g., endpoint titer) and / or one or more product quality attribute values ​​(e.g., chromatographic measurements such as SEC main peak, SEC LMW, and / or SEC HMW).

[0077] In block 506, the predicted attribute values ​​and / or whether the predicted attribute values ​​meet one or more cell line selection criteria (e.g., above or below a certain threshold) are presented to the user via a user interface (e.g., the user interface corresponding to screenshot 410 in Figure 6B) to facilitate the selection of desired cell lines for use in drug product manufacturing. For example, the user may proceed directly from such a display to select a “winning” cell line, or use the displayed information to identify which cell lines should be scaled up in a real-world bioreactor for validation and / or further clonal screening (selection of the winning clone is performed in a subsequent stage).

[0078] In some embodiments, Method 500 includes one or more additional blocks not shown in Figure 7. For example, Method 500 may include two additional blocks, both of which take place before block 502: a first additional block that receives use-case data from the user via a user interface (e.g., the user interface corresponding to screenshot 400 in Figure 6A), and a second additional block in which, based on the use-case data, a machine learning-based regression estimator (each of these estimators is designed / optimized for a different use case) is selected from among several estimators (e.g., from among Model 108). For example, the user input data may represent at least one of one or more attribute values ​​related to a virtual large cell culture, represent the modality of the drug to be produced, and optionally also represent other parameters (e.g., parameters indicating the range of a dataset, such as the CLD and BD datasets discussed above).

[0079] In more specific embodiments and scenarios, the user input data illustrating a use case may include data indicating at least one titer related to a virtual large cell culture, and block 504 may include analyzing multiple attribute values ​​using a decision tree regression estimator, a random forest regression estimator, an xgboost regression estimator, or a linear SVM regression estimator (for example, according to the results discussed above in relation to Figure 4A). In another specific embodiment and scenario, the user input data illustrating a use case may include data indicating at least one chromatographic measurement (e.g., SEC main peak) related to a virtual large cell culture, and block 504 may include analyzing multiple attribute values ​​using an xgboost regression estimator (for example, according to the results discussed above in relation to Figure 4B).

[0080] In embodiments where a machine learning-based regression estimator is selected from among several estimators, method 500 may include an additional block in which, for each estimator, the set of features that best predicts the output of the estimator is determined. In such embodiments, block 502 may include receiving only the attribute values ​​that are included in the most predictive set of features.

[0081] Figure 8 is a simplified block diagram of an exemplary system 800 capable of performing the technique of the second embodiment described herein. The system 800 includes a computing system 802 which is communicably connected to a training server 804 via a network 806. Generally, the computing system 802 is configured to use one or more machine learning (ML) models 808 trained by the training server 804 to determine / predict the ranking of candidate cell lines according to one or more product quality attributes (e.g., specific productivity, titer, and / or cell proliferation) in a virtual small-scale screening culture (e.g., fed-batch culture) based on measurements by a cloning (or cell line) generation and analysis system 850 and measurements in one or more cell pools 810.

[0082] Network 806 may be similar to network 106 in Figure 2, and / or training server 804 may be similar to training server 104. In the illustrated embodiment, the machine learning model 808 is trained by training server 804 and then transferred to computing system 802 via network 806 as needed. However, in other embodiments, one, some, or all of the ML models 808 may be trained on computing system 802 and then uploaded to server 804. In other embodiments, computing system 802 trains and maintains / stores the ML models 808, in which case system 800 may omit both network 806 and training server 804. In yet another embodiment, training server 804 provides access to the models 808 as a web service (for example, computing system 802 provides input data that server 804 uses to make predictions using one or more models 808, and server 804 returns the results to computing system 802).

[0083] Each cell pool 810 may be a pool of genetically modified cells (e.g., Chinese hamster ovary (CHO) cells) in a single container such as a well or vial. Cell pool 810 may be any suitable pool of cells scaled up through successive cell passages in selective growth medium that produce recombinant proteins, and may be of any modality. The cells may be, for example, cells that produce recombinant proteins such as monoclonal antibodies (mAbs) or cells that produce recombinant proteins such as bispecific or other multispecific antibodies. However, generally speaking, not all cells in pool 810 are derived from clones.

[0084] One or more analytical instruments 812 are collectively configured to acquire physical measurements of the cell pool 810, which can be used by the computing system 802 to make predictions, as will be discussed further herein. The analytical instruments 812 can acquire measurements directly and / or indirect or “soft” sensor measurements, or can facilitate their acquisition. As stated above, as used herein, the term “measurement” may refer to a value that is directly measured / sensed (e.g., by one of the instruments 812), a value calculated based on one or more direct measurements, or a value calculated by an instrument other than a measuring instrument (e.g., the computing system 802) based on one or more direct measurements. The analytical instruments 812 may be similar to the analytical instrument 112 in Figure 2, for example, a chromatograph or optical sensor as described herein. The analytical instruments 812 may include, for example, one or more instruments specifically configured to measure cell pool viable cell density (VCD), cell pool viability (VIA), time-integrated viable cell density (IVCD), and cell pool specific productivity.

[0085] The cloning and analysis system 850 can be any suitable (preferably high-throughput) subcloning system. In some embodiments, the cloning and analysis system 850 is a Berkeley Lights Beacon system. As can be seen from Figure 8, the system 850 includes an analysis unit 852 and a cell line generation and proliferation unit 854. The cell line generation and proliferation unit 854 may be a culture chip containing a plurality of physically isolated pens perfused by microfluidic channels. Unit 854 may be, for example, an OptoSelect® Berkeley Lights chip. Each pen can receive transgenic cells from a cell pool using a projection pattern that activates a photoconductor, which manipulates the cells by gently flicking them (for example, as provided by the positioning technology of Berkeley Lights OptoElectro®), and contains the cells (and other generated cells of the cell line) through the cell line generation and analysis process.

[0086] The analysis unit 852 of the cell line generation and analysis system 850 is configured to measure the physical properties of cells in the cloning and proliferation unit 854. The analysis unit 852 may include one or more sensors or instruments for directly acquiring measurements and / or for indirectly acquiring or facilitating the acquisition of "soft" sensor measurements. The instruments of the analysis unit 852 may include fully automated instruments and / or instruments that require human assistance. As just one example, the instruments of the analysis unit 852 (e.g., sensors or other instruments integrated into or interfaced with unit 854) may include one or more imaging devices (e.g., cameras and / or microscopes) and associated software configured to directly or indirectly measure cell number or cell proliferation, and one or more instruments configured to directly or indirectly measure cell productivity by performing secretion assays (e.g., diffusion-based fluorescence assays that bind to antibodies produced by cells on the chip, such as a secretion assay using a Spotlight HuIg2 assay (or Spotlight assay)).

[0087] Computing system 802 may be, for example, a general-purpose computer similar to computing system 102. As can be seen in Figure 8, computing system 802 includes a processing unit 820, a network interface 822, a display 824, a user input device 826, and a memory unit 828. The processing unit 820, network interface 822, display 824, and user input device 826 may be similar to, for example, the processing unit 120, network interface 122, display 124, and user input device 126 in Figure 2.

[0088] Memory unit 828 may be similar to memory unit 128 in Figure 2. Collectively, memory unit 828 may store one or more software applications, data received / used by those applications, and data output / generated by those applications. These applications, when executed by processing unit 820, include a small-scale predictive application 830 that ranks candidate cell lines according to one or more product quality attributes (e.g., specific productivity, titer, and / or cell proliferation) in a virtual small-scale screening culture (e.g., stage 12 in Figure 1) based on measurements obtained by analytical instrument 812 and analytical unit 852, and optionally on other information (e.g., modality, cell pool identifier, etc.). Various units of application 830 will be discussed below, but it will be understood that these units may be distributed across different software applications, and / or the functionality of any one such unit may be divided among two or more software applications.

[0089] In some embodiments, computing system 802, training server 804, and network 806 are computing system 102, training server 104, and network 106, respectively, and memory units (128 and 828) store both small-scale prediction applications 830 and large-scale prediction applications 130. That is, the systems (10 and 800) may be capable of predicting both small-scale and large-scale performance, and Figure 8 represents a different use case than that shown in Figure 2.

[0090] The data acquisition unit 832 of application 830 generally collects values ​​of various attributes related to the cell pool 810 and the cell line generation and proliferation unit 854. For example, the data acquisition unit 832 can receive measurements directly from the analytical instrument 812 and / or the analytical unit 852. In addition or alternatively, the data acquisition unit 832 can receive information stored in a measurement database (not shown in Figure 8) and / or information entered by the user (e.g., via the user input device 826). For example, the data acquisition unit 832 can receive modality, target drug product, drug protein scaffold type and / or any other appropriate information entered by the user and / or stored in the database.

[0091] The prediction unit 834 of application 830 generally operates based on attribute values ​​collected by the data acquisition unit 832, and uses a local machine learning model 836 to predict product quality attribute values ​​of virtual small-scale screening cultures of different candidate cell lines, and uses these predicted values ​​to rank the cell lines. In the illustrated embodiment, the machine learning model 836 is a local copy of model 808 trained by the training server 804, which can be stored, for example, in the RAM of the memory unit 828. However, as mentioned above, the server 804 can utilize / run model 808 in other embodiments, in which case the local copy does not need to reside in the memory unit 828.

[0092] The visualization unit 838 of application 830 generates a user interface that presents the user with a ranking (determined by the prediction unit 834). The visualization unit 838 may also allow the user to interact with the data presented from the prediction unit 834 and / or input parameters for a specific prediction or ranking (e.g., selecting product quality attributes according to which predicted performance should be ranked).

[0093] The operation of System 800 according to one embodiment is described in more detail here for a specific scenario in which Application 830 is used to determine the ranking of one or more cell lines according to one or more small-scale culture product quality attributes. By ranking cell lines in this way, a methodology for selecting top cell lines can be standardized, better selections of cell lines for small-scale screening can be identified, or the small-scale screening stage can be skipped entirely (for example, by skipping directly from stage 11 to stage 14 of Process 10 based on the ranking of various cell lines).

[0094] First, the training server 804 trains the machine learning model 808 using data stored in the training database 840. The machine learning model 808 may include a number of different types of machine learning-based regression estimators (e.g., random forest regression models, extreme gradient boosting (xgboost) regression models, linear regression models, ridge regression models, lasso regression models, principal component analysis (PCA) with linear regression models, partial least squares (PLS) regression, etc.) and possibly one or more non-regression-based models (e.g., neural networks). Furthermore, in some embodiments, the model 808 may include two or more models of any given type (e.g., two or more models of the same type trained on different historical datasets and / or using different feature sets). In addition, different models of the model 808 may be trained to predict values ​​of different product quality attributes (e.g., titer, proliferation, or specific productivity, etc.) to facilitate ranking cell lines according to those different product quality attributes (by the prediction unit 834). Furthermore, the machine learning model 808 can be used to identify which feature (e.g., attribute values ​​from the cell pooling stage and / or cloning and analysis stages) best predicts the relative performance of a candidate cell line for each of one or more small culture product quality attributes. Model 808 can also be trained or retrained using a feature set containing only the most predictive features.

[0095] The training database 840 may include a single database stored in a single memory (e.g., HDD, SSD), multiple databases stored in a single memory, a single database stored in multiple memory locations, or multiple databases stored in multiple memory locations. For each different model within the machine learning model 808, the training database 840 may store corresponding sets of training data (e.g., input / feature data and corresponding labels), which may sometimes overlap between training datasets. To train a model to predict the titer of a virtual small culture, for example, the training database 840 may include numerous training datasets along with their labels, each training dataset containing historical measurements of cell pool titer, cell productivity scores, and / or other measurements made by one or more instruments (e.g., analytical instrument 812, analytical unit 852 instruments, and / or other instruments / sensors). In this example, the labels of each training dataset indicate the titer actually measured for that cell line at a small culture stage.

[0096] In some embodiments, the training server 804 uses additional labeled datasets in the training database 840 to validate the trained machine learning models 808 (for example, to ensure that one given machine learning model 808 provides at least a minimum acceptable accuracy). In some embodiments, the training server 804 also continuously updates / improves one or more machine learning models 808. For example, after the machine learning models 808 have been trained to initially provide a sufficient level of accuracy, additional measurements at the cell pool and subcloning stages (features) as well as the small-scale culture stage (labels) may be used to improve predictive accuracy.

[0097] After model 808 has been sufficiently trained, application 830 can retrieve a specific one of the machine learning models 808 (corresponding to a specific product quality attribute for which a ranking of candidate cell lines is desired) from the training server 804 via network 806 and network interface 822. For example, the product quality attribute may include cell proliferation and the machine learning model may include PLS; or the product quality attribute may include specific productivity and the machine learning model may include PCA; or the product quality attribute may include titer and the machine learning model may include a ridge regression model. The product quality attribute may be indicated by the user via a user interface (e.g., via a user interface generated by user input device 826 and display 824 and visualization unit 838) or based on any other appropriate input. Once the model is retrieved, computing system 802 saves a local copy as local machine learning model 836. In other embodiments, instead of retrieving the model as described above, the input / feature data is sent to the training server 804 (or another server) as needed to use the appropriate model of model 808.

[0098] The data acquisition unit 832 collects the necessary data according to the feature set used for model 836. For example, the data acquisition unit 832 may communicate with the analytical instrument 812 and the analytical unit 852 to collect measurements of titer, pool VCD, pool VIA, cell number, cell productivity score, and / or measurements of other specific attributes of the cell pool 810 and / or the cell line generation and proliferation unit 854. In one such embodiment, the data acquisition unit 832 sends commands to one or more of the analytical instruments 812 and analytical unit 852 to cause one or more of the instruments to automatically collect the desired measurements. In another embodiment, the data acquisition unit 832 collects measurements of the cell pool 810 and the cell line generation and proliferation unit 854 by communicating with a different computing system (not shown in Figure 8) connected to (and optionally controlling) the analytical instrument 812 and / or the analytical unit 852. As described above, the data acquisition unit 832 may also receive information entered by the user (e.g., modality). In some embodiments, application 830 uses some user input information collected by data acquisition unit 832 to select one suitable model 808, and uses other user input information collected by data acquisition unit 832 as one or more features / inputs to the selected model (or to compute features / inputs).

[0099] After the data acquisition unit 832 collects attribute values ​​related to the cell pool 810 and the cell line generation and proliferation unit 854, and attribute values ​​used as inputs / features by the local machine learning model 836, the prediction unit 834 operates the model 836 on these inputs / features to predict the values ​​of the desired product quality attribute (e.g., titer, growth, or specific productivity) for each candidate cell line. The prediction unit 834 then compares the predicted values ​​with each other to order / rank the cell lines from best to worst or worst to best. Importantly, while machine learning models may generally have low accuracy in predicting important product quality attributes in small cultures, nonetheless, certain models (e.g., those discussed herein) have been found to be good at predicting relative values, so that the ranking of candidate cell lines is generally accurate, even if the predicted values ​​used for their ranking have low accuracy.

[0100] The visualization unit 838 may cause the user interface presented on the display 824 to display the determined ranking of cell lines. The above process may be repeated by reading different models of model 808 that have been specially trained on one or more other product quality attributes of interest, collecting the inputs / features used by those models (by the data acquisition unit 832), using the models (e.g., by the prediction unit 834) to predict the other product quality attributes for each candidate cell line, and ranking the candidate cell lines according to those other product quality attributes (e.g., by the prediction unit 834). The visualization unit 838 may then cause the user interface to present all of the cell line rankings (e.g., one for titer, one for cell proliferation, and one for specific productivity) to enable the user to make a more informed choice about which cell lines or multiple cell lines should proceed to (or, in some cases, be bypassed) the small-scale culture stage.

[0101] The prediction unit 834 can store the predictions made by model 836 for each set of candidate cell lines and / or their corresponding rankings in the memory unit 828 or another suitable memory / location. For all candidate cell lines under consideration, predictions and / or rankings are made and stored, and for all desired product quality attributes, the "winning" portion of the candidate cell lines may be selected for advancement to the small-scale culture stage (e.g., stage 12 in Figure 1). The selection of winning cell lines can be fully automated according to several criteria specific to the product quality attributes (e.g., by assigning specific weights to titer, cell proliferation, and specific productivity rankings and then comparing the resulting scores), or it can involve human interaction (e.g., by displaying the predicted rankings to the user via display 824). The winning cell lines can then be advanced to the small-scale cell culture stage (e.g., stage 12 in Figure 1), or, in some embodiments, they can be bypassed from the small-scale cell culture stage and advanced to a later stage (e.g., stage 14 in Figure 1).

[0102] In some embodiments, the computing system 802 is configured to identify which cell lines should be subjected to the procedures discussed above, i.e., which cell lines should be used as "candidate" cell lines. For example, the computing system 802 (e.g., application 830 or another application) may analyze the cell count and diffusion assay results (obtained by the data acquisition unit 832 from the analysis unit 852 of the cell line generation and analysis system 850) to determine which cell lines have the best potential and should be pursued for further cell line development and screening. Cell lines with both high cell productivity scores and high cell counts are considered the best candidates for achieving high performance in small-scale screening cultures. Identification of candidate cell lines may be performed automatically by the processing unit 820 or the prediction unit 834, or in combination with a user manually comparing these factors via the user input device 826. Identification may also be strictly manual, in which case the user evaluates the scores shown on the display 824 via the user input device 826 to select which cell lines should be candidates. Figure 9 shows an exemplary graphic output 860 of display 824, illustrating a plot of cell number versus cell productivity score (Spotlight assay score) for cell line selection. Cell lines that the user wishes to select as candidate cell lines are, for example, enclosed by dashed lines. Various techniques for predicting a given product quality attribute ranking for a hypothetical small-scale screening culture and determining which model is best suited to identifying the most predictive features / inputs for a given model and / or product quality attribute are described here with reference to Figures 10–12G.

[0103] Figure 10 shows an example of a modular and flexible process 900 that provides a data preparation and model selection framework. In particular, process 900 can be used as a framework to identify a well-performing model for predicting values ​​of different product quality attributes and facilitating the ranking of cell lines according to those attributes (e.g., by prediction units 834). At a high level, process 900 includes a stage or step 902 for aggregating data, a stage 910 for data preprocessing, and a stage 920 for defining models. Generally, a model that performs well for a particular attribute value can be identified by training many different models using historical training data generated from previous cell line screening operations and comparing their results. For example, the attribute may include cell proliferation and the machine learning model may include PLS; or the attribute may include specific productivity and the machine learning model may include PCA; or the attribute may include titer and the machine learning model may include a ridge regression model. Various measures can be taken to ensure a robust training dataset (e.g., providing standardized heterogeneous data, removing outliers, attributing missing values, etc.). In some embodiments, special feature engineering techniques are used to extract or derive the best representation of the predictor variables in order to improve the effectiveness of the model. To avoid overfitting, feature reduction can be implemented in some embodiments. The model may be evaluated using metrics such as the root mean square error (RMSE) to measure the accuracy of the predicted values, or the Spearman Rho to measure the correctness of the ranking order.

[0104] In step 902, the training server 804 receives data from the training database 840 or any other suitable database. This step may include inputting user input via the user input device 826, where the user defines possible predictor variables and product quality attribute values ​​predicted by a machine learning regression estimator (model). Predictor variables may include cell pool data and data collected by the cell line generation and analysis system. Other embodiments may use other subcloning systems, but the following discussion refers to an example in which Berkeley Lights' Beacon (hereinafter abbreviated as "BLI") is used as the cell line generation and analysis system. Predicted variables can be defined, for example, as data collected during a clonal batch experiment. First, in step 902, appropriate data is selected from the available historical data. Furthermore, the historical data may include both categorical data such as modality and numerical data such as cell number and titer. Cell pool data may include, for example, data on modality, VCD, pool viability, pool titer, pool ratio productivity and pool time integral VCD. Growth factors such as VCD and viability can be collected periodically over time (e.g., on different days of a 10-day culture). Cell line generation and growth data (BLI data) may include data on cell productivity score, BLI specific productivity, cell number, time-integrated VCD, doubling time, etc. Growth factors measured by BLI, such as cell number, can also be collected periodically over time (e.g., on different days after loading into a cloning and growth unit such as unit 854). If these cell lines are advanced to the next stage of cell line development (e.g., stage 12 in Figure 1), small-scale culture data (e.g., fed-batch cultures) reflecting results such as titer, specific productivity and / or cell growth measurements can serve as labels for various feature sets. A non-limiting list of possible attributes / features for both cell pool datasets (pool data) and cell line generation and analysis datasets (BLI data), as well as fed-batch predictor variables, is shown in Table 2 below.

[0105] In the exemplary process 900, the data preprocessing stage 910 includes steps 912-918. In step 912, the training data is evaluated and cleaned, including handling of missing data and outliers. For example, missing records (e.g., pool VCD data for empty pens), zero values ​​(e.g., values ​​not recorded), incomplete datasets (e.g., for scenarios where data collection was not completed from the cell pool for a cell line to the end of a feed-in batch experiment), outliers, and data from non-conclusive experiments may be removed. In some embodiments, when using combined datasets, some data values ​​may need to be adjusted to compensate for instrument variability.

[0106] In step 914, useful features are extracted or derived from the dataset using special feature engineering techniques to find the best representation of the predictor variables to improve the model's effectiveness. The data may be visualized for underlying relationships to determine which feature engineering steps should be evaluated for performance improvement. For example, the best representation of a predictor variable may be (i) a transformation of the predictor, (ii) an interaction between two or more predictors such as a product or ratio, (iii) a functional relationship between predictors, or (iv) an equal re-representation of the predictors. Assay or growth values ​​may be scaled for cells in the same cohort to give an unbiased view of growth and assay scores. From these observations, features may be computed and added to the predictor dataset (e.g., square of cell number, square of pooled titer).

[0107] Step 914 may include converting categorical variables to numerical values. For example, for the modality categorical variable, monoclonal (mAb) modality may be converted to "10", a specific bispecific modality to "00", and so on. In data preprocessing step 916, the training data may be filtered to include only the features selected in steps 912 and 914 above, and further filtered to defined targets / predictors (e.g., fed batch titer, growth, and specific productivity).

[0108] When training and comparing machine learning models, k-fold cross-validation can be used to measure model performance and select optimal hyperparameters. Therefore, in step 918, the training data can be split into training and test datasets for k-fold cross-validation to avoid training and testing on the same samples. For example, the number of splits can be defined by the number of subcloning projects used in the training dataset (e.g., k=6 means the model is trained and evaluated 6 times across 5 / 1 different partitions of the dataset).

[0109] Stage 920 defines the machine learning model and includes steps 922-928. At a higher level, Stage 920 may include setting up the regressor and scaling method (step 922), training the predictive model by running the preprocessed data from Stage 910 through each model in the model library across a range of hyperparameters (step 924), defining and calculating model performance metrics (step 926), and outputting the final production model (step 928).

[0110] An exemplary step 922 registers the model library and sets how each selected regression model is scaled. Preferably, some or all of the machine learning models selected to be tested in step 922 meet two criteria: (i) they provide quantitative output, and / or (ii) they are interpretable (e.g., by providing coefficient weights or feature importance weights). Machine learning models that can assign weights to input features are generally preferred because they can explain the relative importance of each input feature in predicting the target output. Sparsity-inducible machine learning models are also generally preferred (e.g., models that initially accept many attribute values ​​as features, but require only a small subset of these attribute values ​​as features to make accurate predictions). This characteristic improves interpretability while reducing overfitting by eliminating features that do not significantly affect the target outcome. Regression models / estimators based on decision trees (e.g., random forest regression models, extreme gradient boosting (xgboost) regression models) or other machine learning algorithms (e.g., linear regression models, ridge regression models, lasso regression models, principal component analysis (PCA) with linear regression models, or partial least squares (PLS) regression models) may be particularly well-suited to satisfy both of the above criteria. In some embodiments, one or more neural networks may be selected in step 922, although these are not conventionally considered interpretable. Step 922 may also include setting a range of hyperparameters for the selected regression models.

[0111] In an exemplary step 924, a predictive model is trained. For example, step 924 can train the models selected for inclusion in the library on the entire set of feature data preprocessed in steps 912 and 914 for each target product quality attribute of interest, and cross-validate them over the range of hyperparameters defined in step 922. Step 924 may include performing k-fold validation for each model on the dataset defined in step 918.

[0112] An exemplary step 926 calculates performance metrics using the trained models. For each of the k splits, algorithmic performance metrics such as RMSE (for accuracy in predicting target product quality attributes) and / or Spearman's Rho (for ranking accuracy) may be calculated for each of the predictive models trained in step 924. Each trained model with tuned hyperparameters is then evaluated using one of the splits as a test dataset, and the model with the best metric (e.g., best Spearman's Rho or lowest RMSE) for each predicted product quality attribute is selected. The performance metrics from iterative runs can be saved, and the average of the k splits (e.g., 6 splits) can be calculated to compare model performance. The calculation of the RMSE metric is shown in Equation 2 above. Spearman's Rho can be calculated as follows:

number

[0113] Counterintuitively, as mentioned above, the ability of a particular machine learning model to correctly rank cell lines (according to the relative values ​​of product quality attributes predicted by the model) can far outweigh the ability of those models to accurately predict product quality attributes. For example, a particular machine learning model may have relatively low accuracy when predicting the value of a particular product quality attribute in the feedforward stage, but has been found to perform well in predicting values ​​in a relative sense (e.g., in terms of whether the predicted value is greater or less than the value the model predicts for other cell lines). This ability to accurately rank cell lines may be sufficient, as knowing which cell lines to advance to the next development stage is generally more important than accurately and precisely predicting product quality attributes. Therefore, Spearman's Rho may be a preferred metric to calculate in step 926 (rather than, for example, RMSE).

[0114] In step 928, the “best” model is output / identified as the final production model based on the calculated metrics (e.g., the model with the highest Spearman Rho or the lowest RMSE). If the best model is interpretable, step 928 may include determining the importance of each feature when making predictions. For example, step 928 may include determining feature importance based on coefficient weights (e.g., generated by a Lasso regression model) or feature importance weights (e.g., generated by a tree-based model such as xgboost). The outputs from these interpretable models (e.g., a display of parameters reduced by a Lasso sparsity-induced model or a feature importance plot showing how often each variable was split when training the tree of the xgboost model) are analyzed by the training server 804 or a human reviewer (via the visualization unit 838) to determine the most predictive features (e.g., 2-10 features) for each relative ranking of candidate cell lines according to the predicted product quality attribute values. For example, Figure 11A shows an exemplary output from a Lasso regression model when predicting fed-batch titer, showing that pooled titer predicts fed-batch titer better than the cell productivity score (here, the "Spotlight" assay score), and that the cell productivity score predicts fed-batch titer better than the cell count (which has no predictive ability or only very little predictive ability for fed-batch titer). Similarly, Figure 11B shows an example of a feature importance plot of an xgboost regression model predicting fed-batch titer, showing that pooled titer and cell productivity score (Adj_Au) have strong feature importance compared to other features used. The results show that the model is capable of predicting features based on cell number (e.g., the square of the cell number or "CC"), for example. 2It demonstrates that it works just as well without using ''). Subsequently, a new version of that model, trained using only the winning / best model or the most predictive features, can be used with a much smaller set of features. The model is then saved as a trained model (e.g., by training server 804 to model 808) and can be used to make predictions in new experiments (e.g., by prediction unit 834). Identifying highly predictive features can also be useful for other purposes, such as providing new scientific insights that may give rise to new hypotheses (which may then lead to improvements in bioprocesses).

[0115] Any appropriate attribute can be used for the features discussed above (for example, to train various models initially and, if the features are sufficiently important, to train the final production model). A non-restrictive list of possible attributes / features for both the cell pool dataset (pool data) and the cell line generation and analysis dataset (BLI data) is shown in Table 2 below.

[0116] [Table 5]

[0117] [Table 6]

[0118] Figure 12A is a bar graph 934 showing the performance of the best model (output at step 928 of process 900) against baseline performance for product quality attributes of cell proliferation, specific productivity, and titer, using Spearman's Lometrics (here with cross-validation across 6 divisions). Each attribute was measured at the end of the small cell culture process (here at day 10 of the fed-batch experiment). In this example, specific productivity performance "baseline" is a linear regression in cell productivity score, where a higher cell productivity score corresponds to a higher predicted specific productivity. Similarly, proliferation performance baseline is a linear regression in cell number, where a higher cell number corresponds to a higher predicted proliferation, and titer performance baseline is a linear regression in cell productivity score and cell number, where a higher score in both corresponds to a higher predicted titer.

[0119] As seen in Figure 12A, the predictive ability of the machine learning model identified / output in step 928 of process 900 (see Figures 12B-12G for further discussion) surpasses the baseline performance for ranking candidate cell lines across all three target product quality attributes. The greatest gain was observed in the model predicting proliferation ranking, where the model showed a rank correlation of ρ=0.283 compared to baseline ρ=0 (no predictive ability). The model from step 928 showed only a slight improvement in predicting specific productivity, with the rank correlation increasing from ρ=0.468 to baseline ρ=0.492, which may mean that only the cell productivity score can explain the majority of the differences in order in specific productivity rank. The model from step 928 showed a moderate improvement in the performance predicting titer, with the rank correlation increasing from ρ=0.245 to ρ=0.342.

[0120] Different regression estimators in model library 922 have been found to be better suited to predicting different target product quality attribute values. For example, using the model identification / definition procedure outlined in stage 920, computing system 802 can test multiple regression estimators using the dataset defined in stage 910 and cross-validate each regression model across a range of hyperparameters. Figures 12B–12G show examples of the relative performance of different regression estimators when predicting specific performance attribute values, and the respective selected features used to construct each selected model using the feature reduction method described herein with reference to step 928. The regression estimator exhibiting “best” performance was selected from models with the highest mean Spearman Rho across all cell lines after optimizing the relevant hyperparameters (if any). Mean RMSE is also shown in Figures 12B, 12D, and 12F, but this metric was not used in model selection for reasons described elsewhere herein (i.e., the importance of relative / ranking accuracy over absolute accuracy).

[0121] As seen in Table 936 in Figure 12B, the best regression estimator for predicting titer was found to be ridge regression with a hyperparameter lambda equal to 1.3. Four other models came close to this performance: linear regression, lasso regression with lambda equal to 0.001, PCA with two principal components, and PLS with two principal components. Table 938 in Figure 12C shows the two attributes (pooled titer and cell productivity score (Spotlight assay score)) analyzed by the models selected with feature reduction.

[0122] Table 940 in Figure 12D shows that the best predictor of specific productivity was PCA with two principal components. Table 942 in Figure 12E shows the eight attributes analyzed by the model selected by feature reduction. For the first PCA component, pooled titer, cell productivity score (Spotlight assay score), and specific productivity values ​​in the cell line generation and analysis system are more important, while for the second PCA component, scaled values ​​of these metrics (normalization of the different characteristics of each cell line) are more important.

[0123] Table 944 in Figure 12F shows that the best regression estimator for predicting proliferation was found to be the PLS with one principal component. Table 946 in Figure 12G shows the nine attributes analyzed by the model selected by feature reduction. The model generally gave more weight to pooled data than to data collected by the Berkeley Lights system. In particular, pooled titer, pooled IVCD, and pooled viable cell density on days 6 and 8 were the most important, while cell number received a lower weighting.

[0124] In addition to using Spearman's Rho, other measures or visualizations can be used to determine the ranking accuracy of various models. Such evaluations can be expressed, for example, as a comparison between the ranking determined by the model and the actual rank of the same cell line in a real-world fed-batch experiment. This evaluation can also be assessed by indicating, for example, whether the model's ability to capture the top cell lines (e.g., the top 4 cell lines) for each target product attribute in a real-world fed-batch experiment appears near the top (e.g., within the top 50%) of the cell lines ranked by the model results. Figures 13A–13C show examples of such evaluation results. Each of Figures 13A–13C shows six bar graphs, each representing the evaluation result for one of six evaluated datasets. The top 50% of ranked cell lines are shown as white bars, and the bottom 50% of ranked cell lines are shown as shaded bars. For a model that perfectly predicts the ranking, a given bar graph will have all the white bars located to the left (along the x-axis) of all the shaded bars. The height of each bar represents the relative value of the product quality attribute expressed in real-world small-scale cell cultures for each cell line.

[0125] Referring first to Figure 13A, the exemplary result 950 corresponds to the predicted ranking of cell lines according to the titer of the product quality attribute (in this example, the titer measured on day 10 of fed-batch, small cultures). As seen in Figure 13A, a 50% reduction in extrusion (i.e., of cell lines that have progressed to the fed-batch stage) using this model is too aggressive and causes some of the top real-world cell lines to be eliminated. In this example, at least 38 clones must be extruded from dataset 4 to ensure that all of the top 4 clones are selected.

[0126] Figure 13B shows 952 exemplary results corresponding to the predicted ranking of cell lines according to the specific productivity of a product quality attribute (in this example, the specific productivity (qP) of feed-in batches and small cultures at day 10). The model prediction of specific productivity was promising. For example, even if the number of harvests were halved, only one of the top four clones would be lost across all cell lines. The maximum number of clones required to capture the top four clones (from the predicted ranking) was 31, and datasets 5 and 6 identified all of the top four clones within the top eight clones predicted by the model, respectively.

[0127] Figure 13C shows exemplary results corresponding to the predicted ranking of cell lines according to cell proliferation (IVCD at day 10 of fed-batch, small cultures in this example) of product quality attributes. The proliferation model prediction shows that the best indicator is the pool from which the clones originate, rather than proliferation in the cell line generation and proliferation unit. However, as shown by datasets 3 and 5, this model did not predict that some of the top-growing clones would be in the top 50%. Nevertheless, this information is still valuable when compared to a baseline (as measured in the cell line generation and proliferation unit) which has no predictive ability for cell number. Based on the results from dataset 4, a minimum of 37 clones must be harvested to ensure that the top 4 clones are harvested / advanced.

[0128] Figure 14 is a flowchart of an exemplary method 960 for facilitating the selection of cell lines from among candidate cell lines that produce recombinant proteins to proceed to the next cell line screening stage (e.g., stage 12 in Figure 1). Part or all of method 960 may be executed by one or more processors of the processing unit 820 of the computing system 802 or server 804 (e.g., in the execution of a cloud service), for example, by executing software instructions of application 830 stored in memory unit 828.

[0129] In block 962, a photoelectron cell line generation and analysis system (e.g., system 850 in Figure 8) is used to measure a first set of attribute values ​​for several candidate cell lines. The photoelectron cell line generation and analysis system may, for example, perform optical and assay measurements on the candidate cell lines in block 962. In some embodiments, such measurements are performed by measuring at least the cell count and cell productivity score in several physically isolated pens within the photoelectron cell line generation and analysis system, at least in part. In some of these embodiments, block 962 further includes generating cells of the candidate cell lines by using the photoelectron cell line generation and analysis system to move individual cells to different pens of physically isolated pens having at least one or more photoconductors activated by a light pattern, and housing the individual cells in their respective pens through the cell line generation and analysis process. Furthermore, block 962 may include measuring different values ​​of the first set of attribute values ​​on different days of the cell line generation and analysis process. More generally, the first set of attribute values ​​may include any value of an attribute that can be measured by the analysis unit 852, as discussed elsewhere in this specification, and / or any suitable attribute value that can be measured using the photoelectron cell line generation and analysis system.

[0130] In block 964, a second set of attribute values ​​is obtained for the candidate cell line. The second set of attribute values ​​includes one or more attribute values ​​measured in the cell pool screening stage of the candidate cell line. The attribute values ​​measured in block 964 may include, for example, pool titer, VCD, and / or pool viability. In some embodiments and / or scenarios, other attribute values ​​such as values ​​calculated based on one or more direct measurements (e.g., time-integrated VCD, pool ratio productivity, etc.) or values ​​calculated by an instrument other than the measuring instrument (e.g., computing system 802) based on one or more direct measurements, and / or user-input values ​​(e.g., modality) are obtained in block 964 instead or further. In some embodiments, some of the attribute values ​​obtained in block 964 are measurements taken periodically over time (e.g., on different days). For example, the first attribute value may be the VCD value of the cell pool on day 0, the second attribute value may be the VCD value of the same cell pool on day 3, and so on. More generally, the second set of attribute values ​​may include values ​​for any of the attributes related to the cell pool 810, which can be measured by the analytical instrument 812, or which are discussed elsewhere in this specification, and / or may include values ​​for any other appropriate attributes related to the cell pool.

[0131] In block 966, the candidate cell lines are ranked according to product quality attributes associated with the virtual small-scale screening cultures for each candidate cell line. Block 966 includes predicting the values ​​of the product quality attributes for each candidate cell line by analyzing a first set of attribute values ​​measured in block 962 and a second set of attribute values ​​obtained in block 964 using a machine learning-based regression estimator. Block 968 also includes comparing the predicted values ​​(i.e., to rank the candidate cell lines (e.g., from best to worst with respect to the predicted values)). In some embodiments, the predicted values ​​are predicted values ​​of cell proliferation metrics. In other embodiments, the predicted values ​​are titer, specific productivity metrics, or any other suitable indicator of performance in the virtual small-scale culture screening stage. The machine learning-based regression estimator may be any suitable type of regression estimator (e.g., Ridge, Lasso, PCA, PCS, xgboost, etc.). In other embodiments, other types of machine learning models can be used to make predictions in block 966 (e.g., by prediction unit 834) (e.g., a neural network).

[0132] In some embodiments, block 966 includes at least (i) predicting titer for each of several candidate cell lines by analyzing a first set of attribute values ​​and a second set of attribute values ​​using a machine learning-based regression estimator, and (ii) determining a ranking according to titer by comparing the predicted titers. In some of these embodiments, the first set of attribute values ​​includes values ​​based on a cell productivity score (e.g., the score itself or a value derived from the score), and / or the second set of attribute values ​​includes values ​​based on a cell pool titer (e.g., the cell pool titer itself or a value derived from the score). The machine learning-based regression estimator that analyzes these attributes may be, for example, a ridge regression estimator.

[0133] In other embodiments, block 966 includes at least (i) predicting specific productivity metrics for each of several candidate cell lines by analyzing first and second sets of attribute values ​​using a machine learning-based regression estimator, and (ii) determining a ranking according to specific productivity by comparing the predicted specific productivity metrics. In some of these embodiments, the first set of attribute values ​​includes values ​​based on cell productivity score and values ​​based on cell number, and / or the second set of attribute values ​​includes values ​​based on cell pool titer. The machine learning-based regression estimator that analyzes these attributes may be, for example, a PCA regression estimator having two principal components.

[0134] In yet another embodiment, block 966 includes at least (i) predicting cell growth metrics for each of several candidate cell lines by analyzing first and second sets of attribute values ​​using a machine learning-based regression estimator, and (ii) determining a ranking according to cell growth by comparing the predicted cell growth metrics. In some of these embodiments, the first set of attribute values ​​includes values ​​based on cell number, and the second set of attribute values ​​includes values ​​based on cell pool time integral viable cell density (iVCD), values ​​based on cell pool viable cell density (VCD) on different days, and values ​​based on cell pool viability on different days. The machine learning-based regression estimator that analyzes these attributes may be, for example, a PLS regression estimator having one principal component.

[0135] In block 968, the ranking display (e.g., an ordered list, a bar graph, etc.) is presented to the user via a user interface. For example, block 968 may include generating or displaying a GUI (e.g., by a visualization unit 838) and presenting the GUI on a display (e.g., display 824). In one embodiment, the presentation of the display is triggered by sending data indicating the ranking to another computing device or system, which uses the data to display and present the GUI.

[0136] In some embodiments, Method 960 includes one or more additional blocks not shown in Figure 14. For example, Method 960 may include an additional block (e.g., before block 962) in which the performance of a machine learning-based regression estimator is evaluated by calculating the mean Spearman ranking correlation coefficient for at least the machine learning-based regression estimator (e.g., calculated according to Equation 4). As another example, Method 960 may include further blocks in which, based on the ranking determined in block 966, one or more candidate cell lines are advanced to the next stage of cell line screening (e.g., the fed-batch cell culture stage).

[0137] The embodiments of the present invention include the following:

[0138] Embodiment 1. A method for facilitating the selection of a cell line from among multiple candidate cell lines that produce recombinant protein, comprising: measuring a first set of attribute values ​​for the multiple candidate cell lines using a photoelectron cell line generation and analysis system; obtaining a second set of attribute values ​​for the multiple candidate cell lines using one or more processors, wherein the second set of attribute values ​​includes one or more attribute values ​​measured in a cell pool screening stage of the multiple candidate cell lines; determining a ranking of the multiple candidate cell lines according to product quality attributes associated with a virtual small-scale screening culture for the multiple candidate cell lines using one or more processors, comprising (i) predicting the value of the product quality attribute for each of the multiple candidate cell lines by analyzing the first set of attribute values ​​and the second set of attribute values ​​using a machine learning-based regression estimator, and (ii) comparing the predicted values; and presenting the ranking to the user via a user interface.

[0139] Embodiment 2. The method of Embodiment 1, wherein measuring a first set of attribute values ​​using a photoelectron cell line generation and analysis system comprises performing multiple optical and assay measurements on multiple candidate cell lines.

[0140] Embodiment 3. Performing multiple optical and assay measurements on multiple candidate cell lines comprises measuring at least cell number and cell productivity score in multiple physically isolated pens in a photoelectron cell line generation and analysis system, the method further comprising generating cells of multiple candidate cell lines by using a photoelectron cell line generation and analysis system to move individual cells to different pens of multiple physically isolated pens having at least one or more photoconductors activated by a light pattern, and housing the individual cells in their respective pens through a cell line generation and analysis process, the method of Embodiment 2.

[0141] Embodiment 4. The method of Embodiment 3, wherein measuring a first set of attribute values ​​includes measuring a first attribute value corresponding to a first measurement of an attribute; and a second attribute value corresponding to a second measurement of that attribute, the first and second measurements being performed on different days of the cell line generation and analysis process.

[0142] Embodiment 5. Obtaining a second set of attribute values ​​is any one of the methods in Embodiments 1 to 4, comprising receiving one or more of the measured cell pool titer; the measured cell pool viable cell density (VCD); or the measured cell pool viability.

[0143] Embodiment 6. Obtaining a second set of attribute values ​​is one of the methods in Embodiments 1 to 5, which includes receiving attribute values ​​measured on different days of the cell pool screening stage.

[0144] Embodiment 7. One or more product quality attributes include cell proliferation metrics, one of any one of Embodiments 1 to 6.

[0145] Embodiment 8. One or more product quality attributes include one or more of (i) potency or (ii) relative productivity metrics, according to any one of Embodiments 1 to 6.

[0146] Embodiment 9. A method of any one of Embodiments 1 to 8, wherein determining a ranking comprises at least (i) predicting the titer for each of several candidate cell lines by analyzing a first set of attribute values ​​and a second set of attribute values ​​using a machine learning-based regression estimator, and (ii) determining the ranking according to the titer by comparing the predicted titers; the first set of attribute values ​​includes values ​​based on cell productivity scores; and the second set of attribute values ​​includes values ​​based on cell pool titer.

[0147] Embodiment 10. The method of Embodiment 9, wherein predicting titer includes analyzing a first set of attribute values ​​using a ridge regression estimator.

[0148] Embodiment 11. A method in any one of Embodiments 1 to 8, wherein determining a ranking comprises at least (i) predicting specific productivity metrics for each of several candidate cell lines by analyzing first and second sets of attribute values ​​using a machine learning-based regression estimator, and (ii) determining a ranking according to specific productivity by comparing the predicted specific productivity metrics; the first set of attribute values ​​includes values ​​based on cell productivity score and values ​​based on cell number; and the second set of attribute values ​​includes values ​​based on cell pool titer.

[0149] Embodiment 12. The method of Embodiment 11, wherein predicting relative productivity metrics includes using a principal component analysis (PCA) regression estimator having two principal components.

[0150] Embodiment 13. Determining a ranking comprises at least (i) predicting cell growth metrics for each of several candidate cell lines by analyzing first and second sets of attribute values ​​using a machine learning-based regression estimator, and (ii) determining a ranking according to cell growth by comparing the predicted cell growth metrics; the first set of attribute values ​​includes values ​​based on cell number; and the second set of attribute values ​​includes values ​​based on cell pool titer, values ​​based on cell pool time integral viable cell density (iVCD), values ​​based on cell pool viable cell density (VCD) at different days, and values ​​based on cell pool viability at different days, one of the methods in Embodiments 1 to 8.

[0151] Embodiment 14. The method of Embodiment 13, wherein predicting cell proliferation metrics includes using a partial least squares (PLS) regression estimator having one principal component.

[0152] Embodiment 15. Any one of embodiments 1 to 14, further comprising evaluating the performance of a machine learning-based regression estimator by calculating a Spearman's Rho or a mean Spearman's Rho for at least the machine learning-based regression estimator.

[0153] Embodiment 16. Any one of embodiments 1 to 15, further comprising advancing one or more cell lines of a plurality of candidate cell lines to the next cell line screening stage based on the ranking.

[0154] Apparatus 17. The method of Apparatus 16, wherein the next cell line screening stage is a fed-batch cell culture stage.

[0155] Embodiment 18. One or more non-temporary computer-readable media that, when executed by one or more processors of a computing system, store instructions causing the computing system to perform one of the methods of Embodiments 1 to 15.

[0156] Embodiment 19. A computing system comprising one or more processors; and one or more non-temporary computer-readable media for storing instructions that, when executed by the one or more processors, cause the computing system to perform any one of Embodiments 1 to 15.

[0157] Embodiment 20. A method for facilitating the selection of a master cell line from among candidate cell lines producing recombinant proteins, comprising: receiving, by one or more processors of a computing system, a plurality of attribute values ​​related to a small cell culture for a particular cell line, wherein at least some of the plurality of attribute values ​​are measurements of the small cell culture; predicting, by one or more processors, one or more attribute values ​​related to a virtual large cell culture for a particular cell line by analyzing the plurality of attribute values ​​related to the small cell culture using at least a machine learning-based regression estimator, wherein the predicted one or more attribute values ​​include potency and / or one or more product quality attribute values; and causing, by one or more processors, to present to a user via a user interface, either (i) one or more predicted attribute values, and (ii) an indication of whether the predicted one or more attribute values ​​meet one or more cell line selection criteria, in order to facilitate the selection of a master cell line for use in the manufacture of a drug product.

[0158] Embodiment 21. Analyzing multiple attribute values ​​using a machine learning-based regression estimator is a method of Embodiment 20, which includes analyzing multiple attribute values ​​using a decision tree regression estimator.

[0159] Embodiment 22. Analyzing multiple attribute values ​​using a machine learning-based regression estimator is the method of Embodiment 21, which includes analyzing multiple attribute values ​​using a random forest regression estimator.

[0160] Embodiment 23. Analyzing multiple attribute values ​​using a machine learning-based regression estimator is the method of Embodiment 21, which includes analyzing multiple attribute values ​​using an xgboost regression estimator.

[0161] Embodiment 24. Analyzing multiple attribute values ​​using a machine learning-based regression estimator is a method of Embodiment 20, which includes analyzing multiple attribute values ​​using a linear support vector machine (SVM) regression estimator.

[0162] Embodiment 25. The method of Embodiment 20, wherein analyzing multiple attribute values ​​using a machine learning-based regression estimator includes analyzing multiple attribute values ​​using an elastic net estimator.

[0163] Embodiment 26. One or more predicted attribute values ​​include one or more product quality attributes, using any one of the methods in Embodiments 20 to 25.

[0164] Embodiment 27. The method of Embodiment 26, wherein one or more predicted product quality attribute values ​​include one or more predicted chromatographic measurements.

[0165] Embodiment 28. Any one of Embodiments 20 to 27, further comprising receiving user input data from a user via a user interface, which includes one or more identifiers of a specific cell line, modalities of drugs produced using a specific cell line, instructions for a drug product produced using a specific cell line, or protein scaffold types associated with drugs produced using a specific cell line, and further comprising analyzing the user input data using a machine learning-based regression estimator, which further comprises analyzing the user input data using a machine learning-based regression estimator.

[0166] Embodiment 29. Receiving multiple attribute values ​​related to a small cell culture is any one of embodiments 20 to 28, comprising receiving one or more of the measured titer of the small cell culture; the measured viable cell density of the small cell culture; or the measured viability of the small cell culture.

[0167] Embodiment 30. Receiving multiple attribute values ​​related to a small cell culture is any one of the methods in Embodiments 20 to 29, comprising receiving one or more characteristics of the medium of the small cell culture.

[0168] Embodiment 31. The method of Embodiment 30, wherein receiving one or more characteristics of the culture medium includes receiving a measured glucose concentration of the culture medium.

[0169] Embodiment 32. Receiving multiple attribute values ​​related to a small cell culture includes receiving a first attribute value corresponding to a first measurement of an attribute related to the small cell culture; and a second attribute value corresponding to a second measurement of an attribute related to the small cell culture, wherein the first and second measurements are performed on different days of the small cell culture, in any one of the methods of Embodiments 20 to 31.

[0170] Embodiment 33. Any one of Embodiments 20 to 32, further comprising receiving use case data from a user via a user interface using one or more processors before receiving multiple attribute values ​​related to a small cell culture, and selecting a machine learning-based regression estimator from among multiple estimators using one or more processors based on the use case data, each of the multiple estimators being designed for a different use case.

[0171] Embodiment 34. The method of Embodiment 33, wherein receiving data indicating use cases comprises at least (i) receiving at least one attribute value relating to a virtual large cell culture, and (ii) receiving data indicating the modality of the drug produced.

[0172] Embodiment 35. The method of Embodiment 34, wherein receiving data demonstrating use cases includes receiving data demonstrating titers related to at least a virtual large cell culture; and analyzing multiple attribute values ​​using a machine learning-based regression estimator includes analyzing multiple attribute values ​​using (i) a decision tree regression estimator, (ii) a random forest regression estimator, (iii) an xgboost regression estimator, or (iv) a linear support vector machine (SVM) regression estimator.

[0173] Embodiment 36. The method of Embodiment 34, wherein receiving data demonstrating use cases includes receiving data demonstrating chromatographic measurements related to at least a virtual large cell culture; and analyzing multiple attribute values ​​using a machine learning-based regression estimator includes analyzing multiple attribute values ​​using an xgboost regression estimator.

[0174] Embodiment 37. The method of Embodiment 33, further comprising, for each estimator of a plurality of estimators, one or more processors determining the set of features that best predict the output of the estimator; and receiving a plurality of attribute values ​​related to a small cell culture, which includes receiving only the attribute values ​​that are included in the set of features determined for a machine learning-based regression estimator.

[0175] Embodiment 38. Any one of embodiments 20 to 37, further comprising measuring at least some of a plurality of attribute values ​​related to a small cell culture using one or more analytical instruments.

[0176] Embodiment 39. Receiving multiple attribute values ​​is any one of the methods in Embodiments 20 to 38, which includes receiving measurements from a photoelectron cell line generation and analysis system.

[0177] Embodiment 40. One or more non-temporary computer-readable media that, when executed by one or more processors of a computing system, store instructions causing the computing system to perform any one of the methods of Embodiments 20 to 39.

[0178] Embodiment 41. A computing system comprising one or more processors; and one or more non-temporary computer-readable media for storing instructions that, when executed by the one or more processors, cause the computing system to perform any one of the methods of Embodiments 20 to 39.

[0179] While systems, methods, apparatus, and their components have been described in terms of exemplary embodiments, these systems, methods, apparatus, and their components are not limited to these. The detailed descriptions should be interpreted as illustrative examples only, and since describing all possible embodiments would be impractical, if not impossible, not possible, not all possible embodiments of the present invention are described. Many alternative embodiments can be carried out using either the current art or art developed after the filing date of this patent, but such embodiments still fall within the scope of the claims defining the present invention.

[0180] Those skilled in the art will understand that various modifications, changes, and combinations of the above embodiments can be made without departing from the scope of the present invention, and that such modifications, changes, and combinations are to be interpreted as being within the scope of the concept of the present invention.

Claims

1. A method for easily selecting a master cell line from among candidate cell lines that produce recombinant proteins, One or more processors of a computing system receive, for a particular cell line, a plurality of attribute values ​​related to a small cell culture, wherein at least some of the plurality of attribute values ​​are measurements of the small cell culture; By using one or more of the aforementioned processors, at least a machine learning-based regression estimator is used to analyze the multiple attribute values ​​associated with the small cell culture, Predicting one or more attribute values ​​associated with a hypothetical large-scale cell culture for a particular cell line, wherein the predicted one or more attribute values ​​include titer and / or one or more product quality attribute values, and a machine learning-based regression estimator is trained using historical measurements of one or more attribute values ​​corresponding to past small-scale cell cultures; and The one or more processors may, in order to facilitate the selection of the master cell lines for use in the manufacture of drug products, present to the user via a user interface either or both of the following: (i) one or more predicted attribute values, and (ii) whether one or more of the predicted attribute values ​​meet one or more cell line selection criteria. A method that includes this.

2. The method according to claim 1, wherein analyzing the plurality of attribute values ​​using a machine learning-based regression estimator includes analyzing the plurality of attribute values ​​using a decision tree regression estimator.

3. The method according to claim 2, wherein analyzing the plurality of attribute values ​​using a machine learning-based regression estimator includes analyzing the plurality of attribute values ​​using a random forest regression estimator.

4. The method according to claim 2, wherein analyzing the plurality of attribute values ​​using a machine learning-based regression estimator includes analyzing the plurality of attribute values ​​using an xgboost regression estimator.

5. The method according to claim 1, wherein analyzing the plurality of attribute values ​​using a machine learning-based regression estimator includes analyzing the plurality of attribute values ​​using a linear support vector machine (SVM) regression estimator.

6. The method according to claim 1, wherein analyzing the plurality of attribute values ​​using a machine learning-based regression estimator includes analyzing the plurality of attribute values ​​using an elastic network estimator.

7. The method according to claim 1, wherein the predicted one or more attribute values ​​include the one or more product quality attributes.

8. The method according to claim 7, wherein the predicted one or more product quality attribute values ​​include one or more predicted chromatographic measurements.

9. Through the user interface, the user can access, The identifier of the aforementioned specific cell line, A modality of a drug produced using the aforementioned specific cell line, Instructions for the drug product to be produced using the aforementioned specific cell line, or Protein scaffolds related to the drug produced using the aforementioned specific cell line This further includes receiving user input data that includes one or more of the following: Using the machine learning-based regression estimator, analyzing the multiple attribute values ​​associated with the small cell culture is: The method according to claim 1, further comprising analyzing the user input data using the machine learning-based regression estimator.

10. Receiving the aforementioned multiple attribute values ​​related to the small-scale cell culture means that The measured titer of the aforementioned small cell culture; The measured viable cell density of the aforementioned small-scale cell culture; or Measured viability of the aforementioned small cell culture The method according to claim 1, comprising receiving one or more of the following.

11. The method according to claim 1, wherein receiving the plurality of attribute values ​​related to the small cell culture includes receiving one or more characteristics of the culture medium of the small cell culture.

12. The method according to claim 11, wherein receiving one or more of the characteristics of the culture medium includes receiving the measured glucose concentration of the culture medium.

13. Receiving the aforementioned multiple attribute values ​​related to the small-scale cell culture means that A first attribute value corresponding to the first measurement of the attribute related to the small cell culture; and A second attribute value corresponding to the second measurement of the attribute related to the small cell culture. Including receiving, The method according to claim 1, wherein the first measurement and the second measurement are performed on different days of the small cell culture.

14. Before receiving the multiple attribute values ​​related to the small cell culture, The aforementioned one or more processors receive data indicating usage examples from the user via a user interface; and The process involves one or more processors selecting the machine learning-based regression estimator from among multiple estimators based on the data illustrating the use case. The method according to claim 1, further comprising, each of the plurality of estimators being designed for a different use case.

15. The method of claim 14, wherein receiving data indicating the use case includes at least (i) receiving at least one of the one or more attribute values ​​related to the virtual large cell culture, and (ii) receiving data indicating the modality of the drug to be produced.

16. Receiving data illustrating the aforementioned use cases includes receiving data indicating titers related to at least the aforementioned virtual large-scale cell cultures; and Analyzing the aforementioned multiple attribute values ​​using a machine learning-based regression estimator is (i) Decision tree regression estimator, (ii) Random Forest Regression Estimator, (iii) xgboost regression estimator, or (iv) The method of claim 15, comprising analyzing the plurality of attribute values ​​using a linear support vector machine (SVM) regression estimator.

17. Receiving data demonstrating the aforementioned use cases includes receiving data demonstrating at least the chromatographic measurements related to the aforementioned virtual large-scale cell culture; and The method according to claim 15, wherein analyzing the plurality of attribute values ​​using a machine learning-based regression estimator includes analyzing the plurality of attribute values ​​using an xgboost regression estimator.

18. For each of the plurality of estimators, the one or more processors further determine the set of features that best predict the output of the estimator; and The method of claim 14, wherein receiving the plurality of attribute values ​​related to the small cell culture comprises receiving only the attribute values ​​that are included in the set of features determined for the machine learning-based regression estimator.

19. The method according to claim 1, further comprising measuring at least some of the plurality of attribute values ​​related to the small cell culture using one or more analytical instruments.

20. The method according to claim 1, wherein receiving the plurality of attribute values ​​includes receiving measured values ​​from a photoelectron cell line generation and analysis system.

21. One or more non-temporary computer-readable media that, when executed by one or more processors of a computing system, store instructions causing the computing system to perform the method according to any one of claims 1 to 20.

22. A computing system, One or more processors; and One or more non-temporary computer-readable media that, when executed by the one or more processors, store instructions causing the computing system to perform the method according to any one of claims 1 to 20. A computing system that includes this.