Cell culture medium configuration optimization method, device and equipment and storage medium
By training a target surrogate model and using a component acquisition function to optimize cell culture medium components, the problem of low efficiency in existing technologies is solved, and efficient and accurate cell culture medium component configuration is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUXI BIOLOGICS (SHANGHAI) CO LTD
- Filing Date
- 2024-05-22
- Publication Date
- 2026-04-14
AI Technical Summary
Existing methods for optimizing cell culture media are inefficient, struggle to find the global optimum, and consume significant time and resources.
A predictive model is obtained by training a target surrogate model. The component acquisition function guides the acquisition of component configuration information, gradually approaching the optimal component configuration. The cell culture medium components are optimized by combining a Gaussian process model and a compound kernel function, enabling targeted experiments.
This improved the efficiency and accuracy of cell culture medium optimization, reduced experimental costs and time, and ensured the reliability and accuracy of the final optimization results.
Smart Images

Figure CN121860100A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biopharmaceutical technology, and in particular to a method, apparatus, equipment and storage medium for optimizing cell culture medium preparation. Background Technology
[0002] In cell culture process development, many steps affect the final productivity, among which culture medium optimization is considered one of the most important steps for continuously improving productivity. Optimizing cell culture medium composition is crucial and critical for achieving high productivity in biopharmaceutical products.
[0003] Currently, existing methods for optimizing cell culture media are mainly based on traditional experimental methods and statistical analysis. These methods typically involve a series of trial-and-error experiments, continuously adjusting the combination, concentration, and ratio of components in the culture medium to find the optimal culture conditions. The drawbacks of this approach are its low efficiency, requiring significant time and resources, and often only yielding local optima rather than global optima, making it difficult to accurately describe the complex interactions within the cell culture medium. This results in a low degree of optimization and efficiency in cell culture media. Summary of the Invention
[0004] This invention provides a method, apparatus, equipment, and storage medium for optimizing cell culture medium preparation, thereby achieving comprehensive and efficient optimization of cell culture medium and improving the degree and efficiency of cell culture medium optimization.
[0005] In a first aspect, embodiments of the present invention provide a method for optimizing the preparation of cell culture medium, comprising:
[0006] The target surrogate model is trained based on the current component data set to obtain the current prediction model; the current component data set includes multiple first component configuration information and a first experimental result corresponding to each first component configuration information, the first component configuration information includes a component data set obtained by random sampling of the target component configuration space or a historical component data set; the target component configuration space stores all component types involved in the target culture medium and the value range corresponding to each component type;
[0007] Based on the current prediction model, determine the current component acquisition function;
[0008] Based on the current component acquisition function, configuration information of the target component configuration space is acquired to obtain multiple second component configuration information;
[0009] Based on the current component data set, multiple second component configuration information, and the second experimental results corresponding to each second component configuration information, determine the candidate component data set obtained in this optimization;
[0010] The candidate component data set obtained in the current optimization is added to the current component data set, and the operation of training the target surrogate model based on the training data is returned to execute until the optimal target component data set that meets the preset convergence condition is obtained.
[0011] Secondly, embodiments of the present invention also provide a cell culture medium preparation optimization device, comprising:
[0012] The prediction model determination module is used to train the target surrogate model based on the current component data set to obtain the current prediction model; the current component data set includes multiple first component configuration information and a first experimental result corresponding to each first component configuration information, the first component configuration information includes a component data set obtained by random sampling of the target component configuration space or a historical component data set; the target component configuration space stores all component types involved in the target culture medium and the value range corresponding to each component type;
[0013] The acquisition function determination module is used to determine the acquisition function for the current component based on the current prediction model.
[0014] The configuration information determination module is used to collect configuration information of the target component configuration space based on the current component acquisition function, and obtain multiple second component configuration information.
[0015] The candidate component set determination module is used to determine the candidate component set obtained in the current optimization based on the current component data set, multiple second component configuration information and the second experimental results corresponding to each second component configuration information.
[0016] The target component set determination module is used to add the candidate component data set obtained in the current optimization to the current component data set, and return to execute the operation of training the target surrogate model based on the training data; until the optimal target component data set that meets the preset convergence conditions is obtained.
[0017] Thirdly, embodiments of the present invention also provide an electronic device, the electronic device comprising: at least one processor; and
[0018] A memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the cell culture medium preparation optimization method provided in any embodiment of the present invention.
[0020] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing computer instructions that enable a processor to execute the cell culture medium configuration optimization method provided in any embodiment of the present invention.
[0021] The technical solution of this invention trains a target proxy model based on the current component data set to obtain a current prediction model. This allows for prediction of experimental results with fewer experiments, reducing experimental costs and time. The current component data set includes multiple first component configuration information and a first experimental result corresponding to each first component configuration information. The first component configuration information includes a component data set randomly sampled from the target component configuration space or a historical component data set. The target component configuration space stores all component types involved in the target culture medium and the value range corresponding to each component type. Based on the current prediction model, a current component acquisition function is determined to guide the subsequent acquisition of configuration information. Configuration information is acquired from the target component configuration space based on the current component acquisition function to obtain multiple second component configuration information. This allows for targeted selection of configuration information for experiments, rather than blindly trying all possible combinations, further improving the efficiency of the optimization process. Based on the current component data set, multiple second component configuration information, and the second experimental result corresponding to each second component configuration information, a candidate component data set obtained in this optimization is determined. The candidate component data set obtained in the current optimization is added to the current component data set, and the operation of training the target surrogate model based on the training data is returned. This process continues until the optimal target component data set that meets the preset convergence conditions is obtained. Through continuous iterative optimization, the optimal cell culture medium component configuration can be gradually approximated. If the preset convergence conditions are met, the target component configuration information corresponding to the target cell culture medium is determined based on the last optimized target component data set. This avoids over-optimization and ensures the accuracy and reliability of the final optimization result, thereby achieving comprehensive and efficient optimization of the cell culture medium and improving the degree and efficiency of cell culture medium optimization.
[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart of a cell culture medium preparation optimization method provided in Embodiment 1 of the present invention;
[0025] Figure 2 This is an example diagram illustrating the optimization of cell culture medium preparation according to Embodiment 1 of the present invention;
[0026] Figure 3 This is an example diagram of a farthest point sampling method according to Embodiment 1 of the present invention;
[0027] Figure 4 This is a flowchart of a cell culture medium preparation optimization method provided in Embodiment 2 of the present invention;
[0028] Figure 5 This is a schematic diagram of a cell culture medium preparation optimization device provided in Embodiment 3 of the present invention;
[0029] Figure 6 This is a schematic diagram of the structure of an electronic device for implementing the cell culture medium configuration optimization method of the present invention. Detailed Implementation
[0030] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0031] It should be noted that the terms "target," "current," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0032] Example 1
[0033] Figure 1This is a flowchart illustrating a method for optimizing cell culture medium preparation according to Embodiment 1 of the present invention. This embodiment is applicable to situations where cell culture medium preparation needs to be optimized. Figure 1 As shown, this method can be performed by a cell culture medium preparation optimization device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method specifically includes the following steps:
[0034] S110. Train the target proxy model based on the current component data set to obtain the current prediction model; the current component data set includes multiple first component configuration information and a first experimental result corresponding to each first component configuration information, the first component configuration information includes a component data set or a historical component data set obtained by random sampling of the target component configuration space; the target component configuration space stores all component types involved in the target culture medium and the value range corresponding to each component type.
[0035] The current component dataset can refer to a dataset generated during the current optimization process, containing multiple first-component configuration information and the corresponding first experimental results for each configuration. The target surrogate model can be a mathematical or statistical model that uses an approximate representation of the data to simulate or predict the impact of different component configuration information on cell performance. For example, the target surrogate model could be a Gaussian process model. A Gaussian process model constructs a prior distribution of a Gaussian process and calculates a posterior distribution by combining it with observed data. Then, the posterior distribution is used to predict and infer unknown data. The current prediction model can be a model obtained by training the target surrogate model. Based on the information in the current component dataset, it uses machine learning or statistical learning methods to learn the potential relationships or patterns between component configuration information and experimental results. Configuration information can refer to all components involved in the use of the cell culture medium and the concentration of each component (such as nutrients, growth factors, etc.). For example, configuration information could include various amino acids, vitamins, metal ions, sugars, and various buffers, lipids, nucleotides, etc. Experimental results can include titer (protein yield), Man5 (pentamannose type), peak viable cell density, etc. The target component configuration space refers to the space comprised of all components involved in the use of the target cell culture medium and the range of concentration (proportion) variations of each component (such as nutrients, growth factors, etc.). It can be a multi-dimensional space, where each dimension represents the concentration or proportion of a component in the target cell culture medium. The target cell culture medium can refer to the cell culture medium used to culture target cells that requires optimization.
[0036] Specifically, if historical component configuration information for the target cell culture medium to be optimized exists, multiple first component configuration information is determined from this historical configuration information, and the historical experimental results corresponding to each first component configuration information are used as the first experimental results. If no historical component configuration information for the target cell culture medium to be optimized exists, random sampling is performed from the target component configuration space corresponding to the target cell culture medium to determine multiple first component configuration information, and a first experimental result corresponding to each first component configuration information is generated based on the first component configuration information. The multiple first component configuration information and the first experimental results corresponding to each first component configuration information are merged to generate the current component data set. The target surrogate model is trained using the current component data set. During training, the model learns the patterns and relationships in the data and adjusts its internal parameters to minimize the prediction error. After training, the current prediction model is obtained, which can automatically process large amounts of data and make decisions based on this data, greatly improving efficiency and reducing human error.
[0037] S120. Based on the current prediction model, determine the current component acquisition function.
[0038] Here, the current component acquisition function can refer to the acquisition function defined based on the output of the current prediction model. For example, the current component acquisition function can refer to the expected improvement function.
[0039] Specifically, based on the current prediction model, the parameter types corresponding to the output results of the current prediction model are determined, and an appropriate function is selected as the current component acquisition function according to the parameter type. This allows for more effective use of limited data and information to find the optimal or near-optimal solution.
[0040] For example, S120 may include: determining the optimal first experimental result corresponding to the current component set based on the current component data set; determining the output result of the current prediction model based on the current prediction model; and constructing the current component acquisition function based on the output result and the optimal first experimental result.
[0041] The optimal first experimental result can refer to the best experimental result obtained from the current component dataset according to a certain evaluation criterion (such as efficiency, accuracy, cost, etc.). The output of the current prediction model can refer to the prediction mean and prediction variance. Here, the prediction mean represents the predicted value of the target cell culture medium performance index under a given input (i.e., component configuration information), reflecting the average performance index that the current prediction model believes to be under this component configuration information. The prediction variance represents the uncertainty of the current prediction model's prediction, measuring the possible fluctuation range of the results when the model makes multiple predictions under the same input.
[0042] Specifically, based on the current component dataset, the optimal first experimental result from the current component dataset is extracted. Based on the current prediction model, the predicted mean and predicted variance corresponding to the output of the current prediction model are determined. Combining the output of the current prediction model and the optimal first experimental result, a current component acquisition function is constructed. This current component acquisition function allows for targeted selection of configuration information for experiments, rather than blindly trying all possible combinations, further improving the efficiency of the optimization process.
[0043] For example, the current component sampling function can be an expected boost function, whose main objective is to find an input point that yields the greatest expected boost relative to the best first experimental result. In other words, it determines how sampling should be performed next to most effectively improve the current optimal solution. The expected boost function combines the predicted mean and prediction uncertainty of the objective function. It considers two factors: the predicted mean, which reflects the model's expected output of the objective function at that input point; and prediction uncertainty, which measures the model's confidence in the predicted mean. By balancing these two factors, the expected boost function finds a balance between exploration (finding potentially better solutions) and exploitation (utilizing the currently known best solution, i.e., the optimal first experimental result).
[0044] S130. Based on the current component acquisition function, the configuration information of the target component configuration space is acquired to obtain multiple second component configuration information.
[0045] The second component configuration information can refer to the specific component configuration information collected from the target component configuration space according to the current component acquisition function.
[0046] Specifically, such as Figure 2As shown, the table in the lower right corner is used to fill in the collected configuration information for multiple second-group assignments. The experiment type can be a wet experiment. Configuration information is collected from the target component configuration space based on the current component acquisition function to obtain multiple second-group configuration information. The target cell culture medium is configured based on each second-group configuration information, and experiments are conducted to obtain the second experimental results corresponding to each second-group configuration information. Conducting experiments using each second-group configuration information allows for targeted selection of configuration information for experiments, rather than blindly trying all possible combinations, further improving the efficiency of the optimization process.
[0047] For example, in a wet experiment, cells and extracellular products are first collected from the culture medium. Then, specific biochemical or molecular biological methods are used to measure the concentration of the target protein or compound. Finally, the measured concentration is multiplied by the culture volume to obtain the total amount of the target product in the entire culture system, which is the titer. This titer then serves as the component experimental result corresponding to the component configuration information.
[0048] For example, S130 may include: determining a target number of the second component configuration information based on the current component data set; sampling the target component configuration space based on the target number to obtain a target number of multiple second component configuration information.
[0049] The target quantity can refer to the number of configuration information in the current component dataset.
[0050] Specifically, based on the current component data set, the number of configuration information entries for the current component data set is determined, and this number is set as the target number of second component configuration information entries. Based on the target number of second component configuration information entries, the target component configuration space is sampled to obtain multiple second component configuration information entries that meet the target number. Experiments are conducted based on the second component configuration information entries, and a second experimental result is determined for each second component configuration information entry.
[0051] For example, sampling the target component configuration space includes: sampling component configuration information of the target component configuration space based on the farthest point sampling method to obtain multiple second component configuration information.
[0052] Among them, Farthest Point Sampling (FPS) refers to a commonly used sampling algorithm, mainly used for sampling point cloud data. It selects the point with the largest minimum distance to all points in the currently selected point set and adds it to the set, so that these points can better represent the overall contour of the point cloud. Component configuration information refers to the various component types and the concentration or proportion of each component type that the target cell culture medium may include during use.
[0053] Specifically, a point is randomly selected from the target component configuration space as the first sample point, which represents the first component configuration information. For each point in the target component configuration space (excluding the selected sample point), the minimum distance between the first sample point and all other sample points is calculated. From all unselected points in the target component configuration space, the point with the largest distance to all points in the sample point set is selected as the next sample point (i.e., the next first component configuration information), and this newly selected sample point is added to the sample point set. This process is repeated until the number of sample points in the sample point set meets the required number of first component configuration information points, thus obtaining multiple first component configuration information points. This enhances the generalization ability for different cell culture medium component configurations and provides a more solid data foundation for optimizing cell culture medium configurations.
[0054] For example, the process of the farthest point sampling method is as follows: Figure 3 As shown, firstly, as Figure 3 As shown in subgraph A, a batch of component configuration information is randomly selected from the configuration space as a candidate set for subsequent farthest sampling; furthermore, the maximum global distance and minimum global distance functions are used alternately, as follows: Figure 3 The subgraphs B, C, and D in the diagram show how to select the component configuration information that is furthest from the default configuration from the candidate set to obtain the current component data set.
[0055] S140. Based on the current component data set, multiple second component configuration information and the second experimental results corresponding to each second component configuration information, determine the candidate component data set obtained in this optimization.
[0056] The candidate component data set can refer to the set consisting of the current component data set, multiple second component configuration information, and the second experimental results corresponding to each second component configuration information.
[0057] Specifically, based on the current component data set, multiple second component configuration information, and the second experimental results corresponding to each second component configuration information, the current component data set, multiple second component configuration information, and the second experimental results corresponding to each second component configuration information are merged to obtain the candidate component data set obtained in this optimization.
[0058] S150. Add the candidate component data set obtained in the current optimization to the current component data set, and return to execute the operation of training the target proxy model based on the training data; until the optimal target component data set that meets the preset convergence condition is obtained.
[0059] The preset convergence conditions can refer to pre-set parameters such as the number of times the target surrogate model will be trained based on the current component dataset, the improvement in the optimization result being less than a certain threshold, and the current prediction model meeting the required stability. The target component dataset can refer to the current component dataset obtained after satisfying the preset convergence conditions. The target component configuration information can refer to the component configuration information that is finally determined and used to optimize the cell culture medium configuration.
[0060] Specifically, the candidate component data set obtained in the current optimization is added to the current component data set to obtain a new current component data set. This ensures that subsequent operations are based on the latest current component data set. In other words, steps S110-S150 are executed again based on the latest current component data set to obtain the candidate component data set for the next optimization. The updated current component data set is used to train the target surrogate model. After training, the target surrogate model can be used to predict experimental results under the new component configuration information. Based on these predictions, optimization operations can be performed again to obtain a new candidate component data set, and the above steps are repeated. If the number of times the operation of training the target surrogate model based on the current component data set is performed meets the preset convergence condition, then the final current component data set is determined as the target component data set. By continuously updating the current component data set and training the target surrogate model, the optimal component configuration information can be gradually approximated, thereby improving optimization efficiency and reducing experimental costs.
[0061] The technical solution of this invention trains a target proxy model based on the current component data set to obtain a current prediction model. This allows for prediction of experimental results with fewer experiments, reducing experimental costs and time. The current component data set includes multiple first component configuration information and a first experimental result corresponding to each first component configuration information. The first component configuration information includes a component data set randomly sampled from the target component configuration space or a historical component data set. The target component configuration space stores all component types involved in the target culture medium and the value range corresponding to each component type. Based on the current prediction model, a current component acquisition function is determined to guide the subsequent acquisition of configuration information. Configuration information is acquired from the target component configuration space based on the current component acquisition function to obtain multiple second component configuration information. This allows for targeted selection of configuration information for experiments, rather than blindly trying all possible combinations, further improving the efficiency of the optimization process. Based on the current component data set, multiple second component configuration information, and the second experimental result corresponding to each second component configuration information, a candidate component data set obtained in this optimization is determined. The candidate component data set obtained in the current optimization is added to the current component data set, and the operation of training the target surrogate model based on the training data is returned. This process continues until the optimal target component data set that meets the preset convergence conditions is obtained. Through continuous iterative optimization, the optimal cell culture medium component configuration can be gradually approximated. If the preset convergence conditions are met, the target component configuration information corresponding to the target cell culture medium is determined based on the last optimized target component data set. This avoids over-optimization and ensures the accuracy and reliability of the final optimization result, thereby achieving comprehensive and efficient optimization of the cell culture medium and improving the degree and efficiency of cell culture medium optimization.
[0062] Based on the above scheme, after S150, the method further includes: comparing the third experimental results corresponding to each third component configuration information in the optimal target component data set based on the optimal target component data set, and determining the optimal third experimental result; based on the optimal third experimental result, determining the third component configuration information corresponding to the optimal third experimental result as the target component configuration information corresponding to the target cell culture medium.
[0063] The third component configuration information can refer to the component configuration information in the target component dataset.
[0064] Specifically, based on the final optimized target component dataset, the third experimental results corresponding to each third component configuration in the target component dataset are compared, and the third experimental results are ranked according to experimental requirements. The best third experimental result that meets the experimental requirements is determined as the optimal third experimental result. Based on the optimal third experimental result, the corresponding third component configuration information is determined, and this optimal third component configuration information is used as the target component configuration information for the target cell culture medium. This ensures that the best target component configuration information is selected from the optimized target component dataset, thereby obtaining a high-performance target cell culture medium.
[0065] For example, if the current component acquisition function can be a desired boost function, the principle of using the desired boost function to construct the current component acquisition function is as follows:
[0066] By considering the current optimal objective function value f(X) + The objective function value at the new point X is given by the following formula:
[0067]
[0068] Among them, f t+1 Z is the predicted objective function value at point X, where ξ is a small exploration parameter used to balance exploration and exploitation. μ(X) and σ(X) are the predicted mean and standard deviation at point X, respectively, and Φ and φ are the cumulative distribution function (CDF) and probability density function (PDF) of the standard normal distribution, respectively. Z is defined as:
[0069]
[0070] The principle of using the current component acquisition function to acquire component configuration information of the target component configuration space is as follows:
[0071] Find component configuration information X to maximize the value of EI(X) of the desired improvement function:
[0072]
[0073] Construction of the acquisition function: The expected boost (EI) is constructed as the acquisition function, based on a randomly selected optimizer, to find the configuration point in the configuration space that maximizes the expected boost.
[0074] Parameter update of the acquisition function: Update the current optimal target value (Incumbent) of the expected improvement function based on historical data. This step ensures that the acquisition function reflects the latest experimental results and optimization status.
[0075] Perform configuration optimization: The optimizer searches the configuration space with the goal of maximizing the expected performance improvement. This process involves evaluating multiple configuration points and selecting the configuration with the highest expected performance improvement.
[0076] Select and return the best configuration: From a series of candidate configurations (Challengers) obtained during the optimization process, select a configuration that does not appear in the history as the recommended configuration for the next round of experimentation or evaluation.
[0077] Example 2
[0078] Figure 4 This is a flowchart of a cell culture medium preparation optimization method provided in Embodiment 2 of the present invention. Based on the above embodiments, this embodiment optimizes the step "training the target surrogate model based on the current component data set to obtain the current prediction model". Explanations of terms that are the same as or corresponding to those in the above embodiments will not be repeated here.
[0079] See Figure 4 Another method for optimizing cell culture medium preparation provided in this embodiment specifically includes the following steps:
[0080] S210. Based on the current component data set, the first component configuration information is determined as the input for model training, and the first experimental result corresponding to the first component configuration information is determined as the label for model training.
[0081] Specifically, based on the current component dataset, the first component configuration information and the corresponding first experimental result are determined as the input and label for model training, respectively. By using the first component configuration information from the current component dataset as input and the first experimental result as label, the relationship between the first component configuration information and the first experimental result can be trained.
[0082] S220. The target surrogate model is trained based on the first component configuration information and the first experimental results corresponding to the first component configuration information to obtain the current prediction model; wherein, the target surrogate model is obtained by integrating a Gaussian process model and a composite kernel function, and the composite kernel function is obtained by combining a constant kernel function, a Markov kernel function and a white noise kernel function.
[0083] Composite kernel functions refer to kernel functions obtained by combining multiple basic kernel functions (also known as simple kernel functions or component kernel functions). Basic kernel functions include constant kernel functions, Matérn kernel functions, and white noise kernel functions, among others. In composite kernel functions, each basic kernel function can capture different characteristics or structures of the data. By combining them, more flexible and powerful kernel functions can be constructed to adapt to complex data distributions and task requirements.
[0084] Specifically, the requirements of the surrogate model for cell culture medium configuration optimization are clearly defined, such as prediction accuracy and computational complexity, and the parameters of the Gaussian process model are determined. A composite kernel function is constructed, which can be composed of multiple basic kernel functions (constant kernel function, Markov kernel function, and white noise kernel function). This combination combines the advantages of different kernel functions, thereby improving the performance of the surrogate model. Specifically, each kernel function has its specific applicable scenarios and properties; combining them allows the new composite kernel function to better adapt to complex data structures and problem characteristics. The Gaussian process model and the composite kernel function are integrated to generate the target surrogate model. The target surrogate model is then trained based on the first component configuration information and the corresponding first experimental results to obtain the current prediction model. Based on the current prediction model, and considering the specific application scenario and optimization objective, the current component acquisition function is determined. By clearly defining the input and labels for model training, using the composite kernel function to generate the target surrogate model, and performing model training, the efficiency and accuracy of the cell culture medium configuration optimization process are improved.
[0085] For example, integrating a Gaussian process model and a composite kernel function to obtain a target surrogate model includes: replacing the original kernel function in the Gaussian process model based on the composite kernel function to determine the Gaussian process model after the kernel function replacement; and adjusting the original covariance function in the Gaussian process model after the kernel function replacement based on the composite kernel function to obtain the target surrogate model.
[0086] In this context, the primal kernel function can refer to a function used in a Gaussian process model to measure the similarity between any two points in the input space. The primal covariance function can refer to a function in a Gaussian process model that describes the uncertainty relationship between the model's output values.
[0087] Specifically, the original kernel function in the Gaussian process model is replaced with a predefined composite kernel function; that is, the mathematical expression of the composite kernel function is replaced in the relevant formulas of the Gaussian process model. Since the original kernel function determines the covariance between different input points in the Gaussian process model, the covariance matrix needs to be recalculated after replacing it with the composite kernel function. The covariance function is then adjusted, and the recalculated covariance matrix will be used to calculate the similarity between input points based on the composite kernel function. This ensures that the Gaussian process model integrating the composite kernel function can better adapt to the characteristics of the data and improve the model's predictive performance.
[0088] S230. Based on the current prediction model, determine the current component acquisition function.
[0089] S240. Based on the current component acquisition function, the configuration information of the target component configuration space is acquired to obtain multiple second component configuration information.
[0090] S250. Based on the current component data set, multiple second component configuration information, and the second experimental results corresponding to each second component configuration information, determine the candidate component data set obtained in this optimization.
[0091] S260. Add the candidate component data set obtained in the current optimization to the current component data set, and return to execute the operation of training the target proxy model based on the training data; until the optimal target component data set that meets the preset convergence condition is obtained.
[0092] The technical solution of this invention, based on the current component dataset, determines the first component configuration information as the input for model training and the first experimental result corresponding to the first component configuration information as the label for model training. The target surrogate model is trained based on the first component configuration information and the corresponding first experimental result to obtain the current prediction model. The target surrogate model is obtained by integrating a Gaussian process model and a composite kernel function, where the composite kernel function is a combination of a constant kernel function, a Markov kernel function, and a white noise kernel function. By clearly defining the input and label for model training, using a composite kernel function to generate the target surrogate model, and performing model training, the efficiency and accuracy of the cell culture medium configuration optimization process are improved. While reducing experimental costs and time, this further ensures the reliability and effectiveness of the final target component configuration information.
[0093] For example, the target agent model can be a Gaussian process model, and the model training process is as follows:
[0094] Type and Boundary Determination: Obtain the type and boundaries of the input space. These types and boundaries are used to define the input space of the Gaussian process model.
[0095] Constructing composite kernel functions: The kernel functions used in this invention are combinations of multiple kernel functions to handle different types of input variables. The kernel functions include:
[0096] Constant kernel: Used to define the overall variation range of the output space.
[0097] k constant (x i x j )=σ 2
[0098] Where, x i x j σ represents a sample in the dataset. 2 It is the only parameter of the constant kernel and represents the global variance.
[0099] Maternal kernel: Used to process continuous input variables, often used to address spatial smoothness issues.
[0100]
[0101] Where d(x) i x j ) is sample x i x j The Euclidean distance between them, where l is the length scale parameter, and K ν It is the modified Bessel function, Γ(·) is the gamma function, and σ is the modified Bessel function. 2 It is the variance term.
[0102] White noise kernel: used to simulate observation noise.
[0103]
[0104] in, This is a parameter related to the noise level. This kernel function only applies to x. i x j Non-zero values generated at the same time (i.e., from the same input point) represent noise in the observation process.
[0105] Composite kernel function: The above kernel functions are combined according to the following relationship to serve as a key component of the Gaussian process.
[0106] K(x i x j )=k constant (x i x j )×k Matern (x i x j )+k white(x i x j )
[0107] Gaussian process model construction: Instantiate a Gaussian process model using a defined kernel function and other parameters (such as input space type and bounds). This model is used to learn and predict the behavior of the objective function.
[0108] Gaussian process training: The model is trained on given input (X) and output (Y) data. Training involves learning the parameters of the kernel function to maximize the marginal likelihood function α.
[0109]
[0110] Where K θ It is the covariance matrix determined by the kernel function parameter θ, and n is the number of training samples.
[0111] Prediction: After training, the model can predict the mean and variance of new input points, which will be used in the acquisition function of offline Bayesian optimization.
[0112] For a new input point X * The Gaussian process model predicts the mean μ and variance σ of its target value. 2
[0113] μ(X * )=K(X * ,X)K(X,X) -1 Y
[0114] σ 2 (X * )=K(X * X * )-K(X * ,X)K(X,X) -1 K(X, X) * )
[0115] Wherein K(X) * K(X, X) is the covariance between the new input point and the training data points, and K(X, X) is the covariance between the training data points. * X * ) is the covariance of the new input point itself.
[0116] Optimizing Gaussian process hyperparameters: Iterate through multiple iterations to optimize the hyperparameters of the Gaussian process and find the parameter set θ that maximizes the marginal log-likelihood. *
[0117] θ * =argmax θ α(θ|X,Y)
[0118] Example 3
[0119] Figure 5 This is a schematic diagram of a cell culture medium preparation optimization device provided in Embodiment 3 of the present invention. Figure 5 As shown, the device includes: a prediction model determination module 310, an acquisition function determination module 320, a configuration information determination module 330, a candidate component set determination module 340, and a target component set determination module 350.
[0120] The prediction model determination module 310 is used to train the target proxy model based on the current component data set to obtain the current prediction model. The current component data set includes multiple first component configuration information and a first experimental result corresponding to each first component configuration information. The first component configuration information includes a component data set obtained by random sampling of the target component configuration space or a historical component data set. The target component configuration space stores all component types involved in the target culture medium and the value range corresponding to each component type.
[0121] The acquisition function determination module 320 is used to determine the current component acquisition function based on the current prediction model;
[0122] The configuration information determination module 330 is used to collect configuration information of the target component configuration space based on the current component acquisition function, and obtain multiple second component configuration information.
[0123] The candidate component set determination module 340 is used to determine the candidate component set obtained in the current optimization based on the current component data set, multiple second component configuration information and the second experimental results corresponding to each second component configuration information.
[0124] The target component set determination module 350 is used to add the candidate component data set obtained in the current optimization to the current component data set, and return to execute the operation of training the target surrogate model based on the training data; until the optimal target component data set that meets the preset convergence conditions is obtained.
[0125] The technical solution of this embodiment obtains a current prediction model by training a target proxy model based on the current component data set. This allows for predicting experimental results with fewer experiments, reducing experimental costs and time. The current component data set includes multiple first component configuration information and a first experimental result corresponding to each first component configuration information. The first component configuration information includes a component data set randomly sampled from the target component configuration space or a historical component data set. The target component configuration space stores all component types involved in the target culture medium and the value range corresponding to each component type. Based on the current prediction model, a current component acquisition function is determined to guide the subsequent acquisition of configuration information. Based on the current component acquisition function, configuration information is acquired from the target component configuration space to obtain multiple second component configuration information. This allows for targeted selection of configuration information for experiments, rather than blindly trying all possible combinations, further improving the efficiency of the optimization process. Based on the current component data set, multiple second component configuration information, and the second experimental result corresponding to each second component configuration information, a candidate component data set obtained in this optimization is determined. The candidate component data set obtained in the current optimization is added to the current component data set, and the operation of training the target surrogate model based on the training data is returned. This process continues until the optimal target component data set that meets the preset convergence conditions is obtained. Through continuous iterative optimization, the optimal cell culture medium component configuration can be gradually approximated. If the preset convergence conditions are met, the target component configuration information corresponding to the target cell culture medium is determined based on the last optimized target component data set. This avoids over-optimization and ensures the accuracy and reliability of the final optimization result, thereby achieving comprehensive and efficient optimization of the cell culture medium and improving the degree and efficiency of cell culture medium optimization.
[0126] Optionally, the prediction model determination module 310 includes:
[0127] The training data determination unit is used to determine the first component configuration information as the input for model training based on the current component data set, and to determine the first experimental result corresponding to the first component configuration information as the label for model training.
[0128] The prediction model determination unit is used to train the target surrogate model based on the first component configuration information and the first experimental results corresponding to the first component configuration information to obtain the current prediction model; wherein, the target surrogate model is obtained by integrating a Gaussian process model and a composite kernel function, and the composite kernel function is obtained by combining a constant kernel function, a Markov kernel function and a white noise kernel function.
[0129] Optionally, the prediction model determination unit is specifically used for: replacing the original kernel function in the Gaussian process model based on the composite kernel function to determine the Gaussian process model after the kernel function replacement; and adjusting the original covariance function in the Gaussian process model after the kernel function replacement based on the composite kernel function to obtain the target surrogate model.
[0130] Optionally, the acquisition function determination module 320 is specifically used for: determining the optimal first experimental result corresponding to the current component set based on the current component data set; determining the output result of the current prediction model based on the current prediction model; and constructing the current component acquisition function based on the output result and the optimal first experimental result.
[0131] Optionally, the current acquisition function determination module 320 is specifically used for: replacing the original kernel function in the Gaussian process model based on the composite kernel function to determine the Gaussian process model after the kernel function replacement; and adjusting the original covariance function in the Gaussian process model after the kernel function replacement based on the composite kernel function to obtain the target surrogate model.
[0132] Optionally, the current acquisition function determination module 320 is specifically used for: determining the optimal first experimental result corresponding to the current component set based on the current component data set; determining the output result of the current prediction model based on the current prediction model; and constructing the current component acquisition function based on the output result and the optimal first experimental result.
[0133] Optionally, the configuration information determination module 330 includes:
[0134] The target quantity determination unit is used to determine the target quantity of the second component configuration information based on the current component data set;
[0135] The configuration information determination unit samples the target component configuration space based on the target quantity to obtain multiple second component configuration information for the target quantity.
[0136] Optionally, the configuration information determination unit is specifically used to: sample the component configuration information of the target component configuration space based on the farthest point sampling method to obtain multiple second component configuration information.
[0137] Optionally, the above apparatus further includes: a target configuration information determination module, specifically used for: comparing the third experimental results corresponding to each third component configuration information in the optimal target component data set based on the optimal target component data set, and determining the optimal third experimental result; and determining the third component configuration information corresponding to the optimal third experimental result as the target component configuration information corresponding to the target cell culture medium based on the optimal third experimental result.
[0138] The cell culture medium preparation optimization device provided in this embodiment of the invention can execute the cell culture medium preparation optimization method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.
[0139] Figure 6 A schematic diagram of an electronic device 12 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as desktop computers, workbenches, servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0140] like Figure 6 As shown, the electronic device 12 is represented in the form of a general-purpose computing device. The components of the electronic device 12 may include, but are not limited to: one or more processors or processing units 16, system memory 28, and bus 18 connecting different system components (including system memory 28 and processing unit 16).
[0141] Bus 18 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0142] Electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by electronic device 12, including volatile and non-volatile media, removable and non-removable media.
[0143] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Electronic device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (… Figure 6 Not shown; usually referred to as a "hard drive"). Although Figure 6Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. System memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0144] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in system memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 typically perform the functions and / or methods described in the embodiments of the present invention.
[0145] Electronic device 12 can also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), and with one or more devices that enable a user to interact with electronic device 12, and / or with any device that enables electronic device 12 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via input / output (I / O) interface 22. Furthermore, electronic device 12 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 20. As shown, network adapter 20 communicates with other modules of electronic device 12 via bus 18. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0146] Processing unit 16 executes various functional applications and data processing by running programs stored in system memory 28, such as implementing the steps of a cell culture medium preparation optimization method provided in this embodiment of the invention, the method including:
[0147] The target surrogate model is trained based on the current component data set to obtain the current prediction model; the current component data set includes multiple first component configuration information and a first experimental result corresponding to each first component configuration information, the first component configuration information includes a component data set obtained by random sampling of the target component configuration space or a historical component data set; the target component configuration space stores all component types involved in the target culture medium and the value range corresponding to each component type;
[0148] Based on the current prediction model, determine the current component acquisition function;
[0149] Based on the current component acquisition function, configuration information of the target component configuration space is acquired to obtain multiple second component configuration information;
[0150] Based on the current component data set, multiple second component configuration information, and the second experimental results corresponding to each second component configuration information, determine the candidate component data set obtained in this optimization;
[0151] The candidate component data set obtained in the current optimization is added to the current component data set, and the operation of training the target surrogate model based on the training data is returned to execute until the optimal target component data set that meets the preset convergence condition is obtained.
[0152] Of course, those skilled in the art will understand that the processor can also implement the technical solution of the cell culture medium configuration optimization method provided in any embodiment of the present invention.
[0153] This embodiment provides a computer-readable storage medium storing a computer program thereon. When executed by a processor, the program implements the cell culture medium preparation optimization method steps provided in any embodiment of the present invention, the method comprising:
[0154] The target surrogate model is trained based on the current component data set to obtain the current prediction model; the current component data set includes multiple first component configuration information and a first experimental result corresponding to each first component configuration information, the first component configuration information includes a component data set obtained by random sampling of the target component configuration space or a historical component data set; the target component configuration space stores all component types involved in the target culture medium and the value range corresponding to each component type;
[0155] Based on the current prediction model, determine the current component acquisition function;
[0156] Based on the current component acquisition function, configuration information of the target component configuration space is acquired to obtain multiple second component configuration information;
[0157] Based on the current component data set, multiple second component configuration information, and the second experimental results corresponding to each second component configuration information, determine the candidate component data set obtained in this optimization;
[0158] The candidate component data set obtained in the current optimization is added to the current component data set, and the operation of training the target surrogate model based on the training data is returned to execute until the optimal target component data set that meets the preset convergence condition is obtained.
[0159] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. For example, a computer-readable storage medium can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0160] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0161] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0162] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and Python, as well as conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0163] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computing device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0164] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. A method for optimizing the preparation of cell culture medium, characterized in that, The method includes: The target surrogate model is trained based on the current component data set to obtain the current prediction model; the current component data set includes multiple first component configuration information and a first experimental result corresponding to each first component configuration information, the first component configuration information includes a component data set obtained by random sampling of the target component configuration space or a historical component data set; the target component configuration space stores all component types involved in the target culture medium and the value range corresponding to each component type; Based on the current prediction model, determine the current component acquisition function; Based on the current component acquisition function, configuration information of the target component configuration space is acquired to obtain multiple second component configuration information; Based on the current component data set, multiple second component configuration information, and the second experimental results corresponding to each second component configuration information, determine the candidate component data set obtained in this optimization; The candidate component data set obtained in the current optimization is added to the current component data set, and the operation of training the target surrogate model based on the training data is returned to execute until the optimal target component data set that meets the preset convergence condition is obtained.
2. The method according to claim 1, characterized in that, The step of training the target surrogate model based on the current component dataset to obtain the current prediction model includes: Based on the current component data set, the first component configuration information is determined as the input for model training, and the first experimental result corresponding to the first component configuration information is determined as the label for model training. The target surrogate model is trained based on the first component configuration information and the first experimental results corresponding to the first component configuration information to obtain the current prediction model; wherein, the target surrogate model is obtained by integrating a Gaussian process model and a composite kernel function, and the composite kernel function is obtained by combining a constant kernel function, a Markov kernel function and a white noise kernel function.
3. The method according to claim 2, characterized in that, The target surrogate model is obtained by integrating the Gaussian process model and the composite kernel function, including: Based on the composite kernel function, the original kernel function in the Gaussian process model is replaced to determine the Gaussian process model after the kernel function is replaced. Based on the composite kernel function, the original covariance function in the Gaussian process model after the kernel function is replaced is adjusted to obtain the target surrogate model.
4. The method according to claim 1, characterized in that, The process of determining the current component acquisition function based on the current prediction model includes: Based on the current component data set, determine the optimal first experimental result corresponding to the current component set; Based on the current prediction model, determine the output result of the current prediction model; Based on the output results and the optimal first experimental results, the current component acquisition function is constructed.
5. The method according to claim 1, characterized in that, The step of collecting configuration information of the target component configuration space based on the current component acquisition function to obtain multiple second component configuration information includes: Based on the current component data set, determine the target quantity of the second component configuration information; Based on the target quantity, the target component configuration space is sampled to obtain multiple second component configuration information for the target quantity.
6. The method according to claim 5, characterized in that, The sampling of the target component configuration space includes: The target component configuration space is sampled using the farthest point sampling method to obtain multiple second component configuration information.
7. The method according to claim 1, characterized in that, After obtaining the optimal target component data set that satisfies the preset convergence conditions, the process also includes: Based on the optimal target component data set, the third experimental results corresponding to each third component configuration information in the optimal target component data set are compared to determine the optimal third experimental result; Based on the optimal third experimental result, the third component configuration information corresponding to the optimal third experimental result is determined as the target component configuration information corresponding to the target cell culture medium.
8. A cell culture medium optimization device, characterized in that, include: The prediction model determination module is used to train the target surrogate model based on the current component dataset to obtain the current prediction model; The current component data set includes multiple first component configuration information and a first experimental result corresponding to each first component configuration information. The first component configuration information includes a component data set obtained by random sampling of the target component configuration space or a historical component data set. The target component configuration space stores all component types involved in the target culture medium and the value range corresponding to each component type; The acquisition function determination module is used to determine the acquisition function for the current component based on the current prediction model. The configuration information determination module is used to collect configuration information of the target component configuration space based on the current component acquisition function, and obtain multiple second component configuration information. The candidate component set determination module is used to determine the candidate component set obtained in the current optimization based on the current component data set, multiple second component configuration information and the second experimental results corresponding to each second component configuration information. The target component set determination module is used to add the candidate component data set obtained in the current optimization to the current component data set, and return to execute the operation of training the target surrogate model based on the training data; Until the optimal target component data set that meets the preset convergence conditions is obtained.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the cell culture medium preparation optimization method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the cell culture medium preparation optimization method according to any one of claims 1-7.