Method for predicting serum-free medium component concentrations based on machine learning

By constructing a serum-free culture medium component concentration prediction model using machine learning algorithms, the problem of low efficiency in the research and development of serum-free culture media is solved, enabling rapid and batch component concentration prediction and optimization, and supporting the research and development of serum-free cell cultured meat.

CN115101118BActive Publication Date: 2026-05-12NANJING JOES FUTURE FOOD TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING JOES FUTURE FOOD TECH CO LTD
Filing Date
2022-06-20
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing methods for optimizing serum-free culture medium components are time-consuming and labor-intensive, and cannot meet the needs of simultaneous optimization of multiple components. Existing technologies are inefficient and costly in the development of serum-free culture media, and cannot effectively utilize existing data for rapid development.

Method used

By establishing a sample feature database and utilizing machine learning algorithms such as support vector machines, artificial neural networks, logistic classification, and Naive Bayes classification models, a serum-free culture medium component concentration prediction model is constructed to quickly recommend the optimal formulation and optimize the combination of culture medium components.

Benefits of technology

It enables rapid, batch prediction of serum-free culture medium component concentrations, lowers the development threshold, improves R&D efficiency, provides guidance on serum-free culture medium combinations, and supports serum-free R&D in the cell culture meat process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115101118B_ABST
    Figure CN115101118B_ABST
Patent Text Reader

Abstract

The application discloses a method for predicting serum-free medium component concentration based on machine learning, comprising the following steps: establishing a database of different component characteristic information; model training; establishing a characteristic information database of a group to be predicted; predicting the data of the group to be predicted; database updating. Compared with a basic medium, a serum-free medium needs to screen more components, and existing serum-free medium component optimization methods mostly consume time and labor and cannot meet the demand of simultaneous optimization of multiple components. The application solves the problem of the current technology in the blank of serum-free medium research and development, has the advantages of batch and rapidness, and can provide guidance for serum-free medium research and development.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biotechnology, and in particular relates to a method for optimizing the concentration of serum-free culture medium components based on machine learning. Background Technology

[0002] The conflict between traditional livestock farming methods and resource and environmental factors is becoming increasingly prominent. Traditional livestock farming methods consume a lot of resources and emit large amounts of pollutants. Cell-cultured meat technology is a major innovative technology in meat production that has emerged in the last two decades. This technology can partially replace some traditional livestock farming methods for meat production and is a potential alternative technology for future food production. The development of highly efficient serum-free culture media is one of the major challenges currently facing cell-cultured meat technology.

[0003] Different applications in the field of cultured meat require different cell types, leading to significant differences in the performance characteristics and production needs of cell culture media, including technical difficulty, production processes, and product forms. Compared to basal media, serum-free media require screening of more components, but existing methods for optimizing serum-free media components are mostly time-consuming and labor-intensive and cannot meet the needs of simultaneous optimization of multiple components.

[0004] To address the above issues, this paper explores ways to improve the efficiency of serum-free culture medium development using machine learning methods. The main technical hurdles encountered in practice are as follows:

[0005] 1) Sufficient sample size needs to be accumulated:

[0006] These samples underwent the same wet and dry experimental analysis procedures, and the effective data laid the foundation for model construction.

[0007] 2) Determine the serum-free culture medium development process: The development of serum-free culture medium involves many components, and different components have different information. It is crucial to find the component characteristics related to cell growth and proliferation. These characteristics lay the foundation for predicting the combination of different components of serum-free culture medium based on the sample itself.

[0008] 3) Applying machine learning algorithms to the scenario of predicting the concentration of components in serum-free culture medium with good model performance: There are few reports on using machine learning algorithms to build models for serum-free culture medium data. Summary of the Invention

[0009] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention provides a basic culture medium formulation development method based on machine learning. The aim is to apply machine learning algorithms to the complex formulation optimization process. By constructing a high-quality, sufficiently large database of sample formulations and selecting appropriate machine learning and optimization algorithms, the method can recommend the most likely basic culture medium formulations with good cultivation effects in a short time, thereby lowering the formulation development threshold and solving the technical problems of slow development speed and high development costs caused by the complex composition of existing basic culture media.

[0010] This invention provides a method for predicting the concentration of components in serum-free culture media based on machine learning, which solves the current technological gap in the development of serum-free culture media. It has the advantages of batch processing and speed, and can provide guidance for the development of serum-free culture media.

[0011] To achieve the above objectives, this application proposes the following technical solution:

[0012] (1) Establish a sample feature database with characteristic information of different components:

[0013] The sample feature database includes cell type, concentration of each component of serum-free culture medium, and cell number; the concentration of each component of serum-free culture medium is combined into different levels according to the cell number, preferably 1-10 different categories.

[0014] (2) Model training:

[0015] For the sample feature database obtained in step (1), the concentration of each component in the serum-free culture medium is used as the input feature. The selected machine learning model is trained on the target of the concentration level of each component in the serum-free culture medium divided according to the number of cells. The predicted level corresponding to the given concentration of each component in the serum-free culture medium is output, and the best prediction model is obtained.

[0016] (3) Establish a feature information database for the group to be predicted:

[0017] Based on the original database containing feature information and the component concentration information collected from literature, the concentration of each component is amplified to a fixed range to obtain the feature information database of the group to be predicted.

[0018] (4) Predict the data for the group to be predicted:

[0019] A new dataset is obtained by sampling from the feature information database of the group to be predicted. The number of cells is predicted using the new dataset, and the category level of the formulation is obtained. The selected formulation is then used to conduct new experiments.

[0020] Preferably, it also includes:

[0021] (5) Database update:

[0022] Based on the results of each optimization, the data is compiled and updated to the sample database for use in the next round of optimization. This process is repeated until the selected formula achieves the target cell count or cell proliferation rate.

[0023] Furthermore, the step of obtaining sample feature data in (1) is as follows:

[0024] Cells were seeded into cell culture dishes containing conventional proliferation medium and modified proliferation medium with different components, and passaged for 3 days with medium changes. Cells were then collected for cell counting.

[0025] Furthermore, the classification models used for model training include support vector machine classification models, artificial neural network classification models, logistic classification models, and Naive Bayes classification models, among which artificial neural network classification models are preferred.

[0026] The models that support the training of artificial neural network classification models are:

[0027]

[0028] Exp refers to an exponential function with base e, K represents the number of categories, and w k Let x be the weight coefficient of x, where x is the attribute value of the feature.

[0029] Furthermore, the model supporting the training of the logical classification model is:

[0030]

[0031] Among them, w1, w2, ..., w K These are the parameters of the model. It is a normalization term.

[0032] The parameter W is a matrix, where each row represents the parameters of the classifier for a given category. There are k rows in total, and the K output numbers represent the probabilities of that category, summing to 1. Thus, for a single test sample, the Softmax classification model can obtain probability values ​​for multiple categories, and the model selects the category with the highest probability as the final decision.

[0033] Furthermore, the model trained by the Naive Bayes classification model is:

[0034]

[0035] In the formula, each sample has n features, and the feature output has K categories, defined as C1, C2, ..., C. K y represents the category.

[0036] Furthermore, the multi-class SVM classification model is as follows:

[0037]

[0038] (w ij ) T φ(X t )+b ij ≥1-ξ t ij ,y t =i

[0039] st(w ij ) T φ(X t )+b ij ≥1-ξ t ij ,y t =j (4)

[0040] ξ t ij ≥0

[0041] In the above equation, the parameters with superscripts represent the binary SVM parameters between class i and class j; C is the penalty coefficient, ξ is the slack variable; the subscript t represents the index of the sample in the union of class i and class j; and φ represents the nonlinear mapping from the input space to the feature space.

[0042] The total number of binary SVMs that need to be trained for indices i and j is:

[0043]

[0044] The following formula (6) is used to determine whether the data results of the i-th and j-th classes belong to class i or class j:

[0045] y ij new =sign[(w ij ) T φX new +b ij (6)

[0046] Furthermore, the steps for establishing the feature information database of the group to be predicted in step (3) are as follows:

[0047] First, for each feature attribute, determine the spatial region to be expanded or reduced. The methods and principles include Gaussian distribution and uniform distribution.

[0048] The formula for generating random numbers with mean μ and variance σ using the Gaussian distribution is:

[0049]

[0050] The process of generating uniformly distributed random numbers is as follows:

[0051] RandSeed=(A*RandSeed+B)%M (8)

[0052] Where RandSeed is the random seed, A is the multiplier, B is the increment, and M is the modulus.

[0053] Furthermore, in step (4), a new dataset is obtained by sampling from the prediction group using N-Rooks. The specific steps are as follows:

[0054] 1) Determine the sample size and partition the variable range into equally probable partitions;

[0055] 2) Randomly select a specific number of partitions from the equally probable partitions of each variable;

[0056] 3) Randomly select combinations from specific values ​​of each variable to form a sample;

[0057] 4.) Repeat step 3) until all the specific numbers have been used to obtain the sample set and complete the sampling.

[0058] Furthermore, the newly generated dataset is used as the input value of the model in step (2) above to obtain the performance of the prediction group. Based on the proliferation results of the concentration of each component of the serum-free culture medium obtained in step (4), the result data is organized and updated to the sample feature database described in step (1) for use in the next round of optimization to complete the database update.

[0059] The cultivation effect of this formula can be evaluated by mixing two or more formulas in random or preset proportions to create a new formula.

[0060] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention fills the gap in the current technology in the field of serum-free research and development by establishing a standard serum-free component database and using sample data and sample labels to construct an optimized prediction model based on machine learning algorithms; This invention has wide applications and can predict the performance of different combinations of serum-free culture medium components in batches and quickly, and provide guidance for serum-free research and development in the process of cell culture meat based on the predicted experimental combinations. Attached Figure Description

[0061] Figure 1 This is a flowchart illustrating the optimization of serum-free culture medium component concentrations based on machine learning in this invention.

[0062] Figure 2 This is a data graph of serum-free formulations designed based on machine learning algorithms.

[0063] Figure 3This is a comparison chart of the performance of different algorithms.

[0064] Figure 4 This is a diagram illustrating the design of a serum-free formulation. Detailed Implementation

[0065] The embodiments of the technical solution of the present invention will now be described in detail with reference to the accompanying drawings. These embodiments are merely for illustrating the technical solution of the present invention more clearly.

[0066] It should be noted that, unless otherwise stated, the technical or scientific terms used in this application should have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.

[0067] Existing methods have the following drawbacks: different cell lines require different culture medium formulations, and a culture medium prediction model built for one cell line often fails to meet the required accuracy when predicting the culture medium effects of another cell line. Furthermore, developing a culture medium prediction model for each cell line requires conducting numerous cell culture experiments before building the model from the cell culture data. This fails to achieve the goal of conducting only a small number of experiments when developing culture medium formulations for new cell lines. Existing methods also cannot utilize data from cells with previously developed culture medium formulations, are labor-intensive, and cannot integrate databases.

[0068] The present invention provides a method for developing the concentrations of various components in serum-free culture media based on machine learning, such as... Figure 1 As shown, it includes the following steps:

[0069] (1) Dataset creation and data preprocessing (establishing a sample feature database with different component characteristic information)

[0070] The different grades of the components are categorized as follows: Culture medium formulations with a cell count greater than 1800 per average field of view under a microscope will be labeled as 10; formulations with a cell count greater than 1600 but less than 1800 will be labeled as 9; formulations with a cell count greater than 1400 but less than 1600 will be labeled as 8; formulations with a cell count greater than 1200 but less than 1400 will be labeled as 7; formulations with a cell count greater than 1000 but less than 1200 will be labeled as 6; formulations with a cell count greater than 800 but less than 1000 will be labeled as 5; formulations with a cell count greater than 600 but less than 800 will be labeled as 4; formulations with a cell count greater than 400 but less than 600 will be labeled as 3; formulations with a cell count greater than 200 but less than 400 will be labeled as 2; and formulations with a cell count greater than 0 but less than 200 will be labeled as 1.

[0071] For a given dataset (x1, x2, ..., x...) n (y1,y2,...,y) n )

[0072] X represents the concentration information of each component in different serum-free culture media in the sample pool, and y represents the component grade category corresponding to different concentrations based on the number of cells.

[0073] The training samples were standardized and calculated as follows:

[0074]

[0075] The value obtained after standardization of x′, μ is the sample mean, and s is the standard deviation of the selected sample.

[0076] After standardization, a new training sample set is obtained:

[0077] (x'1,x'2,...,x' n (y1,y2,...,y) n )

[0078] The dataset is divided according to a test set:validation set ratio of 8:2, and then further divided according to a 10-fold cross-validation standard. The dataset is then used for training.

[0079] (2) Model Training

[0080] First, convert the dataset obtained in step (1) from Excel into CSV format. Then, use the pandas module in Python to read the data and specify the values ​​represented by each column. Construct an artificial neural network model with two hidden layers.

[0081] S1: Parameter Initialization

[0082] In this step, the parameters to be trained in the fully connected neural network are initialized, mainly the weight parameters w and the bias parameters b. The initialization method used in this invention initializes these parameters according to a Gaussian distribution, primarily to break the symmetry. If the weights and biases are the same across layers, the outputs of neurons in each layer will not be identical, thus enhancing the learning ability of the neural network. After step S1, the initialized weight parameters and bias parameters are obtained.

[0083] S2: Forward Propagation

[0084] In step S2, the feature values ​​obtained in step S1 are used as inputs, and the results obtained in step S1 are used as model parameters. The output layer then yields the probability levels of the concentrations of each component in the serum-free culture medium.

[0085] The activation function for the hidden layer is:

[0086] ReLU = max(0, w T x+b) ( 10)

[0087] The activation function of the output layer is:

[0088]

[0089] Where z i Let C be the output value of the i-th node, and C be the number of output nodes, i.e. the number of categories.

[0090] S3: Optimize parameters

[0091] In step S3, gradient descent is used to update the model parameters. The optimizer is Adam, and the training batch size is 5. The updated parameters in this step are used to build the model until the model loss function converges to a certain value, yielding the optimal model.

[0092] The loss function used in this invention for multi-class classification is:

[0093]

[0094] The update process for the weight parameter w and the bias parameter b is shown below.

[0095]

[0096] In this step, the output loss value and accuracy value are used to evaluate the model performance.

[0097] In this embodiment, an artificial neural network model is used as an example for illustration. In other embodiments, support vector machine classification models, logistic classification models, and Naive Bayes classification models can all achieve the purpose of the invention through model training. A comparison of the performance of different algorithms is shown in the figure below. Figure 3 As shown.

[0098] (3) Generate the array to be predicted (establish a feature information database for the group to be predicted)

[0099] In this step, the group to be predicted is generated using a Gaussian distribution, and the resulting formula is used in step (4).

[0100] The steps for generating the groups to be predicted using a Gaussian distribution are as follows:

[0101] For each feature attribute, the spatial region to be expanded or reduced is determined, and a random number generation method based on Gaussian distribution is preferred.

[0102] The formula for generating random numbers with mean μ and variance σ using the Gaussian distribution is:

[0103]

[0104] (4) Prediction

[0105] A new dataset is obtained by sampling from the feature information database of the group to be predicted. The sampling steps from the group to be predicted using the N-Rooks method are as follows:

[0106] The input feature values ​​are uniformly distributed into n segments within the interval (a, b), where the i-th segment is represented as follows:

[0107]

[0108]

[0109] The lower and upper limits of the multivariate partitioning are as follows:

[0110]

[0111]

[0112] Define a number η such that 0 ≤ η < 1 as the coefficient for randomly selecting representatives. Then, within the range a... i ≤x i The randomly selected number is:

[0113] (b i -a i )η i +a i =(1-η i )a i +η i b i (16)

[0114] For each interval (layer) of a single variable, there is a random number to represent that interval (layer), thus completing the selection of the representative number for each interval of the single variable.

[0115] Multiple variables can be easily derived from univariate calculations. The random representative numbers for each interval are calculated as follows:

[0116]

[0117] This completes the establishment of the prediction group and the extraction of the concentrations of each component in the serum-free culture medium. The resulting new dataset is used as the input value for the model in step (2) above to obtain the performance of the prediction group, such as... Figure 2 As shown.

[0118] (5) Database update

[0119] The experimental combination predicting performance in the first category was fed into cell culture, where the cell culture steps are as follows:

[0120] ​This embodiment was divided into 11 groups: a control group using conventional proliferation culture medium (allin-19) and experimental groups using 10 different modified proliferation culture media. Porcine muscle stem cells were cultured at 1.5 × 10⁻⁶... 5 / dish was inoculated into 10cm culture dishes containing standard proliferation medium and modified proliferation medium, and cultured with medium changes every two days. On the third day, it was digested with 0.25% trypsin, and the cells were counted using a hemocytometer. Afterwards, the new experimental group ( Figure 4 Data is organized and updated in the database.

[0121] This invention fills a gap in current technology in the field of serum-free research and development by establishing a standard serum-free component database and using sample data and sample labels to construct an optimized prediction model based on machine learning algorithms. This invention has wide applications and can predict the performance of different combinations of serum-free culture medium components in batches and quickly, providing guidance for serum-free research and development in the process of cell cultured meat based on the predicted experimental combinations.

Claims

1. A method for predicting the concentration of components in serum-free culture medium based on machine learning, characterized in that, Includes the following steps: (1) Establish a sample feature database with characteristic information of different components: The sample feature database includes cell type, concentration of each component of serum-free culture medium, and number of cells; the concentration of each component of serum-free culture medium is combined into different levels according to the number of cells, and different levels correspond to different concentrations of each component of serum-free culture medium. (2) Model training: For the component information stored in the sample feature database obtained in step (1), the concentration of each component in the serum-free culture medium is used as the input feature. The selected machine learning model is trained on the target of the concentration level of each component in the serum-free culture medium divided according to the number of cells. The model outputs the predicted level corresponding to the given concentration of each component in the serum-free culture medium and obtains the best prediction model. The classification model used for model training in step (2) is an artificial neural network classification model. The model that supports the training of the artificial neural network classification model is: exp refers to an exponential function with base e, K represents the number of categories, and w k Let x be the weight coefficient of x, where x is the attribute value of the feature; First, convert the data obtained in step (1) into CSV format, read the data and specify the values ​​represented by each column, and construct a two-layer hidden artificial neural network model: S1. Parameter initialization: Initialize the parameters to be trained in the fully connected neural network according to the Gaussian distribution, including the weight parameter w and the bias parameter b. S2, forward propagation, taking the input features obtained in step (1) as input, the results obtained in S1 as model parameters, and the output layer to obtain the level probability of the concentration of each component of the serum-free culture medium. The activation function for the hidden layer is: The activation function of the output layer is: Among them, z i Let C be the output value of the i-th node, and C be the number of categories. S3. Update the model parameters using gradient descent until the model loss function converges to a certain value to obtain the optimal model. (3) Establish a feature information database for the group to be predicted: Based on the original database containing sample characteristic information and the component concentration information obtained from literature collection, the concentration of each serum-free culture medium component in the original database is scaled to a fixed range to expand the database and obtain the characteristic information database of the group to be predicted; the steps for establishing the characteristic information database of the group to be predicted are as follows: For each feature attribute, random numbers are generated using a Gaussian or uniform distribution to determine the spatial region to be expanded or reduced. (4) Predict the data for the group to be predicted: Using N-Rooks to sample the feature information database of the group to be predicted, a new dataset is obtained. The new dataset is input into the model prediction obtained in step (2) to obtain the category level of the formulation and the concentration of each component of the serum-free culture medium.

2. The method according to claim 1, characterized in that... It also includes the following steps: (5) Database update: Based on the proliferation results of the concentration of each component of the serum-free culture medium obtained in step (4), the results data are organized and updated to the sample feature database in step (1) for use in the next round of optimization, thus completing the database update.

3. The method according to claim 1, characterized in that: The steps for establishing the sample feature database in step (1) are as follows: cells are seeded into cell culture dishes containing conventional proliferation medium and modified proliferation medium with different components, and cultured by changing the medium for 3 days. Cells are collected and counted. Based on the number of cells, the concentrations of each component of the serum-free medium are combined into different levels to establish the sample feature database.

4. The method according to claim 1, characterized in that: Different grades correspond to different concentrations of each component in serum-free culture medium. The grades are denoted as 1-10, where: the culture medium formula combination with a cell count greater than 1800 in the average field of view under a microscope is labeled as 10, the formula combination with a cell count greater than 1600 and less than 1800 is labeled as 9, the formula combination with a cell count greater than 1400 and less than 1600 is labeled as 8, the formula combination with a cell count greater than 1200 and less than 1400 is labeled as 7, the formula combination with a cell count greater than 1000 and less than 1200 is labeled as 6, the formula combination with a cell count greater than 800 and less than 1000 is labeled as 5, the formula combination with a cell count greater than 600 and less than 800 is labeled as 4, the formula combination with a cell count greater than 400 and less than 600 is labeled as 3, the formula combination with a cell count greater than 200 and less than 400 is labeled as 2, and the formula combination with a cell count greater than 0 and less than 200 is labeled as 1.

5. The method according to claim 1, characterized in that... The spatial region to be expanded or shrunk is determined using a Gaussian distribution random number generation method, where the Gaussian distribution generates a mean of... The variance is The formula for random numbers is: 。 6. The method according to claim 1, characterized in that, Step (4) specifically includes: 1) Determine the sample size and partition the variable range into equal probability partitions; 2) Randomly select a specific number of partitions from the equally probable partitions of each variable; 3) Randomly select and combine a specific number of variables to form a sample; 4.) Repeat step 3) until all the specific numbers have been used to obtain the sample set, and the sampling is complete.