A medium prediction method and system based on feature transfer meta-learning

By employing a feature transfer meta-learning approach, combining meta-learning and a basic learner, a culture medium formulation prediction model is constructed. This addresses the cost issue of large sample data volumes in artificial intelligence-based culture medium development methods, achieving efficient culture medium formulation prediction and accelerating R&D.

CN116189782BActive Publication Date: 2026-06-02SHENZHEN TAILI BIOTECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN TAILI BIOTECHNOLOGY CO LTD
Filing Date
2023-02-27
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing AI-based methods for developing culture media require a large amount of data on the effects of culture media, resulting in high costs and poor predictive performance when sample data is insufficient.

Method used

A feature transfer meta-learning-based approach is adopted. By selecting suitable cell line sample formulation data and combining it with meta-learning training, a culture medium formulation prediction model is constructed, which reduces the amount of sample data required. The combined model of meta-learner and basic learner is used to predict culture medium formulations.

Benefits of technology

While reducing the amount of culture medium formulation sample data, it improves prediction accuracy and generalization ability, reduces learning and training costs, avoids repeated experiments, improves analysis efficiency, and accelerates research and development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116189782B_ABST
    Figure CN116189782B_ABST
Patent Text Reader

Abstract

The application discloses a culture medium prediction method and system based on feature transfer meta-learning. The method comprises the following steps: (1) training a basic learner for predicting the culture effect of a culture medium formula by using a meta-learner based on feature transfer to obtain a culture medium formula culture effect prediction model; and (2) predicting the culture effect value of the culture medium formula to be predicted by using the culture medium formula culture effect prediction model obtained in step (1). The system comprises a basic learner and a meta-learner; the basic learner comprises a full connection neural network, a convolutional neural network, a recurrent neural network or an attention network model. The application greatly reduces the labor cost, avoids the emergence of inferior formula caused by human factors, improves the accuracy of the test, avoids invalid waste, can reduce repetitive experiments, effectively improves the analysis efficiency, and further speeds up the research and development speed and saves time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biotechnology, and more specifically, relates to a method and system for predicting culture media based on feature transfer meta-learning. Background Technology

[0002] Serum-free, animal-free, and chemically defined culture media consist of carbon sources, amino acids, vitamins, trace metal ions, lipids, buffer reagents, and other additives. Traditional culture medium formulation development methods involve using one or several classic culture media as a base (such as DEME / F12), adding various different components, and employing single-factor experiments or DOE screening experiments to identify key components. Then, various DOE experimental designs, such as response surface methodology, are used to optimize the concentration of each component to obtain the optimal formulation. Alternatively, the formulation can be optimized by analyzing cell metabolism, genomics, and proteomics to identify the changes in each component during cell growth and their impact on the yield and quality of the target product.

[0003] Several artificial intelligence-based methods for developing culture media have been developed, such as Chinese patent documents CN113450882A and CN113450868A. These AI-based methods can, to some extent, shorten the development and testing cycle time of culture media formulations and reduce experimental costs. CN114121161A, through transfer learning, reduces the sample size requirements of the sample database during the development of culture media for different cell types, thereby lowering development costs and training time.

[0004] However, the aforementioned AI-based culture medium development methods require generating a large amount of data on culture medium performance—i.e., a sample database—for the model to learn the initial intelligent learning model. The size of the sample database directly determines the effectiveness of the AI-based culture medium development method. A large amount of data is still needed to build the AI ​​model, thus the existing intelligent learning-based culture medium development methods remain costly. Summary of the Invention

[0005] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention provides a culture medium prediction method and system based on feature transfer meta-learning. Its purpose is to significantly reduce the amount of culture medium formula sample data that needs to be accumulated while maintaining the performance of the culture medium prediction system by selecting suitable cell line sample formulation data and combining it with meta-learning training. This solves the technical problem that existing artificial intelligence culture medium development methods require the accumulation of a large amount of culture medium culture effect data, resulting in high costs.

[0006] To achieve the above objectives, according to one aspect of the present invention, a method for predicting culture media based on feature transfer meta-learning is provided, comprising the following steps:

[0007] (1) A meta-learner based on feature transfer is used to train the basic learner used to predict the culture effect of culture medium formulation, so as to obtain the culture effect prediction model of culture medium formulation.

[0008] Its meta-learning sample data is obtained as follows:

[0009] S1. From the historical cell line sample formulation database, select formulation samples with culture effect indicators including target indicators to form a reference formulation sample database; the formulation samples include: applicable cell lines, culture medium formulations, and culture effect indicators including target indicators.

[0010] S2. For each formula sample in the reference formula sample database obtained in step S1, the target cell line is inoculated and cultured using its culture medium formula, and the culture effect values ​​of the same indicators as the culture effect indicators of the formula sample are collected to form the target cell line sample formula database.

[0011] S3. For the target cell line sample formula database obtained in step S2 and the reference formula sample database obtained in step S1, based on the different culture effects of the same culture medium formula on the target cell line and the cell lines contained in the reference formula sample database, or the culture effects of different culture medium formula on the target cell line and the cell lines contained in the reference formula sample database, determine the similarity between the target cell line and the cell lines contained in the reference formula sample database, and take the cell line with the highest similarity as the reference cell line.

[0012] S4. The culture medium formulation for co-inoculation of the target cell line and the reference cell line, the target index culture effect vector of the target cell line, and the target index culture effect vector of the reference cell line are used to form an experience sample. The experience sample set is collected as a meta-learning experience pool based on feature transfer.

[0013] S5. Randomly sample from the feature transfer-based meta-learning experience pool obtained in step S4 to form a task set, and form a meta-training set from the task set. Heyuan Test Set The training set Heyuan Test Set As meta-learning sample data;

[0014] (2) For the culture medium formulation to be predicted, the culture effect prediction model obtained in step (1) is used to predict the culture effect value.

[0015] Preferably, in the culture medium prediction method based on feature transfer meta-learning, step (1) involves supervised training using the meta-learning sample data, and the loss function used is:

[0016] Loss = metaloss + λ * mmdloss

[0017] Where metaloss is the difference between the predicted value of the culture medium formulation culture effect prediction model for the target cell line and the target index culture effect vector of the target cell line in the meta-learning sample data, mmdloss is the maximum mean difference between the predicted value of the culture medium formulation culture effect prediction model for the target cell line and the target index culture effect vector of the reference cell line in the meta-learning sample data, and λ is the weight.

[0018] Preferably, in the culture medium prediction method based on feature transfer meta-learning, the difference between the predicted value of the culture medium formulation culture effect prediction model for the target cell line and the target index culture effect vector of the target cell line in the meta-learning sample data (metaloss) is calculated according to the following method:

[0019]

[0020] Where Y o ′ bject The culture effect vector obtained by predicting the culture medium formulation of the training sample set of the meta-learning sample data for the culture medium formulation culture effect prediction model is of length dim, where the value of the d-th element, i.e., the d-th target index, is Y. o ′ bject [d];Y object The target cell line's culture performance vector is a training sample set of the meta-learning sample data, with length dim, where the d-th element, i.e., the d-th target indicator, has a value of Y. object [d], where n is the size of the training sample set.

[0021] Preferably, in the culture medium prediction method based on feature transfer learning, step S1, the culture medium formulation of the formulation sample should have good culture effect and universality for the target indicator, preferably screened using ranking or thresholding methods; the culture medium formulation with good culture effect has a culture effect indicator or a weighted value of the culture effect indicator exceeding a preset culture effect threshold; the culture medium formulation with good universality has a culture effect exceeding a preset universality threshold for more than a preset proportion of cell lines; the culture effect threshold can be the value of the culture effect indicator or the weighted value of the culture effect indicator, or a ranking of the two; the ranking universality threshold can also be the value of the culture effect indicator or the weighted value of the culture effect indicator, or a ranking of the two.

[0022] Preferably, in the culture medium prediction method based on feature transfer meta-learning, step S3, the similarity determination, includes the following steps:

[0023] For one or more culture medium formulations, obtain the index values ​​composed of various culture effect indicators of the target cell line and cell lines included in the reference formulation sample database. Evaluate the similarity between the two sets of values ​​for the target cell line and each cell line included in the reference formulation sample database to obtain a similarity score for the one or more culture medium formulations; or

[0024] For each culture effect index, obtain the values ​​of all formulas in the target cell line sample formula database for that culture effect index, and the values ​​of the corresponding formulas in the reference formula sample database for each cell line contained therein. Evaluate the similarity between the two sets of values ​​for the target cell line and each cell line contained in the reference formula sample database to obtain a similarity score for that culture effect index.

[0025] The similarity between the two sets of values ​​is evaluated based on indicators such as n-dimensional cosine similarity or n-dimensional Euclidean distance.

[0026] The similarity scores of all culture effect indicators are used as the similarity between the target cell line and each cell line contained in the reference formula sample database; the similarity between the target cell line and each cell line contained in the reference formula sample database is the average, cumulative, weighted average, or weighted sum of the similarity scores of all culture effect indicators or culture medium formulas.

[0027] Preferably, in the culture medium prediction method based on feature transfer meta-learning, step (1) of the culture medium formulation culture effect prediction model adopts the basic learner and meta-learner adaptation method, wherein the basic learner includes a fully connected neural network, a convolutional neural network, a recurrent neural network, or an attention network model.

[0028] According to another aspect of the present invention, a culture medium prediction system based on feature transfer meta-learning is provided, which includes a base learner and a meta-learner.

[0029] The learning sample data for the meta-learner is obtained according to the following method:

[0030] S1. From the historical cell line sample formulation database, select formulation samples with culture effect indicators including target indicators to form a reference formulation sample database; the formulation samples include: applicable cell lines, culture medium formulations, and culture effect indicators including target indicators.

[0031] S2. For each formula sample in the reference formula sample database obtained in step S1, the target cell line is inoculated and cultured using its culture medium formula, and the culture effect values ​​of the same indicators as the culture effect indicators of the formula sample are collected to form the target cell line sample formula database.

[0032] S3. For the target cell line sample formula database obtained in step S2 and the reference formula sample database obtained in step S1, based on the different culture effects of the same culture medium formula on the target cell line and the cell lines contained in the reference formula sample database, or the culture effects of different culture medium formula on the target cell line and the cell lines contained in the reference formula sample database, determine the similarity between the target cell line and the cell lines contained in the reference formula sample database, and take the cell line with the highest similarity as the reference cell line.

[0033] S4. The culture medium formulation for co-inoculation of the target cell line and the reference cell line, the target index culture effect vector of the target cell line, and the target index culture effect vector of the reference cell line are used to form an experience sample. The experience sample set is collected as a meta-learning experience pool based on feature transfer.

[0034] S5. Randomly sample from the feature transfer-based meta-learning experience pool obtained in step S4 to form a task set, and form a meta-training set from the task set. Heyuan Test Set The training set Heyuan Test Set As meta-learning sample data;

[0035] The basic learner includes fully connected neural networks, convolutional neural networks, recurrent neural networks, or attention network models.

[0036] Preferably, in the culture medium prediction system based on feature transfer meta-learning, the meta-learner is trained under supervision using the meta-learning sample data, and the loss function used is:

[0037] Loss = metaloss + λ * mmdloss

[0038] Where metaloss is the difference between the predicted value of the culture medium formulation culture effect prediction model for the target cell line and the target index culture effect vector of the target cell line in the meta-learning sample data, mmdloss is the maximum mean difference between the predicted value of the culture medium formulation culture effect prediction model for the target cell line and the target index culture effect vector of the reference cell line in the meta-learning sample data, and λ is the weight.

[0039] The difference between the predicted values ​​of the culture medium formulation and the culture effect prediction model for the target cell line and the target index culture effect vector of the target cell line from the meta-learning sample data is calculated by Metaloss using the following method:

[0040]

[0041] Where Y o ′bject The culture effect vector obtained by predicting the culture medium formulation of the training sample set of the meta-learning sample data for the culture medium formulation culture effect prediction model is of length dim, where the value of the d-th element, i.e., the d-th target index, is Y. o ′ bject [d];Y object The target cell line's culture performance vector is a training sample set of the meta-learning sample data, with length dim, where the d-th element, i.e., the d-th target indicator, has a value of Y. object [d], where n is the size of the training sample set.

[0042] Preferably, in the culture medium prediction system based on feature transfer learning, the culture medium formulation of the formulation sample in step S1 should have good culture effect and universality for the target indicator, and is preferably screened by ranking or thresholding methods; the culture medium formulation with good culture effect has a culture effect indicator or a weighted value of the culture effect indicator exceeding a preset culture effect threshold; the culture medium formulation with good universality has a culture effect exceeding a preset universality threshold for more than a preset proportion of cell lines; the culture effect threshold can be the value of the culture effect indicator or the weighted value of the culture effect indicator, or a ranking of the two; the ranking universality threshold can also be the value of the culture effect indicator or the weighted value of the culture effect indicator, or a ranking of the two.

[0043] Preferably, in the culture medium prediction system based on feature transfer meta-learning, step S3, the similarity determination, includes the following steps:

[0044] For one or more culture medium formulations, obtain the index values ​​composed of various culture effect indicators of the target cell line and cell lines included in the reference formulation sample database. Evaluate the similarity between the two sets of values ​​for the target cell line and each cell line included in the reference formulation sample database to obtain a similarity score for the one or more culture medium formulations; or

[0045] For each culture effect index, obtain the values ​​of all culture medium formulations in the target cell line sample formulation database for that culture effect index, and the values ​​of the corresponding culture medium formulations in the reference formulation sample database for each cell line it contains. Evaluate the similarity between the two sets of values ​​for the target cell line and each cell line contained in the reference formulation sample database to obtain a similarity score for that culture effect index.

[0046] The similarity between the two sets of values ​​is evaluated based on indicators such as n-dimensional cosine similarity or n-dimensional Euclidean distance.

[0047] The similarity scores of all culture effect indicators are used as the similarity between the target cell line and each cell line contained in the reference formula sample database; the similarity between the target cell line and each cell line contained in the reference formula sample database is the average, cumulative, weighted average, or weighted sum of the similarity scores of all culture effect indicators or culture medium formulas.

[0048] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects:

[0049] The culture medium formulation prediction method and system based on feature transfer meta-learning provided by this invention selects cell lines similar to the target cells and culture medium formulations with better culture effects and versatility as sample data for meta-learning. In the case of limited data on the effect of culture medium formulations for the target cell lines, it solves the problems of difficult training of basic learners and poor prediction effect of culture medium formulations. While maintaining the accuracy of culture medium formulation prediction, it has good generalization ability and greatly reduces the learning and training cost of current deep learning-based culture medium prediction methods.

[0050] This invention greatly reduces labor costs, avoids the occurrence of inferior formulas due to human factors, improves the accuracy of experiments, avoids ineffective waste, reduces repetitive experiments, effectively improves analytical efficiency, and thus accelerates research and development and saves time. Attached Figure Description

[0051] Figure 1 This is a schematic diagram of the culture medium prediction method based on meta-learning and feature transfer provided in an embodiment of the present invention. Detailed Implementation

[0052] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0053] Currently, most effective culture medium prediction methods are based on deep learning models. With a large amount of training data, deep learning models can fit the culture medium formulation prediction task very well, achieving the required prediction accuracy. However, deep learning models often encounter overfitting problems, resulting in low generalization ability. Statistical learning models require less training data and have good generalization ability, but their task-fitting ability is not as good as deep learning models. Meta-learning provides a framework that combines statistical learning and deep learning. The resulting meta-learning model avoids overfitting while forming a culture medium prediction model with good fitting ability.

[0054] Existing deep learning-based methods for predicting culture media offer high accuracy for in-sample predictions. However, out-of-sample predictions require out-of-sample data collection followed by deep learning for fitting. The core idea is to transform the out-of-sample prediction problem into an in-sample prediction problem as much as possible. However, in practical applications, collecting out-of-sample culture media formulation and culture effect data is difficult due to various constraints. In the field of culture media development, obtaining formulations with high expression levels is challenging. This situation places high demands on the model's generalization ability. The meta-learning framework extends the experience learned by deep learning models in-sample to out-of-sample situations, analyzes the commonalities between out-of-sample and in-sample situations, and extends the experience of in-sample problems to out-of-sample problems, thereby reasonably solving the out-of-sample prediction problem.

[0055] The culture medium prediction method based on feature transfer meta-learning provided by this invention includes the following steps:

[0056] (1) A meta-learner based on feature transfer is used to train the basic learner used to predict the culture effect of culture medium formulation, so as to obtain the culture effect prediction model of culture medium formulation.

[0057] The culture medium formulation culture effect prediction model adopts the basic learner and meta-learner adaptation method. The basic learner includes, but is not limited to, fully connected neural networks, convolutional neural networks, recurrent neural networks, and attention network models.

[0058] Its meta-learning sample data is obtained as follows:

[0059] S1. From the historical cell line sample formulation database, select formulation samples with culture effect indicators including the target indicator to form a reference formulation sample database. The formulation samples include: applicable cell lines, culture medium formulations, and culture effect indicators including the target indicator. The culture medium formulations of the formulation samples should have good culture effect and versatility for the target indicator, preferably selected using ranking or threshold methods. A culture medium formulation with good culture effect has a culture effect indicator or its weighted value exceeding a preset culture effect threshold. A culture medium formulation with good versatility has a culture effect exceeding a preset versatility threshold for cell lines exceeding a preset proportion. The culture effect threshold can be the value of the culture effect indicator, its weighted value, or a ranking of both. The ranking versatility threshold can also be the value of the culture effect indicator, its weighted value, or a ranking of both. The culture effect threshold and the ranking versatility threshold can be the same or different. Generally, a culture medium formulation with good versatility has a culture effect indicator or its weighted value exceeding the preset culture effect threshold on at least two different types of cell lines.

[0060] S2. For each formula sample in the reference formula sample database obtained in step S1, the target cell line is inoculated and cultured using its culture medium formula, and the culture effect values ​​of the same indicators as the culture effect indicators of the formula sample are collected to form the target cell line sample formula database.

[0061] S3. For the target cell line sample formula database obtained in step S2 and the reference formula sample database obtained in step S1, based on the different culture effects of the same culture medium formula on the target cell lines and the cell lines contained in the reference formula sample database, or the culture effects of different culture medium formulas on the target cell lines and the cell lines contained in the reference formula sample database, determine the similarity between the target cell lines and the cell lines contained in the reference formula sample database, and take the cell line with the highest similarity as the reference cell line; the similarity determination includes the following steps:

[0062] For one or more culture medium formulations, obtain the index values ​​composed of various culture effect indicators of the target cell line and cell lines included in the reference formulation sample database. Evaluate the similarity between the two sets of values ​​for the target cell line and each cell line included in the reference formulation sample database to obtain a similarity score for the one or more culture medium formulations; or

[0063] For each culture effect index, obtain the values ​​of all culture medium formulations in the target cell line sample formulation database for that culture effect index, and the values ​​of the corresponding culture medium formulations in the reference formulation sample database for each cell line it contains. Evaluate the similarity between the two sets of values ​​for the target cell line and each cell line contained in the reference formulation sample database to obtain a similarity score for that culture effect index.

[0064] The similarity between the two sets of values ​​can be evaluated using indicators such as n-dimensional cosine similarity and n-dimensional Euclidean distance.

[0065] The similarity scores of all culture effect indicators are used as the similarity between the target cell line and each cell line contained in the reference formula sample database. The similarity between the target cell line and each cell line contained in the reference formula sample database can be the average, cumulative, weighted average, or weighted sum of the similarity scores of all culture effect indicators or culture medium formulas, etc.

[0066] S4. The culture medium formulation for co-inoculation of the target cell line and the reference cell line, the target index culture effect vector of the target cell line, and the target index culture effect vector of the reference cell line are used to form an experience sample. The experience sample set is collected as a meta-learning experience pool based on feature transfer.

[0067] S5. Randomly sample from the feature transfer-based meta-learning experience pool obtained in step S4 to form a task set, and form a meta-training set from the task set. Heyuan Test Set The training set Heyuan Test Set As meta-learning sample data;

[0068] Supervised training is performed using the aforementioned meta-learning sample data, and the loss function used is:

[0069] Loss = metaloss + λ * mmdloss

[0070] Where metaloss represents the difference between the predicted value of the target cell line by the culture medium formulation culture effect prediction model and the target index culture effect vector of the target cell line in the meta-learning sample data; mmdloss represents the maximum mean difference between the predicted value of the target cell line by the culture medium formulation culture effect prediction model and the target index culture effect vector of the reference cell line in the meta-learning sample data; and λ represents the weight. Specifically:

[0071] The difference between the predicted values ​​of the culture medium formulation and the culture effect prediction model for the target cell line and the target index culture effect vector of the target cell line from the meta-learning sample data is calculated by Metaloss using the following method:

[0072]

[0073] Where Y o ′ bject The culture effect vector obtained by predicting the culture medium formulation of the training sample set of the meta-learning sample data for the culture medium formulation culture effect prediction model is of length dim, where the value of the d-th element, i.e., the d-th target index, is Y. o ′ bject [d];Y object The target cell line's culture performance vector is a training sample set of the meta-learning sample data, with length dim, where the d-th element, i.e., the d-th target indicator, has a value of Y. object [d], where n is the size of the training sample set.

[0074] The maximum mean difference (mmdloss) between the predicted values ​​of the target cell line and the target index culture effect vector of the reference cell line in the meta-learning sample data is used to quantitatively evaluate the gap between the distribution of the predicted values ​​of the target cell line and the distribution of the target index culture effect of the reference cell line.

[0075] (2) For the culture medium formulation to be predicted, the culture effect prediction model obtained in step (1) is used to predict the culture effect value.

[0076] In meta-learning, each task corresponds to a subset of samples. It's typically assumed that a probability distribution exists for the task. For sample points that are easily observable (i.e., have a high probability) within the task distribution—the in-distribution task—there is a wealth of past experience to draw upon during training. Conversely, for sample points that are difficult to observe (have a low probability) within the task distribution—the out-of-distribution task—only a limited amount of past experience can be used for training. In out-of-distribution tasks, high-expression data in the training dataset is scarce, making model training difficult. Without collecting more high-expression data, deep learning models need to be trained to complete the task. Meta-learning avoids retraining the base learner; instead, it directly uses past training experience with the base learner, adapting it to the new target cell line to train the base learner for that new target cell line. The base learner is the foundation of the meta-learning model, which is a generalization model combining deep learning models across tasks.

[0077] The following is an example:

[0078] Basic culture medium, fed culture medium, or perfusion culture medium can all be predicted using the culture medium prediction method based on feature transfer element learning provided by this invention. The examples will be specifically introduced using basic culture medium as an example.

[0079] The culture medium prediction method based on feature transfer meta-learning provided in this embodiment, such as... Figure 1 As shown, it includes the following steps:

[0080] (1) Meta-learning based on feature transfer was used to train the initial model for predicting the culture medium formulation culture effect, and the culture medium formulation culture effect prediction model was obtained.

[0081] The culture medium formulation culture effect prediction model adopts the basic learner and meta-learner adaptation method. The basic learner includes, but is not limited to, fully connected neural networks, convolutional neural networks, recurrent neural networks, and attention network models.

[0082] Its meta-learning training specifically includes the following steps:

[0083] S1. From the historical cell line sample formulation database, select formulation samples with culture effect indicators including target indicators to form a reference formulation sample database; the formulation sample includes: applicable cell line, culture medium formulation, and culture effect indicators including target indicators; the culture medium formulation of the formulation sample should have good culture effect and universality for the target indicators. This embodiment describes the sorting or threshold method for screening.

[0084] To select culture medium formulations that demonstrate good culture efficacy and general applicability for the target indicators, one of the following strategies can be adopted:

[0085] Strategy 1: Sorting method, the recipe database is represented as follows (example):

[0086]

[0087] Based on the above formula scores, the formulas are re-sorted, and the required number of formulas from the top-ranked formulas are selected to form a reference formula sample database, as shown in the table below:

[0088] Formula number Formula Score Formula 1 6 Formula 4 7 Formula 3 8 Formula 2 9

[0089] Extreme case: A situation that may occur when the number of historical cell lines is equal to the number of formulations.

[0090]

[0091] Solutions for extreme cases: The first option is to continue accumulating the number of recipes; the second option is to adopt strategy two.

[0092] Strategy 2: Threshold Method

[0093] For the cell lines contained in the historical cell line sample formulation database, the sample formulation database of multiple cell lines is screened according to a set threshold. Those that are above the threshold are defined as having good culture effect (the threshold is usually set as the culture effect value of commercial culture medium), and then the sorting method is used for subsequent operations.

[0094] The above-mentioned screening strategy can be used to select a reference formula sample database from the historical cell line sample formula database.

[0095] S2. For each formulation sample in the reference formulation sample database obtained in step S1, the target cell line is inoculated and cultured using its culture medium formulation. The culture effect values ​​of the same indicators as those of the formulation sample are collected to form a target cell line sample formulation database; wherein:

[0096] X1 to n represent different components (mg / L), Y_cell_density refers to cell density (cells / ml), and Y_iter refers to protein expression level (mg / L).

[0097]

[0098] S3. For the target cell line sample formulation database obtained in step S2 and the reference formulation sample database obtained in step S1, based on the culture effects of different culture medium formulations on the target cell lines and the cell lines contained in the reference formulation sample database, determine the similarity between the target cell lines and the cell lines contained in the reference formulation sample database, and take the cell line with the highest similarity as the reference cell line; the similarity is determined according to the following method:

[0099] For each culture effect index, the values ​​of all culture medium formulations in the target cell line sample formulation database for that culture effect index are obtained, as well as the values ​​of the corresponding culture medium formulations in the formulation sample database for the culture effect index of each cell line they contain. The similarity between the two sets of values ​​is evaluated to obtain a similarity score for that culture effect index. The evaluation of the similarity between the two sets of values ​​can be based on indicators such as n-dimensional cosine similarity and n-dimensional Euclidean distance. In this embodiment, n-dimensional cosine similarity is used for evaluation.

[0100] The similarity scores of all culture effect indicators are combined to determine the similarity between the target cell line and each cell line contained in the reference formula sample database. The similarity between the target cell line and each cell line contained in the reference formula sample database can be the average, cumulative, weighted average, or weighted sum of the similarity scores of all culture effect indicators, etc.

[0101] Where y represents relevant indicators of culture effectiveness, such as cell viability, cell density, and protein expression levels. The formula for calculating cosine similarity is as follows:

[0102]

[0103]

[0104]

[0105] As can be seen from the above examples, reference cell line A is more similar to the target cell line than reference cell line B. Similarly, calculations were performed on all cell lines.

[0106] S4. The culture medium formulation for co-inoculating the target cell line and the reference cell line, the target indicator culture effect vector of the target cell line, and the target indicator culture effect vector of the reference cell line are used to form an experience sample. The experience sample set is collected as a meta-learning experience pool based on feature transfer; as shown in the table below:

[0107]

[0108] Comparison table of reference cell line A and target cell line

[0109] S5. Randomly sample from the feature transfer-based meta-learning experience pool obtained in step S4 to form a task set, and form a meta-training set from the task set. Heyuan Test Set The training set Heyuan Test Set As meta-learning sample data;

[0110] Supervised training is performed using the aforementioned meta-learning sample data, and the loss function used is:

[0111] Loss = metaloss + λ * mmdloss

[0112] Where metaloss represents the difference between the predicted value of the target cell line by the culture medium formulation culture effect prediction model and the target index culture effect vector of the target cell line in the meta-learning sample data; mmdloss represents the maximum mean difference between the predicted value of the target cell line by the culture medium formulation culture effect prediction model and the target index culture effect vector of the reference cell line in the meta-learning sample data; and λ represents the weight. Specifically:

[0113] The difference between the predicted values ​​of the culture medium formulation and the culture effect prediction model for the target cell line and the target index culture effect vector of the target cell line from the meta-learning sample data is calculated by Metaloss using the following method:

[0114]

[0115] Where Y o ′ bject The culture effect vector obtained by predicting the culture medium formulation of the training sample set of the meta-learning sample data for the culture medium formulation culture effect prediction model is of length dim, where the value of the d-th element, i.e., the d-th target index, is Y. o ′ bject [d];Y object The target cell line's culture performance vector is a training sample set of the meta-learning sample data, with length dim, where the d-th element, i.e., the d-th target indicator, has a value of Y. object [d], where n is the size of the training sample set.

[0116] Maximum Mean Discrepancy (MMD) is a term often used in transfer learning, especially in domain adaptation, to measure the distance between two different but related distributions.

[0117] The maximum mean difference (mmdloss) between the predicted values ​​of the target cell line and the target index culture effect vector of the reference cell line in the meta-learning sample data is used to quantitatively evaluate the gap between the distribution of the predicted values ​​of the target cell line and the distribution of the target index culture effect of the reference cell line.

[0118] (2) For the culture medium formulation to be predicted, the culture effect prediction model obtained in step (1) is used to predict the culture effect value.

[0119] The culture medium prediction method based on feature transfer meta-learning provided in this embodiment is used for formulation recommendation:

[0120] ① Verify the culture effect of existing culture medium sample formulations on target cell lines, and select formulations with high cell viability, high cell density, or high protein expression. From the top ten initial culture medium sample formulations with the best culture effect in the target cell line sample formulation database, and the maximum value of the component range given by the formulation designer, form a new search interval for optimal component values.

[0121] ② For the culture medium formulation to be predicted, the meta-learning culture medium formulation culture effect prediction model based on feature transfer obtained in step (1) is used to predict the culture effect of the culture medium formulation for the target cell line type. Based on the predicted culture effect, the optimal culture medium formulation is recommended. Heuristic algorithms such as genetic algorithms can be used to search for the optimal formulation within the search interval.

[0122] The culture medium formulation development method based on the meta-learning model of feature transfer described in this embodiment uses a genetic algorithm to search for culture medium formulations within a search interval to predict the culture effect. The above-mentioned heuristic algorithms include, but are not limited to, simulated annealing algorithm (SA), genetic algorithm (GA), ant colony algorithm (ACA), etc. This embodiment uses genetic algorithm (GA).

[0123] Cell productivity (Qp) refers to the amount of protein expressed per unit cell per unit time.

[0124] The following are the actual effects of recommended formulations obtained for different target cell lines with different culture indicators as the improvement targets.

[0125] Maximum cell density is measured in cells / ml; protein expression level is measured in mg / L; QP is measured in pg / cell / day.

[0126] Example 1: Target cell line 1 aims to improve QP.

[0127]

[0128] Example 2: Target cell line 2 aims to increase maximum cell density.

[0129]

[0130] Example 3: Target cell line 3 aims to increase protein expression levels.

[0131]

[0132] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A medium prediction method based on feature transfer meta-learning, characterized in that, Includes the following steps: (1) A meta-learner based on feature transfer is used to train the basic learner used to predict the culture effect of culture medium formulation, so as to obtain the culture effect prediction model of culture medium formulation. Its meta-learning sample data is obtained as follows: S1. From the historical cell line sample formulation database, select formulation samples with culture effect indicators including target indicators to form a reference formulation sample database; the formulation samples include: applicable cell lines, culture medium formulations, and culture effect indicators including target indicators. S2. For each formula sample in the reference formula sample database obtained in step S1, the target cell line is inoculated and cultured using its culture medium formula, and the culture effect values ​​of the same indicators as the culture effect indicators of the formula sample are collected to form the target cell line sample formula database. S3. For the target cell line sample formula database obtained in step S2 and the reference formula sample database obtained in step S1, based on the different culture effects of the same culture medium formula on the target cell line and the cell lines contained in the reference formula sample database, or the culture effects of different culture medium formula on the target cell line and the cell lines contained in the reference formula sample database, determine the similarity between the target cell line and the cell lines contained in the reference formula sample database, and take the cell line with the highest similarity as the reference cell line. S4. The culture medium formulation for co-inoculation of the target cell line and the reference cell line, the target index culture effect vector of the target cell line, and the target index culture effect vector of the reference cell line are used to form an experience sample. The experience sample set is collected as a meta-learning experience pool based on feature transfer. S5, randomly sampling from the meta-learning experience pool based on feature migration obtained in step S4 to form a task set, and forming a meta-training set from the task set and a meta-test set , the meta-training set and the meta-test set as meta-learning sample data; The meta-learning sample data is used for supervised training, and the loss function used is: : ; wherein a difference between a prediction value of the culture medium formula culture effect prediction model for the target cell strain and a target index culture effect vector of the meta-learning sample data target cell strain, a maximum mean difference between a prediction value of the culture medium formula culture effect prediction model for the target cell strain and a target index culture effect vector of the meta-learning sample data reference cell strain, is a weight; (2) For the culture medium formulation to be predicted, the culture effect prediction model obtained in step (1) is used to predict the culture effect value.

2. The culture medium prediction method based on feature transfer meta-learning as described in claim 1, characterized in that, The culture medium formula culture effect prediction model is used to predict the difference between the target index culture effect vector of the target cell strain and the meta-learning sample data target cell strain This is calculated as follows: ; wherein The culture effect vector of the culture medium formula obtained by predicting the meta-learning sample data training sample set by the culture medium formula culture effect prediction model has a length of , wherein the value of the th element, i.e., the value of the th target index is ; The target index culture effect vector of the target cell strain of the meta-learning sample data training sample set has a length of , wherein the value of the th element, i.e., the value of the th target index is , The size of the training sample set.

3. The medium prediction method based on feature transfer meta-learning according to claim 1, wherein, The culture medium formulation of the formulation sample described in step S1 has good culture effect and versatility for the target indicator, and is screened by sorting or thresholding method; the culture medium formulation with good culture effect has a culture effect indicator or weighted value of culture effect indicator exceeding the preset culture effect threshold; the culture medium formulation with good versatility has a culture effect exceeding the preset versatility threshold for more than a preset proportion of cell lines.

4. The culture medium prediction method based on feature transfer meta-learning as described in claim 3, characterized in that, The cultivation effect threshold is the value of the cultivation effect index, the weighted value of the cultivation effect index, or the ranking of the two; the ranking universality threshold is the value of the cultivation effect index, the weighted value of the cultivation effect index, or the ranking of the two.

5. The culture medium prediction method based on feature transfer meta-learning as described in claim 1, characterized in that, Step S3, the similarity determination, includes the following steps: For one or more culture medium formulations, obtain the index values ​​composed of various culture effect indicators of the target cell line and cell lines included in the reference formulation sample database. Evaluate the similarity between the two sets of values ​​for the target cell line and each cell line included in the reference formulation sample database to obtain a similarity score for the one or more culture medium formulations; or For each culture effect index, obtain the values ​​of all culture medium formulations in the target cell line sample formulation database for that culture effect index, and the values ​​of the corresponding culture medium formulations in the reference formulation sample database for each cell line it contains. Evaluate the similarity between the two sets of values ​​for the target cell line and each cell line contained in the reference formulation sample database to obtain a similarity score for that culture effect index. The similarity between the two sets of values ​​is evaluated based on indicators such as n-dimensional cosine similarity or n-dimensional Euclidean distance. The similarity scores of all culture effect indicators are used as the similarity between the target cell line and each cell line contained in the reference formula sample database; the similarity between the target cell line and each cell line contained in the reference formula sample database is the average, cumulative, weighted average, or weighted sum of the similarity scores of all culture effect indicators or culture medium formulas.

6. The culture medium prediction method based on feature transfer meta-learning as described in claim 1, characterized in that, The culture medium formulation culture effect prediction model in step (1) adopts the basic learner and meta-learner adaptation method. The basic learner includes fully connected neural network, convolutional neural network, recurrent neural network, or attention network model.

7. A culture medium prediction system based on feature transfer meta-learning, characterized in that, Includes basic learners and meta-learners; The learning sample data for the meta-learner is obtained according to the following method: S1. From the historical cell line sample formulation database, select formulation samples with culture effect indicators including target indicators to form a reference formulation sample database; the formulation samples include: applicable cell lines, culture medium formulations, and culture effect indicators including target indicators. S2. For each formula sample in the reference formula sample database obtained in step S1, the target cell line is inoculated and cultured using its culture medium formula, and the culture effect values ​​of the same indicators as the culture effect indicators of the formula sample are collected to form the target cell line sample formula database. S3. For the target cell line sample formula database obtained in step S2 and the reference formula sample database obtained in step S1, based on the different culture effects of the same culture medium formula on the target cell line and the cell lines contained in the reference formula sample database, or the culture effects of different culture medium formula on the target cell line and the cell lines contained in the reference formula sample database, determine the similarity between the target cell line and the cell lines contained in the reference formula sample database, and take the cell line with the highest similarity as the reference cell line. S4. The culture medium formulation for co-inoculation of the target cell line and the reference cell line, the target index culture effect vector of the target cell line, and the target index culture effect vector of the reference cell line are used to form an experience sample. The experience sample set is collected as a meta-learning experience pool based on feature transfer. S5. Randomly sample from the feature transfer-based meta-learning experience pool obtained in step S4 to form a task set, and form a meta-training set from the task set. Heyuan Test Set , the meta-training set Heyuan Test Set As meta-learning sample data; The meta-learner is trained under supervision using the meta-learning sample data, and the loss function used is... for: ; in The difference between the predicted values ​​of the culture medium formulation and culture effect prediction model for the target cell line and the target index culture effect vector of the target cell line in the meta-learning sample data is analyzed. The maximum mean difference between the predicted values ​​of the culture medium formulation and the target index culture effect vector of the reference cell line in the meta-learning sample data is used to predict the culture effect of the culture medium formulation and the target index culture effect vector. As weight; The basic learner includes fully connected neural networks, convolutional neural networks, recurrent neural networks, or attention network models.

8. The culture medium prediction system based on feature transfer meta-learning as described in claim 7, characterized in that, The difference between the predicted values ​​of the culture medium formulation and the culture effect prediction model for the target cell line and the target index culture effect vector of the target cell line based on the meta-learning sample data. Calculate using the following method: ; in The culture effect vector obtained by predicting the culture medium formulation culture effect of the meta-learning sample data training sample set for the culture medium formulation culture effect prediction model has a length of . , of which Element, i.e., the first The target value is ; The target indicator culture effect vector of the target cell line in the training sample set of the meta-learning sample data is of length . , of which Element, i.e., the first The target value is , This represents the size of the training sample set.

9. The culture medium prediction system based on feature transfer meta-learning as described in claim 7, characterized in that, The culture medium formulation of the formulation sample described in step S1 has good culture effect and versatility for the target indicator, and is screened by sorting or thresholding method; the culture medium formulation with good culture effect has a culture effect indicator or weighted value of culture effect indicator exceeding the preset culture effect threshold; the culture medium formulation with good versatility has a culture effect exceeding the preset versatility threshold for more than a preset proportion of cell lines.

10. The culture medium prediction system based on feature transfer meta-learning as described in claim 9, characterized in that, The cultivation effect threshold is the value of the cultivation effect index, the weighted value of the cultivation effect index, or the ranking of the two; the ranking universality threshold is the value of the cultivation effect index, the weighted value of the cultivation effect index, or the ranking of the two.

11. The culture medium prediction system based on feature transfer meta-learning as described in claim 7, characterized in that, Step S3, the similarity determination, includes the following steps: For one or more culture medium formulations, obtain the index values ​​composed of various culture effect indicators of the target cell line and cell lines included in the reference formulation sample database. Evaluate the similarity between the two sets of values ​​for the target cell line and each cell line included in the reference formulation sample database to obtain a similarity score for the one or more culture medium formulations; or For each culture effect index, obtain the values ​​of all formulas in the target cell line sample formula database for that culture effect index, and the values ​​of the corresponding formulas in the reference formula sample database for each cell line contained therein. Evaluate the similarity between the two sets of values ​​for the target cell line and each cell line contained in the reference formula sample database to obtain a similarity score for that culture effect index. The similarity between the two sets of values ​​is evaluated based on indicators such as n-dimensional cosine similarity or n-dimensional Euclidean distance. The similarity scores of all culture effect indicators are used as the similarity between the target cell line and each cell line contained in the reference formula sample database; the similarity between the target cell line and each cell line contained in the reference formula sample database is the average, cumulative, weighted average, or weighted sum of the similarity scores of all culture effect indicators or culture medium formulas.