Method for predicting production stability of clone that produces useful substance, information processing device, program, and prediction model generation method

JPWO2024048079A5Pending Publication Date: 2025-05-12
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024544002
Authority / Receiving Office
JP · JP
Patent Type
Applications
Priority Date
2023-07-07
Filing Date
2023-07-07
Publication Date
2025-05-12

AI Technical Summary

Technical Problem

Current methods for predicting the production stability of clones producing useful substances, such as biopharmaceuticals, are inaccurate and costly, requiring lengthy experimental verification and genetic analysis of numerous clones, which increases costs and reduces the number of high-stability clones identified.

Method used

A method that acquires and analyzes culture data to limit prediction targets, using machine learning to predict production stability based on gene expression data, thereby reducing costs and improving accuracy by focusing on clones with high antibody production potential.

Benefits of technology

This approach allows for high-accuracy prediction of production stability several months into the future, reducing the need for lengthy stability tests and genetic analysis of all clones, thereby lowering costs and efficiently identifying clones with stable production capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2024048079000001
    Figure 2024048079000001
Patent Text Reader

Abstract

Provided are: a method for predicting, at high accuracy and low cost, the production stability of a clone that produces a useful substance; an information processing device; a program; and a prediction model generation method. According to the present invention, at least one processor executes: acquiring at least one piece of clone culture data for a clone that produces a useful substance; analyzing the culture data to limit a clone to be predicted; and predicting the production stability of the useful substance by means of the clone to be predicted, by using data measured for the clone to be predicted. The production stability may be defined by the presence or absence of a change in the production amount of a useful substance at the start of culture and after culture for a predetermined period.
Need to check novelty before this filing date? Find Prior Art

Description

Method for predicting production stability of clones that produce useful substances, information processing device, program, and method for generating a prediction model

[0001] The present disclosure relates to information processing and machine learning techniques for predicting the production stability of clones that produce useful substances.

[0002] In recent years, industrial applications of manufacturing methods that allow cells to produce complex useful substances that are difficult to produce using conventional chemical synthesis have been expanding. One example is biopharmaceuticals, which account for more than half of the top 10 global drug sales and approximately two-thirds of total sales. Compared to traditional small molecule drugs, biopharmaceuticals utilize complex proteins and are extremely difficult to artificially synthesize chemically. For this reason, antibody drugs, an example of biopharmaceuticals, are widely produced by inserting a gene corresponding to a desired human protein into cells such as CHO (Chinese Hamster Ovary) cells, allowing the cells to produce the desired protein through their functions, which is then extracted and purified to produce the antibody drug.

[0003] Because precise control of gene insertion into cells as described above is impossible, it is common to insert genes into a large number of cells simultaneously. Given that the gene insertion position in each cell generated is random, many regulatory authorities require that the cells responsible for antibody production after gene insertion be derived from a single cell and that their properties do not change with subculture, a requirement known as monoclonality, in order to stabilize and guarantee the quality of antibodies as pharmaceuticals.

[0004] Therefore, monoclonality is ensured by extracting a single cell from individual cells with randomly inserted genes, growing that single cell to create a cell clone (hereinafter referred to as a clone), and having this clone produce an antibody. In the present invention, a clone refers to a population of genetically identical cells, or the cells that make up that population.

[0005] On the other hand, industrialization demands clones with high-quality antibody production capabilities. Here, high-quality antibody production capabilities refer to high antibody production capabilities at the current time and stable antibody production capabilities over long-term culture periods. As mentioned above, clones created from individual cells with random gene insertion positions vary in antibody production capabilities, and it is necessary to determine whether each clone has high-quality antibody production capabilities. Whether a clone currently has high antibody production capabilities can be determined by a two-week specification test. However, to determine production stability, i.e., whether antibody production capabilities are stable over long-term culture periods, experimental verification (stability testing) using actual long-term culture for several months is essential.

[0006] Against this background, Patent Document 1 proposes a method for predicting the stability of recombinant protein production of a clone several months into the future from currently available gene expression data of the clone. Also, Non-Patent Document 1 proposes a method for predicting the stability of recombinant protein production at an early stage of clone development by identifying marker genes that can predict stable expression of the recombinant protein at an early stage of clone development.

[0007] International Publication No. 2016 / 075216

[0008] Uros Jamnikar, Petra Nikolic, Ales Belic, Marjanca Blas, Dominik Gaser, Andrej Francky, Holger Laux, Andrej Blejec, Spela Baebler and Kristina Gruden, “Transcriptome study and identification of potential marker genes related to the stable expression of recombinant proteins in CHO clones” BMC Biotechnology volume 15, Article number 98 (2015).

[0009] However, the method described in Patent Document 1 cannot be said to be sufficient in terms of prediction accuracy. Furthermore, because genetic analysis of a large number of clones is generally expensive, there is also the problem that the cost reduction effect obtained by predicting the production stability of a recombinant protein is diminished by the increased costs due to the genetic analysis required for the prediction. While it is conceivable to narrow down the number of clones whose production stability is to be predicted in order to reduce costs, doing so would also reduce the number of clones with high production stability among the predicted clones, resulting in a smaller number of clones with high production stability being obtained, making it difficult to simply narrow down the number of clones to be predicted.

[0010] The first problem to be solved by the present disclosure is to provide a means for predicting the stability of useful substance production in clones with high accuracy, and the second problem is to provide a means for reducing the cost of predicting the stability of useful substance production in clones.

[0011] The present disclosure has been made in consideration of these circumstances, and aims to provide a method, an information processing device, a program, and a method for generating a prediction model that can predict the production stability of clones that produce useful substances with high accuracy and low cost.

[0012] A method according to a first aspect of the present disclosure is a method for predicting the production stability of a clone that produces a useful substance, in which one or more processors acquire culture data of one or more types of clones, analyze the culture data to narrow down the clones to be predicted, and predict the production stability of the useful substance by the clones to be predicted using data measured on the clones to be predicted.

[0013] According to the first aspect, since the production stability is predicted by limiting the prediction target based on information obtained from culture data, it is possible to predict the production stability with higher accuracy than when the target is not limited. Furthermore, since it is only necessary to obtain data necessary for prediction by limiting the target clone, it is possible to reduce costs.

[0014] The predicted production stability may represent the state of the clone several months into the future, similar to the production stability experimentally verified by long-term culture for several months. For example, production stability may be evaluated from the perspective of whether the initial production level is maintained even after long-term culture. According to the first aspect, the results of stability tests requiring long-term culture can be predicted with high accuracy and at low cost.

[0015] The method according to the second aspect of the present disclosure may be configured such that, in the method according to the first aspect, the production stability is defined by whether or not there is a change in the amount of useful substance produced between the start of culture and after a predetermined period of culture.

[0016] A method according to a third aspect of the present disclosure may be configured such that, in the method according to the first or second aspect, one or more processors set an index obtained from the culture data and a threshold value for the index, and limit the prediction target based on the value of the index and the threshold value.

[0017] The method according to the fourth aspect of the present disclosure may be configured such that, in the method according to the third aspect, the threshold value is adjusted so that the prediction accuracy of production stability is higher than when the prediction target is not limited.

[0018] The method according to the fifth aspect of the present disclosure may be configured such that, in the method according to the third or fourth aspect, the threshold is defined using a ranking of the index values. Note that the "rank" may be a ranking when the index values ​​of multiple clones are sorted in descending order or a ranking when the index values ​​are sorted in ascending order. For example, the threshold may be defined as the top 40% of the relative ranking in a population containing multiple clones.

[0019] A method according to a sixth aspect of the present disclosure is the method according to any one of the third to fifth aspects, wherein the prediction target may be a top group of index values.

[0020] A method according to a seventh aspect of the present disclosure is the method according to any one of the third to sixth aspects, wherein the indicator may be the amount of useful substance produced.

[0021] A method according to an eighth aspect of the present disclosure is the method according to any one of the third to sixth aspects, wherein the indicator may be an integrated viable cell density.

[0022] A method according to a ninth aspect of the present disclosure is the method according to any one of the third to sixth aspects, wherein the indicator may be a lactate concentration.

[0023] A method according to a tenth aspect of the present disclosure may be configured such that, in the method according to any one of the first to ninth aspects, the data used to predict production stability includes expression levels of one or more genes.

[0024] The method according to an eleventh aspect of the present disclosure may be configured in the method according to any one of the first to tenth aspects, wherein one or more processors receive input of data to be predicted and predict production stability using a model that performs two-class classification, either stable or unstable.

[0025] A method according to a twelfth aspect of the present disclosure may be the method according to the eleventh aspect, wherein the model is a model trained by machine learning using a plurality of training data sets in which data on training clones with similar limitations to the clone to be predicted is associated with a correct stability label.

[0026] A method according to a thirteenth aspect of the present disclosure may be configured in the method according to the twelfth aspect, wherein the plurality of training data includes training data for a plurality of types of clones that produce different useful substances, and the one or more processors predict the production stability of clones that produce useful substances other than the useful substances used to train the model.

[0027] A method according to a fourteenth aspect of the present disclosure is a method according to any one of the first to thirteenth aspects, wherein the useful substance may be any one of proteins, peptides, and viruses that are pharmaceutical raw materials.

[0028] A method according to a fifteenth aspect of the present disclosure is the method according to any one of the first to fourteenth aspects, wherein the useful substance may be an antibody or an antibody-like protein.

[0029] A method according to a sixteenth aspect of the present disclosure is the method according to any one of the first to fifteenth aspects, wherein the clone may be a cell derived from a vertebrate.

[0030] A method according to a seventeenth aspect of the present disclosure is the method according to any one of the first to fifteenth aspects, wherein the clone may be a cell derived from a mammal.

[0031] An eighteenth method of the present disclosure is the method according to any one of the first to fifteenth aspects, wherein the clone may be a CHO cell or a HEK cell (Human Embryonic Kidney cell).

[0032] An information processing device according to a nineteenth aspect of the present disclosure includes one or more processors and one or more storage devices storing instructions to be executed by the one or more processors, wherein the one or more processors acquire culture data of one or more types of clones that produce useful substances, analyze the culture data to narrow down the clones to be predicted, and predict the stability of production of the useful substance by the clones to be predicted using data measured on the clones to be predicted.

[0033] The information processing device according to the nineteenth aspect may have a configuration including the same aspect as the method according to any one of the second to eighteenth aspects.

[0034] A program according to a twentieth aspect of the present disclosure enables a computer to perform the following functions: acquire culture data of one or more clones that produce useful substances; analyze the culture data to limit the clones to be predicted; and predict the stability of useful substance production by the clones to be predicted using data measured on the clones to be predicted.

[0035] The program according to the twentieth aspect may be configured to include the same aspect as the method according to any one of the second to eighteenth aspects.

[0036] A predictive model generation method according to a twenty-first aspect of the present disclosure is a predictive model generation method for generating a predictive model that enables a computer to predict the production stability of clones that produce useful substances, and includes the steps of: a system including one or more processors acquiring culture data of one or more types of clones; analyzing the culture data to narrow down clones to be predicted; and performing machine learning using a plurality of training data in which measured data for the clones to be predicted are associated with correct stability labels, and training the predictive model so that the output of the predictive model in response to input data approaches the correct stability label.

[0037] The prediction model generation method according to the twenty-first aspect may have a configuration including the same aspects as any one of the methods according to the second to eighteenth aspects.

[0038] According to the present disclosure, the prediction target is appropriately limited based on information obtained by analyzing culture data, making it possible to predict the production stability of clones that produce useful substances with high accuracy. Furthermore, according to the present disclosure, by limiting the prediction target, the cost of predicting production stability can be reduced, making it possible to make predictions at low cost.

[0039] FIG. 1 is an explanatory diagram outlining the production process of antibody pharmaceuticals. FIG. 2 is a graph showing an example of changes in antibody production yield by clone. FIG. 3 is an explanatory diagram outlining the role of stability prediction AI (Artificial Intelligence) realized by this embodiment. FIG. 4 is a conceptual diagram of a machine learning model that predicts production stability based on gene expression data. FIG. 5 is an explanatory diagram outlining a method for predicting clone production stability according to this embodiment. FIG. 6 is a diagram showing an example of a dataset used for model training and evaluation. FIG. 7 is a graph showing an example of target narrowing using a certain index of culture data. FIG. 8 is a diagram showing an example of the number of clones of five types of antibody-producing CHO cells prepared as evaluation samples and the assignment of stability labels. FIG. 9 is a diagram showing an example of the number of clones for each antibody type whose antibody production yield values ​​fall in the top 40% of the relative ranking and the assignment of stability labels. FIG. 10 is a diagram showing an example of the number of clones for each antibody type whose integrated viable cell density values ​​fall in the top 60% of the relative ranking and the assignment of stability labels. FIG. 11 is a chart showing the number of clones for each antibody type whose lactate concentration values ​​fall in the top 40% of the relative ranking and an example of assigning stability labels. FIG. 12 is a block diagram showing the functional configuration of an information processing device according to an embodiment. FIG. 13 is a block diagram showing an example of the hardware configuration of an information processing device. FIG. 14 is a block diagram showing an example of the hardware configuration of a machine learning device that executes machine learning processing for generating a production stability prediction model. FIG. 15 is a flowchart showing an example of a machine learning method executed by the machine learning device. FIG. 16 is a flowchart showing an example of an information processing method executed by the information processing device according to an embodiment.

[0040] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings.

[0041] <<Outline of the antibody drug production process>> Antibody drugs, which are seeing an expanding market among biopharmaceuticals due to their high compatibility between efficacy and safety, are produced using clones of animal cells that can stably produce antibodies, proteins with complex structures. The following explanation uses antibodies as an example of a useful substance. Figure 1 is an explanatory diagram showing an overview of the antibody drug production process. The process leading up to the production of antibody drugs includes [1] the clone production phase, [2] the process development phase, and [3] the GMP (Good Manufacturing Practice) production phase.

[0042] The clone production phase includes the steps of: adding a vector to animal cells suitable for producing antibody pharmaceuticals and carrying out genetic modification to produce multiple clone candidates; and screening these multiple candidates for clones that are superior in terms of antibody production volume, cell proliferation, and quality stability (i.e., cell characteristics that do not change even after repeated proliferation).

[0043] The process development phase is a phase in which the production process (culture conditions, purification conditions, etc.) required for GMP production is developed using the screened clones.

[0044] In the GMP manufacturing phase, the clones are cultured and grown under an established production process, and the clones are made to produce antibodies. These antibodies are then purified and formulated to produce antibody drugs.

[0045] When producing antibodies in clones, it is necessary for productivity to remain stable over a long period of time. To achieve this, as many clones as possible are produced and clones with stable productivity are selected from these. However, this has traditionally required experimental verification over several months of continuous culture, which is a heavy burden.

[0046] Figure 2 is a graph showing an example of changes in antibody production levels depending on the clone. The vertical axis represents antibody productivity, and the horizontal axis represents elapsed time (time point). "Antibody productivity" is expressed as the amount of antibody produced by a clone per unit time.

[0047] FIG. 2 shows graphs plotting changes in the amount of antibody produced by clones over a long period of time (2 to 3 months). Graph G1 is a graph showing changes in antibody production for clones with stable productivity. Graph G2 is a graph showing changes in antibody production for clones with unstable productivity. As shown in graph G1, clones with stable productivity show little change in productivity even after 2 to 3 months have passed from the current time point, and can maintain productivity that is roughly the same as at the current time point. In contrast, as shown in graph G2, clones with unstable productivity gradually decrease in productivity over 2 to 3 months.

[0048] In the present invention, the "current time" refers to the time point at which the two-week standard test is conducted or the time point at which the standard test is completed, i.e., the time point at which the culture for determining the production stability is initiated. Furthermore, the "current antibody productivity" refers to the amount of antibody produced by the clone per unit time in the two-week standard test.

[0049] When antibody-producing cells are created by gene transfer, both stable and unstable clones are produced, as shown in Figure 2. Therefore, in the clone production phase, many types of clones are produced, and from these, clones with stable productivity that exhibit behavior similar to that of graph G1 are selected.

[0050] The productivity behavior shown in Figure 2 varies depending on the type of clone, and in the past, every time the type of antibody produced by a clone changed, an experiment similar to that shown in Figure 2 had to be performed to evaluate the production stability of each clone.

[0051] In contrast, an embodiment of the present disclosure proposes a mechanism for accurately predicting antibody production stability several months into the future based on information obtained from a clone at the current time. Here, "information obtained from the clone at the current time" refers to information obtained from the clone in a two-week specification test. The target variable for prediction, antibody production stability, can be defined as the presence or absence of a change in antibody production volume between the current time and several months into culture. Here, "several months" refers to, for example, a period of two months or more, e.g., two to three months. It may also be the period until a predetermined number of passages are performed. The period may be determined based on the proliferation ability of the clone or the culture period of the clone when actually producing the antibody. The "current time" refers to the initial time point of culture shown on the left side of the graph in Figure 2, i.e., the time when the two-week specification test is completed and the time when culture for determining antibody production stability begins. Productivity is "stable" when there is no change in antibody production volume between the current time and several months into the future. "No change" includes cases where the amount of change is within an acceptable range and can be considered to be substantially unchanged. "Unstable" productivity means that there is a change in the amount of antibody production between now and several months from now, and in many cases, the amount of production will decrease. The threshold value at which it is considered that there is a change in productivity can be set arbitrarily, and may be, for example, ±30% or ±20% of the amount of production at the current time.

[0052] <<Generalization Performance to Unknown Clones>> Figure 3 is an explanatory diagram illustrating the role of the stability prediction AI (Artificial Intelligence) realized by this embodiment. As shown in Figure 3, in the clone production phase, gene transfer is performed to introduce a genetic blueprint of a useful substance to be produced into a host cell. For example, if a blueprint for producing useful substance A is genetically transferred into a host cell, cells that produce useful substance A will be obtained. Since such producing cells are generated stochastically, cells that do not produce useful substance A or cells that produce it in an insufficient amount will also be generated. For this reason, a simple test is first performed at this stage to select high-producing clones that can sufficiently produce useful substance A.

[0053] Conventionally, a stability test would then be conducted for 2 to 3 months, as explained in Figure 2, to confirm whether useful substance A can be continuously produced over several months, and clones with stable production would be selected.

[0054] In this embodiment, a stability prediction AI is constructed as an alternative to conventional stability testing, and the stability prediction AI predicts the state (changes in productivity) two to three months later based on the profile obtained by measuring the current state of the clone, i.e., the state of the clone in a two-week standard test.

[0055] Because the types of useful substances (e.g., antibodies) produced by cells vary widely depending on the purpose, it is desirable to construct a model that can predict the stability of production regardless of the type of useful substance produced by the cells. In other words, a model that can robustly predict the stability of antibody production for unknown antibody species is preferred.

[0056] When training a model to be applied to stability prediction AI, the target useful substance cannot be known in advance, and the type of useful substance used when training the model may be different from the type of useful substance produced by the target clone that the model predicts after training. In other words, a model that can robustly and accurately predict production stability for unknown useful substance species is preferred, and it is preferable to construct a prediction model with domain generalizability, with the useful substance species as the domain.

[0057] Overview of Machine Learning Model for Predicting Production Stability In this embodiment, a stability prediction AI is constructed that, during the clone production phase, estimates (predicts) whether or not there will be a change in productivity two to three months from now based on information about the clone at the current time, i.e., makes it possible to predict the production stability of a useful substance. More specifically, a model is constructed that receives input of the clone's current gene expression data (at the time of specification testing) and outputs a stability label indicating the production stability of a useful substance. More specifically, some of the clones are used for specification testing, and another part is used for genetic analysis to obtain gene expression data, thereby obtaining gene expression data for the clones to be used for specification testing. The stability label can be expressed as two values, "1" indicating "stable" or "0" indicating "unstable." The prediction model for predicting production stability may be a two-class classification model that classifies clones into "stable" and "unstable."

[0058] The gene expression data includes one or more gene levels. The gene expression data used in this embodiment includes data that quantifies the gene expression levels of each of a plurality of genes. The gene expression data can be obtained, for example, by RNA (ribonucleic acid) sequence analysis. The value indicating the gene expression level is, for example, a count value that takes a positive integer, and can be used as a feature value after logarithmic transformation.

[0059] FIG. 4 is a conceptual diagram of a machine learning model MLM that predicts production stability based on gene expression data. An example of a training data set is shown inside the rectangular frame RF1 in FIG. 4. In FIG. 4, the current (specification test) gene expression data for each of multiple clones A to N is visualized as a gene expression pattern GEP using a heat map. The horizontal axis of the gene expression pattern GEP represents the type of gene, and the gene expression level of each of the multiple genes is represented by a two-color gradation (heat map). The number of types of genes a, b, c, d, etc. included in the gene expression data is preferably 300 to 400 types, for example, selected by obtaining all gene expression data for stable clones and unstable clones and using the statistical significance probability between the two groups of stable clones and unstable clones. To further narrow down the number of types of genes, the machine learning model MLM is actually trained using the selected genes while increasing or decreasing the number of types of genes, and the number of types that achieves high prediction performance is searched for, preferably narrowing down to, for example, 50 to 100 types of genes. Although all gene expression data was obtained here, it is not necessary to obtain all gene expression data; some genes may be randomly selected and their gene expression data may be obtained. Due to limitations in the illustration, the colors of the heat map cannot be displayed, so instead, red is displayed as "R," blue as "B," and white as "W." Red (R) represents a relatively high gene expression level, and blue (B) represents a relatively low gene expression level. White (W) represents an intermediate gene expression level.

[0060] Each of the multiple clones A to N has been confirmed to be "stable" or "unstable" based on experimental verification through cultivation for several months after specification testing, and a stability label (correct label) indicating "stable" or "unstable" is assigned to each clone A to N. In this way, a dataset is prepared containing multiple training data in which the current gene expression data of each of the multiple clones A to N is associated (linked) with the correct stability label. Then, the machine learning model MLM is trained using the multiple training data, and the machine learning model MLM learns stable or unstable gene patterns. When the current gene expression data (at the time of specification testing) of an unknown clone X is input to the trained machine learning model MLM, the machine learning model MLM predicts production stability from the input gene expression data and outputs a label of "stable" or "unstable" as the prediction result. Note that Figure 4 shows an example in which the machine learning model MLM predicts that the unknown clone X is "stable."

[0061] Overview of the embodiment: Building a model to predict useful substance production stability by limiting the prediction target Because clones that produce useful substances have various characteristics, it is difficult to predict the production stability of all clones, regardless of type, with high accuracy. In this embodiment, high-accuracy predictions are achieved by limiting the prediction target based on indicators obtained from the culture data of each clone at the current time (at the time of specification testing). Here, the culture data refers to general data that can be measured for clones using a culture device or a dedicated device by sampling a portion of the culture medium containing cells.

[0062] Fig. 5 is an explanatory diagram showing an overview of the method for predicting the stability of production of useful substances from clones according to this embodiment. F5A on the left side of Fig. 5 shows a comparative example in which the prediction target is not limited, and F5B on the right side of Fig. 5 shows an overview of the method according to this embodiment.

[0063] The following describes the case where the prediction target in Figure F5A on the left is not limited. Within the rectangular frame RF2 in Figure F5A on the left, a dataset DSc containing training data for multiple clones producing useful substances A to D is schematically shown. This dataset DSc includes training data for a total of 20 clones, five clones for each of useful substances A to D. The training data here is data in which gene expression data from standard testing for each of the 20 clones is associated (linked) with the correct stability label. The values ​​"9," "7," and "6" displayed below each clone in Figure 5 represent the measured values ​​of certain culture data for each clone during standard testing. Note that instead of the measured values, the relative levels of each clone that can be obtained from the measured values ​​may also be represented. While Figure F5A on the left has been described here, the same applies to Figure F5B on the right.

[0064] In the left diagram F5A, the training data is not limited, and the machine learning model MLMc is trained using all the training data in the dataset DSc, and the learned (trained) model is used to predict the production stability of multiple clones that produce an unknown useful substance X. In this case, the multiple clones that produce the unknown useful substance X to be predicted are not particularly limited, and as shown in the rectangular frame RF3, the production stability is predicted for all five clones that produce the unknown useful substance X. The production stability is predicted by obtaining current (at the time of specification testing) gene expression data for all five clones that produce the unknown useful substance X and inputting it into the learned (trained) model, but the prediction accuracy is low.

[0065] Next, the method according to this embodiment, shown in Figure F5B on the right, will be described. In contrast to the method shown in Figure F5A on the left, which does not limit the prediction target, the method shown in Figure F5B limits the prediction target using the value of certain culture data obtained during specification testing as an index. First, a threshold is determined based on the value of certain culture data, and the population of clones contained in dataset DSd is divided into two groups: those with a relatively large index culture data value relative to the threshold, and those with a small index culture data value. In this example, the threshold is set to "5," and populations with index culture data values ​​of "5" or greater are targeted for training, while populations with culture data values ​​below "5" are excluded. This threshold processing leaves training data for a total of 12 clones, three for each of useful substances A through D, as shown in the rectangular frame RF4. The dataset DSe containing the training data of these limited populations is used to train the machine learning model MLMe. Meanwhile, the training data for the eight clones shown in the dashed rectangular frame RF5, i.e., the training data for clones that do not meet the threshold conditions, is excluded from processing.

[0066] In this way, the machine learning model MLMe is trained using the dataset DSe with a limited target. Then, when predicting the production stability of clones producing an unknown useful substance X using the trained model, the clones to be predicted are limited to those that satisfy the threshold-based restriction conditions (a population with index values ​​higher than the threshold), just like the population of clones in the dataset DSe used to train the model, by applying a threshold to the values ​​of the culture data used as an index. The three types of clones shown in the rectangular frame RF6 represent clones that fall within the prediction target. Furthermore, the two types of clones shown in the dashed rectangular frame RF7 represent clones that are not the prediction target. By limiting the prediction target in this way, high prediction accuracy can be achieved. Furthermore, since the clones not the prediction target shown in the dashed rectangular frame RF7 do not require gene expression data acquisition, the cost of genetic analysis can be reduced.

[0067] <<Example of Dataset Used for Training and Evaluation>> Figure 6 shows an example of a dataset used for training and evaluation of a model. The upper part of Figure 6 shows an example of dataset DSA for clones producing antibody A as a useful substance, and the lower part shows an example of dataset DSB for clones producing antibody B as a useful substance. Although not shown, similar datasets are also available for clones producing other types of antibodies as useful substances.

[0068] The dataset DSA includes culture data measured for each of the multiple clones ACLj during specification testing, gene expression data measured during specification testing, and the correct stability label obtained through stability testing. The subscript j represents an index number identifying the clone. The culture data may include one or more items, such as antibody production, integral viable cell density (IVCD), lactate concentration, and pH. The culture data may be general data that can be measured using a culture device or a dedicated device by sampling a portion of the cell-containing culture medium, and may include, for example, one or more of the total cell count, the amount of cell-secreted substances, the amount of cell-produced substances, the amount of cell metabolites, and the amount of medium components. The letter symbols (symbols with the subscript j) in each cell of the table in Figure 6 represent the value of the corresponding data item.

[0069] The same applies to dataset DSB. The number na of clones ACLj included in dataset DSA and the number nb of clones BCLj included in dataset DSB may be different.

[0070] From the data sets of multiple domains (useful substance types) prepared in this way, the target is narrowed down (limited) by focusing on certain culture data indicators.

[0071] <<Example of Narrowing Down Prediction Targets>> Figure 7 is a graph showing an example of narrowing down prediction targets using a certain index of culture data. The horizontal axis shows multiple types of clones that produce each of multiple useful substances A to E. The vertical axis shows the value of a certain index obtained from culture data during specification testing. The clones shown in Figure 7 are clones used for model training (learning).

[0072] As shown in Figure 7, the distribution range of a certain index obtained from culture data may differ depending on the clones producing different types of useful substances. In this case, as explained in Figure 5, if a threshold value is set for the index value, and the clones are divided into two groups based on their relative magnitude relative to the threshold, and one group of clones is used for training and the other group is not used for training, the number of clones to be trained will vary depending on the type of useful substance produced. For example, if the threshold value for the index is set to 2.5 and the group of clones with values ​​above the threshold are used for training, clones producing useful substance B will not be used for training.

[0073] Therefore, for example, as shown in FIG. 7 , for clones producing each of useful substances A to D, training subjects may be limited to those in the relatively top X% (Top-X%) of a certain index obtained from the culture data. Here, the relative top X% refers to the top X% (Top-X%) when a group of clones producing each of useful substances A to D is sorted in descending order for a certain index obtained from the culture data. The criterion of "X%", which corresponds to the threshold that serves as the limiting condition, is preferably adjusted so that the number of samples from each of useful substances A to E is roughly the same. The relative top X% is an example of a "threshold defined using the ranking of index values" in the present disclosure.

[0074] Even if the training subjects are limited in this way, there may be clones with stable productivity and clones with unstable productivity. When predicting the production stability of clones that produce an unknown useful substance Y using a learned (trained) model, the clones to be predicted are limited to the top X% of clones with respect to a certain index obtained from the culture data, just like the clones used to train the model.

[0075] Here, multiple types of clones producing each of multiple useful substances A to E were used to train the model, but it is not necessary to use multiple types of clones producing different useful substances; for example, only clones producing useful substance A may be used for training. In this case, the method of limiting the population of clones used for training may be to set a threshold by focusing on the value of certain culture data during specification testing and to base the clones' relative magnitude relative to the threshold, or to limit the population of clones to the top X% with respect to the value of a certain index obtained from the culture data during specification testing. Furthermore, although the relative top X% is used, the clones may also be limited to the relative bottom X% depending on a certain index obtained from the culture data.

[0076] The culture data index and threshold for limiting the prediction target may be determined by repeating hypothesis and verification using a prepared data set through trial and error. Alternatively, the culture data index and threshold for limiting the prediction target can be determined by performing exploratory analysis using a prepared data set.

[0077] For example, in the case of a data set of five useful substances A to E (domains) as shown in FIG. 7 , an information processing device including a processor uses a feature selection method, such as a filter method, to evaluate the degree of association between each feature and the target variable (stability label) in each of the five domains, and determines features with high association in, for example, four or more of the five domains as features with high domain universality. The information processing device focuses on a certain index from all data and extracts data that meets specific conditions as a subset, and performs a domain generalization evaluation on the extracted subset based on the number of features with high domain universality. If the number of features with high domain universality is large, the subset is evaluated as having high domain generality. By using data from the subset with high domain generality as training data to learn (train) a prediction model, the learned model can robustly predict production stability for other domains (useful substance species) for a population (subset) limited to the same conditions as during learning.

[0078] In the case of antibody-producing clones, culture data indices that are effective in limiting the target of prediction include, for example, antibody production amount, integral viable cell density, and lactate concentration, and it was confirmed that highly accurate prediction of production stability is possible by targeting the population with the highest values ​​of any of these indices.

[0079] <<Examples of Useful Substances>> The useful substance is not limited to antibodies, but may be antibody-like proteins. The useful substance may be any of proteins, peptides, and viruses, which are pharmaceutical raw materials.

[0080] <<Examples of Clones>> The clones producing useful substances may be cells derived from vertebrates. The clones may be, for example, cells derived from mammals. The clones may be CHO cells or HEK cells.

[0081] Examples 1 to 3, which apply the technology of the present disclosure, are described below. The common configurations of Examples 1 to 3 are as follows: Specifically, the useful substance is an antibody, and the producing cells are CHO cells. Multiple clones of five types of antibody-producing CHO cells were prepared as evaluation samples. 100 gene expression levels were selected from the total gene expression levels measured in a two-week standard test using RNA sequencing (RNA-Seq) analysis and used as explanatory variables. A logistic regression model that classifies genes into two classes, stable or unstable, was used as the learning device. Five-fold cross-validation was performed to train (learn) the prediction model, and performance was evaluated using PRAUC (Area Under the Precision-Recall Curve). In Examples 1 to 3, the number of gene expression levels used as explanatory variables was determined to be 100 by actually training (learning) the prediction model while increasing or decreasing the number of levels using 300 to 400 genes selected using statistical significance probability, and by searching for the number of levels that would result in the highest prediction performance. The standard test was performed by seeding the clone (CHO cells) at a density of 5 x 10^5 cells / mL in a 40 mL flask using suspension culture.

[0082] In the five-fold cross-validation, the dataset was divided into five antibody types, and performance was evaluated using untrained antibody types. That is, the datasets of four antibody types were used as training (learning) data, and the dataset of the remaining antibody type was used as test data for performance evaluation.

[0083] FIG. 8 is a chart showing an example of the number of clones and stability labels assigned to five types of antibody-producing CHO cells prepared as evaluation samples. 182 clones of five types of antibody-producing cells were prepared as evaluation samples, and each clone was assigned a stability label ("stable" or "unstable") by culturing the cells for two months under the same conditions as in the standard test (see FIG. 8 ). For example, there were a total of 24 clones producing antibody A, of which 7 clones were labeled "stable" and 17 clones were labeled "unstable." Furthermore, for each of the 182 clones, culture data and gene expression data were obtained during the standard test, and the gene expression data and stability label were linked for each clone to form training data.

[0084] Example 1 In Example 1, an example of stability prediction will be described in which the prediction target is limited to "clones with relatively high productivity." Here, "clones with relatively high productivity" refers to clones that produce a useful substance in a relatively high amount.

[0085] This section describes a method for limiting the clones to be trained when training a prediction model, when stability prediction is performed by limiting the prediction targets to "clones with relatively high productivity." Note that limiting the clones to be trained corresponds to limiting the clones to be predicted by the prediction model during training, i.e., limiting the clones to be predicted by the prediction model.

[0086] The method for limiting the clones to be trained was to focus on the "antibody production amount" from the culture data of all 182 clones during standard testing, search for a threshold that would improve the predictive performance of the prediction model, and limit the clones to the top 40% of the relative ranking for each antibody type. Here, "antibody production amount" can be, for example, the cumulative amount of antibody production over a two-week (14-day) period during standard testing. Alternatively, it can be the cumulative amount of antibody production over a certain period during standard testing, such as a 10-day period, or it can be divided by the measurement period to obtain the antibody production amount per unit time. "Top 40%" is an example of a threshold. Figure 9 is a chart showing the number of clones for each antibody type whose antibody production amount falls within the top 40% of the relative ranking, and an example of the assignment of stability labels.

[0087] Figure 9 shows an example of a total of 73 clones that fall into the top 40% of relative ranking for each antibody type. Five-fold cross-validation was performed using datasets for each antibody type, including training data linked to gene expression data and stability labels from specification testing for the 73 clones shown in Figure 9. As a result of limiting the training subjects in this way, the predictive performance of the trained prediction model achieved a PRAUC value of 0.743. When predicting the production stability of an unknown useful substance using a trained prediction model, the clones to be predicted are also limited to the top 40% of "antibody production amount" by analyzing culture data from specification testing (current time), as in the case of limiting the training subjects.

[0088] [Comparative Example] In contrast, when similar learning was performed using a dataset containing all data of the 182 clones shown in Figure 8 without limiting the training target, and five-fold cross-validation was performed, the prediction performance of the prediction model according to the comparative example obtained had a PRAUC value of 0.503. Note that the prediction target is not limited to the same target as the training target. It was confirmed that the performance of the prediction model in which the prediction target was limited in Example 1 was more accurate than the prediction model according to the comparative example.

[0089] These results demonstrate that the prediction model generated by the method of Example 1 can be used to make highly accurate predictions for unknown useful substances. At the same time, limiting the selection to clones with relatively high productivity does not pose any obstacle to the process of selecting clones that produce useful substances, and since the selection is limited to targets that can be predicted with high accuracy, the stability predictions disclosed herein are considered to be practical.

[0090] Example 2 In Example 2, an example of stability prediction in which prediction targets are limited to "clones with relatively high cell densities" will be described. First, a method for limiting clones to be trained for training a prediction model when performing stability prediction in which prediction targets are limited to "clones with relatively high cell densities" will be described. As in Example 1, a method was used in which the "integral viable cell density (IVCD)" was focused on from the culture data of all 182 clones during specification testing shown in FIG. 8 , a threshold was searched for to improve the prediction performance of the prediction model, and the clones were limited to the top 60% of the relative ranking for each antibody type. Here, "clones with relatively high cell densities" can be obtained, for example, based on the "integral viable cell density (IVCD)" over a two-week (14-day) period in specification testing. Alternatively, they may be obtained based on the "integral viable cell density (IVCD)" over a certain period during specification testing, for example, 10 days. The "top 60%" is an example of a threshold. FIG. 10 is a chart showing the number of clones for each antibody type whose integral viable cell density values ​​fall in the top 60% of the relative ranking and an example of the stability label assignment.

[0091] Figure 10 shows an example of a total of 109 clones corresponding to the top 60% of relative rankings for each antibody type. Five-fold cross-validation was performed using datasets for each antibody type, including training data linked to gene expression data and stability labels during specification testing for the 109 clones shown in Figure 10. As a result of limiting the training subjects in this way, the predictive performance of the trained prediction model achieved a PRAUC value of 0.647. In other words, it was confirmed that the performance of the prediction model in which the prediction subjects were limited in Example 2 was more accurate than the PRAUC (0.503) of the prediction model in the comparative example in which the subjects were not limited. Furthermore, when predicting the production stability of an unknown useful substance using a trained prediction model, the clones to be predicted were also analyzed using culture data from specification testing (current time), as in the case of limiting the training subjects, and predictions were made by limiting the clones to the top 60% of "integral viable cell density (IVCD)."

[0092] These results demonstrate that the prediction model generated by the method of Example 2 can be used to make highly accurate predictions for unknown useful substances. At the same time, limiting the selection to clones with a relatively high viable cell density does not hinder the process of selecting clones that produce useful substances, and by limiting the selection to targets that can be predicted with high accuracy, the process can be carried out at low cost. Therefore, it is believed that the stability predictions disclosed herein are practical.

[0093] Example 3 In Example 3, an example of stability prediction limited to "clones with relatively high lactate concentrations" is described. First, a method for limiting the clones to be trained for training a prediction model when performing stability prediction limited to "clones with relatively high lactate concentrations" is described. As in Example 1, the "lactate concentration" of the culture medium in which the clones were cultured was focused on from the culture data of the two-week specification test for all 182 clones shown in Figure 8. The "lactate concentration" of each clone was obtained by using the median value of the "lactate concentration" of the culture medium measured at each time point within the two weeks (14 days), for example, every day, as a representative value. A threshold was then searched for to improve the prediction performance of the prediction model, and the method was used to limit the clones to the top 40% of the relative ranking for each antibody type. The "top 40%" is an example of a threshold. Figure 11 is a chart showing the number of clones for each antibody type whose lactate concentration values ​​fall into the top 40% of the relative ranking and an example of the stability label assignment.

[0094] Figure 11 shows examples of 72 clones in total that fall into the top 40% of relative rankings for each antibody type. The reason for the number of clones being one less than in Figure 9 is that data for one clone was missing in the measurement of lactate concentration.

[0095] Five-fold cross-validation was performed using a dataset for each antibody type, including training data linked to gene expression data and stability labels during specification testing for the 72 clones shown in Figure 11. As a result of limiting the target in this way, the predictive performance of the trained prediction model had a PRAUC value of 0.613. In other words, it was confirmed that the performance of the prediction model in which the prediction target was limited in Example 3 was more accurate than the PRAUC (0.503) of the prediction model in the comparative example in which the target was not limited. In addition, when predicting the production stability of an unknown useful substance using a trained prediction model, the culture data at the time of specification testing (current time) was analyzed, as in the case of limiting the training target, and predictions were made by limiting the "lactic acid concentration" to the top 40%.

[0096] These results demonstrate that it is possible to make highly accurate predictions for unknown useful substances. At the same time, limiting the selection to clones with relatively high lactate concentrations does not hinder the process of selecting clones that produce useful substances, and by limiting the selection to targets that can be predicted with high accuracy, it can be carried out at low cost. Therefore, it is believed that the stability predictions disclosed herein are practical.

[0097] <<Configuration Example of Information Processing Device>> Fig. 12 is a block diagram showing the functional configuration of an information processing device 10 according to an embodiment. The information processing device 10 includes a data acquisition unit 12, a prediction target limitation unit 14, a production stability prediction model 16, and a processing result output unit 18. The various functions of the information processing device 10 can be realized by a combination of computer hardware and software. The physical form of the information processing device 10 is not particularly limited, and may be a server computer, a workstation, a personal computer, a tablet terminal, or the like.

[0098] The data acquisition unit 12 acquires various data including culture data and gene expression data of one or more types of clones that produce useful substances.

[0099] The prediction target limiting unit 14 includes a culture data analysis unit 20 and a limiting condition determination unit 22, and analyzes input culture data of one or more types of clones to limit the clones to be predicted. The culture data analysis unit 20 analyzes the culture data. The limiting condition determination unit 22 limits the target using a threshold value based on the analysis results of the culture data. For convenience of explanation, the culture data analysis unit 20 and the limiting condition determination unit 22 are described separately, but the limiting condition determination unit 22 may be included in the culture data analysis unit 20. It may also be understood that the culture data analysis unit 20 functions as the prediction target limiting unit 14.

[0100] The culture data analysis unit 20 may execute a process of determining an index and a threshold for limiting the prediction target from the input data set. The index and threshold that are the limiting conditions for the prediction target may be set based on the analysis results by the culture data analysis unit 20, or may be set in the prediction target limiting unit 14 as known information that is known in advance based on the results of a search process using another information processing device (not shown).

[0101] A machine learning model is applied to the production stability prediction model 16. The production stability prediction model 16 may be a two-class classification model that accepts input of current gene expression data of the clone to be predicted, predicts the production stability of the clone based on the input gene expression data, and outputs a stability label. The production stability prediction model 16 is trained using training data that is limited to a specific target by the method described in the right diagram F5B of Figure 5. The gene expression data input to the production stability prediction model 16 includes one or more gene expression levels. The gene expression data input to the production stability prediction model 16 may also include data on the expression levels of multiple genes. Features used as explanatory variables may be selected by known feature selection methods.

[0102] The processing result output unit 18 outputs the processing result including the prediction result of the production stability prediction model 16. The processing result output unit 18 may be configured to perform at least one of the following processes: displaying the processing result, recording the processing result in a database or the like, and printing the processing result.

[0103] 13 is a block diagram showing an example of the hardware configuration of the information processing device 10. Here, an example will be described in which the processing functions of the information processing device 10 are realized using one computer, but the processing functions of the information processing device 10 may also be realized by a computer system configured using multiple computers.

[0104] The information processing device 10 includes a processor 102, a computer-readable medium 104 which is a non-transitory tangible entity, a communication interface 106, an input / output interface 108, and a bus 110. The processor 102 is connected to the computer-readable medium 104, the communication interface 106, and the input / output interface 108 via the bus 110.

[0105] The processor 102 includes a central processing unit (CPU). The processor 102 may also include a graphics processing unit (GPU). The computer-readable medium 104 includes a memory 112, which is a main storage device, and a storage 114, which is an auxiliary storage device. The computer-readable medium 104 may be, for example, a semiconductor memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination of these. The computer-readable medium 104 is an example of a "storage device" in this disclosure.

[0106] The computer-readable medium 104 includes a data storage area 120 that stores various data, such as culture data and gene expression data for one or more clones. The computer-readable medium 104 also stores a plurality of programs, including a prediction target limitation program 140, a production stability prediction model 16, a processing result output program 180, and a display control program 190, as well as data. The term "program" includes the concept of a program module and includes instructions equivalent to a program. The processor 102 functions as various processing units by executing the instructions of the programs stored in the computer-readable medium 104.

[0107] The prediction target limitation program 140 includes instructions for executing a process of analyzing culture data and limiting prediction targets. The prediction target limitation program 140 may be configured to include a culture data analysis program 142 and a limiting condition determination program 144. The culture data analysis program 142 includes instructions for executing a process of analyzing culture data of one or more types of clones. The culture data analysis program 142 may include instructions for executing a process of searching for indicators and thresholds for narrowing down prediction targets from a dataset.

[0108] The limiting condition determination program 144 includes an instruction to execute a process of limiting the prediction target based on the index and threshold value defined as the limiting condition, using the analysis results of the culture data analysis program 142 .

[0109] The production stability prediction model 16 includes instructions for accepting input of gene expression data of a clone related to a prediction target that satisfies a limiting condition, and for executing a process for predicting production stability.

[0110] The processing result output program 180 includes instructions for executing a process for outputting processing results including the production stability predicted by the production stability prediction model 16. The display control program 190 includes instructions for generating a display signal required for display output on the display device 154 and for executing display control of the display device 154.

[0111] The communication interface 106 performs communication processing with an external device via a wired or wireless connection, and exchanges information with the external device. The information processing device 10 is connected to a communication line (not shown) via the communication interface 106. The communication line may be a local area network, a wide area network, or a combination of these. The communication interface 106 can serve as a data acquisition unit that accepts input data.

[0112] The information processing device 10 may include an input device 152 and a display device 154. The input device 152 may be, for example, a keyboard, a mouse, a multi-touch panel, or other pointing device, or a voice input device, or an appropriate combination thereof. The display device 154 may be, for example, a liquid crystal display, an organic electro-luminescence (OEL) display, a projector, or an appropriate combination thereof. The input device 152 and the display device 154 are connected to the processor 102 via the input / output interface 108. Note that the input device 152 and the display device 154 may be integrated into one unit, such as a touch panel, or the information processing device 10, the input device 152, and the display device 154 may be integrated into one unit, such as a touch panel tablet terminal.

[0113] 14 is a block diagram showing an example of the hardware configuration of a machine learning device 300 that executes machine learning processing to generate a production stability prediction model 16. Here, an example is described in which the processing functions of the machine learning device 300 are realized using one computer, but the processing functions of the machine learning device 300 may also be realized by a computer system configured using multiple computers.

[0114] The machine learning device 300 includes a processor 302, a non-transitory tangible computer-readable medium 304, a communication interface 306, an input / output interface 308, and a bus 310. The computer-readable medium 304 includes a memory 312 and a storage 314. The processor 302 is connected to the computer-readable medium 304, the communication interface 306, and the input / output interface 308 via the bus 310. An input device 352 and a display device 354 are connected to the bus 310 via the input / output interface 308.

[0115] The hardware configuration of the machine learning device 300 may be similar to the corresponding elements of the information processing device 10 described in FIG. 6. The machine learning device 300 may be in the form of a server computer, a personal computer, or a workstation. The machine learning device 300 is an example of a "system including one or more processors" in the present disclosure.

[0116] The machine learning device 300 is connected to a communication line (not shown) via a communication interface 306 and is communicatively connected to external devices such as a data storage unit 550. The data storage unit 550 includes a storage in which a dataset including multiple pieces of training data is stored. The data storage unit 550 may store a dataset including all data from multiple domains as illustrated in FIG. 6, or may store a dataset including only data of samples of a limited target to be predicted. The data storage unit 550 may be built in the storage 314 within the machine learning device 300.

[0117] The computer-readable medium 304 stores a plurality of programs, data, etc., including a prediction target limitation program 320, a learning processing program 330, and a display control program 340. The prediction target limitation program 320 may be similar to the prediction target limitation program 140 described in Fig. 12. The display control program 340 may be similar to the display control program 190 described in Fig. 12.

[0118] The computer-readable medium 304 includes a prediction target data storage area 322. Training data corresponding to limited prediction targets is stored in the prediction target data storage area 322. The corresponding training data may be sampled in a timely manner by the prediction target limitation program 320 from a dataset stored in the data storage unit 550, or a dataset containing only the prediction target may be extracted in advance as a subset.

[0119] The learning processing program 330 includes a data acquisition program 400, a prediction model 410 which is a machine learning model, a loss calculation program 430, and an optimizer 440. The data acquisition program 400 includes instructions for executing a process of acquiring training data from the prediction target data storage area 322. The training data acquired via the data acquisition program 400 is input to the prediction model 410.

[0120] The loss calculation program 430 includes instructions for executing a process of calculating a loss indicating the error between the predicted value of the stability label output from the prediction model 410 and the correct stability label. The optimizer 440 includes instructions for executing a process of calculating an update amount for the parameters of the prediction model 410 from the calculated loss and updating the parameters of the prediction model 410. The optimizer 440 may optimize the parameters using a method such as stochastic gradient descent (SGD).

[0121] <<Flowchart of Machine Learning Method>> Figure 15 is a flowchart showing an example of a machine learning method executed by the machine learning device 300. Here, the description will be made assuming that a dataset to be used for machine learning, such as the one shown in Figure 6, is prepared. In step S102, the processor 302 acquires culture data from the prepared dataset.

[0122] In step S104, the processor 302 analyzes the culture data and limits the training subjects. The processor 302 may select data of a target sample that satisfies the limiting condition or data of a non-target sample that does not satisfy the limiting condition according to a pre-specified index and threshold of the culture data, or may search for an index and threshold that are the limiting condition from the culture data and select data of a target sample from data of a non-target sample.

[0123] In step S106, the processor 302 performs machine learning using only data on clones that satisfy the restriction conditions to train the prediction model 410. That is, the processor 302 inputs gene expression data of samples that satisfy the restriction conditions into the prediction model 410 and calculates a loss indicating the error between the predicted value of the stability label output from the prediction model 410 and the correct stability label. The processor 302 calculates the amount of update for the parameters of the prediction model 410 based on the calculated loss and updates the parameters. In this way, the processor 302 trains the prediction model 410 so that the output (predicted value) from the prediction model 410 for the data input to the prediction model 410 approaches the correct stability label. Note that the parameter update for the prediction model 410 may be performed in mini-batch units.

[0124] In step S108, the processor 302 determines whether to terminate learning. The termination condition for learning may be determined based on the loss value or the number of parameter updates. In a method based on the loss value, for example, the termination condition for learning may be that the loss has converged within a specified range. In a method based on the number of updates, for example, the termination condition for learning may be that the number of updates has reached a specified number. Alternatively, a data set for evaluating the performance of the model may be prepared separately from the training data, and whether to terminate learning may be determined based on an evaluation value using the evaluation data.

[0125] If the determination result in step S108 is No, the processor 302 returns to step S106 and continues the learning process. On the other hand, if the determination result in step S108 is Yes, the processor 302 ends the flowchart in FIG.

[0126] The trained prediction model 410 is incorporated into the information processing device 10 as the production stability prediction model 16. The machine learning method executed by the machine learning device 300 can be understood as a method for generating the production stability prediction model 16, and is an example of a prediction model generation method in the present disclosure.

[0127] 16 is a flowchart showing an example of an information processing method executed by the information processing device 10. In step S202, the processor 102 acquires culture data measured on clones that produce useful substances. The processor 102 may automatically acquire the data from a data storage server (not shown) or may accept input of designated data via a user interface and acquire data on the designated clones.

[0128] In step S204, the processor 102 analyzes the culture data and limits the prediction target. The processor 102 limits the prediction target by applying the same limiting conditions as those used to limit the training target when training the production stability prediction model 16. After the prediction target is limited in step S204, gene expression data is measured for clones corresponding to the prediction target, thereby reducing the workload and costs compared to performing genetic analysis on all clones.

[0129] In step S206, the processor 102 inputs the gene expression data of the clone corresponding to the prediction target into the production stability prediction model 16, and predicts stability using the production stability prediction model 16.

[0130] In step S208, the processor 102 outputs the prediction result output from the production stability prediction model 16. Based on this prediction result of production stability, production clones can be selected.

[0131] After step S208, the processor 102 ends the flowchart of FIG.

[0132] <<Regarding the program that operates a computer>> A program that causes a computer to realize some or all of the processing functions of each device of the information processing device 10 and the machine learning device 300 according to the embodiment can be recorded on a computer-readable medium that is a non-transitory information storage medium such as an optical disk, a magnetic disk, a semiconductor memory, or other tangible object, and the program can be provided through this information storage medium.

[0133] In addition, instead of providing the program by storing it on such a tangible, non-transitory computer-readable medium, it is also possible to provide the program signal as a download service using a telecommunications line such as the Internet.

[0134] Furthermore, some or all of the processing functions of each of the above-mentioned devices may be realized by cloud computing, and may also be provided as SaaS (Software as a Service).

[0135] <<Hardware Configuration of Each Processing Unit>> The hardware configuration of processing units that execute various processes, such as the data acquisition unit 12, the prediction target limitation unit 14, the stability prediction unit including the production stability prediction model 16, the processing result output unit 18, the culture data analysis unit 20, the limiting condition determination unit 22, the learning unit including the prediction model 410 in the machine learning device 300, the loss calculation unit, the parameter update amount calculation unit, and the parameter update unit, in the information processing device 10, is, for example, various processors as shown below.

[0136] The various types of processors include CPUs, which are general-purpose processors that execute programs and function as various processing units, GPUs, programmable logic devices (PLDs) such as FPGAs (Field Programmable Gate Arrays) that are processors whose circuit configuration can be changed after manufacture, and dedicated electrical circuits such as ASICs (Application Specific Integrated Circuits) that are processors with a circuit configuration designed specifically for executing specific processing.

[0137] A single processing unit may be configured with one of these various processors, or may be configured with two or more processors of the same or different types. For example, a single processing unit may be configured with multiple FPGAs, a combination of a CPU and an FPGA, or a combination of a CPU and a GPU. Multiple processing units may also be configured with a single processor. Examples of multiple processing units configured with a single processor include, first, a configuration in which a single processor is configured with a combination of one or more CPUs and software, as typified by client or server computers, and this processor functions as multiple processing units. Second, a configuration in which a processor is used to realize the functions of an entire system including multiple processing units on a single IC (Integrated Circuit) chip, as typified by a system-on-chip (SoC). In this way, the various processing units are configured with one or more of the above-mentioned various processors as a hardware structure.

[0138] Furthermore, the hardware structure of these various processors is, more specifically, an electric circuit made up of a combination of circuit elements such as semiconductor elements.

[0139] Advantages of the Embodiments According to the method for predicting the production stability of a producing clone according to the above-described embodiment and the information processing device 10 for executing the method, the following effects can be obtained.

[0140] [1] Since the clones to be predicted are appropriately limited based on the indicators of the current culture data (at the time of standard testing), the production stability of the clones to be predicted can be predicted with high accuracy.

[0141] [2] Genetic analysis (RNA-Seq analysis) only needs to be performed on the clones to be predicted, which reduces costs compared to performing genetic analysis on all clones.

[0142] [3] By applying the method according to this embodiment instead of conventional stability tests, the development process of producing cells can be shortened and costs reduced.

[0143] Others The present disclosure is not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the technical idea of ​​the present disclosure.

[0144] 10 Information processing device 12 Data acquisition unit 14 Prediction target limitation unit 16 Production stability prediction model 18 Processing result output unit 20 Cultivation data analysis unit 22 Limitation condition determination unit 102 Processor 104 Computer readable medium 106 Communication interface 108 Input / output interface 110 Bus 112 Memory 114 Storage 120 Data storage area 140 Prediction target limitation program 142 Cultivation data analysis program 144 Limitation condition determination program 152 Input device 154 Display device 180 Processing result output program 190 Display control program 300 Machine learning device 302 Processor 304 Computer readable medium 306 Communication interface 308 Input / output interface 310 Bus 312 Memory 314 Storage 320 Prediction target limitation program 322 Prediction target data storage area 330 Learning processing program 340 Display control program 352 Input device 354 Display device 400 Data acquisition program 410 Prediction model 430 Loss calculation program 440 Optimizer 550 Data storage unit DSA, DSB Data set DSc, DSd, DSe Data set F5A Left diagram F5B Right diagram G1 Graph G2 Graph GEP Gene expression pattern MLM Machine learning model MLMc, MLMe Machine learning model RF1 to RF7 Rectangular frames S102 to S108 Steps of machine learning method S202 to S208 Steps of information processing method for predicting production stability

Claims

1. A method for predicting the production stability of clones that produce a useful substance, comprising: one or more processors performing the following steps: acquiring culture data for one or more of the clones; analyzing the culture data to narrow down the clones to be predicted; and predicting the production stability of the useful substance by the clones to be predicted using data measured on the clones to be predicted.

2. The method according to claim 1, wherein the production stability is defined by whether or not there is a change in the amount of the useful substance produced between the start of culture and after a predetermined period of culture.

3. The method of claim 1, wherein the one or more processors: set an index obtained from the culture data and a threshold value for the index; and limit the prediction target based on the value of the index and the threshold value.

4. The method according to claim 3, wherein the threshold is adjusted so that the prediction accuracy of the production stability is higher than when the prediction target is not limited.

5. The method according to claim 3, wherein the threshold is defined using a ranking for the value of the index.

6. The method according to claim 3, wherein the prediction target is a top group of values ​​of the index.

7. The method according to any one of claims 3 to 6, wherein the indicator is the amount of production of the useful substance.

8. The method according to any one of claims 3 to 6, wherein the indicator is integral viable cell density.

9. The method according to any one of claims 3 to 6, wherein the indicator is lactate concentration.

10. The method of any one of claims 1 to 6, wherein the data used to predict production stability includes one or more gene expression levels.

11. The method according to any one of claims 1 to 6, wherein the one or more processors receive the data to be predicted and predict the production stability using a model that performs two-class classification, stable or unstable.

12. The method according to claim 11, wherein the model is a model trained by machine learning using a plurality of training data sets in which the data on training clones with similar restrictions to the clone to be predicted is associated with correct stability labels.

13. The method according to claim 12, wherein the plurality of training data includes training data for a plurality of types of clones that produce different useful substances, and the one or more processors predict production stability for clones that produce useful substances other than the useful substances used to train the model.

14. The method according to any one of claims 1 to 6, wherein the useful substance is any one of proteins, peptides, and viruses that are pharmaceutical raw materials.

15. The method according to any one of claims 1 to 6, wherein the useful substance is an antibody or an antibody-like protein.

16. The method of any one of claims 1 to 6, wherein the clone is a cell derived from a vertebrate.

17. The method of any one of claims 1 to 6, wherein the clone is a mammalian cell.

18. The method of any one of claims 1 to 6, wherein the clone is a CHO cell or a HEK cell.

19. An information processing device comprising one or more processors and one or more storage devices storing instructions to be executed by the one or more processors, wherein the one or more processors acquire culture data of one or more types of clones that produce a useful substance, analyze the culture data to narrow down clones to be predicted, and predict the stability of production of the useful substance by the clones to be predicted using data measured on the clones to be predicted.

20. A program that enables a computer to perform the following functions: acquire culture data of one or more clones that produce a useful substance; analyze the culture data to narrow down the clones to be predicted; and predict the stability of production of the useful substance by the clones to be predicted using data measured on the clones to be predicted.

21. A non-transitory computer-readable recording medium on which the program according to claim 20 is recorded.

22. A method for generating a predictive model that enables a computer to predict the production stability of clones that produce useful substances, the method comprising: a system including one or more processors; acquiring culture data for one or more types of clones; analyzing the culture data to narrow down clones to be predicted; and performing machine learning using a plurality of training data in which measured data for the clones that correspond to the prediction targets is associated with correct stability labels, and training the predictive model so that the output of the predictive model in response to the input data approaches the correct stability labels.