Information processing device, operation method of information processing device, operation program of information processing device, and state prediction model

JPWO2024070543A5Pending Publication Date: 2025-07-03
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024549952
Authority / Receiving Office
JP · JP
Patent Type
Applications
Priority Date
2023-09-06
Filing Date
2023-09-06
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing methods for predicting the concentration of target protein aggregates in biopharmaceutical manufacturing processes lack accuracy and practicality, as they fail to select a reasonable wavenumber band from Raman spectrum data that truly contributes to aggregate prediction.

Method used

An information processing device and method that selects a specific wavenumber band from Raman spectrum data by comparing intensity values of target protein and target component measurements, using a machine learning model to predict the state of target components in biopharmaceutical suspensions with higher accuracy.

Benefits of technology

The solution provides a state prediction model that accurately predicts the concentration of target protein aggregates in biopharmaceutical suspensions, improving prediction accuracy and reliability compared to conventional methods.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

Provided is an information processing device comprising a processor, wherein the processor acquires, in a preparation process for generating a state prediction model that predicts a state of a designated component in a suspension produced in a production process of a biopharmaceutical which uses a target protein as an active ingredient, first spectrum measurement data obtained by measurement of a spectrum of electromagnetic waves that are emitted from the target protein and a second spectrum measurement data obtained by measurement of a spectrum of electromagnetic waves that are emitted from the designated component, and selects a specific wavelength band or a specific frequency band specific to the designated component by comparing the intensity value of the first spectrum measurement data and the intensity value of the second spectrum measurement data.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, operating method of information processing device, operating program of information processing device, and state prediction model

[0001] The technology of the present disclosure relates to an information processing device, an operating method for an information processing device, an operating program for an information processing device, and a state prediction model.

[0002] There are known manufacturing processes for biopharmaceuticals that use target proteins such as antibodies as active ingredients. In these manufacturing processes, a suspension is often produced in which various components, including the target protein, are dispersed in a liquid. Monitoring the state of the target components in this suspension is important for determining the success or failure of the manufacturing process.

[0003] Japanese Patent Application Laid-Open No. 2016-128822 describes a technique for predicting the concentration of aggregates of a target protein as a state of a target component. Specifically, Japanese Patent Application Laid-Open No. 2016-128822 predicts the concentration of aggregates from spectral measurement data obtained by measuring the Raman spectrum of a suspension using a linear model such as a partial least squares (PLS) regression model.

[0004] The technique described in JP 2016-128822 A did not provide a high level of accuracy in predicting the concentration of aggregates, and was therefore of limited practical use. This is thought to be because the technique did not select a wavenumber band from the Raman spectrum measurement data that would contribute to predicting the concentration of aggregates.

[0005] One possible method for selecting wavenumber bands that are likely to contribute to predicting the concentration of aggregates is sparse modeling. However, the wavenumber bands selected by sparse modeling are highly dependent on the Raman spectrum measurement data prepared for the selection. For this reason, it cannot be said with certainty that the wavenumber bands selected by sparse modeling are rational ones that are likely to truly contribute to predicting the concentration of aggregates.

[0006] One embodiment of the technology of the present disclosure provides an information processing device, an operating method of the information processing device, and an operating program of the information processing device that are capable of selecting a wavenumber band or wavelength band of reasonable spectral measurement data that is thought to contribute to predicting the state of a target component in a suspension produced in a manufacturing process of a biopharmaceutical.

[0007] Furthermore, one embodiment of the technology of the present disclosure provides a state prediction model that can predict the state of a target component in a suspension produced in a biopharmaceutical manufacturing process with higher accuracy than conventional models.

[0008] The information processing device disclosed herein includes a processor, and as a preparatory process for generating a state prediction model for predicting the state of a target component in a suspension produced in a manufacturing process of a biopharmaceutical containing a target protein as an active ingredient, the processor acquires first spectral measurement data that measures the spectrum of electromagnetic waves emitted from the target protein and second spectral measurement data that measures the spectrum of electromagnetic waves emitted from the target component, and selects a specific wavenumber band or wavelength band that is specific to the target component by comparing the intensity values ​​of the first spectral measurement data with the intensity values ​​of the second spectral measurement data.

[0009] The state prediction model is preferably generated using a data set consisting of intensity values ​​in a specific wavenumber band or specific wavelength band and correct data on the state of the target component.

[0010] The state of the target component is the concentration of the target component in the suspension, and the concentrations of the target protein and the target component in the suspension from which the data set was derived are preferably both in the range of 0.001 mg / mL to 20 mg / mL.

[0011] The suspension to be used for selection of the specific wavenumber band or specific wavelength band is preferably subjected to a pretreatment that promotes the production of the target component.

[0012] It is preferable that the state prediction model outputs a prediction result of the state of the target component according to the intensity value of a specific wavenumber band or a specific wavelength band of the third spectrum measurement data, which measures the spectrum of electromagnetic waves emitted from a suspension whose state of the target component is unknown.

[0013] The third spectral measurement data is preferably data measured during the manufacturing process.

[0014] The third spectral measurement data is preferably data measured after a virus inactivation treatment or a cation chromatography treatment.

[0015] The first and second spectral measurement data are preferably data measured from a first solution containing a target protein and a second solution containing a target component, which are separated from a suspension using a high-performance liquid chromatography device.

[0016] The target component is preferably an aggregate of a target protein.

[0017] Preferably, the state prediction model is a machine learning model.

[0018] Preferably, the target protein is an antibody.

[0019] Preferably the spectrum is a Raman spectrum.

[0020] The characteristic waveband is 1220 cm -1 ~1260cm -1 range, or 1650 cm -1 ~1690cm -1 It is preferable that the temperature is in at least one of the ranges.

[0021] A method of operating an information processing device disclosed herein includes, as a preparatory process for generating a state prediction model for predicting the state of a target component in a suspension produced in a manufacturing process of a biopharmaceutical containing the target protein as an active ingredient, acquiring first spectral measurement data that measures the spectrum of electromagnetic waves emitted from the target protein and second spectral measurement data that measures the spectrum of electromagnetic waves emitted from the target component, and selecting a specific wavenumber band or wavelength band that is specific to the target component by comparing the intensity values ​​of the first spectral measurement data with the intensity values ​​of the second spectral measurement data.

[0022] The operating program of the information processing device disclosed herein causes a computer to execute processing, as preparatory processing for generating a state prediction model for predicting the state of a target component in a suspension produced in a manufacturing process of a biopharmaceutical containing the target protein as an active ingredient, including acquiring first spectral measurement data that measures the spectrum of electromagnetic waves emitted from the target protein and second spectral measurement data that measures the spectrum of electromagnetic waves emitted from the target component, and selecting a specific wavenumber band or wavelength band that is specific to the target component by comparing the intensity values ​​of the first spectral measurement data with the intensity values ​​of the second spectral measurement data.

[0023] The state prediction model disclosed herein causes a computer to perform a function of outputting a prediction result for the state of a target component based on the intensity value of a specific wavenumber band or wavelength band that is specific to the target component in the suspension, among the intensity values ​​of each wavenumber or wavelength of spectral measurement data obtained by measuring the spectrum of electromagnetic waves emitted from a suspension produced in a manufacturing process for a biopharmaceutical containing a target protein as an active ingredient.

[0024] According to the technology of the present disclosure, it is possible to provide an information processing device, an operating method for an information processing device, and an operating program for an information processing device that are capable of selecting intensity values ​​of wavenumber bands or wavelength bands of reasonable spectral measurement data that are thought to contribute to predicting the state of target components in suspensions produced in the manufacturing process of biopharmaceuticals.

[0025] Furthermore, the technology disclosed herein can provide a state prediction model that can predict the state of a target component in a suspension produced in a biopharmaceutical manufacturing process with higher accuracy than conventional models.

[0026] 1 is a diagram showing an overview of a biopharmaceutical manufacturing process. FIG. 2 is a diagram showing an information processing system. FIG. 3 is a block diagram of computers constituting a selection device, a learning device, and an operation device. FIG. 4 is a diagram showing pretreatment performed on a second purified liquid, a high-performance liquid chromatography device, and data input to the selection device. FIG. 5 is a diagram showing spectral measurement data and Raman spectra. FIG. 6 is a block diagram of a CPU of a computer constituting the selection device. FIG. 7 is a diagram showing a process of identifying first spectral measurement data and second spectral measurement data from a spectral measurement data group based on chromatogram data. FIG. 8 is a diagram showing first spectral measurement data. FIG. 9 is a diagram showing second spectral measurement data. FIG. 10 is a diagram showing a process of calculating difference data between the first spectral measurement data and the second spectral measurement data. FIG. 11 is a diagram showing a process of comparing difference data with a threshold and selecting a characteristic wavenumber band of an aggregate. FIG. 12 is a diagram showing, on a Raman spectrum, a process of comparing difference data with a threshold and selecting a characteristic wavenumber band of an aggregate. FIG. 13 is a block diagram of a CPU of a computer constituting the learning device. FIG. 14 is a diagram showing a neural network constituting a concentration prediction model. FIG. 15 is a diagram showing the formation of a dataset group. FIG. 16 is a diagram showing a process in the learning phase of a concentration prediction model. FIG. 17 is a diagram showing a process in the verification phase of a concentration prediction model. FIG. 18 is a block diagram of a CPU of a computer constituting the operation device. FIG. 1 is a diagram showing the formation of third spectral measurement data. FIG. 2 is a diagram showing a process of generating input data from third spectral measurement data by referring to specific wavenumber band data, inputting the input data to a concentration prediction model, and outputting a concentration prediction result from the concentration prediction model. FIG. 3 is a diagram showing a Raman spectral analysis screen. FIG. 4 is a diagram showing a Raman spectral analysis screen on which a concentration prediction result is displayed. FIG. 5 is a flowchart showing the processing procedure of a selection device. FIG. 6 is a flowchart showing the processing procedure of a learning device. FIG. 7 is a flowchart showing the processing procedure of an operation device. FIG. 8 is a diagram showing another example of the formation of third spectral measurement data. FIG. 9 is a table showing an overview of examples and comparative examples.

[0027] [First Embodiment] As an example, as shown in Figure 1 , a biopharmaceutical manufacturing process 2 is broadly divided into a first process 10, a second process 11, and a third process 12. The first process 10 is a process of incorporating an antibody gene 14 into cells 13 such as Chinese hamster ovary cells (CHO cells) to establish antibody-producing cells 15. The second process is a process of culturing the antibody-producing cells 15 in a culture tank 16.

[0028] The third process 12 is a process for purifying a drug substance 18 of a biopharmaceutical from a culture supernatant 17. The culture supernatant 17 is a solution obtained by removing cells from the culture medium in the culture tank 16 after the second process 11 has been completed. Immunoglobulins, i.e., antibodies 19, produced by the antibody-producing cells 15 are dispersed in the culture supernatant 17. The antibodies 19 are, for example, monoclonal antibodies, and serve as the active ingredient of the biopharmaceutical. Aggregates 20 of the antibodies 19 are also dispersed in the culture supernatant 17. The antibodies 19 are an example of a "target protein" according to the technology of the present disclosure. The aggregates 20 are an example of a "target component" according to the technology of the present disclosure.

[0029] Aggregates 20 are aggregates of antibody 19 itself and / or multiple denatured products of antibody 19 that share 70% or more of the amino acid sequence with antibody 19. Therefore, aggregates 20 have a larger mass than antibody 19. Aggregates 20 also have a larger molecular weight than antibody 19. Specifically, aggregates 20 are substances having a molecular weight 1.2 times or more that of antibody 19. Furthermore, aggregates 20 are preferably substances having a molecular weight 1.5 times or more, more preferably 1.8 times or more, and particularly preferably 1.9 times or more that of antibody 19. Although not shown in the figure, in addition to antibodies 19 and aggregates 20, cell-derived proteins, cell-derived DNA (deoxyribonucleic acid), viruses, and the like are also dispersed in culture supernatant 17.

[0030] In the third process 12, the culture supernatant 17 is continuously or intermittently purified using an immunoaffinity chromatography device 25, a cation chromatography device 26, an anion chromatography device 27, and the like. The culture supernatant 17 is introduced into the immunoaffinity chromatography device 25. The immunoaffinity chromatography device 25 extracts the antibody 19 from the culture supernatant 17 using a column on which a ligand such as protein A, which has affinity for the antibody 19, is immobilized on a carrier, thereby producing a first purified solution 28. The first purified solution 28 is subjected to a virus inactivation treatment 29. The first purified solution 28 is an example of a "suspension" according to the technology of the present disclosure.

[0031] A first purified solution 28 that has been subjected to a virus inactivation treatment 29 is introduced into the cation chromatography device 26. The cation chromatography device 26 extracts antibodies 19 from the first purified solution 28 using a column with a cation exchanger as the stationary phase, thereby producing a second purified solution 30. The second purified solution 30 is an example of a "suspension" according to the technology of the present disclosure.

[0032] A second purified solution 30 is introduced into the anion chromatography device 27. The anion chromatography device 27 extracts the antibody 19 from the second purified solution 30 using a column having an anion exchanger as a stationary phase, thereby producing a third purified solution 31.

[0033] The third purified solution 31 is passed through a filter 32 to remove viruses. The third purified solution 31 is then subjected to concentration and filtration processes using ultrafiltration (UF) and diafiltration (DF) using a filter 33. This results in the production of a drug substance 18 for a biopharmaceutical. By sequentially performing these component separation processes using multiple types of chromatography devices 25-27, contaminants such as aggregates 20 and viruses are gradually removed from the culture supernatant 17, thereby gradually increasing the purity of the antibody 19. A single-pass tangential flow filtration (SPTFF) filter may be provided upstream of the immunoaffinity chromatography device 25.

[0034] 2, an information processing system 40 includes a selection device 41A, a learning device 41B, and an operation device 41C. These devices are interconnected via a network 42 for mutual communication. The network 42 is, for example, a wide area network (WAN) such as the Internet or a public communication network. The selection device 41A, the learning device 41B, and the operation device 41C are, for example, desktop personal computers, notebook personal computers, or tablet terminals.

[0035] The selection device 41A is responsible for selecting a specific wavenumber band specific to the aggregate 20 from among the wavenumbers in the Raman spectrum. The learning device 41B is responsible for training a concentration prediction model 96 (see FIG. 13) that predicts the concentration of the aggregate 20. The operation device 41C is responsible for predicting the concentration of the aggregate 20 using a trained concentration prediction model 96LD (see FIG. 13). The concentration is an example of a "state" according to the technology of the present disclosure. The "state" is an index that represents the physicochemical characteristics of the target component. The selection device 41A, the learning device 41B, and the operation device 41C are also examples of an "information processing device" according to the technology of the present disclosure. In this way, the "information processing device" according to the technology of the present disclosure may be realized across multiple devices.

[0036] 3, the computers constituting the selection device 41A, learning device 41B, and operation device 41C basically have the same configuration, and include storage 45, memory 46, a CPU (Central Processing Unit) 47, a communication unit 48, a display 49, and an input device 50. These are interconnected via a bus line 51.

[0037] The storage 45 is a hard disk drive built into the computers that make up the selection device 41A, learning device 41B, and operation device 41C, or connected via a cable or network. Alternatively, the storage 45 is a disk array with multiple hard disk drives connected in series. The storage 45 stores control programs such as an operating system, various application programs, and various data associated with these programs. Note that a solid state drive may be used instead of a hard disk drive.

[0038] The memory 46 is a work memory for the CPU 47 to execute processing. The CPU 47 loads programs stored in the storage 45 into the memory 46 and executes processing in accordance with the programs. In this way, the CPU 47 comprehensively controls each part of the computer. The CPU 47 is an example of a "processor" according to the technology of the present disclosure. The memory 46 may be built into the CPU 47.

[0039] The communication unit 48 is a network interface that controls the transmission of various information via the network 42, etc. The display 49 displays various screens. The various screens are equipped with operation functions using a GUI (Graphical User Interface). The computers that make up the selection device 41A, learning device 41B, and operation device 41C accept input of operation instructions from an input device 50 via the various screens. The input device 50 is a keyboard, mouse, touch panel, microphone for voice input, etc.

[0040] In the following explanation, the parts of the computer that make up the selection device 41A (storage 45 and CPU 47) are distinguished by adding the suffix "A" to the symbols, the parts of the computer that make up the learning device 41B (storage 45 and CPU 47) are distinguished by adding the suffix "B" to the symbols, and the parts of the computer that make up the operation device 41C (storage 45, CPU 47, and display 49) are distinguished by adding the suffix "C" to the symbols.

[0041] As an example, as shown in FIG. 4 , the specific wavenumber band of aggregates 20 is selected using a first purified solution 28 output from an immunoaffinity chromatography device 25 and subjected to immunoaffinity chromatography. The first purified solution 28 is subjected to a pretreatment 55 for promoting the formation of aggregates 20. Specifically, as shown in Table 56, the pretreatment 55 is a process in which the hydrogen ion exponent (referred to as pH (potential hydrogen)) of the first purified solution 28 is set to 3.0 and the first purified solution 28 is allowed to stand for one week in an environment at a temperature of 24°C. After the pretreatment 55, the first purified solution 28 is introduced into a high performance liquid chromatography device (hereinafter referred to as an HPLC (High Performance Liquid Chromatography) device) 57. The formation of aggregates 20 in the first purified solution 28 may be further promoted by, for example, increasing the temperature to 30°C or higher.

[0042] The HPLC device 57 includes a reservoir 58, a pump 59, an autosampler 60, a column 61, and an ultraviolet detector (hereinafter referred to as a UV (ultraviolet) detector) 62. The reservoir 58 stores a liquid 63 that is a mobile phase. The liquid 63 is, for example, phosphate-buffered saline (PBS). The pump 59 delivers the liquid 63 from the reservoir 58 toward the column 61 at a preset flow rate (for example, 1 mL / min).

[0043] The autosampler 60 is connected between the pump 59 and the column 61. The autosampler 60 automatically injects a preset amount (e.g., several μL to several tens of μL) of the first purified liquid 28 after the pretreatment 55 into the liquid 63 flowing toward the column 61. Note that instead of the autosampler 60, an injector that manually injects the first purified liquid 28 may be used.

[0044] The column 61 contains a packing material (e.g., silica gel, synthetic resin, etc.) as a stationary phase for separating the antibodies 19 and aggregates 20 in the first purified solution 28, and is capable of performing gel filtration chromatography or size exclusion chromatography. The antibodies 19 and aggregates 20 separated by the column 61 are sequentially eluted from the column 61 together with a liquid 63 and reach a UV detector 62. The UV detector 62 irradiates the liquid 63 from the column 61 with detection light and measures the absorbance (amount of light absorbed) of the substances in the liquid 63. The detection light is ultraviolet light and / or visible light of a wavelength matched to the antibodies 19 and aggregates 20 (light with a wavelength of 190 nm to 800 nm, more specifically light with a wavelength of 280 nm).

[0045] The UV detector 62 is connected to the selection device 41A via a computer network such as a local area network (LAN) so as to be able to communicate with the selection device 41A. The UV detector 62 transmits chromatogram data 64, which is the absorbance measurement result, to the selection device 41A.

[0046] A flow cell 65 is connected downstream of the UV detector 62. The liquid 63 that has passed through the UV detector 62 flows through the flow cell 65. A recovery tank 66 for the liquid 63 is connected downstream of the flow cell 65.

[0047] A probe 68 of a Raman spectrometer 67 is connected to the flow cell 65. The Raman spectrometer 67 is an instrument that evaluates substances using the characteristics of Raman scattered light. When excitation light is irradiated onto a substance, the excitation light interacts with the substance, generating Raman scattered light with a wavelength different from that of the excitation light. The wavelength difference between the excitation light and the Raman scattered light corresponds to the energy of the molecular vibrations of the substance. Therefore, Raman scattered light with different wavenumbers can be obtained between substances with different molecular structures. Of the Stokes line and the anti-Stokes line, it is preferable to use the Stokes line for the Raman scattered light. Raman scattered light is an example of an "electromagnetic wave" according to the technology of the present disclosure. Furthermore, the spectrum of Raman scattered light, i.e., the Raman spectrum, is an example of a "spectrum" according to the technology of the present disclosure.

[0048] The Raman spectrometer 67 is composed of a probe 68 and an analyzer 69. The probe 68 emits excitation light from an outlet at its tip toward the liquid 63 flowing through the measurement section 70 of the flow cell 65. Raman scattered light generated by the interaction between the excitation light and the substances in the liquid 63 is received by a light receiving section located at its tip. The probe 68 outputs the received Raman scattered light to the analyzer 69. In this example, laser light was used as the excitation light, with an output of 200 mW, an excitation wavelength of 785 nm, and an irradiation time of 1 second.

[0049] The analyzer 69 resolves the Raman scattered light into wavenumbers and derives the intensity value of the Raman scattered light for each wavenumber, thereby generating spectral measurement data 71. Here, the probe 68 emits excitation light and receives Raman scattered light at preset intervals from time T0, when the autosampler 60 starts to inject the first purified solution 28, until time TN, which is sufficient for the UV detector 62 to measure the absorbance of the antibody 19 and the aggregate 20. The analyzer 69 generates the spectral measurement data 71 each time. Therefore, the spectral measurement data 71 is generated in the form of a plurality of pieces of spectral measurement data 71, including spectral measurement data 71T0 at time T0, spectral measurement data 71T1 at time T1, ..., and spectral measurement data 71TN at time TN.

[0050] The analyzer 69 is connected to the selection device 41A via a computer network such as a LAN so as to be able to communicate with each other, similar to the HPLC device 57. The analyzer 69 transmits a spectrum measurement data group 71G, which is a collection of multiple spectrum measurement data 71, to the selection device 41A.

[0051] As an example, as shown in Fig. 5, the spectrum measurement data 71 is data in which the intensity value of the Raman scattered light for each wave number is registered. In Fig. 5, the spectrum measurement data 71 is -1 ~1800cm -1 The scattered light intensity values ​​in the range from 1 cm -1 5 is a graph in which the intensity values ​​of the spectrum measurement data 71 are plotted for each wave number and connected by a line, that is, the graph represents the Raman spectrum.

[0052] 6, an operating program 75A is stored in the storage 45A of the selection device 41A. The operating program 75A is an application program for causing a computer to function as the selection device 41A. In other words, the operating program 75A is an example of an "operating program for an information processing device" according to the technology of the present disclosure.

[0053] When the operating program 75A is started, the CPU 47A of the computer constituting the selection device 41A works in cooperation with the memory 46, etc. to function as an acquisition unit 80, a read / write control unit (hereinafter referred to as the RW (Read / Write) control unit) 81, and a selection unit 82.

[0054] The acquiring unit 80 acquires the chromatogram data 64 from the HPLC device 57 and the spectrum measurement data group 71G from the Raman spectrometer 67. The acquiring unit 80 outputs the chromatogram data 64 and the spectrum measurement data group 71G to the RW control unit 81.

[0055] The RW control unit 81 controls the storage of various data in the storage 45A and the reading of various data stored in the storage 45A. The RW control unit 81 stores the chromatogram data 64 and the spectrum measurement data group 71G from the acquisition unit 80 in the storage 45A. The RW control unit 81 also reads the chromatogram data 64 and the spectrum measurement data group 71G from the storage 45A and outputs the read chromatogram data 64 and the spectrum measurement data group 71G to the selection unit 82.

[0056] The selection unit 82 selects a characteristic wavenumber band of the aggregate 20 based on the chromatogram data 64 and the spectrum measurement data group 71G. The selection unit 82 generates characteristic wavenumber band data 85 as a result of the selection of the characteristic wavenumber band. The selection unit 82 outputs the characteristic wavenumber band data 85 to the RW control unit 81. The RW control unit 81 stores the characteristic wavenumber band data 85 in the storage 45A.

[0057] 7 , the selection unit 82 identifies first and second spectral measurement data 711 and 712 from among the plurality of spectral measurement data 71 in the spectral measurement data group 71G based on the chromatogram data 64. The first spectral measurement data 711 is data obtained by measuring the Raman spectrum emitted from the antibody 19. The second spectral measurement data 712 is data obtained by measuring the Raman spectrum emitted from the aggregate 20.

[0058] The selection unit 82 derives, from the chromatogram data 64, the time Tan (retention time of antibody 19) at which an absorbance peak indicating antibody 19 appeared, and the time Tag (retention time of aggregate 20) at which an absorbance peak indicating aggregate 20 appeared. The selection unit 82 identifies, as first spectrum measurement data 711, spectrum measurement data 71Tan+α obtained by measuring the Raman spectrum of liquid 63 flowing through measurement unit 70 of flow cell 65 at time Tan. The selection unit 82 also identifies, as second spectrum measurement data 712, spectrum measurement data 71Tag+α obtained by measuring the Raman spectrum of liquid 63 flowing through measurement unit 70 of flow cell 65 at time Tag. Here, the liquid 63 flowing through measurement unit 70 of flow cell 65 at time Tan is an example of a "first solution" according to the technology of the present disclosure. The liquid 63 flowing through the measurement unit 70 of the flow cell 65 at time Tag is an example of the "second solution" according to the technology of the present disclosure. The "+α" in times Tan+α and Tag+α represents the time lag between when the absorbance is measured by the UV detector 62 and when the Raman spectrum is measured by the Raman spectrometer 67 in the measurement unit 70 of the flow cell 65.

[0059] The method for producing the liquid 63 containing the antibody 19 and the liquid 63 containing the aggregate 20 is not limited to the method using the HPLC device 57. For example, the liquid 63 containing the antibody 19 and the liquid 63 containing the aggregate 20 may be separated from the first purified solution 28 using a centrifugal ultrafiltration filter.

[0060] In this way, the spectrum measurement data group 71G includes the first spectrum measurement data 711 and the second spectrum measurement data 712. Therefore, by acquiring the spectrum measurement data group 71G, the acquiring unit 80 acquires the first spectrum measurement data 711 and the second spectrum measurement data 712.

[0061] An example of the first spectrum measurement data 711 is shown in Fig. 8, and an example of the second spectrum measurement data 712 is shown in Fig. 9. As can be seen by comparing Fig. 8 and Fig. 9, the first spectrum measurement data 711 and the second spectrum measurement data 712 are generally the same, but because the former is based on antibody 19 and the latter is based on aggregate 20, the data are slightly different in places.

[0062] 10 , the selection unit 82 calculates difference data 90 between the intensity values ​​at each wavenumber of the first spectrum measurement data 711 and the second spectrum measurement data 712. The difference data 90 is data in which the difference obtained by subtracting the intensity value of the second spectrum measurement data 712 from the intensity value of the first spectrum measurement data 711 is registered for each wavenumber. Prior to calculating the difference data 90, the selection unit 82 normalizes the first spectrum measurement data 711 and the second spectrum measurement data 712 by setting the maximum intensity value to 1 and the minimum intensity value to 0.

[0063] As an example, as shown in FIG. 11 , the selection unit 82 compares the absolute value of the difference of the difference data 90 with a preset threshold value 91. Then, the selection unit 82 selects a wavenumber band in which the absolute value of the difference is equal to or greater than the threshold value as a characteristic wavenumber band of the aggregate 20. In FIG. 11 , the threshold value is set to 0.05, and the characteristic wavenumber band is set to 1220 cm. -1 ~1260cm -1 , and 1650 cm -1 ~1690cm -1 The example shows the case where the specific wave number band is 700 cm. -1 ~1800cm -1 There is no particular limitation as long as it is in the range of 1220 cm -1 ~1690cm -1 It is preferable that the range is 1220 cm -1~1260cm -1 , and 1650 cm -1 ~1690cm -1 It is more preferable that the characteristic wave number band is in the range of 1220 cm -1 ~1260cm -1 , and 1650 cm -1 ~1690cm -1 It is preferable that the range be two or more, such as: A range in which a phenylalanine band appears, a range in which a tryptophan band appears, or a range in which a tyrosine band appears, or the like, may be selected as the characteristic wavenumber band.

[0064] FIG. 12 shows the process shown in FIG. 11 , in which the difference data 90 is compared with a threshold value 91 to select the characteristic wavenumber band of the aggregate, on the Raman spectra of the first spectrum measurement data 711 and the second spectrum measurement data 712.

[0065] In addition, the ratio between the intensity value of each wavenumber of the first spectrum measurement data 711 and the intensity value of each wavenumber of the second spectrum measurement data 712 may be calculated, and the wavenumber band in which the ratio deviates from 1 by more than a threshold value may be selected as the characteristic wavenumber band of the aggregate 20.

[0066] As an example, as shown in FIG. 13 , an operating program 75B is stored in the storage 45B of the learning device 41B. The operating program 75B is an application program for causing a computer to function as the learning device 41B. In other words, the operating program 75B, like the operating program 75A, is an example of an "operating program of an information processing device" according to the technology of the present disclosure. In addition to the operating program 75B, the storage 45B also stores a data set group 95G and a concentration prediction model 96. The concentration prediction model 96 is an example of a "state prediction model" according to the technology of the present disclosure.

[0067] When the operating program 75B is started, the CPU 47B of the computer constituting the learning device 41B functions as the RW control unit 100 and the learning verification unit 101 in cooperation with the memory 46 and the like.

[0068] The RW control unit 100, like the RW control unit 81 of the selection device 41A, controls the storage of various data in the storage 45B and the reading of various data stored in the storage 45B. The RW control unit 100 reads the data set group 95G and the concentration prediction model 96 from the storage 45B, and outputs the read data set group 95G and the concentration prediction model 96 to the learning verification unit 101.

[0069] The learning verification unit 101 performs learning and verification of the concentration prediction model 96 using the data set group 95G. The learning verification unit 101 outputs a trained concentration prediction model 96LD obtained by the learning and verification to the RW control unit 100. The RW control unit 100 stores the concentration prediction model 96LD in the storage 45B.

[0070] As an example, as shown in FIG. 14 , the concentration prediction model 96 is constructed using a neural network 105. Therefore, the concentration prediction model 96 is also an example of a “machine learning model” according to the technology of the present disclosure. As is well known, the neural network 105 has an input layer 106, an intermediate layer (also called a hidden layer) 107, and an output layer 108. The input layer 106, the intermediate layer 107, and the output layer 108 each have a plurality of nodes ND. Coefficients indicating the strength of connections between the nodes ND in the input layer 106 and the nodes ND in the intermediate layer 107, between the nodes ND in the intermediate layer 107, and between the nodes ND in the intermediate layer 107 and the nodes ND in the output layer 108 are set. An appropriate activation function, such as a linear function or a ReLu (Rectified Linear Unit) function, is set for the nodes ND in the output layer 108.

[0071] Each node ND of the input layer 106 receives as input data 130 (see FIG. 20 ) the intensity value of a specific wavenumber band among the intensity values ​​of each wavenumber of the spectrum measurement data 71. The node ND of the output layer 108 outputs a concentration prediction result 115 (see FIG. 18 ), which is a result of predicting the concentration of the aggregate 20.

[0072] 15 , the dataset group 95G includes a plurality of datasets 95. Each dataset 95 is composed of a training or verification intensity value 110 and a correct concentration 111. The training or verification intensity value 110 is obtained by extracting the intensity value of a specific wavenumber band selected by the selection device 41A from the intensity values ​​of each wavenumber in the spectrum measurement data 71LV used to generate the dataset 95. The spectrum measurement data 71LV is data obtained by measuring the Raman spectrum of the second purified solution 30 after the cation chromatography process, output from the cation chromatography device 26, using the flow cell 65 and the Raman spectrometer 67.

[0073] Multiple pieces of spectral measurement data 71LV are measured intermittently from the start to the end of the cation chromatography process by the cation chromatography device 26. Furthermore, multiple pieces of spectral measurement data 71LV are measured by randomly varying the culture conditions of the antibody-producing cells 15, the gradient width, linear flow velocity, and load amount of the cation chromatography device 26, etc. This makes it possible to obtain multiple pieces of spectral measurement data 71LV for the second purified solution 30 having different concentration ratios of antibody 19 and aggregate 20, and thereby obtain multiple learning or verification intensity values ​​110. Note that instead of the illustrated method of measuring the spectral measurement data 71LV in the flow path using a flow cell 65, a method may be employed in which a fraction of the second purified solution 30 flowing out of the flow path outlet is fractionated using a fraction collector and the spectral measurement data 71LV for the fractionated second purified solution 30 is measured.

[0074] The concentrations of antibody 19 and aggregate 20 in second purified solution 30 for measuring spectrum measurement data 71LV are both in the range of 0.001 mg / mL to 20 mg / mL. The concentrations of antibody 19 and aggregate 20 in second purified solution 30 may both be in the range of 0.001 mg / mL to 10,000 mg / mL, preferably in the range of 0.001 mg / mL to 100 mg / mL, and more preferably in the exemplary range of 0.001 mg / mL to 20 mg / mL.

[0075] The correct concentration 111 is a concentration calculated based on the aggregate amount 112 in the second purified solution 30 for which the spectrum measurement data 71LV was measured. The aggregate amount 112 is literally the amount of aggregates 20, and is derived by the mass analysis function of the HPLC device 57. The correct concentration 111 is an example of the "correct data" according to the technology of the present disclosure.

[0076] The learning verification unit 101 performs cross-validation on the concentration prediction model 96 using multiple data sets 95. That is, the learning verification unit 101 designates m of the M data sets 95 as a training data set 95L (see FIG. 16 ), and the remaining M−m data sets as a validation data set 95V (see FIG. 17 ). Then, as shown in FIG. 16 as an example, the learning data set 95L is applied to the concentration prediction model 96 to train the concentration prediction model 96. Also, as shown in FIG. 17 as an example, a validation data set 95V is applied to the concentration prediction model 96 after it has been trained using the training data set 95L, and the prediction accuracy of the concentration of aggregates 20 by the concentration prediction model 96 is verified. The learning verification unit 101 performs such cross-validation a set number of times while changing the configurations of the training data set 95L and the validation data set 95V. Note that m≧M−m, and M−m=1 may also be acceptable.

[0077] 16 , in the learning phase, the learning verification unit 101 inputs a learning or verification intensity value 110 from a learning data set 95L to a concentration prediction model 96, and causes the concentration prediction model 96 to output a learning concentration prediction result 115L. The learning verification unit 101 performs loss calculation for the concentration prediction model 96 using a loss function based on the result of comparison between the correct concentration 111 and the learning concentration prediction result 115L. The learning verification unit 101 updates the coefficients between nodes N and D of the concentration prediction model 96 in accordance with the result of the loss calculation, and updates the concentration prediction model 96 in accordance with the update setting.

[0078] The learning verification unit 101 repeatedly performs the above series of processes, including inputting the learning or verification intensity values ​​110 to the concentration prediction model 96, outputting the learning concentration prediction results 115L from the concentration prediction model 96, calculating the loss, setting the update, and updating the concentration prediction model 96, while changing the training data set 95L. The learning verification unit 101 repeats the above series of processes m times, which is the number of training data sets 95L.

[0079] 17 , in the verification phase, learning verification unit 101 inputs learning or verification intensity values ​​110 from verification data set 95V to concentration prediction model 96, and causes concentration prediction model 96 to output verification concentration prediction result 115V. Learning verification unit 101 verifies the prediction accuracy of the concentration of aggregate 20 by concentration prediction model 96 based on the comparison result between correct concentration 111 and verification concentration prediction result 115V.

[0080] The learning verification unit 101 repeatedly inputs the learning or verification intensity values ​​110 to the concentration prediction model 96, outputs the verification concentration prediction results 115V from the concentration prediction model 96, and verifies the prediction accuracy while changing the verification data set 95V. The learning verification unit 101 repeats the above series of processes Mm times, which is the number of verification data sets 95V.

[0081] The learning verification unit 101 outputs the concentration prediction model 96, for which the above cross-validation has been performed a set number of times, as a concentration prediction model 96LD to the RW control unit 100. The RW control unit 100 stores the concentration prediction model 96LD in the storage 45B.

[0082] 18 , an operating program 75C is stored in the storage 45C of the operating device 41C. The operating program 75C is an application program for causing a computer to function as the operating device 41C. In other words, the operating program 75C, like the operating programs 75A and 75B, is an example of an “operating program for an information processing device” according to the technology of the present disclosure. In addition to the operating program 75C, the storage 45C also stores the specific wavenumber band data 85 from the selection device 41A and the concentration prediction model 96LD from the learning device 41B.

[0083] When the operation program 75C is started, the CPU 47C of the computer constituting the operation device 41C functions as an acquisition unit 120, an RW control unit 121, a prediction unit 122, and a display control unit 123 in cooperation with the memory 46 and the like.

[0084] The acquiring unit 120 acquires third spectrum measurement data 713 from the Raman spectrometer 67. The acquiring unit 120 outputs the third spectrum measurement data 713 to the RW control unit 121.

[0085] The RW control unit 121, like the RW control unit 81 of the selection device 41A and the RW control unit 100 of the learning device 41B, controls the storage of various data in the storage 45C and the reading of various data stored in the storage 45C. The RW control unit 121 stores third spectrum measurement data 713 from the acquisition unit 120 in the storage 45C. The RW control unit 121 also reads the specific wavenumber band data 85, the concentration prediction model 96LD, and the third spectrum measurement data 713 from the storage 45C and outputs the read specific wavenumber band data 85, the concentration prediction model 96LD, and the third spectrum measurement data 713 to the prediction unit 122. The RW control unit 121 also outputs the third spectrum measurement data 713 to the display control unit 123.

[0086] The prediction unit 122 applies the third spectral measurement data 713 to the concentration prediction model 96LD, and causes the concentration prediction model 96LD to output a concentration prediction result 115. The prediction unit 122 outputs the concentration prediction result 115 to the display control unit 123. The concentration prediction result 115 is an example of a “prediction result” according to the technology of the present disclosure.

[0087] The display control unit 123 controls the display of various screens on the display 49 C. For example, the display control unit 123 controls the display of a Raman spectrum analysis screen 135 (see FIG. 21 etc.) on the display 49 C.

[0088] As an example, as shown in FIG. 19 , the third spectral measurement data 713 is data obtained by measuring the Raman spectrum of the second purified solution 30, whose concentration of aggregates 20 is unknown, using a flow cell 65 and a Raman spectrometer 67. The flow cell 65 is installed between the cation chromatography device 26 and the anion chromatography device 27. Therefore, more specifically, the second purified solution 30 is a liquid after the cation chromatography process, which is output from the cation chromatography device 26 during the progress of the manufacturing process 2. That is, the third spectral measurement data 713 is data measured during the progress of the manufacturing process 2. In other words, the third spectral measurement data 713 is data obtained by in-line sensing. Furthermore, the third spectral measurement data 713 is data measured after the cation chromatography process.

[0089] 20 , the prediction unit 122 references the characteristic wavenumber band data 85 and extracts the intensity value of the characteristic wavenumber band from the intensity values ​​of each wavenumber in the third spectrum measurement data 713, thereby generating input data 130. The prediction unit 122 inputs the input data 130 to a concentration prediction model 96LD, and causes the concentration prediction model 96LD to output a concentration prediction result 115. In FIG. 20 , the characteristic wavenumber band is the 1220 cm 2 wavenumber band illustrated in FIG. -1 ~1260cm -1 , and 1650 cm -1 ~1690cm -1 10, a concentration prediction result 115 of 2.485 mg / mL is output.

[0090] In response to an instruction from the user of the operation device 41C, the display control unit 123 displays, on the display 49C, a Raman spectrum analysis screen 135 shown in Fig. 21 as an example. The Raman spectrum analysis screen 135 displays third spectrum measurement data 713.

[0091] An aggregate concentration prediction button 136 is provided at the bottom of the Raman spectrum analysis screen 135. When the aggregate concentration prediction button 136 is pressed, an aggregate concentration prediction instruction is accepted by the CPU 47C of the operation device 41C. Upon receiving the aggregate concentration prediction instruction, the CPU 47C causes the prediction unit 122 to perform the process shown in FIG. 20 and output the concentration prediction result 115 from the concentration prediction model 96LD.

[0092] When the concentration prediction result 115 is input from the prediction unit 122, the display control unit 123 changes the display of the Raman spectrum analysis screen 135 to the screen shown in Fig. 22. In Fig. 22, the Raman spectrum analysis screen 135 displays the concentration prediction result 115 together with the third spectrum measurement data 713.

[0093] Next, the operation of the above configuration will be described with reference to the flowcharts shown in FIGS. 23 to 25 as an example.

[0094] As shown in FIG. 6, the CPU 47A of the selection device 41A functions as an acquisition unit 80, a RW control unit 81, and a selection unit 82 when an operating program 75A is started.

[0095] 23, in the selection device 41A, the acquisition unit 80 acquires chromatogram data 64 from the HPLC device 57 and a spectrum measurement data group 71G from the Raman spectrometer 67, which are measured by the method shown in FIG. 4 (step ST100). The chromatogram data 64 and the spectrum measurement data group 71G are stored in the storage 45A by the RW control unit 81 (step ST110).

[0096] The chromatogram data 64 and the spectrum measurement data group 71G are read from the storage 45A by the RW control unit 81 (step ST120) and output to the selection unit 82. The selection unit 82 first identifies the first spectrum measurement data 711 and the second spectrum measurement data 712 from the spectrum measurement data group 71G based on the chromatogram data 64 (step ST130), as shown in FIG. 7 . Next, as shown in FIG. 10 , differential data 90 between the first spectrum measurement data 711 and the second spectrum measurement data 712 is calculated (step ST140). Finally, as shown in FIG. 11 , the differential data 90 is compared with a threshold 91 to select a characteristic wavenumber band of the aggregate 20 (step ST150). The result of the selection of the characteristic wavenumber band, characteristic wavenumber band data 85, is output from the selection unit 82 to the RW control unit 81, and the RW control unit 81 stores the data in the storage 45A (step ST160).

[0097] As shown in FIG. 13, the CPU 47B of the learning device 41B functions as a RW control unit 100 and a learning verification unit 101 when an operating program 75B is started.

[0098] 15 , a data set group 95G, which is a collection of data sets 95 generated by the method shown in Fig. 15 , and a concentration prediction model 96 are stored in the storage 45B of the learning device 41B. The data set group 95G and the concentration prediction model 96 are read out from the storage 45B by the RW control unit 100 and output to the learning verification unit 101.

[0099] As an example, as shown in FIG. 24 , the learning and verification unit 101 divides the multiple data sets 95 constituting the data set group 95G into m training data sets 95L and M−m verification data sets 95V (step ST200). Then, first, the concentration prediction model 96 is trained using the training data sets 95L. Specifically, as shown in FIG. 16 , the training or verification intensity values ​​110 of the training data sets 95L are input to the concentration prediction model 96, which then outputs a training concentration prediction result 115L (step ST210). Next, the concentration prediction model 96 is updated based on the comparison result between the correct concentration 111 of the training data set 95L and the training concentration prediction result 115L (step ST220). The processes of steps ST210 and ST220 are repeated while the training data set 95L is changed (step ST240) until all of the prepared training data sets 95L have been used (NO in step ST230).

[0100] If all of the prepared training data sets 95L have been used (YES in step ST230), the process proceeds to verifying the prediction accuracy of the concentration prediction model 96 using the verification data set 95V. Specifically, as shown in FIG. 17 , the training or verification intensity values ​​110 of the verification data set 95V are input to the concentration prediction model 96, which then outputs a verification concentration prediction result 115V. Next, the prediction accuracy of the concentration prediction model 96 is verified based on the comparison result between the correct concentration 111 of the verification data set 95V and the verification concentration prediction result 115V (step ST250). Although not shown in the figure, in this verification, as in the case of training, the above series of processes are repeated while the verification data set 95V is changed until all of the prepared verification data set 95V has been used.

[0101] The processes of steps ST200 to ST250 are repeated until the cross-validation is completed a set number of times (NO in step ST260). When the cross-validation is completed a set number of times (YES in step ST260), the concentration prediction model 96 is output from the learning verification unit 101 to the RW control unit 100 as a trained concentration prediction model 96LD. The concentration prediction model 96LD is stored in the storage 45B by the RW control unit 100 (step ST270).

[0102] As shown in FIG. 18, the CPU 47C of the operation device 41C functions as an acquisition unit 120, a RW control unit 121, a prediction unit 122, and a display control unit 123 by starting an operation program 75C.

[0103] The storage 45C of the operation device 41C stores the specific wavenumber band data 85 from the selection device 41A and the concentration prediction model 96LD from the learning device 41B. The specific wavenumber band data 85 and the concentration prediction model 96LD are read from the storage 45C by the RW control unit 121 and output to the prediction unit 122.

[0104] 25, in the operational device 41C, the acquiring unit 120 acquires third spectrum measurement data 713 from the Raman spectrometer 67, which is measured by the method shown in FIG. 19 (step ST300). The third spectrum measurement data 713 is stored in the storage 45C by the RW control unit 121 (step ST310).

[0105] The third spectrum measurement data 713 is read from the storage 45C by the RW control unit 121 (step ST320) and output to the prediction unit 122 and the display control unit 123. Then, as shown in Fig. 21, the display control unit 123 displays the Raman spectrum analysis screen 135 on the display 49C (step ST330).

[0106] The user of the operational device 41C presses the aggregate concentration prediction button 136 to cause the concentration prediction model 96LD to predict the concentration of the aggregate 20 in the second purified solution 30 for which the third spectrum measurement data 713 on the Raman spectrum analysis screen 135 has been measured. This causes the CPU 47C to accept an aggregate concentration prediction instruction (step ST340).

[0107] In response to the aggregate concentration prediction instruction, the prediction unit 122 generates input data 130 from the third spectrum measurement data 713 with reference to the specific wavenumber band data 85, as shown in Fig. 20 (step ST350). The input data 130 is then input to the concentration prediction model 96LD, which then outputs a concentration prediction result 115 (step ST360). The concentration prediction result 115 is output from the prediction unit 122 to the display control unit 123, and as shown in Fig. 22, the display control unit 123 displays the result on the Raman spectrum analysis screen 135 (step ST370).

[0108] The user makes various decisions based on the concentration prediction result 115 on the Raman spectrum analysis screen 135. For example, consider a case where a condition-finding experiment is being conducted using small-scale equipment to determine the culture conditions for antibody-producing cells 15 and / or the purification conditions for the culture supernatant 17. In this case, if the concentration prediction result 115 is worse than the target value, the user decides to stop the current experiment and move on to an experiment using new conditions. Also, consider a case where the condition-finding experiment has been completed and mass production is being conducted using large-scale equipment. In this case, if the concentration prediction result 115 is worse than the target value, the user decides to stop mass production and perform maintenance on the chromatography devices 25 to 27.

[0109] As described above, the CPU 47A of the selection device 41A includes an acquisition unit 80 and a selection unit 82. The acquisition unit 80 and the selection unit 82 perform the following preparatory processing for generating a concentration prediction model 96LD that predicts the concentration of aggregate 20 in second purified solution 30 produced in a biopharmaceutical production process 2 containing antibody 19 as an active ingredient. That is, the acquisition unit 80 acquires first spectrum measurement data 711 obtained by measuring the Raman spectrum emitted from antibody 19 and second spectrum measurement data 712 obtained by measuring the Raman spectrum emitted from aggregate 20. The selection unit 82 selects a characteristic wavenumber band specific to aggregate 20 by comparing the intensity values ​​of the first spectrum measurement data 711 and the second spectrum measurement data 712. This makes it possible to select a reasonable wavenumber band of the spectrum measurement data 71 that is likely to contribute to predicting the concentration of aggregate 20 in second purified solution 30 produced in the biopharmaceutical production process 2.

[0110] 15 to 17 , the concentration prediction model 96LD is generated using a data set 95 configured with learning or validation intensity values ​​110, which are intensity values ​​of a specific wavenumber band, and a correct concentration 111 of the aggregate 20. Therefore, the concentration prediction model 96LD can be a model that outputs a concentration prediction result 115 of the aggregate 20 according to the intensity value of the specific wavenumber band. The concentration prediction model 96LD makes it possible to predict the concentration of the aggregate 20 in the second purified solution 30 produced in the biopharmaceutical production process 2 with higher accuracy than conventional methods.

[0111] Concentration is the most popular index for understanding the physicochemical characteristics of the target component (aggregate 20). Therefore, if the concentration is predicted as the state of the target component, the user can easily understand the physicochemical characteristics of the target component.

[0112] 15, the concentrations of the antibody 19 and the aggregate 20 in the second purified solution 30, which is the basis of the data set 95, are both in the range of 0.001 mg / mL to 20 mg / mL. Therefore, the concentration prediction model 96LD can be a model that can accurately predict relatively low concentrations.

[0113] 4, the first purified solution 28 used for selecting the specific wavenumber band is subjected to a pretreatment 55 that promotes the formation of the aggregate 20. This makes it possible to reliably acquire the second spectrum measurement data 712. Furthermore, since the absorbance peak indicating the aggregate 20 clearly appears in the chromatogram data 64, the second spectrum measurement data 712 can be easily identified.

[0114] 20 , the concentration prediction model 96LD outputs a concentration prediction result 115 of the aggregate 20 in accordance with the intensity value of the specific wavenumber band of the third spectrum measurement data 713 obtained by measuring the Raman spectrum emitted from the second purified solution 30 having an unknown concentration of the aggregate 20. This allows the user to easily know the concentration prediction result 115 of the aggregate 20.

[0115] 19 , the third spectral measurement data 713 is data measured during the progress of the production process 2. This eliminates the need to sample the second purified solution 30 and subject it to a Raman spectrometer 67 that is provided at a location separate from the purification line. Furthermore, the third spectral measurement data 713 can be acquired without interrupting the progress of the production process 2.

[0116] 19 , the third spectrum measurement data 713 is data measured after the cation chromatography process. The second purified solution 30 after the cation chromatography process should have had most of the aggregates 20 removed. Therefore, if the predicted concentration 115 of the aggregates 20 in the second purified solution 30 after the cation chromatography process is high, the user can conclude that the conditions set in the condition finding experiment are inappropriate or that the cation chromatography device 26 is malfunctioning, making it easier for the user to make a decision.

[0117] 7 , the first spectrum measurement data 711 and the second spectrum measurement data 712 are data measured from the liquid 63 containing the antibody 19 and the liquid 63 containing the aggregate 20, which were separated from the second purified solution 30 using the HPLC apparatus 57. Therefore, the first spectrum measurement data 711 is data that clearly represents the characteristics of the antibody 19, and the second spectrum measurement data 712 is data that clearly represents the characteristics of the aggregate 20. Therefore, the characteristic wavenumber band of the aggregate 20 can be accurately selected.

[0118] The target component is an aggregate 20 of antibody 19. The aggregate 20 has adverse effects on biopharmaceuticals, such as causing side effects, and is the cause of a decrease in the efficacy of the biopharmaceutical. Therefore, by setting the target component as aggregate 20 and predicting its state, it is possible to suppress a decrease in the efficacy of the biopharmaceutical.

[0119] 14, the concentration prediction model 96LD is a machine learning model such as a neural network 105. Machine learning models are generally used to predict unknown parameters, and the prediction accuracy can be improved to a certain level through learning. Therefore, the concentration of aggregates 20 can be predicted with higher accuracy than a linear model such as a PLS model.

[0120] Biopharmaceuticals that contain antibody 19 as the target protein are called antibody drugs, and are widely used to treat chronic diseases such as cancer, diabetes, and rheumatoid arthritis, as well as rare diseases such as hemophilia and Crohn's disease. Therefore, using antibody 19 as the target protein can promote the development of antibody drugs that are widely used to treat a variety of diseases.

[0121] A Raman spectrum is likely to reflect information derived from the functional groups of amino acids in a protein, and therefore, by using a Raman spectrum as the spectrum, it is possible to further improve the accuracy of predicting the concentration of aggregate 20, which is a protein.

[0122] As shown in Figures 11 and 12, the characteristic wavenumber band is 1220 cm -1 ~1260cm -1 and 1650 cm-1 ~1690cm -1 It is in the range of 1220 cm -1 ~1260cm -1 The range of 1650 cm is the range in which the band commonly known as amide III, which is attributed to the amide bond of the protein, appears. -1 ~1690cm -1 This range is the range in which the band commonly known as amide I appears. Therefore, a characteristic wavenumber band with high validity can be selected. The characteristic wavenumber band is 1220 cm -1 ~1260cm -1 range, or 1650 cm -1 ~1690cm -1 It is sufficient if the value falls within at least one of the ranges.

[0123] Second Embodiment In the first embodiment, the third spectral measurement data 713 is data measured after the cation chromatography process, but this is not limiting. As an example, as shown in Fig. 26, the third spectral measurement data 713 may be data obtained by measuring the Raman spectrum of the first purified solution 28 after the virus inactivation process 29. In this case, the first purified solution 28 is an example of the "suspension" according to the technology of the present disclosure.

[0124] The first purified solution 28 has a composition closer to that of the culture supernatant 17 than the second purified solution 30. Therefore, if the third spectrum measurement data 713 is data obtained by measuring the Raman spectrum of the first purified solution 28 after the virus inactivation treatment 29 has been performed, when the concentration prediction result 115 is worse than the target value, it can be concluded that the cause lies in the culture conditions of the antibody-producing cells 15, making it easier for the user to make a decision.

[0125] The third spectral measurement data 713 may be data obtained by measuring the Raman spectrum of the third purified solution 31 after the anion chromatography process, output from the anion chromatography device 27. The third spectral measurement data 713 does not have to be data measured during the production process. The third spectral measurement data 713 may be measured by sampling the first purified solution 28 or the second purified solution 30 and subjecting it to a Raman spectrometer 67 provided at a location separate from the purification line.

[0126] Examples and comparative examples of the technology of the present disclosure will be described below.

[0127] In this example, as described in the first embodiment, antibody genes 14 were introduced into cells 13, such as CHO cells, to produce antibody 19, thereby generating a culture supernatant 17 of antibody-producing cells 15. The culture supernatant 17 was then introduced into an immunoaffinity chromatography device 25 for purification, thereby obtaining a first purified solution 28. Next, the first purified solution 28 was subjected to pretreatment 55 under the conditions shown in Table 56 to promote the formation of aggregates 20. The first purified solution 28 was then injected into an HPLC device 57 via an autosampler 60, and chromatogram data 64 was measured using a UV detector 62. The Raman spectrum of the first purified solution 28 was also measured using a flow cell 65 and a Raman spectrometer 67, resulting in the acquisition of a spectral measurement data group 71G.

[0128] The retention time Tan of antibody 19 and the retention time Tag of aggregate 20 were derived from chromatogram data 64, and the first spectrum measurement data 711 and the second spectrum measurement data 712 were thereby identified from spectrum measurement data group 71G. Then, the characteristic wavenumber band of aggregate 20 was selected based on the first spectrum measurement data 711 and the second spectrum measurement data 712.

[0129] Next, a culture supernatant 17 was produced from antibody-producing cells 15 that produced antibody 19 in the same manner as above, and the produced culture supernatant 17 was introduced into an immunoaffinity chromatography device 25 and a cation chromatography device 26 for purification to obtain a second purified solution 30. At this time, the Raman spectrum of the second purified solution 30 was measured using a flow cell 65 and a Raman spectrometer 67 to obtain spectral measurement data 71LV, and the aggregate amount 112 was measured using an HPLC device 57, thereby obtaining a total of nine data sets 95.

[0130] Using the resulting nine data sets 95, cross-validation was performed on the concentration prediction model 96 configured by the neural network 105. Specifically, eight of the nine data sets 95 were used as training data sets 95L and one was used as a validation data set 95V, and cross-validation was performed nine times while changing the configurations of the training data sets 95L and the validation data set 95V.

[0131] Next, during the progress of production process 2, the Raman spectrum of second purified solution 30 after the cation chromatography treatment was measured using flow cell 65 and Raman spectrometer 67 to obtain third spectrum measurement data 713. Then, input data 130 consisting only of intensity values ​​of the wavenumber band specific to aggregate 20 from the third spectrum measurement data 713 was input to concentration prediction model 96LD generated by the above cross-validation, and concentration prediction result 115 was output.

[0132] In Comparative Example 1, the input data 130 of the concentration prediction model 96LD is not limited to the intensity value of the specific wavenumber band of the aggregate 20, but is used for all wavenumber bands 700 cm -1 ~1800cm -1 Comparative Example 2 is an example in which the input data 130 of the concentration prediction model 96LD is set to intensity values ​​of wavenumber bands selected by sparse modeling.

[0133] In Comparative Example 3, similar to JP 2016-128822 A, the concentration prediction model 96LD is a PLS model instead of a neural network 105, and the input data 130 of the concentration prediction model 96LD is set to 800 cm -1 ~1700cm -1 Comparative Example 4 is an example in which input data 130 of concentration prediction model 96LD is set to intensity values ​​of wavenumber bands excluding the specific wavenumber band of aggregate 20.

[0134] As an example, as shown in Table 140 of FIG. 27, the Root-Mean-Square Error (RMSE) of the concentration prediction model 96LD in the example is 0.11, and the R 2The coefficient of determination (RMSE) was 0.87. 2 was 0.81, which means that the prediction accuracy of the concentration prediction model 96LD was slightly worse than in the example. From this result, it was confirmed that the prediction accuracy of the concentration prediction model 96LD was improved by selecting a wavenumber band characteristic of the aggregate 20 and setting the input data 130 of the concentration prediction model 96LD to the intensity value of the wavenumber band characteristic of the aggregate 20.

[0135] Here, Comparative Example 1 has RMSE and R 2 Therefore, at first glance, it can be understood that the prediction accuracy of the concentration prediction model 96LD is good. However, it cannot be denied that there is a possibility that wavenumber bands unrelated to the aggregate 20 are perceived as contributing to the prediction of the concentration of the aggregate 20, that is, that a spurious correlation is occurring. Therefore, it cannot be said that the concentration prediction model 96LD of Comparative Example 1 is reasonable as a model for predicting the concentration of the aggregate 20.

[0136] In addition, in the case of Comparative Example 2, the RMSE was 0.13, and R 2 was 0.81, which was a slight deterioration in the prediction accuracy of the concentration prediction model 96LD compared to the example. This result confirmed that the prediction accuracy of the concentration prediction model 96LD was improved by using intensity values ​​in the wavenumber band specific to the aggregate 20 as input data 130 to the concentration prediction model 96LD, rather than intensity values ​​in the wavenumber band selected by sparse modeling.

[0137] In the case of Comparative Example 3, the RMSE is 0.25, R 2 was 0.55, which significantly deteriorated the prediction accuracy of the concentration prediction model 96LD compared to the example. From this result, it was confirmed that by configuring the concentration prediction model 96LD using the neural network 105 instead of the PLS model and by using the intensity value of the wavenumber band specific to the aggregate 20 as the input data 130 of the concentration prediction model 96LD, the prediction accuracy of the concentration prediction model 96LD is improved compared to the technique described in JP 2016-128822 A.

[0138] In addition, in the case of Comparative Example 4, the RMSE was 0.13, and R 2was 0.82, which means that the prediction accuracy of the concentration prediction model 96LD was slightly worse than in the example. This result confirmed that the prediction accuracy of the concentration prediction model 96LD was improved by using the intensity values ​​of the wavenumber band characteristic of the aggregate 20 as the input data 130 of the concentration prediction model 96LD. This also demonstrated the rationality of the concentration prediction model 96LD, which was generated based on the intensity values ​​of the wavenumber band characteristic of the aggregate 20.

[0139] The target protein is not limited to antibody 19. It may be a cytokine, a hormone, or the like. The target component is not limited to aggregate 20. It may be a cell-derived protein, cell-derived DNA, or the like.

[0140] The spectrum is not limited to a Raman spectrum. It may be an infrared absorption spectrum, a near-infrared absorption spectrum, a nuclear magnetic resonance spectrum, an ultraviolet-visible absorption spectroscopy (UV-Vis) spectrum, or a fluorescence spectrum. In the case of an ultraviolet-visible spectrum or a fluorescence spectrum, a specific wavelength band is selected instead of a specific wavenumber band.

[0141] Even after being downloaded to the operation device 41C, the concentration prediction model 96LD may be trained using the data set 95.

[0142] Although the neural network 105 is exemplified as the concentration prediction model 96LD, the present invention is not limited to this, and may be a decision tree, a random forest, a naive Bayes, a gradient boosting decision tree, or the like.

[0143] The concentration prediction model 96LD is not limited to a machine learning model. It may also be a model generated by multivariate analysis or statistical analysis. Examples of multivariate analysis and statistical analysis include PLS, as described in JP 2016-128822 A, as well as multiple regression, principal component regression, logistic regression, Lasso regression, ridge regression, support vector regression, and Gaussian process regression. In such models generated by multivariate analysis or statistical analysis, determining the coefficients of a regression equation based on at least two datasets 95 corresponds to "generating a state prediction model" according to the technology of the present disclosure using datasets.

[0144] The state of the target component is not limited to the concentration. For example, the density of the target component may be used. Alternatively, two or more states, such as the concentration and the density, may be predicted.

[0145] In the above embodiments, examples have been shown in which the functions of the selection device 41A, the learning device 41B, and the operation device 41C are respectively performed by three computers, but this is not limited to this. The functions of the selection device 41A, the learning device 41B, and the operation device 41C may be performed by a single computer. Furthermore, the functions of the selection device 41A may be performed by a single computer, and the functions of the learning device 41B and the operation device 41C may be performed by a single computer. The functions of the selection device 41A, the learning device 41B, and the operation device 41C may be shared among four or more computers. In this way, the information processing device of the present disclosure may be performed by a single computer or multiple computers.

[0146] In each of the above embodiments, the hardware structure of the processing units that perform various processes, such as the acquisition units 80 and 120, the RW control units 81, 100, and 121, the selection unit 82, the learning verification unit 101, the prediction unit 122, and the display control unit 123, can be any of the various processors listed below. As described above, the various processors include CPUs 47A to 47C, which are general-purpose processors that execute software (operating programs 75A to 75C) and function as various processing units, as well as programmable logic devices (PLDs) that are processors whose circuit configuration can be changed after manufacture, such as FPGAs (Field Programmable Gate Arrays), and dedicated electrical circuits that are processors having a circuit configuration designed specifically for executing specific processing, such as ASICs (Application Specific Integrated Circuits).

[0147] A single processing unit may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs and / or a combination of a CPU and an FPGA).Furthermore, multiple processing units may be configured with a single processor.

[0148] Examples of configuring multiple processing units with a single processor include, first, a form in which one processor is configured with a combination of one or more CPUs and software, as typified by computers such as client and server, and this processor functions as multiple processing units. Second, a form in which a processor is used to realize the functions of the entire system including multiple processing units with a single IC (Integrated Circuit) chip, as typified by systems on chips (SoCs). In this way, various processing units are configured using one or more of the above-mentioned various processors as a hardware structure.

[0149] Furthermore, more specifically, the hardware structure of these various processors can be an electric circuit (circuitry) that combines circuit elements such as semiconductor elements.

[0150] From the above description, the technology described in the following supplementary paragraphs can be understood.

[0151] [Supplementary Item 1] An information processing device including a processor, wherein, as a preparatory process for generating a state prediction model for predicting the state of a target component in a suspension produced in a manufacturing process of a biopharmaceutical containing a target protein as an active ingredient, the processor acquires first spectral measurement data measuring the spectrum of electromagnetic waves emitted from the target protein and second spectral measurement data measuring the spectrum of electromagnetic waves emitted from the target component, and selects a specific wavenumber band or specific wavelength band specific to the target component by comparing intensity values ​​of the first spectral measurement data with intensity values ​​of the second spectral measurement data. [Supplementary Item 2] The information processing device according to Supplementary Item 1, wherein the state prediction model is generated using a dataset consisting of intensity values ​​of the specific wavenumber band or the specific wavelength band and ground truth data for the state of the target component. [Supplementary Item 3] The information processing device according to Supplementary Item 2, wherein the state of the target component is the concentration of the target component in the suspension, and wherein the concentrations of the target protein and the target component in the suspension that is the basis of the dataset are both in the ranges of 0.001 mg / mL to 20 mg / mL. [Supplementary Item 4] The information processing device of any one of Supplementary Items 1 to 3, wherein a suspension used to select the specific wavenumber band or the specific wavelength band is subjected to a pretreatment that promotes production of the target component. [Supplementary Item 5] The information processing device of any one of Supplementary Items 1 to 4, wherein the state prediction model outputs a prediction result for the state of the target component according to an intensity value of the specific wavenumber band or the specific wavelength band of third spectrum measurement data that measures the spectrum of electromagnetic waves emitted from a suspension whose state of the target component is unknown. [Supplementary Item 6] The information processing device of Supplementary Item 5, wherein the third spectrum measurement data is data measured during the progress of the manufacturing process. [Supplementary Item 7] The information processing device of Supplementary Item 5 or Supplementary Item 6, wherein the third spectrum measurement data is data measured after a virus inactivation treatment or a cation chromatography treatment.[Supplementary Item 8] The information processing device according to any one of Supplementary Items 1 to 7, wherein the first spectral measurement data and the second spectral measurement data are data measured from a first solution containing the target protein and a second solution containing the target component, which are separated from the suspension using a high-performance liquid chromatography device. [Supplementary Item 9] The information processing device according to any one of Supplementary Items 1 to 8, wherein the target component is an aggregate of the target protein. [Supplementary Item 10] The information processing device according to any one of Supplementary Items 1 to 9, wherein the state prediction model is a machine learning model. [Supplementary Item 11] The information processing device according to any one of Supplementary Items 1 to 10, wherein the target protein is an antibody. [Supplementary Item 12] The information processing device according to any one of Supplementary Items 1 to 11, wherein the spectrum is a Raman spectrum. [Supplementary Item 13] The characteristic wavenumber band is 1220 cm. -1 ~1260cm -1 range, or 1650 cm -1 ~1690cm -1 13. The information processing device according to claim 12, wherein the information processing device is at least one of the following ranges:

[0152] The technology of the present disclosure can be appropriately combined with the various embodiments and / or various modified examples described above. Furthermore, it is not limited to the above-described embodiments, and various configurations can be adopted without departing from the spirit of the present disclosure. Furthermore, the technology of the present disclosure extends not only to programs but also to storage media that non-temporarily store programs.

[0153] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[0154] In this specification, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed by connecting them with "and / or."

[0155] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

Claims

1. Comprising a processor, The processor is, As a preparatory process for generating a state prediction model for predicting the state of a target component in a suspension produced in the manufacturing process of a biopharmaceutical having a target protein as an active ingredient, Obtaining first spectrum measurement data obtained by measuring the spectrum of electromagnetic waves emitted from the target protein and second spectrum measurement data obtained by measuring the spectrum of electromagnetic waves emitted from the target component, Selecting a specific frequency band or a specific wavelength band specific to the target component by comparing the intensity value of the first spectrum measurement data with the intensity value of the second spectrum measurement data, An information processing apparatus.

2. The information processing apparatus according to claim 1, wherein the state prediction model is generated using a data set composed of the intensity value of the specific frequency band or the specific wavelength band and the correct answer data of the state of the target component.

3. The state of the target component is the concentration of the target component in the suspension, The information processing apparatus according to claim 2, wherein the concentrations of the target protein and the target component in the suspension that is the source of the data set are both in the range of 0.001 mg / mL to 20 mg / mL.

4. The suspension used for selecting the specific frequency band or the specific wavelength band is subjected to a pretreatment for promoting the generation of the target component. The information processing apparatus according to claim 1.

5. The information processing apparatus according to claim 1, wherein the state prediction model outputs a prediction result of the state of the target component according to the intensity value of the specific frequency band or the specific wavelength band of the third spectrum measurement data obtained by measuring the spectrum of electromagnetic waves emitted from a suspension with an unknown state of the target component.

6. The information processing apparatus according to claim 5, wherein the third spectrum measurement data is data measured during the progress of the manufacturing process.

7. The information processing apparatus according to claim 5, wherein the third spectrum measurement data is data measured after virus inactivation treatment or after cation chromatography treatment.

8. The information processing apparatus according to claim 1, wherein the first spectrum measurement data and the second spectrum measurement data are data measured from a first solution containing the target protein and a second solution containing the target component separated from the suspension using a high performance liquid chromatography apparatus.

9. The information processing apparatus according to claim 1, wherein the target component is an aggregate of the target protein.

10. The information processing apparatus according to claim 1, wherein the state prediction model is a machine learning model.

11. The information processing apparatus according to claim 1, wherein the target protein is an antibody.

12. The information processing apparatus according to claim 1, wherein the spectrum is a Raman spectrum.

13. The specific frequency band is 1220 cm -1 to 1260 cm -1 in the range, or 1650 cm -1 to 1690 cm -1 The information processing apparatus according to claim 12, which is in at least one of the ranges.

14. As a preparation process for generating a state prediction model for predicting the state of a target component in a suspension produced in the production process of a biopharmaceutical having a target protein as an active ingredient, acquiring first spectrum measurement data obtained by measuring the spectrum of electromagnetic waves emitted from the target protein and second spectrum measurement data obtained by measuring the spectrum of electromagnetic waves emitted from the target component, and selecting a specific frequency band or specific wavelength band specific to the target component by comparing the intensity value of the first spectrum measurement data and the intensity value of the second spectrum measurement data, A method of operating an information processing apparatus including the above.

15. As a preparation process for generating a state prediction model for predicting the state of a target component in a suspension produced in the production process of a biopharmaceutical having a target protein as an active ingredient, acquiring first spectrum measurement data obtained by measuring the spectrum of electromagnetic waves emitted from the target protein and second spectrum measurement data obtained by measuring the spectrum of electromagnetic waves emitted from the target component, and selecting a specific frequency band or specific wavelength band specific to the target component by comparing the intensity value of the first spectrum measurement data and the intensity value of the second spectrum measurement data, An operation program of an information processing apparatus for causing a computer to execute a process including the above.

16. A state prediction model for causing a computer to execute a function of outputting a prediction result of the state of the target component according to the intensity value of a specific frequency band or specific wavelength band specific to the target component in the suspension among the intensity values of each frequency or each wavelength of spectrum measurement data obtained by measuring the spectrum of electromagnetic waves emitted from a suspension produced in the production process of a biopharmaceutical having a target protein as an active ingredient.