Method for estimating the purification state

The method uses Raman spectroscopy and machine learning to construct a soft sensor for accurate impurity concentration estimation in protein purification, addressing the challenge of quantifying trace impurities and ensuring drug efficacy.

JP2026063012APending Publication Date: 2026-04-10FUJIFILM CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
FUJIFILM CORP
Filing Date
2026-01-08
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing methods for protein purification, such as those described in U.S. Patent Application Publication No. 2020/0062802, fail to accurately quantify trace amounts of impurities in biopharmaceuticals, which can affect drug efficacy, due to the low concentration and small amounts of impurities present in the purification process.

Method used

A method utilizing Raman spectroscopy and machine learning to construct a soft sensor that estimates the concentration of impurities and proteins in a purification process by preprocessing spectral data and selecting highly correlated spectral data for training, allowing for accurate estimation of impurity concentrations even at trace levels.

Benefits of technology

Enables precise estimation of impurity concentrations in protein purification processes, ensuring high accuracy and immediate detection of abnormalities, even at low impurity levels, thereby maintaining drug efficacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026063012000004
    Figure 2026063012000004
  • Figure 2026063012000005
    Figure 2026063012000005
  • Figure 2026063012000006
    Figure 2026063012000006
Patent Text Reader

Abstract

This invention provides a method for estimating the purification state of a protein, which can accurately estimate the concentration of impurities other than protein even when the amount of impurities in the treatment solution after protein purification is trace. [Solution] The method includes quantifying the components contained in a processed liquid obtained by purifying a liquid containing specific proteins and non-protein impurities. The method for estimating the purification state includes obtaining an estimated value of the concentration of impurities based on spectral data showing the intensity for each wavenumber or wavelength of electromagnetic waves irradiated onto the processed liquid and affected by the processed liquid. The concentration of impurities contained in the processed liquid is 20 mg / mL or less, and the weight ratio of impurities to the mixture containing proteins and impurities is 15% or less.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The disclosed technology relates to a method for estimating the purification status of a liquid containing a specific protein after purification treatment. [Background technology]

[0002] The following techniques are known for purifying proteins such as antibodies produced by cells. For example, U.S. Patent Application Publication No. 2020 / 0062802 describes a technique for quantifying a protein purification intermediate during production using in-line Raman spectroscopy. [Overview of the project] [Problems that the invention aims to solve]

[0003] In the manufacture of biopharmaceuticals, proteins such as antibodies, which are biopharmaceutical active pharmaceutical ingredients produced from cultured cells, are purified and formulated. In the protein purification process, the purity of the target protein is increased stepwise by performing purification treatment using multiple different chromatographic methods, such as cation chromatography, anion chromatography, immunoaffinity chromatography, and gel filtration chromatography. It is preferable to monitor the purification state to verify whether the purification treatment is being performed appropriately at each step. In particular, it is important to quantify the impurities separated from the target protein at each step. This is because if impurities other than the target protein are mixed into the pharmaceutical product, even in trace amounts, it may affect the efficacy of the drug. In the purification process, the purity of the target protein is increased stepwise, and the amount of impurities contained in the processing solution processed at each step becomes very small, making it difficult to quantify the impurities. Patent document 1 described above describes a technique for quantifying intermediates of protein purification during manufacturing, but it does not describe the quantification of impurities.

[0004] The disclosed technology was developed in view of the above points, and aims to provide a method for estimating the purification state that can accurately estimate the concentration of impurities other than protein, even when the amount of impurities other than protein in the treatment solution after protein purification is trace. [Means for solving the problem]

[0005] The method for estimating the purified state relating to the disclosed technology is a method for estimating the purified state that includes quantifying the components contained in a processed liquid obtained by subjecting a liquid containing a specific protein and non-protein impurities to a purification treatment, and includes obtaining an estimated value of the concentration of impurities based on spectral data showing the intensity for each wavenumber or wavelength of electromagnetic waves irradiated onto the processed liquid and affected by the processed liquid. The concentration of impurities contained in the processed liquid may be 20 mg / mL or less, and the weight ratio of impurities to the mixture containing the protein and impurities may be 15% or less.

[0006] The method for estimating the purified state relating to the disclosed technology is a method for estimating the purified state that includes quantifying the components contained in a treated solution obtained by subjecting a liquid containing a specific protein and impurities other than the specific protein to a purification treatment, and includes obtaining an estimated value of the concentration of immature glycans that are structurally similar to the protein based on spectral data showing the intensity for each wavenumber or wavelength of electromagnetic waves irradiated onto the treated solution and affected by the treated solution.

[0007] The method for estimating the purification state relating to the disclosed technology may further include obtaining an estimate of the concentration of proteins contained in the processing solution based on spectral data. Specific proteins may be produced from cultured cells. Contaminants may include DNA from cells producing specific antibodies, protein aggregates, protein degradation products, and host cell-derived proteins. The purification process may include chromatographic component separation techniques. The coefficient of determination, indicating the degree of agreement between the estimated concentration of contaminants and the measured value, may be 0.9 or higher. The root mean square error, indicating the degree of deviation between the estimated concentration of contaminants and the measured value, may be 1.2 or lower.

[0008] The method for estimating the purification state relating to the disclosed technology may include constructing a soft sensor that takes spectral data as input and outputs state data by machine learning using multiple combinations of state data indicating the purification state of a liquid containing proteins and impurities and spectral data as training data, and obtaining state data output from the soft sensor by inputting spectral data acquired for the processing liquid into the soft sensor. The state data may include an estimated value of the concentration of impurities contained in the processing liquid.

[0009] The method for estimating the purification state related to the disclosed technology may include preprocessing spectral data and constructing a soft sensor by machine learning using multiple combinations of the preprocessed data and state data as training data. The preprocessing may include selecting the spectral intensity values ​​for each wavenumber or wavelength included in the spectral data to be used as training data. Through selection, the number of intensity values ​​for each wavenumber or wavelength included in the spectral data to be used as training data may be set to 5 or more and less than 1000. Selection may be performed by sparse modeling. The preprocessing may include identifying highly correlated spectral data, which have a relatively high correlation with the state data, as processed data. The preprocessing may include baseline correction of the spectral data.

[0010] The spectral data may represent the spectrum of scattered light when light is irradiated onto a liquid containing proteins and impurities. The state data may include an estimate of the protein concentration in the processed liquid. [Effects of the Invention]

[0011] The disclosed technology provides a method for estimating the purification state, which can accurately estimate the concentration of impurities other than protein even when the amount of impurities other than protein in the treatment solution after protein purification is trace. [Brief explanation of the drawing]

[0012] [Figure 1] This figure shows an example of the antibody purification process according to an embodiment of the disclosed technology. [Figure 2] This figure shows an example of a method for estimating the purification state according to an embodiment of the disclosed technology. [Figure 3] This figure shows an example of a method for obtaining spectral data. [Figure 4] This figure shows an example of training data according to an embodiment of the disclosed technology. [Figure 5] This figure shows an example of a method for estimating the purification state according to an embodiment of the disclosed technology. [Figure 6] This figure shows an example of the hardware configuration of an information processing device according to an embodiment of the disclosed technology. [Figure 7] This figure shows an example of the structure of an estimation model according to an embodiment of the disclosed technology. [Figure 8] This is an example of a functional block diagram showing an example of the functional configuration of an information processing device in the learning phase according to an embodiment of the disclosed technology. [Figure 9] This flowchart shows an example of the flow of the soft sensor construction process according to the disclosed technology. [Figure 10] This is an example of a functional block diagram showing an example of the functional configuration of an information processing device in the operational phase according to an embodiment of the disclosed technology. [Figure 11] It is a flowchart showing an example of the flow of estimation processing according to an embodiment of the disclosed technology. [Figure 12A] It is a graph showing the relationship between the estimated value and the measured value of the impurity concentration. [Figure 12B] It is a graph showing the relationship between the estimated value and the measured value of the impurity concentration. [Figure 12C] It is a graph showing the relationship between the estimated value and the measured value of the impurity concentration. [Figure 13A] It is a graph showing the relationship between the estimated value and the measured value of the antibody concentration. [Figure 13B] It is a graph showing the relationship between the estimated value and the measured value of the antibody concentration. [Figure 13C] It is a graph showing the relationship between the estimated value and the measured value of the antibody concentration. [Figure 14] It is a graph showing the relationship between the estimated value and the measured value of the immature sugar chain concentration.

Embodiments for Carrying Out the Invention

[0013] Hereinafter, an example of an embodiment of the disclosed technology will be described while referring to the drawings. In each drawing, the same or equivalent components and parts are given the same reference numerals, and duplicate explanations are omitted as appropriate.

[0014] The method for estimating the purification state according to the embodiment of the disclosed technology includes quantifying the components contained in the treatment liquid obtained by subjecting a liquid containing a specific protein and impurities other than the protein to a purification treatment. More specifically, it includes obtaining an estimated value of the concentration of impurities contained in the treatment liquid based on the spectral data indicating the intensity for each wave number or wavelength of the electromagnetic wave irradiated to the treatment liquid and affected by the treatment liquid. The method for estimating the purification state according to the disclosed technology is particularly effective when the concentration of impurities contained in the treatment liquid is 20 mg / mL or less and the weight ratio of impurities to the mixture containing the protein and impurities is 15% or less. Further, the method for estimating the purification state according to the disclosed technology may include obtaining an estimated value of the concentration of a specific protein contained in the above treatment liquid.

[0015] Specific proteins may include, for example, immunoglobulins, i.e., antibodies, produced from cultured cells. Impurities include, for example, immature glycans with a similar structure to antibodies, cellular DNA, antibody aggregates, antibody degradation products, and host cell proteins (HCPs). Immature glycans with a similar structure to antibodies are likely to form, for example, when the amount of waste products in the culture medium increases or the oxygen concentration in the culture medium is insufficient during the culture period of antibody-producing cells. Antibody degradation products are formed when antibodies are broken down by degrading enzymes produced during the culture period. Antibody aggregates are likely to form, for example, when the concentration of antibodies produced from cells becomes excessively high or when stress such as heat is applied. DNA released from cells means that the cell membrane of that cell has broken down, i.e., the cell has become dead. Host cell proteins are proteins derived from host cells that are purified along with the antibodies during the antibody purification process. If contaminants like those mentioned above are present in pharmaceuticals that use antibodies produced from cells, even in trace amounts, it can affect the efficacy of the drug. Therefore, it is important to quantify contaminants in the processing solution obtained through the antibody purification process.

[0016] Figure 1 shows an example of an antibody purification process according to an embodiment of the disclosed technology. As shown in Figure 1, the antibody purification process includes an immunoaffinity chromatography purification process P1, a virus inactivation process P2, a cation chromatography purification process P3, an anion chromatography purification process P4, a virus filtering process P5, and a concentration and filtration process P6.

[0017] The immunoaffinity chromatography purification process P1 is a process of extracting antibodies using a column immobilized with a ligand such as protein A that has affinity for the antibody. The virus inactivation process P2 is a process to inactivate the virus contained in the treatment solution obtained by purification process P1. The cation chromatography purification process P3 is a process of extracting antibodies using a column with a cation exchanger as the stationary phase. The anion chromatography purification process P4 is a process of extracting antibodies using a column with an anion exchanger as the stationary phase. The virus filtering process P5 is a process of removing viruses contained in the treatment solution obtained by each of the above processes using a filter. The concentration and filtration process P6 is a concentration and filtration process using UF (Ultrafiltration) and DF (Diafiltration).

[0018] As described above, by performing multiple processes, including component separation methods using multiple different chromatography techniques, in a stepwise manner, impurities are progressively removed and the purity of the antibody is progressively increased. It is preferable to monitor the purification state to verify whether appropriate processing is being performed at each step. The method for estimating the purification state according to the embodiment of the disclosed technology can be used to estimate the purification state of each processing solution obtained in each of the processes P1 to P6 shown in Figure 1. Preferably, the purification state can be estimated for each processing solution obtained in each of the processes P1 to P6, and if there is a subsequent process, the estimated purification state can be used to determine the purification conditions for the subsequent process. Preferably, the purification process can be performed while estimating the antibody concentration, impurity concentration, and immature glycan concentration in process P1. Details of the method for estimating the purification state according to the embodiment of the disclosed technology will be described below.

[0019] A method for estimating the purification state according to an embodiment of the disclosed technology includes constructing a soft sensor that takes spectral data as input and outputs state data by machine learning using multiple combinations of state data indicating the purification state of a liquid containing a specific protein and impurities to be purified, and spectral data indicating the intensity for each wavenumber or wavelength of electromagnetic waves irradiated onto the processed liquid obtained by the purification process and affected by the processed liquid, as training data. A method for estimating the purification state according to an embodiment of the disclosed technology includes obtaining state data output from a soft sensor by inputting spectral data acquired for the processed liquid obtained by the purification process into the soft sensor. The state data includes an estimated value of the concentration of impurities contained in the processed liquid.

[0020] Furthermore, the method for estimating the purification state according to the embodiment of the disclosed technology includes preprocessing spectral data and constructing the soft sensor by machine learning using multiple combinations of the preprocessed data obtained by the preprocessing and state data as training data. The preprocessing method may include dimensional reduction techniques such as sparse modeling, PCA (principal component analysis), LSA (SVD) (latent semantic analysis (singular value decomposition)), LDA (linear discriminant analysis), ICA (independent component analysis), and PLS (partial Least Squares Regression), and may include a process of selecting spectral intensity values ​​for each wavenumber or wavelength included in the spectral data to be used as training data. In this case, the spectral intensity values ​​of the wavenumber or wavelength that remain after selection become the processed data. It is expected that the spectral intensity values ​​for each wavenumber or wavelength that constitute the spectral data will be enormous in number. By selecting the data to be used as training data, it is possible to prevent a decrease in prediction accuracy due to overfitting of the model data. The selection of spectral data can be performed, for example, by sparse modeling. In other words, preprocessing performed on spectral data may include using sparse modeling to exclude data from the spectral data that has a relatively low correlation with the state data, thereby identifying highly correlated spectral data, which has a relatively high correlation with the state data, as processed data. It is preferable to select the number of spectral intensity values ​​for each wavenumber or wavelength included in the spectral data to be used as training data from 5 to less than 1000 through preprocessing, more preferably from 5 to 800, and even more preferably from 5 to 500.

[0021] In this embodiment, sparse modeling refers to selecting explanatory variables (i.e., excluding some explanatory variables) in a regression model where the spectral intensity values ​​for each wavenumber or wavelength included in the spectral data are the explanatory variables and the state data is the dependent variable. As a sparse modeling method, for example, Lasso regression can be used. Lasso regression is a method of selecting explanatory variables so that the cost function calculated by adding a penalty term to the root mean squared error (RMSE) is minimized. In this embodiment, explanatory variables are selected by excluding low-correlation spectral data, which have a relatively low correlation with the state data, from the spectral data. The penalty term may be determined by cross-validation, such as K-fold cross-validation. In the following description, we will explain using the case where the preprocessing performed on the spectral data is a process to identify high-correlation spectral data as an example.

[0022] A liquid containing a specific protein and other impurities can be prepared by known methods, such as culturing cells possessing a gene encoding the specific protein, decellularizing the resulting culture, and purifying it by chromatography. For example, it can be prepared by culturing CHO cells into which the IgG1 antibody gene has been introduced, decellularizing them by filtering, and then purifying them by chromatography with Protein A. By changing the purification conditions such as pH and temperature with Protein A, the ratio of the specific protein and other impurities can be changed. In this embodiment, an aqueous sodium acetate solution was used as the buffer during purification. In protein purification, mainly phosphate-based and acetate-based buffers are used. Since the characteristic wavenumbers of these buffers are well known, it is possible to predict the result regardless of the buffer by removing these wavenumbers. The disclosed technology can be applied regardless of the type of protein. The difference between antibody types lies in the difference in amino acid sequence. Since this difference in amino acid sequence does not appear in spectral differences, the technology can be applied regardless of the type of antibody. The disclosed technology can be applied to immature glycans, regardless of the type of immature glycan.

[0023] As shown in Figure 2, a method for estimating the purification state according to an embodiment of the disclosed technology includes the step of obtaining state data output from a soft sensor 20 by inputting highly correlated spectral data as processed data from spectral data obtained from a processed solution obtained by one of the multiple processes exemplified in Figure 1, which is performed in the antibody purification process. The soft sensor 20 implements the process of outputting state data based on the input highly correlated spectral data using software. The soft sensor 20 is constructed in an information processing device 10 (see Figures 3 and 6) described later.

[0024] In this embodiment, the soft sensor 20 is subjected to an analysis method based on Raman spectroscopy. Specifically, Raman scattered light spectral data is applied as the spectral data input to the soft sensor 20. Raman spectroscopy is a spectroscopic method that evaluates materials using Raman scattered light. When light is shone on a material, the light interacts with the material, generating Raman scattered light with a different wavelength from the incident light. The wavelength difference between the incident light and the Raman scattered light corresponds to the energy of the molecular vibrations of the material, so Raman scattered light with different wavelengths (wavenumbers) can be obtained between materials with different molecular structures. Furthermore, various physical properties such as stress, temperature, electrical properties, orientation, and crystallinity can be estimated using Raman scattered light. Of the Stokes lines and anti-Stokes lines, it is preferable to use Stokes lines for the Raman scattered light. In this embodiment, the Raman spectrum was collected with a laser output of 500 mW, a measurement wavelength of 785 nm, and a laser irradiation time of 1 second.

[0025] Figure 3 shows an example of a method for acquiring spectral data for a processing solution 31 obtained by any of the processes P1 to P6 shown in Figure 1. The spectral data can be acquired using a known Raman spectroscopy probe 40 and analyzer 41. As shown in Figure 3, the tip of the probe 40 is immersed in the processing solution 31 contained in the container 30. Excitation light emitted from a light emitter (not shown) provided at the tip of the probe 40 is irradiated onto the processing solution 31. Raman scattered light generated by the interaction between the excitation light and the processing solution 31 is received by a light receiver (not shown) provided at the tip of the probe 40. The acquired Raman scattered light is decomposed by the analyzer 41 for each wavenumber (reciprocal of wavelength), and spectral data, which is the spectral intensity value for each wavenumber, is generated. The spectral data may also be spectral intensity values ​​for each wavelength. The spectral data is supplied to the information processing device 10.

[0026] The state data output from the soft sensor 20 is data indicating the purification state, which is correlated with the spectral data. The state data includes an estimated concentration of impurities contained in the processing solution 31. The state data may also include an estimated concentration of antibodies contained in the processing solution 31. These state data are not easily monitored in-line by actual measurement. By using the soft sensor 20, it becomes possible to acquire the state data in-line based on spectral data, which is relatively easy to monitor in-line by actual measurement.

[0027] The soft sensor 20 is constructed by machine learning using multiple combinations of spectral data and state data as training data. Figure 4 shows an example of training data 50. Training data 50 is acquired, for example, during the process development stage when purification processing conditions are examined. Training spectral data is acquired, for example, from processing solutions obtained by changing various purification conditions. Purification conditions include, for example, the flow rate when injecting the liquid to be purified into the column, the amount of buffer used when eluting antibodies from the column, and the composition of the buffer.

[0028] The state data for training can be obtained by measurement using conventional sampling methods, with the processed solution from which the training spectral data was acquired as the measurement target. For example, when obtaining the concentration of impurities contained in the processed solution as training state data, it is possible to obtain it using methods such as HPLC (high performance liquid chromatography) for each type of impurity. The training data is acquired for each purification condition, and the training spectral data and training state data for each condition are associated with each other.

[0029] Here, the analyzer 41 receives spectral data, for example, at wavenumber 500 cm⁻¹. -1 From 3000cm -1 Spectral intensity values ​​up to 1 cm -1 The output is generated in increments. Consequently, the number of acquired spectral data points becomes enormous, and if all spectral data is used as training data, the training load becomes excessive, requiring a high-performance processor to perform machine learning. Furthermore, the spectral intensity values ​​of the Raman scattered light that make up the spectral data may include spectral intensity values ​​of wavenumbers that have a low correlation with the state data of the monitored object. For example, spectral intensity values ​​of certain wavenumbers of Raman scattered light may have a low correlation with the concentration of impurities. If the soft sensor 20 is constructed by machine learning using spectral data that includes spectral intensity values ​​of wavenumbers that have a low correlation with the state data of the monitored object as training data, the accuracy of the output values ​​of the soft sensor 20 may decrease.

[0030] Therefore, in this embodiment, as a preprocessing step for spectral data, spectral intensity values ​​of wavenumbers with a relatively high correlation to the state data of the monitoring target are identified as high-correlation spectral data from among the spectral data output from the analyzer 41. Then, in the learning phase, which is the stage in which the soft sensor 20 is constructed by machine learning, the soft sensor 20 is constructed by machine learning using multiple combinations of high-correlation spectral data and state data as learning data. On the other hand, in the operation phase, in which the constructed soft sensor 20 is operated to acquire state data for the processed liquid obtained by the purification process, as shown in Figure 5, state data is obtained from the soft sensor 20 by inputting high-correlation spectral data with a relatively high correlation to the state data of the monitoring target from among the spectral data acquired for the processed liquid obtained by the purification process into the soft sensor 20. The construction of the soft sensor 20 and the acquisition of state data using the soft sensor 20 are performed by the information processing device 10.

[0031] The color of the processing solution from which spectral data is acquired changes depending on the amount of impurities it contains, the type of antibody, and the type of cells that produce the antibody. In addition, external environmental factors such as temperature, humidity, and vibration during spectral data acquisition, as well as fluctuations in the output of the excitation light irradiated onto the processing solution, act as disturbances to the spectral data. These factors may cause fluctuations in the baseline of the spectral data. Baseline fluctuations can lead to a decrease in the accuracy of the output value of the soft sensor 20. Therefore, in this embodiment, baseline correction of the spectral data is performed as a preprocessing step for the spectral data. Baseline correction means removing fluctuations in the baseline of the spectral data caused by disturbances. Baseline correction may be performed, for example, by differentiating the spectral waveform. Alternatively, it may be performed, for example, by removing the baseline obtained by polynomial fitting from the spectral waveform.

[0032] Figure 6 shows an example of the hardware configuration of the information processing device 10. The information processing device 10 includes a CPU (Central Processing Unit) 101, a memory 102 as a temporary storage area, and a non-volatile storage unit 103. The information processing device 10 also includes a display unit 104 such as a liquid crystal display, an input unit 105 such as a keyboard and mouse, a network interface 106 connected to a network, and an external interface 107 to which the analyzer 41 is connected. The CPU 101, memory 102, storage unit 103, display unit 104, input unit 105, network interface 106, and external interface 107 are connected to a bus 108.

[0033] The memory unit 103 is implemented by a storage medium such as an HDD (Hard Disk Drive), SSD (Solid State Drive), or flash memory. The memory unit 103 stores training data 50, an estimation model 60, a soft sensor construction program 70, and an estimation program 80. As shown in Figure 4, the training data 50 is a combination of spectral data and state data.

[0034] Figure 7 shows an example of the structure of the estimated model 60. The estimated model 60 is a neural network that includes an input layer, multiple hidden layers, and an output layer. The input layer of the estimated model 60 receives spectral intensity values ​​for each wavenumber of Raman scattered light, i.e., spectral data. The output layer of the estimated model 60 outputs state data corresponding to the spectral data input to the input layer.

[0035] In the learning phase, the CPU 101 reads the soft sensor construction program 70 from the memory unit 103, loads it into the memory 102, and executes it. In the operation phase, the CPU 101 reads the estimation program 80 from the memory unit 103, loads it into the memory 102, and executes it. An example of the information processing device 10 is a server computer, etc. The CPU 101 is an example of a processor in the disclosed technology.

[0036] Figure 8 is an example of a functional block diagram showing an example of the functional configuration of the information processing device 10 during the learning phase. During the learning phase, the information processing device 10 is configured to include a specific unit 11 and a learning unit 12. The storage unit 103 is assumed to store learning data 50 and an estimated model 60.

[0037] The identification unit 11 performs regression analysis on the training data 50 using Lasso regression, an example of sparse modeling, to identify spectral intensity values ​​with relatively high correlations to state data from among the spectral data included in the training data 50 as high-correlation spectral data. Specifically, the identification unit 11 performs the following processing: The identification unit 11 performs a process to thin out spectral intensity values ​​with randomly determined wavenumbers from the spectral data included in the training data 50, and generates a regression model (regression equation) that shows the relationship between the thinned spectral data and the corresponding state data. The identification unit 11 derives a cost function for the generated regression model by adding a penalty term to the root mean squared error (RMSE). The identification unit 11 repeats each of the above processing a predetermined number of times to generate a regression model for each of multiple spectral data with different wavenumbers that are thinned out, and derives the above cost function for each regression model. The identification unit 11 identifies the fewest number of spectral intensity values ​​that can minimize the above cost function within a predetermined number of iterative calculations as highly correlated spectral data.

[0038] The learning unit 12 trains an estimation model 60 using machine learning, with the combination of highly correlated spectral data identified by the identification unit 11 from the training data 50 and the corresponding state data as training data. This constructs a soft sensor 20 that takes highly correlated spectral data as input and state data as output.

[0039] The learning unit 12 trains the estimation model 60 using the training data 50, following the backpropagation method as an example of machine learning. Specifically, the learning unit 12 extracts highly correlated spectral data identified by the identification unit 11 from the training spectral data included in the training data 50. The learning unit 12 inputs the extracted highly correlated spectral data into the estimation model 60 and obtains state data output from the estimation model 60. The learning unit 12 trains the estimation model 60 so that the difference between the score indicated by the obtained state data and the score indicated by the training state data corresponding to the highly correlated spectral data included in the training data 50 is minimized. The learning unit 12 performs the training of the estimation model 60 using all or part of the combination of highly correlated spectral data and state data included in the training data 50. In addition to backpropagation, other machine learning methods include random forest, linear regression, nonlinear regression (SVM: Support vector machine, Bayesian regression), logistic regression, etc., but backpropagation is preferred.

[0040] Figure 9 is a flowchart showing an example of the flow of the soft sensor construction process performed by the CPU 101 when it executes the soft sensor construction program 70 during the learning phase. The soft sensor construction program 70 is executed, for example, when the user inputs an instruction to execute the soft sensor construction process via the input unit 105.

[0041] In step S1, the identification unit 11 randomly selects spectral intensity values ​​of wavenumbers to be excluded from the spectral data included in the learning data 50 stored in the memory unit 103. That is, the identification unit 11 selects the spectral intensity values ​​of wavenumber 1 cm. -1 The spectral intensity values ​​obtained in increments are thinned out by removing the spectral intensity values ​​for some wavenumbers. The number of wavenumbers to be excluded may be a predetermined number or a randomly determined number. It is preferable that the number of wavenumbers to be excluded is predetermined.

[0042] In step S2, the identification unit 11 generates a regression model (regression equation) that shows the relationship between spectral data composed of spectral intensity values ​​of wavenumbers other than those selected for exclusion in step S1 (i.e., decimated spectral data) and the corresponding state data. Specifically, a regression model is estimated using statistical methods, with the decimated spectral data as explanatory variables and the corresponding state data as the dependent variable. The regression model may be a linear model or a nonlinear model.

[0043] In step S3, the identification unit 11 derives a cost function for the regression model generated in step S2. The cost function is used as an index value indicating the accuracy of the regression model.

[0044] In step S4, the identification unit 11 determines whether the number of repetitions of the process from step S1 to step S3 has reached a predetermined number. The identification unit 11 repeatedly performs the process from step S1 to step S3 until the number of repetitions reaches the predetermined number. As a result, a regression model is generated for each of the multiple thinned spectral data sets with different wavenumbers to be excluded, and a cost function is derived for each of the generated regression models.

[0045] In step S5, the identification unit 11 identifies the downsampled spectral data used to generate the regression model that minimizes the cost function as highly correlated spectral data. The spectral data used to generate the regression model that minimizes the cost function consists of spectral intensity values ​​of wavenumbers that have a relatively high correlation with the state data. In this way, the identification unit 11 identifies spectral data consisting of spectral intensity values ​​of wavenumbers that have a relatively high correlation with the state data as highly correlated spectral data through regression analysis.

[0046] In step S6, the learning unit 12 extracts the highly correlated spectral data identified in step S5 from the spectral data contained in the training data 50 stored in the memory unit 103, and trains the estimation model 60 using machine learning with multiple combinations of the extracted highly correlated spectral data and corresponding state data as training data. Specifically, the learning unit 12 inputs the highly correlated spectral data identified in step S5 into the estimation model 60 and trains the estimation model 60 so that the difference between the score indicated by the state data output from the estimation model 60 and the score indicated by the training state data corresponding to the highly correlated spectral data contained in the training data 50 is minimized. This constructs the soft sensor 20.

[0047] The soft sensor 20 is constructed for each type of state data to be monitored. For example, when the soft sensor 20 is to output an estimated value of the concentration of impurities contained in the processing solution obtained by the purification process as state data, spectral intensity values ​​of wavenumbers that have a high correlation with the concentration of impurities are identified as high-correlation spectral data from among the spectral data. Then, the soft sensor 20 is constructed to output an estimated value of the concentration of impurities based on the high-correlation spectral data by machine learning, using multiple combinations of the identified high-correlation spectral data and state data indicating the concentration of impurities obtained by actual measurement as training data. On the other hand, when the soft sensor 20 is to output an estimated value of the concentration of antibodies as state data, spectral intensity values ​​of wavenumbers that have a high correlation with the concentration of antibodies are identified as high-correlation spectral data from among the spectral data. Then, the soft sensor 20 is constructed to output an estimated value of the concentration of antibodies based on the high-correlation spectral data by machine learning, using multiple combinations of the identified high-correlation spectral data and state data indicating the concentration of antibodies obtained by actual measurement as training data.

[0048] Figure 10 is an example of a functional block diagram showing an example of the functional configuration of the information processing device 10 in the operational phase. In the operational phase, the information processing device 10 is composed of an acquisition unit 13, an extraction unit 14, and an estimation unit 15. The storage unit 103 is assumed to store a trained estimation model 60 that functions as a soft sensor 20.

[0049] A method for estimating the purification state according to an embodiment of the disclosed technology is applied, for example, to quantify the components of a processing solution obtained by a purification process for extracting antibodies. As shown in Figure 3, spectral data is acquired from the processing solution 31 contained in a container 30 by a probe 40 and an analyzer 41.

[0050] The acquisition unit 13 acquires spectral data output from the analyzer 41. The extraction unit 14 extracts highly correlated spectral data from the spectral data acquired by the acquisition unit 13, that is, spectral intensity values ​​of wavenumbers that have a relatively high correlation with the state data of the monitored object, as identified by the identification unit 11.

[0051] The estimation unit 15 reads a trained estimation model 60, which functions as a soft sensor 20, from the storage unit 103, inputs the highly correlated spectral data extracted by the extraction unit 14 to the estimation model 60, and acquires state data output from the estimation model 60. The estimation unit 15 may also control the display unit 104 to display the acquired state data. Alternatively, the estimation unit 15 may store the acquired state data in the storage unit 103.

[0052] Figure 11 is a flowchart showing an example of the flow of estimation processing performed by the CPU 101 when it executes the estimation program 80 during the operation phase. The estimation program 80 is executed, for example, when an instruction to execute estimation processing is input by the user via the input unit 105.

[0053] In step S11, the acquisition unit 13 acquires spectral data output from the analyzer 41. In step S12, the extraction unit 14 extracts highly correlated spectral data from the spectral data acquired by the acquisition unit 13, that is, spectral intensity values ​​of wavenumbers with a relatively high correlation to the state data of the monitoring target, as identified by the identification unit 11. In step S13, the estimation unit 15 reads a trained estimation model 60, which functions as a soft sensor 20, from the storage unit 103, inputs the highly correlated spectral data extracted in step S12 into the read estimation model 60, and acquires state data output from the estimation model 60. The estimation unit 15 controls the display unit 104 to display the acquired state data.

[0054] The estimated concentration of impurities contained in the processed solution obtained by purifying a liquid containing antibodies and impurities was acquired using a soft sensor 20. Figures 12A to 12C are graphs showing the relationship between the estimated impurity concentration obtained using the soft sensor 20 and the measured impurity concentration obtained by sampling. As a comparative example, Figures 12A to 12C also show the relationship between the estimated impurity concentration obtained by analyzing spectral data using PLS, a multivariate analysis method, and the measured value. In Figures 12A to 12C, the processed solution obtained by treatment P1 was used as the liquid containing antibodies and impurities. In Figures 12A to 12C, the case where the estimated impurity concentration was obtained using the soft sensor 20 (example) is shown as a white diamond plot and a solid line, and the case where the estimated impurity concentration was obtained using PLS (comparative example) is shown as a black square plot and a dotted line.

[0055] Proteins, impurities, and proteins containing immature glycans in a liquid containing proteins and other impurities can be measured using known methods. For example, proteins can be measured by subjecting the liquid to protein A chromatography. Impurities can be measured by size exclusion chromatography. Immature glycans can be measured by glycan liberation treatment, fluorescent labeling of the liberated glycans, removal of unreacted substances, and then measurement of the immature glycan concentration by HPLC.

[0056] Figure 12A shows the case where the impurity ratio is 2.5%, Figure 12B shows the case where the impurity ratio is 5%, and Figure 12C shows the case where the impurity ratio is 10%. The impurity ratio is the weight ratio of the impurity to the mixture containing the antibody and the impurity, and is defined by the following equation (1). In equation (1), R C is the ratio of impurities, A is the weight of antibodies contained in the treatment solution, and C is the weight of impurities contained in the treatment solution. R C =C / (A+C)···(1)

[0057] For each of the estimated values ​​related to the example and the estimated values ​​related to the comparative example, the coefficient of determination (R) indicates the degree of agreement with the measured values. 2 The results of calculating the root mean squared error (RMSE), which indicates the degree of deviation from the measured value, are shown in Table 1 below.

[0058] [Table 1]

[0059] As shown in FIGS. 12A to 12C and Table 1, it was confirmed that the estimated value obtained using the soft sensor 20 for the impurity concentration was more accurate than the estimated value obtained using PLS. In particular, by using the soft sensor 20, even when the impurity concentration was 5 mg / mL or less and the impurity ratio was 2.5%, the impurity concentration could be estimated with extremely high accuracy.

[0060] An estimated value of the antibody concentration contained in the treatment liquid obtained by subjecting a liquid containing an antibody and an impurity to a purification treatment was obtained using the soft sensor 20. FIGS. 13A to 13C are graphs showing the relationship between the estimated value of the antibody concentration obtained using the soft sensor 20 and the actually measured value of the antibody concentration obtained by sampling, respectively. In FIGS. 13A to 13C, as a comparative example, the relationship between the estimated value of the antibody concentration obtained by analyzing spectral data using PLS, which is one of the multivariate analysis methods, and the actually measured value is also shown. In FIGS. 13A to 13C, the treatment liquid obtained by treatment P1 was used as the liquid containing an antibody and an impurity. In FIGS. 13A to 13C, the case where the estimated value of the antibody concentration was obtained using the soft sensor 20 (Example) is indicated by an unfilled diamond plot and a solid line, and the case where the estimated value of the antibody concentration was obtained using PLS (Comparative Example) is indicated by a black-filled square plot and a dotted line.

[0061] FIG. 13A shows the case where the antibody ratio is 20%, FIG. 13B shows the case where the antibody ratio is 50%, and FIG. 13C shows the case where the antibody ratio is 80%. The antibody ratio is the weight ratio of the antibody to the mixture containing the antibody and the impurity, and is defined by the following formula (2). In formula (2), R A is the antibody ratio, A is the weight of the antibody contained in the treatment liquid, and C is the weight of the impurity contained in the treatment liquid. R A = A / (A + C) ··· (2)

[0062] For each of the estimated value according to the example and the estimated value according to the comparative example, the coefficient of determination (R 2The results of calculating the root mean square error (RMSE), which indicates the degree of deviation from the measured value, are shown in Table 2 below.

[0063] [Table 2]

[0064] As shown in Figures 13A to 13C and Table 2, it was confirmed that the estimated antibody concentration obtained using the soft sensor 20 was more accurate than the estimated value obtained using PLS. In particular, by using the soft sensor 20, it was possible to estimate the antibody concentration with extremely high accuracy even when the antibody concentration was 5 mg / mL or less and the antibody ratio was 20%.

[0065] The estimated concentration of immature glycans, a type of impurity, contained in the processed solution obtained by purifying a liquid containing antibodies and impurities was acquired using a soft sensor 20. Figure 14 is a graph showing the relationship between the estimated immature glycan concentration acquired using the soft sensor 20 and the measured immature glycan concentration acquired by sampling (Example 1). Figure 14 also shows the relationship between the estimated immature glycan concentration acquired by analyzing spectral data using PLS, one of the multivariate analysis methods, and the measured value (Example 2). In Figure 14, the case where the estimated immature glycan concentration was acquired using the soft sensor 20 (Example 1) is shown as an open diamond plot and a solid line, and the case where the estimated antibody concentration was acquired using PLS (Example 2) is shown as a filled square plot and a dotted line.

[0066] For each of the estimated values ​​related to Example 1 and Example 2, the coefficient of determination (R) indicates the degree of agreement with the measured values. 2 The results of calculating the root mean square error (RMSE), which indicates the degree of deviation from the measured value, are shown in Table 3 below. [Table 3]

[0067] As shown in Figure 14 and Table 3, it was confirmed that the estimated values ​​of immature glycan concentrations obtained using the soft sensor 20 were more accurate than the estimated values ​​obtained using PLS.

[0068] As described above, according to the purification state estimation method of the disclosed technology, even when the amount of impurities other than protein in the processing solution that has undergone purification treatment for extracting a specific protein is trace, the concentration of impurities can be estimated with high accuracy. For example, even when the concentration of impurities in the processing solution is 20 mg / mL or less and the impurity ratio is 15% or less, the concentration of impurities can be estimated with high accuracy.

[0069] Furthermore, since spectral data can be monitored relatively easily in-line through actual measurement, it is possible to estimate the purification state in-line. In addition, the estimated purification state can be obtained immediately. This means that, for example, by applying the disclosed technology to the purification process in pharmaceutical manufacturing, it becomes possible to address any abnormalities that occur during the purification process immediately (e.g., within 10 seconds). Also, by applying the disclosed technology at the process development stage when considering purification conditions, it becomes possible to evaluate the validity of the purification conditions in a short time.

[0070] Furthermore, since highly correlated spectral data, which consists of spectral intensity values ​​of wavenumbers with relatively high correlation to the state data of the monitored object, are used as training data from the spectral data output from the analyzer 41, the training load can be reduced compared to the case where all spectral data output from the analyzer 41 is used as training data, and the accuracy of the output values ​​of the soft sensor 20 can be improved.

[0071] If contaminants are present in pharmaceuticals using antibodies produced from cells, even in trace amounts, it can affect the efficacy of the drug. According to the method for estimating the purification state according to the embodiment of the disclosed technology, an estimated value of the concentration of contaminants can be obtained. Therefore, by applying the disclosed technology to the purification process carried out in the manufacturing process of pharmaceuticals, the quality of the pharmaceuticals can be ensured.

[0072] In this embodiment, the use of Raman scattered light spectra as spectral data is illustrated, but the embodiment is not limited to this configuration. For example, the absorption spectrum of infrared light irradiated onto the purified treatment solution can be used as spectral data. Nuclear magnetic resonance spectra can also be used as spectral data. It is preferable to use Raman scattered light spectra as spectral data.

[0073] Furthermore, in this embodiment, we have illustrated a case where spectral data is preprocessed, and a soft sensor is constructed by machine learning using multiple combinations of the preprocessed data and state data obtained through preprocessing as training data. However, if the training load and the decrease in accuracy of the estimation model due to overfitting are not a problem, unprocessed spectral data may be used as training data.

[0074] Furthermore, in this embodiment, a preprocessing step is given as an example of identifying highly correlated spectral data from the spectral data that have a relatively high correlation with the state data, but the invention is not limited to this. For example, a preprocessing step may be performed to exclude spectral intensity values ​​of predetermined wavenumbers from the spectral data acquired by the analyzer 41 from the training data. Alternatively, a preprocessing step may be performed to group the spectral data acquired by the analyzer 41 so that data with similar wavenumbers belong to the same wavenumber group, and to calculate the average value, standard deviation, median, maximum value, minimum value, etc., of the scattered light intensity for each wavenumber group. In this case, the spectral intensity values ​​for each wavenumber group are used as training data. Additionally, a preprocessing step may be performed to reduce the dimensionality of the training data, which is composed of multiple combinations of spectral data showing intensity for each wavenumber or wavelength and state data.

[0075] An example of how to utilize the purification state estimation method according to the embodiment of the disclosed technology is shown below. For example, in affinity chromatography and cation chromatography included in the manufacturing process of antibody products, the amount of a specific component adsorbed on the column is calculated in advance, and a specified amount of the treatment solution is introduced into the column. By applying the method according to this embodiment to the treatment solution obtained by this purification process to estimate the amount of a specific component, if a specific component is not adsorbed on the column and flows out of the column, such a situation can be immediately detected, and it becomes possible to take measures such as reducing the amount of treatment solution introduced into the column or stopping the introduction of the treatment solution into the column.

[0076] Furthermore, in affinity chromatography and cation exchange chromatography, column degradation can cause contaminants to be adsorbed onto the column, resulting in the elution solution containing no contaminants or containing less contaminants than usual. By applying the method according to this embodiment to the elution solution from the column to estimate the amount of contaminants, such a situation can be immediately detected, allowing for appropriate action such as replacing the column.

[0077] Furthermore, in cation exchange chromatography, a gradient of salt concentrations is applied to elute specific components from the column. The amount of a specific component can be estimated by applying the method according to this embodiment to the treatment solution eluted from the column, and the gradient curve can be controlled according to the concentration of the specific component.

[0078] Furthermore, in the above embodiment, the hardware structure of the processing unit that executes various processes such as the identification unit 11, learning unit 12, acquisition unit 13, extraction unit 14, and estimation unit 15 can be any of the following types of processors. As mentioned above, these types of processors include a CPU, which is a general-purpose processor that executes software (programs) and functions as various processing units, as well as a programmable logic device (PLD), such as an FPGA, whose circuit configuration can be changed after manufacturing, and a dedicated electrical circuit, such as an ASIC (Application Specific Integrated Circuit), which has a circuit configuration specifically designed to execute a particular process.

[0079] A single processing unit may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, multiple processing units may be composed of a single processor.

[0080] Examples of configuring multiple processing units with a single processor include, firstly, a configuration where one or more CPUs and software are combined to form a single processor, as exemplified by client and server computers, and this processor functions as multiple processing units. Secondly, a configuration using a processor that realizes the functions of the entire system, including multiple processing units, on a single IC (Integrated Circuit) chip, as exemplified by System on Chip (SoC). Thus, various processing units are configured as hardware structures using one or more of the above-mentioned various processors. Furthermore, more specifically, the hardware structure of these various processors can utilize electrical circuits (circuitry) that combine circuit elements such as semiconductor elements.

[0081] Furthermore, although the above embodiment describes a configuration in which the soft sensor construction program 70 and the estimation program 80 are pre-stored (installed) in the storage unit 103, the invention is not limited to this configuration. The soft sensor construction program 70 and the estimation program 80 may be provided in the form of recording on a recording medium such as a CD-ROM (Compact Disc Read Only Memory), DVD-ROM (Digital Versatile Disc Read Only Memory), and USB (Universal Serial Bus) memory. Alternatively, the soft sensor construction program 70 and the estimation program 80 may be provided in the form of download from an external device via a network.

[0082] Furthermore, the disclosure of Japanese Patent Application No. 2021-057497, filed on March 30, 2021, is incorporated herein by reference in its entirety. In addition, all documents, patent applications, and technical standards described herein are incorporated herein by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

Claims

1. A method for estimating the state of purification, comprising quantifying the components contained in a treated liquid obtained by subjecting a liquid containing a specific protein and impurities other than the protein to a purification treatment, This includes obtaining an estimated value of the concentration of the impurities based on spectral data showing the intensity for each wavenumber or wavelength of electromagnetic waves irradiated onto the processing liquid and affected by the processing liquid, The concentration of the impurities contained in the processing solution is 20 mg / mL or less, and the weight ratio of the impurities to the mixture containing the protein and the impurities is 15% or less. A method for estimating the state of purification.

2. A method for estimating the state of purification, comprising quantifying the components contained in a treated liquid obtained by subjecting a liquid containing a specific protein and impurities other than the protein to a purification treatment, This includes obtaining an estimated concentration of immature sugar chains structurally similar to the protein, based on spectral data indicating the intensity for each wavenumber or wavelength of electromagnetic waves irradiated onto the processing solution and affected by the processing solution. A method for estimating the state of purification.

3. The process further includes obtaining an estimate of the concentration of the protein contained in the processing solution based on the spectral data. The estimation method according to claim 1 or claim 2.

4. The aforementioned protein is produced from cultured cells. The estimation method according to any one of claims 1 to 3.

5. The impurities include DNA from the protein-producing cells, aggregates of the protein, degradation products of the protein, and host cell-derived proteins. The estimation method according to any one of claims 1 to 4.

6. The purification process includes a component separation method using a chromatography apparatus. The estimation method according to any one of claims 1 to 5.

7. The coefficient of determination, which indicates the degree of agreement between the estimated concentration of the aforementioned impurities and the measured value, is 0.9 or higher. The estimation method according to any one of claims 1 to 6.

8. The root mean square error, which indicates the degree of deviation between the estimated concentration of the aforementioned impurities and the measured value, is 1.2 or less. The estimation method according to any one of claims 1 to 7.

9. A soft sensor is constructed that takes spectral data as input and outputs state data by machine learning using multiple combinations of state data indicating the purification state of the liquid containing the protein and the impurities, and spectral data, as training data. The state data output from the soft sensor is obtained by inputting the spectral data acquired for the processed liquid into the soft sensor. This includes, The aforementioned state data includes an estimated value of the concentration of the impurities contained in the processing liquid. The estimation method according to any one of claims 1 to 8.

10. The spectral data is preprocessed, The soft sensor is constructed by machine learning using multiple combinations of the processed data obtained by the above preprocessing and the state data as training data. The estimation method according to claim 9.

11. The aforementioned preprocessing includes a process of selecting the spectral intensity values ​​for each wavenumber or wavelength included in the spectral data to be used as training data. The estimation method according to claim 10.

12. Through the above selection process, the number of intensity values ​​for each wavenumber or wavelength included in the spectral data that will be used as training data will be set to 5 or more and less than 1000. The estimation method according to claim 11.

13. The aforementioned selection is performed using sparse modeling. The estimation method according to claim 11 or claim 12.

14. The preprocessing includes identifying highly correlated spectral data from the spectral data that have a relatively high correlation with the state data as the processed data. The estimation method according to any one of claims 10 to 13.

15. The preprocessing includes baseline correction of the spectral data. The estimation method according to any one of claims 10 to 14.

16. The spectral data is data showing the spectrum of scattered light when light is irradiated onto the liquid containing the protein and the impurities. The estimation method according to any one of claims 9 to 15.

17. The aforementioned state data includes an estimated value of the concentration of the protein contained in the processing solution. The estimation method according to any one of claims 9 to 16.