Method for estimating refinement state

The method uses spectral data and machine learning to quantify impurities and proteins in protein purification, addressing the challenge of trace contaminant quantification in biopharmaceuticals, ensuring accurate and timely purification state estimation.

JP7803931B2Active Publication Date: 2026-01-21FUJIFILM CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023510649
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-03-30
Filing Date
2022-02-21
Publication Date
2026-01-21
Estimated Expiration
2042-02-21

AI Technical Summary

Technical Problem

Existing methods for protein purification, such as those described in U.S. Patent Application Publication No. 2020/0062802, fail to accurately quantify trace amounts of contaminants during the purification process, which can affect the efficacy of biopharmaceutical products.

Method used

A method for estimating the purification state using spectral data and machine learning to quantify impurities and proteins in treatment solutions, employing Raman spectroscopy and sparse modeling to select highly correlated spectral data for accurate estimation.

Benefits of technology

Accurately estimates impurity concentrations in treatment solutions with trace amounts, enabling immediate response to abnormalities during pharmaceutical manufacturing processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007803931000004
    Figure 0007803931000004
  • Figure 0007803931000005
    Figure 0007803931000005
  • Figure 0007803931000006
    Figure 0007803931000006
Patent Text Reader

Abstract

This method for estimating a purified state comprises: quantifying components contained in a treated liquid obtained by subjecting a liquid containing a specific protein and contaminants other than protein to a purification treatment. The method for estimating a purified state comprises: obtaining an estimated value of the concentration of contaminants on the basis of spectral data indicating the wave number, or intensity for each wavelength, of electromagnetic waves that are emitted to a treated liquid and are then affected by the treated liquid. The concentration of contaminants contained in the treated liquid is at most 20 mg / mL, and the weight ratio of the contaminants with respect to a mixture containing protein and the contaminants is at most 15%.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The disclosed technology relates to a method for estimating the state of purification when a liquid containing a specific protein is subjected to a purification process. [Background technology]

[0002] The following techniques are known for purifying proteins such as antibodies produced by cells: For example, U.S. Patent Application Publication No. 2020 / 0062802 describes a technique for quantifying purified intermediates of proteins during production using in-line Raman spectroscopy. Summary of the Invention [Problem to be solved by the invention]

[0003] In the manufacture of biopharmaceuticals, proteins such as antibodies, which are biopharmaceutical active ingredients produced from cultured cells, are purified and formulated. During the protein purification process, the purity of the target protein is gradually increased by using multiple different chromatographic techniques, such as cation chromatography, anion chromatography, immunoaffinity chromatography, and gel filtration chromatography. Monitoring the purification status is preferable to verify whether the purification process is being carried out appropriately at each step. It is particularly important to quantify the contaminants separated from the target protein at each step. This is because even trace amounts of contaminants other than the target protein can affect the efficacy of a pharmaceutical product. During the purification process, the purity of the target protein is gradually increased, and the amount of contaminants contained in the treatment solution at each step is extremely small, making it difficult to quantify the contaminants. Patent Document 1, cited above, describes a technique for quantifying purified intermediates of proteins during production, but does not describe the quantification of contaminants.

[0004] The disclosed technology has been made in consideration of the above points, and aims to provide a method for estimating the purification state that can accurately estimate the concentration of impurities even when the treatment solution that has undergone a protein purification process contains only trace amounts of impurities other than proteins. [Means for solving the problem]

[0005] A method for estimating a purification state according to the disclosed technology includes quantifying components contained in a treatment solution obtained by performing a purification process on a liquid containing a specific protein and non-protein impurities, and obtaining an estimated value of the concentration of the impurities based on spectral data indicating the intensity for each wave number or wavelength of electromagnetic waves irradiated onto the treatment solution and acted upon by the treatment solution. The concentration of the impurities contained in the treatment solution may be 20 mg / mL or less, and the weight ratio of the impurities to the mixture containing the protein and the impurities may be 15% or less.

[0006] The method for estimating the purification state according to the disclosed technology is a method for estimating the purification state that includes quantifying components contained in a treatment liquid obtained by performing a purification process on a liquid containing a specific protein and impurities other than the specific protein, and includes obtaining an estimated value for the concentration of immature glycans that are structurally similar to the protein based on spectral data that indicates the intensity for each wave number or wavelength of electromagnetic waves that are irradiated onto the treatment liquid and acted upon by the treatment liquid.

[0007] The method for estimating a purification state according to the disclosed technology may further include obtaining an estimate of the concentration of a protein contained in the treatment solution based on the spectral data. The specific protein may be produced from cultured cells. The impurities may include DNA of cells producing a specific antibody, protein aggregates, protein degradation products, and host cell-derived proteins. The purification process may include a component separation technique using chromatography. The coefficient of determination, which indicates the degree of agreement between the estimated value of the impurity concentration and the actual measured value, may be 0.9 or more. The root mean square error, which indicates the degree of deviation between the estimated value of the impurity concentration and the actual measured value, may be 1.2 or less.

[0008] A method for estimating a purification state according to the disclosed technology may include constructing a software sensor that receives spectral data as input and outputs status data through machine learning using a plurality of combinations of spectral data and status data indicating a purification state of a liquid containing proteins and impurities as learning data, and inputting spectral data acquired about the treatment liquid into the software sensor to obtain status data output from the software sensor. The status data may include an estimated value of the concentration of the impurities contained in the treatment liquid.

[0009] The method for estimating the refinement state according to the disclosed technique may include preprocessing the spectral data and constructing a software sensor by machine learning using a plurality of combinations of the processed data obtained by the preprocessing and the state data as training data. The preprocessing may include selecting spectral intensity values ​​for each wavenumber or wavelength included in the spectral data to be used as training data. By the selection, the intensity values ​​for each wavenumber or wavelength included in the spectral data may be selected. valueThe number of spectral data used as training data may be 5 or more and less than 1000. The selection may be performed by sparse modeling. The preprocessing may include identifying, as processed data, highly correlated spectral data that has a relatively high correlation with the state data from the spectral data. The preprocessing may include baseline correction of the spectral data.

[0010] The spectral data may be data representing a spectrum of scattered light of light irradiated onto a liquid containing proteins and impurities, and the status data may include an estimated value of the concentration of proteins contained in the treatment liquid. [Effects of the Invention]

[0011] According to the disclosed technology, a method for estimating the purification state is provided that can accurately estimate the concentration of impurities even when the treatment solution that has undergone a protein purification process contains only trace amounts of impurities other than proteins. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 shows an example of an antibody purification process according to an embodiment of the disclosed technology. [Figure 2] FIG. 1 is a diagram illustrating an example of a method for estimating a purification state according to an embodiment of the disclosed technique. [Figure 3] FIG. 10 is a diagram illustrating an example of a method for acquiring spectral data. [Figure 4] FIG. 10 is a diagram illustrating an example of learning data according to an embodiment of the disclosed technology. [Figure 5] FIG. 1 is a diagram illustrating an example of a method for estimating a purification state according to an embodiment of the disclosed technique. [Figure 6] FIG. 1 is a diagram illustrating an example of a hardware configuration of an information processing device according to an embodiment of the disclosed technology. [Figure 7] FIG. 10 is a diagram illustrating an example of a structure of an estimation model according to an embodiment of the disclosed technology. [Figure 8]FIG. 10 is an example of a functional block diagram illustrating an example of a functional configuration of an information processing device in a learning phase according to an embodiment of the disclosed technology. [Figure 9] 10 is a flowchart illustrating an example of the flow of a soft-sensor construction process according to an embodiment of the disclosed technology. [Figure 10] FIG. 10 is an example of a functional block diagram illustrating an example of a functional configuration of an information processing device in an operation phase according to an embodiment of the disclosed technology. [Figure 11] 10 is a flowchart illustrating an example of the flow of an estimation process according to an embodiment of the disclosed technology. [Figure 12A] 1 is a graph showing the relationship between estimated and measured impurity concentrations. [Figure 12B] 1 is a graph showing the relationship between estimated and measured impurity concentrations. [Figure 12C] 1 is a graph showing the relationship between estimated and measured impurity concentrations. [Figure 13A] 1 is a graph showing the relationship between estimated and measured antibody concentrations. [Figure 13B] 1 is a graph showing the relationship between estimated and measured antibody concentrations. [Figure 13C] 1 is a graph showing the relationship between estimated and measured antibody concentrations. [Figure 14] 1 is a graph showing the relationship between estimated and measured values ​​of immature glycan concentrations. DETAILED DESCRIPTION OF THE INVENTION

[0013] Hereinafter, an example of an embodiment of the disclosed technology will be described with reference to the drawings. In each drawing, the same or equivalent components and parts are given the same reference numerals, and redundant description will be omitted as appropriate.

[0014] A method for estimating a purification state according to an embodiment of the disclosed technology includes quantifying components contained in a treatment solution obtained by performing a purification process on a liquid containing a specific protein and non-protein impurities. More specifically, the method includes obtaining an estimate of the concentration of the impurities contained in the treatment solution based on spectral data indicating the intensity for each wavenumber or wavelength of electromagnetic waves irradiated onto the treatment solution and acted upon by the treatment solution. The method for estimating a purification state according to the disclosed technology is particularly effective when the concentration of the impurities contained in the treatment solution is 20 mg / mL or less and the weight ratio of the impurities to the mixture containing the protein and the impurities is 15% or less. The method for estimating a purification state according to the disclosed technology may also include obtaining an estimate of the concentration of the specific protein contained in the treatment solution.

[0015] The specific protein may be, for example, an immunoglobulin, i.e., an antibody, produced by cultured cells. Contaminants include, for example, immature glycans structurally similar to antibodies, cellular DNA, antibody aggregates, antibody degradation products, and host cell proteins (HCPs). Immature glycans structurally similar to antibodies are likely to form, for example, when the amount of waste products in the culture medium increases or the oxygen concentration in the culture medium becomes insufficient during the culture period of antibody-producing cells. Antibody degradation products are formed by the degradation of antibodies by degradative enzymes produced during the culture period. Antibody aggregates are likely to form, for example, when the concentration of antibodies produced by the cells becomes excessively high or when stress such as heat is applied. DNA excreted from cells indicates that the cell membrane of the cell has been disrupted, i.e., the cell has become a dead cell. Host cell proteins are proteins derived from host cells that are purified together with antibodies during the antibody purification process. If pharmaceuticals using antibodies produced from cells are contaminated with such contaminants, even trace amounts can affect the efficacy of the drug. Therefore, it is important to quantify the contaminants in the treatment solution obtained from the antibody purification process.

[0016] Fig. 1 is a diagram showing an example of an antibody purification process according to an embodiment of the disclosed technology. As shown in Fig. 1, the antibody purification process includes a purification treatment P1 using immunoaffinity chromatography, a virus inactivation treatment P2, a purification treatment P3 using cation chromatography, a purification treatment P4 using anion chromatography, a virus filtration treatment P5, and a concentration and filtration treatment P6.

[0017] The immunoaffinity chromatography purification process P1 is a process for extracting antibodies using a column with a carrier immobilized with a ligand such as protein A that has affinity for antibodies. The virus inactivation process P2 is a process for inactivating viruses contained in the treated liquid obtained by the purification process P1. The cation chromatography purification process P3 is a process for extracting antibodies using a column with a cation exchanger as the stationary phase. The anion chromatography purification process P4 is a process for extracting antibodies using a column with an anion exchanger as the stationary phase. The virus filtration process P5 is a process for removing viruses contained in the treated liquid obtained by each of the above processes using a filter. The concentration and filtration process P6 is a concentration and filtration process using UF (ultrafiltration) and DF (diafiltration).

[0018] As described above, by carrying out a plurality of steps, including a plurality of different chromatographic component separation techniques, impurities are gradually eliminated and the purity of the antibody is gradually increased. It is preferable to monitor the purification state to verify whether the appropriate treatment is being carried out at each step. The method for estimating the purification state according to an embodiment of the disclosed technique can be used to estimate the purification state for each treatment solution obtained in each of the treatments P1 to P6 shown in FIG. 1. Preferably, the purification state is estimated for each treatment obtained in each of the treatments P1 to P6, and if there is a next step, the estimated purification state can be used to determine the purification conditions for the next treatment. Preferably, the antibody concentration, impurity concentration, and immature glycan concentration in treatment P1 are monitored. concentrationThe purification process can be performed while estimating the purification state. Hereinafter, the details of the method for estimating the purification state according to the embodiment of the disclosed technique will be described.

[0019] A method for estimating a purification state according to an embodiment of the disclosed technology includes constructing a software sensor that receives spectral data as input and outputs status data through machine learning using, as training data, multiple combinations of status data indicating the purification state of a liquid containing a specific protein and impurities to be purified and spectral data indicating the intensity per wavenumber or wavelength of electromagnetic waves irradiated onto a treatment liquid obtained by the purification process and acted upon by the treatment liquid. The method for estimating a purification state according to an embodiment of the disclosed technology includes inputting the spectral data acquired about the treatment liquid obtained by the purification process into the software sensor, thereby obtaining status data output from the software sensor. The status data includes an estimated value of the concentration of impurities contained in the treatment liquid.

[0020] Furthermore, a method for estimating a refinement state according to an embodiment of the disclosed technology includes preprocessing spectral data and constructing the soft sensor through machine learning using multiple combinations of the processed data obtained by the preprocessing and the state data as training data. Examples of preprocessing methods include dimensional reduction techniques such as sparse modeling, PCA (principal component analysis), LSA (SVD) (latent semantic analysis (singular value decomposition)), LDA (linear discriminant analysis), ICA (independent component analysis), and PLS (partial least squares regression). The preprocessing method may include selecting spectral intensity values ​​for each wavenumber or wavelength included in the spectral data to be used as training data. In this case, the remaining spectral intensity values ​​for each wavenumber or wavelength are used as processed data. It is expected that the spectral data will contain a huge number of spectral intensity values ​​for each wavenumber or wavelength. Selecting the data to be used as training data can prevent a decrease in prediction accuracy due to overfitting of model data. The selection of spectral data can be performed, for example, by sparse modeling. That is, the preprocessing performed on the spectral data may include identifying, as processed data, highly correlated spectral data that has a relatively high correlation with the state data by excluding data that has a relatively low correlation with the state data from the spectral data using sparse modeling. Spectral intensity values Among these, the number of data to be used as learning data is preferably selected to be 5 or more and less than 1000 by preprocessing, more preferably 5 or more and 800 or less, and even more preferably 5 or more and 500 or less.

[0021] In this embodiment, sparse modeling refers to selecting explanatory variables (i.e., excluding some explanatory variables) for a regression model in which the spectral intensity values ​​for each wavenumber or wavelength included in the spectral data are used as explanatory variables and the state data are used as the objective variable. As a sparse modeling technique, for example, Lasso regression can be used. Lasso regression is a technique for selecting explanatory variables so as to minimize a cost function calculated by adding a penalty term to the root mean squared error (RMSE). In this embodiment, explanatory variables are selected by excluding low-correlation spectral data from the spectral data, which has a relatively low correlation with the state data. The penalty term may be determined, for example, by cross-validation, such as K-fold cross-validation. In the following description, an example is given in which the preprocessing performed on the spectral data is a process for identifying highly correlated spectral data.

[0022] A liquid containing a specific protein and other contaminants can be produced by known methods, such as culturing cells carrying a gene encoding the specific protein, removing the cells from the resulting culture, and purifying it by chromatography. For example, it can be produced by culturing CHO cells transfected with an IgG1 antibody gene, removing the cells by filtration, and then purifying it by chromatography using Protein A. The ratio of the specific protein to the other contaminants can be changed by changing the purification conditions, such as the pH and temperature, used with Protein A. In this embodiment, an aqueous sodium acetate solution was used as the buffer solution for purification. Phosphate- and acetate-based buffer solutions are mainly used in protein purification. Since the characteristic wavenumbers of these buffer solutions are known, removing these wavenumbers makes it possible to predict the wavenumbers regardless of the buffer solution. The disclosed technology can be applied regardless of the type of protein. The difference between antibody species is the difference in amino acid sequence. This difference in amino acid sequence does not appear in the spectrum, so the technology can be applied regardless of the type of antibody. The disclosed technology can be applied to immature glycans regardless of the type of immature glycan.

[0023] As shown in Fig. 2, a method for estimating a purification state according to an embodiment of the disclosed technique includes a step of inputting highly correlated spectral data, which is spectral data acquired for a treatment solution obtained by one of the multiple treatments illustrated in Fig. 1 performed in an antibody purification process, into the soft sensor 20 as processed data, thereby acquiring status data output from the soft sensor 20. The soft sensor 20 uses software to realize a process of outputting status data based on the input highly correlated spectral data. The soft sensor 20 is implemented in an information processing device 10 (see Figs. 3 and 6), which will be described later.

[0024] In this embodiment, the soft sensor 20 employs an analytical technique based on Raman spectroscopy. Specifically, the spectral data input to the soft sensor 20 is Raman scattered light spectral data. Raman spectroscopy is a spectroscopic method for evaluating a substance using Raman scattered light. When a substance is irradiated with light, the light interacts with the substance, generating Raman scattered light with a wavelength different from that of the incident light. The wavelength difference between the incident light and the Raman scattered light corresponds to the energy of the molecular vibrations of the substance, and therefore Raman scattered light with different wavelengths (wavenumbers) can be obtained between substances with different molecular structures. Furthermore, Raman scattered light can be used to estimate various physical properties, such as stress, temperature, electrical properties, orientation, and crystallinity. Of the Stokes and anti-Stokes lines, the Stokes line is preferably used for Raman scattered light. In this embodiment, Raman spectra were collected using a laser output of 500 mW, a measurement wavelength of 785 nm, and a laser irradiation time of 1 second.

[0025] FIG. 3 is a diagram illustrating an example of a method for acquiring spectral data for a processing liquid 31 obtained by any one of processes P1 to P6 shown in FIG. 1 . The spectral data can be acquired using a known Raman spectroscopic probe 40 and analyzer 41. As shown in FIG. 3 , the tip of the probe 40 is immersed in the processing liquid 31 contained in a container 30. Excitation light emitted from a light emitting unit (not shown) provided at the tip of the probe 40 is irradiated onto the processing liquid 31. Raman scattered light generated by the interaction between the excitation light and the processing liquid 31 is received by a light receiving unit (not shown) provided at the tip of the probe 40. The acquired Raman scattered light is resolved into wavenumbers (reciprocals of wavelengths) by the analyzer 41, and spectral data, which are spectral intensity values ​​for each wavenumber, are generated. Note that the spectral data may also be spectral intensity values ​​for each wavelength. The spectral data is supplied to the information processing device 10.

[0026] The status data output from the soft sensor 20 is data that indicates the purification status and has a correlation with the spectral data. The status data includes an estimated value of the concentration of impurities contained in the processing solution 31. The status data may also include an estimated value of the concentration of antibodies contained in the processing solution 31. These status data are not easy to monitor inline by actual measurement. By using the soft sensor 20, it is possible to acquire the status data inline based on the spectral data, which is relatively easy to monitor inline by actual measurement.

[0027] The soft sensor 20 is constructed by machine learning using multiple combinations of spectral data and status data as training data. FIG. 4 shows an example of the training data 50. The training data 50 is obtained, for example, during the process development stage when purification treatment conditions are considered. The training spectral data is obtained, for example, from a treatment liquid obtained by varying various purification conditions. The purification conditions include, for example, the flow rate when the liquid to be purified is injected into the column, the amount and composition of the buffer used when eluting the antibody from the column, etc.

[0028] The learning state data can be obtained by actual measurement using a conventional sampling method, with the processing liquid from which the learning spectrum data was obtained as the measurement target. For example, when acquiring the concentration of impurities contained in the processing liquid as the learning state data, it can be acquired for each type of impurity using a method such as HPLC (high performance liquid chromatography). The learning data is acquired for each purification condition, and the learning spectrum data and learning state data for each condition are associated with each other.

[0029] Here, the analyzer 41 outputs, as spectral data, for example, spectral intensity values ​​in the wavenumber range from 500 cm-1 to 3000 cm-1 in 1 cm-1 increments. Therefore, the amount of acquired spectral data becomes enormous, and if all of the spectral data is used as learning data, the learning load becomes excessive, and a high-performance processor is required to perform machine learning. Furthermore, the spectral intensity values ​​of the Raman scattered light that make up the spectral data may include spectral intensity values ​​of wavenumbers that have low correlation with the status data of the monitoring target. For example, if a specific wavenumber of the Raman scattered light is used as learning data, number of The spectral intensity values ​​are thought to have a low correlation with the concentration of impurities. If the soft-sensor 20 is constructed by machine learning using, as training data, spectral data including spectral intensity values ​​of wavenumbers that have a low correlation with the condition data of the monitoring target, the accuracy of the output value of the soft-sensor 20 may be reduced.

[0030] Therefore, in this embodiment, as a preprocessing step for the spectral data, spectral intensity values ​​of wavenumbers having a relatively high correlation with the status data of the monitoring target are identified as highly correlated spectral data from among the spectral data output from the analyzer 41. Then, in the learning phase, in which the soft sensor 20 is constructed by machine learning, the soft sensor 20 is constructed by machine learning using multiple combinations of highly correlated spectral data and status data as training data. Meanwhile, in the operation phase, in which the constructed soft sensor 20 is operated to acquire status data for the treatment liquid obtained by the purification process, the status data output from the soft sensor 20 is acquired by inputting, into the soft sensor 20, highly correlated spectral data having a relatively high correlation with the status data of the monitoring target from among the spectral data acquired for the treatment liquid obtained by the purification process, as shown in FIG. 5 . The construction of the soft sensor 20 and the acquisition of status data using the soft sensor 20 are performed by the information processing device 10.

[0031] The color of the treatment liquid from which spectral data is acquired varies depending on the amount of impurities contained therein, the type of antibody, and the type of antibody-producing cell. Furthermore, external environmental factors such as temperature, humidity, and vibration during spectral data acquisition, as well as fluctuations in the output of excitation light irradiated onto the treatment liquid, can cause disturbances to the spectral data. These factors can cause fluctuations in the baseline of the spectral data. Baseline fluctuations can reduce the accuracy of the output values ​​of the soft sensor 20. Therefore, in this embodiment, baseline correction of the spectral data is further performed as preprocessing of the spectral data. Baseline correction refers to removing fluctuations in the baseline of the spectral data due to disturbances. Baseline correction may be performed, for example, by differentiating the spectral waveform. Alternatively, baseline correction may be performed, for example, by removing the baseline determined by polynomial fitting from the spectral waveform.

[0032] 6 is a diagram showing an example of the hardware configuration of information processing device 10. Information processing device 10 includes a CPU (Central Processing Unit) 101, a memory 102 as a temporary storage area, and a non-volatile storage unit 103. Information processing device 10 also includes a display unit 104 such as a liquid crystal display, an input unit 105 such as a keyboard and a mouse, a network I / F (Interface) 106 connected to a network, and an external I / F 107 to which analyzer 41 is connected. CPU 101, memory 102, storage unit 103, display unit 104, input unit 105, network I / F 106, and external I / F 107 are connected to a bus 108.

[0033] The storage unit 103 is realized by a storage medium such as a hard disk drive (HDD), a solid state drive (SSD), or a flash memory. The storage unit 103 stores training data 50, an estimation model 60, a soft-sensor construction program 70, and an estimation program 80. As shown in FIG. 4, the training data 50 is a plurality of combinations of spectrum data and state data.

[0034] 7 is a diagram showing an example of the structure of the estimation model 60. The estimation model 60 is a neural network including an input layer, multiple intermediate layers, and an output layer. The input layer of the estimation model 60 receives input of spectral intensity values ​​for each wavenumber of the Raman scattered light, i.e., spectral data. The output layer of the estimation model 60 outputs state data corresponding to the spectral data input to the input layer.

[0035] In the learning phase, the CPU 101 reads the soft-sensor construction program 70 from the storage unit 103, then loads it into the memory 102, and executes it. In the operation phase, the CPU 101 reads the estimation program 80 from the storage unit 103, then loads it into the memory 102, and executes it. An example of the information processing device 10 is a server computer. The CPU 101 is an example of a processor in the disclosed technology.

[0036] 8 is an example of a functional block diagram showing an example of the functional configuration of the information processing device 10 in the learning phase. In the learning phase, the information processing device 10 is configured to include an identification unit 11 and a learning unit 12. It is assumed that learning data 50 and an estimation model 60 are stored in a storage unit 103.

[0037] The identifying unit 11 performs a regression analysis on the training data 50 using Lasso regression, which is an example of sparse modeling, to identify, as highly correlated spectral data, spectral intensity values ​​of wavenumbers that have a relatively high correlation with the state data, among the spectral data included in the training data 50. Specifically, the identifying unit 11 performs the following processing. to The included spectral data is thinned out by processing spectral intensity values ​​of randomly determined wavenumbers, and a regression model (regression equation) is generated that indicates the relationship between the thinned spectral data and the corresponding state data. The identification unit 11 derives a cost function for the generated regression model by adding a penalty term to the root mean squared error (RMSE). The identification unit 11 repeats each of the above processes a predetermined number of times to generate a regression model for each of multiple spectral data that are thinned out at different wavenumbers, and derives the above cost function for each regression model. The identification unit 11 identifies the minimum number of spectral intensity values ​​that can minimize the above cost function as highly correlated spectral data during the predetermined number of repeated calculations. 。

[0038] The learning unit 12 trains the estimation model 60 by machine learning using, as training data, a combination of the highly correlated spectral data identified by the identifying unit 11 from the learning data 50 and the corresponding state data. This constructs a soft sensor 20 that receives the highly correlated spectral data as an input and outputs the state data.

[0039] The learning unit 12 uses the training data 50 to train the estimation model 60 according to backpropagation, an example of machine learning. Specifically, the learning unit 12 extracts the highly correlated spectral data identified by the identifying unit 11 from the training spectral data included in the training data 50. The learning unit 12 inputs the extracted highly correlated spectral data to the estimation model 60 and acquires state data output from the estimation model 60. The learning unit 12 trains the estimation model 60 so as to minimize the difference between the score indicated by the acquired state data and the score indicated by the training state data included in the training data 50 and corresponding to the highly correlated spectral data. The learning unit 12 performs the process of training the estimation model 60 using a combination of all or part of the highly correlated spectral data and the state data included in the training data 50. Note that, in addition to backpropagation, examples of machine learning techniques include random forest, linear regression, nonlinear regression (SVM: Support Vector Machine, Bayesian regression), and logistic regression, with backpropagation being preferred.

[0040] 9 is a flowchart showing an example of the flow of a soft-sensor construction process performed in the learning phase by the CPU 101 executing the soft-sensor construction program 70. The soft-sensor construction program 70 is executed, for example, when a user inputs an instruction to execute the soft-sensor construction process via the input unit 105.

[0041] In step S1, the identifying unit 11 randomly selects the spectral intensity values ​​of wavenumbers to be excluded from the spectral data included in the training data 50 stored in the storage unit 103. That is, the identifying unit 11 randomly selects the spectral intensity values ​​of wavenumbers to be excluded from the spectral data included in the training data 50 stored in the storage unit 103. -1 Among the spectral intensity values ​​acquired at intervals, a process of thinning out the spectral intensity values ​​for some wavenumbers is performed. The number of wavenumbers to be excluded may be a predetermined number or a randomly determined number. It is preferable that the number of wavenumbers to be excluded is a predetermined number.

[0042] In step S2, the identifying unit 11 generates a regression model (regression equation) that indicates the relationship between the spectral data (i.e., the thinned spectral data) configured by the spectral intensity values ​​of the wavenumbers other than the wavenumbers to be excluded selected in step S1 and the corresponding state data. Specifically, the identifying unit 11 estimates a regression model using a statistical method, with the thinned spectral data as an explanatory variable and the corresponding state data as a response variable. The regression model may be a linear model or a nonlinear model.

[0043] In step S3, the specification unit 11 derives a cost function for the regression model generated in step S2. The cost function is used as an index value indicating the accuracy of the regression model.

[0044] In step S4, the identification unit 11 determines whether the number of repetitions of the processes from step S1 to step S3 has reached a predetermined number. The identification unit 11 repeatedly performs the processes from step S1 to step S3 until the number of repetitions reaches the predetermined number. As a result, a regression model is generated for each of the multiple thinned spectral data having different wavenumbers to be excluded, and a cost function is derived for each of the generated regression models.

[0045] In step S5, the identification unit 11 identifies the thinned spectral data used in generating the regression model that minimizes the cost function as highly correlated spectral data. The spectral data used in generating the regression model that minimizes the cost function is configured with spectral intensity values ​​of wavenumbers that have a relatively high correlation with the status data. In this way, the identification unit 11 identifies, through regression analysis, the spectral data configured with spectral intensity values ​​of wavenumbers that have a relatively high correlation with the status data as highly correlated spectral data.

[0046] In step S6, the learning unit 12 extracts the highly correlated spectral data identified in step S5 from the spectral data included in the training data 50 stored in the storage unit 103, and trains the estimation model 60 by machine learning using multiple combinations of the extracted highly correlated spectral data and corresponding state data as training data. Specifically, the learning unit 12 inputs the highly correlated spectral data identified in step S5 to the estimation model 60, and trains the estimation model 60 so as to minimize the difference between the score indicated by the state data output from the estimation model 60 and the score indicated by the training state data included in the training data 50 that corresponds to the highly correlated spectral data. In this way, the soft sensor 20 is constructed.

[0047] The soft sensor 20 is constructed for each type of status data to be monitored. For example, when the soft sensor 20 is caused to output, as status data, an estimated value of the concentration of impurities contained in a treatment solution obtained by a purification process, a spectral intensity value of a wavenumber having a high correlation with the concentration of the impurities among the spectral data is identified as the highly correlated spectral data. Then, by machine learning using, as training data, multiple combinations of the identified highly correlated spectral data and status data indicating the concentrations of the impurities obtained by actual measurement, the soft sensor 20 is constructed to output an estimated value of the concentration of the impurities based on the highly correlated spectral data. On the other hand, when the soft sensor 20 is caused to output, as status data, an estimated value of the concentration of an antibody, a spectral intensity value of a wavenumber having a high correlation with the concentration of the antibody among the spectral data is identified as the highly correlated spectral data. Then, by machine learning using, as training data, multiple combinations of the identified highly correlated spectral data and status data indicating the concentrations of the antibody obtained by actual measurement, the soft sensor 20 is constructed to output an estimated value of the concentration of the antibody based on the highly correlated spectral data.

[0048] 10 is an example of a functional block diagram showing an example of the functional configuration of the information processing device 10 in the operation phase. In the operation phase, the information processing device 10 is configured to include an acquisition unit 13, an extraction unit 14, and an estimation unit 15. It is assumed that a trained estimation model 60 that functions as a soft sensor 20 is stored in the storage unit 103.

[0049] The method for estimating the purification state according to the embodiment of the disclosed technique is applied to, for example, quantifying components of a treatment solution obtained by a purification process for extracting an antibody. As shown in Fig. 3, spectral data is acquired for a treatment solution 31 contained in a container 30 by a probe 40 and an analyzer 41.

[0050] The acquiring unit 13 acquires the spectral data output from the analyzer 41. The extracting unit 14 extracts, from the spectral data acquired by the acquiring unit 13, highly correlated spectral data identified by the identifying unit 11, i.e., spectral intensity values ​​of wavenumbers that have a relatively high correlation with the state data of the monitoring target.

[0051] The estimation unit 15 reads out the trained estimation model 60 functioning as the soft sensor 20 from the storage unit 103, inputs the highly correlated spectral data extracted by the extraction unit 14 to the estimation model 60, and acquires state data output from the estimation model 60. The estimation unit 15 may perform control to display the acquired state data on the display unit 104. The estimation unit 15 may also store the acquired state data in the storage unit 103.

[0052] 11 is a flowchart showing an example of the flow of estimation processing performed in the operation phase by the CPU 101 executing the estimation program 80. The estimation program 80 is executed when, for example, a user inputs an instruction to execute the estimation processing via the input unit 105.

[0053] In step S11, the acquisition unit 13 acquires the spectral data output from the analyzer 41. In step S12, the extraction unit 14 extracts, from the spectral data acquired by the acquisition unit 13, highly correlated spectral data identified by the identification unit 11, i.e., spectral intensity values ​​of wavenumbers having a relatively high correlation with the status data of the monitoring target. In step S13, the estimation unit 15 reads out the trained estimation model 60 that functions as the soft sensor 20 from the storage unit 103, inputs the highly correlated spectral data extracted in step S12 into the read estimation model 60, and acquires the status data output from the estimation model 60. The estimation unit 15 controls the display unit 104 to display the acquired status data.

[0054] Estimated values ​​of impurity concentrations contained in a treated solution obtained by performing a purification process on a liquid containing antibodies and impurities were obtained using soft-sensor 20. FIGS. 12A to 12C are graphs showing the relationship between estimated impurity concentrations obtained using soft-sensor 20 and actual measured impurity concentrations obtained by sampling. As a comparative example, FIGS. 12A to 12C also show the relationship between estimated impurity concentrations obtained by analyzing spectral data using PLS, a multivariate analysis technique, and the actual measured impurity concentrations. In FIGS. 12A to 12C, a treated solution obtained by process P1 was used as the liquid containing antibodies and impurities. In FIGS. 12A to 12C, the case where estimated impurity concentrations were obtained using soft-sensor 20 (Example) is indicated by open diamond plots and a solid line, and the case where estimated impurity concentrations were obtained using PLS (Comparative Example) is indicated by filled square plots and a dotted line.

[0055] The proteins, contaminants, and proteins with immature glycans contained in a liquid containing proteins and non-protein contaminants can be measured by known methods. For example, proteins can be measured by subjecting the liquid to protein A chromatography. Contaminants can be measured by size exclusion chromatography. Immature glycans can be measured by subjecting the liquid to a glycan release treatment, fluorescently labeling the released glycans, removing unreacted material, and then measuring the concentration of the immature glycans by HPLC.

[0056] Figure 12A shows the case where the impurity ratio is 2.5%, Figure 12B shows the case where the impurity ratio is 5%, and Figure 12C shows the case where the impurity ratio is 10%. The impurity ratio is the weight ratio of impurities to a mixture containing antibodies and impurities, and is defined by the following formula (1). In formula (1), R C is the impurity ratio, A is the weight of the antibody contained in the treatment solution, and C is the weight of the impurities contained in the treatment solution. R C =C / (A+C) (1)

[0057] The coefficient of determination (R 2 The root mean squared error (RMSE), which indicates the degree of deviation from the predicted value and the actual measured value, was calculated and the results are shown in Table 1 below.

[0058] [Table 1]

[0059] 12A to 12C and Table 1, it was confirmed that the estimated impurity concentration values ​​obtained using the soft-sensor 20 were more accurate than the estimated values ​​obtained using PLS. In particular, by using the soft-sensor 20, the impurity concentration could be estimated with extremely high accuracy even when the impurity concentration was 5 mg / mL or less and the impurity ratio was 2.5%.

[0060] Estimated values ​​of antibody concentrations contained in a treated solution obtained by performing a purification process on a liquid containing antibodies and impurities were obtained using soft-sensor 20. FIGS. 13A to 13C are graphs showing the relationship between estimated antibody concentrations obtained using soft-sensor 20 and actual measured antibody concentrations obtained by sampling. As a comparative example, FIGS. 13A to 13C also show the relationship between estimated antibody concentrations obtained by analyzing spectral data using PLS, a multivariate analysis technique, and the actual measured antibody concentrations. In FIGS. 13A to 13C, a treated solution obtained by process P1 was used as the liquid containing antibodies and impurities. In FIGS. 13A to 13C, the case where estimated antibody concentrations were obtained using soft-sensor 20 (Example) is indicated by open diamond plots and a solid line, and the case where estimated antibody concentrations were obtained using PLS (Comparative Example) is indicated by filled square plots and a dotted line.

[0061] Figure 13A shows the case where the antibody ratio is 20%, Figure 13B shows the case where the antibody ratio is 50%, and Figure 13C shows the case where the antibody ratio is 80%. body ratio The ratio is the weight ratio of antibody to a mixture containing antibody and impurities, and is defined by the following formula (2): In formula (2), RA is the antibody ratio, A is the weight of antibody contained in the treatment solution, and C is the weight of impurities contained in the treatment solution. RA=A / (A+C) (2)

[0062] The coefficient of determination (R 2 The results of calculating the root mean square error (RMSE), which indicates the degree of deviation from the predicted value and the actual measured value, are shown in Table 2 below.

[0063] [Table 2]

[0064] 13A to 13C and Table 2, it was confirmed that the estimated antibody concentration values ​​obtained using the soft sensor 20 were more accurate than the estimated values ​​obtained using PLS. In particular, by using the soft sensor 20, the antibody concentration could be estimated with extremely high accuracy even when the antibody concentration was 5 mg / mL or less and the antibody ratio was 20%.

[0065] An estimated value of the concentration of immature glycans, which is a type of impurity contained in a treatment liquid obtained by performing a purification process on a liquid containing antibodies and impurities, was obtained using the soft sensor 20. Figure 14 is a graph showing the relationship between the estimated value of immature glycan concentration obtained using the soft sensor 20 and the actual measured value of immature glycan concentration obtained by sampling (Example 1). Figure 14 also shows the relationship between the estimated value of immature glycan concentration obtained by analyzing spectral data using PLS, a multivariate analysis method, and the actual measured value (Example 2). In Figure 14, the case where the estimated value of immature glycan concentration was obtained using the soft sensor 20 (Example 1) is shown by the open diamond plot and solid line, and the case where the estimated value of antibody concentration was obtained using PLS (Example 2) is shown by the filled square plot and dotted line.

[0066] The coefficient of determination (R 2 The results of calculating the root mean square error (RMSE), which indicates the degree of deviation from the predicted value and the actual measured value, are shown in Table 3 below. [Table 3]

[0067] Figure 1 4th grade As shown in Table 3, it was confirmed that the estimated values ​​of immature glycan concentrations obtained using the soft sensor 20 were more accurate than the estimated values ​​obtained using PLS.

[0068] As described above, the method for estimating the purification state according to the embodiment of the disclosed technique allows for accurate estimation of the concentration of impurities even when a treatment solution that has been subjected to a purification process for extracting a specific protein contains trace amounts of impurities other than the protein. For example, the method allows for accurate estimation of the concentration of impurities even when the concentration of the impurities in the treatment solution is 20 mg / mL or less and the impurity ratio is 15% or less.

[0069] Furthermore, because spectral data is relatively easy to monitor inline through actual measurements, it is possible to estimate the purification state inline. Furthermore, the estimated purification state results can be obtained immediately. Therefore, by applying the disclosed technology to purification processes in pharmaceutical manufacturing, for example, it becomes possible to respond immediately (for example, within 10 seconds) if any abnormality occurs during the purification process. Furthermore, by applying the disclosed technology in the process development stage, where purification process conditions are considered, it becomes possible to evaluate the validity of the purification conditions in a short time.

[0070] Furthermore, among the spectral data output from the analyzer 41, highly correlated spectral data composed of spectral intensity values ​​of wavenumbers that have a relatively high correlation with the state data of the monitoring target is used as learning data. Therefore, compared to the case where all the spectral data output from the analyzer 41 is used as learning data, the learning load can be reduced and the accuracy of the output value of the soft sensor 20 can be improved.

[0071] When a pharmaceutical product using an antibody produced from a cell is contaminated with impurities, even a small amount can affect the efficacy of the product. According to a method for estimating the purification state of an embodiment of the disclosed technology, it is possible to obtain an estimated value of the concentration of impurities, and therefore, by applying the disclosed technology to the purification process carried out in the pharmaceutical manufacturing process, the quality of the pharmaceutical product can be ensured.

[0072] In this embodiment, the spectrum of Raman scattered light is used as the spectral data, but the present invention is not limited to this. For example, the absorption spectrum of infrared light irradiated onto a treatment liquid that has been subjected to a purification treatment can also be used as the spectral data. Alternatively, a nuclear magnetic resonance spectrum can also be used as the spectral data. It is preferable to use the spectrum of Raman scattered light as the spectral data.

[0073] In addition, in this embodiment, an example has been given of a case in which spectral data is preprocessed and a soft sensor is constructed by machine learning using multiple combinations of the processed data obtained by the preprocessing and status data as training data. However, if the learning load and the reduction in accuracy of the estimation model due to over-learning are not a problem, spectral data that has not been preprocessed may also be used as training data.

[0074] In addition, in the present embodiment, the preprocessing step is exemplified by identifying highly correlated spectral data from the spectral data that has a relatively high correlation with the status data. However, the preprocessing step is not limited to this. For example, the preprocessing step may be to exclude spectral intensity values ​​of a predetermined wavenumber from the spectral data acquired by the analyzer 41 from the training data. The preprocessing step may also be to group the spectral data acquired by the analyzer 41 so that data with similar wavenumbers belong to the same wavenumber group, and calculate the average, standard deviation, median, maximum, minimum, and other values ​​of the scattered light intensity for each wavenumber group. In this case, the spectral intensity values ​​for each wavenumber group are used as the training data. The preprocessing step may also be to reduce the number of dimensions of training data composed of multiple combinations of spectral data indicating intensity for each wavenumber or wavelength and status data.

[0075] An example of application of the method for estimating the purification state according to an embodiment of the disclosed technology is shown below. For example, in affinity chromatography and cation chromatography, which are included in the manufacturing process of antibody products, the amount of a specific component adsorbed to a column is calculated in advance, and a specified amount of the liquid to be treated is then introduced into the column. By applying the method according to this embodiment to the treated liquid obtained by this purification process to estimate the amount of the specific component, if the specific component is not adsorbed to the column and flows out of the column, such a situation can be immediately detected, and appropriate measures can be taken, such as reducing the amount of the liquid to be treated introduced into the column or stopping the introduction of the liquid to the column.

[0076] Furthermore, in affinity chromatography and cation exchange chromatography, a decrease in column performance can cause impurities to be adsorbed onto the column, resulting in the treatment solution eluted from the column containing no impurities or containing less than normal amounts of impurities. By applying the method according to the present embodiment to the treatment solution eluted from the column to estimate the amount of impurities, it becomes possible to immediately detect such a situation, and appropriate measures such as replacing the column can be taken.

[0077] In cation exchange chromatography, a specific component is eluted from a column using a salt gradient. The amount of the specific component in the treated solution eluted from the column may be estimated by applying the method according to this embodiment, and the gradient curve may be controlled according to the concentration of the specific component.

[0078] Furthermore, in the above embodiment, the following various processors can be used as the hardware structure of processing units that perform various processes, such as the identification unit 11, learning unit 12, acquisition unit 13, extraction unit 14, and estimation unit 15. As described above, the various processors include a CPU, which is a general-purpose processor that executes software (programs) and functions as various processing units, as well as dedicated electrical circuits that are processors having a circuit configuration specifically designed to perform specific processes, such as a programmable logic device (PLD) or an application specific integrated circuit (ASIC), whose circuit configuration can be changed after manufacture, such as an FPGA.

[0079] A single processing unit may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, multiple processing units may be configured with a single processor.

[0080] Examples of configuring multiple processing units with a single processor include, first, a form in which one processor is configured with a combination of one or more CPUs and software, as typified by computers such as client and server, and this processor functions as multiple processing units. Second, a form in which a processor is used to realize the functions of an entire system including multiple processing units with a single IC (Integrated Circuit) chip, as typified by systems on chips (SoCs). In this way, the various processing units are configured as a hardware structure using one or more of the above-mentioned various processors. Furthermore, more specifically, the hardware structure of these various processors can be an electric circuit combining circuit elements such as semiconductor elements.

[0081] In the above embodiment, the soft-sensor construction program 70 and the estimation program 80 are pre-stored (installed) in the storage unit 103, but the present invention is not limited to this. The soft-sensor construction program 70 and the estimation program 80 may be provided in a form recorded on a recording medium such as a CD-ROM (Compact Disc Read Only Memory), a DVD-ROM (Digital Versatile Disc Read Only Memory), or a USB (Universal Serial Bus) memory. The soft-sensor construction program 70 and the estimation program 80 may also be downloaded from an external device via a network.

[0082] The disclosure of Japanese Patent Application No. 2021-057497, filed on March 30, 2021, is incorporated herein by reference in its entirety. In addition, all documents, patent applications, and technical standards described herein are incorporated herein by reference to the same extent as if each individual document, patent application, and technical standard was specifically and individually indicated to be incorporated by reference.

Claims

1. A method for estimating a purification state, comprising quantifying components contained in a treatment solution obtained by subjecting a liquid containing a specific protein and impurities other than the protein to a purification treatment, obtaining an estimated value of the concentration of the impurities based on spectrum data indicating the intensity for each wave number or wavelength of electromagnetic waves that are irradiated to the treatment liquid and acted upon by the treatment liquid; the concentration of the impurities contained in the treatment solution is 20 mg / mL or less, and the weight ratio of the impurities to the mixture containing the protein and the impurities is 15% or less; performing preprocessing to select spectral intensity values ​​for each wavenumber or wavelength included in the spectral data to be used as learning data; constructing a soft sensor by machine learning using a plurality of combinations of the processed data obtained by the preprocessing and state data indicating a purification state of the liquid containing the protein and the impurities as the learning data; The processed data is input to the learned soft sensor to obtain the status data output from the soft sensor. This includes: The preprocessing is a process including identifying, from the spectral data, highly correlated spectral data having a relatively high correlation with the state data as the processed data, and includes a process of thinning out spectral intensity values ​​at randomly determined wavenumbers from the spectral data, generating a regression model showing a relationship between the thinned spectral data and the corresponding state data, and repeatedly performing a process of deriving a cost function for the generated regression model, thereby generating a regression model for each of a plurality of spectral data having different thinned wavenumbers, deriving the cost function for each regression model, and identifying, as the highly correlated spectral data, a spectral intensity value that can minimize the cost function in the repeated calculations. The condition data includes an estimated value of the concentration of the impurities contained in the processing liquid. Methods for estimating refinement state.

2. and obtaining an estimate of the concentration of the protein contained in the treatment solution based on the spectral data. The estimation method according to claim 1 .

3. The protein is produced from cultured cells. The estimation method according to claim 1 or 2.

4. The contaminants include DNA of the cells that produce the protein, aggregates of the protein, degradation products of the protein, and proteins derived from the host cell. The estimation method according to any one of claims 1 to 3.

5. The purification process includes a method for separating components by chromatography. The estimation method according to any one of claims 1 to 4.

6. The coefficient of determination, which indicates the degree of agreement between the estimated value of the concentration of the impurity and the actually measured value, is 0.9 or more. The estimation method according to any one of claims 1 to 5.

7. The root mean square error, which indicates the degree of deviation of the estimated value of the concentration of the impurity from the actually measured value, is 1.2 or less. The estimation method according to any one of claims 1 to 6.

8. By the selection, the number of wave numbers or wavelength intensities included in the spectrum data to be used as the learning data is set to 5 or more and less than 1000. The estimation method according to any one of claims 1 to 7.

9. The pre-processing includes baseline correction of the spectral data. The estimation method according to any one of claims 1 to 8.

10. The spectral data is data indicating a spectrum of scattered light of light irradiated onto the liquid containing the protein and the impurities. The estimation method according to any one of claims 1 to 9.

11. The condition data includes an estimated concentration of the protein contained in the treatment solution. The estimation method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Wavelength selection method, manufacturing method of object substance content estimation device, and object substance content estimation device

    JP2018013418A

  • Sample information acquisition system, data display system including the same, sample information acquisition method, program, and storage medium

    JP2019219419A

  • Real-time monitoring of pharmaceutical purification

    JP2019522802A

  • Information processing device, method for controlling information processing device, and program

    JP2021009135A

  • Method of characterization of visible and / or sub-visible particles in biologics

    WO2020160161A1