Information processing device, operation method for information processing device, and operation program for information processing device
The information processing apparatus uses a prediction model to analyze specific spectral regions of Raman scattered light to accurately predict fragment content in biopharmaceutical suspensions, addressing the accuracy issues in existing methods and improving process control.
Patent Information
- Application Number
- PCT/JP2024/039828
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-25
- Filing Date
- 2024-11-08
- Publication Date
- 2025-07-03
AI Technical Summary
Existing methods for predicting the amount of fragments in suspensions containing biological molecules and their fragments, such as proteins, suffer from reduced accuracy due to similarities in molecular structures, leading to inaccurate monitoring of fragment content in biopharmaceutical production processes.
An information processing apparatus and method that utilizes a prediction model to analyze Raman scattered light spectral data, focusing on specific spectral regions attributed to disulfide bonds, aromatic amino acids, and carbon-carbon, carbon-nitrogen, and carbon-oxygen bonds of proteins, to accurately predict fragment content by performing generation and data input processes that account for structural differences between biological molecules and their fragments.
The solution enables precise prediction of fragment content rates in biopharmaceutical suspensions, enhancing process control and ensuring product quality by improving prediction accuracy compared to conventional methods.
Smart Images

Figure JP2024039828_03072025_PF_FP_ABST
Abstract
Description
Information processing device, operating method for information processing device, and operating program for information processing device
[0001] The technology of the present disclosure relates to an information processing device, an operating method for an information processing device, and an operating program for an information processing device.
[0002] For example, biopharmaceutical manufacturing processes are known, using biological molecules such as proteins, including monoclonal antibodies, as active ingredients. In these manufacturing processes, a suspension is often produced in which various components, including the active ingredient, are dispersed in a liquid. Monitoring the amount of a target component in this suspension is important for the success of the ongoing manufacturing process. Examples of target components include protein fragments (hereinafter simply referred to as fragments), and the amount of a target component includes the content rate and concentration of the fragments. Because the presence of fragments affects the efficacy and safety of biopharmaceuticals, monitoring the amount of the target component as a fragment is particularly important.
[0003] The following technology has attracted attention as a technology for predicting the amount of a target component because it reduces concerns about contamination. Specifically, this technology involves measuring electromagnetic waves such as Raman scattered light emitted from a suspension, and inputting the resulting spectral data, such as Raman spectral data, into a prediction model such as multivariate analysis or a machine learning model, thereby predicting the amount of the target component. For example, Japanese Patent No. 7326307 and JP-A-2021-535739 describe a technology in which fragments are used as target components and a prediction model predicts the amount of the fragments.
[0004] Naturally, fragments originate from proteins, and therefore have the same information about their molecular structure, such as amino acid sequences, as proteins. Therefore, in a suspension containing a mixture of proteins and their fragments, if the fragments were used as target components as in Japanese Patent No. 7326307 and Japanese Translation of PCT International Publication No. 2021-535739, there was a risk that the accuracy of predicting the amount of fragments would be low.
[0005] One embodiment of the technology of the present disclosure provides an information processing device, an operating method of the information processing device, and an operating program of the information processing device that can accurately predict the amount of fragments in a suspension containing a mixture of biological molecules and fragments of biological molecules.
[0006] The information processing device of the present disclosure is an information processing device that predicts the amount of fragments based on spectral data obtained by measuring electromagnetic waves emitted from a suspension in which biological molecules and fragments of biological molecules are dispersed as components in a liquid, and is equipped with a processor. The processor uses a prediction model that performs at least one of a generation process and a data input process according to differences that occur in specific spectral regions of the spectral data of the biological molecules and the spectral data of the fragments, obtains target spectral data of a target suspension in which the amount of fragments is unknown, and inputs the target spectral data into the prediction model to predict the amount of fragments.
[0007] The data input process is preferably a process of selectively inputting data in a specific spectral region from the target spectral data into the prediction model.
[0008] The predictive model is generated using calibration data including reference spectral data of a reference suspension having a known amount of fragments, and the generation process is preferably a process in which spectral data of a sample of fragments artificially generated from biological molecules is included in the calibration data as reference spectral data.
[0009] The data input process is preferably a process of selectively inputting data in a specific spectral region of the reference spectral data into the prediction model.
[0010] Preferably, the biological molecule is a protein and the electromagnetic wave is Raman scattered light.
[0011] The specific spectral region is preferably a region attributable to at least one of a disulfide bond, an aromatic amino acid, a carbon-carbon bond, a carbon-nitrogen bond, and a carbon-oxygen bond of a protein.
[0012] The specific spectral region is the wavenumber 470 cm -1 ~550cm-1 region, wave number 860 cm -1 ~900cm -1 and wavenumber 1060 cm -1 ~1120cm -1 It is preferable that the area is at least one of the above areas.
[0013] Preferably, the protein is an antibody.
[0014] The fragment is preferably at least one of a heavy chain, a light chain, a fragment antigen-binding region, and a fragment crystallizable region of an antibody.
[0015] The target suspension is preferably a purified liquid obtained by passing a culture supernatant of antibody-producing cells through a chromatography device.
[0016] The specific spectral region is preferably a region identified by comparing the spectral data of the biological molecule with the spectral data of a standard sample of a fragment artificially generated from the biological molecule.
[0017] Preferably, the predictive model is a machine learning model.
[0018] The concentration of the fragment in the subject suspension is preferably 0.1 g / L to 5.0 g / L.
[0019] The disclosed method for operating an information processing device is a method for predicting the amount of fragments based on spectral data obtained by measuring electromagnetic waves emitted from a suspension in which biological molecules and fragments of biological molecules are dispersed as components in a liquid, and includes using a prediction model in which at least one of a generation process and a data input process is performed in accordance with differences that occur in specific spectral regions of the spectral data of the biological molecules and the spectral data of the fragments, obtaining target spectral data of a target suspension in which the amount of fragments is unknown, and predicting the amount of fragments by inputting the target spectral data into the prediction model.
[0020] The operating program of an information processing device disclosed herein is an operating program of an information processing device that predicts the amount of fragments based on spectral data obtained by measuring electromagnetic waves emitted from a suspension in which biological molecules and fragments of biological molecules are dispersed as components in a liquid, and causes a computer to execute processes including using a prediction model that performs at least one of a generation process and a data input process according to differences that occur in specific spectral regions of the spectral data of the biological molecules and the spectral data of the fragments, obtaining target spectral data of a target suspension in which the amount of fragments is unknown, and predicting the amount of fragments by inputting the target spectral data into the prediction model.
[0021] According to the technology disclosed herein, it is possible to provide an information processing device, an operating method for an information processing device, and an operating program for an information processing device that are capable of accurately predicting the amount of fragments in a suspension containing a mixture of biological molecules and fragments of biological molecules.
[0022] 1 is a diagram illustrating a measurement system, a culture unit, a purification unit, and an information processing device. FIG. 2 is a diagram illustrating the basic structure of an antibody. FIG. 3 is a diagram illustrating excitation light and Raman scattered light. FIG. 4 is a diagram illustrating Raman spectral data. FIG. 5 is a block diagram of a computer constituting an information processing device. FIG. 6 is a block diagram illustrating a processing unit of a CPU of the information processing device. FIG. 7 is a diagram illustrating processing by a learning unit. FIG. 8 is a diagram illustrating processing by a prediction unit. FIG. 9 is a flowchart illustrating the flow from generation of a fragment preparation to measurement of fragment preparation Raman spectral data. FIG. 10 is a diagram illustrating a specific spectral region identified by comparison of antibody Raman spectral data with fragment preparation Raman spectral data. FIG. 11 is a diagram illustrating generation processing in which a plurality of fragment preparation solutions with different contents are generated and the fragment preparation Raman spectral data and fragment contents are stored as learning data. FIG. 12 is a diagram illustrating a data input screen. FIG. 13 is a diagram illustrating a prediction result display screen. FIG. 14 is a flowchart illustrating the flow of generation processing. FIG. 15 is a flowchart illustrating the flow of processing in the learning phase of a prediction model. FIG. 16 is a flowchart illustrating the flow of processing in an information processing device. FIG. 17 is a graph illustrating the temporal changes in fragment contents according to the prediction model in Example 1 and the fragment contents according to offline analysis. FIG. 18 is a table illustrating evaluation results of the prediction accuracy of fragment contents in Examples and Comparative Examples.
[0023] 1 , a measurement system 10 includes a flow cell 11 and a Raman spectrometer 12. The measurement system 10 is incorporated into a purification unit 15 subsequent to a culture unit 14 in a manufacturing system for a drug substance 13 of a biopharmaceutical, for example.
[0024] The culture unit 14 includes a culture tank and a cell-removing filter. The culture tank contains a cell culture medium (medium). Antibody-producing cells are seeded into the cell culture medium and cultured in the cell culture medium. The culture method may be either perfusion culture or fed-batch culture. The antibody-producing cells are established by incorporating antibody genes into host cells such as Chinese hamster ovary cells (CHO cells). The antibody-producing cells produce immunoglobulins, i.e., antibodies 16, during the culture process. Therefore, not only antibody-producing cells but also antibodies 16 are present in the cell culture medium. The antibodies 16 are, for example, monoclonal antibodies and serve as active ingredients in biopharmaceuticals. The antibodies 16 are examples of the "biological molecules," "components," and "proteins" of the technology disclosed herein.
[0025] The cell removal filter uses, for example, tangential flow filtration (TFF) or alternating tangential flow filtration (ATF) to capture antibody-producing cells in the cell culture solution with a filter membrane and remove the antibody-producing cells from the cell culture solution. The cell removal filter also allows antibodies 16 to pass through. Therefore, the cell culture solution output from the culture unit 14 to the purification unit 15 primarily contains antibodies 16. Hereinafter, the cell culture solution output from the culture unit 14 to the purification unit 15 will be referred to as culture supernatant 17. In addition to antibodies 16, culture supernatant 17 also contains cell-derived proteins, cell-derived DNA (deoxyribonucleic acid), fragments 18 of antibody 16, aggregates 19 of antibody 16, viruses, and the like. These cell-derived proteins, cell-derived DNA, fragments 18, aggregates 19, viruses, etc. are also examples of "components" according to the technology of the present disclosure. Furthermore, culture supernatant 17 is an example of a "suspension" according to the technology of the present disclosure.
[0026] 2, the antibody 16 basically has four polypeptide chains, i.e., two identical heavy chains HC (Heavy Chains) and two identical light chains LC (Light Chains). The antibody 16 has a configuration in which the two heavy chains HC and the two light chains LC are bound by disulfide bonds DB (Disulfide Bonds), and has a symmetrical Y-shape.
[0027] The heavy chain HC is composed of a variable domain VH (Variable domain, Heavy Chain) and constant domains CH (Constant domain, Heavy Chain) 1, CH2, and CH3. The light chain LC is composed of a variable domain VL (Variable domain, Light Chain) and a constant domain CL (Constant domain, Light Chain). The constant domains CH1 and CH2 of the heavy chain HC are connected by a hinge region HR (Hinge Region). The variable domains VH and VL are collectively referred to as the variable region. The constant domains CH1 to CH3 and CL are collectively referred to as the constant region.
[0028] The variable domain VH contains three complementarity-determining regions: CDR-H (Complementary Determining Region, Heavy Chain) 1, CDR-H2, and CDR-H3. The variable domain VL also contains three complementarity-determining regions: CDR-L (Complementary Determining Region, Light Chain) 1, CDR-L2, and CDR-L3. These complementarity-determining regions, CDR-H1 to CDR-H3 and CDR-L1 to CDR-L3, are antigen-binding sites and are also called hypervariable regions.
[0029] The region composed of the variable domains VH and VL and the constant domains CH1 and CL is the fragment antigen-binding region FabR (Fragment antigen-binding region). The region composed of the constant domains CH2 and CH3 and a part of the hinge region HR is the fragment crystallizable region FcR (Fragment crystallizable region). The region composed of the variable domains VH and VL is the fragment variable region FvR (Fragment variable region).
[0030] By causing an enzymatic reaction or a reduction reaction, each portion of the antibody 16, such as the heavy chain HC, the light chain LC, the fragment antigen-binding region FabR, and the fragment crystallizable region FcR, is converted into fragments 18 (heavy chain fragment, light chain fragment, Fab fragment, and Fc fragment). Both the enzymatic reaction and the reduction reaction are reactions that fragment the antibody 16 by digesting or cleaving specific disulfide bonds DB. The enzymatic reaction is catalyzed by an enzyme such as papain. The reduction reaction is catalyzed by a reducing agent 91 (see FIG. 9 ) such as tris(2-carboxyethyl)phosphine (TCEP).
[0031] Returning to FIG. 1 , the purification unit 15 sequentially performs component separation processes on the culture supernatant 17 from the culture unit 14 using multiple chromatography devices, thereby gradually removing contaminants and viruses and gradually increasing the purity of the antibody 16 in the culture supernatant 17. The multiple chromatography devices include, for example, an immunoaffinity chromatography device 20, a cation exchange chromatography device 21, and an anion exchange chromatography device 22, in order from the upstream side. The immunoaffinity chromatography device 20, the cation exchange chromatography device 21, and the anion exchange chromatography device 22 are examples of the "chromatography device" according to the technology of the present disclosure. The chromatography device may also include a size exclusion chromatography device, a hydrophobic interaction chromatography device, etc.
[0032] The culture supernatant 17 is introduced into the immunoaffinity chromatography device 20. The immunoaffinity chromatography device 20 extracts the antibody 16 from the culture supernatant 17 using a column in which a ligand such as protein A, which has affinity for the antibody 16, is immobilized on a carrier, thereby producing a first purified solution 23. The first purified solution 23 is subjected to a virus inactivation treatment 24.
[0033] A first purified solution 23 that has been subjected to a virus inactivation treatment 24 is introduced into a cation exchange chromatography device 21. The cation exchange chromatography device 21 extracts antibodies 16 from the first purified solution 23 using a column with a cation exchanger as the stationary phase, thereby producing a second purified solution 25. The cation exchange chromatography device 21 separates and elutes fragments 18, antibodies 16, and aggregates 19 in this order using a method called gradient elution, in which the salt concentration in a buffer solution that washes the column is increased over time.
[0034] A second purified solution 25 is introduced into the anion exchange chromatography device 22. The anion exchange chromatography device 22 extracts the antibody 16 from the second purified solution 25 using a column having an anion exchanger as a stationary phase, thereby producing a third purified solution 26.
[0035] The third purified solution 26 is passed through a filter 27 to remove viruses. The third purified solution 26 is then subjected to a concentration and filtration process using ultrafiltration (UF) and diafiltration (DF) using a filter 28. The drug substance 13 of the biopharmaceutical is obtained from the third purified solution 26 thus treated. A single-pass tangential flow filtration (SPTFF) filter may be provided upstream of the immunoaffinity chromatography device 20.
[0036] Similar to the culture supernatant 17, the first purified solution 23, the second purified solution 25, and the third purified solution 26 contain a mixture of the antibody 16 and its fragment 18. However, because the purity of the antibody 16 is gradually increased by each of the chromatography devices 20 to 22, the antibody 16 accounts for the majority of the mixture, with only a small amount of the fragment 18. For example, the concentration of the fragment 18 in the second purified solution 25 is 0.1 g / L to 5.0 g / L (0.1 g / L or more and 5 g / L or less). The first purified solution 23, the second purified solution 25, and the third purified solution 26 are examples of the "suspension" and "purified solution" according to the technology of the present disclosure.
[0037] The flow cell 11 is installed in a flow path connecting the cation exchange chromatography device 21 and the anion exchange chromatography device 22. A second purified liquid 25 flows through the flow cell 11 at a preset flow rate from the cation exchange chromatography device 21 to the anion exchange chromatography device 22.
[0038] As an example, as shown in FIG. 3 , the Raman spectrometer 12 is an instrument that evaluates a substance M by utilizing the characteristics of Raman scattered light RSL. When excitation light EL is irradiated onto the substance M, the excitation light EL interacts with the substance M, generating Raman scattered light RSL having a different wavelength from that of the excitation light EL. The wavelength difference between the excitation light EL and the Raman scattered light RSL corresponds to the energy of the molecular vibrations of the substance M. Therefore, Raman scattered light RSL with different wavenumbers can be obtained between substances M with different molecular structures. The Raman scattered light RSL is an example of an "electromagnetic wave" according to the technology of the present disclosure. Of the Stokes line and the anti-Stokes line, it is preferable to use the Stokes line for the Raman scattered light RSL.
[0039] Returning to FIG. 1 , the Raman spectrometer 12 is composed of a sensor unit 35 and an analyzer 36. The sensor unit 35 is connected to the flow cell 11. The sensor unit 35 emits excitation light EL from its tip. The excitation light EL is irradiated onto the second purified liquid 25 flowing through the flow cell 11. Raman scattered light RSL is generated by the interaction between this excitation light EL and components such as the antibody 16 and fragment 18 in the second purified liquid 25. The sensor unit 35 receives the Raman scattered light RSL and outputs the received Raman scattered light RSL to the analyzer 36.
[0040] The analyzer 36 resolves the Raman scattered light RSL into wavenumbers and derives the intensity value of the Raman scattered light RSL for each wavenumber, thereby generating Raman spectrum data 37 representing the spectrum of the Raman scattered light RSL, i.e., the Raman spectrum, as shown in Fig. 4 as an example. The Raman spectrum data 37 is data in which the intensity value of the Raman scattered light RSL for each wavenumber is registered. In this example, the Raman spectrum data 37 is -1 ~1400cm -1 The intensity values of the Raman scattered light RSL in the range from 1 cm -1 4 is data derived at intervals of 1000 s. The graph shown below the Raman spectrum data 37 in Fig. 4 is obtained by plotting the intensity values of the Raman spectrum data 37 for each wavenumber and connecting them with a line. The Raman spectrum data 37 is an example of "spectral data" according to the technology of the present disclosure.
[0041] In this way, the measurement system 10 causes the second purified liquid 25 to flow through the flow cell 11. Then, by irradiating the second purified liquid 25 flowing through the flow cell 11 with excitation light EL via the sensor unit 35, the Raman scattered light RSL of components such as the antibody 16 and fragment 18 in the second purified liquid 25 is measured, and Raman spectrum data 37 is obtained. In other words, the second purified liquid 25 is an example of a "target suspension" according to the technology of the present disclosure.
[0042] Returning to FIG. 1 , the Raman spectrometer 12 is connected to an information processing device 40 via a computer network such as a local area network (LAN). The Raman spectrometer 12 transmits measured Raman spectrum data 37 to the information processing device 40. The Raman spectrum data 37 transmitted from the Raman spectrometer 12 to the information processing device 40 is data for predicting the content of fragments 18 in the second purified liquid 25 (hereinafter referred to as the fragment content). The fragment content is the ratio of the amount of fragments 18 to the amount of a suspension containing fragments 18 as a component, in this case, the second purified liquid 25. For example, if the amount of second purified liquid 25 is 100 g and the amount of fragments 18 therein is 5 g, the fragment content is (5 / 100) × 100 = 5%. The fragment content is an example of the "amount of fragments" according to the technology of the present disclosure. In the following description, the Raman spectrum data 37 transmitted from the Raman spectrometer 12 to the information processing device 40 will be referred to as target Raman spectrum data 37T. The target Raman spectrum data 37T is an example of the “target spectrum data” according to the technology of the present disclosure.
[0043] The information processing device 40 is, for example, a desktop personal computer, and includes a display 41 that displays various screens, and input devices 42 such as a keyboard, a mouse, a touch panel, and / or a microphone for voice input. The information processing device 40 is installed, for example, in a pharmaceutical company that develops biopharmaceuticals, or in an organization that is contracted by a pharmaceutical company to develop biopharmaceuticals, i.e., a contract research organization (CRO). The information processing device 40 is operated by a user U who is involved in the development of biopharmaceuticals at the pharmaceutical company or the contract research organization (hereinafter collectively referred to as a pharmaceutical facility).
[0044] 5, the computer constituting the information processing device 40 includes, in addition to the display 41 and input device 42, a storage 50, a memory 51, a CPU (Central Processing Unit) 52, and a communication unit 53. These are interconnected via a bus line 54.
[0045] The storage 50 is a hard disk drive built into a computer constituting the information processing device 40 or connected via a cable or network. Alternatively, the storage 50 is a disk array consisting of multiple hard disk drives. The storage 50 stores control programs such as an operating system, various application programs, and various data associated with these programs. Note that a solid state drive may be used instead of a hard disk drive.
[0046] The memory 51 is a work memory for the CPU 52 to execute processing. The CPU 52 loads programs stored in the storage 50 into the memory 51 and executes processing in accordance with the programs. In this way, the CPU 52 comprehensively controls each part of the computer. The CPU 52 is an example of a "processor" according to the technology of the present disclosure. The memory 51 may be built into the CPU 52. The communication unit 53 controls the transmission of various information to and from external devices such as the Raman spectrometer 12.
[0047] As an example, as shown in FIG. 6 , an operating program 60 is stored in the storage 50. The operating program 60 is an application program for causing a computer to function as the information processing device 40. In other words, the operating program 60 is an example of an "operating program for an information processing device" according to the technology of the present disclosure. In addition to the operating program 60, the storage 50 also stores a prediction model 61 and a training data group 62. The prediction model 61 is a machine learning model that uses algorithms such as a decision tree, a gradient boosting decision tree, a naive Bayes, a random forest, a support vector machine, or a neural network.
[0048] When the operating program 60 is started, the CPU 52 of the computer constituting the information processing device 40 cooperates with the memory 51 and the like to function as an acquisition unit 65, a read / write (hereinafter abbreviated as RW (Read Write)) control unit 66, an instruction receiving unit 67, a learning unit 68, a prediction unit 69, and a display control unit 70.
[0049] The acquisition unit 65 acquires the target Raman spectrum data 37T from the Raman spectrometer 12. The acquisition unit 65 outputs the target Raman spectrum data 37T to the RW control unit 66.
[0050] The RW control unit 66 controls the storage of various data in the storage 50 and the reading of various data stored in the storage 50. The RW control unit 66 stores the target Raman spectrum data 37T from the acquisition unit 65 in the storage 50. Note that while FIG. 6 depicts the storage 50 as if only one piece of target Raman spectrum data 37T is stored therein, in reality, a plurality of pieces of target Raman spectrum data 37T are stored therein.
[0051] The RW control unit 66 reads out the prediction model 61 and the training data group 62 from the storage 50, and outputs the read-out prediction model 61 and the training data group 62 to the training unit 68. The RW control unit 66 also reads out the target Raman spectrum data 37T and the trained prediction model 61 from the storage 50, and outputs the read-out target Raman spectrum data 37T and the prediction model 61 to the prediction unit 69.
[0052] The instruction receiving unit 67 receives various operation instructions input by the user U via the input device 42. The operation instructions include an instruction to input target Raman spectrum data 37T of the second purified liquid 25 for which the fragment content is to be predicted, an instruction to predict the fragment content, etc.
[0053] The learning unit 68 learns the prediction model 61 using a plurality of pieces of learning data 80 (see FIG. 7 ) that make up the learning data group 62. The learning unit 68 outputs the learned prediction model 61 to the RW control unit 66. The RW control unit 66 stores the learned prediction model 61 from the learning unit 68 in the storage 50.
[0054] The prediction unit 69 inputs the target Raman spectrum data 37T into the prediction model 61, causing the prediction model 61 to predict the fragment content rate, and causes the prediction model 61 to output the prediction result 75. The prediction unit 69 outputs the prediction result 75 to the display control unit 70.
[0055] The display control unit 70 controls the display of various screens on the display 41. For example, the display control unit 70 controls the display 41 to display a data input screen 100 (see FIG. 12 ) for inputting target Raman spectrum data 37T of the second purified solution 25 for which the fragment content is to be predicted and for issuing a fragment content prediction instruction, a prediction result display screen 105 (see FIG. 13 ) for displaying the prediction result 75, and the like.
[0056] As an example, as shown in FIG. 7 , the learning unit 68 uses learning data (also referred to as teacher data or training data) 80 to learn the prediction model 61. The learning data 80 is a set of training Raman spectrum data 37L and a correct fragment content 75CA. The training Raman spectrum data 37L is obtained by measuring Raman scattered light RSL emitted from a reference suspension 81 using a Raman spectrometer 12. The reference suspension 81 is a suspension containing at least fragments 18. The reference suspension 81 includes, for example, a second purified liquid 25 that was produced in the past and has a known fragment content. The correct fragment content 75CA is a fragment content calculated based on the amount of fragments 18 in the reference suspension 81. The amount of fragment 18 in reference suspension 81 can be obtained, for example, by introducing reference suspension 81 into a high performance liquid chromatography (hereinafter referred to as HPLC (High Performance Liquid Chromatography)) device and measuring using an ultraviolet absorption mass spectrometry function provided in the HPLC device. Training data 80 is an example of "calibration data" according to the technology of the present disclosure. Training Raman spectrum data 37L is an example of "reference spectrum data" according to the technology of the present disclosure. Note that, as a method for quantifying the amount of fragment 18, the exemplified HPLC and ultraviolet absorption method are preferred, but are not limited thereto. For example, gel electrophoresis such as polyacrylamide gel electrophoresis (SDS-PAGE: sodium dodecyl sulfate-polyacrylamide gel electrophoresis) or capillary electrophoresis may be used. Furthermore, immunoturbidimetry, colorimetry, fluorescence, etc. may also be used.
[0057] The learning unit 68 performs a data input process on the training Raman spectral data 37L. The data input process selectively retains only data (intensity values) in a specific spectral region from the training Raman spectral data 37L, and sets the data as processed training Raman spectral data 37LP. Selectively retaining only data in a specific spectral region means selectively deleting data outside the specific spectral region.
[0058] The specific spectral regions are the following three regions, the first, second, and third regions. The first region is the wavenumber 470 cm -1 ~550cm -1 (470 cm -1 Above, 550cm -1 The second region is the region between wavenumbers 860 cm and 860 cm. -1 ~900cm -1 (860 cm -1 Above, 900cm -1 The third region is the region of 1060 cm -1 ~1120cm -1 (1060cm -1 Above, 1120cm -1 (See below)
[0059] The first region is a region assigned to the disulfide bond DB of antibody 16. The second region is a region assigned to aromatic amino acids of antibody 16. The third region is a region assigned to carbon-carbon bonds, carbon-nitrogen bonds, and carbon-oxygen bonds of antibody 16.
[0060] The learning unit 68 inputs the processed training Raman spectrum data 37LP into the prediction model 61 and causes the prediction model 61 to output a training prediction result 75L. Next, the learning unit 68 compares the training prediction result 75L with the correct fragment content 75CA and, based on the comparison result, performs a loss calculation for the prediction model 61 using a loss function. The learning unit 68 updates the prediction model 61 in accordance with the result of the loss calculation and updates the prediction model 61 in accordance with the update setting.
[0061] The learning unit 68 repeatedly performs the above series of processes, including data input processing, input of the processed training Raman spectrum data 37LP into the prediction model 61, output of the training prediction result 75L from the prediction model 61, loss calculation, update setting, and update of the prediction model 61, while changing the training data 80. The learning unit 68 terminates the repetition of the above series of processes when the prediction accuracy of the training prediction result 75L for the correct fragment content 75CA reaches a predetermined set level. The learning unit 68 outputs the prediction model 61 whose prediction accuracy has reached the set level to the RW control unit 66 as a trained prediction model 61. Note that learning may be terminated when the above series of processes have been repeated a set number of times, regardless of the prediction accuracy of the training prediction result 75L for the correct fragment content 75CA. Furthermore, learning of the prediction model 61 may continue even after storage in the storage 50.
[0062] 8 , the prediction unit 69 performs the same data input process on the target Raman spectrum data 37T as in the learning unit 68, and converts the target Raman spectrum data 37T into processed target Raman spectrum data 37TP. The prediction unit 69 inputs the processed target Raman spectrum data 37TP into the prediction model 61, and causes the prediction model 61 to output a prediction result 75. The prediction result 75 includes a predicted value of the fragment content.
[0063] How the specific spectral region was found will now be described with reference to FIGS.
[0064] First, as shown in Figure 9 as an example, a preparation 90 of fragment 18 (hereinafter referred to as fragment preparation) is artificially generated from antibody 16. Then, fragment preparation Raman spectral data 37S (see Figure 10) which is Raman spectral data 37 of fragment preparation 90 is obtained. Fragment preparation Raman spectral data 37S is an example of "spectral data of a fragment" and "spectral data of a fragment preparation" according to the technology of the present disclosure.
[0065] More specifically, CHO cells are cultured in the culture section 14 to produce the antibody 16, and a culture supernatant 17 is obtained (step ST10). Next, the culture supernatant 17 is passed through an immunoaffinity chromatography device 20 to obtain a first purified solution 23 (step ST11).
[0066] A reducing agent 91 such as TCEP is added to the first purified solution 23, and the solution is left to stand for one hour in an environment of 37° C. to cause a reduction reaction of the antibodies 16 in the first purified solution 23 (step ST12). Then, gel electrophoresis is used to confirm that the antibodies 16 in the first purified solution 23 have been completely converted into fragments 18 (step ST13).
[0067] The reducing agent 91 is preferably, but is not limited to, TCEP as exemplified above. For example, dithiothreitol, 2-mercaptoethanol, tributylphosphine, 3-mercapto-1,2-propanediol, glutathione, cysteine, or the like may be used as the reducing agent 91. Instead of a reduction reaction, the antibody 16 may be converted into the fragments 18 by an enzymatic reaction. Furthermore, the antibody 16 may be converted into the fragments 18 by changing the temperature and / or the hydrogen ion exponent of the first purified solution 23.
[0068] After the reduction reaction, the first purified solution 23 is subjected to ultrafiltration and diafiltration using a filter 28 to remove the reducing agent 91 (step ST14), thereby obtaining a fragment preparation 90. Alternatively, the reducing agent 91 may be removed from the first purified solution 23 using a chromatography device.
[0069] Next, a fragment preparation solution 95 (see FIG. 11 ) is prepared in which only the obtained fragment preparation 90 is dispersed as a component in a liquid. The Raman scattered light RSL emitted from the fragment preparation solution 95 is measured by the Raman spectrometer 12 to obtain fragment preparation Raman spectrum data 37S (step ST15). Ideally, only the fragment preparation 90 is dispersed in the fragment preparation solution 95, but substances other than the fragment preparation 90 may be mixed in the fragment preparation solution 95 in amounts generally acceptable in the technical field to which the technology of the present disclosure pertains and that do not contradict the spirit of the technology of the present disclosure.
[0070] Meanwhile, the first purified solution 23 obtained in step ST11 is subjected to ultrafiltration and diafiltration in step ST14 without undergoing the reduction reaction in step ST12 or the confirmation of fragmentation by gel electrophoresis in step ST13. Thus, antibody 16 is obtained. Next, an antibody solution containing only the obtained antibody 16 dispersed in a liquid is prepared. The Raman scattered light RSL emitted from the antibody solution is measured using a Raman spectrometer 12 to obtain antibody Raman spectral data 37A (see FIG. 10 ), which is Raman spectral data 37 of the antibody 16. The antibody Raman spectral data 37A is an example of the "spectral data of a biological molecule" according to the technology of the present disclosure. While the antibody solution, like the fragment preparation solution 95, ideally contains only the antibody 16 dispersed therein, it may also contain substances other than the antibody 16 in amounts generally acceptable in the technical field to which the technology of the present disclosure pertains, and within the spirit and scope of the technology of the present disclosure.
[0071] As an example, as shown in FIG. 10, when antibody Raman spectrum data 37A and fragment sample Raman spectrum data 37S are compared, the above-mentioned wavenumber 470 cm -1 ~550cm -1 The first region, wavenumber 860 cm -1 ~900cm -1 The second region, and the wave number 1060 cm -1 ~1120cm -1 It was found that there was a clear difference in intensity value with high reproducibility in the third region of the spectrum. Therefore, the first to third regions were defined as specific spectral regions.
[0072] As an example, as shown in FIG. 11 , in the technique of the present disclosure, a user U uses a fragment preparation 90 generated by steps ST10 to ST14 of FIG. 9 to prepare multiple fragment preparation solutions 95 with different fragment contents as a type of reference suspension 81. Then, the Raman scattered light RSL emitted from each of these multiple fragment preparation solutions 95 is measured using a Raman spectrometer 12 to obtain multiple fragment preparation Raman spectral data 37S. A pair of the thus obtained fragment preparation Raman spectral data 37S and the fragment content of each fragment preparation solution 95 is stored as training data 80 in the training data group 62. In other words, the fragment preparation Raman spectral data 37S of a fragment preparation 90 artificially generated from an antibody 16 is included in the training data 80 as training Raman spectral data 37L. The process shown in FIG. 11 is an example of a “generation process” according to the technique of the present disclosure.
[0073] When the instruction receiving unit 67 receives an instruction from the user U to start the operating program 60, the display control unit 70 controls the display 41 to display, as an example, a data input screen 100 shown in FIG. 12 . The data input screen 100 is provided with an input box 101 for target Raman spectral data 37T and a prediction button 102. A file of the target Raman spectral data 37T can be dragged and dropped into the input box 101. Dragging and dropping the file of the target Raman spectral data 37T into the input box 101 is an instruction to input the target Raman spectral data 37T of the second purified liquid 25 for which the fragment content is to be predicted.
[0074] The user U inputs the desired target Raman spectrum data 37T into the input box 101, and then selects the prediction button 102. When the prediction button 102 is selected, the instruction receiving unit 67 receives an instruction to predict the fragment content rate related to the target Raman spectrum data 37T input into the input box 101. As a result, the prediction unit 69 executes the process shown in FIG. 8, and the prediction unit 69 outputs the prediction result 75 to the display control unit 70.
[0075] Upon receiving the prediction result 75 from the prediction unit 69, the display control unit 70 controls the display 41 to display a prediction result display screen 105, as shown in FIG. 13 as an example. The prediction result display screen 105 displays the fragment content, which is the prediction result 75. In addition, a save button 106 and an OK button 107 are provided at the bottom of the prediction result display screen 105. When the save button 106 is selected, the prediction result 75 is associated with the target Raman spectrum data 37T and stored in the storage 50. When the OK button 107 is selected, the display of the prediction result display screen 105 is cleared. Note that, for example, the fragment content may be predicted using the prediction model 61 at multiple time points at regular intervals to obtain the fragment content at the multiple time points, and the fragment content at the multiple time points may be plotted on a graph with time on the horizontal axis, thereby displaying the temporal progression of the fragment content on the prediction result display screen 105 (see FIG. 17 ).
[0076] Next, the operation of the above configuration will be described with reference to the flowcharts shown in Figures 14, 15, and 16. As shown in Figure 6, the CPU 52 of the information processing device 40 functions as an acquisition unit 65, a RW control unit 66, an instruction acceptance unit 67, a learning unit 68, a prediction unit 69, and a display control unit 70 when the operating program 60 is started.
[0077] First, as shown in FIG. 14 , a fragment preparation 90 is generated through the steps ST10 to ST14 shown in FIG. 9 (step ST100). Using this fragment preparation 90, a plurality of fragment preparation solutions 95 with different fragment contents are prepared as shown in FIG. 11 (step ST110). Next, fragment preparation Raman spectral data 37S for each fragment preparation solution 95 is measured (step ST120). Then, pairs of the fragment preparation Raman spectral data 37S and the fragment contents of the fragment preparation solutions 95 are stored as training data 80 in the training data group 62 (step ST130). As a result, the training data group 62 includes not only the training data 80 consisting of pairs of training Raman spectral data 37L and the correct fragment contents 75CA for various reference suspensions 81, but also training data 80 consisting of pairs of the fragment preparation Raman spectral data 37S and the fragment contents of the fragment preparation solutions 95.
[0078] 15 , in the learning phase of prediction model 61, learning unit 68 performs a data input process on training Raman spectral data 37L (step ST200), as shown in Fig. 7. The data input process selectively retains data in a specific spectral region from training Raman spectral data 37L, and sets the data as processed training Raman spectral data 37LP.
[0079] Next, the processed training Raman spectrum data 37LP is input to the prediction model 61, which outputs a training prediction result 75L (step ST210). The training prediction result 75L is compared with the correct fragment content 75CA, and a loss calculation is performed on the prediction model 61 using a loss function based on the comparison result. The prediction model 61 is then updated according to the result of the loss calculation, and the prediction model 61 is updated in accordance with the updated setting (step ST220).
[0080] The series of processes from steps ST200 to ST220 is repeatedly performed while changing the training data 80 (step ST240) until the prediction accuracy of the training prediction result 75L for the correct fragment content rate 75CA reaches a predetermined set level (NO in step ST230). When the prediction accuracy of the training prediction result 75L for the correct fragment content rate 75CA reaches a predetermined set level (YES in step ST230), the series of processes from steps ST200 to ST220 ends. The trained prediction model 61 is output from the training unit 68 to the RW control unit 66 and stored in the storage 50 under the control of the RW control unit 66 (step ST250).
[0081] 16 , the acquiring unit 65 acquires target Raman spectrum data 37T from the Raman spectrometer 12 (step ST300). The target Raman spectrum data 37T is output from the acquiring unit 65 to the RW control unit 66 and stored in the storage 50 under the control of the RW control unit 66 (step ST310).
[0082] 12 is displayed on the display 41 under the control of the display control unit 70 (step ST320). When the user U inputs desired target Raman spectrum data 37T into the input box 101 on the data input screen 100 and selects the prediction button 102, the instruction receiving unit 67 receives an instruction to predict the fragment content of the target Raman spectrum data 37T input into the input box 101 (step ST330). The prediction instruction is output from the instruction receiving unit 67 to the RW control unit 66.
[0083] Under the control of the RW control unit 66, the target Raman spectrum data 37T for which input and prediction instructions have been given is read from the storage 50 (step ST340). The prediction model 61 is also read from the storage 50. The target Raman spectrum data 37T and the prediction model 61 are output from the RW control unit 66 to the prediction unit 69.
[0084] 8, the prediction unit 69 performs a data input process on the target Raman spectrum data 37T (step ST350). The data input process selectively retains data in a specific spectral region from the target Raman spectrum data 37T, and sets the data as processed target Raman spectrum data 37TP.
[0085] Subsequently, the processed target Raman spectrum data 37TP is input to the prediction model 61, which then outputs a prediction result 75 (step ST360). The prediction result 75 is output from the prediction unit 69 to the display control unit 70.
[0086] As shown in FIG. 13, under the control of the display control unit 70, the prediction result display screen 105 is displayed on the display 41 (step ST370), and the fragment content rate of the prediction result 75 is made available for viewing by the user U.
[0087] The user U makes various decisions depending on the prediction result 75 displayed on the prediction result display screen 105. For example, consider a case where the user U is currently testing conditions for culturing antibody-producing cells using small-scale equipment. In this case, if the prediction result 75 is worse than the target value, the user U may decide to stop the current experiment and move on to an experiment using new conditions. Also, consider a case where the testing has been completed and mass production is underway using large-scale equipment. In this case, if the prediction result 75 is worse than the target value, the user U may decide to stop mass production and perform maintenance on the culture tank in the culture unit 14 or the immunoaffinity chromatography device 20 and / or cation exchange chromatography device 21 in the purification unit 15.
[0088] As described above, the information processing device 40 uses a prediction model 61. The prediction model 61 undergoes generation processing and data input processing in response to differences occurring in specific spectral regions between the antibody Raman spectral data 37A and the fragment preparation Raman spectral data 37S. The information processing device 40 also includes an acquisition unit 65 and a prediction unit 69. The acquisition unit 65 acquires target Raman spectral data 37T of the second purified solution 25, in which the amount of fragment 18 is unknown. The prediction unit 69 predicts the fragment content by inputting the target Raman spectral data 37T into the prediction model 61. By performing generation processing and data input processing in the prediction model 61 in response to differences occurring in specific spectral regions between the antibody Raman spectral data 37A and the fragment preparation Raman spectral data 37S, it becomes possible to more accurately predict the fragment content in the second purified solution 25, in which the antibody 16 and the fragment 18 are mixed, compared to when no countermeasures are taken.
[0089] 8, the data input process is a process of selectively inputting data in a specific spectral region from the target Raman spectral data 37T into the prediction model 61. Also, as shown in FIG. 7, the data input process is a process of selectively inputting data in a specific spectral region from the training Raman spectral data 37L into the prediction model 61.
[0090] The specific spectral region is a region where a difference occurs between the antibody Raman spectral data 37A and the fragment preparation Raman spectral data 37S. Therefore, the data in this specific spectral region represents slight structural differences between the antibody 16 and the fragment preparation 90, and ultimately between the antibody 16 and the fragment 18. Therefore, the data in the specific spectral region is highly useful for predicting the fragment content of the second purified solution 25 containing a mixture of the antibody 16 and the fragment 18. In other words, data outside the specific spectral region does not contribute much to improving the accuracy of the prediction of the fragment content of the second purified solution 25 containing a mixture of the antibody 16 and the fragment 18, and may instead become noise. Therefore, by selectively inputting data in the specific spectral region into the prediction model 61 during the data input process, the accuracy of the prediction of the fragment content by the prediction model 61 can be further improved.
[0091] 7, the prediction model 61 is generated using training data 80 including training Raman spectrum data 37L, which is Raman spectrum data 37 of a reference suspension 81 having a known amount of fragments 18. As shown in FIG. 11, the generation process is a process of including, in the training data 80, fragment preparation Raman spectrum data 37S, which is Raman spectrum data 37 of a fragment preparation 90 artificially generated from an antibody 16, as training Raman spectrum data 37L.
[0092] The fragment preparation Raman spectrum data 37S is data that is purely influenced only by the fragment preparation 90, in other words, data that is not influenced at all by the antibody 16. Therefore, by including the fragment preparation Raman spectrum data 37S, from which the influence of the antibody 16 that may become noise in the training of the prediction model 61 has been eliminated, in the training data 80, the training of the prediction model 61 can be further advanced, and as a result, the prediction accuracy of the fragment content by the prediction model 61 can be further improved. Furthermore, since it is possible to freely prepare fragment preparation solutions 95 with any fragment content and obtain a large amount of fragment preparation Raman spectrum data 37S, this can greatly contribute to enriching the training data 80.
[0093] Although the present embodiment shows an example in which both the generation process and the data input process are performed, the present invention is not limited to this. At least one of the generation process and the data input process may be performed.
[0094] Biopharmaceuticals containing antibody 16 are called antibody drugs and are widely used to treat chronic diseases such as cancer, diabetes, and rheumatoid arthritis, as well as rare diseases such as hemophilia and Crohn's disease. Therefore, this example, in which the biological molecule is a protein and the protein is antibody 16, can promote the development of antibody drugs that are widely used to treat various diseases.
[0095] Raman scattered light RSL is likely to reflect information derived from the functional groups of amino acids in proteins. Therefore, by using electromagnetic waves as Raman scattered light RSL as in this example, it is possible to obtain Raman spectrum data 37 that accurately reflects physical properties such as the fragment content of antibody 16, which is a protein.
[0096] As shown in Figures 7, 8, and 10, the specific spectral regions are the region assigned to the protein disulfide bond DB, the region assigned to aromatic amino acids, and the region assigned to carbon-carbon bonds, carbon-nitrogen bonds, and carbon-oxygen bonds. More specifically, the specific spectral regions are those at a wavenumber of 470 cm -1 ~550cm -1 region, wave number 860 cm -1 ~900cm -1 and wavenumber 1060 cm -1 ~1120cm -1 Therefore, it can be said that this is a reasonable specific spectral region when the biological molecule is a protein.
[0097] The specific spectral region may be a region that belongs to at least one of a protein disulfide bond DB, an aromatic amino acid, a carbon-carbon bond, a carbon-nitrogen bond, and a carbon-oxygen bond. -1 ~550cm -1 region, wave number 860 cm -1 ~900cm -1and wavenumber 1060 cm -1 ~1120cm -1 It is sufficient if the area is at least one of the above.
[0098] More preferably, the specific spectral region is a region attributable to a carbon-carbon bond, a carbon-nitrogen bond, and a carbon-oxygen bond. That is, the specific spectral region is a region attributable to a wavenumber of 1060 cm -1 ~1120cm -1 It is more preferable that the range is:
[0099] As shown in Figure 2, fragment 18 is a heavy chain HC, a light chain LC, a fragment antigen-binding region FabR, and a fragment crystallizable region FcR of antibody 16. These are common fragments 18 of antibody 16. Therefore, the content of common fragments 18 can be predicted with high accuracy.
[0100] The fragment 18 may be at least one of the heavy chain HC, the light chain LC, the fragment antigen-binding region FabR, and the fragment crystallizable region FcR of the antibody 16.
[0101] As shown in FIG. 1 , the target suspension is a second purified solution 25 obtained by passing a culture supernatant 17 of CHO cells, which are antibody-producing cells, through an immunoaffinity chromatography device 20 and a cation exchange chromatography device 21. The immunoaffinity chromatography device 20 and the cation exchange chromatography device 21 serve to capture fragments 18. Therefore, by using the second purified solution 25 as the target suspension, it is possible to monitor whether the immunoaffinity chromatography device 20 and the cation exchange chromatography device 21 are fulfilling their roles. If it is determined based on the prediction result 75 that they are not fulfilling their roles, appropriate measures can be taken, such as performing maintenance on the immunoaffinity chromatography device 20 and / or the cation exchange chromatography device 21.
[0102] As shown in Figure 10, the specific spectral region is a region identified by comparing antibody Raman spectral data 37A with fragment preparation Raman spectral data 37S. Previously, it was believed that the information regarding the molecular structures of antibody 16 and fragment 18 remained unchanged, and that there was no difference between antibody Raman spectral data 37A and fragment preparation Raman spectral data 37S. However, when fragment preparation 90 was artificially generated from antibody 16 and antibody Raman spectral data 37A and fragment preparation Raman spectral data 37S were compared in detail, differences in the data were found in the specific spectral region. The technology of the present disclosure utilizes this valuable knowledge of the specific spectral region to actually improve the prediction accuracy of fragment content using prediction model 61.
[0103] As shown in Figure 6, the prediction model 61 is a machine learning model. Machine learning models are generally used to predict unknown parameters, and their prediction accuracy can be improved to a certain level through learning. Therefore, it is possible to easily generate a prediction model 61 with relatively high prediction accuracy.
[0104] 1, the concentration of fragment 18 in second purified solution 25 is 0.1 g / L to 5.0 g / L. Thus, the technique of the present disclosure can accurately predict the fragment content even when the concentration of fragment 18 is relatively low.
[0105] First, in Example 1, a prediction model 61 was prepared that had been trained by performing the generation process shown in Fig. 11 and the data input process shown in Fig. 7 . Then, as shown in Fig. 1 , target Raman spectrum data 37T of the second purified solution 25 was measured using the measurement system 10. The target Raman spectrum data 37T was measured multiple times at regular time intervals from the start of operation of the cation exchange chromatography device 21 (start of production of the second purified solution 25). Then, the data input process shown in Fig. 8 was performed on the target Raman spectrum data 37T, and the prediction model 61 was made to predict the fragment content.
[0106] On the other hand, a fraction collector was used to collect the second purified solution 25 each time the target Raman spectral data 37T was measured. The collected second purified solution 25 was then subjected to offline analysis using an HPLC device or the like to measure the amount of fragment 18 in the collected second purified solution 25, and the fragment content was calculated based on this.
[0107] The fragment content rate obtained by prediction model 61 in Example 1 and the fragment content rate obtained by offline analysis over time are shown in graph 110 in Figure 17. Graph 110 shows that the fragment content rate obtained by prediction model 61 and the fragment content rate obtained by offline analysis over time both show the same trend. This confirms that prediction model 61 successfully predicts the fragment content rate.
[0108] In addition, the coefficient of determination R of the fragment content rate by the prediction model 61 for the fragment content rate by offline analysis 2 The root mean squared error (RMSE) was calculated to evaluate the degree of deviation of the fragment content rate obtained by the prediction model 61 from the fragment content rate obtained by offline analysis, i.e., the prediction accuracy of the fragment content rate obtained by the prediction model 61. The results are shown in Table 115 of FIG. 18. The coefficient of determination R in Example 1 2 The mean square error (RMSE) was 2.92. These figures show that Example 1 has very high prediction accuracy.
[0109] In Example 2, the data input process shown in Figures 7 and 8 was performed, but the generation process shown in Figure 11 was not performed. 2 The mean square error (RMSE) was 0.89 and the root mean square error (RMSE) was 3.21. These figures show that although the prediction accuracy of Example 2 is inferior to that of Example 1, it can still be said that Example 2 has high prediction accuracy.
[0110] [Comparative Example] The comparative example is a case where the data input process shown in Fig. 7 and Fig. 8 and the generation process shown in Fig. 11 are not performed. In other words, the comparative example can be rephrased as a conventional example. The coefficient of determination R 2was 0.83, and the root mean square error RMSE was 4.05.
[0111] Both Examples 1 and 2 had a higher coefficient of determination R than Comparative Example 1. 2 The root mean square error (RMSE) and the mean square error (RMSE) are improved. Therefore, it was found that by at least performing the data input process, it is possible to improve the prediction accuracy of the fragment content by the prediction model 61 compared to the conventional method. This result can be said to support the idea that data in a specific spectral region is certainly effective in improving the prediction accuracy of the fragment content.
[0112] In addition, Example 1 has a higher coefficient of determination R than Example 2. 2 and the root mean square error (RMSE) improved. Therefore, it was found that performing both the data input process and the generation process can improve the prediction accuracy of the fragment content by the prediction model 61 compared to performing only the data input process. From the above, it was confirmed that the technology disclosed herein makes it possible to accurately predict the fragment content in the second purified solution 25 containing a mixture of antibody 16 and fragment 18.
[0113] Although the fragment content is predicted as the amount of fragments 18, this is not limiting. The amount of fragments 18 itself, or the concentration of fragments 18, etc. may also be predicted. The amounts of two or more types of fragments 18, such as the fragment content and concentration, may also be predicted.
[0114] The antibody 16 may be a bispecific antibody, an antibody-drug conjugate, a low molecular weight antibody, a sugar chain modified antibody, or the like.
[0115] The protein is not limited to antibody 16. It may be a peptide, a nucleic acid (DNA, RNA (Ribonucleic Acid)), a lipid, a virus, a viral subunit, a virus-like particle, or the like. It may also be a cytokine (interferon, interleukin, etc.), a hormone (insulin, glucagon, follicle-stimulating hormone, erythropoietin, etc.), a growth factor (IGF (Insulin-Like Growth Factor)-1, bFGF (Basic Fibroblast Growth Factor), etc.), a blood coagulation factor (factor 7, factor 8, factor 9, etc.), an enzyme (lysosomal enzyme, DNA (Deoxyribonucleic Acid) degrading enzyme, etc.), an Fc (Fragment Crystallizable) fusion protein, a receptor, albumin, or a protein vaccine. Therefore, the fragment 18 is not limited to the heavy chain HC, light chain LC, etc. of the exemplified antibody 16.
[0116] The electromagnetic wave is not limited to Raman scattered light RSL, and therefore the spectrum data is not limited to Raman spectrum data 37. It may be infrared absorption spectrum data, near-infrared absorption spectrum data, nuclear magnetic resonance spectrum data, ultraviolet-visible absorption spectroscopy (UV-Vis) spectrum data, or fluorescence spectrum data.
[0117] The predictive model 61 is not limited to a machine learning model. It may also be a model generated by multivariate analysis or statistical analysis. Examples of multivariate analysis and statistical analysis include linear regression, multiple regression, principal component regression, partial least squares regression, logistic regression, Lasso regression, ridge regression, support vector regression, and Gaussian process regression. Among these, principal component regression is more preferable. In such a model generated by multivariate analysis or statistical analysis, determining the coefficients of a regression equation based on at least two sets of training data 80 is an example of "generation using calibration data" according to the technology of the present disclosure. The predictive model 61 may also be rule-based.
[0118] The target suspension is not limited to the illustrated second purified solution 25. It may be the first purified solution 23 or the third purified solution 26. It may also be the cell culture solution before cells are removed by the cell removal filter in the culture unit 14, or the culture supernatant 17 after cells are removed.
[0119] In the above embodiment, so-called inline sensing is exemplified, but the present disclosure is not limited to this and may be applied to offline sensing.
[0120] The information processing device 40 may be a personal computer installed in the pharmaceutical facility as shown in FIG. 1, or may be a server computer installed in a data center independent of the pharmaceutical facility.
[0121] When the information processing device 40 is configured as a server computer, the target Raman spectrum data 37T is transmitted from a personal computer installed in each pharmaceutical facility to the server computer via a network such as the Internet. The server computer distributes various screens, such as the data input screen 100 and the prediction result display screen 105, to the personal computer in the form of screen data for web distribution created using a markup language such as XML (Extensible Markup Language). The personal computer reproduces the screen to be displayed on the web browser based on the screen data and displays it on the display. Note that other data description languages, such as JSON (Javascript (registered trademark) Object Notation), may be used instead of XML.
[0122] The hardware configuration of the computer constituting the information processing device 40 according to the technology of the present disclosure can be modified in various ways. For example, the information processing device 40 can be configured with multiple computers separated as hardware in order to improve processing power and reliability. For example, the functions of the acquisition unit 65, prediction unit 69, and display control unit 70, as well as the function of the learning unit 68, can be distributed and performed by two computers. In this case, the information processing device 40 is configured with two computers.
[0123] In this way, the hardware configuration of the computer of the information processing device 40 can be changed as appropriate depending on the required performance such as processing power, safety, reliability, etc. Furthermore, not only the hardware, but also application programs such as the operating program 60 can be duplicated or stored in a distributed manner across multiple storage devices in order to ensure safety and reliability.
[0124] In the above embodiment, the following various processors can be used as the hardware structure of processing units that perform various processes, such as the acquisition unit 65, the RW control unit 66, the instruction reception unit 67, the learning unit 68, the prediction unit 69, and the display control unit 70. The various processors include the CPU 52, which is a general-purpose processor that executes software (operation program 60) and functions as various processing units, as described above, as well as dedicated electrical circuits that are processors having a circuit configuration specifically designed to perform specific processes, such as a programmable logic device (PLD) that is a processor whose circuit configuration can be changed after manufacture, such as an FPGA (Field Programmable Gate Array), and an ASIC (Application Specific Integrated Circuit).
[0125] A single processing unit may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs and / or a combination of a CPU and an FPGA).Furthermore, multiple processing units may be configured with a single processor.
[0126] Examples of configuring multiple processing units with a single processor include, first, a form in which one processor is configured with a combination of one or more CPUs and software, as typified by computers such as client and server, and this processor functions as multiple processing units. Second, a form in which a processor is used to realize the functions of the entire system including multiple processing units with a single IC (Integrated Circuit) chip, as typified by systems on chips (SoCs). In this way, various processing units are configured using one or more of the above-mentioned various processors as a hardware structure.
[0127] Furthermore, more specifically, the hardware structure of these various processors can be an electric circuit (circuitry) that combines circuit elements such as semiconductor elements.
[0128] From the above description, the technology described in the following supplementary paragraphs can be understood.
[0129] [Supplementary Item 1] An information processing device that predicts the amount of a fragment based on spectral data obtained by measuring electromagnetic waves emitted from a suspension in which biological molecules and fragments of the biological molecules are dispersed as components in a liquid, the information processing device comprising a processor, wherein the processor uses a prediction model that performs at least one of a generation process and a data input process according to a difference occurring in a specific spectral region between the spectral data of the biological molecules and the spectral data of the fragments, acquires target spectral data of a target suspension in which the amount of the fragments is unknown, and predicts the amount of the fragments by inputting the target spectral data into the prediction model. [Supplementary Item 2] The information processing device of Supplementary Item 1, in which the data input process is a process of selectively inputting data in the specific spectral region of the target spectral data into the prediction model. [Supplementary Item 3] The information processing device of Supplementary Item 1 or Supplementary Item 2, in which the prediction model is generated using calibration data including reference spectral data of a reference suspension in which the amount of the fragments is known, and the generation process is a process of including spectral data of a specimen of the fragment artificially generated from the biological molecules in the calibration data as the reference spectral data. [Supplementary Item 4] The information processing device according to Supplementary Item 3, wherein the data input process is a process of selectively inputting data in the specific spectral region of the reference spectral data into the prediction model. [Supplementary Item 5] The information processing device according to any one of Supplementary Item 1 to Supplementary Item 4, wherein the biological molecule is a protein and the electromagnetic wave is Raman scattered light. [Supplementary Item 6] The information processing device according to Supplementary Item 5, wherein the specific spectral region is a region belonging to at least one of a disulfide bond, an aromatic amino acid, a carbon-carbon bond, a carbon-nitrogen bond, and a carbon-oxygen bond of the protein. [Supplementary Item 7] The specific spectral region is a region belonging to a wavenumber of 470 cm -1 ~550cm -1 region, wave number 860 cm -1 ~900cm -1 and wavenumber 1060 cm -1 ~1120cm -1The information processing device according to Supplementary Item 6, wherein the specific spectral region is at least one of the regions above. [Supplementary Item 8] The information processing device according to any one of Supplementary Items 5 to 7, wherein the protein is an antibody. [Supplementary Item 9] The information processing device according to Supplementary Item 8, wherein the fragment is at least one of a heavy chain, a light chain, a fragment antigen-binding region, and a fragment crystallizable region of the antibody. [Supplementary Item 10] The information processing device according to Supplementary Item 8 or Supplementary Item 9, wherein the target suspension is a purified solution obtained by passing a culture supernatant of cells producing the antibody through a chromatography device. [Supplementary Item 11] The information processing device according to any one of Supplementary Items 1 to 10, wherein the specific spectral region is a region determined by comparing spectral data of the biological molecule with spectral data of a preparation of the fragment artificially generated from the biological molecule. [Supplementary Item 12] The information processing device according to any one of Supplementary Items 1 to 11, wherein the prediction model is a machine learning model. [Supplementary Item 13] The information processing device according to any one of Supplementary Items 1 to 12, wherein the concentration of the fragment in the target suspension is 0.1 g / L to 5.0 g / L.
[0130] The technology of the present disclosure can be appropriately combined with the various embodiments and / or various modified examples described above. Furthermore, it is not limited to the above embodiments, and various configurations can be adopted without departing from the spirit of the present disclosure. Furthermore, the technology of the present disclosure extends not only to programs, but also to storage media that non-temporarily store programs, and computer program products that include programs.
[0131] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[0132] In this specification, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed by connecting them with "and / or."
[0133] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
Claims
1. An information processing apparatus that predicts the amount of a fragment based on spectral data obtained by measuring electromagnetic waves emitted from a suspension in which biological molecules and fragments of the biological molecules are dispersed as components, the apparatus comprising a processor, wherein the processor uses a prediction model in which at least one of generation processing and data input processing corresponding to a difference occurring in a specific spectral region of the spectral data of the biological molecules and the spectral data of the fragments is performed, acquires target spectral data of a target suspension in which the amount of the fragment is unknown, and inputs the target spectral data into the prediction model to predict the amount of the fragment.
2. The information processing apparatus according to claim 1, wherein the data input processing is a process of selectively inputting data of the specific spectral region of the target spectral data into the prediction model.
3. The prediction model is generated using calibration data including reference spectral data of a reference suspension in which the amount of the fragment is known, and the generation processing is a process of including spectral data of a standard of the fragment artificially generated from the biological molecule in the calibration data as the reference spectral data.
4. The information processing apparatus according to claim 3, wherein the data input processing is a process of selectively inputting data of the specific spectral region of the reference spectral data into the prediction model.
5. The information processing apparatus according to claim 1, wherein the biological molecule is a protein and the electromagnetic wave is Raman scattered light.
6. The information processing apparatus according to claim 5, wherein the specific spectral region is a region belonging to at least one of disulfide bonds, aromatic amino acids, carbon-carbon bonds, carbon-nitrogen bonds, and carbon-oxygen bonds of the protein.
7. The specific spectral region is a region with a wave number of 470 cm -1 to 550 cm -1 a region with a wave number of 860 cm -1 to 900 cm -1 a region with a wave number of 1060 cm -1 to 1120 cm -1 The information processing apparatus according to claim 6, which is at least any one of the regions.
8. The information processing apparatus according to claim 5, wherein the protein is an antibody.
9. The information processing apparatus according to claim 8, wherein the fragment is at least one of a heavy chain, a light chain, a fragment antigen-binding region, and a fragment crystallizable region of the antibody.
10. The information processing apparatus according to claim 8, wherein the target suspension is a purified solution obtained by passing a culture supernatant of cells producing the antibody through a chromatograph.
11. The information processing apparatus according to claim 1, wherein the specific spectral region is a region identified by comparing the spectral data of the biological molecule with the spectral data of a standard of the fragment artificially generated from the biological molecule.
12. The information processing apparatus according to claim 1, wherein the prediction model is a machine learning model.
13. The information processing apparatus according to claim 1, wherein the concentration of the fragment in the target suspension is 0.1 g / L to 5.0 g / L.
14. A method of operating an information processing apparatus for predicting the amount of a fragment based on spectral data obtained by measuring electromagnetic waves emitted from a suspension in which a biological molecule and a fragment of the biological molecule are dispersed as components, the method comprising: using a prediction model in which at least one of a generation process and a data input process corresponding to a difference occurring in a specific spectral region of the spectral data of the biological molecule and the spectral data of the fragment is performed; obtaining target spectral data of a target suspension in which the amount of the fragment is unknown; and predicting the amount of the fragment by inputting the target spectral data into the prediction model.
15. An operation program for an information processing apparatus for predicting the amount of a fragment based on spectral data obtained by measuring electromagnetic waves emitted from a suspension in which a biological molecule and a fragment of the biological molecule are dispersed as components, the program causing a computer to execute a process comprising: using a prediction model in which at least one of a generation process and a data input process corresponding to a difference occurring in a specific spectral region of the spectral data of the biological molecule and the spectral data of the fragment is performed; obtaining target spectral data of a target suspension in which the amount of the fragment is unknown; and predicting the amount of the fragment by inputting the target spectral data into the prediction model.
Citation Information
Patent Citations
Use of Raman spectroscopy in downstream purification
JP2021535739A
Multivariate Spectral Analysis and Monitoring of Biological Manufacturing
JP7326307B2
Method for estimating culture state, information processing device, and program
WO2021215179A1
Information processing device, method for operating information processing device, operation program for information processing device, method for generating calibrated state prediction model, and calibrated state prediction model
WO2023090015A1