Presumption method
Patent Information
- Application Number
- JP2025506494
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2025-09-04
- Publication Date
- 2025-11-18
AI Technical Summary
Conventional methods for estimating retention time in chromatography are cumbersome and lack accuracy, requiring complex calculations and failing to provide results suitable for practical use.
A method utilizing machine learning to estimate retention time by processing Hansen Solubility Parameters (HSP) of a target component, stationary phase, and mobile phase, with SMILES notation input into an estimation model to improve accuracy and convenience.
The method significantly enhances the accuracy and convenience of retention time estimation in chromatography, providing reliable results through machine learning processing of HSP and other variables.
Abstract
Description
Estimation method, computer program, and information processing device
[0001] The present invention relates to estimating the retention time of a component of interest in a chromatograph.
[0002] Various studies have been conducted on the estimation of parameters in chromatography. For example, Japanese Patent No. 6110380 (Patent Document 1) discloses a technique for predicting a retention index related to the retention time of a component eluted in chromatography (hereinafter also referred to as a "target component") by using the Kovats retention index (KRI) of the component.
[0003] Japanese Patent No. 6796842 (Patent Document 2) discloses a technique for calculating the Hansen Solubility Parameters (HSP) of a substance from its structural formula. Attempts have also been made to estimate retention times from HSP distances, and software (HSPiP) for estimating HSPs for unknown compounds is commercially available.
[0004] Hiroshi Yamamoto, "YMB Simulator Physical Property Estimation Function," [online], June 1, 2011, searched on December 13, 2022, Internet <URL: https: / / www.pirika.com / Chemistry / JP / TCPE / YMB.html> (Non-Patent Document 1) discloses that HSPiP has a built-in algorithm for inferring physical properties using the YMB method, and can calculate various physical quantities of unknown compounds based on the type and number of functional groups, etc.
[0005] Yamamoto Hiroshi, "High-Performance Liquid Chromatography and Hansen Solubility Parameters (HSP)", [online], September 9, 2009, searched on December 13, 2022, Internet <URL: https: / / www.pirika.com / HSP / JP / Examples / Docs / HPLC.html> (Non-Patent Document 2) discloses that, for software such as that disclosed in Patent Document 2, various physical quantities of the target component are calculated using the SMILES notation of the target component, the HSP is calculated using the physical quantities, and the calculated HSP value is corrected using the molecular volume, thereby increasing the correlation between the estimated retention time and the actual value.
[0006] Patent No. 6110380 Patent No. 6796842
[0007] Hiroshi Yamamoto, "YMB Simulator Physical Property Prediction Function," [online], June 1, 2011, retrieved December 13, 2022, Internet <URL: https: / / www.pirika.com / Chemistry / JP / TCPE / YMB.html> Hiroshi Yamamoto, "High-Performance Liquid Chromatography and Hansen Solubility Parameters (HSP)," [online], September 9, 2009, retrieved December 13, 2022, Internet <URL: https: / / www.pirika.com / HSP / JP / Examples / Docs / HPLC.html>
[0008] However, conventional techniques have been considered to lack convenience. That is, the estimation of HSPs by the method described in Non-Patent Document 1 requires complex calculations because it requires the calculation of various physical quantities of the target components. Furthermore, the retention time prediction method described in Non-Patent Document 2 has low accuracy and has not yet reached a level that can be put into practical use.
[0009] The present invention has been devised in view of the above circumstances, and its purpose is to provide a technique that improves the convenience of estimating the retention time of a target component in chromatography.
[0010] According to one aspect of the present disclosure, there is provided a method for estimating the retention time of a specific target component in chromatography using a specific stationary phase and a specific mobile phase, the method comprising the steps of: acquiring Hansen Solubility Parameters (HSPs) of the specific target component; acquiring the HSPs of the specific stationary phase and the specific mobile phase; and inputting the HSPs of the specific target component, the specific stationary phase, and the specific mobile phase into a machine learning model to acquire an estimated result of an index related to the retention time of the specific target component, wherein the machine learning model has been subjected to machine learning processing using the HSPs of the target component, the stationary phase, and the mobile phase as explanatory variables and the index related to the retention time of the target component as a response variable.
[0011] According to another aspect of the present disclosure, there is provided a method for estimating Hansen Solubility Parameters (HSP) of a substance, the method comprising the steps of acquiring a SMILES notation of the substance and acquiring an estimated result of the HSP of the substance by inputting the SMILES notation of the substance into an estimation model, wherein the estimation model has been subjected to machine learning processing with the SMILES notation as an explanatory variable and the HSP as a target variable.
[0012] According to yet another aspect of the present disclosure, there is provided a computer program that, when executed by a control circuit of a computer, causes the computer to perform the estimation method described above.
[0013] According to yet another aspect of the present disclosure, there is provided an information processing device comprising a control circuit and a memory device storing a computer program executed by the control circuit, wherein the computer program is executed by the control circuit to cause the information processing device to implement the above-mentioned estimation method.
[0014] According to one aspect of the present disclosure, a technique is provided that improves the accuracy of estimating the retention time of a target component in chromatography or that improves the convenience of estimating retention time by easily calculating HSPs.
[0015] 11. A diagram for explaining an overview of retention time estimation according to an embodiment of the present disclosure. A diagram for explaining an example of a hardware configuration of an information processing device 1. A diagram for explaining an example of a configuration of an HSP estimation model 2. A flowchart of an example of processing performed in the information processing device 1. A flowchart of a subroutine of step S20 of FIG. 4. A flowchart of a subroutine of step S40 of FIG. 4. A diagram for explaining a specific example of SMILES notation used in the machine learning processing of the HSP estimation model 2. A diagram for explaining a result of estimation by the HSP estimation model 2. A diagram for explaining a result of estimation by the retention time estimation model 3. A diagram showing an estimation result by the retention time estimation model 3 together with estimation results by other methods. A flowchart of another example of processing performed in the information processing device 1. A flowchart of a subroutine of step S60 of FIG. 11 (estimation of retention time in gradient elution). A diagram showing a specific example of output of retention time estimation results. A flowchart of a subroutine of step S80 of FIG. 11 (report output). A diagram showing an example of an output report. A flowchart of yet another example of processing performed in the information processing device 1.
[0016] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the drawings, the same or corresponding parts are designated by the same reference numerals, and description thereof will not be repeated.
[0017] [Outline of Retention Time Estimation] In this embodiment, a numerical value relating to the retention time of a substance when the substance is eluted in a certain stationary phase using a certain mobile phase in chromatography is estimated. In this specification, the term "target component" may be used to refer to a substance that is a target to be eluted in chromatography.
[0018] 1 is a diagram for explaining an overview of retention time estimation according to one embodiment of the present disclosure, showing an information processing device 1. The information processing device 1 estimates the retention time of a specific target component in a column that uses a specific stationary phase and a specific mobile phase.
[0019] More specifically, the SMILES notation of a specific target component is input to the information processing device 1. The SMILES notation of the target component may be acquired from a database such as PubChem (https: / / pubchem.ncbi.nlm.nih.gov / ).
[0020] The information processing device 1 includes an HSP estimation model 2 and a retention time estimation model 3. The HSP estimation model 2 has undergone machine learning processing for multiple target components, with SMILES notation as explanatory variables and HSP values as response variables. The retention time estimation model 3 has undergone machine learning processing, with the HSP of the target component, the HSP of the stationary phase, and the HSP of the mobile phase as explanatory variables and the logarithm (logk') of the retention coefficient (k') of the target component as response variable. Note that the explanatory variables of the retention time estimation model 3 may include other variables (for example, the molecular volume of the target component).
[0021] The information processing device 1 applies the SMILES representation of a specific target component to the HSP estimation model 2, thereby obtaining an estimation result of the HSP of the specific target component.
[0022] HSPs of a specific target component, HSPs of a specific stationary phase, and HSPs of a specific mobile phase are input to the information processing device 1. Each HSP may be an estimation result output from the HSP estimation model 2, or may be obtained from an information source other than the HSP estimation model 2.
[0023] The information processing device 1 obtains an estimation result of the logarithm of the retention coefficient of the specific target component by applying the HSP of the specific target component, the HSP of the specific stationary phase, and the HSP of the specific mobile phase to the retention time estimation model 3. The information processing device 1 obtains an estimation result of the retention coefficient from the logarithm of the retention coefficient (k'), and then calculates the retention time (t) of the specific target component using the estimation result of the retention coefficient and a given hold-up time (t0).R ) is estimated (calculated).
[0024] 2 is a diagram showing an example of the hardware configuration of the information processing device 1. The information processing device 1 includes a processor 10, which is an example of a control circuit, a memory 20 that functions as a storage unit, and an input / output port 30. A mouse 40, a keyboard 50, and a display device 60 are connected to the input / output port 30. One or more terminal devices may be connected to the input / output port 30 via the Internet, an in-house network, or the like.
[0025] The information processing device 1 is configured, for example, based on a personal computer. The information processing device 1 may also be configured as a server that can be accessed from one or more terminal devices via a network such as the Internet.
[0026] The memory 20 non-temporarily stores an analysis program 200, HSP data 210, and retention time data 220.
[0027] By executing the analysis program 200 , the processor 10 functions as a first model creation unit 201 , a second model creation unit 202 , a first processing unit 203 , a second processing unit 204 , an image processing unit 205 , and an output unit 206 .
[0028] The first model creation unit 201 performs machine learning processing of the HSP estimation model 2. The second model creation unit 202 performs machine learning processing of the retention time estimation model 3. The first processing unit 203 obtains estimation results of the HSP of the target component using the HSP estimation model 2. The second processing unit 204 obtains estimation results of the retention time of the target component using the retention time estimation model 3. The image processing unit 205 creates image data including various estimation results. The output unit 206 outputs a display signal including the image data to the display device 60 via the input / output port 30.
[0029] The HSP data 210 includes an estimation model 211. The estimation model 211 includes data for configuring the HSP estimation model 2 and a set of parameters. At least a part of the set of parameters is updated in the machine learning process.
[0030] The HSP data 210 includes a plurality of training samples. The plurality of training samples are classified into training data 212 and validation data 213. In one implementation example, 80% of the plurality of training samples are classified into training data 212 and 20% are classified into validation data 213. Each of the plurality of training samples includes a SMILES representation of a target component and an HSP value of the target component.
[0031] The retention time data 220 includes an estimation model 221. The estimation model 221 includes data for configuring the retention time estimation model 3 and a set of parameters. At least a part of the set of parameters is updated in the machine learning process.
[0032] The retention time data 220 includes a plurality of training samples. The plurality of training samples are classified into training data 212 and validation data 213. In one implementation example, 80% of the plurality of training samples are classified into training data 212 and 20% are classified into validation data 213. Each of the plurality of training samples includes an HSP of a target component, an HSP of a stationary phase, an HSP of a mobile phase, and the logarithm of the retention time of the target component for the combination of the stationary phase and the mobile phase.
[0033] [Estimation Model for HSPs] Fig. 3 is a diagram showing an example of the configuration of the estimation model 2 for HSPs. The example in Fig. 3 conforms to the SMILES2vec method. HSPs include three types of parameters (a dispersion force term (dD), a polar term (dP), and a hydrogen bond term (dH)). The estimation model 2 for HSPs has settings corresponding to each of the three types of parameters. In other words, the estimation model 2 for HSPs is essentially composed of three models: a model for the dispersion force term, a model for the polar term, and a model for the hydrogen bond term.
[0034] The information processing device 1 provides the SMILES representation of the target component to each of the three models, thereby obtaining the three types of parameters as the HSP of the target component.
[0035] The HSP estimation model 2 may be generated for each column temperature (the set temperature of the column oven). More specifically, the training data used to generate the HSP estimation model 2 may be training data consisting of a set of SMILES notation and parameters classified by column temperature. The parameters classified by column temperature refer to the parameters associated with the column temperature at the time of acquisition. The column temperature may be a value selected from a plurality of predetermined temperature values. More specifically, the column temperature may be selected from a plurality of temperature values (40°C, 35°C, 30°C, 25°C, 20°C, etc.) in 5°C increments. The dispersion force term model, the polarity term model, and the hydrogen bond term model may each be composed of groups corresponding to a plurality of column temperatures. For example, the dispersion force term model group may include models corresponding to five temperatures (a model for a column temperature of 40°C, a model for a column temperature of 35°C, a model for a column temperature of 30°C, a model for a column temperature of 25°C, and a model for a column temperature of 20°C). Similarly, the group of models for polar terms and the group of models for hydrogen bonding terms may each include models corresponding to five different temperatures.
[0036] FIG. 3 shows an example of the configuration of each model constituting the HSP estimation model 2. In the example of FIG. 3, input data (SMILES notation) is input to the HSP estimation model 2. The model to be used is selected based on the parameters to be estimated (a model for dispersion force terms, a model for polar terms, or a model for hydrogen bond terms). The information processing device 1 may accept a setting of the column temperature to be estimated from the user. In this case, the information processing device 1 may select a model corresponding to the column temperature to be estimated as the model to be used. The input data is converted into a one-hot representation in the data processing unit 2A.
[0037] Next, in the word embedding layer 2B, each character is converted into a vector, thereby converting the one-hot representation into two-dimensional data.
[0038] Next, the two-dimensional data is input to a Long Short Term Memory (LSTM) 2D, and further input to an LSTM 2E.
[0039] The data output from the LSTM 2E is converted into one-dimensional data in the data processing unit 2F. The data output from the data processing unit 2F is input to the neural network 2G. A Dropout layer is inserted in the neural network 2G to prevent overlearning.
[0040] The HSP estimation model 2 outputs the estimation result of the HSP based on the output of the neural network 2G.
[0041] The estimation of HSPs using the HSP estimation model 2 differs from the method of calculating the HSP of each component by focusing on the physicochemical properties of the individual components as disclosed in Non-Patent Document 1, in that it uses a model that has been subjected to machine learning processing using the SMILES notation of many components. As a result, the means for estimating HSPs using the HSP estimation model 2 is expected to provide a simpler means while maintaining the same accuracy as conventional methods.
[0042] [Retention Time Estimation Model] The retention time estimation model 3 is configured as a model according to, for example, a random forest algorithm, and is generated by machine learning processing.
[0043] In the machine learning process, the HSPs of the target component, the stationary phase, and the mobile phase are used as explanatory variables, and the logarithm of the retention coefficient of the target component is used as the response variable.
[0044] When the mobile phase is a mixed solution of two components, the "mobile phase HSP" is prepared by combining the HSPs of the two components as shown in formula (1).
[0045] [dDm, dPm, dHm]=[(a*dD1+b*dD2), (a*dP1+b*dP2), (a*dH1+b*dH2)] / (a+b) ... (1) In formula (1), "dD" represents the dispersion force term in HSP. "dP" represents the polar term in HSP. "dH" represents the hydrogen bond term in HSP. The subscript "m" represents the mixed solution. The subscript "1" represents one of the two components, and the subscript "2" represents the other of the two components. "a" and "b" represent the volume fractions of the one component and the other component in the mixed solution. More specifically, the volume fraction of the one component is "a / (a+b)," and the volume fraction of the other component is "b / (a+b)."
[0046] In addition to the HSPs of the target component, the stationary phase, and the mobile phase, other variables may be used as explanatory variables of the retention time estimation model 3. The other variables may be, for example, one or more variables selected from the group consisting of the pH of the mobile phase, the column temperature in column chromatography, and the molecular volume of the target component.
[0047] [Processing Flow] <Main Routine> Fig. 4 is a flowchart of an example of processing performed in the information processing device 1. In one implementation example, the information processing device 1 performs the processing of Fig. 4 by having the processor 10 execute a given program. In one implementation example, the processing of Fig. 4 starts when an operation is performed to start a given application program in the information processing device 1.
[0048] In step S10, the information processing device 1 determines whether an instruction to estimate an HSP has been input. If the information processing device 1 determines that an instruction to estimate an HSP has been input (YES in step S10), the control proceeds to step S20, and if not (NO in step S10), the control proceeds to step S30.
[0049] In step S20, the information processing device 1 estimates the HSP. The processing content of step S20 will be described later with reference to FIG.
[0050] In step S30, the information processing device 1 determines whether an instruction to estimate a retention time has been input. If the information processing device 1 determines that an instruction to estimate a retention time has been input (YES in step S30), the control proceeds to step S40; otherwise (NO in step S30), the control returns to step S10.
[0051] In step S40, the information processing device 1 estimates the retention time. The process of step S40 will be described later with reference to Fig. 6. After that, the information processing device 1 returns control to step S10.
[0052] The information processing device 1 may simultaneously receive input of both an instruction to estimate an HSP and an instruction to estimate a retention time for a specific target component. For example, the information processing device 1 receives input of both an instruction to estimate an HSP and an instruction to estimate a retention time, along with input of the SMILES notation of the specific target component and values other than the HSP of the target component of the explanatory variables of the retention time estimation model 3 (the HSP of a specific stationary phase, the HSP of a specific mobile phase, and other necessary variables).
[0053] <Estimation of HSP> FIG. 5 is a flowchart of a subroutine of step S20 in FIG.
[0054] In step S200, the information processing device 1 acquires data in SMILES notation. In one implementation, the data is input via the keyboard 50.
[0055] In step S202, the information processing device 1 applies the SMILES notation to the HSP estimation model 2 that functions as a model for dispersion force terms (dD model).
[0056] In step S204, the information processing device 1 acquires the estimation result of the parameter dD constituting the HSP.
[0057] In step S206, the information processing device 1 applies the SMILES notation to the HSP estimation model 2 that functions as a model for polar terms (dP model).
[0058] In step S208, the information processing device 1 acquires the estimation result of the parameter dP constituting the HSP.
[0059] In step S210, the information processing device 1 applies the SMILES notation to the estimated model for HSP 2 that functions as a model for hydrogen bond terms (dH model).
[0060] In step S212, the information processing device 1 acquires the estimation result of the parameter dH constituting the HSP.
[0061] In step S214, the information processing device 1 outputs the estimation results of the three types of parameters acquired in steps S204, S208, and S212 as estimation results of the HSP. One example of output is display on the display device 60.
[0062] Thereafter, the information processing device 1 returns the control to Fig. 4. <Estimation of Retention Time> Fig. 6 is a flowchart of a subroutine of step S40 in Fig. 4 .
[0063] In step S400, the information processing device 1 acquires the HSP of the target component (specific target component) for which the retention time is to be estimated. The information processing device 1 may acquire the HSP of the target component by receiving input via the keyboard 50 or the like, or may acquire the HSP as an estimation result of the HSP estimation model 2.
[0064] In step S402, the information processing device 1 acquires the HSP of the stationary phase (specific stationary phase) that is expected to be used for the specific target component. The information processing device 1 may acquire the HSP of the target component by accepting input via the keyboard 50 or the like, or may acquire the HSP as an estimation result of the HSP estimation model 2.
[0065] In step S404, the information processing device 1 acquires the HSP of the mobile phase (specific mobile phase) that is expected to be used for the specific target component. The information processing device 1 may acquire the HSP of the target component by accepting input via the keyboard 50 or the like, or may acquire the HSP as an estimation result of the HSP estimation model 2.
[0066] If the mobile phase is a mixture of multiple solvents, the information processing device 1 may calculate the HSP of the mixture in step S404. More specifically, in step S404, the information processing device 1 receives input of the HSP and volume fraction of each of the multiple solvents, and calculates the HSP of the mixture using the input in accordance with formula (1).
[0067] In step S406, the information processing device 1 acquires, as necessary, values of explanatory variables other than HSP used in the retention time estimation model 3 (pH of a specific mobile phase, column temperature in column chromatography, and / or molecular volume of the target component).
[0068] In step S408, the information processing device 1 applies the data acquired in steps S400, S402, S404 (, S406) to the retention time estimation model 3.
[0069] In step S410, the information processing device 1 acquires an estimated result of the retention coefficient (its logarithm) as a result of applying the data in step S408.
[0070] In step S412, the information processing device 1 derives an estimated result of the retention time of the specific target component using the retention coefficient acquired in step S410.
[0071] In step S414, the information processing device 1 outputs the estimation result derived in step S412. One example of output is displaying the result on the display device 60. Another example is transmitting the data to an external device. Thereafter, the information processing device 1 returns control to FIG. 4.
[0072] In the process described above, the retention time is output as an estimation result. Note that the target to be output as an estimation result may be a retention coefficient or a logarithm of the retention coefficient.
[0073] The information processing device 1 may perform the control of steps S400 to S412 for each of the multiple target components. This derives an estimation result of the retention time for each of the multiple target components in chromatography using a common mobile phase and stationary phase. Then, in step S414, the information processing device 1 may output the estimation results for the multiple target components all at once.
[0074] When the control of steps S400 to S412 is performed for isocratic elution, the estimation results for multiple target components are output all at once, allowing the user to determine whether the elution of the multiple target components should be performed as gradient elution in the actual elution. For example, if the estimation result of isocratic elution shows an appropriate difference in the retention times of the multiple target components, the user can determine that the actual elution can be performed as isocratic elution. On the other hand, if the estimation result of isocratic elution shows no appropriate difference in the retention times of the multiple target components, the user can determine that the actual elution should be performed as gradient elution.
[0075] When outputting the estimation results for multiple target components, the information processing device 1 may determine whether the actual elution can be performed as isocratic elution or whether the actual elution should be performed as gradient elution, as described above, and output the result of the determination together with the estimation results. The information processing device 1 may use a predetermined threshold value related to the interval between retention times for the determination.
[0076] [Example] <Estimation of HSPs> An example of training and use of the HSP estimation model 2 will be described. In this example, SMILES notations of 1,792 types of target components were used as training samples in the machine learning process of the HSP estimation model 2. Fig. 7 is a diagram for explaining a specific example of SMILES notation used in the machine learning process of the HSP estimation model 2. In the data shown in Fig. 7, each training sample is composed of a pair of the SMILES notation and the measured value of one type of parameter constituting the HSP (dispersion force term (dD), polar term (dP), or hydrogen bond term (dH)).
[0077] Of the HSP estimation models 2, 1,792 pieces of data, each consisting of a pair of SMILES notation and the measured value of a dispersion force term parameter, were used as training samples for the machine learning processing of the dispersion force term model. 1,792 pieces of data, each consisting of a pair of SMILES notation and the measured value of a polar term parameter, were used as training samples for the machine learning processing of the polar term model. 1,792 pieces of data, each consisting of a pair of SMILES notation and the measured value of a hydrogen bond term parameter, were used as training samples for the machine learning processing of the hydrogen bond term model.
[0078] Fig. 8 is a diagram for explaining the results of estimation using the HSP estimation model 2. Three graphs G11, G12, and G13 are shown in Fig. 8. The graphs G11, G12, and G13 represent the estimated values of dD, dP, and dH, respectively. Graphs G11, G12, and G13 each show the results for 34 target components (benzylalcohol, phenol, 3-phenylpropanol, p-chlorophenol, acetophenone, benzonitrile, nitrobenzene, methylbenzoate, anisole, benzene, p-nitrotoluene, p-nitrobenzylchloride, toluene, benzophenone, bromobenzene, naphthalene, ethylbenzene, p-xylene, p-dichlorobenzene, propylbenzene, n-butylbenzene, diethylformamide, methylparaben, ethylparaben, propylparaben, butylparaben, acetanilide, propiophenone, butyrophenone, valerophenone, hexanophenone, heptanophenone, octanophenone, and ibuprofen).
[0079] In each of graphs G11, G12, and G13, the horizontal axis represents the estimated values of parameters using the HSP estimation model 2, and the vertical axis represents the calculated values of parameters obtained using the YMB simulator described in Non-Patent Document 1. Each of lines L11, L12, and L13 represents a regression line. The open plots represent estimated values. The black plots represent values on the regression line that correspond to the estimated values.
[0080] As shown in graphs G11, G12, and G13, the estimated values of HSPs by the HSP estimation model 2 were close to the estimated values of HSPs obtained using the YMB simulator.
[0081] Therefore, the estimation means of the HSP estimation model 2 can be said to be a simple means that can replace the YMB simulator, which requires complex physicochemical calculations.
[0082] <Retention Time Estimation> An example of training and use of the retention time estimation model 3 will be described. In this example, a model based on a random forest regression algorithm was used as the retention time estimation model 3. 1,020 types of data sets were used as training samples for the machine learning process of the retention time estimation model 3. The 1,020 types of data sets were composed of combinations of the 34 target components and 30 types of mobile phases. The 30 types of mobile phases were composed of mixed solutions of water and acetonitrile, with acetonitrile mixing ratios varying in 2% increments between 20% and 78%. An ODS (octadecylsilyl) column was assumed, and octadecane was used as the stationary phase.
[0083] Each of the 1,020 data sets included 10 variables (HSP (dD, dP, dH) of the target component, HSP (dD, dP, dH) of the stationary phase, HSP (dD, dP, dH) of the mobile phase, and the molecular volume of the target component) as explanatory variables, and included the logarithm of the retention coefficient of the target substance as the response variable.
[0084] Fig. 9 is a diagram for explaining the results of estimation using the retention time estimation model 3. In the graph shown in Fig. 9, the vertical axis represents the logarithm of the retention coefficient (logk'), which is an example of the results of estimation using the retention time estimation model 3, and the horizontal axis represents the measured logarithm of the retention coefficient. Line L23 represents the regression line.
[0085] In the example of FIG. 9, the root mean square error (RMSE) value is 0.06482, and the coefficient of determination (R 2 ) was 0.99345. This indicates that the estimation by retention time estimation model 3 is highly accurate.
[0086] Fig. 10 is a diagram showing the estimation results using the retention time estimation model 3 together with the estimation results using other methods. Three graphs G21, G22, and G23 are shown in Fig. 10. Graphs G21 and G22 represent the estimation results using other methods, shown for reference. Graph G23 is simply a scaled-down version of the graph shown in Fig. 9. Each of lines L21, L22, and L23 represents a regression line.
[0087] In graphs G21 and G22, the vertical axis represents the measured logarithm of the retention factor, and the horizontal axis represents the estimated logarithm of the retention factor.
[0088] Graph G21 represents the estimation results for the same 34 target components as graph G23, by simple regression analysis according to the following equation (2).
[0089] logk'=C1+C2*molecular volume of target component*(D1-D2) ... (2) In formula (2), "C1" and "C2" represent given constants. "D1" represents the HSP distance between the mobile phase and the target component. "D2" represents the HSP distance between the stationary phase and the target component. Note that the HSP distance (HSP_dis) is calculated according to the following formula (3) when one HSP is represented by the vector (dD1, dP1, dH1) and the other HSP is represented by the vector (dD2, dP2, dH2).
[0090] HSP_dis={4*(dD1-dD2) 2 +(dP1-dP2) 2 +(dH1-dH2) 2} 0.5...(3) Graph G22 represents the estimation results for the same 34 target components as graph G23, by multiple regression analysis according to the following equation (4).
[0091] logk'=C1+C2*molecular volume of target component+C3*D1+C4*D2 ...(4) In formula (4), "C1", "C2", and "C3" represent given constants. "D1", like formula (2), represents the HSP distance between the mobile phase and the target component. "D2", like formula (2), represents the HSP distance between the stationary phase and the target component.
[0092] In the example of graph G21, the coefficient of determination R 2 In the example of graph G22, the coefficient of determination R 2 On the other hand, in the example of graph G23, the coefficient of determination R 2 The value was 0.99345, which was higher than that of graphs G21 and G22.
[0093] In the example of graph G23, the molecular volume of the target component was used as an explanatory variable in addition to the HSPs of the target component, stationary phase, and mobile phase. In this regard, even when the molecular volume was not used as an explanatory variable, when the pH of the mobile phase and / or the column temperature were used as explanatory variables instead of the molecular volume, or when the pH of the mobile phase and / or the column temperature were used as explanatory variables in addition to the molecular volume, the estimation results using retention time estimation model 3 showed similar trends to the estimation results using simple regression analysis and multiple regression analysis.
[0094] As a result, the retention time estimation model 3 can handle the nonlinear relationship between HSPs and retention times, and is thought to be able to estimate retention coefficients (retention times) with higher accuracy than the above-mentioned simple regression analysis and multiple regression analysis.
[0095] Furthermore, when a model based on the random forest algorithm was used as the retention time estimation model 3, particularly high estimation accuracy was obtained.
[0096] Furthermore, SMILES notation can distinguish between isomers of a compound. Therefore, when HSPs are derived using SMILES notation, HSPs can be derived for each isomer of a compound, and thus, the retention time estimation results derived using HSPs can also be derived for each isomer of a compound. By using the retention time estimation results derived in this way, it is possible to identify the type of isomer of a compound based on the difference in retention time in chromatography.
[0097] [Processing Flow (2)] Fig. 11 is a flowchart of another example of processing performed by the information processing device 1. Unlike the example of Fig. 4, the example of Fig. 11 describes the estimation of retention times in isocratic elution and gradient elution separately. The example of Fig. 11 also describes the output of a report on the estimation results. The details of the processing shown in Fig. 11 will be described below.
[0098] The control of each of steps S10 to S40 of the process shown in Fig. 11 corresponds to the control of each of steps S10 to S40 described with reference to Fig. 4. In the example of Fig. 11, the "instruction to estimate retention time" in step S30 means estimation of retention time in isocratic elution. In the example of Fig. 11, if the information processing device 1 determines in step S30 that an instruction to estimate retention time in isocratic elution has been input (YES in step S30), the control proceeds to step S40; otherwise (NO in step S30), the control proceeds to step S50.
[0099] In step S40, the information processing device 1 estimates the retention time in isocratic elution. After that, the information processing device 1 advances the control to step S50.
[0100] In step S50, the information processing device 1 determines whether an instruction to estimate a retention time in gradient elution has been input. If the information processing device 1 determines that an instruction to estimate a retention time in gradient elution has been input (YES in step S50), the control proceeds to step S60; otherwise (NO in step S50), the control proceeds to step S70.
[0101] In step S60, the information processing device 1 estimates retention times in gradient elution. The details of estimating retention times in gradient elution in step S60 will be described later with reference to Fig. 12. Thereafter, the information processing device 1 advances control to step S70.
[0102] In step S70, the information processing device 1 determines whether an instruction to output a report on the estimation result has been input. If the information processing device 1 determines that an instruction to output a report has been input (YES in step S70), the control proceeds to step S80, and if not (NO in step S70), the control returns to step S10.
[0103] In step S80, the information processing device 1 outputs a report. The contents of the report output in step S80 will be described later with reference to Figures 14 and 15. Thereafter, the information processing device 1 returns control to step S10.
[0104] <Estimation of Retention Factor in Gradient Elution> FIG. 12 is a flowchart of the subroutine of step S60 (estimation of retention time in gradient elution) in FIG.
[0105] In step S600, the information processing device 1 acquires the HSP of the target component for which the retention time is to be estimated.
[0106] In step S602, the information processing device 1 acquires the HSP of the stationary phase. In step S604, the information processing device 1 acquires values of explanatory variables other than the HSP used in the retention time estimation model 3, as necessary (such as the pH of a specific mobile phase, the column temperature in column chromatography, and / or the molecular volume of the target component).
[0107] In step S606, the information processing device 1 sets the value of the variable T to "0." The variable T represents the time used for gradient elution. More specifically, the variable T represents the number of unit times corresponding to the length of time elapsed from the start of elution. For example, if the unit time is set to 5 seconds, T=2 represents 10 seconds elapsed from the start of elution. In other words, T=2 represents the timing 10 seconds after the start of elution.
[0108] In gradient elution, the proportions of multiple types of solutions in the mobile phase change over time. Information specifying the change in the proportions of multiple types of solutions in the mobile phase over time is registered in the information processing device 1 as a condition for estimation. This information is input by, for example, a user.
[0109] An example of the information is information that defines the ratio of the first solution to the second solution per unit time when a first solution and a second solution are used as the mobile phase.
[0110] For example, the information specifies that the percentage change begins 10 seconds after the start of elution, the percentage change every 5 seconds is 5%, and the initial percentage of the first solution is 100% and the initial percentage of the second solution is 0%. In this case, for the first 10 seconds after the start of elution, the mobile phase is 100% first solution and 0% second solution. 15 seconds after the start of elution, the mobile phase is 95% first solution and 5% second solution. 20 seconds after the start of elution, the mobile phase is 90% first solution and 10% second solution. 25 seconds after the start of elution, the mobile phase is 85% first solution and 15% second solution. 105 seconds after the start of elution, the mobile phase is 5% first solution and 95% second solution. The breakdown of the mobile phase from 110 seconds after the start of elution is 100% of the first solution and 0% of the second solution.
[0111] In step S608, the information processing device 1 acquires the HSP of the mobile phase at the timing represented by the variable T. More specifically, the information processing device 1 identifies the ratio of the multiple solutions in the mobile phase at the timing represented by the variable T in accordance with the conditions for estimation described above. Then, the information processing device 1 acquires, as the HSP of the mobile phase, the HSP of a mixture in which the multiple solutions are mixed in the identified ratio in accordance with the above-described formula (1).
[0112] In step S610, the information processing device 1 applies the data acquired in steps S600, S602, S604, and S608 to the retention time estimation model 3.
[0113] In step S612, the information processing device 1 acquires an estimated result of the retention coefficient (its logarithm) as a result of applying the data in step S610.
[0114] In step S614, the information processing device 1 derives an estimated result of the retention time of the specific target component using the retention coefficient acquired in step S610.
[0115] In step S616, the information processing device 1 calculates the migration distance of the target component within the column in the latest unit time using the estimation result derived in step S614. More specifically, the information processing device 1 calculates the time required for migration within the column using the retention coefficient, and calculates the migration speed of the retention coefficient from the calculated time and the column length. The information processing device 1 then calculates the product of the migration speed and the unit time as the migration distance.
[0116] In step S618, the information processing device 1 determines whether the cumulative value of the movement distance of the target component is equal to or greater than the length of the column. If the information processing device 1 determines that the cumulative value of the movement distance of the target component is equal to or greater than the length of the column (YES in step S618), the control proceeds to step S622; otherwise (NO in step S618), the control proceeds to step S620.
[0117] In step S620, the information processing device 1 updates the value of the variable T by adding 1, and returns the control to step S606.
[0118] In step S622, the information processing device 1 identifies the time derived as the product of the latest value of the variable T and the unit time as the estimated result of the retention time. Note that the product of the latest value of the variable T and the unit time corresponds to the cumulative value of the unit time when the cumulative value of the migration distance of the target component reaches the length of the column.
[0119] In step S624, information processing device 1 outputs the estimation result, and returns control to FIG.
[0120] In the process of FIG. 12 described above, the estimated retention time for gradient elution is output.
[0121] <Specific Example of Output of Retention Time Estimation Results> Fig. 13 is a diagram showing a specific example of output of retention time estimation results. The information processing device 1 may estimate retention times by performing the control of steps S600 to S622 for each of multiple target components. Then, in step S624, the information processing device 1 may generate display information on a single screen using the retention time estimation results for the multiple target components. The screen 600 shown in Fig. 13 displays the estimation results for the multiple target components. The screen 600 is displayed on, for example, the display device 60. Information input via the mouse 40 or the keyboard 50 is also displayed on the screen 600.
[0122] The screen 600 includes four areas 610, 620, 630, and 640. Area 610 displays information identifying the target component. Area 620 displays the elution conditions used for the estimation. Area 630 displays a chromatograph created using the estimation results. Area 640 displays the numerical values of the estimation results.
[0123] In one implementation example, an instruction to estimate a retention time in gradient elution may be input with the target component input in region 610 and the conditions input in region 620. The information processing device 1 may start the process of Fig. 12 in response to the input of the instruction. Then, the information processing device 1 may output the estimation result by updating the display on the screen 600 to add the estimation result identified in the process of Fig. 12.
[0124] The contents displayed in each of areas 610 to 640 in Figure 13 will be explained below. Area 610 displays the names of five compounds (phenol, benzonitrile, p-chlorophenol, acetophenone, and nitrobenzene) as five target components. Area 610 also displays the HSP (dD, dP, dH) and molecular volume (Volume 3D) values for each target component.
[0125] The conditions displayed in area 620 include the elution mode (Separation Mode: Reversed Phase), the type of stationary phase (Stationary Phase: SB-C18), the types of two solutions used as mobile phases (Mobile Phase A: Acetonitrile, Mobile Phase B: Water), the percentage of solution A during isocratic elution (Mobile Phase A%: 60), the mobile phase flow rate (Flow rate (ml / min): 20), and the column temperature (Temperature (°C): 40). Area 620 may also display conditions for estimation (those mentioned in connection with the explanation of variable T in step S606).
[0126] In addition to a chromatogram including peaks corresponding to each target component (five peaks P11 to P15), a line L11 indicating the change in the proportion of one of the two solutions in the mobile phase is displayed in the region 630. The chromatogram is generated by the information processing device 1 in a simulated manner so as to include peaks representing the retention times identified as the estimated results for each target component.
[0127] In area 640, the retention time of each target component is shown along with the value of the retention coefficient (k').
[0128] <Report Output> FIG. 14 is a flowchart of the subroutine of step S80 (report output) in FIG.
[0129] In step S800, the information processing device 1 acquires information to be output. In one implementation example, the information to be output includes information identifying each of one or more target components, information identifying the mobile phase, information identifying the stationary phase, and conditions used for the estimation.
[0130] In step S802, the information processing device 1 generates a report including the estimation results of the retention times for one or more target components based on the information to be output acquired in step S800.
[0131] In step S804, the information processing device 1 outputs a report as the result of step S802. Thereafter, the information processing device 1 returns control to FIG.
[0132] 15 is a diagram showing an example of an output report. A report 700 includes areas 710, 720, 730, 740, 750, and 760.
[0133] Bibliographical information is displayed in area 710. In one implementation, the bibliographical information includes the name given by the user to the report to be output (Aromatic20230120), the user ID (440275), and the user name (YasuhiroMito).
[0134] Area 720 displays the conditions used to estimate the retention time. In the example of FIG. 15, "Separation Mode" indicates the elution mode, and its value is displayed as "Reversed Phase." "Stationary Phase" indicates the stationary phase, and its value is displayed as "SB-C18." In the example of FIG. 15, additional information about the stationary phase (HSP value (δD: 15.91, δP: 0.1, δH: 0.1), Partial size [μm] = 5, id (mm): 4.6, Length (mm) = 100.0, Inter porosity: 0.6, Intra porosity: 0.6, N: 16000) is also displayed. In particular, "Length" refers to the length of the region packed with the stationary phase, i.e., the length of the column.
[0135] "Mobile Phase A" represents one of the multiple solutions that make up the mobile phase, and its value is displayed as "acetonitrile." In the example of Figure 15, additional information about this solution (HSP values (δD: 15.61, δP: 16.64, δH: 8.32)) is also displayed.
[0136] "Mobile Phase B" represents another type of solution among the multiple types that make up the mobile phase, and its value is displayed as "water." In the example of Figure 15, additional information about this solution (HSP values (δD: 7.58, δP: 7.82, δH: 20.68)) is also displayed.
[0137] "Mobile Phase A%" represents the percentage of solution A in the case of isocratic elution, and its value is displayed as "60." "Flow rate (ml / min)" represents the flow rate of the mobile phase, and its value is displayed as "2." "Temperature (°C)" represents the column temperature, and its value is displayed as "40." "Gradient settings" represents the settings for gradient elution, and its values include the initial percentage of solution A (the type of solution represented by "Mobile Phase A" above) in the mobile phase ("Initial A%: 20"), the final percentage of solution A in the mobile phase ("Final A%: 90"), the start and end times from the start of elution for the change in the percentage of multiple solutions in the mobile phase (Start time (min): 11, End time (min): 20), and the type of variable used for estimation (Compound information [HSP and Volume3D]).
[0138] Area 730 displays a table containing HSP values and molecular volume values for each of one or more target components to be estimated.
[0139] A simulated chromatogram generated using the estimation results is displayed in region 740. This chromatogram is generated to have peaks corresponding to the estimation results of the retention times of one or more target components, similar to the manner described for region 640 in Fig. 13. Region 740 also displays dashed lines representing the change in the proportion of one type of solution in a mobile phase composed of multiple types of solutions.
[0140] The estimation results for each of the one or more target components to be estimated are displayed in area 750. The estimation results include the retention time (RT) and retention coefficient (k'). The area 740 further displays the SMILES notation (Smiles), molecular formula, and molecular weight of each target component.
[0141] Area 760 displays the structural formula of each of the one or more target components to be estimated. As described above, the output report includes a simulated chromatogram generated using the retention times of the one or more target components. This allows the estimation results to be provided to the user in an easy-to-understand manner. The output report includes the structural formula of each target component in addition to the retention times that are the estimation results for the multiple target components. This allows the user to evaluate the differences in the estimation results of retention times between the multiple target components while taking into account the differences in the structural formulas of the multiple target components.
[0142] 14 , the information processing device 1 may use the retention time estimation result stored in advance in the memory 20, or may estimate the retention time itself. When estimating the retention time in the report output, the information processing device 1 estimates the retention time in step S40 and / or step S60 in response to acquiring the information to be output in step S600, i.e., in step S602.
[0143] 14 , the information processing device 1 may acquire both estimation conditions for both isocratic elution and gradient elution as information to be output. In this case, the information processing device 1 generates a report using both the estimated retention times for isocratic elution and the estimated retention times for gradient elution for each of one or more target components. In the report generated in this manner, the conditions for both isocratic elution and gradient elution are displayed in region 720. Furthermore, region 740 displays a chromatogram generated using the retention times for isocratic elution, a chromatogram generated using the retention times for gradient elution, and a gradient curve representing the change in the proportion of one solution in the mobile phase.
[0144] [Processing Flow (3)] Fig. 16 is a flowchart of yet another example of processing performed in the information processing device 1. The processing in Fig. 16 is performed to provide information regarding the validity of each of one or more candidates being an unknown component. In one implementation example, the information processing device 1 starts the processing in Fig. 16 in response to a specific menu being selected while a given application is being executed.
[0145] In step S900, the information processing device 1 acquires the chromatographic analysis results of the unknown component. The analysis results are actual measurements obtained by chromatography. The analysis results may be input to the information processing device 1 by a user via the mouse 40 or the keyboard 50, or may be transmitted to the information processing device 1 from an external device.
[0146] In step S902, the information processing device 1 identifies one or more candidates for compounds that are assumed to be the unknown component. The information processing device 1 may identify the one or more candidates by receiving an input of the one or more candidates from a user and recognizing the input one or more candidates.
[0147] The information processing device 1 may identify one or more candidates by selecting one or more compounds from among a plurality of compounds registered in the memory 20 based on the results of mass analysis of the unknown component. For example, the information processing device 1 acquires a mass spectrogram as a result of the mass analysis, identifies peaks in the mass spectrogram, and identifies m / z values corresponding to the identified peaks. Then, the information processing device 1 selects, from among the plurality of compounds registered in the memory 20, one or more compounds having the identified m / z values as molecular weights as one or more candidates for the unknown component.
[0148] In step S904, the information processing device 1 identifies an estimation result of the retention time for each of the one or more candidates identified in step S902. The estimation result follows the same conditions as the analysis result acquired in step S900 (for example, the same mobile phase and stationary phase). The information processing device 1 may acquire the chromatography conditions under which the analysis result was obtained in addition to the analysis result in step S900. In step S904, the information processing device 1 may perform the estimation of the retention time in step S40 and / or step S60 for each of the one or more candidates.
[0149] In step S906, the information processing device 1 compares the analysis result acquired in step S900 with the estimation result of the retention time of each candidate identified in step S904, and determines, for each candidate, the validity of each candidate being an unknown component. In one example, the information processing device 1 may calculate, for each candidate, the difference between the estimation result and the analysis result acquired in step S900, and determine, as the above-mentioned validity, a ranking according to the smallest difference. In another example, for one or more candidates, the compound with the smallest difference between the estimation result and the analysis result acquired in step S900 may be classified as a final candidate, and the remaining compounds may be classified as non-candidates, thereby determining the above-mentioned validity.
[0150] In step S908, the information processing device 1 outputs the result of the determination of validity in step S906. After that, the information processing device 1 ends the processing in FIG.
[0151] As described above, the process of FIG. 16 outputs the relevance of each of the one or more candidates as an unknown component. The chromatographic analysis results of the unknown component and the estimated retention times of the one or more candidates are used to determine the relevance. The user can refer to the relevance of each of the one or more output candidates to identify the unknown component.
[0152] Aspects It will be understood by those skilled in the art that the exemplary embodiments described above are examples of the following aspects.
[0153] (Item 1) An estimation method according to one aspect is a method for estimating the retention time of a specific target component in chromatography that uses a specific stationary phase and a specific mobile phase, and includes the steps of acquiring an HSP of the specific target component, acquiring HSPs of the specific stationary phase and the specific mobile phase, and acquiring an estimated result of an index related to the retention time of the specific target component by inputting the HSPs of the specific target component, the specific stationary phase, and the specific mobile phase into a machine learning model, wherein the machine learning model may have been subjected to machine learning processing in which the HSPs of the target component, the stationary phase, and the mobile phase are used as explanatory variables, and the index related to the retention time of the target component is used as a response variable.
[0154] The estimation method described in paragraph 1 provides a technique for improving the accuracy of estimating the retention time of a target component in chromatography, thereby improving the convenience of estimating retention times.
[0155] (Item 2) In the estimation method according to item 1, the machine learning model may be a classification model based on a random forest.
[0156] According to the estimation method described in the second paragraph, a technique is provided that more reliably improves the accuracy of estimating the retention time of a target component in chromatography.
[0157] (Item 3) In the estimation method according to item 1 or 2, the index may be a logarithm of a retention coefficient.
[0158] According to the estimation method described in paragraph 3, estimation of retention time can be more easily performed. (paragraph 4) In the estimation method described in any one of paragraphs 1 to 3, the machine learning model is subjected to machine learning processing using a pH of a mobile phase as an explanatory variable, and the estimation method may further include a step of acquiring the pH of the specific mobile phase, and in the step of acquiring the estimation result, the pH of the specific mobile phase may be further input to the machine learning model.
[0159] According to the estimation method described in Section 4, a technique is provided that more reliably improves the accuracy of estimating the retention time of a target component in chromatography.
[0160] (5) In the estimation method described in any one of paragraphs 1 to 4, the machine learning model is subjected to machine learning processing using a column temperature as an explanatory variable, and the estimation method further includes a step of acquiring a column temperature at which the specific stationary phase is used, and in the step of acquiring the estimation result, the column temperature at which the specific stationary phase is used may be further input to the machine learning model.
[0161] According to the estimation method described in Section 5, a technique is provided that more reliably improves the accuracy of estimating the retention time of a target component in chromatography.
[0162] (Item 6) In the estimation method described in any one of Items 1 to 5, the machine learning model is subjected to machine learning processing using a molecular volume of a target component as an explanatory variable, and the estimation method further includes a step of acquiring a molecular volume of the specific target component, and in the step of acquiring an estimation result, the molecular volume of the specific target component may be further input to the machine learning model.
[0163] According to the estimation method described in Section 6, a technique is provided that more reliably improves the accuracy of estimating the retention time of a target component in chromatography.
[0164] (Clause 7) In the estimation method according to any one of clauses 1 to 6, the step of acquiring the HSP of the specific target component includes acquiring a SMILES notation of the specific target component, and inputting the SMILES notation of the specific target component into an estimation model to acquire an estimation result of the HSP of the specific target component, wherein the estimation result of the HSP of the specific target component is used as the HSP of the specific target component, and the estimation model may have been subjected to machine learning processing with the SMILES notation as an explanatory variable and the HSP as a target variable.
[0165] According to the estimation method described in item 7, estimation of HSPs can be easily carried out in estimating the retention time of a target component in chromatography.
[0166] (Clause 8) In the estimation method described in Clause 7, the HSP may include a plurality of parameters, and in the step of acquiring the HSP of the specific target component, each of the plurality of parameters may be estimated by applying settings corresponding to each of the plurality of parameters to the estimation model.
[0167] According to the estimation method described in Section 8, a technique is provided that more reliably improves the accuracy of estimating the retention time of a target component in chromatography.
[0168] (Item 9) In the estimation method of Items 7 or 8, the estimation model may be generated for each of a plurality of column temperatures, the step of acquiring the SMILES representation of the specific target component may include accepting a setting of the column temperature to be estimated, and acquiring the estimation result of the HSP of the specific target component may include selecting an estimation model corresponding to the column temperature to be estimated from the estimation models for the plurality of column temperatures.
[0169] According to the estimation method described in Section 9, a model customized for the column temperature to be estimated is used as the model used for estimating HSP, thereby improving the accuracy of HSP estimation.
[0170] (Clause 10) In the estimation method described in any one of clauses 1 to 9, the step of acquiring HSPs of the specific target component may include further acquiring HSPs of one or more other target components, the step of acquiring an estimation result of an index related to the retention time of the specific target component may include further acquiring an estimation result of the index of the one or more other target components, and the estimation method may further include a step of outputting the estimation results of the index of the specific target component and the one or more other target components.
[0171] According to the estimation method described in Section 10, the estimation results of retention times under common conditions for multiple components are provided.
[0172] (Clause 11) In the estimation method described in Clause 10, the outputting step may include determining the necessity of gradient elution based on the estimation results of the indexes of the specific target component and the one or more other target components, and outputting the result of the determination.
[0173] According to the estimation method described in Section 11, the user is provided with information that assists in determining the necessity of gradient elution.
[0174] (12) In the estimation method according to the 10th or 11th paragraph, the outputting step may include outputting a retention time as an estimation result of the index.
[0175] According to the estimation method described in Section 12, information about chromatography that is easy to understand intuitively is provided to the user.
[0176] (Item 13) In the estimation method according to Item 12, the retention times may include retention times in isocratic elution and retention times in gradient elution.
[0177] According to the estimation method described in Section 13, the user is provided with information to determine whether isocratic elution or gradient elution should be used for eluting multiple target components.
[0178] (14) In the estimation method according to the 12th or 13th paragraph, the retention time may be output in the form of a chromatogram.
[0179] According to the estimation method described in Section 14, the estimation result of the retention time is provided to the user in a format that allows the user to intuitively understand chromatography.
[0180] (Clause 15) In the estimation method described in Clause 14, the outputting step may include outputting conditions for obtaining the estimation result of the index so that the conditions are displayed on the same screen as the retention time.
[0181] According to the estimation method described in Section 15, the user can be provided with information that can be used to evaluate the estimation result together with the estimation result.
[0182] (Item 16) In the estimation method according to any one of Items 12 to 15, the outputting step may include outputting a structural formula of each compound of the specific target component and the one or more other target components.
[0183] According to the estimation method described in Section 16, the user can be provided with the estimation result as well as information that can be used to evaluate the estimation result in a format that is easy to understand intuitively.
[0184] (Item 17) In the estimation method according to any one of Items 1 to 16, the specific mobile phase includes a plurality of solvents, a ratio of each of the plurality of solvents in the specific mobile phase is set for each of a plurality of unit times, the estimation result of the index is a retention time, and the step of acquiring the estimation result of the index of the specific target component includes selecting one unit time from the plurality of unit times, acquiring the estimation result of the index for the selected one unit time, calculating a migration distance of the specific target component in a column based on the estimation result of the index acquired for the selected one unit time, and performing switching of the selected one unit time, and in the step of acquiring the estimation result of the index of the specific target component, the switching may be performed until a cumulative value of the migration distance reaches a predetermined column length, and the cumulative value of the unit time when the cumulative value of the migration distance reaches the predetermined column length may be acquired as a final estimation result.
[0185] The estimation method described in Section 17 provides a specific embodiment for estimating retention time in gradient elution.
[0186] (Item 18) The estimation method described in any one of Items 1 to 17 may further comprise the steps of acquiring a chromatographic analysis result of an unknown component, and identifying one or more candidates for the unknown component as the specific target component, wherein a specific stationary phase and a specific mobile phase are used in the chromatography of the unknown component, and may further comprise the step of identifying the validity of the one or more candidates being the unknown component based on the analysis result and the estimation result of the index corresponding to each of the one or more candidates.
[0187] The estimation method described in paragraph 18 can provide information useful for identifying the unknown component. (Item 19) In the estimation method described in paragraph 18, the step of identifying the one or more candidates may include selecting the one or more candidates based on a result of mass spectrometry of the unknown component.
[0188] According to the estimation method described in paragraph 19, the burden on the user for identifying one or more candidates can be reduced.
[0189] (Clause 20) Another aspect of the estimation method is a method for estimating HSP of a substance, comprising the steps of acquiring a SMILES notation of the substance, and acquiring an estimation result of the HSP of the substance by inputting the SMILES notation of the substance into an estimation model, and the estimation model may have been subjected to machine learning processing with the SMILES notation as an explanatory variable and the HSP as a target variable.
[0190] According to the estimation method described in paragraph 20, estimation of HSPs used for estimating retention times becomes simple, thereby providing a technology that improves the convenience of estimating the retention times of target components in chromatography.
[0191] (Clause 21) In the estimation method described in Clause 20, the HSP may include a plurality of parameters, and in the step of acquiring the HSP of the substance, each of the plurality of parameters may be estimated by applying settings corresponding to each of the plurality of parameters to the estimation model.
[0192] According to the estimation method described in paragraph 21, a technique is provided that more reliably improves the accuracy of estimating the retention time of a target component in chromatography.
[0193] (Clause 22) In the estimation method described in clause 20 or 21, the estimation model may be generated for each column temperature, the step of acquiring the SMILES representation of the substance may include accepting a setting of a temperature to be estimated, and the step of acquiring an estimation result of the HSP of the substance may include selecting an estimation model corresponding to the column temperature to be estimated.
[0194] According to the estimation method described in paragraph 22, a model customized for the temperature to be estimated is used as the model used for estimating HSPs, thereby improving the accuracy of HSP estimation.
[0195] (23rd paragraph) A computer program according to one aspect may cause a computer to implement the estimation method according to any one of the first to 22nd paragraphs when executed by a control circuit of the computer.
[0196] According to the estimation method described in paragraph 23, a technique is provided that improves the accuracy of estimating the retention time of a target component in chromatography, or that improves the convenience of estimating the retention time of a target component in chromatography by simply estimating HSPs.
[0197] (Clause 24) An information processing device according to one aspect is an information processing device comprising a control circuit and a memory device storing a computer program executed by the control circuit, and the computer program may be executed by the control circuit to cause the information processing device to implement the estimation method described in any one of clauses 1 to 22.
[0198] According to the information processing device described in paragraph 24, a technology is provided that improves the accuracy of estimating the retention time of a target component in chromatography or that improves the convenience of estimating the retention time of a target component in chromatography by simply estimating HSPs.
[0199] The embodiments disclosed herein should be considered to be illustrative in all respects and not restrictive. The scope of the present disclosure is defined by the claims, not by the description of the above-described embodiments, and is intended to include all modifications within the meaning and scope of the claims. Furthermore, it is intended that each technique in the embodiments can be implemented alone or, if necessary, in combination with other techniques in the embodiments to the extent possible.
[0200] 1 Information processing device, 2 HSP estimation model, 3 Retention time estimation model, 10 Processor, 20 Memory
Claims
1. 1. A method for estimating the retention time of a specific component of interest in a chromatography utilizing a specific stationary phase and a specific mobile phase, comprising: Obtaining Hansen Solubility Parameters (HSP) for the specific target component; obtaining HSPs of the specific stationary phase and the specific mobile phase, respectively; and inputting the HSPs of the specific target component, the specific stationary phase, and the specific mobile phase into a machine learning model to obtain an estimated result of an index related to the retention time of the specific target component; The machine learning model has been subjected to machine learning processing using the HSPs of the target component, the stationary phase, and the mobile phase as explanatory variables and an index related to the retention time of the target component as a response variable; An estimation method, wherein the machine learning model is a random forest classification model.
2. The estimation method according to claim 1 , wherein the index is a logarithm of the retention coefficient.
3. the machine learning model is subjected to machine learning processing using the pH of the mobile phase as an explanatory variable; acquiring a pH of the specific mobile phase; The estimation method according to claim 1 , wherein the step of obtaining the estimation result further inputs the pH of the specific mobile phase into the machine learning model.
4. the machine learning model is subjected to machine learning processing using a column temperature as an explanatory variable; obtaining a column temperature at which the particular stationary phase is used; The estimation method according to claim 1 , wherein the step of obtaining the estimation result further inputs the column temperature at which the specific stationary phase is used into the machine learning model.
5. the machine learning model is subjected to machine learning processing using the molecular volume of the target component as an explanatory variable; obtaining a molecular volume of the specific component of interest; The estimation method according to claim 1 , wherein the step of obtaining the estimation result further inputs the molecular volume of the specific target component into the machine learning model.
6. The step of obtaining HSP of the specific target component comprises: Obtaining a SMILES representation of the specific target component; and inputting the SMILES representation of the specific target component into a prediction model to obtain a prediction result of the HSP of the specific target component; As the HSP of the specific target component, the estimation result of the HSP of the specific target component is used, The estimation method according to claim 1 , wherein the estimation model is subjected to machine learning processing using SMILES notation as an explanatory variable and HSP as a target variable.
7. The HSP includes a plurality of parameters, The estimation method according to claim 6, wherein in the step of acquiring the HSP of the specific target component, each of the plurality of parameters is estimated by applying a setting corresponding to each of the plurality of parameters to the estimation model.
8. the estimation model is generated for each of a plurality of column temperatures; obtaining the SMILES representation of the specific target component includes receiving a setting of a column temperature to be estimated; The estimation method according to claim 6, wherein obtaining an estimation result of the HSP of the specific target component includes selecting an estimation model corresponding to the column temperature to be estimated from the estimation models for each of the plurality of column temperatures.
9. The step of obtaining HSPs of the specific target component includes further obtaining HSPs of one or more other target components; the step of obtaining an estimated result of an index related to the retention time of the specific target component includes further obtaining an estimated result of the index of the one or more other target components; The estimation method according to claim 1 , further comprising the step of outputting estimation results of the indexes of the specific target component and the one or more other target components.
10. The outputting step includes: Determining the need for gradient elution based on the estimated results of the indicators of the specific target component and the one or more other target components; and outputting a result of said determining.
11. The estimation method according to claim 9 , wherein the outputting step includes outputting a retention time as the estimation result of the index.
12. The estimation method according to claim 11 , wherein the retention times include retention times in isocratic elution and retention times in gradient elution.
13. The estimation method according to claim 11 , wherein the retention times are output in the form of a chromatogram.
14. The estimation method according to claim 11 , wherein the outputting step includes outputting a structural formula of each compound of the specific target component and the one or more other target components.
15. A method for estimating the retention time of a specific target component in a chromatography utilizing a specific stationary phase and a specific mobile phase, comprising: Obtaining Hansen Solubility Parameters (HSP) for the specific target component; obtaining HSPs of the specific stationary phase and the specific mobile phase, respectively; and inputting the HSPs of the specific target component, the specific stationary phase, and the specific mobile phase into a machine learning model to obtain an estimated result of an index related to the retention time of the specific target component; The machine learning model has been subjected to machine learning processing using the HSPs of the target component, the stationary phase, and the mobile phase as explanatory variables and an index related to the retention time of the target component as a response variable; The specific mobile phase comprises a plurality of solvents, a ratio of each of the plurality of solvents in the specific mobile phase is set for each of a plurality of unit times; The estimated result of the index is a retention time, The step of obtaining an estimation result of the index of the specific target component includes: selecting one unit time from the plurality of unit times; Obtaining an estimation result of the index for the selected one unit time; Calculating the migration distance of the specific target component in the column based on the estimation result of the index acquired for the selected one unit time; and performing switching of the selected one unit time, In the step of acquiring an estimation result of the index of the specific target component, The switching is performed until the cumulative value of the movement distance reaches a predetermined column length; An estimation method, wherein an accumulated value for the unit time when the accumulated value of the movement distance reaches a predetermined column length is acquired as a final estimation result.
16. obtaining a chromatographic analysis of the unknown component; and identifying one or more candidate unknown components as the identified component of interest; the chromatography of the unknown component utilizes the specified stationary phase and the specified mobile phase; The estimation method according to claim 1 , further comprising a step of determining the appropriateness of the one or more candidates being the unknown component based on the analysis result and the estimation result of the indicator corresponding to each of the one or more candidates.
17. The method of claim 16 , wherein the step of identifying one or more candidates comprises selecting the one or more candidates based on results of mass spectrometry of the unknown component.
18. 1. A method for estimating the Hansen Solubility Parameters (HSP) of a substance, comprising: obtaining a SMILES representation of the substance; and inputting the SMILES representation of the substance into a prediction model having a Long Short Term Memory (LSTM) to obtain a prediction result of the HSP of the substance, The estimation method is such that the estimation model is subjected to machine learning processing using SMILES notation as an explanatory variable and HSP as a target variable.
19. The HSP includes a plurality of parameters, 19. The estimation method according to claim 18, wherein in the step of acquiring HSPs of the substance, each of the plurality of parameters is estimated by applying settings corresponding to each of the plurality of parameters to the estimation model.
20. the estimation model is generated for each temperature; The step of acquiring the SMILES representation of the substance includes receiving a setting of a temperature to be estimated; The estimation method according to claim 18, wherein the step of obtaining an estimation result of the HSP of the substance includes selecting an estimation model corresponding to a temperature that is a target of the estimation.