Estimation method, computer program, and information processing device

By combining the machine learning model with HSP and SMILES expression, the convenience and accuracy issues of retention time estimation of object components in chromatography were solved, simple and high-precision retention time prediction was achieved, and the estimation effect of chromatography was improved.

CN120813835APending Publication Date: 2025-10-17SHIMADZU GENERAL SERVICES INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202380095680.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-08-24
Filing Date
2023-12-22
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing methods for estimating the retention time of target components in chromatography suffer from inconvenience and low accuracy. In particular, HSP estimation based on non-patent document 1 requires complex calculations, while the retention time prediction method based on non-patent document 2 is not accurate enough for practical use.

Method used

Using a machine learning model, the retention time of the target component is predicted by obtaining the Hansen Solubility Parameters (HSP) of the target component, the stationary phase, and the mobile phase, and estimating them using SMILES expression. This method includes an HSP estimation model and a retention time estimation model, each of which uses machine learning to obtain HSP and retention time estimation results.

Benefits of technology

The accuracy and convenience of estimating the retention time of target components in chromatography are improved, the calculation process of HSP is simplified, and a simpler and more accurate retention time prediction method is provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120813835A_ABST
    Figure CN120813835A_ABST
Patent Text Reader

Abstract

Provided is a technique for improving convenience in estimating the retention time of a target component in chromatography. A method of estimating a retention time of a specific object component, the method comprising the steps of: obtaining an HSP of the specific object component (S400); obtaining the respective HSP of the specific stationary phase and the specific mobile phase (S402, S404); and acquiring an estimation result of an index pertaining to the retention time of the specific target component by inputting the HSP of each of the specific target component, the specific stationary phase, and the specific mobile phase into the machine learning model (S410). The machine learning model performs machine learning processing using the HSP of each of the target component, the stationary phase, and the mobile phase as an explanatory variable, and using an index related to the retention time of the target component as a target variable.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to estimation of retention time of a target component in a chromatograph. BACKGROUND

[0002] In the past, various studies have been made on estimation of parameters in a chromatograph. For example, Japanese Patent No. 6110380 (Patent Literature 1) discloses a technique of predicting a retention index relating to retention time of a component (hereinafter, also referred to as "target component") dissolved in chromatography, using a KRI index (Kovats retention index) of the component.

[0003] Japanese Patent No. 6796842 (Patent Literature 2) discloses a technique of calculating Hansen Solubility Parameters (HSP) of a substance from a structural formula of the substance. In addition, an attempt has been made to estimate retention time from HSP distance, and a software (HSPiP) that estimates HSPs of unknown compounds is being sold.

[0004] Yamamoto Hiroshi, "YMB Simulator Property Estimation Function", [online], June 1, 2013, searched on December 13, 2020, Internet <URL: https: / / www.pirika.com / Chemistry / JP / TCPE / YMB.html> (Non-Patent Literature 1) discloses the following: HSPiP has an algorithm for inferring properties using YMB method built in, and is able to calculate various physical quantities of unknown compounds from the kind, number, and the like of functional groups.

[0005] Yamamoto Hiroshi, "High-Speed Liquid Chromatography and Hansen Solubility Parameters (HSP)", [online], September 9, 2011, searched on December 13, 2020, Internet <URL: https: / / www.pirika.com / HSP / JP / Examples / Docs / HPLC.html> (Non-Patent Literature 2) discloses the following: regarding the software disclosed in Patent Literature 2, various physical quantities of a target component are calculated using a SMILES expression of the target component, HSP is calculated using the physical quantities, and the calculated value of HSP is corrected using a molecular volume, whereby correlation of the estimated retention time with an actual value becomes high.

[0006] PRIOR ART DOCUMENTS

[0007] PATENT LITERATURE

[0008] Patent Literature 1: Japanese Patent No. 6110380

[0009] Patent Literature 2: Japanese Patent No. 6796842

[0010] Non Patent Literature

[0011] Non Patent Literature 1: Yamamoto, Hiroshi, "YMB Simulator Property Estimation Function", [online], June 1, 2013, searched on December 13, 2020, Internet <URL: https: / / www.pirika.com / Chemistry / JP / TCPE / YMB.html>

[0012] Non Patent Literature 2: Yamamoto, Hiroshi, "High-speed Liquid Chromatography and Hansen Solubility Parameter (HSP)", [online], September 9, 2011, searched on December 13, 2020, Internet <URL: https: / / www.pirika.com / HSP / JP / Examples / Docs / HPLC.html > SUMMARY

[0013] PROBLEMS TO BE SOLVED BY THE INVENTION

[0014] However, the existing technology is considered to lack convenience. That is, the estimation of HSP based on the method described in Non Patent Literature 1 requires the calculation of various physical quantities of the object component, and thus requires complex calculation. In addition, the precision of the retention time prediction method based on the method described in Non Patent Literature 2 is low, and does not reach a level that can be put into practical use.

[0015] The present invention was conceived in view of such actual circumstances, and aims to provide a technology that improves the convenience of the estimation of the retention time of an object component in chromatography.

[0016] SOLUTION TO THE PROBLEM

[0017] According to a certain aspect of the present disclosure, there is provided an estimation method that estimates the retention time of a specific object component in chromatography using a specific stationary phase and a specific mobile phase, the estimation method including the steps of: acquiring Hansen solubility parameters (HSP) of the specific object component; acquiring HSPs of the specific stationary phase and the specific mobile phase, respectively; and acquiring an estimation result of an index related to the retention time of the specific object component by inputting the HSPs of the specific object component, the specific stationary phase, and the specific mobile phase, respectively, into a machine learning model, wherein the machine learning model implements a machine learning process with the HSPs of the object component, the stationary phase, and the mobile phase, respectively, as explanatory variables, and with the index related to the retention time of the object component as a target variable.

[0018] According to another aspect of the present disclosure, there is provided an estimation method of estimating Hansen solubility parameters, HSPs, of a substance, the estimation method including the steps of: obtaining a SMILES expression of the substance; and obtaining an estimation result of the HSPs of the substance by inputting the SMILES expression of the substance into an estimation model that implements a machine learning process with the SMILES expression as an explanatory variable and the HSPs as a target variable.

[0019] According to still another aspect of the present disclosure, there is provided a computer program that causes a computer to implement the above-described estimation method by being executed by a control circuit of the computer.

[0020] According to still another aspect of the present disclosure, there is provided an information processing apparatus including: a control circuit; and a storage device that stores a computer program to be executed by the control circuit, wherein the computer program causes the information processing apparatus to implement the above-described estimation method by being executed by the control circuit.

[0021] Effects of Invention

[0022] According to one aspect of the present disclosure, there is provided a technique for improving the convenience of estimation of retention time in a chromatography by improving the accuracy of estimation of retention time of a target component or implementing calculation of HSPs simply. BRIEF DESCRIPTION OF DRAWINGS

[0023] Figure 1 is a diagram for explaining an outline of estimation of retention time according to one embodiment of the present disclosure.

[0024] Figure 2 is a diagram showing an example of a hardware structure of the information processing apparatus 1.

[0025] Figure 3 is a diagram showing an example of a structure of the HSP estimation model 2.

[0026] Figure 4 is a flowchart of an example of a process implemented in the information processing apparatus 1.

[0027] Figure 5 is a flowchart of a subroutine of step S20 of Figure 4

[0028] Figure 6 is a flowchart of a subroutine of step S40 of Figure 4

[0029] Figure 7 is a diagram for explaining a specific example of a SMILES expression utilized in the machine learning process of the HSP estimation model 2.

[0030] Figure 8 ​​is a graph for illustrating a result of estimation by the estimation model 2 for retention time.

[0031] Figure 9 is a graph for illustrating a result of estimation by the estimation model 3 for retention time.

[0032] Figure 10 is a graph in which a result of estimation by the estimation model 3 for retention time is shown together with results of estimation by other methods.

[0033] Figure 11 is a flowchart of another example of the processing implemented in the information processing apparatus 1.

[0034] Figure 12 is a flowchart of a subroutine of step S60 (estimation of retention time in gradient elution) of Figure 11

[0035] Figure 13 is a graph showing a specific example of output of a result of estimation of retention time.

[0036] Figure 14 is a flowchart of a subroutine of step S80 (report output) of Figure 11

[0037] Figure 15 is a graph showing an example of a report output.

[0038] Figure 16 is a flowchart of still another example of the processing implemented in the information processing apparatus 1. DETAILED DESCRIPTION

[0039] Hereinafter, an embodiment of the present disclosure will be explained in detail with reference to the accompanying drawings. Further, the same reference numerals are attached to the same or equivalent portions in the drawings, and the explanation thereof will not be repeated. Figure 1

[0040] [Outline of Estimation of Retention Time]

[0041] In the present embodiment, a value related to a retention time of a substance when the substance is eluted in a certain stationary phase using a certain mobile phase in chromatography is estimated. In the present specification, the term "target component" is sometimes used with respect to a substance that is an object of elution in chromatography.

[0042] Figure 1 is a graph for illustrating an outline of estimation of retention time according to one embodiment of the present disclosure. Figure 1 An information processing apparatus 1 is shown. The information processing apparatus 1 estimates a retention time of a certain target component in a column using a certain stationary phase and a certain mobile phase.

[0043] ​​​More specifically, the SMILES expression of a specific object component is input to the information processing apparatus 1. The SMILES expression of the object component can also be acquired from a database such as PubChem (https: / / pubchem.ncbi.nlm.nih.gov / ).

[0044] The information processing apparatus 1 includes an estimation model 2 for HSP and an estimation model 3 for retention time. The estimation model 2 for HSP performs a machine learning process with the SMILES expression as an explanatory variable and with the value of HSP as a target variable for a plurality of object components. The estimation model 3 for retention time performs a machine learning process with the HSP of the object component, the HSP of the stationary phase, and the HSP of the mobile phase as explanatory variables and with the logarithm of the retention factor (k') of the object component as a target variable. In addition, the explanatory variables of the estimation model 3 for retention time can also include other variables (for example, the molecular volume of the object component).

[0045] The information processing apparatus 1 acquires the estimation result of the HSP of the specific object component by applying the SMILES expression of the specific object component to the estimation model 2 for HSP.

[0046] The HSP of the specific object component, the HSP of the specific stationary phase, and the HSP of the specific mobile phase are input to the information processing apparatus 1. Each HSP can be either the estimation result output from the estimation model 2 for HSP or acquired from an information source other than the estimation model 2 for HSP.

[0047] The information processing apparatus 1 acquires the estimation result of the logarithm of the retention factor of the specific object component by applying the HSP of the specific object component, the HSP of the specific stationary phase, and the HSP of the specific mobile phase to the estimation model 3 for retention time. The information processing apparatus 1 acquires the estimation result of the retention factor from the logarithm of the retention factor (k'), and then derives (calculates) the estimation result of the retention time (t R ) of the specific object component using the estimation result of the retention factor and a given hold-up time (t0).

[0048] [Hardware structure]

[0049] Figure 2 is a diagram showing an example of the hardware structure of the information processing apparatus 1. The information processing apparatus 1 includes a processor 10 as an example of a control circuit, a memory 20 functioning as a storage section, and an input / output port 30. The input / output port 30 is connected to a mouse 40, a keyboard 50, and a display apparatus 60. The input / output port 30 can also be connected to one or a plurality of terminal apparatuses through the Internet or an intranet, or the like.

[0050] The information processing apparatus 1 is configured on the basis of, for example, a personal computer. The information processing apparatus 1 can also be constituted by a server that can be accessed from one or a plurality of terminal apparatuses through a network such as the Internet.

[0051] The analysis program 200, the HSP data 210, and the retention time data 220 are non-transitorily stored in the memory 20.

[0052] By executing the analysis program 200, the processor 10 functions as the first model creation section 201, the second model creation section 202, the first processing section 203, the second processing section 204, the image processing section 205, and the output section 206.

[0053] The first model creation section 201 implements machine learning processing of the HSP estimation model 2. The second model creation section 202 implements machine learning processing of the retention time estimation model 3. The first processing section 203 acquires an estimation result of the HSP of the object component by using the HSP estimation model 2. The second processing section 204 acquires an estimation result of the retention time of the object component by using the retention time estimation model 3. The image processing section 205 creates image data including various estimation results. The output section 206 outputs a display signal including the image data to the display apparatus 60 via the input-output port 30.

[0054] The HSP data 210 includes the estimation model 211. The estimation model 211 includes a data and parameter group for constituting the HSP estimation model 2. At least a part of the parameter group is updated in the machine learning processing.

[0055] The HSP data 210 includes a plurality of learning samples. The plurality of learning samples are classified into training data 212 and validation data 213. In one implementation example, 80% of the plurality of learning samples are classified into the training data 212, and 20% are classified into the validation data 213. Each of the plurality of learning samples includes a SMILES expression of an object component and a value of the HSP of the object component.

[0056] The retention time data 220 includes the estimation model 221. The estimation model 221 includes a data and parameter group for constituting the retention time estimation model 3. At least a part of the parameter group is updated in the machine learning processing.

[0057] The retention time data 220 includes a plurality of learning samples. The plurality of learning samples is classified into the training data 212 and the validation data 213. In one implementation example, 80% of the plurality of learning samples is classified into the training data 212, and 20% is classified into the validation data 213. Each of the plurality of learning samples includes the logarithm of the retention time of the object component in the HSP of the object component, the HSP of the stationary phase, the HSP of the mobile phase, and the combination of the stationary phase and the mobile phase.

[0058] [Estimation model for HSP]

[0059] Figure 3 is a diagram illustrating an example of the structure of the estimation model for HSP 2. Figure 3 The example of the HSP is in accordance with the SMILES2vec method. The HSP includes three parameters (a dispersion force term (dD), a polarity term (dP), and a hydrogen bond term (dH)). The estimation model for HSP 2 has settings corresponding to the three parameters, respectively. That is, the estimation model for HSP 2 is essentially constituted by three models, a model for the dispersion force term, a model for the polarity term, and a model for the hydrogen bond term.

[0060] The information processing apparatus 1 acquires the above three parameters as the HSP of the object component by providing the SMILES expression of the object component to the above three models, respectively.

[0061] Further, the estimation model for HSP 2 can also be generated for each column temperature (the set temperature of the column oven). More specifically, as the training data utilized in the generation of the estimation model for HSP 2, training data constituted by the SMILES expression and the set of parameters classified for each column temperature can also be utilized. The parameters classified for each column temperature refer to those associated with the column temperature at the time of acquisition of the parameters. The column temperature can also be one selected from among a plurality of temperature values decided in advance. More specifically, the column temperature can also be selected from among a plurality of temperature values at every 5°C (40°C, 35°C, 30°C, 25°C, 20°C,...). Each of the model for the dispersion force term, the model for the polarity term, and the model for the hydrogen bond term can also be constituted by a group corresponding to each of a plurality of column temperatures. For example, the group of the model for the dispersion force term can include models corresponding to five temperatures (a model for a column temperature of 40°C, a model for a column temperature of 35°C, a model for a column temperature of 30°C, a model for a column temperature of 25°C, and a model for a column temperature of 20°C). Similarly, the group of the model for the polarity term and the group of the model for the hydrogen bond term can also include models corresponding to five temperatures, respectively.

[0062] An example of the structure of each of the models constituting the estimation model for HSP 2 is illustrated in Figure 3

[0063] In​Figure 3 In the example of FIG. 1, input data (SMILES expression) is input to the HSP estimation model 2. The model to be used is selected based on the parameter (model for dispersion force term, model for polarity term, or model for hydrogen bond term) that is the object of estimation. The information processing apparatus 1 can also accept the setting of the column temperature that is the object of estimation from the user. In this case, the information processing apparatus 1 can also select the model corresponding to the column temperature that is the object of estimation as the model to be used. The input data is converted to a one-hot expression in the data processing section 2A.

[0064] Next, the one-hot expression is converted to a two-dimensional data on a character-by-character basis in the word embedding layer 2B. By this, the one-hot expression described above is converted to a two-dimensional data.

[0065] Next, the two-dimensional data described above is input to the LSTM (Long Short Term Memory) 2D, and further input to the LSTM 2E.

[0066] The data output from the LSTM 2E is converted to one-dimensional data in the data processing section 2F. The data output from the data processing section 2F is input to the neural network 2G. The Dropout layer for preventing overlearning is inserted in the neural network 2G.

[0067] The HSP estimation model 2 outputs the estimation result of the HSP based on the output of the neural network 2G.

[0068] The estimation of the HSP using the HSP estimation model 2 is different from the method of focusing on the physicochemical properties of each component to calculate the HSP of each component as disclosed in Non-Patent Literature 1, and uses a model that performs machine learning processing using SMILES expressions of a large number of components. By this, with respect to the estimation means of the HSP based on the HSP estimation model 2, it is expected to provide a more convenient means while the precision thereof is equivalent to the conventional method or the like.

[0069] [Retention time estimation model]

[0070] The retention time estimation model 3 is configured as a model according to the algorithm of the random forest, for example. The retention time estimation model 3 is generated by machine learning processing.

[0071] In the machine learning processing, the HSP of each of the object component, the stationary phase, and the mobile phase is used as an explanatory variable, and the logarithm of the retention factor of the object component is used as a target variable.

[0072] In the case where the mobile phase is a mixed solution of two components, the HSP of the mobile phase is prepared by combining the HSPs of the two components as shown in Expression (1).

[0073] [dDm, dPm, dHm] = [(a*dD1 + b*dD2), (a*dP1 + b*dP2), (a*dH1 + b*dH2)] / (a+b) (1)

[0074] In Equation (1), "dD" represents a dispersion force term in the HSP. "dP" represents a polarity term in the HSP. "dH" represents a hydrogen bond term in the HSP. The subscript "m" represents a mixed solution. The subscript "1" represents one of the two components, and the subscript "2" represents the other of the two components. "a" and "b" represent the volume fractions of the above one component and the above other component in the mixed solution. More specifically, the volume fraction of the above one component is "a / (a+b)", and the volume fraction of the above other component is "b / (a+b)".

[0075] Further, as the explanatory variables of the estimation model 3 of the retention time, in addition to the HSPs of the object component, the stationary phase, and the mobile phase, other variables can be used. The other variables can be, for example, one or more variables selected from the group consisting of the pH of the mobile phase, the column temperature in the column chromatography, and the molecular volume of the object component.

[0076] [Flow of processing]

[0077] [Main routine]

[0078] Figure 4 A flowchart of one example of the processing implemented in the information processing apparatus 1. In one implementation example, in the information processing apparatus 1, the processing of Figure 4 is implemented by execution of a given program by the processor 10. In one implementation example, the processing of Figure 4 is started by an operation for starting a given application program in the information processing apparatus 1.

[0079] In step S10, the information processing apparatus 1 determines whether or not an instruction of estimation of the HSP has been input. The information processing apparatus 1, when determining that the instruction of estimation of the HSP has been input (YES in step S10), causes the control to proceed to step S20, and otherwise (NO in step S10), causes the control to proceed to step S30.

[0080] In step S20, the information processing apparatus 1 implements the estimation of the HSP. The processing content of step S20 is described later with reference to Figure 5 In step S30, the information processing apparatus 1 determines whether or not an instruction of estimation of the retention time has been input. The information processing apparatus 1, when determining that the instruction of estimation of the retention time has been input (YES in step S30), causes the control to proceed to step S40, and otherwise (NO in step S30), causes the control to proceed to step S50.

[0081] In step S30, the information processing apparatus 1 determines whether or not an instruction of estimation of the retention time has been input. The information processing apparatus 1, when determining that the instruction of estimation of the retention time has been input (YES in step S30), causes the control to proceed to step S40, and otherwise (NO in step S30), causes the control to return to step S10.

[0082] In step S40, the information processing apparatus 1 implements estimation of the retention time. As for the processing content of step S40, refer to Figure 6 described later. Thereafter, the information processing apparatus 1 causes the control to return to step S10.

[0083] The information processing apparatus 1 can also accept input of both the instruction of estimation of the HSP and the instruction of estimation of the retention time with respect to a specific object component. For example, the information processing apparatus 1 accepts input of the SMILES expression of the specific object component, and values of explanatory variables of the estimation model 3 for the retention time other than the HSP of the object component (HSPs of a specific stationary phase, a specific mobile phase, and other variables required), and accepts input of both the instruction of estimation of the HSP and the instruction of estimation of the retention time.

[0084] <Estimation of HSP>

[0085] Figure 5 is Figure 4 the flowchart of the subroutine of step S20.

[0086] In step S200, the information processing apparatus 1 acquires data of the SMILES expression. In one implementation example, the data is input via the keyboard 50.

[0087] In step S202, the information processing apparatus 1 applies the SMILES expression to the estimation model 2 functioning as a model for the dispersion force term (model for dD).

[0088] In step S204, the information processing apparatus 1 acquires an estimation result of the parameter dD constituting the HSP.

[0089] In step S206, the information processing apparatus 1 applies the SMILES expression to the estimation model 2 functioning as a model for the polarity term (model for dP).

[0090] In step S208, the information processing apparatus 1 acquires an estimation result of the parameter dP constituting the HSP.

[0091] In step S210, the information processing apparatus 1 applies the SMILES expression to the estimation model 2 functioning as a model for the hydrogen bond term (model for dH).

[0092] In step S212, the information processing apparatus 1 acquires the estimation result of the parameter dH constituting the HSP.

[0093] In step S214, the information processing apparatus 1 outputs the estimation results of the three parameters acquired in steps S204, S208, and S212 as the estimation result of the HSP. An example of the output is display in the display apparatus 60.

[0094] After that, the information processing apparatus 1 returns the control to Figure 4 .

[0095] <Estimation of Retention Time>

[0096] Figure 6 is a flowchart of a subroutine of step S40 of Figure 4

[0097] In step S400, the information processing apparatus 1 acquires the HSP of the object component (a specific object component) that is the object of the estimation of the retention time. The information processing apparatus 1 can acquire the HSP of the object component by accepting input via the keyboard 50 or the like, or can acquire the HSP of the object component as the estimation result of the estimation model 2 for HSP.

[0098] In step S402, the information processing apparatus 1 acquires the HSP of the stationary phase (a specific stationary phase) that is expected to be used in the specific object component. The information processing apparatus 1 can acquire the HSP of the object component by accepting input via the keyboard 50 or the like, or can acquire the HSP of the object component as the estimation result of the estimation model 2 for HSP.

[0099] In step S404, the information processing apparatus 1 acquires the HSP of the mobile phase (a specific mobile phase) that is expected to be used in the specific object component. The information processing apparatus 1 can acquire the HSP of the object component by accepting input via the keyboard 50 or the like, or can acquire the HSP of the object component as the estimation result of the estimation model 2 for HSP.

[0100] Further, in the case where the mobile phase is a mixed solution of a plurality of solvents, the information processing apparatus 1 can also calculate the HSP of the mixed solution in step S404. More specifically, in step S404, input of the HSP and the volume fraction of each of the plurality of solvents is accepted, and the HSP of the mixed solution is calculated from them according to Equation (1).

[0101] In step S406, the information processing apparatus 1 acquires, as needed, the value of an explanatory variable other than the HSP (the pH of the specific mobile phase, the column temperature in column chromatography, and / or the molecular volume of the object component) that is used in the estimation model 3 for retention time.

[0102] ​In step S408, the information processing apparatus 1 applies the data acquired in steps S400, S402, S404 (, S406) to the retention time estimation model 3.

[0103] In step S410, the information processing apparatus 1 acquires the estimation result of the retention factor (of the logarithm) as a result of the application of the data in step S408.

[0104] In step S412, the information processing apparatus 1 derives the estimation result of the retention time of the specific target component using the retention factor acquired in step S410.

[0105] In step S414, the information processing apparatus 1 outputs the estimation result derived in step S412. An example of the output is a case where it is displayed on the display apparatus 60. Another example is a case where data is transmitted to an external apparatus. After that, the information processing apparatus 1 returns the control to step S400. Figure 4 .

[0106] In the processing described above, the retention time is output as the estimation result. Further, the target output as the estimation result can be the retention factor or the logarithm of the retention factor.

[0107] The information processing apparatus 1 can also perform the control of steps S400 to S412 for each of a plurality of target components. Thereby, the estimation result of the retention time of each of a plurality of target components in the chromatography using a common mobile phase and stationary phase is derived. Also, the information processing apparatus 1 can output the estimation results of a plurality of target components together in step S414.

[0108] In a case where the control of steps S400 to S412 is performed for isocratic elution, by outputting the estimation results of a plurality of target components together, the user can judge whether the elution of the plurality of target components should be performed as gradient elution in actual elution. For example, in a case where there is an appropriate difference in the retention times of a plurality of target components in the estimation result of isocratic elution, the user can judge that the actual elution can be performed as isocratic elution. On the other hand, in a case where there is no appropriate difference in the retention times of a plurality of target components in the estimation result of isocratic elution, the user can judge that the actual elution should be performed as gradient elution.

[0109] The information processing apparatus 1 can also perform the judgment of whether the actual elution can be performed as isocratic elution or should be performed as gradient elution as described above when outputting the estimation results of a plurality of target components, and output the result of the judgment together with the estimation results. The information processing apparatus 1 can also use a threshold value decided in advance in relation to the interval of the retention times in the judgment.

[0110] [EMBODIMENT]

[0111] <Estimation of HSP>

[0112] An embodiment of the training and utilization of the estimation model 2 for HSP will be described. In the machine learning processing of the estimation model 2 for HSP in this example, SMILES expressions of 1792 kinds of object components were utilized as learning samples. Figure 7 is a view for illustrating specific examples of SMILES expressions utilized in the machine learning processing of the estimation model 2 for HSP. In Figure 7 In the data shown in

[0113] In the estimation model 2 for HSP, in the machine learning processing of the model for the dispersion force term, 1792 data each consisting of a group of SMILES expressions and measured values of parameters of the dispersion force term were utilized as learning samples. In the machine learning processing of the model for the polarity term, 1792 data each consisting of a group of SMILES expressions and measured values of parameters of the polarity term were utilized as learning samples. In the machine learning processing of the model for the hydrogen bond term, 1792 data each consisting of a group of SMILES expressions and measured values of parameters of the hydrogen bond term were utilized as learning samples.

[0114] Figure 8 is a view for illustrating results of estimation performed using the estimation model 2 for HSP. In Figure 8 three graphs G11, G12, G13 are shown. Each graph in the graphs G11, G12, G13 represents each of the estimated values of dD, dP, dH. Each graph in the graphs G11, G12, G13 represents results with respect to 34 kinds of object components (benzyl alcohol, phenol, 3-phenylpropanol, p-chlorophenol, acetophenone, benzonitrile, nitrobenzene, methyl benzoate, anisole, benzene, p-nitrotoluene, p-nitrochlorobenzyl, toluene, benzophenone, bromobenzene, naphthalene, ethylbenzene, p-xylene, p-dichlorobenzene, propylbenzene, n-butylbenzene, diethylformamide, methyl p-hydroxybenzoate, ethyl p-hydroxybenzoate, propyl p-hydroxybenzoate, butyl p-hydroxybenzoate, acetanilide, phenylpropanone, butyrophenone, valerophenone, hexanophenone, heptanophenone, octylphenone, isobutylphenylpropanoic acid).

[0115] In each graph in the graphs G11, G12, G13, the horizontal axis represents an estimated value of a parameter based on the estimation model 2 for HSP, and the vertical axis represents a calculated value of the parameter acquired using the YMB simulator described in Non-Patent Literature 1. Lines L11, L12, L13 each represent a regression straight line. A blank plot represents an estimated value. A black plot represents a value on the regression straight line corresponding to the estimated value.

[0116] As shown in the graphs G11, G12, G13, the estimated value of HSP based on the estimation model 2 for HSP becomes a value close to the estimated value of HSP acquired using the YMB simulator.

[0117] Thus, the estimation means of the estimation model 2 for HSP can be said to be a simple means that replaces the YMB simulator that requires complex physicochemical calculations.

[0118] <Estimation of retention time>

[0119] An embodiment of training and utilization of the estimation model 3 for retention time will be described. In this example, as the estimation model 3 for retention time, a model according to an algorithm of random forest regression was utilized. In the machine learning processing of the estimation model 3 for retention time, 1020 kinds of data sets were utilized as learning samples. The 1020 kinds of data sets were composed of combinations of the above-described 34 kinds of object components and 30 kinds of mobile phases. The 30 kinds of mobile phases were mixed solutions of water and acetonitrile, and were composed so that the mixing ratio of acetonitrile was each different by 2% between 20% and 78%. An ODS (octadecylsilyl) column was assumed, and octadecane was used as the stationary phase.

[0120] The 1020 kinds of data sets each contained 10 kinds of variables (HSP (dD, dP, dH) of the object component, HSP (dD, dP, dH) of the stationary phase, HSP (dD, dP, dH) of the mobile phase, and the molecular volume of the object component) as explanatory variables, and contained the logarithm of the retention factor of the object substance as the target variable.

[0121] Figure 9 is a graph for illustrating the result of estimation using the estimation model 3 for retention time. In Figure 9 the graph shown in the graph G21, the vertical axis represents the logarithmic value (log k') of the retention factor that is an example of the estimation result based on the estimation model 3 for retention time, and the horizontal axis represents the measured value of the logarithm of the retention factor. The line L23 represents a regression straight line.

[0122] In the example of Figure 9 , the value of the root mean square error (RMSE) was 0.06482, and the value of the coefficient of determination (R 2 ) was 0.99345. Thus, the estimation of the estimation model 3 for retention time can be said to be an estimation with high precision.

[0123] Figure 10 is a graph in which the estimation result obtained using the estimation model 3 for retention time is shown together with the estimation results of other methods. In Figure 10 , three graphs G21, G22, G23 are shown. The graphs G21, G22 represent the estimation results of other methods shown for reference. The graph G23 isFigure 9 The graph shown is only a graph after being reduced. Lines L21, L22, L23 respectively indicate regression straight lines.

[0124] In graphs G21, G22, the vertical axis indicates measured values of the logarithm of the retention factor, and the horizontal axis indicates estimated values of the logarithm of the retention factor.

[0125] Graph G21 indicates estimated results for 34 kinds of objective components identical to those of graph G23, based on single regression analysis in accordance with the following formula (2).

[0126] logk' = C1 + C2 * molecular volume of objective component * (D1 - D2)... (2)

[0127] In formula (2), "C1" and "C2" indicate given constants. "D1" indicates the HSP distance of the mobile phase from the objective component. "D2" indicates the HSP distance of the stationary phase from the objective component. Further, with respect to the HSP distance (HSP_dis), in a case where one HSP is represented by a vector (dD1, dP1, dH1) and the other HSP is represented by a vector (dD2, dP2, dH2), it is calculated in accordance with the following formula (3).

[0128] HSP_dis = {4 * (dD1 - dD2) 2 + (dP1 - dP2) 2 + (dH1 - dH2) 2} 0.5 ... (3)

[0129] Graph G22 indicates estimated results for 34 kinds of objective components identical to those of graph G23, based on multiple regression analysis in accordance with the following formula (4).

[0130] logk' = C1 + C2 * molecular volume of objective component + C3 * D1 + C4 * D2... (4)

[0131] In formula (4), "C1", "C2", and "C3" indicate given constants. As with formula (2), "D1" indicates the HSP distance of the mobile phase from the objective component. As with formula (2), "D2" indicates the HSP distance of the stationary phase from the objective component.

[0132] In the example of graph G21, the value of the coefficient of determination R 2 was 0.862. In the example of graph G22, the value of the coefficient of determination R 2 was 0.92. On the other hand, in the example of graph G23, the value of the coefficient of determination R 2 was 0.99345, which was higher than graphs G21, G22.

[0133] In addition, in the example of the graph G23, as the explanatory variables, in addition to the HSPs of the object component, the stationary phase, and the mobile phase, the molecular volume of the object component is also utilized. In this regard, even in a case where the molecular volume is not utilized as an explanatory variable, even in a case where the pH and / or the column temperature of the mobile phase is utilized instead of the molecular volume as an explanatory variable, and further even in a case where the pH and / or the column temperature of the mobile phase is utilized as an explanatory variable in addition to the molecular volume, the same tendency is observed in the estimation results obtained by the estimation model 3 for the retention time with respect to the estimation results obtained by the single regression analysis and the multiple regression analysis.

[0134] Thus, it is considered that the estimation model 3 for the retention time is able to cope with the non-linear relationship between the HSP and the retention time, and is able to estimate the retention coefficient (the retention time) with higher precision than the above-mentioned single regression analysis and multiple regression analysis.

[0135] In addition, in a case where the model utilizing the algorithm according to the random forest is utilized as the estimation model 3 for the retention time, particularly, higher estimation precision is obtained.

[0136] In addition, the SMILES expression is able to distinguish isomers of a compound from each other. Thus, in a case where the HSP is derived using the SMILES expression, the HSP is able to be derived for each isomer of a compound, and thus, the estimation results of the retention time derived using the HSP are also able to be derived for each isomer of a compound. By utilizing the estimation results of the retention time derived like this, it is possible to determine the kind of isomers of a compound in a chromatography based on the difference in the retention time.

[0137] [Flow of processing (2)]

[0138] Figure 11 is a flowchart of another example of processing implemented in the information processing apparatus 1. Figure 11 The example of Figure 4 is different from the example of Figure 11 , and describes the estimation of the retention time in the isocratic elution and the estimation of the retention time in the gradient elution, respectively. Figure 11 The example of

[0139] Figure 11 The control of each of the steps S10 to S40 of the processing shown in Figure 4 is equivalent to the control of each of the steps S10 to S40 described with reference to Figure 11 In the example of Figure 11In the example of FIG. 8, in step S30, the information processing apparatus 1, when judging that the instruction of the estimation of the retention time in the isocratic elution has been input (YES in step S30), causes the control to proceed to step S40, and otherwise (NO in step S30), causes the control to proceed to step S50.

[0140] In step S40, the information processing apparatus 1 performs the estimation of the retention time in the isocratic elution. Thereafter, the information processing apparatus 1 causes the control to proceed to step S50.

[0141] In step S50, the information processing apparatus 1 judges whether or not the instruction of the estimation of the retention time in the gradient elution has been input. The information processing apparatus 1, when judging that the instruction of the estimation of the retention time in the gradient elution has been input (YES in step S50), causes the control to proceed to step S60, and otherwise (NO in step S50), causes the control to proceed to step S70.

[0142] In step S60, the information processing apparatus 1 performs the estimation of the retention time in the gradient elution. The content of the estimation of the retention time in the gradient elution in step S60 will be described later. Thereafter, the information processing apparatus 1 causes the control to proceed to step S70. Figure 12 In step S60, the information processing apparatus 1 performs the estimation of the retention time in the gradient elution. The content of the estimation of the retention time in the gradient elution in step S60 will be described later. Thereafter, the information processing apparatus 1 causes the control to proceed to step S70.

[0143] In step S70, the information processing apparatus 1 judges whether or not the instruction of the report output relating to the estimation result has been input. The information processing apparatus 1, when judging that the instruction of the report output has been input (YES in step S70), causes the control to proceed to step S80, and otherwise (NO in step S70), causes the control to return to step S10.

[0144] In step S80, the information processing apparatus 1 performs the report output. The content of the report output in step S80 will be described later. Thereafter, the information processing apparatus 1 causes the control to return to step S10. Figure 14 and Figure 15 In step S80, the information processing apparatus 1 performs the report output. The content of the report output in step S80 will be described later. Thereafter, the information processing apparatus 1 causes the control to return to step S10.

[0145] <Estimation of retention coefficient in gradient elution>

[0146] Figure 12 is Figure 11 a flowchart of a subroutine of step S60 (estimation of retention time in gradient elution) of FIG. 8.

[0147] In step S600, the information processing apparatus 1 acquires the HSP of the object component that is the object of the estimation of the retention time.

[0148] In step S602, the information processing apparatus 1 acquires the HSP of the stationary phase.

[0149] In step S604, the information processing apparatus 1 acquires, as needed, values of explanatory variables other than the HSP used in the retention time estimation model 3 (pH of a specific mobile phase, column temperature in column chromatography, and / or molecular volume of a target component).

[0150] In step S606, the information processing apparatus 1 sets "0" as the value of the variable T. The variable T represents a time for gradient elution. More specifically, the variable T represents the number of units of time corresponding to the length of the elapsed time from the start of elution. For example, in the case where 5 seconds are set as a unit of time, T = 2 indicates that the elapsed time is 10 seconds from the start of elution. That is, T = 2 indicates the timing 10 seconds after the start of elution.

[0151] In gradient elution, the proportions of a plurality of solutions in the mobile phase change with the passage of time. In the information processing apparatus 1, information for determining the change in the proportions of a plurality of solutions in the mobile phase with the passage of time is registered as a condition for estimation. The information is input by a user, for example.

[0152] An example of the above information is information that specifies the proportions of a first solution and a second solution per unit of time in the case where the first solution and the second solution are used as the mobile phase.

[0153] For example, as the above information, it is assumed that the change in the proportions is started 10 seconds after the start of elution, the change in the proportions every 5 seconds is 5%, the initial proportion of the first solution is 100%, and the initial proportion of the second solution is 0%. At this time, from the start of elution to 10 seconds later, the details of the mobile phase are as follows: the first solution is 100% and the second solution is 0%. The details of the mobile phase 15 seconds after the start of elution are as follows: the first solution is 95% and the second solution is 5%. The details of the mobile phase 20 seconds after the start of elution are as follows: the first solution is 90% and the second solution is 10%. The details of the mobile phase 25 seconds after the start of elution are as follows: the first solution is 85% and the second solution is 15%. In addition, the details of the mobile phase 105 seconds after the start of elution are as follows: the first solution is 5% and the second solution is 95%. Furthermore, the details of the mobile phase after 110 seconds from the start of elution are as follows: the first solution is 100% and the second solution is 0%.

[0154] In step S608, the information processing apparatus 1 acquires the HSP of the mobile phase at the timing indicated by the variable T. More specifically, the information processing apparatus 1 determines the proportions of a plurality of solutions in the mobile phase at the timing indicated by the variable T in accordance with the above-described condition for estimation. Furthermore, the information processing apparatus 1 acquires, as the HSP of the mobile phase, the HSP of a mixed solution obtained by mixing a plurality of solutions at the determined proportions in accordance with the above-described formula (1).

[0155] In step S610, the information processing apparatus 1 applies the data acquired in steps S600, S602, S604, and S608 to the retention time estimation model 3.

[0156] In step S612, the information processing apparatus 1 acquires the estimation result of the retention coefficient (of the logarithm) as a result of the application of the data in step S610.

[0157] In step S614, the information processing apparatus 1 derives the estimation result of the retention time of the specific object component using the retention coefficient acquired in step S610.

[0158] In step S616, the information processing apparatus 1 calculates the moving distance of the object component in the column in the latest unit time using the estimation result derived in step S614. More specifically, the information processing apparatus 1 calculates the time related to the movement in the column using the retention coefficient, and calculates the moving speed of the retention coefficient from the calculated time and the length of the column. Then, the information processing apparatus 1 calculates the product of the moving speed and the unit time as the moving distance.

[0159] In step S618, the information processing apparatus 1 determines whether the cumulative value of the moving distance of the object component is equal to or more than the length of the column. The information processing apparatus 1 makes the control proceed to step S622 when it is determined that the cumulative value of the moving distance of the object component is equal to or more than the length of the column (YES in step S618), and makes the control proceed to step S620 otherwise (NO in step S618).

[0160] In step S620, the information processing apparatus 1 updates the value of the variable T by 1, and makes the control return to step S606.

[0161] In step S622, the information processing apparatus 1 determines the time derived as the product of the latest value of the variable T and the unit time as the estimation result of the retention time. Further, the product of the latest value of the variable T and the unit time corresponds to the cumulative value of the unit time when the cumulative value of the moving distance of the object component reaches the length of the column.

[0162] In step S624, the information processing apparatus 1 outputs the estimation result, and makes the control return to Figure 11 .

[0163] In the processing described above Figure 12 , the estimation result of the retention time with respect to the gradient elution is output.

[0164] <Detail of Output of Estimation Result of Retention Time>

[0165] Figure 13is a specific example of an output showing the estimation result of the retention time. The information processing apparatus 1 can also implement the estimation of the retention time by implementing the control of steps S600 to S622 on each of the plurality of object components. Also, the information processing apparatus 1 can also generate the display information of one screen using the estimation results of the retention times of the plurality of object components in step S624. In Figure 13 The estimation results for the plurality of object components are shown in the screen 600. The screen 600 is displayed on the display apparatus 60, for example. In addition, information input via the mouse 40 or the keyboard 50 is displayed in the screen 600.

[0166] The screen 600 includes four regions 610, 620, 630, 640. The region 610 displays information for determining the object components. The region 620 displays the conditions of the elution used in the estimation. The region 630 displays the chromatograph made using the estimation results. The region 640 displays the numerical values of the estimation results.

[0167] In one implementation example, an instruction of the estimation of the retention time in the gradient elution can also be input in a state where the object components are input to the region 610 and the conditions are input to the region 620. The information processing apparatus 1 can also start the process of Figure 12 according to the input of the instruction. Also, the information processing apparatus 1 can update the display of the screen 600 in a manner that the estimation results determined in the process of Figure 12 are added, thereby outputting the estimation results.

[0168] The display contents of each of the regions 610 to 640 of Figure 13 will be described. In the region 610, as the five object components, the names of the five compounds (phenol, benzonitrile, p-chlorophenol, phenylacetone, nitrobenzene) are shown. In the region 610, the values of the HSP (dD, dP, dH) and the molecular volume (Volume 3D) of each object component are also shown.

[0169] The conditions displayed in the area 620 include a mode of separation (Separation Mode: Reversed Phase), a kind of stationary phase (Stationary Phase: SB-C18), kinds of two solutions used as mobile phases (Mobile Phase A: Acetonitrile, Mobile Phase B: Water), a proportion of solution A when isocratic elution (Mobile Phase A%: 60), a flow rate of mobile phase (Flow rate (ml / min): 20), and a column temperature (Temperature (°C): 40). In the area 620, conditions for estimation (conditions mentioned in association with the explanation of the variable T of step S606) can also be displayed.

[0170] In the area 630, in addition to displaying a chromatogram including peaks (five peaks P11 to P15) corresponding to each of the target components, a line L11 representing a change in a proportion of one of the two solutions in the mobile phase is displayed. The chromatogram is generated by the information processing apparatus 1 in a manner that includes a peak representing a retention time determined as a result of estimation of each of the target components.

[0171] In the area 640, a value of a retention coefficient (k') is shown together with a retention time (Retention Time) of each of the target components.

[0172] <Report Output>

[0173] Figure 14 Yes Figure 11 is a flowchart of a subroutine of step S80 (report output) of

[0174] In step S800, the information processing apparatus 1 acquires information of an output target. In one implementation example, the information of the output target includes information for determining each of one or more target components, information for determining a mobile phase, information for determining a stationary phase, and conditions utilized in estimation.

[0175] In step S802, the information processing apparatus 1 generates a report including estimation results of retention times for one or more target components, based on the information of the output target acquired in step S800.

[0176] In step S804, the information processing apparatus 1 outputs the report as a result of step S802. After that, the information processing apparatus 1 returns the control to Figure 11 .

[0177] Figure 15is a drawing showing an example of the output report. The report 700 includes areas 710, 720, 730, 740, 750, 760.

[0178] In the area 710, a directory item is displayed. In one example, the directory item includes a name (Aromatic 20230120) annotated by the user to the output report, a user ID (440275), and a user name (Yasuhiro Mito).

[0179] In the area 720, conditions utilized in the estimation of the retention time are displayed. In the example of Figure 15 "Separation Mode" indicates a mode of elution, and "Reversed Phase" is displayed as a value thereof. "Stationary Phase" indicates a stationary phase, and "SB-C18" is displayed as a value thereof. In the example of Figure 15 In the example, additional information (HSP values (δD: 15.91, δP: 0.1, δH: 0.1), Partical size [μm] = 5, id (mm): 4.6, Length (mm) = 100.0, Inter porosity: 0.6, Intra porosity: 0.6, N: 16000) is also displayed for the stationary phase. In particular, Length refers to a length of a region in which the stationary phase is filled, i.e., a length of a column.

[0180] "Mobile Phase A" indicates one of a plurality of solutions constituting a mobile phase, and "acetonitrile" is displayed as a value thereof. In the example of Figure 15 In the example, additional information (HSP values (δD: 15.61, δP: 16.64, δH: 8.32)) is also displayed for the solution.

[0181] "Mobile Phase B" indicates the other of the plurality of solutions constituting the mobile phase, and "water" is displayed as a value thereof. In the example of Figure 15 In the example, additional information (HSP values (δD: 7.58, δP: 7.82, δH: 20.68)) is also displayed for the solution.

[0182] "Mobile Phase A%" indicates the proportion of Solution A in the case of isocratic elution, and "60" is displayed as its value. "Flow rate (ml / min)" indicates the flow rate of the mobile phase, and "2" is displayed as its value. "Temperature (°C)" indicates the column temperature, and "40" is displayed as its value. "Gradient settings" indicates the settings regarding gradient elution, and as its value, the initial value of the proportion of the mobile phase of Solution A (the kind of solution indicated by "Mobile Phase A" described above) "Initial A%: 20", the final value of the proportion of Solution A in the mobile phase "Final A%: 90", the start and end of the change in the proportion of the plurality of solutions in the mobile phase with respect to the elution start time (Start time (min): 11, End time (min): 20), and the kind of variable utilized in the estimation (Compound information [HSP and Volume 3D]) are displayed.

[0183] A table is displayed in the region 730, which contains the values of the HSP values and the molecular volumes with respect to each of the one or more target components that are the estimation targets.

[0184] A simulated chromatogram generated using the estimation results is displayed in the region 740. This chromatogram is generated as a peak having the estimation result of the retention time of each of the one or more target components, as with the region 640 of Figure 13 A broken line indicating the change in the proportion of one of the solutions in the mobile phase composed of a plurality of solutions is also displayed in the region 740.

[0185] The estimation results with respect to each of the one or more target components that are the estimation targets are displayed in the region 750. The estimation results contain the retention time (RT) and the retention coefficient (k'). The SMILES expression (Smiles), the molecular formula, and the molecular weight of each target component are also displayed in the region 740.

[0186] The structural formula of each of the one or more target components that are the estimation targets is displayed in the region 760.

[0187] As described above, the output report contains a simulated chromatogram generated using the retention times of the one or more target components. Thereby, the estimation results are provided to the user in an easily understandable manner. The output report contains the structural formula of each target component in addition to the retention time as the estimation result with respect to the plurality of target components. Thereby, the user can evaluate the difference in the estimation results of the retention times between the plurality of target components while considering the differences in the structural formulas of the plurality of target components.

[0188] In the report output of the above-described example, the information processing apparatus 1 can utilize the estimation result of the retention time that is saved in advance in the storage 20, or can perform the estimation of the retention time. When the estimation of the retention time is performed in the report output, the information processing apparatus 1 performs the estimation of the retention time in step S40 and / or step S60 in response to the information of the output object acquired in step S600, that is, in step S602. Figure 14

[0189] In the report output of the above-described example, the information processing apparatus 1 can acquire both of the estimation conditions related to both of the isocratic elution and the gradient elution as the information of the output object. In this case, the information processing apparatus 1 generates a report using both of the estimation result of the retention time in the isocratic elution and the estimation result of the retention time in the gradient elution for each of the one or more target components. In the report generated like this, the conditions related to both of the isocratic elution and the gradient elution are displayed in the region 720. Further, a chromatogram generated using the retention time in the isocratic elution, a chromatogram generated using the retention time in the gradient elution, and a gradient curve indicating the change in the proportion of one solution in the mobile phase are displayed in the region 740. Figure 14 [Flow of processing (3)]

[0190]

[0191] Figure 16 This is a flowchart of another example of the processing performed in the information processing apparatus 1. Figure 16 The processing of the above-described example is performed in order to provide information related to the appropriateness of each of the one or more candidates that is an unknown component. In one implementation example, the information processing apparatus 1 starts the processing of the above-described example in response to a specific menu being selected in the execution of a given application program. Figure 16

[0192] In step S900, the information processing apparatus 1 acquires an analysis result of the chromatography of the unknown component. The analysis result is a measured value obtained by the chromatography. The analysis result can be input from the user to the information processing apparatus 1 via the mouse 40 or the keyboard 50, or can be transmitted to the information processing apparatus 1 from an external device.

[0193] In step S902, the information processing apparatus 1 determines one or more candidates of the compound assumed for the unknown component. The information processing apparatus 1 can also accept the input of one or more candidates from the user, and achieve the determination of the one or more candidates by recognizing the input one or more candidates.

[0194] ​​​The information processing apparatus 1 can also achieve the determination of one or more candidates by selecting one or more compounds from among the plurality of compounds registered in the storage 20 based on the result of the mass spectrometric analysis of the unknown component. For example, the information processing apparatus 1 acquires a mass spectrum as the result of the mass spectrometric analysis, determines peaks in the mass spectrum, and determines m / z corresponding to the determined peaks. Then, the information processing apparatus 1 selects one or more compounds having a molecular weight of the determined m / z value from among the plurality of compounds registered in the storage 20 as one or more candidates of the unknown component.

[0195] In step S904, the information processing apparatus 1 determines an estimation result of the retention time of each of the one or more candidates determined in step S902. The estimation result is in the same conditions as the analysis result acquired in step S900 (e.g., the same mobile phase and stationary phase). The information processing apparatus 1 can also acquire the conditions of the chromatography used to obtain the analysis result in step S900 in addition to the analysis result. The information processing apparatus 1 can also perform the estimation of the retention time in step S40 and / or step S60 for each of the one or more candidates in step S904.

[0196] In step S906, the information processing apparatus 1 compares the analysis result acquired in step S900 with the estimation result of the retention time of each of the candidates determined in step S904 to determine the propriety of each of the candidates as the unknown component for each of the candidates. In one example, the information processing apparatus 1 can also calculate the difference between the estimation result and the analysis result acquired in step S900 for each of the candidates as the propriety described above and determine the rank in order from small to large of the difference. In another example, the compound for which the difference between the estimation result and the analysis result acquired in step S900 is smallest can be classified as the final candidate and the remaining compounds can be classified as outside the candidates, whereby the propriety described above can be determined.

[0197] In step S908, the information processing apparatus 1 outputs the result of the determination of the propriety in step S906. Thereafter, the information processing apparatus 1 ends the processing. Figure 16

[0198] As described above, in the processing of Figure 16 , the propriety of each of the one or more candidates as the unknown component is outputted. The analysis result of the chromatography of the unknown component and the estimation result of the retention time of the one or more candidates are utilized in the determination of the propriety. The user can refer to the propriety of each of the one or more candidates outputted to identify the unknown component.

[0199] [MODE]

[0200] ​The person skilled in the art understands that the above-described exemplary embodiments are specific examples of the following modes.

[0201] The estimation method according to the first item can also be a method of estimating a retention time of a specific target component in a chromatography using a specific stationary phase and a specific mobile phase, the estimation method including the steps of: acquiring an HSP of the specific target component; acquiring HSPs of the specific stationary phase and the specific mobile phase, respectively; and acquiring an estimation result of an index related to the retention time of the specific target component by inputting the HSPs of the specific target component, the specific stationary phase, and the specific mobile phase, respectively, into a machine learning model that implements a machine learning process with the HSPs of the target component, the stationary phase, and the mobile phase, respectively, as explanatory variables and with the index related to the retention time of the target component as a target variable.

[0202] According to the estimation method according to the first item, a technique of improving the accuracy of the estimation of the retention time of the target component in the chromatography is provided. Thereby, the convenience of the estimation of the retention time is improved.

[0203] The estimation method according to the second item can also be the estimation method according to the first item in which the machine learning model is a classification model based on a random forest.

[0204] According to the estimation method according to the second item, a technique of more reliably improving the accuracy of the estimation of the retention time of the target component in the chromatography is provided.

[0205] The estimation method according to the third item can also be the estimation method according to the first item or the second item in which the index is a logarithm of a retention coefficient.

[0206] According to the estimation method according to the third item, the estimation of the retention time can be more easily implemented.

[0207] The estimation method according to any one of the first item to the third item can also be the estimation method in which the machine learning model further implements the machine learning process using a pH of the mobile phase as an explanatory variable, the estimation method further including the steps of: acquiring the pH of the specific mobile phase, and in the step of acquiring the estimation result, further inputting the pH of the specific mobile phase into the machine learning model.

[0208] According to the estimation method according to the fourth item, a technique of more reliably improving the accuracy of the estimation of the retention time of the target component in the chromatography is provided.

[0209] (Fifth) In any of the estimation methods described in the first to fourth items, the machine learning model can also perform machine learning processing using a column temperature as an explanatory variable, and the estimation method can further include the step of acquiring a column temperature using the specific stationary phase, and the column temperature using the specific stationary phase can be input to the machine learning model in the step of acquiring the estimation result.

[0210] According to the estimation method described in the fifth item, a technique for more reliably improving the accuracy of estimation of the retention time of the target component in the chromatography is provided.

[0211] (Sixth) In any of the estimation methods described in the first to fifth items, the machine learning model can also perform machine learning processing using a molecular volume of the target component as an explanatory variable, and the estimation method can further include the step of acquiring a molecular volume of the specific target component, and the molecular volume of the specific target component can be input to the machine learning model in the step of acquiring the estimation result.

[0212] According to the estimation method described in the sixth item, a technique for more reliably improving the accuracy of estimation of the retention time of the target component in the chromatography is provided.

[0213] (Seventh) In any of the estimation methods described in the first to sixth items, the step of acquiring the HSP of the specific target component can include the process of acquiring a SMILES expression of the specific target component, and acquiring an estimation result of the HSP of the specific target component by inputting the SMILES expression of the specific target component to an estimation model that performs machine learning processing with a SMILES expression as an explanatory variable and with an HSP as a target variable, and the estimation result of the HSP of the specific target component can be used as the HSP of the specific target component.

[0214] According to the estimation method described in the seventh item, the estimation of the HSP is easily performed in the estimation of the retention time of the target component in the chromatography.

[0215] (Eighth) In the estimation method described in the seventh item, the HSP can include a plurality of parameters, and in the step of acquiring the HSP of the specific target component, each parameter of the plurality of parameters can be estimated by applying a setting corresponding to each parameter of the plurality of parameters to the estimation model.

[0216] According to the estimation method described in the eighth item, a technique for more reliably improving the accuracy of estimation of the retention time of the target component in the chromatography is provided.

[0217] (Ninth) In the estimation method of the seventh or eighth item, the estimation model can also be generated for each of a plurality of column temperatures, and the step of acquiring the SMILES expression of the specific target component can include the following process: receiving a setting of a column temperature that is a target of estimation, and the step of acquiring the estimation result of the HSP of the specific target component can include the following process: selecting an estimation model corresponding to the column temperature that is the target of estimation from among the estimation models for each of the plurality of column temperatures.

[0218] According to the estimation method of the ninth item, as a model used in the estimation of the HSP, a model tailored to the column temperature that is the target of estimation is used. Thus, the accuracy of the estimation of the HSP is improved.

[0219] (Tenth) In the estimation method of any one of the first to ninth items, the step of acquiring the HSP of the specific target component can include the following process: further acquiring HSPs of one or more other target components, the step of acquiring the estimation result of the index related to the retention time of the specific target component can include the following process: further acquiring estimation results of the index of the one or more other target components, and the estimation method can further include the step of outputting the estimation results of the index of each of the specific target component and the one or more other target components.

[0220] According to the estimation method of the tenth item, estimation results of retention times under common conditions are provided for a plurality of components.

[0221] (Eleventh) In the estimation method of the tenth item, the step of outputting can include the following process: determining the necessity of gradient elution based on the estimation results of the index of each of the specific target component and the one or more other target components, and outputting the result of the determination.

[0222] According to the estimation method of the eleventh item, information for assisting the determination of the necessity of gradient elution is provided to the user.

[0223] (Twelfth) In the estimation method of the tenth or eleventh item, the step of outputting can include the following process: outputting the retention time as the estimation result of the index.

[0224] According to the estimation method of the twelfth item, information that is easily and intuitively understood with respect to chromatography is provided to the user.

[0225] (Thirteenth) In the estimation method of the twelfth item, the retention time can include a retention time in isocratic elution and a retention time in gradient elution.

[0226] According to the estimation method described in the thirteenth aspect, information for determining whether isoprocate elution can be adopted as elution of the plurality of target components or gradient elution should be adopted as elution of the plurality of target components is provided to the user.

[0227] (14) In the estimation method described in the twelfth or thirteenth aspect, the retention time can be output in the form of a chromatogram.

[0228] According to the estimation method described in the fourteenth aspect, the estimation result of the retention time is provided to the user in a manner that is easily and intuitively understood with respect to chromatography.

[0229] (15) In the estimation method described in the fourteenth aspect, the step of outputting can include processing of outputting conditions for obtaining the estimation result of the index in the same screen as the retention time.

[0230] According to the estimation method described in the fifteenth aspect, the estimation result and information that can be utilized in evaluation of the estimation result can be provided to the user.

[0231] (16) In any one of the estimation methods described in the twelfth through fifteenth aspects, the step of outputting can include processing of outputting a structural formula of a compound of each of the specific target component and the one or more other target components.

[0232] According to the estimation method described in the sixteenth aspect, the information that can be utilized in evaluation of the estimation result can be provided to the user together with the estimation result in a manner that is easily and intuitively understood.

[0233] (17) In any one of the estimation methods described in the first through sixteenth aspects, the specific mobile phase can include a plurality of solvents, a ratio of each of the plurality of solvents in the specific mobile phase can be set for each of a plurality of unit times, the estimation result of the index can be a retention time, the step of obtaining the estimation result of the index of the specific target component can include processing of selecting one of the plurality of unit times, obtaining the estimation result of the index for the selected one of the plurality of unit times, calculating a moving distance of the specific target component in a column based on the estimation result of the index obtained for the selected one of the plurality of unit times, and implementing switching of the selected one of the plurality of unit times, in the step of obtaining the estimation result of the index of the specific target component, the switching is implemented until a cumulative value of the moving distance reaches a predetermined length of the column, and a cumulative value of the unit time at which the cumulative value of the moving distance reaches the predetermined length of the column is obtained as a final estimation result.

[0234] According to the estimation method according to the seventeenth aspect, a specific manner of estimation of the retention time in the gradient elution is provided.

[0235] (A twenty-first aspect) In the estimation method according to the twentieth aspect, the HSP can include a plurality of parameters, and in the step of acquiring the HSP of the substance, each parameter of the plurality of parameters can be estimated by applying a setting corresponding to each parameter of the plurality of parameters to the estimation model.

[0236] According to the estimation method according to the eighteenth aspect, information useful for identification of the unknown component can be provided.

[0237] (A nineteenth aspect) In the estimation method according to the eighteenth aspect, the step of determining the one or more candidates can include a process of selecting the one or more candidates based on a result of mass spectrometric analysis of the unknown component.

[0238] According to the estimation method according to the nineteenth aspect, a burden on the user for determining the one or more candidates can be alleviated.

[0239] (A twentieth aspect) An estimation method according to another aspect can be a method of estimating HSP of a substance, the estimation method including a step of acquiring a SMILES expression of the substance, and a step of acquiring an estimation result of the HSP of the substance by inputting the SMILES expression of the substance into an estimation model, wherein the estimation model implements a machine learning process with the SMILES expression as an explanatory variable and the HSP as a target variable.

[0240] According to the estimation method according to the twentieth aspect, estimation of the HSP utilized in estimation of the retention time becomes easy, and thus a technique of improving convenience in estimation of the retention time of the target component in the chromatography is provided.

[0241] (A twenty-first aspect) In the estimation method according to the twentieth aspect, the HSP can include a plurality of parameters, and in the step of acquiring the HSP of the substance, each parameter of the plurality of parameters can be estimated by applying a setting corresponding to each parameter of the plurality of parameters to the estimation model.

[0242] According to the estimation method according to the twenty-first aspect, a technique of more reliably improving accuracy in estimation of the retention time of the target component in the chromatography is provided.

[0243] (22) In the estimation method described in (20) or (21), the estimation model can be generated for each column temperature, and the step of acquiring the SMILES representation of the substance can include a process of receiving a setting of a temperature to be an estimation target, and the step of acquiring the estimation result of the HSP of the substance can include a process of selecting an estimation model corresponding to the column temperature to be the estimation target.

[0244] According to the estimation method described in (22), as a model used in the estimation of the HSP, a model customized for a temperature to be an estimation target is used. Thus, the accuracy of the estimation of the HSP is improved.

[0245] (23) A computer program according to one embodiment can also be executed by a control circuit of a computer to cause the computer to implement the estimation method described in any one of (1) to (22).

[0246] According to the estimation method described in (23), a technique for improving the accuracy of the estimation of the retention time of a target component in chromatography or improving the convenience of the estimation of the HSP to improve the convenience of the estimation of the retention time of a target component in chromatography is provided.

[0247] (24) An information processing apparatus according to one embodiment includes a control circuit and a storage device that stores a computer program to be executed by the control circuit, wherein the computer program, when executed by the control circuit, causes the information processing apparatus to implement the estimation method described in any one of (1) to (22).

[0248] According to the information processing apparatus described in (24), a technique for improving the accuracy of the estimation of the retention time of a target component in chromatography or improving the convenience of the estimation of the HSP to improve the convenience of the estimation of the retention time of a target component in chromatography is provided.

[0249] It should be understood that the embodiments disclosed herein are illustrative and not restrictive in all aspects. The scope of the disclosure is not represented by the description of the above-described embodiments but by the claims, and is intended to include all modifications within the meaning and range equivalent to the claims. In addition, each of the technologies in the embodiments can be intended to be able to be implemented alone, and in addition, can be intended to be able to be implemented in combination with other technologies in the embodiments as much as possible as needed.

[0250] Explanation of Reference Signs

[0251] 1: information processing apparatus; 2: estimation model for HSP; 3: estimation model for retention time; 10: processor; 20: memory.

Claims

1. A method for estimating the retention time of a specific target component in a chromatography method using a specific stationary phase and a specific mobile phase, the method comprising the following steps: Obtaining the Hansen Solubility Parameter (HSP) of the specific target component; obtaining HSPs of the specific stationary phase and the specific mobile phase; as well as By inputting the HSP of each of the specific target component, the specific stationary phase, and the specific mobile phase into a machine learning model, an estimation result of an index related to the retention time of the specific target component is obtained, The machine learning model implements a machine learning process using the HSP of the target component, the stationary phase, and the mobile phase as explanatory variables and an index related to the retention time of the target component as a target variable.

2. The estimation method according to claim 1, wherein: The machine learning model is a random forest-based classification model.

3. The estimation method according to claim 1, wherein: The index is the logarithm of the retention coefficient.

4. The estimation method according to claim 1, wherein: The machine learning model also uses the pH of the mobile phase as an explanatory variable to implement machine learning processing. The estimation method further comprises the following steps: obtaining the pH of the specific mobile phase, In the step of obtaining the estimation result, the pH of the specific mobile phase is also input into the machine learning model.

5. The estimation method according to claim 1, wherein: The machine learning model also implements a machine learning process using column temperature as an explanatory variable. The estimation method further comprises the steps of: obtaining the column temperature using the specific stationary phase, In the step of obtaining the estimation result, the column temperature using the specific stationary phase is also input into the machine learning model. The estimation method according to claim 1 , wherein: The machine learning model also uses the molecular volume of the object component as an explanatory variable to implement the machine learning process. The estimation method further comprises the following steps: obtaining the molecular volume of the specific object component, In the step of obtaining the estimation result, the molecular volume of the specific object component is also input into the machine learning model.

7. The estimation method according to claim 1, wherein: The step of obtaining the HSP of the specific target component includes the following processing: Obtaining a SMILES expression of the specific object component; as well as By inputting the SMILES expression of the specific target component into the estimation model, an estimation result of the HSP of the specific target component is obtained, wherein the estimation result of the HSP of the specific target component is used as the HSP of the specific target component, The estimation model implements a machine learning process using SMILES expressions as explanatory variables and HSP as target variables.

8. The estimation method according to claim 7, wherein: The HSP includes multiple parameters, In the step of acquiring the HSP of the specific target component, each of the plurality of parameters is estimated by applying a setting corresponding to each of the plurality of parameters to the estimation model.

9. The estimation method according to claim 7, wherein: The estimation model is generated for a plurality of column temperatures respectively. The process of obtaining the SMILES expression of the specific object component includes the following processes: Accepts the setting of the column temperature to be estimated, The process of acquiring the estimation result of the HSP of the specific target component includes a process of selecting an estimation model corresponding to the column temperature to be estimated from the estimation models for each of the plurality of column temperatures.

10. The estimation method according to claim 1, wherein: The step of acquiring the HSP of the specific target component includes the following steps: further acquiring the HSP of one or more other target components, The step of acquiring the estimation result of the index related to the retention time of the specific target component includes the following processing: further acquiring the estimation results of the index of the one or more other target components, The estimation method further includes the step of outputting an estimation result of the index for the specific target component and each of the one or more other target components.

11. The estimation method according to claim 10, wherein: The step of outputting the output includes the following processing: performing a determination of the necessity of gradient elution based on the estimation result of the index of the specific target component and each of the one or more other target components; as well as The result of the judgment is output.

12. The estimation method according to claim 10, wherein: The step of outputting includes a process of outputting the retention time as an estimation result of the index.

13. The estimation method according to claim 12, wherein: The retention time includes the retention time in isocratic elution and the retention time in gradient elution.

14. The estimation method according to claim 12, wherein: The retention times are output in the form of a chromatogram.

15. The estimation method according to claim 14, wherein: The step of outputting includes a process of outputting conditions for obtaining the estimation result of the index in a manner displayed on the same screen as the retention time.

16. The estimation method according to claim 12, wherein: The step of outputting includes a process of outputting a structural formula of a compound of the specific target component and each of the one or more other target components.

17. The estimation method according to claim 1, wherein: The specific mobile phase includes a variety of solvents, The ratio of each of the plurality of solvents in the specific mobile phase is set for each of a plurality of unit times, The estimated result of the indicator is the retention time, The step of obtaining the estimation result of the index of the specific object component includes the following processing: selecting one unit time among the plurality of unit times; Obtaining an estimation result of the indicator for the selected unit time; calculating a movement distance of the specific target component in the column based on an estimation result of the index acquired for the selected one unit time; as well as implementing the selected switching of the one unit time, In the step of obtaining the estimation result of the index of the specific target component, The switching is performed until the cumulative value of the movement distance reaches a predetermined column length, The cumulative value per unit time when the cumulative value of the moving distance reaches a predetermined column length is acquired as a final estimation result.

18. The estimation method according to claim 1, further comprising the following steps: Obtaining chromatographic analysis results of unknown components; as well as determining one or more candidates of the unknown component as the specific target component, In the chromatography of the unknown component, the specific stationary phase and the specific mobile phase are used, The estimation method further includes the step of determining, for each of the one or more candidates, validity of the one or more candidates being the unknown component based on the analysis result and the estimation result of the indicator corresponding to each of the one or more candidates.

19. The estimation method according to claim 18, wherein: The step of determining the one or more candidates includes a process of selecting the one or more candidates based on a result of mass spectrometry analysis of the unknown component.

20. A method for estimating the Hansen Solubility Parameter (HSP) of a substance, comprising the following steps: Obtaining a SMILES expression of the substance; as well as By inputting the SMILES expression of the substance into the estimation model, the estimation result of the HSP of the substance is obtained, The estimation model implements a machine learning process using SMILES expressions as explanatory variables and HSP as target variables.

21. The estimation method according to claim 20, wherein: The HSP includes multiple parameters, In the step of acquiring the HSP of the substance, each of the plurality of parameters is estimated by applying a setting corresponding to each of the plurality of parameters to the estimation model.

22. The estimation method according to claim 20, wherein: The estimation model is generated for each temperature. The step of obtaining the SMILES expression of the substance includes the following processing: accepting the setting of the temperature to be estimated, The step of acquiring the estimation result of the HSP of the substance includes a process of selecting an estimation model corresponding to the temperature to be estimated. 23 . A computer program that, when executed by a control circuit of a computer, causes the computer to implement the estimation method according to claim 1 .

24. An information processing device comprising: control circuitry; and a storage device storing a computer program executed by the control circuit, in, The computer program causes the information processing device to implement the estimation method according to claim 1 by being executed by the control circuit.

Citation Information

Patent Citations

  • Time division multiplex picture display device

    JP1986010380A