Systems, methods, media, and products for recommending chiral separation conditions
By recommending chiral separation conditions through machine learning algorithms, the problems of time-consuming, labor-intensive, and costly processes in existing technologies are solved, and efficient and accurate chiral separation results are achieved.
Patent Information
- Application Number
- CN202511525307.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-24
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-10-24
AI Technical Summary
Existing technologies rely on manual selection of column type and solvent pair in chiral separation, which is time-consuming, labor-intensive, and costly, and cannot guarantee the success rate of separation. They also ignore the dependence of mobile phase operating parameters on residence time.
The system employs machine learning algorithms to recommend chiral separation conditions. Through modules such as similarity detection, column recommendation, solvent pair recommendation, operating parameter prediction, residence time prediction, and resolution analysis, the system calculates chiral separation conditions by combining machine learning algorithms with chromatographic theory.
It improved the success rate of chiral separation experiments, reduced manual verification time, lowered experimental costs, and improved the accuracy and efficiency of separation results.
Smart Images

Figure CN120998345A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure generally relates to computer systems utilizing computational models, and in particular to systems, methods, media, and products for recommending chiral separation conditions. BACKGROUND
[0002] In the fields of material research and development, drug research and development, and fine chemical industry, the separation of optical enantiomers of chiral molecules is a key but extremely challenging task. Figure 1 are different configurations of an example chiral molecule and the associated chromatogram. Taking a binuclear iridium complex with chirality as an example, for the same structural formula, there are two mirror-symmetric configurations R and S. Different configurations can have different chemical properties, for example, one configuration has drug activity, and the other configuration has physiological toxicity, so it is necessary to separate different enantiomers of the same compound.
[0003] High Performance Liquid Chromatography (HPLC), especially chiral stationary phase chromatography, is currently the most commonly used and most effective means of chiral separation. In the traditional HPLC chiral separation method, the chiral sample to be separated is dissolved in a mobile phase formed by two solvents. When the mobile phase carrying the sample flows through the chiral chromatographic column (stationary phase), because the binding ability of different enantiomers in the sample to the stationary phase is different, the stability of the diastereomeric complex formed is different, resulting in different residence times of the two enantiomers in the chromatographic column, the enantiomer with stronger binding has a longer residence time in the chromatographic column, and the enantiomer with weaker binding has a shorter residence time in the chromatographic column, and finally the two enantiomers flow out from the end of the chromatographic column in turn, which can achieve physical separation. When the separated enantiomers enter the detector, the detector can convert the concentration signal of the enantiomers into a chromatogram. As shown in Figure 1 , the binuclear iridium complex with chirality has two chromatographic peaks in the chromatogram, each chromatographic peak has a corresponding residence time. When the chromatographic peaks can be separated in residence time, it can be considered as successful separation.
[0004] Therefore, the core of the HPLC method is to find the best combination of "chromatographic column-solvent pair (mobile phase)-mobile phase running parameters" for a specific chiral molecule.
[0005] However, because there are a large number of chromatographic column types and solvents on the market, and the mobile phase running parameters of the solvent pair also need to be set artificially. Therefore, for the separation conditions of a specific chiral molecule, it is heavily dependent on the professional knowledge of chemists and a large number of trial-and-error experiments, and the chromatographic column types and solvent pairs and their mobile phase running parameters are tested one by one. This manual testing process not only consumes time and effort, but also requires high cost of chromatographic columns and solutions for experiments, and cannot guarantee the success rate of separation.
[0006] The prior art provides a method of predicting separation conditions relying on retention time, which manually selects a specific column type, then predicts retention time, and determines whether the separation condition meets the chiral separation of the compound by traversing the retention time of all column types. However, due to the diversity of column types, mobile phase types, and mobile phase running parameter combinations, manual selection of these parameters will result in low optimization efficiency. In addition, for a completely new molecule, how to select the first experimental condition (especially the column type and solvent pair) is the biggest difficulty, which often relies on expert experience. Further, the existing method only predicts a single parameter such as retention time, ignoring the dependence between mobile phase running parameters and retention time.
[0007] There is a need in the art for chiral separation condition recommendation techniques that improve on at least one of the above aspects. SUMMARY
[0008] To provide further improved chiral separation condition recommendation techniques using machine learning algorithms to predict column types, solvent pairs, and mobile phase running parameters, the present invention is provided.
[0009] One aspect of the present invention provides a system for recommending chiral separation conditions, comprising: a similarity detection module configured to receive an expression of a compound to be predicted, and invoke a first agent to: based on the expression, calculate a similarity between the compound to be predicted and known compounds; a column recommendation module configured to invoke a second agent to: determine a plurality of candidate column types according to the similarity; a solvent pair recommendation module configured to invoke a third agent to: determine a candidate solvent pair according to the similarity; a running parameter prediction module configured to, for each candidate column type of the plurality of candidate column types, predict a mobile phase running parameter of the candidate solvent pair based on the expression, the candidate column type, and the candidate solvent pair using a first machine learning algorithm; a retention time prediction module configured to, for each candidate column type of the plurality of candidate column types, predict a retention time associated with a chromatographic peak of a chromatogram of the compound to be predicted using the candidate column type, the candidate solvent pair, and the mobile phase running parameter based on the expression, the candidate column type, the candidate solvent pair, and the mobile phase running parameter using a second machine learning algorithm; a resolution analysis module configured to invoke a fourth agent to: perform a resolution analysis associated with the compound to be predicted based on the retention time; and a reordering module configured to reorder the plurality of candidate column types based on a result of the resolution analysis, wherein a recommended chiral separation condition is determined based on the reordered candidate column types and associated candidate solvent pairs and mobile phase running parameters.
[0010] The system as claimed in the above, the similarity detection module is configured to invoke the first agent to: convert the expression of the compound to be predicted into a first molecular descriptor; convert the string of the known compound into a second molecular descriptor; and calculate the similarity between the first molecular descriptor and the second molecular descriptor.
[0011] The system as claimed in the above, the column recommendation module is configured to invoke the second agent to: sort the column types corresponding to the known compound according to the similarity; and determine the plurality of candidate column types based on the result of the sorting.
[0012] The system as claimed in the above, the column recommendation module is further configured to invoke the second agent to: when the similarity of the column type with the highest similarity in the plurality of candidate column types is less than a similarity threshold, generate a predicted column type based on the expression; and replace the column type with the highest similarity with the predicted column type.
[0013] The system as claimed in the above, the solvent pair recommendation module is configured to invoke the third agent to: sort the solvent pairs corresponding to the known compound according to the similarity; and determine the candidate solvent pair based on the result of the sorting.
[0014] The system as claimed in the above, the solvent pair recommendation module is further configured to invoke the third agent to: when the similarity of the candidate solvent pair is less than a similarity threshold, generate a predicted solvent pair based on the expression; and replace the candidate solvent pair with the predicted solvent pair.
[0015] The system as claimed in the above, the running parameter prediction module is configured to: perform feature engineering on the expression, the candidate column type and the candidate solvent pair to splice into a conditional vector; and predict the mobile phase running parameter of the candidate solvent pair using the first machine learning algorithm based on the conditional vector.
[0016] The system as claimed in the above, the running parameter prediction module is configured to: convert the expression into a first molecular descriptor; convert the string of the candidate solvent pair into a second molecular descriptor; convert the candidate column type into one-hot encoding; and splice the first molecular descriptor, the second molecular descriptor and the one-hot encoding into the conditional vector.
[0017] The system as claimed in the preceding paragraph, the mobile phase operating parameters comprising solvent proportions and solvent flow rates, wherein the operating parameter prediction module is configured to invoke a fifth agent to, for each candidate chromatographic column type of the plurality of candidate chromatographic column types, predict the solvent proportions of the candidate solvent pair based on the condition vector, and wherein the operating parameter prediction module is configured to invoke a sixth agent to, for each candidate chromatographic column type of the plurality of candidate chromatographic column types, predict the solvent flow rates of the candidate solvent pair based on the condition vector and the solvent proportions.
[0018] The system as claimed in the preceding paragraph, the retention times comprising a first retention time and a second retention time, wherein the retention time prediction module is configured to use the second machine learning algorithm to predict the first retention time corresponding to a first configuration of the compound to be predicted and the second retention time corresponding to a second configuration of the compound to be predicted, and wherein the resolution analysis module is configured to invoke the fourth agent to perform the resolution analysis based on the first retention time and the second retention time.
[0019] The system as claimed in the preceding paragraph, the retention time prediction module being configured to invoke a seventh agent to, for each candidate chromatographic column type of the plurality of candidate chromatographic column types, predict the first retention time based on the expression, the candidate chromatographic column type, the candidate solvent pair, and the mobile phase operating parameters, and the retention time prediction module being configured to invoke an eighth agent to, for each candidate chromatographic column type of the plurality of candidate chromatographic column types, predict the second retention time based on the expression, the candidate chromatographic column type, the candidate solvent pair, the mobile phase operating parameters, and the first retention time.
[0020] The system as claimed in the preceding paragraph, the resolution analysis module being configured to invoke the fourth agent to, for each candidate chromatographic column type of the plurality of candidate chromatographic column types, calculate a peak resolution based on the retention times.
[0021] The system as claimed in the preceding paragraph, the retention times comprising a first retention time and a second retention time, wherein the resolution analysis module is configured to invoke the fourth agent to, for each candidate chromatographic column type of the plurality of candidate chromatographic column types, calculate the peak resolution based on the first retention time corresponding to a first configuration of the compound to be predicted and a first half-peak width and the second retention time corresponding to a second configuration of the compound to be predicted and a second half-peak width.
[0022] The system as claimed in the preceding paragraph, the known compounds and corresponding column type and solvent pair are obtained from a database, the database is obtained by a large language model by: obtaining literatures related to high performance liquid chromatography; extracting known compounds and corresponding column type and solvent pair from the literatures; storing the known compounds and corresponding column type and solvent pair in the database.
[0023] Another aspect of the present application provides a method for recommending chiral separation conditions, comprising: S100: receiving an expression of a compound to be predicted; S200: calculating a similarity between the compound to be predicted and known compounds based on the expression; S300: determining a plurality of candidate column types and determining a candidate solvent pair according to the similarity; S400: for each candidate column type in the plurality of candidate column types, predicting a mobile phase operating parameter of the candidate solvent pair using a first machine learning algorithm based on the expression, the candidate column type and the candidate solvent pair; S500: for each candidate column type in the plurality of candidate column types, predicting a retention time associated with a chromatographic peak of a chromatogram of the compound to be predicted using a second machine learning algorithm based on the expression, the candidate column type, the candidate solvent pair and the mobile phase operating parameter, wherein the chromatogram is obtained using the candidate column type, the candidate solvent pair and the mobile phase operating parameter; S600: performing a resolution analysis associated with the compound to be predicted based on the retention time; and S700: reordering the plurality of candidate column types based on a result of the resolution analysis, wherein the recommended chiral separation conditions are determined based on the reordered candidate column types and associated candidate solvent pairs and mobile phase operating parameters.
[0024] The method as claimed in the preceding paragraph, the S200 comprises: S201: converting the expression of the compound to be predicted into a first molecular descriptor; S202: converting the string of the known compound into a second molecular descriptor; and S203: calculating a similarity between the first molecular descriptor and the second molecular descriptor.
[0025] The method as claimed in the preceding paragraph, the S300 comprises: S301: according to the similarity, first ordering column types corresponding to the known compounds and second ordering solvent pairs corresponding to the known compounds; S302: determining the plurality of candidate column types based on a result of the first ordering; and S303: determining the candidate solvent pair based on a result of the second ordering.
[0026] The method as described above, the S300 further comprises: S304: when the similarity of the chromatographic column type with the highest similarity among the plurality of candidate chromatographic column types is less than a similarity threshold, generating a predicted chromatographic column type and a predicted solvent pair based on the expression using a third machine learning algorithm; S305: replacing the chromatographic column type with the highest similarity with the predicted chromatographic column type; and S306: replacing the candidate solvent pair with the predicted solvent pair.
[0027] The method as described above, the S400 comprises: S401: performing feature engineering on the expression, the candidate chromatographic column type and the candidate solvent pair to splice into a condition vector; and S402: predicting a mobile phase running parameter of the candidate solvent pair based on the condition vector using the first machine learning algorithm.
[0028] The method as described above, the mobile phase running parameter comprises a solvent ratio and a solvent flow rate, the S402 comprises: S4021: for each candidate chromatographic column type in the plurality of candidate chromatographic column types, predicting the solvent ratio of the candidate solvent pair based on the condition vector using the first machine learning algorithm; S4022: for each candidate chromatographic column type in the plurality of candidate chromatographic column types, predicting the solvent flow rate of the candidate solvent pair based on the condition vector and the solvent ratio using the first machine learning algorithm.
[0029] The method as described above, the retention time comprises a first retention time and a second retention time, the S500 comprises: S501: for each candidate chromatographic column type in the plurality of candidate chromatographic column types, predicting the first retention time corresponding to a first configuration of the compound to be predicted based on the expression, the candidate chromatographic column type, the candidate solvent pair and the mobile phase running parameter using the second machine learning algorithm; S502: for each candidate chromatographic column type in the plurality of candidate chromatographic column types, predicting the second retention time corresponding to a second configuration of the compound to be predicted based on the expression, the candidate chromatographic column type, the candidate solvent pair, the mobile phase running parameter and the first retention time using the second machine learning algorithm, wherein the resolution analysis is performed based on the first retention time and the second retention time.
[0030] The method as described above, the retention time comprises a first retention time and a second retention time, the S600 comprises: for each candidate chromatographic column type in the plurality of candidate chromatographic column types, calculating a peak resolution based on the first retention time and a first half-peak width corresponding to a first configuration of the compound to be predicted and the second retention time and a second half-peak width corresponding to a second configuration of the compound to be predicted.
[0031] The method as described above, the known compounds and corresponding column type and solvent pairs are obtained from a database, the database is obtained by a large language model by: obtaining literatures related to high performance liquid chromatography; extracting known compounds and corresponding column type and solvent pairs from the literatures; storing the known compounds and corresponding column type and solvent pairs in the database.
[0032] Another aspect of the present application provides a computer readable storage medium having stored thereon a computer program which, when executed by a processor, implements the steps of the method according to any one of the preceding aspects.
[0033] Another aspect of the present application provides a computer program product comprising a computer program which, when executed by a processor, implements the steps of the method according to any one of the preceding aspects.
[0034] The system and method according to the present application overcome the limitations of the prior art that cannot recommend column types, by utilizing machine learning algorithms to predict column types, solvent pairs and mobile phase running parameters, improving the experimental success rate of chiral separation conditions. BRIEF DESCRIPTION OF DRAWINGS
[0035] Embodiments of the present application are described in conjunction with the appended drawings.
[0036] Figure 1 are different conformations of an example chiral molecule and associated chromatograms.
[0037] Figure 2 is a block diagram of a system for recommending chiral separation conditions according to some embodiments of the present application.
[0038] Figure 3 is a schematic diagram of a system for recommending chiral separation conditions according to some embodiments of the present application.
[0039] Figure 4 is a schematic diagram of a system for recommending chiral separation conditions according to further embodiments of the present application.
[0040] Figure 5 is a schematic diagram of recommended chiral separation conditions according to some embodiments of the present application.
[0041] Figure 6 is a flowchart of a method for recommending chiral separation conditions according to some embodiments of the present application.
[0042] Figure 7 is a flowchart of a first process associated with a method for recommending chiral separation conditions according to some embodiments of the present application.
[0043] Figure 8 is a flowchart of a second process associated with the method for recommending chiral separation conditions according to some embodiments of the application.
[0044] Figure 9 is a flowchart of a third process associated with the method for recommending chiral separation conditions according to some embodiments of the application.
[0045] Figure 10 is a flowchart of a fourth process associated with the third process of the method for recommending chiral separation conditions according to some embodiments of the application.
[0046] Figure 11 is a flowchart of a fifth process associated with the third process of the method for recommending chiral separation conditions according to some embodiments of the application.
[0047] Figure 12 is a flowchart of a sixth process associated with the method for recommending chiral separation conditions according to some embodiments of the application.
[0048] Figure 13 is a block diagram of a computer readable storage medium according to some embodiments of the application.
[0049] Figure 14 is a block diagram of a computer program product according to some embodiments of the application. DETAILED DESCRIPTION
[0050] In this application, different instances of the same name are distinguished by the use of ordinal numbers, e.g., “first”, “second”, “third”, etc. The ordinal numbers “first”, “second”, “third”, etc. do not indicate relative order in time, space, ordering, and other aspects of the objects indicated.
[0051] In this application, the term “agent” refers to an agent that can perceive an environment and take actions to perform a specific goal. The agent mainly refers to software code. The agent can be executed by the computing resources of a computing device. The agent can call the corresponding model through an Application Programming Interface (API), and call the corresponding tool to interact with various forms of input or implement the corresponding functions.
[0052] According to an aspect of the application, a system for recommending chiral separation conditions is provided.
[0053] Figure 2is a block diagram of a system 100 for recommending chiral separation conditions according to some embodiments of the present application. The system 100 includes a similarity detection module 102, a column recommendation module 104, a solvent pair recommendation module 106, a run parameter prediction module 108, a residence time prediction module 110, a resolution analysis module 112, and a reordering module 114. The modules are described in more detail in connection with Figures 3-4 The system 100 is further described.
[0054] Figure 3 is a schematic diagram of a system 100 for recommending chiral separation conditions according to some embodiments of the present application.
[0055] The similarity detection module 102 can be configured to receive an expression of a compound to be predicted, and invoke the first agent 122 to compute a similarity between the compound to be predicted and known compounds based on the expression.
[0056] For example, a user can input an expression of a compound to be predicted through a user interface, and the similarity detection module 102 can compute a similarity between the expression of the compound to be predicted and a string of known compounds.
[0057] In some embodiments, the similarity detection module 102 can be configured to invoke the first agent 122 to convert the expression of the compound to be predicted to a first molecular descriptor, convert the string of known compounds to a second molecular descriptor, and compute a similarity between the first molecular descriptor and the second molecular descriptor.
[0058] For example, the expression of the compound to be predicted can be a molecular structure diagram, a molecular structure formula, or a SMILES string. The similarity detection module 102 can convert the expression of the compound to be predicted to a first molecular descriptor such as a Morgan fingerprint. The string of known compounds can be a SMILES string. The similarity detection module 102 can convert the string of known compounds to a second molecular descriptor such as a Morgan fingerprint. Then, the similarity detection module 102 can compute a similarity between the first molecular descriptor of the compound to be predicted and the second molecular descriptor of each known compound. As an example, the similarity can be a Tanimoto similarity for measuring similarity of molecular structures.
[0059] The column recommendation module 104 can be configured to invoke the second agent 124 to determine a plurality of candidate column types according to the similarity.
[0060] The optional column types include different column types from different manufacturers, which can be as many as hundreds of column types. For a plurality of known compounds, a corresponding column type is pre-set for each known compound. The column types between different compounds can be the same or different. The column recommendation module 104 can determine a plurality of column types that can be suitable for the to-be-predicted compound from the column types corresponding to all known compounds according to the similarity.
[0061] Since the column type is the most critical factor in the separation experiment, and the number of column types is large, by automatically determining a plurality of candidate column types based on the similarity, the test range of the separation condition can be significantly narrowed, the confirmation efficiency of the column type can be greatly improved, the success rate of the separation experiment can be improved, and manpower and time can be saved.
[0062] In some embodiments, the column recommendation module 104 can be configured to call the second agent 124 to sort the column types corresponding to the known compounds according to the similarity, and determine a plurality of candidate column types based on the sorting result.
[0063] For example, the column recommendation module 104 can sort a plurality of known compounds according to the similarity, and sort their column types accordingly, wherein the same column type corresponding to different known compounds is sorted according to the highest similarity, and then a plurality of candidate column types are determined from the column types with higher rankings. As an example, the top 5 column types from the column types sorted according to the similarity can be selected as the candidate column types.
[0064] In some embodiments, the column recommendation module 104 can also be configured to call the second agent 124 to generate a predicted column type based on the expression of the to-be-predicted compound when the similarity of the column type with the highest similarity in the plurality of candidate column types is less than the similarity threshold, and replace the column type with the highest similarity with the predicted column type.
[0065] For example, when the similarity of the column type with the highest similarity is 0.3, and the similarity threshold is 1, the similarity is much less than the similarity threshold, and the known compound with the highest similarity can be quite different from the to-be-predicted compound. By calling the agent to generate a predicted column type based on the expression of the to-be-predicted compound to replace the candidate column type with a large difference with a similarity of 0.3, an enhanced possibility can be provided to improve the separation effect of the separation experiment. The similarity and the similarity threshold in the above example are only an example, and the user can set the similarity threshold according to actual needs, and the scope of the present disclosure is not limited thereto.
[0066] As an example, the second agent 124 can invoke a classification model to generate a predicted chromatography column type, such as a large language model, a graph neural network model, etc.
[0067] The solvent pair recommendation module 106 can be configured to invoke the third agent 126 to determine a candidate solvent pair according to the similarity.
[0068] The optional solvent pairs include up to dozens of solvent pairs. For a plurality of known compounds, a corresponding solvent pair is pre-set for each known compound. The expression of the solvent pair can be "first solvent / second solvent". The solvent pairs between different compounds can be the same or different. The solvent pair recommendation module 106 can determine a solvent pair that can be suitable for the to-be-predicted compound from the solvent pairs corresponding to all known compounds according to the similarity.
[0069] Since the solvent pair can have an impact on the binding ability of the chiral molecule to the chromatography column, it is also a key factor in the separation experiment. By using an agent to automatically determine the candidate solvent pair, the manual confirmation time of the candidate solvent pair can be reduced, and the success rate of the separation experiment can be improved.
[0070] In some embodiments, the solvent pair recommendation module 106 can be configured to invoke the third agent 126 to sort the solvent pairs corresponding to the known compounds according to the similarity, and determine a candidate solvent pair based on the sorting result.
[0071] For example, the solvent pair recommendation module 106 can sort a plurality of known compounds according to the similarity, and sort their solvent pairs accordingly, and then determine the top-ranked solvent pair as the candidate solvent pair. Alternatively, the solvent pair recommendation module 106 can determine a plurality of solvent pairs with higher rankings as a plurality of candidate solvent pairs to increase the combination of separation conditions and further improve the success rate of the separation experiment.
[0072] In some embodiments, the solvent pair recommendation module 106 can also be configured to invoke the third agent 126 to generate a predicted solvent pair based on the expression when the similarity of the candidate solvent pair is less than a similarity threshold, and replace the candidate solvent pair with the predicted solvent pair.
[0073] For example, when the similarity of the candidate solvent pair is 0.3 and the similarity threshold is 1, the similarity is much less than the similarity threshold, and the known compound with the highest similarity can be quite different from the to-be-predicted compound. By invoking the agent to generate a predicted solvent pair based on the expression of the to-be-predicted compound, the candidate solvent pair with a large difference in similarity of 0.3 can be replaced, which can provide enhanced possibilities for improving the separation effect of the separation experiment. The similarity and the similarity threshold in the above example are only an example, and the user can set the similarity threshold according to actual needs, and the scope of the present disclosure is not limited thereto.
[0074] As an example, the third agent 126 can invoke a classification model to generate a predicted solvent pair, e.g., a large language model, a graph neural network model, etc.
[0075] The running parameter prediction module 108 can be configured to, for each of the plurality of candidate chromatographic column types, predict a mobile phase running parameter of the candidate solvent pair based on the expression of the compound to be predicted, the candidate chromatographic column type, and the candidate solvent pair using a first machine learning algorithm.
[0076] After determining the candidate solvent pair, since the mobile phase running parameter of the candidate solvent pair also has an impact on the separation effect, e.g., affecting the retention time, the peak spacing, the peak width, etc. of the two configurations, an accurate prediction of the mobile phase running parameter of the candidate solvent pair can improve the separation experiment effect.
[0077] As an example, the first machine learning algorithm can be a regression model, e.g., extreme gradient boosting (XGBoost), Tabnet, category boosting (CatBoost), gradient boosting machine (LightGBM), random forest (Random Forest), deep neural network (Deep Neural Network, DNN), graph neural network (Graph Neural Network, GNN), Transformer, etc.
[0078] In some embodiments, the running parameter prediction module 108 can be configured to perform feature engineering on the expression of the compound to be predicted, the candidate chromatographic column type, and the candidate solvent pair to splice into a condition vector, and based on the condition vector, predict the mobile phase running parameter of the candidate solvent pair using the first machine learning algorithm.
[0079] By feature engineering, the expression of the compound to be predicted, the candidate chromatographic column type, and the candidate solvent pair are converted into a condition vector, so that the computing model can read the condition vector.
[0080] In some embodiments, the running parameter prediction module 108 can be configured to convert the expression of the compound to be predicted into a first molecular descriptor, convert the string of the candidate solvent pair into a second molecular descriptor, convert the candidate chromatographic column type into one-hot encoding, and splice the first molecular descriptor, the second molecular descriptor, and the one-hot encoding into a condition vector. For example, the molecular descriptor can include one or more of a MolT5 embedding vector, a Morgan fingerprint, a MACCS Keys molecular fingerprint, etc. for describing molecular structure or related information.
[0081] The retention time prediction module 110 can be configured to, for each of the plurality of candidate chromatographic column types, use the second machine learning algorithm to predict, based on the expression of the compound to be predicted, the candidate chromatographic column type, the candidate solvent pair, and the mobile phase running parameter, a retention time associated with a chromatographic peak of a chromatogram of the compound to be predicted using the candidate chromatographic column type, the candidate solvent pair, and the mobile phase running parameter.
[0082] As an example, the second machine learning algorithm can be a regression model, such as extreme gradient boosting (XGBoost), Tabnet, category boosting (CatBoost), gradient boosting machine (LightGBM), random forest (Random Forest), deep neural network (DNN), graph neural network (GNN), Transformer, etc.
[0083] The separation analysis module 112 can be configured to invoke the fourth agent 128 to perform a separation analysis associated with the compound to be predicted based on the retention time. When the chromatographic peaks of the chromatogram of the compound to be predicted are separated from each other under a particular separation condition, it can be considered that the two configurations corresponding to the compound to be predicted can be separated.
[0084] In some embodiments, the separation analysis module 112 can be configured to invoke the fourth agent 128 to, for each of the plurality of candidate chromatographic column types, calculate a peak separation degree based on the retention time.
[0085] For two chromatographic peaks in a chromatogram under a particular separation condition, although the retention times of the two configurations are different, if the chromatographic peaks are wide, there can still be a case where the two chromatographic peaks overlap. Therefore, by calculating the peak separation degree based on the retention time, it can be further ensured that the recommended separation condition can achieve the separation of the enantiomers.
[0086] In some embodiments, the retention time can include a first retention time and a second retention time. The retention time prediction module 110 can be configured to use the second machine learning algorithm to predict the first retention time corresponding to the first configuration of the compound to be predicted and the second retention time corresponding to the second configuration of the compound to be predicted. The separation analysis module 112 can be configured to invoke the fourth agent 128 to perform a separation analysis based on the first retention time and the second retention time.
[0087] In some embodiments, the retention time can include a first retention time and a second retention time. The separation analysis module 112 can be configured to invoke the fourth agent 128 to, for each of the plurality of candidate chromatographic column types, calculate a peak separation degree based on the first retention time and the first half-peak width corresponding to the first configuration of the compound to be predicted and the second retention time and the second half-peak width corresponding to the second configuration of the compound to be predicted.
[0088] Since the residence time and the half-peak width of each configuration can be considered as a linear relationship, the half-peak width of each configuration can be determined based on the residence time, and then the peak separation degree is calculated based on the residence time and the half-peak width. The calculation formula of the peak separation degree can be:
[0089] ,
[0090] wherein Rs is the peak separation degree, RT1 and RT2 are the first residence time and the second residence time respectively, W_half1 and W_half2 are the first half-peak width and the second half-peak width respectively, and c is a preset coefficient. The half-peak width can be estimated as:
[0091] ,
[0092] wherein W_half is the half-peak width, RT is the residence time, and a and b are preset coefficients. The preset coefficients a, b and c can use the system default parameters, or can be set according to the actual needs of the user in different experimental scenarios.
[0093] By combining the residence time and the half-peak width to calculate the peak separation degree, a more accurate prediction of the separation effect can be provided to ensure the actual separation result in the separation experiment.
[0094] The reordering module 114 can be used to reorder the plurality of candidate chromatographic column types based on the results of the separation degree analysis. The recommended chiral separation condition is determined based on the reordered candidate chromatographic column types and the associated candidate solvent pairs and mobile phase running parameters.
[0095] For example, for a plurality of candidate chromatographic column types previously sorted by similarity and the separation conditions corresponding to each chromatographic column type (e.g., including the mobile phase running parameters of the candidate solvent pairs), it is determined whether the peak separation degree corresponding to each chromatographic column type is greater than or equal to the separation degree threshold. The separation degree threshold can be adjusted according to the user's needs. For one or more chromatographic column types with a peak separation degree greater than or equal to the separation degree threshold, these chromatographic column types and their corresponding separation conditions are sorted by similarity. For one or more chromatographic column types with a peak separation degree less than the separation degree threshold, these chromatographic column types and their corresponding separation conditions are sorted by similarity. The chromatographic column types and the corresponding separation conditions with a peak separation degree greater than or equal to the separation degree threshold can have a higher priority in the subsequent experimental process.
[0096] By combining the machine learning algorithm with the chromatographic theoretical calculation method, i.e., first predicting the intermediate parameter residence time using the machine learning algorithm, then calculating the final performance indicator separation degree using the theoretical formula, and reordering the candidate separation conditions using the separation degree, further optimization of the ordering of the separation conditions guided by the final separation effect can be achieved.
[0097] In some embodiments, the known compounds and corresponding column type and solvent pair can be obtained from a database, which can be obtained by the large language model by: obtaining literatures related to high performance liquid chromatography; extracting the known compounds and corresponding column type and solvent pair from the literatures; storing the known compounds and corresponding column type and solvent pair in the database.
[0098] Figure 4 is a schematic diagram of a system 100 for recommending chiral separation conditions according to further embodiments of the present application. To avoid redundancy, Figure 4 with Figure 3 The specific details of the same elements of
[0099] In some embodiments, the mobile phase operating parameters can include solvent ratio and solvent flow rate. The operating parameter prediction module 108 can be configured to invoke the fifth agent 130 to predict, for each of the plurality of candidate column types, a solvent ratio of the candidate solvent pair based on the condition vector. The operating parameter prediction module 108 can be configured to invoke the sixth agent 132 to predict, for each of the plurality of candidate column types, a solvent flow rate of the candidate solvent pair based on the condition vector and the solvent ratio.
[0100] The solvent ratio can be predicted first based on the condition vector (including the expression of the compound to be predicted, the candidate column type and the candidate solvent pair), for example, the solvent ratio can be the percentage of the solvent with a higher proportion in the candidate solvent pair, and then the solvent flow rate is further predicted based on the condition vector and the predicted solvent ratio, so as to fully consider the dependence of downstream parameters on upstream parameters through the cascading prediction architecture, to accurately simulate the intrinsic logic of the chemical experiment and improve the prediction accuracy.
[0101] In some embodiments, the retention time prediction module 110 can be configured to invoke the seventh agent 134 to predict, for each of the plurality of candidate column types, a first retention time based on the expression, the candidate column type, the candidate solvent pair and the mobile phase operating parameters. The retention time prediction module 110 can be configured to invoke the eighth agent 136 to predict, for each of the plurality of candidate column types, a second retention time based on the expression, the candidate column type, the candidate solvent pair, the mobile phase operating parameters and the first retention time.
[0102] The first retention time can be predicted based on the expression of the compound to be predicted, the candidate chromatographic column type, the candidate solvent pair, and the mobile phase operation parameter, and then the second retention time can be further predicted based on the expression of the compound to be predicted, the candidate chromatographic column type, the candidate solvent pair, the mobile phase operation parameter, and the predicted first retention time, so as to fully consider the internal correlation between the two retention times and improve the prediction accuracy of the retention time.
[0103] After the chiral separation condition recommended by the system 100 is obtained, the user can perform experiments in sequence according to the separation conditions according to the ranking, and separate the compound to be predicted. For example, when the separation experiment is performed according to the first separation condition, the chromatographic peaks in the chromatogram are not completely separated, that is, the separation fails, and another separation experiment is performed according to the next separation condition. When the chromatographic peaks in the chromatogram obtained by the experiment are separated, the confirmation of the separation condition is completed.
[0104] In some embodiments, the first agent 122, the second agent 124, the third agent 126, the fourth agent 128, the fifth agent 130, the sixth agent 132, the seventh agent 134, and the eighth agent 136 can call the corresponding models through an API. As an example, the models can include a classification model, a gradient boosting tree model, a regression model, a deep neural network, a large language model, a multi-modal model, a multi-modal language model, etc. In some embodiments, the multiple models that can be called by the multiple agents can be deployed locally on the system 100. In some embodiments, the multiple models that can be called by the multiple agents can be deployed remotely from the system 100, for example, on the cloud. In some embodiments, some of the multiple models that can be called by the multiple agents can be deployed locally on the system 100, and the other models can be deployed remotely from the system 100. In some embodiments, each of the multiple agents can call various tools to interact with various forms of input or implement corresponding functions.
[0105] Figure 5 is a schematic diagram of a recommended chiral separation condition according to some embodiments of the present application. For the compound to be predicted (target molecule) as shown in the figure, the predicted candidate solvent pair (optimal solvent pair) is hexane and i-PrOH, and the multiple candidate chromatographic column types are ranked according to the similarity as IA, AD, IG, IE, and OD. The user can experiment in sequence according to the ranking (Rank) according to the listed chromatographic column type (Chiral Column), solvent ratio (Ratio), and solvent flow rate (Flow Rate), until the compound to be predicted is successfully separated. When the separation degree threshold is set to 14.5, OD can be ranked first because the separation degree 14.93 of OD is greater than the separation degree threshold. It is verified by experiment that when the chromatographic column type is selected as OD, the compound to be predicted can be successfully separated.
[0106] Some embodiments of the present application achieve fast, accurate and fully automated recommendation of chiral separation conditions based on user inputted target molecular structure, utilizing fast recommendation based on similarity of molecular structure and fine-tuning prediction based on machine learning. Compared with traditional methods which take days or even weeks, some embodiments of the present application can complete high-quality computational simulation and scheme recommendation within minutes, greatly improving R&D efficiency.
[0107] As an example, a system according to some embodiments of the present application establishes a database and prediction model for known compounds based on 33000 samples, and tests on 3700 samples for prediction. Experimental results show that the TOP-2 accuracy is 82%, the TOP-3 accuracy is 90%, and the TOP-5 accuracy is 95%. Experiments prove that the system according to some embodiments of the present application can provide highly accurate recommendation results for chiral separation conditions of the compounds to be predicted. In addition, using the system to predict compounds that have never been reported in the industry, the system according to some embodiments of the present application can recommend separation conditions for multiple unreported compounds, and these compounds can be successfully separated under the recommended two or more separation conditions, further proving the accuracy of the recommendation of chiral separation conditions by the system according to some embodiments of the present application.
[0108] According to another aspect of the present application, a method for recommending chiral separation conditions is provided.
[0109] Figure 6 is a flowchart of a method for recommending chiral separation conditions according to some embodiments of the present application.
[0110] The method can include step S100 of receiving an expression of a compound to be predicted.
[0111] The method can include step S200 of calculating similarity between the compound to be predicted and known compounds based on the expression.
[0112] The method can include step S300 of determining a plurality of candidate column types and determining a candidate solvent pair according to the similarity.
[0113] The method can include step S400 of, for each of the plurality of candidate column types, predicting mobile phase running parameters of the candidate solvent pair based on the expression, the candidate column type and the candidate solvent pair using a first machine learning algorithm.
[0114] The method can comprise a step S500 of predicting, for each of a plurality of candidate chromatographic column types, a retention time associated with a chromatographic peak of a chromatogram of the compound to be predicted using the candidate chromatographic column type, a candidate solvent pair, and a mobile phase running parameter based on the expression, the candidate chromatographic column type, the candidate solvent pair, and the mobile phase running parameter using a second machine learning algorithm.
[0115] The method can comprise a step S600 of performing a resolution analysis associated with the compound to be predicted based on the retention time.
[0116] The method can comprise a step S700 of reordering the plurality of candidate chromatographic column types based on a result of the resolution analysis, wherein the recommended chiral separation condition is determined based on the reordered candidate chromatographic column types and associated candidate solvent pairs and mobile phase running parameters.
[0117] Some embodiments of the present application propose an end-to-end, fully automated prediction system and method from a molecular structure of a compound to be predicted to a complete, quantified, multi-dimensional (including chromatographic column type, solvent pair, mobile phase running parameter, retention time, resolution) chiral separation experimental protocol. By combining fast, coarse-grained recommendations based on molecular structure similarity with slow, fine-grained quantitative predictions based on machine learning, accurate predictions of chiral separation conditions are achieved. Compared with traditional methods that require manual determination of chromatographic column types, some embodiments of the present application can automatically recommend chromatographic column types, helping users more efficiently determine this key factor that has a significant impact on separation results.
[0118] In addition, traditional methods may preferentially recommend separation conditions with longer retention times, but do not consider the resolution between the two chromatographic peaks corresponding to the two configurations of the compound to be predicted, resulting in overlap between the two chromatographic peaks in the experimental results and failure to achieve complete separation. Some embodiments of the present application reorder the separation conditions based on the results of the resolution analysis, ensuring that the preferred separation condition recommended to the user is the one most likely to achieve good separation, greatly improving the success rate of the first experiment and reducing the number and time of experiments.
[0119] Further, some embodiments of the present application provide users with a plurality of chiral separation conditions corresponding to a plurality of candidate chromatographic column types, enabling users to plan separation experiments rather than blindly testing chromatographic column types, solvent pairs, and mobile phase running parameters, providing a faster experimental path to successful separation.
[0120] Moreover, some embodiments of the present application can greatly reduce the consumption of expensive chiral chromatographic columns and high-purity solvents, saving experimental instrument time and experimental labor costs.
[0121] Figure 7This is a flowchart of a first process associated with a method for recommending chiral separation conditions according to some embodiments of the present invention. This first process may be... Figure 6 The specific implementation of step S200 in the method is described, but the scope of the present invention is not limited thereto.
[0122] The first process may include step S201: converting the expression of the compound to be predicted into a first molecular descriptor.
[0123] The first process may include step S202: converting a string of known compounds into a second molecular descriptor.
[0124] The first process may include step S203: calculating the similarity between the first molecular descriptor and the second molecular descriptor.
[0125] Figure 8 This is a flowchart of a second process associated with a method for recommending chiral separation conditions according to some embodiments of the present invention. This second process may be... Figure 6 The specific implementation of step S300 in the method is described, but the scope of the present invention is not limited thereto.
[0126] The second process may include step S301: sorting the column types corresponding to the known compounds in a first sorting based on similarity, and sorting the solvent pairs corresponding to the known compounds in a second sorting.
[0127] The second process may include step S302: determining multiple candidate column types based on the results of the first sorting.
[0128] The second process may include step S303: determining candidate solvent pairs based on the results of the second sorting.
[0129] In some embodiments, the second process may further include step S304: when the similarity of the column type with the highest similarity among multiple candidate column types is less than a similarity threshold, a predicted column type and a predicted solvent pair are generated based on an expression and using a third machine learning algorithm.
[0130] The second process may also include step S305: replacing the column type with the highest similarity with the predicted column type.
[0131] The second process may also include step S306: replacing the candidate solvent pair with the predicted solvent pair.
[0132] Figure 9 This is a flowchart of a third process associated with a method for recommending chiral separation conditions according to some embodiments of the present invention. This third process may be... Figure 6 The specific implementation of step S400 in the method is shown, but the scope of the present invention is not limited thereto.
[0133] The third process can comprise a step S401 of performing feature engineering on the expression, the candidate chromatographic column type, and the candidate solvent pair to concatenate into a condition vector.
[0134] The third process can comprise a step S402 of predicting, based on the condition vector, a mobile phase operating parameter of the candidate solvent pair using a first machine learning algorithm.
[0135] Figure 10 is a flowchart of a fourth process associated with the third process of the method for recommending chiral separation conditions according to some embodiments of the application. The fourth process can be a specific implementation of the step S401 in the third process in Figure 9 The fourth process can comprise a step S4011 of converting the expression into a first molecular descriptor.
[0136] The fourth process can comprise a step S4012 of converting the string of the candidate solvent pair into a second molecular descriptor.
[0137] The fourth process can comprise a step S4013 of converting the candidate chromatographic column type into a one-hot encoding.
[0138] The fourth process can comprise a step S4014 of concatenating the first molecular descriptor, the second molecular descriptor, and the one-hot encoding into the condition vector.
[0139] The fourth process can comprise a step S4015 of predicting, based on the condition vector, a solvent ratio of the candidate solvent pair using the first machine learning algorithm.
[0140] Figure 11 is a flowchart of a fifth process associated with the third process of the method for recommending chiral separation conditions according to some embodiments of the application. The fifth process can be a specific implementation of the step S402 in the third process in Figure 9 The fifth process can comprise a step S4021 of predicting, based on the condition vector, a solvent ratio of the candidate solvent pair using the first machine learning algorithm for each candidate chromatographic column type in the plurality of candidate chromatographic column types.
[0141] The fifth process can comprise a step S4021 of predicting, based on the condition vector, a solvent ratio of the candidate solvent pair using the first machine learning algorithm for each candidate chromatographic column type in the plurality of candidate chromatographic column types.
[0142] The fifth process can comprise a step S4022 of predicting, based on the condition vector and the solvent ratio, a solvent flow rate of the candidate solvent pair using the first machine learning algorithm for each candidate chromatographic column type in the plurality of candidate chromatographic column types.
[0143] Some embodiments of the present application use a cascading prediction architecture to fully consider the physical dependency between experimental parameters, so that the prediction of each step is based on more complete context information, thus the prediction result is more accurate and more practically meaningful than the prediction model that predicts each parameter in isolation.
[0144] In some embodiments, the residence time can include a first residence time and a second residence time, S500 can include: predicting, using the second machine learning algorithm, the first residence time corresponding to the first configuration of the compound to be predicted and the second residence time corresponding to the second configuration of the compound to be predicted, wherein the resolution analysis is performed based on the first residence time and the second residence time.
[0145] Figure 12 is a flowchart of a sixth process associated with the method for recommending chiral separation conditions according to some embodiments of the present application. The sixth process can be Figure 6 The specific implementation of step S500 in the method in
[0146] The sixth process can include step S501: for each candidate column type in the plurality of candidate column types, predicting, using the second machine learning algorithm, a first residence time corresponding to the first configuration of the compound to be predicted based on the expression, the candidate column type, a candidate solvent pair and a mobile phase operating parameter.
[0147] The sixth process can include step S502: for each candidate column type in the plurality of candidate column types, predicting, using the second machine learning algorithm, a second residence time corresponding to the second configuration of the compound to be predicted based on the expression, the candidate column type, a candidate solvent pair, a mobile phase operating parameter and the first residence time.
[0148] Some embodiments of the present application use a cascading prediction method to fully consider the physical dependency between the residence times of the chromatographic peaks of the two configurations, so that the prediction of each step is based on more complete context information, thus the prediction result is more accurate than the prediction model that predicts each parameter in isolation.
[0149] In some embodiments, S600 can include: for each candidate column type in the plurality of candidate column types, calculating a peak resolution based on the residence time.
[0150] In some embodiments, the residence time can include a first residence time and a second residence time, S600 can include: for each candidate chromatographic column type in the plurality of candidate chromatographic column types, calculating a peak resolution based on the first residence time and the first half-peak width corresponding to the first configuration of the compound to be predicted and the second residence time and the second half-peak width corresponding to the second configuration of the compound to be predicted.
[0151] In some embodiments, the known compounds and corresponding chromatographic column types and solvent pairs can be obtained from a database, the database can be obtained by the large language model in the following way: obtaining literatures related to high performance liquid chromatography; extracting known compounds and corresponding chromatographic column types and solvent pairs from the literatures; storing the known compounds and corresponding chromatographic column types and solvent pairs in the database.
[0152] According to another aspect of the present application, a computer readable storage medium is provided.
[0153] Figure 13 is a block diagram of a computer readable storage medium 1300 according to some embodiments of the present application.
[0154] The computer readable storage medium 1300 stores a computer program 1350. The computer program 1350, when executed by a processor, implements the steps of the methods or procedures described above in conjunction with Figures 6-12 the various methods or procedures described above.
[0155] According to another aspect of the present application, a computer program product is provided.
[0156] Figure 14 is a block diagram of a computer program product 1400 according to some embodiments of the present application.
[0157] The computer program product 1400 can include the computer program 1350. The computer program 1350, when executed by a processor, implements the steps of the methods or procedures described above in conjunction with Figures 6-12 the various methods or procedures described above.
[0158] Embodiments of the present application have been described with reference to the accompanying drawings. Each embodiment is illustrative rather than limiting.
Claims
1. A system for recommending chiral separation conditions, characterized in that, The method comprises: a similarity detection module configured to receive an expression of a compound to be predicted, and invoke a first agent to: based on the expression, calculate a similarity between the compound to be predicted and a known compound; a column recommendation module configured to invoke a second agent to: determine a plurality of candidate column types according to the similarity; a solvent pair recommendation module configured to invoke a third agent to: determine a candidate solvent pair according to the similarity; a run parameter prediction module configured to, for each candidate column type in the plurality of candidate column types, predict a mobile phase run parameter of the candidate solvent pair based on the expression, the candidate column type, and the candidate solvent pair, using a first machine learning algorithm; a retention time prediction module configured to, for each candidate column type in the plurality of candidate column types, predict a retention time associated with a chromatographic peak of a chromatogram of the compound to be predicted using the candidate column type, the candidate solvent pair, and the mobile phase run parameter, using a second machine learning algorithm; a resolution analysis module configured to invoke a fourth agent to: perform a resolution analysis associated with the compound to be predicted based on the retention time; and a reordering module configured to reorder the plurality of candidate column types based on a result of the resolution analysis, wherein a recommended chiral separation condition is determined based on the reordered candidate column types and associated candidate solvent pairs and mobile phase run parameters. The similarity detection module is configured to invoke the first agent to:
2. The system of claim 1, wherein, convert the expression of the compound to be predicted into a first molecular descriptor; convert a string of the known compound into a second molecular descriptor; and calculate a similarity between the first molecular descriptor and the second molecular descriptor. The column recommendation module is configured to invoke the second agent to: rank column types corresponding to the known compound according to the similarity; and 3. The system of claim 1, wherein, determine the plurality of candidate column types based on a result of the ranking. The column recommendation module is further configured to invoke the second agent to: when a similarity of a column type having a highest similarity in the plurality of candidate column types is less than a similarity threshold, generate a predicted column type based on the expression; and 4. The system of claim 3, wherein, replace the column type having the highest similarity with the predicted column type. The solvent pair recommendation module is configured to invoke the third agent to: rank solvent pairs corresponding to the known compound according to the similarity; and determine the candidate solvent pair based on a result of the ranking.
5. The system of claim 1, wherein, The solvent pair recommendation module is further configured to invoke the third agent to: when a similarity of the candidate solvent pair is less than a similarity threshold, generate a predicted solvent pair based on the expression; and replace the candidate solvent pair with the predicted solvent pair.
6. The system of claim 5, wherein, The run parameter prediction module is configured to: 7. The system of claim 1, wherein, perform feature engineering on the expression, the candidate chromatographic column type, and the candidate solvent pair to concatenate into a condition vector; and predict, based on the condition vector, the mobile phase operating parameters of the candidate solvent pair using the first machine learning algorithm.
8. The system of claim 7, wherein, The operating parameter prediction module is configured to: convert the expression into a first molecular descriptor; convert a string of the candidate solvent pair into a second molecular descriptor; convert the candidate chromatographic column type into a one-hot encoding; and concatenate the first molecular descriptor, the second molecular descriptor, and the one-hot encoding into the condition vector. The mobile phase operating parameters include a solvent ratio and a solvent flow rate, 9. The system of claim 7, wherein, wherein the operating parameter prediction module is configured to invoke a fifth agent to predict, based on the condition vector, the solvent ratio of the candidate solvent pair for each of the plurality of candidate chromatographic column types; wherein the operating parameter prediction module is configured to invoke a sixth agent to predict, based on the condition vector and the solvent ratio, the solvent flow rate of the candidate solvent pair for each of the plurality of candidate chromatographic column types. The retention time includes a first retention time and a second retention time, 10. The system of claim 1, wherein, wherein the retention time prediction module is configured to predict, using the second machine learning algorithm, the first retention time corresponding to a first configuration of the compound to be predicted and the second retention time corresponding to a second configuration of the compound to be predicted, wherein the separation analysis module is configured to invoke the fourth agent to perform the separation analysis based on the first retention time and the second retention time.
11. The system of claim 10, wherein the retention time prediction module is configured to invoke a seventh agent to predict, based on the expression, the candidate chromatographic column type, the candidate solvent pair, and the mobile phase operating parameters, the first retention time for each of the plurality of candidate chromatographic column types; the retention time prediction module is configured to invoke an eighth agent to predict, based on the expression, the candidate chromatographic column type, the candidate solvent pair, the mobile phase operating parameters, and the first retention time, the second retention time for each of the plurality of candidate chromatographic column types. The separation analysis module is configured to invoke the fourth agent to calculate, based on the retention time, a peak separation for each of the plurality of candidate chromatographic column types.
12. The system of claim 1, wherein, The retention time includes a first retention time and a second retention time, 13. The system of claim 12, wherein, wherein the separation analysis module is configured to invoke the fourth agent to calculate, based on the first retention time corresponding to a first configuration of the compound to be predicted and a first half-peak width and the second retention time corresponding to a second configuration of the compound to be predicted and a second half-peak width, the peak separation for each of the plurality of candidate chromatographic column types. 14. The system of claim 1, wherein, The known compounds and corresponding column type and solvent pairs are obtained from a database, which is obtained by a large language model by: obtaining literature related to high-performance liquid chromatography; extracting known compounds and corresponding column type and solvent pairs from the literature; storing the known compounds and corresponding column type and solvent pairs in the database.
15. A method for recommending chiral separation conditions, characterized in that, comprises: S100: receiving an expression of a compound to be predicted; S200: based on the expression, calculating a similarity between the compound to be predicted and known compounds; S300: determining a plurality of candidate column types and determining a candidate solvent pair according to the similarity; S400: for each candidate column type in the plurality of candidate column types, using a first machine learning algorithm to predict a mobile phase operating parameter of the candidate solvent pair based on the expression, the candidate column type and the candidate solvent pair; S500: for each candidate column type in the plurality of candidate column types, using a second machine learning algorithm to predict a retention time associated with a chromatographic peak of a chromatogram of the compound to be predicted using the candidate column type, the candidate solvent pair and the mobile phase operating parameter based on the expression, the candidate column type, the candidate solvent pair and the mobile phase operating parameter; S600: performing a resolution analysis associated with the compound to be predicted based on the retention time; and S700: reordering the plurality of candidate column types based on a result of the resolution analysis, wherein a recommended chiral separation condition is determined based on the reordered candidate column types and associated candidate solvent pairs and mobile phase operating parameters. The S200 comprises:
16. The method of claim 15, wherein, S201: converting the expression of the compound to be predicted into a first molecular descriptor; S202: converting the string of the known compound into a second molecular descriptor; and S203: calculating a similarity between the first molecular descriptor and the second molecular descriptor. The S300 comprises:
17. The method of claim 15, wherein, S301: according to the similarity, first ordering the column types corresponding to the known compounds and second ordering the solvent pairs corresponding to the known compounds; S302: determining the plurality of candidate column types based on a result of the first ordering; and S303: determining the candidate solvent pair based on a result of the second ordering. The S300 further comprises:
18. The method of claim 17, wherein, S304: when a similarity of a column type with the highest similarity in the plurality of candidate column types is less than a similarity threshold, generating a predicted column type and a predicted solvent pair based on the expression using a third machine learning algorithm; S305: replacing the column type with the highest similarity with the predicted column type; and S306: replacing the candidate solvent pair with the predicted solvent pair. The S400 comprises:
19. The method of claim 15, wherein, S401: performing feature engineering on the expression, the candidate column type and the candidate solvent pair to splice into a condition vector; and S402: inputting the condition vector into the first machine learning algorithm to obtain the mobile phase operating parameter of the candidate solvent pair. S402: predicting, based on the condition vector, the mobile phase operating parameters of the candidate solvent pair using the first machine learning algorithm.
20. The method of claim 19, wherein, The mobile phase operating parameters include solvent ratio and solvent flow rate, and the S402 includes: S4021: predicting, for each candidate chromatographic column type in the plurality of candidate chromatographic column types, the solvent ratio of the candidate solvent pair based on the condition vector using the first machine learning algorithm; S4022: predicting, for each candidate chromatographic column type in the plurality of candidate chromatographic column types, the solvent flow rate of the candidate solvent pair based on the condition vector and the solvent ratio using the first machine learning algorithm.
21. The method of claim 15, wherein, The retention time includes a first retention time and a second retention time, and the S500 includes: S501: predicting, for each candidate chromatographic column type in the plurality of candidate chromatographic column types, the first retention time corresponding to the first configuration of the compound to be predicted based on the expression, the candidate chromatographic column type, the candidate solvent pair, and the mobile phase operating parameters using the second machine learning algorithm; S502: predicting, for each candidate chromatographic column type in the plurality of candidate chromatographic column types, the second retention time corresponding to the second configuration of the compound to be predicted based on the expression, the candidate chromatographic column type, the candidate solvent pair, the mobile phase operating parameters, and the first retention time using the second machine learning algorithm, wherein the resolution analysis is performed based on the first retention time and the second retention time.
22. The method of claim 15, wherein, The retention time includes a first retention time and a second retention time, and the S600 includes: calculating, for each candidate chromatographic column type in the plurality of candidate chromatographic column types, a peak resolution based on the first retention time and the first half-peak width corresponding to the first configuration of the compound to be predicted and the second retention time and the second half-peak width corresponding to the second configuration of the compound to be predicted.
23. The method of claim 15, wherein, The known compounds and corresponding chromatographic column types and solvent pairs are obtained from a database, and the database is obtained by a large language model in the following way: obtaining literature related to high-performance liquid chromatography; extracting known compounds and corresponding chromatographic column types and solvent pairs from the literature; storing the known compounds and corresponding chromatographic column types and solvent pairs in the database.
24. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the method of any one of claims 15-23.
25. A computer program product comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the method of any one of claims 15-23.
Citation Information
Patent Citations
Method for separating and determining florfenicol enantiomer by ultra-performance convergence chromatography
CN114460193A
Multimodal chromatographic separation media and process for using same
US5316680A