Systems, methods, media, and products for recommending chiral separation conditions

By recommending chiral separation conditions through machine learning algorithms, the problems of time-consuming, labor-intensive, and costly processes in existing technologies are solved, achieving efficient and low-cost chiral separation and improving the separation success rate.

CN120998345BActive Publication Date: 2026-01-27SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511525307.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2026-01-27
Estimated Expiration
2045-10-24

AI Technical Summary

Technical Problem

Existing technologies rely on manual selection of column type and solvent pair in chiral separation, which is time-consuming, labor-intensive, and costly, and cannot guarantee the success rate of separation. They also ignore the dependence of mobile phase operating parameters on residence time.

Method used

Machine learning algorithms are used to recommend chiral separation conditions. Through similarity detection, column recommendation, solvent pair recommendation, operating parameter prediction, residence time prediction and resolution analysis modules, chiral separation conditions are calculated by combining machine learning algorithms and chromatographic theory.

Benefits of technology

This improved the success rate of chiral separation experiments, reduced manual verification time, lowered experimental costs, and ensured separation effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998345B_ABST
    Figure CN120998345B_ABST
Patent Text Reader

Abstract

The present application relates to a computer system utilizing a computational model, and discloses a system, method, medium and product for recommending chiral separation conditions. A system for recommending chiral separation conditions comprises: a similarity detection module for calculating similarity between a to-be-predicted compound and a known compound; a column recommendation module for determining a plurality of candidate column types according to the similarity; a solvent pair recommendation module for determining a candidate solvent pair according to the similarity; a running parameter prediction module for predicting a mobile phase running parameter of the candidate solvent pair; a residence time prediction module for predicting a residence time; a resolution analysis module for performing resolution analysis based on the residence time; and a reordering module for reordering the plurality of candidate column types. The system according to the present application overcomes the limitation of the prior art that the column type cannot be recommended, and improves the experimental success rate of chiral separation conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention generally relates to computer systems utilizing computational models, and more specifically to systems, methods, media, and products for recommending chiral separation conditions. Background Technology

[0002] In the fields of materials research and development, drug development and fine chemicals, the optical enantiomer separation of chiral molecules is a crucial but extremely challenging task. Figure 1 These are examples of different configurations of chiral molecules and their associated chromatograms. Taking a chiral binuclear iridium complex as an example, for the same structural formula, there are two mirror-symmetric configurations, R and S. Different configurations can have different chemical properties; for example, one configuration may have pharmaceutical activity while the other has physiological toxicity. Therefore, it is necessary to separate different enantiomers of the same compound.

[0003] High-performance liquid chromatography (HPLC), especially chiral stationary phase chromatography, is currently the most commonly used and effective chiral separation method. In traditional HPLC chiral separation methods, the chiral sample to be separated is dissolved in a mobile phase composed of two solvents. When the mobile phase carrying the sample flows through a chiral chromatographic column (stationary phase), due to the different binding abilities of different enantiomer molecules in the sample to the stationary phase, the stability of the diastereomeric complexes formed varies, resulting in different residence times for the two enantiomers in the column. The more strongly bound enantiomer has a longer residence time in the column, and the less strongly bound enantiomer has a shorter residence time. Ultimately, the two enantiomers elute sequentially from the end of the column, achieving physical separation. When the separated enantiomers enter the detector, the detector converts the enantiomer concentration signal into a chromatogram. Figure 1 As shown, chiral binuclear iridium complexes exhibit two chromatographic peaks in the chromatogram, each with a corresponding residence time. When the chromatographic peaks are distinguishable by their residence times, separation can be considered successful.

[0004] Therefore, the core of the HPLC method lies in finding the optimal combination of "column-solvent pair (mobile phase)-mobile phase operating parameters" for a specific chiral molecule.

[0005] However, due to the abundance of commercially available column types and solvents, and the need for manual setting of mobile phase parameters for solvent pairs, the separation conditions for specific chiral molecules heavily rely on the expertise of chemists and extensive trial-and-error experiments, requiring the testing of each column type, solvent pair, and mobile phase parameter. This manual testing process is not only time-consuming and labor-intensive, but also involves high costs for the required columns and solutions, and cannot guarantee a high success rate for separation.

[0006] Existing techniques provide a method for predicting separation conditions based on residence time. This method involves manually selecting a specific column type and then predicting the residence time. By iterating through the predicted residence times for all column types, it determines whether the separation conditions are suitable for the chiral separation of the compound. However, due to the diverse combinations of column type, mobile phase type, and mobile phase operating parameters, manually selecting these parameters leads to low optimization efficiency. Furthermore, for a novel molecule, selecting the initial experimental conditions (especially the column type and solvent pair) is the biggest challenge, often relying on expert experience. Moreover, existing methods only predict single parameters such as residence time, ignoring the dependency between mobile phase operating parameters and residence time.

[0007] There is a need in this field for chiral separation condition recommendation techniques that can be improved at at least one of the above levels. Summary of the Invention

[0008] This invention is provided to offer a technique for recommending methods to further improve chiral separation conditions by using machine learning algorithms to predict column type, solvent pair, and mobile phase operating parameters.

[0009] One aspect of the present invention provides a system for recommending chiral separation conditions, comprising: a similarity detection module for receiving an expression of a compound to be predicted and invoking a first agent to: calculate a similarity between the compound to be predicted and a known compound based on the expression; a column recommendation module for invoking a second agent to: determine a plurality of candidate column types based on the similarity; a solvent pair recommendation module for invoking a third agent to: determine candidate solvent pairs based on the similarity; a running parameter prediction module for predicting, for each of the plurality of candidate column types, the mobile phase running parameters of the candidate solvent pair using a first machine learning algorithm based on the expression, the candidate column type, and the candidate solvent pair; and a residence time prediction module. The system includes a resolution analysis module for predicting, for each of the plurality of candidate column types, the residence time associated with the chromatographic peaks of the chromatographic peaks obtained by the predicted compound using the candidate column type, the candidate solvent pair, and the mobile phase operating parameters, based on the expression, the candidate column type, the candidate solvent pair, and the mobile phase operating parameters; a resolution analysis module for invoking a fourth agent to perform a resolution analysis associated with the predicted compound based on the residence time; and a reordering module for reordering the plurality of candidate column types based on the results of the resolution analysis, wherein recommended chiral separation conditions are determined based on the reordered candidate column types and associated candidate solvent pairs and mobile phase operating parameters.

[0010] In the system described above, the similarity detection module is used to invoke the first agent to: convert the expression of the compound to be predicted into a first molecular descriptor; convert the string of the known compound into a second molecular descriptor; and calculate the similarity between the first molecular descriptor and the second molecular descriptor.

[0011] In the system described above, the column recommendation module is used to invoke the second agent to: sort the column types corresponding to the known compounds according to the similarity; and determine the plurality of candidate column types based on the sorting results.

[0012] In the system described above, the column recommendation module is further configured to invoke the second agent to: generate a predicted column type based on the expression when the similarity of the column type with the highest similarity among the plurality of candidate column types is less than a similarity threshold; and replace the column type with the highest similarity with the predicted column type.

[0013] In the system described above, the solvent pair recommendation module is used to invoke the third agent to: sort the solvent pairs corresponding to the known compounds according to the similarity; and determine the candidate solvent pairs based on the sorting results.

[0014] In the system described above, the solvent pair recommendation module is further configured to invoke the third agent to: generate a predicted solvent pair based on the expression when the similarity of the candidate solvent pair is less than a similarity threshold; and replace the candidate solvent pair with the predicted solvent pair.

[0015] In the system described above, the operating parameter prediction module is used to: perform feature engineering on the expression, the candidate column type, and the candidate solvent pair to construct a condition vector; and, based on the condition vector, use the first machine learning algorithm to predict the mobile phase operating parameters of the candidate solvent pair.

[0016] In the system described above, the operating parameter prediction module is used to: convert the expression into a first molecular descriptor; convert the string of the candidate solvent pair into a second molecular descriptor; convert the candidate column type into an unique thermal code; and concatenate the first molecular descriptor, the second molecular descriptor, and the unique thermal code into the condition vector.

[0017] In the system described above, the mobile phase operating parameters include solvent ratio and solvent flow rate. The operating parameter prediction module is configured to invoke a fifth agent to predict the solvent ratio of the candidate solvent pair for each of the plurality of candidate column types, based on the condition vector. The operating parameter prediction module is also configured to invoke a sixth agent to predict the solvent flow rate of the candidate solvent pair for each of the plurality of candidate column types, based on the condition vector and the solvent ratio.

[0018] In the system described above, the residence time includes a first residence time and a second residence time, wherein the residence time prediction module is used to: use the second machine learning algorithm to predict the first residence time corresponding to the first configuration of the compound to be predicted and the second residence time corresponding to the second configuration of the compound to be predicted, wherein the separation degree analysis module is used to call the fourth agent to: perform the separation degree analysis based on the first residence time and the second residence time.

[0019] In the system described above, the residence time prediction module is used to invoke a seventh agent to: predict the first residence time for each of the plurality of candidate column types based on the expression, the candidate column type, the candidate solvent pair, and the mobile phase operating parameters; the residence time prediction module is used to invoke an eighth agent to: predict the second residence time for each of the plurality of candidate column types based on the expression, the candidate column type, the candidate solvent pair, the mobile phase operating parameters, and the first residence time.

[0020] In the system described above, the resolution analysis module is used to invoke the fourth agent to calculate peak resolution for each of the plurality of candidate column types based on the residence time.

[0021] In the system described above, the residence time includes a first residence time and a second residence time, wherein the resolution analysis module is used to invoke the fourth agent to: calculate the peak resolution for each of the plurality of candidate column types based on the first residence time and first half-peak width corresponding to the first configuration of the compound to be predicted, and the second residence time and second half-peak width corresponding to the second configuration of the compound to be predicted.

[0022] In the system described above, the known compounds and their corresponding column types and solvent pairs are obtained from a database, which is obtained by a large language model through the following methods: obtaining literature related to high performance liquid chromatography; extracting known compounds and their corresponding column types and solvent pairs from the literature; and storing the known compounds and their corresponding column types and solvent pairs in the database.

[0023] Another aspect of the present invention provides a method for recommending chiral separation conditions, comprising: S100: receiving an expression for a compound to be predicted; S200: calculating a similarity between the compound to be predicted and a known compound based on the expression; S300: determining a plurality of candidate column types and candidate solvent pairs based on the similarity; S400: for each of the plurality of candidate column types, using a first machine learning algorithm to predict the mobile phase operating parameters of the candidate solvent pair based on the expression, the candidate column type, and the candidate solvent pair; S500: for each of the plurality of candidate column types... S600: Based on the expression, the candidate column type, the candidate solvent pair, and the mobile phase operating parameters, a second machine learning algorithm is used to predict the residence time associated with the chromatographic peaks of the chromatogram obtained by the predicted compound using the candidate column type, the candidate solvent pair, and the mobile phase operating parameters; S700: Based on the residence time, a resolution analysis associated with the predicted compound is performed; and S700: Based on the results of the resolution analysis, the plurality of candidate column types are reordered, wherein recommended chiral separation conditions are determined based on the reordered candidate column types and associated candidate solvent pairs and mobile phase operating parameters.

[0024] The method described above includes, in step S200: S201: converting the expression of the compound to be predicted into a first molecular descriptor; S202: converting the string of the known compound into a second molecular descriptor; and S203: calculating the similarity between the first molecular descriptor and the second molecular descriptor.

[0025] The method described above, wherein S300 includes: S301: sorting the column types corresponding to the known compound in a first sorting order based on the similarity, and sorting the solvent pairs corresponding to the known compound in a second sorting order; S302: determining the plurality of candidate column types based on the result of the first sorting order; and S303: determining the candidate solvent pairs based on the result of the second sorting order.

[0026] The method described above further includes, in step S300: S304: when the similarity of the column type with the highest similarity among the plurality of candidate column types is less than a similarity threshold, generating a predicted column type and a predicted solvent pair based on the expression and using a third machine learning algorithm; S305: replacing the column type with the highest similarity with the predicted column type; and S306: replacing the candidate solvent pair with the predicted solvent pair.

[0027] As described above, step S400 includes: S401: performing feature engineering on the expression, the candidate column type, and the candidate solvent pair to construct a condition vector; and S402: using the first machine learning algorithm to predict the mobile phase operating parameters of the candidate solvent pair based on the condition vector.

[0028] As described above, the mobile phase operating parameters include solvent ratio and solvent flow rate. Step S402 includes: S4021: For each of the plurality of candidate column types, based on the condition vector, using the first machine learning algorithm to predict the solvent ratio of the candidate solvent pair; S4022: For each of the plurality of candidate column types, based on the condition vector and the solvent ratio, using the first machine learning algorithm to predict the solvent flow rate of the candidate solvent pair.

[0029] The method described above, wherein the residence time includes a first residence time and a second residence time, and S500 includes: S501: for each of the plurality of candidate column types, based on the expression, the candidate column type, the candidate solvent pair, and the mobile phase operating parameters, using the second machine learning algorithm to predict the first residence time corresponding to the first configuration of the compound to be predicted; S502: for each of the plurality of candidate column types, based on the expression, the candidate column type, the candidate solvent pair, the mobile phase operating parameters, and the first residence time, using the second machine learning algorithm to predict the second residence time corresponding to the second configuration of the compound to be predicted, wherein the resolution analysis is performed based on the first residence time and the second residence time.

[0030] As described above, the residence time includes a first residence time and a second residence time, and step S600 includes: for each of the plurality of candidate column types, calculating peak resolution based on the first residence time and first half-peak width corresponding to the first configuration of the compound to be predicted, and the second residence time and second half-peak width corresponding to the second configuration of the compound to be predicted.

[0031] As described above, the known compounds and their corresponding column types and solvent pairs are obtained from a database, which is obtained by a large language model through the following methods: obtaining literature related to high performance liquid chromatography; extracting known compounds and their corresponding column types and solvent pairs from the literature; and storing the known compounds and their corresponding column types and solvent pairs in the database.

[0032] Another aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in any of the preceding claims.

[0033] Another aspect of the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in any of the above embodiments.

[0034] The system and method of the present invention overcome the limitation of the prior art in recommending column types, and improve the experimental success rate of chiral separation conditions by using machine learning algorithms to predict column type, solvent pair and mobile phase operating parameters. Attached Figure Description

[0035] Various embodiments of the present invention are described in conjunction with the accompanying drawings.

[0036] Figure 1 These are examples of different configurations of chiral molecules and their associated chromatograms.

[0037] Figure 2 This is a block diagram of a system for recommending chiral separation conditions according to some embodiments of the present invention.

[0038] Figure 3 This is a schematic diagram of a system for recommending chiral separation conditions according to some embodiments of the present invention.

[0039] Figure 4 This is a schematic diagram of a system for recommending chiral separation conditions according to other embodiments of the present invention.

[0040] Figure 5 This is a schematic diagram of recommended chiral separation conditions according to some embodiments of the present invention.

[0041] Figure 6 This is a flowchart of a method for recommending chiral separation conditions according to some embodiments of the present invention.

[0042] Figure 7 This is a flowchart of a first process associated with a method for recommending chiral separation conditions according to some embodiments of the present invention.

[0043] Figure 8 This is a flowchart of a second process associated with a method for recommending chiral separation conditions according to some embodiments of the present invention.

[0044] Figure 9 This is a flowchart of a third process associated with a method for recommending chiral separation conditions according to some embodiments of the present invention.

[0045] Figure 10 This is a flowchart of a fourth process associated with a third process of a method for recommending chiral separation conditions, according to some embodiments of the present invention.

[0046] Figure 11 This is a flowchart of a fifth process associated with a third process of a method for recommending chiral separation conditions, according to some embodiments of the present invention.

[0047] Figure 12 This is a flowchart of a sixth process associated with a method for recommending chiral separation conditions according to some embodiments of the present invention.

[0048] Figure 13 This is a block diagram of a computer-readable storage medium according to some embodiments of the present invention.

[0049] Figure 14 This is a block diagram of a computer program product according to some embodiments of the present invention. Detailed Implementation

[0050] In this application, ordinal numbers such as "first," "second," and "third" are used to distinguish different instances of objects with the same name. The ordinal numbers "first," "second," and "third" do not indicate a relative order of the indicated objects in time, space, sequence, or other aspects.

[0051] In this invention, the term "agent" refers to an agent capable of perceiving its environment and taking actions to perform specific goals. An agent primarily refers to software code. Agents can be executed by the computing resources of a computing device. Agents can invoke corresponding models and tools through application programming interfaces (APIs) to interact with various forms of input or implement corresponding functions.

[0052] According to one aspect of the invention, a system for recommending chiral separation conditions is provided.

[0053] Figure 2This is a block diagram of a system 100 for recommending chiral separation conditions according to some embodiments of the present invention. System 100 includes a similarity detection module 102, a column recommendation module 104, a solvent pair recommendation module 106, a running parameter prediction module 108, a residence time prediction module 110, a resolution analysis module 112, and a reordering module 114. Figures 3-4 Further explanation of System 100.

[0054] Figure 3 This is a schematic diagram of a system 100 for recommending chiral separation conditions according to some embodiments of the present invention.

[0055] The similarity detection module 102 can be used to receive the expression of the compound to be predicted and call the first agent 122 to calculate the similarity between the compound to be predicted and known compounds based on the expression.

[0056] For example, a user can input an expression for the compound to be predicted through a user interface, and the similarity detection module 102 can calculate the similarity between the expression for the compound to be predicted and the string of a known compound.

[0057] In some embodiments, the similarity detection module 102 may be used to invoke the first agent 122 to convert the expression of the compound to be predicted into a first molecular descriptor, convert the string of a known compound into a second molecular descriptor, and calculate the similarity between the first molecular descriptor and the second molecular descriptor.

[0058] For example, the expression of the compound to be predicted can be a molecular structure diagram, a molecular structural formula, or a string of SMILES. The similarity detection module 102 can convert the expression of the compound to be predicted into a first molecular descriptor, such as a Morgan fingerprint. The string of known compounds can be a string of SMILES. The similarity detection module 102 can convert the string of known compounds into a second molecular descriptor, such as a Morgan fingerprint. Then, the similarity detection module 102 can calculate the similarity between the first molecular descriptor of the compound to be predicted and the second molecular descriptor of each known compound. As an example, the similarity can be the Tanimoto similarity, which measures the similarity of molecular structures.

[0059] The column recommendation module 104 can be used to invoke the second agent 124 to determine multiple candidate column types based on similarity.

[0060] The available column types include various column types from different manufacturers, potentially numbering in the hundreds. For multiple known compounds, a corresponding column type is pre-set for each known compound. The column types for different compounds may be the same or different. The column recommendation module 104 can determine multiple column types that may be suitable for the compound to be predicted from all known compound column types based on similarity.

[0061] Since column type is the most critical factor in separation experiments, and there are numerous column types, automatically identifying multiple candidate column types based on similarity can significantly narrow the experimental range of separation conditions, greatly improve the efficiency of column type confirmation, increase the success rate of separation experiments, and save manpower and time.

[0062] In some embodiments, the column recommendation module 104 may be used to invoke a second agent 124 to sort the column types corresponding to known compounds according to similarity, and to determine multiple candidate column types based on the sorting results.

[0063] For example, the column recommendation module 104 can sort multiple known compounds based on similarity and correspondingly sort their column types, where the same column type corresponding to different known compounds is sorted by the highest similarity, and then multiple candidate column types are determined from the higher-ranked column types. As an example, the top 5 column types can be selected as candidate column types from the column types sorted by similarity.

[0064] In some embodiments, the column recommendation module 104 may also be used to invoke the second agent 124 so that when the similarity of the column type with the highest similarity among multiple candidate column types is less than a similarity threshold, a predicted column type is generated based on the expression of the compound to be predicted, and the predicted column type is used to replace the column type with the highest similarity.

[0065] For example, when the similarity of the column type with the highest similarity is 0.3, and the similarity threshold is 1, the similarity is much lower than the threshold, meaning the known compound with the highest similarity may be significantly different from the compound to be predicted. By invoking an agent to generate a predicted column type based on the expression of the compound to be predicted, instead of a candidate column type with a large difference and a similarity of 0.3, it is possible to enhance the separation effect of separation experiments. The similarity and similarity threshold in the above example are merely examples; users can set the similarity threshold according to their actual needs, and the scope of this disclosure is not limited thereto.

[0066] As an example, the second agent 124 can invoke a classification model to generate a predicted column type, such as a large language model, a graph neural network model, etc.

[0067] The solvent pair recommendation module 106 can be used to invoke a third agent 126 to determine candidate solvent pairs based on similarity.

[0068] The available solvent pairs include up to dozens. For multiple known compounds, a corresponding solvent pair is pre-set for each known compound. The solvent pair can be expressed as "first solvent / second solvent". Solvent pairs between different compounds may be the same or different. The solvent pair recommendation module 106 can determine solvent pairs that may be suitable for the compound to be predicted from all known compound solvent pairs based on similarity.

[0069] Since solvents can affect the binding ability of chiral molecules to the chromatographic column, they are a key factor in separation experiments. By using an intelligent agent to automatically identify candidate solvent pairs, the time spent on manual verification of candidate solvent pairs can be reduced, thereby improving the success rate of separation experiments.

[0070] In some embodiments, the solvent pair recommendation module 106 may be used to invoke a third agent 126 to sort solvent pairs corresponding to known compounds based on similarity, and to determine candidate solvent pairs based on the sorting results.

[0071] For example, the solvent pair recommendation module 106 can sort multiple known compounds according to similarity, and correspondingly sort their solvent pairs, then determine the highest-ranked solvent pair as a candidate solvent pair. Optionally, the solvent pair recommendation module 106 can determine multiple high-ranking solvent pairs as multiple candidate solvent pairs to increase the combination of separation conditions and further improve the success rate of separation experiments.

[0072] In some embodiments, the solvent pair recommendation module 106 may also be used to invoke a third agent 126 to generate a predicted solvent pair based on an expression when the similarity of the candidate solvent pair is less than a similarity threshold, and to replace the candidate solvent pair with the predicted solvent pair.

[0073] For example, when the similarity of candidate solvent pairs is 0.3 and the similarity threshold is 1, the similarity is much lower than the threshold, meaning the known compound with the highest similarity may differ significantly from the compound to be predicted. By invoking an agent to generate predicted solvent pairs based on the expression of the compound to be predicted, instead of candidate solvent pairs with a similarity of 0.3 that are significantly different, it is possible to enhance the separation effect of separation experiments. The similarity and similarity threshold in the above example are merely illustrations; users can set the similarity threshold according to their actual needs, and the scope of this disclosure is not limited thereto.

[0074] As an example, the third agent 126 can invoke a classification model to generate predicted solvent pairs, such as a large language model, a graph neural network model, etc.

[0075] The operating parameter prediction module 108 can be used to predict the mobile phase operating parameters of a candidate solvent pair for each of a plurality of candidate column types, based on the expression of the compound to be predicted, the candidate column type, and the candidate solvent pair, using a first machine learning algorithm.

[0076] After identifying candidate solvent pairs, the mobile phase operating parameters of these pairs can also affect the separation results, such as the residence time, peak spacing, and peak width of the two configurations. Therefore, accurate prediction of the mobile phase operating parameters of candidate solvent pairs can improve the separation experiment results.

[0077] As an example, the first machine learning algorithm could be a regression model, such as Extreme Gradient Boosting (XGBoost), Tabnet, Category Boosting (CatBoost), Gradient Boosting Machine (LightGBM), Random Forest, Deep Neural Network (DNN), Graph Neural Network (GNN), Transformer, etc.

[0078] In some embodiments, the operating parameter prediction module 108 can be used to perform feature engineering on the expression of the compound to be predicted, the candidate column type, and the candidate solvent pair to form a condition vector, and based on the condition vector, use a first machine learning algorithm to predict the mobile phase operating parameters of the candidate solvent pair.

[0079] Feature engineering converts the expression of the compound to be predicted, candidate column types, and candidate solvent pairs into a condition vector, enabling the computational model to read the condition vector.

[0080] In some embodiments, the runtime parameter prediction module 108 can be used to convert the expression of the compound to be predicted into a first molecular descriptor, convert the string of the candidate solvent pair into a second molecular descriptor, convert the candidate column type into an one-hot encoding, and concatenate the first molecular descriptor, the second molecular descriptor, and the one-hot encoding into a condition vector. For example, the molecular descriptor may include one or more of the following descriptors used to describe molecular structure or related information: MolT5 embedding vector, Morgan fingerprint, MACCS Keys (Molecular Access System Keys) molecular fingerprint, etc.

[0081] The residence time prediction module 110 can be used for each of a plurality of candidate column types to predict the residence time associated with the chromatographic peaks of ...

[0082] As an example, the second machine learning algorithm can be a regression model, such as Extreme Gradient Boosting (XGBoost), Tabnet, Category Boosting (CatBoost), Gradient Boosting Machine (LightGBM), Random Forest, Deep Neural Network (DNN), Graph Neural Network (GNN), Transformer, etc.

[0083] The separation analysis module 112 can be used to invoke the fourth agent 128 to perform a separation analysis associated with the compound to be predicted based on the residence time. When the chromatographic peaks of the chromatogram of the compound to be predicted are separated from each other under specific separation conditions, it can be considered that the two configurations corresponding to the compound to be predicted can be separated.

[0084] In some embodiments, the resolution analysis module 112 may be used to invoke a fourth agent 128 to calculate peak resolution based on residence time for each of a plurality of candidate column types.

[0085] For two chromatographic peaks in a chromatogram under specific separation conditions, even though the residence times of the two configurations are different, overlap may still occur if the peaks are broad. Therefore, calculating peak resolution based on residence time can further ensure that the recommended separation conditions can achieve enantiomer separation.

[0086] In some embodiments, the residence time may include a first residence time and a second residence time. The residence time prediction module 110 may be used to use a second machine learning algorithm to predict the first residence time corresponding to a first configuration of the compound to be predicted and the second residence time corresponding to a second configuration of the compound to be predicted. The separation analysis module 112 may be used to invoke a fourth agent 128 to perform separation analysis based on the first residence time and the second residence time.

[0087] In some embodiments, the residence time may include a first residence time and a second residence time. The resolution analysis module 112 may be used to invoke a fourth agent 128 to calculate peak resolution for each of a plurality of candidate column types, based on a first residence time and a first half-peak width corresponding to a first configuration of the compound to be predicted, and a second residence time and a second half-peak width corresponding to a second configuration of the compound to be predicted.

[0088] Since the residence time and half-width at half-maximum (HWHM) of each configuration can be considered to have a linear relationship, the HWHM of each configuration can be determined based on the residence time, and then the peak separation can be calculated based on the residence time and HWHM. The formula for calculating peak separation is as follows:

[0089] ,

[0090] Where Rs represents peak separation, RT1 and RT2 are the first and second residence times, respectively, W_half1 and W_half2 are the first and second half-peak widths, respectively, and c is a preset coefficient. The half-peak width can be estimated as:

[0091] ,

[0092] Where W_half is the half-maximum width, RT is the residence time, and a and b are preset coefficients. The preset coefficients a, b, and c can use the system default parameters or can be set according to the actual needs of the user in different experimental scenarios.

[0093] By combining residence time and half-peak width to calculate peak resolution, a more accurate prediction of the separation effect can be provided, ensuring the actual separation results in the separation experiment.

[0094] The reordering module 114 can be used to reorder multiple candidate column types based on the results of resolution analysis. Recommended chiral separation conditions are determined based on the reordered candidate column types and associated candidate solvent pairs and mobile phase operating parameters.

[0095] For example, given multiple candidate column types previously sorted by similarity and their corresponding separation conditions (e.g., mobile phase operating parameters including candidate solvent pairs), determine whether the peak resolution for each column type is greater than or equal to a resolution threshold. The resolution threshold can be adjusted according to user needs. For one or more column types with peak resolution greater than or equal to the resolution threshold, sort these column types and their corresponding separation conditions by similarity. For one or more column types with peak resolution less than the resolution threshold, sort these column types and their corresponding separation conditions by similarity. Column types and their corresponding separation conditions with peak resolution greater than or equal to the resolution threshold can have higher priority in subsequent experiments.

[0096] By combining machine learning algorithms with chromatographic theoretical calculation methods—that is, first using machine learning algorithms to predict intermediate parameters such as residence time, then using theoretical formulas to calculate the final performance index of resolution, and using the resolution to reorder candidate separation conditions—further optimization of the ranking of separation conditions can be achieved with the final separation effect as the guide.

[0097] In some embodiments, known compounds and their corresponding column types and solvent pairs can be obtained from a database, which can be obtained by a large language model in the following ways: obtaining literature related to high performance liquid chromatography; extracting known compounds and their corresponding column types and solvent pairs from the literature; and storing the known compounds and their corresponding column types and solvent pairs in the database.

[0098] Figure 4 This is a schematic diagram of a system 100 for recommending chiral separation conditions according to other embodiments of the present invention. To avoid redundancy, Figure 4 and Figure 3 The specific details of the same elements will not be elaborated here.

[0099] In some embodiments, mobile phase operating parameters may include solvent ratio and solvent flow rate. The operating parameter prediction module 108 may be used to invoke a fifth agent 130 to predict the solvent ratio of candidate solvent pairs for each of a plurality of candidate column types, based on a condition vector. The operating parameter prediction module 108 may also be used to invoke a sixth agent 132 to predict the solvent flow rate of candidate solvent pairs for each of a plurality of candidate column types, based on a condition vector and solvent ratio.

[0100] Solvent ratios can be predicted first based on condition vectors (including the expression of the compound to be predicted, candidate column types, and candidate solvent pairs). For example, the solvent ratio can be the percentage of the solvent with a higher proportion in the candidate solvent pair. Then, based on the condition vectors and the predicted solvent ratios, the solvent flow rate can be further predicted. This cascaded prediction architecture fully considers the dependence of downstream parameters on upstream parameters, achieves accurate simulation of the internal logic of chemical experiments, and improves prediction accuracy.

[0101] In some embodiments, the residence time prediction module 110 may be used to invoke a seventh agent 134 to predict a first residence time for each of a plurality of candidate column types, based on an expression, the candidate column type, a candidate solvent pair, and mobile phase operating parameters. The residence time prediction module 110 may also be used to invoke an eighth agent 136 to predict a second residence time for each of a plurality of candidate column types, based on an expression, the candidate column type, a candidate solvent pair, mobile phase operating parameters, and the first residence time.

[0102] The first residence time can be predicted based on the expression of the compound to be predicted, the candidate column type, the candidate solvent pair, and the mobile phase operating parameters. Then, based on the expression of the compound to be predicted, the candidate column type, the candidate solvent pair, the mobile phase operating parameters, and the predicted first residence time, the second residence time can be predicted to fully consider the intrinsic correlation between the two residence times and improve the accuracy of residence time prediction.

[0103] After obtaining the recommended chiral separation conditions from System 100, the user can conduct experiments according to the separation conditions in sequence to separate the compound to be predicted. For example, if the chromatographic peaks in the chromatogram are not completely separated when the separation experiment is conducted according to the first separation condition, i.e., the separation fails, then another separation experiment is conducted according to the next separation condition. When the chromatographic peaks in the obtained chromatogram are separated, the separation conditions are confirmed.

[0104] In some embodiments, the first agent 122, the second agent 124, the third agent 126, the fourth agent 128, the fifth agent 130, the sixth agent 132, the seventh agent 134, and the eighth agent 136 can call corresponding models via APIs. As examples, models may include classification models, gradient boosting tree models, regression models, deep neural networks, large language models, multimodal models, multimodal language models, etc. In some embodiments, multiple models that can be called by multiple agents can be deployed locally on system 100. In some embodiments, multiple models that can be called by multiple agents can be deployed remotely on system 100, for example, in the cloud. In some embodiments, some of the multiple models that can be called by multiple agents can be deployed locally on system 100, while others can be deployed remotely on system 100. In some embodiments, each of the multiple agents can call various tools to interact with various forms of input or to implement corresponding functions.

[0105] Figure 5 This is a schematic diagram of recommended chiral separation conditions according to some embodiments of the present invention. For the compound to be predicted (target molecule) shown in the figure, the predicted candidate solvent pair (optimal solvent pair) is hexane and i-PrOH. Multiple candidate column types are ordered by similarity as IA, AD, IG, IE, and OD. Users can experiment sequentially according to the listed column type, solvent ratio, and flow rate, in order of ranking, until the compound to be predicted is successfully separated. When the resolution threshold is set to 14.5, since the resolution of OD (14.93) is greater than the resolution threshold, OD can be ranked first. Experimental verification shows that when the column type is selected as OD, the compound to be predicted can be successfully separated.

[0106] Some embodiments of this invention, based on user-input target molecular structures, utilize rapid recommendation based on molecular structure similarity and refined prediction based on machine learning to achieve fast, accurate, and fully automated recommendations for chiral separation conditions. Compared to traditional methods that require days or even weeks, some embodiments of this invention can complete high-quality computational simulations and scheme recommendations within minutes, significantly improving R&D efficiency.

[0107] As an example, the system according to some embodiments of the present invention built a database and prediction model for known compounds based on 33,000 samples, and tested it with 3,700 samples as compounds to be predicted. Experimental results showed that the TOP-2 accuracy was 82%, the TOP-3 accuracy was 90%, and the TOP-5 accuracy was 95%. Experiments demonstrate that the system according to some embodiments of the present invention can provide highly accurate recommendations for chiral separation conditions of the compounds to be predicted. Furthermore, by using the system to predict compounds never before reported in the industry, the system according to some embodiments of the present invention can recommend separation conditions for multiple unreported compounds, and these compounds can be successfully separated under two or more recommended separation conditions, further confirming the accuracy of the system's recommendations for chiral separation conditions according to some embodiments of the present invention.

[0108] According to another aspect of the present invention, a method for recommending chiral separation conditions is provided.

[0109] Figure 6 This is a flowchart of a method for recommending chiral separation conditions according to some embodiments of the present invention.

[0110] The method may include step S100: receiving an expression for the compound to be predicted.

[0111] The method may include step S200: calculating the similarity between the compound to be predicted and known compounds based on an expression.

[0112] The method may include step S300: determining multiple candidate column types and candidate solvent pairs based on similarity.

[0113] The method may include step S400: for each of a plurality of candidate column types, using a first machine learning algorithm to predict the mobile phase operating parameters of the candidate solvent pair based on an expression, the candidate column type, and the candidate solvent pair.

[0114] The method may include step S500: for each of a plurality of candidate column types, using a second machine learning algorithm based on an expression, the candidate column type, the candidate solvent pair, and the mobile phase operating parameters, to predict the residence time associated with the chromatographic peak of the compound to be predicted using the candidate column type, the candidate solvent pair, and the mobile phase operating parameters.

[0115] The method may include step S600: performing a separation analysis associated with the compound to be predicted based on the residence time.

[0116] The method may include step S700: reordering multiple candidate column types based on the results of resolution analysis, wherein recommended chiral separation conditions are determined based on the reordered candidate column types and associated candidate solvent pairs and mobile phase operating parameters.

[0117] Some embodiments of this invention propose an end-to-end, fully automated prediction system and method, from the molecular structure of the compound to be predicted to a complete, quantitative, multi-dimensional (including column type, solvent pair, mobile phase operating parameters, residence time, and resolution) chiral separation experimental protocol. By combining rapid, coarse-grained recommendations based on molecular structure similarity with slow, fine-grained quantitative predictions based on machine learning, accurate prediction of chiral separation conditions is achieved. Compared to traditional methods that require manual determination of the column type, some embodiments of this invention can automatically recommend the column type, helping users more efficiently identify this key factor that significantly impacts the separation results.

[0118] Furthermore, traditional methods may prioritize separation conditions with longer residence times, but fail to consider the resolution between the two chromatographic peaks corresponding to the two configurations of the compound being predicted. This results in overlap between the two chromatographic peaks in the experimental results, making complete separation impossible. Some embodiments of the present invention reorder separation conditions based on the results of resolution analysis, ensuring that the preferred separation conditions recommended to the user are the most likely to achieve good separation. This significantly improves the success rate of the first experiment and reduces the number of experiments and time required.

[0119] Furthermore, some embodiments of the present invention provide users with multiple chiral separation conditions corresponding to multiple candidate column types, enabling users to conduct separation experiments in a planned manner, rather than blindly testing column type, solvent pair and mobile phase operating parameters, providing a faster experimental path to achieve successful separation.

[0120] Moreover, some embodiments of the present invention can significantly reduce the consumption of expensive chiral chromatographic columns and high-purity solvents, saving experimental instrument time and labor costs.

[0121] Figure 7This is a flowchart of a first process associated with a method for recommending chiral separation conditions according to some embodiments of the present invention. This first process may be... Figure 6 The specific implementation of step S200 in the method is described, but the scope of the present invention is not limited thereto.

[0122] The first process may include step S201: converting the expression of the compound to be predicted into a first molecular descriptor.

[0123] The first process may include step S202: converting a string of known compounds into a second molecular descriptor.

[0124] The first process may include step S203: calculating the similarity between the first molecular descriptor and the second molecular descriptor.

[0125] Figure 8 This is a flowchart of a second process associated with a method for recommending chiral separation conditions according to some embodiments of the present invention. This second process may be... Figure 6 The specific implementation of step S300 in the method is described, but the scope of the present invention is not limited thereto.

[0126] The second process may include step S301: sorting the column types corresponding to the known compounds in a first sorting based on similarity, and sorting the solvent pairs corresponding to the known compounds in a second sorting.

[0127] The second process may include step S302: determining multiple candidate column types based on the results of the first sorting.

[0128] The second process may include step S303: determining candidate solvent pairs based on the results of the second sorting.

[0129] In some embodiments, the second process may further include step S304: when the similarity of the column type with the highest similarity among multiple candidate column types is less than a similarity threshold, a predicted column type and a predicted solvent pair are generated based on an expression and using a third machine learning algorithm.

[0130] The second process may also include step S305: replacing the column type with the highest similarity with the predicted column type.

[0131] The second process may also include step S306: replacing the candidate solvent pair with the predicted solvent pair.

[0132] Figure 9 This is a flowchart of a third process associated with a method for recommending chiral separation conditions according to some embodiments of the present invention. This third process may be... Figure 6 The specific implementation of step S400 in the method is shown, but the scope of the present invention is not limited thereto.

[0133] The third process may include step S401: performing feature engineering on the expression, the candidate column type, and the candidate solvent pair to splice them into a condition vector.

[0134] The third process may include step S402: using a first machine learning algorithm to predict the mobile phase operating parameters of candidate solvent pairs based on the condition vector.

[0135] Figure 10 This is a flowchart of a fourth process associated with a third process in a method for recommending chiral separation conditions, according to some embodiments of the present invention. This fourth process may be... Figure 9 The specific implementation of step S401 in the third process is described, but the scope of the present invention is not limited thereto.

[0136] The fourth process may include step S4011: converting the expression into a first molecule descriptor.

[0137] The fourth process may include step S4012: converting the strings of candidate solvent pairs into second molecule descriptors.

[0138] The fourth process may include step S4013: converting the candidate column type to an independent thermal code.

[0139] The fourth process may include step S4014: concatenating the first molecular descriptor, the second molecular descriptor, and the one-hot code into a condition vector.

[0140] Figure 11 This is a flowchart of a fifth process associated with a third process in a method for recommending chiral separation conditions, according to some embodiments of the present invention. This fifth process may be... Figure 9 The specific implementation of step S402 in the third process is described herein, but the scope of the invention is not limited thereto. In some embodiments, the mobile phase operating parameters may include solvent ratio and solvent flow rate.

[0141] The fifth process may include step S4021: for each of the multiple candidate column types, using a first machine learning algorithm based on a condition vector to predict the solvent ratio of the candidate solvent pair.

[0142] The fifth process may include step S4022: for each of the multiple candidate column types, using a first machine learning algorithm to predict the solvent flow rate of the candidate solvent pair based on the condition vector and solvent ratio.

[0143] Some embodiments of the present invention first predict the solvent ratio of the solvent pair, and then predict the solvent flow rate based on the solvent ratio. By using a cascaded prediction architecture, the physical dependencies between experimental parameters are fully considered, so that each prediction step is based on more complete contextual information. Therefore, the prediction results are more accurate and have more practical guiding significance than prediction models that predict individual parameters in isolation.

[0144] In some embodiments, the residence time may include a first residence time and a second residence time, and S500 may include: using a second machine learning algorithm to predict the first residence time corresponding to a first configuration of the compound to be predicted and the second residence time corresponding to a second configuration of the compound to be predicted, wherein the separation analysis is performed based on the first residence time and the second residence time.

[0145] Figure 12 This is a flowchart of a sixth process associated with a method for recommending chiral separation conditions according to some embodiments of the present invention. This sixth process may be... Figure 6 The specific implementation of step S500 in the method is described, but the scope of the present invention is not limited thereto.

[0146] The sixth process may include step S501: for each of the multiple candidate column types, using a second machine learning algorithm based on an expression, the candidate column type, the candidate solvent pair, and the mobile phase operating parameters, to predict the first residence time corresponding to the first configuration of the compound to be predicted.

[0147] The sixth process may include step S502: for each of the multiple candidate column types, based on the expression, the candidate column type, the candidate solvent pair, the mobile phase operating parameters, and the first residence time, using a second machine learning algorithm to predict the second residence time corresponding to the second configuration of the compound to be predicted.

[0148] Some embodiments of the present invention first predict the first residence time, and then predict the second residence time based on the first residence time. By using a cascaded prediction method, the physical dependence between the residence times of the chromatographic peaks of the two configurations is fully considered, so that each prediction step is based on more complete contextual information. Therefore, the prediction results are more accurate than the prediction model that predicts each parameter in isolation.

[0149] In some embodiments, S600 may include: calculating peak resolution based on residence time for each of a plurality of candidate column types.

[0150] In some embodiments, the residence time may include a first residence time and a second residence time, and S600 may include: for each of a plurality of candidate column types, calculating peak resolution based on a first residence time and a first half-peak width corresponding to a first configuration of the compound to be predicted and a second residence time and a second half-peak width corresponding to a second configuration of the compound to be predicted.

[0151] In some embodiments, known compounds and their corresponding column types and solvent pairs can be obtained from a database, which can be obtained by a large language model in the following ways: obtaining literature related to high performance liquid chromatography; extracting known compounds and their corresponding column types and solvent pairs from the literature; and storing the known compounds and their corresponding column types and solvent pairs in the database.

[0152] According to another aspect of the present invention, a computer-readable storage medium is provided.

[0153] Figure 13 This is a block diagram of a computer-readable storage medium 1300 according to some embodiments of the present invention.

[0154] A computer-readable storage medium 1300 stores a computer program 1350. When executed by a processor, the computer program 1350 implements the above-mentioned... Figures 6-12 The steps of each method or process described.

[0155] According to another aspect of the present invention, a computer program product is provided.

[0156] Figure 14 This is a block diagram of a computer program product 1400 according to some embodiments of the present invention.

[0157] Computer program product 1400 may include computer program 1350. Computer program 1350, when executed by a processor, implements the above-mentioned... Figures 6-12 The steps of each method or process described.

[0158] Embodiments of the invention have been described with reference to the accompanying drawings. These embodiments are illustrative and not restrictive.

Claims

1. A system for recommending chiral separation conditions, characterized in that, include: A similarity detection module is used to receive an expression of the compound to be predicted and to invoke a first agent to: calculate the similarity between the compound to be predicted and known compounds based on the expression; A column recommendation module is used to invoke a second agent to determine multiple candidate column types based on the similarity. The solvent pair recommendation module is used to invoke a third-party intelligent agent to determine candidate solvent pairs based on the similarity. The operating parameter prediction module is used to predict the mobile phase operating parameters of the candidate solvent pair for each of the plurality of candidate column types, based on the expression, the candidate column type and the candidate solvent pair, using a first machine learning algorithm. The residence time prediction module is used to predict, for each of the plurality of candidate column types, the residence time associated with the chromatographic peaks of the chromatographic peaks of the compound to be predicted using the expression, the candidate column type, the candidate solvent pair, and the mobile phase operating parameters, using a second machine learning algorithm. The separation analysis module is used to invoke a fourth agent to perform separation analysis associated with the compound to be predicted based on the residence time. as well as A reordering module is used to reorder the plurality of candidate column types based on the results of the resolution analysis, wherein recommended chiral separation conditions are determined based on the reordered candidate column types and associated candidate solvent pairs and mobile phase operating parameters.

2. The system as described in claim 1, characterized in that, The similarity detection module is used to invoke the first intelligent agent to: The expression of the compound to be predicted is converted into a first molecule descriptor; Convert the string of the known compound into a second molecule descriptor; as well as Calculate the similarity between the first molecular descriptor and the second molecular descriptor.

3. The system as described in claim 1, characterized in that, The column recommendation module is used to invoke the second agent to: Based on the similarity, the chromatographic column types corresponding to the known compounds are sorted; and The candidate column types are determined based on the sorting results.

4. The system as described in claim 3, characterized in that, The column recommendation module is also used to invoke the second intelligent agent to: When the similarity of the column type with the highest similarity among the multiple candidate column types is less than the similarity threshold, a predicted column type is generated based on the expression. as well as Replace the column type with the highest similarity with the predicted column type.

5. The system as described in claim 1, characterized in that, The solvent recommendation module is used to invoke the third agent to: Based on the similarity, the solvent pairs corresponding to the known compounds are sorted; and The candidate solvent pairs are determined based on the sorting results.

6. The system as described in claim 5, characterized in that, The solvent recommendation module is also used to invoke the third agent to: When the similarity of the candidate solvent pairs is less than a similarity threshold, a predicted solvent pair is generated based on the expression; and The predicted solvent pair is used instead of the candidate solvent pair.

7. The system as described in claim 1, characterized in that, The operating parameter prediction module is used for: Perform feature engineering on the expression, the candidate column type, and the candidate solvent pair to concatenate them into a condition vector; and Based on the conditional vector, the first machine learning algorithm is used to predict the mobile phase operating parameters of the candidate solvent pair.

8. The system as described in claim 7, characterized in that, The operating parameter prediction module is used for: Convert the expression into a first molecule descriptor; Convert the strings of the candidate solvent pairs into second molecule descriptors; Convert the candidate column type to an exclusive heat code; as well as The first molecular descriptor, the second molecular descriptor, and the one-hot code are concatenated to form the condition vector.

9. The system as described in claim 7, characterized in that, The mobile phase operating parameters include solvent ratio and solvent flow rate. The operating parameter prediction module is used to invoke the fifth agent to: predict the solvent ratio of the candidate solvent pair based on the condition vector for each of the plurality of candidate column types. The operating parameter prediction module is used to invoke a sixth agent to predict the solvent flow rate of the candidate solvent pair for each of the plurality of candidate column types, based on the condition vector and the solvent ratio.

10. The system as claimed in claim 1, characterized in that, The dwell time includes a first dwell time and a second dwell time. The residence time prediction module is configured to: use the second machine learning algorithm to predict the first residence time corresponding to the first configuration of the compound to be predicted and the second residence time corresponding to the second configuration of the compound to be predicted. The separation analysis module is used to call the fourth agent to perform the separation analysis based on the first dwell time and the second dwell time.

11. The system as claimed in claim 10, characterized in that, The residence time prediction module is used to invoke the seventh agent to: predict the first residence time for each of the plurality of candidate column types based on the expression, the candidate column type, the candidate solvent pair and the mobile phase operating parameters; The residence time prediction module is used to invoke an eighth agent to predict the second residence time for each of the plurality of candidate column types, based on the expression, the candidate column type, the candidate solvent pair, the mobile phase operating parameters, and the first residence time.

12. The system as claimed in claim 1, characterized in that, The resolution analysis module is used to invoke the fourth agent to calculate the peak resolution for each of the plurality of candidate column types based on the residence time.

13. The system as described in claim 12, characterized in that, The dwell time includes a first dwell time and a second dwell time. The separation analysis module is used to invoke the fourth agent to: calculate the peak resolution for each of the plurality of candidate column types based on the first residence time and first half-peak width corresponding to the first configuration of the compound to be predicted, and the second residence time and second half-peak width corresponding to the second configuration of the compound to be predicted.

14. The system as claimed in claim 1, characterized in that, The known compounds, along with their corresponding column types and solvent pairs, were obtained from a database derived from a large language model using the following methods: Obtain literature related to high performance liquid chromatography; Extract known compounds and their corresponding column types and solvent pairs from the literature; The known compounds, along with their corresponding column types and solvent pairs, are stored in the database.

15. A method for recommending chiral separation conditions, characterized in that, include: S100: Receives the expression for the compound to be predicted; S200: Based on the expression, calculate the similarity between the compound to be predicted and known compounds; S300: Based on the similarity, determine multiple candidate column types and candidate solvent pairs; S400: For each of the plurality of candidate column types, based on the expression, the candidate column type, and the candidate solvent pair, a first machine learning algorithm is used to predict the mobile phase operating parameters of the candidate solvent pair; S500: For each of the plurality of candidate column types, based on the expression, the candidate column type, the candidate solvent pair, and the mobile phase operating parameters, a second machine learning algorithm is used to predict the residence time associated with the chromatographic peaks of the chromatograms obtained by the compound to be predicted using the candidate column type, the candidate solvent pair, and the mobile phase operating parameters; S600: Perform a separation analysis associated with the compound to be predicted based on the residence time; as well as S700: The plurality of candidate column types are reordered based on the results of the resolution analysis, wherein recommended chiral separation conditions are determined based on the reordered candidate column types and associated candidate solvent pairs and mobile phase operating parameters.

16. The method as described in claim 15, characterized in that, S200 includes: S201: Convert the expression of the compound to be predicted into a first molecule descriptor; S202: Convert the string of the known compound into a second molecule descriptor; and S203: Calculate the similarity between the first molecular descriptor and the second molecular descriptor.

17. The method as described in claim 15, characterized in that, The S300 includes: S301: Based on the similarity, sort the column types corresponding to the known compounds in a first sort and sort the solvent pairs corresponding to the known compounds in a second sort. S302: Determine the types of the plurality of candidate chromatographic columns based on the results of the first sorting; and S303: Determine the candidate solvent pairs based on the results of the second sorting.

18. The method as described in claim 17, characterized in that, The S300 also includes: S304: When the similarity of the column type with the highest similarity among the multiple candidate column types is less than the similarity threshold, a predicted column type and a predicted solvent pair are generated based on the expression and using a third machine learning algorithm; S305: Replace the column type with the highest similarity with the predicted column type; and S306: Replace the candidate solvent pair with the predicted solvent pair.

19. The method as described in claim 15, characterized in that, The S400 includes: S401: Perform feature engineering on the expression, the candidate column type, and the candidate solvent pair to concatenate them into a condition vector; and S402: Based on the condition vector, use the first machine learning algorithm to predict the mobile phase operating parameters of the candidate solvent pair.

20. The method as described in claim 19, characterized in that, The mobile phase operating parameters include solvent ratio and solvent flow rate, and S402 includes: S4021: For each of the plurality of candidate column types, based on the condition vector, the first machine learning algorithm is used to predict the solvent ratio of the candidate solvent pair; S4022: For each of the plurality of candidate column types, based on the condition vector and the solvent ratio, the first machine learning algorithm is used to predict the solvent flow rate of the candidate solvent pair.

21. The method as described in claim 15, characterized in that, The dwell time includes a first dwell time and a second dwell time, and S500 includes: S501: For each of the plurality of candidate column types, based on the expression, the candidate column type, the candidate solvent pair, and the mobile phase operating parameters, the second machine learning algorithm is used to predict the first residence time corresponding to the first configuration of the compound to be predicted; S502: For each of the plurality of candidate column types, based on the expression, the candidate column type, the candidate solvent pair, the mobile phase operating parameters, and the first residence time, the second machine learning algorithm is used to predict the second residence time corresponding to the second configuration of the compound to be predicted. The separation analysis is performed based on the first residence time and the second residence time.

22. The method as described in claim 15, characterized in that, The dwell time includes a first dwell time and a second dwell time, and S600 includes: For each of the plurality of candidate column types, peak resolution is calculated based on the first residence time and first half-peak width corresponding to the first configuration of the compound to be predicted, and the second residence time and second half-peak width corresponding to the second configuration of the compound to be predicted.

23. The method as described in claim 15, characterized in that, The known compounds, along with their corresponding column types and solvent pairs, were obtained from a database derived from a large language model using the following methods: Obtain literature related to high performance liquid chromatography; Extract known compounds and their corresponding column types and solvent pairs from the literature; The known compounds, along with their corresponding column types and solvent pairs, are stored in the database.

24. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 15-23.

25. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 15-23.

Citation Information

Patent Citations

  • Method for separating and determining florfenicol enantiomer by ultra-performance convergence chromatography

    CN114460193A

  • Multimodal chromatographic separation media and process for using same

    US5316680A