Systems, methods, and computer program products for chromatographic enantiomeric separation of chiral molecules
The GCNN-based chromatography system addresses the challenge of enantiomer separation by predicting retention times and elution orders, enhancing the efficiency and accuracy of enantiomer identification and separation in chromatography.
Patent Information
- Application Number
- JP2025525080
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-21
- Filing Date
- 2023-11-20
- Publication Date
- 2025-12-16
AI Technical Summary
Existing chromatography methods struggle to accurately distinguish and separate enantiomers due to their identical physical and chemical properties, making it difficult to determine which enantiomer elutes first or second, and current detectors lack the capability to differentiate between them without prior knowledge of the positive or negative signal.
A chromatography system utilizing a graph convolutional neural network (GCNN) trained on datasets of chiral molecules to predict retention times and elution orders of enantiomers based on two-dimensional molecular representations, enabling the identification of peaks on a chromatogram and separation of desired enantiomers.
Enables rapid, cost-effective, and accurate determination of enantiomer elution orders, facilitating efficient selection and extraction of candidate molecules for downstream purification and analysis in drug development, food chemistry, and metabolite analysis.
Smart Images

Figure 2025540585000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This disclosure claims the benefit of the earlier filing date of U.S. Provisional Application No. 63 / 426,830, filed in the USPTO on November 21, 2022, the entire contents of which are incorporated herein by reference.
[0002] FIELD OF THE DISCLOSURE The present disclosure relates to scientific instrumentation, and more particularly to chromatography systems and methods, including devices and methods for identifying and extracting predetermined enantiomers of target molecules. [Background technology]
[0003] Isomers are molecules with the same chemical formula but with atoms arranged differently, exhibiting different properties. Stereoisomers (spatial isomers) are a subclass of isomers that rotate polarized light in different directions. Asymmetric stereoisomers, a subclass of stereoisomers, are optically active or "chiral" and are called "enantiomers." Chiral molecules are optically active and therefore rotate polarized light to the left or right depending on their configuration. Structurally, chiral molecules contain asymmetric centers (chiral atoms or chiral centers) and can therefore exist in two forms (enantiomers) that are mirror images of each other; therefore, regardless of their relative orientation, the molecular structures of the pairs do not overlap. Chiral molecules rotate the plane of polarized light; the degree of this rotation is called the specific optical rotation or optical rotation. Paired enantiomers share the property of rotating the plane of polarized light in different directions; one enantiomer rotates the plane of polarized light to the left, while the other rotates it to the right. In the nomenclature, (+) refers to clockwise rotation and (-) refers to counterclockwise rotation.
[0004] Enantiomers may be distinguished or separated from one another using chromatography, the instrument for which is a "chiral column" (e.g., a glass column packed with silica in an organic solvent). Exemplary commercially available chiral columns used to perform separations are CROWNPAK CR(+) and CROWNPAK CR(-), all manufactured by DAICEL. These chiral columns are available from the Sigma Corporation. These columns contain a chiral crown ether as a chiral selector coated on a 5 μm silica support. To operate these columns under standard conditions, an acidic mobile phase, such as perchloric acid pH 1-2, is used. These columns are reference columns for achieving amino acid separations and have the advantage that the elution order of the enantiomers can be reversed as needed (a CR(-) column gives a reversed elution order compared to a CR(+) column). Non-limiting examples of methods for performing separations are described in further detail, for example, in U.S. Pat. Nos. 4,942,149, 9,145,430, 9,233,355, and 9,409,145, the entire contents of each of which are incorporated herein by reference in their entirety.
[0005] In practice, most biologically active substances that control physiological functions in living organisms are chiral molecules. Various techniques have been developed to determine their absolute configuration, including experimental methods based on diffraction and spectroscopy. X-ray crystallography is one method, but it requires a suitable single crystal, making it less practical. Circular dichroism exciton chirality methods, which combine vibrational circular dichroism analysis with quantum mechanical calculations, are also highly relevant, but are less ideal for interpreting the meaning of the results.
[0006] Chromatography using chiral columns is a common method for determining the optical purity of small organic molecules, especially since nearly 90% of chiral molecules can be resolved by chromatography. However, as recognized by the present inventors, even if chromatographic peak separation of a pair of enantiomers is achieved, it is difficult to distinguish which enantiomer eluted first or second. Because these enantiomers have identical physical and chemical properties, such as melting point, boiling point, density, and thermal conductivity, the only physical difference between the enantiomers is optical activity. Therefore, it is difficult to distinguish between enantiomers using detectors commonly used in chromatography, such as mass spectrometers, ultraviolet-visible detectors, fluorescence detectors, and refractive index detectors.
[0007] On the other hand, circular dichroism detectors can detect the differential absorption of left- and right-handed circularly polarized light, and optical rotation detectors can measure the angle of rotation of plane-polarized light. While these detectors allow for the detection of optical activity, as recognized by the inventors, it is difficult to distinguish the response peaks of a chiral column without a priori knowledge of the positive or negative signal of the desired enantiomer.
[0008] The past decade has seen a rapid acceleration in the development and application of "deep learning" techniques, including applications in chemistry where they can be used to predict physical properties, generate novel molecules, and design synthetic routes.
[0009] As the application of deep learning to molecules progresses, graph neural networks have emerged as an available computer-based platform for molecular analysis. Molecular structures are sometimes input to graph neural networks as a "graph," defined by a set of nodes connected by vertices. For molecules, each atom is considered a node, and the bonds connecting those atoms are vertices. Each atom has several features associated with it (e.g., element type, hybridization state, chirality, etc.). These atomic-level features can be analyzed by neural networks to make predictions about the physical properties of molecules. In addition to graph neural networks, a growing library of graph architectures exists, including message-passing neural networks (MPNNs) and graph attention networks (GANs), among others. Classification and regression tree (CART)-based models have also shown some success in predicting the correct elution order of enantiomers. However, CART-based systems rely heavily on 3D coordinate inputs, requiring knowledge of partial atomic charges and effective polarizabilities using empirical methods.
[0010] For example, the "SMRT dataset" contains liquid chromatography retention times of small molecules and contains over 80,000 entries. As recognized by the inventors, the SMRT dataset does not contain chirality information, but the size of this dataset nonetheless makes it attractive for use with modern deep learning techniques. Summary of the Invention
[0011] According to one non-limiting aspect of the present disclosure, a novel chromatography system includes a non-transitory computer-readable storage device storing computer-executable code, an interface that characterizes two-dimensional (2D) information of an enantiomer of a target molecule and receives input data including information of a particular chromatography column, and processing circuitry. The processing circuitry, upon execution of the computer-executable code, applying input data to a graph convolutional neural network (GCNN), a type of graph neural network, trained on a dataset of a plurality of chiral molecules associated with a particular chromatography column, the dataset labeled with retention times of both enantiomers of each of the plurality of chiral molecules in the dataset, the GCNN being trained to set weights and adjust the weights based on backpropagation loss from the predicted retention times compared to ground truth predicted retention times; collecting the output of the GCNN as predicted retention times of the enantiomers of the target molecule; distinguishing the elution order of the enantiomers of the target molecule based on the relative magnitudes of their predicted retention times; and identifying peaks on a chromatogram for each enantiomer of the target molecule based on the elution order of the enantiomers of the target molecule, the chromatogram being obtained from testing the target molecule on a chromatograph having a particular chromatography column.
[0012] A more complete understanding of the present disclosure and many of the attendant advantages thereof will be readily obtained as the same becomes better understood by reference to the following detailed description when considered in connection with the accompanying drawings. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 is a system-level diagram of a system that employs computational and instrumentation-based characterization of chiral molecules, identification of predicted elution orders of enantiomers of chiral molecules using a trained AI engine, and extraction of desired enantiomers from a supply of target molecules based on the identified elution orders.
[0014] [Figure 2] FIG. 1 is an exemplary chromatogram output from a chiral column showing the reaction peaks for each of a pair of enantiomers.
[0015] [Figure 3]FIG. 1 is a diagram of a computer-based architecture of an artificial intelligence (AI) engine including a trained model used to predict elution order, according to one embodiment.
[0016] [Figure 4A] FIG. 4 is a block diagram of the data extraction network components of the example AI engine shown in FIG.
[0017] [Figure 4B] FIG. 4 is a block diagram of a data analysis network for the example AI engine embodiment shown in FIG.
[0018] [Figure 5A] FIG. 1 is a diagram of an exemplary data structure used for computer analysis and digital communication, the data structure containing 2D molecular description data in the Simplified Molecular Input Line Entry System (SMILES) format.
[0019] [Figure 5B] FIG. 1 is a diagram of an exemplary data structure for information containing output from an AI engine and associations between elution order and retention time of specific enantiomers.
[0020] [Figure 6] FIG. 1 is a diagram of an exemplary transformation process for converting SMILES format 2D information of a target molecule into an adjacency matrix suitable for processing by GCNN.
[0021] [Figure 7] 1 is a flowchart of a method for identifying peaks on a chromatogram obtained from a test of a target molecule on a chromatograph generated from data provided by a particular chromatography column, according to one embodiment of the present disclosure.
[0022] [Figure 8]1 is a flowchart of a method for identifying peaks on a chromatogram after a target molecule has passed through a particular chromatography column and collecting a predetermined enantiomer of the target molecule based on the identified peak response, according to one embodiment of the present disclosure.
[0023] [Figure 9A] 1 is a graph with exemplary data of predicted versus experimental dissolution times for the CROWNPAK-CR-I(+) dataset.
[0024] [Figure 9B] 1 is a graph of exemplary data for predicted versus experimental dissolution times for the CROWNPAK-CR-I(+) dataset.
[0025] [Figure 10] FIG. 1 is a diagram of exemplary amino compounds as target molecules.
[0026] [Figure 11] FIG. 1 is a diagram of an exemplary embodiment of a computer that may be included as part of a distributed computer architecture used to execute the control / analysis processes disclosed herein, such as implementing embodiments of an AI engine disclosed herein. DETAILED DESCRIPTION OF THE INVENTION
[0027] As used herein, elements or steps listed in the singular and preceded by the word "a" or "an" should be understood as not excluding a plurality of elements or steps, unless such exclusion is expressly recited. Furthermore, references to "one embodiment" of the present disclosure are not intended to be interpreted as excluding the existence of additional embodiments that also incorporate the recited features.
[0028] As mentioned above, methods such as the Structure-Dependent Chirality Code (CDCC) use 3D coordinate information to identify elution order. As recognized by the present inventors, several previous approaches have attempted to use 2D information to assess chirality properties, but there has been a lack of understanding of what molecular information is actually required to accurately predict enantiomer separation. The commercial benefit of such a reduction in information is a smaller database size for machine learning modeling.
[0029] In light of these and other limitations of conventional approaches, this document addresses systems and methods that use a 2D representation of the enantiomers of a target molecule, including information about a specific chromatography column. This 2D representation serves as a fingerprint and is presented as a graph representation using SMILES strings, which can then be converted into an adjacency matrix and a list of node attributes. One by-product of this approach is the creation of unique datasets containing retention time information for chiral separations in high-performance liquid chromatography (HPLC). A significant commercial benefit of this approach is the creation of computer-based scientific instruments that identify and correlate the elution order of enantiomers, thereby enabling the efficient selection of candidate molecules for downstream purification and extraction. To this end, in this embodiment, CROWNPAK CR(+) and CR-I(+) columns are used (in this non-limiting example), with the selector being a chiral crown ether ((S)-3,3'-diphenyl-1,1'-binaphthyl)-20-crown-6) moiety. The CROWNPAK CR(+) column was originally designed for separating chiral molecules containing primary amino groups, such as amino acids other than proline, amino alcohols, and amines. While the CROWNPAK CR mobile phase has limitations, the later-released CROWNPAK CR-I column, which is immobilized, does not have such limitations. These columns are also effective for achiral separations, although many of the molecules in the datasets are enantiomeric pairs. As described herein, these datasets are used to train the AI engine GCNN to predict the retention times of chiral molecules, and ultimately use these results to predict the elution order of the enantiomeric pairs.
[0030] Determining the elution order of chiral molecules has important practical applications, particularly in drug development, food chemistry, metabolite analysis, and the like. The techniques described herein can also be used to predict the elution order of unmeasured molecules by learning the elution orders reported by a trained AI engine using GCNN. The disclosed methods are also applicable to other chiral columns. Other chiral columns include "optically active polymer-type chiral columns" whose chiral selectors are polysaccharides or derivatives thereof, optically active poly(meth)acrylamides, optically active poly(amino acids), and / or optically active polyamides; and "optically active polymer-type chiral columns" whose chiral selectors are binaphthalenes. This may include "optically active low molecular weight compound-type chiral columns" which are compounds having a butyl structure or a crown ether structure, proteins or glycoproteins having sugar chain-type chiral columns, nucleic acid-type chiral columns, DNA or RNA-type chiral columns, zwitterion columns, anion exchange columns, and ligand exchange columns.
[0031] FIG. 1 is a system-level diagram of a system employing computer-based characterization of chiral molecules, identification of the predicted elution order of enantiomers of the chiral molecules using a trained AI engine, and optional separation and extraction of the desired enantiomer from a supply of target molecules based on the predicted elution order. Computer 1 (which may be configured as computer 805, described in more detail with respect to FIG. 11 ) receives input data from molecular database 3 via an interface. The input data from the molecular database may optionally characterize molecules with 3D coordinate information, such as CDCC, shown as an element in FIG. 1 . However, the molecular information may optionally be preprocessed and input as two-dimensional (2D) information for the target molecule, including information for a specific chromatographic column. The initial format for the 2D information is the SMILES string format (element 7 in FIG. 1 ), whether it is input in that format or converted to the SMILES string format within computer 1. SMILES is a chemical notation that can be directly analyzed by computer 1. In this embodiment, before performing downstream data processing, the SMILES format data is further converted into an adjacency matrix with node information (described in more detail later) before being placed into data packets 9 and transmitted to another computer over network 13. The downstream processing (e.g., implementation of GCNN for an AI engine) may be performed locally on computer 1.
[0032] In the example of FIG. 1, the molecular information contained in the adjacency matrix is transmitted as a digitized data file and sent to GCNN computer 15 (one or more computers and / or cloud computing resources) via network 13. As described in more detail with respect to FIGS. 2-3B, GCNN computer 15 develops and hosts a trained AI engine to predict the retention times of enantiomers of a target molecule for a particular chromatographic column. Once the predicted retention times are obtained, they are compared with response peaks on the chromatogram (developed within chromatographic processing system 11) of the target molecule to identify the elution order of the particular enantiomer. The comparison and matching may be performed in either (i) computer 1, (ii) chromatographic processing system 11, (iii) GCNN computer 15, or (iv) another device. Once the detection and identification are obtained, they are optionally transmitted to purification system 17, where a feed of the target molecule is passed through a chromatographic column (or other separation system) and the desired enantiomer can be collected at its corresponding retention time in the chromatographic column. In this way, the desired enantiomer of the target molecule can be separated from the undesired enantiomer.
[0033] Before moving on to a description of the AI engine for predicting peak retention times from 2D enantiomer information, let's briefly review the motivation behind the AI engine. To experimentally determine retention times on a chromatography column (e.g., CROWNPAK CR-I(+)), a target molecule with its enantiomeric pair is applied to the chromatography column. The chromatography column selects each enantiomer and provides evidence of the enantiomers in the form of peak responses on the chromatogram.
[0034] FIG. 2 is an exemplary chromatogram showing the voltage output of a chromatography column in millivolts ("Response" on the Y-axis) over time (minutes). When graphed, distinct peak response voltages appear at specific times labeled "Retention Times." The retention time for one of the pair of enantiomers (e.g., "Enantiomer A") appears at 5.94 seconds, while the retention time for the other (e.g., "Enantiomer B") appears at 8.25 seconds. Thus, the elution order for a particular chromatography column is determined by the order of the enantiomers. Enantiomer A precedes enantiomer B. Without further ado, the mystery of which enantiomer is which remains. While experimental systems allow for the evaluation of each enantiomer after extraction, as recognized by the inventors, this is an expensive and time-consuming process when evaluating many different target molecules. In contrast, instrumentation using an AI engine in conjunction with a trained GCNN allows for the detection and identification of the elution order of a vast number of target molecules in a rapid and cost-effective manner.
[0035] Figures 3, 4A, and 4B show a computer-based architecture for developing and applying an artificial intelligence (AI) processing engine that uses a trained GCNN to provide predicted retention times for enantiomers of a target molecule based on input of the target molecule as a machine-readable string. The computer-based system then matches peaks on a chromatogram with each enantiomer of the target molecule based on the elution order of the enantiomers of the target molecule. This matching allows identification of which enantiomer corresponds to which peak.
[0036] We have recognized that using graph-based neural networks can be superior to traditional convolutional neural networks for analyzing intramolecular features. While both have a similar layered structure and use backpropagation, as described, GCNN performs molecular characterization in a non-Euclidean space, while CNN operates in Euclidean space (characterizing spatial features of objects, such as the physical distance between pixels in an image). Furthermore, GCNN "characterizes" intramolecular connections in an arbitrary space, rather than a 3D Euclidean space. The connections within a given molecule are encoded in the adjacency matrix of a graph representing that molecule. For comparison, consider a typical application of CNNs to perform image analysis. In this case, the CNN deploys a set of small fields (kernels) that are "slid" piecewise across an image and then combined (a convolution process) to create new parts / values. However, in non-Euclidean space, the physical space between molecular features, characterized as a 2D abstraction, is less relevant than the mere existence of connections in the graph itself, so "sliding" the kernel across the image is less effective.
[0037] In contrast to CNNs, GCNNs allow arbitrary connections between nodes, and every node has a feature vector. The feature vector (molecular footprint) is processed in a series of layers, which contain weights and (optionally) biases for each element of the feature vector. The output of each layer is aggregated and propagated to the next layer. Layers within GCNNs obtain feature vectors from the neighborhood of a particular node and learn a set of trainable weights based on the feature vectors of that particular node and its neighbors. The final layer outputs the results of "N" adjacent layers of aggregation. To train the AI engine, a "loss" function must exist for the trainable weights used in the layers of the internal neural network. Training can be supervised or unsupervised. In supervised learning, GCNNs are trained with a set of example molecules ("training molecules") containing values for targeted predictive properties. The loss function is used to compare prediction results from GCNN with ground truth values for the example molecules and update the trainable weights using an algorithm to reduce the loss of future predictions from the model. The AI model predicts properties of newly applied molecules based on what it learns from the example molecules.
[0038] Turning to a more specific example, the GCNN of Figure 3 has an input interface 21 that receives a SMILES 2D representation of a target molecule (the input string), which is then optionally converted into an adjacency matrix as described below. The input string is converted into a first feature vector based on a predetermined set of features associated with the atoms contained in the molecule. For a given node, this feature vector is processed by a first feature layer. This process continues as shown in the figure (the top node with multiple connections for the top molecule, the bottom node with multiple connections for the middle molecule, and the bottom node for the third molecule). The process is repeated for each node, as suggested by the shifted positions of the lighter shaded nodes. Application of the feature layer weights (whose initial values are randomly generated) results in a matrix (the "feature map"). The feature map is passed through a rectified linear unit (ReLU). ReLU is an activation function that sets any negative values in the output matrix to 0 and preserves any positive values, which introduces nonlinearity into the neural network. ReLU is just one non-limiting example of an activation function; other examples may be used as well. After the ReLU activation function is applied to the features in the first feature map, the resulting vector is applied to a second feature extraction layer 23 to create a second feature map, and a subsequent further ReLU activation function is applied to the features in the second feature map. Additional featurization layers may also be used before the resulting feature map is applied to one or more pooling layers 24, which gradually reduces the size of the feature map ("flattening" the map) and reduces downstream computations. Two or more pooling layers 24 may be concatenated. One or more fully connected layers 25 (also known as dense layers) then apply final weights to the connected components of each node's feature vector from the final pooling layer to provide a final probability for each candidate label. Since the predictions are molecule-level properties, all node feature vectors are used. For node- or edge-level predictions, only the feature vector associated with that node or edge may be used. The candidate label with the highest probability retention time is then selected as the predicted retention time for the target molecule.
[0039] The AI engine includes at least the GCNN of FIG. 4A , but may also include the structure of FIG. 4B as part of the GCNN. In this embodiment, the AI engine includes both the structure of FIG. 4A as exemplary data extraction network 200 and the structure of FIG. 4B as data analysis network 300. In this configuration, a hosting computer (e.g., computer 1 of FIG. 1 implemented by computer 805 of FIG. 11 , or GCNN computer 15 of FIG. 1 ) implements both data extraction network 200 and data analysis network 300. The data extraction network generates source vectors that are applied as input to data analysis network 300. Additionally, the hosting computer instructs data extraction network 200 to generate source vectors that include apparent (estimated) retention times and, optionally, second features.
[0040] Referring to FIG. 4A, computer 1 (FIG. 1) trains a model for use within an AI engine by first instructing first feature extraction layer 210 to apply at least one first convolution operation to a 2D representation of training target molecules, which are then applied to data extraction network 200 as part of an adjacency matrix, thereby generating at least one target feature map within first extraction layer 210. Computer 1 can then instruct ROI pooling layer 220 to generate one or more ROI pooled feature maps by pooling regions on the target feature map that correspond to ROIs within the 2D representation of the target molecules. These ROIs may not be readily apparent from human observation, but once sufficiently fine features are extracted from the graph and convolved across several layers, a pattern for an indicator or predicted retention time emerges. Computer 1 may then instruct first output layer 230 to generate at least one estimated retention time (as a first estimated feature). Optionally, additional features may also be estimated. That is, the first output layer 230 may perform classification and regression on the subject 2D representation of the target molecule by applying at least one first fully connected (FC) operation to the ROI pooled feature map.
[0041] After such detection of retention times (and optional second features) is complete, the computer can instruct the data vectorization layer 240 to identify the difference between the predicted retention times and the experimentally determined ground truth (GT) retention times and back-propagate the retention time loss to previous layers in the neural network. This back-propagation completes the feature extraction process. The trainable weights in the extraction and pooling layers are adjusted to improve the data vectorization layer 240, which generates source vectors, which are then applied to the data analysis network 300.
[0042] The computer 1 then instructs the data analysis network 300 to calculate a predicted retention time. Herein, the second feature extraction layer 310 of the data analysis network 300 applies a second convolution operation to the source vector to generate at least one source feature map, and the second output layer 320 of the data analysis network 300 performs regression by applying at least one fully-connected operation to the source feature map to thereby calculate an estimated retention time.
[0043] Regarding the training algorithm, in FIG. 4A, the data extraction network 200 may be trained using multiple training examples containing known (ground truth, GT) retention times of 2D molecular structures and enantiomers. These GT retention times are used to determine the difference from predicted retention times to backpropagate losses to upstream neural network layers. Regarding the backpropagation loss algorithm, any of various loss generation algorithms, such as the smoothed L1 loss algorithm and the cross-entropy loss algorithm, may be used. Herein, the data vectorization layer 240 may be implemented using a rule-based algorithm rather than a neural network algorithm. In this case, the data vectorization layer 240 may not need to be trained; it only needs to perform properly using its manually entered settings. As an example, the first feature extraction layer 210, the ROI pooling layer 220, and the first output layer 230 may be obtained by using transfer learning, which repurposes an existing pre-trained neural network model for a new task, as commonly used in object detection networks such as VGG or ResNet.
[0044] Once trained, the AI engine can use the trained model, implemented as the structure of Figure 4A, alone or in conjunction with the second neural network of Figure 4B to predict retention times. Additionally, the AI engine is trained using a training sequence of 2D molecular representations with known retention times and elution orders to adjust the weights in the trained model. The AI engine with the trained model may then be used to process 2D molecular representations of target molecules whose enantiomeric retention times are not known a priori, and the AI engine provides predicted retention times for the target molecules.
[0045] Figure 5A is an exemplary data structure for a chiral molecule including 2D molecular description data in SMILES format. Specifically, an identification (ID) label is included in a first field 511 and is stored as part of the same data structure or in association with the 2D SMILES format molecular description for that molecule in data field 512. The data structure of Figure 5A may be stored in molecule database 3 (Figure 1), as well as other computers in the system of Figure 1.
[0046] Figure 5B is an example of a data structure for an AI output vector containing predicted elution order numbers and retention times. In addition to having a data field for chiral molecule ID 513, it includes a data field for enantiomer ID 514 and associated data fields 515 and 516 for elution order number and retention time, respectively. The data structure of Figure 5B can be used to share the predicted elution order or retention time obtained from GCNN 15 (Figure 1) or to provide this information to a purification system 17. Additionally, it may be used by computer 1 to match chromatographic peaks to specific enantiomers.
[0047] For the training dataset, publicly available data processed using the CROWNPAK CR(+) or CR-I(+) columns may be used.
[0048] In terms of off-the-shelf software architectures, Pytorch Geometric may be used to implement GCNN architectures, and Pytorch may be used to implement MLP-based models and RDKit for various cheminformatics tasks, including molecular characterization. Further explanation is provided with respect to Figure 6.
[0049] For dataset partitioning, the Butina partitioning algorithm may be used to split a dataset into training and validation datasets in any ratio (e.g., 80 / 20 ratio). The Butina algorithm uses Morgan footprints to group similar molecules together according to the molecular Tanimoto similarity index, attempting to fill the training and validation datasets with dissimilar molecules. In this way, the validation dataset provides a better measure of the model's ability to generalize to new molecules than partitioning algorithms that do not consider molecular similarity (such as random partitioning). Because the Butina partitioning algorithm groups similar molecules, enantiomeric pairs naturally fall into the same dataset partition. For molecular characterization, the RDKit package may be used.
[0050] For training the AI Engine, a backpropagation algorithm for gradient calculation and parameter update during training may be used. For example, an AdamW optimizer with a OneCycle learning rate scheduler may be used with an L1 loss function (identical to the mean absolute error loss). All models are then trained for several epochs (e.g., 250 epochs). Hyperparameter optimization may be performed using the OPTUNA package with over 100 iterations for MLP models and 100 or more iterations for GCNN models. It is also possible to monitor the change in the model's loss during training and terminate training when the loss no longer changes after a certain number of epochs, such as five epochs.
[0051] Figure 9A plots an example of predicted dissolution times versus experimental dissolution times for CR-I(+), and Figure 9B plots the CR(+) validation data set. In both plots, a 45-degree (y = x) line is also plotted (as labeled), demonstrating perfect predictive performance by the network. Also plotted as another line in each of Figures 9A and 9B is a linear regression of the data. The regression fixes the intercept through the origin, which is a natural physical choice.
[0052] The above discussion has focused on accurate prediction of retention times. A related task is the prediction of the elution order of enantiomeric molecules. The predicted elution order can be evaluated from the predicted retention times. This evaluation does not require that the elution times themselves be accurate, only that the relative magnitudes of the predicted retention times be accurate.
[0053] Starting with the predicted retention times described above, enantiomer pairs are matched and the relative magnitude of the predicted retention times compared to the relative magnitude of the experimental retention times (training data). If the relative magnitudes of the experimental and predicted elution times match (i.e., both the experimental and predicted retention times for enantiomer A are greater than the respective retention times for enantiomer B), the elution order prediction is accurate. If a discrepancy occurs, the prediction is inaccurate. The final accuracy of the prediction can be calculated as the number of correct predictions divided by the total number of predictions.
[0054] Figure 6 is a flowchart of a process for converting SMILES format input data characterizing 2D information of the enantiomers of a target molecule into an adjacency matrix with an associated list of the molecule's node attributes. The adjacency matrix can be read by GCNN. The process is Starting at step S601, Computer 1 (FIG. 1) receives SMILES-formatted input data from the molecular database (FIG. 1) that characterizes two-dimensional (2D) information of the target molecule's enantiomers and includes information on a specific chromatography column. The process then proceeds to step S602, where the computer imports and loads the Python libraries SPEKTRAL, NETWORKX, and NUMPY. These are non-limiting examples; other off-the-shelf libraries and / or custom algorithms may be used as well. The process then proceeds to step S603, where Computer 1 loads the SMILES-formatted data into a SPEKTRAL object. The process then proceeds to step S604, where Computer 1 converts the SDF object into a Network X object. Finally, an adjacency matrix and a list of node attributes are obtained by a final transformation using the following operations: sdf_adj,sdf_node,_=nx_to_numpy(sdf_nx, nf_keys=['atomic_num'], ef_keys=['type'] )
[0055] The generation of the adjacency matrix in Figure 6 is one example of how the conversion may be performed. Other approaches, such as using the Python-based library Pytorch Geometric, may be used as well. This approach can be performed in three steps: (1) defining a function that maps RDKit atom objects to atomic feature vectors, (2) defining a function that maps RDKit joint objects to joint feature vectors, and (3) defining a function that takes as its input a list of SMILES strings and associated labels, and then uses the functions from (1) and (2) to create as its output a list of labeled Pytorch Geometric graph objects.
[0056] 7 is a flowchart of a method for identifying peaks on a chromatogram obtained from a test of a target molecule on a chromatograph having a specific chromatographic column according to one embodiment of the present disclosure. The method characterizes two-dimensional (2D) information of the enantiomers of the target molecule and can include receiving input data including information of the specific chromatographic column (step S701), applying the input data to a graph convolutional neural network (GCNN) trained with a dataset of multiple chiral molecules associated with the specific chromatographic column, the dataset being labeled with the retention times of both enantiomers of each of the multiple chiral molecules in the dataset (step S702), collecting the output of the GCNN as predicted retention times of the enantiomers of the target molecule (step S703), identifying an elution order of the enantiomers of the target molecule based on the relative magnitudes of the predicted retention times (step S704), and identifying peaks on the chromatogram for each enantiomer of the target molecule based on the elution order of the enantiomers of the target molecule (step S705).
[0057] 8 is a flowchart of a method for identifying peaks on a chromatogram after a target molecule has passed through a specific chromatography column and collecting a predetermined enantiomer of the target molecule based on the identified peak, according to one embodiment of the present disclosure. The method characterizes two-dimensional (2D) information of the enantiomers of the target molecule and includes receiving input data including information of the specific chromatography column (step S801), applying the input data to a graph convolutional neural network (GCNN) trained with a dataset of multiple chiral molecules associated with the specific chromatography column, the dataset being labeled with the retention times of both enantiomers of each of the multiple chiral molecules in the dataset (step S802), collecting the output of the GCNN as predicted retention times of the enantiomers of the target molecule (step S803), identifying the elution order of the enantiomers of the target molecule based on the relative magnitudes of the predicted retention times (step S804), and collecting the elution order of the enantiomers of the target molecule based on the elution order of the enantiomers of the target molecule. The method may include identifying peaks on a chromatogram for each enantiomer (step S805), and collecting a predetermined enantiomer of the target molecule using a chromatograph having a specific chromatographic column based on the identified peaks on the chromatogram after the target molecule has passed through the specific chromatographic column (step S806).
[0058] FIG. 9A is a graph of exemplary predicted versus experimental elution times for the CROWNPAK-CR-I(+) dataset.
[0059] FIG. 9B is a graph of exemplary predicted versus experimental elution times for the CROWNPAK-CR(+) dataset.
[0060] Figure 10 includes exemplary amino compounds that can be used as target molecules. As shown in Figure 10, each of the amino compounds, such as amino acids, amino alcohols, or amines, has at least one chiral atom in its structure, indicated by the symbol *, and therefore has at least two enantiomers.
[0061] The disclosed methods of the present application are also applicable to other target molecules, including compounds with amide, aromatic, carbonyl, nitro, sulfonyl, cyano, hydroxyl, and / or amino groups.
[0062] Aspects of the present disclosure may be embodied as systems, methods, and / or computer program products, which may include a computer-readable storage medium having computer-readable program instructions recorded thereon that can cause one or more processors to perform aspects of the embodiments.
[0063] A computer-readable storage medium may be a tangible device capable of storing instructions for use by an instruction execution device (processor). The computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of these devices. A non-exhaustive list of more specific examples of computer-readable storage media includes each (and suitable combinations) of the following: a floppy disk, a hard disk, a solid-state drive (SSD), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash), a static random access memory (SRAM), a compact disk (CD or CD-ROM), a digital versatile disk (DVD), and a memory card or stick. The computer-readable storage medium used in this disclosure should not be construed as a transitory signal itself, such as an electric wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse passing through a fiber optic cable), or an electrical signal transmitted over a wire.
[0064] The computer-readable program instructions described in this disclosure can be downloaded from a computer-readable storage medium to an appropriate computing or processing device, or to an external computer or external storage device, via a global network (i.e., the Internet), a local area network, a wide area network, and / or a wireless network. The network may include copper transmission lines, fiber optics, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface within each computing or processing device receives the computer-readable program instructions from the network and transmits them to the computer for storage in a computer-readable storage medium within the computing or processing device. Readable program instructions can be transferred.
[0065] Computer-readable program instructions for carrying out operations of the present disclosure may include machine language instructions and / or microcode that may be compiled or interpreted from source code written in any combination of one or more programming languages, including assembly language, Basic, Fortran, Java, Python, R, C, C++, C#, or similar programming languages. The computer-readable program instructions may execute entirely on a user's personal computer, notebook computer, tablet, or smartphone, or entirely on a remote computer or computer server, or any combination of these computing devices. The remote computer or computer server may be connected to one or more of the user's devices via a computer network, including a local area network or a wide area network, or a global network (i.e., the Internet). In some embodiments, electronic circuitry, including, for example, programmable logic circuits, field programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), may execute the computer-readable program instructions by configuring or customizing the electronic circuitry using information from the computer-readable program instructions to carry out aspects of the present disclosure.
[0066] Aspects of the present disclosure are described herein with reference to flow diagrams and block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. Those skilled in the art will understand that each block of the flow diagrams and block diagrams, and combinations of blocks in the flow diagrams and block diagrams, can be implemented by computer-readable program instructions.
[0067] Computer-readable program instructions capable of implementing the systems and methods described in this disclosure can be provided to one or more processors (and / or one or more cores within a processor) of a general-purpose computer, special-purpose computer, or other programmable device to produce a machine such that the instructions, executed via the processor(s) of the computer or other programmable device, create a system for implementing the functions specified in the flow diagrams and block diagrams of this disclosure. These computer-readable program instructions can also be stored on a computer-readable storage medium that can instruct the computer, programmable device, and / or other device to function in a particular way, such that the computer-readable storage medium storing the instructions is an article of manufacture containing instructions that implement aspects of the functions specified in the flow diagrams and block diagrams of this disclosure.
[0068] The computer-readable program instructions may also be loaded onto a computer, other programmable apparatus, or other device to cause the computer, other programmable apparatus, or other device to perform a series of operational steps to produce a computer-implemented process, such that the instructions, executing on the computer, other programmable apparatus, or other device, implement the functions specified in the flow diagrams and block diagrams of this disclosure.
[0069] 11 is an architectural diagram illustrating a networked system 800 of one or more networked computers and servers. In one embodiment, the hardware and software environment illustrated in FIG. 11 can provide an exemplary platform for implementing software and / or methods according to the present disclosure.
[0070] Referring to FIG. 11, a networked system 800 may include, but is not limited to, a computer 805, a network 810, a remote computer 815, a web server 820, a cloud storage server 825, and a computer server 830. In some embodiments, multiple instances of one or more of the functional blocks shown in FIG. 11 may be used.
[0071] Further details of computer 805 are shown in Figure 11. The functional blocks shown in computer 805 are provided only to establish example functionality and are not intended to be exhaustive. Also, although details are not provided for remote computer 815, web server 820, cloud storage server 825, and computer server 830, these other computers and devices may include functionality similar to that shown for computer 805.
[0072] The computer 805 may be a personal computer (PC), a desktop computer, a laptop computer, a tablet computer, a netbook computer, a personal digital assistant (PDA), a smartphone, or any other programmable electronic device capable of communicating with other devices on the network 810.
[0073] Computer 805 may include a processor 835, a bus 837, memory 840, non-volatile storage 845, a network interface 850, a peripherals interface 855, and a display interface 865. Each of these functions may, in some embodiments, be implemented as an individual electronic subsystem (an integrated circuit chip or combination of chips and associated devices), or in other embodiments, some combination of functions may be implemented on a single chip (sometimes called a system-on-chip, or SoC).
[0074] Processor 835 may be one or more single-chip or multi-chip microprocessors, such as those designed and / or manufactured by Intel Corporation, Advanced Micro Devices, Inc. (AMD), Arm Holdings (Arm), Apple Computer, etc. Examples of microprocessors include Celeron, Pentium, Core i3, Core i5, and Core i7 manufactured by Intel Corporation, Opteron, Phenom, Athlon, Turion, and Ryzen manufactured by AMD, and Cortex-A, Cortex-R, and Cortex-M manufactured by Arm.
[0075] Bus 837 may be a proprietary or industry-standard high-speed parallel or serial peripheral interconnect bus such as ISA, PCI, PCI Express (PCI-e), AGP, or the like.
[0076] The memory 840 and the non-volatile storage 845 may be computer-readable storage media. The memory 840 may include any suitable volatile storage device, such as dynamic random access memory (DRAM) and static random access memory (SRAM). The non-volatile storage 845 may include one or more of the following: a floppy disk, a hard disk, a solid-state drive (SSD), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash), a compact disk (CD or CD-ROM), a digital versatile disk (DVD), and a memory card or memory stick.
[0077] Program 848 may be a collection of machine-readable instructions and / or data that is stored in non-volatile storage 845 and used to create, manage, and control certain software functions described in detail elsewhere in this disclosure and illustrated in the figures. In some embodiments, memory 840 may be significantly faster than non-volatile storage 845. In such embodiments, program 848 may be written to non-volatile storage 845 prior to execution by processor 835. The data may be transferred from the device 845 to the memory 840 .
[0078] Computer 805 may be able to communicate and interact with other computers over network 810 via network interface 850. Network 810 may be, for example, a local area network (LAN), a wide area network (WAN) such as the Internet, or a combination of the two, and may include wired, wireless, or fiber optic connections. In general, network 810 may be any combination of connections and protocols that support communication between two or more computers and related devices.
[0079] The peripheral interface 855 may enable the input and output of data between the computer 805 and other devices that may be locally connected. For example, the peripheral interface 855 may provide a connection to an external device 860. The external device 860 may include devices such as a keyboard, a mouse, a keypad, a touchscreen, and / or other suitable input devices. The external device 860 may also include portable computer-readable storage media, such as thumb drives, portable optical or magnetic disks, and memory cards. Software and data, e.g., program 848, used to practice embodiments of the present disclosure may be stored on such portable computer-readable storage media. In such embodiments, software may be loaded directly into the non-volatile storage device 845 or, alternatively, into the memory 840 via the peripheral interface 855. The peripheral interface 855 may connect to the external device 860 using an industry-standard connection, such as RS-232 or Universal Serial Bus (USB).
[0080] Display interface 865 can connect computer 805 to a display 870, which in some embodiments may be used to present a command line or graphical user interface to a user of computer 805. Display interface 865 can connect to display 870 using one or more proprietary or industry standard connections, such as VGA, DVI, DisplayPort, and HDMI.
[0081] As mentioned above, network interface 850 provides for communication with other computing and storage systems or devices external to computer 805. Software programs and data described herein may be downloaded to non-volatile storage 845 via network interface 850 and network 810 from, for example, remote computer 815, web server 820, cloud storage server 825, and computer server 830. Furthermore, the systems and methods described in this disclosure may be performed by one or more computers connected to computer 805 via network interface 850 and network 810. For example, in some embodiments, the systems and methods described in this disclosure may be performed by remote computer 815, computer server 830, or a combination of interconnected computers on network 810.
[0082] Data, datasets, and / or databases used in embodiments of the systems and methods described in this disclosure may be stored and / or downloaded from remote computers 815, web servers 820, cloud storage servers 825, and computer servers 830.
[0083] As used in this application, a circuit is any of the following electronic components (such as semiconductor devices), A networked system may be defined as one or more of a plurality of electronic components, a computer, a network of computing devices, a remote computer, a web server, a cloud storage server, or a computer server, all connected to or interconnected via electronic communications. For example, one or more of a computer, a remote computer, a web server, a cloud storage server, and a computer server may each be encompassed by or include circuitry as a component thereof. In some embodiments, multiple instances of one or more of these components may be used, and each of the multiple instances of one or more of these components may also be encompassed by or include circuitry. In some embodiments, the circuitry represented by a networked system may include a serverless computing system corresponding to a virtualized set of hardware resources. The circuitry represented by a computer may be a personal computer (PC), a desktop computer, a laptop computer, a tablet computer, a netbook computer, a personal digital assistant (PDA), a smartphone, or any other programmable electronic device capable of communicating with other devices on a network. The circuitry may be a general-purpose computer, a special-purpose computer, or other programmable apparatus described herein that includes one or more processors. Each processor may be one or more single-chip or multi-chip microprocessors. A processor is considered a processing circuit or circuitry because it contains transistors and other circuitry within it.A circuit may implement the systems and methods described in this disclosure based on computer-readable program instructions provided to one or more processors (and / or one or more cores within a processor) of one or more general-purpose computers, special-purpose computers, or other programmable devices described in this disclosure to generate a machine, such that the instructions embodied by the circuit or executed via the one or more processors of a programmable device including the circuit create a system for implementing the functions specified in the flow diagrams and block diagrams of this disclosure. Alternatively, a circuit may be a pre-programmed structure such as a programmable logic device, application-specific integrated circuit, or the like, and is considered a circuit whether used alone or in combination with other programmable or pre-programmed circuits.
[0084] Obviously, many modifications and variations of the present disclosure are possible in light of the above teachings. It is therefore to be understood that, within the scope of the appended claims, the present disclosure may be practiced other than as specifically described herein.
Claims
1. 1. A chromatography system comprising: a non-transitory computer-readable storage device having computer-executable code stored therein; an interface that characterizes two-dimensional (2D) information of the enantiomers of the target molecule and receives input data including information of a specific chromatography column; When the computer-executable code is executed, applying the input data to a graph convolutional neural network (GCNN) that has been pre-trained with a dataset of a plurality of chiral molecules associated with the particular chromatography column, the dataset being labeled with retention times of both enantiomers of each of the plurality of chiral molecules in the dataset, the GCNN being trained to set weights and adjust the weights based on backpropagation loss from predicted retention times compared to ground truth predicted retention times; collecting the output of the GCNN as predicted retention times of the enantiomers of the target molecule; discriminating the elution order of the enantiomers of the target molecule based on the relative magnitudes of the predicted retention times; identifying peaks on the chromatogram for each of the enantiomers of the target molecule based on the elution order of the enantiomers of the target molecule; a processing circuit configured to perform Equipped with A chromatography system, wherein the chromatogram is obtained from testing the target molecule on a chromatograph having the specific chromatography column.
2. further comprising the chromatograph having the specific chromatography column; The chromatography system of claim 1 , wherein the chromatograph is implemented by a processing circuit or is separate from the processing circuit.
3. 3. The chromatography system of claim 1, wherein the processing circuitry is further configured to store the elution order of the target molecule in a computer-readable memory and provide the stored elution order to the interface, and the interface provides output data identifying the elution order of the enantiomers of the target molecule associated with the particular chromatography column.
4. The chromatography system according to any one of claims 1 to 3, wherein the target molecule comprises an amino group.
5. 5. The chromatography system of claim 1, wherein the processing circuitry is further configured to determine the content of each enantiomer of the target molecule based on the identified peaks on the chromatogram.
6. 6. The chromatography system according to claim 1, wherein the specific chromatography column comprises one selected from the group consisting of an optically active polymer-type chiral column whose chiral selector is a polysaccharide or a derivative thereof, an optically active poly(meth)acrylamide, an optically active poly(amino acid), and / or an optically active polyamide; an optically active low-molecular-weight compound-type chiral column whose chiral selector is a compound having a binaphthyl structure or a crown ether structure; a protein or glycoprotein having a sugar chain-type chiral column; a nucleic acid-type chiral column; a DNA or RNA-type chiral column; a zwitterion column; an anion exchange column; and a ligand exchange column.
7. 7. The chromatography system of claim 1, wherein the target molecule comprises a compound having an amide group, an aromatic group, a carbonyl group, a nitro group, a sulfonyl group, a cyano group, a hydroxyl group, and / or an amino group.
8. 8. The chromatography system of claim 1, wherein the processing circuitry is configured to implement the GCNN.
9. 9. The chromatography system of claim 1, wherein the GCNN includes at least one pooling layer that extracts chirality features from the input data.
10. 10. The chromatography system of claim 1, wherein the GCNN further comprises a fully connected layer that performs chiral molecule classification based on the extracted chirality features.
11. The chromatography system of any one of claims 1 to 10, further comprising another circuit that implements the GCNN.
12. The chromatography system of claim 11 , wherein the other circuitry comprises a multiprocessor-based computer system.
13. The chromatography system of claim 11 or 12, wherein the other circuitry at least partially comprises cloud computing resources.
14. 14. The chromatography system of claim 1, further comprising a computer programmed to apply the input data to the GCNN via a communication network to other circuits implementing the GCNN.
15. The chromatography system of claim 14 , wherein the communication network comprises, at least in part, internet-based resources.
16. The chromatography system of any one of claims 1 to 15, wherein the input data comprises a characterization of the entire two-dimensional structure of the target molecule.
17. the two-dimensional information is a character string representation of an enantiomer of the target molecule; 17. The chromatography system of claim 1, wherein the processing circuitry is further configured to represent the string representation of the enantiomer of the target molecule in a Simplified Molecular Input Line Entry System (SMILES) compatible format.
18. 18. The chromatography system of claim 17, wherein the processing circuitry is further configured to convert the string representation of the enantiomers of the target molecule in the SMILES-compatible format into an adjacency matrix that is then input as the input data to the GCNN.
19. 1. An identification and purification system comprising: a non-transitory computer-readable storage device having computer-executable code stored therein; an interface that characterizes two-dimensional (2D) information of the enantiomers of the target molecule and receives input data including information of a specific chromatography column; When the computer-executable code is executed, a data set of a plurality of chiral molecules associated with said particular chromatographic column; applying the input data to a graph convolutional neural network (GCNN) trained on a dataset labeled with retention times of both enantiomers of each of the plurality of chiral molecules in the dataset, the GCNN being trained to set weights and adjust the weights based on backpropagation loss from predicted retention times compared to ground truth predicted retention times; collecting the output of the GCNN as predicted retention times of the enantiomers of the target molecule; discriminating the elution order of the enantiomers of the target molecule based on the relative magnitudes of the predicted retention times; identifying peaks on a chromatogram for each enantiomer of the target molecule based on the elution order of the enantiomers of the target molecule, the chromatogram being obtained from passing the target molecule through the particular chromatography column; and a processing circuit configured to: a chromatograph having the specific chromatography column, collecting a predetermined enantiomer of the target molecule based on the identified peak on the chromatogram after the target molecule has passed through the specific chromatography column; An identification and purification system comprising:
20. 20. The identification and purification system of claim 19, wherein the processing circuitry is integrated with or separate from the chromatograph.
21. 21. The identification and purification system of claim 19 or 20, wherein the processing circuitry is further configured to store the elution order of the target molecule in a computer-readable memory and provide the stored elution order to the interface, and the interface provides output data identifying the elution order of the enantiomers of the target molecule associated with the particular chromatography column.
22. The identification and purification system according to any one of claims 19 to 21, wherein the target molecule comprises a compound having an amide group, an aromatic group, a carbonyl group, a nitro group, a sulfonyl group, a cyano group, a hydroxyl group, and / or an amino group.
23. 1. A chromatography method comprising: receiving input data characterizing two-dimensional (2D) information of an enantiomer of a target molecule and including information of a particular chromatography column; applying the input data to a graph convolutional neural network (GCNN) trained with a dataset of a plurality of chiral molecules associated with the particular chromatography column, the dataset labeled with retention times of both enantiomers of each of the plurality of chiral molecules in the dataset, the GCNN being trained to set weights and adjust the weights based on backpropagation loss from predicted retention times compared to ground truth predicted retention times; collecting the output of the GCNN as predicted retention times of the enantiomers of the target molecule; discriminating the elution order of the enantiomers of the target molecule based on the relative magnitudes of the predicted retention times; identifying peaks on the chromatogram for each of the enantiomers of the target molecule based on the elution order of the enantiomers of the target molecule; Including, A chromatography method, wherein the chromatogram is obtained from testing the target molecule on a chromatograph having the specific chromatography column.
24. 24. The chromatography method of claim 23, further comprising storing the elution order of the target molecule in a computer-readable memory and providing the stored elution order to the interface, wherein the interface provides output data identifying the elution order of the enantiomers of the target molecule associated with the particular chromatography column.
25. 25. The chromatographic method of claim 23 or 24, wherein the target molecules comprise compounds having an amide group, an aromatic group, a carbonyl group, a nitro group, a sulfonyl group, a cyano group, a hydroxyl group, and / or an amino group.
26. 1. A method for identification and purification, comprising: receiving input data characterizing two-dimensional (2D) information of an enantiomer of a target molecule and including information of a particular chromatography column; applying the input data to a graph convolutional neural network (GCNN) trained with a dataset of a plurality of chiral molecules associated with the particular chromatography column, the dataset labeled with retention times of both enantiomers of each of the plurality of chiral molecules in the dataset, the GCNN being trained to set weights and adjust the weights based on backpropagation loss from predicted retention times compared to ground truth predicted retention times; collecting the output of the GCNN as predicted retention times of the enantiomers of the target molecule; discriminating the elution order of the enantiomers of the target molecule based on the relative magnitudes of the predicted retention times; identifying peaks on a chromatogram for each enantiomer of the target molecule based on the elution order of the enantiomers of the target molecule, the chromatogram being obtained from passing the target molecule through the particular chromatography column; collecting, by a chromatograph having the specific chromatography column, a predetermined enantiomer of the target molecule based on the identified peak on the chromatogram after the target molecule has passed through the specific chromatography column; The identification and purification method includes:
27. 27. The identification and purification method of claim 26, further comprising storing the elution order of the target molecule in a computer readable memory and providing the stored elution order to the interface, wherein the interface provides output data identifying the elution order of the enantiomers of the target molecule associated with the particular chromatography column.
28. 28. The method of claim 26 or 27, wherein the target molecule comprises a compound having an amide group, an aromatic group, a carbonyl group, a nitro group, a sulfonyl group, a cyano group, a hydroxyl group, and / or an amino group.