Chromatographic outflow curve prediction method and system based on ensemble learning and computational chemistry

Through integrated learning and computational chemical methods, chromatographic experimental parameters and molecular identifiers are processed, combined with machine learning models, the accuracy and efficiency problems of chromatographic efflux curve prediction are solved, and more accurate prediction and environmentally friendly experimental conditions are achieved.

CN120544698APending Publication Date: 2025-08-26SUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510501105.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In the prior art, the prediction method of chromatographic effluent curves depends on empirical formulas or standard substance calibration, and there are problems with limited scope of application and low prediction accuracy, and the operation is cumbersome.

Method used

Using an integrated learning and computational chemistry method, the chromatographic experimental parameters and molecular identifiers are obtained, and the chromatographic outflow curve information is predicted by obtaining chromatographic experimental parameters and molecular identifiers, and by combining machine learning regression models or ensemble learning methods.

Benefits of technology

It realizes a more accurate prediction of chromatographic efflux curve, can reversely deduce the chromatographic experimental conditions, reduce waste of pre-experiment reagents, and is of environmental significance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544698A_ABST
    Figure CN120544698A_ABST
Patent Text Reader

Abstract

The invention discloses a chromatographic outflow curve prediction method and system based on ensemble learning and computational chemistry, and belongs to the field of analytical chemistry and machine learning. The chromatographic outflow curve prediction method based on ensemble learning and computational chemistry comprises the following steps: acquiring chromatographic experimental parameters, and processing the chromatographic experimental parameters to obtain experimental parameter vectors; obtaining a molecular identifier, and processing the molecular identifier to obtain a multi-modal molecular vector; combining the experimental parameter vector and the multi-modal molecular vector to obtain an input vector; and based on the input vector, predicting by adopting a machine learning regression model or an integrated learning method to obtain chromatographic outflow curve information. According to the method, chemical and physical methods are combined to process experimental parameters, a computational chemistry method is adopted to obtain accurate molecule descriptors, a graph neural network is utilized to capture molecular structure information, in addition, the method supports reverse reasoning of proper experimental conditions, and pre-experimental consumption is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of analytical chemistry and machine learning, and particularly relates to a chromatographic elution curve prediction method and system based on integrated learning and computational chemistry. Background Art

[0002] In the fields of medicine, pharmacy, biology, chemistry, environmental science, and materials science, chromatography is an important technique widely used for qualitative and quantitative analysis of molecules. The chromatographic elution curve, as the experimental result of chromatography, contains information such as retention time, peak height, peak width, and peak shape, which can be used for qualitative and quantitative analysis of compounds. Retention time is the primary basis for qualitative analysis. Under the same conditions, the experimental value is compared with the literature value in the database. If the two are the same, the molecule is confirmed. However, the database cannot contain the results of all molecules under all experimental conditions. Therefore, the relevant field is looking for a method to predict the results.

[0003] Traditional non-algorithmic retention time prediction methods often rely on empirical formulas or calibration with standard substances. These methods have problems such as limited applicability, low prediction accuracy, and cumbersome operation. To address this issue, a chromatographic elution curve prediction method based on ensemble learning and computational chemistry was proposed. Summary of the Invention

[0004] In view of the deficiencies in the prior art, the present invention aims to provide a chromatographic elution curve prediction method and system based on integrated learning and computational chemistry, which solves the problems in the prior art.

[0005] The purpose of the present invention can be achieved through the following technical solutions:

[0006] The chromatographic elution curve prediction method based on integrated learning and computational chemistry includes the following steps:

[0007] Obtaining chromatographic experimental parameters and processing them to obtain experimental parameter vectors;

[0008] Obtain molecular identifiers and process them to obtain multimodal molecular vectors;

[0009] Merge the experimental parameter vector and the multimodal molecular vector to obtain the input vector;

[0010] Based on the input vector, the chromatographic elution curve information is predicted using a machine learning regression model or an ensemble learning method.

[0011] Furthermore, the chromatographic experiment parameters include: one or more of: chromatographic column, injection volume, column temperature, column pressure, flow rate, mobile phase type and mobile phase ratio.

[0012] Furthermore, the chromatographic experiment parameters include chromatographic column, injection volume, column temperature, column pressure, flow rate, mobile phase type and mobile phase ratio. The steps of processing the chromatographic experiment parameters to obtain the experimental parameter vector are:

[0013] Encode the columns into column vectors according to their types;

[0014] Combine injection volume, column temperature, column pressure and flow rate into a configuration parameter vector;

[0015] Encode the mobile phase type into an embedded vector and combine it with the mobile phase proportion to form a mobile phase vector;

[0016] Encode the column vector, configuration parameter vector, and mobile phase vector into the original experimental parameter vector;

[0017] The original experimental parameter vector is processed by chemical and physical methods to obtain an optimized experimental parameter vector, which is then merged with the original experimental parameter vector to obtain an experimental parameter vector.

[0018] Furthermore, the step of processing the molecular identifier to obtain the multimodal molecular vector includes:

[0019] The molecular identifiers are converted through RDKit or the PubChem database to obtain: molecular fingerprint vector, molecular descriptor vector and molecular coordinate information;

[0020] The molecular coordinate information is optimized using computational chemistry programs to obtain atomic coordinate tensors;

[0021] Input the atomic coordinate tensor into the computational chemistry program to obtain the computational chemistry vector;

[0022] Input the atomic coordinate tensor into the graph neural network and output the graph neural network vector;

[0023] Molecular fingerprint vectors, molecular descriptor vectors, computational chemistry vectors, and graph neural network vectors are combined into multimodal molecular vectors.

[0024] Furthermore, the graph neural network adopts the MPNN architecture.

[0025] Furthermore, the ensemble learning method is: using a combination of multiple machine learning regression models for prediction.

[0026] Chromatographic elution curve prediction system based on integrated learning and computational chemistry, including:

[0027] Parameter processing module: obtains chromatographic experimental parameters and processes them to obtain experimental parameter vectors;

[0028] Identifier processing module: obtains molecular identifiers and processes them to obtain multimodal molecular vectors;

[0029] Vector merging module: merges the experimental parameter vector and the multimodal molecular vector to obtain the input vector;

[0030] And, prediction module: based on the input vector, a machine learning regression model or an integrated learning method is used to predict the chromatographic elution curve information.

[0031] A computer storage medium stores a readable program, which can execute the above-mentioned chromatographic elution curve prediction method based on integrated learning and computational chemistry when the program is run.

[0032] An electronic device, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus;

[0033] The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the above-mentioned chromatographic elution curve prediction method based on integrated learning and computational chemistry.

[0034] A computer program product includes computer instructions, wherein the computer instructions instruct a computing device to execute operations corresponding to the above-mentioned chromatographic elution curve prediction method based on integrated learning and computational chemistry.

[0035] Beneficial effects of the present invention:

[0036] 1. This invention uses computational chemistry methods to obtain atomic coordinates using precise computational chemistry methods, and then calculates molecular descriptors. Compared with the molecular descriptors obtained by RDKit (an algorithm library mainly based on two-dimensional simplified molecular structures) in the prior art, the results can better reflect the molecular properties and help the model obtain better results.

[0037] 2. The computational chemistry molecular descriptors used in this paper, including the molecular polarity index and the solubility free energy of the molecule in the solvent, directly reflect the three-dimensional molecular properties and their interactions with the mobile phase through quantum mechanical calculations, compared to the descriptors obtained by RDKit used in the prior art. This breaks through the limitations of traditional empirical descriptors in characterizing molecular features, enabling the model to more accurately predict molecular behavior during chromatographic separations.

[0038] 3. The present invention combines chemical and physical methods to process the properties of experimental parameters. The combination of chemical and physical empirical methods is beneficial to the model to obtain better results;

[0039] 4. The present invention uses an embedded vector method to describe nodes and chromatographic columns in graph neural networks, which can more comprehensively reflect their properties compared to manual design in the prior art;

[0040] 5. The method of the present invention can predict peak shape, peak height, and peak width. Compared with the prior art where experiments are limited to retention time, it can reversely deduce chromatographic experimental conditions and reduce the waste of reagents generated in pre-experiments to a certain extent. It has a positive effect on global environmental protection and comprehensive governance, as well as carbon peak and carbon neutrality. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0042] Figure 1 This is a flow chart of the chromatographic elution curve prediction method based on integrated learning and computational chemistry of the present invention;

[0043] Figure 2 This is a graph showing the relationship between loss value and number of training rounds of the present invention;

[0044] Figure 3 is a scatter plot of predicted values ​​and observed values ​​of the present invention;

[0045] Figure 4 It is a deviation scatter plot of the present invention. DETAILED DESCRIPTION

[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0047] Example 1

[0048] like Figure 1 As shown, the chromatographic elution curve prediction method based on integrated learning and computational chemistry includes the following steps:

[0049] S1, obtain the chromatographic experimental parameters and process them to obtain the experimental parameter vector;

[0050] In this embodiment, the chromatographic experiment parameters include: chromatographic column, injection volume, column temperature, column pressure, flow rate, mobile phase type and mobile phase ratio; in other embodiments, the chromatographic experiment parameters include: at least one of chromatographic column, injection volume, column temperature, column pressure, flow rate, mobile phase type and mobile phase ratio;

[0051] Chromatography includes gas chromatography and liquid chromatography; the gas chromatography is any one of packed column gas chromatography, capillary column gas chromatography, headspace gas chromatography, pyrolysis gas chromatography and programmed temperature gas chromatography; the liquid chromatography is any one of normal phase chromatography, reverse phase chromatography, ion exchange chromatography, size exclusion chromatography, affinity chromatography, chiral chromatography, hydrophobic interaction chromatography, hydrophilic interaction chromatography, preparative liquid chromatography and ultra-high performance liquid chromatography.

[0052] In this embodiment, the step of processing the chromatographic experiment parameters to obtain the experimental parameter vector includes:

[0053] S11, encode the chromatographic columns into chromatographic column vectors according to their types;

[0054] S12, combining injection volume, column temperature, column pressure and flow rate into a configuration parameter vector;

[0055] In other embodiments, one of the injection volume, column temperature, column pressure and flow rate can be used as a configuration parameter vector; or multiple injection volume, column temperature, column pressure and flow rate can be combined to obtain a configuration parameter vector.

[0056] S13, encoding the mobile phase type into an embedded vector and combining it with the mobile phase proportion to form a mobile phase vector;

[0057] S14, encodes the column vector, configuration parameter vector, and mobile phase vector into the original experimental parameter vector;

[0058] In other embodiments, the original experimental parameter vector may also be obtained by encoding one or more vectors of: a chromatographic column vector, a configuration parameter vector, and a mobile phase vector;

[0059] S15, processing the original experimental parameter vector by chemical and physical methods to obtain an optimized experimental parameter vector, and merging it with the original experimental parameter vector to obtain an experimental parameter vector;

[0060] In other embodiments, the original experimental parameter vector may be directly used as the experimental parameter vector;

[0061] Among them, the chemical and physical methods include the van Deemter equation:

[0062]

[0063] Wherein, A is the eddy diffusion term, B is the longitudinal diffusion term, C is the mass transfer resistance term, u is the flow velocity, and H is the theoretical plate height. In this embodiment, the van Deemter equation is used to process the original experimental parameter vector of the model, wherein A, B, and C are all vectors obtained by removing the flow velocity u from the original experimental parameter vector of the model and passing it through a two-layer neural network. The vectors A, B, and C are then calculated with the flow velocity u through the van Deemter equation to obtain the final vector H, which is the optimized experimental parameter.

[0064] S2, obtain molecular identifiers and process them to obtain multimodal molecular vectors;

[0065] The molecular identifier is any one of the following: molecular name, structure identifier, and database number;

[0066] The structure identifier is any one of SMILES, InChI, InChIKey and Smarts;

[0067] The database number is any one of: CAS number, PubChem CID number, HSDB number, CCRIS number, NSC number, FEMA number, Caswell number, EINECS number, ChEBI number, EC number, UNII number, UN number, RCRA number and EPA number;

[0068] In this embodiment, the steps of processing the molecular identifier to obtain the multimodal molecular vector include:

[0069] S21, if the molecule identifier is a molecule name or a database number, obtain the structure identifier through the database;

[0070] S22, convert the structure identifier through RDKit or PubChem database to obtain: molecular fingerprint vector, molecular descriptor vector and molecular coordinate information;

[0071] S23, using computational chemistry programs to optimize the molecular coordinate information and obtain the stable conformation atomic coordinate tensor;

[0072] S24, inputting the atomic coordinate tensor into a computational chemistry program to obtain a computational chemistry vector;

[0073] S25, inputting the atomic coordinate tensor into the graph neural network and outputting a graph neural network vector;

[0074] S26, combining the molecular fingerprint vector, the molecular descriptor vector, the computational chemistry vector, and the graph neural network vector into a multimodal molecular vector;

[0075] In other embodiments, S22 obtains one or more of: a molecular fingerprint vector, a molecular descriptor vector, and molecular coordinate information through RDKit conversion or the PubChem database. In other embodiments, the multimodal molecular vector in S26 may further include one or more of: a molecular fingerprint vector, a molecular descriptor vector, a computational chemistry vector, and a graph neural network vector.

[0076] In S22, the molecular fingerprint vector includes one or more combinations of the following: RDKit fingerprint, atom pair fingerprint, topological torsion fingerprint, MACCS bond fingerprint, Morgan fingerprint, pharmacological pharmacophore fingerprint, pattern fingerprint, extended simplified graph fingerprint, molecular hash fingerprint (mhfp_fp), and structural feature fingerprint.

[0077] In S22, the molecular descriptor vector includes one or more combinations of the following: complexity, conformer_id_3d, conformer_rmsd_3d, mmff94_energy_3d, mmff94_partial_charges_3d, multipoles_3d, monoisotopic_mass, pharmacophore_features_3d, shape_selfoverlap_3d, tpsa, volume_3d, xlogp, estate_index series descriptors, qed, sps, heavy_atom_molecular_weight, max_partial_charge, min_partial_charge, bcut2d series descriptors, avg_ipc , balaban_j, bertz_ct, chi0v to chi4v, chi0n to chi4n, hall_kier_alpha, ipc, kappa1 to kappa3, labute_asa, peoe_vsa1 to peoe_vsa14, smr_vsa1 to smr_vsa10, slogp_vsa1 to slogp_vsa12, tpsa, estate_vsa1 to estate_vsa11, vsa_estate1 to vsa_estate10, mol_logp, mol_mr, molecular_weight, exact_mass, coordinate_type, feature_selfoverlap_3d, charge, count series descriptor, num series descriptor, fr series descriptor;

[0078] The descriptor system is generated by chemical informatics tools, and its technical essence is compatible with the RDKit and PubChem descriptor systems, with the same names or only the naming method changed; the "to" is a numerical increment, and the descriptors it represents are in an "or" relationship; the series of descriptors are in an "or" relationship; the estate_index series descriptors include max_abs_estate_index, max_estate_index, min_abs_estate_index, and min_estate_index; the bcut2d series descriptors include bcut2d_mwhi, bcut2 d_mwlow, bcut2d_chghi, bcut2d_chglow, bcut2d_logphi, bcut2d_logplow, bcut2d_mrhi, bcut2d_mrlow; the count series descriptors include atom_stereo_count, bo nd_stereo_count, covalent_unit_count, defined_atom_stereo_count, defined_bond_stereo_count, effective_rotor_count_3d, heavy_atom_c ount, isotope_atom_count, undefined_atom_stereo_count, undefined_bond_stereo_count, heavy_atom_count, nhoh_count, no_count, ring_count, h_bond_acceptor_count, h_bond_donor_count; the num series descriptors include num_aliphatic_carbocycles, num_aliphatic_heterocycles, num_aliphatic_ rings, num_aromatic_carbocycles, num_aromatic_heterocycles, num_aromatic_rings, num_heteroatoms, num_rotatable_bonds, num_saturated_carbocycles, num_saturated_heterocycles, num_saturated_rings; the fr series descriptors include fr_al_coo, fr_al_oh, fr_al_oh_notert, fr_arn, fr_ar_coo,fr_ar_n、fr_ar_nh、fr_ar_oh、fr_coo、fr_coo2、fr_c_o、fr_c_o_nocoo、fr_c_s、fr_hoccn、fr_imine、fr_nh0、fr_nh1、fr_nh2、fr_n_o、fr_ndealkylation1、fr_ndealkylation2、fr_nhpyrrole、fr_sh、fr_aldehyde、fr_alkyl_carbamate、fr_alkyl_halide、fr_allylic_oxid、fr_amide、fr_amidine、fr_aniline、fr_aryl_methyl、fr_azide、fr_azo、fr_barbitur、fr_benzene、fr_benzodiazepine、fr_bicyclic、fr_diazo、fr_dihydropyridine、fr_epoxide、fr_ester、fr_ether、fr_furan、fr_guanido、fr_halogen、fr_hdrzine、fr_hdrzone、fr_imidazole、fr_imide、fr_isocyan、fr_isothiocyan、fr_ketone、fr_ketone_topliss、fr_lactam、fr_lactone、fr_methoxy、fr_morpholine、fr_nitrile、fr_nitro、fr_nitro_arom、fr_nitro_arom_nonortho、fr_nitroso、fr_oxazole、fr_oxime、fr_para_hydroxylation、fr_phenol、fr_phenol_noorthohbond、fr_phos_acid、fr_phos_ester、fr_piperdine、fr_piperzine、fr_priamide、fr_prisulfonamd、fr_pyridine、fr_quatn、fr_sulfide、fr_sulfonamd、fr_sulfone、fr_term_acetylene、fr_tetrazole、fr_thiazole、fr_thiocyan、fr_thiophene、fr_unbrch_alkane、fr_urea。、

[0079] The computational chemistry programs in S23 and S24 are any one or more of the following programs: RDKit, Gaussian, ORCA, or other computational chemistry programs that achieve the same functions.

[0080] The chemical vector calculated in S24 includes one or more of the following descriptors: molecular polarity index, solubility free energy in solvent, dipole moment, molecular polarizability, and molecular surface electrostatic potential extreme value; the molecular polarity index MPI is defined as follows:

[0081]

[0082] Where V is the molecular electrostatic potential, the integral is the integration of the molecular surface S, A is the molecular surface area, and r is the position coordinate vector.

[0083] The graph neural network in S25 adopts the MPNN (message passing neural network) architecture. In the MPNN architecture, molecules are regarded as graph structures, which contain node information and edge information. Atoms are nodes, and randomized or pre-trained embedding vectors are used for initialization.

[0084] Each atom obtains the distance, edge, angle, and dihedral information of the surrounding atoms during message transmission. The message transmission is a total of L rounds. In the lth round, node i obtains the distance information d of the surrounding atoms j, k, and w. ij d jk d jw , angle information α ijk , α jkw , dihedral angle information And each node vector Through the interactive function Get surrounding transmission news That is, formula E1:

[0085]

[0086] Then, the surrounding message aggregation obtains the information of node i That is, formula E2, where N(i) represents the neighbor nodes of node i on the graph:

[0087]

[0088] Next, use the node update function Update node information and enter the next round, that is, formula E3:

[0089]

[0090] After L rounds of transmission, the final node vector of each node is obtained Then use the read function freadout The graph neural network vector described in S25 is obtained, that is, formula E4, where V is the set of all nodes. The set in the formula means that the results of each round of L are input into the readout function, and finally the graph neural network vector is obtained:

[0091]

[0092] Among them, the interaction function Node update function Read out function f readout Both use trainable neural networks.

[0093] S3, merge the experimental parameter vector and the multimodal molecular vector to obtain the input vector;

[0094] S4, based on the input vector, a machine learning regression model or an ensemble learning method is used to predict the chromatographic elution curve information;

[0095] The machine learning regression model used includes: any one of a tree-based ensemble model, a support vector machine, or a neural network.

[0096] The ensemble learning method makes predictions by combining multiple machine learning regression models. Specifically, it uses bagging, boosting, stacking or blending methods to make ensemble predictions by combining two or more models from tree-based ensemble models, support vector machines and neural networks.

[0097] In this embodiment, the chromatographic elution curve information includes: retention time, adjusted retention time, left peak bottom time, adjusted left peak bottom time, right peak bottom time, adjusted right peak bottom time, peak width, peak top height, left half peak height information, right half peak height information, and peak area. In other embodiments, the chromatographic elution curve information includes: retention time, adjusted retention time, left peak bottom time, adjusted left peak bottom time, right peak bottom time, adjusted right peak bottom time, peak width, peak top height, left half peak height information, right half peak height information, and peak area.

[0098] The adjustment refers to: subtracting the dead time from the original time; the dead time refers to: the retention time of the component that does not interact with the stationary phase; the peak height is the peak height at the retention time or the adjusted retention time;

[0099] The left half peak height information is obtained by taking N time points between the left peak bottom time and the retention time, or between the adjusted left peak bottom time and the adjusted retention time, and obtaining the peak height or the ratio of the peak height to the peak top height at each time point;

[0100] The right half peak height information is obtained by taking N time points between the retention time and the right peak bottom time, or between the adjusted right peak bottom time and the adjusted retention time, and obtaining the peak height or the ratio of the peak height to the peak top height at each time point;

[0101] The left half peak height information and the right half peak height information are integers greater than or equal to 1;

[0102] The peak bottom width is equal to the right peak bottom time minus the left peak bottom time.

[0103] Chromatographic elution curve information can be displayed numerically, plotted as a table, or plotted as a graph.

[0104] Chromatographic elution curve information can be used to reversely infer appropriate experimental conditions; where appropriate experimental conditions meet one or both of the following conditions:

[0105] 1) making at least one substance have a good peak shape;

[0106] 2) enabling the mixture to be completely separated; wherein, the basis for complete separation is to satisfy Formula F1:

[0107]

[0108] Where R is the separation, t1 and t2 represent the retention time or adjusted retention time of the two substances, W1 and W2 represent the peak base width of the two substances, and R0 is the minimum separation for complete separation, which is a real number greater than or equal to 1.

[0109] The reverse reasoning method involves an optimization problem in which decision variables are used to make a target value satisfy constraints. The decision variables include one or more combinations of the following: chromatographic column type, injection volume, column temperature, column pressure, flow rate, mobile phase type, and mobile phase ratio. The target value is the resolution R, which is desired to be adjusted to satisfy the constraints. The constraints are to satisfy Formula F1. The optimization problem is solved using gradient descent, simulated annealing, ant colony algorithm, genetic algorithm, particle swarm optimization, differential evolution, or other common optimization problem-solving methods.

[0110] According to S1-S4, a database is constructed for training and the experimental parameters can be obtained.

[0111] Based on similar inventive concepts, an embodiment of the present invention further provides a computer storage medium storing a readable program, which, when run, can execute the above-mentioned chromatographic elution curve prediction method based on integrated learning and computational chemistry.

[0112] Based on similar inventive concepts, an embodiment of the present invention provides an electronic device, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus;

[0113] The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the above-mentioned chromatographic elution curve prediction method based on integrated learning and computational chemistry.

[0114] Based on similar inventive concepts, an embodiment of the present invention further provides a computer program product, comprising computer instructions, which instruct a computing device to execute operations corresponding to the above-mentioned chromatographic elution curve prediction method based on integrated learning and computational chemistry.

[0115] Example 2

[0116] In this embodiment, the technical solution of the present invention is described through specific examples;

[0117] A data set consisting of 3275 data records was constructed. The chromatographic method used was reversed-phase chromatography, the chromatographic columns used were Cosmetics C18-MS-II and Cosmetics C18-AR-II, and the mobile phases used were water, methanol, 20 mmol / L phosphate buffer at pH = 2.5, and 20 mmol / L phosphate buffer at pH = 7. The retention times were all between 2.23 and 17.58 min.

[0118] The original experimental parameter vector of the model includes the chromatographic column, sample concentration and injection volume, mobile phase type and mobile phase ratio, and the chromatographic column is encoded as a 16-dimensional embedding vector.

[0119] The van Deemter equation is used to process the model's original experimental parameter vector, where A, B, and C are the vectors obtained by removing the flow velocity u from the model's original experimental parameter vector and passing it through a two-layer neural network. The vectors A, B, and C are then calculated with the flow velocity u through the van Deemter equation to obtain the final vector H. Each layer of the neural network has 64 hidden values, the loss function is the ReLU function, and the output optimized experimental parameters (i.e., H) are 8-dimensional.

[0120] The SMILES molecular identifier was selected, and the Morgan fingerprint was calculated with a radius of 2 and a bit number of 2048. The following molecular descriptors were obtained using PubChem: complexity, conformer_id_3d, conformer_rmsd_3d, mmff94_energy_3d, mmff94_partial_charges_3d, multipoles_3d, monoisotopic_mass, pharmacophore_features_3d, shape_selfoverlap_3d, tpsa, volume_3d, xlogp, atom_stereo_count, bond_stereo_count, coordinate_type, covalent_unit_count, defined_atom_stereo_count, defined_bond_stereo_count, effective_rotor_count_3d, feature_selfoverlap_3d, heavy_atom_count, isotope_atom_count, undefined_atom_stereo_count, and undefined_atom_stereo_count. _bond_stereo_count, use RDKit to get the following molecular descriptors (S10): max_abs_estate_index, max_estate_index, min_abs_estate_index, min_estate_index, qed, sps, heavy_atom_molecular_weight, max_partial_charge, min_partial_charge, bcut2d_mwhi, bcut2d_mwlow, bcut2d_chghi, bcut2d_chglow, bcut2 d_logphi, bcut2d_logplow, bcut2d_mrhi, bcut2d_mrlow, avg_ipc, balaban_j, bertz_ct, chi0v, chi1v, chi2v, chi3v, chi4v, chi0n, chi1n, chi2n, chi3n, chi4n, hall_kier_alpha, ipc, kappa1, kappa2, kappa3, labute_asa, peoe_vsa1, peoe_vsa2, peoe_vsa3, peoe_vsa4, peoe_vsa5, peoe_vsa6,peoe_vsa7、peoe_vsa8、peoe_vsa9、peoe_vsa10、peoe_vsa11、peoe_vsa12、peoe_vsa13、peoe_vsa14、smr_vsa1、smr_vsa2、smr_vsa3、smr_vsa4、smr_vsa5、smr_vsa6、smr_vsa7、smr_vsa8、smr_vsa9、smr_vsa10、slogp_vsa1、slogp_vsa2、slogp_vsa3、slogp_vsa4、slogp_vsa5、slogp_vsa6、slogp_vsa7、slogp_vsa8、slogp_vsa9、slogp_vsa10、slogp_vsa11、slogp_vsa12、tpsa、estate_vsa1、estate_vsa2、estate_vsa3、estate_vsa4、estate_vsa5、estate_vsa6、estate_vsa7、estate_vsa8、estate_vsa9、estate_vsa10、estate_vsa11、vsa_estate1、vsa_estate2、vsa_estate3、vsa_estate4、vsa_estate5、vsa_estate6、vsa_estate7、vsa_estate8、vsa_estate9、vsa_estate10、mol_logp、mol_mr、charge、heavy_atom_count、nhoh_count、no_count、num_aliphatic_carbocycles、num_aliphatic_heterocycles、num_aliphatic_rings、num_aromatic_carbocycles、num_aromatic_heterocycles、num_aromatic_rings、num_heteroatoms、num_rotatable_bonds、num_saturated_carbocycles、num_saturated_heterocycles、num_saturated_rings、ring_count、fr_al_coo、fr_al_oh、fr_al_oh_notert、fr_arn、fr_ar_coo、fr_ar_n、fr_ar_nh、fr_ar_oh、fr_coo、fr_coo2、fr_c_o、fr_c_o_nocoo, fr_c_s, fr_hoccn, fr_imine, fr_nh0, fr_nh1, fr_nh2, fr_n_o, fr_ndealkylation1, fr_ndealkylation2, fr_nhpyrrole, fr_sh, fr_aldehyde, fr_alkyl_carbamate, fr_alkyl_halide, fr_allylic_oxid, fr_amide, fr_amidine, fr_aniline, fr_aryl_methyl, fr_azide, fr_azo, fr_barbitur, fr_benzene, fr_benzodiazepine, fr_bicyclic, fr_diazo, fr_dihydropyridine, fr_epoxide, fr_ester, fr_ether, fr_furan, fr_guanido, fr_halogen, fr_hdrzine, fr_hdrzone, fr_imidazole, fr_imide, fr_isocyan, fr_isothiocyan, fr_ketone, fr_ketone_topliss, fr_lactam, fr_lactone, fr_methoxy, fr_morpholine, fr_nitrile, fr_nitro, fr_nitro_arom, fr_nitro_arom_nonortho, fr_nitroso, fr_oxazole, fr_oxime, fr_para_hydroxylation, fr_phenol, fr_phenol_noorthohbond, fr_phos_acid, fr_phos_ester, fr_piperdine, fr_piperzine, fr_priamide, fr_prisulfonamd, fr_pyridine, fr_quatn, fr_sulfide, fr_sulfonamd, fr_sulfone, fr_term_acetylene, fr_tetrazole, fr_thiazole, fr_thiocyan, fr_thiophene, fr_unbrch_alkane, fr_urea, the following molecular descriptors that can be obtained by both PubChem and RDKit, obtained by RDKit: molecular_weight, exact_mass, h_bond_acceptor_count,h_bond_donor_count, all the above molecular descriptors, constitute the molecular descriptor vector; the RDKit version number used is 2023.9.5.

[0121] Computational chemistry methods were used to calculate the polarity index and solubility free energy of the molecule in water. Specifically, ORCA5.0.4 software was used for theoretical calculations, and the wave function information of the molecule was calculated at the B3LYP-D3 (BJ) / def2-TZVP / / B97-3c level. The molecular polarity index was calculated using Multiwfn3.8 (dev), which is defined as the ratio of the integral of the molecular surface electrostatic potential to the molecular surface area. The larger the MPI, the greater the overall polarity of the molecule. The solubility free energy was calculated using the SMD implicit solvent model at the M052X / 6-31G * level.

[0122] The graph neural network parameters use the MPNN model, where the interaction function Node update function Both are neural networks with 2 hidden layers, 128 parameters per layer, using sigmoid function as activation function and update function f readout After summing up all node vectors, the neural network is passed through two hidden layers with 128 parameters per layer. The sigmoid function is used as the activation function, and the output is a 64-dimensional graph neural network vector.

[0123] A deep neural network (DNN) is used as a meta-learner. In the DNN parameter setting, the maximum and minimum normalization method is used; the loss function is set to MSE; the learning rate is set to 0.000001; the L2 regularization coefficient is set to 0.000001; the Adam optimizer is used; the network has 2 hidden layers, each containing 8192 parameters; the activation function is ReLU; and the output is 1 value.

[0124] Random Forest (RF) is used as another meta-learner in parallel. The number of trees is set to 100, Bootstrap sampling is used, and the node splitting criterion is MSE.

[0125] Using ensemble learning (S22), DNN and RF are trained independently, and the final prediction result is the weighted average of their outputs, with DNN weight = 0.7 and RF weight = 0.3.

[0126] The model was obtained by using 3111 data (95%) as the training set and 164 data (5%) as the test set, and training for 3000 rounds.

[0127] Figure 2 A graph showing the relationship between the loss value and the number of training rounds of the model trained on the test set of 3111 data items is given, and the training reaches convergence.

[0128] To further illustrate the prediction effect, this embodiment statistically analyzes the error distribution of the 164 sets of predicted values ​​and observed values. Figure 3 A scatter plot of the predicted and observed values ​​of the trained model on the test set of 164 data items is given. Figure 4 The deviation scatter plot is given, and it can be seen that the prediction results are good, which verifies the effectiveness of the method.

[0129] It is particularly noted that the effectiveness of the model is directly related to the number of data items in the data set. Compared with previous inventions, the present invention adjusts the prediction model architecture. The invention content does not include the construction of the data set, but the results generated on the data set can prove its practicality; other implementation methods of replacing the data set should also be considered within the scope of protection of the present invention.

[0130] Example 3

[0131] This embodiment proposes a chromatographic elution curve prediction system based on integrated learning and computational chemistry, including:

[0132] Parameter processing module: obtains chromatographic experimental parameters and processes them to obtain experimental parameter vectors;

[0133] Identifier processing module: obtains molecular identifiers and processes them to obtain multimodal molecular vectors;

[0134] Vector merging module: merges the experimental parameter vector and the multimodal molecular vector to obtain the input vector;

[0135] And, prediction module: based on the input vector, a machine learning regression model or an integrated learning method is used to predict the chromatographic elution curve information.

[0136] The method of the present invention can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CDROM, RAM, floppy disk, hard disk or magneto-optical disk), or as computer code that is originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded over a network and will be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a special-purpose processor or programmable or special-purpose hardware (such as an ASIC or FPGA). It will be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component (e.g., RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by a computer, a processor or hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown here, the execution of the code converts the general-purpose computer into a special-purpose computer for executing the method shown here.

[0137] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention, and such changes and modifications fall within the scope of the invention as claimed.

Claims

1. A chromatographic elution curve prediction method based on ensemble learning and computational chemistry, characterized in that: The following steps are involved: Obtaining chromatographic experimental parameters and processing them to obtain experimental parameter vectors; Obtain molecular identifiers and process them to obtain multimodal molecular vectors; Merge the experimental parameter vector and the multimodal molecular vector to obtain the input vector; Based on the input vector, the chromatographic elution curve information is predicted using a machine learning regression model or an ensemble learning method.

2. The chromatographic elution curve prediction method based on ensemble learning and computational chemistry according to claim 1, characterized in that: The chromatographic experiment parameters include: one or more of: chromatographic column, injection volume, column temperature, column pressure, flow rate, mobile phase type and mobile phase ratio.

3. The chromatographic elution curve prediction method based on ensemble learning and computational chemistry according to claim 1, characterized in that: The chromatographic experiment parameters include chromatographic column, injection volume, column temperature, column pressure, flow rate, mobile phase type and mobile phase ratio. The steps of processing the chromatographic experiment parameters to obtain the experimental parameter vector are as follows: Encode the columns into column vectors according to their types; Combine injection volume, column temperature, column pressure and flow rate into a configuration parameter vector; Encode the mobile phase type into an embedded vector and combine it with the mobile phase proportion to form a mobile phase vector; Encode the column vector, configuration parameter vector, and mobile phase vector into the original experimental parameter vector; The original experimental parameter vector is processed by chemical and physical methods to obtain an optimized experimental parameter vector, which is then merged with the original experimental parameter vector to obtain an experimental parameter vector.

4. The chromatographic elution curve prediction method based on ensemble learning and computational chemistry according to claim 1, characterized in that: The steps of processing the molecular identifier to obtain the multimodal molecular vector include: The molecular identifiers are converted through RDKit or the PubChem database to obtain: molecular fingerprint vector, molecular descriptor vector and molecular coordinate information; The molecular coordinate information is optimized using computational chemistry programs to obtain atomic coordinate tensors; Input the atomic coordinate tensor into the computational chemistry program to obtain the computational chemistry vector; Input the atomic coordinate tensor into the graph neural network and output the graph neural network vector; Molecular fingerprint vectors, molecular descriptor vectors, computational chemistry vectors, and graph neural network vectors are combined into multimodal molecular vectors.

5. According to the chromatographic elution curve prediction method based on integrated learning and computational chemistry according to claim 4, the graph neural network adopts the MPNN architecture.

6. The chromatographic elution curve prediction method based on ensemble learning and computational chemistry according to claim 1, wherein the ensemble learning method comprises: using a combination of multiple machine learning regression models for prediction.

7. A chromatographic elution curve prediction system based on integrated learning and computational chemistry, characterized in that: include: Parameter processing module: obtains chromatographic experimental parameters and processes them to obtain experimental parameter vectors; Identifier processing module: obtains molecular identifiers and processes them to obtain multimodal molecular vectors; Vector merging module: merges the experimental parameter vector and the multimodal molecular vector to obtain the input vector; And, prediction module: based on the input vector, a machine learning regression model or an integrated learning method is used to predict the chromatographic elution curve information.

8. A computer storage medium storing a readable program, characterized in that: When the program is run, the chromatographic elution curve prediction method based on integrated learning and computational chemistry according to any one of claims 1 to 6 can be executed.

9. An electronic device, characterized in that: include: A processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the chromatographic elution curve prediction method based on integrated learning and computational chemistry according to any one of claims 1 to 6.

10. A computer program product comprising computer instructions, characterized in that The computer instructions instruct the computing device to execute operations corresponding to the chromatographic elution curve prediction method based on integrated learning and computational chemistry as described in any one of claims 1 to 6.