Eutectic ligand screening method and system

The method leverages a pre-built database and AI model to efficiently and accurately identify cocrystal partners, addressing inefficiencies in existing screening methods and enhancing drug development by improving the physicochemical properties of pharmaceuticals through cocrystal formation.

CN116206697BActive Publication Date: 2025-07-15BEIJING INSTITUTE OF PETROCHEMICAL TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310084082.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-18
Publication Date
2025-07-15
Estimated Expiration
2043-01-18

AI Technical Summary

Technical Problem

The existing eutectic ligand screening methods are not systematic and comprehensive enough, the screening efficiency is low, and the screening accuracy is poor, so it is impossible to accurately predict eutectic formation.

Method used

Based on the pre-constructed drug eutectic ligand database and expert system, combined with the pre-trained artificial intelligence model, ligand molecules matching the target molecule are screened through functional group matching and supramolecular synthesis sub-rules, and ligand molecules matching the target molecule are trained and verified using machine learning models.

Benefits of technology

It significantly improves the efficiency and accuracy of eutectic ligand screening, reduces screening costs, and can quickly recommend matching ligand molecules without experimental verification, improving the screening efficiency and accuracy of eutectic ligand molecules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116206697B_ABST
    Figure CN116206697B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a method and system for screening eutectic ligands. The method includes: obtaining ligand molecules that match the target molecular structure based on a pre-constructed drug eutectic ligand database and a pre-established expert system, and using the ligand molecules obtained from the drug eutectic ligand database as candidate ligand molecules; screening the candidate ligand molecules based on a pre-trained artificial intelligence model to obtain target ligand molecules; wherein, the drug eutectic ligand database is constructed using a large number of ligand molecules and their molecular properties; the expert system is established based on functional group matching and supramolecular synthon rules and can estimate the binding energy between the target molecule and the ligand molecule; the artificial intelligence model is trained using a machine learning model through eutectic screening examples reported in the literature. The present invention solves the problems of high cost, low screening efficiency, and insufficient systematic and comprehensive screening in the prior art for eutectic ligand screening.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a method and system for screening eutectic ligands. Background Art

[0002] A eutectic is a crystal formed by two or more compounds arranged in an ordered manner in the same crystal lattice with a certain stoichiometric ratio through non-covalent interactions such as hydrogen bonds, π-π interactions, and electrostatic interactions. In the field of chemical drugs, pharmaceutical cocrystals are mainly used to improve the physicochemical properties and biopharmaceutical properties of drugs, such as water solubility, dissolution rate, and bioavailability. Pharmaceutical cocrystals can significantly change the physical or pharmaceutical properties of the active pharmaceutical ingredient (API) without changing its chemical structure. For example, they can improve solubility and dissolution rate, and increase the melting point, which is of great significance for drug development.

[0003] With the development of the pharmaceutical cocrystal technology, how to quickly find suitable eutectic ligands is a key step in cocrystal research. In the existing technologies, the commonly used theoretical calculation methods for predicting and quickly screening pharmaceutical cocrystal ligands are as follows: 1. The virtual cocrystal design method based on molecular surface electrostatic potential (MEPs), which determines the possible intermolecular interaction sites on the molecular surface by calculating the molecular surface electrostatic potential, providing a reliable guiding principle for intermolecular recognition; 2. The design theory based on Hansen solubility parameters (HSP), which predicts the feasibility of forming a cocrystal by calculating the difference in Hansen solubility parameters between the active pharmaceutical ingredient and the eutectic ligand, and realizes the rapid screening of eutectic ligands; 3. The COSMO-RS calculation model based on liquid-phase thermodynamics theory, which uses the screened charge density calculated by first-principles calculation combined with fast statistical thermodynamics to estimate the intermolecular affinity in the solid, providing an efficient method for quickly screening eutectic ligands.

[0004] However, the virtual calculation method based on MEPs ignores the influence of molecular conformation on the distribution of extreme points on MEPs during the calculation process, and the deviation of the calculation results will have a greater impact on the success rate of predicting crystal formation; the HSP method is difficult to estimate strong interactions such as hydrogen bonds between different molecules; the COSMO-RS method can only partially predict the intermolecular affinity, and its accuracy needs to be improved. It cannot accurately predict the formation of a cocrystal and cannot guarantee obtaining a crystal. In fact, when predicting a cocrystal by COSMO, the final result after experimental screening may be a co-amorphous state, or just a physical mixture of these compounds without forming a cocrystal.

[0005] In summary, the existing screening methods are not systematic and comprehensive enough, with low screening efficiency and poor screening accuracy. Summary of the Invention

[0006] To this end, embodiments of the present invention provide a eutectic ligand screening method and system, which can at least partially solve the problems of being insufficiently systematic and comprehensive, having low screening efficiency, and poor screening accuracy in the prior art.

[0007] To achieve the above object, embodiments of the present invention provide the following technical solutions:

[0008] A eutectic ligand screening method, the method comprising:

[0009] Based on a pre-constructed drug eutectic ligand database and a pre-built expert system, obtain ligand molecules that match the target molecular structure, and use the ligand molecules obtained from the drug eutectic ligand database as candidate ligand molecules;

[0010] Based on a pre-trained artificial intelligence model, screen the candidate ligand molecules to obtain target ligand molecules;

[0011] Among them, the drug eutectic ligand database is constructed using a large number of ligand molecules and their molecular properties; the expert system is based on functional group matching and supramolecular synthon rules to match the target molecule, and uses molecular mechanics to calculate the binding energy between the target molecule and the ligand molecule, and sorts them according to the magnitude of the binding energy; the artificial intelligence model is trained using a machine learning model and various properties of the molecule as descriptors through eutectic screening examples reported in the literature.

[0012] In some embodiments, the ligand molecule properties in the drug eutectic ligand database at least include properties reported in the literature, properties obtained or predicted by quantum calculations, and properties estimated by common thermodynamic models.

[0013] In some embodiments, the properties reported in the literature at least include molecular structure, characteristic functional groups, dissociation constant values, and melting points;

[0014] The properties obtained or predicted by quantum calculations at least include molecular surface area and extreme points of the electrostatic potential surface;

[0015] The properties estimated by common thermodynamic models at least include density, cohesive energy, and melting point.

[0016] In some embodiments, constructing the drug eutectic ligand database using a large number of ligand molecules and their molecular properties specifically includes:

[0017] Obtain ligand molecules that meet preset conditions, where the preset conditions include that the selected ligand molecules exist in common foods or food additives and have been proven to have low toxicity and can be ingested in large quantities;

[0018] Quantify and optimize the structure of the ligand molecule to obtain its electron density wave function, and then calculate molecular properties such as molecular surface, molecular shape, molecular volume, and extreme points of molecular surface electrostatic potential; or use it to predict properties such as density and melting point;

[0019] Integrate the ligand molecule and its corresponding ligand molecule properties to form ligand molecule data;

[0020] Based on a large amount of ligand molecule data, form the drug co-crystal ligand database.

[0021] In some embodiments, build the expert system according to the ranking result of the binding possibility between the target molecule and the ligand molecule based on functional group matching and supramolecular synthon rules, specifically including:

[0022] Identify the structural characteristics of the target molecule, and obtain multiple ligand molecules that match in the drug co-crystal ligand database based on the structural characteristics;

[0023] Construct multiple supramolecular synthons according to the structures of the target molecule and all ligand molecules;

[0024] Optimize by molecular mechanics methods, and calculate the binding energy of each supramolecular synthon respectively to obtain multiple binding energies;

[0025] Determine the binding possibility between the target molecule and multiple ligand molecules based on the magnitude of the binding energy to obtain the ranking result of the binding possibility;

[0026] Generate the input file of the machine learning model based on the ranking result.

[0027] In some embodiments, identify the structural characteristics of the target molecule, and obtain multiple ligand molecules that match in the drug co-crystal ligand database based on the structural characteristics, specifically including:

[0028] Determine the functional groups of the target molecule, and screen out ligand molecules with functional groups corresponding to those of the target molecule in the drug co-crystal ligand database to obtain initially selected ligand molecules;

[0029] When a supramolecular structure that can form hydrogen bonds can be constructed between the target molecule and the initially selected ligand molecule, further screen the ligand molecules of the target molecule by calculating whether there is a supramolecular structure that can form strong hydrogen bonds, and exclude the initially selected ligand molecules with supramolecular structures that cannot form strong hydrogen bonds, and use the remaining initially selected ligand molecules as candidate ligand molecules.

[0030] In some embodiments, train the artificial intelligence model by using the machine learning model through the reported co-crystal screening examples, specifically including:

[0031] Obtain the eutectic screening examples reported in the literature, where the eutectic screening examples reported in the literature include positive samples capable of forming eutectics and negative samples incapable of forming eutectics;

[0032] Calculate the electron density wave function for the eutectic screening examples reported in the literature;

[0033] Calculate the molecular property samples based on the electron density wave function;

[0034] Integrate the eutectic screening examples reported in the literature and the molecular property samples into ligand molecule data samples, and form a data set. Split the data set into a training set, a test set, and a validation set;

[0035] Input the data set into a machine learning model, train and update the parameters in the model according to the gap between the prediction result of the training set and the true value of the training set, determine the optimal parameters according to the optimal result on the validation set, and detect the test set data;

[0036] According to the hyperparameter optimization strategy, determine the optimal hyperparameter combination of the model to obtain an optimized artificial intelligence model.

[0037] The present invention also provides a eutectic ligand screening system, and the system includes:

[0038] A ligand molecule acquisition unit, configured to obtain ligand molecules matching the target molecular structure based on a pre-constructed drug eutectic ligand database and a pre-established expert system, and use the ligand molecules obtained from the drug eutectic ligand database as candidate ligand molecules;

[0039] A ligand molecule screening unit, configured to screen the candidate ligand molecules based on a pre-trained artificial intelligence model to obtain target ligand molecules;

[0040] Wherein, the ligand database is constructed using a large number of ligand molecules and their properties; the expert system is built based on functional group matching and supramolecular synthon rules and can estimate the binding energy between the target molecule and the ligand molecule; the artificial intelligence model is obtained by training a machine learning model based on various properties of molecules as descriptors through the eutectic screening examples reported in the literature.

[0041] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above method are implemented.

[0042] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.

[0043] The eutectic ligand screening method and system provided by the present invention are based on a pre-constructed drug eutectic ligand database and a pre-established expert system to obtain ligand molecules that match the target molecular structure, and use the ligand molecules obtained from the drug eutectic ligand database as candidate ligand molecules; based on a pre-trained artificial intelligence model, the candidate ligand molecules are screened to obtain target ligand molecules; wherein, the ligand database is constructed using a large number of ligand molecules and their properties; the expert system is established based on functional group matching and supramolecular synthon rules and can estimate the binding energy between the target molecule and the ligand molecule; the artificial intelligence model is trained using a machine learning model through eutectic screening examples reported in the literature.

[0044] In this way, for the functional groups of the target molecule, the present invention can quickly recommend matching ligand molecules without experimental screening and experimental verification, significantly reducing the screening cost of ligand molecules, improving the screening efficiency and accuracy of eutectic ligand molecules, and being of great help in the fields of drug development, fine chemicals such as pigments, etc., where the physicochemical properties of compounds need to be changed by forming eutectics. The present invention solves the problems of high screening cost, low screening efficiency and lack of systematic and comprehensive screening in the prior art for eutectic ligand screening. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can be obtained according to the provided drawings.

[0046] The structures, proportions, sizes, etc. shown in this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the limited conditions for the implementation of the present invention. Therefore, they do not have a substantial technical meaning. Any modification of the structure, change of the proportional relationship or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.

[0047] Figure 1 One of the flowcharts of the eutectic ligand screening method provided by the present invention;

[0048] Figure 2 Another flowchart of the eutectic ligand screening method provided by the present invention;

[0049] Figure 3 Another flowchart of the eutectic ligand screening method provided by the present invention;

[0050] Figure 4 The fourth flowchart of the eutectic ligand screening method provided by the present invention;

[0051] Figure 5 The diagram of the matching mode of functional groups with higher hydrogen bond energy in the expert system provided by the present invention;

[0052] Figure 6 The diagram of the eutectic hydrogen bond functional groups and their occurrence probabilities in the reference CCDC provided by the present invention;

[0053] Figure 7 The supramolecular structure diagram of the strong hydrogen bond existing between the paracetamol molecule and the ligand molecule screened by the expert system provided by the present invention; wherein, Figure 7 a is combined with nicotinamide; b is combined with glycine;

[0054] Figure 8 The structural schematic diagram of the eutectic ligand screening system provided by the present invention;

[0055] Figure 9 The structural block diagram of a computer device provided by the present invention. Detailed implementation manners

[0056] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0057] Please refer to Figure 1 , the eutectic ligand screening method provided by the present invention includes the following steps:

[0058] S110: Based on a pre-constructed drug eutectic ligand database and a pre-built expert system, obtain ligand molecules that match the target molecular structure, and use the ligand molecules obtained from the drug eutectic ligand database as candidate ligand molecules;

[0059] S120: Based on a pre-trained artificial intelligence model, screen the candidate ligand molecules to obtain target ligand molecules;

[0060] Wherein, the ligand database is constructed using a large number of ligand molecules and their properties; the expert system is built based on functional group matching and supramolecular synthon rules and can estimate the binding energy between the target molecule and the ligand molecule; the artificial intelligence model is obtained by training a machine learning model through eutectic screening examples reported in the literature.

[0061] Specifically, the eutectic ligand screening method provided by the present invention is an intelligent eutectic ligand screening method based on a drug eutectic ligand database and an intelligent eutectic screening system. Before screening, a drug eutectic ligand database and an expert system are pre-constructed, and an artificial intelligence model is pre-trained. During screening, ligand molecular structures are obtained through the pre-established drug eutectic ligand database. Among them, the molecular structures are optimized, ligand molecular properties are collected and calculated, and a ligand database suitable for drug eutectic prediction is constructed. This drug eutectic ligand database can be expanded according to later needs. Then, through the pre-established expert system, the functional groups of the target molecule are automatically identified, and according to the special functional group combinations for forming eutectics, the drug eutectic ligand database is connected, and ligand molecules are automatically recommended for the target molecule according to the possibility of forming eutectics. Subsequently, through the pre-trained artificial intelligence model, the possibility of the target molecule and the ligand molecule recommended by the expert system forming a eutectic is further judged; thus, the prediction model and the expert system are combined to predict the eutectic of the target molecule and the ligand molecule recommended by the expert system, improve the effectiveness of ligand recommendation, and form an efficient eutectic prediction system.

[0062] With the continuous cross-fertilization and penetration between disciplines, machine learning algorithms have gradually been combined with drug eutectic screening and become a new research hotspot. Machine learning is a multi-disciplinary cross-professional field that covers knowledge of probability theory, statistics, approximation theory, and complex algorithms. It uses a computer as a tool and is committed to real-time simulation of the human learning method, and divides the existing content into knowledge structures to effectively improve learning efficiency. The eutectic ligand screening method provided by the present invention solves the problems of insufficient features relied on by the model screening when using machine learning to screen ligands, being restricted in the generalization of intermolecular interactions and the diversity of molecular chemical structures, having no suitable ligand database, not considering the toxicity of ligand molecules, and the eutectics obtained by screening may have high toxicity and low practical value.

[0063] The eutectic ligand screening method provided by the present invention introduces prior knowledge as a feature descriptor on the basis of machine learning, especially molecular properties related to intermolecular forces. The information collected and calculated is different and complementary, making up for the deficiencies of previous machine learning in intermolecular interactions in eutectic ligand screening, and the accuracy of screening eutectic ligands is relatively high. At the same time, the present invention constructs a drug eutectic ligand database, which provides low-toxic or non-toxic ligand molecules and their physical property data for the eutectic screening system. Thus, it has higher practical value, can reduce the cost of drug eutectic ligand screening experiments, and improve the efficiency of drug eutectic ligand screening.

[0064] In some embodiments, as Figure 2 shown, the drug eutectic ligand database is constructed using a large number of ligand molecules and ligand molecular properties, specifically including the following steps:

[0065] S210: Obtain ligand molecules that meet the preset conditions. The preset conditions include that the molecular toxicity is lower than the toxicity threshold and the molecule is inactive. That is to say, the selected ligand molecules are derived from common foods or food additives and have been proven to have low toxicity and can be ingested in large quantities. The collected ligand molecules are non-toxic or have low toxicity and are suitable for the development of drug co-crystals.

[0066] S220: Perform quantum chemical calculations on the ligand molecules and optimize their structures to obtain the electron density wave functions of the ligand molecules. Specifically, perform quantum chemical calculations on the ligand molecules and optimize their structures to obtain their electron density wave functions, and then calculate molecular properties such as molecular surface, molecular shape, molecular volume, and extreme points of the molecular surface electrostatic potential (referred to as the electrostatic potential surface).

[0067] S230: Based on the electron density wave function and other molecular properties, such as density and melting point. Specifically, the properties of the ligand molecules in the drug co-crystal ligand database at least include the properties reported in the literature, the properties obtained or predicted through quantum chemical calculations, and the properties estimated through common thermodynamic models. Among them, the properties reported in the literature at least include molecular structure, characteristic functional groups, dissociation constant values, and melting points. The properties obtained through quantum chemical calculations at least include molecular surface properties and charge distributions. The properties estimated through common thermodynamic models at least include molecular surface area and extreme points of the molecular surface electrostatic potential. The predicted molecular properties at least include density, cohesive energy, and predicted melting point. That is to say, this database contains various properties of molecules, including but not limited to molecular structure, characteristic functional groups, pKa (dissociation constant value), melting point, etc., which may affect the formation of co-crystals. This database also includes the properties obtained or predicted through quantum chemical calculations of molecules, including but not limited to calculating the molecular surface area and extreme points of the electrostatic potential surface based on the results of quantum chemical calculations, and using them to predict the molecular properties that have not been reported, such as density, cohesive energy, melting point, etc.

[0068] S240: Integrate the ligand molecules and their corresponding ligand molecule properties to form ligand molecule data.

[0069] S250: Based on a large amount of ligand molecule data, form the drug co-crystal ligand database.

[0070] Specifically, when constructing a drug co-crystal ligand database suitable for drug co-crystals, for ligand molecules, collect those with low toxicity and no activity, which are suitable for the development of medicinal co-crystals, and ensure that the ligand molecules included in the built database have low or no toxicity and no activity.

[0071] In step S220, the quantization calculation can be performed in ORCA 4.2. First, the structure of the molecule is optimized using the PBEh-3c function and the def2-SVP basis set to obtain the electron density wave function (wfn file) of the molecule. It can also be calculated in other software such as Gaussian, or other density functionals / basis sets can be used. Then, the wfn file is used as the input file for the Multiwfn software to collect and calculate the molecular properties related to the intermolecular interactions for eutectic prediction, and the molecular properties related to the intermolecular and intramolecular force results for solubility prediction. These properties include but are not limited to: boiling point, critical temperature, triple point temperature, enthalpy of fusion, dipole moment, molecular mass, Cp difference between the liquid and solid phases at the triple point, enthalpy of vaporization at the boiling point under temperature conditions, freezing point, Pitzer acentric factor, critical pressure, molar volume at the boiling point, critical volume, critical compressibility factor, molecular volume, density estimated based on mass and volume, pKa value, minimum value of the electrostatic potential surface, maximum value of the electrostatic potential surface, total surface area of the molecular electrostatic potential surface, surface area of the positive region of the molecular electrostatic potential surface, surface area of the negative region of the molecular electrostatic potential surface, overall average value of the electrostatic potential surface, negative average value of the electrostatic potential surface, overall change rate, positive difference, negative difference, charge balance index, νσ^2tot, internal charge separation energy, molecular polarity index, non-polar surface area, polar surface area. Finally, the collected and calculated ligand molecule data are integrated to form a database of more than 280 drug co-crystal ligands with low toxicity and no activity, including maleic acid and citric acid. This ligand database can be increased or decreased according to needs later.

[0072] In some embodiments, as Figure 3 shown, the expert system is built based on the ranking results of the binding possibilities of the target molecule and the ligand molecule according to the functional group matching and supramolecular synthon rules, specifically including the following steps:

[0073] S310: Identify the structural features of the target molecule, and obtain multiple ligand molecules that match based on the structural features in the drug co-crystal ligand database. Among them, the structural features of the target molecule are automatically identified, such as functional groups, hydrogen bond donor and acceptor capabilities, etc., and appropriate ligand molecules in the drug co-crystal ligand database are automatically matched through functional group matching methods, pKa and other rules.

[0074] Specifically, determine the functional groups of the target molecule, and screen out ligand molecules with functional groups corresponding to those of the target molecule in the drug cocrystal ligand database to obtain preliminary selected ligand molecules; when a supramolecular structure capable of forming hydrogen bonds can be constructed between the target molecule and the preliminary selected ligand molecules, further screen the ligand molecules of the target molecule by calculating whether there is a supramolecular structure capable of forming strong hydrogen bonds, and exclude the preliminary selected ligand molecules that cannot form a supramolecular structure with strong hydrogen bonds, and use the remaining preliminary selected ligand molecules as candidate ligand molecules.

[0075] S320: Construct multiple supramolecular synthons according to the structures of the target molecule and all ligand molecules;

[0076] S330: Optimize by molecular mechanics methods and calculate the binding energies of the respective supramolecular synthons to obtain multiple binding energies;

[0077] S340: Determine the binding possibilities between the target molecule and multiple ligand molecules based on the magnitudes of the binding energies to obtain a ranking result of the binding possibilities; automatically construct supramolecular synthons according to the structures of the target molecule and the ligand molecules, optimize by molecular mechanics methods and calculate the binding energies, and rank the ligand molecules according to the possibilities based on the magnitudes of the binding energies.

[0078] S350: Generate an input data set for the machine learning model based on the ranking result.

[0079] Specifically, when building an expert system for automatically identifying the functional groups of a target molecule and automatically recommending ligand molecules, the functional group library can be loaded through the RDKit plugin in the python library; set a molecular fragment virtualizer to slice the molecular formula (smiles format), and then judge the functional groups existing in the target molecule by referring to the functional group library of the RDKit plugin; according to the special functional group combinations for forming cocrystals, connect to the drug cocrystal ligand database constructed in the present invention to screen out ligand molecules with functional groups corresponding to those of the target molecule. Finally, verify whether the target molecule and the recommended ligand molecules can construct a supramolecular structure capable of forming hydrogen bonds, and further screen the ligand molecules of the target molecule by calculating whether there is a supramolecular structure capable of forming strong hydrogen bonds, and exclude the ligand molecules that cannot form such a supramolecular structure.

[0080] To further improve the screening success rate, the artificial intelligence model can be combined with the expert system, as Figure 4 shown, the artificial intelligence model is obtained by training the machine learning model through the cocrystal screening examples reported in the literature, and specifically includes the following steps:

[0081] S410: Obtain the cocrystal screening examples reported in the literature, and the cocrystal screening examples reported in the literature include positive samples capable of forming cocrystals and negative samples incapable of forming cocrystals;

[0082] S420: Calculate the electron density wave function for the eutectic screening examples reported in the literature;

[0083] S430: Calculate the molecular property samples based on the electron density wave function;

[0084] S440: Integrate the eutectic screening examples reported in the literature and the molecular property samples into a ligand molecular data sample, and form a data set. Split the data set into a training set, a test set, and a validation set;

[0085] S450: Input the data set into a machine learning model. Train and update the parameters in the model according to the gap between the predicted results of the training set and the true values of the training set. Determine the optimal parameters according to the optimal results on the validation set, and detect the test set data; The machine learning model preferably selects the Ridge model.

[0086] S460: Determine the optimal hyperparameter combination of the model according to the hyperparameter optimization strategy to obtain an optimized molecular screening model.

[0087] Specifically, as shown in Table 1, the molecular screening model provided by the present invention selects the following molecular properties as descriptors to achieve eutectic ligand screening: including but not limited to boiling point, triple point temperature, dipole moment, molecular mass, molecular volume, difference in liquid and solid phase heat capacities (Cp) at three points, enthalpy of vaporization at the boiling point temperature condition, freezing point, Pizer acentric factor, critical pressure, pKa value, total surface area of the molecular electrostatic potential surface, surface area of the positive region of the molecular electrostatic potential surface, surface area of the negative region of the molecular electrostatic potential surface, overall average value of the electrostatic potential surface, negative average value of the electrostatic potential surface, overall change rate of the electrostatic potential surface, positive difference of the electrostatic potential surface, charge balance index, νσ^2tot (variance σ2tot of the total electrostatic potential on the molecular surface multiplied by the charge balance degree ν), critical compression factor, solid density.

[0088] Table 1 Descriptors selected by the present invention and their descriptions

[0089]

[0090] In a specific usage scenario, building an eutectic prediction model to further screen the eutectic ligands of target molecules specifically includes:

[0091] First, construct an artificial intelligence model suitable for drug eutectic prediction, and a total of 8 artificial intelligence models based on different algorithms are built;

[0092] According to the requirements of eutectic ligand prediction, collect positive samples that can form eutectics and negative samples that cannot form eutectics according to the quantity ratio of 1:1;

[0093] Perform quantum calculations in ORCA 4.2. First, optimize the structure of the molecule using the PBEh-3c function and the def2-SVP basis set to obtain the electronic density wave function (wfn file) of the molecule.

[0094] Use the wfn file as the input file for the Multiwfn software to calculate the molecular properties related to the intermolecular interaction results for eutectic prediction. These properties include, but are not limited to: boiling point, triple point temperature, dipole moment, molecular mass, molecular volume, the difference in liquid and solid Cp at three points (how many times the specification range of the data is the fluctuation range), enthalpy of vaporization at the boiling point under temperature conditions, freezing point, Pitzer acentric factor, critical pressure, pKa value, total surface area of the molecular electrostatic potential surface, surface area of the positive region of the molecular electrostatic potential surface, surface area of the negative region of the molecular electrostatic potential surface, overall average value of the electrostatic potential surface, negative average value of the electrostatic potential surface, overall change rate, positive difference, charge balance index, νσ^2tot (the variance σ2tot of the total electrostatic potential on the molecular surface multiplied by the charge balance degree ν), critical compression factor, solid density.

[0095] Integrate the molecular property data obtained from collection and quantum calculations, and select a specific method to split the data set into a training set, a test set, and a validation set.

[0096] Input the data set into the machine learning model, train and update the parameters in the model according to the gap between the prediction results of the training set and the true values of the training set, determine the optimal parameters according to the best results on the validation set, and detect the test set data.

[0097] Determine the optimal hyperparameter combination of the model according to the hyperparameter optimization strategy, and select the model with the best prediction effect.

[0098] The following uses a specific usage scenario to illustrate the implementation process of the eutectic ligand screening method provided by the present invention.

[0099] Example 1

[0100] In this example, taking the screening of ligand molecules for paracetamol molecules as an example, for the establishment of the drug eutectic ligand database, the specific steps for constructing the database are as follows:

[0101] 1.1 Select molecules with low or no toxicity from the Substances Added to Food (formerly EAFUS) database to form a ligand molecule data set, and a total of more than 280 ligand molecules are obtained.

[0102] 1.2 Optimize the ligand molecules using the Avagadro software under the MMFF94 force field.

[0103] 1.3 Use the optimized molecular structure as the input file, and perform quantum calculations in ORCA 4.2. First, optimize the structure of the molecule using the PBEh-3c function and the def2-SVP basis set to obtain the electron density wave function (wfn file) of the molecule;

[0104] 1.4 Use the wfn file as the input file for the Multiwfn software, and collect and calculate molecular properties related to eutectic prediction, including but not limited to: boiling point, triple point temperature, dipole moment, molecular mass, molecular volume, difference in liquid and solid Cp at three points (how many times the specification range of the data is the fluctuation range), enthalpy of vaporization at the boiling point under temperature conditions, freezing point, Pizer acentric factor, critical pressure, pKa value, total surface area of the molecular electrostatic potential surface, surface area of the positive region of the molecular electrostatic potential surface, surface area of the negative region of the molecular electrostatic potential surface, overall average value of the electrostatic potential surface, negative average value of the electrostatic potential surface, overall change rate, positive difference, charge balance index, νσ^2tot (variance σ2tot of the total electrostatic potential on the molecular surface multiplied by the charge balance degree ν), critical compression factor, solid density;

[0105] 1.5 Integrate the collected and calculated molecular data to form a drug eutectic ligand database.

[0106] When building the expert system, there are the following specific steps:

[0107] 2.1 Set the input format of the expert system to the SMILES format, and load the functional group library through the RDKit plugin in the python library;

[0108] 2.2 Set up a molecular fragment virtualizer, slice the molecular formula in the SMILES format of the target molecule, and judge the existing functional groups by referring to the functional group library of the RDKit plugin;

[0109] 2.3 Organize and store the identified functional groups in an array according to their numbers in the functional group library of the RDKit plugin;

[0110] 2.4 According to the special functional group combinations for forming eutectics, connect to the drug eutectic ligand database constructed in the present invention, and screen out ligand molecules with functional groups corresponding to the functional groups of the target molecule, such as Figure 5 and Figure 6 as shown;

[0111] 2.5 Put the target molecule and the recommended ligand molecules into ABclouster to verify whether a supramolecular structure capable of forming hydrogen bonds can be constructed. Further screen the ligand molecules of the target molecule by calculating whether there is a supramolecular structure capable of forming strong hydrogen bonds, and exclude the ligand molecules that cannot form such a supramolecular structure. Strong hydrogen bond supramolecules are such as Figure 7 as shown.

[0112] When constructing an artificial intelligence model suitable for drug co-crystal ligand screening, there are the following specific steps:

[0113] 3.1 Build a co-crystal prediction model and select 8 common artificial intelligence classification models: BGR (Bagging model), RFR (Random Forest model), DTR (Decision Tree model), LR (Linear model), Ridge (Ridge model), KNR (K-Nearest Neighbor model), SVR (Support Vector Machine model), and MLPR (Multi-Layer Perceptron model). Specifically, in this embodiment, 5 loss functions commonly used in artificial intelligence models are set for evaluating the prediction effect of the model, namely MAE, MSE, MAPE, RMSE, and R^2 function;

[0114] 3.2 According to the requirements of model construction, query CCDC and literature to collect positive samples that can form co-crystals and negative samples that cannot form co-crystals according to the quantity ratio of 1:1; After optimizing the positive sample molecules and negative sample molecules by molecular mechanics and quantum mechanics, use them as inputs and perform quantum chemical calculations in ORCA 4.2. First, optimize the molecular structure using the PBEh-3c function and def2-SVP basis set to obtain the electronic density wave function (wfn file) of the molecule;

[0115] 3.3 Use the wfn file as the input file of the Multiwfn software to calculate the molecular properties related to co-crystal prediction that affect intermolecular interactions: minimum value of the electrostatic potential surface, maximum value of the electrostatic potential surface, total surface area of the molecular electrostatic potential surface, surface area of the positive region of the molecular electrostatic potential surface, surface area of the negative region of the molecular electrostatic potential surface, overall average value of the electrostatic potential surface, positive average value of the electrostatic potential surface, negative average value of the electrostatic potential surface, overall change rate, positive difference, negative difference, charge balance index, νσ^2tot, internal charge separation energy, molecular polarity index, non-polar surface area, polar surface area;

[0116] 3.4 Collect the results of intermolecular interactions related to co-crystal prediction: molecular pKa value, boiling point, critical temperature, triple point temperature, enthalpy of fusion, dipole moment, molecular mass, Cp difference between the liquid and solid phases at the triple point, enthalpy of vaporization at the boiling point under temperature conditions, freezing point, Pitzer acentric factor, critical pressure, critical compressibility factor, molecular volume;

[0117] 3.5 Integrate the collected and calculated molecular property data to form a dataset for training the model, select the random splitting method, and randomly divide it into three sets according to the ratio of the training set, test set, and validation set of 8, 1, 1;

[0118] 3.6 Input the divided data set into the artificial intelligence model. First, analyze the correlation of descriptors, find the descriptors with strong correlation (correlation > 0.8), and then apply them to prediction after screening. Specifically, in this embodiment, there are 26 descriptors with strong correlation (as shown in Table 2), including but not limited to: melting point, boiling point, triple point temperature, enthalpy of fusion, etc.

[0119] Table 2 Descriptors finally selected by the optimal model (Ridge model) of the present invention and their descriptions

[0120]

[0121] 3.6 Exclude the 26 descriptors with strong correlation one by one before model optimization, record the change of the model loss function when each descriptor is deleted, and delete the descriptor that produces the best goodness change after deletion when a round of loop is completed. On this basis, continue to carry out the next round of descriptor screening until the screening stops when no matter which descriptor is deleted, there is no good change. Specifically, in this embodiment, the good change of the model is: the smaller the values of the four loss functions of MAE, MSE, MAPE, and RMSE, the better the model prediction effect, and the closer R^2 is to 1, the better the model effect.

[0122] 3.7 According to the gap between the prediction result of the training set and the true value of the training set, use the iterative update method to optimize the model parameters to train and update the parameters in the artificial intelligence model, determine the optimal model parameters according to the best result on the validation set, detect the data of the test set to determine the optimal hyperparameter combination of the model, and select the prediction model with the best prediction effect. The comparison of model effects is shown in Table 3.

[0123] Table 3 Comparison chart of prediction accuracies of 8 artificial intelligence models constructed by the present invention

[0124] Model Name Prediction Accuracy Ridge Regression Model 0.973 Random Forest Model 0.766 Bagging Model 0.753 Multi-Layer Perceptron Model 0.691 Support Vector Machine Model 0.642 Decision Tree Model 0.591 K-Nearest Neighbor Model 0.562 Linear Model 0.501

[0125] Combine the expert system with the artificial intelligence model to form a eutectic screening system. Based on the drug eutectic ligand database built by the present invention, screen the eutectic ligand molecules for paracetamol molecules. Specifically, there are the following steps:

[0126] 1. Convert the paracetamol molecular structure into the SMILES format and input for query.

[0127] 2. Automatically judge that the paracetamol molecule contains phenolic hydroxyl and amide bonds, which is consistent with the functional groups actually contained in the molecule.

[0128] 3. Through functional group matching by an expert system, 17 ligand molecules that may form cocrystals with paracetamol molecules were automatically recommended from the drug cocrystal ligand database built in step 1, such as nicotinamide, saccharin, glycine, and methyl p-hydroxybenzoate. See Table 4 for details;

[0129] Table 4 Results of ligand molecules recommended by the expert system for paracetamol molecules based on the drug cocrystal ligand database

[0130] Ligand Molecule Nicotinamide Saccharin Glycine Methyl p-Hydroxybenzoate Neotame Cyclohexylsulfamic Acid Vanillic Acid Oxalic Acid Citric Acid Maleic Acid Fumaric Acid Malonic Acid 4-Aminobenzoic Acid Malic Acid Succinic Acid Urea Benzoic Acid

[0131] 4. Verify whether there are strong hydrogen bonds in the supramolecules formed by paracetamol molecules and the ligand molecules screened by the expert system. After verification, paracetamol molecules can form supramolecular structures with strong hydrogen bonds with the 17 ligand molecules screened by the expert system, as Figure 7 shown;

[0132] 5. Perform quantum chemical calculations on paracetamol molecules in ORCA 4.2. First, optimize the structure of the molecule using the PBEh-3c function and the def2-SVP basis set to obtain the electronic density wave function (wfn file) of the molecule;

[0133] 6. Use the wfn file of the paracetamol molecule as the input file for the Multiwfn software, and calculate the molecular properties related to cocrystal prediction that affect intermolecular interactions: minimum value of the electrostatic potential surface, maximum value of the electrostatic potential surface, total surface area of the molecular electrostatic potential surface, positive surface area of the molecular electrostatic potential, negative surface area of the molecular electrostatic potential, overall average value of the molecular electrostatic potential surface, positive average value of the molecular electrostatic potential surface, negative average value of the molecular electrostatic potential surface, overall change rate, positive difference, negative difference, charge balance index, νσ^2tot, internal charge separation energy, molecular polarity index, non-polar surface area, polar surface area;

[0134] 7. Collect the results of intermolecular interactions related to cocrystal prediction: molecular pKa value, boiling point, critical temperature, triple point temperature, enthalpy of fusion, dipole moment, molecular mass, Cp difference between the liquid and solid phases at the triple point, enthalpy of vaporization at the boiling point under temperature conditions, freezing point, Pitzer acentric factor, critical pressure, standard specific gravity at 60 °C, molar volume at the boiling point, critical volume, standard liquid molar volume at 60 °C, critical compression factor, molecular volume, density estimated based on mass and volume;

[0135] 8. Query and return the ligand molecule data of paracetamol recommended by the expert system of the present invention from the drug cocrystal ligand database established by the present invention, and integrate it with the molecular property data of paracetamol obtained by calculation and collection to form the input data of the cocrystal prediction model;

[0136] 9. Predicted by the cocrystal prediction model of the present invention, 15 ligand molecules including glycine and p-Hydroxybenzoic acid are excluded from the recommended queue of paracetamol cocrystal ligand molecules, and the prediction results are shown in Table 5;

[0137] Table 5 Results of further screening of ligand molecules screened by the expert system by the artificial intelligence model

[0138]

[0139] 10. At this time, based on the drug cocrystal ligand database established by the present invention, the intelligent cocrystal screening system of the present invention recommends ligands for paracetamol molecules: nicotinamide and saccharin;

[0140] The intelligent cocrystal ligand screening system based on the drug cocrystal ligand database of the present invention can more accurately judge whether a ligand molecule can form a cocrystal with a target molecule. Taking the paracetamol molecule as an example, the existing high-accuracy prediction model, the deep learning cocrystal prediction model based on graph neural network (CCGNet), judges neotame, cyclamic acid, and vanillic acid as capable of forming a cocrystal with the paracetamol molecule, but these three molecules actually cannot form a cocrystal with the paracetamol molecule. The intelligent cocrystal ligand screening system of the present invention accurately judges that these three molecules cannot form a cocrystal with the paracetamol molecule. Compared with the existing cocrystal prediction model, the intelligent cocrystal ligand screening system of the present invention has higher accuracy.

[0141] The present application has established a drug co-crystal ligand database, which includes more than 280 low-toxic or non-toxic and inactive ligand molecules, including the molecular properties of these ligand molecules related to co-crystal ligand screening, providing a reliable and more convenient data source for drug co-crystal research; the artificial intelligence model for co-crystal ligand screening of the present invention collects and calculates the molecular properties that affect intermolecular forces and the results of intermolecular forces. Compared with the current co-crystal prediction method based on machine learning, it solves the defect that the prediction model lacks important features in the data set; compared with the method of screening co-crystal ligands based on traditional machine learning, the present application integrates an expert system and an artificial intelligence model. First, the ligand molecules that may form co-crystals with the target molecule are reduced from more than 280 (the number of ligand molecules currently recorded in the drug co-crystal ligand database of the present invention) to dozens or even a dozen through the expert system, and then further screened through the artificial intelligence model to increase the prediction accuracy of the target molecule co-crystal ligand to 97%, thereby improving the prediction accuracy.

[0142] In the above specific embodiments, the co-crystal ligand screening method provided by the present invention is based on a pre-constructed drug co-crystal ligand database and a pre-constructed expert system to obtain ligand molecules that match the target molecular structure, and use the ligand molecules obtained in the ligand database as candidate ligand molecules; based on a pre-trained artificial intelligence model, the candidate ligand molecules are screened to obtain the target ligand molecules; wherein the drug co-crystal ligand database is constructed using a large number of ligand molecules and their molecular properties; the expert system is built based on functional group matching and supramolecular synthon rules, and can estimate the binding energy between the target molecule and the ligand molecule; the artificial intelligence model is obtained by training the co-crystal screening examples reported in the literature using a machine learning model. In this way, the present invention can quickly recommend matching ligand molecules for the structure of the target molecule without the need for experimental screening and experimental verification, significantly reducing the screening cost of ligand molecules, and improving the screening efficiency and accuracy of co-crystal ligand molecules, which is of great help in the field of fine chemicals such as drug development and pigments when it is necessary to change the physical and chemical properties of compounds by forming co-crystals. The present invention solves the problems of high screening cost, low screening efficiency and insufficient systematic and comprehensive screening of co-crystal ligands in the prior art.

[0143] In addition to the above method, the present invention also provides a co-crystal ligand screening system, such as Figure 8 As shown, the system comprises:

[0144] The ligand molecule acquisition unit 801 is used to acquire a ligand molecule matching the target molecular structure based on a pre-built drug co-crystal ligand database and a pre-built expert system, and use the ligand molecule obtained in the drug co-crystal ligand database as a candidate ligand molecule;

[0145] The ligand molecule screening unit 802 is configured to screen the candidate ligand molecules based on a pre-trained artificial intelligence model to obtain target ligand molecules;

[0146] Wherein, the pharmaceutical co-crystal ligand database is constructed using a large number of ligand molecules and their molecular properties; the expert system is built based on functional group matching and supramolecular synthon rules and can estimate the binding energy between the target molecule and the ligand molecule; the artificial intelligence model is obtained by training a machine learning model with co-crystal screening examples reported in the literature.

[0147] In some embodiments, the molecular properties of the ligand molecules in the pharmaceutical co-crystal ligand database at least include properties reported in the literature, properties obtained or predicted by quantum calculations, and properties estimated by common thermodynamic models.

[0148] In some embodiments, the properties reported in the literature at least include molecular structure, characteristic functional groups, dissociation constant values, and melting points;

[0149] The properties obtained or predicted by quantum calculations at least include molecular surface area and extreme points of the electrostatic potential surface;

[0150] The properties estimated by common thermodynamic models at least include density, cohesive energy, and predicted melting point.

[0151] In some embodiments, constructing the pharmaceutical co-crystal ligand database using a large number of ligand molecules and ligand molecular properties specifically includes:

[0152] The selected ligand molecules are present in common foods or food additives and have been proven to have low toxicity and can be ingested in large quantities;

[0153] Performing quantum calculations on the ligand molecules and optimizing their structures to obtain their electron density wave functions, and then calculating molecular properties such as molecular surface, molecular shape, molecular volume, and extreme points of the molecular surface electrostatic potential;

[0154] Integrating the ligand molecules and the corresponding ligand molecular properties to form ligand molecule data;

[0155] Based on a large amount of ligand molecule data, the pharmaceutical co-crystal ligand database is formed, and the molecules in the database can be added or deleted as needed.

[0156] In some embodiments, building the expert system based on the ranking results of the binding possibilities between the target molecule and the ligand molecule according to functional group matching and supramolecular synthon rules specifically includes:

[0157] Identifying the structural features of the target molecule and obtaining multiple matching ligand molecules in the pharmaceutical co-crystal ligand database based on the structural features;

[0158] Construct multiple supramolecular synthons according to the structures of the target molecule and all ligand molecules;

[0159] Optimize by molecular mechanics methods and calculate the binding energies of each supramolecular synthon respectively to obtain multiple binding energies;

[0160] Determine the binding possibilities between the target molecule and multiple ligand molecules based on the magnitudes of the binding energies to obtain a ranking result of the binding possibilities;

[0161] Generate an input data set for the machine learning model based on the ranking result.

[0162] In some embodiments, identify the structural features of the target molecule and obtain multiple ligand molecules that match in the drug co-crystal ligand database based on the structural features, specifically including:

[0163] Determine the functional groups of the target molecule, and screen out ligand molecules with functional groups corresponding to those of the target molecule in the drug co-crystal ligand database to obtain initially selected ligand molecules;

[0164] In the case where the target molecule and the initially selected ligand molecules can construct a supramolecular structure capable of forming hydrogen bonds, further screen the ligand molecules of the target molecule by calculating whether there is a supramolecular structure capable of forming strong hydrogen bonds, exclude the initially selected ligand molecules that cannot form a supramolecular structure with strong hydrogen bonds, and use the remaining initially selected ligand molecules as candidate ligand molecules.

[0165] In some embodiments, the artificial intelligence model is obtained by training the machine learning model with the co-crystal screening examples reported in the literature, specifically including:

[0166] Obtain the co-crystal screening examples reported in the literature, where the co-crystal screening examples reported in the literature include positive samples capable of forming co-crystals and negative samples not capable of forming co-crystals;

[0167] Calculate the electron density wave function of the co-crystal screening examples reported in the literature;

[0168] Calculate the sample molecular properties based on the electron density wave function;

[0169] Integrate the co-crystal screening examples reported in the literature and the molecular property samples into ligand molecule data samples, and form a data set. Split the data set into a training set, a test set, and a validation set;

[0170] Input the data set into the machine learning model, train and update the parameters in the model according to the gap between the predicted result of the training set and the true value of the training set, determine the optimal parameters according to the optimal result on the validation set, and detect the test set data;

[0171] According to the hyperparameter optimization strategy, determine the optimal combination of hyperparameters for the model to obtain an optimized artificial intelligence model.

[0172] In the above specific embodiments, the eutectic ligand screening system provided by the present invention is based on a pre-constructed drug eutectic ligand database and a pre-established expert system, obtains ligand molecules that match the target molecular structure, and uses the ligand molecules obtained from the drug eutectic ligand database as candidate ligand molecules; based on a pre-trained artificial intelligence model, screen the candidate ligand molecules to obtain target ligand molecules; wherein, the drug eutectic ligand database is constructed using a large number of ligand molecules and their molecular properties; the expert system is built based on functional group matching and supramolecular synthon rules and can estimate the binding energy between the target molecule and the ligand molecule; the artificial intelligence model is obtained by training a machine learning model through eutectic screening examples reported in the literature.

[0173] In this way, the present invention can quickly recommend matching ligand molecules for the structure of the target molecule without the need for experimental screening and experimental verification, significantly reducing the screening cost of ligand molecules, improving the screening efficiency and accuracy of eutectic ligand molecules, and being of great help in the fields of drug development, pigments and other fine chemical industries where it is necessary to change the physicochemical properties of compounds by forming eutectics. It solves the problems of high screening cost, low screening efficiency and incomplete systematic screening in the prior art for eutectic ligand screening.

[0174] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 9 shown. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and model predictions. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The model predictions of the computer device are used to store static information and dynamic information data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements the steps in the above method embodiments.

[0175] Those skilled in the art can understand that Figure 9 the structure shown in

[0176] Corresponding to the above embodiments, an embodiment of the present invention further provides a computer storage medium, which contains one or more program instructions. Among them, the one or more program instructions are used to be executed by a weight verification system to perform the method as described above.

[0177] The present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the above method.

[0178] In an embodiment of the present invention, the processor may be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0179] It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. The processor reads the information in the storage medium and combines its hardware to complete the steps of the above method.

[0180] The storage medium can be a memory, for example, it can be a volatile memory or a non-volatile memory, or it can include both volatile and non-volatile memories.

[0181] Among them, the non-volatile memory can be a read-only memory (ROM for short), a programmable read-only memory (PROM for short), an erasable programmable read-only memory (EPROM for short), an electrically erasable programmable read-only memory (EEPROM for short), or a flash memory.

[0182] The volatile memory may be a Random Access Memory (RAM) which serves as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DRRAM).

[0183] The storage media described in the embodiments of the present invention are intended to include, but are not limited to, these and any other suitable types of memory.

[0184] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the present invention can be implemented by a combination of hardware and software. When applying software, the corresponding functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes computer storage media and communication media, where the communication media includes any medium that facilitates the transfer of a computer program from one place to another. The storage media can be any available medium accessible by a general or special purpose computer.

[0185] The above specific embodiments further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solution of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for screening eutectic ligands, characterized in that, The method includes: Based on a pre-constructed drug co-crystal ligand database and an expert system for co-crystal screening based on the ligand database, obtaining ligand molecules that match the target molecular structure, and using the ligand molecules obtained from the ligand database as candidate ligand molecules; Based on a pre-trained molecular screening model, screening the candidate ligand molecules to obtain target ligand molecules; the molecular screening model is an artificial intelligence model for co-crystal ligand screening; Among them, the drug co-crystal ligand database is constructed using a large number of ligand molecules and ligand molecule properties; the expert system is built based on functional group matching and supramolecular synthon rules; the molecular screening model is obtained by training a machine learning model with ligand molecule samples.

2. The eutectic ligand screening method according to claim 1, characterized in that, The ligand molecule properties in the ligand database at least include properties obtained by quantum calculation or prediction and properties estimated by a thermodynamic model.

3. The eutectic ligand screening method according to claim 2, wherein The properties obtained by quantum calculation or prediction at least include molecular surface area, volume, molecular polar surface area, non-polar surface area, and the extreme points of molecular surface electrostatic potential; and / or, The properties estimated by the thermodynamic model at least include density, cohesive energy, and melting point.

4. The eutectic ligand screening method according to claim 3, wherein Constructing the drug co-crystal ligand database using a large number of ligand molecules and ligand molecule properties specifically includes: Obtaining ligand molecules that meet preset conditions, where the preset conditions include that the molecular toxicity is lower than the toxicity threshold; Optimizing the structure of the ligand molecules and performing quantum calculations to obtain the electron density wave function of the ligand molecules; Based on the electron density wave function, calculating ligand molecule properties, where the ligand molecule properties at least include molecular properties such as molecular surface, molecular shape, molecular volume, and the extreme points of molecular surface electrostatic potential; Integrating the ligand molecules and the corresponding ligand molecule properties to form ligand molecule data; Based on a large number of ligand molecule data, forming the drug co-crystal ligand database.

5. The eutectic ligand screening method according to claim 1, wherein Building the expert system based on functional group matching and supramolecular synthon rules specifically includes: Identifying the structural characteristics of the target molecule, and obtaining multiple ligand molecules that match in the drug co-crystal ligand database based on the structural characteristics; Constructing multiple supramolecular synthons according to the structures of the target molecule and the ligand molecules; Optimizing by molecular mechanics methods and calculating the binding energies of each supramolecular synthon respectively to obtain multiple binding energies; Determining the binding possibility between the target molecule and multiple ligand molecules based on the magnitude of the binding energy to obtain a ranking result of the binding possibility; Forming the preliminary screening result of the ligand molecules by the expert system based on the ranking result.

6. The eutectic ligand screening method according to claim 1, wherein Obtaining the molecular screening model by training a machine learning model with ligand molecule samples specifically includes: Obtaining ligand molecule samples, where the ligand molecule samples include positive samples that can form co-crystals and negative samples that cannot form co-crystals; Integrating the ligand molecule samples and molecular property samples into ligand molecule data samples, and forming a data set, and splitting the data set into a training set, a test set, and a validation set; Input the dataset into the network model, train and update the parameters in the model according to the gap between the prediction results of the training set and the true values of the training set, determine the optimal parameters according to the optimal results on the validation set, and detect the test set data; According to the hyperparameter optimization strategy, determine the optimal hyperparameter combination of the model to obtain an optimized molecular screening model.

7. A eutectic ligand screening system, characterized in that, The system includes: A ligand molecule acquisition unit, configured to obtain ligand molecules matching the target molecular structure based on a pre-constructed drug cocrystal ligand database and an expert system for cocrystal screening based on the ligand database, and use the ligand molecules obtained from the ligand database as candidate ligand molecules; A ligand molecule screening unit, configured to screen the candidate ligand molecules based on a pre-trained molecular screening model to obtain target ligand molecules; the molecular screening model is an artificial intelligence model for cocrystal ligand screening; Wherein, the drug cocrystal ligand database is constructed using a large number of ligand molecules and ligand molecule properties; the expert system is built based on functional group matching and supramolecular synthon rules; the molecular screening model is obtained by training a machine learning model with ligand molecule samples.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Modeling method and device of compound toxicity prediction model and application of compound toxicity prediction model

    CN110890137A

  • Truncated cartilage-homing peptides and peptide complexes and methods of use thereof

    CN112105375A