Cocrystal formation prediction method and application based on substructure-molecular interaction network

By constructing a substructure-intermolecular interaction network and combining a third-order path algorithm, the existing eutectic prediction methods have solved the problem of high computational complexity and negative sample quality, and efficient and accurate eutectic prediction is achieved, especially suitable for new compounds outside the network.

CN114639449BActive Publication Date: 2025-05-09EAST CHINA UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011479079.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-15
Publication Date
2025-05-09
Estimated Expiration
2040-12-15

AI Technical Summary

Technical Problem

The existing eutectic prediction methods have high computational complexity, low efficiency and negative sample quality problems, making them difficult to apply to large-scale virtual screening and brand-new compound prediction outside the network.

Method used

Develop a method for eutectic formation prediction based on substructure-intermolecular interaction network, and use the construction of intermolecular interaction network, compound-substructure-intermolecular interaction network, and conduct eutectic prediction with third-order path algorithm.

Benefits of technology

It improves the efficiency and success rate of eutectic screening, can predict eutectic ligands for new compounds outside the network, and reduces the computational complexity and dependence on negative samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114639449B_ABST
    Figure CN114639449B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and application for predicting the formation of cocrystals based on a substructure-intermolecular interaction network. The prediction method includes: the construction of an intermolecular interaction network, the construction of a compound-substructure network, the construction of a substructure-intermolecular interaction network, and the prediction based on the intermolecular complementary mechanism, recommending suitable cocrystal ligands for drugs or other types of compounds, thereby developing their cocrystals or salt forms. For a given molecule, a recommended list of cocrystal ligands can be given, and the cocrystal ligands at the top of the recommended list are the dominant ligands of the molecule and are easy to form cocrystals. The introduction of the substructure of the compound enables the structural information of the compound to be utilized by the prediction model and expands its scope of application. The cocrystal ligands can be predicted for compounds inside and outside the network at the same time. Cocrystal design and preparation based on this prediction method can greatly reduce the workload and consumption of chemicals and reagents in cocrystal screening experiments, thereby reducing various resource wastes and environmental pollution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computational crystallography, and in particular to drug cocrystal prediction using a recommendation algorithm based on a complex network. Background Art

[0002] Cocrystals are crystalline materials formed by two or more solid compounds arranged in the same crystal lattice in a certain stoichiometric ratio through non-covalent interactions such as hydrogen bonds, π-π interactions, and electrostatic interactions. In the field of biopharmaceuticals, pharmaceutical cocrystals are mainly used to improve the physicochemical properties and biopharmaceutical properties of drugs such as water solubility, dissolution rate, and bioavailability. Pharmaceutical cocrystals can improve the relevant properties of drugs without changing their chemical structure, and are an extremely attractive solid form of drugs. Due to their superior properties, pharmaceutical cocrystals can also apply for patent protection and be marketed as new drugs. Therefore, the study of pharmaceutical cocrystals has become a new hot spot in the current field of drug research and development. Cocrystals are also widely used in other fields such as organic synthesis, dyes, and energetic materials.

[0003] The methods for preparing cocrystals mainly include solution crystallization, grinding or solution-assisted grinding, and reactive crystallization. Relying on the experience of chemists to develop cocrystals requires screening a large number of cocrystal ligands, which is a time-consuming, labor-intensive and expensive process with a low success rate. A large number of ineffective screening experiments consume a large amount of organic reagents and chemicals that are harmful to the environment, resulting in a waste of resources. Therefore, the use of computational methods to conduct virtual screening of cocrystals and guide researchers to rationally design cocrystals has gradually attracted attention in the pharmaceutical and other fields.

[0004] Existing computational models can be mainly divided into two categories: energy-based and structure-based computational models. Energy-based computational methods usually require quantum chemistry or molecular dynamics simulation calculations, which have high computational complexity and low efficiency, and are difficult to apply to large-scale virtual screening. Structure-based methods rely on machine learning or deep learning to construct a eutectic formation prediction model, but the construction of such models requires high-quality negative samples, which are difficult to obtain in databases and literature and need to be additionally constructed through experiments or random generation. The quality problem of negative samples seriously affects the accuracy of such models. In view of this, the present invention is committed to developing new prediction methods to make up for the shortcomings of current prediction models.

[0005] Many systems in the real world can be well expressed as complex network structures, such as social, biological and information systems. The nodes in the network represent each individual, and the edges in the network represent the interactions or connections between these individuals. The recommendation system based on complex networks has played a huge role in drug research and development fields such as protein-protein interaction prediction, drug target prediction and drug repositioning. The co-crystal system can also be expressed as a network structure, with nodes representing compound molecules and edges representing the interaction relationship between the two compounds that form the co-crystal. In the early stage, we developed a co-crystal recommendation model based on the intermolecular interaction network, but this method can only predict co-crystal ligands for compounds within the intermolecular interaction network, and cannot be directly applied to new compounds outside the network. At the same time, this method only relies on the interaction relationship between molecules to predict co-crystals, ignoring the attribute information of the molecules themselves, such as the structural information of the compound. Therefore, the present invention develops a co-crystal formation prediction model based on the substructure-intermolecular interaction network to overcome the shortcomings of the current method. Summary of the invention

[0006] In view of the shortcomings of the current cocrystal prediction methods, the present invention aims to integrate the two tools of chemical informatics and network science to develop a cocrystal prediction model based on the substructure-intermolecular interaction network. For any compound molecule, the model can be used to generate a list of potential cocrystal ligand recommendations. Focusing on the cocrystal ligands at the top of the recommendation list can reduce the blindness of cocrystal screening and improve the screening success rate. Based on the above objectives, the technical solution of the present invention is as follows:

[0007] In a first aspect of the present invention, a method for predicting cocrystal formation based on a substructure-molecular interaction network is provided, and the main steps are as follows:

[0008] Step 1, constructing an intermolecular interaction network: collecting existing co-crystal data and performing data processing, expressing the intermolecular interaction relationship in the co-crystal system as a network structure, and the network is represented by an adjacency symmetric matrix A;

[0009] Step 2, constructing a compound-substructure network: Calculate different molecular fingerprints of the compound, respectively as the characterization of the compound's substructure, establish multiple compound-substructure networks A', and characterize the structural information of the compound; if the co-crystal ligand is predicted for a new compound outside the intermolecular interaction network, generate its substructure in the same way and incorporate it into the constructed compound-substructure network A';

[0010] Step 3, constructing a substructure-molecular interaction network: coupling the molecular interaction network A and the compound-substructure network A′ described in step 1 and step 2, respectively, to establish a substructure-molecular interaction network A′′;

[0011] Step 4: Based on the substructure-intermolecular interaction network, different third-order path algorithms are designed to mine potential intermolecular interaction relationships in the network and predict cocrystal ligands for existing compounds in the network or new compounds outside the network.

[0012] The specific construction process of each step is described as follows:

[0013] a. The construction of the molecular interaction network described in step 1 specifically includes:

[0014] (1) Collect existing eutectic data from literature or open source databases;

[0015] (2) Data processing: remove eutectic data containing inorganic compounds; remove composite crystals containing liquid and gas components; remove duplicate data;

[0016] (3) Organize the co-crystal data into an edge list structure, where each row represents two compounds that have an interaction and can form a co-crystal, i.e., form an edge in the network;

[0017] (4) The eutectic data of the edge list structure is expressed as the network adjacency matrix A = (a ij ) m×m , characterizing the molecular interactions of the eutectic system. Where m is the number of compound nodes. If nodes i and j form an edge, then a ij =1, otherwise a ij = 0. The adjacency matrix A is a symmetric matrix, that is, a ij =a ji , and a ij=i = 0 to exclude self-loops.

[0018] The intermolecular interaction network A is a unipartite network, which contains only compound nodes, and the compound nodes interact with each other to form intermolecular interaction edges.

[0019] b. In step 2, the process of constructing the compound-substructure network specifically includes:

[0020] (1) Calculate the MACCS fingerprint, PubChem fingerprint, FP4 (Substructure) fingerprint, KR (Klekota-Roth) fingerprint, Estate fingerprint, CDK fingerprint, CDKExt (CDK Extended) fingerprint, Graph (CDKGraph Only) fingerprint, AP (Atom Pairs) fingerprint, ECFP4 (Extended-Connectivity Fingerprint) fingerprint and FCFP4 (Feature-based Connectivity Fingerprint) fingerprint of the compound, and use these 11 molecular fingerprints to characterize the substructure of the compound respectively;

[0021] (2) If the compound that forms the cocrystal contains a substructure, a connection is established between the compound and the substructure to form a compound-substructure network. The 11 substructures produce a total of 11 different compound-substructure networks. At the same time, the compound-substructure connection of the new compound outside the intermolecular interaction network is constructed in the same way.

[0022] With the adjacency matrix A′=(a ij′ ) n×n Each compound-substructure network is characterized in the form of n, where n is the total number of compound nodes and substructure nodes in the network. If compound node i contains substructure j′, then a ij′ =1, otherwise a ij′ = 0. The adjacency matrix A′ is a symmetric matrix, that is, a ij′ =a j′i At the same time, a ij =a ji = 0 and a i′j′ =a j′i′ =0 indicates that there is no edge inside the compound or substructure.

[0023] The compound-substructure network is a bipartite network containing two types of nodes: compounds and compound substructures. There are edges between compounds and substructures, but there are no edges within compounds or substructures.

[0024] c. In step 3, the construction process of the substructure-molecule interaction network is as follows:

[0025] The intermolecular interaction network adjacency matrix A and the 11 compound substructure interaction network adjacency matrix A′ are combined to form a new substructure-intermolecular interaction network adjacency matrix A′′. Among them, A′′=A∪A′, which is a symmetric matrix. If compound nodes i and j form intermolecular interactions, then a ij =a ji =1, otherwise a ij =aji = 0; if compound i contains substructure j′, then a ij′ =a j′i =1, otherwise a ij′ =a j′i =0,a i′j′ =a j′i′ =0 means there is no edge between substructures.

[0026] d. In step 4, the method for constructing a prediction model based on a substructure-molecular interaction network is specifically as follows:

[0027] The molecular pairs mediated by the third-order path have certain structural complementarity and are easy to form cocrystals according to the supramolecular synthon mechanism. The present invention quantifies the tendency of intermolecular interaction based on the number of third-order paths between molecular pairs, thereby predicting the possibility of cocrystal formation. The number of third-order paths between nodes x and y is ∑α xu α uv α vy , if node x is connected to node u, then α xu =1, otherwise α xu =0,α uv and α vy Same reason.

[0028] Large degree nodes in the network, that is, nodes with many intermolecular interaction edges or compound-substructure edges, are more likely to form third-order paths with other compounds, which will bring deviations to the prediction. In order to eliminate the influence of large degree nodes, different degree penalty methods are used and different recommendation scoring algorithms are designed.

[0029] Penalize the degree of the endpoint nodes of the third-order path and design the following three different algorithms:

[0030]

[0031]

[0032]

[0033] Considering the impact of the degree values ​​of the third-order path intermediate nodes u and v on the recommendation score, the following algorithm is designed:

[0034]

[0035] The algorithm assigns more weights to third-order paths mediated by small-degree nodes.

[0036] Considering the influence of the third-order path endpoints and intermediate node degrees, the following algorithm is designed:

[0037]

[0038] Each of the above recommendation algorithms can assign recommendation scores to unconnected pairs of molecules in the network, and give a ranked list of potential ligands for the predicted molecules based on the recommendation scores. The higher the ranking, the greater the possibility that the ligand will form a cocrystal with the predicted molecule.

[0039] The above five algorithms were combined with 11 substructure-molecular interaction networks to construct a total of 55 prediction models. After evaluation through training sets and external validation sets, the better combination of algorithms and substructure-molecular interaction networks was selected as the prediction model for compound cocrystal ligands.

[0040] The second aspect of the present invention provides a eutectic formation prediction model constructed based on the above-mentioned prediction method, which has the following technical characteristics, including: an input display module, a construction module, a storage module, a search module, a eutectic ligand calculation module and a control module.

[0041] Among them, the input display module is used to input existing co-crystal data or compounds to be predicted, and display the prediction results; the construction module constructs an intermolecular interaction network based on the intermolecular interaction relationship in the co-crystal data, and calculates the substructure information of the compounds in the co-crystal data according to the molecular fingerprint, constructs a compound-substructure network, couples the two networks, and constructs the corresponding substructure-molecular interaction network. The storage module is used to store different types of substructure-intermolecular interaction networks and supporting third-order path algorithms; the search module is used to search for the intermolecular interaction relationship and substructure information of the compound to be predicted within the corresponding substructure-intermolecular interaction network; when the search module fails to find the intermolecular interaction relationship and substructure information of the compound to be predicted, that is, the compound to be predicted is a new compound outside the network, the construction module calculates the molecular fingerprint to construct the compound-substructure network of the compound, and incorporates it into the corresponding substructure-intermolecular interaction network; the co-crystal ligand calculation module assigns recommended scores to unconnected molecular pairs in the substructure-intermolecular interaction network according to the supporting third-order path algorithm, gives a ranking list of potential ligands for the predicted molecule according to the recommended scores, and according to preset rules, uses the top-ranked ligands that meet the conditions as the co-crystal ligand prediction results of the target compound; the storage module records and stores the results of the drug pathway relationship calculation module in real time; the control module controls the operation of the input display module, the construction module, the storage module, the search module and the co-crystal ligand calculation module.

[0042] Functions and Effects of the Invention

[0043] Compared with the existing eutectic prediction methods, the present invention has the following main advantages:

[0044] 1) The present invention integrates network technology and chemical informatics tools such as compound substructure, creatively constructs a substructure-molecular interaction network, and designs a third-order path algorithm based on the network for cocrystal prediction. This method enables the structural information of the compound to be used by the prediction model, and expands the application scope of the network-based prediction model. With the help of compound substructure as a bridge, cocrystal ligands can be predicted for new compounds outside the network;

[0045] 2) Compared with the energy-based method, the model developed by the present invention is easy to calculate and can complete a prediction in a few seconds, which is convenient and fast, and can achieve satisfactory prediction performance;

[0046] 3) Testing on standard data sets and external validation sets shows that the present invention has good predictive performance in predicting drug co-crystal ligands;

[0047] 4) Compared with the method based on machine learning, the present invention can construct a prediction model without negative samples. Negative samples are difficult to obtain in databases and literature. If the quality of negative samples generated by random generation or other calculation means is not high, it will affect the prediction performance of the model and reduce its accuracy.

[0048] Therefore, the prediction method developed by the present invention can recommend suitable co-crystal ligands for drugs or other types of compounds, thereby developing their co-crystals or salt forms. For a given molecule, the prediction model can give a recommended list of its co-crystal ligands, and the co-crystal ligands at the top of the recommendation list are considered as its dominant ligands. Paying attention to the predicted ligands at the top of the recommendation list can reduce the blindness of the co-crystal screening experiment and improve the success rate of co-crystal screening. The introduction of compound substructures expands the application scope of network-based methods, which can not only predict co-crystal ligands for compounds within the network, but also predict co-crystal ligands for new compounds outside the network.

[0049] The method developed by the present invention is simple to calculate, has strong generalization ability, and has a high accuracy rate. It can become an effective tool for predicting the formation of cocrystals and improve the efficiency of cocrystal screening. Based on this prediction model, efficient cocrystal design and preparation can greatly reduce the workload of cocrystal screening experiments and the consumption of chemicals and reagents therein, thereby reducing the waste of manpower and material resources and environmental pollution. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a flow chart for constructing a co-crystal formation prediction model based on a substructure-intermolecular interaction network of the present invention.

[0051] Figure 2 It is a schematic diagram of the intermolecular interaction network in Example 1. The dots represent the compounds, and the solid lines represent the cocrystals formed by the compounds.

[0052] Figure 3It is a schematic diagram of the substructure-molecular interaction network in Example 1. Because the complete network is too large to be displayed, this schematic diagram only uses the diclofenac cocrystal in the network under the MACCS substructure as an example for visualization. DETAILED DESCRIPTION

[0053] The present invention is further described below in conjunction with specific embodiments, but the protection scope of the present invention is not limited thereby.

[0054] Embodiment 1:

[0055] The cocrystal formation prediction model based on the substructure-molecular interaction network is constructed by referring to Figure 1 , the specific implementation steps are as follows.

[0056] Eutectic data collection: The training data used is from the literature Cryst. Growth Des. 2016, 16, 6095-6104. The initial data is 1941 binary crystal complex data. After data processing, 1600 eutectic data are finally obtained. Eutectic data is input using the input display module.

[0057] Construction of intermolecular interaction network: After data processing, the intermolecular interaction network was constructed, and only the largest connected subgraph in the network was retained as the intermolecular interaction network. The intermolecular interaction network contains 1486 edges composed of 879 compound nodes, that is, 879 compounds interact with each other to form an interaction entity containing 1486 edges. The remaining 114 pairs of cocrystal molecules only interact with each other and do not have intermolecular interactions with the compounds in this largest connected subgraph. The constructed intermolecular interaction network diagram can be found in Figure 2 .

[0058] Construction of compound-substructure network: PaDEL-Descriptor 2.21 was used to calculate the MACCS fingerprint, PubChem fingerprint, FP4 fingerprint, KR fingerprint, Estate fingerprint, CDK fingerprint, CDKExt, Graph fingerprint and AP fingerprint of the compound; RDKit 2019.03 was used to calculate the ECFP4 fingerprint and FCFP4 fingerprint of the compound. The above 11 molecular fingerprints were used to characterize the substructures of the compound, and the compound and the substructure were connected to construct 11 compound-substructure networks.

[0059] Construction of substructure-molecule interaction network: The intermolecular interaction network and the 11 compound-substructure networks are coupled to establish the 11 substructure-molecule interaction networks. The basic properties of the constructed intermolecular interaction network and the 11 substructure-molecule interaction networks are shown in Table 1. The 11 substructure-molecule interaction networks are constructed using the building module and stored in the storage module.

[0060] Table 1. Intermolecular interaction networks and substructures - basic properties of intermolecular interaction networks

[0061]

[0062]

[0063] a N M , N S , N MI and N MS are the number of compounds, substructures, intermolecular interactions, and compound-substructure edges in the network;

[0064] b The network density (Density, ρ) is defined as in N M The maximum number of edges that can be formed by the interaction of compounds. represents the maximum number of compound-substructure edges, that is, each compound in the network has all N S Obviously, network density represents the ratio of the actual number of edges in the network to the maximum possible number of edges in the network.

[0065] According to Table 1, the density of the intermolecular interaction network is low (0.004), indicating that the network structure is relatively sparse and the limited interaction information will lead to difficulties in prediction. Compared with the intermolecular interaction network, the number of nodes and edges in the substructure-intermolecular interaction network has greatly increased, and the network density under different substructures has increased by more than three times (0.013-0.088). Figure 3 Taking the case of the diclofenac cocrystal under the MACCS substructure as an example, diclofenac only forms two edges with 2,6-dimethylimidazo[2,1-b][1,3,4]thiadiazole and isonicotinamide in the intermolecular interaction network, but the introduction of the compound-substructure network greatly enriches the number of edges in the substructure-intermolecular interaction network. The introduced substructure adds additional structural information to the intermolecular interaction network and significantly increases the network density, which can be used by the prediction model.

[0066] Construction of prediction models: The above five algorithms were combined with 11 seed structure-molecular interaction networks to construct a total of 55 prediction models, namely the cocrystal ligand calculation module.

[0067] Embodiment 2:

[0068] In order to evaluate the prediction performance of each model, a systematic evaluation was performed on the 55 prediction models constructed in Example 1 of the present invention. The specific methods and model performance are as follows.

[0069] The prediction performance of the model was evaluated using 10-fold cross validation: first, the intermolecular interaction edges in the substructure-molecule interaction network were randomly divided into 10 parts; then, each intermolecular interaction edge was deleted from the network in turn and used as a test set, and the remaining 9 parts and all compound-substructure connections were used as training sets. The prediction model was used to generate the test set and the prediction scores for the non-existent edges. By comparing the scores of the test set and the non-existent edges, 10 sets of evaluation index results were obtained. In order to reduce the randomness of the results, the 10-fold cross validation was repeated 10 times, and finally 100 sets of evaluation results were obtained to systematically evaluate the prediction performance of the model.

[0070] The area under the receiver operating characteristic curve (AUC) and ranking score (RS) are used as evaluation indicators. In practical applications, AUC can be interpreted as the probability that the score of a randomly selected edge in the test set is higher than the score of the non-existent edge. In n independent comparisons, if there are n′ times when the score of the test edge is higher than the score of the non-existent edge, and n″ times when the scores are the same, then AUC is defined as:

[0071]

[0072] If all recommendation scores are generated independently and randomly, the AUC value is approximately 0.5. Therefore, the degree to which the AUC exceeds 0.5 represents the predictive ability of the model relative to random guessing. RS is the average ranking of all test edge scores in the entire prediction space, defined as:

[0073]

[0074] Among them, r e The ranking of the scores for a test edge, |E test | is the total number of edges in the test set. If a model has a smaller RS ​​value, the model can assign a higher prediction rank to the test set edges.

[0075] The AUC and RS values ​​of each prediction model are summarized in Table 2 and Table 3, respectively. It can be seen from Table 2 and Table 3 that the prediction performance of each model is relatively close, with AUC values ​​of about 0.9 and RS values ​​of about 0.1. Among the 11 different substructure representation methods, FCFP4 is preferred. Among algorithms 1-5 under the FCFP4 substructure, algorithm 1, algorithm 4 and algorithm 5 are preferred.

[0076] Table 2. AUC values ​​of each prediction model in 10-fold cross validation

[0077]

[0078]

[0079] Table 3. RS values ​​of each prediction model in 10-fold cross validation

[0080] Algorithm 1 Algorithm 2 Algorithm 3 Algorithm 4 Algorithm 5 Estate 0.154±0.015 0.120±0.015 0.128±0.015 0.101±0.014 0.093±0.013 KR 0.149±0.015 0.131±0.015 0.137±0.015 0.101±0.015 0.095±0.014 MACCS 0.116±0.015 0.107±0.015 0.107±0.015 0.100±0.016 0.091±0.014 PubChem 0.130±0.016 0.121±0.016 0.121±0.016 0.105±0.014 0.099±0.015 Sub 0.134±0.017 0.125±0.016 0.127±0.017 0.101±0.016 0.093±0.014 CDK 0.122±0.013 0.111±0.014 0.111±0.013 0.109±0.014 0.099±0.015 CDKE 0.123±0.012 0.113±0.012 0.113±0.012 0.113±0.015 0.102±0.013 CKD 0.127±0.016 0.116±0.015 0.116±0.016 0.109±0.015 0.099±0.014 ECFP4 0.106±0.014 0.108±0.015 0.108±0.014 0.099±0.013 0.094±0.012 FCFP4 0.105±0.015 0.107±0.014 0.107±0.015 0.094±0.015 0.090±0.015 AP 0.140±0.016 0.124±0.017 0.113±0.016 0.113±0.017 0.103±0.015

[0081] Embodiment 3:

[0082] In order to measure the generalization ability of the prediction model, the prediction ability of Algorithm 1_FCFP4, Algorithm 4_FCFP4 and Algorithm 5_FCFP4 models was evaluated using an external validation set. The specific implementation plan is to collect drug cocrystals of marketed drugs as an external validation set, and the data processing process is the same as that of the training set. The external validation set is divided into two parts:

[0083] 1) Category 1: 230 pairs of drug cocrystals, consisting of 40 APIs and 120 cocrystal ligands, both components of which exist within the intermolecular interaction network;

[0084] 2) Category 2: 288 pairs of drug co-crystals, of which 229 pairs of drug co-crystals were composed of 96 drugs (out-of-network compounds) and 90 co-crystal ligands, and 59 pairs of drug co-crystals were composed of 23 APIs and 56 co-crystal ligands (out-of-network compounds).

[0085] The performance of the recommended model in predicting cocrystal ligands for in-network and out-network compounds was evaluated by category 1 and category 2, respectively. The prediction model was used to predict cocrystal ligands for drugs in category 1 and out-network compounds in category 2, and the prediction results were compared with the actual data to evaluate the generalization ability of the model.

[0086] As shown in Table 4, the prediction model achieved a prediction performance comparable to that of 10-fold cross validation in two external validation sets, proving that our model has good generalization ability. The results show that the prediction model of the present invention performs well in predicting co-crystal ligands for new compounds outside the network, indicating that the strategy of coupling compound substructures and intermolecular interaction networks to predict co-crystal ligands for new compounds outside the network is reasonable, making up for the defect that the traditional intermolecular interaction network recommendation model cannot predict co-crystal ligands for new compounds outside the network.

[0087] Table 4. External validation results of the prediction model

[0088]

[0089] Embodiment 4:

[0090] The actual application of the co-crystal formation prediction method of Example 1 in predicting drug co-crystal ligands, the specific results are:

[0091] Diclofenac belongs to Class II drugs in the biopharmaceutics classification system, with low solubility and high permeability. Its water solubility can be improved by means of drug cocrystals. Figure 3 As shown in the figure, the cocrystals formed by diclofenac and its cocrystal ligands 2,6-dimethylimidazo[2,1-b][1,3,4]thiadiazole and isonicotinamide already exist in the intermolecular interaction network. Based on the intermolecular interaction relationship and compound substructure information, algorithms 1_FCFP4, 4_FCFP4 and 5_FCFP4 were used to predict its novel cocrystal ligands respectively, and the top 10 predicted ligands in each model recommendation list were taken as the initial recommended ligands. The three algorithms 1_FCFP4, 4_FCFP4 and 5_FCFP4 with different mechanisms were merged to construct a consensus model, and the intersection of the top 10 cocrystal ligands in the recommendation list of the three models was taken as the final recommended ligand. The prediction results are summarized in Table 5, among which 4,4'-bipyridine has formed a cocrystal with diclofenac (CSD Refcode:GOGGAN01), proving that our prediction model has strong prediction performance.

[0092] Table 5. Prediction results of diclofenac cocrystal ligands

[0093]

[0094] Embodiment 5:

[0095] The actual application of the co-crystal formation prediction method of Example 1 in the prediction of drug co-crystal ligands, the specific results are:

[0096] Apatinib is an effective drug for the treatment of advanced gastric cancer. However, it is poorly soluble in water in the form of free base, so its oral bioavailability is poor. The development of its drug cocrystal is expected to improve its water solubility. Although apatinib cocrystal has been reported recently, it does not exist in the constructed intermolecular interaction network (because the cocrystal data used in the model construction is data before 2016). Therefore, the search module identifies it as a new compound outside the network, and only relies on its structural information to predict cocrystal ligands for it, and verifies the actual performance of the model in predicting cocrystal ligands for new compounds outside the network. The compound-substructure network of apatinib was constructed using FCFP4 fingerprints, and it was connected to the intermolecular interaction network with the help of substructures. Algorithms 1_FCFP4, 4_FCFP4, and 5_FCFP4 were used to predict apatinib cocrystal ligands respectively, and the top 10 predicted ligands in the recommendation list of each model were taken as the initial recommended ligands. Algorithms 1_FCFP4, 4_FCFP4 and 5_FCFP4 with three different mechanisms were merged to construct a consensus model, and the intersection of the top 10 cocrystal ligands in the recommendation lists of the three models was taken as the final recommended ligand.

[0097] The prediction results are summarized in Table 6, among which the cocrystals formed by fumaric acid, succinic acid, p-aminobenzoic acid and p-hydroxybenzoic acid with apatinib have been reported in the document Cryst. Growth Des. 2020, 20, 6820-6830; the cocrystals formed by adipic acid and apatinib have been reported in the document Cryst. Growth Des. 2018, 18, 4701-4714. All five predicted ligands have been experimentally confirmed to form cocrystals with apatinib, indicating that the prediction model constructed by the present invention can be used as an effective tool to guide the experimental screening of cocrystals. For new compounds that have not been found to form cocrystals, suitable cocrystal ligands can be predicted for them based only on their structural information and the existing substructure-molecular interaction network, thereby effectively improving the success rate of cocrystal screening.

[0098] Table 6. Prediction results of apatinib cocrystal ligands

[0099]

[0100]

[0101] The above is only a partial embodiment of the present invention, and does not limit the present invention in any form. Although the present invention has been disclosed as above with preferred embodiments, it is not intended to limit the present invention. Any relevant technical personnel familiar with the art can use the above disclosed methods and technical contents to make many possible changes and modifications to the technical solutions of the present invention, or modify them into equivalent embodiments of equivalent changes without departing from the spirit and technical solutions of the present invention. Therefore, any simple modification, equivalent replacement, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention, still belong to the scope of protection of the technical solution of the present invention. At the same time, it is particularly important that although the above calculation method is carried out with eutectic prediction, it is difficult to strictly distinguish between eutectics and multi-component solid forms such as molecular salts in many cases. The application of the method and model of the present invention to the development of multi-component solid forms such as molecular salts is also within the scope of protection of the present invention.

Claims

1. A method for predicting cocrystal formation based on a substructure-molecular interaction network, characterized in that: The steps include: Step 1: Constructing the intermolecular interaction network: Collecting the existing co-crystal data and processing the data, expressing the intermolecular interaction relationship in the co-crystal system as a network structure, the network is represented by the adjacency symmetric matrix A = (a ij ) m×m express, Where m is the number of compound nodes. If nodes i and j form a eutectic, that is, form an edge in the network, then a ij =1, otherwise a ij =0; Step 2, constructing a compound-substructure network: by calculating different molecular fingerprints of the compound, characterizing the substructure of the compound, establishing multiple compound-substructure networks A′, and associating the structural information of different compounds; if the cocrystal ligand is predicted for a new compound outside the intermolecular interaction network, its substructure is generated in the same way and incorporated into the constructed compound-substructure network A′. Among them, A′=(a ij′ ) n×n is a symmetric matrix, n is the total number of compound nodes and substructure nodes in the network, if compound node i contains substructure j′, then a ij′ =1, otherwise a ij′ =0; and a ij =a j′i′ =0, indicating that there is no edge inside the compound or substructure; Step 3, constructing a substructure-molecular interaction network: coupling the molecular interaction network A and the compound-substructure network A′ described in step 1 and step 2, respectively, to establish multiple substructure-molecular interaction networks A″, Among them, A″=A∪A′, which is a symmetric matrix. If compound nodes i and j form intermolecular interactions, then a ij =1, otherwise a ij = 0; if compound i contains substructure j′, then a ij′ =1, otherwise a ij′ =0; a i′j′ =0 means there is no edge between substructures; Step 4: Construct a prediction model: Based on the substructure-intermolecular interaction network, design a third-order path algorithm according to the intermolecular complementary mechanism, explore potential intermolecular interaction relationships, and predict co-crystal ligands for compounds within the intermolecular interaction network or new compounds outside the network. The algorithm includes any one or more combinations of the following algorithms: Among them, x and y are unconnected molecular pairs to be predicted, u and v are other nodes in the network; ∑α xu α uv α vy is the number of third-order paths between nodes x and y. If node x is connected to node u, then α xu =1, otherwise α xu =0,α uv and α vy Similarly; k x , k y , k u and k v are the degrees of nodes x, y, u and v, that is, the number of edges formed by each node; S xy Represents the recommendation score between nodes x and y, representing the possibility of forming an edge between the two.

2. The method for predicting cocrystal formation based on substructure-intermolecular interaction network according to claim 1, It is characterized in that Among them, in step 1, the specific steps of constructing the intermolecular interaction network are as follows: (1) Collect existing eutectic data from literature or open source databases; (2) Data processing: remove eutectic data containing inorganic compounds, remove composite crystal data containing liquid and gas components, and remove duplicate data; (3) Organize the co-crystal data into an edge list structure, where each row represents two compounds that have an interaction and can form a co-crystal, i.e., form an edge in the network; (4) The eutectic data of the edge list structure is expressed as the adjacency symmetric matrix A of the network = (a ij ) m×m , characterizing the intermolecular interactions of the eutectic system.

3. The method for predicting co-crystal formation based on substructure-intermolecular interaction network according to claim 1, Features: Among them, in step 2, the specific steps of constructing the compound-substructure network are as follows: (1) Calculate the MACCS fingerprint, PubChem fingerprint, FP4 fingerprint, KR fingerprint, Estate fingerprint, CDK fingerprint, CDKExt fingerprint, Graph fingerprint, AP fingerprint, ECFP4 fingerprint and FCFP4 fingerprint of the compound, and use these 11 molecular fingerprints to characterize the substructure of the compound respectively; (2) If the compound contains a substructure, an edge is established between the compound and the substructure to form a compound-substructure network; The 11 substructures generate 11 different compound-substructure networks A′.

4. The method for predicting cocrystal formation based on substructure-intermolecular interaction network according to claim 1, characterized in that: in, In step 3, the construction of the substructure-molecule interaction network is as follows: The intermolecular interaction network A and the 11 compound-substructure networks A′ were merged to generate the 11 substructure-intermolecular interaction networks A″; the new compounds were connected to the intermolecular interaction networks based on their common substructures with the compounds in the network.

5. The method for predicting co-crystal formation based on substructure-intermolecular interaction network according to any one of claims 1 to 4, Features: Among them, in step 4, the algorithm The design basis is as follows: large-degree nodes in the network, that is, nodes with many intermolecular interaction edges or compound-substructure edges, are more likely to form third-order paths with other compounds, which will bring deviations to the prediction; in order to eliminate the influence of large-degree nodes, the above recommendation algorithm is designed by relying on different degree penalty methods; Considering the impact of the degree values ​​of the third-order path intermediate nodes u and v on the recommendation score, design an algorithm The algorithm assigns more weight to third-order paths mediated by small-degree nodes; Considering the influence of the degree values ​​of the endpoints and intermediate nodes of the third-order path, the algorithm is designed According to each of the above recommendation algorithms, a recommendation score is assigned to the molecular pairs that have not formed a cocrystal in the network, and a ranked list of potential ligands for the predicted molecules is given based on the recommendation score. The top-ranked ligands are considered to be the dominant ligand molecules that can form a cocrystal with the predicted molecules.

6. A eutectic formation prediction model constructed based on the eutectic formation prediction method according to claim 5, characterized in that: include: Input display module, construction module, storage module, search module, co-crystal ligand calculation module and control module, Wherein, the input display module is used to input co-crystal data or compounds to be predicted and display prediction results; The construction module constructs an intermolecular interaction network based on the intermolecular interaction relationship in the co-crystal data, calculates the substructure information of the compound in the co-crystal data based on the molecular fingerprint, constructs a compound-substructure network, couples the two networks, and constructs a corresponding substructure-molecular interaction network; The storage module is used to store different types of substructure-molecule interaction networks and supporting third-order path algorithms; The search module is used to search for the intermolecular interaction relationship and substructure information of the compound to be predicted within the corresponding substructure-intermolecular interaction network; When the search module fails to find the molecular interaction relationship and substructure information of the compound to be predicted, that is, the compound to be predicted is a new compound outside the network, the construction module calculates the molecular fingerprint to construct the compound-substructure network of the compound and incorporates it into the corresponding substructure-molecular interaction network; The co-crystal ligand calculation module assigns recommended scores to unconnected molecular pairs in the substructure-intermolecular interaction network according to the supporting third-order path algorithm, gives a ranking list of potential ligands for the predicted molecules according to the recommended scores, and uses the top-ranked ligands that meet the conditions as the co-crystal ligand prediction results of the target compound according to the preset rules; The storage module records and stores the results of the co-crystal ligand calculation module in real time; The control module controls the operation of the input display module, the construction module, the storage module, the search module and the co-crystal ligand calculation module.

Citation Information

Patent Citations

  • Virtual drug screening system for crystal compound

    CN111863120A

  • Eutectic prediction method based on graph neural network and deep learning framework

    CN111882044A