A drug virtual screening method based on capsule dynamic attention mechanism
By constructing a small molecule-protein association network using a capsule dynamic attention mechanism-based virtual drug screening method, the problem of insufficient training samples in new drug development is solved, screening efficiency and accuracy are improved, and interpretable drug screening results are achieved.
Patent Information
- Application Number
- CN202210909693.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-29
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2042-07-29
AI Technical Summary
Existing technologies often suffer from poor model performance in virtual screening of ligands for new drug targets or rare diseases during new drug development due to insufficient training samples, and lack effective data-driven feature extraction methods, resulting in low screening efficiency.
A virtual drug screening method based on capsule dynamic attention mechanism is adopted. By constructing a small molecule-protein association network, protein similarity information is obtained by using similarity algorithm and shortest path algorithm, and interaction is calculated by combining quantum chemical methods. The drug screening model is trained by capsule dynamic attention mechanism algorithm to extract rich data-driven feature vectors.
It improves the efficiency of virtual drug screening, simplifies the screening process, enhances the interpretability of the model, avoids data sparsity problems, and ensures the interpretability and accuracy of screening results.
Smart Images

Figure CN115171811B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence drug research and development, in particular to a drug virtual screening method based on a capsule dynamic attention mechanism. BACKGROUND
[0002] In recent years, new drug development for new drug targets or rare diseases has become a research hotspot in recent years. New drug development needs to first determine the biological activity value of the target and the compound through high-throughput experimental technology to a huge number of compound databases to screen lead compounds. However, the experimental method is time-consuming and labor-intensive, and worse, the number of available compounds is very limited, and not all drug targets are suitable for high-throughput screening experiments. With the rapid development of artificial intelligence, ligand virtual screening method of machine learning is the most important method for screening lead compounds. However, sometimes the training samples of new drug targets or rare disease ligand virtual screening are insufficient or even missing, and good models are often not obtained when they are virtually screened. SUMMARY
[0003] The purpose of the present application is to provide a drug virtual screening method based on a capsule dynamic attention mechanism, which can effectively improve the efficiency of virtual screening by optimizing the method process of drug virtual screening.
[0004] The embodiments of the present application are implemented as follows:
[0005] In a first aspect, the embodiments of the present application provide a drug virtual screening method based on a capsule dynamic attention mechanism, comprising the following steps:
[0006] Step S101: Obtain and utilize small molecule data and small molecule-protein complex data to obtain a benchmark database containing small molecule information and protein information;
[0007] Step S102: Based on the benchmark database, utilize a similarity algorithm to obtain protein similarity information;
[0008] Step S103: Based on the protein similarity information, perform clustering processing on the protein information and the small molecule information in the benchmark database to obtain clustering information;
[0009] Step S104: Based on the clustering information, obtain protein information and small molecule information of the same category from the benchmark database, and obtain heterogeneous association relationship information of the protein information and the small molecule information based on a shortest path algorithm;
[0010] Step S105: Based on a quantum chemistry method, calculate the interaction of the binding site of the small molecule information and the protein information to obtain interaction information;
[0011] Step S106: inputting the heterogeneous association relationship information and the interaction information into a pre-trained drug screening model to obtain target virtual screening information.
[0012] In some embodiments of the present application, the step of obtaining and utilizing the small molecule data and the small molecule-protein complex data to obtain the benchmark database containing small molecule information and protein information specifically comprises:
[0013] Obtaining compound information for target activity testing, and performing clustering processing on the target protein based on the compound information;
[0014] Integrating the compound information according to a preset first integration rule to obtain a small molecule data set;
[0015] Obtaining small molecule-protein complex data, and integrating the small molecule-protein complex data based on a preset second integration rule to obtain a small molecule-protein complex data set;
[0016] Utilizing the small molecule data set and the small molecule-protein complex data set to obtain the benchmark database containing small molecule information and protein information.
[0017] In some embodiments of the present application, the second integration rule comprises retaining a complex structure with an accurate affinity value, retaining a complex structure with a resolution higher than , retaining a structure with an effective graphical representation of a compound, and retaining a structure of a protein that can be successfully mapped to the UniProt library through sequence data.
[0018] In some embodiments of the present application, the step of obtaining protein similarity information based on the benchmark database and utilizing a similarity algorithm specifically comprises:
[0019] Utilizing the formula to obtain Gaussian interaction property kernel similarity information of any two proteins, wherein R(r i ,r j ) is the Gaussian interaction property kernel similarity information of r i and r j , r i and r j represent the interaction spectrum of the i th and j th proteins, n r is the number of proteins, r r is a kernel bandwidth control parameter, and r * is a frequency width parameter.
[0020] Utilizing the formula to obtain expression similarity information of the X group and Y group of proteins, wherein X i and Y irespectively, are expression values of proteins in the protein group X and the protein group Y, and N is the number of components in the expression profile, and respectively, are average values of expression values of proteins in the protein group X and the protein group Y;
[0021] The sequence similarity relationship of the i th and j th proteins is obtained by using the formula wherein NW(i,j) is the score of the i th and j th proteins in the Needleman-Wunsch algorithm.
[0022] Based on the Gaussian interaction profile kernel similarity information, the expression similarity information and the sequence similarity relationship, the protein similarity information is obtained.
[0023] In some embodiments of the present application, the step of obtaining the interaction information by combining the small molecule information and the protein information based on the quantum chemistry method specifically comprises:
[0024] The small molecule-protein complex structure is predicted one by one by using at least one method in Fpocket, GHECOM, ConCavity, POCASA and MetaPocket2.0, and the interaction information of the small molecule-protein complex and the important amino acid residues is obtained through consistency evaluation.
[0025] The interaction of the small molecule and the important amino acid residues is accurately calculated based on the quantum chemistry method, and the non-covalent interaction information including hydrogen bond, hydrophobic interaction, π-stacking, π-cation, salt bridge and halogen bond is obtained.
[0026] In some embodiments of the present application, the training step of the drug screening model specifically comprises:
[0027] The low-level features of the small molecule data and the small molecule-protein complex association relationship data are extracted.
[0028] The joint representation information of the high-level features of the small molecule and the protein is obtained by using the low-level features through the capsule dynamic attention mechanism algorithm, and the target target virtual screening information is output through the full connection layer and the softmax layer.
[0029] In some embodiments of the present application, the low-level features of the small molecule data and the small molecule-protein complex association relationship data include at least one of the convolutional neural network, the restricted Boltzmann machine and the long short-term memory network.
[0030] In a second aspect, the embodiments of the present application provide an electronic device, which includes a memory for storing one or more programs, and a processor. When the one or more programs are executed by the processor, the method in any of the above first aspect is implemented.
[0031] In a third aspect, the embodiments of the present application provide a computer readable storage medium, having stored thereon a computer program, which, when executed by a processor, implements the method of any one of the above first aspect.
[0032] Compared with the prior art, the embodiments of the present application have at least the following advantages or beneficial effects:
[0033] In the embodiments of the present application, the small molecule-protein correlation relationship network is constructed based on the similarity algorithm and the shortest path algorithm, that is, the heterogeneous correlation relationship information of the protein information and the small molecule information can effectively avoid the problem of data sparseness, and then the subsequent drug screening model can fully mine, learn and extract the data-driven feature vector including rich internal feature information from the character sequence data set of the small molecule and the protein. In addition, the interaction of the important amino acids of the small molecule and the protein binding site is accurately analyzed by the quantum chemical method, and the quantitative value is introduced as a data-driven internal feature auxiliary vector. The whole method logic is simple, which can simplify the virtual screening and improve the screening efficiency. In addition, it also has strong interpretability (there is no phenomenon that there is a large "hidden layer" and "neuron" between the input data and the output result in the machine learning model, and each step of input data and output result can be specifically explained), that is, the final output result can be explained, and there is no difficult-to-understand numerical value. BRIEF DESCRIPTION OF DRAWINGS
[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0035] Figure 1 The flowchart of the embodiment of the present application is a drug virtual screening method based on a capsule dynamic attention mechanism.
[0036] Figure 2 The specific flowchart of the step of obtaining and utilizing the small molecule data and the small molecule-protein complex data to obtain the benchmark database containing the small molecule information and the protein information in the embodiment of the present application is shown in the following figure.
[0037] Figure 3 The specific flowchart of the step of obtaining the protein similarity information based on the benchmark database and using the similarity algorithm in the embodiment of the present application is shown in the following figure.
[0038] Figure 4A specific flow chart of the step of obtaining interaction information by calculating the interaction between the small molecule information and the protein information based on the quantum chemistry method in the embodiment of the present application is shown in the following table.
[0039] Figure 5 A structural block diagram of an electronic device provided by the embodiment of the present application is shown in the following table.
[0040] Icon: 1, memory; 2, processor; 3, communication interface. DETAILED DESCRIPTION
[0041] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments of the present application. The components of the embodiments of the present application described and shown in the drawings can be arranged and designed in various different configurations.
[0042] Therefore, the detailed description of the embodiments of the present application provided in the drawings below is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative labor are within the scope of protection of the present application.
[0043] It should be noted that: similar reference numerals and letters represent similar items in the following drawings, therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. Meanwhile, in the description of the present application, if the terms such as first, second, etc. appear, they are only used to distinguish description, and cannot be understood as indicating or implying relative importance.
[0044] It should be noted that: similar reference numerals and letters represent similar items in the following drawings, therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. Meanwhile, in the description of the present application, if the terms such as first, second, etc. appear, they are only used to distinguish description, and cannot be understood as indicating or implying relative importance.
[0045] Some embodiments of the present application will be described in detail with reference to the drawings. In the following description, various embodiments described and embodiments in each of the following examples can be combined with each other, without conflict.
[0046] Embodiments
[0047] Referring to Figure 1 The capsule-based dynamic attention mechanism-based drug virtual screening method comprises the following steps:
[0048] Step S101: Obtain and utilize small molecule data and small molecule-protein complex data to obtain a benchmark database containing small molecule information and protein information.
[0049] Virtual screening is a computational technique for drug discovery, which is used to search small molecule libraries to identify those structures that can bind to drug targets (usually protein receptors or enzymes). Compared with traditional experimental high-throughput screening, virtual screening is a more direct and more reasonable method of drug discovery, and has the advantages of low cost and effective screening. In the above step, by obtaining and utilizing small molecule data and small molecule-protein complex data to obtain a benchmark database containing small molecule information and protein information, raw data support can be provided for subsequent virtual screening.
[0050] Specifically, referring to Figure 2 The above step of obtaining and utilizing small molecule data and small molecule-protein complex data to obtain a benchmark database containing small molecule information and protein information specifically comprises:
[0051] Step S201: Obtain compound information for target activity testing, and perform clustering processing on target proteins based on the compound information.
[0052] In the above step, by clustering the obtained compound information for target activity testing, objects similar to each other can be divided into the same classification cluster, facilitating subsequent processing in different classification clusters. Exemplarily, the target proteins can be four-clustered: enzymes, ion channels, G protein-coupled receptors, and nucleotide receptors. In addition, all compound information for target activity testing can be obtained by downloading ChEMBL20 database.
[0053] Step S202: Integrate the compound information according to the preset first integration rule to obtain a small molecule data set;
[0054] In the above step, by integrating the compound information according to the preset first integration rule, the small molecule data in the compound information can be sorted out to remove unnecessary data, and a more accurate and effective small molecule data set can be obtained.
[0055] For example, the first integration rule may include: removing multi-component compounds, such as mixtures and salts, from the compound information; removing molecules that have isotopic and stereochiral properties; removing molecules that contain one or more inorganic elements; removing compounds that have not been measured and that do not interact with the target; removing compounds with uncertain or ambiguous activity test values; and setting the activity threshold of the compound to 10 μM.
[0056] Step S203: Obtain small molecule-protein complex data and integrate the small molecule-protein complex data based on the preset second integration rule to obtain a small molecule-protein complex dataset.
[0057] In the above steps, the small molecule-protein complex data is integrated using a pre-defined second integration rule. This allows for preliminary screening and integration of the data, improving its purity and facilitating subsequent computational processing. This enhances the efficiency and accuracy of subsequent data processing (by removing unnecessary data and performing data filtering). For example, small molecule-protein complex data can be obtained by downloading complexes from the PDBbind database.
[0058] Specifically, the second integration rule mentioned above includes retaining the structure of each complex with an accurate affinity value and retaining structures with a resolution higher than [missing information]. The integration process preserves the complex structure, retains the structure of compounds with effective graphical representation, and preserves the structure of proteins that can be successfully mapped to the UniProt library through sequence data. Integrating small molecule-protein complex data using the second integration rule described above effectively removes invalid data and improves data validity.
[0059] Step S204: Use the small molecule dataset and the small molecule-protein composite dataset to obtain a benchmark database containing small molecule information and protein information.
[0060] In the above steps, after obtaining the small molecule dataset and the small molecule-protein complex dataset, it can be used to construct a benchmark database containing small molecule information and protein information, thereby providing raw data support for subsequent virtual screening processes.
[0061] Step S102: Based on the benchmark database, obtain protein similarity information using a similarity algorithm;
[0062] In the above step, the protein similarity information can be obtained by using the similarity algorithm, which facilitates the subsequent mining of the correlation between small molecules and proteins, and provides original input data for the final virtual screening. Exemplarily, the protein similarity information can be obtained by using multiple similarity algorithms, so that more accurate and effective similarity data can be obtained.
[0063] Specifically, referring to Figure 3 , the above step of obtaining protein similarity information based on the reference database and using the similarity algorithm specifically includes:
[0064] Step S301: using the formula to obtain the Gaussian interaction property kernel similarity information of any two proteins, wherein R(r i ,r j ) is the Gaussian interaction property kernel similarity information of r i and r j , r i and r j represent the interaction spectrum of the i-th and j-th proteins, respectively, n r is the number of proteins, r r is a kernel bandwidth control parameter, and r * is a frequency bandwidth parameter.
[0065] In the above step, the Gaussian interaction property kernel similarity information of any two proteins can be obtained by calculating the monotonic function of the Euclidean distance between any point r(i) and another point r(j) in the space. Wherein r r is a kernel bandwidth control parameter which can be used to control the radial range of action of the above Gaussian kernel function , by dividing the frequency bandwidth parameter r * by the total number of associated nodes of each node, the kernel vector can be independent of the size of the data set, and the parameters are more accurate and effective. Exemplarily, r * can take a value of 1.
[0066] Step S302: using the formula to obtain the expression similarity information of the X group and Y group of protein groups, wherein X i and Y i are the expression values of the proteins in the X group and Y group of protein groups, respectively, N is the number of components in the expression profile, and are the average expression values of the proteins in the X group and Y group of protein groups, respectively.
[0067] In the above step, the expression similarity information of the X group and the Y group proteomes can be obtained by calculating and processing the expression values of the proteins in the X group and the Y group proteomes in the above expression similarity formula, so that more accurate protein similarity information can be obtained subsequently.
[0068] Step S303: obtaining the sequence similarity relationship of the ith and jth proteins by using the formula obtaining the sequence similarity relationship of the ith and jth proteins, wherein NW(i,j) is the score of the ith and jth proteins in the NEEdleman-Wunsch algorithm.
[0069] In the above step, the NEEdleman-Wunsch algorithm is an algorithm for matching protein sequences or DNA sequences based on the knowledge of bioinformatics, which can be used to compare any two sequences to obtain a corresponding score, so that the sequence similarity relationship of the ith and jth proteins can be obtained by using the formula obtaining the sequence similarity relationship of the ith and jth proteins.
[0070] Step S304: obtaining the protein similarity information based on the Gaussian interaction profile kernel similarity information, the expression similarity information and the sequence similarity relationship.
[0071] In the above step, the protein similarity information can be constructed from three angles based on the Gaussian interaction profile kernel similarity, the expression similarity and the sequence similarity, which can improve the accuracy and effectiveness of data construction.
[0072] Step S103: performing clustering processing on the protein information and the small molecule information in the reference database based on the protein similarity information to obtain clustering information.
[0073] In the above step, the protein information and the small molecule information in the reference database are very large, and if they are directly analyzed, it will be time-consuming and laborious, and typical representatives of the same type cannot be effectively analyzed in depth. That is, by performing clustering processing on the protein information and the small molecule information in the reference database based on the protein similarity information, the protein information and the small molecule information in the same similarity network can be analyzed in depth according to the clustering results subsequently.
[0074] For example, the clustering processing can be performed by methods such as system clustering method, K-means method, fuzzy clustering method, clustering of ordered samples, decomposition method and addition method.
[0075] Step S104: obtaining the protein information and the small molecule information of the same category from the reference database based on the clustering information, and obtaining the heterogeneous association relationship information of the protein information and the small molecule information based on the shortest path algorithm.
[0076] In the above step, by calculating the protein similarity information after clustering through the shortest path algorithm, the heterogeneous association relationship information between the protein information and the small molecule information can be obtained, so that the subsequent similarity feature vectors of the protein information and the small molecule information can be used as one of the input parameters of the drug screening model.
[0077] Step S105: Calculate the interaction between the small molecule information and the protein information binding site based on the quantum chemistry method to obtain interaction information.
[0078] In the above step, by calculating the interaction between the small molecule information and the protein information binding site based on the quantum chemistry method to obtain interaction information, the subsequent small molecule-protein non-covalent interaction features can be extracted, which can be used as one of the input parameters of the drug screening model.
[0079] Specifically, please refer to Figure 4 The above step of calculating the interaction between the small molecule information and the protein information binding site based on the quantum chemistry method to obtain interaction information specifically includes:
[0080] Step S401: Use at least one method of Fpocket, GHECOM, ConCavity, POCASA and MetaPocket2.0 to predict small molecule-protein complex structures one by one, and obtain the interaction information between the small molecule-protein complex and important amino acid residues through consistency evaluation;
[0081] Step S402: Accurately calculate the interaction between the small molecule and the important amino acid residues based on the quantum chemistry method to obtain non-covalent interaction information including hydrogen bond, hydrophobic interaction, π-stacking, π-cation, salt bridge and halogen bond.
[0082] In the above step, the interaction information between the small molecule-protein complex and the important amino acid residues is initially obtained through consistency evaluation, and then the interaction between the small molecule and the important amino acid residues is accurately calculated based on the quantum chemistry method, so that accurate and effective interaction information can be obtained.
[0083] Step S106: Input the heterogeneous association relationship information and the interaction information into the pre-trained drug screening model to obtain target target virtual screening information.
[0084] In the above steps, the drug screening model is trained end-to-end on a large database containing small molecule-target protein relationship information, thereby obtaining the basic knowledge of the interaction between small molecules and proteins. Then the learned drug screening model is migrated to the target task of specific target prediction, and the model is further trained on the data set of the target target, so as to combine the basic knowledge and the unique information features of the target target. Finally, the drug screening model can perform virtual screening based on the heterogeneous association relationship information and interaction information according to the knowledge elements learned in the pre-training and migration training, and output the target target virtual screening information.
[0085] Exemplarily, the drug screening model can be structurally built based on a capsule dynamic attention mechanism. The input of the capsule dynamic attention mechanism algorithm is: a small molecule-protein feature vector combination F, a reference vector h, and an iteration number N; and the output is a coupling coefficient c and a vector output s of the capsule. The specific process can be: (1) initializing mapping matrices W f and W h ; (2) mapping the feature vector combination F and the reference vector h to obtain f i p ∈F p and h proj ; (3) initializing S0, s0←h pyoj ; (4) looping N times, t←0; (5) obtaining a coupling coefficient c i ←softmax(b i ); (6) obtaining a weighted sum feature: (7) updating the capsule output:
[0086] (8) updating the consensus variable: (9) ending the loop; (10) returning the coupling coefficient c and the vector output s of the capsule. That is, it first maps the initialized mapping matrices W f and W h , the reference vector and the feature vector set to the same space dimension, to obtain f i p ∈F p and h proj . Wherein, f i p =W f ·f i , h proj =W h ·h. Then the capsule output s is initialized using the mapped reference h proj vector. After t iterations, the output of the capsule is: s0=h proj 、 wherein, ci is the coupling coefficient, and its updating manner makes the dynamic routing manner of the output updated. First, the drug screening model updates b i through softmax i , b i represents the degree of correlation between the input and output vectors of the capsule, and the drug screening model initializes b i with a zero vector, and then updates the value of b i by accumulating the contents of the input and output of the capsule each time iteration. The specific formula is:
[0087] b i = f i p · s t + b i , The final algorithm returns the obtained coupling coefficient c and the vector output s of the capsule. The drug screening model can visualize the coupling coefficient c, and the capsule output s represents the joint representation of the feature matrix set and the reference vector. The above capsule dynamic attention mechanism algorithm updates its coupling coefficient by dynamic routing, and realizes single-layer multi-step updating of the coupling coefficient. The algorithm greatly reduces the parameter amount of the drug screening model while ensuring the performance of the drug screening model. The traditional single-layer attention algorithm cannot help the model to accurately distinguish the features highly related to the task. The traditional multi-layer attention algorithm helps the model to find the features highly related to the task step by step through multi-layer superposition, but the multi-layer attention algorithm also introduces a large number of parameters, which is not conducive to network lightweight. Unlike the traditional multi-step attention algorithm, the above capsule dynamic attention mechanism algorithm only completes multi-step attention operation in a single layer, updates the coupling coefficient c i through routing update, and obtains the joint representation s of the feature matrix set and the reference vector through multiple accumulations. In the drug screening model, the parameter amount is much less than that of the traditional multi-layer attention algorithm.
[0088] Specifically, the training steps of the above drug screening model specifically include:
[0089] extracting low-level features of small molecule data and small molecule-protein complex association data;
[0090] The capsule dynamic attention mechanism algorithm uses low-level features to obtain the joint representation information of high-level features of small molecules and proteins, and outputs the target virtual screening information through the full connection layer and the softmax layer.
[0091] In the above steps, by extracting low-level features of small molecule data and small molecule-protein complex correlation data first, and then using a capsule dynamic attention mechanism algorithm to obtain joint representation information of small molecules and protein high-level features using the low-level features, information that can be used for output target virtual screening information can be obtained, thereby completing the corresponding drug virtual screening.
[0092] Specifically, the extraction of low-level features of small molecule data and small molecule-protein complex correlation data includes using at least one of a convolutional neural network, a restricted Boltzmann machine, and a long short-term memory network. It should be noted that low-level features of small molecule data and small molecule-protein complex correlation data can also be obtained by other methods, as long as the low-level features of small molecule data and small molecule-protein complex correlation data can be obtained.
[0093] Please refer to Figure 5 , Figure 5 A structural block diagram of an electronic device is provided for the embodiments of the present application. The electronic device includes a memory 1, a processor 2 and a communication interface 3, which are directly or indirectly electrically connected to each other to realize the transmission or interaction of data. For example, these elements can be electrically connected to each other through one or more communication buses or signal lines. The memory 1 can be used to store software programs and modules, such as the program instructions / modules of the drug virtual screening system based on the capsule dynamic attention mechanism provided by the embodiments of the present application. The processor 2 executes the software programs and modules stored in the memory 1 to perform various functional applications and data processing. The communication interface 3 can be used for signaling or data communication with other node devices.
[0094] Among them, the memory 1 can be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.
[0095] The processor 2 can be an integrated circuit chip with signal processing capability. The processor 2 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; or can be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.
[0096] It can be understood that, Figure 5 The structure shown is only schematic, and the electronic device can further include more or fewer components than those shown in the Figure 5 embodiments, or have a different configuration from that shown in the Figure 5 embodiments. Figure 5 The components shown in the embodiments can be implemented in hardware, software or a combination thereof.
[0097] In the embodiments provided by the present application, it should be understood that the disclosed apparatus and method can also be implemented by other ways. The apparatus embodiments described above are only schematic, and the flowcharts and block diagrams in the accompanying drawings show the possible implementation architectures, functions and operation of the apparatus, method and computer program product according to the embodiments of the present application. In this regard, each block in the flowcharts and block diagrams can represent a module, a segment or a portion of code which comprises one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions shown in the blocks can be performed in a different order from that shown in the accompanying drawings. For example, two consecutive blocks can actually be performed substantially concurrently or in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts and combinations of blocks in the block diagrams and / or flowcharts can be implemented by dedicated hardware-based systems which perform the specified function or action, or can be implemented by a combination of dedicated hardware-based systems and computer instructions.
[0098] In addition, each functional module in the various embodiments of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0099] If the above functions are realized in the form of software function modules and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, and various media that can store program codes.
[0100] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
[0101] It is obvious for those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, and the present application can be realized in other specific forms without departing from the spirit or essential characteristics of the present application. Therefore, the embodiments should be regarded as exemplary and non-limiting, and the scope of the present application is defined by the appended claims rather than the above description, and all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application. Any reference signs in the claims should not be regarded as limiting the claims.
Claims
1. A virtual drug screening method based on capsule dynamic attention mechanism, characterized in that, Includes the following steps: Step S101: Obtain and utilize small molecule data and small molecule-protein complex data to obtain a benchmark database containing small molecule information and protein information; Step S102: Based on the benchmark database, obtain protein similarity information using a similarity algorithm; Step S103: Based on protein similarity information, perform clustering processing on protein information and small molecule information in the benchmark database to obtain clustering information; Step S104: Based on clustering information, obtain protein information and small molecule information of the same category from the benchmark database, and obtain heterogeneous association information between protein information and small molecule information based on the shortest path algorithm; Step S105: Calculate the interaction between small molecule information and protein information binding sites based on quantum chemical methods to obtain interaction information; Step S106: Input the heterogeneous association information and interaction information into the pre-trained drug screening model to obtain virtual screening information for target points; The drug screening model is structured based on the capsule dynamic attention mechanism; the input of the capsule dynamic attention mechanism algorithm is: small molecule-protein feature vector combination F, reference vector h, and iteration number N; the output is: coupling coefficient c and the vector output s of the capsule; the specific process is: (1) Initialize the mapping matrix W f and W h (2) Combine the feature vectors F and map them to the reference vector h to obtain f. i p ∈F p and h proj (3) Initialize s0, s0←h proj (4) Repeat N times, t←0; (5) Obtain the coupling coefficient c i ←softmax(b i (6) Obtain the weighted summation characteristic: (7) Update capsule output: (8) Update consensus variables: (9) End the loop; (10) Return the coupling coefficient c and the vector output s of the capsule; that is, it first initializes the mapping matrix W. f and W h The reference vector and the set of eigenvectors are mapped to the same spatial dimension to obtain f. i p ∈F p and h proj ; where f i p =W f ·f i h proj =W h ·h; then use the mapped reference h to output the capsule s. proj The vector is initialized; after t iterations, the capsule's output is: s0 = h proj , Among them, c i It is the coupling coefficient, and its update method is based on the dynamic routing of the output; firstly, the drug screening model updates through an intermediate variable b. i Perform softmax to update c i b i The drug screening model uses a zero vector to initialize b, representing the correlation between the input and output vectors of the capsule. i Then, in each iteration, the input and output of the capsule are accumulated to update b. i The value; the specific formula is: b i =f i p ·s t +b i , The final algorithm returns the coupling coefficient c and the capsule's vector output s; the drug screening model can visualize the coupling coefficient c, and the capsule output s represents the joint representation of the feature matrix set and the reference vector; the above capsule dynamic attention mechanism algorithm uses dynamic routing to update its coupling coefficient, realizing single-layer multi-step updating of the coupling coefficient.
2. The virtual drug screening method based on capsule dynamic attention mechanism as described in claim 1, characterized in that, The steps of acquiring and utilizing small molecule data and small molecule-protein complex data to obtain a benchmark database containing small molecule information and protein information specifically include: Obtain compound information for target activity testing, and cluster target proteins based on the compound information; The compound information is integrated according to the preset first integration rule to obtain a small molecule dataset; Small molecule-protein complex data are acquired and integrated based on a pre-defined second integration rule to obtain a small molecule-protein complex dataset. A benchmark database containing information on small molecules and proteins was obtained using small molecule datasets and small molecule-protein composite datasets.
3. The virtual drug screening method based on capsule dynamic attention mechanism as described in claim 2, characterized in that, The second integration rule includes retaining the structure of each complex with an accurate affinity value and retaining structures with a resolution higher than [missing information]. The complex structure, the structure of the preserved compound with an effective graphical representation, and the structure of the preserved protein that can be successfully mapped to the UniProt library through sequence data.
4. The virtual drug screening method based on capsule dynamic attention mechanism as described in claim 1, characterized in that, The steps for obtaining protein similarity information based on a benchmark database and using a similarity algorithm specifically include: Using formula The obtained Gaussian interaction properties kernel similarity information of any two proteins, where R(r i ,r j ) is r i and r j Gaussian interaction property kernel similarity information, r i and r j Let n represent the interaction spectra of the i-th and j-th proteins, respectively. r r represents the number of proteins. r r is a kernel bandwidth control parameter. * For bandwidth parameters; Using formula We obtained expression similarity information between the proteomes of group X and group Y, where X i and Y i These represent the protein expression values in the proteomes of groups X and Y, respectively, where N is the number of components in the expression profile. and These are the mean expression values of proteins in the proteomes of groups X and Y, respectively. Using formula Obtain the sequence similarity relationship between the i-th and j-th proteins, where NW(i,j) is the score of the i-th and j-th proteins in the Niederman-Onsch algorithm; Protein similarity information is obtained based on Gaussian interaction properties, kernel similarity information, expression similarity information, and sequence similarity relationships.
5. The virtual drug screening method based on capsule dynamic attention mechanism as described in claim 1, characterized in that, The steps for calculating the interaction between small molecule information and protein information binding sites based on quantum chemical methods to obtain interaction information specifically include: The structures of small molecule-protein complexes were predicted one by one using at least one of Fpocket, GHECOM, ConCavity, POCASA and MetaPocket 2.0, and the interaction information between small molecule-protein complexes and important amino acid residues was obtained through consistency evaluation. Based on quantum chemical methods, the interactions between small molecules and important amino acid residues are accurately calculated, and non-covalent interaction information including hydrogen bonds, hydrophobic interactions, π-stacking, π-cations, salt bridges, and halogen bonds is obtained.
6. The virtual drug screening method based on capsule dynamic attention mechanism as described in claim 1, characterized in that, The training steps of the drug screening model specifically include: Extract low-level features from small molecule data and small molecule-protein complex association data; The capsule dynamic attention mechanism algorithm utilizes low-level features to obtain joint representation information of small molecules and high-level protein features, and outputs virtual screening information of target points through fully connected layers and softmax layers.
7. The virtual drug screening method based on capsule dynamic attention mechanism as described in claim 6, characterized in that, The low-level features for extracting small molecule data and small molecule-protein complex association data include at least one of convolutional neural networks, restricted Boltzmann machines, and long short-term memory networks.
8. An electronic device, characterized in that, include: Memory, used to store one or more programs; processor; When the one or more programs are executed by the processor, the method as described in any one of claims 1-7 is implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Small molecule drug virtual screening method based on deep migration learning and application thereof
CN110459274A
Drug-target interaction prediction model method based on deep embedding learning of molecular graph and sequence
CN113327644A