Molecular screening method, training method, device, electronic equipment and storage medium

CN116796282BActive Publication Date: 2026-08-18BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310489389.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-04
Publication Date
2026-08-18
Estimated Expiration
2043-05-04

AI Technical Summary

Benefits of technology

[0010] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method provided according to embodiments of this disclosure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116796282B_ABST
    Figure CN116796282B_ABST
Patent Text Reader

Abstract

The present disclosure provides a molecule screening method, a training method, a device, an electronic device and a storage medium, relates to the technical field of data processing, and particularly relates to the fields of artificial intelligence, big data or medicine technology. The specific implementation scheme is as follows: determining an initial binding conformation subset from a to-be-screened binding conformation set, wherein the initial binding conformation subset includes an initial binding conformation constructed by an initial molecule and a to-be-matched receptor object; performing binding property evaluation on the initial binding conformation to obtain a binding property evaluation result; determining a candidate binding conformation from the initial binding conformation subset according to the binding property evaluation result; processing the candidate binding conformation based on a deep learning algorithm to obtain an affinity detection result; and screening a target molecule matched with the to-be-matched receptor object from the candidate binding conformation according to the affinity detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and more particularly to the fields of artificial intelligence, big data, or pharmaceutical technology. Background Technology

[0002] With the rapid development of science and technology, in the field of pharmaceutical technology, researchers can use computer simulation-based virtual drug screening systems to screen drug molecules that are compatible with receptors such as proteins from a molecular library containing a large amount of data on molecules to be screened, thereby improving the efficiency of drug molecule screening. Summary of the Invention

[0003] This disclosure provides a molecular screening method, training method, apparatus, electronic device, storage medium, and computer program product.

[0004] According to one aspect of this disclosure, a molecular screening method is provided, comprising: determining an initial subset of binding conformations from a set of binding conformations to be screened, wherein the initial subset of binding conformations includes initial binding conformations constructed by an initial molecule and a receptor object to be matched; evaluating the binding properties of the initial binding conformations to obtain binding property evaluation results; determining candidate binding conformations from the initial subset of binding conformations based on the binding property evaluation results; processing the candidate binding conformations based on a deep learning algorithm to obtain affinity detection results; and screening target molecules that match the receptor object to be matched from the candidate binding conformations based on the affinity detection results.

[0005] According to another aspect of this disclosure, a method for training a deep learning model is provided, comprising: acquiring training samples, the training samples including sample binding conformations and affinity detection labels, the sample binding conformation including a sample binding conformation graph composed of sample molecules and sample receptor objects, the sample binding conformation graph including sample atomic nodes of sample molecules, sample atomic edge relationships between sample atomic nodes, sample receptor atomic nodes constituting sample receptor objects, and sample receptor atomic edge relationships between sample receptor atomic nodes; the sample binding conformation graph further includes sample binding edge relationships between sample atomic nodes and sample receptor atomic nodes; and inputting the sample binding conformation graph into a feature of the deep learning model. The feature extraction network outputs sample node features, sample first-side relationship features corresponding to sample atomic edge relationships or sample receptor atomic edge relationships, and sample second-side relationship features corresponding to sample binding edge relationships. The sample node features and sample first-side relationship features are input into the first detection network of the deep learning model, which outputs the sample affinity detection results. The sample node features and sample second-side relationship features are input into the second detection network of the deep learning model, which outputs the sample probability corresponding to the sample binding edge relationship. Finally, the deep learning model is trained using the sample affinity detection results, affinity detection labels, sample probabilities, and sample binding edge relationships to obtain the trained deep learning model.

[0006] According to another aspect of this disclosure, a molecular screening apparatus is provided, comprising: an initial binding conformation subset acquisition module for determining an initial binding conformation subset from a set of binding conformations to be screened, wherein the initial binding conformation subset includes initial binding conformations constructed by an initial molecule and a receptor object to be matched; a binding property evaluation result acquisition module for evaluating the binding properties of the initial binding conformations to obtain a binding property evaluation result; a candidate binding conformation acquisition module for determining candidate binding conformations from the initial binding conformation subset based on the binding property evaluation result; an affinity detection result acquisition module for processing the candidate binding conformations based on a deep learning algorithm to obtain an affinity detection result; and a target molecule acquisition module for screening target molecules that match the receptor object to be matched from the candidate binding conformations based on the affinity detection result.

[0007] According to another aspect of this disclosure, a training apparatus for a deep learning model is provided, comprising: a training sample acquisition module for acquiring training samples, the training samples including sample binding conformations and affinity detection labels, the sample binding conformation including a sample binding conformation graph composed of sample molecules and sample receptors, the sample binding conformation graph including sample atomic nodes of sample molecules, sample atomic edge relationships between sample atomic nodes, sample receptor atomic nodes constituting a sample receptor object, and sample receptor atomic edge relationships between sample receptor atomic nodes; the sample binding conformation graph also includes sample binding edge relationships between sample atomic nodes and sample receptor atomic nodes; and a sample feature extraction module for inputting the sample binding conformation graph into a feature extraction network of the deep learning model, the input... The system comprises: a sample node feature, a first-side relationship feature corresponding to the sample atomic edge relationship or the sample receptor atomic edge relationship, and a second-side relationship feature corresponding to the sample binding edge relationship; a sample affinity detection result acquisition module, which inputs the sample node feature and the first-side relationship feature into the first detection network of the deep learning model and outputs the sample affinity detection result; a sample probability acquisition module, which inputs the sample node feature and the second-side relationship feature into the second detection network of the deep learning model and outputs the sample probability corresponding to the sample binding edge relationship; and a training module, which uses the sample affinity detection result, affinity detection label, sample probability, and sample binding edge relationship to train the deep learning model and obtain the trained deep learning model.

[0008] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method provided according to an embodiment of this disclosure.

[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform a method provided according to an embodiment of this disclosure.

[0010] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method provided according to embodiments of this disclosure.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0012] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0013] Figure 1 This illustration schematically shows an exemplary system architecture to which molecular screening methods and apparatus can be applied according to embodiments of the present disclosure;

[0014] Figure 2 A flowchart illustrating a molecular screening method according to an embodiment of the present disclosure is shown schematically.

[0015] Figure 3 This diagram illustrates an application scenario of the molecular screening method according to embodiments of the present disclosure.

[0016] Figure 4 A schematic diagram of an affinity detection model according to an embodiment of the present disclosure is shown.

[0017] Figure 5 A flowchart illustrating a method for training a deep learning model according to an embodiment of the present disclosure is shown schematically.

[0018] Figure 6 A schematic diagram illustrating the principle of a deep learning model according to an embodiment of the present disclosure is shown.

[0019] Figure 7 A block diagram of a molecular screening apparatus according to an embodiment of the present disclosure is shown schematically;

[0020] Figure 8 A block diagram schematically illustrates a training apparatus for a deep learning model according to an embodiment of the present disclosure; and

[0021] Figure 9 A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0022] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0023] In the technical solution disclosed herein, the acquisition, storage, and application of user personal information comply with the provisions of relevant laws and regulations, necessary confidentiality measures have been taken, and there is no violation of public order and good morals.

[0024] A drug virtual screening system is a data analysis system that uses computer simulation technology to analyze a large-scale compound library to predict the affinity of drug molecules for binding with proteins. Based on this system, potential active molecules that match the target protein molecule can be rapidly screened. However, methods for constructing drug virtual screening systems typically suffer from slow molecular screening speed and low accuracy, making them difficult to apply to screening large-scale libraries of molecules.

[0025] This disclosure provides a molecular screening method, training method, apparatus, electronic device, storage medium, and computer program product. The molecular screening method includes: determining an initial subset of binding conformations from a set of binding conformations to be screened, wherein the initial subset of binding conformations includes initial binding conformations constructed by an initial molecule and a receptor object to be matched; evaluating the binding properties of the initial binding conformations to obtain binding property evaluation results; determining candidate binding conformations from the initial subset of binding conformations based on the binding property evaluation results; processing the candidate binding conformations using a deep learning algorithm to obtain affinity detection results; and screening target molecules that match the receptor object to be matched from the candidate binding conformations based on the affinity detection results.

[0026] Figure 1 An exemplary system architecture for applying molecular screening methods and apparatus according to embodiments of this disclosure is illustrated.

[0027] It is important to note that Figure 1 The examples shown are merely examples of system architectures applicable to embodiments of this disclosure, intended to help those skilled in the art understand the technical content of this disclosure. They do not imply that embodiments of this disclosure cannot be used in other devices, systems, environments, or scenarios. For instance, in another embodiment, an exemplary system architecture to which molecular screening methods and apparatus can be applied may include a terminal device. However, the terminal device can implement the molecular screening methods and apparatus provided in the embodiments of this disclosure without interacting with a server.

[0028] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0029] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (for example only).

[0030] Terminal devices 101, 102, and 103 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0031] Server 105 can be a server that provides various services, such as a backend management server that supports the content browsed by users using terminal devices 101, 102, and 103 (for example only). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0032] It should be noted that the molecular screening method provided in this embodiment can generally be executed by server 105. Correspondingly, the molecular screening device provided in this embodiment can generally be located in server 105. The molecular screening method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the molecular screening device provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.

[0033] Alternatively, the molecular screening method provided in this embodiment of the disclosure can also be executed by terminal devices 101, 102, or 103. Correspondingly, the molecular screening device provided in this embodiment of the disclosure can also be disposed in terminal devices 101, 102, or 103.

[0034] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0035] Figure 2 A flowchart illustrating a molecular screening method according to an embodiment of the present disclosure is shown schematically.

[0036] like Figure 2 As shown, the molecular screening method includes operations S210 to S250.

[0037] In operation S210, an initial binding conformation subset is determined from the set of binding conformations to be screened, wherein the initial binding conformation subset includes the initial binding conformations constructed by the initial molecule and the receptor object to be matched.

[0038] In operation S220, the binding properties of the initial binding conformation are evaluated, and the binding property evaluation results are obtained.

[0039] In operation S230, candidate binding conformations are determined from the initial binding conformation subset based on the binding attribute evaluation results.

[0040] In operation S240, candidate binding conformations are processed based on deep learning algorithms to obtain affinity detection results.

[0041] In the S250 process, target molecules that match the receptor target are selected from candidate binding conformations based on affinity test results.

[0042] According to embodiments of this disclosure, the binding conformations to be screened in the set of binding conformations to be screened can be binding conformations based on the binding of the molecule to be screened and the receptor object. The molecule to be screened can be molecular data in a library of molecules to be screened. The binding conformations to be screened can be obtained, for example, by processing the molecule to be screened and the receptor object through a sampling algorithm, or they can be obtained based on other methods. The embodiments of this disclosure do not limit the specific method of obtaining the binding conformations to be screened.

[0043] According to embodiments of this disclosure, an initial subset of binding conformations can be determined from the set of binding conformations to be screened in any manner. For example, the initial subset of binding conformations can be determined based on random sampling, but it is not limited to this. It can also be determined based on affinity assessment of the binding conformations to be screened. Alternatively, it can be determined based on detecting a screening operation from a user. Embodiments of this disclosure do not limit the specific method by which the initial subset of binding conformations is determined from the set of binding conformations to be screened.

[0044] According to embodiments of this disclosure, the receptor object to be matched may be protein data, or it may be other receptor objects with atomic composition. Embodiments of this disclosure do not limit the specific type of receptor object.

[0045] According to embodiments of this disclosure, evaluating the binding properties of an initial binding conformation can be performed on the docking region between the initial molecule and the receptor in the initial binding conformation. The binding property evaluation can, for example, assess binding properties such as the interaction force and atomic distance of the docking region, obtaining the binding property evaluation results for each initial binding conformation. Therefore, based on the binding property evaluation results, candidate binding conformations that meet the actual requirements can be determined from a subset of initial binding conformations, at least removing some initial binding conformations that do not meet the binding property evaluation requirements, thereby reducing the amount of data processing required for subsequent affinity testing. Simultaneously, it avoids identifying target molecules from initial binding conformations that do not meet the binding property evaluation requirements, improving the accuracy of subsequent target molecule screening.

[0046] It should be noted that the embodiments disclosed herein do not limit the specific evaluation method or the specific type of the combined attribute evaluation, and those skilled in the art can make selections according to actual needs.

[0047] It should be understood that the embodiments of this disclosure do not limit the specific number of candidate binding conformations, and those skilled in the art can design according to actual needs.

[0048] According to embodiments of this disclosure, processing candidate binding conformations based on deep learning algorithms can involve inputting the candidate binding conformations into a deep learning model to obtain affinity detection results corresponding to the candidate binding conformations. The deep learning model can be constructed using any type of deep learning algorithm, such as a multilayer perceptron algorithm or a graph neural network algorithm. Embodiments of this disclosure do not limit the specific type of deep learning algorithm; those skilled in the art can select one according to actual needs.

[0049] According to embodiments of this disclosure, affinity detection results can be used to characterize the affinity level between candidate molecules and receptor objects in candidate binding conformations. Furthermore, through the affinity detection results corresponding to each candidate binding conformation, target candidate binding conformations that meet the requirements can be determined from the candidate binding conformations, and the molecules constituting the target candidate binding conformations can be identified as target molecules.

[0050] According to embodiments of this disclosure, by determining an initial subset of binding conformations from the set of binding conformations to be screened, and further determining candidate binding conformations based on the binding property evaluation results of each of the initial binding conformations, the selected candidate binding conformations can meet the binding property evaluation requirements while reducing the amount of data processing required for subsequent affinity testing, saving computational overhead, and improving molecular screening efficiency. Simultaneously, it can avoid identifying target molecules from initial binding conformations that do not meet the binding property evaluation requirements, meeting users' personalized screening needs and improving the accuracy of target molecule screening.

[0051] The following refers to specific embodiments. Figure 3 and Figure 4 right Figure 2 The method shown will be further explained.

[0052] According to embodiments of this disclosure, determining an initial subset of binding conformations from a set of binding conformations to be screened may include: processing the binding conformations to be screened in the set of binding conformations to be screened based on a scoring function to obtain a screening affinity detection result; and determining an initial subset of binding conformations from the set of binding conformations to be screened based on the screening affinity detection result corresponding to each binding conformation to be screened.

[0053] According to embodiments of this disclosure, the scoring function can be determined in any manner, such as a scoring function determined based on a physical force field (e.g., Qvina2, Vina, Glide, etc.), or a scoring function determined based on a deep learning algorithm (e.g., Gnina, etc.). This disclosure does not limit the specific type of scoring function to embodiments; those skilled in the art can choose according to actual needs.

[0054] According to embodiments of this disclosure, by processing the binding conformation to be screened based on a scoring function, the affinity properties between the molecule to be screened and the receptor object in the binding conformation to be screened can be initially evaluated, thereby screening initial molecules that are initially matched with the receptor object to be matched, reducing the data processing scale of subsequent deep learning algorithms, and improving the screening accuracy of target molecules.

[0055] According to embodiments of this disclosure, evaluating the binding properties of an initial binding conformation to obtain a binding property evaluation result may include: determining candidate binding positions from the initial binding positions between the initial molecule and the receptor object of the initial binding conformation based on preset binding position information; and evaluating the binding properties of the binding property information corresponding to the candidate binding positions to obtain a binding property evaluation result.

[0056] According to embodiments of this disclosure, preset binding site information can be used to characterize the binding position (or docking position) between the initial molecule and the receptor object in the initial binding conformation. The preset binding site can be determined based on preset binding site parameters, such as preset binding site parameters set by the user. Then, candidate binding sites in the initial binding conformation can be determined based on the preset binding site information, thereby enabling user-defined evaluation of the interaction sites between the initial molecule and the receptor object.

[0057] According to embodiments of this disclosure, a scoring function can be used to evaluate the binding attribute information corresponding to candidate binding sites. This allows for a personalized and fine-grained evaluation of the binding attributes of the initial molecule and the receptor at preset candidate binding sites, thereby improving the accuracy of the evaluation of the binding between the molecule and the receptor. This results in the selected target molecule exhibiting better affinity with the receptor at the preset binding sites.

[0058] In one embodiment of this disclosure, the receptor object may be a protein to be matched, and the candidate binding site in the initial binding conformation may be the position where the initial drug molecule and the protein to be matched are connected at a preset local amino acid site.

[0059] According to embodiments of this disclosure, the binding property information may include at least one of the following: hydrogen bond property information, hydrophobicity property information, and binding distance property information.

[0060] According to embodiments of this disclosure, the combined distance attribute information can be the distance between the candidate atoms and the acceptor atoms that are docked to each other in the candidate binding conformation.

[0061] According to embodiments of this disclosure, by evaluating the binding properties of any one or more of the hydrogen bond properties, hydrophobic properties, and binding distance properties, it is possible to evaluate the binding properties of some binding sites in the binding conformation from a microscopic perspective, thereby achieving fine-grained evaluation of the binding conformation, improving the accuracy of the evaluation between the molecule and the acceptor in the binding conformation, and thus improving the accuracy of subsequent target molecule screening.

[0062] Figure 3 The diagram illustrates an application scenario of the molecular screening method according to an embodiment of the present disclosure.

[0063] like Figure 3 As shown, this application scenario may include a molecular screening system 300, which can be constructed based on the molecular screening method provided in the embodiments of this disclosure.

[0064] The molecular screening system 300 may include a first screening module 310, a second screening module 320, a binding property evaluation module 330, and an affinity detection module 340.

[0065] Users can set the receptor object 301 (e.g., receptor protein molecule) to be matched through the client, and can select some or all of the molecules to be screened from the molecule library as the molecule set to be screened 302. The molecular screening system 300 can obtain the receptor object 301 and the molecule set to be screened 302 through the communication interface, and connect the receptor object 301 and the molecules to be screened in the molecule set to be screened 302 according to the first screening module 310 to form a set of conformations to be screened. The first screening module 310 can be a screening module built based on the Qvina2 scoring function algorithm. Based on the first screening module 310, the binding conformations of the conformations to be screened in the set of conformations to be screened are scored for the first time to obtain the first scoring result. For example, the top 5% of the binding conformations to be screened in the first scoring result can be selected from the set of binding conformations to be screened to obtain the first subset of binding conformations to be screened.

[0066] Accordingly, the second screening module 320 can be used to process the first subset of binding conformations to be screened output by the first screening module 310, thereby performing a second binding conformation score on the binding conformations to be screened in the first subset of binding conformations to be screened, and obtaining a second score result. Based on the second score result, an initial subset of binding conformations is determined from the first subset of binding conformations to be screened. For example, the top 50% of the binding conformations to be screened from the first subset of binding conformations to be screened can be selected as the initial binding conformations to obtain the initial subset of binding conformations.

[0067] It should be noted that the screening ratio parameters (e.g., top 5% and top 50% of the sorted data) of the first screening module 310 and the second screening module 320 can be set by the user through the client to meet the user's actual screening needs. By using the first screening module 310 and the second screening module 320 to respectively implement coarse and precise screening of the set of binding conformations to be processed, the user's personalized needs can be met while reducing the amount of data for binding conformation evaluation and improving the screening speed of subsequent target molecules.

[0068] After obtaining the initial binding conformation subset, it can be determined whether the user has selected the binding property evaluation operation. If the determination result is yes, the initial binding conformation subset is input into the binding property evaluation module 330. Based on the preset binding position information corresponding to the binding property evaluation operation, candidate binding positions are determined from the initial binding positions between the initial molecule and the acceptor object in the initial binding conformation. The binding property information, such as hydrophobicity property information and hydrogen bond property information, corresponding to the candidate binding positions are evaluated. This allows the evaluation of the interaction sites between the initial molecule and the acceptor object based on the preset binding position information obtained from the user's custom parameters, thereby improving the evaluation accuracy for the initial binding conformation.

[0069] The binding property evaluation module 330 can evaluate the binding properties (also known as interaction constraint information) of candidate binding conformations based on the Open Drug Discovery Toolkit (ODDT) toolkit, which includes information on hydrophobicity and hydrogen bonding corresponding to the candidate binding positions (amino acid sites).

[0070] Based on the binding attribute evaluation results obtained by the binding attribute evaluation module 330, one or more candidate binding conformations can be identified from the initial subset of binding conformations. These candidate conformations can then be input into the affinity detection module 340, which is constructed based on a graph neural network algorithm. The affinity detection module 340 can further perform affinity detection on the candidate binding conformations and, based on the affinity detection results, determine the top 40% of candidate binding conformations as target binding conformations. This allows the molecule constituting the target binding conformation to be sent to the client as the target molecule after molecular screening, along with the target binding conformation, thus achieving precise screening of the target molecule.

[0071] In another embodiment of this disclosure, the molecular screening system 300 may further include a data preprocessing module, which can preprocess the protein pdb file associated with the receptor protein to be matched to obtain standardized receptor object data, thereby improving the efficiency and accuracy of subsequent molecular screening.

[0072] By selecting the screening step for determining the initial binding conformation subset in the molecular screening method provided in this embodiment, the molecular screening speed can be adaptively improved. For example, the user can adaptively select the first screening module 310 and the second screening module 320 in the molecular screening system 300 to obtain the initial binding conformation subset, thereby improving the computational rate for determining the target molecule. The screening effect of the molecular screening system 300 can be seen in Table 1.

[0073] Table 1

[0074]

[0075] As shown in Table 1, the molecular screening method provided in this disclosure can realize a one-click virtual molecular screening process for molecular-protein docking, affinity conformation detection, and post-screening based on binding property assessment. Furthermore, the molecular screening method provided in this disclosure also supports high-performance distributed parallel computing, enabling molecular docking and activity prediction of ultra-large-scale small molecule virtual screening libraries for proteins with known target structures.

[0076] According to embodiments of this disclosure, the candidate binding conformation includes a candidate binding conformation diagram, which includes candidate atomic nodes of the candidate molecules constituting the candidate binding conformation, candidate atomic edge relationships between candidate atomic nodes, receptor atomic nodes constituting the receptor object, and receptor atomic edge relationships between receptor atomic nodes.

[0077] According to embodiments of this disclosure, processing candidate binding conformations based on deep learning algorithms to obtain affinity detection results may include: extracting features from candidate atomic nodes and acceptor atomic nodes of the candidate binding conformation graph to obtain node features; extracting features from candidate atomic edge relationships and acceptor atomic edge relationships of the candidate binding conformation graph to obtain first edge relationship features; fusing node features and first edge relationship features to obtain first fused features; fusing first edge relationship features and first fused features to obtain second fused features; and performing affinity detection on the candidate binding conformations based on the first fused features and second fused features to obtain affinity detection results.

[0078] According to embodiments of this disclosure, in the candidate binding conformation diagram, candidate atom nodes can be used to characterize atoms in a candidate molecule, and acceptor atom nodes can characterize acceptor atoms constituting the acceptor object. Candidate atom edge relationships can represent the relationship attributes such as distance and chemical bonds between candidate atoms with related relationships in the candidate molecule. Correspondingly, acceptor atom edge relationships can characterize the relationship attributes such as distance and chemical bonds between acceptor atoms with related relationships in the acceptor object.

[0079] According to embodiments of this disclosure, feature extraction can be performed on candidate atomic nodes and receptor atomic nodes of a candidate binding conformation graph based on a neural network algorithm. For example, node features can be extracted based on graph coding embedding network layers. However, it is not limited to this; feature extraction of candidate atomic nodes and receptor atomic nodes can also be achieved based on image coding. The embodiments of this disclosure do not limit the specific method of obtaining node features, and those skilled in the art can choose according to actual needs.

[0080] According to embodiments of this disclosure, the first edge relationship features and node features can be fused based on a neural network algorithm. For example, the first edge relationship features and node features can be fused based on a graph neural network (GNN) algorithm to obtain the first fused feature. The graph neural network algorithm may include graph attention network algorithm, graph convolutional network algorithm, etc., and the embodiments of this disclosure do not limit the specific type of graph neural network algorithm.

[0081] According to embodiments of this disclosure, a first edge relationship feature and a first fusion feature can be fused based on N graph neural network sublayers, where N is a positive integer. For example, a graph neural network layer can be constructed based on multiple sequentially connected graph neural network sublayers to achieve deep feature fusion of the first edge relationship feature and the first fusion feature, thereby extracting the association attributes between nodes and edge relationships in the candidate combination configuration graph, and thus improving the accuracy of subsequent affinity detection results.

[0082] Figure 4 A schematic diagram of an affinity detection model according to an embodiment of the present disclosure is shown.

[0083] like Figure 4 As shown, the candidate binding conformation diagram 401 may include a candidate molecule image region 4011 representing candidate molecules and a receptor object image region 4012 representing receptor objects. The circular nodes in the candidate molecule image region 4011 may be candidate atom nodes, and candidate atom nodes may have candidate atom edge relationships with each other. The square nodes in the receptor object image region 4012 may be receptor atom nodes, and receptor atom nodes may have receptor atom edge relationships with each other.

[0084] like Figure 4 As shown, the affinity detection model 400 may include a feature extraction network 410 and a first detection network 420. The feature extraction network 410 may include a node feature extraction layer 411 and a first edge relationship feature extraction layer 412. The node feature extraction layer 411 and the first edge relationship feature extraction layer 412 may be constructed based on graph embedding network layers. The first detection network 420 may include a first fusion layer 421, a second fusion layer 422, and an affinity detection layer 423.

[0085] Feature extraction is performed on candidate atomic nodes and acceptor atomic nodes of the candidate binding conformation graph, as shown in the figure. This can be achieved by inputting the candidate binding conformation graph 401 into the node feature extraction layer 411 and outputting node features N410. Feature extraction is also performed on candidate atomic edge relationships and acceptor atomic edge relationships of the candidate binding conformation graph. This can be achieved by inputting the candidate binding conformation graph 401 into the first edge relationship feature extraction layer 412 and outputting first edge relationship features B410.

[0086] The first edge relationship feature and node feature can be fused by inputting the first edge relationship feature B410 and node feature N410 into the first fusion layer 421 to obtain the first fused feature R410. The first fusion layer 421 can be constructed based on a graph neural network algorithm.

[0087] The first edge relationship feature and the first fusion feature are fused to obtain the second fusion feature. For example, the first edge relationship feature B410 and the first fusion feature R410 can be input into the second fusion layer 422 to output the second fusion feature R420. The second fusion layer 422 can contain N graph neural network fusion sub-layers, which can be constructed based on the following formulas (1) to (3).

[0088]

[0089]

[0090]

[0091] In formulas (1) to (3), In the candidate combination conformation diagram, e represents the feature corresponding to node v. vu This represents the first edge relationship feature between node v and node u in the candidate combined configuration graph. Let represent the feature corresponding to node u associated with node v, where and It can be the (k-1)th feature output by the (k-1)th graph neural network fusion sublayer of the second fusion layer 422.

[0092] MLP() represents the Multilayer Perceptron Algorithm, Sum Agg () indicates the cumulative aggregation algorithm. The feature associated with node v is represented by the k-th feature output from the k-th fusion sublayer of the graph neural network, and the feature corresponding to node v. `Linear()` indicates a linear regression algorithm. `u` represents a candidate atomic edge relationship or a recipient atomic edge relationship, and `v` represents a candidate atomic node or a recipient atomic node. Here, k is greater than 1 and less than or equal to N.

[0093] It should be understood that, when k=1, the input to the first graph neural network fusion sublayer... and It can be the first fusion feature R410. That is, the first fusion feature R410 can be represented by the feature vector corresponding to the node. When k=N, the k-th feature output by the Nth graph neural network fusion sublayer. The second fusion feature R420 can be included in the output of the second fusion layer 422.

[0094] According to embodiments of this disclosure, affinity detection is performed on candidate binding conformations based on a first fusion feature and a second fusion feature to obtain affinity detection results. This may include processing the first fusion feature and the second fusion feature based on a graph neural network algorithm to obtain affinity detection results.

[0095] likeFigure 4 As shown, the first fusion feature R410 and the second fusion feature R420 can be input into the affinity detection layer 423 constructed based on the graph neural network algorithm, and the affinity detection result 402 can be output.

[0096] According to embodiments of this disclosure, the first fusion feature and the second fusion feature can be input into the affinity detection layer constructed based on the graph neural network algorithm. By fully fusing the node features and the first edge relationship features in the candidate binding conformation, the affinity properties between the candidate molecule and the receptor object in the candidate binding conformation can be accurately detected, thereby improving the accuracy of subsequent screening of target molecules.

[0097] According to embodiments of this disclosure, a virtual screening platform for small molecule drugs targeting direct protein targets can be constructed based on the molecular screening methods provided in the embodiments of this disclosure, so that relevant drug developers can identify drug molecules that match the target protein from a large-scale library of molecules to be screened.

[0098] In one embodiment of this disclosure, a user can establish a communication connection with a small molecule drug virtual screening platform via a client. Matching target ligands (target molecules) are determined for a protein named NIK (NF-KB-inducing kinase). The user can customize the binding mode between the ligand (molecule) to be screened and the NIK protein, such as setting a binding mode that forms hydrogen bonds with specific amino acids in the NIK protein. The candidate binding conformations obtained from the initial screening are then input into the affinity detection module built based on a deep learning algorithm in the small molecule drug virtual screening platform to obtain multiple target molecules.

[0099] Based on the molecular screening method provided in this disclosure, it is possible to identify the top 100 target molecules from a set of 150,000 molecules to be screened. After analysis, it can be determined that 46 active molecules can be successfully recalled from the 100 target molecules, and the molecular screening time is reduced and the molecular screening efficiency is improved.

[0100] In another embodiment of this disclosure, an affinity detection module can be constructed based on a graph neural network algorithm. This module can then determine the top 100 candidate binding conformations from multiple candidate binding conformations as the target binding conformations. By evaluating the enrichment coefficient of the target site position of the target binding conformation, it can be determined that the docking accuracy between the target molecule obtained according to the molecular screening method provided in this disclosure and the receptor object is significantly improved, while ensuring a shorter molecular screening time and improving the overall efficiency of molecular screening.

[0101] Table 2

[0102]

[0103] In Table 2, F0.1 a F1 represents the enrichment coefficient calculated based on the top 0.1% of target binding conformations in the candidate binding conformations. b This represents the enrichment coefficient calculated based on the top 1% of the target binding conformations ranked among the candidate binding conformations.

[0104] The embodiments of this disclosure also provide a method for training a deep learning model. The deep learning model obtained based on the training method provided in the embodiments of this disclosure can be applied to the molecular screening method provided in the above embodiments.

[0105] Figure 5 A flowchart illustrating a method for training a deep learning model according to an embodiment of the present disclosure is shown.

[0106] like Figure 5 As shown, the training method for this deep learning model includes operations S510 to S550.

[0107] During operation S510, training samples are acquired. The training samples include sample binding conformations and affinity detection tags. The sample binding conformations include a sample binding conformation diagram composed of sample molecules and sample receptor objects. The sample binding conformation diagram includes sample atomic nodes of sample molecules, sample atomic edge relationships between sample atomic nodes, sample receptor atomic nodes constituting sample receptor objects, and sample receptor atomic edge relationships between sample receptor atomic nodes. The sample binding conformation diagram also includes sample binding edge relationships between sample atomic nodes and sample receptor atomic nodes.

[0108] In operation S520, the sample-binding conformation diagram is input into the feature extraction network of the deep learning model, and the output is the sample node features, the sample first-side relationship features corresponding to the sample atomic edge relationship or the sample receptor atomic edge relationship, and the sample second-side relationship features corresponding to the sample binding edge relationship.

[0109] In operation S530, the sample node features and the first edge relationship features of the sample are input into the first detection network of the deep learning model, and the sample affinity detection results are output.

[0110] In operation S540, the sample node features and sample second edge relationship features are input into the second detection network of the deep learning model, and the output is the sample probability corresponding to the sample combined with the edge relationship.

[0111] In operating the S550, a deep learning model is trained using the sample affinity detection results, affinity detection labels, sample probabilities, and sample combination edge relationships to obtain the trained deep learning model.

[0112] According to embodiments of this disclosure, the sample binding edge relationship can be an edge relationship that characterizes the binding between a sample molecule and a sample receptor object.

[0113] It should be noted that the technical terms (such as sample molecules, sample receptor objects, etc.) involved in the training method of the deep learning model provided in the embodiments of this disclosure have the same or corresponding technical attributes as the technical terms (such as initial molecules, receptor objects, etc.) involved in the molecular screening method in the above embodiments, and the embodiments of this disclosure will not repeat them here.

[0114] It should be noted that the embodiments of this disclosure do not limit the number of sample binding conformation diagrams. For example, the same sample binding conformation diagram can be used to characterize the internal edge relationships of the sample molecule and the sample receptor object, or the sample binding conformation diagram can include a first sample binding conformation diagram and a second sample binding conformation diagram. The first sample binding conformation diagram can represent the internal molecular structure of the sample molecule and the sample receptor object, and the first sample binding conformation diagram can represent the docking relationship between the sample molecule and the sample receptor object.

[0115] Figure 6 A schematic diagram of a deep learning model according to an embodiment of the present disclosure is shown.

[0116] like Figure 6 As shown, the deep learning model 600 may include a feature extraction network 610, a first detection network 620, and a second detection network 630. The feature extraction network 610 may include a node feature extraction layer 611, a first edge relationship feature extraction layer 612, and a second edge relationship feature extraction layer 613. The first detection network 620 may include a first fusion layer 621, a second fusion layer 622, and an affinity detection layer 623. The second detection network 630 may include a third fusion layer 631, a fourth fusion layer 632, a sample probability detection layer 633, and a mixed density network layer 634.

[0117] The training samples may include a sample binding conformation diagram 601. The sample binding conformation diagram 601 may include a sample molecule image region 6011 representing the sample molecule, and a sample receptor object image region 6012 representing the sample receptor object. Circular nodes in the sample molecule image region 6011 may be sample atom nodes, and these sample atom nodes may have candidate atom edge relationships. Square nodes in the sample receptor object image region 6012 may be sample receptor atom nodes, and these sample receptor atom nodes may have receptor atom edge relationships. The dashed line between the sample molecule image region 6011 and the sample receptor object image region 6012 can represent the sample binding edge relationships between the sample atom nodes and the sample receptor atom nodes.

[0118] likeFigure 6 As shown, the sample combined with the conformational graph is input into the feature extraction network of the deep learning model. For example, the sample combined with the conformational graph 601 can be input into the node feature extraction layer 611, the first edge relationship feature extraction layer 612, and the second edge relationship feature extraction layer 613 respectively, and the sample node feature N610, the sample first edge relationship feature B610, and the sample second edge relationship feature B620 are output.

[0119] like Figure 6 As shown, the sample node features and the sample first edge relationship features are input into the first detection network of the deep learning model. For example, the sample node features N610 and the sample first edge relationship features B610 can be input into the first fusion layer 621 of the first detection network 620, outputting the sample first fusion feature R610. The sample first fusion feature R610 and the sample first edge relationship feature B610 are input into the second fusion layer 622, resulting in the sample second fusion feature R620. The sample second fusion feature R620 and the sample first fusion feature R610 are input into the affinity detection layer 623, outputting the sample affinity detection result.

[0120] like Figure 6 As shown, the sample node features and sample second-side relationship features are input into the second detection network of the deep learning model. For example, the sample node features N610 and sample second-side relationship features B620 can be input into the third fusion layer 631 to output the sample third fusion feature R630. The sample third fusion feature R630 and sample second-side relationship features B620 are input into the fourth fusion layer 632 to output the sample fourth fusion feature R640. The sample third fusion feature R630 and sample fourth fusion feature R640 are input into the sample probability detection layer 633 to output the sample probability.

[0121] According to embodiments of this disclosure, a deep learning model can be jointly trained using the loss value between the sample affinity detection result and the affinity detection label, as well as the fitting result of the sample probability and the sample binding edge relationship. This allows the deep learning model to fully learn the internal structure of the sample molecule, the internal structure of the sample receptor object, and the binding attributes or docking patterns between the sample molecule and the sample receptor object, thereby improving the detection accuracy of the trained deep learning model for affinity detection results.

[0122] According to embodiments of this disclosure, training a deep learning model using sample affinity detection results, affinity detection labels, sample parameters, and sample binding edge relationships may include: processing the sample affinity detection results and affinity detection labels based on a loss function to obtain a loss value; updating the parameters of the current mixing density function based on sample probabilities to obtain an updated mixing density function; processing the sample binding edge relationships based on the updated mixing density function to obtain sample edge distance distribution values; and adjusting the model parameters of the deep learning model according to the loss value and sample edge distance distribution values ​​to obtain a trained deep learning model.

[0123] like Figure 6 As shown, the sample probability output by the sample probability detection layer 633 can be input into the mixed density network layer 634, which outputs the sample edge distance distribution value 603. The mixed density network layer 634 can be constructed based on the Mixture Density Networks (MDN) algorithm.

[0124] By using the sample edge distance distribution value 603 and the loss value between the sample affinity detection result and the affinity detection label, the deep learning model is jointly trained. This allows the deep learning model to fully learn the probability distribution of sample binding edge relationships in the sample binding conformation graph, and further learn the binding attribute information between sample atomic nodes and sample receptor atomic nodes, thereby further improving the accuracy of the trained deep learning model for affinity detection.

[0125] It should be noted that, Figure 6 The second detection network 630 in the deep learning model 600 provided can be constructed using the same or corresponding algorithm as the first detection network 620. For example, the first fusion layer 621 and the third fusion layer 631 can be constructed using the same graph neural network algorithm. The second fusion layer 622 and the fourth fusion layer 632 can be constructed using the N graph neural network fusion sub-layers provided in the above embodiments. The sample probability detection layer 633 and the affinity detection layer can be constructed using graph neural network algorithms.

[0126] The first detection network in the deep learning model provided in this embodiment can be constructed based on the same or corresponding algorithm as the affinity detection model provided in the above embodiment, and the applicant will not elaborate further here.

[0127] Figure 7 A block diagram of a molecular screening apparatus according to an embodiment of the present disclosure is shown schematically.

[0128] like Figure 7As shown, the molecular screening device 700 includes: an initial binding conformation subset acquisition module 710, a binding property evaluation result acquisition module 720, a candidate binding conformation acquisition module 730, an affinity detection result acquisition module 740, and a target molecule acquisition module 750.

[0129] Initial binding conformation subset acquisition module 710 is used to determine an initial binding conformation subset from a set of binding conformations to be screened, wherein the initial binding conformation subset includes initial binding conformations constructed by an initial molecule and a receptor object to be matched.

[0130] The module 720, which combines the property evaluation results, is used to evaluate the binding properties of the initial binding conformation and obtain the binding property evaluation results.

[0131] Candidate binding conformation acquisition module 730 is used to determine candidate binding conformations from an initial subset of binding conformations based on the results of binding property evaluation.

[0132] The affinity detection result acquisition module 740 is used to process candidate binding conformations based on deep learning algorithms to obtain affinity detection results.

[0133] The target molecule acquisition module 750 is used to screen target molecules that match the receptor object to be matched from candidate binding conformations based on affinity detection results.

[0134] According to embodiments of this disclosure, the candidate binding conformation includes a candidate binding conformation diagram, which includes candidate atomic nodes of the candidate molecules constituting the candidate binding conformation, candidate atomic edge relationships between candidate atomic nodes, receptor atomic nodes constituting the receptor object, and receptor atomic edge relationships between receptor atomic nodes.

[0135] The affinity detection result acquisition module includes: a node feature acquisition unit, a first edge relationship feature acquisition unit, a first fusion unit, a second fusion unit, and an affinity detection result acquisition unit.

[0136] The node feature acquisition unit is used to extract features from candidate atomic nodes and acceptor atomic nodes in the candidate binding conformation diagram to obtain node features.

[0137] The first-side relation feature acquisition unit is used to extract features from the candidate atomic side relations and acceptor atomic side relations of the candidate binding conformation diagram to obtain the first-side relation features.

[0138] The first fusion unit is used to fuse node features and first edge relationship features to obtain the first fused feature.

[0139] The second fusion unit is used to fuse the first edge relation feature and the first fusion feature to obtain the second fusion feature.

[0140] The affinity detection result acquisition unit is used to perform affinity detection on candidate binding conformations based on the first fusion feature and the second fusion feature, and obtain the affinity detection result.

[0141] According to embodiments of this disclosure, the affinity detection result acquisition unit includes an affinity detection result acquisition subunit.

[0142] The affinity detection result acquisition sub-unit is used to process the first fusion feature and the second fusion feature based on the graph neural network algorithm to obtain the affinity detection result.

[0143] According to embodiments of this disclosure, the initial binding conformation subset acquisition module includes: a unit for obtaining the affinity detection results to be screened and a unit for obtaining the initial binding conformation subset.

[0144] The unit for obtaining the affinity test results is used to process the binding conformations to be screened in the set of binding conformations to be screened based on a scoring function, and obtain the affinity test results to be screened.

[0145] The initial binding conformation subset acquisition unit is used to determine the initial binding conformation subset from the set of binding conformations to be screened based on the test results of the affinity of each binding conformation to be screened.

[0146] According to embodiments of this disclosure, the module for obtaining the combined attribute evaluation results includes: a candidate combination position obtaining unit and a combination attribute evaluation result obtaining unit.

[0147] The candidate binding site acquisition unit is used to determine candidate binding sites from the initial binding sites between the initial molecule and the receptor object in the initial binding conformation based on preset binding site information.

[0148] The combined attribute evaluation result acquisition unit is used to evaluate the binding attribute information corresponding to the candidate binding position and obtain the binding attribute evaluation result.

[0149] According to embodiments of this disclosure, the binding property information includes at least one of the following: hydrogen bond property information, hydrophobicity property information, and binding distance property information.

[0150] Figure 8 A block diagram of a training apparatus for a deep learning model according to an embodiment of the present disclosure is shown schematically.

[0151] like Figure 8 As shown, the training device 800 for the deep learning model includes: a training sample acquisition module 810, a sample feature extraction module 820, a sample affinity detection result acquisition module 830, a sample probability acquisition module 840, and a training module 850.

[0152] The training sample acquisition module 810 is used to acquire training samples, which include sample binding conformation and affinity detection tags. The sample binding conformation includes a sample binding conformation diagram composed of sample molecules and sample receptors. The sample binding conformation diagram includes sample atomic nodes of sample molecules, sample atomic edge relationships between sample atomic nodes, sample receptor atomic nodes constituting the sample receptor object, and sample receptor atomic edge relationships between sample receptor atomic nodes. The sample binding conformation diagram also includes sample binding edge relationships between sample atomic nodes and sample receptor atomic nodes.

[0153] The sample feature extraction module 820 is used to input the sample combination conformation graph into the feature extraction network of the deep learning model, and output the sample node features, the sample first side relationship features corresponding to the sample atomic edge relationship or the sample receptor atomic edge relationship, and the sample second side relationship features corresponding to the sample combination edge relationship.

[0154] The sample affinity detection result acquisition module 830 is used to input the sample node features and the sample first edge relationship features into the first detection network of the deep learning model and output the sample affinity detection result.

[0155] The sample probability acquisition module 840 is used to input the sample node features and the sample second edge relationship features into the second detection network of the deep learning model, and output the sample probability corresponding to the sample combined edge relationship.

[0156] Training module 850 is used to train a deep learning model using sample affinity detection results, affinity detection labels, sample probabilities, and sample combination edge relationships, resulting in a trained deep learning model.

[0157] According to embodiments of this disclosure, the training module includes: a loss value acquisition unit, an update unit, a sample edge distance distribution value acquisition unit, and a training unit.

[0158] The loss value acquisition unit is used to process the sample affinity detection results and affinity detection labels based on the loss function to obtain the loss value.

[0159] The update unit is used to update the parameters of the current mixture density function based on the sample probability, so as to obtain the updated mixture density function.

[0160] The sample edge distance distribution value acquisition unit is used to process the sample binding edge relationship based on the updated mixing density function to obtain the sample edge distance distribution value.

[0161] The training unit is used to adjust the model parameters of the deep learning model based on the loss value and the sample edge distance distribution value to obtain the trained deep learning model.

[0162] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0163] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described above.

[0164] According to embodiments of the present disclosure, a non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to cause a computer to perform the method described above.

[0165] According to an embodiment of this disclosure, a computer program product includes a computer program that, when executed by a processor, implements the method described above.

[0166] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0167] like Figure 9 As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 902 or a computer program loaded from storage unit 908 into random access memory (RAM) 903. RAM 903 may also store various programs and data required for the operation of device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.

[0168] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0169] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as molecular screening methods and deep learning model training methods. For example, in some embodiments, the molecular screening methods and deep learning model training methods can be implemented as computer software programs, which are tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the molecular screening methods and deep learning model training methods described above can be performed. Alternatively, in other embodiments, the computing unit 901 may be configured by any other suitable means (e.g., by means of firmware) to perform molecular screening methods or deep learning model training methods.

[0170] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0171] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0172] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0173] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0174] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0175] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.

[0176] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0177] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for training a deep learning model, comprising: Training samples are obtained, including sample binding conformation and affinity detection tags. The sample binding conformation includes a sample binding conformation diagram composed of sample molecules and sample receptor objects. The sample binding conformation diagram includes sample atomic nodes of the sample molecules, sample atomic edge relationships between sample atomic nodes, sample receptor atomic nodes constituting the sample receptor object and sample receptor atomic node relationships between sample receptor atomic nodes, and sample binding edge relationships between sample atomic nodes and sample receptor atomic nodes. The sample atomic nodes represent atoms in the sample molecules, the sample receptor atomic nodes represent receptor atoms constituting the sample receptor object, the sample atomic edge relationships represent the distance and chemical bonds between associated atoms, the sample receptor atomic edge relationships represent the distance and chemical bonds between associated receptor atoms, and the sample binding edge relationships represent the binding situation between the sample molecules and the sample receptor objects. The sample molecules are drug molecules, and the sample receptor objects are proteins. The sample is combined with the conformational graph and input into the feature extraction network of the deep learning model. The output is the sample node features, the sample first edge relationship features corresponding to the sample atomic edge relationship or the sample receptor atomic edge relationship, and the sample second edge relationship features corresponding to the sample combination edge relationship. The sample node features and the sample first edge relationship features are input into the first detection network of the deep learning model, and the sample affinity detection result is output. The sample affinity detection result characterizes the affinity level between the sample molecule and the sample receptor object in the sample binding conformation. The sample node features and the sample second edge relationship features are input into the second detection network of the deep learning model, and the sample probability corresponding to the sample combined with the edge relationship is output. The sample affinity detection results and the affinity detection labels are processed based on a loss function to obtain a loss value; The parameters of the current mixing density function are updated based on the sample probabilities to obtain the updated mixing density function; The sample edge relationships are processed based on the updated mixing density function to obtain the sample edge distance distribution values; and The model parameters of the deep learning model are adjusted based on the loss value and the sample edge distance distribution value to obtain the trained deep learning model.

2. A molecular screening method, comprising: An initial binding conformation subset is determined from the set of binding conformations to be screened, wherein the initial binding conformation subset includes an initial binding conformation constructed by an initial molecule and a receptor object to be matched; The binding properties of the initial binding conformation are evaluated to obtain the binding property evaluation results; Based on the binding attribute evaluation results, candidate binding conformations are determined from the initial binding conformation subset; The candidate binding conformations are processed based on a trained deep learning model to obtain affinity detection results, wherein the trained deep learning model is determined based on the method described in claim 1; and Based on the affinity test results, target molecules that match the receptor object to be matched are screened from the candidate binding conformations.

3. The method according to claim 2, wherein, The candidate binding conformation includes a candidate binding conformation diagram, which includes candidate atomic nodes of candidate molecules constituting the candidate binding conformation, candidate atomic edge relationships between the candidate atomic nodes, receptor atomic nodes constituting the receptor object, and receptor atomic edge relationships between the receptor atomic nodes. The process of processing the candidate binding conformations based on a deep learning model to obtain affinity detection results includes: Feature extraction is performed on the candidate atomic nodes and the receptor atomic nodes of the candidate binding conformation diagram to obtain node features; Feature extraction is performed on the candidate atom edge relationships and the acceptor atom edge relationships of the candidate binding conformation diagram to obtain the first edge relationship features; The node features and the first edge relationship features are fused to obtain the first fused feature; By fusing the first edge relationship feature and the first fused feature, a second fused feature is obtained; and Affinity detection is performed on the candidate binding conformations based on the first fusion feature and the second fusion feature to obtain the affinity detection result.

4. The method according to claim 3, wherein, The step of performing affinity detection on the candidate binding conformation based on the first fusion feature and the second fusion feature to obtain the affinity detection result includes: The affinity detection result is obtained by processing the first fusion feature and the second fusion feature based on the graph neural network algorithm.

5. The method according to claim 2, wherein, The process of determining an initial subset of binding conformations from the set of binding conformations to be screened includes: The selection of binding conformations in the set of conformations to be selected is processed based on a scoring function to obtain the affinity detection results; and Based on the affinity test results corresponding to each of the selected binding conformations, the initial binding conformation subset is determined from the set of selected binding conformations.

6. The method according to claim 2, wherein, The binding property evaluation of the initial binding conformation to obtain the binding property evaluation results includes: Based on preset binding site information, candidate binding sites are determined from the initial binding sites between the initial molecule and the receptor object in the initial binding conformation; and The binding attribute information corresponding to the candidate binding position is evaluated to obtain the binding attribute evaluation result.

7. The method according to claim 6, wherein, The combined attribute information includes at least one of the following: Hydrogen bond properties, hydrophobicity properties, and binding distance properties.

8. A molecular screening device, comprising: An initial binding conformation subset acquisition module is used to determine an initial binding conformation subset from a set of binding conformations to be screened, wherein the initial binding conformation subset includes an initial binding conformation constructed by an initial molecule and a receptor object to be matched. The module for obtaining the combined property evaluation results is used to evaluate the binding properties of the initial binding conformation and obtain the binding property evaluation results. A candidate binding conformation acquisition module is used to determine candidate binding conformations from the initial binding conformation subset based on the binding property evaluation results; An affinity detection result acquisition module is used to process the candidate binding conformations based on a trained deep learning model to obtain an affinity detection result, wherein the trained deep learning model is determined based on the method described in claim 1; and The target molecule acquisition module is used to screen out target molecules that match the receptor object to be matched from the candidate binding conformations based on the affinity detection results.

9. The apparatus according to claim 8, wherein, The candidate binding conformation includes a candidate binding conformation diagram, which includes candidate atomic nodes of candidate molecules constituting the candidate binding conformation, candidate atomic edge relationships between the candidate atomic nodes, receptor atomic nodes constituting the receptor object, and receptor atomic edge relationships between the receptor atomic nodes. The affinity detection result acquisition module includes: The node feature acquisition unit is used to extract features from the candidate atomic nodes and the receptor atomic nodes of the candidate binding conformation diagram to obtain node features; The first edge relationship feature acquisition unit is used to extract features from the candidate atomic edge relationships and the receptor atomic edge relationships of the candidate binding conformation diagram to obtain the first edge relationship features; The first fusion unit is used to fuse the node features and the first edge relationship features to obtain the first fused feature; The second fusion unit is used to fuse the first edge relationship feature and the first fusion feature to obtain the second fusion feature; and The affinity detection result acquisition unit is used to perform affinity detection on the candidate binding conformation based on the first fusion feature and the second fusion feature, and obtain the affinity detection result.

10. The apparatus according to claim 9, wherein, The affinity detection result acquisition unit includes: The affinity detection result acquisition subunit is used to process the first fusion feature and the second fusion feature based on the graph neural network algorithm to obtain the affinity detection result.

11. The apparatus according to claim 8, wherein, The initial combination conformation subset acquisition module includes: The unit for obtaining the affinity detection result is used to process the binding conformations to be screened in the set of binding conformations to be screened based on a scoring function to obtain the affinity detection result; and The initial binding conformation subset obtaining unit is used to determine the initial binding conformation subset from the set of binding conformations to be screened based on the test results of the affinity of each of the binding conformations to be screened.

12. The apparatus according to claim 8, wherein, The module for obtaining the combined attribute evaluation results includes: A candidate binding site obtaining unit is configured to determine candidate binding sites from the initial binding sites between the initial molecule and the receptor object in the initial binding conformation, based on preset binding site information; and The combined attribute evaluation result obtaining unit is used to evaluate the binding attribute information corresponding to the candidate binding position and obtain the binding attribute evaluation result.

13. The apparatus according to claim 12, wherein, The combined attribute information includes at least one of the following: Hydrogen bond properties, hydrophobicity properties, and binding distance properties.

14. A training device for a deep learning model, comprising: A training sample acquisition module is used to acquire training samples, which include sample binding conformation and affinity detection tags. The sample binding conformation includes a sample binding conformation diagram composed of sample molecules and sample receptor objects. The sample binding conformation diagram includes sample atomic nodes of the sample molecules, sample atomic edge relationships between the sample atomic nodes, and sample receptor atomic edge relationships between the sample receptor atomic nodes constituting the sample receptor object. The sample binding conformation diagram also includes sample binding edge relationships between the sample atomic nodes and the sample receptor atomic nodes. The sample atomic nodes represent atoms in the sample molecules, the sample receptor atomic nodes represent receptor atoms constituting the sample receptor object, and the sample atomic edge relationships represent the distances and chemical bonds between related atoms. The sample binding edge relationships represent the binding situation between the sample molecules and the sample receptor objects. The sample molecules are drug molecules, and the sample receptor objects represent proteins. The sample feature extraction module is used to input the sample combination conformation graph into the feature extraction network of the deep learning model, and output the sample node features, the sample first edge relationship features corresponding to the sample atomic edge relationship or the sample receptor atomic edge relationship, and the sample second edge relationship features corresponding to the sample combination edge relationship. The sample affinity detection result acquisition module is used to input the sample node features and the sample first edge relationship features into the first detection network of the deep learning model and output the sample affinity detection result, wherein the sample affinity detection result characterizes the affinity level between the sample molecule and the sample receptor object in the sample binding conformation; The sample probability acquisition module is used to input the sample node features and the sample second edge relationship features into the second detection network of the deep learning model, and output the sample probability corresponding to the sample combined edge relationship; as well as The training module is used to train the deep learning model using the sample affinity detection results, the affinity detection labels, the sample probabilities, and the sample binding edge relationships, so as to obtain the trained deep learning model. The training module includes: The loss value acquisition unit is used to process the sample affinity detection result and the affinity detection label based on the loss function to obtain the loss value; The update unit is used to update the parameters of the current mixing density function based on the sample probability, so as to obtain the updated mixing density function; The sample edge distance distribution value acquisition unit is used to process the sample binding edge relationship based on the updated mixing density function to obtain the sample edge distance distribution value; and The training unit is used to adjust the model parameters of the deep learning model according to the loss value and the sample edge distance distribution value to obtain the trained deep learning model.

15. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 7.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 7.

17. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Molecular set generation method and device, terminal and storage medium

    CN114429797A

  • Prediction model training method, binding affinity prediction method, device and equipment

    CN115171776A