Cell-targeted drug preliminary screening model training method, drug preliminary screening method and equipment

By combining molecular fingerprint and topological information based on graph neural network, the data sets are screened and the model is trained, and the problem of insufficient information loss and generalization capabilities in the existing technology is solved, and more efficient initial screening prediction of cellular targeted drugs is achieved.

CN120544722AActive Publication Date: 2025-08-26NANKAI UNIV

Patent Information

Application Number
CN202410210408.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-26
Publication Date
2025-08-26
Estimated Expiration
2044-02-26

AI Technical Summary

Technical Problem

The existing cellular-targeted drug primary screening methods have severe information loss when processing molecules with highly diverse structures and insufficient generalization ability, resulting in low prediction accuracy.

Method used

A regression model based on graph neural network is used to combine molecular fingerprint characteristics and structural topology information, and the data set is screened through distance measurement, and the model is trained to predict the potential cellular-targeted drug activity of the compound.

Benefits of technology

It improves the generalization ability and prediction accuracy of the cell-targeted drug primary screening model, and can more comprehensively retain the molecular characteristics that contribute to drug screening, enhancing the widespread application of the model and the reliability of predicted results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544722A_ABST
    Figure CN120544722A_ABST
Patent Text Reader

Abstract

The invention provides a cell-targeted drug preliminary screening model training method, a drug preliminary screening method and equipment, and the method comprises the steps: obtaining molecular fingerprint feature data corresponding to each compound sample with label information to form a corresponding original data set, and carrying out the data screening of the original data set in a distance measurement mode, and training a regression model based on a graph neural network to learn the structural topology of each compound sample, and training the model as a cell-targeted drug preliminary screening model for predicting whether the compound has potential cell-targeted drug activity or not. According to the application, the comprehensiveness of molecular characterization can be effectively improved in the training process of the cell-targeted drug preliminary screening model, and the generalization ability of the trained cell-targeted drug preliminary screening model can be effectively improved, so that the application universality of the trained cell-targeted drug preliminary screening model can be effectively improved; and the accuracy and the reliability of a potential cell targeted drug activity prediction result predicted by the model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method for training a cell-targeted drug screening model, a drug screening method, and equipment. Background Art

[0002] Compound drug screening is a crucial step in the drug discovery and development process, aiming to identify compounds with therapeutic potential from thousands of potential candidate compounds. Molecular machine learning can be used to improve the efficiency of initial screening for cell-targeted drugs. Looking back at the development of feature engineering for molecular machine learning, early research primarily relied on traditional, hand-crafted features. However, because traditional methods can perform poorly when dealing with highly diverse molecular structures, the demand for more flexible, data-driven approaches has grown. In recent years, with the rise of big data and large models, the field has evolved from initially relying on hand-crafted and heuristically designed features to using more flexible and expressive data-driven features. In recent years, molecular machine learning methods have evolved, using deep learning and self-supervised learning methods to learn data-driven molecular representations from massive amounts of unlabeled molecular data. This is the core of molecular machine learning and currently the mainstream approach for initial screening of cell-targeted drugs.

[0003] However, traditional cell-targeted drug screening methods rely primarily on fixed molecular fingerprints, which can lead to the loss of important information related to drug physiological metabolism, especially when dealing with molecules with highly diverse structures. While data-driven molecular characterization has addressed the shortcomings of molecular fingerprints to some extent, these methods typically focus on modeling molecules based on a single, specific modality, resulting in varying degrees of information loss. Furthermore, traditional cell-targeted drug screening methods, such as self-supervised learning-based molecular characterization methods, lack generalization capabilities.

[0004] Therefore, there is an urgent need to design a training method for cell-targeted drug screening models that can improve the accuracy of cell-targeted drug screening predictions and the generalization ability of the model. Summary of the Invention

[0005] In view of this, the embodiments of the present application provide a cell-targeted drug screening model training method, a drug screening method and an apparatus to eliminate or improve one or more defects in the prior art.

[0006] One aspect of the present application provides a method for training a cell-targeted drug screening model, comprising:

[0007] Acquiring molecular fingerprint feature data corresponding to each compound sample provided with label information to form a corresponding original data set, wherein the label information is used to indicate whether the corresponding compound sample has potential cell-targeted drug activity;

[0008] Using a distance metric to screen each compound sample in the original data set to obtain a corresponding target data set, and dividing a training set from the target data set;

[0009] A graph neural network-based regression model is trained based on the training set so that the graph neural network-based regression model learns the structural topological information of each of the compound samples, and then the graph neural network-based regression model is trained as a cell-targeted drug screening model for predicting whether the compound has potential cell-targeted drug activity based on the molecular fingerprint feature data of the compound.

[0010] In some embodiments of the present application, before obtaining the molecular fingerprint feature data corresponding to each compound sample provided with label information to form a corresponding original data set, the method further includes:

[0011] Obtaining raw information data of each candidate compound used for target cell activation or inhibition and their corresponding label information;

[0012] Performing SMILES string conversion on the original information data of each candidate compound;

[0013] For a candidate compound whose SMILES string conversion is successful and whose generated SMILES string is unique, the SMILES string of the candidate compound is used as a compound sample;

[0014] Candidate compounds that failed to convert SMILES strings were deleted;

[0015] For a candidate compound whose SMILES string conversion is successful but whose generated SMILES string is not unique, one of the SMILES strings of the candidate compound is selected as a compound sample.

[0016] In some embodiments of the present application, the target cells include: tumor cells, immune cells, cardiomyocytes, neuronal cells, endocrine cells, hepatocytes, stem cells or respiratory epithelial cells;

[0017] Wherein, the immune cells include: T cells, B cells, NK cells, myeloid cells or antigen presenting cells.

[0018] In some embodiments of the present application, obtaining molecular fingerprint feature data corresponding to each compound sample provided with label information to form a corresponding original data set includes:

[0019] Based on a preset latent feature space representation method, the molecular fingerprint feature data corresponding to each compound sample with label information is obtained;

[0020] generating an original data set including the molecular fingerprint feature data and the label information corresponding to each of the compound samples;

[0021] The latent feature space representation method includes at least one of an ECFP2 molecular fingerprint representation method, an ECFP4 molecular fingerprint representation method, an ECFP6 molecular fingerprint representation method, and a MACCS Key molecular fingerprint representation method.

[0022] In some embodiments of the present application, the method of using a distance metric to perform data screening on each compound sample in the original data set to obtain a corresponding target data set, and dividing a training set from the target data set, includes:

[0023] Randomly selecting molecular fingerprint feature data of a plurality of compound samples from the original data set as reference anchor points respectively;

[0024] Based on the number of the reference anchor points, a first hyperparameter for representing a radius threshold of a feature space hypersphere, and a second hyperparameter for representing a label difference threshold, clustering the compound samples in the original dataset using a distance metric to screen out heterogeneous samples in the original dataset, thereby obtaining a corresponding target dataset;

[0025] Dividing the target data set into a training set and a test set;

[0026] The distance measurement method includes: a Hamming distance measurement method or a Euclidean distance measurement method.

[0027] In some embodiments of the present application, the training of a graph neural network-based regression model based on the training set, so that the graph neural network-based regression model learns the structural topological information of each of the compound samples, and then training the graph neural network-based regression model into a cell-targeted drug screening model for predicting whether the compound has potential cell-targeted drug activity based on the molecular fingerprint feature data of the compound, comprises:

[0028] Iteratively training a graph neural network-based regression model based on the training set and preset training hyperparameters, so that the graph neural network-based regression model represents nodes used to represent atoms, edges used to represent chemical bonds, and global information containing structural topological information of the compound sample in the form of learnable vectors, and during the iterative training process, screening the potential cell-targeted drug activity prediction results obtained each time by the graph neural network-based regression model according to a preset prediction threshold range;

[0029] The trained graph neural network-based regression model is subjected to a model test based on the test set. If the graph neural network-based regression model passes the model test, the graph neural network-based regression model is used as a cell-targeted drug screening model for predicting whether the compound has potential cell-targeted drug activity based on the molecular fingerprint feature data of the compound.

[0030] In some embodiments of the present application, the graph neural network-based regression model includes: a GCN model and / or an MPNN model.

[0031] Another aspect of the present application provides a method for initial screening of cell-targeted drugs, comprising:

[0032] Obtain molecular fingerprint characteristic data of the target compound;

[0033] The molecular fingerprint feature data of the target compound is input into a preset cell-targeted drug screening model, so that the cell-targeted drug screening model outputs a potential cell-targeted drug activity prediction result indicating whether the target compound has potential cell-targeted drug activity, wherein the cell-targeted drug screening model is pre-trained based on the cell-targeted drug screening model training method.

[0034] In some embodiments of the present application, the cell-targeted drug primary screening method further comprises:

[0035] Outputting a potential cell-targeting drug activity prediction result corresponding to the target compound, so as to perform an in vitro cell experiment on the target compound when or after determining that the target compound has potential cell-targeting drug activity according to the potential cell-targeting drug activity prediction result.

[0036] The third aspect of the present application provides a cell-targeted drug screening model training device, comprising:

[0037] A feature acquisition module is used to acquire molecular fingerprint feature data corresponding to each compound sample with label information to form a corresponding original data set, wherein the label information is used to indicate whether the corresponding compound sample has potential cell-targeted drug activity;

[0038] a data screening module, configured to screen each compound sample in the original data set using a distance metric to obtain a corresponding target data set, and to divide a training set from the target data set;

[0039] A model training module is used to train a graph neural network-based regression model based on the training set, so that the graph neural network-based regression model learns the structural topological information of each of the compound samples, and then trains the graph neural network-based regression model into a cell-targeted drug screening model for predicting whether the compound has potential cell-targeted drug activity based on the molecular fingerprint feature data of the compound.

[0040] The fourth aspect of the present application provides a cell-targeted drug screening device, comprising:

[0041] A data acquisition module is used to obtain molecular fingerprint characteristic data of the target compound;

[0042] A model prediction module is used to input the molecular fingerprint feature data of the target compound into a preset cell-targeted drug screening model, so that the cell-targeted drug screening model outputs a potential cell-targeted drug activity prediction result indicating whether the target compound has potential cell-targeted drug activity, wherein the cell-targeted drug screening model is pre-trained based on the cell-targeted drug screening model training method.

[0043] The fifth aspect of the present application provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for training a model for preliminary screening of cell-targeted drugs is implemented, and / or the method for preliminary screening of cell-targeted drugs is implemented.

[0044] The sixth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it implements the cell-targeted drug screening model training method and / or the cell-targeted drug screening method.

[0045] The seventh aspect of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the cell-targeted drug screening model training method and / or the cell-targeted drug screening method.

[0046] The present application provides a method for training a cell-targeted drug screening model, which obtains molecular fingerprint feature data corresponding to each compound sample with label information to form a corresponding original data set, wherein the label information is used to indicate whether the corresponding compound sample has potential cell-targeted drug activity; a distance metric is used to perform data screening on each compound sample in the original data set to obtain a corresponding target data set, and a training set is divided from the target data set; a graph neural network-based regression model is trained based on the training set, so that the graph neural network-based regression model learns the structural topological information of each compound sample, and then the graph neural network-based regression model is trained as a cell-targeted drug screening model for predicting whether the compound has potential cell-targeted drug activity based on the molecular fingerprint feature data of the compound. The present application provides a method for training a cell-targeted drug screening model, which obtains molecular fingerprint feature data of the compound, and obtains molecular fingerprint feature data of the compound, and obtains molecular fingerprint feature data of the compound. A regression model based on graph neural networks is adopted, and two molecular data structures, molecular fingerprints focusing on molecular functional group information and topological information focusing on compound structure, are adopted at the same time. The information of the two molecular data structures is comprehensively used to retain as many molecular features as possible that may be helpful to the drug screening model, thereby effectively improving the comprehensiveness of molecular characterization during the training process of the cell-targeted drug screening model; by introducing a new and efficient data screening strategy to eliminate compound samples that may cause interference from the original data set, it can overcome the influence of samples with similar characteristics but different label values ​​on the model, and can effectively improve the generalization ability of the trained cell-targeted drug screening model, thereby effectively improving the applicability of the trained cell-targeted drug screening model, and can improve the accuracy and reliability of the model's predicted results of potential cell-targeted drug activity.

[0047] Additional advantages, purposes, and features of the present application will be described in part in the following description and will become apparent to those skilled in the art upon study of the following or may be learned from practice of the present application. The purposes and other advantages of the present application may be achieved and obtained by the structures specifically pointed out in the specification and drawings.

[0048] Those skilled in the art will understand that the purposes and advantages that can be achieved by the present application are not limited to the above specific description, and the above and other purposes that can be achieved by the present application will be more clearly understood based on the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The drawings described herein are intended to provide a further understanding of the present application, constitute a part of the present application, and do not constitute a limitation of the present application. The components in the drawings are not drawn to scale, but are only for the purpose of illustrating the principles of the present application. In order to facilitate the illustration and description of some parts of the present application, the corresponding parts in the drawings may be enlarged, that is, they may become larger than other components in the exemplary device actually manufactured according to the present application. In the drawings:

[0050] Figure 1(a) to Figure 1(d) The following are schematic diagrams of mapping the latent feature spaces extracted by ECFP2, MACCS Key, MolFormer and ImageMol into two-dimensional space using t-SNE technology.

[0051] Figure 2 This is a schematic diagram of the first process of the cell-targeted drug screening model training method in one embodiment of the present application.

[0052] Figure 3 This is a second flow chart of the method for training a cell-targeted drug screening model in one embodiment of the present application.

[0053] FIG4( a ) is a schematic diagram showing the influence of hyperparameters α1=25 and α2=15 on the screening model when the number of reference anchor points K=5 in an example of the present application.

[0054] FIG4( b ) is a schematic diagram showing the influence of hyperparameters α1=25 and α2=40 on the screening model when the number of reference anchor points K=5 in an example of the present application.

[0055] FIG4( c ) is a schematic diagram showing the influence of hyperparameters α1 = 70 and α2 = 40 on the screening model when the number of reference anchor points K = 5 in an example of the present application.

[0056] Figure 5 Schematic diagram of the process of the cell-targeted drug screening method in one embodiment of the present application.

[0057] Figure 6 This is a schematic diagram of the specific algorithm flow of the data screening strategy in the application example of this application.

[0058] Figure 7 This is a schematic diagram of the execution flow of the cell-targeted drug screening model training and verification method in the application example of this application.

[0059] Figure 8 This is a schematic diagram of the model verification and performance evaluation comparison table in the application example of this application. DETAILED DESCRIPTION

[0060] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail in conjunction with the embodiments and drawings. Here, the illustrative embodiments of this application and their descriptions are used to explain this application, but are not intended to limit this application.

[0061] It should also be noted here that in order to avoid obscuring the present application due to unnecessary details, the accompanying drawings only show structures and / or processing steps that are closely related to the scheme according to the present application, while other details that are not closely related to the present application are omitted.

[0062] It should be emphasized that the term "include / comprises" when used herein refers to the existence of features, elements, steps or components, but does not exclude the existence or addition of one or more other features, elements, steps or components.

[0063] It should also be noted that, unless otherwise specified, the term "connection" herein may refer not only to a direct connection but also to an indirect connection involving an intermediate.

[0064] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0065] The process of compound drug screening involves efficiently screening large molecular libraries to identify candidate molecules suitable for treating specific diseases or pathophysiological processes. The specific process involves evaluating the physiological properties of various compounds, including efficacy, pharmacokinetics, and safety. For example, TH17 cells are an important type of T cell and a vital component of the immune system. Drug screening can be used to discover compounds that target Treg or TH17 cells, enabling the exploration of therapeutics for various immune-related diseases.

[0066] As a key branch of molecular machine learning, current mainstream drug screening methods generally employ a feature extractor and classifier model, or a feature extractor and regressor model. Consequently, these methods primarily involve two steps: feature engineering and building and training a classifier. The fundamental challenges of feature engineering are determining which data structure to use to model molecules and how to embed them into an appropriate latent space. In earlier molecular machine learning research, molecular representation methods were typically based on manually extracted features, specifically feature extraction methods that used fixed-length bit vectors to represent the characteristic structure of a molecule.

[0067] With the development of graph neural networks and their related models and optimization techniques, molecular representation based on graphs has become a mainstream approach to molecular feature engineering. Graph neural networks have a natural connection to molecular structure. For example, molecular structure is represented using a topological graph data structure, where atoms are represented as nodes and chemical bonds as edges. In addition to topological graphs, there are two other representative graph-based molecular representation methods: geometric graphs that reflect Euclidean space metrics, and knowledge graphs that can mine relationships between molecules. In addition to these representation methods, some molecular representation learning methods study structured string representations. Other work attempts to bridge the gap between computer vision and molecular representation learning tasks, re-examining the task from the perspective of natural RGB images.

[0068] Here, three of the most representative drug screening methods in recent years are listed: MolCLR, MolFormer and ImageMol.

[0069] Among them, MolCLR, a molecular representation learning framework based on self-supervised learning, models each molecular structure using a topological graph and pre-trains a graph neural network using a self-supervised learning framework based on contrastive learning to extract molecular representation features. It is entirely based on the general self-supervised learning framework SimCLR, in which the encoder network f(·) uses graph neural networks (GNNs) and the projection head g(·) uses multi-layer perceptrons (MLPs). Benefiting from pre-training on massive amounts of data, MolCLR not only achieves state-of-the-art results on multiple challenging benchmark datasets but is also proven to be able to embed molecules into representation features that can distinguish chemical properties. In contrast, MolFormer, a molecular representation learning framework based on self-supervised learning, treats molecular data as SMILES (Simplified Molecular Input Line Entry System) strings and adopts the most widely used masked language model in the field of natural language understanding (NLP) to achieve self-supervised learning tasks, achieving similarly good performance. The final work uses ImageMol, a molecular representation learning framework based on self-supervised learning. This work combines molecular modeling with common RGB images in the field of computer vision, and uses pre-training methods in computer vision, such as image reconstruction and jigsaw puzzle tasks, for molecular representation learning. Overall, existing implementations mainly focus on molecular representation learning and molecular property prediction, using methods such as self-supervised learning to improve model performance, especially when dealing with molecules with challenging structures. Among them, SMILES is a string representation used to represent molecular structure. It is a chemical description language that uses ASCII characters to represent the structure of molecules, allowing computers to easily process and interpret molecular information.

[0070] It can be seen that the existing cell-targeted drug screening model training methods have the following main problems:

[0071] (1) Limitations of molecular characterization techniques: Traditional molecular characterization studies rely primarily on fixed molecular fingerprints. However, this approach can lose important information related to drug physiological metabolism, especially when dealing with molecules with highly diverse structures. Although data-driven molecular characterization has addressed the shortcomings of molecular fingerprints to some extent, these methods typically focus on modeling molecules using a single specific modality, resulting in varying degrees of information loss.

[0072] For example, directly using topological maps for model training only considers the topological relationships in molecular structures, while ignoring the structural and spatial information of functional groups. SMILES, while including topological and functional group structural information, does not cover spatial information. Methods based on general RGB images may ignore prior knowledge in the field of chemistry, such as the chirality of molecules. In studies of cell promotion or inhibition, the lack of molecular information can critically affect the screening of drug models and the discovery of new therapeutic targets. This is because a comprehensive understanding of molecular properties is crucial for precisely regulating cellular responses.

[0073] (2) Generalization issues in specific fields: The effectiveness of existing methods is still limited to public datasets and specific datasets. For example, the three self-supervised learning molecular characterization methods mentioned above have only comprehensively evaluated their methods in some tasks on the MoleculeNet benchmark dataset, which indirectly shows that these methods have problems of poor adaptability and out-of-domain. Among them, the method closest to real-world application value is ImageMol. In this work, not only their models were evaluated on the MoleculeNet benchmark, but also extremely detailed demonstrations were carried out on anti-SARS coronavirus molecular inhibitors, proving that the method can explore potential 3C-like protease inhibitors with the potential to treat COVID-19. However, although existing methods such as ImageMol perform well on molecular compound datasets with the same distribution as the training set, their generalization ability is insufficient on data from unknown domains that do not meet the independent and identically distributed assumptions. As Figure 1(a) to Figure 1(d) As shown in the figure, using t-SNE technology to map the latent feature space extracted by Extended Connectivity FingerPrints (ECFP), MACCSKey (Molecular Access System Keys), MolFormer, and ImageMol into a two-dimensional space, the researchers found that all four methods were not effective in effectively distinguishing inhibitors from promoters. Among them, MACCS Key is a 166-bit fingerprint used to describe the molecular structure.

[0074] Based on this, in order to solve the above-mentioned technical problems existing in the existing methods, the embodiments of the present application respectively provide a cell-targeted drug primary screening model training method, a cell-targeted drug primary screening model training device for executing the cell-targeted drug primary screening model training method, a cell-targeted drug primary screening method, a cell-targeted drug primary screening device for executing the cell-targeted drug primary screening method, computer equipment, computer-readable storage medium and computer program product, which can improve the accuracy of cell-targeted drug primary screening predictions and the model generalization ability.

[0075] The details are described in detail through the following examples.

[0076] Based on this, the embodiment of the present application provides a cell-targeted drug primary screening model training method that can be implemented by a cell-targeted drug primary screening model training device, see Figure 2 The cell-targeted drug screening model training method specifically includes the following contents:

[0077] Step 100: Acquire molecular fingerprint feature data corresponding to each compound sample with label information to form a corresponding original data set, wherein the label information is used to indicate whether the corresponding compound sample has potential cell-targeted drug activity.

[0078] In one or more embodiments of the present application, the label information is used to indicate whether the corresponding compound sample has potential cell-targeted drug activity. Specifically, the expression level of cytokines obtained by biological experiments of the compound corresponding to the compound sample can be used to indicate whether it has potential cell-targeted drug activity. For example, if the expression level of the cytokine of a certain compound is equal to or greater than a preset value, it is determined to have potential cell-targeted drug activity; otherwise, it is determined to not have potential cell-targeted drug activity.

[0079] It can be understood that the compounds corresponding to the compound samples are pre-acquired compounds that are highly correlated with the activation or inhibition of target cells.

[0080] Step 200: Using a distance metric method to perform data screening on each of the compound samples in the original data set to obtain a corresponding target data set, and dividing a training set from the target data set.

[0081] In step 200, a data screening strategy using a distance metric can be used to solve the problem of difficulty in distinguishing between promoters and inhibitors in datasets such as Treg, and can keep the label differences of samples with similar characteristics within an acceptable range to screen out heterogeneous samples in the original dataset.

[0082] Step 300: Training a graph neural network-based regression model based on the training set, so that the graph neural network-based regression model learns the structural topological information of each of the compound samples, and then training the graph neural network-based regression model into a cell-targeted drug screening model for predicting whether the compound has potential cell-targeted drug activity based on the molecular fingerprint feature data of the compound.

[0083] The graph neural network-based regression model provided in the embodiments of this application models molecules. Specifically, the graph neural network uses atoms in the molecule as nodes and chemical bonds as edges, modeling the molecule using a graph data structure. Specifically, the graph neural network-based regression model represents nodes, edges, and global information as learnable vectors, with the structural topology of the compound being part of the global information.

[0084] From the above description, it can be seen that the cell-targeted drug screening model training method provided in the embodiment of the present application adopts a regression model based on a graph neural network, and simultaneously adopts two molecular data structures, namely, molecular fingerprints that focus on molecular functional group information and topological information that focuses on compound structure. The comprehensive use of information from the two molecular data structures can retain as many molecular features as possible that may be helpful to the drug screening model, thereby effectively improving the comprehensiveness of molecular characterization during the training process of the cell-targeted drug screening model; by introducing a new and efficient data screening strategy to eliminate compound samples that may cause interference from the original data set, it can overcome the influence of samples with similar characteristics but different label values ​​on the model, and can effectively improve the generalization ability of the trained cell-targeted drug screening model, thereby effectively improving the applicability of the trained cell-targeted drug screening model, and can improve the accuracy and reliability of the model-predicted potential cell-targeted drug activity prediction results.

[0085] In order to further improve the effectiveness and reliability of the cell-targeted drug screening model training, in a cell-targeted drug screening model training method provided in the embodiment of the present application, see Figure 3 The cell-targeted drug initial screening model training method further includes the following contents before step 100:

[0086] Step 010: Obtaining original information data of each candidate compound for target cell activation or inhibition and their corresponding label information.

[0087] Specifically, the original information data of candidate compounds that are highly correlated with target cell activation or inhibition are first collected, wherein the original information data may include information such as the compound name and number, and the number may be a CAS code (Chemical Abstracts Service Registry Numbers). The expression levels of cytokines measured by biological experiments for the candidate compounds are collected, and label information is formed based on the expression levels of the cytokines of each candidate compound to indicate whether the candidate compound has potential cell-targeted drug activity. Among them, the CAS code is a standardized identity code and a digital identifier used to uniquely identify a chemical substance.

[0088] Step 020: Perform SMILES string conversion on the original information data of each candidate compound; for a candidate compound whose SMILES string conversion is successful and whose generated SMILES string is unique, use the SMILES string of the candidate compound as a compound sample; delete the candidate compound whose SMILES string conversion is unsuccessful; for a candidate compound whose SMILES string conversion is successful but whose generated SMILES string is not unique, select one of the SMILES strings of the candidate compound as a compound sample.

[0089] That is, the raw information data represented by CAS codes is converted into corresponding SMILES strings through the API interface of PubChem, the database of the chemical module. During this conversion process, two special cases may occur:

[0090] (1) If the CAS code of the compound does not have a corresponding SMILES string, it will not be used;

[0091] (2) The CAS code of a compound corresponds to multiple SMILES strings. The first SMILES string in the PubChem database query result can be used as the corresponding conversion result by default.

[0092] The drug screening principle corresponding to the cell-targeted drug screening model training method provided in the embodiment of the present application is based on actual screening data. Therefore, although only two cell types, Treg or Th17, are used as examples in the embodiment, in actual principle, any type of cell can be used.

[0093] Therefore, in order to further improve the applicability and diversity of the initial screening of cell-targeted drugs, in a cell-targeted drug initial screening model training method provided in an embodiment of the present application, the target cells can be applicable to various types of cells. For example, the target cells can be selected from important cell types such as tumor cells, immune cells, cardiomyocytes, neurons, endocrine cells, liver cells, stem cells or respiratory epithelial cells. Among them, the cell types that immune cells can select include but are not limited to: T cells, B cells, NK cells, myeloid cells or antigen presenting cells (such as macrophages or dendritic cells), etc. These cells are all very important cell types for drug targeting.

[0094] In order to further effectively improve the comprehensiveness of molecular characterization during the training process of the cell-targeted drug primary screening model, in a cell-targeted drug primary screening model training method provided in the embodiment of the present application, see Figure 3 Step 100 of the cell-targeted drug screening model training method specifically includes the following:

[0095] Step 110: Based on a preset latent feature space representation method, molecular fingerprint feature data corresponding to each compound sample with label information is obtained.

[0096] Step 120: Generate an original data set containing the molecular fingerprint feature data and the label information corresponding to each of the compound samples; wherein the latent feature space representation method includes: at least one of the ECFP2 molecular fingerprint representation method, the ECFP4 molecular fingerprint representation method, the ECFP6 molecular fingerprint representation method and the MACCS Key molecular fingerprint representation method.

[0097] It is understandable that the embodiments of the present application may give priority to the ECFP2 molecular fingerprint representation method as the latent feature space representation method. On this basis, it is possible to consider expanding the radius range of ECFP, that is, using methods such as ECFP4 or ECFP6 to extract features of compounds, or using MACCS Key or other features as the representation method of compounds.

[0098] In order to further improve the generalization ability of the cell-targeted drug screening model, in a cell-targeted drug screening model training method provided in the embodiment of the present application, see Figure 3 Step 200 of the cell-targeted drug primary screening model training method specifically includes the following:

[0099] Step 210: Randomly select molecular fingerprint feature data of a plurality of compound samples from the original data set to serve as reference anchor points.

[0100] Step 220: Based on the number of the reference anchor points, a preset first hyperparameter for representing the radius threshold of the feature space hypersphere, and a second hyperparameter for representing the label difference threshold, a distance metric is used to cluster the compound samples in the original data set to screen out heterogeneous samples in the original data set, thereby obtaining a corresponding target data set, wherein the distance metric includes: a Hamming distance metric method or a Euclidean distance metric method.

[0101] Step 230: Divide the target data set into a training set and a test set.

[0102] Here, the steps 210 and 220 are described by taking the Hamming distance measurement method as the distance measurement method as an example:

[0103] To address the problem of inseparable promoters and inhibitors in datasets such as Treg, this application proposes a data screening strategy based on the ECFP2 feature. This strategy uses an efficient approximate distance metric (Hamming distance) to keep the label differences of samples with similar features within an acceptable range, thereby filtering out heterogeneous samples in the dataset.

[0104] Where K represents the number of reference anchor points randomly selected from the original dataset, α1 and α2 represent the radius threshold of the feature space hypersphere and the label difference threshold, respectively. Figure 4(a) to Figure 4(c) The visualization of different α1 and α2 values ​​in two-dimensional space when K = 5 is shown. A larger α1 means a larger neighborhood is considered by each anchor point, while a larger α2 means a higher tolerance for sample label discrepancies. This design concept can be understood as clustering samples into K clusters based on anchor points, with each cluster having a radius of α1.

[0105] In order to further effectively improve the comprehensiveness of molecular characterization during the training process of the cell-targeted drug primary screening model, in a cell-targeted drug primary screening model training method provided in the embodiment of the present application, see Figure 3 Step 300 of the cell-targeted drug primary screening model training method specifically includes the following:

[0106] Step 310: Iteratively train the graph neural network-based regression model based on the training set and preset training hyperparameters, so that the graph neural network-based regression model represents the nodes used to represent atoms, the edges used to represent chemical bonds, and the global information containing the structural topological information of the compound sample in the form of learnable vectors, and during the iterative training process, the potential cell-targeted drug activity prediction results obtained each time by the graph neural network-based regression model are screened according to a preset prediction threshold range.

[0107] Step 320: Performing a model test on the trained graph neural network-based regression model based on the test set. If the graph neural network-based regression model passes the model test, the graph neural network-based regression model is used as a preliminary screening model for cell-targeted drugs for predicting whether the compound has potential cell-targeted drug activity based on the molecular fingerprint feature data of the compound.

[0108] Among them, the regression model based on graph neural network includes: graph convolutional neural network GCN (Graph Convolutional Networks) model and / or message passing neural network MPNN (Message Passing Neural Networks) model.

[0109] Specifically, on samples after data screening, this application implements two regression models based on graph neural networks based on the DeepChem library: the GCN model and the MPNN model, which are designed to further analyze the topology of compounds after ECFP functional group screening to achieve drug screening tasks. Among them, the network structure of the two models uses the default settings, and the remaining training hyperparameters are set as follows: batch size is 16, learning rate is 0.001, and the number of training rounds is 100. In order to prevent the model from outputting abnormal prediction scores, during the iterative training stage of the model, the expression levels of cytokines shown in the prediction results of the regression model are clipped, that is, the prediction results are limited to [y min ,y max ], where y min and y max are the minimum and maximum values ​​of the label information in the original dataset, respectively.

[0110] Based on the cell-targeted drug initial screening model training method provided in the above embodiment, the present application also provides an embodiment of a cell-targeted drug initial screening method, see Figure 5 The cell-targeted drug screening method specifically includes the following contents:

[0111] Step 400: Acquire molecular fingerprint feature data of the target compound.

[0112] Step 500: Input the molecular fingerprint feature data of the target compound into a preset cell-targeted drug screening model, so that the cell-targeted drug screening model outputs a potential cell-targeted drug activity prediction result indicating whether the target compound has potential cell-targeted drug activity, wherein the cell-targeted drug screening model is pre-trained based on the cell-targeted drug screening model training method.

[0113] It should be noted that the cell-targeted drug primary screening model training method mentioned in step 500 can be executed with reference to the embodiment of the aforementioned cell-targeted drug primary screening model training method, and will not be described in detail.

[0114] From the above description, it can be seen that the cell-targeted drug screening method provided in the embodiment of the present application can effectively improve the comprehensiveness of molecular characterization during the training process of the cell-targeted drug screening model, and can effectively improve the generalization ability of the trained cell-targeted drug screening model, thereby improving the accuracy and reliability of the potential cell-targeted drug activity prediction results predicted by the model.

[0115] In order to further improve the timeliness and reliability of subsequent processing of cell-targeted drug screening, in the cell-targeted drug screening method provided in this application, see Figure 5 The cell-targeted drug screening method further includes the following steps after step 500:

[0116] Step 600: Outputting the potential cell-targeting drug activity prediction result corresponding to the target compound, so as to perform in vitro cell experiments on the target compound when or after determining that the target compound has potential cell-targeting drug activity according to the potential cell-targeting drug activity prediction result.

[0117] Specifically, the cell-targeted drug screening device outputs the potential cell-targeted drug activity prediction result corresponding to the target compound, so that researchers can conduct in vitro cell experiments on the target compound when or after determining that the target compound has potential cell-targeted drug activity based on the potential cell-targeted drug activity prediction result.

[0118] That is, after initially screening for compounds with potential drug activity, the next step is a more in-depth evaluation to verify the safety and efficacy of these compounds. This stage of evaluation includes more complex and detailed in vitro cell experiments and may even involve the use of animal models. These experiments are designed to simulate the behavior of the compound in vivo and provide more comprehensive information about its pharmacological effects, toxicity, and metabolic properties.

[0119] The purpose of these experiments is to ensure that the compounds are not only effective in theory, but also safe and effective in actual biological systems. If the compounds show good safety and expected efficacy in these further evaluations, they can be selected as candidates for entering biological testing. Biological testing is an important part of the drug development process. It involves further testing of compounds in biological systems to verify their pharmacological effects and safety and lay the foundation for subsequent clinical trials. This stage is a critical transition from laboratory to clinical application, ensuring that only the most promising compounds continue to develop.

[0120] In order to further illustrate the effectiveness of the above-mentioned cell-targeted drug screening model training method and the cell-targeted drug screening method, the present application also provides a cell-targeted drug screening model training and verification method, which is illustrated by taking immune cells as the target cells. Figure 7 , the training and verification method specifically includes the following contents:

[0121] Step 1: Compound database and data preprocessing

[0122] This application example collected 1,400 candidate compounds (compound names, numbers, etc.) that are highly correlated with immune cell activation or inhibition, and collected the expression levels of cytokines measured by biological experiments (label information corresponding to the compounds), and constructed a data set for cell-targeted drug screening. First, the raw data represented by CAS codes are converted into corresponding SMILES strings through the PubChem API interface. During this conversion process, two special cases may occur: (1) the CAS code of the compound does not have a corresponding SMILES string; (2) the CAS code of the compound corresponds to multiple SMILES strings. There are 26 compounds in the Treg dataset that have situation (1), and these compounds were not used in subsequent experiments. For compounds that have situation (2), this application example defaults to the first SMILES string of the PubChem database query result as their corresponding conversion result. The compounds after preprocessing are represented by SMILES strings. Finally, the ECFP2 fingerprints corresponding to these compounds are calculated as the input of the subsequent data screening model.

[0123] Step 2: Data screening strategy

[0124] In order to solve the problem of inseparable promoters and inhibitors in Treg datasets, this application example proposes a data screening strategy based on ECFP2 features. This strategy designs an efficient approximate distance metric method (Hamming distance) to keep the label differences of samples with similar features within an acceptable range to screen out heterogeneous samples in the dataset. The specific algorithm flow of the data screening strategy is as follows Figure 6 As shown in Figure 2, the input of this algorithm is the original dataset, the number of anchor points K, and two threshold-related hyperparameters α1 and α2. The output is a filtered dataset. K represents the number of anchor points randomly selected from the original dataset, while α1 and α2 represent the radius threshold of the feature space hypersphere and the label difference threshold, respectively.

[0125] There are also multiple ways to set the hyperparameters α1 and α2. In one example, this application uses an adaptive dataset setting method. This method uses quantile calculations based on the distribution of the entire dataset to determine the thresholds for the hypersphere radius and label discrepancy tolerance. These two hyperparameters are also pre-set based on experience.

[0126] It should be noted that in the data screening strategy proposed in the application example, the representation method of the latent feature space (ECFP2 is used in the application example of this application) and the distance metric function (Hamming distance is used in the application example of this application) are interchangeable. For example, one can consider expanding the radius range of ECFP, that is, using methods such as ECFP4 or ECFP6 to extract features from compounds, or using MACCS Key or other features as the representation method of compounds. In terms of distance metric, Hamming distance is the most common distance metric method between binary vectors, but in actual operation, Euclidean distance can also play a similar role.

[0127] Step 3: Implement and deploy the regression model

[0128] Based on the samples after data screening, four regression models were designed and verified based on the DeepChem library, including traditional machine learning models SVR and Ridge, and two regression models based on graph neural networks: GCN and MPNN. The purpose is to conduct further topological analysis of compounds after ECFP functional group screening to achieve drug screening tasks.

[0129] exist Figure 7 The MPNN model is omitted in the figure, but in actual execution, the four models mentioned above, including the MPNN model, are trained in step 3. The network structures of all four models use the default settings, and the remaining training hyperparameters are set as follows: batch size of 16, learning rate of 0.001, and number of training epochs of 100. To prevent the model from outputting abnormal prediction scores, the prediction scores of the regression model are manually clipped during the inference phase, limiting them to the range [y_min, y_max], where y_min and y_max are the minimum and maximum labeled values ​​in the original dataset, respectively.

[0130] Among them, the four models (SVR, Ridge, GCN and MPNN) mentioned in the application examples of this application are all independent parallel models and have no correlation during the deployment process. In addition, the reason why this application proposes to use multiple models for training and verification is to reflect the effectiveness of the data screening method proposed in the previous step. In summary, the training process can be one or more. The training and evaluation process of each model is completed independently, and the execution logic sequence and process are the same. Their inputs are all data sets that have been filtered by the data screening strategy, and then these data are further divided into training sets and test sets. The final output is a trained model and the corresponding evaluation results of the model on the test set.

[0131] In addition, the application examples of this application can also train simpler linear regression models, such as linear regression, LASSO regression and ridge regression, or other machine learning models, such as support vector regression, random forest, multi-layer perceptron, etc., which can be used for similar tasks. Even representative works such as MolFormer can fine-tune the model on the data set after data screening to achieve the same effect. However, the effectiveness of these methods still needs to be evaluated. For example, although linear models can fit regression tasks, such methods lose the topological structure information of the compound.

[0132] Step 4: Model Validation and Performance Evaluation

[0133] If the SVR and Ridge modeling methods are used, the corresponding topological information is ignored when modeling molecules (because the feature extraction method of these two methods is still ECFP). Accordingly, compared with the two graph neural network modeling methods GCN and MPNN, the experimental results corresponding to the traditional machine learning models SVR and Ridge are generally poor.

[0134] In order to verify the above content, all experimental results in the application examples of this application are subjected to five-fold cross validation. The application examples of this application not only consider the mean absolute error (MAE) and root mean square error (RMSE) used by the MoleculeNet benchmark dataset, but also introduce two additional evaluation indicators for regression tasks, R2 and Spearman correlation coefficient. These additional evaluation indicators help to more comprehensively evaluate the performance of the model for drug screening tasks. In particular, the Spearman correlation coefficient only considers the ranking based on the variables, not the actual numerical value. It measures the monotonic relationship between two variables and is not affected by outliers. Therefore, when dealing with molecular property data that may have extreme values, it can provide a more robust evaluation.

[0135] Since this application uses five-fold cross validation, the above steps will be repeated five times based on different data partitions when facing the same input data (filtered data set), such as Figure 8 The mean and standard deviation of each indicator of each model in the model verification and performance evaluation comparison table shown are also calculated in this way.

[0136] In the model validation and performance evaluation comparison table, the first two columns of the table are hyperparameters related to the data screening strategy. The value of α1 has two options of 25 and 70, and the value of α2 has three options of 15, 30 and 40. " / " means that the data screening strategy is not used, and the original data set is directly used for the training and verification of the subsequent model (blank control group, in order to illustrate the effectiveness of the data screening strategy proposed in this application). The third column is the number of samples in the screening data set obtained based on different hyperparameters. If the data screening strategy is not used, the number of samples in the entire data set is 1449. The number of samples in the screened data set is much smaller than the original data set size. The fourth column is the four models used, and columns 5-8 are the corresponding indicators of the independent training and testing of the four models under the hyperparameter settings of the current data screening strategy. Among them,

[0137] “#samples” indicates the number of samples, “Model” indicates the model, and “MAE”, “RMSE”, “R2” and “Spearman’s corr.” are different regression model evaluation indicators. MAE (Mean Absolute Error) refers to the mean absolute error, RMSE (Root Mean Square Error) refers to the root mean square error, and R2 (R-Square) refers to the goodness of fit; “Spearman’s corr.” refers to the Spearman correlation coefficient.

[0138] Based on this, this application example aims to overcome the limitations of existing molecular characterization technologies, especially in drug screening models and the discovery of new therapeutic targets. It mainly addresses the following two core issues:

[0139] ① Limitations of single molecule characterization

[0140] Traditional molecular characterization methods typically rely on fixed molecular fingerprints or specific modalities, such as topological maps or SMILES encoding. These methods often fail to fully capture the complexity of molecules, especially the structural and spatial information of functional groups. To address this issue, the application example of this application adopts a fusion method that simultaneously considers molecular functional group information and the topological map of the compound structure. By combining the information of these two data structures, this method can more comprehensively retain molecular features that are beneficial to drug screening models.

[0141] ②The problem of domain generalization

[0142] Existing molecular characterization methods may perform well on specific datasets, but they usually perform poorly when dealing with data from unknown domains and lack generalization capabilities. To address this challenge, this application example proposes a new data screening strategy that clusters samples into K clusters based on random anchor points. By adjusting the neighborhood radius of each cluster and the label difference tolerance of each anchor point for samples, the samples are cleaned, and potentially interfering compound samples are removed from heterogeneous datasets, reducing the negative impact of samples with similar features but different label values ​​on the model. Subsequently, graph neural networks are used to model, train, and cross-validate the screened data, forming an effective drug screening strategy.

[0143] This application example proposes an automated drug screening process. In the first stage, a latent feature space representation method and distance measurement method are used to screen noisy data. In the second stage, another type of molecular modeling method is used to consider multiple molecular data modalities to implement a drug screening method for the field of immunology. This method can efficiently discover potential drugs and quickly and effectively identify potential drugs from thousands of possible candidate molecules. This helps reduce R&D costs and improve the efficiency of the entire drug development process.

[0144] Regarding the limitations of single molecule characterization: The application example of this application simultaneously considers two molecular data structures: molecular fingerprints that focus on molecular functional group information and topological maps that focus on compound structure. By comprehensively using the information of the two molecular data structures, as many molecular features as possible that may be helpful for drug screening models are retained.

[0145] Regarding domain generalization issues: This application example introduces a novel and efficient data screening strategy. This strategy focuses on removing potentially interfering compound samples from heterogeneous datasets, overcoming the impact of samples with similar features but different label values ​​on the model. Graph neural networks are then used to model, train, and cross-validate the filtered data, resulting in an effective initial drug screening strategy.

[0146] In other words, the application example of this application proposes a molecular representation learning method that considers multiple molecular modalities. The proposed data screening strategy can effectively remove the noise sample problem in heterogeneous data sets, and can to a certain extent solve the adaptability problem of graph neural networks represented by GCN and MPNN in the field of molecular machine learning, so that the model can perform generalizable reasoning on molecular compound data sets with unstable labels. In summary, the basic principle of the application example of this application is to improve the accuracy and generalization ability of drug screening models by combining multimodal molecular data structures and efficient data screening methods, thereby opening up new avenues for discovering new therapeutic targets.

[0147] The application example of this application simultaneously considers two molecular data structures: molecular fingerprints that focus on molecular functional group information and topological maps that focus on compound structures. Combining the information of the two molecular data structures can better capture the complexity and diversity of molecules. Compared with the molecular modeling method in the prior art that only considers a single modality, the application example of this application can explore the key features related to physiological metabolism in compounds to a greater extent.

[0148] Data screening strategies address domain adaptability issues: This application example utilizes data screening strategies to improve the robustness, generalization, and adaptability of graph neural network models to specific domains, thereby avoiding significant performance degradation outside of these domains. Furthermore, this application example can effectively process and analyze small datasets for targeted drug screening and effectively ignore noise samples in the dataset, improving the model's generalization capabilities.

[0149] From a software perspective, the present application further provides a cell-targeted drug primary screening model training device for executing all or part of the cell-targeted drug primary screening model training method. The cell-targeted drug primary screening model training device specifically includes the following contents:

[0150] A feature acquisition module is used to acquire molecular fingerprint feature data corresponding to each compound sample with label information to form a corresponding original data set, wherein the label information is used to indicate whether the corresponding compound sample has potential cell-targeted drug activity;

[0151] a data screening module, configured to screen each compound sample in the original data set using a distance metric to obtain a corresponding target data set, and to divide a training set from the target data set;

[0152] A model training module is used to train a graph neural network-based regression model based on the training set, so that the graph neural network-based regression model learns the structural topological information of each of the compound samples, and then trains the graph neural network-based regression model into a cell-targeted drug screening model for predicting whether the compound has potential cell-targeted drug activity based on the molecular fingerprint feature data of the compound.

[0153] The embodiment of the cell-targeted drug screening model training device provided in this application can be specifically used to execute the processing flow of the embodiment of the cell-targeted drug screening model training method in the above embodiment. Its functions will not be described in detail here, and reference can be made to the detailed description of the above embodiment of the cell-targeted drug screening model training method.

[0154] The part of the cell-targeted drug primary screening model training device that performs the cell-targeted drug primary screening model training can be completed in the server or in the client device. The specific selection can be made based on the processing power of the client device and the limitations of the user's usage scenario. This application is not limited to this. If all operations are completed in the client device, the client device may also include a processor for the specific processing of the cell-targeted drug primary screening model training.

[0155] The client device may include a communication module (i.e., a communication unit) that can establish a communication connection with a remote server to implement data transmission with the server. The server may include a server on the task scheduling center side, and in other implementation scenarios, may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, a server cluster consisting of multiple servers, or a server structure of a distributed device.

[0156] The server and the client device may communicate using any suitable network protocol, including network protocols that have not yet been developed as of the filing date of this application. Examples of such network protocols include TCP / IP, UDP / IP, HTTP, and HTTPS. Furthermore, examples of such network protocols include RPC (Remote Procedure Call Protocol) and REST (Representational State Transfer) protocols, which are used on top of the aforementioned protocols.

[0157] From the above description, it can be seen that the cell-targeted drug screening model training device provided in the embodiment of the present application can effectively improve the comprehensiveness of molecular characterization during the training process of the cell-targeted drug screening model, and can effectively improve the generalization ability of the trained cell-targeted drug screening model, thereby effectively improving the applicability of the trained cell-targeted drug screening model, and can improve the accuracy and reliability of the potential cell-targeted drug activity prediction results predicted by the model.

[0158] From a software perspective, the present application further provides a cell-targeted drug screening device for executing all or part of the cell-targeted drug screening method. The cell-targeted drug screening device specifically includes the following:

[0159] A data acquisition module is used to obtain molecular fingerprint characteristic data of the target compound;

[0160] A model prediction module is used to input the molecular fingerprint feature data of the target compound into a preset cell-targeted drug screening model, so that the cell-targeted drug screening model outputs a potential cell-targeted drug activity prediction result indicating whether the target compound has potential cell-targeted drug activity, wherein the cell-targeted drug screening model is pre-trained based on the cell-targeted drug screening model training method.

[0161] It can be seen from the above description that the cell-targeted drug screening device provided in the embodiment of the present application can improve the accuracy and reliability of the prediction results of potential cell-targeted drug activity predicted by the model.

[0162] The present application also provides an electronic device that may include a processor, a memory, a receiver, and a transmitter. The processor is configured to execute the cell-targeted drug screening model training method and / or the cell-targeted drug screening method described in the above embodiments. The processor and the memory may be connected via a bus or other means, with bus connection being an example. The receiver may be connected to the processor and the memory via a wired or wireless means.

[0163] The processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or a combination of the above chips.

[0164] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs, non-transient computer executable programs and modules, such as the cell-targeted drug primary screening model training method and / or the cell-targeted drug primary screening method in the embodiments of the present application. The processor executes various functional applications and data processing of the processor by running the non-transient software programs, instructions and modules stored in the memory, that is, implementing the cell-targeted drug primary screening model training method and / or the cell-targeted drug primary screening method in the above method embodiments.

[0165] The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the processor, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0166] The one or more modules are stored in the memory, and when executed by the processor, perform the cell-targeted drug primary screening model training method and / or cell-targeted drug primary screening method in the embodiment.

[0167] In some embodiments of the present application, the user equipment may include a processor, a memory and a transceiver unit, and the transceiver unit may include a receiver and a transmitter. The processor, memory, receiver and transmitter may be connected through a bus system. The memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to send and receive signals.

[0168] As an implementation method, the functions of the receiver and transmitter in this application can be considered to be implemented through a transceiver circuit or a dedicated transceiver chip, and the processor can be considered to be implemented through a dedicated processing chip, a processing circuit or a general-purpose chip.

[0169] As another implementation method, it is possible to use a general-purpose computer to implement the server provided in the embodiments of the present application. That is, the program code for implementing the functions of the processor, receiver, and transmitter is stored in a memory, and the general-purpose processor implements the functions of the processor, receiver, and transmitter by executing the code in the memory.

[0170] The present application embodiment also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the aforementioned cell-targeted drug primary screening model training method and / or cell-targeted drug primary screening method. The computer-readable storage medium can be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the art.

[0171] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the cell-targeted drug screening model training method and / or the cell-targeted drug screening method.

[0172] It should be understood by those skilled in the art that the various exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software or a combination of the two. Whether it is specifically performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of this application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier.

[0173] It should be understood that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present application.

[0174] In this application, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or replace features of other embodiments.

[0175] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Those skilled in the art will appreciate that various modifications and variations of the present embodiment are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.

Claims

1. A method for training a cell-targeted drug screening model, characterized in that: include: Acquiring molecular fingerprint feature data corresponding to each compound sample provided with label information to form a corresponding original data set, wherein the label information is used to indicate whether the corresponding compound sample has potential cell-targeted drug activity; Using a distance metric to screen each compound sample in the original data set to obtain a corresponding target data set, and dividing a training set from the target data set; A graph neural network-based regression model is trained based on the training set so that the graph neural network-based regression model learns the structural topological information of each of the compound samples, and then the graph neural network-based regression model is trained as a cell-targeted drug screening model for predicting whether the compound has potential cell-targeted drug activity based on the molecular fingerprint feature data of the compound.

2. The cell-targeted drug screening model training method according to claim 1, characterized in that: Before obtaining the molecular fingerprint feature data corresponding to each compound sample provided with label information to form a corresponding original data set, the method further includes: Obtaining raw information data of each candidate compound used for target cell activation or inhibition and their corresponding label information; Performing SMILES string conversion on the original information data of each candidate compound; For a candidate compound whose SMILES string conversion is successful and whose generated SMILES string is unique, the SMILES string of the candidate compound is used as a compound sample; Candidate compounds that failed to convert SMILES strings were deleted; For a candidate compound whose SMILES string conversion is successful but whose generated SMILES string is not unique, one of the SMILES strings of the candidate compound is selected as a compound sample.

3. The cell-targeted drug screening model training method according to claim 2, characterized in that: The target cells include: tumor cells, immune cells, myocardial cells, neuronal cells, endocrine cells, hepatocytes, stem cells or respiratory epithelial cells; Wherein, the immune cells include: T cells, B cells, NK cells, myeloid cells or antigen presenting cells.

4. The cell-targeted drug screening model training method according to claim 1, characterized in that: The step of obtaining molecular fingerprint feature data corresponding to each compound sample provided with label information to form a corresponding original data set includes: Based on a preset latent feature space representation method, the molecular fingerprint feature data corresponding to each compound sample with label information is obtained; generating an original data set including the molecular fingerprint feature data and the label information corresponding to each of the compound samples; The latent feature space representation method includes at least one of an ECFP2 molecular fingerprint representation method, an ECFP4 molecular fingerprint representation method, an ECFP6 molecular fingerprint representation method, and a MACCS Key molecular fingerprint representation method.

5. The cell-targeted drug screening model training method according to claim 1, characterized in that: The method of using a distance metric to screen each compound sample in the original data set to obtain a corresponding target data set, and dividing a training set from the target data set, includes: Randomly selecting molecular fingerprint feature data of a plurality of compound samples from the original data set as reference anchor points respectively; Based on the number of the reference anchor points, a first hyperparameter for representing a radius threshold of a feature space hypersphere, and a second hyperparameter for representing a label difference threshold, clustering the compound samples in the original dataset using a distance metric to screen out heterogeneous samples in the original dataset, thereby obtaining a corresponding target dataset; Dividing the target data set into a training set and a test set; The distance measurement method includes: a Hamming distance measurement method or a Euclidean distance measurement method.

6. The cell-targeted drug screening model training method according to claim 5, characterized in that: The step of training a graph neural network-based regression model based on the training set so that the graph neural network-based regression model learns the structural topological information of each compound sample, and then training the graph neural network-based regression model into a cell-targeted drug screening model for predicting whether the compound has potential cell-targeted drug activity based on the molecular fingerprint feature data of the compound, comprising: Iteratively training a graph neural network-based regression model based on the training set and preset training hyperparameters, so that the graph neural network-based regression model represents nodes used to represent atoms, edges used to represent chemical bonds, and global information containing structural topological information of the compound sample in the form of learnable vectors, and during the iterative training process, screening the potential cell-targeted drug activity prediction results obtained each time by the graph neural network-based regression model according to a preset prediction threshold range; The trained graph neural network-based regression model is subjected to a model test based on the test set. If the graph neural network-based regression model passes the model test, the graph neural network-based regression model is used as a cell-targeted drug screening model for predicting whether the compound has potential cell-targeted drug activity based on the molecular fingerprint feature data of the compound.

7. The method for training a cell-targeted drug screening model according to any one of claims 1 to 6, wherein: The regression model based on graph neural network includes: GCN model and / or MPNN model.

8. A method for primary screening of cell-targeted drugs, characterized in that: include: Obtain molecular fingerprint characteristic data of the target compound; The molecular fingerprint feature data of the target compound is input into a preset cell-targeted drug screening model, so that the cell-targeted drug screening model outputs a potential cell-targeted drug activity prediction result indicating whether the target compound has potential cell-targeted drug activity, wherein the cell-targeted drug screening model is pre-trained based on the cell-targeted drug screening model training method according to any one of claims 1 to 7.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, it implements the cell-targeted drug screening model training method according to any one of claims 1 to 7, and / or implements the cell-targeted drug screening method according to claim 8.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the cell-targeted drug screening model training method according to any one of claims 1 to 7, and / or implements the cell-targeted drug screening method according to claim 8.

Citation Information

Patent Citations

  • Small molecule drug virtual screening method and device based on deep parameter transfer learning

    CN112086146A

  • Method and device for predicting drug susceptibility state and storage medium

    CN112768089A

  • Prediction method and prediction device for drug molecule characteristic attributes

    CN115240781A

  • Drug-induced immune thrombocytopenia toxicity prediction model, method and system

    CN117438090A

  • Artificial intelligence-based drug molecule processing method and apparatus, device, storage medium, and computer program product

    US20230050156A1

Cited By

  • Permeability prediction model construction method, compound screening method and electronic equipment

    CN121938480A