Drug virtual screening method and system based on gene ontology enhanced contrast learning

By constructing a contrastive learning method based on gene ontology enhancement, and combining protein sequence and structural information, the problems of insufficient efficiency and generalization in virtual drug screening are solved, and efficient screening and ranking of new targets are achieved.

CN121331279APending Publication Date: 2026-01-13FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511492709.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing virtual drug screening methods have shortcomings in terms of efficiency, generalization and multimodal information utilization, especially in scenarios lacking binding pockets or high-quality structures, making it difficult to accurately predict new targets.

Method used

We employ a contrastive learning approach based on gene ontology enhancement to construct a multimodal dual encoder framework that supports both sequence and structural inputs. By jointly training contrastive learning and affinity ranking, we obtain a small molecule screening model, integrate the functional semantic information of proteins, and improve the model's generalization ability.

Benefits of technology

It improves the model's generalization ability to novel targets and proteins with incomplete functional annotations, enhances the robustness and applicability of prediction results, and is suitable for screening novel protein targets with no pocket information or unknown structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121331279A_ABST
    Figure CN121331279A_ABST
Patent Text Reader

Abstract

The invention belongs to the cross technical field of computer-aided drug discovery and deep learning, and discloses a drug virtual screening method and system based on gene ontology enhanced contrast learning, and the method comprises the steps: collecting a protein-small molecule historical pairing data set; constructing a multi-mode dual-encoder framework which is composed of a protein encoder and a small molecule encoder and supports sequence and structure dual-mode input; introducing gene ontology information, and performing functional semantic enhancement on the protein encoder through protein-gene ontology and gene ontology-gene ontology contrast learning tasks; a model is trained through the combination of comparative learning and affinity sorting; in the reasoning stage, protein representation independent of gene ontology information is used for small molecule screening, multi-modal prediction results are integrated through an uncertainty perception fusion mechanism, and a more stable sequencing result is obtained. The generalization ability of the model to unknown proteins is improved, and the method is suitable for actual drug discovery scenes.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of computer-aided drug discovery and deep learning, and particularly relates to a drug virtual screening method and system based on gene ontology enhanced contrast learning. BACKGROUND

[0002] Drug virtual screening is an indispensable part of modern drug research and development, and its goal is to efficiently identify small molecule candidates with potential binding capacity to target proteins in a large compound library, thereby significantly reducing experimental costs and accelerating the new drug development process. Traditional virtual screening methods mainly rely on structure-based molecular docking and molecular dynamics calculations, or ligand-based similarity search methods. The former assesses the binding capacity between small molecules and protein pockets through physical and chemical scoring functions, has certain interpretability, but the calculation process is complex and inefficient, and when facing large-scale compound libraries, it is often difficult to meet the actual application requirements; the latter uses the structure or physicochemical characteristics of known active molecules to predict potential candidate molecules, which is simple, but its scope of application and prediction effect are greatly limited when there is a lack of reference molecules or target pocket information.

[0003] In recent years, with the development of deep learning and large-scale biological data, researchers have proposed a series of virtual screening methods based on representation learning. This kind of method models the features of proteins and small molecules through neural networks, and uses large-scale data for end-to-end training, which improves the screening efficiency and prediction accuracy to a certain extent. However, existing methods still have many shortcomings. First, they mostly rely on known or predicted binding pocket information, and when there is a lack of high-quality protein structure or the pocket is difficult to drug, the model performance will decrease significantly. Second, the generalization ability of existing models is limited, and when facing new targets that are quite different from the training set, the prediction effect is not ideal, which is difficult to meet the low-sample or even zero-sample application requirements in new drug research and development. In addition, although gene ontology provides multi-level semantic annotations such as molecular function, cellular component and biological process for protein function, existing methods often only stay at the sequence or structure level of representation learning, and fail to effectively integrate gene ontology semantic information, thereby limiting the model's adaptability to unknown targets. At the same time, proteins usually have both sequence and structure information, and how to reasonably integrate the two, and maintain the robustness of prediction when information is missing or incomplete, is also a key problem that has not been solved.

[0004] Therefore, existing virtual drug screening methods still have significant shortcomings in terms of efficiency, generalization, and multimodal information utilization, especially in scenarios lacking binding pockets or high-quality structures, making it difficult to accurately predict new targets. These challenges urgently require a new method that can combine protein multimodal information with functional semantic knowledge to improve prediction accuracy while enhancing the model's generalization ability, thus better meeting the needs of actual drug development. Summary of the Invention

[0005] To address the problems existing in the prior art, this invention provides a drug virtual screening method and system based on gene ontology-enhanced contrastive learning, which can efficiently screen and sort candidate small molecules when only overall protein information (such as sequence or structure) is provided. It is suitable for novel protein targets with no pocket information or unknown structure.

[0006] To achieve the above objectives, the present invention provides the following solution: A virtual drug screening method based on gene ontology-enhanced contrastive learning, the method comprising: Collect historical datasets of protein-small molecule pairings; A multimodal dual encoder framework supporting sequence and structure dual-modal input was constructed, consisting of a protein encoder and a small molecule encoder. Based on the historical dataset, the multimodal dual encoder framework is jointly trained by contrastive learning and affinity ranking to obtain a small molecule screening model. Protein data is acquired, and candidate small molecules are screened and sorted using the small molecule screening model to obtain the corresponding active compounds.

[0007] Preferred methods for collecting historical datasets of protein-small molecule pairings include: Obtain historical protein data with sequence and three-dimensional structure information from publicly available protein databases; Collect diverse historical small molecule structure data from publicly available compound databases; For each historical protein data point, the corresponding historical small molecule structure data is selected to form a protein-small molecule pairing dataset, thus obtaining the historical dataset.

[0008] Preferably, the multimodal dual encoder framework includes a protein encoder and a small molecule encoder; The small molecule encoder includes: an input layer for receiving atomic type information and three-dimensional coordinate information of the small molecule; a graph neural network encoding layer for extracting local chemical features and spatial geometric relationships of atoms; and an embedding layer for outputting the overall feature representation of the small molecule. The protein encoder includes: a pre-trained language model encoding unit for receiving protein sequences; a spatial geometric feature extraction unit for receiving the three-dimensional structure of proteins; and an embedding layer for outputting a unified protein representation. The outputs of the small molecule encoder and the protein encoder are mapped to a shared high-dimensional feature space for similarity calculation and comparative learning.

[0009] Preferably, the method for obtaining a small molecule screening model by jointly training the multimodal dual encoder framework through contrastive learning and affinity ranking based on the historical dataset includes: A contrastive learning strategy is adopted to construct protein-small molecule positive and negative sample pairs based on the historical dataset and to perform embedding representation alignment training. During training, contrastive learning tasks of protein-gene ontology and gene ontology-gene ontology are introduced to enhance the functional semantics of the protein encoder. The affinity ranking of small molecules corresponding to the same protein is optimized using a ranking loss function. Output the small molecule screening model completed through joint training.

[0010] Preferably, the method for acquiring protein data and using the small molecule screening model to screen and rank candidate small molecules to obtain the corresponding active compounds includes: Input protein sequence, three-dimensional structure, or a combination of both, and use the trained small molecule screening model to calculate its similarity score with candidate small molecules; All candidate small molecules are sorted by similarity score, and molecules with similarity scores above the threshold are identified as potential binding ligands to obtain the corresponding active compounds.

[0011] The present invention also provides a drug virtual screening system based on gene ontology-enhanced contrastive learning. The system is used to implement the aforementioned method and includes: a historical data acquisition module, a dual encoder construction module, a model training module, and a drug virtual screening module. The historical data acquisition module is used to collect historical datasets of protein-small molecule pairings; The dual encoder construction module is used to construct a multimodal dual encoder framework that supports sequence and structure dual-modal input, consisting of a protein encoder and a small molecule encoder. The model training module is used to perform comparative learning and affinity ranking joint training on the multimodal dual encoder framework based on the historical dataset to obtain a small molecule screening model. The drug virtual screening module is used to acquire protein data and use the small molecule screening model to screen and sort candidate small molecules to obtain the corresponding active compounds.

[0012] Preferably, the process of collecting historical datasets of protein-small molecule pairings includes: Obtain historical protein data with sequence and three-dimensional structure information from publicly available protein databases; Collect diverse historical small molecule structure data from publicly available compound databases; For each historical protein data point, the corresponding historical small molecule structure data is selected to form a protein-small molecule pairing dataset, thus obtaining the historical dataset.

[0013] Preferably, the multimodal dual encoder framework includes a protein encoder and a small molecule encoder; The small molecule encoder includes: an input layer for receiving atomic type information and three-dimensional coordinate information of the small molecule; a graph neural network encoding layer for extracting local chemical features and spatial geometric relationships of atoms; and an embedding layer for outputting the overall feature representation of the small molecule. The protein encoder includes: a pre-trained language model encoding unit for receiving protein sequences; a spatial geometric feature extraction unit for receiving the three-dimensional structure of proteins; and an embedding layer for outputting a unified protein representation. The outputs of the small molecule encoder and the protein encoder are mapped to a shared high-dimensional feature space for similarity calculation and comparative learning.

[0014] Preferably, the process of jointly training the multimodal dual encoder framework through contrastive learning and affinity ranking based on the historical dataset to obtain the small molecule screening model includes: A contrastive learning strategy is adopted to construct protein-small molecule positive and negative sample pairs based on the historical dataset and to perform embedding representation alignment training. During training, contrastive learning tasks of protein-gene ontology and gene ontology-gene ontology are introduced to enhance the functional semantics of the protein encoder. The affinity ranking of small molecules corresponding to the same protein is optimized using a ranking loss function. Output the small molecule screening model completed through joint training.

[0015] Preferably, the process of acquiring protein data and using the small molecule screening model to screen and rank candidate small molecules to obtain the corresponding active compounds includes: Input protein sequence, three-dimensional structure, or a combination of both, and use the trained small molecule screening model to calculate its similarity score with candidate small molecules; All candidate small molecules are sorted by similarity score, and molecules with similarity scores above the threshold are identified as potential binding ligands to obtain the corresponding active compounds.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention is the first to introduce a gene ontology-enhanced contrastive learning strategy into a virtual drug screening task. By incorporating functional semantic information into protein representations, it improves the model's generalization ability to novel targets and proteins with incomplete functional annotations. This invention proposes a small molecule screening method combining protein sequence and 3D structure in a dual-modal approach, and introduces an uncertainty-aware fusion mechanism when both sequence and structure are present, effectively alleviating the instability of single-modal predictions and improving the robustness and reliability of prediction results. Through protein-gene ontology and gene ontology-gene ontology contrastive learning enhancement modules, this invention establishes cross-semantic level constraints, enabling protein representations to simultaneously retain structural features and functional semantics, thereby improving the model's applicability in zero-sample and cross-domain scenarios. The proposed method and system can be widely applied to various virtual drug screening scenarios, suitable for situations with only protein sequences, as well as situations with structural information but lacking drug pockets. It can also leverage synergistic advantages when multimodal data is present simultaneously, demonstrating promising application prospects. Attached Figure Description

[0017] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the method flow according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the dual encoder architecture according to an embodiment of the present invention; Figure 3 This is a framework diagram for gene ontology enhancement according to an embodiment of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0021] Example 1 In this embodiment, as Figure 1As shown, a virtual drug screening method based on gene ontology-enhanced contrastive learning includes the following steps: S1. Collect historical datasets of protein-small molecule pairings.

[0022] The method for obtaining the historical dataset includes: acquiring protein data with sequence and three-dimensional structural information from publicly available protein databases; collecting diverse small molecule structural data from publicly available compound databases; and for each historical protein, selecting the corresponding small molecule to form a protein-small molecule pairing dataset, thus obtaining the historical dataset.

[0023] In this embodiment, authoritative databases such as PDBBind and ChEMBL are preferably used to ensure that the proteins cover different families and have reliable structural information. The small molecule portion is obtained from publicly available medicinal chemistry databases, including known active compounds and affinity indices from the same unit and under the same experimental conditions. Through a reasonable pairing strategy (finding affinity values ​​for a protein and many other small molecules from existing literature under the same experimental conditions and from the same unit, which facilitates affinity ranking because affinity values ​​are affected by experimental conditions and there are different evaluation indices), a training set that is both broad and discriminative is constructed to provide diverse inputs for subsequent model learning.

[0024] S2. Construct a dual encoder architecture consisting of a protein encoder and a small molecule encoder.

[0025] In this embodiment, as Figure 2 As shown, the dual encoder comprises a protein encoder and a small molecule encoder. The protein encoder supports two modalities of input: sequence and structure. For the sequence part, contextual semantic information is extracted by a pre-trained large language model, while for the structure part, spatial topological features are extracted by a pre-trained structure-aware large language model, ultimately unified into an embedding space. The small molecule encoder uses a pre-trained structural network, receiving inputs such as atom type, chemical bonds, and three-dimensional coordinates, and outputs a comprehensive small molecule representation. The outputs of the protein and small molecule encoders are mapped to a shared feature space for similarity calculation.

[0026] Furthermore, auxiliary branches are introduced into this architecture to perform protein-gene ontology and gene ontology-gene ontology contrastive learning tasks, thereby enhancing functional semantics and improving the functional consistency and generalization ability of protein representation.

[0027] Gene ontology enhancement includes: Construct protein-gene ontology triples as positive samples; Construct gene ontology-gene ontology triplets as positive samples; By sampling unrelated gene ontology entries within the same semantic branch, difficult negative samples are formed and used for contrastive learning.

[0028] Contrastive learning loss functions include: Protein-small molecule contrast loss is used to maximize the similarity of known binding pairs and minimize the similarity of non-binding pairs. Affinity ranking loss is used to optimize the list-level ranking of candidate small molecules based on known binding strength within the same protein. Gene ontology contrast loss, including protein-gene ontology contrast loss and gene ontology-gene ontology contrast loss, is used to ensure that protein representations are consistent with their corresponding gene ontology entries in the functional semantic space and to distinguish irrelevant gene ontology entries.

[0029] The specific implementation process is as follows: Figure 2 As shown, a dual-encoder architecture consisting of a protein encoder and a small molecule encoder is constructed. The protein encoder supports two modal inputs: sequence and structure. The sequence modal input is... ,in Indicates sequence-level amino acid characteristics, The feature dimension is ; the structural modal input is . ,in Indicates structural level amino acid characteristics, This is the feature dimension. After extraction by the pre-trained model, the overall sequence representation is obtained. Representation of the overall structure And fused into a unified protein embedding: ; The small molecule encoder receives information such as the type of molecules, chemical bonds, and three-dimensional coordinates, which is represented as follows: ,in The structural-level atomic features are represented and encoded into a holistic molecular characterization through a pre-trained structural network: ; Protein and small molecule representations are mapped to a shared feature space, and their matching degree is calculated using cosine similarity. ; Furthermore, an auxiliary branch is introduced into this architecture to perform contrastive learning tasks between protein-gene ontology (P2G) and gene ontology-gene ontology (G2G), thereby enhancing functional semantic representation. Specifically, a P2G triple is defined as... ,in This indicates a protein, and 'g' represents the associated gene ontology (GO) term. Indicates the relationship between the two; a G2G triple is defined as follows: ,in and These are GO terms that have hierarchical or semantic relationships. This relates to the relationship between the two.

[0030] During the negative sampling process, the P2G and G2G triples employ a GO-based tail entity replacement strategy to generate negative samples that are both challenging and functionally consistent. Specifically, for the P2G triples... Maintain protein p and its relationship No change, replace the tail entity Negative GO terminology To ensure functional consistency and increase the difficulty of discrimination, the aforementioned From and Sampling is performed within the GO set of the same semantic dimension, but not within the annotation set of protein p. It should be noted that all GO terms belong to one of the three major semantic dimensions: molecular function, biological process (or cellular component). Sampling within the same dimension maintains ontology consistency and allows the model to learn fine-grained semantic differences. The negative sampling set is defined as follows: ; in, Indicates and A set of GO terms belonging to the same semantic dimension. This represents the set of GO annotations associated with protein p.

[0031] For G2G triplet Its negative sampling is achieved by replacing the tail entity GO terminology. The negative samples... Selected from and Same GO semantic dimension, but original head entities need to be excluded. With tail entity This is to avoid generating obvious negative samples. Its formal definition is as follows: ; In the design of the scoring function, this invention adopts a TransE-based approach: ; in, , and These represent the features of the head entity, tail entity, and relation in a triple. It should be noted that GO terms, whether used as head or tail entities, are modeled using a shared GO encoder to achieve a unified representation of P2G and G2G triples.

[0032] Furthermore, the GO contrastive loss is defined as follows: ; in, For the Sigmoid function, These are the boundary parameters. Through the auxiliary branches, the model can be enhanced at the protein function level, thereby improving the consistency and generalization ability of protein representation.

[0033] S3. Jointly train the dual encoder architecture based on historical datasets to obtain a small molecule screening model.

[0034] In this embodiment, the training process employs a joint optimization strategy, combining contrastive learning, ranking learning, and gene ontology enhancement mechanisms to obtain a small molecule screening model with consistency between structural and functional semantics. Specifically, it includes the following steps: First, positive and negative protein-small molecule sample pairs are constructed based on historical datasets. A contrastive learning loss function is used to maximize the similarity of known bound protein-small molecule pairs in the feature space, while unbound pairs are distinguished, thus ensuring that the representation of proteins and small molecules can capture binding correlations.

[0035] Secondly, for the same protein, candidate small molecules are ranked using known binding strength information, and list-level optimization is achieved using an affinity ranking loss function. By introducing this ranking mechanism during training, the model can learn to distinguish between strongly and weakly binding small molecules, improving the refinement and practicality of the screening results.

[0036] Secondly, a gene ontology enhancement module is introduced, such as... Figure 3 As shown, positive sample triples of protein-gene ontology and gene ontology-gene ontology are constructed, and unrelated entries in the same semantic branch are sampled as difficult negative samples. Contrastive learning loss is used to optimize the protein representation in the functional semantic space, ensuring consistency with the corresponding gene ontology and effectively distinguishing irrelevant gene ontology entries. This mechanism guarantees that the protein representation not only includes sequence and structural features but also incorporates functional semantic information, enhancing the model's generalization ability on new targets and proteins with lost functions.

[0037] Finally, the embeddings of the protein encoder and the small molecule encoder are mapped to a unified high-dimensional feature space. By combining protein-small molecule contrast loss, affinity ranking loss and gene ontology contrast loss, joint training and parameter optimization are performed to obtain a small molecule screening model with excellent generalization performance.

[0038] The specific implementation process is as follows: After the model structure is determined, this embodiment further proposes a joint optimization training method to improve the applicability of the screening model on unknown targets. Specifically, the training objective consists of three parts.

[0039] First, using known protein-small molecule pairings from historical datasets, a contrastive learning loss is employed: ; in, Indicates the embedding similarity between proteins and small molecules. As a positive sample, For negative samples, This is a hyperparameter. This mechanism ensures that positive samples are brought closer together in the feature space, while negative samples are effectively distinguished.

[0040] Secondly, considering the differences in binding strength between different small molecules on the same protein, affinity ranking loss is further introduced: ; in, , All are experimentally measured bonding strengths. and To be compatible with small molecules and Model prediction score. This loss function drives the model to learn to distinguish between strongly and weakly binding molecules, thereby enabling more refined screening.

[0041] Finally, to enhance the functional consistency of protein representation, a gene ontology contrastive learning mechanism is introduced. By employing GO contrastive loss, functional semantics are incorporated into the protein representation, enabling it to maintain reasonable functional discriminative ability even in the absence of structural or experimental labels.

[0042] Considering the above three types of losses, the joint training objective is: ; in, For balancing parameters, For protein-gene ontology contrast loss, The loss is calculated as gene ontology-gene ontology contrastive loss. Through this optimization strategy, the model achieves functional semantic enhancement while learning structure-ligand binding patterns, thus surpassing the generalization performance of existing methods.

[0043] Finally, the embeddings of the protein encoder and the small molecule encoder are mapped to a unified high-dimensional feature space. By combining protein-small molecule contrast loss, affinity ranking loss and gene ontology contrast loss, joint training and parameter optimization are performed to obtain a small molecule screening model with excellent generalization performance.

[0044] S4. Obtain protein data, screen and sort candidate small molecules based on a small molecule screening model, and obtain the corresponding active compounds.

[0045] In this embodiment, the method for screening and sorting candidate small molecules includes: inputting protein sequence, three-dimensional structure, or a combination of both information; using a trained small molecule screening model to calculate the similarity score between the input and the candidate small molecules; sorting all candidate small molecules according to their similarity scores; and using molecules with scores higher than a threshold as potential binding ligands to obtain the corresponding active compounds.

[0046] Specifically, when inputting a protein sequence Three-dimensional structure With small molecules At that time, the sequence encoders obtained during the training phase are first used respectively. Structure encoder and molecular encoder Extracted feature representation: ; ; ; When only a single modality is input, the similarity score is calculated using the corresponding protein representation and small molecule representation; when both sequence and structural information are input simultaneously, the similarity scores for the two modalities are calculated separately. ; ; in, For example, cosine similarity.

[0047] When sequences and structures are input simultaneously, the model integrates the prediction results of the two modalities through an uncertainty-aware fusion mechanism, improving the accuracy and stability of the screening. This method can make reliable predictions even without precise affinity labels and is adaptable to zero-sample or cross-domain scenarios.

[0048] Among them, the uncertainty perception fusion mechanism includes: The similarity calculation unit is used to calculate the small molecule similarity scores output by the protein sequence encoder and the protein structure encoder, respectively, and then sort them from largest to smallest. and ; An uncertainty assessment unit is used to calculate an entropy value based on the ranking results of the similarity scores, to reflect the magnitude of the uncertainty in the prediction: ; The probability is defined as: ; ; The dynamic weighting unit is used to assign weights to sequence prediction and structure prediction based on the magnitude of uncertainty. The weights are inversely proportional to the prediction entropy. The fusion output unit is used to fuse the weighted results and generate the final small molecule ranking score: ; in, To sort by average, This is a hyperparameter used to control sensitivity to uncertainty. The negative sign ensures that a larger score indicates a higher ranking.

[0049] In this embodiment, the applicable scenarios for the method include: When only protein sequences are available but no three-dimensional structures are available, sequence encoders are used for small molecule screening. When a protein's three-dimensional structure exists but a druggable pocket is lacking, a structural encoder is used for small molecule screening. When both sequence and structural information are available, an uncertainty-aware fusion mechanism is used for multimodal prediction. When proteins are novel targets or have incomplete functional annotations, generalization ability can be enhanced by relying on gene ontology enhancement modules, thereby obtaining active small molecules.

[0050] Example 2 In this embodiment, the present invention provides a drug virtual screening system based on gene ontology-enhanced contrastive learning, comprising: a historical data acquisition module, a dual encoder construction module, a model training module, and a drug virtual screening module.

[0051] The historical data acquisition module is used to collect historical datasets of protein-small molecule pairings.

[0052] The workflow of the historical data acquisition module includes: acquiring complete historical protein data with three-dimensional structural information from publicly available protein databases; acquiring diverse historical small molecule structure data from publicly available compound databases; and for each historical protein data set, selecting the corresponding historical small molecule structure data to form a protein-small molecule pairing dataset, thus obtaining the historical dataset.

[0053] The dual encoder building block is used to construct a dual encoder architecture consisting of a protein encoder and a small molecule encoder, and a gene ontology enhancement module is introduced to optimize the functional consistency of protein representation.

[0054] The model training module is used to jointly train the dual encoder architecture based on historical datasets. During the training process, contrastive learning, ranking learning, and gene ontology enhancement mechanisms are introduced to obtain a small molecule screening model with strong generalization ability.

[0055] The workflow of the model training module includes: using a contrastive learning strategy, constructing protein-small molecule positive and negative sample pairs based on the historical dataset, and performing embedding representation alignment training; introducing contrastive learning tasks of protein-gene ontology and gene ontology-gene ontology during the training process to enhance the functional semantics of the protein encoder; using a ranking loss function to optimize the affinity ranking of small molecules corresponding to the same protein; and completing the model training to obtain a jointly trained small molecule screening model.

[0056] The drug virtual screening module is used to input protein data during the inference stage, call the small molecule screening model to screen and sort candidate small molecules, and finally obtain the corresponding active compounds.

[0057] In this embodiment, the multimodal dual encoder framework includes a protein encoder and a small molecule encoder; The small molecule encoder includes: an input layer for receiving atomic type information and three-dimensional coordinate information of the small molecule; a graph neural network encoding layer for extracting local chemical features and spatial geometric relationships of atoms; and an embedding layer for outputting the overall feature representation of the small molecule. The protein encoder includes: a pre-trained language model encoding unit for receiving protein sequences; a spatial geometric feature extraction unit for receiving the three-dimensional structure of proteins; and an embedding layer for outputting a unified protein representation. The outputs of the small molecule encoder and the protein encoder are mapped to a shared high-dimensional feature space for similarity calculation and comparative learning.

[0058] In this embodiment, the process of acquiring protein data and using the small molecule screening model to screen and sort candidate small molecules to obtain the corresponding active compounds includes: Input protein sequence, three-dimensional structure, or a combination of both, and use the trained small molecule screening model to calculate its similarity score with candidate small molecules; All candidate small molecules are sorted by similarity score, and molecules with similarity scores above the threshold are identified as potential binding ligands to obtain the corresponding active compounds.

[0059] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A virtual drug screening method based on gene ontology-enhanced contrastive learning, characterized in that, The method includes: Collect historical datasets of protein-small molecule pairings; A multimodal dual encoder framework supporting sequence and structure dual-modal input was constructed, consisting of a protein encoder and a small molecule encoder. Based on the historical dataset, the multimodal dual encoder framework is jointly trained by contrastive learning and affinity ranking to obtain a small molecule screening model. Protein data is acquired, and candidate small molecules are screened and sorted using the small molecule screening model to obtain the corresponding active compounds.

2. The method according to claim 1, characterized in that, Methods for collecting historical datasets of protein-small molecule pairings include: Obtain historical protein data with sequence and three-dimensional structure information from publicly available protein databases; Collect diverse historical small molecule structure data from publicly available compound databases; For each historical protein data point, the corresponding historical small molecule structure data is selected to form a protein-small molecule pairing dataset, thus obtaining the historical dataset.

3. The method according to claim 1, characterized in that, The multimodal dual encoder framework includes: a protein encoder and a small molecule encoder; The small molecule encoder includes: an input layer for receiving atomic type information and three-dimensional coordinate information of the small molecule; a graph neural network encoding layer for extracting local chemical features and spatial geometric relationships of atoms; and an embedding layer for outputting the overall feature representation of the small molecule. The protein encoder includes: a pre-trained language model encoding unit for receiving protein sequences; a spatial geometric feature extraction unit for receiving the three-dimensional structure of proteins; and an embedding layer for outputting a unified protein representation. The outputs of the small molecule encoder and the protein encoder are mapped to a shared high-dimensional feature space for similarity calculation and comparative learning.

4. The method according to claim 1, characterized in that, Based on the historical dataset, the method for jointly training the multimodal dual encoder framework through contrastive learning and affinity ranking to obtain the small molecule screening model includes: A contrastive learning strategy is adopted to construct protein-small molecule positive and negative sample pairs based on the historical dataset and to perform embedding representation alignment training. During training, contrastive learning tasks of protein-gene ontology and gene ontology-gene ontology are introduced to enhance the functional semantics of the protein encoder. The affinity ranking of small molecules corresponding to the same protein is optimized using a ranking loss function. Output the small molecule screening model completed through joint training.

5. The method according to claim 1, characterized in that, The method for acquiring protein data and using the small molecule screening model to screen and rank candidate small molecules to obtain the corresponding active compounds includes: Input protein sequence, three-dimensional structure, or a combination of both, and use the trained small molecule screening model to calculate its similarity score with candidate small molecules; All candidate small molecules are sorted by similarity score, and molecules with similarity scores above the threshold are identified as potential binding ligands to obtain the corresponding active compounds.

6. A virtual drug screening system based on gene ontology-enhanced contrastive learning, the system being used to implement the method described in any one of claims 1-5, characterized in that, The system includes: a historical data acquisition module, a dual encoder construction module, a model training module, and a drug virtual screening module; The historical data acquisition module is used to collect historical datasets of protein-small molecule pairings; The dual encoder construction module is used to construct a multimodal dual encoder framework that supports sequence and structure dual-modal input, consisting of a protein encoder and a small molecule encoder. The model training module is used to perform comparative learning and affinity ranking joint training on the multimodal dual encoder framework based on the historical dataset to obtain a small molecule screening model. The drug virtual screening module is used to acquire protein data and use the small molecule screening model to screen and sort candidate small molecules to obtain the corresponding active compounds.

7. The system according to claim 6, characterized in that, The process of collecting historical datasets of protein-small molecule pairings includes: Obtain historical protein data with sequence and three-dimensional structure information from publicly available protein databases; Collect diverse historical small molecule structure data from publicly available compound databases; For each historical protein data point, the corresponding historical small molecule structure data is selected to form a protein-small molecule pairing dataset, thus obtaining the historical dataset.

8. The system according to claim 6, characterized in that, The multimodal dual encoder framework includes: a protein encoder and a small molecule encoder; The small molecule encoder includes: an input layer for receiving atomic type information and three-dimensional coordinate information of the small molecule; a graph neural network encoding layer for extracting local chemical features and spatial geometric relationships of atoms; and an embedding layer for outputting the overall feature representation of the small molecule. The protein encoder includes: a pre-trained language model encoding unit for receiving protein sequences; a spatial geometric feature extraction unit for receiving the three-dimensional structure of proteins; and an embedding layer for outputting a unified protein representation. The outputs of the small molecule encoder and the protein encoder are mapped to a shared high-dimensional feature space for similarity calculation and comparative learning.

9. The system according to claim 6, characterized in that, Based on the historical dataset, the process of jointly training the multimodal dual encoder framework through contrastive learning and affinity ranking to obtain the small molecule screening model includes: A contrastive learning strategy is adopted to construct protein-small molecule positive and negative sample pairs based on the historical dataset and to perform embedding representation alignment training. During training, contrastive learning tasks of protein-gene ontology and gene ontology-gene ontology are introduced to enhance the functional semantics of the protein encoder. The affinity ranking of small molecules corresponding to the same protein is optimized using a ranking loss function. Output the small molecule screening model completed through joint training.

10. The system according to claim 6, characterized in that, The process of acquiring protein data and using the small molecule screening model to screen and rank candidate small molecules to obtain the corresponding active compounds includes: Input protein sequence, three-dimensional structure, or a combination of both, and use the trained small molecule screening model to calculate its similarity score with candidate small molecules; All candidate small molecules are sorted by similarity score, and molecules with similarity scores above the threshold are identified as potential binding ligands to obtain the corresponding active compounds.