Molecular property prediction method based on multi-mode gating and comparative learning

By combining multimodal gating and contrastive learning methods with molecular graph structure and fingerprint features, the problem of graph neural networks relying on single-modal data in molecular property prediction is solved, and high accuracy and generalization ability are improved.

CN121075484AActive Publication Date: 2025-12-05LUDONG UNIVERSITY

Patent Information

Application Number
CN202511612007.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2025-12-05
Estimated Expiration
2045-11-06

AI Technical Summary

Technical Problem

Existing graph neural network methods rely on single-modality data and a large number of labeled samples in molecular property prediction, resulting in insufficient generalization ability and failure to fully explore and integrate complementary information between different modalities.

Method used

We employ a multimodal gating and contrastive learning approach, combining molecular graph structure and fingerprint features. Through contrastive learning pre-training and gating fusion mechanisms, we achieve high-precision prediction with a small amount of labeled data. This includes molecular data preprocessing, graph neural network encoder construction, contrastive learning pre-training, joint loss function optimization, and multimodal feature fusion.

Benefits of technology

It improves the model's ability to predict molecular properties, achieves high-precision prediction with a small amount of labeled data, and surpasses the generalization ability of traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121075484A_ABST
    Figure CN121075484A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of bioinformatics, and relates to a molecular property prediction method based on multi-modal gating and comparative learning, which comprises the technologies of comparative learning, graph neural network, cross-modal alignment, gating attention and the like. Firstly, data standardization and graph construction are carried out, and molecular fingerprint embedding is extracted; secondly, a heterogeneous dual-channel graph coding architecture is adopted, one path captures atom short-range interaction through an attention mechanism, the other path integrates a molecular global structure and long-range dependence, and complementary molecular representation is generated; then, a cross-modal attention mechanism is introduced, bidirectional association of graph and fingerprint features is achieved, and modal weights are adaptively and dynamically distributed through a gating fusion module; and finally, a comparison pre-training strategy is adopted, a molecular graph and fingerprints are utilized to construct a sample pair, and discriminative molecular representation is learned on unlabeled data. According to the method, the accuracy of molecular property prediction is remarkably improved, and an efficient and reliable calculation tool is provided for virtual drug screening and lead compound optimization.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of bioinformatics, and relates to a molecular property prediction method based on multi-modal gating and contrast learning, which comprises technologies such as contrast learning, graph neural network, cross-modal alignment and gating attention. BACKGROUND

[0002] Molecular property prediction is a key link in computational chemistry and drug design, and its goal is to accurately infer the physical and chemical properties, biological activity and toxicity of molecules through a computational model. The molecular properties are generated from the complex electronic structure, spatial configuration and interaction force between atoms, so how to effectively extract these implicit features from the molecular representation becomes a determining factor for prediction accuracy. This technology plays an irreplaceable role in reducing the cost of new drug research and development, reducing the dependence on animal experiments, and realizing the optimization of green chemical synthesis path.

[0003] Traditional molecular property prediction methods mainly rely on empirical descriptors or quantum chemical calculations. The former builds a statistical model through manually designed molecular fingerprints or physical and chemical parameters, but the design of descriptors is subjective and difficult to cover complex three-dimensional conformation information; the latter can accurately calculate the electronic structure, but the computational complexity increases exponentially with the number of atoms, which cannot be applied to large molecular systems. With the development of deep learning technology, graph neural networks model molecules as a topological graph of atomic nodes and chemical bond edges, realizing end-to-end molecular representation learning. However, existing graph neural network methods have two major limitations: one is highly dependent on the amount of labeled data, and the generalization ability decreases significantly in data-scarce scenarios; the other is that most models only use a single modality of graph structure, ignoring the complementarity of multi-source features such as molecular fingerprints. SUMMARY

[0004] Traditional molecular property prediction methods are limited in generalization due to their reliance on single modality data and large amounts of labeled samples, and they fail to fully exploit and integrate complementary information between different modalities of data. To solve this problem, the present application proposes a molecular property prediction method based on multi-modal gating and contrast learning, which combines molecular graph structure and fingerprint feature information, uses contrast learning pre-training and gating fusion mechanism, and realizes high-precision prediction with a small amount of labeled data, thereby improving the molecular property prediction ability of the model.

[0005] A molecular property prediction method based on multi-modal gating and contrast learning includes five processes of molecular data preprocessing and feature extraction, graph neural network encoder construction, contrast learning pre-training, joint loss function optimization, and multi-modal feature fusion, and the specific steps are as follows: Step 1, use the chemical information toolkit to parse the molecular SMILES (Simplified Molecular Input Line Entry System) data into a molecular graph structure containing 93-dimensional atomic node features and 11-dimensional chemical bond edge features, while extracting 881-dimensional PubChem fingerprints, 166-dimensional MACCS (Molecular ACCess System) fingerprints, and ErG (Extended Reduced Graph) fingerprints, and splicing them into a 1489-dimensional comprehensive fingerprint feature vector; Step 2, construct two parallel graph neural network encoders, where the graph attention encoder adopts a two-layer graph attention network, the first layer uses 10 attention heads to output 5120-dimensional features, and the second layer uses a single attention head to output 512-dimensional features; The graph isomorphism encoder adopts a three-layer graph isomorphism network, and the output dimensions are 512, 1024 and 512 in turn; And integrate 2 layers of Transformer encoder at the end of each encoder, each layer uses 8 attention heads and 2048-dimensional feedforward network to capture the global semantic dependency relationship of the molecule; Step 3, use the NT-Xent (Normalized Temperature-Scaled Cross-Entropy Loss) contrastive loss function, set the temperature parameter as a learnable variable, input the 512-dimensional features output by the graph attention and graph isomorphism encoders of the same molecule as positive sample pairs, and other molecule features in the current batch as negative samples, and maximize the cosine similarity of the positive sample pairs for self-supervised pre-training; Step 4, calculate the binary cross-entropy classification loss and contrastive learning loss in the downstream task training stage, where the contrastive loss is calculated after projecting the features into a 128-dimensional space, and the contrastive loss weight coefficient is set to 0.1, realizing joint optimization; Step 5, normalize the pre-trained graph structure features, project the fingerprint features through a fully connected network to 512 dimensions, apply a 4-head cross-modal attention mechanism to realize bidirectional interaction between graph features and fingerprint features, use a gated fusion module to dynamically adjust the contribution weight of the three feature sources, and finally input the fused features into a multi-layer perception prediction network containing normalization layer, GELU (Gaussian Error Linear Unit) activation and random inactivation, output the molecular property prediction result.

[0006] A molecular property prediction method based on multi-modal gating and contrastive learning, the implementation process of step 1 is as follows: Firstly, the SMILES string is parsed into a molecular object using the Chemical Information Toolkit, and the atomic feature vector is extracted, including atomic type, degree, number of hydrogen atoms, formal charge, chirality tag, and hybridization state. Secondly, the edge feature vector is extracted, including bond type, stereochemical information, and bond directionality, and an adjacency matrix is constructed to represent the atomic connection relationship. Finally, the PubChem fingerprint, MACCS key fingerprint, and ErG fingerprint are calculated and concatenated into a comprehensive fingerprint vector to provide input features for subsequent multi-modal fusion.

[0007] A molecular property prediction method based on multi-modal gating and contrastive learning, the implementation process of step 2 is as follows: Firstly, a graph attention encoder is constructed to enhance the information transmission between nodes through edge features. Secondly, a graph isomorphism encoder is constructed to aggregate neighbor node information through a multi-layer perceptron and complete feature update combined with edge features. Finally, a Transformer encoder layer is integrated at the end of the two encoders to capture global semantic dependency relationships in molecules using a multi-head self-attention mechanism, forming a parallel processing architecture.

[0008] A molecular property prediction method based on multi-modal gating and contrastive learning, the implementation process of step 3 is as follows: Firstly, the NT-Xent contrastive loss function is used to input the features generated by the graph attention and graph isomorphism encoders of the same molecule as positive sample pairs. Secondly, the features generated by the graph attention and graph isomorphism encoders of other molecules in the current training batch are input as negative samples, and the concentration of the similarity distribution is adjusted through a temperature parameter. Finally, the encoder parameters are optimized to make the positive sample pairs closer to each other in the feature space, and the negative sample pairs further apart, generating a discriminative general molecular representation.

[0009] A molecular property prediction method based on multi-modal gating and contrastive learning, the implementation process of step 4 is as follows: Firstly, the binary cross-entropy classification loss and contrastive learning loss are calculated simultaneously during the downstream task training phase, where the contrastive loss is calculated after projecting the graph features and fingerprint features into the contrastive space through a projection network. Secondly, the contrastive loss weight coefficient is set to dynamically balance the contribution of the two losses. Finally, the special label values are masked to exclude the gradient influence of invalid samples, achieving stable joint optimization training.

[0010] A molecular property prediction method based on multi-modal gating and contrastive learning, the implementation process of step 5 is as follows: Firstly, the graph structure features extracted by the pre-trained graph encoder are normalized, and the molecular fingerprint features are projected to the same dimension through a fully connected network; secondly, a cross-modal attention mechanism is applied to realize the bidirectional interaction between the graph features and the fingerprint features, and the feature enhancement is completed through the multi-head attention of the attention head; finally, a gating fusion module is used to dynamically adjust the contribution weight of the three feature sources, and the fused features are input into a multi-layer perceptron to output the final molecular property prediction result. BRIEF DESCRIPTION OF DRAWINGS

[0011] Figure 1 A molecular property prediction method flow chart based on multi-modal gating and contrast learning.

[0012] Figure 2 A double-encoder pre-training graph.

[0013] Figure 3 A contrast learning graph.

[0014] Figure 4 A multi-modal gating fusion mechanism fusion graph. DETAILED DESCRIPTION

[0015] The present application will be described in detail below in combination with the drawings and examples.

[0016] The purpose of the present application is to propose a molecular property prediction method flow chart based on multi-modal gating and contrast learning, Figure 1 A molecular property prediction method flow chart based on multi-modal gating and contrast learning, including five processes of molecular data preprocessing and feature extraction, graph neural network encoder construction, contrast learning pre-training, joint loss function optimization, and multi-modal feature fusion, and the specific implementation steps of the processes are as follows: Step 1, using the chemical information toolkit, the molecular SMILES data is parsed into a molecular graph structure containing atomic nodes and chemical bond edges, and multi-dimensional molecular fingerprint features such as PubChem fingerprint and MACCS fingerprint are extracted, including the following contents: Firstly, the SMILES string is parsed into a molecular object using the chemical information toolkit, specifically through the Chem.MolFromSmiles() function, and the molecules that fail to parse are filtered out; secondly, the atomic feature vector is extracted, including atomic type (44 elements are encoded by one-of-k, dimension 44), atomic degree (0-10, dimension 11), number of hydrogen atoms (0-10, dimension 11), formal charge (-2 to 2, dimension 5), chirality label (0-3, dimension 4), hybridization state (5 types, dimension 5), aromaticity (Boolean value, dimension 1), and normalized mass (multiplied by 0.01, dimension 1), each atom generates a 93-dimensional feature vector, and all features are normalized (divided by the total sum of features); then, the edge feature vector is extracted, including bond type (single bond, double bond, triple bond, aromatic bond, etc., encoded by one-hot, dimension 6) and stereochemistry information (0-4, encoded by one-hot, dimension 5), the total dimension of edge features is 11, and normalization is performed; at the same time, the molecular fingerprint features are calculated, including PubChem fingerprint (based on 881 SMARTS patterns, dimension 881), MACCS key fingerprint (dimension 166), and ErG fingerprint (dynamic calculation, variable dimension, finally adjusted to fixed dimension), all fingerprints are concatenated into a 1489-dimensional comprehensive vector; finally, the node-level features are aggregated into a molecule-level representation using global average pooling, specifically using the global_mean_pool function, the variable number of node features (dimension 93) are aggregated into a fixed-dimensional molecular graph feature (dimension 93), and these features are integrated into the final input data for the graph neural network, which is used for subsequent model training.

[0017] Step 2, based on the molecular graph in step 1, two parallel graph neural network encoders are constructed, integrating Transformer layers to capture global semantic dependency relationships of molecules; Figure 2 is a dual-encoder contrastive learning pre-trained graph, including the following: Firstly, the graph attention network encoder is constructed, which adopts a two-layer graph convolution structure: the first layer has an input dimension of 93, an output dimension of 512, and 10 attention heads, so the actual output dimension is 512 x 10 = 5120, an ELU activation function and random inactivation (p = 0.2) are used, and the edge features are projected to 93 dimensions through a fully connected layer; the second layer has an input dimension of 5120, an output dimension of 512, and 1 attention head, an ELU activation and random inactivation are used, and the edge features are projected to 512 dimensions; then, a Transformer encoder layer is integrated at the end of the graph attention encoder, specifically using the TransformerEncoder module, which contains 2 layers of encoder layers, each layer has 8 attention heads, a feedforward network dimension of 2048, and an input and output dimension of 512, and captures the global semantic dependency relationship of the molecule through a self-attention mechanism; secondly, the graph isomorphism network encoder is constructed, which adopts a three-layer graph isomorphism convolution structure: the first layer has an input dimension of 93, an output dimension of 512, uses an MLP (linear layer 93→512) for neighbor aggregation, an ELU activation, and the edge features are projected to 93 dimensions through a fully connected layer; the second layer has an input dimension of 512, an output dimension of 1024, an MLP (512→1024), an ELU activation, and the edge features are projected to 512 dimensions; the third layer has an input dimension of 1024, an output dimension of 512, an MLP (1024→512), an ELU activation, and the edge features are projected to 1024 dimensions; then, a Transformer encoder layer is also integrated at the end of the graph isomorphism encoder, with the same parameters as the graph attention encoder; finally, the two encoders process the same molecular graph input in parallel and output 512-dimensional molecular-level feature vectors respectively. These graph-level features are finally used for molecular property prediction tasks to achieve accurate modeling and analysis of molecular properties.

[0018] Step 3, use the NT-Xent contrastive loss function to perform self-supervised pre-training through a dual-encoder architecture to generate general molecular representations. Figure 3 is a contrastive learning graph, which includes the following: first, the NT-Xent contrastive loss function is used, and its mathematical expression is . Wherein, and respectively represent the feature vectors generated by the graph attention and graph isomorphism encoders for the same molecule, sim is the cosine similarity function, is a temperature parameter (default value 0.1, learnable range 0.05-0.3), Nis the batch size (default 128); secondly, construct positive and negative sample pairs: input the features generated by the graph attention and graph isomorphism encoders of the same molecule as positive sample pairs, and the features generated by other molecules in the current training batch as negative samples, excluding self-comparison through a mask matrix; then, adjust the sharpness of the similarity distribution through the temperature parameter, optimize the encoder parameters to maximize the cosine similarity of the positive sample pairs in the feature space and minimize the similarity of the negative sample pairs; finally, save the pre-trained encoder weights for downstream task initialization.

[0019] Step 4, calculate the binary cross-entropy classification loss and the contrastive learning loss simultaneously in the downstream task training phase, and balance the two losses through a weight coefficient for joint optimization. It includes the following contents: first, calculate the binary cross-entropy classification loss in the downstream task training phase, and its mathematical expression is . Wherein, is the true label, is the original output of the model, is the sigmoid function, N is the number of valid samples; secondly, calculate the contrastive learning loss: map the GNN features and fingerprint features to a 128-dimensional contrastive space through a projection network, and the projection network structure is a linear layer (512→256), a normalization layer, a GELU activation, a linear layer (256→128), then use the NT-Xent loss function to calculate the contrastive loss, and the temperature parameter is a learnable variable; then, set the contrastive loss weight coefficient to 0.1 to dynamically balance the two losses, and the total loss formula is ; finally, perform mask processing and only calculate the loss and gradient for valid samples to exclude the influence of invalid samples.

[0020] Step 5, dynamically weight and fuse the pre-trained graph structure features and molecular fingerprint features through cross-modal attention and gating fusion mechanism, and finally input the prediction network to output the molecular property prediction results. Figure 4 is the multi-modal gating fusion mechanism to fuse the graph, including the following contents: first, load the pre-trained graph attention and graph isomorphism encoders as feature extractors; input the molecular graph data in the training set into the graph attention and graph isomorphism encoders respectively to obtain two molecular graph representation vectors and ; secondly, project the molecular fingerprint features (1489 dimensions) to 512 dimensions through a fully connected network, and map them to the same dimension feature space as the graph representation, and the network structure is: linear layer (1489→1024), normalization layer, GELU activation, random inactivation (0.2), linear layer (1024→512) to obtain the fingerprint representation vector ; then, apply the cross-modal attention mechanism to realize the bidirectional interaction of graph features and fingerprint features; the first module takes the graph attention feature For Query, with fingerprint features For Key and Value, output fingerprint information enhanced graph feature representation _attn ; the second module takes fingerprint features For Query, with graph isomorphism features For Key and Value, output GNN information enhanced fingerprint representation _attn ; the original features are connected in residual connection with the enhanced features, to obtain _combined and _combined ; the two enhanced graph representations and fingerprint representations are spliced; then, the contribution weights of the three feature sources (enhanced graph attention, enhanced graph isomorphism, projected fingerprint) are dynamically adjusted using a gated fusion module, which is implemented by the following steps: the three 512-dimensional features are spliced into a 1536-dimensional vector, a weight matrix is generated through a gating network (linear layer 1536→1536, sigmoid activation), and is multiplied element by element with the transformed features (linear layer 1536→1536, normalization layer, GELU), to output a 512-dimensional fusion feature; finally, the fusion feature is input into a prediction network, and the network structure is: linear layer (512→512), normalization layer, GELU activation, random dropout (0.2), linear layer (512→256), normalization layer, GELU, random dropout, linear layer (256→128), GELU, linear layer (128→output dimension), to output the molecular property prediction result; for classification tasks, a sigmoid activation function is used to output the probability, and for regression tasks, the numerical value is directly output.

[0021] In the pre-training stage, the model uses the Adam optimizer with a learning rate of and a weight decay of , and is trained for 200 cycles with a batch size of 128; in the downstream prediction task, a hierarchical learning rate strategy is used, in which the encoder layer learning rate is , the prediction head learning rate is , and the differential weight decay of and is set, and the multi-task joint optimization is realized through a contrast loss weight coefficient of 0.1.

[0022] The initial training stage of the method utilizes 306,000 compounds recorded in ZINC 15. Subsequently, fine-tuning and evaluation are performed on six downstream tasks (such as BBBP, HIV) covering different properties such as blood barrier penetration and side effects, with an average AUC value exceeding previous achievements by 1.67%, reaching 0.8431, which verifies the superior generalization ability.

[0023] The foregoing detailed description of the above examples is further to provide a detailed description of the present application, but it is not to be construed as a limitation to the scope of the present application. Within the concept of the present application, those of ordinary skill in the art can make a number of related simple deductions or substitutions for other examples, all of which are considered to be within the scope of protection of the present application.

Claims

1. A method for molecular property prediction based on multi-modal gating and contrastive learning, characterized in that, The method comprises the following steps: Step 1, using the chemical information toolkit to parse the molecular SMILES data into a molecular graph structure containing 93-dimensional atomic node features and 11-dimensional chemical bond edge features, while extracting 881-dimensional PubChem fingerprints, 166-dimensional MACCS fingerprints and ErG fingerprints, and splicing them into a 1489-dimensional comprehensive fingerprint feature vector; Step 2, constructing two parallel graph neural network encoders, wherein the graph attention encoder adopts a two-layer graph attention network, the first layer uses 10 attention heads to output 5120-dimensional features, and the second layer uses a single attention head to output 512-dimensional features; the graph isomorphism encoder adopts a three-layer graph isomorphism network, and the output dimensions are 512, 1024 and 512 in turn; and a 2-layer Transformer encoder is integrated at the end of each encoder, each layer uses 8 attention heads and a 2048-dimensional feedforward network to capture the global semantic dependency relationship of the molecule; Step 3, using the NT-Xent contrastive loss function, setting the temperature parameter as a learnable variable, inputting the 512-dimensional features output by the graph attention and graph isomorphism encoders of the same molecule as positive sample pairs, and other molecule features in the current batch as negative samples, and maximizing the cosine similarity of the positive sample pairs for self-supervised pre-training; Step 4, in the downstream task training stage, simultaneously calculate the binary cross-entropy classification loss and the contrastive learning loss, wherein the contrastive loss is calculated after the features are mapped to a 128-dimensional space by a projection network, the contrastive loss weight coefficient is set to 0.1, and joint optimization is realized; Step 5, normalizing the pre-trained graph structure features, projecting the fingerprint features to 512-dimensional space through a fully connected network, applying a 4-head cross-modal attention mechanism to realize bidirectional interaction between graph features and fingerprint features, using a gated fusion module to dynamically adjust the contribution weight of the three feature sources, and finally inputting the fused features into a multi-layer perceptron prediction network containing a normalization layer, a GELU activation and a random dropout to output the molecular property prediction results.

2. The method of claim 1, wherein, The specific implementation process of step 1 includes the following steps: first, parse the SMILES string into a molecular object using the chemical information toolkit, specifically through the Chem.MolFromSmiles function, and filter the molecules that fail to parse; second, extract the atomic feature vector, including atomic type, atomic degree, number of hydrogen atoms, formal charge, chirality label, hybridization state, aromaticity, and normalized mass, generate a 93-dimensional feature vector for each atom, and normalize all features; then, extract the edge feature vector, including bond type and stereochemical information, with a total dimension of 11, and normalize; at the same time, calculate the molecular fingerprint features, including PubChem fingerprint, MACCS key fingerprint, and ErG fingerprint, and concatenate all fingerprints into a 1489-dimensional comprehensive vector; finally, aggregate the node-level features into a molecule-level representation through global average pooling, specifically using the global_mean_pool function to aggregate a variable number of node features into a fixed-dimensional molecular graph feature, and integrate these features into the final input data for the graph neural network, which is used for subsequent model training.

3. The multi-modal gating and contrast learning-based molecular property prediction method of claim 1, the specific implementation process of step 2 being as follows: first, construct a graph attention network encoder with a two-layer graph convolution structure: the first layer has an input dimension of 93, an output dimension of 512, and 10 attention heads, so the actual output dimension is 512x10=5120, uses ELU activation function and random inactivation, and projects the edge features to 93 dimensions through a fully connected layer; the second layer has an input dimension of 5120, an output dimension of 512, and 1 attention head, uses ELU activation and random inactivation, and projects the edge features to 512 dimensions; then, integrate a Transformer encoder layer at the end of the graph attention encoder, specifically using the TransformerEncoder module, which contains 2 encoder layers, each with 8 attention heads, a feedforward network dimension of 2048, and an input and output dimension of 512, to capture the global semantic dependency of the molecule through self-attention mechanism; second, construct a graph isomorphism network encoder with a three-layer graph isomorphism convolution structure: the first layer has an input dimension of 93, an output dimension of 512, uses MLP for neighbor aggregation, ELU activation, and projects the edge features to 93 dimensions through a fully connected layer; the second layer has an input dimension of 512, an output dimension of 1024, MLP, ELU activation, and projects the edge features to 512 dimensions; the third layer has an input dimension of 1024, an output dimension of 512, MLP, ELU activation, and projects the edge features to 1024 dimensions; then, also integrate a Transformer encoder layer at the end of the graph isomorphism encoder with the same parameters as the graph attention encoder; finally, the two encoders process the same molecular graph input in parallel and output 512-dimensional molecular-level feature vectors respectively; these graph-level features are finally used for the prediction task of molecular properties, realizing accurate modeling and analysis of molecular properties.

4. The method of claim 1, wherein the step 3 is implemented as follows: first, an NT-Xent contrastive loss function is used, and the mathematical expression thereof is: ; wherein, and denote the feature vectors generated by the graph attention and graph isomorphism encoders for the same molecule, respectively, and sim is the cosine similarity function, is the temperature parameter, N is the batch size; Secondly, construct positive and negative sample pairs: the features generated by the graph attention and graph isomorphism encoder are input into the same molecule as positive sample pairs, and the features generated by other molecules in the current training batch are used as negative samples, and the self-comparison is excluded by the mask matrix; Then, adjust the sharpness of the similarity distribution through the temperature parameter, optimize the encoder parameters to maximize the cosine similarity of the positive sample pairs in the feature space and minimize the similarity of the negative sample pairs; Finally, save the pre-trained encoder weights for downstream task initialization.

5. The method of claim 1, wherein the step 4 is implemented as follows: first, in the downstream task training phase, a binary cross-entropy classification loss is calculated, and the mathematical expression is: ; wherein, For true labels, for model original output, for sigmoid function, N is the number of valid samples; second, calculate the contrastive learning loss: map the GNN features and fingerprint features to a 128-dimensional contrastive space through the projection network, the projection network structure is linear layer, normalization layer, GELU activation, linear layer, and then use the NT-Xent loss function to calculate the contrastive loss, and the temperature parameter is a learnable variable; then, set the contrastive loss weight coefficient to 0.1, dynamically balance the two losses, and the total loss formula is: ; Finally, perform mask processing, only calculate the loss and gradient for valid samples, and exclude the influence of invalid samples.

6. The method according to claim 1, wherein the step 5 is implemented as follows: first, load the pre-trained graph attention and graph isomorphism encoders as feature extractors; input the molecular graph data in the training set into the graph attention and graph isomorphism encoders respectively to obtain two molecular graph representation vectors and ; second, project the molecular fingerprint features to 512 dimensions through a fully connected network to map them to the same feature space as the graph representation, and the network structure is: linear layer, normalization layer, GELU activation, random dropout, linear layer to obtain the fingerprint representation vector; then, apply a cross-modal attention mechanism to realize the bidirectional interaction between the graph features and the fingerprint features; the first module takes the graph attention features as the query, the fingerprint features as the key and the value, and outputs the graph feature representation enhanced by the fingerprint information _attn ; the second module takes the fingerprint features as the query, the graph isomorphism features as the key and the value, and outputs the fingerprint representation enhanced by the GNN information _attn ; connect the original features and the enhanced features in residual to obtain _combined and _ combined ; splice the two enhanced graph representations and fingerprint representations; then, use a gating fusion module to dynamically adjust the contribution weight of the three feature sources, which is implemented as follows: splice the three 512-dimensional features into a 1536-dimensional vector, generate a weight matrix through a gating network, and multiply it with the transformed features element by element to output a 512-dimensional fusion feature; finally, input the fusion feature into the prediction network, and the network structure is: linear layer, normalization layer, GELU activation, random dropout, linear layer, normalization layer, GELU, random dropout, linear layer, GELU, linear layer, and output the molecular property prediction result; for the classification task, use the sigmoid activation function to output the probability, and for the regression task, directly output the numerical value.​​​​

Citation Information

Patent Citations

  • Molecular representation prediction method based on multiple modes

    CN117292764A

  • Molecular property prediction method based on local graph features and global relationship

    CN119339827A

  • Drug target affinity prediction method and system based on multi-scale protein attention mechanism

    CN120783870A

  • Reaction site prediction method and device based on chemical and physical prior driving

    CN120877896A

  • Molecular graph representation learning method based on contrastive learning

    US20230052865A1

Cited By

  • Antibody drug conjugate property prediction method based on multi-modal fusion

    CN121281625A

  • New energy power prediction method and system based on multi-scale state decomposition mechanism

    CN121503819A

  • A new energy power prediction method and system based on a multi-scale state decomposition mechanism

    CN121503819B

  • Method for generating multi-mode synthesizable molecules perceived by chemical reaction

    CN121789834A

  • Reaction optimization method based on double-view comparative learning and double-layer attention

    CN122067649A