Molecular multi-modal representation learning method and device, equipment and storage medium
By using multimodal data and real label data for comparison learning model training, the problem of low accuracy in the prior art comparison learning model is solved, and more accurate molecular properties prediction is achieved.
Patent Information
- Application Number
- CN202510411401.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-08-05
AI Technical Summary
The existing comparison learning model has low accuracy in molecular properties prediction, mainly due to excessive dependence on binary labels, resulting in the sample data being misjudged as negative samples.
Multimodal data is used as the model input, and the comparative learning model is trained in combination with real label data. By predicting the probability value of the multimodal data as a positive or negative sample, the model is optimized until the training termination condition is met.
The accuracy of the comparative learning model is improved, so that it can be accurate to the probability value, and the accuracy of molecular properties prediction is improved.
Smart Images

Figure CN120432036A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of molecular representation learning technology, and in particular to a molecular multimodal representation learning method, apparatus, device and storage medium. Background Art
[0002] With the continuous development of science and technology, the field of molecular representation learning is undergoing a paradigm shift from single modality to multimodal fusion.
[0003] Currently, to integrate multimodal data on molecular structure and molecular phenotype, contrastive learning techniques are often used to binary-weight positive and negative sample pairs. However, this contrastive learning approach, which relies too heavily on binary labels, can lead to sample data (i.e., multimodal data) being misclassified as negative, resulting in low accuracy in contrastive learning models.
[0004] Therefore, how to improve the accuracy of comparative learning models for molecular property prediction is an urgent problem that needs to be solved. Summary of the Invention
[0005] The main purpose of this application is to provide a molecular multimodal representation learning method, device, equipment and storage medium, aiming to improve the accuracy of comparative learning models used for molecular property prediction.
[0006] To achieve the above objectives, the present application provides a molecular multimodal representation learning method, which includes:
[0007] Acquire multimodal data of each sample molecule and true label data of each sample molecule, wherein the true label data represents the true properties of the sample molecule;
[0008] Inputting each of the multimodal data into a contrastive learning model in batches to obtain predicted label data of target data, wherein the target data is the multimodal data input into the contrastive learning model at that time, and the predicted label data represents a probability value of the multimodal data being a positive sample or a negative sample;
[0009] When the contrastive learning model does not meet the preset training termination condition, the contrastive learning model is optimized based on the predicted label data and the real label data to obtain a new contrastive learning model, and the step of batch inputting each multimodal data into the contrastive learning model is returned to execute until the contrastive learning model meets the training termination condition, and molecular property prediction is performed based on the contrastive learning model.
[0010] In one embodiment, the step of batch-inputting each of the multimodal data into the contrastive learning model to obtain predicted label data of the target data includes:
[0011] Inputting each of the multimodal data into a contrastive learning model in batches;
[0012] Extracting features from the target data using the encoder of the contrastive learning model to obtain a target feature vector of the target data;
[0013] De-noising the target feature vector using a denoising module of the contrastive learning model to obtain valid information in the target feature vector;
[0014] The effective information is fused through the expert module of the contrastive learning model to obtain a fusion feature;
[0015] The prediction module of the contrastive learning model performs molecular multimodal representation learning on the fusion features to obtain predicted label data of the target data.
[0016] In one embodiment, the target data includes at least one multimodal data, each multimodal data includes a molecular structure, a molecular cell image, and a molecular gene expression of a sample molecule, and the encoder includes a graph neural network, a first fully connected neural network, and a second fully connected neural network;
[0017] The step of extracting features from the target data using the encoder of the contrastive learning model to obtain a target feature vector of the target data includes:
[0018] Extracting features from the molecular structure using a graph neural network of the contrastive learning model to obtain a first feature vector, wherein the first feature vector represents a characteristic of the molecular structure;
[0019] Performing feature extraction on the molecular cell image using the first fully connected neural network of the contrastive learning model to obtain a second feature vector, wherein the second feature vector represents a cell morphological feature of the molecular cell image;
[0020] performing feature extraction on the molecular gene expression using a second fully connected neural network of the contrastive learning model to obtain a third feature vector, wherein the third feature vector represents a genetic feature of the molecular gene expression;
[0021] The first eigenvector, the second eigenvector, and the third eigenvector are used as target eigenvectors of the target data.
[0022] In one embodiment, the step of denoising the target feature vector using the denoising module of the contrastive learning model to obtain valid information in the target feature vector includes:
[0023] Performing information separation on the target feature vector using the denoising module of the contrastive learning model to obtain effective information and specific information of the target feature vector;
[0024] A Pearson correlation coefficient between the effective information and the specific information is determined, and the Pearson correlation coefficient is used as a denoising loss.
[0025] In one embodiment, after the step of denoising the target feature vector using the denoising module of the contrastive learning model to obtain valid information in the target feature vector, the method further includes:
[0026] Predicting, by a contrastive learning module of the contrastive learning model, alignment probabilities of cross-modal sample pairs in the target data based on the specific information, wherein the cross-modal sample pairs refer to unimodal data of two different modalities;
[0027] A contrastive loss of the cross-modal sample pair is determined based on the alignment probability, the valid information, and the specific information.
[0028] In one embodiment, the method further comprises:
[0029] Predicting, by the generation module of the contrastive learning model, a fourth eigenvector of the molecular cell image in the target data and a fifth eigenvector of the molecular gene expression in the target data based on the molecular structure in the target data;
[0030] In the case where the molecular cell image and / or molecular gene expression is missing in the target data, using the fourth eigenvector and / or the fifth eigenvector as the target eigenvector of the target data;
[0031] In a case where molecular cell images and molecular gene expressions are not missing in the target data, a reconstruction loss is determined based on the fourth eigenvector, the fifth eigenvector, and the target eigenvector.
[0032] In one embodiment, the step of optimizing the contrastive learning model based on the predicted label data and the real label data to obtain a new contrastive learning model includes:
[0033] Determining a cross entropy loss of the contrastive learning model based on the predicted label data and the true label data;
[0034] Determining a property prediction loss of the contrastive learning model based on one or more of the denoising loss, the contrast loss, the reconstruction loss, and the cross entropy loss;
[0035] The model parameters of the contrastive learning model are optimized based on the property prediction loss to obtain a new contrastive learning model.
[0036] In addition, to achieve the above objectives, the present application also provides a molecular multimodal representation learning device, the molecular multimodal representation learning device comprising:
[0037] an acquisition module, configured to acquire multimodal data of each sample molecule and true label data of each sample molecule, wherein the true label data represents the true properties of the sample molecule;
[0038] A prediction module, configured to batch-input each of the multimodal data into the contrastive learning model to obtain predicted label data of target data, wherein the target data is the multimodal data currently input into the contrastive learning model, and the predicted label data represents a probability value of the multimodal data being a positive sample or a negative sample;
[0039] A training termination module is used to optimize the contrastive learning model based on the predicted label data and the real label data to obtain a new contrastive learning model when the contrastive learning model does not meet the preset training termination conditions, and return to execute the step of batch inputting each multimodal data into the contrastive learning model until the contrastive learning model meets the training termination conditions, and perform molecular property prediction based on the contrastive learning model.
[0040] In addition, to achieve the above-mentioned purpose, the present application also provides a storage medium, which is a computer-readable storage medium, and the computer-readable storage medium stores a program for implementing the molecular multimodal representation learning method. The program for implementing the molecular multimodal representation learning method is executed by a processor to implement the steps of the molecular multimodal representation learning method as described above.
[0041] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the molecular multimodal representation learning method as described above.
[0042] The present application provides a molecular multimodal representation learning method, which first obtains the multimodal data of each sample molecule and the true label data of each sample molecule, wherein the true label data represents the true properties of the sample molecule; the multimodal data are input into a contrastive learning model in batches to obtain predicted label data of the multimodal data input into the contrastive learning model at that time, and the predicted label data represents the probability value of the multimodal data being a positive sample or a negative sample; when the contrastive learning model does not meet the preset training termination condition, the contrastive learning model is optimized based on the predicted label data and the true label data to obtain a new contrastive learning model, and the model training step is returned to execute until the contrastive learning model meets the training termination condition. The contrastive learning model obtained at this time can be used for molecular multimodal representation learning.
[0043] In summary, this application uses multimodal data as model input data and the true properties of sample molecules as model training labels to train a contrastive learning model, enabling the contrastive learning model to predict the probability value of molecular multimodal data being a positive sample or a negative sample. Compared to the traditional method of relying on binary labels to classify sample data as positive or negative samples, the contrastive learning model in this application can accurately determine the probability value, thereby improving the accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0045] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0046] Figure 1 This is a flowchart of the first embodiment of the molecular multimodal representation learning method of the present application;
[0047] Figure 2 This is a diagram of the comparative learning model architecture involved in one embodiment of the molecular multimodal representation learning method of the present application;
[0048] Figure 3 A schematic diagram of a molecular multimodal representation learning process involved in an embodiment of the molecular multimodal representation learning method of the present application;
[0049] Figure 4 This is a schematic diagram of the module structure of the molecular multimodal representation learning device of this application;
[0050] Figure 5Schematic diagram of the device structure of the hardware operating environment involved in the molecular multimodal representation learning method in the embodiment of this application.
[0051] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0052] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0053] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0054] The main solution of the present application is: obtaining the multimodal data of each sample molecule and the true label data of each sample molecule, wherein the true label data represents the true properties of the sample molecule; inputting each multimodal data into the contrastive learning model in batches to obtain predicted label data of the target data, wherein the target data is the multimodal data input into the contrastive learning model at that time, and the predicted label data represents the probability value of the multimodal data being a positive sample or a negative sample; when the contrastive learning model does not meet the preset training termination condition, optimizing the contrastive learning model based on the predicted label data and the true label data to obtain a new contrastive learning model, and returning to execute the step of inputting each multimodal data into the contrastive learning model in batches until the contrastive learning model meets the training termination condition, and performing molecular property prediction based on the contrastive learning model.
[0055] Currently, to integrate multimodal data on molecular structure and molecular phenotype, contrastive learning techniques are often used to binary-weight positive and negative sample pairs. However, this contrastive learning approach, which relies too heavily on binary labels, can lead to sample data (i.e., multimodal data) being misclassified as negative, resulting in low accuracy in contrastive learning models.
[0056] Therefore, how to improve the accuracy of comparative learning models for molecular property prediction is an urgent problem that needs to be solved.
[0057] Specifically, breakthroughs in cell phenotypic screening technologies, such as Cell Painting and gene expression profiling (LINCS L1000), can precisely capture compound-induced cellular biological responses, providing critical data support for drug development. By integrating advanced cell imaging and gene expression profiling technologies, researchers can capture the dynamic responses of real cells and gain a deeper understanding of the mechanisms of action of potential drug candidates, thereby facilitating the development of new therapeutic options.
[0058] In order to integrate multimodal data of molecular structure and molecular phenotype, contrastive learning (CL) technology has been widely used in molecular characterization research. However, traditional contrastive learning technology has difficulty in capturing the complex relationship behind various biological and physical effects. Traditional contrastive learning methods use binary weighting to treat positive and negative sample pairs, equally approaching all positive sample pairs and rejecting all negative sample pairs. Although some studies have attempted to improve the contrastive learning framework by increasing the difficulty of similarity measurement, optimizing the sorting of positive sample pairs, or fusing label information, these methods are generally not suitable for molecular phenotypic data because molecular phenotypic data require more sophisticated calibration methods to distinguish experimental noise from key phenotypic variations. Therefore, the development of new contrastive learning methods for the characteristics of molecular phenotypic data is in urgent need of breakthroughs.
[0059] On the other hand, due to technical limitations and the complexity of biological processes, existing multimodal phenotypic datasets often suffer from missing modalities. To fully exploit the value of such data, effective integration is required in the absence of some modalities in order to obtain systematic understanding across biological dimensions. To fully utilize these data, integrating partially missing modalities is crucial for gaining comprehensive insights across biological dimensions. A straightforward strategy for incomplete multimodal learning is to use generative models to synthesize missing data. However, CLOOME and MIGA have the disadvantage of insufficient adaptability to missing modalities, resulting in the abandonment of a large amount of valid training data. Although InfoCORE addresses this problem through zero-value interpolation, it may introduce noise interference and thus affect the credibility of the model. This highlights the urgent need to develop a robust multimodal learning framework - a framework that must be able to effectively learn based on any available phenotypic subset, ultimately advancing our understanding of disease mechanisms and drug responses.
[0060] This application uses multimodal data as model input data and the true properties of sample molecules as model training labels to train a contrastive learning model, enabling the contrastive learning model to predict the probability value of molecular multimodal data being a positive or negative sample. Compared to the traditional method of relying on binary labels to classify sample data as positive or negative samples, the contrastive learning model in this application can accurately calculate the probability value, thereby improving the accuracy of the model.
[0061] It should be noted that the execution entity of the methods in each embodiment of the molecular multimodal representation learning method of this application can be a molecular representation learning system, or a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or a molecular multimodal representation learning device capable of performing the above functions, etc., and this embodiment does not specifically limit this. The following uses the molecular representation learning system as the execution entity as an example to illustrate this embodiment and the following embodiments.
[0062] Based on this, this application proposes a molecular multimodal representation learning method of the first embodiment, please refer to Figure 1 The molecular multimodal representation learning method includes steps S10 to S30:
[0063] Step S10, obtaining multimodal data of each sample molecule and true label data of each sample molecule, wherein the true label data represents the true properties of the sample molecule;
[0064] It should be noted that the molecules used for training are called sample molecules. The true label data of the sample molecules represent the true properties of the sample molecules, which can be understood as the properties of the sample molecules having a certain drug effect.
[0065] Step S20: Input each of the multimodal data into a contrastive learning model in batches to obtain predicted label data of the target data, wherein the target data is the multimodal data input into the contrastive learning model at that time, and the predicted label data represents the probability value of the multimodal data being a positive sample or a negative sample;
[0066] It should be noted that, during the training process of the contrastive learning model, each multimodal data is input into the contrastive learning model in batches, which can be understood as inputting at least one multimodal data into the contrastive learning model for each round of training. One or more multimodal data input into the contrastive learning model during the round of training is referred to as target data for distinction. After the target data is input into the contrastive learning model, the predicted label data of the target data is obtained, and the predicted label data represents the probability value of the multimodal data being a positive sample or a negative sample. Among them, it can be understood that the positive sample represents that the sample molecule corresponding to the target data has the properties in its true label data, and the negative sample represents that the sample molecule corresponding to the target data does not have the properties in its true label data. When the traditional contrastive learning method predicts the properties of molecules, a binary label (1 represents a positive sample, 0 represents a negative sample) is used, which will cause the molecule to actually have properties similar to those in the true label data, but will be classified as a negative sample by the model, and the embodiment of the present application can accurately represent the degree of closeness between the predicted properties of the molecule and the true properties of the molecule by outputting a probability value.
[0067] Step S30: When the contrastive learning model does not meet the preset training termination condition, the contrastive learning model is optimized based on the predicted label data and the real label data to obtain a new contrastive learning model, and the step of batch inputting each multimodal data into the contrastive learning model is returned to execute until the contrastive learning model meets the training termination condition, and molecular property prediction is performed based on the contrastive learning model.
[0068] It should be noted that the training termination condition of the model training can be that the number of training steps reaches a preset number of steps, or the model loss value is less than a preset threshold.
[0069] After obtaining the predicted label data output by the contrastive learning model, determine whether the contrastive learning model meets the preset training termination conditions. If so, the contrastive learning model at this time is used as the trained model for subsequent model testing and model application steps; if not, the contrastive learning model is optimized based on the predicted label data and the real label data to obtain a new contrastive learning model, and return to execute the step of batch inputting each multimodal data into the contrastive learning model, that is, return to execute the model training step, until the contrastive learning model meets the training termination conditions, and molecular multimodal representation learning can be performed based on the contrastive learning model.
[0070] Thus, the present embodiment uses multimodal data as model input data and the true properties of sample molecules as model training labels to train the contrastive learning model, enabling the contrastive learning model to predict the probability of molecular multimodal data being a positive or negative sample. Compared to the traditional method of relying on binary labels to classify sample data as positive or negative samples, the contrastive learning model in the present embodiment can accurately calculate the probability value, thereby improving the accuracy of the model.
[0071] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction and will not be repeated hereafter. On this basis, the step S20 may include:
[0072] Step S201, inputting the multimodal data into a contrastive learning model in batches;
[0073] During the model training process, each multimodal data is input into the contrastive learning model in batches. The embodiment of the present application does not limit the amount of multimodal data input into the contrastive learning model each time.
[0074] Step S202, performing feature extraction on the target data through the encoder of the contrastive learning model to obtain a target feature vector of the target data;
[0075] It should be noted that the contrastive learning model includes an encoder for extracting features from data input to the contrastive learning model to obtain an implicit representation.
[0076] The encoder of the contrast learning model extracts features from the input target data to obtain a feature vector of the target data (hereinafter referred to as the target feature vector for distinction).
[0077] In this embodiment, the target data includes at least one multimodal data, each multimodal data includes a molecular structure, a molecular cell image, and a molecular gene expression of a sample molecule, and the encoder includes a graph neural network, a first fully connected neural network, and a second fully connected neural network;
[0078] The step S202 may include:
[0079] Step A10: extracting features of the molecular structure using a graph neural network of the contrastive learning model to obtain a first feature vector, wherein the first feature vector represents a characteristic of the molecular structure;
[0080] It should be noted that the target data refers to one or more multimodal data input to the contrastive learning model during the round of training, that is, the target data includes at least one multimodal data, and each multimodal data includes the molecular structure, molecular cell image and molecular gene expression of a sample molecule. Among them, it is understandable that the molecular structure is also presented in the form of an image. The encoders in the contrastive learning model include three types, namely a graph neural network and two fully connected neural networks, which process the data of the three modalities in turn to generate a unimodal representation. In this way, the embodiment of the present application constructs a more comprehensive molecular representation through multimodal data, and has a richer understanding of the relationship between molecular features and phenotypes.
[0081] The molecular structure is feature extracted by the graph neural network of the comparative learning model to obtain a feature vector of the molecular structure (hereinafter referred to as the first feature vector for distinction). It should be understood that the first feature vector represents the characteristics of the molecular structure.
[0082] In a feasible implementation, the compound structure (i.e., molecular structure) of the sample molecule can be represented as an attribute graph G = (V, D), also known as a molecular graph. v Denotes the attribute of node V, D uv Represents the properties of edge uv, where |V|=n represents a set of n atoms (nodes), and |δ|=m represents a set of m edges (edges). Graph Neural Network (GNN) f G Learn to embed the molecular graph G into the feature vector h G , where the attributes of nodes and edges are propagated in each iteration. Formally, the lth iteration of GNN is:
[0083]
[0084] in, The representation of node v at layer l, N(v) is the set of all neighbor nodes of node v, X by its atomic properties v Initialize. Represents an aggregate function, Represents the update function. After L graph convolution, h (L) The neighborhood information of its L hops has been captured. Finally, the readout function is used to aggregate all the node representations output by the GNN L-th layer to obtain the representation h of the entire molecular graph. G , which is the first eigenvector:
[0085]
[0086] Step A20, performing feature extraction on the molecular cell image using the first fully connected neural network of the contrastive learning model to obtain a second feature vector, wherein the second feature vector represents a cell morphological feature of the molecular cell image;
[0087] It should be noted that the encoder used for feature extraction of molecular cell images is referred to as the first fully connected neural network for distinction.
[0088] The first fully connected neural network of the contrast learning model is used to extract features from the molecular cell image to obtain a feature vector of the molecular cell image (hereinafter referred to as the second feature vector for distinction). It should be understood that the second feature vector represents the morphological features of the cells in the molecular cell image.
[0089] In one feasible embodiment, the popular software tool CellProfiler (biological image analysis software) is used to extract the morphological features of cells from the manually drawn molecular cell images, and a fully connected neural network f is used to I , the network takes the extracted morphological features as input and encodes each image I into a feature vector called h I , which is the second eigenvector.
[0090] Step A30, performing feature extraction on the molecular gene expression using the second fully connected neural network of the contrastive learning model to obtain a third feature vector, wherein the third feature vector represents the genetic feature of the molecular gene expression;
[0091] It should be noted that the encoder used for feature extraction of molecular gene expression in the contrastive learning model is called the second fully connected neural network for distinction.
[0092] The second fully connected neural network of the comparative learning model is used to extract features of the molecular gene expression to obtain a feature vector of the molecular gene expression (hereinafter referred to as the third feature vector for distinction). It should be understood that the third feature vector represents the genetic characteristics of the molecular gene expression.
[0093] In one possible implementation, in addition to morphological features, gene expression data are also considered. These data are expressed as scalar values and obtained through the L1000 analysis method. To ensure the compatibility and significance of the feature set, the gene expression values are normalized to a range between 0 and 1 and analyzed using another fully connected neural network f E , converting the rescaled gene expression signatures into a new set of genetic signatures f E , which is the third eigenvector.
[0094] Step A40: Using the first eigenvector, the second eigenvector, and the third eigenvector as target eigenvectors of the target data.
[0095] The first eigenvector, the second eigenvector, and the third eigenvector are used as target eigenvectors of the target data.
[0096] Step S203, denoising the target feature vector using a denoising module of the contrastive learning model to obtain valid information in the target feature vector;
[0097] It should be noted that the contrastive learning model also includes a denoising module for removing redundant noise in the target feature vector.
[0098] The target feature vector is denoised by the denoising module of the comparative learning model to obtain denoised information (hereinafter referred to as effective information for distinction). It should be understood that effective information is the feature representing the same molecular properties in the feature vectors corresponding to each of the three modal data.
[0099] Step S204, fusing the effective information through the expert module of the contrastive learning model to obtain a fusion feature;
[0100] It should be noted that the contrastive learning model also includes an expert module for multimodal data fusion.
[0101] The effective information obtained after denoising is fused through the expert module of the comparative learning model to obtain the fusion feature.
[0102] Step S205 , performing molecular multimodal representation learning on the fusion features through the prediction module of the contrastive learning model to obtain predicted label data of the target data.
[0103] It should be noted that the contrastive learning model also includes a prediction module, which is used to perform molecular multimodal representation learning on the fusion features to obtain the probability value of the multimodal data being a positive sample or a negative sample.
[0104] In this embodiment, step S203 may include:
[0105] Step B10, performing information separation on the target feature vector using a denoising module of the contrastive learning model to obtain effective information and specific information of the target feature vector;
[0106] The denoising module of the contrastive learning model is used to separate the target feature vector, obtaining both effective and specific information. It should be noted that specific information refers to features in each single modality that characterize different molecular properties from those in other modalities, and that specific information also includes noise.
[0107] For example, embedding h in a molecular graph G For example, embedding h into the molecular graph G The calculation formula for separating the effective information and specific information in is:
[0108]
[0109] in, Represents the effective information from a single modality, Represents specific information, including experimental noise and significant changes (such as attribute cliffs). is a learnable parameter. σ is the activation function.
[0110] Step B20: determining the Pearson correlation coefficient between the effective information and the specific information, and using the Pearson correlation coefficient as the denoising loss.
[0111] The Pearson correlation coefficient between the independent effective information and the specific information is calculated and used as the loss of the denoising module (hereinafter referred to as denoising loss for distinction).
[0112] For example, embedding h in a molecular graph G For example, after separating the effective information and redundant noise, an independence assumption is introduced to enforce that the effective information and noise components in the feature space remain independent or approximately independent in a statistical sense. To achieve this goal, the embodiment of the present application adds a correlation loss to minimize the Pearson correlation coefficient between the effective information and the noise information, thereby reducing the statistical dependence between them and improving the quality of the learned representation:
[0113]
[0114] in, express and The covariance between them, σ represents the standard deviation, and the exponential function here strengthens the penalty for highly correlated feature components. The total denoising loss function is as follows:
[0115]
[0116] in, is the denoising loss corresponding to the molecular cell image, is the denoising loss corresponding to molecular gene expression.
[0117] In this way, the embodiment of the present application achieves accurate denoising by distinguishing between effective information and specific information in the target feature vector, thereby improving the quality of unimodal embedding.
[0118] In this embodiment, after step S203, the molecular multimodal representation learning method of the present application further includes:
[0119] Step C10: predicting, by a contrastive learning module of the contrastive learning model based on the specific information, an alignment probability of a cross-modal sample pair in the target data, wherein the cross-modal sample pair refers to unimodal data of two different modalities;
[0120] It should be noted that the contrastive learning model also includes a contrastive learning module for predicting the alignment probability of each cross-modal sample pair. A cross-modal sample pair refers to unimodal data of two different modalities. For example, when the target data includes two multimodal data, one multimodal data includes a first molecular structure, a first molecular cell image, and a first molecular gene expression, and the other multimodal data includes a second molecular structure, a second molecular cell image, and a second molecular gene expression, then the first molecular structure and the first molecular cell image, the first molecular structure and the second molecular cell image, the first molecular structure and the first molecular gene expression, and the first molecular structure and the second molecular gene expression are cross-modal sample pairs. Similarly, for the second molecular structure, there are also four pairs of cross-modal sample pairs.
[0121] In one feasible embodiment, in order to make the embeddings of matching molecule-phenotype pairs closer to each other and push the embeddings of mismatched molecule-phenotype pairs apart, traditional contrastive learning methods aim to maximize the lower limit of the mutual information (MI) between the molecular graph and the cell image in the positive sample pair. Among them, the molecule-phenotype pair refers to a cross-modal sample pair consisting of a molecular structure and a molecular cell image or a molecular gene expression. However, when multimodal samples are weakly aligned, these methods have difficulty dealing with cross-modal noise. In this case, multiple molecule-phenotype pairs with similar content may be mislabeled as negative sample pairs, resulting in a decrease in the performance of the model in distinguishing true negative samples from false negative samples. To address these limitations, the embodiment of the present application introduces an adaptive contrastive learning module with a negative sampling calibration function, which can adjust the weights of positive and negative sample pairs, thereby improving the robustness and accuracy of the contrastive learning framework. Specifically, a multi-layer perceptron (MLP) in the contrastive learning module is used to predict the alignment probability of each cross-modal sample pair, γ ijrepresents the alignment probability of the i-th sample and the j-th sample. i ,I j ) as an example, the contrast loss function after alignment probability correction can be expressed as:
[0122]
[0123] Step C20 : determining the contrast loss of the cross-modal sample pair based on the alignment probability, the valid information, and the specific information.
[0124] After obtaining the alignment probability of each cross-modal sample pair, the contrast loss of the cross-modal sample pair is determined based on the alignment probability, the effective information, and the specific information. It can be understood that the feature vector of each modality includes corresponding effective information and specific information.
[0125] For example, taking the alignment probability between molecular structures and molecular cell images as an example, the contrastive learning between molecular graphs and molecular cell images can be expressed as:
[0126]
[0127] Here, (1-γ ij ) emphasizes the importance of correctly labeling negative sample pairs, and γ ij The impact of potential mislabeling is taken into account. represents the contrastive learning loss between molecular structures and molecular cell images, represents the contrastive learning loss between molecular structure and molecular gene expression. and Add them together to get the contrast loss of cross-modal sample pairs.
[0128] Thus, the embodiment of the present application integrates the predicted registration probability into the contrastive learning framework and utilizes and To enhance the robustness to cross-modal pairing and thus achieve better overall performance in multimodal joint learning.
[0129] Thus, the embodiments of the present application propose to calibrate negative sampling in molecular phenotype comparative learning, evaluate the probability of a sample pair representing a true positive / negative association based on the sample pair, and dynamically adjust the processing of the sample pair.
[0130] In this embodiment, the molecular multimodal representation learning method of the present application further includes:
[0131] Step D10, predicting the fourth eigenvector of the molecular cell image and the fifth eigenvector of the molecular gene expression based on the molecular structure in the target data by the generation module of the contrastive learning model;
[0132] It should be noted that the contrastive learning model also includes a generation module for predicting the feature vectors of multimodal data.
[0133] The generation module of the contrastive learning model predicts the eigenvector of the molecular cell image (hereinafter referred to as the fourth eigenvector for distinction) and the eigenvector of the molecular gene expression (hereinafter referred to as the fifth eigenvector for distinction) based on the molecular structure in the target data.
[0134] Step D20, when the molecular cell image and / or molecular gene expression in the target data is missing, using the fourth eigenvector and / or the fifth eigenvector as the target eigenvector of the target data;
[0135] Determine whether the molecular cell image and / or molecular gene expression in the target data is missing. If missing, use the fourth eigenvector and / or the fifth eigenvector as the target eigenvector of the target data. It is understood that when only the molecular cell image is missing, the target eigenvector includes the first eigenvector, the fourth eigenvector, and the third eigenvector; when only the molecular gene expression is missing, the target eigenvector includes the first eigenvector, the second eigenvector, and the fifth eigenvector. That is, when data is missing, the eigenvector of the unimodal data is replaced by the eigenvector predicted by the model.
[0136] For example, taking the missing molecular cell image as an example, the generation module f G2I It will learn and predict the missing modal data based on the molecular graph G The calculation formula is:
[0137]
[0138] Step D30 , when the molecular cell images and molecular gene expressions in the target data are not missing, determining the reconstruction loss based on the fourth eigenvector, the fifth eigenvector and the target eigenvector.
[0139] If it is detected that the molecular cell images and molecular gene expressions in the target data are not missing, the loss between the fourth eigenvector and the second eigenvector is calculated, and the loss between the fifth eigenvector and the third eigenvector is calculated. The sum of these two losses is called the reconstruction loss.
[0140] For example, taking molecular cell images as an example, the loss Defined as a generative modality With the true mode h I The differences between:
[0141]
[0142] Among them, MSE is the mean square error, which is used to measure the average square error between the predicted value and the true value.
[0143] In this embodiment, after generating the missing modal data, the next step is to perform multimodal representation fusion. To achieve this goal, the embodiment of the present application adopts a mixture of experts (MoE) model for modal fusion, which can effectively handle incomplete and noisy multimodal data. The mixture of experts model aims to combine inputs from different modalities based on the correlation between different modalities and tasks. Each modality is represented by a specific expert network Expert = {Expert G ,Expert I ,Expert E Each Expert is a single-layer MLP. The gate control mechanism R(·) determines the expert combination that should be activated for each input sample. The routing weight is calculated using the softmax function, as follows:
[0144]
[0145] in, Represents the weights of the feature vectors of different modal data. Formally, the output of the MOE layer can be expressed as:
[0146]
[0147] Among them, R k Expert k The routing weight of z is . z represents the fusion feature.
[0148] It should be noted that the routing weights of the MoE layer are dynamically adjusted based on the relevance of each modal data. While adding an MoE layer can significantly increase the capacity of the model, it also increases the computational cost. To optimize performance without occupying a large amount of computing resources, the embodiment of the present application balances the pros and cons by varying the number of experts.
[0149] In this way, the embodiment of the present application introduces a hybrid fusion strategy to generate and integrate multimodal representations, so that the model can handle incomplete molecular phenotypic data. This strategy ensures that the model can generate reliable and comprehensive molecular representations even when certain patterns are missing. A large number of experiments have shown that MINER significantly improves the accuracy and robustness of molecular representation learning, and outperforms existing methods in molecular multimodal representation learning and molecular phenotype detection tasks. It is worth noting that this application has been shown to recommend effective candidate drugs for specific diseases recorded in the literature (such as hitting the FDA-approved drug Donepezil in the Alzheimer's disease case), which highlights its potential in advancing drug discovery and development.
[0150] For example, Figure 2 The figure shows the architecture of the contrastive learning model, which mainly includes an encoder, a denoising module, a contrastive learning module and an expert module. The encoder is used to extract features from multimodal data to obtain feature vectors, which can effectively represent various modalities, including molecular structures, cell images and gene expression, thereby generating meaningful representations; the denoising module is used to distinguish between noise and similar information (i.e., effective information) in each modality; then the contrastive learning module is used to dynamically adjust the processing method of sample pairs based on the possibility that the cross-modal sample pairs are called positive sample pairs or negative sample pairs; finally, the expert module is used to generate and fuse multimodal representations to improve the performance of downstream tasks.
[0151] In this embodiment, step S30 may include:
[0152] Step S301, determining the cross entropy loss of the contrastive learning model based on the predicted label data and the true label data;
[0153] Substitute the predicted label data and the true label data into the preset loss function to obtain the loss between the two (hereinafter referred to as cross entropy loss for distinction). Specifically, the calculation formula of cross entropy loss is:
[0154]
[0155] in, Represents the true label y i and the predicted output The cross entropy loss between , N represents the number of samples, and σ R represents the standard deviation of the routing weight of each expert in MoE.
[0156] Step S302, determining a property prediction loss of the contrastive learning model based on one or more of the denoising loss, the contrast loss, the reconstruction loss, and the cross entropy loss;
[0157] In one feasible implementation, the denoising loss, contrast loss, reconstruction loss, and cross entropy loss are added together to obtain the final property prediction loss. Specifically, the final loss function is:
[0158]
[0159] Among them, λ1, λ2, λ3, and λ4 are weights. is the denoising loss used to improve the quality of unimodal representation. is a contrastive loss for aligning multimodal representations with negative sampling calibration. is the reconstruction loss used to generate the missing modality. is the cross entropy loss for accurate prediction of balanced supervised tasks.
[0160] Step S303 : Optimizing the model parameters of the contrastive learning model based on the property prediction loss to obtain a new contrastive learning model.
[0161] After obtaining the property prediction loss, the model parameters of the contrastive learning model are reversely optimized based on the property prediction loss to obtain a new contrastive learning model.
[0162] For example, Figure 3 The figure shows a schematic diagram of the molecular multimodal representation learning process. First, sample data (i.e., multimodal data) and real label data are obtained; the sample data are input into the contrastive learning model in batches to obtain predicted label data; it is determined whether the contrastive learning model meets the training termination conditions; if so, the contrastive learning model at this time is used as the final trained model (hereinafter referred to as the target model for distinction); if not, the property prediction loss is calculated, and the contrastive learning model is optimized based on the property prediction loss to obtain a new contrastive learning model, and the training steps are returned to execute until the contrastive learning model meets the training termination conditions to obtain the target model for molecular multimodal representation learning.
[0163] In this way, the embodiments of the present application learn robust and reliable molecular representations from weakly aligned and incomplete molecular phenotypic data. Specifically, by introducing a dynamic sample processing mechanism, the probability of sample data being a true positive sample or a negative sample is evaluated in real time, adaptive adjustment of sample weights is achieved, and the weight distribution of positive and negative samples is optimized, thereby significantly improving the robustness and accuracy of the comparative learning model.
[0164] In one feasible embodiment, in terms of molecular multimodal representation learning, this application uses two datasets, ChEMBL2K and Broad6k. Specifically, Broad6K contains 6567 molecules, with gene expression and cell morphology loss rates of 50.34% and 1.10%, respectively, while CHEMBL2K contains 2355 molecules, with gene expression and cell morphology loss rates of 75.33% and 0.04%, respectively. This application performed a scaffold split on the two datasets, using a ratio of 0.6:0.15:0.25 for the training set, validation set, and test set. In terms of performance evaluation, the area under the receiver operating characteristic curve (AUROC) is used to measure classification performance. In addition, this application also reports the percentage of tasks with average AUROC thresholds greater than 0.8, 0.85, and 0.9. To ensure consistency, all results are reported in the form of averages and standard deviations of five independent experimental runs. In order to demonstrate the effectiveness of the method of this application, this application conducted experiments on two molecular property datasets, ChEMBL2K and Broad6K. As shown in the table, this application achieves state-of-the-art results in learning multimodal molecular representations, surpassing existing models in most cases. Specifically, compared to the state-of-the-art baseline model InfoCORE, this application achieves substantial improvements, with an average AUROC improvement of 3.657% and an average AUROC improvement of 1.60%. On the ChEMBL2K dataset, AUROC>80% increased by 3.657% and AUROC>80% increased by 1.602%. These gains can be attributed to the integration of multimodal data by our model, particularly through contrastive learning with calibrated negative sampling. By dynamically adjusting the processing of sample pairs, our method generates more powerful and reliable multimodal embeddings.
[0165] Furthermore, these improvements are partly due to our hybrid fusion strategy, which effectively handles incomplete data by integrating available modalities. This is particularly evident for the ChEMBL2K dataset, where the missing gene expression data rate is as high as 75.33%.
[0166] The present application also provides a molecular multimodal characterization learning device, please refer to Figure 4 , the molecular multimodal representation learning device comprises:
[0167] An acquisition module 10 is configured to acquire multimodal data of each sample molecule and true label data of each sample molecule, wherein the true label data represents the true properties of the sample molecule;
[0168] A prediction module 20 is configured to batch-input each of the multimodal data into a contrastive learning model to obtain predicted label data of target data, wherein the target data is the multimodal data currently input into the contrastive learning model, and the predicted label data represents a probability value of the multimodal data being a positive sample or a negative sample;
[0169] The training termination module 30 is used to optimize the comparative learning model based on the predicted label data and the real label data to obtain a new comparative learning model when the comparative learning model does not meet the preset training termination condition, and return to execute the step of batch inputting each multimodal data into the comparative learning model until the comparative learning model meets the training termination condition, and perform molecular property prediction based on the comparative learning model.
[0170] Optionally, the prediction module 20 is further configured to:
[0171] Inputting each of the multimodal data into a contrastive learning model in batches;
[0172] Extracting features from the target data using the encoder of the contrastive learning model to obtain a target feature vector of the target data;
[0173] De-noising the target feature vector using a denoising module of the contrastive learning model to obtain valid information in the target feature vector;
[0174] The effective information is fused through the expert module of the contrastive learning model to obtain a fusion feature;
[0175] The prediction module of the contrastive learning model performs molecular multimodal representation learning on the fusion features to obtain predicted label data of the target data.
[0176] Optionally, the target data includes at least one of the multimodal data, each of the multimodal data includes a molecular structure, a molecular cell image, and a molecular gene expression of a sample molecule, and the encoder includes a graph neural network, a first fully connected neural network, and a second fully connected neural network;
[0177] The prediction module 20 is further configured to:
[0178] Extracting features from the molecular structure using a graph neural network of the contrastive learning model to obtain a first feature vector, wherein the first feature vector represents a characteristic of the molecular structure;
[0179] Performing feature extraction on the molecular cell image using the first fully connected neural network of the contrastive learning model to obtain a second feature vector, wherein the second feature vector represents a cell morphological feature of the molecular cell image;
[0180] performing feature extraction on the molecular gene expression using a second fully connected neural network of the contrastive learning model to obtain a third feature vector, wherein the third feature vector represents a genetic feature of the molecular gene expression;
[0181] The first eigenvector, the second eigenvector, and the third eigenvector are used as target eigenvectors of the target data.
[0182] Optionally, the prediction module 20 is further configured to:
[0183] Performing information separation on the target feature vector using the denoising module of the contrastive learning model to obtain effective information and specific information of the target feature vector;
[0184] A Pearson correlation coefficient between the effective information and the specific information is determined, and the Pearson correlation coefficient is used as a denoising loss.
[0185] Optionally, the molecular multimodal representation learning device further includes a contrast learning module, which is configured to:
[0186] Predicting, by a contrastive learning module of the contrastive learning model, alignment probabilities of cross-modal sample pairs in the target data based on the specific information, wherein the cross-modal sample pairs refer to unimodal data of two different modalities;
[0187] A contrastive loss of the cross-modal sample pair is determined based on the alignment probability, the valid information, and the specific information.
[0188] Optionally, the molecular multimodal representation learning device further includes a generation module, wherein the generation module is configured to:
[0189] Predicting, by the generation module of the contrastive learning model, a fourth eigenvector of the molecular cell image in the target data and a fifth eigenvector of the molecular gene expression in the target data based on the molecular structure in the target data;
[0190] In the case where the molecular cell image and / or molecular gene expression is missing in the target data, using the fourth eigenvector and / or the fifth eigenvector as the target eigenvector of the target data;
[0191] In a case where molecular cell images and molecular gene expressions are not missing in the target data, a reconstruction loss is determined based on the fourth eigenvector, the fifth eigenvector, and the target eigenvector.
[0192] Optionally, the training termination module 30 is further configured to:
[0193] Determining a cross entropy loss of the contrastive learning model based on the predicted label data and the true label data;
[0194] Determining a property prediction loss of the contrastive learning model based on one or more of the denoising loss, the contrast loss, the reconstruction loss, and the cross entropy loss;
[0195] The model parameters of the contrastive learning model are optimized based on the property prediction loss to obtain a new contrastive learning model.
[0196] The molecular multimodal representation learning device provided in the embodiments of this application, employing the molecular multimodal representation learning method of the aforementioned embodiments, can address the technical problem of improving the accuracy of comparative learning models used for molecular property prediction. Compared to the prior art, the beneficial effects of the molecular multimodal representation learning device provided in the embodiments of this application are the same as those of the molecular multimodal representation learning method provided in the aforementioned embodiments. Other technical features of the molecular multimodal representation learning device are the same as those disclosed in the aforementioned embodiments and are not further elaborated here.
[0197] The present application provides a molecular multimodal representation learning device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the molecular multimodal representation learning method in the above-mentioned embodiment one.
[0198] Reference below Figure 5 , which shows a schematic diagram of the structure of a molecular multimodal representation learning device suitable for implementing the embodiments of the present application. The molecular multimodal representation learning device in the embodiments of the present application may include, but is not limited to, mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), and fixed terminals such as digital TVs and desktop computers. Figure 5 The molecular multimodal representation learning device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0199] like Figure 5As shown, the molecular multimodal representation learning device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory 1002 or programs loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the molecular multimodal representation learning device. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, hard disk, etc.; and a communication device 1009. Communication device 1009 can allow the molecular multimodal characterization learning device to communicate wirelessly or wired with other devices to exchange data. Although the figure shows a molecular multimodal characterization learning device with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented or have alternatively.
[0200] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.
[0201] The molecular multimodal representation learning device provided in this application, employing the molecular multimodal representation learning method of the aforementioned embodiment, can address the technical problem of improving the accuracy of comparative learning models used for molecular property prediction. Compared to the prior art, the beneficial effects of the molecular multimodal representation learning device provided in this application are the same as those of the molecular multimodal representation learning method provided in the aforementioned embodiment. Other technical features of the molecular multimodal representation learning device are the same as those disclosed in the aforementioned embodiment and are not further elaborated here.
[0202] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0203] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0204] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, and the computer-readable program instructions are used to execute the molecular multimodal representation learning method in the above-mentioned embodiment.
[0205] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0206] The computer-readable storage medium may be included in the molecular multimodal representation learning device; or it may exist independently without being assembled into the molecular multimodal representation learning device.
[0207] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the molecular multimodal characterization learning device, the molecular multimodal characterization learning device: obtains the multimodal data of each sample molecule and the true label data of each sample molecule, wherein the true label data represents the true properties of the sample molecule; inputs each of the multimodal data into the contrastive learning model in batches to obtain predicted label data of the target data, wherein the target data is the multimodal data input into the contrastive learning model at that time, and the predicted label data represents the probability value of the multimodal data being a positive sample or a negative sample; when the contrastive learning model does not meet the preset training termination condition, the contrastive learning model is optimized based on the predicted label data and the true label data to obtain a new contrastive learning model, and returns to execute the step of inputting each of the multimodal data into the contrastive learning model in batches until the contrastive learning model meets the training termination condition, and molecular property prediction is performed based on the contrastive learning model.
[0208] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0209] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0210] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0211] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., a computer program) for executing the above-mentioned molecular multimodal representation learning method, and can solve the technical problem of how to improve the accuracy of the comparative learning model for molecular property prediction. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the molecular multimodal representation learning method provided in the above-mentioned embodiment, and will not be repeated here.
[0212] An embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the molecular multimodal representation learning method as described above.
[0213] The computer program product provided in this application can improve the accuracy of comparative learning models for molecular property prediction. Compared with the prior art, the beneficial effects of the computer program product provided in the embodiments of this application are the same as those of the molecular multimodal representation learning method provided in the above embodiments, and will not be repeated here.
[0214] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent processing scope of the present application.
Claims
1. A molecular multimodal representation learning method, characterized in that: The molecular multimodal representation learning method includes: Acquire multimodal data of each sample molecule and true label data of each sample molecule, wherein the true label data represents the true properties of the sample molecule; Inputting each of the multimodal data into a contrastive learning model in batches to obtain predicted label data of target data, wherein the target data is the multimodal data input into the contrastive learning model at that time, and the predicted label data represents a probability value of the multimodal data being a positive sample or a negative sample; When the contrastive learning model does not meet the preset training termination condition, the contrastive learning model is optimized based on the predicted label data and the real label data to obtain a new contrastive learning model, and the step of batch inputting each multimodal data into the contrastive learning model is returned to execute until the contrastive learning model meets the training termination condition, and molecular property prediction is performed based on the contrastive learning model.
2. The molecular multimodal representation learning method according to claim 1, wherein: The step of batch-inputting each of the multimodal data into the contrastive learning model to obtain predicted label data of the target data includes: Inputting each of the multimodal data into a contrastive learning model in batches; Extracting features from the target data using the encoder of the contrastive learning model to obtain a target feature vector of the target data; De-noising the target feature vector using a denoising module of the contrastive learning model to obtain valid information in the target feature vector; fusing the effective information through the expert module of the contrastive learning model to obtain a fusion feature; The prediction module of the contrastive learning model performs molecular multimodal representation learning on the fusion features to obtain predicted label data of the target data.
3. The molecular multimodal representation learning method according to claim 2, wherein: The target data includes at least one of the multimodal data, each of the multimodal data includes a molecular structure, a molecular cell image, and a molecular gene expression of a sample molecule, and the encoder includes a graph neural network, a first fully connected neural network, and a second fully connected neural network; The step of extracting features from the target data using the encoder of the contrastive learning model to obtain a target feature vector of the target data includes: Extracting features from the molecular structure using a graph neural network of the contrastive learning model to obtain a first feature vector, wherein the first feature vector represents a characteristic of the molecular structure; Performing feature extraction on the molecular cell image using the first fully connected neural network of the contrastive learning model to obtain a second feature vector, wherein the second feature vector represents a cell morphological feature of the molecular cell image; performing feature extraction on the molecular gene expression using a second fully connected neural network of the contrastive learning model to obtain a third feature vector, wherein the third feature vector represents a genetic feature of the molecular gene expression; The first eigenvector, the second eigenvector, and the third eigenvector are used as target eigenvectors of the target data.
4. The molecular multimodal representation learning method according to claim 2, wherein: The step of denoising the target feature vector by the denoising module of the contrastive learning model to obtain valid information in the target feature vector includes: Performing information separation on the target feature vector using the denoising module of the contrastive learning model to obtain effective information and specific information of the target feature vector; A Pearson correlation coefficient between the effective information and the specific information is determined, and the Pearson correlation coefficient is used as a denoising loss.
5. The molecular multimodal representation learning method according to claim 4, wherein: After the step of denoising the target feature vector using the denoising module of the contrastive learning model to obtain valid information in the target feature vector, the method further includes: Predicting, by a contrastive learning module of the contrastive learning model, alignment probabilities of cross-modal sample pairs in the target data based on the specific information, wherein the cross-modal sample pairs refer to unimodal data of two different modalities; A contrastive loss of the cross-modal sample pair is determined based on the alignment probability, the valid information, and the specific information.
6. The molecular multimodal representation learning method according to claim 5, wherein: The method further comprises: Predicting, by the generation module of the contrastive learning model, a fourth eigenvector of the molecular cell image in the target data and a fifth eigenvector of the molecular gene expression in the target data based on the molecular structure in the target data; In the case where the molecular cell image and / or molecular gene expression is missing in the target data, using the fourth eigenvector and / or the fifth eigenvector as the target eigenvector of the target data; In a case where molecular cell images and molecular gene expressions are not missing in the target data, a reconstruction loss is determined based on the fourth eigenvector, the fifth eigenvector, and the target eigenvector.
7. The molecular multimodal representation learning method according to claim 6, wherein: The step of optimizing the contrastive learning model based on the predicted label data and the real label data to obtain a new contrastive learning model includes: Determining a cross entropy loss of the contrastive learning model based on the predicted label data and the true label data; Determining a property prediction loss of the contrastive learning model based on one or more of the denoising loss, the contrast loss, the reconstruction loss, and the cross entropy loss; The model parameters of the contrastive learning model are optimized based on the property prediction loss to obtain a new contrastive learning model.
8. A molecular multimodal representation learning device, characterized in that: The molecular multimodal representation learning device includes: an acquisition module, configured to acquire multimodal data of each sample molecule and true label data of each sample molecule, wherein the true label data represents the true properties of the sample molecule; A prediction module, configured to batch-input each of the multimodal data into the contrastive learning model to obtain predicted label data of target data, wherein the target data is the multimodal data currently input into the contrastive learning model, and the predicted label data represents a probability value of the multimodal data being a positive sample or a negative sample; A training termination module is used to optimize the contrastive learning model based on the predicted label data and the real label data to obtain a new contrastive learning model when the contrastive learning model does not meet the preset training termination conditions, and return to execute the step of batch inputting each multimodal data into the contrastive learning model until the contrastive learning model meets the training termination conditions, and perform molecular property prediction based on the contrastive learning model.
9. A molecular multimodal representation learning device, characterized in that: The molecular multimodal representation learning device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the molecular multimodal representation learning method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the molecular multimodal representation learning method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Submarine optical cable disturbance detection method and device and computer readable storage medium
CN121479530A