Multi-modal drug molecule prediction method based on deletion modal generation

By constructing the multimodal molecular pre-training model MolBT, the shallow feature representation of missing text modals is generated using the existing modality, which solves the problem of insufficient performance of the multimodal molecular pre-training model in downstream tasks, and achieves stronger generalization capabilities and cross-modal interaction effects.

CN120015169AActive Publication Date: 2025-05-16NORTHEAST FORESTRY UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510077682.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-16
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

The existing multimodal molecular pre-trained models have insufficient performance on downstream tasks, mainly because of the small scale of the molecular description text data set, it is difficult to obtain high-quality molecular description text data, and the model is difficult to fully capture the complex relationship between molecular graph structure and text.

Method used

A multimodal drug molecule prediction method based on missing modal generation is proposed. By constructing a multimodal molecular pre-training model MolBT, the existing modality is used to generate shallow feature representations of missing text modalities, the scale of the multimodal molecular pre-training data set is expanded, and the utilization of multiple singlemodal features is optimized through BridgeLayer, which promotes cross-modal interaction and fusion.

Benefits of technology

The scale of multimodal molecular pre-training data sets has been significantly expanded, and the model's generalization ability in downstream tasks has been improved, especially in molecular-text cross-modal retrieval and molecular property prediction tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015169A_ABST
    Figure CN120015169A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal drug molecule prediction method based on deletion modal generation, belongs to the field of artificial intelligence assisted drug research and development, and particularly relates to a multi-modal drug molecule prediction method. The invention aims to solve the problems that the scale of a multi-modal molecule-text pre-training data set is limited by molecular description text modal deficiency, a model is difficult to fully capture a complex relationship between a molecular graph structure and a text, and the existing multi-modal molecule pre-training model has certain insufficiency in the expression of a downstream task. The method specifically comprises the following steps: 1, constructing a multi-modal molecule pre-training model; 2, pre-training the multi-modal molecule pre-training model to obtain a pre-trained multi-modal molecule pre-training model; 3, based on a downstream task type, performing fine tuning on the pre-trained multi-modal molecule pre-training model to obtain a finely-tuned multi-modal molecule pre-training model; and 4, predicting a downstream task based on the multi-modal molecule pre-training model after fine tuning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence-assisted drug research and development, and specifically relates to a multimodal drug molecule prediction method. Background Art

[0002] Understanding the properties and functions of molecules is crucial for drug development, providing a foundation for the development of targeted therapies, the elucidation of disease mechanisms, and the development of personalized medicine. Traditional drug discovery relies on laboratory experiments, which require a lot of time and economic costs. In recent years, artificial intelligence has played an increasingly important role in assisting drug development, with significant improvements in efficiency and cost-effectiveness. Researchers have applied self-supervised learning strategies to molecular representation learning, training models to understand various molecular representations, such as molecular SMILES strings and molecular graph structures. These models are designed to learn pre-trained models from a large amount of unlabeled data to support various molecular tasks, such as molecular property prediction and molecular generation. According to the different input modalities, current molecular pre-training models can be divided into two categories, namely single-modal pre-training models and multi-modal pre-training models. Compared with single-modal molecular pre-training models, multi-modal molecular pre-training models can fuse information from multiple different modalities, have stronger generalization capabilities, and are suitable for a wider range of downstream tasks. In reality, humans have the ability to understand and learn knowledge from multiple perspectives, and can recognize molecules in different forms by combining molecular graph structures, molecular SMILES strings, and molecular description texts. Similarly, existing research works have explored multimodal molecule-text pre-training models to learn molecular representations in combination with textual knowledge.

[0003] However, the performance of existing multimodal molecular pre-training models on downstream tasks is still insufficient. First, compared with unimodal molecular pre-training datasets, the scale of multimodal molecular-text pre-training datasets is very small. The main challenge is that it is difficult to obtain high-quality molecular description text data. The reason is that the molecular annotation process requires expertise in molecular chemistry, which makes large-scale manual annotation both expensive and cumbersome. Therefore, there is a significant gap in data size between molecular description text data and molecular SMILES strings and molecular graph structures. Some researchers have used large language models (LLMs) to generate pseudo-text data to address the challenge of missing molecular description text data. Although this method has achieved certain results, the generation efficiency of pseudo-text data is low and it requires a certain amount of time and economic cost. Another noteworthy issue is that the existing multimodal molecular graph structure-text pre-training models either ignore and waste the different levels of semantic knowledge contained in different layers of the unimodal encoder, or it is difficult to achieve effective cross-modal alignment. Summary of the invention

[0004] The purpose of the present invention is to solve the problems that the scale of multimodal molecule-text pre-training dataset is limited by the missing text modality of molecular description, the model is difficult to fully capture the complex relationship between molecular graph structure and text, and the performance of existing multimodal molecule pre-training models on downstream tasks has certain deficiencies, and proposes a multimodal drug molecule prediction method based on missing modality generation.

[0005] The specific process of the multimodal drug molecule prediction method based on missing modality generation is as follows:

[0006] Step 1: Construct a multimodal molecular pre-training model MolBT;

[0007] Step 2: pre-train the multimodal molecular pre-training model MolBT to obtain a pre-trained multimodal molecular pre-training model MolBT;

[0008] Step 3: Based on the downstream task type, fine-tune the pre-trained multimodal molecular pre-training model MolBT to obtain the fine-tuned multimodal molecular pre-training model MolBT;

[0009] Step 4: Predict downstream tasks based on the fine-tuned multimodal molecular pre-training model MolBT.

[0010] The beneficial effects of the present invention are:

[0011] The present invention proposes a multimodal molecule pre-training model based on missing modality generation, MolBT. MolBT explores a method to solve the problem of missing text modality of molecular description, that is, to generate shallow feature representation of missing text modality using existing modalities. This method significantly expands the scale of multimodal molecule pre-training datasets. At the same time, we emphasize the use of different levels of semantic knowledge contained in different layers of unimodal encoders. Through BridgeLayer, we establish connections between the top layers of the unimodal encoder and each layer of the cross-modal encoder, thereby optimizing the use of multiple unimodal features and promoting effective bottom-up cross-modal interaction and fusion between the three modalities of molecular graph structure representation, text representation and molecular SMILES string. MolBT has demonstrated excellent generalization ability in a series of downstream tasks. MolBT has shown excellent performance in molecule-text cross-modal retrieval. In the molecular property prediction task, MolBT performed well, highlighting its outstanding performance in physiology and biophysics related classification tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 This is the architecture diagram of the multimodal drug molecule pre-training model based on missing modality generation. DETAILED DESCRIPTION

[0013] Specific implementation method 1: The specific process of the multimodal drug molecule prediction method based on missing modality generation in this implementation method is as follows:

[0014] The present invention constructs a new multimodal molecule-text pre-training model, MolBT. MolBT consists of a molecular graph structure encoder, a sequence encoder and a cross-modal encoder. Specifically, we use Graphormer as the molecular graph structure encoder. Since both the molecular description text and the molecular SMILES string are sequence type data, we use BERT as the molecular sequence encoder to process these two modal data. We have established a cross-modal encoder as a bridge connecting the molecular graph structure encoder and the molecular sequence encoder, so that effective bottom-up cross-modal interaction and fusion are achieved between the molecular graph structure feature representation, the text feature representation and the molecular SMILES feature representation. In addition, a missing modality generation module based on the normalized flow model is introduced in MolBT. The method uses existing modalities (i.e., molecular graph structures and molecular SMILES strings) to generate shallow feature representations of missing text modalities. In the pre-training process, the feature representation of the generated missing text modality is used to greatly expand the scale of the multimodal molecule-text pre-training data set.

[0015] Step 1: Construct a multimodal molecular pre-training model MolBT;

[0016] Step 2: pre-train the multimodal molecular pre-training model MolBT to obtain a pre-trained multimodal molecular pre-training model MolBT;

[0017] Step 3: Based on the downstream task type, fine-tune the pre-trained multimodal molecular pre-training model MolBT to obtain the fine-tuned multimodal molecular pre-training model MolBT;

[0018] Step 4: Predict downstream tasks based on the fine-tuned multimodal molecular pre-training model MolBT.

[0019] Specific implementation method 2: This implementation method is different from the specific implementation method 1 in that the multimodal molecular pre-training model MolBT is constructed in step 1: the specific process is as follows:

[0020] Step 11: input the molecular graph into a molecular graph structure encoder, and the molecular graph structure encoder outputs a molecular graph structure feature representation;

[0021] The molecular graph structure encoder is a Graphormer;

[0022] Step 1 and 2: input the molecular description text into the molecular sequence encoder, and the molecular sequence encoder outputs the molecular description text feature representation;

[0023] The molecular sequence encoder is BERT;

[0024] Step 13: Input the molecular SMILES string into the molecular sequence encoder, and the molecular sequence encoder outputs the molecular SMILES feature representation;

[0025] The molecular sequence encoder is BERT;

[0026] Step 14: Establish a cross-modal encoder, input the molecular graph structure feature representation, the molecular description text feature representation, and the molecular SMILES feature representation into the cross-modal encoder, and the cross-modal encoder outputs the molecular graph structure cross-modal feature representation and the molecular description text cross-modal feature representation;

[0027] Connect the molecular graph structure encoder and the molecular sequence encoder.

[0028] Step 15: Construct a missing modality generation module, input the molecular graph structure feature representation and molecular SMILES feature representation of the missing molecular description text into the missing modality generation module, and the missing modality generation module outputs the corresponding shallow feature representation of the molecular description text.

[0029] The other steps and parameters are the same as those in the first embodiment.

[0030] Specific implementation method three: This implementation method is different from specific implementation method one or two in that in step one, the molecular graph is input into the molecular graph structure encoder, and the molecular graph structure encoder outputs the molecular graph structure feature representation; the specific process is:

[0031] Step 1: Input the molecular graph into the graph feature encoding layer in the molecular graph structure encoder, and the graph feature encoding layer in the molecular graph structure encoder outputs the molecular graph node feature G0.

[0032] in, represents the first atomic feature embedding, represents the second atomic feature embedding, represents the nth atomic feature embedding, R represents a real number, n g Indicates the number of atoms in the molecular graph, d g Represents the feature dimension of the molecular graph structure encoder;

[0033] Step 112: input the molecular graph node feature G0 into the first Graphormer layer, the second Graphormer layer, the third Graphormer layer, the fourth Graphormer layer, the fifth Graphormer layer, the sixth Graphormer layer, the seventh Graphormer layer, the eighth Graphormer layer, the ninth Graphormer layer, the tenth Graphormer layer, the eleventh Graphormer layer, and the twelfth Graphormer layer of the molecular graph structure encoder in sequence, and the twelfth Graphormer layer outputs the molecular graph structure feature representation G 12 ;

[0034] The formula is as follows:

[0035]

[0036] In the formula, K g Indicates the total number of Graphormer layers in the molecular graph structure encoder, K g =12;

[0037] represents the i-th Graphormer layer;

[0038] G i-1 Represents the molecular graph structure features output by the i-1th Graphormer layer;

[0039] G i Represents the molecular graph structural features output by the i-th Graphormer layer.

[0040] The other steps and parameters are the same as those in the first or second embodiment.

[0041] Specific implementation method 4: This implementation method is different from one of the specific implementation methods 1 to 3 in that in step 1 and 2, the molecular description text is input into the molecular sequence encoder, and the molecular sequence encoder outputs the molecular description text feature representation; the specific process is:

[0042] Step 121: Input the molecular description text into the BERT word embedding layer of the molecular sequence encoder, and the BERT word embedding layer of the molecular sequence encoder outputs the molecular description text feature T0;

[0043] Step 122: Input the molecular description text feature T0 into the 1st Transformer layer, 2nd Transformer layer, 3rd Transformer layer, 4th Transformer layer, 5th Transformer layer, 6th Transformer layer, 7th Transformer layer, 8th Transformer layer, 9th Transformer layer, 10th Transformer layer, 11th Transformer layer, and 12th Transformer layer of the molecular sequence encoder in sequence, and the 12th Transformer layer outputs the molecular description text feature representation T 12 ;

[0044] The formula is as follows:

[0045]

[0046] In the formula, K s represents the total number of Transformer layers in the molecular sequence encoder, K s =12;

[0047] represents the i-th Transformer layer;

[0048] T i-1 Represents the molecular description text features output by the i-1th Transformer layer;

[0049] T i Represents the molecular description text features output by the i-th Transformer layer.

[0050] The other steps and parameters are the same as those in Specific Embodiments 1 to 3.

[0051] Specific implementation mode 5: This implementation mode is different from the specific implementation modes 1 to 4 in that in step 13, the molecular SMILES string is input into the molecular sequence encoder, and the molecular sequence encoder outputs the molecular SMILES feature representation; the specific process is:

[0052] Step 131. Input the molecule SMILES string into the BERT word embedding layer of the molecule sequence encoder, and the BERT word embedding layer of the molecule sequence encoder outputs the molecule SMILES feature S0;

[0053] Step 132: Input the molecular SMILES feature S0 into the 1st Transformer layer, the 2nd Transformer layer, the 3rd Transformer layer, the 4th Transformer layer, the 5th Transformer layer, the 6th Transformer layer, the 7th Transformer layer, the 8th Transformer layer, the 9th Transformer layer, the 10th Transformer layer, the 11th Transformer layer, and the 12th Transformer layer of the molecular sequence encoder in sequence, and the 12th Transformer layer outputs the molecular SMILES feature representation S 12 ;

[0054] The formula is as follows:

[0055]

[0056] In the formula,

[0057] K s represents the total number of Transformer layers in the molecular sequence encoder, K s =12;

[0058] represents the i-th Transformer layer;

[0059] S i-1 Represents the molecular SMILES features output by the i-1th Transformer layer;

[0060] S i Represents the molecular SMILES features output by the i-th Transformer layer.

[0061] Since both the molecular SMILES string and the molecular description text are sequence type data, the molecular SMILES string and the molecular description text can share the same encoder.

[0062] The other steps and parameters are the same as those in Specific Embodiments 1 to 4.

[0063] Specific implementation method six: This implementation method is different from specific implementation methods one to five in that a cross-modal encoder is established in step one four, and the molecular graph structure feature representation, the molecular description text feature representation, and the molecular SMILES feature representation are input into the cross-modal encoder, and the cross-modal encoder outputs the molecular graph structure cross-modal feature representation and the molecular description text cross-modal feature representation; the specific process is:

[0064] The output of the 7th layer of the molecular graph encoder G7, the output of the 7th layer of the molecular sequence encoder T7 and S7 are input into the cross-modal encoder, and the output of the cross-modal encoder C G ,C T ;

[0065] The output of the 8th layer of the molecular graph structure encoder G8, the output of the 8th layer of the molecular sequence encoder T8 and S8 are input into the cross-modal encoder, and the output of the cross-modal encoder C G ,C T ;

[0066] The output of the 9th layer of the molecular graph structure encoder G9, the output of the 9th layer of the molecular sequence encoder T9 and S9 are input into the cross-modal encoder, and the output of the cross-modal encoder C G ,C T ;

[0067] The output of the molecular graph structure encoder layer 10 is G 10 , T output from the 10th layer of the molecular sequence encoder 10 and S 10 Input cross-modal encoder, cross-modal encoder output C G ,C T ;

[0068] The output of the molecular graph structure encoder layer 11 is G 11 , T output from the 11th layer of the molecular sequence encoder 11 and S 11 Input cross-modal encoder, cross-modal encoder output C G ,C T ;

[0069] The output of the molecular graph structure encoder layer 12 is G 12 , T output from the 12th layer of the molecular sequence encoder 12 and S 12 Input cross-modal encoder, cross-modal encoder output C G ,C T ;

[0070] The cross-modal encoder has 6 layers, each of which includes a self-attention layer, a common attention layer, and a feed-forward layer FFN in sequence;

[0071] The operation process of each layer in the cross-modal encoder is:

[0072]

[0073] In the formula, CME l represents the lth layer of the cross-modal encoder;

[0074] Each layer of the cross-modal encoder includes a self-attention layer, a common attention layer, and a feed-forward layer FFN in sequence;

[0075] and They represent the molecular graph structure cross-modal feature representation and the molecular text cross-modal feature representation output by the l-th layer of the cross-modal encoder respectively;

[0076] L C is the number of layers of the cross-modal encoder, L C =6;

[0077] Represents the molecular graph structure feature representation of the output of the l-1th layer of the cross-modal encoder;

[0078] Represents the text feature representation of the output of the l-1 layer of the cross-modal encoder;

[0079] Represents the molecular SMILES feature representation of the l-th layer input of the cross-modal encoder;

[0080] A cross-modal encoder with a common attention mechanism based on the Transformer architecture;

[0081] The specific process is:

[0082] 1) Represent the structural features of the molecular graph After BridgeLayer processing, the input of the lth layer of the cross-modal encoder is obtained

[0083] Representing text features After BridgeLayer processing, the input of the lth layer of the cross-modal encoder is obtained

[0084] It is expressed as:

[0085]

[0086] in

[0087]

[0088] In the formula,

[0089] W G W represents the molecular graph structure feature projection matrix (MLP); T Represents the text feature projection matrix (MLP);

[0090] G type Represents the molecular graph structure feature representation G kAfter the type embedding layer, the modal type embedding is obtained; T type Represents the molecular description text feature representation T k After the type embedding layer, the modal type embedding is obtained;

[0091] G k and T k Respectively represent the molecular graph structure feature representation and molecular description text feature representation output by the kth layer of the molecular graph structure encoder and the molecular sequence encoder, k = 7, 8, ..., 12;

[0092] Representation layer normalization LN;

[0093] Representation layer normalization LN;

[0094] 2) Feature representation based on molecular SMILES strings Extract class tags, denoted as

[0095]

[0096] In the formula, Represents a SMILES string feature representation of a molecule S k The first word feature embedding in ; Represents a SMILES string feature representation of a molecule S k The second word feature embedding in ; Represents a SMILES string feature representation of a molecule S k The nth s word-unit feature embedding; Represents a SMILES string feature representation of a molecule S k Class label; n s Indicates the length of the molecule SMILES string; d s Represents the feature dimension of the molecular sequence encoder;

[0097] 3) and The last node feature in is replaced by get and As shown below:

[0098]

[0099] In the formula, Represents the cross-modal feature representation of the molecular graph structure after replacing the features;

[0100] express The first atomic feature embedding in ; express The second atomic feature embedding in ; express The (n-1)th atomic feature embedding in ;

[0101] Represents the cross-modal feature representation of the molecular description text after replacing the features;

[0102] express The first word feature embedding in ; express The second word feature embedding in ; express The (n-1)th word feature embedding in ;

[0103] n g Indicates the number of atoms in the molecular graph; d g Represents the feature dimension of the molecular graph structure encoder; n t Indicates the length of the molecular description text; d t Represents the feature dimension of the text encoder;

[0104] 4) The molecular graph structure part and the molecular text part are input to the lth layer of the cross-modal encoder to obtain and

[0105]

[0106] In the formula, Represents the cross-modal feature representation of the molecular graph structure output of the l-th layer of the cross-modal encoder; Indicates

[0107] Cross-modal feature representation of processed molecular graph structure;

[0108] Represents the text cross-modal feature representation output by the l-th layer of the cross-modal encoder; Indicates Cross-modal feature representation of processed molecular description text;

[0109] Indicates the replaced Cross-modal feature representation of processed molecular graph structure; Indicates the replaced Cross-modal feature representation of processed molecular description text;

[0110] CME l represents the lth layer of the cross-modal encoder;

[0111] Each layer of the cross-modal encoder consists of a self-attention layer, a common attention layer, and a feed-forward layer FFN in sequence.

[0112] The other steps and parameters are the same as those in Specific Implementation Methods 1 to 5.

[0113] Specific implementation method 7: This implementation method is different from any one of specific implementation methods 1 to 6 in that in step 15, a missing modality generation module is constructed, the molecular graph and molecular SMILES string of the missing molecular description text are input into the missing modality generation module, and the missing modality generation module outputs the corresponding molecular description text; the specific process is:

[0114] To reduce the distribution gap between the recovered data and the real data, we construct a missing modality generation module to learn the latent distribution space of each modality, perform cross-modal distribution transfer to estimate the distribution of the missing text modality, and finally restore the shallow feature representation of the missing text modality through a projection network. To transform the cross-modal distribution, we introduce modality-dependent flows and bridge different modalities in the embedded distribution space. To facilitate the distribution transfer, a reversible modality-specific normalization flow is adopted for each modality, a classic generative probabilistic model that maps the features of different modalities into a latent space with a Gaussian distribution, reducing the distribution gap between modalities. Our main idea is to transfer the distribution from the observed modality to the missing modality through cross-flows and produce more confident predictions with higher distribution consistency.

[0115] Step 151, the molecular graph structure feature representation G6 obtained through the first 6 layers of the molecular graph structure encoder, the molecular description text feature representation T6 obtained through the first 6 layers of the molecular sequence encoder, and the molecular SMILES feature representation S6 are recorded as shallow feature representations, i.e., {G6, T6, S6};

[0116] Step 152: Project the shallow feature representations G6, T6, and S6 into the same dimensional space through a one-dimensional convolution layer to obtain Z (m) ∈R T×d ,m∈{t,g,s};

[0117] In the formula, Represents Z (t) , Z (g) or Z (s) ;

[0118] Z (t) The shallow feature represents the output feature of T6 after one one-dimensional convolution layer; Z (g)The shallow feature represents the output feature of G6 after one one-dimensional convolution layer; Z (s) The shallow feature represents the feature output by S6 after passing through one one-dimensional convolutional layer;

[0119] T represents the feature length after the one-dimensional convolution layer; d represents the feature dimension after the one-dimensional convolution layer;

[0120] Step 153: The subsequent missing mode restoration task aims to utilize Z (g) ,Z (s) The two available modalities generate the missing shallow feature representation of the molecular description text modality;

[0121] Assume represents the normalized flow model corresponding to mode m, represents its inverse transformation;

[0122] Z (g) and Z (s) Input into the corresponding normalized flow model respectively (Z (g) Input to the normalized flow model In the (s) Input to the normalized flow model ;), and obtain the latent feature representation with the same Gaussian distribution;

[0123] In the absence of molecular description text modality, the molecular graph structure feature representation G6 obtained from the first 6 layers of the molecular graph structure encoder and the molecular SMILES feature representation S6 obtained from the first 6 layers of the molecular sequence encoder are used to convert Z (g) and Z (s) Input normalized flow model separately and Get the molecular graph structure and the characteristic distribution Y of the molecular SMILES string (g) and Y (s) ;

[0124]

[0125] Where Y (g) represents the distribution of molecular graph structural features obtained by the normalized flow model;

[0126] Y (s) represents the molecular SMILES string feature distribution obtained by the normalized flow model;

[0127] Step 154: Y (g) and Y (s) Perform an averaging operation to sample potential text representations It is expressed as:

[0128]

[0129] Y (g) ~N(μ c ,∑ c ),Y (s) ~N(μ c ,∑ c )

[0130] In the formula, N(μ c ,∑ c ) represents the standard normal distribution of the molecular graph structure features and the molecular SMILES string features after being projected into the shared latent space by the normalized flow model; μ c represents the distribution center; ∑ c Represents covariance; ← represents the generation process of the normalized flow model;

[0131] Step 155: Enter to \ generate The formula is as follows:

[0132]

[0133] In the formula, express Inverse transform, represents the normalized flow model; express The distribution of missing text features generated by the inverse transformation of the normalized flow model;

[0134] Step 156: Input into the projection network MLP to obtain the shallow feature representation of the missing text modality;

[0135]

[0136] In the formula, Proj (t) () represents the feature projection network MLP.

[0137] The other steps and parameters are the same as those in Specific Implementation Methods 1 to 5.

[0138] Specific implementation eight: This implementation differs from any one of specific implementations one to seven in that in step two, the multimodal molecular pre-training model MolBT is pre-trained to obtain a pre-trained multimodal molecular pre-training model MolBT; the specific process is as follows:

[0139] Step 31: Construct the missing modality reconstruction loss, which is defined as follows:

[0140]

[0141] Where, Lrec represents the missing modality reconstruction loss; T k Represents the shallow feature representation of the real text (the molecular graph structure feature representation G6 obtained by the real molecular graph through the first 6 layers of the molecular graph structure encoder, the molecular description text feature representation T6 obtained by the real molecular description text through the first 6 layers of the molecular sequence encoder, and the molecular SMILES feature representation S6 obtained by the real molecular SMILES string through the first 6 layers of the molecular sequence encoder, denoted as shallow feature representation, i.e. {G6, T6, S6};);

[0142] Step 32: Construct a contrast loss function; the specific process is:

[0143] Step 321: Let {(g1,t1),(g2,t2),…,(g N ,t N )} is a batch of molecule-text pairs, N=16;

[0144] Among them, the molecule g i and the corresponding molecular description text t i Constitute a positive pair (g i ,t i );

[0145] Molecule g i Molecular description text with different molecules t j Negative pair (g i ,t j ) i≠j ;

[0146] Molecule g i Represents molecular graphs and molecular SMILES strings;

[0147] Step 322: Use the molecular graph structure encoder and the molecular sequence encoder to obtain the molecular graph structure representation g K , Molecular description text representation K and the SMILES string representation of the molecule s K ;

[0148] g K ,t K and K Input to the cross-modal encoder to obtain cross-modal feature representation and

[0149] Use a multi-layer perceptron (MLP) as a projection network to transform the four representations g K ,t K , Projected into the feature space of the same dimension; the formula is as follows:

[0150]

[0151]

[0152] In the formula, z G represents the structural feature representation of the projected molecular graph; z T represents the feature representation of the molecular description text after projection; c G represents the cross-modal feature representation of the projected molecular graph structure; c T Represents the cross-modal feature representation of the projected molecular description text;

[0153] GraphProj() represents the molecular graph structure feature projection network; TextProj() represents the molecular description text feature projection network;

[0154] GraphProj'() represents the cross-modal feature projection network of the molecular graph structure; TextProj'() represents the cross-modal feature projection network of the molecular description text;

[0155] d g The dimension representing the structural features of the molecular graph; d t The dimension representing the characteristics of the molecular description text;

[0156] The multimodal contrast loss is used to bring samples of different modalities with the same semantic information closer in the feature space, while pushing samples with different semantic information apart;

[0157] The contrast loss used includes three contrast losses between the four feature representations, namely (z G ,z T )、(z G ,c G ) and (c G ,c T );

[0158] Step 3: Use contrast loss to calculate (z G ,z T ) between the contrast loss function l1; the specific process is:

[0159] The characteristic representation of the i-th molecular graph structure after projection is The feature representation of the j-th molecular graph structure after projection is With the i-th molecular graph structure g i The projected feature representation of the molecular description text of the same molecule is With the i-th molecular graph structure g i Molecular description text of different molecules j The projected feature representation of The formula is as follows:

[0160]

[0161] in,

[0162] Represents the molecular graph structure-text contrast loss; represents the text-molecule graph structure contrast loss;

[0163] sim(·,·) represents cosine similarity;

[0164] τ is the temperature hyperparameter, set to 0.1;

[0165] Step 324: Use contrast loss to calculate (z G ,c G ) between the contrast loss function l2;

[0166] The characteristic representation of the i-th molecular graph structure after projection is The feature representation of the j-th molecular graph structure after projection is The cross-modal feature of the molecular graph structure after projection of the i-th molecular graph structure is expressed as With the i-th molecular graph structure g i Different molecules g j The cross-modal feature of the projected molecular graph structure is expressed as The formula is as follows:

[0167]

[0168] in,

[0169] Represents molecular graph structure feature-molecular graph structure cross-modal feature contrast loss; Represents the molecular graph structure cross-modal feature-molecular graph structure feature contrast loss;

[0170] sim(·,·) represents cosine similarity;

[0171] τ is the temperature hyperparameter, set to 0.1;

[0172] Step 325: Compute using contrast loss (c G ,c T ) between the contrast loss function l3;

[0173] The cross-modal feature of the i-th molecular graph structure after projection is expressed as The cross-modal feature of the projected j-th molecular graph structure is expressed as With the i-th molecular graph structure g i The cross-modal feature representation of the projected molecular description text of the same molecule is With the i-th molecular graph structure g i The cross-modal feature representation of the molecular description text after projection of different molecules is: The formula is as follows:

[0174]

[0175] in,

[0176] Represents the contrast loss of molecular graph structure cross-modal features and text cross-modal features; Represents the contrast loss of text cross-modal features-molecular graph structure cross-modal features;

[0177] sim(·,·) represents cosine similarity;

[0178] τ is the temperature hyperparameter, set to 0.1;

[0179] Step 326: Based on the contrast loss function l1, contrast loss function l2, and contrast loss function l3, construct the total contrast loss function L CL , expressed as:

[0180] L CL =(l1+l2+l3) / 3

[0181] Step 3. Reconstruct the loss L based on the missing mode rec And the total contrast loss function L CL , construct the loss function L of the missing modality generation module mmg , expressed as:

[0182] L mmg =αL rec +L CL

[0183] Wherein, α represents the coefficient, α=5;

[0184] Step 34: Train the multimodal molecular pre-training model MolBT to obtain a pre-trained multimodal molecular pre-training model MolBT; the specific process is:

[0185] Step 341. Pre-training phase 1:

[0186] The molecular graph structure encoder, molecular sequence encoder, cross-modal encoder, and missing modality generation module in the multimodal molecular pre-training model MolBT are pre-trained using real molecule-text pair data. The loss function is L mmg ;

[0187] The molecules in the real molecule-text pair data are molecule graphs and molecule SMILES strings;

[0188] The multimodal molecular pre-training model MolBT was pre-trained for 20 rounds using the AdamW optimizer, with the learning rate set to 1e-4 and the weight decay set to 1×10 -5 , the batch size is set to 16;

[0189] Obtaining a multimodal molecular pre-training model MolBT pre-trained in the pre-training stage 1;

[0190] Step 342: Pre-training phase 2:

[0191] Molecular data (molecule graph structures, molecule SMILES strings) using missing molecule description text;

[0192] The molecular graph and molecular SMILES string with missing molecular description text are used to pre-train the molecular graph structure encoder, molecular sequence encoder, cross-modal encoder, and missing modality generation module in the multimodal molecular pre-training model MolBT. The model weights saved in the pre-training stage 1 are applied, and the loss function is L CL ;

[0193] The multimodal molecular pre-training model MolBT was pre-trained for 20 rounds using the AdamW optimizer, with the learning rate set to 1e-4 and the weight decay set to 1×10 -5 , the batch size is set to 16;

[0194] Obtain the multimodal molecular pre-training model MolBT pre-trained in the second pre-training stage;

[0195] Step 343. Pre-training stage 3:

[0196] Use real molecules and text pairs to pre-train the molecular graph encoder, molecular sequence encoder, and cross-modal encoder in the multimodal molecular pre-training model MolBT, and apply the model weights saved in the second pre-training stage. The loss function is L CL ;

[0197] The molecules in the real molecule and text pair data are molecule graphs and molecule SMILES strings;

[0198] The AdamW optimizer was used to pre-train the molecular graph structure encoder, molecular sequence encoder, cross-modal encoder, and missing modality generation module in the multimodal molecular pre-training model MolBT for 20 rounds, with the learning rate set to 1e-4 and the weight decay set to 1×10 -5 , the batch size is set to 16;

[0199] The multimodal molecular pre-training model MolBT pre-trained in the pre-training stage three is obtained (the pre-trained multimodal molecular pre-training model MolBT includes a molecular graph structure encoder, a molecular sequence encoder, a cross-modal encoder and a missing modality generation module).

[0200] The other steps and parameters are the same as those in Specific Embodiments 1 to 7.

[0201] Specific implementation method 9: This implementation method is different from any one of specific implementation methods 1 to 8 in that, in step 3, the pre-trained multimodal molecular pre-training model MolBT is fine-tuned based on the downstream task type to obtain the fine-tuned multimodal molecular pre-training model MolBT; the specific process is:

[0202] 1) When the downstream task is molecular attribute prediction, eight molecular attribute data sets are input into the pre-trained multimodal molecular pre-training model MolBT to obtain a fine-tuned multimodal molecular pre-training model MolBT;

[0203] 8 molecular attribute datasets including BBBP, Tox21, ToxCast, SIDER, ClinTox, MUV, HIV, and BACE;

[0204] 2) When the downstream task is cross-modal retrieval of molecules and texts, the PCDes dataset is input into the pre-trained multimodal molecular pre-training model MolBT to obtain a fine-tuned multimodal molecular pre-training model MolBT.

[0205] The pre-trained multimodal molecular pre-training model MolBT is applied to a wide range of downstream tasks, such as molecular property prediction, molecule-text cross-modal retrieval, etc.

[0206] The other steps and parameters are the same as those in Specific Embodiments 1 to 8.

[0207] Specific implementation ten: This implementation is different from the first to ninth specific implementations in that in step four, downstream tasks are predicted based on the fine-tuned multimodal molecular pre-training model MolBT; the specific process is as follows:

[0208] 1) When the downstream task is molecular property prediction, the molecular properties are predicted based on the multimodal molecular pre-training model MolBT that has been fine-tuned when the downstream task is molecular property prediction; the specific process is:

[0209] The molecular graph to be tested is input into the molecular graph structure encoder of the fine-tuned multimodal molecular pre-training model MolBT, and the molecular graph structure encoder outputs the molecular graph structure feature representation; the molecular graph structure feature representation is input into the prediction head (linear layer), and the prediction head outputs the molecular attribute (used to predict whether the molecule has the target attribute);

[0210] Depending on the prediction task, different prediction heads can be used. Therefore, MolBT can complete binary or multi-classification tasks.

[0211] 2) When the downstream task is cross-modal retrieval of molecular graphs and molecular description texts, cross-modal retrieval of molecular graphs and molecular description texts is performed based on the multimodal molecular pre-training model MolBT that is fine-tuned when the downstream task is cross-modal retrieval of molecules and texts; the specific process is:

[0212] 21) Cross-modal retrieval of molecular graphs and molecular description texts includes two subtasks: M2T and T2M;

[0213] M2T means given a molecular graph, retrieving the molecular description text that matches the molecular graph;

[0214] T2M means given a molecular description text, retrieving a molecular graph that matches the molecular description text;

[0215] 22) In the M2T task, the molecular graph is input into the molecular graph structure encoder of the fine-tuned multimodal molecular pre-training model MolBT, and the molecular graph structure encoder of the fine-tuned multimodal molecular pre-training model MolBT outputs the molecular graph feature representation;

[0216] At the same time, all molecular description texts are input into the molecular sequence encoder of the fine-tuned multimodal molecular pre-training model MolBT, and the molecular sequence encoder of the fine-tuned multimodal molecular pre-training model MolBT outputs all molecular description text feature representations; the downstream tasks do not need the molecular SMILES string, and the molecular SMILES string is only used in the pre-training stage;

[0217] Calculate the cosine similarity between the molecular graph feature representation and all molecular description text feature representations and sort them. The molecular description text with the largest cosine similarity is the best matching molecular description text.

[0218] 23) In the T2M task, the molecular description text is input into the molecular sequence encoder of the fine-tuned multimodal molecular pre-training model MolBT, and the molecular sequence encoder of the fine-tuned multimodal molecular pre-training model MolBT outputs the molecular description text feature representation; the downstream task does not need the molecular SMILES string, and the molecular SMILES string is only used in the pre-training stage;

[0219] At the same time, all molecular graphs are input into the molecular graph structure encoder of the fine-tuned multimodal molecular pre-training model MolBT, and the molecular graph structure encoder of the fine-tuned multimodal molecular pre-training model MolBT outputs all molecular graph feature representations;

[0220] The cosine similarities between the feature representations of the molecular description text and the feature representations of all molecular graphs are calculated and sorted. The molecular graph with the largest cosine similarity is the best matching molecule.

[0221] The other steps and parameters are the same as those in Specific Embodiments 1 to 9.

[0222] The present invention may also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art may make various corresponding changes and modifications based on the present invention, but these corresponding changes and modifications should all fall within the scope of protection of the claims attached to the present invention.

Claims

1. A multimodal drug molecule prediction method based on missing modality generation, characterized by: The specific process of the method is: Step 1: Construct a multimodal molecular pre-training model MolBT; Step 2: pre-train the multimodal molecular pre-training model MolBT to obtain a pre-trained multimodal molecular pre-training model MolBT; Step 3: Based on the downstream task type, fine-tune the pre-trained multimodal molecular pre-training model MolBT to obtain the fine-tuned multimodal molecular pre-training model MolBT; Step 4: Predict downstream tasks based on the fine-tuned multimodal molecular pre-training model MolBT.

2. The multimodal drug molecule prediction method based on missing modality generation according to claim 1, characterized in that: In the step 1, a multimodal molecular pre-training model MolBT is constructed: the specific process is as follows: Step 11: input the molecular graph into a molecular graph structure encoder, and the molecular graph structure encoder outputs a molecular graph structure feature representation; The molecular graph structure encoder is a Graphormer; Step 1 and 2: input the molecular description text into the molecular sequence encoder, and the molecular sequence encoder outputs the molecular description text feature representation; The molecular sequence encoder is BERT; Step 13: Input the molecular SMILES string into the molecular sequence encoder, and the molecular sequence encoder outputs the molecular SMILES feature representation; The molecular sequence encoder is BERT; Step 14: Establish a cross-modal encoder, input the molecular graph structure feature representation, the molecular description text feature representation, and the molecular SMILES feature representation into the cross-modal encoder, and the cross-modal encoder outputs the molecular graph structure cross-modal feature representation and the molecular description text cross-modal feature representation; Step 15: Construct a missing modality generation module, input the molecular graph structure feature representation and molecular SMILES feature representation of the missing molecular description text into the missing modality generation module, and the missing modality generation module outputs the corresponding shallow feature representation of the molecular description text.

3. The multimodal drug molecule prediction method based on missing modality generation according to claim 2, characterized in that: In the step one, the molecular graph is input into a molecular graph structure encoder, and the molecular graph structure encoder outputs a molecular graph structure feature representation; The specific process is: Step 1: Input the molecular graph into the graph feature encoding layer in the molecular graph structure encoder, and the graph feature encoding layer in the molecular graph structure encoder outputs the molecular graph node feature G0. in, represents the first atomic feature embedding, represents the second atomic feature embedding, represents the nth atomic feature embedding, R represents a real number, n g Indicates the number of atoms in the molecular graph, d g Represents the feature dimension of the molecular graph structure encoder; Step 112: input the molecular graph node feature G0 into the first Graphormer layer, the second Graphormer layer, the third Graphormer layer, the fourth Graphormer layer, the fifth Graphormer layer, the sixth Graphormer layer, the seventh Graphormer layer, the eighth Graphormer layer, the ninth Graphormer layer, the tenth Graphormer layer, the eleventh Graphormer layer, and the twelfth Graphormer layer of the molecular graph structure encoder in sequence, and the twelfth Graphormer layer outputs the molecular graph structure feature representation G 12 ; The formula is as follows: In the formula, K g Indicates the total number of Graphormer layers in the molecular graph structure encoder, K g =12; represents the i-th Graphormer layer; G i-1 Represents the molecular graph structure features output by the i-1th Graphormer layer; G i Represents the molecular graph structural features output by the i-th Graphormer layer.

4. The multimodal drug molecule prediction method based on missing modality generation according to claim 3, characterized in that: In the steps 1 and 2, the molecular description text is input into the molecular sequence encoder, and the molecular sequence encoder outputs the molecular description text feature representation; the specific process is: Step 121: Input the molecular description text into the BERT word embedding layer of the molecular sequence encoder, and the BERT word embedding layer of the molecular sequence encoder outputs the molecular description text feature T0; Step 122: Input the molecular description text feature T0 into the 1st Transformer layer, 2nd Transformer layer, 3rd Transformer layer, 4th Transformer layer, 5th Transformer layer, 6th Transformer layer, 7th Transformer layer, 8th Transformer layer, 9th Transformer layer, 10th Transformer layer, 11th Transformer layer, and 12th Transformer layer of the molecular sequence encoder in sequence, and the 12th Transformer layer outputs the molecular description text feature representation T 12 ; The formula is as follows: In the formula, K s represents the total number of Transformer layers in the molecular sequence encoder, K s =12; represents the i-th Transformer layer; T i-1 Represents the molecular description text features output by the i-1th Transformer layer; T i Represents the molecular description text features output by the i-th Transformer layer.

5. The multimodal drug molecule prediction method based on missing modality generation according to claim 4, characterized in that: In the step 1-3, the molecular SMILES character string is input into the molecular sequence encoder, and the molecular sequence encoder outputs the molecular SMILES feature representation; The specific process is: Step 131. Input the molecule SMILES string into the BERT word embedding layer of the molecule sequence encoder, and the BERT word embedding layer of the molecule sequence encoder outputs the molecule SMILES feature S0; Step 132: Input the molecular SMILES feature S0 into the 1st Transformer layer, the 2nd Transformer layer, the 3rd Transformer layer, the 4th Transformer layer, the 5th Transformer layer, the 6th Transformer layer, the 7th Transformer layer, the 8th Transformer layer, the 9th Transformer layer, the 10th Transformer layer, the 11th Transformer layer, and the 12th Transformer layer of the molecular sequence encoder in sequence, and the 12th Transformer layer outputs the molecular SMILES feature representation S 12 ; The formula is as follows: In the formula, K s represents the total number of Transformer layers in the molecular sequence encoder, K s =12; represents the i-th Transformer layer; S i-1 Represents the molecular SMILES features output by the i-1th Transformer layer; S i Represents the molecular SMILES features output by the i-th Transformer layer.

6. The multimodal drug molecule prediction method based on missing modality generation according to claim 5, characterized in that: In the step 1-4, a cross-modal encoder is established, and the molecular graph structure feature representation, the molecular description text feature representation, and the molecular SMILES feature representation are input into the cross-modal encoder, and the cross-modal encoder outputs the molecular graph structure cross-modal feature representation and the molecular description text cross-modal feature representation; the specific process is: The output of the 7th layer of the molecular graph encoder G7, the output of the 7th layer of the molecular sequence encoder T7 and S7 are input into the cross-modal encoder, and the output of the cross-modal encoder C G ,C T ; The output of the 8th layer of the molecular graph structure encoder G8, the output of the 8th layer of the molecular sequence encoder T8 and S8 are input into the cross-modal encoder, and the output of the cross-modal encoder C G ,C T ; The output of the 9th layer of the molecular graph encoder G9, the output of the 9th layer of the molecular sequence encoder T9 and S9 are input into the cross-modal encoder, and the output of the cross-modal encoder C G ,C T ; The output of the molecular graph structure encoder layer 10 is G 10 , T output from the 10th layer of the molecular sequence encoder 10 and S 10 Input cross-modal encoder, cross-modal encoder output C G ,C T ; The output of the molecular graph structure encoder layer 11 is G 11 , T output from the 11th layer of the molecular sequence encoder 11 and S 11 Input cross-modal encoder, cross-modal encoder output C G ,C T ; G output from the 12th layer of the molecular graph structure encoder 12 , T output from the 12th layer of the molecular sequence encoder 12 and S 12 Input cross-modal encoder, cross-modal encoder output C G ,C T ; The cross-modal encoder has 6 layers, each of which includes a self-attention layer, a common attention layer, and a feed-forward layer FFN in sequence; The operation process of each layer in the cross-modal encoder is: In the formula, CME l represents the lth layer of the cross-modal encoder; Each layer of the cross-modal encoder includes a self-attention layer, a common attention layer, and a feed-forward layer FFN in sequence; and They represent the molecular graph structure cross-modal feature representation and the molecular text cross-modal feature representation output by the l-th layer of the cross-modal encoder respectively; L C is the number of layers of the cross-modal encoder, L C =6; Represents the molecular graph structure feature representation of the output of the l-1th layer of the cross-modal encoder; Represents the text feature representation of the output of the l-1 layer of the cross-modal encoder; Represents the molecular SMILES feature representation of the l-th layer input of the cross-modal encoder; The specific process is: 1) Represent the structural features of the molecular graph After BridgeLayer processing, the input of the lth layer of the cross-modal encoder is obtained Representing text features After BridgeLayer processing, the input of the lth layer of the cross-modal encoder is obtained It is expressed as: in In the formula, W G Represents the projection matrix of molecular graph structure features; W T Represents the text feature projection matrix; G type Represents the molecular graph structure feature representation G k After the type embedding layer, the modal type embedding is obtained; T type Represents the molecular description text feature representation T k After the type embedding layer, the modal type embedding is obtained; G k and T k Respectively represent the molecular graph structure feature representation and molecular description text feature representation output by the kth layer of the molecular graph structure encoder and the molecular sequence encoder, k = 7, 8, ..., 12; Representation layer normalization LN; Representation layer normalization LN; 2) Feature representation based on molecular SMILES strings Extract class tags, denoted as In the formula, Represents a SMILES string feature representation of a molecule S k The first word feature embedding in ; Represents a SMILES string feature representation of a molecule S k The second word feature embedding in ; Represents a SMILES string feature representation of a molecule S k The nth s word-unit feature embedding; Represents a SMILES string feature representation of a molecule S k The class tag of n s Indicates the length of the molecule SMILES string; d s Represents the feature dimension of the molecular sequence encoder; 3) and The last node feature in is replaced by get and As shown below: In the formula, Represents the cross-modal feature representation of the molecular graph structure after replacing the features; express The first atomic feature embedding in ; express The second atomic feature embedding in ; express The (n-1)th atomic feature embedding in ; Represents the cross-modal feature representation of the molecular description text after replacing the features; express The first word feature embedding in ; express The second word feature embedding in ; express The (n-1)th word feature embedding in ; n g Indicates the number of atoms in the molecular graph; d g Represents the feature dimension of the molecular graph structure encoder; n t Indicates the length of the molecular description text; d t Represents the feature dimension of the text encoder; 4) The molecular graph structure part and the molecular text part are input to the lth layer of the cross-modal encoder to obtain and In the formula, Represents the cross-modal feature representation of the molecular graph structure output of the l-th layer of the cross-modal encoder; Indicates Cross-modal feature representation of processed molecular graph structure; Represents the text cross-modal feature representation output by the l-th layer of the cross-modal encoder; Indicates Cross-modal feature representation of processed molecular description text; Indicates the replaced Cross-modal feature representation of processed molecular graph structure; Indicates the replaced Cross-modal feature representation of processed molecular description text; CME l represents the lth layer of the cross-modal encoder; Each layer of the cross-modal encoder consists of a self-attention layer, a common attention layer, and a feed-forward layer FFN in sequence.

7. The multimodal drug molecule prediction method based on missing modality generation according to claim 6, characterized in that: In the step 15, a missing modality generation module is constructed, the molecular graph and the molecular SMILES string of the missing molecular description text are input into the missing modality generation module, and the missing modality generation module outputs the corresponding molecular description text; the specific process is: Step 151, the molecular graph structure feature representation G6 obtained through the first 6 layers of the molecular graph structure encoder, the molecular description text feature representation T6 obtained through the first 6 layers of the molecular sequence encoder, and the molecular SMILES feature representation S6 are recorded as shallow feature representations, i.e., {G6, T6, S6}; Step 152: Project the shallow feature representations G6, T6, and S6 into the same dimensional space through a one-dimensional convolution layer to obtain Z (m) ∈R T×d ,m∈{t,g,s}; In the formula, Represents Z (t) , Z (g) or Z (s) ; Z (t) The shallow feature represents the output feature of T6 after one one-dimensional convolution layer; Z (g) The shallow feature represents the output feature of G6 after one one-dimensional convolution layer; Z (s) The shallow feature represents the feature output by S6 after passing through one one-dimensional convolutional layer; T represents the feature length after the one-dimensional convolution layer; d represents the feature dimension after the one-dimensional convolution layer; Step 153: In the absence of molecular description text modality, use the molecular graph structure feature representation G6 obtained from the first 6 layers of the molecular graph structure encoder and the molecular SMILES feature representation S6 obtained from the first 6 layers of the molecular sequence encoder to convert Z (g) and Z (s) Input normalized flow model separately and Get the molecular graph structure and the characteristic distribution Y of the molecular SMILES string (g) and Y (s) ; Where Y (g) represents the distribution of molecular graph structural features obtained by the normalized flow model; Y (s) represents the molecular SMILES string feature distribution obtained by the normalized flow model; Step 154: Y (g) and Y (s) Perform an averaging operation to sample potential text representations It is expressed as: Y (g) ~N(μ c ,∑ c ),Y (s) ~N(μ c ,∑ c ) In the formula, N(μ c ,∑ c ) represents the standard normal distribution of the molecular graph structure features and the molecular SMILES string features after being projected into the shared latent space by the normalized flow model; μ c represents the distribution center; ∑ c Represents covariance; ← represents the generation process of the normalized flow model; Step 155: Input to generate The formula is as follows: In the formula, express Inverse transform, represents the normalized flow model; express The distribution of missing text features generated by the inverse transformation of the normalized flow model; Step 156: Input into the projection network MLP to obtain the shallow feature representation of the missing text modality; In the formula, Proj (t) () represents the feature projection network MLP.

8. The multimodal drug molecule prediction method based on missing modality generation according to claim 7, characterized in that: In the step 2, the multimodal molecular pre-training model MolBT is pre-trained to obtain a pre-trained multimodal molecular pre-training model MolBT; the specific process is: Step 31: Construct the missing modality reconstruction loss, which is defined as follows: Where, L rec represents the missing modality reconstruction loss; T k Represents the real shallow feature representation of text; Step 32: Construct a contrast loss function; the specific process is: Step 321: Let {(g1,t1),(g2,t2),…,(g N ,t N )} is a batch of molecule-text pairs, N=16; in, Molecule g i and the corresponding molecular description text t i Constitute a positive pair (g i ,t i ); Molecule g i Molecular description text with different molecules t j Negative pair (g i ,t j ) i≠j ; Molecule g i Represents molecular graphs and molecular SMILES strings; Step 322: Use the molecular graph structure encoder and the molecular sequence encoder to obtain the molecular graph structure representation g K , Molecular description text representation K and the SMILES string representation of the molecule s K ; g K ,t K and K Input to the cross-modal encoder to obtain cross-modal feature representation and Use a multi-layer perceptron (MLP) as a projection network to transform the four representations g K ,t K , Projected into the feature space of the same dimension; the formula is as follows: In the formula, z G represents the structural feature representation of the projected molecular graph; z T represents the feature representation of the molecular description text after projection; c G represents the cross-modal feature representation of the projected molecular graph structure; c T Represents the cross-modal feature representation of the projected molecular description text; GraphProj() represents the molecular graph structure feature projection network; TextProj() represents the molecular description text feature projection network; GraphProj'() represents the cross-modal feature projection network of the molecular graph structure; TextProj'() represents the cross-modal feature projection network of the molecular description text; d g The dimension representing the structural features of the molecular graph; d t The dimension representing the characteristics of the molecular description text; Step 3: Use contrast loss to calculate (z G ,z T ) between the contrast loss function l1; the specific process is: The characteristic representation of the i-th molecular graph structure after projection is The feature representation of the j-th molecular graph structure after projection is With the i-th molecular graph structure g i The projected feature representation of the molecular description text of the same molecule is With the i-th molecular graph structure g i Molecular description text of different molecules j The projected feature representation of The formula is as follows: in, Represents molecular graph structure-text contrast loss; represents the text-molecule graph structure contrast loss; sim(·,·) represents cosine similarity; τ is the temperature hyperparameter, set to 0.1; Step 324: Use contrast loss to calculate (z G ,c G ) between the contrast loss function l2; The characteristic representation of the i-th molecular graph structure after projection is The feature representation of the j-th molecular graph structure after projection is The cross-modal feature of the molecular graph structure after projection of the i-th molecular graph structure is expressed as With the i-th molecular graph structure g i Different molecules g j The cross-modal feature of the projected molecular graph structure is expressed as The formula is as follows: in, Represents molecular graph structure feature-molecular graph structure cross-modal feature contrast loss; Represents the molecular graph structure cross-modal feature-molecular graph structure feature contrast loss; sim(·,·) represents cosine similarity; τ is the temperature hyperparameter, set to 0.1; Step 325: Use contrast loss calculation (c G ,c T ) between the contrast loss function l3; The cross-modal feature of the i-th molecular graph structure after projection is expressed as The cross-modal feature of the projected j-th molecular graph structure is expressed as With the i-th molecular graph structure g i The cross-modal feature representation of the projected molecular description text of the same molecule is With the i-th molecular graph structure g i The cross-modal feature representation of the molecular description text after projection of different molecules is: The formula is as follows: in, Represents the contrast loss of molecular graph structure cross-modal features and text cross-modal features; Represents the contrast loss of text cross-modal features-molecular graph structure cross-modal features; sim(·,·) represents cosine similarity; τ is the temperature hyperparameter, set to 0.1; Step 326: Based on the contrast loss function l1, contrast loss function l2, and contrast loss function l3, construct the total contrast loss function L CL , expressed as: L CL =(l1+l2+l3) / 3 Step 3. Reconstruct the loss L based on the missing mode rec And the total contrast loss function L CL , construct the loss function L of the missing modality generation module mmg , expressed as: L mmg =αL rec +L CL Wherein, α represents the coefficient, α=5; Step 34: Train the multimodal molecular pre-training model MolBT to obtain a pre-trained multimodal molecular pre-training model MolBT; the specific process is: Step 341. Pre-training phase 1: The molecular graph structure encoder, molecular sequence encoder, cross-modal encoder, and missing modality generation module in the multimodal molecular pre-training model MolBT are pre-trained using real molecule-text pair data. The loss function is L mmg ; The molecules in the real molecule-text pair data are molecule graphs and molecule SMILES strings; The multimodal molecular pre-training model MolBT was pre-trained for 20 rounds using the AdamW optimizer, with the learning rate set to 1e-4 and the weight decay set to 1×10 -5 , the batch size is set to 16; Obtaining a multimodal molecular pre-training model MolBT pre-trained in the pre-training stage 1; Step 342: Pre-training phase 2: Molecular data using missing molecular description text; The molecular graph and molecular SMILES string with missing molecular description text are used to pre-train the molecular graph structure encoder, molecular sequence encoder, cross-modal encoder, and missing modality generation module in the multimodal molecular pre-training model MolBT. The model weights saved in the pre-training stage 1 are applied, and the loss function is L CL ; The multimodal molecular pre-training model MolBT was pre-trained for 20 rounds using the AdamW optimizer, with the learning rate set to 1e-4 and the weight decay set to 1×10 -5 , the batch size is set to 16; Obtain the multimodal molecular pre-training model MolBT pre-trained in the second pre-training stage; Step 343. Pre-training stage 3: Use real molecules and text pairs to pre-train the molecular graph encoder, molecular sequence encoder, and cross-modal encoder in the multimodal molecular pre-training model MolBT, and apply the model weights saved in the second pre-training stage. The loss function is L CL ; The molecules in the real molecule and text pair data are molecule graphs and molecule SMILES strings; The AdamW optimizer was used to pre-train the molecular graph structure encoder, molecular sequence encoder, cross-modal encoder, and missing modality generation module in the multimodal molecular pre-training model MolBT for 20 rounds, with the learning rate set to 1e-4 and the weight decay set to 1×10 -5 , the batch size is set to 16; The multimodal molecular pre-training model MolBT pre-trained in the pre-training stage three is obtained.

9. The multimodal drug molecule prediction method based on missing modality generation according to claim 8, characterized in that: In the step 3, based on the downstream task type, the pre-trained multimodal molecular pre-training model MolBT is fine-tuned to obtain the fine-tuned multimodal molecular pre-training model MolBT; the specific process is: 1) When the downstream task is molecular attribute prediction, eight molecular attribute data sets are input into the pre-trained multimodal molecular pre-training model MolBT to obtain a fine-tuned multimodal molecular pre-training model MolBT; 8 molecular attribute datasets including BBBP, Tox21, ToxCast, SIDER, ClinTox, MUV, HIV, and BACE; 2) When the downstream task is cross-modal retrieval of molecules and texts, the PCDes dataset is input into the pre-trained multimodal molecular pre-training model MolBT to obtain a fine-tuned multimodal molecular pre-training model MolBT.

10. The multimodal drug molecule prediction method based on missing modality generation according to claim 9, characterized in that: In the step 4, the downstream tasks are predicted based on the fine-tuned multimodal molecular pre-training model MolBT; the specific process is as follows: 1) When the downstream task is molecular property prediction, the molecular properties are predicted based on the multimodal molecular pre-training model MolBT that has been fine-tuned when the downstream task is molecular property prediction; the specific process is: The molecular graph to be tested is input into the molecular graph structure encoder of the fine-tuned multimodal molecular pre-training model MolBT, and the molecular graph structure encoder outputs the molecular graph structure feature representation; the molecular graph structure feature representation is input into the prediction head (linear layer), and the prediction head outputs the molecular attributes; 2) When the downstream task is cross-modal retrieval of molecular graphs and molecular description texts, cross-modal retrieval of molecular graphs and molecular description texts is performed based on the multimodal molecular pre-training model MolBT that is fine-tuned when the downstream task is cross-modal retrieval of molecules and texts; the specific process is: 21) Cross-modal retrieval of molecular graphs and molecular description texts includes two subtasks: M2T and T2M; M2T means given a molecular graph, retrieving the molecular description text that matches the molecular graph; T2M means given a molecular description text, retrieving a molecular graph that matches the molecular description text; 22) In the M2T task, the molecular graph is input into the molecular graph structure encoder of the fine-tuned multimodal molecular pre-training model MolBT, and the molecular graph structure encoder of the fine-tuned multimodal molecular pre-training model MolBT outputs the molecular graph feature representation; At the same time, all molecular description texts are input into the molecular sequence encoder of the fine-tuned multimodal molecular pre-training model MolBT, and the molecular sequence encoder of the fine-tuned multimodal molecular pre-training model MolBT outputs the feature representation of all molecular description texts; Calculate the cosine similarity between the molecular graph feature representation and all molecular description text feature representations and sort them. The molecular description text with the largest cosine similarity is the best matching molecular description text. 23) In the T2M task, the molecular description text is input into the molecular sequence encoder of the fine-tuned multimodal molecular pre-training model MolBT, and the molecular sequence encoder of the fine-tuned multimodal molecular pre-training model MolBT outputs the molecular description text feature representation; At the same time, all molecular graphs are input into the molecular graph structure encoder of the fine-tuned multimodal molecular pre-training model MolBT, and the molecular graph structure encoder of the fine-tuned multimodal molecular pre-training model MolBT outputs all molecular graph feature representations; The cosine similarities between the feature representations of the molecular description text and the feature representations of all molecular graphs are calculated and sorted. The molecular graph with the largest cosine similarity is the best matching molecule.

Citation Information

Patent Citations

  • Drug molecule generation method based on adversarial imitation learning

    CN112820361A

  • Multi-modal drug-protein target interaction prediction method and system

    CN115985386A