Compound blood brain barrier permeability prediction method based on multi-modal fusion

By adopting a multimodal fusion method in the prediction of the permeability of compound blood-brain barrier, combining molecular maps, texts, descriptors and three-dimensional features, and using the Kolmogorov-Arnold model for prediction, the problem of insufficient fusion of three-dimensional structures and multimodal features in the existing methods is solved, and the accuracy and generalization ability of prediction are significantly improved.

CN120199362AActive Publication Date: 2025-06-24SICHUAN JINGLANG INTELLECTUAL PROPERTY AGENCY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510314608.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-24
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

Existing methods fail to fully utilize the three-dimensional structural information of the molecular in the prediction of compound blood-brain barrier permeability, resulting in low prediction accuracy; at the same time, data imbalance and insufficient fusion of multimodal features also limit the performance of the model.

Method used

A method for predicting the permeability of compound blood-brain barriers based on multimodal fusion is proposed. By extracting molecular map features, molecular text features, molecular descriptor features and molecular three-dimensional features, and performing feature fusion, the Kolmogorov-Arnold model is used for model training and prediction.

Benefits of technology

Through multimodal feature fusion, the accuracy and generalization ability of blood-brain barrier permeability prediction are improved, data imbalance and feature alignment problems are solved, and the comprehensiveness and accuracy of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199362A_ABST
    Figure CN120199362A_ABST
Patent Text Reader

Abstract

The invention provides a compound blood brain barrier permeability prediction method based on multi-modal fusion, and the method comprises the following specific implementation steps: (1) collecting SMILES expressions of compounds with blood brain barrier permeability labels, and carrying out the data preprocessing of all compounds, so as to obtain an enhanced data set # imgabs0 #; (2) generating a molecular three-dimensional structure feature 3DF, a molecular text feature MNF, a molecular descriptor feature MDF and a molecular graph feature MGF by using an SMILES expression of each compound in # imgabs1 #, and performing normalization processing on the MNF, the 3DF, the MDF and the MGF to obtain a normalized molecular three-dimensional structure feature P-3DF, a normalized text feature P-MNF, a normalized molecular descriptor feature P-MDF and a normalized molecular graph feature P-MGF; (3) constructing a feature fusion network BBBNet, wherein the feature fusion network BBBNet is used for constructing a multi-modal fusion feature vector data set # imgabs2; (4) the # imgabs3 # is divided into a training set and a test set, and the training set is input into a Kolmogov-Arnod model KAN for model training; after the training is completed, performing performance evaluation on the trained KAN model by using a test set; and (5) predicting the blood-brain barrier permeability of the compound by using the trained KAN model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the field of bioinformatics processing, and further relates to a method for predicting the blood-brain barrier permeability of compounds based on multimodal fusion in the field of deep learning technology. The present invention can be used to predict the blood-brain barrier permeability of drugs under research and development. Background Art

[0002] During the research and development of compounds for the central nervous system (CNS), the permeability of the blood-brain barrier (BBB) is a key indicator. As a barrier that protects the brain and its surrounding neurons, the BBB can selectively screen and restrict the entry and exit of substances inside the brain, preventing harmful substances, pathogens, or foreign substances from entering the brain, while maintaining the internal environment stability of brain cells. Therefore, evaluating the BBB permeability of candidate drugs is of great significance in early drug discovery and development, especially in optimizing the efficacy of CNS drugs. However, although traditional clinical experimental methods are accurate, they are often costly and time-consuming.

[0003] In recent years, deep learning and machine learning methods have been widely used in predicting the permeability of the blood-brain barrier (BBB) of compounds, and significant progress has been made in improving the prediction efficiency and accuracy. However, there is still room for performance improvement in existing methods. Most studies focus on molecular text features (SMILES expressions) or molecular graph features, ignoring the role of the three-dimensional structure of molecules. In fact, the three-dimensional structure of molecules plays a crucial role in evaluating their physicochemical properties and interactions with biological targets. Therefore, the lack of full utilization of the three-dimensional structure limits the comprehensiveness and accuracy of the model.

[0004] In addition, the problem of data imbalance remains a challenge. In the current compound dataset, there are relatively more compounds that can penetrate the blood-brain barrier (BBB+), while there are fewer compounds that cannot penetrate the blood-brain barrier (BBB-), which results in poor prediction accuracy for the minority class. Therefore, existing models are difficult to effectively solve this problem when facing actual compound research and development. The training of deep learning models usually requires a large amount of sample data, and in the case of insufficient data volume, the model may not be able to fully learn, resulting in insufficient training, which affects the final prediction result of blood-brain barrier permeability. In addition, the current multimodal feature fusion methods have deficiencies in feature alignment and fusion mechanisms, and fail to fully explore the potential associations between different features, thus further limiting the blood-brain barrier permeability prediction ability and generalization ability of the model. Therefore, how to effectively solve problems such as data imbalance, insufficient data volume, and multimodal feature fusion remains an important challenge in the current prediction of blood-brain barrier permeability of compounds. Summary of the Invention

[0005] The purpose of this section is to outline some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section, as well as in the abstract and title of the present application, to avoid obscuring the purpose of this section, the abstract, and the title. However, such simplifications or omissions shall not be used to limit the scope of the present invention; The object of the present invention is to address the problems existing in the above-mentioned background art and propose a method for predicting the blood-brain barrier permeability of compounds based on multimodal fusion.

[0006] Term Explanation: Multimodal: A modality refers to different types or forms of data, and different modalities represent different data sources or perception channels. Multimodality involves integrating and processing data from multiple modalities for a more comprehensive understanding and reasoning.

[0007] Blood-brain barrier permeability label: A label indicating whether a compound can penetrate the blood-brain barrier. If a compound can penetrate the blood-brain barrier, its blood-brain barrier permeability label is BBB+; otherwise, it is recorded as BBB-.

[0008] Positive sample: A compound sample that can cross the blood-brain barrier, with a blood-brain barrier permeability label of BBB+, also known as a BBB+ sample.

[0009] Negative sample: A compound sample that cannot cross the blood-brain barrier, with a blood-brain barrier permeability label of BBB-, also known as a BBB- sample.

[0010] Minority class sample: In a classification task, a minority class sample refers to a sample of a certain category in a dataset whose quantity is significantly less than that of other categories and usually accounts for a relatively small proportion in the overall distribution.

[0011] ADASYN: An oversampling technique for dealing with class imbalance. By generating new minority class samples, especially more samples in the boundary region, it makes the classifier pay more attention to the minority class samples and improves the model's recognition ability for the minority class.

[0012] Simplified molecular linear input specification: A specification for describing molecular structures using ASCII characters. Its main feature is to represent the atoms and connection methods of compounds through strings, making the input, storage, and computer processing of molecular structures more convenient.

[0013] SMILES expression: A simplified linear input specification for molecules, often used to represent chemical structures. It consists of a series of atomic symbols and connectors, capable of accurately describing the structure of a molecule for computer processing and storage. SMILES expressions can be imported into most molecular editing software and converted into two-dimensional graphs or three-dimensional models of molecules. Based on the main principle of chemoinformatics that similar molecules have similar properties, it can thus be used to predict biochemical properties.

[0014] Three-dimensional molecular structure: The three-dimensional molecular structure refers to the specific shape of a molecule in three-dimensional space, determined by the spatial arrangement of atoms, stereochemical factors, and intermolecular interactions, etc., which has an important impact on the physical, chemical, and biological activities of the molecule.

[0015] Kolmogorov-Arnold model (KAN): Designed based on the Kolmogorov-Arnold representation theorem, it can decompose multi-dimensional input features layer by layer into univariate combinations to approximate complex multi-dimensional functions.

[0016] Chemical toolkit: A set of computer programs for processing chemical information and molecular data. These tools can be used for tasks such as molecular modeling, predicting chemical properties, and analyzing molecular structures. Some commonly used chemical toolkits include RDKit, Open Babel, ChemAxon, etc. These toolkits provide rich functions and algorithms, which can be used to generate molecular fingerprints and descriptors for fields such as molecular similarity calculation, compound screening, and quantitative structure-activity relationship (QSAR) modeling.

[0017] Feature matrix (Node Features): The feature matrix is used to represent the attributes of each node (i.e., atom) in a graph. Each row represents a node, and each column corresponds to a feature of the node, such as atom type, atomic valence, number of bonds, etc.

[0018] Layer Normalization: Normalizes the output of each layer to improve the training speed and stability of the model.

[0019] Query vector space: In the attention mechanism, the query vector space is the multi-dimensional space where the query vector is located. In this space, the query vector is used to match the key vector to calculate the attention weights, thereby determining the importance of each part in the information aggregation process.

[0020] Key vector space: In the attention mechanism, the key vector space is the multi-dimensional space where the key vector is located. The key vector and the query vector are compared in this space to measure the relevance of different information segments to the current query, and then determine the weights for attention allocation.

[0021] Value vector space: In the attention mechanism, the value vector space is a multi-dimensional space where value vectors are located. Each value vector corresponds to a piece of information, and its relationship with the query vector and key vector is adjusted through attention weights.

[0022] Self-Attention Score: It is a value that measures the correlation or similarity between different elements. In the self-attention mechanism, the self-attention score is used to determine the association strength between an element and other elements, thereby determining how the element should be weighted in the final output.

[0023] n-gram model: The n-gram model is a probability-based language model widely used for text modeling in natural language processing. It estimates the probability of a word or character in a sequence by analyzing the occurrence frequencies of consecutive substrings of length n (n-grams).

[0024] n-grams vocabulary: It refers to the set of all possible n-grams included in the n-gram model. N-grams are subsequences composed of consecutive n words or characters, used to represent the local context relationship of the text. The size of the vocabulary depends on the number of all possible n-grams in the text and is an important basis for constructing the n-gram model.

[0025] Substring of length 2: It refers to all possible subsequences composed of two consecutive characters extracted from a string in the order of characters.

[0026] The object of the present invention is directed to the following problems: (1) In terms of feature extraction, existing methods only focus on molecular graph features and molecular text features, and do not fully utilize molecular three-dimensional information, resulting in the inability to comprehensively capture the spatial structure features of molecules and a low accuracy rate in the blood-brain barrier permeability prediction task; (2) Existing methods have the following defects in multi-modal feature fusion: On the one hand, the alignment problem between different modal features is not considered. There are significant differences in feature scales and distributions between different modal data, and it is difficult for traditional weighted average methods to achieve effective cross-modal semantic alignment. On the other hand, the feature fusion strategy with fixed weights cannot adapt to the differences in modalities of different molecular samples, resulting in insufficient sensitivity to key modal information; (3) The models trained and processed based on the blood-brain barrier dataset with insufficient data have the following defects: on the one hand, insufficient data leads to the model's inability to learn the complex relationships of molecules, thus resulting in insufficient generalization ability of the model; on the other hand, when the data volume is small, the gradients of error backpropagation become very small, causing the weight updates of the first few layers of the network to be extremely slow, and finally resulting in the problem of gradient disappearance. In addition, most blood-brain barrier datasets also have the problem of data imbalance, that is, the BBB+ sample data is much more than the BBB- sample data, making the existing methods have a lower accuracy in predicting compounds that cannot penetrate the blood-brain barrier.

[0027] To solve the above problems, the present invention proposes a method for predicting the blood-brain barrier permeability of compounds based on multi-modal fusion. This method extracts molecular graph features, molecular text features, molecular descriptor features, and molecular three-dimensional features, and then fuses these four features. This method can predict molecular properties, understand chemical behaviors, and increase the accuracy of blood-brain barrier permeability prediction; The method for predicting the blood-brain barrier permeability of compounds based on multi-modal fusion has the following steps: S1: Represent each collected compound as a binary tuple , where is the SMILES expression of the compound , and is the blood-brain barrier permeability label of ; perform data preprocessing on all collected compounds to obtain an enhanced dataset ; S2: For any compound , generate molecular three-dimensional structure features 3DF, molecular text features MNF, molecular descriptor features MDF, and molecular graph features MGF; perform normalization processing on 3DF, MNF, MDF, and MGF to obtain the normalized molecular three-dimensional structure features P-3DF, the normalized text features P-MNF, the normalized molecular descriptor features P-MDF, and the normalized molecular graph features P-MGF; S3: Construct a feature fusion network BBBNet for constructing a multi-modal fusion feature vector dataset ; S4: Divide into a training set and a test set, and input the training set into the Kolmogorov-Arnold model KAN for model training; after the training is completed, use the test set to evaluate the performance of the trained KAN model; S5: Use the trained KAN model to predict the blood-brain barrier permeability of compounds.

[0028] As a preferred embodiment of the method for predicting the blood-brain barrier permeability of compounds based on multimodal fusion according to the present invention, the specific steps of the data preprocessing in S1 are as follows: S11: Apply resampling technology to Perform data balancing operations, and the specific steps are as follows: For any binary group of minority class samples , for Use the ADASYN method to generate a new SMILES expression , and then construct a binary group To represent the new minority class sample and add it to ; S12: Apply SMILES enhancement operations to each compound in : (1) For any structurally asymmetric compound , rotate Around the central atom and along the axis perpendicular to the molecular plane to obtain the rotated SMILES expression , construct a binary group , and add To , where the rotation operation is expressed as follows: Among them, Represents the function for performing the rotation operation on , Is a randomly selected rotation angle; (2) For compounds containing more than two closed rings , make the following adjustments to all the closed rings in the compound: Randomly rearrange the connection order of all atoms on the closed ring, and generate a SMILES expression according to the adjusted atomic connection order , construct a binary group , and add To ; (3) For compounds containing double bonds , randomly select a double bond and exchange the stereochemical label of the double bond. If the original stereochemical label is cis, adjust it to trans. If the original stereochemical label is trans, adjust it to cis; Generate the SMILES expression after adjustment , construct a binary group , and add To .

[0029] As a preferred solution of a method for predicting the blood-brain barrier permeability of compounds based on multimodal fusion according to the present invention, the specific steps of S2 are as follows: S21: For any compound , use the RDKit library to generate the normalized molecular three-dimensional structure feature P-3DF, and the specific steps are as follows: (1) The normalized Euclidean distance , between any two atoms in the compound is calculated as follows: (1) Among them, , represent any two atoms in the compound; is 's three-dimensional coordinates, is 's three-dimensional coordinates; represents the normalization operation; (2) The normalized initial three-dimensional feature vector of any atom in the compound, is calculated as follows: (2) Among them, is a One-Hot encoding vector, representing 's atom type; is the vector splicing operation; is 's three-dimensional coordinate vector, representing 's position in three-dimensional space; (3) The normalized initial edge feature vector of any edge in the compound, is calculated as follows: (3) Among them is a One-Hot encoding vector, representing the chemical bond between atom and atom , where , is a binary value, representing the th type of chemical bond, and ; represents the number of all chemical bond types; represents and 's Euclidean distance; (4) Normalized angular feature vector in the compound , The calculation process is as follows: (4) Among them, the vector represents the direction and distance from atom to atom ; the vector represents the direction and distance from to atom ; The function is used to calculate the angle corresponding to the cosine value of the angle between two vectors and ; is the normalized angular feature vector formed by side and .

[0030] S22: For any compound , use the RDKit library to extract information such as the physicochemical properties and topological structure of the compound to generate molecular descriptors; at the same time, construct the molecular graph representation corresponding to the compound; splice the SMILES string of each compound with the generated molecular descriptors to construct the molecular descriptor feature set MDF; at the same time, splice the SMILES expression with the corresponding molecular graph to construct the molecular graph feature set MGF.

[0031] S23: For any compound , extract the SMILES expressions of all compounds to construct the text feature set MNF.

[0032] S24: Normalize MNF, MDF, and MGF to obtain the normalized text feature P-MNF, the normalized molecular descriptor feature P-MDF, and the normalized molecular graph feature P-MGF; As a preferred solution of the method for predicting the blood-brain barrier permeability of compounds based on multimodal fusion described in the present invention, the specific steps of S3 are as follows: S31: Construct the molecular transformation module MolTransNet, and input the P-MGF of each compound into MolTransNet to generate the molecular graph feature vector of the compound; S32: Construct the molecular three-dimensional encoder module GEMM, and input the P-3DF of each compound into GEMM to generate the molecular three-dimensional feature vector of the compound; S33: Input the P-MNF of each compound into the Transformer module to generate the text feature vector of the compound ; S34: Convert the P-MDF of each compound into a molecular descriptor vector , and for the , 、 , of the compound, input them into the attention feature fusion module Afusion together for feature fusion to obtain the feature fusion vector of the compound ; According to the SMILES expression of the compound, blood-brain barrier permeability label and , construct a triple ; Combine the corresponding to all compounds in to form a multi-modal fusion feature vector dataset , where is , the number of triples in

[0033] As a preferred scheme of a method for predicting blood-brain barrier permeability of compounds based on multi-modal fusion according to the present invention, the specific steps of S31 are as follows: S311: For any atom in the compound, solve the one-hop neighbor node set and two-hop neighbor node set of ; Then calculate the atomic feature vector of , is calculated from the following iterative equations (5) to (14): (5) (6) (7) (8) (9) (10) (11) (12) (13) (14) Among them, is a one-hop neighbor node of ; is a one-hop neighbor node of , where represents the set of one-hop neighbor nodes of ; is a two-hop neighbor node of , where represents the set of two-hop neighbor nodes of ; , , are One-Hot encoded vectors, respectively representing , , the atomic types of ; , , are One-Hot encoded vectors, respectively representing , , the degrees of ; , , are One-Hot encoded vectors, respectively representing the number of hydrogen atoms connected to , , ; is an atomic feature vector with a dimension of ; is after rounds of iteration; is an atomic feature vector with a dimension of ; is an atomic feature vector with a dimension of ; is after rounds of iteration; is after rounds of iteration; , , are all weight matrices of size ; is the activation function ReLU; represents the dot product operation, used to calculate the inner product of two vectors or matrices; is The two-hop neighbor node; is the one-hop neighbor node of, where represents the set of one-hop neighbor nodes of; is the two-hop neighbor node of, where represents the set of two-hop neighbor nodes of; , , are One-Hot encoded vectors, representing respectively , , the atomic types of; , , are One-Hot encoded vectors, representing respectively , , the degrees of; , , are One-Hot encoded vectors, representing respectively the number of hydrogen atoms connected to , , ; is a vector with a dimension of ; is an atomic feature vector with a dimension of ; is a vector with a dimension of ; is the atomic feature vector obtained after is rounds of iteration; is the atomic feature vector obtained after is rounds of iteration; is the atomic feature vector obtained after is rounds of iteration; is an atomic feature vector with a dimension of ; is a One-Hot encoded vector, representing the degree of; is a One-Hot encoded vector, representing the number of hydrogen atoms connected to ; is the atomic feature vector obtained after is rounds of iteration; As a preferred solution of a method for predicting the blood-brain barrier permeability of compounds based on multimodal fusion according to the present invention, each compound atom generated in S311 of the one-hop neighbor set and the two-hop neighbor set are as follows: S3111: Find all atoms in the compound that are directly connected by chemical bonds to the atom , and form all the above atoms into ; S3112: For each atom in , find all atoms directly connected by chemical bonds to . If or is not directly connected to , then add to . .

[0034] S312: Based on , calculate the attention feature of the correlation between atoms inside the compound. The specific calculation formula is as follows: (15) where is a randomly initialized weight matrix with a size of ; represents any atom in the compound except ; the query matrix represents the mapping of the atomic feature vector of in the query vector space; the key matrix represents the mapping of in the key vector space; the value matrix represents the mapping of in the value space, used to store ; is an attention weight matrix with a dimension size of , representing the relationship between the atomic feature vectors of and ; S313: Use and to calculate the transformed feature . The specific calculation formula is as follows: (16) (17) where Represents a normalization operation; Represents a residual connection operation; Is an intermediate feature; Is a feed-forward neural network for Performing a non-linear transformation; S314: Average pooling is performed on All atoms of the compound To obtain the molecular graph feature vector of the compound , The calculation formula of (18) Where, Represents the number of atoms in the compound.

[0035] As a preferred solution of a method for predicting the blood-brain barrier permeability of a compound based on multimodal fusion according to the present invention, the specific steps of S32 are as follows: S321: For any atom In the compound, the normalized initial three-dimensional feature vector Of And the edge Of the normalized initial edge feature vector Are extracted from the P-3DF of the compound, where Is Any adjacent node of S322: Calculate The three-dimensional feature vector , Is calculated from the following iterative equation sets (19) to (22): (19) (20) (21) (22) Where, Represents The Euclidean distance between And Is a One-Hot encoded vector representing The chemical bond between And Is After R rounds of iteration to obtain the edge feature vector; Is The number of one-hop neighbor nodes; denote the set of one-hop neighbor nodes of; denote the edge and the normalized angular feature vector formed; is a one-hop neighbor node of; is a One-Hot encoded vector, denoting the atomic type of; is the three-dimensional coordinate vector of, denoting the position in three-dimensional space is for to perform rounds of iteration of the three-dimensional feature vector; is for to perform rounds of iteration of the edge feature vector; is for to perform rounds of iteration of the three-dimensional feature vector; is an angular weighting matrix of size ; is a node weighting matrix of size ; is a vector-level addition operation for adding the elements of two vectors at corresponding positions; is for to perform rounds of iteration of the three-dimensional feature vector; S323: Through operation, fuse the three-dimensional features of all atoms in the compound into the molecular three-dimensional vector , The calculation formula of is as follows: (23).

[0036] As a preferred solution of a method for predicting the blood-brain barrier permeability of a compound based on multimodal fusion according to the present invention, the specific steps of the S33 are as follows: S331: For any atom in the compound, use the n-gram model to construct the bag-of-words vector of, The specific construction formula is as follows: (24) wherein, denotes the The occurrence frequency of an n-gram in a SMILES expression, is the size of the n-grams vocabulary; S332: Generate the position encoding vector of , using to generate the position enhanced bag-of-words vector of , and the specific generation formula is as follows: (25) (26) (27) (28) Among them, represents the index in the SMILES expression; is the dimension index. When is even, use to generate ; when is odd, use to generate ; S333: Use to calculate the text attention feature , and the specific calculation formula is as follows: (29) Among them, is a text weight matrix randomly initialized and generated, with a size of ; represents any atom in the compound except ; the text query matrix represents the mapping of the position enhanced bag-of-words vector of in the query vector space; the text key matrix represents the mapping of in the key vector space; the text value matrix represents and also is used to calculate the text feature correlation between atom and atom ; is a text attention weight matrix with a dimension size of , representing and The relationship between the position-enhanced bag-of-words vectors; The text attention features representing the correlation between atoms within a compound; S334: Use and to calculate the text transformation features , and the specific calculation formula is as follows: (30) (31) where, represents the normalization operation; is the text intermediate feature; is a feed-forward neural network for to perform a non-linear transformation; S335: Perform on all atoms of the compound operation to obtain the molecular text feature vector of the compound , The calculation formula of (32).

[0037] As a preferred scheme of a method for predicting the blood-brain barrier permeability of a compound based on multi-modal fusion according to the present invention, the specific steps of the S34 are as follows: S341: Combine , , , to obtain a vector that combines molecular graph features, molecular three-dimensional features, molecular text features, and molecular descriptor features, The specific calculation formula of (33) S342: Use to calculate the molecular query vector , the molecular bond vector , and the molecular value vector , and the specific formulas are as follows: (34) (35) (36) where, is the modal query weight matrix of size ; represents the mapping in the query vector space; is the modal key weight matrix of size ; represents the mapping in the key vector space; is the modal key weight matrix of size ; represents the mapping in the value vector space; and are used to calculate the similarity between different features in ; S343: Use and to calculate the self-attention weight , and the specific generation formula is as follows: (37) where is the self-attention score; represents the transpose operation of the matrix; is the scaling factor used to prevent from having too large a numerical value; is the normalization function used to convert into a probability distribution; represents the self-attention weight; S344: Use and to calculate the feature fusion vector of the compound, and the specific calculation formula is as follows: (38).

[0038] As a preferred solution of a method for predicting the blood-brain barrier permeability of a compound based on multimodal fusion according to the present invention, the specific steps of S4 are as follows: S41: Divide into a training set and a test set , where is used to train the KAN model, and is used to evaluate the performance of the KAN model; S42: Input any into the KAN model for training to obtain the predicted value of the blood-brain barrier permeability label of the ​​​, The specific calculation formula is as follows: (39) (40) (41) Among them, is the th B-spline basis function; is the th B-spline superposition function, which is obtained by summing ; is the bias term of the th KAN mapping layer; is a linear function, representing the th KAN mapping layer; is a random initial weight, representing the initial weight of the th KAN mapping layer; is the weight obtained after is iterated times, representing the weight coefficient of the th KAN mapping layer; is 's quantity; is the number of triples in; 's value range is {BBB-, BBB+}; S43: Configure the cross-entropy loss function , The specific formula is as follows: (42) Among them, represents the cross-loss function of the th compound; S44: Use the gradient descent method to iterate : Set the change amount of the loss function convergence threshold to ; Calculate in each iteration. When or is satisfied, stop the iteration and output , and obtain the trained KAN model; is calculated from the following iterative equations (43)~(44): (43) (44) Among them, is the fixed learning rate;​ is the weight obtained after rounds of iteration; represents the blood-brain barrier permeability label of the th compound; represents the predicted value of the blood-brain barrier permeability label of the th compound; represents the gradient of ; is a value that can be dynamically adjusted and is set according to the size of different-scale data.

[0039] S45: Use to evaluate the performance of the trained KAN model.

[0040] Compared with the prior art: (1) The present invention designs a data preprocessing method. First, data balancing operation is performed on through resampling technology, solving the problem of data imbalance; subsequently, SMILES augmentation operation is adopted to effectively solve the problems of vanishing gradient and underfitting that occur when the amount of data is insufficient, further improving the training effect of the model.

[0041] (2) The present invention designs a feature fusion network BBBNet to extract four different features: molecular text feature, molecular graph feature, molecular three-dimensional feature, and molecular descriptor feature, solving the problem that molecular information is not fully utilized in traditional methods and improving the comprehensive characterization of complex properties of compounds.

[0042] (3) The present invention also designs a feature fusion module Afusion based on the attention mechanism; this module can efficiently fuse four different modalities of features and enable the model to learn the weights of each modality through the attention mechanism, thus solving the problem of deficiencies in current multi-modal feature fusion and improving the performance of the model in the blood-brain barrier permeability prediction task. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is the overall flowchart of the solution of the present invention; Figure 2 is the flowchart of generating SMILES expression features; Figure 3 is the structural diagram of the feature fusion network BBBNet; Figure 4 is the flowchart of predicting blood-brain barrier permeability labels. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] ​To make the above objects, features, and advantages of the present invention more obvious and understandable, the following provides a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0045] Embodiment This application provides a method that can effectively solve the above-mentioned problems. Next, multiple embodiments will be combined to elaborate in detail how to implement the method for predicting the blood-brain barrier permeability of compounds based on multimodal fusion.

[0046] Figure 1 The flowchart of the method for predicting the blood-brain barrier permeability of compounds based on multimodal fusion is shown, including: S1: Represent each collected compound as a binary tuple , where is the SMILES expression of the compound , and is 's blood-brain barrier permeability label; perform data preprocessing on all compounds to obtain an enhanced data set ; In the embodiments of this application, in order to implement the method for predicting the blood-brain barrier permeability of compounds based on multimodal fusion, a data preprocessing process is constructed, and the execution process of the data preprocessing is as follows: S11: Use resampling technology to perform data balancing operations on , and the specific steps are as follows: For any minority-class sample binary tuple , use the ADASYN method to generate a new SMILES expression for , and then construct the binary tuple to represent the new minority-class sample and add it to ; The specific process of using the ADASYN method to generate new minority-class SMILES expressions is as follows: S111: First, calculate the importance of the minority-class sample , and then calculate the number of new samples according to , and the formula is as follows: where represents The number of the nearest majority class samples, is the maximum number of neighbors in the minority class samples, is the number of generated minority samples; S112: For each minority class sample , use its neighboring majority class samples to generate new minority class samples , and the generation process is as follows: Among them, is a random number in the interval, used to generate new minority class samples between and its neighboring majority class samples. This process is continuously iterated until the required number of minority class samples is generated.

[0047] Exemplarily, the present invention uses the B3DB dataset of McMaster University in Canada. This dataset has a total of 7807 samples, including 4956 BBB+ samples and 2851 BBB- samples respectively. The minority class samples in the dataset refer to BBB- samples. Assume a certain BBB- sample , and use its neighboring BBB+ samples ; assume the number of the nearest BBB+ samples , the maximum number of neighbors , and the total number of BBB- samples generated by is set to ; first calculate the importance of the BBB- sample ; obtain the number of generated BBB- samples ; is a random number with a value in the interval, and takes in the current example; new BBB- samples are generated by . Iterate continuously according to the above method until the th BBB- compound sample is generated. At this time, the total number of BBB- compound samples increases to 4900.

[0048] S12: Apply the SMILES enhancement operation to each compound in BBBD: (1) For any structurally asymmetric compound , rotate around the central atom and along the axis perpendicular to the molecular plane to obtain the rotated SMILES expression , construct the binary group , and Add , where the rotation operation is represented as follows: Among them, represents the function for performing the rotation operation on , is a randomly selected rotation angle; Exemplarily, the SMILES expression of 1,2-dichloroethylene is , which is a compound with an asymmetric structure. Applying the SMILES enhancement operation to 1,2-dichloroethylene specifically involves: selecting the axis perpendicular to the plane of 1,2-dichloroethylene as the rotation axis and choosing the rotation angle to perform the rotation operation. After rotation, the obtained SMILES expression after rotation is .

[0049] (2) For a compound containing more than two closed rings , all the closed rings in the compound are adjusted as follows: randomly rearrange the connection order of all atoms on the closed ring, and generate the SMILES expression according to the adjusted atom connection order, construct the binary tuple , and add to ; Exemplarily, the SMILES expression of naphthalene is , which is a compound with two closed ring structures. Applying the SMILES enhancement operation to naphthalene specifically involves adjusting the two closed rings as follows: randomly rearrange the connection order of all atoms on the closed ring. The new SMILES expression obtained after adjustment is .

[0050] (3) For a compound containing a double bond , randomly select a double bond and exchange the stereochemical label of the double bond. If the original stereochemical label is cis, it will be adjusted to trans; if the original stereochemical label is trans, it will be adjusted to cis; generate the SMILES expression after adjustment, construct the binary tuple , and add to

[0051] Exemplarily, the SMILES expression of an alkene is , which is a compound containing a double bond. Applying the SMILES enhancement operation to the alkene specifically involves: randomly selecting a double bond and exchanging the stereochemical label of the double bond; the first represents that the original stereochemical label is cis; the second Indicates that the stereochemical label is trans; swapping the stereochemical label of the double bond generates a new SMILES expression as . Cis and trans reflect the arrangement of substituents on both sides of the double bond in three-dimensional space.

[0052] S2: For any compound , generate the molecular three-dimensional structure feature 3DF, the molecular text feature MNF, the molecular descriptor feature MDF, and the molecular graph feature MGF; perform normalization on 3DF, MNF, MDF, and MGF to obtain the normalized molecular three-dimensional structure feature P-3DF, the normalized text feature P-MNF, the normalized molecular descriptor feature P-MDF, and the normalized molecular graph feature P-MGF; In the embodiments of this application, in order to implement a method for predicting the blood-brain barrier permeability of compounds based on multimodal fusion, a flowchart for generating SMILES expression features is constructed; Figure 2 Describes each step of generating SMILES expression features, mainly including: generating the molecular three-dimensional structure feature 3DF, the molecular text feature MNF, the molecular descriptor feature MDF, and the molecular graph feature MGF; and performing normalization on 3DF, MNF, MDF, and MGF. The specific steps of feature generation are as follows: S21: For any compound , use the RDKit library to generate the molecular three-dimensional structure feature, and then perform normalization to obtain P-3DF. The specific steps are as follows: (1) The normalized Euclidean distance between any two atoms in the compound , The calculation formula is as follows: (1) Among them, , represent any two atoms in the compound; is 's three-dimensional coordinates, is 's three-dimensional coordinates; represents the normalization operation; (2) The normalized initial three-dimensional feature vector of any atom in the compound, The calculation formula is as follows: (2) Among them, is a One-Hot encoded vector representing 's atom type; is the vector concatenation operation; is The three-dimensional coordinate vector represents the position in three-dimensional space; (3) Any edge in the compound The normalized initial edge feature vector , The calculation formula is as follows: (3) where is a One-Hot encoding vector representing the atom and the atom the chemical bond between them, where , is a binary value representing the type of chemical bond, and ; represents the total number of all chemical bond types; represents and the Euclidean distance between them; (4) The normalized angle feature vector in the compound , The calculation process is as follows: (4) where the vector represents the direction and distance from the atom to the atom , and the vector represents from to the atom the direction and distance; The function is used to calculate the angle corresponding to the cosine value of the angle between the two vectors and ; is the normalized angle feature vector formed by the edge and .

[0053] S22: For any compound , use the RDKit library to extract information such as the physicochemical properties and topological structure of the compound to generate molecular descriptors; at the same time, construct the molecular graph representation corresponding to the compound; concatenate the SMILES string of each compound with the generated molecular descriptors to construct the molecular descriptor feature set MDF; at the same time, concatenate the SMILES expression with the corresponding molecular graph to construct the molecular graph feature set MGF.

[0054] S23: For any compound , extract the SMILES expressions of all compounds to construct the text feature set MNF.

[0055] S24: Normalize MNF, MDF, and MGF to obtain the normalized text feature P-MNF, the normalized molecular descriptor feature P-MDF, and the normalized molecular graph feature P-MGF; Exemplarily, use the normalization method to perform standard scaling on the molecular descriptor feature set MDF. The specific formula is shown as follows: where is a certain sample in the molecular descriptor feature set MDF, represents the mean of the molecular descriptor feature set, represents the standard deviation of the molecular descriptor feature set, is the value after normalization.

[0056] S3: Construct a feature fusion network BBBNet for constructing a multi-modal fusion feature vector data set ; In the embodiments of the present application, in order to implement a multi-modal fusion-based method for predicting the blood-brain barrier permeability of compounds, a feature fusion network BBBNet is designed. The execution process of this model is as Figure 3 shown. The BBBNet construction module performs feature extraction and designs an attention feature fusion module Afusion for feature fusion, and finally obtains a multi-modal fusion feature vector data set ; The feature fusion network BBBNet includes four modules: the MolTransNet module, the GEMM module, the Transfomer module, and the attention feature fusion module Afusion. The MolTransNet module is used to extract the molecular graph feature vector of the compound ; the GEMM module is used to extract the three-dimensional molecular feature vector of the compound ; the Transformer module is used to extract the text feature vector of the compound ; the attention feature fusion module Afusion is used to generate the feature fusion vector of the compound ; S31: Construct a molecular transformation module MolTransNet, input the P-MGF of each compound into the MolTransNet module to generate a molecular graph feature vector ; The construction process of the MolTransNet module is as follows: Using P-MGF as the input of the MolTransNet module, initially extract the initial feature vectors of each atom from the P-MGF, and then gradually update the atomic feature vectors of each atom through multiple rounds of information passing and neighbor feature aggregation; this module further processes these features using a Transformer, and finally obtains the molecular graph feature vector of the compound ; Finally obtain the molecular graph feature vector of the compound as the output

[0057] S311: For any atom in the compound , solve for the set of one-hop neighbor nodes and the set of two-hop neighbor nodes ; then calculate 's atomic feature vector , which is calculated from the following iterative equations (5) - (14): (5) (6) (7) (8) (9) (10) (11) (12) (13) (14) where is 's one-hop neighbor node; is 's one-hop neighbor node, where represents 's set of one-hop neighbor nodes; is 's two-hop neighbor node, where represents 's set of two-hop neighbor nodes; , , is a One-Hot encoded vector, respectively representing , , 's atomic type; , , is a One-Hot encoded vector, respectively representing , , 's degree; , , is a One-Hot encoded vector, respectively representing the number of hydrogen atoms connected to , , ; is an atomic feature vector with a dimension of ; is the atomic feature vector obtained after has been rounds of iteration; is an atomic feature vector with a dimension of ; is an atomic feature vector with a dimension of ; is the atomic feature vector obtained after has been rounds of iteration; is the atomic feature vector obtained after has been rounds of iteration; , , are all weight matrices of size ; is the activation function ReLU; represents the dot product operation, used to calculate the inner product of two vectors or matrices; is 's second-hop neighbor node; is 's first-hop neighbor node, where represents 's first-hop neighbor node set; is 's second-hop neighbor node, where represents 's second-hop neighbor node set; , , are One-Hot encoded vectors, respectively representing , , The atomic type; , , are One - Hot encoded vectors, respectively representing , , 's degree; , , are One - Hot encoded vectors, respectively representing the number of hydrogen atoms connected to , , ; is a vector with a dimension of ; is an atomic feature vector with a dimension of ; is a vector with a dimension of ; is the atomic feature vector obtained after undergoes rounds of iteration; is the atomic feature vector obtained after undergoes rounds of iteration; is the atomic feature vector obtained after undergoes rounds of iteration; is an atomic feature vector with a dimension of ; is a One - Hot encoded vector representing 's degree; is a One - Hot encoded vector representing the number of hydrogen atoms connected to ; is the atomic feature vector obtained after undergoes rounds of iteration; Exemplarily, has a size of , is the number of atomic nodes in each compound; contains 66 - bit information including: 44 bits representing the high - dimensional vector generated by mapping the atomic type through One - Hot encoding, 13 bits representing the atomic degree, and 10 bits representing the number of hydrogen atoms connected to the atom; Calculating requires rounds of update. The feature vectors of all atoms in the compound are aggregated with the feature vectors of all their one - hop neighbors and two - hop neighbors to form a complete round of update; the specific steps are as follows: First, calculate for each atom in the compound , , ; Next, perform feature aggregation: Aggregate with and to form ; At the same time, is aggregated with and to form ; Similarly, aggregate with and to form ; When all atoms in a compound complete one round of feature aggregation, one round of update ends; Then, repeat the above feature aggregation process times, and each round of update is based on the feature vector of the previous round to obtain .

[0058] The construction process of the one-hop neighbor set and the two-hop neighbor set is as follows: S3111: Find all atoms in the compound that are directly connected to the atom by chemical bonds, and form all the above atoms into ; S3112: For each atom in , find all atoms that are directly connected to by chemical bonds. If or is not directly connected to , then add to .

[0059] Exemplarily, assume the SMILES expression of the compound is CC(=O)O, and the compound contains the following atoms: the first carbon atom , the second carbon atom , the oxygen atom in the carbonyl , and the oxygen atom in the hydroxyl . These atoms are connected by chemical bonds to form an undirected graph, where each atom is a node and the chemical bond is an edge in the graph. , , , , , , , The specific calculation process is as follows: First, calculate the set of one-hop neighbor nodes for each atomic node. is directly connected to , so . is directly connected to , and , so . For , it is directly connected to , so . For , it is also directly connected to , so .

[0060] Next, calculate the set of two-hop neighbor nodes for each atom, that is, the set of other atoms that can reach this atom through two chemical bonds. For , its one-hop neighbor node set is ; among the one-hop neighbors of , after removing the atoms that already belong to or are directly connected to or , we get ; for , its one-hop neighbor node set is ; among the directly connected atoms of each neighbor node, after removing the atoms that already belong to or are directly connected to , the result is an empty set, so . For , its one-hop neighbor is ; among the one-hop neighbor set of , after removing the atoms that already belong to or are directly connected to , we get Similarly, for , its one-hop neighbor node set is , and after calculation, we get .

[0061] Finally, the neighbor node set for each atom is obtained as follows: ; ; ; ; ; ; ; .

[0062] S312: Based on , calculate the attention features of the correlation between atoms inside the compound , the specific calculation formula is as follows: (15) Among them, is a randomly initialized weight matrix with a size of ; represents any atom in the compound except ; the query matrix represents the mapping of the atomic feature vector of in the query vector space; the key matrix represents the mapping of in the key vector space; the value matrix represents the mapping of in the value space, used to store ; is an attention weight matrix with a dimension size of , representing the relationship between the atomic feature vectors of and ; S313: Use and to calculate the transformed feature , the specific calculation formula is as follows: (16) (17) Among them, represents the normalization operation; represents the residual connection operation; is the intermediate feature; is a feed-forward neural network used to perform a non-linear transformation on ; S314: Perform average pooling on all atoms of the compound to obtain the molecular graph feature vector of the compound, , The calculation formula of (18) Among them, represents the number of atoms in the compound.

[0063] S32: Construct a molecular three-dimensional encoder module GEMM, input the P-3DF of each compound into GEMM, and generate the molecular three-dimensional feature vector of the compound; In an embodiment of the present application, in order to implement a method for predicting the blood-brain barrier permeability of multimodal fusion compounds, a molecular three-dimensional encoder module GEMM is designed to generate a molecular three-dimensional feature vector of a compound. The specific process of constructing GEMM is as follows: Taking P-3DF as the input of the molecular three-dimensional encoder module GEMM, first extract the three-dimensional feature vector and edge feature vector of each atom from P-3DF, and then gradually optimize the three-dimensional feature vector of each atom through multiple rounds of information transfer and neighbor feature aggregation. Finally, aggregate the three-dimensional features of all atoms in the compound using a pooling operation to obtain a molecular three-dimensional vector. ; S321: For any atom in the compound , extract from the P-3DF of the compound 's normalized initial three-dimensional feature vector and the edge 's normalized initial edge feature vector , where is 's any adjacent node; S322: Calculate 's three-dimensional feature vector , is calculated from the following iterative equations (19) - (22): (19) (20) (21) (22) Among them, represents and 's Euclidean distance; is a One-Hot encoded vector, representing and 's chemical bond; is the edge feature vector obtained after is rounds of iteration; is 's number of one-hop neighbor nodes; represents 's one-hop neighbor node set; represents the edge and 's formed normalized angular feature vector; is a one-hop neighbor node of is a One-Hot encoded vector representing the atomic type of is a three-dimensional coordinate vector of indicating the position of in three-dimensional space is the three-dimensional feature vector after is the edge feature vector after rounds of iteration is the three-dimensional feature vector after rounds of iteration is an angle weighting matrix of size is a node weighting matrix of size is a vector-level addition operation that adds the elements of two vectors at corresponding positions is the three-dimensional feature vector after rounds of iteration One round of iteration consists of two steps as follows (1) Use to calculate ; After the edge feature vectors of all edges in the compound are updated, enter the next stage (2) Use and and also to calculate ; After the three-dimensional feature vectors of all atoms in the compound are updated, one round of iteration is completed (3) Repeat (1) and (2), and after rounds of repeated iteration, obtain .

[0064] S323: Through operation, fuse the three-dimensional features of all atoms in the compound into the molecular three-dimensional vector , and the calculation formula of (23).

[0065] S33: Input the P-MNF of each compound into the Transformer to generate the text feature vector of the compound;​​ The specific steps of the P-MNF input Transformer for each compound are as follows: S331: For any atom in the compound , construct the bag-of-words vector using the n-gram model. The specific construction formula is as follows: (24) where represents the occurrence frequency of the th n-gram in the SMILES expression, and is the size of the n-grams vocabulary; S332: Generate the positional encoding vector , and use to generate the position-enhanced bag-of-words vector . The specific generation formula is as follows: (25) (26) (27) (28) where represents the index in the SMILES expression; is the dimensional index. When is even, use to generate ; when is odd, use to generate ; S333: Use to calculate the text attention feature . The specific calculation formula is as follows: (29) where is a randomly initialized text weight matrix with a size of ; represents any atom in the compound except ; the text query matrix represents the mapping of the position-enhanced bag-of-words vector in the query vector space; the text key matrix represents the Mapping in the key vector space; text value matrix representation Mapping in the value space for storage ; and also for calculating the text feature correlation between atoms and atoms ; is a text attention weight matrix with a dimension size of representing the relationship between and the position-enhanced bag-of-words vectors; Text attention features representing the correlation between atoms inside the compound; S334: Using and to calculate the text transformation features , and the specific calculation formula is as follows: (30) (31) wherein, represents the normalization operation; is the text intermediate feature; is a feed-forward neural network for performing a non-linear transformation on ; S335: Performing operation on all atoms of the compound to obtain the molecular text feature vector of the compound, , and the calculation formula of (32).

[0066] S34: Converting the P-MDF of each compound into a molecular descriptor vector , inputting the , 、 , of the compound into the attention feature fusion module Afusion together for feature fusion to obtain the feature fusion vector of the compound; constructing a triple blood-brain barrier permeability labeland based on the SMILES expression of the compound; corresponding to all compounds in constitute a multi-modal fusion feature vector dataset , where is the number of triples in

[0067] Exemplarily, assume that P-MDF contains the following features: molecular weight (MW), number of hydrogen bond donors, number of hydrogen bond acceptors, LogP, and number of rings. The specific properties of this compound are: the molecular weight is , the number of hydrogen bond donors is , the number of hydrogen bond acceptors is , LogP is , and the number of rings is . Convert the P-MDF of this compound into .

[0068] The present invention designs an attention feature fusion module Afusion for feature fusion; Afusion learns the relative importance of different features through an attention mechanism, thereby effectively fusing the information of each modality and improving the accuracy of blood-brain barrier permeability prediction; the construction process of the attention feature fusion module Afusion is as follows: S341: Combine , , , to obtain a vector that fuses molecular graph features, molecular three-dimensional features, molecular text features, and molecular descriptor features , The specific calculation formula of is as follows: (33) S342: Use to calculate the molecular query vector , molecular bond vector , molecular value vector , and the specific formulas are as follows: (34) (35) (36) where is a modality query weight matrix of size ; represents 's mapping in the query vector space; is a matrix of size The modal key weight matrix; denotes the mapping in the key vector space; is of size the modal key weight matrix; denotes the mapping in the value vector space; and is used to calculate the similarity between different features in Exemplarily, , , the dimension of is set to , and , the value of each matrix element follows a normal distribution with a mean of and a standard deviation of .

[0069] S343: Use and to calculate the self-attention weight , the specific generation formula of is as follows: (37) where is the self-attention score; denotes the transpose operation of the matrix; is the scaling factor, used to prevent the value from being too large; is the normalization function, used to convert into a probability distribution; denotes the self-attention weight; Exemplarily, set the scaling factor to to prevent gradient explosion or numerical instability caused by the value being too large.

[0070] S344: Use and to calculate the feature fusion vector of the compound, and the specific calculation formula is as follows: (38).

[0071] S4: Divide into a training set and a test set, and input the training set into the Kolmogorov - Arnold model KAN for model training; after training is completed, use the test set to evaluate the performance of the trained KAN model; Predicting blood-brain barrier permeability labels using the Kolmogorov-Arnold model KAN; the prediction process of blood-brain barrier permeability labels is as follows Figure 4 as shown. The specific processes of KAN for model training and performance evaluation are as follows: S41: Divide into a training set and a test set , where is used to train the KAN model, is used to evaluate the performance of the KAN model; Exemplarily, the number of samples containing BBB+ and BBB- are respectively and ; Divide into , in the ratio of 70%:30%. There are samples in the training set , among which the number of triples , and contains 7,064 BBB+ samples and BBB- samples; the test set contains samples, among which it contains BBB+ samples and BBB- samples.

[0072] S42: Input any into the KAN model for training to obtain the predicted value of the blood-brain barrier permeability label of the th compound. The specific calculation formula is as follows: (39) (40) (41) where, is the th B-spline basis function; is the th B-spline superposition function, obtained by summing ; is the bias term of the th KAN mapping layer; is a linear function, representing the th KAN mapping layer; is a random initial weight, representing The initial weights of the KAN mapping layer; is after rounds of iteration, representing the weight coefficients of the th KAN mapping layer; is quantity; is the number of triples in; ranges from {BBB-, BBB+}; S43: Configure the cross-entropy loss function , and the specific formula is as follows: (42) where, represents the cross-loss function of the th compound; S44: Use the gradient descent method to iterate: Set the change in the loss function convergence threshold to ; Calculate in each iteration. When or is satisfied, stop the iteration and output , and obtain the trained KAN model; is calculated from the following iterative equations (43)~(44): (43) (44) where, is the fixed learning rate; is after rounds of iteration; represents the blood-brain barrier permeability label of the th compound; represents the predicted value of the blood-brain barrier permeability label of the th compound; represents the gradient of is a value that can be dynamically adjusted and is set according to the size of different-scale data.

[0073] S45: Use to evaluate the performance of the trained KAN model.

[0074] Exemplarily, the initial weights of the KAN model are set to 0.5; Subsequently, Input it into the KAN model for training; to ensure the convergence of the model, set a fixed learning rate and set the change amount of the loss function with a convergence threshold of 0.01. During the training process, the gradient descent method is used to iteratively update the weight parameters of the KAN mapping layer. In each iteration, calculate ; when is satisfied or the maximum number of iteration rounds is reached, stop the iteration and output the finally optimized .

[0075] S5: Use the trained KAN model to predict the blood-brain barrier permeability of compounds.

[0076] Take any compound for actual application as the input of the trained KAN model, and output the predicted value of the blood-brain barrier permeability label of the compound .

[0077] Based on the above, the present invention designs a method for predicting the blood-brain barrier permeability of compounds based on multi-modal fusion, aiming to improve the accuracy and generalization ability of the model for predicting the blood-brain barrier permeability of compounds. Regarding the problem of feature extraction of compounds, the present invention further designs a feature fusion network BBBNet, extracts four features through the network, and uses the Afusion module for feature fusion to construct a multi-modal fusion feature vector dataset , solving the problem that molecules are difficult to be fully characterized in traditional methods. The present invention uses the KAN model to perform the blood-brain barrier permeability prediction task and obtains the blood-brain barrier permeability prediction label. This method effectively extracts and fuses various modal features, significantly improving the performance of the blood-brain barrier permeability prediction task.

[0078] Although the present invention has been described above with reference to the embodiments, various improvements can be made to it and components therein can be replaced with equivalents without departing from the scope of the present invention. In particular, as long as there is no structural conflict, the various features in the embodiments disclosed in the present invention can be combined with each other in any way, and the exhaustive description of these combinations is omitted in this specification only for the sake of saving space and resources. Therefore, the present invention is not limited to the specific embodiments disclosed in the text, but includes all technical solutions falling within the scope of the claims.

Claims

1. A method for predicting compound blood-brain barrier permeability based on multimodal fusion, characterized in that Considering the impact of the imbalanced number of class samples on the prediction results, and using the feature fusion network to improve the accuracy of predicting blood-brain barrier permeability, the method includes the following steps: S1: Represent each collected compound as a binary group ,in For the compound The SMILES expression, for Blood-brain barrier permeability labels; data preprocessing was performed on all collected compounds to obtain an enhanced data set ; S2: For any compound , generate molecular three-dimensional structure feature 3DF, molecular text feature MNF, molecular descriptor feature MDF, and molecular graph feature MGF; normalize 3DF, MNF, MDF, and MGF to obtain normalized molecular three-dimensional structure feature P-3DF, normalized text feature P-MNF, normalized molecular descriptor feature P-MDF, and normalized molecular graph feature P-MGF; S3: Construct a feature fusion network BBBNet to construct a multimodal fusion feature vector dataset ; S4: The training set is divided into a training set and a test set, and the training set is input into the Kolmogorov-Arnold model (KAN) for model training; after the training is completed, the performance of the trained KAN model is evaluated using the test set; S5: Use the trained KAN model to predict the blood-brain barrier permeability of the compound.

2. The method for predicting compound blood-brain barrier permeability based on multimodal fusion according to claim 1, characterized in that: The data preprocessing of S1 includes: S11: Using resampling technology Perform data balancing operation. The specific steps are as follows: For any minority class sample binary ,right Generate new SMILES expressions using the ADASYN method , then construct the tuple Represents new minority class samples and adds them to middle; S12: Yes Apply the SMILES enhancement operation to each compound in: (1) For any asymmetric structural compound ,Will By rotating around the central atom and along an axis perpendicular to the molecular plane, the rotated SMILES expression is obtained. , construct a tuple , and will join in , where the rotation operation is expressed as follows: in, Express Function to perform rotation operation, is a randomly selected rotation angle; (2) Compounds containing more than two closed rings , adjust all closed rings in the compound as follows: rearrange the connection order of all atoms on the closed ring in a random order, and generate a SMILES expression based on the adjusted atomic connection order , construct a tuple , and will join in ; (3) Compounds containing double bonds , randomly select a double bond, exchange the stereo label of the double bond, if the original stereo label is cis, it will be adjusted to trans; if the original stereo label is trans, it will be adjusted to cis; the adjusted SMILES expression is generated , construct a tuple , and will join in .

3. The method for predicting compound blood-brain barrier permeability based on multimodal fusion according to claim 1, characterized in that: The normalized molecular three-dimensional structure feature P-3DF in S2 includes the following information: (1) The normalized Euclidean distance between any two atoms in a compound , The calculation formula is as follows: (1) in, , Represents any two atoms in a compound; yes The three-dimensional coordinates of yes The three-dimensional coordinates of Represents a normalization operation; (2) Any atom in a compound The normalized initial three-dimensional eigenvector of , The calculation formula is as follows: (2) in, Is a One-Hot encoding vector, indicating The atomic type of It is a vector concatenation operation; yes The three-dimensional coordinate vector of Position in three-dimensional space; (3) Any edge in a compound The normalized initial edge eigenvector of , The calculation formula is as follows: (3) in is a One-Hot encoding vector representing the atom and atoms The chemical bonds between , Is a binary value, indicating the Chemical bonds, and ; Represents the number of all chemical bond types; express and The Euclidean distance between (4) Normalized angle eigenvector in compounds , The calculation process is as follows: (4) Among them, the vector Represents from atoms To Atom The direction and distance of a vector Indicates from To Atom direction and distance; Function is used to calculate two vectors and The angle corresponding to the cosine of the angle; It is the edge and The normalized angle feature vector formed.

4. The method for predicting compound blood-brain barrier permeability based on multimodal fusion according to claim 1, characterized in that: The specific steps of constructing the feature fusion network BBBNet in S3 include: S31: Construct the molecular transformation module MolTransNet, input the P-MGF of each compound into MolTransNet, and generate the molecular graph feature vector of the compound ; S32: Construct a molecular three-dimensional encoder module GEMM, input the P-3DF of each compound into GEMM, and generate a molecular three-dimensional feature vector of the compound ; S33: Input the P-MNF of each compound into the Transformer module to generate the text feature vector of the compound ; S34: Convert the P-MDF of each compound into a molecular descriptor vector , the compound , 、 , The components are input together into the attention feature fusion module Afusion for feature fusion to obtain the feature fusion vector of the compound. ; According to the SMILES expression of the compound , Blood-brain barrier permeability label and , construct a triple ;Will All compounds in Construct a multimodal fusion feature vector dataset ,in for The number of triplets in .

5. The method for constructing a BBBNet feature fusion network according to claim 4, characterized in that: The specific process of constructing the molecular transformation module MolTransNet in S31 includes: S311: For any atom in a compound , solve The set of one-hop neighbor nodes and the set of two-hop neighbor nodes ; then calculate The atomic eigenvectors of , It is calculated by the following iterative equations (5)~(14): (5) (6) (7) (8) (9) (10) (11) (12) (13) (14) in, yes One-hop neighbor node; yes One-hop neighbor node of express The set of one-hop neighbor nodes; yes The two-hop neighbor nodes of express The set of two-hop neighbor nodes of , , is a One-Hot encoding vector, representing , , The atomic type of , , is a One-Hot encoding vector, representing , , The degree of , , is a One-Hot encoding vector, representing , , The number of attached hydrogen atoms; is a dimension of The atomic eigenvectors of ; Yes conduct The atomic feature vector obtained after rounds of iterations; is a dimension of The atomic eigenvectors of ; is a dimension of The atomic eigenvectors of ; Yes conduct The atomic feature vector obtained after rounds of iterations; Yes conduct The atomic feature vector obtained after rounds of iterations; , , The size is The weight matrix of is the activation function ReLU; Represents the dot multiplication operation, which is used to calculate the inner product of two vectors or matrices; yes The two-hop neighbor node of yes One-hop neighbor node of express The set of one-hop neighbor nodes; yes The two-hop neighbor nodes of express The set of two-hop neighbor nodes of , , is a One-Hot encoding vector, representing , , The atomic type of , , is a One-Hot encoding vector, representing , , The degree of , , is a One-Hot encoding vector, representing , , The number of attached hydrogen atoms; is a dimension of A vector of is a dimension of The atomic eigenvectors of ; is a dimension of A vector of Yes conduct The atomic feature vector obtained after rounds of iterations; Yes conduct The atomic feature vector obtained after rounds of iterations; Yes conduct The atomic feature vector obtained after rounds of iterations; is a dimension of The atomic eigenvectors of ; Is a One-Hot encoding vector, indicating The degree of Is a One-Hot encoding vector, representing The number of attached hydrogen atoms; Yes conduct The atomic feature vector obtained after rounds of iterations; S312: Based on , which calculates the attention features of the correlations between atoms within the compound , the specific calculation formula is as follows: (15) in, is a randomly initialized weight matrix of size ; Indicates that the compound Any atom other than express The mapping of the atomic feature vectors of in the query vector space; the bond matrix express Mapping in the space of key vectors; value matrices express A mapping in the value space used to store ; The dimension size is The attention weight matrix represents and The relationship between the atomic eigenvectors of ; S313: Use and Calculate transformation features , the specific calculation formula is as follows: (16) (17) in, Represents a normalization operation; Represents the residual connection operation; It is an intermediate feature; is a feed-forward neural network used to Perform nonlinear transformations; S314: All the atoms in the compound Perform average pooling Operation, get the molecular graph feature vector of the compound , The calculation formula is as follows: (18) in, Indicates the number of atoms in a compound.

6. The molecular transformation module MolTransNet according to claim 4, characterized in that: Any atom in the compound generated in S311 The set of one-hop neighbor nodes and the set of two-hop neighbor nodes The specific process includes: S3111: Find all the atoms in the compound Atoms directly connected by chemical bonds make up all the above atoms ; S3112: For Each atom in , find the All atoms connected ,if Or not with Direct connection, then join in .

7. The method for constructing a BBBNet feature fusion network according to claim 4, characterized in that: The specific process of constructing the molecular three-dimensional encoder module GEMM in S32 includes: S321: For any atom in a compound , extracted from the P-3DF of the compound The normalized initial three-dimensional eigenvector of and edge The normalized initial edge eigenvector of ,in yes Any adjacent point of ; S322: Calculation The three-dimensional feature vector of , It is calculated by the following iterative equations (19)~(22): (19) (20) (21) (22) in, express and The Euclidean distance between Is a One-Hot encoding vector, indicating and The chemical bonds between Yes conduct The edge feature vector obtained after rounds of iterations; yes The number of one-hop neighbor nodes; express The set of one-hop neighbor nodes; Represents edge and The normalized angle feature vector formed; yes One-hop neighbor node; Is a One-Hot encoding vector, indicating The atomic type of yes The three-dimensional coordinate vector of Position in three-dimensional space Yes conduct The three-dimensional feature vector after round iteration; Yes conduct The edge feature vector after round iteration; Yes conduct The three-dimensional feature vector after round iteration; Is the size of The angle weighting matrix of Is the size of The node weight matrix of It is a vector-level addition operation, which is used to add the elements of two vectors at corresponding positions; Yes conduct The three-dimensional feature vector after round iteration; S323: Pass Operation, which combines the three-dimensional features of all atoms in the compound into a three-dimensional molecular vector , The calculation formula is as follows: (23)。 8. The method for constructing a BBBNet feature fusion network according to claim 4, characterized in that: In S33, the P-MNF of each compound is input into the Transformer module to generate a text feature vector of the compound. The specific process includes: S331: For any atom in a compound , built using the n-gram model Bag of words vector , the specific construction formula is as follows: (24) in, Indicates The frequency of occurrence of n-grams in SMILES expressions, is the size of the n-grams vocabulary; S332: Generate The position encoding vector ,use generate Position-enhanced bag-of-words vector , the specific generation formula is as follows: (25) (26) (27) (28) in, express Indexes in SMILES expressions; yes Dimension index, when When it is an even number, use generate ;when When it is an odd number, use generate ; S333: Use Calculating text attention features , the specific calculation formula is as follows: (29) in, is the text weight matrix generated by random initialization, with a size of ; Indicates that the compound Any atom other than ; text query matrix express The position of the enhanced word bag vector in the query vector space; the text key matrix express Mapping in key-vector space; text-value matrix express A mapping in the value space used to store ; and besides To count atoms and atoms Correlation of text features between The dimension size is The text attention weight matrix represents and The position of enhances the relationship between bag-of-words vectors; Textual attention features that represent the correlations between atoms within a compound; S334: Use and Compute text transformation features , the specific calculation formula is as follows: (30) (31) in, Represents a normalization operation; It is the middle feature of the text; is a feed-forward neural network used to Perform nonlinear transformations; S335: All atoms of the compound conduct Operation, get the molecular text feature vector of the compound , The calculation formula is as follows: (32)。 9. The method for constructing a BBBNet feature fusion network according to claim 4, characterized in that: The specific process of feature fusion performed by the attention feature fusion module Afusion in S34 includes: S341: , , , Fusion is performed to obtain a vector that combines molecular graph features, molecular three-dimensional features, molecular text features, and molecular descriptor features. , The specific calculation formula is as follows: (33) S342: Use Calculate the molecular query vector , molecular bond vector , numerator value vector , the specific formula is as follows: (34) (35) (36) in, Is the size of The modality query weight matrix of ; express Mapping in query vector space; Is the size of The modal bond weight matrix of ; express Mapping in key vector space; Is the size of The modal bond weight matrix of ; express Mapping in the space of value vectors; and Used for calculation The similarities between different features in S343: Use and Calculating self-attention weights , The specific generation formula is as follows: (37) in, is the self-attention score; Represents the transpose operation of a matrix; is the scaling factor used to prevent The value is too large; is a normalization function used to convert Convert to probability distribution; represents the self-attention weight; S344: Use and Calculate the characteristic fusion vector of the compound , the specific calculation formula is as follows: (38)。 10. The method for predicting compound blood-brain barrier permeability based on multimodal fusion according to claim 1, characterized in that: The S4 will Divide into training set and test set, and input the training set into the Kolmogorov-Arnold model KAN; After the training is completed, the performance of the trained KAN model is evaluated using the test set; The specific steps for training and evaluating the performance of the model include: S41: Divide into training set and test set ,in Used to train the KAN model, Used to evaluate the performance of the KAN model; S42: Any Input into the KAN model for training and get the Prediction of blood-brain barrier permeability signatures for compounds , the specific calculation formula is as follows: (39) (40) (41) in, It is B-spline basis functions; It is A B-spline superposition function, given by indivual The sum is obtained; It is The bias term of the KAN mapping layer; is a linear function, indicating that KAN mapping layer; is a random initial weight, indicating The initial weights of the KAN mapping layers; Yes conduct The weight obtained after rounds of iterations represents the The weight coefficient of each KAN mapping layer; for the number of yes The number of medium triplets; The value range of is {BBB-, BBB+}; S43: Configure the cross entropy loss function , the specific formula is as follows: (42) in, Indicates The cross-loss function of each compound; S44: Use gradient descent to Iterate: Set the loss function change The convergence threshold is ; Calculate in each iteration , when satisfied or When , the iteration stops and outputs , get the trained KAN model; It is calculated by the following iterative equations (43)~(44): (43) (44) in, is a fixed learning rate; Yes conduct The weight obtained after round iteration; Indicates Blood-brain barrier permeability labeling of compounds; Indicates The blood-brain barrier permeability signature prediction values ​​of the compounds; express right The gradient of It is a dynamically adjustable value, set according to the size of data of different scales; S45: Use Perform performance evaluation on the trained KAN model.

Citation Information

Patent Citations

  • Model training method and molecular property information prediction method and device

    CN116524998A

  • Method and device for screening cross-placental transport compound based on multi-modal fusion model and storage medium

    CN118866155A

  • Compound blood brain barrier permeability prediction method based on ternary mixed level fusion convolutional neural network

    CN118918943A

  • Multi-modal fusion-based drug-protein interaction prediction model construction method

    CN119580818A