Metabolite structure prediction method and system based on BERT twinning neural network model
By introducing the BERT twin neural network model in metabolomics data processing, the problem of high-dimensional sparsity processing of metabolomics data in the existing technology is solved, more efficient and accurate metabolites structure prediction is achieved, and the efficiency of drug development is improved.
Patent Information
- Application Number
- CN202510136075.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-07
AI Technical Summary
The existing technology is difficult to effectively handle the high-dimensional sparsity of metabolomics data, and it is impossible to accurately capture key features in the data, resulting in poor metabolites structure analysis effect, and when faced with noise and nonlinear relationships, prediction accuracy and stability need to be improved.
The metabolite structure prediction method based on the BERT twin neural network model is adopted, and a two-way BERT twin neural network model is built and trained through the pre-processing of mass spectrometry data and the calculation of molecular fingerprints, and the mass spectrometry similarity and structural similarity are calculated to achieve efficient identification of metabolite structure.
It improves the accuracy and stability of metabolite structure prediction, enhances the interpretability of data, and can screen a large number of unknown metabolite structures more quickly to meet the needs of drug research and development.
Smart Images

Figure CN120015180A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of metabolite detection in metabolomics, and specifically relates to a metabolite structure prediction method and system based on a BERT twin neural network model, which is a study of a deep learning model for mass spectrometry data processing and detection of metabolites. Background Art
[0002] Metabolomics aims to systematically analyze metabolites in organisms and explore their roles in physiological, pathological and environmental changes. With the development of high-throughput technologies such as mass spectrometry (MS) and nuclear magnetic resonance (NMR), the scale of metabolomics data has increased dramatically, showing high dimensionality, complexity and heterogeneity. How to effectively mine valuable information from it has become a research problem.
[0003] There are many limitations in the existing methods for predicting metabolite structures. Traditional methods based on statistical analysis and dimensionality reduction techniques (such as principal component analysis, t-SNE, etc.) have difficulty in dealing with the high-dimensional sparsity of metabolomics data and cannot accurately capture the key features in the data, resulting in poor analysis of metabolite structures. For example, when analyzing metabolomics data of complex biological samples, these methods may lose important structural information and cannot accurately distinguish metabolites with similar structures.
[0004] Machine learning-based methods, such as support vector machines (SVM), random forests, and K-nearest neighbors (KNN), have improved prediction capabilities to a certain extent, but they are still insufficient when faced with noise and nonlinear relationships in metabolomics data. They are difficult to fully explore the complex relationship between mass spectrometry data and metabolite structure, and the prediction accuracy and stability need to be improved. In practical applications, these methods are easily affected by fluctuations in data quality, resulting in deviations in prediction results.
[0005] In addition, the computational efficiency of existing prediction methods is low on large-scale metabolomics datasets, making it difficult to meet the needs of rapidly identifying a large number of unknown metabolite structures. For example, in drug development, a large number of potential active metabolites need to be rapidly screened, and the efficiency of existing methods limits the development process. Therefore, it is urgent to develop a more efficient and accurate metabolite structure prediction method.
[0006] The present invention proposes a metabolite structure prediction method and system based on the BERT twin neural network model, aiming to perform efficient pattern recognition, anomaly detection and feature extraction based on the similarity relationship of metabolomics data by introducing the BERT twin neural network model. This method can help researchers more accurately identify the potential rules in metabolomics data, improve the interpretability and predictive ability of data, and provide more reliable technical support for the in-depth analysis and application of metabolomics. Summary of the invention
[0007] The purpose of the present invention is to overcome the shortcomings of the above-mentioned prior art and provide a metabolite structure prediction method and system based on the BERT twin neural network model, which provides an effective solution to the problems of sparse mass spectrometry data and high dimensionality, and has been tested on actual data and achieved good results.
[0008] The present invention is implemented in this way. The first object of the present invention is to provide a metabolite structure prediction method based on the BERT twin neural network model, the method comprising: Step S1, obtaining mass spectrum data and SMILES molecular code of the compound; preprocessing the mass spectrum data and converting the SMILES molecular code into a molecular fingerprint of a fixed length; and calculating the structural similarity between the molecular fingerprints; The preprocessed mass spectrometry data were used as a dataset, in which the structural similarity between molecular fingerprints was used as a label; Step S2: Build a BERT twin neural network model and use the data set for training; The BERT twin neural network model includes two parallel sub-networks; The inputs of the two sub-networks are respectively the pre-processed mass spectrum data of different compounds, and the outputs are respectively the first mass spectrum data information features and the second mass spectrum data information features; Calculate mass spectrum similarity using the first mass spectrum data information features and the second mass spectrum data information features output by the two sub-networks; Step S3: Use the trained BERT twin neural network model to identify the metabolite structure.
[0009] Preferably, in step S1, all mass spectrum data are preprocessed and set to one-dimensional array information of the same length, with mass-to-charge ratio as the index and intensity information normalized in the range of 0-1.
[0010] Preferably, the calculation of the structural similarity in step S1 is as follows: in and are the intensities of the i-th peak in the mass spectra of the two compounds, and n is the number of peaks.
[0011] Preferably, in step S2, each sub-network includes a plurality of network units connected in series, and each network unit includes a BERT network and a feedforward network connected in series.
[0012] Preferably, the mass spectrum similarity in step S2 is calculated as follows: in and are the intensities of the i-th peak in the mass spectra of the two compounds, and n is the number of peaks.
[0013] Preferably, the loss function based on the BERT twin neural network model during the training process in step S2 is as follows: Loss = mean(L ) Among them, Loss represents the mean square error loss, is the evenly divided loss between mass spectral similarity and structural similarity, , Represent the predicted mass spectrum similarity and the actual structural similarity, respectively; T represents transposition; The Adam algorithm is used to continue network training until the model reaches loss convergence or the maximum number of iterations is reached.
[0014] Preferably, step S3 specifically includes: 3-1 After preprocessing, the mass spectrometry data of the unknown metabolite to be identified is used to construct mass spectrometry pairs with the mass spectrometry data of known compounds; the constructed mass spectrometry pairs are respectively input into the two sub-networks of the trained BERT-based twin neural network for mass spectrometry similarity prediction; 3-2 Find analogues from known compounds based on the molecular formula or relative molecular mass of unknown metabolites, and calculate the structural similarity between these analogues and known compounds; Based on mass spectrometry similarity and structural similarity, the substances most similar to the unknown metabolites were obtained.
[0015] The second object of the present invention is to provide a metabolite structure prediction system, comprising: The mass spectrometry similarity calculation module is responsible for constructing mass spectrometry pairs between the mass spectrometry data of the unknown metabolites to be identified and the mass spectrometry data of the known compounds after preprocessing; the constructed mass spectrometry pairs are respectively input into the two sub-networks of the trained BERT-based twin neural network for mass spectrometry similarity prediction; The structural similarity calculation module is responsible for finding analogs from known compounds based on the molecular formula or relative molecular mass of unknown metabolites and calculating the structural similarity between these analogs and known compounds; The metabolite structure prediction module is responsible for deriving the substance most similar to the unknown metabolite based on mass spectrometry similarity and structural similarity.
[0016] A third object of the present invention is to provide a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed in a computer, the computer executes the method described.
[0017] A fourth object of the present invention is to provide a computing device, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described is implemented.
[0018] The present invention has the following beneficial effects: 1. Compared with traditional methods for calculating spectrogram similarity (cosine similarity, modified cosine similarity, etc.) and traditional machine learning methods (SVM, random forest, KNN, etc.), the present invention applies a deep learning method that can better extract data features when facing a large amount of high-dimensional and sparse data, thereby improving the accuracy of calculating similarity.
[0019] 2. When processing metabolomics data, the present invention uses the BERT network to mine the potential semantic information of mass spectrometry data through a bidirectional self-attention mechanism, capture long-distance dependencies, and thus improve the accuracy of mass spectrometry similarity calculation, providing a more reliable basis for subsequent metabolite structure identification.
[0020] 3. The metabolite structure identification model based on the BERT twin neural network proposed in the present invention introduces the BERT network to improve the feature extraction ability of the network, better calculate the similarity between mass spectra and reduce the time and space complexity.
[0021] 4. The BERT twin neural network model of the present invention introduces the BERT network, which greatly improves the feature extraction ability of the network. The BERT network can efficiently encode mass spectrometry data, extract key peak information, and reduce unnecessary calculation steps. Compared with other complex models, this model effectively reduces the spatiotemporal complexity while ensuring prediction accuracy. When screening a large number of compounds in drug development, metabolite structure analysis can be completed in a shorter time, saving computing resources and accelerating the development process. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solution of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0023] Figure 1 This is the BERT network model structure provided by an embodiment of the present invention.
[0024] Figure 2 It is the main structure of the BERT twin neural network model provided in an embodiment of the present invention.
[0025] Figure 3This is the application process of metabolite structure identification based on the BERT twin neural network model provided in the embodiment of the present invention DETAILED DESCRIPTION
[0026] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0027] like Figure 1 As shown, a metabolite structure prediction method based on a BERT twin neural network model provided in an embodiment of the present invention includes: Step S1, obtaining mass spectrum data and molecular identifier SMILES molecular code of the compound; The data set used in the present invention is a mass spectrometry data set of some metabolite reactant and product pairs in the KEGG molecular database.
[0028] All mass spectrometry data are preprocessed and set to one-dimensional array information of the same length, with mass-to-charge ratio (m / z) as the index and intensity information normalized to the range of 0-1; the details are as follows: Mass spectrometry data are usually expressed as , where the mass-to-charge ratio in the mass spectrometry data is , the relative intensity of the peak .
[0029] All mass spectrometry data Convert to standard format , The mass-to-charge ratio is defined as The step size of the mass-to-charge ratio is set to 1, and the relative peak intensity corresponding to each mass-to-charge ratio is Normalize the relative peak intensities: Convert SMILES molecular code into a fixed-length molecular fingerprint; SMILES is the chemical identifier of chemical molecules. Using the Rdkit library in PYTHON, SMILES is converted into a molecular fingerprint, which is a fixed-length binary code representing the molecular structure information of the chemical substance.
[0030] Simultaneously calculate the structural similarity between molecular fingerprints; The preprocessed mass spectrometry data is used as a data set, in which the structural similarity between molecular fingerprints is used as a label; the structural similarity is calculated as follows: and are the intensities at the i-th peak in the two mass spectra, respectively.
[0031] n is the number of peaks involved in the calculation (usually pairs of peaks with the same mass-to-charge ratio) Step S2: Build a BERT twin neural network model and use the data set for training; The BERT-based twin neural network model has two input ports and one output port. Each input is mass spectrum data, and the output is the mass spectrum similarity between the mass spectrum data of the two inputs, that is, the cosine similarity between the data.
[0032] See attached Figure 2 , based on the BERT twin neural network model, it includes two parallel sub-networks; Each sub-network includes multiple network units connected in series, and each network unit includes a BERT network and a feed-forward network connected in series; The inputs of the two sub-networks are the pre-processed mass spectrometry data of different compounds. The original mass spectrometry data are encoded through the BERT network, the key peak information in the mass spectrometry data is extracted, and then decoded through the feedforward network. Finally, both sub-networks obtain a one-dimensional vector containing mass spectrometry information features, which are the first mass spectrometry data information features and the second mass spectrometry data information features, respectively.
[0033] See attached Figure 1 , the BERT network includes an input Embedding layer, 6 unit structures, each unit structure includes a multi-head attention mechanism, the first Add & Norm layer, the Feed Forward layer, the second Add & Norm layer, and introduces residual connections.
[0034] Calculate mass spectrum similarity using the first mass spectrum data information features and the second mass spectrum data information features output by the two sub-networks; The mass spectrum similarity is calculated as follows: and are the intensities at the i-th peak in the two mass spectra, respectively.
[0035] n is the number of peaks involved in the calculation (usually pairs of peaks with the same mass-to-charge ratio) The similarity of the mass spectrometry data output by the network is compared with the structural similarity between the molecules represented by the input data, and used for model training and optimization, so that the model learns the correlation between mass spectrometry similarity and molecular structure similarity, which is ultimately used for metabolite structure identification.
[0036] The loss function based on the BERT twin neural network model during training is as follows: Loss = mean(L ) Among them, Loss represents the mean square error loss, is the evenly divided loss between mass spectral similarity and structural similarity, , Represent the predicted mass spectrum similarity and the actual structural similarity, respectively; T represents transposition; Use the Adam algorithm to continue network training until the model reaches loss convergence or reaches the maximum number of iterations; Step (3), see Appendix Figure 3 , using the trained BERT twin neural network model to identify metabolite structures; specifically: 3-1 The mass spectrometry data of the unknown metabolite to be identified are preprocessed and then mass spectrometry pairs are constructed with the mass spectrometry data of the known compounds. The constructed mass spectrometry pairs are respectively input into the two sub-networks of the trained BERT-based twin neural network for mass spectrometry similarity prediction.
[0037] 3-2 Find analogues from known compounds based on the molecular formula or relative molecular mass of unknown metabolites, and calculate the structural similarity between these analogues and known compounds; Based on mass spectrometry similarity and structural similarity, the substances most similar to the unknown metabolites were obtained.
[0038] The higher the similarity between mass spectrum and structure, the better.
[0039] The following is a demonstration of the effect of the present invention: 1. Dataset Selection The metabolite structure identification model based on the BERT twin neural network proposed in the present invention was evaluated on a test data set to verify the identification accuracy of the model.
[0040] The evaluation verification index used in this invention is the Top-K verification index: Top-K accuracy is a common evaluation metric that indicates whether the model contains the true label in the top K items predicted.
[0041] 2. Model verification results The test data set contains a total of 747 metabolite molecules. The experimental results are as follows: Table 1 Prediction results of the model of the present invention on the actual data set Top1 Top3 Top 5 Top 10 Model of the present invention 52.95% 73.91% 83.07% 90.68% The above table verifies that the proposed metabolite structure identification model has achieved a high identification accuracy rate on the actual data set. Therefore, the model further proves that the metabolite structure identification model based on the BERT twin neural network is credible in the task of metabolite molecular structure identification.
[0042] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A metabolite structure prediction method based on the BERT twin neural network model, characterized in that: The method comprises: Step S1, obtaining mass spectrum data and SMILES molecular code of the compound; preprocessing the mass spectrum data and converting the SMILES molecular code into a molecular fingerprint of a fixed length; and calculating the structural similarity between the molecular fingerprints; The preprocessed mass spectrometry data were used as a dataset, in which the structural similarity between molecular fingerprints was used as a label; Step S2: Build a BERT twin neural network model and use the data set for training; The BERT twin neural network model includes two parallel sub-networks; The inputs of the two sub-networks are respectively the pre-processed mass spectrum data of different compounds, and the outputs are respectively the first mass spectrum data information features and the second mass spectrum data information features; Calculate mass spectrum similarity using the first mass spectrum data information features and the second mass spectrum data information features output by the two sub-networks; Step S3: Use the trained BERT twin neural network model to identify the metabolite structure.
2. The method according to claim 1, characterized in that: In step S1, all mass spectrometry data are preprocessed and set to one-dimensional array information of the same length, with mass-to-charge ratio as the index and intensity information normalized in the range of 0-1.
3. The method according to claim 1, characterized in that: The calculation of structural similarity in step S1 is as follows: and are the intensities of the i-th peak in the mass spectra of the two compounds, and n is the number of peaks.
4. The method according to claim 1, characterized in that: In step S2, each sub-network includes multiple network units connected in series, and each network unit includes a BERT network and a feedforward network connected in series.
5. The method according to claim 1, characterized in that: The calculation of mass spectrum similarity in step S2 is as follows: and are the intensities of the i-th peak in the mass spectra of the two compounds, and n is the number of peaks.
6. The method according to claim 1, characterized in that: The loss function based on the BERT twin neural network model during the training process in step S2 is as follows: Loss = mean(L ) Among them, Loss represents the mean square error loss, is the evenly divided loss between mass spectral similarity and structural similarity, , Represent the predicted mass spectrum similarity and the actual structural similarity, respectively; T represents transposition; The Adam algorithm is used to continue network training until the model reaches loss convergence or the maximum number of iterations is reached.
7. The method according to claim 1, characterized in that: Step S3 specifically includes: 3-1 After preprocessing, the mass spectrometry data of the unknown metabolite to be identified is used to construct mass spectrometry pairs with the mass spectrometry data of known compounds; the constructed mass spectrometry pairs are respectively input into the two sub-networks of the trained BERT-based twin neural network for mass spectrometry similarity prediction; 3-2 Find analogues from known compounds based on the molecular formula or relative molecular mass of unknown metabolites, and calculate the structural similarity between these analogues and known compounds; Based on mass spectrometry similarity and structural similarity, the substances most similar to the unknown metabolites were obtained.
8. A metabolite structure prediction system based on the method according to any one of claims 1 to 7, characterized in that include: The mass spectrometry similarity calculation module is responsible for constructing mass spectrometry pairs between the mass spectrometry data of unknown metabolites to be identified and the mass spectrometry data of known compounds after preprocessing; The constructed mass spectrum pairs are respectively input into the two sub-networks of the trained BERT-based twin neural network for mass spectrum similarity prediction; The structural similarity calculation module is responsible for finding analogs from known compounds based on the molecular formula or relative molecular mass of unknown metabolites and calculating the structural similarity between these analogs and known compounds; The metabolite structure prediction module is responsible for deriving the substance most similar to the unknown metabolite based on mass spectrometry similarity and structural similarity.
9. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed in a computer, the computer executes the method according to any one of claims 1 to 7.
10. A computing device comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Metabolite recognition system based on molecular fingerprint prediction and application method thereof
CN112735532A
Drug small molecule property prediction method, device and equipment based on self-supervised learning
CN113707235A
Large-scale metabolome qualitative method based on molecular structure association network
CN114609318A
Analytical method, device and equipment for identifying known and unknown metabolites
CN114923992A
Metabolic reaction pair prediction method based on deep learning and non-targeted metabonomics
CN119339785A