Molecular fingerprint prediction method and system based on CNN and Transform network feature fusion
By combining the feature fusion method of CNN and Transformer networks in molecular fingerprint prediction, the problems of low prediction accuracy and insufficient computing efficiency in the prior art are solved, high-precision and efficient large-scale data processing are achieved, and a new way to annotate unknown metabolites is provided.
Patent Information
- Application Number
- CN202510124830.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-27
AI Technical Summary
When processing mass spectrometry data, existing molecular fingerprint prediction methods are difficult to capture the key features in the data comprehensively and accurately, resulting in low accuracy and reliability of prediction results, and low computational efficiency during large-scale data processing, which cannot meet the needs of practical applications.
Using a molecular fingerprint prediction method based on the fusion of CNN and Transformer network features, the global feature extraction of Transformer and local sub-feature extraction of CNN is combined, and the final molecular fingerprint prediction is performed using a multi-layer perceptron.
It significantly improves the accuracy and robustness of molecular fingerprint prediction, can effectively process large-scale data, meet the needs of rapid screening and analysis of massive compounds in practical applications, and provides a new way to annotate unknown metabolites.
Smart Images

Figure CN120048383A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the cross - field of machine learning and molecular informatics, and particularly relates to a molecular fingerprint prediction method and system based on the feature fusion of CNN and Transformer networks, which are used for encoding molecular structure features and high - precision molecular characterization. This method can be widely applied to fields such as drug design, compound screening, and chemical information retrieval. Through molecular fingerprint prediction, the similarity search in a large - scale compound database can be improved, and then the metabolite identification work can be completed. Background Art
[0002] In fields such as metabolomics and drug research and development, the analysis of small - molecule metabolites in biological samples is extremely crucial. Liquid chromatography - mass spectrometry (LC - MS) technology, with its characteristics of high resolution and sensitivity, has become an important means for analyzing metabolites in biological samples. The mass spectrometry (MS / MS) spectra generated by it contain rich metabolite structure information and are an indispensable tool for metabolite annotation.
[0003] However, the current mass spectrometry libraries have obvious defects, and the coverage only involves a small part of known compounds. In non - targeted metabolomics research, this limitation is particularly prominent. A large number of unknown metabolites are difficult to be annotated by traditional spectral matching methods, seriously hindering the in - depth exploration of metabolic processes in organisms and the R & D process of new drugs.
[0004] To solve the above problems, the molecular fingerprint prediction technology based on computational methods has emerged. However, the existing molecular fingerprint prediction means still have many deficiencies. On the one hand, traditional prediction methods based on rules or simple algorithms cannot effectively process the complex information in mass spectrometry data. For compounds with similar structures, it is difficult to accurately distinguish and predict their molecular fingerprints, resulting in low accuracy and reliability of the prediction results. On the other hand, although some current machine - learning - based prediction methods have improved the prediction performance to a certain extent, due to the high complexity and diversity of mass spectrometry data, these methods are difficult to comprehensively and accurately capture the key features in the data and cannot fully explore the complex relationship between mass spectrometry data and molecular fingerprints. In addition, the existing prediction technologies have low computational efficiency when dealing with large - scale data, unable to meet the requirements of rapid screening and analysis of a large number of compounds in practical applications, and restricting their wide application in fields such as drug design, compound screening, and chemical information retrieval. Summary of the Invention
[0005] The purpose of the present invention is to overcome the above - mentioned deficiencies of the prior art and provide a molecular fingerprint prediction method and system based on the feature fusion of CNN and Transformer networks.
[0006] The present invention is implemented as follows. In a first aspect, the present invention provides a molecular fingerprint prediction method based on the feature fusion of CNN and Transformer networks, including the following steps:
[0007] Step S1: Obtain a data set consisting of the mass spectrometry data, precursor ions, and molecular identifier SMILES of a compound, and use the molecular fingerprint generated from the molecular identifier SMILES as a label; divide the data set into a training set and a test set;
[0008] Step S2: Construct a fusion model, and use the training set and the test set to train and test it;
[0009] Step S3: Use the trained and tested fusion model to predict molecular fingerprints;
[0010] The fusion model includes a Transformer-based feature extraction model, a CNN-based feature extraction model, and a fusion network.
[0011] Preferably, the Transformer-based feature extraction model takes the mass spectrometry data as input and outputs global molecular fingerprint features, specifically a vector consisting of probability values between 0 and 1.
[0012] Preferably, the implementation process of the Transformer-based feature extraction model is as follows:
[0013] Map the mass-to-charge ratio (m / z) and its corresponding peak intensity (intensity) information in the mass spectrometry data into a one-dimensional vector of 5000 bins in size, then input it into the Transformer encoder for further processing, and finally send it to a linear layer for dimensional mapping, and convert the molecular fingerprint output by the linear layer into a probability value between 0 and 1 through the Sigmoid activation function.
[0014] The Transformer encoder includes multiple stacked encoder layers, and each layer contains a multi-head attention layer and a feed-forward feedback layer;
[0015] The multi-head attention layer focuses on different parts of the spectral data in different attention heads to capture the long-range dependencies between peaks, thereby extracting the global features of the mass spectrometry data;
[0016] The feed-forward feedback layer further transforms and processes the global features extracted by the multi-head attention layer, and gradually abstracts the spectral information.
[0017] To stabilize the training process, layer normalization and residual connections are introduced after the multi-head attention layer and the feed-forward feedback layer of each layer.
[0018] Preferably, the input of the CNN-based feature extraction model is mass spectrometry data, and the output is local molecular fingerprint features, specifically a vector composed of probability values between 0 and 1.
[0019] The CNN-based feature extraction model adopts a one-dimensional convolutional neural network (1D CNN) architecture, including two sequentially connected convolutional modules, a Dropout layer, and three fully connected layers. The convolutional module includes a one-dimensional convolutional layer and a pooling layer connected in sequence.
[0020] The implementation process of the CNN-based feature extraction model is as follows:
[0021] The input data first passes through the first convolutional module, extracts primary local features through the one-dimensional convolutional layer, and then performs downsampling through the pooling layer to reduce the feature dimension and improve the robustness of the features. Subsequently, it is input into the second convolutional module, further extracts high-level local features through the one-dimensional convolutional layer, and then continues to compress the feature representation and improve the computational efficiency through the pooling layer. After feature extraction, the pooled features are sent to the Dropout layer to alleviate the overfitting problem of the model and randomly discard some neurons to enhance the generalization ability of the model. Subsequently, the extracted features are flattened and input into three fully connected layers, gradually mapping the features to the dimension of the target molecular fingerprint. Finally, through the Sigmoid activation function, the output of the fully connected layer is converted into the probability value of each site to represent the predicted local molecular fingerprint features.
[0022] Preferably, the fusion network includes a fusion block and a multi-layer perceptron MLP;
[0023] The fusion block multiplies and adds the features output by the Transformer-based feature extraction model and the CNN-based feature extraction model through a learnable per-bit weighting parameter α, specifically as follows:
[0024] F Fused =α·F Transformer +(1-α)F CNN (1)
[0025] where F Fused represents the fused feature, F Transformer represents the global molecular fingerprint feature output by the Transformer-based feature extraction model, and F CNN represents the local molecular fingerprint feature output by the CNN-based feature extraction model;
[0026] The multi-layer perceptron MLP makes the final molecular fingerprint prediction for the fused feature F Fused
[0027] Second aspect, the present invention provides a molecular fingerprint prediction system, including:
[0028] A data acquisition module for acquiring mass spectrometry data;
[0029] A molecular fingerprint prediction module for inputting the mass spectrometry data into a trained and tested fusion model for molecular fingerprint prediction.
[0030] The beneficial effects of the present invention are as follows:
[0031] 1. Improve the accuracy of molecular fingerprint prediction: The present invention combines the feature extraction model of Transformer and the feature extraction model of CNN to capture the key information of mass spectrometry data from both the global and local levels. The combination of the two makes up for the deficiencies of a single model in feature extraction, thus significantly improving the accuracy of molecular fingerprint prediction.
[0032] 2. Suitable for efficient processing of large-scale data: The method of the present invention not only achieves significant improvement in feature extraction, but also has the ability to quickly screen and analyze in a large-scale compound database through an efficient fusion network and model structure. This is of great significance for practical applications in fields such as drug design, compound screening, and chemical information retrieval.
[0033] 3. Provide a new approach for the annotation of unknown metabolites: Currently, in metabolomics research, a large number of unknown metabolites cannot be annotated by traditional methods. Through accurate molecular fingerprint prediction, the present invention provides a new method that can perform efficient similarity searches in a large-scale compound database, quickly locate the possible structures of unknown metabolites, and thus accelerate the identification of metabolites. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions of the present invention, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0035] Figure 1 is the data preprocessing process provided by the embodiment of the present invention.
[0036] Figure 2 is the schematic diagram of the fusion model structure provided by the embodiment of the present invention.
[0037] Figure 3 is the schematic diagram of the Transformer-based feature extraction model structure provided by the embodiment of the present invention.
[0038] Figure 4It is a schematic diagram of the CNN feature extraction model structure provided by the embodiments of the present invention.
[0039] Figure 5 It is the prediction and evaluation process. Specific implementation manners
[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0041] As Figure 2 shown, the embodiments of the present invention provide a molecular fingerprint prediction method, including:
[0042] Step S1: Obtain a data set composed of the mass spectrometry data, precursor ions, and molecular identifier SMILES of a compound, and use the molecular fingerprint generated by the molecular identifier SMILES as a label; divide the data set into a training set and a test set;
[0043] Step S2: Construct a fusion model, and use the training set and the test set to train and test it;
[0044] Step S3: Use the trained and tested fusion model to predict the molecular fingerprint.
[0045] Furthermore, in step S1, the present invention selects the MassBank high-quality public mass spectrometry database. This database covers the mass spectrometry data of compounds from various sources, including plants, microorganisms, animals, etc., and has the characteristics of diverse data sources and diverse experimental conditions, providing rich basic data support for the molecular fingerprint prediction model of the present invention.
[0046] In this data set, the data set of the ESI source type is used to train and test the proposed molecular fingerprint prediction model, because the soft ionization method of ESI can effectively retain the complete structural information of the molecule. Then, metabolites with a mass-to-charge ratio (m / z) between 100 and 1300 are selected as the training targets, and at the same time, mass spectrometry data with SMILES molecular identifiers and capable of effectively generating molecular fingerprints through SMILES are selected. In terms of adduct types, mainly [M + H] + and [M - H] -The mass spectrometry data of two ionic forms are analyzed. In mass spectrometry processing, first, only mass spectra containing at least 5 spectral peaks are considered to ensure that the selected mass spectrometry data has sufficient complexity and rich useful information. Secondly, any data containing negative or zero-valued spectral peaks is excluded to ensure that all spectral peak values are positive and the validity of the spectral peaks. Finally, check and ensure that the m / z values of all spectral peaks do not exceed their relative molecular masses to ensure the validity of m / z. If the m / z values of some data exceed their thresholds, the excess part is selected to be filtered out. Only the mass spectrometry data that meets all the above conditions will be selected and retained to ensure that the provided samples are truly beneficial to model training. The finally selected data mainly includes key information such as mass spectrometry information, precursor ions, and SMILES.
[0047] In the present invention, the molecular fingerprint formed by splicing two molecular fingerprints of MACCS and Morgan is uniformly used as the target molecular fingerprint for training. Specifically, the MACCS molecular fingerprint contains 166 bits and is mainly used to encode the basic structural information of the molecule; the Morgan molecular fingerprint uses topological path encoding with a radius of 2 and contains 1024 bits, which can capture detailed information in the molecular topological structure. The head and tail of the two molecular fingerprints are spliced to finally form a binary molecular fingerprint vector with a length of 1190 bits.
[0048] Figure 1 It is the data preprocessing process provided by the embodiment of the present invention.
[0049] Furthermore, step S2 is specifically:
[0050] In the present invention, the molecular fingerprints predicted by the Transformer and the convolutional neural network (CNN) from two different perspectives of the whole and the local are used as features respectively, and the feature information is integrated through a feature fusion strategy. Then, the fused features are used as inputs and input into a multi-layer perceptron (MLP) for final molecular fingerprint prediction. The model mainly includes the following three parts: a Transformer-based feature extraction model, a CNN-based feature extraction model, and a fusion network. See the appendix Figure 2 。
[0051] (1) Transformer-based feature extraction model
[0052] The Transformer feature extraction model takes mass spectrometry data as input and outputs a fixed-length molecular fingerprint representation, specifically a vector consisting of probability values between 0 and 1. The schematic diagram is shown in Figure 3. Before training, the mass-to-charge ratio (m / z) and peak intensity information of the preprocessed mass spectrometry data are first mapped into a 5000-dimensional one-dimensional vector. Subsequently, these mass spectrometry data are input into the Transformer encoder module for further processing. The Transformer encoder consists of 6 stacked encoder layers, and each layer contains a multi-head attention layer and a feed-forward layer. To stabilize the training process, layer normalization and residual connection are introduced in each encoder layer, which not only improves the convergence speed of the model but also effectively avoids the problems of gradient disappearance and gradient explosion. After being processed by the encoder module, the features are sent to a linear layer for dimensional mapping, and the output of the linear layer is transformed into a vector of probability values between 0 and 1 through the Sigmoid activation function, denoted as y Transform =(y 1 ,y 2 ,...,y n ).
[0053] (2) CNN-based feature extraction model
[0054] The CNN-based feature extraction model adopts a one-dimensional convolutional neural network (1D CNN) architecture, which consists of 8 layers in total, including two one-dimensional convolutional layers, pooling layers, a Dropout layer, and three fully connected layers. The schematic diagram of the model is shown in Figure 4 . The input data of the model first passes through the first convolutional layer to extract primary local features, and then immediately undergoes downsampling through the pooling layer to reduce the feature dimension and improve the robustness of the features. Subsequently, the second convolutional layer further extracts advanced local patterns, and then it is also connected to a pooling layer to continue compressing the feature representation and improving the computational efficiency. After feature extraction, the pooled features are sent to the Dropout layer to alleviate the overfitting problem of the model and randomly discard some neurons to enhance the generalization ability of the model. Subsequently, the extracted features are flattened and input into three fully connected layers to gradually map the features to the dimension of the target molecular fingerprint. Finally, through the Sigmoid activation function, the output of the fully connected layer is transformed into a vector of probability values for each site, denoted as y CNN =(y 1 ,y 2 ,...,y n ).
[0055] (3) Fusion network
[0056] The fusion network includes a fusion block and a multi-layer perceptron (MLP); the multi-layer perceptron (MLP) is selected to complete the final molecular fingerprint prediction task, and the multi-layer perceptron (MLP) has strong flexibility and compatibility. In the feature fusion module, first, the molecular fingerprints predicted by two models, namely Transformer and convolutional neural network (CNN), are used as Feature 1 and Feature 2 respectively. Then, these two features are fused through a bit-by-bit weighted linear fusion strategy to obtain a fused feature, and the fused feature is used as the input to the MLP model for the final molecular fingerprint prediction. The model principle is as Figure 2 shown. In the feature fusion process, the features of Transformer and CNN are multiplied by weights and then added through a learnable bit-by-bit weighting parameter α. α is a learnable parameter, and its initial value is set to 0.5, indicating that the weights of the features extracted by the two models are the same at the beginning. Then, a trained α value will be finally obtained through MLP training to adapt to the model. The fusion formula is:
[0057] F Fused =α·F Transformer +(1-α)F CNN (1)
[0058] where F Fused represents the fused feature, F Transformer represents the feature obtained from the Transformer model, i.e., Feature 1 in the figure, and F CNN is the feature obtained from the CNN model, i.e., Feature 2 in the figure, and α represents the weight parameter. Before and after feature fusion, L2 norm normalization is performed on the features to ensure the consistency of the data distributions from different feature sources, thereby reducing the impact of feature distribution differences on the model performance. The fused feature is input into the multi-layer perceptron (MLP) for further training with the target molecular fingerprint.
[0059] The MLP module consists of 3 fully connected layers, and is combined with the SiLU activation function to enhance the model's fitting ability for non-linear features. Finally, through the Sigmoid activation function, the output of the fully connected layer is converted into a probability value vector between 0 and 1, denoted as y=(y 1 ,y 2 ,...,y n ). Then, through a threshold of 0.5, the probability value of the predicted molecular fingerprint is converted into a binary-form molecular fingerprint, and the specific calculation formulas are as shown in (2) and (3).
[0060]
[0061]
[0062] where, p iis the probability value of the i-th bit, y i is the output of the fully connected layer, f i is the binary of the i-th bit, and the finally generated predicted molecular fingerprint is f = (f 1 , f 2 ,..., f n ).
[0063] (4) Model Training
[0064] Before training, the preprocessed MassBank dataset was divided into a training set and a test set, with a ratio of 9:1. First, the training set was used to separately train the Transformer molecular fingerprint prediction model and the CNN molecular fingerprint prediction model, and their best parameter models were saved. Then, before training the feature fusion model, the trained Transformer and CNN models were separately loaded for feature extraction. After that, the extracted features were first subjected to a feature fusion operation, and the fused features were used as the input of the MLP model and then retrained with the target molecular fingerprint. Finally, the truly desired molecular fingerprint was predicted. The parameter configuration for model training is shown in Table 1.
[0065] Table 1: Parameter Configuration for Model Training
[0066] GPU NVIDIA RTX 2080 Cuda 11.8 Software environment Python 3.9 Optimizer Adam Loss function BCELoss Learning rate 0.001 Number of training epochs 300
[0067] The specific training loss calculation formula is as shown in (4):
[0068]
[0069] where N is the number of samples, y i is the actual label of the sample, is the predicted value.
[0070] (5) Molecular Fingerprint Prediction Evaluation
[0071] The divided test set data is used for molecular fingerprint prediction. The prediction evaluation process is shown in Appendix Figure 5 . First, the precursor ions of each substance in the test set data are used to preliminarily screen compounds to narrow the search range. Then, the spectral data in the test set is input into the fusion model of the present invention to predict the molecular fingerprint. Then, the cosine similarity score is calculated with the standard molecular fingerprint of the candidate substances in the database. The calculation formula is as shown in (5), where A i and B i are the values of vectors A and B in the i-th dimension. The numerator is the dot product of vectors A and B, and the denominator is the product of the Euclidean norms of A and B.
[0072]
[0073] Finally, perform a descending sort according to the scoring results. Based on the predicted results, we calculate the Top-K results, and the calculation formula is shown in (6). Table 2 shows the Top-N results obtained from the test set data on the model of the present invention and other single network models.
[0074]
[0075] Table 2: Top-N prediction results based on the Massbank database
[0076] Rank MLP CNN Transformer CNN + MLP The present invention Top1 47.88% 52.00% 47.38% 50.12% 55.12% Top5 64.25% 68.12% 64.12% 66.12% 71.38% Top10 74.00% 76.12% 75.12% 75.50% 80.12%
[0077] According to the experimental results in Table 2, a detailed comparative analysis was carried out on the Top-1, Top-5, and Top-10 prediction accuracies of different models on the Massbank dataset. The experimental results show that the fusion model based on CNN and Transformer proposed in the present invention performs excellently in all evaluation indicators and is significantly better than other models. In terms of the Top-1 accuracy, the model of the present invention reaches 55.12%, compared with MLP (47.88%), CNN (52.00%), Transformer (47.38%), and CNN+MLP (50.12%), it is improved by 7.24%, 3.12%, 7.74%, and 5.00% respectively. This shows that the model of the present invention demonstrates higher precision in a single prediction task and can better identify the optimal molecular fingerprint. In terms of the Top-5 accuracy, the model of the present invention also performs excellently, reaching 71.38%, compared with MLP (64.25%), CNN (68.12%), Transformer (64.12%), and CNN+MLP (66.12%), it is improved by 7.13%, 3.26%, 7.26%, and 5.26% respectively. This result indicates that the model of the present invention maintains high stability and accuracy in multi-candidate prediction. In addition, in terms of the Top-10 accuracy, the model of the present invention further demonstrates its advantages, with the accuracy reaching 80.12%, compared with MLP (74.00%), CNN (76.12%), Transformer (75.12%), and CNN+MLP (75.50%), it is improved by 6.12%, 4.00%, 5.00%, and 4.62% respectively. This result verifies that the model of the present invention still has strong prediction ability and generalization ability within a larger candidate range.
[0078] Through the above comparative analysis, it can be found that single models (such as MLP, CNN, Transformer) have certain limitations in dealing with the molecular fingerprint prediction task. For example, both MLP and Transformer perform relatively weakly in terms of Top-1 and Top-5 accuracy, which may be due to their deficiencies in local feature extraction or global modeling. However, the model of the present invention that combines the local feature extraction ability of CNN and the global modeling ability of Transformer has achieved a significant improvement in prediction accuracy and generalization ability.
[0079] This embodiment also proposes a molecular fingerprint prediction system based on the above method, including:
[0080] A data acquisition module that acquires mass spectrometry data;
[0081] A molecular fingerprint prediction module that inputs the mass spectrometry data into the trained and tested fusion model for molecular fingerprint prediction.
[0082] In summary, a molecular fingerprint prediction method and system based on the feature fusion of convolutional neural network (CNN) and Transformer proposed by the present invention successfully combines the advantages of the two deep learning architectures and realizes the comprehensive capture of the global and local features of the molecular structure. In view of the complexity and diversity of mass spectrometry data, the Transformer module effectively models long-range dependencies and extracts global context information using the self-attention mechanism; the CNN module captures the local patterns of mass spectrometry data through convolutional operations and explores the detailed connections between peaks. Then, the features of the two are fused to make up for the defects of the two models. Finally, a multi-layer perceptron (MLP) performs a non-linear mapping on the fused features, significantly improving the accuracy and robustness of molecular fingerprint prediction.
[0083] Experimental results show that the proposed method performs excellently in the evaluation of Top-N accuracy on the MassBank dataset, with excellent performance in terms of Top-1, Top-5, and Top-10 accuracy, which fully verifies the effectiveness of the feature fusion strategy in improving the prediction accuracy and practicality of the model.
[0084] In conclusion, the molecular fingerprint prediction method based on the feature fusion of CNN and Transformer proposed by the present invention provides a brand-new and powerful solution for efficiently analyzing mass spectrometry data and realizing the annotation of unknown metabolites.
[0085] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.
Claims
1. A molecular fingerprint prediction method based on CNN and Transformer network feature fusion, characterized in that The following steps are involved: Step S1, obtaining mass spectrum data, precursor ions, and molecular identifier SMILES of the compound to form a data set, and using the molecular identifier SMILES to generate a molecular fingerprint as a label; Divide the dataset into training and testing sets; Step S2: construct a fusion model, and train and test it using a training set and a test set; Step S3, using the trained and tested fusion model to predict molecular fingerprints; The fusion model includes a Transformer-based feature extraction model, a CNN-based feature extraction model, and a fusion network.
2. The method according to claim 1, characterized in that: The Transformer-based feature extraction model takes mass spectrometry data as input and outputs global molecular fingerprint features.
3. The method according to claim 2, characterized in that: The implementation process of the Transformer feature extraction model is as follows: The mass-to-charge ratio and its corresponding peak intensity information in the mass spectrometry data are mapped into a one-dimensional vector, then input into the Transformer encoder for further processing, and finally sent to the linear layer for dimensional mapping. The molecular fingerprint output by the linear layer is converted into a probability value between 0 and 1 through the Sigmoid activation function.
4. The method according to claim 3, characterized in that: The Transformer encoder includes multiple stacked encoder layers, each of which includes a multi-head attention layer and a feedforward feedback layer; The multi-head attention layer focuses on different parts of the spectrogram data in different attention heads, captures the long-distance dependencies between peaks, and thus extracts the global features of the mass spectrometry data; The feedforward feedback layer further transforms and processes the global features extracted by the multi-head attention layer, gradually abstracting the spectrogram information.
5. The method according to claim 4, characterized in that: Layer normalization and residual connection are introduced after the multi-head attention layer and feedforward feedback layer of each layer in the Transformer encoder.
6. The method according to claim 1, characterized in that: The input of the CNN-based feature extraction model is mass spectrometry data, and the output is local molecular fingerprint features.
7. The method according to claim 6, characterized in that: The CNN-based feature extraction model adopts a one-dimensional convolutional neural network, including two convolution modules connected in series, a Dropout layer, and three fully connected layers, wherein the convolution module includes a one-dimensional convolution layer and a pooling layer connected in series.
8. The method according to claim 7, characterized in that: The implementation process of the CNN feature extraction model is as follows: The input data first passes through the first convolution module, extracts primary local features through a one-dimensional convolution layer, and then downsamples through a pooling layer to reduce the feature dimension and improve the robustness of the feature; then, it is input to the second convolution module, further extracts high-level local features through a one-dimensional convolution layer, and then continues to compress feature representation and improve computational efficiency through a pooling layer; after feature extraction, the pooled features are sent to the Dropout layer to alleviate the overfitting problem of the model, and some neurons are randomly discarded to enhance the generalization ability of the model; then, the extracted features are flattened and input into three fully connected layers to gradually map the features to the dimensions of the target molecular fingerprint; finally, through the Sigmoid activation function, the output of the fully connected layer is converted into a probability value for each site to represent the predicted local molecular fingerprint features.
9. The method according to claim 1, characterized in that: The fusion network includes a fusion block and a multi-layer perceptron MLP; The fusion block weights and adds the features output by the Transformer feature extraction model and the CNN feature extraction model through a learnable bit-by-bit weight parameter α, as follows: F Fused =α·F Transformer +(1-α)F CNN (1) Among them, F Fused represents the fused features, F Transformer represents the global molecular fingerprint feature output by the Transformer feature extraction model, F CNN Represents the local molecular fingerprint features output by the CNN feature extraction model; The multi-layer perceptron MLP performs the fusion of the features F Fused Make the final molecular fingerprint prediction.
10. A molecular fingerprint prediction system based on the method according to any one of claims 1 to 9, characterized in that include: A data acquisition module, for acquiring mass spectrometry data; The molecular fingerprint prediction module inputs the mass spectrometry data into the trained and tested fusion model for molecular fingerprint prediction.
Citation Information
Cited By
Non-targeted metabonomics molecular fingerprint intelligent generation method and device
CN120544691A
Biomarker detection method and system based on biological spectrum
CN121656160A