Machine learning based re analysis and fticr ms joint data decoding and optimization method

By constructing a joint data decoding model, the problems of low efficiency and poor accuracy in traditional mass spectrometry analysis are solved, and efficient and accurate mass spectrometry data decoding is achieved, which has significant advantages, especially in the detection of complex protein molecules and junctions.

CN119673289BActive Publication Date: 2025-11-21SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411546779.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-01
Publication Date
2025-11-21
Estimated Expiration
2044-11-01

AI Technical Summary

Technical Problem

Traditional mass spectrometry analysis methods are inefficient and inaccurate, making it difficult to process large-scale biological data. Furthermore, they lack a unified model that integrates reaction sites and mass spectrometry data, which affects decoding performance.

Method used

We construct a joint data decoding method based on machine learning MS-RE decoding model and MS-MS decoding model. We improve decoding accuracy and efficiency through multi-task learning and data preprocessing, including data filtering, feature extraction and model optimization.

Benefits of technology

It improves the accuracy and efficiency of mass spectrometry data decoding, can handle high and low resolution data, has noise resistance and wide adaptability, and is suitable for decoding complex protein molecules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119673289B_ABST
    Figure CN119673289B_ABST
Patent Text Reader

Abstract

The application discloses a RE analysis and FTICRMS combined data decoding and optimization method based on machine learning, which comprises the following steps: obtaining and preprocessing a training data set, constructing a convolutional neural network and decoding data, training a RE data decoding model, training an MS data decoding model, evaluating and optimizing the quality of the data set, constructing an MS-FTICRMS combined algorithm, and testing and evaluating the model. The MS-RE decoding model and the MS-MS decoding model are constructed and combined, so that the model can process high-resolution and low-resolution data at the same time, and the performance of the model is improved through multi-task learning, thereby greatly improving the accuracy and efficiency of mass spectrum data decoding.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of protein mass spectrometry analysis, and particularly relates to a RE analysis and FTICRMS joint data decoding and optimization method based on machine learning. BACKGROUND

[0002] Mass spectrometry (MS) technology is one of the core tools for detecting and identifying molecular mass. In particular, Fourier transform ion cyclotron resonance mass spectrometry (FTICRMS) is widely used in the analysis of complex biological molecules, proteins and metabolites due to its high resolution and high accuracy. On the other hand, reaction site (RE) analysis is crucial for understanding the three-dimensional structure, function and interaction of proteins. In the context of large-scale biological data, how to efficiently and accurately integrate and decode mass spectrometry data and reaction site information has become an important challenge in current data analysis technology. The development of machine learning, especially deep learning technology, is bringing new opportunities to the field, improving the accuracy and efficiency of data decoding through automation and intelligent methods.

[0003] In traditional mass spectrometry analysis, mass spectrometry data is usually processed and decoded by manual or semi-automatic methods. For example, the data output by the mass spectrometry analyzer needs to be manually filtered, impurities removed, and peak values identified through multiple steps, which is time-consuming and prone to human error. The analysis of reaction sites generally relies on structural biology methods such as X-ray crystallography, which can provide accurate three-dimensional structure information, but the data processing is cumbersome and often requires expert experience to complete. With the increasing size of mass spectrometry data and biological data, traditional manual processing methods have been unable to cope with the surge in data volume and the complexity of analysis requirements. In addition, traditional mass spectrometry decoding often ignores the integrated analysis of different data sources such as reaction site data and mass spectrometry data, which limits the comprehensiveness and accuracy of data decoding.

[0004] Although traditional mass spectrometry and reaction site analysis techniques have achieved certain results in the past, they still have many shortcomings in efficiency and accuracy. First, manual decoding of mass spectrometry data is inefficient, and data processing speed is difficult to match the high-throughput requirements of modern scientific research. Second, manual removal of impurities and noise in mass spectrometry data is not only time-consuming, but also prone to bias, leading to a decrease in the accuracy of decoding results. Finally, traditional methods often process mass spectrometry data and reaction site information separately, lacking a unified model to integrate both, thereby affecting the overall decoding effect. SUMMARY

[0005] The main purpose of the present application is to overcome the shortcomings and deficiencies of the prior art, provide a machine learning-based RE analysis and FTICRMS joint data decoding and optimization method, by constructing MS-RE decoding model and MS-MS decoding model, the model can process high resolution and low resolution data at the same time, and the performance of the model is improved through multi-task learning, thereby greatly improving the accuracy and efficiency of mass spectrum data decoding.

[0006] In order to achieve the above purpose, the present application adopts the following technical scheme:

[0007] In a first aspect, the present application provides a machine learning-based RE analysis and FTICRMS joint data decoding and optimization method, comprising the following steps:

[0008] Collecting protein crystal structure data and mass spectrum data, performing preliminary screening and preprocessing on the data, and obtaining a training data set; the training data set includes reaction site data and mass spectrum subgraph data set;

[0009] Constructing a data decoding model, training the data decoding model with reaction site data, automatically extracting and learning the key local features of the reaction site data, obtaining a first data decoding model, and training the data decoding model with the mass spectrum subgraph data set, learning different protein molecular characteristics, and obtaining a second data decoding model;

[0010] Evaluating the first data decoding model through a characteristic curve to obtain a first evaluation result; comprehensively evaluating the second data decoding model through a quality evaluation index to obtain a second evaluation result, the quality evaluation index including resolution and peak characteristics; optimizing the model using the first evaluation result and the second evaluation result respectively to obtain the best first data decoding model and the best second data decoding model;

[0011] Combining the best second data decoding model and the best first data decoding model to obtain a first combined decoding model, combining multiple best second data decoding models to obtain a second combined decoding model, taking the first combined decoding model as the backbone part and the second combined decoding model as the additional module, and constructing a joint decoding model, wherein the output end of the second combined decoding model is connected to the output end of the first combined decoding model;

[0012] Testing the joint decoding model using a given independent test set, and evaluating the decoding performance through actual mass spectrum data and protein structure data, obtaining the best joint decoding model based on the evaluation result, and decoding the mass spectrum data of the protein using the best joint decoding model.

[0013] As a preferred technical scheme, the preliminary screening and preprocessing of the data comprises:

[0014] The reaction site data of the protein is obtained by extracting spatial position information of protein crystal structure data through X-ray crystallography technology;

[0015] The mass spectrum data set is obtained by preliminarily screening the mass spectrum data by comparing the structure of the protein, including a non-ion and solvent molecule set, a non-bridging data set, and a bridging data set, the non-ion and solvent molecule set being a pure protein structure, and the non-bridging data set and the bridging data set having impurities;

[0016] The mass spectrum data set is cut into a fixed size subgraph by using a window algorithm;

[0017] The mass-to-charge ratio of water molecules, solvent molecules, and other interference data in the non-bridging data set and the bridging data set is identified by analyzing the characteristic peak value, and the water molecules, solvent molecules, and other interference data are filtered, as follows:

[0018]

[0019] Wherein, the mass spectrum data is M, the mass-to-charge ratio of the i-th data point of the mass spectrum data is m / z i , M water is the mass of water molecules, ∈ is an allowable error range, and the other interference data includes lipids, salt ions, additives, or other impurity molecules.

[0020] As a preferred technical solution, the data decoding model includes N convolutional blocks and a fully connected layer, the output of the convolutional block is connected to the input of the next convolutional block, and the output of the Nth convolutional block is connected to the fully connected layer; the convolutional block includes a convolutional layer and a pooling layer, and the output of the convolutional layer is connected to the pooling layer.

[0021] As a preferred technical solution, the first data decoding model is evaluated by a characteristic curve, including:

[0022] Each sample in the reaction site data is predicted by different decision thresholds, for each threshold, the true positive rate and the false positive rate are calculated, and the ROC curve is generated according to the true positive rate and the false positive rate of each threshold;

[0023] The robustness of the current model is evaluated based on the AUC value calculated from the ROC curve, and the first evaluation result is generated.

[0024] As a preferred technical solution, the second data decoding model is comprehensively evaluated by a quality evaluation index, including:

[0025] The resolution and signal-to-noise ratio of the data are calculated according to the mass spectrum subgraph data set;

[0026] The peak value characteristics of each subgraph of the mass spectrum are analyzed, sample peak value characteristics are obtained, the matching degree of the sample peak value characteristics and the target peak value characteristics is calculated, and the data with low matching degree is marked as low-quality data;The sample peak value characteristics include peak shape, peak position and peak intensity.

[0027] As a preferred technical solution, the calculation of the matching degree is as follows:

[0028]

[0029] Wherein, I observed is the observed peak intensity, I expected is the expected peak intensity, and N is the number of peaks.

[0030] As a preferred technical solution, the model is optimized using the second evaluation result, specifically:

[0031] The low-quality data is excluded.

[0032] The low-quality data is added to the training data set.

[0033] As a preferred technical solution, the first combined decoding model is used as the backbone part, and the second combined decoding model is used as the additional module to construct the joint decoding model, specifically:

[0034] The sample data is preliminarily decoded by the first combined decoding model, and the sample data is refined by the second combined decoding model, the preliminary decoding result and the refined decoding result are spliced, and the joint data decoding model is obtained;The second combined decoding model has at least one.

[0035] As a preferred technical solution, the decoding performance is evaluated by actual mass spectrum data and protein structure data, specifically:

[0036] The test output result is compared with the real mass spectrum data, the performance of the comparison result is evaluated by the performance index, and the best joint decoding model is selected based on the evaluation result;The performance index ROC curve and its AUC value, accuracy, recall rate and F1 value.

[0037] As a preferred technical solution, the joint decoding model is also optimized by hyperparameter optimization and transfer learning;The hyperparameter optimization includes grid search and Bayesian optimization.

[0038] Compared with the prior art, the present application has the following advantages and beneficial effects:

[0039] (1) The present application can effectively extract local features in the data by constructing a data decoding model for mass spectrometry data and reaction site data, and further enhance the analysis ability of complex data through multi-layer structure. This method can greatly improve the accuracy of mass spectrometry data decoding, especially in the detection of complex protein molecules and connection points.

[0040] (2) The present application can ensure the quality of input data by preprocessing and screening mass spectrometry data, especially removing water molecules and solvent molecules and other interference factors, so that the model can learn important features from pure data. This data quality optimization process lays the foundation for efficient decoding of the model, improving the accuracy and robustness of the model.

[0041] (3) The present application combines MS-RE decoding model and MS-MS decoding model, so that the model can process high-resolution and low-resolution data at the same time, and improve the performance of the model through multi-task learning. This multi-task integration makes the algorithm maintain high decoding performance and consistency on different types of data. At the same time, since the algorithm can adapt to different types of mass spectrometry data decoding tasks through joint decoding model for mass spectrometry data, the scheme is not only suitable for current mass spectrometry decoding, but also can be extended to more extensive mass spectrometry analysis and decoding field, with high application flexibility and universality.

[0042] (4) The present application introduces noise data in the data preprocessing and decoding process, so that the model has stronger anti-noise ability. In actual application scenarios, mass spectrometry data is often disturbed by various interference, and the algorithm can effectively cope with these disturbances to improve the robustness of the model. At the same time, based on the introduction of transfer learning, the model can quickly adapt to new mass spectrometry data tasks, improving its generalization ability. BRIEF DESCRIPTION OF DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0044] Figure 1 The flowchart of the RE analysis and FTICRMS joint data decoding and optimization method based on machine learning of the embodiments of the present application;

[0045] Figure 2 The data preprocessing flowchart of the embodiments of the present application;

[0046] Figure 3A data decoding model structure schematic diagram for an embodiment of the present application adopts a CNN neural network;

[0047] Figure 4 A data set quality evaluation flowchart for an embodiment of the present application;

[0048] Figure 5 A joint decoding model structure diagram for an embodiment of the present application. DETAILED DESCRIPTION

[0049] In order to enable persons skilled in the art to better understand the schemes of the present application, the technical schemes in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by persons skilled in the art without creative labor fall within the scope of protection of the present application.

[0050] In the present application, the phrase "embodiment" means that the specific features, structures or characteristics described in combination with the embodiment can be contained in at least one embodiment of the present application. The appearance of this phrase at various places in the specification does not necessarily mean the same embodiment, nor is it an independent or alternative embodiment to other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described in the present application can be combined with other embodiments.

[0051] Embodiment 1

[0052] Please refer to Figure 1 The present embodiment provides a machine learning-based RE analysis and FTICRMS joint data decoding and optimization method, which includes the following stages: training data set acquisition and preprocessing stage, model construction and training stage, data set quality evaluation and optimization stage, joint model construction stage, and model testing and evaluation stage.

[0053] S1, training data set acquisition and preprocessing stage.

[0054] In the training data set acquisition and preprocessing stage, first, the present embodiment extracts three-dimensional crystal structure data from the PDB protein database as the main source of reaction site (RE) data. The PDB database contains three-dimensional crystal structures of different proteins, which are obtained through techniques such as X-ray crystallography, and can provide spatial position information of proteins. Secondly, the present embodiment obtains mass spectrometry data from the Acpdb database, where the Acpdb database contains mass spectrometry data of various proteins, which are obtained by mass spectrometry analyzers and are a collection of molecular mass and charge ratio information.

[0055] After the data collection is completed, the mass spectrometry data is preliminarily screened by comparing the structure of the protein with the mass spectrometry data to obtain a mass spectrometry data set, including a data set of ions and solvent molecules, a non-bridging data set, and a bridging data set. The data set of ions and solvent molecules represents a pure protein structure without other impurities. The non-bridging data set refers to a pure protein molecule data that can provide a clear mass spectrum and is helpful for accurate training of the model. The bridging data set contains impurities or other molecules combined with the protein. Although it can increase the ability of the model to process complex samples, it may also introduce noise. The non-bridging and bridging data sets distinguish whether there is a connection between different proteins. Such distinction helps to optimize the model training process and ensure the accuracy and reliability of the decoding results.

[0056] Next, as shown in Figure 2 To make these data better adapt to the training of the neural network model, it is necessary to convert the original data into a standardized data form suitable for model input. The mass spectrometry data is preprocessed as follows:

[0057] S11, the mass spectrometry data set is cut into fixed-size subgraphs using a window algorithm.

[0058] Mass spectrometry data is usually a two-dimensional or three-dimensional matrix data. By cutting it into fixed-size subgraphs, it can adapt to the input format of the convolutional neural network. The cutting method can be realized by a window-based algorithm. For example, given a mass spectrometry matrix of size n x n, a sliding window can be used to start from the top left corner and move a fixed step size to extract overlapping or non-overlapping submatrices step by step. Assuming the step size is s and the window size is k x k, the number of generated subgraphs can be determined by the following formula.

[0059] Let the size of the original matrix be m x m, then the number of subgraphs is:

[0060]

[0061] This formula indicates that when extracting subgraphs by sliding window, the number of subgraphs generated in each row and column can be calculated based on the window size k and the step size s, and then the total number is obtained. In actual operation, m represents the row and column dimensions of the mass spectrometry data, k represents the size of the subgraph, and s represents the distance moved each time. In order to improve the training efficiency of the model, the size of the subgraph should be consistent, and a size matching the network structure is usually selected, such as 64 x 64 or 128 x 128.

[0062] S12, identify the mass-to-charge ratio of water molecules, solvent molecules, and other interference data in the non-bridging data set and the bridging data set by analyzing the characteristic peak values, and filter the water molecules, solvent molecules, and other interference data.

[0063] In the process of removing water molecules and solvent molecules, the mass-to-charge ratio corresponding to the water molecules and solvent molecules can be identified by analyzing the characteristic peaks in the mass spectrum data. The peaks of the mass spectrum data represent different fragments in the protein. By analyzing these peaks, specific mass-to-charge ratio intervals can be filtered out to remove water molecules or solvent molecules. Assuming the mass of a water molecule is 18 Da and the mass of a solvent molecule is W Da, the peaks corresponding to this mass-to-charge ratio can be removed by setting the mass-to-charge ratio threshold filtering range in the mass spectrum. Let the mass spectrum data be M, where the mass-to-charge ratio of the i-th data point is m / z i , then it can be screened by the following conditions:

[0064]

[0065] where M water is the mass of the water molecule, and ∈ is an allowable error range. Similarly, the solvent molecules can be screened, and finally the mass spectrum data containing only protein structures is generated.

[0066] Other interfering data includes lipids, salt ions, additives or other impurity molecules. The process of removing these interfering data is usually filtered by setting the allowed mass-to-charge ratio range, combined with signal-to-noise ratio analysis to exclude low mass signals. In addition, by comparing with known mass spectrum database, the unmatched spectrum data is identified and removed, and finally the peaks unrelated to protein structure are confirmed and excluded by feature analysis, so as to ensure the purity of the data set and the accuracy of the analysis.

[0067] S2, model construction and training stage.

[0068] This embodiment establishes a CNN neural network to construct a data decoding model. Convolutional neural network (CNN) is used to process the pre-processed mass spectrum data picture. The input of the network is the pre-processed mass spectrum data image, and the output is the result of data quality optimization. In the construction of CNN model, the key is to determine the suitable network structure, select the number and size of convolution layer, pooling layer and full connection layer. In addition, the selection of optimizer (such as Adam or RMSProp) and the selection of loss function (such as mean square error or cross entropy) also play a crucial role in the convergence and performance of the model. In this process, the hyperparameters of the model also need to be adjusted, such as batch size, learning rate and training iteration number. Through these adjustments, the model can automatically extract features from the mass spectrum data and enhance the ability of data decoding. In order to prevent the model from overfitting, L2 regularization or Dropout technology is usually introduced.

[0069] The structure of the data decoding model constructed in the embodiment includes N convolutional blocks and a fully connected layer. The output of a convolutional block is connected to the input of the next convolutional block, and the output of the Nth convolutional block is connected to the fully connected layer. The convolutional block includes a convolutional layer and a pooling layer, and the output of the convolutional layer is connected to the pooling layer. In a more specific example, as shown in FIG. 2, two convolutional blocks are used here. Figure 3

[0070] After the construction of the basic data decoding model is completed, the basic model will be trained in two levels, namely the RE data decoding model training and the MS data decoding model training.

[0071] For the RE data decoding model training, first, an RE database is constructed to provide sufficient training samples. Second, based on the convolutional neural network (CNN), the model can automatically extract and learn key features from the RE data. Specifically, the convolutional layer in the model extracts local features in the data through a weighted convolution operation, while the pooling layer is used for downsampling to reduce the computational load. Assuming that the input data is x, and the convolution kernel is w, the output y of the convolution operation can be represented as:

[0072] y = f(x * w + b)

[0073] where * represents the convolution operation, b is the bias term, and f is the activation function. In the training process of the convolutional neural network, the model is optimized using the backpropagation algorithm, and the weights and biases of the convolution kernel are continuously adjusted by minimizing the loss function.

[0074] Similar to the RE data decoding model training, the MS data decoding model training is completed by constructing an MS database. In the MS data preprocessing, the MS mass spectrum data is first processed and converted into a picture format suitable for CNN model training. This step ensures that the model can learn different molecular features from the mass spectrum data and build an effective MS data decoding model based on this. Similar to the RE model, the evaluation of the MS model is also performed through the ROC curve and AUC value. In addition, according to different types of mass spectrum data, the cross-validation method can be used to further verify the performance of the model and ensure the generalization ability of the model.

[0075] S3, data set quality evaluation and optimization phase.

[0076] After the training in step S2 is completed, in order to ensure the effectiveness and robustness of the model, the quality evaluation and optimization of the RE model and the MS model will be carried out.

[0077] (1) Quality evaluation and optimization of the RE model.

[0078] ​The quality of the RE model is evaluated by the ROC curve and AUC value in this embodiment. The ROC curve reflects the relationship between the true positive rate and the false positive rate of the model. The ROC curve is generated based on the performance of the model in classifying samples at different thresholds.

[0079] Firstly, the model makes predictions for each sample through different decision thresholds. For each threshold, the corresponding true positive rate (TPR) and false positive rate (FPR) are calculated. The true positive rate (TPR) is defined as the proportion of correctly classified positive samples to the actual positive samples, while the false positive rate (FPR) is defined as the proportion of negative samples incorrectly classified as positive to the actual negative samples. The formulas are as follows:

[0080] True positive rate (TPR):

[0081] False positive rate (FPR):

[0082] Where TP is true positive, FP is false positive, FN is false negative, and TN is true negative.

[0083] In this step, the reaction site data is predicted by setting different decision thresholds. For each threshold, the true positive rate TPR and the false positive rate FPR of the model are calculated. These calculation results are used to draw the receiver operating characteristic curve (ROC curve), which shows the performance of the model at different thresholds and helps to intuitively evaluate the classification performance of the model. By analyzing the ROC curve, the optimal threshold can be determined to optimize the prediction ability of the model.

[0084] Secondly, the AUC value is calculated based on the obtained ROC curve, which is used to determine whether the current trained RE is effective and robust, and the best RE model is selected based on the judgment result. The AUC value is the area under the ROC curve, and its value range is [0, 1]. The closer the AUC value is to 1, the better the performance of the model, which can correctly distinguish between positive and negative samples; on the contrary, the closer the AUC value is to 0, the worse the performance of the model. When the AUC value is 0.5, it means that the model has no classification ability, and its performance is the same as random guessing; when the AUC value is greater than 0.5 and less than 1, the performance of the model is better than random guessing, but the specific degree of superiority needs to be judged according to the size of the AUC value; when the AUC value is equal to 1, it means that the model has perfect classification ability, i.e. all positive examples are correctly predicted as positive examples, and all negative examples are correctly predicted as negative examples.

[0085] (2) Quality evaluation and optimization of MS model.

[0086] As Figure 4As shown, data set quality evaluation and optimization is a key step in mass spectrometry data decoding algorithm, directly related to the accuracy and stability of the model. In the MS mass spectrometry data decoding process, the data is large and complex, and must be strictly evaluated to ensure the effectiveness of the data. The evaluation program mainly analyzes the resolution, peak intensity, signal-to-noise ratio and other indicators of the mass spectrometry data to comprehensively analyze the high-resolution data in the RE database and the MS database.

[0087] First, the resolution of the data. The resolution of the data directly affects the accuracy of the decoding, the higher the resolution, the more detailed information can be captured. Therefore, the evaluation program calculates the resolution of the data to ensure the accuracy of the high-resolution data.

[0088] Second, peak characteristics. The evaluation of peak characteristics, that is, by analyzing the shape, position and intensity of each peak in the mass spectrum, to determine whether it is consistent with the expected biological molecule structure. By calculating the matching degree of the peak characteristics, the program can find abnormal data and mark it as low-quality data, the formula is as follows:

[0089]

[0090] Where I observed is the observed peak intensity, I expected is the expected peak intensity, and N is the number of peaks. Data with low matching degree will be marked as "not similar".

[0091] For these low-quality data, the evaluation program can take two processing methods: first, exclude it from the training data set, thereby improving the training quality of the model. Second, generate and add a certain amount of low-quality data to the training set, so that the model can perform more robustly when facing noisy data. This method enables the model to adapt to complex situations in actual applications, improving decoding accuracy and anti-interference ability. For example, in protein mass spectrometry data, adding a small amount of noise samples can help the model better cope with data anomalies in real-world scenarios.

[0092] S4, joint model construction stage.

[0093] First, in order to solve the main decoding task and the enhancement decoding task, the best RE model and MS model are combined in this embodiment. In order to realize the main decoding task, the RE model and the MS model are combined, the output of the MS model can provide basic information, and the RE model can make inferences or predictions based on these basic information, so as to infer the reactive sites of the molecule through the mass spectrometry information; in order to realize the enhancement decoding task, multiple MS models are combined, and the fragment relationship in the molecular structure is analyzed layer by layer through the mass-to-charge ratio information of the parent ion and the daughter ion. In this way, MS-RE decoding model and MS-MS decoding model are obtained respectively.

[0094] Next, the joint decoding model is constructed. The construction of the MS-FTICRMS joint algorithm is a key step in this scheme. In this embodiment, a joint data decoding model is constructed by taking the MS-RE decoding model as the backbone and multiple MS-MS decoding models as additional modules. This joint decoding model can handle both high-resolution and low-resolution mass spectrometry data and use the MS quality evaluation program to filter the decoding results. Specifically, the data is initially decoded by the MS-RE decoding model, and then further refined using the results of the MS-MS decoding model. "Refinement" refers to the further fine-tuning of the initial decoding results using the results of the MS-MS decoding model, rather than simply integrating them. This process improves the accuracy and detail of the decoding by analyzing and optimizing the output of the MS-MS decoding model. Therefore, Figure 5 "refinement decoding" in should be more explicitly referred to as the optimization of the initial decoding results, rather than the simple integration of the two results. In addition, Figure 5 the two "refinement decodings" in refer to the first optimization of the initial decoding results and the second deeper refinement on the integrated joint decoding model, ensuring the accuracy and completeness of the final decoding.

[0095] By constructing a joint model, effective conversion can be performed on low-resolution data, making it have performance close to high-resolution data. This process combines data feature extraction and quality evaluation tasks to ensure the accuracy and consistency of the decoding results. During training, the joint model is trained through a multi-task learning framework, with the MS-RE decoding task as the backbone and the MS-MS task as an additional module to further enhance the model's performance.

[0096] S5, model testing and evaluation phase.

[0097] As shown in Figure 5 , in the model testing and evaluation phase, the MS-FTICRMS joint decoding model constructed is tested using an independent test set, and the decoding performance is evaluated through actual mass spectrometry data and protein structure data. By comparing the model output results with the real mass spectrometry data, the accuracy and efficiency of the model in decoding mass spectrometry data can be obtained. In addition, by comparing with existing benchmark models, the advantages and disadvantages of the MS-FTICRMS joint decoding model can be comprehensively evaluated. For the performance evaluation of the model, in addition to the ROC curve and AUC value, other evaluation indicators such as accuracy, recall rate, F1 value, etc. can be introduced to obtain more comprehensive performance feedback.

[0098] To ensure the performance of the joint decoding model in practical applications, further hyperparameter optimization and the introduction of transfer learning can be conducted. The core of hyperparameter optimization is to find the optimal combination of model parameters, which can be achieved through grid search or Bayesian optimization. Grid search is an exhaustive method that tries every possible combination of parameters, calculates the performance of the model under each combination, and selects the optimal combination. Assuming there are two hyperparameters: learning rate and batch size, grid search will traverse all combinations within the different value ranges of these parameters and evaluate their performance. Bayesian optimization, on the other hand, gradually approaches the optimal parameters by constructing a probability model, which is more efficient. The formula is as follows:

[0099] x next = argmax(μ(x) + kσ(x))

[0100] where μ(x) is the mean of the objective function, σ(x) is the uncertainty of the objective function, and k is a parameter that balances exploration and exploitation. Through Bayesian optimization, the optimal parameter combination can be found with less computational effort, improving the decoding accuracy of the model.

[0101] In addition, transfer learning techniques can also be used to optimize the training process of the joint decoding model. Transfer learning transfers learned features from other similar datasets to new tasks through pre-training. This not only speeds up the training process, but also maintains the model's generalization ability in cases where data is scarce. Specifically, a preliminary model is first trained on a large-scale mass spectrometry database, and then fine-tuned on a specific task dataset. For example, if the model has been pre-trained on some protein mass spectrometry data, when it is applied to a new task, the model can quickly adapt to different mass spectrometry features, significantly improving accuracy and reducing training time.

[0102] Example 2

[0103] Example 2 is based on Step S1 of Example 1 and provides a more specific example.

[0104] Suppose a set of protein mass spectrometry data is obtained from the Acpdb database, where the mass spectrometry graph of a protein is 256x256 in size, and there are a large number of water molecules and solvent molecules. Through the preprocessing script, we cut the mass spectrometry graph into 64x64 subgraphs, with a moving step size of 32. Then, according to the formula, we can calculate:

[0105]

[0106] Therefore, 81 subgraphs can be extracted from the mass spectrometry data. Next, by filtering the mass-to-charge ratio, we remove the mass range corresponding to water molecules, and use the preprocessed mass spectrometry data for subsequent model training.

[0107] This step can provide a systematic method to process mass spectrum graphs containing interference components by cutting them into small subgraphs and filtering unnecessary water molecules and solvent molecules, thereby enhancing the quality and availability of the data. This process lays the foundation for subsequent model training, improves the model's ability to recognize and decode target proteins, and demonstrates the importance of data preprocessing in mass spectrometry.

[0108] In addition, in order to balance the computational load of decoding and the robustness of decoding, the decoding model of embodiment 2 adopts an MS-RE decoding model and an MS-MS decoding model composed of a two-layer MS model.

[0109] The other steps of embodiment 2 are similar to those of embodiment 1 and will not be described here.

[0110] It should be noted that for the foregoing method embodiments, they are all described as a combination of a series of actions for the sake of simple description, but those skilled in the art should know that the present application is not limited by the order of the described actions, because according to the present application, certain steps can be performed in other orders or simultaneously.

[0111] The technical features of the above embodiments can be combined in any way. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.

[0112] The above embodiments are the preferred embodiments of the present application, but the embodiments of the present application are not limited by the above embodiments, and any changes, modifications, substitutions, combinations, simplifications made without departing from the spirit and principles of the present application shall be equivalent replacement methods and shall be within the scope of protection of the present application.

Claims

1. A method for joint data decoding and optimization based on machine learning-based RE analysis and FTICRMS, characterized in that, Includes the following steps: Protein crystal structure data and mass spectrometry data are collected, and the data are initially screened and preprocessed to obtain a training dataset; the training dataset includes reaction site data and mass spectrometry sub-map datasets; A data decoding model is constructed, and the data decoding model is trained using reaction site data. Key local features of the reaction site data are automatically extracted and learned to obtain the first data decoding model. The data decoding model is then trained using a mass spectrometry sub-map dataset to learn the molecular features of different proteins and obtain the second data decoding model. The first data decoding model is evaluated using feature curves to obtain a first evaluation result; the second data decoding model is comprehensively evaluated using quality evaluation indicators to obtain a second evaluation result, wherein the quality evaluation indicators include resolution and peak features. The model is optimized using the first and second evaluation results respectively to obtain the optimal first and second data decoding models. The optimal second data decoding model is combined with the optimal first data decoding model to obtain a first combined decoding model. Multiple optimal second data decoding models are combined to obtain a second combined decoding model. A joint decoding model is constructed using the first combined decoding model as the backbone and the second combined decoding model as an additional module, wherein the output of the second combined decoding model is connected to the output of the first combined decoding model. Specifically, the construction of the joint decoding model using the first combined decoding model as the backbone and the second combined decoding model as an additional module involves: The sample data is initially decoded using the first combined decoding model, and then further decoded using the second combined decoding model. The initial decoding result and the refined decoding result are then concatenated to obtain the joint data decoding model. The joint decoding model is tested using a given independent test set, and its decoding performance is evaluated using actual mass spectrometry data and protein structure data. Based on the evaluation results, the optimal joint decoding model is obtained, and the optimal joint decoding model is used to decode the protein mass spectrometry data.

2. The method for joint data decoding and optimization based on machine learning RE analysis and FTICRMS as described in claim 1, characterized in that, The preliminary screening and preprocessing of the data includes: X-ray crystallography is used to extract spatial location information from protein crystal structure data to obtain protein reaction site data. Mass spectrometry data are initially screened by comparing protein structures to obtain a mass spectrometry dataset, including sets of ionless and solvent molecules, non-bridging datasets, and bridging datasets. The sets of ionless and solvent molecules contain pure protein structures, while the non-bridging datasets and bridging datasets contain impurities. The mass spectrometry dataset is divided into fixed-size sub-maps using a windowing algorithm. By analyzing characteristic peaks, the mass-charge ratios of water molecules, solvent molecules, and other interfering data in non-bridged datasets and bridged datasets are identified. Water molecules, solvent molecules, and other interfering data are then filtered out, as shown in the following formula: Let the mass spectrometry data be M, and the mass-charge ratio of the i-th data point be m / z. i M water The mass of the water molecule is ∈, which is a permissible error range. Other interfering data include lipids, salt ions, additives, or other impurity molecules.

3. The method for joint data decoding and optimization based on machine learning RE analysis and FTICRMS as described in claim 1, characterized in that, The data decoding model includes N convolutional blocks and fully connected layers. The output of a convolutional block is connected to the input of the next convolutional block, and the output of the Nth convolutional block is connected to the fully connected layer. Each convolutional block includes a convolutional layer and a pooling layer, and the output of the convolutional layer is connected to the pooling layer.

4. The method for joint data decoding and optimization based on machine learning RE analysis and FTICRMS as described in claim 1, characterized in that, The evaluation of the first data decoding model using feature curves includes: Predictions are made for each sample in the response site data using different decision thresholds. For each threshold, the corresponding true positive rate and false positive rate are calculated. Based on the true positive rate and false positive rate of each threshold, an ROC curve is generated. The robustness of the current model is assessed by calculating the AUC value based on the ROC curve, and the first evaluation result is generated.

5. The method for joint data decoding and optimization based on machine learning RE analysis and FTICRMS as described in claim 1, characterized in that, The comprehensive evaluation of the second data decoding model using quality assessment indicators includes: Calculate the resolution and signal-to-noise ratio of the data based on the mass spectrometry subplot dataset; Analyze the peak features of each sub-mass spectrum to obtain sample peak features, calculate the matching degree between sample peak features and target peak features, and mark data with low matching degree as low quality data; the sample peak features include peak shape, peak position and peak intensity.

6. The method for joint data decoding and optimization based on machine learning RE analysis and FTICRMS as described in claim 5, characterized in that, The matching degree is calculated as follows: Among them, I observed For the observed peak intensity, I expected N represents the expected peak intensity, and N represents the number of peaks.

7. The method for joint data decoding and optimization based on machine learning RE analysis and FTICRMS as described in claim 5, characterized in that, This includes optimizing the model using the results of the second evaluation, specifically: Exclude low-quality data; Add low-quality data to the training dataset.

8. The method for joint data decoding and optimization based on machine learning RE analysis and FTICRMS as described in claim 1, characterized in that, The evaluation of decoding performance using actual mass spectrometry data and protein structure data specifically includes: The test output results are compared with real mass spectrometry data, and the performance of the comparison results is evaluated using performance indicators. The best joint decoding model is selected based on the evaluation results. The performance indicators are ROC curves and their AUC values, accuracy, recall, and F1 values.

9. The method for joint data decoding and optimization based on machine learning RE analysis and FTICRMS as described in claim 1, characterized in that, It also includes optimizing the joint decoding model using hyperparameter optimization and transfer learning; the hyperparameter optimization includes grid search and Bayesian optimization.

Citation Information

Patent Citations

  • Multiplexed bead arrays for proteomics

    CN111263886A

  • AI polypeptide structure identification method based on Transform model and application thereof

    CN118298927A