Protein mass spectrum data analysis method and system based on protein stability

By calculating the local noise variance and Bayesian probability correction, combined with adaptive Gaussian filtering and fuzzy membership, the problem of effective peptide signal peak identification of protein mass spectrometry data under non-stationary background noise conditions is solved, and accurate peptide signal peak extraction is achieved.

CN120636550AInactive Publication Date: 2025-09-12唐韵
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510689455.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Under the non-stationary state of background noise, it is difficult to accurately identify the effective peptide signal peaks in protein mass spectrometry data in existing mass spectrometry data analysis, resulting in the generation of false positive peaks and the loss of key proteins.

Method used

By calculating the local noise variance of each data point in protein mass spectrometry data, determining the scale constraint parameters during smoothing, combining Bayesian probability for baseline correction, extracting edge-retained data of stable protein characteristic peaks, and using adaptive Gaussian filtering and fuzzy membership to identify peptide signal peaks.

Benefits of technology

Under the non-stationary background noise state, the effective peptide signal peaks in protein mass spectrometry data can be effectively identified and extracted, avoiding false peak identification and improving the stability and accuracy of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636550A_ABST
    Figure CN120636550A_ABST
Patent Text Reader

Abstract

The invention provides a protein mass spectrum data analysis method and system based on protein stability. The method comprises the following steps: determining scale constraint parameters of a sliding window during smoothing of protein mass spectrum data according to local noise variances at data points in the protein mass spectrum data; carrying out loss constraint on the edge of a protein stability characteristic peak in the protein mass spectrum data based on the scale constraint parameter in combination with a preset sliding window to obtain edge retention data of the protein stability characteristic peak; determining a fuzzy membership degree of each data point belonging to a protein stability characteristic peak according to a plurality of pre-identification peaks in the edge retention data and spatial distribution characteristics of each data point in the edge retention data; and performing baseline correction on the edge retention data based on Bayesian probability in combination with all fuzzy membership degrees, and further extracting peptide fragment signal peaks of the protein. According to the technical scheme provided by the invention, the effective peptide fragment signal peak in the protein mass spectrum data can be analyzed in a non-stationary background noise state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of mass spectrometry data analysis, and more specifically, to a protein mass spectrometry data analysis method and system based on protein stability. Background Art

[0002] Mass spectrometry is a high-precision analytical technique. Its principle is to ionize the sample, separate and detect it according to the mass-to-charge ratio (m / z) of different ions, and thus obtain the molecular weight and structural information of the substance. With the continuous advancement of science and technology, modern mass spectrometers combine multiple ionization methods and mass analyzers, which can handle complex samples and achieve high-throughput analysis. At the same time, mass spectrometry data analysis technology is also becoming increasingly mature. With the help of advanced software algorithms and artificial intelligence applications, the automated processing and in-depth mining of mass spectrometry data have further improved the efficiency and accuracy of identification, providing strong support for systems biology research and providing strong support for scientific research and practical applications.

[0003] In existing mass spectrometry data analysis, sample information is mainly obtained by analyzing the ion peaks in the mass spectrum. In the mass spectrum, first, the main peaks in the mass spectrum are identified. These peaks correspond to different ions in the sample. By calculating the mass-to-charge ratio of the peaks, the molecular weight of the ions can be inferred. Further analysis of the isotope distribution pattern of the peaks can verify the chemical composition of the molecules. However, in protein mass spectrometry data analysis, low-abundance peptides have a signal-to-noise ratio close to the background noise, and the background noise is not completely stable. Due to the random fluctuations of the background noise, "spikes" similar to peptide signals are formed (such as the tail fluctuations of Gaussian noise). When the intensity of these false peaks accidentally approaches the low-abundance true signal, they may be mistakenly identified as valid peaks, resulting in false-positive peptide signal peaks, which in turn leads to the loss of key regulatory proteins. Therefore, how to analyze the valid peptide signal peaks in protein mass spectrometry data under the non-stationary background noise state has become a difficult problem faced by the industry. Summary of the Invention

[0004] The present application provides a protein mass spectrometry data analysis method and system based on protein stability, which can analyze effective peptide signal peaks in protein mass spectrometry data under non-stationary background noise conditions.

[0005] In a first aspect, the present application provides a method for analyzing protein mass spectrometry data based on protein stability, comprising the following steps: Obtaining mass spectrometry data of the protein to be analyzed; calculating the local noise variance at each data point in the protein mass spectrometry data, and determining a scale constraint parameter of a sliding window when smoothing the protein mass spectrometry data according to the local noise variance at each data point; Based on the scale constraint parameter and a preset sliding window, loss constraint is performed on the edge of the protein stable characteristic peak in the protein mass spectrometry data, thereby obtaining edge retention data of the protein stable characteristic peak; Pre-identifying protein stability characteristic peaks on the edge-retained data to obtain a plurality of pre-identified peaks, and determining the fuzzy membership of each data point to the protein stability characteristic peak based on all the pre-identified peaks combined with the spatial distribution characteristics of each data point in the edge-retained data; The edge-retained data are baseline-corrected based on Bayesian probability combined with the fuzzy membership of each data point to a stable characteristic peak of the protein, and then the peptide signal peak of the protein is extracted from the edge-retained data after baseline correction.

[0006] In some embodiments, calculating the local noise variance at each data point in the protein mass spectrometry data specifically includes: Get the preset neighborhood radius; Performing local area division on each data point in the protein mass spectrometry data based on the neighborhood radius, thereby obtaining a neighborhood partition corresponding to each data point; For each data point, the local noise variance at the data point is determined according to the neighborhood partition corresponding to the data point, thereby obtaining the local noise variance at each data point in the protein mass spectrometry data.

[0007] In some embodiments, determining the scale constraint parameter of the sliding window when smoothing the protein mass spectrometry data based on the local noise variance at each data point specifically includes: Normalizing and mapping the local noise variance of each data point to generate a window scale coefficient for each data point; Setting the upper limit and lower limit of the scale constraint of the sliding window when smoothing the protein mass spectrometry data based on all window scale coefficients; The scale constraint parameter of the sliding window when smoothing the protein mass spectrometry data is determined according to the scale constraint upper limit and the scale constraint lower limit.

[0008] In some embodiments, performing loss constraints on the edges of protein stable characteristic peaks in the protein mass spectrometry data based on the scale constraint parameters in combination with a preset sliding window, thereby obtaining edge retention data of the protein stable characteristic peaks specifically includes: Dynamically adjusting the size of a preset sliding window at each data point in the protein mass spectrometry data according to the scale constraint parameter, thereby obtaining an adaptive window corresponding to each data point; Determine a smoothing weight factor corresponding to each data point according to the adaptive window of each data point; Calculating the absolute value of the first derivative of a data point in the protein mass spectrometry data as an edge intensity indicator of whether the data point belongs to the edge of a stable characteristic peak of the protein; generating an edge retention coefficient matrix for performing loss constraints on the edges of protein stable characteristic peaks in the protein mass spectrometry data based on the smoothing weight factor and edge strength index of each data point; Adaptive Gaussian filtering is performed on the protein mass spectrum data using the edge retention coefficient matrix to obtain edge retention data of stable characteristic peaks of the protein.

[0009] In some embodiments, pre-identifying protein stability characteristic peaks on the edge retention data to obtain a plurality of pre-identified peaks specifically comprises: Extracting local maximum points in the edge-preserving data as candidate peak vertices; Calculate the zero-crossing points of the first-order derivatives on the left and right sides of each candidate peak vertex, and then construct the candidate peak corresponding to each candidate peak vertex according to the zero-crossing points of the first-order derivatives on both sides; For each candidate peak, the signal-to-noise ratio of the signal integral value at the peak apex within the candidate peak to the noise intensity of the adjacent baseline region is determined, thereby obtaining the signal-to-noise ratio corresponding to each candidate peak; Multiple pre-identified peaks were extracted based on all signal-to-noise ratios.

[0010] In some embodiments, determining the fuzzy membership of each data point to a protein stabilization characteristic peak based on all pre-identified peaks combined with the spatial distribution characteristics of each data point in the edge-preserving data specifically includes: Extracting the identification peak center of each pre-identified peak; determining a spatial distribution characteristic of each data point in the edge-preserving data; For each data point, the spatial correlation parameter between the data point and the pre-identified peak is determined based on the spatial distribution characteristics of the data point and the identification peak center of the nearest pre-identified peak, thereby obtaining the spatial correlation parameter between each data point and the corresponding pre-identified peak; The fuzzy membership of each data point to the protein stable characteristic peak is determined based on all spatial correlation parameters.

[0011] In some embodiments, mass spectrometry data of the protein to be analyzed is acquired by a mass spectrometer.

[0012] In a second aspect, the present application provides a protein mass spectrometry data analysis system based on protein stability, comprising: An acquisition module, used to acquire mass spectrometry data of proteins to be analyzed; a processing module, configured to calculate the local noise variance at each data point in the protein mass spectrometry data, and determine a scale constraint parameter of a sliding window when smoothing the protein mass spectrometry data according to the local noise variance at each data point; The processing module is further configured to perform loss constraints on the edges of protein stable characteristic peaks in the protein mass spectrometry data based on the scale constraint parameters in combination with a preset sliding window, thereby obtaining edge-retained data of the protein stable characteristic peaks; The processing module is further configured to pre-identify protein stability characteristic peaks on the edge-retained data to obtain a plurality of pre-identified peaks, and determine the fuzzy membership of each data point to the protein stability characteristic peak based on all the pre-identified peaks combined with the spatial distribution characteristics of each data point in the edge-retained data; The execution module is used to perform baseline correction on the edge-retained data based on Bayesian probability combined with the fuzzy membership of each data point to the stable characteristic peak of the protein, and then extract the peptide signal peak of the protein from the edge-retained data after baseline correction.

[0013] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a code, and the processor is configured to obtain the code and execute the above-mentioned protein mass spectrometry data analysis method.

[0014] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned protein mass spectrometry data analysis method is implemented.

[0015] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects: In the protein mass spectrometry data analysis method and system based on protein stability provided in the present application, first, the protein mass spectrometry data to be analyzed is obtained; secondly, the local noise variance at each data point in the protein mass spectrometry data is calculated, and the scale constraint parameter of the sliding window when smoothing the protein mass spectrometry data is determined based on the local noise variance at each data point; further, based on the scale constraint parameter and a preset sliding window, the edge of the protein stability characteristic peak in the protein mass spectrometry data is loss constrained, thereby obtaining edge-retained data of the protein stability characteristic peak; then, the protein stability characteristic peak is pre-identified on the edge-retained data to obtain multiple pre-identified peaks, and the fuzzy membership of each data point to the protein stability characteristic peak is determined based on all the pre-identified peaks and the spatial distribution characteristics of each data point in the edge-retained data; finally, the edge-retained data is baseline-corrected based on the Bayesian probability and the fuzzy membership of each data point to the protein stability characteristic peak, thereby extracting the peptide signal peak of the protein from the baseline-corrected edge-retained data.

[0016] It can be seen that the present application can analyze the effective peptide signal peaks in the protein mass spectrometry data under the non-stationary background noise state; first, the local noise variance at each data point in the protein mass spectrometry data is calculated, and the scale constraint parameters of the sliding window when smoothing the protein mass spectrometry data are determined according to the local noise variance at each data point, which can effectively adjust the smoothing effect of the protein mass spectrometry data to achieve fine-grained retention of high-noise areas and effective noise reduction in low-noise areas, thereby taking into account the integrity of the characteristic peak edges and the smoothness of the background signal as a whole; secondly, based on the scale constraint parameters and the preset sliding window, the edges of the stable characteristic peaks of the protein in the protein mass spectrometry data are loss-constrained, and the edge-retained data of the stable characteristic peaks of the protein are obtained. The edge-retained data not only suppresses the background noise, but also retains the structural characteristics of the characteristic peak mutation area to the greatest extent, providing a basis for subsequent accurate characteristic peak identification and quantitative analysis, thereby avoiding the background noise. Peak recognition deviation caused by random fluctuations; then, the edge-retained data is pre-identified as stable characteristic peaks of the protein to obtain multiple pre-identified peaks, and the fuzzy membership of each data point to the stable characteristic peak of the protein is determined based on all the pre-identified peaks combined with the spatial distribution characteristics of each data point in the edge-retained data, so that each data point can belong to multiple pre-identified peaks to varying degrees at the same time, thereby enhancing the stability and continuity of peak recognition, and then effectively processing mass spectrometry data with low signal-to-noise ratio under complex backgrounds, thereby avoiding identifying false peaks generated when the noise intensity accidentally approaches the low-abundance real signal as valid peaks; finally, the edge-retained data is baseline-corrected based on the Bayesian probability combined with the fuzzy membership of each data point to the stable characteristic peak of the protein, and then the peptide signal peak of the protein is extracted from the edge-retained data after baseline correction; in summary, the technical solution provided by the present application can analyze the effective peptide signal peaks in protein mass spectrometry data under the non-stationary background noise state. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a schematic diagram of an application scenario architecture of a protein mass spectrometry data analysis method based on protein stability according to some embodiments of the present application; Figure 2 is an exemplary flow chart of a protein mass spectrometry data analysis method based on protein stability according to some embodiments of the present application; Figure 3 is an exemplary flow chart of determining local noise variance according to some embodiments of the present application; Figure 4 is an exemplary flow chart of determining a plurality of pre-identified peaks according to some embodiments of the present application; Figure 5is a schematic structural diagram of a protein mass spectrometry data analysis system based on protein stability according to some embodiments of the present application; Figure 6 It is a schematic structural diagram of a computer device for implementing a protein mass spectrometry data analysis method based on protein stability according to some embodiments of the present application. DETAILED DESCRIPTION

[0018] In order to better understand the technical solution of the present application, the technical solution of the present application will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0019] refer to Figure 1 , this figure is a schematic diagram of an application scenario architecture of a rehabilitation training data processing method according to some embodiments of the present application, the application scenario architecture includes a user terminal, a communication network and a server side, the user terminal and the server side are directly or indirectly connected through the communication network, the user terminal can trigger a request for analyzing and processing the protein mass spectrometry data to be analyzed, and the server side responds to the processing request to obtain the protein mass spectrometry data to be analyzed; calculate the local noise variance at each data point in the protein mass spectrometry data, and determine the scale constraint parameter of the sliding window when smoothing the protein mass spectrometry data based on the local noise variance at each data point; based on the scale constraint parameter combined with the preset A sliding window is used to perform loss constraints on the edges of protein stable characteristic peaks in the protein mass spectrometry data, thereby obtaining edge-retained data of protein stable characteristic peaks; protein stable characteristic peaks are pre-identified on the edge-retained data to obtain multiple pre-identified peaks, and the fuzzy membership of each data point to the protein stable characteristic peak is determined based on all pre-identified peaks combined with the spatial distribution characteristics of each data point in the edge-retained data; the edge-retained data is baseline-corrected based on Bayesian probability combined with the fuzzy membership of each data point to the protein stable characteristic peak, and the peptide signal peak of the protein is extracted from the baseline-corrected edge-retained data, and the peptide signal peak is returned to the user terminal.

[0020] refer to Figure 2 , which is an exemplary flow chart of a protein mass spectrometry data analysis method based on protein stability according to some embodiments of the present application. The protein mass spectrometry data analysis method 100 based on protein stability mainly includes the following steps: In step 101, mass spectrometry data of a protein to be analyzed is obtained.

[0021] In a specific implementation, the mass spectrometry data of the protein to be analyzed can be obtained by a mass spectrometer, and the protein mass spectrometry data is generated by the mass spectrometer performing ionization, mass separation and detection on a protein sample.

[0022] It should be noted that the protein mass spectrometry data in this application represents two-dimensional data reflecting the signal intensity of peptides in proteins at different mass-to-charge ratios. Specifically, the protein mass spectrometry data is mass spectrum data with mass-to-charge ratio as the horizontal axis and signal intensity as the vertical axis, which is used to reflect the abundance distribution of peptides in proteins at different mass-to-charge ratios.

[0023] In step 102, the local noise variance at each data point in the protein mass spectrometry data is calculated, and the scale constraint parameter of the sliding window when smoothing the protein mass spectrometry data is determined according to the local noise variance at each data point.

[0024] It should be noted that the local noise variance in this application represents the variance value reflecting the degree of signal intensity fluctuation in the small range of each data point in the protein mass spectrometry data. It is used to measure the noise level in the local area of ​​the data point, and can reflect the uncertainty and fluctuation of the mass spectrometry signal at this point, thereby providing a basis for the subsequent adaptive adjustment of the smoothing window scale and improving the accuracy of characteristic peak identification.

[0025] In some embodiments, reference Figure 3 As shown in FIG. 1 , this figure is an exemplary flow chart of determining the local noise variance according to some embodiments of the present application. In this embodiment, the local noise variance at each data point in the protein mass spectrometry data can be calculated by the following steps: First, in step 1021, a preset neighborhood radius is obtained; Then, in step 1022, each data point in the protein mass spectrum data is divided into a local area based on the neighborhood radius, thereby obtaining a neighborhood partition corresponding to each data point; Finally, in step 1023, for each data point, the local noise variance at the data point is determined according to the neighborhood partition corresponding to the data point, thereby obtaining the local noise variance at each data point in the protein mass spectrometry data.

[0026] In a specific implementation, first, a preset neighborhood radius is obtained. The neighborhood radius can be set to a size of 5-10 data points according to the data resolution of the protein mass spectrometry data. For example, in the present application, the neighborhood radius can be set to a size of 6 data points. In addition, in other embodiments, it can also be set to a domain radius of other sizes, which is not limited here; then, for each data point in the protein mass spectrometry data, the range of the neighborhood radius divided with the data point as the center is used as the domain partition corresponding to the data point, and then the neighborhood partition corresponding to each data point is obtained; finally, for each data point, the local noise variance at the data point is determined according to the neighborhood partition corresponding to the data point, that is: for each data point, the variance of the data points in the neighborhood partition corresponding to the data point is calculated as the local noise variance at the data point, and then the local noise variance at each data point in the protein mass spectrometry data is obtained, that is, the noise level is characterized by the degree of discreteness of the data in the neighborhood partition.

[0027] It should be noted that the neighborhood radius in this embodiment represents the parameter value for dividing the neighborhood area; the neighborhood partition in this embodiment represents a continuous data interval divided with each data point in the protein mass spectrometry data as the center, which is used to analyze the noise fluctuation around the data point.

[0028] In some embodiments, determining the scale constraint parameter of the sliding window when smoothing the protein mass spectrometry data based on the local noise variance at each data point can be specifically implemented by the following steps, namely: Normalizing and mapping the local noise variance of each data point to generate a window scale coefficient for each data point; Setting the upper limit and lower limit of the scale constraint of the sliding window when smoothing the protein mass spectrometry data based on all window scale coefficients; The scale constraint parameter of the sliding window when smoothing the protein mass spectrometry data is determined according to the scale constraint upper limit and the scale constraint lower limit.

[0029] It should be noted that the scale constraint parameter in this application represents the window scale used for constrained smoothing processing for each data point. By determining the scale constraint parameter, it is possible to effectively achieve fine-grained retention of high-noise areas and effective noise reduction in low-noise areas, thereby taking into account the integrity of the characteristic peak edge and the smoothness of the background signal as a whole.

[0030] In a specific implementation, first, the local noise variance of each data point is normalized and mapped to generate the window scale coefficient of each data point, that is, the local noise variance of each data point is normalized to between 0 and 1 by minimum-maximum normalization, and the normalized local noise variance is used as the window scale coefficient of the corresponding data point, thereby obtaining the window scale coefficient of each data point; then, the minimum window scale coefficient and the maximum window scale coefficient are extracted from all the window scale coefficients, and the scale constraint upper limit of the sliding window when the protein mass spectrometry data is smoothed is set according to the maximum window scale coefficient, and the scale constraint lower limit of the sliding window when the protein mass spectrometry data is smoothed is set according to the minimum window scale coefficient. For example, when the maximum window scale coefficient is higher than 0.5, the protein mass spectrometry is set. When the data is smoothed, the upper limit of the scale constraint of the sliding window is 15 data points. When the maximum window scale coefficient is lower than 0.5, the upper limit of the scale constraint of the sliding window when the protein mass spectrometry data is smoothed is set to 12 data points. When the minimum window scale coefficient is higher than 0.5, the lower limit of the scale constraint of the sliding window when the protein mass spectrometry data is smoothed is set to 8 data points. When the minimum window scale coefficient is lower than 0.5, the upper limit of the scale constraint of the sliding window when the protein mass spectrometry data is smoothed is set to 6 data points. Finally, the range span between the scale constraint upper limit and the scale constraint is used as the scale constraint parameter of the sliding window when the protein mass spectrometry data is smoothed, and the range span is the difference between the scale constraint upper limit and the scale constraint.

[0031] It should be noted that, in this embodiment, the window scale coefficient represents a mapping value of the sliding window size dynamically generated by the local noise variance of each data point; in this embodiment, the scale constraint upper limit represents the maximum constrained size of the sliding window when the protein mass spectrometry data is smoothed; in this embodiment, the scale constraint lower limit represents the minimum constrained size of the sliding window when the protein mass spectrometry data is smoothed. By determining the scale constraint upper limit and the scale constraint lower limit, the sliding window at different data points can be more effectively and dynamically adjusted, so as to effectively achieve fine-grained retention of high-noise areas and effective noise reduction in low-noise areas.

[0032] In step 103, loss constraints are performed on the edges of the protein stable characteristic peaks in the protein mass spectrometry data based on the scale constraint parameters and a preset sliding window, thereby obtaining edge-preserved data of the protein stable characteristic peaks.

[0033] It should be noted that in this application, the protein stable characteristic peak refers to a stable characteristic peak with a significant signal mutation edge in the protein mass spectrometry data, and the edge-retained data refers to the data obtained after retaining the edge of the protein stable characteristic peak in the protein mass spectrometry data. Specifically, the edge-retained data refers to the data set containing the boundary information of the protein stable characteristic peak retained by dynamically adjusting the sliding window according to the scale constraint parameters during the smoothing process of the protein mass spectrometry data, and detecting and protecting the signal changes at the edge of the characteristic peak. The edge-retained data not only suppresses the background noise, but also retains the structural characteristics of the characteristic peak mutation area to the greatest extent, providing a basis for subsequent accurate characteristic peak identification and quantitative analysis.

[0034] In some embodiments, based on the scale constraint parameter and a preset sliding window, the edge of the protein stable characteristic peak in the protein mass spectrometry data is subjected to loss constraint, thereby obtaining the edge retention data of the protein stable characteristic peak. Specifically, the following steps can be used, namely: Dynamically adjusting the size of a preset sliding window at each data point in the protein mass spectrometry data according to the scale constraint parameter, thereby obtaining an adaptive window corresponding to each data point; Determine a smoothing weight factor corresponding to each data point according to the adaptive window of each data point; Calculating the absolute value of the first derivative of a data point in the protein mass spectrometry data as an edge intensity indicator of whether the data point belongs to the edge of a stable characteristic peak of the protein; generating an edge retention coefficient matrix for performing loss constraints on the edges of protein stable characteristic peaks in the protein mass spectrometry data based on the smoothing weight factor and edge strength index of each data point; Adaptive Gaussian filtering is performed on the protein mass spectrum data using the edge retention coefficient matrix to obtain edge retention data of stable characteristic peaks of the protein.

[0035] In a specific implementation, first, the minimum scale of the sliding window and the window scale coefficient corresponding to each data point are obtained, and the window scale coefficient is consistent with that in step 102. An adjustment function for adjusting the preset sliding window is constructed according to the minimum scale and the scale constraint parameter, and the window scale coefficient corresponding to each data point is input into the adjustment function as an input parameter. The result output by the adjustment function is used as the adaptive window size corresponding to each data point. The preset sliding window is adjusted by the adaptive window size, and then the adaptive window corresponding to each data point is obtained. According to the principle that a small window is used to retain details in a low-noise area and a large window is used to suppress interference in a high-noise area, the adjustment function is constructed as follows: adjustment function = minimum scale + window scale coefficient × scale constraint parameter. In addition, the scale constraint lower limit in step 102 can be used as the minimum scale, and the preset sliding window is a window of any size. Secondly, for each data point, the size of the adaptive window corresponding to the data point can be normalized to between 0 and 1 by minimum-maximum normalization, and the normalized value is used as the smoothing weight factor corresponding to the data point, that is, the larger the adaptive window, the larger the smoothing weight factor. Further, the protein quality is calculated. The absolute value of the first-order derivative of the data point in the spectral data is used as the edge strength index of the data point belonging to the edge of the protein stable characteristic peak, and then the edge strength index corresponding to each data point is obtained; then, for each data point, the smoothing weight factor of the data point and the normalized value of the edge strength index are weighted multiplied (for example, edge retention coefficient = smoothing weight factor × (1-normalized value of edge strength index)), where (1-normalized value of edge strength index) means that when the edge strength index is larger, the smoothing weight factor is reduced, and the result of the weighted product is used as the edge retention coefficient corresponding to the data point, and each data point is weighted. The corresponding edge retention coefficients are arranged in chronological order to obtain an edge retention coefficient matrix; finally, in the process of adaptive Gaussian filtering of the protein mass spectrometry data by Gaussian filtering, for each data point, the intensity of the Gaussian filter kernel is adjusted based on the edge retention coefficient of the data point (for example, the maximum standard deviation of the Gaussian filter (such as σ_max=3) represents the maximum smoothing intensity, and the adjustment is performed by: Gaussian filter kernel=maximum smoothing intensity×(1-edge retention coefficient)+maximum smoothing intensity×edge retention coefficient), and the data obtained by Gaussian filtering is used as the edge retention data of the stable characteristic peak of the protein.

[0036] It should be noted that, in this embodiment, the adaptive window represents the optimal sliding window for each data point in the protein mass spectrometry data; in this embodiment, the smoothing weight factor represents the degree to which each data point in the protein mass spectrometry data is affected by the smoothing operation, and the smoothing weight factor essentially reflects the intensity of smoothing the local area of ​​the data point; in this embodiment, the edge strength index represents a quantitative index for measuring the intensity of the signal change at each data point, and the edge strength index reflects the degree of mutation of the data at this position. The larger the value, the steeper the signal change, and the more likely it is to be at the edge position of the stable characteristic peak of the protein. The smaller the value, the gentler the signal change at this position, which may belong to the background or the internal area of ​​the characteristic peak; in this embodiment, the edge retention coefficient matrix represents a matrix composed of multiple edge retention coefficients, wherein the edge retention coefficient represents a weighted adjustment factor constructed for retaining the edge signal of the characteristic peak and smoothing and denoising the non-edge area.

[0037] It should also be noted that the loss constraint in the present application represents a process of retaining the edge of a stable characteristic peak of a protein in protein mass spectrometry data, wherein the edge of the stable characteristic peak of a protein in the protein mass spectrometry data is subjected to loss constraint based on the scale constraint parameter in combination with a preset sliding window, that is: the size of the preset sliding window at each data point in the protein mass spectrometry data is dynamically adjusted according to the scale constraint parameter, thereby obtaining an adaptive window corresponding to each data point; the smoothing weight factor corresponding to each data point is determined according to the adaptive window of each data point; the absolute value of the first-order derivative of the data point in the protein mass spectrometry data is calculated as an edge strength index of the data point belonging to the edge of the stable characteristic peak of the protein; the edge retention coefficient matrix when the loss constraint is performed on the edge of the stable characteristic peak of the protein in the protein mass spectrometry data is generated by the smoothing weight factor and the edge strength index of each data point; the protein mass spectrometry data is adaptively Gaussian filtered by the edge retention coefficient matrix to obtain edge retention data of the stable characteristic peak of the protein, that is, the edge retention data is used as the result of the loss constraint, thereby completing the loss constraint on the edge of the stable characteristic peak of the protein in the protein mass spectrometry data.

[0038] In step 104, the edge-retained data is pre-identified for protein stability characteristic peaks to obtain multiple pre-identified peaks, and the fuzzy membership of each data point to the protein stability characteristic peak is determined based on all the pre-identified peaks combined with the spatial distribution characteristics of each data point in the edge-retained data.

[0039] It should be noted that the pre-identification peak in this application refers to a signal area with potential protein stability characteristic peak attributes that is preliminarily screened out through local maximum search. The pre-identification peak usually has characteristics such as significant peak height, symmetrical peak shape, and clear edges, and is easy to identify. By determining the pre-identification peak, more obvious protein stability characteristic peak attributes can be obtained in advance, and then the missed or hidden protein stability characteristic peaks can be further identified from the protein mass spectrometry data based on the pre-identification peak.

[0040] In some embodiments, reference Figure 4 As shown in FIG. 1 , this figure is an exemplary flow chart for determining multiple pre-identified peaks according to some embodiments of the present application. In this embodiment, the edge retention data is pre-identified as protein stable characteristic peaks, and the multiple pre-identified peaks can be obtained by the following steps: First, in step 1041, local maximum points in the edge-preserving data are extracted as candidate peak vertices; Next, in step 1042, the first-order derivative zero-crossing points on the left and right sides of each candidate peak vertex are calculated, and then the candidate peak corresponding to each candidate peak vertex is constructed based on the first-order derivative zero-crossing points on both sides; Then, in step 1043, for each candidate peak, the signal-to-noise ratio of the signal integral value at the peak vertex within the candidate peak to the noise intensity of the adjacent baseline region is determined, thereby obtaining the signal-to-noise ratio corresponding to each candidate peak; Finally, in step 1044, multiple pre-identified peaks are extracted based on all signal-to-noise ratios.

[0041] In the specific implementation, first, all local maximum points are scanned in the edge-retained data (defined as the point whose amplitude is greater than all points within the radius of its left and right neighborhood, for example, the radius is 5 data points), and the local maximum points are used as candidate peak vertices; secondly, for each candidate peak vertex, the mass spectrum curve in the edge-retained data is searched to the left and right sides for points where the first-order derivative is zero (calculated by the central difference method, such as the zero-crossing point of the derivative is determined as the peak boundary, for example, the position where the left derivative changes from positive to negative is the peak starting point, and the right derivative changes from negative to positive is the ending point), and the area where the first-order derivatives on both sides cross zero is combined as the candidate peak corresponding to the candidate peak vertex, thereby obtaining the candidate peak corresponding to each candidate peak vertex; then, for each candidate peak vertex, Peak, determine the signal-to-noise ratio of the signal integral value of the peak apex in the candidate peak and the noise intensity of the adjacent baseline area, that is: for each candidate peak, take the ratio of the signal integral value of the peak apex in the candidate peak to the noise intensity of the adjacent baseline area as the signal-to-noise ratio, and then obtain the signal-to-noise ratio corresponding to each candidate peak, wherein the adjacent baseline area is the area of ​​6 data points on the left and right outside the candidate peak boundary, and the noise intensity is the signal standard deviation of the data points in the adjacent baseline area; finally, compare the signal-to-noise ratio corresponding to each candidate peak with the signal-to-noise ratio threshold, and extract candidate peaks with a signal-to-noise ratio greater than the signal-to-noise ratio threshold as pre-identified peaks. The signal-to-noise ratio threshold can be set according to actual needs or based on expert knowledge and is not limited here.

[0042] It should be noted that in this embodiment, the candidate peak vertex represents the top of the preliminarily selected protein signal peak; in this embodiment, the candidate peak represents the preliminarily selected protein signal peak region; in this embodiment, the signal-to-noise ratio represents the ratio between the signal intensity and the noise intensity, which is used to measure the significance of the useful signal in the background noise and is a key indicator in signal detection and analysis.

[0043] In some embodiments, the fuzzy membership of each data point to a protein stability characteristic peak is determined based on all pre-identified peaks combined with the spatial distribution characteristics of each data point in the edge-preserving data, specifically by the following steps: Extracting the identification peak center of each pre-identified peak; determining a spatial distribution characteristic of each data point in the edge-preserving data; For each data point, the spatial correlation parameter between the data point and the pre-identified peak is determined based on the spatial distribution characteristics of the data point and the identification peak center of the nearest pre-identified peak, thereby obtaining the spatial correlation parameter between each data point and the corresponding pre-identified peak; The fuzzy membership of each data point to the protein stable characteristic peak is determined based on all spatial correlation parameters.

[0044] It should be noted that the fuzzy membership in this application represents an indicator for measuring the likelihood that a data point in edge-retaining data spatially belongs to a stable characteristic peak of a protein. The value range is usually 0 to 1. The coefficient is calculated by the Gaussian membership function based on the spatial distance between the data point and the center of the characteristic peak. The fuzzy membership reflects a fuzzy attribution relationship, so that each data point can belong to multiple pre-identified peaks at the same time to varying degrees, thereby enhancing the stability and continuity of peak identification, and effectively processing mass spectrometry data with a low signal-to-noise ratio in a complex background.

[0045] In a specific implementation, first, for each pre-identified peak, the peak vertex of the pre-identified peak is extracted as the identification peak center; secondly, the spatial distribution characteristics of each data point in the edge-retained data are determined, that is, the spatial position of the data point in the edge-retained data can be used as the spatial distribution characteristics of the data point, thereby obtaining the spatial distribution characteristics of each data point; then, for each data point, the spatial correlation parameter between the data point and the pre-identified peak is determined based on the spatial distribution characteristics of the data point and the identification peak center of the nearest pre-identified peak, thereby obtaining the spatial correlation parameters of each data point and the corresponding pre-identified peak, that is, for each data point, the relative distance between the spatial distribution characteristics of the data point and the position of the identification peak center of the nearest pre-identified peak is calculated. The distance is used as the spatial correlation parameter between the data point and the pre-identified peak, and then the spatial correlation parameter between each data point and the corresponding pre-identified peak is obtained; finally, the fuzzy membership of each data point to the protein stable characteristic peak is determined based on all the spatial correlation parameters, that is: the spatial correlation parameter corresponding to each data point is obtained, each spatial correlation parameter is normalized to between 0 and 1 by minimum-maximum normalization, and each normalized spatial correlation parameter is used as the input variable of the Gaussian membership function, and the output result of the Gaussian membership function is used as the fuzzy membership of the corresponding data point, wherein the adjustment parameter in the Gaussian membership function can be set according to actual needs, for example, the half-peak width mean of all pre-identified peaks can be used as the adjustment parameter.

[0046] It should be noted that, in this embodiment, the center of the identified peak represents the peak top position of the pre-identified peak; in this embodiment, the spatial distribution feature represents the spatial distribution position of the data point in the edge-preserving data; in this embodiment, the spatial correlation parameter represents a numerical indicator for measuring the degree of spatial correlation between the data point and the pre-identified peak in the protein mass spectrometry data. The spatial correlation parameter reflects the proximity of the data point to the center of the pre-identified peak and is the core intermediate variable for calculating the fuzzy membership of the point to the peak.

[0047] In step 105, the edge-preserving data is baseline-corrected based on the Bayesian probability combined with the fuzzy membership of each data point to the stable characteristic peak of the protein, and then the peptide signal peak of the protein is extracted from the baseline-corrected edge-preserving data.

[0048] In some embodiments, baseline correction of the edge-preserving data based on Bayesian probability combined with the fuzzy membership of each data point to a protein stability characteristic peak can be performed by the following steps: The fuzzy membership of each data point to the protein stable characteristic peak is converted into the prior probability of the non-protein stable characteristic peak region; Based on the prior probability corresponding to each data point, the posterior probability of each data point belonging to the non-protein stable characteristic peak region is calculated by the Bayesian probability; Performing weighted regression fitting on the mass spectrum curve corresponding to the edge retention data using the normalized posterior probability corresponding to each data point as a fitting weight to obtain a mass spectrum baseline curve; The baseline component is removed from the edge-retained data based on the mass spectrum baseline curve, thereby completing the baseline correction process for the edge-retained data.

[0049] In the specific implementation, first, the fuzzy membership of each data point belonging to the protein stable characteristic peak is converted into a priori probability of the non-protein stable characteristic peak region, that is: since the protein stable characteristic peak region and the non-protein stable characteristic peak region are complementary to each other, (1-fuzzy membership) can be used to represent the possibility that the data point belongs to the non-protein stable characteristic peak region, and (1-fuzzy membership) is used as the priori probability of the data point belonging to the non-protein stable characteristic peak region, and then the fuzzy membership of each data point belonging to the protein stable characteristic peak is converted into a priori probability of the non-protein stable characteristic peak region; secondly, based on the prior probability corresponding to each data point, the posterior probability of each data point belonging to the non-protein stable characteristic peak region is calculated by the Bayesian probability, that is: for each data point, the total probability of the observation value corresponding to the data point in the marginal retention data is extracted, and the total probability and the prior probability corresponding to the data point are input as input parameters into the In the Bayesian probability function, the result output by the Bayesian probability function is used as the posterior probability that the data point belongs to the non-protein stable characteristic peak region, thereby obtaining the posterior probability that each data point belongs to the non-protein stable characteristic peak region; then, the mass spectrum curve corresponding to the edge-retained data is weightedly regressed and fitted using the normalized posterior probability corresponding to each data point as the fitting weight, and the fitted curve is used as the mass spectrum baseline curve, wherein the weighted regression fitting adopts the existing weighted polynomial fitting, which is not described here; finally, the baseline component is removed from the edge-retained data based on the mass spectrum baseline curve to complete the baseline correction processing of the edge-retained data, that is: each baseline value in the mass spectrum baseline curve is removed from the edge-retained data, thereby completing the baseline correction processing of the edge-retained data; in addition, in this embodiment, the prior probability needs to be iteratively optimized through the existing expectation maximization (EM) algorithm, which is not described here.

[0050] It should be noted that the prior probability in this embodiment represents the preliminary judgment probability that the data point in the edge-retained data belongs to the non-protein stable characteristic peak region; the posterior probability in this embodiment represents the updated probability judgment that the data point in the edge-retained data belongs to the non-protein stable characteristic peak region; the mass spectrometry baseline curve in this embodiment represents the trend line that does not contain the protein stable characteristic peak signal component in the protein mass spectrometry data, that is, it represents the changing trend of instrument background noise, chemical noise or other non-target signal components. In actual protein mass spectrometry data, ideal protein mass spectrometry data should appear as clear characteristic peaks (representing the ion abundance of peptides or proteins), but due to the influence of factors such as experimental conditions, ionization efficiency, and detection sensitivity, the entire protein mass spectrometry data is often superimposed with a fluctuating background signal. This background signal forms a "baseline" in the figure, which is not completely zero and may fluctuate with the scanning range or time. Therefore, the baseline needs to be accurately calibrated.

[0051] In specific implementation, the peptide signal peaks of proteins can be extracted from the edge-retained data after baseline correction based on the open source software tool (XCMS) for processing liquid chromatography-mass spectrometry data in the existing peak identification tools. I will not go into details here. XCMS is mainly used for the identification, alignment and quantitative analysis of peptide signal peaks in metabolomics and proteomics. XCMS is developed based on R language and is widely used for automated analysis and processing of mass spectrometry data. In addition, other peak identification tools can also be used to extract protein peptide signal peaks from the edge-retained data after baseline correction, which is not limited here.

[0052] It should be noted that the peptide signal peaks in this application represent the mass-intensity characteristic peaks finally identified in the protein mass spectrometry data. These signal peaks represent the presence of specific peptides, and their mass information can be used to infer the amino acid sequence and then identify the original protein.

[0053] In addition, another aspect of the present application, in some embodiments, the present application provides a protein mass spectrometry data analysis system based on protein stability, referring to Figure 5 , which is a schematic structural diagram of a protein mass spectrometry data analysis system based on protein stability according to some embodiments of the present application. The protein mass spectrometry data analysis system 200 based on protein stability includes: an acquisition module 201, a processing module 202, and an execution module 203, which are described as follows: Acquisition module 201, in this application, acquisition module 201 is mainly used to acquire protein mass spectrometry data to be analyzed; Processing module 202, in the present application, is mainly used to calculate the local noise variance at each data point in the protein mass spectrometry data, and determine the scale constraint parameter of the sliding window when smoothing the protein mass spectrometry data based on the local noise variance at each data point; The processing module 202 is further configured to perform loss constraints on the edges of protein stable characteristic peaks in the protein mass spectrometry data based on the scale constraint parameters in combination with a preset sliding window, thereby obtaining edge-preserved data of the protein stable characteristic peaks; In addition, the processing module 202 is further configured to pre-identify protein stability characteristic peaks on the edge-retained data to obtain a plurality of pre-identified peaks, and determine the fuzzy membership of each data point to the protein stability characteristic peak based on all the pre-identified peaks combined with the spatial distribution characteristics of each data point in the edge-retained data; Execution module 203, in this application, execution module 203 is mainly used to perform baseline correction on the edge-retained data based on Bayesian probability combined with the fuzzy membership of each data point to the stable characteristic peak of the protein, and then extract the peptide signal peak of the protein from the edge-retained data after baseline correction.

[0054] In addition, the present application also provides a computer device, which includes a memory and a processor, wherein the memory stores code, and the processor is configured to obtain the code and execute the above-mentioned protein mass spectrometry data analysis method based on protein stability.

[0055] In some embodiments, reference Figure 6 , which is a schematic diagram of the structure of a computer device for implementing a protein mass spectrometry data analysis method based on protein stability according to some embodiments of the present application. The protein mass spectrometry data analysis method based on protein stability in the above embodiment can be performed by Figure 6 The computer device 300 shown in FIG. 1 is implemented as shown in FIG. 1 , and the computer device 300 includes at least one processor 301 , a communication bus 302 , a memory 303 , and at least one communication interface 304 .

[0056] The processor 301 can be a general-purpose central processing unit (CPU), or an application-specific integrated circuit (ASIC) or one or more processors for controlling the execution of the protein mass spectrometry data analysis method based on protein stability in the present application.

[0057] The communication bus 302 may be used to transmit information between the aforementioned components.

[0058] Memory 303 may be, but is not limited to, a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, a random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, an optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer. Memory 303 may be independent and connected to processor 301 via communication bus 302. Memory 303 may also be integrated with processor 301.

[0059] Memory 303 is used to store program code for executing the solution of the present application, and is controlled by processor 301 for execution. Processor 301 is used to execute the program code stored in memory 303. The program code may include one or more software modules. The determination of the protein mass spectrometry data analysis method based on protein stability in the above embodiment can be implemented by processor 301 and one or more software modules in the program code in memory 303.

[0060] The communication interface 304 uses any device such as a transceiver to communicate with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area network (WLAN), etc.

[0061] In a specific implementation, as an example, a computer device may include multiple processors, each of which may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. A processor herein may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0062] The aforementioned computer device can be a general-purpose computer device or a dedicated computer device. In a specific implementation, the computer device can be a desktop computer, a portable computer, a network server, a personal digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device, or an embedded device. The embodiments of this application do not limit the type of computer device.

[0063] In addition, the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned protein mass spectrometry data analysis method based on protein stability.

[0064] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0065] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A protein mass spectrometry data analysis method based on protein stability, characterized in that: The steps include: Obtaining mass spectrometry data of the protein to be analyzed; calculating the local noise variance at each data point in the protein mass spectrometry data, and determining a scale constraint parameter of a sliding window when smoothing the protein mass spectrometry data according to the local noise variance at each data point; Based on the scale constraint parameter and a preset sliding window, loss constraint is performed on the edge of the protein stable characteristic peak in the protein mass spectrometry data, thereby obtaining edge retention data of the protein stable characteristic peak; Pre-identifying protein stability characteristic peaks on the edge-retained data to obtain a plurality of pre-identified peaks, and determining the fuzzy membership of each data point to the protein stability characteristic peak based on all the pre-identified peaks combined with the spatial distribution characteristics of each data point in the edge-retained data; The edge-retained data are baseline-corrected based on Bayesian probability combined with the fuzzy membership of each data point to a stable characteristic peak of the protein, and then the peptide signal peak of the protein is extracted from the edge-retained data after baseline correction.

2. The method according to claim 1, wherein Calculating the local noise variance at each data point in the protein mass spectrometry data specifically includes: Get the preset neighborhood radius; Performing local area division on each data point in the protein mass spectrometry data based on the neighborhood radius, thereby obtaining a neighborhood partition corresponding to each data point; For each data point, the local noise variance at the data point is determined according to the neighborhood partition corresponding to the data point, thereby obtaining the local noise variance at each data point in the protein mass spectrometry data.

3. The method according to claim 1, wherein Determining the scale constraint parameters of the sliding window when smoothing the protein mass spectrometry data based on the local noise variance at each data point specifically includes: Normalizing and mapping the local noise variance of each data point to generate a window scale coefficient for each data point; Setting the upper limit and lower limit of the scale constraint of the sliding window when smoothing the protein mass spectrometry data based on all window scale coefficients; The scale constraint parameter of the sliding window when smoothing the protein mass spectrometry data is determined according to the scale constraint upper limit and the scale constraint lower limit.

4. The method according to claim 1, wherein Performing loss constraints on the edges of protein stable characteristic peaks in the protein mass spectrometry data based on the scale constraint parameters in combination with a preset sliding window, thereby obtaining edge retention data of the protein stable characteristic peaks specifically includes: Dynamically adjusting the size of a preset sliding window at each data point in the protein mass spectrometry data according to the scale constraint parameter, thereby obtaining an adaptive window corresponding to each data point; Determine a smoothing weight factor corresponding to each data point according to the adaptive window of each data point; Calculating the absolute value of the first derivative of a data point in the protein mass spectrometry data as an edge intensity indicator of whether the data point belongs to the edge of a stable characteristic peak of the protein; generating an edge retention coefficient matrix for performing loss constraints on the edges of protein stable characteristic peaks in the protein mass spectrometry data based on the smoothing weight factor and edge strength index of each data point; Adaptive Gaussian filtering is performed on the protein mass spectrum data using the edge retention coefficient matrix to obtain edge retention data of stable characteristic peaks of the protein.

5. The method according to claim 1, wherein Pre-identifying protein stable characteristic peaks on the edge retention data to obtain multiple pre-identified peaks specifically includes: Extracting local maximum points in the edge-preserving data as candidate peak vertices; Calculate the zero-crossing points of the first-order derivatives on the left and right sides of each candidate peak vertex, and then construct the candidate peak corresponding to each candidate peak vertex according to the zero-crossing points of the first-order derivatives on both sides; For each candidate peak, the signal-to-noise ratio of the signal integral value at the peak apex within the candidate peak to the noise intensity of the adjacent baseline region is determined, thereby obtaining the signal-to-noise ratio corresponding to each candidate peak; Multiple pre-identified peaks were extracted based on all signal-to-noise ratios.

6. The method according to claim 1, wherein Determining the fuzzy membership of each data point to a protein stability characteristic peak based on all pre-identified peaks combined with the spatial distribution characteristics of each data point in the edge-retained data specifically includes: Extracting the identification peak center of each pre-identified peak; determining a spatial distribution characteristic of each data point in the edge-preserving data; For each data point, the spatial correlation parameter between the data point and the pre-identified peak is determined based on the spatial distribution characteristics of the data point and the identification peak center of the nearest pre-identified peak, thereby obtaining the spatial correlation parameter between each data point and the corresponding pre-identified peak; The fuzzy membership of each data point to the protein stable characteristic peak is determined based on all spatial correlation parameters.

7. The method according to claim 1, wherein The mass spectrometer is used to obtain mass spectrometry data of the protein to be analyzed.

8. A protein mass spectrometry data analysis system based on protein stability, characterized in that: The system includes: An acquisition module, used to acquire mass spectrometry data of proteins to be analyzed; a processing module, configured to calculate the local noise variance at each data point in the protein mass spectrometry data, and determine a scale constraint parameter of a sliding window when smoothing the protein mass spectrometry data according to the local noise variance at each data point; The processing module is further configured to perform loss constraints on the edges of protein stable characteristic peaks in the protein mass spectrometry data based on the scale constraint parameters in combination with a preset sliding window, thereby obtaining edge-retained data of the protein stable characteristic peaks; The processing module is further configured to pre-identify protein stability characteristic peaks on the edge-retained data to obtain a plurality of pre-identified peaks, and determine the fuzzy membership of each data point to the protein stability characteristic peak based on all the pre-identified peaks combined with the spatial distribution characteristics of each data point in the edge-retained data; The execution module is used to perform baseline correction on the edge-retained data based on Bayesian probability combined with the fuzzy membership of each data point to the stable characteristic peak of the protein, and then extract the peptide signal peak of the protein from the edge-retained data after baseline correction.

9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores codes, and the processor is configured to acquire the codes and execute the protein mass spectrometry data analysis method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the protein mass spectrometry data analysis method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Flow cytometer signal background dynamic acquisition method based on window sliding screening

    CN121049138A