Mass spectrometry data processing method, system, device and storage medium thereof

By acquiring the Gaussian distribution of the signal intensity of the mass spectrometry data for data sampling and adjustment, the problem of insufficient number and diversity of mass spectrometry data is solved, and the data amplification and model training effect is improved.

CN118861619BActive Publication Date: 2025-08-15AFFILIATED HUSN HOSPITAL OF FUDAN UNIV +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410836217.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-25
Publication Date
2025-08-15
Estimated Expiration
2044-06-25

AI Technical Summary

Technical Problem

The prior art is difficult to obtain sufficient quantity and variety of mass spectrometry data in sample detection, resulting in poor model training results.

Method used

By acquiring the Gaussian distribution of the signal intensity of the mass spectrometry data, data sampling and adjustment are carried out to generate multiple different and reliable new mass spectrometry data, achieving data amplification and diversity improvement.

Benefits of technology

The number of mass spectrometry data is increased, while the diversity and reliability of the data is improved, and the training effect and analysis accuracy of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118861619B_ABST
    Figure CN118861619B_ABST
Patent Text Reader

Abstract

A method for processing mass spectrometry data, and its system, device, and storage medium, includes: obtaining raw mass spectrometry data, the raw mass spectrometry data including signal intensity that varies with mass-to-charge ratio; obtaining a Gaussian distribution of the signal intensity; sampling the Gaussian distribution of the signal intensity, and adjusting the distribution of the signal intensity of the raw mass spectrometry data to varying degrees based on the sampling results, thereby obtaining a plurality of different newly added mass spectrometry data corresponding to the raw mass spectrometry data. Because the Gaussian distribution of signal intensity can characterize the normal fluctuation pattern of signal intensity, sampling the signal intensity based on the Gaussian distribution of signal intensity makes the newly added mass spectrometry data more reliable and achieves data amplification of the raw mass spectrometry data, thereby increasing the amount of data while improving the diversity of the data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of data processing technology, and in particular to a method for processing mass spectrometry data and a system, device, and storage medium thereof. Background Art

[0002] Sample testing technology has been widely used in fields such as medicine, biology, environment and food science. By testing the samples to be tested, test data with analytical significance is obtained.

[0003] However, as sample analysis methods become more complex, the requirements for the amount of sample test data also increase. For example, with the development of artificial intelligence, technologies such as machine learning and deep learning have gradually been used to implement data analysis, so more test data is needed to meet the requirements of model training. Summary of the Invention

[0004] The problem solved by the embodiments of the present invention is to provide a method for processing mass spectrometry data and its system, device and storage medium, which increase the amount of data while improving the diversity of data.

[0005] To solve the above problems, an embodiment of the present invention provides a method for processing mass spectrometry data, including: obtaining original mass spectrometry data, wherein the original mass spectrometry data includes signal intensity that varies with mass-to-charge ratio; obtaining a Gaussian distribution of the signal intensity; performing data sampling on the Gaussian distribution of the signal intensity, and adjusting the distribution of the signal intensity of the original mass spectrometry data to different degrees based on the sampling results, to obtain multiple different new mass spectrometry data corresponding to the original mass spectrometry data.

[0006] Correspondingly, an embodiment of the present invention also provides a mass spectrometry data processing system, including: an original data acquisition module for acquiring original mass spectrometry data, wherein the original mass spectrometry data includes signal intensity that changes with the mass-to-charge ratio; a Gaussian distribution acquisition module for acquiring the Gaussian distribution of the signal intensity; a data amplification module for performing data sampling on the Gaussian distribution of the signal intensity, and adjusting the distribution of the signal intensity of the original mass spectrometry data to different degrees based on the sampling results, to obtain multiple different newly added mass spectrometry data corresponding to the original mass spectrometry data.

[0007] Accordingly, an embodiment of the present invention also provides a device comprising at least one memory and at least one processor, wherein the memory stores one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method for processing mass spectrometry data described in any one of the embodiments of the present invention.

[0008] Correspondingly, an embodiment of the present invention further provides a storage medium, wherein the storage medium stores one or more computer instructions, and the one or more computer instructions are used to implement the method for processing mass spectrometry data described in any one of the embodiments of the present invention.

[0009] Compared with the prior art, the technical solution of the embodiment of the present invention has the following advantages:

[0010] In the mass spectrometry data processing method provided by an embodiment of the present invention, the Gaussian distribution of signal intensity is utilized, and through sampling, the distribution of the signal intensity of the original mass spectrometry data is adjusted to different degrees based on the sampling results to obtain multiple different new mass spectrometry data corresponding to the original mass spectrometry data. Among them, since the Gaussian distribution of signal intensity can characterize the normal fluctuation law of signal intensity, the signal intensity is sampled based on the Gaussian distribution of signal intensity, so that the reliability of the new mass spectrometry data is higher, and data amplification of the original mass spectrometry data is achieved, thereby increasing the amount of data while improving the diversity of the data.

[0011] In the mass spectrometry data processing system provided by an embodiment of the present invention, based on the Gaussian distribution of signal intensity provided by the Gaussian distribution acquisition module, the data amplification module utilizes the Gaussian distribution of signal intensity, and through sampling, adjusts the distribution of the signal intensity of the original mass spectrometry data to different degrees based on the sampling results, thereby obtaining multiple different new mass spectrometry data corresponding to the original mass spectrometry data. Among them, since the Gaussian distribution of signal intensity can characterize the normal fluctuation law of signal intensity, the signal intensity is sampled based on the Gaussian distribution of signal intensity, so that the reliability of the new mass spectrometry data is higher, and data amplification of the original mass spectrometry data is realized, thereby increasing the amount of data while improving the diversity of data. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 1 is a flow chart of an embodiment of a method for processing mass spectrometry data according to the present invention;

[0013] Figure 2 yes Figure 1 Schematic diagram of raw mass spectrum data of an embodiment corresponding to step S10;

[0014] Figure 3 is a schematic diagram of an embodiment of a Gaussian distribution of signal intensity in a method for processing mass spectrometry data of the present invention;

[0015] Figure 4 yes Figure 1 Schematic diagram of newly added mass spectrum data in one embodiment corresponding to step S20;

[0016] Figure 5A t-SNE distribution diagram of an embodiment obtained based on original mass spectrum data, and a t-SNE distribution diagram of an embodiment obtained based on mass spectrum data obtained by the mass spectrum data processing method according to an embodiment of the present invention;

[0017] Figure 6 This is a comparison diagram of the confusion matrix of the machine learning model accuracy;

[0018] Figure 7 This is a comparison diagram of the confusion matrix of the deep learning model accuracy;

[0019] Figure 8 1 is a schematic structural diagram of an embodiment of a mass spectrometry data processing system of the present invention;

[0020] Figure 9 This is a hardware structure diagram of a device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0021] As can be seen from the background technology, in order to improve the training effect of the model, the requirement for the amount of data is getting higher and higher.

[0022] However, due to the limitation of the number of actual tests on the samples to be tested, it is difficult to obtain diverse data (for example, in hospitals, the amount of mass spectrometry data obtained from testing a single strain is usually in the single digit), resulting in the actual amount of collected data being difficult to meet the needs of model training.

[0023] In order to solve the above technical problems, an embodiment of the present invention provides a method for processing mass spectrometry data. Figure 1 , which shows a flow chart of an embodiment of a method for processing mass spectrometry data of the present invention.

[0024] In this embodiment, the method for processing mass spectrometry data includes the following basic steps:

[0025] Step S10: acquiring raw mass spectrum data, wherein the raw mass spectrum data includes signal intensity that varies with mass-to-charge ratio;

[0026] Step S20: Obtaining Gaussian distribution of the signal strength;

[0027] Step S30: performing data sampling on the Gaussian distribution of the signal intensity, and adjusting the distribution of the signal intensity of the original mass spectrum data to varying degrees based on the sampling results, to obtain a plurality of different newly added mass spectrum data corresponding to the original mass spectrum data.

[0028] The embodiment of the present invention utilizes the Gaussian distribution of signal intensity, and through sampling, adjusts the distribution of the signal intensity of the original mass spectrum data to different degrees based on the sampling results to obtain multiple different new mass spectrum data corresponding to the original mass spectrum data. Among them, since the Gaussian distribution of signal intensity can characterize the normal fluctuation law of signal intensity, the signal intensity is sampled based on the Gaussian distribution of signal intensity, so that the reliability of the new mass spectrum data is higher, and data amplification of the original mass spectrum data is achieved, thereby increasing the amount of data while improving the diversity of the data.

[0029] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0030] Combined with reference Figure 1 and Figure 2 , Figure 2 3 is a schematic diagram of an embodiment of raw mass spectrum data. Step S10 is executed to obtain raw mass spectrum data, where the raw mass spectrum data includes signal intensity that varies with mass-to-charge ratio.

[0031] Subsequently, the signal intensity distribution of the original mass spectrometry data is adjusted to varying degrees to obtain adjusted signal intensities, and new mass spectrometry data different from the original mass spectrometry data is obtained accordingly, thereby obtaining a larger number of new mass spectrometry data based on the original mass spectrometry data to achieve data amplification.

[0032] It should be noted that, in the mass spectrum data, the abscissa is the mass-to-charge ratio (m / z), and the ordinate is the signal intensity.

[0033] Mass spectrometry data is: the sample is dissociated, and different flight times are generated according to the mechanism of different molecular sizes, forming different fragmentation mass spectrometry characteristics.

[0034] Therefore, the mass spectrometry data is amplified by adjusting the distribution of signal intensity to different degrees. Since the test of mass spectrometry data usually has a certain degree of volatility, the mass spectrometry data processing method of this embodiment can generate fitting mass spectrometry data caused by signal intensity fluctuations.

[0035] It should also be noted that the raw mass spectrum data is obtained by testing a preset number of tested samples.

[0036] like Figure 2 As shown, Figure 2 The horizontal axis represents the mass-to-charge ratio, and the vertical axis represents the signal intensity. In this embodiment, the signal intensity of the mass spectrometry data is the peak intensity. For example, the sample to be tested may be a pathogen, and the mass spectrometry data is obtained by detecting the pathogen.

[0037] Mass spectrometry is a technique used to analyze the chemical composition of molecules in a sample. By using a mass spectrometer to examine the sample, the mass-to-charge ratio (m / z) of the molecule and the intensity at each mass-to-charge ratio can be determined. Therefore, mass spectrometry data contains intensities that vary with the mass-to-charge ratio. Mass spectrometry data is often used in fields such as medicine, biology, environment, and food science.

[0038] Correspondingly, the mass-to-charge ratio of the mass spectrum data is the position of the mass-to-charge ratio corresponding to each molecule, and the signal intensity is the intensity at the mass-to-charge ratio.

[0039] To further improve the reliability of sample analysis, the fusion or splicing of mass spectrometry data with data from other modalities (e.g., Raman data) is increasingly being used. For example, when using machine learning or deep learning models for sample analysis, both the mass spectrometry data and Raman data of the sample to be tested can be input into the model, and the sample analysis can be performed by fusing or splicing the data, thereby improving the accuracy of the sample analysis.

[0040] Specifically, the raw mass spectrum data is obtained by testing the sample using a detection device, for example, a mass spectrometer.

[0041] It should be noted that the number of the original mass spectrum data may be one or more.

[0042] Continue to refer Figure 1 , execute step S20 to obtain the Gaussian distribution of the signal strength.

[0043] The Gaussian distribution of signal intensity can characterize the normal fluctuation pattern of signal intensity. Therefore, the Gaussian distribution of the signal intensity is first obtained so that the signal intensity can be sampled based on the Gaussian distribution to generate simulated mass spectrum data as new mass spectrum data.

[0044] Moreover, the signal intensity is sampled based on the Gaussian distribution of the signal intensity, which makes the newly added mass spectrometry data more reliable.

[0045] In this embodiment, in the step of obtaining the Gaussian distribution of the signal intensity, the Gaussian distribution of the signal intensity detectable by a detection device is obtained, and the detection device is used to obtain the original mass spectrum data.

[0046] When the detection equipment detects the sample to be tested, the detected signal intensity usually fluctuates to a certain extent (for example, the fluctuation of the detection result of the detection equipment may be caused by human operation, sample preparation or the detection equipment itself), and the random distribution of the signal intensity of the statistical mass spectrometry data satisfies the Gaussian distribution.

[0047] Moreover, fluctuations in the detection results of the detection equipment will usually have similar or identical effects on different mass-to-charge ratios of the same sample being tested. In other words, it can be considered that different mass-to-charge ratios are subject to the same Gaussian distribution, and even have similar or identical effects on different samples being tested. In other words, it can be considered that different samples being tested are subject to the same Gaussian distribution. Therefore, selecting a Gaussian distribution of signal intensity that can be detected by the detection equipment is beneficial to improving the universality of the Gaussian distribution. Accordingly, compared with the solution of setting an independent Gaussian distribution for each mass-to-charge ratio, this embodiment is beneficial to reducing the complexity of the mass spectrometry data processing method.

[0048] In addition, when using the newly added mass spectrometry data for model training, feature extraction is usually performed, and the signal intensity Gaussian distribution used to characterize the stability of the detection results of the detection equipment is used to indicate that the fluctuation cannot reflect the characteristics of the sample being tested itself, and the fluctuation is caused by the detection equipment. This is conducive to reducing the attention to the fluctuation caused by the detection equipment during model training. Therefore, the newly added mass spectrometry data obtained by the mass spectrometry data processing method described in this embodiment is used for model training. While increasing the training samples, it can reduce the probability of negative impact on the training effect, making the trained model more accurate.

[0049] Specifically, with reference to Figure 3 , Figure 3 It is a schematic diagram of the Gaussian distribution of signal intensity provided by one embodiment of the present invention. In this embodiment, the Gaussian distribution is a Gaussian distribution of the signal intensity offset percentage, and the signal intensity offset percentage refers to: the offset percentage of the signal intensity relative to the signal intensity of the original mass spectrometry data.

[0050] like Figure 3 As shown, Figure 3 The horizontal axis represents the signal strength deviation percentage, and the vertical axis represents the quantity. The higher the quantity of the vertical axis, the higher the probability corresponding to the deviation percentage.

[0051] The offset percentage is used to characterize the relative intensity, so that in subsequent sampling, different mass-to-charge ratios or different tested samples can share the same Gaussian distribution data (ie, the offset percentage of the signal intensity), thereby reducing the complexity of the data processing.

[0052] It should be noted that before subsequently adjusting the signal intensity distribution of the original mass spectrum data to different degrees, it also includes: executing step S25, performing a first data amplification on the original mass spectrum data by copying the data, and obtaining a plurality of copy data corresponding to the original mass spectrum data.

[0053] The copied data is obtained by copying the original mass spectrum data, that is, the copied data is the same as the corresponding original mass spectrum data. Therefore, the subsequent adjustment of the distribution of the signal intensity of the copied data to different degrees is equivalent to adjusting the distribution of the signal intensity of the original mass spectrum data to different degrees.

[0054] It should be noted that by duplicating data, after subsequent data sampling, it is convenient to use multiple sampling results to adjust each duplicate data separately at the same time, thereby improving the efficiency of data processing.

[0055] It is understandable that in other embodiments, the step of copying data may not be performed, and the distribution of the signal intensity of the original mass spectrum data may be adjusted according to different sampling results, and the adjusted data may be stored as newly added mass spectrum data.

[0056] It should also be noted that, when there are multiple original mass spectrum data, the number of replicated data corresponding to each original mass spectrum data can be determined according to data requirements.

[0057] In this embodiment, in the first data amplification performed on the original mass spectrum data by copying the data, the higher the quality of the original mass spectrum data, the greater the number of corresponding copy data.

[0058] The higher the quality of the original mass spectrometry data, the higher the accuracy of the original mass spectrometry data. Accordingly, the more features it embodies, the greater the effect on data analysis (for example, it is conducive to realizing data calculation or classification). Therefore, for original mass spectrometry data of higher quality, the more data amplification is performed, which is more conducive to improving the reliability of the data analysis results. Accordingly, when training the model, increasing the number of reliable training samples is conducive to improving the training effect of the model, making the trained model more accurate.

[0059] As an example, the method for obtaining the number of copy data corresponding to each of the original mass spectrum data includes: setting the weight corresponding to each of the original mass spectrum data based on the quality of the original mass spectrum data, and the higher the quality of the original mass spectrum data, the greater the corresponding weight; based on the target total number of mass spectrum data and the weight of each of the original mass spectrum data, obtaining the number of copy data corresponding to each of the original mass spectrum data.

[0060] The weights are set based on the quality, thereby increasing the proportion of high-quality mass spectrometry data in the overall data, that is, improving the validity of the data information.

[0061] In this embodiment, the quality of the original mass spectrum data is determined by comparing the original mass spectrum data with standard mass spectrum data. The higher the matching degree, the higher the quality of the original mass spectrum data.

[0062] The degree of matching with standard mass spectrometry data is used as the evaluation criterion, which reduces the complexity of quality judgment.

[0063] Specifically, the method of comparing the original mass spectrum data with the standard mass spectrum data includes: comparing the position distribution of the mass-to-charge ratio of the original mass spectrum data with the position distribution of the mass-to-charge ratio of the standard mass spectrum data. The higher the matching degree of the position distribution, the higher the quality of the original mass spectrum data.

[0064] The signal intensities obtained from testing different samples are often different, and the signal intensity has little significance for quality evaluation. In a single original mass spectrum, the position of each mass-to-charge ratio should theoretically be fixed. Therefore, the quality of the original mass spectrum data is comprehensively evaluated through the position distribution of the mass-to-charge ratio to obtain the overall quality of the single original mass spectrum.

[0065] It should be noted that, in other embodiments, other quality assessment methods may be selected according to other situations such as the specific type of the original mass spectrometry data, or actual needs.

[0066] refer to Figure 1 , execute step S30, perform data sampling on the Gaussian distribution of the signal intensity, and adjust the distribution of the signal intensity of the original mass spectrum data to different degrees based on the sampling results to obtain multiple different new mass spectrum data corresponding to the original mass spectrum data.

[0067] Since the Gaussian distribution of signal intensity can characterize the normal fluctuation pattern of signal intensity, sampling the signal intensity based on the Gaussian distribution of signal intensity makes the newly added mass spectrometry data more reliable and realizes data amplification of the original mass spectrometry data, thereby increasing the amount of data while improving the diversity of data; accordingly, model training based on richer and more diverse data is conducive to improving the accuracy of the model.

[0068] In addition, by performing data amplification on the original mass spectrometry data, in actual operation, when the original mass spectrometry data needs to be fused or spliced with data from other modalities, the data volume of the mass spectrometry data of the modality described in this embodiment can match that of the data from other modalities, thereby being compatible with the data from other modalities.

[0069] Since several replicated data corresponding to the original mass spectrum data are first obtained by copying the data, data sampling is performed on the Gaussian distribution of the signal intensity, and the distribution of the signal intensity of the original mass spectrum data is adjusted to different degrees based on the sampling results, including: adjusting the signal intensities of the several replicated data based on different sampling results to achieve second data amplification and obtain multiple different new mass spectrum data.

[0070] That is, for a single original mass spectrum data, a plurality of newly added mass spectrum data different from the original mass spectrum data are obtained through the second data amplification, and the newly added mass spectrum data are also different from each other.

[0071] In this embodiment, data sampling is performed on the Gaussian distribution of the signal intensity, and based on the sampling results, the distribution of the signal intensity of the original mass spectrum data is adjusted to different degrees. From the Gaussian distribution of the offset percentage, an offset percentage is selected for each mass-to-charge ratio, and the signal intensity corresponding to the mass-to-charge ratio in the original mass spectrum data is adjusted based on the selected offset percentage.

[0072] When there are multiple original mass spectrum data, since each original mass spectrum data shares the same Gaussian distribution data, and the Gaussian distribution is a Gaussian distribution of signal intensity offset percentage, for each mass-to-charge ratio in each original mass spectrum data, the original value of the signal intensity of the mass-to-charge ratio can be adjusted based on the sampled selected offset percentage.

[0073] For example, for each mass-to-charge ratio, the signal intensity can be adjusted using the formula I2=I1*(1+m%); where I2 represents the signal intensity of the mass-to-charge ratio in the newly added mass spectrum data, I1 represents the signal intensity of the mass-to-charge ratio in the original mass spectrum data, and m% represents the offset percentage of the signal intensity, which can be zero, a positive value, or a negative value. In other words, for each mass-to-charge ratio, after determining the signal intensity offset percentage, the original signal intensity of the mass-to-charge ratio is adjusted to obtain a new signal intensity.

[0074] It can be seen that although the signal intensities of different mass-to-charge ratios may be different, and the signal intensity distributions of different tested samples may also be different, by using a Gaussian distribution of relative intensity, the same calculation method can be used to adjust the original value of the signal intensity of each mass-to-charge ratio for different original mass spectrometry data and different mass-to-charge ratios, thereby reducing the complexity of data amplification.

[0075] In this embodiment, in the same newly added mass spectrum data, the adjustment degrees of the signal intensity of each mass-to-charge ratio relative to the signal intensity of the mass-to-charge ratio in the original mass spectrum data are different.

[0076] Specifically, since several replicated data corresponding to the original mass spectrum data are first obtained by replicating the data, the signal intensities of the several replicated data are adjusted respectively based on different sampling results, and the signal intensities of each mass-to-charge ratio in the same replicated data are adjusted to different degrees.

[0077] By varying the degree of change in the signal intensity of each mass-to-charge ratio relative to the original value within the same newly added mass spectrum data, more diverse data can be obtained. Furthermore, when the data obtained using the method described in this embodiment is analyzed, even if data normalization is performed, the diversity of the normalized data can still be maintained.

[0078] For example, for a certain mass-to-charge ratio, the offset of the signal intensity in the newly added mass spectrum data relative to the signal intensity in the original mass spectrum data is +3% (i.e., an increase of 3%), while for another mass-to-charge ratio, the offset of the signal intensity in the newly added mass spectrum data relative to the signal intensity in the original mass spectrum data is -2% (i.e., a decrease of 2%).

[0079] refer to Figure 4 , Figure 4 is a schematic diagram of newly added mass spectrum data according to an embodiment of the present invention, and in order to show the difference between the original mass spectrum data and the newly added mass spectrum data, Figure 4 In the Figure 2 The original mass spectrum data in are shown in the same coordinate.

[0080] Among them, Figure 4 In the figure, the darker data represents the original mass spectrum data, and the lighter data represents the newly added mass spectrum data. Therefore, by performing data amplification, newly added mass spectrum data different from the original mass spectrum data can be obtained.

[0081] Combined with reference Figure 5 , Figure 5 The t-SNE distribution diagram based on mass spectrometry data is shown, where Figure 5 (a) shows a t-SNE distribution diagram of an embodiment obtained based on the original mass spectrum data, Figure 5 (b) shows a t-SNE distribution diagram of an embodiment obtained based on the mass spectrometry data obtained by the method described in this embodiment.

[0082] t-SNE (t-Distributed Stochastic Neighbor Embedding) is an unsupervised nonlinear technique that is mainly used to reduce the dimensionality of data and enable the data to retain the information it carries in the high-dimensional space in the low-dimensional space.

[0083] like Figure 5 As shown, taking the analysis of pathogens as an example, each point represents a data, and the data points in the same circle, ellipse, square or trapezoidal circle belong to the same pathogen category. For example, the pathogens include: Mycobacterium abscessus (M.abscessus), Mycobacterium fortuitum (M.fortuitum), Mycobacterium ulcerans (M.ulcerans), Mycobacterium peregrinum (M.peregrinum), Mycobacterium phlei (M.phlei) and Mycobacterium chelonae (M.chelonae).

[0084] It should be noted that for the convenience of illustration, Figure 5 In the figure, the points in the circle represent Mycobacterium chelonae, the points in the square circle represent Mycobacterium fortuitum, the points in the trapezoidal circle represent Mycobacterium abscessus, the points in the solid oval circle represent Mycobacterium ulcerans, the points in the dotted oval circle represent Mycobacterium peregrinum, and the points in the double-dashed oval circle represent Mycobacterium phlei.

[0085] from Figure 5 As can be seen from (a), the t-SNE distribution diagram obtained by analyzing the samples only with the original mass spectrometry data shows that the number of data points under the same pathogen category is relatively small. Figure 5 As can be seen from (b), when the mass spectrometry data obtained by the method described in this embodiment are used for sample analysis, the number of data points under the same pathogen category increases and the distribution is stable. Moreover, the distribution of the data points does not show serious deformation, and the categories can still be distinguished well.

[0086] refer to Figure 6 , Figure 6 This is a comparative diagram of the accuracy confusion matrix obtained based on the machine learning model. The darker the color of the grid where the number is located, the higher the accuracy.

[0087] in, Figure 6 (a) represents the accuracy confusion matrix of an embodiment of a machine learning model trained based on raw mass spectrometry data, Figure 6 (b) represents an accuracy confusion matrix of an embodiment of a machine learning model obtained by training the mass spectrometry data obtained by the mass spectrometry data processing method described in this embodiment.

[0088] As an example, the model is used to analyze pathogens, and in the confusion matrix, True represents the true result, Predicted represents the predicted result, and the types of pathogens include: Mycobacterium abscessus (M.abscessus), Mycobacterium fortuitum (M.fortuitum), Mycobacterium ulcerans (M.ulcerans), Mycobacterium peregrinum (M.peregrinum), Mycobacterium phlei (M.phlei) and Mycobacterium chelonae (M.chelonae).

[0089] For example, taking True as Mycobacterium chelonae (M. chelonae) as an example, when the machine learning model trained based on the original mass spectrometry data is used for analysis, the probability of predicting an accurate result is 0.83 (or 83%), and there is a 0.17 (or 17%) probability of predicting it as Mycobacterium ulcerans (M. ulcerans).

[0090] like Figure 6 As shown in (a), after training based on the original mass spectrometry data, the overall prediction accuracy of the machine learning model for pathogen categories is 69.40%, as shown in Figure 6 As shown in (b), after training with the mass spectrometry data obtained based on the mass spectrometry data processing method described in this embodiment, the overall prediction accuracy of the machine learning model for pathogen categories is 78.43%, and a higher prediction accuracy can be achieved.

[0091] refer to Figure 7 , Figure 7 This is a comparative diagram of the accuracy confusion matrix obtained based on the deep learning model.

[0092] in, Figure 7 (a) represents the accuracy confusion matrix of an embodiment of a deep learning model trained based on raw mass spectrometry data, Figure 7 (b) represents the accuracy confusion matrix of an embodiment of a deep learning model obtained by training the mass spectral data obtained by the mass spectral data processing method according to an embodiment of the present invention.

[0093] As an example, the model is used to analyze pathogens, and in the confusion matrix, True represents the true result, and Predicted represents the predicted result.

[0094] like Figure 7 As shown in (a), after training based on the original mass spectrometry data, the overall prediction accuracy of the deep learning model for pathogen categories is 75.38%, as shown in Figure 7 As shown in (b), after training with the mass spectrometry data obtained based on the mass spectrometry data processing method described in this embodiment, the overall prediction accuracy of the deep learning model for pathogen categories is 90.33%, and a higher prediction accuracy can be achieved.

[0095] Correspondingly, the present invention also provides a mass spectrometry data processing system. Figure 8 It is a structural diagram of an embodiment of a system for processing mass spectrometry data of the present invention.

[0096] refer to Figure 8 , and combined with reference Figures 2 to 7 The mass spectrum data processing system includes: a raw data acquisition module 10, used to acquire raw mass spectrum data, wherein the raw mass spectrum data includes a signal intensity that varies with the mass-to-charge ratio; a Gaussian distribution acquisition module 20, used to acquire the Gaussian distribution of the signal intensity; and a data amplification module 30, used to perform data sampling on the Gaussian distribution of the signal intensity, and adjust the distribution of the signal intensity of the raw mass spectrum data to different degrees based on the sampling results, so as to obtain a plurality of different newly added mass spectrum data corresponding to the raw mass spectrum data.

[0097] The mass spectrometry data processing system adjusts the signal intensity distribution of the original mass spectrometry data to varying degrees to obtain adjusted signal intensities, and accordingly obtains new mass spectrometry data that is different from the original mass spectrometry data, thereby obtaining a larger amount of new mass spectrometry data based on the original mass spectrometry data to achieve data amplification.

[0098] Among them, since the Gaussian distribution of signal intensity can characterize the normal fluctuation law of signal intensity, the signal intensity is sampled based on the Gaussian distribution of signal intensity, so that the reliability of the newly added mass spectrometry data is higher. In addition, data amplification of the original mass spectrometry data is achieved, thereby increasing the amount of data while improving the diversity of data; accordingly, model training based on richer and more diverse data is conducive to improving the accuracy of the model.

[0099] In addition, by performing data amplification on the original mass spectrometry data, in actual operation, when the original mass spectrometry data needs to be fused or spliced with data from other modalities, the data volume of the mass spectrometry data of the modality described in this embodiment can match that of the data from other modalities, thereby being compatible with the data from other modalities.

[0100] It should be noted that, in the mass spectrum data, the abscissa is the mass-to-charge ratio, and the ordinate is the signal intensity.

[0101] It should also be noted that the raw mass spectrum data is obtained by testing a preset number of tested samples.

[0102] like Figure 2 As shown, Figure 2The horizontal axis represents the position of the mass-to-charge ratio, and the vertical axis represents the signal intensity. In this embodiment, the signal intensity of the mass spectrometry data is the peak intensity. For example, the sample to be tested can be a pathogen, and the mass spectrometry data is obtained by detecting the pathogen.

[0103] Mass spectrometry is a technique used to analyze the chemical composition of molecules in a sample. By testing the sample with a mass spectrometer, the mass-to-charge ratio (m / z) of the molecule and the intensity at each mass-to-charge ratio can be obtained. Therefore, mass spectrometry data contains intensity that varies with the mass-to-charge ratio. Mass spectrometry data is often used in fields such as medicine, biology, environment, and food science.

[0104] Correspondingly, the mass-to-charge ratio of the mass spectrum data is the position of the mass-to-charge ratio corresponding to each molecule, and the signal intensity is the intensity at the mass-to-charge ratio.

[0105] To further improve the reliability of sample analysis, the fusion or splicing of mass spectrometry data with data from other modalities (e.g., Raman data) is increasingly being used. For example, when using machine learning or deep learning models for sample analysis, both the mass spectrometry data and Raman data of the sample to be tested can be input into the model, and the sample analysis can be performed by fusing or splicing the data, thereby improving the accuracy of the sample analysis.

[0106] Specifically, the raw mass spectrum data is obtained by testing the sample using a detection device, for example, a mass spectrometer.

[0107] The Gaussian distribution of signal intensity can characterize the normal fluctuation pattern of signal intensity. Therefore, the Gaussian distribution of signal intensity is first obtained so that the data amplification module 30 samples the signal intensity based on the Gaussian distribution, thereby generating simulated mass spectrum data as new mass spectrum data.

[0108] Moreover, the signal intensity is sampled based on the Gaussian distribution of the signal intensity, which makes the newly added mass spectrometry data more reliable.

[0109] In this embodiment, the Gaussian distribution acquisition module 20 is used to obtain the Gaussian distribution of the signal intensity that can be detected by a detection device, and the detection device is used to obtain the original mass spectrum data.

[0110] When the detection equipment detects the sample to be tested, the detected signal intensity usually fluctuates to a certain extent (for example, the fluctuation of the detection results of the detection equipment may be caused by human operation, sample preparation or the detection equipment itself), and the random distribution of the signal intensity of the statistical mass spectrometry data often satisfies the Gaussian distribution.

[0111] Moreover, fluctuations in the detection results of the detection equipment will usually have similar or identical effects on different mass-to-charge ratios of the same sample being tested. In other words, it can be considered that different mass-to-charge ratios are subject to the same Gaussian distribution, and even have similar or identical effects on different samples being tested. In other words, it can be considered that different samples being tested are subject to the same Gaussian distribution. Therefore, selecting a Gaussian distribution of signal intensity that can be detected by the detection equipment is beneficial to improving the universality of the Gaussian distribution. Accordingly, compared with the solution of setting an independent Gaussian distribution for each mass-to-charge ratio, this embodiment is beneficial to reducing the complexity of the mass spectrometry data processing method.

[0112] In addition, when using the newly added mass spectrometry data for model training, feature extraction is usually performed, and the signal intensity Gaussian distribution used to characterize the stability of the detection results of the detection equipment is used to indicate that the fluctuation cannot reflect the characteristics of the sample being tested itself, and the fluctuation is caused by the detection equipment. This is conducive to reducing the attention to the fluctuation caused by the detection equipment during model training. Therefore, the newly added mass spectrometry data obtained by the mass spectrometry data processing method described in this embodiment is used for model training. While increasing the training samples, it can reduce the probability of negative impact on the training effect, making the trained model more accurate.

[0113] Combined with reference Figure 3 In this embodiment, the Gaussian distribution is a Gaussian distribution of signal intensity offset percentage, and the signal intensity offset percentage refers to: the offset percentage of the signal intensity relative to the signal intensity of the original mass spectrum data. Figure 3 As shown, Figure 3 The horizontal axis represents the signal strength deviation percentage, and the vertical axis represents the quantity. The higher the quantity of the vertical axis, the higher the probability corresponding to the deviation percentage.

[0114] The offset percentage is used to characterize the relative intensity, so that in subsequent sampling, different mass-to-charge ratios or different tested samples can share the same Gaussian distribution data (ie, the offset percentage of the signal intensity), thereby reducing the complexity of the data processing.

[0115] It should be noted that the mass spectrometry data processing system also includes: a data replication module 25, which is used to perform a first data amplification on the original mass spectrometry data by replicating the data before subsequently adjusting the distribution of the signal intensity of the original mass spectrometry data to different degrees, and obtain a number of replicated data corresponding to the original mass spectrometry data.

[0116] The copied data is obtained by copying the original mass spectrum data, that is, the copied data is the same as the corresponding original mass spectrum data. Therefore, the subsequent adjustment of the distribution of the signal intensity of the copied data to different degrees is equivalent to adjusting the distribution of the signal intensity of the original mass spectrum data to different degrees.

[0117] It should be noted that by duplicating data, after subsequent data sampling, it is convenient to use multiple sampling results to adjust each duplicate data separately at the same time, thereby improving the efficiency of data processing.

[0118] It is understandable that in other embodiments, the data replication module may be omitted, and the data amplification module 30 adjusts the distribution of the signal intensity of the original mass spectrum data according to different sampling results, and stores the adjusted data as newly added mass spectrum data.

[0119] It should also be noted that, when there are multiple original mass spectrum data, the number of replicated data corresponding to each original mass spectrum data can be determined according to data requirements.

[0120] In this embodiment, the higher the quality of the original mass spectrum data, the greater the amount of corresponding replicated data.

[0121] The higher the quality of the original mass spectrometry data, the higher the accuracy of the original mass spectrometry data. Accordingly, the more features it embodies, the greater the effect on data analysis (for example, it is conducive to realizing data calculation or classification). Therefore, for original mass spectrometry data with higher quality, the more data amplification is performed, which is more conducive to improving the reliability of the data analysis results. Accordingly, when training the model, the number of reliable training samples can be increased, which is conducive to improving the training effect of the model and making the trained model more accurate.

[0122] As an example, the data replication module 25 includes: a weight setting unit, used to set the weight corresponding to each of the original mass spectrum data based on the quality of the original mass spectrum data, and the higher the quality of the original mass spectrum data, the greater the corresponding weight; a quantity setting unit, used to obtain the number of replicated data corresponding to each of the original mass spectrum data based on the target total number of mass spectrum data and the weight of each of the original mass spectrum data.

[0123] The weights are set based on the quality, thereby increasing the proportion of high-quality mass spectrometry data in the overall data, that is, improving the validity of the data information.

[0124] In this embodiment, the quality of the original mass spectrum data is determined by comparing the original mass spectrum data with standard mass spectrum data. The higher the matching degree, the higher the quality of the original mass spectrum data.

[0125] The degree of matching with standard mass spectrometry data is used as the evaluation criterion, which reduces the complexity of quality judgment.

[0126] Specifically, the method of comparing the original mass spectrum data with the standard mass spectrum data includes: comparing the position distribution of the mass-to-charge ratio of the original mass spectrum data with the position distribution of the mass-to-charge ratio of the standard mass spectrum data. The higher the matching degree of the position distribution, the higher the quality of the original mass spectrum data.

[0127] The signal intensities obtained from testing different samples are often different, and the signal intensity has little significance for quality evaluation. In a single original mass spectrum, the position of each mass-to-charge ratio should theoretically be fixed. Therefore, the quality of the original mass spectrum data is comprehensively evaluated through the position distribution of the mass-to-charge ratio to obtain the overall quality of the single original mass spectrum.

[0128] It should be noted that, in other embodiments, depending on other situations such as the specific type of the raw mass spectrometry data, or actual needs, the quality of the raw mass spectrometry data can also be obtained through other quality assessment methods.

[0129] In this embodiment, since several replicated data corresponding to the original mass spectrum data are first obtained by copying the data, the data amplification module 30 adjusts the signal strength of the several replicated data based on different sampling results to achieve second data amplification and obtain multiple different newly added mass spectrum data.

[0130] That is, for a single original mass spectrum data, a plurality of newly added mass spectrum data different from the original mass spectrum data are obtained through the second data amplification, and the newly added mass spectrum data are also different from each other.

[0131] When there are multiple original mass spectrum data, since each original mass spectrum data shares the same Gaussian distribution data, and the Gaussian distribution is a Gaussian distribution of signal intensity offset percentage, for each mass-to-charge ratio in each original mass spectrum data, the original value of the signal intensity of the mass-to-charge ratio can be adjusted based on the sampled selected offset percentage.

[0132] For example, for each mass-to-charge ratio, the data amplification module 30 adjusts the signal intensity based on the formula I2=I1*(1+m%); wherein I2 represents the signal intensity of the mass-to-charge ratio in the newly added mass spectrum data, I1 represents the signal intensity of the mass-to-charge ratio in the original mass spectrum data, and m% represents the offset percentage of the signal intensity, which can be zero, a positive value, or a negative value. In other words, for each mass-to-charge ratio, after determining the signal intensity offset percentage, the original signal intensity of the mass-to-charge ratio is adjusted to obtain a new signal intensity.

[0133] It can be seen that although the signal intensities of different mass-to-charge ratios may be different, and the signal intensity distributions of different tested samples may also be different, by using a Gaussian distribution of relative intensity, the same calculation method can be used to adjust the original value of the signal intensity of each mass-to-charge ratio for different original mass spectrometry data and different mass-to-charge ratios, thereby reducing the complexity of data amplification.

[0134] In this embodiment, in the same newly added mass spectrum data, the adjustment degrees of the signal intensity of each mass-to-charge ratio relative to the signal intensity of the mass-to-charge ratio in the original mass spectrum data are different.

[0135] Specifically, since a plurality of replicated data corresponding to the original mass spectrum data are obtained by replicating the data, the data amplification module 30 adjusts the signal intensity of each mass-to-charge ratio in the same replicated data to different degrees.

[0136] By varying the degree of change in the signal intensity of each mass-to-charge ratio relative to the original value within the same newly added mass spectrum data, more diverse data can be obtained. Furthermore, when the data obtained using the method described in this embodiment is analyzed, even if data normalization is performed, the diversity of the normalized data can still be maintained.

[0137] For example, for a certain mass-to-charge ratio, the offset of the signal intensity in the newly added mass spectrum data relative to the signal intensity in the original mass spectrum data is +3% (i.e., an increase of 3%), while for another mass-to-charge ratio, the offset of the signal intensity in the newly added mass spectrum data relative to the signal intensity in the original mass spectrum data is -2% (i.e., a decrease of 2%).

[0138] refer to Figure 4 , Figure 4 is a schematic diagram of newly added mass spectrum data according to an embodiment of the present invention, and in order to show the difference between the original mass spectrum data and the newly added mass spectrum data, Figure 4 In the Figure 2 The original mass spectrum data in are shown in the same coordinate.

[0139] Among them, Figure 4 In the figure, the darker data represents the original mass spectrum data, and the lighter data represents the newly added mass spectrum data. Therefore, by performing data amplification, newly added mass spectrum data different from the original mass spectrum data can be obtained.

[0140] Combined with reference Figure 5 , Figure 5 The t-SNE distribution diagram based on mass spectrometry data is shown, where Figure 5 (a) shows a t-SNE distribution diagram of an embodiment obtained based on the original mass spectrum data, Figure 5 (b) shows a t-SNE distribution diagram of an embodiment obtained based on the mass spectrometry data obtained by the method described in this embodiment.

[0141] like Figure 5 As shown, taking the analysis of pathogens as an example, each point represents a data, and the data points in the same circle, ellipse, square or trapezoidal circle belong to the same pathogen category. For example, the pathogens include: Mycobacterium abscessus (M.abscessus), Mycobacterium fortuitum (M.fortuitum), Mycobacterium ulcerans (M.ulcerans), Mycobacterium peregrinum (M.peregrinum), Mycobacterium phlei (M.phlei) and Mycobacterium chelonae (M.chelonae).

[0142] It should be noted that for the convenience of illustration, Figure 5 In the figure, the points in the circle represent Mycobacterium chelonae, the points in the square circle represent Mycobacterium fortuitum, the points in the trapezoidal circle represent Mycobacterium abscessus, the points in the solid oval circle represent Mycobacterium ulcerans, the points in the dotted oval circle represent Mycobacterium peregrinum, and the points in the double-dashed oval circle represent Mycobacterium phlei.

[0143] from Figure 5 As can be seen from (a), the t-SNE distribution diagram obtained by analyzing the samples only with the original mass spectrometry data shows that the number of data points under the same pathogen category is relatively small. Figure 5 As can be seen from (b), when the mass spectrometry data obtained by the method described in this embodiment are used for sample analysis, the number of data points under the same pathogen category increases and the distribution is stable. Moreover, the distribution of the data points does not show serious deformation, and the categories can still be distinguished well.

[0144] refer to Figure 6 , Figure 6 This is a comparative diagram of the accuracy confusion matrix obtained based on the machine learning model. The darker the color of the grid where the number is located, the higher the accuracy.

[0145] in, Figure 6 (a) represents the accuracy confusion matrix of an embodiment of a machine learning model trained based on raw mass spectrometry data, Figure 6 (b) represents an accuracy confusion matrix of an embodiment of a machine learning model obtained by training the mass spectrometry data obtained by the mass spectrometry data processing method described in this embodiment.

[0146] As an example, the model is used to analyze pathogens, and in the confusion matrix, True represents the true result, Predicted represents the predicted result, and the types of pathogens include: Mycobacterium abscessus (M.abscessus), Mycobacterium fortuitum (M.fortuitum), Mycobacterium ulcerans (M.ulcerans), Mycobacterium peregrinum (M.peregrinum), Mycobacterium phlei (M.phlei) and Mycobacterium chelonae (M.chelonae).

[0147] For example, taking True as Mycobacterium chelonae (M. chelonae) as an example, when the machine learning model trained based on the original mass spectrometry data is used for analysis, the probability of predicting an accurate result is 0.83 (or 83%), and there is a 0.17 (or 17%) probability of predicting it as Mycobacterium ulcerans (M. ulcerans).

[0148] like Figure 6 As shown in (a), after training based on the original mass spectrometry data, the overall prediction accuracy of the machine learning model for pathogen categories is 69.40%, as shown in Figure 6 As shown in (b), after training with the mass spectrometry data obtained based on the mass spectrometry data processing method described in this embodiment, the overall prediction accuracy of the machine learning model for pathogen categories is 78.43%, and a higher prediction accuracy can be achieved.

[0149] refer to Figure 7 , Figure 7 This is a comparative diagram of the accuracy confusion matrix obtained based on the deep learning model.

[0150] in, Figure 7 (a) represents the accuracy confusion matrix of an embodiment of a deep learning model trained based on raw mass spectrometry data, Figure 7 (b) represents the accuracy confusion matrix of an embodiment of a deep learning model obtained by training the mass spectral data obtained by the mass spectral data processing method according to an embodiment of the present invention.

[0151] As an example, the model is used to analyze pathogens, and in the confusion matrix, True represents the true result, and Predicted represents the predicted result.

[0152] like Figure 7 As shown in (a), after training based on the original mass spectrometry data, the overall prediction accuracy of the deep learning model for pathogen categories is 75.38%, as shown in Figure 7 As shown in (b), after training with the mass spectrometry data obtained based on the mass spectrometry data processing method described in this embodiment, the overall prediction accuracy of the deep learning model for pathogen categories is 90.33%, and a higher prediction accuracy can be achieved.

[0153] It should be noted that, in this embodiment, the mass spectrometry data processing system is used to implement the mass spectrometry data processing method described in the aforementioned embodiment. The specific description of the mass spectrometry data processing system can be combined with the relevant records in the aforementioned embodiment.

[0154] Correspondingly, an embodiment of the present invention further provides a device, which can implement the mass spectrometry data processing method provided by the embodiment of the present invention by loading the above-mentioned mass spectrometry data processing method in the form of a program.

[0155] refer to Figure 9 , which shows a hardware structure diagram of a device provided by an embodiment of the present invention. The device of this embodiment includes: at least one processor 01, at least one communication interface 02, at least one memory 03 and at least one communication bus 04.

[0156] In this embodiment, the number of each of the processor 01 , the communication interface 02 , the memory 03 and the communication bus 04 is at least one, and the processor 01 , the communication interface 02 and the memory 03 communicate with each other via the communication bus 04 .

[0157] The communication interface 02 may be an interface of a communication module for network communication, such as an interface of a GSM module.

[0158] The processor 01 may be a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the method described in this embodiment.

[0159] Memory 03 may include high-speed RAM memory, or may also include non-volatile memory, such as at least one disk storage device. Memory 03 stores one or more computer instructions, which are executed by processor 01 to implement the mass spectrometry data processing method provided in the aforementioned embodiment.

[0160] It should be noted that the above-mentioned implementation device may also include other devices (not shown) that may not be necessary for understanding the contents disclosed in the embodiments of the present invention; since these other devices may not be necessary for understanding the contents disclosed in the embodiments of the present invention, the embodiments of the present invention will not introduce them one by one.

[0161] An embodiment of the present invention further provides a storage medium storing one or more computer instructions, wherein the one or more computer instructions are used to implement the mass spectrometry data processing method provided in the aforementioned embodiment.

[0162] The embodiments of the present invention described above are combinations of elements and features of the present invention. Unless otherwise mentioned, elements or features may be considered as optional. Each element or feature may be put into practice without being combined with other elements or features. In addition, embodiments of the present invention may be constructed by combining some elements and / or features. The order of operations described in the embodiments of the present invention may be rearranged. Some configurations of any one embodiment may be included in another embodiment and may be replaced by the corresponding configuration of another embodiment. It is obvious to those skilled in the art that claims that do not have a clear reference relationship to each other in the appended claims may be combined into embodiments of the present invention, or may be included as new claims in amendments after submitting this application.

[0163] The embodiments of the present invention may be implemented by various means such as hardware, firmware, software, or a combination thereof. In a hardware configuration, the method according to the exemplary embodiment of the present invention may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, etc.

[0164] In a firmware or software configuration, the embodiments of the present invention may be implemented in the form of modules, procedures, functions, and the like. Software codes may be stored in a memory unit and executed by a processor. The memory unit may be located inside or outside the processor and may send and receive data to and from the processor via various known means.

[0165] The above description of the disclosed embodiments will enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein, but is to be construed in the widest possible manner consistent with the principles and novel features disclosed herein.

[0166] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention should be based on the scope defined by the claims.

Claims

1. A method for processing mass spectrometry data, characterized in that: include: Acquiring raw mass spectrum data, wherein the raw mass spectrum data includes signal intensity that varies with mass-to-charge ratio; Obtaining a Gaussian distribution of the signal intensity, wherein the Gaussian distribution is a Gaussian distribution of a signal intensity offset percentage, and the signal intensity offset percentage refers to: an offset percentage of the signal intensity relative to the signal intensity of the original mass spectrum data; Data sampling is performed on the Gaussian distribution of the signal intensity, and based on the sampling results, the distribution of the signal intensity of the original mass spectrum data is adjusted to different degrees to obtain a plurality of different newly added mass spectrum data corresponding to the original mass spectrum data; data sampling is performed on the Gaussian distribution of the signal intensity, and based on the sampling results, the distribution of the signal intensity of the original mass spectrum data is adjusted to different degrees, and an offset percentage is selected for each mass-to-charge ratio from the Gaussian distribution of the offset percentage, and the signal intensity corresponding to the mass-to-charge ratio in the original mass spectrum data is adjusted based on the selected offset percentage.

2. The method for processing mass spectrum data according to claim 1, wherein: Before adjusting the signal intensity distribution of the raw mass spectrometry data to varying degrees, the method further includes: performing a first data amplification on the original mass spectrum data by copying the data to obtain a plurality of copy data corresponding to the original mass spectrum data; Data sampling is performed on the Gaussian distribution of the signal intensity, and based on the sampling results, the distribution of the signal intensity of the original mass spectrum data is adjusted to different degrees, including: based on different sampling results, the signal intensity of the several replicated data is adjusted respectively to achieve second data amplification and obtain multiple different new mass spectrum data.

3. The method for processing mass spectrum data according to claim 2, wherein: In performing the first data amplification on the original mass spectrum data by copying the data, the higher the quality of the original mass spectrum data is, the greater the number of corresponding copy data is.

4. The method for processing mass spectrum data according to claim 3, wherein: The method of judging the quality of the original mass spectrum data includes: comparing the original mass spectrum data with standard mass spectrum data; the higher the matching degree, the higher the quality of the original mass spectrum data.

5. The method for processing mass spectrum data according to claim 4, wherein: The method of comparing the original mass spectrum data with the standard mass spectrum data includes: comparing the position distribution of the mass-to-charge ratio of the original mass spectrum data with the position distribution of the mass-to-charge ratio of the standard mass spectrum data. The higher the matching degree of the position distribution, the higher the quality of the original mass spectrum data.

6. The method for processing mass spectrum data according to claim 3, wherein: Methods for obtaining the number of replicated data corresponding to each of the original mass spectrometry data include: Based on the quality of the original mass spectrum data, setting a weight corresponding to each of the original mass spectrum data, wherein the higher the quality of the original mass spectrum data, the greater the corresponding weight; Based on the target total number of mass spectrum data and the weight of each of the original mass spectrum data, the number of replicated data corresponding to each of the original mass spectrum data is obtained.

7. The method for processing mass spectrum data according to claim 1, wherein: In obtaining the Gaussian distribution of the signal intensity, a Gaussian distribution of the signal intensity detectable by a detection device is obtained, and the detection device is used to obtain the original mass spectrum data.

8. The method for processing mass spectrum data according to any one of claims 1 to 7, wherein: In each newly added mass spectrum data, the adjustment degree of the signal intensity of each mass-to-charge ratio relative to its signal intensity in the original mass spectrum data is different.

9. The method for processing mass spectrum data according to any one of claims 1 to 7, wherein: The signal intensity of the mass spectrum data is the spectrum peak intensity.

10. A mass spectrometry data processing system, characterized in that: include: A raw data acquisition module, configured to acquire raw mass spectrum data, wherein the raw mass spectrum data includes signal intensity that varies with mass-to-charge ratio; a Gaussian distribution acquisition module, configured to acquire a Gaussian distribution of the signal intensity, wherein the Gaussian distribution is a Gaussian distribution of a signal intensity offset percentage, and the signal intensity offset percentage refers to: an offset percentage of the signal intensity relative to the signal intensity of the original mass spectrum data; A data amplification module is used to perform data sampling on the Gaussian distribution of the signal intensity, and based on the sampling results, adjust the distribution of the signal intensity of the original mass spectrum data to different degrees to obtain multiple different new mass spectrum data corresponding to the original mass spectrum data; perform data sampling on the Gaussian distribution of the signal intensity, and based on the sampling results, adjust the distribution of the signal intensity of the original mass spectrum data to different degrees, select an offset percentage for each mass-to-charge ratio from the Gaussian distribution of the offset percentage, and adjust the signal intensity corresponding to the mass-to-charge ratio in the original mass spectrum data based on the selected offset percentage.

11. A device, characterized in that The method comprises at least one memory and at least one processor, wherein the memory stores one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method for processing mass spectrometry data according to any one of claims 1 to 9.

12. A storage medium, characterized in that: The storage medium stores one or more computer instructions, and the one or more computer instructions are used to implement the method for processing mass spectrometry data according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Mass spectrum data analysis method

    CN107818329A