LC-MS peak detection method, device, equipment and medium
By acquiring the raw data set of the LC-MS instrument, adjusting the peak detection parameters and constructing a lightweight multi-task convolutional neural network, the problem that the existing model cannot directly process the raw signal is solved, and efficient and real-time LC-MS peak detection is achieved, which is adaptable to different mass spectrometry platforms and improves detection accuracy and efficiency.
Patent Information
- Application Number
- CN202510488190.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-09-12
AI Technical Summary
The existing LC-MS peak detection model cannot directly process the original signal, making it difficult to meet real-time requirements. The model needs to be retrained for different mass spectrometry platforms, which limits the universality of the technology.
By acquiring the original data set collected by the LC-MS instrument, extracting and adjusting the peak detection parameters, using the conditional generative adversarial network to generate synthetic peak region feature parameters, and constructing a lightweight multi-task convolutional neural network to achieve peak region classification and boundary definition, adapting to real-time detection needs.
Significantly reduce the false positive rate, improve processing efficiency, enhance cross-instrument generalization capabilities, and achieve high-precision and high-efficiency fully automated mass spectrometry data analysis.
Smart Images

Figure CN120629385A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of analytical chemistry, and in particular to an LC-MS peak detection method, device, equipment and medium. Background Art
[0002] LC-MS (Liquid Chromatograph Mass Spectrometer) technology has been widely used in fields such as environmental monitoring, metabolomics, and drug analysis. Although deep learning technologies have been explored for peak detection (such as DeepIso's feature recognition), these models cannot directly process raw LC-MS signals, making them difficult to meet the real-time requirements of practical scenarios. Furthermore, the models must be retrained for different mass spectrometry platforms, limiting the technology's universal applicability. Summary of the Invention
[0003] The present invention provides an LC-MS peak detection method, device, electronic device and medium to address the technical problems that traditional peak detection models cannot directly process raw LC-MS signals, have difficulty meeting the real-time requirements of actual scenarios, and require retraining the model for different mass spectrometry platforms, limiting the universality of the technology.
[0004] In a first aspect, a LC-MS peak detection method is provided, comprising:
[0005] Obtain raw data sets of environmental samples collected by LC-MS instruments;
[0006] Extracting peak detection parameters from the original data set, adjusting the peak detection parameters, and determining a plurality of peak regions and characteristic parameters of each peak region based on the adjusted peak detection parameters;
[0007] Based on the characteristic parameters of each peak region, preset instrument noise, and preset chromatographic conditions, a conditional generative adversarial network is used to generate a synthetic peak region that is consistent with the characteristic parameter distribution of the real peak region, and the synthetic characteristic parameters of each synthetic peak region are obtained;
[0008] Based on the characteristic parameters and synthetic characteristic parameters, a LC-MS peak detection model was constructed using a lightweight multi-task convolutional neural network. The LC-MS peak detection model includes a peak region classification model for outputting the probability that each peak region belongs to the noise category, the valid peak category, or the category requiring review, and a peak boundary definition model for outputting the probability that each scan point within the peak region belongs to the single peak region and the overlapping region.
[0009] The test data of the sample to be tested is input into the LC-MS peak detection model to obtain the test results.
[0010] In a second aspect, an LC-MS peak detection device is provided, comprising:
[0011] An acquisition module, used to obtain the raw data set of environmental samples collected by the LC-MS instrument;
[0012] an adjustment module for extracting peak detection parameters from the original data set and adjusting the peak detection parameters; a determination module for determining a plurality of peak regions and characteristic parameters of each peak region based on the adjusted peak detection parameters;
[0013] The first generation module is used to generate a synthetic peak region that is consistent with the characteristic parameter distribution of the real peak region through a conditional generative adversarial network based on the characteristic parameters of each peak region, preset instrument noise and preset chromatographic conditions, and obtain synthetic characteristic parameters of each synthetic peak region;
[0014] A construction module for constructing an LC-MS peak detection model using a lightweight multi-task convolutional neural network based on the feature parameters and the synthetic feature parameters, wherein the LC-MS peak detection model includes a peak region classification model for outputting the probability that each peak region belongs to a noise category, a valid peak category, or a category requiring review, and a peak boundary definition model for outputting the probability that each scan point within the peak region belongs to a single peak region and an overlapping region;
[0015] The second generation module is used to input the test data of the test sample into the LC-MS peak detection model to obtain the test results.
[0016] In a third aspect, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned LC-MS peak detection method when executing the computer program.
[0017] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned LC-MS peak detection method are implemented.
[0018] In the scheme implemented by the above-mentioned LC-MS peak detection method, device, electronic device and storage medium, the original data set collected by the environmental sample LC-MS instrument is first obtained, the peak detection parameters are extracted and adjusted to determine multiple peak regions and their characteristic parameters, and then the conditional generative adversarial network is used to generate synthetic peak regions and synthetic characteristic parameters that are consistent with the distribution of the true peak region characteristic parameters. Then, based on the characteristic parameters and synthetic characteristic parameters, a lightweight multi-task convolutional neural network is used to construct an LC-MS peak detection model including a peak region classification model and a peak boundary definition model. Finally, the sample data to be tested is input into the model to obtain the detection result. Compared with traditional deep learning methods (such as DeepIso), there is a problem that they rely on pre-processed features (such as mass spectrum images or peak tables) and cannot directly parse the original LC-MS signal, resulting in delays in the analysis process. The present application directly extracts peak detection parameters from the original data and dynamically adjusts them, and combines the improved lightweight multi-task convolutional network to simultaneously complete the classification of peak regions and peak boundary integration, adapting to real-time detection needs, achieving high-precision detection of true positive peaks and significantly reducing the amount of noise. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0020] Figure 1 1 is a schematic flow chart of an LC-MS peak detection method according to one embodiment of the present invention;
[0021] Figure 2 This is one of the schematic diagrams of the LC-MS peak detection model architecture in one embodiment of the present invention;
[0022] Figure 3 This is a second schematic diagram of the LC-MS peak detection model architecture in one embodiment of the present invention;
[0023] Figure 4 1 is a schematic flow chart of an LC-MS peak detection method according to one embodiment of the present invention;
[0024] Figure 5 1 is a schematic diagram of a confidence-driven interactive peak detection interface in one embodiment of the present invention;
[0025] Figure 6 Schematic diagram of the structure of an LC-MS peak detection device in one embodiment of the present invention. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. It should be understood that the drawings in the present invention are only for the purpose of illustration and description and are not used to limit the scope of protection of the present invention.
[0027] In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in the present invention illustrate operations implemented according to some embodiments of the present invention. It should be understood that the operations in the flowcharts may be implemented out of sequence, and steps that do not have a logical contextual relationship may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the present disclosure, may add one or more other operations to the flowcharts, or may remove one or more operations from the flowcharts.
[0028] In addition, the embodiments described in the present invention are only some of the embodiments of the present invention, rather than all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.
[0029] It should be noted that the term "comprising" will be used in the embodiments of the present invention to indicate the presence of the features subsequently claimed, but does not preclude the addition of other features. It should also be noted that similar reference numerals and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures.
[0030] The following is a detailed description of this case with reference to the relevant drawings in the specification.
[0031] In the embodiments of this specification, the wide application of liquid chromatography-mass spectrometry (LC-MS) in the fields of environmental monitoring, metabolomics, drug analysis, etc. has put forward higher requirements for data processing. Traditional peak detection algorithms (such as XCMS, MZmine2) rely on fixed thresholds and manual parameter adjustment, which makes it difficult to effectively distinguish low-intensity signals from noise, resulting in a high false positive rate (usually more than 20%), especially in complex samples where the overlapping peak integration error is significant. In addition, these tools are highly sensitive to instrument type and chromatographic conditions, and require repeated optimization of parameters based on expert experience. They have a low degree of automation and are time-consuming to process (about 1 hour for a single file), which seriously restricts the efficiency of high-throughput analysis. In recent years, although deep learning technology has been attempted for peak detection (such as feature recognition of DeepIso), its models are mostly designed for pre-processed data, cannot directly process raw LC-MS signals, and rely on high-computing power equipment (such as GPU), making it difficult to meet the real-time requirements of actual scenarios. The shortcomings of existing methods in noise robustness, peak shape generalization ability and end-to-end automation have become bottlenecks in high-resolution mass spectrometry data analysis.
[0032] Furthermore, existing deep learning solutions have two major limitations: first, the model architecture is complex (such as U-Net), the computational cost is high, and it is difficult to adapt to edge devices; second, there is a lack of a credibility assessment mechanism for the detection results, which cannot assist in manual review. For example, although the noise filtering algorithm proposed by Kantz et al. can reduce false positives, it relies on upstream tools to complete peak detection and cannot solve the problems of missed detection or misintegration in the original data. At the same time, existing methods generally ignore cross-instrument compatibility, and the model needs to be retrained for different mass spectrometry platforms, which limits the universality of the technology. Therefore, there is an urgent need for a peak detection method that takes into account high precision, high efficiency and strong generalization capabilities to break through the inherent defects of traditional algorithms and existing deep learning solutions and meet the automation requirements of complex LC-MS data analysis.
[0033] Based on the above problems, this application proposes an LC-MS peak detection and integration method and system based on dynamic parameter optimization and lightweight multi-task learning, aiming to significantly reduce the false positive rate, improve processing efficiency and enhance cross-instrument generalization capabilities. First, a dynamic PA (peak area) detection algorithm is designed to automatically adjust key parameters (such as the maximum number of consecutive zero points and the minimum PA length) through a real-time noise evaluation module to solve the problems of missed detection of low-intensity peaks or misjudgment of noise caused by traditional fixed thresholds. The algorithm combines sliding window statistics with adaptive filtering technology to dynamically optimize detection sensitivity according to the noise characteristics of different mass spectrometers to ensure robustness in complex data scenarios.
[0034] Secondly, a lightweight multi-task convolutional neural network (CNN) architecture is constructed to integrate PA classification and peak integration into a joint learning task. Compared with the traditional U-Net, the present invention uses a deep separable convolution module to replace the standard convolution layer, which reduces the computational complexity by more than 90% while ensuring accuracy, and is suitable for CPU or edge device deployment. The network synchronously outputs the classification probability (noise / peak / uncertainty) and integration boundary by sharing the underlying feature extraction layer, and introduces a hybrid loss function of intersection over union (IoU) and weighted cross entropy to achieve collaborative optimization of classification and integration. Furthermore, based on the Bayesian deep learning framework, the model generates a confidence score for each detected peak, quantifies the uncertainty of the results, and prioritizes pushing low-confidence peaks to the interactive interface for manual review, forming an efficient quality control process with "machine as the main and manual as the auxiliary".
[0035] In addition, this application simulates the noise distribution and peak shape characteristics of various mass spectrometry instruments (such as LC-Q-TOF, Orbitrap) through generative adversarial networks (GANs), constructs a cross-platform synthetic data set, and combines the curriculum learning strategy to progressively train the model from simple to complex samples, significantly improving the generalization ability of unseen instrument data. Experiments show that the false positive rate of this solution can be reduced to below 3%, the single file processing time is shortened to within 30 seconds, and it can be adapted to mainstream high-resolution mass spectrometry equipment without retraining, breaking through the platform dependence limitations of existing tools. Ultimately, the present invention provides a high-precision, high-efficiency, and highly generalized fully automated solution for LC-MS data analysis in the fields of environmental monitoring, metabolomics, and clinical diagnosis.
[0036] See also Figure 1 This embodiment provides an LC-MS peak detection method, which specifically includes the following steps:
[0037] S10: Obtain the raw data set of environmental samples collected by LC-MS instrument.
[0038] It is understood that the execution subject of the present invention may be an LC-MS peak detection device, or a terminal or a server, which is not limited herein. The embodiment of the present invention is described by taking a server as the execution subject as an example.
[0039] LC-MS (Liquid Chromatography-Mass Spectrometry) is a commonly used analytical chemistry technique that combines the separation capabilities of liquid chromatography with the detection capabilities of mass spectrometry. During an experiment, an LC-MS instrument collects a large amount of information about the compounds in an environmental sample, which is stored in an instrument-specific raw data format.
[0040] Specifically, the raw data set involves raw data from multiple LC-MS instruments across different experiments and platforms. The data set includes unprocessed data such as chromatograms (retention time), mass spectra (mass-to-charge ratio, ion intensity), instrument model and serial number, acquisition timestamp, and operator information.
[0041] Through the above method, a large number of original LC-MS data sets covering different categories of environmental samples and experimental conditions are collected to ensure the generalization ability of the model.
[0042] S20: extracting peak region detection parameters from the original data set, adjusting the peak region detection parameters, and determining a plurality of peak regions and characteristic parameters of each peak region based on the adjusted peak region detection parameters.
[0043] In this step, the raw LC-MS data set contains not only key parameters involved in peak region detection (such as mass-to-charge ratio and retention time), but also a large number of other parameters or information not directly involved in peak region detection (such as instrument model and serial number, acquisition timestamp, operator information, etc.). Peak detection parameters for peak regions are extracted from the large amount of raw data.
[0044] Furthermore, the extracted peak region detection parameters are dynamically adjusted to optimize and reduce false positive / false negative peaks, improve subsequent detection sensitivity, and ensure robustness in complex data scenarios. Ultimately, the adjusted peak region detection parameters are used to identify multiple peak regions in the mass spectrometry data (i.e., specific regions in the mass spectrum that may contain compound signals) and the characteristic parameters of each peak region. The characteristic parameters can then be used as label information for each peak region and as training data for subsequent neural network model training.
[0045] In one embodiment of the present application, a specific solution for identifying possible peak regions is provided. In S20, peak detection parameters are extracted from the original data set, the peak detection parameters are adjusted, and multiple peak regions and characteristic parameters of each peak region are determined based on the adjusted peak detection parameters. The solution specifically includes the following steps S21-S24:
[0046] S21: Based on the target format, the original data set is converted to obtain a converted target data set.
[0047] In this step, different LC-MS instruments generate raw data files in specific formats. For example, ThermoFisher instruments often generate .raw files. These files usually contain the raw mass spectrometry data collected by the instrument, but they may be in a proprietary format of the instrument manufacturer, which is not conducive to subsequent general analysis. To solve this problem, this application proposes using a specialized data conversion tool (such as ProteoWizard) to convert the raw data format into an open format such as mzML to obtain the converted target data set.
[0048] S22: extracting peak detection parameters from the target data set, wherein the peak detection parameters include a mass-to-charge ratio matrix and a retention time matrix. The mass-to-charge ratio matrix records the signal intensities of different mass-to-charge ratios corresponding to each scanning point, and the retention time matrix records the retention time corresponding to each scanning point.
[0049] In this step, mass spectrometry data that can be used for subsequent peak detection is extracted from the target dataset converted to an open format as peak detection parameters. Mass spectrometry data is typically stored as a two-dimensional matrix, with one dimension representing the mass-to-charge ratio (m / z) and the other representing the retention time (RT). The mass-to-charge ratio matrix and retention time matrix of the mass spectrometry signal are extracted from the mass spectrometry data. Mass-to-charge ratio and retention time are two important parameters in LC-MS data analysis. The mass-to-charge ratio can be used to determine the molecular weight of a compound, while the retention time reflects the retention characteristics of the compound on the chromatographic column and is related to the compound's chemical properties and structure.
[0050] In the examples of this application, the mass-to-charge ratio matrix and retention time matrix of the mass spectrometry signal are extracted from the target data set by programming or using professional data analysis software as detection parameters for peak region detection that may exist in the raw data. The mass-to-charge ratio matrix records the signal intensities of different mass-to-charge ratios corresponding to each scan point; the retention time matrix records the retention time corresponding to each scan point.
[0051] S23: Determine a retention time axis based on the retention time matrix, and adjust parameters of the retention time axis.
[0052] In this step, the retention time axis represents the passage of time during the liquid chromatography separation process, with its scale corresponding to the duration from the start of sample injection to each time point. In LC-MS analysis, the retention time axis marks the time when the mass spectrometry scan occurs, representing the entire analytical process in the temporal dimension. All retention time data in the retention time matrix is determined based on the retention time axis. The retention time of each peak in the sample or scan is recorded at the corresponding time point on the retention time axis.
[0053] During LC-MS analysis, retention times may drift due to factors such as instrument stability, changes in mobile phase composition, and temperature fluctuations. To eliminate the impact of this drift on data analysis, retention time correction is required to ensure accurate compound identification.
[0054] In one embodiment of the present application, a specific retention time axis parameter adjustment solution is provided. In S23, the parameters of the retention time axis are adjusted, which specifically includes the following steps:
[0055] Setting the sliding window length and performing sliding window sampling on the retention time axis to obtain multiple retention time windows;
[0056] Acquire multiple signal intensities of multiple scanning points in each retention time window;
[0057] Based on the multiple signal intensities, the median intensity of each retention time window is determined;
[0058] Determine the baseline noise region for each retention time window based on multiple signal intensities and median intensity;
[0059] Get the number and standard deviation of noise points in the baseline noise area;
[0060] Each retention time window is adjusted based on the number of noise points and standard deviation.
[0061] In this embodiment, first, the retention time axis is divided into a sliding window of fixed size (for example, every 10 seconds is a window) to analyze noise piece by piece. Continuous scanning points are included in each retention time window, and in each retention time window, the signal intensity of each scanning point is obtained, and based on multiple signal intensities, the median intensity of the retention time window is determined. Thereafter, in each retention time window, the scanning points whose signal intensity is higher than three times of the global median intensity are eliminated to remove obvious high-intensity signals (possibly true peaks), and only low-intensity regions are retained for noise assessment, and the remaining scanning points are the baseline noise regions of the window.
[0062] Furthermore, the number of noise points (i.e., scanning points) contained in the baseline noise region of each retention time window and the signal intensity of each noise point are obtained, and the standard deviation is calculated using the standard deviation calculation formula based on the number of noise points and the signal intensity.
[0063] The formula for calculating standard deviation is:
[0064]
[0065] Where σ is the standard deviation; N is the number of noise points in the baseline noise region of the retention time window; I iis the signal intensity of the i-th noise point; μ is the average signal intensity in the baseline noise area.
[0066] Furthermore, for any retention time window, if the number of noise points in the baseline noise area of the window is less than the preset noise point number threshold (such as N < 5), it means that there are insufficient effective noise points in the window. The statistical characteristics calculated based on these small amounts of data may not accurately reflect the actual noise situation, resulting in estimation bias. At this time, the size of the window is gradually expanded according to certain rules. Increasing the size of the window can include more noise points, thereby increasing the number of samples. For example, the width of the window is increased by a fixed proportion or a fixed number of points each time. In practical applications, it is necessary to determine the appropriate expansion method and expansion amplitude based on the characteristics of the data and the needs of the analysis. This application does not make specific restrictions here. In addition, when there are insufficient effective noise points in the current window, the information of the adjacent windows can also be used for interpolation estimation. Adjacent windows usually have similar noise characteristics, so the mean of the noise standard deviation σ of the adjacent windows can be calculated, and then this mean can be used to replace the σ value of the current window that cannot be accurately calculated due to data sparseness. Specifically, first, determine the range of the adjacent windows, such as selecting one or more windows on the left and right sides of the current window. Thereafter, calculate the value of the noise standard deviation σ of these adjacent windows and find their mean. Finally, this mean is used as the noise standard deviation estimate for the current window. For example, for a one-dimensional data sequence, the current window is the i-th window, and the i-1-th and i+1-th windows are selected, and the average of their σ values is calculated as the σ estimate for the i-th window.
[0067] Furthermore, based on the standard deviation of the baseline intensity within each retention time window, the formula Determines the maximum number of consecutive zero points allowed, thereby reducing the generation of false positive peaks in high noise areas and preserving weak signals in low noise areas. zero ≥1, to avoid peak region breakage caused by too small parameters; the formula parameters 0.3 and 4 are the coefficients that minimize the sum of the false positive rate (FPR) and the false negative rate (FNR) (verified by grid search technology on LC-MS data sets).
[0068] In the above manner, by correcting the retention time axis, the retention times in different samples can be unified to a relatively accurate scale, so that the chromatographic peaks of the same compound in different samples can be more accurately aligned, thereby improving the accuracy of peak matching.
[0069] S24: Based on the mass-to-charge ratio matrix and the adjusted retention time matrix, the optimized centWave algorithm is used to detect the peak region to obtain multiple peak regions and characteristic parameters corresponding to each peak region.
[0070] In this step, centWave algorithm is an algorithm commonly used in peak detection in LC-MS data, searches for local maxima in mass-to-charge ratio matrix and retention time matrix dimensions based on continuous wavelet transform (CWT), and determines the position and boundary of chromatographic peak in combination with the constraints of peak shape. However, this algorithm may have some limitations when processing complex mass spectrometry data, such as unsatisfactory processing of mass deviation and baseline drift. In order to improve the accuracy and reliability of peak detection, the present application embodiment proposes to dynamically expand the scanning range of centWave algorithm and to carry out quadratic polynomial fitting to peak boundary, solves the peak splitting problem caused by mass deviation, and eliminates the influence of baseline drift on integration.
[0071] Specifically, in mass spectrometry, due to factors such as the precision limitation of the instrument, the interference of impurities in the sample, and the uncertainty of the ionization process, the m / z value actually measured may have a certain deviation from the theoretical value. This mass deviation may cause the peak m / z dimension originally belonging to the same compound to split, causing the algorithm to mistakenly identify it as multiple different peaks, thereby affecting subsequent quantitative and qualitative analysis. In order to solve the peak splitting problem caused by mass deviation and make the identification of peaks more accurate, the embodiment of the present application dynamically expands the scanning range to the m / z interval of adjacent ± 0.005Da, and the algorithm can search for peaks that may belong to the same compound in a wider range. In this way, even if there is a certain mass deviation, the peaks that were originally split are more likely to be merged into a complete peak, improving the accuracy and reliability of peak detection. For example, when the theoretical m / z value of a compound is 100.000Da, but the actual measured value may fluctuate between 99.995-100.005Da, after expanding the scanning range, the algorithm can capture all signals within this fluctuation range and process them as a whole. By dynamically expanding the scanning range, false positive and false negative results can be effectively reduced, and the accuracy of analysis can be improved.
[0072] Furthermore, baseline drift is a common problem in LC-MS analysis. It can be caused by factors such as chromatographic column aging, changes in mobile phase composition, and instrument instability. Baseline drift causes the baseline of the chromatographic peak to no longer be a straight line, but rather to exhibit a certain tilt or curvature. This can affect the peak area integration and thus the quantitative analysis results of the compound. To eliminate the impact of baseline drift on integration and improve the accuracy of quantitative analysis, the present embodiment proposes that during peak boundary processing, a smoother and more accurate baseline model can be obtained by fitting the scan points at the peak region boundary with a quadratic polynomial. Once the fitted baseline model is obtained, it can be subtracted from the original signal, eliminating the impact of baseline drift. In this way, when performing peak area integration, the result obtained more accurately reflects the actual content of the compound. For example, in a chromatographic peak with baseline drift, direct integration may also include the baseline drift portion, resulting in an overestimation of the peak area. However, after quadratic polynomial fitting and baseline correction, the integration result more accurately reflects the actual content of the compound.
[0073] Through the above method, the optimized centWave algorithm is used to identify possible peak areas and output the characteristic parameters of each peak area to ensure the accuracy of the analysis.
[0074] S30: Based on the characteristic parameters of each peak region, preset instrument noise, and preset chromatographic conditions, a conditional generative adversarial network is used to generate synthetic characteristic parameters that are consistent with the characteristic parameter distribution of the peak region.
[0075] In this step, in the actual data collection process, due to limitations such as experimental conditions and instrument performance, the characteristic parameters of the real peak area obtained often have certain limitations, and the diversity of the data is insufficient. In order to enable the model to access a wider range of data features, a conditional generative adversarial network is used. With real characteristic parameters as input, the generator simulates different preset instrument noises (such as random noise of Q-TOF, Gaussian noise of Orbitrap, etc.) and preset chromatographic conditions (such as peak tailing caused by gradient delay) to output synthetic characteristic parameters consistent with the real data distribution. On the one hand, it can supplement the missing sample categories in the real data, thereby greatly enriching the diversity of the data and enabling subsequent models to access a wider range of data features; on the other hand, in many cases, obtaining a large amount of real data may face problems such as high cost and long time. Generating synthetic characteristic parameters can quickly expand the data scale without increasing a large amount of experimental cost and time, providing richer training materials for subsequent deep learning models, helping the model to learn more comprehensive features and laws, thereby improving the generalization ability and performance of the model.
[0076] Through the above methods, the original feature parameters are combined, transformed or calculated to obtain new features, so as to mine deeper information in the data and help improve the performance of the model.
[0077] S40: Based on the feature parameters and the synthetic feature parameters, an LC-MS peak detection model is constructed through a lightweight multi-task convolutional neural network, wherein the LC-MS peak detection model includes a peak region classification model for outputting the probability that each peak region belongs to a noise category, a valid peak category, and a category that requires review, and a peak boundary definition model for outputting the probability that each scanning point in the peak region belongs to a single peak region and an overlapping region.
[0078] In this step, the real feature parameters and the synthetic feature parameters are combined as training data, and the LC-MS peak detection model is constructed through a lightweight multi-task convolutional neural network architecture. Traditional convolutional neural networks (CNN) may perform well when processing a single task, but when faced with multiple tasks, if a network is designed separately for each task, the amount of computation and storage requirements will increase significantly. Therefore, the present application proposes the use of a lightweight convolutional neural network architecture, by sharing the feature extraction layer, and after the shared layer, designing a specific branch network for each task to learn task-related features, so that the performance of each task can be improved while ensuring multi-tasking capabilities, multi-tasking can be achieved, and the amount of computation can be greatly reduced. By constructing a lightweight multi-task convolutional neural network (CNN) architecture, peak area classification and peak integration are integrated into a joint learning task. Compared with the traditional U-Net, the present embodiment uses a depth-separable convolution module to replace the standard convolution layer, which reduces the amount of computation by more than 90% while ensuring accuracy, and is suitable for CPU or edge device deployment.
[0079] Specifically, the LC-MS peak detection model includes a peak region classification model and a peak boundary definition model. Among them, the peak classification model is used to output the probability that each peak region identified belongs to the noise category, the probability of belonging to the effective peak category, and the probability of needing to review the category. In practical applications, LC-MS data are often very complex and contain a large amount of noise and overlapping peaks. In order to ensure the quality of LC-MS data detection, this application proposes to classify according to various characteristic parameters of the peak (such as peak height, peak width, retention time, mass-to-charge ratio, peak symmetry, peak area ratio, etc.), and divide each detected peak region into one of three categories: Category 1: The peak region does not contain peaks, but only contains noise; Category 2: The peak region contains one or more peaks; Category 3: The peak region contains some peaks, but requires the special attention of experts, that is, experts need to perform manual review. Among them, the effective peak represents the real chemical composition in the sample, and accurately identifying and analyzing it is the key to subsequent research; the noise peak is caused by factors such as instrument noise and background interference, and it needs to be excluded; and for some peak regions with unclear characteristics and difficult to judge, they are marked as needing to review categories so that further analysis and judgment can be performed by professionals later. The most difficult problem is separating small noise peaks (i.e., category 2) from signals that are too loud, too small, or look strange to be attributed to peaks (i.e., category 3). This application uses the inherent ability of CNN to predict the peak shape of the identified peak area. The output of CNN for peak classification is the probability (from 0 to 1) calculated for assigning the peak area to each of the three categories. The sum of the three probabilities is equal to 1. In practical applications, the peak area is usually classified into the category with the highest probability. By constructing a peak classification model to perform classification based on peak shape features, peaks with noise shapes can be identified.
[0080] Furthermore, traditional peak detection methods are prone to misjudgment when dealing with complex peak shapes. In order to locate the boundaries of the peaks more accurately, this application uses a second convolutional neural network (CNN) to locate the peak boundaries based on the prediction of the peak area itself. Its basic structural idea draws on the U-net. The U-net is a network structure commonly used for image segmentation, with a unique contraction path (downsampling) and expansion path (upsampling). In the contraction path, the network extracts high-level features of the image through convolution and pooling operations to reduce the data dimension; in the expansion path, the feature map is restored to its original size through deconvolution or upsampling operations, and the corresponding feature maps in the contraction path are spliced to retain more detail information. This structure can achieve fast and accurate image segmentation. By constructing a peak boundary definition model, a separation area is added, wherein the separation area refers to the area between adjacent peaks. Accurately identifying the separation area helps to more clearly define the range of each peak and avoid mutual interference between peaks, thereby improving the accuracy of peak boundary determination. Specifically, after the peak classification model predicts a valid peak, the peak boundary definition model accurately defines the peak boundary through upsampling and feature concatenation. It then outputs the probability of each scan point belonging to the peak region, as well as the probability of belonging to the overlapping desegmented region. When the probability of a scan point belonging to the peak region exceeds a set threshold, it is included in the peak range, thereby more accurately defining the peak boundary and avoiding missed or false peak detections.
[0081] Using the above method, two consecutive artificial neural networks were constructed and trained. First, the peak areas were classified into three categories: noise, valid peaks, and those requiring review. The boundaries of each valid peak were further determined to facilitate the integration of the area, ensuring the accuracy and reliability of peak detection, thereby realizing the detection of high-resolution LC-MS data.
[0082] In one embodiment of the present application, a specific model training scheme is provided. In S40, based on the feature parameters and the synthetic feature parameters, a LC-MS peak detection model is constructed by a lightweight multi-task convolutional neural network, which specifically includes the following steps S41-S46:
[0083] S41: Divide the retention time series of each peak area into a number of points of preset dimensions through linear interpolation method, normalize the signal intensity of each peak area based on the maximum signal intensity of each peak area, add a label to each peak area based on the model training target, and combine the feature parameters to obtain the feature vector of each peak area.
[0084] S42: Divide the retention time series of each synthetic peak region into a number of points of preset dimensions through a linear interpolation method, normalize the signal intensity of each synthetic peak region based on the maximum signal intensity of each synthetic peak region, add a label to each synthetic peak region based on the model training objective, and combine the feature parameters to obtain a synthetic feature vector for each synthetic peak region.
[0085] For steps S41-S42, in deep learning, neural networks usually require that the input data have a fixed dimension. However, the number of scanning points of the retention time series of each peak area actually detected may be different, which will bring difficulties to the training and processing of the neural network. In addition, there may be large differences in the signal intensity of different peak areas, and this difference will affect the training effect of the neural network. Based on the above problems, the present application proposes to perform standardization and interpolation processing on the detected peak areas. Specifically, the retention time series of each peak area is unified into 256 scanning points through linear interpolation, which can ensure the dimensional consistency of the input neural network, so that the neural network can process these data more efficiently. Furthermore, based on the maximum signal intensity of the peak area, the overall intensity is scaled to the [0, 1] interval. In addition, in mass spectrometry data, low-intensity peaks often contain important information, but due to their low intensity, they are easily ignored in the overall data. Therefore, the 1% area with signal intensity lower than the maximum value is logarithmically transformed (I′=log10(I+1)) to enhance the feature expression of low-intensity peaks.
[0086] Furthermore, supervised learning requires explicit labels to guide the learning of the characteristics and patterns of each peak region, enabling accurate identification and classification in subsequent predictions. Therefore, labels are added to each peak region based on the model training objectives. These labels include labels indicating whether each peak region belongs to the noise category, valid peak category, or category requiring review; as well as labels indicating whether each scan point within the peak region belongs to a single peak region or an overlapping region. Furthermore, characteristic parameters include the peak region's retention time series, mass-to-charge ratio, peak height, peak width, peak area, peak boundary, and peak region type.
[0087] Optionally, the peak regions are classified, wherein the peak regions include single peak regions and overlapping peak regions containing multiple peaks. When processing overlapping peak regions, in order to more accurately analyze and identify the characteristics of each sub-peak, these sub-peaks need to be segmented and assigned corresponding labels for supervised learning in the training phase.
[0088] Furthermore, the labels are combined with the feature parameters to form the feature vector of each peak area, so as to integrate various information of the peak area and facilitate subsequent processing and analysis by machine learning and deep learning models.
[0089] Through the above method, real data and synthetic data are preprocessed to form real feature vectors and synthetic feature vectors as model training data, which increases the diversity of data and enables the model to be exposed to more different types of data during the training process, thereby improving the generalization ability of the model and enabling it to make more accurate predictions when facing unknown data.
[0090] S43: Acquire peak types of the peak region and the composite peak region, wherein the peak types include single peaks and overlapping peaks.
[0091] S44: Classify the feature vectors and the synthesized feature vectors based on the peak type, and generate a first training data set according to a first preset ratio based on the feature vectors and the synthesized feature vectors corresponding to the single peak, and generate a second training data set according to a second preset ratio based on the feature vectors and the synthesized feature vectors corresponding to the overlapping peaks.
[0092] For steps S43-S44, the unimodal region represents the signal containing only one compound in the region, and its shape is usually more regular, and the peak shape is single, which is easy to identify and analyze. Overlapping peak region means that the signals of multiple compounds overlap each other in the region, and the peak shape is complex, and there may be situations such as peak fusion and shoulder peak. In the true peak region, the eigenvectors of the unimodal region and the overlapping peak region are classified. At the same time, in the synthetic peak region, the synthetic eigenvectors of the unimodal region and the overlapping peak region are classified, and the eigenvectors and synthetic eigenvectors of the unimodal type are summarized as training data for the initial stage of the model. Further, the eigenvectors and synthetic eigenvectors of the overlapping peak type are summarized as training data for the later stage of the model. The peak shape regularity and signal of the unimodal region are relatively easy to identify and analyze. In the initial stage of model training, the eigenvectors and synthetic eigenvectors of the unimodal type are first used for training, which can allow the model to quickly grasp the basic peak feature patterns and laws, and establish a preliminary understanding and recognition ability to peak data. Just like humans learn knowledge, they start with simple content and gradually accumulate experience and skills to lay the foundation for subsequent processing of more complex tasks, thereby improving the generalization ability of the model and enabling it to make more accurate predictions and judgments when faced with unknown peak data.
[0093] Alternatively, in the early stages of model training, the model's understanding of data characteristics is very limited, requiring a large amount of data to learn basic patterns and regularities. The controllable nature of synthetic data can provide a rich and diverse set of training samples, helping the model quickly grasp the characteristics of different data types. Therefore, in the initial stages of model training, the first preset ratio can be 70% synthetic feature vectors. This allows the model to quickly be exposed to a large amount of peak data with diverse variations, accelerating learning and improving its generalization ability. As training progresses, the model matures and requires data that is more realistic to further optimize performance. While real data may have some issues, it contains the complexities and noise of real applications and is the real-world scenario the model will ultimately face. Therefore, in the second preset ratio, 90% real feature vectors are used. Increasing the proportion of real data allows the model to better adapt to real-world scenarios and improves its accuracy and reliability in real-world applications. This gradual increase in the proportion of real data aligns with human learning patterns. Just like when learning new knowledge, people typically begin with theoretical and simulated examples to grasp basic concepts and methods, and then consolidate and deepen their understanding through real-world examples. The same is true for the model during training. First, basic recognition capabilities are established through synthetic data, and then fine-tuned through real data to better improve the performance of the model.
[0094] S45: Inputting the first training data set and the second training data set into the lightweight multi-task convolutional neural network model, performing backpropagation and gradient calculation on the model parameters based on the classification task cross entropy loss function and the integration task intersection-over-union loss function, and using the AdamW optimizer to iteratively update the parameters of the lightweight multi-task convolutional neural network model according to the gradient calculation results;
[0095] S46: When the change in the damage function is less than a preset threshold or the number of iterations is equal to the maximum number of iterations, a lightweight multi-task convolutional neural network model is output as an LC-MS peak detection model.
[0096] For steps S45-S46, the training phase employs a phased optimization strategy. Initial pre-training is performed on a mixed set of 70% synthetic data and 30% real data. A curriculum learning strategy is used to gradually transition from single-peak samples to complex overlapping-peak samples. The loss function is designed as a weighted combination of the cross entropy of the classification task and the intersection over union (IoU) of the integration task (weight ratio 1:2). KL divergence is introduced to constrain the parameter distribution of the Bayesian neural network to quantify prediction uncertainty. AdamW is used as the optimizer, with an initial learning rate of 3e-4. A cosine annealing scheduler is used to dynamically adjust the learning rate to avoid local optima. After each round of training, model accuracy and generalization are evaluated on an independent validation set. If classification accuracy does not improve for five consecutive rounds, an early stopping mechanism is automatically triggered to save the optimal weights. This early stopping mechanism prevents model overfitting, preventing the model from overlearning on the training set while performance on the validation set no longer improves. Saving the optimal weights ensures that the best model obtained during training is retained for subsequent testing and application.
[0097] Optionally, the lightweight multi-task convolutional neural network model includes: a shared feature extraction layer, a classification branch layer, and an integration branch layer, wherein the classification branch layer is used to output the probability that the peak area belongs to the noise category, the effective peak category, and the category that needs to be reviewed; the integration branch layer is used to output the probability that each detection point belongs to the peak area and the overlapping area. Specifically, Figure 2 and Figure 3 As shown in the figure, it is a schematic diagram of the LC-MS peak detection model architecture, where: Figure 2 Schematic diagram of the peak region classification model architecture. Figure 3 Schematic diagram of the model architecture for defining peak boundaries. The constructed dual-path network consists of a shared feature extraction layer, a classification branch, and an integration branch. This design allows for simultaneous processing of different tasks. The shared feature extraction layer extracts common features from the input data, while the classification and integration branches respectively perform classification and peak region localization based on these shared features. This enables multi-task parallel processing, improving network efficiency and performance. The shared layer utilizes a modified MobileNetV3 module, replacing traditional convolutions with depthwise separable convolutions and embedding a channel-wise attention mechanism. This reduces the number of parameters by 90% while preserving key peak shape features. The classification branch, consisting of a global average pooling layer and a fully connected layer, determines whether a peak is noise, a valid peak, or a category requiring review. Furthermore, the integration branch leverages the skip connection structure of the U-Net to accurately locate peak boundaries through upsampling and feature concatenation. It then outputs a probability mask indicating whether each scan point belongs to a peak region or an overlap segment. This probability mask is then used to determine the peak boundary of the peak region using a code.
[0098] Through this approach, by dynamically adjusting the mix of synthetic and real data, initial training is based on 70% synthetic data, gradually transitioning to 90% real data, achieving the desired learning effect. Ultimately, the preprocessed data is fed into the downstream neural network in tensor form (dimension: number of channels × 256), completing the end-to-end peak detection and integration tasks. This allows the model to gradually learn more advanced and complex features, improving its generalization capabilities and enabling it to better cope with complex and changing data in the real world.
[0099] In one embodiment of the present application, a specific training data set selection scheme is provided, that is, the model training process specifically includes the following steps:
[0100] During model training, get the current training round;
[0101] If the current training round is less than or equal to the preset training round threshold, training the model based on the first training data set;
[0102] If the current training round is greater than the preset training round threshold, the model is trained based on the second training data set.
[0103] In this embodiment, during the model training process, each completed pass through the entire training dataset is called a training epoch. The first training dataset is relatively simple and regular, containing less noise and relatively simple patterns or features. In the early stages of training, the model's parameters are randomly initialized, and its ability to understand the data and extract features is relatively weak. Using a simple, single-peaked first training dataset allows the model to quickly learn some basic patterns and features, establish a preliminary cognitive framework, avoid getting lost in complex data, reduce the difficulty and complexity of training, and help the model converge quickly. The second training dataset is often more complex and diverse, and may contain more noise and more complex patterns or features. After a certain number of rounds of training, after the model has learned basic features and patterns on the first training dataset, its parameters have been optimized to a certain extent, and it has a certain degree of learning and adaptability. At this time, introducing a more complex second training dataset allows the model to further learn more advanced and complex features, improve the model's generalization ability, and enable it to better cope with complex and changing data in the real world.
[0104] By using data sets of different difficulty levels in stages, we can avoid the problem of using complex data at the beginning, which may cause the model to have difficulty converging or fall into a local optimal solution. This allows the model to learn effective features in a shorter time and speed up training.
[0105] S50: Inputting the test data of the test sample into the LC-MS peak detection model to obtain the test result.
[0106] In this step, the data to be tested is the LC-MS data of the sample to be tested collected by the LC-MS instrument. These data contain the characteristic information of various chemical components in the sample. The data to be tested is input into the trained LC-MS peak detection model. The model performs characteristic detection on the sample to be tested through three steps: peak area detection, peak area classification, and effective peak integration. First, possible peak areas are identified from the complex LC-MS data to determine the approximate starting and ending positions of each peak. After the peak area is detected, the model needs to classify these peak areas to determine whether each peak belongs to the noise category, the effective peak category, or the category that needs to be reviewed. Subsequently, the peak boundary definition model is used to detect the scanning points within the effective peak, and the probability of each scanning point belonging to the peak area and the probability of belonging to the overlapping area are predicted, so as to more accurately define the peak boundary and avoid missed or false peaks.
[0107] In one embodiment of the present application, a specific data detection scheme is provided. In S50, the test data of the test sample is input into the LC-MS peak detection model to obtain the detection result, which specifically includes the following steps S51-S54:
[0108] S51: Input the data to be tested into the peak region classification model, detect the peak region contained in the data to be tested, and the characteristic parameters of the peak region, and based on the characteristic parameters of the peak region, determine the probability that the peak region belongs to the noise category, the effective peak category, and the category that needs to be reviewed.
[0109] In this step, the LC-MS data to be tested is fed into a trained peak region classification model. The model identifies possible peak regions based on the data's characteristic patterns, thereby obtaining the peak regions contained in the data and their characteristic parameters. Based on these characteristic parameters, the model quantifies the likelihood of the peak regions belonging to different categories, calculating the probability of the peak regions belonging to the noise category, the valid peak category, and the category requiring manual review.
[0110] S52: Determine the category of the peak region based on multiple probabilities.
[0111] In this step, the sum of the probabilities of the peak region belonging to different categories is 1. The maximum probability principle can be adopted, that is, the category with the highest probability is selected as the final category of the peak region. For example, if the probability of the peak region belonging to the valid peak category is the highest, it is determined as the valid peak category.
[0112] S53: If the peak region belongs to the valid peak category, the characteristic parameters of the peak region are input into the peak boundary definition model, and the probability of each scanning point belonging to the peak region and the overlapping region is output.
[0113] S54: If the peak region belongs to the category requiring review, the characteristic parameters of the peak region are sent to the terminal of the R&D personnel.
[0114] In steps S53-S54, the peak boundary definition model learns the features and boundary information of the peak region. It can then accurately locate the peak boundary through upsampling and feature concatenation, determining the probability mask for each scan point in the valid peak to belong to the peak region or the overlapping region segmentation area. These probabilities can help determine the start and end positions of the peak, as well as its overlap with other peaks. The probability mask for each scan point in the peak region to belong to the peak region or the overlapping region segmentation area is then input into a preset code to obtain the peak boundary of the peak region.
[0115] Furthermore, if the peak area belongs to the category that requires manual review, it means that the peak area is relatively complex and the model classification result is relatively uncertain. At this time, the characteristic parameters of the peak area are sent to the developer, and manual judgment is made with the help of the developer's professional knowledge to improve the accuracy of the classification.
[0116] In one embodiment of the present application, Figure 4 As shown, a specific data preprocessing solution is provided, which specifically includes the following steps:
[0117] Step 1: Convert raw LC-MS data into standard mzML files;
[0118] Step 2: Calculate the standard deviation of each retention time window using the sliding window noise algorithm;
[0119] Step 3: Dynamically adjust peak detection parameters based on standard deviation;
[0120] Step 4: Use the improved peak detection algorithm to generate peak regions;
[0121] Step 5: Normalize and interpolate the peak area to obtain the standardized peak area.
[0122] In an embodiment of the present application, first, the original LC-MS data is converted into a standard mzML format, and the mass-to-charge ratio (m / z) and retention time (RT) matrix of the mass spectrometry signal are extracted. Secondly, the retention time (RT) axis is divided into sliding windows of fixed size (for example, one window every 10 seconds), and each window contains continuous mass spectrometry scanning points. In each window, signal points with an intensity 3 times higher than the global median intensity are eliminated to avoid residual peak signals interfering with noise estimation, and the remaining points are regarded as baseline noise areas. The standard deviation of the signal intensity values in the baseline noise area is calculated. If there are insufficient effective noise points in the window (such as N<5), the window range is expanded or the σ mean interpolation of adjacent windows is used to avoid estimation bias caused by local data sparsity. Based on the standard deviation σ of the baseline intensity in each RT window, according to the formula Determines the maximum number of consecutive zero points allowed, thereby reducing the generation of false positive peaks in high noise areas and preserving weak signals in low noise areas.zero ≥1 to avoid peak region splitting due to excessively small parameters; the formula parameters 0.3 and 4 are coefficients that minimize the sum of the false positive rate (FPR) and the false negative rate (FNR), which can be verified using grid search technology on LC-MS datasets. Subsequently, an improved centWave algorithm was used to detect peak regions. Its core improvements include: dynamically expanding the scan range to adjacent m / z intervals of ±0.005 Da to avoid peak splitting due to mass deviation; and fitting a quadratic polynomial to the peak region boundaries to eliminate the impact of baseline drift on integration.
[0123] Furthermore, the detected peak regions are normalized and interpolated. The retention time series of each peak region is unified into 256 data points through linear interpolation to ensure the dimensional consistency of the input neural network. The signal intensity adopts a piecewise normalization strategy: based on the maximum signal intensity value of the peak region, the overall signal intensity is scaled to the [0, 1] interval, and the 1% area with intensity lower than the maximum value is logarithmically transformed (I′=log10(I+1)) to enhance the feature expression of low-intensity peaks. For peak regions containing multiple peaks, a method combining local maximum detection and Gaussian fitting is used to automatically segment sub-peaks and generate corresponding labels for supervised learning in the training phase.
[0124] In actual application scenarios, the LC-MS peak detection method provided in this application deploys a quantized lightweight model to the target hardware (CPU / FPGA). For FPGA devices, 8-bit integer quantization and instruction set parallelization technology are used to increase the model inference speed to 50 peak areas per second. For each detected peak area, the model synchronously outputs the category probability (such as "valid peak" probability> 85%, it automatically passes) and the confidence score (based on the Monte Carlo Dropout sampling variance calculation), and the low confidence peak (confidence <70%) will be pushed to the interactive interface and highlighted for manual review. In addition, it also supports real-time processing and offline batch processing modes, and integrates with the laboratory environment sample information management system through a RESTful API to achieve fully automated pipeline analysis from raw data to peak tables. Furthermore, the system deployment scheme in the embodiment of the present application optimizes hardware adaptation and software architecture for different application scenarios, and is specifically implemented as follows: First, the edge computing device is accelerated. In order to meet real-time requirements, an FPGA and CPU collaborative computing architecture is adopted. The lightweight multi-task CNN model was converted to an 8-bit integer quantization format using the TensorRT toolchain and deployed to the customized logic unit of an FPGA. Parallel pipelining was used to achieve high-speed inference (single PA processing time ≤ 2ms). The FPGA communicated with the main CPU via the PCIe interface, dynamically allocating computing tasks: dense PA regions were processed by the FPGA, while sparse regions were processed by CPU multithreading, achieving an overall throughput of 50 PAs per second. For environments without an FPGA, the model was further optimized to the OpenVINO-supported IR format and accelerated on an Intel CPU using the AVX-512 instruction set, achieving single-file processing time under 30 seconds. Secondly, a cloud-edge collaborative architecture was adopted, supporting a distributed deployment model. Real-time peak detection was performed at the edge, while model updates and big data analysis were handled by the cloud. Edge devices encryptedly uploaded raw data of low-confidence peaks (confidence <70%) and peak region features to the cloud via the MQTT protocol, triggering online model fine-tuning based on reinforcement learning. The optimized weight increments were then distributed to the edge nodes. At the same time, a cross-laboratory peak database is built on the cloud, multi-source data is aggregated through a federated learning framework, and global model versions are periodically generated to ensure continuous evolution of the algorithm. Furthermore, the new system is integrated with the laboratory environment sample information management system. The following functions are achieved through seamless docking capabilities provided by standardized RESTful APIs: Data input: supports streaming uploads in formats such as mzML and mzXML, and automatically parses instrument metadata (such as LC gradient programs, mass spectrometry resolution); Result output: Generates a peak table in CSV / JSON format, including m / z, RT, area and confidence fields, and comes with a visual map (PNG / SVG); Quality control feedback: The manual review results are reverse-annotated to the database to trigger model retraining.The system is compatible with mainstream platforms such as OpenLIMS and LabWare, and provides Python / Java SDK to simplify secondary development. In addition, the system has an interactive visual interface, such as. Figure 5 Figure 1 shows a schematic diagram of the confidence-driven interactive peak detection interface. The "File Data Advanced" section is an advanced feature for file data in the current interface. This section displays detailed file information. "Opened files" shows the list of files currently being operated on / loaded; "Detected features" displays detected features; and "Calculated cossims" displays calculated cosine similarities. Furthermore, in the chromatogram coordinate system shown in the figure, the horizontal axis represents retention time in minutes, and the vertical axis represents peak signal intensity. A browser-based 2D chromatogram rendering engine, built using WebGL technology, implements the following features: dynamic confidence mapping: peak detection confidence is overlaid as a heat map, with low-confidence areas automatically highlighted; multi-dimensional filtering: peak list filtering by m / z range, RT interval, and confidence threshold; and manual correction tools: interactive controls such as a peak segmentation brush and baseline correction slider are provided, with user correction results synchronized to the model fine-tuning queue in real time. The interface is mobile-friendly and uses WebAssembly technology to enable offline operation, meeting the needs of network-restricted scenarios such as field work and cleanrooms. Finally, the system deployment also considered cross-platform security and compatibility. The system is packaged as a Docker image, supporting one-click deployment on Windows, Linux, and macOS. A lightweight blockchain module is built in for audit tracking, ensuring data immutability. The communication layer utilizes a dual-certificate system (device certificate + user certificate) and the national SM4 encryption algorithm, meeting GLP / GMP compliance requirements. Through an abstract instrument driver layer, it can be expanded to connect to mainstream mass spectrometers such as Thermo, Agilent, and Waters, achieving plug-and-play cross-platform compatibility.
[0125] It can be seen that in the above scheme, the original data set collected by the LC-MS instrument of the environmental sample is first obtained, the peak detection parameters are extracted and adjusted to determine multiple peak regions and their characteristic parameters, and then the conditional generative adversarial network is used to generate a synthetic peak region and synthetic characteristic parameters that are consistent with the distribution of the real peak region characteristic parameters. Then, based on the characteristic parameters and synthetic characteristic parameters, a lightweight multi-task convolutional neural network is used to construct an LC-MS peak detection model including a peak region classification model and a peak boundary definition model. Finally, the sample data to be tested is input into the model to obtain the test results. Compared with traditional deep learning methods (such as DeepIso), there is a problem that they rely on pre-processed features (such as mass spectrometry images or peak tables) and cannot directly parse the original LC-MS signal, resulting in a delay in the analysis process. This application directly extracts peak detection parameters from the original data and dynamically adjusts them, and combines the improved lightweight multi-task convolutional network to simultaneously complete the classification of peak regions and peak boundary integration, adapting to real-time detection needs, achieving high-precision detection of true positive peaks and significantly reducing the amount of noise.
[0126] In one embodiment, a LC-MS peak detection device is provided, which corresponds one-to-one with the LC-MS peak detection method in the above embodiment. Figure 6 As shown, the LC-MS peak detection device 100 includes: an acquisition module 101, an adjustment module 102, a determination module 103, a first generation module 104, a construction module 105, and a second generation module 106. The functional modules are described in detail as follows:
[0127] An acquisition module 101 is used to acquire an original data set of environmental samples collected by an LC-MS instrument;
[0128] An adjustment module 102 is configured to extract peak detection parameters from the original data set and adjust the peak detection parameters. A determination module 103 is configured to determine a plurality of peak regions and characteristic parameters of each peak region based on the adjusted peak detection parameters.
[0129] The first generation module 104 is configured to generate, based on the characteristic parameters of each peak region, preset instrument noise, and preset chromatographic conditions, a synthetic peak region having a distribution consistent with the characteristic parameters of the real peak region through a conditional generative adversarial network, and obtain synthetic characteristic parameters of each synthetic peak region;
[0130] A construction module 105 is configured to construct an LC-MS peak detection model using a lightweight multi-task convolutional neural network based on the characteristic parameters and the synthetic characteristic parameters, wherein the LC-MS peak detection model includes a peak region classification model for outputting the probability that each peak region belongs to a noise category, a valid peak category, or a category requiring review, and a peak boundary definition model for outputting the probability that each scan point within the peak region belongs to a single peak region and an overlapping region;
[0131] The second generating module 106 is used to input the test data of the test sample into the LC-MS peak detection model to obtain the detection result.
[0132] In one embodiment, the adjustment module 102 is specifically configured to:
[0133] Based on the target format, the original data set is converted to obtain the converted target data set;
[0134] Extract peak detection parameters from the target data set, where the peak detection parameters include a mass-to-charge ratio matrix and a retention time matrix. The mass-to-charge ratio matrix records the signal intensities of different mass-to-charge ratios corresponding to each scanning point, and the retention time matrix records the retention time corresponding to each scanning point.
[0135] Based on the retention time matrix, the retention time axis is determined, and parameters of the retention time axis are adjusted.
[0136] In one embodiment, the determination module 103 is specifically configured to:
[0137] Based on the mass-to-charge ratio matrix and the adjusted retention time matrix, the optimized centWave algorithm is used to detect the peak region to obtain multiple peak regions and the characteristic parameters corresponding to each peak region.
[0138] In one embodiment, the adjustment module 102 is further configured to:
[0139] Setting the sliding window length and performing sliding window sampling on the retention time axis to obtain multiple retention time windows;
[0140] Acquire multiple signal intensities of multiple scanning points in each retention time window;
[0141] Based on the multiple signal intensities, the median intensity of each retention time window is determined;
[0142] Determine the baseline noise region for each retention time window based on multiple signal intensities and median intensity;
[0143] Get the number and standard deviation of noise points in the baseline noise area;
[0144] Each retention time window is adjusted based on the number of noise points and standard deviation.
[0145] In one embodiment, the construction module 105 is specifically configured to:
[0146] The retention time series of each peak region is divided into a number of points of a preset dimension by linear interpolation. The signal intensity of each peak region is normalized based on the maximum signal intensity of each peak region. A label is added to each peak region based on the model training target, and the feature parameters are combined to obtain the feature vector of each peak region.
[0147] The retention time series of each synthetic peak region is divided into a number of points of a preset dimension by a linear interpolation method. The signal intensity of each synthetic peak region is normalized based on the maximum signal intensity of each synthetic peak region. A label is added to each synthetic peak region based on the model training target, and the feature parameters are combined to obtain a synthetic feature vector for each synthetic peak region.
[0148] Obtaining peak types of peak regions and composite peak regions, wherein the peak types include single peak regions and overlapping peak regions;
[0149] Classifying the eigenvectors and the synthesized eigenvectors based on the peak type, generating a first training data set according to a first preset ratio based on the eigenvectors and the synthesized eigenvectors corresponding to the single peaks, and generating a second training data set according to a second preset ratio based on the eigenvectors and the synthesized eigenvectors corresponding to the overlapping peaks;
[0150] The first and second training datasets are input into a lightweight multi-task convolutional neural network model, and the model parameters are back-propagated and gradient calculated based on the cross-entropy loss function of the classification task and the intersection-over-union loss function of the integration task. The parameters of the lightweight multi-task convolutional neural network model are iteratively updated using the AdamW optimizer according to the gradient calculation results.
[0151] When the change in the damage function is less than a preset threshold or the number of iterations is equal to the maximum number of iterations, the output is a lightweight multi-task convolutional neural network model as the LC-MS peak detection model.
[0152] In one embodiment, the acquisition module 101 is further used to obtain the current training round during the model training process.
[0153] In one embodiment, the apparatus further comprises:
[0154] Training modules, specifically for:
[0155] If the current training round is less than or equal to the preset training round threshold, training the model based on the first training data set;
[0156] If the current training round is greater than the preset training round threshold, the model is trained based on the second training data set.
[0157] In one embodiment, a lightweight multi-task convolutional neural network model includes: a shared feature extraction layer, a classification branch layer, and an integration branch layer, wherein the classification branch layer is used to output the probability that the peak area belongs to the noise category, the valid peak category, or the category that needs to be reviewed; the integration branch layer is used to output the probability that each detection point belongs to the peak area and the overlapping area.
[0158] In one embodiment, the second generating module 106 is specifically configured to:
[0159] Inputting the data to be tested into the peak region classification model, detecting the peak regions contained in the data to be tested and the characteristic parameters of the peak regions, and determining the probability of the peak regions belonging to the noise category, the valid peak category, and the category requiring review based on the characteristic parameters of the peak regions;
[0160] Based on multiple probabilities, the category of each peak area is determined;
[0161] If the peak region belongs to the valid peak category, the model is defined by the peak boundary, and the probability of each scan point in the peak region belonging to the peak region and the overlapping region is output;
[0162] If the peak region belongs to the category requiring review, the characteristic parameters of the peak region are sent to the terminal of the R&D personnel.
[0163] The present invention provides an LC-MS peak detection device 100, which first obtains the original data set collected by the environmental sample LC-MS instrument, extracts and adjusts the peak detection parameters to determine multiple peak regions and their characteristic parameters, and then uses a conditional generative adversarial network to generate a synthetic peak region and synthetic characteristic parameters that are consistent with the distribution of the real peak region characteristic parameters. Then, based on the characteristic parameters and the synthetic characteristic parameters, a lightweight multi-task convolutional neural network is used to construct an LC-MS peak detection model including a peak region classification model and a peak boundary definition model. Finally, the sample data to be tested is input into the model to obtain the detection result. Compared with traditional deep learning methods (such as DeepIso), there is a problem that they rely on pre-processed features (such as mass spectrum images or peak tables) and cannot directly parse the original LC-MS signal, resulting in a delay in the analysis process. The present application directly extracts peak detection parameters from the original data and dynamically adjusts them, and combines an improved lightweight multi-task convolutional network to simultaneously complete the classification of peak regions and peak boundary integration, adapting to real-time detection needs, achieving high-precision detection of true positive peaks and significantly reducing the amount of noise.
[0164] The specific definition of the LC-MS peak detection device can be found in the definition of the LC-MS peak detection method above and will not be repeated here. The various modules in the above-mentioned LC-MS peak detection device can be implemented in whole or in part by software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the electronic device in hardware form, or can be stored in the memory of the electronic device in software form, so that the processor can call and execute the corresponding operations of the above-mentioned modules.
[0165] In one embodiment, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0166] Obtain raw data sets of environmental samples collected by LC-MS instruments;
[0167] Extracting peak detection parameters from the original data set, adjusting the peak detection parameters, and determining a plurality of peak regions and characteristic parameters of each peak region based on the adjusted peak detection parameters;
[0168] Based on the characteristic parameters of each peak region, preset instrument noise, and preset chromatographic conditions, a conditional generative adversarial network is used to generate a synthetic peak region that is consistent with the characteristic parameter distribution of the real peak region, and the synthetic characteristic parameters of each synthetic peak region are obtained;
[0169] Based on the characteristic parameters and synthetic characteristic parameters, a LC-MS peak detection model was constructed using a lightweight multi-task convolutional neural network. The LC-MS peak detection model includes a peak region classification model for outputting the probability that each peak region belongs to the noise category, the valid peak category, or the category requiring review, and a peak boundary definition model for outputting the probability that each scan point within the peak region belongs to the single peak region and the overlapping region.
[0170] The test data of the sample to be tested is input into the LC-MS peak detection model to obtain the test results.
[0171] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0172] Obtain raw data sets of environmental samples collected by LC-MS instruments;
[0173] Extracting peak detection parameters from the original data set, adjusting the peak detection parameters, and determining a plurality of peak regions and characteristic parameters of each peak region based on the adjusted peak detection parameters;
[0174] Based on the characteristic parameters of each peak region, preset instrument noise, and preset chromatographic conditions, a conditional generative adversarial network is used to generate a synthetic peak region that is consistent with the characteristic parameter distribution of the real peak region, and the synthetic characteristic parameters of each synthetic peak region are obtained;
[0175] Based on the characteristic parameters and synthetic characteristic parameters, a LC-MS peak detection model was constructed using a lightweight multi-task convolutional neural network. The LC-MS peak detection model includes a peak region classification model for outputting the probability that each peak region belongs to the noise category, the valid peak category, or the category requiring review, and a peak boundary definition model for outputting the probability that each scan point within the peak region belongs to the single peak region and the overlapping region.
[0176] The test data of the sample to be tested is input into the LC-MS peak detection model to obtain the test results.
[0177] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or electronic device can be referred to the relevant descriptions on the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0178] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0179] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0180] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A LC-MS peak detection method, characterized in that include: Obtain raw data sets of environmental samples collected by LC-MS instruments; Extracting peak detection parameters from the original data set, adjusting the peak detection parameters, and determining a plurality of peak regions and characteristic parameters of each peak region based on the adjusted peak detection parameters; Based on the characteristic parameters of each peak region, preset instrument noise, and preset chromatographic conditions, generating a synthetic peak region consistent with the characteristic parameter distribution of the real peak region through a conditional generative adversarial network, and obtaining synthetic characteristic parameters of each synthetic peak region; Based on the characteristic parameters and the synthesized characteristic parameters, an LC-MS peak detection model is constructed using a lightweight multi-task convolutional neural network, wherein the LC-MS peak detection model includes a peak region classification model for outputting the probability that each peak region belongs to a noise category, a valid peak category, and a category requiring review, and a peak boundary definition model for outputting the probability that each scanning point in the peak region belongs to a single peak region and an overlapping region; The test data of the test sample is input into the LC-MS peak detection model to obtain the test results.
2. The method according to claim 1, characterized in that The steps of extracting peak detection parameters from the original data set, adjusting the peak detection parameters, and determining a plurality of peak regions and characteristic parameters of each peak region based on the adjusted peak detection parameters specifically include: Based on the target format, the original data set is format converted to obtain a converted target data set; Extracting the peak detection parameters from the target data set, wherein the peak detection parameters include a mass-to-charge ratio matrix and a retention time matrix, the mass-to-charge ratio matrix records the signal intensities of different mass-to-charge ratios corresponding to each scanning point, and the retention time matrix records the retention time corresponding to each scanning point; Determining a retention time axis based on the retention time matrix, and adjusting parameters of the retention time axis; Based on the mass-to-charge ratio matrix and the adjusted retention time matrix, the peak region is detected using the optimized centWave algorithm to obtain multiple peak regions and characteristic parameters corresponding to each peak region.
3. The method according to claim 2, characterized in that The step of adjusting parameters of the retention time axis specifically includes: Setting a sliding window length and performing sliding window sampling on the retention time axis to obtain a plurality of retention time windows; Acquire multiple signal intensities of multiple scanning points in each retention time window; determining a median intensity of each retention time window based on the multiple signal intensities; determining a baseline noise region for each retention time window based on the plurality of signal intensities and the median intensity; Obtaining the number and standard deviation of noise points in the baseline noise area; Each retention time window is adjusted based on the number of noise points and the standard deviation.
4. The method according to claim 1, wherein The step of constructing an LC-MS peak detection model based on the characteristic parameters and the synthetic characteristic parameters through a lightweight multi-task convolutional neural network specifically includes: The retention time series of each peak region is divided into a number of points of a preset dimension by linear interpolation. The signal intensity of each peak region is normalized based on the maximum signal intensity of each peak region. A label is added to each peak region based on the model training target, and the feature parameters are combined to obtain the feature vector of each peak region. The retention time series of each synthetic peak region is divided into a number of points of a preset dimension by a linear interpolation method. The signal intensity of each synthetic peak region is normalized based on the maximum signal intensity of each synthetic peak region. A label is added to each synthetic peak region based on the model training target, and the feature parameters are combined to obtain a synthetic feature vector for each synthetic peak region. Obtaining peak types of the peak region and the composite peak region, wherein the peak types include single peak regions and overlapping peak regions; Classifying the feature vectors and the synthesized feature vectors based on the peak type, generating a first training data set according to a first preset ratio based on the feature vectors and the synthesized feature vectors corresponding to the single peak, and generating a second training data set according to a second preset ratio based on the feature vectors and the synthesized feature vectors corresponding to the overlapping peaks; Inputting the first training data set and the second training data set into a lightweight multi-task convolutional neural network model, performing backpropagation and gradient calculation on the model parameters based on the classification task cross entropy loss function and the integration task intersection-over-union loss function, and using the AdamW optimizer to iteratively update the parameters of the lightweight multi-task convolutional neural network model according to the gradient calculation results; When the change in the damage function is less than a preset threshold or the number of iterations is equal to the maximum number of iterations, the output is a lightweight multi-task convolutional neural network model as the LC-MS peak detection model.
5. The method according to claim 4, characterized in that Also includes: During model training, get the current training round; If the current training round is less than or equal to a preset training round threshold, training the model based on the first training data set; If the current training round is greater than the preset training round threshold, the model is trained based on the second training data set.
6. The method according to claim 4, characterized in that The lightweight multi-task convolutional neural network model includes: a shared feature extraction layer, a classification branch layer and an integration branch layer, wherein the classification branch layer is used to output the probability that the peak area belongs to the noise category, the valid peak category and the category that needs to be reviewed; the integration branch layer is used to output the probability that each detection point belongs to the peak area and the overlapping area.
7. The method according to any one of claims 1 to 6, characterized in that The step of inputting the test data of the test sample into the LC-MS peak detection model to obtain the test result specifically includes: Inputting the data to be tested into the peak region classification model, detecting the peak regions contained in the data to be tested and the characteristic parameters of the peak regions, and determining the probability of the peak regions belonging to the noise category, the valid peak category, and the category requiring review based on the characteristic parameters of the peak regions; Based on multiple probabilities, the category of each peak area is determined; If the peak region belongs to the valid peak category, the model is defined by the peak boundary, and the probability of each scan point in the peak region belonging to the peak region and the overlapping region is output; If the peak region belongs to the category requiring review, the characteristic parameters of the peak region are sent to the terminal of the R&D personnel.
8. An LC-MS peak detection device, characterized in that include: The acquisition module is used to obtain the raw data set of environmental samples collected by the LC-MS instrument; An adjustment module is used to extract peak detection parameters from the original data set and adjust the peak detection parameters; a determination module is used to determine multiple peak regions and characteristic parameters of each peak region based on the adjusted peak detection parameters; A first generation module is configured to generate, based on the characteristic parameters of each peak region, preset instrument noise, and preset chromatographic conditions, a synthetic peak region having a distribution consistent with the characteristic parameters of the real peak region through a conditional generative adversarial network, and obtain synthetic characteristic parameters of each synthetic peak region; A construction module is configured to construct an LC-MS peak detection model based on the characteristic parameters and the synthesized characteristic parameters using a lightweight multi-task convolutional neural network, wherein the LC-MS peak detection model includes a peak region classification model for outputting the probability that each peak region belongs to a noise category, a valid peak category, and a category requiring review, and a peak boundary definition model for outputting the probability that each scanning point in the peak region belongs to a single peak region and an overlapping region; The second generating module is used to input the test data of the test sample into the LC-MS peak detection model to obtain the detection result.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the LC-MS peak detection method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the LC-MS peak detection method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Comprehensive two-dimensional mass spectrum data non-targeted screening method and device, medium and computer equipment
CN121191637A
A chromatographic data analysis method and system based on intelligent sensors
CN122545734A