Chromatographic signal data processing method, device, equipment, medium and product
The baseline is fitted through the Savitzky-Golay filter denoising and cubic spline interpolation, and the peak identification is combined with the symmetric zero area method, which solves the problems of noise interference and baseline drift in chromatographic signal processing, and improves the accuracy of chromatographic peak recognition.
Patent Information
- Application Number
- CN202510413512.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-11
AI Technical Summary
Noise interference and baseline drift in chromatographic signal processing lead to difficulty in identifying signal peak characteristics, affecting the accuracy of the analysis results.
The Savitzky-Golay filter was used to filter out high-frequency noise, the baseline was fitted by cubic spline interpolation, and peak recognition was performed by symmetric zero area method.
Effectively remove noise and baseline drift, improve the accuracy of chromatographic peak recognition and ensure the reliability of analysis results.
Smart Images

Figure CN120294227A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to a method, apparatus, device, medium, and product for processing chromatographic signal data. Background Art
[0002] Chromatography is an important analytical method widely used in fields such as chemistry, biology, pharmacy, and environmental science. By detecting the retention time and signal intensity of each component in a sample, qualitative and quantitative analysis of complex mixtures can be achieved. However, during the chromatographic signal processing, two major challenges are often faced: one is that the presence of noise interference makes it difficult to identify the characteristics of signal peaks; the other is that the baseline drift phenomenon will seriously affect the accuracy of peak area and peak height, and thus affect the reliability of the analysis results.
[0003] Traditional chromatographic signal processing methods mainly rely on techniques such as polynomial fitting or moving average for baseline correction, and use simple threshold methods or derivative methods for peak identification. However, when facing complex baselines, low signal-to-noise ratios, or overlapping peaks, these methods are prone to inaccurate identification problems. Summary of the Invention
[0004] The purpose of this application is to provide a method, apparatus, device, medium, and product for processing chromatographic signal data, which can improve the accuracy of chromatographic signal identification.
[0005] To achieve the above purpose, this application provides the following solutions:
[0006] In the first aspect, this application provides a method for processing chromatographic signal data, where the method for processing chromatographic signal data includes:
[0007] Obtain the original chromatographic signal;
[0008] Use a Savitzky-Golay filter to filter out high-frequency noise in the original chromatographic signal to obtain a denoised signal;
[0009] Use cubic spline interpolation to fit the baseline of the denoised signal;
[0010] Deduct the baseline from the denoised signal to obtain a corrected signal;
[0011] Perform peak identification on the corrected signal through the symmetric zero area method to generate an identification result.
[0012] In the second aspect, this application provides a device for processing chromatographic signal data, characterized in that the device for processing chromatographic signal data includes:
[0013] An original chromatographic signal acquisition module, configured to obtain an original chromatographic signal;
[0014] A denoising signal acquisition module, configured to use a Savitzky-Golay filter to filter out high-frequency noise in the original chromatographic signal to obtain a denoised signal;
[0015] A baseline acquisition module, configured to use cubic spline interpolation to fit the baseline of the denoised signal;
[0016] A corrected signal acquisition module, configured to subtract the baseline from the denoised signal to obtain a corrected signal;
[0017] An identification module, configured to perform peak identification on the corrected signal by the symmetric zero area method to generate an identification result.
[0018] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the steps of the chromatographic signal data processing method described in any one of the above.
[0019] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the chromatographic signal data processing method described in any one of the above are implemented.
[0020] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the chromatographic signal data processing method described in any one of the above are implemented.
[0021] According to the specific embodiments provided by the present application, the following technical effects are disclosed in the present application:
[0022] In this solution, a Savitzky-Golay filter is used for denoising processing. This filter can, while removing noise, maximize the preservation of the original waveform of the signal, and can achieve high-fidelity denoising even in the case of low signal-to-noise ratio, effectively avoiding waveform distortion caused by improper denoising and laying a good foundation for subsequent analysis. Then, cubic spline interpolation technology is used to correct the non-linear baseline drift. In this way, even when the baseline is complex, the baseline can be accurately fitted, effectively reducing the baseline correction error, and enabling subsequent analysis to be carried out on a more accurate baseline. Finally, combined with the peak feature recognition technology based on symmetric zero area, the characteristic points corresponding to the peaks in the chromatographic data are identified, and even in the case of overlapping peaks, the characteristic points corresponding to the peaks can be accurately located, solving the problem of inaccurate identification of characteristic points in traditional methods. Since the noise and baseline drift phenomena in the chromatographic signal can be effectively eliminated, the true chromatographic peak information can be accurately retained, and the characteristic points in the chromatographic data can be accurately identified, thereby achieving the purpose of improving the accuracy of chromatographic peak identification. Description of the Drawings
[0023] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0024] Figure 1 is a flowchart of a method for processing chromatographic signal data shown according to an exemplary embodiment;
[0025] Figure 2 is a chromatogram of the original signal shown according to an exemplary embodiment;
[0026] Figure 3 is a denoised signal shown according to an exemplary embodiment;
[0027] Figure 4 is a corrected signal shown according to an exemplary embodiment;
[0028] Figure 5 are the peak start point, peak vertex and peak end point shown according to an exemplary embodiment;
[0029] Figure 6 is a schematic diagram of the functional modules of a device for processing chromatographic signal data shown according to an exemplary embodiment;
[0030] Figure 7 is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0031] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0032] To make the above objects, features and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0033] Figure 1 is a flowchart of a method for processing chromatographic signal data shown according to an exemplary embodiment. As Figure 1 shown, the method includes the following steps S101 - S105:
[0034] In step S101, an original chromatographic signal is acquired.
[0035] A chromatographic signal refers to the measurable electrical signal or other types of signals generated by a detector for the separated components during chromatographic analysis.
[0036] It is usually presented in the form of a chromatogram, such as Figure 2 The figure shows the chromatogram of the original signal. The ordinate of the chromatogram is the response signal of the detector, that is, the signal intensity, and the abscissa is time.
[0037] If a large amount of historical chromatographic data is stored in a dedicated database, the required original chromatographic signal can be obtained through the corresponding database management software and programming language interfaces. For example, use SQL statements to query chromatographic signal data that meets specific conditions (such as sample name, experiment date, etc.) from a relational database.
[0038] If the original chromatographic signal data is stored in common formats such as ordinary text files, CSV files, etc., it can be directly read through file reading functions.
[0039] In a laboratory environment, professional chromatographic analysis instruments are used, such as gas chromatographs, liquid chromatographs, etc. When a sample is injected into the instrument, the detector inside the instrument will record the response signal of the sample components changing with time in real time.
[0040] In step S102, a Savitzky-Golay filter is used to filter out the high-frequency noise in the original chromatographic signal to obtain a denoised signal.
[0041] The Savitzky-Golay filter achieves smoothing by performing polynomial fitting on local data points of the signal. It can effectively remove high-frequency noise while better retaining the characteristics of the signal (such as peaks, valleys, etc.). For the Figure 2 original chromatographic signal graph in, after passing through the Savitzky-Golay filter, a denoised signal as shown in Figure 3 is obtained.
[0042] In one embodiment, using a Savitzky-Golay filter in step S102 to filter out the high-frequency noise in the original chromatographic signal to obtain a denoised signal includes the following sub-steps S1021 - S1024:
[0043] S1021. Obtain the frequency characteristics of the original chromatographic signal.
[0044] The frequency characteristics can reflect the distribution of different frequency components in the signal. To obtain the frequency characteristics of the original chromatographic signal, the Fourier transform is usually used to convert the signal from the time domain to the frequency domain, so that the energy distribution of the signal at different frequencies can be visually seen.
[0045] The frequency characteristics of the original chromatographic signal can be obtained through discrete Fourier transform (DFT) or fast Fourier transform (FFT).
[0046] S1022. Obtain the noise level in the original chromatographic signal.
[0047] The estimation of the noise level helps to determine the parameters of the filter. A common method is to estimate the standard deviation of the noise by analyzing the statistical characteristics of the signal.
[0048] In the original chromatographic signal, some relatively stable regions can usually be found. The fluctuations in these regions are mainly caused by noise. The data in these stable segments can be selected, and its standard deviation can be calculated as the estimated value of the noise level.
[0049] S1023. Obtain the window width and polynomial order of the Savitzky-Golay filter according to the frequency characteristics and noise level to obtain the target Savitzky-Golay filter.
[0050] The performance of the Savitzky-Golay filter mainly depends on these two parameters: the window width and the polynomial order. The window width determines the number of data points considered when the filter performs polynomial fitting, and the polynomial order determines the complexity of the fitting polynomial.
[0051] According to experience, the window width is usually selected as an odd number and is generally between 5 and 25; the polynomial order is usually selected between 2 and 4. Tentative selection can be made within this range according to the frequency characteristics and noise level, and then the optimal parameters can be determined by observing the filtering effect.
[0052] Some optimization algorithms (such as grid search, genetic algorithm, etc.) can be used to automatically search for the optimal window width and polynomial order.
[0053] S1024. Use the target Savitzky-Golay filter to filter out the high-frequency noise in the original chromatographic signal.
[0054] After determining the window width and polynomial order of the target Savitzky-Golay filter, the original chromatographic signal can be filtered using this filter to remove the high-frequency noise in it.
[0055] Through the above steps, the task of filtering out the high-frequency noise in the original chromatographic signal using the Savitzky-Golay filter can be completed.
[0056] In step S103, cubic spline interpolation fitting is used to fit the baseline of the denoised signal.
[0057] In chromatographic signal processing, the baseline refers to the background level of the signal when no target substance elutes. Due to the influence of various factors, the baseline of the original chromatographic signal may exhibit drift, fluctuations, etc., which can interfere with subsequent operations such as peak identification and quantitative analysis. Therefore, it is necessary to perform baseline fitting on the filtered signal to accurately determine the true baseline of the signal. In this disclosure, the method of cubic spline interpolation is used to fit the baseline of the denoised signal. This method can better approximate the true baseline while ensuring the smoothness of the interpolation curve.
[0058] In one embodiment, step S103 includes the following sub-steps S1031 - S1032:
[0059] S1031. Starting from the starting point of the denoised signal, sequentially traverse the entire denoised signal with the first sliding window to obtain the baseline points in each first sliding window.
[0060] The sliding window is a commonly used signal processing method. It slides a window of a fixed length over the signal and analyzes and processes the data within the window. In this step, the first sliding window is used to sequentially traverse the entire denoised signal with the aim of finding the points within each window that can represent the baseline level of the local region, i.e., the baseline points.
[0061] In one embodiment, in step S1031, starting from the starting point of the denoised signal, sequentially traverse the entire denoised signal with the first sliding window to obtain the baseline points in each first sliding window, including performing the following sub-steps S10311 - S10313 on the chromatographic signal intensity within each first sliding window:
[0062] S10311. Obtain the minimum value of the chromatographic signal intensity within the current first sliding window.
[0063] To determine the minimum value of the chromatographic signal intensity within the current first sliding window is to find the lowest level of the signal within this local region, and this lowest level can largely represent the baseline situation of this region.
[0064] In one embodiment, obtaining the minimum value of the chromatographic signal intensity within the current first preset sliding window in step S10311 includes the following sub-steps S103111 - S103112:
[0065] S103111. Perform the determination of the minimum value condition on each chromatographic signal intensity within the current first sliding window.
[0066] In actual operation, for a given current first sliding window, it contains a certain number of chromatographic signal intensity data points, and these data points need to be processed to obtain the minimum value.
[0067] To more accurately determine the minimum value, instead of simply comparing magnitudes, a minimum value condition is set, and this condition is judged for each data point of the chromatographic signal intensity within the current first sliding window. Doing so can avoid misjudging some special cases (such as abnormally small values under noise interference) as the minimum value, making the determined minimum value better represent the true baseline level.
[0068] The minimum value condition is:
[0069] Y(t i ) < Y(t i-1 ) ; Y(t i ) < Y(t i+1 ) ;
[0070] Y′(t i ) ≈ 0 ;
[0071] Y(t i ) < ∈ ;
[0072] Where Y(t i ) represents the chromatographic signal intensity corresponding to time point t i ), Y(t i-1 ) represents the chromatographic signal intensity corresponding to time point t i-1 ), Y(t i+1 ) represents the chromatographic signal intensity corresponding to time point t i+1 ), ∈ represents the upper limit of the signal amplitude of the baseline point, Y′(t i ) represents the derivative of the chromatographic signal intensity corresponding to time point t i ), and i is an integer greater than 0.
[0073] Starting from the first data point of the window, the judgment of the minimum value condition is sequentially performed on each data point.
[0074] S103112. Determine that the chromatographic signal intensity satisfying the minimum value condition is the minimum value of the chromatographic signal intensity within the current first sliding window.
[0075] S10312. Obtain the time interval between the time point corresponding to the minimum value and the adjacent time point.
[0076] After determining the minimum value of the chromatographic signal intensity within the current first sliding window, it is necessary to further analyze the characteristics of this minimum value in the time series. By obtaining the time interval between the time point corresponding to this minimum value and the adjacent time points, the distribution of this minimum value in the time dimension can be understood, which helps to judge whether this minimum value is caused by noise or other abnormal factors, or truly represents the characteristics of the baseline.
[0077] After finding the minimum value of the chromatographic signal intensity within the first sliding window, it is necessary to determine the specific time point corresponding to this minimum value. Since the chromatographic signal changes with time, each chromatographic signal intensity value corresponds to a time point. Find this time point t i After that, look at its adjacent time points (if any). One is the time point t i-1 before it, and one is the time point t i+1 after it. Then calculate the time interval t i -t i-1 between the time point corresponding to this minimum value and the previous time point, and the time interval t i+1 -t i with the subsequent time point.
[0078] S10313. If the time interval meets the preset time range, determine the minimum value as the baseline point in the first sliding window.
[0079] The present disclosure will first stipulate a suitable time range [Δt min , Δt max , which is set according to the characteristics of chromatographic analysis and our requirements. After calculating the time intervals between the time point corresponding to the minimum value and the adjacent time points, check whether these time intervals are within the preset time range. If these time intervals are all within the preset time range, that is, Δt min ≤t i+1 -t i ≤Δt max , and Δt min ≤t i -t i-1 ≤Δt max , it means that the situation where this minimum value is located meets the judgment criteria for the baseline point, and it can be determined that the point corresponding to this minimum value is the baseline point in the first sliding window. If it does not meet the preset time range, mark this minimum value as an abnormal point and eliminate it, so as to ensure that the spacing between the baseline points is uniform and the distribution of the points is reasonable.
[0080] S1032. Perform cubic spline interpolation on each baseline point to obtain the baseline of the denoised signal.
[0081] In signal processing scenarios such as chromatographic analysis, the filtered signal may still have irregular baseline fluctuations. Through the previous steps, a series of baseline points representing the baseline characteristics have been identified. The purpose of the step of "performing cubic spline interpolation on each baseline point to obtain the baseline of the denoised signal" is to construct a continuous, smooth curve that can accurately reflect the overall baseline trend of the denoised signal based on these discrete baseline points, so as to lay a foundation for more accurate analysis of the effective components (such as chromatographic peaks) in the signal in the future.
[0082] In the previous steps, multiple baseline points have been identified. Each baseline point contains two key pieces of information: one is the time corresponding to this point, and the other is the intensity of the chromatographic signal at that time point. These baseline points can be imagined as landmarks on a map, where time is the longitude of the landmark and signal intensity is the latitude. Arrange these baseline points in order for subsequent processing.
[0083] Cubic spline interpolation is a mathematical method used to construct a smooth curve between known discrete data points. Its core idea is to construct a cubic polynomial function between every two adjacent baseline points and require these polynomial functions to satisfy certain continuity conditions at the connection points to ensure that the entire interpolated curve is smooth.
[0084] Perform cubic spline interpolation on the detected baseline points (t i , Y i ) to construct the baseline function H(t), that is:
[0085] H(t) = a i (t - t i ) 3 + b i (t - t i ) 2 + c i (t - t i ) + d i , t i ≤ t ≤ t i+1 ;
[0086] where t i is the time series, and a i , b i , c i , d i are interpolation coefficients that can be calculated through boundary conditions and baseline points.
[0087] In step S104, subtract the baseline from the denoised signal to obtain the corrected signal.
[0088] Subtract the fitted baseline H(t) from the denoised signal to obtain the corrected signal b(t), as Figure 4 shown, which is the corrected signal corresponding to the denoised signal in Figure 3 .
[0089] The original chromatographic signal often contains various noises. To analyze the signal characteristics more clearly, tools such as the Savitzky-Golay filter are first used to process the original chromatographic signal to remove the high-frequency noises. The relatively clean and smooth signal obtained is the denoised signal. It still contains chromatographic peak information and baseline fluctuation information. The denoised signal can be imagined as a curve containing various undulations (both the chromatographic peaks we are concerned about and the up-and-down fluctuations of the baseline).
[0090] By further analyzing the denoised signal and using methods such as cubic spline interpolation, a curve representing the baseline trend in the signal, that is, the baseline, is fitted according to specific rules (such as first determining the baseline points and then interpolating the baseline points). This baseline reflects the background fluctuation situation in the signal except for the effective components such as chromatographic peaks.
[0091] Since both the denoised signal and the baseline are a series of numerical points that change with time, for each time point in the denoised signal, find the intensity value of the denoised signal corresponding to that time point, and at the same time find the intensity value corresponding to the same time point on the baseline. Then subtract the intensity value of the baseline at the same time point from the intensity value of the denoised signal at that time point.
[0092] For example, at a certain moment t1, the intensity of the denoised signal is S1, and the intensity of the baseline at this moment is B1. Then the intensity of the corrected signal at this moment is S1 - B1. Perform such subtraction operations for each time point in the denoised signal, thereby obtaining a series of new intensity values, and these new intensity values constitute the corrected signal.
[0093] After the above operation of subtracting the baseline, the influence of the baseline part in the corrected signal is greatly eliminated. At this time, the corrected signal mainly retains the effective components such as chromatographic peaks and possibly a small amount of low-frequency noises. In this way, when analyzing the signal subsequently, such as identifying peaks in the corrected signal by the symmetric zero-area method, the position, height, area and other characteristics of the chromatographic peaks can be detected and analyzed more accurately, avoiding the interference of baseline fluctuations on the peak identification results, improving the accuracy and reliability of the analysis, and making the analysis results of sample components and the like more precise.
[0094] In step S105, peak identification is performed on the corrected signal by the symmetric zero-area method to generate an identification result.
[0095] In the chromatographic signal processing flow, after a series of operations such as obtaining the original chromatographic signal, filtering out high-frequency noises (obtaining the denoised signal), fitting the baseline and subtracting the baseline (obtaining the corrected signal), the next task is to perform peak identification on the corrected signal to determine the relevant characteristics of the chromatographic peaks, such as the peak apex, peak start point and peak end point, as Figure 5 shown, isFigure 4 The corresponding peak start point, peak vertex point, and peak end point in the corrected signal.
[0096] In one embodiment, step S105 performs peak identification on the corrected signal by the symmetric zero area method, and the generated identification result includes:
[0097] Starting from the starting point of the corrected signal, sequentially traverse the entire corrected signal with a second sliding window, and perform the following steps A1 - A4 for each data point corresponding to the corrected signal in each second sliding window:
[0098] The symmetric zero area method is used to process the corrected signal to accurately identify each key position of the chromatographic peak in the corrected signal, and finally generate an identification result containing this information. This process will provide a basis for subsequent further analysis of the chromatographic peak (such as calculating parameters such as peak area, peak height, and peak width, thereby inferring sample components and contents, etc.).
[0099] Starting from the starting point of the corrected signal, use a second sliding window of a specific size and slide and traverse sequentially over the entire corrected signal. For each data point of the corrected signal within the second sliding window, the following four steps (A1 - A4) need to be performed to determine whether this point is the peak vertex point, peak start point, or peak end point of the chromatographic peak.
[0100] A1. Obtain the area difference between the two sides of each data point corresponding to the corrected signal in the current second sliding window.
[0101] Within the current second sliding window, for each data point of the corrected signal, it is necessary to calculate the areas on both sides of this point. Here, the "area" generally refers to the area enclosed by the corrected signal curve within a certain range on the left and right sides (the specific range is related to the sliding window) of this data point and a certain reference line (such as the zero baseline). By calculating the left - hand area and the right - hand area, and then finding the difference between the two, the area difference corresponding to each data point is obtained. This area difference reflects the imbalance of the signal distribution on both sides of this data point and is an important basis for subsequent judgment.
[0102] Specifically, set the initial width L of the second sliding window and calculate the left - right area difference ΔI ω (j), that is:
[0103]
[0104] Where represents the weight coefficient, b(t) is the corrected signal, and j represents the position of the current data point in the second sliding window.
[0105] A2. If the area difference is zero and the data point corresponding to the zero area difference is a local maximum, then determine that the data point corresponding to the zero area difference is the peak vertex.
[0106] When the area difference of a certain data point is zero, it indicates that the signal distributions on both sides of this point reach a balanced state in terms of area. At the same time, if this data point is the maximum value of the signal intensity within a certain range around it (usually referring to the adjacent data points within the sliding window), that is, a local maximum, then it can be considered that this data point is the position of the peak vertex of the chromatographic peak. Because at the top of the chromatographic peak, the signal intensity reaches the maximum, and the signal intensity gradually decreases from the peak vertex to both sides. In an ideal situation, the areas on both sides of the peak vertex are equal (i.e., the area difference is zero).
[0107] Specifically, for peak vertex determination: The points that meet the following two conditions are the peak vertices.
[0108] Condition 1: The area difference ΔI ω (j) = 0 (or within the allowable range);
[0109] Condition 2: The corrected signal b(t) is a local maximum at this point, that is: b'(t) = 0, b”(t) < 0.
[0110] A3. If the area difference is zero and the area difference of the previous data point adjacent to the zero area difference is negative, determine the peak start point of the data point corresponding to the zero area difference.
[0111] When the area difference of a data point is zero and the area difference of the data point immediately before this one is negative, it indicates that before this data point, the area enclosed by the signal curve and the baseline is smaller on the left side than on the right side (the area difference is negative), and at this data point, the area reaches balance (the area difference is zero). This situation conforms to the characteristic of the start of the chromatographic peak rising, that is, starting from this point, the signal intensity begins to rise to form a chromatographic peak. Therefore, it can be determined that this data point with a zero area difference is the start point of the chromatographic peak.
[0112] That is, ΔI ω (j) changes from negative to zero, indicating that the signal changes from asymmetric to symmetric.
[0113] A4. If the area difference is zero and the area difference of the next data point adjacent to the zero area difference is positive, determine the peak end point of the data point corresponding to the zero area difference.
[0114] When the area difference of a data point is zero and the area difference corresponding to the data point immediately following this data point is positive, it means that after this data point, the area enclosed by the signal curve and the reference line is larger on the left side than on the right side (the area difference is positive), while the area is balanced at this data point (the area difference is zero). This is consistent with the characteristic that the chromatographic peak starts to decline after reaching the top, that is, from this point, the signal intensity starts to decline. Therefore, it can be determined that this data point with an area difference of zero is the end point of the chromatographic peak.
[0115] That is, ΔI ω (j) the point where it changes from zero to a positive value, indicating the end symmetry of the signal.
[0116] By sequentially performing the above steps A1 - A4 on each data point in the corrected signal, after traversing the entire corrected signal, the key position information such as the peak apex, peak start point, and peak end point of all chromatographic peaks in the corrected signal can be determined. Finally, a peak recognition result containing this information is generated, providing accurate data support for subsequent chromatographic analysis.
[0117] In traditional chromatographic data analysis methods, there are many limitations. On the one hand, a single algorithm often has difficulty achieving a good balance between processing speed and accuracy, resulting in the inability to meet the requirements of complex scenarios in practical applications. On the other hand, traditional methods are prone to problems such as peak shape distortion, baseline correction errors, and inaccurate identification of characteristic points. In response to these pain points, the present disclosure proposes a comprehensively optimized solution.
[0118] First, this solution uses a Savitzky - Golay filter for denoising. This filter can, while removing noise, maximize the preservation of the original waveform of the signal, achieving high - fidelity denoising and effectively avoiding waveform distortion caused by improper denoising, laying a good foundation for subsequent analysis. Then, cubic spline interpolation technology is used to correct the non - linear baseline drift. In this way, the baseline can be accurately fitted, effectively reducing baseline correction errors, enabling subsequent analysis to be carried out on a more accurate baseline. Finally, combined with the peak feature recognition technology based on symmetric zero area, the characteristic points corresponding to the peaks in the chromatographic data are identified. This technology can accurately locate the characteristic points corresponding to the peaks, solving the problem of inaccurate identification of characteristic points in traditional methods, and this technology can also well identify the characteristic points corresponding to the peaks even in the face of overlapping peaks.
[0119] Through the organic combination of the above - mentioned series of technologies, this solution ensures the processing speed while significantly improving the analysis accuracy of complex chromatographic data. Whether facing the interference of high noise, complex baseline conditions, or the challenge of overlapping peaks, it can provide effective coping strategies, bringing new ideas and methods to the field of chromatographic data analysis.
[0120] The present application also provides an application scenario, which applies the above-mentioned method for processing chromatographic signal data. Specifically: The method for processing chromatographic signal data provided in this embodiment can be applied to the separation and identification of compounds. In chemical experiments such as organic synthesis and natural product research, various compounds in a complex mixture can be separated through chromatographic analysis, and the compounds can be identified by combining data such as retention time and peak shape with techniques such as mass spectrometry to determine their structures and compositions. It can also be applied to the detection of pesticide and veterinary drug residues: Detecting pesticide and veterinary drug residues in food is an important part of ensuring food safety. Chromatographic analysis can accurately detect trace amounts of pesticide and veterinary drug residues in food, ensure that the food meets safety standards, and protect the health of consumers.
[0121] Based on the same inventive concept, the embodiment of the present application also provides a device for processing chromatographic signal data for implementing the above-mentioned method for processing chromatographic signal data. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the device for processing chromatographic signal data provided below can refer to the limitations on the method for processing chromatographic signal data in the above text, and will not be elaborated here.
[0122] In an exemplary embodiment, as Figure 6 shown, a device for processing chromatographic signal data is provided, including:
[0123] An original chromatographic signal acquisition module 11, configured to acquire an original chromatographic signal;
[0124] A denoised signal acquisition module 12, configured to filter out high-frequency noise in the original chromatographic signal by using a Savitzky-Golay filter to obtain a denoised signal;
[0125] A baseline acquisition module 13, configured to fit the baseline of the denoised signal by using cubic spline interpolation;
[0126] A corrected signal acquisition module 14, configured to subtract the baseline from the denoised signal to obtain a corrected signal;
[0127] An identification module 15, configured to perform peak identification on the corrected signal by using the symmetric zero-area method to generate an identification result.
[0128] In one embodiment, the denoised signal acquisition module 12 is specifically configured to:
[0129] Acquire the frequency characteristics in the original chromatographic signal;
[0130] Acquire the noise level in the original chromatographic signal;
[0131] Obtain the window width and polynomial order of the Savitzky-Golay filter according to the frequency characteristic and the noise level to obtain a target Savitzky-Golay filter;
[0132] Use the target Savitzky-Golay filter to filter out high-frequency noise in the original chromatographic signal.
[0133] In one embodiment, the baseline acquisition module 13 is specifically configured to:
[0134] Starting from the starting point of the denoised signal, traverse the entire denoised signal successively with a first sliding window, and obtain the baseline points in each first sliding window;
[0135] Perform cubic spline interpolation on each of the baseline points to obtain the baseline of the denoised signal.
[0136] In one embodiment, the baseline acquisition module 13 is specifically configured to:
[0137] Perform the following steps on the chromatographic signal intensity within each first sliding window:
[0138] Obtain the minimum value of the chromatographic signal intensity within the current first sliding window;
[0139] Obtain the time interval between the time point corresponding to the minimum value and the adjacent time point;
[0140] If the time interval satisfies a preset time range, determine that the minimum value is the baseline point in the first sliding window.
[0141] In one embodiment, the baseline acquisition module 13 is specifically configured to:
[0142] Perform a determination of the minimum value condition on each chromatographic signal intensity within the current first sliding window;
[0143] Determine that the chromatographic signal intensity that satisfies the minimum value condition is the minimum value of the chromatographic signal intensity within the current first sliding window;
[0144] The minimum value condition is:
[0145] Y(t i ) < Y(t i-1 ) ; Y(t i ) < Y(t i+1 ) ;
[0146] Y'(t i ) ≈ 0 ;
[0147] Y(t i ) < ∈ ;
[0148] Among them, Y(t i ) represents the chromatographic signal intensity corresponding to the time point t i ), Y(t i-1 ) represents the chromatographic signal intensity corresponding to the time point t i-1 ), Y(t i+1 ) represents the chromatographic signal intensity corresponding to the time point t i+1 ), where ∈ represents the upper limit of the baseline point signal amplitude, and Y′(t i ) represents the derivative of the chromatographic signal intensity corresponding to the time point t i .
[0149] In one embodiment, the recognition module 15 is specifically configured to:
[0150] Starting from the starting point of the corrected signal, traverse the entire corrected signal successively with a second sliding window, and perform the following steps for each data point corresponding to the corrected signal in each second sliding window:
[0151] Obtain the area difference between both sides of each data point corresponding to the corrected signal in the current second sliding window;
[0152] If the area difference is zero and the data point corresponding to the zero area difference is a local maximum, determine the data point corresponding to the zero area difference as the peak vertex;
[0153] If the area difference is zero and the previous area difference adjacent to the zero area difference is negative, determine the data point corresponding to the zero area difference as the peak start point;
[0154] If the area difference is zero and the next area difference adjacent to the zero area difference is positive, determine the data point corresponding to the zero area difference as the peak end point.
[0155] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 7As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for processing chromatographic signal data.
[0156] Those skilled in the art can understand that Figure 7 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0157] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0158] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0159] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0160] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0161] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0162] The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0163] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity in description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0164] Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for processing chromatographic signal data, characterized in that, The processing method of the chromatographic signal data includes: Obtain the original chromatographic signal; Use a Savitzky-Golay filter to filter out the high-frequency noise in the original chromatographic signal to obtain a denoised signal; Use cubic spline interpolation to fit the baseline of the denoised signal; Subtract the baseline from the denoised signal to obtain a corrected signal; Perform peak identification on the corrected signal by the symmetric zero-area method to generate an identification result.
2. The processing method of chromatographic signal data according to claim 1, characterized in that, Input the original chromatographic signal into a Savitzky-Golay filter to filter out the high-frequency noise in the original chromatographic signal, including: Obtain the frequency characteristics in the original chromatographic signal; Obtain the noise level in the original chromatographic signal; Obtain the window width and polynomial order of the Savitzky-Golay filter according to the frequency characteristics and the noise level to obtain a target Savitzky-Golay filter; Use the target Savitzky-Golay filter to filter out the high-frequency noise in the original chromatographic signal.
3. The processing method of chromatographic signal data according to claim 1, characterized in that The using cubic spline interpolation to fit the baseline of the denoised signal includes: Starting from the starting point of the denoised signal, sequentially traverse the entire denoised signal with a first sliding window to obtain the baseline points in each first sliding window; Perform cubic spline interpolation on each baseline point to obtain the baseline of the denoised signal.
4. The processing method of chromatographic signal data according to claim 3, characterized in that, The obtaining the baseline points in each first sliding window includes: Perform the following steps on the chromatographic signal intensity in each first sliding window: Obtain the minimum value of the chromatographic signal intensity in the current first sliding window; Obtain the time interval between the time point corresponding to the minimum value and the adjacent time point; If the time interval satisfies a preset time range, determine the minimum value as the baseline point in the first sliding window.
5. The method for processing chromatographic signal data according to claim 4, wherein, The obtaining the minimum value of the chromatographic signal intensity in the current first preset sliding window includes: Perform a determination of the minimum value condition on each chromatographic signal intensity in the current first sliding window; Determine the chromatographic signal intensity that satisfies the minimum value condition as the minimum value of the chromatographic signal intensity in the current first sliding window; The minimum value condition is: Y(t i )<Y(t i-1 );Y(t i )<Y(t i+1 ); Y'(t i )≈0; Y(t i )<∈; Among them, Y(t i ) represents the chromatographic signal intensity corresponding to the time point t i ; Y(t i-1 ) represents the chromatographic signal intensity corresponding to the time point t i-1 ; Y(t i+1 ) represents the chromatographic signal intensity corresponding to the time point t i+1 . The ∈ represents the upper limit of the baseline point signal amplitude, and Y'(t i ) represents the derivative of the chromatographic signal intensity corresponding to the time point t i . i is an integer greater than 0.
6. The processing method of chromatographic signal data according to claim 1, characterized in that, The performing peak identification on the corrected signal by the symmetric zero-area method to generate an identification result includes: Starting from the starting point of the corrected signal, sequentially traverse the entire corrected signal with a second sliding window, and perform the following steps on each data point corresponding to the corrected signal in each second sliding window: Obtain the area difference between the two sides of each data point corresponding to the corrected signal in the current second sliding window; If the area difference is zero and the data point corresponding to the zero area difference is a local maximum, determine the data point corresponding to the zero area difference as the peak vertex; If the area difference is zero and the previous area difference adjacent to the zero area difference is negative, determine the data point corresponding to the zero area difference as the peak starting point; If the area difference is zero and the next area difference adjacent to the area difference being zero is positive, determine the peak end point of the data point corresponding to the area difference being zero.
7. A processing device for chromatographic signal data, characterized in that The processing device for the chromatographic signal data includes: An original chromatographic signal acquisition module, configured to acquire an original chromatographic signal; A denoised signal acquisition module, configured to filter out high-frequency noise in the original chromatographic signal by using a Savitzky-Golay filter to obtain a denoised signal; A baseline acquisition module, configured to fit the baseline of the denoised signal by using cubic spline interpolation; A corrected signal acquisition module, configured to subtract the baseline from the denoised signal to obtain a corrected signal; An identification module, configured to perform peak identification on the corrected signal by using the symmetric zero area method to generate an identification result.
8. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method for processing chromatographic signal data according to any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for processing chromatographic signal data according to any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for processing chromatographic signal data according to any one of claims 1-6.
Citation Information
Cited By
Coal combustion spectrum detection method and system
CN120690319A
A method and system for detecting coal combustion spectrometry
CN120690319B
Mass spectrum gas source analysis method and system based on K-means clustering algorithm
CN120929865A