Chromatographic baseline correction method, device, equipment, medium and product
By constructing polynomial models and iterative optimization techniques, the problem of excessive or insufficient correction in traditional baseline correction methods is solved, and the accuracy and accuracy of the chromatographic signal are significantly improved.
Patent Information
- Application Number
- CN202510349032.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-05-30
AI Technical Summary
Traditional baseline correction methods may experience problems such as over-correction or inadequate correction, resulting in the accuracy and accuracy of the chromatographic signal.
By acquiring the original chromatographic data and preprocessing, a polynomial model is constructed to react to the chromatographic signal change trend, and the optimal order of the model is determined by K-fold cross-validation. Then, a second polynomial model is constructed based on the optimal order for fitting, obtaining the initial fit baseline, and the final fit baseline is obtained through iterative optimization, thereby performing baseline correction.
It effectively avoids underfitting or overfitting problems caused by improper order selection, significantly improves the accuracy and applicability of baseline correction, suppresses the interference of strong peaks on baseline fitting, and improves the accuracy of chromatographic signals.
Smart Images

Figure CN120064541A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of baseline correction, and particularly to a chromatographic baseline correction method, device, equipment, medium and product. Background Art
[0002] Chromatographic analysis is a separation and analysis technique widely used in many fields such as chemistry, medicine and environment. Its core lies in the accurate interpretation of chromatographic signals. Chromatographic signals consist of characteristic peaks of target components and baselines. However, due to experimental conditions, instrument drift and sample complexity, there are often noise and baseline drift problems in chromatographic signals. These interferences not only reduce the analysis accuracy, but also may affect qualitative and quantitative results. Especially when detecting low-concentration target components, the correction of the signal baseline is particularly important.
[0003] Traditional baseline correction methods include manual adjustment, moving average and wavelet transform, etc. However, traditional baseline correction methods may have problems of overcorrection or undercorrection. Summary of the Invention
[0004] The purpose of this application is to provide a chromatographic baseline correction method, device, equipment, medium and product, which can improve the accuracy of chromatographic baseline correction.
[0005] To achieve the above purpose, this application provides the following solutions:
[0006] In the first aspect, this application provides a chromatographic baseline correction method, including:
[0007] Obtain original chromatographic data, and preprocess the original chromatographic data to obtain chromatographic data, where the chromatographic data includes a time point t and a chromatographic signal in one-to-one mapping, and the chromatographic signal is used to represent the signal intensity at the time point t, and all the chromatographic data constitutes a chromatographic data set;
[0008] Construct a first polynomial model for reflecting the change trend of chromatographic signals in the chromatographic data set, and determine the optimal order of the first polynomial model through the K-fold cross-validation method;
[0009] Construct a second polynomial model according to the optimal order of the first polynomial model, and fit the chromatographic data through the second polynomial model to obtain an initial fitted baseline;
[0010] Iteratively optimize the initial fitted baseline to obtain a final fitted baseline, and subtract the corresponding final estimated value of the final fitted baseline from the chromatographic signal to obtain the chromatographic signal after baseline correction.
[0011] In the second aspect, this application provides a chromatographic baseline correction device, including:
[0012] An acquisition module, configured to acquire original chromatographic data and preprocess the original chromatographic data to obtain chromatographic data, where the chromatographic data includes a time point t and a chromatographic signal in one-to-one mapping, and the chromatographic signal is used to represent the signal intensity at the time point t, and all the chromatographic data constitutes a chromatographic data set;
[0013] A construction module, configured to construct a first polynomial model for reflecting the change trend of the chromatographic signal in the chromatographic data set, and determine the optimal order of the first polynomial model by using a K-fold cross-validation method;
[0014] A fitting module, configured to construct a second polynomial model according to the optimal order of the first polynomial model, and fit the chromatographic data by using the second polynomial model to obtain an initial fitting baseline;
[0015] An optimization module, configured to iteratively optimize the initial fitting baseline to obtain a final fitting baseline, and subtract the corresponding final estimated value of the final fitting baseline from the chromatographic signal to obtain the chromatographic signal after baseline correction.
[0016] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the steps of the chromatographic baseline correction method described in any one of the above.
[0017] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the chromatographic baseline correction method described in any one of the above are implemented.
[0018] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the chromatographic baseline correction method described in any one of the above are implemented.
[0019] According to the specific embodiments provided by the present application, the following technical effects are disclosed in the present application:
[0020] The present application provides a chromatographic baseline correction method, apparatus, device, medium and product. The method determines the optimal order of the first polynomial model through the K-fold cross-validation method, which can effectively avoid the underfitting or overfitting problems caused by improper order selection, thereby significantly improving the accuracy and applicability of baseline correction. A second polynomial model is constructed according to the optimal order of the first polynomial model, and the chromatographic data is fitted by the second polynomial model to obtain an initial fitted baseline; the initial fitted baseline is iteratively optimized to obtain a final fitted baseline, and the corresponding final estimated value of the final fitted baseline is subtracted from the chromatographic signal to obtain the chromatographic signal after baseline correction. The baseline fitting effect is gradually improved according to the iterative optimization process, effectively suppressing the interference of strong peaks on baseline fitting, thereby improving the accuracy of chromatographic baseline correction. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 It is a schematic flowchart of a chromatographic baseline correction method provided by an embodiment of the present application;
[0023] Figure 2 It is the original chromatogram provided by an embodiment of the present application;
[0024] Figure 3 It is the chromatogram after preprocessing provided by an embodiment of the present application;
[0025] Figure 4 It is the final fitted baseline diagram provided by an embodiment of the present application;
[0026] Figure 5 It is the chromatographic signal after baseline correction provided by an embodiment of the present application;
[0027] Figure 6 It is a schematic diagram of the functional modules of a chromatographic baseline correction device provided by an embodiment of the present application;
[0028] Figure 7 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0030] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0031] Original chromatographic data is usually a set of signal intensity values recorded by a detector at different time points. These data can exist in the form of a table, which contains two columns of information:
[0032] Time: The time stamp corresponding to each data point, which reflects the time interval from the start of sample injection to the detection of the signal, and the unit is usually minutes (min).
[0033] Signal Intensity: The response value recorded by the detector at the corresponding time point, which may be a physical quantity such as voltage, current, or other forms, depending on the detection technology used.
[0034] For example, a simple fragment of original chromatographic data is shown in Table 1:
[0035] Table 1
[0036] Time (minutes) Signal intensity (mv) 0.0 5 0.1 5 ... ... 3.5 150 ... ... 7.8 86 ... ...
[0037] A chromatogram is a graphical representation drawn using the above original chromatographic data, which is used to visually display the changes of each component in the sample over time. In the chromatogram, the x-axis represents time, marking the time required for each component to pass through the chromatographic column and reach the detector, and the y-axis represents signal intensity, showing the response degree of the detector to each component. Each peak represents one or more compounds, and its position (i.e., the appearance time) is called the retention time, which reflects the characteristic time of the compound under specific conditions.
[0038] In an exemplary embodiment, as Figure 1 shown, a chromatographic baseline correction method is provided, including the following steps S101 to step S104. Among them:
[0039] In step S101, the original chromatographic data is obtained, and the original chromatographic data is preprocessed to obtain chromatographic data, which includes a one-to-one mapped time point t and a chromatographic signal, and the chromatographic signal is used to represent the signal intensity at the time point t, and all the chromatographic data constitutes a chromatographic data set.
[0040] In one embodiment, Figure 2 is the original chromatogram, Figure 3 is the chromatogram after pretreatment; the "obtaining original chromatographic data and performing pretreatment on the original chromatographic data to obtain chromatographic data" in step S101 further includes the following sub-steps S1011 - S1013:
[0041] S1011. Perform mean normalization on the original chromatographic data.
[0042] The original chromatographic data includes time data and signal intensity data. The time data is an independent variable, representing the time required for a certain component to pass through the detector from the start of injection, mainly used to determine the position of characteristic peaks, and directly reflects the speed of each component in the sample passing through the chromatographic column. The signal intensity refers to the measured value of the detector's response to the substance flowing through it. The signal intensity directly reflects the concentration or quantity of each component in the sample. When a certain compound passes through the detector, it will cause a change in the signal intensity, forming a peak. The signal intensity may vary greatly due to various factors (such as instrument sensitivity, sample concentration, etc.), so normalization processing is required to ensure the comparability of data between different samples.
[0043] Performing pretreatment on the chromatographic data mainly refers to processing the signal intensity data in the chromatographic data. The chromatographic data can also be understood as a sequence of signal intensity values. The signal intensity in the chromatographic data may form differences due to factors such as sample concentration and detector sensitivity. Therefore, through normalization processing, the differences caused by different experimental conditions can be eliminated. By scaling the signal intensity to a specific range (such as 0 - 1), the data of different samples can be made comparable and not affected by the original scale.
[0044] For example, perform normalization on the original chromatographic data y = [y 1 ,y 2 ,...,y n , and the normalization is as shown in the following formula:
[0045]
[0046] where, y i represents each element in the original chromatographic data, max(y) and min(y) respectively represent the maximum and minimum values in the original chromatographic data, and y’ i represents each element in the normalized data set.
[0047] S1012. Perform standardization on the original chromatographic data after mean normalization.
[0048] For the normalized data y’ = [y’ 1 ,y’2 ,..., y' n Perform Z - score normalization. The purpose is to make the mean of the data 0 and the variance 1. The normalization process is shown in the following formula:
[0049]
[0050] where μ' represents the mean of the original chromatographic data after normalization, σ' represents the standard deviation of the original chromatographic data after normalization, and z i represents the finally normalized data.
[0051] S1013. Perform noise smoothing on the original chromatographic data after normalization to obtain chromatographic data.
[0052] The purpose of noise smoothing is to reduce random noise in the signal, including but not limited to the moving average method or the Savitzky - Golay filtering method.
[0053] The Savitzky - Golay filtering method is a smoothing method based on local polynomial fitting. In each sliding window, a polynomial is used to fit the data in the window, and the value of the fitting polynomial at the center point of the window is used to replace the original data point, thereby achieving a smoothing effect.
[0054] For the normalized data z i , assuming the data length N = 20, the polynomial order is selected as 2, and the sliding window size is 10. Therefore, each time data is processed, 10 consecutive data points will be selected, and the window will slide point by point to cover the entire data sequence. Use the quadratic polynomial P(x) = ax 2 + bx + c to fit the data in the window. Starting from the first point of the data, slide the window point by point and perform polynomial fitting on the data in each window.
[0055] In the first window, the data points covered by the window are z 1 , z 2 ,..., z 10 . Use the quadratic polynomial P(x) = ax 2 + bx + c to fit these 10 data points, and calculate the coefficients a, b, and c of the polynomial by the least - squares method. At the center point of the window (the 5th point), calculate the value of the polynomial P(5) and use P(5) to replace the original data point z 5 .
[0056] In the second window, the window slides one data point to the right, and the data points covered by the window are z 2 , z 3 ,..., z 11 . Use the quadratic polynomial P(x) = ax 2Fit these 10 data points with \(y = ax^{2}+bx + c\), and calculate the coefficients \(a\), \(b\), and \(c\) of the polynomial by the least squares method. Calculate the value \(P(6)\) of the polynomial at the center point of the window (the 6th point), and replace the original data point \(z\) with \(P(6)\). 6 .
[0057] Continue to slide the window until all data points are processed. After the above processing, the value at the center point of each window is replaced by the polynomial fitting value. The finally obtained smoothed data sequence is the result of Savitzky-Golay filtering.
[0058] In step S102, construct a first polynomial model for reflecting the change trend of chromatographic signals in the chromatographic data set, and determine the optimal order of the first polynomial model by the K-fold cross-validation method.
[0059] K-fold cross-validation is a method for evaluating the performance of a model. It divides the data set into K parts (or "folds"), then uses K - 1 parts for training, and the remaining one part for validation, repeating this process K times, each time selecting a different part as the validation set. This can ensure that each data point is used for training and validation, thus obtaining a more reliable evaluation of the model performance.
[0060] Specifically, the above step S102 includes the following sub-steps S1021 - S1024:
[0061] S1021: Randomly divide the chromatographic data set into K mutually exclusive subsets.
[0062] Mutually exclusive subsets mean that the elements in each subset do not overlap with the elements in other subsets. In other words, each data point is only assigned to one subset, rather than multiple subsets. The purpose of doing this is to ensure that during the cross-validation process, the model can encounter data that it has not seen before each time it is validated, thus more accurately evaluating the generalization ability of the model. When dividing the chromatographic data set, make each subset contain the same number of data points, aiming to ensure that each subset can represent the characteristics of the entire chromatographic data set and can fairly evaluate the model performance during the training and validation of the model.
[0063] If the size of the dataset is not divisible by K, for example, if a dataset has 105 samples and 5-fold cross-validation is required (K = 5), it is impossible to divide the dataset into 5 subsets of exactly equal size. In this case, sufficient samples can be randomly drawn from the dataset so that the remaining number of samples is divisible by K. For example, if the dataset has 105 samples, 5 samples can be randomly drawn, leaving 100 samples, which can then be divided evenly by 5. Usually, unequal subset sizes are allowed, but they should be as close as possible. For example, 105 samples can be divided into 5 subsets, with 4 subsets having 21 samples each and the last subset having 19 samples.
[0064] S1022. Sequentially select one subset from K mutually exclusive subsets as the validation set, and the remaining K - 1 subsets as the training set.
[0065] S1023. For each candidate order, perform the K-fold cross-validation loop steps. In each loop, use the training set to fit the first polynomial model, calculate the model error on the validation set, and obtain the average error of the first polynomial model based on the model error.
[0066] S1024. Compare the K average errors corresponding to the K candidate orders, and select the candidate order corresponding to the smallest average error as the optimal order.
[0067] Specifically, for each candidate polynomial order, select the i-th subset as the validation set, and the remaining k - 1 subsets as the training set. Select an order n of a polynomial, and use the training set to fit the first polynomial model. The general form of the polynomial model is: y = a 0 +a 0 x+a 2 x 2 +...+a n x n . The coefficients a 0 , a 1 ,...a n of the polynomial model can be calculated by the least squares method.
[0068] For each candidate polynomial order, evaluate the model performance using the validation set. For example, evaluate the model performance by calculating the mean squared error. Among them, the mean squared error evaluated with the i-th subset as the validation set is expressed by the following formula:
[0069]
[0070] where y pred represents the model prediction value, y text represents the true value; N represents the number of samples in each subset; i represents the current fold index, used to distinguish different validation sets in cross-validation.
[0071] The average of the k - fold validation errors is taken as the final evaluation index of the model performance. The calculation formula for the cross - validation error is as follows:
[0072]
[0073] k represents the total number of folds in cross - validation, that is, the total number of validation sets. Common values are 5 or 10, which means the dataset will be divided into 5 or 10 parts. Each time, one of the parts is used as the validation set, and the rest are used as the training set. Compare the average errors of different polynomial orders, and select the polynomial order with the smallest average error as the optimal order n.
[0074] In a specific example, assume there is a chromatographic dataset containing 100 samples, and each sample contains the intensity value of the chromatographic signal. Select K = 5 for 5 - fold cross - validation, that is, randomly divide the dataset into 5 parts, with each part containing 20 samples.
[0075] Set the candidate orders d = 1, 2, 3, namely the linear model (1st order), quadratic model (2nd order), and cubic model (3rd order). For each order, perform 5 loops. Each time, select a different subset as the validation set, and the remaining 4 subsets as the training set.
[0076] When d = 2:
[0077] Model form: y = ax 2 + bx + c;
[0078] In the first loop: Select the 1st subset (20 samples) as the validation set, and the remaining 2 - 5th subsets (80 samples) as the training set. Fit a 2nd - order polynomial model on the training set, and calculate the coefficients a = 0.5, b = 1.2, c = 0.1 by the least - squares method.
[0079] For the 20 samples, each sample has a predicted value y pred and an actual value y text , and then calculate the mean - square error of the model on the validation set
[0080] In the second loop:
[0081] Training set: The 1st, 3 - 5th subsets (80 samples);
[0082] Validation set: The 2nd subset (20 samples);
[0083] Coefficients: a = 0.6, b = 1.1, c = 0.2.
[0084]
[0085] Repeat the above steps until all 5 subsets have been used as the validation set once.
[0086] The third cycle:
[0087] Training set: subsets 1, 2, 4 - 5 (80 samples);
[0088] Validation set: subset 3 (20 samples);
[0089] Coefficients: a = 0.4, b = 1.3, c = 0.3.
[0090]
[0091] The fourth cycle:
[0092] Training set: subsets 1, 2, 3, 5 (80 samples);
[0093] Validation set: subset 4 (20 samples);
[0094] Coefficients: a = 0.7, b = 1.0, c = 0.4.
[0095]
[0096] The fifth cycle:
[0097] Training set: subsets 1 - 4 (80 samples);
[0098] Validation set: subset 5 (20 samples);
[0099] Coefficients: a = 0.55, b = 1.25, c = 0.35.
[0100]
[0101] For d = 2:
[0102] Compare the average errors for d = 1, 2, 3 and select the order with the minimum average error as the optimal order.
[0103] For d = 1, the average error E cv = 0.65;
[0104] For d = 2, the average error E cv = 0.35;
[0105] For d = 3, the average error E cv = 0.40;
[0106] From the above calculation results, it can be seen that d = 2 (quadratic model) has the smallest average error. Therefore, a second-order polynomial is selected as the optimal model order.
[0107] In step S103, a second polynomial model is constructed according to the optimal order of the first polynomial model, and the chromatographic data is fitted by the second polynomial model to obtain an initial fitted baseline.
[0108] In one embodiment, the above step S103 includes the following sub-steps S1031 - S1033:
[0109] S1031. Construct a second polynomial model according to the chromatographic data set and the optimal order.
[0110] Specifically, given the denoised chromatographic data Y = {(x 1 , y 1 ), (x 2 , y 2 ), …, (x N , y N )} of N points, where x i represents the time point, y i represents the signal intensity, and the optimal order n obtained through step S1024. The second polynomial model is represented by the following formula:
[0111] P(x) = c 0 + c 1 x + … + c n-1 x n-1 + c n x n ;
[0112] where c i (i = 0, 1, 2, …, n) are the coefficients of the polynomial, and n is the optimal order of the polynomial.
[0113] To simplify the calculation, the second polynomial model can also be represented by the following formula:
[0114] P(x i ) = AC;
[0115] where A represents an N×(n + 1) coefficient matrix; C represents a polynomial coefficient vector.
[0116] For each data point x i , calculate the power terms of each data point x i from x 0 to x n to form each row of the A matrix. The matrix A is represented in the following way:
[0117]
[0118] The polynomial coefficient vector C is represented as follows:
[0119]
[0120] S1032. Construct an error function according to the chromatographic data set and the second polynomial model.
[0121] Specifically, define the error function E(c) as the sum of the squares of the differences between the data points y i and the polynomial P(x i ), that is:
[0122]
[0123] After simplification:
[0124] E(c) = ||y - AC|| 2 ;
[0125] The goal of constructing the error function is to find a set of polynomial coefficient vectors C such that the sum of the squares of the errors between the polynomial predicted values and the actual observed values is minimized. ||y - AC|| 2 represents the square of the Euclidean distance between the actual observed value vector y and the polynomial predicted value AC, that is, the sum of the squares of the errors.
[0126] S1033. Based on the error function, obtain the optimal coefficients of the second polynomial, and obtain the initial fitting baseline according to the optimal coefficients of the second polynomial.
[0127] Specifically, to minimize the deviation E(c) between the fitted chromatographic data and the original chromatographic data, using the least squares method, take the partial derivatives of the right side of E(c) = ||y - AC|| 2 with respect to c j (j = 0, 1, 2, ……, n) and set them equal to 0, and solve to get:
[0128] C = (A T A) -1 A T y;
[0129] where C represents the optimal coefficients of the fitted polynomial. Finally, use the obtained coefficients to construct the following polynomial model:
[0130]
[0131] P(x) is the initial fitting baseline.
[0132] Exemplarily, for the denoised chromatographic data set Y = {(x 1 , y 1 ), (x 2, y 2 ), …, (x N , y N )} for fifth-order polynomial fitting, where the basic form of the polynomial fitting is:
[0133] P(x) = c 0 + c 1 x + … + c 5 x 5 (1);
[0134] Among them, c i (i = 0, 1, 2, …, 5) are the coefficients of the polynomial.
[0135] Define the error function E(c) as the sum of the squares of the differences between the data points y i and the polynomial P(x i ), that is:
[0136]
[0137] Express the polynomial in matrix form: P(x i ) = AC;
[0138] Among them, represents the coefficient matrix, represents the polynomial coefficients.
[0139] Rewrite formula (2) as:
[0140] E(c) = ||y - AC|| 2 (3);
[0141] To minimize the deviation E(c) between the fitted chromatographic data and the original chromatographic data, using the least squares method, take the partial derivatives of the right side of equation (3) with respect to c j (j = 0, 1, 2, …, 5) and set them equal to 0, and solve to get:
[0142] C = (A T A) -1 A T y (4);
[0143] Among them, C is the optimal coefficient of the fitted polynomial. Use the obtained coefficients to construct a polynomial model:
[0144]
[0145] This polynomial model is the best fit for the denoised chromatographic data points.
[0146] In step S104, the initial fitting baseline is iteratively optimized to obtain the final fitting baseline, and the corresponding final estimated value of the final fitting baseline is subtracted from the chromatographic signal to obtain the chromatographic signal after baseline correction.
[0147] In one embodiment, "iteratively optimizing the initial fitting baseline to obtain the final fitting baseline" in step S104 includes: performing multiple iterative optimization operations until a predetermined convergence condition is reached.
[0148] The iterative optimization operation includes the following sub-steps A1 - A4:
[0149] A1. Based on the target fitting baseline, for each time point t, determine a corresponding fitting baseline value; wherein, the fitting baseline value is used to represent the estimated signal intensity value at time point t; the target fitting baseline in the first iterative optimization operation is the initial fitting baseline, and the target fitting baseline in any other iterative optimization operation is the new fitting baseline obtained in the previous iterative optimization operation.
[0150] A2. Based on the chromatographic data, obtain the actual signal intensity value corresponding to each time point t.
[0151] A3. Generate a new sequence according to the comparison result between the estimated signal intensity value and the actual signal intensity value at each time point t.
[0152] A4. Recalculate the new fitting baseline based on the new sequence.
[0153] Wherein, the new fitting baseline obtained when the predetermined convergence condition is reached is the final fitting baseline.
[0154] In one embodiment, generating a new sequence according to the comparison result between the estimated signal intensity value and the actual signal intensity value at each time point t includes:
[0155] If the estimated signal intensity value at a certain time point t is greater than the actual signal intensity value, then retain the actual signal intensity value and assign the actual signal intensity value to the corresponding position in the new sequence.
[0156] If the actual signal intensity value at a certain time point t is greater than the estimated signal intensity value, then retain the estimated signal intensity value and assign the estimated signal intensity value to the corresponding position in the new sequence.
[0157] Specifically, the chromatographic signal after denoising processing is assigned as the initial data to the sequence y 0 (t), where each time point t has a corresponding actual signal intensity value. Perform polynomial curve fitting on y 0 (t) to generate a fitting curve E(t) for estimating the background baseline, that is, the initial fitting baseline.
[0158] For each time point t, calculate its corresponding fitted value E(t) (i.e., the estimated signal strength) and the actual signal strength value y 0 (t), and generate a new sequence Y 1 (t) according to the comparison result:
[0159] If E(t) > y 0 (t), then retain the actual signal strength value y 0 (t), and assign this value to the corresponding position in the new sequence Y 1 (t). Because if the fitted value is larger, it may mean that this point is on the characteristic peak, or it may be because high background values that may exist but actually do not exist are considered in the fitting process. Therefore, the actual measured value is selected to protect the characteristic peak, so as to protect the characteristic information such as the characteristic peak from being misjudged as part of the baseline, resulting in incorrect removal or weakening.
[0160] If Y 0 (t) > E(t), then retain the fitted value E(t), and assign this value to the corresponding position in the new sequence Y 1 (t). Because if the actual signal strength value is large, it indicates that this point is closer to the background noise or the baseline part. Therefore, the fitted value is selected to better reflect the background baseline.
[0161] This step is actually correcting the part that may be misidentified as a signal to make it closer to the true baseline. Use Y 1 (t) instead of the original y 0 (t) to perform the next round of baseline fitting. This is because Y 1 (t) has already partially removed the part considered to be a signal, so it is more suitable for estimating the background baseline.
[0162] Use the updated new sequence Y 1 (t) as the basis for the next round of iteration, and perform the polynomial fitting and comparison update steps again. After each round of iteration, calculate the difference between the new and old sequences until the change amount of these differences is less than a preset threshold (such as the standard deviation). At this time, it is considered that the convergence state has been reached and the iteration ends.
[0163] When the iteration process meets the convergence condition, the final fitted baseline Y final (t) is obtained, that is, the sequence generated in the last iteration. The final fitted baseline graph corresponding to the sequence generated in the last iteration is as shown in Figure 4 . Subtract Y 0 (t) from the sequence y final (t) to obtain the chromatographic signal L(t) after baseline correction, L(t) = y 0 (t) - Y final (t). From Figure 5It can be seen that the signal L(t) after baseline correction shows the chromatographic data after baseline correction, with characteristic peaks being more clearly visible and background interference being effectively removed, which makes subsequent quantitative analysis, qualitative analysis, and characteristic peak identification more accurate and reliable.
[0164] It should be noted that the initial fitted baseline E(t) may not be completely accurate because it is based on the original chromatographic signal y 0 (t), which contains the true signal peak as well as noise or non-ideal baseline drift. By comparing y 0 (t) and E(t) and generating a new sequence (such as Y 1 (t)), the baseline estimate can be adjusted according to the actual signal so that it better reflects the true background signal rather than containing potential signal peaks. The new sequence generated in each iteration (such as Y 1 (t), Y 2 (t),...) is optimized based on the result of the previous iteration. These new sequences gradually reduce the inclusion of the true signal peak and more reflect the background information. Therefore, refitting the baseline based on these new sequences can obtain a more accurate background estimate. Through multiple iterations, the algorithm can gradually reduce the error until a preset convergence criterion is reached (for example, the change amount between consecutive iterations is less than a certain threshold). The purpose of doing this is to ensure that the final baseline estimate is as close as possible to the true background signal, thereby effectively removing background interference without affecting the true signal peak.
[0165] Suppose a new sequence Y 1 (t) is obtained after the first iteration. At this time, use Y 1 (t) to replace y 0 (t) for a new baseline fit because it is desired that the baseline is closer to the background rather than the signal peak. As the number of iterations increases, the new sequence will get closer and closer to the true baseline and farther away from the signal peak region, and finally an accurate baseline estimate is obtained. Using the new sequence to calculate the new fitted baseline is to gradually optimize the baseline estimate to make it more accurately reflect the background signal, thereby effectively performing baseline correction. This method relies on the principle of iterative improvement and finally achieves the ideal result through continuous adjustment and optimization.
[0166] The judgment process of the convergence condition can be achieved in the following way:
[0167] For three iterations, the initial fitted baseline is E(t). In the first iteration, a new sequence Y 1 (t) is generated from the initial fitted baseline E(t), and a new fitted baseline E’(t) is calculated based on the new sequence Y 1 (t). In the second iteration, a new sequence Y 2 (t) is generated from E’(t), and based on the new sequence Y2 (t) Calculate the new fitted baseline E”(t). In the third iteration, a new sequence Y 3 (t) is generated from E”(t), and based on the new sequence Y 3 (t), calculate the new fitted baseline E”’(t).
[0168] The difference between the newly generated fitted baseline E’(t) (or E”(t), E”’(t),...) after each iteration and the fitted baseline E(t) obtained in the previous iteration can be regarded as an error, and this difference can be measured in various ways. The most common is to use the root mean square error or the mean absolute error.
[0169] Assume that the fitted baseline after the i-th iteration is E i (t), and the fitted baseline after the (i + 1)-th iteration is E i+1 (t). For each time point t, the difference between the two iterations is expressed by the following formula:
[0170] Difference(t) = |E i+1 (t) - E i (t)|;
[0171] The root mean square error of the difference is calculated by the following formula:
[0172]
[0173] The mean absolute error of the difference is calculated by the following formula:
[0174]
[0175] If the change in error (such as the root mean square error or the mean absolute error E) in several consecutive iterations is lower than a preset threshold, it is considered that the algorithm has converged and the iteration stops. This threshold can be adjusted according to the requirements of the specific application and the characteristics of the data.
[0176] Based on the same inventive concept, the embodiments of the present application also provide a chromatographic baseline correction device for implementing the above-mentioned ones. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the chromatographic baseline correction device provided below can refer to the limitations on the chromatographic baseline correction method in the above text, and will not be repeated here.
[0177] In an exemplary embodiment, as Figure 6 shown, a chromatographic baseline correction device is provided, including:
[0178] An acquisition module 610, configured to acquire original chromatographic data, perform preprocessing on the original chromatographic data to obtain chromatographic data, where the chromatographic data includes a time point t and a chromatographic signal in one-to-one mapping, the chromatographic signal is used to represent the signal intensity at the time point t, and all the chromatographic data constitutes a chromatographic data set;
[0179] A construction module 620, configured to construct a first polynomial model for reflecting the change trend of the chromatographic signal in the chromatographic data set, and determine the optimal order of the first polynomial model by using a K-fold cross-validation method;
[0180] A fitting module 630, configured to construct a second polynomial model according to the optimal order of the first polynomial model, and fit the chromatographic data by using the second polynomial model to obtain an initial fitting baseline;
[0181] An optimization module 640, configured to iteratively optimize the initial fitting baseline to obtain a final fitting baseline, and subtract the corresponding final estimated value of the final fitting baseline from the chromatographic signal to obtain the chromatographic signal after baseline correction.
[0182] As an optional implementation manner, in the aspect of performing preprocessing on the original chromatographic data to obtain chromatographic data, the acquisition module 610 specifically includes:
[0183] Performing mean normalization processing on the original chromatographic data;
[0184] Performing standardization processing on the mean-normalized original chromatographic data;
[0185] Performing noise smoothing processing on the standardized original chromatographic data to obtain the chromatographic data.
[0186] As an optional implementation manner, the construction module 620 is specifically configured to:
[0187] Randomly divide the chromatographic data set into K mutually exclusive subsets;
[0188] Successively select one subset of the K mutually exclusive subsets as a validation set, and the remaining K - 1 subsets as training sets;
[0189] For each candidate order, perform K cross-validation loop steps. In each loop, use the training set to fit the first polynomial model, calculate the model error on the validation set, and obtain the average error of the first polynomial model based on the model error;
[0190] Compare the K average errors corresponding to the K candidate orders, and select the candidate order corresponding to the smallest average error among them as the optimal order.
[0191] As an alternative implementation, the fitting module 630 is specifically configured to:
[0192] Construct a second polynomial model according to the chromatographic data set and the optimal order;
[0193] Construct an error function according to the chromatographic data set and the second polynomial model;
[0194] Based on the error function, obtain the optimal coefficients of the second polynomial, and obtain the initial fitting baseline according to the optimal coefficients of the second polynomial.
[0195] As an alternative implementation, in terms of iteratively optimizing the initial fitting baseline to obtain the final fitting baseline, the optimization module 640 is specifically configured to:
[0196] Perform multiple iterative optimization operations until a predetermined convergence condition is reached;
[0197] The iterative optimization operation includes:
[0198] Based on the target fitting baseline, for each time point t, determine a corresponding fitting baseline value; wherein, the fitting baseline value is used to represent the signal intensity estimation value at time point t; the target fitting baseline in the first iterative optimization operation is the initial fitting baseline, and the target fitting baseline in any other iterative optimization operation is the new fitting baseline obtained in the previous iterative optimization operation;
[0199] Based on the chromatographic data, obtain the actual signal intensity value corresponding to each time point t;
[0200] Generate a new sequence according to the comparison result between the signal intensity estimation value and the actual signal intensity value at each time point t;
[0201] Recalculate a new fitting baseline based on the new sequence;
[0202] Wherein, the new fitting baseline obtained when the predetermined convergence condition is reached is the final fitting baseline.
[0203] As an alternative implementation, in terms of generating a new sequence according to the comparison result between the signal intensity estimation value and the actual signal intensity value at each time point t, the optimization module 640 is further configured to:
[0204] If the signal intensity estimation value at a certain time point t is greater than the actual signal intensity value, then retain the actual signal intensity value and assign the actual signal intensity value to the corresponding position in the new sequence;
[0205] If the actual signal strength value at a certain time point t is greater than the signal strength estimated value, then retain the signal strength estimated value and assign the signal strength estimated value to the corresponding position in the new sequence.
[0206] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 7 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store chromatographic data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements a chromatographic baseline correction method.
[0207] Those skilled in the art can understand that Figure 7 the structure shown in
[0208] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0209] In an exemplary embodiment, a computer-readable storage medium is also provided, storing a computer program, which, when executed by a processor, implements the steps in the above method embodiments.
[0210] In an exemplary embodiment, a computer program product is provided, including a computer program, which, when executed by a processor, implements the steps in the above method embodiments.
[0211] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0212] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0213] The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., and are not limited thereto.
[0214] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0215] In this article, specific examples are used to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present application.
Claims
1. A chromatographic baseline correction method, characterized in that: The chromatographic baseline correction method comprises: Acquire original chromatographic data, and preprocess the original chromatographic data to obtain chromatographic data, wherein the chromatographic data includes a one-to-one mapping of a time point t and a chromatographic signal, wherein the chromatographic signal is used to represent the signal intensity at the time point t, and all the chromatographic data constitute a chromatographic data set; Constructing a first polynomial model for reflecting the variation trend of the chromatographic signal in the chromatographic data set, and determining the optimal order of the first polynomial model by a K-fold cross-validation method; constructing a second polynomial model according to the optimal order of the first polynomial model, and fitting the chromatographic data by using the second polynomial model to obtain an initial fitting baseline; The initial fitted baseline is iteratively optimized to obtain a final fitted baseline, and a final estimated value corresponding to the final fitted baseline is subtracted from the chromatographic signal to obtain the chromatographic signal after baseline correction.
2. The chromatographic baseline correction method according to claim 1, characterized in that: The raw chromatographic data is preprocessed to obtain chromatographic data, including: Performing mean normalization processing on the original chromatographic data; Standardizing the original chromatographic data after the mean normalization process; The normalized original chromatographic data is subjected to noise smoothing processing to obtain the chromatographic data.
3. The chromatographic baseline correction method according to claim 1, characterized in that: The step of constructing a first polynomial model for reflecting the variation trend of the chromatographic signal in the chromatographic data set, and determining the optimal order of the first polynomial model by a K-fold cross validation method, comprises: Randomly dividing the chromatographic data set into K mutually exclusive subsets; Select one of the K mutually exclusive subsets as a validation set, and the remaining K-1 subsets as training sets; For each candidate order, performing K cross-validation cycle steps, in each cycle, fitting the first polynomial model using the training set, calculating the model error on the validation set, and obtaining an average error of the first polynomial model based on the model error; The K average errors corresponding to the K candidate orders are compared, and the candidate order corresponding to the smallest average error is selected as the optimal order.
4. The chromatographic baseline correction method according to claim 1, characterized in that: The step of constructing a second polynomial model according to the optimal order of the first polynomial model, and fitting the chromatographic data by using the second polynomial model to obtain an initial fitting baseline comprises: constructing a second polynomial model according to the chromatographic data set and the optimal order; constructing an error function based on the chromatographic data set and the second polynomial model; Based on the error function, optimal coefficients of a second polynomial are obtained, and an initial fitting baseline is obtained according to the optimal coefficients of the second polynomial.
5. The chromatographic baseline correction method according to claim 1, characterized in that: The iterative optimization of the initial fitting baseline to obtain a final fitting baseline includes: Perform multiple iterative optimization operations until a predetermined convergence condition is reached; The iterative optimization operation includes: Based on the target fitting baseline, for each time point t, a corresponding fitting baseline value is determined; wherein the fitting baseline value is used to represent the signal intensity estimation value at the time point t; the target fitting baseline in the first iterative optimization operation is the initial fitting baseline, and the target fitting baseline in any other iterative optimization operation is the new fitting baseline obtained in the previous iterative optimization operation; Based on the chromatographic data, obtaining an actual signal intensity value corresponding to each time point t; Generate a new sequence based on the comparison result of the signal strength estimation value and the actual signal strength value at each time point t; recalculating a new fitted baseline based on the new sequence; The new fitting baseline obtained when the predetermined convergence condition is reached is the final fitting baseline.
6. The chromatographic baseline correction method according to claim 5, characterized in that: The generating a new sequence according to the comparison result of the signal strength estimation value and the actual signal strength value at each time point t comprises: If the signal strength estimation value at a certain time point t is greater than the actual signal strength value, retaining the actual signal strength value and assigning the actual signal strength value to the corresponding position in the new sequence; If the actual signal strength value at a certain time point t is greater than the signal strength estimation value, the signal strength estimation value is retained and assigned to a corresponding position in the new sequence.
7. A chromatographic baseline correction device, characterized in that: The chromatographic baseline correction device comprises: an acquisition module, used for acquiring raw chromatographic data and preprocessing the raw chromatographic data to obtain chromatographic data, wherein the chromatographic data includes a one-to-one mapping of a time point t and a chromatographic signal, wherein the chromatographic signal is used to represent the signal intensity at the time point t, and all the chromatographic data constitute a chromatographic data set; A construction module, used for constructing a first polynomial model for reflecting the variation trend of the chromatographic signal in the chromatographic data set, and determining the optimal order of the first polynomial model by a K-fold cross validation method; A fitting module, used for constructing a second polynomial model according to the optimal order of the first polynomial model, and fitting the chromatographic data by means of the second polynomial model to obtain an initial fitting baseline; The optimization module is used to iteratively optimize the initial fitting baseline to obtain a final fitting baseline, and subtract a final estimated value corresponding to the final fitting baseline from the chromatographic signal to obtain the chromatographic signal after baseline correction.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the chromatographic baseline correction method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the chromatographic baseline correction method described in any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the chromatographic baseline correction method described in any one of claims 1 to 6 are implemented.