Data processing method for component content detection

Adjusting the peak sensitivity of the ALS algorithm through adaptive weights, the problem of baseline distortion in the baseline correction of Cordyceps sinensis chromatography in traditional ALS algorithm is solved, and a more accurate detection of Cordyceps sinensis concentration is achieved.

CN120294228AInactive Publication Date: 2025-07-11SHAANXI EVERGREEN HERBAL BIOTECH CO LTD

Patent Information

Application Number
CN202510779584.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the baseline correction of Cordyceps Cordyceps chromatography in Cordyceps Cordyceps, the use of fixed peak sensitivity parameters leads to baseline correction distortion, affecting the detection accuracy of Cordyceps content.

Method used

The peak sensitivity of the ALS algorithm is adjusted using adaptive weights, and the peak sensitivity is dynamically adjusted by calculating the near-peak possibility and smoothing requirements of each data point, and baseline correction is performed.

Benefits of technology

It improves the accuracy of baseline correction, ensures the accuracy of chromatographic data, and provides a reliable basis for the detection of Cordyceps concentration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120294228A_ABST
    Figure CN120294228A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a data processing method for component content detection.The method comprises the steps that a chromatogram of a cordyceps militaris extract is obtained through high performance liquid chromatography, the abscissa of the chromatogram is retention time, and the ordinate of the chromatogram is an absorbance value; and performing baseline correction on the chromatogram by using an ALS algorithm, in the correction process, adjusting the peak sensitivity of the ALS algorithm based on the adaptive weight of each data point in the chromatogram, and performing baseline correction so as to detect the cordycepin concentration in the cordyceps militaris extract based on the chromatogram after baseline correction. According to the method, the baseline correction accuracy of the chromatographic data of the cordyceps militaris extract can be improved, so that accurate detection of the concentration of cordycepin in cordyceps militaris is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing. More specifically, the present invention relates to a data processing method for detecting component contents. Background Art

[0002] Cordyceps militaris is a medicinal fungus rich in various bioactive components. Its core active ingredient, cordycepin, has important pharmacological effects such as antioxidant and immune regulation, and is a key indicator for measuring the quality and medicinal value of Cordyceps militaris. Detecting the content of cordycepin is not only a necessary means to evaluate its medicinal efficacy, but also an important quality control link to ensure the safety and effectiveness of related products (such as health products and drugs).

[0003] In the prior art, high performance liquid chromatography (HPLC) has been widely used in the detection and analysis of cordycepin content due to its high separation efficiency and detection sensitivity. However, the phenomenon of chromatographic baseline drift (caused by factors such as mobile phase gradient elution program, column temperature fluctuation, or detector response drift) will significantly affect the peak recognition accuracy, resulting in distortion of the integration area of the target component, and then causing systematic deviation of the content determination result.

[0004] In the chromatographic data preprocessing stage, baseline correction is a key technical means to eliminate noise interference and ensure the accuracy of quantitative analysis. The traditional asymmetric least squares method (ALS algorithm) realizes baseline fitting by alternately performing baseline smoothing and signal correction. However, it has inherent defects in practical applications: its fixed-width window parameter cannot adaptively match the different chemical characteristics of wide peaks (such as the main cordycepin peak) and narrow peaks (such as coexisting interfering components) in the chromatogram of Cordyceps militaris extract, resulting in baseline lift artifacts in the strong peak region due to excessive smoothing, and loss of effective signals in the weak peak region due to insufficient noise suppression, leaving noise residues, thus affecting the effect of baseline fitting, resulting in low accuracy of the obtained chromatographic data, and then seriously affecting the detection accuracy of cordycepin content. Summary of the Invention

[0005] In order to solve the problem of baseline correction distortion caused by using a fixed peak sensitivity parameter in the cordycepin chromatographic baseline correction of Cordyceps militaris by the asymmetric least squares method (ALS algorithm), the present invention provides a data processing method for detecting component contents. The method includes: Using high performance liquid chromatography to obtain a chromatogram of Cordyceps militaris extract, where the abscissa of the chromatogram is the retention time and the ordinate is the absorbance value; Using the ALS algorithm to perform baseline correction on the chromatogram. During the correction process, based on the adaptive weights of each data point in the chromatogram, adjust the peak sensitivity of the ALS algorithm and perform baseline correction to detect the cordycepin concentration in the Cordyceps militaris extract based on the chromatogram after baseline correction; Among them, the method for obtaining the adaptive weight of each data point includes: taking the maximum value points in the chromatogram as peak data points, and calculating the near-peak possibility of each data point. The near-peak possibility represents the proximity of the corresponding data point to the peak data points with low isolation and high absorbance values; Based on the near-peak possibility and the fluctuation degree of the absorbance values within the neighborhood range of each data point, calculate the adaptive weight of each data point. The adaptive weight is negatively correlated with the near-peak possibility of the corresponding data point and the fluctuation degree of the absorbance values within the neighborhood range of the corresponding data point.

[0006] The present invention uses the adaptive weight of each data point to adjust the peak sensitivity of the ALS algorithm, which can significantly reduce the peak sensitivity of the data points in the strong peak region (the region where the data points with high near-peak possibility are located) and the noise interference region (the region where the data points with high smoothing requirement are located). Therefore, the baseline fitting will not be interfered by the peaks or noise, thus avoiding the problems of abnormal elevation of the baseline in the strong peak region and noise residue. This ensures the accuracy of the corrected baseline, and further improves the accuracy of the corrected chromatographic data. Finally, it provides an accurate data basis for realizing the detection of the cordycepin concentration in the cordyceps militaris extract.

[0007] Preferably, calculating the near-peak possibility of each data point satisfies the following relational expression: ; In the formula, is the near-peak possibility of the th data point in the chromatogram; is the absorbance value of the th peak data point in the chromatogram; is the maximum absorbance value of all peak data points in the chromatogram; is the isolation value of the th peak data point in the chromatogram. The isolation value represents the degree of isolation of the position and absorbance value of the peak data point within the neighborhood range; is the sampling order difference between the th data point in the chromatogram and the th peak data point in the chromatogram; is the membership symbol; is the natural exponential function; is the function with the maximum value as the return value; is the set composed of all peak data points.

[0008] When calculating the near-peak possibility of each data point in the present invention, by measuring the isolation of each maximum value point in terms of position and absorbance value, the interference of noise points can be eliminated, thus ensuring the accuracy of the near-peak possibility.

[0009] Preferably, the method for obtaining the isolation value includes: Count the number of peak data points within the neighboring range of each peak data point in the statistical chromatogram, take the opposite number, use the opposite number as the power of the exponential function, and perform the power operation to obtain the first index of each peak data point; Calculate the difference in absorbance values between each peak data point and each data point within the neighboring range, sum and take the average, and use the normalized value of the obtained average as the second index of each peak data point; Calculate the isolation value of each peak data point based on the first index and the second index, and the isolation value is positively correlated with both the second index and the first index.

[0010] Preferably, the degree of fluctuation is the relative degree of fluctuation, and the method for obtaining the relative degree of fluctuation includes: Select the target neighboring range, calculate the difference in absorbance values between each data point within the target neighboring range and the adjacent data points before and after, sum them, and perform another summation operation on the obtained cumulative sums to obtain the local fluctuation value of the absorbance value within the target neighboring range; Calculate the difference in absorbance values of the data points at both ends within the target neighboring range to obtain the overall fluctuation value of the absorbance value within the target neighboring range, and normalize the ratio of the local fluctuation value to the overall fluctuation value to obtain the relative degree of fluctuation of the absorbance value within the target neighboring range.

[0011] The present invention can accurately evaluate the possibility that each data point is in the noise interference area, providing an accurate data basis for the subsequent calculation of the adaptive weight.

[0012] Preferably, the relative degree of fluctuation is used as the smoothing requirement degree of the corresponding data point, and the smoothing requirement degree satisfies the following relational expression: ; In the formula, is the smoothing requirement degree of the th data point in the chromatogram; is the preset neighboring radius; , and are the absorbance values of the th data point, the th data point, and the th data point respectively; , are the absorbance values of the data points at both ends within the neighboring range of the th data point in the chromatogram; is the preset hyperparameter; is the normalization function; is the absolute value symbol.

[0013] Preferably, the adaptive weight of each data point satisfies the following relational expression:

[0014] In the formula, is the adaptive weight of the th data point in the chromatogram; is the near-peak possibility of the th data point in the chromatogram; is the smoothing requirement of the th data point in the chromatogram; is the hyperbolic tangent function.

[0015] The present invention reduces the peak sensitivity of the ALS algorithm when performing baseline fitting on strong peak regions and noise interference regions, thereby avoiding baseline fitting to strong peak regions and noise interference regions.

[0016] Preferably, adjusting the peak sensitivity of the ALS algorithm based on the adaptive weights of several data points in the chromatogram includes: Obtaining the preset peak sensitivity of the ALS algorithm when performing baseline correction on the chromatogram, and performing a multiplication operation on the adaptive weight of each data point and the preset peak sensitivity to obtain the adaptive peak sensitivity.

[0017] Preferably, the process of performing baseline correction based on the adaptive peak sensitivity includes: Taking the adaptive peak sensitivity as the sample weight of the corresponding data point, and performing low-pass filtering on the chromatographic data in the chromatogram, so as to use the obtained filtering result as the initial baseline in the ALS algorithm; Based on the sample weight and the initial baseline, constructing an objective function for baseline correction by the ALS algorithm, and updating the initial baseline with the goal of minimizing the objective function. After the update operation, the corrected baseline is obtained.

[0018] The present invention can improve the accuracy of the baseline correction result of the ALS algorithm.

[0019] Preferably, the objective function satisfies the following relational expression: ; In the formula, is the objective function value; is the sample weight of the th data point in the chromatogram; is the absorbance value of the th data point in the chromatogram; is the baseline intensity of the th data point in the initial baseline; is the total number of data points in the chromatogram; is the preset regularization parameter; is the second-order difference representation symbol.

[0020] Preferably, the method for obtaining the maximum points among all data points includes: Screen each data point greater than the adjacent data points before and after from the chromatogram to obtain several maximum points.

[0021] The present invention has the following effects: The present invention adaptively adjusts the peak sensitivity of the ALS algorithm through the near-peak possibility and smoothness requirement degree of each data point in the chromatogram, and can set a smaller peak sensitivity for the data points in the strong peak region and the noise interference region of the chromatogram. This enables the baseline to remain smooth in these two regions, avoiding the problems of abnormal uplift of the baseline in the strong peak region and noise residue after baseline fitting. Therefore, the accuracy of the baseline correction result is improved, and thus accurate chromatographic data can be obtained, providing a reliable basis for realizing the accurate detection of cordycepin concentration. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] By referring to the detailed description below with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein: Figure 1 is a schematic flow chart of the steps of a data processing method for detecting component content according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0023] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0024] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0025] Referring to Figure 1 , a data processing method for detecting component content includes steps S1 - S3, specifically as follows: S1: Using high performance liquid chromatography, obtain the chromatogram of the cordyceps militaris extract, where the abscissa of the chromatogram is the retention time and the ordinate is the absorbance value.

[0026] It should be noted that the Cordyceps militaris extract contains various components such as cordycepin, polysaccharides, and proteins, and the matrix is complex. By utilizing the distribution difference between the non-polar stationary phase and the polar mobile phase and dynamically changing the elution strength, cordycepin can be completely separated from other components within a reasonable retention time. Therefore, high-performance liquid chromatography can be used to determine the retention time of cordycepin in the chromatographic data of the extract, so as to determine the concentration of cordycepin in Cordyceps militaris based on the retention.

[0027] Specifically, high-performance liquid chromatography (HPLC) can be used to real-time collect the absorbance signals (mAU) of the extract within the retention time range, such as 0 - 15 min, at a sampling frequency of 10 Hz, and construct a chromatogram with the retention time as the abscissa and the absorbance value as the ordinate. Each data point in the chromatogram is a sampling point.

[0028] S2: Use the ALS algorithm to perform baseline correction on the chromatogram. During the correction process, based on the adaptive weights of each data point in the chromatogram, adjust the peak sensitivity of the ALS algorithm and perform baseline correction.

[0029] It should be noted that the detector is the core component of the high-performance liquid chromatography (HPLC) system, which is used to identify and quantify the chemical components in the sample, usually by measuring the absorbance of the sample at a specific wavelength to generate signals. When the absorption intensities of the two phases (such as acetonitrile and aqueous solution) in the mobile phase at the detection wavelength are different, the dynamic change in the proportion of the mobile phase during gradient elution will cause the overall rise or fall of the baseline signal, and this phenomenon is called baseline drift.

[0030] The ALS (Adaptive Least Squares) algorithm is a baseline correction method, and its core is to iteratively optimize the separation baseline and peak signals. The traditional ALS algorithm usually uses a global peak sensitivity parameter (that is, the same sensitivity setting is used for all sampling points), resulting in the baseline may be abnormally elevated near the peak with a high absorbance value, and the weak peaks in the low absorbance region may be covered by noise due to insufficient sensitivity, resulting in the weak peaks not being accurately identified. Therefore, the present invention improves the traditional ALS algorithm, and the improvement content is: calculate the adaptive weights of each data point through the near-peak possibility and smoothness requirement degree of each data point in the chromatogram, so as to dynamically adjust the peak sensitivity of the ALS algorithm, and use the ALS algorithm with adaptive peak sensitivity to perform baseline correction.

[0031] Specifically, the determination of the adaptive weights of each data point in the chromatogram can be achieved through the following steps: Step 1: Take the maximum value points in the chromatogram as peak data points, and calculate the near-peak possibility of each data point. The near-peak possibility represents the proximity of the corresponding data point to the peak data point with low isolation and high absorbance value.

[0032] In an exemplary embodiment of the present invention, the determination of each maximum point in the chromatogram can be achieved through the following steps: Select each data point in the chromatogram that is greater than the adjacent data points before and after, to obtain a number of maximum points.

[0033] Exemplarily, for the th data point in the chromatogram, if the absorbance value of this data point is greater than the absorbance value of the th data point, and at the same time greater than the absorbance value of the th data point, then this data point is taken as a maximum point, so that all the maximum points in the chromatogram can be obtained.

[0034] It should be noted that, in order to avoid isolated noise points being misjudged as peak data points, the present invention utilizes the characteristics that noise has outlier properties both in position and absorbance value, and calculates the isolation value of each peak data point, so as to accurately evaluate the proximity of each data point to the true peak data point with a high absorbance value.

[0035] In an exemplary embodiment of the present invention, the determination of the isolation value of each data point can be achieved through the following steps: (1) Count the number of peak data points within the neighborhood range of each peak data point in the chromatogram, and take the opposite number, and use the opposite number as the power of the exponential function to perform power operation to obtain the first index of each peak data point; Optionally, for the th peak data point, a data segment formed by expanding R data points to both sides with this peak data point as the center, from the th data point to the th data point, can be used as the neighborhood range of this peak data point. Among them, , are functions that return the maximum value and the minimum value respectively; is the total number of data points in the chromatogram. In this embodiment .

[0036] Specifically, the first index of each peak data point satisfies the relational expression: ; in the formula, is the first index of the th peak data point; is the number of peak data points within the neighborhood range of this peak data point; is the natural exponential function, where the natural exponential function is an exponential function with the natural constant e as the base.

[0037] Optionally, when has a large value, it indicates that the There are fewer peak data points within the neighborhood of a peak data point, indicating a higher degree of isolation of the position of the peak data point, and a lower isolation value corresponding to the peak data point.

[0038] (2) Calculate the difference in absorbance values between each peak data point and each data point within the neighborhood, sum and average them, and use the normalized value of the obtained average as the second index for each peak data point; Among them, the second index refers to a parameter that can reflect the degree of isolation of the absorbance value of each peak data point within the neighborhood.

[0039] Specifically, the second index of each peak data point satisfies the following relational expression: ; In the formula, is the second index of the th peak data point; is the absorbance value of the th peak data point; is the absorbance value of the th data point within the neighborhood of the th peak data point; is the preset neighborhood radius, which is in this embodiment; is the absolute value symbol; is the normalization function.

[0040] Among them, the larger the value of, the more inconsistent the absorbance value of the th peak data point is with the surrounding data points, indicating a higher degree of isolation of the absorbance value of the peak data point, and a higher isolation value corresponding to the peak data point.

[0041] (3) Calculate the isolation value of each peak data point according to the first index and the second index. The isolation value is positively correlated with both the second index and the first index.

[0042] Specifically, the isolation value of each peak data point satisfies the relational expression: ; In the formula, is the isolation value of the th peak data point; , are respectively the first index and the second index of the th peak data point.

[0043] Among them, when the isolation value of any peak data point is large, it indicates that the peak data point is likely to be noise. Therefore, when evaluating the near-peak possibility of each data point, the influence of the peak data point needs to be reduced.

[0044] In another embodiment, the isolation values of each peak data point can also be calculated through other calculation formulas. For example, instead of using a calculation formula based on multiplication, a calculation formula based on summation can be used to determine the isolation values.

[0045] Furthermore, after determining the isolation values of each peak data point, the near-peak probability of each data point in the chromatogram can be calculated based on the sampling order difference between each data point and each peak data point in the chromatogram, the absorbance value of each peak data point, and the isolation value.

[0046] Specifically, the near-peak probability of each data point in the chromatogram satisfies the following relational expression: ; In the formula, is the near-peak probability of the th data point in the chromatogram; is the absorbance value of the th peak data point in the chromatogram; is the maximum absorbance value of all peak data points in the chromatogram; is the isolation value of the th peak data point in the chromatogram; is the sampling order difference between the th data point in the chromatogram and the th peak data point in the chromatogram; is the belonging symbol; is the natural exponential function; is the function that returns the maximum value; is the set composed of all peak data points.

[0047] Among them, when the smaller the value, it indicates that the th data point is closer to the th peak data point. At this time, if the isolation value of the th peak data point is smaller and the absorbance value is larger, it indicates that the th data point is closer to the peak with a higher true absorbance value in the chromatogram, and the corresponding near-peak probability of this data point is larger.

[0048] Step 2: Based on the near-peak probability and the fluctuation degree of the absorbance value within the near neighborhood of each data point, calculate the adaptive weight of each data point. The adaptive weight is negatively correlated with the near-peak probability of the corresponding data point and the fluctuation degree of the absorbance value within the near neighborhood of the corresponding data point.

[0049] It should be noted that during the acquisition of chromatographic data, it may be interfered by noise, resulting in random fluctuations caused by noise in the change trend of absorbance values in the chromatogram. This kind of random fluctuation is easily misjudged as a pseudo-peak (i.e., a small peak) during the baseline correction by the ALS algorithm, thus leading to noise residue. Therefore, when determining the adaptive peak sensitivity in the improved ALS algorithm of the present invention, the fluctuation degree of absorbance values within the neighboring range of each data point is also evaluated to determine which data points need to be strengthened in smoothing processing, so as to effectively eliminate noise interference while retaining the true peaks.

[0050] In an exemplary embodiment of the present invention, the fluctuation degree of absorbance values within the neighboring range of each data point is the relative fluctuation degree, and the determination of the relative fluctuation degree of absorbance values within any neighboring range can be achieved through the following steps: (1) Select a target neighboring range, calculate the difference in absorbance values between each data point within the target neighboring range and its adjacent data points before and after, and sum them. Then, perform a summation operation on the obtained cumulative sums again to obtain the local fluctuation value of absorbance values within the target neighboring range; Among them, the target neighboring range refers to the neighboring range of a randomly selected data point.

[0051] Specifically, the local fluctuation value of absorbance values within the target neighboring range satisfies the relational expression: ; in the formula, , and are respectively the absorbance values of the th data point, the th data point, and the th data point; is the absolute value symbol; is the preset neighboring radius; is the summation symbol.

[0052] In another embodiment, the local fluctuation value of absorbance values within the neighboring range of each data point can also be calculated through the relational expression:

[0053] It should be noted that the true peaks in the chromatogram usually have continuous and gradually changing absorbance changes (such as rising first and then falling), while the noise shows random and abrupt fluctuations. By analyzing the change in absorbance values (such as difference or ratio) between adjacent points before and after each data point, the true signal and noise can be more accurately distinguished.

[0054] (2) Calculate the difference in absorbance values between the data points at both ends within the target neighboring range to obtain the overall fluctuation value of absorbance values within the target neighboring range, and perform normalization processing on the ratio of the local fluctuation value to the overall fluctuation value to obtain the relative fluctuation degree of absorbance values within the target neighboring range.

[0055] In an exemplary embodiment of the present invention, the relative fluctuation degree of the absorbance values within each neighboring range is used as the smoothing requirement degree of the corresponding data point. Specifically, the smoothing requirement degree of each data point satisfies the following relational expression: ; In the formula, is the smoothing requirement degree of the th data point in the chromatogram; is the preset neighboring radius; , and are respectively the absorbance values of the th data point, the th data point, and the th data point; , are the absorbance values of the data points at both ends within the neighboring range of the th data point in the chromatogram; is a preset hyperparameter, and this value is an empirical value used to avoid the denominator being zero; is a normalization function; is the absolute value symbol.

[0056] Among them, reflects the overall change trend of the absorbance values within the neighboring range of the th data point. The smaller this value is, the flatter the change in this neighboring range, which may be the noise or baseline region. At this time, if the local fluctuation is large, significant smoothing is required to remove noise, and the corresponding smoothing requirement degree is large; if this value is large, it indicates that this neighboring range may be the rising or falling edge of a real peak. At this time, if the local fluctuation is small, the smoothing requirement of this data point is small, and the corresponding smoothing requirement degree is small.

[0057] Furthermore, after determining the near-peak possibility and smoothing requirement degree of each data point, the adaptive weight of each data point can be calculated. Specifically, the adaptive weight of each data point satisfies the following relational expression: ; In the formula, is the adaptive weight of the th data point in the chromatogram; is the near-peak possibility of the th data point in the chromatogram; is the smoothing requirement degree of the th data point in the chromatogram; is the hyperbolic tangent function.

[0058] Among them, when or is larger, it indicates that the The more likely a data point is to be in the true peak region or the noise interference region, at this time, setting a smaller adaptive weight can reduce the peak sensitivity during baseline correction by ALS, avoid excessive elevation of the baseline in the true peak region, and fit the baseline to the noise.

[0059] In another exemplary embodiment, it can also be through the relational expression: ; Calculate the adaptive weight of each data point. In the formula, is the normalization function.

[0060] In an exemplary embodiment of the present invention, the determination of the adaptive peak sensitivity can be achieved through the following steps: Obtain the preset peak sensitivity when the ALS algorithm performs baseline correction on the chromatogram, and perform a multiplication operation on the adaptive weight of each data point and the preset peak sensitivity to obtain the adaptive peak sensitivity.

[0061] Specifically, the adaptive peak sensitivity satisfies the following relational expression: ; In the formula, is the adaptive peak sensitivity of the th data point in the chromatogram; is the adaptive weight of the th data point in the chromatogram; is the preset peak sensitivity, in this embodiment .

[0062] In an exemplary embodiment of the present invention, the determination of performing baseline correction based on the adaptive peak sensitivity can be achieved through the following steps: Take the adaptive peak sensitivity as the sample weight corresponding to the data point, and perform low-pass filtering on the chromatographic data in the chromatogram, and use the obtained filtering result as the initial baseline in the ALS algorithm; based on the sample weight and the initial baseline, construct an objective function when the ALS algorithm performs baseline correction, and update the initial baseline with the goal of minimizing the objective function, and obtain the corrected baseline after the update operation.

[0063] Among them, the sample weight reflects the influence degree of the data point on the baseline fitting. Both the sample weight and the initial baseline are professional terms in the ALS algorithm, and are not elaborated in detail in this embodiment.

[0064] Specifically, the constructed objective function satisfies the following relational expression: ; In the formula, is the objective function value; is the sample weight of the th data point in the chromatogram; is the absorbance value of the th data point in the chromatogram; is the baseline intensity of the th data point in the initial baseline; is the total number of data points in the chromatogram; is a preset regularization parameter, and this value is an empirical value. In this embodiment, ; is the symbol for the second-order difference representation.

[0065] Among them, reflects the weighted residual between the initial baseline and the absorbance value of the data points in the chromatogram; is the regularization term, which is used to constrain the second-order derivative smoothing of the baseline.

[0066] It should be noted that when the traditional ALS algorithm uses a fixed peak sensitivity for baseline correction, due to the high signal intensity in the strong peak region, the residual is relatively large, and the algorithm will erroneously raise the baseline to reduce the residual. In the present invention, the closer the data point is to the true strong peak region, the smaller the corresponding adaptive peak sensitivity. At this time, the sample weight tends to zero, so that the algorithm can ignore the residual in the strong peak region and avoid the abnormal rise of the baseline in the strong peak region.

[0067] Moreover, using a fixed peak sensitivity will also perform baseline fitting on the noise interference region, resulting in the objective function paying too much attention to the fitting residual in the noise interference region and causing noise residue. In the present invention, when the data point is in the noise interference region, the corresponding adaptive peak sensitivity is small. At this time, the sample weight tends to zero, so that the objective function does not pay attention to, or pays less attention to, the fitting residual in the noise interference region, enabling the regularization term to be used to force the smoothing of the noise interference region and avoid noise residue.

[0068] Optionally, the objective function can be solved by the least squares method to update the initial baseline. When the difference between the baselines of two adjacent times is less than a preset threshold, such as , the iteration is terminated and the corrected baseline is output.

[0069] It should be noted that the process of constructing the objective function of the ALS algorithm for baseline correction based on the sample weight and the initial baseline is a prior art, and this embodiment will not be described in detail here.

[0070] S3: Based on the chromatogram after baseline correction, detect the cordycepin concentration in the cordyceps militaris extract.

[0071] Among them, the chromatographic data in the chromatogram after baseline correction is the value obtained by subtracting the corrected baseline point by point from the chromatographic data before baseline correction.

[0072] Further, after obtaining the chromatogram with the corrected baseline, the precise retention time window of the cordycepin characteristic peak in the cordyceps militaris chromatographic data can be determined by using the pre-stored cordycepin standard chromatogram. Then, the peak area of the target chromatographic peak within the corresponding retention time range of the cordyceps militaris chromatographic data after baseline correction is integrated, and quantitative analysis is performed based on the pre-established cordycepin chromatographic peak area-concentration standard curve, so as to obtain the actual concentration of cordycepin in the cordyceps militaris extract.

[0073] In the description of this specification, the meanings of "a plurality of" and "several" are at least two, such as two, three or more, etc., unless otherwise specifically and clearly defined.

[0074] Although this specification has shown and described multiple embodiments of the present invention, it is obvious to those skilled in the art that such embodiments are provided only by way of example. Those skilled in the art will think of many changes, alterations and alternative ways without departing from the spirit and idea of the present invention. It should be understood that various alternative solutions to the embodiments of the present invention described herein can be adopted in the process of practicing the present invention.

Claims

1. A data processing method for detecting component content, characterized in that, Including: Using high performance liquid chromatography to obtain a chromatogram of the Cordyceps militaris extract, where the abscissa of the chromatogram is the retention time and the ordinate is the absorbance value; Using the ALS algorithm to perform baseline correction on the chromatogram. During the correction process, based on the adaptive weights of each data point in the chromatogram, adjust the peak sensitivity of the ALS algorithm and perform baseline correction to detect the cordycepin concentration in the Cordyceps militaris extract based on the chromatogram after baseline correction; Among them, the method for obtaining the adaptive weights of each data point includes: taking the maximum value points in the chromatogram as peak data points, and calculating the near-peak possibility of each data point. The near-peak possibility represents the proximity of the corresponding data point to a peak data point with low isolation and high absorbance value; Based on the near-peak possibility and the fluctuation degree of the absorbance values within the near-neighbor range of each data point, calculate the adaptive weights of each data point. The adaptive weights are negatively correlated with the near-peak possibility of the corresponding data point and the fluctuation degree of the absorbance values within the near-neighbor range of the corresponding data point.

2. The data processing method for detecting component content according to claim 1, wherein, The calculation of the near-peak possibility of each data point satisfies the following relational expression: ; In the formula, is the near-peak possibility of the th data point in the chromatogram; is the absorbance value of the th peak data point in the chromatogram; is the maximum absorbance value of all peak data points in the chromatogram; is the isolation value of the th peak data point in the chromatogram. The isolation value represents the degree of isolation of the position and absorbance value of the peak data point within the neighboring range; is the sampling order difference between the th data point in the chromatogram and the th peak data point in the chromatogram; is the membership symbol; is the natural exponential function; is the function that returns the maximum value; is the set composed of all peak data points.

3. A data processing method for ingredient content detection according to claim 2, characterized in that, The method for obtaining the isolation value includes: Count the number of peak data points within the near-neighbor range of each peak data point in the chromatogram, take the opposite number, and use the opposite number as the power of the exponential function for power operation to obtain the first index of each peak data point; Calculate the difference in absorbance values between each peak data point and each data point within the near-neighbor range, sum and take the average, and use the normalized value of the obtained average as the second index of each peak data point; According to the first index and the second index, calculate the isolation value of each peak data point. The isolation value is positively correlated with the second index and the first index.

4. A data processing method for detecting component content according to claim 1, characterized in that, The fluctuation degree is the relative fluctuation degree. The method for obtaining the relative fluctuation degree includes: Select a target near-neighbor range, calculate the difference in absorbance values between each data point within the target near-neighbor range and the adjacent data points before and after, sum them, and perform a summation operation on the obtained cumulative sums again to obtain the local fluctuation value of the absorbance values within the target near-neighbor range; Calculate the difference in absorbance values of the data points at both ends within the target near-neighbor range to obtain the overall fluctuation value of the absorbance values within the target near-neighbor range, and perform a normalization process on the ratio of the local fluctuation value to the overall fluctuation value to obtain the relative fluctuation degree of the absorbance values within the target near-neighbor range.

5. A data processing method for detecting component content according to claim 4, characterized in that, Use the relative fluctuation degree as the smoothing requirement degree of the corresponding data point. The smoothing requirement degree satisfies the following relational expression: ; Wherein, is the smoothing requirement degree of the th data point in the chromatogram; is the preset neighbor radius; , and are the absorbance values of the th data point, the th data point, and the th data point respectively; , are the absorbance values of the data points at both ends within the neighbor range of the th data point in the chromatogram; is the preset hyperparameter; is the normalization function; is the absolute value symbol.

6. The data processing method for detecting component content according to claim 2 or 5, characterized in that, The adaptive weights of each data point satisfy the following relational expression: In the formula, is the adaptive weight of the th data point in the chromatogram; is the near-peak possibility of the th data point in the chromatogram; is the smoothing requirement degree of the th data point in the chromatogram; is the hyperbolic tangent function.

7. A data processing method for detecting component content according to claim 1, characterized in that, The adjustment of the peak sensitivity of the ALS algorithm based on the adaptive weights of several data points in the chromatogram includes: Obtain the preset peak sensitivity when the ALS algorithm performs baseline correction on the chromatogram, and perform a multiplication operation on the adaptive weights of each data point and the preset peak sensitivity to obtain the adaptive peak sensitivity.

8. A data processing method for detecting component content according to claim 7, characterized in that, The process of performing baseline correction based on the adaptive peak sensitivity includes: Take the adaptive peak sensitivity as the sample weight of the corresponding data point, and perform low-pass filtering on the chromatographic data in the chromatogram to use the obtained filtering result as the initial baseline in the ALS algorithm; Based on the sample weight and the initial baseline, construct an objective function for baseline correction by the ALS algorithm, and update the initial baseline with the goal of minimizing the objective function. After the update operation, the corrected baseline is obtained.

9. A data processing method for detecting component content according to claim 8, characterized in that The objective function satisfies the following relational expression: ; In the formula, is the objective function value; is the sample weight of the -th data point in the chromatogram; is the absorbance value of the -th data point in the chromatogram; is the baseline intensity of the -th data point in the initial baseline; is the total number of data points in the chromatogram; is the preset regularization parameter; is the symbol for the second-order difference representation.

10. A data processing method for detecting component content according to claim 1, characterized in that, The method for obtaining the maximum value points among all the data points includes: Screen each data point greater than the adjacent data points before and after from the chromatogram to obtain several maximum value points.

Citation Information

Patent Citations

  • Method for investigating quality of characterization natural cordyceps sinensis character

    CN101000331A

  • Method for analyzing content of adenosine and cordycepin in cordyceps militaris by virtue of high performance liquid chromatography (HPLC)

    CN102928526A

  • Method for determining cordycepin content in cordyceps sinensis by virtue of high performance liquid chromatography

    CN110470754A

  • Hydrogen peak baseline estimation method in carbon-oxygen ratio energy spectrum

    CN114722339A

  • Full-automatic spectrum baseline correction method for spectrum

    CN117007184A

Cited By

  • Hydroxytyrosol content detection method and system for food quality inspection

    CN121540829A

  • Method for detecting active ingredients of spleen-tonifying traditional Chinese medicine based on high performance liquid chromatography

    CN121577816A