A spectral data classification method for ultraviolet analyzer
By calculating the local relative fluctuation degree and shape matching degree and dynamically adjusting the moving average window size, the problem of improper window size selection in traditional baseline correction methods is solved, more accurate baseline correction and characteristic peak identification are achieved, and the accuracy of spectral data classification is improved.
Patent Information
- Application Number
- CN202510845920.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-06-24
AI Technical Summary
When the baseline changes rapidly, the traditional moving average method will cause inaccurate baseline correction due to the window size selection, which will affect the accuracy of the absorbance value of the spectral data, and further lead to inaccurate characteristic peak selection and reduced spectral data classification accuracy.
By calculating the local relative fluctuation degree of each wavelength point, the high fluctuation area is screened out and the peak area is identified. Combining the characteristics of baseline drift and real gas absorption peak, the ideal gas absorption peak model and shape matching are used to dynamically adjust the moving average window size for baseline correction.
It significantly improves the accuracy of the absorbance values of spectral data, enhances the accuracy of characteristic peak identification, and improves the efficiency and reliability of spectral data analysis.
Smart Images

Figure CN120354247B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of spectrum analysis technology, and in particular to a spectrum data classification method for an ultraviolet analyzer. Background Art
[0002] Ultraviolet (UV) analyzers are used in chemical analysis, environmental monitoring, biomedicine, and other fields to detect and analyze the absorption spectra of samples. Baseline correction, as a core step in the preprocessing phase, plays a vital role. However, spectral data often suffers from problems such as baseline drift and noise interference, which can affect the accuracy of spectral data and the reliability of subsequent analysis. Traditional baseline correction and feature extraction methods have limitations. For example, the moving average method can smooth out important spectral details, resulting in the misidentification or omission of true peaks.
[0003] One common baseline correction method is the moving average method. This method smoothes the spectral curve by applying a moving average to the spectral data, thereby removing high-frequency noise and baseline fluctuations. However, the moving average method has significant limitations. The selection of the moving average window size significantly affects the correction results, especially in cases of rapidly varying baselines. While a larger window size can more effectively smooth baseline fluctuations, it also smooths out detailed spectral data, potentially leading to inaccurate baseline correction.
[0004] Inaccurate baseline correction can lead to deviations in the absorbance values of spectral data. This deviation directly affects subsequent characteristic peak selection, as accurate identification of characteristic peaks relies on precise absorbance data. Furthermore, this deviation can reduce the accuracy of spectral data classification and affect the final analysis results. Summary of the Invention
[0005] In order to solve the technical problems existing in the existing moving average method for baseline correction, namely, when the baseline changes rapidly, the moving average method causes inaccurate baseline correction due to the window size selection, which in turn affects the accuracy of the absorbance values of the spectral data, resulting in inaccurate characteristic peak selection and reduced spectral data classification accuracy, the present invention provides solutions in the following aspects.
[0006] A spectral data classification method for an ultraviolet analyzer comprises: acquiring spectral data and performing preprocessing; analyzing the preprocessed spectral data, calculating the local relative fluctuation degree of each wavelength point, calculating the second-order difference of the mean of the light intensity in the local window of the wavelength point and the mean of the adjacent windows, performing normalization processing, obtaining the probability of each wavelength point belonging to a high fluctuation region, and screening out the wavelength points belonging to the high fluctuation region and the peak regions corresponding to the wavelength points based on the probability and a set probability threshold; distinguishing based on the different characteristics of baseline drift and real gas absorption peaks, analyzing the characteristics of the high fluctuation region, and identifying the probabilities of peak regions belonging to baseline fluctuations and peak regions that may be real peaks; measuring the shape matching degree based on the Pearson correlation coefficient between the ideal gas absorption peak model and the normalized peak vector of the corresponding peak region, and calculating the probability that the peak region belongs to a real peak in combination with the edge steepness of the peak region, and judging whether the peak region is a real peak; dynamically adjusting the window size of the moving average based on the probability that the peak region belongs to a real peak to achieve more accurate baseline correction and perform visual display.
[0007] By acquiring spectral data and preprocessing it, the local relative fluctuation degree of each wavelength point is calculated, and the wavelength points belonging to the high fluctuation area and their corresponding peak areas are screened out. Further, based on the different characteristics of baseline drift and real gas absorption peaks, a distinction is made, and the probability of peak areas belonging to baseline fluctuations and possible real peaks is identified. By combining shape matching and edge steepness, the probability of a peak area belonging to a real peak is calculated, and the window size of the moving average is dynamically adjusted accordingly to achieve more accurate baseline correction. Ultimately, by visually displaying an efficient and accurate spectral data classification method, the accuracy of the spectral data absorbance values is significantly improved, the accuracy of characteristic peak identification is enhanced, and the efficiency and reliability of spectral data analysis are improved.
[0008] Preferably, the local relative fluctuation degree includes:
[0009] Perform first-order difference on the spectral data, and take the average of the absolute values of the first-order difference results of all wavelength points as the indicator of the overall spectral baseline change rate;
[0010] Take any wavelength point as the target wavelength point, take the wavelength points in the data range on both sides of the target wavelength point as the local window of the target wavelength point, calculate the mean value of the light intensity in the local window of the target wavelength point, and calculate the mean value of the light intensity in the local window corresponding to the adjacent wavelength points on both sides of the target wavelength point respectively;
[0011] Calculate the second-order difference between the mean of the light intensity in the local window of the target wavelength point and the mean of the adjacent windows. The absolute value of the second-order difference result is divided by the standard deviation of the window corresponding to the target wavelength point for normalization to obtain the local relative fluctuation degree of the target wavelength point.
[0012] By calculating the first-order difference of the spectral data, an indicator of the rate of change of the overall spectral baseline is obtained, thereby quantifying the overall fluctuation of the spectral data. Furthermore, with each wavelength point as the center, the mean of the light intensity within its local window and the second-order difference of the means of adjacent windows are calculated, and the local relative fluctuation degree of each wavelength point is obtained through normalization. By effectively identifying high-fluctuation regions in the spectral data, it provides an important basis for subsequent peak region screening and baseline correction, significantly improving the accuracy and reliability of spectral data analysis.
[0013] Preferably, the probability that the wavelength point belongs to the high fluctuation area includes:
[0014] Taking any wavelength point as the target wavelength point, calculate the sum of the preset basic sensitivity, global calibration coefficient, and overall spectral baseline change rate indicators respectively, calculate the ratio between the local relative fluctuation degree of the target wavelength point and the sum, and multiply it by the scaling factor to obtain the weighted fluctuation index;
[0015] Use a negative exponential function to exponentially decay the weighted volatility index, add 1 to the attenuated result and take the inverse to obtain the probability that the target wavelength point belongs to the high volatility area.
[0016] By calculating the weighted fluctuation index of the target wavelength and performing exponential decay using a negative exponential function, the probability that the target wavelength belongs to the high-fluctuation region is ultimately determined. This method not only comprehensively considers the effects of local relative fluctuation, basic sensitivity, global calibration coefficient, and the rate of change of the overall spectral baseline, but also more accurately assesses the fluctuation significance of each wavelength point, effectively distinguishing between high-fluctuation and non-high-fluctuation regions, significantly improving the accuracy and reliability of spectral data analysis.
[0017] Preferably, the step of screening out wavelength points belonging to a high fluctuation region includes:
[0018] If the probability is greater than a preset probability threshold, the wavelength point is considered to belong to the high fluctuation area; otherwise, if the probability is less than or equal to the preset probability threshold, the wavelength point is considered not to belong to the high fluctuation area.
[0019] Preferably, obtaining the peak region comprises the steps of:
[0020] The density clustering algorithm is used to cluster the wavelength points belonging to the high fluctuation area to identify the high fluctuation area, where the identification conditions are: wavelength points Can be with wavelength point The conditions for clustering into one category are ,in, is the spectral resolution;
[0021] The connected domain analysis algorithm is used to connect adjacent high-fluctuation regions to merge adjacent high-fluctuation regions, and the merged high-fluctuation regions are further divided into independent peak regions.
[0022] By using a density clustering algorithm to cluster wavelength points belonging to high-fluctuation regions, high-fluctuation regions are identified based on the distance between the wavelength points. A connected domain analysis algorithm is then used to connect adjacent high-fluctuation regions and then divide the merged high-fluctuation regions into independent peak regions. This method effectively identifies and merges high-fluctuation regions, ensuring the integrity and independence of the peak regions, thereby significantly improving the accuracy and efficiency of spectral data analysis.
[0023] Preferably, the probability of the peak region possibly being a true peak includes:
[0024] The Gaussian function is used as the ideal gas absorption peak model, which is expressed as ,in, represents the wavelength, represents the central wavelength of the peak, Indicates the width of the peak;
[0025] Peak shape area The spectral data within are normalized to obtain the normalized absorption intensity , the normalized absorption intensity With the ideal Gaussian shape The Pearson correlation coefficient between them is used as the shape matching degree;
[0026] The shape matching degree is added by 1 and then multiplied by the weight factor to obtain the normalized shape matching degree. The average value of the normalized shape matching degree and the edge steepness of the peak region is taken as the probability that the peak region belongs to a true peak.
[0027] By using a Gaussian function as the ideal gas absorption peak model, the spectral data within the peak region is normalized and the Pearson correlation coefficient between the normalized absorption intensity and the ideal Gaussian shape is calculated to obtain the shape matching degree. The shape matching degree is then increased by 1 and multiplied by a weighting factor to obtain the normalized shape matching degree. This is then combined with the edge steepness of the peak region to calculate the probability that the peak region is a true peak. This method can comprehensively consider the shape characteristics and edge steepness of the peak region to more accurately identify true peaks and baseline fluctuations, avoiding misidentification, thereby significantly improving the accuracy of characteristic peak selection and providing a more reliable basis for subsequent spectral data classification.
[0028] Preferably, the edge steepness of the peak region includes:
[0029] Taking any peak shape area as the target peak shape area, the radian value of the angle between the peak top in the target peak shape area and the starting point of the peak shape area is calculated to obtain the starting point angle, the radian value of the angle between the peak top in the target peak shape area and the ending point of the peak shape area is calculated to obtain the ending point angle, the starting point angle and the ending point angle are normalized respectively, the normalized angle values are added and then normalized to obtain the edge steepness of the target peak shape area.
[0030] Preferably, the determining whether the peak region is a true peak comprises:
[0031] If the probability that the peak region belongs to a true peak is greater than a preset probability threshold, the peak region is considered to belong to a true peak. Otherwise, if it is less than or equal to the preset probability threshold, the peak region is considered not to belong to a true peak.
[0032] Preferably, the dynamically adjusting the moving average window size comprises the steps of:
[0033] Taking any peak shape region as the target peak shape region, calculating the variation range of the preset maximum window size and the preset minimum window size in the target peak shape region;
[0034] Calculate the probability that the target peak region does not belong to the true peak, divide the probability that it does not belong to the true peak by the preset probability threshold, and obtain the normalized result;
[0035] The product of the variation range and the normalized result is added to the preset minimum window size to obtain the corrected window size of the target peak area.
[0036] Preferably, the visual display comprises the steps of:
[0037] The pre-processed and baseline-corrected spectral data is displayed graphically, with the horizontal axis representing wavelength and the vertical axis representing absorbance. The automatically identified peak regions are marked on the spectral data graph, and different colors or markers are used to distinguish between true peaks and non-true peaks.
[0038] Based on manual inspection, the automatically identified peak shape is checked to see if it conforms to the expected gas absorption peak characteristics. For the confirmed characteristic peaks, the wavelength, absorbance, half-peak width and area parameters are recorded and used as the feature vector for subsequent classification;
[0039] The recorded characteristic peak parameters are compared with the ultraviolet absorption spectrum database of known gas molecules to establish a characteristic peak database for subsequent qualitative analysis and multi-component classification.
[0040] The present invention has the following effects:
[0041] 1. The present invention optimizes the baseline correction process according to the probability that the peak area belongs to the true peak by dynamically adjusting the window size of the moving average, which can significantly improve the accuracy of the absorbance value of the spectral data. It not only improves the accuracy of the baseline correction, but also enhances the accuracy of the characteristic peak identification, thereby improving the efficiency and reliability of the spectral data analysis.
[0042] 2. This method calculates the probability that a peak region is a true peak by combining shape matching and edge steepness, thereby more accurately identifying true peaks from baseline fluctuations. This method considers not only the shape characteristics of the peak region but also the steepness of the edge, effectively distinguishing true peaks from baseline fluctuations and avoiding misidentification. This significantly improves the accuracy of characteristic peak selection and, consequently, the precision of spectral data classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 This is a method flow chart of steps S1 to S5 in a spectral data classification method for an ultraviolet analyzer according to an embodiment of the present invention. DETAILED DESCRIPTION
[0044] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, but not all of the embodiments.
[0045] Reference Figure 1 A spectrum data classification method for an ultraviolet analyzer includes steps S1 to S5, specifically as follows:
[0046] S1: Acquire spectral data and perform preprocessing.
[0047] The sample is irradiated with an ultraviolet light source. The substances in the sample absorb the ultraviolet light energy and undergo electronic transitions. Photons of specific wavelengths are then absorbed, forming an absorption peak. The intensity of the transmitted light is then recorded using an array CCD detector to generate two-dimensional spectral data, in which one dimension is wavelength and the other is light intensity.
[0048] Wavelet transform or Kalman filtering are used to remove electronic noise and other interference signals from the spectral data. The above algorithms are all existing technologies and will not be described in detail here.
[0049] To further illustrate, the local fluctuations in spectral data are used to distinguish between baseline fluctuations and true peaks, thereby identifying regions of high fluctuation and determining whether these regions are true peaks. Baseline drift typically manifests as a low-frequency, wide-band, slowly varying signal, while true gas absorption peaks are characterized by high frequencies, narrow bandwidths, and sharp transitions (consistent with the narrow-band absorption characteristics of the Lambert-Beer law). By quantifying the local fluctuations to identify regions of high variation, the corresponding peak regions are then verified as true peaks using shape matching and edge steepness, ensuring that true peaks are not smoothed out.
[0050] S2: Analyze the preprocessed spectral data, calculate the local relative fluctuation degree of each wavelength point, calculate the second-order difference between the mean of the light intensity in the local window of the wavelength point and the mean of the adjacent window, and normalize it to obtain the probability that each wavelength point belongs to the high fluctuation area, and filter out the wavelength points belonging to the high fluctuation area and the peak area corresponding to the wavelength point based on the probability and the set probability threshold.
[0051] Local relative volatility, including:
[0052] Perform first-order difference on the spectral data, and take the average of the absolute values of the first-order difference results of all wavelength points as the indicator of the overall spectral baseline change rate;
[0053] Specifically, the index of the overall spectral baseline change rate satisfies the following relationship:
[0054] ;
[0055] Where, An indicator of the rate of change of the overall spectral baseline, Indicates spectral data at wavelength The change in light intensity at Indicates the total number of spectral data points.
[0056] In other words, a larger index of the rate of change of the overall spectral baseline reflects a faster overall change of the baseline, and vice versa, it indicates a slower overall change of the baseline.
[0057] Take any wavelength point as the target wavelength point, take the wavelength points in the data range on both sides of the target wavelength point as the local window of the target wavelength point, calculate the mean value of the light intensity in the local window of the target wavelength point, and calculate the mean value of the light intensity in the local window corresponding to the adjacent wavelength points on both sides of the target wavelength point respectively;
[0058] Calculate the second-order difference between the mean of the light intensity in the local window of the target wavelength point and the mean of the adjacent windows. The absolute value of the second-order difference result is divided by the standard deviation of the window corresponding to the target wavelength point for normalization to obtain the local relative fluctuation degree of the target wavelength point.
[0059] Specifically, the local relative fluctuation degree satisfies the following relationship:
[0060] ;
[0061] Where, Indicates wavelength point The local relative fluctuation degree, Indicated by wavelength point The mean light intensity in the local window centered at Indicated by wavelength point The mean light intensity in the local window centered at Indicated by wavelength point The mean light intensity in the local window centered at Indicated by wavelength point The standard deviation of the light intensity within the local window centered at .
[0062] That is to say, by dividing by the standard deviation of the window corresponding to the target wavelength point, it is used for normalization to make the result dimensionless. By calculating the second-order difference between the mean of the current window and the mean of the adjacent window, it reflects the acceleration of the mean change. The greater the average difference between the average value of the corresponding light intensity in the local window of the current window and the average value of the corresponding light intensity in the corresponding local windows of the adjacent wavelength points, if the difference between the mean value of the current window and the mean value of the adjacent window is large, it indicates that there is obvious fluctuation at that position (possibly a peak or baseline mutation).
[0063] The probability that a wavelength point belongs to the high fluctuation area includes:
[0064] Taking any wavelength point as the target wavelength point, calculate the sum of the preset basic sensitivity, global calibration coefficient, and overall spectral baseline change rate indicators respectively, calculate the ratio between the local relative fluctuation degree of the target wavelength point and the sum, and multiply it by the scaling factor to obtain the weighted fluctuation index;
[0065] Use a negative exponential function to exponentially decay the weighted volatility index, add 1 to the attenuated result and take the inverse to obtain the probability that the target wavelength point belongs to the high volatility area.
[0066] Specifically, the probability that a wavelength point belongs to the high fluctuation region satisfies the following relationship:
[0067] ;
[0068] Where, is the wavelength point The probability of belonging to the high volatility area, Indicates wavelength point The local relative fluctuation degree, An indicator of the rate of change of the overall spectral baseline, represents the scaling factor, Indicates the basic sensitivity, Represents the global calibration coefficient.
[0069] That is, when the local volatility The larger the value, the more drastic the fluctuation of the spectrum data. The greater the probability of belonging to the high volatility area. When the overall change is faster, that is, The larger the denominator, The larger the wavelength, the This is because when the overall change is fast, local fluctuations are more likely to be caused by baseline drift rather than characteristic peaks, and vice versa. It is the sigmoid function, which indicates normalization.
[0070] Filter out wavelength points belonging to the high fluctuation area, including:
[0071] If the probability is greater than a preset probability threshold, the wavelength point is considered to belong to the high fluctuation area; otherwise, if the probability is less than or equal to the preset probability threshold, the wavelength point is considered not to belong to the high fluctuation area.
[0072] Exemplarily, the preset probability threshold is 0.85.
[0073] To further illustrate, the high-fluctuation regions in the spectral data are effectively identified and divided to further analyze the true peaks. Specifically, through the density clustering algorithm and the connected domain analysis algorithm, the wavelength points with significant local fluctuations are aggregated into high-fluctuation regions, and independent peak regions are further extracted to provide clear structured data for subsequent characteristic peak verification and classification. The specific steps are as follows:
[0074] The density clustering algorithm is used to cluster the wavelength points belonging to the high fluctuation area to identify the high fluctuation area, where the identification conditions are: wavelength points Can be with wavelength point The conditions for clustering into one category are ,in, is the spectral resolution, exemplarily, .
[0075] In this way, wavelength points with significant local fluctuations are clustered into several high-fluctuation regions, which may contain real absorption peaks or baseline fluctuations.
[0076] The connected domain analysis algorithm is used to connect adjacent high-fluctuation areas to merge adjacent high-fluctuation areas; the merged high-fluctuation areas are further divided into independent peak areas, and each peak area is Perform analysis to verify whether it is a real peak.
[0077] Through connected domain analysis, adjacent high-fluctuation areas are merged into larger areas, thereby reducing fragmented high-fluctuation areas and improving the efficiency and accuracy of subsequent analysis.
[0078] Any peak area is recorded as ,in, Indicates peak shape area The starting wavelength point, Indicates peak shape area The end wavelength point of any peak area Conduct analysis.
[0079] In this way, each peak region is clearly defined, facilitating subsequent peak verification and analysis. Characteristics such as shape matching and edge steepness are used to determine whether the peak region conforms to the characteristics of a true peak (e.g., high frequency, narrow bandwidth, abrupt transitions, etc.). This ensures that the extracted peak region contains true peaks, not baseline fluctuations or other noise signals.
[0080] S3: Based on the different characteristics of baseline drift and real gas absorption peaks, the characteristics of high fluctuation areas are analyzed to identify the probability of peak areas belonging to baseline fluctuations and peak areas that may be real peaks.
[0081] It should be noted that baseline drift refers to the phenomenon in which the baseline (i.e., background signal or noise) in spectral data slowly changes with wavelength or time. This change is typically low-frequency, broad, and relatively gradual. True gas absorption peaks are peaks in spectral data formed by the absorption of light of specific wavelengths by gas molecules. These peaks typically have specific shapes and positions, consistent with physical and chemical principles.
[0082] Distinguishing baseline drift from true gas absorption peaks is crucial in spectral data analysis. Baseline drift typically manifests as a low-frequency, broad, and slowly varying signal, while true gas absorption peaks appear as high-frequency, narrow-band, and abruptly varying signals. By analyzing these characteristics, it is possible to effectively distinguish baseline drift from true gas absorption peaks, thereby improving the accuracy of spectral data analysis.
[0083] The Gaussian function is used as the ideal gas absorption peak model, which is expressed as ,in, represents the wavelength, represents the central wavelength of the peak, Indicates the width of the peak;
[0084] Peak shape area The spectral data within are normalized to obtain the normalized absorption intensity , the normalized absorption intensity With the ideal Gaussian shape The Pearson correlation coefficient between ;
[0085] The shape matching degree is added by 1 and then multiplied by the weight factor to obtain the normalized shape matching degree. The average value of the normalized shape matching degree and the edge steepness of the peak region is taken as the probability that the peak region belongs to a true peak.
[0086] Specifically, the probability of belonging to a true peak satisfies the following relationship:
[0087] ;
[0088] Where, Indicates peak shape area The probability of belonging to a true peak, Indicates peak shape area The shape matching degree, Indicates peak shape area The edge steepness, Represents the weight factor.
[0089] For example, the weight factor is 0.5, which can be adjusted according to specific circumstances; The value of between, so The value range becomes , multiplied by the weight factor so that the range becomes between.
[0090] Normalized absorption intensity With the ideal Gaussian shape The Pearson correlation coefficient between them is a well-known technology and will not be described in detail here.
[0091] Edge steepness of the peak shape area, including:
[0092] Taking any peak shape area as the target peak shape area, the radian value of the angle between the peak top in the target peak shape area and the starting point of the peak shape area is calculated to obtain the starting point angle, the radian value of the angle between the peak top in the target peak shape area and the ending point of the peak shape area is calculated to obtain the ending point angle, the starting point angle and the ending point angle are normalized respectively, the normalized angle values are added and then normalized to obtain the edge steepness of the target peak shape area.
[0093] Specifically, the edge steepness satisfies the following relationship:
[0094] ;
[0095] Where, Indicates peak shape area The edge steepness, Indicates the horizontal distance between the peak top and the starting point of the peak area, represents the Euclidean distance between the peak top and the starting point of the peak area, Indicates the horizontal distance between the peak top and the end point of the peak shape area, represents the Euclidean distance between the peak top and the end point of the peak region, represents the inverse cosine function, Represents the normalization function.
[0096] It should be noted that when the steepness angle is 90°, the edge of the peak region is very steep, nearly vertical. In this case, the edge steepness value is close to its maximum value, indicating that the edge of the peak region has a very steep edge. This highly steep edge usually indicates that the spectral feature of the peak region is very significant and is likely a true absorption or emission peak, rather than a false peak caused by baseline drift or noise.
[0097] In spectral analysis, true peaks typically have steep edges, while baseline fluctuations are relatively gentle. Therefore, when the steepness angle is 90°, the peak region is more likely to be a true characteristic peak. This highly steep edge can serve as a key characteristic of a characteristic peak, useful for subsequent qualitative analysis and multi-component classification.
[0098] In addition to using the inverse cosine function to calculate angles, several other techniques can be used to assess the edge steepness of a peak region: First, by calculating the first-order derivative of the spectral data in the peak region, the slope of the spectral curve can be determined. A larger absolute value of the slope indicates a steeper edge. Second, calculating the second-order derivative of the spectral data can reflect the curvature of the spectral curve. A larger absolute value of the curvature also indicates a steeper edge. Finally, a gradient operator (such as the Sobel or Prewitt operator) can be used to calculate the edge gradient of the peak region. A larger modulus of the gradient indicates a steeper edge. Each of these methods has its own advantages, and the appropriate technique for assessing the edge steepness of a peak region can be selected based on the specific application scenario and data characteristics.
[0099] S4: The shape matching degree is measured based on the Pearson correlation coefficient between the ideal gas absorption peak model and the normalized peak shape vector of the corresponding peak shape area. Combined with the edge steepness of the peak shape area, the probability that the peak shape area belongs to a true peak is calculated, and whether the peak shape area is a true peak is determined.
[0100] It should be noted that the ideal gas absorption peak model is a theoretical model used to describe the light absorption characteristics of gas molecules within a specific wavelength band, based on the vibrational and rotational transitions of ideal gas molecules. The formation of the absorption peak is closely related to the quantized energy level transitions of the molecules. Its characteristics include center frequency, half-width (FWHM), and absorption intensity, which are affected by factors such as temperature, pressure, and gas composition. This model has extensive applications in fields such as spectral analysis, atmospheric remote sensing, and industrial process monitoring. By analyzing the absorption peak, parameters such as gas composition, concentration, and temperature can be determined, making it an important tool for studying the interaction between gas molecules and light.
[0101] If the probability that the peak region belongs to a true peak is greater than a preset probability threshold, the peak region is considered to belong to a true peak. Otherwise, if it is less than or equal to the preset probability threshold, the peak region is considered not to belong to a true peak.
[0102] For example, the preset probability threshold is , can be adjusted according to specific conditions.
[0103] S5: Based on the probability that the peak region belongs to a true peak, the window size of the moving average is dynamically adjusted to achieve more accurate baseline correction and perform visual display.
[0104] Taking any peak shape region as the target peak shape region, calculating the variation range of the preset maximum window size and the preset minimum window size in the target peak shape region;
[0105] Calculate the probability that the target peak region does not belong to the true peak, divide the probability that it does not belong to the true peak by the preset probability threshold, and obtain the normalized result;
[0106] The product of the variation range and the normalized result is added to the preset minimum window size to obtain the corrected window size of the target peak area.
[0107] Specifically, the corrected window size satisfies the following relationship:
[0108] ;
[0109] Where, Indicates peak shape area The corrected window size, Indicates the preset minimum window size. Indicates the preset maximum window size. Indicates peak shape area The probability of belonging to a true peak, Indicates the preset probability threshold.
[0110] Exemplarily, the preset minimum window size is 3, and the preset maximum window size is 15, which can be adjusted according to actual conditions.
[0111] Visual presentation, including the following steps:
[0112] The pre-processed and baseline-corrected spectral data is displayed graphically, with the horizontal axis representing wavelength and the vertical axis representing absorbance. The automatically identified peak regions are marked on the spectral data graph, and different colors or markers are used to distinguish between true peaks and non-true peaks.
[0113] Based on manual inspection, the automatically identified peak shape is checked to see if it conforms to the expected gas absorption peak characteristics. For the confirmed characteristic peaks, the position (wavelength), intensity (absorbance), half-peak width and area parameters are recorded and used as the feature vector for subsequent classification;
[0114] The recorded characteristic peak parameters are compared with the ultraviolet absorption spectrum database of known gas molecules to establish a characteristic peak database for subsequent qualitative analysis and multi-component classification.
[0115] For example, three peak regions are automatically identified After manual verification, confirm is the true peak, It is not a real peak. The characteristic parameters of the sample are obtained and compared with the absorption peaks of known gas molecules to ultimately determine the gas components and their concentrations contained in the sample.
[0116] It should be noted that those skilled in the art may make various modifications and improvements without departing from the scope of the present invention, and these modifications and improvements fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be based on the appended claims.
Claims
1. A spectrum data classification method for ultraviolet analyzer, characterized in that: include: Acquire spectral data and perform preprocessing; The pre-processed spectral data is analyzed to calculate the local relative fluctuation degree of each wavelength point. The second-order difference between the mean value of the light intensity in the local window of the wavelength point and the mean value of the adjacent windows is calculated and normalized to obtain the probability of each wavelength point belonging to the high fluctuation area. The wavelength points belonging to the high fluctuation area and the peak areas corresponding to the wavelength points are screened out based on the probability and the set probability threshold. Obtaining peak shape areas includes the following steps: The density clustering algorithm is used to cluster the wavelength points belonging to the high fluctuation area to identify the high fluctuation area, where the identification conditions are: wavelength points Can be with wavelength point The conditions for clustering into one category are ,in, is the spectral resolution; The connected domain analysis algorithm is used to connect adjacent high-fluctuation regions to merge adjacent high-fluctuation regions, and the merged high-fluctuation regions are further divided into independent peak regions; Based on the different characteristics of baseline drift and real gas absorption peaks, the characteristics of high fluctuation areas are analyzed to identify the probability of peak areas belonging to baseline fluctuations and peak areas that may be real peaks; The shape matching degree is measured by the Pearson correlation coefficient between the ideal gas absorption peak model and the normalized peak shape vector of the corresponding peak shape area. The edge steepness of the peak shape area is combined to calculate the probability that the peak shape area belongs to a real peak and determine whether the peak shape area is a real peak. According to the probability that the peak area belongs to the true peak, the window size of the moving average is dynamically adjusted, including the steps of: taking any peak area as the target peak area, calculating the variation range of the preset maximum window size and the preset minimum window size in the target peak area; calculating the probability that the target peak area does not belong to the true peak, dividing the probability that the target peak area does not belong to the true peak by the preset probability threshold to obtain a normalized result; adding the product between the variation range and the normalized result to the preset minimum window size to obtain the corrected window size of the target peak area, so as to achieve more accurate baseline correction and perform visual display.
2. The spectral data classification method for ultraviolet analyzer according to claim 1, characterized in that: The local relative fluctuation degree includes: Perform first-order difference on the spectral data, and take the average of the absolute values of the first-order difference results of all wavelength points as the indicator of the overall spectral baseline change rate; Take any wavelength point as the target wavelength point, take the wavelength points in the data range on both sides of the target wavelength point as the local window of the target wavelength point, calculate the mean value of the light intensity in the local window of the target wavelength point, and calculate the mean value of the light intensity in the local window corresponding to the adjacent wavelength points on both sides of the target wavelength point respectively; Calculate the second-order difference between the mean of the light intensity in the local window of the target wavelength point and the mean of the adjacent windows. The absolute value of the second-order difference result is divided by the standard deviation of the window corresponding to the target wavelength point for normalization to obtain the local relative fluctuation degree of the target wavelength point.
3. The spectrum data classification method for ultraviolet analyzer according to claim 1, characterized in that: The probability that the wavelength point belongs to the high fluctuation area includes: Taking any wavelength point as the target wavelength point, calculate the sum of the preset basic sensitivity, global calibration coefficient, and overall spectral baseline change rate indicators respectively, calculate the ratio between the local relative fluctuation degree of the target wavelength point and the sum, and multiply it by the scaling factor to obtain the weighted fluctuation index; Use a negative exponential function to exponentially decay the weighted volatility index, add 1 to the attenuated result and take the inverse to obtain the probability that the target wavelength point belongs to the high volatility area.
4. The spectrum data classification method for ultraviolet analyzer according to claim 1, characterized in that: The step of screening out wavelength points belonging to the high fluctuation region includes: If the probability is greater than a preset probability threshold, the wavelength point is considered to belong to the high fluctuation area; otherwise, if the probability is less than or equal to the preset probability threshold, the wavelength point is considered not to belong to the high fluctuation area.
5. The spectrum data classification method for ultraviolet analyzer according to claim 1, characterized in that: The probability of the peak region possibly being a true peak includes: The Gaussian function is used as the ideal gas absorption peak model, which is expressed as ,in, represents the wavelength, represents the central wavelength of the peak, Indicates the width of the peak; Peak shape area The spectral data within are normalized to obtain the normalized absorption intensity , the normalized absorption intensity With the ideal Gaussian shape The Pearson correlation coefficient between them is used as the shape matching degree; The shape matching degree is added by 1 and then multiplied by the weight factor to obtain the normalized shape matching degree. The average value of the normalized shape matching degree and the edge steepness of the peak region is taken as the probability that the peak region belongs to a true peak.
6. The spectrum data classification method for ultraviolet analyzer according to claim 5, characterized in that: The edge steepness of the peak area includes: Taking any peak shape area as the target peak shape area, the radian value of the angle between the peak top in the target peak shape area and the starting point of the peak shape area is calculated to obtain the starting point angle, the radian value of the angle between the peak top in the target peak shape area and the ending point of the peak shape area is calculated to obtain the ending point angle, the starting point angle and the ending point angle are normalized respectively, the normalized angle values are added and then normalized to obtain the edge steepness of the target peak shape area.
7. The spectrum data classification method for ultraviolet analyzer according to claim 1, characterized in that: The determining whether the peak region is a real peak comprises: If the probability that the peak region belongs to a true peak is greater than a preset probability threshold, the peak region is considered to belong to a true peak. Otherwise, if it is less than or equal to the preset probability threshold, the peak region is considered not to belong to a true peak.
8. The spectrum data classification method for ultraviolet analyzer according to claim 1, characterized in that: The visual display comprises the steps of: The pre-processed and baseline-corrected spectral data is displayed graphically, with the horizontal axis representing wavelength and the vertical axis representing absorbance. The automatically identified peak regions are marked on the spectral data graph, and different colors or markers are used to distinguish between true peaks and non-true peaks. Based on manual inspection, the automatically identified peak shape is checked to see if it conforms to the expected gas absorption peak characteristics. For the confirmed characteristic peaks, the wavelength, absorbance, half-peak width and area parameters are recorded and used as the feature vector for subsequent classification; The recorded characteristic peak parameters are compared with the ultraviolet absorption spectrum database of known gas molecules to establish a characteristic peak database for subsequent qualitative analysis and multi-component classification.
Citation Information
Patent Citations
Nitrogen and sulfur in-situ measurement method and system based on differential ultraviolet spectrum technology
CN118032695A
Food pesticide residue detection system and method
CN119198665A